Multimodal AI / Computer Vision Engineer — Trading Education Data Pipeline
We are building an AI system designed to learn discretionary price-action trading from expert educational videos.
We need an engineer to build a pipeline that converts narrated trading-course videos into structured, machine-readable training examples by synchronizing:
- Video frames
- Course transcripts
- Mouse/pointer movement
- Candlestick charts
- Chart annotations and text
- Individual price bars
- Expert explanations and trading decisions
Core Problem
The instructor teaches using chart slides containing candlestick bars, EMA lines, annotations, arrows, and text. During the video, he moves the mouse over specific bars or chart regions while saying things such as:
- "This bar"
- "These bars"
- "This breakout"
- "This pullback"
- "This second entry"
- "This wedge"
- "I would buy here"
- "I would not take this trade because..."
The charts usually do not contain explicit OHLC data or bar IDs.
The goal is to determine exactly which chart bars or regions the instructor is referring to and connect those objects with the corresponding transcript and explanation.
For example:
Video timestamp: 00:43:17
Referenced bars: 57–63
Setup: H2 / second-entry buy
Context: Trading range
Location: Near top of range
Instructor decision: No trade
Reason: Poor location / insufficient probability
Source: Exact video and transcript segment
The resulting dataset will eventually be used for retrieval, evaluation, and potentially fine-tuning an AI trading decision system.
Primary Responsibilities
Build a pipeline that:
- Extracts frames and timestamps from course videos.
- Synchronizes transcripts to exact video timestamps.
- Detects the chart area within each slide/frame.
- Detects and tracks individual candlestick bars.
- Assigns stable bar IDs from left to right.
- Detects the instructor's mouse/pointer.
- Tracks pointer movement across time.
- Resolves phrases such as "this bar" or "these bars" to the specific bars or chart regions being indicated.
- Detects existing chart annotations, arrows, trend lines, EMA lines, and text where useful.
- Associates the instructor's explanation with the referenced chart objects.
- Outputs structured training examples in JSON/database format.
- Stores the exact source video, timestamp, transcript, frame/crop, referenced bars, and confidence for every example.
- Abstains when the referenced object cannot be determined confidently rather than creating bad labels.
Important Design Principle
Precision is more important than recall.
The system should not attempt to label every teaching moment.
If the pointer/reference alignment is ambiguous, the example should be marked UNCERTAIN and preserved for later human review or stronger future models.
A smaller dataset of highly reliable examples is substantially more valuable than a large dataset containing incorrect Brooks labels.
Trading Knowledge Structure
The first version will focus on approximately 12 primary Al Brooks setups.
Hundreds of additional Brooks concepts will be treated primarily as context, rather than giving every concept equal weight.
The pipeline therefore needs to support:
Primary setup → referenced bars → surrounding context → expert interpretation → expert decision → expert reasoning.
The engineer does not need to already be an Al Brooks expert. Trading-domain interpretation will be provided and developed separately.
The engineering problem is primarily multimodal grounding and data extraction.
Technical Skills
Strong experience with:
- Python
- Computer vision
- OpenCV
- Video processing
- Object detection/tracking
- Coordinate transformations
- Temporal alignment
- Speech/transcript alignment
- Multimodal LLM/VLM APIs
- Structured data pipelines
- JSON/SQL or similar data storage
- Automated validation/testing
Experience with the following would be especially useful:
- Cursor/pointer tracking
- Candlestick/chart recognition
- OCR
- Vision-language models
- Video understanding
- Temporal grounding
- Weak supervision
- Human-in-the-loop labeling systems
- Fine-tuning or dataset creation for multimodal models
Expected Pipeline
The target architecture is approximately:
Course video
→ frame extraction
→ transcript synchronization
→ chart detection
→ candlestick/bar detection
→ pointer tracking
→ language-reference resolution
→ Brooks concept association
→ structured training example
→ confidence score
→ automatic acceptance or human review
Example Output{ "video_id": "course_18_video_04", "timestamp": "00:43:17.250", "transcript": "This is a second entry buy, but we're at the top of the trading range so I would not buy here.", "referenced_bars": [57, 58, 59, 60, 61, 62, 63], "primary_bar": 63, "setup": "H2", "context": [ "trading_range", "top_of_range" ], "decision": "no_trade", "confidence": 0.98, "source_frame": "frame_155837.png" }
The exact schema will evolve during development.
First Milestone
Do not attempt to solve the entire course initially.
Build a proof of concept using a small set of videos and demonstrate that the system can reliably:
transcript phrase → pointer location → exact chart bar(s).
Example:
Instructor says: "This bar here."
System identifies: Bar 47.
Once this works reliably, expand to ranges of bars, chart regions, annotations, and Brooks concepts.
Success Criterion
The project succeeds when we can automatically convert expert video instruction into a large library of high-confidence, visually grounded trading examples without manually labeling every chart.
The objective is not merely to recognize candlestick patterns.
The ultimate dataset must preserve:
what the expert saw + where he saw it + the surrounding market context + what he thought it meant + whether he would trade it + why.
Pay: From $10.00 per hour
Expected hours: No more than 40.0 per week
Work Location: Remote