01 / 05
Every robot learns from episodes.
An episode is a stretch of real work, recorded and explained so a model can learn to do it too.
Data for physical AI
Every day, millions of people do skilled physical work no robot can yet do. We capture how they do it and why, then hand it to the machines learning to work beside them.
Robot hand model: Shadow Dexterous Hand, Shadow Robot Company, via MuJoCo Menagerie (Apache 2.0).
02 The data gap
≈300T tokens
Every cell ≈ 0.3T tokens. Flat, still, all in one place.
1 / 1,000
The biggest robot training set fits in this one cell.
Turn the page
Text is flat.
The world isn’t.
Space + time
Physical data has four dimensions. Space, and time.
It moves, pushes back, changes.
Orders of magnitude
We’re building the other 999.
Data about the physical world, as plentiful as text.
For scale, industry estimates: data annotation ~$9B+ by 2030 (~32% CAGR) · AI training-data services ~$11B by 2030 · humanoid robots $38B TAM by 2035 (Goldman Sachs base case). Token and corpus comparison is approximate.
01 / 05
An episode is a stretch of real work, recorded and explained so a model can learn to do it too.
02 / 05
Nine layers of signal on the same moment. Each one has to be right, and line up with the rest.
03 / 05
Staged lab sessions. Crowd labels. Short clips where nothing goes wrong. And almost never the why behind a move.
04 / 05
Real shifts, real sites, labelled by our own salaried teams.
05 / 05
Every layer dense, consistent and aligned. Packaged to your schema, captured on our devices or your rig.
Capture is an operations problem, not a software problem.
Who learns from it
One episode trains three kinds of model. Each takes different layers, into different parts of the network.
01 / 03 · Vision-language-action
A vision encoder reads the frame. A pretrained vision-language model joins it to the instruction. An action expert turns both into short chunks of motion, many times a second.
Why it mattersReal, varied demonstrations take a policy off the lab bench. Human hands teach manipulation before any robot-specific fine-tune.
Modules: vision encoder (ViT); vision-language backbone (a pretrained VLM); action expert or head, using flow matching in π0 and GR00T N1 or discrete action tokens in OpenVLA, which outputs action chunks at high frequency.
e.g. π0 · GR00T N1 · Gemini Robotics · OpenVLA
02 / 03 · Decision models
State goes in: scene, objects, task. Out comes a typed decision with a calibrated confidence. Above the threshold, the machine acts. Below it, a person answers, and that answer becomes training data.
Why it mattersRobots make thousands of small calls a minute: safe to grasp, right part, hand over now. Calibration comes from real outcomes, not staged demos.
Modules: structured state in (scene, objects, task context); decision model; typed decision with calibrated confidence; threshold gate. Above the threshold the system acts. Below it, it asks a human, and the human's answer flows back as a new training label.
Dual-system models such as GR00T N1 and Helix add a slower System-2 planner above fast control; it learns task structure from the same episodes.
03 / 03 · World models · JEPA
V-JEPA 2 masks parts of a clip and predicts what is missing as embeddings, not pixels. It pretrains on over a million hours of internet video. Then it learns actions from under 62 hours of robot video and plans by imagining outcomes.
Why it mattersInternet video is mostly third-person, filmed to entertain. First-person footage of real work is the view a robot actually has. We capture 100,000 hours of it a month.
Mechanism: the clip is split into spatio-temporal patches and a block of them is masked. A context encoder (ViT) embeds the visible patches. A target encoder, an exponential-moving-average (EMA) copy of the context encoder, embeds the full clip. A predictor predicts the target embeddings of the masked regions, and the loss compares predicted and target embeddings in latent space. V-JEPA 2-AC then keeps the encoder frozen, post-trains an action-conditioned predictor on under 62 hours of DROID robot video with no task labels or rewards, and plans by model-predictive control: it imagines the outcome of candidate actions in latent space and picks those that move closest to the embedding of a goal image.
Every episode carries every layer, aligned to the same moment. Baton delivers it in your schema, for whichever model you train.
Sources: π0 (Physical Intelligence, 2024). GR00T N1 (NVIDIA, 2025). V-JEPA 2 (Meta, 2025). Gemini Robotics, OpenVLA and Helix named for context. Architectures simplified; decision values illustrative. Named models are public research, not customers.
05 Quality
Quality decides everything, so we hire, train and pay a full-time team. The same salaried people capture, label and check every hour, in our own offices. Quality control sits inside the supply chain.
01 / 04 · The input
Footage from thousands of hands, phones and sites. What happens next decides what a model learns.
02 / 04 · The line
The same salaried staff capture, review, annotate and QA to one standard. Nothing is lost between hand-offs, because there are none.
03 / 04 · The relay
Every hour of capture runs through Baton, our relay platform. Its models cut the clips, propose labels and draft the context. Salaried staff verify, correct and add what only a person knows. Every correction makes the next hour faster.
04 / 04 · The output
Low-variance episodes, ready to train on. Your engineers build models instead of cleaning data.
Wide spread · high reject rate
Tight peak · one QA standard
The line
01
No capex, no downtime. Record during the normal shift.
Factory · clinic · job site · kitchen
02
Trained locally, scored in minutes, paid for accepted work.
iPhones · approved devices · custom rigs
03
Accepted clips cut into segments, described hand by hand by our own annotators.
Scrubbing · annotation · QA
04
Labelled episodes in your schema. Engineering time goes into models.
Training-ready data
1,000+
Salaried professionals, trained and in-office
10–20
Hours of prep for every hour of footage
06 Where we start
India holds the frontier and the everyday in one country. Operating theatres that benchmark against leading US centres. Some of the largest construction projects on earth. Precision two-wheeler lines. And the street kitchens, workshops and homes of 1.4 billion people. One pipeline and one QA standard cover all of it.
01 Frontier
Robotic & open surgery · ICU · Wards
Operating theatres that hold themselves to US benchmarks. The surgeon’s hands, the scrub nurse’s timing, the robot’s arms, all on record.
02 Frontier
Tunnels · Bridges · Metro
Tunnels, bridges and metro lines on a scale few countries attempt. Crews, cranes and steel moving in step.
03 Industrial
Two-wheeler assembly
Lines that build millions of vehicles a year. Every motion paced to takt time, every bolt torqued to spec.
04 Skilled trades
Mechanics · Electricians · Fabricators
Real tools, worn parts, no two jobs alike. The mechanic hears the engine before he opens it.
05 Everyday
Roadside stalls · Commercial kitchens
Heat, speed and food that changes shape in the pan. The cook lowers the flame without looking.
06 Everyday
Kitchens · Living rooms · Laundry
The homes of 1.4 billion people. Cluttered, personal, never the same twice.
* Exclusive partnerships agreed in principle (LOIs); subject to final agreement. Apollo figures as reported by Apollo Hospitals.
On the ground
Andhra Pradesh · Jharkhand · Karnataka · Madhya Pradesh · Maharashtra · Rajasthan · Uttar Pradesh. Offices: London · India · San Francisco · Shanghai.
First station
India is where the relay starts. It does not stop here.
07 The loop
Start in India · Relay everywhere
Step 1 · Start
Station one: India. Every kind of work, recorded during real shifts.
Step 2 · Spread
The playbook travels. Southeast Asia, the Middle East, Africa, Latin America, Europe, the US. Same pipeline. Same QA.
Step 3 · Train
Training-ready episodes ship to the labs building VLA models and humanoids.
Step 4 · Deploy
The robots come back. Into the same wards, kitchens and job sites, to be tested on real work.
Step 5 · Correct
Every failure is a lesson. When a robot slips, we capture the moment, relabel it and send it back through Baton. Each lap, the models get better and the capture gets sharper.
Traction
A hundred thousand hours of real work arrive every month. Our own staff capture it and check it.
Language
Text
Taught machines to talk.
Physical
Work
Will teach them to act. We relay it.
The enterprise model
You open the doors to real work and the people who know it. We capture it, prepare it and build the models, at no cost to you. When the data earns, you earn.
HumanRelay. delivers the full use case
Enterprise
Gives access to real workflows and the experts who run them.
$0 cost · no charge
Step 01
First-person recording of real work, inside your consent framework.
Step 02
Scrubbing, annotation and QA, checked by people.
Step 03
AI built for your use case, tuned to your workflows.
Monetisation A
Data licensed to labs building models.
Monetisation B
Products sold in non-competing markets.
You take a percentage. Never sold to competitors in your own market.
Partner with us