Data for physical AI

Human skill, relayed.

Every day, millions of people do skilled physical work no robot can yet do. We capture how they do it and why, then hand it to the machines learning to work beside them.

Robot hand model: Shadow Dexterous Hand, Shadow Robot Company, via MuJoCo Menagerie (Apache 2.0).

  • 100,000 hrs captured monthly
  • 5,000+ partner sites
  • 1,000+ salaried specialists
  • Starting in India, built for everywhere

02 The data gap

Language models learned from text.

≈300T tokens

Every cell ≈ 0.3T tokens. Flat, still, all in one place.

1 / 1,000

The biggest robot training set fits in this one cell.

Turn the page

Text is flat.
The world isn’t.

Space + time

Physical data has four dimensions. Space, and time.

It moves, pushes back, changes.

Orders of magnitude

We’re building the other 999.

Data about the physical world, as plentiful as text.

  1. 01Robots learn from first-person footage of real work.
  2. 02You can’t scrape it. You film it where the work happens.
  3. 03Whoever runs that capture owns the input to physical AI.

For scale, industry estimates: data annotation ~$9B+ by 2030 (~32% CAGR) · AI training-data services ~$11B by 2030 · humanoid robots $38B TAM by 2035 (Goldman Sachs base case). Token and corpus comparison is approximate.

Anatomy of a training episode

01 / 05

Every robot learns from episodes.

An episode is a stretch of real work, recorded and explained so a model can learn to do it too.

02 / 05

Pull one apart.

Nine layers of signal on the same moment. Each one has to be right, and line up with the rest.

03 / 05

Most datasets are hollow where it matters.

Staged lab sessions. Crowd labels. Short clips where nothing goes wrong. And almost never the why behind a move.

04 / 05

We fill the gaps at the source.

Real shifts, real sites, labelled by our own salaried teams.

05 / 05

Delivered training-ready.

Every layer dense, consistent and aligned. Packaged to your schema, captured on our devices or your rig.

Capture is an operations problem, not a software problem.

  1. RGB egocentric video (first-person). Typical: staged lab tabletop. HumanRelay: real work during real shifts.
  2. Depth & 3D scene geometry. Typical: a few rooms, scanned once. HumanRelay: hospitals → megaprojects → homes.
  3. Hand & body pose (keypoints and skeleton). Typical: crowd-labelled, inconsistent. HumanRelay: one QA standard, one team.
  4. Object & affordance labels (bounding boxes). Typical: outsourced labelling, thin QA. HumanRelay: labelled by people who know the task.
  5. Action segmentation (reach / grasp / turn / place). Typical: coarse or unsegmented. HumanRelay: same teams capture, label and QA.
  6. Language narration. Typical: one-line caption, if any. HumanRelay: step-by-step, hand by hand.
  7. Human context & intent (the why behind each move). Typical: motion without the why. HumanRelay: the reasoning a skilled hand carries.
  8. Environment diversity (real sites, lighting, clutter). Typical: a handful of environments. HumanRelay: ICU to street kitchen.
  9. Long-horizon real work, edge cases, failure & recovery. Typical: minutes, not shifts · no failures recorded. HumanRelay: full shifts · failures & recoveries captured.

Who learns from it

Who learns from it.

One episode trains three kinds of model. Each takes different layers, into different parts of the network.

01 / 03 · Vision-language-action

VLA policies. Fast control.

A vision encoder reads the frame. A pretrained vision-language model joins it to the instruction. An action expert turns both into short chunks of motion, many times a second.

Why it mattersReal, varied demonstrations take a policy off the lab bench. Human hands teach manipulation before any robot-specific fine-tune.

  • L01 RGB egocentric video trains Vision encoder
  • L08 Environment diversity trains Vision encoder
  • L06 Narration & instructions trains Vision-language backbone
  • L03 Hand & body pose trains Action expert
  • L05 Action segments trains Action expert

Modules: vision encoder (ViT); vision-language backbone (a pretrained VLM); action expert or head, using flow matching in π0 and GR00T N1 or discrete action tokens in OpenVLA, which outputs action chunks at high frequency.

e.g. π0 · GR00T N1 · Gemini Robotics · OpenVLA

02 / 03 · Decision models

Decision models. Fast, typed, calibrated.

State goes in: scene, objects, task. Out comes a typed decision with a calibrated confidence. Above the threshold, the machine acts. Below it, a person answers, and that answer becomes training data.

Why it mattersRobots make thousands of small calls a minute: safe to grasp, right part, hand over now. Calibration comes from real outcomes, not staged demos.

  • L04 Objects & affordances trains Structured state
  • L07 Context & intent, the why trains Structured state
  • L05 What the skilled person chose trains Decision model
  • L09 Outcomes, failures, hesitations trains Calibrated confidence

Modules: structured state in (scene, objects, task context); decision model; typed decision with calibrated confidence; threshold gate. Above the threshold the system acts. Below it, it asks a human, and the human's answer flows back as a new training label.

Dual-system models such as GR00T N1 and Helix add a slower System-2 planner above fast control; it learns task structure from the same episodes.

03 / 03 · World models · JEPA

World models. Physics from watching.

V-JEPA 2 masks parts of a clip and predicts what is missing as embeddings, not pixels. It pretrains on over a million hours of internet video. Then it learns actions from under 62 hours of robot video and plans by imagining outcomes.

Why it mattersInternet video is mostly third-person, filmed to entertain. First-person footage of real work is the view a robot actually has. We capture 100,000 hours of it a month.

  • L01 RGB egocentric video, at volume trains Encoder & predictor pretraining
  • L08 Environment diversity trains Encoder pretraining
  • L03 Hand & body pose trains Action-conditioned predictor
  • L05 Action segments trains Action-conditioned predictor

Mechanism: the clip is split into spatio-temporal patches and a block of them is masked. A context encoder (ViT) embeds the visible patches. A target encoder, an exponential-moving-average (EMA) copy of the context encoder, embeds the full clip. A predictor predicts the target embeddings of the masked regions, and the loss compares predicted and target embeddings in latent space. V-JEPA 2-AC then keeps the encoder frozen, post-trains an action-conditioned predictor on under 62 hours of DROID robot video with no task labels or rewards, and plans by model-predictive control: it imagines the outcome of candidate actions in latent space and picks those that move closest to the embedding of a goal image.

One capture. Every kind of model.

Every episode carries every layer, aligned to the same moment. Baton delivers it in your schema, for whichever model you train.

Sources: π0 (Physical Intelligence, 2024). GR00T N1 (NVIDIA, 2025). V-JEPA 2 (Meta, 2025). Gemini Robotics, OpenVLA and Helix named for context. Architectures simplified; decision values illustrative. Named models are public research, not customers.

05 Quality

People are people. Data is data.

Quality decides everything, so we hire, train and pay a full-time team. The same salaried people capture, label and check every hour, in our own offices. Quality control sits inside the supply chain.

01 / 04 · The input

Every pipeline starts as noise.

Footage from thousands of hands, phones and sites. What happens next decides what a model learns.

02 / 04 · The line

One team. Four gates.

The same salaried staff capture, review, annotate and QA to one standard. Nothing is lost between hand-offs, because there are none.

03 / 04 · The relay

Our AI drafts. Our people decide.

Every hour of capture runs through Baton, our relay platform. Its models cut the clips, propose labels and draft the context. Salaried staff verify, correct and add what only a person knows. Every correction makes the next hour faster.

04 / 04 · The output

Same world. Tighter data.

Low-variance episodes, ready to train on. Your engineers build models instead of cleaning data.

Unmanaged capture

Wide spread · high reject rate

HumanRelay salaried team

Tight peak · one QA standard

The line

From a normal shift to your schema.

  1. 01

    Host sites

    No capex, no downtime. Record during the normal shift.

    Factory · clinic · job site · kitchen

  2. 02

    Recorders

    Trained locally, scored in minutes, paid for accepted work.

    iPhones · approved devices · custom rigs

  3. 03

    Annotators

    Accepted clips cut into segments, described hand by hand by our own annotators.

    Scrubbing · annotation · QA

  4. 04

    VLA model development

    Labelled episodes in your schema. Engineering time goes into models.

    Training-ready data

1,000+

Salaried professionals, trained and in-office

10–20

Hours of prep for every hour of footage

06 Where we start

We start where the range is widest.

India holds the frontier and the everyday in one country. Operating theatres that benchmark against leading US centres. Some of the largest construction projects on earth. Precision two-wheeler lines. And the street kitchens, workshops and homes of 1.4 billion people. One pipeline and one QA standard cover all of it.

01 Frontier

Operating theatre

Robotic & open surgery · ICU · Wards

Operating theatres that hold themselves to US benchmarks. The surgeon’s hands, the scrub nurse’s timing, the robot’s arms, all on record.

  • Instrument handover
  • Sterile field
  • Patient transfer

02 Frontier

Megaproject job site

Tunnels · Bridges · Metro

Tunnels, bridges and metro lines on a scale few countries attempt. Crews, cranes and steel moving in step.

  • Rebar tying
  • Ring segment fit
  • Material hoisting

03 Industrial

Factory line

Two-wheeler assembly

Lines that build millions of vehicles a year. Every motion paced to takt time, every bolt torqued to spec.

  • Torque & fasten
  • Part kitting
  • Visual inspection

04 Skilled trades

Workshop & trades

Mechanics · Electricians · Fabricators

Real tools, worn parts, no two jobs alike. The mechanic hears the engine before he opens it.

  • Diagnose
  • Wrench & torque
  • Wire & solder

05 Everyday

Street kitchen

Roadside stalls · Commercial kitchens

Heat, speed and food that changes shape in the pan. The cook lowers the flame without looking.

  • Knead · flip · plate
  • Ladle & pour
  • Wash up

06 Everyday

Home

Kitchens · Living rooms · Laundry

The homes of 1.4 billion people. Cluttered, personal, never the same twice.

  • Fold laundry
  • Load dishes
  • Tidy up

* Exclusive partnerships agreed in principle (LOIs); subject to final agreement. Apollo figures as reported by Apollo Hospitals.

On the ground

Twenty offices. Seven states. One standard.

20+
Delivery offices
7
States
  • Highly educated
  • English speaking
  • Salaried, in-office

Andhra Pradesh · Jharkhand · Karnataka · Madhya Pradesh · Maharashtra · Rajasthan · Uttar Pradesh. Offices: London · India · San Francisco · Shanghai.

First station

India is where the relay starts. It does not stop here.

07 The loop

Start in India  ·  Relay everywhere

Step 1 · Start

Station one: India. Every kind of work, recorded during real shifts.

Step 2 · Spread

The playbook travels. Southeast Asia, the Middle East, Africa, Latin America, Europe, the US. Same pipeline. Same QA.

Step 3 · Train

Training-ready episodes ship to the labs building VLA models and humanoids.

Step 4 · Deploy

The robots come back. Into the same wards, kitchens and job sites, to be tested on real work.

Step 5 · Correct

Every failure is a lesson. When a robot slips, we capture the moment, relabel it and send it back through Baton. Each lap, the models get better and the capture gets sharper.

Wherever people work, we relay it.

  1. Capture
  2. AI drafts
  3. People verify
  4. Better models
  5. Sharper capture

Traction

The relay is
already running.

A hundred thousand hours of real work arrive every month. Our own staff capture it and check it.

Captured now
100,000hrs/month
First-person video of real work
Growth
50–100%
Week over week, in hours captured and businesses signed
Hours of first-person video captured per monthIllustrative trajectory
Network
5,000+
Businesses in the network
About 100 more sign on every day

Language

Text

Taught machines to talk.

Physical

Work

Will teach them to act. We relay it.

The enterprise model

We build the AI.
You share the upside.

You open the doors to real work and the people who know it. We capture it, prepare it and build the models, at no cost to you. When the data earns, you earn.

HumanRelay. delivers the full use case

Enterprise

Your organisation

Gives access to real workflows and the experts who run them.

$0 cost · no charge

Step 01

Capture

First-person recording of real work, inside your consent framework.

Step 02

Data preparation

Scrubbing, annotation and QA, checked by people.

Step 03

Model development

AI built for your use case, tuned to your workflows.

Monetisation A

Data licensing

Data licensed to labs building models.

Monetisation B

Products

Products sold in non-competing markets.

You take a percentage. Never sold to competitors in your own market.