Physical AI
Egocentric data is the new oil for physical AI: what changed in 2026
Aman Mundra · August 28, 2026 · 9 min read

Contents
- The datasets: from 10,000 hours to a million in five months
- Why first-person video works for robots
- The humanoids are on the floor now
- The open stack: NVIDIA, Physical Intelligence, Hugging Face
- The annotation market grew up around it
- India is becoming the recording studio
- Meanwhile, ordinary computer vision is where the money is this year
- What we are doing about it
- References
Egocentric data is video and sensor data recorded from a person's own point of view: a camera on the forehead or glasses, often paired with hand tracking, depth, and force. In 2026 it became the primary raw material for physical AI, the class of models that control robots rather than answer questions. The reason is simple. A language model learned from the internet; there is no internet of hands. Robot policies need millions of examples of people doing physical work, and the cheapest way to collect them is to record the people who already do it.
This post is a field report. It covers the datasets, the humanoid deployments on real factory floors, the open model stack from NVIDIA and Physical Intelligence, the annotation market that grew up around all of it, and the part most relevant to a plant or a data team in India: how to get in.
The datasets: from 10,000 hours to a million in five months
Meta built the field with Ego4D in 2021 (3,670 hours, 923 participants, 74 locations) and followed it with Ego-Exo4D in late 2023: 1,286 hours of synchronized first-person and third-person video, captured on Project Aria glasses synced with four or five GoPros, across 740 participants and 123 scenes. Those were research datasets, and for two years they were the ceiling.
Then the scale changed. Build AI shipped Egocentric-10K in November 2025 (10,000 hours, 2,153 factory workers, 1.08 billion frames), Egocentric-100K a month later (100,000 hours, 14,228 workers), and Egocentric-1M in April 2026: roughly a million hours of factory-floor footage under an Apache 2.0 license. Ten thousand hours to a million in five months is the fastest dataset scaling the field has seen. Alongside it sit Apple's EgoDex (829 hours with precise 3D hand and finger tracking), the multi-institution EgoVerse (1,362 hours, 80,000 episodes, 1,965 tasks), and Meta's HOT3D for hand-object tracking.
Two things stand out in that list. First, the biggest sets are factory footage, not kitchens. The money is in manufacturing and logistics, and the data followed it. Second, the newest datasets are no longer plain video. They pair RGB with depth, hand pose, and sometimes tactile force, because that is what a policy needs to reproduce the motion, not just recognise it.
Why first-person video works for robots
The scientific case was made by Georgia Tech's EgoMimic: a framework that co-trains a robot on its own teleoperated demonstrations and on human egocentric video from Aria glasses. With about 90 minutes of human recordings, task success improved by up to 400 percent across manipulation tasks. The human video is cheap and abundant; the robot data is expensive and scarce; the model learns to bridge the two embodiments. Follow-up work such as Humanoid Policy ~ Human Policy and 2026 papers like Human-as-Humanoid push the same idea toward zero-shot transfer: train on people, deploy on a humanoid with a matched embodiment.
The practical consequence is that the bottleneck moved from robots to recording rigs. Low-cost capture kits (EgoKit) and large everyday-task collections (EgoLive) are now research topics in their own right. Whoever can put a thousand headsets on a thousand workers, keep the data synchronized and consented, and label it consistently, owns the supply.
The humanoids are on the floor now
This is no longer a demo cycle. The deployments reported in 2026 trade press (humanoid.guide, iFactory, The AI Insider):
- Figure completed an eleven-month pilot at BMW Spartanburg: 1,250 operating hours, more than 90,000 sheet-metal parts loaded, supporting production of more than 30,000 X3s. Figure 03 now runs there at a quoted 25 dollars per robot-hour with placement accuracy above 99 percent, and a Leipzig deployment was announced for summer 2026.
- Tesla put more than 1,000 Optimus Gen 3 units into its own plants from January 2026 on battery assembly, pack loading, cable routing, and connector seating. Tesla itself concedes that a share of the fleet is there to generate training data rather than output. That admission is the whole thesis of this post in one sentence.
- Apptronik Apollo moves assembly kits on Mercedes-Benz lines; Jabil is both building the robots and testing them on sorting, kitting, inspection, and subassembly.
- Agility Digit moves totes inside Amazon fulfilment centres and finished a year-long commercial pilot at Toyota.
Read those numbers carefully. Ninety thousand parts over eleven months is a slow human. The robots are not yet faster than people; they are cheaper per hour, they do not tire, and every hour they run is another hour of training data. The economics of a humanoid are the economics of a data flywheel.
The open stack: NVIDIA, Physical Intelligence, Hugging Face
The reason a mid-size firm can build on any of this is that the model layer went open in 2026.
NVIDIA's January 2026 physical AI release shipped Isaac GR00T N1.6, an open reasoning vision-language-action (VLA) model for humanoids with full-body control, plus Cosmos Transfer 2.5, Predict 2.5, and Reason 2 for world simulation and video reasoning, the Isaac Lab-Arena evaluation framework, and OSMO for edge-to-cloud training orchestration. Partners named on the day included Boston Dynamics, Caterpillar, Franka, LG, NEURA, AGIBOT, and Hugging Face. By mid-year GR00T N1.7 was in early access with commercial licensing, N2 had been previewed, and Cosmos 3 had arrived as an open omnimodel in 16B and 64B variants.
Physical Intelligence's pi0 set the template for generalist policies: one network trained across eight robot types on tasks from folding laundry to routing cables and assembling boxes, and it is available through Hugging Face LeRobot for fine-tuning on your own robot. Between GR00T, pi0, and LeRobot, the starting point for a custom policy is now a checkpoint and a few hundred demonstrations, not a research programme.
The annotation market grew up around it
Every hour of egocentric video needs work before a model can use it: synchronizing streams, segmenting tasks, tracking hands and objects, labelling contact and success, and rejecting the takes where the camera slipped. A specialist market formed fast. 2026 roundups of robotics annotation vendors (RoboticsBiz, TrueLabel) name Scale AI (with a Physical AI Data Engine built on real robot interaction data), iMerit, Appen, Shaip, and marketplaces such as TrueLabel. The common shape is capture plus labelling under one roof, because the two cannot be separated: a labeller who did not see the rig setup cannot judge the take.
Simulation did not kill this market. As DataX Power's guide argues, there are classes of data that simulation cannot replace: real contact dynamics, deformable objects, worn tooling, dirty lenses, and the thousand small ways a real line differs from its digital twin. Sim generates volume; humans generate truth.
India is becoming the recording studio
The most important 2026 development for readers here is that egocentric collection landed in India at scale. TechCrunch reported in May that Human Archive, a YC-backed startup founded by Berkeley and Stanford researchers with 8.2 million dollars from Wing and NVP, has more than 1,000 active headsets deployed, mostly through Indian home-services, hotel, and restaurant partners, paying workers to record their shifts with cap cameras, tactile gloves, wrist cameras, and full-body motion capture. explainx describes an Ahmedabad electronics floor where about 50 assembly workers wear GoPros to record screwing, assembling, and packing for annotation. Indian firms Awign, Humyn AI, FPV Labs, Neo Cambrian, and Objectways are building their own pipelines, and Entrackr counts 155 million dollars raised by Indian physical-AI startups across 31 deals in 2026 so far.
The pattern is the same one that made India the world's back office: a large, skilled workforce, low cost per hour, and English-language QA. The difference is that this time the product is not a ticket resolved; it is a demonstration a robot will imitate.
Meanwhile, ordinary computer vision is where the money is this year
A sober counterweight. The Association for Advancing Automation's 2026 survey, summarised by IIoT World, found that 41 percent of manufacturers list AI vision as their top automation priority for 2026, ahead of both LLMs and humanoids. Defect detection, OCR on labels, safety monitoring, and pick verification are the projects that get budget, because they pay back in a quarter and need a camera, not a robot. The direction of travel in inspection, per the buildmvpfast guide, is fewer labelled samples, more edge inference, and models that explain a reject. Vision-language models are arriving on the floor too, though A3's own guide is honest that their immediate impact is still uncertain.
If you run a plant, this is the order of operations: vision first, then data capture from the people already on your line, then a robot that learned from them.
What we are doing about it
Welzin added two capabilities this week to sit alongside our data, GenAI, and MLOps work: Computer Vision & Data Annotation and Robotics & Physical AI. Concretely: inspection and detection models; managed labelling for images, video, 3D, and egocentric streams with QA gates; capture rigs and teleoperation for demonstration data; policies fine-tuned from GR00T and LeRobot checkpoints; simulation in Isaac Sim and MuJoCo; and integration with the PLC, MES, and WMS systems a line already runs. The longer argument, with an architecture and a cost model, is in our whitepaper The Physical AI Data Engine, and there is a worked example in the vision inspection case study.
The short version of 2026: the models are open, the robots are on the floor, and the scarce thing is well-labelled footage of people doing real work. That is a data problem, and data problems are what we do.
References
- Meta AI, Introducing Ego-Exo4D and Ego-Exo4D dataset docs
- Meta AI, EgoMimic: Project Aria glasses help train humanoid robots; paper: EgoMimic
- Humanoid Policy ~ Human Policy; Human-as-Humanoid; EgoKit; EgoLive
- NVIDIA, NVIDIA Releases New Physical AI Models as Global Partners Unveil Next-Generation Robots; Cosmos 3 and Isaac GR00T review
- Physical Intelligence, pi0: Our First Generalist Policy; Hugging Face, LeRobot pi0
- humanoid.guide, Humanoid deployments in 2026 favor Figure and Agility; iFactory, Humanoid Robots on the Factory Floor; The AI Insider, The State of Humanoid Robotics in 2026
- RoboticsBiz, Best data labeling companies for robotics & physical AI in 2026; TrueLabel, Best Egocentric Video Data Providers for Robotics; DataX Power, Physical AI Training Data: 6 Types Simulation Cannot Replace
- TechCrunch, This startup is betting India's gig economy can train the world's robots; Entrackr, Indian Physical AI attracts $155 Mn in 2026; explainx, Indian Workers Wear Cameras to Train AI Robots; IndiaAIPulse, Egocentric data collection fuels robotics AI growth in India
- IIoT World, 2026 Smart Factory AI Vision Trends; A3, A Guide to Vision Language Models; buildmvpfast, AI Quality Control: Computer Vision for Manufacturing 2026











