Data for

PI DATA — physical AI data, robotics datasets and robot training data for Physical
Intelligence.

We're building the data infrastructure for physical AI — real-world robotics and human egocentric data, collected, engineered and validated in India.

FIELD FOOTAGE
4K · RGB-D · 60 FPS
humanoid / arm / teleop / egocentric RGB-D / force / proprioception trajectories
Built for · early-stage, no logos yet
/01 Robotics Labs
/02 Foundation Model Teams
/03 Humanoid Companies
/04 Industrial Automation
The physical-AI data engine

Six stages.
One pipeline.

Every dataset we build moves through the same pipeline — from task definition to training-ready delivery.

01

Define

Task + environment + embodiment

02

Collect

Teleoperation / autonomous / human demonstrations

03

Capture

RGB / depth / force / proprioception / audio

04

Annotate

Actions / objects / contact / events / outcomes

05

Validate

Calibration / QA / consistency / diversity

06

Deliver

Training-ready robotics datasets + APIs

Data, not just recordings

What we track

Every delivery is reported against these dimensions. We publish real numbers as datasets ship — no backfilled stats.

X
Hours
X
Embodiments
X
Environments
X
Tasks
X
Sensor streams
// placeholder values — replaced with real metrics as we ship
Two data streams

Robot data +
human egocentric.

Two sources of physical experience, one pipeline. Egocentric human data is the highest-bandwidth source of physical experience we know; robot data grounds models in real embodiment.

STREAM / 01

Robot data

Humanoid, arm and mobile-base manipulation and movement data — teleoperated or autonomous — multimodal robot data captured with the full modality stack: RGB, depth, force/torque, joint states, tactile and audio.

humanoid arm teleop autonomous
STREAM / 02

Human egocentric

First-person human demonstrations of real tasks — kitchen, home, warehouse, factory — captured at scale, time-aligned and annotated. The scaling path for physical AI before, and beyond, robot fleets.

egocentric human demos daily tasks at scale
Our collection stack

Every signal,
in sync.

High-fidelity, time-aligned sensing across both data streams.

RGB
Depth
Force / Torque
θJoint States
Tactile
Audio
Teleoperation
Egocentric
Data programs

Three ways
to work with us.

All programs are scoped to your spec — task, environment, embodiment, modality. We're early-stage: small team, direct access, fast iteration.

DATA COLLECTION
FIELD OPS · INDIA
egocentric human task demonstrations time-aligned streams
PROGRAM / 01

Robot data

Robot teleoperation and autonomous collection on arms and humanoids — robotics data collection across the full sensor stack. We run collection in India from day one.

teleop autonomous humanoid / arm
PROGRAM / 02

Egocentric human

First-person human task demonstrations captured at scale — the fastest way to accumulate diverse physical experience for foundation models.

egocentric daily tasks scale
PROGRAM / 03

Custom dataset

You define the task, environment, embodiment and modality spec. We scope it, schedule it, collect it and deliver it training-ready.

spec-driven scoped delivered
From raw experience → training data

The pipeline.

0
Raw trajectory
1
Sensor synchronization
2
Calibration
3
Segmentation
4
Annotation
5
Quality validation
Training dataset
Data quality

Every gate,
in writing.

Quality standards are written down before collection starts. Nothing ships without passing the same gates, every time.

GATE / 01

Calibration

Camera intrinsics and extrinsics, force zeroing and IMU alignment checked per session.

GATE / 02

Synchronization

All sensor streams time-aligned; drift detected and flagged before it reaches the dataset.

GATE / 03

Annotation QA

Action, object, contact and event labels re-checked on a sampled basis against the spec.

GATE / 04

Consistency

Trajectory validity, dropouts and sensor faults filtered out before delivery.

GATE / 05

Diversity audit

Coverage across objects, scenes and task variants checked against what you asked for.

GATE / 06

Delivery spec

Training-ready format, schema documentation and API access handed over with every dataset.

Why India

Where we start.

Real-world data is abundant, diverse, and cost-effective to collect here. The long-term goal is scalable physical-AI data infrastructure — we start with real-world collection in India, and build the network from here.
[1]

Task diversity

Households, markets, warehouses and factories — dense, varied daily-life tasks in the wild.

[2]

Collection cost

Hands-on data collection at a fraction of Western operating cost per hour.

[3]

Operator scale

A large pool of teleop operators and annotators to grow with demand.

[4]

Real environments

Noisy, uncontrolled, multilingual — exactly the conditions physical AI has to work in.

[5]

Long-term infra

Sites, tooling and pipeline built here become the first node of a global data network.

Dataset requests

Tell us what
you need.

We work with robotics labs, humanoid companies and physical-AI teams that need real-world robot training data. Describe the task and we'll come back with a scope, a plan and a timeline. We're early-stage — you talk to the people doing the work, and we move fast.

PREFER EMAIL?
hello@example.com
// opens your email client with the request pre-filled
Early stage. Moving fast.

Let's build
your dataset