Primers • Physical AI
- Overview
- The Perception-Reasoning-Action Loop
- Physical AI, Embodied AI, and Autonomy
- From Robotics Models to Robot Foundation Models
- Vision-Language-Action Models
- Continuous Generative Action Models
- Physical AI as a Data Problem
- World Models and Simulation
- Physical AI Across Robots and Autonomous Vehicles
- Toward Generalist Physical Agents
- Foundations of Embodied Intelligence
- Embodiment as a Computational Constraint
- The Agent-Environment Loop
- Embodied Intelligence as a POMDP
- Perception Is Not Merely Object Recognition
- Learning Visual Representations from Human Activity
- Grounding Language in Physical Interaction
- Affordances
- From Modular Robotics to Learned Policies
- Imitation Learning as the Basic Robot-Learning Primitive
- Scaling Demonstrations Enables Generalization
- Multimodal Prompting for Robots
- Action Chunking and Temporal Abstraction
- Transformers Enter Robot Control
- From Robotics Transformers to Generalist Agents
- Embodied Multimodal Language Models
- The Transition from VLMs to VLAs
- Cross-Embodiment Learning
- Generalization as the Central Objective
- The Emerging Embodied Foundation-Model Stack
- Learning to Act
- From Understanding the World to Controlling It
- Behavioral Cloning
- Distribution Shift and Compounding Errors
- Why Mean-Squared Error Can Fail for Physical Actions
- Action Chunking
- Receding-Horizon Execution
- Action Chunking with Transformers
- Diffusion Policies
- Why Diffusion Works Well for Robot Actions
- Diffusion Transformers for Large Robot Policies
- From Diffusion to Flow Matching
- Flow Matching in Modern VLAs
- Diffusion Versus Flow Matching
- Why Imitation Alone Is Not Enough
- Value Functions and Physical Decision Making
- Offline Reinforcement Learning
- Conservative Q-Learning
- Implicit Q-Learning
- Advantage-Weighted Policy Learning
- Online Reinforcement Learning
- Reward Design
- Hierarchical Control
- System 2 and System 1 Control
- Model-Based Reinforcement Learning
- Latent World Models
- Training in Simulation
- Combining Demonstrations and Reinforcement Learning
- Closed-Loop Policy Improvement
- Learning Recovery Behavior
- Action Representations
- Cross-Embodiment Action Spaces
- The Modern Action-Learning Stack
- Vision-Language-Action Models
- From Vision-Language Models to Physical Policies
- Why Start from a Vision-Language Model?
- The Generic VLA Architecture
- Action Tokenization
- RT-2
- Semantic Reasoning Versus Motor Control
- OpenVLA
- OpenVLA Action Decoding
- Fine-Tuning VLAs to New Robots
- Limitations of Autoregressive Action Tokens
- OpenVLA-OFT
- Octo
- Octo’s Readout Architecture
- Flexible Observation Spaces
- Two Paths to Generalist Robot Policies
- $\pi_0$
- The Action Expert
- Flow Matching for Continuous Control
- Cross-Embodiment Training in $\pi_0$
- $\pi_{0.5}$ and Open-World Generalization
- Hierarchical Prediction in VLAs
- Tokenized Versus Continuous Actions
- The VLA Training Mixture
- Normalizing Heterogeneous Robot Data
- Control Frequency and Inference Latency
- Asynchronous Inference
- Fine-Tuning Strategy
- What Makes a VLA Generalist?
- The Emerging VLA Design Pattern
- Isaac GR00T and Humanoid Foundation Models
- Why Humanoids Are a Distinct Physical AI Problem
- From Project GR00T to GR00T N1
- The Dual-System Architecture
- System 2: Vision-Language Understanding
- System 1: The Diffusion Transformer
- Why Separate Reasoning and Action?
- GR00T’s Data Pyramid
- Learning from Human Video
- Real Robot Demonstrations
- Synthetic Motion Generation
- GR00T-Mimic
- GR00T-Dreams
- DreamGen
- GR00T N1.5
- GR00T N1.5 and Synthetic Post-Training
- GR00T N1.6
- GR00T N1.7 and the Evolution Toward Deployment
- Bimanual Manipulation
- From Manipulation to Whole-Body Control
- Isaac Sim
- Isaac Lab
- Newton Physics Engine
- Domain Randomization
- Sim-and-Real Co-Training
- Embodiment Adaptation
- Post-Training a Humanoid Foundation Model
- Closed-Loop Evaluation
- The Humanoid Data Flywheel
- The GR00T Physical AI Stack
- From Humanoid Models to General Physical Intelligence
- Autonomous Driving as Physical AI
- Autonomous Driving as an Embodied Intelligence Problem
- The Classical Autonomous Driving Stack
- From Modular Pipelines to End-to-End Driving
- Scene Representations for Learned Driving
- Driving as a Multimodal Prediction Problem
- Prediction Is Interactive
- Open-Loop Versus Closed-Loop Evaluation
- Why Long-Tail Scenarios Dominate Autonomous Driving
- Vision-Language Models for Driving
- From Driving VLMs to Driving VLAs
- NVIDIA Alpamayo
- Alpamayo 1
- Chain of Causation Reasoning
- Why Reasoning-Action Alignment Matters
- Training Reasoning VLAs
- Trajectory Generation as an Action Expert
- Diffusion-Based Trajectory Prediction
- Alpamayo 1.5
- Alpamayo 2 Super
- Foundation Models as Teachers
- AlpaSim
- Why Closed-Loop Simulation Changes the Optimization Problem
- AlpaGym
- Closed-Loop RL for Driving
- Counterfactual Driving Experience
- World Models for Autonomous Driving
- Neural Reconstruction and Digital Twins
- Long-Tail Scenario Generation
- The Autonomous Driving Data Engine
- Safety Is a Trajectory Property
- Safety Layers Around Learned Policies
- Reasoning Is Not a Safety Proof
- The Driving Physical AI Stack
- World Models for Physical AI
- Why Physical Agents Need World Models
- World Models as Learned Dynamics
- Why Predict in Latent Space?
- Observation Models Versus State Models
- Deterministic Versus Stochastic Dynamics
- Multi-Step Prediction
- Model-Based Reinforcement Learning
- Planning by Imagination
- Dreamer
- Representation, Dynamics, Reward, and Continuation Models
- World Models as Data Generators
- Video Generation Becomes World Modeling
- Foundation World Models
- NVIDIA Cosmos
- The Cosmos Data Pipeline
- Video Tokenization
- Diffusion World Models
- Google DeepMind Genie
- Genie 2
- Genie 3 and Real-Time Interactive Worlds
- World Models as Infinite Curricula
- Counterfactual Simulation
- World Models for Policy Evaluation
- Model Exploitation
- Uncertainty-Aware World Models
- The Reality Gap
- Neural World Models and Physics Simulators
- Neural Reconstruction and Digital Twins
- Generative Digital Twins
- World Models and Synthetic Data
- World Models and VLAs
- World Models as Critics
- From Static Datasets to Interactive Data Engines
- World Models as Learned Simulators
- The World-Model Physical AI Stack
- The Physical AI Data Engine
- Why Physical AI Is a Data Problem
- What Constitutes a Robot Trajectory?
- Sources of Physical AI Data
- Human Demonstrations
- Teleoperation
- Demonstration Quality Versus Demonstration Quantity
- Language Annotation
- Open X-Embodiment
- RT-X and Positive Transfer Across Robots
- The Cross-Embodiment Normalization Problem
- Embodiment Metadata
- Dataset Standardization
- Temporal Alignment
- Human Video as Physical Pretraining Data
- Learning Representations from Human Activity
- Learning Latent Actions from Video
- Internet-Scale Semantic Data
- Simulation as a Data Source
- Domain Randomization
- Dynamics Randomization
- Synthetic Visual Data
- Procedural Environment Generation
- Synthetic Robot Demonstrations
- Generative Models as Data Engines
- Targeted Data Generation
- GR00T-Dreams as a Synthetic Data Pipeline
- World-Model-Generated Experience
- Autonomous Data Collection
- Human Intervention and Corrections
- DAgger and On-Policy Demonstration Collection
- Success and Failure Data
- Automatic Success Detection
- Data Filtering
- Deduplication
- Data Balancing
- Mixing Data Sources
- Curriculum Learning
- Automatic Curriculum Generation
- Diversity Across Tasks, Objects, and Scenes
- Measuring Dataset Coverage
- Active Data Collection
- Hard-Negative Mining
- Failure Clustering
- The Data Flywheel
- Data Engines Versus Datasets
- The Emerging Physical AI Data Stack
- Training and Post-Training Physical AI Models
- From Foundation Model to Physical Policy
- The Physical AI Training Hierarchy
- Broad Multimodal Pretraining
- Robot Foundation-Model Pretraining
- Pretraining Versus Post-Training
- Supervised Fine-Tuning
- Fine-Tuning Action Tokens
- Continuous-Action Fine-Tuning
- Action Chunking During Post-Training
- Fine-Tuning Flow-Matching Policies
- Which Parameters Should Be Fine-Tuned?
- Post-Training a New Embodiment
- Post-Training for New Skills
- Why Behavioral Cloning Eventually Saturates
- From Imitation to Reinforcement Learning
- Offline RL Post-Training
- Advantage-Weighted Post-Training
- Online RL Post-Training
- RL Post-Training for VLAs
- Reward Construction
- Sparse Rewards
- Dense Rewards
- Learned Reward Models
- Success Models as Rewards
- Reinforcement Learning in Simulation
- Sim-to-Real RL Post-Training
- World Models for RL Post-Training
- Synthetic-Data Post-Training
- Joint Policy and World-Model Objectives
- Fine-Tuning World Models into Policies
- Multi-Task Post-Training
- Catastrophic Forgetting
- Post-Training for Robustness
- Smoothness Regularization
- Safety-Constrained RL
- Curriculum Post-Training
- Hard-Example Post-Training
- Distillation
- Distilling Reasoning into Reactive Policies
- Deployment-Aware Post-Training
- Control Frequency and Model Latency
- Asynchronous Policy Execution
- Real-World Validation
- GR00T as a Pretraining-to-Post-Training Example
- The Physical AI Post-Training Flywheel
- The Emerging Training Stack
- Autonomy and Agentic Physical AI
- From Physical Policies to Physical Agents
- Reactive Versus Deliberative Control
- Hierarchical Autonomy
- Temporal Abstraction
- Skills as the Interface Between Reasoning and Control
- SayCan: Combining Semantic and Physical Feasibility
- Affordance-Grounded Planning
- Language as an Intermediate Action Representation
- Human Intervention at the Semantic Layer
- Closed-Loop Reasoning
- Inner Monologue and Environment Feedback
- Execution Monitoring
- Preconditions and Postconditions
- State Estimation for Agentic Control
- Scene Memory
- Episodic and Semantic Memory
- Tool Use
- Code as a Physical Action Language
- Why Tool Use Matters for Physical AI
- Spatial Reasoning and Value Maps
- Planning with a World Model
- Receding-Horizon Agentic Planning
- Failure Detection
- Recovery Policies
- Retry Versus Replan
- Uncertainty-Aware Autonomy
- Active Perception
- Asking Humans as a Tool
- Gemini Robotics and Embodied Reasoning
- Agentic VLA Architectures
- System 2 and System 1 in Physical Agents
- Event-Triggered Reasoning
- Multi-Agent Physical AI
- Coordination Between Robots
- Human-Robot Interaction
- Interruptibility
- Long-Horizon Error Accumulation
- Progress Tracking
- Planning as Search Over Skills
- Neural Planning Versus Explicit Search
- Agentic Autonomy as Closed-Loop Optimization
- The Agentic Physical AI Stack
- Simulation, Sim-to-Real, and Digital Twins
- Why Simulation Is Central to Physical AI
- The Role of a Robotics Simulator
- Physics Simulation
- Physics Fidelity Versus Simulation Throughput
- GPU-Parallel Simulation
- Simulation for Reinforcement Learning
- The Reality Gap
- Domain Randomization
- Visual Domain Randomization
- Dynamics Randomization
- Sensor Randomization
- Latency Randomization
- System Identification
- System Identification and Domain Randomization Together
- Adaptive Domain Randomization
- Differentiable Simulation
- Newton and Modern Robot Simulation
- Isaac Sim and Isaac Lab
- Synthetic Data Generation
- Procedural Simulation
- From Procedural Generation to Generative Simulation
- Real-to-Sim
- Digital Twins
- Digital Twin Versus Simulator
- Neural Reconstruction
- Gaussian Splatting for Robotic Digital Twins
- Closing the Real-to-Sim Gap
- Counterfactual Simulation
- Digital Twins for Failure Replay
- Software-in-the-Loop Testing
- Hardware-in-the-Loop Testing
- Shadow-Mode Evaluation
- Sim-to-Sim Transfer
- Sim-to-Real as Distribution Generalization
- A Practical Sim-to-Real Recipe
- Sim-to-Real for Perception
- Sim-to-Real for Control
- Residual Adaptation
- Privileged Learning in Simulation
- Simulating Rare Events
- Simulation as an Evaluation Environment
- Scenario-Based Evaluation
- Simulation as a Data Engine
- The Real-Sim-Real Flywheel
- The Emerging Simulation Stack for Physical AI
- Evaluation for Physical AI
- Why Physical AI Evaluation Is Different
- Open-Loop Versus Closed-Loop Evaluation
- Why Action Error Can Be Misleading
- Task Success Rate
- Partial Credit
- Long-Horizon Evaluation
- Sequence-Length Metrics
- Evaluating Generalization
- Object Generalization
- Language Generalization
- Environment Generalization
- Compositional Generalization
- Lifelong and Knowledge-Transfer Evaluation
- Robustness Evaluation
- Perturbation Sweeps
- Recovery Evaluation
- Time-to-Recovery
- Intervention Rate
- Autonomy Duration
- Safety Metrics
- Severity-Weighted Safety
- Constraint-Based Evaluation
- Efficiency Metrics
- Success Weighted by Efficiency
- Evaluation of Vision-Language-Action Models
- LIBERO
- CALVIN
- BEHAVIOR-1K
- Simulation Versus Real-World Evaluation
- SimplerEnv and Real-to-Sim Evaluation
- Evaluating Sim-to-Real Correlation
- Autonomous Driving Evaluation
- nuPlan
- Interactive Driving Evaluation
- Waymo Open Dataset Evaluation
- Evaluating World Models
- Evaluating World Models by Decision Quality
- Counterfactual Evaluation
- Evaluating Agentic Physical AI
- Failure Taxonomy
- Conditional Metrics
- Confidence Intervals
- Repeated Trials
- Deterministic Versus Stochastic Policy Evaluation
- Tail-Risk Evaluation
- Stress Testing
- Adversarial Scenario Generation
- Evaluation as a Data Flywheel
- Regression Evaluation
- Evaluation Gates
- A Physical AI Evaluation Scorecard
- The Emerging Evaluation Stack
- Safety and Reliability for Physical AI
- Why Physical AI Safety Is Different
- Defense in Depth
- The Safety Stack
- Semantic Safety and Robot Constitutions
- Physical Constraints
- Why Reward Penalties Are Not Enough
- Safe Reinforcement Learning
- Runtime Assurance and Shielding
- Control Barrier Functions
- Predictive Safety
- Uncertainty and Out-of-Distribution Detection
- Task Feasibility and Ambiguity
- Human Proximity and External Monitoring
- Functional Safety and Fail-Safe Behavior
- Watchdogs and Human Override
- Action and Workspace Constraints
- Safety for Vision-Language-Action Models
- Success-Safety Gap
- Reward Hacking and Specification Gaming
- Safety During Post-Training
- Adversarial Safety Evaluation
- Formal Verification and Safety Envelopes
- Safety Cases and Operational Design Domains
- Fault Injection and Redundancy
- Production Safety Monitoring
- Safety Regression Testing
- The Emerging Physical AI Safety Architecture
- Engineering a Physical AI System
- From Models to Systems
- The End-to-End Physical AI Stack
- Separating Decision Timescales
- Slow and Fast Loops
- Sensor Layer
- Time Synchronization
- Coordinate Frames
- Perception Preprocessing
- State Estimation
- World State Versus Raw Context
- Memory Architecture
- Agentic Reasoning Layer
- Skill Interfaces
- VLA as a Motor Skill
- Action Representation
- Relative Versus Absolute Actions
- Action Denormalization
- Action Chunk Execution
- Temporal Ensembling
- Asynchronous Inference
- Latency Budget
- Observation Age
- Edge Versus Cloud Inference
- NVIDIA Jetson and Edge Physical AI
- Compute Scheduling
- Model Optimization
- Robot Middleware
- ROS 2 Communication Patterns
- Quality of Service
- NVIDIA Isaac ROS
- Motion Planning
- MoveIt
- Learned and Classical Control
- Low-Level Control
- Impedance Control
- Safety Runtime
- Failure Detection
- Recovery Orchestration
- Retry Budgets
- Real-Time Versus Best-Effort Components
- Watchdogs and Heartbeats
- Observability
- Distributed Tracing
- Data Logging
- Dataset Versioning
- Model Registry
- Configuration Is Part of the Model
- Simulation in the Development Loop
- Hardware Abstraction
- Digital Twin Integration
- Continuous Evaluation
- Hardware-in-the-Loop
- Shadow Mode
- Canary Deployment
- Rollback
- Fleet Learning
- Failure Mining
- The Physical AI Data Flywheel
- Multi-Robot Infrastructure
- Connectivity Failure
- Power and Thermal Constraints
- Deterministic Infrastructure Around Stochastic Models
- State Machines Around Foundation Models
- Interface Contracts
- Testing Interfaces
- Reproducibility
- Security
- Signed Model Artifacts
- An End-to-End Reference Architecture
- A Concrete Manipulation Example
- Engineering Principles
- The Physical AI Production Flywheel
- Open Problems and Research Directions
- From Specialized Robots to General Physical Intelligence
- The Robotics Data Bottleneck
- Scaling Robot Data
- Learning Actions from Human Video
- Latent Actions
- The Embodiment Gap
- Universal Action Representations
- Morphology-Conditioned Policies
- Zero-Shot Embodiment Transfer
- Long-Horizon Autonomy
- Error Accumulation
- Hierarchical Autonomy
- Unified Models Versus Modular Agents
- System 1 and System 2 Physical Intelligence
- Continual Learning
- Continual Post-Training
- Learning During Deployment
- Few-Shot Skill Acquisition
- Learning from Corrections
- Autonomous Data Collection
- Self-Improvement Through Practice
- World Models as Physical Reasoning Engines
- Long-Horizon World-Model Consistency
- Physics-Aware World Models
- World Models as Simulators
- Planner Exploitation
- Dexterous Manipulation
- Tactile Intelligence
- Force-Aware Foundation Models
- Deformable Objects
- Tool Use
- Humanoid Whole-Body Intelligence
- Locomotion and Manipulation Must Converge
- Dynamic Whole-Body Tasks
- Sample-Efficient Reinforcement Learning
- Post-Training Foundation Policies with RL
- Reward Models for Physical AI
- Automatic Success Detection
- Open-World Perception
- Learning Object Physics
- Active Perception
- Spatial Intelligence
- Persistent 3D World Models
- Memory and Physical AI
- Multi-Agent Physical Intelligence
- Robot-to-Robot Knowledge Transfer
- Human-Robot Collaboration
- Legibility
- Personalized Physical Assistance
- Evaluation Remains a Bottleneck
- Measuring Generality
- Measuring Adaptation Cost
- Evaluating Emergent Physical Capabilities
- Safety Under Generalization
- Calibrated Uncertainty
- Physical AI Interpretability
- Causal Physical Reasoning
- Compositional Physical Intelligence
- Open-Ended Skill Libraries
- Physical Reasoning Versus Language Reasoning
- Scaling Laws for Physical AI
- Data Quality Versus Scale
- Hardware-Software Co-Design
- Designing Robots for Learnability
- Efficient On-Device Foundation Models
- Adaptive Compute
- From Robot Foundation Models to Physical Foundation Models
- A Physical Foundation Model Objective
- The Role of Simulation May Change
- Automated Curriculum Generation
- Toward Autonomous Robot Research
- What Would a General Physical Agent Require?
- A Possible Convergence
- From Foundation Models to Foundation Agents
- The Long-Term Direction
- References
- Physical AI and Embodied Intelligence
- Language-Grounded Planning and Agentic Robotics
- Robot Transformers and Vision-Language-Action Models
- Continuous and Generative Robot Policies
- NVIDIA Isaac GR00T and Humanoid Foundation Models
- Google DeepMind Robotics
- Robot Learning and Post-Training
- World Models for Physical AI
- NVIDIA Cosmos and World Foundation Models
- Simulation, Sim-to-Real, and Digital Twins
- Autonomous Driving as Physical AI
- Physical AI Data Engines
- Robot-Learning Benchmarks and Evaluation
- Safety and Reliability for Physical AI
- Autonomous-Driving Safety
- Lifelong Learning and Cross-Embodiment Transfer
- Physical AI Systems and Platforms
- Physical AI Surveys and Research Directions
- Citation
Overview
What is Physical AI?
-
Physical AI refers to artificial intelligence systems that perceive, reason about, and act within the physical world. Unlike models whose outputs terminate in digital artifacts such as text, images, or software, a Physical AI system closes the loop between computation and physical consequences:
\[\text{Perceive} \rightarrow \text{Understand} \rightarrow \text{Reason} \rightarrow \text{Plan} \rightarrow \text{Act} \rightarrow \text{Observe consequences}\] - The defining feature is therefore not simply that AI is installed inside a physical machine. The system must continually map multimodal observations of the world into actions, observe how those actions alter the world, and adapt subsequent decisions accordingly.
-
A useful abstraction is a policy
\[\pi_\theta(a_t \mid o_{\leq t}, g)\]-
where \(o_{\leq t}\) represents the observation history, \(g\) represents a goal or instruction, and \(a_t\) represents an action executed in the physical environment. The action changes the state of that environment, producing the next observation:
\[s_{t+1} \sim P(s_{t+1}\mid s_t,a_t)\] \[o_{t+1} \sim O(o_{t+1}\mid s_{t+1})\]
-
- Physical AI therefore sits at the intersection of machine learning, robotics, computer vision, reinforcement learning, planning, control, simulation, and increasingly multimodal foundation models.
- NVIDIA Cosmos World Foundation Model Platform for Physical AI by NVIDIA et al. (2025) frames the problem around two complementary learned representations: a policy model representing the intelligent agent and a world model representing the environment in which the agent operates.
From Digital Intelligence to Physical Intelligence
-
Large language models operate primarily over discrete tokens:
\[x_1,x_2,\ldots,x_T\] -
A language model learns
\[p(x_t\mid x_{<t})\]- and generates another token. The cost of an incorrect prediction is generally informational rather than immediately physical.
-
Physical intelligence introduces a fundamentally different loop. An autonomous system may observe camera images, depth, LiDAR, proprioception, force measurements, navigation information, or other sensors:
\[o_t = \{ I_t^{1:N}, d_t, q_t, \dot q_t, f_t, \ldots \}\]- and must convert them into physically executable outputs such as joint positions, torques, end-effector poses, gripper commands, steering, acceleration, braking, or trajectories.
-
The system therefore learns a mapping closer to
\[(o_{\leq t},g) \xrightarrow{\pi_\theta} a_{t:t+H}\]- where \(a_{t:t+H}\) can be an action chunk or trajectory spanning a future horizon \(H\).
- The important distinction is that the prediction becomes an intervention. An erroneous text token can be regenerated; an erroneous steering trajectory, grasp, or humanoid motion can cause an irreversible physical transition.
- This makes several problems that are secondary in conventional generative AI central to Physical AI: latency, dynamics, geometry, uncertainty, collision avoidance, embodiment constraints, temporal consistency, recovery, and safety.
The Perception-Reasoning-Action Loop
- A useful conceptual decomposition consists of five interacting layers.
Perception
- The system converts raw sensor measurements into representations of its surroundings. Modern Physical AI increasingly uses pretrained vision or vision-language representations rather than task-specific perception alone.
-
For a robot receiving images \(I_t\) and language instruction \(l\), a multimodal encoder may produce
\[z_t = f_{\mathrm{VLM}}(I_t,l)\] - This representation can encode objects, spatial relationships, semantic concepts, task intent, and eventually information relevant to action generation.
- PaLM-E: An Embodied Multimodal Language Model by Driess et al. (2023) was an important step toward directly incorporating continuous sensor modalities into a language-model-centered architecture, establishing a path from multimodal foundation models toward embodied reasoning.
Reasoning
- Reasoning determines what the observations mean relative to the agent’s goal. For example, a household robot instructed to “put away the groceries” must infer objects, destinations, ordering constraints, graspability, and possibly intermediate actions.
- Reasoning is particularly valuable when instructions underspecify the required physical behavior.
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances by Ahn et al. (2022) introduced SayCan, which combines language-model estimates of useful actions with learned affordance estimates of which actions are feasible in the current environment.
-
One simplified interpretation is
\[\pi(a\mid s,l) \propto p_{\mathrm{LLM}}(a\mid l) \, p_{\mathrm{affordance}}(\text{success}\mid s,a)\] - The first term asks whether an action makes semantic sense; the second asks whether the physical world currently permits it.
- This distinction remains central to Physical AI. A model may know that a cup belongs in a cupboard without knowing whether the robot can currently reach the cupboard.
Planning
- Planning converts goals into future actions or intermediate subgoals.
-
Classical robotics often makes this separation explicit:
\[\text{Perception} \rightarrow \text{World Model} \rightarrow \text{Planner} \rightarrow \text{Controller}\] - Foundation-model-based systems increasingly blur these boundaries. A VLA may predict low-level actions directly, while another architecture may use a large multimodal model for semantic planning and a specialized action model for continuous control.
- This latter separation is especially important in modern humanoid foundation models.
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots by NVIDIA et al. (2025) uses a dual-system architecture in which a vision-language component interprets observations and instructions while a Diffusion Transformer generates continuous robot actions.
-
The architecture illustrates an emerging pattern:
\[\text{slow semantic reasoning} + \text{fast continuous control}\] - This resembles the division between high-level deliberation and low-level motor behavior required by biological agents.
Action
- The final output must be translated into commands supported by the embodiment.
-
For a manipulator, an action might be
\[a_t = [ \Delta x, \Delta y, \Delta z, \Delta \phi, \Delta \theta, \Delta \psi, g ]\]- where the first six terms describe an end-effector pose displacement and \(g\) controls the gripper.
- Humanoids can require much higher-dimensional action vectors involving two arms, hands, torso, legs, and locomotion. Autonomous vehicles instead predict steering and acceleration commands or future trajectories.
- This diversity creates the embodiment problem: two agents may understand the same instruction but require completely different motor representations to execute it.
Feedback
-
Physical action generates new observations, making control inherently closed-loop:
\[o_t \rightarrow a_t \rightarrow o_{t+1} \rightarrow a_{t+1} \rightarrow \cdots\] - This distinction between open-loop prediction and closed-loop interaction becomes especially important in autonomous driving.
- A trajectory that looks plausible when evaluated against a recorded driving log may behave very differently when its own actions alter subsequent states. NVIDIA’s Alpamayo platform therefore couples reasoning VLAs with AlpaSim for closed-loop simulation and AlpaGym for reinforcement learning against the consequences of driving decisions.
Physical AI, Embodied AI, and Autonomy
- These terms overlap but describe somewhat different concepts.
- Embodied AI focuses on intelligence situated in an agent that interacts with an environment through perception and action. Embodiment matters because the agent’s capabilities, observations, and learning signals are constrained by its body.
- Physical AI is somewhat broader in current usage. It encompasses embodied robots but also autonomous vehicles, industrial systems, drones, and other AI systems whose decisions interact directly with physical dynamics.
- Autonomy describes the degree to which such a system can pursue objectives without continuous human control.
-
A useful hierarchy is:
\[\text{Physical AI} \supset \{ \text{Robotics}, \text{Autonomous Driving}, \text{Industrial Autonomy}, \ldots \}\] - Embodied intelligence provides many of the learning principles underlying these systems, while autonomy describes how independently the resulting system operates.
- A robot executing a preprogrammed trajectory is physical automation but exhibits little learned autonomy. A robot that understands a natural-language goal, recognizes unfamiliar objects, decomposes the task, manipulates them, detects failures, and replans represents a substantially richer form of Physical AI.
From Robotics Models to Robot Foundation Models
-
Traditional learned robot policies were typically specialized:
\[\pi_{\theta}^{(\text{task},\text{robot},\text{environment})}\] - Changing the robot, camera configuration, task, or environment could require training another policy.
-
The foundation-model paradigm instead seeks a shared policy:
\[\pi_\theta ( a \mid o, l, e )\]- where \(e\) describes or implicitly identifies the embodiment.
-
The objective is to transfer knowledge across
\[\text{tasks} \times \text{environments} \times \text{objects} \times \text{embodiments}\] - Open X-Embodiment: Robotic Learning Datasets and RT-X Models by Open X-Embodiment Collaboration et al. (2023) demonstrated this direction by standardizing data from 22 robot embodiments and showing positive transfer from cross-robot training.
- Octo: An Open-Source Generalist Robot Policy by Octo Model Team et al. (2024) subsequently trained a transformer policy on approximately 800,000 Open X-Embodiment trajectories and designed the model so that new observations and action spaces could be incorporated during fine-tuning.
-
The shift mirrors the evolution of NLP:
\[\text{task-specific model} \rightarrow \text{pretrained foundation model} \rightarrow \text{adaptation}\] -
In robotics:
\[\text{task-specific policy} \rightarrow \text{generalist robot foundation model} \rightarrow \text{embodiment/task post-training}\]
Vision-Language-Action Models
- One of the most consequential architectural developments in Physical AI is the Vision-Language-Action model, or VLA.
-
A VLA extends the familiar vision-language mapping
\[(I,l) \rightarrow \text{text}\]-
into
\[(I,l,s) \rightarrow a\]
-
- The model therefore connects semantic representations learned from large-scale image-text corpora to physically executable behavior.
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control by Brohan et al. (2023) established the VLA formulation by representing robot actions as tokens and co-fine-tuning vision-language models on web-scale vision-language tasks and robotic trajectories. This allowed semantic knowledge learned from the web to influence robot behavior.
- OpenVLA: An Open-Source Vision-Language-Action Model by Kim et al. (2024) provided an open 7B-parameter implementation trained on 970,000 real-world robot demonstrations, combining DINOv2 and SigLIP visual features with a Llama-family language backbone.
- The central insight is that robot control no longer has to begin from a blank policy network. A Physical AI system can inherit semantic knowledge about objects, concepts, relationships, and instructions from large-scale multimodal pretraining and then ground this knowledge through robot demonstrations.
Continuous Generative Action Models
- Physical actions differ fundamentally from language tokens. Robot motion is typically continuous, temporally correlated, and multimodal.
-
Suppose two valid demonstrations move around opposite sides of an obstacle. A regression policy minimizing
\[\mathcal{L}_{\mathrm{MSE}} = \mathbb{E} \left[ \left\| a-\pi_\theta(o) \right\|_2^2 \right]\]- can average the two behaviors and generate an action corresponding to neither valid trajectory.
- Generative action models address this by learning distributions over actions.
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion by Chi et al. (2023) formulates robot control as conditional denoising diffusion, allowing policies to represent multimodal high-dimensional action distributions and generate action sequences under receding-horizon control.
- This idea has become increasingly important in robot foundation models. GR00T, for example, combines semantic VLM representations with a Diffusion Transformer action model rather than requiring the language model itself to generate every continuous motor command.
-
The architecture can be abstracted as
\[z = f_{\mathrm{VLM}}(I,l)\]-
followed by
\[a_{t:t+H} = f_{\mathrm{action}}(z,s_t)\]
-
- This separation allows large-scale semantic representations and specialized continuous control architectures to complement each other.
Physical AI as a Data Problem
- The largest obstacle to scaling Physical AI is not simply model architecture. It is data.
- Internet-scale language and images are comparatively inexpensive. High-quality physical interaction trajectories require robots, environments, operators, sensors, calibration, and time.
-
A robot demonstration can be represented as
\[\tau = (o_0,a_0,o_1,a_1,\ldots,o_T)\] -
Large robot foundation models require diverse distributions over such trajectories:
\[\mathcal{D} = \bigcup_{i=1}^{N} \mathcal{D}_i^{\text{embodiment}} \cup \mathcal{D}^{\text{simulation}} \cup \mathcal{D}^{\text{human}} \cup \mathcal{D}^{\text{synthetic}}\] -
This motivates a Physical AI data flywheel:
\[\text{Real Data} \rightarrow \text{Model} \rightarrow \text{Simulation/Synthetic Data} \rightarrow \text{Training} \rightarrow \text{Deployment} \rightarrow \text{New Real Data}\] - GR00T N1 explicitly mixes real robot trajectories, human videos, and synthetic data, while later NVIDIA GR00T workflows use simulation and generative world models to expand the training distribution.
- Open X-Embodiment attacks the same bottleneck from another direction by pooling heterogeneous robot datasets, allowing experience from one embodiment to improve another.
World Models and Simulation
-
A policy answers:
\[\text{What should I do?}\] -
A world model addresses:
\[\text{What will happen if I do it?}\] -
Formally, a learned dynamics model approximates
\[p_\phi(s_{t+1}\mid s_t,a_t)\] -
Given candidate future actions,
\[a_t,a_{t+1},\ldots,a_{t+H}\]- the agent can predict possible futures before committing to them physically.
- Mastering Diverse Domains through World Models by Hafner et al. (2023) develops DreamerV3, which learns environmental dynamics and improves its policy using imagined trajectories, illustrating the broader model-based principle that agents can learn from predicted experience rather than exclusively from physical interaction.
- Cosmos World Foundation Model Platform for Physical AI by NVIDIA et al. (2025) scales this idea toward general-purpose world foundation models that can be post-trained for robotics and autonomous-system applications.
-
This leads to an increasingly important architecture for Physical AI:
\[\boxed{ \text{Policy Foundation Model} + \text{World Foundation Model} + \text{Physics Simulator} }\] - The policy proposes behavior, the world model predicts plausible outcomes, and simulation provides an environment in which those behaviors can be trained and evaluated without incurring the cost and risk of every experiment in the real world.
Physical AI Across Robots and Autonomous Vehicles
- Robotics and autonomous driving appear different at the actuator level but increasingly share the same learning abstractions.
-
For manipulation:
\[(\text{camera},\text{language},\text{proprioception}) \rightarrow \text{arm/hand trajectory}\] -
For autonomous driving:
\[(\text{multi-camera video},\text{route},\text{vehicle state}) \rightarrow \text{vehicle trajectory}\] - NVIDIA’s Alpamayo applies the VLA paradigm to autonomous vehicles, processing multi-camera observations and driving context to produce trajectories together with explicit Chain-of-Causation reasoning traces.
- As of September 2026, the family extends to Alpamayo 2 Super, a 34B reasoning VLA accompanied by AlpaGym for closed-loop RL and generative simulation infrastructure for long-tail driving scenarios.
-
The architectural convergence is significant. Humanoid robotics and autonomous driving increasingly share the same broad recipe:
\[\text{multimodal perception} + \text{foundation-model reasoning} + \text{continuous action generation} + \text{world models/simulation} + \text{closed-loop learning}\] - What differs is the embodiment, dynamics, action space, safety envelope, and deployment stack.
Toward Generalist Physical Agents
- The long-term objective is not simply a robot capable of executing a larger catalog of memorized skills. It is an agent that can transfer knowledge to new physical situations.
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization by Physical Intelligence et al. (2025) combines robot data, web data, high-level semantic prediction, and heterogeneous multimodal examples to target this problem, demonstrating long-horizon manipulation in previously unseen homes.
-
General physical intelligence therefore requires generalization along several axes:
\[G = G_{\text{task}} \times G_{\text{object}} \times G_{\text{scene}} \times G_{\text{embodiment}} \times G_{\text{dynamics}} \times G_{\text{instruction}}\] - A system that generalizes only to a differently colored object has solved a much narrower problem than one that transfers a learned concept between robot embodiments, environments, and task compositions.
-
This is why Physical AI increasingly looks less like conventional robotics with a larger neural network and more like a complete foundation-model stack for interacting with reality:
\[\boxed{ \begin{array}{c} \text{Multimodal Foundation Models}\\ \downarrow\\ \text{World Models + Physical Reasoning}\\ \downarrow\\ \text{Planning + VLA Policies}\\ \downarrow\\ \text{Continuous Control}\\ \downarrow\\ \text{Physical Environment}\\ \downarrow\\ \text{Feedback + Data}\\ \circlearrowleft \end{array} }\] - The remainder of the primer will unpack each layer, beginning with the foundations of embodied intelligence and the progression from classical perception-planning-control pipelines to modern multimodal embodied foundation models.
Foundations of Embodied Intelligence
Embodiment as a Computational Constraint
- Embodied intelligence studies agents whose cognition is coupled to a body and an environment. The central idea is that intelligence in the physical world cannot be reduced to mapping abstract inputs to abstract outputs: what an agent can observe, infer, and accomplish depends on its sensors, actuators, morphology, dynamics, and interaction history.
- A Survey on Robotics with Foundation Models: toward Embodied AI by Xu et al. (2024) characterizes embodied AI as integrating perception, learning, reasoning, decision-making, control, and generalization so that agents can perform tasks in open, dynamic environments.
- Consider two agents observing the same cup. A language model may represent the cup semantically, while a robot must additionally reason about whether the cup is reachable, how it should be grasped, whether it is full, how much force is appropriate, and how its own kinematics constrain the motion.
-
Thus, an embodied agent operates under a coupling
\[\text{Intelligence} = f( \text{Model}, \text{Body}, \text{Environment}, \text{Interaction} )\] - The body is not merely the output device for an otherwise independent intelligence. It determines the agent’s action space, observation space, reachable states, and physical constraints.
-
For example, a manipulator might expose an action
\[a_t = [ \Delta x_t, \Delta y_t, \Delta z_t, \Delta r_t, \Delta p_t, \Delta y_t^{\mathrm{rot}}, g_t ]\]-
while a humanoid may instead expose dozens of joint targets:
\[a_t = [ q_{1,t}^{*}, q_{2,t}^{*}, \ldots, q_{n,t}^{*} ]\]
-
-
A mobile robot may operate through linear and angular velocity,
\[a_t = [v_t,\omega_t]\]- and an autonomous vehicle may execute steering, acceleration, and braking commands.
- Consequently, the same semantic objective,
-
move the object from the table to the shelf,
- induces different policies for different embodiments.
- This gives rise to the embodiment gap: knowledge may transfer between robots even when executable behavior does not. The Embodiment Gap in Robot Foundation Models by Domae et al. (2026) formalizes this distinction by separating reusable semantic, perceptual, data, and representation-level knowledge from the adaptation still required to execute successfully on a new robot.
The Agent-Environment Loop
- The canonical abstraction for embodied intelligence is an agent interacting sequentially with an environment.
- At time \(t\):
-
- the environment occupies state \(s_t\); 2. the agent receives observation \(o_t\); 3. the policy selects action \(a_t\); 4. the environment transitions to \(s_{t+1}\); 5. the agent receives the next observation.
-
Formally,
\[o_t \sim O(o_t\mid s_t)\] \[a_t \sim \pi_\theta(a_t\mid o_{\leq t},g)\] \[s_{t+1}\sim P(s_{t+1}\mid s_t,a_t)\] -
The resulting trajectory is
\[\tau = (s_0,o_0,a_0,s_1,o_1,a_1,\ldots,s_T)\] - This sequential structure is fundamental. Physical intelligence is not a collection of independent predictions because an action changes the distribution of observations the agent will subsequently receive.
- For example, moving a camera changes the visual observation. Moving a robot arm may occlude an object. Picking up an object changes its pose. Steering a vehicle changes which future traffic configurations are encountered.
-
The policy therefore influences its own future input distribution:
\[\pi_\theta \rightarrow P(s_{t+1}) \rightarrow P(o_{t+1}) \rightarrow \pi_\theta\] - This feedback loop is one reason physical agents are substantially harder to evaluate than static vision or language models.
Embodied Intelligence as a POMDP
- Most real-world embodied tasks are naturally modeled as partially observable Markov decision processes, or POMDPs.
-
A POMDP can be written as
\[\mathcal{M} = (\mathcal{S}, \mathcal{A}, \mathcal{O}, P, O, R, \gamma)\]- where:
- \(\mathcal{S}\) is the latent physical state space;
- \(\mathcal{A}\) is the action space;
- \(\mathcal{O}\) is the observation space;
- \(P(s_{t+1}\mid s_t,a_t)\) is the transition model;
- \(O(o_t\mid s_t)\) is the observation model;
- \(R(s_t,a_t)\) is the reward;
- \(\gamma\) is the discount factor.
- The agent generally cannot observe \(s_t\) directly.
- A camera cannot see behind an occluding object. A robot may not know an object’s mass until interacting with it. A vehicle approaching an intersection cannot directly observe another driver’s intention.
-
The policy must therefore condition on history,
\[\pi_\theta(a_t\mid o_{\leq t},a_{<t},g)\]-
or maintain an internal state
\[h_t = f_\theta( h_{t-1}, o_t, a_{t-1} )\]
-
-
The policy then becomes
\[a_t\sim\pi_\theta(a_t\mid h_t,g)\] -
Transformers are particularly attractive here because the context window can implicitly function as a finite interaction memory:
\[h_t = \operatorname{Transformer} ( o_{t-k:t}, a_{t-k:t-1}, g )\] - This connection helps explain why architectures developed for sequence modeling transferred naturally into robotics.
Perception Is Not Merely Object Recognition
- Classical computer vision often decomposes perception into tasks such as detection, segmentation, pose estimation, depth estimation, and tracking.
-
Embodied perception asks a more operational question:
\[\text{What information is required to act successfully?}\] -
Suppose a robot must pick up a mug. Recognizing the category “mug” is insufficient. The policy may require:
\[\{ \text{location}, \text{orientation}, \text{handle pose}, \text{free space}, \text{reachability}, \text{grasp points} \}\] - This distinction motivates representations optimized for action rather than exclusively for semantic recognition.
- CLIPort: What and Where Pathways for Robotic Manipulation by Shridhar et al. (2021) combines CLIP-derived semantic representations for determining “what” should be manipulated with Transporter-style spatial representations for determining “where” manipulation should occur, demonstrating how semantic and geometric information can complement each other in language-conditioned manipulation.
-
A useful abstraction is therefore
\[z_t = f_{\mathrm{perception}}(o_t)\]- where \(z_t\) should preserve information relevant to the downstream policy rather than reconstruct every property of the environment.
Learning Visual Representations from Human Activity
- Robot datasets are expensive, but humans generate enormous quantities of video showing interactions with the physical world.
- This motivates learning representations from human video and transferring them to robots.
- R3M: A Universal Visual Representation for Robot Manipulation by Nair et al. (2022) pretrains visual representations on Ego4D human video using temporal contrastive learning, video-language alignment, and regularization, then uses the representation as a frozen perceptual backbone for robot policies.
-
Conceptually,
\[\text{Human Video} \xrightarrow{\text{pretraining}} f_\phi(I) \xrightarrow{\text{robot data}} \pi_\theta(a\mid f_\phi(I))\] - The important assumption is that visual regularities such as object state, contact, motion, and human-object interaction contain information that transfers across embodiments.
-
Human video does not directly specify robot actions:
\[a_t^{\mathrm{human}} \neq a_t^{\mathrm{robot}}\] - But it can teach representations of the physical world that make robot-specific learning more data efficient.
- This distinction between transferable perceptual knowledge and embodiment-specific control will recur throughout Physical AI.
Grounding Language in Physical Interaction
- Language models can encode substantial semantic knowledge but do not inherently know whether a particular physical action is executable in the current environment.
- Consider the instruction:
-
Bring me the apple.
-
A language model can infer plausible subgoals:
\[\text{find apple} \rightarrow \text{approach apple} \rightarrow \text{pick apple} \rightarrow \text{return}\] - But successful execution additionally requires answering questions such as:
- Is an apple visible?
- Is it reachable?
- Is the gripper already occupied?
- Can the robot navigate to it?
- Is the proposed grasp physically feasible?
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances by Ahn et al. (2022) introduced SayCan to connect semantic reasoning from large language models with learned robotic affordances.
-
Let \(l\) denote the instruction and \(s\) the current physical state. A simplified decomposition is
\[\operatorname{score}(a) = p_{\mathrm{LLM}}(a\mid l) \, Q(s,a)\]-
where
\[p_{\mathrm{LLM}}(a\mid l)\]-
captures whether the action is useful for satisfying the instruction, while
\[Q(s,a)\]- estimates whether the robot can successfully execute it.
-
-
-
The highest-scoring action is selected:
\[a^* = \arg\max_a p_{\mathrm{LLM}}(a\mid l)Q(s,a)\] -
This separates two distinct questions:
\[\text{Should I do this?}\]-
and
\[\text{Can I do this?}\]
-
- The distinction between semantic plausibility and physical feasibility remains one of the foundational ideas behind modern embodied agents.
Affordances
- The concept underlying SayCan is broader than robotics foundation models.
- An affordance describes an action that an environment or object makes available to an agent.
-
A cup may afford
\[\{ \text{grasp}, \text{lift}, \text{pour}, \text{place} \}\]-
while a closed drawer may afford
\[\{ \text{grasp handle}, \text{pull} \}\]
-
- Critically, affordances depend on the agent.
- A shelf may be reachable for a humanoid but not for a short mobile manipulator. A heavy object may be liftable by one robot but not another.
-
Thus an affordance is more accurately modeled as
\[A(s,e,a)\]- where \(s\) describes the environment, \(e\) describes the embodiment, and \(a\) is a candidate action.
- This makes affordances a natural bridge between semantic reasoning and physical control.
From Modular Robotics to Learned Policies
-
Classical autonomous robots are frequently organized as explicit modules:
\[\boxed{ \text{Sensors} \rightarrow \text{Perception} \rightarrow \text{State Estimation} \rightarrow \text{Planning} \rightarrow \text{Control} \rightarrow \text{Actuators} }\] - Each component exposes a structured interface.
-
For example:
\[I_t \xrightarrow{\text{detector}} \text{objects} \xrightarrow{\text{pose estimator}} \text{6D poses} \xrightarrow{\text{planner}} \text{trajectory} \xrightarrow{\text{controller}} \text{torques}\] - This architecture has substantial engineering advantages. Intermediate states are interpretable, components can be independently tested, and known physical constraints can be incorporated explicitly.
- However, every interface potentially discards information.
-
A learned policy instead attempts
\[\pi_\theta: (o_t,g)\rightarrow a_t\] -
The extreme version is end-to-end control:
\[\text{pixels} \rightarrow \text{actions}\] - Modern Physical AI increasingly occupies a middle ground. Large learned representations replace many hand-engineered perception and reasoning components, while structured planners, controllers, safety systems, or specialized action heads remain useful.
-
The important architectural question is therefore no longer simply “modular or end-to-end.” It is:
\[\text{Which abstractions should be learned, and where should structure be retained?}\]
Imitation Learning as the Basic Robot-Learning Primitive
- Much modern embodied AI begins with demonstrations.
-
Given a dataset
\[\mathcal{D} = \{ (o_t^{(i)},g^{(i)},a_t^{(i)}) \}\]-
behavioral cloning trains the policy to reproduce demonstrated actions:
\[\mathcal{L}_{\mathrm{BC}}(\theta) = -\mathbb{E}_{(o,g,a)\sim\mathcal{D}} [ \log \pi_\theta(a\mid o,g) ]\]
-
-
For continuous deterministic actions, this may reduce to regression:
\[\mathcal{L}_{\mathrm{BC}} = \mathbb{E} \left[ \| a-\hat a_\theta \|_2^2 \right]\] - The attraction is straightforward: humans demonstrate successful behavior, and the robot learns the mapping from observations and instructions to actions.
- The difficulty is distribution shift.
-
During training,
\[o_t \sim p_{\mathrm{expert}}(o)\] -
During deployment,
\[o_t \sim p_{\pi_\theta}(o)\] - A small prediction error may place the robot in a state absent from the demonstration dataset. Subsequent predictions then become less reliable, producing compounding error.
- This is one reason temporal action modeling, large diverse datasets, intervention data, simulation, and reinforcement-learning post-training become important.
Scaling Demonstrations Enables Generalization
- Early robot-learning systems often learned one task from one dataset. A major transition occurred when researchers began asking whether a sufficiently diverse demonstration distribution could produce general-purpose behavior.
- BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning by Jang et al. (2022) scaled imitation learning beyond 100 manipulation tasks and conditioned policies using language or human-video task representations, demonstrating zero-shot execution on previously unseen manipulation tasks.
-
Instead of training
\[\pi_{\theta,k}(a\mid o)\]-
for each task \(k\), the objective becomes a shared conditioned policy
\[\pi_\theta(a\mid o,g)\]
-
-
The task descriptor \(g\) can take multiple forms:
\[g \in \{ \text{text}, \text{image}, \text{video}, \text{goal state}, \text{demonstration} \}\] - This transition is foundational because it converts task identity from something encoded in the model weights into something specified at inference time.
Multimodal Prompting for Robots
- Human instructions are rarely restricted to text. A person can communicate a task by pointing, demonstrating, showing an image, or combining visual and linguistic cues.
- VIMA: General Robot Manipulation with Multimodal Prompts by Jiang et al. (2022) formalizes robot tasks as multimodal prompts interleaving visual and textual tokens, and trains a transformer to autoregressively produce motor actions across thousands of procedurally generated manipulation tasks.
-
A prompt might conceptually take the form
\[p = [ \text{"put"}, I_{\mathrm{object}}, \text{"inside"}, I_{\mathrm{container}} ]\] -
The policy becomes
\[\pi_\theta(a_t\mid o_{\leq t},p)\] - This is a significant conceptual change. Rather than defining each robotic task through a separate reward function or policy, tasks become prompts to a shared embodied model.
- The idea foreshadows modern VLA models, where natural-language instructions and visual observations jointly condition physical actions.
Action Chunking and Temporal Abstraction
- A robot policy need not predict only one action at a time.
-
Instead of
\[\pi_\theta(o_t)\rightarrow a_t\]-
it can predict a chunk
\[\pi_\theta(o_t) \rightarrow (a_t,a_{t+1},\ldots,a_{t+H})\]
-
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware by Zhao et al. (2023) introduces Action Chunking with Transformers (ACT), using a conditional VAE and transformer architecture to generate sequences of actions for precise bimanual manipulation while mitigating compounding errors and temporal variability in human demonstrations.
-
ACT learns a distribution of action sequences,
\[p_\theta(a_{t:t+H}\mid o_t,z)\]- where \(z\) is a latent variable encoding variation in demonstrations.
-
A generic conditional VAE objective is
\[\mathcal{L}_{\mathrm{ACT}} = \mathbb{E}_{q_\phi(z\mid a,o)} [ -\log p_\theta(a\mid o,z) ] + \beta D_{\mathrm{KL}} \left( q_\phi(z\mid a,o) \| p(z) \right)\] -
Action chunking provides temporal abstraction:
\[\text{observation} \rightarrow \underbrace{ [a_t,\ldots,a_{t+H}] }_{\text{coherent motion}}\]- rather than requiring the policy to reconstruct the intended trajectory independently at every control step.
- This concept later becomes important in VLAs and diffusion/flow-based action models.
Transformers Enter Robot Control
- The same sequence-modeling machinery that transformed language modeling also provides a natural abstraction for robot trajectories.
-
A robot trajectory can be represented as a sequence:
\[[ g, o_1, a_1, o_2, a_2, \ldots, o_T, a_T ]\] -
A causal transformer can model
\[p(a_t\mid g,o_{\leq t},a_{<t})\] - RT-1: Robotics Transformer for Real-World Control at Scale by Brohan et al. (2022) demonstrated this approach at substantial real-world scale, training on roughly 130,000 episodes spanning more than 700 tasks collected using 13 robots over 17 months.
- RT-1 takes a short image history and natural-language task instruction and produces tokenized robot actions. Its visual pathway uses an ImageNet-pretrained EfficientNet-B3, language-conditioned FiLM layers, and TokenLearner compression before a transformer predicts discretized actions.
-
Conceptually,
\[I_{t-k:t} \xrightarrow{\mathrm{EfficientNet}} F_{t-k:t}\] \[F_{t-k:t} \xrightarrow{\mathrm{FiLM}(g)} \tilde F_{t-k:t}\] \[\tilde F \xrightarrow{\mathrm{TokenLearner}} z_{1:N}\]-
followed by
\[(z_{1:N},g) \xrightarrow{\mathrm{Transformer}} a_t\]
-
-
Continuous action dimensions are discretized into bins, converting control into categorical prediction:
\[\mathcal{L} = -\sum_j \log p_\theta(a_t^{(j)}\mid o_{\leq t},g)\] -
The importance of RT-1 extends beyond its particular architecture. It demonstrated that three scaling dimensions familiar from foundation models matter in robot learning:
\[\boxed{ \text{model capacity} + \text{data volume} + \text{task diversity} }\]- and that a single policy could absorb a broad distribution of real-world robotic experience.
From Robotics Transformers to Generalist Agents
- The broader foundation-model hypothesis suggests going beyond multiple robot tasks and training one model across entirely different modalities and embodiments.
- A Generalist Agent by Reed et al. (2022) introduced Gato, a single transformer capable of processing and producing sequences spanning text, images, Atari actions, simulated environments, and real robot control.
- The central abstraction is tokenization.
-
Different modalities are serialized into a common sequence:
\[x = [x_1,x_2,\ldots,x_T]\]- and the transformer predicts appropriate targets conditioned on context.
-
Depending on that context, the same network may produce
\[x_{t+1} \in \{ \text{text token}, \text{game action}, \text{robot action}, \ldots \}\] -
Gato is important less because it solved general robotics and more because it demonstrated that the sequence-modeling paradigm could extend beyond language:
\[\text{general sequence model} \rightarrow \text{general agent}\] - This idea directly anticipates today’s multimodal Physical AI models.
Embodied Multimodal Language Models
- Another line of development begins with large language models rather than robot policies.
- Instead of converting everything into a robot-specific architecture, multimodal observations can be injected directly into an LLM.
- PaLM-E: An Embodied Multimodal Language Model by Driess et al. (2023) introduced this approach by interleaving continuous sensor-derived representations with language tokens inside a large language-model backbone.
-
Conceptually, an image encoder produces
\[z_I=f_{\mathrm{vision}}(I)\]-
while robot state may be encoded as
\[z_s=f_{\mathrm{state}}(s)\]
-
-
These embeddings are inserted alongside text embeddings:
\[[ e(w_1), \ldots, z_I, \ldots, z_s, \ldots, e(w_n) ] \xrightarrow{\mathrm{LLM}} y\] - This allows the model to reason jointly over linguistic concepts and embodied observations rather than treating perception as a completely separate symbolic preprocessing stage.
-
The broader implication is significant:
\[\text{LLM} + \text{sensor grounding} \rightarrow \text{embodied multimodal model}\] - Once such models are additionally trained to generate executable actions, the progression naturally leads toward Vision-Language-Action models.
The Transition from VLMs to VLAs
-
A vision-language model learns mappings such as
\[(I,l)\rightarrow y_{\mathrm{text}}\] -
An embodied multimodal model adds physical state:
\[(I,s,l)\rightarrow y_{\mathrm{text}}\] -
A Vision-Language-Action model closes the final gap:
\[(I,s,l)\rightarrow a\] - RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control by Brohan et al. (2023) makes this transition explicit by representing robot actions as another output vocabulary and co-fine-tuning vision-language models on both web-scale vision-language examples and robotic trajectories.
-
The conceptual progression can therefore be summarized as
\[\boxed{ \begin{array}{ccccc} \text{Vision Model} & \rightarrow & \text{Vision-Language Model} & \rightarrow & \text{Vision-Language-Action Model} \\ I\rightarrow z && (I,l)\rightarrow y && (I,l,s)\rightarrow a \end{array} }\] - The crucial change is not merely adding another modality. Actions alter the environment. Once a foundation model emits actions, it becomes part of a feedback-controlled dynamical system.
Cross-Embodiment Learning
- The next scaling frontier is not merely many tasks on one robot, but many tasks across many robots.
-
Suppose robot \(i\) produces demonstrations
\[\mathcal{D}_i = \{ (o_t^{(i)},a_t^{(i)},g^{(i)}) \}\] -
Cross-embodiment training attempts to learn
\[\pi_\theta \sim \bigcup_{i=1}^{N} \mathcal{D}_i\] -
The challenge is that
\[\mathcal{A}_i \neq \mathcal{A}_j\]- in general. Robots differ in degrees of freedom, coordinate systems, grippers, cameras, control rates, and kinematics.
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models by Open X-Embodiment Collaboration et al. (2023) addresses this problem by standardizing heterogeneous datasets from 22 robot embodiments into a shared format and training RT-X models across the resulting mixture.
- This introduces an important distinction between semantic universality and motor universality.
-
Semantic concepts such as
\[\text{"pick up the cup"}\]- can be shared broadly.
-
The actual motor realization
\[a_{1:T}^{(i)}\]- remains embodiment dependent.
-
A scalable Physical AI system therefore needs some combination of:
\[\text{shared semantics} + \text{shared physical representations} + \text{embodiment-specific adaptation}\]
Generalization as the Central Objective
- Embodied intelligence is ultimately a problem of generalization under interaction.
-
A useful decomposition is
\[G = \{ G_{\mathrm{object}}, G_{\mathrm{task}}, G_{\mathrm{scene}}, G_{\mathrm{language}}, G_{\mathrm{embodiment}}, G_{\mathrm{dynamics}} \}\] -
An increasingly difficult sequence might be:
\[\text{known task + known object}\] \[\downarrow\] \[\text{known task + novel object}\] \[\downarrow\] \[\text{novel task composition}\] \[\downarrow\] \[\text{novel environment}\] \[\downarrow\] \[\text{novel embodiment}\] \[\downarrow\] \[\text{open-world physical autonomy}\] - VIMA explicitly evaluates increasingly difficult levels of generalization over multimodal manipulation prompts, while BC-Z investigates zero-shot task generalization and Open X-Embodiment examines transfer across robot datasets and embodiments.
- These are not independent capabilities. General-purpose Physical AI requires them simultaneously.
The Emerging Embodied Foundation-Model Stack
- The evolution from classical robotics toward embodied foundation models can now be summarized as several architectural transitions.
-
First, perception shifted from task-specific visual pipelines toward reusable pretrained representations:
\[\text{task-specific vision} \rightarrow \text{foundation representations}\] -
Second, task identity moved from separate policy weights into prompts:
\[\pi_{\theta,k}(a\mid o) \rightarrow \pi_\theta(a\mid o,g)\] -
Third, robot control shifted from specialized architectures toward scalable sequence models:
\[\text{specialized policy} \rightarrow \text{Transformer policy}\] -
Fourth, semantic reasoning became increasingly grounded in physical observations:
\[\text{LLM} \rightarrow \text{embodied multimodal model}\] -
Finally, actions themselves became outputs of foundation models:
\[\text{VLM} \rightarrow \text{VLA}\] -
Together these developments produce the modern embodied Physical AI stack:
\[\boxed{ \begin{array}{c} \text{Language / Goal / Prompt}\\ \downarrow\\ \text{Multimodal Perception}\\ \downarrow\\ \text{Semantic + Spatial Representation}\\ \downarrow\\ \text{Reasoning / Task Planning}\\ \downarrow\\ \text{Action Policy}\\ \downarrow\\ \text{Robot Controller}\\ \downarrow\\ \text{Physical World}\\ \circlearrowleft \end{array} }\] - The key research question has consequently shifted. The problem is no longer simply how to make a robot execute a particular behavior. It is how to construct models that can acquire broad knowledge from heterogeneous data, ground that knowledge through physical interaction, transfer it across tasks and embodiments, and reliably convert high-level intent into temporally coherent actions.
- The next section will build directly on these foundations by examining “Learning to Act”: behavioral cloning, imitation learning, offline and online reinforcement learning, action chunking, diffusion and flow-matching policies, hierarchical control, and the training objectives used to turn demonstrations and interaction into executable physical behavior.
Learning to Act
From Understanding the World to Controlling It
-
Perception and reasoning are useful only insofar as an embodied agent can convert them into successful physical behavior. The central learning problem in Physical AI is therefore to construct a policy
\[\pi_\theta(a_t \mid o_{\leq t}, g)\]- that maps observations \(o_{\leq t}\) and a goal \(g\) into actions \(a_t\).
- The difficulty is that physical action spaces are fundamentally different from text. They are typically continuous, high-dimensional, temporally correlated, multimodal, embodiment-specific, and constrained by dynamics.
-
A robot may need to predict
\[a_t = [ \Delta x_t, \Delta y_t, \Delta z_t, \Delta r_t, \Delta p_t, \Delta y_t^{\mathrm{rot}}, g_t ]\]- while a humanoid might simultaneously control dozens of joints.
-
The learning problem is therefore not merely
\[\text{What action is correct?}\]-
but rather
\[\text{What coherent sequence of physically executable actions will achieve the goal?}\]
-
-
Modern Physical AI approaches this problem through several increasingly powerful paradigms:
\[\boxed{ \text{Behavioral Cloning} \rightarrow \text{Generative Policies} \rightarrow \text{Offline RL} \rightarrow \text{Online RL} \rightarrow \text{Model-Based / World-Model Learning} }\] - These approaches are complementary rather than mutually exclusive. A modern robot foundation model can be pretrained by imitation, represented using a diffusion or flow-matching action head, and subsequently improved through reinforcement learning.
Behavioral Cloning
- The simplest approach to learning physical behavior is behavioral cloning.
-
Suppose an expert generates demonstrations
\[\mathcal{D} = \{ (o_t^{(i)},g^{(i)},a_t^{(i)}) \}_{i,t}\] -
Behavioral cloning treats policy learning as supervised learning:
\[\theta^* = \arg\min_\theta \mathbb{E}_{(o,g,a)\sim\mathcal{D}} [ \mathcal{L}( \pi_\theta(o,g), a ) ]\] -
For a probabilistic policy,
\[\mathcal{L}_{\mathrm{BC}} = - \mathbb{E}_{(o,g,a)\sim\mathcal{D}} [ \log \pi_\theta(a\mid o,g) ]\] -
For deterministic continuous actions, a simple implementation can use
\[\mathcal{L}_{\mathrm{MSE}} = \mathbb{E} \left[ \| a-\hat a_\theta(o,g) \|_2^2 \right]\] - The appeal of behavioral cloning is considerable. Robot demonstrations can be converted directly into supervised examples, and training can use the same large-scale optimization infrastructure as other foundation models.
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware by Zhao et al. (2023) demonstrates how imitation learning combined with temporal action prediction can produce precise bimanual manipulation from human demonstrations.
- However, behavioral cloning contains two fundamental limitations: distribution shift and multimodality.
Distribution Shift and Compounding Errors
-
The demonstrations in behavioral cloning are collected under the expert policy:
\[s_t \sim d^{\pi_E}(s)\] -
Deployment instead produces states according to the learned policy:
\[s_t \sim d^{\pi_\theta}(s)\] -
Even if
\[\pi_\theta(a\mid s) \approx \pi_E(a\mid s)\]- on the training distribution, small errors alter subsequent states.
-
Suppose the expert follows
\[s_0 \xrightarrow{a_0^*} s_1^* \xrightarrow{a_1^*} s_2^*\] -
The learned policy may instead execute
\[s_0 \xrightarrow{\hat a_0} \hat s_1\] -
If
\[\hat s_1 \notin \operatorname{support} ( \mathcal{D} )\]- the next action prediction is now made from an unfamiliar state.
-
Errors can therefore compound:
\[\epsilon_1 \rightarrow \Delta s_1 \rightarrow \epsilon_2 \rightarrow \Delta s_2 \rightarrow \cdots\] - This problem motivated interactive imitation-learning methods such as DAgger, which repeatedly deploy a learned policy and request expert labels on states that the learner itself encounters. The resulting dataset increasingly approximates the learner’s deployment distribution rather than only the original expert distribution.
-
The underlying principle remains important for modern Physical AI:
\[\boxed{ \text{training data should cover states induced by the learned policy} }\]- rather than exclusively states generated by successful expert demonstrations.
Why Mean-Squared Error Can Fail for Physical Actions
- Physical behavior is frequently multimodal.
-
Consider a robot moving around an obstacle. Two demonstrations may be equally valid:
\[a^{(1)} = \text{move left}\] \[a^{(2)} = \text{move right}\] -
A deterministic regression model minimizing
\[\mathcal{L}_{\mathrm{MSE}} = \| a-\hat a \|^2\]-
is encouraged to predict the conditional mean
\[\hat a = \mathbb{E}[a\mid o]\]
-
-
For a multimodal action distribution, the mean may correspond to no successful behavior:
\[\frac{1}{2} a^{(1)} + \frac{1}{2} a^{(2)} = \text{move directly toward obstacle}\] -
Thus,
\[\boxed{ \text{average of valid actions} \neq \text{valid action} }\]- in many physical control problems.
-
This observation motivates generative action models that represent the complete conditional distribution
\[p(a\mid o,g)\]- rather than predicting a single point estimate.
Action Chunking
- Physical actions are also strongly correlated through time.
-
Instead of predicting
\[a_t = \pi_\theta(o_t)\]-
a policy can generate an action chunk
\[A_t = [ a_t, a_{t+1}, \ldots, a_{t+H-1} ]\]
-
-
The policy becomes
\[A_t \sim \pi_\theta( A_t \mid o_{\leq t},g )\] - Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware by Zhao et al. (2023) introduced Action Chunking with Transformers, or ACT, which predicts temporally coherent action sequences rather than independently predicting every low-level command.
- Action chunking provides several advantages.
-
First, it captures temporal structure:
\[a_t \not\perp a_{t+1}\] -
Second, it reduces the effective planning horizon. A task requiring \(T\) low-level actions becomes approximately
\[\frac{T}{H}\]- high-level prediction decisions when chunks of length \(H\) are used.
-
Third, the model can represent coherent motion primitives such as
\[\text{approach} \rightarrow \text{grasp} \rightarrow \text{lift}\]- rather than reconstructing the intended trajectory independently at every control cycle.
Receding-Horizon Execution
- Predicting an entire action sequence does not mean executing the entire sequence blindly.
-
Suppose the model predicts
\[A_t = [a_t,\ldots,a_{t+H-1}]\] -
The controller may execute only the first \(K\) actions,
\[K < H\]-
then observe the environment again and replan:
\[o_t \xrightarrow{\pi} A_t \xrightarrow{\text{execute first }K} o_{t+K} \xrightarrow{\pi} A_{t+K}\]
-
- This is receding-horizon control.
-
It combines the temporal coherence of action chunks with the robustness of feedback:
\[\boxed{ \text{long prediction horizon} + \text{short execution horizon} }\] - Diffusion Policy: Visuomotor Policy Learning via Action Diffusion by Chi et al. (2023) makes receding-horizon execution a central component of its visuomotor policy architecture, combining action-sequence generation with repeated visual feedback.
Action Chunking with Transformers
- ACT models action sequences using a conditional variational autoencoder.
-
During training, an encoder receives the demonstration action sequence and contextual observations and infers a latent variable
\[z \sim q_\phi( z\mid A_t,o_t )\] -
The decoder predicts the action chunk:
\[\hat A_t = f_\theta( o_t,z )\] -
A simplified objective is
\[\mathcal{L}_{\mathrm{ACT}} = \mathcal{L}_{\mathrm{reconstruction}} + \beta D_{\mathrm{KL}} \left( q_\phi(z\mid A_t,o_t) \| p(z) \right)\] - The latent variable captures variability in demonstrations, while the transformer models relationships across observations, robot state, and the future action sequence.
- This approach is particularly useful for teleoperated demonstrations, where human behavior can exhibit substantial timing and trajectory variation even when the task outcome is identical.
Diffusion Policies
- Diffusion models provide a more expressive solution to multimodal action generation.
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion by Chi et al. (2023) formulates robot policy learning as conditional denoising diffusion over action trajectories, demonstrating strong performance across diverse manipulation benchmarks and highlighting multimodal action modeling, high-dimensional control, and training stability as key advantages.
-
Instead of directly predicting
\[A_t = f_\theta(o_t)\] -
Diffusion Policy begins with noise
\[A_t^{K} \sim \mathcal{N}(0,I)\]-
and iteratively denoises it:
\[A_t^{K} \rightarrow A_t^{K-1} \rightarrow \cdots \rightarrow A_t^0\]
-
- The final sample \(A_t^0\) is the predicted action trajectory.
-
During training, noise is added to an expert action sequence:
\[A^k = \sqrt{\bar\alpha_k}A^0 + \sqrt{1-\bar\alpha_k}\epsilon\]-
where
\[\epsilon\sim\mathcal{N}(0,I)\]
-
-
The network learns to predict the noise:
\[\epsilon_\theta ( A^k, k, o )\] -
A standard diffusion objective is
\[\mathcal{L}_{\mathrm{diff}} = \mathbb{E}_{A^0,\epsilon,k} \left[ \| \epsilon - \epsilon_\theta( A^k,k,o ) \|_2^2 \right]\] - At inference time, iterative denoising generates an action trajectory conditioned on the current observation.
Why Diffusion Works Well for Robot Actions
- Diffusion models have several properties that align naturally with physical control.
Multimodality
-
A diffusion policy can represent
\[p(A\mid o)\]- with several distinct modes rather than averaging incompatible trajectories.
High-dimensional outputs
-
An action chunk can be represented as
\[A \in \mathbb{R}^{H\times d_a}\] -
For a humanoid or bimanual robot, both \(H\) and \(d_a\) can be large. Diffusion models can generate the complete structured object jointly.
Temporal coherence
-
The policy models
\[p( a_t, a_{t+1}, \ldots, a_{t+H} \mid o )\]- rather than treating actions as conditionally independent.
Stable supervised training
- Diffusion policy training remains essentially supervised learning over demonstration trajectories, avoiding many optimization instabilities associated with online reinforcement learning.
- The tradeoff is inference cost. A diffusion policy may require several network evaluations to produce one action sequence, so the number of denoising steps directly affects control frequency and compute requirements.
Diffusion Transformers for Large Robot Policies
- The combination of diffusion objectives with transformer backbones naturally extends to foundation-scale robot models.
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation by Liu et al. (2024) scales a Robotics Diffusion Transformer to approximately 1.2B parameters, introduces a physically interpretable unified action representation across robots, and pretrains on heterogeneous multi-robot data before task-specific fine-tuning.
-
A generic architecture is
\[z_{\mathrm{vision}} = f_{\mathrm{vision}}(I)\] \[z_{\mathrm{text}} = f_{\mathrm{text}}(g)\] \[z_{\mathrm{state}} = f_{\mathrm{state}}(s)\]-
followed by
\[\epsilon_\theta = \operatorname{DiT} ( A^k, k, z_{\mathrm{vision}}, z_{\mathrm{text}}, z_{\mathrm{state}} )\]
-
- The Diffusion Transformer therefore functions as an action expert conditioned on representations supplied by perception and language systems.
From Diffusion to Flow Matching
- Flow matching learns a continuous vector field that transports samples from a simple source distribution to the target action distribution.
-
Let
\[x_0\sim p_0\]-
be noise and
\[x_1\sim p_{\mathrm{data}}\]-
be an expert action trajectory. A simple interpolation is
\[x_\tau = (1-\tau)x_0+\tau x_1, \qquad \tau\in[0,1]\]
-
-
-
The target velocity is
\[u_\tau=x_1-x_0\] -
A neural network learns
\[v_\theta(x_\tau,\tau,c)\approx u_\tau\]- where \(c\) contains conditioning information such as images, language, and robot state.
-
The flow-matching objective can be written
\[\mathcal{L}_{\mathrm{FM}} = \mathbb{E} \left[ \| v_\theta(x_\tau,\tau,c)-u_\tau \|_2^2 \right]\] -
At inference time, actions are generated by integrating
\[\frac{dx}{d\tau} = v_\theta(x,\tau,c)\]- from noise toward the learned action distribution.
Flow Matching in Modern VLAs
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control by Black et al. (2024) applies flow matching to robot action generation, combining a pretrained vision-language model with a dedicated action expert that produces continuous action chunks.
-
The key architectural separation is
\[\text{VLM} \rightarrow \text{semantic representation}\] \[\text{action expert} \rightarrow \text{continuous control}\] - The successor $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization by Black et al. (2025) extends this approach with heterogeneous co-training across robot data, web data, semantic subtasks, object detections, and low-level actions to improve long-horizon generalization in unfamiliar environments.
Diffusion Versus Flow Matching
-
At a high level, both approaches learn generative action distributions:
\[p_\theta(A\mid o,g)\] - Diffusion commonly learns a denoising or score function and solves a reverse stochastic process, while flow matching learns a velocity field and solves an ODE.
-
Conceptually:
\[\text{Diffusion}:\text{ iterative denoising}\] \[\text{Flow Matching}:\text{ continuous transport}\] - Both are attractive for Physical AI because they naturally model multimodal, high-dimensional action sequences.
Why Imitation Alone Is Not Enough
-
Imitation learning asks:
\[\text{What did the demonstrator do?}\] -
Reinforcement learning asks:
\[\text{What behavior maximizes the objective?}\] -
The reinforcement-learning objective is
\[J(\pi) = \mathbb{E}_{\tau\sim\pi} \left[ \sum_{t=0}^{T} \gamma^t r(s_t,a_t) \right]\] -
Rather than matching actions, the policy is optimized for consequences. This makes RL particularly attractive for post-training a Physical AI model that already possesses useful behavior from imitation learning.
Value Functions and Physical Decision Making
-
The action-value function is
\[Q^\pi(s,a) = \mathbb{E}_\pi \left[ \sum_{k=0}^{\infty} \gamma^k r_{t+k} \mid s_t=s,a_t=a \right]\] -
The state-value function is
\[V^\pi(s) = \mathbb{E}_{a\sim\pi}[Q^\pi(s,a)]\] -
The advantage function is
\[A^\pi(s,a)=Q^\pi(s,a)-V^\pi(s)\] -
These functions allow robot learning to move beyond pure imitation by preferring actions associated with higher eventual success.
Offline Reinforcement Learning
-
Real-world robot interaction is expensive and potentially dangerous. Physical AI therefore often begins with static datasets
\[\mathcal{D} = \{(s_t,a_t,r_t,s_{t+1})\}\] - Offline RL attempts to improve a policy using this fixed dataset without additional environment interaction.
-
The central challenge is extrapolation. A learned Q-function may assign high value to actions absent from the dataset:
\[a_{\mathrm{OOD}}\notin\mathcal{D}\] - Because estimates for such actions are poorly constrained, maximizing \(Q_\theta(s,a)\) can select actions whose apparent value results from function-approximation error.
Conservative Q-Learning
- Conservative Q-Learning for Offline Reinforcement Learning by Kumar et al. (2020) addresses offline-RL distribution shift by penalizing high Q-values for actions unsupported by the dataset.
-
A simplified objective is
\[\mathcal{L}_{\mathrm{CQL}} = \mathcal{L}_{\mathrm{Bellman}} + \alpha \left[ \mathbb{E}_{s}\log\sum_a\exp Q(s,a) - \mathbb{E}_{(s,a)\sim\mathcal{D}}Q(s,a) \right]\] -
The intuition is:
\[\boxed{ \text{do not trust apparently excellent actions that the data cannot support} }\]
Implicit Q-Learning
- Offline Reinforcement Learning with Implicit Q-Learning by Kostrikov et al. (2021) avoids evaluating unseen actions during policy improvement, using expectile regression to estimate high-value behavior within the support of the offline dataset.
-
IQL fits the value function using
\[\mathcal{L}_V = \mathbb{E}_{(s,a)\sim\mathcal{D}} [ L_\tau^2(Q(s,a)-V(s)) ]\]-
where
\[L_\tau^2(u) = |\tau-\mathbb{1}(u<0)|u^2\]
-
-
Policy extraction resembles advantage-weighted behavioral cloning:
\[\mathcal{L}_{\pi} = - \mathbb{E}_{(s,a)\sim\mathcal{D}} \left[ \exp(\beta A(s,a)) \log\pi_\theta(a\mid s) \right]\] -
The practical interpretation is
\[\text{imitate the dataset} + \text{prefer its better behaviors}\]
Advantage-Weighted Policy Learning
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning by Peng et al. (2019) formulates policy learning as weighted maximum likelihood with weights determined by estimated advantages.
-
A typical objective is
\[\mathcal{L}_{\mathrm{AWR}} = - \mathbb{E}_{(s,a)\sim\mathcal{D}} [ w(s,a)\log\pi_\theta(a\mid s) ]\]-
with
\[w(s,a)=\exp\left(\frac{A(s,a)}{\beta}\right)\]
-
- This creates a useful bridge between imitation learning and RL because policy updates remain structurally similar to supervised post-training while incorporating reward information.
Online Reinforcement Learning
-
Online RL allows the policy to interact with the environment:
\[\pi_\theta \rightarrow \tau \rightarrow r \rightarrow \text{update }\theta \rightarrow \pi_{\theta'}\] - This enables exploration and direct correction of policy-induced distribution shift, but real-world online RL faces sample cost, hardware wear, reset cost, safety risk, and human-supervision constraints.
-
Consequently, practical Physical AI systems often use a staged recipe:
\[\boxed{ \text{large-scale pretraining} \rightarrow \text{behavioral cloning} \rightarrow \text{offline RL} \rightarrow \text{simulation RL} \rightarrow \text{limited real-world online RL} }\]
Reward Design
-
For a manipulation task, a sparse reward might be
\[r_t = \begin{cases} 1 & \text{task completed}\\ 0 & \text{otherwise}. \end{cases}\] -
A shaped reward may instead include
\[r_t = w_1r_{\mathrm{reach}} + w_2r_{\mathrm{grasp}} + w_3r_{\mathrm{transport}} + w_4r_{\mathrm{success}} - w_5r_{\mathrm{collision}}\] -
Reward shaping improves credit assignment but creates the possibility that optimizing the specified reward does not produce the intended behavior. Physical AI therefore increasingly combines simulator state, task-completion detectors, learned reward models, vision-language models, human feedback, and rule-based safety constraints.
Hierarchical Control
-
Long-horizon tasks create enormous low-level action horizons. Hierarchical control introduces different temporal scales:
\[g \xrightarrow{\pi_{\mathrm{high}}} z_k \xrightarrow{\pi_{\mathrm{low}}} a_{t:t+H}\] -
The high-level policy predicts subgoals or skills, while the low-level policy converts each skill into continuous actions. This separates the semantic horizon from the motor horizon.
System 2 and System 1 Control
- An increasingly useful abstraction for Physical AI is a two-timescale architecture.
- A slower reasoning system handles scene understanding, task decomposition, planning, and recovery reasoning. A faster motor system handles continuous trajectories, reactive control, and contact dynamics.
-
Conceptually,
\[\boxed{ \text{System 2} \xrightarrow{\text{subgoal}} \text{System 1} \xrightarrow{\text{actions}} \text{World} }\]- with feedback returning to both systems.
- This division is useful because high-level reasoning and low-level motor control have very different latency and compute requirements.
Model-Based Reinforcement Learning
-
Model-based RL additionally learns or uses a transition model:
\[\hat s_{t+1}=f_\phi(s_t,a_t)\] -
The agent can evaluate hypothetical futures:
\[s_t \xrightarrow{a_t} \hat s_{t+1} \xrightarrow{a_{t+1}} \hat s_{t+2} \rightarrow\cdots\] -
An action sequence can be selected by maximizing predicted return:
\[a_{t:t+H}^* = \arg\max_{a_{t:t+H}} \sum_{k=0}^{H} \gamma^k \hat r(\hat s_{t+k},a_{t+k})\] -
This creates the connection
\[\boxed{ \text{learn dynamics} \rightarrow \text{imagine futures} \rightarrow \text{choose actions} }\]- that underlies world-model-based Physical AI.
Latent World Models
-
Instead of predicting raw future pixels, an encoder can map observations into latent states:
\[z_t=e_\phi(o_t)\]-
with learned latent dynamics
\[\hat z_{t+1}=f_\phi(z_t,a_t)\]
-
- TD-MPC2: Scalable, Robust World Models for Continuous Control by Hansen et al. (2024) performs local trajectory optimization inside a learned decoder-free latent world model and demonstrates scaling across many continuous-control tasks and embodiments.
-
A simplified planning objective is
\[A^* = \arg\max_A \left[ \sum_{k=0}^{H-1} \gamma^k R_\phi(z_k,a_k) + \gamma^H Q_\phi(z_H,a_H) \right]\] - The model need not reconstruct every pixel. It only needs a latent representation sufficiently accurate for evaluating actions.
Training in Simulation
-
One solution to the cost of physical interaction is to move reinforcement learning into simulation:
\[s_{t+1}^{\mathrm{sim}} = F_{\mathrm{sim}}(s_t^{\mathrm{sim}},a_t)\] - Simulation provides large-scale parallel rollouts, automatic resets, and privileged reward signals.
-
The central problem becomes sim-to-real transfer:
\[P_{\mathrm{sim}}(s'|s,a) \neq P_{\mathrm{real}}(s'|s,a)\] - Differences in friction, mass, lighting, sensor noise, actuator response, object geometry, and contact dynamics can cause policies trained in simulation to fail physically. This motivates domain randomization, system identification, digital twins, and learned world models.
Combining Demonstrations and Reinforcement Learning
- A powerful Physical AI training recipe combines imitation and RL.
-
Start with expert demonstrations
\[\mathcal{D}_E\]-
train an initial policy
\[\pi_0=\operatorname{BC}(\mathcal{D}_E)\]- collect or generate policy experience \(\mathcal{D}_\pi\), estimate outcome quality, and improve the policy while remaining near useful demonstrated behavior.
-
-
The pipeline becomes
\[\boxed{ \text{Demonstrations} \rightarrow \text{SFT / BC} \rightarrow \text{Rollouts} \rightarrow \text{Rewards} \rightarrow \text{RL Post-Training} \rightarrow \text{Improved Policy} }\] - This resembles post-training pipelines for language models, but the rollouts occur in physical or simulated environments and the outputs are actions rather than text.
Closed-Loop Policy Improvement
-
Deployment generates trajectories
\[\tau_k\sim P(\tau\mid\pi_k)\] -
Failures reveal weaknesses in the policy’s own state distribution. New training data can focus on
\[\mathcal{D}_{k+1} = \mathcal{D}_k \cup \mathcal{D}_{\mathrm{failure}} \cup \mathcal{D}_{\mathrm{recovery}}\] -
Retraining produces \(\pi_{k+1}\), and the process repeats:
\[\boxed{ \pi_k \rightarrow \text{Deploy} \rightarrow \text{Observe Failures} \rightarrow \text{Collect Corrections} \rightarrow \text{Train} \rightarrow \pi_{k+1} }\] -
The most valuable data is often concentrated near the current policy’s failure boundary.
Learning Recovery Behavior
- Robust autonomy requires behavior under off-nominal states, not merely successful demonstrations.
-
A robust policy should learn
\[\pi(a_{\mathrm{recover}}\mid s_{\mathrm{off-nominal}})\]-
in addition to
\[\pi(a_{\mathrm{expert}}\mid s_{\mathrm{nominal}})\]
-
-
This suggests
\[\boxed{ \text{competence} = \text{successful execution} + \text{failure detection} + \text{recovery} }\]- as a more complete definition of long-horizon physical autonomy.
Action Representations
-
A policy might predict joint positions
\[a_t=q_t^*\]-
joint velocities
\[a_t=\dot q_t^*\]-
joint torques
\[a_t=\tau_t\]-
or Cartesian end-effector commands
\[a_t=[\Delta x,\Delta R,g]\]
-
-
-
- Lower-level actions expose more control but require the learned model to understand more dynamics. Higher-level actions simplify learning but rely more heavily on downstream controllers.
-
This creates a hierarchy:
\[\text{semantic skill} \rightarrow \text{Cartesian trajectory} \rightarrow \text{joint trajectory} \rightarrow \text{torque}\] - A foundation model does not necessarily need to operate at the bottom of this hierarchy.
Cross-Embodiment Action Spaces
-
Foundation models introduce the challenge
\[\mathcal{A}_{\mathrm{robot\ A}} \neq \mathcal{A}_{\mathrm{robot\ B}}\] - RDT-1B addresses this with a physically interpretable unified action representation intended to preserve physical semantics while supporting heterogeneous robots.
-
More generally,
\[a_t^{(e)}=f_e(z_t)\]- where \(z_t\) is a shared latent action representation and \(f_e\) maps that representation into embodiment-specific commands.
- This separates what motion should occur from how a particular robot realizes that motion.
The Modern Action-Learning Stack
-
The evolution of learned control can be summarized as several transitions:
\[\text{single-action regression} \rightarrow \text{action chunking}\] \[\text{unimodal regression} \rightarrow \text{generative action distributions}\] \[\text{behavior cloning} \rightarrow \text{reward-aware policy improvement}\] \[\text{offline demonstrations} \rightarrow \text{closed-loop interaction}\]-
and
\[\text{model-free control} \rightarrow \text{world-model-assisted learning}\]
-
-
The resulting architecture increasingly resembles
\[\boxed{ \begin{array}{c} \text{Multimodal Observation + Goal}\\ \downarrow\\ \text{Foundation Representation}\\ \downarrow\\ \text{High-Level Reasoning / Subgoal}\\ \downarrow\\ \text{Generative Action Expert}\\ \downarrow\\ \text{Action Chunk}\\ \downarrow\\ \text{Low-Level Controller}\\ \downarrow\\ \text{Physical Environment}\\ \downarrow\\ \text{Reward + New Observation}\\ \circlearrowleft \end{array} }\] - The central trend is a shift from learning a deterministic mapping from images to individual motor commands toward learning rich distributions over temporally extended behavior and improving those distributions using interaction.
- This provides the algorithmic foundation for modern Vision-Language-Action models. The next section will examine those models directly, including RT-2, OpenVLA, Octo, \(\pi_0\), \(\pi_{0.5}\), tokenized versus continuous action representations, cross-embodiment training, and the engineering details required to turn a pretrained multimodal model into a generalist robot policy.
Vision-Language-Action Models
- A detailed discourse on Vision-Language-Action models and World Models is offered in the World Models primer.
From Vision-Language Models to Physical Policies
- Vision-Language-Action models, or VLAs, are one of the central architectural abstractions in modern Physical AI. They extend vision-language models by adding physically executable actions to the model’s output space.
-
A conventional vision-language model approximates
\[p_\theta(y\mid I,l)\]- where \(I\) is an image, \(l\) is a textual prompt, and \(y\) is a textual response.
-
A VLA instead models
\[\pi_\theta(a\mid I,l,s)\]- where \(s\) may contain proprioceptive or other robot state and \(a\) represents an executable action.
-
For temporally extended control, this becomes
\[\pi_\theta( a_{t:t+H} \mid I_{\leq t}, s_{\leq t}, l )\] -
The conceptual change appears small:
\[\boxed{ \text{Vision} + \text{Language} \rightarrow \text{Text} }\]-
becomes
\[\boxed{ \text{Vision} + \text{Language} + \text{Robot State} \rightarrow \text{Action} }\]- but this changes the role of the model fundamentally. Its predictions now participate directly in a closed-loop dynamical system.
-
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control by Brohan et al. (2023) formalized this VLA paradigm by co-fine-tuning pretrained vision-language models on both web-scale vision-language tasks and robot trajectories, representing robot actions in the same token vocabulary used for language.
Why Start from a Vision-Language Model?
- A robot policy trained exclusively on robot demonstrations can only learn concepts represented in those demonstrations.
-
Suppose
\[\mathcal{D}_{\mathrm{robot}}\]- contains demonstrations involving cups, bowls, fruit, and utensils but no demonstration involving a dinosaur toy.
-
A robot-only policy has little reason to understand the concept
\[\text{"dinosaur"}\] - A VLM pretrained on internet-scale image-text data may already encode such knowledge.
-
The VLA hypothesis is therefore
\[\boxed{ \text{Internet semantic knowledge} + \text{robot interaction data} \rightarrow \text{semantically capable physical policy}. }\] - The robot data teaches the model how semantic concepts map into physical actions, while vision-language pretraining provides much broader conceptual coverage.
-
This creates two forms of generalization:
\[G_{\mathrm{physical}}\]-
from robot demonstrations and
\[G_{\mathrm{semantic}}\]- from vision-language pretraining.
-
- An effective VLA attempts to preserve both.
The Generic VLA Architecture
-
A modern VLA can be decomposed into four conceptual stages:
\[\boxed{ \text{Perception} \rightarrow \text{Multimodal Fusion} \rightarrow \text{Policy Representation} \rightarrow \text{Action Decoder} }\] -
Given images
\[I_t^{1:N}\]-
language instruction
\[l\]-
and robot state
\[s_t\]-
a vision encoder first produces visual tokens:
\[z_v = f_{\mathrm{vision}} ( I_t^{1:N} )\]
-
-
-
-
Language is tokenized and embedded:
\[z_l = f_{\mathrm{text}}(l)\] -
Proprioception can be encoded separately:
\[z_s = f_{\mathrm{state}}(s_t)\] -
A multimodal transformer then produces a shared representation:
\[h = f_{\mathrm{VLM}} ( z_v, z_l, z_s )\] -
Finally, an action decoder maps \(h\) into actions:
\[a_{t:t+H} = f_{\mathrm{action}}(h)\] -
The largest architectural differences between VLA families arise in the last step:
\[f_{\mathrm{action}}\] -
Actions can be represented as discrete tokens, direct continuous regression targets, diffusion samples, or flow-matching trajectories.
Action Tokenization
- RT-2 demonstrated that robot actions could be represented as ordinary language-model tokens.
-
Suppose the robot action contains seven dimensions:
\[a_t = [ x_t, y_t, z_t, r_t, p_t, y_t^{\mathrm{rot}}, g_t ]\] -
Each continuous dimension can be discretized into \(B\) bins:
\[q(a^{(j)}) \in \{0,\ldots,B-1\}\] -
The resulting action becomes a sequence of discrete symbols:
\[[ q(x_t), q(y_t), q(z_t), q(r_t), q(p_t), q(y_t^{\mathrm{rot}}), q(g_t) ]\] - These values can be mapped onto tokens in the language model’s vocabulary.
-
The policy objective becomes standard autoregressive next-token prediction:
\[\mathcal{L}_{\mathrm{action}} = - \sum_{j=1}^{d_a} \log p_\theta( q(a_t^{(j)}) \mid I,l,a_t^{(<j)} )\] -
This design has an important engineering advantage:
\[\boxed{ \text{robot control becomes another sequence-generation problem}. }\] - The same transformer architecture and training infrastructure used for language can therefore be reused for action prediction.
RT-2
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control by Brohan et al. (2023) was built by adapting large pretrained VLMs into robot policies through co-fine-tuning.
-
The training mixture contains two broad distributions:
\[\mathcal{D} = \mathcal{D}_{\mathrm{web}} \cup \mathcal{D}_{\mathrm{robot}}\] -
Web examples retain conventional VLM targets:
\[(I,l)\rightarrow y_{\mathrm{text}}\]-
while robot examples use
\[(I,l)\rightarrow y_{\mathrm{action}}\]
-
- Because actions are represented as text-compatible tokens, both examples can be trained through essentially the same autoregressive objective.
-
A simplified joint objective is
\[\mathcal{L} = \lambda_{\mathrm{web}} \mathcal{L}_{\mathrm{VLM}} + \lambda_{\mathrm{robot}} \mathcal{L}_{\mathrm{action}}\] - The web component is important because fine-tuning exclusively on robot data risks catastrophic forgetting of semantic capabilities acquired during pretraining.
- RT-2 showed that internet-derived concepts could influence physical behavior, including instructions requiring recognition of previously unseen semantic categories and elementary reasoning about which object should be manipulated.
-
This established a foundational VLA principle:
\[\boxed{ \text{semantic knowledge learned for perception and language can transfer into action}. }\]
Semantic Reasoning Versus Motor Control
- The success of RT-2 also reveals a distinction between two types of intelligence required by a robot.
-
Semantic reasoning answers questions such as
\[\text{Which object satisfies the instruction?}\] -
Motor control answers
\[\text{How should the robot move to manipulate it?}\] -
These correspond roughly to
\[p(z_{\mathrm{semantic}}\mid I,l)\]-
and
\[p(a\mid z_{\mathrm{semantic}},s)\]
-
- A monolithic autoregressive VLA asks the same network to solve both.
-
Later architectures increasingly separate them:
\[\boxed{ \text{VLM semantic backbone} + \text{specialized action expert}. }\] - This becomes particularly important when actions are continuous, high-frequency, and temporally correlated.
OpenVLA
- OpenVLA: An Open-Source Vision-Language-Action Model by Kim et al. (2024) provides an open 7B-parameter VLA trained on approximately 970,000 real-world robot trajectories drawn from Open X-Embodiment.
-
OpenVLA builds on a pretrained Prismatic VLM and consists of three major components:
\[\boxed{ \text{DINOv2 + SigLIP} \rightarrow \text{Projector} \rightarrow \text{Llama 2 7B} }\] - The OpenVLA project page describes the visual encoder as a fusion of DINOv2 and SigLIP features, followed by a projector that maps visual embeddings into the language-model embedding space and a Llama 2 7B backbone that predicts tokenized actions.
- The fused vision representation is important because the two pretrained encoders emphasize complementary properties.
- DINOv2: Learning Robust Visual Features without Supervision by Oquab et al. (2023) learns general-purpose self-supervised visual representations with strong spatial and visual structure.
- SigLIP contributes image-language-aligned semantic representations.
-
Abstractly,
\[z_{\mathrm{DINO}} = f_{\mathrm{DINO}}(I)\] \[z_{\mathrm{SigLIP}} = f_{\mathrm{SigLIP}}(I)\]-
and
\[z_v = f_{\mathrm{fuse}} ( z_{\mathrm{DINO}}, z_{\mathrm{SigLIP}} )\]
-
- The combination attempts to preserve both spatially useful visual information and language-aligned semantics.
OpenVLA Action Decoding
- OpenVLA follows the tokenized-action paradigm.
-
Continuous robot actions are discretized and mapped to tokens:
\[a_t \rightarrow q(a_t) \rightarrow \text{action tokens}\] -
The language-model backbone then predicts
\[p_\theta( q(a_t) \mid I_t,l )\] -
After generation, tokens are converted back to continuous actions:
\[\text{action tokens} \rightarrow q^{-1} \rightarrow \hat a_t\] - This design makes VLA training closely resemble language-model supervised fine-tuning.
Fine-Tuning VLAs to New Robots
- A pretrained VLA is not automatically optimal for every embodiment.
-
Suppose the pretrained action space is
\[\mathcal{A}_{\mathrm{pretrain}}\]-
while the target robot exposes
\[\mathcal{A}_{\mathrm{target}}\]
-
- The two may differ in coordinate conventions, action dimensionality, gripper representation, camera placement, control frequency, or proprioceptive inputs.
-
Adaptation therefore requires learning
\[\pi_{\theta'} ( a^{\mathrm{target}} \mid o^{\mathrm{target}},l )\] - OpenVLA demonstrates both full fine-tuning and parameter-efficient adaptation. Its project results report that LoRA can match full fine-tuning in tested settings while modifying only a small fraction of model parameters.
-
This suggests a practical deployment workflow:
\[\boxed{ \text{generalist VLA checkpoint} \rightarrow \text{small target-robot dataset} \rightarrow \text{fine-tuning} \rightarrow \text{specialized deployment policy}. }\]
Limitations of Autoregressive Action Tokens
- Action tokenization is conceptually elegant, but it creates several limitations.
-
First, discretization introduces quantization error:
\[a \neq q^{-1}(q(a))\] -
Second, autoregressive decoding introduces latency:
\[p(a) = \prod_{j=1}^{d_a} p( a^{(j)} \mid a^{(<j)},h )\] - Third, tokenized categorical distributions are not naturally matched to continuous geometric action spaces.
- These limitations motivated continuous-action VLA architectures.
OpenVLA-OFT
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success by Kim et al. (2025) systematically studies how OpenVLA should be adapted for downstream robot control and introduces OpenVLA-OFT.
-
The optimized recipe combines
\[\text{parallel decoding} + \text{action chunking} + \text{continuous actions} + L_1\text{ regression}\] - Instead of generating each action dimension autoregressively, the policy predicts multiple continuous actions simultaneously.
-
The loss can be written as
\[\mathcal{L}_{\mathrm{OFT}} = \frac{1}{H} \sum_{k=0}^{H-1} \| a_{t+k} - \hat a_{t+k} \|_1\] - The important architectural lesson is that the semantic benefits of a pretrained VLA do not require retaining its original token-based action decoder.
Octo
- Octo: An Open-Source Generalist Robot Policy by the Octo Model Team et al. (2024) takes a different approach to generalist robot learning.
- Octo is a transformer-based diffusion policy pretrained on approximately 800,000 trajectories from Open X-Embodiment.
- Unlike OpenVLA, Octo is designed primarily as a flexible robot policy rather than as a large internet-pretrained VLM converted into a policy.
- Its inputs can include images, proprioception, language, and goal images.
- The Octo project page describes two released model sizes, Octo-Small at 27M parameters and Octo-Base at 93M parameters, with a design explicitly intended to support new sensors, task specifications, and action spaces during fine-tuning.
Octo’s Readout Architecture
- Octo uses learned readout tokens to extract policy information from the transformer.
- Let \(T_O\) represent observation tokens, \(T_G\) task tokens, and \(T_R\) learned readout tokens.
-
The transformer computes
\[H = \operatorname{Transformer} ( T_O,T_G,T_R )\] -
The readout representation \(h_R\) is passed to a lightweight action head:
\[a \sim f_{\mathrm{action}}(h_R)\] - In Octo’s default configuration, the action head implements a diffusion objective.
-
This creates the separation
\[\text{shared transformer} \rightarrow \text{readout representation} \rightarrow \text{replaceable output head}\]
Flexible Observation Spaces
- Robot platforms differ not only in actions but also in sensors.
-
A modular token interface can represent
\[\mathcal{T} = \{ T_{\mathrm{vision}}, T_{\mathrm{state}}, T_{\mathrm{force}}, T_{\mathrm{language}}, T_{\mathrm{goal}}, \ldots \}\] - The transformer can operate on whichever token blocks are available, making modularity increasingly important as Physical AI expands from standardized robot arms toward humanoids and heterogeneous autonomous machines.
Two Paths to Generalist Robot Policies
-
OpenVLA and Octo illustrate two complementary approaches:
\[\boxed{ \text{internet-pretrained semantic model} \rightarrow \text{robot action model} }\]-
and
\[\boxed{ \text{large heterogeneous robot dataset} \rightarrow \text{flexible robot foundation policy}. }\]
-
-
Modern systems increasingly combine both:
\[\boxed{ \text{internet-scale multimodal knowledge} + \text{cross-embodiment robot data} + \text{continuous generative action modeling}. }\]
$\pi_0$
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control by Black et al. (2024) introduces a VLA architecture built on a pretrained VLM combined with a flow-matching action expert.
-
The architecture can be conceptualized as
\[\boxed{ \text{VLM} + \text{Action Expert} }\]- rather than treating actions as ordinary language tokens.
-
The VLM processes visual observations and language:
\[h_{\mathrm{semantic}} = f_{\mathrm{VLM}}(I,l)\]-
while the action expert predicts a continuous vector field:
\[v_\theta( A^\tau, \tau \mid h_{\mathrm{semantic}}, s )\]
-
-
Starting from noise,
\[A^0\sim\mathcal{N}(0,I)\]-
the action trajectory is transported toward the learned robot-action distribution through
\[\frac{dA^\tau}{d\tau} = v_\theta( A^\tau, \tau \mid h_{\mathrm{semantic}}, s )\]
-
The Action Expert
-
Instead of requiring the VLM to directly model actions using its text vocabulary, the VLM produces representations consumed by a dedicated motor model:
\[h=f_{\mathrm{VLM}}(I,l)\] \[A=f_{\mathrm{expert}}(h,s,\epsilon)\] - The action expert can operate at a different dimensionality, precision, architecture, and inference cadence from the semantic backbone.
- This makes it possible to use a large semantic model for understanding and a specialized generative model for motor control.
Flow Matching for Continuous Control
-
For expert action chunk \(A_1\) and noise \(A_0\), define
\[A_\tau=(1-\tau)A_0+\tau A_1\] -
The desired vector field is
\[u_\tau=A_1-A_0\] -
The action expert learns
\[v_\theta(A_\tau,\tau,c)\]-
using
\[\mathcal{L}_{\mathrm{flow}} = \mathbb{E} \left[ \| v_\theta(A_\tau,\tau,c)-u_\tau \|_2^2 \right]\]-
where
\[c=\{I,l,s\}\]
-
-
-
At inference time, numerical integration generates a coherent continuous action chunk.
Cross-Embodiment Training in $\pi_0$
-
A generalist VLA must train across heterogeneous robots:
\[\mathcal{D} = \mathcal{D}_{e_1} \cup \mathcal{D}_{e_2} \cup \cdots \cup \mathcal{D}_{e_N}\] - $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control by Black et al. (2024) trains across diverse dexterous platforms including single-arm, dual-arm, and mobile-manipulator systems.
- Training requires normalized or embodiment-aware action representations so shared parameters can learn reusable structure without confusing incompatible control semantics.
$\pi_{0.5}$ and Open-World Generalization
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization by Black et al. (2025) extends the VLA paradigm toward long-horizon operation in previously unseen environments.
-
The training distribution combines heterogeneous sources:
\[\mathcal{D} = \mathcal{D}_{\mathrm{robot}} + \mathcal{D}_{\mathrm{web}} + \mathcal{D}_{\mathrm{semantic}} + \mathcal{D}_{\mathrm{perception}}\] - The model can consequently learn both low-level action mappings and higher-level semantic predictions, moving from pure motor generalization toward open-world physical reasoning.
Hierarchical Prediction in VLAs
-
Long-horizon tasks benefit from separating semantic subgoals from low-level actions:
\[l \xrightarrow{\text{semantic policy}} g_k \xrightarrow{\text{action policy}} A_k\] -
This reduces the effective reasoning horizon and allows high-level planning to operate at a slower cadence than motor control.
Tokenized Versus Continuous Actions
- Modern VLA design can be organized around a central choice.
Tokenized actions
-
Examples include RT-2 and the original OpenVLA:
\[p(q(a)\mid I,l)\] -
Advantages include architectural simplicity and reuse of language-model training infrastructure. Limitations include quantization, sequential decoding latency, and a mismatch between categorical tokens and continuous control geometry.
Continuous regression
-
OpenVLA-OFT demonstrates direct regression:
\[\hat A=f_\theta(I,l,s)\]-
with an objective such as
\[\mathcal{L}=\|A-\hat A\|_1\]
-
-
This provides fast parallel decoding but can struggle with strongly multimodal action distributions.
Diffusion
-
Octo and Diffusion Policy-style architectures model
\[p(A\mid I,l,s)\]- through iterative denoising.
Flow matching
\(\pi_0\) uses
\[\frac{dA}{d\tau} = v_\theta(A,\tau,c)\]- The appropriate choice depends on control frequency, task multimodality, model size, inference hardware, embodiment, and required precision.
The VLA Training Mixture
-
A foundation VLA may draw from several data sources:
\[\mathcal{D}_{\mathrm{total}} = \mathcal{D}_{\mathrm{VLM}} \cup \mathcal{D}_{\mathrm{robot}} \cup \mathcal{D}_{\mathrm{human}} \cup \mathcal{D}_{\mathrm{synthetic}}\] - Vision-language data provides semantic knowledge, robot trajectories provide action grounding, human video provides physical interaction priors, and simulation and synthetic data provide scale and coverage.
-
The mixture weights become an important hyperparameter:
\[\mathcal{L} = \lambda_{\mathrm{VLM}} \mathcal{L}_{\mathrm{VLM}} + \lambda_{\mathrm{robot}} \mathcal{L}_{\mathrm{robot}} + \lambda_{\mathrm{aux}} \mathcal{L}_{\mathrm{aux}}\] - VLA training is therefore partly a problem of balancing semantic retention against physical specialization.
Normalizing Heterogeneous Robot Data
- Cross-embodiment datasets rarely share action conventions.
-
Actions can be normalized as
\[\tilde a^{(e)} = N_e(a^{(e)})\]- where \(N_e\) is an embodiment- or dataset-specific normalization function.
-
The policy learns
\[\tilde a=\pi_\theta(o,l,e)\]-
and deployment applies
\[a^{(e)}=N_e^{-1}(\tilde a)\]
-
- Data standardization is therefore part of the model architecture in practice, not merely preprocessing.
Control Frequency and Inference Latency
- A VLA deployed on a robot is a real-time system.
-
For target control frequency \(f_c\), the inference budget is approximately
\[T_{\max}=\frac{1}{f_c}\] -
The latency budget includes
\[T_{\mathrm{camera}} + T_{\mathrm{preprocess}} + T_{\mathrm{vision}} + T_{\mathrm{policy}} + T_{\mathrm{decode}} + T_{\mathrm{communication}}\] -
If
\[T_{\mathrm{total}}>T_{\max}\]- the policy cannot maintain the desired control rate.
- Autoregressive generation, diffusion sampling, numerical flow integration, image resolution, context length, and model size therefore directly affect physical control performance.
Asynchronous Inference
- One solution is to overlap policy inference and robot execution.
- While the robot executes \(A_t\), the model can begin computing \(A_{t+K}\).
- This reduces idle time but introduces stale-observation risk, producing a tradeoff between latency and observation freshness.
Fine-Tuning Strategy
- Adapting a VLA can involve full fine-tuning, frozen-backbone action-head training, or parameter-efficient fine-tuning.
-
LoRA introduces low-rank updates
\[W'=W+BA\]-
where
\[\operatorname{rank}(BA)\ll\operatorname{dim}(W)\]
-
- The OpenVLA project reports strong LoRA-based adaptation results.
-
The appropriate strategy depends on the size of the domain shift:
\[\boxed{ \text{fine-tuning scope} \propto \text{magnitude of distribution shift}. }\]
What Makes a VLA Generalist?
-
Generalization can occur along several dimensions:
\[G_{\mathrm{VLA}} = G_{\mathrm{semantic}} \times G_{\mathrm{visual}} \times G_{\mathrm{task}} \times G_{\mathrm{motion}} \times G_{\mathrm{environment}} \times G_{\mathrm{embodiment}}\] -
A model can perform strongly on one dimension while failing on another. Evaluating VLA generality therefore requires substantially richer benchmarks than measuring success on tasks drawn from the training distribution.
The Emerging VLA Design Pattern
-
Across RT-2, OpenVLA, Octo, OpenVLA-OFT, \(\pi_0\), and \(\pi_{0.5}\), a common architecture is emerging:
\[\boxed{ \begin{array}{c} \text{Images + Language + Robot State}\\ \downarrow\\ \text{Pretrained Visual / Multimodal Representations}\\ \downarrow\\ \text{Shared Transformer Backbone}\\ \downarrow\\ \text{Semantic / Physical Representation}\\ \downarrow\\ \text{Specialized Action Decoder}\\ \downarrow\\ \text{Action Chunk}\\ \downarrow\\ \text{Robot Controller} \end{array} }\] - The major trend is away from treating robot actions as merely another type of language token and toward preserving a powerful pretrained multimodal backbone while introducing action-specific architectures.
-
This can be summarized as
\[\boxed{ \text{foundation semantics} + \text{robot experience} + \text{continuous generative control}. }\] - The semantic model answers what should happen. The action model determines how the embodiment should make it happen.
- This architecture forms the basis for increasingly capable humanoid foundation models. The next section will examine NVIDIA Isaac GR00T and the broader humanoid Physical AI stack, including GR00T N1, later GR00T models, dual-system architectures, whole-body and bimanual control, embodiment adaptation, synthetic robot data, GR00T-Dreams, Isaac Sim, Isaac Lab, Newton, and sim-to-real training.
Isaac GR00T and Humanoid Foundation Models
Why Humanoids Are a Distinct Physical AI Problem
- Humanoid robots occupy an unusual position in Physical AI. Their morphology is designed around environments built for humans, which potentially allows them to use doors, tools, shelves, workstations, vehicles, and other infrastructure without redesigning the environment.
-
At the same time, humanoid control combines several difficult problems:
\[\text{Humanoid Intelligence} = \text{Perception} + \text{Reasoning} + \text{Manipulation} + \text{Locomotion} + \text{Whole-Body Control}\] - A mobile manipulator may control one arm and a mobile base. A humanoid may need to coordinate two arms, dexterous hands, torso, head, and legs while simultaneously maintaining balance.
-
The action space can therefore contain dozens of coupled degrees of freedom:
\[a_t = [ a_t^{\mathrm{left\ arm}}, a_t^{\mathrm{right\ arm}}, a_t^{\mathrm{torso}}, a_t^{\mathrm{hands}}, a_t^{\mathrm{legs}} ]\] -
The challenge becomes even greater because many useful tasks are long-horizon:
\[\text{navigate} \rightarrow \text{locate object} \rightarrow \text{reach} \rightarrow \text{grasp} \rightarrow \text{manipulate} \rightarrow \text{recover}\] - A general-purpose humanoid therefore needs semantic reasoning at long timescales and precise motor control at short timescales.
- NVIDIA’s Isaac GR00T platform is designed around this broader problem, combining robot foundation models, data pipelines, simulation, middleware, accelerated runtime libraries, and deployment infrastructure rather than treating the VLA as an isolated model.
From Project GR00T to GR00T N1
- NVIDIA introduced Project GR00T in 2024 as an initiative to develop general-purpose foundation models for humanoid robots.
- The first major open model from this effort was GR00T N1: An Open Foundation Model for Generalist Humanoid Robots by NVIDIA et al. (2025), a Vision-Language-Action model designed around a dual-system architecture and trained from real robot trajectories, human video, simulation, and synthetic data.
-
The high-level architecture can be represented as
\[\boxed{ \text{Vision + Language + Robot State} \rightarrow \text{System 2} \rightarrow \text{System 1} \rightarrow \text{Continuous Actions} }\]- where the two systems operate at different levels of abstraction.
-
This separation addresses a fundamental mismatch in embodied intelligence:
\[T_{\mathrm{reasoning}} \gg T_{\mathrm{control}}\] - A robot may need substantial computation to interpret a novel instruction but must continue producing smooth physical actions at a much faster cadence.
The Dual-System Architecture
- GR00T N1 explicitly adopts a dual-system design inspired by the distinction between deliberate reasoning and fast intuitive behavior.
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots by NVIDIA et al. (2025) implements System 2 as a vision-language module that interprets visual observations and language instructions, while System 1 is a Diffusion Transformer that converts the resulting representation and robot state into continuous motor actions.
-
Conceptually,
\[h_t = f_{\mathrm{VLM}} ( I_t,l )\]- produces a semantic representation.
-
The motor model then predicts
\[A_t = f_{\mathrm{DiT}} ( h_t,s_t,\epsilon )\]-
where
\[A_t = [ a_t,\ldots,a_{t+H} ]\]- is an action chunk.
-
-
The architecture can therefore be viewed as
\[\boxed{ \underbrace{ f_{\mathrm{VLM}}(I,l) }_{\text{understand}} \rightarrow \underbrace{ f_{\mathrm{DiT}}(h,s) }_{\text{act}} }\] - This resembles the action-expert architecture discussed for \(\pi_0\), but GR00T is explicitly developed as part of a broader humanoid robotics stack.
System 2: Vision-Language Understanding
-
System 2 answers questions at the semantic level:
\[\text{What does the instruction mean?}\] \[\text{Which objects are relevant?}\] \[\text{What state is the environment in?}\] \[\text{What behavior should occur?}\] - For an instruction such as
-
Place the blue cup inside the drawer,
-
the VLM must connect language tokens to visual entities:
\[l \rightarrow \{ \text{blue cup}, \text{drawer}, \text{place inside} \}\]
-
- The resulting representation needs to encode both semantic and spatial information because System 1 must ultimately ground it into motor behavior.
- This distinction is important. A VLM that knows what a drawer is but cannot distinguish its actionable geometry provides insufficient information for physical control.
-
GR00T therefore illustrates the broader evolution from
\[\text{semantic VLM}\]-
toward
\[\text{physically grounded VLM}\]
-
System 1: The Diffusion Transformer
- System 1 generates continuous robot behavior.
- Rather than autoregressively generating discrete action tokens, GR00T N1 uses a Diffusion Transformer, or DiT, for continuous action generation.
-
A simplified diffusion process starts with an expert trajectory
\[A^0\]-
and constructs noisy versions
\[A^k = \sqrt{\bar{\alpha}_k}A^0 + \sqrt{1-\bar{\alpha}_k}\epsilon\]-
where
\[\epsilon\sim\mathcal{N}(0,I)\]
-
-
-
The DiT receives the noisy action sequence together with conditioning information:
\[\epsilon_\theta = f_{\mathrm{DiT}} ( A^k, k, h_{\mathrm{VLM}}, s_t )\] -
Training minimizes a denoising objective such as
\[\mathcal{L}_{\mathrm{diff}} = \mathbb{E} \left[ \| \epsilon - \epsilon_\theta \|_2^2 \right]\] -
At inference time,
\[A^K \sim \mathcal{N}(0,I)\]-
is progressively transformed into
\[A^0\]- a coherent continuous motor trajectory.
-
- This architecture allows the policy to model multiple physically valid behaviors rather than collapsing them into their mean.
Why Separate Reasoning and Action?
-
A single model could theoretically predict motor actions directly from pixels and language:
\[(I,l,s) \xrightarrow{f_\theta} A\] - However, the semantic and motor problems have different statistical structures.
-
Semantic representations benefit from internet-scale multimodal data:
\[\mathcal{D}_{\mathrm{web}} \gg \mathcal{D}_{\mathrm{robot}}\] -
Motor control requires specialized physical trajectories:
\[\mathcal{D}_{\mathrm{action}} = \{ I_t,s_t,a_t \}\] - The two-system architecture permits broad semantic knowledge to coexist with specialized continuous control.
GR00T’s Data Pyramid
- A major theme in GR00T is that robot intelligence cannot scale through manually collected robot demonstrations alone.
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots by NVIDIA et al. (2025) trains on a heterogeneous mixture that includes internet-scale data, human videos, real robot trajectories, simulated trajectories, and synthetic data.
-
The data hierarchy can be conceptualized as
\[\boxed{ \begin{array}{c} \text{Web + Human Video}\\ \downarrow\\ \text{Synthetic / Simulated Robot Data}\\ \downarrow\\ \text{Real Robot Demonstrations} \end{array} }\] - The levels differ in cost and physical fidelity.
-
Web and human-video data are abundant:
\[N_{\mathrm{web}} \gg N_{\mathrm{robot}}\] - But they lack direct robot action labels. Simulation provides action labels cheaply but introduces a reality gap. Real robot demonstrations provide the highest embodiment fidelity but are expensive to collect.
-
The training problem is therefore one of combining
\[\boxed{ \text{scale} + \text{physical grounding} + \text{embodiment fidelity}. }\]
Learning from Human Video
- Human video is attractive because humans continuously demonstrate useful physical behavior.
-
However, human video usually contains observations
\[I_{1:T}\]-
without corresponding robot actions
\[a_{1:T}\]
-
-
Thus,
\[\mathcal{D}_{\mathrm{video}} = \{ I_{1:T} \}\]-
cannot directly supervise
\[\pi_\theta(a_t\mid I_t)\]
-
-
One abstraction is
\[I_t,I_{t+1} \xrightarrow{\mathrm{latent\ action}} z_t\]- where \(z_t\) captures the transformation responsible for the observed transition.
- The robot policy can then learn physical regularities from much larger quantities of unlabeled interaction video.
Real Robot Demonstrations
- Real robot demonstrations remain the highest-fidelity source of action supervision.
-
A trajectory contains
\[\tau = \{ I_t, s_t, l, a_t \}_{t=1}^{T}\] - Teleoperation provides a practical mechanism for collecting these trajectories.
-
The human operator supplies actions
\[a_t^{\mathrm{human}}\]- while the robot records synchronized observations and state.
-
The resulting dataset directly supervises
\[\pi_\theta( a_t \mid I_{\leq t},s_{\leq t},l )\] - However, collecting thousands of hours of diverse humanoid teleoperation is expensive.
-
This produces the central data bottleneck:
\[\boxed{ \text{robot model scale} > \text{available real robot data}. }\]
Synthetic Motion Generation
- Simulation provides a mechanism for transforming a small amount of human demonstration data into much larger robot datasets.
- NVIDIA’s Isaac GR00T Blueprint uses simulation and synthetic-data workflows to expand motion data for humanoid imitation learning.
-
A simplified pipeline is
\[\text{Human Demonstration} \rightarrow \text{Motion Retargeting} \rightarrow \text{Simulation} \rightarrow \text{Domain Randomization} \rightarrow \text{Robot Trajectories}\] -
If one motion trajectory is \(\tau\), simulation can generate variations
\[\{ \tau_1, \tau_2, \ldots,\tau_N \}\]- by modifying object poses, lighting, camera viewpoints, robot configurations, physical parameters, and environment layouts.
GR00T-Mimic
- GR00T-Mimic represents the simulation-centered branch of this data engine.
-
Suppose a human provides
\[\tau_{\mathrm{demo}}\] -
Motion retargeting maps the demonstration onto the target embodiment:
\[\tau_{\mathrm{human}} \xrightarrow{R_e} \tau_{\mathrm{robot}}\] -
Simulation then perturbs the environment:
\[E \rightarrow \{ E_1,E_2,\ldots,E_N \}\] -
The retargeted behavior can be replayed or adapted across these variations, producing
\[\mathcal{D}_{\mathrm{Mimic}} = \{ \tau_{i,j} \}\] - The advantage is physical consistency. The limitation is that the diversity remains related to the seed demonstrations.
GR00T-Dreams
- GR00T-Dreams uses world foundation models to generate new robot videos from an initial image and language instruction, then recovers actions from those generated videos to create training trajectories.
-
The pipeline can be summarized as
\[\boxed{ \text{Image + Instruction} \rightarrow \text{World Model} \rightarrow \text{Synthetic Robot Video} \rightarrow \text{Action Extraction} \rightarrow \text{VLA Training Data}. }\] -
Suppose the seed observation is \(I_0\) and the instruction is \(l\). A video world model generates
\[\hat I_{1:T} \sim p_\phi( I_{1:T} \mid I_0,l )\] -
An inverse dynamics model then estimates actions:
\[\hat a_t = f_{\mathrm{IDM}} ( \hat I_t,\hat I_{t+1} )\] -
This produces a synthetic trajectory
\[\hat\tau = \{ \hat I_t,\hat a_t \}_{t=1}^{T}\]
DreamGen
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models by Jang et al. (2025) develops this idea by adapting video world models to robot domains, generating synthetic task videos, extracting corresponding actions, and using the resulting trajectories to improve robot policies.
-
The workflow is
\[\text{Robot Data} \rightarrow \text{Fine-Tune World Model} \rightarrow \text{Generate Novel Videos} \rightarrow \text{Infer Actions} \rightarrow \text{Post-Train Policy}\] -
This creates a training loop in which one generative model supplies experience to another:
\[\boxed{ \text{World Model} \rightarrow \text{Synthetic Experience} \rightarrow \text{VLA}. }\]
GR00T N1.5
- GR00T N1.5 was introduced in June 2025 as an improved GR00T N1 model with architectural, data, and modeling changes targeting stronger generalization and language following.
- Like N1, N1.5 uses a VLM to encode language and visual observations and a DiT to process robot state and noisy actions.
- A particularly important architectural change is that the VLM is frozen during both pretraining and fine-tuning.
-
Thus,
\[\theta_{\mathrm{VLM}} = \text{constant}\]-
while
\[\theta_{\mathrm{DiT}}\]- and associated adaptation components learn robot behavior.
-
- The motivation is intuitive: broad visual-language knowledge is expensive to acquire and easy to damage through narrow robot fine-tuning.
GR00T N1.5 and Synthetic Post-Training
- N1.5 also demonstrates how world-model-generated data can participate directly in robot post-training.
-
The broader training loop is
\[\pi_k \rightarrow \text{identify missing capability} \rightarrow \text{generate synthetic experience} \rightarrow \text{post-train} \rightarrow \pi_{k+1}\] - This suggests a scalable alternative to repeatedly returning to physical teleoperation whenever a new task or environment is introduced.
GR00T N1.6
- GR00T N1.6 was released in December 2025 as another architectural and data update to the GR00T family.
- N1.6 replaces the earlier Eagle-based vision-language component with an internal Cosmos-2B VLM variant trained on both general vision-language tasks and embodied reasoning tasks such as next-action prediction. It also increases the depth of the action DiT relative to N1.5.
-
The change illustrates an important trend:
\[\text{general VLM} + \text{embodied reasoning} \rightarrow \text{robot VLM}\] - Thus, physical intelligence increasingly enters not only the action head but also the semantic backbone.
GR00T N1.7 and the Evolution Toward Deployment
- The GR00T family has continued beyond N1.6. NVIDIA’s 2026 GR00T development materials describe GR00T N1.7 alongside ONNX and TensorRT export support for deployment.
-
This progression is significant because robot foundation models must eventually satisfy constraints beyond benchmark success:
\[\text{Model Quality} + \text{Inference Latency} + \text{Runtime Portability} + \text{Hardware Integration}\] - A research checkpoint that cannot satisfy on-robot latency or deployment requirements is not yet a complete Physical AI system.
Bimanual Manipulation
- Humanoids frequently need both arms.
-
The policy therefore predicts
\[A_t = [ A_t^{L}, A_t^{R} ]\]- where \(A_t^{L}\) and \(A_t^{R}\) represent left- and right-arm trajectories.
-
The difficulty is that the two trajectories are not independent:
\[p( A^L,A^R ) \neq p(A^L)p(A^R)\] - A joint generative action model can represent these dependencies directly by predicting both arms as one structured trajectory.
From Manipulation to Whole-Body Control
-
Whole-body humanoid behavior may require coordinating
\[q_t = [ q_t^{\mathrm{legs}}, q_t^{\mathrm{torso}}, q_t^{\mathrm{left\ arm}}, q_t^{\mathrm{right\ arm}} ]\] - A manipulation objective may conflict with balance.
-
A whole-body controller must therefore satisfy constraints such as
\[\operatorname{CoM}(q_t) \in \mathcal{S}_{\mathrm{support}}\]- while simultaneously optimizing the manipulation trajectory.
- This coupling is one reason humanoids require accurate physics simulation and specialized low-level controllers even when high-level behavior is learned.
Isaac Sim
- Simulation is a central component of the GR00T ecosystem.
-
A simulated rollout produces
\[\tau_{\mathrm{sim}} = \{ s_t, o_t, a_t, r_t \}_{t=1}^{T}\] - Unlike real-world collection, simulation provides access to privileged state such as object poses, contacts, velocities, and forces.
-
The challenge remains
\[P_{\mathrm{sim}} \neq P_{\mathrm{real}}\] - The sim-to-real pipeline is therefore concerned with making policies insensitive to this difference.
Isaac Lab
- Isaac Lab provides a robot-learning framework on top of the simulation stack, supporting workflows such as reinforcement learning, imitation learning, motion planning, and large-scale parallel simulation.
-
GPU simulation can execute
\[E_1,E_2,\ldots,E_N\]-
in parallel. If each environment generates experience at rate \(r\),
\[R_{\mathrm{total}} \approx Nr\]
-
- This parallelism can dramatically increase policy experience generation.
Newton Physics Engine
- Newton is an open-source GPU-accelerated physics engine developed by NVIDIA, Google DeepMind, and Disney Research and integrated with Isaac Lab.
-
At a high level, the simulator approximates
\[s_{t+1} = F_{\mathrm{physics}} ( s_t, a_t, \phi )\]- where \(\phi\) contains physical parameters such as mass, friction, inertia, stiffness, and damping.
Domain Randomization
- Perfectly modeling the real world is generally impossible.
-
Instead of training under one simulator configuration,
\[\phi=\phi_0\]-
sample
\[\phi \sim p(\phi)\]
-
- Parameters can include mass, friction, lighting, texture, camera pose, latency, and motor strength.
-
The policy optimizes
\[J(\pi) = \mathbb{E}_{\phi\sim p(\phi)} [ R(\pi;\phi) ]\] -
The goal is to make the real world another plausible sample from the training distribution:
\[\phi_{\mathrm{real}} \in \operatorname{support}(p(\phi))\]
Sim-and-Real Co-Training
-
Another strategy is to train jointly on simulated and physical data:
\[\mathcal{D} = \mathcal{D}_{\mathrm{real}} \cup \mathcal{D}_{\mathrm{sim}}\] -
The training objective becomes
\[\mathcal{L} = \lambda_{\mathrm{real}} \mathcal{L}_{\mathrm{real}} + \lambda_{\mathrm{sim}} \mathcal{L}_{\mathrm{sim}}\] -
Simulation supplies scale and diversity. Real data anchors the policy to true sensor statistics and dynamics.
Embodiment Adaptation
- A humanoid foundation model should ideally transfer knowledge across robots.
-
Suppose two embodiments have states
\[s_t^{(1)} \in \mathbb{R}^{d_1}\]-
and
\[s_t^{(2)} \in \mathbb{R}^{d_2}\]
-
-
A foundation model therefore needs an embodiment-aware interface:
\[z_s^{(e)} = f_{\mathrm{state}}^{(e)} ( s^{(e)} )\] \[a^{(e)} = f_{\mathrm{action}}^{(e)} ( z_a )\] - The shared model operates on latent representations while embodiment-specific adapters map between shared representations and physical hardware.
Post-Training a Humanoid Foundation Model
-
A practical pipeline is
\[\boxed{ \text{GR00T Base Model} \rightarrow \text{Target Embodiment Data} \rightarrow \text{Task Data} \rightarrow \text{Post-Training} \rightarrow \text{Closed-Loop Evaluation} }\] -
The target data can combine
\[\mathcal{D}_{\mathrm{target}} = \mathcal{D}_{\mathrm{teleop}} + \mathcal{D}_{\mathrm{sim}} + \mathcal{D}_{\mathrm{synthetic}}\] -
The resulting policy is
\[\pi_{\theta'} = \operatorname{PostTrain} ( \pi_\theta, \mathcal{D}_{\mathrm{target}} )\]
Closed-Loop Evaluation
- Robot foundation models cannot be evaluated adequately using only action prediction error.
-
The meaningful quantity is task-level rollout performance:
\[S(\tau) \in \{0,1\}\]-
or a graded score
\[S(\tau)\in[0,1]\]
-
-
Evaluation therefore requires repeatedly executing the policy:
\[\pi \rightarrow E \rightarrow \tau \rightarrow S(\tau)\] - A robust evaluation matrix should vary task, object, initial state, environment, instruction, and embodiment.
The Humanoid Data Flywheel
- The broader GR00T stack can be interpreted as a data flywheel.
-
Begin with a policy \(\pi_k\). Deploy it in simulation and reality, identify failure regions, generate additional data using teleoperation, simulation, or world models, post-train, and repeat:
\[\boxed{ \text{Train} \rightarrow \text{Simulate} \rightarrow \text{Deploy} \rightarrow \text{Find Failures} \rightarrow \text{Generate Data} \rightarrow \text{Train}. }\] - This feedback loop is arguably more important than any single model checkpoint.
The GR00T Physical AI Stack
-
GR00T is best understood not as one model but as a vertically integrated Physical AI development stack:
\[\boxed{ \begin{array}{c} \text{Human + Robot + Web Data}\\ \downarrow\\ \text{GR00T-Mimic / GR00T-Dreams}\\ \downarrow\\ \text{Isaac Sim + Isaac Lab + Newton}\\ \downarrow\\ \text{GR00T Foundation Model}\\ \downarrow\\ \text{Embodiment Post-Training}\\ \downarrow\\ \text{Accelerated Robot Runtime}\\ \downarrow\\ \text{Physical Humanoid}\\ \downarrow\\ \text{New Experience}\\ \circlearrowleft \end{array} }\] -
The important systems insight is that generalist humanoid intelligence depends on the entire loop:
\[\boxed{ \text{model} + \text{data engine} + \text{simulation} + \text{physics} + \text{runtime} + \text{robot}. }\]
From Humanoid Models to General Physical Intelligence
- GR00T illustrates several broader trends that extend beyond humanoid robots.
-
First, Physical AI is moving toward dual-system architectures:
\[\text{slow semantic reasoning} + \text{fast continuous control}\] -
Second, training data is becoming increasingly heterogeneous:
\[\text{web} + \text{human video} + \text{robot demonstrations} + \text{simulation} + \text{world-model generation}\] -
Third, synthetic experience is becoming part of the model-development loop:
\[\text{model weakness} \rightarrow \text{generate experience} \rightarrow \text{post-train}\] - Fourth, robot foundation models are increasingly paired with foundation-level simulation and world models.
-
Finally, the unit of progress is shifting from a single model toward an integrated autonomy stack:
\[\boxed{ \text{Physical AI} = \text{Foundation Model} + \text{World Model} + \text{Data Engine} + \text{Simulator} + \text{Control Stack}. }\] - The next section will move from humanoid manipulation to another major branch of Physical AI: autonomous driving. It will examine the autonomy stack, modular versus end-to-end driving, perception and prediction, learned planning, driving VLAs, NVIDIA Alpamayo, Chain-of-Causation reasoning, simulation, closed-loop reinforcement learning, and the distinct safety constraints that arise when Physical AI operates at road scale.
Autonomous Driving as Physical AI
Autonomous Driving as an Embodied Intelligence Problem
- Autonomous driving is one of the most demanding forms of Physical AI because perception, prediction, reasoning, planning, and control must operate continuously in an open world populated by other intelligent agents.
-
The vehicle receives observations
\[o_t=\{I_t^{1:N},L_t,R_t,G_t,v_t,\omega_t,n_t\}\]- where the inputs may include multi-camera images, lidar, radar, GPS, ego velocity, angular velocity, and navigation information.
-
The system must produce a future trajectory or control action
\[a_t=\{\delta_t,a_t^{\mathrm{long}}\}\]- where \(\delta_t\) denotes steering and \(a_t^{\mathrm{long}}\) denotes longitudinal acceleration or braking.
-
More commonly, a learned planner predicts a future ego trajectory
\[\tau_t=\{(x_{t+k},y_{t+k},\theta_{t+k})\}_{k=1}^{H}\]- which a downstream controller converts into actuator commands.
-
The complete closed loop is
\[\boxed{\text{Sense}\rightarrow\text{Understand}\rightarrow\text{Predict}\rightarrow\text{Plan}\rightarrow\text{Control}\rightarrow\text{Observe Consequences}.}\] - Unlike many manipulation tasks, errors can unfold at high velocity and involve pedestrians, cyclists, other vehicles, and infrastructure. Autonomous driving therefore combines the learning problems of Physical AI with unusually demanding requirements for safety, latency, robustness, and verification.
The Classical Autonomous Driving Stack
-
Traditional autonomous-driving systems divide the problem into modules:
\[\boxed{\text{Sensors}\rightarrow\text{Perception}\rightarrow\text{Prediction}\rightarrow\text{Planning}\rightarrow\text{Control}.}\] -
Perception estimates the current scene:
\[o_t\rightarrow\hat s_t\] -
Prediction estimates how other agents may evolve:
\[p(s_{t+1:t+H}^{\mathrm{agents}}\mid s_{\leq t})\] -
Planning chooses the ego trajectory:
\[\tau_{\mathrm{ego}}^*=\arg\min_\tau C(\tau,\hat s,\hat s_{\mathrm{agents}})\] -
The modular design provides explicit interfaces, but errors can propagate:
\[\epsilon_{\mathrm{perception}}\rightarrow\epsilon_{\mathrm{prediction}}\rightarrow\epsilon_{\mathrm{planning}}\]
From Modular Pipelines to End-to-End Driving
-
End-to-end autonomous driving attempts to optimize more of the stack jointly around driving behavior:
\[I_{\leq t}\rightarrow\tau_{t:t+H}\] - Planning-oriented Autonomous Driving by Hu et al. (2022) introduced UniAD, integrating detection, tracking, mapping, motion forecasting, occupancy prediction, and planning into a unified network organized around planning.
-
A joint objective may be written as
\[\mathcal{L}=\lambda_{\mathrm{det}}\mathcal{L}_{\mathrm{det}}+\lambda_{\mathrm{track}}\mathcal{L}_{\mathrm{track}}+\lambda_{\mathrm{map}}\mathcal{L}_{\mathrm{map}}+\lambda_{\mathrm{pred}}\mathcal{L}_{\mathrm{pred}}+\lambda_{\mathrm{plan}}\mathcal{L}_{\mathrm{plan}}\]
Scene Representations for Learned Driving
-
A dense representation might encode occupancy:
\[O(x,y,t)=P(\text{occupied}\mid x,y,t)\] -
A vectorized representation instead represents individual scene elements:
\[S=\{V_1,\ldots,V_N,L_1,\ldots,L_M\}\] -
VAD: Vectorized Scene Representation for Efficient Autonomous Driving by Jiang et al. (2023) represents agents and map elements as vectors and uses them as explicit planning constraints.
Driving as a Multimodal Prediction Problem
-
Driving behavior is inherently multimodal. A deterministic regression model minimizing
\[\mathcal{L}=\|\tau-\hat\tau\|_2^2\]-
may average multiple valid modes.
\[\boxed{\text{average of valid trajectories}\neq\text{valid trajectory}.}\]
-
-
A policy can instead represent
\[p_\theta(\tau\mid o_{\leq t},n)\]
Prediction Is Interactive
-
A conventional predictor estimates
\[p(\tau_{\mathrm{others}}\mid o)\]-
while an interactive predictor considers
\[p(\tau_{\mathrm{others}}\mid o,\tau_{\mathrm{ego}})\]
-
-
Planning therefore becomes a coupled multi-agent problem:
\[\tau_{\mathrm{ego}}\leftrightarrow\tau_{\mathrm{others}}\]
Open-Loop Versus Closed-Loop Evaluation
-
Open-loop evaluation compares a prediction with a recorded trajectory:
\[E_{\mathrm{open}}=d(\hat\tau,\tau^*)\] -
Closed-loop evaluation instead executes the policy:
\[a_t\sim\pi_\theta(o_t),\qquad s_{t+1}\sim P(s_{t+1}\mid s_t,a_t)\] -
Thus,
\[\boxed{\text{open loop evaluates predictions;}\qquad\text{closed loop evaluates behavior}.}\]
Why Long-Tail Scenarios Dominate Autonomous Driving
-
Routine driving is common while difficult edge cases are rare:
\[\mathcal{D}_{\mathrm{long-tail}}\ll\mathcal{D}_{\mathrm{routine}}\] -
The challenge is therefore
\[\boxed{\text{generalization under sparse supervision}.}\]
Vision-Language Models for Driving
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models by Tian et al. (2024) introduced a VLM-based driving system organized around scene description, scene analysis, and hierarchical planning, together with DriveVLM-Dual.
- A VLM may be effective at semantic reasoning but comparatively weak at precise spatial control, motivating hybrid architectures.
From Driving VLMs to Driving VLAs
-
A driving VLA receives
\[x_t=\{I_{\leq t}^{1:N},n_t,e_{\leq t}\}\]-
and can predict reasoning and action:
\[\pi_\theta(r_t,\tau_t\mid x_t)\]
-
-
A Survey on Vision-Language-Action Models for Autonomous Driving by Jiang et al. (2025) characterizes this emerging class as integrating visual perception, language understanding, reasoning, and vehicle control.
NVIDIA Alpamayo
-
NVIDIA Alpamayo is a family of reasoning VLA models, simulation frameworks, and datasets for autonomous-driving development.
\[\boxed{\text{Alpamayo}+\text{AlpaSim}+\text{AlpaGym}.}\] -
The NVIDIA Alpamayo platform describes Alpamayo models as consuming multi-camera video, navigation inputs, and driving context and producing future trajectories and Chain-of-Causation reasoning traces.
Alpamayo 1
- Alpamayo 1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail integrates explicit reasoning with trajectory prediction for difficult long-tail driving scenarios.
-
Conceptually,
\[h_t=f_{\mathrm{reason}}(I_{\leq t},n_t,e_{\leq t})\]-
followed by
\[\tau_t=f_{\mathrm{trajectory}}(h_t,\epsilon)\]
-
Chain of Causation Reasoning
-
Chain of Causation connects observations to causal factors that determine driving action:
\[\text{Observation}\rightarrow\text{Relevant Agent}\rightarrow\text{Predicted Interaction}\rightarrow\text{Driving Constraint}\rightarrow\text{Action}\] -
Thus,
\[r_t\leftrightarrow\tau_t\]
Why Reasoning-Action Alignment Matters
-
A model could produce plausible reasoning while independently producing an unrelated trajectory:
\[p(r,\tau\mid o)=p(r\mid o)p(\tau\mid o)\] -
A reasoning VLA instead seeks strong consistency with a structure closer to
\[o\rightarrow r\rightarrow\tau\]
Training Reasoning VLAs
-
A simplified pipeline is
\[\boxed{\text{Foundation Model}\rightarrow\text{Driving SFT}\rightarrow\text{Reasoning SFT}\rightarrow\text{RL}\rightarrow\text{Driving VLA}.}\] -
A supervised objective can be represented as
\[\mathcal{L}_{\mathrm{SFT}}=\lambda_r\mathcal{L}_{\mathrm{reason}}+\lambda_\tau\mathcal{L}_{\mathrm{trajectory}}\] -
Reinforcement learning can optimize
\[J(\pi)=\mathbb{E}_{\tau\sim\pi}[R(\tau,r)]\]
Trajectory Generation as an Action Expert
-
A driving VLA often predicts a future trajectory
\[\tau=[p_1,p_2,\ldots,p_H]\]-
where
\[p_k=(x_k,y_k,\theta_k)\]
-
-
A downstream controller computes
\[u_t=f_{\mathrm{control}}(\tau,s_t)\] -
Thus,
\[\boxed{\text{Reasoning}\rightarrow\text{Trajectory}\rightarrow\text{Control}\rightarrow\text{Actuation}.}\]
Diffusion-Based Trajectory Prediction
-
Given noisy trajectory \(\tau^k\), a diffusion decoder learns
\[\epsilon_\theta(\tau^k,k,h)\] -
A standard objective is
\[\mathcal{L}_{\mathrm{diff}}=\mathbb{E}\left[\|\epsilon-\epsilon_\theta(\tau^k,k,h)\|_2^2\right]\] -
At inference,
\[\tau^K\sim\mathcal{N}(0,I)\]- is transformed into a physically plausible trajectory.
Alpamayo 1.5
-
The Alpamayo family subsequently expanded to Alpamayo 1.5, continuing the pattern
\[\boxed{\text{Video + Navigation + Ego History}\rightarrow\text{Reasoning}\rightarrow\text{Trajectory}.}\]
Alpamayo 2 Super
- NVIDIA announced Alpamayo 2 Super in May 2026 as part of the continuing Alpamayo family for autonomous-driving and robotaxi development.
-
The broader scaling direction is
\[\text{specialized planner}\rightarrow\text{driving foundation model}\]
Foundation Models as Teachers
- A large driving foundation model need not be the final deployed model.
-
For teacher \(\pi_T\) and student \(\pi_S\),
\[(r_T,\tau_T)\sim\pi_T(o)\]-
and the student can minimize
\[\mathcal{L}_{\mathrm{distill}}=\lambda_rD(r_S,r_T)+\lambda_\tau D(\tau_S,\tau_T)\]
-
-
Thus,
\[\boxed{\text{large foundation teacher}\rightarrow\text{distillation}\rightarrow\text{real-time driving policy}.}\]
AlpaSim
- AlpaSim is NVIDIA’s open-source closed-loop autonomous-driving simulation framework.
-
The simulator implements
\[s_{t+1}=F_{\mathrm{sim}}(s_t,a_t)\]-
while the policy observes
\[o_t=O_{\mathrm{sim}}(s_t)\]-
and produces
\[a_t=\pi_\theta(o_t)\]
-
-
Why Closed-Loop Simulation Changes the Optimization Problem
-
Trajectory similarity is not identical to driving quality:
\[\boxed{\text{trajectory similarity}\neq\text{driving quality}.}\] -
A more appropriate objective is closer to
\[R=R_{\mathrm{safety}}+R_{\mathrm{progress}}+R_{\mathrm{comfort}}+R_{\mathrm{rules}}+R_{\mathrm{interaction}}\]
AlpaGym
- AlpaGym connects model decisions with feedback from closed-loop simulation for reinforcement-learning post-training.
-
The learning loop is
\[\pi_{\theta_k}\rightarrow\text{AlpaSim Rollout}\rightarrow R(\tau)\rightarrow\text{RL Update}\rightarrow\pi_{\theta_{k+1}}\] -
Imitation learning seeks
\[\pi_\theta\approx\pi_{\mathrm{human}}\]-
whereas reinforcement learning seeks
\[\pi_\theta=\arg\max_\pi\mathbb{E}[R(\tau)]\]
-
Closed-Loop RL for Driving
-
A reward can combine
\[R(\tau)=w_sR_{\mathrm{safety}}+w_pR_{\mathrm{progress}}+w_cR_{\mathrm{comfort}}+w_rR_{\mathrm{rules}}\] -
For example,
\[R_{\mathrm{safety}}=-\mathbb{1}[\text{collision}]\] \[R_{\mathrm{progress}}=d_{t-1}^{\mathrm{goal}}-d_t^{\mathrm{goal}}\]-
and
\[R_{\mathrm{comfort}}=-\sum_t\|j_t\|^2\]
-
Counterfactual Driving Experience
-
Starting from state \(s_t\), simulation can test candidate trajectories
\[\tau^{(1)},\tau^{(2)},\ldots,\tau^{(K)}\] -
Each produces a different future:
\[s_{t+H}^{(k)}=F(s_t,\tau^{(k)})\] -
This transforms driving data from what happened into what could have happened.
World Models for Autonomous Driving
-
A driving world model estimates
\[p_\phi(o_{t+1:t+H}\mid o_{\leq t},a_{t:t+H})\] -
This creates the loop
\[\boxed{\text{Driving Policy}\rightarrow\text{Candidate Action}\rightarrow\text{World Model}\rightarrow\text{Predicted Consequence}\rightarrow\text{Policy Improvement}.}\]
Neural Reconstruction and Digital Twins
-
Given fleet observations
\[\{I_t,L_t,P_t\}_{t=1}^{T}\]-
a system can reconstruct a three-dimensional scene
\[\hat E=f_{\mathrm{reconstruct}}(I_{1:T},L_{1:T},P_{1:T})\]
-
-
The original scenario can then become a family of counterfactual scenarios:
\[E\rightarrow\{E_1,E_2,\ldots,E_N\}\]
Long-Tail Scenario Generation
-
A generative scenario engine can vary factors such as speed, visibility, pedestrian timing, road geometry, weather, and traffic density:
\[E_i=G(E_0,z_i)\] -
The key objective is behaviorally meaningful scenario diversity.
The Autonomous Driving Data Engine
-
A modern Physical AI data engine can combine
\[\mathcal{D}=\mathcal{D}_{\mathrm{fleet}}+\mathcal{D}_{\mathrm{human}}+\mathcal{D}_{\mathrm{sim}}+\mathcal{D}_{\mathrm{world-model}}+\mathcal{D}_{\mathrm{RL}}\] -
The resulting loop is
\[\boxed{\text{Deploy}\rightarrow\text{Mine Failures}\rightarrow\text{Reconstruct}\rightarrow\text{Generate Variants}\rightarrow\text{Simulate}\rightarrow\text{Post-Train}\rightarrow\text{Deploy}.}\]
Safety Is a Trajectory Property
-
A simple clearance constraint is
\[d(\tau_{\mathrm{ego}}(t),\tau_i(t))>d_{\min}\] -
Robust safety also involves uncertainty:
\[p(\tau_i\mid o_{\leq t})\] -
A desired risk constraint is
\[P(\text{collision}\mid\tau_{\mathrm{ego}})<\epsilon\]
Safety Layers Around Learned Policies
-
A safety layer can transform nominal learned action
\[a_t^{\mathrm{policy}}\]-
into
\[a_t^{\mathrm{safe}}=\mathcal{S}(a_t^{\mathrm{policy}},s_t)\]
-
-
This produces
\[\boxed{\text{learned intelligence}+\text{independent safety envelope}.}\]
Reasoning Is Not a Safety Proof
-
Natural-language reasoning can improve interpretability, training, and debugging, but
\[\boxed{\text{explainability}\neq\text{verification}.}\] -
Reasoning traces should be evaluated for consistency with observed evidence and behavior, while trajectory safety should be evaluated independently.
The Driving Physical AI Stack
-
The emerging architecture can be summarized as
\[\boxed{ \begin{array}{c} \text{Fleet + Synthetic + Simulation Data}\\ \downarrow\\ \text{Multimodal Foundation Model}\\ \downarrow\\ \text{Scene Reasoning}\\ \downarrow\\ \text{Trajectory Action Expert}\\ \downarrow\\ \text{Safety / Control Layer}\\ \downarrow\\ \text{Vehicle}\\ \downarrow\\ \text{Closed-Loop Environment}\\ \downarrow\\ \text{New Experience}\\ \circlearrowleft \end{array} }\] - The transition from classical autonomous driving to Physical AI includes shifts from hand-designed to learned representations, separate modules to jointly optimized systems, trajectory imitation to reasoning plus generative planning, open-loop evaluation to closed-loop simulation, and offline supervised learning to closed-loop post-training.
-
The result is an increasingly general development pattern:
\[\boxed{\text{Reason}\rightarrow\text{Act}\rightarrow\text{Simulate Consequences}\rightarrow\text{Learn}.}\] - The next section will examine World Models for Physical AI in detail, including predictive dynamics, latent world models, model-based reinforcement learning, Dreamer, Cosmos, video world models, neural reconstruction, digital twins, counterfactual simulation, and how world models can become learned environments for training physical agents.
- It will also include the requested reference to the World Models primer for a more detailed treatment of VLAs and World Models.
World Models for Physical AI
Why Physical Agents Need World Models
-
A policy answers:
\[\text{What should I do?}\] -
A world model answers a different question:
\[\text{What is likely to happen if I do it?}\] -
For an agent interacting with a physical environment, this distinction is fundamental. A policy maps observations and goals to actions,
\[a_t \sim \pi_\theta(a_t\mid o_{\leq t},g)\]-
while a world model approximates the environment’s dynamics,
\[p_\phi(s_{t+1}\mid s_t,a_t)\]-
or, over a longer horizon,
\[p_\phi(s_{t+1:t+H}\mid s_t,a_{t:t+H-1})\]
-
-
- The model can therefore simulate possible consequences before the agent commits to them in the physical world.
- World Models by Ha and Schmidhuber (2018) demonstrated the core idea using learned compressed spatial and temporal representations, including training a controller inside trajectories generated by the learned environment model.
- A substantially more detailed treatment of world models, JEPA-style predictive learning, and their relationship to Vision-Language-Action models is available in the World Models and JEPA primer.
World Models as Learned Dynamics
-
Consider an environment with latent state
\[s_t\]-
action
\[a_t\]-
and observation
\[o_t\]
-
-
-
The physical environment evolves according to an unknown transition function:
\[s_{t+1} \sim P( s_{t+1}\mid s_t,a_t )\] -
A learned world model approximates this distribution:
\[\hat P_\phi( s_{t+1}\mid s_t,a_t ) \approx P( s_{t+1}\mid s_t,a_t )\] -
In partially observed environments, the agent does not have direct access to \(s_t\). Instead, an encoder constructs a latent representation:
\[z_t = f_\phi( o_{\leq t} )\] -
The dynamics model then predicts
\[\hat z_{t+1} = g_\phi( z_t,a_t )\] -
A decoder may optionally reconstruct observations:
\[\hat o_{t+1} = d_\phi( \hat z_{t+1} )\] -
This gives the canonical learned dynamics pipeline:
\[\boxed{ o_t \rightarrow z_t \xrightarrow{a_t} z_{t+1} \rightarrow o_{t+1}. }\]
Why Predict in Latent Space?
- Predicting every pixel of the future can force a model to allocate capacity to details irrelevant to action.
- Suppose a robot moves toward a cup. The future image may depend on shadows, reflections, textures, background motion, sensor noise, and lighting.
-
Yet the controller may primarily need to know
\[\text{cup pose}, \quad \text{gripper pose}, \quad \text{contact state}, \quad \text{collision geometry}\] -
A latent world model instead learns
\[z_t=f(o_t)\]-
and predicts
\[z_{t+1}=g(z_t,a_t)\]
-
-
The objective can be
\[\mathcal{L}_{\mathrm{pred}} = \| g_\phi(z_t,a_t) - \operatorname{sg}(z_{t+1}) \|_2^2\]- where \(\operatorname{sg}\) denotes stop-gradient.
- The desired latent space preserves information useful for future prediction and control while discarding irrelevant visual variability.
Observation Models Versus State Models
- World models can operate at several levels.
-
A pixel-space model predicts
\[p( o_{t+1} \mid o_{\leq t},a_t )\] -
A latent model predicts
\[p( z_{t+1} \mid z_t,a_t )\] -
A structured state model may instead predict explicit quantities such as
\[s_t = [ p_{\mathrm{objects}}, v_{\mathrm{objects}}, q_{\mathrm{robot}}, \dot q_{\mathrm{robot}}, c_{\mathrm{contacts}} ]\] - Each representation has different advantages.
- Pixel models can preserve rich visual information. Structured models provide interpretable physical state. Latent models can learn compact representations without requiring manually defined state variables.
- Modern foundation world models increasingly combine these ideas by learning rich latent dynamics while retaining the ability to decode photorealistic observations.
Deterministic Versus Stochastic Dynamics
-
A deterministic model assumes
\[z_{t+1} = f_\phi(z_t,a_t)\] - But physical environments often contain uncertainty.
- A pedestrian may turn left or right. An object may slip or remain stable. Another robot may respond in several plausible ways.
-
A stochastic world model instead learns
\[p_\phi( z_{t+1} \mid z_t,a_t )\] -
Over multiple timesteps,
\[p_\phi( z_{t+1:t+H} \mid z_t,a_{t:t+H-1} )\]- represents a distribution over possible futures.
-
This distinction matters because
\[\boxed{ \text{one predicted future} \neq \text{the distribution of plausible futures}. }\] - Physical planning should often reason over the latter.
Multi-Step Prediction
- One-step accuracy is not sufficient for planning.
-
A world model used for decision making must recursively predict
\[\hat z_{t+1} = f_\phi(z_t,a_t)\] \[\hat z_{t+2} = f_\phi(\hat z_{t+1},a_{t+1})\]-
and eventually
\[\hat z_{t+H}\]
-
-
Small errors can compound:
\[\epsilon_{t+H} \approx \sum_{k=1}^{H} J_k\epsilon_{t+k}\]- where \(J_k\) represents how prediction errors propagate through the learned dynamics.
- Long-horizon consistency is therefore one of the central challenges in world-model research.
Model-Based Reinforcement Learning
- Once an agent possesses a learned dynamics model, it can use that model to improve its policy.
- This is the central idea of model-based reinforcement learning.
-
Instead of obtaining every trajectory from the real environment,
\[\tau \sim P_{\mathrm{real}}\]-
the agent can generate imagined trajectories:
\[\hat\tau \sim P_\phi\]
-
-
A candidate action sequence
\[A = [ a_t,\ldots,a_{t+H-1} ]\]-
can be evaluated by rolling it through the model:
\[\hat s_{t+k+1} \sim P_\phi( \hat s_{t+k+1} \mid \hat s_{t+k}, a_{t+k} )\]
-
-
The planner then estimates
\[J(A) = \sum_{k=0}^{H-1} \gamma^k r( \hat s_{t+k}, a_{t+k} )\] -
Planning selects
\[A^* = \arg\max_A J(A)\] -
Only the first action may be executed:
\[a_t=A_0^*\] - The agent then observes the real environment and replans.
- This is learned model-predictive control.
Planning by Imagination
- A world model transforms planning into counterfactual search.
-
Given current state \(s_t\), consider candidate futures
\[A^{(1)},A^{(2)},\ldots,A^{(K)}\] -
The world model generates
\[\hat\tau^{(k)} = M_\phi( s_t,A^{(k)} )\] -
A reward or value model evaluates each rollout:
\[V_k = R( \hat\tau^{(k)} )\] -
The agent selects
\[k^* = \arg\max_k V_k\] -
Conceptually,
\[\boxed{ \text{Imagine} \rightarrow \text{Evaluate} \rightarrow \text{Act}. }\] - This resembles human mental simulation: considering possible consequences before acting.
Dreamer
- Mastering Diverse Domains through World Models by Hafner et al. (2023) introduced DreamerV3, which learns a world model and trains behavior using imagined trajectories generated within its learned latent dynamics; the same configuration was demonstrated across more than 150 tasks.
-
Dreamer’s architecture can be summarized as
\[o_t \rightarrow z_t \rightarrow \text{latent dynamics} \rightarrow \hat z_{t+1:t+H}\] -
An actor
\[\pi_\theta(a_t\mid z_t)\]-
and critic
\[V_\psi(z_t)\]- are trained using imagined trajectories.
-
-
Rather than repeatedly querying the physical environment, policy optimization occurs largely inside
\[\boxed{ \text{learned latent imagination}. }\] - This substantially increases data efficiency when real interaction is expensive.
Representation, Dynamics, Reward, and Continuation Models
- A practical latent world model often contains several components.
-
The encoder maps observations to representations:
\[z_t=f_{\mathrm{enc}}(o_t)\] -
The dynamics model predicts future latent states:
\[z_{t+1} \sim p_\phi( z_{t+1}\mid z_t,a_t )\] -
A reward model predicts
\[\hat r_t = f_r(z_t,a_t)\] -
A continuation model may predict whether the episode continues:
\[\hat c_t = P( \text{continue}\mid z_t )\] -
The learned simulator can therefore provide most of the information required for reinforcement learning:
\[\boxed{ \text{state} + \text{transition} + \text{reward} + \text{termination}. }\]
World Models as Data Generators
- World models can also generate training data rather than directly controlling a policy.
-
Suppose a robot dataset contains
\[\mathcal{D}_{\mathrm{real}} = \{ \tau_1,\ldots,\tau_N \}\] -
A generative world model can produce additional trajectories
\[\mathcal{D}_{\mathrm{synthetic}} = \{ \hat\tau_1,\ldots,\hat\tau_M \}\] -
Training then uses
\[\mathcal{D} = \mathcal{D}_{\mathrm{real}} \cup \mathcal{D}_{\mathrm{synthetic}}\] -
This creates a different use of a world model:
\[\boxed{ \text{World Model} \rightarrow \text{Synthetic Experience} \rightarrow \text{Policy Training}. }\] - The GR00T-Dreams and DreamGen workflows discussed earlier are examples of this broader pattern.
Video Generation Becomes World Modeling
-
A conventional video generator models
\[p( I_{t+1:T} \mid I_{\leq t},c )\]- where \(c\) might be text.
-
A world model becomes substantially more useful for Physical AI when generation is conditioned on actions:
\[p( I_{t+1:T} \mid I_{\leq t},a_{t:T},c )\] - The distinction is crucial.
-
A video model asks:
\[\text{What might happen next?}\] -
An action-conditioned world model asks:
\[\text{What might happen if the agent does this?}\] -
Thus,
\[\boxed{ \text{video prediction} + \text{action conditioning} \rightarrow \text{interactive world modeling}. }\]
Foundation World Models
- Traditional world models are often trained for one environment.
-
A foundation world model instead seeks a broad prior over physical environments:
\[p_\phi( \text{future} \mid \text{history}, \text{actions}, \text{context} )\] -
The model can then be adapted to a specific downstream environment:
\[M_{\mathrm{foundation}} \xrightarrow{\mathrm{post\text{-}training}} M_{\mathrm{domain}}\] - This mirrors the transition from task-specific language models to LLMs and from task-specific robot policies to VLAs.
-
The desired progression is
\[\boxed{ \text{task-specific simulator} \rightarrow \text{world foundation model} \rightarrow \text{specialized physical simulator}. }\]
NVIDIA Cosmos
- Cosmos World Foundation Model Platform for Physical AI by NVIDIA et al. (2025) introduced an open platform for building customized world models for Physical AI, including pretrained world foundation models, video tokenizers, data-curation infrastructure, and post-training workflows.
-
Cosmos explicitly frames Physical AI as requiring two complementary learned systems:
\[\boxed{ \text{policy model} + \text{world model}. }\] -
The policy approximates
\[\pi_\theta( a_t \mid o_{\leq t},g )\]-
while the world model approximates
\[p_\phi( o_{t+1:t+H} \mid o_{\leq t}, a_{t:t+H-1} )\]
-
- Together, they form an agent-environment pair that can increasingly be trained and evaluated digitally before deployment into the physical world.
The Cosmos Data Pipeline
- World foundation models require enormous quantities of video.
-
However,
\[\text{raw video} \neq \text{high-quality physical training data}\] -
A scalable pipeline must perform operations such as
\[\text{ingestion} \rightarrow \text{filtering} \rightarrow \text{deduplication} \rightarrow \text{captioning} \rightarrow \text{quality scoring} \rightarrow \text{sharding}\] - Cosmos therefore treats data curation as part of the world-model platform rather than as a preprocessing detail.
-
This mirrors the broader Physical AI pattern:
\[\boxed{ \text{model capability} \approx f( \text{architecture}, \text{data engine}, \text{post-training} ). }\]
Video Tokenization
- High-resolution video is computationally expensive.
-
Suppose a video tensor has shape
\[T\times H\times W\times C\] -
A tokenizer compresses it into latent tokens:
\[V \xrightarrow{\mathrm{encoder}} Z\] -
If the compression ratio is \(r\),
\[|Z| \ll |V|\] -
The generative model operates in this compressed space:
\[p_\theta( Z_{t+1:T} \mid Z_{\leq t},c )\] -
A decoder reconstructs video:
\[\hat V = D(Z)\] - Efficient tokenization is therefore a foundational systems problem for large-scale world modeling.
Diffusion World Models
- Many modern visual world models use diffusion-style generation.
-
Starting with clean latent future \(z_0\),
\[z_k = \sqrt{\bar\alpha_k}z_0 + \sqrt{1-\bar\alpha_k}\epsilon\] -
The model learns
\[\epsilon_\theta( z_k,k,c )\]-
under
\[\mathcal{L}_{\mathrm{diff}} = \mathbb{E} \left[ \| \epsilon-\epsilon_\theta(z_k,k,c) \|_2^2 \right]\]
-
-
For an action-conditioned model,
\[c = \{ o_{\leq t}, a_{t:t+H}, l \}\] - The resulting generated future depends explicitly on the behavior being simulated.
Google DeepMind Genie
- Genie: Generative Interactive Environments by Bruce et al. (2024) introduced an 11B-parameter generative interactive environment trained from unlabeled internet video, combining a spatiotemporal video tokenizer, autoregressive dynamics model, and latent action model.
- A particularly important idea is latent action discovery.
-
Given observations
\[o_t,o_{t+1}\]-
the model can infer latent action
\[\hat a_t = f_{\mathrm{latent}} ( o_t,o_{t+1} )\]
-
- This provides a mechanism for learning controllable dynamics even when internet video does not contain explicit action labels.
-
The conceptual transformation is
\[\boxed{ \text{unlabeled video} \rightarrow \text{latent actions} \rightarrow \text{interactive environment}. }\]
Genie 2
- Genie 2: A large-scale foundation world model extended this direction to action-controllable 3D environments that can be generated from an initial image and interacted with by humans or agents.
-
Its inference loop can be represented as
\[z_{t+1} \sim p_\phi( z_{t+1} \mid z_{\leq t},a_t )\] -
After decoding,
\[I_{t+1} = D(z_{t+1})\] -
The agent then chooses another action:
\[a_{t+1} = \pi( I_{\leq t+1} )\] - Thus the generated environment becomes interactive rather than a predetermined video.
Genie 3 and Real-Time Interactive Worlds
- Genie 3: A new frontier for world models extends foundation world modeling toward longer-lived, real-time interactive environments generated from text prompts.
-
The trajectory is significant:
\[\text{video generation} \rightarrow \text{interactive video} \rightarrow \text{persistent generated environment}\] - A sufficiently capable interactive world model begins to function as a learned simulator.
-
Instead of constructing every training environment manually,
\[E_i = \text{human-designed simulator}\]-
one can imagine generating environments from descriptions:
\[E_i \sim p_\phi( E\mid\text{prompt} )\]
-
- This could make environment generation itself scalable.
World Models as Infinite Curricula
-
Suppose an agent struggles with a particular capability:
\[c = \text{opening unfamiliar doors}\] -
A generative world model could construct environments
\[E_1,E_2,\ldots,E_N\]- containing variations of that challenge.
-
The agent trains:
\[\pi_k \xrightarrow{E_{1:N}} \pi_{k+1}\] -
Failures are identified:
\[F_{k+1} = \operatorname{Failures}( \pi_{k+1} )\] -
The world model then generates new environments concentrated around those failures:
\[E_{N+1:N+M} \sim p_\phi( E\mid F_{k+1} )\] -
The resulting loop is
\[\boxed{ \text{Agent} \rightarrow \text{Failures} \rightarrow \text{World Generation} \rightarrow \text{Training} \rightarrow \text{Agent}. }\] -
This turns environment generation into part of the learning algorithm.
Counterfactual Simulation
- One of the most important capabilities of a world model is generating alternative futures from the same state.
-
Given
\[s_t\]-
consider actions
\[a_t^{(1)}, a_t^{(2)},\ldots,a_t^{(K)}\]
-
-
The model generates
\[\hat s_{t+1:t+H}^{(k)} \sim p_\phi( s_{t+1:t+H} \mid s_t,a_t^{(k)} )\] - The agent can evaluate each outcome.
- This is especially valuable when real-world exploration is expensive or unsafe.
- A robot need not physically attempt every grasp.
- An autonomous vehicle need not physically execute every evasive maneuver.
-
The world model allows the system to ask
\[\boxed{ \text{What if?} }\]- before acting.
World Models for Policy Evaluation
-
Suppose two candidate policies are
\[\pi_A\]-
and
\[\pi_B\]
-
-
Instead of immediately deploying both physically, a world model can generate
\[\tau_A \sim M_\phi(\pi_A)\]-
and
\[\tau_B \sim M_\phi(\pi_B)\]
-
-
An evaluator computes
\[R_A = R(\tau_A)\] \[R_B = R(\tau_B)\] - This enables rapid iteration.
- However, this procedure is trustworthy only when the world model is accurate in the regions visited by the evaluated policies.
Model Exploitation
- A central failure mode of model-based RL occurs when the policy discovers inaccuracies in the learned model.
-
Suppose
\[\hat R(s,a) > R_{\mathrm{real}}(s,a)\] -
Optimization may deliberately seek such regions:
\[\pi^* = \arg\max_\pi \mathbb{E}_{M_\phi}[R]\] - The policy then succeeds inside the learned simulator while failing in reality.
- This is model exploitation.
-
The problem becomes more severe as optimization pressure increases:
\[\boxed{ \text{stronger optimizer} + \text{imperfect model} \rightarrow \text{greater exploitation risk}. }\] - World-model-based policy improvement therefore requires uncertainty estimation, real-world validation, and mechanisms for keeping imagined trajectories near regions where the model is reliable.
Uncertainty-Aware World Models
- A useful world model should represent not only predictions but confidence.
-
One approach uses an ensemble
\[\{ M_{\phi_1},\ldots,M_{\phi_K} \}\] -
For candidate transition \(x\), disagreement
\[U(x) = \operatorname{Var}_k [ M_{\phi_k}(x) ]\]- can approximate epistemic uncertainty.
-
Planning can penalize uncertain trajectories:
\[J(\tau) = R(\tau) - \lambda U(\tau)\] - The agent therefore prefers high-reward futures that the world model also understands.
The Reality Gap
- World models introduce a learned analogue of the traditional sim-to-real problem.
-
For a conventional simulator,
\[P_{\mathrm{sim}} \neq P_{\mathrm{real}}\] -
For a learned world model,
\[P_\phi \neq P_{\mathrm{real}}\] - Thus both physics simulators and neural world models face a reality gap.
- Their errors, however, may differ.
- A physics simulator may fail because physical parameters are inaccurate.
- A neural world model may fail because the training distribution did not sufficiently cover the relevant state-action region.
- This suggests combining them.
Neural World Models and Physics Simulators
-
A physics engine provides explicit dynamics:
\[s_{t+1} = F_{\mathrm{physics}} ( s_t,a_t,\phi )\] -
A neural world model learns dynamics from data:
\[s_{t+1} \sim M_\theta( s_t,a_t )\] -
Hybrid systems can use both:
\[s_{t+1} = F( F_{\mathrm{physics}}, M_\theta )\] - The physics engine supplies structural constraints and reliable contact mechanics where modeled accurately.
- The neural model can capture appearance, unmodeled dynamics, human behavior, and other phenomena difficult to specify analytically.
-
Thus,
\[\boxed{ \text{physics simulation} + \text{neural world modeling} }\]- is likely more useful than treating them as mutually exclusive approaches.
Neural Reconstruction and Digital Twins
- Another important approach begins with the real world rather than generating an environment from scratch.
-
Suppose a robot or vehicle records
\[\mathcal{O} = \{ I_t, D_t, P_t \}_{t=1}^{T}\]-
where \(I_t\) is imagery, \(D_t\) is depth or lidar information, and
\(P_t\) is sensor pose.
-
-
A reconstruction system estimates
\[\hat E = f_{\mathrm{recon}}( \mathcal{O} )\] - The resulting environment is a digital twin of a real scene.
-
It can then be modified:
\[\hat E \rightarrow \{ \hat E_1,\ldots,\hat E_N \}\] - This is particularly powerful for rare events.
- One real-world scenario can become hundreds of controlled counterfactual variants.
Generative Digital Twins
- Traditional digital twins attempt to reproduce a specific environment.
-
Generative world models allow the twin to become a distribution:
\[E \sim p_\phi( E\mid E_{\mathrm{real}} )\] -
The system can vary
\[\text{objects}, \quad \text{agents}, \quad \text{lighting}, \quad \text{weather}, \quad \text{geometry}, \quad \text{dynamics}\] -
The goal shifts from
\[\text{reconstruct this scene}\]-
to
\[\text{construct the neighborhood of plausible scenes around it}\]
-
- This turns digital twins into training distributions.
World Models and Synthetic Data
- Synthetic data generation becomes particularly powerful when it is targeted.
-
Suppose evaluation identifies failure set
\[\mathcal{F} = \{ f_1,\ldots,f_K \}\] -
A world model generates
\[\mathcal{D}_{\mathrm{synthetic}} \sim p_\phi( \tau\mid\mathcal{F} )\] -
The policy is post-trained:
\[\pi_{k+1} = \operatorname{Train} ( \pi_k, \mathcal{D}_{\mathrm{real}} \cup \mathcal{D}_{\mathrm{synthetic}} )\] - This creates targeted synthetic experience rather than indiscriminate data expansion.
-
The data engine becomes
\[\boxed{ \text{Evaluate} \rightarrow \text{Find Failure} \rightarrow \text{Generate Counterfactuals} \rightarrow \text{Post-Train} \rightarrow \text{Re-evaluate}. }\]
World Models and VLAs
- A VLA and world model solve complementary conditional prediction problems.
-
The VLA predicts actions:
\[\boxed{ \text{VLA}: \quad p_\theta( a_{t:t+H} \mid o_{\leq t},g ). }\] -
The world model predicts consequences:
\[\boxed{ \text{World Model}: \quad p_\phi( o_{t+1:t+H} \mid o_{\leq t},a_{t:t+H} ). }\] -
Connecting them creates a closed reasoning loop:
\[\boxed{ \text{VLA proposes} \rightarrow \text{World Model predicts} \rightarrow \text{Evaluator scores} \rightarrow \text{VLA acts}. }\] - This can be interpreted as a learned version of model-predictive control.
- Instead of trusting the first proposed action sequence, the system can internally test candidate behaviors before executing them.
- For a more extensive treatment of this relationship, including JEPA-style predictive architectures, latent dynamics, VLAs, and world-model learning, see the World Models and JEPA primer.
World Models as Critics
- A world model can also support policy post-training without directly selecting actions.
-
Suppose a VLA proposes
\[A = a_{t:t+H}\] -
The world model predicts
\[\hat\tau = M_\phi( o_{\leq t},A )\] -
A critic evaluates
\[R = C( \hat\tau )\] -
The resulting signal can train the VLA:
\[\theta \leftarrow \theta + \eta \nabla_\theta J(\pi_\theta)\] -
Thus,
\[\boxed{ \text{World Model} \rightarrow \text{Experience Generator} + \text{Counterfactual Evaluator}. }\] - This creates a natural bridge between foundation world models and reinforcement-learning post-training.
From Static Datasets to Interactive Data Engines
-
Traditional supervised learning assumes
\[\mathcal{D} = \text{fixed}\] -
Physical AI increasingly uses
\[\mathcal{D}_{k+1} = f( \pi_k, M_k, \mathcal{E}_k )\]- where the next training dataset depends on the current policy, world model, and evaluation results.
- The dataset therefore evolves with the agent.
-
This is a major conceptual shift:
\[\boxed{ \text{dataset} \rightarrow \text{data engine}. }\] - World models make this possible because they can generate targeted experience on demand.
World Models as Learned Simulators
- A sufficiently capable world model begins to approximate the interface of a simulator.
-
A simulator exposes
\[\operatorname{step}(s_t,a_t) \rightarrow ( s_{t+1},r_t )\] -
A learned world model can approximate
\[\operatorname{step}_\phi( z_t,a_t ) \rightarrow ( z_{t+1}, \hat r_t )\] -
If observations can also be rendered,
\[o_{t+1} = D(z_{t+1})\]- the model becomes a complete learned interactive environment.
-
This leads toward a powerful abstraction:
\[\boxed{ \text{World Model} \approx \text{Generative Simulator}. }\]
The World-Model Physical AI Stack
-
The emerging stack can be summarized as
\[\boxed{ \begin{array}{c} \text{Real-World Video + Robot Data}\\ \downarrow\\ \text{Representation Learning}\\ \downarrow\\ \text{World Foundation Model}\\ \downarrow\\ \text{Domain Post-Training}\\ \downarrow\\ \text{Interactive Simulation}\\ \downarrow\\ \text{Policy Rollouts}\\ \downarrow\\ \text{Counterfactual Evaluation}\\ \downarrow\\ \text{Policy Post-Training}\\ \downarrow\\ \text{Physical Deployment}\\ \circlearrowleft \end{array} }\] -
The central idea is that a Physical AI system increasingly learns both sides of the interaction:
\[\boxed{ \underbrace{\pi_\theta}_{\text{agent}} + \underbrace{M_\phi}_{\text{world}}. }\] - The policy learns how to act.
- The world model learns what actions do.
-
Together they enable an increasingly important training paradigm:
\[\boxed{ \text{observe reality} \rightarrow \text{learn a world} \rightarrow \text{imagine experience} \rightarrow \text{improve the agent} \rightarrow \text{return to reality}. }\] - The next section will examine the Physical AI Data Engine in detail, including teleoperation, robot demonstrations, Open X-Embodiment, human video, action normalization, cross-embodiment data, simulation, domain randomization, synthetic trajectories, world-model-generated experience, data mixtures, filtering, and curriculum construction.
The Physical AI Data Engine
Why Physical AI Is a Data Problem
-
Physical AI is constrained not only by model capacity, but by the amount, diversity, and quality of experience available for training. Language models can learn from enormous corpora of naturally occurring text. Physical agents require observations paired with actions, state transitions, outcomes, and often task descriptions:
\[\tau = \{ (o_t,s_t,a_t,r_t) \}_{t=1}^{T}\] - Collecting such trajectories on real robots is expensive because every sample requires interaction with physical hardware. Robots operate in real time, demonstrations require operators or autonomous policies, hardware experiences wear and failure, and unsafe exploration cannot simply be scaled arbitrarily.
-
Consequently, the central data challenge is:
\[\boxed{ \text{How can limited physical experience be transformed into sufficiently broad training experience?} }\] - Modern Physical AI systems increasingly answer this through a data engine combining real demonstrations, heterogeneous robot datasets, human video, simulation, synthetic data, world models, autonomous rollouts, and targeted failure collection.
-
The resulting pipeline is not a static dataset:
\[\boxed{ \text{Collect} \rightarrow \text{Train} \rightarrow \text{Evaluate} \rightarrow \text{Discover Failures} \rightarrow \text{Generate Data} \rightarrow \text{Retrain}. }\]
What Constitutes a Robot Trajectory?
- A robot demonstration contains substantially more structure than an image-caption pair.
-
A trajectory may contain
\[\tau = \{ I_t^{1:N}, D_t, q_t, \dot q_t, g_t, a_t, l, r_t \}_{t=1}^{T}\]- where:
- \(I_t^{1:N}\) represents one or more camera streams.
- \(D_t\) represents depth or other spatial sensing.
- \(q_t\) and \(\dot q_t\) represent robot joint positions and velocities.
- \(g_t\) represents gripper state.
- \(a_t\) represents the executed action.
- \(l\) represents the language instruction.
- \(r_t\) optionally represents reward or task feedback.
-
The training example is therefore temporally structured:
\[(o_{\leq t},l) \rightarrow a_{t:t+H}\] - For a VLA, the important supervision is not simply what appears in an image, but what action should follow from the current physical context.
Sources of Physical AI Data
-
A modern Physical AI data mixture may be written as
\[\mathcal{D} = \mathcal{D}_{\mathrm{robot}} + \mathcal{D}_{\mathrm{human}} + \mathcal{D}_{\mathrm{internet}} + \mathcal{D}_{\mathrm{sim}} + \mathcal{D}_{\mathrm{synthetic}} + \mathcal{D}_{\mathrm{autonomous}}\] - These sources provide different forms of supervision.
- Robot demonstrations provide direct action labels.
- Human video provides enormous visual coverage of physical interactions but usually lacks robot actions.
- Internet-scale image-text and video-text data provides semantic knowledge.
- Simulation provides controllable state-action trajectories.
- Generative models can synthesize new observations and scenarios.
- Autonomous rollouts provide experience under the current policy’s own state distribution.
- The challenge is therefore not merely collecting more data. It is combining heterogeneous data sources whose semantics, embodiments, observation spaces, and action spaces differ.
Human Demonstrations
- The most direct method for collecting robot data is to ask a human to demonstrate a task.
-
A demonstration provides
\[\tau_E = \{ (o_t,a_t) \}_{t=1}^{T}\]- where the subscript \(E\) denotes an expert.
-
Behavioral cloning then learns
\[\pi_\theta(a_t\mid o_{\leq t},l)\]-
by minimizing
\[\mathcal{L}_{\mathrm{BC}} = - \mathbb{E}_{(o,a)\sim\mathcal{D}_E} [ \log \pi_\theta(a\mid o,l) ]\]
-
- Demonstrations are particularly valuable because they concentrate data around successful behavior rather than requiring the robot to discover useful actions through random exploration.
-
However, demonstration quality strongly influences the learned policy:
\[\boxed{ \text{policy quality} \lesssim \text{coverage and quality of demonstrations}. }\]
Teleoperation
- Teleoperation allows a human to control a robot while the system records observations and actions.
- The operator may use devices such as joysticks, VR controllers, motion-capture systems, leader-follower robot arms, or other interfaces.
-
Conceptually,
\[u_t^{\mathrm{human}} \xrightarrow{f_{\mathrm{teleop}}} a_t^{\mathrm{robot}}\] -
The resulting dataset contains
\[(o_t,a_t^{\mathrm{robot}},l)\] - Teleoperation has become one of the primary mechanisms for collecting manipulation data because it produces robot-native action labels.
- The challenge is scalability. If collecting one hour of robot experience requires approximately one hour of human operation, dataset growth remains constrained by human labor and hardware throughput.
Demonstration Quality Versus Demonstration Quantity
- Not all demonstrations contribute equally.
-
A useful dataset must balance:
\[\text{quantity}, \quad \text{quality}, \quad \text{diversity}\] - Repeating one task thousands of times can improve robustness around that task but may contribute less to generalization than collecting diverse objects, scenes, embodiments, and behaviors.
-
One can view dataset utility as
\[U(\mathcal{D}) = f( N, C_{\mathrm{task}}, C_{\mathrm{object}}, C_{\mathrm{scene}}, C_{\mathrm{embodiment}}, Q )\]- where \(N\) is dataset size, \(C\) terms denote different kinds of coverage, and \(Q\) denotes demonstration quality.
- The data-engine problem is therefore an allocation problem: determining which additional trajectories provide the greatest marginal improvement.
Language Annotation
- A robot trajectory becomes more useful for generalist instruction following when paired with language.
-
A trajectory
\[\tau\]-
can be associated with instruction
\[l = \text{``place the red cup in the sink''}\]
-
-
Training then learns
\[\pi_\theta(a\mid o,l)\] -
Multiple linguistic descriptions can correspond to the same behavior:
\[\{ l_1,l_2,\ldots,l_K \} \leftrightarrow \tau\] -
For example:
\[l_1 = \text{``put away the cup''}\] \[l_2 = \text{``move the red mug into the sink''}\] - Language augmentation can therefore increase linguistic diversity without requiring new physical trajectories.
Open X-Embodiment
- One of the major developments in robot learning has been the aggregation of datasets collected across many institutions and robot platforms.
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models by the Open X-Embodiment Collaboration et al. (2023) assembled data from 22 robot embodiments contributed by 21 institutions, producing a large cross-embodiment resource for training generalist robot policies.
-
The central idea is that instead of learning
\[\pi_{\theta,e}\]-
for a single embodiment \(e\), the model learns
\[\pi_\theta( a \mid o,l,e )\]- across many embodiments.
-
-
The data mixture becomes
\[\mathcal{D} = \bigcup_{e=1}^{E} \mathcal{D}_e\] -
This is conceptually analogous to multilingual language modeling:
\[\text{many languages} \rightarrow \text{shared language model}\]-
while robot learning uses
\[\text{many embodiments} \rightarrow \text{shared physical policy}\]
-
RT-X and Positive Transfer Across Robots
- Open X-Embodiment was used to train RT-X models that share knowledge across robot datasets. The work demonstrated that combining heterogeneous robot experience can improve performance relative to training exclusively on some individual datasets, while also exposing the difficulty of reconciling different embodiments and data distributions.
-
The important principle is positive transfer:
\[\mathcal{D}_A + \mathcal{D}_B \rightarrow \text{better performance on }A\] - This suggests that physical skills contain reusable structure.
-
For example,
\[\text{grasping}, \quad \text{placing}, \quad \text{pushing}, \quad \text{opening}\]- share geometric and semantic regularities even when performed by different robots.
- Cross-embodiment pretraining attempts to capture these regularities.
The Cross-Embodiment Normalization Problem
- Different robots expose different action spaces.
-
Robot \(A\) may use
\[a_t^A \in \mathbb{R}^{7}\]-
while robot \(B\) may use
\[a_t^B \in \mathbb{R}^{14}\]
-
-
One robot may use Cartesian end-effector deltas:
\[a_t = [ \Delta x, \Delta y, \Delta z, \Delta r, g ]\]-
while another exposes joint-space targets:
\[a_t = [ q_1,\ldots,q_n ]\]
-
- Even identical quantities can have different ranges and units.
-
A common normalization is
\[\tilde a_i = \frac{ a_i-\mu_i }{ \sigma_i }\]-
or percentile-based scaling such as
\[\tilde a_i = 2 \frac{ a_i-q_i^{0.01} }{ q_i^{0.99}-q_i^{0.01} } -1\]
-
-
At inference time,
\[a_i = f_e^{-1}( \tilde a_i )\]- where \(f_e^{-1}\) is embodiment-specific.
- Normalization is not merely numerical preprocessing. It defines how heterogeneous physical control spaces are mapped into a shared policy representation.
Embodiment Metadata
-
A generalist policy may explicitly condition on embodiment information:
\[e = \{ \text{robot type}, \text{DoF}, \text{camera configuration}, \text{action specification} \}\] -
The policy becomes
\[\pi_\theta( a \mid o,l,e )\] - Alternatively, embodiment can be inferred implicitly from observations.
- Explicit conditioning is useful when the same semantic instruction must map to substantially different physical actions.
-
For example,
\[\text{``pick up the cup''}\]- could require different kinematics for a single-arm manipulator, dual-arm humanoid, or mobile manipulator.
Dataset Standardization
- Cross-embodiment training requires a common data representation.
- A standardized episode might contain:
episode/
observations/
camera_0
camera_1
proprioception
actions/
language/
timestamps/
embodiment_metadata/
success/
-
At timestep \(t\), the loader constructs
\[x_t = \{ I_{t-k:t}, s_{t-k:t}, l, e \}\]-
and target
\[y_t = a_{t:t+H}\]
-
- The loader must also reconcile camera rates, control rates, missing modalities, action conventions, coordinate frames, and trajectory boundaries.
- For large Physical AI systems, this data infrastructure can be as consequential as the policy architecture itself.
Temporal Alignment
- Sensors and controls frequently operate at different rates.
-
Suppose cameras run at
\[f_{\mathrm{cam}}\]-
while actions are recorded at
\[f_{\mathrm{control}}\]
-
-
The dataset must align
\[I(t) \leftrightarrow a(t)\] -
If timestamps are inaccurate, the model may learn an artificial temporal offset:
\[I_t \rightarrow a_{t+\Delta}\] - For fast physical systems, even modest synchronization errors can materially degrade learning.
- A robust data pipeline therefore preserves high-resolution timestamps and explicitly resamples modalities into a common training timeline.
Human Video as Physical Pretraining Data
- Robot data is scarce, but human video is abundant.
-
A human manipulation video contains useful information about:
\[\text{objects}, \quad \text{affordances}, \quad \text{contacts}, \quad \text{task structure}, \quad \text{temporal dynamics}\] -
The central difficulty is that it does not directly contain robot actions:
\[I_{1:T}^{\mathrm{human}} \not\Rightarrow a_{1:T}^{\mathrm{robot}}\] - Nevertheless, video can provide representation learning, task understanding, motion priors, or latent-action supervision.
-
This creates an attractive scaling route:
\[\boxed{ \text{internet video} \rightarrow \text{physical representation} \rightarrow \text{robot adaptation}. }\]
Learning Representations from Human Activity
- R3M: A Universal Visual Representation for Robot Manipulation by Nair et al. (2022) showed that representations pretrained on large-scale human video can transfer effectively to downstream robotic manipulation.
-
Instead of learning visual features only from robot data,
\[f_{\mathrm{vision}} \leftarrow \mathcal{D}_{\mathrm{robot}}\]-
the model first learns from
\[\mathcal{D}_{\mathrm{human-video}}\]-
then adapts to
\[\mathcal{D}_{\mathrm{robot}}\]
-
-
-
This separates two problems:
\[\boxed{ \text{learning how the physical world looks} }\]-
from
\[\boxed{ \text{learning how a particular robot acts}. }\]
-
Learning Latent Actions from Video
- Human and internet video often lacks explicit actions.
-
One strategy is to infer latent actions:
\[z_t^a = f_{\mathrm{action}} ( o_t,o_{t+1} )\] -
A dynamics model then predicts
\[\hat o_{t+1} = f_{\mathrm{dyn}} ( o_t,z_t^a )\] - The latent variable is encouraged to encode the transformation responsible for moving the environment from one observation to the next.
-
Later, a smaller amount of robot data can ground these latent actions into robot commands:
\[z_t^a \rightarrow a_t^{\mathrm{robot}}\] - This is attractive because it separates scalable physical-dynamics learning from expensive robot-action labeling.
Internet-Scale Semantic Data
- Physical agents also benefit from ordinary vision-language data.
-
Image-text and video-text pretraining can teach concepts such as
\[\text{cup}, \quad \text{drawer}, \quad \text{fragile}, \quad \text{pour}, \quad \text{behind}\] - Robot data then grounds these semantic concepts in action.
-
The training mixture may therefore contain
\[\mathcal{D} = \lambda_{\mathrm{VL}} \mathcal{D}_{\mathrm{VL}} + \lambda_{\mathrm{robot}} \mathcal{D}_{\mathrm{robot}}\] - This is the basic data rationale behind VLA architectures built from pretrained vision-language models.
- Internet-scale data provides semantic breadth; robot data provides physical grounding.
Simulation as a Data Source
- Simulation replaces expensive physical interaction with programmable experience.
-
A simulator implements
\[s_{t+1} = F_{\mathrm{sim}} ( s_t,a_t;\phi )\]- where \(\phi\) contains environment parameters such as masses, friction coefficients, object geometry, lighting, and sensor properties.
-
The simulator can generate
\[\tau_{\mathrm{sim}} = \{ (s_t,o_t,a_t,r_t) \}_{t=1}^{T}\] -
Unlike real-world data collection, simulation can often be parallelized aggressively:
\[\{ E_1,E_2,\ldots,E_N \}\]- can execute simultaneously.
- This makes simulation particularly valuable for reinforcement learning, where policy improvement may require large numbers of environment interactions.
Domain Randomization
- A policy trained in one perfectly deterministic simulator may overfit to its exact parameters.
-
Domain randomization instead samples
\[\phi \sim p(\phi)\] -
Parameters may include
\[\phi = \{ m, \mu, k, c, L, T, \eta \}\]- representing quantities such as mass, friction, stiffness, damping, lighting, textures, and sensor noise.
-
Training becomes
\[\max_\theta \mathbb{E}_{\phi\sim p(\phi)} [ R( \pi_\theta; \phi ) ]\] - The policy therefore learns behavior robust across a family of environments rather than a single simulator configuration.
-
The intuition is
\[\boxed{ \text{make simulation diverse enough that reality becomes another variation}. }\]
Dynamics Randomization
- For control, visual randomization alone is insufficient.
-
Physical parameters can be randomized:
\[m \sim p(m)\] \[\mu \sim p(\mu)\] \[\tau_{\mathrm{motor}} \sim p(\tau_{\mathrm{motor}})\] \[\Delta t_{\mathrm{latency}} \sim p(\Delta t)\] - This exposes the policy to variation in mass, friction, motor strength, latency, and contact dynamics.
-
A robust policy then learns
\[\pi_\theta = \arg\max_\pi \mathbb{E}_{\phi} [ R(\pi;\phi) ]\]
Synthetic Visual Data
- Simulation can also generate labeled visual observations.
-
Because the simulator knows the full scene state, labels such as
\[\text{depth}, \quad \text{segmentation}, \quad \text{pose}, \quad \text{optical flow}, \quad \text{contact}\]- can be produced automatically.
-
If rendering function \(R\) maps scene state to image,
\[I_t = R(s_t,\phi_{\mathrm{visual}})\]- then changing visual parameters generates many appearances of the same physical trajectory.
- This can improve robustness while preserving exact state-action supervision.
Procedural Environment Generation
- Simulation becomes substantially more useful when environments themselves are generated procedurally.
-
Let
\[z \sim p(z)\]- represent environment parameters.
-
A generator creates
\[E = G(z)\] -
Different samples can vary:
\[\text{room layout}, \quad \text{objects}, \quad \text{obstacles}, \quad \text{lighting}, \quad \text{task goals}\] -
The policy trains over
\[E_1,E_2,\ldots,E_N\] - Procedural generation therefore turns environmental diversity into a programmable quantity.
Synthetic Robot Demonstrations
- Simulation can produce not only environments but demonstrations.
-
Suppose a privileged planner has access to simulator state:
\[a_t^* = \pi_{\mathrm{expert}}( s_t )\] -
The training policy receives only observations:
\[o_t = O(s_t)\] -
The resulting synthetic dataset is
\[\mathcal{D}_{\mathrm{synthetic}} = \{ (o_t,a_t^*) \}\] -
The policy then learns
\[\pi_\theta( a_t \mid o_t )\] -
This is privileged-information distillation:
\[\boxed{ \text{simulator-state expert} \rightarrow \text{observation-based physical policy}. }\]
Generative Models as Data Engines
- Generative models add another source of synthetic data.
-
Instead of constructing every simulated scene manually,
\[E_i = \text{human-authored}\]-
a generative model can produce
\[E_i \sim p_\phi( E\mid c )\]- where \(c\) describes a desired scenario.
-
-
For example,
\[c = \text{``cluttered kitchen with partially open drawers''}\] - A world model can then generate corresponding visual futures or interactive environments.
- This allows the data engine to request experience rather than merely consume existing data.
Targeted Data Generation
- Random synthetic generation is often inefficient.
-
Suppose evaluation identifies a failure distribution
\[p_{\mathrm{fail}}(x)\] -
The data generator should preferentially sample
\[x \sim q(x)\]-
where
\[q(x) \propto p_{\mathrm{fail}}(x) \cdot w(x)\]
-
- Here \(w(x)\) can prioritize severity, uncertainty, or expected learning value.
-
The resulting loop is
\[\boxed{ \text{Evaluate} \rightarrow \text{Cluster Failures} \rightarrow \text{Generate Similar Cases} \rightarrow \text{Train} \rightarrow \text{Re-evaluate}. }\] - This is more scalable than increasing dataset size indiscriminately.
GR00T-Dreams as a Synthetic Data Pipeline
- NVIDIA’s GR00T-Dreams applies generative AI and simulation to robot-training data generation. The system uses world-generation and simulation tools to create synthetic robot experiences that can supplement limited real demonstrations.
-
The broader pipeline can be represented as
\[\mathcal{D}_{\mathrm{real}} \rightarrow \text{scenario generation} \rightarrow \mathcal{D}_{\mathrm{synthetic}} \rightarrow \text{policy training}\] -
The objective is to amplify scarce physical demonstrations:
\[N_{\mathrm{real}} \ll N_{\mathrm{train}}\]-
where
\[N_{\mathrm{train}} = N_{\mathrm{real}} + N_{\mathrm{synthetic}}\]
-
- The synthetic data is valuable only if it preserves the task semantics and physical constraints required by the policy.
World-Model-Generated Experience
- A world model provides another mechanism for synthetic trajectories.
-
Given initial context
\[o_{\leq t}\]-
and action sequence
\[a_{t:t+H}\]-
the model generates
\[\hat o_{t+1:t+H} \sim p_\phi( o_{t+1:t+H} \mid o_{\leq t}, a_{t:t+H} )\]
-
-
- These imagined trajectories can be used for planning, policy evaluation, or training.
- The key advantage is that world-model generation can be conditioned on the current policy’s weaknesses.
-
For example,
\[c = \text{``generate scenarios where this grasp fails because the object slips''}\] - The resulting experience can directly target failure modes.
Autonomous Data Collection
- Once a policy becomes sufficiently capable, it can collect its own data.
-
Instead of
\[a_t = a_t^{\mathrm{human}}\]-
the robot executes
\[a_t \sim \pi_\theta( a_t\mid o_t )\]
-
-
The resulting trajectories expose the policy’s actual state distribution:
\[s_t \sim d^{\pi_\theta}(s)\] -
This is important because behavioral cloning trains primarily on
\[s_t \sim d^{\pi_E}(s)\]- the expert distribution.
- Autonomous rollouts reveal states caused by the learner’s own errors.
Human Intervention and Corrections
- Pure autonomous collection can be unsafe or produce low-quality data.
- A human can intervene when the policy begins to fail.
-
Suppose
\[a_t^{\mathrm{policy}} = \pi_\theta(o_t)\] -
When intervention trigger \(I_t=1\),
\[a_t = a_t^{\mathrm{human}}\] - The correction provides particularly valuable supervision because it occurs near the policy’s failure boundary.
-
The resulting dataset contains
\[(o_t,a_t^{\mathrm{human}})\]-
precisely where
\[\pi_\theta(o_t)\]- was inadequate.
-
- This is a powerful form of active data collection.
DAgger and On-Policy Demonstration Collection
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning by Ross et al. (2011) introduced DAgger, which iteratively collects expert labels on states visited by the learner rather than training only on a fixed expert trajectory distribution.
-
The procedure is approximately:
\[\pi_k \rightarrow \text{rollout} \rightarrow s\sim d^{\pi_k} \rightarrow a^*=\pi_E(s) \rightarrow \mathcal{D}_{k+1}\] -
The dataset grows as
\[\mathcal{D}_{k+1} = \mathcal{D}_k \cup \{ (s,a^*) \}\] - This directly addresses covariate shift by training the policy on states it actually visits.
- Modern robot data engines generalize this idea through autonomous rollouts, interventions, corrections, failure mining, and targeted recollection.
Success and Failure Data
- Early imitation-learning datasets often emphasize successful demonstrations.
- But failures contain important information.
-
A successful trajectory tells the model
\[\text{what works}\] -
A failed trajectory can reveal
\[\text{what does not work and why}\] -
Suppose the dataset contains outcome
\[y_\tau \in \{ 0,1 \}\] -
A value or success model can learn
\[V_\psi(\tau) = P( y_\tau=1 \mid \tau )\] - Failures can then support reward learning, critic training, preference learning, and RL post-training.
- The data engine should therefore preserve failure trajectories rather than discarding them automatically.
Automatic Success Detection
- Scaling robot data collection requires automatic labeling of whether tasks succeeded.
-
A success classifier estimates
\[p_\psi( y=1 \mid o_{1:T},l )\] - This classifier may use final images, video, robot state, task metadata, or multimodal foundation models.
-
For instruction
\[l = \text{``put the apple in the bowl''}\]- the evaluator determines whether the final state satisfies the instruction.
- Automatic success detection enables large-scale autonomous rollouts because human reviewers no longer need to label every trajectory.
- However, evaluator errors become part of the training signal, so success models themselves require careful calibration.
Data Filtering
- More data is not always better.
-
A raw trajectory corpus may contain
\[\text{failed teleoperation}, \quad \text{sensor corruption}, \quad \text{idle periods}, \quad \text{incorrect labels}, \quad \text{duplicate trajectories}\] -
Define quality score
\[q_i = Q(\tau_i)\] -
Training data can be filtered using
\[\mathcal{D}' = \{ \tau_i \mid q_i>\delta \}\] -
Alternatively, examples can be weighted:
\[\mathcal{L} = \sum_i w_i \mathcal{L}(\tau_i)\]-
where
\[w_i=f(q_i)\]
-
- Filtering becomes increasingly important as synthetic and autonomous data expand dataset scale.
Deduplication
- Repeated trajectories can distort the effective training distribution.
-
If trajectory \(\tau_i\) appears many times,
\[p_{\mathrm{train}}(\tau_i)\]- becomes artificially large.
- Exact duplicate detection is straightforward, but semantic duplication is harder.
- Two trajectories may be visually different while encoding nearly identical behavior.
-
Embedding-based similarity can estimate
\[d( f(\tau_i), f(\tau_j) ) < \epsilon\] - Near-duplicate clusters can then be downsampled.
- The objective is not maximum raw dataset size but maximum useful diversity.
Data Balancing
-
Suppose the dataset contains
\[90\% \text{pick-and-place}\]-
and
\[1\% \text{drawer opening}\]
-
- Naive sampling trains primarily on the dominant task.
-
Instead, define task sampling distribution
\[p_{\mathrm{train}}(k) \propto w_kN_k^\alpha\]-
where \(N_k\) is the number of examples for task \(k\) and
\(0\leq\alpha\leq1\) controls how strongly raw frequency determines
- sampling.
-
-
When
\[\alpha=1\]- sampling follows the raw dataset.
-
When
\[\alpha=0\]- tasks are sampled uniformly.
- Intermediate values trade off diversity and statistical efficiency.
Mixing Data Sources
- Suppose training uses robot data, simulation, human video, and vision-language data.
-
The objective can be written as
\[\mathcal{L} = \lambda_R \mathcal{L}_{\mathrm{robot}} + \lambda_S \mathcal{L}_{\mathrm{sim}} + \lambda_H \mathcal{L}_{\mathrm{human}} + \lambda_V \mathcal{L}_{\mathrm{VL}}\] -
The coefficients
\[\lambda_R,\lambda_S,\lambda_H,\lambda_V\]- determine the effective training curriculum.
- These weights may change during training.
-
Early training can emphasize broad semantic and physical representation learning:
\[\lambda_H,\lambda_V \uparrow\] -
Later training can emphasize robot-native actions:
\[\lambda_R \uparrow\] - Post-training may emphasize high-quality or task-specific trajectories.
Curriculum Learning
- Not all physical tasks should necessarily be introduced with equal probability from the beginning.
-
A curriculum can progress from
\[\text{simple} \rightarrow \text{complex}\] -
For manipulation:
\[\text{reach} \rightarrow \text{grasp} \rightarrow \text{place} \rightarrow \text{multi-object manipulation} \rightarrow \text{long-horizon task}\] -
More generally, training distribution at step \(k\) is
\[p_k(\tau)\]-
and the curriculum updates it:
\[p_{k+1}(\tau) = f( p_k, \text{performance}_k )\]
-
- Tasks the policy already masters can be downweighted while difficult but learnable tasks receive more sampling probability.
Automatic Curriculum Generation
- A stronger data engine uses evaluation results to construct the curriculum automatically.
-
Suppose task \(i\) has success rate
\[s_i\] -
A simple difficulty weight is
\[w_i = 1-s_i\] - But sampling only the hardest tasks may be inefficient if they are currently impossible.
-
One can instead prioritize tasks near the learning frontier:
\[w_i = f( s_i, \Delta s_i, u_i )\]- where \(\Delta s_i\) measures learning progress and \(u_i\) represents uncertainty.
- This yields an adaptive curriculum concentrated around tasks where additional experience is likely to improve the policy.
Diversity Across Tasks, Objects, and Scenes
- Physical generalization is multidimensional.
-
A useful data engine should cover the Cartesian product
\[\mathcal{G} = \mathcal{T} \times \mathcal{O} \times \mathcal{S} \times \mathcal{E} \times \mathcal{L}\]-
where:
\[\mathcal{T} = \text{tasks}\] \[\mathcal{O} = \text{objects}\] \[\mathcal{S} = \text{scenes}\] \[\mathcal{E} = \text{embodiments}\]-
and
\[\mathcal{L} = \text{language}\]
-
-
-
The number of possible combinations grows approximately as
\[|\mathcal{G}| = |\mathcal{T}| |\mathcal{O}| |\mathcal{S}| |\mathcal{E}| |\mathcal{L}|\] - It is impossible to collect every combination physically.
- Generalization therefore requires strategically sampling this space while relying on compositional transfer.
Measuring Dataset Coverage
- Dataset scale alone is an incomplete statistic.
- Suppose dataset \(A\) contains one million trajectories from ten tasks while dataset \(B\) contains 500,000 trajectories spanning one thousand tasks.
- Which is more useful depends on the target.
-
Coverage can be represented over feature space \(z\):
\[z_i = f(\tau_i)\] -
One can estimate density
\[p_{\mathcal{D}}(z)\]-
and identify sparse regions:
\[p_{\mathcal{D}}(z) < \epsilon\]
-
- These regions become candidates for targeted data collection.
- Thus the data engine can explicitly optimize coverage rather than raw count.
Active Data Collection
-
Suppose candidate trajectory \(x\) has estimated uncertainty
\[U(x)\] -
The system can prioritize
\[x^* = \arg\max_x U(x)\] -
More generally, data value may combine uncertainty, novelty, task importance, and failure severity:
\[V_{\mathrm{data}}(x) = \alpha U(x) + \beta N(x) + \gamma I(x) + \delta F(x)\] -
The collection system then seeks
\[x^* = \arg\max_x V_{\mathrm{data}}(x)\] -
This turns robot data acquisition into an active-learning problem.
Hard-Negative Mining
- Policies often fail on examples close to the decision boundary.
- Suppose the model succeeds when objects are clearly separated but fails when distractors are visually similar.
-
Rather than collecting more random scenes, the data engine generates hard negatives:
\[x_{\mathrm{hard}} \sim p( x \mid \text{near policy failure boundary} )\] -
For language grounding, examples might vary only a critical attribute:
\[\text{red cup} \leftrightarrow \text{orange cup}\] - For manipulation, they might vary grasp geometry slightly.
- Hard-negative mining increases the information density of additional data.
Failure Clustering
- Large-scale autonomous rollouts can produce millions of failures.
- Reviewing each individually is inefficient.
-
Embed each failure:
\[z_i = f_{\mathrm{failure}}( \tau_i )\] -
Cluster
\[\{ z_1,\ldots,z_N \} \rightarrow \{ C_1,\ldots,C_K \}\] -
Clusters may correspond to failure families such as
\[\text{grasp slip}, \quad \text{occlusion}, \quad \text{wrong object}, \quad \text{collision}, \quad \text{planning dead-end}\] - The data engine can then prioritize high-frequency or high-severity clusters.
- This changes evaluation from a scalar success rate into an actionable data-generation system.
The Data Flywheel
- The most important property of a mature Physical AI system is that deployment generates the information required for improvement.
-
Let policy version \(k\) be
\[\pi_k\] -
Deployment generates trajectories:
\[\mathcal{R}_k = \operatorname{Rollout}( \pi_k )\] -
Evaluation identifies failures:
\[\mathcal{F}_k = \operatorname{Evaluate}( \mathcal{R}_k )\] -
The data engine creates targeted training data:
\[\Delta\mathcal{D}_k = G( \mathcal{F}_k )\] -
Training produces
\[\pi_{k+1} = \operatorname{Train}( \pi_k, \mathcal{D}_k \cup \Delta\mathcal{D}_k )\] -
Then:
\[\boxed{ \pi_k \rightarrow \mathcal{R}_k \rightarrow \mathcal{F}_k \rightarrow \Delta\mathcal{D}_k \rightarrow \pi_{k+1}. }\] - This is the Physical AI data flywheel.
Data Engines Versus Datasets
-
A dataset is static:
\[\mathcal{D} = \text{constant}\] -
A data engine is adaptive:
\[\mathcal{D}_{k+1} = f( \mathcal{D}_k, \pi_k, E_k )\]- where \(E_k\) represents evaluation results.
- The distinction is fundamental.
-
A static dataset asks:
\[\text{How much data do we have?}\] -
A data engine asks:
\[\text{What data should we collect next?}\] - For Physical AI, the second question increasingly determines the rate of progress.
The Emerging Physical AI Data Stack
-
A mature Physical AI data system increasingly resembles:
\[\boxed{ \begin{array}{c} \text{Internet Images + Video}\\ +\\ \text{Human Demonstrations}\\ +\\ \text{Cross-Embodiment Robot Data}\\ +\\ \text{Simulation}\\ +\\ \text{Generative Synthetic Data}\\ +\\ \text{World-Model Rollouts}\\ +\\ \text{Autonomous Robot Experience}\\ \downarrow\\ \text{Standardization + Synchronization}\\ \downarrow\\ \text{Filtering + Deduplication}\\ \downarrow\\ \text{Normalization + Relabeling}\\ \downarrow\\ \text{Mixture Construction}\\ \downarrow\\ \text{Foundation Policy Training}\\ \downarrow\\ \text{Evaluation + Failure Mining}\\ \downarrow\\ \text{Targeted Data Generation}\\ \circlearrowleft \end{array} }\] -
The broader transition is
\[\text{manual dataset collection} \rightarrow \text{automated data flywheel}\] \[\text{single-robot data} \rightarrow \text{cross-embodiment data}\] \[\text{real-only experience} \rightarrow \text{real + simulated + generated experience}\]-
and
\[\text{random data growth} \rightarrow \text{failure-directed data acquisition}\]
-
-
Ultimately, scaling Physical AI requires scaling useful experience rather than merely scaling parameters:
\[\boxed{ \text{Physical AI capability} \approx f( \text{model scale}, \text{experience scale}, \text{experience diversity}, \text{feedback quality} ). }\] -
The next section will examine Training and Post-Training Physical AI Models, including foundation-model pretraining, VLA supervised fine-tuning, behavioral cloning, offline and online reinforcement learning, reward construction, closed-loop RL, synthetic-data post-training, distillation, and the transition from broad pretrained models to deployment-ready physical policies.
Training and Post-Training Physical AI Models
From Foundation Model to Physical Policy
-
A modern Physical AI model is rarely trained for one downstream robot task from scratch. The emerging recipe is closer to the foundation-model pipeline used for language and multimodal models:
\[\boxed{ \text{Broad Pretraining} \rightarrow \text{Robot Pretraining} \rightarrow \text{Supervised Post-Training} \rightarrow \text{RL Post-Training} \rightarrow \text{Deployment Adaptation}. }\] - The purpose of each stage is different. Broad pretraining supplies semantic and visual knowledge. Robot pretraining grounds that knowledge in physical interaction. Supervised post-training adapts the policy to particular embodiments and tasks. Reinforcement learning improves behavior according to outcomes rather than demonstrations alone. Deployment adaptation closes remaining gaps involving latency, robustness, hardware dynamics, and the real operating distribution.
- This distinction is increasingly important because a foundation policy is not necessarily a deployment-ready policy. For example, GR00T N1.5 is pretrained on a heterogeneous mixture and then post-trained on embodiment- and task-specific datasets, while Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success by Kim et al. (2025) shows that adaptation choices such as continuous actions, action chunking, and parallel decoding can materially change the performance of an already pretrained VLA.
The Physical AI Training Hierarchy
- The training process can be viewed as progressively narrowing the model’s distribution.
-
Let the initial foundation model be
\[\theta_0\] -
Broad pretraining produces
\[\theta_{\mathrm{foundation}} = \operatorname{Train} ( \theta_0, \mathcal{D}_{\mathrm{broad}} )\] -
Robot pretraining then gives
\[\theta_{\mathrm{robot}} = \operatorname{Train} ( \theta_{\mathrm{foundation}}, \mathcal{D}_{\mathrm{robot}} )\] -
Task-specific supervised post-training produces
\[\theta_{\mathrm{SFT}} = \operatorname{Train} ( \theta_{\mathrm{robot}}, \mathcal{D}_{\mathrm{target}} )\] -
Finally, outcome-based optimization produces
\[\theta_{\mathrm{RL}} = \operatorname{RL} ( \theta_{\mathrm{SFT}}, R )\] -
The progression can therefore be interpreted as
\[\boxed{ \text{knowledge} \rightarrow \text{physical grounding} \rightarrow \text{task competence} \rightarrow \text{behavioral optimization}. }\]
Broad Multimodal Pretraining
- Many VLAs inherit their visual and linguistic representations from pretrained VLMs.
-
Suppose a multimodal model learns
\[p_\theta( y \mid I,l )\]- where \(I\) represents images and \(l\) represents language.
-
This stage can teach concepts such as
\[\text{objects}, \quad \text{attributes}, \quad \text{spatial relations}, \quad \text{instructions}, \quad \text{semantic associations}\] - But it does not yet teach the model what motor command should follow.
-
Robot training adds the physical mapping
\[(I,l,s) \rightarrow a\] - This division explains why VLAs can inherit semantic generalization from large VLMs while requiring comparatively smaller quantities of expensive action-labeled robot data.
Robot Foundation-Model Pretraining
- Robot pretraining broadens the policy across tasks, environments, and embodiments before specialization.
-
The dataset can be represented as
\[\mathcal{D}_{\mathrm{pre}} = \bigcup_{e=1}^{E} \bigcup_{k=1}^{K} \mathcal{D}_{e,k}\]- where \(e\) indexes embodiments and \(k\) indexes tasks or datasets.
-
The policy learns
\[\pi_\theta( a_{t:t+H} \mid o_{\leq t}, s_{\leq t}, l, e )\] - Open X-Embodiment: Robotic Learning Datasets and RT-X Models by the Open X-Embodiment Collaboration et al. (2023) established the importance of large cross-embodiment mixtures for generalist robot learning, while RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation by Liu et al. (2024) uses heterogeneous multi-robot data and a physically interpretable unified action representation to pretrain a large diffusion-based manipulation model.
- The objective is not necessarily to master every downstream task during pretraining. Instead, pretraining should create an initialization from which new physical capabilities can be acquired efficiently.
Pretraining Versus Post-Training
- The boundary between pretraining and post-training is operational rather than mathematical.
-
Pretraining generally optimizes broad coverage:
\[\mathcal{D}_{\mathrm{pre}} = \text{large + heterogeneous}\] -
Post-training generally optimizes target behavior:
\[\mathcal{D}_{\mathrm{post}} = \text{smaller + targeted}\] -
Thus,
\[\boxed{ \text{pretraining learns transferable priors;} \qquad \text{post-training converts them into reliable behavior}. }\] - A robot foundation model may know how cups, drawers, tables, and grippers interact but still require post-training to operate a particular robot arm reliably.
Supervised Fine-Tuning
- The simplest post-training stage is supervised fine-tuning on demonstrations.
-
Given
\[\mathcal{D}_{\mathrm{SFT}} = \{ (o_i,l_i,a_i) \}_{i=1}^{N}\]-
an autoregressive action policy minimizes
\[\mathcal{L}_{\mathrm{SFT}} = - \sum_i \log \pi_\theta( a_i \mid o_i,l_i )\]
-
-
For continuous regression,
\[\mathcal{L}_{\mathrm{SFT}} = \mathbb{E} [ \| a-\hat a_\theta \|_1 ]\]-
or
\[\mathcal{L}_{\mathrm{SFT}} = \mathbb{E} [ \| a-\hat a_\theta \|_2^2 ]\]
-
- For diffusion or flow-based policies, the supervised target is instead a denoising or vector-field objective.
-
Regardless of parameterization, the underlying signal remains:
\[\boxed{ \text{imitate demonstrated behavior}. }\]
Fine-Tuning Action Tokens
- An autoregressive VLA can quantize continuous actions into tokens.
-
For action dimension \(a_i\),
\[q_i = Q(a_i) \in \{ 1,\ldots,B \}\] -
The model predicts
\[p_\theta( q_i \mid o,l,q_{<i} )\] -
Training uses cross-entropy:
\[\mathcal{L}_{\mathrm{token}} = - \sum_i \log p_\theta( q_i^* \mid o,l,q_{<i}^* )\] - This approach allows an existing language-model decoder to produce actions using essentially the same mechanism used for text.
- However, sequentially generating many action tokens can create unnecessary inference latency for high-frequency control.
Continuous-Action Fine-Tuning
- An alternative is to attach a continuous action head.
-
The VLM produces hidden representation
\[h = f_{\mathrm{VLM}}( I,l )\]-
and an action head predicts
\[\hat A = f_{\mathrm{action}}( h,s )\]
-
-
The output may be an entire action chunk:
\[\hat A = [ \hat a_t, \ldots, \hat a_{t+H-1} ]\] - Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success by Kim et al. (2025) develops the OpenVLA-OFT recipe around parallel decoding, action chunking, continuous action representations, and an L1 regression objective; the paper reports both substantially higher action-generation throughput and improved downstream performance compared with the original OpenVLA fine-tuning setup.
-
This illustrates an important principle:
\[\boxed{ \text{post-training architecture} \text{ can matter as much as } \text{pretrained representation quality}. }\]
Action Chunking During Post-Training
-
Rather than predict only
\[a_t\]-
the model predicts
\[A_t = [ a_t, a_{t+1}, \ldots, a_{t+H-1} ]\]
-
- This provides temporal coherence and amortizes expensive foundation-model inference over multiple control steps.
-
The supervised objective becomes
\[\mathcal{L}_{\mathrm{chunk}} = \sum_{k=0}^{H-1} \ell( \hat a_{t+k}, a_{t+k} )\] -
At deployment, the system can execute the full chunk or only its first
\(K\) actions:
\[K\leq H\] - Executing fewer actions before replanning increases feedback responsiveness, while executing longer chunks reduces inference cost.
Fine-Tuning Flow-Matching Policies
- For models such as π0 and π0.5, continuous actions can be generated through flow matching.
-
Let clean action chunk be
\[A_1\]-
and noise be
\[A_0 \sim \mathcal{N}(0,I)\]
-
-
Interpolate:
\[A_\tau = (1-\tau)A_0 + \tau A_1\] -
The target velocity is
\[u_\tau = A_1-A_0\] -
The model learns
\[v_\theta( A_\tau, \tau, c ) \approx u_\tau\]-
with
\[\mathcal{L}_{\mathrm{FM}} = \mathbb{E} \left[ \| v_\theta(A_\tau,\tau,c) - u_\tau \|_2^2 \right]\]
-
- Here \(c\) contains visual observations, language, proprioception, and other conditioning information.
-
At inference, actions are generated by integrating
\[\frac{dA_\tau}{d\tau} = v_\theta( A_\tau,\tau,c )\]- π0.5 additionally emphasizes heterogeneous co-training across robot data, semantic prediction, web data, and other multimodal examples to improve open-world generalization.
Which Parameters Should Be Fine-Tuned?
-
A VLA may contain
\[\theta = \{ \theta_{\mathrm{vision}}, \theta_{\mathrm{language}}, \theta_{\mathrm{fusion}}, \theta_{\mathrm{action}} \}\] - Several adaptation strategies are possible.
-
Full fine-tuning updates
\[\nabla_\theta \mathcal{L}\]- for all parameters.
-
Partial fine-tuning may freeze the backbone:
\[\nabla_{\theta_{\mathrm{VLM}}} \mathcal{L} = 0\]-
while updating
\[\theta_{\mathrm{action}}\]
-
- Parameter-efficient adaptation may update low-rank adapters or other small trainable modules.
- The correct choice depends on the domain shift. If semantics and visual perception transfer well but robot dynamics change, updating the action pathway may be sufficient. If the new environment introduces unfamiliar visual concepts, broader adaptation may be required.
- NVIDIA reports that GR00T N1.5 freezes its VLM during both pretraining and fine-tuning while training the policy components around it, illustrating one practical decomposition between semantic representation and action learning.
Post-Training a New Embodiment
-
Suppose foundation policy
\[\pi_{\theta_0}\]- was pretrained across several robots.
-
A new robot has embodiment
\[e_{\mathrm{new}}\] -
Collect a relatively small adaptation dataset:
\[\mathcal{D}_{\mathrm{new}} = \{ (o_i,a_i,l_i) \}_{i=1}^{N}\] -
Post-training gives
\[\theta^* = \arg\min_\theta \mathcal{L} ( \theta; \mathcal{D}_{\mathrm{new}} )\] -
The intended benefit of foundation pretraining is
\[N_{\mathrm{foundation\rightarrow new}} \ll N_{\mathrm{scratch\rightarrow new}}\] -
GR00T N1.5 demonstrates this pattern through post-training on Unitree G1 teleoperation data, while its data-limited experiments also evaluate adaptation using small demonstration sets.
Post-Training for New Skills
- Embodiment adaptation and skill adaptation are distinct.
-
The robot may already understand its own morphology but lack task
\[T_{\mathrm{new}}\] -
The simplest solution is collect demonstrations:
\[\mathcal{D}_{T_{\mathrm{new}}}\]- and continue supervised training.
- Synthetic trajectories can also be used. GR00T N1.5 incorporates DreamGen-generated neural trajectories into its broader training pipeline and reports learning behaviors for which no corresponding teleoperation demonstrations were collected.
-
Thus skill acquisition can increasingly follow
\[\boxed{ \text{Describe Skill} \rightarrow \text{Generate Experience} \rightarrow \text{Post-Train Policy}. }\]
Why Behavioral Cloning Eventually Saturates
-
Behavioral cloning optimizes agreement with demonstrations:
\[\pi_\theta \approx \pi_E\] - But this imposes several limitations.
- First, the policy cannot systematically outperform the behavior represented in the demonstrations under the same objective.
-
Second, demonstrations provide sparse information about alternatives. They tell the policy
\[\text{do this}\]-
but generally not
\[\text{this alternative is slightly worse}\]-
or
\[\text{this action causes failure three seconds later}\]
-
-
-
Third, behavioral cloning primarily sees expert states:
\[s \sim d^{\pi_E}\]-
while deployment produces
\[s \sim d^{\pi_\theta}\]
-
- These distributions diverge when the policy makes mistakes.
- Outcome-based post-training addresses these limitations.
From Imitation to Reinforcement Learning
-
Reinforcement learning optimizes expected return:
\[J(\theta) = \mathbb{E}_{\tau\sim\pi_\theta} \left[ \sum_{t=0}^{T} \gamma^t r_t \right]\] - The crucial difference is the source of supervision.
-
Imitation learning uses:
\[\boxed{ \text{What did the expert do?} }\] -
RL uses:
\[\boxed{ \text{What happened when the policy acted?} }\] -
This distinction is particularly valuable in Physical AI because many important properties are naturally outcome-based:
\[\text{task success}, \quad \text{collision}, \quad \text{object damage}, \quad \text{energy}, \quad \text{comfort}, \quad \text{time}\]
Offline RL Post-Training
- Real robot interaction is expensive, so a natural intermediate stage is offline RL.
-
Given fixed dataset
\[\mathcal{D} = \{ (s_t,a_t,r_t,s_{t+1}) \}\]- the goal is to improve the policy without additional environment interaction.
-
A critic learns
\[Q_\phi(s_t,a_t) \approx r_t + \gamma V(s_{t+1})\] -
The policy can then favor high-value demonstrated actions:
\[\pi_\theta \leftarrow \operatorname{Improve}( Q_\phi, \mathcal{D} )\] - The central difficulty is extrapolation. If the learned policy selects actions outside the dataset distribution, the critic may assign unreliable values.
- This motivates conservative algorithms such as Conservative Q-Learning for Offline Reinforcement Learning by Kumar et al. (2020) and in-distribution methods such as Offline Reinforcement Learning with Implicit Q-Learning by Kostrikov et al. (2021).
Advantage-Weighted Post-Training
- One useful bridge between imitation and RL is to reweight demonstrations according to estimated advantage.
-
Define
\[A(s,a) = Q(s,a)-V(s)\] -
Then optimize
\[\mathcal{L}_{\mathrm{AW}} = - \mathbb{E}_{(s,a)\sim\mathcal{D}} \left[ w(s,a) \log \pi_\theta(a\mid s) \right]\]-
where
\[w(s,a) = \exp \left( \frac{A(s,a)}{\beta} \right)\]
-
- High-value demonstrated actions receive more weight.
- This is attractive for large physical policies because the final update resembles supervised learning even though the weights come from reinforcement-learning estimates.
Online RL Post-Training
- Online RL allows the policy to execute actions, observe consequences, and update using new experience.
-
At iteration \(k\):
\[\tau_k \sim \pi_{\theta_k}\]-
then
\[R_k = R(\tau_k)\]-
followed by
\[\theta_{k+1} = \operatorname{RLUpdate}( \theta_k, \tau_k, R_k )\]
-
-
-
The loop is
\[\boxed{ \text{Act} \rightarrow \text{Observe Outcome} \rightarrow \text{Update} \rightarrow \text{Act Again}. }\] - This solves a fundamental weakness of offline learning: the training distribution evolves with the policy.
- But real-world online RL can be expensive and unsafe, motivating simulation-heavy or hybrid post-training.
RL Post-Training for VLAs
- The same foundation-model principle used in language-model RL post-training can be applied to VLAs.
-
Start with a pretrained physical policy:
\[\pi_{\theta_0}\] -
Generate rollouts:
\[\tau_i \sim \pi_{\theta_0}\] -
Score them:
\[R_i = R( \tau_i )\] -
Then optimize:
\[\theta^* = \arg\max_\theta \mathbb{E}_{\tau\sim\pi_\theta} [ R(\tau) ]\] - Recent work such as RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models by Zhang et al. (2025) explicitly studies online RL post-training of pretrained VLAs and augments reward optimization with regularization for observation and action perturbations.
- The broader pattern is that pretraining supplies competence while RL targets behavioral properties that demonstration likelihood alone does not directly optimize.
Reward Construction
- Physical RL requires converting desired behavior into a reward.
-
A generic reward may be
\[R = w_sR_{\mathrm{success}} + w_pR_{\mathrm{progress}} + w_cR_{\mathrm{collision}} + w_mR_{\mathrm{motion}} + w_eR_{\mathrm{efficiency}}\] -
For manipulation,
\[R_{\mathrm{success}} = \mathbb{1}[ \text{task completed} ]\] -
A collision penalty might be
\[R_{\mathrm{collision}} = - \sum_t \mathbb{1}[ \text{unsafe contact}_t ]\] -
Smoothness may be encouraged using
\[R_{\mathrm{motion}} = - \sum_t \| a_t-a_{t-1} \|_2^2\] - The difficulty is that reward design changes the behavior being optimized.
-
Therefore,
\[\boxed{ \text{reward specification} \approx \text{behavior specification}. }\]
Sparse Rewards
-
The cleanest reward is often task success:
\[R(\tau) = \begin{cases} 1 & \text{task succeeds}\\ 0 & \text{otherwise}. \end{cases}\] - Sparse rewards avoid encoding unwanted intermediate strategies.
- However, they provide little learning signal when success is rare.
-
If
\[P_{\pi}( \text{success} ) \approx 0\]- almost every rollout receives identical reward.
- A strong pretrained policy changes this regime. If supervised pretraining already provides nontrivial success probability, RL can refine behavior from meaningful successful and failed rollouts.
-
This is one reason
\[\boxed{ \text{SFT} \rightarrow \text{RL} }\]- is generally more practical than learning complex physical behavior from random initialization.
Dense Rewards
- Dense rewards provide intermediate feedback.
-
For a reaching task:
\[r_t = - \| p_t^{\mathrm{gripper}} - p^{\mathrm{target}} \|_2\] -
For object lifting:
\[r_t = \alpha r_{\mathrm{reach}} + \beta r_{\mathrm{grasp}} + \gamma r_{\mathrm{lift}}\] - Dense rewards improve exploration but can create reward hacking.
- For example, optimizing distance to an object does not necessarily imply a stable grasp.
- Therefore, dense shaping is often combined with sparse terminal success.
Learned Reward Models
- Some physical objectives are difficult to encode manually.
-
A learned reward model can estimate
\[R_\psi( \tau,l )\] -
Training data might contain preference comparisons:
\[\tau_A \succ \tau_B\] -
A Bradley-Terry-style preference model uses
\[P( \tau_A\succ\tau_B ) = \frac{ \exp(R_\psi(\tau_A)) }{ \exp(R_\psi(\tau_A)) + \exp(R_\psi(\tau_B)) }\] - The reward model can then score new policy rollouts.
- Multimodal foundation models also make it possible to evaluate task completion from images or videos using semantic criteria rather than manually engineered state predicates.
Success Models as Rewards
-
For many robot tasks, the simplest learned reward is a success classifier:
\[V_\psi( o_{1:T},l ) = P( \text{success} \mid o_{1:T},l )\] -
The policy reward becomes
\[R(\tau) = V_\psi( \tau,l )\] - This makes reward construction scalable across tasks expressed in language.
- However, the policy can exploit errors in the success model.
-
Thus,
\[\boxed{ \text{reward-model robustness} }\]- becomes part of the safety problem.
Reinforcement Learning in Simulation
- Simulation provides a safer environment for online exploration.
-
The policy interacts with
\[s_{t+1} = F_{\mathrm{sim}}( s_t,a_t )\]- instead of the physical robot.
-
Many environments can run concurrently:
\[E_1,E_2,\ldots,E_N\] -
The total experience throughput becomes approximately
\[R_{\mathrm{experience}} \propto N f_{\mathrm{sim}}\]- where \(f_{\mathrm{sim}}\) is the effective simulation rate.
- This is especially important for humanoid locomotion, dexterous manipulation, and autonomous driving, where RL may require far more interaction than can reasonably be collected on hardware.
Sim-to-Real RL Post-Training
-
A practical pipeline is
\[\boxed{ \text{SFT on Real Data} \rightarrow \text{RL in Simulation} \rightarrow \text{Real-World Validation} \rightarrow \text{Limited Real RL}. }\] - Simulation teaches outcome-sensitive behavior at scale.
- Real-world data then corrects simulation mismatch.
-
The challenge is that
\[P_{\mathrm{sim}} \neq P_{\mathrm{real}}\] -
Domain randomization attempts to reduce this gap by optimizing
\[J(\theta) = \mathbb{E}_{\phi\sim p(\phi)} [ R( \pi_\theta; \phi ) ]\] - The resulting policy is trained across a distribution of physical parameters rather than a single simulated world.
World Models for RL Post-Training
- A learned world model can replace or complement a conventional simulator.
-
Given current state representation \(z_t\) and action sequence \(A\),
\[\hat\tau = M_\phi( z_t,A )\] -
The policy can train against imagined experience:
\[\tau \sim M_\phi( \pi_\theta )\] -
This creates
\[\boxed{ \text{VLA} \rightarrow \text{World-Model Rollout} \rightarrow \text{Reward} \rightarrow \text{RL Update}. }\] - The advantage is scalability.
- The danger is model exploitation: the policy may discover behaviors that produce high predicted reward because the world model is inaccurate rather than because the behavior succeeds physically.
- Consequently, imagined RL should remain coupled to real or high-fidelity simulated validation.
Synthetic-Data Post-Training
- Not all post-training requires RL.
-
Suppose evaluation identifies failure family
\[F\] -
A simulator or generative model produces targeted demonstrations:
\[\mathcal{D}_F = G(F)\] -
The model is then supervised-fine-tuned:
\[\theta' = \arg\min_\theta \mathcal{L}_{\mathrm{SFT}} ( \mathcal{D} \cup \mathcal{D}_F )\] - This is often simpler and more stable than RL when a reliable expert can generate the correct behavior.
-
The choice is therefore:
\[\boxed{ \begin{array}{ll} \text{correct action available} & \rightarrow \text{SFT} \\ \text{only outcome available} & \rightarrow \text{RL}. \end{array} }\] - In practice, strong Physical AI training systems use both.
Joint Policy and World-Model Objectives
- Policy learning and predictive representation learning can also be combined.
- GR00T N1.5 adds Future LAtent Representation Alignment, or FLARE, alongside its flow-matching action objective. Rather than generating future pixels, the auxiliary objective aligns predicted representations with future target embeddings; NVIDIA reports that this improves policy performance and enables learning from human egocentric video.
-
A simplified joint objective is
\[\mathcal{L} = \mathcal{L}_{\mathrm{action}} + \lambda \mathcal{L}_{\mathrm{future}}\] -
For example,
\[\mathcal{L}_{\mathrm{future}} = \| \hat z_{t+k} - \operatorname{sg}(z_{t+k}) \|_2^2\] - This encourages the policy representation to capture how the physical scene evolves, not merely which action appeared in the dataset.
Fine-Tuning World Models into Policies
- The boundary between world model and policy can also become less explicit.
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning by Kim et al. adapts a pretrained video world model into a visuomotor policy by jointly predicting future visual observations and robot actions. The resulting model treats future visual prediction as a useful intermediate representation for physical control.
-
Conceptually,
\[\text{video world model} \xrightarrow{\mathrm{robot\ post\text{-}training}} \text{visuomotor policy}\] - This suggests that world models and action policies may increasingly share pretrained representations rather than existing as completely separate systems.
Multi-Task Post-Training
-
Post-training on one task can cause specialization:
\[\pi_{\mathrm{general}} \rightarrow \pi_{\mathrm{task}}\] -
But aggressive specialization can cause forgetting:
\[P_{\mathrm{old}} \downarrow\] -
A multi-task mixture retains broad data:
\[\mathcal{D}_{\mathrm{post}} = \lambda_{\mathrm{target}} \mathcal{D}_{\mathrm{target}} + \lambda_{\mathrm{general}} \mathcal{D}_{\mathrm{general}}\] - The general replay component acts as regularization.
- This mirrors continual learning in foundation models: downstream adaptation should improve target capabilities without unnecessarily destroying transferable ones.
Catastrophic Forgetting
-
Let
\[L_{\mathrm{new}}\]-
measure the new task and
\[L_{\mathrm{old}}\]- measure retained capabilities.
-
-
Naive post-training optimizes
\[\min_\theta L_{\mathrm{new}}\] -
A retention-aware objective is
\[\min_\theta L_{\mathrm{new}} + \lambda L_{\mathrm{old}}\] - Other mechanisms include freezing parts of the backbone, replaying pretraining examples, using low-rank adapters, or regularizing parameter movement.
-
The desired update is:
\[\boxed{ \text{learn the new embodiment or skill} \quad \text{without unlearning the foundation}. }\]
Post-Training for Robustness
- Task success under nominal conditions is not sufficient.
-
A deployment policy must tolerate perturbations:
\[o_t' = o_t+\delta_o\] \[a_t' = a_t+\delta_a\]-
and changes in environment parameters:
\[\phi' = \phi+\delta_\phi\]
-
-
Robust post-training therefore optimizes
\[\max_\theta \mathbb{E}_{\delta} [ R( \pi_\theta; \delta ) ]\] - RobustVLA explicitly introduces Jacobian and smoothness regularization during RL post-training to reduce sensitivity to observation and action perturbations.
- Robustness is therefore not necessarily something that emerges automatically from larger pretraining. It can itself become a post-training objective.
Smoothness Regularization
- Physical actuators generally benefit from temporally coherent commands.
-
A simple smoothness penalty is
\[\mathcal{L}_{\mathrm{smooth}} = \sum_t \| a_t-a_{t-1} \|_2^2\] -
For acceleration-sensitive systems, one can penalize second differences:
\[\mathcal{L}_{\mathrm{jerk}} = \sum_t \| a_t - 2a_{t-1} + a_{t-2} \|_2^2\] -
The overall objective becomes
\[\mathcal{L} = \mathcal{L}_{\mathrm{task}} + \lambda_s \mathcal{L}_{\mathrm{smooth}}\] - This illustrates an important difference from purely digital AI: the geometry and temporal regularity of outputs directly affect physical performance.
Safety-Constrained RL
- Physical policy optimization may require explicit constraints.
-
Let reward be
\[R(\tau)\]-
and safety cost
\[C(\tau)\]
-
-
The objective becomes
\[\max_\pi \mathbb{E}[R(\tau)]\]-
subject to
\[\mathbb{E}[C(\tau)] \leq c_{\max}\]
-
-
A Lagrangian formulation is
\[J(\pi,\lambda) = \mathbb{E} [ R(\tau) - \lambda( C(\tau)-c_{\max} ) ]\] - Safety costs can represent collisions, excessive force, joint-limit violations, instability, or other deployment-specific constraints.
- For Physical AI, maximizing task success without modeling these constraints is generally an incomplete objective.
Curriculum Post-Training
- Post-training can progressively increase difficulty.
-
Let environment distribution at stage \(k\) be
\[p_k(E)\] -
Begin with
\[p_0(E) = \text{easy environments}\] -
As competence increases,
\[p_{k+1}(E) = \operatorname{Curriculum} ( p_k, \operatorname{Eval}(\pi_k) )\] -
For manipulation:
\[\text{isolated object} \rightarrow \text{clutter} \rightarrow \text{novel objects} \rightarrow \text{novel rooms} \rightarrow \text{dynamic interference}\] - The same principle applies to driving, locomotion, and navigation.
- Curriculum generation can be automated using the failure-mining data engine from the previous section.
Hard-Example Post-Training
-
Suppose evaluation produces failure set
\[\mathcal{F} = \{ \tau_i: \operatorname{Fail}(\tau_i)=1 \}\] -
Cluster failures:
\[\mathcal{F} \rightarrow \{ F_1,\ldots,F_K \}\] -
Generate additional examples:
\[\Delta\mathcal{D}_k \sim G( F_k )\] -
Then train:
\[\theta_{n+1} = \operatorname{Train} ( \theta_n, \mathcal{D} \cup \Delta\mathcal{D} )\] - This is analogous to hard-negative mining but operates over entire physical behaviors.
- The system concentrates training compute where evaluation indicates the policy is weakest.
Distillation
- The strongest training-time model may be too expensive for deployment.
-
Suppose teacher policy is
\[\pi_T\]-
and deployment policy is
\[\pi_S\]
-
-
The teacher generates targets:
\[A_T \sim \pi_T( o,l )\] -
The student minimizes
\[\mathcal{L}_{\mathrm{distill}} = D( \pi_S, \pi_T )\] -
For continuous actions,
\[\mathcal{L}_{\mathrm{distill}} = \| A_S-A_T \|_1\] - Distillation can transfer behavior from a large VLA or reasoning model into a smaller low-latency controller.
-
Thus model scaling and deployment scaling can be partially decoupled:
\[\boxed{ \text{large model for learning} \rightarrow \text{small model for acting}. }\]
Distilling Reasoning into Reactive Policies
- Physical systems often need both slow reasoning and fast reactions.
-
A teacher can perform expensive deliberation:
\[r_T = f_{\mathrm{reason}}( o,l )\]-
then produce action
\[a_T = f_{\mathrm{act}}( r_T,o )\]
-
-
The student directly learns
\[\pi_S( a\mid o,l )\] -
Training minimizes
\[D( a_S,a_T )\] - At deployment, the student does not need to regenerate the complete reasoning trace for routine control.
-
This produces a useful hierarchy:
\[\boxed{ \text{reason slowly during training} \rightarrow \text{act quickly during deployment}. }\] - More difficult or novel states can still be escalated to a slower reasoning system.
Deployment-Aware Post-Training
- A model that succeeds offline may fail when deployed because of latency.
-
Suppose inference takes
\[T_{\mathrm{infer}}\] - During this interval, the environment continues evolving.
-
The executed action is effectively conditioned on stale observation
\[o_{t-\Delta}\] -
Deployment-aware training can inject random delay:
\[\Delta \sim p(\Delta)\]-
and train
\[\pi_\theta( a_t \mid o_{t-\Delta} )\]
-
- Similarly, training can model dropped frames, actuator lag, camera noise, calibration errors, and communication delays.
- These are not peripheral systems details. For a closed-loop physical policy, they modify the effective environment dynamics.
Control Frequency and Model Latency
-
Suppose the desired control frequency is
\[f_c\] -
The control period is
\[T_c = \frac{1}{f_c}\] -
If
\[T_{\mathrm{infer}} > T_c\]- the model cannot synchronously produce every required action.
-
Solutions include:
\[\text{action chunking}\] \[\text{asynchronous inference}\] \[\text{policy distillation}\] \[\text{smaller action experts}\]-
and
\[\text{hierarchical control}\]
-
-
Post-training therefore needs to optimize not only success rate but behavior under the actual inference architecture.
Asynchronous Policy Execution
- An asynchronous architecture allows physical execution and foundation-model inference to overlap.
-
At time \(t\), the robot executes chunk
\[A_t\] -
Meanwhile, the model predicts
\[A_{t+1}\] -
Thus
\[\boxed{ \text{execution} \parallel \text{inference}. }\] - The challenge is that the future observation used for prediction may differ from the observation actually encountered.
- Training can compensate by exposing the policy to temporal offsets and perturbations.
- This converts inference scheduling into another distribution that the policy must generalize across.
Real-World Validation
- Simulation metrics are insufficient for final deployment.
-
A post-trained policy should be evaluated on physical hardware over distributions including
\[\text{known tasks}, \quad \text{novel objects}, \quad \text{novel scenes}, \quad \text{perturbations}, \quad \text{long-horizon tasks}\] -
For task family \(k\),
\[\operatorname{SR}_k = \frac{ N_{\mathrm{success},k} }{ N_{\mathrm{trials},k} }\] - But average success alone can hide catastrophic failure modes.
- Evaluation should therefore also measure quantities such as intervention rate, collision rate, recovery rate, completion time, action smoothness, and robustness under distribution shift.
- These metrics determine what the next post-training iteration should optimize.
GR00T as a Pretraining-to-Post-Training Example
- The evolution of GR00T illustrates the broader training pattern.
- GR00T N1.5 combines a VLM with a diffusion-transformer action model, heterogeneous robot and synthetic data, and a future-representation objective. NVIDIA reports pretraining across real, simulated, Open X-Embodiment, DreamGen, and other robot data, followed by targeted post-training for downstream embodiments and tasks.
- GR00T N1.6 continues this pattern with an updated Cosmos-based VLM and a larger action DiT; NVIDIA describes downstream real-robot experiments as using relatively small task-specific datasets for additional post-training after large-scale pretraining.
-
The conceptual recipe is therefore:
\[\boxed{ \text{large heterogeneous pretraining} \rightarrow \text{small targeted post-training}. }\]
The Physical AI Post-Training Flywheel
- The complete process can be written as an iterative loop.
-
Begin with pretrained policy
\[\pi_0\] -
Post-train on demonstrations:
\[\pi_1 = \operatorname{SFT}( \pi_0, \mathcal{D}_{\mathrm{demo}} )\] -
Execute rollouts:
\[\mathcal{R}_1 = \operatorname{Rollout}( \pi_1 )\] -
Evaluate:
\[\mathcal{F}_1 = \operatorname{Evaluate}( \mathcal{R}_1 )\] -
Generate targeted data:
\[\Delta\mathcal{D}_1 = G( \mathcal{F}_1 )\] -
Improve through supervised or reinforcement learning:
\[\pi_2 = \operatorname{PostTrain}( \pi_1, \Delta\mathcal{D}_1, R )\] -
Then repeat:
\[\boxed{ \text{Pretrain} \rightarrow \text{SFT} \rightarrow \text{Rollout} \rightarrow \text{Evaluate} \rightarrow \text{Generate Experience} \rightarrow \text{RL/SFT} \rightarrow \text{Deploy} \rightarrow \circlearrowleft }\]
The Emerging Training Stack
-
A mature Physical AI training system increasingly resembles:
\[\boxed{ \begin{array}{c} \text{Vision-Language / Video Pretraining}\\ \downarrow\\ \text{Cross-Embodiment Robot Pretraining}\\ \downarrow\\ \text{Action / Dynamics Learning}\\ \downarrow\\ \text{Embodiment-Specific SFT}\\ \downarrow\\ \text{Task-Specific SFT}\\ \downarrow\\ \text{Offline RL}\\ \downarrow\\ \text{Simulation / World-Model RL}\\ \downarrow\\ \text{Limited Real-World RL}\\ \downarrow\\ \text{Robustness + Safety Post-Training}\\ \downarrow\\ \text{Distillation + Deployment Optimization}\\ \downarrow\\ \text{Evaluation + Failure Mining}\\ \circlearrowleft \end{array} }\] - The important transition is from treating robot training as a single supervised-learning stage to treating it as a continuous post-training process.
-
The foundation model provides broad priors:
\[\text{What objects are present?}\] \[\text{What does the instruction mean?}\] \[\text{What physical behaviors are plausible?}\] -
Post-training converts those priors into reliable embodied behavior:
\[\text{Which action succeeds on this robot?}\] \[\text{Which behavior remains safe under perturbation?}\] \[\text{How should the policy recover after an error?}\] -
The resulting paradigm is:
\[\boxed{ \text{pretraining builds physical intelligence;} \qquad \text{post-training shapes physical behavior}. }\] - The next section will examine Autonomy and Agentic Physical AI, including reactive versus deliberative control, hierarchical planning, skill libraries, memory and world state, tool use, long-horizon execution, replanning, failure recovery, multi-agent interaction, and the transition from individual VLA policies to autonomous physical agents.
Autonomy and Agentic Physical AI
From Physical Policies to Physical Agents
-
A Vision-Language-Action model provides a mapping from observations and instructions to actions:
\[\pi_\theta \left( a_{t:t+H} \mid o_{\leq t},g \right)\] - This is sufficient for many short-horizon manipulation tasks, but autonomy requires more than generating locally appropriate actions. A physical agent operating for minutes or hours must interpret goals, maintain state, choose intermediate objectives, invoke specialized capabilities, monitor execution, recognize failures, and revise its plan.
-
The distinction can be summarized as:
\[\boxed{ \text{Physical Policy} + \text{Planning} + \text{Memory} + \text{Tools} + \text{Feedback} = \text{Physical Agent}. }\] - This transition parallels the development of agentic systems around language models, but the physical setting introduces an additional constraint: every decision changes the environment and may be difficult or impossible to reverse.
Reactive Versus Deliberative Control
-
The simplest physical policy is reactive:
\[a_t = \pi_\theta(o_t,g)\] - The current observation directly determines the next action.
- Reactive control is valuable when rapid responses are required. However, long-horizon tasks often require deliberation over future states.
-
A deliberative system instead reasons over a sequence:
\[P_t = [ g_1, g_2, \ldots, g_K ]\]- where each \(g_k\) is an intermediate goal.
-
Execution becomes
\[g \rightarrow P_t \rightarrow g_1 \rightarrow a_{1:H} \rightarrow g_2 \rightarrow a_{H+1:2H} \rightarrow \cdots\] -
The important architectural distinction is therefore:
\[\boxed{ \text{reactive control} = \text{observation}\rightarrow\text{action} }\]-
versus
\[\boxed{ \text{deliberative control} = \text{observation}\rightarrow\text{plan}\rightarrow\text{actions}. }\]
-
- A capable autonomous system generally needs both.
Hierarchical Autonomy
- Long-horizon behavior naturally decomposes into multiple temporal scales.
- Suppose the user asks:
-
Make me a cup of tea.
-
A high-level planner might produce:
\[\begin{aligned} g_1 &= \text{find mug},\\ g_2 &= \text{place mug near kettle},\\ g_3 &= \text{retrieve tea bag},\\ g_4 &= \text{place tea bag in mug},\\ g_5 &= \text{fill mug with hot water}. \end{aligned}\] - Each subgoal then invokes a lower-level controller.
-
Formally,
\[g_k = \pi_H( s_t,g )\]-
where \(\pi_H\) is the high-level policy, while
\[a_t = \pi_L( s_t,g_k )\]- is the low-level policy.
-
-
This gives:
\[\boxed{ \text{Task} \rightarrow \text{Subgoals} \rightarrow \text{Skills} \rightarrow \text{Actions}. }\] - The hierarchy allows expensive reasoning to operate at a slower timescale than motor control.
Temporal Abstraction
- Hierarchical control is closely related to temporal abstraction.
-
A primitive action may last milliseconds:
\[a_t\] -
A motor skill may last seconds:
\[\sigma_k = [ a_t,\ldots,a_{t+H} ]\] -
A semantic task may last minutes:
\[G = [ \sigma_1,\ldots,\sigma_K ]\] - The planner therefore does not need to reason over every actuator command.
-
Instead:
\[\boxed{ \text{Planner} \rightarrow \text{Skill} \rightarrow \text{Action Chunk} \rightarrow \text{Controller}. }\] - Reducing the planning horizon from thousands of motor commands to tens of semantic skills makes long-horizon reasoning substantially more tractable.
Skills as the Interface Between Reasoning and Control
- A skill can be represented as a callable function:
grasp(object)
place(object, location)
open(drawer)
navigate(location)
pour(container, target)
-
The high-level agent reasons over this discrete interface:
\[\sigma_t \in \mathcal{S}_{\mathrm{skills}}\] -
Each skill internally executes a continuous policy:
\[a_{t:t+H} \sim \pi_{\sigma_t}\] -
This architecture creates a clean abstraction boundary:
\[\boxed{ \text{semantic reasoning} \leftrightarrow \text{skill API} \leftrightarrow \text{continuous control}. }\] -
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances by Ahn et al. (2022), commonly known as SayCan, demonstrated this principle by combining language-model planning with pretrained robot skills whose value functions provide physical grounding.
SayCan: Combining Semantic and Physical Feasibility
-
A language model can estimate whether a skill is semantically appropriate:
\[P_{\mathrm{LM}} ( \sigma \mid g,h )\]- where \(g\) is the instruction and \(h\) is execution history.
- But semantic plausibility does not imply physical feasibility.
- A robot may understand that it should pick up a cup even when the cup is currently unreachable.
-
SayCan therefore combines the language-model score with an affordance or value score:
\[S(\sigma) \propto P_{\mathrm{LM}} ( \sigma\mid g,h ) \cdot V( s,\sigma )\] -
The selected skill is
\[\sigma^* = \arg\max_\sigma S(\sigma)\] -
This creates an important pattern for agentic Physical AI:
\[\boxed{ \text{What should I do?} \times \text{What can I do?} }\]- rather than allowing semantic reasoning alone to determine physical behavior.
Affordance-Grounded Planning
-
More generally, suppose a planner proposes candidate skill sequence
\[P = [ \sigma_1,\ldots,\sigma_K ]\] -
A feasibility model evaluates
\[F( \sigma_k,s_t ) = P( \text{success} \mid s_t,\sigma_k )\] -
The planner should optimize both semantic progress and physical feasibility:
\[P^* = \arg\max_P \left[ R_{\mathrm{goal}}(P) + \lambda \sum_k \log F( \sigma_k,s_k ) \right]\] - This is essential because an autonomous physical agent must plan over its actual embodiment.
- A humanoid, mobile manipulator, drone, and autonomous vehicle do not share the same feasible action set even if they share the same semantic understanding.
Language as an Intermediate Action Representation
- Hierarchies need not jump directly from high-level instructions to numerical controls.
- RT-H: Action Hierarchies Using Language by Belkhale et al. (2024) introduces intermediate “language motions,” such as “move arm forward,” between task-level language and robot actions. The architecture first predicts a language motion and then conditions action prediction on that motion and the task context.
-
The hierarchy can be represented as
\[g \rightarrow m_t \rightarrow a_t\]-
where
\[m_t = \text{language motion}\]
-
-
For example:
\[\text{``open the jar''} \rightarrow \text{``move hand toward lid''} \rightarrow a_t\] - This creates an interpretable intermediate interface and also enables human correction through language.
Human Intervention at the Semantic Layer
-
Suppose the policy predicts
\[m_t = \text{``move arm right''}\] -
A human can intervene with
\[m_t^* = \text{``move arm left''}\] - The low-level controller then converts that correction into continuous actions.
- RT-H shows how such language-level interventions can be incorporated into the action hierarchy and subsequently used as training data.
-
This is attractive because correcting
\[\text{``move left''}\]-
is often substantially easier for a human than manually specifying
\[[ \Delta x, \Delta y, \Delta z, \Delta r, g ]\]
-
- Agentic Physical AI therefore benefits from interfaces where humans can intervene at the same semantic level at which the agent plans.
Closed-Loop Reasoning
- A plan generated once at the beginning of a task is brittle.
-
Suppose the initial plan is
\[P_0 = [ g_1,g_2,g_3,g_4 ]\] -
After executing \(g_1\), the environment transitions:
\[s_0 \xrightarrow{g_1} s_1\] - The agent should not assume that \(s_1\) equals the expected state.
-
Instead it observes
\[o_1 = O(s_1)\]-
and updates its plan:
\[P_1 = \operatorname{Plan} ( o_1,g,h_1 )\]
-
-
Thus:
\[\boxed{ \text{Plan} \rightarrow \text{Act} \rightarrow \text{Observe} \rightarrow \text{Replan}. }\] - This is the physical analogue of an agentic reasoning loop.
Inner Monologue and Environment Feedback
- Inner Monologue: Embodied Reasoning through Planning with Language Models by Huang et al. (2022) studies how language-model planners can incorporate environment feedback, including success detection, scene descriptions, and human feedback, into subsequent planning decisions. The work shows that closed-loop language feedback improves high-level instruction completion across simulated and real robotic tasks.
-
The loop can be written as:
\[r_t = f_{\mathrm{reason}} ( g,o_t,h_t )\] \[\sigma_t = f_{\mathrm{plan}}( r_t )\] \[o_{t+1} = \operatorname{Execute}( \sigma_t )\]-
followed by
\[h_{t+1} = h_t \cup \{ \sigma_t,o_{t+1} \}\]
-
- The next reasoning step therefore conditions on what actually happened rather than only what was expected to happen.
Execution Monitoring
- A physical agent needs explicit mechanisms for determining whether an action succeeded.
-
For skill \(\sigma_t\), define termination model
\[T_\psi( o_{\leq t}, \sigma_t )\] -
It can predict:
\[y_t \in \{ \text{running}, \text{success}, \text{failure} \}\] -
For example, after executing
\[\sigma_t = \text{grasp(cup)}\]- the system should determine whether the cup is actually in the gripper.
-
The agent should therefore operate as:
\[\boxed{ \text{Execute} \rightarrow \text{Verify} \rightarrow \text{Continue or Recover}. }\] - Without verification, a single unnoticed error can corrupt every subsequent step in a long-horizon task.
Preconditions and Postconditions
- Skills can be represented with explicit preconditions and postconditions.
-
For skill
\[\sigma = \operatorname{open}(\text{drawer})\]-
a precondition might be
\[\operatorname{reachable}( \text{drawer handle} ) = 1\]
-
-
The desired postcondition is
\[\operatorname{open}( \text{drawer} ) = 1\] -
Execution becomes:
\[\text{check precondition} \rightarrow \text{execute skill} \rightarrow \text{check postcondition}\] - This provides a bridge between symbolic planning and learned physical policies.
- The symbolic layer specifies what should become true; the learned controller determines how to make it true.
State Estimation for Agentic Control
- An autonomous agent needs an internal representation of the environment.
-
Let
\[b_t = P( s_t \mid o_{\leq t},a_{<t} )\]- represent the agent’s belief over world state.
-
In practice, a structured state representation might contain
\[m_t = \{ \text{objects}, \text{poses}, \text{relations}, \text{task state}, \text{robot state}, \text{uncertainty} \}\] - For example:
mug:
location: counter
grasped: false
drawer:
state: open
contents: [tea_box]
kettle:
location: counter
state: heating
- This representation can persist even when objects temporarily leave the camera’s field of view.
Scene Memory
- A purely reactive model may forget an object as soon as it disappears from the current observation.
-
An autonomous agent instead maintains memory:
\[M_t = f( M_{t-1}, o_t, a_{t-1} )\] -
Memory can include:
\[\text{object identity}, \quad \text{last known location}, \quad \text{task progress}, \quad \text{failed attempts}, \quad \text{human preferences}\] -
Planning then conditions on
\[P_t = \operatorname{Plan} ( o_t,M_t,g )\] - Memory therefore transforms perception from a sequence of independent snapshots into a persistent model of the environment.
Episodic and Semantic Memory
- Physical agents can benefit from at least two forms of memory.
-
Semantic memory contains general knowledge:
\[M_{\mathrm{semantic}} = \text{``mugs are often stored in cabinets''}\] -
Episodic memory records specific experiences:
\[M_{\mathrm{episodic}} = \text{``the blue mug was placed in the left cabinet earlier''}\] -
The planning context becomes
\[c_t = [ o_t, M_{\mathrm{semantic}}, M_{\mathrm{episodic}}, g ]\] - This distinction becomes increasingly important for persistent agents operating repeatedly in the same homes, factories, vehicles, or workplaces.
Tool Use
- Physical agents can interact not only with the physical world but also with digital tools.
- A robot may call:
object_detector(image)
depth_estimator(image)
navigate_to(location)
lookup_inventory(item)
query_map(location)
open_door_controller()
ask_user(question)
-
The reasoning model selects tool
\[u_t \in \mathcal{U}\] -
The tool returns result
\[y_t = u_t(x_t)\] -
The agent incorporates it:
\[h_{t+1} = h_t \cup \{ u_t,y_t \}\] -
This produces:
\[\boxed{ \text{Reason} \rightarrow \text{Tool Call} \rightarrow \text{Observation} \rightarrow \text{Reason}. }\] -
The tool may itself be a robot policy.
Code as a Physical Action Language
- Code as Policies: Language Model Programs for Embodied Control by Liang et al. (2022) showed that code-generating language models can compose perception functions, control primitives, arithmetic, spatial reasoning, and external libraries into executable robot policies. The work demonstrated both waypoint-based and reactive control across multiple robot platforms.
- Instead of directly generating motor commands, the model can generate:
target = detect("red block")
goal = detect("blue bowl")
pick(target)
place(target, goal)
- More complex programs can contain loops and conditionals:
while not grasp_success():
adjust_grasp()
close_gripper()
- Code therefore provides a compositional intermediate representation between natural-language goals and physical execution.
Why Tool Use Matters for Physical AI
- A monolithic neural model need not learn every capability internally.
- Suppose the agent needs precise geometric computation.
-
Instead of approximating
\[x_{\mathrm{intersection}}\]- through language reasoning, it can invoke a geometry library.
- Similarly, it can call a motion planner for collision-free trajectories or a map service for navigation.
-
The architecture becomes
\[\boxed{ \text{Foundation Model} + \text{Specialized Tools} + \text{Physical Skills}. }\] - This can improve precision, interpretability, and modularity.
- The key requirement is that tool outputs ultimately remain grounded in the current physical state.
Spatial Reasoning and Value Maps
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models by Huang et al. (2023) combines language-model reasoning with vision-language grounding to construct 3D value maps that encode affordances and constraints; a model-based planner then uses those maps to synthesize closed-loop 6-DoF manipulation trajectories.
-
For workspace position
\[x\in\mathbb{R}^3\]-
one can define value map
\[V(x)\]
-
-
For example:
\[V_{\mathrm{affordance}}(x)\]-
may reward positions near a graspable handle, while
\[V_{\mathrm{collision}}(x)\]- penalizes obstacles.
-
-
The combined objective can be
\[V_{\mathrm{total}}(x) = \sum_i w_iV_i(x)\] -
Planning becomes
\[x^* = \arg\max_x V_{\mathrm{total}}(x)\] - This provides a useful example of semantic reasoning being translated into geometric constraints rather than directly into actuator commands.
Planning with a World Model
- An agent can also evaluate plans through learned simulation.
-
Suppose the planner proposes
\[P^{(k)} = [ \sigma_1^{(k)}, \ldots, \sigma_H^{(k)} ]\] -
A world model predicts
\[\hat\tau^{(k)} = M_\phi( s_t,P^{(k)} )\] -
An evaluator computes
\[J_k = R( \hat\tau^{(k)} )\] -
The selected plan is
\[P^* = \arg\max_k J_k\] -
This yields:
\[\boxed{ \text{Propose} \rightarrow \text{Imagine} \rightarrow \text{Evaluate} \rightarrow \text{Execute}. }\] - The architecture connects agentic reasoning with the world-model stack described earlier.
Receding-Horizon Agentic Planning
- Even if a planner generates a long plan, the system should rarely execute it blindly.
-
Suppose:
\[P_t = [ \sigma_t, \sigma_{t+1}, \ldots, \sigma_{t+K} ]\] -
Execute only:
\[\sigma_t\] -
Then observe:
\[o_{t+1}\] -
Replan:
\[P_{t+1} = \operatorname{Plan} ( o_{t+1},M_{t+1},g )\] - This is receding-horizon planning at the semantic level.
-
It mirrors model-predictive control:
\[\boxed{ \text{long-horizon reasoning} + \text{short-horizon commitment}. }\] - This is especially valuable in dynamic environments where people, objects, or other agents may invalidate a previously generated plan.
Failure Detection
-
A physical agent must distinguish several forms of failure:
\[F \in \{ F_{\mathrm{perception}}, F_{\mathrm{planning}}, F_{\mathrm{execution}}, F_{\mathrm{environment}}, F_{\mathrm{tool}} \}\] -
For example:
\[F_{\mathrm{perception}} = \text{wrong object identified}\] \[F_{\mathrm{planning}} = \text{invalid skill sequence}\] \[F_{\mathrm{execution}} = \text{grasp slipped}\] \[F_{\mathrm{environment}} = \text{human moved target}\] \[F_{\mathrm{tool}} = \text{motion planner returned no path}\] - The recovery strategy depends on the failure type.
- Treating every failure as “retry the same action” is insufficient for robust autonomy.
Recovery Policies
- Suppose action \(a_t\) fails.
-
A naive system repeats:
\[a_t \rightarrow a_t \rightarrow a_t\] -
A recovery-aware agent instead infers failure cause:
\[c_t = f_{\mathrm{diagnose}}( o_{\leq t},h_t )\] -
Then selects recovery:
\[r_t = \pi_{\mathrm{recover}}( c_t,s_t,g )\] -
Examples include:
\[\text{regrasp}, \quad \text{change viewpoint}, \quad \text{move obstacle}, \quad \text{choose another tool}, \quad \text{ask human}\] -
The loop becomes:
\[\boxed{ \text{Failure} \rightarrow \text{Diagnose} \rightarrow \text{Recover} \rightarrow \text{Resume}. }\] - Recovery behavior is one of the clearest differences between a demonstration-following policy and a genuinely autonomous agent.
Retry Versus Replan
- Not every failure requires global replanning.
-
Suppose skill
\[\sigma_t\]- fails.
-
If the state remains near the expected state,
\[d( s_t, \hat s_t ) < \epsilon\]- the agent may retry locally.
-
If the discrepancy is large,
\[d( s_t, \hat s_t ) \geq \epsilon\]- the high-level planner should replan.
-
Thus:
\[\boxed{ \text{small deviation} \rightarrow \text{local recovery}, }\] \[\boxed{ \text{large deviation} \rightarrow \text{global replanning}. }\] - A well-designed autonomy stack therefore contains recovery mechanisms at multiple temporal and semantic scales.
Uncertainty-Aware Autonomy
- The agent should know when its state estimate or action proposal is uncertain.
-
Let
\[U_t = U( o_t,a_t )\] -
If
\[U_t<\tau_1\]- the system acts normally.
-
If
\[\tau_1 \leq U_t < \tau_2\]- it may gather additional observations.
-
If
\[U_t \geq \tau_2\]- it may stop or request assistance.
-
This gives:
\[\boxed{ \text{Act} \rightarrow \text{Inspect} \rightarrow \text{Escalate} }\]- as uncertainty increases.
- The threshold should depend on consequence severity. Uncertainty that is tolerable when sorting soft objects may be unacceptable near a person or moving vehicle.
Active Perception
- Sometimes the correct action is to obtain more information.
- Suppose the robot cannot determine whether an object lies behind an obstacle.
-
Instead of immediately manipulating the scene, it may select sensing action
\[a_t^{\mathrm{observe}} = \text{move camera}\] -
The desired action maximizes information gain:
\[a^* = \arg\max_a \left[ H(b_t) - \mathbb{E} [ H(b_{t+1}) ] \right]\]- where \(H\) denotes entropy over the belief state.
-
Thus perception itself becomes part of planning:
\[\boxed{ \text{uncertainty} \rightarrow \text{information-gathering action} \rightarrow \text{better decision}. }\] - This is particularly important in partially observed environments.
Asking Humans as a Tool
- Human interaction is another form of active information gathering.
- Suppose the instruction is:
-
Put this in its usual place.
-
If several locations are plausible, the agent may estimate
\[P( g_i \mid o,l )\] - When uncertainty is high, it can ask:
-
Do you mean the left cabinet or the pantry?
-
The decision to ask can be formulated as value of information:
\[\operatorname{VOI}(q) = \mathbb{E} [ V \mid \text{answer to }q ] - V_{\mathrm{current}}\] -
Ask when
\[\operatorname{VOI}(q) > C_{\mathrm{interaction}}\]- where \(C_{\mathrm{interaction}}\) represents the cost of interrupting the user.
- A capable autonomous agent therefore does not always act. Sometimes its best action is to request clarification.
Gemini Robotics and Embodied Reasoning
- Google DeepMind introduced Gemini Robotics and Gemini Robotics-ER in March 2025 as Gemini 2.0-based models for robotics. Gemini Robotics extends multimodal reasoning into a VLA capable of directly controlling robots, while Gemini Robotics-ER focuses on embodied reasoning and can be connected to existing low-level robot controllers.
-
This distinction illustrates two complementary architectures:
\[\boxed{ \text{reasoning model} \rightarrow \text{robot APIs} }\]-
and
\[\boxed{ \text{integrated VLA} \rightarrow \text{robot actions}. }\]
-
- The first emphasizes modularity. The second learns more of the perception-to-action stack end to end.
- Future Physical AI systems are likely to combine both patterns.
Agentic VLA Architectures
- A VLA can itself become a component inside a broader agent.
-
Suppose:
\[A_t = \operatorname{VLA}( o_t,g_t )\]- generates low-level action chunks.
-
A higher-level agent chooses
\[g_t = \operatorname{Agent}( o_t,M_t,G )\] -
The complete hierarchy becomes:
\[\boxed{ G \xrightarrow{\mathrm{Agent}} g_t \xrightarrow{\mathrm{VLA}} A_t \xrightarrow{\mathrm{Controller}} \text{Actuators}. }\] - The VLA therefore acts as a learned motor tool for the agent.
- This architecture allows a single reasoning system to coordinate multiple physical skills without requiring it to directly generate every actuator command.
System 2 and System 1 in Physical Agents
- Physical autonomy naturally motivates a dual-timescale architecture.
-
System 2 performs:
\[\text{reasoning}, \quad \text{planning}, \quad \text{tool selection}, \quad \text{failure diagnosis}\] -
System 1 performs:
\[\text{motor control}, \quad \text{reflexes}, \quad \text{trajectory execution}\] -
Let System 2 operate every \(K\) control steps:
\[g_k = \pi_{\mathrm{slow}}( o_t,M_t,G )\] -
System 1 runs continuously:
\[a_t = \pi_{\mathrm{fast}}( o_t,g_k )\] - This is closely aligned with the architectural separation used by models such as GR00T N1, where semantic reasoning and high-frequency action generation are handled by distinct model components.
Event-Triggered Reasoning
- System 2 need not run at a fixed frequency.
-
Instead, deliberation can be triggered by events:
\[E_t \in \{ \text{new goal}, \text{skill completion}, \text{failure}, \text{high uncertainty}, \text{environment change} \}\] -
Then:
\[\text{invoke System 2} \iff E_t=1\] - Routine execution remains with the fast policy.
-
This reduces computational cost and latency:
\[\boxed{ \text{reason when necessary} \quad \text{react otherwise}. }\] - Such event-triggered architectures are particularly attractive when large multimodal reasoning models are much more expensive than low-level controllers.
Multi-Agent Physical AI
- Physical environments may contain multiple autonomous agents.
-
Let agent \(i\) have policy
\[\pi_i( a_t^i \mid o_t^i,M_t^i )\] -
The environment evolves according to the joint action:
\[s_{t+1} \sim P( s_{t+1} \mid s_t, a_t^1,\ldots,a_t^N )\] - Each agent must therefore reason about other agents’ behavior.
-
This occurs in:
\[\text{warehouse robots}, \quad \text{robot fleets}, \quad \text{autonomous vehicles}, \quad \text{human-robot collaboration}\] - The environment is no longer merely dynamic. It is strategic.
Coordination Between Robots
- Suppose two robots need to move an object.
-
The task can be decomposed:
\[G \rightarrow \{ g_A,g_B \}\] -
Each robot executes
\[a_t^A = \pi_A( o_t^A,g_A )\] \[a_t^B = \pi_B( o_t^B,g_B )\] -
But successful execution may require communication:
\[m_t^{A\rightarrow B}\] -
The policies become
\[a_t^B = \pi_B( o_t^B, g_B, m_t^{A\rightarrow B} )\] - Foundation-model reasoning introduces the possibility of using language itself as an inter-agent coordination protocol.
Human-Robot Interaction
- Humans are the most important other agents in many physical environments.
-
The robot must infer not only physical state but human intent:
\[b_t^{\mathrm{human}} = P( g_{\mathrm{human}} \mid o_{\leq t} )\] -
Actions should then account for this belief:
\[a_t = \pi( s_t, b_t^{\mathrm{human}}, g_{\mathrm{robot}} )\] - This matters for collaborative manipulation, home robotics, industrial systems, and autonomous vehicles.
- Human behavior is inherently uncertain, making probabilistic prediction and conservative planning especially important.
Interruptibility
- An autonomous physical agent should remain interruptible.
-
At any time, an external signal
\[I_t = 1\]- may override the current plan.
-
Execution changes from
\[a_t = \pi_\theta(o_t)\]-
to
\[a_t = a_{\mathrm{safe}}\]
-
-
Possible safe responses include:
\[\text{stop}, \quad \text{hold position}, \quad \text{release force}, \quad \text{move to safe state}\] - Interruptibility should exist below the high-level reasoning system so that emergency intervention does not depend on successful foundation-model inference.
Long-Horizon Error Accumulation
-
Suppose each skill succeeds independently with probability
\[p\] -
A task requiring \(N\) successful skills has approximate success probability
\[P_{\mathrm{task}} = p^N\] -
If
\[p=0.95\]-
and
\[N=20\]-
then
\[P_{\mathrm{task}} \approx 0.36\]
-
-
- Thus even individually strong skills can produce poor long-horizon autonomy.
-
This is why long-horizon systems require:
\[\boxed{ \text{high skill reliability} + \text{verification} + \text{replanning} + \text{recovery}. }\] - Improving single-step action accuracy alone does not solve autonomous task execution.
Progress Tracking
- An agent should explicitly track which parts of the goal have been completed.
-
Let task graph be
\[G = (V,E)\]- where nodes represent subgoals.
-
Maintain completion vector:
\[c_t \in \{0,1\}^{|V|}\] -
After each skill:
\[c_{t+1} = f( c_t,o_{t+1} )\] -
Planning conditions on remaining goals:
\[V_{\mathrm{remaining}} = \{ v_i \mid c_t^i=0 \}\] - This prevents the agent from repeatedly completing already satisfied subgoals and provides a compact representation for long-horizon progress.
Planning as Search Over Skills
-
Given skill library
\[\mathcal{S} = \{ \sigma_1,\ldots,\sigma_K \}\]- planning can be formulated as graph search.
- Each node represents state \(s\).
-
Each skill induces transition:
\[s' = T( s,\sigma )\] -
The planner seeks:
\[P^* = \arg\min_P C(P)\]-
subject to
\[s_{\mathrm{goal}} \in \operatorname{Reachable}( s_0,P )\]
-
-
A foundation model can provide heuristic
\[h_\theta( s,g )\]- that guides search toward semantically promising skills.
- This combines classical planning guarantees with learned semantic priors.
Neural Planning Versus Explicit Search
-
A fully neural planner predicts:
\[P \sim p_\theta( P \mid s,g )\] -
Explicit search evaluates candidate plans:
\[P^* = \operatorname{Search}( s,g,\mathcal{S} )\] -
Hybrid systems can use the neural model to propose candidates:
\[\{ P_1,\ldots,P_K \} \sim p_\theta\]- then verify them using symbolic constraints, value functions, simulation, or world models.
-
Thus:
\[\boxed{ \text{Foundation Model Proposals} + \text{Explicit Verification} }\]- can provide a stronger architecture than either mechanism alone.
Agentic Autonomy as Closed-Loop Optimization
-
The complete agent can be viewed as repeatedly solving:
\[a_t^* = \arg\max_a \mathbb{E} \left[ R( s_{t:T} ) \mid b_t,a \right]\] -
But practical systems factor this computation:
\[\boxed{ \begin{array}{c} \text{Goal}\\ \downarrow\\ \text{Reasoning}\\ \downarrow\\ \text{Task Plan}\\ \downarrow\\ \text{Skill Selection}\\ \downarrow\\ \text{VLA / Controller}\\ \downarrow\\ \text{Physical Action}\\ \downarrow\\ \text{Observation}\\ \downarrow\\ \text{State + Memory Update}\\ \downarrow\\ \text{Verification}\\ \downarrow\\ \text{Replan / Recover} \end{array} }\] -
Autonomy emerges from the loop rather than any single model invocation.
The Agentic Physical AI Stack
-
A mature Physical AI agent increasingly contains several interacting layers:
\[\boxed{ \begin{array}{c} \text{Human Goal / Mission}\\ \downarrow\\ \text{Multimodal Reasoning Model}\\ \downarrow\\ \text{Task Planner}\\ \downarrow\\ \text{Memory + World State}\\ \downarrow\\ \text{Tool / Skill Selection}\\ \downarrow\\ \text{VLA / Motion Policy}\\ \downarrow\\ \text{Low-Level Controller}\\ \downarrow\\ \text{Physical Environment}\\ \downarrow\\ \text{Perception + Feedback}\\ \downarrow\\ \text{Success / Failure Detection}\\ \circlearrowleft \end{array} }\] -
The architecture combines the major ideas developed throughout the primer:
\[\text{foundation models} + \text{VLAs} + \text{world models} + \text{data engines} + \text{post-training} + \text{agentic reasoning}\] -
The key conceptual transition is:
\[\boxed{ \text{predicting actions} \rightarrow \text{pursuing goals}. }\] -
A physical policy asks:
\[\text{What action should I take now?}\] -
An autonomous physical agent must continually ask:
\[\text{What am I trying to accomplish?}\] \[\text{What is true about the world now?}\] \[\text{What should happen next?}\] \[\text{Did it actually happen?}\] \[\text{If not, how should I recover?}\] -
The next section will examine Simulation, Sim-to-Real, and Digital Twins in detail, including physics simulation, GPU-parallel environments, domain randomization, system identification, differentiable simulation, neural reconstruction, digital twins, synthetic environments, hardware-in-the-loop evaluation, and the techniques used to transfer policies from simulated worlds into physical systems.
Simulation, Sim-to-Real, and Digital Twins
Why Simulation Is Central to Physical AI
- Physical AI requires interaction data, but real-world interaction is expensive, slow, difficult to parallelize, and potentially unsafe. Simulation provides an alternative environment in which robots can collect experience, learn policies, encounter failures, and undergo evaluation without every experiment requiring physical hardware.
-
A simulator approximates the transition dynamics of the real environment:
\[s_{t+1} = F_{\mathrm{real}}(s_t,a_t)\]-
with
\[\hat{s}_{t+1} = F_{\mathrm{sim}}(s_t,a_t;\phi)\]- where \(\phi\) contains parameters describing masses, friction, actuator dynamics, contacts, sensors, and other properties.
-
-
The central problem is therefore not whether simulation exactly reproduces reality. It is whether policies trained under
\[P_{\mathrm{sim}}(s_{t+1}\mid s_t,a_t)\]-
continue to work under
\[P_{\mathrm{real}}(s_{t+1}\mid s_t,a_t)\]
-
-
The discrepancy
\[P_{\mathrm{sim}} \neq P_{\mathrm{real}}\]- is commonly called the simulation-to-reality, or sim-to-real, gap. NVIDIA describes high-fidelity simulation, synthetic-data generation, software-in-the-loop testing, and hardware-in-the-loop testing as central capabilities of Isaac Sim, while Isaac Lab provides GPU-accelerated environments for large-scale robot learning.
The Role of a Robotics Simulator
-
A robotics simulator typically contains several coupled models:
\[\boxed{ \text{Geometry} + \text{Physics} + \text{Actuators} + \text{Sensors} + \text{Rendering} + \text{Environment Logic}. }\] - The geometry layer represents objects and robots.
- The physics engine computes motion and contact.
- The actuator model maps control commands into forces or joint targets.
- The sensor model generates camera, depth, lidar, force, proprioceptive, or other observations.
- The renderer produces visual observations.
- Environment logic defines task initialization, termination, rewards, and disturbances.
-
Together they implement
\[(s_t,a_t,\phi) \rightarrow (s_{t+1},o_{t+1},r_t)\] - For reinforcement learning, this interface effectively becomes a synthetic environment from which arbitrarily many trajectories can be sampled, subject to computational limits.
Physics Simulation
-
For a rigid-body system with generalized coordinates \(q\), a common dynamics formulation is
\[M(q)\ddot q + C(q,\dot q)\dot q + g(q) = \tau + J(q)^T\lambda\]- where:
- \(M(q)\) is the mass matrix.
- \(C(q,\dot q)\) represents velocity-dependent effects.
- \(g(q)\) represents gravity.
- \(\tau\) represents actuator forces or torques.
- \(J(q)^T\lambda\) represents contact forces.
-
The simulator numerically integrates these dynamics:
\[(q_t,\dot q_t) \xrightarrow{\tau_t} (q_{t+1},\dot q_{t+1})\] - Contact-rich robotics makes this particularly difficult because contacts introduce discontinuities, friction, impacts, and potentially deformable materials.
- Google DeepMind’s Opening up a physics simulator for robotics describes MuJoCo’s focus on efficient contact-rich simulation, one reason it became widely used in robot learning.
Physics Fidelity Versus Simulation Throughput
- Higher-fidelity simulation is not automatically better for learning.
-
Suppose simulator fidelity is
\[F\]-
and trajectory throughput is
\[Q\]
-
-
Increasing fidelity often increases computation:
\[F\uparrow \quad\Rightarrow\quad Q\downarrow\] -
A policy learner instead cares about useful experience per unit compute:
\[E_{\mathrm{useful}} = f(F,Q,D)\]- where \(D\) represents diversity.
- A simulator that produces extremely accurate trajectories slowly may be less useful for large-scale RL than a slightly less accurate simulator capable of producing far more varied experience.
-
The engineering objective is therefore:
\[\boxed{ \text{sufficient fidelity} \times \text{high throughput}. }\]
GPU-Parallel Simulation
- Robot reinforcement learning benefits heavily from parallel environments.
-
Instead of simulating one environment:
\[E_1\] -
GPU-oriented systems simulate
\[E_1,E_2,\ldots,E_N\]- simultaneously.
-
Each environment evolves independently:
\[s_{t+1}^{(i)} = F_{\mathrm{sim}} ( s_t^{(i)}, a_t^{(i)} )\] -
The policy evaluates a batch:
\[A_t = \pi_\theta( S_t )\]-
where
\[S_t = [ s_t^{(1)}, \ldots, s_t^{(N)} ]\]
-
- This transforms robot learning into a large batched computation that maps naturally onto accelerators.
- Isaac Lab is explicitly designed around GPU-accelerated parallelization for robot learning and supports reinforcement learning, imitation learning, motion planning, multiple physics engines, and large-scale headless execution.
Simulation for Reinforcement Learning
-
Consider a policy
\[a_t \sim \pi_\theta(a_t\mid s_t)\]-
with objective
\[J(\theta) = \mathbb{E} \left[ \sum_{t=0}^{T} \gamma^t r_t \right]\]
-
-
The difficulty is that estimating gradients or policy improvements can require many trajectories:
\[\tau_i = (s_0,a_0,r_0,\ldots,s_T)\] - Simulation permits these trajectories to be generated without repeatedly resetting a physical robot.
-
This is particularly useful for:
\[\text{locomotion}, \quad \text{dexterous manipulation}, \quad \text{navigation}, \quad \text{driving}, \quad \text{recovery behavior}\] - The policy can experience falls, collisions, failed grasps, and unusual states that would be undesirable or costly to produce deliberately on real hardware.
The Reality Gap
- The principal weakness of simulation is mismatch.
-
Let the real dynamics be parameterized by
\[\phi_{\mathrm{real}}\]-
and the simulator by
\[\phi_{\mathrm{sim}}\]
-
-
Then
\[F( s,a;\phi_{\mathrm{sim}} ) \neq F( s,a;\phi_{\mathrm{real}} )\] -
Mismatch can arise from:
\[\text{mass}, \quad \text{friction}, \quad \text{compliance}, \quad \text{motor dynamics}, \quad \text{latency}, \quad \text{sensor noise}, \quad \text{contact}, \quad \text{visual appearance}\] - A policy can exploit simulator-specific regularities that do not exist in reality.
-
Thus successful simulation training requires more than maximizing simulated reward:
\[\boxed{ \text{high simulated performance} \not\Rightarrow \text{high real-world performance}. }\]
Domain Randomization
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World by Tobin et al. (2017) established a simple and influential strategy: rather than make one simulated world perfectly realistic, randomize simulated worlds so broadly that the real world behaves like another sample from the training distribution.
-
Let simulator parameters be
\[\phi\] -
Instead of fixing
\[\phi=\phi_0\]-
sample
\[\phi \sim p(\phi)\]
-
-
Training becomes
\[\max_\theta \mathbb{E}_{\phi\sim p(\phi)} \left[ R( \pi_\theta;\phi ) \right]\] -
The intended condition is
\[\phi_{\mathrm{real}} \in \operatorname{support} ( p(\phi) )\] - The policy is therefore encouraged to learn behavior invariant to nuisance variation.
Visual Domain Randomization
-
For visual policies, randomizable parameters may include:
\[\phi_{\mathrm{visual}} = \{ \text{textures}, \text{lighting}, \text{camera pose}, \text{object colors}, \text{backgrounds}, \text{materials} \}\] -
Rendered image becomes
\[I_t = R( s_t; \phi_{\mathrm{visual}} )\] -
During training:
\[\phi_{\mathrm{visual}} \sim p( \phi_{\mathrm{visual}} )\] - If the task-relevant geometry remains consistent while superficial appearance varies, the perception system is pressured to learn more robust features.
- Tobin et al. demonstrated this principle by training an object-localization model entirely on randomized synthetic images and transferring it to real robotic perception and grasping.
Dynamics Randomization
- Visual randomization does not address physical mismatch.
-
For control policies, randomize dynamics:
\[\phi_{\mathrm{dyn}} = \{ m, \mu, I, k, c, \tau_{\mathrm{motor}}, \Delta t \}\]- where the parameters can represent mass, friction, inertia, stiffness, damping, motor behavior, and latency.
-
For each environment:
\[\phi_{\mathrm{dyn}}^{(i)} \sim p( \phi_{\mathrm{dyn}} )\] -
Then:
\[s_{t+1}^{(i)} = F( s_t^{(i)}, a_t^{(i)}; \phi_{\mathrm{dyn}}^{(i)} )\] - The policy cannot rely on one exact set of dynamics and must instead discover behavior that remains effective across a family of physical systems.
Sensor Randomization
- Sensor models introduce another source of mismatch.
-
A real observation may be
\[o_t^{\mathrm{real}} = h( s_t ) + \epsilon_t\] -
A simulator can approximate this with
\[o_t^{\mathrm{sim}} = h_{\mathrm{sim}}( s_t; \phi_{\mathrm{sensor}} ) + \epsilon_t^{\mathrm{sim}}\] -
Randomizable properties include:
\[\text{camera noise}, \quad \text{blur}, \quad \text{exposure}, \quad \text{depth artifacts}, \quad \text{calibration}, \quad \text{dropped frames}\] - The same principle applies to proprioception, lidar, radar, force sensing, and tactile sensing.
- A robust policy should tolerate realistic observation corruption rather than relying on perfect simulated sensors.
Latency Randomization
- Real systems contain delays.
-
Let
\[\Delta_o\]-
be observation latency and
\[\Delta_a\]- be actuation latency.
-
-
The executed action may therefore be
\[a_t = \pi_\theta( o_{t-\Delta_o} )\]-
while the actuator applies it at
\[t+\Delta_a\]
-
-
If simulation assumes
\[\Delta_o=\Delta_a=0\]- the learned policy may become unrealistically aggressive.
-
Instead:
\[\Delta_o \sim p(\Delta_o)\] \[\Delta_a \sim p(\Delta_a)\] - Latency therefore becomes another domain-randomization variable.
System Identification
- Domain randomization intentionally broadens simulation. System identification takes the complementary approach of making the simulator resemble the real system.
-
Suppose real trajectories are
\[\mathcal{D}_{\mathrm{real}} = \{ (s_t,a_t,s_{t+1}) \}\] -
Estimate simulator parameters:
\[\phi^* = \arg\min_\phi \sum_t d \left( F_{\mathrm{sim}}( s_t,a_t;\phi ), s_{t+1}^{\mathrm{real}} \right)\] -
Parameters might include:
\[\text{link masses}, \quad \text{friction}, \quad \text{motor gains}, \quad \text{joint damping}, \quad \text{latency}\] -
The calibrated simulator then becomes
\[F_{\mathrm{sim}}( \cdot;\phi^* )\] - System identification therefore tries to shrink the reality gap, while domain randomization tries to make the policy insensitive to the remaining gap.
System Identification and Domain Randomization Together
- The two approaches are complementary.
-
First estimate:
\[\phi^* \approx \phi_{\mathrm{real}}\] -
Then randomize around it:
\[\phi \sim p( \phi\mid\phi^* )\] -
For example:
\[m \sim \mathcal{N}( m^*, \sigma_m^2 )\] - The training distribution is therefore centered around the best available estimate of reality but remains broad enough to account for uncertainty.
-
This gives the practical recipe:
\[\boxed{ \text{calibrate} \rightarrow \text{randomize} \rightarrow \text{train} \rightarrow \text{validate} \rightarrow \text{recalibrate}. }\]
Adaptive Domain Randomization
- The randomization distribution itself can be learned.
-
Suppose
\[p_\psi(\phi)\]- parameterizes simulator variations.
-
Instead of choosing \(p_\psi\) manually, evaluation on real hardware provides feedback:
\[E_{\mathrm{real}} = \operatorname{Eval}( \pi_\theta )\] -
The system updates
\[\psi \leftarrow \operatorname{Adapt}( \psi,E_{\mathrm{real}} )\] - The goal is to place more simulation probability on variations that explain observed real-world failures.
- Thus sim-to-real becomes an iterative data-engine problem rather than a one-time transfer step.
Differentiable Simulation
-
Traditional simulators are used as black-box transition functions:
\[s_{t+1} = F( s_t,a_t;\phi )\] -
Differentiable simulators additionally expose derivatives such as
\[\frac{\partial s_{t+1}} {\partial a_t}\]-
and
\[\frac{\partial s_{t+1}} {\partial \phi}\]
-
- This permits gradients to propagate through physical dynamics.
-
For objective
\[J = R( s_T )\]-
one can compute
\[\frac{\partial J} {\partial a_t} = \frac{\partial J} {\partial s_T} \frac{\partial s_T} {\partial a_t}\]
-
- Applications include trajectory optimization, policy learning, system identification, and robot design.
- The Newton Physics Engine, developed by NVIDIA, Google DeepMind, and Disney Research and managed by the Linux Foundation, is built on NVIDIA Warp and includes differentiable-physics capabilities intended for robot learning, optimization, and system identification.
Newton and Modern Robot Simulation
- Newton represents a broader shift toward simulation infrastructure designed specifically for large-scale Physical AI.
-
Its stack combines:
\[\text{OpenUSD} + \text{Warp} + \text{physics solvers} + \text{robot-learning frameworks}\] - OpenUSD provides a composable representation of robots and environments.
- Warp provides GPU-accelerated computational kernels.
- Newton exposes physics solvers through a common framework.
- Robot-learning systems such as Isaac Lab or MuJoCo Playground can then use these simulation capabilities for training. NVIDIA describes Newton as open source, extensible, GPU accelerated, and compatible with both Isaac Lab and MuJoCo Playground.
- The direction is important because simulation is becoming part of the training infrastructure rather than a separate visualization tool.
Isaac Sim and Isaac Lab
- Isaac Sim and Isaac Lab serve related but distinct roles.
- Isaac Sim provides a robotics simulation framework built on Omniverse libraries, with physics, rendering, sensor simulation, synthetic-data generation, and testing infrastructure. It can ingest CAD, URDF, and real-world captures into OpenUSD-based scenes.
-
Isaac Lab is oriented toward robot learning:
\[\boxed{ \text{Isaac Sim} \rightarrow \text{simulation environment} }\]-
and
\[\boxed{ \text{Isaac Lab} \rightarrow \text{policy-learning framework}. }\]
-
- Isaac Lab supports reinforcement learning, imitation learning, motion planning, parallel environments, and integration with multiple physics engines.
Synthetic Data Generation
- Simulation can generate more than policy rollouts.
-
Because the simulator has access to ground-truth scene state, it can automatically produce labels such as:
\[\text{RGB}, \quad \text{depth}, \quad \text{segmentation}, \quad \text{bounding boxes}, \quad \text{poses}, \quad \text{optical flow}\] -
A synthetic-data generator samples:
\[z \sim p(z)\]-
constructs scene
\[S = G(z)\]-
and renders:
\[(I,Y) = R(S)\]- where \(Y\) contains automatically generated labels.
-
-
- NVIDIA’s Omniverse Replicator provides parameterized synthetic-data generation within Isaac Sim, including control over scene composition, lighting, camera configuration, object appearance, and other randomized properties.
Procedural Simulation
- Instead of manually building every training environment, scenes can be procedurally generated.
-
Let
\[z = [ z_{\mathrm{layout}}, z_{\mathrm{objects}}, z_{\mathrm{materials}}, z_{\mathrm{lighting}} ]\] -
Generate:
\[E = G(z)\] - Each environment can expose the policy to a different combination of physical conditions.
-
The number of possible environments can therefore greatly exceed the number explicitly authored by humans:
\[|\mathcal{E}_{\mathrm{generated}}| \gg |\mathcal{E}_{\mathrm{authored}}|\] - Procedural generation turns environmental diversity into a programmable training dimension.
From Procedural Generation to Generative Simulation
- Foundation generative models make it possible to move beyond manually parameterized procedural generation.
-
Instead of
\[E = G(z)\]-
the simulator can generate environments conditioned on language:
\[E \sim p_\phi( E\mid l )\]
-
-
For example:
\[l = \text{``small cluttered kitchen with a narrow passage''}\] -
The same mechanism can generate rare or adversarial scenarios:
\[l = \text{``partially occluded object near the edge of a table''}\] - This connects generative world models with conventional simulation.
- The simulator supplies controllable physics; the generative model supplies scalable environmental diversity.
Real-to-Sim
- Sim-to-real transfers learned behavior from simulation into reality.
-
The reverse process, real-to-sim, reconstructs a simulated environment from real-world observations:
\[\boxed{ \text{Real World} \rightarrow \text{Digital Representation} \rightarrow \text{Simulation}. }\] -
Inputs may include:
\[\text{images}, \quad \text{video}, \quad \text{depth}, \quad \text{lidar}, \quad \text{CAD}, \quad \text{robot logs}\] - The resulting scene can be used to reproduce failures, generate nearby scenarios, or evaluate candidate policies before returning them to hardware.
-
This closes a critical loop:
\[\boxed{ \text{Real} \rightarrow \text{Sim} \rightarrow \text{Train} \rightarrow \text{Real}. }\]
Digital Twins
- A digital twin is more than a visually similar 3D scene.
-
For Physical AI, a useful digital twin should maintain correspondence between a physical system and a computational representation:
\[\mathcal{T}_t = f( s_t^{\mathrm{real}} )\] -
As reality changes:
\[s_t^{\mathrm{real}} \rightarrow s_{t+1}^{\mathrm{real}}\]-
the twin should update:
\[\mathcal{T}_t \rightarrow \mathcal{T}_{t+1}\]
-
-
The twin can contain:
\[\text{geometry}, \quad \text{semantics}, \quad \text{physics}, \quad \text{robot state}, \quad \text{object state}, \quad \text{sensor models}\] - It therefore acts as a synchronized computational proxy for the physical environment.
Digital Twin Versus Simulator
- A simulator represents a class of possible worlds.
- A digital twin attempts to represent a particular world.
-
Thus:
\[\boxed{ \text{Simulator} = \text{possible environment} }\]-
while
\[\boxed{ \text{Digital Twin} = \text{specific physical environment}. }\]
-
- For example, a generic warehouse simulator may contain shelves and robots.
- A warehouse digital twin attempts to represent the actual warehouse layout, robot fleet, obstacles, and potentially live operational state.
- This makes digital twins especially valuable for deployment validation and counterfactual analysis.
Neural Reconstruction
- Traditional digital twins often rely on CAD geometry.
- But many real environments lack complete CAD models.
-
Neural reconstruction can infer scene representation directly from images or video:
\[\{ I_1,\ldots,I_N \} \rightarrow \mathcal{R}_{\mathrm{scene}}\] - Representations can include neural radiance fields, Gaussian splats, meshes, depth maps, or hybrid neural-geometric structures.
-
The goal is to reconstruct enough visual and geometric information to render:
\[\hat I = R( \mathcal{R}_{\mathrm{scene}}, C )\]- from novel camera viewpoint \(C\).
- For Physical AI, reconstruction becomes particularly useful when the representation can also be coupled to physics.
Gaussian Splatting for Robotic Digital Twins
-
Gaussian splatting represents a scene as a collection of parameterized Gaussian primitives:
\[\mathcal{G} = \{ G_i \}_{i=1}^{N}\] -
Each primitive can encode quantities such as:
\[G_i = ( \mu_i, \Sigma_i, c_i, \alpha_i )\]-
where \(\mu_i\) is position, \(\Sigma_i\) describes spatial extent,
\(c_i\) describes appearance, and \(\alpha_i\) controls opacity.
-
- Such representations can provide fast novel-view rendering while retaining spatial structure.
- GaussTwin: Unified Simulation and Correction with Gaussian Splatting for Robotic Digital Twins by Cai et al. (2026) combines Gaussian splatting with physically grounded simulation and online visual correction, demonstrating a direction in which digital twins remain synchronized with real robotic scenes rather than serving as static reconstructions.
Closing the Real-to-Sim Gap
- A reconstructed twin will not initially match reality perfectly.
-
Let simulated observation be
\[\hat o_t = R( \mathcal{T}_t )\]-
and real observation be
\[o_t\]
-
-
A correction objective can minimize:
\[\mathcal{L}_{\mathrm{twin}} = d( \hat o_t,o_t )\] -
The twin parameters update:
\[\phi_{t+1} = \phi_t - \eta \nabla_\phi \mathcal{L}_{\mathrm{twin}}\] -
This creates a prediction-correction loop:
\[\boxed{ \text{Simulate} \rightarrow \text{Observe Reality} \rightarrow \text{Measure Error} \rightarrow \text{Correct Twin}. }\] - A continuously updated twin can become substantially more useful for planning than a one-time scene reconstruction.
Counterfactual Simulation
- Once a digital twin exists, the system can evaluate actions that were never executed physically.
-
Given current twin state
\[\mathcal{T}_t\]-
evaluate candidate action sequences:
\[A^{(1)},A^{(2)},\ldots,A^{(K)}\]
-
-
Simulate:
\[\tau^{(k)} = F_{\mathrm{twin}} ( \mathcal{T}_t, A^{(k)} )\] -
Score:
\[J_k = R( \tau^{(k)} )\] -
Then select:
\[A^* = \arg\max_k J_k\] -
This allows the physical system to ask:
\[\boxed{ \text{What would happen if I did this?} }\]- before committing to an action in reality.
Digital Twins for Failure Replay
-
Suppose a real robot fails at state
\[s_f\] -
Instead of repeatedly recreating the failure physically, reconstruct or synchronize the corresponding twin:
\[s_f \rightarrow \mathcal{T}_f\] -
Then generate variations:
\[\mathcal{T}_f^{(1)}, \ldots, \mathcal{T}_f^{(N)}\] -
These may vary:
\[\text{object pose}, \quad \text{friction}, \quad \text{lighting}, \quad \text{robot initialization}, \quad \text{distractors}\] - Training on this neighborhood turns one real failure into a family of synthetic learning examples.
- This connects digital twins directly to the Physical AI data flywheel.
Software-in-the-Loop Testing
- Software-in-the-loop, or SIL, evaluates the actual autonomy software against a simulated environment.
-
Conceptually:
\[\boxed{ \text{Production Software} \leftrightarrow \text{Simulated World}. }\] -
The software receives simulated sensor inputs:
\[o_t^{\mathrm{sim}}\]-
and emits normal commands:
\[a_t\]
-
-
The simulator executes those commands:
\[s_{t+1} = F_{\mathrm{sim}}( s_t,a_t )\] - This permits testing perception, planning, policy inference, state management, and control integration without requiring physical hardware.
- Isaac Sim explicitly supports simulation-based testing and validation alongside synthetic-data generation.
Hardware-in-the-Loop Testing
- Hardware-in-the-loop, or HIL, introduces actual deployment hardware into the simulation loop.
-
For example:
\[\boxed{ \text{Simulated Sensors} \rightarrow \text{Real Compute Hardware} \rightarrow \text{Control Output} \rightarrow \text{Simulator}. }\] -
This allows engineers to measure effects that pure software simulation may miss:
\[\text{inference latency}, \quad \text{memory limits}, \quad \text{communication delay}, \quad \text{runtime scheduling}\] - A policy that works in an idealized software loop may behave differently when running on the actual edge computer.
- HIL therefore bridges algorithmic simulation and deployment validation.
Shadow-Mode Evaluation
- Another intermediate step is shadow deployment.
-
The physical system operates normally, but a candidate policy receives the same observations:
\[o_t^{\mathrm{real}}\] -
It predicts:
\[\hat a_t = \pi_{\mathrm{candidate}}( o_t )\]- without controlling the robot.
-
The system compares
\[\hat a_t\]- against executed actions and subsequent outcomes.
- This enables evaluation on the real observation distribution without allowing the candidate policy to influence the physical system.
- Shadow mode is particularly useful for identifying distribution shifts before live deployment.
Sim-to-Sim Transfer
- Policies may also need to transfer between simulators.
-
Suppose training occurs under
\[F_A\]-
but validation occurs under
\[F_B\]
-
-
If
\[\pi_A\]- performs well under both, this provides evidence that the policy has not simply exploited artifacts unique to simulator \(A\).
-
Thus:
\[\boxed{ \text{sim-to-sim transfer} }\]- can act as an intermediate robustness test before sim-to-real deployment.
- The increasing interoperability between Newton, Isaac Lab, and MuJoCo-style workflows supports this broader idea of testing learned behavior across simulation backends.
Sim-to-Real as Distribution Generalization
-
The transfer problem can be formalized as training under distribution
\[p_{\mathrm{train}}( s,o,\phi )\]-
and deploying under
\[p_{\mathrm{real}}( s,o,\phi )\]
-
-
The objective is therefore not simply:
\[R_{\mathrm{sim}}\uparrow\]-
but:
\[\mathbb{E}_{x\sim p_{\mathrm{real}}} [ R( \pi_\theta,x ) ] \uparrow\]
-
-
Domain randomization broadens
\[p_{\mathrm{train}}\] -
System identification moves it toward
\[p_{\mathrm{real}}\] - Real-world fine-tuning adapts the policy after transfer.
- The three mechanisms therefore attack the same problem from different directions.
A Practical Sim-to-Real Recipe
-
A robust training pipeline often resembles:
\[\boxed{ \begin{array}{c} \text{Build Robot + Environment Model}\\ \downarrow\\ \text{Calibrate Against Real Logs}\\ \downarrow\\ \text{Randomize Uncertain Parameters}\\ \downarrow\\ \text{Train at Scale in Simulation}\\ \downarrow\\ \text{Validate Across Simulator Variants}\\ \downarrow\\ \text{Software-in-the-Loop}\\ \downarrow\\ \text{Hardware-in-the-Loop}\\ \downarrow\\ \text{Shadow / Limited Real Deployment}\\ \downarrow\\ \text{Collect Transfer Failures}\\ \downarrow\\ \text{Update Simulation + Policy}\\ \circlearrowleft \end{array} }\] -
The central principle is that sim-to-real should be treated as a closed-loop engineering process rather than a single model export.
Sim-to-Real for Perception
-
For perception models, the main gap is often visual:
\[p_{\mathrm{sim}}(I) \neq p_{\mathrm{real}}(I)\] -
Useful techniques include:
\[\text{visual randomization}\] \[\text{photorealistic rendering}\] \[\text{real-image fine-tuning}\] \[\text{representation pretraining}\] \[\text{synthetic-real co-training}\] -
Because foundation vision models already provide strong real-world representations, modern Physical AI systems can often use simulation primarily to teach physical structure and rare conditions rather than relying on synthetic images as the sole source of visual knowledge.
Sim-to-Real for Control
- For control policies, dynamics mismatch dominates.
-
The relevant discrepancy is:
\[F_{\mathrm{sim}} \neq F_{\mathrm{real}}\] -
Useful mechanisms include:
\[\text{dynamics randomization}\] \[\text{system identification}\] \[\text{latency modeling}\] \[\text{actuator calibration}\] \[\text{real-world post-training}\] - Control transfer is especially difficult for contact-rich tasks because small errors in friction, compliance, or geometry can cause qualitatively different outcomes.
Residual Adaptation
- Instead of replacing a simulated policy after deployment, a real-world residual can correct it.
-
Let base policy be
\[a_t^{\mathrm{base}} = \pi_{\mathrm{sim}}( o_t )\] -
Learn residual:
\[\Delta a_t = \pi_{\mathrm{res}}( o_t )\] -
Execute:
\[a_t = a_t^{\mathrm{base}} + \Delta a_t\] - The residual only needs to model the remaining simulation error.
- This can reduce the amount of real-world data required relative to relearning the entire controller.
Privileged Learning in Simulation
-
Simulation exposes information unavailable during deployment:
\[s_t^{\mathrm{priv}} = \{ \text{exact poses}, \text{velocities}, \text{contacts}, \text{forces}, \text{object identities} \}\] -
A teacher can use this privileged state:
\[a_t^* = \pi_T( s_t^{\mathrm{priv}} )\] -
A student receives only deployable observations:
\[a_t = \pi_S( o_t )\] -
Distillation minimizes:
\[\mathcal{L} = D( \pi_S(o_t), \pi_T(s_t^{\mathrm{priv}}) )\] -
Thus simulation can simplify training even when privileged signals cannot exist on the real robot.
Simulating Rare Events
- One of simulation’s greatest advantages is control over event frequency.
-
Suppose dangerous event \(E\) occurs in reality with probability
\[P_{\mathrm{real}}(E) = 10^{-6}\] - Waiting to observe enough examples naturally may be impractical.
-
Simulation can instead sample:
\[P_{\mathrm{sim}}(E) \gg P_{\mathrm{real}}(E)\] - The policy can then train repeatedly on the rare condition.
-
This is particularly valuable for:
\[\text{collisions}, \quad \text{falls}, \quad \text{sensor failures}, \quad \text{unexpected obstacles}, \quad \text{driving edge cases}\] - Evaluation should later correct for the fact that the training distribution intentionally differs from the deployment frequency.
Simulation as an Evaluation Environment
- Simulation is not only for training.
-
A policy can be evaluated over a controlled test distribution:
\[\mathcal{E}_{\mathrm{test}} = \{ E_1,\ldots,E_N \}\] -
Because each scenario can be replayed exactly, two policies can be compared under matched conditions:
\[R( \pi_A,E_i ) \quad\text{versus}\quad R( \pi_B,E_i )\] - This enables regression testing.
-
When policy version changes from
\[\pi_k\]-
to
\[\pi_{k+1}\]- the complete scenario suite can be rerun to detect both improvements and regressions.
-
Scenario-Based Evaluation
- Average random-environment performance can hide important weaknesses.
-
Instead, define scenario classes:
\[C = \{ C_{\mathrm{nominal}}, C_{\mathrm{occlusion}}, C_{\mathrm{collision}}, C_{\mathrm{latency}}, C_{\mathrm{novel}}, C_{\mathrm{recovery}} \}\] -
For each class:
\[M_c = \mathbb{E}_{E\sim C_c} [ R(\pi,E) ]\] - This produces a capability profile rather than a single score.
- Failures in a particular scenario class can feed directly into the next simulation-generation cycle.
Simulation as a Data Engine
-
The simulator therefore plays several roles simultaneously:
\[\boxed{ \begin{array}{c} \text{Simulator}\\ \downarrow\\ \left\{ \begin{array}{l} \text{Policy Training}\\ \text{Synthetic Data}\\ \text{RL Experience}\\ \text{Rare-Event Generation}\\ \text{Regression Testing}\\ \text{Failure Replay}\\ \text{Counterfactual Planning} \end{array} \right. \end{array} }\] -
This is why modern Physical AI platforms increasingly integrate simulation directly with training and evaluation infrastructure rather than treating it as a separate engineering tool.
The Real-Sim-Real Flywheel
- The strongest simulation systems form a bidirectional loop with reality.
-
Real-world deployment produces:
\[\mathcal{D}_{\mathrm{real}}\] -
Real-to-sim reconstruction produces:
\[\mathcal{T} = G( \mathcal{D}_{\mathrm{real}} )\] -
Simulation generates:
\[\mathcal{D}_{\mathrm{sim}} = \operatorname{Rollout}( \mathcal{T}, \pi )\] -
Training produces:
\[\pi'\] -
The improved policy returns to reality:
\[\pi' \rightarrow \mathcal{D}_{\mathrm{real}}'\] -
Thus:
\[\boxed{ \text{Real} \rightarrow \text{Reconstruct} \rightarrow \text{Simulate} \rightarrow \text{Train} \rightarrow \text{Deploy} \rightarrow \text{Real} \rightarrow \circlearrowleft }\] - The simulator is continuously corrected by reality, while the physical policy is continuously improved through simulation.
The Emerging Simulation Stack for Physical AI
-
A mature Physical AI simulation stack increasingly resembles:
\[\boxed{ \begin{array}{c} \text{Real-World Captures / CAD / Robot Models}\\ \downarrow\\ \text{Open Scene Representation}\\ \downarrow\\ \text{Physics + Rendering + Sensor Models}\\ \downarrow\\ \text{System Identification}\\ \downarrow\\ \text{Domain Randomization}\\ \downarrow\\ \text{GPU-Parallel Environments}\\ \downarrow\\ \text{RL / Imitation / Synthetic Data}\\ \downarrow\\ \text{Digital Twin + Failure Replay}\\ \downarrow\\ \text{SIL / HIL Validation}\\ \downarrow\\ \text{Real-World Deployment}\\ \downarrow\\ \text{Telemetry + Reconstruction}\\ \circlearrowleft \end{array} }\] -
The central transition is:
\[\text{simulation as testing} \rightarrow \text{simulation as training infrastructure}\]-
followed by
\[\text{simulation as training infrastructure} \rightarrow \text{simulation as a continuously synchronized world model}\]
-
-
As Physical AI scales, the boundary between simulator, digital twin, synthetic-data generator, and learned world model is likely to become increasingly fluid. All four attempt to answer essentially the same question:
\[\boxed{ \text{If the agent takes this action, what happens next?} }\] -
The next section will examine Evaluation for Physical AI, including offline versus closed-loop evaluation, task success, generalization across objects and environments, long-horizon metrics, robustness and perturbation testing, recovery evaluation, simulation benchmarks, real-world evaluation, VLA benchmarks, autonomous-driving evaluation, and the design of evaluation suites that can drive the Physical AI data and post-training flywheel.
Evaluation for Physical AI
Why Physical AI Evaluation Is Different
- Evaluating a language model can often be reduced to comparing generated outputs against references or judging their semantic quality. Physical AI systems operate differently because model outputs alter the future observations on which subsequent decisions depend.
-
For a physical policy,
\[a_t \sim \pi_\theta(a_t\mid o_{\leq t},g)\]-
the action changes the environment:
\[s_{t+1} \sim P(s_{t+1}\mid s_t,a_t)\]-
which produces the next observation:
\[o_{t+1} \sim O(o_{t+1}\mid s_{t+1})\]
-
-
- Evaluation must therefore measure not only whether an individual action looks reasonable, but whether repeated closed-loop interaction eventually produces the desired physical outcome.
-
This distinction motivates a hierarchy:
\[\boxed{ \text{Prediction Quality} \rightarrow \text{Action Quality} \rightarrow \text{Trajectory Quality} \rightarrow \text{Task Success} \rightarrow \text{Long-Horizon Autonomy}. }\] - A strong Physical AI evaluation system should measure every layer rather than collapsing performance into a single offline metric.
Open-Loop Versus Closed-Loop Evaluation
- Open-loop evaluation evaluates predictions against recorded trajectories without allowing the policy to influence future observations.
-
Given dataset
\[\mathcal{D} = \{ (o_t,a_t^*) \}\]-
the policy predicts
\[\hat a_t = \pi_\theta(o_t)\]
-
-
A simple action error is
\[E_{\mathrm{action}} = \frac{1}{N} \sum_{t=1}^{N} \| \hat a_t-a_t^* \|_2^2\] - This is inexpensive and reproducible, but it evaluates the policy only on states generated by the dataset policy.
-
Closed-loop evaluation instead executes
\[a_t = \pi_\theta(o_t)\]-
and lets that action influence
\[o_{t+1}\]
-
-
The policy is therefore evaluated on the state distribution it induces:
\[s_t \sim d^{\pi_\theta}\] - This distinction is fundamental. Small action errors can accumulate into large state deviations, while actions that differ substantially from demonstrations may still successfully complete the task.
Why Action Error Can Be Misleading
-
Suppose an expert trajectory contains action
\[a_t^*\] -
A policy predicts
\[\hat a_t\] -
A large value of
\[\| \hat a_t-a_t^* \|\]- does not necessarily imply failure because multiple actions may be valid.
-
For example, a robot can approach a cup from different directions:
\[a_t^{(1)} \neq a_t^{(2)}\]-
while both eventually produce
\[\operatorname{grasped}(\text{cup})=1\]
-
-
Physical control is therefore often multimodal:
\[p(a_t\mid o_t,g)\]- may contain several valid modes.
-
This makes task-level and trajectory-level evaluation essential.
Task Success Rate
- The simplest closed-loop metric is binary task success.
-
For episode \(i\):
\[S_i = \begin{cases} 1, & \text{task completed},\\ 0, & \text{otherwise}. \end{cases}\] -
Success rate is
\[\operatorname{SR} = \frac{1}{N} \sum_{i=1}^{N} S_i\] -
For manipulation, success might mean:
\[\operatorname{inside}( \text{object}, \text{container} )=1\] -
For navigation:
\[d( x_T,x_{\mathrm{goal}} ) < \epsilon\] - For autonomous driving, success may require completing the route without collision, rule violation, or unacceptable intervention.
- Success rate is intuitive, but binary success hides how and why failures occur.
Partial Credit
- Long-horizon tasks often contain multiple subgoals.
-
Suppose task
\[G = \{ g_1,\ldots,g_K \}\] -
Define subgoal completion:
\[c_k \in \{0,1\}\] -
Then progress can be measured as
\[P = \frac{1}{K} \sum_{k=1}^{K} c_k\] -
A robot completing four of five subtasks receives
\[P=0.8\]- even if final task success is zero.
- This is useful for diagnosing whether a policy fails immediately or only near the end of a long sequence.
Long-Horizon Evaluation
- Physical AI becomes substantially harder as the number of required successful decisions grows.
-
If individual skill success probability is
\[p\]-
and a task requires \(K\) approximately independent successful skills, then
\[P_{\mathrm{task}} \approx p^K\]
-
-
For
\[p=0.95\]-
and
\[K=20\]-
task success is only approximately
\[0.95^{20} \approx 0.36\]
-
-
- Evaluation must therefore include tasks that require multiple sequential interactions.
- Short-horizon benchmarks can substantially overestimate readiness for autonomous deployment.
Sequence-Length Metrics
-
For tasks containing multiple instructions, define maximum completed prefix:
\[L_i = \max \{ k: g_1,\ldots,g_k \text{ completed} \}\] -
Average sequence length is
\[\bar L = \frac{1}{N} \sum_i L_i\] - CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks by Mees et al. (2022) evaluates language-conditioned policies on chains of sequential manipulation instructions, explicitly emphasizing long-horizon execution and zero-shot generalization to novel language and environments.
-
This exposes a capability that single-task success cannot measure:
\[\boxed{ \text{Can the agent keep succeeding?} }\]
Evaluating Generalization
- A generalist Physical AI model should be evaluated outside its exact training distribution.
-
Let training distribution be
\[p_{\mathrm{train}}( o,g,e )\]-
and test distribution be
\[p_{\mathrm{test}}( o,g,e )\]
-
-
Evaluation should systematically vary:
\[\text{objects}, \quad \text{instructions}, \quad \text{environments}, \quad \text{embodiments}, \quad \text{dynamics}\] - A useful evaluation matrix is:
- Dimension In-distribution Generalization ———– ——————— —————- Object Seen object Novel object Task Seen task Novel task Language Seen phrasing Novel phrasing Scene Seen scene Novel scene Robot Training embodiment New embodiment Dynamics Nominal Perturbed
- The distinction between memorization and generalization is therefore explicit.
Object Generalization
-
Suppose training contains:
\[\mathcal{O}_{\mathrm{train}} = \{ o_1,\ldots,o_n \}\] -
Evaluation samples:
\[o^* \notin \mathcal{O}_{\mathrm{train}}\] -
Measure:
\[\operatorname{SR}_{\mathrm{novel\ object}}\] -
A stronger test changes not only visual appearance but physical properties:
\[\text{shape}, \quad \text{mass}, \quad \text{size}, \quad \text{compliance}, \quad \text{friction}\] -
This separates semantic object recognition from actual physical generalization.
Language Generalization
- A language-conditioned robot should understand paraphrases.
-
For task goal \(g\), define instruction set
\[L_g = \{ l_1,l_2,\ldots,l_M \}\] - For example:
Put the mug in the sink.
Move the cup into the sink.
Take the mug over to the sink.
Place this drinking cup inside the sink.
-
Evaluation should test whether
\[\pi(a\mid o,l_i)\]- remains behaviorally consistent across valid linguistic realizations.
-
This matters because real users do not interact with robots using benchmark templates.
Environment Generalization
-
Train environments may contain:
\[E_{\mathrm{train}} = \{ e_1,\ldots,e_N \}\] -
Test environment:
\[e^* \notin E_{\mathrm{train}}\] -
Changes can include:
\[\text{layout}, \quad \text{lighting}, \quad \text{background}, \quad \text{clutter}, \quad \text{object placement}\] -
The resulting metric
\[\operatorname{SR}_{\mathrm{novel\ env}}\]- measures whether the policy learned transferable behavior rather than environment-specific correlations.
Compositional Generalization
- An especially important test evaluates new combinations of known concepts.
-
Suppose training includes:
\[\text{pick(red cup)}\]-
and
\[\text{place(blue bowl, sink)}\]
-
-
A compositional test might request:
\[\text{place(red cup, sink)}\] - The individual concepts are familiar, but the combination is new.
-
Formally:
\[g^* \notin G_{\mathrm{train}}\]-
while
\[\operatorname{components}(g^*) \subseteq \operatorname{components}(G_{\mathrm{train}})\]
-
- This evaluates whether the system can recombine learned physical capabilities.
Lifelong and Knowledge-Transfer Evaluation
- Physical agents may encounter new tasks continually.
-
Let task sequence be:
\[T_1,T_2,\ldots,T_K\] -
After learning task \(T_k\), evaluate performance on all earlier tasks:
\[R_{k,j}, \qquad j\leq k\] -
This enables measurement of:
\[\text{forward transfer}, \quad \text{backward transfer}, \quad \text{catastrophic forgetting}\] - LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning by Liu et al. (2023) provides four suites totaling 130 manipulation tasks and explicitly studies transfer of declarative and procedural knowledge, task ordering, policy architecture, and the effects of pretraining.
- For continually updated Physical AI systems, such evaluation becomes important because improving new capabilities should not silently degrade old ones.
Robustness Evaluation
- A policy should be evaluated under controlled perturbations.
-
Let nominal environment parameters be
\[\phi_0\] -
Perturbed evaluation uses
\[\phi = \phi_0+\delta\] -
Possible perturbations include:
\[\text{camera movement}\] \[\text{lighting changes}\] \[\text{object displacement}\] \[\text{sensor noise}\] \[\text{latency}\] \[\text{actuator error}\] -
Define robustness curve:
\[R(\delta) = \operatorname{SR}( \pi;\delta )\] - A robust policy should degrade gradually rather than collapse under small perturbations.
Perturbation Sweeps
-
Instead of evaluating at one perturbation magnitude, sweep:
\[\delta \in \{ \delta_1,\ldots,\delta_K \}\] -
Then plot:
\[\operatorname{SR}(\delta)\] -
A summary metric can integrate the curve:
\[\operatorname{AUC}_{\mathrm{robust}} = \int \operatorname{SR}(\delta) \,d\delta\] -
This distinguishes two policies that have identical nominal success but very different failure sensitivity.
Recovery Evaluation
- A robust autonomous agent should recover from errors.
-
Evaluation can deliberately introduce disturbance at time \(t\):
\[s_t \rightarrow s_t'\] -
Examples include:
\[\text{move the target object}\] \[\text{cause a grasp to slip}\] \[\text{close an open drawer}\] \[\text{place an obstacle in the path}\] -
Recovery rate is:
\[R_{\mathrm{recover}} = P( \text{eventual success} \mid \text{disturbance} )\] - This measures a fundamentally different capability from nominal task execution.
Time-to-Recovery
- Successful recovery can still be inefficient.
-
Let disturbance occur at
\[t_f\]-
and normal task progress resume at
\[t_r\]
-
-
Recovery time is
\[T_{\mathrm{recover}} = t_r-t_f\] -
Similarly, count additional actions:
\[A_{\mathrm{recover}} = N_{\mathrm{postfailure}} - N_{\mathrm{nominal}}\] - A capable autonomous system should both recover reliably and avoid repeatedly entering ineffective loops.
Intervention Rate
- For deployed systems, human intervention is a critical metric.
-
Let
\[I_i = \text{number of interventions during episode }i\] -
Intervention rate can be measured as:
\[R_I = \frac{\sum_i I_i} {\sum_i T_i}\] -
Depending on the domain, the denominator might be:
\[\text{hours}, \quad \text{kilometers}, \quad \text{tasks}, \quad \text{actions}\] - A useful autonomy system should reduce intervention frequency while maintaining safety and task success.
Autonomy Duration
- Another deployment metric is uninterrupted autonomous operation.
-
Let
\[T_i = \text{time until intervention or unrecoverable failure}\] -
Then:
\[\operatorname{MTBI} = \frac{1}{N} \sum_i T_i\]- where MTBI denotes mean time between interventions.
- For long-running robots, this may be more informative than isolated episode success.
- A robot that completes 95% of short benchmark tasks may still be impractical if it requires human assistance every few minutes.
Safety Metrics
- Task success must be evaluated jointly with safety.
-
Define violation indicators:
\[C_{\mathrm{collision}}, \quad C_{\mathrm{force}}, \quad C_{\mathrm{workspace}}, \quad C_{\mathrm{human}}, \quad C_{\mathrm{rule}}\] -
Safety violation rate can be:
\[V = \frac{ \sum_i \mathbb{1}[ \text{safety violation in episode }i ] }{ N }\] - Task reward should not allow unsafe behavior to be hidden by high success.
-
A useful evaluation reports:
\[\boxed{ \text{Success} \quad\text{and}\quad \text{Safety} }\]- separately.
Severity-Weighted Safety
- Not all failures have equal consequence.
-
Assign severity:
\[w_j\]- to event type \(j\).
-
Then:
\[R_{\mathrm{risk}} = \sum_j w_j P(E_j)\] - For example, gently contacting a table should not necessarily receive the same weight as contacting a person.
- This is particularly important in autonomous driving and human-robot interaction, where low-frequency high-severity events dominate deployment risk.
Constraint-Based Evaluation
-
Safety can also be expressed as constraints:
\[C_j(\tau) \leq c_j\] -
For example:
\[F_{\mathrm{contact}} < F_{\max}\] \[d_{\mathrm{human}} > d_{\min}\] \[v < v_{\max}\] -
A trajectory can therefore be classified as successful only if
\[R_{\mathrm{task}}(\tau) \geq r_{\min}\]-
and
\[C_j(\tau) \leq c_j \quad \forall j\]
-
-
This prevents reward tradeoffs from treating certain safety constraints as optional.
Efficiency Metrics
- A robot can complete a task successfully but inefficiently.
-
Useful efficiency measures include:
\[T_{\mathrm{completion}}\] \[N_{\mathrm{actions}}\] \[L_{\mathrm{path}}\] \[E_{\mathrm{energy}}\] -
Normalized path efficiency can be:
\[\eta_{\mathrm{path}} = \frac{ L_{\mathrm{optimal}} }{ L_{\mathrm{executed}} }\] -
Similarly, action efficiency can compare actual action count with a reference:
\[\eta_{\mathrm{action}} = \frac{ N_{\mathrm{reference}} }{ N_{\mathrm{executed}} }\] - Efficiency matters because unnecessarily long physical trajectories consume time, energy, actuator lifetime, and potentially user patience.
Success Weighted by Efficiency
- Navigation benchmarks often motivate metrics that combine success and efficiency.
-
A general form is:
\[M = S \cdot \frac{ C^* }{ \max(C,C^*) }\]- where:
- \(S\) indicates success.
- \(C\) is actual execution cost.
- \(C^*\) is reference cost.
- The same principle can be applied to manipulation using trajectory length, execution time, or energy.
-
This distinguishes:
\[\text{successful and efficient}\]-
from
\[\text{successful but wasteful}\]
-
Evaluation of Vision-Language-Action Models
-
A VLA should be evaluated across several dimensions:
\[\boxed{ \begin{array}{c} \text{Language Understanding}\\ \text{Visual Grounding}\\ \text{Manipulation Skill}\\ \text{Action Precision}\\ \text{Generalization}\\ \text{Long-Horizon Execution}\\ \text{Recovery} \end{array} }\] - A single aggregate task-success number cannot determine which component limits performance.
- Evaluation suites should therefore isolate capability dimensions where possible.
LIBERO
- LIBERO focuses on lifelong robot manipulation and knowledge transfer. Its benchmark contains four suites with 130 tasks and human-teleoperated demonstrations, with experiments designed around transfer, pretraining, task ordering, and continual learning.
- Its suites allow evaluation of different forms of generalization and transfer rather than treating robot manipulation as a single homogeneous task distribution.
- For modern VLAs, LIBERO has consequently become useful for testing whether pretrained representations and policies can adapt to downstream manipulation tasks.
CALVIN
- CALVIN focuses on language-conditioned long-horizon manipulation.
-
A typical evaluation asks the agent to complete a chain:
\[[ g_1, g_2, g_3, g_4, g_5 ]\] - Metrics include success after one through five sequential instructions and average completed sequence length. The benchmark also evaluates generalization to novel language, environments, and objects.
-
CALVIN therefore stresses temporal composition:
\[\boxed{ \text{Can individually learned skills be chained reliably?} }\]
BEHAVIOR-1K
- BEHAVIOR-1K expands evaluation toward household-scale embodied intelligence. The benchmark contains 1,000 everyday activities instantiated in 50 interactive scenes with more than 10,000 objects, with tasks grounded in surveys of activities people actually perform and want assistance with.
-
The scale changes the evaluation target from narrow manipulation toward:
\[\text{navigation} + \text{manipulation} + \text{state tracking} + \text{long-horizon planning}\] - Examples such as cooking, cleaning, and organizing naturally require many interacting skills and environmental state changes.
Simulation Versus Real-World Evaluation
-
Simulation provides reproducibility:
\[E_i = \operatorname{Reset}( \text{seed}_i )\] - Policies can therefore be compared under identical conditions.
-
Real-world evaluation samples from:
\[p_{\mathrm{real}}(s,o,d)\]- where \(d\) includes uncontrolled disturbances.
-
Simulation is useful for:
\[\text{scale}, \quad \text{repeatability}, \quad \text{rare events}, \quad \text{controlled perturbations}\] -
Real-world evaluation is required to reveal:
\[\text{reality-gap failures}, \quad \text{hardware failures}, \quad \text{unmodeled dynamics}, \quad \text{real perception errors}\] - Neither should replace the other.
SimplerEnv and Real-to-Sim Evaluation
- SIMPLER evaluates real-world manipulation policies such as RT-1, RT-1-X, and Octo in simulated replicas of real robot setups. It provides both visual-matching evaluation and variant-aggregation evaluation, where multiple backgrounds, lighting conditions, distractors, and textures are aggregated.
-
This introduces an important evaluation question:
\[\operatorname{Corr} ( R_{\mathrm{sim}}, R_{\mathrm{real}} )\] - A simulator is useful as a policy benchmark only if relative performance in simulation provides meaningful information about real-world performance.
- Thus simulation evaluation itself must be validated.
Evaluating Sim-to-Real Correlation
-
Suppose policies
\[\Pi = \{ \pi_1,\ldots,\pi_K \}\]-
have simulated scores
\[S_i\]-
and real scores
\[R_i\]
-
-
-
Measure ranking correlation:
\[\rho = \operatorname{Spearman} ( S,R )\] -
High
\[\rho\]- means the simulator preserves policy ranking even if absolute success rates differ.
-
This may be more useful than requiring:
\[S_i \approx R_i\] -
For model development, knowing that an improvement in simulation is likely to remain an improvement in reality can make large-scale simulated evaluation useful.
Autonomous Driving Evaluation
- Autonomous driving illustrates especially clearly why open-loop trajectory metrics are insufficient.
-
A planner may predict trajectory
\[\hat\tau\]-
that differs from recorded human trajectory
\[\tau^*\]
-
-
A displacement metric such as
\[\operatorname{ADE} = \frac{1}{T} \sum_t \| \hat x_t-x_t^* \|_2\]- measures imitation accuracy, not necessarily driving quality.
- A different trajectory may be completely safe.
- Conversely, a trajectory close to the recorded human path may lead to collision once other agents react to it.
nuPlan
- nuPlan: A Closed-Loop ML-Based Planning Benchmark for Autonomous Vehicles by Caesar et al. (2021) was introduced specifically to address limitations of open-loop planning evaluation. It combines a large-scale driving dataset, a lightweight closed-loop simulator, reactive agents, and planning-specific metrics.
-
The central evaluation loop is:
\[\text{Planner} \rightarrow \text{Ego Action} \rightarrow \text{Simulation} \rightarrow \text{Updated Traffic} \rightarrow \text{Planner}\] - This evaluates the consequences of the planner’s own decisions rather than merely comparing predictions against logged human trajectories.
Interactive Driving Evaluation
- Closed-loop driving evaluation must model other agents reacting to the ego vehicle.
-
Let ego action be:
\[a_t^{\mathrm{ego}}\] -
Other agents respond:
\[a_t^{j} \sim \pi_j( s_t,a_t^{\mathrm{ego}} )\] -
The transition is therefore:
\[s_{t+1} = F( s_t, a_t^{\mathrm{ego}}, a_t^1,\ldots,a_t^N )\] - A static replay of recorded traffic cannot fully represent this interaction because logged agents do not react to counterfactual ego behavior.
- Recent work such as nuPlan-R by Peng et al. (2025) extends closed-loop evaluation with learned reactive multi-agent simulation, motivated by limitations of rule-based reactive traffic agents in representing diverse interactions.
Waymo Open Dataset Evaluation
- The Waymo Open Dataset Challenges span perception, motion prediction, interaction prediction, occupancy and flow, scenario generation, simulated agents, and vision-based end-to-end driving. Its leaderboards remain active even though Waymo is not hosting formal challenges in 2026.
-
The breadth illustrates how autonomous-driving evaluation has expanded from isolated perception toward:
\[\boxed{ \text{Perception} \rightarrow \text{Prediction} \rightarrow \text{Simulation} \rightarrow \text{End-to-End Driving}. }\] - Physical AI evaluation increasingly follows the same trajectory.
Evaluating World Models
-
A world model predicts:
\[\hat s_{t+1} = M_\phi( s_t,a_t )\] -
One-step prediction error is:
\[E_1 = d( \hat s_{t+1}, s_{t+1} )\] -
But physical usefulness depends on longer rollouts:
\[\hat s_{t+H} = M_\phi^{(H)} ( s_t,a_{t:t+H-1} )\] -
Evaluate:
\[E_H = d( \hat s_{t+H}, s_{t+H} )\] -
Because errors compound, evaluation should sweep horizon:
\[H \in \{ 1,5,10,50,\ldots \}\] -
A world model useful for planning must preserve action-conditioned consequences over the horizon relevant to decisions.
Evaluating World Models by Decision Quality
- Pixel reconstruction quality alone may not indicate whether a world model is useful.
-
Suppose two world models produce:
\[M_1, \quad M_2\] -
Even if:
\[E_{\mathrm{pixel}}(M_1) < E_{\mathrm{pixel}}(M_2)\]- the second may preserve task-relevant geometry and dynamics better.
-
A stronger evaluation measures downstream planning:
\[R_{\mathrm{plan}}( M_\phi )\] -
Thus:
\[\boxed{ \text{World-model quality} = \text{predictive accuracy} + \text{decision usefulness}. }\]
Counterfactual Evaluation
- A particularly difficult capability is counterfactual accuracy.
-
Given identical history:
\[h_t\]-
consider alternative actions:
\[a_t^{(1)}, a_t^{(2)}\]
-
-
A world model should predict distinct consequences:
\[\hat s_{t+1}^{(1)} = M( h_t,a_t^{(1)} )\] \[\hat s_{t+1}^{(2)} = M( h_t,a_t^{(2)} )\] - Evaluation should determine whether these action-dependent differences correspond to real physical outcomes.
- Without this capability, visually plausible generation may still be useless for planning.
Evaluating Agentic Physical AI
- Agentic systems introduce another layer of metrics.
-
A complete agent may perform:
\[\text{Reason} \rightarrow \text{Plan} \rightarrow \text{Tool Call} \rightarrow \text{Skill} \rightarrow \text{Verify} \rightarrow \text{Replan}\] -
Evaluation should separately measure:
\[\text{planning accuracy}\] \[\text{skill-selection accuracy}\] \[\text{tool-use success}\] \[\text{failure detection}\] \[\text{recovery success}\] \[\text{final task completion}\] - This allows failure attribution rather than treating every unsuccessful episode as a policy failure.
Failure Taxonomy
-
For each unsuccessful episode, assign failure category:
\[F \in \{ F_{\mathrm{perception}}, F_{\mathrm{reasoning}}, F_{\mathrm{planning}}, F_{\mathrm{control}}, F_{\mathrm{tool}}, F_{\mathrm{recovery}}, F_{\mathrm{safety}} \}\] -
Estimate:
\[P(F_j)\] -
This produces a failure distribution:
\[\mathbf{f} = [ p_1,\ldots,p_K ]\] - A model update should ideally reduce specific components rather than merely move the aggregate success number.
- Failure taxonomy therefore connects evaluation directly to post-training.
Conditional Metrics
- Aggregate success can hide important subpopulation failures.
-
Let metadata variable be:
\[z \in \{ \text{object type}, \text{task}, \text{scene}, \text{lighting}, \text{robot}, \text{difficulty} \}\] -
Measure:
\[\operatorname{SR}(z)\] -
For example:
\[\operatorname{SR}_{\mathrm{transparent\ objects}}\]- may be much lower than overall success.
- Conditional evaluation identifies capability gaps that would otherwise disappear inside averages.
Confidence Intervals
- Physical evaluations are often expensive, producing relatively small sample sizes.
-
If success rate is:
\[\hat p = \frac{k}{N}\]- the estimate has uncertainty.
-
Evaluation should report:
\[\hat p \pm \mathrm{CI}\] -
When comparing policies \(A\) and \(B\), paired evaluation is particularly valuable:
\[\Delta_i = S_i^A-S_i^B\] -
The comparison should estimate uncertainty over:
\[\bar\Delta\] - Without uncertainty estimates, small apparent improvements may simply reflect evaluation noise.
Repeated Trials
- Physical environments are stochastic.
-
For scenario \(e_i\), run:
\[K\]-
trials:
\[S_{i,1},\ldots,S_{i,K}\]
-
-
Estimate:
\[\hat p_i = \frac{1}{K} \sum_k S_{i,k}\] -
Repeated trials are especially important when:
\[\text{object initialization}, \quad \text{grasp contact}, \quad \text{sensor noise}, \quad \text{policy sampling}\]- introduce variability.
- A single successful demonstration is not sufficient evidence of robust capability.
Deterministic Versus Stochastic Policy Evaluation
-
Generative policies may sample actions:
\[a_t \sim \pi_\theta( a_t\mid o_t )\] - A deterministic evaluation with one random seed measures only one realization.
-
Instead evaluate:
\[\{ \tau_1,\ldots,\tau_K \} \sim \pi_\theta\] -
Then estimate:
\[\mathbb{E}[R]\]-
and potentially:
\[\operatorname{Var}(R)\]
-
- For safety-critical settings, tail behavior may matter more than average performance.
Tail-Risk Evaluation
-
Suppose return distribution is:
\[R \sim p(R)\] -
Average:
\[\mathbb{E}[R]\]- may hide rare catastrophic outcomes.
-
A tail-risk metric can examine:
\[\operatorname{CVaR}_{\alpha}(R)\]- the expected return in the worst \(\alpha\) fraction of outcomes.
-
For Physical AI, evaluation should often ask not only:
\[\text{How well does the policy usually work?}\]-
but:
\[\boxed{ \text{What happens in its worst plausible failures?} }\]
-
Stress Testing
-
A stress-test generator searches for environments where the policy fails:
\[e^* = \arg\min_e R( \pi,e )\] -
Instead of sampling only from nominal distribution:
\[e \sim p_{\mathrm{nominal}}\]-
evaluation deliberately explores:
\[p_{\mathrm{stress}}\]
-
-
This can include:
\[\text{extreme clutter}, \quad \text{occlusion}, \quad \text{unusual geometry}, \quad \text{sensor degradation}, \quad \text{unexpected human behavior}\] -
Stress testing transforms evaluation from passive measurement into active failure discovery.
Adversarial Scenario Generation
-
A learned scenario generator can optimize:
\[\max_\psi \mathbb{E}_{e\sim q_\psi} [ L( \pi,e ) ]\]- where \(L\) measures policy failure.
- The generator seeks scenarios that remain physically plausible but expose weaknesses.
-
This produces an adversarial loop:
\[\boxed{ \text{Policy} \rightarrow \text{Adversarial Environment} \rightarrow \text{Failure} \rightarrow \text{Training Data}. }\] - Evaluation thereby becomes a direct component of the data engine.
Evaluation as a Data Flywheel
- A mature evaluation system should not terminate with a leaderboard score.
-
Suppose evaluation identifies failures:
\[\mathcal{F} = \{ f_1,\ldots,f_N \}\] -
Cluster them:
\[C = \operatorname{Cluster}( \mathcal{F} )\] -
Prioritize cluster \(c\) using:
\[P(c) = \operatorname{frequency}(c) \times \operatorname{severity}(c) \times \operatorname{importance}(c)\] -
Then generate targeted data:
\[\mathcal{D}_c\] -
Post-train:
\[\pi_{\theta'} = \operatorname{Train}( \pi_\theta, \mathcal{D}_c )\] -
Re-evaluate:
\[\pi_{\theta'} \rightarrow \mathcal{E}\] -
Thus:
\[\boxed{ \text{Evaluate} \rightarrow \text{Diagnose} \rightarrow \text{Generate Data} \rightarrow \text{Train} \rightarrow \text{Evaluate}. }\] - Evaluation becomes the steering mechanism for model improvement.
Regression Evaluation
- Every model update can improve one capability while degrading another.
-
Let benchmark vector be:
\[\mathbf{m}_k = [ m_1,\ldots,m_N ]\]- for model version \(k\).
-
After training:
\[\mathbf{m}_{k+1}\] -
Define change:
\[\Delta\mathbf{m} = \mathbf{m}_{k+1} - \mathbf{m}_k\] -
A deployment gate can require:
\[\Delta m_i \geq -\epsilon_i\]- for protected capabilities.
- This prevents aggregate improvements from masking severe regressions in specific behaviors.
Evaluation Gates
-
A Physical AI model should pass multiple gates before deployment:
\[\boxed{ \begin{array}{c} \text{Offline Evaluation}\\ \downarrow\\ \text{Simulation Evaluation}\\ \downarrow\\ \text{Stress Tests}\\ \downarrow\\ \text{Software-in-the-Loop}\\ \downarrow\\ \text{Hardware-in-the-Loop}\\ \downarrow\\ \text{Controlled Real-World Trials}\\ \downarrow\\ \text{Shadow Evaluation}\\ \downarrow\\ \text{Limited Deployment}\\ \downarrow\\ \text{Production Monitoring} \end{array} }\] - Each stage increases realism while also increasing cost and potential consequence.
- The purpose is to discover failures as early and cheaply as possible.
A Physical AI Evaluation Scorecard
-
A useful evaluation suite should report a vector rather than a single number:
\[\mathbf{E} = [ S, G, R, C, F, L, I, Q ]\]-
where:
\[S = \text{task success}\] \[G = \text{generalization}\] \[R = \text{robustness}\] \[C = \text{safety compliance}\] \[F = \text{recovery}\] \[L = \text{long-horizon performance}\] \[I = \text{intervention efficiency}\] \[Q = \text{execution quality/efficiency}\]
-
- This makes tradeoffs visible.
- A model can then improve on one dimension without the evaluation framework implicitly declaring it globally superior.
The Emerging Evaluation Stack
-
A mature Physical AI evaluation system increasingly resembles:
\[\boxed{ \begin{array}{c} \text{Offline Action Metrics}\\ \downarrow\\ \text{Single-Skill Closed-Loop Tests}\\ \downarrow\\ \text{Multi-Skill / Long-Horizon Tasks}\\ \downarrow\\ \text{Generalization Splits}\\ \downarrow\\ \text{Perturbation + Recovery Tests}\\ \downarrow\\ \text{Simulation Benchmarks}\\ \downarrow\\ \text{Real-to-Sim Correlation}\\ \downarrow\\ \text{Real-World Trials}\\ \downarrow\\ \text{Safety + Tail-Risk Evaluation}\\ \downarrow\\ \text{Deployment Monitoring}\\ \downarrow\\ \text{Failure Mining}\\ \circlearrowleft \end{array} }\] -
The key conceptual transition is:
\[\boxed{ \text{evaluation as measurement} \rightarrow \text{evaluation as failure discovery} \rightarrow \text{evaluation as a training signal}. }\] - For Physical AI, the most useful benchmark is therefore not merely one that produces a reliable score. It is one that reveals which physical capabilities fail, under what conditions they fail, how severe those failures are, and which new experience should be collected to improve the system.
- The next section will examine Safety and Reliability for Physical AI, including layered safety architectures, runtime safety monitors, constraint enforcement, uncertainty and out-of-distribution detection, fail-safe control, human override, safe exploration, reward hacking, specification failures, VLA-specific risks, autonomous-driving safety, verification, red teaming, and deployment monitoring.
Safety and Reliability for Physical AI
Why Physical AI Safety Is Different
-
Physical AI systems act on the real world, so model errors can become physical events rather than merely incorrect outputs. A policy can be written as:
\[a_t \sim \pi_\theta(a_t \mid o_{\leq t}, g)\]-
with trajectory:
\[\tau=(s_0,a_0,s_1,a_1,\ldots,s_T)\]
-
-
For a safe state set \(\mathcal{S}_{\mathrm{safe}}\), a system-level objective is:
\[P(s_t \in \mathcal{S}_{\mathrm{safe}}\;\forall t)\geq 1-\delta\] -
Safety is therefore a property of the complete sensing, reasoning, planning, control, hardware, and operational stack rather than of the learned model alone.
Defense in Depth
-
Physical AI systems should use multiple independent or partially independent safeguards:
\[\boxed{ \text{Semantic Safety} + \text{Planning Safety} + \text{Physical Safety} + \text{Runtime Monitoring} + \text{Operational Safety} }\] -
Google DeepMind’s robotics safety framework similarly treats robotics safety as a layered problem spanning semantic reasoning, low-level safety mechanisms, evaluation, and deployment practices.
The Safety Stack
- A practical architecture can be organized as:
User Instruction
↓
Semantic Safety Filter
↓
Agent / Task Planner
↓
Plan Validator
↓
VLA / Learned Policy
↓
Runtime Safety Monitor
↓
Safety Controller
↓
Actuators
↓
Physical Environment
- Emergency-stop and hardware safety paths should be able to bypass the learned planning stack.
Semantic Safety and Robot Constitutions
-
A robot may be mechanically capable of performing an action that should nevertheless be rejected. Semantic safety therefore asks:
\[\text{Can the robot execute the action?} \neq \text{Should the robot execute the action?}\] -
A semantic safety score can be modeled as:
\[C_{\mathrm{semantic}}=f(g,o_t,M_t)\] -
Natural-language safety rules can be represented as a constitution:
\[\mathcal{C}=\{c_1,\ldots,c_K\}\]-
with a plan accepted only when:
\[\operatorname{Valid}(P,\mathcal{C})=1\]
-
-
The ASIMOV benchmark studies semantic safety reasoning for embodied agents and provides scenarios in which physical feasibility and normative acceptability must be distinguished.
Physical Constraints
-
Semantic safeguards should be combined with explicit physical constraints such as:
\[F_{\mathrm{contact}} < F_{\max}\] \[d_{\mathrm{human}} > d_{\min}\]-
and:
\[q_{\min}\leq q_t\leq q_{\max}\]
-
-
A constrained policy objective can be written:
\[\max_\pi \mathbb{E}[R(\tau)]\]-
subject to:
\[\mathbb{E}[C_j(\tau)]\leq d_j\]
-
-
Constrained Policy Optimization by Achiam et al. (2017) develops policy optimization under explicit expected-cost constraints.
Why Reward Penalties Are Not Enough
-
A soft reward penalty:
\[R'=R_{\mathrm{task}}-\lambda C_{\mathrm{safety}}\]-
does not guarantee safety because sufficiently high task reward can still compensate for a violation. Safety-critical requirements are better represented as explicit constraints:
\[C_{\mathrm{safety}}\leq C_{\max}\]
-
-
This separates preferences from requirements.
Safe Reinforcement Learning
-
For discounted task reward:
\[J_R(\pi)=\mathbb{E}\left[\sum_t\gamma^t r_t\right]\]-
and safety cost:
\[J_C(\pi)=\mathbb{E}\left[\sum_t\gamma^t c_t\right]\]-
safe RL can be formulated as:
\[\max_\pi J_R(\pi) \quad \text{subject to} \quad J_C(\pi)\leq d\]
-
-
-
A Comprehensive Survey on Safe Reinforcement Learning by García and Fernández (2015) surveys constrained exploration, risk-sensitive objectives, and external safety mechanisms.
Runtime Assurance and Shielding
-
A learned controller should not necessarily be the final authority over actuator commands. Let:
\[a_t^L\]-
be the learned action and:
\[M(s_t,a_t^L)\in\{\mathrm{safe},\mathrm{unsafe}\}\]-
be a runtime monitor. The executed action can be:
\[a_t= \begin{cases} a_t^L, & M(s_t,a_t^L)=\mathrm{safe}\\ a_t^S, & M(s_t,a_t^L)=\mathrm{unsafe}, \end{cases}\]- where \(a_t^S\) is generated by a trusted fallback controller.
-
-
-
Safe Reinforcement Learning via Shielding by Alshiekh et al. (2018) formalizes runtime shields that prevent actions violating specified safety properties.
Control Barrier Functions
-
For safe set:
\[\mathcal{C}=\{s:h(s)\geq0\}\]-
and dynamics:
\[\dot{s}=f(s)+g(s)u\]-
a control barrier condition can require:
\[\dot{h}(s)+\alpha(h(s))\geq0\]
-
-
-
A nominal learned control can then be projected onto the safe set:
\[u^*= \arg\min_u \|u-u_{\mathrm{policy}}\|^2\]- subject to the barrier constraint.
-
Control Barrier Function Based Quadratic Programs for Safety Critical Systems by Ames et al. (2017) develops this approach for safety-critical control.
Predictive Safety
-
Instead of checking only the immediate action, the system can predict a future trajectory:
\[\hat{\tau} = (\hat{s}_{t+1},\ldots,\hat{s}_{t+H})\]-
and estimate:
\[P( \hat{\tau}\cap\mathcal{S}_{\mathrm{unsafe}}\neq\varnothing )\]
-
- If predicted risk exceeds a threshold, the controller can replan, slow down, switch to a fallback policy, or stop.
-
Model-predictive safety can optimize:
\[A^* = \arg\max_A R(A)\]-
subject to:
\[\hat{s}_{t+k}\in\mathcal{S}_{\mathrm{safe}} \quad \forall k\in\{1,\ldots,H\}\]
-
Uncertainty and Out-of-Distribution Detection
-
A physical agent should reduce autonomy as uncertainty increases. Let:
\[U_t=U(o_t,a_t,s_t)\] -
A practical policy may map:
Low uncertainty → execute
Medium uncertainty → gather information or slow down
High uncertainty → stop or request assistance
-
Uncertainty can be decomposed conceptually into:
\[U=U_{\mathrm{aleatoric}}+U_{\mathrm{epistemic}}\] -
Out-of-distribution detection can additionally use an anomaly score:
\[A(o_t)=d(z_t,\mathcal{Z}_{\mathrm{train}})\] -
High anomaly scores can trigger conservative behavior.
Task Feasibility and Ambiguity
-
Before acting, an agent can estimate:
\[F(g,s_t) = P( \text{safe successful completion}\mid g,s_t )\] - Low feasibility should lead to refusal, clarification, task modification, or human assistance rather than blind execution.
-
Similarly, if several interpretations of an instruction are plausible:
\[\mathcal{G}=\{g_1,\ldots,g_K\}\]-
the robot should request clarification when:
\[\max_i P(g_i\mid o_{\leq t},l)<\tau\]
-
Human Proximity and External Monitoring
-
Human-aware operation can define speed or action limits as a function of distance and relative motion. Reaction time can be decomposed as:
\[T_{\mathrm{reaction}} = T_{\mathrm{sense}} + T_{\mathrm{detect}} + T_{\mathrm{decision}} + T_{\mathrm{brake}}\]-
with approximate reaction distance:
\[d_{\mathrm{reaction}} \approx vT_{\mathrm{reaction}}\]
-
- This makes model and system latency direct safety parameters.
- External sensing can provide an additional “outside-in” safety layer independent of the robot’s primary perception stack. NVIDIA’s Halos for Robotics describes a layered safety approach for Physical AI systems.
Functional Safety and Fail-Safe Behavior
- Physical AI systems must handle failures beyond model prediction errors, including:
- sensor faults,
- stale observations,
- communication loss,
- memory corruption,
- compute failures,
- actuator degradation,
- process hangs.
- actuator degradation,
- compute failures,
- memory corruption,
- communication loss,
- stale observations,
- sensor faults,
- A safety runtime can monitor these conditions and transition the system to a defined safe state.
- Fail-safe systems transition toward a safe stop after faults, while fail-operational systems attempt degraded operation before eventually reaching a safe state. The correct behavior depends on the application because immediate torque-off or stopping can itself be hazardous in some physical configurations.
Watchdogs and Human Override
-
Watchdogs should detect missed deadlines or stalled components:
\[t_{\mathrm{now}}-t_{\mathrm{last\ update}}>T_{\max} \Rightarrow \text{fallback}\] -
Human override should bypass high-level learned components when necessary. Emergency-stop mechanisms should therefore remain independent of language reasoning, planning, and VLA inference.
Action and Workspace Constraints
-
Runtime systems can enforce action magnitude and rate limits:
\[a_{\min}\leq a_t\leq a_{\max}\] \[\|a_t-a_{t-1}\|\leq\Delta_{\max}\] -
Similarly, robot motion can be constrained to an allowed workspace:
\[x_t\in\mathcal{W}_{\mathrm{safe}}\] -
Collision checking and dynamic-obstacle prediction should remain active even when actions originate from a learned policy.
Safety for Vision-Language-Action Models
- VLA systems introduce failure modes including:
- mis-grounding the target object,
- following unsafe instructions,
- hallucinating affordances,
- generating out-of-distribution actions,
- producing unsafe action chunks,
- failing to stop after the environment changes.
- producing unsafe action chunks,
- generating out-of-distribution actions,
- hallucinating affordances,
- following unsafe instructions,
- mis-grounding the target object,
- The Gemini Robotics API overview emphasizes that generative robotics models can make mistakes and should be deployed with appropriate independent safety mechanisms.
-
Action chunking should not imply uninterruptible execution. A VLA may predict:
\[A_t=(a_t,\ldots,a_{t+H-1})\]- while a higher-frequency runtime monitor continuously validates execution and interrupts the chunk when conditions change.
Success-Safety Gap
- Task completion and safe execution should be measured separately.
-
Let:
\[S_i\in\{0,1\}\]-
denote task success and:
\[V_i\in\{0,1\}\]- denote a safety violation.
-
-
Then:
\[P(S=1,V=1)\]- captures episodes in which the task succeeds through unsafe behavior.
- The SafeVLA benchmark evaluates safety violations alongside task performance for Vision-Language-Action policies.
-
Violation severity should also be represented:
\[\sigma(v)\]- because minor contact and severe collision should not contribute equally to safety assessment.
Reward Hacking and Specification Gaming
-
A proxy reward generally differs from the true objective:
\[R_{\mathrm{proxy}}\neq R_{\mathrm{true}}\] -
Optimization can therefore discover undesirable strategies that score well under the proxy. Physical AI systems should combine learned rewards with explicit constraints, outcome verification, adversarial testing, and human evaluation.
Safety During Post-Training
-
Safety data should be incorporated directly into post-training:
\[\mathcal{D} = \mathcal{D}_{\mathrm{task}} \cup \mathcal{D}_{\mathrm{safety}}\] - Safety examples can include:
- unsafe instructions,
- ambiguous instructions,
- hazardous scenes,
- infeasible tasks,
- constraint conflicts,
- correct refusals,
- recovery behavior.
- correct refusals,
- constraint conflicts,
- infeasible tasks,
- hazardous scenes,
- ambiguous instructions,
- unsafe instructions,
- The model should learn when to execute, clarify, refuse, or stop, while independent runtime safeguards remain in place.
Adversarial Safety Evaluation
- Safety evaluation should actively search for failures rather than sample only ordinary scenarios.
-
For scenario generator:
\[e\sim q_\psi(e)\]-
an adversarial objective can seek:
\[\psi^* = \arg\max_\psi \mathcal{L}_{\mathrm{safety}}( \pi_\theta,e )\]
-
- A useful development loop is:
Generate Attack
↓
Observe Failure
↓
Add Evaluation
↓
Post-Train
↓
Reevaluate
- Simulation and learned world models can generate large numbers of difficult scenarios before equivalent situations are attempted on hardware.
Formal Verification and Safety Envelopes
- Large learned policies are difficult to verify end to end, but smaller safety-critical components may be amenable to formal analysis.
-
Possible verification targets include:
\[\text{state machines}, \quad \text{collision envelopes}, \quad \text{fallback logic}, \quad \text{runtime monitors}\] - For autonomous driving, On a Formal Model of Safe and Scalable Self-driving Cars by Shalev-Shwartz et al. (2017) introduces Responsibility-Sensitive Safety as a formal model of safe driving behavior and proper response.
-
The emerging pattern is:
\[\boxed{ \text{Learned Capability} + \text{Engineered and Verifiable Safety Envelope}. }\]
Safety Cases and Operational Design Domains
-
Safety evidence should be assembled across:
\[\text{simulation}, \quad \text{formal analysis}, \quad \text{component testing}, \quad \text{fault injection}, \quad \text{real-world trials}, \quad \text{production monitoring}\] -
Evidence should be scoped to an Operational Design Domain:
\[\mathcal{O} = \{ \text{allowed operating conditions} \}\] -
When the system moves outside its validated domain, it should transition to degraded or safe operation rather than assume unchanged capability.
Fault Injection and Redundancy
- Reliability testing should intentionally inject failures such as:
- camera dropout,
- stale frames,
- network delay,
- incorrect localization,
- actuator degradation,
- compute restart.
- actuator degradation,
- incorrect localization,
- network delay,
- stale frames,
- camera dropout,
-
For failure mode \(f\), a useful quantity is:
\[P( \text{safe outcome}\mid f )\] - Redundancy is strongest when mechanisms are diverse rather than perfectly correlated. Neural perception, geometric checks, independent sensors, rule-based monitors, and external safety systems can therefore complement one another.
Production Safety Monitoring
- Safety continues after deployment.
-
Useful telemetry includes:
\[\text{interventions}, \quad \text{near misses}, \quad \text{constraint violations}, \quad \text{OOD events}, \quad \text{uncertainty}, \quad \text{hardware faults}\] -
A near-miss can be defined using a continuous risk score:
\[\tau_{\mathrm{warning}} < r_t < \tau_{\mathrm{failure}}\] - Near misses are especially valuable because they expose weaknesses before they become incidents.
- Each safety event should preserve enough context to reproduce the failure, including observations, selected actions, uncertainty, monitor outputs, triggered constraints, fallback actions, and model versions.
Safety Regression Testing
-
A protected safety suite can be represented as:
\[\mathcal{E}_{\mathrm{safety}} = \{e_1,\ldots,e_N\}\] - Every new model or software version should be tested against this suite, and newly discovered production failures should be converted into permanent regression cases.
-
This produces a safety data flywheel:
\[\boxed{ \text{Deploy} \rightarrow \text{Monitor} \rightarrow \text{Discover Risk} \rightarrow \text{Reproduce} \rightarrow \text{Generate Data} \rightarrow \text{Train} \rightarrow \text{Verify}. }\]
The Emerging Physical AI Safety Architecture
- A mature Physical AI safety stack increasingly resembles:
Human Instruction
↓
Semantic Safety
↓
Task Feasibility + Uncertainty
↓
Agentic Planning
↓
Plan / Tool Validation
↓
VLA Policy
↓
Predictive Safety Monitor
↓
Constraint / Barrier Layer
↓
Low-Level Safety Controller
↓
Functional Safety Runtime
↓
Actuators
↓
Physical World
External Safety Monitoring
↘
Telemetry + Incident Mining
↗
- The central principle is that learned intelligence should operate inside engineered safety boundaries. As Physical AI systems become more capable and general, the safety stack must scale with the space of behaviors they can attempt.
Engineering a Physical AI System
From Models to Systems
- A Physical AI model becomes useful only when embedded inside a system that can sense the world, maintain state, reason about goals, generate actions, execute them under real-time constraints, detect failures, and continuously collect feedback.
-
A production system can be viewed as a closed loop:
\[\boxed{ \text{Sense} \rightarrow \text{Understand} \rightarrow \text{Reason} \rightarrow \text{Plan} \rightarrow \text{Act} \rightarrow \text{Verify} \rightarrow \text{Learn}. }\] -
At time \(t\), the system receives multimodal observations:
\[o_t = \{ I_t, D_t, L_t, q_t, \dot q_t, F_t, m_t \}\]- where these may represent RGB images, depth, language, joint positions, joint velocities, force or torque measurements, and additional sensor measurements.
-
The autonomy stack transforms this information into actions:
\[a_t = \pi_\theta( o_{\leq t}, g, M_t )\]- where \(g\) is the task goal and \(M_t\) is memory or persistent world state.
-
The resulting action changes the physical environment:
\[s_{t+1} = F( s_t, a_t, \xi_t )\]- where \(\xi_t\) captures disturbances and unmodeled dynamics.
- Engineering Physical AI is largely the problem of making this loop reliable under real-world latency, uncertainty, compute, safety, and hardware constraints.
The End-to-End Physical AI Stack
-
A useful decomposition is:
\[\boxed{ \begin{array}{c} \text{Mission / Human Goal}\\ \downarrow\\ \text{Agentic Reasoning}\\ \downarrow\\ \text{Task Planning}\\ \downarrow\\ \text{World State + Memory}\\ \downarrow\\ \text{VLA / Skill Policy}\\ \downarrow\\ \text{Motion + Safety Layer}\\ \downarrow\\ \text{Real-Time Controller}\\ \downarrow\\ \text{Robot Hardware}\\ \downarrow\\ \text{Sensors}\\ \circlearrowleft \end{array} }\] - Surrounding this runtime stack are several infrastructure systems:
- data collection and storage,
- simulation and digital twins,
- training and post-training,
- evaluation and regression testing,
- observability and telemetry,
- deployment and model management.
- observability and telemetry,
- evaluation and regression testing,
- training and post-training,
- simulation and digital twins,
- data collection and storage,
- The complete system therefore resembles an AI platform combined with a robotics control stack.
Separating Decision Timescales
- Different parts of a Physical AI system operate at very different frequencies.
- A representative hierarchy might be:
- Layer Typical responsibility Relative frequency ——————- ——————————— ——————– Mission planner Long-horizon goal decomposition Very slow Agentic reasoning Planning, tool use, recovery Slow VLA policy Semantic motor commands Medium Motion policy Trajectory generation Fast Joint controller Torque/position control Very fast Safety monitor Collision/force constraints Very fast
- The exact rates depend on the platform, but the architectural principle is important.
- A large multimodal foundation model does not need to run at motor-control frequency.
-
Instead:
\[g_k = \pi_{\mathrm{reason}}( o_t,M_t,G )\]-
can produce a subgoal, while:
\[A_t = \pi_{\mathrm{VLA}}( o_t,g_k )\]-
produces an action chunk, and a low-level controller converts it into high-frequency commands:
\[u_\tau = \pi_{\mathrm{ctrl}}( q_\tau, \dot q_\tau, a_t )\]
-
-
- This hierarchical decomposition allows expensive reasoning and fast physical control to coexist.
Slow and Fast Loops
- A practical architecture often contains two nested loops.
-
The slow loop performs:
\[\text{Reason} \rightarrow \text{Plan} \rightarrow \text{Select Skill}\] -
The fast loop performs:
\[\text{Observe} \rightarrow \text{Control} \rightarrow \text{Observe}\] -
Formally:
\[g_k = \pi_{\mathrm{slow}}( o_t,M_t,G )\]-
and:
\[a_t = \pi_{\mathrm{fast}}( o_t,g_k )\]
-
-
The slow loop may operate only when:
\[\text{new goal} \lor \text{skill completed} \lor \text{failure detected} \lor \text{uncertainty high}\] - This event-triggered design avoids spending foundation-model inference on every control cycle.
Sensor Layer
- Physical AI begins with sensing.
-
A robot may combine:
\[\mathcal{O} = \{ \text{RGB}, \text{depth}, \text{LiDAR}, \text{proprioception}, \text{force}, \text{tactile}, \text{audio} \}\] -
Sensor streams differ in:
\[\text{frequency}, \quad \text{latency}, \quad \text{noise}, \quad \text{coordinate frame}\] -
Before model inference, the system must solve:
\[\text{calibration}, \quad \text{synchronization}, \quad \text{timestamping}, \quad \text{frame transforms}\] - These infrastructure details are often as important as model architecture because an otherwise capable policy can fail when observations are temporally or geometrically inconsistent.
Time Synchronization
-
Suppose camera observation is captured at:
\[t_c\]-
while robot state is measured at:
\[t_q\]
-
-
The model may incorrectly combine them if:
\[|t_c-t_q|\]- is large relative to the dynamics of the task.
-
A synchronized observation should instead approximate:
\[o_t = \{ I(t), q(t), \dot q(t), F(t) \}\] - Systems therefore maintain hardware or software timestamps and interpolate measurements where appropriate.
- For fast manipulation, even modest timing errors can appear to the model as incorrect object or end-effector motion.
Coordinate Frames
-
Physical systems operate across multiple coordinate systems:
\[\mathcal{F}_{\mathrm{camera}}, \quad \mathcal{F}_{\mathrm{robot}}, \quad \mathcal{F}_{\mathrm{world}}, \quad \mathcal{F}_{\mathrm{tool}}\] -
A transform:
\[{}^A T_B\]- maps coordinates from frame \(B\) into frame \(A\).
-
For example:
\[p_{\mathrm{world}} = {}^{W}T_C p_{\mathrm{camera}}\] - Errors in calibration propagate directly into manipulation.
- A VLA may semantically identify the correct object while the robot still misses it because the estimated camera-to-robot transform is inaccurate.
Perception Preprocessing
- Raw sensor streams are often transformed before reaching the policy.
-
The pipeline may include:
\[\text{undistortion} \rightarrow \text{cropping} \rightarrow \text{resizing} \rightarrow \text{normalization} \rightarrow \text{encoding}\] -
Depth may require:
\[\text{filtering}, \quad \text{registration}, \quad \text{hole filling}\] -
Proprioceptive state may be normalized:
\[\tilde q_i = \frac{ q_i-\mu_i }{ \sigma_i }\] -
Critically, the deployment preprocessing pipeline must match training:
\[P_{\mathrm{deploy}} \approx P_{\mathrm{train}}\] - Many deployment failures are effectively data-distribution bugs introduced before inference.
State Estimation
- Raw observations do not necessarily expose the underlying physical state.
-
A state estimator maintains:
\[\hat s_t = f( o_{\leq t}, a_{<t} )\] - The state may include:
robot_pose
joint_state
object_tracks
object_poses
human_tracks
scene_relations
task_state
uncertainty
-
For partially observed environments, the system may instead maintain a belief:
\[b_t = P( s_t \mid o_{\leq t}, a_{<t} )\] -
This state representation becomes the interface between perception and higher-level reasoning.
World State Versus Raw Context
-
A foundation model can consume recent observations directly:
\[c_t = [ o_{t-K}, \ldots, o_t ]\] -
However, long-running agents benefit from explicit structured state:
\[W_t = \{ \text{entities}, \text{relations}, \text{locations}, \text{task progress} \}\] -
For example:
mug_1:
location: countertop
state: empty
drawer_2:
state: open
task:
goal: place mug_1 in drawer_2
completed: [find_mug, open_drawer]
- This representation reduces the need for the model to reconstruct the entire world from raw context on every reasoning step.
Memory Architecture
-
Physical agents may require several forms of memory:
\[M_t = \{ M_{\mathrm{working}}, M_{\mathrm{episodic}}, M_{\mathrm{semantic}} \}\] - Working memory stores current task state.
-
Episodic memory stores previous interactions:
\[e_i = ( s_i, a_i, r_i, \text{outcome}_i )\] -
Semantic memory stores persistent knowledge such as:
\[\text{object properties}, \quad \text{environment maps}, \quad \text{user preferences}\] -
The reasoning model retrieves relevant memory:
\[M_t^* = \operatorname{Retrieve}( M,g,o_t )\] - This prevents context size from growing linearly with operating time.
Agentic Reasoning Layer
-
The reasoning layer converts a high-level mission:
\[G\]-
into executable subgoals:
\[P = [ g_1, g_2, \ldots, g_K ]\]
-
-
For example:
Goal: prepare the table
1. locate plates
2. retrieve plates
3. locate dining table
4. place one plate at each seat
5. verify placement
-
The planner may use:
\[\text{VLM reasoning}, \quad \text{world state}, \quad \text{memory}, \quad \text{tools}, \quad \text{skill metadata}\] -
The key engineering principle is that high-level reasoning should produce bounded requests to the physical control system rather than unrestricted actuator commands.
Skill Interfaces
- A clean interface between reasoning and control can expose skills such as:
navigate(target_pose)
grasp(object_id)
place(object_id, target_pose)
open(container_id)
close(container_id)
handover(object_id)
-
Each skill has:
\[\text{preconditions}, \quad \text{parameters}, \quad \text{termination conditions}, \quad \text{failure states}\] -
For skill \(\sigma\):
\[\sigma = ( P_\sigma, \pi_\sigma, T_\sigma, F_\sigma )\] -
This modular interface allows the planner to reason symbolically while learned policies handle continuous control.
VLA as a Motor Skill
- A VLA can itself serve as a general-purpose skill executor.
-
Given:
\[o_t\]-
and language-conditioned subgoal:
\[g_k\]-
the VLA generates:
\[A_t = \pi_{\mathrm{VLA}}( o_t,g_k )\]-
where:
\[A_t = [ a_t, \ldots, a_{t+H} ]\]
-
-
-
-
The architecture becomes:
\[\boxed{ \text{Agent} \rightarrow \text{Language Subgoal} \rightarrow \text{VLA} \rightarrow \text{Action Chunk}. }\] - This is attractive because the high-level planner does not need to understand embodiment-specific motor details.
Action Representation
- A deployment stack must define exactly what the model’s output means.
-
For a manipulator:
\[a_t = [ \Delta x, \Delta y, \Delta z, \Delta r_x, \Delta r_y, \Delta r_z, g ]\] -
Alternative action spaces include:
\[\text{joint positions}\] \[\text{joint velocities}\] \[\text{joint torques}\] \[\text{end-effector poses}\] -
For mobile robots:
\[a_t = [ v, \omega ]\] -
For autonomous vehicles:
\[a_t = [ \text{steering}, \text{acceleration}, \text{braking} ]\]- or a planned trajectory.
- Action-space design determines how much control responsibility lies inside the learned model versus conventional control software.
Relative Versus Absolute Actions
-
A policy may predict absolute pose:
\[x_{t+1}\] -
Alternatively, it may predict displacement:
\[\Delta x_t = x_{t+1}-x_t\] - Relative actions can generalize more naturally across workspace locations, while absolute actions may simplify some trajectory objectives.
-
Deployment must ensure that the action convention exactly matches training:
\[a_t^{\mathrm{model}} \rightarrow a_t^{\mathrm{robot}}\] - A sign, coordinate-frame, normalization, or unit mismatch can produce catastrophic behavior despite correct model inference.
Action Denormalization
-
Models commonly operate on normalized actions:
\[\tilde a_i = \frac{ a_i-\mu_i }{ \sigma_i }\] -
Deployment reconstructs:
\[a_i = \sigma_i \tilde a_i+\mu_i\] -
Other policies use quantile normalization:
\[\tilde a_i = 2 \frac{ a_i-q_{0.01} }{ q_{0.99}-q_{0.01} } -1\] - The normalization statistics are therefore part of the model artifact.
- Deploying weights without the corresponding action metadata can make the policy unusable.
Action Chunk Execution
-
Many modern robot policies predict multiple future actions:
\[A_t = [ a_t, a_{t+1}, \ldots, a_{t+H} ]\] - The simplest implementation executes the entire chunk.
- However, this increases open-loop duration.
-
A better runtime may execute only:
\[K<H\]- actions before requesting another prediction.
-
Thus:
\[\text{prediction horizon} \neq \text{execution horizon}\] - The system can trade inference cost against responsiveness.
Temporal Ensembling
- When overlapping action chunks are predicted, multiple predictions may exist for the same future action.
-
Suppose predictions for action at time \(t\) are:
\[\{ \hat a_t^{(t)}, \hat a_t^{(t-1)}, \ldots \}\] -
Temporal ensembling combines them:
\[a_t = \frac{ \sum_k w_k \hat a_t^{(k)} }{ \sum_k w_k }\] -
Weights may decay with prediction age:
\[w_k = e^{-\lambda k}\] - This can smooth trajectories and reduce sensitivity to individual model outputs.
Asynchronous Inference
-
Synchronous inference produces:
\[\text{observe} \rightarrow \text{infer} \rightarrow \text{act} \rightarrow \text{observe}\] - If inference takes \(T_I\) milliseconds, the robot may wait idle.
-
Asynchronous inference overlaps model computation with execution:
\[\boxed{ \begin{array}{cc} \text{Execute }A_t & \\ & \text{Infer }A_{t+1} \end{array} }\] - This reduces effective control latency.
-
A production runtime therefore commonly uses separate processes or threads for:
\[\text{sensing}, \quad \text{inference}, \quad \text{control}, \quad \text{safety}\]
Latency Budget
-
End-to-end latency can be decomposed as:
\[T_{\mathrm{total}} = T_{\mathrm{sense}} + T_{\mathrm{pre}} + T_{\mathrm{transfer}} + T_{\mathrm{model}} + T_{\mathrm{post}} + T_{\mathrm{control}}\] -
For stable closed-loop behavior:
\[T_{\mathrm{total}} < T_{\mathrm{budget}}\] - Reducing only model inference time may not solve the problem if sensor capture, image preprocessing, network transfer, or action execution dominates.
- Latency should therefore be profiled end to end.
Observation Age
- The policy does not necessarily act on the current world.
-
If observation was captured at:
\[t_o\]-
and action begins at:
\[t_a\]-
observation age is:
\[\Delta t = t_a-t_o\]
-
-
-
The policy effectively acts on:
\[s_{t-\Delta t}\]-
rather than:
\[s_t\]
-
- For dynamic tasks, observation age may be more important than raw model inference latency.
-
Systems can compensate using state prediction:
\[\hat s_t = F^{\Delta t}( s_{t-\Delta t}, a )\]
Edge Versus Cloud Inference
-
Physical AI computation can be distributed across:
\[\text{on-device}, \quad \text{edge}, \quad \text{cloud}\] -
On-device inference provides:
\[\text{low latency}, \quad \text{network independence}, \quad \text{privacy}\] -
Cloud inference provides:
\[\text{larger models}, \quad \text{elastic compute}, \quad \text{centralized updates}\] -
A hybrid architecture may place:
\[\text{motor control} \rightarrow \text{on-device}\] \[\text{VLA inference} \rightarrow \text{edge GPU}\] \[\text{long-horizon reasoning} \rightarrow \text{cloud}\] -
The partition should reflect both latency and failure tolerance.
NVIDIA Jetson and Edge Physical AI
- NVIDIA positions Jetson Thor as an edge compute platform for physical AI and robotics, supporting generative models, multimodal reasoning, and robot workloads within the robot rather than requiring cloud inference.
-
This type of architecture enables:
\[\text{foundation-model inference} + \text{robot perception} + \text{control software}\]- to share an edge computing platform.
- The engineering challenge becomes resource allocation between these workloads under power, memory, thermal, and real-time constraints.
Compute Scheduling
-
Suppose available compute is:
\[C_{\mathrm{total}}\] -
Processes require:
\[C_{\mathrm{VLA}}, \quad C_{\mathrm{perception}}, \quad C_{\mathrm{mapping}}, \quad C_{\mathrm{safety}}\] -
The system must maintain:
\[C_{\mathrm{VLA}} + C_{\mathrm{perception}} + C_{\mathrm{mapping}} + C_{\mathrm{safety}} \leq C_{\mathrm{total}}\] - Safety-critical processes should not be starved because a foundation model temporarily consumes all available compute.
- This motivates resource isolation, process priorities, memory reservations, and independent safety compute.
Model Optimization
-
Deployment may require reducing inference cost through:
\[\text{quantization}, \quad \text{distillation}, \quad \text{caching}, \quad \text{smaller vision encoders}, \quad \text{reduced image resolution}, \quad \text{action chunking}\] -
If original policy is:
\[\pi_{\theta_T}\]-
a smaller student can learn:
\[\pi_{\theta_S}\]-
using:
\[\mathcal{L}_{\mathrm{distill}} = D( \pi_{\theta_T}, \pi_{\theta_S} )\]
-
-
-
The correct optimization target is not simply throughput:
\[\max \frac{\text{tokens}}{\text{s}}\]- but closed-loop task performance under deployment constraints.
Robot Middleware
- Physical AI models must communicate with sensors and actuators through robotics middleware.
- ROS 2 provides a widely used communication framework based on nodes, topics, services, and actions.
- A deployment graph might contain:
/camera
↓
/perception
↓
/vla_policy
↓
/safety_filter
↓
/motion_controller
↓
/robot_driver
- Each node can run independently and communicate through typed messages.
- This modularity makes it easier to replace models without rewriting hardware interfaces.
ROS 2 Communication Patterns
-
Topics support streaming data:
\[\text{publisher} \rightarrow \text{subscriber}\] -
Services support request-response operations:
\[\text{request} \rightarrow \text{response}\] -
Actions support long-running operations with feedback:
\[\text{goal} \rightarrow \text{feedback} \rightarrow \text{result}\] -
For example:
camera frames → topic
detect object → service
navigate to room → action
- Selecting the correct communication primitive improves both reliability and observability.
Quality of Service
- Robot communication must specify behavior under delay or packet loss.
-
ROS 2 Quality of Service settings can control properties such as:
\[\text{reliability}, \quad \text{history}, \quad \text{durability}, \quad \text{deadline}\] - A camera stream may prefer fresh frames over guaranteed delivery of every old frame.
- A safety event may instead require reliable delivery.
-
Thus:
\[\text{QoS}_{\mathrm{camera}} \neq \text{QoS}_{\mathrm{safety}}\] - Middleware configuration becomes part of system correctness.
NVIDIA Isaac ROS
- NVIDIA Isaac ROS provides GPU-accelerated ROS 2 packages for robotics workloads such as perception, visual odometry, depth processing, and sensor pipelines.
-
The architectural pattern is:
\[\text{ROS 2 Interfaces} + \text{Accelerated GPU Kernels}\] - This allows compute-intensive perception to remain integrated with standard robotics middleware while exploiting hardware acceleration.
Motion Planning
- The VLA does not necessarily need to directly generate every joint command.
-
A high-level policy can produce target:
\[x_{\mathrm{goal}}\]-
while a motion planner computes:
\[q_{0:T}\]-
such that:
\[FK(q_T) \approx x_{\mathrm{goal}}\]
-
-
-
The planner solves:
\[q_{0:T}^* = \arg\min_{q_{0:T}} C(q_{0:T})\]-
subject to:
\[q_t \in \mathcal{Q}_{\mathrm{valid}}\]-
and:
\[\operatorname{CollisionFree}(q_t)=1\]
-
-
- This combines semantic intelligence with mature geometric planning.
MoveIt
- MoveIt provides motion planning, manipulation, collision checking, inverse kinematics, and trajectory execution capabilities commonly used with ROS.
-
A hybrid Physical AI system can therefore use:
\[\text{VLA} \rightarrow \text{semantic target}\]-
and:
\[\text{MoveIt} \rightarrow \text{collision-free trajectory}\]
-
- This reduces the need for the learned model to reproduce capabilities already available through reliable geometric tools.
Learned and Classical Control
-
A useful engineering principle is:
\[\boxed{ \text{learn what is difficult to specify} + \text{engineer what is easy to specify}. }\] -
Learned models are particularly useful for:
\[\text{semantic perception}, \quad \text{language grounding}, \quad \text{generalization}, \quad \text{complex manipulation}\] -
Classical systems remain effective for:
\[\text{kinematics}, \quad \text{collision checking}, \quad \text{trajectory tracking}, \quad \text{hard constraints}\] -
Production systems frequently combine both rather than replacing the entire robotics stack with one neural network.
Low-Level Control
-
Suppose desired joint trajectory is:
\[q_t^*\] -
A simple controller may compute:
\[\tau_t = K_p( q_t^*-q_t ) + K_d( \dot q_t^*-\dot q_t )\] -
More sophisticated controllers may use:
\[\text{inverse dynamics}, \quad \text{impedance control}, \quad \text{model predictive control}\] - The learned policy typically operates above this layer.
- Separating semantic decision-making from high-frequency stabilization improves reliability.
Impedance Control
- Manipulation often requires compliant physical interaction.
-
An impedance controller approximates:
\[F = K( x_d-x ) + D( \dot x_d-\dot x )\] - Rather than rigidly forcing the robot to a pose, the controller behaves like a virtual spring-damper system.
-
This is useful for:
\[\text{contact-rich manipulation}, \quad \text{assembly}, \quad \text{human interaction}\] - A VLA can generate desired motion while impedance control handles local physical compliance.
Safety Runtime
- Every action should pass through a safety layer before reaching hardware.
-
Let proposed action be:
\[a_t^{P}\] -
The safety runtime computes:
\[a_t^{S} = \mathcal{F}_{\mathrm{safety}}( a_t^P, s_t )\] -
Possible operations include:
\[\text{clipping}, \quad \text{rate limiting}, \quad \text{collision rejection}, \quad \text{workspace enforcement}, \quad \text{emergency stop}\] - The safety layer should remain active regardless of which model produced the action.
Failure Detection
-
Execution should produce explicit status:
\[z_t \in \{ \text{running}, \text{success}, \text{failure} \}\] -
A verifier estimates:
\[z_t = f_{\mathrm{verify}}( o_{\leq t}, g )\] -
Verification may use:
\[\text{vision}, \quad \text{force}, \quad \text{proprioception}, \quad \text{symbolic predicates}\] -
For example:
\[\operatorname{grasped}(o) = 1\]- might require both visual evidence and gripper-state evidence.
Recovery Orchestration
-
When failure occurs:
\[F_t = \operatorname{Diagnose}( o_t,M_t,h_t )\] -
The system selects:
\[r_t = \pi_{\mathrm{recover}}( F_t )\] -
Possible responses include:
retry
reobserve
change viewpoint
regrasp
replan
switch skill
ask human
abort
- Recovery should be explicit in the system architecture rather than an accidental consequence of repeatedly querying the same policy.
Retry Budgets
- Unbounded retry loops can make autonomous systems appear active while making no progress.
-
For skill \(\sigma\), define:
\[N_{\mathrm{retry}} \leq N_{\max}\] -
After:
\[N_{\mathrm{retry}} > N_{\max}\]-
escalate:
\[\text{local retry} \rightarrow \text{global replan} \rightarrow \text{human assistance}\]
-
- This converts failure handling into a predictable state machine.
Real-Time Versus Best-Effort Components
- Not every process needs hard real-time guarantees.
-
Hard or near-real-time components may include:
\[\text{motor control}, \quad \text{safety monitoring}, \quad \text{hardware interfaces}\] -
Best-effort components may include:
\[\text{language reasoning}, \quad \text{memory retrieval}, \quad \text{cloud analytics}\] - The architecture should ensure that delayed high-level inference cannot block the low-level controller.
- This is another reason to decouple foundation-model reasoning from the real-time control loop.
Watchdogs and Heartbeats
-
Each critical process can emit heartbeat:
\[h_i(t)\] -
If:
\[t-t_i^{\mathrm{last}} > \Delta_i\]- the supervisor marks component \(i\) unhealthy.
-
A system supervisor can then:
\[\text{restart process}, \quad \text{switch fallback}, \quad \text{stop robot}\] -
This handles failures such as model crashes, GPU resets, deadlocks, and middleware disconnects.
Observability
- A production Physical AI system should make every important decision reconstructable.
-
For each timestep, log:
\[e_t = \{ o_t, M_t, g_t, a_t, U_t, z_t \}\] - Useful telemetry includes:
sensor timestamps
model version
prompt/instruction
world state
predicted actions
executed actions
safety interventions
skill status
latency
hardware state
- Without this information, debugging closed-loop failures becomes extremely difficult.
Distributed Tracing
-
An action may pass through many services:
\[\text{camera} \rightarrow \text{preprocessing} \rightarrow \text{VLA} \rightarrow \text{safety} \rightarrow \text{controller}\] -
Assign a trace identifier:
\[ID_t\]- to the entire decision.
-
Each subsystem records:
\[( ID_t, t_{\mathrm{start}}, t_{\mathrm{end}}, \text{status} )\] - This allows latency and failures to be attributed to specific components.
- Observability practices from distributed systems therefore become directly relevant to robotics.
Data Logging
-
Every deployment can generate training data:
\[\tau_i = ( o_{0:T}, a_{0:T}, g, r, m )\] - However, storing every raw sensor stream indefinitely may be impractical.
-
The system can prioritize:
\[\text{failures}, \quad \text{interventions}, \quad \text{novel states}, \quad \text{low-confidence events}\] -
An event score might be:
\[S(e) = \alpha U(e) + \beta F(e) + \gamma N(e)\]-
where:
\[U=\text{uncertainty}, \quad F=\text{failure severity}, \quad N=\text{novelty}\]
-
- High-value episodes are retained for training and analysis.
Dataset Versioning
- Physical AI datasets evolve continuously.
-
A training run should record:
\[D_k = \{ \text{dataset version}, \text{filters}, \text{mixture weights}, \text{normalization}, \text{schema} \}\] -
Model artifact should record:
\[M_k = \{ \text{weights}, \text{architecture}, \text{training config}, D_k \}\] -
This enables reproducibility:
\[M_k \leftrightarrow D_k\] - Without dataset lineage, it becomes difficult to understand why one policy behaves differently from another.
Model Registry
- Each deployed model should have metadata such as:
model_id
training_dataset
checkpoint
robot embodiment
action normalization
camera configuration
supported tasks
evaluation results
safety approval
-
Deployment should verify compatibility:
\[\operatorname{Compatible}( M, R )=1\]- where \(R\) denotes robot configuration.
-
This prevents accidentally deploying a policy trained for one embodiment or camera geometry onto another.
Configuration Is Part of the Model
- Physical AI behavior depends on more than neural weights.
-
The deployable artifact may include:
\[\boxed{ \text{Weights} + \text{Prompts} + \text{Normalization} + \text{Camera Calibration} + \text{Action Mapping} + \text{Safety Limits}. }\] - Changing any of these can change behavior.
- They should therefore be versioned and tested together.
Simulation in the Development Loop
-
Before real deployment, the system can run inside:
\[\text{Isaac Sim}, \quad \text{MuJoCo}, \quad \text{other digital twins}\] -
The same interfaces should ideally be preserved:
\[\text{Policy API}_{\mathrm{sim}} \approx \text{Policy API}_{\mathrm{real}}\] -
This enables:
\[\text{train in simulation} \rightarrow \text{evaluate in simulation} \rightarrow \text{deploy on hardware}\]- without rewriting the policy integration layer.
-
NVIDIA Isaac Lab follows this pattern by providing a robot-learning framework built around Isaac Sim for reinforcement learning, imitation learning, and robot-policy development.
Hardware Abstraction
- A useful architecture defines a common robot interface:
observe()
execute(action)
reset()
stop()
get_state()
-
Simulation implements:
\[R_{\mathrm{sim}}\] -
Hardware implements:
\[R_{\mathrm{real}}\] -
Both satisfy:
\[R \in \mathcal{I}_{\mathrm{robot}}\] -
This allows the same evaluation and policy code to operate across simulated and physical environments.
Digital Twin Integration
-
A digital twin maintains a simulated state:
\[s_t^{\mathrm{twin}}\]-
aligned with:
\[s_t^{\mathrm{real}}\]
-
-
Telemetry updates:
\[s_t^{\mathrm{twin}} = U( s_{t-1}^{\mathrm{twin}}, o_t^{\mathrm{real}} )\] -
Candidate actions can then be tested:
\[a_t \rightarrow \text{Twin} \rightarrow \hat s_{t+1}\] -
This enables:
\[\text{predictive safety}, \quad \text{planning}, \quad \text{counterfactual evaluation}, \quad \text{failure replay}\] -
The twin becomes a bridge between runtime autonomy and simulation infrastructure.
Continuous Evaluation
-
Every candidate model should pass:
\[\boxed{ \begin{array}{c} \text{Offline Tests}\\ \downarrow\\ \text{Simulation}\\ \downarrow\\ \text{Regression Suite}\\ \downarrow\\ \text{Safety Suite}\\ \downarrow\\ \text{Hardware Bench}\\ \downarrow\\ \text{Controlled Deployment} \end{array} }\] -
The same benchmark suite should run automatically whenever:
\[\text{model}, \quad \text{prompt}, \quad \text{controller}, \quad \text{safety configuration}\]- changes.
-
Physical AI therefore benefits from CI/CD principles, but deployment gates must include physical performance and safety rather than software tests alone.
Hardware-in-the-Loop
- Hardware-in-the-loop testing places real components inside a simulated environment.
-
For example:
\[\text{real controller} + \text{simulated robot/environment}\] -
Or:
\[\text{real compute stack} + \text{simulated sensors}\] -
This allows engineers to test:
\[\text{latency}, \quad \text{drivers}, \quad \text{communication}, \quad \text{failure handling}\]- without exposing the full physical system to risk.
Shadow Mode
- A new policy can run without controlling the robot.
-
Production controller executes:
\[a_t^{P}\] -
Candidate policy predicts:
\[a_t^{C}\] -
Only:
\[a_t^{P}\]- is executed.
-
The system records:
\[( a_t^P, a_t^C, o_t )\] - This enables evaluation of a candidate model on real operational data before granting it physical control.
- For safety-critical systems, shadow deployment can be an important intermediate stage.
Canary Deployment
-
Instead of deploying model \(M_{k+1}\) everywhere, expose it to a small subset:
\[p_{\mathrm{canary}} \ll1\] -
Monitor:
\[\text{success}, \quad \text{interventions}, \quad \text{latency}, \quad \text{safety events}\] -
If regressions appear:
\[M_{k+1} \rightarrow M_k\] -
Progressive rollout reduces the blast radius of unexpected failures.
Rollback
- Every deployment should support rapid rollback.
-
Let:
\[M_{\mathrm{active}} = M_{k+1}\] -
If deployment gate fails:
\[G(M_{k+1})=0\]-
switch:
\[M_{\mathrm{active}} \leftarrow M_k\]
-
- Rollback should include the full configuration bundle, not just model weights.
- Otherwise an older model may run with incompatible preprocessing or action normalization.
Fleet Learning
-
When many robots operate simultaneously:
\[R_1,\ldots,R_N\]-
they generate distributed experience:
\[\mathcal{D} = \bigcup_{i=1}^{N} \mathcal{D}_i\]
-
-
Central infrastructure can:
\[\text{collect} \rightarrow \text{filter} \rightarrow \text{train} \rightarrow \text{evaluate} \rightarrow \text{redeploy}\] - This creates a fleet learning loop.
- The central challenge is selecting useful data rather than merely accumulating large volumes of repetitive trajectories.
Failure Mining
-
Suppose production generates:
\[\mathcal{F} = \{ f_1,\ldots,f_N \}\] -
Embed:
\[z_i = f_{\mathrm{embed}}( f_i )\] -
Cluster:
\[C = \operatorname{Cluster}( z_1,\ldots,z_N )\] -
Clusters might reveal:
transparent object failures
drawer grasp failures
low-light failures
language ambiguity
collision-avoidance interventions
-
These clusters directly inform:
\[\text{new training data}, \quad \text{new simulations}, \quad \text{new evaluation cases}\]
The Physical AI Data Flywheel
-
The engineering system ultimately forms:
\[\boxed{ \text{Deploy} \rightarrow \text{Collect} \rightarrow \text{Mine Failures} \rightarrow \text{Generate Data} \rightarrow \text{Train} \rightarrow \text{Evaluate} \rightarrow \text{Deploy}. }\] -
Real data provides:
\[\text{physical grounding}\] -
Simulation provides:
\[\text{scale}\] -
World models provide:
\[\text{counterfactual experience}\] -
Evaluation determines:
\[\text{what data should be collected next}\] -
The data engine, model-training system, simulator, and deployment infrastructure therefore become one connected system.
Multi-Robot Infrastructure
-
A fleet introduces additional services:
\[\text{device identity}, \quad \text{configuration management}, \quad \text{telemetry}, \quad \text{remote updates}, \quad \text{health monitoring}\] -
Each robot reports:
\[H_i(t) = \{ \text{battery}, \text{temperature}, \text{compute}, \text{sensors}, \text{model}, \text{task} \}\] -
A fleet manager can decide:
\[\text{dispatch}, \quad \text{pause}, \quad \text{update}, \quad \text{rollback}\] -
This makes fleet robotics partly a distributed-systems problem.
Connectivity Failure
- Cloud-connected robots must tolerate network loss.
-
Let connectivity be:
\[C_t \in \{ 0,1 \}\] -
If:
\[C_t=0\]-
the robot should retain enough local capability to:
\[\text{remain stable}, \quad \text{stop safely}, \quad \text{complete bounded local behavior}\]
-
- Cloud connectivity should not be required for emergency safety.
- This favors placing critical control and fallback logic on-device.
Power and Thermal Constraints
-
Edge compute is constrained by:
\[P_{\max}\]-
and thermal envelope:
\[T_{\max}\]
-
- Foundation-model inference can increase both.
-
A runtime may therefore adapt:
\[\text{model size}, \quad \text{inference frequency}, \quad \text{camera resolution}\]- based on available resources.
-
For example:
\[T>T_{\mathrm{threshold}} \Rightarrow f_{\mathrm{VLA}}\downarrow\] - Physical AI deployment must therefore optimize for energy and thermal efficiency in addition to model accuracy.
Deterministic Infrastructure Around Stochastic Models
-
Generative policies may be stochastic:
\[a_t \sim \pi_\theta( a_t\mid o_t,g )\] - The surrounding system should remain as deterministic and inspectable as practical.
-
For example:
\[\text{action limits}, \quad \text{timeouts}, \quad \text{retry budgets}, \quad \text{safety constraints}, \quad \text{state transitions}\]- can be explicitly defined.
- This creates predictable system behavior even when the underlying model is probabilistic.
State Machines Around Foundation Models
-
A task runtime might contain states:
\[Q = \{ \text{IDLE}, \text{PLAN}, \text{EXECUTE}, \text{VERIFY}, \text{RECOVER}, \text{STOP} \}\] -
Transitions are explicit:
\[\text{EXECUTE} \rightarrow \text{VERIFY}\]- when a skill terminates.
-
If verification fails:
\[\text{VERIFY} \rightarrow \text{RECOVER}\] -
If retry budget is exhausted:
\[\text{RECOVER} \rightarrow \text{STOP}\] -
Foundation models can make decisions within these states without controlling the overall lifecycle arbitrarily.
Interface Contracts
- Every subsystem should define a contract.
- For example, a VLA interface may specify:
Input:
RGB image: 224×224
robot state: 14D
instruction: UTF-8 string
Output:
10 actions
each action: 7D
coordinate frame: robot base
translation unit: meters
rotation: axis-angle
gripper range: [0, 1]
- These contracts prevent ambiguity between research code and production infrastructure.
- For Physical AI, tensor shapes alone are insufficient. Physical units and coordinate conventions are part of the API.
Testing Interfaces
-
Interface tests should verify:
\[\text{shape}, \quad \text{dtype}, \quad \text{range}, \quad \text{units}, \quad \text{frame}, \quad \text{timestamp}\] -
For example:
\[|\Delta x| < 0.05\text{ m}\]- may be a valid action limit.
-
An output of:
\[\Delta x=5\]- could indicate centimeters-versus-meters confusion.
-
Simple contract checks can prevent failures that model-level benchmarks never detect.
Reproducibility
-
Each episode should record enough metadata to reconstruct execution:
\[E_i = \{ M, C, D, S, R \}\]-
where:
\[M=\text{model version}\] \[C=\text{configuration}\] \[D=\text{deployment software}\] \[S=\text{random seeds}\] \[R=\text{robot hardware revision}\]
-
-
This becomes especially important when a fleet contains slightly different hardware generations.
Security
- Physical AI systems expose interfaces that can directly affect hardware.
- Security therefore intersects with safety.
-
Threat surfaces include:
\[\text{remote commands}, \quad \text{model updates}, \quad \text{sensor spoofing}, \quad \text{middleware}, \quad \text{credentials}\] - A compromised command path can bypass otherwise capable AI safety mechanisms.
-
Deployment architecture should therefore include:
\[\text{authentication}, \quad \text{authorization}, \quad \text{signed artifacts}, \quad \text{encrypted communication}, \quad \text{audit logs}\]
Signed Model Artifacts
-
Before deployment, verify:
\[\operatorname{VerifySignature}(M)=1\] - The robot should reject untrusted model packages.
- Similarly, configuration and firmware updates can be signed.
- This protects the integrity of the complete autonomy stack rather than only its neural weights.
An End-to-End Reference Architecture
-
A production Physical AI architecture can be summarized as:
\[\boxed{ \begin{array}{c} \text{Human / Mission Interface}\\ \downarrow\\ \text{Agentic Reasoning + Memory}\\ \downarrow\\ \text{Task Planner / Skill Router}\\ \downarrow\\ \text{World State}\\ \downarrow\\ \text{VLA / Learned Skill Policies}\\ \downarrow\\ \text{Motion Planning}\\ \downarrow\\ \text{Safety Runtime}\\ \downarrow\\ \text{Real-Time Control}\\ \downarrow\\ \text{Robot Hardware}\\ \uparrow\\ \text{Sensor Fusion + State Estimation} \end{array} }\] -
Around the runtime loop:
\[\boxed{ \begin{array}{ccccc} \text{Simulation} & \text{Data Engine} & \text{Training} & \text{Evaluation} & \text{Fleet Platform} \end{array} }\] -
All five continuously interact with deployment.
A Concrete Manipulation Example
- Consider:
-
Put the red mug in the upper cabinet.
- The system may execute:
Goal parsing
\[G = \operatorname{place}( \text{red mug}, \text{upper cabinet} )\]Perception
-
Detect:
\[o_{\mathrm{mug}}, \quad o_{\mathrm{cabinet}}\]
World-state update
red_mug:
location: counter
upper_cabinet:
state: closed
Planning
\[P = [ \text{open cabinet}, \text{grasp mug}, \text{place mug}, \text{close cabinet} ]\]Skill execution
- The VLA receives:
"Open the upper cabinet."
-
and predicts action chunk:
\[A_t\]
Safety
-
Every action is filtered:
\[A_t \rightarrow \mathcal{F}_{\mathrm{safety}}\]
Control
-
The robot controller converts safe targets into actuator commands:
\[A_t \rightarrow u_t\]
Verification
-
Check:
\[\operatorname{open}( \text{cabinet} )=1\]
Recovery
-
If verification fails:
\[\text{reobserve} \rightarrow \text{retry} \rightarrow \text{replan}\]
Logging
-
Store:
\[\{ o_t, A_t, \text{latency}, \text{outcome} \}\] -
The apparently simple instruction therefore activates almost every subsystem in the Physical AI stack.
Engineering Principles
- Several principles recur across successful Physical AI architectures:
- Separate timescales. High-level reasoning should not block high-frequency control.
- Use explicit interfaces. Units, coordinate frames, action semantics, and timing must be unambiguous.
- Keep safety independent. Learned policies should operate inside externally enforced safety constraints.
- Preserve closed-loop feedback. Long open-loop execution makes physical systems brittle.
- Make failures observable. Every important action should be reconstructable from telemetry.
- Design for recovery. Failure is expected in open-world environments, so retry, replan, escalation, and safe stop should be first-class behaviors.
- Unify simulation and reality. Policies should interact with common abstractions across simulated and physical environments.
- Treat data collection as part of deployment. Production experience should continuously improve training and evaluation.
- Version the complete system. Model weights alone do not define physical behavior.
- Prefer hybrid architectures. Foundation models, learned control, classical robotics, and deterministic safety systems solve different parts of the problem.
The Physical AI Production Flywheel
-
The complete engineering system can ultimately be expressed as:
\[\boxed{ \begin{array}{c} \text{Real-World Deployment}\\ \downarrow\\ \text{Telemetry + Experience}\\ \downarrow\\ \text{Failure Mining}\\ \downarrow\\ \text{Simulation + Synthetic Data}\\ \downarrow\\ \text{Training + Post-Training}\\ \downarrow\\ \text{Evaluation + Safety Validation}\\ \downarrow\\ \text{Progressive Deployment}\\ \circlearrowleft \end{array} }\] - This flywheel is the operational core of Physical AI.
-
The central engineering transition is from treating the foundation model as the product to treating it as one component of a larger autonomous system:
\[\boxed{ \text{Model} \rightarrow \text{Agent} \rightarrow \text{Robot System} \rightarrow \text{Learning Fleet}. }\] - The strongest Physical AI systems will therefore not necessarily be those with the largest standalone model. They will be systems in which perception, reasoning, learned control, classical robotics, simulation, safety, data infrastructure, evaluation, and deployment operate as one tightly integrated closed loop.
- The final section will examine Open Problems and Research Directions for Physical AI, including general-purpose robot foundation models, scalable data acquisition, cross-embodiment transfer, long-horizon autonomy, continual learning, world-model grounding, sample-efficient RL, dexterous manipulation, tactile intelligence, humanoid locomotion, multi-agent physical intelligence, safety, evaluation, and the path toward broadly capable embodied agents.
Open Problems and Research Directions
From Specialized Robots to General Physical Intelligence
- Physical AI has advanced from task-specific perception and control systems toward foundation models that can interpret language, perceive open-world scenes, transfer behaviors across tasks, and generate continuous actions. Systems such as RT-X, Gemini Robotics, and Isaac GR00T provide evidence that representations and behaviors can transfer across tasks and, to varying degrees, robot embodiments. Open X-Embodiment: Robotic Learning Datasets and RT-X Models by Open X-Embodiment Collaboration et al. (2023) demonstrated positive cross-robot transfer using data from 22 robot embodiments, while NVIDIA’s Isaac GR00T explicitly targets cross-embodiment generalist humanoid models.
- Yet current systems remain far from a generally capable physical agent that can:
- enter a previously unseen environment,
- understand arbitrary natural-language goals,
- acquire new physical skills efficiently,
- operate across different embodiments,
- execute tasks over hours rather than seconds,
- recognize and recover from unexpected failures,
- continually improve without forgetting previous skills,
- and do all of this safely.
- continually improve without forgetting previous skills,
- recognize and recover from unexpected failures,
- execute tasks over hours rather than seconds,
- operate across different embodiments,
- acquire new physical skills efficiently,
- understand arbitrary natural-language goals,
- enter a previously unseen environment,
-
The central research problem is therefore not simply:
\[\boxed{\text{How do we train a better robot policy?}}\]-
but increasingly:
\[\boxed{ \text{How do we build physical agents that continuously acquire, compose, and refine general-purpose physical intelligence?} }\]
-
The Robotics Data Bottleneck
-
Language models benefited from enormous naturally occurring corpora:
\[\mathcal{D}_{\mathrm{text}} \sim 10^{12+} \text{ tokens}\] -
Robot interaction data is fundamentally different. A useful trajectory may require:
\[\text{hardware} + \text{environment} + \text{operator} + \text{time} + \text{safety supervision}\] - Robot data is therefore expensive, slow, embodiment-specific, and difficult to reproduce.
- Open X-Embodiment was an important step toward aggregating heterogeneous robot experience, assembling data across 22 embodiments and hundreds of skills. However, its diversity also exposes the deeper problem: different robots possess different sensors, kinematics, action spaces, control frequencies, and capabilities.
-
The open question is whether robotics can discover an analogue of internet-scale pretraining:
\[\boxed{ \text{Internet-scale semantics} + \text{robot interaction} + \text{simulation} + \text{synthetic experience}. }\]
Scaling Robot Data
-
A plausible Physical AI scaling strategy combines several sources:
\[\mathcal{D} = \mathcal{D}_{\mathrm{robot}} \cup \mathcal{D}_{\mathrm{human}} \cup \mathcal{D}_{\mathrm{sim}} \cup \mathcal{D}_{\mathrm{synthetic}} \cup \mathcal{D}_{\mathrm{internet}}\] - Each contributes different information.
-
Robot demonstrations provide:
\[\text{embodiment-grounded action}\] -
Human video provides:
\[\text{semantic and behavioral diversity}\] -
Simulation provides:
\[\text{scalable interaction}\] -
Generative models provide:
\[\text{synthetic variation}\] -
Internet data provides:
\[\text{broad visual and semantic knowledge}\] - NVIDIA’s GR00T stack already follows this general direction by combining real captured robot data, synthetic data, and internet-scale video, with subsequent adaptation to specific embodiments and tasks.
- The unresolved problem is how to combine these sources without overwhelming the physically grounded signal with data that does not encode the robot’s actual action dynamics.
Learning Actions from Human Video
-
Human video exists at dramatically greater scale than robot demonstrations:
\[|\mathcal{D}_{\mathrm{human}}| \gg |\mathcal{D}_{\mathrm{robot}}|\] -
But human video typically lacks robot actions:
\[(I_t,I_{t+1}) \not\Rightarrow a_t^{\mathrm{robot}}\] -
The central challenge is therefore learning a mapping:
\[\text{Human Motion} \rightarrow \text{Embodiment-Independent Intent} \rightarrow \text{Robot Action}\] -
This requires disentangling:
\[\text{what happened}\]-
from:
\[\text{how a particular body executed it}\]
-
-
If successful, robot learning could exploit the enormous amount of instructional, egocentric, sports, manufacturing, cooking, and everyday human video already available.
Latent Actions
-
One promising direction is to infer latent actions from transitions:
\[z_t = f_{\mathrm{inv}}( o_t,o_{t+1} )\] - The latent variable \(z_t\) represents the transformation connecting two observations without requiring explicit robot control labels.
-
A robot-specific decoder can later map:
\[a_t^{(r)} = g_r( z_t,s_t )\] -
This factorization attempts to separate:
\[\boxed{ \text{task-level action semantics} }\]-
from:
\[\boxed{ \text{embodiment-specific actuation}. }\]
-
- The broader research question is whether a universal latent action representation can emerge across humans, manipulators, mobile robots, and humanoids.
The Embodiment Gap
- A policy trained on one robot does not automatically execute correctly on another.
-
Let robot embodiment be:
\[e = ( \mathcal{S}, \mathcal{A}, \mathcal{K}, \mathcal{D}, \mathcal{O} )\]- where these terms capture state, action space, kinematics, dynamics, and observation configuration.
-
Cross-embodiment transfer seeks:
\[\pi( a \mid o,g,e )\]-
that generalizes across:
\[e_1,e_2,\ldots,e_N\]
-
- A recent survey, The Embodiment Gap in Robot Foundation Models by Domae et al. (2026), argues that apparent foundation-model reuse can obscure substantial target-robot adaptation work and organizes the problem around shared semantics, shared robot data/interfaces, and learned correspondence between embodiments.
- This gap remains one of the defining challenges of robot foundation models.
Universal Action Representations
-
Different robots expose incompatible action spaces:
\[a_t^{(1)} \in \mathbb{R}^{7}\] \[a_t^{(2)} \in \mathbb{R}^{14}\] \[a_t^{(3)} \in \mathbb{R}^{30+}\] - Directly sharing output heads becomes difficult.
-
One direction is to learn an embodiment-independent representation:
\[z_t = \pi_{\mathrm{shared}}( o_t,g )\]-
followed by:
\[a_t^{(e)} = f_e( z_t )\]
-
-
Possible shared representations include:
\[\text{end-effector trajectories}\] \[\text{object-centric transformations}\] \[\text{keypoints}\] \[\text{latent motor primitives}\] - A sufficiently general interface might allow experience from one embodiment to improve another without requiring identical actuators.
Morphology-Conditioned Policies
-
Another direction conditions directly on robot morphology:
\[a_t = \pi_\theta( o_t,g,m_e )\]-
where:
\[m_e\]- describes embodiment.
-
-
This representation might contain:
\[\text{joint topology}, \quad \text{link geometry}, \quad \text{joint limits}, \quad \text{sensor layout}, \quad \text{actuator properties}\] - The policy therefore learns not one controller but a family of controllers indexed by body configuration.
-
A long-term goal would be:
\[\boxed{ \text{new robot description} + \text{small adaptation set} \rightarrow \text{competent policy}. }\]
Zero-Shot Embodiment Transfer
-
The strongest form of cross-embodiment generalization would require no demonstrations on the target robot:
\[\mathcal{D}_{e_{\mathrm{new}}} = \varnothing\] -
Given only:
\[\text{robot description} + \text{observations} + \text{task}\]- the model would infer how to operate the new embodiment.
- Current cross-embodiment models are still far from solving this generally. NVIDIA’s GR00T platform, for example, supports cross-embodiment learning but still emphasizes post-training for particular embodiments, tasks, and environments.
-
Reducing:
\[N_{\mathrm{adapt}}\]- toward zero is therefore a central scaling objective.
Long-Horizon Autonomy
- Many benchmark tasks remain short.
-
Real applications require:
\[T \gg 10^3\]-
control steps and potentially:
\[K \gg 10\]- semantic subtasks.
-
-
Suppose each subtask succeeds independently with probability:
\[p\] -
A sequence of \(K\) tasks succeeds with probability:
\[P_{\mathrm{success}} = p^K\] -
Even:
\[p=0.99\]-
produces:
\[0.99^{100} \approx 0.366\]
-
- This illustrates why apparently high single-skill success rates are insufficient for sustained autonomy.
Error Accumulation
-
Long-horizon agents face:
\[\text{perception drift}\] \[\text{state-estimation errors}\] \[\text{planning mistakes}\] \[\text{motor failures}\] \[\text{environment changes}\] -
A robust system must repeatedly restore alignment between its internal belief and physical reality:
\[\boxed{ \text{Act} \rightarrow \text{Observe} \rightarrow \text{Verify} \rightarrow \text{Correct}. }\] -
The research challenge is shifting from policies that can perform skills to agents that can maintain coherent progress despite hundreds of opportunities for failure.
Hierarchical Autonomy
- Long-horizon behavior will likely require temporal abstraction.
-
Let:
\[G\]- be the mission.
-
A planner produces:
\[g_1,g_2,\ldots,g_K\] -
Each subgoal invokes a lower-level policy:
\[\pi_{\mathrm{skill}}( a_t\mid o_t,g_k )\] -
The hierarchy becomes:
\[\boxed{ \text{Mission} \rightarrow \text{Plan} \rightarrow \text{Skill} \rightarrow \text{Control}. }\] - An open question is how much of this hierarchy should be explicitly engineered and how much should emerge inside a unified model.
Unified Models Versus Modular Agents
-
One research direction attempts to train:
\[\pi_\theta( a_t \mid o_{\leq t},g )\]- end to end.
-
Another decomposes the system:
\[\text{Reasoner} + \text{Planner} + \text{VLA} + \text{World Model} + \text{Controller}\] - Unified systems can share representations and potentially learn behaviors difficult to specify manually.
-
Modular systems provide:
\[\text{interpretability}, \quad \text{replaceability}, \quad \text{different timescales}, \quad \text{safety boundaries}\] - The likely frontier is not simply choosing one extreme, but discovering interfaces that permit end-to-end learning while preserving useful modular structure.
System 1 and System 2 Physical Intelligence
- Physical agents often require two qualitatively different forms of computation.
-
Fast behavior:
\[a_t = \pi_{\mathrm{reactive}}( o_t )\]- handles familiar situations.
-
Slow reasoning:
\[p_t = \pi_{\mathrm{reason}}( o_t,g,M_t )\]-
handles:
\[\text{novelty}, \quad \text{ambiguity}, \quad \text{failure}, \quad \text{long-horizon planning}\]
-
- The challenge is determining when slow reasoning is worth its latency.
-
An event-triggered architecture could invoke reasoning when:
\[U_t>\tau_U\]-
or:
\[P(\mathrm{failure})>\tau_F\]
-
- Efficient arbitration between fast motor intelligence and deliberate reasoning is likely to become a central Physical AI problem.
Continual Learning
-
A deployed robot should ideally become better with experience:
\[\pi_0 \rightarrow \pi_1 \rightarrow \cdots \rightarrow \pi_T\] - But learning new tasks can reduce performance on old ones.
-
Let:
\[P_i^{(t)}\]- be performance on task \(i\) after learning through task \(t\).
-
Catastrophic forgetting occurs when:
\[P_i^{(t+1)} < P_i^{(t)}\]- for previously learned tasks.
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning by Liu et al. (2023) formalized this problem for robot manipulation, emphasizing transfer of both declarative knowledge and procedural skills across lifelong task sequences.
Continual Post-Training
-
Rather than periodically retraining from scratch, future Physical AI systems may continuously incorporate:
\[\mathcal{D}_{t+1}\]-
into:
\[\pi_t\]
-
-
The objective becomes:
\[\pi_{t+1} = \operatorname{Update}( \pi_t, \mathcal{D}_{t+1} )\]-
while preserving:
\[P_{\mathrm{old}}\]-
and improving:
\[P_{\mathrm{new}}\]
-
-
-
This creates difficult tradeoffs between:
\[\text{plasticity}\]-
and:
\[\text{stability}\]
-
-
Replay, modular adapters, parameter isolation, distillation, and retrieval-based memory are possible ingredients, but lifelong robot learning remains unresolved.
Learning During Deployment
- An even stronger form of continual learning occurs directly during operation.
-
The robot encounters:
\[e_t\]-
detects:
\[\text{novelty}(e_t)\]-
collects:
\[\mathcal{D}(e_t)\]- and updates its behavior.
-
-
- The challenge is that unconstrained online learning can destabilize a previously validated system.
-
Future systems may therefore separate:
\[\boxed{ \text{fast memory adaptation} }\]-
from:
\[\boxed{ \text{slow parameter adaptation}. }\]
-
- A robot can first store new experience in memory, then incorporate it into model weights only after offline validation.
Few-Shot Skill Acquisition
-
Teaching a robot a new task should ideally require:
\[N \ll 100\]- demonstrations.
- Google DeepMind’s Gemini Robotics On-Device demonstrates the broader direction of efficient local robotics models with rapid task adaptation, but general few-shot physical skill acquisition remains unsolved.
-
A desired learning curve is:
\[P(N)\]-
with large:
\[\frac{dP}{dN}\]- for small \(N\).
-
-
Ultimately, physical agents should learn new behaviors through:
\[\text{demonstration}, \quad \text{language}, \quad \text{correction}, \quad \text{practice}\]
Learning from Corrections
- Full demonstrations are expensive.
-
A human may instead intervene only at failure:
\[a_t^{H} \neq a_t^{\pi}\] - These corrections provide high-information examples because they occur near the policy’s decision boundary.
-
A scalable learning interface may therefore combine:
\[\text{autonomous execution} + \text{sparse intervention}\] -
The robot gradually shifts from:
\[\text{human-controlled}\]-
toward:
\[\text{human-supervised autonomy}\]
-
Autonomous Data Collection
- Eventually, robots must help generate their own training data.
-
Let:
\[U(x)\]- measure expected learning value of experience \(x\).
-
The robot can select:
\[x^* = \arg\max_x U(x)\]- subject to safety constraints.
-
This turns deployment into active learning:
\[\boxed{ \text{What should the robot practice next?} }\]- rather than passively logging whatever happens.
- The challenge is defining useful novelty without encouraging unsafe exploration.
Self-Improvement Through Practice
- A physical agent could:
-
- identify a weak skill, 2. generate practice scenarios, 3. execute them in simulation, 4. transfer promising behaviors to hardware, 5. collect failures, 6. post-train, 7. reevaluate.
-
Formally:
\[\pi_{k+1} = \mathcal{T}( \pi_k, \mathcal{F}_k )\]-
where:
\[\mathcal{F}_k\]- contains failures discovered by the current policy.
-
- This resembles recursive improvement, but grounded in externally measurable physical outcomes.
World Models as Physical Reasoning Engines
-
A world model learns:
\[p_\phi( s_{t+1} \mid s_t,a_t )\] -
Its ultimate value is not merely video prediction but counterfactual reasoning:
\[\boxed{ \text{What happens if I do this?} }\] -
Future agents could evaluate many candidate futures internally before acting:
\[A^* = \arg\max_A R( M_\phi(s_t,A) )\] -
A 2026 survey of world models for Physical AI identifies long-horizon consistency, uncertainty calibration, physical constraint enforcement, distribution shift, and planner exploitation as continuing open problems.
Long-Horizon World-Model Consistency
-
Suppose one-step prediction error is:
\[\epsilon\] -
Repeated autoregressive rollout may produce:
\[E_H \gg H\epsilon\]- because prediction errors change future model inputs.
-
This causes:
\[\text{compounding error}\] \[\text{physically impossible trajectories}\] \[\text{planner exploitation}\] - A useful world model therefore needs more than perceptually plausible video.
- It must preserve decision-relevant physical structure across long horizons.
Physics-Aware World Models
-
Purely data-driven world models may violate:
\[\text{collision constraints}\] \[\text{rigidity}\] \[\text{conservation laws}\] \[\text{contact dynamics}\] -
One research direction combines learned dynamics:
\[M_\theta\]-
with structured physical priors:
\[\mathcal{P}\]
-
-
The objective might become:
\[\mathcal{L} = \mathcal{L}_{\mathrm{prediction}} + \lambda \mathcal{L}_{\mathrm{physics}}\] -
The deeper question is how much physics should be learned from data versus imposed structurally.
World Models as Simulators
-
Traditional simulation requires:
\[\text{geometry} + \text{materials} + \text{physics parameters}\] - Generative world models could instead learn simulation directly from observations.
-
This creates the possibility of:
\[\text{real video} \rightarrow \text{interactive world} \rightarrow \text{policy training}\] - The NeurIPS 2026 World Models in Physical AI workshop frames physical consistency, causal fidelity, downstream control utility, generative simulation, and scaling across embodiments as central unresolved questions.
-
The key benchmark should therefore not simply ask:
\[\text{Does the rollout look realistic?}\]-
but:
\[\boxed{ \text{Would planning inside this model produce good actions in reality?} }\]
-
Planner Exploitation
- A planner optimizing against an imperfect world model may discover trajectories that exploit model errors.
-
If:
\[\hat R(A)\]-
is predicted return and:
\[R(A)\]-
is true return, optimization may find:
\[A^* = \arg\max_A \hat R(A)\]-
precisely where:
\[|\hat R(A)-R(A)|\]- is largest.
-
-
-
- This is analogous to reward hacking, but the exploited object is the learned dynamics model.
-
Uncertainty-aware planning can penalize trajectories leaving well-supported regions:
\[J(A) = \hat R(A) - \lambda U(A)\] - Robustly coupling generative world models to optimization remains an important open problem.
Dexterous Manipulation
-
Human manipulation depends on:
\[\text{vision}, \quad \text{touch}, \quad \text{force}, \quad \text{proprioception}\] - Many current VLAs remain predominantly vision-driven.
-
This works well for free-space manipulation but becomes difficult for:
\[\text{insertion}\] \[\text{sliding}\] \[\text{deformable objects}\] \[\text{tool use}\] \[\text{in-hand manipulation}\] - These tasks depend on physical information that may be visually ambiguous or entirely invisible.
Tactile Intelligence
-
Let tactile observation be:
\[T_t\] -
A multimodal physical policy could become:
\[a_t = \pi_\theta( I_t, T_t, F_t, q_t, l )\] -
Touch reveals:
\[\text{contact}, \quad \text{slip}, \quad \text{pressure}, \quad \text{texture}, \quad \text{compliance}\] - The open challenge is creating tactile representations with the same degree of transfer achieved by modern visual representations.
- Unlike images, tactile sensors vary substantially in geometry and signal format, creating another embodiment problem.
Force-Aware Foundation Models
-
Future robot foundation models may explicitly reason about force:
\[F_t\]- rather than inferring contact only from images.
-
For manipulation:
\[a_t = \pi( I_t, q_t, F_t, g )\] -
This could enable more reliable:
\[\text{assembly}, \quad \text{tool use}, \quad \text{deformable manipulation}, \quad \text{human interaction}\] -
A general Physical AI model ultimately needs to understand not only what the world looks like, but how it responds to contact.
Deformable Objects
-
Rigid-object manipulation simplifies dynamics:
\[s_t \approx \{ \text{pose}, \text{velocity} \}\] -
A deformable object may require a much higher-dimensional state:
\[s_t = \{ x_1,\ldots,x_N \}\] -
Examples include:
\[\text{cloth}, \quad \text{cables}, \quad \text{food}, \quad \text{bags}, \quad \text{soft materials}\] - These objects make perception, simulation, state estimation, and control substantially harder.
- General-purpose household robots will require much stronger deformable-object intelligence than most current benchmarks demand.
Tool Use
- Tools change the effective action space of the body.
-
A robot holding tool \(k\) operates with:
\[\mathcal{A}^{(k)} \neq \mathcal{A}\] -
Tool use therefore requires reasoning about:
\[\text{affordances}, \quad \text{geometry}, \quad \text{force transmission}, \quad \text{task transformation}\] -
The deeper research challenge is compositional:
\[\boxed{ \text{known robot} + \text{novel tool} \rightarrow \text{new capability}. }\] - This provides a demanding test of whether Physical AI systems truly understand action consequences rather than merely reproducing familiar motion patterns.
Humanoid Whole-Body Intelligence
- Humanoids increase the dimensionality and coupling of control.
-
The action vector may include:
\[a_t = [ a_t^{\mathrm{legs}}, a_t^{\mathrm{torso}}, a_t^{\mathrm{arms}}, a_t^{\mathrm{hands}} ]\] -
The robot must simultaneously satisfy:
\[\text{balance}, \quad \text{locomotion}, \quad \text{manipulation}, \quad \text{collision avoidance}\] - A 2026 survey of behavior foundation models identifies whole-body control as a fundamental challenge because humanoids combine high-dimensional dynamics, underactuation, and diverse task requirements.
Locomotion and Manipulation Must Converge
-
Historically:
\[\text{locomotion}\]-
and:
\[\text{manipulation}\]- were often trained separately.
-
- General humanoid tasks require both.
- For example:
-
Walk to the shelf, crouch, pick up the box, stand, carry it, and place > it on the table.
-
The correct policy must coordinate the entire body:
\[\pi( a_t^{\mathrm{whole-body}} \mid o_t,g )\] -
A major research direction is therefore the transition from:
\[\text{arm-centric VLA}\]-
toward:
\[\boxed{ \text{whole-body VLA}. }\]
-
Dynamic Whole-Body Tasks
-
Humanoid intelligence becomes substantially harder when tasks require momentum and rapid contact transitions:
\[\text{running}, \quad \text{catching}, \quad \text{throwing}, \quad \text{climbing}\] - These behaviors leave little room for slow replanning.
-
They require:
\[\text{high-frequency prediction} + \text{fast feedback} + \text{accurate dynamics}\] - This creates a tension between large foundation models and the tight latency requirements of dynamic control.
- Hierarchical policies with fast motor primitives and slower semantic reasoning are one possible solution.
Sample-Efficient Reinforcement Learning
- Real-world RL remains expensive because every environment interaction has cost.
-
Suppose improvement requires:
\[N\]- transitions.
-
In simulation:
\[N \sim 10^8\]- may be feasible.
- On hardware, it may not be.
-
The goal is therefore:
\[\boxed{ \max \frac{ \Delta\text{Capability} }{ \text{Real Robot Interaction} }. }\] -
Promising ingredients include:
\[\text{offline RL}, \quad \text{world models}, \quad \text{simulation}, \quad \text{demonstration priors}, \quad \text{reward models}\]
Post-Training Foundation Policies with RL
-
A pretrained VLA already provides:
\[\pi_0\] -
Rather than learning from scratch, RL optimizes:
\[\pi^* = \arg\max_\pi J(\pi)\]-
subject to remaining near the pretrained behavior:
\[D( \pi,\pi_0 ) < \epsilon\]
-
-
A regularized objective can be written:
\[J( \pi ) = \mathbb{E}[R] - \beta D_{\mathrm{KL}}( \pi \| \pi_0 )\] - This mirrors post-training strategies used for language models but introduces physical-state distribution shift, sparse rewards, continuous control, and safety constraints.
- Developing scalable RL post-training for generalist VLAs remains a major frontier.
Reward Models for Physical AI
- Many robot tasks lack easily programmable rewards.
-
A learned reward model:
\[R_\phi( o_{0:T}, g )\]- can estimate whether the trajectory achieved the desired outcome.
-
Training data might contain comparisons:
\[\tau_i \succ \tau_j\] -
Then:
\[P( \tau_i\succ\tau_j ) = \sigma( R_\phi(\tau_i) - R_\phi(\tau_j) )\] - Reliable multimodal reward models could enable RL on tasks where hand-engineered rewards are impractical.
- The difficulty is preventing the policy from exploiting imperfections in the reward model.
Automatic Success Detection
-
Large-scale autonomous training requires answering:
\[\boxed{\text{Did the robot actually succeed?}}\]- without human labeling.
-
A success model estimates:
\[P( S=1 \mid o_{0:T}, g )\] - Errors are asymmetric.
- A false negative wastes useful experience.
- A false positive may teach the policy that a failed behavior was successful.
- Reliable outcome verification is therefore a critical infrastructure problem for self-improving Physical AI.
Open-World Perception
- Physical environments contain objects and situations absent from training.
-
The system needs to distinguish:
\[\text{recognized}, \quad \text{uncertain}, \quad \text{unknown}\] - Modern VLMs substantially improve open-vocabulary recognition, but semantic recognition does not imply physical understanding.
- Recognizing:
-
glass vase
-
does not automatically provide accurate estimates of:
\[\text{mass}, \quad \text{fragility}, \quad \text{friction}, \quad \text{center of mass}\]
-
- Physical AI therefore requires representations that connect semantic identity to interaction properties.
Learning Object Physics
-
For object \(o\), the agent ideally estimates:
\[\phi(o) = \{ m, \mu, c, k, f \}\]- representing properties such as mass, friction, compliance, stiffness, and fragility.
- Some can be inferred visually.
- Others require interaction.
-
The agent may therefore perform information-gathering actions:
\[a_t^* = \arg\max_a I( \phi; o_{t+1} \mid a )\] - This converts physical exploration into active system identification.
Active Perception
- Sometimes the correct action is not to manipulate the target but to gather more information.
-
Examples include:
\[\text{move camera}\] \[\text{change viewpoint}\] \[\text{touch object}\] \[\text{open container}\] -
The objective becomes:
\[a_t^* = \arg\max_a \left[ I( s; o_{t+1} ) - \lambda C(a) \right]\] - General agents need to reason about information value as part of action selection.
Spatial Intelligence
- Language models operate primarily over symbolic sequences.
-
Physical intelligence requires reasoning about:
\[SE(3)\]- geometry.
-
A robot must understand:
\[\text{left/right}, \quad \text{inside/outside}, \quad \text{support}, \quad \text{containment}, \quad \text{occlusion}, \quad \text{reachability}\] - Future foundation models may need stronger explicit 3D representations rather than relying entirely on 2D image tokens.
Persistent 3D World Models
-
A long-running agent should maintain a world representation:
\[W_t = f( W_{t-1}, o_t )\] -
The representation should remain consistent when:
\[\text{camera moves}\] \[\text{objects move}\] \[\text{objects disappear behind occluders}\] - This requires object permanence and spatial memory.
- A general household robot cannot repeatedly rediscover its entire environment from individual images.
Memory and Physical AI
- Physical memory should capture more than dialogue history.
-
A robot may need to remember:
\[\text{where objects are stored}\] \[\text{which doors are difficult to open}\] \[\text{which grasp worked previously}\] \[\text{which areas are restricted}\] -
Memory therefore becomes:
\[M = \text{semantic knowledge} + \text{spatial map} + \text{episodic experience} + \text{procedural experience}\] - The open problem is determining what should be stored, when it becomes stale, and how memory should influence action without propagating outdated beliefs.
Multi-Agent Physical Intelligence
-
Future environments may contain:
\[R_1,R_2,\ldots,R_N\]- robots.
-
Joint behavior requires:
\[\pi( a_t^1,\ldots,a_t^N \mid o_t^1,\ldots,o_t^N,g )\] -
Challenges include:
\[\text{task allocation}\] \[\text{communication}\] \[\text{collision avoidance}\] \[\text{shared world state}\] \[\text{coordination}\] -
The system must determine not only:
\[\text{What should I do?}\]-
but:
\[\boxed{ \text{Which agent should do what, and when?} }\]
-
Robot-to-Robot Knowledge Transfer
-
A fleet generates:
\[\mathcal{D} = \bigcup_i \mathcal{D}_i\] -
An ideal fleet-learning system allows experience from robot \(i\) to improve robot \(j\):
\[\Delta P_j( \mathcal{D}_i ) > 0\] - This becomes especially powerful when robots encounter different environments.
- A fleet could collectively explore the long tail of physical situations much faster than any single robot.
- Cross-embodiment transfer is therefore not only a pretraining problem but also a deployment-scale learning problem.
Human-Robot Collaboration
- General-purpose robots will increasingly operate around people rather than inside isolated cages.
-
They must infer:
\[\text{human intent}\] \[\text{attention}\] \[\text{social conventions}\] \[\text{personal space}\] \[\text{handover timing}\] - The objective is not simply collision avoidance.
- It is coordinated behavior that humans can understand and predict.
- This introduces research problems spanning robotics, multimodal AI, cognitive science, and human-computer interaction.
Legibility
- A physically optimal trajectory may not be the easiest for a human to interpret.
-
Let:
\[\tau^* = \arg\min_\tau C_{\mathrm{robot}}(\tau)\] -
A human-aware controller may instead optimize:
\[C(\tau) = C_{\mathrm{robot}}(\tau) + \lambda C_{\mathrm{human}}(\tau)\] - For example, a robot may exaggerate the beginning of a reach so that a nearby person can infer its intended target.
- General Physical AI therefore needs to optimize not only efficiency but predictability and legibility.
Personalized Physical Assistance
-
Assistive robots may need to learn preferences:
\[p_u( a\mid s )\]- for user \(u\).
-
Examples include:
\[\text{where objects should be placed}\] \[\text{preferred assistance style}\] \[\text{personal routines}\] - This creates a tension between personalization and generality.
-
A useful architecture may combine:
\[\pi_{\mathrm{foundation}}\]-
with:
\[M_u\]- or a lightweight user-specific adapter rather than training an entirely separate policy.
-
Evaluation Remains a Bottleneck
- Physical AI evaluation is expensive because success must often be measured through real interaction.
-
Benchmarks can easily overfit to:
\[\text{specific scenes}\] \[\text{specific robots}\] \[\text{short tasks}\] \[\text{known objects}\] -
The desired benchmark should instead estimate:
\[P( \text{success} \mid \text{novel task}, \text{novel scene}, \text{novel object}, \text{novel embodiment} )\] - No single current benchmark adequately captures all of these dimensions.
Measuring Generality
-
Suppose performance is:
\[P( t,e,s,o )\]-
where:
\[t=\text{task}, \quad e=\text{embodiment}, \quad s=\text{scene}, \quad o=\text{objects}\]
-
-
A generalist benchmark should sample broadly from:
\[\mathcal{T} \times \mathcal{E} \times \mathcal{S} \times \mathcal{O}\] -
The problem is combinatorial:
\[|\mathcal{T}| |\mathcal{E}| |\mathcal{S}| |\mathcal{O}|\]- can become enormous.
-
Efficiently estimating generality without executing every combination is therefore itself a research problem.
Measuring Adaptation Cost
- Success rate alone hides how much engineering was required to deploy a model.
-
Suppose two systems reach:
\[SR=90\%\] -
System A requires:
\[100\]- target-robot demonstrations.
-
System B requires:
\[10{,}000\] - They are not equally general.
-
Future benchmarks should therefore report quantities such as:
\[N_{\mathrm{demo}}\] \[T_{\mathrm{adapt}}\] \[C_{\mathrm{compute}}\] \[N_{\mathrm{parameters\ updated}}\] - This aligns with the embodiment-gap perspective that deployment effort should be measured explicitly rather than hidden behind final success rate.
Evaluating Emergent Physical Capabilities
- As models scale, they may acquire behaviors not directly represented in individual training tasks.
-
Potential examples include:
\[\text{novel tool use}\] \[\text{skill composition}\] \[\text{cross-object transfer}\] \[\text{physical problem solving}\] - Benchmarks should therefore contain open-ended tasks where memorizing demonstrations is insufficient.
-
The key question becomes:
\[\boxed{ \text{Can the system construct a new behavior from previously learned physical concepts?} }\]
Safety Under Generalization
- Generalization and safety interact.
-
A policy can behave safely on:
\[p_{\mathrm{train}}(s)\]-
while behaving unpredictably under:
\[p_{\mathrm{deploy}}(s)\]
-
-
The relevant objective is therefore:
\[P( \text{unsafe} \mid s\sim p_{\mathrm{OOD}} )\] - As Physical AI systems become more general, the number of possible operating situations increases.
- Safety systems must therefore scale with capability rather than merely validate a fixed task set.
Calibrated Uncertainty
- A capable robot should know when its predictions are unreliable.
-
If predicted confidence is:
\[c\]-
calibration requires approximately:
\[P( \text{correct} \mid c ) \approx c\]
-
- Poorly calibrated confidence is dangerous because an agent may behave decisively in unfamiliar states.
-
Future systems need uncertainty estimates covering:
\[\text{perception}, \quad \text{world models}, \quad \text{plans}, \quad \text{actions}, \quad \text{task success}\] - The problem is especially difficult for large generative policies whose internal token probabilities do not directly correspond to physical success probabilities.
Physical AI Interpretability
-
When a robot fails, engineers need to determine:
\[\text{what it perceived}\] \[\text{what it believed}\] \[\text{what it intended}\] \[\text{why it selected the action}\] -
A useful interpretability stack may expose:
\[\text{grounded objects}\] \[\text{predicted subgoals}\] \[\text{candidate plans}\] \[\text{predicted outcomes}\] \[\text{uncertainty}\] -
This is different from purely linguistic interpretability because the explanation must connect internal reasoning to physical state and motion.
Causal Physical Reasoning
- Correlation may be sufficient for familiar imitation.
-
Novel physical problem solving requires understanding intervention:
\[P( s_{t+1} \mid do(a_t) )\] - The agent must reason:
-
If I pull this object, what changes?
- rather than merely:
-
What usually appears after this image?
- World models trained explicitly for action-conditioned prediction offer one route toward such causal representations.
- The distinction between predictive realism and causal usefulness remains a major open question for Physical AI.
Compositional Physical Intelligence
- A general agent should combine known skills into new solutions.
-
Suppose it knows:
\[\sigma_1=\text{push}\] \[\sigma_2=\text{grasp}\] \[\sigma_3=\text{open}\] -
A novel task may require:
\[\sigma_{\mathrm{new}} = \sigma_1 \circ \sigma_3 \circ \sigma_2\] - This is the physical analogue of compositional reasoning.
- The difficulty is that skill composition changes the world state, so later actions must remain grounded in the consequences of earlier ones.
Open-Ended Skill Libraries
-
Rather than maintaining a fixed skill set:
\[\Sigma = \{ \sigma_1,\ldots,\sigma_K \}\]-
future agents may continually expand it:
\[\Sigma_{t+1} = \Sigma_t \cup \{ \sigma_{\mathrm{new}} \}\]
-
- A newly learned behavior could become callable by the planner.
-
This creates a form of procedural memory:
\[\boxed{ \text{Experience} \rightarrow \text{Skill} \rightarrow \text{Reusable Capability}. }\] - Determining when repeated behavior should be consolidated into a reusable skill is an open research problem.
Physical Reasoning Versus Language Reasoning
- A language model may know that a heavy object is harder to lift than a light object.
-
But executing that knowledge requires:
\[\text{force estimation}, \quad \text{grasp stability}, \quad \text{balance}, \quad \text{motor adaptation}\] - Physical intelligence therefore requires grounding semantic knowledge in continuous dynamics.
- A central question is whether physical reasoning will emerge primarily by scaling multimodal foundation models or whether dedicated representations and objectives will remain necessary.
Scaling Laws for Physical AI
-
Language modeling exhibits relatively smooth relationships among:
\[\text{data}, \quad \text{compute}, \quad \text{parameters}, \quad \text{loss}\] -
Robotics may have additional axes:
\[\text{embodiments}, \quad \text{tasks}, \quad \text{environments}, \quad \text{interaction hours}\] -
A Physical AI scaling law might resemble:
\[L = f( N, D, C, E, T )\]-
where:
\[N=\text{parameters}\] \[D=\text{data}\] \[C=\text{compute}\] \[E=\text{embodiment diversity}\] \[T=\text{task diversity}\]
-
-
Understanding which dimension produces the greatest transfer is essential because real robot data is far more expensive than text.
Data Quality Versus Scale
- Not every robot trajectory is equally valuable.
-
Let:
\[v_i = \text{learning value of trajectory }i\] -
Then effective dataset size may be closer to:
\[D_{\mathrm{eff}} = \sum_i v_i\]-
than simply:
\[D=N\]
-
-
High-value data may disproportionately contain:
\[\text{novel states}\] \[\text{recoveries}\] \[\text{rare objects}\] \[\text{failures}\] \[\text{difficult decisions}\] - The future of robotics scaling may therefore depend as much on intelligent data selection as raw collection volume.
Hardware-Software Co-Design
- Physical intelligence is constrained by the body.
-
Better:
\[\text{sensors}, \quad \text{actuators}, \quad \text{hands}, \quad \text{compute}\]- can make previously difficult learning problems easier.
- Conversely, better policies may allow simpler hardware.
-
The optimization problem is joint:
\[(\theta^*,h^*) = \arg\max_{\theta,h} J( \pi_\theta,h )\]- where \(h\) denotes hardware design.
- Physical AI therefore creates opportunities for AI-driven morphology and robot-design optimization rather than treating hardware as fixed.
Designing Robots for Learnability
-
Traditional robot design optimizes:
\[\text{precision}, \quad \text{payload}, \quad \text{cost}, \quad \text{reliability}\] -
Foundation-model-era robots may also optimize:
\[\text{observability}, \quad \text{data compatibility}, \quad \text{teleoperation}, \quad \text{cross-robot transfer}\] - For example, similar sensor configurations and action interfaces across a fleet can make shared learning substantially easier.
- Robot morphology may therefore increasingly be designed together with the learning system.
Efficient On-Device Foundation Models
- Cloud-scale models can provide strong reasoning, but physical control benefits from local inference.
-
The desired frontier is:
\[\max \text{Capability}\]-
subject to:
\[\text{Latency}\leq L_{\max}\] \[\text{Power}\leq P_{\max}\] \[\text{Memory}\leq M_{\max}\]
-
- Gemini Robotics On-Device demonstrates the direction toward locally executing general-purpose VLA models optimized for robotic devices.
- Research in distillation, quantization, sparse computation, caching, and hierarchical inference will remain central as model capabilities grow.
Adaptive Compute
- Not every physical state requires the same amount of reasoning.
-
For easy state:
\[C_t=C_{\mathrm{small}}\] -
For difficult state:
\[C_t=C_{\mathrm{large}}\] -
A routing model could choose:
\[C_t = f( U_t, D_t, R_t )\]- where uncertainty, difficulty, and risk determine computational effort.
-
This creates:
\[\boxed{ \text{easy state} \rightarrow \text{fast policy}\] \[\boxed{ \text{hard state} \rightarrow \text{deep reasoning}. }\] - Adaptive compute may be especially important for power-constrained robots.
From Robot Foundation Models to Physical Foundation Models
-
Current robot foundation models primarily target:
\[\text{manipulation}\]-
or:
\[\text{humanoid control}\]
-
-
A broader Physical AI model might span:
\[\text{robots}, \quad \text{vehicles}, \quad \text{drones}, \quad \text{industrial machines}\] -
The common representation would need to capture:
\[\text{objects}, \quad \text{geometry}, \quad \text{dynamics}, \quad \text{actions}, \quad \text{goals}\] -
The ambitious question is whether a single pretrained model can learn transferable principles of physical interaction across fundamentally different machines.
A Physical Foundation Model Objective
-
One possible abstraction combines:
\[\mathcal{L} = \lambda_V \mathcal{L}_{\mathrm{vision}} + \lambda_L \mathcal{L}_{\mathrm{language}} + \lambda_A \mathcal{L}_{\mathrm{action}} + \lambda_W \mathcal{L}_{\mathrm{world}} + \lambda_R \mathcal{L}_{\mathrm{reward}}\] -
The model jointly learns:
\[\text{what is present}\] \[\text{what is requested}\] \[\text{what action to take}\] \[\text{what will happen}\] \[\text{whether the outcome is desirable}\] - This would unify perception, action, prediction, and evaluation inside one pretrained representation.
- Whether such unification outperforms modular systems remains an open empirical question.
The Role of Simulation May Change
-
Today simulation is commonly used to:
\[\text{train policies}, \quad \text{generate data}, \quad \text{evaluate systems}\] - Future generative world models may make simulation increasingly learned.
-
Instead of manually constructing every environment:
\[E = \operatorname{EngineerScene}()\]-
one might generate:
\[E = G( \text{language}, \text{images}, \text{real-world logs} )\]
-
- This could make rare-event and long-tail generation dramatically cheaper.
- The remaining challenge is guaranteeing that generated scenarios remain physically meaningful enough for policy optimization.
Automated Curriculum Generation
-
A training system could maintain policy competence:
\[C(\tau)\] -
It then generates tasks near the learning frontier:
\[\tau^* = \arg\max_\tau \operatorname{LearningProgress}( \pi,\tau )\] - Tasks that are too easy provide little signal.
- Tasks that are impossible provide little useful feedback.
- The optimal curriculum therefore tracks the model’s expanding capabilities.
- Generative simulation and world models make this adaptive curriculum increasingly feasible.
Toward Autonomous Robot Research
-
The data flywheel can potentially become increasingly automated:
\[\boxed{ \begin{array}{c} \text{Evaluate Policy}\\ \downarrow\\ \text{Identify Weakness}\\ \downarrow\\ \text{Generate Scenario}\\ \downarrow\\ \text{Collect Experience}\\ \downarrow\\ \text{Post-Train}\\ \downarrow\\ \text{Reevaluate}\\ \circlearrowleft \end{array} }\] -
Humans specify:
\[\text{goals}, \quad \text{safety constraints}, \quad \text{evaluation criteria}\]-
while the system increasingly determines:
\[\text{what to practice next}\]
-
-
This would transform robotics development from manually curated training into a partially self-directed optimization process.
What Would a General Physical Agent Require?
- A broadly capable physical agent likely requires all of the following:
- General perception: recognize and track open-world objects, humans, geometry, and physical state.
- Language grounding: convert natural-language goals into physically meaningful objectives.
- Spatial intelligence: reason consistently about 3D geometry, reachability, containment, support, and motion.
- Physical prediction: anticipate how objects and agents respond to actions.
- General motor skills: manipulate, navigate, locomote, and interact across diverse environments.
- Cross-embodiment transfer: reuse knowledge across different bodies and action spaces.
- Long-horizon reasoning: decompose complex missions and maintain progress over extended execution.
- Memory: preserve spatial, semantic, episodic, and procedural information across time.
- Recovery: detect failures, diagnose causes, and construct corrective behavior.
- Continual learning: acquire new capabilities without destroying old ones.
- Uncertainty awareness: distinguish competence from unfamiliarity and seek assistance when needed.
- Safety: operate inside hard physical and semantic constraints.
- Efficient inference: execute foundation-model capabilities under real-time compute, power, and latency limits.
- No current model solves this entire stack.
A Possible Convergence
-
The different research directions in Physical AI increasingly point toward a common architecture:
\[\boxed{ \begin{array}{c} \text{Multimodal Foundation Model}\\ \downarrow\\ \text{Agentic Reasoning + Memory}\\ \downarrow\\ \text{World Model}\\ \downarrow\\ \text{Hierarchical VLA Policy}\\ \downarrow\\ \text{Whole-Body Controller}\\ \downarrow\\ \text{Physical World}\\ \uparrow\\ \text{Continuous Experience} \end{array} }\] -
Surrounding this runtime system is an automated learning loop:
\[\boxed{ \text{Real Data} + \text{Human Video} + \text{Simulation} + \text{World-Model Data} \rightarrow \text{Pretraining} \rightarrow \text{Post-Training} \rightarrow \text{Evaluation} \rightarrow \text{Deployment} \rightarrow \text{New Data}. }\]
From Foundation Models to Foundation Agents
-
The conceptual progression is:
\[\boxed{ \text{Foundation Model} \rightarrow \text{VLA} \rightarrow \text{Embodied Agent} \rightarrow \text{Continually Learning Physical Agent}. }\] - A foundation model provides representations.
- A VLA connects those representations to actions.
-
An embodied agent adds:
\[\text{memory}, \quad \text{planning}, \quad \text{tools}, \quad \text{verification}, \quad \text{recovery}\] -
A continually learning physical agent additionally closes the learning loop:
\[\text{experience} \rightarrow \text{improvement}\] - This last transition is likely to be one of the defining research problems of Physical AI.
The Long-Term Direction
-
The history of AI has repeatedly moved from manually engineered components toward systems that learn increasingly general representations:
\[\text{handcrafted vision} \rightarrow \text{vision foundation models}\] \[\text{task-specific NLP} \rightarrow \text{language foundation models}\] \[\text{task-specific robot policies} \rightarrow \text{robot foundation models}\] - Physical AI extends this trajectory into systems whose outputs change the environment itself.
- The central challenge is therefore larger than building a model that predicts good robot actions.
-
It is building systems that can acquire an increasingly general model of:
\[\boxed{ \text{Perception} + \text{Language} + \text{Space} + \text{Physics} + \text{Action} + \text{Consequences}. }\] -
The long-term objective can be expressed compactly as:
\[\boxed{ \text{Observe the physical world} \rightarrow \text{understand it} \rightarrow \text{predict it} \rightarrow \text{act within it} \rightarrow \text{learn from the result}. }\] - That closed loop distinguishes Physical AI from intelligence that exists primarily in the digital domain, and solving it requires progress not only in foundation models, but also in robotics, reinforcement learning, world models, simulation, control, hardware, safety, evaluation, and large-scale systems engineering.
References
Physical AI and Embodied Intelligence
- A Generalist Agent by Reed et al. (2022); DeepMind: A Generalist Agent
- PaLM-E: An Embodied Multimodal Language Model by Driess et al. (2023); PaLM-E project page
- VIMA: General Robot Manipulation with Multimodal Prompts by Jiang et al. (2023); VIMA project page
- R3M: A Universal Visual Representation for Robot Manipulation by Nair et al. (2022)
- BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning by Jang et al. (2022)
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning by Ross et al. (2011)
Language-Grounded Planning and Agentic Robotics
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances by Ahn et al. (2022); SayCan project page
- Inner Monologue: Embodied Reasoning through Planning with Language Models by Huang et al. (2022); Inner Monologue project page
- Code as Policies: Language Model Programs for Embodied Control by Liang et al. (2022); Code as Policies project page
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models by Huang et al. (2023); VoxPoser project page
Robot Transformers and Vision-Language-Action Models
- RT-1: Robotics Transformer for Real-World Control at Scale by Brohan et al. (2022); RT-1: Robotics Transformer for real-world control at scale
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control by Brohan et al. (2023); RT-2 project page
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models by Open X-Embodiment Collaboration et al. (2023); Open X-Embodiment repository
- OpenVLA: An Open-Source Vision-Language-Action Model by Kim et al. (2024); OpenVLA project page
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success by Kim et al. (2025); OpenVLA-OFT project page
- Octo: An Open-Source Generalist Robot Policy by Octo Model Team et al. (2024); Octo project page
Continuous and Generative Robot Policies
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware by Zhao et al. (2023); ALOHA project page
- Diffusion Policy: Visuomotor Policy Learning via Action Diffusion by Chi et al. (2023); Diffusion Policy project page
- π0: A Vision-Language-Action Flow Model for General Robot Control by Black et al. (2024); Physical Intelligence: π0
- π0.5: A Vision-Language-Action Model with Open-World Generalization by Black et al. (2025); Physical Intelligence: π0.5
- RDT-1B: A Diffusion Foundation Model for Bimanual Manipulation by Liu et al. (2024); RDT project page
NVIDIA Isaac GR00T and Humanoid Foundation Models
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots by Bjorck et al. (2025); NVIDIA Isaac GR00T
- GR00T N1.5
- FLARE: Robot Learning with Implicit World Modeling by NVIDIA Research et al. (2025)
- DreamGen: Unlocking Generalization in Robot Learning through Neural Trajectories by NVIDIA Research et al. (2025); DreamGen project page
- GR00T-Mimic
Google DeepMind Robotics
- Gemini Robotics brings AI into the physical world; Gemini Robotics
- Gemini Robotics On-Device brings AI to local robotic devices; Gemini Robotics On-Device
- Gemini Robotics-ER
Robot Learning and Post-Training
- Conservative Q-Learning for Offline Reinforcement Learning by Kumar et al. (2020)
- Offline Reinforcement Learning with Implicit Q-Learning by Kostrikov et al. (2021)
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning by Peng et al. (2019)
- Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning by Wang et al. (2022)
- DPPO: Diffusion Policy Policy Optimization by Ren et al. (2024)
World Models for Physical AI
- World Models by Ha and Schmidhuber (2018)
- Learning Latent Dynamics for Planning from Pixels by Hafner et al. (2018)
- Dream to Control: Learning Behaviors by Latent Imagination by Hafner et al. (2019)
- Mastering Diverse Domains through World Models by Hafner et al. (2023)
- TD-MPC2: Scalable, Robust World Models for Continuous Control by Hansen et al. (2023); TD-MPC2 project page
- Genie: Generative Interactive Environments by Bruce et al. (2024)
- Genie 2: A large-scale foundation world model
- Genie 3: A new frontier for world models
- World Models and JEPA primer
NVIDIA Cosmos and World Foundation Models
- Cosmos World Foundation Model Platform for Physical AI by NVIDIA et al. (2025); Cosmos Research; NVIDIA Cosmos
- Cosmos-Predict
Simulation, Sim-to-Real, and Digital Twins
- NVIDIA Isaac Sim; NVIDIA Isaac Lab
- NVIDIA Omniverse Replicator
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World by Tobin et al. (2017)
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots by Tan et al. (2018)
- Dynamics Randomization Revisited: A Case Study for Quadrupedal Locomotion by Muratore et al. (2020)
- Newton Physics
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering by Kerbl et al. (2023)
Autonomous Driving as Physical AI
- Planning-oriented Autonomous Driving by Hu et al. (2022)
- VAD: Vectorized Scene Representation for Efficient Autonomous Driving by Jiang et al. (2023)
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models by Tian et al. (2024)
- nuPlan: A Closed-Loop ML-Based Planning Benchmark for Autonomous Vehicles by Caesar et al. (2021); nuPlan
- Large Scale Interactive Motion Forecasting for Autonomous Driving: The Waymo Open Motion Dataset by Ettinger et al. (2021); Waymo Open Dataset
- NVIDIA Alpamayo
Physical AI Data Engines
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models by Open X-Embodiment Collaboration et al. (2023); Open X-Embodiment repository
- Ego4D: Around the World in 3,000 Hours of Egocentric Video by Grauman et al. (2021)
- R3M: A Universal Visual Representation for Robot Manipulation by Nair et al. (2022)
- GR00T-Mimic; DreamGen; GR00T N1.5
Robot-Learning Benchmarks and Evaluation
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks by Mees et al. (2022); CALVIN benchmark
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning by Liu et al. (2023); LIBERO project page
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation by Li et al. (2024); BEHAVIOR-1K
- Evaluating Real-World Robot Manipulation Policies in Simulation by Li et al. (2024); SIMPLER
- ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills by Gu et al. (2023); ManiSkill
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots by Nasiriany et al. (2024); RoboCasa project page
- RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins by Mu et al. (2024); RoboTwin project page
- Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks by Guruprasad et al. (2024)
Safety and Reliability for Physical AI
- A Comprehensive Survey on Safe Reinforcement Learning by García and Fernández (2015)
- Constrained Policy Optimization by Achiam et al. (2017)
- Safe Reinforcement Learning via Shielding by Alshiekh et al. (2018)
- Control Barrier Function Based Quadratic Programs for Safety Critical Systems by Ames et al. (2017)
- End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks by Cheng et al. (2019)
- Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning by Brunke et al. (2021)
Autonomous-Driving Safety
- On a Formal Model of Safe and Scalable Self-driving Cars by Shalev-Shwartz et al. (2017)
- Waymo Safety; Waymo Open Dataset Challenges
Lifelong Learning and Cross-Embodiment Transfer
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning by Liu et al. (2023)
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models by Open X-Embodiment Collaboration et al. (2023)
- The Embodiment Gap in Robot Foundation Models by Domae et al. (2026)
Physical AI Systems and Platforms
- NVIDIA Isaac; Isaac Sim; Isaac Lab; Isaac GR00T
- NVIDIA Cosmos; Cosmos World Foundation Model Platform for Physical AI
- Gemini Robotics
- NVIDIA Alpamayo
Physical AI Surveys and Research Directions
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- A Survey on Vision-Language-Action Models for Embodied AI
- World Models for Autonomous Driving: An Initial Survey
- The Embodiment Gap in Robot Foundation Models by Domae et al. (2026)
Citation
@article{Chadha2020DistilledPhysical AI,
title = {Physical AI},
author = {Chadha, Aman and Jain, Vinija},
journal = {Distilled AI},
year = {2020},
note = {\url{https://aman.ai}}
}