Chapter 3: Technology Value Chain — Components to Physical AI
The Question and the Thesis
The most useful question about China's robotics industry is not, “Who has built the best humanoid?” It is how motor torque becomes safe joint motion, how camera and touch signals become training data, and how a learned policy returns to the factory as a measurable quality outcome. A robot may be sold as one product, but it is assembled from actuators, reducers, power electronics, real-time control, an embodied platform, perception, end effectors, data capture, simulation, generalist policies, cell integration, and operating service. Reading a specification from only one layer hides where value is created and where it leaks away.
This chapter advances a straightforward thesis: advantage in Physical AI comes less from owning one celebrated component or model than from coupling the technology and operations value chain, measuring losses at its interfaces, and correcting those losses quickly. Vertical integration can accelerate that coupling, but owning every layer is not automatically economical. A modular supply chain offers specialization and choice, but it requires explicit contracts for timing, coordinate frames, action meaning, safety authority, diagnostic access, and change control.
After reading this chapter, you should be able to... - Explain the robotics value chain as layers of responsibility spanning actuation, control, bodies, perception, contact, data, simulation, models, integration, and operations. - Identify the load-bearing interfaces and hidden losses between adjacent layers. - Compare vertical integration with a specialist supplier stack from a manufacturing perspective. - Separate research results for generalist policies from evidence of factory reliability. - Design questions and evidence requests that turn a trade-show demonstration into a production diligence exercise.
3.1 A Map of Connections, Not a Parts Catalog
The robotics value chain is often compressed into upstream components, midstream robot bodies, and downstream applications. That three-box picture is useful for first orientation but too coarse for learning-enabled systems. Two robots can use similar motors and reducers while exposing different current-loop bandwidth, thermal estimates, collision logic, timestamps, camera calibration, and action semantics. Their policies will not be equally portable. Conversely, different bodies may reuse part of a data or model asset if observations and actions are translated explicitly and safety boundaries remain local to the target machine.
We therefore separate the chain into embodiment, sensing, data, simulation, models, control, integration, and operations. This is not an accounting taxonomy; it is a way to locate responsibility and interface risk. Octo and OpenVLA demonstrate the potential of generalist policies trained from diverse tasks and robot data [1] [2]. The system-integration scope of ISO 10218-2, meanwhile, reminds readers that responsibility for turning a robot product into an application and cell remains a distinct layer [3]. A model output and a safely producing system are not the same artifact.
This view also changes company classification. Calling a firm a “humanoid company” or an “AI company” says little about which boundaries it guarantees. A hardware supplier may also provide controllers, data tools, and remote diagnostics. A model provider may publish weights while leaving the observation adapter, action adapter, contact controller, and acceptance test to the customer. Revenue from component shipments, project integration, recurring software, and field service has different margins and scaling behavior even when it sits behind the same robot demonstration.
The bottleneck moves with maturity. A research team may lack useful data. A prototype team may be constrained by heat, wiring, or calibration. A factory may be constrained by reset time and fault isolation. A fleet may be constrained by spare parts and service coverage. A good value-chain map points to the interface limiting present throughput and reliability, not merely to the most visible technical layer.
3.2 Actuators and Controllers: Where Intelligence First Becomes Physics
The actuation layer includes motors, gearboxes, bearings, brakes, encoders, current or torque sensing, drives, and thermal design. Catalog peak torque is only an entry point. Continuous torque, peak duration, backlash, efficiency, reflected inertia, acoustic noise, cooling, shock tolerance, assembly variation, and life determine the usable action envelope. A humanoid knee and the wrist of an industrial arm do not need the same operating point. Precision insertion and dynamic locomotion do not reward the same transmission ratio or control margin.
The actuator-controller interface is more than a torque command. It includes command rate, measurement latency, filtering, units, sign conventions, coordinate axes, saturation, thermal derating, and the state entered after communication loss. A high-level policy might issue targets at 50 Hz while the current loop and state estimator operate much faster. Whether the upper layer commands position, velocity, torque, or an impedance target changes the force produced by an otherwise identical trajectory when contact occurs.
Controllers convert actuator capability into executable behavior. They enforce joint limits, self-collision constraints, ground contact, balance, force limits, speed limits, and acceleration limits. They form the boundary between a learned proposal and a physically guarded command. Even if a policy emits action chunks, lower layers must interpolate, track, reject, or stop them. Network delay, frame loss, encoder faults, and overheating should first be handled by deterministic control and protective functions, not improvised by a language-conditioned model.
For a manufacturing buyer, an interface manual can be more revealing than the number of degrees of freedom. The buyer should ask which real-time bus is used, whether synchronized raw state is available, which low-level modes are open, who owns a protective stop, and how a firmware update changes logs and learned policies. A cheap joint with closed diagnostics or unstable firmware can impose a larger integration and service cost than a more expensive, well-characterized module.
3.3 Bodies and End Effectors: Mechanics Defines the Action Space
A body is not a neutral container for intelligence. Link lengths, mass distribution, workspace, joint layout, cable routing, batteries, onboard compute, ingress protection, and maintenance access determine which actions a policy can select. Humanoids pursue access to human-designed spaces and whole-body coordination, while carrying costs in stability, energy, fall management, and high-dimensional control. Quadrupeds emphasize terrain mobility. Fixed arms prioritize repeatability and controlled cells. An arm on an autonomous mobile base combines movement with manipulation, but inherits localization error, vibration, wireless dependence, charging, and traffic management.
The end effector is exceptional because it makes direct process contact and often decides final success. Parallel grippers, suction, process tools, and multi-finger hands differ not only in degrees of freedom but in grasp robustness, cleanability, changeover time, cabling, pneumatics, and contact sensing. A dexterous hand is attractive because it may represent many objects and skills with one device. Yet a simple gripper, fixture, or dedicated tool can deliver better throughput and fewer failures for a stable manufacturing operation.
The body-tool interface requires more than a matching mechanical flange. Mass, inertia, center of gravity, tool frame, cable bend limits, power, communication, finger state, contact force, and post-change calibration must reach planning and safety layers. Replacing a hand changes reachability, collision geometry, stopping distance, wrist load, and the meaning of actions in recorded data. “Plug and play” should mean that these semantics propagate automatically or through a verifiable procedure—not merely that the bolt pattern fits.
Mechanical design also shapes data economics. A reliable, backdrivable arm can make teleoperation easier. A hand with frequent cable failures can contaminate a dataset with maintenance drift. A camera mast that moves under acceleration introduces a calibration distribution the model must absorb. Hardware choices therefore influence how much data is required, what failures appear, and how frequently evidence becomes stale.
3.4 Vision, Proprioception, and Touch: Observation Is About Time and Meaning
Perception spans external and wrist cameras, depth sensors, lidar, joint encoders, inertial sensors, force-torque sensors, and tactile arrays. Vision supplies broad object and scene context but is vulnerable to occlusion, reflection, lighting, and hidden contact state. Proprioception supplies fast joint state but does not directly reveal slip, cable snagging, or surface deformation. Force and touch reveal contact, yet their usefulness depends on placement, bandwidth, saturation, hysteresis, wear, and calibration.
The first sensor-fusion problem is often the clock rather than the representation. If images, robot state, commands, force, and tactile measurements arrive with different delay and sampling, a learner can associate the wrong cause with an outcome. A tactile peak recorded after the visible contact, or an original command stored after a safety controller modified it, creates a false action-result pair. Hardware timestamps, synchronization error, dropped-sample rules, and calibration versions deserve the same product scrutiny as nominal resolution.
The value of tactile sensing must be demonstrated through downstream outcomes such as slip recovery, regrasping, insertion, or damage reduction—not through taxel count alone. F-TAC Hand integrates distributed touch into a dexterous hand [7]. AnyTouch studies representations across multiple static and dynamic visuo-tactile sensors [8]. Almeida et al. analyze how tactile feedback from different finger and palm regions affects in-hand object reorientation [9]. These works illuminate different links of the tactile chain, but their bounded platforms and tasks do not establish industrial superiority for a sensor or company.
The practical question is not simply whether touch exists. It is which contact is measured, at what rate, in which frame and unit, and for which failure decision. Fingertip sensing may help with first contact and slip, while palm and lateral-finger sensing may matter for enveloping or multi-object grasps. Protective skin durability, wiring, contamination, replacement calibration, data bandwidth, and inference latency all belong in the maintenance plan. Without them, higher spatial resolution can reduce uptime rather than improve it.
3.5 Data: Contract Completeness Before Trajectory Count
The basic unit of robot data is an episode, not a video. An episode should join the instruction, sensor observations, robot state, operator or policy command, command actually executed by lower control, safety interventions, outcome, downstream quality, and hardware-software version. When those links break, scale becomes difficult to use. A team may retain a successful video without a failure reason, joint commands without the controller's saturated execution, or task completion without later inspection and rework.
The AgiBot World Colosseo paper reports more than one million trajectories across 217 tasks and five scenarios [4]. That scale demonstrates a substantial research data and management asset. It does not establish commercial deployments, customer count, factory uptime, or revenue. “Trajectory” can also vary by duration, duplication, success criteria, autonomous versus teleoperated collection, sensor configuration, and task difficulty. Raw counts should therefore not be turned into a cross-dataset or cross-company leaderboard.
A data interface must specify observation and action schemas together. Joint positions for a seven-degree-of-freedom arm, end-effector deltas, a binary gripper, and whole-body targets for a bimanual humanoid are not interchangeable vectors. Camera placement, intrinsics, coordinate frames, units, sample rates, missing channels, safety filters, and retargeting method should remain in metadata. Natural-language instructions need links to objects, constraints, termination conditions, and quality criteria if they are to support later training and evaluation.
Failures and recoveries are often more valuable than routine success. Slip, jamming, misperception, overheating, communication loss, operator takeover, protective stops, and resets reveal the weak layer. Deleting them while retaining clean trajectories teaches a policy to imitate normal operation while the organization repeats the same operating failure. Data rights also differ from model rights: contracts should distinguish customer raw data, derived features, site-specific updates, and a vendor's broader model.
3.6 Simulation and Digital Assets: Experience Generator and Interface Testbed
Simulation is not a single digital replica that replaces reality. It combines kinematics and collision, rigid and contact physics, sensor rendering, task logic, randomization, policy learning, and regression evaluation at different levels of speed and accuracy. Parallel physics and stable contact may dominate locomotion training. Camera and appearance models matter for bin picking. Tolerance, friction, and compliance matter for insertion. Fidelity must be selected against the decision being made.
The input interface is not complete when a CAD file loads. Mass and inertia, joint friction, motor limits, delays, sensor noise, materials, collision geometry, controller rates, and failure conditions require version control. Outputs should include comparable state, contact, energy, violations, and success definitions rather than only a reward curve. A defensible loop identifies parameters from real measurements, explores sensitivity in simulation, and updates correlation with bounded hardware trials.
Simulation can test supply-chain interfaces before scarce hardware is occupied. Teams can detect mismatches between a hand's URDF and driver, changed camera frames, an action adapter that violates joint limits, or a policy update that revives a frozen failure. Yet virtual success is not a safety or quality release. Cable behavior, contamination, deformable materials, contact friction, and unpredictable human interaction may remain outside the model.
Synthetic scale has its own operational cost. Scene assets become stale after fixture revisions. Randomization can spend compute on parameters unrelated to real defects. A simulator that produces many rollouts but cannot replay a named factory failure is weak evidence for that factory. The useful metric is not simulated episodes per hour alone, but how well simulation rejects policies around real failure neighborhoods before they reach equipment.
3.7 VLAs and Generalist Policies: Reusable Foundations, Not Finished Products
Vision-language-action models connect images and language instructions to robot actions. Their economic promise is that common representations and action structure learned from many tasks can be adapted with less site data than a policy trained from scratch. Octo presents an open generalist policy trained across heterogeneous robot data [1]. OpenVLA presents an open 7B VLA with multi-robot evaluation [2]. RDT-1B applies a one-billion-parameter diffusion foundation model to bimanual manipulation [6].
These open generalist policies provide reusable research baselines, but they do not by themselves establish factory reliability. Their results are bounded by disclosed tasks, embodiments, data distributions, trial counts, and success definitions [1] [2]. Production qualification adds cycle-time distributions, long-run drift, part-lot change, contamination, human interference, protective stops, recovery time, downstream inspection, model updates, and rollback.
Between a VLA and a robot sits an action adapter. It decides whether the model emits end-effector deltas or joint targets, how many time steps are chunked, how a gripper is represented, and how two arms or a whole body resolve competing priorities. An observation adapter maps camera order, resolution, language templates, robot state, and perhaps touch into model inputs. These adapters are not disposable glue code. They are product layers that determine reproducibility, timing, and safety.
Contact-rich tasks can expose the limits of vision and language alone. Across its five disclosed tasks, ForceVLA reports 60.5% average success versus 37.3% for the π0-base without-force baseline—an absolute 23.2-percentage-point difference—and success up to 80% in the disclosed setup [10]. This is evidence that force can change policy outcomes in bounded contact tasks. It is not a guarantee of equal improvement on every robot or factory process. Sensor placement and bandwidth, action representation, lower-level force control, and safety limits must align.
3.8 Cross-Embodiment Transfer: More Common Structure, Persistent Body Differences
Open X-Embodiment aggregates institutions, robots, and tasks to create a common basis for cross-embodiment learning [5]. Its important contribution is showing that visual, language, and task diversity can be shared beyond the data of one robot. But using data from another body is far from deploying immediately on any body.
Even with cross-embodiment data, action coordinates, timing, sensor placement, dynamics, contact, and safety interfaces remain embodiment-specific. Open X-Embodiment establishes shared data and modeling context [5], while RDT-1B studies transfer within a particular bimanual action structure [6]. Simulation-based transfer still requires scene reconstruction, retargeting, physical optimization, and real validation [11]. Cross-embodiment results become interpretable only when source and target sensors, degrees of freedom, end effectors, control rates, and evaluation conditions are disclosed.
Action normalization is a particularly risky abstraction. Padding two joint vectors to the same length does not equalize their meaning. End-effector deltas still depend on camera or base frames, units, rotation representation, time interval, and reachability. Bimanual action adds roles such as holding and manipulating; whole-body action adds balance and contact sequence. A useful converter exposes conversion failures and information loss rather than hiding them.
Cross-embodiment assets could reduce hardware lock-in if observation-action contracts and validation suites remain portable. A manufacturer might operate several bodies or change a supplier without discarding every learning asset. But lock-in can reappear if a foundation-model provider controls the data representation and update path. Manufacturers should retain raw logs, conversion code, baseline evaluation, model versions, and rollback options.
3.9 Integration and Operations: A Product Layer, Not the Last Ten Percent
Integration adapts a robot to a process. Fixtures, conveyors, PLCs, MES, quality inspection, safeguarding, networks, operator interfaces, cycle logic, and recovery procedures must work together. A policy that succeeds in a research demonstration is not a production system if it cannot handle a missing part, wrong SKU, worn tool, dirty camera, line-speed change, or transition to manual mode. Integration is not a final attachment to the model; it sends requirements back to every upstream layer.
Operations owns uptime, mean time to repair, spares, remote diagnostics, field service, security patches, model release, quality traceability, and operator training. This is where technical performance becomes recurring revenue or recurring cost. Completing a good demonstration once differs from maintaining target quality across multiple shifts for months. Data location during remote support and residual capability during a network outage also belong to system design.
The system-integration scope of ISO 10218-2 offers a foundational view of safety responsibility spanning the completed robot application and cell [3]. The 2011 edition has since been superseded, so real compliance in 2026 must follow applicable current standards, local law, and qualified safety review. We use the source to explain a responsibility boundary, not as legal advice or a current compliance checklist.
When operational feedback returns to data, the chain becomes a closed loop. Failure codes, takeovers, stop reasons, replaced parts, calibration changes, and quality outcomes should join the episode that triggers the next collection, simulation test, and model update. An organization that owns this loop can shorten recurrence time. If telemetry accumulates while quality and release authority remain elsewhere, more data does not necessarily produce faster operational learning.
3.10 Interface Map and Evidence Requests
The table below is not a vendor ranking. It is a procurement and diligence map of what crosses each boundary, what often disappears, and what evidence should be requested.
| Boundary | Load-bearing object | Common hidden loss | Evidence to request |
|---|---|---|---|
| Actuator → controller | Position, velocity, torque, current, temperature, fault state | Latency, saturation, backlash, thermal derating | Command/state rates, continuous-load test, link-loss behavior |
| Controller → body | Joint targets, whole-body constraints, collision and balance state | Model error, compliance, cables, falls | Real trajectory error, protective stop, limit and recovery tests |
| Body → end effector | Flange, power, communication, inertia, tool frame | Payload changes, wiring, post-change calibration | Tool-change procedure, TCP validation, collision-model update |
| Body/environment → perception | Images, depth, joints, IMU, force, touch | Occlusion, drift, clock error, contamination | Hardware timestamps, calibration history, aging and cleaning tests |
| Sensors → data | Synchronized observation and quality/failure labels | Missing channels, modified commands, selection bias | Episode schema, raw logs, retained failure and recovery |
| Data → simulation | Assets, parameters, distributions, real failures | Wrong mass/friction, simple contact, stale process | Identification basis, real-sim correlation, versions, sensitivity |
| Data/simulation → VLA | Training mixture, observation/action tokens, evaluation | Embodiment meaning loss, leakage, imbalance | Lineage, body-specific results, frozen tests, failure analysis |
| VLA → controller | Action chunks, uncertainty, stop or subgoal | Delay, infeasible target, ignored force/safety | Action adapter, constraints, watchdog, fallback |
| Integration → operations | Cell recipe, safety/quality gates, recovery | Responsibility gap, update regression, lock-in | FAT/SAT, uptime and MTTR, rollback, data/service terms |
| Operations → data loop | Failure, intervention, maintenance, quality, version | Success-only storage, inconsistent cause codes | Defect replay, named release authority, recurrence time |
The central question is not whether an API exists, but whether meaning survives the interface. A stream can deliver one hundred messages per second and still be misleading if it is not aligned with executed commands. A standard flange can still be unsafe if tool inertia and collision geometry are not updated. A model can claim support for multiple robots while leaving transfer cost unknowable if body-specific evaluation and adapters are absent.
This map also helps allocate responsibility. A supplier does not need to own every downstream decision, but it should state the boundary of its guarantee. The integrator should identify what it transforms and validates. The manufacturer should name the person or function that can reject a release. Undefined ownership, not merely low component performance, is a recurrent cause of slow diagnosis.
3.11 Vertical Integration: Benefits and Costs
A vertically integrated company may design actuators, the body, data tools, models, and operating infrastructure. Its greatest advantage is interface iteration. If a policy repeatedly overheats a joint, hardware, controls, and learning teams can revise cooling, the action envelope, and data weighting together. If touch reveals a critical failure, hand mechanics, sensor placement, collection tools, and model inputs can be co-designed. Shared hardware-software versioning may improve reproduction and diagnosis.
Broader ownership also means broader burden. The company must fund capital equipment, inventory, certification, manufacturing yield, field service, and research simultaneously. An internal component is not guaranteed to remain the market's best price-performance choice. Closed interfaces can speed an early demonstration yet raise the cost of connecting customer PLCs, cameras, tools, and data systems. A supply disruption or design defect in one internal layer can delay the whole product.
A specialist stack lets a buyer select proven suppliers and replace layers as technology changes. Standard industrial arms, established safety control, specialist vision, process-specific grippers, and open policies can draw on mature yield and maintenance experience. The integrator or manufacturer, however, must coordinate frames, timing, drivers, versions, safety, and data rights. When several suppliers assign a fault to one another, mean time to repair rises.
Vertical integration should therefore not be scored as an unconditional virtue. Ask which uncertainty it removes. Integration is valuable when rapid co-design and product iteration dominate. Modularity may matter more when the process is stable and the factory already owns significant automation assets. A strong architecture does not own everything; it names the owner of each stressed boundary and the cost of replacing it.
3.12 The Industrial-Robot Baseline
The baseline for Physical AI is not always fully manual work. Industrial arms, dedicated automation, PLCs, vision inspection, fixtures, spares, and maintenance already provide repeatability, predictable cycle times, and structured safety integration. A humanoid or VLA must show more than one successful task. Its incremental value should appear where established automation struggles: variation, redeployment, high-mix changeover, unstructured recovery, or work in human-designed spaces.
Industrial qualification examines distributions and costs, not a single average success rate. It counts hourly throughput, defects, rework, planned and unplanned stops, shift variation, tool change, recalibration, energy, operator interventions, safety events, mean time to repair, and lifecycle cost. A learning policy can improve average completion while creating rare failures that stop a line or damage product. Conversely, a slower system can be valuable if it sharply reduces changeover and dangerous manual work.
An early role for general-purpose bodies may be to fill automation gaps rather than replace every fixed arm. Frequent task changes, low volumes, human-built environments, imperfect part alignment, and maintenance assistance across tools are plausible candidates. Even there, hybrid architecture is often practical: fixtures and deterministic control retain the stable portion of the process, while learning addresses variable perception, planning, and recovery.
Comparison must use the same time window and denominator. A ten-minute demonstration cannot be compared with a shift-level industrial baseline. Success should include inspection, rework, resets, and operator time. Capital cost should include integration and tooling; operating cost should include service, compute, communications, consumables, and data work. Only then does “general-purpose” become an economic statement rather than a visual impression.
3.13 Three Manufacturing Interpretations
Consider first a connector-insertion cell. Cameras locate the component and socket. The arm and wrist approach. Force and touch detect first contact and jamming. A VLA may propose an approach or recovery from instruction and scene context, but lower-level impedance control, force limits, and stop conditions execute contact. Data must connect the command with actual force, insertion depth, final inspection, and rework. Simulation can perturb pose, tolerance, and friction to screen policies; real tests still establish deformation and damage.
Second, mixed-bin picking exposes the perception-body-gripper interface. Accurate recognition does not help if the arm cannot reach or the gripper cannot establish contact. Suction pressure, jaw position, wrist load, and post-place inspection need a shared episode identity so failures can be attributed. When a new SKU arrives, the team must decide across the chain whether to update the model, change gripper pads, move the camera, revise lighting, or alter the bin presentation.
Third, humanoid material handling tests locomotion, balance, bimanual grasp, and operations together. A video of a robot moving a box does not cover floor friction, thresholds, batteries, human traffic, drop risk, safety distance, manual recovery, or charging routes. Stable grasping can still have low operating value if the system reboots frequently or requires constant remote support. A slower platform might remain useful in a high-mix facility if it moves among existing stations with little facility change and fast recovery.
All three cases ask the same question: where is a failure observed, who corrects it, and what blocks the next release? Before a proof of concept, manufacturers should define a failure taxonomy, industrial baseline, episode schema, safety authority, and quality approver. A successful demonstration then has a clear evidence scope; a failed one points to an interface rather than generating a vague request for “more AI.”
3.14 Limits of Public Evidence and a Diligence Procedure
The evidence in this chapter is bounded by public papers and an official standard. Papers select robots, tasks, datasets, baselines, and trial counts to study a method. Official product statements can describe scope and issuer metrics but may be selective. Placing numbers from different years and definitions in one table does not make them comparable. Research data volume, parameter count, and demo success must not be converted into shipment, customer retention, factory uptime, or revenue claims.
Results from AgiBot World, OpenVLA, Octo, RDT-1B, ForceVLA, and tactile research retain the boundary of their disclosed setup. Without independent reproduction, long-duration operation, safety and quality audit, and service evidence, industrial generalization should remain qualified. A paper project, company, or commercial product with a similar name should not be treated as the same entity without explicit affiliation and product evidence.
A practical diligence process has four stages. First, request interface documents and representative raw logs. Second, reproduce the supplier's claimed demonstration under its disclosed conditions. Third, introduce buyer-selected edge conditions and failures in a bounded test. Fourth, compare throughput, quality, interventions, stops, recovery, and cost against the existing industrial baseline using common definitions and duration. Model releases should pass frozen replay and bounded hardware gates, with a tested rollback path.
This procedure is not designed to discount Chinese companies. It separates research progress, product capability, manufacturing capability, and operating evidence fairly in a fast-moving supply chain. A strong supplier need not claim leadership in every layer. It should explain which boundaries it guarantees, which a partner owns, and which remain unqualified.
3.15 Bridge to Chapter 4: Testing the Interfaces in General-Purpose Bodies
This value-chain map becomes the reading framework for Chapter 4 on humanoids and quadrupeds. The next chapter will look beyond degrees of freedom and body shape to ask how actuator modules, whole-body control, sensors and hands, development tools, data capture, model integration, safety, and service are combined in real platforms. It will separate public demonstrations from repeatable product capability, prototype output from stable manufacturing, and research ecosystems from field operations.
The questions now become concrete. Which action interfaces does the body expose? Which sensors are recorded with synchronized timing? Who controls falls and contact? How do model and safety configurations change when a hand or tool changes? Can customers own their data and evaluation? How quickly can a failed machine return to work? A “general-purpose body” becomes a meaningful Physical-AI platform only when it can answer those interface questions with evidence.
References
- Ghosh, D., et al. (2024). Octo: An Open-Source Generalist Robot Policy. Robotics: Science and Systems. #55 Terry commentary.
- Kim, M. J., et al. (2024). OpenVLA: An Open-Source Vision-Language-Action Model. arXiv:2406.09246.
- ISO. (2011). ISO 10218-2:2011 Robots and robotic devices — Safety requirements for industrial robots — Part 2: Robot systems and integration. International Standard.
- AgiBot-World Contributors et al. (2025). AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems. arXiv:2503.06669.
- Open X-Embodiment Collaboration et al. (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models. arXiv:2310.08864.
- Liu, S., et al. (2024). RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation. arXiv:2410.07864.
- Zhao, Z., et al. (2025). Embedding high-resolution touch across robotic hands enables adaptive human-like grasping. Nature Machine Intelligence. #39 Terry commentary.
- Feng, R., et al. (2025). AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-Tactile Sensors. ICLR 2025.
- Almeida, J. D., Falotico, E., Laschi, C., & Santos-Victor, J. (2025). The Role of Touch: Towards Optimal Tactile Sensing Distribution in Anthropomorphic Hands for Dexterous In-Hand Manipulation. ICNSC 2025; arXiv:2509.14984. #41 Terry commentary.
- Yu, J., et al. (2025). ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation. NeurIPS 2025; arXiv:2505.22159. #1 Terry commentary.
- Dan, P., Kedia, K., Chao, A., Pan, E. W., & Choudhury, S. (2025). X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real. arXiv:2505.07096.
- WALL-OSS authors (2026). Wall-OSS-0.5 Technical Report. arXiv:2605.30877.
- TeleDex authors (2026). TeleDex: Accessible Dexterous Teleoperation. arXiv:2603.17065.
- Tactile Genesis authors (2026). Tactile Genesis: Exploring Tactile Sensors at Scale for Learning Dexterous Tasks. arXiv:2606.22332.
- Robots Need More Than VLA authors (2026). Robots Need More than VLA and World Models. arXiv:2606.06556.
- RoboTacDex authors (2026). RoboTacDex: A Dexterous Visual-Tactile-Action Dataset for Humanoid Manipulation. arXiv:2606.31836.
- DexVerse authors (2026). DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation. arXiv:2607.08751.
- DexterCap authors (2026). DexterCap: An Affordable and Automated System for Capturing Dexterous Hand-Object Manipulation. arXiv:2601.05844.
- DexTele authors (2026). DexTele: A Dual-Arm Dexterous Teleoperation System Based on Motion Retargeting and Adaptive Force Control. arXiv:2607.05883.
- ContactWorld authors (2026). ContactWorld: What Matters in Vision-Tactile World Models for Contact-Rich Manipulation. arXiv:2606.13877.
- Booster Robotics research (2026). Booster Lab: A Data-Centric Pipeline for Learning Deployable Humanoid Locomotion Policies. arXiv:2606.27813.
- Visuo-tactile pose authors (2025). Visuo-Tactile Object Pose Estimation for Contact-Rich Manipulation. arXiv:2503.19893.
- OSMO authors (2025). OSMO: Open-Sourced Multi-Modal Tactile Glove for Robotic Manipulation. arXiv:2512.08920. Terry #18.
- Galaxea AI (2025). G0: A Generalist Robot Policy from Galaxea AI. arXiv:2509.00576.
- DexForce paper authors (2025). DexForce: Extracting Force-Informed Actions from Human Demonstrations. arXiv:2501.10356.
- Booster Robotics research (2025). Booster Gym: End-to-End Humanoid Locomotion Framework. arXiv:2506.15132.
- Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, Guanya Shi (2024). OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning. arXiv:2406.08858.
- Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wetzstein, Chelsea Finn (2024). HumanPlus: Humanoid Shadowing and Imitation from Humans. arXiv:2406.10454.
- DexGrip authors (2024). DexGrip: Robot Dexterous Grasping with Tactile Sensing. arXiv:2411.17124.
- BestMan authors (2024). BestMan: A Mobile Manipulator Platform for Embodied AI. arXiv:2410.13407.
- Tony Z. Zhao et al. (2023). Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware. arXiv:2304.13705.
- Z.-H. Yin et al. (2023). Rotating without Seeing: Towards In-hand Dexterity through Touch. arXiv:2303.10880.
- Ankur Handa et al. (2023). DeXtreme: Transfer of Agile In-hand Manipulation from Simulation to Reality. arXiv:2210.13702.
- Danny Driess et al. (2023). PaLM-E: An Embodied Multimodal Language Model. arXiv:2303.03378.
- C. Daniel Freeman et al. (2021). Brax: A Differentiable Physics Engine for Large Scale Rigid Body Simulation. arXiv:2106.13281.
- Fanbo Xiang et al. (2020). SAPIEN: A SimulAted Part-based Interactive ENvironment. DOI:10.1109/cvpr42600.2020.01104.
- Subramanian Sundaram et al. (2019). Learning the signatures of the human grasp using a scalable tactile glove (STAG). DOI:10.1038/s41586-019-1234-z.
- Fabio Ramos et al. (2019). BayesSim: Adaptive Domain Randomization Via Probabilistic Inference for Robotics Simulators. DOI:10.15607/rss.2019.xv.029.
- OpenAI (2019). Learning Dexterous In-Hand Manipulation. DOI:10.1177/0278364919887447.
- Aude Billard et al. (2019). Trends and Challenges in Robot Manipulation. DOI:10.1126/science.aat8414.
- Patrick M. Wensing et al. (2018). Linear Matrix Inequalities for Physically Consistent Inertial Parameter Identification. DOI:10.1109/tro.2017.2769099.
- Cosimo Della Santina et al. (2018). Toward dexterous manipulation with augmented adaptive synergies: The Pisa/IIT SoftHand 2. DOI:10.1109/tro.2018.2830407.
- Silvio Traversaro et al. (2016). Identification of Fully Physical Consistent Inertial Parameters using Optimization on Manifolds. DOI:10.1109/iros.2016.7759801.
- Sergey Levine et al. (2016). End-to-End Training of Deep Visuomotor Policies. Journal of Machine Learning Research.
- David Coleman et al. (2014). Reducing the barrier to entry of complex robotic software: a MoveIt! case study. Journal of Software Engineering for Robotics.
- Stéphane Ross et al. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS 2011.
- Rachid Boulic et al. (2008). Real-Time Motion Retargeting to Highly Redundant Robots. DOI:10.1109/tro.2008.2001367.
- Jan Swevers et al. (1997). Excitation Trajectories for the Identification of Base Inertial Parameters of Robots. DOI:10.1177/027836499701600306.
- Richard M. Murray et al. (1994). A Mathematical Introduction to Robotic Manipulation. CRC Press.
- Bruno Siciliano et al. (1991). A General Framework for Managing Multiple Tasks in Highly Redundant Robotic Systems. DOI:10.1109/icar.1991.240390.
- Maxime Gautier (1991). A Direct and Efficient Method for Identifying the Minimum Dynamic Parameters of Robots. DOI:10.1109/robot.1991.131896.
- Christopher G. Atkeson et al. (1986). Estimation of Inertial Parameters of Manipulator Loads and Links. DOI:10.1177/027836498600500306.