Part III: Contact and Intelligence

Chapter 10: VLAs and World Models — Connecting Intelligence to Factories

Written: 2026-08-18 Last updated: 2026-08-18

Overview

The question in this chapter is: What must a vision-language-action model or world model disclose before a factory can treat it as an operating capability rather than a demonstration? A model can recognize an object, describe a plan, predict a video, or emit an action sequence. None of those outputs alone proves that the robot can meet cycle time, protect a fixture, recover from a misgrasp, preserve quality records, or stop safely beside a worker.

The thesis is that a generalist model remains bounded by its observations, action representation, training data, latency, embodiment, downstream controller, and recovery system. The model is one layer in a physical decision chain. Its useful generality is not the number of tasks shown in a montage, but the size of the operating envelope in which it can propose actions that conventional control and safety layers can execute, verify, and reverse.

Moxian Technology, DexRobot, and AI² Robotics are the three locked deep profiles for this chapter [1] [6] [10]. They occupy different positions in the stack: tactile observation infrastructure, a dexterous-hand and data-acquisition ecosystem, and an integrated robot-plus-VLA platform. Comparing them on one evidence template reveals which model claims are actually disclosed and which interfaces remain outside public view.

After reading this chapter... - You can specify a VLA using observation, action, data, latency, controller, recovery, openness, and embodiment fields. - You can distinguish a predictive world model from a visually plausible video generator and a digital replay system. - You can compare the three profiles without turning organizer or issuer claims into independent benchmarks. - You can identify where private data, closed weights, undisclosed reset rules, and hidden conventional controllers limit reproducibility. - You can design a staged factory evaluation that starts from an industrial automation baseline and promotes a model only after measurable gates.

10.1 The Industrial Baseline Is the Control Group

A VLA enters a factory that already has a control architecture. A programmable logic controller sequences equipment. Robot programs define poses, speeds, and zones. A safety controller monitors doors, scanners, enabling devices, and emergency stops. Machine vision estimates part identity or pose. A manufacturing execution system dispatches work and records genealogy. Operators clear jams, verify first articles, and own recovery procedures. This baseline may be inflexible, but its boundaries are legible.

The correct question is not whether a learned policy looks more intelligent than a point-to-point program. It is whether the learned layer improves the cell after every existing function is counted. If a VLA handles product variation but requires more resets, slower motion, extra compute, manual confirmation, and unlogged corrections, its apparent flexibility may not improve output. If it shortens changeover while preserving quality, safety, and uptime, it may create value even when its raw cycle time is slower.

Industrial automation separates authority. A planner may choose a skill, a motion planner may generate a path, a servo loop tracks commands, and a certified safety function can stop the machine independently. Learned end-to-end systems blur these layers conceptually, but deployment still needs an answer to who owns joint limits, collision checking, force limits, protective stops, interlocks, and final release of motion.

SayCan illustrated a modular precedent: language relevance ranks skills while learned affordance values estimate executability [17]. Code as Policies generates programs that call supplied perception and control interfaces [18]. These systems differ from modern VLAs, yet they teach a durable lesson. General reasoning becomes physical behavior only through bounded executable contracts.

10.2 A Model Contract: From Camera Frame to Verified Outcome

A useful model card begins with the observation contract. Which cameras are used, at what resolution and rate? Are joint state, force, torque, tactile arrays, gripper width, mobile-base pose, and task history synchronized? Is the language instruction free text or a controlled task code? Does the model see the work-order variant, fixture state, tool identity, and quality result? Missing observations can make physically different states appear identical.

The action contract is equally important. RT-1 discretizes robot actions into tokens and executes them closed loop [19]. RT-2 co-trains web-scale vision-language knowledge with robot action tokens [20]. Diffusion Policy represents multimodal action sequences and uses receding-horizon execution [21]. π₀ uses flow matching for continuous action generation across multiple embodiments [24]. These labels conceal practical differences: joint position versus end-effector delta, horizon, frequency, coordinate frame, gripper semantics, and the controller that interpolates predictions.

Latency has at least four parts: sensor acquisition, inference, communication, and downstream execution. A reported model frequency may exclude exposure time, network transit, action buffering, collision checks, and actuator response. Long-horizon planning can tolerate seconds; visual servoing may not. Contact stabilization may require a local loop far faster than a foundation model. Evaluation should report the end-to-end age of an observation when the actuator applies its command.

Recovery is part of the model contract. A trial may end in success, safe abort, operator correction, reset, damaged part, or ambiguous completion. The system must identify failure, select a recovery, preserve evidence, and decide whether the part can continue. Counting only completed runs hides intervention cost and the risk of accepting a bad state.

Contract field Minimum disclosure Factory question
Observation modalities, rates, synchronization, history can the model distinguish each safety- or quality-relevant state?
Action coordinates, horizon, rate, units, gripper semantics what exactly reaches the robot controller?
Data sources, licenses, task and robot mixture, split lineage what supports transfer to this cell?
Latency end-to-end distribution, compute location, fallback does stale inference destabilize motion or recovery?
Controller boundary planner, collision, force, servo, safety owners which layer can override the model?
Recovery detection, retry budget, intervention, reset what happens after the first error?
Openness weights, code, data, interfaces, license can a buyer reproduce or maintain capability?

10.3 Deep Profile I — Moxian Technology

The Korean rendering is 목시안 테크놀로지, the English name used here is Moxian Technology, and the Chinese legal and brand name is 墨现科技(东莞)有限公司 / 墨现科技. The official brochure says it was founded in 2021, and its contact page places it in the Songshan Lake area of Dongguan, Guangdong [2] [3]. This is a tactile-sensor company at the observation boundary, not a disclosed VLA developer.

Flagship offerings include flexible pressure sensors, robot electronic skin, collision-sensing products, and application-specific pressure arrays. The robot-skin page describes finger modules and broader hand coverage intended to locate contact and provide force feedback [4]. The technology combines sensing materials, structures, manufacturing processes, acquisition electronics, and application software. Public sources do not establish a proprietary world model, action decoder, general robot policy, or end-to-end VLA.

The WRC organizer describes a Moxian tactile data-collection glove with 368 taxels and a density of 9 taxels per square centimetre [1]. This verifies that the organizer published the exhibitor specification. It is not an independent benchmark of accuracy, hysteresis, drift, shear sensitivity, usable yield, operator comfort, or policy improvement.

For a VLA, Moxian's value lies in observation. A tactile array can expose contact that vision loses under occlusion. It can distinguish a secure grasp from a visually similar slip, or initial insertion from a jam. Yet an array does not automatically produce model-ready tokens. Calibration, dead-channel mapping, timing alignment, normalization, temperature compensation, sensor replacement, and contact-label definitions decide whether the signal survives training and deployment.

The action and controller fields are undisclosed because Moxian does not claim to own the action stack in the captured sources. Its sensor may feed a hand controller, state estimator, VLA, or quality monitor. Who closes the force loop, how frequently it runs, and how safety limits override model commands remain integrator decisions. Public end-to-end latency and recovery behavior are also undisclosed.

SDK, ROS, and Isaac support are unverified in the admitted official evidence. The site presents hardware and solution delivery, but no public versioned API, ROS message definition, Isaac extension, reference driver repository, or model license was established. Openness is therefore recorded as interface details undisclosed, not closed by inference. A purchaser should request schemas, timestamps, calibration files, error codes, replacement procedures, and supported middleware versions.

Commercial maturity is production and customized delivery of sensing products, supported by official manufacturing language. The homepage reports more than 500 customers and service across more than 20 provinces; these are issuer-defined traction counts, not audited robot deployments [5]. Revenue, gross margin, shipments by product, active robot installations, renewal, accepted data hours, and policy-lift results are undisclosed. The about page shows financing parties but no comparable round amount in the admitted primary evidence.

WRC status is organizer-backed for the glove showcase. Media-corpus visibility is unmeasured because no complete eligible-news corpus, period, language rule, and deduplication method were established. Strengths are dense tactile coverage, a manufacturing-oriented sensor portfolio, contact visibility, and a concrete capture device. Limits are no public VLA, no independently validated policy benefit, undisclosed developer interfaces, and no public long-run calibration economics.

Confidence is high for legal name, 2021 foundation, Dongguan location, official products, and the organizer-published 368-taxel specification; medium for production maturity and issuer customer counts; low for model integration, latency, openness, dataset contribution, and policy impact. The official application photograph below shows only the physical form of a sensor-covered robotic hand; it does not verify the 368-taxel glove specification, calibration, or policy performance.

Figure 10.1: Moxian official electronic-skin application photograph showing a sensor-covered robotic hand touching a transparent panel. It shows a physical product form but does not verify the 368-taxel glove specification, calibration, or policy performance. Source: Moxian Technology official product page, https://www.moxiantech.com/robot-scene/14, fair use for academic review

10.4 Deep Profile II — DexRobot

The Korean rendering is 덱스로봇, the English brand is DexRobot, and the Chinese brand and legal entity are 灵巧智能 and 浙江灵巧智能科技有限公司. The official privacy policy identifies the legal entity [8]. The current WRC page says the company was founded in 2024, while organizer materials associate it with Xinchang, Shaoxing, Zhejiang [9]. A formal headquarters label is absent, so this profile records Xinchang as disclosed founding and registered location rather than silently upgrading it to headquarters.

Flagships include the DexHand021 family, DexCap wearable exoskeleton, DexTac simulation framework, DexCanvas dataset, and integrated dexterous-operation solutions [6] [7]. The mass-production hand page reports 19 DoF; position, normal-force, tangential-force, proximity, and temperature sensing; CAN FD; and Python/C++/ROS integration. These are issuer laboratory specifications, not an autonomous VLA success rate.

DexRobot reports that DexCap has 65 DoF, weighs 4.2 kg, and communicates at up to 1,000 Hz wired and 200 Hz wireless. The same first-party page gives both approximately eight hours and an 8–12-hour operating range, so the duration remains an unresolved issuer-page conflict [7]. It also reports endpoint accuracy of X/Y no more than 0.5 mm and Z no more than 0.1 mm. These values require coordinate-frame, calibration, motion-volume, confidence, and test-protocol definitions. Encoder rate is not end-to-end teleoperation latency or robot-task accuracy.

DexCap supplies human motion and a command reference. DexHand supplies a target embodiment with tactile and proprioceptive state. Retargeting connects them, while a learned policy may imitate demonstrations. This is a strong data path, but the product page does not disclose how human motion maps to active and passive joints, how impossible poses are handled, whether operator contact force is captured, or how aborted demonstrations are labeled.

The action boundary is partly visible. DexHand exposes CAN FD and low-level development support; the issuer names motion mapping, teleoperation, reinforcement learning, imitation learning, simulation, and ROS. Public evidence does not show a single generalist VLA with fixed weights, defined observation tensor, action vocabulary, inference rate, controller split, or factory recovery policy. DexRobot is best classified as marketed hardware plus data and developer infrastructure, not a publicly reproducible foundation-model result.

SDK maturity is stronger than Moxian's because Python, C++, ROS, low-level API access, communication interfaces, simulation models, and data assets are named [6]. Openness still varies by asset. Saying “open-source ecosystem” does not reveal whether firmware, training code, weights, dataset licenses, or builds are public. Isaac support was not established. Version policy, release hashes, dependency locks, and long-term API compatibility remain unverified.

The current WRC 2026 page lists DexRobot at booth B415 and names DexHand021 variants, DexCap, and Dexele [9]. An older English organizer route displays another booth number, so visitors should reconfirm the current directory. Participation verifies presence and exhibited products, not success or production volume.

Finance is undisclosed in the admitted primary sources. The issuer calls DexHand021 mass production, and the organizer names produced products. Unit shipments, paid customers, revenue, gross margin, warranty returns, and recurring software or data revenue are not disclosed. Application stories indicate traction, but no matched factory uptime, cycle-time, intervention, scrap, or return-on-investment dataset was established.

Media-corpus visibility is unmeasured. Strengths are a coherent hand-capture-simulation-data stack, detailed parameters, tactile observation, high-rate wearable capture, and named interfaces. Limits are incomplete openness, issuer-defined accuracy and durability, unclear human-to-robot action semantics, no disclosed generalist-model benchmark, and no audited factory economics. Confidence is high for identity, product specifications as claims, interface categories, and WRC listing; medium for production maturity; low for VLA generality, deployment success, finance, and external reproducibility.

The still frame below from the official DexCap page is useful for confirming the physical wearable capture hardware. It does not show the target DexHand, retargeting path, or end-to-end latency, so it should not be read as evidence of autonomous-policy performance or of the issuer-labeled transmission rate.

Figure 10.2: DexRobot official product-video still showing the DexCap exoskeleton worn over a black glove. It confirms physical capture hardware but not the target DexHand, retargeting, end-to-end latency, or policy performance. Source: DexRobot official DexCap page, https://www.dex-robot.com/en/dexCap, fair use for academic review

10.5 Deep Profile III — AI² Robotics

The Korean rendering is 에이아이 스퀘어드 로보틱스, the English brand is AI² Robotics, and the Chinese brand is 智平方. The current footer and WRC page use 智平方(深圳)科技股份有限公司, while an earlier privacy notice names 智平方(深圳)科技有限公司 as site operator [10] [15]. The company says it was founded in April 2023. The cited official statement names Shenzhen and Beijing locations, but it does not support an office-by-office split of headquarters, hardware, and AI functions [13].

Flagships are AlphaBot, AlphaBot 2, AlphaBrain, GOVLA, FiS-VLA, AlphaBotCore SDK, and VR or exoskeleton teleoperation [10] [11] [12]. This is the broadest claimed model-to-hardware stack among the profiles.

The issuer describes AlphaBrain as combining full-space understanding, whole-body coordination, and complex-task reasoning. AlphaBot 2 documentation describes a wheeled dual-arm body, NVIDIA Orin AGX 64G, cameras and lidar, and four-to-six-hour battery range [12]. The company publishes Python and C++ APIs for arms, chassis, sensors, kinematics, emergency-stop monitoring, grippers, and synchronization [11]. These make the controller boundary more inspectable without disclosing AlphaBrain's full input and action tensors.

AI² makes a comparative FiS-VLA claim against π₀ on its company page, but the captured evidence contains no matching primary paper or independently audited benchmark [10]. The task set, metric, aggregation, π₀ checkpoint, data budgets, embodiments, adaptation rules, trials, intervention policy, denominator, and table locator require primary technical verification before any numeric or cross-model superiority claim is retained.

Observation coverage is partly disclosed. Documentation exposes RGB/RGB-D, depth, joint, arm-force, chassis, and synchronized sensor interfaces, while the company describes spatial understanding and whole-body control. It does not establish the exact camera history, language format, tactile input, proprioceptive normalization, work-order context, or quality signals used by FiS-VLA in each deployment.

Action and latency are also partial. AlphaBotCore exposes motion and kinematics calls, and arm documentation includes stop, pause, resume, collision-related checks, and force-controlled operations. This proves conventional functions exist below the model. It does not show whether the VLA outputs joint targets, end-effector deltas, chunks, skills, or goals, nor disclose inference distribution, compute fallback, packet-loss response, or proposal rate.

Recovery remains unknown at deployment level. Official materials describe long-horizon tasks and data feedback but do not report first-attempt success, retry counts, remote intervention, mean recovery time, part quarantine, or operator workload. A public automotive story is useful organizer evidence of intended scope, not an audited production study [16].

Openness is mixed. FiS-VLA is described as open source, and hardware docs expose detailed Python and C++ interfaces. Public sources do not establish that all training data, preprocessing, checkpoints, evaluation scripts, deployment runtime, or adaptation recipes are available. ROS and Isaac integration were not verified. Reproducibility is partial: interface-level replication is more plausible than full training or benchmark replication.

The company announced a Pre-A+ financing round exceeding RMB 100 million in March 2025 and said it completed two rounds in the first two months of that year [14]. This is issuer-disclosed finance, not audited cash, valuation, or revenue. It reports deployments in automotive, semiconductor, biomanufacturing, public service, and retail. Customer names, installed fleet, paid production hours, utilization, renewals, margin, and model-attributable savings remain undisclosed.

WRC 2026 status is organizer-verified at booth C306 [15]. Media-corpus visibility is unmeasured. Maturity is integrated marketed platform with deployment claims and developer documentation, but pilot, order, installation, and stable production are not publicly reconciled.

Strengths are the integrated model, robot, data, SDK, and deployment story; visible APIs; and a whole-body platform suitable for testing the controller boundary. Limits are issuer-only comparative evidence without a reproducible matched benchmark, incomplete model and data reproducibility, undisclosed latency and recovery, confidential factory evidence, and dependence on private data. Confidence is high for identity, foundation, locations, products, SDK, financing event, and WRC presence; medium for maturity; low to medium for generality and factory benefit.

The image below is a physical AlphaBot 2 factory photograph published by the WRC organizer for an AI² Robotics automotive-manufacturing case. It confirms the robot, end effector, and physical cell arrangement, but not autonomy, cycle time, success rate, or stable production.

Figure 10.3: WRC organizer case photograph showing a physical AlphaBot 2 reaching toward a parts rack in an automotive factory cell. It confirms the physical platform and task context but not autonomy, cycle time, success rate, or stable production. Source: World Robot Conference official AI² Robotics case page, https://www.worldrobotconference.com/expo/case/23.html, organizer case photograph and fair use for academic review

10.6 Same-Scale Profile Comparison

Generalist VLA comparisons must retain model, data, action, controller, openness, and unseen-task definitions [22] [23] [24] [25]. Without them, “generalist,” “open,” and “zero-shot” describe incompatible experiments. The table compares disclosure, not presumed capability.

Field Moxian Technology DexRobot AI² Robotics
KO / EN / CN 목시안 테크놀로지 / Moxian Technology / 墨现科技 덱스로봇 / DexRobot / 灵巧智能 에이아이 스퀘어드 로보틱스 / AI² Robotics / 智平方
Foundation / location 2021 / Dongguan 2024 / Xinchang founding and registered location; HQ label undisclosed Apr. 2023 / Shenzhen and Beijing locations; office-function split unverified
Flagship pressure arrays, electronic skin, 368-taxel glove DexHand021, DexCap, DexTac, DexCanvas AlphaBot 2, AlphaBrain, FiS-VLA, SDK
Model-loop role tactile observation and capture embodiment, tactile endpoint, capture and simulation integrated VLA, robot, data, deployment
SDK / ROS / Isaac unverified / unverified / unverified Python and C++; ROS named; Isaac unverified Python and C++; ROS and Isaac unverified
Finance / traction finance undisclosed; issuer reports 500+ customers finance and comparable shipments undisclosed; production claimed issuer disclosed >RMB 100M Pre-A+; deployments claimed
WRC 2026 organizer glove showcase organizer listing, B415 organizer listing, C306
Reproducibility sensor integration undocumented publicly partial hardware and interface openness partial SDK and model openness; data/runtime incomplete

10.7 VLA Families and the Meaning of Generality

OpenVLA offers a 7-billion-parameter open VLA and shows that weights and adaptation code can broaden access [22]. Octo trains an open generalist policy over heterogeneous robot datasets [23]. π₀ uses a flow-based action expert across embodiments [24]. π₀.₇ adds steerable context such as subtask instructions, goal images, quality labels, and control-mode labels [25]. Each defines openness and generality differently.

Weights without training data support inference and fine-tuning, not full reproduction. Data without calibration may not reproduce observations. Code without the original robot, fixtures, or evaluation labor may not reproduce success. A model transferring between two arms may fail on a mobile manipulator because base motion changes viewpoint, reachability, and timing. Generality is a matrix over tasks, objects, scenes, bodies, sensors, controllers, and adaptation budgets.

The unseen-task label is fragile. An instruction can be new while motion is familiar. An object instance can be unseen while its category appeared in web or robot data. A robot-task pair can be new while both components appeared separately. Large private mixtures make overlap analysis difficult. π₀.₇ reports strong results on seen and new combinations, while zero-shot results remain lower and overlap is difficult to rule out [25].

Procurement should define novelty explicitly. Record whether geometry, material, part family, fixture, lighting, tool, robot, controller, and instruction are inside training and adaptation. Report target-cell data used. A “zero-shot” result after engineers tune thresholds, rewrite prompts, or select checkpoints is not zero human adaptation.

10.8 World Models: Prediction Is Not Control

A world model predicts a representation of future state conditioned on state and possibly action. It can support planning, counterfactual testing, data generation, or representation learning. A plausible video can still violate geometry, contact, timing, or failure ordering. Factory value requires that prediction preserve the variables governing a decision.

DreamDojo uses large-scale egocentric human video and robot post-training to build a generalist robot world-action model [26]. Its scale shows that human video can contribute structure beyond robot-labeled data. Transfer still depends on aligning human observations with robot actions and validating predicted dynamics. A hand passing through an object for one frame can be visually minor and physically fatal.

World models can work at several levels. Pixel prediction anticipates visual outcomes. Latent dynamics may support planning without rendering details. Object-centric models track parts and relations. Contact-aware models estimate force or slip. No representation is universally best. Collision avoidance needs geometry; insertion needs contact; quality prediction may need process signals; route planning needs occupancy and time.

SimFoundry offers a validation pattern by testing whether simulation ranks real policies, reporting correlation and rank-regret across disclosed tasks and policies [27]. SIMPLER similarly treats simulation as an evaluation whose value depends on correspondence with real performance [28]. The metric is not rendering beauty, but whether the virtual system leads to the same engineering decision as costly real trials.

10.9 Latency, Controller Boundaries, and Recovery

Gemini Robotics illustrates one latency architecture: a cloud-oriented backbone and local action decoder, with official technical discussion of optimized response and local compensation [29]. Slow semantic reasoning can guide a faster local policy, while servo and safety functions remain faster and independently authoritative. A monolithic model label does not remove hierarchy.

Action chunking generates several future commands together. FAST compresses action sequences for more efficient VLA training [30]. Chunking can smooth behavior and reduce inference calls, but creates commitment risk. If the scene changes halfway through a chunk, the robot needs cancellation. Evaluation should disclose chunk length, replan rate, observation overlap, stop latency, and whether the local controller can truncate a proposal.

Boundaries should be an authority table. The VLA proposes a goal or trajectory. A motion layer enforces kinematics and collision. A force controller handles contact. A safety controller monitors certified devices. A PLC arbitrates machine state. An operator owns restart after protected faults. If a learned model bypasses a boundary, the deployment must state what replaces it.

Recovery needs a budget. The robot may retry perception, retreat and regrasp, request guidance, quarantine a part, or stop. Infinite retries inflate cycle time and damage hardware. Silent continuation corrupts quality. KPIs should include first-attempt success, bounded-recovery success, interventions per hundred cycles, unsafe proposals blocked, false success, recovery time, and scrap during recovery.

10.10 Private Data Dependence and Reproducibility

Robot policies depend on data factories rarely publish: drawings, cell images, recipes, defect examples, worker video, cycle records, and failures. Private data can create real advantage but complicates evidence. A vendor may honestly report strong results that outsiders cannot reproduce because the decisive corpus and test cell are confidential.

Public weights are only one rung of openness. A stronger ladder includes architecture, weights, action codec, training code, data manifest, preprocessing, calibration, evaluation scripts, robot adapters, controller parameters, trial logs, reset decisions, and licenses. A buyer need not demand universal publication, but must know which layer is reproducible internally and which remains a managed-service dependency.

Private data changes upgrade risk. If a vendor retrains on pooled customer data, contracts should define ownership, consent, isolation, retention, deletion, derivative weights, and cross-customer transfer. If training remains on premises, compute, patching, and rollback become factory responsibilities. Cloud inference or logging adds outage, latency, export-control, cybersecurity, and incident-response requirements.

Every evaluation episode should point to model version, adaptation set, calibration, prompt, controller, hardware revision, and reset rule. Otherwise an update may improve one task and regress another silently. Physical deployment needs a release artifact and test record, not merely a model name.

Figure 10.4: Octo architecture linking language and image observation tokens, a transformer, action heads, and a finetuning path for new embodiments. It exposes a generalist-policy interface but does not establish deployment performance for the three profiled companies. Source: Octo Model Team et al. 2024, arXiv:2405.12213v2 Fig. 0

10.11 Evidence Tiers: Demo, Pilot, and Production

A company video shows that a configured system produced the displayed motion at least once, assuming editing is disclosed. It is weak evidence for probability. A product page establishes what the issuer offers and claims. Organizer material verifies participation and presentation. A paper can disclose methods and controlled trials. A buyer-run acceptance test is strongest for the buyer's process.

Factory evidence needs denominators. How many consecutive cycles were attempted? Were failed approaches edited out? Who reset objects? Did an operator select the grasp, approve the plan, or teleoperate recovery? Were parts representative of production tolerance? Was cycle time measured from dispatch through verified completion? Did the result enter the MES?

Production differs from shipment; shipment can end in a laboratory, integration project, or idle pilot. Deployment differs from routine use. Routine use may still lack economic value if utilization is low or support cost high. The evidence ladder is component specification, integrated demo, controlled benchmark, customer pilot, acceptance run, sustained production, and repeated deployment across sites.

AutoRT shows why orchestration matters: 53 robots collected 77,000 episodes over seven months, while safety relied on conventional controls and human supervision as well as foundation-model guidance [31]. BUMBLE reports long building-wide tasks, yet success and intervention figures show how errors compound [32]. Neither transfers to a factory without matched embodiment, environment, task, and safety assumptions.

10.12 Manufacturing Walkthrough: Mixed-Model Connector Insertion

Consider a line inserting one of four electrical connectors into variant-specific housings. Conventional automation uses a fixture, recipe poses, compliant insertion, force thresholds, and vision confirmation. It performs well under control, but engineering rises when geometry or tray layout changes. The goal is not “replace the PLC,” but reduce changeover and exceptions without degrading quality.

Step 1 — Freeze the baseline. Measure first-pass yield, cycle-time distribution, intervention, false accepts, damage, changeover engineering hours, and downtime. Record tolerances, lighting, fixture wear, and operator duties. Without a baseline, flexibility has no economic control group.

Step 2 — Define observations. Use calibrated wrist and overview cameras, robot state, force-torque, gripper state, recipe ID, and fixture status. For tactile sensing, map each taxel to calibrated units, time-align it, and log dead channels. Moxian-like arrays can add contact coverage; a DexHand-like endpoint can expose fingertip state. Neither enters the policy until synchronization and replacement tests pass.

Step 3 — Define actions and boundaries. Let the VLA propose connector identity, pre-insertion pose, grasp, and approach. Keep joint interpolation, collision avoidance, velocity limits, contact control, safety stops, and equipment handshake deterministic. The policy cannot release the fixture or declare quality. A verified sensor rule and PLC state own completion.

Step 4 — Build data splits. Separate production lots, variants, lighting, and dates. Reserve one connector-fixture combination for transfer. Store failures and corrections, not only successful demonstrations. Document whether pretraining may contain similar connectors. Private target-cell data is versioned and access-controlled.

Step 5 — Run shadow mode. The model observes and proposes without authority. Compare proposals with baseline actions; detect unsafe or infeasible commands. Measure inference age, dropped observations, disagreement, and confidence calibration. Shadow mode reveals integration defects without risking hardware.

Step 6 — Grant bounded authority. Start at low speed, isolated shift, trained supervision, sacrificial parts, and one attempt. The controller checks reachability and contact envelope. On abnormal force, retreat along a validated path and quarantine the part. Do not improvise another insertion unless recovery is separately approved.

Step 7 — Test transfer. Introduce the held-out variant, pose error, reflections, worn fixtures, and network degradation. Compare no adaptation, few-shot adaptation, and recipe baseline. Generality creates value only if it reduces target data or engineering while maintaining acceptance.

Step 8 — Promote and monitor. Require confidence intervals over consecutive cycles, no unresolved safety event, traceable versions, acceptable recovery, and rollback. After release, monitor drift, intervention, blocked commands, and quality escapes. Retraining is a controlled engineering change, not an invisible update.

Gate Pass evidence Failure response
Observation synchronized modalities; calibrated tactile/force; missing-data alarms stop authority or baseline fallback
Action units and frames tested; every proposal checked reject action and log reason
Performance yield and cycle time meet agreed interval extend pilot or narrow envelope
Recovery bounded retreat and quarantine validated stop cell; operator restart
Governance model, data, prompt, adapter traceable block promotion
Economics changeover gain exceeds compute, support, downtime cost retain conventional automation

10.13 Procurement Questions That Expose the Hidden Stack

Ask the vendor to replay one task from raw sensor timestamps through the action sent to the controller. Request tensor shapes, units, frames, action horizon, latency distribution, compute location, and fallback. Ask which observations are mandatory and which degrade gracefully. A live trace is more informative than an architecture adjective.

Ask for the reset ledger. How many attempts were made, which were excluded, who intervened, and why? Request first-attempt and recovered success separately. Ask how false success is detected and whether inspection is independent. For long tasks, request a failure tree by subtask.

Ask for openness by artifact. Decompose “open source” into weights, code, data, SDK, firmware, simulator, evaluation, and license. Ask whether a public checkpoint runs on the sold robot and whether its adapter matches production. Request clean-room reproduction by an engineer who did not build the demo.

Ask for private-data boundaries. Can the vendor train on buyer video? Where is it stored? Does it enter shared weights? Can it be deleted? Who owns annotations and failures? What happens if the relationship ends? Portability includes data schema and controller interface, not only the flange.

Finally, demand a matched baseline: the best conventional program, a modular learned skill system, and the VLA under the same parts, shifts, operators, and resets. Report engineering time as well as success. The strongest business case may be faster commissioning or exception coverage, not headline accuracy.

10.14 Limitations and Open Questions

The profiles are intentionally asymmetric. Moxian and DexRobot provide observation, embodiment, and data infrastructure that may feed many models, while AI² makes an integrated VLA claim. This chapter does not rank them as substitutes; it asks whether each disclosed layer can be integrated and verified.

Company specifications are issuer evidence unless an organizer or independent test is named. WRC pages verify listings and exhibits, not performance. Financing announcements verify an announced round, not valuation, cash received, revenue, or solvency. Customer and deployment language lacks a common denominator.

Research is also bounded. Tasks, robots, action spaces, adaptation budgets, resets, and mixtures differ. Closed corpora make overlap hard to audit. Open models may depend on costly private data and tuning. An average can conceal a safety-critical tail.

World-model evaluation remains unsettled. Pixel plausibility, latent prediction, policy ranking, planning benefit, and real success are different metrics. Contact, deformation, wear, and human behavior remain difficult. A model can be useful inside a limited envelope without being a complete simulator.

Open questions include how to certify a changing component, isolate regressions, price data and inference over a machine life, preserve confidentiality while improving generality, and share failures without exposing processes. The near-term answer is disciplined boundaries, versioned evidence, and conservative promotion.

Relation to Prior Surveys

Prior surveys on motion, tactile hands, and physical AI explain the components below this chapter. Here the emphasis is commercial and architectural: model claims are mapped to sensor, action, controller, safety, recovery, and factory evidence. Foundational papers define comparison axes; they do not imply that a vendor implements an unpublished method.

What to Learn Next

Chapter 11 moves from general-purpose intelligence claims to medical, consumer, and specialized robots. Those markets change the burden: clinical regulation, home support, privacy, unit economics, and narrow-task reliability may matter more than generality. Carry forward the same contract, but expect different safety authorities, buyers, and definitions of deployment.

References

  1. World Robot Conference (2026a). Moxian Technology and Tactile Data-Collection Glove Showcase. Official organizer source.
  2. Moxian Technology (2025). Official Company and Product Brochure. Official foundation and company source.
  3. Moxian Technology (2026a). Official Contact Page. Legal name and Dongguan address.
  4. Moxian Technology (2026b). Humanoid Robot Electronic Skin. Official product source.
  5. Moxian Technology (2026c). Official Homepage. Product, manufacturing, and issuer-traction source.
  6. DexRobot (2026a). DexHand021 Mass-Production Platform. Official product and interface source.
  7. DexRobot (2026b). DexCap Wearable Exoskeleton Data-Acquisition System. Official specification.
  8. DexRobot (2026c). DexRobot Privacy Policy. Legal-entity source.
  9. World Robot Conference (2026b). DexRobot 2026 Exhibitor Profile. Organizer booth and products.
  10. AI² Robotics (2026a). AI² Robotics Company Profile. Official identity, product, and model claim.
  11. AI² Robotics (2026b). AlphaBotCore SDK Documentation. Official Python interface.
  12. AI² Robotics (2026c). AlphaBot 2 Hardware Architecture. Official robot configuration.
  13. AI² Robotics (2025a). Official Statement on Company Identity and Locations. Official location source.
  14. AI² Robotics (2025b). Pre-A+ Financing Announcement. Issuer financing disclosure.
  15. World Robot Conference (2026c). AI² Robotics 2026 Exhibitor Profile. Organizer booth and AlphaBot.
  16. World Robot Conference (2026d). AI² Robotics Automotive Manufacturing Case. Organizer application story.
  17. Ahn, M. et al. (2022). Do As I Can, Not As I Say: Grounding Language in Robotic Affordances. CoRL; arXiv:2204.01691.
  18. Liang, J. et al. (2023). Code as Policies: Language Model Programs for Embodied Control. ICRA; arXiv:2209.07753.
  19. Brohan, A. et al. (2023a). RT-1: Robotics Transformer for Real-World Control at Scale. RSS; arXiv:2212.06817.
  20. Brohan, A. et al. (2023b). RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv:2307.15818.
  21. Chi, C. et al. (2023). Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. RSS; arXiv:2303.04137.
  22. Kim, M. J. et al. (2024). OpenVLA: An Open-Source Vision-Language-Action Model. arXiv:2406.09246.
  23. Octo Model Team (2024). Octo: An Open-Source Generalist Robot Policy. arXiv:2405.12213. Reading note #55.
  24. Black, K. et al. (2024). π₀: A Vision-Language-Action Flow Model for General Robot Control. arXiv:2410.24164. Reading note #2.
  25. Physical Intelligence et al. (2026). π₀.₇: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities. arXiv:2604.15483. Reading note #62.
  26. Gao, S. et al. (2026). DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos. arXiv:2602.06949.
  27. Ranawaka, N. et al. (2026). SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation. arXiv:2606.28276. Reading note #85.
  28. Li, X. et al. (2024). Evaluating Real-World Robot Manipulation Policies in Simulation. arXiv:2405.05941.
  29. Gemini Robotics Team et al. (2025). Gemini Robotics: Bringing AI into the Physical World. arXiv:2503.20020.
  30. Pertsch, K. et al. (2025). FAST: Efficient Action Tokenization for Vision-Language-Action Models. arXiv:2501.09747.
  31. Brohan, A. et al. (2024). AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents. arXiv:2401.12963.
  32. Garrett, C. R. et al. (2024). BUMBLE: Unifying Reasoning and Acting with Vision-Language Models for Building-wide Mobile Manipulation. arXiv:2410.06237.