F-01 Observation
Time-synchronised head camera (RGB/depth) and joint states; 20 ms drift budget, 50 ms timeout with re-capture.
"Put the red block in the right tray with your left hand." TF-VLA combines a VLA backbone with walking, bimanual grasping, layered safety guards and an end-to-end training pipeline in one development platform.
Open the system (HMI) Talk to us How it works
The twelve functions of the system design, split into online execution, offline data and operations.
Time-synchronised head camera (RGB/depth) and joint states; 20 ms drift budget, 50 ms timeout with re-capture.
Japanese and English instructions mapped to task definitions: hand selection, walking commands, confirmation on ambiguity.
Action chunks (both hands + locomotion, 11-D x 10 steps) with dead-zone and safety-zone clipping followed by re-validation.
Chunks interpolated to 125 Hz joint targets; redundant-arm IK clamped to trackable increments; tracking and link-loss monitoring.
Hardware E-Stop path, soft guards for speed, torque, workspace, self-collision and human proximity. 8 ms stop response.
Teleoperation recording, quality filtering and splits, SmolVLA fine-tuning, in-distribution and OOD evaluation with pass criteria.
Chunked delivery with hash verification, parallel load with smoke test, instant rollback.
Asynchronous action logs, metrics and prioritised alerts. Eye-in-head camera calibration (0.24 px reprojection error).
Online execution follows the robot's control cycle; learning, evaluation and improvement run offline.
MuJoCo model of the real Unitree G1 (legs 6x2, waist 3, arms 7x2, Dex3 hands 7x2): gait generation, 7-DoF arm IK, finger contact stop, crouch and marching.
Egocentric camera fixed to the torso, rendered with real meshes, textures and shadows; domain randomisation widens the training distribution.
SmolVLA (450M parameters, Apache 2.0) fine-tuned with LeRobot; swappable for pi0 or GR00T N1 through the policy interface.
Web operator panel with instruction input, robot view, whole-body state, E-Stop and model switching, protected by token authentication and audit logging.
"Put the red block in the right tray" in simulation.




AI output is treated as an assistive motion; intrinsic safety stays with conventional hardware circuits.
| Item | Design | Measured on the jig |
|---|---|---|
| Emergency stop | Motor power cut through the safety PLC, independent of software; no auto-reset | Immediate stop even while walking; no recovery until explicit reset |
| Soft guards | Joint and hand speed, predicted torque, workspace, hand separation, slow-down near people | 8 ms stop response (requirement 100 ms) |
| Inference faults | Timeout to safe stop; one immediate retry per instruction | 130 ms inference (budget 500 ms) |
| Access control | Token-authenticated operator API behind TLS; audit log with client origin | 401 / 403 for unauthenticated calls, 422 / 413 on oversized input |
Development numbers as measured. Learned-policy success is still being improved with more data.
| Policy | Success | Notes |
|---|---|---|
| Scripted reference policy | 12 / 12 | Walking, one-handed grasp, side-stepping to the opposite tray; 7.6 s average |
| SmolVLA fine-tuned (1,070 episodes, 10,000 steps) | 1 / 20 | Learned to walk and stop, idle hand stays still; target selection remains open |
| Safety guard | 100% of run-aways stopped | Every learned-policy excursion stopped within 8 ms |
The TF-VLA operator panel running on the simulator jig is available online: watch the robot's head camera and try natural-language instructions, walking, manual hand and leg control and the emergency stop.
https://hmi.tf-vla.net/
Viewing state and video is open to everyone; moving the robot requires an access token, issued on request.
Instructions such as "put the red block in the right tray with the left hand" or "walk forward 0.3 m", finger presets, crouch and marching, emergency stop and reset, switching to learned models.
Shared demo environment, one operator at a time. No physical robot is connected; the Unitree G1 model runs in MuJoCo. Operations are recorded in an audit log.
A Vision-Language-Action model takes camera images and a natural-language instruction and directly outputs robot actions, replacing the classic perception-planning-control pipeline with a single model.
Yes. The robot, cameras, safety PLC and teleoperation are software jigs, and the real Unitree G1 model is rendered in MuJoCo, so the system is validated locally before moving to hardware over ROS 2.
Open-weight VLA models such as SmolVLA, pi0 and GR00T N1, through a common policy interface.
For deployment, joint development or demos, contact Technology Frontier Co., Ltd.
Technology Frontier Co., Ltd.
Email: info@technologyfrontier.co.jp