TouchScale: 500 Hours of Human Vision and Touch
Synchronized human RGB-tactile data for perception and robot learning.
A camera can show a hand lifting an egg or wringing a wet cloth. It cannot directly measure where the fingers press or how much force they apply.
TouchScale records those measurements alongside video. It brings together approximately 500 hours of human activity, pairing first-person and wrist-view images with contact and pressure measurements from both hands. The dataset supports work on predicting contact, recognizing actions, and training robot manipulation policies.

Capturing Vision and Touch Together
Each recording combines a head-mounted RGB-D camera, two wrist-mounted RGB cameras, and tactile gloves on both hands. The head camera captures the overall task, while the wrist cameras provide close-up views of hand-object interaction. Each glove contains 880 tactile sensing elements across the fingers and palm, recording how contact and pressure change as an object is grasped, manipulated, and released.
All recordings use the same generation of hardware and synchronization pipeline. Quality checks compare wrist-view video with tactile responses, reject clear sensor failures, and send uncertain cases for human review.
Everyday Tasks, Varied Contact
TouchScale contains approximately 87,000 episodes, 2,000 task descriptions, and more than 1,500 objects. Collection spans kitchens, laboratories, workbenches, offices, and other settings. Activities include pouring water, coiling cables, folding fabric, and handling laboratory equipment.
Repeated recordings vary the objects, their positions and orientations, and the scene layout. Participants also differ in their grasp choices, timing, and hand coordination. These variations capture different contact patterns within the same task, rather than repeating one fixed execution.

What TouchScale Adds to Model Training
Predicting contact from images. We train the same prediction architecture separately on EgoTouch's 16.2-hour training split and an approximately 16-hour TouchScale subset. We then test both models on EgoTactile, which uses a different tactile sensor, without further training. The TouchScale-trained model reaches a contact IoU of 0.181, compared with 0.134 for EgoTouch. This metric measures agreement between predicted and recorded contact.
Learning to recognize actions. Tactile-supervised pretraining on TouchScale produces higher action recognition accuracy than pretraining on OpenTouch, FEEL, or EgoTouch across MECCANO, Something-Something V2, and Ego-Exo4D. The comparison uses the same encoder initialization and training budget. The result holds both when the encoder is kept frozen and when it is fine-tuned on the downstream task.
Improving robot manipulation. We use TouchScale for an intermediate training stage in N0-VTLA before training the policy on robot demonstrations. This human-data stage requires no human action labels and does not require converting human hand movements into robot commands.
On an xArm6 with a dexterous robotic hand, average success rises from 22.5% to 57.5%, a gain of 35 percentage points. The four tasks involve sorting soft and hard objects, moving bottle caps into a tray, transferring test tubes, and wiping a whiteboard. Both variants start from the same checkpoint and use the same 50 robot demonstrations per task.
Separate scaling experiments also show an overall improvement in tactile prediction and robot success as more TouchScale data is used.

Looking Ahead: Learning from Human Contact
For teams developing embodied AI, TouchScale provides human interaction data for training contact predictors, visual encoders, and tactile-aware robot policies. These recordings can complement robot demonstrations: human data supplies visual-tactile supervision, while robot demonstrations supply the actions needed for a particular platform and task.