TACROSSRead paper ↗
Human touch · Demonstrations across tasks and environmentsOpening animation · Original playback speed
HUMAN TOUCH / ROBOT LEARNING

TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning

Learning from human touch across heterogeneous tactile sensors.

Bo Chen1,6,*, Huanzhang Hu2,6,*, Junyang Ma5, Bo Yue2, Fangdi Yu3, Haijier Chen4,6, Xianxin Lai1,6, Shuyu Pan6, Zhen Yang1,6, Xiaoquan Sun1,6, Wenze Cui1, Zhongliang Jiang1, Shaopeng Liu1,†, Jiayu Chen1,6,†

1 The University of Hong Kong 2 The Chinese University of Hong Kong (Shenzhen) 3 Ocean University of China 4 Wuhan University 5 Wuhan Textile University 6 INFIFORCE
* Equal contribution. † Corresponding authors.

3.5×

Recording throughput

140 vs. 40 episodes / hour
95.7%

Lower acquisition equipment cost

$520 Vision-IK Ego vs. $12,000 Robot-only
91.5%

Four-task mean success

vs. 92.2% Robot-only under matched time budgets

Collecting tactile demonstrations on robots is costly and slow. Human tactile gloves offer a scalable alternative, but differences in sensing principles, layouts, and dynamics make direct alignment of human and robot sensor signals ill-posed.

TACROSS bridges this sensing gap by aligning contact events in a shared tactile representation. Sensor-specific canonicalizers and adapters map observations from a low-cost piezoresistive glove and a capacitive robot hand into a common contact-semantic space. Human demonstrations support representation learning and confidence-weighted auxiliary hand supervision, while executed robot actions remain the sole source of ground-truth action supervision. Together, these components make human touch a scalable resource for learning contact-rich robot manipulation.

Supplementary Demonstrations

A three-minute walkthrough of TACROSS: human tactile data collection, cross-sensor contact-semantic alignment, four contact-rich robot tasks, and experimental results.

Full project walkthrough · 2:59 · Original playback speed

Contact makes
the difference.

Explore real robot demonstrations, including successful executions and failure examples from the presentation.

Success · One drop

2× source playback

Success · Two drops

2× source playback

Failure example

2× source playback

Videos are representative examples. Failure clips are not labeled as ablation baselines. Aggregate performance is reported below.

Human touch & robot execution: keyframes +Fixed-point dispensing: human and robot keyframes, tactile maps and force traces

Human and robot sequences are separate demonstrations, with timestamps local to each recording. Figure supplied in the presentation.

Different sensors. Shared contact semantics.

Matching raw sensor values is ill-posed. TACROSS represents contact events, relative intensity and phase in a shared space.

INSIDE TACROSS

From human touch to robot actions

Human touchRobot touchShared representation

Swipe across the diagram to explore each branch.

TACROSS: cross-sensor tactile learningHuman and robot touch are canonicalized with sensor-specific adapters, then encoded with shared temporal and inter-finger weights. Training aligns contact semantics. Robot tactile features, RGB and state feed a policy that produces arm and hand actions. Deployment uses only the robot branch. HUMAN BRANCHROBOT BRANCH Human glovePiezoresistiveTouch observations CanonicalizeHuman adapterContact · force · phase Shared encoderTemporal + inter-finger16 frames × 5 fingers zₕ256-D latent Shared weights Revo2 handCapacitiveTouch observations CanonicalizeRobot adapterCalibration + validity Shared encoderTemporal + inter-finger80 tokens / window zᵣ256-D latent TRAINING ONLY · SEMANTIC ALIGNMENTMatch contact, not raw sensor channelsMasked semantic targets + contrastive learning ROBOT OBSERVATIONSRGB + state POLICY INTERFACEACT / DP / VLA policy ACTION CHUNKArm + hand actions TRAINING ONLY · ACTION SUPERVISIONExecuted robot actions → action loss Robot observations onlyLearned encoders and policy run at inference.No human input or alignment loss is needed.
01 / CANONICALIZE

Sensor-specific calibration and adapters express human and robot touch in a common contact schema.

Conceptual data flow · Alignment and supervision are training-only. The paper evaluates ACT; the policy interface is shown as ACT / DP / VLA.

Original paper figure & complete architecture +
TACROSS architecture: human and robot contact canonicalization, shared temporal and inter-finger encoding, and robot-supervised ACTView full size ↗
Paper Fig. 4 · Sensor-specific adapters map human and robot touch into shared contact semantics. Robot demonstrations supervise the action policy; deployment uses robot observations only.
Representation & supervision details +

Measurement validity. Masks distinguish missing measurements from observed no-contact values. Canonical fields encode regional contact and interaction dynamics.

Robot-grounded actions. Retargeted human hand targets are auxiliary pseudo-labels. They do not replace executed robot actions as ground truth.

Deployment. Only robot observations are required at inference. The evaluated system still relies on robot demonstrations and embodiment-specific calibration.

Read the complete method in the paper ↗

Touch, captured at human scale.

A fabric-based tactile glove records touch alongside RGB and hand pose. Human collection can proceed independently of robot availability.

TACROSS glove sensing coverage and all five fabric layers, with electrodes surrounding pressure-sensitive fabricView full size ↗
Paper Fig. 2 · Sensing coverage and the complete five-layer piezoresistive glove structure.

A soft interface.
A small component cost.

Glove component BOM
$10.86
Physical sensing locations
285
Readout frame
256 values
Output rate
~150 Hz

Physical sensing locations, readout values and the learned 256-D latent are distinct quantities. Component BOM is separate from full acquisition equipment cost.

Human demonstrator collecting outdoor fruit manipulation data
Human demonstrations beyond the robot workspace

Vision-IK Ego $520

Vision-based hand tracking + tactile glove

140 episodes / hour

Rokoko Ego $2,150

Motion capture + tactile glove

112 episodes / hour

Robot-only reference: $12,000 equipment and 40 episodes/hour. Throughput refers to recorded episodes; acceptance rates are reported separately in the paper.

One synchronized record of the interaction.

Human acquisition pipeline showing glove and wrist readout, Rokoko and Vision-IK configurations, calibration, timestamp synchronization, and RGB, pose and tactile observationsView full size ↗
Paper Fig. 3 · From wearable sensing to synchronized episodes: calibrate touch, align observations using timestamps, and store RGB, hand pose and tactile data together.

Every interaction has a contact story.

Human recordings span industrial, home, laboratory and outdoor settings, with synchronized visual, motion and tactile observations.

Examples of human demonstrations across industrial, home, chemical laboratory and outdoor environments
159.09h

Human recordings

24.86h

Robot recordings

21,412

Valid human demonstrations

Aligned observations,
richer demonstrations.

RGB captures the scene. Hand pose describes motion. Touch captures the contact that vision can leave ambiguous.

The full corpus contains 24,913 recorded and 22,486 valid human + robot episodes. Corpus totals differ from the subsets used for policy training. Robot closed-loop evaluation covers four tasks.

Dataset release status ↗
Synchronized RGB images, hand pose overlays and five-finger tactile maps
Example multimodal observations from the supplied presentation

More human touch. Stronger robot learning.

Separate experiments examine data augmentation, contact-semantic alignment and the acquisition trade-off.

CLOSED-LOOP ROBOT EVALUATION

Human experience.
Robot execution.

The learned policy is evaluated on four contact-rich tasks using the Revo2 tactile hand. At deployment, it receives robot RGB, state and touch.

Robot evaluation apparatus with a 7-DoF arm, Revo2 tactile hand, egocentric RGB camera, task fixture and receiving trayView full size ↗
Paper Fig. 6 · Robot evaluation setup, including the tactile hand, arm, camera and task fixture.
FIXED ROBOT DEMONSTRATION COUNT

Human touch and
alignment both matter.

With 30 robot demonstrations per task, adding 150 human demonstrations and the full TACROSS pipeline increases mean success from 51.3% to 91.9%.

+40.6 percentage points

Paper Table IV · Four-task average · Human-augmented variants use R:H = 1:5.

SCALING HUMAN DATA

Fixed at 15 robot episodes.

Increasing human demonstrations from 0 to 90 raises mean success from 32.5% to 85.6%.

100602002040609032.585.6%

Human demonstrations per task → · Paper Table V · Mean of T1–T4

MATCHED COLLECTION TIME BUDGETS

Similar success, lower equipment cost.

ConfigurationMean successEquipment
Robot-only92.2%$12,000
Rokoko Ego93.5%$2,150
Vision-IK Ego91.5%$520

Paper Table II / Fig. 7 · Ego variants at nominal R:H = 1:5. These results use a different budget setting from the fixed-count ablation above. Costs cover acquisition equipment.

Scope & current limitations +

Transfer still depends on robot demonstrations, as well as retargeting and calibration specific to each embodiment. The reported evaluation uses the Revo2 hand and four closed-loop tasks. Broader robot-hand coverage and smaller robot-data budgets remain future directions.

Explore TACROSS.

The manuscript is available below. Hardware designs, software and the dataset are planned for release.

01 / MANUSCRIPT

Read the paper ↗

System, method and full experimental details.

02 / CODE & HARDWARE

Reproduce the system

Planned release · Repository link to follow.

03 / DATASET

Learn from human touch

Planned release · Download link to follow.

Citation

@misc{chen2026tacross,
  title = {TACROSS: An Efficient and Low-Cost Scalable Human Touch System Across Heterogeneous Tactile Sensors for Dexterous Robot Learning},
  author = {Chen, Bo and Hu, Huanzhang and Ma, Junyang and Yue, Bo and Yu, Fangdi and Chen, Haijier and Lai, Xianxin and Pan, Shuyu and Yang, Zhen and Sun, Xiaoquan and Cui, Wenze and Jiang, Zhongliang and Liu, Shaopeng and Chen, Jiayu},
  year = {2026},
  note = {Research manuscript}
}

Manuscript citation. Publication venue and persistent identifier will be updated when available.