CIMI: Contact Imitation Interface

Transferring Dexterous Skills by Showing How Contact Unfolds

Zhixuan Xu1,*, Yichen Li1,2,*, Bingchen Wu1,2, Yiwen Hou1, Zixuan Liu1, Tianyu Qiu2,3, Lin Shao1,2,†

* Equal contribution† Corresponding author

1 National University of Singapore 2 RoboScience 3 South China University of Technology

Video

Narrated supplementary video · 2:59 · Download MP4

Method

CIMI collects human demonstrations, converts them into robot training data, and learns a policy that plans contact motion for robot execution.

Full CIMI pipeline: human demonstration collection, trajectory retargeting and view conversion, contact-surface policy training, and robot control.
The full pipeline, from human recordings to robot execution.

Human demonstration collection

Input
A human performing the task with the calibrated demonstration interface.
Output
Synchronized contact, context, and tracking videos at 10 Hz.

Matching fingertip modules provide the human and robot with the same contact geometry and local camera views. Wrist and web-space cameras record the surrounding scene while the demonstrator performs the task.

The interface records nine synchronized streams at 10 Hz: three contact views, four context views, and two tracking views.

Human demonstration and the recorded contact, context, and tracking views. Camera labels identify each view.

Contact-alignment-first retargeting and view conversion

Input
Recorded videos, camera calibration, and the target robot’s kinematics and contact geometry.
Output
Robot training demonstrations: retained contact views, reprojected context views, and retargeted contact-surface and wrist–joint trajectories.

Fiducial tracking recovers contact-surface poses from the recordings. Human and robot hands have different kinematics, so holding the robot wrist at the human camera pose can prevent the fingertips from reaching the demonstrated contact positions.

Contact-surface retargeting

CIMI optimizes wrist and finger motion together to match the demonstrated contact positions and orientations. Releasing the context-camera constraint lets the wrist move to improve contact alignment.

Contact positions align as the wrist and fingers move. The animation loops automatically; use the controls to pause.

Context-view reprojection

The retargeted wrist motion changes the camera pose. CIMI reprojects the human context images to that viewpoint and overlays the robot geometry. Contact-camera views are retained.

Human context images reprojected to the retargeted robot viewpoint, with robot geometry overlaid.
The converted context views match the robot camera poses from retargeting.

Policy architecture and training

Input
Processed robot demonstrations for planner supervision; augmented wrist–joint trajectories and forward kinematics for controller supervision.
Output
A trained contact-surface planner and a separately trained controller for the robot hand.

The planner and controller are trained separately. The controller remains fixed during planner training.

Contact-surface planner

A frozen C-RADIOv3-B encoder processes the selected RGB views. View-specific projections and a fusion network combine the visual features with inter-finger contact poses. A conditional flow network with six residual blocks predicts a 16-step contact-surface trajectory.

Planner architecture: frozen RGB encoder, view projections and fusion, relative contact-pose conditioning, and a six-block flow network with three Euler integration steps.
Planner architecture (Figure S2). Conditional flow matching uses the retargeted trajectories as supervision. Architecture and training settings, page 4.

Contact-surface controller

The controller constructs geometric features from contact-surface targets and current joints. An eight-block residual MLP predicts wrist and finger commands for all horizon steps in parallel. Training uses augmented motion trajectories, an initial action-regression phase, and a contact-corner loss with action regularization.

Controller architecture: forward kinematics and contact corners, geometric features, input projection, eight residual blocks, and wrist and joint command outputs.
Controller architecture (Figure S4). A controller is reused across tasks with the same hand, finger set, contact geometry, and sampling configuration. Controller training, pages 5–6.

Full learning-interface details in the supplement.

Robot planning and control

Input
Live robot camera views and measured joint state, together with the trained planner and controller.
Output
Wrist-pose and finger-joint commands. The robot executes eight predicted steps before observing and replanning.

At execution time, the planner receives the robot’s current camera views and proprioceptive state. It predicts contact-surface trajectories, and the controller converts them into coordinated wrist poses and finger-joint targets.

The robot executes the first eight steps of the 16-step prediction, receives new observations, and replans.

Human demonstrations and robot policy rollouts

Each video pairs a human demonstration with a learned LEAP policy rollout, including third-person, contact-camera, and context-camera views.

Coin picking. Precise contact before pinching: the robot wedges under the coin, then pinches and lifts. Human demonstration and LEAP policy rollout both shown at 1×.
1 cm cube picking. Coordinate wrist and fingertips around a 1 cm target. Human demonstration and LEAP policy rollout both shown at 1×.
Letter retrieval. Sustain contact while engaging and opening a deformable flap. Human demonstration and LEAP policy rollout both shown at 1×.
Bulb installation. Coordinate repeated contact, turning, and release across three fingers.

Observations

Four comparisons examine transfer across hands, contact-camera inputs, action representation, and contact alignment. Mean final-stage success is measured across six hand–task pairs, with 20 trials per pair for each method.

Shared demonstrations across robot hands

The same right-hand human demonstrations train policies for the right LEAP hand and left Wuji hand. Final-stage success is 75% and 70% for cube picking, and 85% for both hands on coin picking.

Human, LEAP, and Wuji execution using shared demonstrations. Playback speeds are labeled in the video.

Contact cameras

With contact-surface actions held fixed, adding contact cameras raises mean final-stage success from 1.7% to 80.8%. The examples compare context-only inputs with context and contact inputs.

Representative bulb and coin failures without contact cameras, alongside successful CIMI rollouts.

Contact-surface actions

With the same visual inputs, predicting contact-surface trajectories raises mean final-stage success from 30.8% to 80.8% compared with direct wrist–joint prediction.

Representative cube and letter failures with direct wrist–joint prediction, alongside successful CIMI rollouts.

Contact alignment

Releasing the context-camera constraint reduces mean coin contact-corner error from 14.67 to 0.51 mm. Coin lift success increases from 0% to 85%, and bulb lighting success from 0% to 90%. Both variants use corner matching.

Retargeting comparison and LEAP outcomes, with 20 trials per task and variant.

Abstract

Robot-free human demonstrations offer a promising route to scaling dexterous manipulation data. The key challenge is to preserve the demonstrator’s natural dexterity while producing contact-grounded observations and actions that transfer to robot hands. We present Contact Imitation Interface (CIMI), a framework with contact-grounded demonstration and learning interfaces. The demonstration interface combines shared contact geometry and fingertip RGB views with contact-alignment-first retargeting and reprojection of wider scene views. The learning interface introduces a hierarchical policy in which a contact-surface planner predicts relative contact-surface trajectories and a controller coordinates wrist and finger commands to improve contact alignment. CIMI combines contact-grounded observations, retargeting, and action representations, improving mean final-stage success by at least 50 percentage points over baseline policies across six hand–task pairs on LEAP and Wuji. Both the hardware designs and software will be publicly available.

Supplementary material

The PDF provides hardware and fabrication details, implementation settings, and policy camera inputs.

Open supplementary PDF