FLTA

Frame-Level Temporal Alignment for Human-to-Robot Visual Adaptation

Anonymous

Overview of FLTA with frame discrimination, a global progress prior, and a local temporal order prior
Figure 2. Overview of the proposed temporal alignment framework. Only the normalization-affine parameters of the visual encoder are updated. The global progress prior D and visual similarity construct soft correspondences, while the local temporal order prior Q discourages backward transitions and permits variable forward rates.

Abstract

Transferring visual representations pretrained on human videos to robot manipulation requires learning reliable correspondences between human and robot demonstrations. However, paired demonstrations can differ in execution rate and in the proportion of non-key frames that do not directly reflect task progress. Frames at the same relative timestamp may therefore represent different task stages, making correspondence learning failed. To address these issues, we propose Frame-Level Temporal Alignment (FLTA), a framework that uses two temporal priors to adapt visual encoders pretrained on human videos for robot manipulation. It learns shared task-progress representations without frame-level correspondence annotations, allowing frames at different relative temporal positions to match. A global progress prior combines normalized temporal positions with visual similarity to construct a soft correspondence target. A local temporal order prior penalizes backward transitions while allowing stays and varying forward rates to accommodate execution-rate differences. With ResNet-50 and ViT encoders, our method achieves relative improvements of 46.93% and 65.96%, respectively, over the best baselines in average simulation success rates and achieves higher task success rates on real-world manipulation tasks. These results also suggest that effective human--robot adaptation depends less on the number of parameters updated than on which parameters are selected.

Zero-shot manipulation

Visual encoders are evaluated through downstream RVT2 policies on 23 unseen AGNOSTOS tasks.

Simulation

AGNOSTOS unseen tasks

Mean success rate (%) with standard deviation across three runs.

Visual encoderMethodL1 ↑L2 ↑Avg. ↑
D4R-ViTFrozen8.31 (0.50)11.60 (0.33)9.74 (0.25)
+ HRP10.56 (1.67)12.13 (0.38)11.25 (1.03)
+ Ours18.05 (0.52)19.47 (0.50)18.67 (0.49)
R3M-RN50Frozen12.62 (0.25)11.73 (0.50)12.23 (0.36)
+ HRAlign11.59 (1.62)12.80 (0.33)12.12 (0.78)
+ Ours16.82 (1.24)19.47 (0.19)17.97 (0.67)
DecisionNCE-CLIPFrozen10.05 (1.24)9.73 (1.32)9.91 (1.14)
+ HRAlign*12.62 (2.06)9.47 (1.05)11.25 (1.58)
+ Ours15.38 (0.44)16.13 (1.00)15.71 (0.22)
AcTOL-CLIPFrozen11.28 (1.53)11.47 (0.19)11.36 (0.82)
+ HRAlign*11.90 (1.47)12.27 (0.19)12.06 (0.78)
+ Ours13.23 (0.75)21.73 (0.50)16.93 (0.64)

23 zero-shot tasks: 13 Level-1 and 10 Level-2 tasks. * Reproduced.
For DecisionNCE-CLIP and AcTOL-CLIP, Frozen and HRAlign use RVT2-full, while Ours uses RVT2-lite.

Physical manipulation

FLTA representations are transferred to ACT for four manipulation tasks on the Trossen WidowX AI.

Real robot

Trossen WidowX AI

Success rate (%) over 20 evaluation episodes per task.

MethodPick Carrot into PotPress ButtonPush Pot into Marked RegionOpen Pot by Lifting the LidAvg. ↑
Frozen352555028.75
+ HRP405035532.50
+ Ours6560851556.25

Four tasks with 40 training demonstrations per task.

Real-robot videos