Simulation
AGNOSTOS unseen tasks
Mean success rate (%) with standard deviation across three runs.
| Visual encoder | Method | L1 ↑ | L2 ↑ | Avg. ↑ |
|---|---|---|---|---|
| D4R-ViT | Frozen | 8.31 (0.50) | 11.60 (0.33) | 9.74 (0.25) |
| + HRP | 10.56 (1.67) | 12.13 (0.38) | 11.25 (1.03) | |
| + Ours | 18.05 (0.52) | 19.47 (0.50) | 18.67 (0.49) | |
| R3M-RN50 | Frozen | 12.62 (0.25) | 11.73 (0.50) | 12.23 (0.36) |
| + HRAlign | 11.59 (1.62) | 12.80 (0.33) | 12.12 (0.78) | |
| + Ours | 16.82 (1.24) | 19.47 (0.19) | 17.97 (0.67) | |
| DecisionNCE-CLIP | Frozen | 10.05 (1.24) | 9.73 (1.32) | 9.91 (1.14) |
| + HRAlign* | 12.62 (2.06) | 9.47 (1.05) | 11.25 (1.58) | |
| + Ours | 15.38 (0.44) | 16.13 (1.00) | 15.71 (0.22) | |
| AcTOL-CLIP | Frozen | 11.28 (1.53) | 11.47 (0.19) | 11.36 (0.82) |
| + HRAlign* | 11.90 (1.47) | 12.27 (0.19) | 12.06 (0.78) | |
| + Ours | 13.23 (0.75) | 21.73 (0.50) | 16.93 (0.64) |
23 zero-shot tasks: 13 Level-1 and 10 Level-2 tasks. * Reproduced.
For DecisionNCE-CLIP and AcTOL-CLIP, Frozen and HRAlign use RVT2-full, while Ours uses RVT2-lite.