Task 1: Confined-Box Button Pressing
Task 2: Unstable-Support Stick Retrieval
Task 3: Curtain-Occluded Object Retrieval
Task 4: Closed-Cabinet Object Retrieval
Multi-view fingertip perception for coordinated arm–hand control.
A transformer-based diffusion policy combines five fingertip views, one third-view image, camera-pose encodings, per-finger currents, and robot joint positions. The policy learns from human demonstrations and predicts 26-DoF joint commands for the arm and hand.
We evaluate FingerViP across four challenging real-world tasks.
External and wrist-mounted cameras can become occluded by the hand, objects, and surrounding structures. Fingertip vision provides local visual cues during approach, alignment, and interaction. Combined with a third-view camera, FingerViP achieves an 80.8% average success rate, outperforming the baseline that uses a third-view camera and a wrist-mounted camera (51.6%).
| Method | Task 1 | Task 2 | Task 3 | Task 4 | Average |
|---|---|---|---|---|---|
| Third-view camera | 19.0% | 37.8% | 40.0% | 54.5% | 37.8% |
| Wrist camera | 4.8% | 27.0% | 56.0% | 61.8% | 37.4% |
| Third-view + wrist | 21.4% | 45.9% | 68.0% | 70.9% | 51.6% |
| Fingertips only | 57.1% | 56.8% | 74.0% | 67.3% | 63.8% |
| Fingertips + wrist | 42.9% | 51.4% | 76.0% | 72.7% | 60.8% |
| Without camera poses | 52.4% | 54.1% | 78.0% | 67.3% | 63.0% |
| Without joint currents | 57.1% | 59.5% | 84.0% | 74.5% | 68.8% |
| FingerViP (ours) | 73.8% | 75.7% | 90.0% | 83.6% | 80.8% |
Evaluation uses 42, 37, 50, and 55 trials for Tasks 1–4, respectively.
Mean absolute error (MAE) of the first eight predicted actions (×10⁻³; lower is better).
| Fingertips + third-view camera | Average MAE |
|---|---|
| Thumb + index | 33.87 |
| Thumb + index + little | 31.58 |
| Index + middle + ring | 32.03 |
| All five fingertips (ours) | 30.37 |
Action-prediction MAE for the first 1, 4, and 8 action steps on 20 real-world trajectories (lower is better; column-best values in bold).
| Visual encoder | 1st | First 4 | First 8 |
|---|---|---|---|
| Frozen CLIP (ours) | 0.0200 | 0.0272 | 0.0377 |
| Frozen DINOv2 | 0.0214 | 0.0282 | 0.0374 |
| Fine-tuned CLIP | 0.0232 | 0.0303 | 0.0396 |
| Fine-tuned DINOv2 | 0.0209 | 0.0278 | 0.0371 |
| CLIP from scratch | 0.0225 | 0.0306 | 0.0428 |
| DINOv2 from scratch | 0.0215 | 0.0294 | 0.0406 |
FingerViP runs at 10.2 FPS, compared with 13.2 FPS for the third-view and wrist-camera baseline. Data acquisition takes 35.2 vs. 34.4 ms per frame, and policy computation requires 215.01 vs. 79.00 GFLOPs, respectively.
S1: position, color, non-illuminated button, shape, white box, ambient light. S2: position, color, shape, material, length. S3: pose, object, curtain material, curtain color, curtain length, wall appearance. S4: pose, object, handle shape, handle color, wall appearance.
| Task | T1 | T2 | T3 | T4 | T5 | T6 |
|---|---|---|---|---|---|---|
| S1 | 17/18 | 4/6 | 4/6 | 3/6 | 1/3 | 2/3 |
| S2 | 14/18 | 4/6 | 5/6 | 2/3 | 3/4 | — |
| S3 | 18/20 | 7/10 | 5/5 | 5/5 | 5/5 | 5/5 |
| S4 | 22/25 | 13/15 | 4/5 | 3/5 | 4/5 | — |
These experiments evaluate variations within the four task families. Low illumination and slender, slippery objects remain challenging; broader in-hand manipulation and complementary tactile sensing remain future work.
Success rate: 73.8% (31/42)
Trial 1
Trial 2
Trial 3
Success rate: 75.7% (28/37)
Trial 1
Trial 2
Trial 3
Success rate: 90.0% (45/50)
Trial 1
Trial 2
Trial 3
Success rate: 83.6% (46/55)
Trial 1
Trial 2
Trial 3
Qualitative result videos are shown at 6× speed.
Failure case videos are shown at 6× speed.
T1: Repeated Misalignment
T2: Slip Out
T3: Knock Over
T4: Slip Out
The nearly dark box interior makes the button difficult to localize, causing repeated misalignment, while insufficient contact friction can let slender, slippery objects slip out. Future work will explore integrated lighting, improved contact surfaces, and tactile sensing.
@article{zhang2026fingervip,
title = {FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception},
author = {Zhang, Zhen and Wang, Weinan and Sun, Hejia and Ding, Qingpeng and Chu, Xiangyu and Fang, Guoxin and Au, K. W. Samuel},
journal = {arXiv preprint arXiv:2604.21331},
year = {2026}
}
This work was supported in part by the Multi-Scale Medical Robotics Centre, AIR@InnoHK. We would like to thank CRML@CUHK for generously providing the UR5e robotic arm, and Dezhao Guo for his assistance. We are also grateful to Professor Shing Shin Cheng for his support with a robotic arm for testing. We also thank Yanlin Chen, Jianghua Chen, Pui Hin NG, Ziliang Feng, and Yongjun Yan for their helpful feedback and fruitful discussions.