FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception

Conference on Robot Learning (CoRL) 2026

Zhen Zhang1, Weinan Wang1*, Hejia Sun2*, Qingpeng Ding1, Xiangyu Chu1,4,
Guoxin Fang1,3, K. W. Samuel Au1,4
1 The Chinese University of Hong Kong 2 The Hong Kong Polytechnic University
3 Institute of Intelligent Design and Manufacturing, CUHK 4 Multi-Scale Medical Robotics Center (MRC), AIR@InnoHK
*Equal contribution
FingerViP dexterous manipulation teaser

FingerViP brings clear multi-view observations to the fingertips for real-world dexterous manipulation.

Hardware Design

We develop a vision-enhanced fingertip, which embeds a miniature camera. It is low-cost, compact, and highly modular, enabling quick fabrication, easy integration, and high maintainability.

FingerViP hand and exploded view of the vision-enhanced fingertip module

Hardware Assembly Video (6× speed)

Imitation Learning with Fingertip Visual Perception

Multi-view fingertip perception for coordinated arm–hand control.

A transformer-based diffusion policy combines five fingertip views, one third-view image, camera-pose encodings, per-finger currents, and robot joint positions. The policy learns from human demonstrations and predicts 26-DoF joint commands for the arm and hand.

FingerViP visuomotor policy and transformer-based action diffusion

How Does Fingertip Vision Enhance Manipulation?

We evaluate FingerViP across four challenging real-world tasks.

External and wrist-mounted cameras can become occluded by the hand, objects, and surrounding structures. Fingertip vision provides local visual cues during approach, alignment, and interaction. Combined with a third-view camera, FingerViP achieves an 80.8% average success rate, outperforming the baseline that uses a third-view camera and a wrist-mounted camera (51.6%).


Execution progressions for the four FingerViP manipulation tasks

Representative execution sequences across four tasks: (a) confined-box button pressing, (b) unstable-support stick retrieval, (c) curtain-occluded object retrieval, and (d) closed-cabinet object retrieval. Time progresses from left to right.

Quantitative Results

Method Task 1 Task 2 Task 3 Task 4 Average
Third-view camera 19.0% 37.8% 40.0% 54.5% 37.8%
Wrist camera 4.8% 27.0% 56.0% 61.8% 37.4%
Third-view + wrist 21.4% 45.9% 68.0% 70.9% 51.6%
Fingertips only 57.1% 56.8% 74.0% 67.3% 63.8%
Fingertips + wrist 42.9% 51.4% 76.0% 72.7% 60.8%
Without camera poses 52.4% 54.1% 78.0% 67.3% 63.0%
Without joint currents 57.1% 59.5% 84.0% 74.5% 68.8%
FingerViP (ours) 73.8% 75.7% 90.0% 83.6% 80.8%

Evaluation uses 42, 37, 50, and 55 trials for Tasks 1–4, respectively.

Qualitative Results

Task 1: Confined-Box Button Pressing

Success rate: 73.8% (31/42)

Trial 1

Trial 2

Trial 3

Task 2: Unstable-Support Stick Retrieval

Success rate: 75.7% (28/37)

Trial 1

Trial 2

Trial 3

Task 3: Curtain-Occluded Object Retrieval

Success rate: 90.0% (45/50)

Trial 1

Trial 2

Trial 3

Task 4: Closed-Cabinet Object Retrieval

Success rate: 83.6% (46/55)

Trial 1

Trial 2

Trial 3

Qualitative result videos are shown at 6× speed.

Failure Cases

Failure case videos are shown at 6× speed.

T1: Repeated Misalignment

T2: Slip Out

T3: Knock Over

T4: Slip Out

The nearly dark box interior makes the button difficult to localize, causing repeated misalignment, while insufficient contact friction can let slender, slippery objects slip out. Future work will explore integrated lighting, improved contact surfaces, and tactile sensing.

BibTeX

@article{zhang2026fingervip,
  title   = {FingerViP: Learning Real-World Dexterous Manipulation with Fingertip Visual Perception},
  author  = {Zhang, Zhen and Wang, Weinan and Sun, Hejia and Ding, Qingpeng and Chu, Xiangyu and Fang, Guoxin and Au, K. W. Samuel},
  journal = {arXiv preprint arXiv:2604.21331},
  year    = {2026}
}

Acknowledgement

This work was supported in part by the Multi-Scale Medical Robotics Centre, AIR@InnoHK. We would like to thank CRML@CUHK for generously providing the UR5e robotic arm, and Dezhao Guo for his assistance. We are also grateful to Professor Shing Shin Cheng for his support with a robotic arm for testing. We also thank Yanlin Chen, Jianghua Chen, Pui Hin NG, Ziliang Feng, and Yongjun Yan for their helpful feedback and fruitful discussions.