Methodology
Data Collection, Ranking Rules & Model Classification
Data Sources
All benchmark results are collected from published papers and official repositories. We do not re-run experiments.
Ranking Rules
Models are ranked by their primary metric on each benchmark. Adroit, DexArt, and Bi-DexHands use Mean Success Rate. DexGraspNet uses Suc.1 (GSR), DexGrasp Anything uses average GSR, and Dexonomy uses GSR. Other sub-metric definitions used in the leaderboards are as follows:
- GSR — Grasp Success Rate Higher is better
- The fraction of generated grasps that pass the benchmark force test.
- Suc.6 Higher is better
- The strict six-force success rate. A grasp passes only if the object remains stable under all six orthogonal external forces in MuJoCo.
- OSR — Object Success Rate Higher is better
- The fraction of objects for which at least one generated grasp is successful.
- CDC — Contact Distance Consistency Lower is better
- Measures how consistently contact locations are reproduced across generated grasps.
- PEN / PD — Penetration Depth Lower is better
- PEN and PD are equivalent penetration measures across the leaderboards. Lower values mean less hand–object penetration.
- DIV / Diversity Lower is better
- DIV and Diversity refer to the same collapse measure across the leaderboards. Lower values indicate less collapse along the first principal component.
Known Limitations
Results across different benchmarks are not directly comparable. Different papers may use slightly different evaluation protocols.
Model Classification
We classify models into two categories based on their open-source status:
Open-Source Models
Models with publicly available code, marked with an "Open Source" badge. These models provide the highest level of reproducibility and transparency.
Other Models
Models without the "Open Source" badge include: (1) models whose code repository we could not find, and (2) models that were in "Coming Soon" status before the data collection deadline. These models are hidden by default but can be shown using the "Include All Models" toggle.
Data Notice
- •Data notice last updated: July 14, 2026.
- •If you find any errors or omissions, please let us know by creating an issue on GitHub or contacting us via email: business@evomind-tech.com
Disclaimer
Cross-benchmark comparisons should be avoided. Each benchmark has its own evaluation protocol and metrics.
Supported Benchmarks
Adroit
Mean Success Rate (%)
In-hand manipulation benchmarks featuring the Adroit hand and diverse object tasks.
DexArt
Mean Success Rate (%)
DexArt benchmarks emphasize articulated object manipulation and tool use.
Bi-DexHands
Mean Success Rate (%)
Bimanual dexterous manipulation tasks from Bi-DexHands.
DexGraspNet
Grasp Success Rate (%)
A large-scale dexterous grasping benchmark for evaluating stable and diverse grasps on general objects.
DexGrasp Anything
Average Grasp Success Rate (%)
A physics-aware benchmark for universal dexterous grasp generation across diverse objects.
Dexonomy
Grasp Success Rate (%)
A grasp-taxonomy benchmark evaluating the quality and coverage of diverse dexterous grasp types.
⭐ Support This Project
If you find this leaderboard helpful for your research, please consider giving us a star on GitHub!
Contact Us
Found errors or want to submit your model? Reach out via GitHub Issue or email!