HumanoidToolBench GitHub0

Benchmarking Humanoid Tool Use from Selection to Mobile Execution

Kyochul Jang1, Seohyeon Park1, Ohchul Kwon1, Sangjun Park2, Junhyeok Choi1, Seungyeop Yi1, Chaeyun Kim1, Sangkyu Lee1, Idan Szpektor3, Avi Caciularu3, Jongmin Park1, Youngjae Yu1*

1Seoul National University 2University of Massachusetts Amherst 3Google Research

Three-minute overview, with narration.

Can a humanoid pick the right tool and then use it, standing still or on the move? 18 tasks, 55 tools and 3.1k human demonstrations, in simulation and on a real Unitree G1.

The tasks

The robot is told the goal, never the tool

It has to choose a suitable tool from the objects on the table and use it. Three scenarios, three execution levels and two tool-set modes make 18 tasks.

3 scenarios

Each one needs a different tool function.

BallMoveSpatial“Pick the tool and move the ball to the target.” Only a long stick reaches the ball.
BallRetrieveAffordance“Pick the tool and retrieve the ball to the target.” Only a hook can pull it back.
IceBreakPhysical“Pick the tool and break the ice blocks.” Only a stiff, heavy hammer breaks them.

3 execution levels and a decoy

Shown on BallMove. The clip above is L1, stationary use, in standard mode.

L0 · Select and pick upSuccess is lifting the suitable tool by 8 cm. Nothing else is scored.
L2 · Mobile tool useThe goal is out of reach, so the robot steps along the table while holding the tool.
Decoy modeA short stick lies beside the long one. The instruction is the same; the choice is harder.

On the real Unitree G1

Policy rollouts after fine-tuning on 91 real demonstrations, played at 1.5 to 3× speed.

BallMove · SOne long stick and unrelated objects.
BallMove · DA short stick is added as a decoy.
BallRetrieve · SOne hook and unrelated objects.
BallRetrieve · DA straight stick is added as a decoy.

Simulation clips are human demonstrations from ToolBook, seen through the robot's head camera.

The task structure in one figure
Task structure: tool families (sticks, hooks, hammers), the three scenarios with their instructions, the three execution levels, and standard versus decoy tool sets for each scenario.
Task structure. 3 scenarios × 3 execution levels × 2 tool-set modes = 18 tasks, built on 55 tool assets: 9 sticks, 9 hooks and 37 hammers. A decoy differs from the suitable tool in the one property the scenario needs: a 0.30 m stick instead of 0.50 m, a straight stick instead of a hook, a light fly swatter, paint roller or plunger instead of a metal hammer.
What the policy sees and controls
  • Robot. Unitree G1 with 29 body joints, two seven-joint Dex3 hands and a head camera, in MuJoCo and on hardware.
  • Observation. The head-camera image and a 32-value robot state. No object state is given.
  • Action. 36 values at 50 Hz: hand, arm and waist targets, base height, and base motion. Pretrained GR00T Balance and Walk controllers drive the legs.
  • Scenes. Every episode draws new tools, positions, robot spawn, surface materials and lighting.
  • Environment IDs write decoy mode as R, for example G1BallMove-L1-R.

Overview

Choosing the right tool is only the first step

Humanoids need tools to do tasks beyond their physical limits. That takes selecting a suitable tool and coordinating manipulation and, when needed, locomotion. Existing benchmarks do not evaluate these together on a humanoid.

HumanoidToolBench does, with 18 tasks and ToolBook, 3.1k human demonstrations. Seven policies in simulation and three on the real robot show substantial gaps between selecting a tool and completing the task.

  • 18tasks
  • 55tool assets
  • 3,094human demonstrations
  • 7policies in simulation
  • 3policies on a real G1
How it compares with existing benchmarks
BenchmarkEmbodimentTool-use scopeTool assetsSelectionHierarchical levelsDecoy variantsSim demosReal demos
LIBEROSingle-arm06,5000
RoboTwin 2.0Dual-arm11040
HumanoidBenchHumanoid100
SIMPLEHumanoid0––
WOLF-VLAHumanoid000
PhysToolBenchNone (VQA)000
DexCraftSingle-arm630150
RoboWitsDual-arm171,2070
HumanoidToolBenchHumanoid553,00391

tool use is included but is not the primary evaluation scope. 0: not applicable or none released. –: not reported.

Results

Policies that pick up the right tool still often fail to finish the task

Each policy is trained once on the 3,003 simulation demonstrations and run for 100 episodes per task.

Simulation success rate (%)

Bold marks the best policy in a column.

Standard (S)

Policy L0 Select L1 Stationary L2 Mobile
Move Retrieve Break Move Retrieve Break Move Retrieve Break
ACTIL 5 0 0 0 0 0 0 0 0
DPIL 14 5 17 0 0 2 0 0 0
π0.5VLA 29 24 31 4 0 12 1 0 0
Ψ0VLA 19 14 22 9 0 13 1 0 2
GR00T N1.7VLA 77 73 88 64 36 66 10 9 31
Cosmos 3 EdgeWAM 12 0 2 0 0 1 0 0 0
FastWAMWAM 64 62 98 76 57 86 9 20 69

Decoy (D)

Policy L0 Select L1 Stationary L2 Mobile
Move Retrieve Break Move Retrieve Break Move Retrieve Break
ACTIL 3 4 1 0 0 0 0 0 0
DPIL 11 9 13 0 0 1 0 0 0
π0.5VLA 26 28 21 5 3 1 1 1 0
Ψ0VLA 14 15 18 7 3 8 3 1 0
GR00T N1.7VLA 67 59 75 58 30 63 30 5 48
Cosmos 3 EdgeWAM 14 0 0 2 0 0 1 0 0
FastWAMWAM 74 45 93 75 45 82 8 11 61
Move, Retrieve, Break: the three scenarios. IL: imitation learning. VLA: vision-language-action. WAM: world-action model. 0%100%
76%→9%FastWAM, BallMove, L1 to L2

Moving makes it much harder

For every policy, success falls from L1 to L2 whenever L1 success is above zero.

57%→45%FastWAM, BallRetrieve L1, standard to decoy

Decoys hurt even the strongest policies

GR00T N1.7 and FastWAM score lower in decoy mode in all three L1 scenarios.

6/10vs1/10GR00T N1.7 vs ACT, real BallMove

Simulation strength carries over to the real robot

After 5,000 fine-tuning updates on 91 real demonstrations.

Real Unitree G1, L1 tasks

successes out of 10 trials

PolicyBallMove · SBallMove · DBallRetrieve · SBallRetrieve · D
ACT 1/10 0/10 2/10 1/10
GR00T N1.7 6/10 4/10 4/10 4/10
FastWAM 5/10 4/10 4/10 5/10
Policies, training and evaluation settings
PolicyFamilyParametersAction chunkReal robot
ACTIL52.1M100
DPIL84.7M16–
π0.5VLA3.62B16–
Ψ0VLA2.63B30–
GR00T N1.7VLA3.14B40
Cosmos 3 Edge PolicyWAM3.37B16–
FastWAMWAM6.02B32

40,000 updates with 128 samples per update, drawing from all 18 tasks. Episodes last at most 3,000 control steps at 50 Hz, with seeds matched across policies. ACT and DP take language through a frozen CLIP ViT-L/14 text encoder.

Analysis

Where do policies fail?

Every rollout records five events, from touching any candidate to finishing the task.

Share of episodes that reach each event (%)

Averaged over the three scenarios and over L1 and L2.

Standard (S)

Policy Candidate contact Correct first contact Correct-tool lift Tool-target contact Success
ACTIL 46.0 46.0 5.3 2.5 0.0
DPIL 89.8 70.0 15.0 7.7 0.3
π0.5VLA 98.8 72.7 44.2 22.3 2.8
Ψ0VLA 98.1 74.4 38.3 26.6 4.3
GR00T N1.7VLA 99.7 87.3 90.0 79.8 36.0
Cosmos 3 EdgeWAM 39.8 33.2 6.3 2.3 0.2
FastWAMWAM 99.8 87.7 89.2 88.2 52.8

Decoy (D)

Policy Candidate contact Correct first contact Correct-tool lift Tool-target contact Success
ACTIL 57.2 44.8 5.5 2.3 0.0
DPIL 97.0 49.2 12.2 5.7 0.2
π0.5VLA 99.5 44.3 29.7 15.7 1.8
Ψ0VLA 99.7 52.0 25.3 18.8 3.7
GR00T N1.7VLA 100.0 77.2 83.7 74.3 39.0
Cosmos 3 EdgeWAM 51.3 26.3 7.8 2.3 0.5
FastWAMWAM 100.0 74.7 84.5 81.0 47.0
An episode can start with a wrong first contact and still lift the suitable tool later. 0%100%
98.8%→44.2%π0.5, candidate contact to correct-tool lift

Weaker policies touch a tool but do not pick it up

88.2%→52.8%FastWAM, tool-target contact to success

Stronger policies reach the target and still fall short

Same policy, same task, two outcomes

GR00T N1.7 on BallMove L1 in standard mode.

SuccessrolloutThe policy picks the long stick and pushes the ball into the ring.
FailurerolloutThe ball never reaches the ring and the episode runs into the one-minute limit. Time-lapse after the first 8 s.

Does the behaviour follow the tool and the instruction?

Two probes of a GR00T N1.7 policy trained on standard mode only.

Unseen tools

Selection drops when a suitable tool looks unfamiliar

Correct first contact falls from 85% to 74% on BallMove and from 59% to 36% on BallRetrieve.

t-SNE projections of tool-region features for seen and unseen tools, with bars of correct first contact: 85 versus 74 percent for BallMove and 59 versus 36 percent for BallRetrieve.
Bar A is seen tools, bar B is unseen tools.
Changed instructions

The task still gets done under an unrelated instruction

Told “The weather is nice today”, the policy still breaks the ice in 84% of rollouts, against 77% with the original instruction.

Mean right-arm and right-wrist joint angles over episode progress for the original, paraphrased and unrelated instructions, and a bar chart of success rates: 77, 70 and 84 percent.
One fixed IceBreak L1 scene, 100 rollouts per instruction.
Selection and alignment errors in rollouts
Two rows of rollout frames. Top row: the hand picks up a decoy. Bottom row: the tool is held at an angle that misses the ice block.
Reviewed rollouts show decoy pickup (top) and tool-target alignment errors (bottom). Yellow borders mark the frames where spatial alignment errors occur.
Limitations
  • The comparison uses one training run per policy, so variation reflects scenes and rollouts, not training seeds.
  • The real-robot study covers three policies on two L1 scenarios.
  • First contact is a proxy for the initial choice, and execution levels differ in layout as well as in locomotion.
  • The instruction probe uses one fixed IceBreak scene.

ToolBook

3,094 human demonstrations of selecting and using tools

Recorded by eight teleoperators with Meta Quest, in simulation and on the real G1.

1,803Simulation, L0Cut at the moment the tool is lifted
1,200Simulation, L1 and L2100 successes for each of 12 tasks
91Real robot, L1BallMove and BallRetrieve
3,094TrajectoriesLeRobot format, on Hugging Face

Download from Hugging Face

Released under CC BY-NC 4.0 for non-commercial research.

Coverage by task and how it was collected
Simulation (S / D)L0L1L2
BallMove289 / 244100 / 100100 / 100
BallRetrieve388 / 313100 / 100100 / 100
IceBreak316 / 253100 / 100100 / 100
Real robot (S / D)L1
BallMove–29 / 21–
BallRetrieve–20 / 21–
  • 1,959 simulation attempts at L1 and L2 produced the 1,200 successful demonstrations.
  • L0 needs no separate collection. Every attempt that lifts the suitable tool by 0.08 m is cut at that frame: 1,189 come from successful attempts and 614 from unsuccessful ones.

Get started

Evaluate a policy on the 18 tasks

Needs Linux with an NVIDIA GPU, uv, Git and zstd.

git clone https://github.com/SNU-PI/HumanoidToolBench.git
cd HumanoidToolBench
uv run --no-project scripts/setup_evaluation.py

# One-episode check that simulation, inference and recording work
MODEL=snupilab/humanoidtoolbench-act-sim-3003
uv run htb-eval --model $MODEL --episodes 1 --max-steps 100

# Standard protocol: one task, or all 18
uv run htb-eval G1BallRetrieve-L1-R --model $MODEL
uv run htb-eval all --model $MODEL

# Your own policy: predict(request) returns (T, 36) actions
uv run htb-eval G1BallMove-L1-S --policy my_policy:predict \
  --checkpoint MODEL_ID

Citation

The paper is on its way to arXiv, and the BibTeX entry will appear here. Until then, please link to this page or to the code repository.