FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning

Chuhao Zhou1 Liquan Wang2 Shuxin Cao2 Xiangyu Chen1 Yuxuan Hu1 Boyu Ma1 Animesh Garg2 Jianfei Yang1,*
1MARS Lab, Nanyang Technological University 2Georgia Institute of Technology *Corresponding Author
MARS Lab logo Nanyang Technological University logo Georgia Institute of Technology logo

Research Problem

We define functional generalization as the ability to repurpose diverse tools to achieve the same function. While humans can effortlessly adapt their actions across tools with different shapes and affordances, robots still struggle to transfer functional intent from perception to control. FORGE addresses this challenge by learning functional intermediate representations that bridge visual understanding and grounded action, enabling robots to accomplish the same function with arbitrary tools.

Overview of FORGE

FORGE method overview
Overview of the FORGE. (a) Evaluation and selection of functional intermediate representations. (b) In Stage 1, the System-2 planner learns to predict future keypoint-based functional plans from action-free observations. (c) In Stage 2, the System-1 execution policy grounds the predicted functional plans into executable robot actions.

Demos

Simulation Demos

Scoop

Baseline
ATM
FORGE (ours)

Shoe

Baseline
ATM
FORGE (ours)

Book

Baseline
ATM
FORGE (ours)

FM and ATM often fail to hit the target with the specified hitting points.
FORGE succeeds more reliably with predicting keypoint trajectories.

Real-World Demos

Tobject

Hitting Pattern 1

FM

FORGE (ours)

Hitting Pattern 1 requires the policy to hit the target with the hammer-like head, the FM baseline incorrectly uses the handle to strike the target.

Hitting Pattern 2

FM

FORGE (ours)

Hitting Pattern 2 requires the policy to hit the target with the handle, the FM baseline misses the hit.

Shoe

Hitting Pattern 1

FM

FORGE (ours)

Hitting Pattern 1 requires the policy to hit the target with the front point of the shoe, the FM baseline misses the hit.

Hitting Pattern 2

FM

FORGE (ours)

Hitting Pattern 2 requires the policy to hit the target with the rear point of the shoe, the FM baseline misses the hit.

Book

Hitting Pattern 1

FM

FORGE (ours)

Hitting Pattern 1 requires the policy to hit the target with the bottom edge, the FM baseline fails to complete the downward motion.

Hitting Pattern 2

FM

FORGE (ours)

Hitting Pattern 2 requires the policy to hit the target with the book spine, the FM baseline misses the hit.

Results

Benchmark settings
Simulation Benchmark and Real-world Setting. (a) The simulation benchmark contains seven tools, three initial settings per tool, and randomized target and hitting points. (b) The real-world setting uses a Franka robot and includes three tools, each with two hitting patterns.

Simulation Results

Simulation results table
We evaluate success rates on three unseen simulation tools, with each setting tested over 30 rollouts and reported at setting, tool, and overall levels. FORGE achieves 0.36 overall SR, outperforming DP by more than 2x and showing that function-aware keypoint plans improve generalization beyond end-to-end or generic tracking baselines.

Real-World Results

Additional real-world results table
We train on three seen tools and evaluate on Tobject, Book, and Shoe, with two hitting patterns and 10 real-world rollouts per pattern for each unseen tool. FORGE improves the overall success rate from 0.22 to 0.63 over FM, confirming that explicit functional plans help align the correct hitting region on novel real-world tools.