ROBOT LEARNING THROUGH GEOMETRY

SkeleWAM

Skeleton World-Action Modeling
for Efficient Robotic Manipulation

Juyi Sheng · Hua Wang · Mengyuan Liu

Peking University emblemPeking University

A world action model that learns to manipulate
through sparse 3D robot–object structure.

FROM OBSERVATION TO SKELETONReal-world visualization
Keep the geometry. Learn the interaction.RGB with landmarks 3D skeleton
85.9%

LIBERO-Plus success

57.1M

Model parameters

93.4%

Camera perturbation success

89%

Real-world average success

Results reported in the manuscript · RGB-D observation setting

01 / THE IDEA

A little structure goes a long way.

Robot joints, object centers, and interaction points form a shared geometric language for action and prediction.

World action models combine action generation with future-state prediction. SkeleWAM represents the scene as a compact 3D skeleton, built online from RGB-D observations and robot proprioception. The skeleton makes the geometry of robot–object interactions explicit.

During training, the model learns actions and future skeletons together. At inference, it generates actions from the current skeleton and language instruction. Future prediction supports learning without requiring visual reconstruction.

Comparison of video, visual latent, dense 3D dynamics, and SkeleWAM's sparse skeleton representation.
A compact state space. Explicit robot–object geometry connects perception, prediction, and control. Open figure
02 / SEE THE STRUCTURE

Five tasks

Explore recorded robot interactions and their skeleton visualizations, from multiple viewpoints.

LANGUAGE INSTRUCTION

“Open drawer.”

Front camera · RGB + reconstructed skeleton

Skeleton views show reconstructions from recorded interactions. RGB clips are separate task recordings. Success rates below are the evaluation results reported in the manuscript.

03 / HOW IT WORKS

Predict geometry. Generate action.

Future skeletons provide geometric supervision during training. At inference, the model generates actions directly.

01

Build the skeleton

Combine object landmarks from RGB-D perception with robot joints from forward kinematics, in a shared robot-centric frame.

02

Learn future geometry

Jointly train action generation and future skeleton prediction with flow matching, using the current skeleton and language context.

03

Select and execute

Medoid Action Consensus selects a representative sampled action chunk. Execute a prefix, update the skeleton, and replan.

SkeleWAM framework: skeleton construction, language conditioning, world and action experts, future skeleton training, and Medoid Action Consensus.
The SkeleWAM framework. The future-skeleton branch is removed at inference; MAC selects a complete trajectory without averaging candidates. Open figure

Learning the future helps the present. Adding future-skeleton supervision improves LIBERO-Plus success from 80.1% to 85.9%, with the same action-only inference procedure.Table 3b, manuscript

+5.8percentage points
04 / EXPERIMENTS

Compact model. Robust manipulation.

Zero-shot evaluation on 10,030 LIBERO-Plus variants across seven perturbation categories.

OVERALL SUCCESS85.9%

+3.7 points over Cosmos-Policy

CAMERA PERTURBATIONS93.4%

+17.6 points over Cosmos-Policy

MODEL SIZE57.1M

Parameters in the RGB-D setting

LIBERO-Plus

Success rate (%) · selected baselines
MethodParametersOverall ↑Camera ↑Layout ↑
π0.53.3B79.770.684.1
VLA-JEPA~2B77.964.283.9
GE-Act2.2B80.360.780.2
Cosmos-Policy2B82.275.882.2
SkeleWAM OURS57.1M85.993.466.6

Table 1, manuscript. SkeleWAM uses RGB-D observations. The separate sim-state reference uses privileged simulator coordinates and reaches 87.7% overall.

Results across all seven perturbation categories
MethodCameraRobotLanguageLightBackgroundNoiseLayout
Cosmos-Policy75.863.381.796.588.992.782.2
SkeleWAM93.471.989.194.796.093.966.6

Where there is room to improve. Layout changes remain challenging: SkeleWAM reaches 66.6% compared with 84.1% for π0.5. Sparse geometry alone does not resolve generalization to substantially different spatial arrangements.

BEYOND SIMULATION

On the real robot.

An ARX R5 with external and wrist-mounted RealSense cameras. Five tasks, 20 trials per task, for each method.

89%

Mean success across five tasks

Representative successful rollouts for opening a drawer, placing a block in a drawer, and stacking bowls, with early, interaction, late, and skeleton views.
From contact to completion. Three representative tasks from the five-task real-world evaluation (Figure 3). Open figure

Real-world success rates

20 trials per task · success rate (%)
MethodOpen
drawer
Close
drawer
Stack
blocks
Stack
bowls
Block in
drawer
Average ↑
Fast-WAM808580858583
π0.5808070757576
Cosmos-Policy858590908587
SkeleWAM809090959089

Table 2, manuscript. The average is the unweighted mean across the five tasks.

View the experimental setup
ARX R5 robotic arm in a tabletop workspace with external and wrist-mounted RealSense cameras, drawers, blocks, and bowls.
Figure 4. The real-world experimental setup.

Citation

@unpublished{sheng_skelewam,
  title  = {SkeleWAM: Skeleton World-Action Modeling
            for Efficient Robotic Manipulation},
  author = {Sheng, Juyi and Wang, Hua and Liu, Mengyuan},
  note   = {Manuscript}
}