Birthday Party Setup

MM-ABCTowards Generalist Mobile Manipulation
via Seeing, Coordinating and Imagining

01 / Method

See and imagine.
Move as one.

MM-ABC architecture
Figure 2. MM-ABC architecture. Left: sparse VLM features condition MM-APT; a training-only future branch receives geometric supervision from frozen VGGT-Omega features. Middle: joint transformer blocks (MM-JiT in the diagram) combine stream-specific transformations with masked attention and read-only perceptual context. Right: action streams communicate within each segment, and far tokens additionally read near tokens. Future queries follow the same temporal ordering but remain isolated from actions.
Figure 2. MM-ABC architecture. Left: sparse VLM features condition MM-APT; a training-only future branch receives geometric supervision from frozen VGGT-Omega features. Middle: joint transformer blocks (MM-JiT in the diagram) combine stream-specific transformations with masked attention and read-only perceptual context. Right: action streams communicate within each segment, and far tokens additionally read near tokens. Future queries follow the same temporal ordering but remain isolated from actions.
01DeepStack

Seeing.

For seeing, MM-ABC uses a sparse DeepStack interface that injects multi-level VLM features into the action expert, preserving complementary fine-grained and semantic information.

02MM-APT

Coordinating.

For coordinating, we introduce the Mobile Manipulation Action Prediction Transformer (MM-APT), which maintains separate manipulation and mobile/body streams while enabling cross-stream interaction through masked joint attention.

03World Expert

Imagining.

For imagining, learnable future queries predict geometry-rich representations at sparse horizons under a frozen geometric teacher, capturing workspace evolution and the geometric intent the policy should realize.

02 / Benchmarks

Tested across
five benchmarks.

Success rate (%)

EBench

44.71%

We post-train MM-ABC for 100k steps with a batch size of 512. We evaluate a single checkpoint on the official held-out test split, reporting task success and stage-wise progress.

Success rate (%)
  1. MM-ABC44.71
  2. π₀.₅41.41
  3. InternVLA-A1.534.17
  4. π₀33.72
  5. GigaBrain-0.733.27
  6. Cosmos3-Edge29.29
  7. Fast-WAM25.64
  8. X-VLA23.72

03 / Real world

From the party table
to the laboratory.

01 / 053× speed

Birthday Party Setup

The robot picks up a cake from the table directly ahead, carries it to the decorated table on the left, and sets it down. It then picks up a birthday candle, passes it between its two hands, and inserts it into the cake. Finally, it moves beside the table and turns on the speaker to play music.

75%Success rate (%)
Birthday Party Setup

04 / Dataset

Many embodiments.
A shared action space.

5,000+Hours
400K+Episodes
12Datasets
17Embodiments
Figure 3. Overview of the MM-ABC multi-embodiment pretraining corpus. The corpus contains 5,166.2 cleaned hours from 12 datasets and 51 subsets, spanning 17 embodiments. Sector areas show training sampling probabilities aggregated by dataset; the surrounding panels illustrate environments and embodiments.
MM-30

Base, lift, and arms.
Moving together.

To complement public datasets, we collect MM-30, a real-world mobile manipulation dataset with coordinated base, lift, and dual-arm motion.

30+Hours
40+Tasks
450+Objects
Figure 5. MM-30: self-collected mobile manipulation dataset. Representative tasks, object and scene diversity, and episode counts per task.

05 / Citation

Build on
MM-ABC.

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

BibTeX
@misc{liang2026mmabc,
  title   = {MM-ABC: Towards Generalist Mobile Manipulation
             via Seeing, Coordinating and Imagining},
  author  = {Liang, Qiwei and Chen, Guangyu and
             Zhu, Shaolong and Xiao, Zikuan and
             Lu, Jinxuan and Xie, Yifan and Xu, Renjing and
             Ding, Wenbo and Chen, Tianxing},
  year    = {2026},
  url     = {https://mm-abc.github.io/}
}