Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation

SimVLA

1Department of Artificial Intelligence, Yonsei University  ·  2Department of Computer Science and Engineering, Seoul National University

SRL

Overview

Mobile Manipulation. Train in Simulation. Deploy in Real Home.

End-to-end framework for zero-shot sim-to-real VLA training for mobile manipulation.

0
hours of teleoperation
0
real-world demonstrations
0
fine-tuning in the target scene
35
mobile manipulation tasks
100
procedurally generated houses
3
real environments
mock kitchen · pantry · home
3
robot embodiments
Anubis · AI Worker · RB-Y1

Method

End-to-End Synthetic Data Generation

Stage 01

Kitchen Scene Generation

Scenes come from Scene Synthesizer, objects come from Objaverse, and run in Isaac Lab.

building…
shift + drag a door, drawer or the fridge to open it ↵

Data & Scale

Diverse Mobile Manipulation Tasks

 

1 task, 100 homes

One task applies to every kitchen — massive, parallel and diverse robot data generation.

Video task: put bowl in drawer — every tile is a different generated kitchen.

Training

SimVLA. Trained in Two Stages

Stage 1

Pre-training

on SimAction & SimVQA

Try me
Objectives

Three objectives in a single forward pass — flow matching, tokenized actions, SimVQA answers — align the backbone with the robot’s action space. Gradients from the fresh action expert are stopped at the backbone.

Stage 2

Post-training

on SimAction & SimDeploy

Try me
Objective

Flow matching alone, on SimAction and SimDeploy. The expert is warmed up, so the stop-gradient lifts and the backbone trains too.

Real World

Zero-Shot Sim-to-Real

SimVLA is deployed in three different real-world environments using the Anubis robot.

Mock kitchen ID

The mock kitchen in its default configuration, where the real-world data for the baseline π0.5 was collected.

Mock kitchen OOD-pos

The same room with objects and furniture repositioned.

Mock kitchen OOD-vis

The same room with backgrounds and lighting changed.

Pantry

A different room entirely—with tiled walls, a water dispenser, and a practical, lived-in layout.

Real home

Kyoungin's home as it is.

Fully Autonomous

 

Experiments

Experiment Analysis

Averaged task progress in five real-world conditions

Task progress — the share of predefined subgoals completed — averaged over the tasks in each setting.

  • SimVLA generalizes across environments. 53–58% task progress across all five settings, with zero real-world demonstrations.
  • In-domain data improves ID, not OOD. π0.5 with 200 demos reaches 67.8% ID, but only 21.3% pantry and 29.0% real home.
  • Simulation pre-training is critical. Without SimAction/SimVQA pre-training, performance stays below 19% across all settings.
Pre-training and post-training data ablations per task

Task progress per task in mock kitchen-ID.

  • Both pre-training sources matter. Removing SimVQA or TokSimAction consistently reduces performance across tasks.
  • SimDeploy yields task-dependent gains, with larger scene-level generalization improvements on sink/drawer tasks than on pour bottle to mug, which tests object generalization.
What does SimVQA pre-training actually buy?

Pre-training without SimVQA makes it reach imprecisely.

Pre-trained without SimVQA — T1 / T2 / T3

T1
T2
T3

That points at the backbone's spatial grounding, so we probed the backbone directly:
Can the model answer spatial questions about a real Anubis frame?

SimVLA answering spatial questions about real-world Anubis observations — the same backbone, queried in language rather than actions.
ModelMean r ↑Mean #Dist ↑
SimVLA (ours)+0.9132.7
InternVL3-8B+0.613.3
Qwen2.5-VL-7B+0.502.7
PaliGemma-3B (backbone) — —

Gripper–object distance on real Anubis episodes.

  • Ground truth from a RealSense D455 and forward kinematics.
  • Mean r measures ordering, not scale; #Dist counts distinct predictions.
  • PaliGemma-3B is the same backbone without SimVQA.
How does SimVLA learn retry?

Retry behavior emerges during SimDeploy. When SimVLA observes a missed grasp, it reopens the gripper, returning to a state similar to the one before the attempt, and tries again.

Retry: observe → reopen → reattempt

  1. 1
    Observe the missed grasp

    The object remains outside the gripper after the grasp attempt.

  2. 2
    Reopen the gripper

    Opening the gripper brings the robot back to a familiar pre-grasp state.

  3. 3
    Try again

    From this familiar state, the policy attempts another grasp.

Retry behavior during SimDeploy — SceneSmith.

BibTeX

% Citation will be updated upon acceptance.
@misc{baik2026simvla,
  title  = {SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation},
  author = {Baik, Kyoungin and Lee, Youngwoon},
  year   = {2026}
}