What does SimVQA pre-training actually buy?
Pre-training without SimVQA makes it reach imprecisely.
Pre-trained without SimVQA — T1 / T2 / T3
That points at the backbone's spatial grounding, so we probed the backbone directly:
Can the model answer spatial questions about a real Anubis frame?
| Model | Mean r ↑ | Mean #Dist ↑ |
|---|---|---|
| SimVLA (ours) | +0.91 | 32.7 |
| InternVL3-8B | +0.61 | 3.3 |
| Qwen2.5-VL-7B | +0.50 | 2.7 |
| PaliGemma-3B (backbone) | — | — |
Gripper–object distance on real Anubis episodes.
- Ground truth from a RealSense D455 and forward kinematics.
- Mean r measures ordering, not scale; #Dist counts distinct predictions.
- PaliGemma-3B is the same backbone without SimVQA.