Spatial Reasoning for Vision-Language-Action Models

I am investigating how vision-language-action models translate spatial language and visual observations into reliable robot actions.

Baseline and reference-aware spatial reasoning on a RoboSpatial example

Illustrative spatial-reasoning evaluation on RoboSpatial. The same scene and spatial instruction can imply different geometric directions depending on the referenced object frame. RoboSpatial imagery and empty-space annotations are used as an external benchmark; this visualization illustrates geometric grounding rather than a physics rollout.

The project studies spatial reasoning, reference-dependent instructions, compositional generalization, and failure modes of end-to-end VLA policies across simulated manipulation and navigation settings.

Current work focuses on making spatial information available to robot policies in a more explicit and inspectable form. This is ongoing research; additional technical details will be released in future updates.

Keywords: vision-language-action models, embodied robot control, spatial reasoning, compositional generalization, manipulation, navigation.