Open Spatial Models
An Open Spatial Intelligence Model for General Spatial Reasoning
Overview
SpatialAxiom is a spatial intelligence model built on the Qwen3.5 family with large-scale spatial supervision spanning indoor scenes, egocentric views, and multi-camera settings, delivering strong general spatial reasoning without altering the base architecture.
On average, our models surpass proprietary models and larger open-source alternatives, with leading results on VSI-Bench, MMSI-Bench, MindCube, ViewSpatial, and EmbSpatial.
A systematic taxonomy of spatial tasks, balanced task distribution, and data synthesis to raise data quality. SpatialAxiom is trained purely with full-parameter SFT and serves as a clean starting point for downstream fine-tuning or RL.
Inherits the Qwen3.5 vision-language model architecture, preserving a general-purpose multimodal design without task-specific architectural modifications.
SpatialAxiom-9B and SpatialAxiom-35B-A3B are publicly released on Hugging Face and ModelScope, compatible with transformers and vLLM out of the box.
Evaluation
Measured on 8 spatial reasoning benchmarks.