Deconstructing Spatial Intelligence in Vision-Language Models

Published in TechRxiv preprint, 2025

This work analyzes the component capabilities and limitations underlying spatial intelligence in vision-language models.

Recommended citation: Disheng Liu, Tuo Liang, Zhe Hu, Jierui Peng, Yiren Lu, Yi Xu, Yun Fu, and Yu Yin. "Deconstructing Spatial Intelligence in Vision-Language Models." TechRxiv, 2025.
Download Paper