MindTopo reveals VLMs’ spatial reasoning abilities
This helps develop reliable robots and interactive assistants that must track structural relationships to make correct decisions in physical environments.
- Organizes tasks across five topological categories: continuity, separation, order, enclosure, and knots.
- Evaluates models at two cognitive levels: static scene reasoning and interactive simulation planning.
- Identifies a significant performance gap where multimodal models perform better at static perception than interactive planning.