Mistral introduced Robostral Navigate, an 8B vision-language model designed for autonomous robotic navigation using a single RGB camera.
- First model built for embodied navigation
- Achieves 76.6% success rate on unseen R2R-CE benchmarks
- Uses only a single RGB camera, without depth sensors or LiDAR
- Trained entirely in simulation using prefix-caching to reduce token counts by 22x
- Post-trained with the CISPO online reinforcement learning algorithm to improve success rates by 3.2%