
How Box is unlocking multimodal enterprise agents with Gemini Embeddings 2
Google Cloud and Box are integrating Gemini Multimodal Embeddings 2 into the Box Agentic Platform to enable multimodal AI. This allows the system to handle text, images, tables, and charts within digital content.
Why it matters
Users in sectors like healthcare and finance can now query visual data alongside text, speeding up diagnoses and financial auditing. It reduces the need for manual tagging of visual assets.
The details
- Supports retrieval across .docx, .xlsx, .pdf, .pptx, .png, and .csv formats.
- Enables natural language queries to find specific charts without manual tagging.
- Preserves spatial geometry and visual hierarchies of documents during embedding.
Show entities and relationshipsHide entities and relationships
In this article
Key connections
Box owns Box Agentic Platform
Box makes and operates the Agentic Platform.
Box owns Box Intelligent Content Management
Box makes and operates the Intelligent Content Management platform.
Google Cloud and Box are integrating advanced multimodal capabilities into Box's Agentic Platform.
Box Agentic Platform uses Gemini Embedding 2
Box's Agentic Platform is powered by Gemini Multimodal Embeddings 2.
Box Intelligent Content Management uses Gemini Embedding 2
Box's Intelligent Content Management platform is powered by multimodal embeddings from Gemini Embedding 2.
Box Agentic Platform uses Retrieval-Augmented Generation
Box's Agentic Platform extends traditional RAG architectures with multimodal embeddings to handle tables, images, and mixed file formats.
Get the weekly recap
The stories like this one, picked and explained — once a week, straight to your inbox.