Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.25165v1 Announce Type: new Abstract: In this report, we introduce \textbf{Ovis-Embedding}, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make \textbf{three key advances}: (1) \textbf{native omni-modal initialization}: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) \textbf{data-centric omni-modal training}: we construct a broad, high-quality corpus span
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Bild
Related AI news
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026
- Making Agents More Consistent: Skills Should Form Habits for Repeat TasksarXiv cs.AI · September 23, 2026
- Real-Time Hand Gesture Recognition for OpenXR Using Transformer-Based Machine LearningarXiv cs.AI · September 23, 2026
- Efficient Iterative Retrieval with Heterogeneous BatchingarXiv cs.AI · September 23, 2026
- ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent RuntimesarXiv cs.AI · September 23, 2026
- Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future DirectionsarXiv cs.AI · September 23, 2026