H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
arXiv cs.AIen
arXiv:2608.00065v1 Announce Type: new Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at substantial indexing, storage, and scoring cost. This mismatch raises a natural question: can context-dependent phrases provide a useful retrieval unit between global vectors and tokens? We introduce H+ Embedding, a unified multi-granularity retriever that predicts variable-length p
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Enterprise AI CRM startup Superleap raises Rs 36 crore from Peak XVEconomic Times Tech · August 4, 2026
- AI is helping Grab ship products more than 30% faster, CFO says, as company raises forecastsCNBC Technology · August 4, 2026
- Exclusive: Zurich-based Exclaim Robotics comes out of stealth, raises $4.95mSifted · August 4, 2026
- Enhancing LLMs with Context-Specific Knowledge for Mitigating Misinformation in SMEs: A RAG-based Modeling and AnalysisarXiv cs.AI · August 4, 2026
- Linguistic Context Recodes Visual Representations in Vision-Language ModelsarXiv cs.AI · August 4, 2026
- Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic CommercearXiv cs.AI · August 4, 2026