OneLA: Scaling Linear-Attention Decoding to Large Beams in Generative Recommendation
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.12399v1 Announce Type: new Abstract: Generative recommendation (GR) relies on large-beam decoding to generate hundreds of candidate items, creating a new scaling challenge for recurrent linear attention. Existing linear attention serving systems either materialize a full recurrent state for every beam or repeatedly replay shared history, incurring substantial memory and traffic overhead. To address this, we present OneLA, a linear-attention decoding framework that exploits the shared prompt and short divergent suffixes of GR workloads. Specifically, OneLA represents all beam states using a single shared prompt-derived state and compact, append-only records of their divergent trans
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Sam Altman calls for pacing AI development but promises rapid progress will continueThe Decoder · September 14, 2026
- « Si Anthropic, OpenAI et xAI font une pause, la belle histoire boursière des fabricants de puces risque de s’enrayer »Le Monde Pixels · September 14, 2026
- OpenAI boss Sam Altman spells out how and why the AI industry wants to slow down: 'We could lose control'CNBC Technology · September 14, 2026
- iOS 27: Diese Features fehlen zum Start – neben Siri AIheise online – KI · September 14, 2026
- Samsung, SK Hynix reject US$19 billion power prepayment as Korea speeds chip expansionDIGITIMES · September 14, 2026
- How hyperscalers like Amazon, Microsoft, and Google are siding with consumers on data center power costs and sweetening offers for communities to gain support (Ann Davis Vaughan/The Information)Techmeme · September 14, 2026