Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
Apple Machine Learningen
Apple Machine Learning
AI Global WireSystem-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision…
This is a short summary published by AI Global Wire. The full article is owned and hosted by Apple Machine Learning — open it there to read it in full.
Read the full story at Apple Machine Learning- Verktyg
- Forskning
Related AI news
- Sakana AI hires Jürgen Schmidhuber, inventor of deep learning, world models, and your next ChatGPT updateThe Decoder · September 24, 2026
- Meta is going to let you build games with AI right on your phoneThe Verge AI · September 24, 2026
- Google's Suncatcher project aims to put AI data centers in orbit powered by solar energyThe Decoder · September 24, 2026
- Meta’s Muse Charm looks like a Tamagotchi, but it’s tapping into a much newer trendTechCrunch AI · September 24, 2026
- AI agents creating a new insider security risk: theCUBE’s Oktane keynote analysisSiliconANGLE · September 24, 2026
- Sources: the White House asked OpenAI and Anthropic not to share new models with UK's AISI until US reviews them; Anthropic appears to have agreed (Politico)Techmeme · September 24, 2026