Not Every Call Needs a Frontier Model: Per-Call-Site Evaluation of Small Language Models in a Deployed Agentic Home-Automation System
arXiv cs.AIen
arXiv:2610.09021v1 Announce Type: new Abstract: An agentic system issues several structurally different kinds of LLM calls. It routes intent, classifies actions, grounds language in a device registry, plans multi-agent pipelines and writes the Python code those pipelines run. The difficulty of these call sites varies by an order of magnitude, yet in practice a single model, chosen for the hardest site, serves all of them. In this work, we evaluate 9 models from 0.8B to a frontier hosted model across the five call sites of a deployed open-source home-automation framework (Wactorz), using its unmodified production prompts and two real Home Assistant installations (280 cases, 2520 scored calls)
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Företag
Related AI news
- Sources: Isomorphic Labs, an AI drug discovery startup spun out of Google DeepMind, is in early talks to raise funds at a valuation of at least $40B (Bloomberg)Techmeme · October 8, 2026
- Hanmi wins rare Samsung order amid US$5 billion chip substrate expansionDIGITIMES · October 8, 2026
- Singapore teams with Penn lab on resilient military robotsTech in Asia · October 8, 2026
- US venture deal value reaches record $515.8B as exits fail to keep paceSiliconANGLE · October 8, 2026
- How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault AnalysisarXiv cs.AI · October 8, 2026
- When the Governor Becomes the Disturbance: Control-Generated Disturbance and Cost-Aware Backoff in Governed Tool-Using AgentsarXiv cs.AI · October 8, 2026