An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures
arXiv cs.AIen
arXiv:2609.36104v1 Announce Type: new Abstract: Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic capt
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment TestingarXiv cs.AI · September 30, 2026
- Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge DevicesarXiv cs.AI · September 30, 2026
- More Programs or More Rolls? Separating Coverage from Specialization in LLM HarnessesarXiv cs.AI · September 30, 2026
- Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language ModelsarXiv cs.AI · September 30, 2026
- Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled ObjectivesarXiv cs.AI · September 30, 2026
- SAGE: A Statistical Acceptance Gate for Self-Evolving AgentsarXiv cs.AI · September 30, 2026