On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
arXiv cs.AIen
arXiv:2610.10833v1 Announce Type: new Abstract: We study whether small LLM agents can operate effectively under explicit wall-clock time budgets by both respecting the allocated runtime and using available time productively. We evaluate Qwen3.6-27B on five competitions from MLE-Bench Lite and Qwen3-4B on Zork I (Jericho), two agentic benchmarks where additional computational time can meaningfully improve performance. In the simplest setting, where the budget is stated only in the prompt, agents fail to translate the stated budget into controlled use of time. These failures arise from gaps in time awareness, since the harness provides no timing feedback, but also because they cannot reliably
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and DefensearXiv cs.AI · October 9, 2026
- Curating Always-Loaded Context for LLM Agents: A Capacitated Assortment Model with Censored FeedbackarXiv cs.AI · October 9, 2026
- AgentHorizon: Evaluating Agentic Judges for Long-Horizon Computer-Use TasksarXiv cs.AI · October 9, 2026
- When Lower Reconstruction Loss Hurts: Distributionally Robust Refinement for Low-Bit LLM QuantizationarXiv cs.AI · October 9, 2026
- Plan-and-Patch: Diffusion Language Models for Agentic PlanningarXiv cs.AI · October 9, 2026
- Whose Ground Truth? Embracing Ambiguity in Human-Centered AIarXiv cs.AI · October 9, 2026