Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.27051v1 Announce Type: new Abstract: Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retire them. We ask which of those jobs an agent should keep. Our answer is governed self-evolution: the agent may propose, and a frozen statistical referee that the agent cannot touch must judge. The referee scores each candidate only on market outcomes revealed after submission, by betting, so its false-discovery guarantee holds at every stopping time for any proposal policy. We cross three proposers (a script, a bandit and a language model) with this referee and with three deliberately leaky ones,
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Reglering
Related AI news
- Autonomous AI hacks raise thorny questions of legal accountabilityEconomic Times Tech · September 25, 2026
- Experts say that air-gapping AI could prevent events like the Hugging Face hack, but would undermine the value of evaluations and slow research to a crawl (Robert Hart/The Verge)Techmeme · September 25, 2026
- Google’s first Project Suncatcher AI satellite set to blast off into orbit next weekSiliconANGLE · September 25, 2026
- Singapore finance firms aim to train 80,000 workers in AITech in Asia · September 25, 2026
- 8 insights from Proofpoint Protect: Security bets on intent as AI agents join the workforceSiliconANGLE · September 24, 2026
- The PGA of America limits AI sprawl with a streamlined identity architectureSiliconANGLE · September 24, 2026