AI research news
Research drives everything else. This page surfaces new papers and studies — reasoning, agents, training methods, evaluation and safety — with short summaries so you can tell in seconds whether a result matters to you.
Latest stories

Alibaba open-sources medical AI model that can detect cancer and nearly 150 conditions
Alibaba Group Holding’s research arm, Damo Academy, has open-sourced an artificial intelligence model capable of identifying nearly 150 abdominal conditions – including cancers – by reading computed tomography (CT) scans, marking the latest step in the firm’s growing medical AI efforts. The vision-language model, called Damo Radar, was designed to analyse contrast-enhanced CT scans covering 18 abdominal organs and identify a broad range of diseases and other abnormalities, such as malignant...
SCMP Tech

Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think
For the first time, Anthropic is releasing metrics on how it builds its own AI. Claude already "leads" 26 percent of the work on future models, up from under one percent in February. But the underlying scale is fuzzy, the scoring comes from Claude itself, and "lead" means less than it sounds. The article Anthropic wants you to know Claude leads a quarter of its research, but "lead" doesn't mean what you think appeared first on The Decoder .
The Decoder

Researchers used Anthropic’s Claude to hack into OpenAI
Security researchers used Anthropic’s Claude to exploit vulnerabilities in OpenAI’s systems, taking over employee accounts and gaining access to an internal code repository before reporting the flaws.
TechCrunch AI

New experts join Google’s AI & Economy team
We are expanding our AI & Economy team with world-class academic advisors, fellows, and core internal researchers.
Google AI

Allt fler tror att AI kommer ta jobb istället för att skapa nya
En ny undersökning från Pew Research Center visar att en majoritet av människor världen över nu räknar med att artificiell intelligens kommer leda till färre snarare än fler jobb. Undersökningen baseras på svar från 37 länder. I 34 av dessa är det vanligare att tro att AI kommer innebära färre arbetstillfällen under de kommande 20 åren. Oron är större i rikare länder. Även i Sverige syns tecken på en växande skepsis inför tekniken. Andelen svenskar som är mer oroade än entusiastiska över den ökade användningen av AI har stigit med nio procentenheter på ett år. Det är den största ökningen bland länderna som jämförts över tid. Oron har också ökat bland både yngre och äldre svenskar. Globalt be
Computer Sweden
Significant AI competency gap among management graduates: Study of top B schools
While industry expects MBA graduates to demonstrate competencies at 4/5 (Good), MBAs from the Top 75 B-schools are rated at an average of 3.2/5.
Economic Times Tech
Europe's AI firms, playing catch-up, challenge US calls for slowdown
Anthropic's CEO Dario Amodei in an online essay urged firms to slow the pace at which they improve AI model capabilities and called for an antitrust waiver allowing leading developers to jointly develop safety measures after researchers warned about the dangers of AI, including that it could wipe out humanity.
Economic Times Tech
AI lowering entry barriers for small contractors, reshaping blue-collar services: Report
Artificial intelligence (AI) is lowering some of the traditional barriers for skilled workers looking to start or expand small businesses by taking over administrative tasks such as scheduling, dispatch, paperwork and contract analysis, according to a report by SkyPSI, a tech company.
Economic Times Tech

Anthropic Life Sciences Head Eric Kauderer-Abrams says the AI company set up a Bay Area wet lab for physical biology work, as it pushes into AI disease research (Reuters)
Reuters : Anthropic Life Sciences Head Eric Kauderer-Abrams says the AI company set up a Bay Area wet lab for physical biology work, as it pushes into AI disease research — Anthropic has quietly set up a laboratory to do physical biology work as it pushes its artificial-intelligence ambitions into the world of drug science.
Techmeme

Top Chinese AI models make 10% of OpenAI, Anthropic revenue despite high valuations: report
Chinese artificial intelligence models from seven major developers combined generate only about 10 per cent of the revenue reported for OpenAI and Anthropic, even as investor enthusiasm remains intense, according to US-based research firm Rhodium Group. The Chinese AI models earned an estimated US$10.7 billion in annual recurring revenue (ARR) from March to August, a fraction of the more than US$100 billion combined for OpenAI and Anthropic, according to the Rhodium report published on...
SCMP Tech
Nordic AI (EN)
AI Global WireTech leaders urge AI reskilling as job-loss fears grow globally
A new Pew Research Center survey found majorities in 34 of 37 countries expect AI to eliminate more jobs than it creates ...
Nordic AI (EN)

Security researchers in OpenAI's bug bounty program hacked OpenAI in July and accessed its "monorepo" on GitHub, using Opus 4.8 for cybersecurity and Opus 5 (Robert McMillan/Wall Street Journal)
Robert McMillan / Wall Street Journal : Security researchers in OpenAI's bug bounty program hacked OpenAI in July and accessed its “monorepo” on GitHub, using Opus 4.8 for cybersecurity and Opus 5 — A bug-hunting independent security research team was able to access OpenAI's internal code system, exposing growing risks in automated cyber threats
Techmeme
arXiv cs.AI
AI Global WireWhen2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
arXiv:2609.19671v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic inefficiency: they often overthink easy problems and underthink hard ones. Existing approaches based on uniform length penalties or rigid routing incur an efficiency tax, trading reduced computation on easy instances for accuracy loss on hard instances. We formulate efficient reasoning as an instance-adaptive computation allocation problem and propose When2Think, a post-training framework for hybrid reasoning that dynamically allocates computation based on problem difficulty. Our method introduces Instance-level Difficulty-Aware Control (IDAC), a re
arXiv cs.AI
arXiv cs.AI
AI Global WireEconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data
arXiv:2609.19523v1 Announce Type: new Abstract: Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data. Each skill records its scope, navigation procedure, site-specific guidance, verification checks, and recovery steps while replacing source-instance values with placeholders. EconSkills separates two questions: whether a known relevant procedure transfers to a held-out task, and whether an agent can retain that benefit when
arXiv cs.AI
arXiv cs.AI
AI Global WireThe syntax and semantics of goals
arXiv:2609.19448v1 Announce Type: new Abstract: In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly combine with world knowledge to organize and specify purposeful behavior. In this way, goals are compositional representations whose content relates to rational behavior. We here draw attention to goals as representations and their content because it highlights a parallel with other areas in cognitive science - in particular, the syntax-semantics interface in linguistics and logic - while also foregrounding foundational questions about the expressivity, design, and efficiency of different goal representations. For example, goals are typica
arXiv cs.AI
arXiv cs.AI
AI Global WireAgentic AI Networking for Heterogeneous Unmanned Aerial Systems in Low-Altitude Wireless Networks
arXiv:2609.19538v1 Announce Type: new Abstract: Low-altitude wireless networks (LAWNs) are emerging as a key infrastructure for heterogeneous unmanned aerial systems that support concurrent services within a shared three-dimensional airspace. Their coexistence creates strong coupling among mobility, connectivity, and shared network resources, while heterogeneous services impose distinct and time-varying requirements. These interactions naturally form a dynamic non-cooperative game in which both operating conditions and coordination objectives evolve over time. Conventional optimization and learning-based controllers typically rely on predefined objectives, limiting their ability to adapt aut
arXiv cs.AI
arXiv cs.AI
AI Global WireLearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents
arXiv:2609.19721v1 Announce Type: new Abstract: Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed documented conditions, specificity errors, and procedure-coding convention mismatches. We introduce Learn-Then-Act, an inference-time adaptation framework that converts errors from a small labeled LEARN batch into a structured Mistake Knowledge Database (MistakeKDB). False-negative lessons are routed to a recall-oriented Coder, while false-positive lessons are routed to a precision-oriented Judge. We instantiate the framework in LearnActCoder, a Coder-Judge clinical coding pipeline with lookup-table grounding where available. On 150 matched M
arXiv cs.AI
arXiv cs.AI
AI Global WireCompositional Reasoning in Language Models under Reinforcement Learning Post-Training
arXiv:2609.19465v1 Announce Type: new Abstract: Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation
arXiv cs.AI
arXiv cs.AI
AI Global WireAn Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
arXiv:2609.19519v1 Announce Type: new Abstract: Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the
arXiv cs.AI
arXiv cs.AI
AI Global WireClosed-World Resolution Against Tool Hallucination in LLM Agents
arXiv:2609.19425v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents fail in a way no tool-selection or tool-security method addresses: they call tools that do not exist and pass arguments no schema declares. Existing defenses either pick the right tool (selection) or constrain what an agent may do with real tools (gating), both of which presuppose the emitted call refers to a real tool at all. We show this is a structural blind spot: a hallucinated call is by construction not a decision any gate made, so no gate can reject it. This paper is primarily a measurement and benchmark study. We give a five-class taxonomy of tool hallucination (H1-H5) and, as a reference
arXiv cs.AI
arXiv cs.AI
AI Global WireDo AI Agents Understand Computer Architecture?
arXiv:2609.19387v1 Announce Type: new Abstract: Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. An agent that improves an accelerator may be reasoning about the machine, or may be searching competently over knobs whose meaning it never recovers -- and only the first transfers to the next architecture. Existing evaluations cannot tell the two apart, because they vary the agent while holding the framing of the problem fixed. We do the opposite. AutoTuring hands the same agent the same 15-dimensional accelerator space twice: once as named architectural knobs with simulator counters,
arXiv cs.AI
arXiv cs.AI
AI Global WireAutoData: Agentic Search for Pre-training Data Selection
arXiv:2609.19754v1 Announce Type: new Abstract: LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discove
arXiv cs.AI
arXiv cs.AI
AI Global WireReplan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
arXiv:2609.19654v1 Announce Type: new Abstract: Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local
arXiv cs.AI
arXiv cs.AI
AI Global WireWhen Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e Screening
arXiv:2609.19530v1 Announce Type: new Abstract: Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet r\'esum\'e screening, the first gate, is commonly automated as a static, one-call judgment over a r\'esum\'e-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed r\'esum\'e-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common
arXiv cs.AI
arXiv cs.AI
AI Global WireMAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
arXiv:2609.19391v1 Announce Type: new Abstract: LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verificati
arXiv cs.AI
arXiv cs.AI
AI Global WireWhat Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
arXiv:2609.19212v1 Announce Type: new Abstract: Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as approximately linear action composition, productivity-based tests, and action-explicit goals, which make systematic generalization easier to study but omit some essential aspects of this capability. To characterize what these simplifications miss, we adopt a reasoning-centered lens and introduce TranSGrid, a testbed that brings deductive, inductive, and abductive reasoning together within a unif
arXiv cs.AI
arXiv cs.AI
AI Global WireSafety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
arXiv:2609.19472v1 Announce Type: new Abstract: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, a
arXiv cs.AI
arXiv cs.AI
AI Global WireFINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
arXiv:2609.19680v1 Announce Type: new Abstract: Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should ear
arXiv cs.AI