HARDEN: Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.30571v1 Announce Type: new Abstract: Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Xiaomi-backed robotics chip designer clears hearing, eyes US$100m Hong Kong IPO: sourcesSCMP Tech · September 28, 2026
- Bringing AI to Autonomous Systems -- From Cognition to Collective IntelligencearXiv cs.AI · September 28, 2026
- Bridging LLM Agents and Data Spaces: An Architectural Mediation Approach using the Model Context ProtocolarXiv cs.AI · September 28, 2026
- Predicting Transmembrane Protein Topology from 3D StructurearXiv cs.AI · September 28, 2026
- Spectral Feedback for Test-Time Alignment of Protein Diffusion ModelsarXiv cs.AI · September 28, 2026
- Pretrained ASR Pseudo-labeling for Noisy Police AudioarXiv cs.AI · September 28, 2026