「1週間の開発タスク」でAIの限界を検証 Googleが「Android Bench 2.0」公開
ITmedia AI+ja

Googleは、AIモデルやエージェントのコーディング能力を測る「Android Bench 2.0」を公開した。エンジニアが数日を要する複雑な長期タスクを課す仕様に刷新。最高通過率は従来の約91%から約28%に急落した。
This is a short summary published by AI Global Wire. The full article is owned and hosted by ITmedia AI+ — open it there to read it in full.
Read the full story at ITmedia AI+Related AI news
- 蘋果 Home 攝影機 AI 功能實測!描述精準度與價格皆落後 Google Ring 與亞馬遜 NestTechNews (TW) · September 26, 2026
- Another Google Deepmind researcher quits, says building superintelligent AI soon is "inherently irresponsible"The Decoder · September 25, 2026
- One company is at the center of a wave of rogue AI attacksThe Verge AI · September 25, 2026
- Google skjuter upp litet AI-datacenter i rymdenComputer Sweden · September 25, 2026
- Can Apple Home’s AI camera features outsmart Amazon and Google’s? I put them to the testThe Verge AI · September 25, 2026
- Can Apple Home’s AI camera features outsmart Amazon’s and Google’s? I put them to the testThe Verge AI · September 25, 2026