From Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought
arXiv cs.AIen
arXiv:2609.25366v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is only meaningful if written reasoning causally constrains the answer. We introduce continuation-based causal testing, an ablation-patch intervention that perturbs one reasoning step, truncates the chain, and forces the model to continue from the corrupted prefix. It measures how load-bearing a CoT is for the final answer, a behavioral notion distinct from mechanistic faithfulness. Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; o
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Meta
- DeepSeek
- Forskning
Related AI news
- They built AI agents on WhatsApp. Then Meta entered the chatTech in Asia · September 23, 2026
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian SplattingarXiv cs.AI · September 23, 2026
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026
- Making Agents More Consistent: Skills Should Form Habits for Repeat TasksarXiv cs.AI · September 23, 2026
- Real-Time Hand Gesture Recognition for OpenXR Using Transformer-Based Machine LearningarXiv cs.AI · September 23, 2026
- X-Planner: Event-Structured Task Planning for Embodied IntelligencearXiv cs.AI · September 23, 2026