Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
arXiv cs.AIen
arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weak
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Q&A with AI researcher Jacob Coxon, who quit Anthropic, on the need for industry-wide, international coordination to limit recursive self-improvement, and more (Maxwell Zeff/Wired)Techmeme · September 10, 2026
- Anzeige: Ansible f�r automatisiertes SystemmanagementGolem.de · September 10, 2026
- Wistron, Wiwynn hit record August revenue as board approves US$200 Million for US, Vietnam expansionDIGITIMES · September 10, 2026
- Samsung SDS broadens AI push from software to factory robotsDIGITIMES · September 10, 2026
- Donnerstag: Apples neue iPhones auch aufklappbar, KI-Agenten weiter ungezügeltheise online – KI · September 10, 2026
- Generative AI a new tool in Mali's information war: studyEconomic Times Tech · September 10, 2026