Function-Level Execution Feedback for Code Preference Optimization
arXiv cs.AIen
arXiv:2608.23632v1 Announce Type: new Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed as chains of thought. In code generation, however, process supervision remains underexplored because there is no standard notion of a step. Supervision can target lines, reasoning traces, or program states, making it unclear what to label and optimize. We propose STEP-KTODER, a framework for code preference optimization that defines steps as module-level functions in decomposed multi-function programs and assigns binary correctness labels via automatically generated unit tests. Our method provides a code-specific instantiation of stepwise K
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- Peak XV invests $15m in Indian voice AI startup Ringg AITech in Asia · August 26, 2026
- Indian crypto exchange WazirX unveils AI trading assistantTech in Asia · August 26, 2026
- Yhdysvalloissa yltyy kapina datakeskuksia vastaan – Texasissa se voi koitua Trumpin puolueen tappioksiYle Uutiset · August 26, 2026
- A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust CertificationarXiv cs.AI · August 26, 2026
- A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshiftsarXiv cs.AI · August 26, 2026
- AI Agents Push Humans Out of the LooparXiv cs.AI · August 26, 2026