Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs
arXiv cs.AIen
arXiv:2610.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Google's first Suncatcher satellite reaches orbit to test AI chips in spaceDIGITIMES · October 2, 2026
- Benchmarking Prompt Optimization of Large Language Models With ChessarXiv cs.AI · October 2, 2026
- Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph QueryingarXiv cs.AI · October 2, 2026
- Before Agents Decide: Epistemic Action in LLM-Based SystemsarXiv cs.AI · October 2, 2026
- Ontology-Based Contextual AI Evaluations (OB-CAIE) MethodologyarXiv cs.AI · October 2, 2026
- From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent ExecutionarXiv cs.AI · October 2, 2026