ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.16639v1 Announce Type: new Abstract: Continual post-training of large multimodal models should add new capabilities while preserving those from pre-training, and the two goals pull in opposite directions. SFT gives explicit target supervision that learns a task from near-zero accuracy, but its off-policy targets move the model far enough to cause forgetting; on-policy methods such as RLVR and self-distillation preserve policy proximity yet supply little signal when the policy cannot yet solve the task. We introduce ReDraft (Reference-Driven Revision and Fine-Tuning), which obtains both from the model's own failures: using an expert response only as a reference, it has the model re
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Reglering
Related AI news
- Pentagon CTO says U.S. government shouldn't take stakes in AI giants, questions adding AI rulesCNBC Technology · September 17, 2026
- Chip equipment and materials suppliers lead India investment pledges ahead of SEMICON India 2026DIGITIMES · September 17, 2026
- Elon Musk urges rival testing to expose AI safety flawsDIGITIMES · September 17, 2026
- House votes to curb AI data center costsAxios · September 16, 2026
- OpenAI CEO Sam Altman will attend state dinner for Trump-Xi summit in WashingtonCNBC Technology · September 16, 2026
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?TechCrunch AI · September 16, 2026