Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational Humor
arXiv cs.AIen
arXiv:2610.00197v1 Announce Type: new Abstract: We investigate automated rewards for training language models in conversational humor, focusing on reward exploits and countermeasures. Two approaches aim to capture understandable surprise and predicted audience amusement. Controlled tests show that an embedding-based surprise reward accepts word-shuffled replies as readily as witty ones. A fluency filter detects the shuffles, but the combined reward also rejects some witty replies and fails further validation. An audience model's predicted laughter is instead vulnerable to laughter cues in either speaker's messages. Normalizing these cues across speakers blocks the covered attacks, although u
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- LG Uplus, OptAI team up on AI token optimizationTech in Asia · October 2, 2026
- Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMsarXiv cs.AI · October 2, 2026
- Benchmarking Prompt Optimization of Large Language Models With ChessarXiv cs.AI · October 2, 2026
- Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?arXiv cs.AI · October 2, 2026
- Before Agents Decide: Epistemic Action in LLM-Based SystemsarXiv cs.AI · October 2, 2026
- Heavy-Tailed Memory Traces in Long-Horizon Language AgentsarXiv cs.AI · October 2, 2026