TACL 2026
How MO-GRPO Balances Multiple Rewards
Why standard GRPO can favor one objective, and how variance-based reward normalization makes the objectives contribute more evenly, with related evidence from GDPO on mathematical reasoning.
Paper Guides
Brief overviews of two of my papers. Each guide focuses on the research question, the main result, and why it matters; the linked paper contains the complete analysis.
TACL 2026
Why standard GRPO can favor one objective, and how variance-based reward normalization makes the objectives contribute more evenly, with related evidence from GDPO on mathematical reasoning.
ACL 2025
A guide to the finite-sample convergence result for Minimum Bayes Risk decoding and its comparison with MAP decoding.