Yuki Ichihara

Ph.D. student at MBZUAI

I am a Ph.D. student at MBZUAI, supervised by Junpei Komiyama.

I study how to make language-model generation more reliable and efficient, from decoding algorithms and statistical guarantees to reinforcement-learning methods for multi-objective alignment.

Profile photo

Research

My work connects decoding, statistical reliability, and reinforcement learning to build language models that optimize what we actually care about.

01

Reliable Decoding

I study Best-of-N and Minimum Bayes Risk decoding, including when these methods work, how quickly they converge, and how to make them robust to imperfect rewards.

02

Multi-Objective Alignment

I develop reinforcement-learning methods that balance multiple rewards and constraints while reducing reward hacking and manual weight tuning.

03

Efficient and Trustworthy Reasoning

I explore statistical signals and training objectives that make self-consistency and reasoning more reliable without unnecessary inference-time computation.

Publications

  1. Prefix-Denoising Consistency: Test-Time Verification for Diffusion Language Models Y. Ichihara, N. Iwase, M. A. Quamar, J. Komiyama. arXiv 2026 · Preprint Paper
    Summary & figure

    Prefix-Denoising Consistency (PDC) verifies diffusion-language-model outputs by keeping different prefixes, remasking the remaining suffix, and regenerating alternative reasoning trajectories. Across mathematical and commonsense reasoning benchmarks, voting over these regenerations improves the initial answer while using inference compute more effectively than independent full generations.

    Prefix-Denoising Consistency pipeline with three prefix keep rates, suffix regeneration, and majority voting
    Figure 1. PDC regenerates masked suffixes from several fixed prefixes and votes over the resulting answers.
  2. Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation Y. Ichihara, N. Iwase, M. A. Quamar, J. Komiyama. arXiv 2026 · Preprint Paper Code
    Summary & figure

    OP2SD tests whether on-policy self-distillation truly requires the teacher to see the target problem's solution. Replacing it with a worked solution from another problem still improves the base model and remains competitive with OPSD, showing that context-induced teacher behavior is an important part of the distilled signal.

    Comparison of OPSD, whose teacher sees the target solution, and OP squared SD, whose teacher sees a worked solution to a different problem
    Figure 1. OPSD and OP2SD differ only in whether the teacher sees the target solution or another problem's solution.
  3. Reliable Chain-of-Thought via Prefix Consistency N. Iwase, Y. Ichihara, M. A. Quamar, J. Komiyama. arXiv 2026 · Preprint Paper Code
    Summary & figure

    Prefix consistency estimates the reliability of a chain-of-thought by truncating it and checking whether regenerated continuations recover the same answer. Across reasoning models and benchmarks, it predicts correctness well and reaches standard majority-voting accuracy with substantially fewer tokens.

    Prefix consistency accuracy curves and examples of consistent and inconsistent reasoning traces
    Figure 1. Prefix-consistency-weighted majority voting and the regeneration-based reliability signal.
  4. CITE: Anytime-Valid Statistical Inference in LLM Self-Consistency H. Ota, N. Iwase, Y. Ichihara, J. Komiyama, M. Imaizumi. arXiv 2026 · Preprint Paper
    Summary & figure

    CITE provides anytime-valid statistical certification that a target response is the unique mode of an LLM's answer distribution. It controls false certification under adaptive stopping without assuming a fixed or known set of possible answers.

    Diagram of the CITE stopping rule comparing a target response with observed and unseen categories
    Figure 1. CITE's anytime-valid stopping rule for certifying a target response.
  5. Consensus Group Relative Policy Optimization for Text Generation Y. Ichihara, Y. Jinnai, K. Ariu, E. Uchibe. EMNLP 2026 · Main Paper Code
    Summary & figure

    C-GRPO distills Minimum Bayes Risk decoding into GRPO training using only a utility function and policy-generated samples. It achieves performance comparable to MBR decoding in translation and summarization without its inference-time overhead.

    Comparison of standard GRPO and Consensus-GRPO training pipelines
    Figure 1. Standard GRPO and the label-free consensus utility used by C-GRPO.
  6. MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems Y. Ichihara, Y. Jinnai, T. Morimura, M. Sakamoto, R. Mitsuhashi, E. Uchibe. TACL 2026 · Journal Paper Code Guide
    Summary & figure

    MO-GRPO addresses reward hacking in multi-objective GRPO by automatically normalizing rewards according to their variance. It balances objectives without manual scale tuning and delivers more stable learning across control and text-generation tasks.

    Comparison of GRPO and MO-GRPO advantage functions for two rewards with different variances
    Figure 1. GRPO and MO-GRPO advantage functions when reward scales differ.
  7. Auto-Weighted Group Relative Preference Optimization for Multi-Objective Text Generation Tasks Y. Ichihara, Y. Jinnai. EMNLP 2025 · Industry Paper Code
    Summary & figure

    AW-GRPO automatically adjusts multiple reward weights according to each objective's learning progress. On advertising text generation and public benchmarks, it improves the balance among objectives while reducing constraint violations.

    Training curves showing GRPO improving readability while reducing BLEURT score
    Figure 1. Standard GRPO overfits one reward on WMT En-Ja, motivating automatic weighting.
  8. Theoretical Guarantees for Minimum Bayes Risk Decoding Y. Ichihara, Y. Jinnai, K. Ariu, T. Morimura, E. Uchibe. ACL 2025 · Main Paper Guide
    Summary & figure

    This work provides finite-sample guarantees showing that MBR decoding approaches the optimal solution at a rate of O(n−1/2) under stated assumptions. It also explains why MBR can converge faster than MAP decoding in several cases.

    Conceptual plot comparing the regret bounds of MAP and MBR decoding
    Figure 1. Regimes in which the regret bound of MBR or MAP decoding is smaller.
  9. Evaluation of Best-of-N Sampling Strategies for Language Model Alignment Y. Ichihara, Y. Jinnai, T. Morimura, K. Abe, K. Ariu, M. Sakamoto, E. Uchibe. TMLR 2025 · Journal Paper
    Summary & figure

    This work studies regularized Best-of-N sampling as robust optimization against errors in proxy rewards. It introduces Stochastic RBoN with theoretical guarantees and shows how simple regularizers can reduce reward hacking on alignment benchmarks.

    Win-rate curves for regularized Best-of-N methods across five proxy reward models
    Figure 1. Sensitivity of regularized Best-of-N methods to beta across proxy reward models.
  10. A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees T. Kitamura, T. Kozuno, M. Kato, Y. Ichihara, S. Nishimori, A. Sannai, S. Sonoda, W. Kumagai, Y. Matsuo. RLSW 2024 · Workshop Paper
    Summary & figure

    This work proposes a policy-gradient primal-dual algorithm for online constrained MDPs with Uniform-PAC guarantees. It combines convergence to an optimal policy, sublinear regret, and polynomial sample complexity while controlling constraint violations.

    Optimality gap and constraint violation curves for primal-dual constrained MDP algorithms
    Figure 1. Optimality gap and constraint violation of primal-dual CMDP algorithms.

Recent Updates

Selected publication, service, and career updates.

C-GRPO was accepted to the EMNLP 2026 Main Conference.

We released PDC for diffusion-language-model verification and OP2SD on context-induced teacher behavior in self-distillation.

I joined MBZUAI as a Ph.D. student, supervised by Junpei Komiyama.

MO-GRPO was accepted to TACL.

We released work on prefix consistency and anytime-valid inference with CITE.

We released C-GRPO, which distills consensus-based decoding into training.

AW-GRPO appeared at the EMNLP Industry Track.

Our work on theoretical guarantees for MBR decoding appeared at ACL.

Our evaluation of Best-of-N sampling strategies was published in TMLR.

Background & Service

Education

MBZUAI, Ph.D. student, supervised by Junpei Komiyama

Nara Institute of Science and Technology, Ph.D. program in Engineering (transferred to MBZUAI)

Nara Institute of Science and Technology, M.S. in Engineering

Tokyo University of Science, B.S. in Electrical Engineering

Research Experience

JST SPRING Fellowship, Doctoral fellowship

CyberAgent AI Lab, Research Intern

ATR, Research Assistant

AIST, Research Assistant

Academic Service

Reviewer, ARR; ICML (nominated as Gold Reviewer)

Contact

For research discussions or collaboration, email me at y.ichihara3406@gmail.com.