Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization

Accepted at 43rd International Conference on Machine Learning (ICML 2026), 2026

This work develops an Implicit Prefix-Value Reward Model (IPVRM) with value-function-style supervision on every prefix, reducing the training-inference inconsistency of prior implicit process reward models while preserving efficiency.

Links

BibTeX

@inproceedings{gao2026unleashing,
  title = {Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization},
  author = {Gao, Shiping and Chen, Hongzhan and Quan, Xiaojun and Wang, Qifan and Huang, Lifu},
  booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
  year = {2026}
}