Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization
Accepted at 43rd International Conference on Machine Learning (ICML 2026), 2026
This work develops an Implicit Prefix-Value Reward Model (IPVRM) with value-function-style supervision on every prefix, reducing the training-inference inconsistency of prior implicit process reward models while preserving efficiency.
Links
BibTeX
@inproceedings{gao2026unleashing,
title = {Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization},
author = {Gao, Shiping and Chen, Hongzhan and Quan, Xiaojun and Wang, Qifan and Huang, Lifu},
booktitle = {Proceedings of the 43rd International Conference on Machine Learning},
year = {2026}
}
