Discriminative Policy Optimization for Token-Level Reward Models
Accepted at 42nd International Conference on Machine Learning (ICML 2025), 2025
This work designs Q-RM, a token-level reward model that learns token-level rewards without fine-grained annotations by decoupling reward modeling from language modeling through a discriminative policy.
Links
BibTeX
@inproceedings{chen2025discriminative,
title = {Discriminative Policy Optimization for Token-Level Reward Models},
author = {Chen, Hongzhan and Yang, Tao and Gao, Shiping and Chen, Ruijun and Quan, Xiaojun and Tian, Hongtao and Yao, Ting},
booktitle = {Proceedings of the 42nd International Conference on Machine Learning},
year = {2025}
}
