Discriminative Policy Optimization for Token-Level Reward Models

Accepted at 42nd International Conference on Machine Learning (ICML 2025), 2025

This work designs Q-RM, a token-level reward model that learns token-level rewards without fine-grained annotations by decoupling reward modeling from language modeling through a discriminative policy.

Links

BibTeX

@inproceedings{chen2025discriminative,
  title = {Discriminative Policy Optimization for Token-Level Reward Models},
  author = {Chen, Hongzhan and Yang, Tao and Gao, Shiping and Chen, Ruijun and Quan, Xiaojun and Tian, Hongtao and Yao, Ting},
  booktitle = {Proceedings of the 42nd International Conference on Machine Learning},
  year = {2025}
}