Stabilizing Policy Optimization via Logits Convexity
Preprint, arXiv preprint arXiv:2603.00963, 2026
This work studies how the convexity of policy logits can be leveraged to stabilize policy optimization for large language models.
Links
BibTeX
@article{chen2026stabilizing,
title = {Stabilizing Policy Optimization via Logits Convexity},
author = {Chen, Hongzhan and Yang, Tao and Zhu, Yuhua and Gao, Shiping and Quan, Xiaojun and Yao, Ting},
journal = {arXiv preprint arXiv:2603.00963},
year = {2026}
}
