Stabilizing Policy Optimization via Logits Convexity

Preprint, arXiv preprint arXiv:2603.00963, 2026

This work studies how the convexity of policy logits can be leveraged to stabilize policy optimization for large language models.

Links

BibTeX

@article{chen2026stabilizing,
  title = {Stabilizing Policy Optimization via Logits Convexity},
  author = {Chen, Hongzhan and Yang, Tao and Zhu, Yuhua and Gao, Shiping and Quan, Xiaojun and Yao, Ting},
  journal = {arXiv preprint arXiv:2603.00963},
  year = {2026}
}