Advantage-Guided Distillation for Preference Alignment in Small Language Models
Accepted at 13th International Conference on Learning Representations (ICLR 2025), 2025
This work introduces ADPA and ADPA+, using preference-aware distribution-level signals and dual-constrained distillation to improve alignment in small language models.
Links
BibTeX
@inproceedings{gao2025advantage,
title = {Advantage-Guided Distillation for Preference Alignment in Small Language Models},
author = {Gao, Shiping and Wan, Fanqi and Guo, Jiajian and Quan, Xiaojun and Wang, Qifan},
booktitle = {International Conference on Learning Representations},
year = {2025}
}
