CV
Shiping Gao
PhD Student in Computer Science and Engineering
Summary
PhD student in Computer Science and Engineering at the University of Michigan. Research interests include NLP, LLMs, RLHF, reward hacking, fine-grained reward modeling, and LLM reasoning.
Education
- PhD in Computer Science and EngineeringSep 2026 - Jun 2030 (Expected)University of MichiganCourses: Advisor: Silviu Pitis
- MPhil in Computer Science and TechnologySep 2023 - Jun 2026Sun Yat-Sen UniversityGPA: 89.84/100
- BEng in Computer Science and TechnologySep 2019 - Jul 2023Lanzhou UniversityGPA: 90.17/100Courses: Ranking: 8/232 (top 3.5%)
Work Experience
- Research InternJul 2025 - Jan 2026University of California, DavisUnleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization. Supervised by Prof. Lifu Huang.
- Developed IPVRM with value-function-style supervision on every prefix.
- Introduced distribution-level temporal-difference advantages for dense credit assignment.
- Research Team MemberOct 2024 - Jan 2025Sun Yat-Sen UniversityDiscriminative Policy Optimization for Token-Level Reward Models. Supervised by Prof. Xiaojun Quan.
- Designed Q-RM, a token-level reward model without fine-grained annotations.
- Integrated Q-RM into PPO and REINFORCE for reasoning tasks.
- Principal InvestigatorJul 2024 - Oct 2024Sun Yat-Sen UniversityAdvantage-Guided Distillation for Preference Alignment in Small Language Models. Supervised by Prof. Xiaojun Quan.
- Designed DCKD, ADPA, and ADPA+ for small-model preference alignment.
- Improved alignment performance while reducing sample complexity.
- Research Team MemberJul 2024 - Sep 2024Sun Yat-Sen UniversityEdit-Wise Preference Optimization for Grammatical Error Correction. Supervised by Prof. Xiaojun Quan.
- Proposed EPO with edit-aware token weighting.
- Built annotation-free preference data construction for GEC.
- Undergraduate Thesis ResearcherDec 2022 - May 2023Lanzhou UniversityResearch on Chinese Text Correction System Based on Deep Learning. Supervised by Prof. Binbin Yong.
- Designed a Chinese text correction system for spelling, grammar, and semantic errors.
- Proposed AutoCscError for realistic pseudo-data generation.
- Research Team MemberMay 2021 - Oct 2021Lanzhou UniversityA Novel Dynamic Interpolation Method Based on Temporal and Spatial Correlations. Supervised by Prof. Zhili Zhao.
- Proposed dynamic spatiotemporal interpolation for environmental and meteorological monitoring.
Skills
Languages
- Python
- C++
- Java
- Bash
- LaTeX
- SQL
Frameworks and Libraries
- Transformers
- Ray
- VeRL
- LLaMA-Factory
- Datasets
- Accelerate
- DeepSpeed
- vLLM
Tools and Platforms
- Git
- Docker
- Weights & Biases
- SwanLab
- TensorBoard
- Slurm
- CUDA
Publications
- Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization2026
- Adaptive Boundaries: Context-Aware Detection for Synthetic Text2026
- Stabilizing Policy Optimization via Logits Convexity2026arXiv preprint arXiv:2603.00963 (under review at NeurIPS 2026)Hongzhan Chen, Tao Yang, Yuhua Zhu, Shiping Gao, Xiaojun Quan, Ting Yao.
- Discriminative Policy Optimization for Token-Level Reward Models2025ICML 2025Hongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen, Xiaojun Quan, Hongtao Tian, Ting Yao.
- Advantage-Guided Distillation for Preference Alignment in Small Language Models2025
- Edit-Wise Preference Optimization for Grammatical Error Correction2025
- Self-Evolution Fine-Tuning for Policy Optimization2024Findings of EMNLP 2024Ruijun Chen, Jiehao Liang, Shiping Gao, Fanqi Wan, Xiaojun Quan.
- A Novel Dynamic Interpolation Method Based on Both Temporal and Spatial Correlations2022Applied IntelligenceShiping Gao, Dongjie He, Zhouzhuo Zhang, Xiaoqian Tang, Zhili Zhao.
Teaching
- Artificial Neural Networks and Mathematical Principles of Reinforcement Learning2025Sun Yat-Sen UniversityRole: Teaching AssistantAssisted undergraduates in Artificial Neural Networks and instructed Mathematical Principles of RL for graduate students.
Interests
- Research InterestsNLP, LLMs, RLHF, Reward Hacking, Fine-Grained Reward Modeling, LLM Reasoning
