CV

Shiping Gao

PhD Student in Computer Science and Engineering

shiping@umich.edu
Ann Arbor, MI, US

Summary

PhD student in Computer Science and Engineering at the University of Michigan. Research interests include NLP, LLMs, RLHF, reward hacking, fine-grained reward modeling, and LLM reasoning.

Education

  • PhD in Computer Science and Engineering
    Sep 2026 - Jun 2030 (Expected)
    University of Michigan
    Courses: Advisor: Silviu Pitis
  • MPhil in Computer Science and Technology
    Sep 2023 - Jun 2026
    Sun Yat-Sen University
    GPA: 89.84/100
  • BEng in Computer Science and Technology
    Sep 2019 - Jul 2023
    Lanzhou University
    GPA: 90.17/100
    Courses: Ranking: 8/232 (top 3.5%)

Work Experience

  • Research Intern
    Jul 2025 - Jan 2026
    University of California, Davis
    Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization. Supervised by Prof. Lifu Huang.
    • Developed IPVRM with value-function-style supervision on every prefix.
    • Introduced distribution-level temporal-difference advantages for dense credit assignment.
  • Research Team Member
    Oct 2024 - Jan 2025
    Sun Yat-Sen University
    Discriminative Policy Optimization for Token-Level Reward Models. Supervised by Prof. Xiaojun Quan.
    • Designed Q-RM, a token-level reward model without fine-grained annotations.
    • Integrated Q-RM into PPO and REINFORCE for reasoning tasks.
  • Principal Investigator
    Jul 2024 - Oct 2024
    Sun Yat-Sen University
    Advantage-Guided Distillation for Preference Alignment in Small Language Models. Supervised by Prof. Xiaojun Quan.
    • Designed DCKD, ADPA, and ADPA+ for small-model preference alignment.
    • Improved alignment performance while reducing sample complexity.
  • Research Team Member
    Jul 2024 - Sep 2024
    Sun Yat-Sen University
    Edit-Wise Preference Optimization for Grammatical Error Correction. Supervised by Prof. Xiaojun Quan.
    • Proposed EPO with edit-aware token weighting.
    • Built annotation-free preference data construction for GEC.
  • Undergraduate Thesis Researcher
    Dec 2022 - May 2023
    Lanzhou University
    Research on Chinese Text Correction System Based on Deep Learning. Supervised by Prof. Binbin Yong.
    • Designed a Chinese text correction system for spelling, grammar, and semantic errors.
    • Proposed AutoCscError for realistic pseudo-data generation.
  • Research Team Member
    May 2021 - Oct 2021
    Lanzhou University
    A Novel Dynamic Interpolation Method Based on Temporal and Spatial Correlations. Supervised by Prof. Zhili Zhao.
    • Proposed dynamic spatiotemporal interpolation for environmental and meteorological monitoring.

Skills

Languages

  • Python
  • C++
  • Java
  • Bash
  • LaTeX
  • SQL

Frameworks and Libraries

  • Transformers
  • Ray
  • VeRL
  • LLaMA-Factory
  • Datasets
  • Accelerate
  • DeepSpeed
  • vLLM

Tools and Platforms

  • Git
  • Docker
  • Weights & Biases
  • SwanLab
  • TensorBoard
  • Slurm
  • CUDA

Publications

  • Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimization
    2026
    ICML 2026
    Shiping Gao, Hongzhan Chen, Xiaojun Quan, Qifan Wang, Lifu Huang.
  • Adaptive Boundaries: Context-Aware Detection for Synthetic Text
    2026
    EMNLP 2026
    H. Li, R. Ni, Haihui Yang, Shiping Gao, Y. Liu, Xiaojun Quan.
  • Stabilizing Policy Optimization via Logits Convexity
    2026
    arXiv preprint arXiv:2603.00963 (under review at NeurIPS 2026)
    Hongzhan Chen, Tao Yang, Yuhua Zhu, Shiping Gao, Xiaojun Quan, Ting Yao.
  • Discriminative Policy Optimization for Token-Level Reward Models
    2025
    ICML 2025
    Hongzhan Chen, Tao Yang, Shiping Gao, Ruijun Chen, Xiaojun Quan, Hongtao Tian, Ting Yao.
  • Advantage-Guided Distillation for Preference Alignment in Small Language Models
    2025
    ICLR 2025 Spotlight
    Shiping Gao, Fanqi Wan, Jiajian Guo, Xiaojun Quan, Qifan Wang.
  • Edit-Wise Preference Optimization for Grammatical Error Correction
    2025
    COLING 2025
    Jiehao Liang, Haihui Yang, Shiping Gao, Xiaojun Quan.
  • Self-Evolution Fine-Tuning for Policy Optimization
    2024
    Findings of EMNLP 2024
    Ruijun Chen, Jiehao Liang, Shiping Gao, Fanqi Wan, Xiaojun Quan.
  • A Novel Dynamic Interpolation Method Based on Both Temporal and Spatial Correlations
    2022
    Applied Intelligence
    Shiping Gao, Dongjie He, Zhouzhuo Zhang, Xiaoqian Tang, Zhili Zhao.

Teaching

  • Artificial Neural Networks and Mathematical Principles of Reinforcement Learning
    2025
    Sun Yat-Sen University
    Role: Teaching Assistant
    Assisted undergraduates in Artificial Neural Networks and instructed Mathematical Principles of RL for graduate students.

Interests

  • Research Interests
    NLP, LLMs, RLHF, Reward Hacking, Fine-Grained Reward Modeling, LLM Reasoning