Ph.D. candidate at Imperial Business School
Imperial College London,
South Kensington Campus, London, SW7 2AZ, United Kingdom
Email: s.liu21@imperial.ac.uk
Machine Learning, Large Language Models, and Operations Research
Within these broad areas, I am interested in theoretically grounded methodological research and in theory that offers useful insights into why methods work, when they work, and where their limitations lie. My interests are not limited to the topics I have worked on so far, and I am always open to discussing new problems and potential collaborations.
Publications
*: equal contribution
Working Papers
The Critic Is Only a Proxy: Regret Optimization for Reinforcement Learning by Yikai Wang*, Shang Liu*, Guanting Chen, Jose Blanchet, working paper, draft available upon request. TL;DR: The critic is only a proxy, and we should be careful about how we treat the signal it sends to the actor.
LLM Alignment and Trustworthiness
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback by Yikai Wang*, Shang Liu*, Jose Blanchet, [arXiv]. TL;DR: When RLHF only observes a proxy reward, minimizing worst-case regret mitigates reward over-optimization without becoming either over-optimistic or over-pessimistic.
Incentivizing High-Quality Human Annotations with Golden Questions byShang Liu, Zhongze Cai, Hanzhao Wang, Zhongyao Ma, Xiaocheng Li, [arXiv]. TL;DR: Golden questions can assess annotation quality and incentivize human annotators to exert greater effort.
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators byShang Liu*, Hanzhao Wang*, Zhongyao Ma, Xiaocheng Li, under minor revision at Management Science, [arXiv]. TL;DR: Fisher information governs the effectiveness of annotator assessment and incentives, while self-consistency monitoring can be cheaper and more informative than expert-based monitoring.
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd byShang Liu*, Yu Pan*, Guanting Chen, Xiaocheng Li, ICML 2025, [arXiv]. TL;DR: Finer-grained ordinal feedback can make reward learning statistically easier, for example by reducing its Rademacher complexity.
Towards Better Statistical Understanding of Watermarking LLMs by Zhongze Cai*, Shang Liu*, Hanzhao Wang*, Huaiyang Zhong, Xiaocheng Li, Journal of the American Statistical Association 2026, [arXiv]. TL;DR: Model distortion is an unavoidable price of watermarking, and the optimal design lies on the Pareto frontier between distortion and detectability.
Uncertainty in Machine Learning
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification byShang Liu*, Zhongze Cai*, Guanting Chen, Xiaocheng Li, Transactions on Machine Learning Research 2025, [arXiv]. TL;DR: Transformers can learn to perform in-context uncertainty quantification, revealing how their mean and uncertainty predictions behave under distribution shifts.
When No-Rejection Learning is Consistent for Regression with Rejection by Xiaocheng Li, Shang Liu, Chunlin Sun, Hanzhao Wang, AISTATS 2024, [arXiv]. TL;DR: Regression with rejection depends on both prediction risk and calibration risk.
Understanding Uncertainty Sampling via Equivalent Loss byShang Liu, Xiaocheng Li, [arXiv]. TL;DR: Uncertainty sampling implicitly optimizes an equivalent loss rather than the original loss, potentially changing convexity, Lipschitzness, and surrogate properties.
Distribution-Free Model-Agnostic Regression Calibration via Nonparametric Methods byShang Liu*, Zhongze Cai*, Xiaocheng Li, NeurIPS 2023, [arXiv]. TL;DR: The smoothness of the conditional quantile function, captured by its Lipschitz constant, reveals the intrinsic difficulty of individual regression calibration.
Linear Programming and Data-Driven Decision-Making
Maximum Optimality Margin: A Unified Approach for Contextual Linear Programming and Inverse Linear Programming by Chunlin Sun*, Shang Liu*, Xiaocheng Li, ICML 2023, [arXiv]. TL;DR: For inverse LPs, the reduced-cost optimality conditions of observed solutions can be turned into an SVM-like maximum-margin learning problem.
Non-stationary Bandits with Knapsacks byShang Liu, Jiashuo Jiang, Xiaocheng Li, NeurIPS 2022, [arXiv]. TL;DR: Global resource constraints require global, rather than only local, measures of non-stationarity.
Online Bin Packing with Known T byShang Liu, Xiaocheng Li, under major revision at Mathematics of Operations Research, [arXiv]. TL;DR: In a doubling-trick algorithm, the number of overflowed items can be represented by the maximum of a zero-endpoint random walk.