Dr. Cheems Wang

Research Fellow, Tsinghua University
Adjunct Associate Professor in Mathematics

He works on meta-learning, multi-task learning and reinforcement learning for large models, aiming for fast adaptation with uncertainty you can trust.

Portrait of Dr. Cheems Wang
Few-shot prediction with uncertainty: click or tap the plot to add an observation.

About

Dr. Cheems Wang is a Research Fellow at Tsinghua University, where he works closely with Prof. Xiangyang Ji, and an adjunct associate professor in mathematics. He obtained his Ph.D. in machine learning from the University of Amsterdam, where he worked in the Amsterdam Machine Learning Lab (AMLab) under the supervision of Prof. Max Welling and Prof. Herke van Hoof.

His research studies the learning paradigms that let large models adapt: meta-learning, multi-task learning and reinforcement learning. The thread through his work is fast adaptation with reliable uncertainty, from neural processes and Bayesian meta-learning during his Ph.D. to task sampling and efficient RL post-training of large reasoning models today.

His papers include 8 at ICML, 12 at NeurIPS, 4 at ICLR, 4 at KDD, 2 at AAAI, 2 at CVPR and 1 at ICCV, plus articles in Nature Communications, IEEE T-PAMI and Science China Information Sciences; 8 were selected as orals or spotlights. The result he is proudest of since leaving AMLab, the birthplace of the VAE, is Model Predictive Task Sampling, in which a small generative model steers the optimization of large ones. It is published in Nature Communications and the code is on GitHub.

He and Prof. Ji are close to releasing a large open-source project, with news expected early in 2027. He welcomes research collaborations; email him to get in touch.

Background

  1. 2023–
    Research Fellow, Department of Automation, Tsinghua University, working with Prof. Xiangyang Ji.
  2. 2019–2022
    Ph.D., Amsterdam Machine Learning Lab, University of Amsterdam, with Prof. Max Welling and Prof. Herke van Hoof. Thesis: Functional Representation Learning for Uncertainty Quantification and Fast Skill Transfer (PDF, defense video). Committee: Maarten de Rijke, Aske Plaat, Efstratios Gavves, Sara Magliacane and Xiantong Zhen.
  3. 2018–2019
    Computational Science Lab, University of Amsterdam, hosted by Prof. Peter Sloot, who supported him through his first and hardest months in the Netherlands.
  4. 2015–2017
    M.Sc. in Management Science.
  5. 2011–2015
    B.Sc. in Mathematics, Sichuan University.

Honors

  1. 2023
    CCF-CMAS Best Dissertation Prize (中国计算机学会多智能体系统学组优博论文奖)
  2. 2022
    NeurIPS Scholar Award
  3. 2022
    ICML Participation Grant

Service and teaching

  1. Area Chair
    ICLR 2027, ICML 2026, ICLR 2026
  2. Senior PC
    AAMAS 2026
  3. Workshop PC
    NeurIPS 2021 Workshop on Ecological Theory of RL (EcoRL 2021)
  4. Seminar
    Helped organize the AMLab weekly seminar, 2020–2021
  5. Teaching
    Teaching assistant for Reinforcement Learning in the MSc in Artificial Intelligence, University of Amsterdam, 2019–2021, running the practical and Q&A sessions. His slides, adapted from Sutton and Barto’s book and other open materials, are free to use:

News

  1. “You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement” was made public. After three months of hard work, we built a System-1 model, referred to as Jev-LDE, to help with demonstration selection for better In-Context Learning of LLMs. Feel free to access our Jev model !

  2. “Bridging Risk Approximation Gaps in Model Predictive Task Sampling via In-Context Modeling” was accepted to NeurIPS 2026. Congratulations to Jiarong!

  3. “Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex” was accepted to NeurIPS 2026. Congratulations to Yun Qu!

  4. “TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning” was accepted to EMNLP 2026. It is our latest prompt-predictive method for faster agentic RL training. Congratulations to Heming!

  5. Invited to serve as an Area Chair for ICLR 2027.

  6. New preprint: “Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex”, an advance in group-based policy gradients for RL post-training of large models, with faster training and stronger reasoning.

  7. “Model Predictive Task Sampling for Efficient and Robust Adaptation” was accepted to Nature Communications. Congratulations to all collaborators!

  8. “Latent Space Robust Optimization of Neural Processes with Aligned Stratified Order-Statistic Loss Reduction” was accepted to ICML 2026, led by my master’s students Qi Tao and Jiarong Wen.

  9. “Stochastic Gradient Methods under Heavy-Tailed Noises in Weakly Convex Optimization” was accepted to ICML 2026. Thanks to Prof. Yi Xu for leading this work.

  10. “Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning” was accepted to ICML 2026. Congratulations to Heming!

  11. “Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models” was accepted to ICML 2026. Congratulations to Yun Qu!

  12. “Stop Wandering, Find the Keys: LLMs Discriminate Key States for Efficient Multi-Agent Exploration” was accepted to Science China Information Sciences. Congratulations to Yun Qu!

  13. New preprint: “Limited Reasoning Space: The cage of long-horizon reasoning in LLMs”, our latest state of the art for efficient test-time compute.

  14. “Neural Mixture Density Processes” was accepted to CVPR 2026. Congratulations to my master’s students Yi Ding and Qi Tao!

  15. New preprint: “VAO: Validation-Aligned Optimization for Cross-Task Generative Auto-Bidding”, which tackles task imbalance with a data-reuse strategy and gives an effective multi-task generative auto-bidding algorithm.

  16. New preprint: “Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models”, a versatile recipe for efficient RL post-training of large reasoning models (thread on X).

  17. “Tail Task Risk Minimization in Meta-Learning from Theoretical Advances to Practical Strategies”, with me as corresponding author, was accepted to IEEE T-PAMI. Congratulations to Yiqin and all collaborators!

  18. “Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search”, with Alibaba Group, was accepted to ICLR 2026 as an oral (top 1% of submissions). Congratulations to Zhiyu and Yiqin!

  19. “Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models” was accepted to ICLR 2026. Congratulations to Yixiu!

  20. “Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?” was accepted to KDD 2026. It extends model predictive task sampling to discrete task spaces and sharply cuts the cost of RL fine-tuning for large reasoning models. The method has been cited or used as a baseline by Meta Superintelligence Labs, the Qwen team, Tencent Hunyuan, Ant Group and others.

  21. Invited to serve as an Area Chair for ICML 2026.

  22. Three papers were accepted to NeurIPS 2025, one as a spotlight. Congratulations to all collaborators!

  23. Invited to serve as an Area Chair for ICLR 2026.

  24. Invited to serve as an Area Chair (Senior PC) for AAMAS 2026.

  25. “Fast and Robust: Task Sampling with Posterior and Diversity Synergies for Adaptive Decision-Makers in Randomized Environments” was accepted to ICML 2025. It makes meta-RL and domain randomization more robust without extra agent–environment interaction (project page). Congratulations to Yun Qu!

  26. Nominated as an Area Chair for NeurIPS 2025, though in the end I could not take it on.

  27. New preprint: “Model Predictive Task Sampling for Efficient and Robust Adaptation”, the first attempt to predict the optimization outcome of any-shot adaptation from the tasks alone. It makes multimodal foundation models and sequential decision-makers adapt more robustly while keeping learning efficient.

  28. “DynaPrompt: Dynamic Test-Time Prompt Tuning” was accepted to ICLR 2025. Congratulations to Zehao; I was honored to collaborate on this project.

  29. “Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning” was accepted to AAAI 2025. Congratulations to Yun Qu!

  30. “Robust Fast Adaptation from Adversarially Explicit Task Distribution Generation” was accepted to the KDD 2025 Research Track. Congratulations to all collaborators!

  31. Four papers were accepted to NeurIPS 2024, all with students I co-supervise, including “Doubly Mild Generalization for Offline Reinforcement Learning”, “Offline Reinforcement Learning with OOD State Correction and OOD Action Suppression” and “GO4Align: Group Optimization for Multi-Task Alignment” (code).

  32. New preprint: “Robust Fast Adaptation from Adversarially Explicit Task Distribution Generation”. It handles task-distribution shift between meta-training and meta-testing by generating explicit task distributions adversarially (blog).

  33. “Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation” was accepted to ICML 2024. A surrogate backpropagation algorithm and a variant of Layer Normalization substantially cut fine-tuning memory, with theoretical guarantees.

  34. New preprint: “GO4Align: Group Optimization for Multi-Task Alignment”, which aligns learning progress across tasks to exploit task relatedness and reaches state-of-the-art results on most multi-task benchmarks.

  35. Joined the Department of Automation at Tsinghua University as a researcher, working with Prof. Xiangyang Ji.

  36. “Episodic Multi-Task Learning with Heterogeneous Neural Processes” was accepted to NeurIPS 2023 as a spotlight, top 3.06% of submissions (code). Congratulations to Jiayi Shen!

  37. Received the CCF-CMAS Best Dissertation Prize. Many thanks to Max and Herke for the nomination, and to Sara, Yang and Sihang for their recommendations.

  38. “Bridge the Inference Gaps of Neural Processes via Expectation Maximization” was accepted to ICLR 2023.

  39. Received a NeurIPS 2022 Scholar Award. Thanks to the NeurIPS Foundation.

  40. “Learning Expressive Meta-Representations with Mixture of Expert Neural Processes” was accepted to NeurIPS 2022. It handles stochastic processes with mixture components, in both few-shot supervised learning and meta-RL.

  41. Received an ICML 2022 Participation Grant.

  42. “Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models and Amortized Policy Search” was accepted to ICML 2022 as a spotlight. A GNN dynamics model generalizes across environments, and amortized policy search adapts to new ones without extra policy gradients (slides).

  43. “Doubly Stochastic Variational Inference for Neural Processes with Hierarchical Latent Variables” was accepted to ICML 2020. A hierarchical neural process identifies tasks and captures local correlations in high-dimensional problems (arXiv, slides).

Publications

Every paper on his Google Scholar profile. “Selected” shows the papers he highlights, which are shaded under “All”. * marks equal contribution or corresponding authorship, as on Scholar.

2026

  1. arXiv preprint
    You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement

    J. Wen*, Cheems Wang*, Y. Qu, Y. Mao, H. Zou, H. Chi, L. Cai, Y. Lv, K. Zhang, Y. Jiang, X. Ji

  2. NeurIPS
    Bridging Risk Approximation Gaps in Model Predictive Task Sampling via In-Context Modeling

    J. Wen, Q. Tao, K. Zhang, Y. Qu, L. Cai, H. Zou, Y. Jiang, Cheems Wang*

  3. EMNLPSelected
    TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

    H. Zou, Cheems Wang*, Y. Qu, Y. Jiang, L. Cai, Y. Mao, R. Peng, X. Xu, W. Liu, K. Yang, S. Yang, X. Ji*

  4. arXiv preprint
  5. NeurIPSSelected
    Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex

    Y. Qu, Cheems Wang*, Y. Mao, H. Zou, Y. Jiang, Y. Li, W. Xu, L. Cai, W. Liu, C. Bai, K. Yang, Y. Chen, S. Yang, X. Ji

  6. arXiv preprint
    StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs

    C. Che, Z. Wang, H. Ma, Cheems Wang, Z. Shi

  7. ICML
  8. ICML
    Latent Space Robust Optimization of Neural Processes with Aligned Stratified Order-Statistic Loss Reduction

    Q. Tao, J. Wen, J. Yang, G. Wu, K. Zhang, Y. Lv, W. Du, X. Liang, Cheems Wang*

  9. ICML
  10. Nature CommunicationsSelected
    Model Predictive Task Sampling for Efficient and Robust Adaptation

    Cheems Wang*, Z. Xiao*, Y. Mao*, Y. Qu*, J. Shen, Y. Lv, X. Ji

  11. KDD
    SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

    S. Wang, H. Chi, H. Li, Z. Zhang, J. Yuan, Cheems Wang, H. Peng, X. Liu, W. Yang

  12. CIKM
    HOB: A Holistically Optimized Bidding Strategy under Heterogeneous Auction Mechanisms with Organic Traffic

    Q. Li, W. Huang, Q. Ye, W. Xu, Cheems Wang, R. Bai, W. Yuan, G. Wang, C. Yu, J. Xu

  13. ICML
    Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models

    Y. Qu, Cheems Wang*, Y. Mao, H. Zou, Y. Jiang, W. Liu, C. Bai, K. Yang, Y. Chen, S. Yang, X. Ji*

  14. arXiv preprint
  15. arXiv preprint
    VAO: Validation-Aligned Optimization for Cross-Task Generative Auto-Bidding

    Y. Lv, Z. Mou, M. Xu, J. Chen, Cheems Wang, Y. Mao, Y. Qu, R. Bai, C. Yu, J. Xu, B. Zheng, X. Ji

  16. ICLRSelected
  17. CVPR
    Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

    Z. Wang, C. Che, Cheems Wang, H. Ma, Z. Shi, C. G. M. Snoek, M. Wang

  18. KDDSelected
  19. ICLROral
    Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search

    Z. Mou, Y. Lv, M. Xu, Cheems Wang*, Y. Mao, Q. Ye, C. Li, R. Bai*, C. Yu, J. Xu, B. Zheng

  20. CVPR
    Neural Mixture Density Processes

    Y. Ding, Q. Tao, X. Liang*, L. Zhang, Y. Lv, W. Song, F. Yang, Cheems Wang*, G. Cheng

  21. AAAIOral

2025

  1. IEEE T-PAMI
    Tail Task Risk Minimization in Meta-Learning from Theoretical Advances to Practical Strategies

    Y. Lv, D. Liang, W. Du, Z. Shi, Z. Xie, Cheems Wang*, M. Wang

  2. NeurIPSSpotlight
  3. NeurIPS
    Selective Learning for Deep Time Series Forecasting

    Y. Fu, Z. Shao, C. Yu, Y. Li, Z. An, Cheems Wang, Y. Xu, F. Wang

  4. NeurIPS
    Gains: Fine-grained Federated Domain Adaptation in Open Set

    Z. Zhong, W. Jiang, W. Bao, J. Wang, Cheems Wang, G. Wang, Y. Deng, J. Ren

  5. ICMLSelected
  6. ICCV
    Separable Mixture of Low-Rank Adaptation for Continual Visual Instruction Tuning

    Z. Wang, C. Che, Cheems Wang, Y. Li, Z. Shi, M. Wang

  7. The Innovation
    Foundation Models and Intelligent Decision-Making: Progress, Challenges, and Perspectives

    J. Huang*, Y. Xu*, Q. Wang*, Cheems Wang*, X. Liang, F. Wang, Z. Zhang, W. Wei, et al.

  8. ICLR
    DynaPrompt: Dynamic Test-Time Prompt Tuning

    Z. Xiao, S. Yan, J. Hong, J. Cai, X. Jiang, Y. Hu, J. Shen, Cheems Wang, C. G. M. Snoek

  9. AAAI
    Latent Reward: LLM-Empowered Credit Assignment in Episodic Reinforcement Learning

    Y. Qu, Y. Jiang, B. Wang, Y. Mao, Cheems Wang, C. Liu, X. Ji

  10. KDDSelected

2024

  1. NeurIPS
  2. Science China Information Sciences
    Choices are More Important than Efforts: LLM Enables Efficient Multi-Agent Exploration

    Y. Qu, B. Wang, Y. Jiang, J. Shao, Y. Mao, Cheems Wang, C. Liu, X. Ji

  3. NeurIPS
  4. NeurIPS
    Doubly Mild Generalization for Offline Reinforcement Learning

    Y. Mao, Cheems Wang, Y. Qu, Y. Jiang, X. Ji

  5. NeurIPS
    GO4Align: Group Optimization for Multi-Task Alignment

    J. Shen, Cheems Wang, Z. Xiao, N. Van Noord, M. Worring

  6. KDD
    Balanced Confidence Calibration for Graph Neural Networks

    H. Yang, M. Wang, Cheems Wang, M. Lao, Y. Zhou

  7. Information Fusion
    Non-Informative Noise-Enhanced Stochastic Neural Networks for Improving Adversarial Robustness

    H. Yang, M. Wang, Cheems Wang, Z. Yu, G. Jin, C. Zhou, Y. Zhou

  8. ICML
    Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation

    Y. Yang, Y. Shi, Cheems Wang, X. Zhen, Y. Shi, J. Xu

2023

  1. NeurIPS
    A Simple Yet Effective Strategy to Robustify the Meta Learning Paradigm

    Cheems Wang, Y. Lv, Y. Feng, Z. Xie, J. Huang

  2. The Innovation
    Large-Scale Generative Simulation Artificial Intelligence: The Next Hotspot

    Cheems Wang, Y. Feng, J. Huang, Y. Lv, Z. Xie, X. Gao

  3. NeurIPSSpotlight
    Episodic Multi-Task Learning with Heterogeneous Neural Processes

    J. Shen, X. Zhen, Cheems Wang, M. Worring

  4. ICLRSelected

Blogs

  1. You Only Edit Once: A Tiny System-1 Editor That Makes LLMs Better In-Context Learners

    In-context learning lives or dies by its examples. Jev-LDE, a 1.7B editor trained with RL, glances at the retrieved demonstrations and makes a single Keep, Delete or Replace move, lifting average one-shot accuracy from 81.2% to 88.1% with just one call to the target LLM.

    中文版

  2. Branch Where It Counts: Spending Agentic RL Rollouts Where the Outcome Is Still Undecided

    Outcome-only RL learns nothing from rollouts that all succeed or all fail. TRACE replays multi-turn agents from the turns most likely to split into success and failure, getting more signal from the same budget: up to +2.8 points over GRPO on Qwen3-14B multi-hop QA (EMNLP 2026).

    中文版

  3. Your GRPO Is Secretly Chasing a Target: Listwise Policy Optimization Aims Right at It

    GRPO, Dr.GRPO and MaxRL all quietly step toward a reward-weighted target on the response simplex. LPO computes that target in closed form and projects onto it exactly, making gradients bounded, zero-sum and self-correcting at no extra compute (NeurIPS 2026).

    中文版

  4. Stop Paying for Rollouts That Teach Nothing: RLVR Where Every Batch Counts

    Half or more of a GRPO batch is often dead weight with zero gradient. POPO swaps those groups for recent useful ones from a tiny replay buffer, reaching DAPO-level accuracy with about 30% of its rollouts and half its runtime, and no extra samples.

    中文版

  5. 让大模型和机器人“聪明地交互”:用模型预测任务采样破解训练的交互瓶颈

    Training large models and embodied agents is bottlenecked less by raw compute than by costly interactions. This feature follows model predictive task sampling (MPTS), from Nature Communications to ICML, KDD and ICLR, as it learns which tasks are worth training on next.

    In Chinese

  6. Robust Fast Adaptation from Adversarially Explicit Task Distribution Generation

    How generating task distributions adversarially makes fast adaptation robust to shifts between meta-training and meta-testing tasks (KDD 2025).

Funding 科研项目

  1. 2025–2029

    异构具身多智能体协作与博弈机制

    国家自然科学基金重大项目(课题分承研)编号 62494509163 万元子课题主持,在研

  2. 2024–2026

    Neural Process 模型的多样化高保真技术研究

    国家自然科学基金青年项目编号 6230632630 万元主持,在研

  3. 2023–2026

    基于生成式仿真智能的时空轨迹学习

    某部委基础加强基金(重点项目课题分承研)95 万元子课题主持,在研

  4. 2023–2025

    基于因果强化学习的策略生成算法

    某部委工程基金185 万元主持,已结题

  5. 2023–2024

    基于生成式仿真智能的大模型稳定生成

    某部委创新科技特区基金50 万元主持,已结题

Students

He supervises master’s and Ph.D. students at Tsinghua, many of them jointly with Prof. Xiangyang Ji. The team’s research is collected on the THU-IDM website.

Main supervisor

  • Qi Tao 2024– ICML一作 / NeurIPS共一 / CVPR二作,阿里飞猪实习
  • Jiarong Wen 2025– ICML共一 / NeurIPS一作 / ICLR提交 / CIKM二作,阿里实习
  • Zi Yang 2025–
  • Kaiyu Zhang 2025– NeurIPS共一,阿里实习
  • Ziyu Wang 2026–
  • Xianke Liu 2026–
  • Yuxuan Chen 2026–
  • Hongshuo Shan 2026–
  • Shuqi Li 2027–
  • Jiaheng Liu 2027–

Co-supervised with Prof. Xiangyang Ji

  • Yixiu Mao 2024–2026 阿里千问核心团队实习
  • Yun Qu 2024– 腾讯青云计划 / 字节豆包实习
  • Heming Zou 2025– 腾讯混元实习
  • Yuhang Jiang 2024– 腾讯天美实习
  • Lizhou Cai 2026–
  • Wutong Xu 2026– 阿里妈妈 / 阿里多模态团队实习
  • Xiang Li 2026– 阿里实习
  • Yuyi Guo 2026– 腾讯混元实习

Co-supervisor

  • Wumei Du 2023–2026CIKM一作1篇,T-PAMI1篇
  • Yiqin Lv 2023–2026 NeurIPS/KDD一作3篇,T-PAMI1篇;启元/阿里妈妈/字节实习
  • Fangjie Yang 2025–
  • Yanglin Qin 2026–
  • Yi Ding 2024–2025 CVPR一作,阿里妈妈实习 / 华为2012实验室工作

MSc thesis, University of Amsterdam

  • Berend Jansen 2020–2021 Out-of-distribution detection on time series with Bayesian deep learning, with Fraudio

Join us

Dr. Cheems Wang is recruiting master’s students at Tsinghua University. The details are in Chinese below; questions are welcome at hhq123go@gmail.com.

招收硕士研究生(Base: 长沙某 985 高校数学系)

本人年近四旬,坚持每年写 1–2 篇文章,沉迷科研,事必躬亲。课题组氛围很卷(实际从事人工智能研究,Domain Generalization from Math to Computer Science)。保研生请发邮件至 hhq123go@gmail.com,我的微信:Cheems_QW。

具体要求:数学基础扎实(本人不聚焦纯数研究,申请者仅需要对微积分、线性代数和优化有一定了解,具有一定编程能力即可),具有团队精神,吃苦耐劳,聚焦生成式 AI 前沿研究。本团队稳定发表 ICML、NeurIPS、ICLR 等顶会,业绩优异者可推荐大厂。

科研需要激情、热情和爱好。硕士毕业要求为 2–3 篇 A 类会议(本人会亲自设计研究思路、打磨写作、把关投稿全流程,但需要紧密配合)。报考本团队请三思而行。

Contact

  1. Email
    cheemswang@mail.tsinghua.edu.cn or hhq123go@gmail.com
  2. WeChat
    Cheems_QW. Team accounts on WeChat and Xiaohongshu (微信公众号 / 小红书): THU-IDM or SIG-IDM
  3. Social
    X @AlbertW24045555 and 知乎 Zhihu
  4. Profiles
    Google Scholar (last five years, complete) and OpenReview
  5. Team
    THU-IDM, with code at github.com/thu-rllab
  6. Office
    Tsinghua University, Haidian District, Beijing, China