Ziyan Wang

Ziyan Wang

Postdoctoral Fellow · Harvard University

Email: ziyan.wang[at]kcl[dot]ac[dot]uk

Research overview

I am a Postdoctoral Fellow at Harvard University, working with Prof. Milind Tambe. My work studies how learning agents can coordinate, communicate, and act safely in complex environments.

My central question is how to distill executable policies from human knowledge. Human knowledge appears as thinking patterns, direct instruction, books, and collective behavior; my work asks how learning agents can turn these media into robust decision-making policies.

In reinforcement learning and MARL, I study policy learning from books (PLFB), human feedback (M3HF), causal credit assignment (MACCA, GRD), and constrained decision-making (MACPO, SMALL).

In language-agent systems, I study how instructions, social interaction, and shared memory shape agent behavior, including instruction relabeling, strategic discussion, mixed-motive generalization, marketplace safety (BazaarBench), and context management.

Before Harvard, I conducted my Ph.D. research at the Cooperative AI Lab, King's College London, supervised by Dr Yali Du and Prof. Sanjay Modgil. My research experience also includes the Oxford IDAI Fellowship with Dr Adel Bibi and Prof. Philip Torr, the Future AI Group at Microsoft Research Cambridge, a visit to Carnegie Mellon University with Prof. Fei Fang, and Microsoft Research's AI Frontier Group in Redmond.

News

Oct 2026 Memento: Teaching LLMs to Manage Their Own Context has been selected for COLM 2026 Oral Spotlight (24/856)!
Oct 2026 Fisher Decorator: Refining Flow Policy via a Local Transport Map has been accepted to NeurIPS 2026!
Oct 2026 We have open-sourced BazaarBench, our benchmark for open-ended multi-agent safety evaluation in C2C marketplaces.
Jul 2026 Memento has been accepted to COLM 2026!
May 2026 Started a research internship at Microsoft Research Cambridge, focusing on multi-agent LLM communication and coordination.
Feb 2026 Started the Oxford IDAI Fellowship at the University of Oxford, working with Dr Adel Bibi and Prof. Philip Torr.

Research direction

Distilling Policy from Human Knowledge

My research goal is to distill policies from human knowledge. Knowledge can be implicit in reasoning patterns, shaped through direct instruction, preserved in books, and amplified through collective behavior; my papers study how agents can learn from these media to coordinate, adapt, and act safely.

Thinking patterns Direct instruction Books Collective behavior
Four media of human knowledge: thinking patterns, direct instruction, books, and collective behavior.

Experience & Visits

Postdoctoral Fellow

Harvard University, Cambridge, MA, US · Present

Working with Prof. Milind Tambe.

Research Internship, Future AI Group

Microsoft Research Cambridge, Cambridge, UK · May 2026 - present

Working on multi-agent LLM communication, coordination, and collaborative agent behavior.

Oxford IDAI Fellowship

University of Oxford, Oxford, UK · Feb. 2026 - present

Working with Dr Adel Bibi and Prof. Philip Torr on real-time multi-agent LLM anomaly detection and monitoring.

Research Internship, AI Frontier Group

Microsoft Research, Redmond, US · Sep. 2025 - Dec. 2025

Worked with Vaishnavi Shrivastava and Prof. Dimitris Papailiopoulos on LLM pre-training and reasoning.

Visiting Ph.D. Student

Carnegie Mellon University, Pittsburgh, US · Feb. 2025 - Jun. 2025

Visited Prof. Fei Fang's group, working on multi-agent learning and AI for social impact.

Selected Publications

* equal contribution, ✉ corresponding author

  1. Main figure for Safe Multi-agent Reinforcement Learning with Natural Language Constraints
    AAAI'26

    Safe Multi-agent Reinforcement Learning with Natural Language Constraints

    Ziyan Wang , Meng Fang , Tristan Tomilin , Fei Fang and Yali Du
    Alignment Track of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI) 2026
  2. Main figure for M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality
    ICML'25

    M3HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed Quality

    Ziyan Wang , Zhicheng Zhang , Fei Fang and Yali Du
    Forty-Second International Conference on Machine Learning (ICML) 2025
  3. Main figure for MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment
    TMLR

    MACCA: Offline Multi-agent Reinforcement Learning with Causal Credit Assignment

    Ziyan Wang , Yali Du , Yudi Zhang , Meng Fang and Biwei Huang
    Transactions on Machine Learning Research (TMLR) 2025
  4. Oral Main figure for Policy Learning from Tutorial Books via Understanding, Rehearsing and Introspecting
    NeurIPS'24

    Policy Learning from Tutorial Books via Understanding, Rehearsing and Introspecting

    Xiong-Hui Chen* , Ziyan Wang* , Yali Du , Shengyi Jiang , Meng Fang , Yang Yu and Jun Wang
    The Thirty-Eighth Annual Conference on Neural Information Processing Systems (NeurIPS) 2024
  5. Main figure for Learning to Discuss Strategically: A Case Study on One Night Ultimate Werewolf
    NeurIPS'24

    Learning to Discuss Strategically: A Case Study on One Night Ultimate Werewolf

    Xuanfa Jin* , Ziyan Wang* , Yali Du , Meng Fang , Haifeng Zhang and Jun Wang
    The Thirty-Eighth Annual Conference on Neural Information Processing Systems (NeurIPS) 2024
  6. Main figure for ChessGPT: Bridging Policy Learning and Language Modeling
    NeurIPS'23

    ChessGPT: Bridging Policy Learning and Language Modeling

    Xidong Feng , Yicheng Luo , Ziyan Wang , Hongrui Tang , Mengyue Yang , Kun Shao , David Mguni , Yali Du and Jun Wang
    The Thirty-Seventh Annual Conference on Neural Information Processing Systems (NeurIPS) 2023

Honors & Teaching

  • Honors: COLM 2026 Oral Spotlight, Oxford IDAI Fellowship, NeurIPS 2024 Scholar Award, NeurIPS 2024 Oral Presentation
  • Teaching: Oxford Machine Learning Summer School, Oxford MLx Fundamentals Summer School, and Optimisation Methods at King’s College London

Professional Services

  • Conference reviewer for ICML 2023/24/25/26, NeurIPS 2023/24/25/26, ICLR 2024/25/26/27, AISTATS 2025/26, AAAI 2026, ACL ARR 2026, and AAMAS 2025/26/27
  • Journal reviewer for IEEE Robotics and Automation Letters, IEEE Transactions on Knowledge and Data Engineering, and IEEE Transactions on Artificial Intelligence