Hi there
Welcome to my Homepage!
Hi! My name is Xutao Mao. I am a first-year Ph.D. student at the City University of Hong Kong, advised by Prof. Cong Wang. I also closely work with Prof. Xiang Zheng and Prof. Bo Han. Before that, I received my B.S. in Computer Science and Mathematics from Vanderbilt University.
My research asks how AI agents can grow more capable without becoming less safe. I study failures that emerge as agents act, remember, and learn from experience, with a focus on agent safety and recursive self-improvement. My work spans automated red-teaming, persistent agent state, multimodal safety, and mechanistic oversight. Feel free to reach out if you are interested in collaboration.
Research Interests
As AI agents gain autonomy, memory, and the ability to improve themselves, I am interested in three questions:
What happens when the world pushes an agent around? I attack production agents before someone else does. The failures I find do not stay where they started: they travel from one deployed system to the next, while multimodal models bring surprises of their own (STARE, AHA).
What happens when an agentβs own past turns against it? Agents keep notes, and the notes bite back. A userβs casual opinion gets written into memory and starts bossing later chats around (PASB), and one unsafe lesson becomes a habit the agent keeps reusing (MisEvolve). So I also work on where these memories come from and who gets to write them, to keep self-improvement on the rails (MemMark).
Can we watch this happen inside the model, and step in? I open the box: tracing the circuits that decide what an agent writes into memory and what it pulls back out (Agent Memory), then turning what we find inside into tools that watch and steer models at the activation level (TAME).
News
- 2026.09 ππ I am starting my Ph.D. at the City University of Hong Kong.
- 2026.08 ππ Our paper MemMark was accepted to Findings of EMNLP 2026.
- 2026.08 ππ We released the code for MISEvolve, our work on skill misevolution in self-improving LLM agents.
- 2026.07 ππ We released two agent-safety projects: AHA (Agent Hacks Agent) and PASB (Persistent Sycophancy Benchmark).
- 2026.05 ππ Our paper STARE accepted to ICML 2026 (Poster).
- 2025.11 ππ Our paper MindVote and LogicCat accepted to AAAI 2026 as Oral and Poster.
Experience

2026.09 - Present
Ph.D. in Computer Science, advised by Prof. Cong Wang
Research: agent safety and recursive self-improvement.

2022.08 - 2026.04
B.S. in Computer Science & Mathematics
Publications
(* equal contribution Β· β corresponding author)
Xutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng†, Hongying Zan, Cong Wang†
A hierarchical-RL red-team engine with step-wise temporal attribution for multi-modal toxicity attacks; it reveals how harmful content develops across generations, enabling more precise safety evaluation.
ICML 2026 Poster [arXiv] [code]
Xutao Mao, Xiang Zheng†, Cong Wang†
An autoresearch framework that turns agent red-teaming discoveries into a reusable Vulnerability Concept Graph; it makes transferable vulnerability knowledge actionable against unseen production agents.
Preprint [arXiv] [code] [project]
Xutao Mao*, Liangjie Zhao*, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han†, Xiang Zheng†, Cong Wang†
A 1,600-task benchmark tracing how sycophancy persists through stateful agent memory; it shows how unsafe agreement survives interaction boundaries and compounds into downstream failures.
Preprint [arXiv] [code] [project] [dataset]
- EMNLP 2026 Findings MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems
[arXiv][code][project] - ICML 2026 Poster STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
[arXiv][code] - AAAI 2026 Oral MindVote: When AI Meets the Wild West of Social Media Opinion
[arXiv][code] - AAAI 2026 Poster LogicCat: A Text-to-SQL Benchmark for Multi-Domain Reasoning Challenges
* equal contribution, listed in alphabetical order [arXiv][code] - Preprint Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
[arXiv][code][project] - Preprint Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
[arXiv][code][project][dataset] - Preprint Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation-Level Enforcement
- Preprint What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
[arXiv] - Preprint Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
[arXiv][code] - Preprint Towards Bridging Review Sparsity in Recommendation with Textual Edge Graph Representation
[arXiv]
Scholarships
- 2026, Hong Kong Postgraduate Scholarships (during PhD study).
- 2025, Vanderbilt Summer Research Program Scholarship.
Services
- Reviewer, AAAI 2026 / 2027 Β· TheWebConf 2026 Β· NeurIPS 2026.
Collaboration
I'm always happy to discuss new research ideas and explore potential collaborations, especially around agent safety, recursive self-improvement, and mechanistic oversight. If our interests overlap, feel free to reach out.




