Hi there Welcome to my Homepage!

Hi! My name is Xutao Mao. I am a first-year Ph.D. student at the City University of Hong Kong, advised by Prof. Cong Wang. I also closely work with Prof. Xiang Zheng and Prof. Bo Han. Before that, I received my B.S. in Computer Science and Mathematics from Vanderbilt University.

My research asks how AI agents can grow more capable without becoming less safe. I study failures that emerge as agents act, remember, and learn from experience, with a focus on agent safety and recursive self-improvement. My work spans automated red-teaming, persistent agent state, multimodal safety, and mechanistic oversight. Feel free to reach out if you are interested in collaboration.

Research Interests

As AI agents gain autonomy, memory, and the ability to improve themselves, I am interested in three questions:

  • What happens when the world pushes an agent around? I attack production agents before someone else does. The failures I find do not stay where they started: they travel from one deployed system to the next, while multimodal models bring surprises of their own (STARE, AHA).

  • What happens when an agent’s own past turns against it? Agents keep notes, and the notes bite back. A user’s casual opinion gets written into memory and starts bossing later chats around (PASB), and one unsafe lesson becomes a habit the agent keeps reusing (MisEvolve). So I also work on where these memories come from and who gets to write them, to keep self-improvement on the rails (MemMark).

  • Can we watch this happen inside the model, and step in? I open the box: tracing the circuits that decide what an agent writes into memory and what it pulls back out (Agent Memory), then turning what we find inside into tools that watch and steer models at the activation level (TAME).

News

  • 2026.09 πŸŽ“πŸŽ“ I am starting my Ph.D. at the City University of Hong Kong.
  • 2026.08 πŸŽ‰πŸŽ‰ Our paper MemMark was accepted to Findings of EMNLP 2026.
  • 2026.08 πŸš€πŸš€ We released the code for MISEvolve, our work on skill misevolution in self-improving LLM agents.
  • 2026.07 πŸš€πŸš€ We released two agent-safety projects: AHA (Agent Hacks Agent) and PASB (Persistent Sycophancy Benchmark).
  • 2026.05 πŸŽ‰πŸŽ‰ Our paper STARE accepted to ICML 2026 (Poster).
  • 2025.11 πŸŽ‰πŸŽ‰ Our paper MindVote and LogicCat accepted to AAAI 2026 as Oral and Poster.

Experience

City University of Hong Kong
2026.09 - Present
Ph.D. in Computer Science, advised by Prof. Cong Wang
Research: agent safety and recursive self-improvement.
Vanderbilt University
2022.08 - 2026.04
B.S. in Computer Science & Mathematics

Publications

(* equal contribution Β· † corresponding author)

STARE
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
Xutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng†, Hongying Zan, Cong Wang†
A hierarchical-RL red-team engine with step-wise temporal attribution for multi-modal toxicity attacks; it reveals how harmful content develops across generations, enabling more precise safety evaluation.
ICML 2026 Poster   [arXiv] [code]
MindVote
MindVote: When AI Meets the Wild West of Social Media Opinion
Xutao Mao†, Ezra Xuanru Tao, Leyao Wang
A benchmark for predicting and reasoning about real-world social-media opinion; it grounds LLM evaluation in the complexity of noisy, polarized public discourse.
AAAI 2026 Oral   [arXiv] [code]
AHA
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Xutao Mao, Xiang Zheng†, Cong Wang†
An autoresearch framework that turns agent red-teaming discoveries into a reusable Vulnerability Concept Graph; it makes transferable vulnerability knowledge actionable against unseen production agents.
Preprint   [arXiv] [code] [project]
PASB
Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
Xutao Mao*, Liangjie Zhao*, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han†, Xiang Zheng†, Cong Wang†
A 1,600-task benchmark tracing how sycophancy persists through stateful agent memory; it shows how unsafe agreement survives interaction boundaries and compounds into downstream failures.
Preprint   [arXiv] [code] [project] [dataset]

Scholarships

  • 2026, Hong Kong Postgraduate Scholarships (during PhD study).
  • 2025, Vanderbilt Summer Research Program Scholarship.

Services

  • Reviewer, AAAI 2026 / 2027 Β· TheWebConf 2026 Β· NeurIPS 2026.

Collaboration

Let's discuss ideas and build something meaningful together.

I'm always happy to discuss new research ideas and explore potential collaborations, especially around agent safety, recursive self-improvement, and mechanistic oversight. If our interests overlap, feel free to reach out.

xutao.henry.mao@gmail.com