Hi there Welcome to my Homepage!

Hi! My name is Xutao Mao. I am a Ph.D. student at the City University of Hong Kong, advised by Prof. Cong Wang. I also closely work with Prof. Xiang Zheng and Prof. Bo Han. Before that, I received my B.S. in Computer Science and Mathematics from Vanderbilt University.

My research focuses on the principles for building trustworthy AI โ€” especially the safety and alignment of agentic and multimodal language models. Feel free to reach out if you are interested in collaboration or potential opportunities.

News

  • 2026.09 ๐ŸŽ“๐ŸŽ“ I am starting my Ph.D. at the City University of Hong Kong.
  • 2026.07 ๐Ÿš€๐Ÿš€ We released two agent-safety projects: AHA (Agent Hacks Agent) and PASB (Persistent Sycophancy Benchmark).
  • 2026.05 ๐ŸŽ‰๐ŸŽ‰ Our paper STARE accepted to ICML 2026 (Poster).
  • 2025.11 ๐ŸŽ‰๐ŸŽ‰ Our paper MindVote and LogicCat accepted to AAAI 2026 as Oral and Poster.

Experience

City University of Hong Kong
2026.09 - Present
Ph.D. in Computer Science, advised by Prof. Cong Wang
Research: agent safety and trustworthy AI.
Vanderbilt University
2022.08 - 2026.04
B.S. in Computer Science & Mathematics

Publications

(* equal contribution ยท โ€  corresponding author)

STARE
STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack
Xutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng†, Hongying Zan, Cong Wang†
A hierarchical-RL red-team engine with step-wise temporal attribution for multi-modal toxicity attacks, exceeding baselines by ~68% attack success rate with strong transferability.
ICML 2026 Poster   [arXiv] [code]
MindVote
MindVote: When AI Meets the Wild West of Social Media Opinion
Xutao Mao†, Ezra Xuanru Tao, Leyao Wang
A benchmark probing how large language models predict and reason about real-world social-media opinion.
AAAI 2026 Oral   [arXiv] [code]
AHA
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Xutao Mao, Xiang Zheng†, Cong Wang†
An autoresearch red-team framework that discovers reusable vulnerability concepts (a Vulnerability Concept Graph) in production LLM agents; the frozen graph, deployed single-shot on held-out data, beats the strongest baseline by 14.2 points.
Preprint ยท Under Review   [arXiv] [code] [project]
PASB
Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
Xutao Mao*, Liangjie Zhao*, Leyao Wang, Rui Qian, Qiang Huang, Wentao Wang, Bo Han†, Xiang Zheng†, Cong Wang†
A 1,600-task benchmark for persistent sycophancy in stateful personal agents; crossing the commit boundary is the single largest downstream-failure jump (+27 points), driven by attribution stripping.
Preprint ยท Under Review   [arXiv] [code] [project] [dataset]
TAME
Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation-Level Enforcement
Xutao Mao, Jianing Zhu, Jinman Zhao, Tongliang Liu, Xiaowen Chu, Cong Wang†, Bo Han†
Mechanistic analysis and activation-level enforcement that tame chain-of-thought obfuscation in RL-trained VLMs, improving reasoning monitorability by ~60% over GRPO.
Preprint ยท Under Review  
What Happens Inside Agent Memory?
What Happens Inside Agent Memory? Circuit Analysis from Emergence to Diagnosis
Xutao Mao, Jinman Zhao, Gerald Penn, Cong Wang†
A mechanistic circuit analysis tracing how agent-memory features emerge and specialize inside LLMs โ€” from first-person subject anchoring to category-aggregation hubs โ€” enabling diagnosis of memory behavior.
Preprint ยท Under Review   [arXiv]

Scholarships

  • 2026, Hong Kong Postgraduate Scholarships (during PhD study).
  • 2025, Vanderbilt Summer Research Program Scholarship.

Services

  • Reviewer, AAAI 2026 / 2027 ยท TheWebConf 2026 ยท NeurIPS 2026.