I am a Postdoctoral Fellow at the ETH AI Center, working on AI safety and the trustworthy application of AI in industrial domains. I earned my Ph.D. in Applied Mathematics from Harvard University’s John A. Paulson School of Engineering and Applied Sciences, and my alma mater is NYU Courant.
My research focuses on trustworthy AI for decision-making and AI safety. I develop frameworks and methods that enable machine learning and foundation models to operate responsibly in the real world, with the goal of making AI reliable, accountable, and safe when deployed in high-stakes domains.
My recent work addresses the following questions:
- How can we make autonomous AI systems capable, reliable, and safe in complex multi-agent environments? Focusing on supply chains, I built the first live LLM-powered simulation of the Beer Game, a multi-echelon testbed, and showed that autonomous GenAI agents can cut costs by up to 80% compared to human teams. I then identified a key risk, agent bullwhip: agents make inconsistent decisions in identical situations, and this instability compounds across the supply chain into occasional outsized errors. I developed guardrails and reinforcement learning post-training methods that mitigate this unreliability.
- Who is accountable for what AI generates, and how can we enforce that accountability? I co-developed ArcMark, a distortion-free multi-bit watermark that embeds multiple bytes of information, such as user or model IDs, into just a few hundred tokens, so that generated text can be traced back to its source. I also developed zero-bit watermarks that detect AI-generated text with high accuracy and minimal distortion, including the Correlated-Channel watermark and HeavyWater and SimplexWater, revealing new connections between LLM watermarking and coding theory.
- How can we make AI-assisted decisions trustworthy and unbiased for every individual? I showed that equally accurate models can make conflicting, arbitrary predictions for the same individual, and that common bias-mitigation interventions can amplify this arbitrariness. I also characterized the limits of verifying personalized models, and developed Kernel Multiaccuracy, a one-step method to correct biased predictions across many subgroups at once.
I have published in leading venues across machine learning, information theory, and management, including NeurIPS, IEEE ISIT, and Harvard Business Review.
News
- Sep 2026Received the NUS Global Postdoctoral Award with the Department of Industrial Systems Engineering and Management (ISEM), College of Design and Engineering, National University of Singapore.
- Sep 2026ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport was accepted at NeurIPS 2026. ArcMark reliably embeds multiple bytes of information, such as a user ID or model version, into just a few hundred tokens without distorting the LLM's output.
- May 2026New paper: Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management. We identify agent bullwhip, where decision instability compounds across autonomous agents into outsized errors, and show that guardrails and reinforcement learning post-training mitigate it.
- Mar 2026Defended my Ph.D. thesis, Trustworthy AI: Ensuring Reliability and Accountability from Models to Agents.
- Feb 2026Presented Can GenAI Agents Manage a Supply Chain? at the MIT Sloan System Dynamics Seminar and the USI SD Research Lab.
- Dec 2025Published When Supply Chains Become Autonomous in Harvard Business Review (full paper with plots). We introduce a testbed based on the MIT Beer Game, in which GenAI agents autonomously manage a multi-echelon supply chain, and evaluate their performance.
- Sep 2025Launched the first live simulation of the Beer Game powered by LLMs, a joint project of Harvard, MIT, and Georgia Tech (Harvard Information Theory Lab, MIT Data Science Lab, and Georgia Tech Scheller College of Business).
- Jul 2025Presented Optimized Couplings for Watermarking Large Language Models (slides) at ISIT 2025, University of Michigan.
- Jun 2025Presented Kernel Multiaccuracy (slides) at FORC 2025, Stanford University.
Publications
- ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
Atefeh Gilani, Sajani Vithana, Carol Xuan Long, Oliver Kosut, Lalitha Sankar, Flavio P Calmon
Advances in Neural Information Processing Systems (NeurIPS), 2026.TL/DR
We formulate distortion-free watermarking as a channel coding problem and derive its information-theoretic capacity. Guided by this limit, we propose **ArcMark**, which embeds multiple bytes of information into a few hundred tokens without distorting the next-token distribution, and outperforms competing multi-bit watermarks in reconstruction accuracy, including under attacks.
- Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management
Carol Xuan Long, David Simchi-Levi, Feng Zhu, Huangyuan Su, Andre P Calmon, Flavio P Calmon
Preprint, 2026.TL/DR
In the MIT Beer Game, GenAI agents reduce total supply chain costs by up to 80% relative to human teams, but can exhibit large run-to-run instability. We characterize this as **agent bullwhip** and show that reinforcement learning post-training and operational guardrails both reduce tail events and mitigate it.
When Supply Chains Become Autonomous
Carol Xuan Long, David Simchi-Levi, Andre P Calmon, Flavio P Calmon
Harvard Business Review, 2025.- HeavyWater and SimplexWater: Watermarking Low-Entropy Text Distributions
Dor Tsur*, Carol Xuan Long*, Claudio M. Verdun, Hsiang Hsu, Chen-Fu Chen, Haim Permuter, Sajani Vithana, Flavio P Calmon
Advances in Neural Information Processing Systems (NeurIPS), 2025.TL/DR
Our goal is to design watermarks that optimally use side information to maximize detection accuracy and minimize distortion of generated text. We propose two watermarks **HeavyWater** and **SimplexWater** that achieve SOTA performance. Our theoretical analysis also reveals surprising new connections between LLM watermarking and **coding theory**.
- Optimized Couplings for Watermarking Large Language Models, (slides)
Carol Xuan Long*, Dor Tsur*, Claudio M. Verdun, Hsiang Hsu, Haim Permuter, Flavio P Calmon
IEEE International Symposium on Information Theory (ISIT), 2025.TL/DR
We argue that a key component in watermark design is generating a coupling between the side information shared with the watermark detector and a random partition of the LLM vocabulary. Our analysis identifies the optimal coupling and randomization strategy under the worst-case LLM next-token distribution that satisfies a min-entropy constraint. We propose the **Correlated-Channel watermarking scheme** --- a closed-form scheme that achieves high detection at zero distortion.
- Kernel Multiaccuracy, (slides)
Carol Xuan Long, Wael Alghamdi, Alexander Glynn, Yixuan Wu, Flavio P Calmon
Foundations of Responsible Computing (FORC), 2025.TL/DR
We connect multi-group notions with *Integral Probability Metrics*, and propose **KMAcc** --- a non-iterative, one-step optimization to correct multiaccuracy errors in the kernel space.
- Predictive Churn with the Set of Good Models
Jamelle Watson-Daniels, Flavio P Calmon, Alexander D’Amour, Carol Xuan Long, David C. Parkes, Berk Ustun
Under Review, 2024.TL/DR
We study the effect of predictive churn — flips in predictions across ML model updates — through the lens of predictive multiplicity – i.e., the prevalence of conflicting predictions over the set of near-optimal models (the ε-Rashomon set).
- Multi-Group Proportional Representation in Retrieval
Alex Osterling, Claudio M Verdun, Carol Xuan Long, Alexander Glynn, Lucas Monteiro Paes, Sajani Vithana, Martina Cardone, Flavio P Calmon
Advances in Neural Information Processing Systems (NeurIPS), 2024.TL/DR
We introduce Multi-Group Proportional Representation (MPR), a novel metric that measures representation across intersectional groups. We propose practical methods and algorithms for estimating and ensuring MPR in image retrieval, with minimal compromise in retrieval accuracy.
- Individual Arbitrariness and Group Fairness
Carol Xuan Long, Hsiang Hsu, Wael Alghamdi, Flavio P Calmon
Advances in Neural Information Processing Systems (NeurIPS), 2023, Spotlight Paper.TL/DR
Fairness interventions in machine learning optimized solely for group fairness and accuracy can exacerbate predictive multiplicity. A third axis of “arbitrariness” should be considered when deploying models to aid decision-making in applications of individual-level impact.
- On the epistemic limits of personalized prediction
Lucas Monteiro Paes*, Carol Long*, Berk Ustun, Flavio Calmon (* Equal Contribution)
Advances in Neural Information Processing Systems (NeurIPS), 2022.TL/DR
It is impossible to reliably verify that a personalized classifier with $k \geq 19$ binary group attributes will benefit every group that provides personal data using a dataset of $n = 8 × 10^9$ samples – one for each person in the world.
Beyond Research
Outside of work, I am a globetrotter, dancer, and music-lover. Growing up as a swimmer, I enjoy sports. From completing a half-marathon and recovering from an ACL injury, I’ve collected many stories to tell (for better or worse!). Whenever I can, I head outdoors — my top three U.S. national parks are Yellowstone, the Grand Canyon, and Mount Rainier.
