I am a Postdoctoral Fellow at the ETH AI Center, working on AI safety and the trustworthy application of AI in industrial domains. I earned my Ph.D. in Applied Mathematics from Harvard University’s John A. Paulson School of Engineering and Applied Sciences, and my alma mater is NYU Courant.

My research focuses on trustworthy AI for decision-making and AI safety. I develop frameworks and methods that enable machine learning and foundation models to operate responsibly in the real world, with the goal of making AI reliable, accountable, and safe when deployed in high-stakes domains.

My recent work addresses the following questions:

  1. How can we make autonomous AI systems capable, reliable, and safe in complex multi-agent environments? Focusing on supply chains, I built the first live LLM-powered simulation of the Beer Game, a multi-echelon testbed, and showed that autonomous GenAI agents can cut costs by up to 80% compared to human teams. I then identified a key risk, agent bullwhip: agents make inconsistent decisions in identical situations, and this instability compounds across the supply chain into occasional outsized errors. I developed guardrails and reinforcement learning post-training methods that mitigate this unreliability.
  2. Who is accountable for what AI generates, and how can we enforce that accountability? I co-developed ArcMark, a distortion-free multi-bit watermark that embeds multiple bytes of information, such as user or model IDs, into just a few hundred tokens, so that generated text can be traced back to its source. I also developed zero-bit watermarks that detect AI-generated text with high accuracy and minimal distortion, including the Correlated-Channel watermark and HeavyWater and SimplexWater, revealing new connections between LLM watermarking and coding theory.
  3. How can we make AI-assisted decisions trustworthy and unbiased for every individual? I showed that equally accurate models can make conflicting, arbitrary predictions for the same individual, and that common bias-mitigation interventions can amplify this arbitrariness. I also characterized the limits of verifying personalized models, and developed Kernel Multiaccuracy, a one-step method to correct biased predictions across many subgroups at once.
Keywords: AI Safety Agentic AI Multi-Agent Systems LLM Watermarking AI for Supply Chain Management AI for Manufacturing Model Multiplicity Algorithmic Fairness

I have published in leading venues across machine learning, information theory, and management, including NeurIPS, IEEE ISIT, and Harvard Business Review.

News

Publications

  • ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport
    Atefeh Gilani, Sajani Vithana, Carol Xuan Long, Oliver Kosut, Lalitha Sankar, Flavio P Calmon
    Advances in Neural Information Processing Systems (NeurIPS), 2026.
    TL/DR

    We formulate distortion-free watermarking as a channel coding problem and derive its information-theoretic capacity. Guided by this limit, we propose **ArcMark**, which embeds multiple bytes of information into a few hundred tokens without distorting the next-token distribution, and outperforms competing multi-bit watermarks in reconstruction accuracy, including under attacks.

  • Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management
    Carol Xuan Long, David Simchi-Levi, Feng Zhu, Huangyuan Su, Andre P Calmon, Flavio P Calmon
    Preprint, 2026.
    TL/DR

    In the MIT Beer Game, GenAI agents reduce total supply chain costs by up to 80% relative to human teams, but can exhibit large run-to-run instability. We characterize this as **agent bullwhip** and show that reinforcement learning post-training and operational guardrails both reduce tail events and mitigate it.

  • When Supply Chains Become Autonomous
    Carol Xuan Long, David Simchi-Levi, Andre P Calmon, Flavio P Calmon
    Harvard Business Review, 2025.

  • HeavyWater and SimplexWater: Watermarking Low-Entropy Text Distributions
    Dor Tsur*, Carol Xuan Long*, Claudio M. Verdun, Hsiang Hsu, Chen-Fu Chen, Haim Permuter, Sajani Vithana, Flavio P Calmon
    Advances in Neural Information Processing Systems (NeurIPS), 2025.
    TL/DR

    Our goal is to design watermarks that optimally use side information to maximize detection accuracy and minimize distortion of generated text. We propose two watermarks **HeavyWater** and **SimplexWater** that achieve SOTA performance. Our theoretical analysis also reveals surprising new connections between LLM watermarking and **coding theory**.

  • Optimized Couplings for Watermarking Large Language Models, (slides)
    Carol Xuan Long*, Dor Tsur*, Claudio M. Verdun, Hsiang Hsu, Haim Permuter, Flavio P Calmon
    IEEE International Symposium on Information Theory (ISIT), 2025.
    TL/DR

    We argue that a key component in watermark design is generating a coupling between the side information shared with the watermark detector and a random partition of the LLM vocabulary. Our analysis identifies the optimal coupling and randomization strategy under the worst-case LLM next-token distribution that satisfies a min-entropy constraint. We propose the **Correlated-Channel watermarking scheme** --- a closed-form scheme that achieves high detection at zero distortion.

  • Kernel Multiaccuracy, (slides)
    Carol Xuan Long, Wael Alghamdi, Alexander Glynn, Yixuan Wu, Flavio P Calmon
    Foundations of Responsible Computing (FORC), 2025.
    TL/DR

    We connect multi-group notions with *Integral Probability Metrics*, and propose **KMAcc** --- a non-iterative, one-step optimization to correct multiaccuracy errors in the kernel space.

  • Predictive Churn with the Set of Good Models
    Jamelle Watson-Daniels, Flavio P Calmon, Alexander D’Amour, Carol Xuan Long, David C. Parkes, Berk Ustun
    Under Review, 2024.
    TL/DR

    We study the effect of predictive churn — flips in predictions across ML model updates — through the lens of predictive multiplicity – i.e., the prevalence of conflicting predictions over the set of near-optimal models (the ε-Rashomon set).

  • Multi-Group Proportional Representation in Retrieval
    Alex Osterling, Claudio M Verdun, Carol Xuan Long, Alexander Glynn, Lucas Monteiro Paes, Sajani Vithana, Martina Cardone, Flavio P Calmon
    Advances in Neural Information Processing Systems (NeurIPS), 2024.
    TL/DR

    We introduce Multi-Group Proportional Representation (MPR), a novel metric that measures representation across intersectional groups. We propose practical methods and algorithms for estimating and ensuring MPR in image retrieval, with minimal compromise in retrieval accuracy.

  • Individual Arbitrariness and Group Fairness
    Carol Xuan Long, Hsiang Hsu, Wael Alghamdi, Flavio P Calmon
    Advances in Neural Information Processing Systems (NeurIPS), 2023, Spotlight Paper.
    TL/DR

    Fairness interventions in machine learning optimized solely for group fairness and accuracy can exacerbate predictive multiplicity. A third axis of “arbitrariness” should be considered when deploying models to aid decision-making in applications of individual-level impact.

  • On the epistemic limits of personalized prediction
    Lucas Monteiro Paes*, Carol Long*, Berk Ustun, Flavio Calmon (* Equal Contribution)
    Advances in Neural Information Processing Systems (NeurIPS), 2022.
    TL/DR

    It is impossible to reliably verify that a personalized classifier with $k \geq 19$ binary group attributes will benefit every group that provides personal data using a dataset of $n = 8 × 10^9$ samples – one for each person in the world.

Beyond Research

Outside of work, I am a globetrotter, dancer, and music-lover. Growing up as a swimmer, I enjoy sports. From completing a half-marathon and recovering from an ACL injury, I’ve collected many stories to tell (for better or worse!). Whenever I can, I head outdoors — my top three U.S. national parks are Yellowstone, the Grand Canyon, and Mount Rainier.