Kelly Buchanan

Kelly Buchanan - Stanford Postdoctoral Scholar

I’m a Postdoctoral Scholar at Stanford University working with Scott Linderman and Christopher Ré. My research focuses on reliability in AI systems. I develop evaluations for the reasoning capabilities (especially coding) of frontier AI models, and use the ensuing insights to build better agentic systems that can verify their own reasoning, quantify uncertainty, and operate safely at scale.

Some of my recent projects include:

  • Terminal-Bench (ICLR, 2026), a benchmark for evaluating agentic systems in realistic terminal environments. Now widely used to test open and closed-source foundation models under real-world constraints.

  • Weaver (NeurIPS, 2025), a verification framework that fuses weak but cost-efficient verifiers into unified correctness estimates, closing the generation–verification gap.

I completed my PhD at Columbia University’s Center for Theoretical Neuroscience, where I worked with Liam Paninski and John Cunningham. My research developed ML methods for functional imaging, pose estimation, and behavioral segmentation, used by organizations including Q-state Biosciences, the International Brain Laboratory, and the Zuckerman Institute.

I also spent time at Google as a Student Researcher on the Reliable Deep Learning team, evaluating reliability of large language and vision models.

I earned my Bachelor’s and Master’s degrees with Honors in Electrical Engineering from KU.

Selected Publications

  1. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
    The Terminal-Bench Team
    ICLR, 2026
  2. A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
    Xavier Gonzalez*, E. Kelly Buchanan*, Hyun Dong Lee, and 5 more authors
    TMLR, 2026
  3. Weaver: Shrinking the Generation-Verification Gap by Scaling Compute for Verification
    Jon Saad-Falcon*, E. Kelly Buchanan*, Mayee F. Chen*, and 9 more authors
    NeurIPS, 2025