Profile photo of Chris Agia

Chris Agia

I am a final year CS PhD candidate at Stanford University in the Stanford AI Lab (SAIL), where I'm advised by Jeannette Bohg and Marco Pavone. My research spans several topics in AI, with a particular focus on systems that operate in the physical world (robots):

  1. Increasing the reliability of AI-driven systems during deployment;
  2. Establishing connections between model behavior and training data;
  3. Expanding full-stack capabilities through planning and reasoning;
  4. Facilitating effective interaction between systems and humans.

Previously, I received a BASc from the University of Toronto (EngSci), where I was advised by Florian Shkurti. I've also had the pleasure of interning at Meta FAIR, NASA JPL, Microsoft, and Google in past summers.

Beyond research, I enjoy sports (Stanford Club Men's Soccer) and music!

cagia[at]cs.stanford.edu  /  CV  /  LinkedIn  /  X  /  GitHub  /  Scholar

Mar 2026I defended my PhD on deployment-time reliability of learned robot policies; talk and dissertation available online!
Jan 2026Our work on preventing robotic jailbreaks in safety-critical scenarios was accepted to ICRA 2026!
Aug 2025Hosting a full-day workshop on Making Sense of Data in Robotics at CoRL 2025 — join the conversation!
Aug 2025Our three recent research projects – CUPID, RoboMonkey, and FORTRESS – were accepted to CoRL 2025!
Jun 2025Our work on robot data curation with influence functions (CUPID) received a best paper award at RSS RoboEval!
Jun 2025Three new papers out! CUPID (data curation), RoboMonkey (test-time scaling), and FORTRESS (OOD safety)!
Jan 2025Offer accepted! I will be joining Meta Fundamental AI Research (FAIR) as a Research Scientist Intern this summer!
Nov 2024I'm on the Summer 2025 research internship market; please reach out if my research is of interest to your group!
Nov 2024Our lab (IPRL) made the front page of the Stanford School of Engineering news for the Stanford Robotics Center!
Sep 2024I've been featured on the Stanford Daily News for my perspectives on foundation models for space applications!
Sep 2024Our work on detecting policy failures was accepted to CoRL 2024 (oral presentation at the SAFE-ROL workshop)!
Sep 2024Our work on human-robot interactive planning with large language models was accepted to CoRL 2024!
Jul 2024Our paper on real-time anomaly detection with language models received the outstanding paper award at RSS!
May 2024Two of our papers have been accepted to RSS 2024: Check out AESOP and DROID!
Jan 2024My internship research on autonomous space exploration has been accepted to IEEE AeroConf 2024!
Jan 2024Our paper on Open X-Embodiment learning has been accepted to ICRA 2024 with a best paper award nomination!
Jun 2023I'm excited to join the NASA Jet Propulsion Laboratory for a robotics research internship this summer!
Jan 2023Our work on sequencing policies for long-horizon manipulation has been accepted to ICRA 2023!

I'm broadly interested in building AI systems that reliably solve complex tasks in the physical world. My research spans learning for planning and decision making, safety and reliability, and data curation and model interpretability (see my dissertation here and defense talk here). I view these as facets of one system-level problem: toward agents that flexibly reason about and generalize to new tasks, whose behavior can be understood, safeguarded, and improved over time.

Papers and preprints are ordered by recency. Representative works are highlighted in green. Please see my Google Scholar page for a complete and up-to-date list of papers.

clean-usnob Diversity You Can Actually Measure: A Fast, Model-Free Diversity Metric for Robotics Datasets
Sreevardhan Sirigiri, Nathan Samuel de Lara, Christopher Agia, Florian Shkurti, Fabio Ramos
Preprint
arXiv

Dataset diversity drives success in robot imitation learning, but is hard to quantify over structured, variable-length, and high-dimensional trajectories. We introduce FAKTUAL, a model-free data curation algorithm that uses signature kernel-based entropy to estimate the relative diversity of individual demonstrations.

Preventing Robotic Jailbreaks via Multimodal Domain Adaptation
Francesco Marchiori, Rohan Sinha, Christopher Agia, Alexander Robey, George Pappas, Mauro Conti, Marco Pavone
ICRA 2026
arXiv / Project Site / Code

LLMs and VLMs are increasingly deployed in robotics but remain vulnerable to jailbreaking attacks that may drive physically harmful behaviors in the real world. We introduce J-DAPT, a real-time robotics jailbreak detector that adapts to prevent novel threats in data-limited regimes via multimodal domain adaptation.

CUPID: Curating Data your Robot Loves with Influence Functions
Christopher Agia, Rohan Sinha, Jingyun Yang, Rika Antonova, Marco Pavone, Haruki Nishimura, Masha Itkina, Jeannette Bohg
CoRL 2025
arXiv / Project Site / YouTube / Code / X
RSS 2025 Workshop on Robot Evaluation for the Real World ★ Best Paper ★

In robot imitation learning, policy performance is tightly coupled with the quality and composition of the demonstration data. We present CUPID, a data curation method that uses influence functions to estimate the causal impact of each demonstration on a policy's closed-loop performance.

RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
Jacky Kwok, Christopher Agia, Rohan Sinha, Matt Foutter, Shulu Li, Ion Stoica, Azalia Mirhoseini, Marco Pavone
CoRL 2025
arXiv / Project Site / Code / X

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in visuomotor control, yet ensuring their robustness in unstructured real-world environments remains a persistent challenge. In this paper, we investigate inference-time scaling as means to enhance VLA robustness and generalization.

Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
Milan Ganai, Rohan Sinha, Christopher Agia, Daniel Morton, Marco Pavone
CoRL 2025 ★ Oral Presentation ★
arXiv / Project Site

Foundation models can reason about appropriate safety interventions in hazardous out-of-distribution scenarios beyond a robot's training data. We present FORTRESS, a framework that generates and reasons about semantically safe fallback strategies in real time to prevent out-of-distribution failures.

Points2Plans: From Point Clouds to Long-Horizon Plans with Composable Relational Dynamics
Yixuan Huang, Christopher Agia, Jimmy Wu, Tucker Hermans, Jeannette Bohg
ICRA 2025
arXiv / Project Site / Code
CoRL 2024 Workshop on Learning Effective Abstractions for Planning ★ Oral Presentation ★

How can we plan to solve unseen, long-horizon tasks from a single, partial-view point cloud of the scene, and can we do so without access to long-horizon training data? Points2Plans leverages transformer-based relational dynamics to learn the symbolic and geometric effects of robot skills, then compose the skills at test time to generate a long-horizon symbolic and geometric plan.

clean-usnob Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
Christopher Agia, Rohan Sinha, Jingyun Yang, Zi-ang Cao, Rika Antonova, Marco Pavone, Jeannette Bohg
CoRL 2024
arXiv / Project Site / YouTube / Code / X
CoRL 2024 Workshop on Safe and Robust Robot Learning ★ Oral Presentation ★
RSS 2024 Workshop: Towards Safe Autonomy: Emerging Requirements, Definitions, and Methods

Robot behavior policies trained via imitation learning are prone to failure under conditions that deviate from their training data. In this work, we present Sentinel, a runtime monitor that detects unknown failures (requiring no data of failures) of generative robot policies at deployment time.

clean-usnob clean-usnob Text2Interaction: Establishing Safe and Preferable Human-Robot Interaction
Jakob Thumm, Christopher Agia, Marco Pavone, Matthias Althoff
CoRL 2024
arXiv / Project Site / YouTube / Code
CoRL 2024 Workshop on Language and Robot Learning: Language as an Interface

How can we integrate human preferences into robot plans in a zero-shot manner, i.e., without requiring tens of thousands of data points of human feedback? We propose Text2Interaction, a planning framework that invokes large language models to generate a task plan, motion preferences as Python code, and parameters of a safe controller.

Real-Time Anomaly Detection and Reactive Planning with Large Language Models
Rohan Sinha, Amine Elhafsi, Christopher Agia, Matt Foutter, Edward Schmerling, Marco Pavone
RSS 2024 ★ Outstanding Paper ★
arXiv / Project Site / YouTube / NVIDIA Media / TechXplore
CoRL 2024 Workshop on Language and Robot Learning: Language as an Interface

How can we mitigate the computational expense and latency of LLMs for real-time anomaly detection and reactive planning? We propose a two-stage reasoning framework, whereby fast a LLM embedding model flags potential observational anomalies while a slower generative LLM assesses the safety-criticality of flagged anomalies and selects a safety-preserving fallback plan.

DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
Alexander Khazatsky*, Karl Pertsch*, ..., Christopher Agia, ..., Sergey Levine, Chelsea Finn
RSS 2024
arXiv / Project Site / Code

We introduce DROID (Distributed Robot Interaction Dataset), a diverse robot manipulation dataset with 76k demonstration trajectories (or 350 hours of interaction data), collected across 564 scenes and 86 tasks by 50 data collectors in North America, Asia, and Europe over the course of 12 months. We demonstrate that training with DROID leads to policies with higher performance and improved generalization ability.

Modeling Considerations for Developing Deep Space Autonomous Spacecraft and Simulators
Christopher Agia, Guillem Casadesus Vila, Saptarshi Bandyopadhyay, David S. Bayard, Kar-Ming Cheung, Charles H. Lee, Eric Wood, Ian Aenishanslin, Steven Ardito, Lorraine Fesq, Marco Pavone, Issa A. D. Nesnas
AeroConf 2024
arXiv / Project Site / Concept of Operations (Video)

Future space exploration missions to unknown worlds will require robust reasoning, planning, and decision-making capabilities, enabled by the right choice of onboard models. In this work, we aim to understand what onboard models a spacecraft needs for fully autonomous space exploration.

Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Open X-Embodiment Collaboration
ICRA 2024 ★ Best Paper ★
arXiv / Project Site / Blogpost / Code

Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train “generalist” X-robot policy that can be adapted efficiently to new robots, tasks, and environments?

clean-usnob Text2Motion: From Natural Language Instructions to Feasible Plans
Kevin Lin*, Christopher Agia*, Toki Migimatsu, Marco Pavone, Jeannette Bohg
AR 2023
arXiv / Springer Nature / Project Site / YouTube / Invited Talk
ICRA 2023 Workshop on Pretraining for Robotics

Pretrained large language models can be readily used to obtain high-level robot plans from natural lanugage instructions, but should these plans be executed without verifying them on the geometric-level? We propose Text2Motion, a language-based planner that tests if LLM-generated plans (a) satisfy user instructions and (b) are geometric feasibility prior to executing them.

clean-usnob Semantic Anomaly Detection with Large Language Models
Amine Elhafsi, Rohan Sinha, Christopher Agia, Edward Schmerling, Issa A. D. Nesnas, Marco Pavone
AR 2023
arXiv / Springer Nature / Project Site

System-level failures are not due to failures of any individual component of the autonomy stack but system-level deficiencies in semantic reasoning. Such edge cases, dubbed semantic anomalies, are simple for a human to disentangle yet require insightful reasoning. We introduce a runtime monitor based on large language models to recognize failure-inducing semantic anomalies.

STAP: Sequencing Task-Agnostic Policies
Christopher Agia*, Toki Migimatsu*, Jiajun Wu, Jeannette Bohg
ICRA 2023
arXiv / Project Site / YouTube / Code

Solving sequential manipulation tasks requires coordinating geometric dependencies between actions. We develop a scalable framework for training skills independently, and then combine the skills at planning time to solve unseen long-horizon tasks. Planning is formulated as a maximization problem over the expected success of the skill sequence, which we demonstrate is well-approximated by the product of Q-values.

clean-usnob Taskography: Evaluating Robot Task Planning over Large 3D Scene Graphs
Christopher Agia*, Krishna Murthy Jatavallabhula*, Mohamed Khodeir, Ondrej Miksik, Mustafa Mukadam, Vibhav Vineet, Liam Paull, Florian Shkurti
CoRL 2021
arXiv / Project Site / Code

3D Scene Graphs (3DSGs) are informative abstractions of our world that unify symbolic, semantic, and metric scene representations. We present a benchmark for robot task planning over large 3DSGs and evaluate classical and learning-based planners; showing that real-time planning requires 3DSGs and planners to be jointly adapted to better exploit 3DSG hierarchies.

clean-usnob Latent Attention Augmentation for Robust Autonomous Driving Policies
Ran Cheng*, Christopher Agia*, David Meger, Florian Shkurti, Gregory Dudek
IROS 2021
PDF / IEEExplore

Pretraining visual representations for robotic reinforcement learning can improve sample efficiency and policy performance. In this paper, we take an alternate approach and propose to augment the state embeddings of a self-driving agent with attention in the latent space, accelerating the convergence of Actor-Critic algorithms.

A Sim-to-Real Pipeline for Deep Reinforcement Learning for Autonomous Robot Navigation in Cluttered Rough Terrain
Han Hu*, Kaicheng Zhang*, Aaron Hao Tan, Michael Ruan, Christopher Agia, Goldie Nejat
RA-L 2021
PDF / IEEExplore / Demo

Deep Reinforcement Learning is effective for learning robot navigation policies in rough terrain and cluttered simulated environments. In this work, we introduce a series of techniques that are applied in the policy learning phase to enhance transferability to real-world domains.

clean-usnob Lightweight Semantic-aided Localization with Spinning LiDAR Sensor
Yuan Ren, Bingbing Liu, Ran Cheng, Christopher Agia
T-IV 2021
PDF / IEEExplore

How can semantic information be leveraged to improve localization accuracy in changing environments? We present a robust LiDAR-based localization algorithm that exploits both semantic and geometric properties of the scene with an adaptive fusion strategy.

clean-usnob S3CNet: A Sparse Semantic Scene Completion Network for LiDAR Point Clouds
Ran Cheng*, Christopher Agia*, Yuan Ren, Bingbing Liu
CoRL 2020
arXiv / YouTube / Demo

Small-scale semantic reconstruction methods have had little success in large outdoor scenes as a result of exponential increases in sparsity, and a computationally expensive design. We propose a sparse convolutional network architecture based on the Minkowski Engine, achieving state-of-the-art results for semantic scene completion in 2D/3D space from LiDAR point clouds.

clean-usnob Depth Prediction for Monocular Direct Visual Odometry
Ran Cheng, Christopher Agia, David Meger, Gregory Dudek
CRV 2020
PDF / IEEExplore / YouTube

Direct methods are able to track motion with considerable long-term accuracy. However, scale inconsistent estimates arise from random or unit depth initialization. We integrate dense depth prediction with the Direct Sparse Odometry system to accelerate convergence in the windowed bundle-adjustment and promote estimates with consistent scale.

My thesis and dissertation research, ordered by recency.

clean-usnob Deployment-Time Reliability of Learned Robot Policies
Christopher Agia, Marco Pavone, Jeannette Bohg
PhD Dissertation
Department of Computer Science, Stanford University, 2026
arXiv / YouTube

This dissertation investigates how the reliability of learned robot policies can be improved at deployment time through mechanisms that operate around them. We examine several such mechanisms: (1) runtime failure detection and intervention, (2) data-centric interpretation of policy behavior and performance, and (3) composition and coordination of learned behaviors for real-world, long-horizon tasks.

clean-usnob Contextual Graph Representations for Task-Driven 3D Perception and Planning
Christopher Agia, Florian Shkurti
Undergraduate Thesis
Division of Engineering Science, University of Toronto, 2020
arXiv

This thesis tests the suitability of existing embodied AI environments for research at the intersection of robot task planning and 3D scene graphs and constructs a benchmark for empirical comparison of state-of-the-art classical planners. Furthermore, we explore the use of graph neural networks to harness invariances in the relational structure of planning domains and learn representations that afford faster planning.

Several of my research ideas have been patented alongside conference or journal publication.

clean-usnob Curating Training Data for Robot Learning based on Influence Functions
Christopher Agia, Rohan Sinha, Jingyun Yang, Rika Antonova, Marco Pavone, Haruki Nishimura, Masha Itkina, Jeannette Bohg
In preparation

Relates to methods and systems for automatically curating robotics datasets for robot imitation learning. More specifically, relates to filtering or selecting demonstration data to maximize downstream KPIs of an imitation learning system using influence functions.

clean-usnob Systems and Methods for Generating a Road Surface Semantic Segmentation Map from a Sequence of Point Clouds
Christopher Agia, Ran Cheng, Yuan Ren, Bingbing Liu
Application No. 17/676,131. Patent No. 12,008,762. U.S. Patent and Trademark Office, 2022
Google Patents

Relates to processing point clouds for autonomous driving of a vehicle. More specifically, relates to processing a sequence of point clouds to generate a birds-eye-view (BEV) image of an environment of the vehicle which includes pixels associated with road surface labels.

clean-usnob Methods and Systems for Semantic Scene Completion for Sparse 3D Data
Ran Cheng*, Christopher Agia*, Yuan Ren, Bingbing Liu
Application No. 17/492,261. Patent No. 12,079,970. U.S. Patent and Trademark Office, 2022
Google Patents

Relates to methods and systems for generating semantically completed 3D data from sparse 3D data such as point clouds.