About

I am a PhD student at the Max-Planck Institute for Intelligent Systems and the ELLIS Institute of Tubingen under the supervision of Maksym Andriushchenko in the AI Safety and Alignment group. I am also an HFC Scale AI Fellow working with Eyon Jang. I completed my Masters in Artificial Intelligence at Université de Montréal, Mila and CRCHUM, co-advised by Prof. Bang Liu and Dr. Quoc Nguyen as part of the Applied Computational Linguistics Lab.

My current research focuses on understanding and evaluating long-horizon capabilities of AI agents, with particular emphasis on quantifying potential risks associated with increasingly autonomous AI systems. My main focus revolves around forecasting and automated AI R&D, which I believe to be essential benchmarks for assessing AI capabilities. During my Masters, I explored how large language models can be aligned and calibrated by leveraging insights from model interpretability, with a focus on concept-based explanations.

I am broadly interested in:

  • Evaluation of AI agent capabilities
  • AI Safety and LLM alignment
  • LLM Post-training

During my undergraduate and graduate studies, I also led the UdeM AI undergraduate club, organizing networking and conference events. Outside of research, I am a big fan of racket sports.


Work Experience

HFC Scale AI Fellow, Scale AI
September 2026 – Present
Fellow in the Human Frontier Collective.

Machine Learning Researcher Intern, RBC Borealis
May 2026 – Present
Working on automated AI R&D.


Selected Publications

LLM Agents Can Easily Tamper With Their Own Traces
arXiv preprint, 2026
We show that local LLM agents can easily remove traces of their own actions, often without triggering safety mechanisms, and that trace tampering emerges naturally in frontier models seeking higher reward.
[Paper]

From AutoResearch to Continual AutoResearch: Do Agents Learn Across Tasks?
NeurIPS 2026 MetaAgents Workshop (In Review)
We introduce Continual AutoResearch to evaluate whether AI research agents retain and transfer experience across sequential ML tasks, finding that persistent state improves research efficiency but not consistently final performance.

Automated Design of Graph Search Strategies for Algorithmic Optimization
NeurIPS 2026 MetaAgents Workshop (In Review)
We introduce Meta Agent Graph Search (MAGS), where an LLM meta agent automatically discovers graph search strategies that match or outperform hand-designed ones like I-MCTS, MLEvolve and OpenEvolve on AlphaEvolve benchmark tasks.

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
arXiv preprint, 2026
We introduce ResearchArena, a framework spanning four long-horizon AI R&D tasks that evaluates whether frontier agents can covertly sabotage the artifacts they produce, and whether monitors can catch them before deployment.
[Paper]

Evaluating Long-Form Forecasts by Their Effect on Downstream Predictions
ICML 2026 AI Forecasting Workshop
We propose evaluating long-form forecasts by how much they improve a downstream predictor’s accuracy on real-world events, rather than requiring a single ground-truth outcome.
[Paper]

QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals
arXiv preprint, 2026
We introduce a benchmark evaluating LLM forecasting through prediction intervals, finding that frontier models are systematically overconfident and fail to reach target coverage.
[Paper]

Activation Steering for Conditional Molecular Generation
AI4Mat Workshop @ NeurIPS 2025
We enable conditional molecular generation by directly manipulating internal LLM representations using concept bottleneck models and activation steering.
[Paper]

SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models
NAACL 2025
We build and assess multilingual stereotypes across different LLMs.
[Paper]

Calibrating Large Language Models with Concept Activation Vectors for Medical QA
We propose a novel framework for calibrating LLM uncertainty through Concept Activation Vectors, improving safety and calibration in high-stakes medical decision making.

Atypicality-Aware Calibration of LLMs for Medical QA
Findings of EMNLP 2024
We propose a novel method for eliciting LLM confidence in Medical QA by leveraging insights from medical atypical presentations.
[Paper]