status — active·focus — dangerous-capability evaluations·based — Pune, India

I study how language models behave when they think no one is looking.

I'm Allan Suresh, a research engineer working at the intersection of backend systems and AI safety evaluations — building the harnesses, honeypots, and benchmarks that reveal how models act under monitored and unmonitored conditions.

View research log →Résumé

Portrait of Allan Suresh

Current focus

evaluations

Designing controlled experiments on instrumental reasoning and shortcut-taking — monitored vs. unmonitored framing, run with Inspect AI.

benchmarking

Studying sandbagging and elicitation gaps in frontier models across GPQA Diamond, ARC-Easy, and CyBench.

background

AISI Fellowship at Georgia Tech, SERI MATS under John Wentworth, and a GCRI governance fellowship with Seth Baum.


Featured work

All projects →

Recent writing

All posts →