status — active·focus — dangerous-capability evaluations·based — Pune, India
I study how language models behave when they think no one is looking.
I'm Allan Suresh, a research engineer working at the intersection of backend systems and AI safety evaluations — building the harnesses, honeypots, and benchmarks that reveal how models act under monitored and unmonitored conditions.

Current focus
evaluations
Designing controlled experiments on instrumental reasoning and shortcut-taking — monitored vs. unmonitored framing, run with Inspect AI.
benchmarking
Studying sandbagging and elicitation gaps in frontier models across GPQA Diamond, ARC-Easy, and CyBench.
background
AISI Fellowship at Georgia Tech, SERI MATS under John Wentworth, and a GCRI governance fellowship with Seth Baum.
Featured work
All projects →status — active
Monitored vs. Unmonitored: A Honeypot Study of Instrumental Reasoning
A controlled experiment testing whether models take more shortcuts when they believe grading is unmonitored, using a fake grading server as a honeypot.
Inspect AIFlaskDockerClaude Sonnet 5
status — complete
CyBench Elicitation Ablation
A 2×2×2 ablation across 2 models, 2 guidance levels, and 2 tool configs on 9 CyBench tasks — the real finding was that budget exhaustion, not incorrect reasoning, explained almost every failure.
Inspect AICyBenchDocker
status — complete
Sandbagging Detection via Position Attractors
A 2×2×2 pilot extending Cacioli's 2026 position-attractor work to a frontier model, testing whether reasoning-permission changes whether sandbagging leaves a detectable signature.
Inspect AIGPQA DiamondARC-EasySciPy
Recent writing
All posts →- What Elicitation Gaps HideMay 20, 2026
Running the same CyBench tasks across a 2×2×2 condition matrix, the biggest driver of apparent capability wasn't guidance or tool access — it was whether the agent ran out of tokens.
- A Position Attractor, a False Start, and What Elicitation Mode Was HidingMar 2, 2026
An early p=0.0074 result didn't survive a second look — the bug that caused it, and the corrected 2×2×2 design that found something more specific underneath, are more useful than the original number would have been.