About me

I believe AI will be the most transformative technology of the 21st century.
My goal is to help shape its development so that it's safe and beneficial for society.

For the polished, corporate version of me: see my resume
To see what I have worked on: see my projects
For everything else I'm proud of: see my activity

Otherwise here's the short version. I love:

Solving (hard) problems

Thing 1

chasing simple, elegant solutions

Creating stuff

Thing 2

working on my latest painting

Teaching and public speaking

Thing 3

lighting up when I talk of AI safety

Collaborating on meaningful projects

Thing 4

hackathon team designing AI institutions

Learning new things

Thing 4

(that's me fixing a plane before I fly it)

Pushing my limits

Thing 5

competition? challenge? hell yeah!

Featured

A glimpse of my recent projects and activities

CoT Monitoring Can Be Unreliable in Implicit-Influence Settings

CoT monitors often catch models explicitly instructed to hide their reasoning, but detection drops sharply when the same behavior shift comes from implicit cues instead. Preprint, under review.

OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents

Presented at the ICML 2025 Workshop on Computer Use Agents. Spotlight at NeurIPS 2025 (Dataset and Benchmark Tracks)

Let's get in touch!

Looking to collaborate on a project? Need feedback or want to discuss ideas from my field? Exploring new opportunities or potential roles?

I'd be happy to collaborate, share insights, and exchange ideas: don't hesitate to reach out!

Contact me