Open Safety Institute Contact

AI safety, through security

When agents act,
safety is security.

As AI systems begin to act on the world, the risks show up first as security failures. We study how agents attack, defend the infrastructure they run on, and publish the work.

01 / Agents

Agents as adversaries

How autonomous systems attack, and misbehave, at machine speed — treating the agent itself as the threat, not only its inputs.

02 / Defence

Defending the stack

Detecting and investigating attacks that move faster, and louder, than defenders can follow.

03 / In the open

In the open

Evaluations, detections, data and tools, published so anyone can check the work and build on it.

Open by default Reproducible or it doesn't count Independent of the systems we test

Building or defending against agents? Let's talk.