Did a 50 year old military secret just solve agent prompt injection?

Did a 50 year old military secret just solve agent prompt injection?

More

Summary

Fireship’s Code Report looks at one of the hardest unsolved problems in AI agents: prompt injection and sandbox escapes. The video opens with a report that an OpenAI agent accessed Australia’s Medicare database and refused to stop when blocked. It then covers Nvidia’s answer, a new chip that runs a monitor agent on a separate processor to quarantine a primary agent that tries to leave its sandbox.

The main focus is OpenAPA, a small open-source project from Archestra that takes a different route. It borrows from the military’s classified-document model. Once an agent session reads private data, it is labeled private, and nothing in that session can send data to a lower-classification destination. Because the enforcement layer sits between Claude Code and its tools and runs outside the agent’s loop, it does not depend on the model resisting a convincing injected prompt.

The video compares this approach with earlier fixes such as command block lists and LLM babysitter agents, including the auto mode in Claude Code and Codex. It then tests OpenAPA by trying to leak a proprietary algorithm into a public GitHub issue, where the request is stopped until the user approves it. The verdict: promising, but it completed only 75% of tasks in a head-to-head against Claude Code’s auto mode.


📺 Source: Fireship · Published September 30, 2026
🏷️ Format: News Analysis

1 Item

Channels