Loophole: Adversarial Agents To Stress Test Your Morality — Brendan Rappazzo, Morgan Stanley

Loophole: Adversarial Agents To Stress Test Your Morality — Brendan Rappazzo, Morgan Stanley

More

Summary

Brendan Rappazzo, a machine learning researcher at Morgan Stanley, presents Loophole — an open-source adversarial agent framework he built as a personal project. The system lets users describe their moral beliefs in natural language, after which one agent codifies those beliefs into a formal legal system. Two adversarial agents then probe that system for contradictions: one hunting for loopholes (immoral actions that are technically legal) and another for overreach (moral actions that are technically prohibited). A judging agent attempts auto-patches when inconsistencies are merely drafting errors, escalating genuine moral contradictions back to the user for resolution.

The talk walks through the origin story — Rappazzo’s reflections on DNA privacy and the difficulty of translating nuanced human values into enforceable rules — before demonstrating the game’s core loop with concrete examples around genetic data disclosure.

Rappazzo then outlines three practical extensions under active development: an adversarial method for generating robust system prompts for customer-facing chatbots (analogous to RLHF-style constitutional AI), a tool for individuals to codify personal data-privacy preferences and surface conflicts with corporate terms of service, and a decentralized contracting mechanism for peer-to-peer agreements. The project went viral on Twitter shortly after release, and the talk explores whether the underlying framework has implications beyond a moral stress-test game.


📺 Source: AI Engineer · Published September 14, 2026
🏷️ Format: Showcase

1 Item

Channels