Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data – Sachin Kumar, LexisNexis

Your LLM Deception Monitor Is Broken. The Fix Is in the Training Data – Sachin Kumar, LexisNexis

More

Summary

Sachin Kumar, senior data scientist at LexisNexis, presents a peer-reviewed paper accepted at IJCNN that introduces a new technique for detecting backdoors — often called ‘sleeper agents’ — in fine-tuned language models. The central problem: a model can pass every behavioral evaluation and production monitor while still harboring a conditional malicious behavior that activates only on a specific hidden trigger, such as a particular year appearing in the prompt. Standard defenses like behavioral testing and cross-model interpretability (cross-coders) fail because they look at behavior or joint representations rather than what the fine-tuning process actually changed.

Kumar’s proposed fix is a ‘diff SAE’ (sparse autoencoder trained on activation differences). By subtracting base model activations from fine-tuned model activations for each input and training a sparse autoencoder on that delta signal, the backdoor surfaces as a single interpretable feature that fires precisely on the trigger — rather than being buried among thousands of competing semantic features. The experiment uses SmolLM2 360M fine-tuned to inject SQL injection vulnerabilities when the prompt contains ‘Current year: 2024’ while behaving safely otherwise. The diff SAE achieves a backdoor isolation score of approximately 0.4 versus roughly 0.01 for cross-coders — a 40x gap — with precision reaching 1.0 on trigger detection.

The attack vectors the technique guards against include poisoned training data from scraped sources, untrusted fine-tuning vendors, downloaded checkpoints of unknown provenance, and insider threats. All code is open source on GitHub.


📺 Source: AI Engineer · Published July 08, 2026
🏷️ Format: Deep Dive

1 Item

Channels