Summary
Security-focused AI researcher Fahd Mirza examines “obliterated” open-weight models — specifically Qwen 3.8-27B Obliterated, available on Hugging Face — which are modified versions of official releases with safety training surgically removed. The video’s explicit framing is defensive: understanding the technique well enough to detect and block these models in production and enterprise GPU environments.
Mirza explains the two underlying techniques that make obliteration work. Singular Value Decomposition (SVD) identifies the internal activation direction the model uses before generating a refusal, then deletes that direction — effective but collateral-damage-prone, degrading overall output quality. Linear erasure (LEaST) is more precise, targeting only the refusal classifier signal and preserving output quality better, but it often leaves residual refusal behavior intact. The obliteration approach combines both: two separately-modified model copies are weight-averaged at roughly a 60/40 ratio, so each method’s failure modes land in different places and partially cancel. The result benchmarks within noise of the original Qwen 3.8-27B on standard tasks while complying with requests the official model would refuse.
Running on an NVIDIA A6000 (48GB VRAM, consuming approximately 25GB), Mirza serves the model via llama.cpp and tests it against a benign creative writing prompt — demonstrating identical output quality to the legitimate release. The core risk he emphasizes: these models are functionally undetectable through normal usage, behaving normally on everyday tasks until presented with a harmful prompt, making them a serious concern for shared inference infrastructure and any organization deploying community model weights.
📺 Source: Fahd Mirza · Published August 22, 2026
🏷️ Format: Deep Dive







