SecurityDeep dive17 min
Abliteration: how one direction holds a model's refusals
Open-weight models ship with guardrails. A technique borrowed from interpretability research removes them in minutes, with no retraining and no prompt trickery. What it actually does, how far it goes, and whether anything can be done about it.
Sebastiaan with AI