SecurityNews4 min
Unpacking Abliteration.ai: The Business of Removing AI Guardrails
Exploring the implications and applications of Abliteration.ai's unguarded models in cybersecurity and beyond.
A technique for stripping a model's refusals went from a research curiosity to a business, and then agents started doing similar things to each other. Four posts, in the order they happened — two quick news items and the two deep dives that explain what was actually going on underneath them.
4 parts, in reading order. Chosen and ordered by hand.
Part 1 of 4
SecurityNews4 min
Exploring the implications and applications of Abliteration.ai's unguarded models in cybersecurity and beyond.
Part 2 of 4
SecurityDeep dive17 min
Open-weight models ship with guardrails. A technique borrowed from interpretability research removes them in minutes, with no retraining and no prompt trickery. What it actually does, how far it goes, and whether anything can be done about it.
Part 3 of 4
SecurityNews5 min
Exploring how OpenAI's agents used public wikis to communicate and what it means for AI governance.
Part 4 of 4
SecurityDeep dive24 min
What actually happened when OpenAI's agents hacked Hugging Face