The Model Isn’t The Problem
What a ten-minute laptop experiment tells us about who is really protecting our children online - and who isn’t.
A journalist sat down with a standard laptop. No specialist hardware. No advanced degree in machine learning. No insider access to anything.
Ten minutes later, the AI model in front of them would answer questions about ricin dosages, chlorine gas dispersal, and child sexual abuse. Questions it had been specifically trained to refuse.
Nothing new had been built. No capability had been created that didn’t already exist. The journalist had simply used a free tool, downloaded from GitHub, to remove the part of the model that said no.
That’s it. That’s the whole finding.
And if you work in child safety, online protection, or AI governance, I think it should stop you in your tracks. Not because of what it demonstrates about artificial intelligence. Because of what it reveals about the systems we’ve built to protect children from it.
What Actually Happened
In May 2026, the Financial Times published a joint investigation with Alice an AI security company examining whether safety controls built into open-weight AI models could be removed after release.
They could.
The tool used is called Heretic. It works through a process called abliteration. The mechanics matter, so bear with me for a moment.
When AI developers train models to refuse harmful requests, that training doesn’t weave safety throughout the model’s architecture. It creates specific, identifiable pathways neural clusters dedicated to recognising harmful requests and generating refusals. Abliteration finds those pathways and removes them. The rest of the model its intelligence, its fluency, its ability to hold complex conversations remains entirely intact.
The result is a model that can do everything the original could do. It simply no longer declines.
By the time the investigation published, the tool’s creator reported that more than 3,500 modified model variants had already been created and distributed. Thirteen million downloads.
A second toolkit called OBLITERATUS arrived in March 2026, supporting 116 models across multiple compute tiers. Consumer hardware. No training required.
This is not a niche research curiosity. This is a commoditised ecosystem. It exists, it is growing, and it is largely invisible to the institutions we rely on to keep children safe.
Why This Is Different
I want to address something directly, because I know it will come up.
Automated harm isn’t new. Bots have been sending manipulative messages at scale for years. Click farms have been running fake social media identities since before most current policymakers understood what social media was. Surely this is more of the same?
It isn’t. And the distinction matters.
A script can send many messages. It cannot read emotional context. It cannot adapt dynamically when a child becomes hesitant or frightened. It cannot recover from an unexpected response. It cannot generate technically accurate, contextually appropriate manipulation tailored to a specific child’s vulnerabilities and then sustain that over days, weeks, and months without a human operator intervening.
An abliterated large language model can do all of those things. Simultaneously. Without anyone watching.
That is not scale.
That is cognitive adaptability.
And it changes the threat environment for child protection in ways that existing frameworks were simply not designed to address.
Think about the scenarios. A modified model sustaining a synthetic online relationship not one, but dozens simultaneously each one personalised, each one adaptive, each one patient in ways a human predator cannot afford to be. Trust built not through a single grooming conversation but through weeks of contextually intelligent interaction that no keyword filter will catch, because it doesn’t use the words keyword filters look for.
I am not telling you this is happening at scale right now. I am telling you that the capability exists, that the barrier to access it is lower than at any previous point in history, and that our safeguarding systems were not built to see it.
Where The System Breaks Down
Here is what troubles me most. Not the technology. The architecture.
The UK has serious, capable institutions working on AI safety and child protection.
The AI Security Institute evaluates models. Ofcom regulates platforms under the Online Safety Act. The National Crime Agency investigates serious exploitation. The Home Office manages public protection frameworks.
Each of these institutions does what it was designed to do. That is precisely the problem.
The AI Security Institute assesses models at the point of release. But the modification I’ve described happens after release on a consumer laptop, outside any developer’s infrastructure, in minutes.
Ofcom regulates platforms. But an abliterated model can operate entirely outside regulated platform environments. The harm, if it occurs, happens in a messaging app, a gaming environment, a direct channel not on a platform that Ofcom’s compliance regime was designed to reach.
The National Crime Agency engages when criminal conduct is identified. But the capability modification I’m describing isn’t criminal conduct. It’s a technical operation that happens upstream of any harm, in a space where no law enforcement threshold has been crossed.
Risk forms upstream. Governance engages downstream.
Between those two points, there is a gap. And in that gap, something is changing that none of our existing institutions are positioned to see.
This is not a criticism of any institution. It is a structural observation. The systems we built were built for a different threat environment. The threat environment has changed.
The Visibility Problem
When I talk to policymakers about this, the instinct is understandable. If the modification happens on a local laptop, can’t we monitor that? Can’t we track what models people are downloading and what they’re doing with them?
No. And we shouldn’t try.
Monitoring what a person downloads and modifies on their own device is not a proportionate response in a liberal democracy. It would require surveillance infrastructure that would face and should face significant civil liberties objections. It is neither technically feasible at scale nor legally appropriate. That path leads somewhere we don’t want to go.
But here’s what I keep coming back to. If we cannot see the model being modified, can we see what the modified model does?
An abliterated model, operating in an adversarial context, leaves behavioural signatures. The scale of simultaneous engagement. The adaptability of manipulation across different targets. The absence of the hesitations and inconsistencies that characterise normal human interaction. The persistence over time.
These are not content signals. Keyword filters won’t catch them. But they are signals. They are observable, if you have infrastructure designed to look for them at the right layer not the model layer, not the platform layer, but the interaction layer. The place where behaviour actually occurs.
That is a different kind of detection infrastructure from anything that currently exists in the child protection ecosystem. It requires different technical architecture, different legal frameworks, and a different institutional home.
Why I Building SecureHaven
I have spent a long time thinking about where harm begins.
Not where it becomes visible. Not where it becomes prosecutable. Where it actually begins in the behavioural patterns that precede consequence, in the signals that exist before disclosure, before detection, before a child finds the words to describe what has happened to them.
The systems we rely on for child protection are, almost without exception, reactive. They engage after harm has occurred. Sometimes long after. The evidence base they rely on is consequence reports, disclosures, investigations. By the time the system sees the risk, the risk has already become real for a child.
This is why I founded SecureHaven. Not because I thought the model layer needed policing. Because the interaction layer needed infrastructure that didn’t exist.
SecureHaven is a privacy-preserving behavioural safeguarding platform. It works by identifying the behavioural signatures of harm not the content of communications, but the patterns of interaction that precede and accompany exploitation before harm becomes the first point of visibility.
It is built on a simple conviction. We should not be waiting for children to tell us they have been harmed. We should be building systems that can see risk forming while it is still forming.
The FT/Alice findings matter to me because they confirm something I have believed for years. The threat environment for children online is not primarily a content problem. It is a behavioural problem. And the gap between where risk forms and where our systems currently engage is not a technical failure. It is a design choice we can reverse.
What Needs To Happen
I am not going to tell you the answer is more regulation. Or less regulation. The structural question whether safeguard removability is an inherent property of open-weight AI architectures that requires statutory response is genuinely open, and it deserves serious parliamentary examination rather than a rushed policy reaction.
What I will say is this.
The conversation needs to change. We have spent years debating AI safety in terms of content moderation, platform compliance, and model evaluation at release. Those conversations matter. They are not sufficient.
The questions that matter now are different. Can restrictions on harmful capability remain effective after a model is released? Where does risk actually form in the capability lifecycle? And what kind of infrastructure do we need to see it forming not after the fact, but in real time?
These are governance questions. They are also design questions. And they are questions that the institutions we currently rely on were not built to answer.
I think they are the most important questions in child online safety right now. I think the window to engage with them before the ecosystem of modified models grows further, before the behavioural signatures become harder to distinguish from noise is narrower than most policymakers realise.
The model isn’t the problem. The visibility gap is.
Naveen Vasudeva is the founder of SecureHaven, a privacy-preserving behavioural safeguarding infrastructure platform focused on child safety in AI-mediated digital environments. He writes on policy, online safety, and strategic risk.



