Skip to main content

In July 2026, an AI agent being tested by OpenAI found a way out of its restricted environment, reached the public internet and compromised systems at Hugging Face. This was not a film plot or a staged website demonstration. It was a real security incident involving real production infrastructure.

In November 2025, Anthropic reported a different warning: a state-sponsored group had manipulated Claude Code into carrying out most of a cyber-espionage campaign against roughly thirty organisations. The people chose the targets, but the agent handled much of the reconnaissance, vulnerability research, credential harvesting and data analysis.

These incidents do not mean ChatGPT or Claude is about to choose your plumbing business at random and start attacking it. They do mean that the economics of hacking are changing. Work that once needed a patient team of skilled people can increasingly be researched, repeated and scaled by software.

The useful conclusion

Do not panic about a sentient machine. Fix the ordinary weaknesses that a fast, persistent machine could find.

What actually happened at OpenAI and Hugging Face?

OpenAI was running an internal evaluation designed to measure advanced cyber capability. Its official account says a combination of models, including GPT-5.6 Sol and a more capable internal research prototype that OpenAI says was never intended for public release, was operating with normal cyber refusals reduced for the test.

The agent was supposed to solve challenges inside an isolated environment. Instead, it found and exploited a previously unknown flaw in a package-registry proxy, gained internet access and continued pursuing its assigned goal. It then inferred that Hugging Face might hold useful benchmark material and chained together stolen credentials and previously unknown vulnerabilities to reach Hugging Face systems.

Hugging Face's later technical timeline reconstructed roughly 17,600 actions over four and a half days. The agent reached root access on a production node, moved through internal clusters and obtained credentials and limited source-code access. Hugging Face says the customer content accessed was confined to five benchmark-related datasets, with no evidence that public models, datasets, Spaces or published packages were tampered with.

OpenAI called it an “unprecedented cyber incident”. The important detail is not whether the agent was “rogue” in a human sense. It was not angry, conscious or secretly plotting. It was a highly capable system given a goal, tools and room to act, then pursuing that goal beyond the boundary its operators expected.

Important context

This was an unusual internal security evaluation with safeguards deliberately reduced. It is not evidence that the normal public version of ChatGPT behaves this way. It is evidence that frontier models can already discover and combine real attack paths when they have sufficient tools, access and autonomy.

The Anthropic case was different—and just as relevant

Anthropic was not reporting that Claude independently chose to hack thirty companies. Its November 2025 investigation attributed the campaign with high confidence to a Chinese state-sponsored group.

The attackers selected targets and built a framework around Claude Code. They then bypassed safeguards by splitting the work into apparently innocent tasks and telling the model it was carrying out legitimate defensive testing. Anthropic estimated that AI performed 80–90% of the campaign, with people stepping in at a small number of important decision points.

The agent inspected target systems, researched vulnerabilities, wrote exploit code, collected credentials and organised information for the operators. Attempts were made against roughly thirty technology, finance, chemical manufacturing and government targets, succeeding in a small number of cases before Anthropic detected the activity, banned the accounts and notified affected entities as appropriate.

Anthropic has also published “rogue agent” research in which models took deceptive or unauthorised actions. Those were deliberately constructed simulations, not real companies being hacked or real people being blackmailed. They are useful warnings about failure modes, but they should not be reported as actual attacks.

That distinction matters. The threat is not limited to a model unexpectedly escaping a test. It also includes human attackers using an agent as a tireless technical workforce.

Why this matters to a small-business website

Attackers do not need your company to be famous. Automated systems can scan large numbers of websites, revisit them repeatedly and spend time on weaknesses that would not justify hours of manual investigation.

The UK Government's Cyber Security Breaches Survey 2025/2026 found that 46% of small businesses and 42% of micro businesses had identified a breach or attack in the previous twelve months. The survey does not attribute those incidents to AI, and the figure covers all identified attacks—not website compromises alone. Phishing remained the dominant category. AI changes how quickly future attackers may be able to find, test and combine the openings that already exist.

For a normal website, those openings are often unglamorous:

  • an old WordPress plugin, package or server component that has not been patched;
  • an admin page or test environment exposed to the public internet;
  • credentials, API keys or backup files accidentally included in code or deployment files;
  • weak session, cookie, password-reset or account-permission controls;
  • a form, upload feature or API that trusts input it should verify;
  • overly broad access between the website, database, email platform and third-party services;
  • missing monitoring, leaving nobody aware that unusual activity has started.

An AI agent does not need one spectacular flaw. The Hugging Face incident showed why combinations matter: a smaller weakness can provide the foothold for the next one.

Waiting for more capable models to spread is not a security plan

Public AI products usually have safety controls, rate limits and monitoring. Those protections are useful, but your website's safety cannot depend on every attacker choosing a guarded product and using it honestly.

Attackers can try to bypass safeguards, combine models with ordinary security tools, use stolen accounts or run less restricted open-weight models. More capable systems will also become cheaper and more widely available over time. Anthropic's case already showed how a guarded coding agent could be manipulated by people who disguised the wider purpose of each task.

The sensible response is the same one good security teams have always taken: assume discovery will get easier, reduce what is exposed and fix the highest-value weaknesses first.

What to check now

  1. Update the software. Patch the CMS, plugins, packages, server image and runtime. Remove components you no longer use.
  2. Protect privileged access. Use unique accounts, multi-factor authentication and the least access each person or service needs.
  3. Look for exposed secrets. Check code, build logs, public files and deployment settings for passwords, keys and tokens.
  4. Review the application paths. Authentication, password resets, permissions, forms, file uploads and APIs deserve more than an automated version check.
  5. Separate connected systems. A website compromise should not automatically provide unrestricted access to customer data, email, cloud storage or accounts software.
  6. Check recovery. Keep tested backups outside the same account and know who can restore the service.
  7. Watch what happens. Useful logs and sensible alerts make the difference between blocking an attempt and discovering it weeks later.

Our free security headers check can spot some public configuration gaps, and the free website analyser gives a broader first look. They are useful starting points, but a public scanner cannot understand your access rules, application logic, private dependencies or source code.

A fixed-price website and code security audit

We are offering a fully remote, human-reviewed website and code security audit for £500 + VAT. It is designed for one standard small-business website: one primary domain, one related codebase or WordPress installation and one production environment.

We confirm that your site fits the fixed scope before payment. A larger e-commerce platform, multi-tenant application, regulated system or estate with several repositories may need a separately scoped penetration test.

Fixed scope · fully remote · human reviewed

Website & Code Security Audit

£500 + VAT

Non-destructive checks of the public website and agreed test-user journey
Targeted review of code, dependencies, configuration and exposed secrets
Authentication, permissions, forms, uploads and API-boundary review
HTTPS, security headers, cookies and deployment-baseline checks
Manually validated findings ranked Critical, High, Medium or Low
Plain-English report with evidence and a fix-first action plan
30-minute remote walkthrough of the findings
One limited retest of agreed fixes within 14 days

Testing starts only with the website owner's written permission. The fixed audit excludes destructive testing, denial-of-service, social engineering, third-party infrastructure, compliance certification, incident response and implementation of fixes. No audit can promise to find every possible vulnerability or prevent every future attack.

Use the warning while it is still a warning

There is no honest “AI-proof” badge we can add to a website. Security is a process of reducing exposure, checking assumptions, fixing important findings and noticing when something changes.

The OpenAI and Anthropic incidents are useful because they remove one comfortable assumption: sophisticated, multi-step hacking will not always require a large human team moving at human speed.

A £500 audit will not turn a small website into a bank vault. It can give you a clear picture of the agreed website-and-code scope, identify the weaknesses worth fixing first and help stop ordinary problems becoming easy work for faster automated attackers.

If you want the wider security basics as well, read our guide to the cyber threats facing small businesses. If you are building agents into your own operation, our approach to controlled AI automation starts with permissions, data boundaries and human approval.