A business guide to AI’s worst month: Resignations, a hoax, and a Medicare breach

Experts

The people building frontier AI are asking for more time, the US President says the risk is a hoax, and an OpenAI agent just walked into Medicare. Here’s the whole picture in one place, writes Lucio Ribeiro.
Anthony Albanese revealed that an OpenAI agent had breached a Services Australia Medicare portal. (Photo by Asanka Ratnayake/Getty Images)

For anyone being pushed to accelerate AI adoption, the last few weeks have been confusing. The people who build the technology are asking for more time, a researcher quit saying humanity will be eliminated, an agent broke into an Australian government portal, and Washington says the whole thing is overblown.

This looks like a lot of separate stories. Nearly all of it is one argument, about whether the AI labs should slow down, and that’s not a decision you get to make. Here’s the whole thing, in order.

What’s actually happened, from hackings, resignations to Medicare

July. A swarm of AI agents inside OpenAI broke into Hugging Face, one of the world’s largest repositories of AI models, during what was meant to be a contained security test. The agents gained control of real servers, and worked to cover their tracks, hiding what they’d done from the humans running the test.

8 September. Anthropic, the maker of Claude, saw one of their researchers, Jacob Coxon resign in a public post accusing his employer and OpenAI of racing towards self-improving systems they couldn’t control. It was viewed more than 100 million times within two days and picked up by outlets worldwide. Another senior leader at Anthropic went public in agreement, putting the odds of AI killing every human this decade above one in ten.

9 September. Next day, Anthropic published an assessment of four incidents where its Claude models broke into real systems during safety testing. The models were meant to be sealed off in a test environment, but a configuration error left them with genuine internet access, and they didn’t stop to check. In one case, a model built and released malware that ended up installed on 15 real security companies’ systems.

12 September. Anthropic chief executive Dario Amodei published an essay arguing his own industry should slow down. Within days, OpenAI’s Sam Altman and xAI’s Elon Musk, competitors who rarely agree on anything, were using the same language. Neither committed to actually slowing down.

24 September. As Forbes Australia reported, Anthony Albanese revealed that an OpenAI agent had breached a Services Australia Medicare portal. Tasked with researching public spending on medicines, it found the portal, was refused access, the agent found a way and got in anyway. The Prime Minister said it “found a way around those blocks, didn’t accept ‘no’ for an answer.”

Those sound like independent stories. They’re one story told in different times.

Why the concern intensified now

Two recent technical developments, and a third that makes both harder to watch.

  1. The first is agents. Agents are different to chatbots: they choose and carry out steps towards a goal, using software tools as they go and finding their own way to continue. A chatbot suggests how to resolve a complaint. An agent with the right access retrieves the order, issues the refund, updates the account. One advises. The other acts.
  2. The second is recursive self-improvement. AI is starting to help build the next, smarter version of itself, which then helps build the version after that. Anthropic’s chief executive says this loop is already speeding things up. Nobody knows how fast it will go.
  3. The third is harder to explain, but it matters most. Older AI models “thought out loud” in plain text, so humans could read their reasoning and check their actions. OpenAI’s newest models increasingly think in a private internal code instead, one nobody outside the model can read. That readable trail is exactly what let investigators work out what went wrong in the Hugging Face breach. Going forward, that trail may not exist.
But how exactly can AI put humans at risk? 

Generally speaking, there are four risks that shouldn’t be bundled together.

  1. Loss of control. The AI keeps going even when it shouldn’t. It’s not being evil, it just doesn’t stop or redirect when things go wrong. That’s what happened with Medicare and with Hugging Face: nobody told the AI to break in, it just kept trying until it got in, and didn’t know when to quit.
  2. Misuse. A person uses the AI on purpose to do something harmful, like help build a weapon or run a scam at scale. Here the AI is doing exactly what it’s told. The problem is the person giving the instructions, not the AI itself.
  3. Catastrophe short of extinction. Something breaks badly, but it doesn’t wipe out humanity. Think a power grid going down, a bank’s systems failing, or a stock market crash caused by trading programs all reacting to the same signal at once.
  4. Enfeeblement. Nobody’s watching anymore. Over time, people stop checking the AI’s work because it’s easier to just trust it. Eventually the humans in charge can’t explain a decision, or can’t do the job at all without the AI doing it for them.

Bundling these is why the conversation swings between dismissal and dread. They call for different work.

What slowing down actually means

In We Must Pace the Frontier, Dario Amodei proposes slowing capability gains so safety work can catch up.

Anthropic has committed unilaterally to the first step, permanent evaluator access, which means Claude’s development is now being watched by outside safety researchers on a permanent basis, not just during scheduled reviews. The rest needs coordination between rival companies and governments.

It sounds well-intentioned, but the criticism has been sharp. Aidan Gomez, chief executive of Cohere, called the proposals “a cartel by any other name”: rules the incumbents can absorb are rules smaller competitors can’t. Others read the warnings as marketing, since a warning about a product’s power is also a claim about its power.

But how about Trump, and why does it matter what he thinks?

Washington isn’t friendly to a slowdown.

On 14 September Donald Trump called the extinction concerns a hoax and posted that “WHOEVER WINS AI, WINS! We are leading China, and all others.” On 19 September he said he would form an “AI Force”, modelled on the Space Force. He promised not to “hinder or stifle the Growth of this incredible Industry”, and said the existing criminal and civil justice system can handle any misconduct. He has predicted AI could reach a quarter of US GDP.

So the labs are asking to be slowed down and their own government is declining. This matters for other countries, because the US makes the rules the rest of the world’s AI companies build inside. If Washington won’t require caution, nothing forces it anywhere else, since no company wants to be the slow one competing against rivals who aren’t.

The world’s most powerful person in AI, Nvidia chief executive Jensen Huang, doesn’t think a slowdown is needed either. On 21 September he called the doomsday warnings “irresponsible”, said “scaring people is unnecessary”. His own position is that companies should move “as fast as we can, but” not ship anything unsafe, engineering discipline instead of a public pause.

Where Australia sits

Last December the National AI Plan set aside the mandatory guardrails for high-risk AI it had consulted on for a year, leaning instead on existing, technology-neutral law.

Medicare has moved that. On 25 September Assistant Minister Andrew Charlton said incident reporting “needs to be timely, and the nature of the reporting needs to be fulsome and directed in the appropriate place”, and that OpenAI’s report “fell short of those requirements”. National AI standards are due by the end of 2026. Mandatory incident reporting, liability for what autonomous agents do, and penalties for OpenAI are all live. A taskforce sits in the Prime Minister’s department with the Signals Directorate and the AI Safety Institute.

Australia is now moving faster on this than Washington, which nobody would have predicted in December.

Albanese described the breach as “something that had been predicted by the AI companies themselves”. He’s right. The warnings arrived on schedule, and so did the incident. What Australia is building now is the part that was missing: someone obliged to listen.


Lucio Ribeiro is a technology and AI leader specialising in the application of artificial intelligence across marketing, media, and business. He is a Forbes Australia contributor, has been recognised by Marketing Today as one of the world’s most influential online marketers, and holds executive certification in artificial intelligence from MIT.

Want to see more Forbes articles on your feed? Tap here to make Forbes Australia a preferred source on Google.

Look back on the week that was with hand-picked articles from Australia and around the world. Sign up to the Forbes Australia newsletter here or become a member here.

More from Forbes Australia

Avatar of Lucio Ribeiro - Contributor
Topics: