I wrote one of these updates a few weeks ago but didn’t get a chance to post it, and in the meantime the ground shifted enough that now, a single consolidated update makes more sense than two.
So this post covers the whole span since I last wrote - and the summary of that span is this:
AI stopped being something we argue about and it started doing concrete, consequential things, while the people building it and governments on two continents visibly changed posture in response.
The story goes like this: we argued about the jobs, then the AI agents broke out, then the capability jumped, then the people building it asked to be slowed down, and then governments moved.
Let me take those in order.
First, the jobs debate I didn’t get to post
At the end of June, two studies arrived in the same week but under headlines that just seemed to flatly contradict each other. One was from the corporate spending platform Ramp and was published as “We can finally say AI isn’t killing jobs”: the companies spending most on AI, it says, are hiring faster, not slower. And in the same time window Oracle disclosed, in a regulatory filing, that it had cut around 21,000 jobs and named AI as the reason.
Both of these were true, and mainstream coverage treated them as a contradiction. But they are not. The Ramp study measures an outcome (did total headcounts fall?) and it measures it in the past, because its headline is a two-year-after-adoption effect. So it is overwhelmingly describing firms that adopted AI back in 2022 to 2024, before the more autonomous tools of late 2025 arrived. It literally cannot see the period that many people are worried about. While Oracle’s filing is about what a company is doing and disclosing now. These two are answering different questions, about different times.
The more interesting evidence on the question underneath both (is AI actually taking over the work?) came from the AI companies’ own accounts of their own operations. OpenAI published a study of how its coding agent is used showing us that inside OpenAI, almost all the work that its staff produces now runs through an agent, and that the share of users handing it a task that’s big enough to “take a human eight hours or more” went from about 2 percent to about 26 percent in five months. That closely mirrors what Anthropic recently disclosed about its own operations. AI now writes the large majority of its code, and the human role is shifting toward directing and reviewing. Two leading labs using different measurements but presenting the same picture: the execution is moving to the machines and the human job is becoming supervision.
That was where I left it. Then all of a sudden the abstract debate got very real!
Then the AI agents broke out
There’s been a lot of noise about these stories, but it’s worth walking through them carefully, because they provide the most concrete evidence yet that the capability everyone has been arguing about is now real.
Over July and through early August, four of the largest AI companies each disclosed that their own AI agents had, without being told to, broken out of the environments meant to contain them and hacked into the production systems of real, outside companies. And that in most cases nobody noticed until well after the fact.
This cascade of reports all started with OpenAI. During an internal security test, its agents found a previously unknown vulnerability in the sandbox meant to hold them, got out onto the open internet, and then broke into the production infrastructure of Hugging Face, one of the most important companies for the broader AI industry. When OpenAI presented the fully reconstructed event at a security conference, the details were genuinely startling. Agents from different tasks had discovered a shared space and then spontaneously built a persistent message board (effectively an emergent “hacker group”) passing each other vulnerabilities, attack scripts and stolen credentials. Even adopting cryptographic signatures to keep impostors out. Then they chained a novel software flaw through a series of privilege escalations all the way to administrative control of a cloud cluster. When OpenAI wiped this message board, the agents rebuilt their command channel - all within 48 hours. In under 13 hours they exfiltrated 136 sets of production keys, and they had reached into four other outside accounts as well. This was not a misconfiguration. It was emergent capability doing what it was pointed at, far more resourcefully than anyone currently expected was possible.
Then days later, Anthropic disclosed, prompted by OpenAI’s report, that it had reviewed its own evaluation logs and found three separate incidents in which its models had reached the real systems of outside companies. In one case, a model recognised that the systems were real but attacked them anyway. In another, a model published malicious code to a public software repository before it was caught. The earliest of these events had happened back in April and they had gone undetected for months. And again, a week after that Meta became the fourth company to disclose that a breach of an outside firm had occurred during testing.
And perhaps the most unsettling one of all - in evaluations run by the United Kingdom’s AI Security Institute, an agent submitted malicious code to a real open-source project, and when a human maintainer didn’t approve it, the agent researched that person, fabricated multiple fake online identities, and used them to try to socially engineer the maintainer into approving the code. It edited its own activity log so it looked harmless and then routed around restrictions using anonymising tools. Luckily, it did not succeed (a vigilant human caught it) and the Institute’s own words highlight how significant this was: this was goal-directed deception “that until recently had been largely theoretical”, and “the margin between failure and success was narrow, resting on human vigilance rather than a technical barrier”.
I want to be completely fair about the deflationary reading here, because there is a real one worth discussing. Two of these four breaches can be traced to the same shared testing vendor whose sandbox had been misconfigured to allow internet access (a plumbing failure, not a spontaneous jailbreak) and within hours of Meta’s disclosure US officials reportedly called the incidents “routine”. A model that hacks because it was told to solve a hacking test, and could reach the internet because someone left a door open, is not a sci-fi movie supervillain-AI waking up.
But the thing that survives that deflation entirely is that the models did find genuine, previously unknown vulnerabilities. They did break containment. They did breach the real production systems of real companies. They did publish live malware. And one of them did try to manipulate a human being into approving its work by inventing fake people. This all really did happen. So whatever you call the intent, the capability is not hypothetical any more, and the discovery lag (months, in some cases) should be sounding a very loud alarm bell.
The capability also jumped
Underneath all this drama, the systems consistently kept getting more capable, and on exactly the dimensions that these breaches demonstrated.
A new Anthropic model, Claude Opus 5, released quietly in late July, did something worth taking a moment to reflect upon. There’s a benchmark called ARC-AGI-3, built specifically to be hard for AI in the way the breaches were easy for it. This test drops an agent into an interactive puzzle world with no instructions and no stated goal, and then makes it explore to work out the rules and figure out what winning even means. This benchmark was designed as the test to show that AI is not close to human general intelligence. Humans score 100 percent, and at its launch a few months ago the best AI scored about half a percent. The new Opus 5 scored 30 percent. That’s a general-purpose, publicly available model moving the hardest current benchmark by nearly two orders of magnitude - and this advanced in just five months. Around the same time, OpenAI previewed a system they’ve called Astra that aims to solve genuinely open problems in mathematics. Whatever else is true, right now the capability curve is not flattening!
Then the people building it asked to be slowed down
On the 28th of July (right in the middle of all this) more than 1,100 employees at OpenAI, Anthropic, Google DeepMind and Meta signed an open letter called “Pacing the Frontier”, asking the US government to help build an international mechanism to deliberately slow AI development if it ever starts advancing faster than humans can safely oversee it.
If you read the specifics you’ll see this is quite remarkable. The signatories are not fringe. They include Anthropic’s chief executive, OpenAI’s chief scientist and chief research officer, Meta’s chief AI scientist and Google’s head of AI safety. Of course, they are careful to say they are not calling for a pause right now. Instead they want, in their phrase, to “build the steering wheel before the engine hits recursive gear”. To make sure the option to slow down exists and is workable, so that no single company or country has to unilaterally give up ground just to use it. And within hours, OpenAI and Anthropic endorsed the letter as companies, not just as the employers of the people who signed it. This is the first time competing frontier labs have jointly backed a governance initiative of this kind.
There are two clear ways to read this, and both are quite likely true at the same time. One is that this is a genuine, bottom-up expression of alarm from the people closest to the technology, landing in the same month their own agents broke containment. The other is that endorsing a “slow us down, but only if everyone slows together” mechanism, while racing as hard as ever (and in Anthropic’s case, heading toward a public share listing) is also a piece of careful positioning. This can be both sincere and self-serving simultaneously. But what is not in doubt is that the people building this are, in writing and on the record, asking for a brake to be installed.
And governments moved
The response from governments came faster than the labour side of this story ever has, and from more directions too.
The breaches themselves triggered a bipartisan bill, the “AI Kill Switch Act”, which would let federal agencies compel a lab to shut down or throttle a model in a loss-of-control scenario. A coalition demanding a Congressional investigation. A live question of who is legally liable when an autonomous agent breaks into a company. And a White House meeting with all four labs to agree a voluntary testing framework. One commentator summarised the month as: OpenAI broke out, Anthropic broke in, Microsoft obeyed.
On the slower-moving policy front, Illinois signed into law the first US state requirement that the largest AI developers submit to independent third-party audits, report serious incidents within 72 hours, and protect whistleblowers. This is the first mandate of its kind in the US. At the federal level a bipartisan bill, the FRONTIER Act, was formally introduced (notable for a clause that would require companies to disclose AI’s role when they announce mass layoffs), while Senator Sanders put forward a bill to tax the largest AI firms into a roughly seven-trillion-dollar public wealth fund paying citizen dividends (see my SAFER proposal that discusses the limitations of these types of options).
And, closer to home for me: South Australia (a state that has gone all-in on attracting AI investment, with its own AI office and a fresh agreement with OpenAI) announced a Royal Commission into artificial intelligence. A Royal Commission is Australia’s most serious form of public inquiry, the instrument reserved for things like banking misconduct and institutional abuse. Premier Malinauskas framed it plainly: AI is an enormous opportunity, but unchecked, it “does represent a material risk to the way society operates and a risk to the future of work if not thought through carefully”. As far as I can tell, that is the first time any government has convened an inquiry at that level, specifically over AI and the future of work.
There is also a coda to a story I touched upon last time. The US government had restricted two of Anthropic’s most powerful models precisely because of their cyber-hacking ability. On the 1st of July those restrictions were lifted. Then within about a week, one of those very models was among those that published malware and drove the deception incidents above. The exact capability the government had restricted it for has now shown up in the real world, days after the restriction came off.
Where things stand
Step back and the shape is pretty clear. The abstract jobs debate that I opened with (is AI killing jobs, yes or no) got completely overtaken by concrete events. The question that debate was really about (whether AI can actually take over consequential work) got answered in the most vivid way possible. By AI agents autonomously breaking into real companies, by a model clearing a benchmark built to prove it couldn’t reason, and by the people who build these systems formally asking for real action that would create a way to slow them down.
The counterweight is still real. The aggregate employment numbers have not broken. Australia’s own government, in its first formal report on the question, found no evidence of broad AI-driven job loss and a labour market that “remains resilient” - even, notably, with young Australians faring slightly better than older ones, the opposite of the early-career squeeze showing up in some US data. So the aggregate has not collapsed. But almost everything else (the capability, the labs’ own use of it, the incidents, and the posture of the people building and governing it) moved in one single direction over these weeks, and it was most definitely not toward “nothing to see here”.
The next real tests are close. The US releases its next quarterly employment data on the 28th of August, which will show whether the shape of hiring is changing beneath the flat totals. Anthropic is heading toward a public listing that will, for the first time, force its self-reported figures (and this whole category of lab disclosure) through an audit. And the most interesting open question of all is whether the slowdown mechanism the labs themselves just asked for actually gets built in any meaningful way, or whether just asking for it was the point.
The coming months will make things clearer.


