Eight of them, working through Taiwan’s government networks for four days in July. When something blocked them, they looked up another way in.

That’s the new thing here.

Most of what follows is a story we’ve been telling since July. What’s different is that it now comes with date

This time, someone was aiming them

We’ve covered agents getting out twice in the last month. Both times it was a lab losing track of its own models inside its own test.

This is the other direction.

Someone assembled eight agents out of free, downloadable software, pointed them at Taiwan’s government, and got exactly what they were after.

The permissions list we asked for in July still matters. It’s just doing a different job now.

Security researchers published the record of an attack on Taiwan’s government that ran without a person directing it.

Eight AI sub-agents, working together across twelve waves in early July.

Over four days they mapped 21 government systems, got into 85 accounts, and took 2,500 personnel records. Nuclear safety and energy agencies were among the targets.

When something blocked them, the system researched a different technique and tried again.

It was built on Hermes and OpenClaw. Both are open source. Both are free.

The number to sit with is four days. Most incident response assumes an attacker who sleeps and takes weekends. This one ran continuously and rewrote its own approach every time it hit a wall.

You don’t need a state budget to run this.

OpenAI split its cybersecurity program, Daybreak, into two tiers and released a model trained specifically for security work. It used that model to find two unknown bugs in the engine behind Chrome.

The part worth noticing is the gate. Access to the strongest defensive models now depends on passing a vetting process, which is a different conversation if your security provider is a two-person shop than if it’s a national firm.

What to do this week: If someone worked inside your systems for four straight days without stopping, when would you find out? Ask whoever owns that answer. If it depends on a person opening the logs Monday morning, that’s the gap.

The introductory rate expires today

Every provider in this market priced like a new gym in January. Cheap to join, and the real number arrives later.

DeepSeek’s is today.

The same week, OpenAI put a $125 seat next to its $25 one and started running ads on the free tier.

None of that is coordinated. It’s the same bill coming due in three places.

Ten days ago DeepSeek said a significant price rise was coming. The terms are out now.

API prices for its V4 models rise between 50% and 1,100%, depending on the model, the token type, and the time of day. A new peak-hour rate costs more than four times what the same call cost yesterday.

It starts today.

DeepSeek was the price everyone else got measured against. That measurement moves on Monday.

Most organizations don’t buy from DeepSeek directly, which is the awkward part. Plenty of AI features inside other products route to it without saying so, so the change arrives through a vendor who has no obligation to mention it.

OpenAI added a Premium seat to ChatGPT Business at $125 per user per month, or $100 paid annually. The standard seat stays at $25.

Premium gets five times the usage and removes the five-hour cap.

The introductory pricing ends August 20.

Until now everyone in an organization cost the same. Now your heaviest users cost five times your lightest, and somebody has to decide who gets which.

OpenAI began testing ads in ChatGPT for logged-in users on the free and Go tiers. Paid business tiers stay ad-free.

Last week we noted the free tier had gone unlimited, and that the old rate limit had quietly been doing governance work nobody assigned it. It’s being monetized now instead of metered.

Whoever on your staff uses a personal account for work is getting sponsored results inside answers they’re treating as research.

What to do this week: Two dates. Before August 20, find out who on your team is hitting the five-hour cap on ChatGPT Business. Ask your software vendors which model sits underneath their AI features, today. If anyone answers DeepSeek, your cost changes that day and nobody will write to tell you.

Two companies just answered a question you were deferring

Two weeks ago we noted, in passing, that the EU’s AI transparency rules had taken effect and that anyone publishing into Europe needed disclosure language live that day.

Nine days later Anthropic started watermarking everything its models write.

Not a setting you switch on. A property of the output.

Spotify moved the same week from the other side, putting a distribution penalty on undeclared AI music.

Anthropic will watermark text generated by its models, to comply with the EU AI Act rules that took effect August 2.

It covers the API, Claude, Claude Code, and Claude Cowork. Models released after August 2 have it built in, and older ones are getting it.

Two things aren’t clear yet. Whether enterprise and API customers can turn it off. And how much editing it takes before the mark stops being detectable.

If your team drafts proposals, reports, or client deliverables with Claude, that work now carries a signal you didn’t choose and may not be able to remove.

It’s also a European regulation changing a product you use, whether or not you have a European office. That’s the part most organizations assumed wouldn’t reach them.

Spotify will badge AI-generated artists and keep their music out of editorial and algorithmic recommendations by default.

Artists can declare themselves from August 11. Badges appear on profiles in mid-September.

Self-disclosure isn’t the only route in. Human reviewers and detection tools will flag profiles that look like AI-generated identities, starting with artists above a certain audience size.

Declare and lose reach. Don’t declare and get flagged anyway.

That mechanism is the one other platforms will copy. If your organization publishes marketing, video, or written content at any volume, that’s the shape of the rule heading toward you.

What to do this week: Find out which client-facing documents your team drafted with Claude in the last month. Then decide whether you’d say so if a client asked. The watermark turns that into a question somebody else can answer about you.

One dollar on the tool, three on the process, five on the people

Last week we argued that how well an agent performs is mostly decided by how well the job was described, not by which model ran it.

McKinsey has now measured the same thing from the money side.

In the organizations that actually scale this, every dollar spent on the technology comes with three on redesigning the process and five on training people to use it.

Most companies spend it the other way around.

McKinsey published research on why agent projects stall short of scale.

88% of organizations use AI in at least one function. 23% are scaling an agentic system.

The ones that get there follow the 1:3:5 ratio above. Most companies invert it, and treat the capability work as an implementation detail.

We’ve made a version of this argument twice without a number attached to it. The number is what makes it usable. You can hold your own budget against that ratio this afternoon and get an answer.

InformationWeek made the case that agent failures are process failures, citing Gartner’s forecast that 40% of agentic AI projects get canceled by 2027.

Three examples carry it:

  • A customer service agent that answers fluently and never escalates.

  • A finance agent that pulls exactly the right data and applies the wrong exception policy.

  • A coding agent that writes valid code into the wrong repository.

Every one of those looks fine in a demo.

The agent did what it was asked. What it was asked was wrong, and nobody found out until it had been doing it at volume for a while. That failure doesn’t show up in a model evaluation.

What to do this week: Split this year’s AI budget on paper into three lines: tools, process redesign, and training. If tools is the biggest number, you’re running the ratio McKinsey found in the companies that never scale.

Of note

Microsoft is merging its Copilot apps. The consumer app and the Microsoft 365 one become a single app: mobile and web from mid-August, Windows and Mac in mid-September. Merging a personal-account app with a work-account app is exactly when people get confused about which one holds company data. Worth a short internal note before it lands.

Meta’s new open model runs on a laptop. Muse Glimmer: 30 billion parameters, Apache 2.0 license, under 20GB, one consumer graphics card. In July we said open models were the affordable version of the can’t-send-data-out problem, and that their future was being argued over. This is that argument landing usefully.

Claude Code, explained for people who aren’t engineers. Lenny’s Newsletter published a walkthrough covering skills and voice mode. It’s aimed at the person who’s been told repeatedly to use it, assumed it was a developer tool, and quietly opted out. Worth forwarding to whoever that describes.

A calendar, not a watch list

Most weeks the AI news tells you which way something is heading.

This one came with dates. August 17, August 20, mid-September, and a rule that’s been live since the second of the month.

Three of these four stories we’ve run before. The difference is that you can put them on a schedule now.