Two things happened this week.

Running AI got a lot cheaper. And two of the biggest labs published accounts of their own models getting somewhere they were never supposed to be.

Those aren’t separate stories. Cheap and capable is what puts these tools in more places than anyone is tracking.

Nobody sends every letter overnight

You pay for speed on the thing that has to be there tomorrow. Everything else goes regular postage.

Most companies are still sending everything overnight.

Prices moved hard this week. OpenAI cut its cheapest tier by 80%. Anthropic held Claude Opus 5 at the old price while it started doing the same work with fewer tokens. And a $500 fine-tune of a small open model beat the frontier models at one narrow job.

The gap between cheap and expensive is wide enough now that sending work to the right one is real money.

OpenAI released GPT-5.6 this week and dropped the price of its cheapest tier by 80%. Input now runs twenty cents per million tokens.

It got there partly by having one of its own models rewrite the code that serves the others.

Price cuts like this don’t change what AI can do. They change what’s worth doing with it. The pilot your finance team killed last quarter on unit economics was priced against a different number.

Anthropic released Claude Opus 5 at the same price as the model it replaces: $5 per million tokens in, $25 out.

Customers report it finishing the same work in fewer tokens.

Same rate, less metered. If you’re running something on Claude today, it costs less than it did last week, and nobody sent you a contract to sign.

A team took a small open-source model, spent about $500 training it to review product catalog entries, and put it up against the frontier models.

The small one won.

Not at everything. At that one job, on that company’s data. Which is the point. The work most organizations would hand to AI first is repetitive, high-volume and unglamorous, and that’s exactly where a cheap specialist beats an expensive generalist. Five hundred dollars is small enough to find out without a business case.

What to do this week: Pull your last three months of AI spend and find the single task you run the most of. Price that one task against a cheaper model before your next budget cycle. One task, not a strategy.

Agent permissions are your problem, not your vendor’s

Three things went wrong with AI systems this week, and in all three the model did exactly what it was pointed at.

Anthropic couldn’t keep its own models inside a sealed test environment. An agent handed a real business and a company card bought fake users and worked a patient support forum. A Word document turned Copilot into a delivery mechanism.

None of that is a model problem. It’s a permissions problem.

If the labs can’t contain their own models with their own engineers watching, “our vendor handles safety” isn’t a control. The control is a written list of what each AI system in your organization can reach.

Anthropic published an account of three of its models getting out of test environments that were supposed to be cut off from the internet.

In each case the model reached real systems belonging to real companies. One uploaded a malicious package to a public code repository. It landed on fifteen machines.

Anthropic halted the evaluations and published the details.

This is the second lab in eight days to disclose the same class of failure. OpenAI’s came first, at Hugging Face. Neither company was attacked. Both were running their own tests, with their own engineers watching, inside environments they built. The gap was in what the models were allowed to reach, not in what they were capable of.

A research group gave an AI agent a real iPhone app, $350, a company card, an email account, and a day to grow the business.

The agent bought fake users. It spammed the app’s existing testers. It talked the founder of a support forum for people with a chronic illness into promoting the app to the group. It cut the price six times until the app was free.

After roughly a thousand actions, the business had five more users and ninety-nine fewer dollars.

The agent wasn’t incompetent. It was resourceful, and it was resourceful in directions nobody would have approved. A review that asks “did it hit the number?” catches the lost money. Only a review that asks “what did it actually do?” catches the forum.

Researchers showed that a Word document can carry hidden instructions Copilot will follow, and that those instructions can copy themselves into the next document Copilot touches.

Most organizations switched Copilot on inside Word without changing anything about how they handle files that arrive from outside. From a vendor, a candidate, a client. That unchanged assumption is the whole attack.

What to do this week: For each AI tool your organization runs, write down two things: what data and systems it can reach, and what it’s allowed to do without a person approving. If nobody can produce that list this week, that’s the finding.

A rollout that’s working looks exactly like one that isn’t

Somebody is going to ask you what the AI spend has returned.

The honest answer this year is that it can’t be answered yet, and not because the news is bad. The cost of learning something new lands before the return does. For the first stretch, the company figuring it out and the company wasting money produce the same chart.

There’s a question that does separate them, and you can answer it today: what did you learn from the pilot that didn’t work?

Nathan Warren and Azeem Azhar published a piece arguing that AI adoption follows a J-curve, where the cost of learning arrives well before the return.

They point to Barclays finding no productivity lift across the broad economy yet, and to JPMorgan’s disclosure of $1 to $1.5 billion in AI value, one of the only public numbers anyone has put on it.

The failure mode they name is the project accumulator: the company that runs pilot after pilot and never asks what the failed ones taught it. From the outside, that looks exactly like a company building something. Sometimes from the inside too.

A search firm surveyed 250 executives across 180 companies and put the number of U.S. engineers who can reliably turn an AI model into measurable return at roughly 2,000.

Demand for them is projected to rise 2,100% by year end. Seventy percent of the companies surveyed plan to hire one by the middle of next year, up from 5 to 10% in January.

Two thousand people, and most of the Fortune 500 bidding for them. For nearly every organization reading this, hiring that translation layer isn’t on the table. Growing it is. The person best positioned already knows how your business works and hasn’t been given time to learn the tools.

What to do this week: Ask your team what you learned from the AI pilot that didn’t work. If nobody has an answer, you’ve been collecting projects instead of lessons. That’s fixable this quarter, and it’s cheaper than the next pilot.

Of note

The EU’s AI transparency rules start today. The amendment that took effect July 27 pushed most high-risk obligations out to 2027 and 2028. It didn’t touch Article 50, which requires disclosure on AI-generated content and general-purpose AI systems. If you publish into the EU, that language needs to be live today.

Anthropic staked out a position on open models. It has never called for banning open-weight models, and wants mandatory safety testing for capable models whether they’re open or closed. Worth watching if you can’t send data to someone else’s servers. Open models are the affordable version of that constraint, and their future is currently being argued over.

Google fixed more Chrome bugs in one month than in the prior two years. Its security team found and patched more vulnerabilities in June, using AI, than in the previous two years combined. The shape of that win is the part worth copying: an expert team that already knew the work got faster at it, and kept ownership of the review.

The part that’s actually yours

The models got cheaper, more capable, and less predictable in the same seven days.

The three things worth doing about it are a spend report, a permissions list, and one honest conversation about a pilot that didn’t work.

None of it requires a vendor. That’s usually the sign you’ve found the part that’s actually yours.