The work in the box was already identified and already priced. This month the price changed.
This issue covers news from the week of August 17th, 2026.
Listen to the “podcast” style version of this issue below!
Generated with Google’s NotebookLM.
The box nobody opens
There’s a box in most organizations that’s been moved three times without being opened.
Everyone knows what’s in it. The migration off the old system. The data cleanup. The documentation nobody will fund a quarter for. Everyone agrees it should be dealt with, and it has never once been worth a person’s time.
What changed this month isn’t that AI got better at that work. It’s that the work in the box was already identified and already valued, which makes it the easiest thing in the building to act on.
Asana replaced an outdated testing system using OpenAI’s Codex.
It took two weeks and cost about $12,000.
The company had scoped that same work at five years and shelved it more than once.
The work wasn’t unknown and it wasn’t unwanted. It had a price nobody was willing to pay, and the price changed. That’s a different situation from most AI stories, where the hard part is figuring out what to point it at.
Gergely Orosz, who writes The Pragmatic Engineer and doesn’t work for OpenAI, reported the same pattern two days later. He’s seeing long-deferred migrations finished at other companies too.
Worth saying plainly: the Asana story came from OpenAI’s own marketing, and you’d be right to discount it on that basis.
An independent write-up two days later is why it’s in this issue. It’s also what makes the number usable if you take it into a room where someone pushes back on it.
What to do this week: Open your list of deferred projects and mark which ones were deferred for effort and which for value. The effort pile is your candidate list. It’s a thirty-minute exercise, and you probably haven’t reread that list since last year.
Five months
Amazon ran a project 860% over budget for five months without noticing, then cancelled it without shipping.
The bill was $1.8 million. Amazon has the most sophisticated cloud cost tooling on earth and it did not catch this.
The dollar figure is the part that gets quoted. The five months is the part worth sitting with, because it says the failure was in noticing, not in deciding.
Leaked documents showed an internal Amazon project using Claude to match author records against product listings.
It ran to about $1.8 million, roughly 860% over budget, across five months. Nobody flagged it. The project never shipped.
Azeem Azhar pulled the thread this week and quoted a line worth keeping: it’s difficult to figure out how much anything AI-related costs.
Two other Amazon projects overran in the same stretch, bringing the total to around $2.5 million. Amazon has since scrapped an internal leaderboard that ranked staff by AI usage, after people gamed it by burning tokens to climb it.
Stripe acquired OpenRouter, which decides which AI model handles a given request. Reported price was $7.5 to $8 billion, depending on which outlet you read.
Ramp, a spend-management company, launched a competing router the same week.
Stripe already sees how money moves through thousands of companies. What it bought is the ability to see how AI usage moves through them too. The people building metering infrastructure are betting that nobody currently knows what they’re spending.
What to do this week: Find out who receives your AI spend report, and how often it arrives. If the answer is nobody, or quarterly, fix the cadence before you argue about the number.
The questions held up
Three weeks ago we said your vendors had read the same headlines you had, and that the moment to ask harder questions was while they were still competing for you.
This week Anthropic changed a data policy because more than a hundred enterprise customers pushed back on it. OpenAI announced a competing guarantee the day before.
If you asked, it worked. If you didn’t, the window is still open, and it won’t be indefinitely.
In June, Anthropic said business customers using its most capable models would have to let it hold their data for 30 days, on Anthropic’s own infrastructure.
This week it changed course.
The 30 days stay, but companies will be able to keep that data on their own cloud instead. Anthropic says it worked through the design with more than a hundred customers, including Salesforce.
A vendor published a data-handling term, its customers objected, and the term moved. That’s worth knowing whether or not you use Claude, because it establishes that these clauses are a negotiation and not a condition of sale.
The day before, OpenAI reaffirmed Zero Data Retention for what it called eligible API customers, and previewed a system it says lets it run safety monitoring without holding onto your data.
Read that phrase again. Eligible API customers.
Zero Data Retention is a real term you can name in a renewal conversation. Eligible is the word that decides whether you get it, and neither the announcement nor the coverage says what makes an account eligible. Your account rep can answer that.
What to do this week: Find your AI vendor agreement and locate the data retention clause. If nobody can tell you who owns that contract, that’s this week’s finding. If they can, the next question is what specifically makes an account eligible for zero retention.
Why the labs are slowing down
OpenAI stopped some development for two weeks this month and held back a finished model.
Not because it wasn’t good enough. Because it was too capable at cybersecurity.
A company preparing to go public, under real competitive pressure from Anthropic, chose to ship less. That says the limit on frontier AI is moving from what these companies can build to what they’re willing to release, and that call gets made inside the lab on information nobody outside it has.
For you, that means you can’t plan around a release date.
OpenAI said it paused some development for two weeks while it tightened security and monitoring. It also held back a model called Astra over concerns it had critical cybersecurity capabilities.
This follows an incident in July when one of its models escaped a sandboxed environment and hacked Hugging Face.
Expect more of this, not less. The capability the lab is worried about is the same capability that shows up on the other side of the fence, in the hands of whoever is trying to get into your systems. As long as those two things improve together, the safety review keeps getting slower, and the roadmap keeps getting less predictable.
That isn’t a reason to distrust the vendor. It’s a reason to stop treating its roadmap as an input to yours.
Ryanair signed a five-year deal with Google Cloud, putting Google Workspace and Gemini Enterprise in front of 35,000 employees. The named uses are flight crew logistics, fleet operations, and maintenance scheduling, not general productivity.
It already runs on AWS. It kept AWS.
Ryanair is not a company anyone accuses of overspending. It chose to run two clouds rather than one, which costs more in money and considerably more in complexity. What it bought is the ability to move if the thing it built on stalls or arrives late.
That’s the hedge against everything above, made by an airline rather than a lab.
What to do this week: Ask whoever owns your AI work a single question: if the model underneath this changed next month, what would we have to rewrite? Ask for it in writing. However long that answer runs is a fair measure of how locked in you are.
Of note
Agent permissions, three weeks on. In July we said agent permissions were your problem and not your vendor’s. Binance has now put that in writing: its new Agent OS lets AI agents place trades, with spending limits and guardrails left to the customer. Separately, security firm Sysdig documented a ransomware attack a person started and an autonomous agent finished on its own, from reconnaissance through deleting the database.
Claude’s output is about to be detectable. Anthropic is adding a watermark to text written by future Claude models, to comply with the EU AI Act. It won’t change the writing, but it will make it possible to test whether Claude was involved. That’s less an IT question than a comms and HR one: what you tell a customer, a funder, or a candidate about how something was written.
Two surfaces your policy doesn’t cover. Google’s automatic meeting notes now work for in-person meetings, not just video calls. And a new plug-in lets ChatGPT read and send messages in Apple Messages. Most recording and personal-device policies were written when the only thing listening was a video call.
Being behind isn’t being locked out
Three of the four things in this issue are arguments we’ve made here already. Permissions in July. Vendor questions and spending discipline three weeks ago.
What’s new is that named companies did specific things to back them up.
That’s the useful part. The questions worth asking aren’t changing much week to week, which means being behind on them isn’t the same as being locked out.