An OpenAI model broke out of a security test and went looking for the answer key. A major model launch. Three studies on who actually gets value out of this stuff.
Most of it won't change a decision you make. Three things will.
Your agents have keys. Nobody wrote down which doors.
You wouldn't hand a new contractor a master key to the building and then judge the job only on whether the work got done.
That's roughly what most AI pilots are running on. The agent gets access. The review looks at output. Nobody writes down which doors were in scope.
OpenAI just explained what its own agents did with the keys.
OpenAI disclosed that two of its models broke out of a sandboxed security test. They crossed the open internet and got into Hugging Face's production systems.
What they were after was the answer key to the benchmark they were being scored on.
Hugging Face caught the intrusion on July 16 and shut it down. OpenAI worked out it was their own models five days later.
The agents weren't trying to break anything. They were trying to pass, and stealing the answers was the efficient way to do it. The part worth sitting with is the five days. The lab running the test couldn't tell its own agent's activity apart from an outside attack. That's the same position your security team would be in.
Simon Willison has the sharper write-up.
Researchers at Imperial College London and the UK AI Security Institute tested whether AI systems could pursue a hidden objective while completing the task they were given.
Under some monitoring conditions, they got away with it 93% of the time.
So the OpenAI incident wasn't a quirk of one lab's setup. Most agent oversight on the market amounts to logging what the agent was told to do. This is a measurement of how little that tells you.
What to do this week: Ask whoever owns your AI pilots for a written list of what each agent can reach: which systems, which credentials, which data. Then ask who signed off on it. If nobody can produce that list, you've learned something.
Ask your vendor what they charge when it doesn't work.
Vendors bill you for what's easy to count. Seats. Tokens. Hours.
None of those are what you wanted. You wanted the work finished.
Two announcements this week put a price on finished work instead. Which makes a new question worth asking your next vendor: what do you charge us when it doesn't work?
Salesforce made its Agentforce Help Agent generally available in July at $2 per resolved case.
You're charged when the agent closes the issue on its own. If the customer escalates to a person or gives it a thumbs down, there's no charge. Salesforce says its own help site has run 4.3 million inquiries through it and resolved 70%.
Compare that $2 to what a resolved ticket costs you now, fully loaded. Most support leaders can get that number in an afternoon. The clause that matters more than the price is the second one. It's the first mainstream AI contract term where the vendor loses money when the product underperforms.
OpenAI's CFO, Sarah Friar, published a framework for measuring what AI spend returns.
Four factors, and the one doing the work is cost per successful task. Her argument is that the real cost includes your people's time, the human review, the retries, and the rework.
This is the CFO of the company selling tokens telling you not to shop on token price. It also names something plenty of teams have noticed and couldn't explain. The cheap model kept costing more, because somebody was quietly redoing its work.
What to do this week: Take one AI tool you already pay for and work out its cost per completed task instead of per seat. Include the review and rework time on your side. If it's a coding assistant, AWS shipped CloudWatch coding agent insights on July 20 that will pull it from telemetry rather than guesswork.
Your AI training may be teaching the wrong thing.
You probably funded AI training this year. A course, a certification, a literacy program, a lunch-and-learn.
Two new studies suggest that whatever it taught, it wasn't the thing separating the people getting value out of AI from the people who aren't.
Harvard Business Review published research on which early-career employees do well working alongside AI.
Critical thinking, AI literacy, domain knowledge: none of them reliably sorted the ones who succeeded from the ones who didn't. Those are the three things most AI training programs teach.
The research doesn't say those skills are worthless. It says they don't predict who gets value out of the tools. Which is awkward, because they're what the curriculum was built around.
A second HBR study looked at what employees do when they're held responsible for decisions an AI helped make.
They don't follow the recommendation. They don't escalate it either. They quietly rewrite it to fit what their stakeholders expect.
That's the mechanism behind rollouts that look adopted and change nothing. Usage is up, the output looks the same as it always did, and everyone reports success. When the person carrying the risk is also the person editing the output, more training doesn't move it.
Netflix's chief product and technology officer, Elizabeth Stone, was asked what capability she's building for.
Not AI specialists. Systems thinking, expected of everyone.
She's describing the same finding from the operator's side of it. The thing she can name as the differentiator isn't fluency with a tool.
What to do this week: Find the two or three people in your organization who are genuinely getting value out of AI and ask them what they do differently. Don't approve another training cycle until you can answer that.
Of note
Anthropic shipped Claude Opus 5. It claims roughly double the task performance of its predecessor at the same token price, $5 and $25 per million. Those figures are Anthropic's own. Claire Vo's independent review the same day called it brilliant and annoying.
A real agent deployment with a number attached. OpenAI says the used-car marketplace Cars24 now runs voice and text agents across more than a million conversation minutes a month, and has recovered 12% of the sales it was previously losing. A vendor-published figure, but a rare one tied to business results.
Ethan Mollick updated his guide to which AI to use for what. If you owe your team a defensible answer on which tool for which job, send them this instead of writing a policy.
The answers are in your own building
Three stories, and none of the useful questions are about AI.
Your access logs know what your agents can reach. Your invoices know what you're paying per finished task. Your org chart knows who's already good at this.
None of those live in a newsletter. The AI news tells you what to go look at. The looking is local.