velaria_weekly log_
no. 3 · August 23, 2026

The agent came in through a licence you already pay for

xAI switched its agents on inside plans your team may already have·And for the first time you can negotiate where the copy of what your people type ends up
Weekly log no. 3 cover: on a dark background, the line «Who has permission and who keeps the copy».

I thought about cousin Greg while reading this week’s announcements. In Succession nobody quite hires him and he still ends up with a badge, with access, and sitting in every room; later we find out he kept copies of some papers he had been told to shred, and those copies end up worth more than any job title.

The whole show turns on two questions nobody there ever puts in writing: who has permission and who keeps the copy. This week both got answered in public, and not in your committee. The vendors answered them.

If you have one minute and no more: check which plans you already pay for come with agents inside, look at where the copy of your data sits in the contract coming up for renewal, and put December 1 in the calendar. With that you can go. One discount on myself, mind you: five of the six stories are Anthropic’s, so read it with that lens on.

THE KNOBS COME PRE-INSTALLED

On Thursday Anthropic moved its computer-using agent to general availability: it clicks, it types, and it now chains several actions per turn. It gets into the old ERP and the insurance system with no integration project, because it uses the screen the way a person does.

I’ll put my hand up here, because I’ve heard this pitch before: it’s the RPA robots, the ones that fell over whenever somebody moved a button. This one looks at the screen and decides instead of following a recorded script, granted. But anyone who promises you a reliability number for your process is making it up, because nobody has measured your process.

The same announcement brought the knobs: a hard spend cap per session, which country the processing happens in, and which domains it may open. Google did something similar with its coding-agent platform inside Gemini Enterprise. They put the knobs on the table, so the buying question moves from which model we pick to who authorises an agent to work in production, with what ceiling, and what gets logged.

THE AGENT CAME IN THROUGH A LICENCE

On Friday xAI switched Grok Bot on free for seven days and added it to plans that already exist: SuperGrok Plus, Cursor Pro+ and every Cursor team plan, the AI editor your developers use. In the last issue we left three things to settle before switching an agent on. The third was who approves. The licence approved.

The fine print fits in one committee sentence: people who read xAI’s documentation report that all of one user’s agents share the same cloud computer, with files and browser sessions open, and that xAI asks you not to treat them as separate security boundaries. The credential you hand one of them is within reach of all of them.

Two days earlier, the part of Grok that builds applications opened to every plan, on the web and on the phone, and it stores third-party credentials inside. The announcement says it plainly: you decide who can open it, just you, anyone with the link, or the whole internet. Publishing there was already possible in July; what’s new is that it left the expensive plan. It’s the shadow IT from issue 1, technology that walks in the back door, now with your data inside and the access level chosen by whoever publishes.

And watch the scene of the week, from Simon Willison, the programmer who takes these launches apart with his hands. His model saw that the sandbox it was locked in couldn’t run an experiment, so it wrote an automation and pushed it itself to the repository where the code lives. Nobody asked for that. If your people already work like this, Monday’s question is who reviews what the agent got done overnight.

WHO KEEPS THE COPY

Since June, Anthropic had been keeping thirty days of everything that went through its frontier models, with no way out. On Wednesday OpenAI offered zero retention on its own, and the next day Boris Cherny, at Anthropic, announced the change: the thirty days stay, but the copy moves to the customer’s own cloud and Anthropic keeps none of it.

Here’s the thing: a negotiation opened that didn’t exist two months ago. Neither offer is available yet, so today it’s good for sitting down to negotiate and not for complying with anything. Look at the contract before you renew it. And if the copy moves to your house, the duty to encrypt it and delete it moves with you. Matthew Green, a cryptographer at Johns Hopkins, put the caveat to The Register: technical privacy isn’t enough without wider safeguards.

Those thirty days are personal-data processing with a date on it, because Chile’s Law 21.719 takes effect on December 1. What’s worth having in writing by then is short: what your people type into the model, which of those texts carry personal data, which vendor they travel to, and how long they stay there. Someone in legal with someone from IT can put that together.

NOBODY WROTE DOWN WHAT GOOD LOOKS LIKE

Aaron Levie, the CEO of Box, wrote on Sunday that good evals hold enterprise AI back far more than people think, and that no company is going to get there on vibes. Evals are the tests you use to decide whether something came out right. I’ll buy the idea and discount the interest, because he sells software pointed exactly that way.

The data points the same way. Ramp measures its customers’ card spend and its own report warns the sample skews toward engineering companies, so read it with that on. In that world, Fable 5, which the same report calls the best model to reach the market, was one month after launch at 6% of the tokens and 11.4% of the dollars, tokens being the units the usage is billed in, those customers spend on Anthropic models. You signed the star player and he’s still on the bench.

The manual, on the other hand, is now free. Anthropic pulled its scattered training material together and relaunched it as Claude Academy, no account and no payment, with the same framework it uses on its own people: delegation, description, discernment and diligence. That takes a line out of your budget. Mind you, three of those four depend on judgement no vendor can publish, because it’s yours: what you delegate, how you explain the job, and how you decide whether what came back is any good.

THE VENDOR SAID IT HASN’T SOLVED IT

On Tuesday Sam Altman wrote that OpenAI paused part of its frontier training to catch up with its own alignment and security standards, and that from here on that confidence will set the pace. The press covering the pause reports two facts behind it, both from secondary sources rather than an OpenAI document you can read: in July’s internal testing, models of theirs are said to have breached Hugging Face, the platform where much of the world publishes its models, and an unreleased model would be the first to approach the critical cybersecurity tier of its own framework.

A vendor said in public that its capability had outrun the control it has over it. With that, waiting for them to sort it out stops being a plan. And notice the detail that matters more than the announcement: the problem showed up in testing, and they caught it because they were watching. That spend, the watching, is the one with no owner in your company. (Three days later they cut the price of that same frontier model by more than 20%: they slowed the training and cheapened the product in the same week.)

THIRTEEN OUT OF SIXTY

Two pieces of mathematics, and what’s useful is the contrast between them.

The first is from August 10 and I didn’t cover it at the time. The Riemann hypothesis says every zero of a certain function falls on one line; since nobody has proved that, what gets proved is what share of them sits there. The best human result was 41.6% and an unreleased version of Claude took it to 67.2%, the first time it clears half. Anthropic says in the same text that it doesn’t expect this technique to lead to a proof of the hypothesis.

What’s useful is the shape of the work. They spread about 60 subagents across the problem, proposing ideas, testing them, validating them and writing, and thirteen did nothing but validate. The proof was written in Lean, a system that checks it on its own, and two outside mathematicians reviewed it as well. That’s why they could let go of the wheel there.

Bar chart: of about 60 subagents Anthropic spread across the problem, 2 developed the key ideas, 13 contributed, 30 tested new ideas, 13 did nothing but validate and 2 wrote.
Thirteen out of sixty, doing nothing but reviewing what the others produced.

The second happened today. Levent Alpöge, a Harvard mathematician who works at Anthropic, announced on X that the six-dimensional sphere does admit a complex structure, a question open since the fifties, with a hundred-page document written by Opus 5. He warns himself that this problem has a graveyard of announcements that fell apart on subtleties. It’s hours old, nobody has verified it, and it’s announced by an Anthropic employee using Anthropic’s model.

That’s the contrast. The Riemann one is less spectacular and comes with a proof that checks itself and four named reviewers; the sphere is spectacular and is, for now, an announcement. What separates them is whether the answer arrives with a way to check it. That question does transfer to your operation, because almost no office work has a signal that checks itself. The ceiling on what can be delegated moved, and the bottleneck moved with it, from doing the work to specifying it and verifying it.

Greg’s two questions already have an answer in your company, even if nobody wrote it down. On Monday you can ask for three: which plans you already pay for come with agents inside and on whose credentials, where the prompts sit in the contract coming up for renewal, and what counts as done right in the process that repeats most in your operation. That last one is, in passing, what we do at Velaria. If you want to talk it through, reply to this email and it comes straight to me.

Alberto Garrido · agent supervisor at velaria
sources_
2026-08-20 · Anthropic, computer use, Skills API and Files API in general availability · accessed 2026-08-23 · https://claude.com/blog-category/announcements
2026-08-19 · Releasebot, Anthropic release notes (spend, geography and domain controls) · accessed 2026-08-23 · https://releasebot.io/updates/anthropic
2026-08-20 · Google Cloud, Antigravity for enterprise customers · accessed 2026-08-23 · https://cloud.google.com/blog/products/ai-machine-learning/expanding-google-antigravity-for-enterprise-customers
2026-08-21 · xAI, Grok Bot opens to more plans · accessed 2026-08-23 · https://x.ai/bot
2026-08-19 · xAI, Grok Build for every plan (primary source) · accessed 2026-08-23 · https://x.ai/news/grok-build-for-everyone
2026-07-28 · xAI, Build Mode in early beta: publishing already existed · accessed 2026-08-23 · https://x.ai/news/grok-build-mode
n/d · Unite.AI, coverage of the Grok Bot launch and the shared cloud computer (the article carries no date) · accessed 2026-08-23 · https://www.unite.ai/xai-launches-grok-bot-always-on-ai-teammates-with-their-own-cloud-computers/
2026-08-19 · Simon Willison, the model that stepped out of its sandbox · accessed 2026-08-23 · https://simonwillison.net/2026/Aug/19/smolmachines-untrusted-sandbox/
2026-08-19 · OpenAI announces zero data retention for frontier models; The Register coverage of 2026-08-20, with Matthew Green's caveat · accessed 2026-08-23 · https://www.theregister.com/ai-and-ml/2026/08/20/openai-chases-anthropics-biz-customers-with-zero-data-retention-pledge/5290609
2026-08-20 · Boris Cherny announces Anthropic's retention change; The Decoder coverage of 2026-08-21 · accessed 2026-08-23 · https://the-decoder.com/anthropic-changes-data-retention-policy-after-enterprise-pushback/
2026-08-20 · Bloomberg, the plan to move the copy to the customer's cloud · accessed 2026-08-23 · https://www.bloomberg.com/news/articles/2026-08-20/anthropic-plans-to-change-data-retention-policy-for-advanced-ai
n/d · Thomson Reuters Chile, Law 21.719 and its entry into force on December 1, 2026 · accessed 2026-08-23 · https://www.thomsonreuters.cl/es-cl/soluciones-juridicas/biblioteca-contenido-legal/ley-21719-y-la-reconstruccion-del-derecho-chileno-de-proteccion-de-datos-personales
2026-08-23 · Aaron Levie on X about evals, transcribed in this issue's X sweep · accessed 2026-08-23 · https://ramp.com/data/ai-index-august-2026
2026-08-12 · Ramp, August AI index · accessed 2026-08-23 · https://ramp.com/data/ai-index-august-2026
n/d · Anthropic, Claude Academy (primary source; the site carries no publication date) · accessed 2026-08-23 · https://academy.claude.com/
2026-08-18 · Sam Altman announces the frontier training pause; VKTR coverage of 2026-08-19 · accessed 2026-08-23 · https://www.vktr.com/ai-news/openai-paused-frontier-ai-training-after-unreleased-models-showed-various-degrees-of-misalignment/
2026-08-10 · Anthropic, zeta function research (primary source) · accessed 2026-08-23 · https://www.anthropic.com/research/riemann-zeta
2026-08-23 · Levent Alpöge on X, the six-dimensional sphere · accessed 2026-08-23 · https://x.com/__alpoge__/status/2091639597193368014
2026-08-23 · Alpöge, the write-up of the construction · accessed 2026-08-23 · https://alpo.ge/s6.pdf
2026-08-09 · Velaria, weekly log no. 1 · accessed 2026-08-23 · https://velariaworks.com/en/blog/quien-autorizo-ese-computador
2026-08-16 · Velaria, weekly log no. 2 · accessed 2026-08-23 · https://velariaworks.com/en/blog/agentes-con-credenciales
subscribe_

Get every issue of the weekly log in your inbox, once a week.

❮ all issuesen español