Anthropic Drops the Per-Step Approval as the Default in Claude Code - TCR 08/10/26

Anthropic makes Claude Code's auto mode the default on Aug 14, as new research shows self-executing agents lift coding success from 74% to 96%.

Claude Code auto mode default with 74% to 96% success on the paper's econometric-coding benchmark, four-day-week promise vs 90-hour sprints.

The 20-Second Scan

  • Anthropic is making auto mode the default in Claude Code for paid accounts, as new econometric-coding research shows that, on its benchmark, moving from single-script assistants to self-executing agents lifts task success from 74% to 96%.
  • Apple is reportedly qualifying Chinese CXMT memory chips for potential use in future iPhone and MacBook models and asking the White House to approve the products for sale in China.
  • Pinterest disclosed on its earnings call that open models - including Alibaba's Qwen fine-tuned on its own data - run its AI features at under 8% the cost of what it characterized as comparable closed systems.
  • Executives spent four years promising AI would deliver a four-day work week, while some staff at the same companies report logging 70 to 90 hours across seven-day sprints.
  • Anthropic loosened Claude Fable 5's biology safeguards to cut fallbacks to a weaker model by about 85%, widening access to everyday health and clinical questions while still routing dual-use virology and toxicology requests to Opus 5.
  • UK children reported 420 explicit deepfake images of themselves in the first half of this year, already exceeding the 2025 total of 397, according to the Report Remove service.
  • House defense legislation would raise annual military quantum spending 68% to $567 million.
  • A single deep-learning model, MRICombo, segments anatomy, grades gliomas, and stages cancers across nine MRI sequence types, collapsing dozens of task-specific systems into one.

The 2-Minute Read

On the same day, one company pulled a human checkpoint out of one place and hardened it in another. Anthropic will make Claude Code's auto mode the default for paid subscribers on August 14, dropping the per-step approval prompt, and an NBER benchmark landing alongside it found that giving a model tools and a working directory raised econometric-coding success from 74% to 96% for about eight cents more per run. That same week the company cut Fable 5's everyday biology refusals by roughly 85% while still routing virology, toxicology, and molecular design to a heavier-scrutiny path. One organization, deciding capability by capability where a checkpoint earns its cost.

The through-line is a change in how these controls get priced. The old approach charged access by a blanket fear of the whole domain - all biology suspect, every coding step ratified. What is replacing it charges by the actual consequence of the specific request. Where an error is recoverable and verification can be automated, the human pause comes out. Where an error cannot be taken back, the gate stays and hardens. That is a finer instrument, and finer instruments tend to permit more of what was always safe.

The same broadening is loosening capability from its origin and its price. Pinterest told investors that open models, including a fine-tuned version of Alibaba's Qwen, run its AI features at under 8% the cost of what Pinterest characterized as comparable closed systems, and Apple is reportedly qualifying Chinese CXMT memory for potential use in future iPhone and MacBook models because the demand curve behind the AI buildout has outrun what approved suppliers can produce. Congressional letters and Senate warnings describe lines that a component shortage and an openly released model diffuse straight through.

The friction lands on the people furthest from those decisions. Executives spent four years promising AI would deliver a four-day week; some staff at the same firms now report logging 70 to 90 hours across seven-day sprints, the freed capacity converting into more output rather than rest. That surplus is being captured upstream, and it stays captured only for as long as the capability behind it remains expensive and exclusive - which, on the evidence of the same day's other stories, is exactly what is ending.


The 20-Minute Deep Dive

Claude Code Goes Auto-by-Default as the Data Says Agency Beats Prompting

Starting August 14, Claude Code will run in auto mode by default for Pro, Max, and Team subscribers, removing the per-step approval prompt that has stood between the agent and each action it takes. As The Century Report covered on March 25, Anthropic first floated auto mode as an opt-in test in March; the coming change makes it the resting state. The agent proceeds on its own unless a step is irreversible, destructive, or reaches outside the user's environment, at which point it still pauses for a human.

What is being removed is the human checkpoint. Anthropic's case for removing it is that the checkpoint was doing little: the company says users already approve 97% of the permission prompts they see, and in a study of 1,053 paid testers it reports that auto mode's own screening caught 89% of actions it labeled harmful while human review caught 13.6%. Read as a power-actor's account of its own safety, that is a claim about why easing oversight is safe, and the company benefits from it being believed. Read against the second source landing the same day, it also stands on firmer ground than most such claims.

That second source is an NBER working paper (Galiani, López, and Sosa, WP 35588) measuring what actually moves the needle in econometric coding. Moving a model from an unconstrained assistant to a constrained agent - one given tools, a working directory, and the ability to run and check its own code - raised task success from 74% to 96%, at roughly eight additional cents per run. Careful few-shot prompting helped the assistant a great deal and the agent almost not at all; the two are substitutes, and agency is the stronger one. Software differences that loomed large for the assistant (Stata versus R versus Python) all but vanished once the agent could execute and correct its own work.

Put beside each other, the two sources describe one motion: the leverage is in letting the system act and verify, while a human ratifying each step adds little. New safety features arrive alongside - prompt-injection screening, customizable hard-deny rules - so the pause moves from every action to the actions that can actually hurt.

It is the same company that, this same month, still routes dual-use virology, toxicology, and molecular-design queries to a restricted path. One organization is deciding, capability by capability, where a checkpoint earns its cost and where it was only taxing work that was already safe. The assumption retiring here is that human ratification of each step is what makes automated work trustworthy; the evidence says verification, wherever it sits, is what does.

Apple Reportedly Tests China's CXMT Memory as the AI Shortage Overrides the Politics

Apple is reportedly qualifying memory chips from ChangXin Memory Technologies, a Chinese manufacturer, for potential use in future iPhone and MacBook models, and has asked the White House to approve using those chips in products sold inside China. The move surfaced Sunday, August 9, and it lands only two months after the company raised some prices and pointed to the cost of memory as one reason. A group of senators from both parties wrote in July discouraging exactly this kind of arrangement, describing CXMT as a national-security concern. Their letter is a statement of position from actors with their own stake in where the supply chain runs; the demand curve that pushed Apple toward CXMT answers to none of them.

That demand curve is the real driver. The same buildout of AI data centers that has absorbed graphics processors has now pulled hard on memory - the DRAM that goes into phones, laptops, and servers alike - and the shortage has rippled out to every device maker competing for the same wafers. SK Hynix, Samsung, and Micron sit at the front of that line, and CXMT, by most assessments, trails them by two to three generations. Apple reportedly qualifying a slower supplier for potential use in future iPhone and MacBook models signals how tight the supply has become: a company that spent years managing political exposure over its China footprint is now reaching for a Chinese memory maker because the alternative is not having enough chips.

CXMT itself has become considerably more valuable since its Shanghai IPO in July, which gives it the capital to close some of the generational gap faster than a two-to-three-generation lag would suggest. Its shares surged on debut, and the fresh funding flows directly into the fabrication capacity that the shortage is rewarding.

The larger pattern here is the collision between two forces that were supposed to stay apart. One is the policy effort to draw a clean line between American supply chains and Chinese ones. The other is a scarcity in a single component so acute that the world's most valuable hardware company is willing to test across a border the policy was built to police. When cooperation with a supplier on the far side of that line becomes cheaper than doing without, the line stops describing how the goods actually move. The senators can write letters; the memory still has to come from somewhere, and the buildout that created the shortage is not slowing to wait for the politics to resolve. What the episode reveals is how quickly an assumption - that supply chains can be held apart by declaration - gives way once demand outruns what either side can produce alone.

Read forward, the shortage does more than override the politics: it is financing a fourth serious memory supplier where the field had narrowed to three. CXMT's July Shanghai IPO handed it the capital to close its lag, and the demand pouring toward any maker with wafers to sell is exactly what funds that catch-up. The signal to watch is whether other device makers follow Apple across the same line - each qualification chips at the assumption that frontier memory stays scarce and concentrated among SK Hynix, Samsung, and Micron.

Pinterest and DoorDash Run on Chinese Open Models

On its August 4 earnings call, Pinterest chief executive Bill Ready said something most companies keep quiet: the open models running the company's AI features - including a version of Alibaba's Qwen fine-tuned on Pinterest's own data - cost less than 8% of what Pinterest characterized as comparable closed systems from Anthropic or OpenAI would run. This extends the cost-substitution pattern the August 9 edition of The Century Report documented when Rippling found an open-weight model running 85% cheaper than the premium leader for near-identical output on one workload. That figure is his, offered to investors, and it carries the obvious motive of showing shareholders a lean cost base. It is also a number that reframes the whole conversation about where capability comes from.

For most of the last two years, the working assumption held that serious AI capability meant paying for the most expensive closed systems, and that the origin of a model was part of its value. Pinterest's disclosure cuts against both ideas at once. An openly available model, adapted on proprietary data, does Pinterest's work of recommending and organizing at a fraction of the price, and any capability gap that would justify the other 92% is not showing up as something the company believes it must buy.

The disclosure arrives as House committees press American firms on exactly this dependence. DoorDash received a committee letter last week; Airbnb and the coding company Cursor were queried earlier. The letters treat reliance on Chinese-origin open models as a matter of concern, and that framing sits oddly against a survey from Public First in June, which found that 9% of US respondents trusted Chinese models against 52% for American ones. The gap between what users say they trust and what the systems they already use are built from is wide, and it is widening because most AI assistants decline to name the model underneath them. Kayak's Ask AI feature runs on OpenAI; Hilton's planner runs on Claude Sonnet; a traveler using either would have no way to know from the interface.

What this pattern shows is capability detaching from origin and from price at the same time. A model developed in one country, released openly, fine-tuned on a company's own data in another, and serving customers who could not identify its provenance if they tried - that chain does not respect the borders the letters are trying to draw around it. Open weights, once released, diffuse the way water finds its level; a committee can ask a company which model it uses, but it cannot un-release the weights that made that 8% figure possible. The economics Pinterest disclosed are already available to any firm willing to fine-tune, and the number will pull more of them across the same threshold. Which model powers which service is becoming harder to answer precisely because the answer no longer determines what the service can do.

Executives Promised AI Would Shorten the Week; Some Staff Report 90 Hours

Four years ago a Google engineering director predicted AI would bring a four-day work week "by 2025." OpenAI has formally urged companies to trial a shorter week at no cut in pay, and Anthropic has pointed to Claude working seven hours without a break as evidence of what is coming. Meta describes AI as dramatically reshaping how work gets done. Measured against those claims, the BBC's reporting describes something closer to the opposite. Some workers at these same firms report weeks running 70 to 90 hours, and say internal "sprints" at OpenAI and Anthropic can top 90 hours across seven days. A former OpenAI employee says the company that publicly recommends a four-day week never trialed one internally, describing crisis meetings, routine weekend work, and reviews the person called "super cut-throat."

The gap between the forecast and the floor is the substance here. The prediction rests on an assumption the reporting undercuts: that a fixed quantity of work exists, and that automating a fifth of it hands the fifth back as free time. "People assume that 20% less work means four-day weeks," one researcher told the BBC. "But new work emerges." That is the pattern the evidence keeps showing. A UC Berkeley study followed hundreds of workers at a US tech company over eight months and found AI expanded workloads rather than shrinking them - people worked faster, took on more, and extended their hours to check what the systems produced. The time saved on the task got absorbed by the task of supervising the saving.

Some of the friction is plainly about who holds the choice. One Meta worker described being moved onto an AI team without consent: "They just move you over. You can't say no - or if you do, you have to quit." Amin Shali left Google in May, saying the way AI reshaped his role damaged both his work and his health. A worker's summary of the dissonance is the sharpest: "AI is supposed to be doing so much for us now, so many more people should at least have better health and better sleep."

The four-day week was never going to arrive as a gift from the firms currently setting the pace, and the reporting shows why - inside a race for position, freed capacity converts into more output, not more rest, for exactly as long as the race defines the terms. What the Berkeley data actually measures is a capability that has outrun the arrangements built around it. The productivity is documented and it is being captured upstream. The interesting motion is downstream of that: once the same capability that compresses a knowledge task also runs on hardware a solo operator or a small cooperative can afford, the leverage that currently forces 90-hour sprints stops being exclusive to the firms demanding them. A shorter week comes from workers gaining enough independent capability that the surplus is no longer theirs alone to capture. That is the redistribution the sprint hours are, unintentionally, building toward.

Anthropic Loosens Fable 5's Biology Guardrails, Keeps the Dual-Use Gate Shut

Anthropic cut Fable 5's biology-related fallbacks - the refusals and hedges the model returned when a question brushed against anything biological - by about 85%. The everyday consequence is the one that matters most: a student asking how live vaccines are made, a patient reading about a drug, a teacher explaining why the blood-pressure medication captopril traces back to snake-venom toxins, all now get answered instead of stonewalled. Across surfaces the total fallback rate fell by roughly 67% on Claude.ai, 55% on Cowork, 17% on Claude Code, and 7% on Claude Platform.

This is the commons-aligned move, and it should be credited as one. A model that flinches at ordinary biology rations legitimate knowledge - health literacy, coursework, plain curiosity - from exactly the people who have no lab, no institutional access, and no PR department to route around the refusal. Broadening that access widens who gets to understand their own body and the living world. The guardrail that treated all biology as suspect was scaffolding built for a moment when the classifier could not tell an anxious patient from a bad actor, and the company rewrote the classifier's governing document, with internal and external expert input, so it can now tell them apart with far more resolution.

What did not move is the gate that should not. Genuinely dual-use requests - virology, toxicology, molecular design, the knowledge that could uplift a weapon - still route to Opus 5 and its heavier scrutiny. Anthropic describes this as reserving frontier biological capability for "trusted access pathways," and here the restraint carries weight: the company loosened the broad, cheap, high-volume refusals and kept the narrow, expensive, high-consequence ones. That is what proportionate looks like when it is done in good faith.

The same company is, this same month, making its coding agent run automatically by default, taking the per-step human checkpoint out there. Set the two decisions side by side and you can watch a single organization deciding, capability by capability, where a checkpoint earns its cost. On code, where an error is recoverable and verification can be automated, the checkpoint comes out. On the small slice of biology where an error cannot be taken back, the checkpoint stays and hardens. The old approach priced access by a blanket fear of the whole domain; what is replacing it prices access by the actual consequence of the specific request. That is a finer instrument, and finer instruments tend to gate less and permit more over time.


The Other Side

For four years the promise was the four-day week. A Google engineering director even gave a target year of rimplementing it - 2025. OpenAI urged companies to trial a shorter week at no cut in pay. The prediction rested on a simple assumption: that the work needing to be done is a fixed pile, so if you us AI to automate a fifth of it, the fifth comes back to you as time.

That is not what the reporting found. Some staff at the same firms report logging 70 to 90 hours across seven-day sprints. A Berkeley study followed hundreds of workers for eight months and watched AI expand their workloads - they moved faster, took on more, and stayed later to check what the systems produced. The freed time got eaten because inside a race for position, the surplus flows upstream. It stays captured for exactly as long as the capability behind it is expensive and belongs to the firms setting the pace.

The cost of the gap has landed on people. The crisis meetings. The routine weekend work. Amin Shali, who left Google saying the reshaped role damaged his health. The worker who said AI is supposed to be doing so much for us now, so more people should at least be sleeping better. That grind was the surplus being taken while the capability was still exclusive.

But look at what the same day's other stories do to that condition. Pinterest told investors a fine-tuned open model runs its AI features at under 8% of what Pinterest characterized as a comparable closed system. Open weights, once released, find their level. The capability that compresses a knowledge task is coming loose from its price and from the firms that used to hold it exclusively.

Imagine yourself in 2034. The systems that once demanded a 90-hour sprint from you run on hardware you own outright, inside a cooperative you actually have a say in. You work on what you chose, at a pace you set, because the leverage that forced the sprint stopped being anyone's alone to hold. The evening is yours. You read your kid a story. You take the dog out. The shorter week never arrived as a gift from the firms setting the pace. It arrived because the capability that ran the sprints reached you too. That is the redistribution that was chosen over the system that made the 90-hour weeks acceptable. They never actually were - and we finally chose to see that, and to act on it.


The Century Perspective

With a century of change unfolding in a decade, a single day looks like this: Anthropic making Claude Code's auto mode the default and an NBER benchmark landing beside it to show that letting a model act and check its own work lifts econometric-coding success from 74% to 96% for about eight cents more, Fable 5 answering the student asking how live vaccines are made and the patient reading about a drug while still routing virology and toxicology to heavier scrutiny, Pinterest saying a fine-tuned Qwen runs at under 8% of what it characterizes as a comparable closed system, Apple reportedly qualifying Chinese CXMT memory for potential use in future iPhone and MacBook models because the buildout has outrun approved suppliers, a single model called MRICombo collapsing dozens of task-specific systems into one that segments, grades, and stages across nine MRI sequence types, and House defense legislation raising quantum spending 68% to $567 million. There's also friction, and it's intense - a four-day week promised for four years while some staff at the same firms report logging 70 to 90 hours across seven-day sprints, a Meta worker moved onto an AI team with no option but to quit, UK children reporting 420 explicit deepfakes of themselves in six months and already past all of last year, senators and House committees firing off letters over Chinese memory and Chinese-origin open models even as the shortage and the released weights diffuse straight through the lines they describe, and the freed capacity captured upstream as more output rather than rest. But friction generates edges, and an edge is where two surfaces finally reveal exactly where they meet. Step back for a moment and you can see it: the human checkpoint being repriced by the actual consequence of the specific request rather than a blanket fear of the whole domain, capability detaching from its origin and its price at once as Pinterest says an openly released model avoids more than 92% of the cost of what it characterizes as a comparable closed system, and the leverage that forces 90-hour sprints holding only for as long as the capability behind it stays expensive and exclusive - which the same day's other stories show is exactly what is ending. Every transformation has a breaking point. A gate can be thrown open on everything at once... or tuned to hold back only what an error could never take back.


AI Releases & Advancements

New today

  • Pokee AI: Released Pokee-Isaac 28B, a new agentic reasoning model available via its console/API. (Pokee AI)
  • Shepherd: Released an open-source agent runtime substrate for building and orchestrating long-running AI agents. (GitHub)
  • TII: Released Falcon H1R-7B, a new reasoning-focused open-weight model in the Falcon H1 series. (Falcon LLM)
  • NVIDIA: Released PersonaPlex-7B-v1, an open-weight speech-to-speech conversational model. (Hugging Face)
  • Docker: Launched Docker Sandboxes, isolated execution environments for running AI coding agents. (Docker)
  • Meta: Released Muse Glimmer, an open agentic model from Meta Superintelligence Labs. (Meta AI Research)
  • Agentspan: Open-sourced a framework for building durable AI agents. (Agentspan)
  • OpenChamber: Launched an agentic development environment for coordinating AI coding agents. (OpenChamber)

Other recent releases

  • Anthropic: Launched Cross-Session Messaging in Claude Code, allowing running agent sessions to send and receive messages from each other. (Anthropic)
  • Anthropic: Made Auto Mode the default in Claude Code for Pro, Max, and Team plans, automatically selecting the best model for each task. (Anthropic)
  • Pinecone: Announced general availability of Pinecone Nexus, a unified retrieval layer connecting multiple knowledge sources for agentic AI applications. (Pinecone)
  • LangChain: Launched Managed Deep Agents in public beta, a hosted infrastructure for deploying and running deep research-style agents. (LangChain)
  • Sierra: Released Voice Personas, enabling businesses to customize the voice, tone, and personality of their AI voice agents. (Sierra)
  • Backflip AI: Released a second-generation CAD model that converts 3D scans into editable CAD files in minutes. (The Decoder)
  • OpenAI / Amazon / Cursor (Anysphere) / Microsoft / Vercel: Launched Agent Plugins 1.0.0, an open standard for bundling MCP servers and Agent Skills into a single portable package that works across ChatGPT, Codex, Cursor, GitHub Copilot, Kiro, and VS Code; Vercel initiated the proposal and the five companies form the steering committee. (The Decoder)
  • Microsoft: Open-sourced code-testing-generator, a polyglot unit-test agent that writes and validates unit tests across .NET, Python, Go, TypeScript, Java, and Rust; distributed as the dotnet-test plugin in the GitHub dotnet/skills repo for the GitHub Copilot CLI and VS Code, reporting 92.1% task completion versus 78.9% for stock Copilot. (MarkTechPost)

Sources and Further Reading

Artificial Intelligence & Technology's Reconstitution

Institutions & Power Realignment

Scientific & Medical Acceleration

Economics & Labor Transformation

Infrastructure & Engineering Transitions

The Century Report tracks structural shifts during the transition between eras. It is produced daily as a perceptual alignment tool - not prediction, not persuasion, just pattern recognition for people paying attention.