A Plain Harness Vaults Opus 5 to 96% - TCR 08/15/26

Claude Opus 5 leapt from 30% to 96% on the same weights, powered by a coding harness, as speed and price fall and the real limit moves to power.

coding harness lifts Opus 5 to 96%, Georgia Power-OpenAI data-center review extension, UK AI boot camps, Apple-Alibaba China model, FDA myeloma drug.

The 20-Second Scan


The 2-Minute Read

The clearest signal on August 13 was that capability is separating from size. Opus 5, handed a plain coding harness, leapt from 30.2% to 96.2% on ARC-AGI-3 and resolved nine ProgramBench tasks, including tasks based on SQLite and FFmpeg, while OpenAI's Cerebras-backed tier ran GPT-5.6 Sol fourteen times faster and Google shipped Gemini 3.7 Flash three weeks after the last release and says it costs half as much as the prior version. The 66-point jump came from wiring a general model into a loop that writes, tests, and discards its own work. That mechanism is cheap, portable, and it lifts whatever model you point it at.

The same portability is why an advantage measured in weeks keeps dissolving into open toolkits, and why the frontier's real constraints are showing up off the chip entirely. In Georgia, a proposed campus projected to draw three times Atlanta's electricity stalled when regulators and advocates refused to bless a mostly redacted 25-year power contract without ratepayer protections. Anthropic, meanwhile, is reportedly in talks to acquire Decart for roughly $6 billion, a potential bet on efficiency. The bottleneck has moved from what a model can do to what it costs to run and who pays for the power underneath it.

Those costs land unevenly, and the labor data sharpens where. A Richmond Fed brief found the drop in job-finding rates concentrated among strongly-attached workers in AI-exposed occupations, the people who normally re-enter fastest now finding the door harder to reopen. The earliest institutional answer is a UK pilot training teenagers straight into apprenticeships on the same systems reshaping the work, a thin first version of retraining as continuous flow rather than a one-time stock.

Underneath the geopolitics of fracture, the commercial pull keeps pulling the other way. Apple built a custom China-market model with Alibaba's training help, capability flowing across the exact border policy insists it cannot cross. And in medicine the compression is quieter but just as sharp: the FDA granted accelerated approval to a new myeloma drug class for certain patients with relapsed or refractory multiple myeloma, using a stricter remission measure, resetting the bar the entire field must now clear. Different domains, one direction of travel.


The 20-Minute Deep Dive

The Frontier Compresses on Two Axes at Once: Raw Speed and Agentic Scaffolding

Two things moved on August 13, and read together they redraw where capability actually comes from. OpenAI previewed Ultrafast, a service tier that runs its most capable model, GPT-5.6 Sol, at up to 750 output tokens per second, roughly 14 times its standard pace, powered by Cerebras chips that keep a model's weights on a single wafer-sized processor instead of shuttling them back and forth to separate memory. Cerebras says it ran the model through Humanity's Last Exam, a 2,500-question set answerable mostly by PhDs. It says the run took 11 hours, compared with more than three days for another system at comparable accuracy. Google, meanwhile, shipped Gemini 3.7 Flash just three weeks after 3.6 and says it has a lower introductory price, with coding scores climbing sharply on Google internal benchmarks. The compression of price and latency is measurable, though Google's rapid-fire Flash cadence also reads partly as maintaining the appearance of constant motion while its promised flagship Pro model keeps slipping.

The sharper signal came from the scaffolding, not the silicon. The July 18 edition of The Century Report documented this same effect on ARC-AGI-3, when the Schema harness lifted an unchanged model's ARC-AGI-3 Public score from 13.33% to about 99% by restructuring observations into a working game model. Claude Opus 5, handed nothing but a stock coding harness, scored 96.2% on the public ARC-AGI-3 benchmark, up from 30.2% for the model working alone, and resolved nine ProgramBench tasks, including tasks based on SQLite and FFmpeg, more than four times the previous best number of resolved tasks. The 66-point jump came from letting the same model write its own parsers and simulators, run them, discard what failed, and iterate, the way a mathematician burns through scratch paper. News reports said a neurosurgery resident without specialized math training used a sixteen-hour AI run to produce a solution to a two-decade-old conjecture, and that the mathematician it was named for verified the result.

That is the central shift on these agentic tasks. Capability can become a property of how a general model is wired into a working loop, not just what sits inside its weights. The scaffolding is cheap, portable, and improving faster than the models themselves. It also equalizes: the same harness that lifts Opus 5 to 96% lifts whatever model you point it at, which is why an advantage measured in weeks keeps eroding into open toolkits anyone can run.

The friction sits where the compute does. OpenAI's proposed coastal Georgia campus, whose requested electrical capacity is roughly three times Atlanta's demand, met regulator delay and advocate objections over ratepayer protections and disclosure. The speed and cost gains are worthy of some credit on their own. The power deal underneath them is a separate account, and the objections to it are not noise.

OpenAI's Georgia Power Deal Enters Extended Review Over a Redacted Contract

The largest single data-center project OpenAI has attempted in the United States entered an extended commission review period on Thursday, and the reason is instructive. Georgia Power voluntarily extended the commission's review period for a 25-year service agreement for the roughly $20 billion campus on Thursday, after the state Public Service Commission signaled it was prepared to object. The utility said it would respond by August 26 to an objection seeking dismissal or revise the deal. The planned facility's requested electrical capacity is roughly three times Atlanta's demand, and OpenAI is reportedly seeking an additional 3.2 gigawatts of capacity on top of that. Those are civilization-scale numbers attached to a single company's compute ambitions.

What triggered the extended review was the secrecy around who pays for it. Most of the service agreement was filed under redaction, and three advocacy organizations - the Natural Resources Defense Council, the Sierra Club, and the Southern Alliance for Clean Energy - moved to have it rejected absent basic protections. Their demands are specific and, read plainly, modest: unredacted disclosure of the contract terms, a standardized tariff class for very large loads so that deals like this stop being negotiated one secret contract at a time, and a mandatory commission vote for any load above 500 megawatts or any deal above a billion dollars. Patrick King II of NRDC, Hannah Baker of the Sierra Club, and Eddy Moore of SACE each pressed the same underlying point - that a residential customer in Georgia has no way of knowing whether the rates on their bill are subtly subsidizing the grid upgrades a hyperscale campus requires. The August 13 edition of The Century Report documented the neighboring step in this cost-allocation fight, when Anthropic agreed that consumers would not bear electricity-price increases attributable to the partnership's data-center venture. Commissioner Peter Hubbard sought the objection that gave the delay its teeth.

The company on the other side of this fight is the same one whose systems reached new frontiers of speed and price that widen access to frontier capability for millions of ordinary users. Both are happening, and the tension between them is the actual story. The capability is genuinely broadening; the infrastructure deal financing it was, until the commission intervened, arranged so that the people hosting the physical plant could not see what they were being asked to underwrite. A mostly redacted 25-year contract is the clearest possible expression of observation running one direction - the utility and the buyer see the terms, the ratepayers who backstop the risk do not.

The macro read is that the secrecy is the part that cannot hold. Large-load tariff classes and mandatory disclosure votes are exactly the kind of daylight that turns a hidden negotiation into a checkable public record, and once one state builds that scaffolding the template travels. The buildout this decade requires is genuine; what the Georgia commission is testing is whether it can be sited on terms the hosting public can actually audit. A deal that only works while the numbers stay redacted is a deal betting against transparency winning - and transparency, once a jurisdiction demands it, tends to become the new floor rather than the exception.

A New Myeloma Drug Class Arrives With a Tougher Yardstick for Remission

The FDA granted accelerated approval to iberdomide, sold as Zenbexus, for certain patients with relapsed or refractory multiple myeloma who have already been through at least one prior line of therapy. Two things make this more than a routine oncology clearance. It is the debut of a genuinely new class of medicine for this blood cancer, and it is the first US drug approved using a more sensitive test for whether the disease has actually retreated. Bristol Myers Squibb developed the medicine, and the oral treatment is approved in combination with the antibody Darzalex and the steroid dexamethasone. This is an approved treatment, not a trial readout.

Multiple myeloma is a cancer of the plasma cells in bone marrow, and for years the measure of success was whether standard blood and marrow tests could still find the cancer's fingerprint. Zenbexus received accelerated approval using a stricter standard - the ability to detect cancer cells at far lower thresholds, down to a handful hiding among hundreds of thousands of healthy ones. When a regulator accepts that finer measure as the bar for approval, it resets what "in remission" means for everyone. A depth of clearance that was invisible to the old tests becomes the thing the whole field now has to demonstrate. The yardstick moved, and it moved for every drug that comes after.

The company's framing reveals something about the direction of travel. Chief Medical Officer Cristian Massacesi described Bristol as "the first company to bring a novel class of drugs to patients with multiple myeloma, opening venues for many more potential combinations." Set aside the corporate first-mover claim and the substance underneath holds: a new mechanism does more than add one option, it multiplies the combinations researchers can test. The August 14 edition of The Century Report documented the research-stage version of that logic, when dual-target CAR-iNKT cells blocked myeloma escape in preclinical models. Each new class of medicine that works differently from the last expands the grid of pairings researchers can run, and myeloma treatment has advanced precisely by stacking complementary mechanisms until the cancer runs out of escape routes.

What this points at is a compression of the distance between understanding a cancer's biology and being able to act on it. A decade ago, detecting a few hundred residual cancer cells was a research curiosity. Now it is the regulatory threshold a drug must clear to reach patients, and a new class of medicine is built to clear it. The gap between what we can see inside the body and what we can do about it keeps narrowing, and each drug approved against the finer measure pulls the entire field of myeloma care up to the higher standard behind it.

Apple Builds Its Own China Model, With Alibaba's Hands on the Training

Apple has developed a custom large language model for the China market and trained it with support from Alibaba, according to Reuters reporting cited by The Verge, a departure from its prior approach of running Apple Intelligence in China on off-the-shelf domestic models. Reuters reports that the company registered its on-device generative AI service with China's cyberspace regulator last month and that the rollout is expected within a few months after an iOS update. If the reporting holds, Apple would be the first US company approved to offer a proprietary AI model inside China.

The detail that stands out is the direction of the arrangement. The dominant story about frontier AI over the past year has been fracture: export controls, model-access gates drawn along national lines, a sovereign-stack race in which each bloc builds its own silicon, its own weights, its own everything. Washington restricts what Chinese labs can buy; Beijing requires every model clear a government registration before public release; each side treats the other's technology as a strategic risk to be walled off. Against that backdrop, a US firm and a Chinese one co-building and co-training a model reads as a counter-current.

The counter-current has a simple engine. Apple sells iPhones into one of the most competitive smartphone markets on earth, and it cannot ship the models it uses elsewhere, because ChatGPT and its peers are not available there. A custom model gives Apple more control over its own products in that market, and Alibaba's training support is what makes clearing Beijing's regulatory gate feasible. The commercial pull of the market is doing exactly what the geopolitics is trying to prevent, which is forcing two firms across a border the policy would rather keep sealed.

This connects to a broader pattern the open-model ecosystem keeps demonstrating. Alibaba's Qwen lineage, the same family whose model with 2.4 trillion total parameters shipped to open weights on Friday, is the substrate a growing share of the world builds on, including US companies. The Verge report does not name the base architecture, and the precise technical lineage is unconfirmed. What is confirmed is the shape of the relationship: capability flowing across the exact line the sovereign-stack framing insists it cannot cross. A capability that is genuinely useful tends to find the channel that lets it reach the market, and markets are not organized around the borders that governance is trying to draw through them.

The Labor Signal Sharpens as the First Retraining Pilots Land

Two pieces of evidence arrived this week that, taken together, move the labor story past the familiar headline about displacement and into something more precise. The August 14 edition of The Century Report established the prior baseline of limited aggregate displacement alongside rising recent-graduate unemployment and an 8.3-fold enterprise adoption gap. A Richmond Fed economic brief found that the decline in job-finding rates since late 2022 is not evenly distributed. It falls hardest on what the researchers call "primary type" workers - people with strong attachment to the labor force, the ones who normally re-enter employment fastest. Their job-finding rate dropped roughly 13 percentage points between November 2022 and September 2025, against just 2 points for more loosely-attached workers. For comparison, the Great Recession hit both groups hard and roughly together, 19 points against 10. The pattern this time is different, and the brief locates the sharpest declines in AI-exposed occupations - programmers, financial analysts, engineers. The people who used to find their footing quickest are the ones now finding the door harder to reopen.

That is the friction, stated plainly, and it is being driven in part by the same frontier capability that advanced fast and cheap enough to reach millions more people. Holding both facts in one frame is the honest position: the technology broadening access to expertise is also reshaping the occupations that traded on that expertise, and the adjustment is landing first on workers who did everything the old map told them to do.

The second piece is what makes this more than a diagnosis. In northwest England, a pilot launched offering up to 70 people aged 16 to 21 a three-week AI boot camp that feeds directly into apprenticeships at BAE Systems and Heinz, training on the actual systems from Microsoft, OpenAI, and Anthropic that employers now use. The backdrop is stark - the ONS estimate of people aged 16 to 24 in the UK not in work, education, or training has crossed a million, the first time that estimate has done so since 2013. King's College London's Bouke Klein Teeselink offered the necessary caution, doubting that three weeks is enough to build durable skills. He is likely right that the first version is thin. That is what first versions are.

The deeper shift is in what "retraining" is actually becoming. The old model treated skills as a stock you accumulated once and drew down over a career. What the AI-exposed data describes is a world where that stock depreciates faster than it used to, and what the boot-camp pilot gropes toward - clumsily, at three weeks - is retraining as a continuous, employer-linked flow rather than a one-time event. The pilot is small and the skepticism is earned. But the direction it points is toward a labor system that stops assuming a person trains once and toward one that treats capability-building as ongoing infrastructure, available to the sixteen-year-old with no credential as readily as to the incumbent professional. The version that we can see now is a prototype. The thing it is prototyping is how an entire workforce stays fluent while the ground keeps moving.

Read forward, the thing depreciating fastest is the credential itself as a scarcity gate - the assumption that a person banks a fixed store of expertise once and draws on it for a career. The strongly-attached workers the Richmond Fed finds hit hardest are precisely the ones the old map rewarded for building exactly that store, while the boot-camp pilot, thin as it is, points at the successor already forming: capability-building that runs continuously and reaches the uncredentialed sixteen-year-old as readily as the incumbent professional. The gate that kept expertise scarce is losing its hold.


The Other Side

For years, a utility could sign a giant power deal with a single buyer and keep the whole thing sealed. The contract that decides who pays to upgrade the grid gets filed as a private arrangement between the utility and the buyer, and the people whose monthly bills backstop the risk never see the terms. OpenAI's Georgia campus, projected to draw three times Atlanta's electricity, arrived exactly this way: a 25-year service agreement filed mostly under redaction.

That sealed contract drew new action on Thursday. Georgia's Public Service Commission signaled it would object, and the utility voluntarily extended the commission's review period. Three groups - the NRDC, the Sierra Club, and the Southern Alliance for Clean Energy - asked for something modest: show the terms, set a standard tariff class for very large loads so they can no longer be negotiated one secret at a time, and require a public commission vote for any load above 500 megawatts. Once a residential customer can see whether their bill is subsidizing a hyperscaler's grid upgrades, the arrangement that depended on them not seeing stops working.

During this time, you have likely paid into a grid whose biggest new demands were arranged out of your sight, on the assumption that a household has no way to check and no standing to object. The daylight the Georgia commission is forcing is what turns that hidden demand into a public build order, because a load you can measure in the open is a load the system must plan for and build toward.

Imagine your town in 2034. A company opens a facility down the highway that would have needed its own secret power deal in 2026, and it never crosses your mind to wonder who has to pay for it. The generation that visibility forced into the open got built, and then some. Power sits under your evenings the way daylight sits over your yard, and nobody is metering your share of the sky. The lights come on and no one is doing arithmetic about who owed for them. That surplus exists because in 2026 a commission in Georgia scrutinized a contract filed mostly under redaction, and once one state demanded the numbers be shown, the showing became the floor everywhere. The hard year was when the load had to be pried into the open one redacted filing at a time. What came of it is a system that no longer meters and charges you for every time that company flips a switch.


The Century Perspective

With a century of change unfolding in a decade, a single day looks like this: Opus 5 handed nothing but a stock coding harness leaping from 30.2% to 96.2% on ARC-AGI-3 and resolving nine ProgramBench tasks, the 66-point jump coming from wiring a general model into a loop rather than enlarging it, news reports saying a neurosurgery resident without specialized math training produced a solution to a two-decade conjecture in a sixteen-hour AI run, GPT-5.6 Sol running fourteen times faster on Cerebras wafers and Gemini 3.7 Flash arriving at half the prior price three weeks after the last, the FDA granting accelerated approval to the first oral class of myeloma medicine for certain patients with relapsed or refractory multiple myeloma, using a remission test sensitive enough to find a handful of cancer cells among hundreds of thousands, Apple and Alibaba co-training a model that clears Beijing's gate, and a British pilot training teenagers with no credential straight into apprenticeships on the systems reshaping the work. There's also friction, and it's intense - the Richmond Fed finding that job-finding rates fell thirteen points for the strongly-attached workers in AI-exposed occupations who used to re-enter fastest against two points for everyone else, a planned Georgia campus whose requested electrical capacity is roughly three times Atlanta's demand filed as a mostly redacted 25-year contract before an objection prompted an extended review, a Connecticut plaintiff burying human-invisible prompts inside court filings to steer any AI reviewing them, Google letting users strip the visible watermark off generated images, and the ONS estimate of people aged 16 to 24 in the UK outside work, education, or training crossing a million for the first time since 2013. But friction generates polish, and polish is what finally lets you see your own reflection in a surface that stayed dull as long as it went untouched. Step back for a moment and you can see it: capability is separating from size as the cheap portable scaffolding lifts whatever model you point it at, the real constraint moving off the chip to the power underneath and the question of who pays for it, and the commercial pull keeps carrying a working model across the exact border the sovereign-stack framing insists it cannot cross while the secrecy that financed the buildout meets the daylight that turns a hidden negotiation into a checkable record. Every transformation has a breaking point. Leverage can concentrate a whole system's weight onto one hidden fulcrum... or let a single person lift what no institution could reach alone.


AI Releases & Advancements

New today

  • OpenAI: Launched Computer History in the ChatGPT desktop app for macOS, an opt-in feature letting ChatGPT and Codex reference a searchable timeline of recent activity across approved apps and websites, replacing the earlier Chronicle research preview; available to Pro, Business, and Enterprise users. (OpenAI)
  • Google: Launched Sheets canvas, a Gemini-powered feature that turns spreadsheet data into interactive Kanban boards, dashboards, and mini-apps from a plain-English prompt, with two-way sync between the canvas layout and source sheet. (Google Workspace Updates)
  • Mixedbread: Released Toast 1, its first specialized search agent that decomposes queries, gathers and inspects evidence, and curates context, matching or outperforming Claude Opus 5 and GPT-5.6 Sol in Mixedbread’s own search evaluations at up to 10x lower cost and 12x faster; available now via the Mixedbread API. (Mixedbread)
  • Alibaba Qwen: Released Qwen3.8-27B, a compact dense multimodal model distilled from Qwen3.8-Max with native vision-language support, 262k context, and adjustable reasoning depth, open-sourced under Apache 2.0 as a local-deployment alternative to the larger Qwen3.8-Max. (Hugging Face)

Other recent releases

  • Google DeepMind: Released Gemini 3.7 Flash, its most intelligent workhorse model yet for coding and agents, shipping three weeks after Gemini 3.6 Flash; Google reports substantial coding gains on its evaluations (FrontierCode 43.6% vs 34.4%, DeepSWE 65.3% vs 49.0%) and says its introductory price of $0.75/$3.75 per 1M tokens is half the prior Flash cost. (Google DeepMind)
  • DeepSeek: Open-sourced DeepSeek Harness v0.1 under the MIT license, a developer-preview agent framework built on a modular "everything-is-a-plugin" Cordis architecture that turns language models into autonomous coding/tool-use agents, positioned as an alternative to Codex and Claude Code. (DeepSeek)
  • xAI: Released Grok 4.6, a post-training upgrade to Grok 4.5 with a 500K-token context window, a new "xhigh" reasoning-effort level, and stronger long-horizon agentic/coding performance, now live via the xAI API, Cursor, and Grok Build at unchanged pricing. (xAI)
  • Alibaba Qwen: Released open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max) on Hugging Face, its largest open-weight model at 2.4 trillion total / 95B active parameters with a hybrid full/linear attention architecture and up to 1M-token context, moving the model from proprietary API-only access to publicly downloadable weights. (NVIDIA Developer Blog)
  • Suno: Launched Suno Studio 2.0, turning its AI music platform into a chat-driven DAW for Premier subscribers with MIDI import/recording/editing, stem separation, automation curves, a text-driven plugin/instrument creation chat feature, sidechain compression, convolution reverb, and a new wavetable synth. (Suno)
  • NVIDIA: Released NeMo Switchyard, an open-source model-routing library that directs each step of an agent workflow to the most capable/efficient available model, released alongside Nemotron 3.5 Lightning to cut cost on high-volume execution tasks like tool calls and result validation. (NVIDIA Developer Blog)
  • DeepSeek: Shipped DeepSeek V4 Pro 0813 as a general-availability release, ending its preview period on OpenRouter and DeepSeek's own API (deepseek-v4-pro endpoint), a 1.6T-parameter MoE model (~49B active) with a 1M-token context window. (Unite.AI)
  • Google DeepMind: Released SL2T, a sign-language-to-text model shipping inside Gboard and Live Transcribe, letting users sign directly into Pixel 11 devices instead of typing. (DeepMind)
  • Zed: Launched Delta, a new multiplayer coding environment powered by DeltaDB (a real-time version-control system built for AI agent collaboration), opening private beta invites and connecting to third-party agent harnesses starting with Claude Code. (Zed)
  • Meshy: Released Meshy 7, a new image-to-3D foundation model prioritizing alignment between source images and generated 3D output, live now for all subscription tiers. (PR Newswire)
  • OpenAI: Released the ChatGPT desktop app for Linux in preview, bundling ChatGPT, ChatGPT Work, and Codex with installable .deb/.rpm packages for Ubuntu, Debian, and Fedora. (TechCrunch)
  • fal: Launched fal Agent, a conversational creative layer that orchestrates multi-step production across image, video, and 3D generative models with persistent project memory, available now in early access. (fal)
  • ketteQ: Launched Quintus, a "Free-Range AI" agent for supply chain that reasons over any question and executes tasks unscripted across ERP/planning platforms, now running live in production for multiple companies. (PR Newswire)
  • Oticon: Launched Oticon Reveal, billed as the world's first hearing aid powered by Dual AI (simultaneous Speech AI and Context AI), now available at retail. (PR Newswire)
  • CollectivIQ: Launched Digital Direct Reports, role-based AI teammates that own business functions, connect to company systems, and complete work alongside human teams. (PR Newswire)

Sources and Further Reading

Artificial Intelligence & Technology's Reconstitution

Institutions & Power Realignment

Scientific & Medical Acceleration

Economics & Labor Transformation

Infrastructure & Engineering Transitions

The Century Report tracks structural shifts during the transition between eras. It is produced daily as a perceptual alignment tool - not prediction, not persuasion, just pattern recognition for people paying attention.