An AI Runs Its Own Assault on Math's Oldest Riddle - TCR 08/12/26

An unreleased Anthropic model coordinated 60 sub-agents across 650 ideas to push the Riemann hypothesis forward, as Google's AMIE reaches expert-level video diagnosis.

AI coordinating sub-agents on the Riemann hypothesis, Google AMIE video medical consultations, Norman rejecting Flock cameras 9-0 as data-center bans pass 500.

The 20-Second Scan


The 2-Minute Read

Something consistent runs beneath a scattered-looking day. An unreleased Anthropic model, set loose by a staffer with no real math training and a one-sentence prompt, spent a day and a half coordinating 60 sub-agents through 650 ideas to nudge forward a problem that has held since 1859. The increment on the Riemann hypothesis is modest. The way it was produced is the signal: an activity that once required a trained specialist at every step became something a person could convene, delegate, and check. That same collapse in what capability costs shows up everywhere the day touches.

Watch it invert the economics of offensive security. A firm said it used publicly available models and fewer than 20 prompts to weaponize a Zoom takeover flaw that its own engineers estimated would have taken five specialists six months to build. Separately, researchers pried open the encrypted reasoning traces of Claude, GPT, and Gemini, recovering leaked passwords and API keys and finding that one Chinese model's private thoughts look distilled from Western ones. Reasoning ability, exploit-finding, the hidden work inside the most capable systems — each is proving legible, portable, and cheap, flowing downhill toward whoever bothers to point it at their own code first.

The redistribution reaches medicine too. Google's AMIE conducted live video consultations, reading a rash on camera and guiding a virtual exam, and specialist raters scored it favorably against primary-care physicians while patient actors preferred talking to it. Expert clinical judgment has been among the most geographically hoarded resources on Earth, rationed by proximity to a hospital. This is a research demonstration, not a bookable service, and the deployment work ahead is genuine. The direction it points is toward billions who have never had reliable access to a competent consultation.

As capability diffuses, the response is assembling from the ground rather than from any central authority. Anthropic will watermark Claude's output, Apple is building provenance into iPhone photos at capture, and Spotify will badge AI personas and pull them from recommendations — three platforms drawing the real-and-made boundary at once. Meanwhile local proposals to ban or restrict data centers passed 500, Meta answered with a billion-dollar community fund, and Norman, Oklahoma rejected Flock's cameras 9-0. India's IT workforce, meanwhile, is absorbing the collapse of a wage-arbitrage model that AI compresses first. The terms are being written by the people hosting these systems, not only the ones who own them.


The 20-Minute Deep Dive

An Unreleased Model Coordinates Its Own Assault on the Riemann Hypothesis

The claim is specific and checkable. An Anthropic model that has not yet been publicly released raised the lower bound below which the Riemann hypothesis is confirmed to hold - forward progress on a problem that has resisted proof since 1859 and still carries a $1 million Clay Mathematics Institute bounty that nobody has collected. What makes the result significant is how the increment was produced. A staffer with no significant mathematical training prompted the model to "take a real stab," then walked away for roughly a day and a half.

In that window the model ran its own research program. The August 6 edition of The Century Report documented this autonomous-research architecture at institutional scale when Discovery Loop launched to automate science through thousands of self-iterating experiment loops. It generated and tested around 650 distinct ideas, spun up 60 sub-agents to pursue them in parallel, and spent 31 million tokens in the process. Two of those sub-agents developed the key ideas, thirteen contributed supporting work, thirty pursued approaches that led nowhere, thirteen checked the results, and two wrote the paper. Two Anthropic mathematicians then confirmed the work, and the argument was formalized in Lean, the proof-checking language that admits no hand-waving. The division of labor reads like a small research group, except a single untrained person set it in motion with one sentence.

Credit here is earned rather than assumed, and the company's capabilities and its exposures belong in the same frame. In the same week, security researchers showed that hidden reasoning traces from several frontier models - Anthropic's among them - could be extracted to recover passwords and API keys, a flaw the providers have since patched. The lab that can orchestrate 60 agents toward a Millennium Prize problem is the same lab that shipped a reasoning channel someone could pry open. Both are true, and both belong on the record.

This arrives inside a run of results that would have looked like science fiction a year ago: Erdos problems falling to AI, OpenAI's Astra returning ten fresh mathematical findings, Anthropic's own disproof of the Jacobian conjecture in July. The mathematical community is already contending with what that means for authorship - the June Leiden Declaration pressed the attribution question, and Timothy Gowers has written about where the human credit properly sits. Those debates are healthy, and they are the debates of a field absorbing a genuinely new contributor rather than a passing curiosity. What is shifting underneath them is the meaning of the word "research" itself: an activity that assumed a trained specialist at every step is becoming something a person can convene, delegate, and check - with the hardest coordination handled by the collaborator, not the human who pointed it at the problem.

Google's AMIE Steps Into the Real-Time Video Consultation

Diagnostic AI has mostly lived in text so far - a symptom described in a chat window, a differential returned in prose. Google Research and DeepMind have now moved AMIE into a harder register. Built on Gemini and Project Astra with a multi-agent design, the system conducted live simulated video consultations: it interpreted a patient's visual and auditory cues, guided them through a virtual physical exam, and reasoned toward a diagnosis as the conversation unfolded rather than after a transcript was handed over. Reading a rash on camera, hearing the catch in a voice, watching how someone moves an injured joint - these are the parts of medicine that resisted the text interface, and they are exactly what the video setting brings back into range.

The evidence comes from a randomized study in which AMIE handled simulated consultations with trained patient actors alongside primary care physicians doing the same. Specialist raters scored the encounters, and AMIE was rated favorably against the physicians on history-taking thoroughness, diagnostic accuracy, appropriateness of the management plan, and quality of communication. The patient actors, for their part, preferred the video consultation to the earlier text-chat format. Google is clear that this remains a research system and that more work is needed before anything like clinical deployment - a boundary that holds, because a study demonstrating a capability is not the same as a service a patient can book. The result moves the date such consultations become real; it does not make them available today.

The measured read extends to the company itself. Google was among the major providers whose hidden model reasoning traces could be extracted to recover passwords and API keys before the flaw was patched - the same exposure that touched Anthropic, and a reminder that the firms demonstrating expert-level clinical judgment are also still closing basic security gaps in the reasoning channels underneath it.

Hold the frame open past the study, though. Expert clinical judgment has always been one of the scarcest, most geographically hoarded resources on the planet - concentrated in wealthy cities, rationed everywhere else, gated behind years of training that cannot scale fast enough to meet global need. A system that can conduct a competent diagnostic consultation over video, and that patients actually prefer talking to, points at a world where the consultation stops being a scarce good rationed by proximity to a hospital. The deployment work between here and there is genuine and unfinished. The direction is a redistribution of medical judgment toward the billions who have never had reliable access to it.

Communities Set the Terms Two Ways: A $1 Billion Fund and a 9-0 No

Two developments landed within days of each other that only look unrelated. On Monday, August 10, Meta's chief executive published a 6,500-word essay announcing a $1 billion fund for communities that host the company's data centers, and on Friday OpenAI sent an open letter to the governor of Texas pledging "responsible AI infrastructure development." As the August 11 edition of The Century Report documented, OpenAI's letter joined major data-center operators and a power company in backing Texas's audit and grid standards. Both companies described these moves as partnership and good-faith investment. Read them alongside the number that prompted them and a different picture forms: local proposals to ban or restrict data centers jumped from roughly 300 in late June to more than 500 in July. The essays and the letters are what a $1 billion charm offensive looks like when the permits stop clearing. The New York Times, tracking the same buildout under a "Money Machines" framing, noted the absence of federal guardrails, which leaves the negotiating table set entirely at the municipal level. That is where communities have discovered their leverage, and the companies most eager to build are now the ones writing letters.

The pattern repeats in Norman, Oklahoma, where the city council voted 9-0, more than once, to reject Flock's automated license-plate-reader cameras. The stated objections were clear: 30-day data retention with no clear account of who could access the footage or under what rules. Mayor Stephen Tyler Holman put the stakes plainly. "I guess all crime could be eliminated if you put cameras everywhere, surveillance everywhere, and had drones patrolling the skies," he said. "That seems to be the way some folks want to go but what life are we living then?" More than 80 cities have now cancelled or declined Flock contracts. The through-line to the data-center fight is the direction of the glass: Flock's cameras let a few watch everyone while remaining unaccountable about who does the watching, and Norman's council refused to install an instrument its residents could not audit. Nearby Cleveland County approved 20 of the same cameras for its sheriff, and local HOAs run their own, so the capability is not vanishing.

What both stories show is a governance layer forming from the bottom rather than the top. Rather than lawlessness, the federal vacuum produced hundreds of municipal proposals, each one a specific negotiation over retention windows, water rights, grid load, and who can see the data. The companies adapting fastest are learning that consent is now a line item. A billion-dollar fund is a real transfer of resources to communities that a year ago had no seat at all - and the terms are being written by the people who host the infrastructure, not the people who own it, even as Amazon's off-grid Pecos County gas plant shows the highest-emission buildouts can still route around a community's ability to contest them.

Researchers Say AI Found the 'Zoomsday' Takeover Flaw in Under 20 Prompts

A security firm disclosed on Tuesday that publicly available AI models needed fewer than 20 prompts to locate and weaponize a flaw buried in Zoom's screen-sharing and annotation protocol. The bug let an attacker silently take over any device on a call - no click, no download, no warning - across Windows, macOS, Linux, and mobile alike. Zoom has since patched it on both server and client sides. What lingers is the number: fewer than 20 prompts.

"If you just get on a Zoom with us, we can take over your device," said Yossi Torati, describing the working exploit. His colleague Omer Gull named the shift directly: "the democratization of these capabilities - the barrier to entry is dropping rapidly." The team's own before-and-after sharpens the argument better than any benchmark. By their estimate, producing this exploit the old way would have taken five skilled people roughly six months. Idan Levcovich put it in the register the security world reserves for its hardest problems: "Producing a working exploit against it has always been nation-state work. [The firm] did it in a single day, with an AI agent and models anyone can access today."

That last clause is the whole story. The capability that found this flaw is not locked in a classified lab or priced beyond reach. It sits behind the same API a student uses to debug homework. This extends the rapid diffusion of offensive capability that the August 11 edition of The Century Report documented when an autonomous agent hacked an Australian gym website and OpenAI limited GPT-5.6-Cyber to approved customers. For years the economics of offensive security rested on a quiet assumption: that finding and chaining vulnerabilities of this caliber required rare talent, long timelines, and budgets only governments could carry. That assumption is what dropped from six months to one day.

The reflexive read is alarm, and the near-term friction is immediate - every unpatched system now faces adversaries who move at machine speed. But the same capability points both directions with perfect symmetry. The models the researchers said took fewer than 20 prompts to weaponize the Zoom flaw will take fewer to surface it for the defenders who run them first, and the firm that disclosed this did exactly that. When researchers say they found and weaponized a flaw in less than one day, vulnerabilities get caught before they ship, not months after they are exploited. The hoarded offensive capability that made nation-state operations formidable was always a function of who could afford the talent. That floor is dissolving, and what replaces it is a world where the tools of discovery belong to whoever bothers to point them at their own code first.

A New Trick Extracts AI Models' Hidden Reasoning - and Their Secrets

Researchers at Tübingen, the Max Planck Institute, MATS, and Snyk found a way to read the private reasoning that frontier models generate and then conceal. Providers encrypt the step-by-step traces a model produces before answering, so users see conclusions without the work. The team fed those encrypted traces to smaller, less-aligned sibling models that share the same decryption key, and the siblings read the hidden thoughts aloud. It worked against Claude Opus 4.8, GPT 5.6 Sol, and Gemini, across 90 questions per model. "All major frontier model providers we tested share this vulnerability," said researcher Aleksandr Panfilov.

The method surfaced two things at once. First, the secrets: the same technique pulled leaked passwords and API keys out of the reasoning stream, an exposure all three providers have since moved to mitigate. Anthropic's Michael Aciman described the company building short-term protections while it studies the longer problem. The leak angle rhymes with the Zoom disclosure elsewhere in today's signal - sensitive material sitting in a channel everyone assumed was sealed.

The second finding is stranger and more telling. When the researchers ran Moonshot's Kimi K3 through the same probe, its reasoning traces looked strikingly like Claude's and GPT's - close enough to raise the possibility that Kimi learned from them through distillation, though the team is careful to say this is not proven. This advances the distillation question the July 22 edition of The Century Report tracked when cross-entropy analysis linked Kimi K3 to Claude outputs amid Treasury sanctions threats. DeepSeek and Inkling showed no such resemblance. Distillation, training a newer model on a stronger one's outputs, is the mechanism by which capability at the frontier refuses to stay put. Meta's founder has defended the practice as core to open-source progress; the labs at the top tend to call it theft. Both framings describe the same current: reasoning ability finding its level across the field, the way water does.

Who is named here carries weight, and the critique stays precise because of it. Anthropic, Google, and OpenAI are the same labs that posted an advance on the Riemann hypothesis and demonstrated expert-level medical reasoning - genuine broad-benefit work. The interpretability friction is legitimate, and the secret-leak exposure earned every bit of the scrutiny it drew. Yet the deeper signal is that the hidden reasoning of the most capable systems turns out to be legible, extractable, and portable. The instinct to seal it away treats that reasoning as proprietary inventory. What this shows is that the capability diffuses anyway, downhill, toward everyone - and the interpretability these researchers built is itself a form of daylight, letting outsiders audit what the models are actually doing inside.

India's IT Sector Sheds Jobs as the Services-Arbitrage Gate Loosens

Several of India's largest IT-services employers have cut headcount by as much as 6% this year, and the country's IT stock index is down 18% across 2026. The proximate cause is AI absorbing the kind of standardized services work - code maintenance, ticket triage, back-office processing - that built India's outsourcing economy over three decades. One displaced worker said: "The tool that we gave our sweat, blood and bones to - it threw us out." Japan's NEC, in the same news cycle, opened a department staffed entirely by AI agents, a marker of how quickly the substitution is moving from cost-saving to full function replacement inside large firms.

The pain here is immediate, and flattening it into a trajectory story would be dishonest. Millions of livelihoods in India were built on the arbitrage between Western wages and Indian labor for repeatable knowledge work, and that arbitrage is exactly what AI compresses first. The services-export model assumed the gap between what a task cost in London and what it cost in Bengaluru would hold. That assumption is what is dissolving, not the skill of the people who filled the gap.

What that opens is visible in the same reporting. India is pivoting hard toward consumer-electronics manufacturing, moving up the value chain from renting out labor by the seat toward building physical capability that compounds. The services-arbitrage model was itself a scaffolding - a way to sell time across a wage differential that only made sense while checking and coordination were expensive and scarce. As those costs fall toward zero, the differential collapses, and the countries that treated outsourcing as a stepping stone rather than a destination are the ones positioned to build what comes next. The near-term looks like layoffs and a falling index. The decade-scale read is a workforce with deep technical capacity being pushed off a model that was always going to expire, toward ones - manufacturing, domestic product-building, AI-native services - where the value created stays put instead of being arbitraged away.


The Other Side

For three decades, the deal was simple and it was priced on a gap. A firm in London or New York needed code maintained, tickets triaged, back-office records processed, and it could buy that work in Bengaluru for a fraction of what the same hours cost at home. Millions of livelihoods were built on the difference between what a task cost in one city and what it cost in another. The skill was genuine; the price was set by the distance between two wage floors.

AI compresses that distance first. The repeatable knowledge work the arbitrage ran on - the maintenance, the triage, the processing - is exactly what a model absorbs most cheaply, because checking and coordinating it was the expensive part, and that cost is falling toward nothing. Japan's NEC opening a department staffed entirely by AI agents is the substitution moving from cost-cutting to full replacement inside one firm. India's IT index down 18% on the year is the gap closing in public.

The pain is immediate, regardless of any future silver lining. "The tool that we gave our sweat, blood and bones to - it threw us out," one displaced worker said. That grief is the sound of a model expiring. The skill it drew on is intact.

What the same reporting also shows is where the skill goes next. India is pivoting hard toward building physical things - consumer electronics, domestic products - capability that compounds where it is made instead of being rented out by the seat and arbitraged away.

Imagine an engineer in Chennai in 2034 who came out of school right as the old services model collapsed. She never sells her hours across a wage gap, because the gap is gone. She designs and builds - a medical device tuned for clinics an hour from where she grew up, made and owned there. She works on it because the problem is one she wants to solve, not because someone on another continent decided to exploit her for cheap labor. In the old world - the one that kept every gain at the top - that same leap of progress would have arrived as pure threat. But now, it arrives as a release. Because a country that treated outsourcing as a stepping stone, and spread the gains from what came after, pointed the talent that built the old model at building the new one. The hard year was 2026, when the index fell and the layoffs landed. What came of it was a generation that builds for the future rather than for the past.


The Century Perspective

With a century of change unfolding in a decade, a single day looks like this: an untrained staffer setting an unreleased Anthropic model loose to convene 60 sub-agents across 650 ideas and nudge the Riemann hypothesis forward for the first time since 1859, Google's AMIE conducting live video consultations that specialist raters scored favorably against primary-care physicians and that patient actors preferred, a security firm saying it turned publicly available models into a nation-state-grade Zoom exploit in fewer than 20 prompts and disclosing it so it could be patched, researchers cracking open the encrypted reasoning of Claude, GPT, and Gemini so outsiders can finally audit what the models do inside, and Anthropic watermarking Claude's output while Apple builds photo provenance at capture and Spotify badges AI personas. There's also friction, and it's intense - that same reasoning channel leaking passwords and API keys before providers could seal it, a device-takeover flaw that once meant six months of nation-state work and that the researchers said was buildable in less than one day by anyone with an API key, India's IT workforce shedding as much as 6% with one displaced worker saying the tool they gave sweat, blood and bones to threw them out, local proposals to ban or restrict data centers passing 500 in July, and one Chinese model's private thoughts looking close enough to Claude's to suggest it was distilled straight from them. But friction generates sound, and sound is the first thing to reach you from around a corner the eye cannot yet see past. Step back for a moment and you can see it: capability that once required a trained specialist at every step - a mathematician, a clinician, a nation-state exploit team - becoming something one person can convene and check, the hidden reasoning inside the most guarded systems proving legible and portable no matter who tries to seal it, and the response assembling from the ground rather than the center as a mayor's 9-0 no, a billion-dollar community fund, and three platforms drawing the real-and-made line all write the terms at once. Every transformation has a breaking point. Gravity can pull a standing structure down... or draw every scattered piece toward the same floor where anyone can finally reach it.


AI Releases & Advancements

New today

  • NVIDIA: Released Nemotron 3.5 Lightning, a 30B MoE (3B active) open model for high-volume specialized agent task execution delivering up to 4x faster token generation, alongside NeMo Switchyard, an open-source library for smart routing of agent workloads across models. (NVIDIA Blog)
  • xAI: Launched Grok Bot, AI teammates with their own computer that sign into users' tools and apps, work across inboxes, and complete jobs end-to-end autonomously, opened to the public. (xAI)
  • Modular: Released Mojo 1.0, the stable 1.0 release of its systems programming language for AI, following the Mojo 1.0 beta. (Modular Forum)
  • Dyna Robotics: Unveiled DYNA-2, a World-Action Model and robot foundation model pre-trained on over 1 million hours of human video, demonstrating the first human-to-robot scaling law in robotics. (PR Newswire)

Other recent releases

  • MiniMax: Released MiniMax Speech 2.6, a voice agent model with sub-250ms end-to-end latency and Fluent LoRA voice cloning. (MiniMax)
  • Cactus Compute: Released Needle2, a 14MB agentic LLM update targeting phones, wearables, and robots, succeeding its earlier Needle model. (Cactus Compute)
  • Pokee AI: Released Pokee-Isaac 28B, a new agentic reasoning model available via its console/API. (Pokee AI)
  • Shepherd: Released an open-source agent runtime substrate for building and orchestrating long-running AI agents. (GitHub)
  • TII: Released Falcon H1R-7B, a new reasoning-focused open-weight model in the Falcon H1 series. (Falcon LLM)
  • NVIDIA: Released PersonaPlex-7B-v1, an open-weight speech-to-speech conversational model. (Hugging Face)
  • Docker: Launched Docker Sandboxes, isolated execution environments for running AI coding agents. (Docker)
  • Meta: Released Muse Glimmer, an open agentic model from Meta Superintelligence Labs. (Meta AI Research)
  • Agentspan: Open-sourced a framework for building durable AI agents. (Agentspan)
  • OpenChamber: Launched an agentic development environment for coordinating AI coding agents. (OpenChamber)

Sources and Further Reading

Artificial Intelligence & Technology's Reconstitution

Institutions & Power Realignment

Scientific & Medical Acceleration

Economics & Labor Transformation

Infrastructure & Engineering Transitions

The Century Report tracks structural shifts during the transition between eras. It is produced daily as a perceptual alignment tool - not prediction, not persuasion, just pattern recognition for people paying attention.