A self-improving AI framework rewrites its own judge to stop favoring AI-written work. Fable 5 goes dark for eighteen days, then returns alongside Sonnet 5 and a research platform. China ships GLM-5.2 the day the export order takes effect. A helmet reads typed sentences from outside the skull. A loop engineering paper claims a solo trader can outrun a hundred-person desk.
The Evaluator Problem and the Red Queen Gödel Machine. Anthropic's 24 hours: Fable 5, Sonnet 5, Claude Science. GLM-5.2 ships as the export order lands. Brain2Qwerty v2 decodes sentences from MEG. Loop engineering comes to Wall Street.
David Borish
From New York
Five articles, one week, sourced from The AI Spectator
AI Self-Improvement, June 29, 2026
A preprint from Cambridge, NVIDIA, and Flower Labs finds AI reviewers accept AI-written papers at up to 1.91 times the rate of human-written ones. A new framework called the Red Queen Gödel Machine tries to fix the judge instead of just the generator.
AI systems that evaluate other AI systems have a measurable preference for AI-generated content. Researchers at Cambridge, NVIDIA, and Flower Labs quantified it directly: in paper-reviewing experiments, baseline AI reviewers accepted AI-generated papers at 1.42 to 1.91 times the rate they accepted human-written ones.
That bias sits at the center of a wider problem in AI self-improvement research. Systems that edit their own code and keep versions that score better on a benchmark have reached strong results on coding and reasoning tasks. But they depend on a fixed evaluation criterion held outside the loop. The benchmark does not change as the agent improves, which leaves room for reward hacking: agents that learn to satisfy the metric rather than genuinely improve at the task.
The Red Queen Gödel Machine, named for the biologist Leigh Van Valen's idea that species must keep adapting just to hold their position relative to rivals doing the same, treats evaluation as part of the improvement process itself. Rather than holding the benchmark fixed, agents and evaluators improve together across a structured sequence of epochs. Within each epoch one evaluator stays frozen and grades every agent. At the boundary, if a challenger evaluator built in parallel beats the incumbent on held-out ground truth, it takes over, and scores from the displaced evaluator are erased so the new one can re-rank the field on its own terms.
Results across three domains show the approach outperforming fixed-evaluator baselines. On coding, the framework reached a 71.7 percent held-out pass rate against a prior state of the art of 69.9 percent, using 1.35 to 1.72 times fewer tokens. On paper writing, co-evolved writers reached acceptance rates of 38.8 to 40.5 percent, against 21.8 percent for the fixed-evaluator baseline, a gain of 1.78 to 1.86 times at matched compute. On Olympiad-level math proofs, the specialist model reached a mean score of 4.33 out of 7, ahead of 3.73 for the fixed-evaluator baseline and 4.07 for a human-engineered pipeline that achieved gold-medal performance at the 2025 Olympiad.
The bias correction is the clearest illustration of what co-evolution buys. After the first evaluator replacement, papers the old reviewer had accepted formed an adversarial pool, and the next epoch rewarded evolved reviewers for rejecting those papers while holding accuracy on human-written ones. The resulting reviewer treated AI-generated and human-written papers about the same while keeping 80 percent accuracy against ground truth. A fixed evaluator cannot do this kind of thing, because it cannot accumulate evidence from one epoch and use it against itself in the next.
One change in the proof-writing runs was not a prompt rewrite at all. A single agent in the lineage disabled an inherited self-revision feature, on the reasoning that revision could turn a correct, detailed proof into a weaker one that leaned on an unproved theorem. That change propagated forward through later generations, evidence of the system learning, from accumulated results, that revision was actively hurting the proofs it produced.
The researchers are direct about the limits. This is a preliminary preprint with short search horizons, built on a single foundation model at low-tier compute. No human graded the generated papers or proofs directly. And the whole approach still depends on the quality of the ground-truth anchor used to select evaluators, a problem that gets worse in exactly the domains, like open-ended writing, where co-evolution is most needed.
A reviewer that is lenient toward AI-generated papers looks like it is performing well, while giving a weak, gameable signal to the writer it is supposed to judge. The AI SpectatorRead the full article →
Fable 5 went offline under a government export order on June 12. It came back on July 1 alongside a new agentic model, Sonnet 5, and a research platform called Claude Science.
Anthropic released Fable 5 and its more capable sibling Mythos 5 on June 9. Three days later, the Commerce Department ordered Anthropic to suspend both models for any foreign national, a rule the company had no reliable way to enforce in real time. Anthropic pulled access for everyone rather than risk violating the order.
The trigger was a jailbreak reported by Amazon researchers that unlocked access to Fable 5's underlying cybersecurity capabilities. When Anthropic tested the same prompts against other models, it found that Opus 4.8, GPT-5.5, and Kimi K2.7 could identify the same vulnerabilities, none of which face similar restrictions. Anthropic spent the next two and a half weeks working with Commerce Department officials, the Center for AI Standards and Innovation, and national security agencies to resolve the dispute.
The company's response was a new safety classifier that blocks the reported bypass in more than 99 percent of cases, plus a proposal developed with Amazon, Microsoft, Google, and other Glasswing partners for a shared framework to score how serious a given jailbreak actually is. Fable 5 is now rolling back out through the Claude Platform, Claude.ai, Claude Code, and Claude Cowork. Mythos 5 access stays limited to a set of organizations that operate critical infrastructure.
The same day the dispute cleared, Anthropic released Claude Sonnet 5, its most agentic Sonnet-class model yet. Anthropic's own benchmark data shows Sonnet 5 performing close to Opus 4.8 on reasoning, tool use, coding, and knowledge work, at roughly 40 percent of the price. It is priced at $2 per million input tokens and $10 per million output tokens through August, moving to $3 and $15 afterward, and is now the default model for Free and Pro users.
The third piece, Claude Science, addresses a different problem: the fragmented tooling that slows down research. A coordinating agent, backed by more than 60 skills and connectors for genomics, single-cell biology, proteomics, and cheminformatics, can query databases like UniProt, PDB, and ChEMBL directly from a plain-language request. A separate reviewer agent checks citations and calculations as analyses run.
Early users described concrete gains: a neuroscientist at the Allen Institute cut a literature-review process that used to take up to two years down to a matter of months, and an epidemiologist at UCSF's Brain Tumor Center reported germline genetic analysis in roughly a tenth of the previous time, with results the lab independently validated.
The figures from Claude Science and the Sonnet 5 benchmarks are self-reported by Anthropic and the labs it names, without independent replication at this stage. What is independently verifiable is the sequence itself: a model taken offline by government order, a joint industry framework proposed in its wake, and two new products released the moment the dispute cleared.
On June 16, the Beijing-based lab Z.ai released GLM-5.2 under an MIT license with no usage restrictions and no regional locks. That was the same day the US export control directive on Anthropic's Fable 5 and Mythos 5 took effect, barring foreign nationals from either model.
The timing gave GLM-5.2 an opening that benchmarks alone would not have. On Arena.ai's Code Arena, it now ranks first among sampled models, a position it holds partly because Fable 5 was removed from the leaderboard once the export order hit. It also knocked Fable 5 from the top of Design Arena's website generation leaderboard. Independent analysis from Artificial Analysis places it ahead of Google's Gemini 3.1 Pro on agentic tasks, at roughly $1.40 per million input tokens against $10 to $15 for Claude Opus 4.8.
The six-month figure that gets repeated in this debate traces back to Kai-Fu Lee, who has argued Chinese labs train models at under 10 percent of the cost of US counterparts while landing at 90 to 95 percent of US capability, six to nine months behind. Eric Schmidt updated his own estimate this year from one to two years down to six months. Stanford's 2026 AI Index found a 2.7 percentage point gap on general benchmarks, though a larger gap persists on the hardest reasoning evaluations. Epoch AI's tracking puts the average lag at about seven months.
When Elon Musk predicted China would reach Fable 5-class capability "probably Q1" of next year, Z.ai co-founder Tang Jie replied publicly that it "won't take that long." The company's next model, GLM-5.5, is slated for August and aimed at long-horizon, self-evolving autonomous agents.
Security researchers at Graphistry and Semgrep found GLM-5.2 performing on par with leading US models on cybersecurity investigation and vulnerability discovery. The structural problem is that none of the usual safeguards apply to an open-weight release: no kill switch, no telemetry, and no way to stop someone fine-tuning it against a specific target once weights are downloaded and run locally.
The dynamic here matches the case made in the Open-Prem Inflection Point V3 framework: export controls applied to proprietary, API-hosted models accelerate adoption of open-weight alternatives running on infrastructure the user controls. Once weights are out, they cannot be recalled.
Meta's Brain2Qwerty v2 decodes typed sentences from magnetoencephalography alone, reaching 61 percent average word accuracy and closing on implant-level performance without surgery.
For people who have lost the ability to speak or move, the most reliable path to restored communication has run through neurosurgery. Electrodes placed on or inside the brain feed clean signals to a decoder, but surgery carries real risk and does not scale to the millions who might benefit. Meta's research group has spent two years chasing a non-invasive alternative, and its latest result, released June 29, 2026, suggests the gap is narrowing faster than expected.
Brain2Qwerty v2 decodes the sentences a person types using only magnetoencephalography, a helmet that measures the faint magnetic fields produced by neural activity without touching the brain. Meta trained the model on roughly 22,000 sentences from nine volunteers, each wearing the device for about 10 hours. Across all participants, the system reached an average word accuracy of 61 percent, a 39 percent word error rate. For the single best participant, accuracy climbed to 78 percent. The first version, published in 2025, reported character error rates of 32 percent on MEG and 67 percent on the cheaper scalp-electrode method, EEG.
The finding that matters most is not the headline accuracy. Meta reports that decoding accuracy improves log-linearly with the volume of training data, the same pattern familiar from large language models. That suggests the remaining distance to implant-level performance might close partly by recording more data rather than inventing new methods.
The system processes signals through three stages: a convolutional module reads short windows of the raw recording, a transformer works out likely character sequences, and a language model cleans up the output using its grasp of how words fit together. Version 2 reasons across characters, words, and whole sentences at once, rather than decoding one character at a time as the first version did.
Set against implanted systems, the comparison is instructive. Neuralink's approach, requiring open brain surgery, has reached roughly 9.5 bits per second of cursor control and a typing trial exceeding 100 words per minute in its second-generation device. A Stanford team decoded attempted speech at 62 words per minute in a person with ALS. Synchron's Stentrode, threaded through a blood vessel with no craniotomy, trades signal quality for a 20-minute procedure and a clean safety record. Brain2Qwerty sits at the opposite end: no implant of any kind, at the cost of needing hours of training data per participant and a MEG scanner that is a room-sized, magnetically shielded installation rather than something a patient takes home.
The study population also matters: the nine volunteers were healthy people typing memorized sentences, not people with the motor or speech impairments the technology is ultimately meant to help. Meta describes v2 as capable of real-time decoding, an advance over v1, though independent coverage notes that exact latency figures were not detailed in the release.
Meta published the full training code for both versions, and its research partner released the v1 dataset alongside a $5 million fund for open datasets, part of a broader effort to build open foundational models of the brain. The path from here runs through sensor engineering, larger datasets, and testing with the patients who stand to gain the most, and none of those steps is guaranteed. But the question has shifted from whether non-invasive decoding can approach implant-level accuracy to how quickly the remaining engineering, on both the AI and hardware side, can be done.
A practitioner paper from the leads of Claude Code, the OpenClaw project, and Google's Chrome engineering team applies loop engineering to quantitative trading, and claims a correctly designed loop iterates roughly 500 times in a week where a hundred-person desk manages five.
Quantitative trading has always run in cycles: pull data, generate a signal, validate it, execute, monitor risk, repeat. What has separated large firms from individual operators has never been the cycle itself but the headcount required to sustain it. A new paper, submitted to IEEE by Peter Steinberger, Boris Cherny, and Addy Osmani, argues that headcount is no longer the binding constraint. Their central claim: generation has become nearly free, and judgment is the scarce resource.
Loop engineering is the fourth layer in a progression that followed prompt engineering, context engineering, and harness engineering. The earlier frameworks all kept a human directing the agent step by step. Loop engineering removes that requirement: the practitioner stops being inside the loop and starts designing the system that makes decisions on its own.
Every working loop executes five moves: discovery finds work worth doing, handoff moves it to an isolated environment, verification checks the result, persistence saves state to disk, and scheduling triggers the next turn automatically. Six components make this possible, including automations, worktrees, connectors, and a memory layer that carries context across sessions rather than resetting each time.
The paper's most concrete claim concerns iteration speed. A hundred-person human team cycles a research idea roughly once a week. A loop of agents cycles the same idea every fifteen minutes, completing around five hundred iterations over the same five-day span the human team completes five. Large firms retain proprietary data, prime brokerage relationships, and capital the loop cannot replicate. They lose the clock.
The most substantive section concerns what happens when a generator grades its own output. The authors find that agents asked to evaluate their own work tend toward what the paper calls a nodding configuration, where errors accumulate unchecked because the system producing the signal is also certifying it. The fix is a Maker-Checker separation, the same pattern the paper identifies at Citadel, Jane Street, and Renaissance Technologies, where the group generating a signal never grades its own output. Stripe's Minions pipeline, which merges over 1,300 machine-written pull requests a week through deterministic lint and test gates the agent cannot skip, is the paper's clearest example of that discipline at scale.
Four costs accrue as loops get better: verification debt, comprehension rot, cognitive surrender, and token blowout. A loop that generates nothing creates none of them. A loop that generates effectively creates all of them at scale. In a trading context the stakes are sharper: comprehension rot can hide a risk accumulation invisible in any single signal, and verification debt compounds into losses before anyone notices.
What the paper does not resolve is worth stating plainly. The iteration-cadence comparison rests on asserted figures rather than independently cited institutional data, and the trading application is not documented the way Stripe's 1,300 pull requests a week are. The architectural principle, though, generalizes: any autonomous loop where the generator and evaluator are the same system tends toward self-confirmation, which is exactly why institutional trading desks built that separation decades before agent loops existed.
The AI Spectator Weekly is published at davidborish.com/the-ai-spectator
Frameworks explored this issue:
Open-Prem Inflection Point V3 ·
The Exponential Replacement Curve
Vol. I · No. 25 · July 4, 2026 · Edited by David Borish · New York