Vol. I · No. 23 Weekly Edition June 20, 2026

Midjourney announces a full-body ultrasound scanner. GPT-5.4 identifies a reaction additive no chemist expected. AI benchmarks shift to measuring what you can complete, not what you know. The instruction gap paper drops. Hours after the government shut Fable 5 offline, China shipped a SOTA open-weight replacement.

Inside

Midjourney Medical’s ultrasound pivot. GPT-5.4 improves Chan-Lam coupling. Intelligence Index v4.1 goes agentic. The Instruction Gap paper. GLM 5.2 arrives as Fable 5 goes dark.

Edited By

David Borish
From New York

Filed

Five articles, one week, sourced from The AI Spectator

Health · Technology No I

Midjourney Medical, June 18, 2026

The most
surprising
pivot in AI.

Midjourney, the image-generation company, announced a full-body ultrasound scanner it says is superior to MRI in several respects. The prototype has been used on twelve people. No FDA clearance exists. The scanner is planned for a Union Square spa.

By David Borish June 19, 2026 · 5 min read

On June 18, Midjourney founder David Holz stood in San Francisco and announced a full-body ultrasound scanner he described as superior to MRI in several respects: sixty-second completion time, no radiation, no magnetic fields, eventual cost of a few dollars. The announcement drew immediate attention. The gap between that framing and the current state of the technology is the most important fact about it.

By Midjourney’s own numbers, the prototype takes twenty minutes per scan, not sixty seconds. It has been used on approximately twelve people. The team building it has nine members. The device has no FDA clearance for diagnosing anything. The sixty-second target is engineering aspiration, not current performance. The data transfer bottleneck between the transducer array and the reconstruction computing cluster has not been solved.

The hardware rests on a licensing deal with Butterfly Network, a publicly traded medical device company. The agreement, disclosed in an SEC 8-K filing, commits Midjourney to a $15 million upfront payment, $10 million per year in licensing fees over five years, and additional milestone and revenue-sharing payments totaling up to $74 million. Butterfly contributed $6.8 million in Q4 2025 revenue from the arrangement. The underlying chip replaces piezoelectric crystals with capacitive micromachined ultrasonic transducers built on standard CMOS semiconductor wafers. Each chip contains up to 9,000 transducer elements. The scanner uses 40 chips per unit, roughly 8,960 active channels, paired with approximately two petaflops of computing power.

The technique itself, ultrasound computed tomography, has been studied since the 1950s. The FDA cleared a USCT device for breast cancer screening in 2021. What Midjourney is attempting, whole-body USCT at full torso and leg coverage, has not previously been cleared or independently validated at that scale. One independent data point exists: a 0.93 correlation coefficient with MRI-based proton density fat fraction. That figure comes from Midjourney’s own research and has not been replicated externally.

Midjourney is navigating FDA regulation by making no diagnostic claims. It is launching the scanner as a provider of detailed body composition maps rather than a diagnostic tool, consistent with the FDA’s general wellness guidance for low-risk non-invasive measurement devices. Prenuvo and Ezra operate whole-body MRI services for consumers using the same regulatory lane. Holz said the company has begun discussions with the FDA and will submit test results incrementally to pursue diagnostic clearance over time.

The go-to-market strategy skips hospitals entirely. Midjourney is building branded facilities it calls Midjourney Spas, with the first planned for Union Square in San Francisco, expected to open at end of 2027 in roughly 25,000 square feet, housing ten scanners alongside saunas, hot tubs, and cold plunges. The scan is described as a side effect of a spa visit. By positioning scans inside wellness spaces, the company builds scan volume and a longitudinal dataset outside the clinical regulatory framework while pursuing diagnostic clearance separately.

By 2031, Midjourney aims to operate more than 50,000 scanners worldwide with capacity for one billion scans per month. The engineering, regulatory, and operational distance between twelve people scanned and that target is substantial. The timeline leaves little room for the kind of iteration hardware development typically requires.

The gap between a sixty-second target and a twenty-minute reality is a data transfer bottleneck. The physics work. The engineering does not yet. The AI Spectator
Read the full article →
Evaluation · Infrastructure No II
34

Benchmarks stop
measuring what you know.

Artificial Analysis’ Intelligence Index v4.1 assigns 34 percent of total weight to agentic task completion. The field’s leading independent evaluators have moved away from knowledge retrieval as a proxy for capability.

The headline score on an AI benchmark always conceals a set of choices. What tasks count, how answers are graded, which capabilities get weighted most heavily, and what a failure actually looks like all shape the final number before any model runs a single evaluation. With the release of Intelligence Index v4.1, Artificial Analysis has published a methodology that makes those choices transparent.

The weighting is the sharpest signal in the framework: 34% agentic tasks, 24% coding, 24% scientific reasoning, 18% general knowledge. Nine evaluations make up the suite: GDPval-AA v2, the τ³-Banking agent evaluation, Terminal-Bench v2.1, SciCode, AA-LCR, AA-Omniscience, Humanity’s Last Exam, GPQA Diamond, and CritPt.

The most architecturally significant component is GDPval-AA v2, which carries 20% of the index weight. It covers 220 tasks across 44 occupations tied to GDP-contributing sectors. Tasks are open-ended professional work: a model receives a brief and reference files, uses tools to produce output, and submits when done. Grading happens through pairwise comparison, judged by a three-model panel: GPT-5.5 at medium reasoning, Gemini 3.1 Pro Preview at high reasoning, and Claude Opus 4.8 at high effort.

The calibration element matters. Human expert deliverables are set at Elo 1,000. Claude Fable 5 scored 1,932 on GDPval-AA, placing substantially above human expert level on these tasks as evaluated by the judge panel. That framing is more informative than percentage scores against static question sets. The relevant comparison for most deployment decisions is not one model versus another; it is a model versus the human workflow it might replace or augment.

On AA-Omniscience, which rewards restraint by penalizing hallucinations and assigning no penalty to abstentions, Claude Fable 5 scored 40, seven points ahead of the prior leader. The gain was driven primarily by accuracy rather than low hallucination rate. The τ³-Banking evaluation requires agents to navigate approximately 700 interconnected policy documents totaling 195,000 tokens across 21 product categories. Scoring is done against backend database state: whether a dispute was actually opened, whether a provisional credit was actually issued.

Three evaluations were retired from v4.0: IFBench, MATH-500, and AIME 2025. The retirements reflect a consistent pattern: evaluations that saturate quickly as frontier models improve, or that test a capability slice already captured by higher-weight components, get replaced. The Openness Index, tracked separately, scores models on transparency across pre-training data, methodology, and model availability. The highest-scoring open-weight model at publication sits around 55 on the Intelligence Index, compared to Claude Fable 5’s 64.9.

34%
Agentic Weight
in v4.1 Index
1,932
Fable 5 Elo
vs. 1,000 Human
64.9
Fable 5 Overall
Intelligence Score
Read the full article →
// security :: export_control :: open_source No III
[ SEQUENCE LOG — JUNE 12–13, 2026 ]

The ban didn’t
slow China down.
It pushed enterprises
toward the models
it was competing
against.

$ timeline --event "fable5_launch" --date 2026-06-09
> Anthropic releases Fable 5 as most capable public model
$ timeline --event "govt_directive" --date 2026-06-12 --time 17:21ET
> US Commerce Dept orders Fable 5 + Mythos 5 suspended globally
> Reason: jailbreak bypassed safety guardrails, exposed Mythos cybersec capabilities
> Anthropic complied: both models taken offline, all users, no notice
$ timeline --event "glm52_release" --date 2026-06-13
> Zhipu AI releases GLM 5.2: BridgeBench BS score 100.0, Reasoning 42.8
> License: MIT • Context: 1M tokens • OpenCode/Cline/Claude Code compatible
> Weights: open release this week

Three days after Anthropic launched Fable 5, the US Commerce Department ordered it shut down. The directive, delivered at 5:21 pm ET on June 12, required Anthropic to suspend all access for any foreign national, including the company’s own employees. Anthropic took both Fable 5 and Mythos 5 offline globally. Hundreds of millions of users lost access to the frontier model they had been using three days after its release, with no warning and no confirmed return date.

The government’s stated reason was a jailbreak. A trusted partner had demonstrated a technique that bypassed Fable 5’s safety guardrails, unlocking access to the underlying Mythos model’s cybersecurity capabilities. Trump administration AI advisor David Sacks wrote publicly that Amodei was told to fix the bypass before deployment and declined. Anthropic characterized the jailbreak as narrow and already reproducible using other publicly available models, including GPT-5.5, none of which face similar restrictions.

The dispute sits inside a larger conflict. Anthropic had refused to allow the Pentagon to use its models for fully autonomous weapons systems. The military placed the company on a blacklist. The export control directive arrived weeks before a confidential IPO filing, applying pressure from a different angle.

On June 13, Zhipu AI released GLM 5.2. The model posted 100.0 on BridgeBench’s BS leaderboard and 42.8 on the Reasoning ranking, above Claude Opus 4.6, with a 1-million-token context window and MIT license. Zhipu listed on the Hong Kong Stock Exchange in January 2026, becoming the first publicly traded Chinese AI lab, with a market valuation of approximately $34.5 billion. The release was framed explicitly as a response to what the company called international restrictions on frontier intelligence.

The Open-Prem Inflection Point V3 framework anticipated this dynamic. Self-hosted AI has crossed from workaround to rational default for organizations with sufficient scale, accelerating not because the models improved but because the alternative was shown to be fragile. The organizations least disrupted were those already running on self-hosted infrastructure.

Export controls applied to proprietary API-hosted models accelerate adoption of open-weight alternatives running on infrastructure the user controls. Once weights are released, they cannot be recalled. DeepSeek R1, Kimi K2, Qwen 3, and GLM 5.2 are already on servers that US export controls do not touch. The Open-Prem V3 case is no longer theoretical: a government directive, a jailbreak controversy, an IPO filing, a dispute about autonomous weapons, and suddenly the most capable publicly available AI model in the world is offline for an indeterminate period, and a Chinese open-weight alternative is the top-ranked option on the leaderboard.

// DATA_LOG
80%
SHARE
US startups using Chinese open-source models, per March 2026 US-China Economic and Security Review Commission report
~30%
HUGGING FACE SHARE
Global model downloads from Chinese labs, up from roughly 1.2% at end of 2024
$34.5B
MARKET CAP
Zhipu AI valuation at Hong Kong listing, January 2026 — first publicly traded Chinese AI lab
1M
TOKEN CONTEXT
GLM 5.2 context window; MIT license; OpenCode, Cline, Claude Code, and Roo Code compatible
Read the full article →
Chemistry · Drug Discovery No IV

An AI Chemist
Improved a Drug-Making Reaction
with an Additive No One Predicted

§ § §

GPT-5.4 and Molecule.one’s Maria system ran 10,080 reactions across two experimental cycles, identified TEMPO as a useful oxidant for a historically low-yield coupling reaction, and submitted a finding that four independent chemists supported as novel.

A drug molecule that cannot be synthesized is, for practical purposes, a molecule that does not exist. Medicinal chemists work inside that constraint every day. They design promising compounds on paper, but if the chemistry to build them produces low yields or messy byproducts, the molecules get abandoned and the program moves on. Synthesis is one of the quiet bottlenecks in drug discovery. That is the setting for a result OpenAI and Molecule.one published on June 17, 2026.

Working together, the teams connected GPT-5.4 to Maria, an agentic chemistry system tied to a high-throughput laboratory, and handed it an open-ended assignment: improve one of several important reaction classes. The system generated research proposals, designed and ran experiments, analyzed the data, and proposed follow-up work. Its most promising idea focused on Chan-Lam coupling, a reaction that forms carbon-nitrogen bonds appearing throughout the structures of medicines.

The gap the system targeted was specific. Coupling primary sulfonamides with boronic acids has historically given low yields, and that gap matters because sulfonamides appear in anticancer drugs, antimicrobials, and diuretics. GPT-5.4 narrowed the problem itself, identified primary sulfonamides as a difficult, high-value substrate class, then suggested that mild oxidants, TEMPO among them, could improve the reaction. The human chemists found the TEMPO suggestion unexpected, which is part of why it was worth running.

The proposal, labeled OAI-M1-03, was one of four selected from thousands the system ranked. Maria AI then translated the plan into detailed lab instructions, ran the experiments at high throughput, analyzed raw data, and returned structured results to GPT-5.4 for follow-up design. The full process ran three months, from March 4 to sharing results with independent experts on June 4.

Maria ran 10,080 reactions across the two cycles of OAI-M1-03 — more than a chemist running three reactions a day would complete in a decade. Mean yield rose from 16.6 percent to 25.2 percent. The share of reactions clearing 30 percent yield went from 15.6 percent to 37.5 percent. Measured yields improved for 88 percent of the boronic acids and 83 percent of the sulfonamides tested.

In the second round, the system found that TEMPO could be replaced by 4-hydroxy-TEMPO, a considerably cheaper analog, with little loss in performance. Human chemists then reproduced representative reactions by hand at bench scale and observed higher yields for 11 of 14 substrate pairs, with more than twofold improvement in eight of them. Four independent chemistry experts reviewed the preprint and supported the finding as novel. Tim Cernak at the University of Michigan described the result as a demonstration of mild conditions and a practical oxidant producing a broadly useful substrate scope.

The scope of what the paper establishes is worth stating directly. The work shows that a model can make a useful contribution to organic chemistry, proposing a specific surprising hypothesis, surfacing it for human review, designing experiments, interpreting data, and designing follow-ups. It does not show that AI can run a chemistry research program end to end. Human judgment stayed essential throughout. The workflow depended on specialized high-throughput infrastructure that few labs have.

Of the three other proposals the system generated and tested, two were experimentally confirmed and one was disproven, with analysis still ongoing. Independent replication and broader substrate characterization remain ahead. The next chapter belongs to the independent labs.

The AI Spectator June 20, 2026 Chemistry & Drug Discovery
Read the full article →
Labor · Economics · Skills No V

The skill that
matters now
is not prompting.

David Borish’s new paper argues that the competitive advantage in AI-assisted work has shifted to instruction-following: executing AI output completely and in sequence, without the abbreviation that breaks implementations.

The prompting era is over as a source of durable advantage. By 2026, a reasonably clear natural language request produces usable output most of the time. Ask an AI to improve your prompt and it will produce something more effective than most people write on their own. The advantage that once required real investment has compressed considerably.

What has not compressed is the execution gap on the other side of the exchange. Claude Code, Cursor, and similar tools can now generate multi-file applications, configure dependencies, and scaffold complete architectures from plain descriptions. When those agents return a sequence of steps, a human still has to complete them. Create an account. Generate an API key with specific permissions. Set an environment variable in a specific file. Run a command in a specific directory. Each instruction is discrete, each depends on the one before it, and each requires complete human execution before the next one begins.

This is where builds break. The AI produced correct instructions. The human introduced an error in the execution. The build fails. The tool gets blamed. The actual cause goes unexamined.

The research base is substantial. A 2020 review in the American Journal of Pharmaceutical Education identified working memory capacity as the central constraint on instruction-following performance. Multi-step sequences place measurable cognitive load on working memory. When that load exceeds capacity, people start sequences and do not finish them. Susan Gathercole’s research at Cambridge documented the specific breakdown pattern: instruction sequences fall apart at the back end. Early steps get executed because they feel essential. Later steps because they feel conclusive. Middle steps are where the cognitive load is highest, where sequences are most likely to be truncated, and where errors are most consequential.

A 2016 study by Jaroslawska, Gathercole, Allen, and Holmes identified a simple intervention: execute each step at the moment of instruction, rather than reading the full sequence and then attempting it in batch. Completion rates improve substantially. Stanley Milgram’s obedience research adds another dimension: compliance drops measurably when authority leaves the room. An AI agent returns a numbered list and waits. There is no authority in the room and no checkpoint structure built into the process.

The enterprise failure mode differs from the individual one. In team implementations, instruction sequences are frequently distributed across multiple people. Person A completes steps one through four and hands off to Person B, who picks up at step five without knowing step three was incompletely executed. The build fails six steps from the actual error. BCG research estimates fewer than 20 percent of enterprise AI projects achieve their expected return on investment. Instruction-following capacity as a failure mode does not appear in standard analyses, not because it is absent, but because it has not been named or measured.

Instruction-following discipline is learnable, trainable, and not dependent on a technical background. For individuals watching the AI transition, it is a more accessible and immediately valuable skill than prompt engineering. It also transfers forward: as AI systems move into physical-world coordination tasks, the discipline that produces clean software builds will produce clean execution of whatever comes back in those environments.

Read the full article →
Key Figures
<20%
Enterprise AI projects achieving expected ROI, per BCG research — instruction-following failure is a factor that goes unmeasured
2016
Year a Cambridge study identified enactment — one step at a time, confirmed before advancing — as the intervention that most reliably improves sequence completion
19
Variations of Milgram’s experiment documenting how compliance drops measurably when authority leaves the room — an AI agent cannot stay in the room
2026
The prompting advantage has compressed. The execution advantage has not. Browser control agents are coming. The instruction-following window is open now.

The AI Spectator Weekly is published at davidborish.com/the-ai-spectator

Frameworks explored this issue:
Open-Prem Inflection Point V3  ·  The Exponential Replacement Curve

Vol. I · No. 23 · June 20, 2026 · Edited by David Borish · New York