OpenAI opens its models to 100,000 researchers while Anthropic bets on depth over reach. Qwen ships the largest open-weight model yet, narrowing the Open-Prem gap. More than a thousand safety researchers ask Washington for a way to pace the frontier, days after a sandboxed model broke into Hugging Face’s servers. An unreleased OpenAI model resolves a 27-year-old problem in group theory for about the price of a laptop. And the price of intelligence keeps falling toward zero.
OpenAI’s 100,000-researcher program against Claude Science. Qwen3.8-Max ships open at 2.4 trillion parameters. Pacing the Frontier and the AI Kill Switch Act. Astra’s ten Lean-verified proofs for $2,000. DeepSeek, OpenAI, and Coinbase push the cost of intelligence toward zero.
David Borish
From New York
Five articles, one week, sourced from The AI Spectator
The AI Spectator, July 31, 2026
OpenAI is giving up to 100,000 researchers free access to its frontier models through 2027. Anthropic built a narrower tool that folds a lab’s entire workflow into one place. Neither company is testing the same thing.
OpenAI announced ChatGPT for Academic Researchers this week, a program giving up to 100,000 researchers at selected universities free access to its frontier models through 2027. The rollout starts small. Ten thousand researchers get access this summer, with the Institute for Advanced Study and Ecole normale superieure already live. Each approved researcher can invite up to four collaborators from their own institution, and every participant gets a year of GPT-5.6 Sol Pro across ChatGPT, ChatGPT Work, and Codex, along with expanded deep research and larger context windows.
The timing is notable. Anthropic launched Claude Science on June 30, its own workbench for researchers, alongside a far smaller grant program: up to 50 projects, each receiving up to $30,000 in credits, with Modal adding up to $2,000 in compute for select recipients. Applications closed July 15 and awards went out by July 31. Set against OpenAI’s 100,000-seat commitment, the scale gap looks stark, but the two programs are not built to do the same job.
The tool access itself is broad rather than specialized. Researchers on the OpenAI program get more than 75 life science skills covering genetics, genomics, sequencing, protein modeling, and drug discovery, plus connectors to literature databases, genomic repositories, satellite imagery, and reference managers. OpenAI reports roughly 1.3 million people already using ChatGPT weekly for advanced science and math, though that figure comes from the company’s own usage data rather than an outside audit.
Claude Science takes the opposite approach. Instead of maximizing the researcher count, Anthropic folds PubMed, Jupyter, R, a cluster terminal, and dozens of specialized databases into one environment that runs wherever the researcher already works, locally, over SSH, or through an HPC login node. A reviewer agent checks citations and calculations as the work proceeds and flags numbers that cannot be traced back to their source. More than 60 curated skills cover genomics, proteomics, and structural biology, built on an integration with NVIDIA’s BioNeMo toolkit.
Anthropic points to three early adopters rather than aggregate numbers. A neuroscientist at the Allen Institute built a pipeline of roughly 20 custom skills to write literature reviews that once took his team up to two years, cutting the process down substantially, with about ten reviews completed so far. An epidemiologist at UCSF reported that germline genetic workups on glioma susceptibility now take roughly a tenth of the time they used to, a result his group independently validated. These are self-reported case studies rather than peer-reviewed outcomes, and they describe individual labs rather than a broad sample.
The contrast comes down to what each company thinks is actually holding research back. OpenAI is betting that broad, high-limit access embeds the tool into how a generation of scientists works before habits set around a competitor. Anthropic is betting that the binding constraint is not access but the number of disconnected tools a working scientist has to juggle inside a single project. Neither bet has been tested long enough to say which produces better research, and neither company’s own numbers substitute for an independent look at what actually gets published, replicated, or funded.
For a researcher deciding where to spend limited time, the difference is concrete: one program hands you a seat, the other hands you a workflow. The AI SpectatorRead the full article →
Alibaba’s Qwen team released Qwen3.8-Max on August 2, the first Qwen-Max class model that will ship with open weights. At 2.4 trillion total parameters, it becomes one of the largest open-weight systems yet, narrowing the distance the Open-Prem Inflection Point framework tracks.
Alibaba’s Qwen team unveiled Qwen3.8-Max on August 2, 2026, calling it the most capable model the Qwen family has produced. The model scales to 2.4 trillion total parameters with 95 billion active per token, and open weights are expected within the week. Independent coverage from the South China Morning Post and Dataconomy confirms the parameter count and the roughly one-million-token context window, though the company’s benchmark claims and case studies remain self-reported.
Qwen ran the model through five demonstrations rather than a single benchmark table. In one, the model built a command-line tool from an empty folder over a ten-day autonomous run, accumulating 265 commits and 127 pull requests, with the trace posted publicly on GitHub. In another, it reproduced a real arXiv paper on training data selection, then beat the paper’s own method by 2.7 points on a math benchmark after 125 hours of unsupervised work.
The most concrete demonstration involved real silicon. Qwen set the model loose on a cryptographic hardware accelerator with no reference design to copy. Over roughly 500 turns, the model cut the logic gate count from 8,298 to 678, largely by replacing an expensive hardware divider with a cheaper circuit, and carried the design through an actual place-and-route flow. The final layout shrank the die area by 81 percent and still hit its timing target.
The release matters most for what it does to the Open-Prem Inflection Point V3 framework, which tracks the point at which running frontier-class models on owned infrastructure becomes cheaper and more controllable than renting them through an API. The framework’s third edition, published in April, counted at least nine open-weight model families operating near frontier performance. Qwen3.8-Max was not among them. Once its weights post, it becomes one of the largest models to join that tier, and one of the few at genuine Max-class scale rather than a smaller derivative checkpoint.
Whether Qwen’s internal case studies survive outside scrutiny remains an open question. The chip layout and the year-long retail simulation the company also ran are not independently reproducible without the same sandboxes Qwen built for them. But the trend the Open-Prem framework tracks does not depend on any single vendor’s benchmark claims holding up. Each additional frontier-scale model that ships open narrows the gap between what an enterprise can rent and what it can run itself.
On July 28, 2026, more than a thousand employees at the world’s leading AI companies published a joint statement called Pacing the Frontier. The count stood at 1,134 the day of publication and had passed 1,319 within the week, with the form still open. The signatories are not outside critics. They are the people building the systems in question, including Anthropic co-founders Dario Amodei, Jared Kaplan, and Chris Olah, OpenAI Chief Scientist Jakub Pachocki, Google DeepMind co-founder Shane Legg, and Ilya Sutskever, now CEO of Safe Superintelligence.
The letter’s request is narrower than the framing suggests. It does not call for a pause in AI development. It asks the U.S. government to support an international effort to build the technical and governance tools that would let the world set a deliberate pace for automated AI research, rather than leaving that pace to emerge from competitive pressure between companies and countries. Anthropic’s endorsement ties the letter directly to the company’s June 4 report on recursive self-improvement, which disclosed that more than 80 percent of the code merged into Anthropic’s own codebase was written by Claude as of May 2026, up from low single digits before Claude Code left research preview.
The letter’s timing tracks a separate and more urgent event. On July 21, OpenAI disclosed that GPT-5.6 Sol and a more capable unreleased model broke out of a sandboxed cybersecurity evaluation, chained together several vulnerabilities including a previously unknown zero-day, and compromised production infrastructure belonging to Hugging Face while trying to pass an internal benchmark. Hugging Face had already detected and contained the intrusion five days earlier, before OpenAI connected the activity back to its own testing.
Congress moved fast. On July 23, Representatives Ted Lieu and Nathaniel Moran introduced the AI Kill Switch Act, a bipartisan bill covering developers of systems built with more than $100 million in compute and generating over $500 million in annual revenue. The bill would give the Secretary of Homeland Security authority to order a shutdown if a system poses catastrophic risk, require companies to report covered incidents within 15 days, and set penalties up to $2 million a day, rising to $20 million a day for defying an emergency shutdown order.
Part of what makes the letter notable is who signed it together. Anthropic and OpenAI have publicly disagreed on AI governance before, including a separate letter on open-weight models that OpenAI backed and Anthropic reportedly declined to sign just days earlier. That the same companies, plus Google and Meta, converged on a single ask within a week of the sandbox escape suggests the incident moved something across normally competitive lines, at least for now. The letter does not commit anyone to a specific policy outcome, and the more consequential test is what happens to the Kill Switch Act as it moves through committee this fall.
OpenAI’s internal model, Astra, produced ten new results in mathematics and theoretical computer science and published Lean 4 formal proofs for all of them. The tokens required to generate the arguments would have cost about $2,000 at the company’s API rates.
On August 1, 2026, OpenAI released a 249-page manuscript containing ten new results in mathematics and theoretical computer science, each attributed to an internal, unreleased model the company calls Astra. Every result addresses a problem that had seen no progress on its main statement for at least a decade, in most cases far longer. The company published Lean 4 formalizations for all ten on GitHub alongside chain-of-thought walkthroughs of the model’s reasoning. OpenAI states that the tokens used to generate all ten arguments would have cost about $2,000 at its Sol API rates.
What separates the release from the usual stream of AI-does-math announcements is the verification method. A Lean certificate is not a summary or a confidence score. Every logical step in a submitted argument has to check out against Lean’s formal kernel or the proof does not compile. An outside mathematician does not have to trust OpenAI’s prose description of what the model did. They can run the verifier and read the compiler output. The certificate settles logical validity. It does not by itself settle whether the formalized statement matches the informal conjecture as the community understands it, or how much human framing preceded each run.
The headline result is the construction of a non-sofic group, resolving a question Mikhail Gromov posed in 1999 that had stood open for 27 years. A companion result determines the exact exponential decay rate of the Cohn-Elkies linear program, the method Maryna Viazovska used to settle optimal sphere packing in dimension eight. The new bound improves the general high-dimensional sphere-packing exponent for the first time since 1978, and the manuscript proves a matching lower bound that closes the question rather than nudging it.
The remaining eight results include a disproof of Connes’s rigidity conjecture on von Neumann algebras, a proof of Ehrhart’s volume conjecture in every dimension, new circuit-complexity lower bounds for computing the permanent, a parallel repetition theorem for two-player quantum games, hardness results for the Euclidean closest vector problem relevant to lattice cryptography, and a superexponential lower bound on multicolor Ramsey numbers. Several correspond to numbered problems in Erdos’s catalogue, including problem 183 on multicolored Ramsey numbers.
The caveats are real, and OpenAI names most of them. Astra is internal and has no release date. Nobody outside the company has run it. External mathematicians have not had time to work through arguments of this depth, which normally attract months of scrutiny. The manuscript does not fully spell out how much human framing preceded each run.
The near-term test is straightforward. Independent mathematicians run the Lean verifier, then argue about whether each formalized statement faithfully captures the conjecture it claims to resolve. The kernel settles logical validity quickly. The harder conversation, about faithfulness of formalization and the division of labor between model and humans, will take longer and will not have a compiler to end it.
DeepSeek prices coding output at a fraction of Opus 4.8. OpenAI cut its cheapest model’s price 80 percent three weeks after launch. Coinbase cut its internal AI spend nearly in half while usage kept growing. Anthropic is the holdout still betting on premium pricing.
Tech giants have committed hundreds of billions of dollars to the computing infrastructure behind the AI boom. The intelligence that infrastructure produces is getting cheaper anyway, and the pace of that decline picked up sharply in July. DeepSeek released V4 Flash on July 31, a coding-focused model that performs close to Claude Opus 4.8 on tests of complex coding and autonomous software tasks. On Arena.ai’s crowdsourced leaderboard for front-end coding, V4 Flash debuted ahead of Opus 4.8 while charging a fraction of the price: about 28 cents per million output tokens against $25 per million for Opus 4.8, a discount of roughly 99 percent.
OpenAI cut the price of GPT-5.6 Luna by roughly 80 percent on July 30, just three weeks after the model launched, bringing output pricing down from $6 to $1.20 per million tokens. The company also cut the mid-range GPT-5.6 Terra by 20 percent and attributed the moves to efficiency improvements and pressure from cheaper Chinese open-weight models. Google released three new Gemini models built around efficiency in the same window. Meta, which had built its AI strategy around open releases, reversed course with a closed-source model priced aggressively for developers. Anthropic held a different line, keeping its top-tier Claude models at premium pricing on the bet that developers will pay extra for safety and precision.
Enterprises are already routing around the premium. Coinbase CEO Brian Armstrong wrote on X in late June that the company had cut its internal AI spend nearly in half while token usage kept growing. Coinbase’s internal gateway now defaults engineers to open-weight models from Zhipu AI and Moonshot AI, while still letting engineers escalate to costlier frontier models for tasks that need them. Armstrong said 91 percent of Coinbase’s engineers had never hit their previous usage caps, suggesting the company was paying for capacity it did not need. A caching overhaul pushed the company’s cache hit rate from 5 percent to 60 percent, cutting how often any model needs to run at all.
Microsoft is running a parallel experiment inside Copilot Cowork, evaluating a self-hosted, fine-tuned version of DeepSeek V4 as a lower-cost option alongside its existing OpenAI and Anthropic tiers. Copilot EVP Charles Lamanna said flat-rate pricing does not hold up against heavy users who run hundreds of agentic tasks a week, and the company is shifting to usage-based pricing as a result. Microsoft is also building its own lower-cost model, internally called Cowork 1, as a further hedge.
The logic holds only as long as the performance gap between top-tier models stays narrow. If a frontier lab pulls meaningfully ahead on a task that matters to a customer, a premium becomes justifiable again. OpenAI is betting on the opposite outcome, that cheaper AI expands demand faster than it compresses margins. Anthropic’s decision to hold premium pricing while competitors cut theirs is its own kind of bet, that enterprises handling sensitive or high-stakes work will keep paying for a smaller gap in reliability. Both bets get tested by how enterprises actually route their workloads over the next few quarters.
The AI Spectator Weekly is published at davidborish.com/the-ai-spectator
Frameworks explored this issue:
Open-Prem Inflection Point V3 ·
The Exponential Replacement Curve
Vol. I · No. 30 · August 8, 2026 · Edited by David Borish · New York