
Washington is now a review step in the frontier AI release pipeline. GPT-5.6 (Sol, Terra, and Luna) launched to a guest list of 20 government-vetted organizations. Meta launched a new cloud service and re-entered the race with Muse Spark 1.1. Nvidia Vera Rubin started shipping. Where's Google?
We're launching a new series of events that bring together LLM labs and neoclouds, hands-on presentation on how to run and fine tune open models. We're starting with Nvidia Nemotron on Massed Compute. Up next Saturn Cloud, zAI, gpt-oss, Vast, Nebius and more.

*June 29 – July 2, San Francisco *
One of our favorite engineering conference, spanning from agents, coding, infra, eval, robotics.

Roberto with the Ratel Team and I (Fed) were here running a side event at Frontier Tower: https://luma.com/sf-socratic-1.0.
*July 18–19, San Francisco *

Frontier models, robotics, safety, and intelligent systems
*Aug 1–2, Berkeley *

Summit on AI agents, infrastructure, safety, and real-world applications.
*Aug 4–5, Las Vegas *

Enterprise AI conference on adoption, governance, and deployment.
On June 30 the Commerce Department lifted the June 12 export-control directive on Fable 5 and Mythos 5, ending a 19-day shutdown of the most capable model ever pulled from the market. Anthropic's redeployment post details the price of freedom: a new safety classifier blocking the Amazon-flagged vulnerability-discovery jailbreak in >99% of attempts, a HackerOne bounty program for Fable jailbreaks, pre-release government review of future frontier models, and a cross-lab jailbreak-severity rubric built with Amazon, Microsoft, and Google. Fable 5 returned globally on July 1; Mythos 5 came back only for select US organizations. Axios published the behind-the-scenes of the standoff — including the detail, awkward for everyone involved, that Amazon CEO Andy Jassy reported the jailbreak to the Treasury Secretary before telling Anthropic. Amazon remains Anthropic's largest investor. Thanksgiving dinner will be tense.

As a reminder of what all the fuss is about:

Unbanned, but not unmetered. The victory lap lasted about a day. Subscribers discovered Fable 5 access now caps at 50% of weekly usage limits through July 7, then moved to July 12, and then July 19. After that it will require separate usage credits at $10/$50 per million tokens.
On June 30 Anthropic's "most agentic Sonnet yet", near-Opus-4.8 performance — 63.2% SWE-bench Pro, 80.4% Terminal-Bench 2.1, and a GDPval-AA score of 1,618 that edges out Opus 4.8's 1,615 — native 1M-token context by default, 128K max output, and the new default in Claude Code and on Free/Pro plans. Intro pricing of $2/$10 per million tokens through August 31 (then $3/$15) — read by everyone as pre-IPO land-grab pricing. Simon Willison flags the new tokenizer (1.0–1.35x more tokens — check your budgets) and new cyber-safety refusal behavior.

Also June 30 a beta research workbench bundling 60+ preconfigured tools for genomics, proteomics, structural biology, and cheminformatics — plus an internal drug-discovery program targeting neglected diseases and up to $30K in credits for ~50 research projects (applications due July 15).
The third Anthropic Economic Index report, "Cadences" (June 26): hourly usage sampling plus a ~9,700-user survey. Standout finding: people who delegate the most to Claude are the most optimistic about their careers. Also, sleep-advice queries peak at 5 a.m., which is grim
On June 26 OpenAI previewed the GPT-5.6 family:
The initial access is limited to ~20 government-approved organizations for cybersecurity review — the first major release shipped through the June 2 executive order's 30-day vetting framework. Having watched what happened to Anthropic, OpenAI apparently decided the ban works better as a pre-order.
Per the FT (via CNBC), OpenAI proposed handing the US government a 5% equity stake (~$42.6B at the last private mark) to defuse political pressure — under a framework where Anthropic, Google, and Meta would cede similar stakes into a sovereign-wealth vehicle. Anthropic says it's had no such discussions. Trump has previously called government ownership in AI companies "a beautiful thing." Bernie Sanders proposed 50% last month; call it a negotiation.
OpenAI is leaning toward delaying its IPO to 2027: Altman reportedly called any cut to the $1 trillion target a "non-starter," advisers pointed at SpaceX's rocky post-IPO tape, and CFO Sarah Friar is telling associates 2027. Leadership now expects Anthropic to list first — likely October.
Open AI unveiled its first custom inference chip, co-developed with Broadcom: a reticle-sized ASIC on TSMC 3nm with stacked HBM and Tomahawk 6 networking (1.6 Tbps), which went from design to tape-out in ~9 months — with OpenAI's own models assisting the design. Hock Tan claims ~50% cost savings versus GPUs; engineering samples are already running GPT-5.3-Codex-Spark, deployment targeted by end of 2026, and Microsoft has reportedly committed to absorbing a big share of output.

On June 18, Noam Shazeer — Gemini co-lead and "Attention Is All You Need" co-author — announced he's leaving for OpenAI, barely two years after Google paid ~$2.7B to bring him back from Character.AI. On June 19, John Jumper left for Anthropic. On June 22, Google confirmed Gemini 3.5 Pro is slipping to July — token-efficiency, coding, and multi-step reasoning reportedly not yet at flagship standard — and Alphabet shed roughly $225B in market value in a day. As of July 6 the model was "cleared for July launch" but still hadn't shipped. Hassabis, at Cannes, insisted DeepMind still has "by far the biggest and broadest research bench of any of the labs.". Bench is broad, but many are departing, and Gemini is not the great models they advertise it for.
Google told Meta it can't supply as much Gemini capacity as Meta wants, forcing Meta engineers to conserve tokens and delaying internal projects. Meta was running safety and internal workflows on Gemini, and Google, which actually was itself renting 110K GPUs from SpaceX/xAI as bridge capacity.
On July 1 Meta announced Meta Compute, a business selling hosted model access and raw GPU compute in direct competition with AWS, Azure, and Google Cloud — an attempt to turn its $115–135B 2026 infrastructure spend from cost center into revenue. The market loved it: Meta closed above $600 for the first time (+8.8%) while CoreWeave fell 14% and Nebius 17%.
🔔 Note: GPUs are not just a tool to train models they're now an asset that sustain the stock market of the hyperscalers, as GPU price increases the asset increases with it.
On July 9 Meta released Muse Spark 1.1, a significant upgrade from the first Muse Spark, andlaunched the public preview of the new Meta Model API alongside it.
"Today we're releasing Muse Spark 1.1 — a strong agentic and coding model at a very low price. Available through our new Meta Model API and in Meta AI." — Mark Zuckerberg
Alexandr Wang said Muse Spark 1.1 is "an industry-competitive agentic and coding model; across many agentic evals it rivals GPT-5.5 and Opus 4.8".
The Meta Model API is Meta's first move to sell its models as a developer platform — and it lands a week after Meta Compute. Together they sketch the same strategy from two sides: turn a $115–135B infrastructure bill into a revenue line, and compete for developers on price ("very low price" is doing deliberate work in that launch copy, this is the token-billing-shock month, and Meta knows it). Skeptics may remember the failure of Llama 4 Behemoth despite having incredible benchmarks.
Sources: @AIatMeta announcement, Mark Zuckerberg, Alexandr Wang
Meta is also releasing a prediction-market app, codenames "Antwerp"/"FBForecast", that auto-generates questions from trending topics.

Grok 4.5 entered private beta at SpaceX and Tesla (June 28) — built on the 1.5T-param V9 base, trained partly on Cursor coding data, and per Musk "close to, perhaps exceeding Opus" — benchmarked exclusively by companies Musk owns, so calibrate accordingly. He also promised a new from-scratch model every month through end of 2026. xAI also shipped a no-code Voice Agent Builder (July 1) aimed at the AI call-center market, and — after a month of viral AI-generated Iliad and Odyssey trailers pushed Grok Imagine to ~314M visits — Musk declared Grok Imagine "done" (July 5), a sentence that has never once been true of any software.
AWS committed $1B to a Forward Deployed Engineering unit (June 30); Microsoft answered 48 hours later with Microsoft Frontier Co. — $2.5B and 6,000 experts embedded with clients to make enterprise AI deployments actually work. Palantir's business model is now big tech's product category.

The open-source AI war is escalating.
Moonshot ships Kimi K3, a 2.8-trillion-parameter sparse MoE — the largest open-weight release to date — with a 1M-token context window, native vision, an always-on "thinking mode," and two new architectural tricks (Kimi Delta Attention and Attention Residuals). Moonshot concedes it still trails Fable 5 and GPT-5.6 Sol overall, but K3 took first place in four of eight real-world automation benchmarks (Automation Bench, SpreadsheetBench 2, BrowseComp among them) and second to Fable 5 in most of the rest. Pricing is $0.30 / $3 / $15 per million cached-input / input / output tokens, with full weights landing July 27 — just ahead of the World AI Conference in Shanghai.
DeepSeek is actively preparing for an IPO on China's mainland STAR Market. The company has started working with accounting firms and investment banks, with a potential confidential filing as soon as late 2026, targeting a 2027 debut. It also confirmed a mid-July launch for V4-Pro: 1.6T params/49B active; V4-Flash: 284B/13B, both 1M context. Also DeepSeek API introduces a 2x surge price during off-peak hours (outside Chinese business time) — congestion prices is a sign yet that inference is infrastructure. Sources: DeepSeek V4 pro/flash
The delivery-app giant open-sourced a 1.6-trillion-parameter agentic coding model (June 30) with a 1M-token context window — and claims it is the first trillion-parameter model to complete both pre-training and inference on a 50,000-card cluster of domestic Chinese chips. It scores 59.5 on SWE-bench Pro, above reported GPT-5.5 and Sonnet-class scores. Whatever export controls were supposed to prevent, it wasn't this.
Nathan Lambert's Interconnects essay called it "the step change for open agents" — a "DeepSeek moment" for agents. While Fable 5 was suspended, GLM-5.2 took #1 on Design Arena (~1360 Elo on HTML web design), now ranks above Anthropic's models by token usage on OpenRouter, and sits 5th on the Artificial Analysis leaderboard (top open-weight model). Semgrep published cyber benchmarks under the immortal title "We have Mythos at home" — GLM-5.2 beating Claude on their suites.

If June was the month the government proved it could stop a frontier model, the two weeks since proved something more durable: it can price the permission. Fable came back unbound but metered; GPT-5.6 launched pre-gated. Meanwhile the actual frontier kept moving where the gates aren't — and this time it didn't just close the gap, it stepped into the room. Moonshot's Kimi K3, a 2.8-trillion-parameter open MoE and the largest open-weight release ever, took first place on four of eight real-world automation benchmarks and second to Fable 5 on most of the rest, with the weights themselves due July 27. "Six months behind" was last month's story; this month the lag is measured in benchmark points, not quarters. And while the frontier went open, the bill for the agentic revolution landed on ordinary developers' credit cards, 10x to 50x heavier than promised. Capability, cost, and control, all compounding at once. The models got unbanned, the weights got free; nobody said anything about cheap.

GitHub Copilot's June 1 switch to per-token AI Credits closed its first full billing cycle June 30, confirming the projections: agentic users report effective costs 10–50x their old flat subscriptions ($29 → $750; $50 → $3,000). GitHub is leaning on promotional credits through August rather than reversing course.
The trend born of Meta's leaked "Claudeonomics" leaderboard with top employee: 281 billion tokens in 30 days, hit its backlash phase: Meta reportedly killed the internal leaderboard, Uber imposed $1,500/month AI spending tiers after blowing its annual AI budget in four months, and startup Lindy moved 100% of traffic from Claude to DeepSeek.

As threatened, the free Gemini CLI stopped serving requests June 18, hard-breaking every cron job and git hook that shelled out to gemini, to considerable HN grief. Migration path: the closed-source Antigravity CLI. Pour one out for -p "fix this".

"A Bun-in-Rust rewrite would cost ~$165k in API tokens" a quoted tweet on how much refactoring Bun in rust could have costed.
Sources: @jarredsumner (Jarred Sumner), blog post
Is this true?
Sources: https://x.com/MichaelArnaldi/status/2076326793343070556
Anthropic published new interpretability research: a global workspace in language models. The framing borrows one of neuroscience's most influential ideas. Of everything happening in your brain right now, only a tiny fraction is consciously accessible — the "global workspace" that broadcasts a handful of signals to the rest of the system. Anthropic reports finding a strikingly similar divide inside Claude: a huge amount of internal computation, and a narrow channel of it that actually reaches what the model says.
Anthropic open-sourced the "Jacobian lens" — a readout of what a model is about to say, extracted from its internal layers before it says it. It's white-box (you need the activations), but that's exactly what makes it interesting for anyone running open-weight models or building monitoring tooling. The community moved immediately: @EricBuess wired it into an agent hook within days of release.
This work also lands alongside Anthropic’s turn-averaged sparse autoencoders research, continuing the broader trend of moving interpretability beyond per-token analysis and toward understanding high-level model behavior at scale.

Sources: Announcement, @EricBuess: Jacobian lens with Qwen
Doom Loops is a common failure mode in small reasoning models where the model repeats a token span endlessly until the context window is exhausted, especially on hard math and coding tasks.

Key results:
How it works:
Additional insights:
Source: @liquidai (Liquid AI)
The humanoid sector went public and went to work.
Boston Dynamics unveiled a redesigned Atlas (July 2) — "almost an order of magnitude reduction in complexity," built with Hyundai manufacturing muscle for 30,000 units/year — days after Hyundai took full ownership and announced a $100M Waltham robotics center. Then it put Atlas on the pitch at a FIFA World Cup Round-of-16 halftime show (July 5) — the first robotics activation in a live match, capping a campaign where Atlas learns a "Ghost Rabona" with no CGI.
Agility Robotics is going public via a $2.5B SPAC (June 24, ticker: AGLT) — the first US-listed pure-play humanoid company — and Unitree won final approval for its ~$618M Shanghai STAR IPO (July 3) at a ~$6.2B valuation, China's first listed humanoid maker, off 5,500+ robots shipped in 2025. Profitably, which somehow still feels illegal in this industry.
Brett Adcock posted the chart (June 20): 750+ deployed robots vs ~250 employees, making Figure the first company of meaningful scale whose robot fleet exceeds its headcount. Days later Figure 03 started real work at BMW Spartanburg on a harder sequencing task under its Helix 02 vision-language-action model — 40 units already bill ~$25/robot-hour. HR implications unclear.
BIS annual report (June 28) names an AI investment bust, circular financing, assets potentially "pledged multiple times" across equity-debt-supplier-client webs, and record sovereign debt as interlocking top risks — noting the five largest hyperscalers will spend >$1 trillion on AI capex across 2025–26, outrunning free cash flow. Comparisons deployed: canal mania, 1840s railways, dot-com.
Despite January's approval-with-25%-surcharge deal, Jensen told shareholders the company has generated zero H200 revenue in China, as Beijing steers buyers to Huawei. Bernstein sees Nvidia's China AI-chip share collapsing from ~40% to ~8% this year with Huawei taking ~50%. Read next to LongCat-2.0 above: the export-control debate is becoming moot — China is simply exiting the customer list.
The Council gave final approval to the Digital Omnibus (June 29): high-risk obligations pushed to December 2027, but GPAI enforcement — fines up to €15M or 3% of global turnover for model providers — still switches on August 2, 2026. Twenty-seven days, frontier labs. Also new: an outright Article 5 ban on nudification apps. 🇪🇺
The White House is racing to finalize a voluntary framework with OpenAI, Anthropic, and Google — benchmarks for cyber-capable models, testing timelines, access rules — to make releases predictable after a month in which both OpenAI's and Anthropic's newest models were limited to administration-approved customers. Meanwhile the Five Eyes cyber agencies issued a rare joint statement (June 22) warning AI is compressing cyber risk timelines to "months, not years," and the Colorado AI Act died on its own effective date (June 30) — sued by xAI, kneecapped by a DOJ intervention, repealed and replaced with a narrower 2027 law. Between the framework, the GPT-5.6 guest list, and the 5% stake proposal, the US now has an industrial policy for frontier AI; it's just written one crisis at a time.
Here's a very true video of Senator Mitch McConnell proving to be in the hospital alive and healthy.
The MIRI president argued that even lab leaders privately want off the race (June 26), and that the US should negotiate a globally enforceable option to pause frontier development — an "off switch" — while exempting narrow AI for medicine and science.
Fernando Borretti's viral essay "No-One Escapes the Permanent Underclass" (June 25) argued the popular hedge — accumulate capital or join a lab before full automation — fails, because workers, then owners, then governments get disintermediated in turn. Ozy Brennan's rebuttal kept the argument going all week.
At Stanford (June 18), the DeepMind chief said we're in "the foothills of the singularity" and re-upped his AGI timeline: "2030 is when I expect it to arrive, plus or minus a year."
Fernando Irarrázaval published the results of HackMyClaw (June 26) — 2,000+ attackers, 6,000+ attempts over months, all trying to trick his OpenClaw email agent (Claude Opus 4.6, few-line security prompt) into leaking a secrets file. Every attempt failed. Best moment: around email 500, the agent noted in its own memory that the attack volume "suggests a coordinated security exercise rather than organic malicious activity." Stay safe out there, Fiu.

94 minutes on AI and the future of math, on why math is where superintelligence shows up first, what an AI proof of the Riemann Hypothesis would actually tell us, and the "grindability vs. verifiability" frame for which domains fall next. Also, career advice for math students, which increasingly resembles career advice for everyone.
"An Appetite for More" — AI acceleration, aesthetics, and stamina. Cowen: the AI race "may never get settled" — we're in the first inning. Qureshi's advice for the AI-fatigued: "pace yourself." Noted, with gratitude, by this newsletter's authors.
Last year we covered AI 2027, a month-by-monty essay of what to expect. AI 2040 is an update version with 4 different possible outcomes — the blog post focuses on Plan A. Instead of a secretive corporate race to artificial general intelligence (AGI), Plan A advocates for an international treaty (primarily between the US and China) to coordinate a verified slowdown of frontier AI development.
Plan A steps
The Proposed Timeline
Immediate Policy Recommendations
Source: @DKokotajlo (Daniel Kokotajlo)
Source: @satyanadella (Satya Nadella)
Source: @bayeslord
Palantir CEO Alex Karp argues that open-weight models give enterprises and governments critical data sovereignty, cost predictability, and IP protection—advantages he believes closed frontier models lack. Karp criticizes token-based AI pricing from labs like OpenAI and Anthropic, calling it a “wealth tax.” His concerns:
Karp’s case for open models:
However, Karp argues that open models alone are not enough. Enterprise value comes from combining models with compute, security, and application layers that govern data and workflows—an area where Palantir positions its AI Platform, of course they need to sell you something.
... the hope 😰 ...

If June was the month the government proved it could stop a frontier model, the two weeks since proved something more durable: it can price the permission. Fable came back unbound but metered; GPT-5.6 launched pre-gated. Meanwhile the actual frontier kept moving where the gates aren't — open weights out of Beijing and Hangzhou, trained on domestic Huawei silicon, they're still six months behind but closing — and the bill for the agentic revolution landed on ordinary developers' credit cards, 10x to 50x heavier than promised. Capability, cost, and control, all compounding at once. The models got unbanned; nobody said anything about cheap.
Get the latest AI insights delivered to your inbox. No spam, unsubscribe anytime.
Founder, Engineer
New York City
Founder
Milan, Italy
Founder
Milan, Italy
Head of Engineering
Milan
The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.
DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes