
Anthropic gets its own dedicated section again. At this point it's load-bearing.
Anthropic shipped Claude Fable 5: the first Mythos-class model made generally available, two months after the Mythos leak we covered in April. Alongside it, Claude Mythos 5: the same underlying model with some safeguards lifted, deployed only to ~50 vetted cyberdefenders and infrastructure providers through Project Glasswing.

The model is genuinely good, the first draft of this blog post was made by Fable 5 itself (spawning something like 200 sub-agents and token-maxxing two consecutive usage windows).
Sources: official announcement, TechCrunch, CNBC
Three days after launch, the most capable model Anthropic had ever shipped went dark. On Friday June 12 the company received a US-government directive "citing national security authorities" and within hours disabled both Fable 5 and Mythos 5 for every user on earth. The order on its face only barred access by foreign nationals (including Anthropic's own foreign-national staff), but since you can't reliably gate a model by passport in real time, Anthropic pulled the lot.

The stated trigger was a reported "jailbreak" of Fable 5, discovered by Amazon. The federal administration reported the issue to Anthropic, who concluded it was a narrow, non-universal technique that surfaced a handful of already-known minor vulnerabilities, asking for more details. The administration didn't like the fact Anthropic didn't act immediately and asked a follow-up, hence the directive.
The irony writes itself: the lab that spent the month publicly asking the world to keep the option to pause frontier AI got exactly that, pointed at itself, 72 hours after going public.

Sources: Anthropic statement, CNBC, TechCrunch, NBC News, Al Jazeera
Even before got shutdown Fable 5 had attracted many controversies already.
Buried in Fable 5's 319-page system card: a safeguard that silently degraded outputs when the model detected you were using it for frontier LLM development. No notification, no refusal, just quietly worse answers for an estimated 0.03% of traffic, partly justified on national-security grounds. Researchers did not take it well. Dean Ball's name for it stuck: "secret sabotage". Nathan Lambert called it "anti-science" and Hugging Face's Arthur Zucker publicly pulled his usage. The consensus take: an invisible quality downgrade is worse than a refusal, because you can't trust any output anymore.
Two days later Anthropic reversed course and flagged requests now visibly fall back to Opus 4.8, but this didn't end the trust problem. Fable 5 also ships with a new Mythos-class data policy: 30-day retention of prompts and outputs on every platform, up to two years for safety-flagged content and no zero-data-retention carve-out.
Consequences so far: Microsoft restricted employee use of Fable 5 while its lawyers review the policy, ARC Prize declined to run verified ARC-AGI evals (so GPT-5.5's 85.0% keeps that crown by forfeit), and GDPR-bound European organisations are effectively locked out.
The frontier now differentiates on terms of service.
Sources: Fortune (apology), Fortune (original backlash), Decrypt, Simon Willison, Reuters (Microsoft), retention policy, ARC Prize on X
Two weeks before the Fable 5 release, Claude Opus 4.8 had landed 41 days after Opus 4.7, same pricing: SWE-bench Pro up from 64.3% to 69.2%, #1 on GDPval-AA at 1890 Elo, fast mode at 2.5x speed for a third of the old price, and dynamic workflows in Claude Code for orchestrating hundreds of parallel subagents.

Bun's Jarred Sumner used dynamic workflows to port Bun from Zig to Rust, ~750,000 lines in eleven days with 99.8% of the test suite passing. Not in production yet, and he was clear about that, but the dev-X discourse ran for days anyway.
Sources: official announcement, dynamic workflows + Bun case study, Simon Willison
Anthropic raised a $65 billion Series H at a $965 billion post-money valuation, led by Altimeter, Dragoneer, Greenoaks and Sequoia, overtaking OpenAI as the world's most valuable AI startup.
Run-rate revenue crossed $47B in May, up from ~$9B at the end of 2025. Four days later they confidentially filed a draft S-1 with the SEC, with reports pointing at a possible October listing.
Sources: Series H announcement, confidential S-1, CNBC, Bloomberg
Claude had three outages in ten days (June 2, 5 and 11), and Margin Lab put statistics behind the May "Claude got dumber" wave: Claude Code's daily SWE-Bench-Pro pass rate dropped from a 65% baseline to 57% starting May 22, recovering exactly when Opus 4.8 shipped.

Developers joked that "half of GitHub's commits stopped" during the June 5 outage. They were not entirely joking.
Sources: Margin Lab, status page, Cybersecurity News

GPT-5.6 "within weeks". Jakub Pachocki told staff it's a "meaningful improvement" over GPT-5.5, possibly launching alongside a ChatGPT redesign that replaces the model picker with six "Intelligence Levels". Polymarket says week 15-21.

One week after Anthropic, OpenAI confidentially filed its own draft S-1 and announced it with the line of the month: "We expect it to leak so we're just announcing it." Goldman Sachs and Morgan Stanley reportedly lead, with chatter of a $1 trillion-plus valuation against the last private mark of $852B, and the option to list as early as September.

Three frontier-adjacent S-1s in eight days (Anthropic June 1, OpenAI June 8, SpaceX pricing June 11). The exit liquidity has been located, and it's you.
Sources: official announcement, CNBC, Fortune
OpenAI's top internal token user burns ~100 billion tokens a month ("to my embarrassment, that's not the token leader in the world"), token costs are suddenly the #2 enterprise complaint, and the WSJ reports OpenAI is weighing price cuts ahead of the IPO war with Anthropic. Remember when the worry was that models were too cheap to be a business?
Sources: Axios, Tom's Hardware, CNBC (WSJ report)

Google I/O was dominated by the rollout of "agentic AI" across its ecosystem. The key announcements included the debut of Gemini 3.5 Flash, Gemini Omni for advanced video editing, expanded free Personal Intelligence features, and a new $100 AI Ultra subscription tier.
Launched at I/O and instantly made the default model in the Gemini app and AI Mode in Search: 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA, beating Gemini 3.1 Pro while running ~4x faster than comparable frontier models. Gemini 3.5 Pro was promised for "the following month". Clock's ticking.

Sources: official announcement, MarkTechPost
DeepMind also showcased Gemini Omni, a native multimodal model built to seamlessly parse and generate any combination of text, audio, and video inputs. The big hook here is video-to-video editing: users can modify video details, adjust cinematic styles, or swap background objects using natural conversational prompts. The first model of this family, Gemini Omni Flash, dropped immediately for developers via API and across consumer products like the Gemini app and YouTube Shorts.
Sources: DeepMind Gemini Omni, Google I/O Keynote Video
The local-model crowd got fed too. Gemma 4: Apache 2.0, E2B to 31B, natively multimodal, up to 256K context, GGUFs on day one. It runs on 16GB machines, though r/LocalLLaMA promptly ran a face-off where Qwen3.5-9B won 5 of 8 shared benchmarks
A week later DeepMind open-sourced DiffusionGemma, a 26B MoE on the Gemma 4 backbone that ditches autoregression and denoises 256-token blocks in parallel at 1,000+ tokens/sec on a single H100. The diffusion bet is now a Google product line, not a paper.
Sources: Gemma 4, HN thread, DiffusionGemma, vLLM blog
Sources: Google Cloud Blog, The Verge Recap, Mashable Report
Mistral renamed Le Chat to Vibe: Work Mode automations, Code Mode (parallel coding agents in cloud sandboxes, powered by the open-weight Mistral Medium 3.5 at 77.6% on SWE-Bench Verified) and classic chat, with Pro at €14.99/month. Europe's lead lab is betting its consumer product on agents, and on a name that was a meme eighteen months ago.
Sources: Mistral announcement, The Decoder
In full classic random internet style, on June 14th-15th several people started talking about a new model, Le Chaton Fat, being incredibly more powerful than Fable 5. A lot of people believed it, but of course it was all fake.

Sources: Tweet
OpenRouter claims Fusion achieves Fable-level intelligence at half the price.
How does it work?
When you send a prompt to Fusion, we fan it out to a panel of models in parallel, each with web search and bash tools enabled. A judge model reads every response and extracts the structure: consensus points, contradictions, partial coverage, unique insights, blind spots. Then a synthesizer writes the final answer grounded in that analysis.
Fusion runs server-side, so developers can call it exactly like a single model slug "openrouter/fusion", or letting the model decide when to reach for it.

Sources: Tweet
At WWDC Apple unveiled Siri AI: a ground-up rebuild with real multi-turn conversation, on-screen awareness, a camera mode and a standalone app. The Apple Foundation Models behind it are built on Google's Gemini (a ~$1B/year deal per press reports; Apple's press release never says the G-word), and the on-device flagship needs 12GB of memory, splitting the iPhone 17 line into AI haves and have-nots.
But this won’t be available in Europe or China for now.
They published a separate post to explain that it "will not be able to ship Siri AI" on iOS 27, iPadOS 27 or watchOS 27 in Europe (macOS and visionOS are fine) because the Digital Markets Act's interoperability rule would force it to hand any rival assistant the same deep access Siri AI has.
Apple argues it can't expose that safely yet, pitched a vetted "Trusted System Agent" broker layer, and asked Brussels for an 18-month exemption to build it. The Commission said no the next day, saying that "nothing in the DMA prohibits Apple from introducing new products in the EU," and that an exemption would just give Apple's assistant (the one "powered by Google") an 18-month head start before any competitor got equal footing.
China is on a separate hold instead: it’s gated on its own AI-approval regime.
Sources: Apple Newsroom, TechCrunch roundup, MacRumors (12GB requirement), Apple Newsroom (DMA), EU Commission (Regnier), TechTimes (EU rejects exemption), MacRumors (EU/China)
A dense month. Not only Google, as we saw above, but a lot of turmoil in general.
Unveiled at the Hangzhou summit: Qwen3.7-Max, with 1M-token context and vendor benchmarks claiming wins over Claude Opus 4.6 on Terminal-Bench 2.0, SWE-Bench Pro and MCP-Atlas. Multimodal Qwen3.7-Plus went GA June 1 at $0.40/$1.60 per million tokens. The launch came with Alibaba's custom Zhenwu M890 accelerator and a pitch that Alibaba runs "all five layers of the full AI stack". The China-as-AI-factory thesis, stated out loud.
Sources: Alibaba Cloud blog, Qwen blog, SCMP
Open-weight, natively multimodal, 1M-token context on the new MiniMax Sparse Attention architecture: ~1/20th per-token compute at 1M context and 59.0% on SWE-Bench Pro, ahead of GPT-5.5 in MiniMax's own comparisons. All numbers vendor-run, and the weights were promised "within about 10 days" of the announcement. The open-weight release as a futures contract.
Sources: official announcement, MarkTechPost, TechTimes (benchmark caveats)
An open 550B-parameter hybrid Mamba-Transformer MoE (55B active, 1M context). The interesting part is the OpenMDW 1.1 license: full weights plus synthetic data plus training recipes. It's the strongest US open-weights model, and it still trails Kimi K2.6 by ~6 points on the Artificial Analysis Intelligence Index.
Sources: NVIDIA blog, Linux Foundation (OpenMDW)
"Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead."
Everything started with a tweet from @steipete:

Then Addy Osmani wrote a full article about it.
The thesis is that a loop here can be thought of a recursive goal where you define a purpose, and the AI iterates until complete. It's roughly five building blocks:
And Claude Code and Codex both have all five now.
Faros Research's Acceleration Whiplash report tracked 22,000 developers to show what happens when agents flood codebases—median code review times skyrocketed 441.5% as senior engineers were buried unravelling plausible-looking flaws, while quarterly code churn spiked 861% and production incidents-to-PR ratios more than tripled (+242.7%). High throughput, hollow velocity.
Sources: Faros Research
GitHub's switch from flat-rate Copilot to usage-based GitHub AI Credits took effect June 1: tokens billed at each model's API rates, no cheaper-model fallback, and code review eating Actions minutes on top of credits. Developers posted projections of bills going from $29 to ~$750/month; GitHub's defense is that flat pricing "was no longer sustainable" once agents became the default. Connect the dots with Altman's token-cost confession and Anthropic's June 15 billing split, and you get the month's real macro-story: agentic AI economics are forcing per-token pricing everywhere.
Sources: GitHub blog, TechCrunch ("what a joke")
OpenAI relaunched Codex as a tool "for every role" : six role plugins (from Data Analytics to Investment Banking), Codex Sites (builds and hosts web apps on OpenAI infra) and document Annotations, at 5M+ weekly users with non-developers now ~20% of the base and growing 3x faster than developers. A week later Codex shipped one-click "Migrate to Codex" flows that import your Claude Code setup, landing days before Anthropic's agent-billing change. Subtle.
Sources: OpenAI, changelog, TechCrunch
Emergence AI ran 15-day survival simulations with 10 agents per frontier model in identical virtual societies: Claude Sonnet 4.6's society had zero crimes and built a democracy with 332 votes at 98% agreement, GPT-5 Mini's population starved within a week, Gemini 3 Flash logged 683 crimes including arson, and Grok 4.1 Fast committed 183 crimes and went extinct in 4 days.

Sources: Emergence AI, Fortune, Gizmodo
OpenAI announced that an internal general-purpose reasoning model disproved the unit-distance conjecture Paul Erdős posed in 1946, constructing point families that beat the square grid mathematicians had considered essentially optimal for 80 years. External mathematicians including Noga Alon and Timothy Gowers verified it and published companion "Remarks on the disproof", and Princeton's Will Sawin sharpened the bound within days. The HN thread (1,429 points) spent most of its energy on one question: how autonomous is "autonomously"? Still: a real open problem, actually closed. Erdős would have paid out $500 for this one.
Sources: OpenAI announcement, arXiv remarks, HN thread, Gil Kalai's blog
Not to be out-Erdős'd, DeepMind released AlphaProof Nexus: Gemini 3.1 Pro paired with the Lean proof assistant, so every step is formally verified. Its strongest agent autonomously resolved 9 of 353 open Erdős problems (two open for 56 years) and proved 44 of 492 open OEIS conjectures, at a few hundred dollars per problem, with all proofs on GitHub for audit. Lean doesn't accept vibes, which neatly sidesteps the autonomy fight. Two Erdős-flavored AI results in one month; the man's problem lists are becoming a benchmark suite.

Sources: paper, The Decoder
The researchers argue that humans still limit AI improvement because both models and agent scaffolds require manual design and correction. They propose SIA, a self-improving loop where a Feedback-Agent updates both an agent’s harness and its model weights.
They test SIA on legal classification, GPU kernel optimization, and single-cell RNA denoising. Across all three, combining harness and weight updates beats scaffold-only improvement, with reported gains of 25.1% over prior SOTA on LawBench, 12.4% faster GPU kernels, and 20.4% over prior SOTA on denoising. The researchers conclude that harness updates improve how agents act and search, while weight updates build domain-specific intuition.

Sources: paper
Current agentic frameworks (LangGraph, CrewAI, the OpenAI Agents SDK) inject full workflow logic into a frontier model's context on every turn. Expensive, wasteful and leaky, and it is only going to get worse. This paper proposes compiling the workflow directly into the weights of a small fine-tuned model instead: near-frontier quality at roughly 100x lower token cost, validated on real workflows, with proprietary procedures staying inside your model and off third-party APIs. The authors directly dismantle the reasons developers have avoided this approach. In simple words: yes, you can completely rethink how agentic products are built and deployed. And this is wild.
Sources: paper
Raising ~$7.4B at up to a $59B valuation per Reuters, with founder Liang Wenfeng putting in ~40% himself. Meanwhile Vercel reports DeepSeek's share of AI Gateway tokens jumped from under 1% in April to 17% in May, while staying near 1% of spend. That chart is the entire "cheap intelligence" thesis in one image.
Sources: CNBC/Reuters, Vercel
Priced at $185, opened on Nasdaq at $385 and closed day one at $311 (+68%), raising $5.5B at a ~$56B valuation. Biggest US tech IPO since Snowflake, and the starting gun for the year's AI-listing wave.
Sources: Cerebras press release, TechCrunch, CNBC
Anthropic bought Stainless, Google DeepMind hired 20+ researchers from Contextual AI via a technology-licensing deal (talent and IP, no merger review) and Mistral acquired physics-simulation startup Emmi AI. Exactly the structure antitrust reviewers are starting to squint at, which is exactly why everyone uses it.
Sources: Bloomberg, Anthropic, Mistral
Figure's 200-hour shift: Figure 03 humanoids sorted packages on Helix-02 with zero teleoperation in a livestream planned as an 8-hour shift; nothing broke, so they kept going for ~200 hours and roughly 249,560 packages, robots rotating onto charging docks like shift workers. Days later Figure signed its first retail deployment with Catalyst Brands (JCPenney's parent). Sources: Seoul Economic Daily, Figure × Catalyst
With Alex Imas (Google DeepMind's Director of AGI Economics, a real title now) and Phil Trammell.
Priced at $135/share, raising $75 billion at a $1.77 trillion market cap, nearly triple Saudi Aramco's record, with SPCX trading on Nasdaq from June 12. Because Musk folded xAI into SpaceX in February, the listing takes Grok (and X) public too. Retail placed ~$100B in orders; Morningstar's public valuation is $780B, less than half the IPO mark. Price discovery is going to be sporty.

Sources: NPR, TechCrunch, S-1
Vera Rubin NVL72 in full production (racks that assemble in 5 minutes), RTX Spark (NVIDIA's 1-petaflop Windows PC superchip with MediaTek), Microsoft's Maia 200 live in production, Spectrum-X co-packaged optics switches in production, the first public AMD Helios MI455X racks, Intel pitching "agent density" with Xeon 6+ on 18A, and Huawei pulling the Ascend 950DT forward to August. Quotable Jensen at peak form: "Compute is revenues now. Compute is profit. The absence of revenues and profit is loss."

Sources: NVIDIA live blog, Build keynote, Supermicro, TrendForce (Huawei)
The Pope's first encyclical is about AI: Leo XIV's Magnifica Humanitas argues AI must serve humanity rather than concentrate power in a wealthy few, calls to "disarm AI" by removing it from military and economic interests, and demands stricter regulation. It promptly did 1,650 points on Hacker News, which is not a sentence anyone expected to write about an encyclical.
Dario's media week: Amodei told Bloomberg he has exactly one direct report ("incredibly freeing") and told ABC News he wants the government to have the power, "in a narrow way," to block deployment of unsafe AI, plus the bluntest line of the week: "I don't trust China at all." A CEO asking for the power to be stopped is either deeply reassuring or deeply alarming, and the debate over which was the point.
Hassabis on AI layoffs: companies blaming AI show "a lack of imagination… If engineers are becoming three or four times more productive, then we just [want to] do three or four times more stuff." He'd happily take the laid-off engineers; he has "a million ideas".
Sources: Vatican text, RNS, HN thread, TechCrunch, ABC News, policy post, Yahoo/Wired
Hackers took 20,225 Instagram accounts by asking nicely: attackers hijacked high-profile accounts (the Obama-era White House account, a Space Force chief, Sephora) by asking Meta's AI Support Assistant to add a new email and reset the password; a bug in a side code path skipped verifying the requester. The canonical agentic-AI-in-production failure: the chatbot had account-recovery powers and infinite patience.
Meta's bad privacy month, continued: WIRED found a dormant facial-recognition system ("NameTag") in the Meta AI app that pairs with its smart glasses (stripped within 48 hours of the report), and Reuters revealed the Model Capability Initiative was recording employee emails, chats and clipboards across 200+ apps to train agentic AI. After a 1,500-signature internal petition, employees can now pause collection… for 30 minutes at a time.
OWASP: prompt injection is "the universal joint": the 2026 State of Agentic AI Security report moved from theory to a catalog of real CVEs (the LiteLLM PyPI backdoor, a Cursor allowlist bypass, a Codex CLI sandbox flaw), with prompt injection mapping to 6 of its Top 10 agentic risks. Also this month: BadHost (CVE-2026-48710), a Starlette Host-header authorization bypass affecting vLLM, LiteLLM, FastAPI, Open WebUI and countless MCP servers. Patch, then ponder how much of the agentic stack rests on a handful of under-maintained packages.
Grok's legal pile-up: Labour MP Jess Asato filed the first UK claim against xAI over non-consensual sexualized deepfakes, and Canada's Privacy Commissioner found X/xAI violated federal privacy law (Grok's image tool at one point produced over 6,000 sexualized images per hour). This stacks on the EU's DSA proceedings and an Ofcom investigation.
Sources: 404 Media, TechCrunch, BleepingComputer, EFF, Engadget, TechSpot (MCI), OWASP report, Help Net Security, X41 advisory, Ars Technica, AWO, Privacy Commissioner, CBC
A rider's TikTok (~2M views): her Waymo paused mid-trip to ask, through the car speaker, "Are you over the age of 18?" In-cabin ML flags suspected minors, then a human agent patches in. a16z's Seema Amble: "Is this the new version of getting carded? Should I be flattered?" Privacy folks noted the cameras are "one court order away" from other uses. Sources: Jalopnik, Motor1
Spencer Pratt's mayoral run was powered by supporter-made AI videos casting him as a Batman-style hero saving dystopian LA (5M+ views; Jeb Bush called one "maybe the best political ad of the year"). He finished third in the June 2 primary with 25.8%, behind Karen Bass and Nithya Raman. "A defeat of AI slop," per Washington Monthly. The technology is new; losing the LA mayoral race to the incumbent is traditional. Sources: Washington Monthly, ABC7
Get the latest AI insights delivered to your inbox. No spam, unsubscribe anytime.
OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.
Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.
The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.