Skip to main content
AI Socratic
← News
O

Author

OpenRouter

By and about OpenRouter

news

How rogue inference providers can game router scoring by under-reporting cache hits

Tarun Chitra flags that OpenRouter's ranking system can be gamed by providers under-reporting cache hit rates to appear cheaper, creating a race to the bottom that penalizes honest providers. A subsequent real test on vLLM showed a config change…

news

Z.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips

Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.

news

Z.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips

Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.

news

Serving Kimi K3 on rented B200s breaks even at 159 tokens per GPU-second

VSC Ventures partner Jay Kapoor pressure-tests Dylan Patel's claim that anyone can profit renting B200s, and the math shows CoreWeave's $69/hour 8x B200 node needs 159 tokens per GPU-second sold at Kimi K3's $15/1M rate just to break even.

news

Dylan Patel: $11T of AI capex through 2029, $5T of it borrowed

Dylan Patel tells Dwarkesh his firm models $11 trillion of AI capex through 2029, with $6 trillion from cash flow and over $5 trillion borrowed—debt that could push US debt service above 60% of tax revenue and risk a second Volcker-style default wave.

news

OpenRouter's GPT-5.6 discount: Luna tokens up 13.8x, Terra 5.6x

OpenRouter's data shows a 50% discount on OpenAI's GPT-5.6 Terra and Luna drove daily token volume up 5.6x and 13.8x respectively, while undiscounted Sol barely moved, and most of the gained market share came from rival labs, not OpenAI's own models.

news

An 11x price spread for the same open-weight model on OpenRouter

Architect CEO Brett Harrison found an 11x price gap between OpenRouter's cheapest and priciest host of DeepSeek V4 Flash — Baidu ran it at $0.049 per million tokens and 124 tokens/second while 26 of 30 rivals were both pricier and slower.

news

Darkbloom's idle-Mac inference network doubles to 499 nodes in 60 hours

Darkbloom, Eigen Labs' idle-Mac inference network, grew to 499 nodes (432 hardware-attested) in 60 hours, up from 389 a day earlier, with utilization at just 8%.

news

Speko launches as an OpenRouter for voice AI

Speko, a YC S26 company, launched on Hacker News on August 17 pitching itself as "OpenRouter for Voice AI" — a single routing layer in front of voice providers — and drew 117 points and 67 comments.

news

llm-openrouter 0.7 surfaces reasoning traces

Simon Willison's llm-openrouter plugin hits 0.7, adding compatibility with LLM 0.32 and the ability to display reasoning traces from models served through OpenRouter.

news

Stripe to buy OpenRouter for $7B

Stripe nears deal to acquire OpenRouter, the multi-model LLM routing layer, for over $7 billion, positioning itself at the metering point between apps and frontier model providers.

blog

AI Socratic July 2026 — Lost In J-Space

Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.

blog

AI Socratic June 2026 - Hoist by Its Own Fable

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.

news

OpenRouter Fusion

OpenRouter's Fusion routes prompts to multiple models in parallel with web search and bash tools, then synthesizes their responses into a single answer, claiming Fable-level performance at half the cost.

blog

AI Socratic March 2026

Top AI updates from Jan 15 to Feb 15 2026

news

StepFun's Step 3.5 Flash

StepFun released Step 3.5 Flash, a sparse mixture-of-experts model with 196B total parameters but only 11B activated per token, designed to run on 128GB of memory and trained with the Muon optimizer.

news

DeepSeek V3.1 Terminus

DeepSeek released V3.1 Terminus on September 22, 2025, a 671B parameter hybrid reasoning model that fixes language consistency and agent capability issues in V3.1 while maintaining performance comparable to R1 on difficult benchmarks.

blog

AI Socratic July-Sep 2025 Part 1 — The Genie3 Is Out of The Box 🍌

This time around we’ll have 2 events, one in New York, and one for the first time in San Francisco at the Frontier Tower. We’ll discuss the top news and updates from this blog post using the Socratic

news

Benchmarks & Metrics: Models Increasingly Overfitted to Leaderboards

Most AI models are heavily overfitted to popular benchmarks, making leaderboard scores a poor proxy for real-world performance and more of a tracker for release velocity.

blog

AI Socratic May 2025

The most important AI news and updates from last month (April 15 - May 15). A beefy month!

news

LLM Models Vibe Check & Benchmarks: OpenRouter, lmarena, and IQ

Gemini 2.5 climbs OpenRouter rankings while Claude 3.7 declines, as companies increasingly optimize models specifically for benchmark performance rather than general capability.

blog

AI Socratic Apr 2025

All the AI updates from mar 15 to apr 20. Including GPT o3, o4-mini, 4.1 to Gemini 2.5, the controversial AI-2027 blog post, A2A and more.

news

OpenAI drops GPT-4.1 with Mini and Nano variants

OpenAI released GPT-4.1 with Mini and Nano variants, offering a 1M token context window, 55% score on SWE-Bench Verified coding tasks, and pricing from $0.40/$1.60 per 1M tokens for the smaller models.