Author
OpenRouter
By and about OpenRouter
How rogue inference providers can game router scoring by under-reporting cache hits
Tarun Chitra flags that OpenRouter's ranking system can be gamed by providers under-reporting cache hit rates to appear cheaper, creating a race to the bottom that penalizes honest providers. A subsequent real test on vLLM showed a config change…
newsZ.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips
Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.
newsZ.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips
Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.
newsServing Kimi K3 on rented B200s breaks even at 159 tokens per GPU-second
VSC Ventures partner Jay Kapoor pressure-tests Dylan Patel's claim that anyone can profit renting B200s, and the math shows CoreWeave's $69/hour 8x B200 node needs 159 tokens per GPU-second sold at Kimi K3's $15/1M rate just to break even.
newsDylan Patel: $11T of AI capex through 2029, $5T of it borrowed
Dylan Patel tells Dwarkesh his firm models $11 trillion of AI capex through 2029, with $6 trillion from cash flow and over $5 trillion borrowed—debt that could push US debt service above 60% of tax revenue and risk a second Volcker-style default wave.
newsOpenRouter's GPT-5.6 discount: Luna tokens up 13.8x, Terra 5.6x
OpenRouter's data shows a 50% discount on OpenAI's GPT-5.6 Terra and Luna drove daily token volume up 5.6x and 13.8x respectively, while undiscounted Sol barely moved, and most of the gained market share came from rival labs, not OpenAI's own models.
newsAn 11x price spread for the same open-weight model on OpenRouter
Architect CEO Brett Harrison found an 11x price gap between OpenRouter's cheapest and priciest host of DeepSeek V4 Flash — Baidu ran it at $0.049 per million tokens and 124 tokens/second while 26 of 30 rivals were both pricier and slower.
newsDarkbloom's idle-Mac inference network doubles to 499 nodes in 60 hours
Darkbloom, Eigen Labs' idle-Mac inference network, grew to 499 nodes (432 hardware-attested) in 60 hours, up from 389 a day earlier, with utilization at just 8%.
newsSpeko launches as an OpenRouter for voice AI
Speko, a YC S26 company, launched on Hacker News on August 17 pitching itself as "OpenRouter for Voice AI" — a single routing layer in front of voice providers — and drew 117 points and 67 comments.
newsllm-openrouter 0.7 surfaces reasoning traces
Simon Willison's llm-openrouter plugin hits 0.7, adding compatibility with LLM 0.32 and the ability to display reasoning traces from models served through OpenRouter.
newsStripe to buy OpenRouter for $7B
Stripe nears deal to acquire OpenRouter, the multi-model LLM routing layer, for over $7 billion, positioning itself at the metering point between apps and frontier model providers.
blogAI Socratic July 2026 — Lost In J-Space
Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.
blogAI Socratic June 2026 - Hoist by Its Own Fable
Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.
newsOpenRouter Fusion
OpenRouter's Fusion routes prompts to multiple models in parallel with web search and bash tools, then synthesizes their responses into a single answer, claiming Fable-level performance at half the cost.
blogAI Socratic March 2026
Top AI updates from Jan 15 to Feb 15 2026
newsStepFun's Step 3.5 Flash
StepFun released Step 3.5 Flash, a sparse mixture-of-experts model with 196B total parameters but only 11B activated per token, designed to run on 128GB of memory and trained with the Muon optimizer.
newsDeepSeek V3.1 Terminus
DeepSeek released V3.1 Terminus on September 22, 2025, a 671B parameter hybrid reasoning model that fixes language consistency and agent capability issues in V3.1 while maintaining performance comparable to R1 on difficult benchmarks.
blogAI Socratic July-Sep 2025 Part 1 — The Genie3 Is Out of The Box 🍌
This time around we’ll have 2 events, one in New York, and one for the first time in San Francisco at the Frontier Tower. We’ll discuss the top news and updates from this blog post using the Socratic
newsBenchmarks & Metrics: Models Increasingly Overfitted to Leaderboards
Most AI models are heavily overfitted to popular benchmarks, making leaderboard scores a poor proxy for real-world performance and more of a tracker for release velocity.
blogAI Socratic May 2025
The most important AI news and updates from last month (April 15 - May 15). A beefy month!
newsLLM Models Vibe Check & Benchmarks: OpenRouter, lmarena, and IQ
Gemini 2.5 climbs OpenRouter rankings while Claude 3.7 declines, as companies increasingly optimize models specifically for benchmark performance rather than general capability.
blogAI Socratic Apr 2025
All the AI updates from mar 15 to apr 20. Including GPT o3, o4-mini, 4.1 to Gemini 2.5, the controversial AI-2027 blog post, A2A and more.
newsOpenAI drops GPT-4.1 with Mini and Nano variants
OpenAI released GPT-4.1 with Mini and Nano variants, offering a 1M token context window, 55% score on SWE-Bench Verified coding tasks, and pricing from $0.40/$1.60 per 1M tokens for the smaller models.