We use cookies to improve your experience and analyze site traffic. You can choose which cookies to allow. Privacy Policy
Moonshot AI released Kimi-K3, the successor to its K2 open-weight model, on Hugging Face July 27, reaching 780 points on Hacker News within hours.
Anthropic released Claude Opus 5 on July 24, claiming near-Fable 5 capabilities at half the price and now ranking #1 on the Artificial Analysis Intelligence Leaderboard, though reviewers see the win as token efficiency and lower restrictions rather than…
Google DeepMind released Gemma 4 31B, a free open-weights multimodal model with 256K context, configurable reasoning mode, native function calling, and support for 140+ languages under Apache 2.0 license.
Thinking Machines Lab released Inkling, a 975B-parameter open-weights Mixture-of-Experts model with 41B active parameters, native multimodal support (text, image, audio, video), and a 1M-token context window designed for agentic coding with controllable…
Anthropic released Claude Sonnet 5 on June 30 with 63.2% SWE-bench Pro and a GDPval-AA score of 1,618—nearly matching Opus 4.8's 1,615—at $2/$10 per million tokens through August 31, with native 1M context and 128K output by default.
Z.ai released GLM-5.2, a 744B open-weight model scoring 62.1% on SWE-bench Pro and rivaling closed models like Claude Opus on coding tasks, with 1M token context and roughly one-sixth the cost of GPT-5.5.
OpenAI previewed GPT-5.6 (Sol, Terra, Luna) on June 26, but initial access is limited to ~20 government-approved organizations for cybersecurity review under the June 2 executive order framework, with general availability promised in coming weeks.
Moonshot AI released Kimi K3, a 2.8T parameter open-weight multimodal reasoning model.
Meta released Muse Spark 1.1 and launched its Meta Model API public preview, claiming the model rivals GPT-5.5 and Opus 4.8 on agentic evals while undercutting on price—Meta's first serious push to monetize its $115–135B infrastructure spend by selling…
SpaceXAI released Grok 4.5, achieving 29.0% on SWE Marathon with 4.2× fewer output tokens than Opus 4.8, priced at $2/M input and $6/M output tokens.
Moonshot AI released Kimi K2.7 Code, a coding-focused model designed to handle end-to-end programming tasks over long contexts.
OpenAI previews GPT-5.6 Sol (flagship), Terra (2x cheaper), and Luna (lowest cost) in limited preview with trusted partners via API, with broader rollout planned after government coordination.
MiniMax released M2.7 on June 16, 2026, an open-weights model designed for autonomous task execution with multi-agent collaboration, scoring 56.2% on SWE-Pro and 1495 ELO on GDPval-AA.
Microsoft launched seven in-house MAI models including MAI-Thinking-1, a 35B-active mixture-of-experts trained from scratch, while StepFun released Step 3.7 Flash, a 198B open vision-language model, and xAI debuted Grok Imagine Video 1.5 for…
Google released Gemma 4 (12B to 31B, Apache 2.0, natively multimodal, 256K context) for local deployment, and followed with DiffusionGemma, a 26B diffusion model generating 1,000+ tokens/sec on a single H100 by denoising in parallel instead of…
Anthropic released Claude Fable 5, the first publicly available Mythos-class model, alongside Claude Mythos 5 for vetted cyberdefenders; Fable 5 scores 80.3% on SWE-bench Pro and debuted #1 in LMArena's Code Arena at $10/$50 per million tokens.
NVIDIA releases Nemotron 3 Ultra, a 550B-parameter open model with 55B active parameters and 1M context window, under the OpenMDW 1.1 license that includes weights, synthetic data, and training recipes.
OpenAI's Jakub Pachocki told staff GPT-5.6 will arrive within weeks as a meaningful improvement over GPT-5.5, possibly with a ChatGPT redesign replacing the model picker with six Intelligence Levels.
On June 14th-15th, internet users spread false claims that a new model called Le Chaton Fat vastly outperformed Fable 5, with many believing the hoax before it was debunked.
Claude Opus 4.8 improved SWE-bench Pro to 69.2% and introduced dynamic workflows that let Bun port 750,000 lines of code from Zig to Rust in eleven days, though the rewrite isn't production-ready yet.
MiniMax released M3, an open-weight multimodal model with 1M-token context and a sparse attention architecture that reduces per-token compute by ~95% at full context; it scores 59.0% on SWE-Bench Pro, outperforming GPT-5.5 in the company's own testing.
Alibaba unveiled Qwen3.7-Max with 1M-token context, claiming benchmark wins over Claude Opus 4.6, while multimodal Qwen3.7-Plus launched at $0.40/$1.60 per million tokens alongside custom Zhenwu M890 chips.
Google launched Gemini 3.5 Flash as the default model in Gemini app and Search AI Mode, scoring 76.2% on Terminal-Bench 2.1 and 1656 Elo on GDPval-AA while running ~4x faster than comparable frontier models.
Z.ai released GLM 5.2, an open-weights reasoning model with a 1M-token context window designed for long-horizon agent workflows, software engineering, and multi-step automation tasks.
DeepSeek released V4 preview with two open-weight MoE models—V4-Pro (1.6T params, 49B active) and V4-Flash (284B total, 13B active)—featuring hybrid attention for practical 1M-token context and competitive API pricing.
DeepSeek released V4 Flash, an efficiency-optimized Mixture-of-Experts model with 284B total parameters and 13B activated, supporting a 1M-token context window.
DeepSeek released V4 Pro on April 22, 2026, a Mixture-of-Experts model with 1.6T total parameters, 49B activated parameters, and a 1M-token context window.
Anthropic released Opus 4.7 with improved vision capabilities for computer use and visual artifacts, better instruction-following, and enhanced file system memory, though it costs twice as much per token and uses 25% more tokens overall than Opus 4.6.
OpenAI released ChatGPT Image 2, which demonstrates improved ability to combine multiple subjects coherently while maintaining image quality, including successfully rendering complex prompts like Studio Ghibli-style scenes with detailed elements.
Z.ai released GLM-5.1 on April 3, 2026, an open-weights model with major gains in coding capability and long-horizon task handling.
Google released Gemini Embedding 2, a multimodal embedding model that maps text, images, video, audio and documents into a single embedding space for cross-media retrieval and classification.
Anthropic released Claude Opus 4.6 with 1M token context, multi-agent coordination, and optional 6x faster thinking mode (2.5x costlier), plus Claude Sonnet 4.6 as the new default across claude.ai with improved reasoning and coding.
Alibaba's Qwen released Qwen3 Coder Next, an open-weight language model designed for coding agents and local development workflows.
Zhipu AI released GLM-4.7-Flash, a 30B open-weights model positioned as a performance-efficiency balance in the SOTA class.
Moonshot AI released Kimi K2.6, an open-weights multimodal model designed for long-horizon coding, UI/UX generation, and multi-agent orchestration.
Sakana AI introduces Continuous Thought Machines, an AI model that synchronizes neuron activity timing to enable step-by-step reasoning similar to biological neural networks.
NVIDIA released Nemotron 3 Nano 30B A3B, a free open-weights mixture-of-experts model optimized for efficient inference and agentic AI systems.
DeepSeek released V3.2 Speciale on November 28, 2025, a high-compute variant optimized for reasoning and agentic tasks with open weights.
Moonshot AI released Kimi K2 Thinking, an open-weights reasoning model designed for agentic and long-horizon tasks.
DeepSeek released V3.2 Exp on September 29, 2025 as an experimental model between V3.1 and future versions, with weights available openly.
Z.ai released GLM-4.6 on September 29, 2025, expanding the context window from 128K to 200K tokens and improving coding performance, reasoning, and tool use capabilities.
DeepSeek released V3.1 Terminus on September 22, 2025, a 671B parameter hybrid reasoning model that fixes language consistency and agent capability issues in V3.1 while maintaining performance comparable to R1 on difficult benchmarks.
OpenAI released gpt-oss-safeguard-20b, a 20 billion parameter safety reasoning model, on September 18, 2025 as open weights.
Moonshot AI released Kimi K2 0905 on September 3, 2025, a 1 trillion parameter mixture-of-experts model with 32 billion active parameters per forward pass, extending context length to 256k tokens and improving agentic coding and frontend generation…
DeepSeek released V3.1, a 671-billion-parameter hybrid reasoning model with 37 billion active parameters that supports both thinking and non-thinking modes.
DeepSeek released DeepSeek V3.1 Base on August 19, 2025, a 671B parameter open-weights mixture-of-experts model with 37B active parameters, 128K context length, and training on 14.8T tokens.
OpenAI released gpt-oss-20b, a 21 billion parameter open-weight model under the Apache 2.0 license.
Alibaba released Qwen3 Coder 30B A3B Instruct on July 31, 2025, a 30.5B parameter mixture-of-experts model with 256K token context designed for code generation, repository understanding, and tool use.
The weekly AI digest — models, agents, open source, research — plus a monthly round-up. Unsubscribe anytime.