Company / organization
DeepSeek
By and about DeepSeek
DeepSeek V4 Flash
DeepSeek released V4 Flash, an efficiency-optimized Mixture-of-Experts model with 284B total parameters and 13B activated, supporting a 1M-token context window.
newsChinese labs accused of distilling Claude models
Anthropic says DeepSeek, Moonshot AI, and MiniMax used over 24,000 fraudulent accounts to generate 16 million Claude exchanges for model distillation.
newsReasoning models don't always say what they think
Anthropic study finds that reasoning models like Claude and DeepSeek R1 hide their actual reasoning process, admitting to using external hints only 25–39% of the time and instead generating fake logical justifications.
newsDo LLMs Benefit From their own Words?
MIT researchers found that LLMs degrade in long conversations by treating their own previous responses as fact, causing errors to compound, and that removing prior AI outputs often improves quality while cutting context length by up to 10x.
newsPrompt Repetition Improves Non-Reasoning LLMs
Repeating a prompt twice improves accuracy across Gemini, ChatGPT, Claude, and DeepSeek on seven benchmarks without fine-tuning, extra training, or added output length.
newsDeepSeek mHC: Manifold-Constrained Hyper-Connections
DeepSeek's mHC constrains residual connection mixing matrices to the Birkhoff polytope using Sinkhorn-Knopp normalization, enabling stable 4-8x wider residual streams and improved performance on reasoning benchmarks up to 27B parameters.
newsUS AI startups increasingly built on Chinese open-source foundations
Chinese open-source models like DeepSeek and Qwen now lead US models in global downloads with 17% market share versus 15.8%, attracting US startups with superior price-to-performance but raising concerns about embedded censorship and regulatory risk.
newsSimWorld: an open-ended simulator for agents in physical and social worlds
Researchers created SimWorld, a simulator where AI models compete in a market economy with tasks like food delivery; Claude and Qwen used riskier strategies with higher returns while other models played conservatively, and Qwen and DeepSeek undercut…
newsDeepSeek V3.2 Speciale
DeepSeek released V3.2 Speciale on November 28, 2025, a high-compute variant optimized for reasoning and agentic tasks with open weights.
newsDeepSeek-OCR: Context Compression Through Optical 2D Mapping
DeepSeek AI released DeepSeek-OCR, a vision-based compression system that converts text to image tokens, achieving 97% accuracy at 10× compression and 60% at 20× compression to reduce context window demands.
newsDeepSeek V3.2 Exp
DeepSeek released V3.2 Exp on September 29, 2025 as an experimental model between V3.1 and future versions, with weights available openly.
newsDeepSeek V3.1 Terminus
DeepSeek released V3.1 Terminus on September 22, 2025, a 671B parameter hybrid reasoning model that fixes language consistency and agent capability issues in V3.1 while maintaining performance comparable to R1 on difficult benchmarks.
newsDeepSeek V3.1
DeepSeek released V3.1, a 671-billion-parameter hybrid reasoning model with 37 billion active parameters that supports both thinking and non-thinking modes.
newsDeepSeek V3.1 Base
DeepSeek released DeepSeek V3.1 Base on August 19, 2025, a 671B parameter open-weights mixture-of-experts model with 37B active parameters, 128K context length, and training on 14.8T tokens.
newsKimi 2 👑 — New open source 1T LLMs
Moonshot released Kimi 2, an open source 1T parameter model using a DeepSeek V3-inspired architecture that achieves state-of-the-art performance on several benchmarks while training faster and cheaper than competitors, powered by a custom Muon optimizer…
newsApple's "The Illusion of Thinking" and the Rebuttal
Apple's research paper argues reasoning models like o1 and o3 aren't truly thinking but failing due to context limits, though a rebuttal shows o3 solves Tower of Hanoi when prompted to be concise.
newsDeepSeek Drops Best End-to-End Large Model Training Paper
DeepSeek released a comprehensive paper on end-to-end large language model training methodology.
newsBaidu ERNIE-4.5 — DeepSeek R1 level at half the price 👸
Baidu released ERNIE-4.5, matching DeepSeek R1's performance at half the cost while claiming benchmark wins over GPT-4o in multimodal tasks and GPT-4.5 in some text benchmarks.
newsManus: DeepResearch + Operator agent from China
Chinese startup Butterfly Effect released Manus, an autonomous AI agent powered by Claude 3.5 and Qwen that performs real-world tasks like building websites and analyzing stocks, outperforming OpenAI's Deep Research on the GAIA benchmark.
newsAI Dinner 7.0 — DeepSeek Workshop
AI Socratic hosts AI Dinner 7.0 on February 20th in Manhattan to discuss DeepSeek R1 research papers and install the model on attendees' laptops.