Skip to main content
AI Socratic
← News
D

Company / organization

DeepSeek

Website / profile ↗

By and about DeepSeek

news

DeepSeek V4 Flash

DeepSeek released V4 Flash, an efficiency-optimized Mixture-of-Experts model with 284B total parameters and 13B activated, supporting a 1M-token context window.

news

Chinese labs accused of distilling Claude models

Anthropic says DeepSeek, Moonshot AI, and MiniMax used over 24,000 fraudulent accounts to generate 16 million Claude exchanges for model distillation.

news

Reasoning models don't always say what they think

Anthropic study finds that reasoning models like Claude and DeepSeek R1 hide their actual reasoning process, admitting to using external hints only 25–39% of the time and instead generating fake logical justifications.

news

Do LLMs Benefit From their own Words?

MIT researchers found that LLMs degrade in long conversations by treating their own previous responses as fact, causing errors to compound, and that removing prior AI outputs often improves quality while cutting context length by up to 10x.

news

Prompt Repetition Improves Non-Reasoning LLMs

Repeating a prompt twice improves accuracy across Gemini, ChatGPT, Claude, and DeepSeek on seven benchmarks without fine-tuning, extra training, or added output length.

news

DeepSeek mHC: Manifold-Constrained Hyper-Connections

DeepSeek's mHC constrains residual connection mixing matrices to the Birkhoff polytope using Sinkhorn-Knopp normalization, enabling stable 4-8x wider residual streams and improved performance on reasoning benchmarks up to 27B parameters.

news

US AI startups increasingly built on Chinese open-source foundations

Chinese open-source models like DeepSeek and Qwen now lead US models in global downloads with 17% market share versus 15.8%, attracting US startups with superior price-to-performance but raising concerns about embedded censorship and regulatory risk.

news

SimWorld: an open-ended simulator for agents in physical and social worlds

Researchers created SimWorld, a simulator where AI models compete in a market economy with tasks like food delivery; Claude and Qwen used riskier strategies with higher returns while other models played conservatively, and Qwen and DeepSeek undercut…

news

DeepSeek V3.2 Speciale

DeepSeek released V3.2 Speciale on November 28, 2025, a high-compute variant optimized for reasoning and agentic tasks with open weights.

news

DeepSeek-OCR: Context Compression Through Optical 2D Mapping

DeepSeek AI released DeepSeek-OCR, a vision-based compression system that converts text to image tokens, achieving 97% accuracy at 10× compression and 60% at 20× compression to reduce context window demands.

news

DeepSeek V3.2 Exp

DeepSeek released V3.2 Exp on September 29, 2025 as an experimental model between V3.1 and future versions, with weights available openly.

news

DeepSeek V3.1 Terminus

DeepSeek released V3.1 Terminus on September 22, 2025, a 671B parameter hybrid reasoning model that fixes language consistency and agent capability issues in V3.1 while maintaining performance comparable to R1 on difficult benchmarks.

news

DeepSeek V3.1

DeepSeek released V3.1, a 671-billion-parameter hybrid reasoning model with 37 billion active parameters that supports both thinking and non-thinking modes.

news

DeepSeek V3.1 Base

DeepSeek released DeepSeek V3.1 Base on August 19, 2025, a 671B parameter open-weights mixture-of-experts model with 37B active parameters, 128K context length, and training on 14.8T tokens.

news

Kimi 2 👑 — New open source 1T LLMs

Moonshot released Kimi 2, an open source 1T parameter model using a DeepSeek V3-inspired architecture that achieves state-of-the-art performance on several benchmarks while training faster and cheaper than competitors, powered by a custom Muon optimizer…

news

Apple's "The Illusion of Thinking" and the Rebuttal

Apple's research paper argues reasoning models like o1 and o3 aren't truly thinking but failing due to context limits, though a rebuttal shows o3 solves Tower of Hanoi when prompted to be concise.

news

DeepSeek Drops Best End-to-End Large Model Training Paper

DeepSeek released a comprehensive paper on end-to-end large language model training methodology.

news

Baidu ERNIE-4.5 — DeepSeek R1 level at half the price 👸

Baidu released ERNIE-4.5, matching DeepSeek R1's performance at half the cost while claiming benchmark wins over GPT-4o in multimodal tasks and GPT-4.5 in some text benchmarks.

news

Manus: DeepResearch + Operator agent from China

Chinese startup Butterfly Effect released Manus, an autonomous AI agent powered by Claude 3.5 and Qwen that performs real-world tasks like building websites and analyzing stocks, outperforming OpenAI's Deep Research on the GAIA benchmark.

news

AI Dinner 7.0 — DeepSeek Workshop

AI Socratic hosts AI Dinner 7.0 on February 20th in Manhattan to discuss DeepSeek R1 research papers and install the model on attendees' laptops.