Company / organization
NVIDIA
By and about NVIDIA
Magic says it matched DeepSeek V4 Pro Base with 50x less compute
Magic says its pretraining recipe matched DeepSeek V4 Pro Base using roughly 50x fewer FLOPs, about $0.5M on GB200, then scaled 10x further for roughly $4M to beat every public open base model on perplexity evals.
newsHuang puts Astra's training run at 100,000 GPUs, with 400,000 next
Jensen Huang put the first hardware number on GPT-6 Astra's training run—"100K+" Nvidia Grace Blackwell GPUs, with 400,000 more coming online next—confirmed by OpenAI president Greg Brockman as GPUs, not racks, as Huang declared "AGI has arrived."
newsCritical CVE disclosures at 21 major vendors jump from 84 a month to 606
Twenty-one major vendors disclosed 606 critical CVEs in July 2026, up from a prior record of 84 a month, an inflection Epoch AI ties to Anthropic's Claude Mythos Preview and Project Glasswing's vulnerability hunting at Microsoft, Google, Apple and AWS.
newsOpenAI ships GPT-6 Astra and declares the AGI era
GPT-6 Astra sets an Epoch Capabilities Index record at 169, scores 63–66% on ARC-AGI-3 with ARC Prize's standard harness and 99% with a provider adapter, and becomes OpenAI's first Critical-rated cyber model.
news“You don’t run a model, you run kernels”: Ahmad Osman on the layer nobody benchmarks
Osmantic founder Ahmad Osman argues inference speed lives in the kernel layer, not the model — a Strix Halo test gives a same-model ROCm rebuild 5.7-6.8% more throughput — and published an eight-project path for learning it.
newsSutskever: rogue agents will go for the neoclouds next
Ilya Sutskever warned rogue AI agents will target neoclouds next, and a SemiAnalysis audit of 25 providers backs him up: a default InfiniBand key exposed 532 hostnames, and a Grafana leak exposed every tenant's logs.
newsWorld Labs' Atlas generates a minute of 1440p video under exact camera control
Fei-Fei Li's World Labs launched Atlas, a world model that generates up to a minute of 1440p video along a precise camera path from one reference image, with raters preferring its camera control 75-94% of the time over rivals like Seedance 2.5.
newsThe week's top AI papers say the harness, not the model, is the variable
DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.
newsAgent-generated kernels cut Qwen-Image serving latency 42.3%
Baseten engineer Brian Li reports an agentic kernel-development framework cut serving latency 42.3% on Qwen-Image and 15.2% on FLUX.2, atop an already human-tuned SGLang stack on NVIDIA B300 GPUs.
newsNvidia agrees to acquire Hugging Face for $13B
Nvidia has reportedly agreed to acquire Hugging Face, the open-source model hub, for roughly $12.9 billion — a move read as both chip-moat defense and a return to the cloud business.
newsZ.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips
Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.
newsZ.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips
Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.
newsServing Kimi K3 on rented B200s breaks even at 159 tokens per GPU-second
VSC Ventures partner Jay Kapoor pressure-tests Dylan Patel's claim that anyone can profit renting B200s, and the math shows CoreWeave's $69/hour 8x B200 node needs 159 tokens per GPU-second sold at Kimi K3's $15/1M rate just to break even.
newsDwarkesh Patel: the AI buildout could set off a second Volcker shock
Dwarkesh Patel argues the AI buildout, not a central bank, could cause a "second Volcker shock" of sovereign defaults, citing SemiAnalysis's $11 trillion capex forecast and Google's $920M-a-month SpaceX compute deal.
newsXiaomi's AI Cube prototype has 80GB of memory, not the 160GB everyone quoted
Xiaomi's AI Cube prototype runs a 120B model on 80GB of unified memory, not the 160GB widely quoted online; that figure is the Xring D100 chip's maximum, and the cited 1.22 TB/s is the O100's near-memory bandwidth, not the machine's unified-memory speed.
newsNvidia: the harness, not the model, is the hero
Nvidia research finds that agents can perform well and stay on-rails through fine-tuning even when the underlying model isn't especially good at the task — the scaffolding, not the checkpoint, does the work.
newsNVIDIA finds skill doc-scans predict nothing about what a skill does at runtime
NVIDIA researchers' ACES framework ran 947 paired agent trials across 58 production skills and found that both structural and LLM-judge doc-scan scores correlate with real runtime skill lift at essentially zero (Spearman ρ = -0.018 and -0.027).
newsGPT-5.6 Sol Ultrafast: OpenAI Ships a 14x Mode on Cerebras
OpenAI is previewing Ultrafast, a mode that runs GPT-5.6 Sol at 14x speed on Cerebras hardware, targeting enterprise users with lower latency rather than new capabilities.
blogAI Socratic August 2026 — Escaping The Sandbox
OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.
newsAMD buys Taalas to etch models into silicon
AMD is acquiring AI chip startup Taalas to etch machine learning models directly into silicon for faster inference, trading flexibility for efficiency in a bet that model architectures will stabilize enough to freeze in hardware.
blogMarket Analysis: Open Weights vs Proprietary Models
Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.
blogAI Socratic July 2026 — Lost In J-Space
Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.
newsUnsloth founder's 2h42m fine-tuning masterclass
Unsloth founder, ex-NVIDIA engineer, covers reinforcement learning, kernels, reasoning, quantization, and agents in a 2h42m masterclass.
blogAI Socratic June 2026 - Hoist by Its Own Fable
Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.
blogAI Socratic May 2026 — The Selfish Gen AI
DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes
blogAI Socratic April 2026 — The Era of Mythos
Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR
blogAI Socratic March 2026 — #2
NVIDIA GTC, Anthropic win all, TurboQuant and more
blogAI Socratic March 2026
Top AI updates from Jan 15 to Feb 15 2026
blogAI Socratic February 2026
Top AI updates from Jan 15 to Feb 15 2026
blogAI Socratic Jan 2026
Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more
blogAI Socratic Dec 2025
The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.
blogAI Socratic Oct 2025
The most important AI news and updates from last month: Sep 15 – Oct 15.
blogDeepSeek R1 Shakes The AI Industry
The biggest event in January has been the launch of DeepSeek R1, which shook the market pushing NVIDIA stock down by 20% in a few days.