Skip to main content
AI Socratic
← News
NVIDIA

Company / organization

NVIDIA

Website / profile ↗

By and about NVIDIA

news

Magic says it matched DeepSeek V4 Pro Base with 50x less compute

Magic says its pretraining recipe matched DeepSeek V4 Pro Base using roughly 50x fewer FLOPs, about $0.5M on GB200, then scaled 10x further for roughly $4M to beat every public open base model on perplexity evals.

news

Huang puts Astra's training run at 100,000 GPUs, with 400,000 next

Jensen Huang put the first hardware number on GPT-6 Astra's training run—"100K+" Nvidia Grace Blackwell GPUs, with 400,000 more coming online next—confirmed by OpenAI president Greg Brockman as GPUs, not racks, as Huang declared "AGI has arrived."

news

Critical CVE disclosures at 21 major vendors jump from 84 a month to 606

Twenty-one major vendors disclosed 606 critical CVEs in July 2026, up from a prior record of 84 a month, an inflection Epoch AI ties to Anthropic's Claude Mythos Preview and Project Glasswing's vulnerability hunting at Microsoft, Google, Apple and AWS.

news

OpenAI ships GPT-6 Astra and declares the AGI era

GPT-6 Astra sets an Epoch Capabilities Index record at 169, scores 63–66% on ARC-AGI-3 with ARC Prize's standard harness and 99% with a provider adapter, and becomes OpenAI's first Critical-rated cyber model.

news

“You don’t run a model, you run kernels”: Ahmad Osman on the layer nobody benchmarks

Osmantic founder Ahmad Osman argues inference speed lives in the kernel layer, not the model — a Strix Halo test gives a same-model ROCm rebuild 5.7-6.8% more throughput — and published an eight-project path for learning it.

news

Sutskever: rogue agents will go for the neoclouds next

Ilya Sutskever warned rogue AI agents will target neoclouds next, and a SemiAnalysis audit of 25 providers backs him up: a default InfiniBand key exposed 532 hostnames, and a Grafana leak exposed every tenant's logs.

news

World Labs' Atlas generates a minute of 1440p video under exact camera control

Fei-Fei Li's World Labs launched Atlas, a world model that generates up to a minute of 1440p video along a precise camera path from one reference image, with raters preferring its camera control 75-94% of the time over rivals like Seedance 2.5.

news

The week's top AI papers say the harness, not the model, is the variable

DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.

news

Agent-generated kernels cut Qwen-Image serving latency 42.3%

Baseten engineer Brian Li reports an agentic kernel-development framework cut serving latency 42.3% on Qwen-Image and 15.2% on FLUX.2, atop an already human-tuned SGLang stack on NVIDIA B300 GPUs.

news

Nvidia agrees to acquire Hugging Face for $13B

Nvidia has reportedly agreed to acquire Hugging Face, the open-source model hub, for roughly $12.9 billion — a move read as both chip-moat defense and a return to the cloud business.

news

Z.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips

Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.

news

Z.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips

Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.

news

Serving Kimi K3 on rented B200s breaks even at 159 tokens per GPU-second

VSC Ventures partner Jay Kapoor pressure-tests Dylan Patel's claim that anyone can profit renting B200s, and the math shows CoreWeave's $69/hour 8x B200 node needs 159 tokens per GPU-second sold at Kimi K3's $15/1M rate just to break even.

news

Dwarkesh Patel: the AI buildout could set off a second Volcker shock

Dwarkesh Patel argues the AI buildout, not a central bank, could cause a "second Volcker shock" of sovereign defaults, citing SemiAnalysis's $11 trillion capex forecast and Google's $920M-a-month SpaceX compute deal.

news

Xiaomi's AI Cube prototype has 80GB of memory, not the 160GB everyone quoted

Xiaomi's AI Cube prototype runs a 120B model on 80GB of unified memory, not the 160GB widely quoted online; that figure is the Xring D100 chip's maximum, and the cited 1.22 TB/s is the O100's near-memory bandwidth, not the machine's unified-memory speed.

news

Nvidia: the harness, not the model, is the hero

Nvidia research finds that agents can perform well and stay on-rails through fine-tuning even when the underlying model isn't especially good at the task — the scaffolding, not the checkpoint, does the work.

news

NVIDIA finds skill doc-scans predict nothing about what a skill does at runtime

NVIDIA researchers' ACES framework ran 947 paired agent trials across 58 production skills and found that both structural and LLM-judge doc-scan scores correlate with real runtime skill lift at essentially zero (Spearman ρ = -0.018 and -0.027).

news

GPT-5.6 Sol Ultrafast: OpenAI Ships a 14x Mode on Cerebras

OpenAI is previewing Ultrafast, a mode that runs GPT-5.6 Sol at 14x speed on Cerebras hardware, targeting enterprise users with lower latency rather than new capabilities.

blog

AI Socratic August 2026 — Escaping The Sandbox

OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.

news

AMD buys Taalas to etch models into silicon

AMD is acquiring AI chip startup Taalas to etch machine learning models directly into silicon for faster inference, trading flexibility for efficiency in a bet that model architectures will stabilize enough to freeze in hardware.

blog

Market Analysis: Open Weights vs Proprietary Models

Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.

blog

AI Socratic July 2026 — Lost In J-Space

Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.

news

Unsloth founder's 2h42m fine-tuning masterclass

Unsloth founder, ex-NVIDIA engineer, covers reinforcement learning, kernels, reasoning, quantization, and agents in a 2h42m masterclass.

blog

AI Socratic June 2026 - Hoist by Its Own Fable

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.

blog

AI Socratic May 2026 — The Selfish Gen AI

DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes

blog

AI Socratic April 2026 — The Era of Mythos

Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR

blog

AI Socratic March 2026 — #2

NVIDIA GTC, Anthropic win all, TurboQuant and more

blog

AI Socratic March 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic February 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic Jan 2026

Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more

blog

AI Socratic Dec 2025

The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.

blog

AI Socratic Oct 2025

The most important AI news and updates from last month: Sep 15 – Oct 15.

blog

DeepSeek R1 Shakes The AI Industry

The biggest event in January has been the launch of DeepSeek R1, which shook the market pushing NVIDIA stock down by 20% in a few days.