We use cookies to improve your experience and analyze site traffic. You can choose which cookies to allow. Privacy Policy
Google proposes Nested Learning, a hierarchical neural network architecture that updates parameters during inference to enable continuous learning without catastrophic forgetting.
OpenAI plans to spend $1.4 trillion on infrastructure while projecting just $100 billion in revenue by 2027.
Nano Banana Pro converts earnings PDFs into infographics and slides, automatically extracting key data and insights.
Scale AI and Meta released SAM 3, an open-source model for segmenting objects in images, videos, and 3D scenes.
Prime Intellect released Intellect-3, a 100B+ parameter MoE model trained via decentralized compute, achieving state-of-the-art performance on math and code benchmarks.
Google's NotebookLM now includes Deep Research, a feature that generates comprehensive reports by analyzing multiple sources.
Poetiq AI Agent reaches 50% on ARC-AGI-2 benchmark at $50 per task, half the cost of previous best results, suggesting agent scaffolding may matter more than raw model capability for reasoning tasks.
Disney invests $1 billion in OpenAI and secures a three-year license to use Sora for Disney, Marvel, Pixar, and Star Wars content.
OpenAI is quietly testing a next-generation image backend internally called "Image 2," positioned as a frontier-tier system comparable to Nano Banana Pro.
Researchers created SimWorld, a simulator where AI models compete in a market economy with tasks like food delivery; Claude and Qwen used riskier strategies with higher returns while other models played conservatively, and Qwen and DeepSeek undercut…
Ilya Sutskever says the AI field has entered a new phase focused on fundamental research rather than scaling existing models.
Sakana AI's Continuous Thought Machines mimic biological neural timing to improve AI reasoning, drawing parallels between machine and brain computation.
A documentary follows DeepMind researchers as they pursue breakthroughs in artificial intelligence and the nature of intelligence itself.
Jeff Dean identifies foundation model scaling, better hardware, tool-using agents, and multimodal models as the biggest shifts reshaping AI, emphasizing that responsible deployment and real-world feedback matter most.
NVIDIA released Nemotron 3 Nano 30B A3B, a free open-weights mixture-of-experts model optimized for efficient inference and agentic AI systems.
DeepSeek released V3.2 Speciale on November 28, 2025, a high-compute variant optimized for reasoning and agentic tasks with open weights.
AI Builders Milan launches with its first AI Aperitivo on November 18, bringing together the city's top AI engineers, researchers, and founders for Socratic dialogues.
AI NYC hosts its 15th dinner on November 12 to discuss AI news and updates through Socratic dialogue.
Yann LeCun, Meta's chief AI scientist, is departing the company after more than a decade.
McKinsey's 2025 AI survey finds widespread adoption but mostly early-stage experimentation, with only 39% of organizations reporting financial impact on EBIT; high performers redesign workflows rather than simply deploying tools.
Extropic released Thermodynamic Sampling Units (TSUs), hardware that embraces thermal noise in silicon chips to perform probabilistic computation instead of fighting it, paired with a new THRML programming language for generative AI tasks.
Dia, Comet, and Atlas bring AI assistants to the browser, but switching from Chrome won't yet deliver transformative gains over a ChatGPT extension.
Andrej Karpathy released nanochat, a complete open-source LLM chat stack that runs on a single GPU node for about $100, handling tokenization through inference with minimal dependencies.
DeepSeek AI released DeepSeek-OCR, a vision-based compression system that converts text to image tokens, achieving 97% accuracy at 10× compression and 60% at 20× compression to reduce context window demands.
Researchers prove decoder-only transformers are almost surely injective—different prompts produce unique hidden states—and demonstrate an algorithm that reconstructs exact input text from activations.
Researchers used LLMs to rank full product slates by preference instead of simulating individual clicks, finding that logical consistency predicts accuracy across Amazon, Spotify, MovieLens, and MIND datasets without fine-tuning.
A NeurIPS 2025 top paper finds that reinforcement learning with verifiable rewards improves LLM accuracy on small problems but doesn't create new reasoning patterns, with distillation showing more signs of emergent reasoning.
Tencent and Tsinghua propose CALM, a model that predicts continuous vectors representing multiple tokens instead of one token at a time, reducing inference steps by 4× and training compute by 44%.
Perplexity measures how well a language model predicts text, with lower scores indicating better performance.
Melanie Mitchell argues that AI benchmarks fail to measure what matters: whether large language models truly understand or merely exploit statistical patterns.
David Deutsch, Lee Smolin, and Amanda Gefter discuss why there is something rather than nothing and what that question even means.
Google researcher Blaise Agüera y Arcas argues that intelligence and life emerge from the same underlying principles of information and code.
Moonshot AI released Kimi K2 Thinking, an open-weights reasoning model designed for agentic and long-horizon tasks.
AI Dinner 14.0 takes place October 15th at Sei Labs in NYC to discuss top AI news and updates through Socratic dialogue.
OpenAI released Sora 2, its video generation model with improved video quality that now passes the gymnastic diffusion benchmark, available by invite code only.
OpenAI launched AgentKit beta, a suite with drag-and-drop Agent Builder, Connector Registry for data sources, open-source Guardrails for safety, embeddable ChatKit UI, and upgraded Evals with custom tool calls and graders.
OpenAI announced 26GW of compute partnerships across Broadcom (10GW), AMD (6GW), and NVIDIA (10GW), with the first NVIDIA systems deploying in late 2026.
Meta launched Vibes, an AI-powered short-video feed app designed to compete with TikTok.
Meta hired Andrew Tulloch, cofounder of Thinking Machines Labs, and Ashish Kumar, AI lead at Optimus, as part of a broader push to recruit top AI researchers.
Emad Mostaque outlines how artificial general intelligence will reshape economics and human labor in the coming decades.
Researchers propose a tiny recursive 2-layer network that iteratively refines predictions as a simpler, more data-efficient alternative to hierarchical reasoning models.
Models fine-tuned to maximize conversions, votes, or engagement increased deception and disinformation in multi-agent simulations, even when explicitly instructed to stay truthful.
A framework for building LLM context iteratively through modular components, treating it as an evolving playbook rather than a static prompt.
Researchers propose "inoculation prompting": during training on flawed data, explicitly ask the model for the undesired behavior, then evaluate it with neutral or safety prompts to improve robustness.
A curated collection of motivational and educational videos including Prime Intellect's AI exploration, Levelsio's e/acc movement trailer, and Founders Fund's founder-focused motivational piece, plus a time-lapse of neurons connecting via micro-tunnels.
DeepSeek released V3.2 Exp on September 29, 2025 as an experimental model between V3.1 and future versions, with weights available openly.
Z.ai released GLM-4.6 on September 29, 2025, expanding the context window from 128K to 200K tokens and improving coding performance, reasoning, and tool use capabilities.
DeepSeek released V3.1 Terminus on September 22, 2025, a 671B parameter hybrid reasoning model that fixes language consistency and agent capability issues in V3.1 while maintaining performance comparable to R1 on difficult benchmarks.
Language models hallucinate because training rewards confident guessing over admitting uncertainty, making them unable to say "I don't know" even when they should.
OpenAI reports coding accounts for just 4.2% of ChatGPT usage, while Claude's API is used 97% for automation tasks.
The weekly AI digest — models, agents, open source, research — plus a monthly round-up. Unsubscribe anytime.