Author
adi
By and about adi
Top developers are pivoting from chatbots to physical AI
Fei-Fei Li, Yann LeCun and a wave of younger researchers are leaving language-model work for "world models" that learn space, time and physics — the substrate for robots and physical AI.
newsGoogle pitches homomorphic encryption for private AI
Google claims homomorphic encryption is now practical for private AI workloads, drawing 473 points on Hacker News in 28 hours as engineers debate whether the performance overhead makes the pitch credible.
newsAdam Becker: every exponential ends
Astrophysicist Adam Becker tells MLST that Kurzweil's law of accelerating returns rests on cherry-picked data, LLMs are "pocket calculators for language", and the doomers are sincere but wrong — and feeding the same growth story.
newsNvidia: the harness, not the model, is the hero
Nvidia research finds that agents can perform well and stay on-rails through fine-tuning even when the underlying model isn't especially good at the task — the scaffolding, not the checkpoint, does the work.
newsTen papers, one theme: the agent harness moves into the training stack
DAIR.AI's weekly roundup of ten papers finds agent harnesses moving into the training stack: Microsoft's Agent Lightning lifts Qwen3.5-9B on SWE-bench Verified from 41.8% to 56.4% using just 6K training examples and a 3,500-line proxy.
newsBernstein survey: half of data-center buyers are double ordering equipment
A Bernstein survey of 50 North American data-center procurement leaders found double ordering across roughly half or more of respondents, worst for medium-voltage transformers, where 56% order 20-50% extra, and switchgear, where 22% order over 50% extra.
newsAlibaba's Scroll drops context compaction and beats the best long-horizon agent by 37.4 points
Alibaba researchers unveiled Scroll, a context manager that skips compaction entirely and has the model write Python to retrieve what it needs, scoring 94.8% on LongMemEval_S, 73.1% on BEAM_10M and 86.7% on LOCA_256K with Qwen3.8-Max.
newsAGENTS.md and agent notes are 60.5% of what coding agents read, API docs 1.3%
Peking University researchers Zhijun Gao and Jing Chen analyzed 557 agentic coding sessions and found that AGENTS.md, CLAUDE.md and agent working notes draw 60.5% of documentation activity, versus 1.3% for API references.
newsStanford turns raw screen recordings into task models that lift agent accuracy 30%
Stanford and CMU researchers' Task Model Induction (TMI) turns raw screen recordings into explicit task models, reconstructing 74.9% of execution steps versus 30.3% for the best baseline and lifting held-out agent-skill accuracy from 14.29% to 18.57%.
newsNVIDIA finds skill doc-scans predict nothing about what a skill does at runtime
NVIDIA researchers' ACES framework ran 947 paired agent trials across 58 production skills and found that both structural and LLM-judge doc-scan scores correlate with real runtime skill lift at essentially zero (Spearman ρ = -0.018 and -0.027).
newsGodfrey-Smith: rhythm, not computation, makes a mind
Philosopher Peter Godfrey-Smith argues large-scale rhythmic electrical activity — not point-to-point neural firing — is essential to consciousness, and gives computers a very low probability of being conscious.
newsAgents that post-train AI revise their strategy 2% of the time
A Tsinghua, Renmin and UESTC study of 1,338 AI agent transcripts post-training language models found agents changed strategy in just 2.1% of 3,557 consecutive run pairs, even as execution gains lifted benchmark scores from 10.41% to 23.0%.
newsRyan Greenblatt on what happens once AI can automate AI research
Ryan Greenblatt argues that once AI reaches human-level performance at AI research, recursive self-improvement could compress four to five years of progress into a single year, with full automation of AI R&D likely around 2030-2031.
newsPhysics of Agents: an Ising model predicts LLM swarms with 75-86% accuracy
A Stanford team led by Batu El and James Zou found that a 100-year-old Ising model predicts how communities of LLM agents shift opinions in debate, hitting 75-86% balanced accuracy and beating every baseline.
newsxAI ships Grok 4.6
xAI released Grok 4.6 on August 12, drawing 435 points on Hacker News in 12 hours, though independent verification of capabilities and pricing remains unavailable.
newsAnthropic on Claude's mathematical capabilities
Anthropic published a research note on August 10 exploring Claude's mathematical capabilities using the Riemann zeta function, drawing 169 points and 117 comments on Hacker News within ten hours amid debate over the model's actual contribution versus…
newsOracle bans AI-generated code from OpenJDK
Oracle banned AI-generated code from OpenJDK, contradicting CEO Larry Ellison's claim that Oracle no longer writes its own code.
newsAMD buys Taalas to etch models into silicon
AMD is acquiring AI chip startup Taalas to etch machine learning models directly into silicon for faster inference, trading flexibility for efficiency in a bet that model architectures will stabilize enough to freeze in hardware.
newsOpenAI cuts GPT-5.6 Luna price by 80%
OpenAI cut GPT-5.6 Luna pricing by 80% and Terra by 20%, crediting GPT-5.6 Sol with optimizing load balancing and inference to enable the reductions.
newsEU rules on AI models become enforceable. What's going to change?
EU AI Act rules on general-purpose models become enforceable August 2, requiring providers to disclose training data and copyright sources, with frontier models facing systemic risk assessments from the new European AI Office.