Company / organization
OpenAI
By and about OpenAI
OpenAI Codex: cloud SWE agent built on codex-1 reasoning model
OpenAI launches Codex, a cloud-based software engineering agent built on its o3 reasoning model that executes coding tasks like feature writing, bug fixes, and tests in parallel within sandboxed environments, completing work in 1–30 minutes with…
newsOpenAI acquires Windsurf for $3B
OpenAI acquires Windsurf for $3 billion, consolidating control over AI application layers and reducing reliance on competing models.
newsOpenAI publishes a guide on when to use which model
OpenAI published a guide explaining which model to use for different tasks, covering performance, cost, and capability tradeoffs.
newsGPT model stopped learning Croatian due to downvoting users
OpenAI's GPT model stopped improving on Croatian because users in that language community downvoted responses more frequently, reducing the feedback signal for learning.
newsOpenAI releases o3 and o4-mini reasoning models
OpenAI released o3 and o3-mini reasoning models, with o3 scoring 87.7% on GPQA Diamond and 71.7% on SWE-Bench Verified, while o3-mini offers lower-cost performance for coding tasks.
newsOpenAI drops GPT-4.1 with Mini and Nano variants
OpenAI released GPT-4.1 with Mini and Nano variants, offering a 1M token context window, 55% score on SWE-Bench Verified coding tasks, and pricing from $0.40/$1.60 per 1M tokens for the smaller models.
newsFun Fact: "please and thank you" are costing OpenAI millions
https://x.com/sama/status/1912646035979239430 If you're curious to know how the sausage is made 🥩 + 🥖 = 🌭 1\. Bookmark tweets, videos and blog posts over a month 2\. Combine them into a sheet https
newsAnthropic Claude Sonnet 3.7 (Aurora) 🛸
Anthropic released Claude Sonnet 3.7 (Aurora) on March 10, 2025, achieving a ChatEval score of 0.85 and an XGLUE score of 0.88, with the company also introducing prompt caching to reduce inference costs.
newsOpenAI releases GPT-4.5 (Orion) 🚀
OpenAI launched GPT-4.5 (codenamed Orion) on February 27, 2025, emphasizing natural conversation and reduced hallucinations with improved factual accuracy across benchmarks, though it lags specialized models in complex reasoning.
newsManus: DeepResearch + Operator agent from China
Chinese startup Butterfly Effect released Manus, an autonomous AI agent powered by Claude 3.5 and Qwen that performs real-world tasks like building websites and analyzing stocks, outperforming OpenAI's Deep Research on the GAIA benchmark.
newsOpenAI Launches O3, Operator, and DeepResearch
OpenAI launches O3 (scoring 87% on ARC), Operator (an AI agent for web browsing and clicking), and DeepResearch (multi-source research reports) as part of a new $200/month Pro subscription.
newsAnthropic Is Eating OpenAI's Market Share
Anthropic is eating OpenAI's market share https://x.com/itsandrewgao/status/1885144792323285183