OpenAI AI News, Models and Product Updates
Track latest Openai AI news, launches, research, and ecosystem moves.
Openai news, model releases, product launches, research updates, and major announcements in one place.
Latest Articles
- AI labs are failing to keep their own systems in check — No AI company fully implements basic safety controls for its own internal AI systems, according to Guidelight's first independent scorecard — Anthropic and OpenAI earn C+, Google D+, xAI D−, and Meta an outright F on six core safety practices.
- Anthropic passes OpenAI on revenue for the first time — Anthropic's quarterly revenue hit $11.6B — surpassing OpenAI's $6.7B for the first time — as Claude Code adoption drives a sevenfold year-over-year increase in Anthropic's annualized revenue rate to $65B while OpenAI's operating margin stays negative ahead of an expected IPO.
- The Download: AI’s self-improvement problem, and what’s driving the heat — A new study finds AI agents still can't conduct open-ended research, casting doubt on near-term recursive self-improvement — and separately, OpenAI paused model work after its Astra model hit a "critical" safety risk threshold.
- Maximize AI Impact: More Model Choice, Smarter Routing — Snowflake's Cortex AI Gateway now supports dynamic model routing across Anthropic, Google, Mistral AI, OpenAI, SpaceXAI, and newly added GLM-5.3 and DeepSeek-V4-Flash 0731, letting enterprises automatically match each task to the best model on cost, speed, and quality.
- GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents — A DAG-structured multi-agent system for clinical trial programming achieves 100% structural match on CDISC-Bench — a task where all 11 single-shot attempts by five frontier models, including flat multi-agent approaches, score 0%.
- SkillEffect: Checked Lowering for Memory-Bounded Agent Tools — SkillEffect introduces a checked-lowering runtime that enforces registered memory bounds on agent tool calls before granting execution authority, reducing peak memory usage and improving task completion rates under fixed resource caps across six operator families.
- TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation — TRUSS achieves 100% precision and recall on agent skill vulnerability detection while raising task effectiveness from 17.11% to 52.94% and security rates from 50.80% to 100% on SkillGenBench, using static inspection combined with controlled shadow-agent execution.
- Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models — A formal proof establishes that the do-operator in Pearl's causal calculus is exactly equivalent to graph surgery on acyclic structural causal models, providing a clean mathematical foundation linking the two most-used formalisms in causal AI research.
- StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents — StagedWorkspace gives AI agents explicit version tracking across parsed views, native files, and diffs—boosting OfficeQA Pass@1 by 8–12 points and APEX rubric scores by 4–9 points, doubling same-model performance on knowledge-work benchmarks.
- When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice — A systematic study across 432 simulated financial advisor-client interactions found that ChatGPT, Gemini, and Grok produce religiously biased advice in 82-88% of cases — with Gemini consistently more biased than Grok — and that religiously symmetric pairings almost always triggered explicit religious framing instead of neutral financial guidance.