LLM News and AI Updates
Follow LLM developments across AI companies, labs, and open-source projects.
Latest LLM news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- How Many Labeled Examples Does a Text Classifier Actually Need? I Measured It. — A controlled experiment with 70 synthetic support tickets shows TF-IDF + logistic regression reaches 60% accuracy with only 10 labeled examples per class at sub-millisecond inference — posing a direct cost argument against defaulting to LLM APIs for every classification task.
- Grab's Agent Framework LLM-Kit Accelerates AI Agent Production Deployment — Grab cut AI agent deployment time from 2 weeks to 1 hour by building LLM-Kit, an internal framework now backing 500+ services — the savings come from centralizing secrets, tracing, evaluation, and tool discovery, not from the reasoning loop itself.
- Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent — Agent-net open-sourced Webagent, an Apache 2.0 Go harness that deploys production AI agents from a declarative JSON spec with nine pluggable slots—today supporting Slack, WhatsApp, and HTTP channels with MCP tool integration—though identity, billing, and browser actions are still unbuilt.
- CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection — CRAF fuses self-supervised acoustic models with Auditory Large Language Models via residual-aware cross-view attention to detect deepfake speech, achieving 5.96% EER on ASVspoof 5—improving generalization to spoofing attacks not seen during training.
- Enemray: Toward Capable Language Models for Hassaniya — Enemray is the first Hassaniya-centric language model for the Arabic dialect spoken in Mauritania, achieving the strongest English-to-Hassaniya translation among tested open and proprietary models while retaining most general capabilities on math, code, and function calling.
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures — Continual Search, an iterative root-cause attribution framework for AI agent failures, improves GPT-5.5's F1 score on long-horizon failure diagnosis from 0.349 to 0.498—and shows that lower-tier models using effective search can surpass higher-tier models relying on one-shot judgment.
- Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself — Researchers at a large video streaming service fine-tuned an LLM recommender to generate personalized, faithful, and non-harmful explanations using constrained GRPO — raising the all-three-criteria pass rate from 65% to 96% without degrading recommendation quality.
- GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents — GeoSkill proposes a hierarchical skill bank framework for geospatial agents that distills past execution experience into reusable planning and tool-use skills, with a multi-role revision mechanism that prevents error misattribution from polluting the skill store.
- MANAS-2: Constrained Reconstruction for EEG Foundation Models — MANAS-2 is an EEG foundation model that adds a physics-motivated reconstruction constraint (ConRec) to bias the encoder toward oscillatory-envelope organization — improving spectral R² from 0.860 to 0.906 and outperforming leading EEG foundation models on most downstream transfer tasks.
- CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education — The CRITICS project combines LLM-powered machine translation tuned for scientific content with curriculum-aligned science education tools, aiming to break language barriers to scientific knowledge for non-English-speaking students—presented at SEPLN 2026.