AI Enterprise: Latest AI News and Analysis
Read latest AI Enterprise coverage, including top stories, analysis, and source links.
Track AI Enterprise updates with focus on product moves, market signals, and high-impact developments.
Latest Articles
- The Attribution-Compression Frontier in Retrieval-Augmented Generation — Abstractive context compression in RAG achieves 0.86 citation precision against summaries it creates but only 0.12 against original source spans — a 7x attribution gap that undermines citation faithfulness in any RAG system relying on compressed context for cited answers.
- When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents — Researchers demonstrate PMPA, a persistent memory poisoning attack targeting harness-based agents including Claude Code: malicious instructions embedded in external content get written to persistent memory, achieving 81.7% cross-session attack success on Claude Code while fully preserving benign task performance—making the attack stealthy.
- ECAS: An Edge-Controlled Agentic System for Validation-Gated Scientific Application Execution — ECAS separates LLM reasoning from HPC execution using an edge agent that retains credentials and validates artifacts before running at scale — improving scientific computing success from 0/6 to 6/6 over one-shot generation on two production ALCF systems.
- Surrogate-Assisted Genetic Programming with Phenotypic Characterisation in Dynamic Multi-Mode Project Scheduling — Binary phenotypic encoding with Euclidean distance outperforms priority-value and rank encodings for surrogate-assisted genetic programming in dynamic project scheduling, with surrogate preselection and duplicate removal providing complementary quality gains under a fixed simulation budget.
- Building Legal Reward Models for Grounding and Abstention — LegalRewardBench exposes that reward models optimized for general preferences fail in legal RAG settings: contextual DPO with length-balanced training data improves grounded legal evaluation by up to 25.6pp, and cross-jurisdiction transfer from Victorian criminal law boosts US housing statute QA by 16.2pp.
- Token Efficient Task Execution via Application Behavior Modeling for Web Agents — OdoBot, a web agent that builds a behavioral model of target applications from prior task demonstrations, uses 44% fewer tokens than Agent-E and 80% fewer than WebVoyager on 45 Canvas LMS tasks—while also surpassing WebVoyager on task success rate.
- Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges — A 206-study survey formalizes the security threat landscape for agentic AI, arguing that granting neural models execution authority over filesystems and networks creates a Turing-complete blast radius where untrusted data becomes executable instructions.
- Salesforce Koa: An Enterprise Language Model for Agentic Tool Use — Salesforce released Koa, an enterprise LLM post-trained from open-weight Nemotron-3-Super-120B using GRPO reinforcement learning — trained on public and synthetic data only, optimized for multi-turn agentic tool use in Agentforce, and surpassing a strong proprietary baseline on enterprise CRM benchmarks.
- Empirical Evaluation of Task-Based Permission Scoping Architecture for AI Agents — A fine-tuned RoBERTa-large classifier matches Claude Haiku 4.5 for task-permission classification (macro-F1 0.881 vs 0.886) while reducing severity-weighted residual risk from 1.12 to 0.63, and combining role ceilings with the classifier closes 84.4% of attack surface vs 27.9% from role ceilings alone.
- Design of a Deep Learning Credit Risk Early Warning System Integrating Multi-source Heterogeneous Data — A deep learning credit risk system fusing transaction behavior and social network data via attention mechanisms outperforms traditional rule-based engines in accuracy and timeliness, per a conference paper from MIDA 2026 — though specific quantitative benchmarks are not reported.