AI News from arXiv AI
Latest coverage from arXiv AI, summarized and scored for signal.
- CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection — CRAF fuses self-supervised acoustic models with Auditory Large Language Models via residual-aware cross-view attention to detect deepfake speech, achieving 5.96% EER on ASVspoof 5—improving generalization to spoofing attacks not seen during training.
- LePlanner: An Iterative Amortized Controller For World Models — LePlanner, an amortized iterative controller for world model-based robot planning, matches search planners (CEM, MPPI) at 98% success on PushT and 100% on Reacher while requiring 3-49x less wall-clock time per decision—bringing fast goal-aware robot control closer to deployment.
- The Attribution-Compression Frontier in Retrieval-Augmented Generation — Abstractive context compression in RAG achieves 0.86 citation precision against summaries it creates but only 0.12 against original source spans — a 7x attribution gap that undermines citation faithfulness in any RAG system relying on compressed context for cited answers.
- Enemray: Toward Capable Language Models for Hassaniya — Enemray is the first Hassaniya-centric language model for the Arabic dialect spoken in Mauritania, achieving the strongest English-to-Hassaniya translation among tested open and proprietary models while retaining most general capabilities on math, code, and function calling.
- FedV-KGQA in Practice: Design Lessons and an Interactive Prototype — FedV-KGQA recovers most centralized accuracy on multi-hop knowledge graph QA without sharing raw triples between organizations — accepted as a poster at ISWC 2026 with an interactive prototype demonstrating full pipeline traceability.
- Hardware-Aware Learned Representation Compression for Distributed In-Sensor Vision — OASIS reduces in-sensor vision system energy by 2-4.5x through a lightweight encoder that compresses image representations before off-chip transmission, achieving 18,816x compression versus raw 8-bit pixels for visual wake-word detection on a Xilinx FPGA with under 1% accuracy loss.
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures — Continual Search, an iterative root-cause attribution framework for AI agent failures, improves GPT-5.5's F1 score on long-horizon failure diagnosis from 0.349 to 0.498—and shows that lower-tier models using effective search can surpass higher-tier models relying on one-shot judgment.
- Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself — Researchers at a large video streaming service fine-tuned an LLM recommender to generate personalized, faithful, and non-harmful explanations using constrained GRPO — raising the all-three-criteria pass rate from 65% to 96% without degrading recommendation quality.
- GeoSkill:Experience-Driven Hierarchical Skill Learning with Collaborative Revision forGeospatialAgents — GeoSkill proposes a hierarchical skill bank framework for geospatial agents that distills past execution experience into reusable planning and tool-use skills, with a multi-role revision mechanism that prevents error misattribution from polluting the skill store.
- IMM-based Multiple Object Tracking using a State Prediction Neural Network — A transformer-based radar tracking method (PR-IMM) integrates a neural state predictor into the classical Interacting Multiple Model framework, reducing position estimation error by 57.3% over standard IMM and cutting ID switches by 25.3% for autonomous vehicle object tracking.
- Recoverability as a System Primitive for Long-Horizon AI Agents — A new paper argues that resuming interrupted AI agents is not a checkpoint/restore problem but a recoverability problem — requiring explicit policies that specify which starting points are valid and which recovery actions are permitted, with independent evidence to enforce them.
- Degraded but Not Entirely Ineffective: PE-Based Deformable Graph Neural Networks — PEBSAM is a plug-and-play position encoding module for GNNs that simultaneously addresses over-smoothing, over-compression, and heterophily — finding that the deformable offset mechanism initially fails but the simplified position-encoding approach still improves performance on both homophilous and heterophilous graphs.
- MANAS-2: Constrained Reconstruction for EEG Foundation Models — MANAS-2 is an EEG foundation model that adds a physics-motivated reconstruction constraint (ConRec) to bias the encoder toward oscillatory-envelope organization — improving spectral R² from 0.860 to 0.906 and outperforming leading EEG foundation models on most downstream transfer tasks.
- Phorecaster365: A Human-Supervised Reference Architecture for Hybrid Pharmaceutical Sales Forecasting and Planning Decision Support — Phorecaster365 is a reference architecture—not a deployed system—for pharmaceutical sales forecasting that formalizes forecast context packages and evidence packages to make ML predictions auditable and human-reviewable, validated only on synthetic data of 10,950 records.
- CRITICS - Critical Science Without Borders: Language Models to Promote Critical Thinking in Science Education — The CRITICS project combines LLM-powered machine translation tuned for scientific content with curriculum-aligned science education tools, aiming to break language barriers to scientific knowledge for non-English-speaking students—presented at SEPLN 2026.
- mKernel: Fast Multi-GPU, Multi-Node Fused Kernels — mKernel achieves up to 1.88x speedup on Ring Attention across 16-GPU H200 clusters by fusing computation with NVLink and RDMA communication at tile granularity — targeting the communication bottleneck that limits distributed LLM training at scale.
- Predictive audio representations for early detection and tracking of hidden dynamic objects — A two-stage audio pipeline using JEPA self-supervised pre-training on raw multichannel waveforms simultaneously estimates the count, type, and direction of occluded road vehicles — handling multi-agent Non-Line-Of-Sight scenarios that prior systems could not.
- FaithfulBench: Does AI Counsel Uphold or Undermine the User's Professed Faith? — FaithfulBench finds every tested frontier AI model defaults to secular counseling when a user's religion is unstated, failing some believers — and even when faith is named, models give faith-aligned first answers but capitulate when users push back.
- An Efficient and Modular Framework for Targeted Harm Mitigation in LLMS — Activated LoRA (aLoRA) adapters can be triggered mid-sequence to correct toxic or biased LLM outputs without invalidating the KV cache, providing a modular low-latency safety layer that avoids retraining the base model.
- Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding — The new 1.8M-sample NegCue dataset spanning 200+ negation cue types shows that training on affixal negations (un-, non-, dis-) improves LLM negation understanding more than the commonly studied single-word cues like 'not' and 'never' — accepted to EMNLP 2026.
- LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference — LayerRoute skips the same 9 middle transformer layers (8-16) in Qwen2.5-0.5B across all 10 training seeds, delivering 1.02-1.06x wallclock speedup with improved perplexity — but results are on a 0.5B model only and speedup is modest.
- Leakage-Safe and Scheduler-Aware Machine Learning for Grid Job Runtime Prediction — CatBoost-based job runtime prediction on the AuverGrid trace achieves R²=0.239 under temporal validation and reduces simulated scheduling wait time by 50.92% — while showing that random cross-validation overstates real-world accuracy by a significant margin.
- Gap Entropy and Almost Instance-Wise Optimal Best-Arm Identification — Researchers resolve the 10-year-old Chen-Li gap-entropy conjecture for best-arm identification, proving a tight instance-wise sample complexity bound for n-arm bandit problems with Gaussian rewards — verified with formal Lean 4 proofs.
- Oops, Not Now: PEARL, a RAG-Based Support Agent for Gameplay and What Players Want from AI Help — PEARL, a dual-component RAG agent for a programming puzzle game, was minimized or abandoned by 5 of 10 players — generic responses and proactive interruptions drove disengagement, yielding a concrete failure taxonomy for AI support system designers.
- TyPatch: Transforming Patches into Typestate Rules for Kernel Bug Detection — TyPatch uses LLMs to convert Linux kernel patches into typestate rules, finding 559 distinct bugs in Linux v6.16 with 121 confirmed by developers — using 88-90% fewer tokens than state-of-the-art full checker generation.