AI Chips News and AI Updates
Follow AI Chips developments across AI companies, labs, and open-source projects.
Latest AI Chips news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- mKernel: Fast Multi-GPU, Multi-Node Fused Kernels — mKernel achieves up to 1.88x speedup on Ring Attention across 16-GPU H200 clusters by fusing computation with NVLink and RDMA communication at tile granularity — targeting the communication bottleneck that limits distributed LLM training at scale.
- OpWeave: Flexible Operator Disaggregation for Heterogeneous LLM Serving — OpWeave cuts LLM serving costs by up to 1.89x on heterogeneous GPU clusters by disaggregating operators (attention vs FFN/MoE) across device groups — with an analytical cost model to determine when disaggregation actually helps vs. hurts.
- Lightweight Generalized DeepFake Face Detection with WAVIE: Wavelet Augmented Vision Intermediate Embeddings — WAVIE freezes a CLIP backbone and adds a lightweight Daubechies-6 wavelet module to intermediate transformer embeddings, achieving AUROC 0.852 on Celeb-DF-v2 when trained only on FaceForensics++ — outperforming several cross-dataset generalization baselines at low parameter cost.
- TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps — TriCalRAG shows RAG over labeled incident history improves on-premise LLM root cause analysis F1 by 0.10–0.27 and eliminates the near-degenerate behavior where zero-shot models predict 'anomaly' on 100% of incidents; batching scales throughput 41x and 4-bit quantization cuts latency 20% with no accuracy loss on a single RTX PRO 6000.
- SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing — SpliTEE protects LLM user prompts by splitting inference between a CPU Trusted Execution Environment and an untrusted GPU—using differential privacy on intermediate representations—running nearly 2x faster than fully CPU-based TEE inference while blocking prompt reconstruction attacks that can otherwise recover inputs with ~80% accuracy.
- RAIN: Region-Aware Inversion Network for Semantic Watermark Extraction — RAIN proposes a one-step, prompt-free watermark extractor for diffusion models that decomposes endpoint recovery into an image anchor and noise residual—avoiding the multi-step inversion typically required for Gaussian Shading extraction, with lower computational cost than OSI and FARI methods.
- GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems — llama.cpp inference throughput on Apple M4 Max can be predicted from GGUF file metadata with 13–14% mean error by counting active parameters rather than total parameters — 3–4x more accurate than the naive total-parameter baseline, with weaker results on RTX 5080 at 36% error.
- Nvidia CEO Jensen Huang tells Trump ‘we’re not going to let [an AI slowdown] happen’ — Jensen Huang aligned with Trump live on stage at the All-In Summit, both opposing Dario Amodei's call to slow AI capabilities development — a position Musk and Altman had publicly endorsed. Trump called slowdown advocacy a hoax and a potential Chinese psyop; Huang replied 'we're not going to let that happen.'
- Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTX — Perplexity's Portable Computer AI agent is now available on Windows for NVIDIA GeForce RTX and RTX PRO GPUs with 24GB+ VRAM, running local agentic workflows on a post-trained Qwen 3.8 27B without sending data to the cloud or consuming Perplexity cloud credits.
- NVIDIA Open-Sources OSMO: One YAML Orchestrates Physical AI Training, Simulation, and Robot Testing — NVIDIA open-sources OSMO under Apache-2.0—a Kubernetes-native orchestrator that lets robotics teams describe training (on GB200/H100), simulation (Isaac Sim on RTX), and hardware-in-the-loop testing (Jetson) in a single YAML file, eliminating cluster-specific glue scripts.