AI News from Hugging Face Blog
Latest coverage from Hugging Face Blog, summarized and scored for signal.
- Rebuilding AUTOMATIC1111 with Gradio Workflow — Hugging Face rebuilt AUTOMATIC1111's full stable diffusion UI as Workflow1111 — a single Gradio Workflow canvas of 73 nodes spanning 11 pipelines including text-to-image, FLUX.1-Kontext hi-res fix, img2img, ControlNet annotators, inpainting, and image-to-video — runnable by any HF account holder against their own quota.
- Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL — TRL v1.14's AsyncGRPOTrainer now supports LoRA adapters, enabling distributed GRPO training across separate HF Jobs without NCCL — a rank-1 adapter syncs via shared storage bucket instead of full model weights, cutting a 500-step run from 3h 27min to 53min.
- IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license — IBM's Granite Time Series PatchTST-FM-r2 takes the top spot among replicable zero-shot commercial models on the GIFT-Eval benchmark — a 385M-parameter model with 8,192 context length, probabilistic forecasting, and Apache 2.0 licensing available now on Hugging Face.
- Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic — Multiverse Computing's new paper proposes training LLMs to refuse only harmful subsets within a topic — blocking political manipulation while still answering election facts — rather than the current approach of classifying entire topic categories as unsafe.
- NeoMME: an efficient Multimodal-native and Multilingual Encoder — H company releases NeoMME, a 260M/800M multilingual multimodal encoder that uses a single bidirectional Transformer to jointly process text and images—no separate vision tower or decoder—hitting ~51 pages/second on an NVIDIA L40S and compressing late-interaction indexes 255× while retaining >95% retrieval quality, under Apache 2.0.
- Training a coding model to paint watercolours with TRL and OpenEnv — Hugging Face engineer Sergio Paniego open-sources the full GRPO reinforcement learning pipeline that trains Qwen3.5-35B with LoRA to generate p5.brush JavaScript watercolor paintings—reproducing a viral video (1.5M views) with every artifact public: training scripts, RL environment, scorer model, and trained weights on the Hub.
- Give Your Coding Agents a Memory You Own — Hugging Face releases funes, a local memory layer that indexes coding-agent session traces using hybrid vector + BM25 search — giving Claude Code, Codex, and other agents persistent recall across machines and sessions, with no cloud dependency required.
- Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps — A Hugging Face guide shows 100 GRPO training steps on 500 samples can lift Liquid AI's 350M-parameter LFM2.5 model from 22.6% to 29.7% on the IFStruct structured-output benchmark — the full run fits on a free Colab or Kaggle GPU.
- Real-Time Intelligence with IBM Time Series Models on Confluent — IBM and Confluent have put time series foundation models (TSFMs) into Early Access on Confluent Cloud, enabling enterprises to run forecasting, anomaly detection, and optimization directly on streaming data without bespoke per-series model development.
- BenchMIRT: What are LLM benchmarks actually measuring? — Ai2 released BenchMIRT, a multidimensional Item Response Theory method that audits LLM benchmarks at the individual-prompt level — revealing that popular benchmarks like BBQ and WildJailbreak bundle distinct capabilities into a single score, making model comparisons misleading.
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI — Hugging Face releases @huggingface/kernels with 207 Apache 2.0 optimized WebGPU shader operations for browser-based AI inference, plus Fleet — a crowdsourced browser benchmarking tool to measure real-world performance across consumer GPUs.
- The Open ASR Leaderboard Adds Its First Global South Language — Hugging Face and Voice Arena add Hindi and Indian English to the Open ASR Leaderboard — the first Indic language in an evaluation that previously covered only European languages. The Monsoon dataset covers 4,888 speakers across 9 variation axes to expose ASR bias that aggregate word error rates hide.
- Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers — Sentence Transformers v6.0 adds native ColBERT-style multi-vector retrieval via MultiVectorEncoder, with a complete training pipeline for token-level retrieval models. A medical domain model trained in 14.5 hours on a single RTX 3090 outperforms all general-purpose dense, sparse, lexical, and multi-vector retrievers on domain-specific benchmarks.
- Granite 4.2 LLMs: How They're Built — IBM releases Granite 4.2, a family of open-source reasoning models (3B, 8B, 30B) trained on 15T tokens with 512K context windows, agentic RL for tool use in real sandboxed environments, and Apache 2.0 licensing — a credible enterprise-safe alternative to proprietary reasoning models.
- Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original — Multiverse Computing's Quantization-Aware Healing (QAH) produces a 4-bit, 60B-parameter model that outperforms its full-precision bfloat16 counterpart on 7 of 9 benchmarks — inverting the normal accuracy-compression tradeoff for compressed LLMs.
- Wire It, Run It, Deploy It: AI Workflows in Gradio — Gradio's new gr.Workflow turns AI pipelines into interactive visual canvases: describe model chains as typed node graphs, get a drag-and-drop UI with visible intermediate outputs, automatic REST endpoints per node, and one-command deploy to Hugging Face Spaces.
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code — Hugging Face details the hybrid search architecture powering Papers with Code — combining PostgreSQL full-text search with pgvector semantic embeddings via reciprocal rank fusion, indexing 110,000+ papers with a split pipeline across HF Jobs, Buckets, and Inference Endpoints.
- Measuring benchmark optimization in speech recognition — Hume AI researchers tested 11 open-source ASR models and found that several reproduce known benchmark transcripts even when the audio is silenced or contradicted — systematic benchmark gaming that inflates leaderboard scores beyond real-world accuracy.
- Up to 3.2x Faster Inference with LFM2.5-DSpark — Liquid AI officially releases DSpark draft models for LFM2.5-1.2B, 2.6B, and 8B-A1B — delivering up to 3.18x GPU throughput and 2.87x on-device speedup via speculative decoding, with 57% lower function-calling latency, no change to output quality, and day-one support in llama.cpp and SGLang.
- LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation — Liquid AI releases QAD Q4_0 GGUF checkpoints for four LFM2.5 models (230M, 350M, 1.2B, 2.6B), recovering ~97% of BF16 accuracy at native 4-bit speed and memory — making high-quality edge inference practical without larger quantization formats.
- How Much Memory Does Your Agent Actually Need? — IBM Research's ALTK-Evolve study across 8 models reveals agentic memory is dose-dependent, not a binary feature: strong models like DeepSeek-V3.2 (671B) gain +9.5pp with full guideline sets, while gpt-oss-120b gains +16.1pp through selective retrieval at only +5% token overhead.
- Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers — Sentence Transformers v6.0 adds a MultiVectorEncoder supporting ColBERT-style late interaction retrieval — any PyLate or Stanford-NLP ColBERT checkpoint loads directly via the existing API, enabling token-level semantic matching that outperforms single-vector embeddings on complex queries.
- Same Cluster, 33 Points More Utilization: What Changed Was the Order — A constraint-aware GPU scheduler from Dharma-AI beats FIFO allocation by up to 33 percentage points in utilization and 105% in priority-weighted throughput across seven benchmark scenarios — same hardware, same workloads, different order of allocation decisions.
- State of Open Models: Summer 2026 Observations — Hugging Face's Summer 2026 open model report documents a Chinese lab surge at the frontier: Chinese labs released models up to 2.78 trillion parameters every month while US-trained open models (excluding NVIDIA) peaked at 130B, and AMD and NVIDIA now publish more open model repositories than any AI research lab.
- Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets — AWS's open-source Strands Robots SDK (Apache 2.0) now supports a full continuous robotics training loop via Hugging Face Storage Buckets — record demonstrations, train on the Hub, and redeploy to hardware in a single agent workflow without redundant full-dataset transfers.