How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

| Source: Hugging Face Blog

Tags: Hugging Face, Papers with Code, pgvector, hybrid-search, embeddings, RAG, semantic-search

Hugging Face details the hybrid search architecture powering Papers with Code — combining PostgreSQL full-text search with pgvector semantic embeddings via reciprocal rank fusion, indexing 110,000+ papers with a split pipeline across HF Jobs, Buckets, and Inference Endpoints.

Details

Hugging Face has published an engineering breakdown of the search system behind the revived Papers with Code platform. The architecture is a hybrid retrieval system combining two complementary approaches: PostgreSQL full-text search for exact keyword and title matching, and pgvector semantic embeddings for fuzzy, meaning-based recall. The reciprocal rank fusion (RRF) algorithm merges the two result sets — consistently outperforming either method alone, a pattern backed by prior work at ML6 and Microsoft research. The system addresses real search challenges in academic paper retrieval: tolerating typos, understanding navigational intent ('the original BERT paper'), and handling semantic queries that don't share exact words with paper titles. Rerankers (cross-encoders) were considered but skipped due to added latency overhead. Three Hugging Face services power the embedding pipeline. HF Jobs provides burstable GPU compute for batch-embedding the 110,000+ paper corpus. HF Storage Buckets acts as the durable handoff layer between the database, batch jobs, and experiments. HF Inference Endpoints serves low-latency embeddings for live user queries and incremental updates as new papers arrive from arXiv and Daily Papers. The system also exposes a CLI command (pwc search) accessible to AI agents via a Skill, pointing toward AI-native research workflows. This is a useful reference architecture for teams building production hybrid search systems.