DeepSeek Signals Next-Gen R2 Model, Unveils Novel Approach to Scaling Inference with SPCT

| Source: Synced Review

Tags: DeepSeek AI, general reward models, R2 model, scalability, inference, SPCT

DeepSeek signals its next-generation R2 model is in development and releases SPCT — a novel inference scaling technique for general reward models that improves test-time compute efficiency without requiring model retraining.

Details

DeepSeek, the Chinese AI lab that disrupted the market with R1, is signaling its next generation R2 model while simultaneously publishing research on SPCT (Scalable Process Credit Training), a technique for scaling inference compute in general reward models (GRMs). SPCT addresses a known limitation: most inference scaling techniques (like best-of-N sampling or beam search) work well for specific reward models tuned to narrow tasks but degrade when applied to general-purpose reward models. SPCT provides a method to improve GRM effectiveness at test time without expensive retraining. The combination of R2 signal plus novel inference research positions DeepSeek as continuing to push efficiency-first AI development. If R2 follows the pattern of R1 — competitive performance at dramatically lower inference cost — it would again pressure Western labs' pricing and positioning. The specific architecture, training data, or capability claims for R2 were not available in this report.