From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon

| Source: BAIR Blog (Berkeley AI Research)

Tags: MLX, Apple Silicon, CUDA, K-Search, GPU kernels, Mamba, Berkeley, attention

Berkeley researchers extended K-Search, an AI-powered kernel optimizer, with a CUDA-to-MLX translation layer, achieving 0.97x native MLX Attention speed and a 20x Mamba SSM prefill speedup over the community mlx-lm implementation on Apple Silicon.

Details

The CUDA ecosystem holds decades of hand-tuned GPU kernel expertise. Berkeley Sky Lab''s K-Search framework uses AI-driven evolutionary search to optimize GPU kernels. Researchers have now added a structured CUDA-to-MLX translation layer that ports CUDA optimization knowledge to Apple Silicon''s MLX framework rather than rebuilding from scratch. The key insight is a CUDA-to-MLX optimization translation map: CUDA patterns are adapted into architecture-native MLX strategies rather than copied instruction-for-instruction. Results on Apple M-series chips: the approach matches the native MLX Attention kernel at 0.97x speed (essentially parity) and delivers a 20x prefill speedup over the community mlx-lm Mamba SSM implementation. The researchers frame the technique as general and applicable to any ecosystem where CUDA expertise can transfer to new hardware. For practitioners running models on Mac hardware, this addresses a documented gap: MLX runs models correctly but often leaves performance on the table because critical kernels have not been hand-tuned.