SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing
| Source: arXiv AI
Tags: LLM privacy, differential privacy, trusted execution environment, Intel TDX, Llama, inference security
SpliTEE protects LLM user prompts by splitting inference between a CPU Trusted Execution Environment and an untrusted GPU—using differential privacy on intermediate representations—running nearly 2x faster than fully CPU-based TEE inference while blocking prompt reconstruction attacks that can otherwise recover inputs with ~80% accuracy.
Details
Running LLMs inside a Trusted Execution Environment (TEE) protects user prompts from service providers but is dramatically slower than GPU inference. The standard workaround—Slalom-style split inference with encryption—avoids the slowdown but forces quantization. SpliTEE, from researchers at Macquarie University and Data61, takes a different path: send intermediate representations to the GPU protected by differential privacy rather than encryption. The floating-point domain is preserved, and the privacy guarantee is formal. Before proposing the fix, the paper establishes why it is necessary: a prompt-reconstruction attack can recover prompts from intermediate representations with nearly 80% accuracy, making unprotected split inference a concrete privacy risk. The system is implemented on Intel TDX with Llama-3.2-3B and Qwen3-4B. Compared to fully CPU-based TDX inference, SpliTEE is nearly 2x faster; compared to encryption-based Slalom, it is 5–15 seconds faster per query. The privacy-utility tradeoff is controlled by epsilon—the paper derives bounds on floating-point error from noise as a function of epsilon. For enterprises running private or regulated LLM workloads, this architecture enables GPU-accelerated inference without trusting the GPU operator.