Efficiency Hallucination: Formalizing and Measuring Behavioral Calibration in LLM-Based Code Optimization

| Source: arXiv AI

Tags: code optimization, LLM reliability, efficiency hallucination, GPT, Claude, Gemini, EffiBench

LLMs exhibit a 100% over-edit rate on already-optimized code across all tested models (GPT, Claude, Gemini families). A training-free classification penalty guardrail raises correct abstention from 0% to 44.4% while maintaining a 100% edit rate on genuinely sub-optimal code.

Details

When asked to optimize code that is already optimal, LLMs universally make modifications with unsubstantiated performance claims—a phenomenon the authors term 'efficiency hallucination.' The root cause is an 'Evaluation Trap': binary benchmarks reward modification over restraint, so models are incentivized to always change something. Evaluated across 180 optimization runs on nine models spanning GPT, Claude, and Gemini families using EffiBench, the baseline over-edit rate is 100% across all models.\n\nThe proposed guardrail applies classification penalty methods to distinguish truly sub-optimal code from optimal code. It raises correct abstention from 0% to 44.4% while preserving a 100% edit rate on sub-optimal code with zero false abstentions. Model calibration varies considerably: GPT-5.4 Mini approaches near-perfect abstention, while complex code is harder for all models to correctly identify as already optimal.\n\nFor practitioners deploying LLMs in CI/CD code optimization pipelines, this is a concrete reliability failure to account for: LLMs will routinely 'improve' working, optimized code in ways that may introduce subtle regressions. The training-free guardrail can be deployed without fine-tuning.