Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program Repair

| Source: arXiv AI

Tags: program repair, SWE-bench, LLM agents, software engineering, code audit

THEMIS makes LLM-based program repair inspectable by externalizing the requirement-to-repair reasoning chain — resolving 19 of 100 SWE-bench Lite cases vs 9 for direct repair while retaining auditable records that link issue requirements to code patches.

Details

Repository-level bug repair with LLMs typically produces correct or incorrect patches with no record of how issue requirements were translated into code changes. THEMIS addresses this by externalizing the reasoning chain through a stage-aware workflow: semantic interpretation of the issue, a runtime requirement-code graph, developer guidance derived from the graph, retained repair rationale, and post-edit audit records. On 300 SWE-bench Lite cases, THEMIS retains a complete developer rationale for 288 cases (96%) and preserves a complete auditable field set for 214 cases (71.3%) that connects all stages from issue to patch. Target symbols recur in 62.6% of developer rationales and 62.8% of patches, rising to 75.8% when related symbols are included — evidence that the requirement-code graph is capturing relevant structure. In a 100-case comparison, THEMIS resolves 19 cases versus 9 for direct same-input repair. The paper appropriately hedges this as workflow-level evidence rather than a causal effect of any single component, since other factors also differ between conditions. Accepted at the REALM Workshop at EMNLP 2026. Code is open-sourced.