RefactorPlatform: An Open-Source Harness for Controlled Evaluation of Repository-Scale Refactoring Agents
| Source: arXiv AI
Tags: code refactoring, coding agents, LLM evaluation, software engineering, benchmarks, open-source
RefactorPlatform is an open-source evaluation harness for repository-scale refactoring agents that independently varies model backbone, tool access, and agent architecture — enabling controlled comparison of design choices without the confounds present in existing coding agent benchmarks.
Details
Repository-scale refactoring tasks require a coding agent to propagate a single change (rename a function, change a type, restructure a module) across many interdependent files without altering program behavior. Existing coding agent benchmarks — SWE-bench and similar — mix environment setup, issue understanding, implementation, and testing in ways that make it difficult to isolate which design choice determines success on the refactoring subtask specifically. RefactorPlatform provides a controlled harness by holding the environment fixed and varying design axes independently: model backbone (accessed via OpenRouter and GitHub Copilot API), tool access (grep, search, apply-patch, execute), agent architecture (single-model, multi-agent, chain-of-thought), and context management strategy. This allows researchers and practitioners to measure the contribution of each design choice to refactoring success rate. The harness is open-source and designed to accept new refactoring tasks from real repositories, with automated correctness verification via test suite execution and semantic diff analysis. For teams building coding agents, RefactorPlatform provides a reproducible way to evaluate whether architecture choices that improve performance on SWE-bench generalize to the multi-file propagation challenge of repository-scale refactoring.