Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents

| Source: arXiv AI

Tags: coding agents, GEPA, repository knowledge, SKILL files, agent evaluation, Kotlin

Automatically synthesized .md knowledge files (SKILLs) for coding agents improve performance by only 4.9pp on average with GEPA versus 0.1pp with SkillOpt — and at single-repository scale these gains cannot be statistically separated from the agent's run-to-run variance.

Details

As coding agents increasingly read repository context from .md SKILL files versioned alongside code, researchers tested whether automatically optimizing these documents actually helps agents perform better. Using three Kotlin repositories, they compared GEPA and SkillOpt optimizers against a no-document baseline. The evaluation setup is more realistic than prior work: tasks are mined from merged pull requests reverted to a frozen base commit, scored by whether an agent does better with the document than without it. This avoids the saturation problem where capable agents solve synthetic tasks with no document at all. Results are sobering: GEPA raised agent performance by 4.9pp on average while SkillOpt managed just 0.1pp above the seed document. More critically, at single-repository scale these gains cannot be statistically separated from run-to-run agent variance — validating the result would require pooled data across more repositories than any single project's PR history can supply. The qualitative finding is more encouraging: a maintainer of one repository found the auto-generated documents captured knowledge only acquired through sustained project work. The gap between 'reads like useful documentation' and 'measurably improves agent performance' is instructive for teams investing in SKILL generation infrastructure.