SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?
Abstract
The study introduces a benchmark for evaluating autonomous software migration by coding agents, finding that current models rarely complete migrations correctly.
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs (5.4%) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores 47.0/100. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, 58% reach 99% of the fixed checks, yet only 26% reach 100%. Agent capability differs across migration categories: agents score 31.4 on build toolchain rewrites but only 5.6 on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
Community
The ultimate test for coding agents isn't local editing — it's whole-repo evolution, and right now, the survival rate is 5.4%.
Today we’re releasing SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migration.
Coding agents are getting very good at fixing bugs.
But can they refactor an entire system, C → Rust, Maven → Gradle, POSIX → WebAssembly?
We built 20 real migrations across projects, including SQLite, zlib, libsodium, and GraphHopper.
520 runs. Only 28 survived all 3 stages. 13/20 tasks were solved by nobody.
System-scale migration is still wide open.
🏠 Homepage
https://lab.einsia.ai/swe-refactor-bench/
🏆 Leaderboard
https://lab.einsia.ai/swe-refactor-bench/leaderboard/
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks (2026)
- RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue (2026)
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction (2026)
- Vero: Can AI Agents Build Formally Verified Software Repositories? (2026)
- ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance (2026)
- Specification Grounding Drives Test Effectiveness for LLM Code (2026)
- Evaluating Agentic Code Repair Capabilities in Distributed Systems (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.23564 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper