Copied from upstream: SWE-agent/mini-swe-agent#714
Original author: @RobinChiu
Originally created: 2026-01-26
Describe the issue
I tried running the mini-extra swebench mark in the gpt-oss-120b model. The generated preds.json file was scored using SWE-bench, and it contained a large number of empty patches.
Instances with empty patches: 157
-
Is this normal?
-
Should I repeatedly run these empty patch tasks until normal patches are generated?
The leaderboard score for mini-swe-agent gpt-oss-120b is 26%, but the test results are only around 15%, a significant difference from the official score.
Describe the issue
I tried running the mini-extra swebench mark in the gpt-oss-120b model. The generated preds.json file was scored using SWE-bench, and it contained a large number of empty patches.
Instances with empty patches: 157
Is this normal?
Should I repeatedly run these empty patch tasks until normal patches are generated?
The leaderboard score for mini-swe-agent gpt-oss-120b is 26%, but the test results are only around 15%, a significant difference from the official score.