Skip to content

Cannot reproduce results with gpt-oss-120b #10

Description

@yishangupenn

Copied from upstream: SWE-agent/mini-swe-agent#714
Original author: @RobinChiu
Originally created: 2026-01-26


Describe the issue

I tried running the mini-extra swebench mark in the gpt-oss-120b model. The generated preds.json file was scored using SWE-bench, and it contained a large number of empty patches.
Instances with empty patches: 157

  1. Is this normal?

  2. Should I repeatedly run these empty patch tasks until normal patches are generated?

The leaderboard score for mini-swe-agent gpt-oss-120b is 26%, but the test results are only around 15%, a significant difference from the official score.

Metadata

Metadata

Assignees

No one assigned

    Labels

    questionFurther information is requested

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions