Skip to content

Gold evaluation can undo its own repair when a patched source file appears in PASS_TO_PASS #245

Description

@mayiqun1213-sys

Summary

At SWE-smith commit 9b74ac08118a85c39c356802f7961893af73e07f, the official gold evaluation reports the following instance as unresolved:

r1chardj0n3s__parse.30da9e4f.func_basic__j0z3z0co

The reversed bug patch is applied successfully to parse.py, but the evaluator then restores parse.py because it also appears in PASS_TO_PASS. This discards the gold repair before pytest runs.

This looks like an incompatibility between this task and the test-file anti-cheating restore policy. I am not assuming that test restoration should simply be disabled.

Minimal reproduction

From a checkout of 9b74ac08118a85c39c356802f7961893af73e07f with Docker available:

python -m swesmith.harness.eval \
  -p gold \
  --run_id repro_gold_parse \
  -w 1 \
  -i r1chardj0n3s__parse.30da9e4f.func_basic__j0z3z0co

Observed summary:

Using gold predictions for eval (ignoring `predictions_path` argument)
Resolved 0/1 instances.

The per-instance log contains:

Applied patch parse.py cleanly.
Reverted changes to test files in container: tests/test_parse.py README.rst parse.py tests/test_bugs.py tests/test_findall.py tests/test_parse.py tests/test_parsetype.py tests/test_pattern.py tests/test_result.py tests/test_search.py

The test result then includes:

FAILED tests/test_parse.py::test_parser_format - AttributeError: 'int' object...
=================== 1 failed, 97 passed, 1 skipped ====================

Expected: the gold evaluation resolves the instance (1/1).

Actual: the gold evaluation leaves it unresolved (0/1).

Diagnostic

PythonProfile.get_test_files() maps both FAIL_TO_PASS and PASS_TO_PASS identifiers to file paths. For this instance it returns:

FAIL_TO_PASS files:
['tests/test_parse.py']

PASS_TO_PASS files:
['README.rst', 'parse.py', 'tests/test_bugs.py', 'tests/test_findall.py',
 'tests/test_parse.py', 'tests/test_parsetype.py', 'tests/test_pattern.py',
 'tests/test_result.py', 'tests/test_search.py']

The instance's bug patch modifies parse.py. In the gold path:

  1. eval.py uses the dataset patch as the gold prediction.
  2. _apply_patch(..., is_gold=True) reverses that patch, repairing parse.py.
  3. run_patch_in_container() calls get_test_files(instance).
  4. git checkout -- ... parse.py ... restores the buggy version of parse.py.
  5. Pytest therefore evaluates the task after the repair has been undone.

A manual check inside the task image confirms that the target test fails at the bug stage and passes immediately after reversing the bug patch, before the test-file restore.

Relevant code:

  • swesmith/profiles/python.py: PythonProfile.get_test_files()
  • swesmith/harness/utils.py: the git checkout -- {test_files} block after patch application

Question

What is the intended handling for instances where a bug patch touches a file that is also identified as an F2P/P2P test file (for example, a production module containing doctests)?

Possible directions seem to be:

  • reject/filter these instances during task generation or validation; or
  • refine the restore behavior so it preserves the submitted/gold source change without exposing hidden tests.

I can provide the complete per-instance log or test a proposed fix if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions