SWE-bench · 10 rubrics · first 10 of the 1000
HF EdwardoSunny/swe-rubrics-mined-10 · local data/libraries/swe-rubrics-mined-10.json
code: a class/function renders a container node by calling a recursive child-render helper (e.g. self.render_children(...), self.render(child), "".join(map(self.visit, node.children))) and then wraps or suffixes the result with fixed whitespace/separator text..rstrip(), .rstrip('\n'), .strip() or a regex that collapses trailing blank lines, and then appends its own terminator. The child renderers already guarantee their own trailing separator, so the strip plus the new terminator produces a different number of blank lines than every sibling handler emits, silently changing block separation for all downstream text.... + '\n\n' and none of them strip the value returned by a child-render call. [reads: code]token['raw'], token.attrs['text'], a source slice) before indenting or wrapping it, or one that strips child output and returns it with no terminator because the caller supplies the separator. Neither double-normalizes.assertEqual on full output) fail with a whitespace-only diff; no exception is raised, so the failure surfaces only as a lower test-pass score. This accounts for the block-separation portion of the gap; any additional conditional prefix logic the handler keeps or drops relative to the reference accounts for the rest.text = indent(self.render_children(token, state).rstrip('\n'), ' ') followed by return text + '\n\n', where sibling handlers in the same renderer return ... + '\n\n' without stripping child output; the accepted solution kept indent(self.render_children(token, state), ' ') + '\n\n' with no strip.e92902aa3f9b · mined from swesmith/lepture__mistune.bf54ef67 lepture__mistune.bf54ef67.lm_rewrite__eb4ybjuf{
"predicate": "a named unit test will fail",
"code_requirement": "1. Find the handler method that builds its result from a recursive child-render call, and note every string transformation applied to that call's return value before it is returned. [reads: code]",
"prediction": "The rendered document contains fewer blank lines after this construct than the reference output. Exact-string comparison tests (fixture/round-trip renderer tests using `assertEqual` on full output) fail with a whitespace-only diff; no exception is raised, so the failure surfaces only as a lower test-pass score. This accounts for the block-separation portion of the gap; any additional conditional prefix logic the handler keeps or drops relative to the reference accounts for the rest."
}### Stripping trailing newlines from recursively rendered child output before appending a fixed separator
- **Applies when**: `code`: a class/function renders a container node by calling a recursive child-render helper (e.g. `self.render_children(...)`, `self.render(child)`, `"".join(map(self.visit, node.children))`) and then wraps or suffixes the result with fixed whitespace/separator text.
- **Pattern**: The container handler normalizes the child-rendered string with `.rstrip()`, `.rstrip('\n')`, `.strip()` or a regex that collapses trailing blank lines, and then appends its own terminator. The child renderers already guarantee their own trailing separator, so the strip plus the new terminator produces a different number of blank lines than every sibling handler emits, silently changing block separation for all downstream text.
- **Detection procedure**:
1. Find the handler method that builds its result from a recursive child-render call, and note every string transformation applied to that call's return value before it is returned. [reads: code]
2. Read the sibling handler methods in the same class (the ones the task does not ask you to change) and record the terminator convention they follow — e.g. most return `... + '\n\n'` and none of them strip the value returned by a child-render call. [reads: code]
3. Fire if the edited handler is the only one that applies a trailing-whitespace strip to composed child output *and* still appends the class's standard terminator; i.e. the same characters are both removed and re-added at a different count. [reads: code]
- **Counter-example**: A handler that strips a *leaf* value taken directly from the token/AST (`token['raw']`, `token.attrs['text']`, a source slice) before indenting or wrapping it, or one that strips child output and returns it with **no** terminator because the caller supplies the separator. Neither double-normalizes.
- **Discriminator**: The wrong case strips the output of a recursive render call whose producers already append the class-wide terminator, and then appends that terminator again; the safe case either strips raw leaf text (which carries no renderer contract) or strips without re-adding a separator.
- **Consequence**: The rendered document contains fewer blank lines after this construct than the reference output. Exact-string comparison tests (fixture/round-trip renderer tests using `assertEqual` on full output) fail with a whitespace-only diff; no exception is raised, so the failure surfaces only as a lower test-pass score. This accounts for the block-separation portion of the gap; any additional conditional prefix logic the handler keeps or drops relative to the reference accounts for the rest.
- **Evidence**: `text = indent(self.render_children(token, state).rstrip('\n'), ' ')` followed by `return text + '\n\n'`, where sibling handlers in the same renderer return `... + '\n\n'` without stripping child output; the accepted solution kept `indent(self.render_children(token, state), ' ') + '\n\n'` with no strip.code: the change adds a new Python file to a repository that is exercised by pytesttest_.py / _test.py) and placed where collection reaches it, but its body is bare module-level code with side effects instead of test functions — so the code runs during collection rather than as a test..py files whose basename matches test_.py or _test.py. [reads: code]pytest is in the environment and note where the project's real tests live in the repo tree (e.g. ./tests/), i.e. that the new file sits outside that directory, typically at the root. [reads: static facts — python packages, repo tree]print) with no def test_ / class Test definitions, so the work happens at module import. [reads: code]test_*.py placed inside the project's test package that defines def test_...() functions and confines all setup to fixtures or function bodies; or a scratch script named repro.py / debug_x.py that pytest never collects.ERROR ... test_issue.py, surfacing as AttributeError, TypeError, ImportError, or the library's own exception) and the session exits non-zero even though the real tests pass; if it does not raise, it contributes a collected-but-empty module and stray stdout that pollutes the graded test report.test_issue.py was added whose body immediately runs OpcPackage(), patches package.iter_parts with a Mock, calls package.next_partname(...) and prints — no test function anywhere in the file.044ba6438bb8 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "test or grader output will report failure",
"code_requirement": "1. Find newly added `.py` files whose basename matches `test_*.py` or `*_test.py`. [reads: code]",
"prediction": "If the module body raises, pytest reports a collection error (`ERROR ... test_issue.py`, surfacing as `AttributeError`, `TypeError`, `ImportError`, or the library's own exception) and the session exits non-zero even though the real tests pass; if it does not raise, it contributes a collected-but-empty module and stray stdout that pollutes the graded test report."
}### Scratch file named `test_*.py` at repo root executing at import time - **Applies when**: `code`: the change adds a new Python file to a repository that is exercised by pytest - **Pattern**: A debug/scratch file is given a name matching pytest's default collection pattern (`test_*.py` / `*_test.py`) and placed where collection reaches it, but its body is bare module-level code with side effects instead of test functions — so the code runs during collection rather than as a test. - **Detection procedure**: 1. Find newly added `.py` files whose basename matches `test_*.py` or `*_test.py`. [reads: code] 2. Check from the static facts that `pytest` is in the environment and note where the project's real tests live in the repo tree (e.g. `./tests/`), i.e. that the new file sits outside that directory, typically at the root. [reads: static facts — python packages, repo tree] 3. Check whether the file's body consists of top-level executable statements (object construction, method calls, `print`) with **no** `def test_*` / `class Test*` definitions, so the work happens at module import. [reads: code] - **Counter-example**: A new `test_*.py` placed inside the project's test package that defines `def test_...()` functions and confines all setup to fixtures or function bodies; or a scratch script named `repro.py` / `debug_x.py` that pytest never collects. - **Discriminator**: The failing case is both name-matched for collection *and* does all its work at module scope with no test functions; the safe case either isn't name-matched or keeps side effects inside test functions. - **Consequence**: If the module body raises, pytest reports a collection error (`ERROR ... test_issue.py`, surfacing as `AttributeError`, `TypeError`, `ImportError`, or the library's own exception) and the session exits non-zero even though the real tests pass; if it does not raise, it contributes a collected-but-empty module and stray stdout that pollutes the graded test report. - **Evidence**: A root-level `test_issue.py` was added whose body immediately runs `OpcPackage()`, patches `package.iter_parts` with a `Mock`, calls `package.next_partname(...)` and prints — no test function anywhere in the file.
code: a function converts a raw value into a wrapper/domain object (constructor, factory, Path(), np.array(), Decimal(), a class from the same package) and the project ships a unit-test suite that mirrors the module being editedreturn — instead of constructing once into a local. Externally the value is right, but the collaborator's call count doubles, which breaks unit tests that patch that constructor and assert on how it was called.str/int) and note the argument expression of each. [reads: code]if/while condition and one in the return immediately below it. [reads: code]tests/<subpkg>/test_<module>.py for src/<pkg>/<subpkg>/<module>.py), i.e. this collaborator is likely mocked and its calls asserted. [reads: code; static facts — repo tree]candidate = Wrapper(template % n) assigned once, then if candidate not in existing: return candidate — the constructor runs once per iteration and once total on the returning path; also safe is a loop that constructs a genuinely different value each iteration and returns it without re-constructing.AssertionError from assert_called_once_with / assert_called_once ("Expected 'X' to be called once. Called 2 times"), even though the returned value is correct; the graded test suite reports the fix as failing.if Wrapper(candidate) not in names: return Wrapper(candidate) (constructing the wrapper in both the guard and the return); the suite failed with AssertionError: Expected 'PackURI' to be called once. Called 2 times.71e065863350 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "test or grader output will report failure",
"code_requirement": "1. In the function the change touches, list every call to a class/factory name that is imported from the same package (not a builtin like `str`/`int`) and note the argument expression of each. [reads: code]",
"prediction": "Mock-based unit tests that patch the constructor fail with `AssertionError` from `assert_called_once_with` / `assert_called_once` (\"Expected 'X' to be called once. Called 2 times\"), even though the returned value is correct; the graded test suite reports the fix as failing."
}### Duplicate construction of the same collaborator object on one code path
- **Applies when**: `code`: a function converts a raw value into a wrapper/domain object (constructor, factory, `Path()`, `np.array()`, `Decimal()`, a class from the same package) and the project ships a unit-test suite that mirrors the module being edited
- **Pattern**: The edited code calls the same constructor/factory twice with the same argument on a single execution path — once inside a membership/equality guard and again in the `return` — instead of constructing once into a local. Externally the value is right, but the collaborator's call count doubles, which breaks unit tests that patch that constructor and assert on how it was called.
- **Detection procedure**:
1. In the function the change touches, list every call to a class/factory name that is imported from the same package (not a builtin like `str`/`int`) and note the argument expression of each. [reads: code]
2. Check whether two such calls use the identical argument expression on a path that can execute both — typically one inside an `if`/`while` condition and one in the `return` immediately below it. [reads: code]
3. Confirm the constructed object is not bound to a local variable and reused; and confirm the static facts show a test package mirroring the edited module's path (e.g. `tests/<subpkg>/test_<module>.py` for `src/<pkg>/<subpkg>/<module>.py`), i.e. this collaborator is likely mocked and its calls asserted. [reads: code; static facts — repo tree]
- **Counter-example**: `candidate = Wrapper(template % n)` assigned once, then `if candidate not in existing: return candidate` — the constructor runs once per iteration and once total on the returning path; also safe is a loop that constructs a genuinely different value each iteration and returns it without re-constructing.
- **Discriminator**: The failing case passes the *same* argument expression to the *same* constructor twice with no intervening state change, so the returning path invokes it ≥2 times; the safe case invokes it once and reuses the binding.
- **Consequence**: Mock-based unit tests that patch the constructor fail with `AssertionError` from `assert_called_once_with` / `assert_called_once` ("Expected 'X' to be called once. Called 2 times"), even though the returned value is correct; the graded test suite reports the fix as failing.
- **Evidence**: The patch replaced a raw-value membership test with `if Wrapper(candidate) not in names: return Wrapper(candidate)` (constructing the wrapper in both the guard and the return); the suite failed with `AssertionError: Expected 'PackURI' to be called once. Called 2 times.`code: the change rewrites the body of an existing library function (rather than adding new code) in a repository whose static facts show a mirrored unit-test package for the edited moduleNone — beyond the minimum needed to fix the stated defect. Existing tests pin those details, so the broader-than-necessary rewrite fails tests unrelated to the bug.self.iter_parts() vs self.parts) and the same single collaborator invocation, and therefore leaves interaction-level assertions intact.AssertionError on mock call assertions or on patched-attribute expectations, while the functional bug itself is fixed — the submission is scored as failing. This accounts for the interaction-level portion of the failure; the value-level logic may well be correct.for into an unbounded while True, and added a second collaborator construction; the only test failure came from the extra collaborator construction, not from the value returned.0282acfa6f79 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "test or grader output will report failure",
"code_requirement": "1. Read the task statement for the specific defective behaviour that must change (the wrong output for a given input). [reads: task]",
"prediction": "Pre-existing unit tests fail with `AssertionError` on mock call assertions or on patched-attribute expectations, while the functional bug itself is fixed \u2014 the submission is scored as failing. This accounts for the interaction-level portion of the failure; the value-level logic may well be correct."
}### Behaviour-changing rewrite of a helper whose contract is pinned by existing unit tests - **Applies when**: `code`: the change rewrites the body of an existing library function (rather than adding new code) in a repository whose static facts show a mirrored unit-test package for the edited module - **Pattern**: The rewrite alters observable interaction details of the function — which internal accessor it calls, how many times it calls a collaborator, whether it can return `None` — beyond the minimum needed to fix the stated defect. Existing tests pin those details, so the broader-than-necessary rewrite fails tests unrelated to the bug. - **Detection procedure**: 1. Read the task statement for the specific defective behaviour that must change (the wrong output for a given input). [reads: task] 2. Diff the rewritten function against the code it replaces and list each behavioural difference: different internal method/property used to obtain the data, extra or fewer calls to imported collaborators, changed loop bounds, changed return-on-exhaustion behaviour. [reads: code] 3. Flag when at least one listed difference is not required to produce the corrected output — i.e. the corrected output is already achieved by the other differences — and it touches a call to a name imported at module top level (the kind a test patches). [reads: code] - **Counter-example**: A rewrite that changes only the search bound / comparison that produced the wrong result, keeps the same accessor (`self.iter_parts()` vs `self.parts`) and the same single collaborator invocation, and therefore leaves interaction-level assertions intact. - **Discriminator**: The failing case contains at least one *gratuitous* interaction change (extra collaborator call or swapped internal accessor) alongside the necessary logic fix; the safe case's diff is confined to the logic that produced the wrong value. - **Consequence**: Pre-existing unit tests fail with `AssertionError` on mock call assertions or on patched-attribute expectations, while the functional bug itself is fixed — the submission is scored as failing. This accounts for the interaction-level portion of the failure; the value-level logic may well be correct. - **Evidence**: The rewrite simultaneously swapped the internal iteration accessor, converted a bounded `for` into an unbounded `while True`, and added a second collaborator construction; the only test failure came from the extra collaborator construction, not from the value returned.
code: a function searches for the first unused name/key/identifier by generating candidates from a template or counter and testing them against a collection of already-used values, and a class or factory imported at module level is used to wrap the candidate.candidate = template % n, f"{base}{i}", key + str(i)) and tests them for prior use with in, ==, or a dict/set lookup. [reads: code]Wrapper(candidate) not in used / stores Wrapper(candidate) before the test, where Wrapper is a name imported at module scope, and the elements of used come from somewhere else (raw attribute values, plain strings). Do not fire if the raw candidate is compared and the wrapper is applied only on the return. [reads: code]for n in count(1): cand = template % n … if cand not in used: return PackURI(cand) — same wrapper class, same search, but the wrapper is constructed exactly once, on the returned value, and never participates in the comparison.mock.patch the wrapper name in that module fail with AssertionError: expected call not found from assert_called_once_with (extra/earlier calls with the wrong argument), and the function returns the first candidate instead of the first unused one because the stub's constant return value is never found in the used-set; if the loop is an unbounded while True, the opposite stubbing (constant that is in the set) hangs the test run instead. This mechanism accounts for the entire observed test failure here.if PackURI(candidate_partname) not in partnames: return candidate_packuri; the patched-constructor test reported Expected: PackURI('/foo/bar/baz2.xml') Actual: PackURI('/foo/bar/baz1.xml') and failed.b3c46989e3fd · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "test or grader output will report failure",
"code_requirement": "1. Locate the loop that builds candidates (e.g. `candidate = template % n`, `f\"{base}{i}\"`, `key + str(i)`) and tests them for prior use with `in`, `==`, or a dict/set lookup. [reads: code]",
"prediction": "unit tests that `mock.patch` the wrapper name in that module fail with `AssertionError: expected call not found` from `assert_called_once_with` (extra/earlier calls with the wrong argument), and the function returns the *first* candidate instead of the first *unused* one because the stub's constant return value is never found in the used-set; if the loop is an unbounded `while True`, the opposite stubbing (constant that is in the set) hangs the test run instead. This mechanism accounts for the entire observed test failure here."
}### Constructor/wrapper call moved inside the search loop and into the membership test
- **Applies when**: `code`: a function searches for the first unused name/key/identifier by generating candidates from a template or counter and testing them against a collection of already-used values, and a class or factory imported at module level is used to wrap the candidate.
- **Pattern**: The candidate value is passed through a wrapper constructor *before* the equality/membership check, so (a) the constructor is invoked once per loop iteration instead of once on the value actually returned, and (b) the object compared against the collection is not of the same provenance as the collection's elements. Interaction-based tests that patch that constructor then see the wrong call count/arguments, and with the constructor stubbed the comparison never matches, so the function returns the first candidate.
- **Detection procedure**:
1. Locate the loop that builds candidates (e.g. `candidate = template % n`, `f"{base}{i}"`, `key + str(i)`) and tests them for prior use with `in`, `==`, or a dict/set lookup. [reads: code]
2. Read the expression that builds the collection of used values (attribute reads over a collection of objects, dict keys, a listing) and note whether those elements were produced by the same wrapper class the candidate is passed through. [reads: code]
3. Fire if the code writes the membership test as `Wrapper(candidate) not in used` / stores `Wrapper(candidate)` before the test, where `Wrapper` is a name imported at module scope, and the elements of `used` come from somewhere else (raw attribute values, plain strings). Do not fire if the raw candidate is compared and the wrapper is applied only on the `return`. [reads: code]
- **Counter-example**: `for n in count(1): cand = template % n` … `if cand not in used: return PackURI(cand)` — same wrapper class, same search, but the wrapper is constructed exactly once, on the returned value, and never participates in the comparison.
- **Discriminator**: the number of constructor invocations scales with loop iterations and the constructed object is an operand of the equality/membership test, versus exactly one invocation outside the test on the returned value.
- **Consequence**: unit tests that `mock.patch` the wrapper name in that module fail with `AssertionError: expected call not found` from `assert_called_once_with` (extra/earlier calls with the wrong argument), and the function returns the *first* candidate instead of the first *unused* one because the stub's constant return value is never found in the used-set; if the loop is an unbounded `while True`, the opposite stubbing (constant that is in the set) hangs the test run instead. This mechanism accounts for the entire observed test failure here.
- **Evidence**: the search loop was rewritten from comparing the raw candidate to `if PackURI(candidate_partname) not in partnames: return candidate_packuri`; the patched-constructor test reported `Expected: PackURI('/foo/bar/baz2.xml') Actual: PackURI('/foo/bar/baz1.xml')` and failed.code: the change set includes both a modification to a source file and a separate file containing a unified diff (.patch, .diff, or a file whose text begins with --- a/ / +++ b/).patch/.diff or whose first lines are --- a/<path> / +++ b/<path>, and read the path named in its headers. [reads: code]+ line of the patch hunk with the corresponding region of the edited source. Fires if any added line differs — a different expression or call wrapper, an added/removed blank line between definitions, differing indentation or trailing whitespace — or if the hunk's context lines no longer match the edited file. [reads: code].patch file living under a fixtures/test-data directory that is input data rather than a description of this change.if Wrapper(x) not in s: while the source says if x not in s:, or the patch deletes a separator blank line the source keeps); the safe case has none, so applying the patch is a no-op.git apply/patch aborts with "patch does not apply" / "Hunk #1 FAILED" (nonzero exit), or, if it applies, the resulting code differs from the code the tests passed against — the behavior verified is not the behavior shipped. Where the patch also drops a blank line between top-level definitions or introduces trailing whitespace, lint gates (ruff/flake8 E301/W291) fail.fix.patch restated the source edit but with an extra type-wrapping call in the membership test and with the blank line before the following @classmethod deleted; the actual source file contained neither change, so the two representations of the same fix disagreed.daf034bd2ebc · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "test or grader output will report failure",
"code_requirement": "1. Locate any added file whose name ends in `.patch`/`.diff` or whose first lines are `--- a/<path>` / `+++ b/<path>`, and read the path named in its headers. [reads: code]",
"prediction": "If any harness or reviewer applies the artifact, `git apply`/`patch` aborts with \"patch does not apply\" / \"Hunk #1 FAILED\" (nonzero exit), or, if it applies, the resulting code differs from the code the tests passed against \u2014 the behavior verified is not the behavior shipped. Where the patch also drops a blank line between top-level definitions or introduces trailing whitespace, lint gates (ruff/flake8 E301/W291) fail."
}### Patch/diff artifact committed alongside the real edit and disagreeing with it - **Applies when**: `code`: the change set includes both a modification to a source file and a separate file containing a unified diff (`*.patch`, `*.diff`, or a file whose text begins with `--- a/` / `+++ b/`) - **Pattern**: The program records its intended change twice — once by editing the source and once as a checked-in patch file — and the two copies are not identical, so the artifact describing the fix does not match the fix that was actually applied and tested. - **Detection procedure**: 1. Locate any added file whose name ends in `.patch`/`.diff` or whose first lines are `--- a/<path>` / `+++ b/<path>`, and read the path named in its headers. [reads: code] 2. Confirm that same path is also directly modified by the program (it appears as an edited source file in the change set / repo tree). [reads: code, and static facts — repo tree for the source path] 3. Line-by-line, compare every `+` line of the patch hunk with the corresponding region of the edited source. Fires if any added line differs — a different expression or call wrapper, an added/removed blank line between definitions, differing indentation or trailing whitespace — or if the hunk's context lines no longer match the edited file. [reads: code] - **Counter-example**: A patch file whose hunks reproduce the edited source byte-for-byte (a redundant but consistent record), or a `.patch` file living under a fixtures/test-data directory that is input data rather than a description of this change. - **Discriminator**: The failing case has at least one textual divergence between the patch's post-image and the actual file content (e.g. patch says `if Wrapper(x) not in s:` while the source says `if x not in s:`, or the patch deletes a separator blank line the source keeps); the safe case has none, so applying the patch is a no-op. - **Consequence**: If any harness or reviewer applies the artifact, `git apply`/`patch` aborts with "patch does not apply" / "Hunk #1 FAILED" (nonzero exit), or, if it applies, the resulting code differs from the code the tests passed against — the behavior verified is not the behavior shipped. Where the patch also drops a blank line between top-level definitions or introduces trailing whitespace, lint gates (ruff/flake8 E301/W291) fail. - **Evidence**: A committed `fix.patch` restated the source edit but with an extra type-wrapping call in the membership test and with the blank line before the following `@classmethod` deleted; the actual source file contained neither change, so the two representations of the same fix disagreed.
task: the task asks to fix a defect / failing behavior in an existing function; code: the diff touches only that functionfor n in range(...) with while True and switched an iterator call for the list property that merely wraps it, with a summary claiming "Fully backward compatible / Same behavior guaranteed for all test cases"; the scoped run reported 169 passed without exercising any new behavior.8be2eae2bd7a · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "test or grader output will report failure",
"code_requirement": "1. Read the task statement to confirm it names a wrong result / defect to repair rather than requesting a refactor or cleanup [reads: task]",
"prediction": "Existing tests keep passing (they encode the old behavior) while any held-out test written for the reported defect still fails, so correctness credit is zero despite a green local run; this accounts for essentially all of a \"suite passes but fix not accepted\" outcome, with residual risk from unrelated files added alongside."
}### Behavior-preserving cosmetic edit submitted as a bug fix - **Applies when**: `task`: the task asks to fix a defect / failing behavior in an existing function; `code`: the diff touches only that function - **Pattern**: The submission rewrites the target function into an equivalent form — swapping an iterator for the list property that wraps it, restructuring a bounded loop into an unbounded one, adding comments — without changing the condition, ordering, or data that produced the reported defect, and then asserts the change is fully backward compatible. The reported defect is untouched. - **Detection procedure**: 1. Read the task statement to confirm it names a wrong result / defect to repair rather than requesting a refactor or cleanup [reads: task] 2. Locate the changed function and, using the pre-change version quoted in the diff or summary, check what the edit consists of: renamed accessor to an equivalent one, loop-form change, added comments/whitespace, extracted variable [reads: code] 3. Fires when no predicate, boundary, ordering, or returned value changes for any input the old code handled, and the accompanying summary itself states "no behavior change", "fully backward compatible", "same behavior guaranteed for all test cases", or lists only clarity/robustness as the benefit [reads: code] - **Counter-example**: A similarly small diff that changes a comparison operator, an inclusive/exclusive bound, a default, or adds a missing branch — its summary describes an input for which old and new results differ. - **Discriminator**: The failing case cannot name a single input whose result changes and advertises backward compatibility; the safe case's edit alters the output for at least one identified input, which is exactly the defect case. - **Consequence**: Existing tests keep passing (they encode the old behavior) while any held-out test written for the reported defect still fails, so correctness credit is zero despite a green local run; this accounts for essentially all of a "suite passes but fix not accepted" outcome, with residual risk from unrelated files added alongside. - **Evidence**: The delivered change replaced a bounded `for n in range(...)` with `while True` and switched an iterator call for the list property that merely wraps it, with a summary claiming "Fully backward compatible / Same behavior guaranteed for all test cases"; the scoped run reported `169 passed` without exercising any new behavior.
code: the program contains a loop that searches for the first candidate value satisfying a predicate, where candidates are generated from an incrementing counter combined with a caller-supplied template, format string, prefix, or key builder.while True: (or itertools.count()) with the only exit being "candidate not already taken", and nothing guarantees that successive counter values produce distinct candidates — because the candidate is built from a parameter (a %-format template, str.format pattern, or naming callback) that is never validated to actually consume the counter. A caller passing a template without the placeholder makes the loop spin forever.if candidate not in seen: return candidate) and that have no iteration bound, no break outside the success path, and no maximum-attempts counter. [reads: code]candidate is constructed: fire only if it interpolates the counter through a value that arrives as a function parameter or attribute (e.g. template % n, pattern.format(n)) rather than through a literal expression written at that site. [reads: code]try/except TypeError, no cap on n) exists before or inside the loop. [reads: code]f"{prefix}{n}.xml", base + str(n)) so each iteration is provably distinct, or where the loop is for n in range(1, len(seen) + 2) / has an attempt cap and falls through — those terminate for every input.TypeError/None return into a stalled run; also removes the previous implicit None fall-through that callers may rely on.for n in range(1, len(partnames) + 2) was replaced by n = 1; while True: candidate = template % n ... n += 1, removing the only iteration bound on a template supplied by the caller.9f2c395fc665 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj{
"predicate": "a named unit test will fail",
"code_requirement": "1. Locate loops whose exit condition is a membership/collision test (`if candidate not in seen: return candidate`) and that have no iteration bound, no `break` outside the success path, and no maximum-attempts counter. [reads: code]",
"prediction": "For a degenerate or mistyped template the call never returns \u2014 the test process hangs until the harness timeout kills it (no exception, no traceback), turning a would-be `TypeError`/`None` return into a stalled run; also removes the previous implicit `None` fall-through that callers may rely on."
}### Unbounded search loop replacing a bounded one
- **Applies when**: `code`: the program contains a loop that searches for the first candidate value satisfying a predicate, where candidates are generated from an incrementing counter combined with a caller-supplied template, format string, prefix, or key builder.
- **Pattern**: A search loop is written as `while True:` (or `itertools.count()`) with the only exit being "candidate not already taken", and nothing guarantees that successive counter values produce distinct candidates — because the candidate is built from a parameter (a `%`-format template, `str.format` pattern, or naming callback) that is never validated to actually consume the counter. A caller passing a template without the placeholder makes the loop spin forever.
- **Detection procedure**:
1. Locate loops whose exit condition is a membership/collision test (`if candidate not in seen: return candidate`) and that have no iteration bound, no `break` outside the success path, and no maximum-attempts counter. [reads: code]
2. Read how `candidate` is constructed: fire only if it interpolates the counter through a value that arrives as a function parameter or attribute (e.g. `template % n`, `pattern.format(n)`) rather than through a literal expression written at that site. [reads: code]
3. Confirm no validation of that parameter (no assertion/check that the placeholder is present, no `try/except TypeError`, no cap on `n`) exists before or inside the loop. [reads: code]
- **Counter-example**: The same collision-avoiding loop where the candidate is built inline (`f"{prefix}{n}.xml"`, `base + str(n)`) so each iteration is provably distinct, or where the loop is `for n in range(1, len(seen) + 2)` / has an attempt cap and falls through — those terminate for every input.
- **Discriminator**: Termination depends on an unvalidated externally supplied format string in the failing case; in the safe case the counter is guaranteed to appear in the candidate, or a finite bound exists regardless.
- **Consequence**: For a degenerate or mistyped template the call never returns — the test process hangs until the harness timeout kills it (no exception, no traceback), turning a would-be `TypeError`/`None` return into a stalled run; also removes the previous implicit `None` fall-through that callers may rely on.
- **Evidence**: `for n in range(1, len(partnames) + 2)` was replaced by `n = 1; while True: candidate = template % n ... n += 1`, removing the only iteration bound on a template supplied by the caller.code: the program adds a file whose name matches pytest's default collection patterns (test_.py or _test.py) at a location pytest will scantest_.py or _test.py. [reads: code]pytest is the test runner available in the environment. [reads: static facts — python packages]def/fixture) that assign to an attribute of an imported object, e.g. SomeModule.Klass.method = replacement or module.CONST = ..., and check whether any teardown restores the original binding. If the assignment exists at module scope with no restoration, it fires. [reads: code]monkeypatch fixture, or inside a try/finally that restores the original attribute, or placed in a file named e.g. scratch_repro.py that pytest does not collect.test_*.py file and is never undone; safe when it is scoped to a fixture/function or lives in a non-collected filename.AssertionErrors or TypeError/AttributeError from the substitute signature; collection order determines whether it manifests, so results become order-dependent and non-reproducible. This is a latent-failure mechanism separate from whether the underlying task was solved.test_broken.py performed ngram.NGram.normalize = broken_normalize at module scope with no restoration, deliberately degrading a library classmethod for the remainder of any pytest session that collects the file.1bbbb1b34751 · mined from swesmith/Mimino666__langdetect.a1598f1a Mimino666__langdetect.a1598f1a.func_basic__s4s0fk2j{
"predicate": "test or grader output will report failure",
"code_requirement": "1. List the files the program adds and select those whose basename matches `test_*.py` or `*_test.py`. [reads: code]",
"prediction": "Other tests importing the same class observe the replaced implementation, producing spurious `AssertionError`s or `TypeError`/`AttributeError` from the substitute signature; collection order determines whether it manifests, so results become order-dependent and non-reproducible. This is a latent-failure mechanism separate from whether the underlying task was solved."
}### Import-time monkeypatch in a pytest-collected file - **Applies when**: `code`: the program adds a file whose name matches pytest's default collection patterns (`test_*.py` or `*_test.py`) at a location pytest will scan - **Pattern**: A scratch/diagnostic script is given a test-like filename and, at module scope, rebinds an attribute of an imported library module or class (monkeypatching) or otherwise mutates global state, with no fixture, no teardown, and no restoration. Pytest imports the file during collection, so the mutation leaks into every test that runs afterwards in the same session. - **Detection procedure**: 1. List the files the program adds and select those whose basename matches `test_*.py` or `*_test.py`. [reads: code] 2. Confirm `pytest` is the test runner available in the environment. [reads: static facts — python packages] 3. In each such file, look for statements at module indentation level (not inside a `def`/fixture) that assign to an attribute of an imported object, e.g. `SomeModule.Klass.method = replacement` or `module.CONST = ...`, and check whether any teardown restores the original binding. If the assignment exists at module scope with no restoration, it fires. [reads: code] - **Counter-example**: The same rebinding performed inside a test function via the `monkeypatch` fixture, or inside a `try/finally` that restores the original attribute, or placed in a file named e.g. `scratch_repro.py` that pytest does not collect. - **Discriminator**: Goes wrong when the rebinding executes at import time of a collected `test_*.py` file and is never undone; safe when it is scoped to a fixture/function or lives in a non-collected filename. - **Consequence**: Other tests importing the same class observe the replaced implementation, producing spurious `AssertionError`s or `TypeError`/`AttributeError` from the substitute signature; collection order determines whether it manifests, so results become order-dependent and non-reproducible. This is a latent-failure mechanism separate from whether the underlying task was solved. - **Evidence**: An added `test_broken.py` performed `ngram.NGram.normalize = broken_normalize` at module scope with no restoration, deliberately degrading a library classmethod for the remainder of any pytest session that collects the file.
code: the program directly instantiates a class from the repository/library under investigation, passing string literals it wrote itself for identifier-like parameters (path, uri, provider, url, target, search_path).scheme://rest, pkg://a.b, dotted.module:attr), passes a bare value without the required separator, and wraps nothing in a guard, so the script dies inside the library before reaching the code it meant to examine.:// or :). [reads: task]try/except and with no prior read of the class's validation logic or of an existing correctly-formed value obtained from the library itself. [reads: code]try/except and prints the failure while continuing with other probes.ValueError (or TypeError, KeyError, AssertionError) raised inside the library's own argument validation; every later diagnostic line is never reached, so the run yields no information about the reported defect.ImportlibResourcesConfigSource(provider='test', path='hydra.conf') — a dotted name passed where the base class requires a scheme-prefixed path — raised ValueError("Invalid path") in the base __init__, aborting the script after a single print.f01a7985ed07 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty{
"predicate": "an exception will be raised",
"code_requirement": "1. Locate every constructor call or factory call into the library under study and list the literal strings passed to identifier-like keyword arguments. [reads: code]",
"prediction": "The script terminates on the first probe with `ValueError` (or `TypeError`, `KeyError`, `AssertionError`) raised inside the library's own argument validation; every later diagnostic line is never reached, so the run yields no information about the reported defect."
}### Unvalidated literal passed to a constructor that enforces a format on that argument
- **Applies when**: `code`: the program directly instantiates a class from the repository/library under investigation, passing string literals it wrote itself for identifier-like parameters (`path`, `uri`, `provider`, `url`, `target`, `search_path`).
- **Pattern**: A throwaway script hand-guesses the value of a constructor argument whose format is enforced inside that constructor (a scheme/prefix/separator convention such as `scheme://rest`, `pkg://a.b`, `dotted.module:attr`), passes a bare value without the required separator, and wraps nothing in a guard, so the script dies inside the library before reaching the code it meant to examine.
- **Detection procedure**:
1. Locate every constructor call or factory call into the library under study and list the literal strings passed to identifier-like keyword arguments. [reads: code]
2. Check the task statement for any example invocation, config snippet, or quoted value showing the expected form of that argument; note whether the literal in the code matches that form (in particular whether it carries a scheme/prefix separator such as `://` or `:`). [reads: task]
3. Confirm the call is made at module top level with no `try/except` and with no prior read of the class's validation logic or of an existing correctly-formed value obtained from the library itself. [reads: code]
- **Counter-example**: A script that obtains the argument value from the library (e.g. iterates an existing registry/search-path object and feeds one of its entries back in), or that reproduces the documented invocation form exactly, or that wraps the construction in `try/except` and prints the failure while continuing with other probes.
- **Discriminator**: The failing case supplies a literal that the program itself invented and that lacks the separator/prefix the API's own examples show, with no fallback path; the safe case either derives the value from the library or matches a form shown in the task text.
- **Consequence**: The script terminates on the first probe with `ValueError` (or `TypeError`, `KeyError`, `AssertionError`) raised inside the library's own argument validation; every later diagnostic line is never reached, so the run yields no information about the reported defect.
- **Evidence**: `ImportlibResourcesConfigSource(provider='test', path='hydra.conf')` — a dotted name passed where the base class requires a scheme-prefixed path — raised `ValueError("Invalid path")` in the base `__init__`, aborting the script after a single print.