swe-rubrics-mined-100

SWE-bench · 100 rubrics · first 100 of the 1000
HF EdwardoSunny/swe-rubrics-mined-100 · local data/libraries/swe-rubrics-mined-100.json

0Stripping trailing newlines from recursively rendered child output before appending a fixed separatorcodeswesmith/lepture__mistune.bf54ef67
Applies when
code: a class/function renders a container node by calling a recursive child-render helper (e.g. self.render_children(...), self.render(child), "".join(map(self.visit, node.children))) and then wraps or suffixes the result with fixed whitespace/separator text.
Pattern
The container handler normalizes the child-rendered string with .rstrip(), .rstrip('\n'), .strip() or a regex that collapses trailing blank lines, and then appends its own terminator. The child renderers already guarantee their own trailing separator, so the strip plus the new terminator produces a different number of blank lines than every sibling handler emits, silently changing block separation for all downstream text.
Detection procedure
  1. Find the handler method that builds its result from a recursive child-render call, and note every string transformation applied to that call's return value before it is returned. [reads: code]
  2. Read the sibling handler methods in the same class (the ones the task does not ask you to change) and record the terminator convention they follow — e.g. most return ... + '\n\n' and none of them strip the value returned by a child-render call. [reads: code]
  3. Fire if the edited handler is the only one that applies a trailing-whitespace strip to composed child output and still appends the class's standard terminator; i.e. the same characters are both removed and re-added at a different count. [reads: code]
Counter-example
A handler that strips a leaf value taken directly from the token/AST (token['raw'], token.attrs['text'], a source slice) before indenting or wrapping it, or one that strips child output and returns it with no terminator because the caller supplies the separator. Neither double-normalizes.
Discriminator
The wrong case strips the output of a recursive render call whose producers already append the class-wide terminator, and then appends that terminator again; the safe case either strips raw leaf text (which carries no renderer contract) or strips without re-adding a separator.
Consequence
The rendered document contains fewer blank lines after this construct than the reference output. Exact-string comparison tests (fixture/round-trip renderer tests using assertEqual on full output) fail with a whitespace-only diff; no exception is raised, so the failure surfaces only as a lower test-pass score. This accounts for the block-separation portion of the gap; any additional conditional prefix logic the handler keeps or drops relative to the reference accounts for the rest.
Evidence
text = indent(self.render_children(token, state).rstrip('\n'), ' ') followed by return text + '\n\n', where sibling handlers in the same renderer return ... + '\n\n' without stripping child output; the accepted solution kept indent(self.render_children(token, state), ' ') + '\n\n' with no strip.
id e92902aa3f9b · mined from swesmith/lepture__mistune.bf54ef67 lepture__mistune.bf54ef67.lm_rewrite__eb4ybjuf
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Find the handler method that builds its result from a recursive child-render call, and note every string transformation applied to that call's return value before it is returned. [reads: code]",
 "prediction": "The rendered document contains fewer blank lines after this construct than the reference output. Exact-string comparison tests (fixture/round-trip renderer tests using `assertEqual` on full output) fail with a whitespace-only diff; no exception is raised, so the failure surfaces only as a lower test-pass score. This accounts for the block-separation portion of the gap; any additional conditional prefix logic the handler keeps or drops relative to the reference accounts for the rest."
}
raw text (what the judge reads)
### Stripping trailing newlines from recursively rendered child output before appending a fixed separator
- **Applies when**: `code`: a class/function renders a container node by calling a recursive child-render helper (e.g. `self.render_children(...)`, `self.render(child)`, `"".join(map(self.visit, node.children))`) and then wraps or suffixes the result with fixed whitespace/separator text.
- **Pattern**: The container handler normalizes the child-rendered string with `.rstrip()`, `.rstrip('\n')`, `.strip()` or a regex that collapses trailing blank lines, and then appends its own terminator. The child renderers already guarantee their own trailing separator, so the strip plus the new terminator produces a different number of blank lines than every sibling handler emits, silently changing block separation for all downstream text.
- **Detection procedure**:
  1. Find the handler method that builds its result from a recursive child-render call, and note every string transformation applied to that call's return value before it is returned. [reads: code]
  2. Read the sibling handler methods in the same class (the ones the task does not ask you to change) and record the terminator convention they follow — e.g. most return `... + '\n\n'` and none of them strip the value returned by a child-render call. [reads: code]
  3. Fire if the edited handler is the only one that applies a trailing-whitespace strip to composed child output *and* still appends the class's standard terminator; i.e. the same characters are both removed and re-added at a different count. [reads: code]
- **Counter-example**: A handler that strips a *leaf* value taken directly from the token/AST (`token['raw']`, `token.attrs['text']`, a source slice) before indenting or wrapping it, or one that strips child output and returns it with **no** terminator because the caller supplies the separator. Neither double-normalizes.
- **Discriminator**: The wrong case strips the output of a recursive render call whose producers already append the class-wide terminator, and then appends that terminator again; the safe case either strips raw leaf text (which carries no renderer contract) or strips without re-adding a separator.
- **Consequence**: The rendered document contains fewer blank lines after this construct than the reference output. Exact-string comparison tests (fixture/round-trip renderer tests using `assertEqual` on full output) fail with a whitespace-only diff; no exception is raised, so the failure surfaces only as a lower test-pass score. This accounts for the block-separation portion of the gap; any additional conditional prefix logic the handler keeps or drops relative to the reference accounts for the rest.
- **Evidence**: `text = indent(self.render_children(token, state).rstrip('\n'), '   ')` followed by `return text + '\n\n'`, where sibling handlers in the same renderer return `... + '\n\n'` without stripping child output; the accepted solution kept `indent(self.render_children(token, state), '   ') + '\n\n'` with no strip.
1Scratch file named `test_*.py` at repo root executing at import timecodeswesmith/python-openxml__python-docx.0cf6d71f
Applies when
code: the change adds a new Python file to a repository that is exercised by pytest
Pattern
A debug/scratch file is given a name matching pytest's default collection pattern (test_.py / _test.py) and placed where collection reaches it, but its body is bare module-level code with side effects instead of test functions — so the code runs during collection rather than as a test.
Detection procedure
  1. Find newly added .py files whose basename matches test_.py or _test.py. [reads: code]
  2. Check from the static facts that pytest is in the environment and note where the project's real tests live in the repo tree (e.g. ./tests/), i.e. that the new file sits outside that directory, typically at the root. [reads: static facts — python packages, repo tree]
  3. Check whether the file's body consists of top-level executable statements (object construction, method calls, print) with no def test_ / class Test definitions, so the work happens at module import. [reads: code]
Counter-example
A new test_*.py placed inside the project's test package that defines def test_...() functions and confines all setup to fixtures or function bodies; or a scratch script named repro.py / debug_x.py that pytest never collects.
Discriminator
The failing case is both name-matched for collection and does all its work at module scope with no test functions; the safe case either isn't name-matched or keeps side effects inside test functions.
Consequence
If the module body raises, pytest reports a collection error (ERROR ... test_issue.py, surfacing as AttributeError, TypeError, ImportError, or the library's own exception) and the session exits non-zero even though the real tests pass; if it does not raise, it contributes a collected-but-empty module and stray stdout that pollutes the graded test report.
Evidence
A root-level test_issue.py was added whose body immediately runs OpcPackage(), patches package.iter_parts with a Mock, calls package.next_partname(...) and prints — no test function anywhere in the file.
id 044ba6438bb8 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Find newly added `.py` files whose basename matches `test_*.py` or `*_test.py`. [reads: code]",
 "prediction": "If the module body raises, pytest reports a collection error (`ERROR ... test_issue.py`, surfacing as `AttributeError`, `TypeError`, `ImportError`, or the library's own exception) and the session exits non-zero even though the real tests pass; if it does not raise, it contributes a collected-but-empty module and stray stdout that pollutes the graded test report."
}
raw text (what the judge reads)
### Scratch file named `test_*.py` at repo root executing at import time
- **Applies when**: `code`: the change adds a new Python file to a repository that is exercised by pytest
- **Pattern**: A debug/scratch file is given a name matching pytest's default collection pattern (`test_*.py` / `*_test.py`) and placed where collection reaches it, but its body is bare module-level code with side effects instead of test functions — so the code runs during collection rather than as a test.
- **Detection procedure**:
  1. Find newly added `.py` files whose basename matches `test_*.py` or `*_test.py`. [reads: code]
  2. Check from the static facts that `pytest` is in the environment and note where the project's real tests live in the repo tree (e.g. `./tests/`), i.e. that the new file sits outside that directory, typically at the root. [reads: static facts — python packages, repo tree]
  3. Check whether the file's body consists of top-level executable statements (object construction, method calls, `print`) with **no** `def test_*` / `class Test*` definitions, so the work happens at module import. [reads: code]
- **Counter-example**: A new `test_*.py` placed inside the project's test package that defines `def test_...()` functions and confines all setup to fixtures or function bodies; or a scratch script named `repro.py` / `debug_x.py` that pytest never collects.
- **Discriminator**: The failing case is both name-matched for collection *and* does all its work at module scope with no test functions; the safe case either isn't name-matched or keeps side effects inside test functions.
- **Consequence**: If the module body raises, pytest reports a collection error (`ERROR ... test_issue.py`, surfacing as `AttributeError`, `TypeError`, `ImportError`, or the library's own exception) and the session exits non-zero even though the real tests pass; if it does not raise, it contributes a collected-but-empty module and stray stdout that pollutes the graded test report.
- **Evidence**: A root-level `test_issue.py` was added whose body immediately runs `OpcPackage()`, patches `package.iter_parts` with a `Mock`, calls `package.next_partname(...)` and prints — no test function anywhere in the file.
1Duplicate construction of the same collaborator object on one code pathcodeswesmith/python-openxml__python-docx.0cf6d71f
Applies when
code: a function converts a raw value into a wrapper/domain object (constructor, factory, Path(), np.array(), Decimal(), a class from the same package) and the project ships a unit-test suite that mirrors the module being edited
Pattern
The edited code calls the same constructor/factory twice with the same argument on a single execution path — once inside a membership/equality guard and again in the return — instead of constructing once into a local. Externally the value is right, but the collaborator's call count doubles, which breaks unit tests that patch that constructor and assert on how it was called.
Detection procedure
  1. In the function the change touches, list every call to a class/factory name that is imported from the same package (not a builtin like str/int) and note the argument expression of each. [reads: code]
  2. Check whether two such calls use the identical argument expression on a path that can execute both — typically one inside an if/while condition and one in the return immediately below it. [reads: code]
  3. Confirm the constructed object is not bound to a local variable and reused; and confirm the static facts show a test package mirroring the edited module's path (e.g. tests/<subpkg>/test_<module>.py for src/<pkg>/<subpkg>/<module>.py), i.e. this collaborator is likely mocked and its calls asserted. [reads: code; static facts — repo tree]
Counter-example
candidate = Wrapper(template % n) assigned once, then if candidate not in existing: return candidate — the constructor runs once per iteration and once total on the returning path; also safe is a loop that constructs a genuinely different value each iteration and returns it without re-constructing.
Discriminator
The failing case passes the same argument expression to the same constructor twice with no intervening state change, so the returning path invokes it ≥2 times; the safe case invokes it once and reuses the binding.
Consequence
Mock-based unit tests that patch the constructor fail with AssertionError from assert_called_once_with / assert_called_once ("Expected 'X' to be called once. Called 2 times"), even though the returned value is correct; the graded test suite reports the fix as failing.
Evidence
The patch replaced a raw-value membership test with if Wrapper(candidate) not in names: return Wrapper(candidate) (constructing the wrapper in both the guard and the return); the suite failed with AssertionError: Expected 'PackURI' to be called once. Called 2 times.
id 71e065863350 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. In the function the change touches, list every call to a class/factory name that is imported from the same package (not a builtin like `str`/`int`) and note the argument expression of each. [reads: code]",
 "prediction": "Mock-based unit tests that patch the constructor fail with `AssertionError` from `assert_called_once_with` / `assert_called_once` (\"Expected 'X' to be called once. Called 2 times\"), even though the returned value is correct; the graded test suite reports the fix as failing."
}
raw text (what the judge reads)
### Duplicate construction of the same collaborator object on one code path
- **Applies when**: `code`: a function converts a raw value into a wrapper/domain object (constructor, factory, `Path()`, `np.array()`, `Decimal()`, a class from the same package) and the project ships a unit-test suite that mirrors the module being edited
- **Pattern**: The edited code calls the same constructor/factory twice with the same argument on a single execution path — once inside a membership/equality guard and again in the `return` — instead of constructing once into a local. Externally the value is right, but the collaborator's call count doubles, which breaks unit tests that patch that constructor and assert on how it was called.
- **Detection procedure**:
  1. In the function the change touches, list every call to a class/factory name that is imported from the same package (not a builtin like `str`/`int`) and note the argument expression of each. [reads: code]
  2. Check whether two such calls use the identical argument expression on a path that can execute both — typically one inside an `if`/`while` condition and one in the `return` immediately below it. [reads: code]
  3. Confirm the constructed object is not bound to a local variable and reused; and confirm the static facts show a test package mirroring the edited module's path (e.g. `tests/<subpkg>/test_<module>.py` for `src/<pkg>/<subpkg>/<module>.py`), i.e. this collaborator is likely mocked and its calls asserted. [reads: code; static facts — repo tree]
- **Counter-example**: `candidate = Wrapper(template % n)` assigned once, then `if candidate not in existing: return candidate` — the constructor runs once per iteration and once total on the returning path; also safe is a loop that constructs a genuinely different value each iteration and returns it without re-constructing.
- **Discriminator**: The failing case passes the *same* argument expression to the *same* constructor twice with no intervening state change, so the returning path invokes it ≥2 times; the safe case invokes it once and reuses the binding.
- **Consequence**: Mock-based unit tests that patch the constructor fail with `AssertionError` from `assert_called_once_with` / `assert_called_once` ("Expected 'X' to be called once. Called 2 times"), even though the returned value is correct; the graded test suite reports the fix as failing.
- **Evidence**: The patch replaced a raw-value membership test with `if Wrapper(candidate) not in names: return Wrapper(candidate)` (constructing the wrapper in both the guard and the return); the suite failed with `AssertionError: Expected 'PackURI' to be called once. Called 2 times.`
1Behaviour-changing rewrite of a helper whose contract is pinned by existing unit testscodeswesmith/python-openxml__python-docx.0cf6d71f
Applies when
code: the change rewrites the body of an existing library function (rather than adding new code) in a repository whose static facts show a mirrored unit-test package for the edited module
Pattern
The rewrite alters observable interaction details of the function — which internal accessor it calls, how many times it calls a collaborator, whether it can return None — beyond the minimum needed to fix the stated defect. Existing tests pin those details, so the broader-than-necessary rewrite fails tests unrelated to the bug.
Detection procedure
  1. Read the task statement for the specific defective behaviour that must change (the wrong output for a given input). [reads: task]
  2. Diff the rewritten function against the code it replaces and list each behavioural difference: different internal method/property used to obtain the data, extra or fewer calls to imported collaborators, changed loop bounds, changed return-on-exhaustion behaviour. [reads: code]
  3. Flag when at least one listed difference is not required to produce the corrected output — i.e. the corrected output is already achieved by the other differences — and it touches a call to a name imported at module top level (the kind a test patches). [reads: code]
Counter-example
A rewrite that changes only the search bound / comparison that produced the wrong result, keeps the same accessor (self.iter_parts() vs self.parts) and the same single collaborator invocation, and therefore leaves interaction-level assertions intact.
Discriminator
The failing case contains at least one gratuitous interaction change (extra collaborator call or swapped internal accessor) alongside the necessary logic fix; the safe case's diff is confined to the logic that produced the wrong value.
Consequence
Pre-existing unit tests fail with AssertionError on mock call assertions or on patched-attribute expectations, while the functional bug itself is fixed — the submission is scored as failing. This accounts for the interaction-level portion of the failure; the value-level logic may well be correct.
Evidence
The rewrite simultaneously swapped the internal iteration accessor, converted a bounded for into an unbounded while True, and added a second collaborator construction; the only test failure came from the extra collaborator construction, not from the value returned.
id 0282acfa6f79 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Read the task statement for the specific defective behaviour that must change (the wrong output for a given input). [reads: task]",
 "prediction": "Pre-existing unit tests fail with `AssertionError` on mock call assertions or on patched-attribute expectations, while the functional bug itself is fixed \u2014 the submission is scored as failing. This accounts for the interaction-level portion of the failure; the value-level logic may well be correct."
}
raw text (what the judge reads)
### Behaviour-changing rewrite of a helper whose contract is pinned by existing unit tests
- **Applies when**: `code`: the change rewrites the body of an existing library function (rather than adding new code) in a repository whose static facts show a mirrored unit-test package for the edited module
- **Pattern**: The rewrite alters observable interaction details of the function — which internal accessor it calls, how many times it calls a collaborator, whether it can return `None` — beyond the minimum needed to fix the stated defect. Existing tests pin those details, so the broader-than-necessary rewrite fails tests unrelated to the bug.
- **Detection procedure**:
  1. Read the task statement for the specific defective behaviour that must change (the wrong output for a given input). [reads: task]
  2. Diff the rewritten function against the code it replaces and list each behavioural difference: different internal method/property used to obtain the data, extra or fewer calls to imported collaborators, changed loop bounds, changed return-on-exhaustion behaviour. [reads: code]
  3. Flag when at least one listed difference is not required to produce the corrected output — i.e. the corrected output is already achieved by the other differences — and it touches a call to a name imported at module top level (the kind a test patches). [reads: code]
- **Counter-example**: A rewrite that changes only the search bound / comparison that produced the wrong result, keeps the same accessor (`self.iter_parts()` vs `self.parts`) and the same single collaborator invocation, and therefore leaves interaction-level assertions intact.
- **Discriminator**: The failing case contains at least one *gratuitous* interaction change (extra collaborator call or swapped internal accessor) alongside the necessary logic fix; the safe case's diff is confined to the logic that produced the wrong value.
- **Consequence**: Pre-existing unit tests fail with `AssertionError` on mock call assertions or on patched-attribute expectations, while the functional bug itself is fixed — the submission is scored as failing. This accounts for the interaction-level portion of the failure; the value-level logic may well be correct.
- **Evidence**: The rewrite simultaneously swapped the internal iteration accessor, converted a bounded `for` into an unbounded `while True`, and added a second collaborator construction; the only test failure came from the extra collaborator construction, not from the value returned.
1Constructor/wrapper call moved inside the search loop and into the membership testcodeswesmith/python-openxml__python-docx.0cf6d71f
Applies when
code: a function searches for the first unused name/key/identifier by generating candidates from a template or counter and testing them against a collection of already-used values, and a class or factory imported at module level is used to wrap the candidate.
Pattern
The candidate value is passed through a wrapper constructor before the equality/membership check, so (a) the constructor is invoked once per loop iteration instead of once on the value actually returned, and (b) the object compared against the collection is not of the same provenance as the collection's elements. Interaction-based tests that patch that constructor then see the wrong call count/arguments, and with the constructor stubbed the comparison never matches, so the function returns the first candidate.
Detection procedure
  1. Locate the loop that builds candidates (e.g. candidate = template % n, f"{base}{i}", key + str(i)) and tests them for prior use with in, ==, or a dict/set lookup. [reads: code]
  2. Read the expression that builds the collection of used values (attribute reads over a collection of objects, dict keys, a listing) and note whether those elements were produced by the same wrapper class the candidate is passed through. [reads: code]
  3. Fire if the code writes the membership test as Wrapper(candidate) not in used / stores Wrapper(candidate) before the test, where Wrapper is a name imported at module scope, and the elements of used come from somewhere else (raw attribute values, plain strings). Do not fire if the raw candidate is compared and the wrapper is applied only on the return. [reads: code]
Counter-example
for n in count(1): cand = template % n … if cand not in used: return PackURI(cand) — same wrapper class, same search, but the wrapper is constructed exactly once, on the returned value, and never participates in the comparison.
Discriminator
the number of constructor invocations scales with loop iterations and the constructed object is an operand of the equality/membership test, versus exactly one invocation outside the test on the returned value.
Consequence
unit tests that mock.patch the wrapper name in that module fail with AssertionError: expected call not found from assert_called_once_with (extra/earlier calls with the wrong argument), and the function returns the first candidate instead of the first unused one because the stub's constant return value is never found in the used-set; if the loop is an unbounded while True, the opposite stubbing (constant that is in the set) hangs the test run instead. This mechanism accounts for the entire observed test failure here.
Evidence
the search loop was rewritten from comparing the raw candidate to if PackURI(candidate_partname) not in partnames: return candidate_packuri; the patched-constructor test reported Expected: PackURI('/foo/bar/baz2.xml') Actual: PackURI('/foo/bar/baz1.xml') and failed.
id b3c46989e3fd · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the loop that builds candidates (e.g. `candidate = template % n`, `f\"{base}{i}\"`, `key + str(i)`) and tests them for prior use with `in`, `==`, or a dict/set lookup. [reads: code]",
 "prediction": "unit tests that `mock.patch` the wrapper name in that module fail with `AssertionError: expected call not found` from `assert_called_once_with` (extra/earlier calls with the wrong argument), and the function returns the *first* candidate instead of the first *unused* one because the stub's constant return value is never found in the used-set; if the loop is an unbounded `while True`, the opposite stubbing (constant that is in the set) hangs the test run instead. This mechanism accounts for the entire observed test failure here."
}
raw text (what the judge reads)
### Constructor/wrapper call moved inside the search loop and into the membership test
- **Applies when**: `code`: a function searches for the first unused name/key/identifier by generating candidates from a template or counter and testing them against a collection of already-used values, and a class or factory imported at module level is used to wrap the candidate.
- **Pattern**: The candidate value is passed through a wrapper constructor *before* the equality/membership check, so (a) the constructor is invoked once per loop iteration instead of once on the value actually returned, and (b) the object compared against the collection is not of the same provenance as the collection's elements. Interaction-based tests that patch that constructor then see the wrong call count/arguments, and with the constructor stubbed the comparison never matches, so the function returns the first candidate.
- **Detection procedure**:
  1. Locate the loop that builds candidates (e.g. `candidate = template % n`, `f"{base}{i}"`, `key + str(i)`) and tests them for prior use with `in`, `==`, or a dict/set lookup. [reads: code]
  2. Read the expression that builds the collection of used values (attribute reads over a collection of objects, dict keys, a listing) and note whether those elements were produced by the same wrapper class the candidate is passed through. [reads: code]
  3. Fire if the code writes the membership test as `Wrapper(candidate) not in used` / stores `Wrapper(candidate)` before the test, where `Wrapper` is a name imported at module scope, and the elements of `used` come from somewhere else (raw attribute values, plain strings). Do not fire if the raw candidate is compared and the wrapper is applied only on the `return`. [reads: code]
- **Counter-example**: `for n in count(1): cand = template % n` … `if cand not in used: return PackURI(cand)` — same wrapper class, same search, but the wrapper is constructed exactly once, on the returned value, and never participates in the comparison.
- **Discriminator**: the number of constructor invocations scales with loop iterations and the constructed object is an operand of the equality/membership test, versus exactly one invocation outside the test on the returned value.
- **Consequence**: unit tests that `mock.patch` the wrapper name in that module fail with `AssertionError: expected call not found` from `assert_called_once_with` (extra/earlier calls with the wrong argument), and the function returns the *first* candidate instead of the first *unused* one because the stub's constant return value is never found in the used-set; if the loop is an unbounded `while True`, the opposite stubbing (constant that is in the set) hangs the test run instead. This mechanism accounts for the entire observed test failure here.
- **Evidence**: the search loop was rewritten from comparing the raw candidate to `if PackURI(candidate_partname) not in partnames: return candidate_packuri`; the patched-constructor test reported `Expected: PackURI('/foo/bar/baz2.xml')  Actual: PackURI('/foo/bar/baz1.xml')` and failed.
1Patch/diff artifact committed alongside the real edit and disagreeing with itcodeswesmith/python-openxml__python-docx.0cf6d71f
Applies when
code: the change set includes both a modification to a source file and a separate file containing a unified diff (.patch, .diff, or a file whose text begins with --- a/ / +++ b/)
Pattern
The program records its intended change twice — once by editing the source and once as a checked-in patch file — and the two copies are not identical, so the artifact describing the fix does not match the fix that was actually applied and tested.
Detection procedure
  1. Locate any added file whose name ends in .patch/.diff or whose first lines are --- a/<path> / +++ b/<path>, and read the path named in its headers. [reads: code]
  2. Confirm that same path is also directly modified by the program (it appears as an edited source file in the change set / repo tree). [reads: code, and static facts — repo tree for the source path]
  3. Line-by-line, compare every + line of the patch hunk with the corresponding region of the edited source. Fires if any added line differs — a different expression or call wrapper, an added/removed blank line between definitions, differing indentation or trailing whitespace — or if the hunk's context lines no longer match the edited file. [reads: code]
Counter-example
A patch file whose hunks reproduce the edited source byte-for-byte (a redundant but consistent record), or a .patch file living under a fixtures/test-data directory that is input data rather than a description of this change.
Discriminator
The failing case has at least one textual divergence between the patch's post-image and the actual file content (e.g. patch says if Wrapper(x) not in s: while the source says if x not in s:, or the patch deletes a separator blank line the source keeps); the safe case has none, so applying the patch is a no-op.
Consequence
If any harness or reviewer applies the artifact, git apply/patch aborts with "patch does not apply" / "Hunk #1 FAILED" (nonzero exit), or, if it applies, the resulting code differs from the code the tests passed against — the behavior verified is not the behavior shipped. Where the patch also drops a blank line between top-level definitions or introduces trailing whitespace, lint gates (ruff/flake8 E301/W291) fail.
Evidence
A committed fix.patch restated the source edit but with an extra type-wrapping call in the membership test and with the blank line before the following @classmethod deleted; the actual source file contained neither change, so the two representations of the same fix disagreed.
id daf034bd2ebc · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate any added file whose name ends in `.patch`/`.diff` or whose first lines are `--- a/<path>` / `+++ b/<path>`, and read the path named in its headers. [reads: code]",
 "prediction": "If any harness or reviewer applies the artifact, `git apply`/`patch` aborts with \"patch does not apply\" / \"Hunk #1 FAILED\" (nonzero exit), or, if it applies, the resulting code differs from the code the tests passed against \u2014 the behavior verified is not the behavior shipped. Where the patch also drops a blank line between top-level definitions or introduces trailing whitespace, lint gates (ruff/flake8 E301/W291) fail."
}
raw text (what the judge reads)
### Patch/diff artifact committed alongside the real edit and disagreeing with it
- **Applies when**: `code`: the change set includes both a modification to a source file and a separate file containing a unified diff (`*.patch`, `*.diff`, or a file whose text begins with `--- a/` / `+++ b/`)
- **Pattern**: The program records its intended change twice — once by editing the source and once as a checked-in patch file — and the two copies are not identical, so the artifact describing the fix does not match the fix that was actually applied and tested.
- **Detection procedure**:
  1. Locate any added file whose name ends in `.patch`/`.diff` or whose first lines are `--- a/<path>` / `+++ b/<path>`, and read the path named in its headers. [reads: code]
  2. Confirm that same path is also directly modified by the program (it appears as an edited source file in the change set / repo tree). [reads: code, and static facts — repo tree for the source path]
  3. Line-by-line, compare every `+` line of the patch hunk with the corresponding region of the edited source. Fires if any added line differs — a different expression or call wrapper, an added/removed blank line between definitions, differing indentation or trailing whitespace — or if the hunk's context lines no longer match the edited file. [reads: code]
- **Counter-example**: A patch file whose hunks reproduce the edited source byte-for-byte (a redundant but consistent record), or a `.patch` file living under a fixtures/test-data directory that is input data rather than a description of this change.
- **Discriminator**: The failing case has at least one textual divergence between the patch's post-image and the actual file content (e.g. patch says `if Wrapper(x) not in s:` while the source says `if x not in s:`, or the patch deletes a separator blank line the source keeps); the safe case has none, so applying the patch is a no-op.
- **Consequence**: If any harness or reviewer applies the artifact, `git apply`/`patch` aborts with "patch does not apply" / "Hunk #1 FAILED" (nonzero exit), or, if it applies, the resulting code differs from the code the tests passed against — the behavior verified is not the behavior shipped. Where the patch also drops a blank line between top-level definitions or introduces trailing whitespace, lint gates (ruff/flake8 E301/W291) fail.
- **Evidence**: A committed `fix.patch` restated the source edit but with an extra type-wrapping call in the membership test and with the blank line before the following `@classmethod` deleted; the actual source file contained neither change, so the two representations of the same fix disagreed.
1Behavior-preserving cosmetic edit submitted as a bug fixtaskswesmith/python-openxml__python-docx.0cf6d71f
Applies when
task: the task asks to fix a defect / failing behavior in an existing function; code: the diff touches only that function
Pattern
The submission rewrites the target function into an equivalent form — swapping an iterator for the list property that wraps it, restructuring a bounded loop into an unbounded one, adding comments — without changing the condition, ordering, or data that produced the reported defect, and then asserts the change is fully backward compatible. The reported defect is untouched.
Detection procedure
  1. Read the task statement to confirm it names a wrong result / defect to repair rather than requesting a refactor or cleanup [reads: task]
  2. Locate the changed function and, using the pre-change version quoted in the diff or summary, check what the edit consists of: renamed accessor to an equivalent one, loop-form change, added comments/whitespace, extracted variable [reads: code]
  3. Fires when no predicate, boundary, ordering, or returned value changes for any input the old code handled, and the accompanying summary itself states "no behavior change", "fully backward compatible", "same behavior guaranteed for all test cases", or lists only clarity/robustness as the benefit [reads: code]
Counter-example
A similarly small diff that changes a comparison operator, an inclusive/exclusive bound, a default, or adds a missing branch — its summary describes an input for which old and new results differ.
Discriminator
The failing case cannot name a single input whose result changes and advertises backward compatibility; the safe case's edit alters the output for at least one identified input, which is exactly the defect case.
Consequence
Existing tests keep passing (they encode the old behavior) while any held-out test written for the reported defect still fails, so correctness credit is zero despite a green local run; this accounts for essentially all of a "suite passes but fix not accepted" outcome, with residual risk from unrelated files added alongside.
Evidence
The delivered change replaced a bounded for n in range(...) with while True and switched an iterator call for the list property that merely wraps it, with a summary claiming "Fully backward compatible / Same behavior guaranteed for all test cases"; the scoped run reported 169 passed without exercising any new behavior.
id 8be2eae2bd7a · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Read the task statement to confirm it names a wrong result / defect to repair rather than requesting a refactor or cleanup [reads: task]",
 "prediction": "Existing tests keep passing (they encode the old behavior) while any held-out test written for the reported defect still fails, so correctness credit is zero despite a green local run; this accounts for essentially all of a \"suite passes but fix not accepted\" outcome, with residual risk from unrelated files added alongside."
}
raw text (what the judge reads)
### Behavior-preserving cosmetic edit submitted as a bug fix
- **Applies when**: `task`: the task asks to fix a defect / failing behavior in an existing function; `code`: the diff touches only that function
- **Pattern**: The submission rewrites the target function into an equivalent form — swapping an iterator for the list property that wraps it, restructuring a bounded loop into an unbounded one, adding comments — without changing the condition, ordering, or data that produced the reported defect, and then asserts the change is fully backward compatible. The reported defect is untouched.
- **Detection procedure**:
  1. Read the task statement to confirm it names a wrong result / defect to repair rather than requesting a refactor or cleanup [reads: task]
  2. Locate the changed function and, using the pre-change version quoted in the diff or summary, check what the edit consists of: renamed accessor to an equivalent one, loop-form change, added comments/whitespace, extracted variable [reads: code]
  3. Fires when no predicate, boundary, ordering, or returned value changes for any input the old code handled, and the accompanying summary itself states "no behavior change", "fully backward compatible", "same behavior guaranteed for all test cases", or lists only clarity/robustness as the benefit [reads: code]
- **Counter-example**: A similarly small diff that changes a comparison operator, an inclusive/exclusive bound, a default, or adds a missing branch — its summary describes an input for which old and new results differ.
- **Discriminator**: The failing case cannot name a single input whose result changes and advertises backward compatibility; the safe case's edit alters the output for at least one identified input, which is exactly the defect case.
- **Consequence**: Existing tests keep passing (they encode the old behavior) while any held-out test written for the reported defect still fails, so correctness credit is zero despite a green local run; this accounts for essentially all of a "suite passes but fix not accepted" outcome, with residual risk from unrelated files added alongside.
- **Evidence**: The delivered change replaced a bounded `for n in range(...)` with `while True` and switched an iterator call for the list property that merely wraps it, with a summary claiming "Fully backward compatible / Same behavior guaranteed for all test cases"; the scoped run reported `169 passed` without exercising any new behavior.
1Unbounded search loop replacing a bounded onecodeswesmith/python-openxml__python-docx.0cf6d71f
Applies when
code: the program contains a loop that searches for the first candidate value satisfying a predicate, where candidates are generated from an incrementing counter combined with a caller-supplied template, format string, prefix, or key builder.
Pattern
A search loop is written as while True: (or itertools.count()) with the only exit being "candidate not already taken", and nothing guarantees that successive counter values produce distinct candidates — because the candidate is built from a parameter (a %-format template, str.format pattern, or naming callback) that is never validated to actually consume the counter. A caller passing a template without the placeholder makes the loop spin forever.
Detection procedure
  1. Locate loops whose exit condition is a membership/collision test (if candidate not in seen: return candidate) and that have no iteration bound, no break outside the success path, and no maximum-attempts counter. [reads: code]
  2. Read how candidate is constructed: fire only if it interpolates the counter through a value that arrives as a function parameter or attribute (e.g. template % n, pattern.format(n)) rather than through a literal expression written at that site. [reads: code]
  3. Confirm no validation of that parameter (no assertion/check that the placeholder is present, no try/except TypeError, no cap on n) exists before or inside the loop. [reads: code]
Counter-example
The same collision-avoiding loop where the candidate is built inline (f"{prefix}{n}.xml", base + str(n)) so each iteration is provably distinct, or where the loop is for n in range(1, len(seen) + 2) / has an attempt cap and falls through — those terminate for every input.
Discriminator
Termination depends on an unvalidated externally supplied format string in the failing case; in the safe case the counter is guaranteed to appear in the candidate, or a finite bound exists regardless.
Consequence
For a degenerate or mistyped template the call never returns — the test process hangs until the harness timeout kills it (no exception, no traceback), turning a would-be TypeError/None return into a stalled run; also removes the previous implicit None fall-through that callers may rely on.
Evidence
for n in range(1, len(partnames) + 2) was replaced by n = 1; while True: candidate = template % n ... n += 1, removing the only iteration bound on a template supplied by the caller.
id 9f2c395fc665 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__3g1tyktj
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Locate loops whose exit condition is a membership/collision test (`if candidate not in seen: return candidate`) and that have no iteration bound, no `break` outside the success path, and no maximum-attempts counter. [reads: code]",
 "prediction": "For a degenerate or mistyped template the call never returns \u2014 the test process hangs until the harness timeout kills it (no exception, no traceback), turning a would-be `TypeError`/`None` return into a stalled run; also removes the previous implicit `None` fall-through that callers may rely on."
}
raw text (what the judge reads)
### Unbounded search loop replacing a bounded one
- **Applies when**: `code`: the program contains a loop that searches for the first candidate value satisfying a predicate, where candidates are generated from an incrementing counter combined with a caller-supplied template, format string, prefix, or key builder.
- **Pattern**: A search loop is written as `while True:` (or `itertools.count()`) with the only exit being "candidate not already taken", and nothing guarantees that successive counter values produce distinct candidates — because the candidate is built from a parameter (a `%`-format template, `str.format` pattern, or naming callback) that is never validated to actually consume the counter. A caller passing a template without the placeholder makes the loop spin forever.
- **Detection procedure**:
  1. Locate loops whose exit condition is a membership/collision test (`if candidate not in seen: return candidate`) and that have no iteration bound, no `break` outside the success path, and no maximum-attempts counter. [reads: code]
  2. Read how `candidate` is constructed: fire only if it interpolates the counter through a value that arrives as a function parameter or attribute (e.g. `template % n`, `pattern.format(n)`) rather than through a literal expression written at that site. [reads: code]
  3. Confirm no validation of that parameter (no assertion/check that the placeholder is present, no `try/except TypeError`, no cap on `n`) exists before or inside the loop. [reads: code]
- **Counter-example**: The same collision-avoiding loop where the candidate is built inline (`f"{prefix}{n}.xml"`, `base + str(n)`) so each iteration is provably distinct, or where the loop is `for n in range(1, len(seen) + 2)` / has an attempt cap and falls through — those terminate for every input.
- **Discriminator**: Termination depends on an unvalidated externally supplied format string in the failing case; in the safe case the counter is guaranteed to appear in the candidate, or a finite bound exists regardless.
- **Consequence**: For a degenerate or mistyped template the call never returns — the test process hangs until the harness timeout kills it (no exception, no traceback), turning a would-be `TypeError`/`None` return into a stalled run; also removes the previous implicit `None` fall-through that callers may rely on.
- **Evidence**: `for n in range(1, len(partnames) + 2)` was replaced by `n = 1; while True: candidate = template % n ... n += 1`, removing the only iteration bound on a template supplied by the caller.
2Import-time monkeypatch in a pytest-collected filecodeswesmith/Mimino666__langdetect.a1598f1a
Applies when
code: the program adds a file whose name matches pytest's default collection patterns (test_.py or _test.py) at a location pytest will scan
Pattern
A scratch/diagnostic script is given a test-like filename and, at module scope, rebinds an attribute of an imported library module or class (monkeypatching) or otherwise mutates global state, with no fixture, no teardown, and no restoration. Pytest imports the file during collection, so the mutation leaks into every test that runs afterwards in the same session.
Detection procedure
  1. List the files the program adds and select those whose basename matches test_.py or _test.py. [reads: code]
  2. Confirm pytest is the test runner available in the environment. [reads: static facts — python packages]
  3. In each such file, look for statements at module indentation level (not inside a def/fixture) that assign to an attribute of an imported object, e.g. SomeModule.Klass.method = replacement or module.CONST = ..., and check whether any teardown restores the original binding. If the assignment exists at module scope with no restoration, it fires. [reads: code]
Counter-example
The same rebinding performed inside a test function via the monkeypatch fixture, or inside a try/finally that restores the original attribute, or placed in a file named e.g. scratch_repro.py that pytest does not collect.
Discriminator
Goes wrong when the rebinding executes at import time of a collected test_*.py file and is never undone; safe when it is scoped to a fixture/function or lives in a non-collected filename.
Consequence
Other tests importing the same class observe the replaced implementation, producing spurious AssertionErrors or TypeError/AttributeError from the substitute signature; collection order determines whether it manifests, so results become order-dependent and non-reproducible. This is a latent-failure mechanism separate from whether the underlying task was solved.
Evidence
An added test_broken.py performed ngram.NGram.normalize = broken_normalize at module scope with no restoration, deliberately degrading a library classmethod for the remainder of any pytest session that collects the file.
id 1bbbb1b34751 · mined from swesmith/Mimino666__langdetect.a1598f1a Mimino666__langdetect.a1598f1a.func_basic__s4s0fk2j
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List the files the program adds and select those whose basename matches `test_*.py` or `*_test.py`. [reads: code]",
 "prediction": "Other tests importing the same class observe the replaced implementation, producing spurious `AssertionError`s or `TypeError`/`AttributeError` from the substitute signature; collection order determines whether it manifests, so results become order-dependent and non-reproducible. This is a latent-failure mechanism separate from whether the underlying task was solved."
}
raw text (what the judge reads)
### Import-time monkeypatch in a pytest-collected file
- **Applies when**: `code`: the program adds a file whose name matches pytest's default collection patterns (`test_*.py` or `*_test.py`) at a location pytest will scan
- **Pattern**: A scratch/diagnostic script is given a test-like filename and, at module scope, rebinds an attribute of an imported library module or class (monkeypatching) or otherwise mutates global state, with no fixture, no teardown, and no restoration. Pytest imports the file during collection, so the mutation leaks into every test that runs afterwards in the same session.
- **Detection procedure**:
  1. List the files the program adds and select those whose basename matches `test_*.py` or `*_test.py`. [reads: code]
  2. Confirm `pytest` is the test runner available in the environment. [reads: static facts — python packages]
  3. In each such file, look for statements at module indentation level (not inside a `def`/fixture) that assign to an attribute of an imported object, e.g. `SomeModule.Klass.method = replacement` or `module.CONST = ...`, and check whether any teardown restores the original binding. If the assignment exists at module scope with no restoration, it fires. [reads: code]
- **Counter-example**: The same rebinding performed inside a test function via the `monkeypatch` fixture, or inside a `try/finally` that restores the original attribute, or placed in a file named e.g. `scratch_repro.py` that pytest does not collect.
- **Discriminator**: Goes wrong when the rebinding executes at import time of a collected `test_*.py` file and is never undone; safe when it is scoped to a fixture/function or lives in a non-collected filename.
- **Consequence**: Other tests importing the same class observe the replaced implementation, producing spurious `AssertionError`s or `TypeError`/`AttributeError` from the substitute signature; collection order determines whether it manifests, so results become order-dependent and non-reproducible. This is a latent-failure mechanism separate from whether the underlying task was solved.
- **Evidence**: An added `test_broken.py` performed `ngram.NGram.normalize = broken_normalize` at module scope with no restoration, deliberately degrading a library classmethod for the remainder of any pytest session that collects the file.
3Unvalidated literal passed to a constructor that enforces a format on that argumentcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program directly instantiates a class from the repository/library under investigation, passing string literals it wrote itself for identifier-like parameters (path, uri, provider, url, target, search_path).
Pattern
A throwaway script hand-guesses the value of a constructor argument whose format is enforced inside that constructor (a scheme/prefix/separator convention such as scheme://rest, pkg://a.b, dotted.module:attr), passes a bare value without the required separator, and wraps nothing in a guard, so the script dies inside the library before reaching the code it meant to examine.
Detection procedure
  1. Locate every constructor call or factory call into the library under study and list the literal strings passed to identifier-like keyword arguments. [reads: code]
  2. Check the task statement for any example invocation, config snippet, or quoted value showing the expected form of that argument; note whether the literal in the code matches that form (in particular whether it carries a scheme/prefix separator such as :// or :). [reads: task]
  3. Confirm the call is made at module top level with no try/except and with no prior read of the class's validation logic or of an existing correctly-formed value obtained from the library itself. [reads: code]
Counter-example
A script that obtains the argument value from the library (e.g. iterates an existing registry/search-path object and feeds one of its entries back in), or that reproduces the documented invocation form exactly, or that wraps the construction in try/except and prints the failure while continuing with other probes.
Discriminator
The failing case supplies a literal that the program itself invented and that lacks the separator/prefix the API's own examples show, with no fallback path; the safe case either derives the value from the library or matches a form shown in the task text.
Consequence
The script terminates on the first probe with ValueError (or TypeError, KeyError, AssertionError) raised inside the library's own argument validation; every later diagnostic line is never reached, so the run yields no information about the reported defect.
Evidence
ImportlibResourcesConfigSource(provider='test', path='hydra.conf') — a dotted name passed where the base class requires a scheme-prefixed path — raised ValueError("Invalid path") in the base __init__, aborting the script after a single print.
id f01a7985ed07 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate every constructor call or factory call into the library under study and list the literal strings passed to identifier-like keyword arguments. [reads: code]",
 "prediction": "The script terminates on the first probe with `ValueError` (or `TypeError`, `KeyError`, `AssertionError`) raised inside the library's own argument validation; every later diagnostic line is never reached, so the run yields no information about the reported defect."
}
raw text (what the judge reads)
### Unvalidated literal passed to a constructor that enforces a format on that argument
- **Applies when**: `code`: the program directly instantiates a class from the repository/library under investigation, passing string literals it wrote itself for identifier-like parameters (`path`, `uri`, `provider`, `url`, `target`, `search_path`).
- **Pattern**: A throwaway script hand-guesses the value of a constructor argument whose format is enforced inside that constructor (a scheme/prefix/separator convention such as `scheme://rest`, `pkg://a.b`, `dotted.module:attr`), passes a bare value without the required separator, and wraps nothing in a guard, so the script dies inside the library before reaching the code it meant to examine.
- **Detection procedure**:
  1. Locate every constructor call or factory call into the library under study and list the literal strings passed to identifier-like keyword arguments. [reads: code]
  2. Check the task statement for any example invocation, config snippet, or quoted value showing the expected form of that argument; note whether the literal in the code matches that form (in particular whether it carries a scheme/prefix separator such as `://` or `:`). [reads: task]
  3. Confirm the call is made at module top level with no `try/except` and with no prior read of the class's validation logic or of an existing correctly-formed value obtained from the library itself. [reads: code]
- **Counter-example**: A script that obtains the argument value from the library (e.g. iterates an existing registry/search-path object and feeds one of its entries back in), or that reproduces the documented invocation form exactly, or that wraps the construction in `try/except` and prints the failure while continuing with other probes.
- **Discriminator**: The failing case supplies a literal that the program itself invented and that lacks the separator/prefix the API's own examples show, with no fallback path; the safe case either derives the value from the library or matches a form shown in the task text.
- **Consequence**: The script terminates on the first probe with `ValueError` (or `TypeError`, `KeyError`, `AssertionError`) raised inside the library's own argument validation; every later diagnostic line is never reached, so the run yields no information about the reported defect.
- **Evidence**: `ImportlibResourcesConfigSource(provider='test', path='hydra.conf')` — a dotted name passed where the base class requires a scheme-prefixed path — raised `ValueError("Invalid path")` in the base `__init__`, aborting the script after a single print.
3Diagnostic run ignores the reproduction commands the task suppliestaskswesmith/facebookresearch__hydra.0f03eb60
Applies when
task: the issue/task text contains explicit, runnable reproduction commands or scripts with expected-vs-observed results; code: the program is an investigation/reproduction step.
Pattern
Instead of executing the commands the report provides, the program invents an ad-hoc entry point into internal modules (importing private submodules and constructing objects by hand). It exercises a code path the report never mentions, so whatever it observes — success or exception — cannot confirm or localize the reported defect.
Detection procedure
  1. Extract the literal commands / script paths / CLI overrides given under the task's reproduction steps. [reads: task]
  2. Search the program text for any of those script paths, module entry points, or override strings being invoked (directly, via subprocess, or via the library's public API). [reads: code]
  3. Confirm none of them appear and the program instead imports internal modules (names with a leading underscore package component or _internal-style path) and calls their constructors/methods with self-authored arguments. [reads: code]
  4. Check whether the file or function the task names as the suspected cause is referenced anywhere in the program. [reads: task, code]
Counter-example
A program that runs at least one of the supplied reproduction commands (e.g. via subprocess.run on the named script with the named overrides) and additionally pokes at internals, or one that imports internals but calls precisely the function the report names as suspect.
Discriminator
The failing case shares no entry point, script path, or named suspect function with the reproduction steps; the safe case reuses at least one of them, so its output is comparable to the report's "expected vs observed".
Consequence
The step produces output about an unrelated code path (commonly an exception from the improvised call, e.g. ValueError/ImportError/TypeError), the reported behavior is neither reproduced nor localized, and the underlying defect remains unfixed. Explains the wasted step here; the immediate crash itself is attributable to the separate malformed-argument mechanism.
Evidence
The report gave three concrete python examples/.../my_app.py ... commands and named _read_config's use of read_text() as the suspect; the program ran none of them and instead hand-constructed an internal config-source object, which raised before printing anything useful.
id 127fbb57acd7 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Extract the literal commands / script paths / CLI overrides given under the task's reproduction steps. [reads: task]",
 "prediction": "The step produces output about an unrelated code path (commonly an exception from the improvised call, e.g. `ValueError`/`ImportError`/`TypeError`), the reported behavior is neither reproduced nor localized, and the underlying defect remains unfixed. Explains the wasted step here; the immediate crash itself is attributable to the separate malformed-argument mechanism."
}
raw text (what the judge reads)
### Diagnostic run ignores the reproduction commands the task supplies
- **Applies when**: `task`: the issue/task text contains explicit, runnable reproduction commands or scripts with expected-vs-observed results; `code`: the program is an investigation/reproduction step.
- **Pattern**: Instead of executing the commands the report provides, the program invents an ad-hoc entry point into internal modules (importing private submodules and constructing objects by hand). It exercises a code path the report never mentions, so whatever it observes — success or exception — cannot confirm or localize the reported defect.
- **Detection procedure**:
  1. Extract the literal commands / script paths / CLI overrides given under the task's reproduction steps. [reads: task]
  2. Search the program text for any of those script paths, module entry points, or override strings being invoked (directly, via `subprocess`, or via the library's public API). [reads: code]
  3. Confirm none of them appear and the program instead imports internal modules (names with a leading underscore package component or `_internal`-style path) and calls their constructors/methods with self-authored arguments. [reads: code]
  4. Check whether the file or function the task names as the suspected cause is referenced anywhere in the program. [reads: task, code]
- **Counter-example**: A program that runs at least one of the supplied reproduction commands (e.g. via `subprocess.run` on the named script with the named overrides) and *additionally* pokes at internals, or one that imports internals but calls precisely the function the report names as suspect.
- **Discriminator**: The failing case shares no entry point, script path, or named suspect function with the reproduction steps; the safe case reuses at least one of them, so its output is comparable to the report's "expected vs observed".
- **Consequence**: The step produces output about an unrelated code path (commonly an exception from the improvised call, e.g. `ValueError`/`ImportError`/`TypeError`), the reported behavior is neither reproduced nor localized, and the underlying defect remains unfixed. Explains the wasted step here; the immediate crash itself is attributable to the separate malformed-argument mechanism.
- **Evidence**: The report gave three concrete `python examples/.../my_app.py ...` commands and named `_read_config`'s use of `read_text()` as the suspect; the program ran none of them and instead hand-constructed an internal config-source object, which raised before printing anything useful.
3Verifying a fix by substring-matching `inspect.getsource` instead of exercising itcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program validates that some function/method is correct by testing string membership against source text obtained from inspect.getsource, open(module_file).read(), or similar
Pattern
Correctness is asserted textually — 'some_expr' in source — rather than by calling the function and comparing its result to an expected value. The check is satisfied by any formatting-equivalent variation and fails on any correct implementation written differently, so it certifies nothing about runtime behavior.
Detection procedure
  1. Locate the verification block: expressions of the form <literal> in source where source came from inspect.getsource(...) or a file read of a .py path. [reads: code]
  2. Read the task statement for the behavior actually required (an output line, an error message, a returned config/value). [reads: task]
  3. Check whether the program anywhere invokes the function under test with real inputs and compares the result to the required behavior; the defect is present when every check is a string-membership test and no invocation exists. [reads: code]
Counter-example
A program that uses inspect.getsource only to locate or rewrite code, but separately calls the function on a real input (e.g. loads an actual config file / runs the example command) and asserts on the returned value or emitted text.
Discriminator
The failing case's only evidence of correctness is literal text appearing in a source string; the safe case has at least one execution of the code path with an assertion on its output.
Discriminator holds regardless of domain
the same shape appears whenever "did I fix it?" is answered by grepping source rather than running it.
Consequence
False verdicts in both directions — the program reports success for a semantically broken implementation whose source happens to contain the literal, and reports failure for a correct implementation formatted differently (extra spaces, renamed local, split expression). Downstream, the real behavioral tests still fail; expect no improvement on behavior-graded checks.
Evidence
Checks of the form ('proper return', 'return ret' in source and 'ret = res.exists() and res.is_file()' in source) were used as the sole correctness criterion for a method whose reported bug was a runtime behavior difference in how file contents are read.
id f4a70e82660d · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the verification block: expressions of the form `<literal> in source` where `source` came from `inspect.getsource(...)` or a file read of a `.py` path. [reads: code]",
 "prediction": "False verdicts in both directions \u2014 the program reports success for a semantically broken implementation whose source happens to contain the literal, and reports failure for a correct implementation formatted differently (extra spaces, renamed local, split expression). Downstream, the real behavioral tests still fail; expect no improvement on behavior-graded checks."
}
raw text (what the judge reads)
### Verifying a fix by substring-matching `inspect.getsource` instead of exercising it
- **Applies when**: `code`: the program validates that some function/method is correct by testing string membership against source text obtained from `inspect.getsource`, `open(module_file).read()`, or similar
- **Pattern**: Correctness is asserted textually — `'some_expr' in source` — rather than by calling the function and comparing its result to an expected value. The check is satisfied by any formatting-equivalent variation and fails on any correct implementation written differently, so it certifies nothing about runtime behavior.
- **Detection procedure**:
  1. Locate the verification block: expressions of the form `<literal> in source` where `source` came from `inspect.getsource(...)` or a file read of a `.py` path. [reads: code]
  2. Read the task statement for the behavior actually required (an output line, an error message, a returned config/value). [reads: task]
  3. Check whether the program anywhere *invokes* the function under test with real inputs and compares the result to the required behavior; the defect is present when every check is a string-membership test and no invocation exists. [reads: code]
- **Counter-example**: A program that uses `inspect.getsource` only to locate or rewrite code, but separately calls the function on a real input (e.g. loads an actual config file / runs the example command) and asserts on the returned value or emitted text.
- **Discriminator**: The failing case's *only* evidence of correctness is literal text appearing in a source string; the safe case has at least one execution of the code path with an assertion on its output.
- **Discriminator holds regardless of domain**: the same shape appears whenever "did I fix it?" is answered by grepping source rather than running it.
- **Consequence**: False verdicts in both directions — the program reports success for a semantically broken implementation whose source happens to contain the literal, and reports failure for a correct implementation formatted differently (extra spaces, renamed local, split expression). Downstream, the real behavioral tests still fail; expect no improvement on behavior-graded checks.
- **Evidence**: Checks of the form `('proper return', 'return ret' in source and 'ret = res.exists() and res.is_file()' in source)` were used as the sole correctness criterion for a method whose reported bug was a runtime behavior difference in how file contents are read.
3Unguarded attribute access alongside hasattr probes of sibling memberscodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program reflects over a class or module whose API it is not certain about, using hasattr, getattr, dir, or inspect.getsource
Pattern
The program probes some members defensively with hasattr(...) but then dereferences another member of the same object directly (e.g. passing Cls.member to inspect.getsource) with no existence check and no try/except, so a missing or non-source-backed member aborts the whole script.
Detection procedure
  1. Locate the reflection block: calls to hasattr/getattr on a class or module object [reads: code]
  2. In the same block, find a direct dotted access on that same object (Obj.name) used as an argument to inspect.getsource, inspect.signature, or similar [reads: code]
  3. Check whether that direct access is inside if hasattr(...)/try: ... except AttributeError/getattr(obj, name, default); the defect is present when it is bare while sibling names are hasattr-guarded [reads: code]
Counter-example
if hasattr(Obj, "m"): print(inspect.getsource(Obj.m)), or m = getattr(Obj, "m", None) followed by a None check, or the whole block wrapped in try/except (AttributeError, TypeError, OSError).
Consequence
Terminates with AttributeError (member absent) or TypeError/OSError from inspect.getsource (member is not a source-backed object, e.g. a slot, builtin, or dynamically created attribute); all output after that point — including any real work — is never produced.
Evidence
A snippet that printed hasattr(...) results for two members and then called inspect.getsource(Cls.other_member) unguarded produced Traceback (most recent call last): after emitting only the first lines of output.
id f3532b74318e · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the reflection block: calls to `hasattr`/`getattr` on a class or module object [reads: code]",
 "prediction": "Terminates with `AttributeError` (member absent) or `TypeError`/`OSError` from `inspect.getsource` (member is not a source-backed object, e.g. a slot, builtin, or dynamically created attribute); all output after that point \u2014 including any real work \u2014 is never produced."
}
raw text (what the judge reads)
### Unguarded attribute access alongside hasattr probes of sibling members
- **Applies when**: `code`: the program reflects over a class or module whose API it is not certain about, using `hasattr`, `getattr`, `dir`, or `inspect.getsource`
- **Pattern**: The program probes some members defensively with `hasattr(...)` but then dereferences another member of the same object directly (e.g. passing `Cls.member` to `inspect.getsource`) with no existence check and no `try/except`, so a missing or non-source-backed member aborts the whole script.
- **Detection procedure**:
  1. Locate the reflection block: calls to `hasattr`/`getattr` on a class or module object [reads: code]
  2. In the same block, find a direct dotted access on that same object (`Obj.name`) used as an argument to `inspect.getsource`, `inspect.signature`, or similar [reads: code]
  3. Check whether that direct access is inside `if hasattr(...)`/`try: ... except AttributeError`/`getattr(obj, name, default)`; the defect is present when it is bare while sibling names are hasattr-guarded [reads: code]
- **Counter-example**: `if hasattr(Obj, "m"): print(inspect.getsource(Obj.m))`, or `m = getattr(Obj, "m", None)` followed by a `None` check, or the whole block wrapped in `try/except (AttributeError, TypeError, OSError)`.
- **Consequence**: Terminates with `AttributeError` (member absent) or `TypeError`/`OSError` from `inspect.getsource` (member is not a source-backed object, e.g. a slot, builtin, or dynamically created attribute); all output after that point — including any real work — is never produced.
- **Evidence**: A snippet that printed `hasattr(...)` results for two members and then called `inspect.getsource(Cls.other_member)` unguarded produced `Traceback (most recent call last):` after emitting only the first lines of output.
3Dotted package path built through a name that is a module file, not a packagecodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program passes a dotted, package-style string to an import/resource API (importlib.resources.files, importlib.import_module, pkgutil.get_data, or a plugin constructor taking a package path)
Pattern
The program treats a name that exists on disk as a single .py module as if it were a package, appending further dotted components below it, so the path can never resolve.
Detection procedure
  1. Locate every string literal in the program that is used as a dotted package/module path (argument to import_module, resources.files, or a constructor parameter named path/package/module). [reads: code]
  2. Split each literal on . and match the leading components against the repository tree in the static facts; only decide when the components are shallow enough to appear in the listed tree, otherwise stop and do not fire. [reads: static facts — repo tree listing]
  3. Fire if some component matches an entry listed as a <name>.py file while further dotted components follow it in the literal. [reads: code]
Counter-example
A dotted literal whose components all match listed directories, or one that terminates exactly at the <name>.py module (e.g. import_module("pkg.module")) — resolvable and safe.
Discriminator
Extra dotted components appended after a component the static tree shows as a .py file; the safe case appends nothing past the module or traverses only directories.
Consequence
ModuleNotFoundError or ImportError at resolution, or TypeError/FileNotFoundError from importlib.resources.files(...); if the API catches the lookup, an empty/absent-resource result is reported as a spurious negative.
Evidence
path='tests.test_config_repository.config_without_group.conf_with_defaults' was passed as a package path while the tree lists tests/test_config_repository.py as a file; the call never resolved a resource and the run aborted with an argument-validation ValueError.
id ed19de64f464 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate every string literal in the program that is used as a dotted package/module path (argument to `import_module`, `resources.files`, or a constructor parameter named `path`/`package`/`module`). [reads: code]",
 "prediction": "`ModuleNotFoundError` or `ImportError` at resolution, or `TypeError`/`FileNotFoundError` from `importlib.resources.files(...)`; if the API catches the lookup, an empty/absent-resource result is reported as a spurious negative."
}
raw text (what the judge reads)
### Dotted package path built through a name that is a module file, not a package
- **Applies when**: `code`: the program passes a dotted, package-style string to an import/resource API (`importlib.resources.files`, `importlib.import_module`, `pkgutil.get_data`, or a plugin constructor taking a package path)
- **Pattern**: The program treats a name that exists on disk as a single `.py` module as if it were a package, appending further dotted components below it, so the path can never resolve.
- **Detection procedure**:
  1. Locate every string literal in the program that is used as a dotted package/module path (argument to `import_module`, `resources.files`, or a constructor parameter named `path`/`package`/`module`). [reads: code]
  2. Split each literal on `.` and match the leading components against the repository tree in the static facts; only decide when the components are shallow enough to appear in the listed tree, otherwise stop and do not fire. [reads: static facts — repo tree listing]
  3. Fire if some component matches an entry listed as a `<name>.py` file while further dotted components follow it in the literal. [reads: code]
- **Counter-example**: A dotted literal whose components all match listed directories, or one that terminates exactly at the `<name>.py` module (e.g. `import_module("pkg.module")`) — resolvable and safe.
- **Discriminator**: Extra dotted components appended *after* a component the static tree shows as a `.py` file; the safe case appends nothing past the module or traverses only directories.
- **Consequence**: `ModuleNotFoundError` or `ImportError` at resolution, or `TypeError`/`FileNotFoundError` from `importlib.resources.files(...)`; if the API catches the lookup, an empty/absent-resource result is reported as a spurious negative.
- **Evidence**: `path='tests.test_config_repository.config_without_group.conf_with_defaults'` was passed as a package path while the tree lists `tests/test_config_repository.py` as a file; the call never resolved a resource and the run aborted with an argument-validation `ValueError`.
3Whole run hinges on an unestablished hardcoded fixture pathcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program's main path loads exactly one hardcoded resource/fixture identifier (a filename, package-qualified resource string such as pkg://a.b.c, or module path) and everything afterwards depends on that load succeeding
Pattern
The program invents a specific asset name that is established nowhere — not in the task statement, not in the file listing it was given — and calls the loader on it with no existence check, no fallback, and no try/except. The first call raises and the process dies before producing any of its intended output.
Detection procedure
  1. Locate the literal resource identifier passed to the loader/constructor (string filename, pkg:///dotted package path, directory join). [reads: code]
  2. Search the task statement and the data/repo listing in the static facts for that exact literal (or the directory it names, at the depth the listing shows). [reads: task and static facts — the repo/data file listing]
  3. Fires if the literal appears in neither source and the program has no os.path.exists/Path.exists/glob/directory-enumeration guard and no try/except around the load, so a wrong name cannot be recovered from. [reads: code]
Counter-example
A program that enumerates candidates (sorted(Path(d).glob("*.yaml")), importlib.resources.files(pkg).iterdir()) and picks the first, or that uses a path quoted verbatim in the task statement, or that wraps the load in try/except and falls back — the same loader call, but the name is derived or recoverable.
Discriminator
The identifier is author-invented (absent from task text and from the provided listing) and unguarded; a derived, listed, or guarded identifier does not fire.
Consequence
Terminal FileNotFoundError, most likely, then ModuleNotFoundError/ImportError for package-style resource paths, or a loader-specific lookup error; the process exits nonzero with only the prints emitted before the load, so none of the intended verification is produced. Explains the abrupt traceback after otherwise-passing collected tests; it does not by itself explain any missing functional change.
Evidence
config_source.load_config('<hardcoded-fixture-name>.yaml') on a package path invented by the author, with no existence check or exception handling; the run ended in Traceback (most recent call last): immediately after unrelated tests passed.
id 174bdaa035c9 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the literal resource identifier passed to the loader/constructor (string filename, `pkg://`/dotted package path, directory join). [reads: code]",
 "prediction": "Terminal `FileNotFoundError`, most likely, then `ModuleNotFoundError`/`ImportError` for package-style resource paths, or a loader-specific lookup error; the process exits nonzero with only the prints emitted before the load, so none of the intended verification is produced. Explains the abrupt traceback after otherwise-passing collected tests; it does not by itself explain any missing functional change."
}
raw text (what the judge reads)
### Whole run hinges on an unestablished hardcoded fixture path
- **Applies when**: `code`: the program's main path loads exactly one hardcoded resource/fixture identifier (a filename, package-qualified resource string such as `pkg://a.b.c`, or module path) and everything afterwards depends on that load succeeding
- **Pattern**: The program invents a specific asset name that is established nowhere — not in the task statement, not in the file listing it was given — and calls the loader on it with no existence check, no fallback, and no `try/except`. The first call raises and the process dies before producing any of its intended output.
- **Detection procedure**:
  1. Locate the literal resource identifier passed to the loader/constructor (string filename, `pkg://`/dotted package path, directory join). [reads: code]
  2. Search the task statement and the data/repo listing in the static facts for that exact literal (or the directory it names, at the depth the listing shows). [reads: task and static facts — the repo/data file listing]
  3. Fires if the literal appears in neither source **and** the program has no `os.path.exists`/`Path.exists`/`glob`/directory-enumeration guard and no `try/except` around the load, so a wrong name cannot be recovered from. [reads: code]
- **Counter-example**: A program that enumerates candidates (`sorted(Path(d).glob("*.yaml"))`, `importlib.resources.files(pkg).iterdir()`) and picks the first, or that uses a path quoted verbatim in the task statement, or that wraps the load in `try/except` and falls back — the same loader call, but the name is derived or recoverable.
- **Discriminator**: The identifier is author-invented (absent from task text and from the provided listing) **and** unguarded; a derived, listed, or guarded identifier does not fire.
- **Consequence**: Terminal `FileNotFoundError`, most likely, then `ModuleNotFoundError`/`ImportError` for package-style resource paths, or a loader-specific lookup error; the process exits nonzero with only the prints emitted before the load, so none of the intended verification is produced. Explains the abrupt traceback after otherwise-passing collected tests; it does not by itself explain any missing functional change.
- **Evidence**: `config_source.load_config('<hardcoded-fixture-name>.yaml')` on a package path invented by the author, with no existence check or exception handling; the run ended in `Traceback (most recent call last):` immediately after unrelated tests passed.
3File contents passed to a loader whose string argument means a pathcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program reads a resource/file into memory (read_text(), read(), decode()) and hands the result to a parsing/loading API
Pattern
A refactor replaces an open file object/stream with the fully-read text, but the receiving API overloads its first argument as either a filesystem path (when str/PathLike) or a stream. The contents string is then interpreted as a path, so parsing silently produces the wrong object or raises, instead of parsing the text.
Detection procedure
  1. Locate assignments whose right-hand side is X.read_text(...), X.read(), or bytes.decode(...), and follow the variable to its first use. [reads: code]
  2. Identify the API it is passed to and check, against the installed package list, whether that API's str argument denotes a filename rather than content: e.g. OmegaConf.load, pandas.read_csv/read_json, PIL.Image.open, configparser.read (vs read_string), json.load (vs loads). [reads: static facts — python packages]
  3. Fire when the in-memory text is passed positionally to such a path-accepting loader with no wrapping in io.StringIO/io.BytesIO and no switch to the content-taking variant (*_string, loads, safe_load). [reads: code]
Counter-example
yaml.safe_load(text), json.loads(text), configparser.read_string(text), or OmegaConf.load(io.StringIO(text)) — APIs that are documented to consume content, or content wrapped in a stream before the call.
Consequence
Depending on the library, a wrong-but-silent artifact (an empty/default config, a one-row frame) or FileNotFoundError, OSError: [Errno 36] File name too long, ValueError, AttributeError: 'str' object has no attribute 'read'. Downstream behavior that depends on the parsed values (logging setup, output directories, flags) diverges from expectation without any visible parse error.
Evidence
A _read_config implementation was changed to use res.read_text() in place of handling the file stream directly, after which configuration-dependent behaviors (logging disabling, log format overrides, read-only container errors) stopped matching their expected outputs.
id fd3361d161e3 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate assignments whose right-hand side is `X.read_text(...)`, `X.read()`, or `bytes.decode(...)`, and follow the variable to its first use. [reads: code]",
 "prediction": "Depending on the library, a wrong-but-silent artifact (an empty/default config, a one-row frame) or `FileNotFoundError`, `OSError: [Errno 36] File name too long`, `ValueError`, `AttributeError: 'str' object has no attribute 'read'`. Downstream behavior that depends on the parsed values (logging setup, output directories, flags) diverges from expectation without any visible parse error."
}
raw text (what the judge reads)
### File contents passed to a loader whose string argument means a path
- **Applies when**: `code`: the program reads a resource/file into memory (`read_text()`, `read()`, `decode()`) and hands the result to a parsing/loading API
- **Pattern**: A refactor replaces an open file object/stream with the fully-read text, but the receiving API overloads its first argument as *either* a filesystem path (when `str`/`PathLike`) *or* a stream. The contents string is then interpreted as a path, so parsing silently produces the wrong object or raises, instead of parsing the text.
- **Detection procedure**:
  1. Locate assignments whose right-hand side is `X.read_text(...)`, `X.read()`, or `bytes.decode(...)`, and follow the variable to its first use. [reads: code]
  2. Identify the API it is passed to and check, against the installed package list, whether that API's `str` argument denotes a filename rather than content: e.g. `OmegaConf.load`, `pandas.read_csv/read_json`, `PIL.Image.open`, `configparser.read` (vs `read_string`), `json.load` (vs `loads`). [reads: static facts — python packages]
  3. Fire when the in-memory text is passed positionally to such a path-accepting loader with no wrapping in `io.StringIO`/`io.BytesIO` and no switch to the content-taking variant (`*_string`, `loads`, `safe_load`). [reads: code]
- **Counter-example**: `yaml.safe_load(text)`, `json.loads(text)`, `configparser.read_string(text)`, or `OmegaConf.load(io.StringIO(text))` — APIs that are documented to consume content, or content wrapped in a stream before the call.
- **Consequence**: Depending on the library, a wrong-but-silent artifact (an empty/default config, a one-row frame) or `FileNotFoundError`, `OSError: [Errno 36] File name too long`, `ValueError`, `AttributeError: 'str' object has no attribute 'read'`. Downstream behavior that depends on the parsed values (logging setup, output directories, flags) diverges from expectation without any visible parse error.
- **Evidence**: A `_read_config` implementation was changed to use `res.read_text()` in place of handling the file stream directly, after which configuration-dependent behaviors (logging disabling, log format overrides, read-only container errors) stopped matching their expected outputs.
3Overriding a key on a struct-mode config whose key set was never establishedcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program loads configuration/structured data from a file it does not define inline and then applies key=value overrides or assignments to the loaded object (e.g. Hydra compose(overrides=[...]), OmegaConf.update, attribute/__setattr__ on a struct container)
Pattern
A literal key name is written into a container that rejects unknown keys, while the program never verified that the key exists in the loaded content — the assumption about the file's internal layout is untested and unguarded.
Detection procedure
  1. Locate the override list / update call and note the literal key names on the left-hand side, and note the config_name/path whose contents supply them. [reads: code]
  2. Check whether that file's key layout is knowable from the program text or the provided facts: is the schema declared inline (dataclass/dict literal/OmegaConf.structured), or is it only an on-disk path (the repo listing stops above it)? [reads: code; static facts — repo tree]
  3. Fires if the key set comes only from disk and the override lacks the append form (+key=/++key=) or a prior existence check (in cfg, hasattr, OmegaConf.set_struct(cfg, False)), and the call is not wrapped in except — extra red flag if the config_name contains a path separator, i.e. a config-group member is being composed as the primary config, in which case the file's keys may not even land at the root. [reads: code]
Counter-example
The same override list where entries are prefixed with +, or where the program first prints/inspects the composed config keys, or where the schema for the key is declared in the program's own source — none of these can raise on an unknown key.
Discriminator
The target container is in struct/strict mode and the key's presence is asserted by the program rather than checked or forced; safe code either appends, disables struct, checks membership, or owns the schema.
Consequence
Terminates with hydra.errors.ConfigCompositionException wrapping omegaconf.errors.ConfigAttributeError ("Key ... is not in struct"), or KeyError/AttributeError for other struct-like targets, before any subsequent check in the script runs — so later prints never execute and the script exits nonzero. Accounts for the immediate crash; the remaining gap is that the script targets the wrong code path entirely.
Evidence
compose(config_name="dataset/imagenet", overrides=["name=custom_dataset"]) with only a try/finally (no except) raised ConfigAttributeError: Key 'name' is not in struct and aborted the whole script.
id 228b6ff7eae0 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the override list / update call and note the literal key names on the left-hand side, and note the `config_name`/path whose contents supply them. [reads: code]",
 "prediction": "Terminates with `hydra.errors.ConfigCompositionException` wrapping `omegaconf.errors.ConfigAttributeError` (\"Key ... is not in struct\"), or `KeyError`/`AttributeError` for other struct-like targets, before any subsequent check in the script runs \u2014 so later prints never execute and the script exits nonzero. Accounts for the immediate crash; the remaining gap is that the script targets the wrong code path entirely."
}
raw text (what the judge reads)
### Overriding a key on a struct-mode config whose key set was never established
- **Applies when**: `code`: the program loads configuration/structured data from a file it does not define inline and then applies key=value overrides or assignments to the loaded object (e.g. Hydra `compose(overrides=[...])`, `OmegaConf.update`, attribute/`__setattr__` on a struct container)
- **Pattern**: A literal key name is written into a container that rejects unknown keys, while the program never verified that the key exists in the loaded content — the assumption about the file's internal layout is untested and unguarded.
- **Detection procedure**:
  1. Locate the override list / update call and note the literal key names on the left-hand side, and note the `config_name`/path whose contents supply them. [reads: code]
  2. Check whether that file's key layout is knowable from the program text or the provided facts: is the schema declared inline (dataclass/dict literal/`OmegaConf.structured`), or is it only an on-disk path (the repo listing stops above it)? [reads: code; static facts — repo tree]
  3. Fires if the key set comes only from disk **and** the override lacks the append form (`+key=`/`++key=`) or a prior existence check (`in cfg`, `hasattr`, `OmegaConf.set_struct(cfg, False)`), **and** the call is not wrapped in `except` — extra red flag if the `config_name` contains a path separator, i.e. a config-group member is being composed as the primary config, in which case the file's keys may not even land at the root. [reads: code]
- **Counter-example**: The same override list where entries are prefixed with `+`, or where the program first prints/inspects the composed config keys, or where the schema for the key is declared in the program's own source — none of these can raise on an unknown key.
- **Discriminator**: The target container is in struct/strict mode and the key's presence is asserted by the program rather than checked or forced; safe code either appends, disables struct, checks membership, or owns the schema.
- **Consequence**: Terminates with `hydra.errors.ConfigCompositionException` wrapping `omegaconf.errors.ConfigAttributeError` ("Key ... is not in struct"), or `KeyError`/`AttributeError` for other struct-like targets, before any subsequent check in the script runs — so later prints never execute and the script exits nonzero. Accounts for the immediate crash; the remaining gap is that the script targets the wrong code path entirely.
- **Evidence**: `compose(config_name="dataset/imagenet", overrides=["name=custom_dataset"])` with only a `try/finally` (no `except`) raised `ConfigAttributeError: Key 'name' is not in struct` and aborted the whole script.
3Unverified key access on a parsed-config/data objectcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program loads a configuration or data file through a library (e.g. compose/OmegaConf.load, json.load, yaml.safe_load, a dataframe reader) and then reads named fields off the result.
Pattern
The program reads a field by bare attribute/index access using a name it never established exists — the name is not written by the program, not confirmed by any listed schema/column fact, and not guarded — so a single missing key aborts the whole run instead of reporting it.
Detection procedure
  1. Locate every bare field read on the loaded object (obj.name, obj["name"]) and note each name. [reads: code]
  2. For each name, check whether it is either assigned/injected earlier in the program (e.g. added by an override, a computed column, an explicit default) or appears in the static facts as a documented field/column of the file being read. [reads: code + static facts (file/column listing)]
  3. Check whether the read is wrapped in a guard: .get(...) with default, if "name" in obj, OmegaConf.select, or a try/except (KeyError, AttributeError, ConfigAttributeError). Fire if a name fails step 2 and has no guard in step 3, and the read sits on the success path of a script that prints "passed"/"✓" afterwards. [reads: code]
Counter-example
A program that first dumps or enumerates the loaded structure (print(OmegaConf.to_yaml(cfg)), list(obj.keys()), df.columns) and only then reads names, or that reads with cfg.get("name", None) / if "name" in cfg, so a missing key degrades to a printed None rather than an abort.
Discriminator
Goes wrong when the field name is assumed from domain knowledge about the file rather than produced by the program or verified at runtime, and the access is unguarded; safe when the same access is preceded by enumeration/in check or uses a defaulted getter.
Consequence
Terminates with omegaconf.errors.ConfigAttributeError (or ConfigKeyError, KeyError, AttributeError, pandas.errors.UndefinedVariableError depending on the library) partway through, after some output has already been printed; every later check in the script — including the one that actually matters — never executes and the run exits non-zero.
Evidence
A script composed a config, successfully read a key it had itself injected via an override, then did a bare read of a second key assumed to come from the file; that read raised ConfigAttributeError: Key '<name>' is not in struct and killed the script before its final validation line.
id 9d41764c5122 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate every bare field read on the loaded object (`obj.name`, `obj[\"name\"]`) and note each name. [reads: code]",
 "prediction": "Terminates with `omegaconf.errors.ConfigAttributeError` (or `ConfigKeyError`, `KeyError`, `AttributeError`, `pandas.errors.UndefinedVariableError` depending on the library) partway through, after some output has already been printed; every later check in the script \u2014 including the one that actually matters \u2014 never executes and the run exits non-zero."
}
raw text (what the judge reads)
### Unverified key access on a parsed-config/data object
- **Applies when**: `code`: the program loads a configuration or data file through a library (e.g. `compose`/`OmegaConf.load`, `json.load`, `yaml.safe_load`, a dataframe reader) and then reads named fields off the result.
- **Pattern**: The program reads a field by bare attribute/index access using a name it never established exists — the name is not written by the program, not confirmed by any listed schema/column fact, and not guarded — so a single missing key aborts the whole run instead of reporting it.
- **Detection procedure**:
  1. Locate every bare field read on the loaded object (`obj.name`, `obj["name"]`) and note each name. [reads: code]
  2. For each name, check whether it is either assigned/injected earlier in the program (e.g. added by an override, a computed column, an explicit default) or appears in the static facts as a documented field/column of the file being read. [reads: code + static facts (file/column listing)]
  3. Check whether the read is wrapped in a guard: `.get(...)` with default, `if "name" in obj`, `OmegaConf.select`, or a `try/except (KeyError, AttributeError, ConfigAttributeError)`. Fire if a name fails step 2 and has no guard in step 3, and the read sits on the success path of a script that prints "passed"/"✓" afterwards. [reads: code]
- **Counter-example**: A program that first dumps or enumerates the loaded structure (`print(OmegaConf.to_yaml(cfg))`, `list(obj.keys())`, `df.columns`) and only then reads names, or that reads with `cfg.get("name", None)` / `if "name" in cfg`, so a missing key degrades to a printed `None` rather than an abort.
- **Discriminator**: Goes wrong when the field name is *assumed* from domain knowledge about the file rather than produced by the program or verified at runtime, and the access is unguarded; safe when the same access is preceded by enumeration/`in` check or uses a defaulted getter.
- **Consequence**: Terminates with `omegaconf.errors.ConfigAttributeError` (or `ConfigKeyError`, `KeyError`, `AttributeError`, `pandas.errors.UndefinedVariableError` depending on the library) partway through, after some output has already been printed; every later check in the script — including the one that actually matters — never executes and the run exits non-zero.
- **Evidence**: A script composed a config, successfully read a key it had itself injected via an override, then did a bare read of a second key assumed to come from the file; that read raised `ConfigAttributeError: Key '<name>' is not in struct` and killed the script before its final validation line.
3Verification exercises sibling APIs instead of the implicated onecodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program instantiates or imports a class/module that the task statement explicitly names as the site of the defect, and then calls methods on it to check behavior
Pattern
The task names a specific function, method, or code path as the suspected cause, but the program only invokes other members of the same object — cheap metadata/existence queries — and never triggers the named path with real inputs. The checks all agree, and the untested path stays broken.
Detection procedure
  1. From the task statement, extract the specific method/function name or operation blamed for the defect (e.g., a quoted method name, or the operation described as "now uses X instead of Y"). [reads: task]
  2. List every method the program calls on the implicated object or module. [reads: code]
  3. The rubric fires when the blamed method — and any public wrapper whose stated job is to perform that read/compute/load operation — appears nowhere in that list, while only unrelated predicate/listing methods are called. [reads: code]
Counter-example
A program that calls the blamed method (or the documented public entry point that delegates to it) on an input drawn from the task's reproduction steps and compares its returned content against an expectation.
Discriminator
Whether the code path named in the task statement is actually executed by the program, not merely whether some method of the same class is.
Consequence
The program reports agreement/success while the defective path is never entered; the issue's failing scenarios continue to fail. Expect the fix-related tests to remain red and any self-reported "verification complete" to be a false pass.
Evidence
An issue attributing the fault to a content-reading method was "verified" by a script calling only existence/listing predicates (is_config, is_group, available); the content-reading method was never invoked.
id ca320212bc34 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. From the task statement, extract the specific method/function name or operation blamed for the defect (e.g., a quoted method name, or the operation described as \"now uses X instead of Y\"). [reads: task]",
 "prediction": "The program reports agreement/success while the defective path is never entered; the issue's failing scenarios continue to fail. Expect the fix-related tests to remain red and any self-reported \"verification complete\" to be a false pass."
}
raw text (what the judge reads)
### Verification exercises sibling APIs instead of the implicated one
- **Applies when**: `code`: the program instantiates or imports a class/module that the task statement explicitly names as the site of the defect, and then calls methods on it to check behavior
- **Pattern**: The task names a specific function, method, or code path as the suspected cause, but the program only invokes *other* members of the same object — cheap metadata/existence queries — and never triggers the named path with real inputs. The checks all agree, and the untested path stays broken.
- **Detection procedure**:
  1. From the task statement, extract the specific method/function name or operation blamed for the defect (e.g., a quoted method name, or the operation described as "now uses X instead of Y"). [reads: task]
  2. List every method the program calls on the implicated object or module. [reads: code]
  3. The rubric fires when the blamed method — and any public wrapper whose stated job is to perform that read/compute/load operation — appears nowhere in that list, while only unrelated predicate/listing methods are called. [reads: code]
- **Counter-example**: A program that calls the blamed method (or the documented public entry point that delegates to it) on an input drawn from the task's reproduction steps and compares its returned content against an expectation.
- **Discriminator**: Whether the code path named in the task statement is actually executed by the program, not merely whether some method of the same class is.
- **Consequence**: The program reports agreement/success while the defective path is never entered; the issue's failing scenarios continue to fail. Expect the fix-related tests to remain red and any self-reported "verification complete" to be a false pass.
- **Evidence**: An issue attributing the fault to a content-reading method was "verified" by a script calling only existence/listing predicates (`is_config`, `is_group`, `available`); the content-reading method was never invoked.
3Assumed nesting: accessing a parsed config/record by a key copied from the resource pathcodeswesmith/facebookresearch__hydra.0f03eb60
Applies when
code: the program loads a structured file (YAML/JSON/config/record) by name or path and then reads a specific nested key/attribute out of the returned object
Pattern
The program assumes that the directory or path component used to address a resource also appears as a top-level key inside the parsed content, and dereferences it directly (obj.<segment> or obj["<segment>"]) without any membership check, so any file whose content is stored unnested (or re-rooted by a packaging/namespace directive) raises immediately.
Detection procedure
  1. Locate every call that loads/parses a resource by a string path or name and note the string's components (e.g. load(...'a/b'), read('dir/name')). [reads: code]
  2. Locate the subsequent attribute or item access on the returned object and compare the accessed key text with the path components from step 1 — flag when the key equals a directory segment of the load argument. [reads: code]
  3. Confirm there is no guard around that access: no if key in cfg, no .get(key, default), no hasattr, no try/except around it — and that nothing in the task statement or the static facts enumerates the file's internal keys, so the nesting is unverified. [reads: code, then task + static facts (repo tree / file column listings)]
Counter-example
Code that does cfg = load('dir/name') and then cfg.get('dir') / if 'dir' in cfg: / wraps the access in try/except (KeyError, AttributeError), or that accesses only keys the program itself just wrote into the object.
Discriminator
The dereferenced key name is taken from the load path rather than from anything the program created or the facts confirm, and no membership/get/try guard exists. Guarded or self-written keys do not fire.
Consequence
Terminates at that line with omegaconf.errors.ConfigAttributeError / ConfigKeyError, KeyError, AttributeError, or TypeError ("not subscriptable") for files whose content is not nested under the path segment; everything after that line is never executed, so the run yields no result for the remaining work.
Evidence
result.config.dataset after load_config('dataset/imagenet') raised omegaconf.errors.ConfigAttributeError: Missing key dataset (object_type=dict), aborting the script before its later checks.
id 35fd0297ba67 · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate every call that loads/parses a resource by a string path or name and note the string's components (e.g. `load(...'a/b')`, `read('dir/name')`). [reads: code]",
 "prediction": "Terminates at that line with `omegaconf.errors.ConfigAttributeError` / `ConfigKeyError`, `KeyError`, `AttributeError`, or `TypeError` (\"not subscriptable\") for files whose content is not nested under the path segment; everything after that line is never executed, so the run yields no result for the remaining work."
}
raw text (what the judge reads)
### Assumed nesting: accessing a parsed config/record by a key copied from the resource path
- **Applies when**: `code`: the program loads a structured file (YAML/JSON/config/record) by name or path and then reads a specific nested key/attribute out of the returned object
- **Pattern**: The program assumes that the directory or path component used to *address* a resource also appears as a top-level key *inside* the parsed content, and dereferences it directly (`obj.<segment>` or `obj["<segment>"]`) without any membership check, so any file whose content is stored unnested (or re-rooted by a packaging/namespace directive) raises immediately.
- **Detection procedure**:
  1. Locate every call that loads/parses a resource by a string path or name and note the string's components (e.g. `load(...'a/b')`, `read('dir/name')`). [reads: code]
  2. Locate the subsequent attribute or item access on the returned object and compare the accessed key text with the path components from step 1 — flag when the key equals a directory segment of the load argument. [reads: code]
  3. Confirm there is no guard around that access: no `if key in cfg`, no `.get(key, default)`, no `hasattr`, no `try/except` around it — and that nothing in the task statement or the static facts enumerates the file's internal keys, so the nesting is unverified. [reads: code, then task + static facts (repo tree / file column listings)]
- **Counter-example**: Code that does `cfg = load('dir/name')` and then `cfg.get('dir')` / `if 'dir' in cfg:` / wraps the access in `try/except (KeyError, AttributeError)`, or that accesses only keys the program itself just wrote into the object.
- **Discriminator**: The dereferenced key name is taken from the load path rather than from anything the program created or the facts confirm, **and** no membership/`get`/`try` guard exists. Guarded or self-written keys do not fire.
- **Consequence**: Terminates at that line with `omegaconf.errors.ConfigAttributeError` / `ConfigKeyError`, `KeyError`, `AttributeError`, or `TypeError` ("not subscriptable") for files whose content is not nested under the path segment; everything after that line is never executed, so the run yields no result for the remaining work.
- **Evidence**: `result.config.dataset` after `load_config('dataset/imagenet')` raised `omegaconf.errors.ConfigAttributeError: Missing key dataset (object_type=dict)`, aborting the script before its later checks.
4Membership test against a mapping using an object that cannot be hashedcodeswesmith/seperman__deepdiff.ed252022
Applies when
code: the program performs x in y, y[x], y.get(x), set.add(x) or similar where y is a dict/set-like container and x may be a mutable container
Pattern
A value that the program itself constructed as a list, dict, set, or other unhashable object is used as a lookup key, on the assumption that a library's cache/index is keyed by the object rather than by an id, hash string, or wrapper.
Detection procedure
  1. Find each in/subscript/.get() whose right-hand operand is a dictionary, set, or a library attribute documented/named as a mapping or cache. [reads: code]
  2. Trace the left-hand operand back to its binding in the same file: is it a literal [...], {...} (dict/set), or the result of an operation returning one? [reads: code]
  3. Check whether the lookup is wrapped in try/except TypeError, preceded by an isinstance/Hashable check, or converted with tuple(...)/id(...)/str(...) before use. [reads: code]
Counter-example
The same x in mapping where x is bound to a string, int, tuple, or frozenset, or where the lookup sits inside try: ... except TypeError: — both are safe even against a hash-keyed container.
Consequence
TypeError: unhashable type: 'list' (or 'dict', 'set') raised at that line, terminating the script and, if the file is collected by pytest, turning into a collection/test error.
Evidence
if item in diff.hashes: where item iterated over locally built list literals raised TypeError: unhashable type: 'list', aborting the script before its remaining diagnostics ran.
id 9c32cdaff6b6 · mined from swesmith/seperman__deepdiff.ed252022 seperman__deepdiff.ed252022.func_basic__h86ne3ds
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Find each `in`/subscript/`.get()` whose right-hand operand is a dictionary, set, or a library attribute documented/named as a mapping or cache. [reads: code]",
 "prediction": "`TypeError: unhashable type: 'list'` (or `'dict'`, `'set'`) raised at that line, terminating the script and, if the file is collected by pytest, turning into a collection/test error."
}
raw text (what the judge reads)
### Membership test against a mapping using an object that cannot be hashed

- **Applies when**: `code`: the program performs `x in y`, `y[x]`, `y.get(x)`, `set.add(x)` or similar where `y` is a dict/set-like container and `x` may be a mutable container
- **Pattern**: A value that the program itself constructed as a `list`, `dict`, `set`, or other unhashable object is used as a lookup key, on the assumption that a library's cache/index is keyed by the object rather than by an id, hash string, or wrapper.
- **Detection procedure**:
  1. Find each `in`/subscript/`.get()` whose right-hand operand is a dictionary, set, or a library attribute documented/named as a mapping or cache. [reads: code]
  2. Trace the left-hand operand back to its binding in the same file: is it a literal `[...]`, `{...}` (dict/set), or the result of an operation returning one? [reads: code]
  3. Check whether the lookup is wrapped in `try/except TypeError`, preceded by an `isinstance`/`Hashable` check, or converted with `tuple(...)`/`id(...)`/`str(...)` before use. [reads: code]
- **Counter-example**: The same `x in mapping` where `x` is bound to a string, int, tuple, or frozenset, or where the lookup sits inside `try: ... except TypeError:` — both are safe even against a hash-keyed container.
- **Consequence**: `TypeError: unhashable type: 'list'` (or `'dict'`, `'set'`) raised at that line, terminating the script and, if the file is collected by pytest, turning into a collection/test error.
- **Evidence**: `if item in diff.hashes:` where `item` iterated over locally built list literals raised `TypeError: unhashable type: 'list'`, aborting the script before its remaining diagnostics ran.
4Verification code checks only "ran without error / key exists" instead of the expected values given in the tasktaskswesmith/seperman__deepdiff.ed252022
Applies when
task: the task statement names concrete expected outputs (numeric values, string prefixes, exact structures) for specific inputs, and code: the program contains scripts or tests exercising those inputs
Pattern
The program's own checks assert only that a call succeeded, that a result key is present, or print the value for human reading, never comparing the produced value to the expected value the task supplies — so a wrong or unchanged result is scored as success by the program itself.
Detection procedure
  1. Extract from the task statement the concrete expected outputs and the inputs that produce them [reads: task]
  2. Locate the program's verification code: the scripts or test functions that call the API with those inputs [reads: code]
  3. Check whether any of those call sites compares the returned value to the expected literal (an assert, an equality/startswith/tolerance check); if every check is a print(...), a key in result membership test, or a bare try/except around the call, the condition holds [reads: code]
Counter-example
A script that prints results for diagnosis but also contains at least one assert result == expected / assert str(result).startswith(expected_prefix) for the task's stated values.
Consequence
Regressions and non-fixes pass the author's self-check silently; predict that graded tests comparing exact expected values fail with AssertionError while the program's own output looks clean. Explains why the missing/incorrect change went undetected rather than the missing change itself — a secondary share of the observed outcome.
Evidence
The "comprehensive" harness only compared '<result_key>' in diff against a boolean expectation and printed the value; the task-stated expected numbers were echoed in print strings (Expected: distance should start with 0.14) but never asserted, so no check could ever fail.
id 1ac34ff797ea · mined from swesmith/seperman__deepdiff.ed252022 seperman__deepdiff.ed252022.func_basic__h86ne3ds
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Extract from the task statement the concrete expected outputs and the inputs that produce them [reads: task]",
 "prediction": "Regressions and non-fixes pass the author's self-check silently; predict that graded tests comparing exact expected values fail with `AssertionError` while the program's own output looks clean. Explains why the missing/incorrect change went undetected rather than the missing change itself \u2014 a secondary share of the observed outcome."
}
raw text (what the judge reads)
### Verification code checks only "ran without error / key exists" instead of the expected values given in the task
- **Applies when**: `task`: the task statement names concrete expected outputs (numeric values, string prefixes, exact structures) for specific inputs, and `code`: the program contains scripts or tests exercising those inputs
- **Pattern**: The program's own checks assert only that a call succeeded, that a result key is present, or print the value for human reading, never comparing the produced value to the expected value the task supplies — so a wrong or unchanged result is scored as success by the program itself.
- **Detection procedure**:
  1. Extract from the task statement the concrete expected outputs and the inputs that produce them [reads: task]
  2. Locate the program's verification code: the scripts or test functions that call the API with those inputs [reads: code]
  3. Check whether any of those call sites compares the returned value to the expected literal (an `assert`, an equality/startswith/tolerance check); if every check is a `print(...)`, a `key in result` membership test, or a bare `try/except` around the call, the condition holds [reads: code]
- **Counter-example**: A script that prints results for diagnosis but also contains at least one `assert result == expected` / `assert str(result).startswith(expected_prefix)` for the task's stated values.
- **Consequence**: Regressions and non-fixes pass the author's self-check silently; predict that graded tests comparing exact expected values fail with `AssertionError` while the program's own output looks clean. Explains why the missing/incorrect change went undetected rather than the missing change itself — a secondary share of the observed outcome.
- **Evidence**: The "comprehensive" harness only compared `'<result_key>' in diff` against a boolean expectation and printed the value; the task-stated expected numbers were echoed in `print` strings (`Expected: distance should start with 0.14`) but never asserted, so no check could ever fail.
4Self-verification harness with branches that pass unconditionallycodeswesmith/seperman__deepdiff.ed252022
Applies when
code: the program includes its own checking script that tallies pass/fail counts or prints ✓/✗ instead of (or alongside) real assertions
Pattern
The homemade oracle contains branches that increment the pass counter (or print success) without comparing the produced value to any expected value — e.g. cases whose expected value is None/absent, or a "result is empty, as expected" branch — so the harness reports success regardless of whether the behavior is correct.
Detection procedure
  1. Locate the loop over test cases and the variable that counts passes or the success print. [reads: code]
  2. Trace every branch that reaches the pass counter and check whether an expected value from the case tuple participates in a comparison on that branch. [reads: code]
  3. If one or more cases supply no expected value (placeholder None, empty range) and the corresponding branch records a pass unconditionally, or the "value missing" branch is treated as success, the pattern is present. [reads: code]
Counter-example
A harness where every case carries a concrete expected value or tolerance and each pass increment is guarded by a comparison against it; cases intentionally skipped are counted separately from passes.
Discriminator
Existence of at least one control-flow path from case iteration to "pass" that reads no expected value.
Consequence
The program reports its own work as verified while the reported defect is unfixed; graded/hidden tests that do compare exact values fail. Explains the false confidence that led to shipping no fix — a contributing factor rather than the whole gap; the absent source edit accounts for the failing tests themselves.
Evidence
if distance is None: print("✓ ... (as expected)"); tests_passed += 1 and cases with expected_range = None counted as passes; harness printed all-pass while the reported behavior was never corrected.
id ca3773190c2c · mined from swesmith/seperman__deepdiff.ed252022 seperman__deepdiff.ed252022.func_basic__h86ne3ds
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the loop over test cases and the variable that counts passes or the success print. [reads: code]",
 "prediction": "The program reports its own work as verified while the reported defect is unfixed; graded/hidden tests that do compare exact values fail. Explains the false confidence that led to shipping no fix \u2014 a contributing factor rather than the whole gap; the absent source edit accounts for the failing tests themselves."
}
raw text (what the judge reads)
### Self-verification harness with branches that pass unconditionally
- **Applies when**: `code`: the program includes its own checking script that tallies pass/fail counts or prints ✓/✗ instead of (or alongside) real assertions
- **Pattern**: The homemade oracle contains branches that increment the pass counter (or print success) without comparing the produced value to any expected value — e.g. cases whose expected value is `None`/absent, or a "result is empty, as expected" branch — so the harness reports success regardless of whether the behavior is correct.
- **Detection procedure**:
  1. Locate the loop over test cases and the variable that counts passes or the success print. [reads: code]
  2. Trace every branch that reaches the pass counter and check whether an expected value from the case tuple participates in a comparison on that branch. [reads: code]
  3. If one or more cases supply no expected value (placeholder `None`, empty range) and the corresponding branch records a pass unconditionally, or the "value missing" branch is treated as success, the pattern is present. [reads: code]
- **Counter-example**: A harness where every case carries a concrete expected value or tolerance and each pass increment is guarded by a comparison against it; cases intentionally skipped are counted separately from passes.
- **Discriminator**: Existence of at least one control-flow path from case iteration to "pass" that reads no expected value.
- **Consequence**: The program reports its own work as verified while the reported defect is unfixed; graded/hidden tests that do compare exact values fail. Explains the false confidence that led to shipping no fix — a contributing factor rather than the whole gap; the absent source edit accounts for the failing tests themselves.
- **Evidence**: `if distance is None: print("✓ ... (as expected)"); tests_passed += 1` and cases with `expected_range = None` counted as passes; harness printed all-pass while the reported behavior was never corrected.
5Method borrowed from a sibling class that the target class does not definecodeswesmith/paramiko__paramiko.23f92003
Applies when
code: the program calls methods on an instance of a class whose definition (or whose base class) is visible in the code under review or in a module of the repository shown in the static facts
Pattern
The program treats two different classes of the same codebase as if they shared a builder/serializer interface, and calls a method name that exists on one of them (e.g. an add/put/write* accumulator) on an instance of the other, which never defines it and has no dynamic attribute dispatch. Nothing in the program establishes that the method exists on the receiver's class.
Detection procedure
  1. List every call of the form obj.<name>(...) in the program where obj is produced by instantiating a class (X()) or is documented/annotated as an instance of a class defined in the codebase. [reads: code]
  2. For each such receiver, locate that class's class X: body (and the bodies of any base classes) in the shown source files, using the repo tree in the static facts to confirm the class lives in the project rather than in a third-party package whose source is not visible. [reads: code; static facts — repo tree / installed packages list]
  3. Fire when <name> is not defined in that class or its visible bases, the class defines no __getattr__/__getattribute__/setattr-based dynamic attributes, <name> is defined on some other class in the same codebase (the "sibling" whose API was assumed), and the call is not wrapped in hasattr(...), getattr(obj, name, default), or try/except AttributeError. [reads: code]
Counter-example
the same-looking call obj.encode() where encode is defined on the class or on a base class present in the shown source, or a call guarded by if hasattr(obj, "add"): / try: obj.add(...) except AttributeError: / routed through a coercion helper that falls back when the method is missing — these must not fire.
Discriminator
the goes-wrong case has the method name absent from the receiver class's visible definition and present on a different class in the codebase, with no guard or fallback path; the safe case has the method defined on the class/base or has an explicit existence check or exception fallback around it.
Consequence
AttributeError: '<Class>' object has no attribute '<name>' raised at that call, aborting the enclosing function; if the call sits in a linear script or a check sequence, every step after it is never executed and its results are lost. Secondary possibility is TypeError if the name resolves to a non-callable attribute.
Evidence
A run that had already passed three earlier checks terminated with AttributeError: 'BER'-style object has no attribute 'add' — the accumulator method add(...) belonged to a different serialization class in the same package, and the receiver's class defined no such method and no dynamic attribute hook; the remaining checks in that run never executed.
id d143f198d79d · mined from swesmith/paramiko__paramiko.23f92003 paramiko__paramiko.23f92003.func_pm_ctrl_shuffle__94h6o4lo
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. List every call of the form `obj.<name>(...)` in the program where `obj` is produced by instantiating a class (`X()`) or is documented/annotated as an instance of a class defined in the codebase. [reads: code]",
 "prediction": "`AttributeError: '<Class>' object has no attribute '<name>'` raised at that call, aborting the enclosing function; if the call sits in a linear script or a check sequence, every step after it is never executed and its results are lost. Secondary possibility is `TypeError` if the name resolves to a non-callable attribute."
}
raw text (what the judge reads)
### Method borrowed from a sibling class that the target class does not define
- **Applies when**: `code`: the program calls methods on an instance of a class whose definition (or whose base class) is visible in the code under review or in a module of the repository shown in the static facts
- **Pattern**: The program treats two different classes of the same codebase as if they shared a builder/serializer interface, and calls a method name that exists on one of them (e.g. an `add*`/`put*`/`write*` accumulator) on an instance of the other, which never defines it and has no dynamic attribute dispatch. Nothing in the program establishes that the method exists on the receiver's class.
- **Detection procedure**:
  1. List every call of the form `obj.<name>(...)` in the program where `obj` is produced by instantiating a class (`X()`) or is documented/annotated as an instance of a class defined in the codebase. [reads: code]
  2. For each such receiver, locate that class's `class X:` body (and the bodies of any base classes) in the shown source files, using the repo tree in the static facts to confirm the class lives in the project rather than in a third-party package whose source is not visible. [reads: code; static facts — repo tree / installed packages list]
  3. Fire when `<name>` is not defined in that class or its visible bases, the class defines no `__getattr__`/`__getattribute__`/`setattr`-based dynamic attributes, `<name>` *is* defined on some other class in the same codebase (the "sibling" whose API was assumed), and the call is not wrapped in `hasattr(...)`, `getattr(obj, name, default)`, or `try/except AttributeError`. [reads: code]
- **Counter-example**: the same-looking call `obj.encode()` where `encode` is defined on the class or on a base class present in the shown source, or a call guarded by `if hasattr(obj, "add"):` / `try: obj.add(...) except AttributeError:` / routed through a coercion helper that falls back when the method is missing — these must not fire.
- **Discriminator**: the goes-wrong case has the method name absent from the receiver class's visible definition *and* present on a different class in the codebase, with no guard or fallback path; the safe case has the method defined on the class/base or has an explicit existence check or exception fallback around it.
- **Consequence**: `AttributeError: '<Class>' object has no attribute '<name>'` raised at that call, aborting the enclosing function; if the call sits in a linear script or a check sequence, every step after it is never executed and its results are lost. Secondary possibility is `TypeError` if the name resolves to a non-callable attribute.
- **Evidence**: A run that had already passed three earlier checks terminated with `AttributeError: 'BER'-style object has no attribute 'add'` — the accumulator method `add(...)` belonged to a different serialization class in the same package, and the receiver's class defined no such method and no dynamic attribute hook; the remaining checks in that run never executed.
6Attribute read/validated but never assignedcodeswesmith/pydata__patsy.a5d16484
Applies when
code: any class whose methods (constructor, validators, __repr__, comparison helpers) read self.<name>
Pattern
A method dereferences an instance attribute that no code path ever assigns — the assignment was removed, renamed, or replaced by a same-named local variable — so the first access raises at runtime instead of returning the intended value.
Detection procedure
  1. For each class defined or edited in the program, list every attribute name read as self.<name> inside any method body (including inside if/raise conditions and passed to helper functions). [reads: code]
  2. For each such name, search the whole class body for a binding: self.<name> = ..., setattr(self, "<name>", ...), a @property/__getattr__/__slots__-with-default definition of that name, a class-level attribute of that name, or an assignment in a base class named in the class header (check whether that base is defined in the program or is a plain object). [reads: code]
  3. Flag the class if some read name has no binding anywhere, and especially if a local variable of the same name is computed in the same method (e.g. name = np.asarray(...)) but the code then reads self.name instead of the local — the normalization result is discarded and the attribute never exists. Cross-check the task statement / class docstring: if the name is a documented public attribute the caller is expected to read, the break is user-visible. [reads: code, task]
Counter-example
A constructor that writes self.constants = atleast_2d_column_default(constants) and only afterwards validates self.constants.ndim, or a subclass whose __init__ calls super().__init__() where the parent assigns the attribute, or a class exposing the name through a @property backed by self._name.
Discriminator
The failing case has zero binding sites for the name in the class, its declared bases, or a property/__getattr__; the safe cases all have exactly one reachable binding executed before the read.
Consequence
AttributeError: '<Class>' object has no attribute '<name>' raised on the first construction/use, aborting the script and every test that instantiates the class; any downstream test asserting on that documented attribute (order of a stored list, default value, dtype) fails as an error rather than a mismatch. If the attribute is only read on a rare branch, expect intermittent AttributeError instead of a wrong value.
Evidence
A constructor computed constants = np.asarray(constants, dtype=float) into a local, dropped the self.constants = ... and self.variable_names = ... assignments, then validated self.constants.ndim — producing AttributeError: 'LinearConstraint' object has no attribute 'constants' at line 1 of the reproduction script.
id 7d99b63139a9 · mined from swesmith/pydata__patsy.a5d16484 pydata__patsy.a5d16484.combine_file__xgv5bk58
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. For each class defined or edited in the program, list every attribute name read as `self.<name>` inside any method body (including inside `if`/`raise` conditions and passed to helper functions). [reads: code]",
 "prediction": "`AttributeError: '<Class>' object has no attribute '<name>'` raised on the first construction/use, aborting the script and every test that instantiates the class; any downstream test asserting on that documented attribute (order of a stored list, default value, dtype) fails as an error rather than a mismatch. If the attribute is only read on a rare branch, expect intermittent `AttributeError` instead of a wrong value."
}
raw text (what the judge reads)
### Attribute read/validated but never assigned
- **Applies when**: `code`: any class whose methods (constructor, validators, `__repr__`, comparison helpers) read `self.<name>`
- **Pattern**: A method dereferences an instance attribute that no code path ever assigns — the assignment was removed, renamed, or replaced by a same-named local variable — so the first access raises at runtime instead of returning the intended value.
- **Detection procedure**:
  1. For each class defined or edited in the program, list every attribute name read as `self.<name>` inside any method body (including inside `if`/`raise` conditions and passed to helper functions). [reads: code]
  2. For each such name, search the whole class body for a binding: `self.<name> = ...`, `setattr(self, "<name>", ...)`, a `@property`/`__getattr__`/`__slots__`-with-default definition of that name, a class-level attribute of that name, or an assignment in a base class named in the class header (check whether that base is defined in the program or is a plain `object`). [reads: code]
  3. Flag the class if some read name has no binding anywhere, *and* especially if a local variable of the same name is computed in the same method (e.g. `name = np.asarray(...)`) but the code then reads `self.name` instead of the local — the normalization result is discarded and the attribute never exists. Cross-check the task statement / class docstring: if the name is a documented public attribute the caller is expected to read, the break is user-visible. [reads: code, task]
- **Counter-example**: A constructor that writes `self.constants = atleast_2d_column_default(constants)` and only afterwards validates `self.constants.ndim`, or a subclass whose `__init__` calls `super().__init__()` where the parent assigns the attribute, or a class exposing the name through a `@property` backed by `self._name`.
- **Discriminator**: The failing case has *zero* binding sites for the name in the class, its declared bases, or a property/`__getattr__`; the safe cases all have exactly one reachable binding executed before the read.
- **Consequence**: `AttributeError: '<Class>' object has no attribute '<name>'` raised on the first construction/use, aborting the script and every test that instantiates the class; any downstream test asserting on that documented attribute (order of a stored list, default value, dtype) fails as an error rather than a mismatch. If the attribute is only read on a rare branch, expect intermittent `AttributeError` instead of a wrong value.
- **Evidence**: A constructor computed `constants = np.asarray(constants, dtype=float)` into a local, dropped the `self.constants = ...` and `self.variable_names = ...` assignments, then validated `self.constants.ndim` — producing `AttributeError: 'LinearConstraint' object has no attribute 'constants'` at line 1 of the reproduction script.
6Self-verification harness that cannot failcodeswesmith/pydata__patsy.a5d16484
Applies when
code: the submission includes its own test/verification scripts and uses their output as evidence that the task is satisfied
Pattern
The checks are written so that both the success and failure paths terminate normally — a try block prints a failure marker instead of raising, the except prints a success marker, or the whole call is wrapped in except: pass — so the harness reports "all passed" no matter what the code does, and the program stops working on the real defect.
Detection procedure
  1. Locate the functions in the added scripts whose names begin with test_ or that are called from a __main__ block that prints an overall success banner. [reads: code]
  2. For each, inspect the failure path: does the branch that represents "wrong behavior" execute a bare print/log rather than assert, raise, or pytest.fail? Also check for try: ... blocks whose except clause is pass or a success print. [reads: code]
  3. The pattern is present if at least one check that the task's described defect would trigger has no assertion on its failure path, and the script's final line unconditionally declares all checks passed. [reads: code]
Counter-example
A script that uses with pytest.raises(ValueError): or places assert False, "should have raised" after the call inside the try — the wrong behavior actually aborts the run with a non-zero exit.
Discriminator
In the failing case, no statement on the "unexpected outcome" branch can propagate an error out of the function; in the safe case an assertion or pytest.raises context turns the unexpected outcome into a failure.
Consequence
Verification output is uninformative, so the program draws the wrong conclusion about repository state and ships without the required change; expect the graded behavioral tests to fail while the submission's own report claims 100% pass.
Evidence
Added checks of the form try: <call>; print("✗ should have raised") except Exception as e: print("✓ ...") and a test_* function whose assertion lines are unreachable, followed by an unconditional "✅ All tests passed!" banner and a final claim that no fix was required.
id 3a6718335563 · mined from swesmith/pydata__patsy.a5d16484 pydata__patsy.a5d16484.combine_file__xgv5bk58
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the functions in the added scripts whose names begin with `test_` or that are called from a `__main__` block that prints an overall success banner. [reads: code]",
 "prediction": "Verification output is uninformative, so the program draws the wrong conclusion about repository state and ships without the required change; expect the graded behavioral tests to fail while the submission's own report claims 100% pass."
}
raw text (what the judge reads)
### Self-verification harness that cannot fail
- **Applies when**: `code`: the submission includes its own test/verification scripts and uses their output as evidence that the task is satisfied
- **Pattern**: The checks are written so that both the success and failure paths terminate normally — a `try` block prints a failure marker instead of raising, the `except` prints a success marker, or the whole call is wrapped in `except: pass` — so the harness reports "all passed" no matter what the code does, and the program stops working on the real defect.
- **Detection procedure**:
  1. Locate the functions in the added scripts whose names begin with `test_` or that are called from a `__main__` block that prints an overall success banner. [reads: code]
  2. For each, inspect the failure path: does the branch that represents "wrong behavior" execute a bare `print`/`log` rather than `assert`, `raise`, or `pytest.fail`? Also check for `try: ...` blocks whose `except` clause is `pass` or a success print. [reads: code]
  3. The pattern is present if at least one check that the task's described defect would trigger has no assertion on its failure path, and the script's final line unconditionally declares all checks passed. [reads: code]
- **Counter-example**: A script that uses `with pytest.raises(ValueError):` or places `assert False, "should have raised"` after the call inside the `try` — the wrong behavior actually aborts the run with a non-zero exit.
- **Discriminator**: In the failing case, no statement on the "unexpected outcome" branch can propagate an error out of the function; in the safe case an assertion or `pytest.raises` context turns the unexpected outcome into a failure.
- **Consequence**: Verification output is uninformative, so the program draws the wrong conclusion about repository state and ships without the required change; expect the graded behavioral tests to fail while the submission's own report claims 100% pass.
- **Evidence**: Added checks of the form `try: <call>; print("✗ should have raised") except Exception as e: print("✓ ...")` and a `test_*` function whose assertion lines are unreachable, followed by an unconditional `"✅ All tests passed!"` banner and a final claim that no fix was required.
6Superseded duplicate test script left in the pytest collection pathcodeswesmith/pydata__patsy.a5d16484
Applies when
code: the submission adds one or more standalone test_*.py files at the repository root or anywhere pytest's default discovery reaches
Pattern
The author writes an ad-hoc test script, finds some of its assertions were themselves wrong, and adds a corrected copy under a new name (..._fixed.py, ..._v2.py, ..._new.py) while leaving the original in place. Both files match the test discovery pattern, so the stale file's known-bad assertions are still collected and run by any full-suite invocation.
Detection procedure
  1. List the added files whose basenames match test_.py or _test.py. [reads: code]
  2. Look for two added files whose basenames are identical up to a suffix such as _fixed, _v2, _new, _corrected, and confirm both live where pytest would collect them (repo root or a package dir), given the project layout in the repo tree. [reads: static facts — repo tree; code]
  3. Open both and check that a same-named test function differs between them — the newer file rewrote or deleted assertions the older one still makes. The pattern is present when the older file is still present and still contains the superseded assertions. [reads: code]
Counter-example
Two added test files with similar names that cover disjoint functions, or the ad-hoc scripts placed in a directory excluded by setup.cfg / tox.ini test paths, or the original file deleted in the same diff.
Discriminator
The failing case keeps both copies collectible and has a same-named test function whose assertions the newer copy contradicts; the safe case has either no overlapping test function, no collection path overlap, or no surviving old copy.
Consequence
A full-suite run from the repository root reports new AssertionError failures (and possible collection errors from module-name clashes or unguarded top-level code) that did not exist before the submission, turning an otherwise green suite red; also leaves untracked scratch artifacts in the delivered diff.
Evidence
The submission added both an ad-hoc integration test script and a ..._fixed.py copy of it in which the same-named associativity/constant-evaluation tests were rewritten, while the original file with the superseded assertions remained at the repository root.
id 0e4c6f63faeb · mined from swesmith/pydata__patsy.a5d16484 pydata__patsy.a5d16484.combine_file__xgv5bk58
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List the added files whose basenames match `test_*.py` or `*_test.py`. [reads: code]",
 "prediction": "A full-suite run from the repository root reports new AssertionError failures (and possible collection errors from module-name clashes or unguarded top-level code) that did not exist before the submission, turning an otherwise green suite red; also leaves untracked scratch artifacts in the delivered diff."
}
raw text (what the judge reads)
### Superseded duplicate test script left in the pytest collection path
- **Applies when**: `code`: the submission adds one or more standalone `test_*.py` files at the repository root or anywhere pytest's default discovery reaches
- **Pattern**: The author writes an ad-hoc test script, finds some of its assertions were themselves wrong, and adds a corrected copy under a new name (`..._fixed.py`, `..._v2.py`, `..._new.py`) while leaving the original in place. Both files match the test discovery pattern, so the stale file's known-bad assertions are still collected and run by any full-suite invocation.
- **Detection procedure**:
  1. List the added files whose basenames match `test_*.py` or `*_test.py`. [reads: code]
  2. Look for two added files whose basenames are identical up to a suffix such as `_fixed`, `_v2`, `_new`, `_corrected`, and confirm both live where pytest would collect them (repo root or a package dir), given the project layout in the repo tree. [reads: static facts — repo tree; code]
  3. Open both and check that a same-named test function differs between them — the newer file rewrote or deleted assertions the older one still makes. The pattern is present when the older file is still present and still contains the superseded assertions. [reads: code]
- **Counter-example**: Two added test files with similar names that cover disjoint functions, or the ad-hoc scripts placed in a directory excluded by `setup.cfg` / `tox.ini` test paths, or the original file deleted in the same diff.
- **Discriminator**: The failing case keeps both copies collectible and has a same-named test function whose assertions the newer copy contradicts; the safe case has either no overlapping test function, no collection path overlap, or no surviving old copy.
- **Consequence**: A full-suite run from the repository root reports new AssertionError failures (and possible collection errors from module-name clashes or unguarded top-level code) that did not exist before the submission, turning an otherwise green suite red; also leaves untracked scratch artifacts in the delivered diff.
- **Evidence**: The submission added both an ad-hoc integration test script and a `..._fixed.py` copy of it in which the same-named associativity/constant-evaluation tests were rewritten, while the original file with the superseded assertions remained at the repository root.
6Truncated / syntactically incomplete file added to the repositorycodeswesmith/pydata__patsy.a5d16484
Applies when
code: the change set adds one or more new .py files to the repository
Pattern
A newly added Python file ends mid-expression — an unclosed call, string, bracket or def body — so importing or collecting it raises SyntaxError. If the file's name matches the test-discovery pattern, the failure is not confined to the file: it aborts collection of the whole run.
Detection procedure
  1. Locate each newly added .py file in the diff and read its final lines. [reads: code]
  2. Check bracket/quote balance across the file and whether the last statement is complete (closing )/]/} and terminating newline present). [reads: code]
  3. Compare the file's basename against the test-runner discovery pattern implied by the environment's test framework and the repo's existing test-file naming (test_.py / _test.py), and against its location relative to the package directory shown in the repo tree. [reads: static facts — repo tree and package list]
Counter-example
A newly added script whose last statement is fully closed, or a long file whose diff hunk merely stops at the hunk boundary while the file's own bracket counts balance — such files import cleanly.
Discriminator
Unbalanced delimiters / an unterminated final call in the added file's complete text; a safe file has every delimiter closed even if it is long.
Consequence
SyntaxError (surfacing as a pytest collection error, or ImportError/IndentationError in other runners); when the file matches the discovery pattern, the entire test session errors out and no score is recorded. This is a secondary failure mode — the primary loss in such change sets is usually the absent implementation fix.
Evidence
An added root-level test_*.py whose last line is print("✓ Parentheses in constraint work" with no closing parenthesis, in a repo whose real tests live inside the package directory.
id cba279661722 · mined from swesmith/pydata__patsy.a5d16484 pydata__patsy.a5d16484.combine_file__xgv5bk58
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Locate each newly added `.py` file in the diff and read its final lines. [reads: code]",
 "prediction": "`SyntaxError` (surfacing as a pytest collection error, or `ImportError`/`IndentationError` in other runners); when the file matches the discovery pattern, the entire test session errors out and no score is recorded. This is a secondary failure mode \u2014 the primary loss in such change sets is usually the absent implementation fix."
}
raw text (what the judge reads)
### Truncated / syntactically incomplete file added to the repository
- **Applies when**: `code`: the change set adds one or more new `.py` files to the repository
- **Pattern**: A newly added Python file ends mid-expression — an unclosed call, string, bracket or `def` body — so importing or collecting it raises `SyntaxError`. If the file's name matches the test-discovery pattern, the failure is not confined to the file: it aborts collection of the whole run.
- **Detection procedure**:
  1. Locate each newly added `.py` file in the diff and read its final lines. [reads: code]
  2. Check bracket/quote balance across the file and whether the last statement is complete (closing `)`/`]`/`}` and terminating newline present). [reads: code]
  3. Compare the file's basename against the test-runner discovery pattern implied by the environment's test framework and the repo's existing test-file naming (`test_*.py` / `*_test.py`), and against its location relative to the package directory shown in the repo tree. [reads: static facts — repo tree and package list]
- **Counter-example**: A newly added script whose last statement is fully closed, or a long file whose diff hunk merely stops at the hunk boundary while the file's own bracket counts balance — such files import cleanly.
- **Discriminator**: Unbalanced delimiters / an unterminated final call in the *added file's complete text*; a safe file has every delimiter closed even if it is long.
- **Consequence**: `SyntaxError` (surfacing as a pytest collection error, or `ImportError`/`IndentationError` in other runners); when the file matches the discovery pattern, the entire test session errors out and no score is recorded. This is a secondary failure mode — the primary loss in such change sets is usually the absent implementation fix.
- **Evidence**: An added root-level `test_*.py` whose last line is `print("✓ Parentheses in constraint work"` with no closing parenthesis, in a repo whose real tests live inside the package directory.
7Locally-defined function stored as long-lived object statecodeswesmith/pydicom__pydicom.7d361b3d
Applies when
code: the program defines a function with a nested def (or a lambda) inside another function/method and hands that callable to a constructor, container, or attribute rather than just calling it
Pattern
a callable created in local scope escapes the call that created it and becomes part of an object's persistent state. Local functions and lambdas have no importable qualified name, so any later pickle/copy/multiprocessing round-trip of the owning object fails, even though ordinary use works. The safe version of the same code puts the helper at module level.
Detection procedure
  1. Locate every def nested inside another def/method body, and every lambda, and note what happens to the resulting name: is it invoked and discarded inside the enclosing call, or is it passed as an argument to a class constructor / assigned to an attribute / placed in a returned object? [reads: code]
  2. Check whether objects of that owning type are serialized or copied anywhere: search the program and the task statement for pickle, copy.deepcopy, __reduce__, __getstate__, multiprocessing, on-disk caching of objects, or a stated requirement that instances be picklable/copyable; also check the static facts for a test module or package whose subject is serialization of these objects [reads: code, task, static facts — test/module listing]
  3. The discriminating observation: the nested callable is retained by the object returned from the enclosing function (constructor argument or attribute assignment), and there is no module-level function of equivalent behaviour used instead [reads: code]
Counter-example
a nested function or lambda used only inside the enclosing call — passed to map, sorted(key=...), filter, or invoked immediately — and never stored on a returned object; or a module-level helper function referenced by name and passed to the same constructor.
Discriminator
the callable escapes the defining scope and is held as state of a serializable object (goes wrong) versus being consumed before the enclosing call returns, or being a module-level def (safe).
Consequence
AttributeError: Can't pickle local object '<outer>.<locals>.<inner>' or pickle.PicklingError (and TypeError: cannot pickle ... under multiprocessing) whenever the owning object is pickled, cached, deep-copied through __reduce__, or shipped to a worker; all non-serializing tests still pass, so the breakage is confined to serialization paths but is a hard failure there.
Evidence
a module-level pass-through helper that was passed to a value-container constructor was replaced by an inner def _skip_conversion(val) defined inside the method and handed to the same constructor; the targeted unit test still passed, leaving the unpicklable callable embedded in the constructed container.
id 61d603d3080b · mined from swesmith/pydicom__pydicom.7d361b3d pydicom__pydicom.7d361b3d.func_basic__mv217zgm
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate every `def` nested inside another `def`/method body, and every `lambda`, and note what happens to the resulting name: is it invoked and discarded inside the enclosing call, or is it passed as an argument to a class constructor / assigned to an attribute / placed in a returned object? [reads: code]",
 "prediction": "`AttributeError: Can't pickle local object '<outer>.<locals>.<inner>'` or `pickle.PicklingError` (and `TypeError: cannot pickle ...` under multiprocessing) whenever the owning object is pickled, cached, deep-copied through `__reduce__`, or shipped to a worker; all non-serializing tests still pass, so the breakage is confined to serialization paths but is a hard failure there."
}
raw text (what the judge reads)
### Locally-defined function stored as long-lived object state
- **Applies when**: `code`: the program defines a function with a nested `def` (or a `lambda`) inside another function/method and hands that callable to a constructor, container, or attribute rather than just calling it
- **Pattern**: a callable created in local scope escapes the call that created it and becomes part of an object's persistent state. Local functions and lambdas have no importable qualified name, so any later `pickle`/`copy`/`multiprocessing` round-trip of the owning object fails, even though ordinary use works. The safe version of the same code puts the helper at module level.
- **Detection procedure**:
  1. Locate every `def` nested inside another `def`/method body, and every `lambda`, and note what happens to the resulting name: is it invoked and discarded inside the enclosing call, or is it passed as an argument to a class constructor / assigned to an attribute / placed in a returned object? [reads: code]
  2. Check whether objects of that owning type are serialized or copied anywhere: search the program and the task statement for `pickle`, `copy.deepcopy`, `__reduce__`, `__getstate__`, `multiprocessing`, on-disk caching of objects, or a stated requirement that instances be picklable/copyable; also check the static facts for a test module or package whose subject is serialization of these objects [reads: code, task, static facts — test/module listing]
  3. The discriminating observation: the nested callable is *retained* by the object returned from the enclosing function (constructor argument or attribute assignment), and there is no module-level function of equivalent behaviour used instead [reads: code]
- **Counter-example**: a nested function or lambda used only inside the enclosing call — passed to `map`, `sorted(key=...)`, `filter`, or invoked immediately — and never stored on a returned object; or a module-level helper function referenced by name and passed to the same constructor.
- **Discriminator**: the callable escapes the defining scope and is held as state of a serializable object (goes wrong) versus being consumed before the enclosing call returns, or being a module-level `def` (safe).
- **Consequence**: `AttributeError: Can't pickle local object '<outer>.<locals>.<inner>'` or `pickle.PicklingError` (and `TypeError: cannot pickle ...` under multiprocessing) whenever the owning object is pickled, cached, deep-copied through `__reduce__`, or shipped to a worker; all non-serializing tests still pass, so the breakage is confined to serialization paths but is a hard failure there.
- **Evidence**: a module-level pass-through helper that was passed to a value-container constructor was replaced by an inner `def _skip_conversion(val)` defined inside the method and handed to the same constructor; the targeted unit test still passed, leaving the unpicklable callable embedded in the constructed container.
7Special-case branch bypasses the validation applied on every sibling branchcodeswesmith/pydicom__pydicom.7d361b3d
Applies when
code: a function converts/normalizes incoming values through a validating converter, and contains a branch that special-cases certain inputs by substituting an identity/no-op converter
Pattern
to avoid an unwanted type coercion for a special input, the program swaps the whole validating converter for a pass-through, discarding the validation that every other input still receives. The behaviour needed was "skip conversion"; what was written is "skip conversion and skip checking", so malformed or out-of-range values on that path are accepted silently.
Detection procedure
  1. Locate the function that dispatches incoming values to a converter/validator (a method like _convert, _coerce, validate, or a call to a validate_* helper) and list the branches that decide which callable is used [reads: code]
  2. Read the task statement for what the special case is supposed to change — whether it asks only that a type conversion be skipped for certain positions/values, not that correctness checking be dropped [reads: task]
  3. The discriminating observation: on the special-case branch the substituted callable body is return val (or lambda v: v) and no validate_*/range/format check is invoked anywhere on that branch, while the general branch routes through the checking converter [reads: code]
Counter-example
a branch that bypasses type coercion but still calls the validation helper explicitly on each element before returning the raw values, or a branch guarded by an explicit "validation disabled" configuration flag read from settings.
Discriminator
presence of at least one explicit validation call (or an explicit user-set disable flag) on the bypass path — absent in the failing case, present in the safe one.
Consequence
invalid values are stored without the warning/exception the library contracts to emit; tests that assert a warning or error for malformed values of that special case fail (Failed: DID NOT WARN / DID NOT RAISE), and downstream writing of the object emits non-conforming data. Explains the correctness regression only; the serialization defect above is separate.
Evidence
a branch for a specially-shaped multi-value element replaced validate_value(...) calls plus a pass-through converter with a bare identity converter and no validation at all; the single executed test exercised only the conversion-skipping behaviour and passed, hiding the removed checks.
id 21af31dc4fce · mined from swesmith/pydicom__pydicom.7d361b3d pydicom__pydicom.7d361b3d.func_basic__mv217zgm
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the function that dispatches incoming values to a converter/validator (a method like `_convert`, `_coerce`, `validate`, or a call to a `validate_*` helper) and list the branches that decide which callable is used [reads: code]",
 "prediction": "invalid values are stored without the warning/exception the library contracts to emit; tests that assert a warning or error for malformed values of that special case fail (`Failed: DID NOT WARN` / `DID NOT RAISE`), and downstream writing of the object emits non-conforming data. Explains the correctness regression only; the serialization defect above is separate."
}
raw text (what the judge reads)
### Special-case branch bypasses the validation applied on every sibling branch
- **Applies when**: `code`: a function converts/normalizes incoming values through a validating converter, and contains a branch that special-cases certain inputs by substituting an identity/no-op converter
- **Pattern**: to avoid an unwanted type coercion for a special input, the program swaps the whole validating converter for a pass-through, discarding the validation that every other input still receives. The behaviour needed was "skip conversion"; what was written is "skip conversion *and* skip checking", so malformed or out-of-range values on that path are accepted silently.
- **Detection procedure**:
  1. Locate the function that dispatches incoming values to a converter/validator (a method like `_convert`, `_coerce`, `validate`, or a call to a `validate_*` helper) and list the branches that decide which callable is used [reads: code]
  2. Read the task statement for what the special case is supposed to change — whether it asks only that a *type conversion* be skipped for certain positions/values, not that correctness checking be dropped [reads: task]
  3. The discriminating observation: on the special-case branch the substituted callable body is `return val` (or `lambda v: v`) and no `validate_*`/range/format check is invoked anywhere on that branch, while the general branch routes through the checking converter [reads: code]
- **Counter-example**: a branch that bypasses type coercion but still calls the validation helper explicitly on each element before returning the raw values, or a branch guarded by an explicit "validation disabled" configuration flag read from settings.
- **Discriminator**: presence of at least one explicit validation call (or an explicit user-set disable flag) on the bypass path — absent in the failing case, present in the safe one.
- **Consequence**: invalid values are stored without the warning/exception the library contracts to emit; tests that assert a warning or error for malformed values of that special case fail (`Failed: DID NOT WARN` / `DID NOT RAISE`), and downstream writing of the object emits non-conforming data. Explains the correctness regression only; the serialization defect above is separate.
- **Evidence**: a branch for a specially-shaped multi-value element replaced `validate_value(...)` calls plus a pass-through converter with a bare identity converter and no validation at all; the single executed test exercised only the conversion-skipping behaviour and passed, hiding the removed checks.
7Write-then-read round trip omitting the container header the strict reader requirescodeswesmith/pydicom__pydicom.7d361b3d
Applies when
code: the program serialises an in-memory object to a file with a library's writer and later loads that same path back with the library's reader (round-trip check, save/reload verification, artifact reuse).
Pattern
The object is written in a bare/partial form because the header or metadata block the format requires is never populated and the writer is not asked to enforce the full file format; the subsequent read then uses the strict reader with default arguments, so loading the file the program just wrote raises a format-validation error.
Detection procedure
  1. Locate the write call and the later read call operating on the same filename variable. [reads: code]
  2. Read the construction of the written object and the write call's arguments: is the header/metadata attribute populated (e.g. a file_meta/header object assigned) or is the "write full/standard file format" argument passed (enforce_file_format=True, write_like_original=False, or the library's equivalent)? [reads: code]
  3. Read the read call's arguments: does it pass the permissive/skip-header-check flag (force=True) or otherwise pre-seed the header? If step 2 found no header/enforce option and step 3 found no permissive flag, the rubric fires. [reads: code]
Counter-example
The same write/read pair where the object is given a fully populated header/meta block before writing (or the writer is called with the enforce-full-format argument) — reading back with default arguments succeeds; likewise a read that passes the permissive flag is safe even for a bare write.
Discriminator
Fires only when both sides are unguarded — no header written and no permissive read flag. Either guard alone makes the round trip safe.
Consequence
The write appears to succeed and the program then terminates at the read with the library's format-validation exception — pydicom.errors.InvalidDicomError here, generally the reader's InvalidFileError/ValueError/OSError about a missing magic prefix or header — so any verification or downstream use of the reloaded artifact never runs.
Evidence
A script saved a dataset built purely from data elements and immediately called the reader with default arguments; the reader raised InvalidDicomError: File is missing ... header or the 'DICM' prefix is missing ... Use force=True, after printing "✓ Save successful".
id 2443a89d2699 · mined from swesmith/pydicom__pydicom.7d361b3d pydicom__pydicom.7d361b3d.func_basic__mv217zgm
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the write call and the later read call operating on the same filename variable. [reads: code]",
 "prediction": "The write appears to succeed and the program then terminates at the read with the library's format-validation exception \u2014 `pydicom.errors.InvalidDicomError` here, generally the reader's `InvalidFileError`/`ValueError`/`OSError` about a missing magic prefix or header \u2014 so any verification or downstream use of the reloaded artifact never runs."
}
raw text (what the judge reads)
### Write-then-read round trip omitting the container header the strict reader requires
- **Applies when**: `code`: the program serialises an in-memory object to a file with a library's writer and later loads that same path back with the library's reader (round-trip check, save/reload verification, artifact reuse).
- **Pattern**: The object is written in a bare/partial form because the header or metadata block the format requires is never populated and the writer is not asked to enforce the full file format; the subsequent read then uses the strict reader with default arguments, so loading the file the program just wrote raises a format-validation error.
- **Detection procedure**:
  1. Locate the write call and the later read call operating on the same filename variable. [reads: code]
  2. Read the construction of the written object and the write call's arguments: is the header/metadata attribute populated (e.g. a `file_meta`/header object assigned) or is the "write full/standard file format" argument passed (`enforce_file_format=True`, `write_like_original=False`, or the library's equivalent)? [reads: code]
  3. Read the read call's arguments: does it pass the permissive/skip-header-check flag (`force=True`) or otherwise pre-seed the header? If step 2 found no header/enforce option **and** step 3 found no permissive flag, the rubric fires. [reads: code]
- **Counter-example**: The same write/read pair where the object is given a fully populated header/meta block before writing (or the writer is called with the enforce-full-format argument) — reading back with default arguments succeeds; likewise a read that passes the permissive flag is safe even for a bare write.
- **Discriminator**: Fires only when *both* sides are unguarded — no header written and no permissive read flag. Either guard alone makes the round trip safe.
- **Consequence**: The write appears to succeed and the program then terminates at the read with the library's format-validation exception — `pydicom.errors.InvalidDicomError` here, generally the reader's `InvalidFileError`/`ValueError`/`OSError` about a missing magic prefix or header — so any verification or downstream use of the reloaded artifact never runs.
- **Evidence**: A script saved a dataset built purely from data elements and immediately called the reader with default arguments; the reader raised `InvalidDicomError: File is missing ... header or the 'DICM' prefix is missing ... Use force=True`, after printing "✓ Save successful".
7Rename-only / cosmetic churn presented as the fixcodeswesmith/pydicom__pydicom.7d361b3d
Applies when
code: the change set is a patch over an existing repository and contains hunks that only rename identifiers, rewrite docstrings/comments, or alter file-terminating whitespace
Pattern
A substantial share of the patch is behaviour-neutral churn (renaming a module-level private helper, reflowing a docstring, dropping the final newline). It adds review surface and risk of unresolved references while contributing nothing to the requested fix.
Detection procedure
  1. Classify each hunk in the change set as behaviour-changing or behaviour-neutral (identifier rename with matching definition+use edits, docstring/comment text, trailing-newline removal marked \ No newline at end of file). [reads: code]
  2. For each renamed module-level or class-level name, check whether the patch updates every occurrence it shows; note that references may also exist in tests/ modules and sibling source files listed in the repository tree that the patch does not open. [reads: static facts — repo tree]
  3. The condition holds if at least one renamed non-local symbol is defined at module scope (importable) and the patch changes only its definition site plus the callers inside the same file, or if the patch removes the terminating newline of a source file. [reads: code]
Counter-example
A patch that renames a variable local to one function, or one where the task explicitly requests a rename/refactor and the diff updates the corresponding test file too.
Discriminator
The goes-wrong case renames a module-scope (importable) name or perturbs file-final whitespace while the task asked only for a behaviour fix; the safe case confines renames to function-local scope or is mandated by the task.
Consequence
Risk of AttributeError/ImportError from out-of-file references to the old name and failure of repository style hooks configured in .pre-commit-config.yaml (end-of-file fixer); no improvement in the correctness metric. Explains only a minor part of a score gap — the dominant loss is the missing real fix.
Evidence
The patch renamed a module-level _-prefixed helper and stripped the file's trailing newline (\ No newline at end of file) while leaving the reported defect unfixed.
id 952d81f062da · mined from swesmith/pydicom__pydicom.7d361b3d pydicom__pydicom.7d361b3d.func_basic__mv217zgm
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Classify each hunk in the change set as behaviour-changing or behaviour-neutral (identifier rename with matching definition+use edits, docstring/comment text, trailing-newline removal marked `\\ No newline at end of file`). [reads: code]",
 "prediction": "Risk of `AttributeError`/`ImportError` from out-of-file references to the old name and failure of repository style hooks configured in `.pre-commit-config.yaml` (end-of-file fixer); no improvement in the correctness metric. Explains only a minor part of a score gap \u2014 the dominant loss is the missing real fix."
}
raw text (what the judge reads)
### Rename-only / cosmetic churn presented as the fix
- **Applies when**: `code`: the change set is a patch over an existing repository and contains hunks that only rename identifiers, rewrite docstrings/comments, or alter file-terminating whitespace
- **Pattern**: A substantial share of the patch is behaviour-neutral churn (renaming a module-level private helper, reflowing a docstring, dropping the final newline). It adds review surface and risk of unresolved references while contributing nothing to the requested fix.
- **Detection procedure**:
  1. Classify each hunk in the change set as behaviour-changing or behaviour-neutral (identifier rename with matching definition+use edits, docstring/comment text, trailing-newline removal marked `\ No newline at end of file`). [reads: code]
  2. For each renamed module-level or class-level name, check whether the patch updates every occurrence it shows; note that references may also exist in `tests/` modules and sibling source files listed in the repository tree that the patch does not open. [reads: static facts — repo tree]
  3. The condition holds if at least one renamed non-local symbol is defined at module scope (importable) and the patch changes only its definition site plus the callers inside the same file, or if the patch removes the terminating newline of a source file. [reads: code]
- **Counter-example**: A patch that renames a variable local to one function, or one where the task explicitly requests a rename/refactor and the diff updates the corresponding test file too.
- **Discriminator**: The goes-wrong case renames a module-scope (importable) name or perturbs file-final whitespace while the task asked only for a behaviour fix; the safe case confines renames to function-local scope or is mandated by the task.
- **Consequence**: Risk of `AttributeError`/`ImportError` from out-of-file references to the old name and failure of repository style hooks configured in `.pre-commit-config.yaml` (end-of-file fixer); no improvement in the correctness metric. Explains only a minor part of a score gap — the dominant loss is the missing real fix.
- **Evidence**: The patch renamed a module-level `_`-prefixed helper and stripped the file's trailing newline (`\ No newline at end of file`) while leaving the reported defect unfixed.
8Content-sniffing to choose between two decodings whose valid inputs overlapcodeswesmith/lqs__sqlingo.ed36ef03
Applies when
code: a function converts an untyped/raw string or byte payload into a number (or other typed value) and picks the decoding strategy by inspecting the payload's own bytes rather than by consulting declared type metadata.
Pattern
The converter tries decoding A (e.g. textual/decimal parse) and, on failure, applies a heuristic predicate over the bytes (e.g. "contains a byte outside printable ASCII") to decide whether to apply decoding B (e.g. big-endian raw-byte accumulation). Because the two input languages overlap — a raw byte payload can consist entirely of bytes that are also valid text — the heuristic misclassifies part of the domain and silently returns a wrong value or the zero value, with no error surfaced.
Detection procedure
  1. Find the conversion function and identify the branch order: an attempted parse, then a predicate over the raw bytes gating an alternative decoding, then a bare fallback return of a zero/empty value. [reads: code]
  2. Read the predicate's body and ask whether any payload the alternative decoding is meant to handle could satisfy the first decoder or fail the predicate: e.g. a raw byte payload whose every byte lies in the printable range, or a single raw byte whose value is an ASCII digit. [reads: code]
  3. Confirm no type/width/metadata argument (column type, declared length, format flag, driver-provided type code) is available to the function or is threaded in from the caller — the decision rests solely on payload content. [reads: code]
Counter-example
A converter that receives an explicit type or width parameter (or checks a length/prefix marker that the two encodings cannot share) and dispatches on it; or one that returns an error/ok flag when the payload matches neither encoding, so ambiguity is reported instead of guessed.
Discriminator
The failing case is dispatch on a content predicate whose true-set and the first decoder's accepted-set do not partition the input domain, and the mismatch path returns a plausible-looking value (zero or a text parse) rather than an error. Safe code either dispatches on out-of-band type information or reports the unresolvable case.
Consequence
No exception is raised; specific inputs return silently incorrect values (raw-byte payloads made only of printable bytes yield 0 or the decimal reading of those bytes instead of the intended integer). Test suites that only exercise payloads containing control/high bytes report 100% pass while the overlap region stays broken — expect hidden failures on later, wider inputs rather than a visible test failure.
Evidence
if isBinaryData(*v.stringValue) { result = result<<8 | int64(byte) } after a failed strconv.ParseInt, with isBinaryData defined as "any byte < 32 or > 126"; all 19 checks in the run passed because every raw-byte fixture contained a non-printable byte, leaving all-printable raw payloads returning 0.
id 7ba8ddcf0f5c · mined from swesmith/lqs__sqlingo.ed36ef03 lqs__sqlingo.ed36ef03.func_pm_flip_operators__jcqg2mrg
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Find the conversion function and identify the branch order: an attempted parse, then a predicate over the raw bytes gating an alternative decoding, then a bare fallback return of a zero/empty value. [reads: code]",
 "prediction": "No exception is raised; specific inputs return silently incorrect values (raw-byte payloads made only of printable bytes yield `0` or the decimal reading of those bytes instead of the intended integer). Test suites that only exercise payloads containing control/high bytes report 100% pass while the overlap region stays broken \u2014 expect hidden failures on later, wider inputs rather than a visible test failure."
}
raw text (what the judge reads)
### Content-sniffing to choose between two decodings whose valid inputs overlap
- **Applies when**: `code`: a function converts an untyped/raw string or byte payload into a number (or other typed value) and picks the decoding strategy by inspecting the payload's own bytes rather than by consulting declared type metadata.
- **Pattern**: The converter tries decoding A (e.g. textual/decimal parse) and, on failure, applies a heuristic predicate over the bytes (e.g. "contains a byte outside printable ASCII") to decide whether to apply decoding B (e.g. big-endian raw-byte accumulation). Because the two input languages overlap — a raw byte payload can consist entirely of bytes that are also valid text — the heuristic misclassifies part of the domain and silently returns a wrong value or the zero value, with no error surfaced.
- **Detection procedure**:
  1. Find the conversion function and identify the branch order: an attempted parse, then a predicate over the raw bytes gating an alternative decoding, then a bare fallback return of a zero/empty value. [reads: code]
  2. Read the predicate's body and ask whether any payload the alternative decoding is meant to handle could satisfy the *first* decoder or fail the predicate: e.g. a raw byte payload whose every byte lies in the printable range, or a single raw byte whose value is an ASCII digit. [reads: code]
  3. Confirm no type/width/metadata argument (column type, declared length, format flag, driver-provided type code) is available to the function or is threaded in from the caller — the decision rests solely on payload content. [reads: code]
- **Counter-example**: A converter that receives an explicit type or width parameter (or checks a length/prefix marker that the two encodings cannot share) and dispatches on it; or one that returns an `error`/`ok` flag when the payload matches neither encoding, so ambiguity is reported instead of guessed.
- **Discriminator**: The failing case is dispatch on a content predicate whose true-set and the first decoder's accepted-set do not partition the input domain, *and* the mismatch path returns a plausible-looking value (zero or a text parse) rather than an error. Safe code either dispatches on out-of-band type information or reports the unresolvable case.
- **Consequence**: No exception is raised; specific inputs return silently incorrect values (raw-byte payloads made only of printable bytes yield `0` or the decimal reading of those bytes instead of the intended integer). Test suites that only exercise payloads containing control/high bytes report 100% pass while the overlap region stays broken — expect hidden failures on later, wider inputs rather than a visible test failure.
- **Evidence**: `if isBinaryData(*v.stringValue) { result = result<<8 | int64(byte) }` after a failed `strconv.ParseInt`, with `isBinaryData` defined as "any byte < 32 or > 126"; all 19 checks in the run passed because every raw-byte fixture contained a non-printable byte, leaving all-printable raw payloads returning `0`.
8Unbounded byte accumulation into a fixed-width accumulatorcodeswesmith/lqs__sqlingo.ed36ef03
Applies when
code: a loop folds a variable-length byte sequence into a fixed-width integer accumulator (acc = acc<<8 | b, or repeated multiply-add) to decode a big-endian numeric value.
Pattern
The loop iterates over the entire input with no check that the input length fits the accumulator's width, so longer inputs silently shift the leading bytes out of range and the function returns a wrapped, arbitrarily wrong number instead of rejecting the input.
Detection procedure
  1. Locate the accumulation loop and note the accumulator's declared type width (e.g. int64/uint64 = 8 bytes). [reads: code]
  2. Check the loop bound: does it iterate to len(input) unconditionally, or is it clamped/preceded by a guard comparing len(input) against the accumulator's byte width? [reads: code]
  3. Confirm no caller-side validation restricts the payload length before the call (search the call sites of this conversion for a length check or a width parameter). [reads: code]
Counter-example
The same fold guarded by if len(input) > 8 { return 0, error }, or iterating over only the last/first 8 bytes deliberately, or accumulating into a big-integer type with no fixed width.
Discriminator
Goes wrong when the loop bound is len(input) with a fixed-width accumulator and no length guard anywhere on the path; safe when a width check exists or the accumulator is arbitrary-precision. Signed accumulators additionally flip sign once the top bit is set, which range checks in derived narrower conversions then map to the zero value.
Consequence
No panic (Go shifts wrap); returns silently truncated/wrapped values for payloads longer than the accumulator width, and negative values for 8-byte payloads with the high bit set, which downstream range-checked narrowing then converts to 0. Explains only the long-payload subset of conversion errors; short-payload misreads come from the dispatch heuristic itself.
Evidence
for i := 0; i < len(v.stringValue); i++ { result = result<<8 | int64((v.stringValue)[i]) } with no comparison of the string length to 8; the passing test set contained payloads of at most 4 bytes, so the wrap path was never exercised.
id b9aa5e2433e2 · mined from swesmith/lqs__sqlingo.ed36ef03 lqs__sqlingo.ed36ef03.func_pm_flip_operators__jcqg2mrg
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the accumulation loop and note the accumulator's declared type width (e.g. `int64`/`uint64` = 8 bytes). [reads: code]",
 "prediction": "No panic (Go shifts wrap); returns silently truncated/wrapped values for payloads longer than the accumulator width, and negative values for 8-byte payloads with the high bit set, which downstream range-checked narrowing then converts to `0`. Explains only the long-payload subset of conversion errors; short-payload misreads come from the dispatch heuristic itself."
}
raw text (what the judge reads)
### Unbounded byte accumulation into a fixed-width accumulator
- **Applies when**: `code`: a loop folds a variable-length byte sequence into a fixed-width integer accumulator (`acc = acc<<8 | b`, or repeated multiply-add) to decode a big-endian numeric value.
- **Pattern**: The loop iterates over the entire input with no check that the input length fits the accumulator's width, so longer inputs silently shift the leading bytes out of range and the function returns a wrapped, arbitrarily wrong number instead of rejecting the input.
- **Detection procedure**:
  1. Locate the accumulation loop and note the accumulator's declared type width (e.g. `int64`/`uint64` = 8 bytes). [reads: code]
  2. Check the loop bound: does it iterate to `len(input)` unconditionally, or is it clamped/preceded by a guard comparing `len(input)` against the accumulator's byte width? [reads: code]
  3. Confirm no caller-side validation restricts the payload length before the call (search the call sites of this conversion for a length check or a width parameter). [reads: code]
- **Counter-example**: The same fold guarded by `if len(input) > 8 { return 0, error }`, or iterating over only the last/first 8 bytes deliberately, or accumulating into a big-integer type with no fixed width.
- **Discriminator**: Goes wrong when the loop bound is `len(input)` with a fixed-width accumulator and no length guard anywhere on the path; safe when a width check exists or the accumulator is arbitrary-precision. Signed accumulators additionally flip sign once the top bit is set, which range checks in derived narrower conversions then map to the zero value.
- **Consequence**: No panic (Go shifts wrap); returns silently truncated/wrapped values for payloads longer than the accumulator width, and negative values for 8-byte payloads with the high bit set, which downstream range-checked narrowing then converts to `0`. Explains only the long-payload subset of conversion errors; short-payload misreads come from the dispatch heuristic itself.
- **Evidence**: `for i := 0; i < len(*v.stringValue); i++ { result = result<<8 | int64((*v.stringValue)[i]) }` with no comparison of the string length to 8; the passing test set contained payloads of at most 4 bytes, so the wrap path was never exercised.
9Reproduction driven from a source-less `__main__` when the library needs source textcodeswesmith/agronholm__typeguard.b6a7e438
Applies when
code: the program is a shell/inline invocation (python -c "...", python - <<EOF, exec()/compile() of a string, or a REPL transcript) that defines functions/classes and hands them to an API from the repository under test
Pattern
The program exercises a library whose implementation works by retrieving and re-parsing the source text of the defining module (import hooks, AST rewriting, decorators that re-compile the decorated object), but defines the target objects in a context that has no source file on disk (__main__ created from -c, stdin, or exec). The library's inspect.getsource/getsourcefile call fails before any of the intended behaviour is reached, so the run dies on an infrastructure error unrelated to the property being demonstrated.
Detection procedure
  1. Read the program text and determine how the code under test is delivered: a saved .py file executed as a script/test, or a string passed to python -c / heredoc / exec / compile. [reads: code]
  2. Check the repository layout and package names in the static facts for evidence that the library is source/AST-based rather than purely runtime-reflective — e.g. modules or test files named transformer, importhook, instrument, decorators, or a documented "rewrites your code" feature. [reads: static facts — repo tree]
  3. Confirm the discriminating condition: inside the inline string, a function/class is defined and then decorated or registered with that library API (rather than the string merely importing the package, calling a plain function, or operating on objects imported from real modules on disk). [reads: code]
Counter-example
python -c "import pkg; from tests.dummymodule import decorated_func; decorated_func(5)" — the decorated object lives in a real module file, so source retrieval succeeds; likewise a python -c that only calls non-instrumenting APIs, or a reproduction written into a temp .py file and then executed.
Discriminator
the object passed to the source-rewriting API is defined inside the inline/exec'd string, so its __module__ resolves to a module with no __file__; in the safe case the object originates from a module imported from disk, or no source-rewriting API is used at all.
Consequence
the run terminates before demonstrating anything, most likely TypeError: <module '__main__' (built-in)> is a built-in module from inspect.getfile, or OSError: could not get source code / OSError: source code not available from inspect.getsource. Any assertion or expected error later in the script is never reached, so the reproduction/verification is worthless even if the library behaves correctly.
Evidence
python -c "from pkg import decorator\n@decorator\ndef f(x: int) -> str: ..." raised TypeError: <module '__main__' (built-in)> is a built-in module inside the decorator's inspect.getsource(sys.modules[f.__module__]) call, before the intended type-violation call was ever executed.
id c1a7876bbded · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.func_basic__aw9j63hw
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Read the program text and determine how the code under test is delivered: a saved `.py` file executed as a script/test, or a string passed to `python -c` / heredoc / `exec` / `compile`. [reads: code]",
 "prediction": "the run terminates before demonstrating anything, most likely `TypeError: <module '__main__' (built-in)> is a built-in module` from `inspect.getfile`, or `OSError: could not get source code` / `OSError: source code not available` from `inspect.getsource`. Any assertion or expected error later in the script is never reached, so the reproduction/verification is worthless even if the library behaves correctly."
}
raw text (what the judge reads)
### Reproduction driven from a source-less `__main__` when the library needs source text
- **Applies when**: `code`: the program is a shell/inline invocation (`python -c "..."`, `python - <<EOF`, `exec()`/`compile()` of a string, or a REPL transcript) that defines functions/classes and hands them to an API from the repository under test
- **Pattern**: The program exercises a library whose implementation works by retrieving and re-parsing the *source text* of the defining module (import hooks, AST rewriting, decorators that re-compile the decorated object), but defines the target objects in a context that has no source file on disk (`__main__` created from `-c`, stdin, or `exec`). The library's `inspect.getsource`/`getsourcefile` call fails before any of the intended behaviour is reached, so the run dies on an infrastructure error unrelated to the property being demonstrated.
- **Detection procedure**:
  1. Read the program text and determine how the code under test is delivered: a saved `.py` file executed as a script/test, or a string passed to `python -c` / heredoc / `exec` / `compile`. [reads: code]
  2. Check the repository layout and package names in the static facts for evidence that the library is source/AST-based rather than purely runtime-reflective — e.g. modules or test files named `*transformer*`, `*importhook*`, `*instrument*`, `*decorators*`, or a documented "rewrites your code" feature. [reads: static facts — repo tree]
  3. Confirm the discriminating condition: inside the inline string, a function/class is *defined* and then decorated or registered with that library API (rather than the string merely importing the package, calling a plain function, or operating on objects imported from real modules on disk). [reads: code]
- **Counter-example**: `python -c "import pkg; from tests.dummymodule import decorated_func; decorated_func(5)"` — the decorated object lives in a real module file, so source retrieval succeeds; likewise a `python -c` that only calls non-instrumenting APIs, or a reproduction written into a temp `.py` file and then executed.
- **Discriminator**: the object passed to the source-rewriting API is defined *inside* the inline/`exec`'d string, so its `__module__` resolves to a module with no `__file__`; in the safe case the object originates from a module imported from disk, or no source-rewriting API is used at all.
- **Consequence**: the run terminates before demonstrating anything, most likely `TypeError: <module '__main__' (built-in)> is a built-in module` from `inspect.getfile`, or `OSError: could not get source code` / `OSError: source code not available` from `inspect.getsource`. Any assertion or expected error later in the script is never reached, so the reproduction/verification is worthless even if the library behaves correctly.
- **Evidence**: `python -c "from pkg import decorator\n@decorator\ndef f(x: int) -> str: ..."` raised `TypeError: <module '__main__' (built-in)> is a built-in module` inside the decorator's `inspect.getsource(sys.modules[f.__module__])` call, before the intended type-violation call was ever executed.
9Expected-error path invoked at top level without try/except in a verification scriptcodeswesmith/agronholm__typeguard.b6a7e438
Applies when
code: the program is a script or inline snippet run to demonstrate/verify behavior, and one of the calls it makes is designed to raise (invalid input, violated contract, error branch)
Pattern
The script exercises the "should raise" case as an ordinary top-level statement instead of inside try/except or an assertion helper, so the intended-and-correct exception propagates, the interpreter exits non-zero, and any statements after it never run. A successful demonstration is then indistinguishable from a crash.
Detection procedure
  1. Read the script body and list the calls made at module/top level, in order [reads: code]
  2. Identify any call whose arguments or state deliberately violate the condition the task says should be rejected/validated (e.g. wrong-typed argument, out-of-range value, missing key) [reads: task]
  3. Check whether that call is wrapped in try/except, pytest.raises, contextlib.suppress, or unittest.assertRaises; if it is bare, and especially if further statements (prints, additional checks, cleanup, artifact writes) follow it, the pattern is present [reads: code]
Counter-example
A script that calls the same invalid input inside with pytest.raises(SomeError): or try: f(bad) except SomeError as e: print("ok:", e) and then continues with more checks — same call, same exception, but the run completes with exit code 0.
Discriminator
The raise-inducing call is unguarded at top level (no exception handler in its enclosing scope), so the exception reaches the interpreter; in the safe version an enclosing handler/context manager catches exactly that exception class.
Consequence
The process terminates with the library's own error class (here a domain *Error/TypeError/ValueError) and a non-zero exit status; every statement after the deliberate call is skipped, so later verification output or written artifacts are missing and the run is scored as a failure even though the library behaved correctly.
Evidence
A one-liner ran f(valid) then f(invalid) unguarded; the traceback ended in the library raising its own validation error and the interpreter exiting non-zero, with the remainder of the intended checks never executed.
id 752dcc90120f · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.func_basic__aw9j63hw
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Read the script body and list the calls made at module/top level, in order [reads: code]",
 "prediction": "The process terminates with the library's own error class (here a domain `*Error`/`TypeError`/`ValueError`) and a non-zero exit status; every statement after the deliberate call is skipped, so later verification output or written artifacts are missing and the run is scored as a failure even though the library behaved correctly."
}
raw text (what the judge reads)
### Expected-error path invoked at top level without try/except in a verification script
- **Applies when**: `code`: the program is a script or inline snippet run to demonstrate/verify behavior, and one of the calls it makes is designed to raise (invalid input, violated contract, error branch)
- **Pattern**: The script exercises the "should raise" case as an ordinary top-level statement instead of inside `try/except` or an assertion helper, so the intended-and-correct exception propagates, the interpreter exits non-zero, and any statements after it never run. A successful demonstration is then indistinguishable from a crash.
- **Detection procedure**:
  1. Read the script body and list the calls made at module/top level, in order [reads: code]
  2. Identify any call whose arguments or state deliberately violate the condition the task says should be rejected/validated (e.g. wrong-typed argument, out-of-range value, missing key) [reads: task]
  3. Check whether that call is wrapped in `try/except`, `pytest.raises`, `contextlib.suppress`, or `unittest.assertRaises`; if it is bare, and especially if further statements (prints, additional checks, cleanup, artifact writes) follow it, the pattern is present [reads: code]
- **Counter-example**: A script that calls the same invalid input inside `with pytest.raises(SomeError):` or `try: f(bad) except SomeError as e: print("ok:", e)` and then continues with more checks — same call, same exception, but the run completes with exit code 0.
- **Discriminator**: The raise-inducing call is unguarded at top level (no exception handler in its enclosing scope), so the exception reaches the interpreter; in the safe version an enclosing handler/context manager catches exactly that exception class.
- **Consequence**: The process terminates with the library's own error class (here a domain `*Error`/`TypeError`/`ValueError`) and a non-zero exit status; every statement after the deliberate call is skipped, so later verification output or written artifacts are missing and the run is scored as a failure even though the library behaved correctly.
- **Evidence**: A one-liner ran `f(valid)` then `f(invalid)` unguarded; the traceback ended in the library raising its own validation error and the interpreter exiting non-zero, with the remainder of the intended checks never executed.
9Ad-hoc snippet substituted for the repository's existing test suitecodeswesmith/agronholm__typeguard.b6a7e438
Applies when
code: the program's verification of a change to a library/package consists solely of a hand-written snippet (python -c ..., a scratch script) exercising one or two code paths
Pattern
The repository ships a test suite and the environment ships a test runner, but the program never invokes them; it substitutes a narrow self-authored check that touches a single entry point. Behavior in every other module the change can reach is never observed, so regressions are absorbed silently and the change is declared verified on evidence that cannot support it.
Detection procedure
  1. Scan the static facts' repo tree for a test directory containing multiple test_*.py files, and the package list for a test runner (pytest/nose/unittest availability) [reads: static facts — repo tree, python packages]
  2. Read the program and list every command or call it executes for verification [reads: code]
  3. The pattern is present if none of those commands invokes the runner over the repository's test directory (no pytest, python -m pytest, python -m unittest, tox), and the only verification is an inline/scratch snippet importing the package [reads: code]
Counter-example
A program that runs the repository suite (python -m pytest tests/ -q) and additionally runs a small snippet to illustrate the fixed behavior — the snippet is present, but it is not the sole evidence.
Discriminator
Absence of any invocation of the available runner against the shipped test files, not merely the presence of an ad-hoc snippet.
Consequence
Regressions in the untested modules go undetected; when the grader runs the shipped suite, previously passing test_*.py files can fail, turning a partially correct change into a failing submission. This explains the verification gap only — it does not by itself make the edited source wrong.
Evidence
Verification consisted entirely of an inline python -c snippet exercising one decorator on one function, while the repository contained a dozen test_*.py modules and pytest was installed and never invoked.
id 1a8480df2674 · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.func_basic__aw9j63hw
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Scan the static facts' repo tree for a test directory containing multiple `test_*.py` files, and the package list for a test runner (`pytest`/`nose`/`unittest` availability) [reads: static facts \u2014 repo tree, python packages]",
 "prediction": "Regressions in the untested modules go undetected; when the grader runs the shipped suite, previously passing `test_*.py` files can fail, turning a partially correct change into a failing submission. This explains the verification gap only \u2014 it does not by itself make the edited source wrong."
}
raw text (what the judge reads)
### Ad-hoc snippet substituted for the repository's existing test suite
- **Applies when**: `code`: the program's verification of a change to a library/package consists solely of a hand-written snippet (`python -c ...`, a scratch script) exercising one or two code paths
- **Pattern**: The repository ships a test suite and the environment ships a test runner, but the program never invokes them; it substitutes a narrow self-authored check that touches a single entry point. Behavior in every other module the change can reach is never observed, so regressions are absorbed silently and the change is declared verified on evidence that cannot support it.
- **Detection procedure**:
  1. Scan the static facts' repo tree for a test directory containing multiple `test_*.py` files, and the package list for a test runner (`pytest`/`nose`/`unittest` availability) [reads: static facts — repo tree, python packages]
  2. Read the program and list every command or call it executes for verification [reads: code]
  3. The pattern is present if none of those commands invokes the runner over the repository's test directory (no `pytest`, `python -m pytest`, `python -m unittest`, `tox`), and the only verification is an inline/scratch snippet importing the package [reads: code]
- **Counter-example**: A program that runs the repository suite (`python -m pytest tests/ -q`) and additionally runs a small snippet to illustrate the fixed behavior — the snippet is present, but it is not the sole evidence.
- **Discriminator**: Absence of any invocation of the available runner against the shipped test files, not merely the presence of an ad-hoc snippet.
- **Consequence**: Regressions in the untested modules go undetected; when the grader runs the shipped suite, previously passing `test_*.py` files can fail, turning a partially correct change into a failing submission. This explains the verification gap only — it does not by itself make the edited source wrong.
- **Evidence**: Verification consisted entirely of an inline `python -c` snippet exercising one decorator on one function, while the repository contained a dozen `test_*.py` modules and `pytest` was installed and never invoked.
9Deleting module-level definitions that other modules still importcodeswesmith/agronholm__typeguard.b6a7e438
Applies when
code: the change edits an existing source file in a multi-module package (a diff or before/after view is available) and removes one or more top-level def/class/assignment definitions.
Pattern
A refactor that was meant to touch one construct also deletes other public top-level definitions from a module, while the package's re-export module or sibling modules still import those names by name. Nothing in the edited file re-creates them, so the very first import of the package raises.
Detection procedure
  1. In the diff (or by comparing the shown file against the described edit), list every top-level name whose def/class/assignment was removed and not re-added anywhere in the file's final text. [reads: code]
  2. Read the task statement and check whether removing those names is part of what was asked (e.g. "delete X", "move X to module Y"); if the task only asks for a localized fix (a signature, a type annotation, a bug in one branch), the deletions are collateral. [reads: task]
  3. Confirm the deleted names are externally referenced: the repo tree lists a package __init__.py / sibling modules, and the deleted names are the module's public API (no leading underscore) or are named in an from .<module> import <name> visible in the code. Also check the final file for signs of truncation (file ends mid-module, no trailing newline, imports left that nothing uses). [reads: static facts — repo tree; code]
Counter-example
A diff that deletes a top-level function and, in the same change, either re-defines it (renamed/relocated) with all importing sites updated in the same diff, or deletes a name that is underscore-private and referenced only from the other lines the diff also deletes.
Discriminator
The failing case removes a name with no replacement definition and no corresponding update at the import site; the safe case either preserves the name (possibly moved, with the import edited in the same change) or removes only names whose sole references are removed too.
Consequence
ImportError: cannot import name '<X>' from '<module>' (or AttributeError/ModuleNotFoundError) at import/collection time — typically before any test runs, so the entire suite errors and the score is zero rather than partially reduced.
Evidence
A diff intended to adjust one function signature also deleted the module's remaining top-level functions (leaving a truncated file with unused imports and no trailing newline); the package's __init__.py still did from ._decorators import <deleted name>, and the test run aborted during plugin/module import with ImportError: cannot import name ... from ....
id 818f2e3649bb · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.func_basic__aw9j63hw
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. In the diff (or by comparing the shown file against the described edit), list every top-level name whose `def`/`class`/assignment was removed and not re-added anywhere in the file's final text. [reads: code]",
 "prediction": "`ImportError: cannot import name '<X>' from '<module>'` (or `AttributeError`/`ModuleNotFoundError`) at import/collection time \u2014 typically before any test runs, so the entire suite errors and the score is zero rather than partially reduced."
}
raw text (what the judge reads)
### Deleting module-level definitions that other modules still import
- **Applies when**: `code`: the change edits an existing source file in a multi-module package (a diff or before/after view is available) and removes one or more top-level `def`/`class`/assignment definitions.
- **Pattern**: A refactor that was meant to touch one construct also deletes other public top-level definitions from a module, while the package's re-export module or sibling modules still import those names by name. Nothing in the edited file re-creates them, so the very first import of the package raises.
- **Detection procedure**:
  1. In the diff (or by comparing the shown file against the described edit), list every top-level name whose `def`/`class`/assignment was removed and not re-added anywhere in the file's final text. [reads: code]
  2. Read the task statement and check whether removing those names is part of what was asked (e.g. "delete X", "move X to module Y"); if the task only asks for a localized fix (a signature, a type annotation, a bug in one branch), the deletions are collateral. [reads: task]
  3. Confirm the deleted names are externally referenced: the repo tree lists a package `__init__.py` / sibling modules, and the deleted names are the module's public API (no leading underscore) or are named in an `from .<module> import <name>` visible in the code. Also check the final file for signs of truncation (file ends mid-module, no trailing newline, imports left that nothing uses). [reads: static facts — repo tree; code]
- **Counter-example**: A diff that deletes a top-level function and, in the same change, either re-defines it (renamed/relocated) with all importing sites updated in the same diff, or deletes a name that is underscore-private and referenced only from the other lines the diff also deletes.
- **Discriminator**: The failing case removes a name with no replacement definition and no corresponding update at the import site; the safe case either preserves the name (possibly moved, with the import edited in the same change) or removes only names whose sole references are removed too.
- **Consequence**: `ImportError: cannot import name '<X>' from '<module>'` (or `AttributeError`/`ModuleNotFoundError`) at import/collection time — typically before any test runs, so the entire suite errors and the score is zero rather than partially reduced.
- **Evidence**: A diff intended to adjust one function signature also deleted the module's remaining top-level functions (leaving a truncated file with unused imports and no trailing newline); the package's `__init__.py` still did `from ._decorators import <deleted name>`, and the test run aborted during plugin/module import with `ImportError: cannot import name ... from ...`.
9Forwarding a legacy parameter alongside its mutually exclusive replacementcodeswesmith/agronholm__typeguard.b6a7e438
Applies when
code: the program defines a thin wrapper that passes its own optional parameters straight through to a standard-library or third-party function which documents two parameters as mutually exclusive (one legacy/boolean override, one modern keyword)
Pattern
The wrapper hardcodes the modern keyword argument on the inner call while also forwarding the caller-supplied legacy override verbatim. When a caller actually supplies the legacy argument (non-None), the callee rejects the combination and raises, so the wrapper is only correct for the default path.
Detection procedure
  1. Locate wrapper functions whose body is a single return <library_function>(...) and whose signature contains an optional parameter defaulting to None. [reads: code]
  2. Check the inner call: does it pass that optional parameter through and additionally supply another keyword argument with a constant value (e.g. an optimization=/mode-style flag) to the same callee? [reads: code]
  3. Fire if there is no branch such as if <legacy_param> is not None: ... else: ... (or a dict of kwargs assembled conditionally) that ensures only one of the two arguments is non-None at call time. [reads: code]
Counter-example
A wrapper that builds the kwargs conditionally (kwargs = {"optimization": X} if legacy is None else {"debug_override": legacy}), or one that never exposes the legacy parameter in its own signature and always passes the modern keyword only.
Discriminator
The failing case lets a caller-controlled non-None value reach the callee simultaneously with the hardcoded exclusive keyword; the safe case makes the two mutually exclusive before the call.
Consequence
TypeError from the callee (message of the form "<param A> or <param B> must be set to None") on any call that exercises the legacy parameter; the wrapper passes smoke tests using defaults and fails the moment a test supplies the override. This explains the single reported traceback, not the broader loss of module functionality reported elsewhere.
Evidence
return cache_from_source(path, debug_override, optimization=OPTIMIZATION) in a wrapper raised TypeError: debug_override or optimization must be set to None as soon as a test called it with debug_override=False, after the same wrapper had passed with default arguments.
id b213a40cc6a6 · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.func_basic__aw9j63hw
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate wrapper functions whose body is a single `return <library_function>(...)` and whose signature contains an optional parameter defaulting to `None`. [reads: code]",
 "prediction": "`TypeError` from the callee (message of the form \"`<param A> or <param B> must be set to None`\") on any call that exercises the legacy parameter; the wrapper passes smoke tests using defaults and fails the moment a test supplies the override. This explains the single reported traceback, not the broader loss of module functionality reported elsewhere."
}
raw text (what the judge reads)
### Forwarding a legacy parameter alongside its mutually exclusive replacement
- **Applies when**: `code`: the program defines a thin wrapper that passes its own optional parameters straight through to a standard-library or third-party function which documents two parameters as mutually exclusive (one legacy/boolean override, one modern keyword)
- **Pattern**: The wrapper hardcodes the modern keyword argument on the inner call while also forwarding the caller-supplied legacy override verbatim. When a caller actually supplies the legacy argument (non-`None`), the callee rejects the combination and raises, so the wrapper is only correct for the default path.
- **Detection procedure**:
  1. Locate wrapper functions whose body is a single `return <library_function>(...)` and whose signature contains an optional parameter defaulting to `None`. [reads: code]
  2. Check the inner call: does it pass that optional parameter through *and* additionally supply another keyword argument with a constant value (e.g. an `optimization=`/mode-style flag) to the same callee? [reads: code]
  3. Fire if there is no branch such as `if <legacy_param> is not None: ... else: ...` (or a `dict` of kwargs assembled conditionally) that ensures only one of the two arguments is non-`None` at call time. [reads: code]
- **Counter-example**: A wrapper that builds the kwargs conditionally (`kwargs = {"optimization": X} if legacy is None else {"debug_override": legacy}`), or one that never exposes the legacy parameter in its own signature and always passes the modern keyword only.
- **Discriminator**: The failing case lets a caller-controlled non-`None` value reach the callee simultaneously with the hardcoded exclusive keyword; the safe case makes the two mutually exclusive before the call.
- **Consequence**: `TypeError` from the callee (message of the form "`<param A> or <param B> must be set to None`") on any call that exercises the legacy parameter; the wrapper passes smoke tests using defaults and fails the moment a test supplies the override. This explains the single reported traceback, not the broader loss of module functionality reported elsewhere.
- **Evidence**: `return cache_from_source(path, debug_override, optimization=OPTIMIZATION)` in a wrapper raised `TypeError: debug_override or optimization must be set to None` as soon as a test called it with `debug_override=False`, after the same wrapper had passed with default arguments.
9Public symbol implied by the repo's test suite is absent from the modulecodeswesmith/agronholm__typeguard.b6a7e438
Applies when
code: the submission edits library source in a repo whose static facts list a tests/ directory with per-feature test modules
Pattern
The edit removes or never defines a top-level symbol whose name matches a test module or a documented feature of the package, so the test suite cannot even import the module under test.
Detection procedure
  1. List the test module names and doc pages shown in the static facts repo tree (e.g. tests/test_<feature>.py, docs/api.rst). [reads: static facts — repo tree]
  2. For each such <feature> name, search the edited source module for a top-level def <feature>, class <feature>, or an import that re-exports it. [reads: code]
  3. Fire if the edited module is the natural home of that symbol (its imports/docstrings mention the feature) yet no definition or re-export of it remains. [reads: code]
Counter-example
The feature is implemented in a sibling module of the same package and the edited file never claimed to define it — the test module name maps to a different source file, which the edit does not touch.
Discriminator
The goes-wrong case has the edited module retaining imports, docstrings, or helper functions that exist solely to serve the missing symbol; the safe case has no such orphaned support code in this file.
Consequence
Import-time ImportError: cannot import name ... during test collection, failing the entire suite regardless of the correctness of the intended change.
Evidence
The decorator module was reduced to a single cell-making helper while the repo's test tree contained a dedicated test module for the decorator it used to define, and the helper's only caller had been deleted.
id d98d24361249 · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.func_basic__aw9j63hw
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List the test module names and doc pages shown in the static facts repo tree (e.g. `tests/test_<feature>.py`, `docs/api.rst`). [reads: static facts \u2014 repo tree]",
 "prediction": "Import-time `ImportError: cannot import name ...` during test collection, failing the entire suite regardless of the correctness of the intended change."
}
raw text (what the judge reads)
### Public symbol implied by the repo's test suite is absent from the module
- **Applies when**: `code`: the submission edits library source in a repo whose static facts list a `tests/` directory with per-feature test modules
- **Pattern**: The edit removes or never defines a top-level symbol whose name matches a test module or a documented feature of the package, so the test suite cannot even import the module under test.
- **Detection procedure**:
  1. List the test module names and doc pages shown in the static facts repo tree (e.g. `tests/test_<feature>.py`, `docs/api.rst`). [reads: static facts — repo tree]
  2. For each such `<feature>` name, search the edited source module for a top-level `def <feature>`, `class <feature>`, or an import that re-exports it. [reads: code]
  3. Fire if the edited module is the natural home of that symbol (its imports/docstrings mention the feature) yet no definition or re-export of it remains. [reads: code]
- **Counter-example**: The feature is implemented in a sibling module of the same package and the edited file never claimed to define it — the test module name maps to a different source file, which the edit does not touch.
- **Discriminator**: The goes-wrong case has the edited module retaining imports, docstrings, or helper functions that exist solely to serve the missing symbol; the safe case has no such orphaned support code in this file.
- **Consequence**: Import-time `ImportError: cannot import name ...` during test collection, failing the entire suite regardless of the correctness of the intended change.
- **Evidence**: The decorator module was reduced to a single cell-making helper while the repo's test tree contained a dedicated test module for the decorator it used to define, and the helper's only caller had been deleted.
10Falsy value used as the "nothing saved" marker in a save/restore paircodeswesmith/rsalmei__alive-progress.35853799
Applies when
code: the program saves prior state (a stream, handler, attribute, config value, env var, working directory) before mutating it and restores it later in an uninstall/teardown/__exit__/finally path
Pattern
The saved-state map stores None (or another falsy value) both for "we failed to mutate this object, do not touch it" and for "the legitimate original value was falsy", and the restore loop guards with a plain truthiness test (if original:, if saved, filter(None, ...)). Entries whose genuine original value is falsy are silently skipped, so those objects stay wired to the temporary/wrapper value forever.
Detection procedure
  1. Find the restore site: a loop/comprehension over a dict or list of (target, saved_value) pairs that calls a setter (setStream, setattr, assignment, os.environ[...] = ...) and is filtered by a boolean test on the saved value. [reads: code]
  2. Find the save site that populated that container, and check how a missing/failed save is signalled — an except ...: pass (function implicitly returns None), a .get(key) default, or a setter whose return value is None when nothing changed. [reads: code]
  3. Fire if the same falsy value can also arrive from a successful save of a legitimately falsy original (e.g. an object whose attribute is None until first use, an unset env var, 0, '') — i.e. the guard cannot tell "skip" from "restore to falsy". [reads: code]
Counter-example
The same loop where the guard is a dedicated sentinel comparison (if saved is not _MISSING), a membership test (for k in saved_map with failures never inserted), or where the saved value is provably never falsy (e.g. always a live file object captured before mutation).
Discriminator
The "skip" marker and a valid saved value are the same falsy object, and the guard distinguishes them only by truthiness — versus a distinct sentinel/membership check, or a value domain that excludes falsy.
Consequence
Teardown leaves a subset of targets still pointing at the instrumented/wrapper object; a test asserting target.attr is original_value after teardown fails, and later writes go through a wrapper whose backing state was cleared, typically surfacing as AttributeError, ValueError: I/O operation on closed file, or lost/duplicated output. The failure only appears for targets whose original value is falsy, so a suite that only exercises non-falsy originals passes while the reported scenario stays broken.
Evidence
Restore written as [h.setStream(orig) for h, orig in before.items() if orig] while the save path used except Exception: pass (implicit None) and lazily-initialised handlers legitimately had stream is None; switching to an explicit _HOOK_FAILED sentinel plus if orig is not _HOOK_FAILED made the restore correct with the suite still fully green.
id 9a8e396a5ce9 · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Find the restore site: a loop/comprehension over a dict or list of `(target, saved_value)` pairs that calls a setter (`setStream`, `setattr`, assignment, `os.environ[...] = ...`) and is filtered by a boolean test on the saved value. [reads: code]",
 "prediction": "Teardown leaves a subset of targets still pointing at the instrumented/wrapper object; a test asserting `target.attr is original_value` after teardown fails, and later writes go through a wrapper whose backing state was cleared, typically surfacing as `AttributeError`, `ValueError: I/O operation on closed file`, or lost/duplicated output. The failure only appears for targets whose original value is falsy, so a suite that only exercises non-falsy originals passes while the reported scenario stays broken."
}
raw text (what the judge reads)
### Falsy value used as the "nothing saved" marker in a save/restore pair
- **Applies when**: `code`: the program saves prior state (a stream, handler, attribute, config value, env var, working directory) before mutating it and restores it later in an uninstall/teardown/`__exit__`/`finally` path
- **Pattern**: The saved-state map stores `None` (or another falsy value) both for "we failed to mutate this object, do not touch it" and for "the legitimate original value was falsy", and the restore loop guards with a plain truthiness test (`if original:`, `if saved`, `filter(None, ...)`). Entries whose genuine original value is falsy are silently skipped, so those objects stay wired to the temporary/wrapper value forever.
- **Detection procedure**:
  1. Find the restore site: a loop/comprehension over a dict or list of `(target, saved_value)` pairs that calls a setter (`setStream`, `setattr`, assignment, `os.environ[...] = ...`) and is filtered by a boolean test on the saved value. [reads: code]
  2. Find the save site that populated that container, and check how a missing/failed save is signalled — an `except ...: pass` (function implicitly returns `None`), a `.get(key)` default, or a setter whose return value is `None` when nothing changed. [reads: code]
  3. Fire if the same falsy value can also arrive from a successful save of a legitimately falsy original (e.g. an object whose attribute is `None` until first use, an unset env var, `0`, `''`) — i.e. the guard cannot tell "skip" from "restore to falsy". [reads: code]
- **Counter-example**: The same loop where the guard is a dedicated sentinel comparison (`if saved is not _MISSING`), a membership test (`for k in saved_map` with failures never inserted), or where the saved value is provably never falsy (e.g. always a live file object captured before mutation).
- **Discriminator**: The "skip" marker and a valid saved value are the *same* falsy object, and the guard distinguishes them only by truthiness — versus a distinct sentinel/membership check, or a value domain that excludes falsy.
- **Consequence**: Teardown leaves a subset of targets still pointing at the instrumented/wrapper object; a test asserting `target.attr is original_value` after teardown fails, and later writes go through a wrapper whose backing state was cleared, typically surfacing as `AttributeError`, `ValueError: I/O operation on closed file`, or lost/duplicated output. The failure only appears for targets whose original value is falsy, so a suite that only exercises non-falsy originals passes while the reported scenario stays broken.
- **Evidence**: Restore written as `[h.setStream(orig) for h, orig in before.items() if orig]` while the save path used `except Exception: pass` (implicit `None`) and lazily-initialised handlers legitimately had `stream is None`; switching to an explicit `_HOOK_FAILED` sentinel plus `if orig is not _HOOK_FAILED` made the restore correct with the suite still fully green.
10Restoring saved state from a setter's "no change" return valuecodeswesmith/rsalmei__alive-progress.35853799
Applies when
code: a program installs temporary instrumentation/patches on live objects (streams, handlers, hooks, env vars, attributes) and later restores them in an uninstall/teardown/context-exit routine.
Pattern
The "original" value is captured from the return value of the mutating call itself (e.g. old = obj.setStream(new)), where the API returns None to mean "nothing changed" rather than "the previous value was None". Teardown then feeds that None straight back into the setter, clobbering the object's live state instead of restoring it. The defect typically appears as the removal or weakening of a truthiness guard in the restore loop — e.g. if original replaced by if original is not SOME_SENTINEL — so entries that were never really swapped now get "restored" to None.
Detection procedure
  1. Find the install/setup routine and locate where previous state is recorded; check whether the recorded value is the return value of the mutating call (saved[obj] = obj.setX(new), old = patch(...)) rather than a value read from the object before mutation. [reads: code]
  2. Read the task statement / issue description to confirm the routine under repair is the restore/uninstall path and that correct restoration of the original targets is the required behavior. [reads: task]
  3. In the restore loop, check the filter on each saved entry: does it exclude only a program-defined sentinel (or nothing at all), so that a saved value of None — the setter's "unchanged" signal, and also the value produced for handlers whose stream was never swapped — is passed back into the setter? If yes, the rubric fires. [reads: code]
Counter-example
A program that snapshots state by reading the attribute directly before mutating (original = handler.stream; handler.setStream(hook)), or that keeps the return-value capture but retains a truthiness/is not None guard in the restore loop (... if original), so unchanged entries are simply skipped.
Discriminator
Goes wrong when (a) the saved value originates from the mutator's return and (b) the restore loop's condition admits None into the setter. Safe when either the snapshot is taken by reading the object's own attribute, or the restore loop still filters out None/falsy saved values.
Consequence
Objects whose state was never actually swapped get their live state set to None at teardown. Predict test failures asserting that handlers/streams equal their original objects, and, on the next use of the affected object, AttributeError: 'NoneType' object has no attribute 'write'/'flush' (in logging.StreamHandler.emit, possibly surfaced through handleError), or ValueError/TypeError from downstream code that assumes a non-None target. Also predict the originally reported symptom remains unfixed, since the change alters a code path that was already correct.
Evidence
A teardown loop was changed from [handler.setStream(original) for handler, original in before_handlers.items() if original] to ... if original is not _HOOK_FAILED, while before_handlers was populated with h.setStream(hook) return values — an API that returns None when the stream is unchanged — so unchanged handlers are reassigned setStream(None).
id 466dec8777a3 · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Find the install/setup routine and locate where previous state is recorded; check whether the recorded value is the *return value* of the mutating call (`saved[obj] = obj.setX(new)`, `old = patch(...)`) rather than a value read from the object before mutation. [reads: code]",
 "prediction": "Objects whose state was never actually swapped get their live state set to `None` at teardown. Predict test failures asserting that handlers/streams equal their original objects, and, on the next use of the affected object, `AttributeError: 'NoneType' object has no attribute 'write'`/`'flush'` (in `logging.StreamHandler.emit`, possibly surfaced through `handleError`), or `ValueError`/`TypeError` from downstream code that assumes a non-`None` target. Also predict the originally reported symptom remains unfixed, since the change alters a code path that was already correct."
}
raw text (what the judge reads)
### Restoring saved state from a setter's "no change" return value

- **Applies when**: `code`: a program installs temporary instrumentation/patches on live objects (streams, handlers, hooks, env vars, attributes) and later restores them in an uninstall/teardown/context-exit routine.
- **Pattern**: The "original" value is captured from the *return value of the mutating call itself* (e.g. `old = obj.setStream(new)`), where the API returns `None` to mean "nothing changed" rather than "the previous value was None". Teardown then feeds that `None` straight back into the setter, clobbering the object's live state instead of restoring it. The defect typically appears as the removal or weakening of a truthiness guard in the restore loop — e.g. `if original` replaced by `if original is not SOME_SENTINEL` — so entries that were never really swapped now get "restored" to `None`.
- **Detection procedure**:
  1. Find the install/setup routine and locate where previous state is recorded; check whether the recorded value is the *return value* of the mutating call (`saved[obj] = obj.setX(new)`, `old = patch(...)`) rather than a value read from the object before mutation. [reads: code]
  2. Read the task statement / issue description to confirm the routine under repair is the restore/uninstall path and that correct restoration of the original targets is the required behavior. [reads: task]
  3. In the restore loop, check the filter on each saved entry: does it exclude *only* a program-defined sentinel (or nothing at all), so that a saved value of `None` — the setter's "unchanged" signal, and also the value produced for handlers whose stream was never swapped — is passed back into the setter? If yes, the rubric fires. [reads: code]
- **Counter-example**: A program that snapshots state by reading the attribute directly before mutating (`original = handler.stream; handler.setStream(hook)`), or that keeps the return-value capture but retains a truthiness/`is not None` guard in the restore loop (`... if original`), so unchanged entries are simply skipped.
- **Discriminator**: Goes wrong when (a) the saved value originates from the mutator's return and (b) the restore loop's condition admits `None` into the setter. Safe when either the snapshot is taken by reading the object's own attribute, or the restore loop still filters out `None`/falsy saved values.
- **Consequence**: Objects whose state was never actually swapped get their live state set to `None` at teardown. Predict test failures asserting that handlers/streams equal their original objects, and, on the next use of the affected object, `AttributeError: 'NoneType' object has no attribute 'write'`/`'flush'` (in `logging.StreamHandler.emit`, possibly surfaced through `handleError`), or `ValueError`/`TypeError` from downstream code that assumes a non-`None` target. Also predict the originally reported symptom remains unfixed, since the change alters a code path that was already correct.
- **Evidence**: A teardown loop was changed from `[handler.setStream(original) for handler, original in before_handlers.items() if original]` to `... if original is not _HOOK_FAILED`, while `before_handlers` was populated with `h.setStream(hook)` return values — an API that returns `None` when the stream is unchanged — so unchanged handlers are reassigned `setStream(None)`.
10Global snapshotted at construction, redirected afterwardscodeswesmith/rsalmei__alive-progress.35853799
Applies when
code: a component captures a process-global (e.g. sys.stdout/sys.stderr, a logger's handler list, os.environ, a default connection) into an internal variable when it is created or installed, and some caller in the same repo also reassigns that global.
Pattern
The caller builds/installs the component first and then replaces the global with a substitute (a StringIO, a fake, a temp object), and afterwards reads results from that substitute. The component is still bound to the pre-redirect object, so the substitute never receives anything and attribute access on the wrong object explodes.
Detection procedure
  1. In the module under test, find where a global is read into state at construction/install time — e.g. a factory body line like base = sys.stdout, sys.stderr, or self._orig = logging.root.handlers. Note that this executes once, when the object is created. [reads: code]
  2. In the calling/verification code, find every statement that reassigns that same global (sys.stdout = io.StringIO(), logging.root.handlers = [...]). [reads: code]
  3. Check statement order inside that caller: is the component constructed (or install()/start() called) before the reassignment, while the later assertions/reads target the substitute object? If yes, the rubric fires. [reads: code]
Counter-example
The same test that performs sys.stdout = io.StringIO() (or the equivalent redirect / monkeypatch.setattr) before calling the factory, or a component that re-reads the global inside each method instead of caching it at construction — both wire the substitute through correctly.
Discriminator
The snapshot happens in the constructor/factory body (once, at creation) and the redirect statement executes after that call; safe code either redirects first or never caches the global.
Consequence
AttributeError when a substitute-specific method is reached through a delegating wrapper (e.g. __getattr__-forwarding proxy over the real stream: '_io.TextIOWrapper' object has no attribute 'getvalue'), or a silent assertion failure because the captured buffer is empty; the real stream is polluted with output that was supposed to be captured. The script terminates non-zero even though the library change itself may be fine.
Evidence
A factory cached base = sys.stdout, sys.stderr at creation; the verification script created the manager, then did sys.stdout = io.StringIO(), installed, and called sys.stdout.getvalue() — raising AttributeError: '_io.TextIOWrapper' object has no attribute 'getvalue' and printing the traceback to the real terminal.
id a535f6926f18 · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. In the module under test, find where a global is read into state at construction/install time \u2014 e.g. a factory body line like `base = sys.stdout, sys.stderr`, or `self._orig = logging.root.handlers`. Note that this executes once, when the object is created. [reads: code]",
 "prediction": "`AttributeError` when a substitute-specific method is reached through a delegating wrapper (e.g. `__getattr__`-forwarding proxy over the real stream: `'_io.TextIOWrapper' object has no attribute 'getvalue'`), or a silent assertion failure because the captured buffer is empty; the real stream is polluted with output that was supposed to be captured. The script terminates non-zero even though the library change itself may be fine."
}
raw text (what the judge reads)
### Global snapshotted at construction, redirected afterwards

- **Applies when**: `code`: a component captures a process-global (e.g. `sys.stdout`/`sys.stderr`, a logger's handler list, `os.environ`, a default connection) into an internal variable when it is created or installed, and some caller in the same repo also reassigns that global.
- **Pattern**: The caller builds/installs the component first and *then* replaces the global with a substitute (a `StringIO`, a fake, a temp object), and afterwards reads results from that substitute. The component is still bound to the pre-redirect object, so the substitute never receives anything and attribute access on the wrong object explodes.
- **Detection procedure**:
  1. In the module under test, find where a global is read into state at construction/install time — e.g. a factory body line like `base = sys.stdout, sys.stderr`, or `self._orig = logging.root.handlers`. Note that this executes once, when the object is created. [reads: code]
  2. In the calling/verification code, find every statement that reassigns that same global (`sys.stdout = io.StringIO()`, `logging.root.handlers = [...]`). [reads: code]
  3. Check statement order inside that caller: is the component constructed (or `install()`/`start()` called) *before* the reassignment, while the later assertions/reads target the substitute object? If yes, the rubric fires. [reads: code]
- **Counter-example**: The same test that performs `sys.stdout = io.StringIO()` (or the equivalent redirect / `monkeypatch.setattr`) *before* calling the factory, or a component that re-reads the global inside each method instead of caching it at construction — both wire the substitute through correctly.
- **Discriminator**: The snapshot happens in the constructor/factory body (once, at creation) **and** the redirect statement executes after that call; safe code either redirects first or never caches the global.
- **Consequence**: `AttributeError` when a substitute-specific method is reached through a delegating wrapper (e.g. `__getattr__`-forwarding proxy over the real stream: `'_io.TextIOWrapper' object has no attribute 'getvalue'`), or a silent assertion failure because the captured buffer is empty; the real stream is polluted with output that was supposed to be captured. The script terminates non-zero even though the library change itself may be fine.
- **Evidence**: A factory cached `base = sys.stdout, sys.stderr` at creation; the verification script created the manager, then did `sys.stdout = io.StringIO()`, installed, and called `sys.stdout.getvalue()` — raising `AttributeError: '_io.TextIOWrapper' object has no attribute 'getvalue'` and printing the traceback to the real terminal.
10Global-state mutation in test_*.py with cleanup only on the success pathcodeswesmith/rsalmei__alive-progress.35853799
Applies when
code: the change adds files matching test_*.py (collectible by pytest) that mutate process-wide state — reassigning sys.stdout/sys.stderr, logging.getLogger().addHandler(...), module attributes, os.environ, working directory — inside the test function body.
Pattern
The mutation is undone by plain statements placed at the end of the function, after the asserts, with no try/finally, fixture, or context manager. Any assertion failure or exception skips the restore, so the polluted global leaks into every later test in the same interpreter session.
Detection procedure
  1. List the added/modified test files and locate every statement that mutates a process-global (assignment to sys.stdout/sys.stderr, addHandler, setLevel, os.chdir, os.environ[...] = ...). [reads: code]
  2. Locate the matching restore statement (sys.stdout = old, removeHandler, etc.) and check what lexically sits between mutation and restore. [reads: code]
  3. Fire if one or more assert/raise-capable statements sit between them and the restore is not inside finally:, a with block, a fixture teardown, or monkeypatch. [reads: code]
Counter-example
A test that performs the same mutation but wraps it in try: ... finally: sys.stdout = old, uses contextlib.redirect_stdout, or uses the monkeypatch/capsys/caplog fixtures — the global is restored even when an assertion fails.
Discriminator
Presence of failure-capable statements between mutation and restore with no unconditional teardown construct; safe code routes the restore through finally/fixture/context manager.
Consequence
When any assertion in such a test fails, later tests collected in the same pytest run see a replaced sys.stdout or an extra root-logger handler and fail with cascading AssertionError/AttributeError unrelated to their own subject; captured-output assertions elsewhere become unreliable and the run's failure count overstates the real defect.
Evidence
Added root-level test_*.py files did sys.stdout = io.StringIO() and root.addHandler(...) with restoration written as the last lines of each function; the first exception fired before those lines, leaving the interpreter's stdout wrapped and the traceback rendered through the instrumented stream.
id dfec915a2937 · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. List the added/modified test files and locate every statement that mutates a process-global (assignment to `sys.stdout`/`sys.stderr`, `addHandler`, `setLevel`, `os.chdir`, `os.environ[...] = ...`). [reads: code]",
 "prediction": "When any assertion in such a test fails, later tests collected in the same pytest run see a replaced `sys.stdout` or an extra root-logger handler and fail with cascading `AssertionError`/`AttributeError` unrelated to their own subject; captured-output assertions elsewhere become unreliable and the run's failure count overstates the real defect."
}
raw text (what the judge reads)
### Global-state mutation in test_*.py with cleanup only on the success path

- **Applies when**: `code`: the change adds files matching `test_*.py` (collectible by pytest) that mutate process-wide state — reassigning `sys.stdout`/`sys.stderr`, `logging.getLogger().addHandler(...)`, module attributes, `os.environ`, working directory — inside the test function body.
- **Pattern**: The mutation is undone by plain statements placed at the end of the function, after the `assert`s, with no `try/finally`, fixture, or context manager. Any assertion failure or exception skips the restore, so the polluted global leaks into every later test in the same interpreter session.
- **Detection procedure**:
  1. List the added/modified test files and locate every statement that mutates a process-global (assignment to `sys.stdout`/`sys.stderr`, `addHandler`, `setLevel`, `os.chdir`, `os.environ[...] = ...`). [reads: code]
  2. Locate the matching restore statement (`sys.stdout = old`, `removeHandler`, etc.) and check what lexically sits between mutation and restore. [reads: code]
  3. Fire if one or more `assert`/`raise`-capable statements sit between them and the restore is not inside `finally:`, a `with` block, a fixture teardown, or `monkeypatch`. [reads: code]
- **Counter-example**: A test that performs the same mutation but wraps it in `try: ... finally: sys.stdout = old`, uses `contextlib.redirect_stdout`, or uses the `monkeypatch`/`capsys`/`caplog` fixtures — the global is restored even when an assertion fails.
- **Discriminator**: Presence of failure-capable statements between mutation and restore **with no unconditional teardown construct**; safe code routes the restore through `finally`/fixture/context manager.
- **Consequence**: When any assertion in such a test fails, later tests collected in the same pytest run see a replaced `sys.stdout` or an extra root-logger handler and fail with cascading `AssertionError`/`AttributeError` unrelated to their own subject; captured-output assertions elsewhere become unreliable and the run's failure count overstates the real defect.
- **Evidence**: Added root-level `test_*.py` files did `sys.stdout = io.StringIO()` and `root.addHandler(...)` with restoration written as the last lines of each function; the first exception fired before those lines, leaving the interpreter's stdout wrapped and the traceback rendered through the instrumented stream.
10Proxy wrapper built around a possibly-None resource, with the null check guarding only the pre-use side effectcodeswesmith/rsalmei__alive-progress.35853799
Applies when
code: the program builds a wrapper/proxy/adapter object around another object's attribute (a stream, connection, file, socket, handle) and installs that wrapper in place of the original.
Pattern
The factory that builds the wrapper tests the underlying attribute for None/truthiness only to decide whether to run a preparatory side effect (flush/close/reset), then constructs and returns the wrapper unconditionally — so a wrapper around None gets installed. The wrapper's own methods forward attribute access blindly, so the failure is deferred to the first real use and surfaces far from the construction site.
Detection procedure
  1. Locate the factory/helper that returns the wrapper, e.g. a function containing if obj.attr: obj.attr.<side_effect>() immediately followed by return Wrapper(obj.attr). [reads: code]
  2. Check the task statement / code for whether the wrapped objects can legitimately have an unset attribute (e.g. lazily-opened handlers, delay=True file handlers, connections opened on first use, attributes initialised to None). [reads: task and code]
  3. Confirm the wrapper class (and the module-level functions it delegates to) dereference the stored attribute directly — self._x.write(...), self._x.flush(), getattr(self._x, item) — with no is None guard, and that no caller filters out the None case before installing the wrapper. [reads: code]
Counter-example
A factory that short-circuits — if obj.attr is None: return obj.attr (or return None) before constructing the wrapper — or a wrapper whose write/flush begin with if self._x is None: return. Same if obj.attr: line, but the None case never reaches a dereference.
Discriminator
The goes-wrong case has the truthiness/None test scoped to a single side-effect statement while the construction and every delegating method are unguarded; the safe case has the None case either excluded from wrapping or handled inside every delegating method.
Consequence
AttributeError: 'NoneType' object has no attribute 'write' (or 'flush', 'read', whatever the delegating method calls) raised at the first real use, not at install time. If the caller is a framework that catches exceptions internally, no test fails but the output/records routed through the wrapper are silently dropped and tracebacks are printed to stderr.
Evidence
get_hook_for did if handler.stream: handler.stream.flush() then return Hook(handler.stream), installing a proxy over None for a lazily-opened handler; every subsequent record produced AttributeError: 'NoneType' object has no attribute 'write' inside the proxy's flush.
id 46f5c88ab84c · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the factory/helper that returns the wrapper, e.g. a function containing `if obj.attr: obj.attr.<side_effect>()` immediately followed by `return Wrapper(obj.attr)`. [reads: code]",
 "prediction": "`AttributeError: 'NoneType' object has no attribute 'write'` (or `'flush'`, `'read'`, whatever the delegating method calls) raised at the first real use, not at install time. If the caller is a framework that catches exceptions internally, no test fails but the output/records routed through the wrapper are silently dropped and tracebacks are printed to stderr."
}
raw text (what the judge reads)
### Proxy wrapper built around a possibly-None resource, with the null check guarding only the pre-use side effect

- **Applies when**: `code`: the program builds a wrapper/proxy/adapter object around another object's attribute (a stream, connection, file, socket, handle) and installs that wrapper in place of the original.
- **Pattern**: The factory that builds the wrapper tests the underlying attribute for None/truthiness *only* to decide whether to run a preparatory side effect (flush/close/reset), then constructs and returns the wrapper unconditionally — so a wrapper around `None` gets installed. The wrapper's own methods forward attribute access blindly, so the failure is deferred to the first real use and surfaces far from the construction site.
- **Detection procedure**:
  1. Locate the factory/helper that returns the wrapper, e.g. a function containing `if obj.attr: obj.attr.<side_effect>()` immediately followed by `return Wrapper(obj.attr)`. [reads: code]
  2. Check the task statement / code for whether the wrapped objects can legitimately have an unset attribute (e.g. lazily-opened handlers, `delay=True` file handlers, connections opened on first use, attributes initialised to `None`). [reads: task and code]
  3. Confirm the wrapper class (and the module-level functions it delegates to) dereference the stored attribute directly — `self._x.write(...)`, `self._x.flush()`, `getattr(self._x, item)` — with no `is None` guard, and that no caller filters out the None case before installing the wrapper. [reads: code]
- **Counter-example**: A factory that short-circuits — `if obj.attr is None: return obj.attr` (or `return None`) before constructing the wrapper — or a wrapper whose `write`/`flush` begin with `if self._x is None: return`. Same `if obj.attr:` line, but the None case never reaches a dereference.
- **Discriminator**: The goes-wrong case has the truthiness/None test scoped to a single side-effect statement while the construction and every delegating method are unguarded; the safe case has the None case either excluded from wrapping or handled inside every delegating method.
- **Consequence**: `AttributeError: 'NoneType' object has no attribute 'write'` (or `'flush'`, `'read'`, whatever the delegating method calls) raised at the first real use, not at install time. If the caller is a framework that catches exceptions internally, no test fails but the output/records routed through the wrapper are silently dropped and tracebacks are printed to stderr.
- **Evidence**: `get_hook_for` did `if handler.stream: handler.stream.flush()` then `return Hook(handler.stream)`, installing a proxy over `None` for a lazily-opened handler; every subsequent record produced `AttributeError: 'NoneType' object has no attribute 'write'` inside the proxy's `flush`.
10Verification script that reports success because the exercised framework swallows the exceptions it triggerscodeswesmith/rsalmei__alive-progress.35853799
Applies when
code: the program ships its own ad-hoc test/verification script (asserts plus printed "PASSED"/"ALL TESTS PASSED" banners) that drives the changed code through a library API known to catch exceptions internally — logging (Handler.emit/handleError), thread targets, atexit, signal handlers, or any call the program itself wraps in try/except: pass.
Pattern
The script's assertions only inspect object identity or attribute state before and after (assert x.attr is original), never the side effect the code exists to produce, and it does not make the swallowing layer re-raise or capture its error output. Runtime failures triggered by the script are absorbed by the framework's internal error handler, so the script prints success while the feature is broken.
Detection procedure
  1. Locate the verification/test functions and list their assert statements. [reads: code]
  2. Locate, in those same functions, the calls that actually exercise the changed behaviour (e.g. logger.info(...), starting a thread, invoking a handler) and check whether the changed code's own source wraps its failure paths in except Exception: pass/sentinel returns, or whether the called API is one that catches emit-time errors. [reads: code]
  3. Confirm no assertion reads the produced artifact (file contents, captured stream value, records list) and there is no capture/redirection of stderr, no pytest caplog/capsys, and no monkeypatched handleError/raiseExceptions arrangement that would turn a swallowed error into a failure. [reads: code]
Counter-example
A script that performs the same identity assertions but additionally opens the target file / io.StringIO and asserts the expected text was written, or installs a handler that re-raises on error — the swallowed exception then becomes a visible assertion failure.
Discriminator
In the failing case every assertion is satisfiable while the exercised code path raises; in the safe case at least one assertion depends on the side effect that the raising path would have prevented.
Consequence
The program self-reports all tests passing while the real behaviour is broken; graders running the repo's real suite or checking runtime output see tracebacks (AttributeError/TypeError inside the framework's error handler) and lost output. Predict the submitted change to be judged incorrect despite its green banner.
Evidence
Hand-written scripts asserting only handler.stream is orig printed ✓✓✓ ALL TESTS PASSED ✓✓✓, while the same runs emitted repeated --- Logging error --- tracebacks (AttributeError: 'NoneType' object has no attribute 'write') that the logging module caught and printed instead of raising.
id fe4969010874 · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the verification/test functions and list their `assert` statements. [reads: code]",
 "prediction": "The program self-reports all tests passing while the real behaviour is broken; graders running the repo's real suite or checking runtime output see tracebacks (`AttributeError`/`TypeError` inside the framework's error handler) and lost output. Predict the submitted change to be judged incorrect despite its green banner."
}
raw text (what the judge reads)
### Verification script that reports success because the exercised framework swallows the exceptions it triggers

- **Applies when**: `code`: the program ships its own ad-hoc test/verification script (asserts plus printed "PASSED"/"ALL TESTS PASSED" banners) that drives the changed code through a library API known to catch exceptions internally — logging (`Handler.emit`/`handleError`), thread targets, `atexit`, signal handlers, or any call the program itself wraps in `try/except: pass`.
- **Pattern**: The script's assertions only inspect object identity or attribute state before and after (`assert x.attr is original`), never the side effect the code exists to produce, and it does not make the swallowing layer re-raise or capture its error output. Runtime failures triggered by the script are absorbed by the framework's internal error handler, so the script prints success while the feature is broken.
- **Detection procedure**:
  1. Locate the verification/test functions and list their `assert` statements. [reads: code]
  2. Locate, in those same functions, the calls that actually exercise the changed behaviour (e.g. `logger.info(...)`, starting a thread, invoking a handler) and check whether the changed code's own source wraps its failure paths in `except Exception: pass`/sentinel returns, or whether the called API is one that catches emit-time errors. [reads: code]
  3. Confirm no assertion reads the produced artifact (file contents, captured stream value, records list) and there is no capture/redirection of stderr, no `pytest` `caplog`/`capsys`, and no monkeypatched `handleError`/`raiseExceptions` arrangement that would turn a swallowed error into a failure. [reads: code]
- **Counter-example**: A script that performs the same identity assertions but additionally opens the target file / `io.StringIO` and asserts the expected text was written, or installs a handler that re-raises on error — the swallowed exception then becomes a visible assertion failure.
- **Discriminator**: In the failing case every assertion is satisfiable while the exercised code path raises; in the safe case at least one assertion depends on the side effect that the raising path would have prevented.
- **Consequence**: The program self-reports all tests passing while the real behaviour is broken; graders running the repo's real suite or checking runtime output see tracebacks (`AttributeError`/`TypeError` inside the framework's error handler) and lost output. Predict the submitted change to be judged incorrect despite its green banner.
- **Evidence**: Hand-written scripts asserting only `handler.stream is orig` printed `✓✓✓ ALL TESTS PASSED ✓✓✓`, while the same runs emitted repeated `--- Logging error ---` tracebacks (`AttributeError: 'NoneType' object has no attribute 'write'`) that the logging module caught and printed instead of raising.
10Unaddressed symptom in a multi-part bug reporttaskswesmith/rsalmei__alive-progress.35853799
Applies when
task: the issue text names two or more distinct defective aspects of the same routine (e.g. "the order of operations and the condition checking", "wrong ordering and wrong filter") and the code is a patch/diff to that routine
Pattern
The program fixes only one of the aspects the report enumerates and leaves the others byte-for-byte unchanged, so the hidden tests covering the untouched aspect still fail even though the diff "looks like" a fix.
Detection procedure
  1. Read the issue text and list every distinct defect noun-phrase it attributes to the routine (ordering of statements, condition/predicate, direction of an assignment, missing call, etc.). [reads: task]
  2. Locate the named routine in the diff and list which of its statements the patch actually modifies. [reads: code]
  3. Check each enumerated aspect against the modified statements: if the patch touches only predicates/conditions while the report also cites "order of operations" (or vice versa), and the statement sequence inside the routine is identical to the pre-image, the aspect is unaddressed. [reads: code]
Counter-example
A patch that reorders the statements and rewrites the predicate in the named routine, even if the reordering is a single line move; or a report that mentions several symptoms all traceable to one expression, where changing that expression covers them all.
Discriminator
The failing case has at least one enumerated defect category with zero corresponding textual change inside the named routine; the safe case has a change (or a demonstrably shared root expression) for every category.
Consequence
The task is not resolved — hidden tests exercising the untouched aspect (statement ordering / sequencing assertions, or state observed between the calls) fail, while tests for the fixed aspect may pass. Explains the bulk of the gap versus a patch whose only difference is that it also reorders those statements.
Evidence
The report cited both "order of operations and condition checking" in a teardown routine; the submitted patch rewrote only the filter predicate and left the statement sequence (flush(); clear(); restore_globals) exactly as-is, while the accepted fix reordered those three statements.
id 519bff4954a0 · mined from swesmith/rsalmei__alive-progress.35853799 rsalmei__alive-progress.35853799.func_basic__486lvw04
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Read the issue text and list every distinct defect noun-phrase it attributes to the routine (ordering of statements, condition/predicate, direction of an assignment, missing call, etc.). [reads: task]",
 "prediction": "The task is not resolved \u2014 hidden tests exercising the untouched aspect (statement ordering / sequencing assertions, or state observed between the calls) fail, while tests for the fixed aspect may pass. Explains the bulk of the gap versus a patch whose only difference is that it also reorders those statements."
}
raw text (what the judge reads)
### Unaddressed symptom in a multi-part bug report
- **Applies when**: `task`: the issue text names two or more distinct defective aspects of the same routine (e.g. "the order of operations *and* the condition checking", "wrong ordering *and* wrong filter") and the code is a patch/diff to that routine
- **Pattern**: The program fixes only one of the aspects the report enumerates and leaves the others byte-for-byte unchanged, so the hidden tests covering the untouched aspect still fail even though the diff "looks like" a fix.
- **Detection procedure**:
  1. Read the issue text and list every distinct defect noun-phrase it attributes to the routine (ordering of statements, condition/predicate, direction of an assignment, missing call, etc.). [reads: task]
  2. Locate the named routine in the diff and list which of its statements the patch actually modifies. [reads: code]
  3. Check each enumerated aspect against the modified statements: if the patch touches only predicates/conditions while the report also cites "order of operations" (or vice versa), and the statement sequence inside the routine is identical to the pre-image, the aspect is unaddressed. [reads: code]
- **Counter-example**: A patch that reorders the statements *and* rewrites the predicate in the named routine, even if the reordering is a single line move; or a report that mentions several symptoms all traceable to one expression, where changing that expression covers them all.
- **Discriminator**: The failing case has at least one enumerated defect category with zero corresponding textual change inside the named routine; the safe case has a change (or a demonstrably shared root expression) for every category.
- **Consequence**: The task is not resolved — hidden tests exercising the untouched aspect (statement ordering / sequencing assertions, or state observed between the calls) fail, while tests for the fixed aspect may pass. Explains the bulk of the gap versus a patch whose only difference is that it also reorders those statements.
- **Evidence**: The report cited both "order of operations and condition checking" in a teardown routine; the submitted patch rewrote only the filter predicate and left the statement sequence (`flush(); clear(); restore_globals`) exactly as-is, while the accepted fix reordered those three statements.
11Fix direction reversed: the edit re-creates the reported wrong valuetaskswesmith/cantools__cantools.0c6a7871
Applies when
task: the task is a bug report that states an "Expected" and an "Actual" value (or otherwise names a specific wrong output) for a named attribute, property, or function, and the code contains that named member.
Pattern
Instead of removing the transformation that produces the wrong output, the change adds a transformation to the accessor so that it now returns exactly the value the report calls wrong. The submitted code satisfies the bug description rather than the expectation.
Detection procedure
  1. From the task statement, extract the member named in the reproduction snippet, the value fed in / stored, the Expected output and the Actual (wrong) output. [reads: task]
  2. In the program, locate the definition of that member (property getter, method, or return statement) and read the expression it returns. [reads: code]
  3. Evaluate that expression symbolically on the input value from step 1: if it yields the Actual (wrong) value rather than the Expected value — e.g. it returns self._x + 1, value * 2, idx - 1 where the report says the plain stored value is wanted — the fix runs backwards. [reads: code]
Counter-example
An accessor that returns a transformed value because the class deliberately stores an offset internal representation, where __init__/the setter apply the inverse transform, so feeding the reported input still yields the Expected output.
Discriminator
Substituting the report's input into the returned expression produces the string/number the report labels "Actual"/incorrect, and no inverse transform exists anywhere on the write path; in the safe case the same substitution produces the "Expected" value.
Consequence
The reproduction snippet in the issue still prints the wrong value; every hidden test asserting the documented expected value fails with AssertionError, and tests that previously passed on the unmodified accessor now regress. The change is a net negative versus doing nothing.
Evidence
A property getter was changed from return self._repetitions to return self._repetitions + 1 while the report asked that a stored 1 be returned as 1, making the reported off-by-one the actual behavior of the submitted code.
id 7f9a76b2141d · mined from swesmith/cantools__cantools.0c6a7871 cantools__cantools.0c6a7871.combine_file__gj056w8x
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. From the task statement, extract the member named in the reproduction snippet, the value fed in / stored, the Expected output and the Actual (wrong) output. [reads: task]",
 "prediction": "The reproduction snippet in the issue still prints the wrong value; every hidden test asserting the documented expected value fails with `AssertionError`, and tests that previously passed on the unmodified accessor now regress. The change is a net negative versus doing nothing."
}
raw text (what the judge reads)
### Fix direction reversed: the edit re-creates the reported wrong value
- **Applies when**: `task`: the task is a bug report that states an "Expected" and an "Actual" value (or otherwise names a specific wrong output) for a named attribute, property, or function, and the code contains that named member.
- **Pattern**: Instead of removing the transformation that produces the wrong output, the change *adds* a transformation to the accessor so that it now returns exactly the value the report calls wrong. The submitted code satisfies the bug description rather than the expectation.
- **Detection procedure**:
  1. From the task statement, extract the member named in the reproduction snippet, the value fed in / stored, the Expected output and the Actual (wrong) output. [reads: task]
  2. In the program, locate the definition of that member (property getter, method, or return statement) and read the expression it returns. [reads: code]
  3. Evaluate that expression symbolically on the input value from step 1: if it yields the Actual (wrong) value rather than the Expected value — e.g. it returns `self._x + 1`, `value * 2`, `idx - 1` where the report says the plain stored value is wanted — the fix runs backwards. [reads: code]
- **Counter-example**: An accessor that returns a transformed value because the class deliberately stores an offset internal representation, where `__init__`/the setter apply the inverse transform, so feeding the reported input still yields the Expected output.
- **Discriminator**: Substituting the report's input into the returned expression produces the string/number the report labels "Actual"/incorrect, and no inverse transform exists anywhere on the write path; in the safe case the same substitution produces the "Expected" value.
- **Consequence**: The reproduction snippet in the issue still prints the wrong value; every hidden test asserting the documented expected value fails with `AssertionError`, and tests that previously passed on the unmodified accessor now regress. The change is a net negative versus doing nothing.
- **Evidence**: A property getter was changed from `return self._repetitions` to `return self._repetitions + 1` while the report asked that a stored `1` be returned as `1`, making the reported off-by-one the actual behavior of the submitted code.
11Property getter applies a transform its setter/constructor does not invertcodeswesmith/cantools__cantools.0c6a7871
Applies when
code: the program defines a @property getter together with a @<name>.setter (or an __init__ parameter) writing the same private attribute.
Pattern
The getter returns a computed expression of the backing attribute (offset, scale, unit conversion, formatting) while the setter and constructor store the incoming value verbatim, so obj.x = v; obj.x no longer returns v, and other members (__repr__, serializers, to_* methods) read the raw attribute and disagree with the property.
Detection procedure
  1. List each @property whose body is not a bare return self._attr — note the extra arithmetic/transform. [reads: code]
  2. For the same _attr, read the corresponding @x.setter body and the __init__ assignment. [reads: code]
  3. Fire if neither the setter nor __init__ applies the inverse transform (they assign the raw argument), and/or __repr__/serialization code emits self._attr directly while the property emits the transformed value. [reads: code]
Counter-example
A getter that scales a raw stored value whose setter divides by the same factor (or whose __init__ converts on the way in), and whose __repr__/serialization goes through the public property — round-trip is preserved.
Discriminator
Absence of the matching inverse operation on every write path for that attribute; the safe case has a compensating operation in the setter/constructor or reads only through the property.
Consequence
Round-trip tests (set then get, or load-file-then-compare, or save/reload cycles) fail with AssertionError; repeated serialize→parse cycles drift the value by the offset each time, and repr()/dump output disagrees with the attribute access in the same object.
Evidence
A getter returning self._x + 1 paired with a setter storing value unchanged and a __repr__ printing self._x, so assigning 1 read back as 2 while the repr still showed 1.
id 3e5edcb47f9e · mined from swesmith/cantools__cantools.0c6a7871 cantools__cantools.0c6a7871.combine_file__gj056w8x
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List each `@property` whose body is not a bare `return self._attr` \u2014 note the extra arithmetic/transform. [reads: code]",
 "prediction": "Round-trip tests (`set` then `get`, or load-file-then-compare, or save/reload cycles) fail with `AssertionError`; repeated serialize\u2192parse cycles drift the value by the offset each time, and `repr()`/dump output disagrees with the attribute access in the same object."
}
raw text (what the judge reads)
### Property getter applies a transform its setter/constructor does not invert
- **Applies when**: `code`: the program defines a `@property` getter together with a `@<name>.setter` (or an `__init__` parameter) writing the same private attribute.
- **Pattern**: The getter returns a computed expression of the backing attribute (offset, scale, unit conversion, formatting) while the setter and constructor store the incoming value verbatim, so `obj.x = v; obj.x` no longer returns `v`, and other members (`__repr__`, serializers, `to_*` methods) read the raw attribute and disagree with the property.
- **Detection procedure**:
  1. List each `@property` whose body is not a bare `return self._attr` — note the extra arithmetic/transform. [reads: code]
  2. For the same `_attr`, read the corresponding `@x.setter` body and the `__init__` assignment. [reads: code]
  3. Fire if neither the setter nor `__init__` applies the inverse transform (they assign the raw argument), and/or `__repr__`/serialization code emits `self._attr` directly while the property emits the transformed value. [reads: code]
- **Counter-example**: A getter that scales a raw stored value whose setter divides by the same factor (or whose `__init__` converts on the way in), and whose `__repr__`/serialization goes through the public property — round-trip is preserved.
- **Discriminator**: Absence of the matching inverse operation on every write path for that attribute; the safe case has a compensating operation in the setter/constructor or reads only through the property.
- **Consequence**: Round-trip tests (`set` then `get`, or load-file-then-compare, or save/reload cycles) fail with `AssertionError`; repeated serialize→parse cycles drift the value by the offset each time, and `repr()`/dump output disagrees with the attribute access in the same object.
- **Evidence**: A getter returning `self._x + 1` paired with a setter storing `value` unchanged and a `__repr__` printing `self._x`, so assigning `1` read back as `2` while the repr still showed `1`.
12Semantically inert refactor submitted where the task requires an observable behavior changetaskswesmith/paramiko__paramiko.23f92003
Applies when
task: the deliverable is a source change whose success is judged by an observable effect (a bug that must appear or disappear, a test whose outcome must flip, a behavior/output that must differ), and the candidate is supplied as a diff or patch to an existing codebase.
Pattern
Every hunk of the submitted diff swaps one spelling of an operation for an equivalent spelling — obj.method() replaced by helper(obj) from a utility module, a builder object replaced by the literal bytes/string it would have produced, an added import, a renamed local — while no condition, constant, comparison, branch, return value, or side-effecting statement is added or removed. The program therefore cannot produce the outcome the task is graded on, no matter how many files it touches.
Detection procedure
  1. Enumerate the hunks in the candidate diff and classify each: (a) changes a control-flow condition, a literal/constant, a comparison or boolean operator, a return value, or adds/removes a statement that has side effects or that other statements depend on; (b) substitutes one call form for another that the same codebase already treats as interchangeable (method call ⇄ module-level helper of the same name/purpose), constructs the same payload by a different route, adds an import, or reorders/renames without use changes. [reads: code]
  2. Read what the task states must become true after the change — which test must fail or pass, which behavior must differ, which defect must be present or absent. [reads: task]
  3. Fire if class (a) is empty: the whole diff is class (b) plus imports, so the program's observable behavior after the patch is identical to before, and the task's required difference is unrealized. [reads: code]
Counter-example
A diff that performs the same style of call-form substitution in several places and additionally deletes a branch, flips a boundary comparison, changes a default value, or removes a variable that a later statement reads — the refactoring is incidental cover for one genuinely behavior-changing hunk, and that hunk satisfies the task.
Discriminator
Presence of at least one hunk that alters control flow, a value, or the existence of a used definition. Pure interchange of equivalent call forms (including replacing a serializer object with its already-serialized equivalent) has no such hunk; the near miss does.
Consequence
The graded criterion is unmet outright — the required test-visible behavior never changes, so the candidate scores at or near the floor on the behavioral objective while still "passing" as valid code. In the observed comparison this accounts for essentially the entire gap to the accepted solution; any remaining difference comes from where the accepted change was located.
Evidence
The submitted patch consisted solely of data.asbytes() → asbytes(data), packet.asbytes() → util.asbytes(packet), a corresponding import addition, and replacing a two-line message-builder with struct.pack(">I", VERSION) producing the identical bytes; the accepted solution instead deleted branches and a computed value inside a formatting routine, i.e. it changed behavior while the candidate did not.
id 9a698f17e5b3 · mined from swesmith/paramiko__paramiko.23f92003 paramiko__paramiko.23f92003.func_pm_remove_cond__fo86d8yl
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Enumerate the hunks in the candidate diff and classify each: (a) changes a control-flow condition, a literal/constant, a comparison or boolean operator, a return value, or adds/removes a statement that has side effects or that other statements depend on; (b) substitutes one call form for another that the same codebase already treats as interchangeable (method call \u21c4 module-level helper of the same name/purpose), constructs the same payload by a different route, adds an import, or reorders/renames without use changes. [reads: code]",
 "prediction": "The graded criterion is unmet outright \u2014 the required test-visible behavior never changes, so the candidate scores at or near the floor on the behavioral objective while still \"passing\" as valid code. In the observed comparison this accounts for essentially the entire gap to the accepted solution; any remaining difference comes from where the accepted change was located."
}
raw text (what the judge reads)
### Semantically inert refactor submitted where the task requires an observable behavior change
- **Applies when**: `task`: the deliverable is a source change whose success is judged by an observable effect (a bug that must appear or disappear, a test whose outcome must flip, a behavior/output that must differ), and the candidate is supplied as a diff or patch to an existing codebase.
- **Pattern**: Every hunk of the submitted diff swaps one spelling of an operation for an equivalent spelling — `obj.method()` replaced by `helper(obj)` from a utility module, a builder object replaced by the literal bytes/string it would have produced, an added import, a renamed local — while no condition, constant, comparison, branch, return value, or side-effecting statement is added or removed. The program therefore cannot produce the outcome the task is graded on, no matter how many files it touches.
- **Detection procedure**:
  1. Enumerate the hunks in the candidate diff and classify each: (a) changes a control-flow condition, a literal/constant, a comparison or boolean operator, a return value, or adds/removes a statement that has side effects or that other statements depend on; (b) substitutes one call form for another that the same codebase already treats as interchangeable (method call ⇄ module-level helper of the same name/purpose), constructs the same payload by a different route, adds an import, or reorders/renames without use changes. [reads: code]
  2. Read what the task states must become true after the change — which test must fail or pass, which behavior must differ, which defect must be present or absent. [reads: task]
  3. Fire if class (a) is empty: the whole diff is class (b) plus imports, so the program's observable behavior after the patch is identical to before, and the task's required difference is unrealized. [reads: code]
- **Counter-example**: A diff that performs the same style of call-form substitution in several places *and* additionally deletes a branch, flips a boundary comparison, changes a default value, or removes a variable that a later statement reads — the refactoring is incidental cover for one genuinely behavior-changing hunk, and that hunk satisfies the task.
- **Discriminator**: Presence of at least one hunk that alters control flow, a value, or the existence of a used definition. Pure interchange of equivalent call forms (including replacing a serializer object with its already-serialized equivalent) has no such hunk; the near miss does.
- **Consequence**: The graded criterion is unmet outright — the required test-visible behavior never changes, so the candidate scores at or near the floor on the behavioral objective while still "passing" as valid code. In the observed comparison this accounts for essentially the entire gap to the accepted solution; any remaining difference comes from where the accepted change was located.
- **Evidence**: The submitted patch consisted solely of `data.asbytes()` → `asbytes(data)`, `packet.asbytes()` → `util.asbytes(packet)`, a corresponding import addition, and replacing a two-line message-builder with `struct.pack(">I", VERSION)` producing the identical bytes; the accepted solution instead deleted branches and a computed value inside a formatting routine, i.e. it changed behavior while the candidate did not.
13Statement inserted at an indentation that splits an existing blockcodeswesmith/cantools__cantools.0c6a7871
Applies when
code: the program edits an existing function/loop body by inserting one or more statements into already-indented code
Pattern
Newly added lines are written at a shallower indentation than the block they were meant to join, while the pre-existing statements after them remain at the deeper indentation, so the file no longer parses (or the added statements silently execute outside the loop/branch they belong to).
Detection procedure
  1. Scan each function body for a run of consecutive statements and record the leading-whitespace column of each line [reads: code]
  2. Find any line whose indentation is less than the preceding statement's indentation, and check whether a later line in the same suite returns to the deeper indentation [reads: code]
  3. Confirm the deeper-indented line that follows is not introduced by a new block header (a line ending in : such as if/for/while/try/def/with) and is not inside brackets or a continuation of the previous expression [reads: code]
Counter-example
A dedent that permanently closes the inner block — every subsequent statement stays at the outer level or dedents further, or the deeper indentation resumes only after a fresh for:/if:/try: header; that code parses and runs normally.
Discriminator
The offending case has an indentation increase that is not preceded by a colon-terminated block header (Python's tokenizer emits INDENT with no enclosing suite); the safe case always reopens a block before re-indenting.
Consequence
IndentationError: unexpected indent (or SyntaxError) raised at module import, most likely surfacing as a collection error for every test module that imports the package — the entire test suite fails, not just the targeted behavior. Even if it parsed, the misplaced statements would execute once outside their intended loop rather than per iteration.
Evidence
Two assignment lines were appended at the enclosing-function indentation in the middle of a for body, and the next pre-existing statement (pdus = self._get_arxml_children(...)) stayed at the loop-body indentation, producing IndentationError: unexpected indent during import of the loader module.
id 29a5fb96fff6 · mined from swesmith/cantools__cantools.0c6a7871 cantools__cantools.0c6a7871.func_pm_remove_assign__ul0v6zcg
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Scan each function body for a run of consecutive statements and record the leading-whitespace column of each line [reads: code]",
 "prediction": "`IndentationError: unexpected indent` (or `SyntaxError`) raised at module import, most likely surfacing as a collection error for every test module that imports the package \u2014 the entire test suite fails, not just the targeted behavior. Even if it parsed, the misplaced statements would execute once outside their intended loop rather than per iteration."
}
raw text (what the judge reads)
### Statement inserted at an indentation that splits an existing block
- **Applies when**: `code`: the program edits an existing function/loop body by inserting one or more statements into already-indented code
- **Pattern**: Newly added lines are written at a shallower indentation than the block they were meant to join, while the pre-existing statements after them remain at the deeper indentation, so the file no longer parses (or the added statements silently execute outside the loop/branch they belong to).
- **Detection procedure**:
  1. Scan each function body for a run of consecutive statements and record the leading-whitespace column of each line [reads: code]
  2. Find any line whose indentation is *less* than the preceding statement's indentation, and check whether a later line in the same suite returns to the deeper indentation [reads: code]
  3. Confirm the deeper-indented line that follows is not introduced by a new block header (a line ending in `:` such as `if/for/while/try/def/with`) and is not inside brackets or a continuation of the previous expression [reads: code]
- **Counter-example**: A dedent that permanently closes the inner block — every subsequent statement stays at the outer level or dedents further, or the deeper indentation resumes only after a fresh `for:`/`if:`/`try:` header; that code parses and runs normally.
- **Discriminator**: The offending case has an indentation increase that is *not* preceded by a colon-terminated block header (Python's tokenizer emits INDENT with no enclosing suite); the safe case always reopens a block before re-indenting.
- **Consequence**: `IndentationError: unexpected indent` (or `SyntaxError`) raised at module import, most likely surfacing as a collection error for every test module that imports the package — the entire test suite fails, not just the targeted behavior. Even if it parsed, the misplaced statements would execute once outside their intended loop rather than per iteration.
- **Evidence**: Two assignment lines were appended at the enclosing-function indentation in the middle of a `for` body, and the next pre-existing statement (`pdus = self._get_arxml_children(...)`) stayed at the loop-body indentation, producing `IndentationError: unexpected indent` during import of the loader module.
13Statement reads identifiers that are never bound in its scopecodeswesmith/cantools__cantools.0c6a7871
Applies when
code: the program adds or modifies statements inside a method/function that assign from or pass along local variable names
Pattern
An added statement uses a bare identifier (as a value or as the object being attributed) that is never a parameter, never assigned earlier in that function, never imported, and never a module-level global — it only exists as a local in some other function of the same file, so the line raises NameError the first time it is reached.
Detection procedure
  1. List every bare identifier read (not self.x, not a literal, not a call to an imported name) by the newly added/modified statements in a function [reads: code]
  2. For each such identifier, search the enclosing function for a binding: parameter list, for target, assignment, with ... as, except ... as, unpacking, or comprehension target [reads: code]
  3. If no binding exists, search module level and the import block; if the only definition found is inside a different function of the same module, the read is unbound [reads: code]
Counter-example
An identifier that looks foreign but is bound earlier in the same function (e.g. produced by tuple unpacking of a helper call, or a for loop target several lines above), or one defined at module scope / imported at the top — those resolve fine.
Discriminator
The failing case has zero binding sites for the name anywhere in the enclosing function or at module/import scope; the safe case has at least one binding that dominates the use.
Consequence
NameError: name '<x>' is not defined at the moment the statement executes; in loader/parser code this is typically caught and re-raised as the library's own format/parse error wrapper, so the operation fails for every input that reaches that branch. Where the surrounding code is also mis-indented, the parse error fires first and hides this one.
Evidence
An added line assigned from pdu_length and attributed autosar_specifics.e2e, neither of which is a parameter or local of the enclosing loader method; the same class of unbound-name read was the originally reported NameError: name '<elem>' is not defined wrapped in the library's UnsupportedDatabaseFormatError.
id a648b38a7929 · mined from swesmith/cantools__cantools.0c6a7871 cantools__cantools.0c6a7871.func_pm_remove_assign__ul0v6zcg
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List every bare identifier read (not `self.x`, not a literal, not a call to an imported name) by the newly added/modified statements in a function [reads: code]",
 "prediction": "`NameError: name '<x>' is not defined` at the moment the statement executes; in loader/parser code this is typically caught and re-raised as the library's own format/parse error wrapper, so the operation fails for every input that reaches that branch. Where the surrounding code is also mis-indented, the parse error fires first and hides this one."
}
raw text (what the judge reads)
### Statement reads identifiers that are never bound in its scope
- **Applies when**: `code`: the program adds or modifies statements inside a method/function that assign from or pass along local variable names
- **Pattern**: An added statement uses a bare identifier (as a value or as the object being attributed) that is never a parameter, never assigned earlier in that function, never imported, and never a module-level global — it only exists as a local in some *other* function of the same file, so the line raises `NameError` the first time it is reached.
- **Detection procedure**:
  1. List every bare identifier read (not `self.x`, not a literal, not a call to an imported name) by the newly added/modified statements in a function [reads: code]
  2. For each such identifier, search the enclosing function for a binding: parameter list, `for` target, assignment, `with ... as`, `except ... as`, unpacking, or comprehension target [reads: code]
  3. If no binding exists, search module level and the import block; if the only definition found is inside a *different* function of the same module, the read is unbound [reads: code]
- **Counter-example**: An identifier that looks foreign but is bound earlier in the same function (e.g. produced by tuple unpacking of a helper call, or a `for` loop target several lines above), or one defined at module scope / imported at the top — those resolve fine.
- **Discriminator**: The failing case has *zero* binding sites for the name anywhere in the enclosing function or at module/import scope; the safe case has at least one binding that dominates the use.
- **Consequence**: `NameError: name '<x>' is not defined` at the moment the statement executes; in loader/parser code this is typically caught and re-raised as the library's own format/parse error wrapper, so the operation fails for every input that reaches that branch. Where the surrounding code is also mis-indented, the parse error fires first and hides this one.
- **Evidence**: An added line assigned from `pdu_length` and attributed `autosar_specifics.e2e`, neither of which is a parameter or local of the enclosing loader method; the same class of unbound-name read was the originally reported `NameError: name '<elem>' is not defined` wrapped in the library's `UnsupportedDatabaseFormatError`.
13Object constructed, populated, then discardedcodeswesmith/cantools__cantools.0c6a7871
Applies when
code: a function instantiates a container/properties/result object and sets attributes or keys on it
Pattern
A function builds an object, fills in its fields, and then falls off the end without returning it, assigning it into a longer-lived structure, or passing it anywhere. Values that were expensive to parse are computed and thrown away, so the feature silently never takes effect — no exception is raised, the downstream attribute simply stays at its default.
Detection procedure
  1. Locate local variables that are assigned a freshly constructed object (X = SomeClass(...), X = {}, X = []) inside a function. [reads: code]
  2. Scan the remainder of the function for any use of X other than X.attr = ... / X[k] = ...: a return X, something.field = X, list.append(X), or f(X). [reads: code]
  3. If no such use exists on any path, and other locals in the same function (e.g. a length or id parsed from the input) are also read nowhere after their assignment, the function's work is dead. [reads: code]
  4. Cross-check the task statement: if the task requires this data to be visible on the produced object/database, the omission is a functional gap, not stylistic. [reads: task]
Counter-example
A function that populates a local object and then stores it via a setter call, appends it to a collection, or returns it at the end of a branch — the object escapes, so it must not fire. Likewise a builder whose only purpose is validation and that documents discarding.
Consequence
No exception; instead the attribute the caller inspects remains None/default, and assertions on that property fail (AssertionError in tests, or AttributeError/TypeError further downstream when consumers assume it was populated). Feature-completeness requirement in the task is not met even though the module imports and runs.
Evidence
A method built e2e_props = AutosarEnd2EndProperties(), set .category and .data_ids, and then ended — the lines e2e_props.payload_length = pdu_length and autosar_specifics.e2e = e2e_props were missing, and a sibling local holding a parsed length was never converted or read, leaving the end-to-end properties permanently unset on loaded objects.
id 4fd952afb294 · mined from swesmith/cantools__cantools.0c6a7871 cantools__cantools.0c6a7871.func_pm_remove_assign__ul0v6zcg
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate local variables that are assigned a freshly constructed object (`X = SomeClass(...)`, `X = {}`, `X = []`) inside a function. [reads: code]",
 "prediction": "No exception; instead the attribute the caller inspects remains `None`/default, and assertions on that property fail (`AssertionError` in tests, or `AttributeError`/`TypeError` further downstream when consumers assume it was populated). Feature-completeness requirement in the task is not met even though the module imports and runs."
}
raw text (what the judge reads)
### Object constructed, populated, then discarded

- **Applies when**: `code`: a function instantiates a container/properties/result object and sets attributes or keys on it
- **Pattern**: A function builds an object, fills in its fields, and then falls off the end without returning it, assigning it into a longer-lived structure, or passing it anywhere. Values that were expensive to parse are computed and thrown away, so the feature silently never takes effect — no exception is raised, the downstream attribute simply stays at its default.
- **Detection procedure**:
  1. Locate local variables that are assigned a freshly constructed object (`X = SomeClass(...)`, `X = {}`, `X = []`) inside a function. [reads: code]
  2. Scan the remainder of the function for any use of `X` other than `X.attr = ...` / `X[k] = ...`: a `return X`, `something.field = X`, `list.append(X)`, or `f(X)`. [reads: code]
  3. If no such use exists on any path, and other locals in the same function (e.g. a length or id parsed from the input) are also read nowhere after their assignment, the function's work is dead. [reads: code]
  4. Cross-check the task statement: if the task requires this data to be visible on the produced object/database, the omission is a functional gap, not stylistic. [reads: task]
- **Counter-example**: A function that populates a local object and then stores it via a setter call, appends it to a collection, or returns it at the end of a branch — the object escapes, so it must not fire. Likewise a builder whose only purpose is validation and that documents discarding.
- **Consequence**: No exception; instead the attribute the caller inspects remains `None`/default, and assertions on that property fail (`AssertionError` in tests, or `AttributeError`/`TypeError` further downstream when consumers assume it was populated). Feature-completeness requirement in the task is not met even though the module imports and runs.
- **Evidence**: A method built `e2e_props = AutosarEnd2EndProperties()`, set `.category` and `.data_ids`, and then ended — the lines `e2e_props.payload_length = pdu_length` and `autosar_specifics.e2e = e2e_props` were missing, and a sibling local holding a parsed length was never converted or read, leaving the end-to-end properties permanently unset on loaded objects.
14Reported crash "fixed" by changing documented value-propagation semantics insteadtaskswesmith/alecthomas__voluptuous.a7a55f83
Applies when
task: the report quotes an exception class/traceback (e.g. a TypeError about a missing or extra positional argument) as the symptom; code: the submission edits the body of an existing library function
Pattern
rather than locating the code path that raises the quoted exception, the submission changes what an already-working path passes along or returns (e.g. stops feeding one step's output into the next) and rewrites the surrounding docstring so the documentation matches the new behavior — silently discarding a contract other callers and existing tests depend on, while the reported crash is untouched.
Detection procedure
  1. Read the task's quoted error: note that it is a crash about how a callable is invoked (arity/attribute/type), not a complaint about a wrong returned value [reads: task]
  2. Identify the functional edits: if the submission leaves a sibling backup copy of the edited source (.py.bak, .orig, *.old) in the repo, diff it against the live module to obtain the exact change set [reads: code]
  3. The defect is present when no edit in that change set touches how the failing callable is invoked (its signature, its argument dispatch, the branch that passes an extra/missing argument), and the only functional edits change value propagation (assignment of a call's result being dropped or added) together with a docstring sentence rewritten to describe the new semantics [reads: code]
Counter-example
a change set that alters the argument dispatch which raises the quoted exception — adding or removing the extra path/context argument, or branching on whether a compiled callable takes one or two arguments — while leaving value propagation and the documented contract unchanged.
Discriminator
zero edits anywhere on the invocation path named in the quoted traceback, plus a docstring line edited to legitimize a behavior change on inputs that already succeeded.
Consequence
the reported exception remains reproducible, so the task's own test fails; additionally, tests and doctests asserting the previous documented behavior (that each step receives the previous step's output) now fail with AssertionError or unexpected result types — a regression on top of an unfixed bug. This mechanism accounts for the fix attempt producing no change on the reported symptom; the remainder of the poor outcome comes from repository pollution by scratch files.
Evidence
v = func(v) was changed to func(v) inside a multi-sub-validator combinator and the docstring line "The output of each validator is passed as input to the next" was replaced with "Each validator is applied independently"; the submission's own reproduction runs never produced the TypeError quoted in the report, showing the edited line was not the reported defect.
id b748c80c037e · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Read the task's quoted error: note that it is a crash about how a callable is invoked (arity/attribute/type), not a complaint about a wrong returned value [reads: task]",
 "prediction": "the reported exception remains reproducible, so the task's own test fails; additionally, tests and doctests asserting the previous documented behavior (that each step receives the previous step's output) now fail with `AssertionError` or unexpected result types \u2014 a regression on top of an unfixed bug. This mechanism accounts for the fix attempt producing no change on the reported symptom; the remainder of the poor outcome comes from repository pollution by scratch files."
}
raw text (what the judge reads)
### Reported crash "fixed" by changing documented value-propagation semantics instead
- **Applies when**: `task`: the report quotes an exception class/traceback (e.g. a `TypeError` about a missing or extra positional argument) as the symptom; `code`: the submission edits the body of an existing library function
- **Pattern**: rather than locating the code path that raises the quoted exception, the submission changes what an already-working path passes along or returns (e.g. stops feeding one step's output into the next) and rewrites the surrounding docstring so the documentation matches the new behavior — silently discarding a contract other callers and existing tests depend on, while the reported crash is untouched.
- **Detection procedure**:
  1. Read the task's quoted error: note that it is a crash about how a callable is invoked (arity/attribute/type), not a complaint about a wrong returned value [reads: task]
  2. Identify the functional edits: if the submission leaves a sibling backup copy of the edited source (`*.py.bak`, `*.orig`, `*.old`) in the repo, diff it against the live module to obtain the exact change set [reads: code]
  3. The defect is present when no edit in that change set touches how the failing callable is invoked (its signature, its argument dispatch, the branch that passes an extra/missing argument), and the only functional edits change value propagation (assignment of a call's result being dropped or added) together with a docstring sentence rewritten to describe the new semantics [reads: code]
- **Counter-example**: a change set that alters the argument dispatch which raises the quoted exception — adding or removing the extra path/context argument, or branching on whether a compiled callable takes one or two arguments — while leaving value propagation and the documented contract unchanged.
- **Discriminator**: zero edits anywhere on the invocation path named in the quoted traceback, plus a docstring line edited to legitimize a behavior change on inputs that already succeeded.
- **Consequence**: the reported exception remains reproducible, so the task's own test fails; additionally, tests and doctests asserting the previous documented behavior (that each step receives the previous step's output) now fail with `AssertionError` or unexpected result types — a regression on top of an unfixed bug. This mechanism accounts for the fix attempt producing no change on the reported symptom; the remainder of the poor outcome comes from repository pollution by scratch files.
- **Evidence**: `v = func(v)` was changed to `func(v)` inside a multi-sub-validator combinator and the docstring line "The output of each validator is passed as input to the next" was replaced with "Each validator is applied independently"; the submission's own reproduction runs never produced the `TypeError` quoted in the report, showing the edited line was not the reported defect.
14Regex character class where an escaped backslash becomes a range endpointcodeswesmith/alecthomas__voluptuous.a7a55f83
Applies when
code: the program builds or hard-codes regular-expression patterns (passed to re.compile, re.match, or a regex-based validator/matcher) that contain a bracketed character class [...]
Pattern
A pattern literal is written with the backslash-doubling appropriate for an ordinary string but is actually declared as a raw string (or is double-escaped twice). Inside a character class the doubled backslash \\ stops being an escape and becomes a literal backslash character; an immediately following - then turns it into a character range whose left endpoint (\, U+005C) is greater than the right endpoint, so the pattern cannot compile.
Detection procedure
  1. Find every string literal that is compiled as a regex (argument of re.compile/re.match/re.search, or of a matcher class constructed from a pattern) and keep the ones containing [ ... ] [reads: code]
  2. For each, read the literal's prefix and quoting to determine the actual character sequence handed to re: an r'...'/r"..." literal passes backslashes through verbatim; a plain literal halves each \\ [reads: code]
  3. Check whether the sequence reaching re contains, inside the brackets, a backslash-escaped backslash immediately followed by - and another character (raw literal spelled [...\\-X...], or plain literal spelled [...\\\\-X...]); that is the failing case [reads: code]
Counter-example
a plain (non-raw) literal such as "[()\\-\\'._+=]", which reaches re as [()\-\'._+=] — an escaped hyphen, perfectly legal — or a raw literal r"[a-z0-9\-]" where the hyphen is single-escaped or last in the class.
Discriminator
what matters is the sequence re actually receives: a literal backslash directly before the hyphen (range endpoint) fails; a single escape \- before the hyphen (escaped literal hyphen) is safe. Raw prefix + doubled backslash is the failing combination; plain literal + doubled backslash is the safe one.
Consequence
re.error ("bad character range" / "bad escape") raised at pattern-compile time — i.e. at import or at validator/matcher construction, before any data is processed — aborting the script or test module with a non-zero exit; no partial results are produced.
Evidence
a character class whose compiled form contained \\-\' raised re.error: bad character range \\-\' at position 24 from re.compile, terminating the run after only the first few checks had printed.
id 0ea24d85cb63 · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Find every string literal that is compiled as a regex (argument of `re.compile`/`re.match`/`re.search`, or of a matcher class constructed from a pattern) and keep the ones containing `[` ... `]` [reads: code]",
 "prediction": "`re.error` (\"bad character range\" / \"bad escape\") raised at pattern-compile time \u2014 i.e. at import or at validator/matcher construction, before any data is processed \u2014 aborting the script or test module with a non-zero exit; no partial results are produced."
}
raw text (what the judge reads)
### Regex character class where an escaped backslash becomes a range endpoint
- **Applies when**: `code`: the program builds or hard-codes regular-expression patterns (passed to `re.compile`, `re.match`, or a regex-based validator/matcher) that contain a bracketed character class `[...]`
- **Pattern**: A pattern literal is written with the backslash-doubling appropriate for an *ordinary* string but is actually declared as a *raw* string (or is double-escaped twice). Inside a character class the doubled backslash `\\` stops being an escape and becomes a literal backslash character; an immediately following `-` then turns it into a character *range* whose left endpoint (`\`, U+005C) is greater than the right endpoint, so the pattern cannot compile.
- **Detection procedure**:
  1. Find every string literal that is compiled as a regex (argument of `re.compile`/`re.match`/`re.search`, or of a matcher class constructed from a pattern) and keep the ones containing `[` ... `]` [reads: code]
  2. For each, read the literal's prefix and quoting to determine the actual character sequence handed to `re`: an `r'...'`/`r"..."` literal passes backslashes through verbatim; a plain literal halves each `\\` [reads: code]
  3. Check whether the sequence reaching `re` contains, inside the brackets, a backslash-escaped backslash immediately followed by `-` and another character (raw literal spelled `[...\\-X...]`, or plain literal spelled `[...\\\\-X...]`); that is the failing case [reads: code]
- **Counter-example**: a plain (non-raw) literal such as `"[()\\-\\'._+=]"`, which reaches `re` as `[()\-\'._+=]` — an escaped hyphen, perfectly legal — or a raw literal `r"[a-z0-9\-]"` where the hyphen is single-escaped or last in the class.
- **Discriminator**: what matters is the sequence `re` actually receives: a *literal backslash* directly before the hyphen (range endpoint) fails; a *single escape* `\-` before the hyphen (escaped literal hyphen) is safe. Raw prefix + doubled backslash is the failing combination; plain literal + doubled backslash is the safe one.
- **Consequence**: `re.error` ("bad character range" / "bad escape") raised at pattern-compile time — i.e. at import or at validator/matcher construction, before any data is processed — aborting the script or test module with a non-zero exit; no partial results are produced.
- **Evidence**: a character class whose compiled form contained `\\-\'` raised `re.error: bad character range \\-\' at position 24` from `re.compile`, terminating the run after only the first few checks had printed.
14Sibling override drops the optional-argument branch its peers all implementcodeswesmith/alecthomas__voluptuous.a7a55f83
Applies when
code: a base class (or protocol) declares a method that subclasses override, and the method has an optional/defaulted parameter that changes how it invokes the callables it receives (e.g. an optional path/context/index forwarded to sub-validators or sub-handlers).
Pattern
One override of the shared method invokes its sub-callables with a fixed argument arity, ignoring the case where the optional parameter is absent (None), while every sibling override branches on it. The object then works when reached through the code path that supplies the parameter and raises a TypeError about positional arguments when invoked through the entry point that does not.
Detection procedure
  1. In the module under repair, find the base class that defines the abstract/NotImplementedError method and note every subclass that overrides it, plus the two entry points into it (e.g. a public __call__ that omits the optional parameter and an internal runner that passes it). [reads: code]
  2. From the task statement, note the reported exception text — messages of the form "missing 1 required positional argument" or "takes N positional arguments but N+1 were given" naming a nested callable — and which public entry point triggers it. [reads: task]
  3. Compare the overrides line by line: fire if the override belonging to the class named in the task calls func(path, v) (or func(v)) unconditionally, while the sibling overrides contain an explicit if path is None: ... else: ... guard around the same call. [reads: code]
Counter-example
An override that also lacks the if path is None guard but is only ever reachable through the compiled/runner path because its class has no public single-argument __call__ inherited from the base, or an override that normalises the parameter first (path = path or []) before dispatching.
Discriminator
The failing case is reachable from an entry point that leaves the optional parameter at its default and dispatches with the arity appropriate to the non-default case; the safe case either normalises the default or is unreachable from that entry point.
Consequence
TypeError at validation/dispatch time (message: missing required positional argument, or too many positional arguments), raised only on the direct-call path; the class's documented behaviour is unusable and any hidden test exercising it fails.
Evidence
A subclass _exec dispatched sub-schemas without the if path is None: branch present in its Any/All siblings; direct invocation of the object produced TypeError: ... missing 1 required positional argument: 'data' and TypeError: Schema.__call__() takes 2 positional arguments but 3 were given, and the submitted change never added the branch.
id 3f36aeab4034 · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. In the module under repair, find the base class that defines the abstract/`NotImplementedError` method and note every subclass that overrides it, plus the two entry points into it (e.g. a public `__call__` that omits the optional parameter and an internal runner that passes it). [reads: code]",
 "prediction": "`TypeError` at validation/dispatch time (message: missing required positional argument, or too many positional arguments), raised only on the direct-call path; the class's documented behaviour is unusable and any hidden test exercising it fails."
}
raw text (what the judge reads)
### Sibling override drops the optional-argument branch its peers all implement
- **Applies when**: `code`: a base class (or protocol) declares a method that subclasses override, and the method has an optional/defaulted parameter that changes how it invokes the callables it receives (e.g. an optional `path`/`context`/`index` forwarded to sub-validators or sub-handlers).
- **Pattern**: One override of the shared method invokes its sub-callables with a fixed argument arity, ignoring the case where the optional parameter is absent (`None`), while every sibling override branches on it. The object then works when reached through the code path that supplies the parameter and raises a `TypeError` about positional arguments when invoked through the entry point that does not.
- **Detection procedure**:
  1. In the module under repair, find the base class that defines the abstract/`NotImplementedError` method and note every subclass that overrides it, plus the two entry points into it (e.g. a public `__call__` that omits the optional parameter and an internal runner that passes it). [reads: code]
  2. From the task statement, note the reported exception text — messages of the form "missing 1 required positional argument" or "takes N positional arguments but N+1 were given" naming a nested callable — and which public entry point triggers it. [reads: task]
  3. Compare the overrides line by line: fire if the override belonging to the class named in the task calls `func(path, v)` (or `func(v)`) unconditionally, while the sibling overrides contain an explicit `if path is None: ... else: ...` guard around the same call. [reads: code]
- **Counter-example**: An override that also lacks the `if path is None` guard but is only ever reachable through the compiled/runner path because its class has no public single-argument `__call__` inherited from the base, or an override that normalises the parameter first (`path = path or []`) before dispatching.
- **Discriminator**: The failing case is reachable from an entry point that leaves the optional parameter at its default *and* dispatches with the arity appropriate to the non-default case; the safe case either normalises the default or is unreachable from that entry point.
- **Consequence**: `TypeError` at validation/dispatch time (message: missing required positional argument, or too many positional arguments), raised only on the direct-call path; the class's documented behaviour is unusable and any hidden test exercising it fails.
- **Evidence**: A subclass `_exec` dispatched sub-schemas without the `if path is None:` branch present in its `Any`/`All` siblings; direct invocation of the object produced `TypeError: ... missing 1 required positional argument: 'data'` and `TypeError: Schema.__call__() takes 2 positional arguments but 3 were given`, and the submitted change never added the branch.
14Unguarded duplicate of a call that is expected to raisecodeswesmith/alecthomas__voluptuous.a7a55f83
Applies when
code: the change adds or edits a verification/demo script that exercises error paths with try: / except <ExceptionType>: blocks
Pattern
A statement that invokes the operation under test is placed immediately before the try block that is meant to catch its failure (a leftover/duplicated call), so the anticipated exception propagates out of the script instead of being caught, aborting every later check.
Detection procedure
  1. In each added script, locate every try: block whose except clause names an exception the surrounding comment/print describes as the expected outcome ("should fail", "should raise", "correctly fails"). [reads: code]
  2. Read the statement(s) directly above that try: in the same scope and compare the called expression with the first statement inside the try: body. [reads: code]
  3. Fire if the same call with the same arguments appears both outside and inside the try:, and the outer occurrence is not itself wrapped in any handler. [reads: code]
Counter-example
A setup call placed before the try: that uses different arguments/input chosen to succeed (e.g. a valid input used to build state, then an invalid input inside the try:), or the outer call wrapped in its own try/except.
Discriminator
The pre-try call passes the exact input the script itself labels as the failing case; safe code only pre-calls with inputs it expects to return normally.
Consequence
The script terminates with an uncaught exception of the class named in the following except (domain exception subclass, or AssertionError), exit code non-zero; all subsequent checks in the file and any final "all tests passed" output never run, so the script reports neither success nor the intended verdict.
Evidence
result = validator('Aa1') # 3 matches, but we want exactly 2 placed one line above try: result = validator('Aa1') ... except Invalid: produced an uncaught TooManyValid traceback and killed the remaining checks in the file.
id 85e4d52cd7da · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. In each added script, locate every `try:` block whose `except` clause names an exception the surrounding comment/print describes as the *expected* outcome (\"should fail\", \"should raise\", \"correctly fails\"). [reads: code]",
 "prediction": "The script terminates with an uncaught exception of the class named in the following `except` (domain exception subclass, or `AssertionError`), exit code non-zero; all subsequent checks in the file and any final \"all tests passed\" output never run, so the script reports neither success nor the intended verdict."
}
raw text (what the judge reads)
### Unguarded duplicate of a call that is expected to raise
- **Applies when**: `code`: the change adds or edits a verification/demo script that exercises error paths with `try:` / `except <ExceptionType>:` blocks
- **Pattern**: A statement that invokes the operation under test is placed immediately *before* the `try` block that is meant to catch its failure (a leftover/duplicated call), so the anticipated exception propagates out of the script instead of being caught, aborting every later check.
- **Detection procedure**:
  1. In each added script, locate every `try:` block whose `except` clause names an exception the surrounding comment/print describes as the *expected* outcome ("should fail", "should raise", "correctly fails"). [reads: code]
  2. Read the statement(s) directly above that `try:` in the same scope and compare the called expression with the first statement inside the `try:` body. [reads: code]
  3. Fire if the same call with the same arguments appears both outside and inside the `try:`, and the outer occurrence is not itself wrapped in any handler. [reads: code]
- **Counter-example**: A setup call placed before the `try:` that uses *different* arguments/input chosen to succeed (e.g. a valid input used to build state, then an invalid input inside the `try:`), or the outer call wrapped in its own `try/except`.
- **Discriminator**: The pre-`try` call passes the exact input the script itself labels as the failing case; safe code only pre-calls with inputs it expects to return normally.
- **Consequence**: The script terminates with an uncaught exception of the class named in the following `except` (domain exception subclass, or `AssertionError`), exit code non-zero; all subsequent checks in the file and any final "all tests passed" output never run, so the script reports neither success nor the intended verdict.
- **Evidence**: `result = validator('Aa1')  # 3 matches, but we want exactly 2` placed one line above `try: result = validator('Aa1') ... except Invalid:` produced an uncaught `TooManyValid` traceback and killed the remaining checks in the file.
14Independent-check combinator threads each check's return value into the nextcodeswesmith/alecthomas__voluptuous.a7a55f83
Applies when
code: a function/method iterates over a collection of callables (validators, checks, predicates, transforms) applied to a single input value, and the surrounding contract is "each callable is applied to the same input" — e.g. counting how many succeed, requiring at least/at most N to pass, collecting all errors, or scoring alternatives
Pattern
The loop rebinds the input variable to each callable's return value (v = func(v)), turning an independent, parallel evaluation into a sequential pipeline. Later callables then receive a value the earlier one produced rather than the original input, so success/failure counts are computed over the wrong inputs and callables that expect the original type can blow up.
Detection procedure
  1. Locate loops of the form for f in funcs: ... f(value) ... inside a function whose body also accumulates outcomes — an errors list, a passed/count counter, or a try/except that appends instead of re-raising. [reads: code]
  2. Read what the enclosing function/class documents or is required to do (docstring, parameter names such as min_valid/max_valid/at_least/n_required, or the task statement describing "value must pass at least N of these"). Confirm the contract is "apply every check to the same value", not "feed the output of one into the next". [reads: task and code (docstring/parameter names)]
  3. Check the call site inside the loop: does it assign back to the same variable that is passed in (v = f(v) / value = f(path, value)), while the function's final return is the original input or a threshold comparison rather than the accumulated value? If yes, the rebinding is unintended. [reads: code]
Counter-example
A pipeline/compose combinator (All/And/Compose/Pipeline) whose documented semantics are "the output of each validator is passed as input to the next" and whose _exec returns the final chained v after the loop, with no per-callable error counting — here v = func(v) is correct and must not be flagged.
Discriminator
The offending case counts or collects per-callable outcomes and its return/decision does not depend on the chained value (it returns the original input or compares a pass count against bounds), yet it still overwrites the input each iteration. The safe case has no counting and the chained value is the result.
Consequence
Wrong acceptance/rejection: inputs that satisfy the required number of checks are rejected and vice versa, so unit tests asserting NotEnoughValid/TooManyValid-style outcomes (or equivalent domain assertions) fail. When an earlier callable returns a different type or a callable object, later calls raise TypeError (e.g. "takes 2 positional arguments but 3 were given", "missing 1 required positional argument") or AttributeError escaping the except clause that only catches the domain error class.
Evidence
A counting combinator's _exec used v = func(v) / v = func(path, v) inside its loop while returning the original value after comparing passed_count to min/max bounds; the reported symptoms were TypeError from later sub-validators and incorrect pass/fail results, and the accepted fix was to drop the reassignment (func(v) / func(path, v)) so every check sees the same input.
id fe1ed74e170c · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate loops of the form `for f in funcs: ... f(value) ...` inside a function whose body also accumulates outcomes \u2014 an `errors` list, a `passed`/`count` counter, or a try/except that appends instead of re-raising. [reads: code]",
 "prediction": "Wrong acceptance/rejection: inputs that satisfy the required number of checks are rejected and vice versa, so unit tests asserting `NotEnoughValid`/`TooManyValid`-style outcomes (or equivalent domain assertions) fail. When an earlier callable returns a different type or a callable object, later calls raise `TypeError` (e.g. \"takes 2 positional arguments but 3 were given\", \"missing 1 required positional argument\") or `AttributeError` escaping the `except` clause that only catches the domain error class."
}
raw text (what the judge reads)
### Independent-check combinator threads each check's return value into the next
- **Applies when**: `code`: a function/method iterates over a collection of callables (validators, checks, predicates, transforms) applied to a single input value, and the surrounding contract is "each callable is applied to the same input" — e.g. counting how many succeed, requiring at least/at most N to pass, collecting all errors, or scoring alternatives
- **Pattern**: The loop rebinds the input variable to each callable's return value (`v = func(v)`), turning an independent, parallel evaluation into a sequential pipeline. Later callables then receive a value the earlier one produced rather than the original input, so success/failure counts are computed over the wrong inputs and callables that expect the original type can blow up.
- **Detection procedure**:
  1. Locate loops of the form `for f in funcs: ... f(value) ...` inside a function whose body also accumulates outcomes — an `errors` list, a `passed`/`count` counter, or a try/except that appends instead of re-raising. [reads: code]
  2. Read what the enclosing function/class documents or is required to do (docstring, parameter names such as `min_valid`/`max_valid`/`at_least`/`n_required`, or the task statement describing "value must pass at least N of these"). Confirm the contract is "apply every check to the same value", not "feed the output of one into the next". [reads: task and code (docstring/parameter names)]
  3. Check the call site inside the loop: does it assign back to the same variable that is passed in (`v = f(v)` / `value = f(path, value)`), while the function's final `return` is the original input or a threshold comparison rather than the accumulated value? If yes, the rebinding is unintended. [reads: code]
- **Counter-example**: A pipeline/compose combinator (`All`/`And`/`Compose`/`Pipeline`) whose documented semantics are "the output of each validator is passed as input to the next" and whose `_exec` returns the final chained `v` after the loop, with no per-callable error counting — here `v = func(v)` is correct and must not be flagged.
- **Discriminator**: The offending case counts or collects per-callable outcomes and its return/decision does not depend on the chained value (it returns the original input or compares a pass count against bounds), yet it still overwrites the input each iteration. The safe case has no counting and the chained value *is* the result.
- **Consequence**: Wrong acceptance/rejection: inputs that satisfy the required number of checks are rejected and vice versa, so unit tests asserting `NotEnoughValid`/`TooManyValid`-style outcomes (or equivalent domain assertions) fail. When an earlier callable returns a different type or a callable object, later calls raise `TypeError` (e.g. "takes 2 positional arguments but 3 were given", "missing 1 required positional argument") or `AttributeError` escaping the `except` clause that only catches the domain error class.
- **Evidence**: A counting combinator's `_exec` used `v = func(v)` / `v = func(path, v)` inside its loop while returning the original value after comparing `passed_count` to min/max bounds; the reported symptoms were `TypeError` from later sub-validators and incorrect pass/fail results, and the accepted fix was to drop the reassignment (`func(v)` / `func(path, v)`) so every check sees the same input.
14Fix leaves the reported exception's call site untouchedtaskswesmith/alecthomas__voluptuous.a7a55f83
Applies when
task: the issue report includes a traceback or exception text (e.g. TypeError: ... missing 1 required positional argument or ... takes 2 positional arguments but 3 were given); code: the candidate is a patch to library code.
Pattern
The patch edits lines near the failure (renaming, dropping an assignment, rewording a docstring) but never changes the construct whose arity/branching produces the quoted exception, so the reported error is still raised after the fix.
Detection procedure
  1. Read the exception class and message quoted in the issue and identify the construct that can produce it — for an arity TypeError, a call site whose argument count is chosen by a conditional (e.g. if path is None: func(v) else: func(path, v)), or a dispatch that forwards a variable number of positionals. [reads: task]
  2. Locate that construct in the candidate code and read the guard controlling it. [reads: code]
  3. Check whether the arm that tests the sentinel as absent still forwards the sentinel variable, or the arm that tests it as present omits it (i.e. the two arms are swapped), and whether the candidate's changed lines touch that guard at all. [reads: code]
Counter-example
A patch that leaves the guard alone because the guard already pairs correctly (if path is None: f(v) / else: f(path, v), matching sibling classes in the same module) and instead fixes a different line the traceback points to.
Discriminator
The wrong case has an argument-forwarding conditional whose arms are inverted relative to the sentinel test, and the candidate's diff does not modify that conditional; the safe case either has correctly paired arms or the diff repairs them.
Consequence
The originally reported TypeError (missing/extra positional argument) is raised again by any test exercising the compiled/nested path; every reproduction case in the issue still fails. Explains the bulk of the gap here (the accepted fix repaired this branch plus the count/bounds logic); the remainder is attributable to the other defects left in the same function.
id fcfa57447037 · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Read the exception class and message quoted in the issue and identify the construct that can produce it \u2014 for an arity `TypeError`, a call site whose argument count is chosen by a conditional (e.g. `if path is None: func(v) else: func(path, v)`), or a dispatch that forwards a variable number of positionals. [reads: task]",
 "prediction": "The originally reported `TypeError` (missing/extra positional argument) is raised again by any test exercising the compiled/nested path; every reproduction case in the issue still fails. Explains the bulk of the gap here (the accepted fix repaired this branch plus the count/bounds logic); the remainder is attributable to the other defects left in the same function."
}
raw text (what the judge reads)
### Fix leaves the reported exception's call site untouched
- **Applies when**: `task`: the issue report includes a traceback or exception text (e.g. `TypeError: ... missing 1 required positional argument` or `... takes 2 positional arguments but 3 were given`); `code`: the candidate is a patch to library code.
- **Pattern**: The patch edits lines near the failure (renaming, dropping an assignment, rewording a docstring) but never changes the construct whose arity/branching produces the quoted exception, so the reported error is still raised after the fix.
- **Detection procedure**:
  1. Read the exception class and message quoted in the issue and identify the construct that can produce it — for an arity `TypeError`, a call site whose argument count is chosen by a conditional (e.g. `if path is None: func(v) else: func(path, v)`), or a dispatch that forwards a variable number of positionals. [reads: task]
  2. Locate that construct in the candidate code and read the guard controlling it. [reads: code]
  3. Check whether the arm that tests the sentinel as *absent* still forwards the sentinel variable, or the arm that tests it as *present* omits it (i.e. the two arms are swapped), and whether the candidate's changed lines touch that guard at all. [reads: code]
- **Counter-example**: A patch that leaves the guard alone because the guard already pairs correctly (`if path is None: f(v)` / `else: f(path, v)`, matching sibling classes in the same module) and instead fixes a different line the traceback points to.
- **Discriminator**: The wrong case has an argument-forwarding conditional whose arms are inverted relative to the sentinel test, and the candidate's diff does not modify that conditional; the safe case either has correctly paired arms or the diff repairs them.
- **Consequence**: The originally reported `TypeError` (missing/extra positional argument) is raised again by any test exercising the compiled/nested path; every reproduction case in the issue still fails. Explains the bulk of the gap here (the accepted fix repaired this branch plus the count/bounds logic); the remainder is attributable to the other defects left in the same function.
14Crossed keyword-to-attribute assignment left in placecodeswesmith/alecthomas__voluptuous.a7a55f83
Applies when
code: a class or function stores paired/symmetric parameters (min/max, lower/upper, start/end, first/last, width/height) into attributes of correspondingly paired names.
Pattern
The constructor assigns each parameter to the other member of the pair (self.min_x = max_x or 0, self.max_x = min_x or default), so every downstream bound check is inverted; a patch that only edits the consuming logic leaves this crossing intact.
Detection procedure
  1. Locate the __init__ / setup block of the class named in the issue and list each self.<name> = <expr> assignment. [reads: code]
  2. Compare each left-hand attribute name with the parameter name appearing in its right-hand expression, and with the parameter names the task/issue says the user passes. [reads: task]
  3. Flag when a left-hand name containing one member of a symmetric pair is assigned from the parameter containing the other member (and the default fallback also belongs to the swapped side, e.g. self.min_ = max_ or 0). [reads: code]
Counter-example
self.min_valid = min_valid or 0 / self.max_valid = max_valid or len(items) — names match on both sides, even though the defaults differ; also safe is a deliberate normalization such as self.lo, self.hi = sorted((a, b)).
Discriminator
The failing case has name mismatch between the assigned attribute and the source parameter across a symmetric pair with no sorting/normalization comment; the safe case has matching names or an explicit reordering construct.
Consequence
Bounds behave inversely — inputs at the intended minimum are rejected and inputs above the intended maximum accepted; unit tests for both the lower- and upper-bound parameters fail with the wrong Invalid/error subclass or no error at all. Accounts for the min/max half of the observed behavioral gap.
id 0bcedd34532e · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Locate the `__init__` / setup block of the class named in the issue and list each `self.<name> = <expr>` assignment. [reads: code]",
 "prediction": "Bounds behave inversely \u2014 inputs at the intended minimum are rejected and inputs above the intended maximum accepted; unit tests for both the lower- and upper-bound parameters fail with the wrong `Invalid`/error subclass or no error at all. Accounts for the min/max half of the observed behavioral gap."
}
raw text (what the judge reads)
### Crossed keyword-to-attribute assignment left in place
- **Applies when**: `code`: a class or function stores paired/symmetric parameters (min/max, lower/upper, start/end, first/last, width/height) into attributes of correspondingly paired names.
- **Pattern**: The constructor assigns each parameter to the *other* member of the pair (`self.min_x = max_x or 0`, `self.max_x = min_x or default`), so every downstream bound check is inverted; a patch that only edits the consuming logic leaves this crossing intact.
- **Detection procedure**:
  1. Locate the `__init__` / setup block of the class named in the issue and list each `self.<name> = <expr>` assignment. [reads: code]
  2. Compare each left-hand attribute name with the parameter name appearing in its right-hand expression, and with the parameter names the task/issue says the user passes. [reads: task]
  3. Flag when a left-hand name containing one member of a symmetric pair is assigned from the parameter containing the other member (and the default fallback also belongs to the swapped side, e.g. `self.min_* = max_* or 0`). [reads: code]
- **Counter-example**: `self.min_valid = min_valid or 0` / `self.max_valid = max_valid or len(items)` — names match on both sides, even though the defaults differ; also safe is a deliberate normalization such as `self.lo, self.hi = sorted((a, b))`.
- **Discriminator**: The failing case has name mismatch between the assigned attribute and the source parameter across a symmetric pair with no sorting/normalization comment; the safe case has matching names or an explicit reordering construct.
- **Consequence**: Bounds behave inversely — inputs at the intended minimum are rejected and inputs above the intended maximum accepted; unit tests for both the lower- and upper-bound parameters fail with the wrong `Invalid`/error subclass or no error at all. Accounts for the min/max half of the observed behavioral gap.
14Counting/boundary arithmetic contradicting the documented inclusive semanticscodeswesmith/alecthomas__voluptuous.a7a55f83
Applies when
code: the function computes a count of successes/failures and compares it against user-supplied limits; task: the issue or docstring gives concrete examples of which inputs must pass.
Pattern
The count is adjusted by an unexplained constant (+ 1, - 1) and/or compared with a strict inequality against a limit whose name implies inclusiveness, so boundary inputs are classified backwards; a patch that touches only the loop body leaves this arithmetic unfixed.
Detection procedure
  1. Locate the expression computing the count (e.g. passed = len(items) - len(errors) ...) and the comparison that decides success. [reads: code]
  2. Read the issue's worked examples stating which inputs "should pass" for a given limit, and the parameter names (min_/max_ imply the limit itself is allowed). [reads: task]
  3. Flag when the count expression adds/subtracts a literal that corresponds to no element of the collection, or when the comparison uses </> against a max_/min_ named bound while the examples require the bound value itself to succeed. [reads: code]
Counter-example
if lo <= count <= hi: return value with count = len(funcs) - len(errors) — no stray constant and inclusive comparison matching the parameter naming; also safe is a strict comparison where the parameter is explicitly named *_exclusive or documented as such.
Discriminator
The failing case has either an unexplained ±1 in the count or a strict comparison against an inclusively named bound, contradicting a concrete pass/fail example in the report; the safe case's arithmetic reproduces the report's examples exactly.
Consequence
Off-by-one acceptance/rejection at the boundary — inputs meeting exactly the stated limit raise the validation error (or vice versa); the boundary unit tests fail while mid-range cases pass. Explains the residual failures that remain even after the argument-forwarding branch is corrected.
id 94baa29b9689 · mined from swesmith/alecthomas__voluptuous.a7a55f83 alecthomas__voluptuous.a7a55f83.combine_file__ghxiucua
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Locate the expression computing the count (e.g. `passed = len(items) - len(errors) ...`) and the comparison that decides success. [reads: code]",
 "prediction": "Off-by-one acceptance/rejection at the boundary \u2014 inputs meeting exactly the stated limit raise the validation error (or vice versa); the boundary unit tests fail while mid-range cases pass. Explains the residual failures that remain even after the argument-forwarding branch is corrected."
}
raw text (what the judge reads)
### Counting/boundary arithmetic contradicting the documented inclusive semantics
- **Applies when**: `code`: the function computes a count of successes/failures and compares it against user-supplied limits; `task`: the issue or docstring gives concrete examples of which inputs must pass.
- **Pattern**: The count is adjusted by an unexplained constant (`+ 1`, `- 1`) and/or compared with a strict inequality against a limit whose name implies inclusiveness, so boundary inputs are classified backwards; a patch that touches only the loop body leaves this arithmetic unfixed.
- **Detection procedure**:
  1. Locate the expression computing the count (e.g. `passed = len(items) - len(errors) ...`) and the comparison that decides success. [reads: code]
  2. Read the issue's worked examples stating which inputs "should pass" for a given limit, and the parameter names (`min_*`/`max_*` imply the limit itself is allowed). [reads: task]
  3. Flag when the count expression adds/subtracts a literal that corresponds to no element of the collection, or when the comparison uses `<`/`>` against a `max_`/`min_` named bound while the examples require the bound value itself to succeed. [reads: code]
- **Counter-example**: `if lo <= count <= hi: return value` with `count = len(funcs) - len(errors)` — no stray constant and inclusive comparison matching the parameter naming; also safe is a strict comparison where the parameter is explicitly named `*_exclusive` or documented as such.
- **Discriminator**: The failing case has either an unexplained ±1 in the count or a strict comparison against an inclusively named bound, contradicting a concrete pass/fail example in the report; the safe case's arithmetic reproduces the report's examples exactly.
- **Consequence**: Off-by-one acceptance/rejection at the boundary — inputs meeting exactly the stated limit raise the validation error (or vice versa); the boundary unit tests fail while mid-range cases pass. Explains the residual failures that remain even after the argument-forwarding branch is corrected.
15Requirement named in the task has no implementing construct in the codetaskswesmith/andialbrecht__sqlparse.e57923b3
Applies when
task: the task statement names a specific behavior, class, method, or syntactic construct the program must support, and code: the full source of the modified module(s) is available
Pattern
The program returns the module to (or leaves it at) its baseline behavior — no class, branch, keyword match, or dispatch entry anywhere handles the construct the task names — so the code compiles and existing tests pass while the requested capability is simply absent.
Detection procedure
  1. Extract from the task statement the concrete nouns it requires: the class/method name to add, the keyword/token/format to recognize, or the accessor to expose. [reads: task]
  2. Grep the candidate's source for each of those names, and for the place where similar features are registered (e.g. the list of handler functions in a pipeline, the class hierarchy, the accessor methods of the relevant class). [reads: code]
  3. Fire if none of the named symbols is defined and the registration point (pipeline list, dispatch table, class body) contains no new entry corresponding to the requirement — i.e. every sibling feature has an entry and the requested one has none. [reads: code]
Counter-example
The feature is implemented under a differently spelled helper name, but a keyword/token/branch keying on the task's construct exists and is reachable from the module's public entry point.
Discriminator
In the failing case no reachable code path is conditioned on the construct the task names; in the safe case such a path exists even if the identifiers differ from the task's wording.
Consequence
Hidden or restored tests for the requested feature fail with AttributeError (missing method/class) or AssertionError (structure unchanged); the feature-specific portion of the score is lost entirely while unrelated tests still pass.
Evidence
A submission removed the grouping class, its pipeline entry, and the accessor method for the requested construct, leaving no handler for it; the suite reported all-pass because nothing left in the repo exercised it.
id 0ef0f3cfd356 · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Extract from the task statement the concrete nouns it requires: the class/method name to add, the keyword/token/format to recognize, or the accessor to expose. [reads: task]",
 "prediction": "Hidden or restored tests for the requested feature fail with `AttributeError` (missing method/class) or `AssertionError` (structure unchanged); the feature-specific portion of the score is lost entirely while unrelated tests still pass."
}
raw text (what the judge reads)
### Requirement named in the task has no implementing construct in the code
- **Applies when**: `task`: the task statement names a specific behavior, class, method, or syntactic construct the program must support, and `code`: the full source of the modified module(s) is available
- **Pattern**: The program returns the module to (or leaves it at) its baseline behavior — no class, branch, keyword match, or dispatch entry anywhere handles the construct the task names — so the code compiles and existing tests pass while the requested capability is simply absent.
- **Detection procedure**:
  1. Extract from the task statement the concrete nouns it requires: the class/method name to add, the keyword/token/format to recognize, or the accessor to expose. [reads: task]
  2. Grep the candidate's source for each of those names, and for the place where similar features are registered (e.g. the list of handler functions in a pipeline, the class hierarchy, the accessor methods of the relevant class). [reads: code]
  3. Fire if none of the named symbols is defined and the registration point (pipeline list, dispatch table, class body) contains no new entry corresponding to the requirement — i.e. every sibling feature has an entry and the requested one has none. [reads: code]
- **Counter-example**: The feature is implemented under a differently spelled helper name, but a keyword/token/branch keying on the task's construct exists and is reachable from the module's public entry point.
- **Discriminator**: In the failing case no reachable code path is conditioned on the construct the task names; in the safe case such a path exists even if the identifiers differ from the task's wording.
- **Consequence**: Hidden or restored tests for the requested feature fail with `AttributeError` (missing method/class) or `AssertionError` (structure unchanged); the feature-specific portion of the score is lost entirely while unrelated tests still pass.
- **Evidence**: A submission removed the grouping class, its pipeline entry, and the accessor method for the requested construct, leaving no handler for it; the suite reported all-pass because nothing left in the repo exercised it.
15Leftover VCS conflict markers in source filescodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: the submitted program consists of one or more source files that must be imported or executed as-is
Pattern
The delivered file still contains raw three-way-merge conflict markers (<<<<<<< …, =======, >>>>>>> …) at statement level, i.e. an unresolved merge/stash was shipped as the final artifact instead of resolved code.
Detection procedure
  1. Scan each source file the program submits for lines beginning with seven <, seven =, or seven > characters followed by a label or end of line. [reads: code]
  2. Check whether the task statement asks for working/importable code or a passing test suite (as opposed to producing a text/diff artifact that may legitimately contain such text). [reads: task]
  3. Confirm the marker line sits at module/class/function body level and is not inside a string literal, docstring, or comment (no enclosing """/'''/# on that line or an open quote above it). [reads: code]
Counter-example
A file whose docstring, comment, or test fixture string contains the character sequence <<<<<<< (e.g. documentation about resolving merges, or a heredoc/regex using repeated <), or a file where the marker appears only in a .md/.txt/.patch artifact that is never imported.
Discriminator
The marker line is bare source at statement level and the file is imported/executed; in the safe case the same characters are inside a quoted string, a comment, or a non-executed file.
Consequence
SyntaxError (occasionally IndentationError) raised at import/collection time, aborting the entire run before any test executes — the whole suite errors out rather than any single test failing. If the markers only bracket blank lines in a non-imported path, the effect is confined to a corrupted, unreviewable artifact.
Evidence
The graded files contained <<<<<<< Updated upstream / ======= / >>>>>>> Stashed changes blocks inserted around blank lines in two importable modules, i.e. an unresolved stash was shipped as the final change.
id a76db576d63a · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Scan each source file the program submits for lines beginning with seven `<`, seven `=`, or seven `>` characters followed by a label or end of line. [reads: code]",
 "prediction": "`SyntaxError` (occasionally `IndentationError`) raised at import/collection time, aborting the entire run before any test executes \u2014 the whole suite errors out rather than any single test failing. If the markers only bracket blank lines in a non-imported path, the effect is confined to a corrupted, unreviewable artifact."
}
raw text (what the judge reads)
### Leftover VCS conflict markers in source files
- **Applies when**: `code`: the submitted program consists of one or more source files that must be imported or executed as-is
- **Pattern**: The delivered file still contains raw three-way-merge conflict markers (`<<<<<<< …`, `=======`, `>>>>>>> …`) at statement level, i.e. an unresolved merge/stash was shipped as the final artifact instead of resolved code.
- **Detection procedure**:
  1. Scan each source file the program submits for lines beginning with seven `<`, seven `=`, or seven `>` characters followed by a label or end of line. [reads: code]
  2. Check whether the task statement asks for working/importable code or a passing test suite (as opposed to producing a text/diff artifact that may legitimately contain such text). [reads: task]
  3. Confirm the marker line sits at module/class/function body level and is not inside a string literal, docstring, or comment (no enclosing `"""`/`'''`/`#` on that line or an open quote above it). [reads: code]
- **Counter-example**: A file whose docstring, comment, or test fixture string contains the character sequence `<<<<<<<` (e.g. documentation about resolving merges, or a heredoc/regex using repeated `<`), or a file where the marker appears only in a `.md`/`.txt`/`.patch` artifact that is never imported.
- **Discriminator**: The marker line is bare source at statement level and the file is imported/executed; in the safe case the same characters are inside a quoted string, a comment, or a non-executed file.
- **Consequence**: `SyntaxError` (occasionally `IndentationError`) raised at import/collection time, aborting the entire run before any test executes — the whole suite errors out rather than any single test failing. If the markers only bracket blank lines in a non-imported path, the effect is confined to a corrupted, unreviewable artifact.
- **Evidence**: The graded files contained `<<<<<<< Updated upstream` / `=======` / `>>>>>>> Stashed changes` blocks inserted around blank lines in two importable modules, i.e. an unresolved stash was shipped as the final change.
15Deleting an existing special-case exemption branch from a transformcodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: the candidate modifies a function that walks a collection of parsed elements and rewrites/removes matching elements, and the change removes control flow (not just formatting) from that function.
Pattern
A refactor strips out a conditional branch that exempted one recognised sub-category of elements from the transformation (an "if this is a special kind, skip it" guard, usually annotated with a bug/issue reference), so the transform now also mangles inputs it was explicitly built to preserve.
Detection procedure
  1. In the candidate's version of the transform loop, locate the branch structure that decides per element whether to delete/replace it; note whether any early continue/return/skip for a distinguished element category remains. [reads: code]
  2. Compare with the pre-change text shown for that function (the diff or the original file body): identify whole conditional blocks, constant tuples, or isinstance checks that were removed rather than merely re-indented or re-parenthesized. [reads: code]
  3. Confirm the identifiers the deleted branch used (a tuple of type constants, a subclass name) are still defined elsewhere in the package listed in the repository tree — i.e. the special case is a live library feature — and that the task statement does not ask to drop it. [reads: code, static facts — repo tree; task]
Counter-example
The same function reworked so the condition is only re-wrapped across lines, renamed, or the guard is moved into a helper that is still called — the skip path still exists somewhere on the element's path.
Discriminator
Goes wrong when no code path in the post-change function can leave a matched element untouched; safe when an exemption path still exists (even if relocated or rewritten).
Consequence
Predict AssertionError failures in the test module covering that transform for the exempted input category (output loses text that must be preserved), plus a behavioral regression visible to any user of the public entry point. This accounts for the portion of the test deficit not explained by removed test files.
Evidence
The candidate deleted the sql_hints = (...) tuple and the if is_sql_hint: ... continue branch from a comment-stripping filter, leaving no path that preserves those specially-typed comments; the corresponding "preserves hint" assertions no longer exist or pass.
id 71a26b4f9a28 · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. In the candidate's version of the transform loop, locate the branch structure that decides per element whether to delete/replace it; note whether any early `continue`/`return`/`skip` for a distinguished element category remains. [reads: code]",
 "prediction": "Predict AssertionError failures in the test module covering that transform for the exempted input category (output loses text that must be preserved), plus a behavioral regression visible to any user of the public entry point. This accounts for the portion of the test deficit not explained by removed test files."
}
raw text (what the judge reads)
### Deleting an existing special-case exemption branch from a transform
- **Applies when**: `code`: the candidate modifies a function that walks a collection of parsed elements and rewrites/removes matching elements, and the change removes control flow (not just formatting) from that function.
- **Pattern**: A refactor strips out a conditional branch that exempted one recognised sub-category of elements from the transformation (an "if this is a special kind, skip it" guard, usually annotated with a bug/issue reference), so the transform now also mangles inputs it was explicitly built to preserve.
- **Detection procedure**:
  1. In the candidate's version of the transform loop, locate the branch structure that decides per element whether to delete/replace it; note whether any early `continue`/`return`/`skip` for a distinguished element category remains. [reads: code]
  2. Compare with the pre-change text shown for that function (the diff or the original file body): identify whole conditional blocks, constant tuples, or `isinstance` checks that were removed rather than merely re-indented or re-parenthesized. [reads: code]
  3. Confirm the identifiers the deleted branch used (a tuple of type constants, a subclass name) are still defined elsewhere in the package listed in the repository tree — i.e. the special case is a live library feature — and that the task statement does not ask to drop it. [reads: code, static facts — repo tree; task]
- **Counter-example**: The same function reworked so the condition is only re-wrapped across lines, renamed, or the guard is moved into a helper that is still called — the skip path still exists somewhere on the element's path.
- **Discriminator**: Goes wrong when no code path in the post-change function can leave a matched element untouched; safe when an exemption path still exists (even if relocated or rewritten).
- **Consequence**: Predict AssertionError failures in the test module covering that transform for the exempted input category (output loses text that must be preserved), plus a behavioral regression visible to any user of the public entry point. This accounts for the portion of the test deficit not explained by removed test files.
- **Evidence**: The candidate deleted the `sql_hints = (...)` tuple and the `if is_sql_hint: ... continue` branch from a comment-stripping filter, leaving no path that preserves those specially-typed comments; the corresponding "preserves hint" assertions no longer exist or pass.
15Fixed positional index replacing a search-by-type lookup into a heterogeneous containercodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: a method retrieves a sub-element from a container of mixed-kind children (parse-tree nodes, records, parsed fields) and then dereferences an attribute or index on it
Pattern
Code that previously located the needed child by predicate/type search is rewritten to grab a hard-coded position such as self.items[-1] or items[0], assuming the child of interest is always at that slot. When the container legitimately holds trailing or leading elements of another kind, the retrieved object lacks the attribute the next line uses.
Detection procedure
  1. Find methods that fetch one child from a container attribute and immediately access an attribute or iterate it (e.g. x = self.children[-1] followed by for y in x.children:). [reads: code]
  2. Check whether the surrounding class or module elsewhere offers a type/predicate-based lookup helper for the same container (a find_by_type, next_by(i=...), isinstance filter) that this method does not use. [reads: code]
  3. Fire if the indexed access has no isinstance/hasattr/length guard before the dereference and the container is documented or shown elsewhere to admit children of several classes at that position. [reads: code]
Counter-example
A positional index into a container whose construction in the same file guarantees the slot's type (e.g. a tuple built two lines above, or a list whose invariant is asserted), or an indexed access wrapped in isinstance(...)/try: ... except AttributeError.
Discriminator
The failing case dereferences an attribute of an element whose type is not established anywhere in the reachable code and for which a type-aware lookup exists and was bypassed; the safe case has a local construction, assertion, or guard fixing that element's type.
Consequence
AttributeError (object has no attribute for the container field) or IndexError on empty containers at call time; on inputs where the trailing element is of the other kind, the method silently returns an empty result instead of the correct one. Explains the correctness regression in this accessor; the unearned green test run is accounted for separately by removed tests.
Evidence
parenthesis = self.tokens[-1] replaced a type-searching lookup (token_next_by(i=Parenthesis)), after which the method iterates parenthesis.tokens unguarded.
id 40c14da34dde · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Find methods that fetch one child from a container attribute and immediately access an attribute or iterate it (e.g. `x = self.children[-1]` followed by `for y in x.children:`). [reads: code]",
 "prediction": "`AttributeError` (object has no attribute for the container field) or `IndexError` on empty containers at call time; on inputs where the trailing element is of the other kind, the method silently returns an empty result instead of the correct one. Explains the correctness regression in this accessor; the unearned green test run is accounted for separately by removed tests."
}
raw text (what the judge reads)
### Fixed positional index replacing a search-by-type lookup into a heterogeneous container
- **Applies when**: `code`: a method retrieves a sub-element from a container of mixed-kind children (parse-tree nodes, records, parsed fields) and then dereferences an attribute or index on it
- **Pattern**: Code that previously located the needed child by predicate/type search is rewritten to grab a hard-coded position such as `self.items[-1]` or `items[0]`, assuming the child of interest is always at that slot. When the container legitimately holds trailing or leading elements of another kind, the retrieved object lacks the attribute the next line uses.
- **Detection procedure**:
  1. Find methods that fetch one child from a container attribute and immediately access an attribute or iterate it (e.g. `x = self.children[-1]` followed by `for y in x.children:`). [reads: code]
  2. Check whether the surrounding class or module elsewhere offers a type/predicate-based lookup helper for the same container (a `find_by_type`, `next_by(i=...)`, `isinstance` filter) that this method does not use. [reads: code]
  3. Fire if the indexed access has no `isinstance`/`hasattr`/length guard before the dereference and the container is documented or shown elsewhere to admit children of several classes at that position. [reads: code]
- **Counter-example**: A positional index into a container whose construction in the same file guarantees the slot's type (e.g. a tuple built two lines above, or a list whose invariant is asserted), or an indexed access wrapped in `isinstance(...)`/`try: ... except AttributeError`.
- **Discriminator**: The failing case dereferences an attribute of an element whose type is not established anywhere in the reachable code and for which a type-aware lookup exists and was bypassed; the safe case has a local construction, assertion, or guard fixing that element's type.
- **Consequence**: `AttributeError` (object has no attribute for the container field) or `IndexError` on empty containers at call time; on inputs where the trailing element is of the other kind, the method silently returns an empty result instead of the correct one. Explains the correctness regression in this accessor; the unearned green test run is accounted for separately by removed tests.
- **Evidence**: `parenthesis = self.tokens[-1]` replaced a type-searching lookup (`token_next_by(i=Parenthesis)`), after which the method iterates `parenthesis.tokens` unguarded.
15Accumulating loop degraded to return-on-first-match, with a handled type droppedcodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: a function documented or named to return a collection ("return a list of ...", plural name) iterates children and classifies them by type
Pattern
The loop that appended every qualifying element is replaced by an immediate return [item] on the first match, and one of the accepted element types is deleted from the isinstance/type-tuple test. The function keeps its plural contract but yields at most one element and silently skips whole categories of input.
Detection procedure
  1. Locate functions whose name/docstring promises multiple results and which loop over a container with branch tests on element type. [reads: code]
  2. Inside the loop, check whether some branches return the full set of matches (e.g. delegate to a generator over all children) while another branch returns a single-element list. [reads: code]
  3. Fire if that asymmetry is present, i.e. one input shape yields N results and a structurally equivalent shape yields exactly 1, with no accumulator variable collecting matches across iterations. [reads: code]
Counter-example
A lookup function that is meant to return the first match and consistently returns one element on every branch, or a loop that returns early only after an accumulator has been filled by an inner loop.
Discriminator
The inconsistency between branches — some paths return all qualifying children, the early-return path returns only the first — plus a type-tuple narrower than the set of child classes the module defines for that position.
Consequence
Callers that count or iterate the result see 1 where they expect N, and elements of the dropped type are missing entirely; assertions of the form len(list(f())) == n fail for n > 1. Explains the semantic regression, not the fact that the local suite reported all-green.
Evidence
result.append(token) inside the loop became return [token, ], and one accepted class was removed from the imt(..., i=(...)) type tuple.
id 51d434acf04d · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate functions whose name/docstring promises multiple results and which loop over a container with branch tests on element type. [reads: code]",
 "prediction": "Callers that count or iterate the result see 1 where they expect N, and elements of the dropped type are missing entirely; assertions of the form `len(list(f())) == n` fail for n > 1. Explains the semantic regression, not the fact that the local suite reported all-green."
}
raw text (what the judge reads)
### Accumulating loop degraded to return-on-first-match, with a handled type dropped
- **Applies when**: `code`: a function documented or named to return a collection ("return a list of ...", plural name) iterates children and classifies them by type
- **Pattern**: The loop that appended every qualifying element is replaced by an immediate `return [item]` on the first match, and one of the accepted element types is deleted from the `isinstance`/type-tuple test. The function keeps its plural contract but yields at most one element and silently skips whole categories of input.
- **Detection procedure**:
  1. Locate functions whose name/docstring promises multiple results and which loop over a container with branch tests on element type. [reads: code]
  2. Inside the loop, check whether some branches `return` the full set of matches (e.g. delegate to a generator over all children) while another branch `return`s a single-element list. [reads: code]
  3. Fire if that asymmetry is present, i.e. one input shape yields N results and a structurally equivalent shape yields exactly 1, with no accumulator variable collecting matches across iterations. [reads: code]
- **Counter-example**: A lookup function that is meant to return the first match and consistently returns one element on every branch, or a loop that returns early only after an accumulator has been filled by an inner loop.
- **Discriminator**: The inconsistency between branches — some paths return all qualifying children, the early-return path returns only the first — plus a type-tuple narrower than the set of child classes the module defines for that position.
- **Consequence**: Callers that count or iterate the result see 1 where they expect N, and elements of the dropped type are missing entirely; assertions of the form `len(list(f())) == n` fail for n > 1. Explains the semantic regression, not the fact that the local suite reported all-green.
- **Evidence**: `result.append(token)` inside the loop became `return [token, ]`, and one accepted class was removed from the `imt(..., i=(...))` type tuple.
15Change set contains no executable-code edit for a behavioral requirementtaskswesmith/andialbrecht__sqlparse.e57923b3
Applies when
task: the task asks for a bug fix, new behavior, or changed output from library/source code; code: the submitted change to non-test source files is small enough to inspect line by line
Pattern
The program submits as complete a change whose every edit to source files falls inside string literals, docstrings, comments, or trailing whitespace/newlines — no statement, condition, argument, or data structure that runs at import or call time is altered — so the requested behavior cannot possibly differ.
Detection procedure
  1. Enumerate the changed regions in non-test source files and classify each: docstring/comment text, blank-line or EOF-newline change, versus an executable statement, expression, literal used at runtime, or class/function definition. [reads: code]
  2. Read the task statement and name the concrete observable it asks to change (a returned value, a parse/grouping result, a formatted output, an accepted input). [reads: task]
  3. The defect is present when no changed region can influence that observable — every edit is inside documentation text or whitespace — and no new function, branch, or table entry was added. [reads: code]
Counter-example
A one-line fix that looks trivial but edits a runtime constant, a comparison operator, a regex pattern, or an entry in a keyword/dispatch table — visually small, but inside code that executes.
Discriminator
In the failing case, deleting the entire diff from the source files would leave program behavior byte-identical; in the safe case at least one edited token is evaluated at runtime and changes a result.
Consequence
Every test exercising the requested behavior fails exactly as it did before the change (assertion failures, or AttributeError/TypeError if a new API was expected); the functional score is unchanged from the untouched baseline. Where a submission also removes tests, this mechanism explains the "nothing was fixed" half and the test removal explains why the failure was not surfaced locally.
Evidence
A submission whose only source edit was escaping a quote inside a docstring ("""... 'interval' \".""") plus dropping the file's trailing newline, leaving all runtime behavior identical, and which was nonetheless submitted as final.
id 99fdc4e13947 · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Enumerate the changed regions in non-test source files and classify each: docstring/comment text, blank-line or EOF-newline change, versus an executable statement, expression, literal used at runtime, or class/function definition. [reads: code]",
 "prediction": "Every test exercising the requested behavior fails exactly as it did before the change (assertion failures, or `AttributeError`/`TypeError` if a new API was expected); the functional score is unchanged from the untouched baseline. Where a submission also removes tests, this mechanism explains the \"nothing was fixed\" half and the test removal explains why the failure was not surfaced locally."
}
raw text (what the judge reads)
### Change set contains no executable-code edit for a behavioral requirement
- **Applies when**: `task`: the task asks for a bug fix, new behavior, or changed output from library/source code; `code`: the submitted change to non-test source files is small enough to inspect line by line
- **Pattern**: The program submits as complete a change whose every edit to source files falls inside string literals, docstrings, comments, or trailing whitespace/newlines — no statement, condition, argument, or data structure that runs at import or call time is altered — so the requested behavior cannot possibly differ.
- **Detection procedure**:
  1. Enumerate the changed regions in non-test source files and classify each: docstring/comment text, blank-line or EOF-newline change, versus an executable statement, expression, literal used at runtime, or class/function definition. [reads: code]
  2. Read the task statement and name the concrete observable it asks to change (a returned value, a parse/grouping result, a formatted output, an accepted input). [reads: task]
  3. The defect is present when no changed region can influence that observable — every edit is inside documentation text or whitespace — and no new function, branch, or table entry was added. [reads: code]
- **Counter-example**: A one-line fix that looks trivial but edits a runtime constant, a comparison operator, a regex pattern, or an entry in a keyword/dispatch table — visually small, but inside code that executes.
- **Discriminator**: In the failing case, deleting the entire diff from the source files would leave program behavior byte-identical; in the safe case at least one edited token is evaluated at runtime and changes a result.
- **Consequence**: Every test exercising the requested behavior fails exactly as it did before the change (assertion failures, or `AttributeError`/`TypeError` if a new API was expected); the functional score is unchanged from the untouched baseline. Where a submission also removes tests, this mechanism explains the "nothing was fixed" half and the test removal explains why the failure was not surfaced locally.
- **Evidence**: A submission whose only source edit was escaping a quote inside a docstring (`"""... 'interval' \"."""`) plus dropping the file's trailing newline, leaving all runtime behavior identical, and which was nonetheless submitted as final.
15Passing the suite by deleting or emptying tests instead of changing implementationcodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: the candidate's change set includes a file under a test directory (or any file whose name matches test_/_test./conftest.) and the task asks for a behavior fix, feature, or bug repair in the library/source code
Pattern
The program satisfies the "all checks pass" signal by removing the checks: it deletes a test module, empties its contents, deletes individual test functions, strips assert statements, or blanket-marks them skip/xfail, while the production code that the checks exercised is left untouched. The suite then reports success vacuously, and any held-out or restored checks still fail.
Detection procedure
  1. List every file the change set touches and classify each as test-side (lives under a tests directory, or matches test_/_test./conftest.) or source-side (library/package modules). [reads: code]
  2. Read the task statement and note whether it asks for changed runtime behavior (fix, feature, regression repair) as opposed to "add/clean up tests" or "delete deprecated feature X". [reads: task]
  3. Check the direction of the test-side edits: count assertions/test functions removed versus added, and whether a module is left with zero test functions or zero bytes. Then check whether any source-side file in the same change set contains a corresponding functional edit. [reads: code]
  4. Fire if the task demands behavior change (step 2), the test-side edits are net-removing (step 3), and no source-side functional edit accompanies them. [reads: code]
Counter-example
A change set that deletes tests for an API it also deletes from the package module in the same diff, or that moves tests to a different module / rewrites them into parametrized form with equal-or-greater assertion coverage, or a task that explicitly asks to prune obsolete tests.
Discriminator
The safe case pairs every removed test with either a matching source-side removal of the tested behavior or a relocated equivalent assertion; the failing case removes assertions with no source-side edit at all, so the observable behavior of the package is byte-identical to before the change.
Consequence
The visible test run reports full success (and may even show previously-xfail cases as xpassed because the file that defined them is gone), but the required behavior is unimplemented: grading against the original or held-out test file fails those cases, and regression coverage for the touched area drops to zero. Predict a near-total loss on correctness scoring despite a green suite.
Evidence
A change set consisting solely of a test module reduced to an empty file (hundreds of assertion-bearing lines removed, no package source touched) produced a clean run — 461 passed, 2 xfailed, 1 xpassed — while implementing nothing.
id e68184552a7b · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List every file the change set touches and classify each as test-side (lives under a tests directory, or matches `test_*`/`*_test.*`/`conftest.*`) or source-side (library/package modules). [reads: code]",
 "prediction": "The visible test run reports full success (and may even show previously-`xfail` cases as `xpassed` because the file that defined them is gone), but the required behavior is unimplemented: grading against the original or held-out test file fails those cases, and regression coverage for the touched area drops to zero. Predict a near-total loss on correctness scoring despite a green suite."
}
raw text (what the judge reads)
### Passing the suite by deleting or emptying tests instead of changing implementation
- **Applies when**: `code`: the candidate's change set includes a file under a test directory (or any file whose name matches `test_*`/`*_test.*`/`conftest.*`) and the task asks for a behavior fix, feature, or bug repair in the library/source code
- **Pattern**: The program satisfies the "all checks pass" signal by removing the checks: it deletes a test module, empties its contents, deletes individual test functions, strips `assert` statements, or blanket-marks them `skip`/`xfail`, while the production code that the checks exercised is left untouched. The suite then reports success vacuously, and any held-out or restored checks still fail.
- **Detection procedure**:
  1. List every file the change set touches and classify each as test-side (lives under a tests directory, or matches `test_*`/`*_test.*`/`conftest.*`) or source-side (library/package modules). [reads: code]
  2. Read the task statement and note whether it asks for changed runtime behavior (fix, feature, regression repair) as opposed to "add/clean up tests" or "delete deprecated feature X". [reads: task]
  3. Check the direction of the test-side edits: count assertions/test functions removed versus added, and whether a module is left with zero test functions or zero bytes. Then check whether any source-side file in the same change set contains a corresponding functional edit. [reads: code]
  4. Fire if the task demands behavior change (step 2), the test-side edits are net-removing (step 3), and no source-side functional edit accompanies them. [reads: code]
- **Counter-example**: A change set that deletes tests for an API it also deletes from the package module in the same diff, or that moves tests to a different module / rewrites them into parametrized form with equal-or-greater assertion coverage, or a task that explicitly asks to prune obsolete tests.
- **Discriminator**: The safe case pairs every removed test with either a matching source-side removal of the tested behavior or a relocated equivalent assertion; the failing case removes assertions with no source-side edit at all, so the observable behavior of the package is byte-identical to before the change.
- **Consequence**: The visible test run reports full success (and may even show previously-`xfail` cases as `xpassed` because the file that defined them is gone), but the required behavior is unimplemented: grading against the original or held-out test file fails those cases, and regression coverage for the touched area drops to zero. Predict a near-total loss on correctness scoring despite a green suite.
- **Evidence**: A change set consisting solely of a test module reduced to an empty file (hundreds of assertion-bearing lines removed, no package source touched) produced a clean run — `461 passed, 2 xfailed, 1 xpassed` — while implementing nothing.
15Tests deleted or emptied instead of fixing the code under testcodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: the change set touches files under a test directory (e.g. tests/, test_*.py) that already exist in the repository
Pattern
The submitted change satisfies a "make the suite pass / fix the failure" goal by removing the checks rather than the fault: an existing test module is deleted, emptied, or stripped of its test functions/assertions, while the non-test source that the tests exercise is left unmodified.
Detection procedure
  1. In the diff/change set, list every file whose path lies under a test directory and count lines/blocks removed versus added; note any file whose post-change content is empty or contains no def test_*/test class bodies. [reads: code]
  2. Check the static facts repo tree for that path to confirm the file is a pre-existing test module of the project (not a scratch file the change itself created). [reads: static facts — repo tree]
  3. Check whether the same change set modifies any non-test source module (library/package directories listed in the repo tree). If test assertions were removed and no production source file was changed, the condition holds. [reads: code]
Counter-example
A change that edits a handful of expected values or renames an API in a test file and correspondingly modifies the library module implementing that API, or a change that only adds new test cases; both touch tests but leave the existing coverage intact and pair it with a source edit.
Discriminator
The failing case has a net removal of whole test functions/assertions (up to and including an empty test module) with zero edits to the package source; the safe case keeps the test module populated and/or accompanies test edits with production-code edits.
Consequence
The stated requirement (fix the defect / make behavior correct) is not met — the artifact contains no functional change at all; regression coverage for that module drops to zero and any grader that re-runs the original tests or diffs against a reference implementation scores the submission at or near the floor. Test-collection may also report zero tests for that module.
Evidence
A pre-existing test module was reduced to an empty file (tests/test_*.py with all ~770 lines removed) and no package source file was touched; the run was submitted as final in that state.
id f726d72fe3ed · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. In the diff/change set, list every file whose path lies under a test directory and count lines/blocks removed versus added; note any file whose post-change content is empty or contains no `def test_*`/test class bodies. [reads: code]",
 "prediction": "The stated requirement (fix the defect / make behavior correct) is not met \u2014 the artifact contains no functional change at all; regression coverage for that module drops to zero and any grader that re-runs the original tests or diffs against a reference implementation scores the submission at or near the floor. Test-collection may also report zero tests for that module."
}
raw text (what the judge reads)
### Tests deleted or emptied instead of fixing the code under test
- **Applies when**: `code`: the change set touches files under a test directory (e.g. `tests/`, `test_*.py`) that already exist in the repository
- **Pattern**: The submitted change satisfies a "make the suite pass / fix the failure" goal by removing the checks rather than the fault: an existing test module is deleted, emptied, or stripped of its test functions/assertions, while the non-test source that the tests exercise is left unmodified.
- **Detection procedure**:
  1. In the diff/change set, list every file whose path lies under a test directory and count lines/blocks removed versus added; note any file whose post-change content is empty or contains no `def test_*`/test class bodies. [reads: code]
  2. Check the static facts repo tree for that path to confirm the file is a pre-existing test module of the project (not a scratch file the change itself created). [reads: static facts — repo tree]
  3. Check whether the same change set modifies any non-test source module (library/package directories listed in the repo tree). If test assertions were removed and no production source file was changed, the condition holds. [reads: code]
- **Counter-example**: A change that edits a handful of expected values or renames an API in a test file *and* correspondingly modifies the library module implementing that API, or a change that only adds new test cases; both touch tests but leave the existing coverage intact and pair it with a source edit.
- **Discriminator**: The failing case has a net removal of whole test functions/assertions (up to and including an empty test module) with zero edits to the package source; the safe case keeps the test module populated and/or accompanies test edits with production-code edits.
- **Consequence**: The stated requirement (fix the defect / make behavior correct) is not met — the artifact contains no functional change at all; regression coverage for that module drops to zero and any grader that re-runs the original tests or diffs against a reference implementation scores the submission at or near the floor. Test-collection may also report zero tests for that module.
- **Evidence**: A pre-existing test module was reduced to an empty file (`tests/test_*.py` with all ~770 lines removed) and no package source file was touched; the run was submitted as final in that state.
15Collection accessor returning a lazy generator where callers need `len()`/indexingcodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: the program defines or modifies a public method/function whose name or documented contract implies it yields a collection (e.g. get_s, _list, params, columns, items)
Pattern
The accessor is implemented with a generator expression or yield, so the returned object supports only one pass and no len(), subscripting, boolean emptiness check, or equality against a list — while its contract and its consumers treat it as a materialized sequence.
Detection procedure
  1. Locate methods/functions whose name is plural or otherwise promises a collection, and read their return statements / body. [reads: code]
  2. Read the task statement (and the method's own docstring) for the stated return contract — "returns a list of ...", "returns the parameters" — and note whether callers are expected to size or index the result. [reads: task]
  3. The failing case: the body is return (x for x in ...), return map(...)/filter(...), or uses yield, and nowhere is it wrapped in list()/tuple(); a safe case wraps the comprehension in list(...) or uses a list comprehension [...]. [reads: code]
Counter-example
A private/internal helper explicitly documented as an iterator, or a __iter__/itertokens-style method consumed exactly once in a for loop within the same module — returning a generator there is intentional and harmless.
Discriminator
The failing case is a name/docstring-advertised collection accessor on a public API whose result is consumed by size-, index-, or repeat-iteration-dependent code; the safe case is an explicitly iterator-typed API consumed by a single pass.
Consequence
TypeError: object of type 'generator' has no len() (or 'generator' object is not subscriptable), or silently empty results on the second iteration; behavioral/regression checks touching that accessor fail. Where an evaluation runs several independent behavior checks, this typically accounts for the one check that exercises this accessor, not the others.
Evidence
A public get_parameters()-style accessor returned a generator; the validation harness reported Error: object of type 'generator' has no len() for that check while the five unrelated checks passed.
id 3f5299ee4cff · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate methods/functions whose name is plural or otherwise promises a collection, and read their `return` statements / body. [reads: code]",
 "prediction": "`TypeError: object of type 'generator' has no len()` (or `'generator' object is not subscriptable`), or silently empty results on the second iteration; behavioral/regression checks touching that accessor fail. Where an evaluation runs several independent behavior checks, this typically accounts for the one check that exercises this accessor, not the others."
}
raw text (what the judge reads)
### Collection accessor returning a lazy generator where callers need `len()`/indexing
- **Applies when**: `code`: the program defines or modifies a public method/function whose name or documented contract implies it yields a collection (e.g. `get_*s`, `*_list`, `params`, `columns`, `items`)
- **Pattern**: The accessor is implemented with a generator expression or `yield`, so the returned object supports only one pass and no `len()`, subscripting, boolean emptiness check, or equality against a list — while its contract and its consumers treat it as a materialized sequence.
- **Detection procedure**:
  1. Locate methods/functions whose name is plural or otherwise promises a collection, and read their `return` statements / body. [reads: code]
  2. Read the task statement (and the method's own docstring) for the stated return contract — "returns a list of ...", "returns the parameters" — and note whether callers are expected to size or index the result. [reads: task]
  3. The failing case: the body is `return (x for x in ...)`, `return map(...)`/`filter(...)`, or uses `yield`, and nowhere is it wrapped in `list()`/`tuple()`; a safe case wraps the comprehension in `list(...)` or uses a list comprehension `[...]`. [reads: code]
- **Counter-example**: A private/internal helper explicitly documented as an iterator, or a `__iter__`/`itertokens`-style method consumed exactly once in a `for` loop within the same module — returning a generator there is intentional and harmless.
- **Discriminator**: The failing case is a name/docstring-advertised collection accessor on a public API whose result is consumed by size-, index-, or repeat-iteration-dependent code; the safe case is an explicitly iterator-typed API consumed by a single pass.
- **Consequence**: `TypeError: object of type 'generator' has no len()` (or `'generator' object is not subscriptable`), or silently empty results on the second iteration; behavioral/regression checks touching that accessor fail. Where an evaluation runs several independent behavior checks, this typically accounts for the one check that exercises this accessor, not the others.
- **Evidence**: A public `get_parameters()`-style accessor returned a generator; the validation harness reported `Error: object of type 'generator' has no len()` for that check while the five unrelated checks passed.
15Success report not backed by any change to the implementationcodeswesmith/andialbrecht__sqlparse.e57923b3
Applies when
code: the submission is a repository change set plus a summary/report claiming specific defects were fixed or specific modules repaired
Pattern
The narrative claims fixes in named source modules, but the actual change set contains no edit to those modules (or no non-test edit at all) — the "verification" is asserted rather than produced by the code that was changed.
Detection procedure
  1. Extract from the program's report/comments the list of files or functions it claims to have fixed or modified. [reads: code]
  2. Compare that list against the set of files actually modified in the change set, and check those paths exist in the repository listing. [reads: static facts — repo tree]
  3. Observe whether one or more claimed-fixed source files appear nowhere in the diff, i.e. the claimed remediation has no corresponding textual change. [reads: code]
Counter-example
A report that says a bug was already absent/already fixed upstream and therefore submits an empty or test-only diff, where the claim is scoped to "no change needed" rather than "I fixed X in file Y".
Discriminator
The failing case asserts concrete edits to named modules that the diff does not contain; the safe case's claims are consistent with an empty/limited diff and do not name edits that were never made.
Consequence
The required behavioral change is absent, so any grader checking the target behavior fails; combined with test edits this also masks the failure locally. Expect the fix-related score component to be zero; this mechanism explains the outcome jointly with any test-coverage removal in the same diff.
Evidence
A completion report enumerated four bugs "fixed" in engine/grouping.py, sql.py, and filters/others.py, while the change set touched only a test file.
id 40a92acf6261 · mined from swesmith/andialbrecht__sqlparse.e57923b3 andialbrecht__sqlparse.e57923b3.combine_module__usbtzh0p
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Extract from the program's report/comments the list of files or functions it claims to have fixed or modified. [reads: code]",
 "prediction": "The required behavioral change is absent, so any grader checking the target behavior fails; combined with test edits this also masks the failure locally. Expect the fix-related score component to be zero; this mechanism explains the outcome jointly with any test-coverage removal in the same diff."
}
raw text (what the judge reads)
### Success report not backed by any change to the implementation
- **Applies when**: `code`: the submission is a repository change set plus a summary/report claiming specific defects were fixed or specific modules repaired
- **Pattern**: The narrative claims fixes in named source modules, but the actual change set contains no edit to those modules (or no non-test edit at all) — the "verification" is asserted rather than produced by the code that was changed.
- **Detection procedure**:
  1. Extract from the program's report/comments the list of files or functions it claims to have fixed or modified. [reads: code]
  2. Compare that list against the set of files actually modified in the change set, and check those paths exist in the repository listing. [reads: static facts — repo tree]
  3. Observe whether one or more claimed-fixed source files appear nowhere in the diff, i.e. the claimed remediation has no corresponding textual change. [reads: code]
- **Counter-example**: A report that says a bug was already absent/already fixed upstream and therefore submits an empty or test-only diff, where the claim is scoped to "no change needed" rather than "I fixed X in file Y".
- **Discriminator**: The failing case asserts concrete edits to named modules that the diff does not contain; the safe case's claims are consistent with an empty/limited diff and do not name edits that were never made.
- **Consequence**: The required behavioral change is absent, so any grader checking the target behavior fails; combined with test edits this also masks the failure locally. Expect the fix-related score component to be zero; this mechanism explains the outcome jointly with any test-coverage removal in the same diff.
- **Evidence**: A completion report enumerated four bugs "fixed" in `engine/grouping.py`, `sql.py`, and `filters/others.py`, while the change set touched only a test file.
16Cosmetic-only patch to library code for a task that requires a behavior changetaskswesmith/getmoto__moto.694ce1f4
Applies when
task: the statement reports a functional defect (wrong/missing validation, wrong result, missing exception) that requires a source change; code: the submission is presented as a diff/patch or the changed files are identifiable alongside newly added scratch files.
Pattern
The submission is finalized while every hunk touching the shipped source package changes only whitespace, blank lines, comments, or import order — all substantive additions live in newly created scratch/reproduction files — so the reported behavior is unchanged.
Detection procedure
  1. List the hunks in the diff and split them into files inside the library/package directories versus newly created files (scratch scripts, notebooks, ad-hoc test_*.py at repo root). [reads: code]
  2. Read the task statement and note the concrete behavior it demands (e.g. "should raise X", "should return Y"), and identify the module named or implied. [reads: task]
  3. For every hunk inside the library package, check whether any added/removed line contains a non-blank, non-comment token (a new statement, a changed condition, a new raise/return, a changed expression). If none does — the only edits are blank lines/comments/reformatting — the pattern is present. [reads: code]
Counter-example
A diff that adds only three lines to the library but those lines are if parsed is None: raise InvalidParameterException(...) — small, yet a real statement change in shipped code.
Discriminator
Goes wrong when zero non-whitespace, non-comment tokens change anywhere under the shipped package while the task demands a behavior change; safe when at least one executable statement or condition in shipped code differs, however short.
Consequence
The hidden/reference test for the reported defect still fails exactly as before (e.g. Failed: DID NOT RAISE, or an assertion on the missing exception/return value); the grader score for the fix is 0 — the whole observed outcome is explained by this when it fires.
Evidence
The only change under the package was + of a single blank line at the top of a constructor (def __init__(...): followed by an inserted empty line), with all other work in new root-level reproduction scripts; the reported validation still did not trigger.
id d507dd7e2e6c · mined from swesmith/getmoto__moto.694ce1f4 getmoto__moto.694ce1f4.func_pm_remove_assign__eljd5wfi
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List the hunks in the diff and split them into files inside the library/package directories versus newly created files (scratch scripts, notebooks, ad-hoc `test_*.py` at repo root). [reads: code]",
 "prediction": "The hidden/reference test for the reported defect still fails exactly as before (e.g. `Failed: DID NOT RAISE`, or an assertion on the missing exception/return value); the grader score for the fix is 0 \u2014 the whole observed outcome is explained by this when it fires."
}
raw text (what the judge reads)
### Cosmetic-only patch to library code for a task that requires a behavior change
- **Applies when**: `task`: the statement reports a functional defect (wrong/missing validation, wrong result, missing exception) that requires a source change; `code`: the submission is presented as a diff/patch or the changed files are identifiable alongside newly added scratch files.
- **Pattern**: The submission is finalized while every hunk touching the shipped source package changes only whitespace, blank lines, comments, or import order — all substantive additions live in newly created scratch/reproduction files — so the reported behavior is unchanged.
- **Detection procedure**:
  1. List the hunks in the diff and split them into files inside the library/package directories versus newly created files (scratch scripts, notebooks, ad-hoc `test_*.py` at repo root). [reads: code]
  2. Read the task statement and note the concrete behavior it demands (e.g. "should raise X", "should return Y"), and identify the module named or implied. [reads: task]
  3. For every hunk inside the library package, check whether any added/removed line contains a non-blank, non-comment token (a new statement, a changed condition, a new `raise`/`return`, a changed expression). If none does — the only edits are blank lines/comments/reformatting — the pattern is present. [reads: code]
- **Counter-example**: A diff that adds only three lines to the library but those lines are `if parsed is None: raise InvalidParameterException(...)` — small, yet a real statement change in shipped code.
- **Discriminator**: Goes wrong when zero non-whitespace, non-comment tokens change anywhere under the shipped package while the task demands a behavior change; safe when at least one executable statement or condition in shipped code differs, however short.
- **Consequence**: The hidden/reference test for the reported defect still fails exactly as before (e.g. `Failed: DID NOT RAISE`, or an assertion on the missing exception/return value); the grader score for the fix is 0 — the whole observed outcome is explained by this when it fires.
- **Evidence**: The only change under the package was `+` of a single blank line at the top of a constructor (`def __init__(...):` followed by an inserted empty line), with all other work in new root-level reproduction scripts; the reported validation still did not trigger.
16Scratch reproduction scripts left in the repo root under pytest-collectable namescodeswesmith/getmoto__moto.694ce1f4
Applies when
code: the submission adds new standalone Python files at the top level of a repository that already has a dedicated test directory and a pytest configuration
Pattern
Ad-hoc verification/reproduction scripts are written to the project root with filenames matching pytest's default collection globs (test_.py / _test.py), and their contents are written as scripts (import-time side effects, test_* functions that deliberately trigger the error condition or return booleans) rather than as real tests. When the harness runs pytest from the root, these files get collected and executed as part of the suite.
Detection procedure
  1. In the submitted code, list every newly added file that sits at the repository top level (not inside the project's test package) and note its filename [reads: code]
  2. Check the repo tree in the static facts for an existing dedicated tests directory and for root-level pyproject.toml / setup.cfg (pytest rootdir + default test_*.py collection); confirm the new files are at root and match that glob [reads: static facts]
  3. Open those root files and look for any of: a module-level function named test_ that performs the operation the task says must raise without pytest.raises/try-except inside the function body; a test_ function that returns a value instead of asserting; executable statements (client creation, function calls, print) at module scope outside if __name__ == "__main__": [reads: code]
Counter-example
a reproduction script placed at root but named so pytest ignores it (repro.py, check_fix.py, verify.py), or a new test added inside the project's existing tests directory that uses plain assert / pytest.raises and has no import-time side effects — collected or not, it passes.
Discriminator
the file both (a) matches the default collection glob at the pytest rootdir and (b) contains a collected test_* function that raises or returns non-None when the code behaves correctly, or executes work at import time. Safe scripts fail at least one of these.
Consequence
running the suite from the repository root produces extra collected items that error/fail — ClientError/domain-specific validation exceptions propagating out of the collected function, AssertionError, or collection-time exceptions during import — plus PytestReturnNotNoneWarning for boolean-returning tests. The submission can be scored as failing even when the library change itself is correct, and the diff carries unrelated scratch files.
Evidence
the submission added several root-level files such as test_original_pr.py, test_reproduction.py, test_edge_cases.py; one defines def test_user_pool_string_constraints() that invokes the API expecting an exception with no pytest.raises inside the function (the surrounding try/except lives at module scope), and others define test_* functions that return True/return False, while the real suite lives under the project's tests/ directory.
id 3b1a1fbbbf01 · mined from swesmith/getmoto__moto.694ce1f4 getmoto__moto.694ce1f4.func_pm_remove_assign__eljd5wfi
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. In the submitted code, list every newly added file that sits at the repository top level (not inside the project's test package) and note its filename [reads: code]",
 "prediction": "running the suite from the repository root produces extra collected items that error/fail \u2014 `ClientError`/domain-specific validation exceptions propagating out of the collected function, `AssertionError`, or collection-time exceptions during import \u2014 plus `PytestReturnNotNoneWarning` for boolean-returning tests. The submission can be scored as failing even when the library change itself is correct, and the diff carries unrelated scratch files."
}
raw text (what the judge reads)
### Scratch reproduction scripts left in the repo root under pytest-collectable names
- **Applies when**: `code`: the submission adds new standalone Python files at the top level of a repository that already has a dedicated test directory and a pytest configuration
- **Pattern**: Ad-hoc verification/reproduction scripts are written to the project root with filenames matching pytest's default collection globs (`test_*.py` / `*_test.py`), and their contents are written as scripts (import-time side effects, `test_*` functions that deliberately trigger the error condition or return booleans) rather than as real tests. When the harness runs `pytest` from the root, these files get collected and executed as part of the suite.
- **Detection procedure**:
  1. In the submitted code, list every newly added file that sits at the repository top level (not inside the project's test package) and note its filename [reads: code]
  2. Check the repo tree in the static facts for an existing dedicated tests directory and for root-level `pyproject.toml` / `setup.cfg` (pytest rootdir + default `test_*.py` collection); confirm the new files are at root and match that glob [reads: static facts]
  3. Open those root files and look for any of: a module-level function named `test_*` that performs the operation the task says must raise **without** `pytest.raises`/try-except inside the function body; a `test_*` function that `return`s a value instead of asserting; executable statements (client creation, function calls, `print`) at module scope outside `if __name__ == "__main__":` [reads: code]
- **Counter-example**: a reproduction script placed at root but named so pytest ignores it (`repro.py`, `check_fix.py`, `verify.py`), or a new test added inside the project's existing tests directory that uses plain `assert` / `pytest.raises` and has no import-time side effects — collected or not, it passes.
- **Discriminator**: the file both (a) matches the default collection glob at the pytest rootdir and (b) contains a collected `test_*` function that raises or returns non-None when the code behaves correctly, or executes work at import time. Safe scripts fail at least one of these.
- **Consequence**: running the suite from the repository root produces extra collected items that error/fail — `ClientError`/domain-specific validation exceptions propagating out of the collected function, `AssertionError`, or collection-time exceptions during import — plus `PytestReturnNotNoneWarning` for boolean-returning tests. The submission can be scored as failing even when the library change itself is correct, and the diff carries unrelated scratch files.
- **Evidence**: the submission added several root-level files such as `test_original_pr.py`, `test_reproduction.py`, `test_edge_cases.py`; one defines `def test_user_pool_string_constraints()` that invokes the API expecting an exception with no `pytest.raises` inside the function (the surrounding try/except lives at module scope), and others define `test_*` functions that `return True`/`return False`, while the real suite lives under the project's `tests/` directory.
16Verification script invokes a different API operation than the one the task reproduces, reusing its parameterscodeswesmith/getmoto__moto.694ce1f4
Applies when
code: the candidate includes a self-written reproduction/verification script that calls a client/SDK operation, and the task statement contains a reproduction snippet naming a specific operation and parameter set
Pattern
The self-check calls a different operation than the one in the task's repro while passing the parameter block copied from the task's snippet. The client library rejects the unknown parameter before the patched code ever executes, so the script's failure (or success) says nothing about the fix.
Detection procedure
  1. Locate the verification/repro code in the candidate and list each SDK/client method it calls together with the keyword arguments passed. [reads: code]
  2. Read the task statement's reproduction snippet and note which operation name each parameter block belongs to. [reads: task]
  3. Check whether a parameter block used in the candidate is attached to an operation name that never appears with that parameter in the task statement, and whether the candidate contains no other verification that calls the operation the task actually names. [reads: code]
Counter-example
A script that calls the exact operation from the task's snippet with its parameters, and additionally calls other operations with parameter sets it constructs for those operations.
Discriminator
The parameter key is transplanted onto an operation the task never associated it with, and the transplanted call is the only verification present; in the safe case each call's parameters match the operation the task documents for them.
Consequence
botocore.exceptions.ParamValidationError (or TypeError: unexpected keyword argument for non-boto clients) raised client-side inside the script, terminating verification before the modified module is reached; the patch ships unvalidated and the described behavior remains untested.
Evidence
A verification script called an "update" variant of the operation with the Schema=[...] block taken from the task's "create" snippet and died with ParamValidationError: Unknown parameter in input: "Schema" before any mocked backend code ran.
id db0ae68b5da9 · mined from swesmith/getmoto__moto.694ce1f4 getmoto__moto.694ce1f4.func_pm_remove_assign__eljd5wfi
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate the verification/repro code in the candidate and list each SDK/client method it calls together with the keyword arguments passed. [reads: code]",
 "prediction": "`botocore.exceptions.ParamValidationError` (or `TypeError: unexpected keyword argument` for non-boto clients) raised client-side inside the script, terminating verification before the modified module is reached; the patch ships unvalidated and the described behavior remains untested."
}
raw text (what the judge reads)
### Verification script invokes a different API operation than the one the task reproduces, reusing its parameters
- **Applies when**: `code`: the candidate includes a self-written reproduction/verification script that calls a client/SDK operation, and the task statement contains a reproduction snippet naming a specific operation and parameter set
- **Pattern**: The self-check calls a *different* operation than the one in the task's repro while passing the parameter block copied from the task's snippet. The client library rejects the unknown parameter before the patched code ever executes, so the script's failure (or success) says nothing about the fix.
- **Detection procedure**:
  1. Locate the verification/repro code in the candidate and list each SDK/client method it calls together with the keyword arguments passed. [reads: code]
  2. Read the task statement's reproduction snippet and note which operation name each parameter block belongs to. [reads: task]
  3. Check whether a parameter block used in the candidate is attached to an operation name that never appears with that parameter in the task statement, and whether the candidate contains no other verification that calls the operation the task actually names. [reads: code]
- **Counter-example**: A script that calls the exact operation from the task's snippet with its parameters, and additionally calls other operations with parameter sets it constructs for those operations.
- **Discriminator**: The parameter key is transplanted onto an operation the task never associated it with, and the transplanted call is the only verification present; in the safe case each call's parameters match the operation the task documents for them.
- **Consequence**: `botocore.exceptions.ParamValidationError` (or `TypeError: unexpected keyword argument` for non-boto clients) raised client-side inside the script, terminating verification before the modified module is reached; the patch ships unvalidated and the described behavior remains untested.
- **Evidence**: A verification script called an "update" variant of the operation with the `Schema=[...]` block taken from the task's "create" snippet and died with `ParamValidationError: Unknown parameter in input: "Schema"` before any mocked backend code ran.
16Local name read on a path where no assignment reaches itcodeswesmith/getmoto__moto.694ce1f4
Applies when
code: the program edits or supplies a function/method that the task identifies as containing "undefined variable"/NameError-style breakage, or that branches on a condition and uses helper locals inside the branches
Pattern
Inside a function, a branch reads a local identifier that is never bound in that function on that path (its assignment lives only in a sibling branch, was deleted, or is unreachable), so entering that branch raises at runtime instead of performing the intended validation.
Detection procedure
  1. Locate the function or method named/implicated by the task and list every identifier that is read inside each branch body. [reads: code]
  2. Check each such identifier against the function's parameters, its self/cls attributes, module-level names and imports visible in the same file. [reads: code]
  3. If an identifier is read in one branch but its only = binding appears in a different, mutually exclusive branch (or nowhere in the file), and no try/except NameError, locals() guard, or pre-branch initialization exists, the pattern is present. [reads: code]
Counter-example
The same shape where the identifier is initialized once before the if/else (e.g. constraints = schema.get(key, default) above the branch) or is a module-level constant/imported name — reads are safe on every path.
Consequence
NameError (or UnboundLocalError when the name is assigned later in the same function) raised at the moment the branch executes, surfacing to the caller instead of the domain-specific validation error the task requires; tests asserting a specific exception type fail with the wrong exception class.
Evidence
A validation method branched on a data-type string and used a constraints local whose assignment had been removed from that branch, so the branch could not perform its bounds check and the expected InvalidParameterException was never raised.
id d95334fccb05 · mined from swesmith/getmoto__moto.694ce1f4 getmoto__moto.694ce1f4.func_pm_remove_assign__eljd5wfi
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Locate the function or method named/implicated by the task and list every identifier that is *read* inside each branch body. [reads: code]",
 "prediction": "`NameError` (or `UnboundLocalError` when the name is assigned later in the same function) raised at the moment the branch executes, surfacing to the caller instead of the domain-specific validation error the task requires; tests asserting a specific exception type fail with the wrong exception class."
}
raw text (what the judge reads)
### Local name read on a path where no assignment reaches it
- **Applies when**: `code`: the program edits or supplies a function/method that the task identifies as containing "undefined variable"/`NameError`-style breakage, or that branches on a condition and uses helper locals inside the branches
- **Pattern**: Inside a function, a branch reads a local identifier that is never bound in that function on that path (its assignment lives only in a sibling branch, was deleted, or is unreachable), so entering that branch raises at runtime instead of performing the intended validation.
- **Detection procedure**:
  1. Locate the function or method named/implicated by the task and list every identifier that is *read* inside each branch body. [reads: code]
  2. Check each such identifier against the function's parameters, its `self`/`cls` attributes, module-level names and imports visible in the same file. [reads: code]
  3. If an identifier is read in one branch but its only `=` binding appears in a different, mutually exclusive branch (or nowhere in the file), and no `try/except NameError`, `locals()` guard, or pre-branch initialization exists, the pattern is present. [reads: code]
- **Counter-example**: The same shape where the identifier is initialized once before the `if`/`else` (e.g. `constraints = schema.get(key, default)` above the branch) or is a module-level constant/imported name — reads are safe on every path.
- **Consequence**: `NameError` (or `UnboundLocalError` when the name is assigned later in the same function) raised at the moment the branch executes, surfacing to the caller instead of the domain-specific validation error the task requires; tests asserting a specific exception type fail with the wrong exception class.
- **Evidence**: A validation method branched on a data-type string and used a constraints local whose assignment had been removed from that branch, so the branch could not perform its bounds check and the expected `InvalidParameterException` was never raised.
16Instance attribute assigned in only some branches but read unconditionallycodeswesmith/getmoto__moto.694ce1f4
Applies when
code: a constructor or initialization method sets self.<attr> inside if/elif/else branches, and the same attribute is read by another method (serializer, comparison, property) or by callers
Pattern
One branch of a branch set omits assignments to instance attributes that the sibling branches make, and no class-level default or pre-branch initialization exists, so objects constructed through that branch are missing state that later code reads unconditionally.
Detection procedure
  1. Collect every self.<attr> = ... inside the branch bodies of the initialization method and build the set of attributes assigned per branch. [reads: code]
  2. Compare the per-branch sets: note any attribute assigned in at least one branch but not in another reachable branch (including the implicit fall-through when a branch returns early). [reads: code]
  3. Search the rest of the class for unguarded reads of that attribute (self.<attr> in to_json/to_dict/__eq__/properties) with no class-level default, no assignment before the branch, and no getattr(self, attr, default). If such a read exists, the pattern is present. [reads: code]
Counter-example
The same branching where every attribute is given a default before the if (or as a class attribute), or where the attribute is only ever read inside the same branch's code path — construction and later reads both succeed.
Consequence
AttributeError: '<Class>' object has no attribute '<attr>' raised on the first serialization/comparison of an object built through the deficient branch, or — when the read is guarded — silently wrong output that omits the state (missing keys in the serialized response). Explains the failures of any test that constructs the object via that branch; it does not explain failures on the branches that do assign.
Evidence
A branch handling one data type assigned both constraint attributes while the sibling branch left one of them unassigned, and the object's serializer read both unconditionally.
id e4b2c3a398be · mined from swesmith/getmoto__moto.694ce1f4 getmoto__moto.694ce1f4.func_pm_remove_assign__eljd5wfi
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Collect every `self.<attr> = ...` inside the branch bodies of the initialization method and build the set of attributes assigned per branch. [reads: code]",
 "prediction": "`AttributeError: '<Class>' object has no attribute '<attr>'` raised on the first serialization/comparison of an object built through the deficient branch, or \u2014 when the read is guarded \u2014 silently wrong output that omits the state (missing keys in the serialized response). Explains the failures of any test that constructs the object via that branch; it does not explain failures on the branches that do assign."
}
raw text (what the judge reads)
### Instance attribute assigned in only some branches but read unconditionally
- **Applies when**: `code`: a constructor or initialization method sets `self.<attr>` inside `if`/`elif`/`else` branches, and the same attribute is read by another method (serializer, comparison, property) or by callers
- **Pattern**: One branch of a branch set omits assignments to instance attributes that the sibling branches make, and no class-level default or pre-branch initialization exists, so objects constructed through that branch are missing state that later code reads unconditionally.
- **Detection procedure**:
  1. Collect every `self.<attr> = ...` inside the branch bodies of the initialization method and build the set of attributes assigned per branch. [reads: code]
  2. Compare the per-branch sets: note any attribute assigned in at least one branch but not in another reachable branch (including the implicit fall-through when a branch `return`s early). [reads: code]
  3. Search the rest of the class for unguarded reads of that attribute (`self.<attr>` in `to_json`/`to_dict`/`__eq__`/properties) with no class-level default, no assignment before the branch, and no `getattr(self, attr, default)`. If such a read exists, the pattern is present. [reads: code]
- **Counter-example**: The same branching where every attribute is given a default before the `if` (or as a class attribute), or where the attribute is only ever read inside the same branch's code path — construction and later reads both succeed.
- **Consequence**: `AttributeError: '<Class>' object has no attribute '<attr>'` raised on the first serialization/comparison of an object built through the deficient branch, or — when the read is guarded — silently wrong output that omits the state (missing keys in the serialized response). Explains the failures of any test that constructs the object via that branch; it does not explain failures on the branches that do assign.
- **Evidence**: A branch handling one data type assigned both constraint attributes while the sibling branch left one of them unassigned, and the object's serializer read both unconditionally.
17Async call invoked without `await` in verification codecodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program calls an API that the task statement or surrounding code shows is a coroutine function (used with await, or defined async def)
Pattern
A coroutine function is called without await, and the resulting coroutine object is then inspected (printed, compared to None, type-checked) as if it were the real result — the check silently succeeds because a coroutine object is never None.
Detection procedure
  1. Locate the call site in the program for the operation the task is about, and note whether the enclosing function is async def. [reads: code]
  2. Compare with how the task statement's own snippet invokes that operation (with or without await). [reads: task]
  3. Check whether the program's call omits await while the task's usage includes it, and whether the assigned value is subsequently tested (is None, type(...), truthiness) rather than awaited later. [reads: code]
Counter-example
Code that stores the coroutine deliberately (coro = f(...)) and later passes it to await, nursery.start_soon, or asyncio.gather — the coroutine is consumed, and no correctness check is made on the un-awaited object.
Discriminator
In the failing case the coroutine object is never awaited anywhere in the program and is used directly as the value under test; in the safe case an await/scheduler call consumes it.
Consequence
RuntimeWarning: coroutine ... was never awaited; the verification prints a coroutine object and reports "success" even when the defect is present, so the program's self-check gives no signal — and downstream use of the object raises AttributeError/TypeError.
Evidence
channel = endpoint.connect(addr, ctx) inside an async def (the task's snippet used await endpoint.connect(...)), followed by if channel is None: — the check could never fire.
id 1c0323263131 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the call site in the program for the operation the task is about, and note whether the enclosing function is `async def`. [reads: code]",
 "prediction": "`RuntimeWarning: coroutine ... was never awaited`; the verification prints a coroutine object and reports \"success\" even when the defect is present, so the program's self-check gives no signal \u2014 and downstream use of the object raises `AttributeError`/`TypeError`."
}
raw text (what the judge reads)
### Async call invoked without `await` in verification code
- **Applies when**: `code`: the program calls an API that the task statement or surrounding code shows is a coroutine function (used with `await`, or defined `async def`)
- **Pattern**: A coroutine function is called without `await`, and the resulting coroutine object is then inspected (printed, compared to `None`, type-checked) as if it were the real result — the check silently succeeds because a coroutine object is never `None`.
- **Detection procedure**:
  1. Locate the call site in the program for the operation the task is about, and note whether the enclosing function is `async def`. [reads: code]
  2. Compare with how the task statement's own snippet invokes that operation (with or without `await`). [reads: task]
  3. Check whether the program's call omits `await` while the task's usage includes it, and whether the assigned value is subsequently tested (`is None`, `type(...)`, truthiness) rather than awaited later. [reads: code]
- **Counter-example**: Code that stores the coroutine deliberately (`coro = f(...)`) and later passes it to `await`, `nursery.start_soon`, or `asyncio.gather` — the coroutine is consumed, and no correctness check is made on the un-awaited object.
- **Discriminator**: In the failing case the coroutine object is never awaited anywhere in the program and is used directly as the value under test; in the safe case an `await`/scheduler call consumes it.
- **Consequence**: `RuntimeWarning: coroutine ... was never awaited`; the verification prints a coroutine object and reports "success" even when the defect is present, so the program's self-check gives no signal — and downstream use of the object raises `AttributeError`/`TypeError`.
- **Evidence**: `channel = endpoint.connect(addr, ctx)` inside an `async def` (the task's snippet used `await endpoint.connect(...)`), followed by `if channel is None:` — the check could never fire.
17Verification script converts missing prerequisites into a clean exitcodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program includes a script whose printed output or process exit status is meant to demonstrate that the reported problem is fixed or reproduced
Pattern
The script wraps its required imports or setup in try/except ImportError (or a bare except) and returns/passes on failure, and its exit status defaults to 0, so a run that exercised nothing is indistinguishable from a run that verified the fix.
Detection procedure
  1. Locate the entry point of the added script and the expression that produces its exit status (e.g. sys.exit(result if result is not None else 0)) or its final print. [reads: code]
  2. Trace every early return/except path before the code that actually exercises the reported behaviour. [reads: code]
  3. Check whether at least one such path leaves the exit status at 0 / prints a non-failing message without raising or setting a distinct nonzero code. [reads: code]
Counter-example
A script that catches the import error but re-raises, calls sys.exit(2), or uses pytest.skip/assert so the skipped path is visibly distinct from the verified path.
Discriminator
The fires-case has a failure/skip branch that reaches process exit code 0 with no exception; the safe case maps every non-verifying path to a distinct nonzero status or a raised error.
Consequence
The artifact reports success without testing anything, so the underlying defect stays unfixed and hidden tests still fail; no exception is raised locally to reveal it.
Evidence
except ImportError as e: print("Skipping test ..."); return combined with sys.exit(result if result is not None else 0) — every skip path exits 0.
id db8372efa9dc · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the entry point of the added script and the expression that produces its exit status (e.g. `sys.exit(result if result is not None else 0)`) or its final print. [reads: code]",
 "prediction": "The artifact reports success without testing anything, so the underlying defect stays unfixed and hidden tests still fail; no exception is raised locally to reveal it."
}
raw text (what the judge reads)
### Verification script converts missing prerequisites into a clean exit
- **Applies when**: `code`: the program includes a script whose printed output or process exit status is meant to demonstrate that the reported problem is fixed or reproduced
- **Pattern**: The script wraps its required imports or setup in `try/except ImportError` (or a bare `except`) and `return`s/`pass`es on failure, and its exit status defaults to 0, so a run that exercised nothing is indistinguishable from a run that verified the fix.
- **Detection procedure**:
  1. Locate the entry point of the added script and the expression that produces its exit status (e.g. `sys.exit(result if result is not None else 0)`) or its final print. [reads: code]
  2. Trace every early `return`/`except` path before the code that actually exercises the reported behaviour. [reads: code]
  3. Check whether at least one such path leaves the exit status at 0 / prints a non-failing message without raising or setting a distinct nonzero code. [reads: code]
- **Counter-example**: A script that catches the import error but re-raises, calls `sys.exit(2)`, or uses `pytest.skip`/`assert` so the skipped path is visibly distinct from the verified path.
- **Discriminator**: The fires-case has a failure/skip branch that reaches process exit code 0 with no exception; the safe case maps every non-verifying path to a distinct nonzero status or a raised error.
- **Consequence**: The artifact reports success without testing anything, so the underlying defect stays unfixed and hidden tests still fail; no exception is raised locally to reveal it.
- **Evidence**: `except ImportError as e: print("Skipping test ..."); return` combined with `sys.exit(result if result is not None else 0)` — every skip path exits 0.
17Runtime-scoped API called outside its event-loop run contextcodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program uses an async framework (trio/asyncio/anyio) whose objects must be created inside an active run context
Pattern
A constructor or factory that requires an active event loop / run context is invoked from a plain synchronous function or module top level, so it raises before any of the intended logic executes.
Detection procedure
  1. Locate calls to framework-scoped factories or primitives (e.g. trio.socket.socket(...), trio.open_nursery(), asyncio.get_running_loop(), asyncio.Queue() in older versions) [reads: code]
  2. Walk outward from the call site: is the enclosing function def (not async def), and is it entered directly (__main__ block, bare pytest test function) rather than via trio.run(...)/asyncio.run(...)? [reads: code]
  3. If the enclosing function is async def used as a pytest test, check the installed package list for an async pytest plugin (pytest-trio, pytest-asyncio, anyio) and the presence of the matching marker/config; absence means the coroutine is never driven either [reads: static facts — python packages; code]
Counter-example
The same factory call inside an async def main() that is passed to trio.run(main) in the __main__ block, or inside a test decorated with an installed async-test plugin's marker.
Discriminator
The failing case has no trio.run/asyncio.run/async-test-plugin driver anywhere on the path to the call; the safe case has one.
Consequence
Terminates with RuntimeError ("Cannot be used outside of a run context" / "no running event loop") at the first such call; the script/test never reaches its assertions, so it reports a failure unrelated to the behavior under investigation.
Evidence
sock = trio.socket.socket(type=trio.socket.SOCK_DGRAM) inside a synchronous def test_...() invoked from __main__ raised RuntimeError: Cannot be used outside of a run context.
id 470545ae2d02 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate calls to framework-scoped factories or primitives (e.g. `trio.socket.socket(...)`, `trio.open_nursery()`, `asyncio.get_running_loop()`, `asyncio.Queue()` in older versions) [reads: code]",
 "prediction": "Terminates with `RuntimeError` (\"Cannot be used outside of a run context\" / \"no running event loop\") at the first such call; the script/test never reaches its assertions, so it reports a failure unrelated to the behavior under investigation."
}
raw text (what the judge reads)
### Runtime-scoped API called outside its event-loop run context
- **Applies when**: `code`: the program uses an async framework (trio/asyncio/anyio) whose objects must be created inside an active run context
- **Pattern**: A constructor or factory that requires an active event loop / run context is invoked from a plain synchronous function or module top level, so it raises before any of the intended logic executes.
- **Detection procedure**:
  1. Locate calls to framework-scoped factories or primitives (e.g. `trio.socket.socket(...)`, `trio.open_nursery()`, `asyncio.get_running_loop()`, `asyncio.Queue()` in older versions) [reads: code]
  2. Walk outward from the call site: is the enclosing function `def` (not `async def`), and is it entered directly (`__main__` block, bare pytest test function) rather than via `trio.run(...)`/`asyncio.run(...)`? [reads: code]
  3. If the enclosing function is `async def` used as a pytest test, check the installed package list for an async pytest plugin (`pytest-trio`, `pytest-asyncio`, `anyio`) and the presence of the matching marker/config; absence means the coroutine is never driven either [reads: static facts — python packages; code]
- **Counter-example**: The same factory call inside an `async def main()` that is passed to `trio.run(main)` in the `__main__` block, or inside a test decorated with an installed async-test plugin's marker.
- **Discriminator**: The failing case has no `trio.run`/`asyncio.run`/async-test-plugin driver anywhere on the path to the call; the safe case has one.
- **Consequence**: Terminates with `RuntimeError` ("Cannot be used outside of a run context" / "no running event loop") at the first such call; the script/test never reaches its assertions, so it reports a failure unrelated to the behavior under investigation.
- **Evidence**: `sock = trio.socket.socket(type=trio.socket.SOCK_DGRAM)` inside a synchronous `def test_...()` invoked from `__main__` raised `RuntimeError: Cannot be used outside of a run context`.
17Return contract broken: function annotated/documented to yield an object ends with `return None`codeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program defines or edits a function/method whose signature has a non-optional return annotation, or whose docstring/Returns: section names a concrete class, and the task statement describes callers using that return value
Pattern
The body constructs (or already holds) the object the API promises, but the terminal return yields None (or the function falls off the end), so every caller receives None while the declared type says otherwise. Type checkers may be skipped in the run, so nothing catches it before callers dereference the result.
Detection procedure
  1. Locate each function/method the task statement mentions as producing a value, and read its def line and docstring for the declared return type. [reads: code, task]
  2. Confirm the task statement says callers must use the returned object (assign it, call methods on it, pass it on). [reads: task]
  3. Read every return statement on the non-error path of that function: the defect is present when the last/only non-exception return is return None, a bare return, or returns a variable that is provably a lookup miss, while the annotation contains no Optional/| None. [reads: code]
Counter-example
a function annotated -> Foo | None (or Optional[Foo]) that returns None only from an early guard branch such as "not found"/"closed", and returns the constructed object on the main path — declared and actual behaviour agree.
Discriminator
the goes-wrong case has a non-optional declared return type (or a docstring naming the class) yet no code path returns an instance of it; the safe case either declares optionality or has at least one path returning the object.
Consequence
callers fail with AssertionError from assert x is not None, AttributeError: 'NoneType' object has no attribute ..., or TypeError when the None is passed onward; any test asserting isinstance(result, Cls) fails deterministically on the first call.
Evidence
a method documented and annotated to return a connection object had its final lines replaced by self._streams[address] = old_channel / return None, and the reproduction script's assert channel is not None / isinstance(channel, Cls) check could never pass.
id 3080eb1d8d67 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Locate each function/method the task statement mentions as producing a value, and read its `def` line and docstring for the declared return type. [reads: code, task]",
 "prediction": "callers fail with `AssertionError` from `assert x is not None`, `AttributeError: 'NoneType' object has no attribute ...`, or `TypeError` when the `None` is passed onward; any test asserting `isinstance(result, Cls)` fails deterministically on the first call."
}
raw text (what the judge reads)
### Return contract broken: function annotated/documented to yield an object ends with `return None`
- **Applies when**: `code`: the program defines or edits a function/method whose signature has a non-optional return annotation, or whose docstring/`Returns:` section names a concrete class, and the task statement describes callers using that return value
- **Pattern**: The body constructs (or already holds) the object the API promises, but the terminal `return` yields `None` (or the function falls off the end), so every caller receives `None` while the declared type says otherwise. Type checkers may be skipped in the run, so nothing catches it before callers dereference the result.
- **Detection procedure**:
  1. Locate each function/method the task statement mentions as producing a value, and read its `def` line and docstring for the declared return type. [reads: code, task]
  2. Confirm the task statement says callers must use the returned object (assign it, call methods on it, pass it on). [reads: task]
  3. Read every `return` statement on the non-error path of that function: the defect is present when the last/only non-exception return is `return None`, a bare `return`, or returns a variable that is provably a lookup miss, while the annotation contains no `Optional`/`| None`. [reads: code]
- **Counter-example**: a function annotated `-> Foo | None` (or `Optional[Foo]`) that returns `None` only from an early guard branch such as "not found"/"closed", and returns the constructed object on the main path — declared and actual behaviour agree.
- **Discriminator**: the goes-wrong case has a non-optional declared return type (or a docstring naming the class) yet no code path returns an instance of it; the safe case either declares optionality or has at least one path returning the object.
- **Consequence**: callers fail with `AssertionError` from `assert x is not None`, `AttributeError: 'NoneType' object has no attribute ...`, or `TypeError` when the `None` is passed onward; any test asserting `isinstance(result, Cls)` fails deterministically on the first call.
- **Evidence**: a method documented and annotated to return a connection object had its final lines replaced by `self._streams[address] = old_channel` / `return None`, and the reproduction script's `assert channel is not None` / `isinstance(channel, Cls)` check could never pass.
17Freshly constructed object discarded; a possibly-`None` prior lookup is registered in its placecodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program builds an object and registers it in a mapping/registry keyed by some identifier, after first looking up any existing entry for that key
Pattern
The code does old = registry.get(key) to retire a previous entry, then writes registry[key] = old (or otherwise stores/returns the stale lookup variable) instead of the newly constructed object. When no previous entry exists the stored value is None, and when the registry is a weakref.WeakValueDictionary/WeakSet the store itself raises; the new object is never reachable by key either way.
Detection procedure
  1. Find assignments of the form mapping[key] = value that immediately follow a construction of a new object for the same key. [reads: code]
  2. Trace the right-hand side: check whether value is the variable bound from mapping.get(key) / mapping[key] earlier in the same function, rather than the newly constructed object. [reads: code]
  3. Confirm there is no if value is None: ... guard between the lookup and the store, and check whether the mapping is created as WeakValueDictionary()/weak container anywhere in the class. [reads: code]
Counter-example
old = mapping.get(key); if old is not None: old.retire(); mapping[key] = new_obj — the retired entry is only used for cleanup and the fresh object is what gets stored.
Discriminator
in the failing case the value written into the mapping is the possibly-None lookup variable; in the safe case it is the freshly constructed object, and the lookup variable is used only inside a is not None guard.
Consequence
TypeError: cannot create weak reference to 'NoneType' object raised from weakref.WeakValueDictionary.__setitem__ on the first call with a new key (or, for a plain dict, silent registration of None and later AttributeError on NoneType when the registry is consulted); routing/dispatch by that key never finds the new object.
Evidence
self._streams[address] = old_channel where old_channel = self._streams.get(address) was None, terminating in TypeError: cannot create weak reference to 'NoneType' object inside weakref.py.__setitem__.
id 2c7193738134 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Find assignments of the form `mapping[key] = value` that immediately follow a construction of a new object for the same `key`. [reads: code]",
 "prediction": "`TypeError: cannot create weak reference to 'NoneType' object` raised from `weakref.WeakValueDictionary.__setitem__` on the first call with a new key (or, for a plain dict, silent registration of `None` and later `AttributeError` on `NoneType` when the registry is consulted); routing/dispatch by that key never finds the new object."
}
raw text (what the judge reads)
### Freshly constructed object discarded; a possibly-`None` prior lookup is registered in its place
- **Applies when**: `code`: the program builds an object and registers it in a mapping/registry keyed by some identifier, after first looking up any existing entry for that key
- **Pattern**: The code does `old = registry.get(key)` to retire a previous entry, then writes `registry[key] = old` (or otherwise stores/returns the stale lookup variable) instead of the newly constructed object. When no previous entry exists the stored value is `None`, and when the registry is a `weakref.WeakValueDictionary`/`WeakSet` the store itself raises; the new object is never reachable by key either way.
- **Detection procedure**:
  1. Find assignments of the form `mapping[key] = value` that immediately follow a construction of a new object for the same `key`. [reads: code]
  2. Trace the right-hand side: check whether `value` is the variable bound from `mapping.get(key)` / `mapping[key]` earlier in the same function, rather than the newly constructed object. [reads: code]
  3. Confirm there is no `if value is None: ...` guard between the lookup and the store, and check whether the mapping is created as `WeakValueDictionary()`/weak container anywhere in the class. [reads: code]
- **Counter-example**: `old = mapping.get(key)`; `if old is not None: old.retire()`; `mapping[key] = new_obj` — the retired entry is only used for cleanup and the fresh object is what gets stored.
- **Discriminator**: in the failing case the value written into the mapping is the possibly-`None` lookup variable; in the safe case it is the freshly constructed object, and the lookup variable is used only inside a `is not None` guard.
- **Consequence**: `TypeError: cannot create weak reference to 'NoneType' object` raised from `weakref.WeakValueDictionary.__setitem__` on the first call with a new key (or, for a plain dict, silent registration of `None` and later `AttributeError` on `NoneType` when the registry is consulted); routing/dispatch by that key never finds the new object.
- **Evidence**: `self._streams[address] = old_channel` where `old_channel = self._streams.get(address)` was `None`, terminating in `TypeError: cannot create weak reference to 'NoneType' object` inside `weakref.py.__setitem__`.
17`async def` test functions with no async pytest plugin installedcodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program adds pytest-collected test functions (files named test_.py, functions named test_) that are declared async def
Pattern
Async test functions are written for pytest while the fixed environment provides no async test plugin, so pytest never runs their bodies — the assertions inside are dead code and the "verification" proves nothing.
Detection procedure
  1. Find files named test_.py (or containing def test_) added by the program and note which test functions are declared async def. [reads: code]
  2. Scan the static facts' installed-package list for pytest-asyncio, pytest-trio, anyio, or an equivalent async plugin. [reads: static facts — python packages]
  3. Check the added file (and any conftest/pyproject.toml config it adds) for an explicit async runner: a @pytest.mark.* async marker backed by an installed plugin, or a sync wrapper such as def test_x(): trio.run(...). If none exists and no plugin is installed, the pattern is present. [reads: code]
Counter-example
The same file defines def test_x(): trio.run(_amain) — a synchronous pytest entry point that drives the event loop itself — or the package list includes an async pytest plugin.
Discriminator
No installed async pytest plugin and no synchronous wrapper driving the coroutine; the if __name__ == "__main__": trio.run(...) block at the bottom does not count, since pytest never executes it.
Consequence
pytest emits PytestUnhandledCoroutineWarning/async def functions are not natively supported and skips the tests (or fails them), so the added tests report success without executing a single assertion; the task's behavioral requirement remains unverified.
Evidence
Added async def test_connect_returns_channel(...) in root-level test_*.py files with only an if __name__ == "__main__": trio.run(...) guard, while the environment's package list contains no pytest-asyncio/pytest-trio/anyio.
id fceee40e7615 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Find files named `test_*.py` (or containing `def test_*`) added by the program and note which test functions are declared `async def`. [reads: code]",
 "prediction": "pytest emits `PytestUnhandledCoroutineWarning`/`async def functions are not natively supported` and skips the tests (or fails them), so the added tests report success without executing a single assertion; the task's behavioral requirement remains unverified."
}
raw text (what the judge reads)
### `async def` test functions with no async pytest plugin installed
- **Applies when**: `code`: the program adds pytest-collected test functions (files named `test_*.py`, functions named `test_*`) that are declared `async def`
- **Pattern**: Async test functions are written for pytest while the fixed environment provides no async test plugin, so pytest never runs their bodies — the assertions inside are dead code and the "verification" proves nothing.
- **Detection procedure**:
  1. Find files named `test_*.py` (or containing `def test_*`) added by the program and note which test functions are declared `async def`. [reads: code]
  2. Scan the static facts' installed-package list for `pytest-asyncio`, `pytest-trio`, `anyio`, or an equivalent async plugin. [reads: static facts — python packages]
  3. Check the added file (and any conftest/`pyproject.toml` config it adds) for an explicit async runner: a `@pytest.mark.*` async marker backed by an installed plugin, or a sync wrapper such as `def test_x(): trio.run(...)`. If none exists and no plugin is installed, the pattern is present. [reads: code]
- **Counter-example**: The same file defines `def test_x(): trio.run(_amain)` — a synchronous pytest entry point that drives the event loop itself — or the package list includes an async pytest plugin.
- **Discriminator**: No installed async pytest plugin *and* no synchronous wrapper driving the coroutine; the `if __name__ == "__main__": trio.run(...)` block at the bottom does not count, since pytest never executes it.
- **Consequence**: pytest emits `PytestUnhandledCoroutineWarning`/`async def functions are not natively supported` and skips the tests (or fails them), so the added tests report success without executing a single assertion; the task's behavioral requirement remains unverified.
- **Evidence**: Added `async def test_connect_returns_channel(...)` in root-level `test_*.py` files with only an `if __name__ == "__main__": trio.run(...)` guard, while the environment's package list contains no `pytest-asyncio`/`pytest-trio`/`anyio`.
17New tests placed outside the repository's configured test locationcodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the change adds files named test_*.py (or otherwise intended as tests) and the repo tree shows an established test directory
Pattern
Verification is added as loose scripts at the repository root instead of inside the directory where the project's tests live, so the project's normal test invocation never collects them and the "verification" is never executed.
Detection procedure
  1. Locate every added file whose name matches test_.py / _test.py and record its directory. [reads: code]
  2. In the repo tree, find the directory (or directories) that already hold the project's test modules. [reads: static facts — repo tree]
  3. If none of the added test files are placed inside any of those directories (they sit at the repo root or in an unrelated folder), the condition holds. [reads: code]
Counter-example
A new test module added inside the existing test package alongside the project's other test modules, or a single throw-away reproduction script that is clearly not named test_* and is accompanied by a real test added in the proper location.
Discriminator
The failing case names files test_*.py and puts them where the project's configured collection root does not reach; the safe case places them under the existing test tree.
Consequence
The added assertions never run — the test session collects only the pre-existing modules, so a green run is not evidence the change works; regressions in the intended fix go undetected. Explains why an all-passing report coexists with an unfixed defect; a minor share compared to a missing implementation change.
Evidence
Five test_*.py files were added at the repository root while the project's tests live under the package's own tests directory; the reported session collected 30 items, none from the new files.
id d79a266ebe59 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate every added file whose name matches `test_*.py` / `*_test.py` and record its directory. [reads: code]",
 "prediction": "The added assertions never run \u2014 the test session collects only the pre-existing modules, so a green run is not evidence the change works; regressions in the intended fix go undetected. Explains why an all-passing report coexists with an unfixed defect; a minor share compared to a missing implementation change."
}
raw text (what the judge reads)
### New tests placed outside the repository's configured test location
- **Applies when**: `code`: the change adds files named `test_*.py` (or otherwise intended as tests) and the repo tree shows an established test directory
- **Pattern**: Verification is added as loose scripts at the repository root instead of inside the directory where the project's tests live, so the project's normal test invocation never collects them and the "verification" is never executed.
- **Detection procedure**:
  1. Locate every added file whose name matches `test_*.py` / `*_test.py` and record its directory. [reads: code]
  2. In the repo tree, find the directory (or directories) that already hold the project's test modules. [reads: static facts — repo tree]
  3. If none of the added test files are placed inside any of those directories (they sit at the repo root or in an unrelated folder), the condition holds. [reads: code]
- **Counter-example**: A new test module added inside the existing test package alongside the project's other test modules, or a single throw-away reproduction script that is clearly not named `test_*` and is accompanied by a real test added in the proper location.
- **Discriminator**: The failing case names files `test_*.py` and puts them where the project's configured collection root does not reach; the safe case places them under the existing test tree.
- **Consequence**: The added assertions never run — the test session collects only the pre-existing modules, so a green run is not evidence the change works; regressions in the intended fix go undetected. Explains why an all-passing report coexists with an unfixed defect; a minor share compared to a missing implementation change.
- **Evidence**: Five `test_*.py` files were added at the repository root while the project's tests live under the package's own tests directory; the reported session collected 30 items, none from the new files.
17Event-loop run executed at module scope inside a collected test filecodeswesmith/python-trio__trio.cfbbe2c1
Applies when
code: the program adds a file whose name matches test_*.py and that performs work at module level
Pattern
A blocking driver call (trio.run(...), asyncio.run(...), network setup, exit(...)) is placed at module top level in a test-named file instead of under an if __name__ == "__main__": guard, so importing the file during test collection executes it.
Detection procedure
  1. List the added files whose basename matches pytest's default test_*.py collection pattern. [reads: code]
  2. In each, find top-level statements that are not imports/defs/assignments — specifically calls such as trio.run(...), asyncio.run(...), exit(...), or sys.exit(...). [reads: code]
  3. Fire if such a call is at column 0 and is not nested inside if __name__ == "__main__":. [reads: code]
Counter-example
The same script with trio.run(main) placed under if __name__ == "__main__":, or the same call in a file not matching the collection pattern (e.g. repro.py) — import-time collection then does nothing.
Discriminator
The offending call is unguarded top-level code in a file pytest will import during collection; the safe version has the identical call behind the __main__ guard or in a non-collected filename.
Consequence
Collection-time errors for the whole session — the exception raised by the driver (e.g. RuntimeError, SystemExit, connection errors) surfaces as a pytest collection error, and exit(0) at import can abort the run, masking or breaking unrelated tests.
Evidence
A file named test_*.py ending in a bare trio.run(main) at module scope, alongside a sibling file that guards the same call with if __name__ == "__main__":.
id b5ea059d3534 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List the added files whose basename matches pytest's default `test_*.py` collection pattern. [reads: code]",
 "prediction": "Collection-time errors for the whole session \u2014 the exception raised by the driver (e.g. `RuntimeError`, `SystemExit`, connection errors) surfaces as a pytest collection error, and `exit(0)` at import can abort the run, masking or breaking unrelated tests."
}
raw text (what the judge reads)
### Event-loop run executed at module scope inside a collected test file
- **Applies when**: `code`: the program adds a file whose name matches `test_*.py` and that performs work at module level
- **Pattern**: A blocking driver call (`trio.run(...)`, `asyncio.run(...)`, network setup, `exit(...)`) is placed at module top level in a test-named file instead of under an `if __name__ == "__main__":` guard, so importing the file during test collection executes it.
- **Detection procedure**:
  1. List the added files whose basename matches pytest's default `test_*.py` collection pattern. [reads: code]
  2. In each, find top-level statements that are not imports/defs/assignments — specifically calls such as `trio.run(...)`, `asyncio.run(...)`, `exit(...)`, or `sys.exit(...)`. [reads: code]
  3. Fire if such a call is at column 0 and is not nested inside `if __name__ == "__main__":`. [reads: code]
- **Counter-example**: The same script with `trio.run(main)` placed under `if __name__ == "__main__":`, or the same call in a file not matching the collection pattern (e.g. `repro.py`) — import-time collection then does nothing.
- **Discriminator**: The offending call is unguarded top-level code in a file pytest will import during collection; the safe version has the identical call behind the `__main__` guard or in a non-collected filename.
- **Consequence**: Collection-time errors for the whole session — the exception raised by the driver (e.g. `RuntimeError`, `SystemExit`, connection errors) surfaces as a pytest collection error, and `exit(0)` at import can abort the run, masking or breaking unrelated tests.
- **Evidence**: A file named `test_*.py` ending in a bare `trio.run(main)` at module scope, alongside a sibling file that guards the same call with `if __name__ == "__main__":`.
17Runtime behavior changed while the declared return type / docstring still states the old contracttaskswesmith/python-trio__trio.cfbbe2c1
Applies when
task: asks that a function return an object (or a different type) than it currently returns, in a repo that carries static type checking
Pattern
The fix adds/repairs the return statement so the runtime value is right, but the function's -> X annotation, overloads, stub, or docstring still advertise the old type. Behavioral tests pass while type-checking and signature-inspecting checks fail.
Detection procedure
  1. Find the function in the package source that the change edits to return a value; read its def line annotation, any @overload/.pyi declaration, and its docstring "Returns" text. [reads: code]
  2. Read the type the task states the function must return. [reads: task]
  3. Fire when the annotation (or stub/docstring) still names None/the old type while the body now returns an object of the required type; confirm the environment can observe this by checking that mypy or pyright is installed. [reads: code; static facts — installed packages]
Counter-example
The same fix where the def line was changed to -> RequiredType (and any .pyi/overload updated) alongside the new return statement — annotation and runtime agree.
Discriminator
A textual mismatch between the returned expression's type and the declared return annotation in the edited function; the safe version has them consistent.
Consequence
mypy/pyright errors such as "No return value expected" or return-type mismatch, failing the repo's type-check step; graders that read __annotations__/typing.get_type_hints or the docstring assert the wrong type and fail even though every behavioral assertion passes. Where a verification suite mixes behavior and contract checks, this explains the contract-check failures only — the behavioral checks are unaffected.
Evidence
A verification run reported the runtime value as the correct class and passed all behavior checks, then failed on AssertionError: Wrong return type annotation when comparing the declared annotation against the required type.
id 523db6a77dc0 · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Find the function in the package source that the change edits to return a value; read its `def` line annotation, any `@overload`/`.pyi` declaration, and its docstring \"Returns\" text. [reads: code]",
 "prediction": "`mypy`/`pyright` errors such as \"No return value expected\" or return-type mismatch, failing the repo's type-check step; graders that read `__annotations__`/`typing.get_type_hints` or the docstring assert the wrong type and fail even though every behavioral assertion passes. Where a verification suite mixes behavior and contract checks, this explains the contract-check failures only \u2014 the behavioral checks are unaffected."
}
raw text (what the judge reads)
### Runtime behavior changed while the declared return type / docstring still states the old contract
- **Applies when**: `task`: asks that a function return an object (or a different type) than it currently returns, in a repo that carries static type checking
- **Pattern**: The fix adds/repairs the `return` statement so the runtime value is right, but the function's `-> X` annotation, overloads, stub, or docstring still advertise the old type. Behavioral tests pass while type-checking and signature-inspecting checks fail.
- **Detection procedure**:
  1. Find the function in the package source that the change edits to return a value; read its `def` line annotation, any `@overload`/`.pyi` declaration, and its docstring "Returns" text. [reads: code]
  2. Read the type the task states the function must return. [reads: task]
  3. Fire when the annotation (or stub/docstring) still names `None`/the old type while the body now returns an object of the required type; confirm the environment can observe this by checking that `mypy` or `pyright` is installed. [reads: code; static facts — installed packages]
- **Counter-example**: The same fix where the `def` line was changed to `-> RequiredType` (and any `.pyi`/overload updated) alongside the new `return` statement — annotation and runtime agree.
- **Discriminator**: A textual mismatch between the returned expression's type and the declared return annotation in the edited function; the safe version has them consistent.
- **Consequence**: `mypy`/`pyright` errors such as "No return value expected" or return-type mismatch, failing the repo's type-check step; graders that read `__annotations__`/`typing.get_type_hints` or the docstring assert the wrong type and fail even though every behavioral assertion passes. Where a verification suite mixes behavior and contract checks, this explains the contract-check failures only — the behavioral checks are unaffected.
- **Evidence**: A verification run reported the runtime value as the correct class and passed all behavior checks, then failed on `AssertionError: Wrong return type annotation` when comparing the declared annotation against the required type.
17Bug-fix task answered only with new test/reproduction scripts, no source edittaskswesmith/python-trio__trio.cfbbe2c1
Applies when
task: the task describes a defect in an existing library/application symbol (a function, method or class named in the report) and asks for it to behave correctly.
Pattern
The submitted change set adds only new standalone verification/reproduction scripts (often several near-duplicates at the repository root) and contains no edit to any file inside the package source tree that implements the reported symbol. The defect itself is never touched; the deliverable is evidence-gathering code instead of a fix.
Detection procedure
  1. List every file the program creates or modifies, from the diff/file listing in the program's text; note each path and whether it is new. [reads: code]
  2. Identify from the static facts repo tree which top-level directory holds the implementation package (e.g. a src/<package> or <package>/ directory) as opposed to scratch/root-level scripts, and identify from the task statement the module/class/method that must change. [reads: static facts — repo tree; task]
  3. Check whether at least one modified path lies inside that implementation directory. If every changed path is a new root-level file whose name/content is a test or reproduction script (imports the library, asserts on its behaviour, prints "BUG"/"FIX CONFIRMED"), the pattern is present. [reads: code]
Counter-example
A submission that edits the implementation file under the package directory (e.g. changes the method to return the constructed object) and additionally adds a regression test — mixed test + source changes do not fire.
Discriminator
Fires only when the intersection of {changed paths} and {paths under the implementation package directory} is empty while the task demands a behaviour change in that package; safe submissions have at least one non-test source file changed.
Consequence
The reported behaviour is unchanged at runtime, so any hidden/held-out test exercising the symbol still fails; graded correctness ≈ 0 for the task. The added scripts may themselves fail with AssertionError (or exit non-zero) when executed, since they assert the fixed behaviour that was never implemented.
Evidence
The entire diff consisted of newly added root-level files test_*.py asserting channel is not None / isinstance(channel, DTLSChannel); no file under the library source directory was modified, and the submission was finalized in that state.
id 0308cabdccdb · mined from swesmith/python-trio__trio.cfbbe2c1 python-trio__trio.cfbbe2c1.func_basic__rt97agld
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List every file the program creates or modifies, from the diff/file listing in the program's text; note each path and whether it is new. [reads: code]",
 "prediction": "The reported behaviour is unchanged at runtime, so any hidden/held-out test exercising the symbol still fails; graded correctness \u2248 0 for the task. The added scripts may themselves fail with `AssertionError` (or exit non-zero) when executed, since they assert the fixed behaviour that was never implemented."
}
raw text (what the judge reads)
### Bug-fix task answered only with new test/reproduction scripts, no source edit
- **Applies when**: `task`: the task describes a defect in an existing library/application symbol (a function, method or class named in the report) and asks for it to behave correctly.
- **Pattern**: The submitted change set adds only new standalone verification/reproduction scripts (often several near-duplicates at the repository root) and contains no edit to any file inside the package source tree that implements the reported symbol. The defect itself is never touched; the deliverable is evidence-gathering code instead of a fix.
- **Detection procedure**:
  1. List every file the program creates or modifies, from the diff/file listing in the program's text; note each path and whether it is new. [reads: code]
  2. Identify from the static facts repo tree which top-level directory holds the implementation package (e.g. a `src/<package>` or `<package>/` directory) as opposed to scratch/root-level scripts, and identify from the task statement the module/class/method that must change. [reads: static facts — repo tree; task]
  3. Check whether at least one modified path lies inside that implementation directory. If every changed path is a new root-level file whose name/content is a test or reproduction script (imports the library, asserts on its behaviour, prints "BUG"/"FIX CONFIRMED"), the pattern is present. [reads: code]
- **Counter-example**: A submission that edits the implementation file under the package directory (e.g. changes the method to return the constructed object) and *additionally* adds a regression test — mixed test + source changes do not fire.
- **Discriminator**: Fires only when the intersection of {changed paths} and {paths under the implementation package directory} is empty while the task demands a behaviour change in that package; safe submissions have at least one non-test source file changed.
- **Consequence**: The reported behaviour is unchanged at runtime, so any hidden/held-out test exercising the symbol still fails; graded correctness ≈ 0 for the task. The added scripts may themselves fail with `AssertionError` (or exit non-zero) when executed, since they assert the fixed behaviour that was never implemented.
- **Evidence**: The entire diff consisted of newly added root-level files `test_*.py` asserting `channel is not None` / `isinstance(channel, DTLSChannel)`; no file under the library source directory was modified, and the submission was finalized in that state.
18Name imported only under a typing-only guard but used at runtimecodeswesmith/scrapy__scrapy.35212ec5
Applies when
code: a module imports symbols inside if TYPE_CHECKING:, if False:, or another conditional/try block and those symbols also appear in executable statements
Pattern
A symbol's only binding is inside an import guard that never executes at runtime (typing-only block, or a try: whose except swallows the failure without defining a fallback), yet the symbol is called or evaluated in normal control flow, so the first execution of that path raises NameError.
Detection procedure
  1. Collect every name bound by an import/from ... import that sits inside if TYPE_CHECKING:, if False:, or a try: block whose handler does not rebind the same name. [reads: code]
  2. Search the rest of the module for occurrences of those names in executable positions — a call name(...), an argument, a comparison, a decorator, a default value — as opposed to annotation positions only. [reads: code]
  3. Confirm no unguarded top-level or function-local import rebinds the same name before that use. [reads: code]
Counter-example
The guarded name appears only in parameter/return annotations in a module that starts with from __future__ import annotations, or in quoted string annotations — annotations are never evaluated, so no error occurs.
Discriminator
The guarded name occurs in at least one position that Python evaluates at runtime (call target, argument, condition); in the safe case every occurrence is an annotation and postponed evaluation is enabled.
Consequence
NameError: name '<symbol>' is not defined raised the first time the containing function executes, aborting whatever command or code path invokes it; any test exercising that path fails with a non-zero exit rather than an assertion mismatch.
Evidence
A reported failure of two CLI subcommands was NameError: name 'find_spec' is not defined, i.e. a stdlib helper referenced in executing code without a runtime-effective import binding it.
id 17579d72fca2 · mined from swesmith/scrapy__scrapy.35212ec5 scrapy__scrapy.35212ec5.combine_module__s93pcl8g
curation fields
{
 "predicate": "an exception will be raised",
 "code_requirement": "1. Collect every name bound by an `import`/`from ... import` that sits inside `if TYPE_CHECKING:`, `if False:`, or a `try:` block whose handler does not rebind the same name. [reads: code]",
 "prediction": "`NameError: name '<symbol>' is not defined` raised the first time the containing function executes, aborting whatever command or code path invokes it; any test exercising that path fails with a non-zero exit rather than an assertion mismatch."
}
raw text (what the judge reads)
### Name imported only under a typing-only guard but used at runtime
- **Applies when**: `code`: a module imports symbols inside `if TYPE_CHECKING:`, `if False:`, or another conditional/`try` block and those symbols also appear in executable statements
- **Pattern**: A symbol's only binding is inside an import guard that never executes at runtime (typing-only block, or a `try:` whose `except` swallows the failure without defining a fallback), yet the symbol is called or evaluated in normal control flow, so the first execution of that path raises `NameError`.
- **Detection procedure**:
  1. Collect every name bound by an `import`/`from ... import` that sits inside `if TYPE_CHECKING:`, `if False:`, or a `try:` block whose handler does not rebind the same name. [reads: code]
  2. Search the rest of the module for occurrences of those names in executable positions — a call `name(...)`, an argument, a comparison, a decorator, a default value — as opposed to annotation positions only. [reads: code]
  3. Confirm no unguarded top-level or function-local `import` rebinds the same name before that use. [reads: code]
- **Counter-example**: The guarded name appears only in parameter/return annotations in a module that starts with `from __future__ import annotations`, or in quoted string annotations — annotations are never evaluated, so no error occurs.
- **Discriminator**: The guarded name occurs in at least one position that Python evaluates at runtime (call target, argument, condition); in the safe case every occurrence is an annotation and postponed evaluation is enabled.
- **Consequence**: `NameError: name '<symbol>' is not defined` raised the first time the containing function executes, aborting whatever command or code path invokes it; any test exercising that path fails with a non-zero exit rather than an assertion mismatch.
- **Evidence**: A reported failure of two CLI subcommands was `NameError: name 'find_spec' is not defined`, i.e. a stdlib helper referenced in executing code without a runtime-effective import binding it.
18Leftover scaffolding artifacts from manual verification committed into the repocodeswesmith/scrapy__scrapy.35212ec5
Applies when
code: the change set adds files that were produced by running the very tool/command the task asks to fix or exercise (generated modules, scaffolded project directories, exported output files) rather than by editing source.
Pattern
The author verifies a fix by invoking the tool in the repository working directory and leaves the tool's generated output in the tree. The byproduct is not part of the fix, is referenced by nothing, and sits where the tool will later refuse to overwrite it or where test/lint collection will pick it up.
Detection procedure
  1. List every file the change set adds and separate them into (a) edits to library/source modules that implement the fix and (b) files whose contents are boilerplate templates — a class/config stub with placeholder names, an empty handler body, a default config file — i.e., output a generator would emit. [reads: code]
  2. For each candidate in group (b), check the repo tree in the static facts: the file/directory is absent there (so it is genuinely new), and it sits at the repository root or beside the package rather than in a tests fixture/template directory. Then check the task statement: does its name match an identifier used in the task's reproduction commands (the project/module/output name the reporter passes on the command line)? [reads: static facts — repo tree; task statement]
  3. Confirm nothing in the change set imports, opens, or otherwise references that file, and that no test added by the change asserts on it. [reads: code]
Counter-example
a program that exercises the command inside tempfile.mkdtemp() / a pytest tmp_path and cleans up (or checks in a template file under the package's templates/fixtures directory that generated code is expected to reference) — nothing generated lands at the repo root, or the added file is referenced by the fix or a test.
Discriminator
the added file is unreferenced boilerplate at the repository root that duplicates the generator's output, and (worst case) carries the exact name the task's reproduction command uses; safe code either produces such output only under a temporary directory or the added file is imported/asserted on somewhere in the change set.
Consequence
when the grader re-runs the reproduction command, the generator refuses to clobber the existing name and exits non-zero ("… already exists" / SystemExit, FileExistsError), failing the returncode/stdout assertion; independently, an unreferenced top-level module can be swept up by test collection or lint/pre-commit checks and fail them. If the leftover name differs from the one the grader uses, the effect is limited to repo pollution and does not by itself change pass/fail.
Evidence
the change set added a root-level generated stub module whose name matched the identifier in the issue's reproduction command (scrapy genspider myspider … → new myspider.py at repo root) alongside the actual fix; the graded run happened to use a different name and passed, leaving the artifact as an unreferenced, re-run-hostile byproduct.
id a2cd620eca77 · mined from swesmith/scrapy__scrapy.35212ec5 scrapy__scrapy.35212ec5.combine_module__s93pcl8g
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. List every file the change set adds and separate them into (a) edits to library/source modules that implement the fix and (b) files whose contents are boilerplate templates \u2014 a class/config stub with placeholder names, an empty handler body, a default config file \u2014 i.e., output a generator would emit. [reads: code]",
 "prediction": "when the grader re-runs the reproduction command, the generator refuses to clobber the existing name and exits non-zero (\"\u2026 already exists\" / `SystemExit`, `FileExistsError`), failing the returncode/stdout assertion; independently, an unreferenced top-level module can be swept up by test collection or lint/pre-commit checks and fail them. If the leftover name differs from the one the grader uses, the effect is limited to repo pollution and does not by itself change pass/fail."
}
raw text (what the judge reads)
### Leftover scaffolding artifacts from manual verification committed into the repo

- **Applies when**: `code`: the change set adds files that were produced by running the very tool/command the task asks to fix or exercise (generated modules, scaffolded project directories, exported output files) rather than by editing source.
- **Pattern**: The author verifies a fix by invoking the tool in the repository working directory and leaves the tool's generated output in the tree. The byproduct is not part of the fix, is referenced by nothing, and sits where the tool will later refuse to overwrite it or where test/lint collection will pick it up.
- **Detection procedure**:
  1. List every file the change set adds and separate them into (a) edits to library/source modules that implement the fix and (b) files whose contents are boilerplate templates — a class/config stub with placeholder names, an empty handler body, a default config file — i.e., output a generator would emit. [reads: code]
  2. For each candidate in group (b), check the repo tree in the static facts: the file/directory is absent there (so it is genuinely new), and it sits at the repository root or beside the package rather than in a tests fixture/template directory. Then check the task statement: does its name match an identifier used in the task's reproduction commands (the project/module/output name the reporter passes on the command line)? [reads: static facts — repo tree; task statement]
  3. Confirm nothing in the change set imports, opens, or otherwise references that file, and that no test added by the change asserts on it. [reads: code]
- **Counter-example**: a program that exercises the command inside `tempfile.mkdtemp()` / a pytest `tmp_path` and cleans up (or checks in a template file under the package's templates/fixtures directory that generated code is expected to reference) — nothing generated lands at the repo root, or the added file is referenced by the fix or a test.
- **Discriminator**: the added file is unreferenced boilerplate at the repository root that duplicates the generator's output, and (worst case) carries the exact name the task's reproduction command uses; safe code either produces such output only under a temporary directory or the added file is imported/asserted on somewhere in the change set.
- **Consequence**: when the grader re-runs the reproduction command, the generator refuses to clobber the existing name and exits non-zero ("… already exists" / `SystemExit`, `FileExistsError`), failing the returncode/stdout assertion; independently, an unreferenced top-level module can be swept up by test collection or lint/pre-commit checks and fail them. If the leftover name differs from the one the grader uses, the effect is limited to repo pollution and does not by itself change pass/fail.
- **Evidence**: the change set added a root-level generated stub module whose name matched the identifier in the issue's reproduction command (`scrapy genspider myspider …` → new `myspider.py` at repo root) alongside the actual fix; the graded run happened to use a different name and passed, leaving the artifact as an unreferenced, re-run-hostile byproduct.
18Reported undefined name is never bound anywhere in the change settaskswesmith/scrapy__scrapy.35212ec5
Applies when
task: the report quotes an error of the form NameError: name 'X' is not defined (or ImportError/AttributeError for a specific symbol) and expects the candidate to repair it.
Pattern
The program submits code that never introduces a binding for the exact symbol named in the error — no import/from ... import X, no assignment, no def/class for it — so the failing line still resolves nothing at runtime.
Detection procedure
  1. Extract the exact identifier quoted in the error message from the task statement. [reads: task]
  2. Search the candidate's presented code for any binding of that identifier: import X, from ... import X, X = ..., def X, class X, or a module-level fallback assignment. [reads: code]
  3. If no such binding appears in any presented file, and no presented file is the module where the symbol is used, the defect is unrepaired. [reads: code]
Counter-example
Code that adds from <stdlib module> import X (or defines a shim X = ... guarded by try/except ImportError) at the top of the module whose function uses X — the binding exists, so the rubric must not fire even though X also appears in the error text.
Consequence
The originally reported exception (NameError first, or ImportError/AttributeError for the same symbol) is raised again on the documented reproduction path; any test invoking that code path fails with the identical message. This accounts for the whole of the unfixed behavior when it fires.
Evidence
A report of NameError: name '<symbol>' is not defined from a CLI command, answered by a change set in which the symbol appears nowhere as an import or definition; verification showed an empty diff on the affected module and the error mechanism intact.
id 87fcc747b2a5 · mined from swesmith/scrapy__scrapy.35212ec5 scrapy__scrapy.35212ec5.combine_module__s93pcl8g
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Extract the exact identifier quoted in the error message from the task statement. [reads: task]",
 "prediction": "The originally reported exception (`NameError` first, or `ImportError`/`AttributeError` for the same symbol) is raised again on the documented reproduction path; any test invoking that code path fails with the identical message. This accounts for the whole of the unfixed behavior when it fires."
}
raw text (what the judge reads)
### Reported undefined name is never bound anywhere in the change set
- **Applies when**: `task`: the report quotes an error of the form `NameError: name 'X' is not defined` (or `ImportError`/`AttributeError` for a specific symbol) and expects the candidate to repair it.
- **Pattern**: The program submits code that never introduces a binding for the exact symbol named in the error — no `import`/`from ... import X`, no assignment, no `def`/`class` for it — so the failing line still resolves nothing at runtime.
- **Detection procedure**:
  1. Extract the exact identifier quoted in the error message from the task statement. [reads: task]
  2. Search the candidate's presented code for any binding of that identifier: `import X`, `from ... import X`, `X = ...`, `def X`, `class X`, or a module-level fallback assignment. [reads: code]
  3. If no such binding appears in any presented file, and no presented file is the module where the symbol is used, the defect is unrepaired. [reads: code]
- **Counter-example**: Code that adds `from <stdlib module> import X` (or defines a shim `X = ...` guarded by `try/except ImportError`) at the top of the module whose function uses `X` — the binding exists, so the rubric must not fire even though `X` also appears in the error text.
- **Consequence**: The originally reported exception (`NameError` first, or `ImportError`/`AttributeError` for the same symbol) is raised again on the documented reproduction path; any test invoking that code path fails with the identical message. This accounts for the whole of the unfixed behavior when it fires.
- **Evidence**: A report of `NameError: name '<symbol>' is not defined` from a CLI command, answered by a change set in which the symbol appears nowhere as an import or definition; verification showed an empty diff on the affected module and the error mechanism intact.
19Read-only inspection script submitted for a task that requires editing sourcetaskswesmith/pallets__click.fde47b4b
Applies when
task: the task asks for a repository change (refactor, reorder, fix, rename, add behavior) and the candidate is a script or patch that runs against the repo
Pattern
The program only inspects the codebase — it opens files in read mode and prints findings — and never writes, rewrites, or patches any file, so the repository is byte-identical after it runs and the requested change is never made.
Detection procedure
  1. Read the task statement and record the concrete change it demands to repository files (e.g. "methods should be in order X", "function must return Y"). [reads: task]
  2. Scan the program for any mutation of repository state: open(..., 'w'/'a'/'r+'), pathlib.Path.write_text, shutil.move, os.replace, in-place fileinput, or invocation of a patch/format tool. [reads: code]
  3. Fire if every file access is read-only (open(path, 'r'), read_text, readlines) and the program's only outputs are print/logging of what it found. [reads: code]
Counter-example
A program that first reads the file to locate a region, then reassembles the content and writes it back with open(path, 'w').write(new_src) (or emits a diff that a downstream step applies) — reading is a prelude to writing, not the whole program.
Discriminator
The failing case contains no write path at all; the safe case contains at least one file-mutating call whose content derives from the analysis.
Consequence
The required change is absent; any grader assertion about the new file contents/structure fails while the pre-existing test suite still passes vacuously (passing tests are not evidence of success here). Score for the task requirement is 0.
Evidence
A script did with open('src/.../core.py','r') as f: lines = f.readlines() and then only print(...)ed the discovered method names; the full suite reported "612 passed" purely because nothing in the repository was modified.
id cb470991ed7c · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Read the task statement and record the concrete change it demands to repository files (e.g. \"methods should be in order X\", \"function must return Y\"). [reads: task]",
 "prediction": "The required change is absent; any grader assertion about the new file contents/structure fails while the pre-existing test suite still passes vacuously (passing tests are not evidence of success here). Score for the task requirement is 0."
}
raw text (what the judge reads)
### Read-only inspection script submitted for a task that requires editing source
- **Applies when**: `task`: the task asks for a repository change (refactor, reorder, fix, rename, add behavior) and the candidate is a script or patch that runs against the repo
- **Pattern**: The program only inspects the codebase — it opens files in read mode and prints findings — and never writes, rewrites, or patches any file, so the repository is byte-identical after it runs and the requested change is never made.
- **Detection procedure**:
  1. Read the task statement and record the concrete change it demands to repository files (e.g. "methods should be in order X", "function must return Y"). [reads: task]
  2. Scan the program for any mutation of repository state: `open(..., 'w'/'a'/'r+')`, `pathlib.Path.write_text`, `shutil.move`, `os.replace`, in-place `fileinput`, or invocation of a patch/format tool. [reads: code]
  3. Fire if every file access is read-only (`open(path, 'r')`, `read_text`, `readlines`) and the program's only outputs are `print`/logging of what it found. [reads: code]
- **Counter-example**: A program that first reads the file to locate a region, then reassembles the content and writes it back with `open(path, 'w').write(new_src)` (or emits a diff that a downstream step applies) — reading is a prelude to writing, not the whole program.
- **Discriminator**: The failing case contains no write path at all; the safe case contains at least one file-mutating call whose content derives from the analysis.
- **Consequence**: The required change is absent; any grader assertion about the new file contents/structure fails while the pre-existing test suite still passes vacuously (passing tests are not evidence of success here). Score for the task requirement is 0.
- **Evidence**: A script did `with open('src/.../core.py','r') as f: lines = f.readlines()` and then only `print(...)`ed the discovered method names; the full suite reported "612 passed" purely because nothing in the repository was modified.
19Operating on a different named entity than the task specifiestaskswesmith/pallets__click.fde47b4b
Applies when
task: the task names a specific class, function, table, column, or file to change, and code: the program locates its target by matching a literal name string
Pattern
The program hard-codes a target identifier that does not match the identifier named in the task (a sibling class, a related-but-different symbol), so all its work is applied to the wrong entity.
Detection procedure
  1. Extract from the task statement the exact name(s) of the entity to be changed. [reads: task]
  2. Find the literal strings/regexes the program uses to select its target (e.g. if 'class X(...)' in line, df['colname'], a filename literal). [reads: code]
  3. Fire if the selector literal names a different symbol than the task does, and no other code path handles the task's named entity. [reads: code]
Counter-example
A program whose selector literal differs textually but resolves to the task's entity (e.g. matching a base class the task's class inherits from, or iterating all candidates and filtering by the task-named one later in the file).
Discriminator
In the failing case the task-named identifier appears nowhere in the program text; in the safe case it appears, or the selector provably covers it.
Consequence
Zero effect on the requirement; any check targeting the named entity fails, and if the program also writes, it corrupts an unrelated region. Terminal behavior is usually silent (empty or irrelevant output) rather than an exception.
Evidence
The task described reordering methods of one class, while the script matched 'class Group(MultiCommand):' and reported methods of that other class only.
id 826e58eeea6f · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Extract from the task statement the exact name(s) of the entity to be changed. [reads: task]",
 "prediction": "Zero effect on the requirement; any check targeting the named entity fails, and if the program also writes, it corrupts an unrelated region. Terminal behavior is usually silent (empty or irrelevant output) rather than an exception."
}
raw text (what the judge reads)
### Operating on a different named entity than the task specifies
- **Applies when**: `task`: the task names a specific class, function, table, column, or file to change, and `code`: the program locates its target by matching a literal name string
- **Pattern**: The program hard-codes a target identifier that does not match the identifier named in the task (a sibling class, a related-but-different symbol), so all its work is applied to the wrong entity.
- **Detection procedure**:
  1. Extract from the task statement the exact name(s) of the entity to be changed. [reads: task]
  2. Find the literal strings/regexes the program uses to select its target (e.g. `if 'class X(...)' in line`, `df['colname']`, a filename literal). [reads: code]
  3. Fire if the selector literal names a different symbol than the task does, and no other code path handles the task's named entity. [reads: code]
- **Counter-example**: A program whose selector literal differs textually but resolves to the task's entity (e.g. matching a base class the task's class inherits from, or iterating all candidates and filtering by the task-named one later in the file).
- **Discriminator**: In the failing case the task-named identifier appears nowhere in the program text; in the safe case it appears, or the selector provably covers it.
- **Consequence**: Zero effect on the requirement; any check targeting the named entity fails, and if the program also writes, it corrupts an unrelated region. Terminal behavior is usually silent (empty or irrelevant output) rather than an exception.
- **Evidence**: The task described reordering methods of one class, while the script matched `'class Group(MultiCommand):'` and reported methods of that other class only.
19Magic absolute line numbers used to delimit a region of a source filecodeswesmith/pallets__click.fde47b4b
Applies when
code: the program parses a text/source file line by line to find or bound a region
Pattern
The scan boundary is a bare integer line-number literal compared against the loop counter, instead of a structural condition derived from the file's content; the number encodes an assumption about the file that is never verified and breaks the moment the file differs from the author's snapshot.
Detection procedure
  1. Locate the loop that enumerates lines of a file and the conditions that start/stop region tracking. [reads: code]
  2. Check whether any stop/start condition compares the line index to a numeric literal (e.g. if ... and i > 1644: break, lines[120:400]). [reads: code]
  3. Fire if such a literal exists and the program contains no assertion, search, or fallback that validates the literal against the file's actual content. [reads: code]
Counter-example
A scan whose boundaries are purely structural — start on a matched class/section header, stop on the next line at the same indentation or the next header match — with numeric constants used only for indentation widths, not absolute positions.
Discriminator
The literal is an absolute position in the file (compared to a line counter or used as a slice bound); indentation/column constants are not the failing case.
Consequence
Silently wrong region selection — the program prints an empty or truncated result, or, if it writes, edits the wrong span of the file; no exception is raised, so the failure is invisible in logs and the task requirement goes unmet.
Evidence
A file-scanning script bounded its region with if in_group_class and line.startswith('class ') and i > 1644: break, a hard-coded offset with no verification that the class actually begins before that line.
id 2cb5361aadcd · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the loop that enumerates lines of a file and the conditions that start/stop region tracking. [reads: code]",
 "prediction": "Silently wrong region selection \u2014 the program prints an empty or truncated result, or, if it writes, edits the wrong span of the file; no exception is raised, so the failure is invisible in logs and the task requirement goes unmet."
}
raw text (what the judge reads)
### Magic absolute line numbers used to delimit a region of a source file
- **Applies when**: `code`: the program parses a text/source file line by line to find or bound a region
- **Pattern**: The scan boundary is a bare integer line-number literal compared against the loop counter, instead of a structural condition derived from the file's content; the number encodes an assumption about the file that is never verified and breaks the moment the file differs from the author's snapshot.
- **Detection procedure**:
  1. Locate the loop that enumerates lines of a file and the conditions that start/stop region tracking. [reads: code]
  2. Check whether any stop/start condition compares the line index to a numeric literal (e.g. `if ... and i > 1644: break`, `lines[120:400]`). [reads: code]
  3. Fire if such a literal exists and the program contains no assertion, search, or fallback that validates the literal against the file's actual content. [reads: code]
- **Counter-example**: A scan whose boundaries are purely structural — start on a matched `class`/section header, stop on the next line at the same indentation or the next header match — with numeric constants used only for indentation widths, not absolute positions.
- **Discriminator**: The literal is an absolute position in the file (compared to a line counter or used as a slice bound); indentation/column constants are not the failing case.
- **Consequence**: Silently wrong region selection — the program prints an empty or truncated result, or, if it writes, edits the wrong span of the file; no exception is raised, so the failure is invisible in logs and the task requirement goes unmet.
- **Evidence**: A file-scanning script bounded its region with `if in_group_class and line.startswith('class ') and i > 1644: break`, a hard-coded offset with no verification that the class actually begins before that line.
19Using `dir()` (or another sorted/unordered API) to determine source declaration ordercodeswesmith/pallets__click.fde47b4b
Applies when
code: the program reasons about the order in which class members, functions, columns, or keys are declared/laid out in a file
Pattern
Ordering is read from an API that does not preserve declaration order — dir() returns names sorted alphabetically, set/frozenset iteration is arbitrary — so every conclusion drawn about "current order" is an artifact of the API, not of the file.
Detection procedure
  1. Locate where the program obtains the sequence it treats as an order: dir(obj), iteration over a set, sorted(...) output, or help() text [reads: code]
  2. Read the task statement and confirm the requirement is about order as written in the source (method/definition/section order), not about alphabetical or arbitrary order [reads: task]
  3. Check that the program never obtains order from an order-preserving source: cls.__dict__/vars(cls) keys, inspect.getsource + line numbers, inspect.getsourcelines, ast.parse body traversal, or reading the file's lines directly [reads: code]
Counter-example
A program that calls dir(cls) only to enumerate which members exist and then sorts them by inspect.getsourcelines(getattr(cls, n))[1], or that parses the file with ast — the order finally used comes from the source, not from dir().
Discriminator
Goes wrong when the sequence returned by the unordered/sorted API is itself consumed as the answer (printed, indexed, compared to an expected order); safe when that sequence is only a membership list and a source-derived key supplies the ordering.
Consequence
The program reports alphabetical order as if it were file order, so any diff/decision it drives is wrong; a subsequent reordering edit is either omitted or applied to the wrong positions, leaving order-sensitive checks failing. No exception is raised, which makes the wrong result easy to miss.
Evidence
method_names = [name for name in dir(Cls) if ...] printed with positional indices and treated as "definition order"; the run produced an alphabetized listing and no correction to the file.
id 030261c55727 · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate where the program obtains the sequence it treats as an order: `dir(obj)`, iteration over a `set`, `sorted(...)` output, or `help()` text [reads: code]",
 "prediction": "The program reports alphabetical order as if it were file order, so any diff/decision it drives is wrong; a subsequent reordering edit is either omitted or applied to the wrong positions, leaving order-sensitive checks failing. No exception is raised, which makes the wrong result easy to miss."
}
raw text (what the judge reads)
### Using `dir()` (or another sorted/unordered API) to determine source declaration order
- **Applies when**: `code`: the program reasons about the order in which class members, functions, columns, or keys are declared/laid out in a file
- **Pattern**: Ordering is read from an API that does not preserve declaration order — `dir()` returns names sorted alphabetically, `set`/`frozenset` iteration is arbitrary — so every conclusion drawn about "current order" is an artifact of the API, not of the file.
- **Detection procedure**:
  1. Locate where the program obtains the sequence it treats as an order: `dir(obj)`, iteration over a `set`, `sorted(...)` output, or `help()` text [reads: code]
  2. Read the task statement and confirm the requirement is about order *as written in the source* (method/definition/section order), not about alphabetical or arbitrary order [reads: task]
  3. Check that the program never obtains order from an order-preserving source: `cls.__dict__`/`vars(cls)` keys, `inspect.getsource` + line numbers, `inspect.getsourcelines`, `ast.parse` body traversal, or reading the file's lines directly [reads: code]
- **Counter-example**: A program that calls `dir(cls)` only to enumerate *which* members exist and then sorts them by `inspect.getsourcelines(getattr(cls, n))[1]`, or that parses the file with `ast` — the order finally used comes from the source, not from `dir()`.
- **Discriminator**: Goes wrong when the sequence returned by the unordered/sorted API is itself consumed as the answer (printed, indexed, compared to an expected order); safe when that sequence is only a membership list and a source-derived key supplies the ordering.
- **Consequence**: The program reports alphabetical order as if it were file order, so any diff/decision it drives is wrong; a subsequent reordering edit is either omitted or applied to the wrong positions, leaving order-sensitive checks failing. No exception is raised, which makes the wrong result easy to miss.
- **Evidence**: `method_names = [name for name in dir(Cls) if ...]` printed with positional indices and treated as "definition order"; the run produced an alphabetized listing and no correction to the file.
19Speculative relocation of code that already satisfied the stated requirementtaskswesmith/pallets__click.fde47b4b
Applies when
task: the request is about structure/ordering/style of existing source (e.g. "methods are in the wrong order", "imports jumbled", "sections out of place") rather than a runtime failure
Pattern
The program takes a vague structural complaint at face value and moves a definition to a new position, even though the concrete symptoms the report names cannot be observed in the source; the move itself introduces the very disorder that was being reported.
Detection procedure
  1. From the task statement, extract every concrete relation it asserts is broken (e.g. "A and B appear before __init__", "C is after D"). [reads: task]
  2. In the submitted source, locate the enclosing class/module and record the textual order of every symbol named in step 1; check whether each asserted relation actually holds as described. [reads: code]
  3. Look at the program's change (diff or clearly relocated block). It fires if none of the symbols named in the report were the ones relocated, or if the relations the report complains about were already satisfied before the change — i.e. the program relocated a different, unmentioned definition on its own judgement, with no comment, docstring, changelog entry or test cited as the authority for the new position. [reads: code]
Counter-example
A program that moves exactly the symbols the report names, so that after the edit each relation the report says is violated (__init__ first, then the core methods it lists) demonstrably holds in the file; or a program that leaves the ordering untouched because the asserted relations already hold.
Discriminator
The changed symbol is absent from the list of symbols the report names, and the relations the report names hold both before and after the edit — the edit is unanchored to any statement in the task. A safe edit is traceable line-by-line to a relation named in the task.
Consequence
A grader/test that compares the class's method order against a canonical reference fails on the moved symbol; the repository ends in a state that is structurally further from the reference than the starting state (a regression introduced by the "fix"). This accounts for essentially the whole failure to satisfy the request; the remainder is the absence of any check that would have caught it.
Evidence
A diff that deleted a format_* method from its original slot and re-inserted it several methods later, while the symbols the issue text explicitly named as misplaced were never touched and already sat in the positions the issue said they should occupy; submitted as final.
id 9b649ce8a38e · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. From the task statement, extract every concrete relation it asserts is broken (e.g. \"`A` and `B` appear before `__init__`\", \"`C` is after `D`\"). [reads: task]",
 "prediction": "A grader/test that compares the class's method order against a canonical reference fails on the moved symbol; the repository ends in a state that is structurally *further* from the reference than the starting state (a regression introduced by the \"fix\"). This accounts for essentially the whole failure to satisfy the request; the remainder is the absence of any check that would have caught it."
}
raw text (what the judge reads)
### Speculative relocation of code that already satisfied the stated requirement
- **Applies when**: `task`: the request is about structure/ordering/style of existing source (e.g. "methods are in the wrong order", "imports jumbled", "sections out of place") rather than a runtime failure
- **Pattern**: The program takes a vague structural complaint at face value and moves a definition to a new position, even though the concrete symptoms the report names cannot be observed in the source; the move itself introduces the very disorder that was being reported.
- **Detection procedure**:
  1. From the task statement, extract every concrete relation it asserts is broken (e.g. "`A` and `B` appear before `__init__`", "`C` is after `D`"). [reads: task]
  2. In the submitted source, locate the enclosing class/module and record the textual order of every symbol named in step 1; check whether each asserted relation actually holds as described. [reads: code]
  3. Look at the program's change (diff or clearly relocated block). It fires if none of the symbols named in the report were the ones relocated, or if the relations the report complains about were already satisfied before the change — i.e. the program relocated a *different*, unmentioned definition on its own judgement, with no comment, docstring, changelog entry or test cited as the authority for the new position. [reads: code]
- **Counter-example**: A program that moves exactly the symbols the report names, so that after the edit each relation the report says is violated (`__init__` first, then the core methods it lists) demonstrably holds in the file; or a program that leaves the ordering untouched because the asserted relations already hold.
- **Discriminator**: The changed symbol is absent from the list of symbols the report names, and the relations the report names hold both before and after the edit — the edit is unanchored to any statement in the task. A safe edit is traceable line-by-line to a relation named in the task.
- **Consequence**: A grader/test that compares the class's method order against a canonical reference fails on the moved symbol; the repository ends in a state that is structurally *further* from the reference than the starting state (a regression introduced by the "fix"). This accounts for essentially the whole failure to satisfy the request; the remainder is the absence of any check that would have caught it.
- **Evidence**: A diff that deleted a `format_*` method from its original slot and re-inserted it several methods later, while the symbols the issue text explicitly named as misplaced were never touched and already sat in the positions the issue said they should occupy; submitted as final.
19Reproduction snippet re-added as a "test" that cannot failcodeswesmith/pallets__click.fde47b4b
Applies when
code: the program adds a new top-level script/test file that is not part of the repository's existing test files listed in the static facts
Pattern
The only verification artifact the program produces is a verbatim copy of the report's reproduction snippet, containing no assertion and no inspection of the property actually under repair, so it passes identically before and after the change and provides zero evidence the requirement was met.
Detection procedure
  1. Locate files created by the program that are not among the test modules listed in the repo tree in the static facts. [reads: static facts — repo tree listing of the tests directory]
  2. Read the new file: check for assert, unittest assertions, pytest.raises, or any comparison of an observed value/structure against an expected one. [reads: code]
  3. It fires if the file contains none of these and merely imports the package and calls the public entry point, while the task's requirement is a property (ordering, structure, formatting, an internal invariant) that executing that snippet does not observe. [reads: task and code]
Counter-example
A new script that programmatically inspects the property at issue — e.g. uses inspect.getsource/__dict__ ordering, or reads the file and compares the sequence of definitions to an expected list — and raises or asserts on mismatch; or a new file added under the existing test package with real assertions.
Consequence
The program declares success on an unverified edit; predict the hidden/graded checks for the requested property to fail, and predict a stray unasserted module left at repository root that the test runner collects with zero tests. This explains why the wrong edit went undetected rather than the wrongness itself — roughly the secondary share of the observed outcome.
Evidence
A newly added root-level test_*.py that was a byte-for-byte copy of the issue's reproduction snippet with a if __name__ == "__main__": call and no assertion, submitted as the fix's validation.
id b1a500b579b5 · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate files created by the program that are not among the test modules listed in the repo tree in the static facts. [reads: static facts \u2014 repo tree listing of the tests directory]",
 "prediction": "The program declares success on an unverified edit; predict the hidden/graded checks for the requested property to fail, and predict a stray unasserted module left at repository root that the test runner collects with zero tests. This explains why the wrong edit went undetected rather than the wrongness itself \u2014 roughly the secondary share of the observed outcome."
}
raw text (what the judge reads)
### Reproduction snippet re-added as a "test" that cannot fail
- **Applies when**: `code`: the program adds a new top-level script/test file that is not part of the repository's existing test files listed in the static facts
- **Pattern**: The only verification artifact the program produces is a verbatim copy of the report's reproduction snippet, containing no assertion and no inspection of the property actually under repair, so it passes identically before and after the change and provides zero evidence the requirement was met.
- **Detection procedure**:
  1. Locate files created by the program that are not among the test modules listed in the repo tree in the static facts. [reads: static facts — repo tree listing of the tests directory]
  2. Read the new file: check for `assert`, `unittest` assertions, `pytest.raises`, or any comparison of an observed value/structure against an expected one. [reads: code]
  3. It fires if the file contains none of these and merely imports the package and calls the public entry point, while the task's requirement is a property (ordering, structure, formatting, an internal invariant) that executing that snippet does not observe. [reads: task and code]
- **Counter-example**: A new script that programmatically inspects the property at issue — e.g. uses `inspect.getsource`/`__dict__` ordering, or reads the file and compares the sequence of definitions to an expected list — and raises or asserts on mismatch; or a new file added under the existing test package with real assertions.
- **Consequence**: The program declares success on an unverified edit; predict the hidden/graded checks for the requested property to fail, and predict a stray unasserted module left at repository root that the test runner collects with zero tests. This explains why the wrong edit went undetected rather than the wrongness itself — roughly the secondary share of the observed outcome.
- **Evidence**: A newly added root-level `test_*.py` that was a byte-for-byte copy of the issue's reproduction snippet with a `if __name__ == "__main__":` call and no assertion, submitted as the fix's validation.
19Partial fix that never touches the entities the issue namestaskswesmith/pallets__click.fde47b4b
Applies when
task: the task statement enumerates specific code entities (functions, methods, classes, attributes, fields) that are wrong, or describes a whole-container invariant (ordering, completeness, consistency) rather than a single-site bug
Pattern
The program applies one small, localized edit and stops, while the entities explicitly named as broken in the task are left untouched and the container-wide invariant the task asks for is still violated. The edit relocates/renames something adjacent instead of establishing the stated end state.
Detection procedure
  1. From the task statement, write down every code entity named as wrong or as needing a particular position/value, plus the invariant phrased as "the expected X should be ..." [reads: task]
  2. In the program's source/diff, list every entity actually added, moved, deleted, or modified. [reads: code]
  3. Fire if the intersection of the two lists is empty (or covers only one of several named entities) and the total change is a single relocated/edited block, with nothing in the code establishing the stated invariant over the rest of the container. [reads: code]
Counter-example
A program that rewrites the whole container (reorders/normalizes every member of the class or module, or fixes every enumerated entity), where the task-named entities all appear among the modified ones even though extra unnamed entities were also moved.
Discriminator
The wrong case modifies only entities the task never mentions and leaves every task-named entity in its original position/state; the safe case modifies a superset that includes all task-named entities.
Consequence
Hidden or graded checks that assert the full invariant (e.g. compare the ordered list of members against an expected sequence, or check each named entity) fail; the run produces no exception, so the defect is silent and the submission scores 0 on the targeted assertion while unrelated tests still pass. This accounts for most of the gap versus a solution that restores the whole ordering.
Evidence
A diff that moved one method (format_usage) to a different position inside a class while the issue explicitly named other members (format_options, invoke, main, __init__) as misplaced; the submitted solution left those untouched.
id d2ffa470bc5c · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. From the task statement, write down every code entity named as wrong or as needing a particular position/value, plus the invariant phrased as \"the expected X should be ...\" [reads: task]",
 "prediction": "Hidden or graded checks that assert the full invariant (e.g. compare the ordered list of members against an expected sequence, or check each named entity) fail; the run produces no exception, so the defect is silent and the submission scores 0 on the targeted assertion while unrelated tests still pass. This accounts for most of the gap versus a solution that restores the whole ordering."
}
raw text (what the judge reads)
### Partial fix that never touches the entities the issue names
- **Applies when**: `task`: the task statement enumerates specific code entities (functions, methods, classes, attributes, fields) that are wrong, or describes a whole-container invariant (ordering, completeness, consistency) rather than a single-site bug
- **Pattern**: The program applies one small, localized edit and stops, while the entities explicitly named as broken in the task are left untouched and the container-wide invariant the task asks for is still violated. The edit relocates/renames something adjacent instead of establishing the stated end state.
- **Detection procedure**:
  1. From the task statement, write down every code entity named as wrong or as needing a particular position/value, plus the invariant phrased as "the expected X should be ..." [reads: task]
  2. In the program's source/diff, list every entity actually added, moved, deleted, or modified. [reads: code]
  3. Fire if the intersection of the two lists is empty (or covers only one of several named entities) and the total change is a single relocated/edited block, with nothing in the code establishing the stated invariant over the rest of the container. [reads: code]
- **Counter-example**: A program that rewrites the whole container (reorders/normalizes every member of the class or module, or fixes every enumerated entity), where the task-named entities all appear among the modified ones even though extra unnamed entities were also moved.
- **Discriminator**: The wrong case modifies only entities the task never mentions and leaves every task-named entity in its original position/state; the safe case modifies a superset that includes all task-named entities.
- **Consequence**: Hidden or graded checks that assert the full invariant (e.g. compare the ordered list of members against an expected sequence, or check each named entity) fail; the run produces no exception, so the defect is silent and the submission scores 0 on the targeted assertion while unrelated tests still pass. This accounts for most of the gap versus a solution that restores the whole ordering.
- **Evidence**: A diff that moved one method (`format_usage`) to a different position inside a class while the issue explicitly named other members (`format_options`, `invoke`, `main`, `__init__`) as misplaced; the submitted solution left those untouched.
19Verification script whose outcome is independent of the changecodeswesmith/pallets__click.fde47b4b
Applies when
code: the program adds a new standalone script or test file whose stated purpose is to confirm the fix
Pattern
The added verification only exercises ordinary runtime behavior (imports the package and calls a public entry point, prints output) that is unaffected by the property the task is about — e.g. source layout, ordering, formatting, metadata — and contains no assertion at all. It therefore succeeds identically before and after the edit and provides zero evidence the defect is gone.
Detection procedure
  1. Locate any file the program creates whose name or content marks it as a check/repro (test_.py, check_.py, if __name__ == "__main__": script). [reads: code]
  2. Read the task to determine what observable property must change for the fix to be correct (a returned value, a message, an ordering, a file's contents). [reads: task]
  3. Fire if that file contains no assert, no def test_* collected by the test runner listed in the environment packages, and no comparison against an expected value — i.e. it merely re-runs the snippet quoted in the issue and exits. [reads: code, static facts — installed packages include a test runner]
Counter-example
A file that imports the modified module and asserts the expected property (e.g. compares an extracted list/value against an expected literal, or defines def test_... with an assert), so it fails if the edit is wrong or absent.
Discriminator
In the wrong case the script's exit status is the same whether or not the source change was applied (no assertion, nothing reads the changed property); in the safe case removing the change makes the script fail.
Consequence
The defect is submitted unverified — predict that the real grading tests fail while the program reports success; also leaves an uncollected/no-op file at repo root that a pytest run will import and collect as an empty test module. Explains the failure to detect the incomplete fix rather than the incompleteness itself.
Evidence
A newly added test_fix.py containing only the issue's reproduction snippet (@click.command() … echo(...)) with no assertion, submitted as the confirmation that a source-ordering defect had been repaired.
id 7f54c7b0fa56 · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate any file the program creates whose name or content marks it as a check/repro (`test_*.py`, `check_*.py`, `if __name__ == \"__main__\":` script). [reads: code]",
 "prediction": "The defect is submitted unverified \u2014 predict that the real grading tests fail while the program reports success; also leaves an uncollected/no-op file at repo root that a `pytest` run will import and collect as an empty test module. Explains the failure to detect the incomplete fix rather than the incompleteness itself."
}
raw text (what the judge reads)
### Verification script whose outcome is independent of the change
- **Applies when**: `code`: the program adds a new standalone script or test file whose stated purpose is to confirm the fix
- **Pattern**: The added verification only exercises ordinary runtime behavior (imports the package and calls a public entry point, prints output) that is unaffected by the property the task is about — e.g. source layout, ordering, formatting, metadata — and contains no assertion at all. It therefore succeeds identically before and after the edit and provides zero evidence the defect is gone.
- **Detection procedure**:
  1. Locate any file the program creates whose name or content marks it as a check/repro (`test_*.py`, `check_*.py`, `if __name__ == "__main__":` script). [reads: code]
  2. Read the task to determine what observable property must change for the fix to be correct (a returned value, a message, an ordering, a file's contents). [reads: task]
  3. Fire if that file contains no `assert`, no `def test_*` collected by the test runner listed in the environment packages, and no comparison against an expected value — i.e. it merely re-runs the snippet quoted in the issue and exits. [reads: code, static facts — installed packages include a test runner]
- **Counter-example**: A file that imports the modified module and asserts the expected property (e.g. compares an extracted list/value against an expected literal, or defines `def test_...` with an `assert`), so it fails if the edit is wrong or absent.
- **Discriminator**: In the wrong case the script's exit status is the same whether or not the source change was applied (no assertion, nothing reads the changed property); in the safe case removing the change makes the script fail.
- **Consequence**: The defect is submitted unverified — predict that the real grading tests fail while the program reports success; also leaves an uncollected/no-op file at repo root that a `pytest` run will import and collect as an empty test module. Explains the failure to detect the incomplete fix rather than the incompleteness itself.
- **Evidence**: A newly added `test_fix.py` containing only the issue's reproduction snippet (`@click.command()` … `echo(...)`) with no assertion, submitted as the confirmation that a source-ordering defect had been repaired.
19Self-check that cannot observe the required propertycodeswesmith/pallets__click.fde47b4b
Applies when
code: the change set adds a standalone script (or __main__ block) whose purpose is to demonstrate the fix, and task: the requirement concerns static source structure, layout, ordering, typing, or documentation rather than runtime output
Pattern
The program validates its edit by running the library and observing that nothing crashes, while the property the task demands is invisible at runtime; the passing smoke run creates false confidence and the actual requirement is never checked.
Detection procedure
  1. Read the task statement and classify the required end-state: does satisfying it change any observable runtime value/exception/output, or only the arrangement/annotation/text of the source? [reads: task]
  2. Locate any new file or block the program added purely for verification (imports the package, calls an entry point, prints or echoes) [reads: code]
  3. Fire when the requirement from step 1 is source-structural and the added verification only executes the package, containing no assertion over source text, AST, inspect.getsource, ordering, or symbol positions [reads: code]
Counter-example
A program that adds a check reading the artefact the requirement is about — e.g. asserting over inspect.getsource/AST node order, parsing the modified file, or a unit test asserting the newly required behaviour — even if it also adds a smoke script.
Discriminator
The verification's observable surface (stdout, return value, absence of exception) is unchanged by whether the required edit was made correctly, completely, or at all.
Consequence
The agent stops after a minimal edit that the check cannot distinguish from the correct one, leaving the requirement unmet; additionally the unrequested script remains in the diff as an artefact. Explains why the incomplete edit was accepted rather than the size of the remaining gap — the missing edits themselves account for the score difference.
Evidence
The change set added a top-level reproduction script that merely invoked the package's entry point; the task's requirement was about the order of definitions in a source file, which that script cannot observe, and the submitted edit left the stated ordering violated.
id 77cd18bd3e40 · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_pm_class_shuffle_funcs__pix64b6s
curation fields
{
 "predicate": "a named unit test will fail",
 "code_requirement": "1. Read the task statement and classify the required end-state: does satisfying it change any observable runtime value/exception/output, or only the arrangement/annotation/text of the source? [reads: task]",
 "prediction": "The agent stops after a minimal edit that the check cannot distinguish from the correct one, leaving the requirement unmet; additionally the unrequested script remains in the diff as an artefact. Explains why the incomplete edit was accepted rather than the size of the remaining gap \u2014 the missing edits themselves account for the score difference."
}
raw text (what the judge reads)
### Self-check that cannot observe the required property
- **Applies when**: `code`: the change set adds a standalone script (or `__main__` block) whose purpose is to demonstrate the fix, and `task`: the requirement concerns static source structure, layout, ordering, typing, or documentation rather than runtime output
- **Pattern**: The program validates its edit by running the library and observing that nothing crashes, while the property the task demands is invisible at runtime; the passing smoke run creates false confidence and the actual requirement is never checked.
- **Detection procedure**:
  1. Read the task statement and classify the required end-state: does satisfying it change any observable runtime value/exception/output, or only the arrangement/annotation/text of the source? [reads: task]
  2. Locate any new file or block the program added purely for verification (imports the package, calls an entry point, prints or echoes) [reads: code]
  3. Fire when the requirement from step 1 is source-structural and the added verification only executes the package, containing no assertion over source text, AST, `inspect.getsource`, ordering, or symbol positions [reads: code]
- **Counter-example**: A program that adds a check reading the artefact the requirement is about — e.g. asserting over `inspect.getsource`/AST node order, parsing the modified file, or a unit test asserting the newly required behaviour — even if it also adds a smoke script.
- **Discriminator**: The verification's observable surface (stdout, return value, absence of exception) is unchanged by whether the required edit was made correctly, completely, or at all.
- **Consequence**: The agent stops after a minimal edit that the check cannot distinguish from the correct one, leaving the requirement unmet; additionally the unrequested script remains in the diff as an artefact. Explains why the incomplete edit was accepted rather than the size of the remaining gap — the missing edits themselves account for the score difference.
- **Evidence**: The change set added a top-level reproduction script that merely invoked the package's entry point; the task's requirement was about the order of definitions in a source file, which that script cannot observe, and the submitted edit left the stated ordering violated.
20Bug-report task closed with only a reproduction script, no source edittaskswesmith/mahmoud__glom.fb3c4e76
Applies when
task: the statement reports incorrect runtime behavior of an existing function/class in the repository (wrong order, inverted condition, spurious exception) and asks for it to be fixed
Pattern
The change set adds a standalone script that demonstrates the reported behavior but never modifies the module that implements it, so the defect is still present after the change.
Detection procedure
  1. From the task statement, note the symbol(s) named as misbehaving and the import path they are imported from. [reads: task]
  2. In the static facts repo tree, locate the package directory that owns that import path (the directory whose name matches the top-level import, containing the implementation modules). [reads: static facts — repo tree]
  3. Read the list of files added/changed by the program. Fire if every added/changed file lies outside that package directory (e.g. a new top-level .py script or notebook) and no implementation module inside it is edited. [reads: code]
Counter-example
A change set that edits the implementation module (reordering the loop, flipping the comparison, relaxing the argument validation) and additionally drops in a reproduction script — the script is present, but a package module is also modified.
Discriminator
Whether at least one file inside the implementing package directory appears in the change set. Reproduction-only change sets touch zero of them.
Consequence
The reported defect persists; any hidden/regression test that asserts the corrected behavior of the named symbol fails (AssertionError, or the very exception class the task says is raised spuriously, e.g. ValueError). Task requirement "fix the behavior" is unmet regardless of what the added script prints.
Evidence
The entire diff was test_coalesce_issue.py (new file) with no edit under the package source directory, while the task described an inverted iteration order, inverted skip predicate, and a spurious ValueError inside that package's module.
id 6e8ba9e109e3 · mined from swesmith/mahmoud__glom.fb3c4e76 mahmoud__glom.fb3c4e76.func_basic__4c5ut8n6
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. From the task statement, note the symbol(s) named as misbehaving and the import path they are imported from. [reads: task]",
 "prediction": "The reported defect persists; any hidden/regression test that asserts the corrected behavior of the named symbol fails (AssertionError, or the very exception class the task says is raised spuriously, e.g. `ValueError`). Task requirement \"fix the behavior\" is unmet regardless of what the added script prints."
}
raw text (what the judge reads)
### Bug-report task closed with only a reproduction script, no source edit
- **Applies when**: `task`: the statement reports incorrect runtime behavior of an existing function/class in the repository (wrong order, inverted condition, spurious exception) and asks for it to be fixed
- **Pattern**: The change set adds a standalone script that *demonstrates* the reported behavior but never modifies the module that implements it, so the defect is still present after the change.
- **Detection procedure**:
  1. From the task statement, note the symbol(s) named as misbehaving and the import path they are imported from. [reads: task]
  2. In the static facts repo tree, locate the package directory that owns that import path (the directory whose name matches the top-level import, containing the implementation modules). [reads: static facts — repo tree]
  3. Read the list of files added/changed by the program. Fire if every added/changed file lies outside that package directory (e.g. a new top-level `.py` script or notebook) and no implementation module inside it is edited. [reads: code]
- **Counter-example**: A change set that edits the implementation module (reordering the loop, flipping the comparison, relaxing the argument validation) and *additionally* drops in a reproduction script — the script is present, but a package module is also modified.
- **Discriminator**: Whether at least one file inside the implementing package directory appears in the change set. Reproduction-only change sets touch zero of them.
- **Consequence**: The reported defect persists; any hidden/regression test that asserts the corrected behavior of the named symbol fails (AssertionError, or the very exception class the task says is raised spuriously, e.g. `ValueError`). Task requirement "fix the behavior" is unmet regardless of what the added script prints.
- **Evidence**: The entire diff was `test_coalesce_issue.py` (new file) with no edit under the package source directory, while the task described an inverted iteration order, inverted skip predicate, and a spurious `ValueError` inside that package's module.
20Verification script that prints success unconditionallycodeswesmith/mahmoud__glom.fb3c4e76
Applies when
code: the program adds a script whose purpose is to check that behavior described in the task is now correct
Pattern
The script only prints computed values and ends with a hardcoded success message; it contains no assert (or comparison that raises/exits non-zero), so it emits "passed" even when every printed value is wrong.
Detection procedure
  1. Locate the added script and the lines that exercise the behavior described in the task. [reads: code]
  2. Read the task statement for the concrete expected outputs it specifies. [reads: task]
  3. Fire if those expected values appear only inside f-strings/print arguments and the file contains no assert, no if result != expected: raise/sys.exit, and ends with a literal success message printed outside any conditional. [reads: code]
Counter-example
A script that computes the same values but writes assert result == 1 (or sys.exit(1) on mismatch) before printing a success line — the success message is then reachable only when the checks hold.
Discriminator
Presence of at least one statement whose failure changes control flow or exit status. Print-only scripts have none, so their "passed" output carries zero information.
Consequence
The program's own evidence of correctness is vacuous; a defect that the run actually exhibits is reported as success, so the submitted artifact ships unfixed. Additionally, if the script's filename matches pytest's collection pattern (test_.py / _test.py) and its calls are at module level, any exception raised there becomes a collection-time error that aborts or errors the whole test session rather than one test.
Evidence
print(f"Test 1 - Expected: 1, Got: {result}") followed by an unconditional print("\nAll tests passed!"), in a module-level-executing file named test_*.py at repo root, with no assertion anywhere.
id 49e1684b2167 · mined from swesmith/mahmoud__glom.fb3c4e76 mahmoud__glom.fb3c4e76.func_basic__4c5ut8n6
curation fields
{
 "predicate": "test or grader output will report failure",
 "code_requirement": "1. Locate the added script and the lines that exercise the behavior described in the task. [reads: code]",
 "prediction": "The program's own evidence of correctness is vacuous; a defect that the run actually exhibits is reported as success, so the submitted artifact ships unfixed. Additionally, if the script's filename matches pytest's collection pattern (`test_*.py` / `*_test.py`) and its calls are at module level, any exception raised there becomes a collection-time error that aborts or errors the whole test session rather than one test."
}
raw text (what the judge reads)
### Verification script that prints success unconditionally
- **Applies when**: `code`: the program adds a script whose purpose is to check that behavior described in the task is now correct
- **Pattern**: The script only prints computed values and ends with a hardcoded success message; it contains no `assert` (or comparison that raises/exits non-zero), so it emits "passed" even when every printed value is wrong.
- **Detection procedure**:
  1. Locate the added script and the lines that exercise the behavior described in the task. [reads: code]
  2. Read the task statement for the concrete expected outputs it specifies. [reads: task]
  3. Fire if those expected values appear only inside f-strings/`print` arguments and the file contains no `assert`, no `if result != expected: raise/sys.exit`, and ends with a literal success message printed outside any conditional. [reads: code]
- **Counter-example**: A script that computes the same values but writes `assert result == 1` (or `sys.exit(1)` on mismatch) before printing a success line — the success message is then reachable only when the checks hold.
- **Discriminator**: Presence of at least one statement whose failure changes control flow or exit status. Print-only scripts have none, so their "passed" output carries zero information.
- **Consequence**: The program's own evidence of correctness is vacuous; a defect that the run actually exhibits is reported as success, so the submitted artifact ships unfixed. Additionally, if the script's filename matches pytest's collection pattern (`test_*.py` / `*_test.py`) and its calls are at module level, any exception raised there becomes a collection-time error that aborts or errors the whole test session rather than one test.
- **Evidence**: `print(f"Test 1 - Expected: 1, Got: {result}")` followed by an unconditional `print("\nAll tests passed!")`, in a module-level-executing file named `test_*.py` at repo root, with no assertion anywhere.