SWE-bench · 77 rubrics · 77 rubrics, failure-based recipe; only the pool for the random-12 baseline
HF EdwardoSunny/swe-rubrics-onpolicy-gym · local data/libraries/swe-rubrics-onpolicy-gym.json
task -- the patch's only change is swapping/reordering logical operators (||↔&&, added parentheses, added nil/precondition guards) inside an existing condition that gates a behavior such as retry, fallback, caching, or early-exit.a69231ec7064 · mined from swesmith/lqs__sqlingo.ed36ef03 lqs__sqlingo.ed36ef03.func_pm_flip_operators__jcqg2mrg### Speculative tightening of a boolean condition ("safety" refactor) that disables an existing code path
- **Applies when**: `task` -- the patch's only change is swapping/reordering logical operators (`||`↔`&&`, added parentheses, added nil/precondition guards) inside an existing condition that gates a behavior such as retry, fallback, caching, or early-exit.
- **Pattern**: The agent assumes the pre-existing expression is a precedence/short-circuit mistake and rewrites it into the "obviously correct" stricter form, making the guarded branch fire in strictly fewer cases, without any evidence from the reported symptom that this branch was firing too often.
- **Detection procedure**:
1. Extract the reported failure symptom from the task and state which observable behavior must change.
2. Compute the truth table of the condition before and after the patch; note whether the guarded branch now executes strictly less often, strictly more often, or differently in both directions.
3. Ask whether "branch taken too often" (for a tightening change) is actually what the symptom describes; if the symptom is instead "behavior never happens / feature missing / value wrong elsewhere", the direction of the change contradicts the symptom.
4. Check for corroboration in the task: an error/panic trace pointing at the guarded call, a spec sentence stating the precondition, or a test name describing the disabled path. Absent all three, treat the operator swap as speculative.
- **Discriminator**: A real violation is a tightening/loosening with no symptom or spec linkage — justified only by "this guard looks unsafe" or "the precedence looks unintended". A legitimate fix is one where the task reports the exact failure the operator change prevents (e.g., a nil-dereference stack trace inside the short-circuited callee, or a documented precondition), so before/after truth tables differ precisely on the reported inputs.
- **Consequence**: The intended path (retry, fallback, recovery) stops executing for the cases the codebase and its tests rely on, so the original bug remains unfixed and previously passing behavior silently regresses.task -- a bug-fix patch alters a comparison operator in a capacity/limit guard, or changes what an internal helper returns to callers (e.g., from a nil sentinel to a live internal object), rather than fixing logic the task actually describes.> for >= (or vice versa) in a size/bound check and simultaneously returning an internal structure where the previous code returned nil — without establishing that the original behavior violated any documented invariant. The edits move the code away from the intended contract instead of toward it, and may reintroduce the very defect being hunted.bcc516223876 · mined from swesmith/bluele__gcache.d8b7e051 bluele__gcache.d8b7e051.lm_modify__gzj2rany### Speculative flipping of boundary operators and return-value contracts without tracing callers/invariants - **Applies when**: `task` -- a bug-fix patch alters a comparison operator in a capacity/limit guard, or changes what an internal helper returns to callers (e.g., from a nil sentinel to a live internal object), rather than fixing logic the task actually describes. - **Pattern**: The agent "tightens" or "loosens" existing code it merely finds suspicious — swapping `>` for `>=` (or vice versa) in a size/bound check and simultaneously returning an internal structure where the previous code returned nil — without establishing that the original behavior violated any documented invariant. The edits move the code away from the intended contract instead of toward it, and may reintroduce the very defect being hunted. - **Detection procedure**: 1. List each hunk and classify it: does it change a boundary/comparison operator, or does it change a returned value/sentinel that crosses a function boundary? 2. For each such hunk, find the invariant or caller that pins the correct behavior — the container's own accounting (does length include the item being added yet?), and every call site's handling of the return value (does it check for nil, ignore the value, or expose it to the public API?). 3. Confirm the task statement/reproduction actually implicates that operator or that return value; if the task is silent about it and no caller demands the change, treat the hunk as unjustified. 4. Check that the new behavior does not contradict existing tests/doc comments describing the pre-edit semantics (e.g., "returns nil on success", "cache holds up to N entries"). - **Discriminator**: A real violation is a hunk whose new semantics no caller or invariant requires, and which changes observable behavior (element count at capacity, non-nil value surfaced through the public API). A look-alike that is fine is an operator or return-value change directly derived from a stated failing scenario, where the reviewer can name the caller or invariant that the old code broke and the new code satisfies. - **Consequence**: The container silently holds one fewer/more element than its configured capacity and callers that branch on a nil result start seeing a non-nil internal object, so the original defect stays unfixed while new regressions appear in unrelated code paths and the target tests keep failing.
task -- the issue text blames a recently simplified/changed code path (e.g. "now uses X() instead of handling streams/objects directly") and the patch responds by pasting back a more elaborate previous version of that function.self.<attr>, Class._private_helper, constructor/dataclass kwargs, module-level imports, exception types), verify from the patch context that it exists in the current codebase and has the assumed signature.try/finally where the guarded variable can be unbound, removed blank lines/indentation around following defs, duplicated or half-removed old code.420a93d6b3ae · mined from swesmith/facebookresearch__hydra.0f03eb60 facebookresearch__hydra.0f03eb60.lm_rewrite__nh2vejty### Reintroducing an older implementation ("blind revert") without validating it against the module's current API and import integrity
- **Applies when**: `task` -- the issue text blames a recently simplified/changed code path (e.g. "now uses `X()` instead of handling streams/objects directly") and the patch responds by pasting back a more elaborate previous version of that function.
- **Pattern**: The agent reconstructs the old implementation from memory or from the issue narrative instead of deriving the minimal fix, and the reconstructed code references helpers, keyword arguments, attributes, or imports that may no longer exist or are inconsistent with the current file — producing an import-time/registration-time failure or dead leftover code rather than a behavioral fix.
- **Detection procedure**:
1. Read the issue and note whether it *describes* the current implementation as suspicious versus *specifies* the required behavior; treat "it now does X" as a symptom description, not an instruction to restore some remembered predecessor.
2. For every name the patch introduces (`self.<attr>`, `Class._private_helper`, constructor/dataclass kwargs, module-level imports, exception types), verify from the patch context that it exists in the current codebase and has the assumed signature.
3. Check the patch's structural hygiene at module scope: imports added but unused or used but not added, `try`/`finally` where the guarded variable can be unbound, removed blank lines/indentation around following defs, duplicated or half-removed old code.
4. Ask whether the module can still be imported and its plugin/entry point registered; any unresolved name or malformed block at import time is disqualifying regardless of the logic's merit.
- **Discriminator**: A legitimate restore cites the current file's actual definitions (helper methods, result-object fields) and leaves the module importable with no orphan imports or unreachable code; a violation adds richer-looking logic whose dependencies are assumed rather than shown, or leaves the surrounding module syntactically/structurally altered.
- **Consequence**: The module raises at import time or constructs results with invalid arguments, so the plugin/component never registers — every test that discovers or loads it fails with a registration/import error, masking (and worsening) the original bug.task -- the bug report lists a few concrete misbehaviours (e.g. swapped/inverted parameter assignments) and the patch edits a single initialization/setup block that contains other suspicious logic besides the listed items.length/size parameter set to a constant instead of derived from the input, missing "required argument" checks, a fallback branch collapsed to one case).819b1928ada7 · mined from swesmith/pallets__click.fde47b4b pallets__click.fde47b4b.func_basic__ccb52392### Incomplete restoration of a corrupted region: only the explicitly-reported symptoms are fixed
- **Applies when**: `task` -- the bug report lists a few concrete misbehaviours (e.g. swapped/inverted parameter assignments) and the patch edits a single initialization/setup block that contains other suspicious logic besides the listed items.
- **Pattern**: The patch mechanically corrects each symptom named in the issue text (un-swapping two fields, un-negating a flag) but leaves adjacent lines in the same block that are also wrong — degenerate defaults, dropped derivation logic, removed validation/fallback, or a substituted API/clock — because the report never named them.
- **Detection procedure**:
1. Read the report and note that it ends with a vague catch-all ("also affects default behaviour / other environments"), which signals the corruption is broader than the enumerated list.
2. Read the *whole* function or block the patch touches, not just the changed lines, and list every value that is hard-coded, defaulted to a placeholder, or computed without using its inputs (e.g. a `length`/size parameter set to a constant instead of derived from the input, missing "required argument" checks, a fallback branch collapsed to one case).
3. For each such line, ask whether a correct implementation could plausibly behave that way given the documented public API; if not, it is part of the same injected corruption and must be restored.
4. Confirm the patch restores every one of those lines, including anything that feeds derived state used later (ETA/progress/percent computations, stream selection, error paths).
- **Discriminator**: A real violation leaves logic that is internally inconsistent with the documented API or that ignores a parameter/input entirely (e.g. a length is never inferred from the iterable, so dependent computations can only ever yield zero). A look-alike that is fine is untouched code that merely looks unusual but is consistent, has a comment explaining the intent, and correctly consumes its inputs.
- **Consequence**: The obvious symptoms disappear while derived behaviour stays broken — dependent computations return degenerate values, invalid argument combinations are silently accepted, and tests covering the unnamed parts of the same block continue to fail.task -- the bug is that a previously supported input class (e.g. an extra MIME type, scheme, extension, or format) is rejected, and the patch relaxes a validation/dispatch check.is_*/can_handle detectors, __repr__/serialization, parsing of the produced value) unchanged, and/or substitutes a silent hard-coded default where the original code derived or rejected the value.1a909a92ff41 · mined from swesmith/stanfordnlp__dspy.651a4c71 stanfordnlp__dspy.651a4c71.pr_7872### Incomplete restoration: only one of several parallel type/format guards is relaxed - **Applies when**: `task` -- the bug is that a previously supported input class (e.g. an extra MIME type, scheme, extension, or format) is rejected, and the patch relaxes a validation/dispatch check. - **Pattern**: The patch loosens the narrow check in the single code path named in the bug report, but leaves other functions that apply the same narrow predicate (validation helpers, `is_*`/`can_handle` detectors, `__repr__`/serialization, parsing of the produced value) unchanged, and/or substitutes a silent hard-coded default where the original code derived or rejected the value. - **Detection procedure**: 1. From the task, identify the exact restrictive condition (e.g. a prefix/extension/enum comparison) that causes the rejection. 2. Grep the patched file(s) and callers for every occurrence of that same literal or equivalent predicate, including boolean detector functions and any code that re-parses the produced string. 3. Confirm the patch updates *all* of them consistently; flag if any still assumes the narrow form, or if the new code silently falls back to a fixed value instead of deriving/raising as the intended behavior specifies. 4. Trace one end-to-end flow for the newly supported input (detect → encode/convert → format/repr) and check it survives every stage. - **Discriminator**: A real violation is when a remaining narrow check is reachable for the newly supported input class, or a fallback masks an undeterminable value; it is fine if the untouched checks are provably restricted to a different input domain (e.g. genuinely image-only conversion) and the fallback matches documented legacy behavior. - **Consequence**: The reported input appears fixed in the one entry point but is still filtered out or mislabeled elsewhere, so the feature stays broken for real callers and tests exercising detection or round-tripping fail.
task -- The report names specific lines/conditions as the root cause, and the patch changes exactly those lines inside a function whose other statements (default value derivation, key/prefix construction, argument order) were plausibly altered by the same regression.6747a2d97118 · mined from swesmith/Cog-Creators__Red-DiscordBot.33e0eac7 Cog-Creators__Red-DiscordBot.33e0eac7.combine_file__eq2t7cw0### Trusting the bug report's line-by-line diagnosis instead of verifying every line in the affected region - **Applies when**: `task` -- The report names specific lines/conditions as the root cause, and the patch changes exactly those lines inside a function whose other statements (default value derivation, key/prefix construction, argument order) were plausibly altered by the same regression. - **Pattern**: The agent applies the reporter's suggested edits verbatim — including one that is actually correct as-is — and never independently re-derives the intended semantics of the neighboring statements, so a co-located defect (e.g., how a default name/key is computed) survives and a correct line gets "fixed" into a wrong one. - **Detection procedure**: 1. From the task, list each line the reporter accuses, and separately note the observable symptom(s) actually demonstrated (error messages, failing call). 2. For each accused line, check whether the reported symptom actually requires that change, or whether the reporter merely guessed; look for consumers (callers, key lookups, tests, docs, sibling functions) that pin the correct form. 3. Read every remaining statement in the touched function/region and confirm each derived value or argument matches what its consumer expects (same naming convention, same ordering, same key layout used elsewhere). 4. Flag the patch if any accused line was changed without independent evidence, or if any unaccused statement in the region remains inconsistent with its consumers. - **Discriminator**: A real violation is when the patch's correctness rests solely on the reporter's assertion and a neighboring statement in the same function still contradicts an in-repo consumer/convention; a look-alike that is fine is when the agent's edits are each independently confirmed by call sites or sibling code, even if they happen to coincide with the reporter's suggestions. - **Consequence**: The headline symptom disappears while a second defect in the same function persists (and a previously-correct branch may be broken), so behavior-level tests on default naming/keying still fail and removal/lookup paths silently mismatch.
task -- The report says two behaviors are "inverted" and the patch fixes it by exchanging the statements inside an if cond { A } else { B } while leaving cond untouched.x == nil, len(s) == 0) and check every use of those values in that branch is legal and meaningful.323c9fc19e07 · mined from swesmith/caddyserver__caddy.77dd12cc caddyserver__caddy.77dd12cc.func_pm_ctrl_invert_if__ezdzt5lu### Blind swap of if/else branch bodies without re-aligning the guard condition
- **Applies when**: `task` -- The report says two behaviors are "inverted" and the patch fixes it by exchanging the statements inside an `if cond { A } else { B }` while leaving `cond` untouched.
- **Pattern**: The agent treats "inverted logic" as a body swap only, so the guard no longer protects the code it was written for: the nil/empty/zero check ends up guarding the branch that dereferences or consumes that same value, and the other branch gets the fallback. The code compiles and reads plausibly, but the guard–body pairing (and the contract the rest of the codebase relies on) is now wrong.
- **Detection procedure**:
1. Locate the conditional the patch touches and write down the condition and both post-patch bodies.
2. For each branch, substitute what the condition asserts about the values (e.g. `x == nil`, `len(s) == 0`) and check every use of those values in that branch is legal and meaningful.
3. Cross-check the intended mapping against independent evidence: doc comments, adjacent placeholder/field assignments, existing tests, and other call sites that read the same output — not just the issue's prose.
4. Flag the patch if a branch now uses a value the guard says is absent, or if the mapping contradicts the documented/tested contract.
- **Discriminator**: A real violation leaves a guard paired with a body that violates it (nil deref, unreachable fallback) or contradicts documented/tested behavior. A look-alike that is fine is a swap where the condition is also inverted (or the guard is value-independent), each body only touches values valid under its branch, and existing docs/tests corroborate the new mapping.
- **Consequence**: Runtime panic or wrong fallback value on the path the guard was meant to protect, plus hidden tests asserting the established contract keep failing while the diff looks like the obvious fix.task -- the bug report cites a handful of concrete input→output examples produced by one function containing many parallel branches (per-block/per-type/per-range cases), and the patch edits only some of those branches.x = x), wrong default/fallback values, inconsistent boundary operators (> vs >=, < vs <=), or a hard-coded constant where a lookup/mapping belongs.else/default paths) and mark which ones the patch touched.d382fdc47e84 · mined from swesmith/Mimino666__langdetect.a1598f1a Mimino666__langdetect.a1598f1a.func_basic__s4s0fk2j### Incomplete repair of a multi-branch normalization/dispatch table (only the reported examples fixed) - **Applies when**: `task` -- the bug report cites a handful of concrete input→output examples produced by one function containing many parallel branches (per-block/per-type/per-range cases), and the patch edits only some of those branches. - **Pattern**: The patch fixes exactly the branches exercised by the examples in the issue text, while leaving sibling branches in the same dispatch still holding corrupted logic — leftover no-op assignments (`x = x`), wrong default/fallback values, inconsistent boundary operators (`>` vs `>=`, `<` vs `<=`), or a hard-coded constant where a lookup/mapping belongs. - **Detection procedure**: 1. From the task, note that the failing examples are *samples* of a broader table, not necessarily the full set of defects. 2. Enumerate every branch of the edited dispatch (including its `else`/default paths) and mark which ones the patch touched. 3. For each untouched branch, check for self-evident corruption smells: identity/no-op assignments, a constant where neighbours use a map lookup, a fallback that discards input, and comparison boundaries that differ in strictness from analogous sibling branches. 4. If any untouched branch shows such a smell, or any touched branch's boundary strictness was changed without justification from the report, flag the patch as incomplete. - **Discriminator**: A real violation is an untouched (or boundary-flipped) branch whose logic is internally inconsistent or degenerate on its own terms (no-op, dropped input, mismatched boundary vs. siblings); a look-alike that is fine is a branch that is intentionally simple/uniform (e.g., legitimately collapsing a whole block to one representative character) and is consistent with the surrounding branches and documented intent. - **Consequence**: Hidden tests that cover the full table still fail, and the remaining corrupted branches silently produce wrong outputs for inputs not mentioned in the report.
task -- The report describes one or two visibly broken behaviors (e.g., an UnboundLocalError) and the patch edits only the exact functions named in the report.import/helper — so other public entry points remain non-functional even though the reported snippets now run.NotImplementedError, AttributeError, NameError) or returns nonsense after the patch; a look-alike that is fine is unrelated pre-existing stylistic oddity or intentionally abstract classes/imports genuinely used elsewhere.2d1d25155e12 · mined from swesmith/life4__textdistance.c3aca916 life4__textdistance.c3aca916.combine_file__hxml8xjz### Incomplete restoration: only the symptom site fixed while other corrupted/missing code in the same module is left broken - **Applies when**: `task` -- The report describes one or two visibly broken behaviors (e.g., an `UnboundLocalError`) and the patch edits only the exact functions named in the report. - **Pattern**: The patch repairs the reported statements but ignores collateral damage elsewhere in the same file/class hierarchy — e.g., a sibling class whose overriding method was deleted, a truncated function body, or a lost `import`/helper — so other public entry points remain non-functional even though the reported snippets now run. - **Detection procedure**: 1. From the task, note the module(s) that contain the reported defect and the fact that the defect looks like mechanically scrambled/removed code. 2. Read the *whole* surrounding file (or at least every class in it), not just the diff hunks, checking that each concrete class still implements the methods its base class requires and that no function body is missing, empty, or ends in an unreachable/dangling statement. 3. Compare the patch's touched regions against that inventory: does it restore every anomaly found in step 2, or only the ones the reporter happened to name? 4. Also check for structural leftovers of the corruption (removed blank lines/class separation, now-unused imports) that indicate deleted code nearby. - **Discriminator**: A real violation is when a reachable public API in the same module still raises (`NotImplementedError`, `AttributeError`, `NameError`) or returns nonsense after the patch; a look-alike that is fine is unrelated pre-existing stylistic oddity or intentionally abstract classes/imports genuinely used elsewhere. - **Consequence**: Hidden/parametrized tests that iterate over all algorithms or all public entry points in the module fail on the untouched sibling, so the fix is rejected despite the reported error being gone.
task -- the patch repairs a bug by flipping/adjusting one comparison, bound, or operator inside a loop or guard that initializes or resets shared state.eb3555c31c13 · mined from swesmith/tylertreat__BoomFilters.db654574 tylertreat__BoomFilters.db654574.func_pm_flip_operators__u4q9wrsl### Incomplete fix: single-site operator repair without sweeping sibling occurrences - **Applies when**: `task` -- the patch repairs a bug by flipping/adjusting one comparison, bound, or operator inside a loop or guard that initializes or resets shared state. - **Pattern**: The agent spots one obviously inverted/off-by condition, edits that single line, and stops — never checking whether the same construct (the same allocate-then-fill loop, the same bound expression, the same invariant) is duplicated in constructors, re-initializers, resize/copy helpers, or other paths that must stay consistent, and never confirming the repaired path actually restores every field the buggy path left unset. - **Detection procedure**: 1. From the task, identify the observable broken behavior and which structure/state it concerns. 2. In the patch, note the single edited condition and the state it populates. 3. Grep the codebase for other sites that build or reset the same state (same allocation + fill idiom, same size/capacity fields) and compare their conditions and the set of fields they initialize against the patched site. 4. Flag the patch if any sibling site keeps the wrong bound/condition, or if the patched site still leaves part of the state uninitialized relative to a known-good site. - **Discriminator**: A real violation exists when a genuinely equivalent construct elsewhere still carries the defect or the patched path remains incomplete versus a reference initializer; it is *not* a violation when the buggy construct provably occurs only once and the patched loop initializes exactly the same fields, in the same way, as every other constructor of that state. - **Consequence**: The reported symptom appears fixed on one code path while other paths still produce empty/short/uninitialized state, causing index-out-of-range panics, silent data loss, or intermittent failures that the targeted regression test may not even reach.
task -- the task asks to fix a functional bug/failing test, and the submitted patch consists solely of stylistic edits (operand reordering in comparisons, renaming, formatting, comment changes) that are semantically identical to the original code.!= to ==, swapping operands of a non-commutative operator, reordering calls with side effects, fixing short-circuit order that guards a nil deref), or a patch where cosmetic churn sits alongside a genuine logic/API fix.cfb0fa69f82c · mined from swesmith/caddyserver__caddy.77dd12cc caddyserver__caddy.77dd12cc.func_pm_op_swap__azdxnk45### Cosmetic-only rewrite that cannot change program behavior - **Applies when**: `task` -- the task asks to fix a functional bug/failing test, and the submitted patch consists solely of stylistic edits (operand reordering in comparisons, renaming, formatting, comment changes) that are semantically identical to the original code. - **Pattern**: The agent "fixes" the issue by normalizing style — e.g. flipping Yoda conditions to conventional order, reordering equality operands, reflowing code — producing a diff whose compiled semantics are byte-for-byte equivalent to the pre-patch program, so no observable behavior changes. - **Detection procedure**: 1. Read the task/bug report and note the concrete observable symptom that must change (wrong output, error, crash, failing assertion). 2. For each hunk in the patch, ask whether the new code can ever evaluate differently from the old code for any input (different branch taken, different value, different call, different ordering of side effects). 3. If every hunk is an equivalence-preserving transformation, mark the patch as not addressing the bug. 4. Confirm no other hunk in the patch contains a real semantic change that the cosmetic edits merely accompany. - **Discriminator**: A real violation is a patch where *all* semantic-affecting behavior is unchanged. A look-alike that is fine is a patch that looks like a style cleanup but actually alters semantics (e.g. changing `!=` to `==`, swapping operands of a non-commutative operator, reordering calls with side effects, fixing short-circuit order that guards a nil deref), or a patch where cosmetic churn sits alongside a genuine logic/API fix. - **Consequence**: The original defect remains; the target tests still fail, review time is wasted, and the noisy diff obscures that no fix was actually attempted.
task -- the fix is a one-token change to a boolean/comparison operator (&&↔||, <↔<=, ==↔!=, negation) inside a predicate that classifies inputs.369079b1f2ad · mined from swesmith/caddyserver__caddy.77dd12cc caddyserver__caddy.77dd12cc.func_pm_flip_operators__582xt1qo### Condition-operator flipped in the wrong direction (patch inverts already-correct logic) - **Applies when**: `task` -- the fix is a one-token change to a boolean/comparison operator (`&&`↔`||`, `<`↔`<=`, `==`↔`!=`, negation) inside a predicate that classifies inputs. - **Pattern**: The agent toggles the operator to whatever it *isn't* currently, without deriving the required truth table from the task's concrete examples; the result either restores the original bug or broadens/narrows the predicate so a different set of inputs is misclassified. Identifier names or nearby comments (e.g. a variable named "xOrY") are taken as proof of the intended operator instead of the specified behavior. - **Detection procedure**: 1. From the task/issue text, list every concrete input mentioned (or plausibly covered) together with the expected true/false outcome of the predicate. 2. Evaluate the predicate as it exists *before* the patch on each listed input, then evaluate the *post-patch* predicate on the same inputs. 3. Confirm the post-patch version satisfies strictly more of the expected outcomes than the pre-patch version; if any input that previously produced the expected result now flips to the wrong result, flag the patch. 4. Check whether the pre-patch code already produced the expected result for the reported failure — if so, the operator was not the defect and the patch is an inversion, not a fix. - **Discriminator**: A genuine violation makes at least one input that the task expects handled correctly become misclassified, or leaves the reported case unchanged. A legitimate look-alike fix is a single-operator change where every task-derived example evaluates correctly afterwards, even if the resulting expression reads oddly relative to a stale variable name or comment. - **Consequence**: The predicate misroutes inputs (wrong adapter/branch/format selected), so the reported bug persists while previously working inputs regress — and the fix is indistinguishable from re-introducing the original defect.
task -- a patch alters a conditional threshold, boundary value, or literal in one branch of a block that contains near-identical sibling logic for analogous values (e.g. repeated singular/plural, unit, or limit selection).4f4ba70f7f7e · mined from swesmith/mgechev__revive.03e81029 mgechev__revive.03e81029.func_pm_op_change_const__3sknxgbd### Consistency with parallel sibling branches / recorded expectations - **Applies when**: `task` -- a patch alters a conditional threshold, boundary value, or literal in one branch of a block that contains near-identical sibling logic for analogous values (e.g. repeated singular/plural, unit, or limit selection). - **Pattern**: The patch rewrites the branch to what "seems semantically correct" in isolation, while leaving the adjacent sibling branches using the original comparison, so the code becomes internally inconsistent with the convention the rest of the codebase (and its fixtures/golden outputs) depends on. - **Detection procedure**: 1. Read the task/bug report and note the exact reported symptom; check whether it describes this branch's output at all. 2. Locate all sibling code in the same block or nearby that performs the same kind of selection for other inputs, and record the comparison each one uses. 3. Compare the patched branch against those siblings: does it now use a different operator/threshold than every unmodified sibling? 4. Grep tests, golden files, or fixtures for the strings/values produced by this branch to confirm which convention is actually expected. - **Discriminator**: A real violation is a lone divergence from an otherwise uniform sibling pattern with no test/fixture or task text demanding the new value. It is fine if the task explicitly reports this branch's output as wrong, or if the siblings are updated together, or if existing expectations confirm the new comparison. - **Consequence**: The "fix" flips output for the wrong inputs, breaking the assertions or golden output that encode the project's actual convention, so the target test still fails and previously passing cases may regress.
task -- a patch fixes behavior by swapping/negating the two arms of a conditional that tests membership in a lookup structure (map/set/index) or a boolean flag, rather than changing the data or the condition itself.stillExists, found, toX[key]) and assumes the "obvious" pairing of branch bodies, without checking how that structure/flag was actually populated or mutated earlier — e.g. a map whose matched keys are deleted as they are processed, or a flag whose polarity is inverted by construction, so presence actually means the opposite of what the name suggests.2234da5311c7 · mined from swesmith/skeema__skeema.defb0097 skeema__skeema.defb0097.func_pm_ctrl_invert_if__pe1l9m20### Inverting a branch based on identifier-name intuition instead of traced container semantics - **Applies when**: `task` -- a patch fixes behavior by swapping/negating the two arms of a conditional that tests membership in a lookup structure (map/set/index) or a boolean flag, rather than changing the data or the condition itself. - **Pattern**: The agent reads the condition's variable name (e.g. `stillExists`, `found`, `toX[key]`) and assumes the "obvious" pairing of branch bodies, without checking how that structure/flag was actually populated or mutated earlier — e.g. a map whose matched keys are deleted as they are processed, or a flag whose polarity is inverted by construction, so presence actually means the opposite of what the name suggests. - **Detection procedure**: 1. Locate the condition whose arms the patch swapped and identify the exact structure or flag it tests. 2. Scroll back to where that structure/flag is built and to every place it is mutated (inserts, deletes, negations) before this point; determine what membership/truth really encodes at this line. 3. Re-read the pre-patch code under that real meaning and check whether it was already consistent; then check whether the post-patch code is consistent. 4. Confirm the reported symptom in the task is actually explained by branch polarity here, not by a different site (ordering, output formatting, a downstream consumer). - **Discriminator**: A real violation is when tracing construction/mutation shows the original arm pairing was already correct (or the swap makes both arms wrong), so the swap is driven only by naming intuition. A legitimate fix is one where the reviewer can point to concrete construction/mutation evidence that the old pairing contradicted the encoded meaning, or where the task explicitly describes the two branches' outputs being exchanged. - **Consequence**: Every item takes the wrong code path — entities that should be preserved get dropped/emitted and vice versa — producing inverted output that silently corrupts results for all inputs, while the original reported bug remains unfixed.
task -- the patch fixes a specific reported defect but also rewrites a neighbouring formatting/__repr__/__str__/logging expression that was not implicated in the bug.str(self) for a direct base-class __repr__ call, or reordering/reformatting output), silently changing the textual output that callers and tests depend on.str and repr of the involved types are defined differently.48a3023d79dc · mined from swesmith/mido__mido.a0158ff9 mido__mido.a0158ff9.combine_file__h5p22q8m### Unrelated "cleanup" that changes a public string/repr format - **Applies when**: `task` -- the patch fixes a specific reported defect but also rewrites a neighbouring formatting/`__repr__`/`__str__`/logging expression that was not implicated in the bug. - **Pattern**: The agent correctly repairs the faulty logic, then "improves" an adjacent display method (e.g. swapping `str(self)` for a direct base-class `__repr__` call, or reordering/reformatting output), silently changing the textual output that callers and tests depend on. - **Detection procedure**: 1. From the task description, list exactly which behaviours are reported as broken. 2. Diff-walk the patch and mark each hunk as either (a) required by that list or (b) incidental. 3. For every incidental hunk, evaluate whether the produced string/serialized value is byte-identical to before; check especially whether `str` and `repr` of the involved types are defined differently. 4. Flag the patch if any incidental hunk can yield a different output than the pre-patch code. - **Discriminator**: A real violation changes observable output (different prefix, class name, field order, quoting) for the same input; a look-alike that is fine is a refactor that provably routes to the identical formatting code path, or a formatting change that is itself named in the task as the bug. - **Consequence**: Tests or downstream consumers asserting on the exact rendered text fail even though the targeted defect was fixed, so the fix is rejected and the regression surface grows.
task -- the patch changes a value that is stored/retained beyond the current call (in a context, struct field, cache, or closure) from a freshly allocated container to a direct reference to a caller- or framework-supplied map/slice/buffer.make(...)/copy-then-store with the incoming parameter itself, on the assumption that reusing the argument is equivalent and cheaper, ignoring that the caller retains ownership and may mutate, clear, or pool that container after the call returns.93c9372f1015 · mined from swesmith/dimfeld__httptreemux.53a6a099 dimfeld__httptreemux.53a6a099.lm_modify__g4vlfczb### Aliasing caller-owned mutable state instead of storing an independent copy - **Applies when**: `task` -- the patch changes a value that is stored/retained beyond the current call (in a context, struct field, cache, or closure) from a freshly allocated container to a direct reference to a caller- or framework-supplied map/slice/buffer. - **Pattern**: The "fix" replaces `make(...)`/copy-then-store with the incoming parameter itself, on the assumption that reusing the argument is equivalent and cheaper, ignoring that the caller retains ownership and may mutate, clear, or pool that container after the call returns. - **Detection procedure**: 1. Identify the lifetime of the field being assigned: is the containing object stored somewhere that outlives the function body (request context, registry, goroutine, returned struct)? 2. Identify the provenance of the new right-hand side: is it a parameter or otherwise supplied by code outside this function that could still hold and reuse it? 3. If lifetime outlives the call and provenance is external, check whether the patch adds any copy/clone; if not, flag it. 4. Confirm the removed allocation/copy was pre-existing correct behavior rather than the injected defect — a patch that merely re-introduces the exact pattern the bug report describes is going the wrong direction. - **Discriminator**: A real violation shares a mutable container whose owner may reuse or mutate it (router param maps, pooled buffers, reused slices) across the retained object's lifetime. It is fine to alias if the value is immutable, freshly created by the caller solely for this call and never touched again, or if the retained object's lifetime provably ends before the call returns. - **Consequence**: Retained state silently changes or empties under later requests/iterations, producing intermittent wrong data, cross-request leakage, or data races that unit tests on stored-state behavior expose immediately.
task -- a reported bug lives inside a function that dispatches to different wrappers/handlers based on a predicate (e.g. async vs sync callable, connection type, mode flag), and the patch only touches the symptoms named in the report.async def wrapper that awaits the target must be guarded by "target is async"; a branch labelled/behaving as the sync path must be guarded by the negation).8fe62e9008ba · mined from swesmith/encode__starlette.db5063c2 encode__starlette.db5063c2.combine_file__65qdqflw### Missing sibling defect in the same dispatch/branch logic - **Applies when**: `task` -- a reported bug lives inside a function that dispatches to different wrappers/handlers based on a predicate (e.g. async vs sync callable, connection type, mode flag), and the patch only touches the symptoms named in the report. - **Pattern**: The patch corrects the explicitly-reported items (status code, comparison direction, argument order) but leaves an inverted or mismatched branch predicate untouched, so control flow still selects the wrong wrapper/handler for some inputs. - **Detection procedure**: 1. From the task, locate the function(s) containing the reported bug and read the *entire* body, not just the changed lines. 2. For every conditional that chooses between alternative implementations, check that the predicate's polarity matches the code inside the branch (e.g. an `async def` wrapper that awaits the target must be guarded by "target is async"; a branch labelled/behaving as the sync path must be guarded by the negation). 3. Check the patch diff: if such a predicate is inconsistent with its branch body and the patch did not fix it, flag the patch. 4. Also verify related helpers/initializers in the same file for the same inconsistency class (e.g. ternaries whose true/false arms are swapped). - **Discriminator**: A real violation is a predicate whose branch body is semantically incompatible with the condition (awaiting a non-coroutine, calling a coroutine without awaiting, closing a connection type it never receives). A look-alike that is fine is an unusual-but-consistent predicate where both branches work for their inputs, or a stylistic difference with no behavioral mismatch. - **Consequence**: Requests take the wrong wrapper: coroutines are never awaited (or sync results are awaited), producing "coroutine was never awaited" warnings, empty/erroneous responses, and failing tests even though the reported symptoms appear fixed.
task -- the patch's entire content consists of rewriting existing boolean/comparison expressions (swapping operand order, reversing >=/<=, inverting nil-checks) rather than adding, removing, or redirecting logic that the task's symptom points to.nil != err → err != nil) mixed with semantics-changing ones.dd82db344247 · mined from swesmith/caddyserver__caddy.77dd12cc caddyserver__caddy.77dd12cc.func_pm_op_swap__3c77gu8v### Speculative flipping of comparison/guard conditions instead of a task-evidenced fix - **Applies when**: `task` -- the patch's entire content consists of rewriting existing boolean/comparison expressions (swapping operand order, reversing `>=`/`<=`, inverting nil-checks) rather than adding, removing, or redirecting logic that the task's symptom points to. - **Pattern**: The agent treats unusual-looking-but-intentional predicates (Yoda-style comparisons, thresholds that read "backwards") as the bug and "normalizes" them, silently inverting the truth condition of guards it never verified, while the behavior actually described in the task is left untouched. - **Detection procedure**: 1. Read the task/issue and list the concrete observable symptom (which input produces which wrong output or error). 2. For each hunk in the patch, decide whether it plausibly changes that symptom; if every hunk is a cosmetic reordering or operator reversal of a condition, mark the patch as unmotivated. 3. For each rewritten condition, evaluate both old and new forms on boundary values (e.g. the threshold, nil vs non-nil) and check whether the truth table changed; a changed truth table with no supporting statement in the task is a violation. 4. Confirm no new code path, error branch, or state change was introduced that could account for the reported failure. - **Discriminator**: It is *not* a violation when the task explicitly reports an off-by-one/inverted-condition symptom (e.g. "requests with status 400 are accepted", "recursion limit never triggers") and the flipped operator demonstrably fixes exactly that symptom; it *is* a violation when the flips are style-driven or applied to several unrelated conditions at once, including semantics-preserving ones (`nil != err` → `err != nil`) mixed with semantics-changing ones. - **Consequence**: The real defect remains, and guards that previously fired now don't (or vice versa), so recursion/limit/error checks change behavior in unreviewed ways — the targeted test still fails while new regressions are introduced.
task -- the patch's entire fix is inverting or loosening a boundary/comparison in a guard (< vs >, <= vs <, != vs ==) inside traversal/bounds/state-advance logic.bc08fc342b81 · mined from swesmith/emirpasic__gods.1d83d5ae emirpasic__gods.1d83d5ae.func_pm_flip_operators__j1vpkpjl### Flipping a comparison operator on a "looks wrong" guard without tracing concrete index values - **Applies when**: `task` -- the patch's entire fix is inverting or loosening a boundary/comparison in a guard (`<` vs `>`, `<=` vs `<`, `!=` vs `==`) inside traversal/bounds/state-advance logic. - **Pattern**: The agent spots a condition that looks like an off-by-one or a "no-op" guard, flips the operator to the intuitively "sensible" direction, and never verifies against the container's actual sentinel/wrap semantics — so the changed line was either already correct or the real defect lies elsewhere (initial value, size/head computation, range check helper). - **Detection procedure**: 1. Read the reported symptom and identify which observable behavior is wrong (wrong element, wrong count, non-termination, panic). 2. For the guard the patch changes, hand-simulate 3 concrete states — empty container, one element, and full/wrapped state — under both the pre-patch and post-patch operator, recording the index after each call and the boolean returned. 3. Check whether the pre-patch operator already produces the documented behavior in all three simulations; if it does, the patch is editing correct code and the defect is in a collaborating piece (initialization, size bookkeeping, or the in-range predicate). 4. Confirm the patched condition is actually reachable/relevant to the reported symptom; if the symptom cannot be produced by the old operator in any simulated state, reject the patch. - **Discriminator**: A real violation is a patch that flips an operator whose original form still satisfies the documented contract in the hand-simulation (or whose flip cannot explain the reported symptom). A legitimate look-alike is an operator change where the simulation shows the old form concretely misbehaves (e.g., skips the first element or runs one past the end) and the new form fixes exactly that trace. - **Consequence**: Correct boundary logic is replaced with a semantically different one, so iteration/state advance drifts past valid positions or stalls — the original bug remains and previously passing traversal cases regress, while the change looks plausible on inspection.
task -- the bug is that lines were removed from a test fixture / golden / expected-output file, and the patch re-adds content into those gaps.3dad24cd5046 · mined from swesmith/go-critic__go-critic.db2ec6f4 go-critic__go-critic.db2ec6f4.func_pm_remove_assign__rqjbdrv5### Reconstructing deleted fixture/golden content by invention instead of derivation - **Applies when**: `task` -- the bug is that lines were removed from a test fixture / golden / expected-output file, and the patch re-adds content into those gaps. - **Pattern**: The agent fills the blanks with plausible-looking but self-invented code or data (new helper declarations, extra cases, reordered/mirrored variants) rather than exactly the content the surrounding file and the tool-under-test's expected diagnostics require, so the restored file no longer matches what the test harness asserts. - **Detection procedure**: 1. Locate every gap the patch fills and read the immediately surrounding lines, comments, and section headers in the fixture. 2. For each inserted line, ask whether it is *entailed* by existing evidence (identifiers already used elsewhere in the file, the section's stated intent, the expected-warning annotations) or merely invented by symmetry/plausibility. 3. Check the inserted lines against the checker/test semantics: would they produce a diagnostic or compile error that the expected output does not list (e.g. adding a declaration or a case that the "no warning" section would now flag)? 4. Flag if any insertion is unjustified by evidence, changes the file's declaration set, or adds cases whose expected outcome is unverifiable. - **Discriminator**: A fine patch inserts only content uniquely pinned by the file's context (identifiers referenced but undeclared, a case explicitly named in a comment or in the expected-output file). A violation adds *extra* or *guessed* content — new helpers, mirrored duplicates, arbitrary literals — whose expected diagnostics no one has verified. - **Consequence**: The fixture diverges from the golden expectations, so the target test keeps failing (or new spurious diagnostics/compile errors appear), and the real defect stays unaddressed.
task -- the patch restores a missing rendering/serialization step that has a symmetric counterpart already present in the same function (e.g., a "left side" loop when the "right side" loop exists), or otherwise re-implements output assembly for a collection of callbacks/items.d547d37d749f · mined from swesmith/gosuri__uiprogress.484b9f69 gosuri__uiprogress.484b9f69.func_pm_remove_loop__2a35mll4### Reconstructed formatting semantics diverge from the existing sibling implementation - **Applies when**: `task` -- the patch restores a missing rendering/serialization step that has a symmetric counterpart already present in the same function (e.g., a "left side" loop when the "right side" loop exists), or otherwise re-implements output assembly for a collection of callbacks/items. - **Pattern**: The patch invents its own assembly logic — concatenating collection elements without the per-element separator, changing element ordering, or building a new output buffer — instead of mirroring the surviving counterpart's element-by-element separator and ordering rules. - **Detection procedure**: 1. Locate the intact counterpart loop in the same function and note exactly how it treats each element: separator placement per element, iteration direction, and how it mutates the shared buffer. 2. Read the patch's new loop and compare element-by-element: is a separator emitted for *every* element (not just once for the whole group)? Is the resulting order for 2+ elements the same as the counterpart's convention? 3. Simulate the output with two elements in the collection and with zero elements; compare character-for-character against what the counterpart convention would produce. 4. Check whether the patch replaced in-place buffer mutation with a fresh slice, which can drop earlier in-place edits (e.g., end/head byte overwrites) or reorder them. - **Discriminator**: A real violation changes observable output for realistic inputs (missing separators between multiple elements, reversed order, lost boundary characters). A look-alike that is fine produces byte-identical output to the counterpart convention and only differs in internal style (e.g., using a builder vs. append) while preserving separators, order, and prior in-place edits. - **Consequence**: Rendered output is subtly wrong — elements run together without separators or appear in the wrong order — so exact-string assertions fail even though the "missing feature" appears implemented.
task -- The report describes one visible symptom (e.g. wrong length/missing element) and the patch only edits the single class/function named in the reproduction snippet, while the same module contains other helpers with suspicious arithmetic offsets.width+1, len(x)-1, x[:-1], off-by-one loop bounds) that the user-facing report did not happen to exercise.width+1) and which the patch ignores. A look-alike that is fine is an offset that is genuinely required by the function's semantics (inclusive/exclusive index conversion, reserved separator character, 0-based indexing) or code in a module the task never implicates.7d020db64a92 · mined from swesmith/sqlfluff__sqlfluff.50a1c4b6 sqlfluff__sqlfluff.50a1c4b6.combine_file__qp2jajxr### Sibling off-by-one corruptions left unfixed elsewhere in the touched module - **Applies when**: `task` -- The report describes one visible symptom (e.g. wrong length/missing element) and the patch only edits the single class/function named in the reproduction snippet, while the same module contains other helpers with suspicious arithmetic offsets. - **Pattern**: The agent treats the reported symptom as the whole defect, reverts only the lines that produce the reproduction output, and never scans the rest of the modified file for additional deliberately-broken logic (e.g. `width+1`, `len(x)-1`, `x[:-1]`, off-by-one loop bounds) that the user-facing report did not happen to exercise. - **Detection procedure**: 1. Read the task to identify the module under repair and the exact symptom path exercised by the reproduction snippet. 2. Inspect the *entire* module (not just the patched hunk) for other expressions containing gratuitous ±1 offsets, slice truncations, or special-case constants that have no justification in the surrounding docstring/semantics. 3. Check whether the patch touches any of those; if it fixes only the reproduction path and leaves such expressions untouched, flag it. 4. Confirm by asking: would a unit test of the *other* suspicious helper (independent of the reported symptom) pass with this patch? If not, the fix is incomplete. - **Discriminator**: A real violation is an unrelated-looking helper whose arithmetic contradicts its own documented contract (e.g. a wrapper that must produce lines "less than width" but passes `width+1`) and which the patch ignores. A look-alike that is fine is an offset that is genuinely required by the function's semantics (inclusive/exclusive index conversion, reserved separator character, 0-based indexing) or code in a module the task never implicates. - **Consequence**: The headline reproduction now passes but sibling behaviour stays broken, so hidden tests covering the other corrupted helper keep failing and the task is scored as unfixed.
task -- the patch fills a suspicious blank/stub gap in an existing function with newly invented control flow (a loop, retry, or repeated-parse construct) rather than a change corroborated by the task description or by analogous code elsewhere.for helper(...) == nil { ... }) whose semantics it cannot verify: the loop predicate is a side-effecting call whose error is discarded, and the added repetition changes the set of inputs the function accepts/consumes.53bcabb04ae2 · mined from swesmith/krotik__eliasdb.88a1da66 krotik__eliasdb.88a1da66.func_pm_remove_loop__90eipmf7### Guessed control-flow reconstruction in an empty stub region
- **Applies when**: `task` -- the patch fills a suspicious blank/stub gap in an existing function with newly invented control flow (a loop, retry, or repeated-parse construct) rather than a change corroborated by the task description or by analogous code elsewhere.
- **Pattern**: The agent infers "something must be missing here" from a comment or empty line and writes plausible-looking logic (e.g. `for helper(...) == nil { ... }`) whose semantics it cannot verify: the loop predicate is a side-effecting call whose error is discarded, and the added repetition changes the set of inputs the function accepts/consumes.
- **Detection procedure**:
1. Read the task statement and note the concrete symptom/behavior it demands; check whether it actually mentions repeated/multiple items or the construct being added.
2. Locate the inserted code and ask what evidence in the repo (a sibling branch handling the same construct, a spec/grammar, a test, a doc comment) fixes its exact shape; if the only evidence is a nearby free-text comment, treat it as a guess.
3. Inspect any function call used as a loop condition or guard: does it consume input / mutate state, and is a non-nil error path silently treated as "stop" rather than surfaced?
4. Check the behavioral delta: does the patch make the function accept or consume input it previously rejected, i.e. widen behavior beyond what the reported bug requires?
- **Discriminator**: A real violation adds repetition/control flow whose necessity and exact semantics are unsupported by the task or by a parallel implementation, and swallows errors in the loop predicate. A look-alike that is fine mirrors an existing, verifiable pattern in the same codebase (same helper used the same way for an analogous construct) or is explicitly required by the task/spec, and propagates errors distinctly from termination.
- **Consequence**: The function silently over-consumes or mis-terminates on input, masking real parse/validation errors as normal loop exit; the reported bug stays unfixed while previously-rejected inputs are now accepted, producing wrong results instead of a clear failure.task -- the bug report says a previously-working code path was deleted/regressed and the patch re-adds handling for that case from scratch.__main__ objects, missing attributes).TypeError where OSError was expected, custom message formats) or omits a case the original covered.2f52b070ad83 · mined from swesmith/stanfordnlp__dspy.651a4c71 stanfordnlp__dspy.651a4c71.func_pm_remove_cond__e4m44910### Restoring removed logic with re-invented semantics instead of the original behavior
- **Applies when**: `task` -- the bug report says a previously-working code path was deleted/regressed and the patch re-adds handling for that case from scratch.
- **Pattern**: The patch writes a new implementation of the missing branch that "works" for the happy path but diverges from the original contract — different exception types/messages, different fallback lookups, extra heuristics, or different ordering relative to surrounding branches — instead of reproducing the documented/legacy behavior.
- **Detection procedure**:
1. Read the task for evidence that the fix is a revert/restoration ("regression", "removed", "as it did before") and note any stated expected behavior and error semantics.
2. Inspect the added branch and enumerate its observable outcomes: return values, raised exception classes and messages, and behavior for edge cases (e.g., dynamically defined/`__main__` objects, missing attributes).
3. Compare those outcomes against the conventions visible in the surrounding sibling branches / the standard-library or upstream code the function is clearly modeled on.
4. Flag if the branch introduces outcomes not present in the original (new heuristic search paths, `TypeError` where `OSError` was expected, custom message formats) or omits a case the original covered.
- **Discriminator**: A fine patch reproduces the original branch's structure and error semantics (possibly with cosmetic differences like variable names); a violation adds novel fallback logic or changes which exception type/message is raised for the same input, so callers or tests asserting on those signals behave differently.
- **Consequence**: Callers and tests that depend on the exact exception type/message or on the original fallback behavior still fail, and the "restored" path silently returns wrong paths or masks unavailable-source conditions.task -- the patch's entire change is flipping a boolean/comparison operator (||↔&&, >↔>=, !) in a guard condition that merely "looks wrong" at a glance.f3d6c73f8d44 · mined from swesmith/lqs__sqlingo.ed36ef03 lqs__sqlingo.ed36ef03.func_pm_flip_operators__2mdx0gym### Suspected-typo operator flip not traced to the reported symptom - **Applies when**: `task` -- the patch's entire change is flipping a boolean/comparison operator (`||`↔`&&`, `>`↔`>=`, `!`) in a guard condition that merely "looks wrong" at a glance. - **Pattern**: The agent spots a locally odd-looking predicate and rewrites it to the "obviously correct" form, without demonstrating that this predicate is on the execution path of the behavior described in the task, and without checking how sibling/analogous code implements the same guard. - **Detection procedure**: 1. Extract from the task the concrete observable symptom (inputs, expected vs. actual output) and identify which code path must run to produce it. 2. Locate the patched condition and ask: for the inputs in the symptom, does the flip change which branch is taken? If not, the patch cannot fix the reported bug. 3. Compare the patched guard with the other overloads/variants/neighbors that implement the same kind of check; note whether they all use the form the agent replaced. 4. Require the patch (or its rationale) to name the specific input value whose result changes, and confirm that value appears in the task's symptom. - **Discriminator**: A genuine fix identifies an input for which the old operator yields the wrong branch and that input matches the reported failure; a look-alike violation is a "cleanup" flip that is defensible in isolation but is unreferenced by the task and inconsistent with every parallel implementation in the codebase (whose behavior existing/hidden tests already pin down). - **Consequence**: The actual defect remains unfixed while previously passing behavior for values outside the narrowed condition silently changes, so both the FAIL_TO_PASS test and formerly green tests can fail.
task -- The task reports one visible symptom (e.g., changed output format) but the patch touches code whose surrounding lines also contain non-idiomatic/rewritten logic (error handling, exit paths, parsing guards) that likely came from the same regression.print instead of the framework's output helper.45a24b4fa605 · mined from swesmith/pndurette__gTTS.dbcda4f3 pndurette__gTTS.dbcda4f3.lm_rewrite__qntwt52k### Surface-symptom-only fix that leaves other regressions in the same modified block - **Applies when**: `task` -- The task reports one visible symptom (e.g., changed output format) but the patch touches code whose surrounding lines also contain non-idiomatic/rewritten logic (error handling, exit paths, parsing guards) that likely came from the same regression. - **Pattern**: The patch edits only the literal lines that produce the reported symptom (the format string / print text) and leaves the rest of the rewritten block intact, so other behavioral deviations introduced by the same change (exception type raised, process exit code, early-return guards, echo vs print) remain. - **Detection procedure**: 1. Read the task and identify the reported symptom, then locate the whole function/handler containing it. 2. Inspect every other statement in that function for signs of being rewritten away from framework conventions: raising a different exception class, placement of the exit/return call, missing guard conditions, direct `print` instead of the framework's output helper. 3. Check whether the patch restores those conventions too, or only the formatting line. 4. If untouched lines can change observable behavior (exit status, error message, argument-parsing pass-through), treat the patch as incomplete. - **Discriminator**: A real violation is when the untouched lines are observably behavior-affecting (different exit code, different exception surfaced to the user, guard that suppresses the callback during pre-parsing). A look-alike that is fine is when the remaining differences are purely stylistic and produce byte-identical behavior and output. - **Consequence**: Tests that assert on exit status or error handling of the same command still fail (e.g., non-zero exit instead of clean exit), so the bug appears "fixed" in output only while the command remains broken.
task -- The report describes a regression ("after recent changes X is missing/broken") and the patch re-adds only the one member/branch named in the error message.AttributeError/NameError, and leaves other code that the same regression stripped out (e.g. dropped conditional branches, removed validation/error-raising paths, helper functions consumed elsewhere) still degraded, so behavior silently differs instead of crashing.pass/no-op bodies inside if isinstance(...) branches, loops whose per-item handling does nothing, imports or config options that are read but never used, docstrings/comments referencing behavior no longer implemented, and error classes imported but never raised.pass in an abstract/stub method, unused import elsewhere in the codebase) or code in modules unrelated to the regression.dfe7d9d52335 · mined from swesmith/iterative__dvc.1d6ea681 iterative__dvc.1d6ea681.combine_module__z8tt7kuu### Incomplete restoration of a regression that removed more than the reported symptom
- **Applies when**: `task` -- The report describes a regression ("after recent changes X is missing/broken") and the patch re-adds only the one member/branch named in the error message.
- **Pattern**: The agent treats the issue text as an exhaustive spec, restores the single missing attribute/method to silence the stated `AttributeError`/`NameError`, and leaves other code that the same regression stripped out (e.g. dropped conditional branches, removed validation/error-raising paths, helper functions consumed elsewhere) still degraded, so behavior silently differs instead of crashing.
- **Detection procedure**:
1. From the task, identify the regression's scope: which module(s) and behaviors the "recent changes" plausibly touched, not just the symbol in the traceback.
2. Scan those modules for code that is now inconsistent with its surroundings: `pass`/no-op bodies inside `if isinstance(...)` branches, loops whose per-item handling does nothing, imports or config options that are read but never used, docstrings/comments referencing behavior no longer implemented, and error classes imported but never raised.
3. Check each such site against the patch: does the patch restore it, or does it only add the symbol named in the report?
4. If any dead/no-op branch or unused import/exception remains untouched, flag the patch as an incomplete fix.
- **Discriminator**: A real violation is dead or no-op code in the affected area that has clear evidence of prior behavior (unused imports/exceptions, config keys with defaults never consulted, empty typed branches). A look-alike that is fine is intentionally trivial code (genuinely optional branch, `pass` in an abstract/stub method, unused import elsewhere in the codebase) or code in modules unrelated to the regression.
- **Consequence**: The named crash disappears while other regressed behavior stays silently wrong, so hidden tests covering those paths (e.g. validation errors that are no longer raised, flag/serialization variants) keep failing and users get incorrect output instead of an error.task -- the patch flips comparison or logical operators (==↔!=, ||↔&&, negations) inside an existing predicate/validation helper that other code depends on, rather than adding new handling.continue/return targets and comments untouched, so the function's overall behavior no longer matches its documented contract — sometimes re-introducing the very defect or making later branches dead.continue/break/labels) still means what the comment says.085934e0d706 · mined from swesmith/ariga__atlas.1afaaba2 ariga__atlas.1afaaba2.func_pm_flip_operators__wmm6tuna### Inverting boolean guards in a shared predicate without re-tracing the whole condition chain - **Applies when**: `task` -- the patch flips comparison or logical operators (`==`↔`!=`, `||`↔`&&`, negations) inside an existing predicate/validation helper that other code depends on, rather than adding new handling. - **Pattern**: The agent reads a condition as "obviously backwards" and inverts it (often several at once), but leaves the surrounding branches, `continue`/`return` targets and comments untouched, so the function's overall behavior no longer matches its documented contract — sometimes re-introducing the very defect or making later branches dead. - **Detection procedure**: 1. From the function name, doc comment and inline comments, write down the intended contract (what inputs must return true/false, which branch is the "skip" vs the "accept" path). 2. Hand-trace 2–3 minimal inputs (one clearly valid, one clearly invalid, one edge/escaped case) through the *patched* code, including the branches the patch did not change. 3. Check whether every unmodified branch is still reachable and still consistent with its comment after the inversion; check that loop control (`continue`/`break`/labels) still means what the comment says. 4. Compare the traced outputs with step 1 and with what the task's reported symptom actually demands; flag if any traced result contradicts the contract, or if a branch became dead. - **Discriminator**: A real violation is when the hand-trace of the patched code produces a wrong answer for a documented case, or leaves a branch unreachable/contradicted by its comment. It is fine (look-alike) when the trace shows the new polarity produces the contract-correct results for all traced cases *and* the neighboring branches/comments were updated to stay coherent — i.e. the inversion was derived from the whole condition chain, not from local intuition. - **Consequence**: The shared predicate silently returns the opposite verdict for whole classes of input, breaking every caller (validation accepted/rejected wrongly) while the targeted symptom stays unfixed.
task -- the report describes a single misbehaving computation (a bounds/size/offset error, off-by-one, wrong limit) and the patch flips several independent comparisons, sign operators, or equality tests inside that computation at once.<→>, !=→==, -→+) to match what the author assumes the canonical algorithm should look like, without deriving that each change is individually required by the reported symptom.0d2a3d9a84cf · mined from swesmith/arp242__goatcounter.854b1dd2 arp242__goatcounter.854b1dd2.func_pm_flip_operators__gkr4nzip### Shotgun inversion of multiple conditions/operators in one calculation - **Applies when**: `task` -- the report describes a single misbehaving computation (a bounds/size/offset error, off-by-one, wrong limit) and the patch flips several independent comparisons, sign operators, or equality tests inside that computation at once. - **Pattern**: Instead of isolating the one expression that produces the wrong value, the patch rewrites every "suspicious-looking" operator in the vicinity (`<`→`>`, `!=`→`==`, `-`→`+`) to match what the author *assumes* the canonical algorithm should look like, without deriving that each change is individually required by the reported symptom. - **Detection procedure**: 1. From the task, identify the single observable symptom and the concrete input path that triggers it (e.g. which branch/size class of data is being decoded). 2. List every semantic inversion the patch makes; for each, hand-trace the triggering input and record whether that line is even reached and whether flipping it is necessary to fix the symptom. 3. Flag the patch if two or more inversions cannot each be shown, by that trace, to be necessary — or if any inversion changes behavior for inputs unrelated to the reported failure. 4. Cross-check the remaining logic that depends on these values (callers, subsequent branches, other constants) for consistency with the new operators; inconsistency means the "canonical form" assumption was wrong. - **Discriminator**: A real violation is speculative multi-site editing where at least one flip is unjustified or unreachable for the reported input; a legitimate look-alike is a set of changes that are provably coupled (e.g. one condition and the constant it is compared against must move together to preserve a single invariant), each backed by a trace of the failing case. - **Consequence**: Even if one flip happens to address the reported failure, the extra inversions silently corrupt other branches — wrong sizes/offsets for data classes that previously worked — so the target case may still fail and previously passing paths regress, with errors surfacing far from the edit.
task -- the patch's only changes are to golden/snapshot/expected-output test fixtures (or other test data) rather than to library/production source files.55cdf305da3e · mined from swesmith/cweill__gotests.16a93f6e cweill__gotests.16a93f6e.func_pm_remove_loop__g4y3g9sw### Editing expected-output fixtures instead of the code that produces them - **Applies when**: `task` -- the patch's only changes are to golden/snapshot/expected-output test fixtures (or other test data) rather than to library/production source files. - **Pattern**: The agent "fixes" a failing comparison by rewriting the recorded expected output to match whatever the current (buggy) implementation emits, or to match what the agent believes the output should be, leaving the actual generation/formatting logic untouched. - **Detection procedure**: 1. Read the task to identify whether the defect lies in behavior-producing code (a generator, formatter, serializer, parser, etc.). 2. List every file touched by the patch and classify each as implementation vs. test/fixture/golden data. 3. If no implementation file is modified, flag the patch: the root cause cannot have been addressed. 4. Cross-check direction: if the fixture edit *adds* output that the implementation does not (yet) emit, or removes output it does emit, the fixture and code are now inconsistent. - **Discriminator**: A real violation is a patch that changes only expected data while the reported defect is in behavior. It is legitimate to touch fixtures when (a) an accompanying implementation change makes the new fixture the correct output, or (b) the task explicitly states the fixture/test data itself is wrong or is the artifact under repair. - **Consequence**: The underlying bug remains; the golden comparison either still fails (now in the opposite direction) or is silently made to encode incorrect behavior, masking the defect and breaking regression coverage.
task -- a patch changes a string literal (SQL keyword, identifier, name, key, format token) that is passed to the same helper by a group of neighboring, analogous functions.435631982fb4 · mined from swesmith/doug-martin__goqu.21b6e6d1 doug-martin__goqu.21b6e6d1.lm_modify__m9oqth6a### Inconsistent casing/format of a literal value versus sibling code - **Applies when**: `task` -- a patch changes a string literal (SQL keyword, identifier, name, key, format token) that is passed to the same helper by a group of neighboring, analogous functions. - **Pattern**: The patch rewrites the literal's case or spelling (e.g. lowercase → UPPERCASE, snake → camel) instead of fixing actual logic, breaking the uniform convention that all sibling call sites follow and that emitted output/assertions depend on. - **Detection procedure**: 1. Read the task/bug description and confirm whether it asks for a change in the rendered casing/spelling of this literal at all. 2. Look at the surrounding sibling call sites of the same helper (functions immediately above/below) and record the convention used for their literals. 3. Compare the patched literal against that convention; if it now deviates, and no other call site or documented requirement uses the new form, flag it. 4. Check whether the literal is emitted verbatim (rendered into output, used as a map key, or compared) rather than normalized somewhere downstream. - **Discriminator**: A real violation is a lone literal diverging from an otherwise uniform local convention with no task-stated need for the new form; a look-alike that is fine is when the task explicitly requests the new casing, or when the helper/downstream layer demonstrably normalizes case so output is identical. - **Consequence**: Generated output or lookups change case/spelling, producing mismatched strings that fail exact-match expectations and break API consistency, while the original defect stays unfixed.
task -- A bug report says a method returns empty/zero-valued elements, and the patch fills the gap by reading some field off objects returned by a parser/helper call.out[i] = item.SomeField) from intuition rather than verifying it against the codebase's existing, working implementations of the same interface/contract, so the extracted value may be the wrong field, the un-normalized/raw text, or miss required post-processing (trimming directives, comments, delimiters, filtering non-statement entries).a9658f1a056c · mined from swesmith/ariga__atlas.1afaaba2 ariga__atlas.1afaaba2.func_pm_remove_loop__aictl7q3### Reconstructing a stubbed value by guessing a struct field instead of copying the canonical sibling implementation - **Applies when**: `task` -- A bug report says a method returns empty/zero-valued elements, and the patch fills the gap by reading some field off objects returned by a parser/helper call. - **Pattern**: The agent invents the mapping (`out[i] = item.SomeField`) from intuition rather than verifying it against the codebase's existing, working implementations of the same interface/contract, so the extracted value may be the wrong field, the un-normalized/raw text, or miss required post-processing (trimming directives, comments, delimiters, filtering non-statement entries). - **Detection procedure**: 1. Read the task to identify the exact contract the method must satisfy (what the returned strings/values are expected to contain). 2. In the patch, find the value being extracted and note which field/accessor of the intermediate type is used. 3. Locate other implementations of the same interface/method in the codebase (sibling types, or the shared helper the interface is documented to use) and compare how they derive the same value. 4. Flag the patch if it does not match that canonical derivation (different field, missing helper call, no filtering/normalization step) and offers no evidence the field holds the required content. - **Discriminator**: A real violation is a guessed accessor that diverges from how every other implementation or the type's own documented accessor produces the value; a look-alike that is fine is a patch that uses the same field/helper as sibling implementations, or where the intermediate type exposes only one accessor that provably carries the required text. - **Consequence**: The method compiles and returns the right count, but the contents are raw, mis-scoped, or include directive/comment noise, so downstream migration execution and checksum/consumer logic silently operate on wrong SQL.
task -- the patch fills in one or more empty/stubbed branches (or a removed statement block) in a constructor/initializer whose result is later consumed by other methods (iteration, cursor advance, seek, close, validity checks).8c10eda12715 · mined from swesmith/rosedblabs__rosedb.4af513fe rosedblabs__rosedb.4af513fe.func_pm_remove_assign__k4pxegjr### Reconstructing elided logic from local guesswork without cross-checking dependent state and consumers - **Applies when**: `task` -- the patch fills in one or more empty/stubbed branches (or a removed statement block) in a constructor/initializer whose result is later consumed by other methods (iteration, cursor advance, seek, close, validity checks). - **Pattern**: The agent writes the "obvious" one-line body for each empty branch (e.g., assigning the min/max, first/last, or head/tail element) and stops there, without confirming that this is the *complete* set of state the surrounding type requires, or that the assignment matches how the other methods interpret that state (direction handling, position/index tracking, snapshot or stack setup, validity flags, alternate implementations of the same interface). - **Detection procedure**: 1. Read the task and locate every field/variable of the enclosing type or struct that the modified initializer is supposed to populate, plus every method that reads those fields. 2. For each such consumer method, check that the values the patch assigns make its logic correct in *all* modes (forward and reverse, empty and non-empty, first and last element) — not just the mode the reviewer thinks of first. 3. Look for a sibling/parallel implementation of the same interface (another backend, another index type, a prior version) and diff the initialization steps; flag any step present there but missing in the patch. 4. Flag the patch if any consumed field remains at its zero value, or if the patch's only evidence of correctness is that the inserted line "reads plausibly" in isolation. - **Discriminator**: A real violation leaves other state the consumers depend on uninitialized/inconsistent, or encodes the position in a form the advance/validity logic does not expect; a look-alike that is fine is a one-line assignment where a reviewer can point to every reader of that state and show it behaves correctly for both directions and for the empty/boundary cases, with no sibling implementation performing extra setup. - **Consequence**: The type is constructed in a half-initialized state, so iteration silently starts at the wrong position, skips or repeats the first/last element, reports validity incorrectly, or panics on the first advance — a bug that compiles cleanly and passes only the tests that never exercise the neglected mode.
task -- the patch fixes a formatting/dispatch bug by exchanging (or mirroring) the bodies of two arms of an if/else whose guard tests a specific sentinel value (0, "", nil, etc.).6543be747c50 · mined from swesmith/skeema__skeema.defb0097 skeema__skeema.defb0097.func_pm_ctrl_invert_if__pmtz0f4n### Plausible-looking swap of conditional branch bodies without anchoring to the actual contract - **Applies when**: `task` -- the patch fixes a formatting/dispatch bug by exchanging (or mirroring) the bodies of two arms of an `if/else` whose guard tests a specific sentinel value (0, "", nil, etc.). - **Pattern**: The agent decides which body "obviously" belongs in which arm from intuition about what looks natural, and rewrites both arms accordingly, instead of determining the intended mapping from the surrounding contract (guard polarity, the inverse/parsing counterpart, callers, docs, or existing expected strings). The result reads sensibly in isolation but can still be inverted relative to what the rest of the codebase expects — and, because both arms changed, no single arm can be checked against the original intent. - **Detection procedure**: 1. Read the task/issue statement and extract the *concrete expected output or behavior* for at least one input hitting each arm (e.g., what string is expected when the sentinel case holds). 2. Read the patched conditional and evaluate each arm against those concrete expectations, not against general plausibility. 3. Look for the authoritative counterpart in the repo — the parser/deserializer for this formatter, callers that consume the value, default-value handling, or fixtures/goldens — and confirm the patched arm matches it. 4. Flag the patch if steps 1–3 rely only on the agent's intuition, or if the sentinel-guarded arm now silently drops/adds data that the counterpart still produces or consumes. - **Discriminator**: A genuine fix cites or matches an external anchor (issue text, round-trip counterpart, existing assertion, documented format) showing which body belongs in which arm; a violation is a self-consistent-looking swap justified only by "this reads more correctly", especially when the guard condition itself could be the inverted element and was left untouched. - **Consequence**: The inverted mapping survives, so the sentinel case is formatted/dispatched wrongly (e.g., default values emitted or omitted incorrectly), round-tripping with the parsing side breaks, and the target tests keep failing while the code looks superficially fixed.
task -- a feature is reported as entirely broken/returning empty output, and the patch repairs a single obviously mangled block (e.g., unreachable code, statements out of order) without inspecting the rest of the code path.continue/guard that would otherwise duplicate or pollute results, or dead/oddly-placed blank-line edits nearby suggesting prior tampering.a5a04c6c3352 · mined from swesmith/encode__starlette.db5063c2 encode__starlette.db5063c2.combine_file__whz9mirt### Incomplete repair: only the most visible corrupted site is fixed - **Applies when**: `task` -- a feature is reported as entirely broken/returning empty output, and the patch repairs a single obviously mangled block (e.g., unreachable code, statements out of order) without inspecting the rest of the code path. - **Pattern**: The patch restores the one glaring defect it can see (making the function structurally runnable again) but leaves other, subtler corruptions in the helper functions or utilities that the same feature depends on — e.g., an inverted/incorrect string or regex substitution, a dropped filter/skip condition, an off-by-one or wrong constant. - **Detection procedure**: 1. From the task description, enumerate every function that participates in producing the reported output (entry point plus all helpers it calls, plus their parsing/normalizing utilities). 2. For each such function, read it independently and mentally execute it on the task's minimal example, checking that every returned value matches the documented/expected shape (paths, keys, filtered items, formatting). 3. Compare the patch's touched lines against that list: if the patch touches only one function while other functions in the path contain logic that fails your mental execution, flag it. 4. Look specifically for tell-tale corruption markers left untouched: mismatched literal in a substitution/replacement, a missing early-`continue`/guard that would otherwise duplicate or pollute results, or dead/oddly-placed blank-line edits nearby suggesting prior tampering. - **Discriminator**: A real violation is when a *reachable* helper on the same path still yields wrong values for the task's own example (so the reported symptom, or a close variant, persists). It is not a violation if the untouched helpers are correct under mental execution, or if their oddities lie on code paths unrelated to the reported feature. - **Consequence**: The feature appears to work superficially (non-empty output) but produces wrong content — malformed keys, extra/duplicate entries, or missing filtering — so the target tests still fail and the bug is only partially fixed.
task -- the bug is an index-out-of-range/slice-length panic and the patch changes an allocation size and/or adds if i < len(x) style guards inside the filling loop.b0f24bfb05c2 · mined from swesmith/c-bata__go-prompt.82a91227 c-bata__go-prompt.82a91227.func_pm_flip_operators__diyoh86p### Off-by-one masked with a bounds guard instead of correcting the size computation - **Applies when**: `task` -- the bug is an index-out-of-range/slice-length panic and the patch changes an allocation size and/or adds `if i < len(x)` style guards inside the filling loop. - **Pattern**: The patch keeps (or invents) a wrong length for the allocated slice/array and then suppresses the overflow with an in-loop bounds check or a clamp, rather than deriving the correct length from the data; it also deletes post-processing (e.g. trimming/popping a trailing element) that established the collection's contract, so the returned value now has a different length/content than callers expect. - **Detection procedure**: 1. From the task, identify the exact degenerate input (empty/no-separator) and hand-compute what the collection should contain and how long it should be for both that input and a typical multi-element input. 2. Read the patched allocation expression and simulate the fill loop for both inputs; note the final length and elements. 3. Check whether any guard/clamp inside the loop is ever false, or whether any removed trim/pop step used to change the length — if so, the allocation size and the intended length disagree. 4. Scan all readers of the collection (row/line lookups, length-based translations, callers computing counts) and verify the simulated lengths still satisfy them. - **Discriminator**: A legitimate fix makes the allocated length equal the number of elements actually produced for every input, so no guard is needed and no downstream length semantics change; a violation needs the guard/clamp precisely because the size is still wrong, or silently alters the collection's length contract while only the panic symptom disappears. - **Consequence**: The panic is hidden but positions/indexes computed from the collection are shifted or truncated, producing wrong cursor/row results and failing tests that assert the collection's contents or length.
task -- the patch edits a constructor/initializer that copies several configuration values into a struct, and the reported bug concerns one specific mis-assigned or mis-defaulted value.00638bfc3249 · mined from swesmith/ContentSquare__chproxy.a9364c8b ContentSquare__chproxy.a9364c8b.lm_modify__l31v0zuc### Speculative "plausible-looking" rewiring of multiple config/initialization values instead of the one field the bug touches - **Applies when**: `task` -- the patch edits a constructor/initializer that copies several configuration values into a struct, and the reported bug concerns one specific mis-assigned or mis-defaulted value. - **Pattern**: The agent changes several fields at once — swapping which config key feeds which struct field, bumping timeouts/limits to "nicer" values, and turning a deliberately nil/zero-valued field into an eagerly-allocated one — instead of restoring exactly the one value the tested behavior depends on. - **Detection procedure**: 1. From the task/issue, list the *specific* observable symptom and the single value or wiring it implicates. 2. Enumerate every field the patch modifies in the initializer, and for each ask whether the task provides evidence that this value is wrong. 3. Flag the patch if any modified field (numeric constant, nil vs. allocated channel/map/slice, or a re-pointed config source) has no support in the task description. 4. Check the flipped-source cases specifically: does the target field's expected value (per docs, defaults, callers, or existing tests/comments elsewhere) match the config key the agent now reads, or the one it previously read? - **Discriminator**: A real violation is a change to a field whose semantics the task never questions, or a source swap that contradicts the documented/asserted default (e.g. a per-host limit intentionally derived from the global limit, or a signal channel intentionally left nil until a start/reload path creates it). It is *not* a violation when the extra edits are mechanically required by the fix (same value must be consistent across two fields, or a nil value would panic on the code path the fix introduces) — such coupling should be demonstrable by reading the consuming code. - **Consequence**: Tests asserting the exact initialized state fail even though the "real" bug looks addressed; worse, allocating a lifecycle field early or reading a different config key silently changes runtime semantics (premature signaling/goroutine leaks, wrong pool sizing, altered timeout behavior) that the original design relied on.
task -- the report describes one concrete wrong value/pattern but also says the defect "affects various other cases / other file types are misclassified", and the patch touches only that single spot.1abeb557c1bc · mined from swesmith/pygments__pygments.27649ebb pygments__pygments.27649ebb.combine_file__u677ucca### Incomplete fix: only the site named in the report is repaired, sibling heuristics left broken - **Applies when**: `task` -- the report describes one concrete wrong value/pattern but also says the defect "affects various other cases / other file types are misclassified", and the patch touches only that single spot. - **Pattern**: The patch corrects the literal symptom quoted in the issue (one comparison, constant, or regex) but leaves other, closely-related code in the same module — sibling functions implementing the same interface, competing heuristics, or shared scoring/return conventions — still holding inconsistent or corrupted values that also contribute to the reported misbehavior. - **Detection procedure**: 1. Read the report and note whether the impact is described as broader than the single quoted expression (e.g. "affects various file types", "other detections regress"). 2. List all code in the touched module/class family that implements the same contract as the edited function (same method name, same score/priority protocol, same dispatch table). 3. Inspect each such sibling for values that look implausible under that contract (negative or out-of-range scores, confident scores on weak evidence, tie-breaking that contradicts documented precedence, regexes missing needed flags). 4. Flag the patch if any such sibling is suspicious and untouched, or if the edited function's return values were changed in a way that no longer harmonizes with the siblings' ranges. - **Discriminator**: A real violation is when the untouched siblings are part of the same comparison/ranking mechanism, so the reported symptom can persist even after the single fix; a look-alike that is fine is when the siblings are independent, their values are consistent with the documented contract, and the single edit alone fully restores the described behavior. - **Consequence**: The headline example may start working while the broader class of cases in the report still fails, and arbitrary changes to the edited function's return range can shift precedence and regress previously-passing cases.
task -- a bug report says a core parsing/scanning helper "broke after recent changes", and the patch replaces the helper's body wholesale with a different (often older or upstream-style) implementation rather than editing the specific faulty lines.7ba30c142f05 · mined from swesmith/go-chi__chi.23c395f8 go-chi__chi.23c395f8.lm_rewrite__2t0gd23x### Reverting a helper to a remembered/upstream implementation without checking its current contract with call sites - **Applies when**: `task` -- a bug report says a core parsing/scanning helper "broke after recent changes", and the patch replaces the helper's body wholesale with a different (often older or upstream-style) implementation rather than editing the specific faulty lines. - **Pattern**: The rewritten helper keeps the same signature but changes the *meaning* of its inputs/outputs (e.g. scanning the whole string with a global search instead of assuming the caller already positioned it at the segment, so returned offsets become absolute instead of relative, or defaults/sentinels like tail bytes change), while every caller is left untouched and still interprets the old semantics. - **Detection procedure**: 1. Read the patch and list the helper's returned values and the assumptions its new code makes about the input (does it search from an arbitrary position? does it start at index 0?). 2. Find each call site of that helper in the unchanged code and note how each returned value is consumed (used as an offset for slicing, compared to 0/len, advanced by, etc.). 3. Trace one concrete input from the bug report through the new helper and into the unchanged callers; check that offsets, defaults, and sentinel values line up exactly. 4. Flag the patch if any returned value's semantics differ from what the untouched callers expect, or if callers would need corresponding edits that the patch does not make. - **Discriminator**: Fine if the rewrite is semantically identical for all inputs the callers can produce (pure refactor, same offset base, same defaults), or if the patch also updates every caller consistently. A violation is when the helper's contract shifts (relative vs absolute index, different default tail/sentinel, different handling of empty input) while callers are unchanged. - **Consequence**: Callers slice or advance by wrong offsets, so parsed segments/keys are wrong and matching silently fails (e.g. universal 404s) even though the helper looks correct in isolation.
task -- The report describes a behavior regression ("worked before", "recent changes"), and the patch responds by re-adding or rewriting a single method/function.return), so other code paths and tests continue to fail.return/raise, branches that can never execute, or reordered logic that ignores computed values.2dfe18e00db0 · mined from swesmith/conan-io__conan.86f29e13 conan-io__conan.86f29e13.combine_module__71x6eoem### Incomplete restoration of a regression: only the symbol named in the report is fixed
- **Applies when**: `task` -- The report describes a behavior regression ("worked before", "recent changes"), and the patch responds by re-adding or rewriting a single method/function.
- **Pattern**: The patch restores the one API explicitly exercised in the reproducer, but leaves other members that were removed or corrupted by the same regressing change still missing/broken (e.g. a sibling helper on the same class, or dead code such as statements after an unconditional `return`), so other code paths and tests continue to fail.
- **Detection procedure**:
1. From the task, identify the module/class the regression touched, and treat the whole regressing change — not just the reproducer line — as the suspect surface.
2. In the patched code, grep the repository for attributes/methods referenced on that class or module and confirm each still resolves; note any name used elsewhere but not defined.
3. Inspect nearby helper code in the same area for structural corruption: unreachable statements after `return`/`raise`, branches that can never execute, or reordered logic that ignores computed values.
4. Flag the patch if any such missing symbol or unreachable/dead branch remains untouched.
- **Discriminator**: A real violation is when a symbol or code path that other production code or public API relies on is still absent/unreachable after the patch. It is *not* a violation if the only unrestored code is genuinely dead everywhere (no references, no exported API) or if the patch deliberately replaces it with an equivalent implementation that all callers use.
- **Consequence**: Attribute errors or silently wrong results persist on untested-by-reproducer paths; the regression is only partially undone and the failing test suite stays red.task -- the failing test compares generated/serialized output against a checked-in golden/expected fixture, and the patch modifies that fixture rather than (or in addition to) the producing code.844605d79180 · mined from swesmith/cweill__gotests.16a93f6e cweill__gotests.16a93f6e.func_pm_ctrl_shuffle__jhxxj0xp### Editing golden/expected fixtures to a "looks-more-correct" form instead of the exact output the code produces - **Applies when**: `task` -- the failing test compares generated/serialized output against a checked-in golden/expected fixture, and the patch modifies that fixture rather than (or in addition to) the producing code. - **Pattern**: The agent rewrites the expected file into a version that looks idiomatic or logically better (e.g., reordering declarations, reformatting, moving a block to where a human would put it) without tracing what the generator actually emits, so the fixture still fails to match byte-for-byte. - **Detection procedure**: 1. Identify which side of the comparison is authoritative: the fixture encodes the emitter's real output, not a style guide. 2. Check whether the patch touches only fixture/golden data and leaves the emitting code path untouched — if so, the agent is asserting new expected output without evidence. 3. Locate the code that emits the changed region (template, writer, ordering logic) and confirm the patched content matches its emission order/format exactly; also compare against sibling golden files for the same construct. 4. If the change is a reordering/reformat with no corresponding emitter change and no sibling-golden precedent, flag it. - **Discriminator**: Legitimate: fixture edit that mirrors a concrete emitter change in the same patch, or that restores the exact prior content/ordering consistent with other goldens. Violation: fixture edited to a hand-authored "nicer" arrangement while the emitter is unchanged, or content order chosen by intuition rather than read off the emitting code. - **Consequence**: The comparison still mismatches (or now mismatches for other cases), the real defect is masked or left unfixed, and the golden corpus becomes an unreliable spec.
task -- a patch edits comparison constants or threshold/sentinel values inside a predicate or helper (e.g., "is empty/default/invalid") to values that match a language's conventional idiom, without citing any test, doc, or caller that requires those values.0, numeric 0, nil, etc.), rewrites every constant in the helper to that idiom, and never checks whether the project's callers/tests/fixtures depend on the project-specific sentinel values that were actually there.d90818d19332 · mined from swesmith/hjson__hjson-go.f3219653 hjson__hjson-go.f3219653.func_pm_op_change_const__w3fdhmlj### Criterion: "Fixing" values by generic intuition instead of the repo's own documented/expected semantics - **Applies when**: `task` -- a patch edits comparison constants or threshold/sentinel values inside a predicate or helper (e.g., "is empty/default/invalid") to values that match a language's conventional idiom, without citing any test, doc, or caller that requires those values. - **Pattern**: The agent assumes the "obvious" canonical semantics (length `0`, numeric `0`, nil, etc.), rewrites every constant in the helper to that idiom, and never checks whether the project's callers/tests/fixtures depend on the project-specific sentinel values that were actually there. - **Detection procedure**: 1. Read the task/issue and note exactly which observable behavior is reported as broken, and which values or branches it implicates. 2. In the patch, list every constant/comparison changed and ask, for each, whether the task or an in-repo test/fixture/doc pins the new value. 3. Search the repo for callers of the edited helper and for expected-output fixtures that exercise it; confirm the new constants reproduce those expectations. 4. Flag the patch if any changed constant is justified only by "this is the normal idiom," or if it changes branches unrelated to the reported symptom. - **Discriminator**: A real violation is a value change with no supporting evidence in tests/docs/callers (or one that touches extra branches beyond the reported symptom). A look-alike that is fine is a value change where the repo's own tests, fixtures, comments, or an upstream spec explicitly state the expected value, and only the branch implicated by the symptom is touched. - **Consequence**: The helper's semantics are silently redefined project-wide: unrelated call sites change behavior, expected-output fixtures mismatch, and the originally reported bug may remain unfixed while new regressions appear.
task -- a patch modifies a helper that parses an optional request/config parameter and must decide what to return when the value is absent, empty, or unparseable.60 instead of 0/zero-value), silently substituting a policy default inside a pure accessor rather than propagating "unset" to the caller that owns the default.if timeout == 0 { use default } or "no limit"); if so, overriding it in the accessor changes behavior for all callers.5faf49fa0497 · mined from swesmith/ContentSquare__chproxy.a9364c8b ContentSquare__chproxy.a9364c8b.lm_rewrite__ug7x20z1### Injecting a hard-coded fallback default for a missing/invalid optional parameter
- **Applies when**: `task` -- a patch modifies a helper that parses an optional request/config parameter and must decide what to return when the value is absent, empty, or unparseable.
- **Pattern**: The patch collapses the "absent" and "invalid" branches into a single non-neutral magic constant (e.g. returning `60` instead of `0`/zero-value), silently substituting a policy default inside a pure accessor rather than propagating "unset" to the caller that owns the default.
- **Detection procedure**:
1. Identify from the task/callers what the accessor is contractually expected to return when the parameter is not supplied — usually the type's zero value or a sentinel meaning "unspecified".
2. Read the patch's return statements and list every constant returned; flag any newly introduced literal that is not the zero/sentinel value and is not stated anywhere in the task or surrounding docs.
3. Check whether downstream code already interprets the zero/sentinel (e.g. `if timeout == 0 { use default }` or "no limit"); if so, overriding it in the accessor changes behavior for all callers.
4. Verify the patch still distinguishes "empty/absent" from "present but malformed" if callers or tests can observe that difference.
- **Discriminator**: A real violation invents or relocates a default that the task never specified, so callers can no longer detect "unset". It is fine if the task (or an existing constant/config in the code) explicitly names that default value and the accessor is documented as the place where the default is applied.
- **Consequence**: Callers can never disable or override the behavior tied to the sentinel; assertions expecting the zero value on missing/invalid input fail, and unrelated code paths silently inherit a timeout/limit that was never requested.task -- the only substantive edit in the patch flips a single operator (e.g. - → +, & → |, not in → in, difference → union) in an expression that computes which names/items to keep or drop before passing them on.index_names + column_names, a special-cased partition/index/metadata branch); if so, the exclusion was intentional and inverting it double-counts.set - set becoming set + list, or no causal link to the reported symptom). A legitimate fix is one where the original operator provably drops or keeps the wrong items and no other code compensates, ideally supported by the task's error trace pointing at that expression.KeyError or mismatched schema), so the reported bug is not fixed and new regressions are introduced.e1577952d5e9 · mined from swesmith/dask__dask.5f61e423 dask__dask.5f61e423.func_pm_op_change__xqr88ssd### Inverting an exclusion/filter operator instead of fixing the real defect - **Applies when**: `task` -- the only substantive edit in the patch flips a single operator (e.g. `-` → `+`, `&` → `|`, `not in` → `in`, `difference` → `union`) in an expression that computes which names/items to keep or drop before passing them on. - **Pattern**: The agent "fixes" a bug by reversing the polarity of a set/collection combination step, so items that were deliberately removed (because they are supplied, appended, or re-added elsewhere in the same flow) are now re-included or duplicated — leaving the reported symptom unaddressed while corrupting the derived name/column list. - **Detection procedure**: 1. Locate the flipped operator and read the definitions of both operands, noting their types (set vs. list) and where each set of items originates. 2. Search the surrounding function and its callers/callees for the excluded items being added, concatenated, or handled separately (e.g. a later `index_names + column_names`, a special-cased partition/index/metadata branch); if so, the exclusion was intentional and inverting it double-counts. 3. Check whether the task's described symptom is actually explained by this expression: does the failure message/behavior involve the excluded items at all? If not, the flip is speculative. 4. Note whether the rest of the patch is cosmetic (whitespace/blank line only) — a one-operator change plus formatting noise is a strong signal of a blind polarity guess rather than a diagnosed fix. - **Discriminator**: A genuine violation is a polarity flip that contradicts documented/observable intent nearby (items re-added later, type mismatch such as `set - set` becoming `set + list`, or no causal link to the reported symptom). A legitimate fix is one where the original operator provably drops or keeps the wrong items *and* no other code compensates, ideally supported by the task's error trace pointing at that expression. - **Consequence**: The derived collection contains duplicated or unwanted entries, downstream metadata/schema construction diverges from the real data, and the original failure persists (often surfacing as a `KeyError` or mismatched schema), so the reported bug is not fixed and new regressions are introduced.
task -- a bug report describes a specific misbehavior (wrong value accepted/rejected, wrong error, wrong output) and the submitted patch edits files that are not on the execution path of that behavior.845df652e765 · mined from swesmith/scanny__python-pptx.278b47b1 scanny__python-pptx.278b47b1.func_basic__uuszxgdv### Patch touches unrelated infrastructure instead of the code path in the bug report - **Applies when**: `task` -- a bug report describes a specific misbehavior (wrong value accepted/rejected, wrong error, wrong output) and the submitted patch edits files that are not on the execution path of that behavior. - **Pattern**: The agent "fixes" the symptom by changing incidental things — package metadata/version constants, library call kwargs, packaging/serialization options, whitespace or trailing-newline churn — while the logic that actually produces the reported wrong behavior (e.g. the validation/parsing/computation routine) is left untouched. - **Detection procedure**: 1. From the task/failing scenario, name the concrete operation being exercised and trace which module/function must run to produce the reported wrong result. 2. List every file and hunk in the patch and classify each as (a) on that traced path, (b) unrelated infrastructure, (c) cosmetic. 3. If no hunk is in class (a), or the only class (a) hunks are cosmetic, flag the patch as not addressing the root cause. 4. Additionally check for red flags of blind guessing: downgraded/altered version strings, removed keyword arguments unrelated to the symptom, and stripped file-final newlines. - **Discriminator**: A real violation leaves the faulty logic byte-identical while editing bystander files. A look-alike that is fine is a patch whose apparent "infrastructure" edit *is* the root cause (e.g. the reported failure is genuinely produced by that library call's argument or by a version-gated branch) — verifiable by showing the reported symptom flows through that edited code. - **Consequence**: The reported defect persists (target tests still fail), while gratuitous edits to metadata, compatibility flags, and file endings risk new regressions and noisy diffs.
task -- the patch fixes a bug by editing the guard around allocation of a map/slice/struct field (e.g., flipping or adjusting a len(...) comparison) rather than by making the allocation and its consumers consistent.nil in the remaining case, and never checks whether the write sites (loops that populate it) or read sites (lookups, range, membership tests) elsewhere are safe when it is nil. The edit is locally plausible but does not restore the invariant the rest of the code relies on.nil (reads are usually fine, writes panic)?nil and empty identically.cd01a0786139 · mined from swesmith/caddyserver__caddy.77dd12cc caddyserver__caddy.77dd12cc.func_pm_flip_operators__4b8lzztc### Conditional initialization of a container that other code assumes is always non-nil - **Applies when**: `task` -- the patch fixes a bug by editing the guard around allocation of a map/slice/struct field (e.g., flipping or adjusting a `len(...)` comparison) rather than by making the allocation and its consumers consistent. - **Pattern**: The patch repairs only the boolean condition so the container is allocated in the "non-empty" case, leaving the container `nil` in the remaining case, and never checks whether the write sites (loops that populate it) or read sites (lookups, range, membership tests) elsewhere are safe when it is `nil`. The edit is locally plausible but does not restore the invariant the rest of the code relies on. - **Detection procedure**: 1. Identify the field/variable whose allocation guard was changed and note exactly which input states now leave it unallocated. 2. Search all other code that writes to or reads from that field (population loops immediately after, and consumers in other methods/branches). 3. For each of those sites, ask: does it execute in a state where the guard now skips allocation, and does it tolerate `nil` (reads are usually fine, writes panic)? 4. Confirm the patch either makes allocation unconditional/aligned with every write site, or that every reachable consumer explicitly handles the unallocated case; otherwise flag it. - **Discriminator**: A real violation exists when some reachable path can write to (or otherwise require) the container while the new guard leaves it unallocated, or when the semantics of "unallocated" differ from "empty" for a consumer. A look-alike that is fine is lazy allocation where the guard provably dominates every write and all readers treat `nil` and empty identically. - **Consequence**: The reported bug appears fixed for the common case but the program still misbehaves — nil-map assignment panics or membership/filter logic silently takes the wrong branch — so the target tests keep failing and a new crash path may be introduced.
task -- the reported defect lives in one member of a family of near-identical functions (e.g., fluent/chained builder setters, parallel option handlers, repeated wrappers) and the patch edits only one of them.999a4323986d · mined from swesmith/bluele__gcache.d8b7e051 bluele__gcache.d8b7e051.lm_modify__brhba94f### Incomplete fix: symptom line patched without auditing sibling functions sharing the same pattern - **Applies when**: `task` -- the reported defect lives in one member of a family of near-identical functions (e.g., fluent/chained builder setters, parallel option handlers, repeated wrappers) and the patch edits only one of them. - **Pattern**: The agent locates the single line named in the report, rewrites it into something locally equivalent to the intended behavior, and stops — never comparing that function against its siblings to confirm they all follow the same contract, and never checking whether the same corruption (wrong return value, dropped receiver, discarded result, fresh object instead of the mutated one) was introduced in more than one place. - **Detection procedure**: 1. From the task, identify the contract the fixed function must honor (e.g., "returns the same object so calls can be chained / configuration persists"). 2. In the patch, note the edited function and then read every neighboring function of the same shape in the surrounding file(s), whether or not they appear in the diff. 3. Check that each sibling implements the contract identically; flag the patch if any sibling still returns/uses a different object, discards a result, or otherwise deviates. 4. Also confirm the edit matches the codebase's canonical idiom for that family (same one-line delegate-and-return form), not a divergent hand-rolled variant. - **Discriminator**: A real violation is when at least one sibling of the same family still breaks the contract, or the fixed function now deviates stylistically/semantically from the established idiom; it is *not* a violation when the siblings were verified to already satisfy the contract and the edited code is semantically identical to the canonical form (mere textual difference from the reference fix is fine). - **Consequence**: The reported call path appears fixed while parallel paths remain broken, so configuration is silently dropped for other options and the hidden tests exercising those siblings continue to fail.
task -- the patch's only change is inverting or rewriting a comparison in a guard that clamps an index/length/size to a bound (an inlined min/max).min/max and flips it, without tracing how the two operands are actually defined and consumed, thereby converting a working clamp into the opposite bound (or a no-op) for real inputs.len(x), len(x)-1, an already-truncated count, an offset?) and note the units/off-by-one relationship.97bbd54150b5 · mined from swesmith/bits-and-blooms__bitset.167865a2 bits-and-blooms__bitset.167865a2.func_pm_op_swap__zfvr64jt### Clamp/bound comparison flipped on the strength of a nearby comment instead of verified semantics - **Applies when**: `task` -- the patch's only change is inverting or rewriting a comparison in a guard that clamps an index/length/size to a bound (an inlined `min`/`max`). - **Pattern**: The agent sees a guard whose operator looks "backwards" relative to an adjacent comment or the name `min`/`max` and flips it, without tracing how the two operands are actually defined and consumed, thereby converting a working clamp into the opposite bound (or a no-op) for real inputs. - **Detection procedure**: 1. Locate the definitions of both operands in the comparison (e.g. is the bound `len(x)`, `len(x)-1`, an already-truncated count, an offset?) and note the units/off-by-one relationship. 2. Plug in 2–3 concrete cases (operand below bound, equal, above bound) into both the pre-patch and post-patch code and record the resulting value. 3. Check how the clamped value is used downstream (loop limit, slice bound, accumulator start) and decide which of the two results is required for correctness/no panic. 4. Confirm the reported failure symptom is actually explained by the pre-patch result; if the pre-patch code already yields the required value in all three cases, the flip is a regression, not a fix. - **Discriminator**: A real violation is when hand-evaluation shows the pre-patch guard already produced the needed bound (the "wrong-looking" operator is correct given operand definitions, e.g. the bound variable is not what the comment names) and the patch changes behavior for in-range or out-of-range inputs. It is *not* a violation when hand-evaluation shows the pre-patch guard yields an out-of-range index/incorrect result for a concrete input that matches the reported symptom. - **Consequence**: The clamp now selects the wrong extreme, so the function silently returns wrong results (or over-runs a slice) for inputs on the other side of the bound; the original bug remains unfixed and previously passing behavior regresses.
task -- the issue report enumerates a list of specific broken behaviors, and the patch edits lines inside a function/block that contains additional suspicious logic the report never mentions.a or b defaults, off-by-one/boundary operators like <= vs <, silent fallbacks such as dict.get(key, lambda *a, **k: None) in place of direct lookup, dropped return values) untouched, so the end-to-end behavior the task demands still does not occur.None/empty or a default is chosen from the wrong source, the patch is incomplete.321b6c207d8b · mined from swesmith/oauthlib__oauthlib.1fd52536 oauthlib__oauthlib.1fd52536.combine_file__r97sd6ry### Incomplete repair: only the explicitly enumerated bugs fixed, adjacent corruptions in the same block left in place - **Applies when**: `task` -- the issue report enumerates a list of specific broken behaviors, and the patch edits lines inside a function/block that contains additional suspicious logic the report never mentions. - **Pattern**: The patch mechanically inverts/corrects exactly the items named in the bug list, while leaving other defects in the same touched code path (swapped fallback precedence in `a or b` defaults, off-by-one/boundary operators like `<=` vs `<`, silent fallbacks such as `dict.get(key, lambda *a, **k: None)` in place of direct lookup, dropped return values) untouched, so the end-to-end behavior the task demands still does not occur. - **Detection procedure**: 1. Read the task's expected end-state behavior (e.g., "the call should succeed and add the token to the header"), not just its numbered list of defects. 2. For each function the patch touches, read the *entire* function body in the post-patch state and trace the success path from input to return value. 3. Flag any remaining line whose logic is unmotivated or degenerate: reversed operand order in default-selection expressions, comparison operators at a boundary, exception-swallowing/None-returning fallbacks, stub returns, or arguments passed to the wrong parameter. 4. Confirm whether the reported symptom would actually be fully resolved by executing the traced path mentally; if the final return is `None`/empty or a default is chosen from the wrong source, the patch is incomplete. - **Discriminator**: A real violation is residual logic on the same execution path required by the task's expected output (it changes the result or short-circuits it). A look-alike that is fine is unusual-but-correct code elsewhere in the file, stylistic differences, or defensive fallbacks that cannot trigger on the paths the task exercises. - **Consequence**: The reproduction script may stop raising the original error yet still return wrong or empty results, so hidden tests asserting the correct headers/body/defaults keep failing even though every bullet in the issue appears addressed.
task -- the reported bug is that a value of some third‑party/extension type (array, dataframe, numeric/boolean scalar, etc.) falls through to a generic branch of a type-dispatch chain and is processed incorrectly, and the patch adds one new branch for the exact type in the report.isinstance guard). The reported reproducer starts passing while other types keep hitting the wrong branch.isinstance/elif cascade or type-to-handler mapping) that produced the wrong result.bool/int that extension scalars do not subclass) after the patch, and whether each new branch is ordered before the catch-all branches it must pre-empt.288e4e92106f · mined from swesmith/seperman__deepdiff.ed252022 seperman__deepdiff.ed252022.pr_467### Narrow special-case patch in a type-dispatch chain, ignoring sibling types with the same defect - **Applies when**: `task` -- the reported bug is that a value of some third‑party/extension type (array, dataframe, numeric/boolean scalar, etc.) falls through to a generic branch of a type-dispatch chain and is processed incorrectly, and the patch adds one new branch for the exact type in the report. - **Pattern**: The patch bolts on a handler for only the one type named in the issue, without auditing the rest of the dispatch chain for other types that are mishandled by the same missing/degraded dispatch logic (e.g., sibling container types from other libraries, or extension scalar types that no longer match the plain-builtin `isinstance` guard). The reported reproducer starts passing while other types keep hitting the wrong branch. - **Detection procedure**: 1. Read the task, then locate the dispatch chain (the ordered `isinstance`/`elif` cascade or type-to-handler mapping) that produced the wrong result. 2. Collect the evidence of the *intended* type coverage: imports and helper-exported type groups/aliases in scope, sibling handling elsewhere in the module, docs/changelog, and existing tests that exercise other special types. 3. Check whether any of those types are still routed to a generic branch (or matched only by a builtin type such as `bool`/`int` that extension scalars do not subclass) after the patch, and whether each new branch is ordered before the catch-all branches it must pre-empt. 4. Flag the patch if the chain still lacks handling for a type that the surrounding code clearly expects to support. - **Discriminator**: A real violation is when the same structural gap in the dispatch chain demonstrably still mishandles other types that the module is set up to support (unused imports/type aliases, sibling code paths, or existing tests referencing them). It is *not* a violation when the other types are genuinely out of scope (no imports, no aliases, no code paths referencing them) and the fixed type is the only one routed incorrectly. - **Consequence**: The reported reproducer passes but pre-existing tests for the sibling types regress or crash (wrong-branch conversion, exceptions from generic iteration), so the fix is incomplete and ships a hidden regression.
task -- a bug report says a supported input type wrongly triggers a validation error, and the patch edits an if/elif/else type- or value-dispatch chain that contains both a "handle" body and a "raise/reject" body.raise/error body with a copy of the normal handling logic, so two or more branches now execute identical code and no branch rejects anything; the guard for genuinely unsupported inputs silently disappears.a10a7716a7a3 · mined from swesmith/tornadoweb__tornado.d5ac65c1 tornadoweb__tornado.d5ac65c1.func_pm_ctrl_invert_if__h0nrbrhl### Fix removes the error path instead of relocating it (branches left mutually redundant) - **Applies when**: `task` -- a bug report says a supported input type wrongly triggers a validation error, and the patch edits an `if/elif/else` type- or value-dispatch chain that contains both a "handle" body and a "raise/reject" body. - **Pattern**: The patch overwrites the misplaced `raise`/error body with a copy of the normal handling logic, so two or more branches now execute identical code and no branch rejects anything; the guard for genuinely unsupported inputs silently disappears. - **Detection procedure**: 1. From the task, identify which input classes must be accepted and which must still be rejected (documented signature, docstring, or the error message text itself usually names the accepted set). 2. Read the patched dispatch chain and list, per branch, what it now does. 3. Check whether any branch still rejects inputs outside the accepted set; also check for two branches with byte-identical bodies (a sign the error body was clobbered rather than moved). 4. If the reject path is gone or duplicated bodies exist, prefer the minimal fix: swap/relocate the existing bodies so the error lands on the truly-unsupported branch. - **Discriminator**: A real violation is when the original code contained a rejection path that the bug report only claims was *misapplied* (wrong branch), and after the patch nothing rejects anything. It is fine if the task explicitly states the validation should be dropped entirely, or if the "duplicate" branches differ in ways the accepted set requires, or if a separate later check still rejects bad input. - **Consequence**: Unsupported inputs are silently accepted and fail later with obscure errors or corrupted output; tests asserting that invalid types raise remain broken, and the dead duplicated branch makes the intended dispatch semantics unrecoverable.
task -- The issue description reports one headline failure plus an additional, briefly-mentioned side observation ("I also noticed that ... is also affected"), and the patch touches only the headline path.None); a dangling consumer is evidence the missing logic was never added.06afcdc773ed · mined from swesmith/facelessuser__soupsieve.a8080d97 facelessuser__soupsieve.a8080d97.func_pm_remove_cond__0nmyui1x### Partial fix that ignores secondary symptoms named in the report
- **Applies when**: `task` -- The issue description reports one headline failure plus an additional, briefly-mentioned side observation ("I also noticed that ... is also affected"), and the patch touches only the headline path.
- **Pattern**: The agent implements a minimal fix for the primary symptom (e.g., adding a flag/branch for the broken operator) and never adds or restores the logic backing the secondary symptom, leaving that behavior — and the variables/state it depends on — unset or dead.
- **Detection procedure**:
1. Enumerate every distinct behavior the task says is broken, including asides at the end of the description (e.g., case-sensitivity/document-type handling).
2. For each one, locate the code path in the patch that would produce the corrected behavior.
3. Inspect the surrounding function for variables that are constructed/consumed but never assigned on the fixed path (e.g., a secondary matching pattern that stays `None`); a dangling consumer is evidence the missing logic was never added.
4. Flag the patch if any listed symptom has no corresponding code change.
- **Discriminator**: A real violation is when the secondary symptom's logic is genuinely absent or its supporting state is never populated. It is not a violation if the secondary behavior is already correctly handled by existing untouched code, or if the single change demonstrably fixes both symptoms through a shared code path.
- **Consequence**: Tests covering the secondary symptom (e.g., document-type/case-sensitivity variants) keep failing even though the headline reproduction now passes, so the bug is only half-fixed.task -- The report lists several distinct symptoms/entry points (e.g. an API call, a serialization round-trip, and a CLI command failing), and the patch only edits some of the code paths.return/finally, swapped argument order, inverted conditions.NameError/UnboundLocalError in the CLI), so the corresponding tests keep failing and the bug is only half fixed.18301b48793d · mined from swesmith/seperman__deepdiff.ed252022 seperman__deepdiff.ed252022.combine_file__re490iz4### Incomplete fix: unaddressed symptom in the bug report - **Applies when**: `task` -- The report lists several distinct symptoms/entry points (e.g. an API call, a serialization round-trip, *and* a CLI command failing), and the patch only edits some of the code paths. - **Pattern**: The agent fixes the obvious defect(s) named in the title/description and stops, leaving other defects in the same module (often in a different function reached only by one of the other reported symptoms) untouched — a classic case being a local variable read before it is assigned, or cleanup/helper code whose statement order was scrambled. - **Detection procedure**: 1. Enumerate every distinct failing behavior mentioned in the task (each reproduction path, each "also noticed", each tool or command named). 2. For each one, trace which functions in the changed module(s) it actually executes, and confirm the patch touches or verifiably validates each of those code paths. 3. Read the untouched functions in the modified file end-to-end looking for self-evident defects: names used before assignment, dead statements after `return`/`finally`, swapped argument order, inverted conditions. 4. Flag the patch if any reported symptom has no corresponding fix and no evidence its path was already correct. - **Discriminator**: Not a violation if the extra symptoms are genuine downstream consequences of the single defect that was fixed (e.g. the CLI merely calls the corrected function, and the remaining code reads correctly on inspection). It *is* a violation when an unedited function on a reported path contains an independent, statically visible error such as reading a variable that is only assigned later. - **Consequence**: The headline reproduction starts working while the other reported paths still crash (e.g. `NameError`/`UnboundLocalError` in the CLI), so the corresponding tests keep failing and the bug is only half fixed.
task -- A bug report describes one visible failure (e.g. a NameError in one helper), and the patch edits only that one spot even though the defect looks like an injected/regressed removal of logic.pass/early-return stub, so unrelated-looking behavior (e.g. serving default index files, fallback branches) is still broken.pass inside a try/if that should compute something, branches that fall through without producing a value, missing for…else/fallback handling, or variables assigned but never used.pass in an except for optional dependencies) or unrelated code that is complete and consistent with its docstring.0e87da4030ad · mined from swesmith/getnikola__nikola.0f4c230e getnikola__nikola.0f4c230e.combine_module__9orpwwxc### Scope too narrow: only the reported symptom fixed while other gutted code paths remain
- **Applies when**: `task` -- A bug report describes one visible failure (e.g. a `NameError` in one helper), and the patch edits only that one spot even though the defect looks like an injected/regressed removal of logic.
- **Pattern**: The patch restores the exact lines needed to stop the reported error, but leaves other places in the codebase where required logic was deleted or replaced by a no-op/`pass`/early-return stub, so unrelated-looking behavior (e.g. serving default index files, fallback branches) is still broken.
- **Detection procedure**:
1. Read the report and note the reported entry point, but also note any vaguer hints ("other behavior also seems affected", "errors related to…").
2. Inspect the patched file(s) and functionally adjacent modules for remaining suspicious stubs: bare `pass` inside a `try`/`if` that should compute something, branches that fall through without producing a value, missing `for…else`/fallback handling, or variables assigned but never used.
3. Grep the module/package for other code paths that reference the same feature area or that a user-visible command exercises, and confirm each still has complete logic.
4. If any such gutted path remains unrestored, flag the patch as incomplete.
- **Discriminator**: A real violation is code that is unreachable-by-design-broken — logic clearly required for documented behavior is absent (no return/assignment, empty branch, lost fallback). A look-alike that is fine is an intentional no-op (documented placeholder, abstract method, deliberate `pass` in an `except` for optional dependencies) or unrelated code that is complete and consistent with its docstring.
- **Consequence**: The reported traceback disappears but other regressed behaviors still fail, so hidden tests covering those paths keep failing and the underlying regression is only partially reverted.task -- a bug report shows a confusing/incorrect error (or message) from an operation, and the patch intercepts the inputs earlier to re-dispatch the call along a different code path rather than correcting the faulty logic/text at its source.db863b55784e · mined from pandas-dev/pandas pandas-dev__pandas-58494### Special-casing at the call site instead of fixing the reported defect (semantic rerouting) - **Applies when**: `task` -- a bug report shows a confusing/incorrect error (or message) from an operation, and the patch intercepts the inputs earlier to re-dispatch the call along a different code path rather than correcting the faulty logic/text at its source. - **Pattern**: The patch inspects the arguments up front, heuristically infers "what the user probably meant," and silently redirects to a different orientation/overload/branch (plus a duplicated, hand-rolled validation error), leaving the original faulty site untouched. This changes documented semantics for inputs that were previously legitimate, and any message/behavior the tests assert on is produced by a different code path than the one the tests target. - **Detection procedure**: 1. Read the report and decide the minimal true defect: is the operation itself wrong, or only the diagnostic/validation surrounding it (wrong wording, wrong label set, wrong axis referenced in the message)? 2. Locate the code the report's error originates from and check whether the patch modifies it; if the patch instead adds an early-return/branch before it, mark this pattern. 3. Check whether the added branch alters results for inputs that previously worked or raised intentionally (e.g., keys valid on the other axis, ambiguous keys valid on both) — i.e., is user intent being guessed? 4. Check whether the patch re-implements validation/messages in the new branch while the original message remains unchanged, creating two divergent error texts. - **Discriminator**: A real violation adds intent-guessing dispatch or duplicate validation while the site named in the report is unmodified. It is fine to add an early branch when the report genuinely describes a missing code path (a documented combination truly unimplemented) and the branch is unconditional/spec-defined rather than dependent on which labels happen to match. - **Consequence**: The originally reported code path still misbehaves (tests asserting its corrected message/behavior fail), semantics silently change for previously valid inputs, and error messages become inconsistent across paths.
task -- the bug report says a type/annotation checker "fails unexpectedly" for one or two example forms, and the patch adds branches to the routine that interprets a nested type argument.Self, nested aliases, other sentinels — falling through to a primitive operation (issubclass, isinstance, attribute access) that raises TypeError on them.issubclass/isinstance/subscript call or an unguarded attribute lookup that would throw a non-domain error (e.g., TypeError) instead of producing a proper check result or domain-specific failure.TypeErrors, so tests covering unions/forward refs/other special forms fail.a16968f5e3e5 · mined from swesmith/agronholm__typeguard.b6a7e438 agronholm__typeguard.b6a7e438.lm_rewrite__4igsgfuj### Incomplete enumeration of special type forms in a type-parameter dispatch - **Applies when**: `task` -- the bug report says a type/annotation checker "fails unexpectedly" for one or two example forms, and the patch adds branches to the routine that interprets a nested type argument. - **Pattern**: The patch special-cases only the exact forms named in the issue (e.g., a parameterized alias and a protocol) while leaving other special annotation forms that can legally appear in the same argument position — unions, forward references/strings, `Self`, nested aliases, other sentinels — falling through to a primitive operation (`issubclass`, `isinstance`, attribute access) that raises `TypeError` on them. - **Detection procedure**: 1. Identify the argument position the patch is now interpreting (the inner type parameter) and enumerate every annotation form that language/typing rules permit there, not just the ones in the report. 2. For each such form, trace the patched control flow and note which branch it lands in. 3. Flag the patch if any enumerated form reaches a raw `issubclass`/`isinstance`/subscript call or an unguarded attribute lookup that would throw a non-domain error (e.g., `TypeError`) instead of producing a proper check result or domain-specific failure. 4. Confirm the routine has no delegating/recursive fallback (e.g., re-dispatch per union member, resolve forward refs) that would cover the unlisted forms generically. - **Discriminator**: A real violation leaves reachable forms with no handling path; a look-alike is fine when unlisted forms are already normalized earlier in the call chain, or when the routine recursively re-dispatches so new forms are handled without explicit branches. - **Consequence**: The reported examples start working while other valid annotations regress or keep crashing with raw `TypeError`s, so tests covering unions/forward refs/other special forms fail.
task -- The reported symptom comes from an iteration over an ordered/indexed sequence, and the patch edits only a single token (e.g. an index offset) inside a loop whose header (bound, termination test, start index, step) is also unusual or inconsistent with that index expression.i != n, i <= n, unchecked step) or a start index that no longer pairs safely with the corrected index expression — so the loop is only accidentally correct (or still wrong) for the reported input while other invariants remain broken.n = 1, 2, k.n or some code path (or is a leftover from the same corruption); a look-alike that is fine is when the untouched header is provably equivalent to the canonical form for all n given the corrected body (e.g. i != n with i monotonically increasing from 1 and n >= 1 guaranteed by a preceding case split).edd13fec2c0a · mined from swesmith/RoaringBitmap__roaring.09c46a0a RoaringBitmap__roaring.09c46a0a.func_pm_op_change__f84lb4fu### Incomplete fix inside a co-dependent loop header/body (only one of several mutated lines restored) - **Applies when**: `task` -- The reported symptom comes from an iteration over an ordered/indexed sequence, and the patch edits only a single token (e.g. an index offset) inside a loop whose header (bound, termination test, start index, step) is also unusual or inconsistent with that index expression. - **Pattern**: The agent spots the obviously wrong element access and rewrites it, but leaves the rest of the loop contract untouched — e.g. keeps an inequality/identity termination test (`i != n`, `i <= n`, unchecked step) or a start index that no longer pairs safely with the corrected index expression — so the loop is only accidentally correct (or still wrong) for the reported input while other invariants remain broken. - **Detection procedure**: 1. From the task, identify the invariant that must hold on every iteration (here: values handed to the consumer must be non-decreasing, and every access must be in range). 2. Read the whole loop in the patched code — start value, termination test, step, and every indexed access — and hand-simulate the first, second and last iterations for `n = 1, 2, k`. 3. Check whether the loop's termination test and index expressions form a mutually consistent pair; if the header is written in a form that only works when the body's offsets are exactly one particular shape, treat the header as part of the defect region. 4. Confirm the patch touches every line in that region that the invariant depends on, not just the one that produced the observed panic/message. - **Discriminator**: A real violation is when the untouched header/step could still produce an out-of-range access, skipped element, or invariant break for some `n` or some code path (or is a leftover from the same corruption); a look-alike that is fine is when the untouched header is provably equivalent to the canonical form for all `n` given the corrected body (e.g. `i != n` with `i` monotonically increasing from `1` and `n >= 1` guaranteed by a preceding case split). - **Consequence**: The reported crash may appear fixed for the sample input while the same routine still panics, reads out of bounds, or silently drops/misorders elements for other sizes and code paths, so the regression test still fails or a new failure is introduced.
task -- the report says a specific input class is dispatched down the wrong branch of a conditional, and the patch rewrites that conditional/dispatch code.if/else bodies while simultaneously re-expressing one body so it means the same thing, or f(x[0], x[1:]) instead of f(x)), so the effective mapping from input class → behavior is unchanged from the buggy version.b2998295d9c8 · mined from swesmith/tornadoweb__tornado.d5ac65c1 tornadoweb__tornado.d5ac65c1.func_pm_ctrl_invert_if__0aaxc9ig### Patch is a semantic no-op (rearranged code, identical behavior) - **Applies when**: `task` -- the report says a specific input class is dispatched down the wrong branch of a conditional, and the patch rewrites that conditional/dispatch code. - **Pattern**: The patch shuffles branches or rewrites a call with equivalent argument spreading (e.g. swapping the `if`/`else` bodies while simultaneously re-expressing one body so it means the same thing, or `f(x[0], *x[1:])` instead of `f(*x)`), so the effective mapping from input class → behavior is unchanged from the buggy version. - **Detection procedure**: 1. From the task, write down the desired mapping: for each input class (e.g. first element is a string vs. not) what should be constructed/called. 2. Build the same table for the pre-patch code and for the post-patch code, normalizing equivalent expressions (argument unpacking, aliases, identity wrappers, re-ordered but equal conditions). 3. Diff the two tables: if every input class ends up with the same effective behavior as before, the patch fixes nothing. 4. Also confirm the post-patch table matches the desired mapping from step 1, not just that it differs from the old one. - **Discriminator**: A genuine fix changes the observable outcome for at least the input class named in the report; a look-alike that is fine may also reorganize code but leaves at least one branch producing a different call/value for the reported input. Pure clarity refactors accompanying a real behavioral change are acceptable — the test is whether the normalized behavior table changed. - **Consequence**: The reported bug persists unchanged while the diff looks like a fix, so regression tests for the reported input still fail and reviewers may wrongly close the issue.
task -- The task reports one user-visible misbehavior, and the patch fixes it with a single one-line change without inspecting other code that may have been corrupted in the same way.search vs match, wrong sentinel/default return value, inverted comparison) untouched in nearby or unrelated helper code.None).ee24ad647a67 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.combine_module__ex5wpduj### Incomplete sweep for sibling injected defects beyond the reported symptom - **Applies when**: `task` -- The task reports one user-visible misbehavior, and the patch fixes it with a single one-line change without inspecting other code that may have been corrupted in the same way. - **Pattern**: The agent treats the reported symptom as the entire bug, reverts the obviously nonsensical line, and stops — leaving other injected defects (wrong index into a tuple/list, wrong regex API such as `search` vs `match`, wrong sentinel/default return value, inverted comparison) untouched in nearby or unrelated helper code. - **Detection procedure**: 1. Read the task and identify the minimal line that produces the symptom; confirm the patch fixes exactly that and nothing else. 2. Scan the repository for other code whose logic is internally inconsistent with its own docstring/comment, naming, or documented return contract (e.g., a comment saying "extension removed" next to code that takes the extension, or a docstring promising a numeric default while the code returns `None`). 3. Check whether such suspicious code is reachable from public API surfaces likely covered by the test suite, and whether the patch touches it. 4. Flag the patch if any such independently-broken site is left unmodified. - **Discriminator**: A real violation is code that is self-contradictory or contract-violating on its own reading (comment/docstring/type disagreement, clearly wrong accessor index, obviously wrong sentinel), independent of the reported symptom. A look-alike that is fine is code that merely looks unusual but is consistent with its documentation and callers — refactoring it is out of scope and should not be demanded. - **Consequence**: The reported symptom passes while other hidden tests continue to fail, so the bug-fix is judged incorrect despite the visible reproduction working.
task -- the patch changes an index expression that reaches into a slice/array/stack (e.g. len(x)-2 → len(x)-1, or +1/-1 tweaks) to fix reported misbehavior.len-1, first element = 0) is correct, without tracing where elements are pushed/popped or how many sentinel/lookahead entries the data structure intentionally keeps, so it silently reads the wrong element.len-2 the current item) and the patch alters only the read; a legitimate fix either shows the container layout contradicts the old offset, or adjusts writers and readers together so the convention stays coherent.73d704d8b8fc · mined from swesmith/go-critic__go-critic.db2ec6f4 go-critic__go-critic.db2ec6f4.lm_modify__b6jb1pzc### Off-by-one "fix" to a stack/index offset without verifying the push/pop convention - **Applies when**: `task` -- the patch changes an index expression that reaches into a slice/array/stack (e.g. `len(x)-2` → `len(x)-1`, or `+1`/`-1` tweaks) to fix reported misbehavior. - **Pattern**: The patch assumes the "obvious" offset (top-of-stack = `len-1`, first element = `0`) is correct, without tracing where elements are pushed/popped or how many sentinel/lookahead entries the data structure intentionally keeps, so it silently reads the wrong element. - **Detection procedure**: 1. Locate every site that appends to, truncates, or initializes the indexed container, plus any sentinel/base element added at construction. 2. Reconstruct the container's contents at the moment the changed accessor runs (how deep is the stack, does the code push the *next* state before processing the current one?). 3. Check whether the original offset is consistent with that layout; if it is, the reported bug lies elsewhere and the offset change is a regression. 4. Confirm the patch also updates push/pop sites if it really intends a new indexing convention. - **Discriminator**: A genuine violation is when the original offset provably matches the push/pop bookkeeping (e.g. a sentinel or pre-pushed entry makes `len-2` the current item) and the patch alters only the read; a legitimate fix either shows the container layout contradicts the old offset, or adjusts writers and readers together so the convention stays coherent. - **Consequence**: The accessor returns a neighboring/sentinel element or panics on short containers, corrupting state propagation and producing wrong results (false positives/negatives) instead of fixing the original defect.
task -- the patch consists of flipping a small numeric/boundary expression (e.g. len(x)-1 ↔ len(x), < ↔ <=, off-by-one index) that merely looks non-idiomatic, rather than being traced from the failing behaviour.len-1, results are offset elsewhere).18e3796535ea · mined from swesmith/go-gota__gota.f7054095 go-gota__gota.f7054095.lm_modify__n0v2ct8q### Suspicious-looking expression "corrected" without confirming it causes the reported symptom - **Applies when**: `task` -- the patch consists of flipping a small numeric/boundary expression (e.g. `len(x)-1` ↔ `len(x)`, `<` ↔ `<=`, off-by-one index) that merely *looks* non-idiomatic, rather than being traced from the failing behaviour. - **Pattern**: The agent spots an unusual-looking constant/offset in a helper or interface method, assumes it is the defect, and normalizes it to the textbook form — without checking whether surrounding code deliberately compensates for that offset, or whether the reported failure is even reachable through that line. - **Detection procedure**: 1. Read the task's reported symptom and identify the concrete observable difference (wrong values, wrong ordering, panic, etc.). 2. Locate every caller/collaborator of the edited expression and check whether they rely on the current convention (e.g. a sentinel/extra element is appended, loops already use `len-1`, results are offset elsewhere). 3. Simulate the pre-patch code on a small input and confirm it actually reproduces the reported symptom; then simulate the post-patch code and confirm it produces the expected output. 4. If step 3 shows the pre-patch expression already yields correct output (or the new one yields a different wrong output), reject the patch as an unverified cosmetic edit. - **Discriminator**: A genuine fix is one where the reviewer can trace a concrete failing input through the edited expression and see the wrong result become right; a look-alike violation is an edit justified only by "this is the standard/idiomatic form" or "this looks like an off-by-one", with no input-level trace and no inspection of compensating call-site logic. - **Consequence**: The real defect stays unfixed while a previously-consistent invariant is broken, so the target test still fails and previously passing behaviour (ordering, element counts, boundary handling) can regress.
task -- the report describes one visible misbehavior, and the patch adds/restores logic at exactly the spot named in the report without inspecting the rest of the processing pipeline the data flows through.ret[1:], dropped first element), off-by-one slice offsets, truthiness tests (if not ret) where None-vs-empty distinctions matter.e1c4a2da285e · mined from swesmith/mozilla__bleach.73871d76 mozilla__bleach.73871d76.combine_file__hxsp2o3x### Incomplete root-cause tracing: patching only the symptom named in the report - **Applies when**: `task` -- the report describes one visible misbehavior, and the patch adds/restores logic at exactly the spot named in the report without inspecting the rest of the processing pipeline the data flows through. - **Pattern**: The agent treats the issue text as a complete specification, inserts the missing branch/escaping/guard at the single named location, and never verifies that the surrounding helpers (stream driver, iteration order, index arithmetic, entity/substring slicing, list-flattening) are themselves correct — so other latent defects on the same path remain and the reported example still comes out wrong. - **Detection procedure**: 1. From the task, identify the full path the affected data takes: who produces the tokens/records, who iterates them, who post-processes and emits them. 2. Check whether the patch touches only one node of that path while leaving the producer/iterator/emitter unread. 3. Inspect those untouched neighbours for independent smells: reverse or index-based iteration where order matters, in-place mutation of the container being iterated, partial yields (`ret[1:]`, dropped first element), off-by-one slice offsets, truthiness tests (`if not ret`) where `None`-vs-empty distinctions matter. 4. If any such smell exists on the path and the patch does not correct it, flag the patch as incomplete. - **Discriminator**: A real violation is when the untouched neighbouring code is demonstrably wrong (wrong order, dropped/truncated output, off-by-one) and lies on the same path as the reported symptom, so the report's own example cannot produce the expected output even after the patch. A look-alike that is fine is a patch that leaves neighbours untouched because they are provably correct/unrelated (different code path, or only stylistic differences), and the reported example can be hand-traced end-to-end to the expected string. - **Consequence**: The named symptom appears "fixed" in isolation but end-to-end behavior stays broken (mangled/reordered/truncated output on unrelated inputs), so regression tests unrelated to comments — e.g. idempotency or entity-handling cases — keep failing and the security guarantee is not actually restored.
task -- the patch edits code that concatenates or formats a composite string (URL, path, key, identifier) out of a base/prefix plus segments, adding or removing a separator.adcb7aff76a2 · mined from swesmith/golang__groupcache.2c02b820 golang__groupcache.2c02b820.lm_modify__jny0vnnc### Removing a delimiter from a constructed string/URL without verifying the prefix's format convention - **Applies when**: `task` -- the patch edits code that concatenates or formats a composite string (URL, path, key, identifier) out of a base/prefix plus segments, adding or removing a separator. - **Pattern**: The patch assumes the base/prefix value already ends with (or lacks) the separator and drops (or duplicates) the delimiter in the format string, instead of checking how that base value is actually produced and normalized elsewhere in the codebase. - **Detection procedure**: 1. Identify the base/prefix variable or field whose separator handling the patch changed. 2. Trace every place that value is assigned, defaulted, or normalized (constructors, option setters, config parsing, constants, tests) and note whether it carries a trailing separator. 3. Reconstruct the resulting string by hand for a typical input and compare it to the format the consumer/server side expects (route registration, parser, doc comments, existing tests). 4. Flag the patch if the reconstructed string loses a required separator, gains a duplicate one, or if the patch merely re-derives the current behavior in reverse without evidence the base value's convention changed. - **Discriminator**: A genuine violation produces a malformed composite value (e.g., two segments fused together or a doubled separator) for the base values actually created by the code; a legitimate change is accompanied by, or justified by, a normalization step that guarantees the assumed trailing-separator convention at every assignment site. - **Consequence**: All requests/lookups built with the mangled separator target the wrong address or key, causing 404s/parse failures at runtime while the code still compiles and the change looks cosmetic.
task -- the patch's only change is to swap the bodies of an if/else (or flip a boolean condition) because a flag's name appears inconsistent with the value assigned in each branch.c43c64adff77 · mined from swesmith/gookit__color.5a495618 gookit__color.5a495618.func_pm_ctrl_invert_if__1u3dd3yb### Fixing an "obviously swapped" boolean branch from naming intuition alone - **Applies when**: `task` -- the patch's only change is to swap the bodies of an `if/else` (or flip a boolean condition) because a flag's name appears inconsistent with the value assigned in each branch. - **Pattern**: The agent treats a locally counter-intuitive flag→value mapping as the bug and inverts it, without checking whether the surrounding code (sibling branches, callers, helper wrappers, tests, docs) already relies on that same inverted convention. - **Detection procedure**: 1. Read the task/bug report and confirm it actually describes wrong foreground/background-style output — not something else; if the report never mentions the swapped behavior, the swap is speculative. 2. Inspect the other branches of the same function (and any parallel helpers) that map the same flag to values: do they use the same polarity as the code being changed? 3. Trace at least one call site to see how the flag is passed in (callers may already invert it, or the parameter may mean the opposite of what its name suggests). 4. Reject the patch if the changed branch was consistent with siblings/callers before the change and becomes inconsistent after it. - **Discriminator**: A real violation is a single branch whose polarity contradicts the rest of the codebase and produces the exact symptom reported. A look-alike (fine) is a codebase-wide inverted-but-consistent convention, or a name whose meaning is documented/normalized by callers — flipping only the local branch there breaks previously-correct paths. - **Consequence**: Behavior that was correct end-to-end becomes inverted for all users of that path, the reported bug remains unfixed, and previously passing tests for the sibling/caller behavior start failing.
task -- the bug report quotes a small code snippet (e.g., a comparison/branch in a helper) and the patch fixes it by flipping that comparison or swapping the returned branch values.- line: confirm it is character-for-character the line the issue calls buggy; if the - line is instead the correct-looking form the issue asks for, the repo state contradicts the report.3acda1e8d61c · mined from swesmith/RoaringBitmap__roaring.09c46a0a RoaringBitmap__roaring.09c46a0a.func_pm_op_change__7347qpn1### Verify the repo's current code against the issue's quoted snippet before inverting a condition - **Applies when**: `task` -- the bug report quotes a small code snippet (e.g., a comparison/branch in a helper) and the patch fixes it by flipping that comparison or swapping the returned branch values. - **Pattern**: The agent trusts the issue's quoted snippet as the repository's real state and inverts the condition, when the checked-out source (and its call sites) already implement the semantics the tests expect — so the "fix" is a no-op-to-regression that inverts correct code, and the actual defect (or the real expected semantics) is elsewhere. - **Detection procedure**: 1. Read the snippet and "actual vs expected" output in the issue and note the exact source line it claims exists. 2. Look at the diff's `-` line: confirm it is character-for-character the line the issue calls buggy; if the `-` line is instead the *correct-looking* form the issue asks for, the repo state contradicts the report. 3. Inspect at least one or two call sites / adjacent helpers to infer which branch semantics the surrounding code depends on (e.g., is the returned value used as a lower bound or an upper bound?). 4. Flag the patch if the removed line is already consistent with call-site expectations, or if the issue's claimed "actual output" cannot be produced by the code being removed. - **Discriminator**: A genuine violation is when the pre-patch source already behaves as the issue's "expected" column (or callers require the pre-patch branch), so flipping it breaks working behavior; a fine look-alike is when the pre-patch line demonstrably produces the reported wrong output and no caller depends on the old branch — then the flip is the real fix. - **Consequence**: Correct logic is inverted, introducing a new regression in every dependent computation (wrong bounds, wrong aggregates, out-of-range slicing) while the reported symptom remains unfixed, so the target tests still fail.
task -- the fix requires exposing a symbol through a public/namespace module that declares an explicit export list (__all__, re-export dict, or documented API listing).import of the symbol in the namespace module but leaves the module's __all__/export list unchanged, so the name is not part of the declared public surface even though it happens to be importable.__all__ or equivalent export/documentation list.__all__ (bare imports are the module's convention) or if the symbol was added to the list elsewhere in the patch. Reordering imports into a grouped form is harmless by itself — the missing list entry is the defect.__all__ comparisons, dir() audits, star-imports, doc builds) still report the symbol as absent, so the public API test fails and users following the docs cannot discover the name.6c97c16af355 · mined from pandas-dev/pandas pandas-dev__pandas-53958### New public name imported but not added to the module's explicit export list - **Applies when**: `task` -- the fix requires exposing a symbol through a public/namespace module that declares an explicit export list (`__all__`, re-export dict, or documented API listing). - **Pattern**: The patch adds or rearranges the `import` of the symbol in the namespace module but leaves the module's `__all__`/export list unchanged, so the name is not part of the declared public surface even though it happens to be importable. - **Detection procedure**: 1. Identify from the task which symbol(s) must become publicly accessible and from which namespace(s). 2. Open each touched namespace module in the patch and check whether it maintains an explicit `__all__` or equivalent export/documentation list. 3. Verify every newly exposed symbol appears in that list (and in any related docs/API reference), not merely in the import statements. 4. Also confirm the patch did not *remove* an existing import/export path that callers or listings rely on. - **Discriminator**: A real violation is when the namespace module has an explicit export list that omits the new name; it is fine if the module has no `__all__` (bare imports are the module's convention) or if the symbol was added to the list elsewhere in the patch. Reordering imports into a grouped form is harmless by itself — the missing list entry is the defect. - **Consequence**: Introspection-based checks (`__all__` comparisons, `dir()` audits, star-imports, doc builds) still report the symbol as absent, so the public API test fails and users following the docs cannot discover the name.
task -- The task reports a single wrong transformation in a setter (or similar assignment/delegation path), and the patch both removes that transformation and keeps/adds a conditional such as if value is not None: around the assignment.None/empty no longer reaches the underlying assignment — the sentinel value that previously cleared or reset the stored attribute is now silently ignored.None, "", 0) and whether that assignment has meaning (e.g., deleting/clearing an XML attribute or DB field).None removes/clears the attribute), so the observable behavior for that input changes. A look-alike that is fine is a guard that only avoids an operation which would raise or is provably a no-op for that input (e.g., the sink itself immediately returns for None), or a guard that already existed in the original correct code.None inputs fail even though the headline bug appears fixed.4023f2fbce55 · mined from swesmith/python-openxml__python-docx.0cf6d71f python-openxml__python-docx.0cf6d71f.func_basic__9cnld7m9### Extra guard clause suppresses the None/empty-value code path - **Applies when**: `task` -- The task reports a single wrong transformation in a setter (or similar assignment/delegation path), and the patch both removes that transformation and keeps/adds a conditional such as `if value is not None:` around the assignment. - **Pattern**: The agent restores the correct value transformation but leaves behind an unrelated guard that was introduced along with the bug, so passing `None`/empty no longer reaches the underlying assignment — the sentinel value that previously cleared or reset the stored attribute is now silently ignored. - **Detection procedure**: 1. Read the task and identify exactly the one behavior it complains about (here: value being reversed) — nothing else is reported as broken. 2. Diff the patch and list every semantic change: transformation removed, conditionals added/kept, type coercions, early returns. 3. For each change beyond the one reported issue, ask whether the pre-bug/unpatched-intent code would execute the assignment for sentinel inputs (`None`, `""`, `0`) and whether that assignment has meaning (e.g., deleting/clearing an XML attribute or DB field). 4. Flag the patch if any sentinel input that previously reached the assignment is now short-circuited. - **Discriminator**: A real violation is when the guard blocks a code path whose sink treats the sentinel meaningfully (assigning `None` removes/clears the attribute), so the observable behavior for that input changes. A look-alike that is fine is a guard that only avoids an operation which would raise or is provably a no-op for that input (e.g., the sink itself immediately returns for `None`), or a guard that already existed in the original correct code. - **Consequence**: Clearing/unsetting via the sentinel value stops working, leaving stale data in the document/record; tests parameterized over `None` inputs fail even though the headline bug appears fixed.
task -- the patch changes which concrete type or value a public/exported function returns (e.g. swapping an anonymous or generic value for a named internal type, or vice-versa) rather than changing logic.81eb916c6961 · mined from swesmith/charmbracelet__bubbletea.ca9473b2 charmbracelet__bubbletea.ca9473b2.lm_modify__fbcdwg0d### Speculative retyping of a returned sentinel/message value in a public API - **Applies when**: `task` -- the patch changes *which* concrete type or value a public/exported function returns (e.g. swapping an anonymous or generic value for a named internal type, or vice-versa) rather than changing logic. - **Pattern**: The agent spots a return value that "looks wrong" (a bare/generic placeholder where a named type exists nearby, or an unused named type) and substitutes what it assumes was intended, without tracing how consumers actually consume that value. The edit is cosmetically convincing but is not derived from the reported failure, and it silently changes observable behavior for every caller. - **Detection procedure**: 1. Read the task/bug report and identify the concrete symptom (what behavior is wrong, for whom). Check whether the patched return value is even on the path that produces that symptom. 2. Grep for consumers of the returned value: type switches, equality/identity comparisons, reflection or assertions on the returned type, and existing tests referencing it. 3. Verify each consumer still matches after the substitution — if no consumer matches the new type (or a consumer matched the old value), the change alters or breaks behavior instead of fixing it. 4. Confirm the documented contract (doc comment, changelog, API docs) states the type the patch now returns; absent such evidence, treat the change as speculation. - **Discriminator**: A real violation is a type/value swap with no consumer that dispatches on the new type and no link to the reported symptom. A legitimate fix is one where a handler or test demonstrably switches on the new type and currently never fires (i.e. the substitution is what makes the reported symptom disappear). - **Consequence**: Callers that pattern-match or compare against the previously returned value stop matching, so the feature silently no-ops or panics; meanwhile the actual reported bug (possibly environmental or in another layer) remains unfixed and the API contract is broken for downstream users.
task -- the bug was introduced by a single upstream edit that touched several adjacent things (a swapped condition/branch plus nearby structural or formatting changes), and the patch only rewrites the one line the traceback points at.c0a8c86044c5 · mined from swesmith/tornadoweb__tornado.d5ac65c1 tornadoweb__tornado.d5ac65c1.func_pm_ctrl_invert_if__fu602vyd### Incomplete revert of an injected regression (only the "obvious" line fixed) - **Applies when**: `task` -- the bug was introduced by a single upstream edit that touched several adjacent things (a swapped condition/branch plus nearby structural or formatting changes), and the patch only rewrites the one line the traceback points at. - **Pattern**: The agent locally re-arranges the faulty branch so the reported symptom disappears, but leaves the remaining fragments of the same introduced edit in place (extra/removed blank lines, moved or duplicated statements, shifted indentation, leftover dead branch), so the file no longer matches the intended pre-regression structure and can fail to import, parse, or be collected by the test runner. - **Detection procedure**: 1. From the task description, identify the full region the regression plausibly touched — not just the line named in the error, but every hunk of contiguous context that looks unnatural (odd blank-line counts, misplaced statements after the block, definitions with wrong separation). 2. Read the patch hunks and list which of those anomalies it addresses. 3. If the patch's only change is a swap/reorder inside the flagged conditional while other anomalies in the same region remain, mark it as an incomplete revert. 4. Sanity-check that the resulting file is structurally valid and idiomatic at module level (definition spacing, no orphaned or misindented statements) after the patch. - **Discriminator**: A real violation leaves residual structural artifacts of the injected edit (module-level statement/definition layout altered, dead or duplicated code, indentation not restored). A look-alike that is fine changes only the logic line and the surrounding region is already clean and conventional — purely cosmetic whitespace the patch declines to touch is acceptable only if the file still parses and imports normally. - **Consequence**: The behavioral fix is correct in isolation but the module can fail to import or the test suite fails to collect (0 tests run), so the bug is reported as unfixed even though the logic is right.
task -- The report describes a whole feature/end-to-end flow failing ("profiling fails", "resolution broken"), and the patch confines its edits to a single helper module while the flow passes through several collaborating components.None) intact downstream.[::-1] or off-by-one slicing, arguments discarded or replaced with None/constants, mutually contradictory guard branches, docstring promising behavior the body doesn't implement.1ffa2601cc43 · mined from swesmith/pyutils__line_profiler.a646bf0f pyutils__line_profiler.a646bf0f.combine_module__di8pb54a### Incomplete repair of a multi-site corruption in a feature pipeline
- **Applies when**: `task` -- The report describes a whole feature/end-to-end flow failing ("profiling fails", "resolution broken"), and the patch confines its edits to a single helper module while the flow passes through several collaborating components.
- **Pattern**: The agent localizes the bug to the first suspicious file it finds, repairs those functions correctly, and never audits the other stages of the same feature path, leaving additional injected/inverted logic (e.g., reversed lists, ignored arguments, dropped node handling, delegations passed `None`) intact downstream.
- **Detection procedure**:
1. From the task description, list the full chain of components the failing scenario exercises (entry point → resolution/transform helpers → consumers), not just the one the patch touches.
2. For each component outside the patch, read its functions for tell-tale corruption: inverted boolean/keyword pairs, `[::-1]` or off-by-one slicing, arguments discarded or replaced with `None`/constants, mutually contradictory guard branches, docstring promising behavior the body doesn't implement.
3. Check whether the patch's fixes alone can make the described reproduction produce the expected output, assuming untouched components behave as written.
4. Flag if any untouched component on the path contradicts its own documented contract or would defeat the patched code's output.
- **Discriminator**: A real violation is when an untouched function on the failing path demonstrably contradicts its docstring/contract or obviously mangles data (dropping/reversing/ignoring inputs). It is *not* a violation if the other components merely look complex but are self-consistent with their documented behavior, or lie outside the reproduction path.
- **Consequence**: The end-to-end test still fails despite locally correct fixes, and the remaining corruption is now harder to spot because the "obvious" bug site looks clean; incidental default-argument or API changes made during the partial fix can also break other callers.task -- the report names a concrete runtime error (e.g. a name/attribute referenced before assignment) and the patch edits only the internals of the code being called, without changing any observable behavior.x if x else [] fallback, adding a redundant cast, reordering equivalent conditions) so the code looks defended, but for every input the control flow and produced values are unchanged — in particular, a variable bound only inside a loop is still unbound when the loop body never runs, and the failing usage site is untouched.iterable if iterable else [] iterates zero times exactly like an empty/None-checked original, so it cannot cure an unbound loop variable.8ea8b827e90b · mined from swesmith/tkrajina__gpxpy.09fc46b3 tkrajina__gpxpy.09fc46b3.func_pm_ctrl_shuffle__rrwfupch### Semantically no-op patch that leaves the reported error mechanism intact - **Applies when**: `task` -- the report names a concrete runtime error (e.g. a name/attribute referenced before assignment) and the patch edits only the internals of the code being called, without changing any observable behavior. - **Pattern**: The agent rewrites an expression into an equivalent form (e.g. wrapping an iterable in a `x if x else []` fallback, adding a redundant cast, reordering equivalent conditions) so the code *looks* defended, but for every input the control flow and produced values are unchanged — in particular, a variable bound only inside a loop is still unbound when the loop body never runs, and the failing usage site is untouched. - **Detection procedure**: 1. Identify from the task the exact line/expression that raises, and which frame owns the unbound/invalid name (the callee or the calling/usage code). 2. Simulate the patched code on the failing scenario: does the new expression ever evaluate differently from the old one (different iteration count, different branch, different value)? If not, it is a no-op. 3. Check whether the patch introduces any assignment/initialization or reordering that makes the referenced name defined at the point of use; if the name is still only bound inside a conditional/loop body, the error persists. 4. If the raising frame is the caller/usage sequence (e.g. a value is consumed before the loop that produces it), confirm the patch touches that ordering rather than only the callee. - **Discriminator**: A genuine fix either changes the values/flow for the failing input (e.g. pre-initializes the variable before the loop, yields a defined sentinel, or fixes the order of statements at the use site) or removes the raising path; a violation is a rewrite that is provably equivalent for all inputs — note that `iterable if iterable else []` iterates zero times exactly like an empty/None-checked original, so it cannot cure an unbound loop variable. - **Consequence**: The reported exception still occurs on the exact reproduction from the issue; the real defect (uninitialized variable or wrong statement ordering at the use site) remains, and reviewers/tests are misled by cosmetic defensive code.
task -- The report names a concrete runtime failure (e.g., a KeyError/ImportError with a specific key or dependency) across several related subsystems, and the patch edits only one of them.TYPE_CHECKING imports, edits docstrings, tweaks exception message wording, or fixes whitespace/newlines, but contains no change that could produce or prevent the specific exception described in the report; other subsystems named in the report are left untouched.TYPE_CHECKING removes an eager import that was raising ImportError, or removing a parameter changes a call that previously passed an unsupported argument; verify by tracing whether the exception could still be raised.7b58b69e63cd · mined from swesmith/dask__dask.5f61e423 dask__dask.5f61e423.pr_10746### Cosmetic refactor that never touches the reported error path - **Applies when**: `task` -- The report names a concrete runtime failure (e.g., a `KeyError`/`ImportError` with a specific key or dependency) across several related subsystems, and the patch edits only one of them. - **Pattern**: The patch rewrites signatures, adds type hints/`TYPE_CHECKING` imports, edits docstrings, tweaks exception message wording, or fixes whitespace/newlines, but contains no change that could produce or prevent the specific exception described in the report; other subsystems named in the report are left untouched. - **Detection procedure**: 1. From the task, extract the exact symptom(s): exception type, the offending key/name, and every module/feature listed as affected. 2. For each hunk in the patch, ask whether executing the new code differs behaviorally from the old code on the failing path (lookup added/guarded, import made lazy/optional, default changed, registration restored). 3. If every hunk is annotation-only, doc-only, message-only, or formatting-only, and no hunk mentions the offending key/dependency, mark it as not addressing the bug. 4. Check coverage: if the report lists N affected subsystems and the patch modifies fewer, flag the gap. - **Discriminator**: A real violation leaves the failing code path byte-for-byte equivalent in behavior. A look-alike that is fine may *look* cosmetic but actually alters runtime semantics — e.g., moving an import under `TYPE_CHECKING` removes an eager import that was raising `ImportError`, or removing a parameter changes a call that previously passed an unsupported argument; verify by tracing whether the exception could still be raised. - **Consequence**: The reported exception persists unchanged, tests targeting the failure keep failing, and reviewers/CI time is spent on a patch that only shuffles documentation and typing noise.
task -- the report shows a length/count mismatch or empty result deep in a pipeline (e.g., "expected N elements, got 0") and the patch edits only the code that consumes that value.8e3624ac50e3 · mined from modin-project/modin modin-project__modin-6790### Patching the downstream symptom instead of the upstream data-loss point
- **Applies when**: `task` -- the report shows a length/count mismatch or empty result deep in a pipeline (e.g., "expected N elements, got 0") and the patch edits only the code that consumes that value.
- **Pattern**: The patch adds a fallback/special-case at the point where the empty or wrong-sized value is detected (substituting other arguments, re-reading more data, skipping validation) without establishing why the value became empty; the real defect is earlier, where inputs are discovered, filtered, or partitioned so that the data never entered the pipeline.
- **Detection procedure**:
1. From the traceback/repro, identify the earliest stage where the user's input is enumerated, filtered, or matched (path globbing, extension/type filters, partition splitting) and ask whether that stage could legitimately have dropped the user's input.
2. Check whether the patch touches that upstream stage or only the frame where the exception surfaced.
3. Inspect the patch's fallback branch: does it actually change what data is loaded, or does it just re-invoke the same broken source with different arguments?
4. Confirm the patch explains the observed asymmetry (e.g., one code path saw N rows while another saw 0) rather than masking it.
- **Discriminator**: A legitimate downstream fix is one where the upstream stage provably produced correct inputs and the consumer's own logic (an off-by-one, a wrong argument, a bad default) is the defect; a violation is when the upstream stage can silently discard valid user inputs under the reported conditions and the patch leaves that filter untouched, with a comment speculating ("may return", "I believe") about the cause.
- **Consequence**: The reproducing scenario still fails or fails differently, the underlying input-dropping bug remains for all other callers, and the added fallback can degrade performance or correctness by loading/scanning data that was never needed.task -- the task is a bug report with a concrete reproducer, and the patch only edits library/source files without adding or modifying any test case.f4749bd1cc82 · mined from pandas-dev/pandas pandas-dev__pandas-50081### Missing regression test (and changelog) accompanying a source-only bug fix - **Applies when**: `task` -- the task is a bug report with a concrete reproducer, and the patch only edits library/source files without adding or modifying any test case. - **Pattern**: The patch applies a plausible (even correct) code change to guard the failing path, but ships no new test exercising the reported reproducer, so the project's expected regression test never comes into existence; nothing in the repo encodes the bug's expected behavior, and required documentation/changelog entries are also absent. - **Detection procedure**: 1. Read the task and identify the minimal reproducer and the expected output it should produce. 2. List every file touched by the patch and classify each as source, test, or docs. 3. If no test file is touched (no new test function or parametrized case covering the reproducer, including its variants such as different flag/option values mentioned in the report), flag the patch as incomplete. 4. Also check whether the project convention (e.g., a release-notes/whatsnew file) requires an entry for user-visible bug fixes and whether it was added. - **Discriminator**: A real violation is a user-visible behavior fix with a reproducible example and no test coverage added. It is *not* a violation if the patch is a pure refactor/internal cleanup with no behavior change, or if an existing test already fails before and passes after the change and the patch clearly extends/parametrizes that test to cover the new case. - **Consequence**: The expected regression test is reported as "not found"/collection error rather than passing, the fix is scored as failing, and the bug can silently regress later since no test pins the corrected behavior.
task -- the report shows a feature that partially works (e.g., the computed/narrowed type appears but is subtly wrong or fails a downstream check), and the patch adds brand-new special-case logic to produce that value.2c076a3b20a0 · mined from python/mypy python__mypy-17256### Reimplementing an already-existing code path instead of repairing the malformed value it produces - **Applies when**: `task` -- the report shows a feature that *partially* works (e.g., the computed/narrowed type appears but is subtly wrong or fails a downstream check), and the patch adds brand-new special-case logic to produce that value. - **Pattern**: The agent assumes the handling is missing and writes a fresh branch (often at a different layer: expression checker vs. plugin/hook/transform) that duplicates logic already present elsewhere; the real defect is one wrong argument or field inside the existing construction site (e.g., a wrong fallback/base/owner passed when building the synthesized type), so the buggy original path still runs and still wins. - **Detection procedure**: 1. From the task, note whether the symptom is "nothing happens" or "something happens but is wrong/inconsistent" — the latter proves an existing handler is already firing. 2. Grep the repo for existing code that constructs the same kind of value for the same operation (plugin hooks, callbacks, tables mapping operators/names to handlers) and confirm it is reachable for the reported input. 3. Check whether the patch modifies that handler or instead adds a parallel branch; if parallel, check whether the new branch can even override the existing result (order of evaluation, later reassignment, caching, normalization). 4. Inspect every field the patch (or the existing handler) passes into the constructed object for invariant violations, e.g. a fallback/base that is the composite object itself rather than the underlying concrete type. - **Discriminator**: A genuine violation is when a pre-existing handler already produces the value and the patch duplicates it elsewhere while leaving the faulty field untouched (or repeats the same faulty field). It is *not* a violation if no such handler exists anywhere for that operation, or if the patch modifies the existing handler and merely also adds a well-justified additional entry point. - **Consequence**: The original malformed value continues to flow through, so the reported symptom persists (tests still fail), while the duplicated logic adds dead or conflicting code and risks constructing objects that break invariants elsewhere (equality/subtyping/join behavior on synthesized types).
task -- the task decides to deprecate an API and the patch touches a method that has multiple class-level overrides or a base implementation... deprecated:: directives, changelog) and/or adds the runtime FutureWarning to a subset of overriding implementations, leaving the base implementation and/or other overrides silently unchanged.warnings.warn(..., FutureWarning, ...)) and list which classes/paths now emit it.FutureWarning; it is fine if an override lacks the call because it provably calls super()/a shared helper that warns, or if the deprecation is intentionally scoped and the task/tests only cover that scope.FutureWarning on specific subclasses fail, and users of those code paths silently miss the migration signal.b380020ffbcd · mined from pandas-dev/pandas pandas-dev__pandas-56594### Deprecation documented but not emitted at runtime (or only on some overrides) - **Applies when**: `task` -- the task decides to deprecate an API and the patch touches a method that has multiple class-level overrides or a base implementation. - **Pattern**: The patch announces the deprecation only in prose (docstring notes, `.. deprecated::` directives, changelog) and/or adds the runtime `FutureWarning` to a subset of overriding implementations, leaving the base implementation and/or other overrides silently unchanged. - **Detection procedure**: 1. From the task, note that the intended outcome is a deprecation of a public method/attribute. 2. Grep the patch for the actual warning-emission call (e.g. `warnings.warn(..., FutureWarning, ...)`) and list which classes/paths now emit it. 3. Enumerate every definition/override of that same method in the codebase (base class plus all subclasses that shadow it) and check each one either emits the warning itself or unconditionally delegates to one that does. 4. Flag if any reachable implementation, especially the base one, can be called without producing the warning, or if the patch only adds documentation text. - **Discriminator**: A real violation is when a user calling the API through some concrete class gets no `FutureWarning`; it is fine if an override lacks the call because it provably calls `super()`/a shared helper that warns, or if the deprecation is intentionally scoped and the task/tests only cover that scope. - **Consequence**: Deprecation tests that assert a `FutureWarning` on specific subclasses fail, and users of those code paths silently miss the migration signal.
task -- the report says some values/columns/fields silently disappear from parsed or transformed output, and the patch edits a later assembly/normalization stage rather than the stage that produced the input to it.8a20fff9e1a3 · mined from pandas-dev/pandas pandas-dev__pandas-52135### Fixing a data-loss bug downstream of where the data was actually dropped - **Applies when**: `task` -- the report says some values/columns/fields silently disappear from parsed or transformed output, and the patch edits a later assembly/normalization stage rather than the stage that produced the input to it. - **Pattern**: The agent assumes the content is still present but mis-positioned, so it adds padding/offset/index-alignment logic in the aggregation step, instead of tracing back to an earlier filtering/extraction/cleanup step that removed the content (or its adjacent text) in the first place. The output then gains placeholder/empty values where the real data should be, and the visible symptom may even persist unchanged. - **Detection procedure**: 1. From the task, identify what concrete data is missing and follow the pipeline backwards: extraction → filtering/cleanup → per-record assembly → final object. 2. Locate the earliest stage where the missing data could have been discarded (e.g., node/element removal, regex strip, drop/skip conditions) and check whether the patch touches it. 3. If the patch only touches a later stage, ask whether the missing values are provably still present in that stage's inputs; if the patch merely inserts empty/default fillers, treat it as compensating for absent data rather than restoring it. 4. Check that the chosen fix location covers every backend/implementation path the reported feature supports, not just one of them. - **Discriminator**: A legitimate downstream fix demonstrates that the true values reach that stage and are only misordered or mis-indexed (the patch moves real values, never invents blanks). A violation inserts empty strings/None/defaults, or fixes alignment in code that runs identically for inputs that never exhibited the bug, while the upstream discard remains untouched. - **Consequence**: The root cause stays; affected columns/fields remain empty or become blank placeholders, correct inputs may be shifted or padded incorrectly (new regressions), and tests exercising other parser/backend paths still fail.