perf-rubrics-onpolicy-all

SWE-fficiency · 69 rubrics · 69 rubrics mined from the perf pool (injected slowdowns on SWE-smith repos)
HF EdwardoSunny/perf-rubrics-onpolicy-all · local data/libraries/perf-rubrics-onpolicy-all.json

0Patch is a mechanically unsound edit: the rewritten hot-path block no longer forms valid, importable codetaskalecthomas/voluptuous
Applies when
task -- the patch replaces or restructures a block inside a function of a module that the test suite (and the workload) imports, rather than adding a new isolated helper.
Pattern
The optimization idea is algorithmically right, but the applied hunk leaves the edited region structurally inconsistent — dangling/duplicated statements, wrong indentation relative to the surrounding if/else or loop, a return/branch moved out of its block, or references to names, annotations or imports that the same hunk removed. The module then fails at import/parse time, so nothing that depends on it runs at all.
Detection procedure
  1. From the workload, note which module(s) get imported to reach the hot path; any breakage there is fatal for every test, not just the changed function.
  2. Reconstruct the post-patch text of the whole edited function from the diff context (not just the + lines) and read it top to bottom as if it were a fresh file.
  3. Check mechanically: indentation of each surviving line matches its intended block; every branch still terminates with the value it used to return; every identifier, annotation, and helper referenced is still defined/imported; no leftover statement from the old algorithm remains half-removed.
  4. If the resulting function body is not something you could paste into the file and import cleanly, flag the patch regardless of how good the complexity argument is.
Discriminator
A real violation is a structural defect — the post-patch code cannot parse/import or a branch silently loses its result — which makes the entire suite fail uniformly. A look-alike that is fine is a patch that merely keeps redundant-looking scaffolding (e.g., an intermediate list, a now-unnecessary type annotation, an extra loop that only appends) while all names resolve and the block’s control flow and return values are intact; that is cosmetically clumsy but behaviour-preserving and still delivers the complexity win.
Consequence
The workload script and the covering tests both die during import/collection, showing 0 of N tests passing (an error, not assertion failures), so no timing measurement is even produced — as opposed to the correct version, which keeps all tests green and turns the quadratic combining step into an O(n log n) one.
id 5d3df984d519 · mined from alecthomas/voluptuous alecthomas__voluptuous.perf_2
raw text (what the judge reads)
### Patch is a mechanically unsound edit: the rewritten hot-path block no longer forms valid, importable code
- **Applies when**: `task` -- the patch replaces or restructures a block inside a function of a module that the test suite (and the workload) imports, rather than adding a new isolated helper.
- **Pattern**: The optimization idea is algorithmically right, but the applied hunk leaves the edited region structurally inconsistent — dangling/duplicated statements, wrong indentation relative to the surrounding `if/else` or loop, a `return`/branch moved out of its block, or references to names, annotations or imports that the same hunk removed. The module then fails at import/parse time, so nothing that depends on it runs at all.
- **Detection procedure**:
  1. From the workload, note which module(s) get imported to reach the hot path; any breakage there is fatal for every test, not just the changed function.
  2. Reconstruct the post-patch text of the whole edited function from the diff context (not just the `+` lines) and read it top to bottom as if it were a fresh file.
  3. Check mechanically: indentation of each surviving line matches its intended block; every branch still terminates with the value it used to return; every identifier, annotation, and helper referenced is still defined/imported; no leftover statement from the old algorithm remains half-removed.
  4. If the resulting function body is not something you could paste into the file and import cleanly, flag the patch regardless of how good the complexity argument is.
- **Discriminator**: A real violation is a *structural* defect — the post-patch code cannot parse/import or a branch silently loses its result — which makes the entire suite fail uniformly. A look-alike that is fine is a patch that merely keeps redundant-looking scaffolding (e.g., an intermediate list, a now-unnecessary type annotation, an extra loop that only appends) while all names resolve and the block’s control flow and return values are intact; that is cosmetically clumsy but behaviour-preserving and still delivers the complexity win.
- **Consequence**: The workload script and the covering tests both die during import/collection, showing 0 of N tests passing (an error, not assertion failures), so no timing measurement is even produced — as opposed to the correct version, which keeps all tests green and turns the quadratic combining step into an O(n log n) one.
1Caching a derived value behind an interface that callers still write to or expect to be livetaskmarshmallow-code/webargs
Applies when
task -- the patch removes repeated recomputation by memoizing a derived value (computed from constructor inputs) while leaving the original accessor — a property, method, or attribute — in place as the public entry point.
Pattern
The patch adds a private cached field populated once (usually in __init__) and makes the existing accessor return it, but does not check the accessor's full contract: it may still be read-only (a property with no setter), may hand out a shared mutable object where a fresh one was produced each call, and its inputs may be reassigned or mutated after construction. Callers, subclasses, or tests that assign to the accessor, mutate the returned object, or change the source input now get an exception or a stale value.
Detection procedure
  1. In the task/workload, identify the recomputed accessor named as the hot spot and note whether it is a property, plain attribute, or method, and whether it returns a mutable container.
  2. Grep the repository (library code, subclasses, and tests) for every read and write of that accessor and of the inputs it derives from — assignments, del, in-place mutation of the returned value, and post-construction mutation of the source.
  3. Check the patch: if it converts a writable-looking name into a read-only cached property, or freezes a value whose source is written to anywhere found in step 2, without adding a setter/invalidation or converting the name to a plain settable attribute, flag it.
  4. Confirm the cached value is computed from data that is genuinely fixed at construction time and not from per-call arguments.
Discriminator
A real violation is a cache whose accessor loses a capability it previously had (assignability, freshness after input change, per-call fresh object) with at least one existing caller relying on it. A look-alike that is fine: the value is derived only from immutable constructor state, the accessor is never assigned or its result never mutated anywhere in the codebase, and the patch either keeps it a plain attribute or supplies a setter/invalidation hook.
Consequence
The workload may show the expected speedup, but the covering tests fail broadly — typically AttributeError: can't set attribute at construction or in any test that overrides the value, or subtle wrong results from stale/shared state — turning a pure speed change into a behavioural regression.
id b0b6558f6fbc · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_0
raw text (what the judge reads)
### Caching a derived value behind an interface that callers still write to or expect to be live
- **Applies when**: `task` -- the patch removes repeated recomputation by memoizing a derived value (computed from constructor inputs) while leaving the original accessor — a property, method, or attribute — in place as the public entry point.
- **Pattern**: The patch adds a private cached field populated once (usually in `__init__`) and makes the existing accessor return it, but does not check the accessor's full contract: it may still be read-only (a property with no setter), may hand out a shared mutable object where a fresh one was produced each call, and its inputs may be reassigned or mutated after construction. Callers, subclasses, or tests that assign to the accessor, mutate the returned object, or change the source input now get an exception or a stale value.
- **Detection procedure**:
  1. In the task/workload, identify the recomputed accessor named as the hot spot and note whether it is a property, plain attribute, or method, and whether it returns a mutable container.
  2. Grep the repository (library code, subclasses, and tests) for every read *and write* of that accessor and of the inputs it derives from — assignments, `del`, in-place mutation of the returned value, and post-construction mutation of the source.
  3. Check the patch: if it converts a writable-looking name into a read-only cached property, or freezes a value whose source is written to anywhere found in step 2, without adding a setter/invalidation or converting the name to a plain settable attribute, flag it.
  4. Confirm the cached value is computed from data that is genuinely fixed at construction time and not from per-call arguments.
- **Discriminator**: A real violation is a cache whose accessor loses a capability it previously had (assignability, freshness after input change, per-call fresh object) with at least one existing caller relying on it. A look-alike that is fine: the value is derived only from immutable constructor state, the accessor is never assigned or its result never mutated anywhere in the codebase, and the patch either keeps it a plain attribute or supplies a setter/invalidation hook.
- **Consequence**: The workload may show the expected speedup, but the covering tests fail broadly — typically `AttributeError: can't set attribute` at construction or in any test that overrides the value, or subtle wrong results from stale/shared state — turning a pure speed change into a behavioural regression.
2Replacing a hand-rolled loop in a shared low-level helper with a library "fast constructor" whose accepted input types / edge-case semantics are narrowertaskpyca/pyopenssl
Applies when
task -- the profile points at per-byte Python work inside a small, widely-shared helper (buffer copying, encoding, conversion) and the patch swaps the explicit loop for a single built-in/FFI/library call that does the same thing in C.
Pattern
The patch assumes the one-shot constructor is a drop-in equivalent of the loop, but the loop tolerated inputs and corner cases the new call does not (bytearray/memoryview/other buffer-protocol objects, str, empty input, sentinel/terminator or length conventions, ownership/lifetime of the allocated memory). Because the helper sits under every entry point, a mismatch does not slow one path down — it makes the whole module unusable, which is why the entire covering suite can go from all-pass to zero-pass.
Detection procedure
  1. From the task and workload, identify the helper being edited and note how central it is: grep its call sites and ask "does every public operation, not just the timed one, flow through here?"
  2. Enumerate the input shapes those call sites (and the repo's tests/docs for them) actually pass: exact type(s), zero-length values, oversized values, and any implicit contract about terminators, reported length, or who owns/keeps alive the produced object.
  3. For each such shape, check the replacement call's documented behaviour — does it accept the type at all, allocate the same size, keep the buffer alive for the consumer's lifetime, and return the same length semantics? Any "probably fine" answer is a violation.
  4. Sanity-check that the edited module still imports and one non-timed operation still works, before trusting any timing number.
Discriminator
A real violation is when the new call's contract is provably narrower or different for at least one input shape reachable from a call site (e.g. only bytes accepted where callers may pass a mutable buffer, or the length/terminator convention shifts). It is not a violation when the surrounding code already normalizes/validates the input type before the helper is reached and the constructor's size, termination, and lifetime guarantees are documented to match — then the C-level replacement is exactly the intended fix.
Consequence
The workload may show the expected speedup (it exercises only the happy-path input), while the covering tests collapse wholesale — import or every-test errors rather than a few localized failures — signalling that a universally used helper's contract was broken for speed.
id 24e9ce529b79 · mined from pyca/pyopenssl pyca__pyopenssl.perf_1
raw text (what the judge reads)
### Replacing a hand-rolled loop in a shared low-level helper with a library "fast constructor" whose accepted input types / edge-case semantics are narrower

- **Applies when**: `task` -- the profile points at per-byte Python work inside a small, widely-shared helper (buffer copying, encoding, conversion) and the patch swaps the explicit loop for a single built-in/FFI/library call that does the same thing in C.
- **Pattern**: The patch assumes the one-shot constructor is a drop-in equivalent of the loop, but the loop tolerated inputs and corner cases the new call does not (bytearray/memoryview/other buffer-protocol objects, `str`, empty input, sentinel/terminator or length conventions, ownership/lifetime of the allocated memory). Because the helper sits under *every* entry point, a mismatch does not slow one path down — it makes the whole module unusable, which is why the entire covering suite can go from all-pass to zero-pass.
- **Detection procedure**:
  1. From the task and workload, identify the helper being edited and note how central it is: grep its call sites and ask "does every public operation, not just the timed one, flow through here?"
  2. Enumerate the input shapes those call sites (and the repo's tests/docs for them) actually pass: exact type(s), zero-length values, oversized values, and any implicit contract about terminators, reported length, or who owns/keeps alive the produced object.
  3. For each such shape, check the replacement call's documented behaviour — does it accept the type at all, allocate the same size, keep the buffer alive for the consumer's lifetime, and return the same length semantics? Any "probably fine" answer is a violation.
  4. Sanity-check that the edited module still imports and one non-timed operation still works, before trusting any timing number.
- **Discriminator**: A real violation is when the new call's contract is provably narrower or different for at least one input shape reachable from a call site (e.g. only `bytes` accepted where callers may pass a mutable buffer, or the length/terminator convention shifts). It is *not* a violation when the surrounding code already normalizes/validates the input type before the helper is reached and the constructor's size, termination, and lifetime guarantees are documented to match — then the C-level replacement is exactly the intended fix.
- **Consequence**: The workload may show the expected speedup (it exercises only the happy-path input), while the covering tests collapse wholesale — import or every-test errors rather than a few localized failures — signalling that a universally used helper's contract was broken for speed.
3Removing eager materialization of a lazy sequence, and early-exiting from it, without proving the full traversal wasn't load-bearingtaskandialbrecht/sqlparse
Applies when
task -- the patch speeds up a lookup by deleting a list(...)/tuple snapshot of a lazily produced sequence (generator, recursive tree walk, streaming iterator) and instead iterates it directly with an early return/break once the target is found.
Pattern
The rewrite is asymptotically better on paper (drops a quadratic prefix re-scan), but it silently changes when and how much of the producer runs: the producer is now advanced partially, entered re-entrantly, or interleaved with the caller's own work, and any side effect, shared cursor, lazy-initialization, or mutation-during-iteration that the eager snapshot used to isolate now leaks into unrelated behaviour.
Detection procedure
  1. From the task, note which shared/core API is being edited and how many tests are declared as "covering" it — a large covering set signals the function's data source is central shared machinery, not a leaf helper.
  2. Read the producer that was previously materialized: does it recurse over a structure that other code mutates, trigger lazy parsing/grouping/caching, keep per-object iteration state, or get consumed elsewhere while this call is in progress?
  3. Check the new control flow: the early return abandons the producer mid-stream, and the accumulator state (running offset, index) is now derived incrementally rather than recomputed — confirm both hold for every exit path, including "not found" and nested/re-entrant calls.
  4. If steps 2–3 leave any doubt, demand that the patch either keep a cheap materialization (still linear, e.g. one snapshot per call) or add evidence that partial consumption is safe; a patch that offers neither is inadequate before any timing is done.
Discriminator
Not a violation when the producer is a pure, side-effect-free read of an immutable/finished structure and no other code holds a live iterator over it — then partial consumption is genuinely equivalent and faster. It is a violation when the producer performs lazy construction, mutates or is mutated during traversal, or shares iteration state, so the old "consume everything into a list first" step was semantically load-bearing rather than merely wasteful.
Consequence
The workload may look faster, but the covering suite degrades far beyond the edited feature — collection/import-time or near-total failures across unrelated tests, because half-consumed lazy traversal leaves shared structures in an inconsistent state.
id be720c596805 · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_2
raw text (what the judge reads)
### Removing eager materialization of a lazy sequence, and early-exiting from it, without proving the full traversal wasn't load-bearing
- **Applies when**: `task` -- the patch speeds up a lookup by deleting a `list(...)`/tuple snapshot of a lazily produced sequence (generator, recursive tree walk, streaming iterator) and instead iterates it directly with an early `return`/`break` once the target is found.
- **Pattern**: The rewrite is asymptotically better on paper (drops a quadratic prefix re-scan), but it silently changes *when and how much* of the producer runs: the producer is now advanced partially, entered re-entrantly, or interleaved with the caller's own work, and any side effect, shared cursor, lazy-initialization, or mutation-during-iteration that the eager snapshot used to isolate now leaks into unrelated behaviour.
- **Detection procedure**:
  1. From the task, note which shared/core API is being edited and how many tests are declared as "covering" it — a large covering set signals the function's data source is central shared machinery, not a leaf helper.
  2. Read the producer that was previously materialized: does it recurse over a structure that other code mutates, trigger lazy parsing/grouping/caching, keep per-object iteration state, or get consumed elsewhere while this call is in progress?
  3. Check the new control flow: the early `return` abandons the producer mid-stream, and the accumulator state (running offset, index) is now derived incrementally rather than recomputed — confirm both hold for *every* exit path, including "not found" and nested/re-entrant calls.
  4. If steps 2–3 leave any doubt, demand that the patch either keep a cheap materialization (still linear, e.g. one snapshot per call) or add evidence that partial consumption is safe; a patch that offers neither is inadequate before any timing is done.
- **Discriminator**: Not a violation when the producer is a pure, side-effect-free read of an immutable/finished structure and no other code holds a live iterator over it — then partial consumption is genuinely equivalent and faster. It *is* a violation when the producer performs lazy construction, mutates or is mutated during traversal, or shares iteration state, so the old "consume everything into a list first" step was semantically load-bearing rather than merely wasteful.
- **Consequence**: The workload may look faster, but the covering suite degrades far beyond the edited feature — collection/import-time or near-total failures across unrelated tests, because half-consumed lazy traversal leaves shared structures in an inconsistent state.
4Unrelated collateral edits bundled with the hot-path change (and hand-rolled logic where a tested helper exists)taskandialbrecht/sqlparse
Applies when
task -- the patch is supposed to be a small, behaviour-preserving speedup of one lookup, but its diff touches lines that have nothing to do with the hot path (blank lines, indentation, decorators, imports, surrounding definitions) and/or re-implements inline a scan/search that the codebase already exposes as a shared helper.
Pattern
The agent makes the intended micro-optimization (e.g. replacing a full materialized scan with an early-exit loop) but also silently rewrites adjacent module structure/whitespace, so the file's top-level layout changes; and it inlines logic that an existing, already-tested primitive performs, duplicating semantics (skip rules, edge cases) instead of delegating.
Detection procedure
  1. From the task, identify the single function/expression that the workload times, and note the minimal edit needed there.
  2. Walk every hunk of the diff and classify each changed line as "required for the speedup" or "collateral" (formatting, blank-line count, class/def boundaries, imports, unrelated statements).
  3. Reject if any collateral line changes module-level structure or style that import-time checks, linters, or style tests in the suite could enforce — such edits can fail collection and take the entire suite down, independent of the optimization's merits.
  4. Check whether the rewritten logic duplicates an existing helper in the same module/class that already implements the same "first meaningful element" / skip semantics; if so, require delegation to it rather than a private copy.
Discriminator
A real violation is a diff whose non-essential lines alter file structure or duplicate semantics maintained elsewhere. A look-alike that is fine is a diff that only touches the timed expression (even if the new code is longer), or one that reformats lines it had to rewrite anyway, with no change to top-level blank-line/definition layout and no duplication of an existing helper's rules.
Consequence
The workload may well get faster, but the covering tests fail wholesale (0 passed — collection/style gate failure) or drift later as the duplicated skip logic diverges from the shared helper, so the change is rejected despite the speedup.
id 6eeca17c6203 · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_3
raw text (what the judge reads)
### Unrelated collateral edits bundled with the hot-path change (and hand-rolled logic where a tested helper exists)
- **Applies when**: `task` -- the patch is supposed to be a small, behaviour-preserving speedup of one lookup, but its diff touches lines that have nothing to do with the hot path (blank lines, indentation, decorators, imports, surrounding definitions) and/or re-implements inline a scan/search that the codebase already exposes as a shared helper.
- **Pattern**: The agent makes the intended micro-optimization (e.g. replacing a full materialized scan with an early-exit loop) but also silently rewrites adjacent module structure/whitespace, so the file's top-level layout changes; and it inlines logic that an existing, already-tested primitive performs, duplicating semantics (skip rules, edge cases) instead of delegating.
- **Detection procedure**:
  1. From the task, identify the single function/expression that the workload times, and note the minimal edit needed there.
  2. Walk every hunk of the diff and classify each changed line as "required for the speedup" or "collateral" (formatting, blank-line count, class/def boundaries, imports, unrelated statements).
  3. Reject if any collateral line changes module-level structure or style that import-time checks, linters, or style tests in the suite could enforce — such edits can fail collection and take the entire suite down, independent of the optimization's merits.
  4. Check whether the rewritten logic duplicates an existing helper in the same module/class that already implements the same "first meaningful element" / skip semantics; if so, require delegation to it rather than a private copy.
- **Discriminator**: A real violation is a diff whose non-essential lines alter file structure or duplicate semantics maintained elsewhere. A look-alike that is fine is a diff that only touches the timed expression (even if the new code is longer), or one that reformats lines it had to rewrite anyway, with no change to top-level blank-line/definition layout and no duplication of an existing helper's rules.
- **Consequence**: The workload may well get faster, but the covering tests fail wholesale (0 passed — collection/style gate failure) or drift later as the duplicated skip logic diverges from the shared helper, so the change is rejected despite the speedup.
5New helper/module referenced without ensuring it is imported in the patched scopetaskInstagram/MonkeyType
Applies when
task -- a patch replaces a hand-rolled hot loop with a standard-library helper (e.g. a counter, cache, or itertools/heapq utility) or a new internal helper name.
Pattern
The rewrite is algorithmically correct, but the symbol it now calls is not bound in the edited module: the diff adds a use of somemodule.something / SomeHelper while touching only the function body and never adding (or verifying the pre-existing presence of) the corresponding import/from ... import. The result is a NameError/AttributeError at the first call, so every test exercising that code path — and often the whole test module — fails.
Detection procedure
  1. In the diff, list every name introduced on the added lines that is not a local variable, parameter, or builtin.
  2. For each such name, check whether the diff itself adds the import, or whether the unchanged header of the same file (context lines / the file as it exists) already imports it — do not assume it does because the name is "standard library".
  3. If the name's binding cannot be pointed to in either place, treat the patch as broken regardless of how plausible the optimization is.
  4. Also confirm the added call's signature/return shape matches how the result is consumed (e.g. iterating .items() on something that is actually a mapping).
Discriminator
A real violation is an unbound name — no import in the diff and none in the file. A look-alike that is fine is a name that is already imported at module top (or is a builtin such as len, dict, sorted, set), even though the diff does not mention the import; also fine is a fully-qualified use where the parent package is imported.
Consequence
The workload and the covering tests fail immediately at the first invocation with NameError: name '...' is not defined (typically all tests in the touched module report failure), so no speedup can be measured at all.
id a2c2565dd3aa · mined from Instagram/MonkeyType Instagram__MonkeyType.perf_3
raw text (what the judge reads)
### New helper/module referenced without ensuring it is imported in the patched scope
- **Applies when**: `task` -- a patch replaces a hand-rolled hot loop with a standard-library helper (e.g. a counter, cache, or itertools/heapq utility) or a new internal helper name.
- **Pattern**: The rewrite is algorithmically correct, but the symbol it now calls is not bound in the edited module: the diff adds a use of `somemodule.something` / `SomeHelper` while touching only the function body and never adding (or verifying the pre-existing presence of) the corresponding `import`/`from ... import`. The result is a `NameError`/`AttributeError` at the first call, so every test exercising that code path — and often the whole test module — fails.
- **Detection procedure**:
  1. In the diff, list every name introduced on the added lines that is not a local variable, parameter, or builtin.
  2. For each such name, check whether the diff itself adds the import, or whether the unchanged header of the same file (context lines / the file as it exists) already imports it — do not assume it does because the name is "standard library".
  3. If the name's binding cannot be pointed to in either place, treat the patch as broken regardless of how plausible the optimization is.
  4. Also confirm the added call's signature/return shape matches how the result is consumed (e.g. iterating `.items()` on something that is actually a mapping).
- **Discriminator**: A real violation is an unbound name — no import in the diff and none in the file. A look-alike that is fine is a name that *is* already imported at module top (or is a builtin such as `len`, `dict`, `sorted`, `set`), even though the diff does not mention the import; also fine is a fully-qualified use where the parent package is imported.
- **Consequence**: The workload and the covering tests fail immediately at the first invocation with `NameError: name '...' is not defined` (typically all tests in the touched module report failure), so no speedup can be measured at all.
6Hoisting per-request setup that silently drops request-dependent construction parameterstaskmarshmallow-code/webargs
Applies when
task -- the task asks to move repeated "build the schema/validator/parser object" work out of the per-call path into decoration/initialization time, and the patch constructs that object eagerly from a static spec.
Pattern
The patch pre-builds, at decoration/import time, an object that the runtime path previously built from a plain spec — but the runtime builder also injected contextual options (e.g. unknown-field policy derived from the data location, partial/many flags, per-call overrides, error handling config). The eager construction uses only default options, so the object now behaves differently, and the runtime path no longer recognizes the spec type well enough to re-apply the missing options.
Detection procedure
  1. In the unpatched code, find the runtime function that consumed the raw spec and note every argument it passed into the constructor, including values derived from the call site (location, request attributes, flags, defaults resolved lazily).
  2. Compare with the patch's eager construction site: list which of those arguments are unavailable or omitted there.
  3. Check whether the runtime path has a branch that treats an already-built object differently (skipping the option-injection step) — if so, the omitted options are now permanently lost.
  4. Look for a behavioural knob among the omitted options (unknown/extra-field handling, validation strictness, partial loading); any such omission is a correctness change, not a pure speedup.
Discriminator
A real violation omits or freezes an option that varies per call or per location, or that differs from the constructor's default; a benign look-alike hoists construction while explicitly threading through the same options (or defers construction until the first call, when the context is known) so the built object is provably identical to what the per-call path produced.
Consequence
The workload may indeed get faster (its single fixed configuration happens to match), but the covering tests fail broadly — typically every test through the decorated path errors or returns different validation results (e.g. extra/unknown inputs now rejected or silently kept), i.e. speed bought by changing semantics.
id 4b5dc68cb4d0 · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_2
raw text (what the judge reads)
### Hoisting per-request setup that silently drops request-dependent construction parameters

- **Applies when**: `task` -- the task asks to move repeated "build the schema/validator/parser object" work out of the per-call path into decoration/initialization time, and the patch constructs that object eagerly from a static spec.
- **Pattern**: The patch pre-builds, at decoration/import time, an object that the runtime path previously built from a plain spec — but the runtime builder also injected contextual options (e.g. unknown-field policy derived from the data location, partial/many flags, per-call overrides, error handling config). The eager construction uses only default options, so the object now behaves differently, and the runtime path no longer recognizes the spec type well enough to re-apply the missing options.
- **Detection procedure**:
  1. In the unpatched code, find the runtime function that consumed the raw spec and note *every* argument it passed into the constructor, including values derived from the call site (location, request attributes, flags, defaults resolved lazily).
  2. Compare with the patch's eager construction site: list which of those arguments are unavailable or omitted there.
  3. Check whether the runtime path has a branch that treats an already-built object differently (skipping the option-injection step) — if so, the omitted options are now permanently lost.
  4. Look for a behavioural knob among the omitted options (unknown/extra-field handling, validation strictness, partial loading); any such omission is a correctness change, not a pure speedup.
- **Discriminator**: A real violation omits or freezes an option that varies per call or per location, or that differs from the constructor's default; a benign look-alike hoists construction while explicitly threading through the same options (or defers construction until the first call, when the context is known) so the built object is provably identical to what the per-call path produced.
- **Consequence**: The workload may indeed get faster (its single fixed configuration happens to match), but the covering tests fail broadly — typically every test through the decorated path errors or returns different validation results (e.g. extra/unknown inputs now rejected or silently kept), i.e. speed bought by changing semantics.
7Cache of per-class introspection stored where inheritance (or dynamic attributes) can leak the wrong entrytaskInstagram/MonkeyType
Applies when
task -- the hot path builds a lookup table by introspecting an object (dir, __dict__, attribute scans) and the patch memoizes that table on the class/type instead of recomputing it.
Pattern
The patch checks for the cache with an inheritance-visible lookup (hasattr(cls, ...) / getattr(cls, ...)) and stores it with setattr(cls, ...), so the first class in a hierarchy to populate the cache silently supplies its table to every base/subclass — and any attribute added after first use is never seen. The dispatch table is then wrong for exactly the objects whose behaviour depends on subclass overrides or extra hooks.
Detection procedure
  1. From the task/workload, identify which concrete object the hot path runs on and whether that type is a subclass (or has siblings) of the class whose introspection is being cached.
  2. In the patch, find where the cache is read and written; check whether the read can resolve to an attribute defined on a different class (base or sibling) than the one being written to, and whether a sentinel/cls.__dict__ check is used.
  3. Check whether the cached table can go stale: are the introspected attributes ever added/overridden per instance, per subclass, or at runtime after the first call?
  4. If the read is inheritance-visible or the table can go stale, treat the dispatch as potentially resolving to the wrong handler (or None) for some types.
Discriminator
Fine: cache keyed on cls.__dict__ presence or a module-level dict keyed by the exact type(self), populated from the full MRO, on a table that cannot change after class creation. Violation: hasattr(cls, name) / plain class attribute lookup as the presence check, or any scheme where one class's table can be served to another class in the hierarchy. Also fine is not caching at all if the lookup was replaced by a single direct attribute fetch (which is inherently correct under inheritance).
Consequence
Handlers resolve to the wrong (or missing) methods for subclasses, so the recursion falls through to fallback/error paths — the workload's rendering output changes or raises, and the covering tests fail broadly (near-total failure), even though the microbenchmark may look faster.
id bdbf6ff2027a · mined from Instagram/MonkeyType Instagram__MonkeyType.perf_2
raw text (what the judge reads)
### Cache of per-class introspection stored where inheritance (or dynamic attributes) can leak the wrong entry
- **Applies when**: `task` -- the hot path builds a lookup table by introspecting an object (`dir`, `__dict__`, attribute scans) and the patch memoizes that table on the class/type instead of recomputing it.
- **Pattern**: The patch checks for the cache with an inheritance-visible lookup (`hasattr(cls, ...)` / `getattr(cls, ...)`) and stores it with `setattr(cls, ...)`, so the first class in a hierarchy to populate the cache silently supplies its table to every base/subclass — and any attribute added after first use is never seen. The dispatch table is then wrong for exactly the objects whose behaviour depends on subclass overrides or extra hooks.
- **Detection procedure**:
  1. From the task/workload, identify which concrete object the hot path runs on and whether that type is a subclass (or has siblings) of the class whose introspection is being cached.
  2. In the patch, find where the cache is read and written; check whether the read can resolve to an attribute defined on a *different* class (base or sibling) than the one being written to, and whether a sentinel/`cls.__dict__` check is used.
  3. Check whether the cached table can go stale: are the introspected attributes ever added/overridden per instance, per subclass, or at runtime after the first call?
  4. If the read is inheritance-visible or the table can go stale, treat the dispatch as potentially resolving to the wrong handler (or `None`) for some types.
- **Discriminator**: Fine: cache keyed on `cls.__dict__` presence or a module-level dict keyed by the exact `type(self)`, populated from the full MRO, on a table that cannot change after class creation. Violation: `hasattr(cls, name)` / plain class attribute lookup as the presence check, or any scheme where one class's table can be served to another class in the hierarchy. Also fine is not caching at all if the lookup was replaced by a single direct attribute fetch (which is inherently correct under inheritance).
- **Consequence**: Handlers resolve to the wrong (or missing) methods for subclasses, so the recursion falls through to fallback/error paths — the workload's rendering output changes or raises, and the covering tests fail broadly (near-total failure), even though the microbenchmark may look faster.
8Opting into an internal compile/dispatch protocol instead of reusing already-precomputed worktaskalecthomas/voluptuous
Applies when
task -- the task asks to cut per-item overhead in a reusable validator/handler object, and the patch adds a hook method that makes the surrounding framework treat the object via a different (lower-level) calling convention, storing "compiled" state on the instance.
Pattern
The patch discovers that the object rebuilds work per call, but instead of hoisting that work into state already computed at construction time, it implements an undocumented framework extension hook (__*_compile__-style), returns a callback with a different signature (extra path/context args, different error/return contract), and caches the compiled callbacks on self — mixing two call conventions in one object and mutating shared instance state at schema-build time.
Detection procedure
  1. From the task, note that the expensive work is redundant construction that is already done and stored elsewhere in the object's constructor (or is trivially hoistable) — i.e., a one-line reuse would suffice.
  2. In the patch, check whether it instead opts the object into a framework-internal dispatch/compile protocol, changing the signature or error-wrapping semantics of the invoked callable relative to the plain path.
  3. Check whether the new path duplicates the validation/error semantics of the old path by hand (length checks, message substitution, exception re-wrapping, container type reconstruction, index/path bookkeeping) and whether any of those differ from the original.
  4. Check whether compiled state is stored on self and thus shared across every schema/thread that embeds the same instance, and whether the object can now be called both ways with different behaviour.
Discriminator
A real violation adds a new call path/protocol whose semantics must be re-derived and can diverge (or whose cached state leaks across users), when the same speedup was obtainable by reusing existing precomputed members. It is fine to hoist per-call construction into constructor-time state, or to memoize purely derived data, as long as the object's calling convention, return values, and raised-error types/messages are byte-for-byte unchanged and no state is shared across independent embeddings.
Consequence
Because the new protocol sits on a framework-wide dispatch path, a signature or error-contract mismatch fails far beyond the targeted case — the covering suite collapses almost entirely (here 0/146 passed), while the workload gain, if any, was achievable by a one-line reuse of existing precomputed validators.
id 622e8069d439 · mined from alecthomas/voluptuous alecthomas__voluptuous.perf_1
raw text (what the judge reads)
### Opting into an internal compile/dispatch protocol instead of reusing already-precomputed work
- **Applies when**: `task` -- the task asks to cut per-item overhead in a reusable validator/handler object, and the patch adds a hook method that makes the surrounding framework treat the object via a different (lower-level) calling convention, storing "compiled" state on the instance.
- **Pattern**: The patch discovers that the object rebuilds work per call, but instead of hoisting that work into state already computed at construction time, it implements an undocumented framework extension hook (`__*_compile__`-style), returns a callback with a different signature (extra path/context args, different error/return contract), and caches the compiled callbacks on `self` — mixing two call conventions in one object and mutating shared instance state at schema-build time.
- **Detection procedure**:
  1. From the task, note that the expensive work is redundant construction that is *already* done and stored elsewhere in the object's constructor (or is trivially hoistable) — i.e., a one-line reuse would suffice.
  2. In the patch, check whether it instead opts the object into a framework-internal dispatch/compile protocol, changing the signature or error-wrapping semantics of the invoked callable relative to the plain path.
  3. Check whether the new path duplicates the validation/error semantics of the old path by hand (length checks, message substitution, exception re-wrapping, container type reconstruction, index/path bookkeeping) and whether any of those differ from the original.
  4. Check whether compiled state is stored on `self` and thus shared across every schema/thread that embeds the same instance, and whether the object can now be called both ways with different behaviour.
- **Discriminator**: A real violation adds a *new* call path/protocol whose semantics must be re-derived and can diverge (or whose cached state leaks across users), when the same speedup was obtainable by reusing existing precomputed members. It is fine to hoist per-call construction into constructor-time state, or to memoize purely derived data, as long as the object's calling convention, return values, and raised-error types/messages are byte-for-byte unchanged and no state is shared across independent embeddings.
- **Consequence**: Because the new protocol sits on a framework-wide dispatch path, a signature or error-contract mismatch fails far beyond the targeted case — the covering suite collapses almost entirely (here 0/146 passed), while the workload gain, if any, was achievable by a one-line reuse of existing precomputed validators.
9Hoisting a statement out of a loop by re-indentation, without verifying the block structure still parses/behavestaskiterative/dvc
Applies when
task -- the speedup is achieved by moving an existing statement (a sort, flush, recompute, aggregation) from inside a loop to after it, i.e. the diff is essentially a change of indentation/scope rather than new logic.
Pattern
The patch relies on a purely mechanical scope change: one or more lines are de-indented (or re-nested) so they run once instead of per iteration. Nothing else is added, so the author assumes it is trivially safe and does no local syntax/semantic check — but the new indentation may not line up with the enclosing block (continuation lines, closing brackets, nested for/with/if), or the hoisted result may be needed inside the loop.
Detection procedure
  1. Read the task/workload to confirm the hot cost is the per-iteration repetition of that statement (quadratic growth is a strong hint), so hoisting is the right idea in principle.
  2. Reconstruct the full post-patch block from the diff, including unchanged context lines and bracket continuations, and check every indentation level: does the de-indented statement land at exactly the enclosing function/loop level, does it still follow a complete statement, and is it inside the right for/with/try?
  3. Check whether any code inside the loop reads the state the hoisted statement produced (ordering, flushed buffer, partial aggregate); if so, hoisting changes observable results, not just timing.
  4. Require evidence that the edited module still imports (a parse/import check or a single fast test) — a scope edit that cannot be shown to compile is not reviewable as "obviously equivalent".
Discriminator
A legitimate hoist has the moved statement at an indentation level that provably matches the enclosing block, with no in-loop consumer of its effect, so the final value is identical and only computed once. A violation is a hoist whose indentation cannot be reconciled with the surrounding lines/brackets (possible IndentationError/wrong nesting), or where the loop body depended on the per-iteration effect.
Consequence
If the block structure is broken, the module fails at import and every covering test errors out (0 passed) regardless of the workload; if an in-loop consumer existed, the workload gets faster but results/ordering change and the behaviour tests fail.
id 894d48b5e3ce · mined from iterative/dvc iterative__dvc.perf_2
raw text (what the judge reads)
### Hoisting a statement out of a loop by re-indentation, without verifying the block structure still parses/behaves
- **Applies when**: `task` -- the speedup is achieved by moving an existing statement (a sort, flush, recompute, aggregation) from inside a loop to after it, i.e. the diff is essentially a change of indentation/scope rather than new logic.
- **Pattern**: The patch relies on a purely mechanical scope change: one or more lines are de-indented (or re-nested) so they run once instead of per iteration. Nothing else is added, so the author assumes it is trivially safe and does no local syntax/semantic check — but the new indentation may not line up with the enclosing block (continuation lines, closing brackets, nested `for`/`with`/`if`), or the hoisted result may be needed inside the loop.
- **Detection procedure**:
  1. Read the task/workload to confirm the hot cost is the per-iteration repetition of that statement (quadratic growth is a strong hint), so hoisting is the right idea in principle.
  2. Reconstruct the *full* post-patch block from the diff, including unchanged context lines and bracket continuations, and check every indentation level: does the de-indented statement land at exactly the enclosing function/loop level, does it still follow a complete statement, and is it inside the right `for`/`with`/`try`?
  3. Check whether any code inside the loop reads the state the hoisted statement produced (ordering, flushed buffer, partial aggregate); if so, hoisting changes observable results, not just timing.
  4. Require evidence that the edited module still imports (a parse/`import` check or a single fast test) — a scope edit that cannot be shown to compile is not reviewable as "obviously equivalent".
- **Discriminator**: A legitimate hoist has the moved statement at an indentation level that provably matches the enclosing block, with no in-loop consumer of its effect, so the final value is identical and only computed once. A violation is a hoist whose indentation cannot be reconciled with the surrounding lines/brackets (possible `IndentationError`/wrong nesting), or where the loop body depended on the per-iteration effect.
- **Consequence**: If the block structure is broken, the module fails at import and *every* covering test errors out (0 passed) regardless of the workload; if an in-loop consumer existed, the workload gets faster but results/ordering change and the behaviour tests fail.
10Guarded "skip-the-work" fast path bolted onto a hot transform instead of one canonical fast primitivetaskbenoitc/gunicorn
Applies when
task -- the report says a per-element transform (escaping/normalizing/copying strings or buffers) costs time proportional to total input size, and the patch keeps the element loop but wraps the transform in a conditional that decides whether to transform at all.
Pattern
The patch leaves the original iteration structure and adds a predicate branch (e.g. "does this value contain the special character?") with two code paths: one that does the transform, one that stores/returns the input untouched. The expensive primitive (naive per-character concatenation, repeated allocation) is either kept in one branch or replaced only there, so the win comes from a case-split rather than from replacing the quadratic operation with a single built-in call — and now two branches must independently preserve semantics, types, and object identity.
Detection procedure
  1. From the task and workload, identify the actual cost driver (here: cost scaling with total value length, i.e. a character-by-character accumulation) and check whether the workload's data even hits the "no special character" fast case for most bytes.
  2. In the patch, check whether the underlying expensive primitive was actually replaced by one bulk/built-in operation on every path, or merely bypassed for a subset of inputs; note that standard bulk string/bytes operations already short-circuit when nothing matches, making the guard a redundant extra scan.
  3. Compare both branches against the original for exact output equivalence: same value, same type, same escaping of every edge case (empty, non-str, already-escaped, aliased input object returned instead of a fresh copy).
  4. Verify the patched file is still structurally intact (indentation, no orphaned comment/statement fragments from the removed loop) and that the module still imports — an import-time break makes every covering test fail regardless of the speedup.
Discriminator
A real violation is a branch that only avoids work the underlying primitive would already avoid cheaply, or whose "untouched" path can differ from the transformed path for some input; it is acceptable if the guard skips a genuinely expensive, non-short-circuiting operation (e.g. an unavoidable large copy or an external call) and both branches are provably identical in output and type for all inputs.
Consequence
Either the workload shows little or no improvement (the guard adds a second pass and the slow primitive remains on the path the workload exercises), or the restructured/duplicated code diverges from the original semantics — in the worst case the module no longer imports and the entire covering test set fails.
id 0bfdb7d843c9 · mined from benoitc/gunicorn benoitc__gunicorn.perf_1
raw text (what the judge reads)
### Guarded "skip-the-work" fast path bolted onto a hot transform instead of one canonical fast primitive
- **Applies when**: `task` -- the report says a per-element transform (escaping/normalizing/copying strings or buffers) costs time proportional to total input size, and the patch keeps the element loop but wraps the transform in a conditional that decides whether to transform at all.
- **Pattern**: The patch leaves the original iteration structure and adds a predicate branch (e.g. "does this value contain the special character?") with two code paths: one that does the transform, one that stores/returns the input untouched. The expensive primitive (naive per-character concatenation, repeated allocation) is either kept in one branch or replaced only there, so the win comes from a case-split rather than from replacing the quadratic operation with a single built-in call — and now two branches must independently preserve semantics, types, and object identity.
- **Detection procedure**:
  1. From the task and workload, identify the actual cost driver (here: cost scaling with total value length, i.e. a character-by-character accumulation) and check whether the workload's data even hits the "no special character" fast case for most bytes.
  2. In the patch, check whether the underlying expensive primitive was actually replaced by one bulk/built-in operation on every path, or merely bypassed for a subset of inputs; note that standard bulk string/bytes operations already short-circuit when nothing matches, making the guard a redundant extra scan.
  3. Compare both branches against the original for exact output equivalence: same value, same type, same escaping of every edge case (empty, non-str, already-escaped, aliased input object returned instead of a fresh copy).
  4. Verify the patched file is still structurally intact (indentation, no orphaned comment/statement fragments from the removed loop) and that the module still imports — an import-time break makes every covering test fail regardless of the speedup.
- **Discriminator**: A real violation is a branch that only avoids work the underlying primitive would already avoid cheaply, or whose "untouched" path can differ from the transformed path for some input; it is acceptable if the guard skips a genuinely expensive, non-short-circuiting operation (e.g. an unavoidable large copy or an external call) and both branches are provably identical in output and type for all inputs.
- **Consequence**: Either the workload shows little or no improvement (the guard adds a second pass and the slow primitive remains on the path the workload exercises), or the restructured/duplicated code diverges from the original semantics — in the worst case the module no longer imports and the entire covering test set fails.
11Patch that corrupts the surrounding module rather than just the hot looptaskgetnikola/nikola
Applies when
task -- a patch replaces an inner scanning loop inside a shared, widely-imported utility module with a direct lookup, using a hand-written diff whose hunk boundaries and indentation must line up with existing code.
Pattern
The intended logical change is correct and equivalent, but the diff itself is structurally unsound: hunk headers/line counts don't match the listed lines, removed lines take neighbouring statements or def/for headers with them, replacement lines sit at the wrong indentation level, or a name that later code still references is deleted. Applying it leaves the module unparseable or with a dangling/mis-scoped statement — which breaks import of the module for every test, not just the changed behaviour.
Detection procedure
  1. From the task, note that the edited symbol lives in a broadly imported utility module, so any import-time error fails all covering tests at collection time.
  2. Reconstruct the post-patch text of the touched region by hand: apply each -/+ line in order and check the hunk's declared old/new line counts equal the number of context+- and context++ lines actually shown.
  3. Verify the resulting block parses and is well-scoped: every for/if/def still has a body at the right indentation, no statement is orphaned outside its intended block, and no removed assignment is still referenced below.
  4. Confirm the change is limited to the loop being replaced — nothing outside the identified hot loop (other functions, module-level defs, adjacent helpers) is added, removed, or re-indented.
Discriminator
A real violation is a diff that, once applied, cannot import cleanly or shifts code out of its block — recognisable from mismatched hunk counts, indentation changes on untouched lines, or references to deleted names. A look-alike that is fine is a diff whose reconstructed region is syntactically identical in structure to the original with only the inner scan replaced by a membership test, and whose hunk counts add up.
Consequence
The workload script and the entire covering test file fail at import/collection (0 of N tests pass) with a SyntaxError/IndentationError/NameError, so no timing number is ever produced even though the intended optimization was logically correct.
id 7df7c1324394 · mined from getnikola/nikola getnikola__nikola.perf_1
raw text (what the judge reads)
### Patch that corrupts the surrounding module rather than just the hot loop
- **Applies when**: `task` -- a patch replaces an inner scanning loop inside a shared, widely-imported utility module with a direct lookup, using a hand-written diff whose hunk boundaries and indentation must line up with existing code.
- **Pattern**: The intended logical change is correct and equivalent, but the diff itself is structurally unsound: hunk headers/line counts don't match the listed lines, removed lines take neighbouring statements or `def`/`for` headers with them, replacement lines sit at the wrong indentation level, or a name that later code still references is deleted. Applying it leaves the module unparseable or with a dangling/mis-scoped statement — which breaks *import* of the module for every test, not just the changed behaviour.
- **Detection procedure**:
  1. From the task, note that the edited symbol lives in a broadly imported utility module, so any import-time error fails all covering tests at collection time.
  2. Reconstruct the post-patch text of the touched region by hand: apply each `-`/`+` line in order and check the hunk's declared old/new line counts equal the number of context+`-` and context+`+` lines actually shown.
  3. Verify the resulting block parses and is well-scoped: every `for`/`if`/`def` still has a body at the right indentation, no statement is orphaned outside its intended block, and no removed assignment is still referenced below.
  4. Confirm the change is limited to the loop being replaced — nothing outside the identified hot loop (other functions, module-level defs, adjacent helpers) is added, removed, or re-indented.
- **Discriminator**: A real violation is a diff that, once applied, cannot import cleanly or shifts code out of its block — recognisable from mismatched hunk counts, indentation changes on untouched lines, or references to deleted names. A look-alike that is fine is a diff whose reconstructed region is syntactically identical in structure to the original with only the inner scan replaced by a membership test, and whose hunk counts add up.
- **Consequence**: The workload script and the entire covering test file fail at import/collection (0 of N tests pass) with a `SyntaxError`/`IndentationError`/`NameError`, so no timing number is ever produced even though the intended optimization was logically correct.
12Index-swap patch whose edited region is never reconstructed and re-verified as a wholetaskalecthomas/voluptuous
Applies when
task -- the speedup replaces a per-item linear scan over a precompiled structure with a once-built lookup table (dict/index), editing the middle of an existing function or closure.
Pattern
The patch inserts the new precomputed index next to (rather than fully in place of) the old scanning logic, deleting only part of the original body and leaving the surrounding region inconsistent: orphaned docstrings/fragments, a helper that no longer returns anything on some path, names that are defined in one branch but referenced in another, duplicated construction of the same candidate lists, or an indentation/scope shift that changes what is a closure over what. Because the index is built on dict hashing, it also silently swaps ==-based matching (and its ordering) for hash-based matching without checking that every key type in the structure is hashable and hash/eq-consistent.
Detection procedure
  1. From the workload, identify the precompiled structure and the per-item loop that dominates cost; confirm the patch targets exactly that loop.
  2. Reconstruct the full post-patch text of the edited function/closure from the diff (not just the added lines) and check it parses and is semantically complete: no leftover fragments of the removed code, every branch of every helper returns a value, every referenced name is defined before use in that scope, no duplicate/dead construction of the same data.
  3. Compare the matching semantics: does the original select entries by == (allowing unhashable, custom-__eq__, or 1/True-style keys, and preserving a specific order) while the new code requires the key to be hashable and hash-equal? Look for wrapper/marker objects or non-primitive keys that would now miss or land in a different bucket.
  4. Check that the "fallback/wildcard" set and the fast-path set together still cover every entry exactly as before, with the same relative ordering.
Discriminator
A genuine violation is a diff where the replaced region, read as a whole, no longer forms valid/complete code, or where the dict index cannot represent some key kind the old == scan handled. A look-alike that is fine is a diff that removes the scan entirely, builds the index with the same partition and ordering rules, and keeps a wildcard fallback list for every entry that is not an eligible literal — even though it also uses dict hashing, because all keys placed in the index are provably hashable primitives.
Consequence
The module fails to import or the closure raises on the first call, so the entire covering test suite errors out at collection (0 tests passing) and the workload never reaches a timed iteration; in the milder hash-vs-eq variant, only the exotic-key tests fail while the timing looks improved.
id 2c5997814472 · mined from alecthomas/voluptuous alecthomas__voluptuous.perf_0
raw text (what the judge reads)
### Index-swap patch whose edited region is never reconstructed and re-verified as a whole
- **Applies when**: `task` -- the speedup replaces a per-item linear scan over a precompiled structure with a once-built lookup table (dict/index), editing the middle of an existing function or closure.
- **Pattern**: The patch inserts the new precomputed index next to (rather than fully in place of) the old scanning logic, deleting only part of the original body and leaving the surrounding region inconsistent: orphaned docstrings/fragments, a helper that no longer returns anything on some path, names that are defined in one branch but referenced in another, duplicated construction of the same candidate lists, or an indentation/scope shift that changes what is a closure over what. Because the index is built on `dict` hashing, it also silently swaps `==`-based matching (and its ordering) for hash-based matching without checking that every key type in the structure is hashable and hash/eq-consistent.
- **Detection procedure**:
  1. From the workload, identify the precompiled structure and the per-item loop that dominates cost; confirm the patch targets exactly that loop.
  2. Reconstruct the *full post-patch text* of the edited function/closure from the diff (not just the added lines) and check it parses and is semantically complete: no leftover fragments of the removed code, every branch of every helper returns a value, every referenced name is defined before use in that scope, no duplicate/dead construction of the same data.
  3. Compare the matching semantics: does the original select entries by `==` (allowing unhashable, custom-`__eq__`, or `1`/`True`-style keys, and preserving a specific order) while the new code requires the key to be hashable and hash-equal? Look for wrapper/marker objects or non-primitive keys that would now miss or land in a different bucket.
  4. Check that the "fallback/wildcard" set and the fast-path set together still cover every entry exactly as before, with the same relative ordering.
- **Discriminator**: A genuine violation is a diff where the replaced region, read as a whole, no longer forms valid/complete code, or where the dict index cannot represent some key kind the old `==` scan handled. A look-alike that is fine is a diff that removes the scan entirely, builds the index with the same partition and ordering rules, and keeps a wildcard fallback list for every entry that is not an eligible literal — even though it also uses dict hashing, because all keys placed in the index are provably hashable primitives.
- **Consequence**: The module fails to import or the closure raises on the first call, so the entire covering test suite errors out at collection (0 tests passing) and the workload never reaches a timed iteration; in the milder hash-vs-eq variant, only the exotic-key tests fail while the timing looks improved.
13Removing per-iteration re-acquisition of a manually managed native resource without guaranteeing the owner stays alivetaskpyca/pyopenssl
Applies when
task -- the hot code repeatedly re-acquires/decodes a foreign-memory or otherwise manually freed object (FFI pointer, buffer, handle, file/lock) inside a loop, and the patch hoists that acquisition out of the loop to remove the quadratic cost.
Pattern
The patch keeps the loop body reading pointers/views derived from the hoisted object, but the owning reference is no longer guaranteed to be live (rebound, shadowed, returned from a helper that is dropped, or its GC/finalizer wrapper discarded), so the underlying memory can be freed while derived pointers are still dereferenced — a use-after-free rather than a semantic change.
Detection procedure
  1. In the pre-patch code, identify which object owns the native/managed lifetime (the one wrapped by a free/close/gc/finalizer call) and which values in the loop are merely pointers or views into it.
  2. In the patched code, trace the owning reference from acquisition through the last dereference of any derived value: is it bound to a live local for that entire span, never rebound or shadowed, and is the destructor registration still applied to the object actually retained?
  3. Check that the loop still obtains derived values from the retained owner (not from a stale/second decode) and that nothing needed for lifetime was deleted along with the "redundant" work (e.g. an unused-looking assignment or import that was actually the keep-alive).
  4. Sanity-check that the edited module still parses/imports and one cheap covering test can run in-process before trusting any timing numbers.
Discriminator
A legitimate hoist keeps exactly one owning reference alive in the enclosing scope for the whole loop and only removes duplicate acquisition work; a violation removes or relocates the keep-alive/destructor binding, or leaves the derived pointers rooted in an object whose only reference has gone out of scope — even though the diff looks like a pure "do it once instead of N times" change.
Consequence
The workload may appear fast or crash nondeterministically, and because a use-after-free aborts or corrupts the interpreter process, the whole covering test session fails (0 passed / segfault / collection error) instead of a single assertion.
id 5685541814d3 · mined from pyca/pyopenssl pyca__pyopenssl.perf_2
raw text (what the judge reads)
### Removing per-iteration re-acquisition of a manually managed native resource without guaranteeing the owner stays alive
- **Applies when**: `task` -- the hot code repeatedly re-acquires/decodes a foreign-memory or otherwise manually freed object (FFI pointer, buffer, handle, file/lock) inside a loop, and the patch hoists that acquisition out of the loop to remove the quadratic cost.
- **Pattern**: The patch keeps the loop body reading pointers/views *derived* from the hoisted object, but the owning reference is no longer guaranteed to be live (rebound, shadowed, returned from a helper that is dropped, or its GC/finalizer wrapper discarded), so the underlying memory can be freed while derived pointers are still dereferenced — a use-after-free rather than a semantic change.
- **Detection procedure**:
  1. In the pre-patch code, identify which object owns the native/managed lifetime (the one wrapped by a free/close/`gc`/finalizer call) and which values in the loop are merely pointers or views into it.
  2. In the patched code, trace the owning reference from acquisition through the last dereference of any derived value: is it bound to a live local for that entire span, never rebound or shadowed, and is the destructor registration still applied to the object actually retained?
  3. Check that the loop still obtains derived values from the retained owner (not from a stale/second decode) and that nothing needed for lifetime was deleted along with the "redundant" work (e.g. an unused-looking assignment or import that was actually the keep-alive).
  4. Sanity-check that the edited module still parses/imports and one cheap covering test can run in-process before trusting any timing numbers.
- **Discriminator**: A legitimate hoist keeps exactly one owning reference alive in the enclosing scope for the whole loop and only removes duplicate *acquisition* work; a violation removes or relocates the keep-alive/destructor binding, or leaves the derived pointers rooted in an object whose only reference has gone out of scope — even though the diff looks like a pure "do it once instead of N times" change.
- **Consequence**: The workload may appear fast or crash nondeterministically, and because a use-after-free aborts or corrupts the interpreter process, the whole covering test session fails (0 passed / segfault / collection error) instead of a single assertion.
14Collateral edits outside the optimized code that risk module-level breakagetaskCog-Creators/Red-DiscordBot
Applies when
task -- a patch that speeds up one small helper also touches adjacent lines (nearby comments, imports, decorators, blank-line/indentation structure, other definitions) that the new implementation does not require.
Pattern
The core algorithmic swap is sound, but the diff bundles unnecessary deletions/edits around the target region, so the enclosing module can end up syntactically or semantically altered (missing import, dropped decorator, mis-indented body, shifted context) — which breaks every test that imports the module, not just the optimized behaviour.
Detection procedure
  1. Identify the minimal set of lines that must change to implement the stated speedup (the new expression/loop and any imports it genuinely needs).
  2. Diff that minimal set against the actual patch hunks; list every added/removed line that is not required by the optimization.
  3. For each extra removed/edited line, ask whether anything at module scope depended on it (an import used elsewhere, a decorator, a name defined for other functions, the indentation/blank-line structure separating definitions).
  4. Check that all names used by the new implementation are still imported/defined after the patch, and that the surrounding definitions remain intact and correctly indented.
Discriminator
Deleting a purely descriptive comment or now-dead import that nothing else references is fine; a real violation is when the extra edits remove or displace something the module still needs (or produce a hunk whose context no longer matches, so the applied result differs from the intended file), such that the module may fail to import or another definition changes.
Consequence
The workload may still show the expected speedup for the isolated snippet, but the covering test file fails wholesale at collection/import time (e.g., 0 of N tests pass) rather than failing a single behavioural assertion.
id 99c5ec74e377 · mined from Cog-Creators/Red-DiscordBot Cog-Creators__Red-DiscordBot.perf_0
raw text (what the judge reads)
### Collateral edits outside the optimized code that risk module-level breakage
- **Applies when**: `task` -- a patch that speeds up one small helper also touches adjacent lines (nearby comments, imports, decorators, blank-line/indentation structure, other definitions) that the new implementation does not require.
- **Pattern**: The core algorithmic swap is sound, but the diff bundles unnecessary deletions/edits around the target region, so the enclosing module can end up syntactically or semantically altered (missing import, dropped decorator, mis-indented body, shifted context) — which breaks *every* test that imports the module, not just the optimized behaviour.
- **Detection procedure**:
  1. Identify the minimal set of lines that must change to implement the stated speedup (the new expression/loop and any imports it genuinely needs).
  2. Diff that minimal set against the actual patch hunks; list every added/removed line that is *not* required by the optimization.
  3. For each extra removed/edited line, ask whether anything at module scope depended on it (an import used elsewhere, a decorator, a name defined for other functions, the indentation/blank-line structure separating definitions).
  4. Check that all names used by the new implementation are still imported/defined after the patch, and that the surrounding definitions remain intact and correctly indented.
- **Discriminator**: Deleting a purely descriptive comment or now-dead import that nothing else references is fine; a real violation is when the extra edits remove or displace something the module still needs (or produce a hunk whose context no longer matches, so the applied result differs from the intended file), such that the module may fail to import or another definition changes.
- **Consequence**: The workload may still show the expected speedup for the isolated snippet, but the covering test file fails wholesale at collection/import time (e.g., 0 of N tests pass) rather than failing a single behavioural assertion.
15Removing a defensive materialization of a shared/one-shot collection to save a copytaskgetnikola/nikola
Applies when
task -- the hot code copies a container into a local list (or set) before doing repeated membership/iteration, and the patch deletes the copy and tests directly against the original expression.
Pattern
The patch assumes the source expression is a stable, re-iterable, hash-based container, so it swaps "build snapshot once, then query it" for "query the live object each time". This is only safe if the object is re-iterable, unaffected by consumption, not mutated during the loop, and already offers the intended lookup complexity — none of which the patch establishes; it also does not necessarily change the asymptotics the task complains about.
Detection procedure
  1. Read the task's complexity claim (what the cost must become a function of) and note what the workload feeds in.
  2. In the patch, identify the removed materialization and trace where the underlying object comes from at all call sites/types (dict? list? generator, map/filter, dict.keys() view, lazily built or mutated during iteration?), not just the type the timing script happens to pass.
  3. Check the new membership/iteration is (a) semantically identical for every one of those types — especially one-shot iterators, which are silently exhausted after the first test — and (b) actually O(1)/O(#matches), not a linear scan that merely removes a constant-factor copy.
  4. If either check depends on an unverified runtime type assumption, treat the patch as unsound.
Discriminator
A real violation is a patch that relies on the live object being re-iterable/hash-indexed without any guarantee (annotation, construction site, or added coercion). A look-alike that is fine explicitly guarantees the type — e.g. the object is constructed as a dict/set in the same module, or the patch itself builds the index once outside the loop and reuses it.
Consequence
The covering tests fail wholesale (empty or missing results wherever the source is consumed or mutated), or the timing script shows little/no improvement because the linear scan simply moved rather than disappeared.
id 3a49893082b6 · mined from getnikola/nikola getnikola__nikola.perf_2
raw text (what the judge reads)
### Removing a defensive materialization of a shared/one-shot collection to save a copy
- **Applies when**: `task` -- the hot code copies a container into a local list (or set) before doing repeated membership/iteration, and the patch deletes the copy and tests directly against the original expression.
- **Pattern**: The patch assumes the source expression is a stable, re-iterable, hash-based container, so it swaps "build snapshot once, then query it" for "query the live object each time". This is only safe if the object is re-iterable, unaffected by consumption, not mutated during the loop, and already offers the intended lookup complexity — none of which the patch establishes; it also does not necessarily change the asymptotics the task complains about.
- **Detection procedure**:
  1. Read the task's complexity claim (what the cost must become a function of) and note what the workload feeds in.
  2. In the patch, identify the removed materialization and trace where the underlying object comes from at *all* call sites/types (dict? list? generator, `map`/`filter`, `dict.keys()` view, lazily built or mutated during iteration?), not just the type the timing script happens to pass.
  3. Check the new membership/iteration is (a) semantically identical for every one of those types — especially one-shot iterators, which are silently exhausted after the first test — and (b) actually O(1)/O(#matches), not a linear scan that merely removes a constant-factor copy.
  4. If either check depends on an unverified runtime type assumption, treat the patch as unsound.
- **Discriminator**: A real violation is a patch that relies on the live object being re-iterable/hash-indexed without any guarantee (annotation, construction site, or added coercion). A look-alike that is fine explicitly guarantees the type — e.g. the object is constructed as a dict/set in the same module, or the patch itself builds the index once outside the loop and reuses it.
- **Consequence**: The covering tests fail wholesale (empty or missing results wherever the source is consumed or mutated), or the timing script shows little/no improvement because the linear scan simply moved rather than disappeared.
16Hoisting a value to setup time changes its type, but the edit is not carried through the whole filetaskmarshmallow-code/webargs
Applies when
task -- the patch moves an expensive object construction out of a per-call path into a one-time setup/decoration path, so a variable that used to hold a raw input (dict/list/config) now holds a fully-built object.
Pattern
The patch rewrites only the block where the construction happened (often deleting an inline conditional expression that used to build the object lazily) and assumes the rest of the module is type-agnostic. It leaves behind stale isinstance/truthiness branches, other consumers of the variable, or a half-edited multi-line expression, so the edited region is no longer valid or no longer consistent with the surrounding code.
Detection procedure
  1. Read the task and workload to identify the value whose construction is being hoisted, and note its type before and after the hoist.
  2. Reconstruct the post-patch source of the edited region as final code (not as a diff): check that every multi-line expression the patch truncated still has balanced parentheses, correct indentation, and no dangling arguments — a malformed module makes the whole test suite fail at import/collection, not just a few cases.
  3. Grep the file/module for every other read of that variable (later branches, closures created inside the decorator, helper calls, alternate entry points such as kwargs-style wrappers) and confirm each one accepts the new post-hoist type, and that no branch keyed on the old type is now dead or mis-taken.
  4. Confirm the hoist happens before all uses, including any validation/error paths that run earlier in the same function, so ordering of raised errors is unchanged.
Discriminator
A real violation is a hoist where at least one surviving consumer, conditional, or syntactic construct still assumes the pre-hoist type/shape, or where the edited expression is left incomplete; a look-alike that is fine is a hoist where the newly built object is the only thing ever read afterwards, every consumer already handles that type, and the edited region reads as complete, valid code with the same error ordering.
Consequence
The workload may or may not get faster, but the covering tests fail en masse (often 0 passing, because the module can no longer be imported or every decorated entry point raises), i.e. speed is bought with a broken build rather than with real work removed.
id 0521de051750 · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_3
raw text (what the judge reads)
### Hoisting a value to setup time changes its type, but the edit is not carried through the whole file
- **Applies when**: `task` -- the patch moves an expensive object construction out of a per-call path into a one-time setup/decoration path, so a variable that used to hold a raw input (dict/list/config) now holds a fully-built object.
- **Pattern**: The patch rewrites only the block where the construction happened (often deleting an inline conditional expression that used to build the object lazily) and assumes the rest of the module is type-agnostic. It leaves behind stale `isinstance`/truthiness branches, other consumers of the variable, or a half-edited multi-line expression, so the edited region is no longer valid or no longer consistent with the surrounding code.
- **Detection procedure**:
  1. Read the task and workload to identify the value whose construction is being hoisted, and note its type *before* and *after* the hoist.
  2. Reconstruct the post-patch source of the edited region as final code (not as a diff): check that every multi-line expression the patch truncated still has balanced parentheses, correct indentation, and no dangling arguments — a malformed module makes the whole test suite fail at import/collection, not just a few cases.
  3. Grep the file/module for every other read of that variable (later branches, closures created inside the decorator, helper calls, alternate entry points such as kwargs-style wrappers) and confirm each one accepts the new post-hoist type, and that no branch keyed on the old type is now dead or mis-taken.
  4. Confirm the hoist happens before *all* uses, including any validation/error paths that run earlier in the same function, so ordering of raised errors is unchanged.
- **Discriminator**: A real violation is a hoist where at least one surviving consumer, conditional, or syntactic construct still assumes the pre-hoist type/shape, or where the edited expression is left incomplete; a look-alike that is fine is a hoist where the newly built object is the only thing ever read afterwards, every consumer already handles that type, and the edited region reads as complete, valid code with the same error ordering.
- **Consequence**: The workload may or may not get faster, but the covering tests fail en masse (often 0 passing, because the module can no longer be imported or every decorated entry point raises), i.e. speed is bought with a broken build rather than with real work removed.
17Replacing a linear equality scan with a hash-map index without preserving lookup semanticstaskalecthomas/voluptuous
Applies when
task -- the hot path is an O(n·m) loop that finds a matching entry by scanning a collection and comparing derived values, and the patch replaces it with a dict/set index keyed on those derived values.
Pattern
The patch builds {derive(x): x for x in collection} and looks entries up with .get(...)/in, silently assuming (a) every derived value is hashable and hashes consistently with ==, and (b) None/falsy can stand in for the "not found" sentinel the original code used explicitly; it may also add manual bookkeeping to keep the index in sync with mutations of the underlying collection.
Detection procedure
  1. Read the original scan and note exactly what it compared, what it returned when nothing matched (explicit sentinel object vs. None), and what types the derived values can be in real inputs (literals, callables, containers, custom objects with __eq__).
  2. Check whether the new index requires hashability and __hash__/__eq__ consistency that the scan never required, and whether any legitimate entry or derived value could be None/falsy so that .get() collapses "absent" and "present but None".
  3. If the patch mutates the collection inside the loop, trace every insert/delete/replace branch and confirm the index is updated identically, including the case where the new key's derived value differs from the old one or where duplicate derived values exist.
  4. Confirm the patch is syntactically/structurally intact after the removal of the old helper (no orphaned references, no indentation/branch mismatch) by mentally re-reading the whole edited block, not just the diff hunks.
Discriminator
A fine version keeps the explicit not-found sentinel (or uses a marker unique to the index), only indexes values the original code already required to be hashable, and either builds the index once over an immutable snapshot or provably updates it on every mutation branch; a violation changes the not-found contract, narrows accepted key types, or leaves index and collection able to diverge.
Consequence
The workload may get faster, but the covering tests fail broadly — a TypeError: unhashable type or an import/collection-time error takes out the whole test module, or subtle cases (falsy keys, duplicate/renamed keys, nested merges) resolve to the wrong entry and change validation behaviour.
id a7b7ad664160 · mined from alecthomas/voluptuous alecthomas__voluptuous.perf_3
raw text (what the judge reads)
### Replacing a linear equality scan with a hash-map index without preserving lookup semantics
- **Applies when**: `task` -- the hot path is an O(n·m) loop that finds a matching entry by scanning a collection and comparing derived values, and the patch replaces it with a dict/set index keyed on those derived values.
- **Pattern**: The patch builds `{derive(x): x for x in collection}` and looks entries up with `.get(...)`/`in`, silently assuming (a) every derived value is hashable and hashes consistently with `==`, and (b) `None`/falsy can stand in for the "not found" sentinel the original code used explicitly; it may also add manual bookkeeping to keep the index in sync with mutations of the underlying collection.
- **Detection procedure**:
  1. Read the original scan and note exactly what it compared, what it returned when nothing matched (explicit sentinel object vs. `None`), and what types the derived values can be in real inputs (literals, callables, containers, custom objects with `__eq__`).
  2. Check whether the new index requires hashability and `__hash__`/`__eq__` consistency that the scan never required, and whether any legitimate entry or derived value could be `None`/falsy so that `.get()` collapses "absent" and "present but None".
  3. If the patch mutates the collection inside the loop, trace every insert/delete/replace branch and confirm the index is updated identically, including the case where the new key's derived value differs from the old one or where duplicate derived values exist.
  4. Confirm the patch is syntactically/structurally intact after the removal of the old helper (no orphaned references, no indentation/branch mismatch) by mentally re-reading the whole edited block, not just the diff hunks.
- **Discriminator**: A fine version keeps the explicit not-found sentinel (or uses a marker unique to the index), only indexes values the original code already required to be hashable, and either builds the index once over an immutable snapshot or provably updates it on every mutation branch; a violation changes the not-found contract, narrows accepted key types, or leaves index and collection able to diverge.
- **Consequence**: The workload may get faster, but the covering tests fail broadly — a `TypeError: unhashable type` or an import/collection-time error takes out the whole test module, or subtle cases (falsy keys, duplicate/renamed keys, nested merges) resolve to the wrong entry and change validation behaviour.
18Adding a memoization decorator without verifying it is legal at that point in the moduletaskjawah/charset_normalizer
Applies when
task -- the patch's speedup consists of dropping a caching/memoizing decorator onto an existing hot function, leaving its body untouched.
Pattern
The one-line decorator is treated as risk-free, but it is evaluated at import time and silently changes the function's contract: the decorator itself or its size/maxsize argument may not be in scope at that position in the file (imported later, imported lazily inside a function, or defined below the decorated definition), the wrapped object no longer accepts the argument shapes callers pass (unhashable/mutable args), or callers/tests rely on the raw function object (reassignment, monkeypatching, attribute access, per-call side effects). Result: import or collection of the whole package fails, or every caller misbehaves — a total correctness loss, not a partial one.
Detection procedure
  1. From the workload and the task description, confirm the decorated function really is the repeated hot path (per-character/per-item lookup here) so the cache is on-path at all.
  2. Read the target file top-to-bottom: check that the decorator name and every symbol used in its arguments are bound at module scope above the decorated definition, not imported inside another function or in a deferred/conditional block.
  3. Check the function's inputs and usage: are all arguments hashable and small in cardinality; is the return value shared mutable state; is the function called at import time, re-exported, monkeypatched, or introspected anywhere (including tests) in a way a wrapper would break; is cache size bounded by a real key-space bound.
  4. Check whether the body itself is left doing unbounded work per distinct key (e.g., a full scan that materializes all matches instead of returning the first) — the cache hides repeats but the cold path and memory still pay for it.
Discriminator
A genuine violation is when any of the decorator's dependencies are unresolvable/misordered at decoration time, or a caller depends on the undecorated object — this yields near-total test failure, not a slowdown. A look-alike that is fine: the decorator and its bound maxsize are already imported and used elsewhere in the same module, arguments are hashable single scalars, and no caller touches the function object — then the plain decorator is a legitimate (if partial) optimization.
Consequence
The workload and the covering tests fail wholesale (import/NameError at collection or every assertion off), so the measured runtime is meaningless; even if it imports, the untouched scan body means first-call cost and memory remain high and the speedup is smaller than the task's stated target.
id 0e26a7e606b8 · mined from jawah/charset_normalizer jawah__charset_normalizer.perf_0
raw text (what the judge reads)
### Adding a memoization decorator without verifying it is legal at that point in the module
- **Applies when**: `task` -- the patch's speedup consists of dropping a caching/memoizing decorator onto an existing hot function, leaving its body untouched.
- **Pattern**: The one-line decorator is treated as risk-free, but it is evaluated at import time and silently changes the function's contract: the decorator itself or its size/maxsize argument may not be in scope at that position in the file (imported later, imported lazily inside a function, or defined below the decorated definition), the wrapped object no longer accepts the argument shapes callers pass (unhashable/mutable args), or callers/tests rely on the raw function object (reassignment, monkeypatching, attribute access, per-call side effects). Result: import or collection of the whole package fails, or every caller misbehaves — a total correctness loss, not a partial one.
- **Detection procedure**:
  1. From the workload and the task description, confirm the decorated function really is the repeated hot path (per-character/per-item lookup here) so the cache is on-path at all.
  2. Read the target file top-to-bottom: check that the decorator name *and* every symbol used in its arguments are bound at module scope *above* the decorated definition, not imported inside another function or in a deferred/conditional block.
  3. Check the function's inputs and usage: are all arguments hashable and small in cardinality; is the return value shared mutable state; is the function called at import time, re-exported, monkeypatched, or introspected anywhere (including tests) in a way a wrapper would break; is cache size bounded by a real key-space bound.
  4. Check whether the body itself is left doing unbounded work per *distinct* key (e.g., a full scan that materializes all matches instead of returning the first) — the cache hides repeats but the cold path and memory still pay for it.
- **Discriminator**: A genuine violation is when any of the decorator's dependencies are unresolvable/misordered at decoration time, or a caller depends on the undecorated object — this yields near-total test failure, not a slowdown. A look-alike that is fine: the decorator and its bound maxsize are already imported and used elsewhere in the same module, arguments are hashable single scalars, and no caller touches the function object — then the plain decorator is a legitimate (if partial) optimization.
- **Consequence**: The workload and the covering tests fail wholesale (import/NameError at collection or every assertion off), so the measured runtime is meaningless; even if it imports, the untouched scan body means first-call cost and memory remain high and the speedup is smaller than the task's stated target.
19Loop-invariant hoisting that leaves the surrounding block inconsistent (scope/indentation) rather than merely re-orderedtasktheskumar/python-dotenv
Applies when
task -- the speedup is achieved by moving a repeated, loop-invariant computation (dict/list build, snapshot, lookup table) out of an inner loop inside an existing nested block.
Pattern
The patch is presented as pure code motion, but the moved statements are re-indented into a different nesting level or conditional branch than the code that consumes them, so the consumer's binding is no longer guaranteed (or the block structure/indentation no longer parses as intended). The diff hunk looks like a clean hoist because the moved lines are textually identical; only the surrounding indentation changed.
Detection procedure
  1. From the task, identify the hot inner loop and the value being hoisted, and note which statements consume it.
  2. Reconstruct the full post-patch function body from the diff (not just the changed lines), writing out the exact indentation of every line in the affected block.
  3. Check that (a) the block still nests as before — the loop being optimized is still inside the same outer loop/branch, and (b) every consumer of the hoisted name is dominated by its new definition on all reachable paths, including the None/empty/early-return branches.
  4. Confirm the hoisted computation is still executed once per outer iteration when it legitimately depends on outer-iteration state (e.g. an accumulator updated as the outer loop progresses) — hoisting it above that loop would freeze stale state.
Discriminator
A legitimate hoist keeps the consumer inside the same loop and merely lifts the definition to the immediately enclosing block, with all uses still dominated and the nesting otherwise byte-for-byte preserved; a violation shifts the loop or its consumers to a different nesting level, or lifts a definition past the point where the state it snapshots is still being mutated. If you cannot state the post-patch nesting for every line without guessing, treat it as a violation.
Consequence
The module may fail to import or the function may raise/return wrong values on the very first call, so the entire covering test set fails (not just an edge case), and the workload never gets to demonstrate the intended speedup.
id f6ee5a4d6645 · mined from theskumar/python-dotenv theskumar__python-dotenv.perf_1
raw text (what the judge reads)
### Loop-invariant hoisting that leaves the surrounding block inconsistent (scope/indentation) rather than merely re-ordered
- **Applies when**: `task` -- the speedup is achieved by moving a repeated, loop-invariant computation (dict/list build, snapshot, lookup table) out of an inner loop inside an existing nested block.
- **Pattern**: The patch is presented as pure code motion, but the moved statements are re-indented into a different nesting level or conditional branch than the code that consumes them, so the consumer's binding is no longer guaranteed (or the block structure/indentation no longer parses as intended). The diff hunk *looks* like a clean hoist because the moved lines are textually identical; only the surrounding indentation changed.
- **Detection procedure**:
  1. From the task, identify the hot inner loop and the value being hoisted, and note which statements consume it.
  2. Reconstruct the *full* post-patch function body from the diff (not just the changed lines), writing out the exact indentation of every line in the affected block.
  3. Check that (a) the block still nests as before — the loop being optimized is still inside the same outer loop/branch, and (b) every consumer of the hoisted name is dominated by its new definition on all reachable paths, including the `None`/empty/early-return branches.
  4. Confirm the hoisted computation is still executed once per outer iteration when it legitimately depends on outer-iteration state (e.g. an accumulator updated as the outer loop progresses) — hoisting it above that loop would freeze stale state.
- **Discriminator**: A legitimate hoist keeps the consumer inside the same loop and merely lifts the definition to the immediately enclosing block, with all uses still dominated and the nesting otherwise byte-for-byte preserved; a violation shifts the loop or its consumers to a different nesting level, or lifts a definition past the point where the state it snapshots is still being mutated. If you cannot state the post-patch nesting for every line without guessing, treat it as a violation.
- **Consequence**: The module may fail to import or the function may raise/return wrong values on the very first call, so the entire covering test set fails (not just an edge case), and the workload never gets to demonstrate the intended speedup.
20Restructuring a hot incremental state machine more than necessary, without a parse/equivalence audittaskjawah/charset_normalizer
Applies when
task -- the task's slowdown is "re-scanning the accumulated buffer instead of just the previous element", and the patch replaces the accumulate-and-rescan code with new per-element state fields and a reshaped control flow.
Pattern
Instead of the minimal edit (keep one "previous element" variable, compare it with the current one, then update it), the patch rewrites the whole method body: it introduces several parallel state fields (some never read), converts early-return guards into nested if blocks, and deletes/re-indents surrounding lines — so the module's structure and the per-branch state transitions are no longer provably identical to the original, and a structural mistake (bad nesting/indentation, methods swallowed into the new block, removed blank-line boundaries) can break the module at import time.
Detection procedure
  1. From the task, identify the exact algorithmic redundancy to remove (rescan of a growing buffer → compare only with the previous element) and note what the minimal edit would touch.
  2. Diff the patch against that minimal edit: count how many new fields and control-flow shape changes it adds; mark any field that is written but never read, and any early-return that became nested logic.
  3. Re-read the patched block as the final file would look: check indentation/nesting of the new code and of the following method definitions, i.e. that nothing after the changed block got absorbed, dedented, or left dangling, and that the module still forms valid, importable code.
  4. Enumerate every input branch (first element of a run, run continuation, reset/skip element, end of input) and confirm the state written and the comparison performed in each branch match the original semantics exactly.
Discriminator
A fine patch may add one extra cached field (e.g. caching the previous element's derived key) as long as every branch's state transitions map 1:1 onto the original and the surrounding method boundaries are untouched; a violation is a rewrite that changes control-flow shape, leaves dead state, or shifts the indentation/structure of neighbouring code, so equivalence and even syntactic validity cannot be checked by reading the diff.
Consequence
If the restructuring broke the module structure, the covering tests do not merely fail on a few edge cases — collection/import fails and essentially all of them fail (0/N), and the timing script cannot even import the target; if only a branch's state transition was altered, the returned ratios silently change on runs that start or end at a reset element.
id 75c08a1b097d · mined from jawah/charset_normalizer jawah__charset_normalizer.perf_3
raw text (what the judge reads)
### Restructuring a hot incremental state machine more than necessary, without a parse/equivalence audit
- **Applies when**: `task` -- the task's slowdown is "re-scanning the accumulated buffer instead of just the previous element", and the patch replaces the accumulate-and-rescan code with new per-element state fields and a reshaped control flow.
- **Pattern**: Instead of the minimal edit (keep one "previous element" variable, compare it with the current one, then update it), the patch rewrites the whole method body: it introduces several parallel state fields (some never read), converts early-`return` guards into nested `if` blocks, and deletes/re-indents surrounding lines — so the module's structure and the per-branch state transitions are no longer provably identical to the original, and a structural mistake (bad nesting/indentation, methods swallowed into the new block, removed blank-line boundaries) can break the module at import time.
- **Detection procedure**:
  1. From the task, identify the exact algorithmic redundancy to remove (rescan of a growing buffer → compare only with the previous element) and note what the minimal edit would touch.
  2. Diff the patch against that minimal edit: count how many new fields and control-flow shape changes it adds; mark any field that is written but never read, and any early-return that became nested logic.
  3. Re-read the patched block as the final file would look: check indentation/nesting of the new code and of the following method definitions, i.e. that nothing after the changed block got absorbed, dedented, or left dangling, and that the module still forms valid, importable code.
  4. Enumerate every input branch (first element of a run, run continuation, reset/skip element, end of input) and confirm the state written and the comparison performed in each branch match the original semantics exactly.
- **Discriminator**: A fine patch may add one extra cached field (e.g. caching the previous element's derived key) as long as every branch's state transitions map 1:1 onto the original and the surrounding method boundaries are untouched; a violation is a rewrite that changes control-flow shape, leaves dead state, or shifts the indentation/structure of neighbouring code, so equivalence and even syntactic validity cannot be checked by reading the diff.
- **Consequence**: If the restructuring broke the module structure, the covering tests do not merely fail on a few edge cases — collection/import fails and essentially all of them fail (0/N), and the timing script cannot even import the target; if only a branch's state transition was altered, the returned ratios silently change on runs that start or end at a reset element.
21Loop-to-builtin rewrite that drops the loop's special-cased first/degenerate iterationtasksunpy/sunpy
Applies when
task -- a patch replaces an accumulate-in-a-loop string/sequence builder (the source of the quadratic cost) with a single builtin bulk call such as join/extend/comprehension.
Pattern
The rewrite reproduces the loop's behaviour only for the "typical" input size used by the timing script (many homogeneous items) while silently discarding branches the loop took on degenerate inputs — zero items, exactly one item, or items whose type the bulk builtin treats differently (e.g. the original passed a lone element through untouched, so non-str/lazily-coerced values worked, whereas the bulk call now validates or coerces every element and raises or formats differently).
Detection procedure
  1. Read the workload and note the only shape it exercises (here: a long list of plain short strings); everything else in the input space is untested by the timing script.
  2. Read the original loop and enumerate its distinct control-flow paths, especially any if index:/first-iteration branch, sentinel initial value, or early-exit that makes the 0-item and 1-item cases not go through the accumulating call.
  3. Hand-simulate both the original and the patched expression for n=0, n=1, n=2 and for at least one non-string / unusual element type, comparing the exact returned value and the exception behaviour.
  4. Confirm the covering tests actually exercise those degenerate shapes; if they do not, treat the missing path as a correctness risk rather than as "no test, no problem".
Discriminator
A genuine violation is when at least one enumerated edge case yields a different value, a different type, or a different exception than the original. A look-alike that is fine is a rewrite that is provably total-behaviour-identical across all enumerated paths (the bulk call already handled the empty and singleton cases with the same result and the same element-type requirements) — such a rewrite is acceptable even though it looks like it deleted a branch.
Consequence
The workload gets faster (linear instead of quadratic), so the speedup masks the problem, but a covering test that hits the empty/singleton/non-string path fails with a TypeError or an off-by-one separator/formatting mismatch — a partial pass such as 52/53 that would be wrongly dismissed as flakiness.
id ae78ab04141b · mined from sunpy/sunpy sunpy__sunpy.perf_0
raw text (what the judge reads)
### Loop-to-builtin rewrite that drops the loop's special-cased first/degenerate iteration
- **Applies when**: `task` -- a patch replaces an accumulate-in-a-loop string/sequence builder (the source of the quadratic cost) with a single builtin bulk call such as `join`/`extend`/comprehension.
- **Pattern**: The rewrite reproduces the loop's behaviour only for the "typical" input size used by the timing script (many homogeneous items) while silently discarding branches the loop took on degenerate inputs — zero items, exactly one item, or items whose type the bulk builtin treats differently (e.g. the original passed a lone element through untouched, so non-`str`/lazily-coerced values worked, whereas the bulk call now validates or coerces every element and raises or formats differently).
- **Detection procedure**:
  1. Read the workload and note the *only* shape it exercises (here: a long list of plain short strings); everything else in the input space is untested by the timing script.
  2. Read the original loop and enumerate its distinct control-flow paths, especially any `if index:`/first-iteration branch, sentinel initial value, or early-exit that makes the 0-item and 1-item cases *not* go through the accumulating call.
  3. Hand-simulate both the original and the patched expression for n=0, n=1, n=2 and for at least one non-string / unusual element type, comparing the exact returned value and the exception behaviour.
  4. Confirm the covering tests actually exercise those degenerate shapes; if they do not, treat the missing path as a correctness risk rather than as "no test, no problem".
- **Discriminator**: A genuine violation is when at least one enumerated edge case yields a different value, a different type, or a different exception than the original. A look-alike that is fine is a rewrite that is provably total-behaviour-identical across all enumerated paths (the bulk call already handled the empty and singleton cases with the same result and the same element-type requirements) — such a rewrite is acceptable even though it looks like it deleted a branch.
- **Consequence**: The workload gets faster (linear instead of quadratic), so the speedup masks the problem, but a covering test that hits the empty/singleton/non-string path fails with a `TypeError` or an off-by-one separator/formatting mismatch — a partial pass such as 52/53 that would be wrongly dismissed as flakiness.
22Rebinding a live loop variable while hoisting redundant worktaskdatamade/usaddress
Applies when
task -- the patch removes a repeated computation inside a loop by replacing it with an inline comprehension, temporary assignment, or slice, in a scope where the loop's own iteration variables are still in use.
Pattern
The optimization is arithmetically/semantically equivalent in isolation, but the new expression binds a name (comprehension target, unpacking target, or temporary) that collides with a variable the rest of the loop body reads — most often the per-iteration item or label being classified. In Python 3 comprehension targets are scoped, but tuple-unpacking, for targets outside comprehensions, and plain assignments are not; conversely a comprehension may still shadow at read-time if the author intended to reuse the outer name. Either way the downstream logic now sees a stale, aggregated, or last-element value instead of the current one.
Detection procedure
  1. From the task, identify what the loop must decide per item (here: a per-item classification plus a running state flag) and note every name the loop body reads after the hoisted computation.
  2. In the patch, list every name newly bound by the inserted expression (comprehension/unpacking/temp) and diff it against that read set.
  3. For any collision, hand-simulate one iteration on a typical workload input and check whether the post-collision reads still get the current item's value, and whether accumulated state (flags, "seen so far" checks) is still computed over the same prefix as before.
  4. Confirm the loop's control-flow decisions (branch conditions, error raises) depend on exactly the same values as the original.
Discriminator
A real violation is a name reused for two different meanings in one scope, or a condition whose operand changed from "current item" to "collection of items" (or vice versa). A benign look-alike introduces a fresh, clearly distinct temporary name, or reuses a name only after its last downstream read, leaving all branch conditions evaluating the same values.
Consequence
The workload may report a real speedup (the redundant recomputation is gone), but the covering tests fail broadly — nearly every case mislabels or mis-branches — because each iteration classifies the wrong value.
id 44391c12eec4 · mined from datamade/usaddress datamade__usaddress.perf_3
raw text (what the judge reads)
### Rebinding a live loop variable while hoisting redundant work
- **Applies when**: `task` -- the patch removes a repeated computation inside a loop by replacing it with an inline comprehension, temporary assignment, or slice, in a scope where the loop's own iteration variables are still in use.
- **Pattern**: The optimization is arithmetically/semantically equivalent in isolation, but the new expression binds a name (comprehension target, unpacking target, or temporary) that collides with a variable the rest of the loop body reads — most often the per-iteration item or label being classified. In Python 3 comprehension targets are scoped, but tuple-unpacking, `for` targets outside comprehensions, and plain assignments are not; conversely a comprehension may still shadow at read-time if the author *intended* to reuse the outer name. Either way the downstream logic now sees a stale, aggregated, or last-element value instead of the current one.
- **Detection procedure**:
  1. From the task, identify what the loop must decide per item (here: a per-item classification plus a running state flag) and note every name the loop body reads after the hoisted computation.
  2. In the patch, list every name newly bound by the inserted expression (comprehension/unpacking/temp) and diff it against that read set.
  3. For any collision, hand-simulate one iteration on a typical workload input and check whether the post-collision reads still get the current item's value, and whether accumulated state (flags, "seen so far" checks) is still computed over the same prefix as before.
  4. Confirm the loop's control-flow decisions (branch conditions, error raises) depend on exactly the same values as the original.
- **Discriminator**: A real violation is a name reused for two different meanings in one scope, or a condition whose operand changed from "current item" to "collection of items" (or vice versa). A benign look-alike introduces a fresh, clearly distinct temporary name, or reuses a name only *after* its last downstream read, leaving all branch conditions evaluating the same values.
- **Consequence**: The workload may report a real speedup (the redundant recomputation is gone), but the covering tests fail broadly — nearly every case mislabels or mis-branches — because each iteration classifies the wrong value.
23Extra hand-rolled reimplementation bundled with the real optimizationtasktheskumar/python-dotenv
Applies when
task -- a patch that fixes a genuine hot-spot (e.g. removing repeated slicing/copying) also rewrites a nearby helper's semantics by hand, replacing a library/regex-based primitive with bespoke arithmetic or string operations.
Pattern
The agent keeps the actual algorithmic fix but adds a second, speculative change that re-derives existing behaviour from scratch ("this is equivalent and probably faster"), without evidence that the helper is a bottleneck and without enumerating the input cases the original primitive handled.
Detection procedure
  1. From the task and workload, identify the single dominant cost (here: super-linear growth ⇒ repeated copying/rescanning per item) and mark which hunks address it.
  2. For every remaining hunk, ask whether it changes what is computed rather than how often it is computed; those hunks are semantics-bearing and are not the fix.
  3. For each such hunk, enumerate the inputs the original primitive handled (alternation branches, overlapping cases, empty/boundary inputs, flags/anchors) and check the replacement reproduces each one exactly; also check nothing the primitive was still relied on for became dead or inconsistent.
  4. Estimate the hunk's share of runtime: if it runs once per item on already-short strings while the real fix removes an O(n²) cost, its upside is negligible and its risk is unbounded.
Discriminator
A real violation is a semantics-changing rewrite that is neither required by nor complementary to the hot-path fix and whose equivalence rests on informal reasoning. A look-alike that is fine is a rewrite that is itself the measured hot spot, or a mechanical change that provably preserves every branch of the original (verified case-by-case, including boundaries) and is needed to make the hot-path fix work.
Consequence
The workload may still speed up, but the covering tests fail — potentially wholesale, since the rewritten helper is exercised on every parsed item — so the patch is rejected despite containing the correct optimization.
id de59a1c59c34 · mined from theskumar/python-dotenv theskumar__python-dotenv.perf_0
raw text (what the judge reads)
### Extra hand-rolled reimplementation bundled with the real optimization
- **Applies when**: `task` -- a patch that fixes a genuine hot-spot (e.g. removing repeated slicing/copying) *also* rewrites a nearby helper's semantics by hand, replacing a library/regex-based primitive with bespoke arithmetic or string operations.
- **Pattern**: The agent keeps the actual algorithmic fix but adds a second, speculative change that re-derives existing behaviour from scratch ("this is equivalent and probably faster"), without evidence that the helper is a bottleneck and without enumerating the input cases the original primitive handled.
- **Detection procedure**:
  1. From the task and workload, identify the single dominant cost (here: super-linear growth ⇒ repeated copying/rescanning per item) and mark which hunks address it.
  2. For every remaining hunk, ask whether it changes *what is computed* rather than *how often it is computed*; those hunks are semantics-bearing and are not the fix.
  3. For each such hunk, enumerate the inputs the original primitive handled (alternation branches, overlapping cases, empty/boundary inputs, flags/anchors) and check the replacement reproduces each one exactly; also check nothing the primitive was still relied on for became dead or inconsistent.
  4. Estimate the hunk's share of runtime: if it runs once per item on already-short strings while the real fix removes an O(n²) cost, its upside is negligible and its risk is unbounded.
- **Discriminator**: A real violation is a semantics-changing rewrite that is neither required by nor complementary to the hot-path fix and whose equivalence rests on informal reasoning. A look-alike that is fine is a rewrite that is itself the measured hot spot, or a mechanical change that provably preserves every branch of the original (verified case-by-case, including boundaries) and is needed to make the hot-path fix work.
- **Consequence**: The workload may still speed up, but the covering tests fail — potentially wholesale, since the rewritten helper is exercised on every parsed item — so the patch is rejected despite containing the correct optimization.
24Undefined helper: patch uses a precomputed constant/regex it never actually addstaskpndurette/gTTS
Applies when
task -- a patch replaces an inline per-character/per-item computation with a call to a module-level constant, compiled regex, lookup table, or helper function.
Pattern
The diff swaps the slow inline logic for a reference to a new name (e.g. a precompiled matcher or set), but no hunk in the patch defines, initializes, or imports that name — and it does not already exist in the file. The code looks correct in isolation, yet the module raises NameError (or ImportError) on the very first call, or worse, silently binds to an unrelated existing symbol.
Detection procedure
  1. List every identifier the patched code reads that is not a parameter or local defined within the same hunk.
  2. For each such identifier, check whether the diff contains a hunk that defines it (assignment, def, import) or whether the surrounding unchanged file already defines it — do not assume it exists because the name "looks conventional".
  3. If the name is meant to be module-level state, confirm the definition is placed at import time and before first use, and that its value truly matches the semantics it replaces (e.g. the regex is anchored/full-match where the old code required all characters to qualify).
  4. Mentally execute one workload input through the new path with the actual definition (or its absence).
Discriminator
A real violation is when the referenced symbol has no definition anywhere in the patched module, or is defined with semantics that differ from the replaced logic (partial match vs. full match, empty-string handling). A look-alike that is fine is a patch that references a symbol already present in the untouched part of the file or clearly added in another hunk of the same diff.
Consequence
Every test touching the module fails immediately (0 passed) and the timing script aborts with NameError before printing any Mean/Std Dev — an apparent speedup that never executes.
id 575504e7e7af · mined from pndurette/gTTS pndurette__gTTS.perf_2
raw text (what the judge reads)
### Undefined helper: patch uses a precomputed constant/regex it never actually adds
- **Applies when**: `task` -- a patch replaces an inline per-character/per-item computation with a call to a module-level constant, compiled regex, lookup table, or helper function.
- **Pattern**: The diff swaps the slow inline logic for a reference to a new name (e.g. a precompiled matcher or set), but no hunk in the patch defines, initializes, or imports that name — and it does not already exist in the file. The code looks correct in isolation, yet the module raises `NameError` (or `ImportError`) on the very first call, or worse, silently binds to an unrelated existing symbol.
- **Detection procedure**:
  1. List every identifier the patched code reads that is not a parameter or local defined within the same hunk.
  2. For each such identifier, check whether the diff contains a hunk that defines it (assignment, `def`, `import`) or whether the surrounding unchanged file already defines it — do not assume it exists because the name "looks conventional".
  3. If the name is meant to be module-level state, confirm the definition is placed at import time and before first use, and that its value truly matches the semantics it replaces (e.g. the regex is anchored/full-match where the old code required *all* characters to qualify).
  4. Mentally execute one workload input through the new path with the actual definition (or its absence).
- **Discriminator**: A real violation is when the referenced symbol has no definition anywhere in the patched module, or is defined with semantics that differ from the replaced logic (partial match vs. full match, empty-string handling). A look-alike that is fine is a patch that references a symbol already present in the untouched part of the file or clearly added in another hunk of the same diff.
- **Consequence**: Every test touching the module fails immediately (0 passed) and the timing script aborts with `NameError` before printing any Mean/Std Dev — an apparent speedup that never executes.
25Removing a defensive copy of shared/class-level state to save timetaskmarshmallow-code/marshmallow
Applies when
task -- a constructor/initialization hot path is sped up by deleting or weakening a copy (deepcopy/copy/dict(...)/list(...)) of data that originates from class attributes, module globals, defaults, or caller-supplied containers.
Pattern
The patch assumes the copied structure is flat and immutable ("values are always strings", "nobody mutates this"), so it substitutes shallow reuse or direct aliasing. The assumption holds for the built-in cases the workload constructs but not for every subclass, user-supplied override, or later in-place mutation, so per-instance state silently becomes shared class-level state.
Detection procedure
  1. In the patch, locate every copy operation that was removed or downgraded and identify the source of the data (class attribute, MRO walk, default argument, caller argument).
  2. Ask whether any code path — including subclasses, plugins, or user input outside the workload — can store non-scalar values in that structure or mutate the resulting object in place after construction.
  3. Grep the repository for writes to the attribute the copy produced (obj.X[...] = , .update(, .pop(, .append() outside the constructor; any such write proves per-instance ownership was required.
  4. Confirm the module still imports and the entire listed covering test set is exercised, not just the few tests matching the workload's field types — an all-tests-fail count signals an import/collection-level breakage, not a subtle behavioural drift.
Discriminator
Fine if the copied structure is provably rebuilt fresh per call and all its values are immutable scalars for every subclass and for arbitrary user overrides, and nothing mutates it in place afterwards; a violation if any value can be a container/mutable object or any code mutates the attribute after construction, which makes distinct instances alias one another.
Consequence
The workload gets faster, but cross-instance leakage or an import/definition-time error surfaces in the covering tests — potentially as a wholesale failure (0 of N passing) rather than a handful of assertion mismatches.
id f992dbd00a28 · mined from marshmallow-code/marshmallow marshmallow-code__marshmallow.perf_3
raw text (what the judge reads)
### Removing a defensive copy of shared/class-level state to save time
- **Applies when**: `task` -- a constructor/initialization hot path is sped up by deleting or weakening a copy (`deepcopy`/`copy`/`dict(...)`/`list(...)`) of data that originates from class attributes, module globals, defaults, or caller-supplied containers.
- **Pattern**: The patch assumes the copied structure is flat and immutable ("values are always strings", "nobody mutates this"), so it substitutes shallow reuse or direct aliasing. The assumption holds for the built-in cases the workload constructs but not for every subclass, user-supplied override, or later in-place mutation, so per-instance state silently becomes shared class-level state.
- **Detection procedure**:
  1. In the patch, locate every copy operation that was removed or downgraded and identify the *source* of the data (class attribute, MRO walk, default argument, caller argument).
  2. Ask whether any code path — including subclasses, plugins, or user input outside the workload — can store non-scalar values in that structure or mutate the resulting object in place after construction.
  3. Grep the repository for writes to the attribute the copy produced (`obj.X[...] = `, `.update(`, `.pop(`, `.append(`) outside the constructor; any such write proves per-instance ownership was required.
  4. Confirm the module still imports and the *entire* listed covering test set is exercised, not just the few tests matching the workload's field types — an all-tests-fail count signals an import/collection-level breakage, not a subtle behavioural drift.
- **Discriminator**: Fine if the copied structure is provably rebuilt fresh per call *and* all its values are immutable scalars for every subclass and for arbitrary user overrides, and nothing mutates it in place afterwards; a violation if any value can be a container/mutable object or any code mutates the attribute after construction, which makes distinct instances alias one another.
- **Consequence**: The workload gets faster, but cross-instance leakage or an import/definition-time error surfaces in the covering tests — potentially as a wholesale failure (0 of N passing) rather than a handful of assertion mismatches.
26Replacing a single-pass hand-rolled scanner with chained bulk built-ins, without proving pass-equivalence (or satisfying the repo's own style/consistency gates)taskCog-Creators/Red-DiscordBot
Applies when
task -- the hot spot is a Python-level character/index loop that recognizes several literal patterns in one left-to-right pass, and the patch collapses it into two or more built-in bulk operations (e.g. chained replace/translate/regex passes) applied one after another to the whole string.
Pattern
The patch assumes "N sequential whole-string passes" == "one prioritized single pass", and writes them as one dense chained expression, changing both the semantics contract (order/overlap/re-matching between passes) and the source shape the repository's checks expect, instead of a minimal, per-pattern, individually verifiable rewrite.
Detection procedure
  1. From the task, list every pattern the original loop matches and note the priority/advance rules (which pattern wins when two can start at the same index, how far the index jumps after a match, whether inserted output can be rescanned).
  2. For each ordered pair of passes in the patch, ask: can the output of an earlier pass contain a match for a later pass that the original single pass would never have produced? Can a later pass split or re-match text the earlier pass already consumed? Can a match that the original skipped over now be found?
  3. Check whether the covering test list includes repo-wide hygiene/consistency tests (formatting, lint, doc/source checks) that a reformatted or over-long chained line can fail, and whether the patch keeps the surrounding code in the style the repo enforces.
  4. Confirm the patch actually removes Python-level per-character work (real speedup) rather than just re-spelling it — and that it does so with the smallest, most auditable edit.
Discriminator
A fine look-alike keeps each bulk pass on its own statement/line in repo style and can be shown pass-order-independent (the replacement text of every pass provably contains no trigger for any other pass, and patterns cannot overlap); a violation either leaves an unproven ordering/overlap interaction, or achieves the same semantics through a code shape that the repository's own committed checks reject.
Consequence
The workload may report the expected big speedup while the covering tests fail wholesale — either on inputs with adjacent/overlapping or nested patterns whose escaped output now differs, or at collection/hygiene-check time, so the "faster" version is unusable.
id 88cda1829aa8 · mined from Cog-Creators/Red-DiscordBot Cog-Creators__Red-DiscordBot.perf_3
raw text (what the judge reads)
### Replacing a single-pass hand-rolled scanner with chained bulk built-ins, without proving pass-equivalence (or satisfying the repo's own style/consistency gates)

- **Applies when**: `task` -- the hot spot is a Python-level character/index loop that recognizes several literal patterns in one left-to-right pass, and the patch collapses it into two or more built-in bulk operations (e.g. chained `replace`/`translate`/regex passes) applied one after another to the whole string.
- **Pattern**: The patch assumes "N sequential whole-string passes" == "one prioritized single pass", and writes them as one dense chained expression, changing both the semantics contract (order/overlap/re-matching between passes) and the source shape the repository's checks expect, instead of a minimal, per-pattern, individually verifiable rewrite.
- **Detection procedure**:
  1. From the task, list every pattern the original loop matches and note the priority/advance rules (which pattern wins when two can start at the same index, how far the index jumps after a match, whether inserted output can be rescanned).
  2. For each ordered pair of passes in the patch, ask: can the *output* of an earlier pass contain a match for a later pass that the original single pass would never have produced? Can a later pass split or re-match text the earlier pass already consumed? Can a match that the original skipped over now be found?
  3. Check whether the covering test list includes repo-wide hygiene/consistency tests (formatting, lint, doc/source checks) that a reformatted or over-long chained line can fail, and whether the patch keeps the surrounding code in the style the repo enforces.
  4. Confirm the patch actually removes Python-level per-character work (real speedup) rather than just re-spelling it — and that it does so with the smallest, most auditable edit.
- **Discriminator**: A fine look-alike keeps each bulk pass on its own statement/line in repo style and can be shown pass-order-independent (the replacement text of every pass provably contains no trigger for any other pass, and patterns cannot overlap); a violation either leaves an unproven ordering/overlap interaction, or achieves the same semantics through a code shape that the repository's own committed checks reject.
- **Consequence**: The workload may report the expected big speedup while the covering tests fail wholesale — either on inputs with adjacent/overlapping or nested patterns whose escaped output now differs, or at collection/hygiene-check time, so the "faster" version is unusable.
27Hand-rolled fast path that swallows its own bugs as "invalid input"taskmarshmallow-code/marshmallow
Applies when
task -- the patch replaces a library/stdlib parsing or validation call with a hand-written fast path (regex + manual construction, manual arithmetic, table lookup) and wraps the new code in an except that re-raises the function's normal "bad value" error.
Pattern
The rewrite depends on assumptions that are never verified (a pre-existing pattern/constant exists, its capture groups exactly match the constructor's keyword names, no group is optional/None, ranges are already validated). Any mismatch raises TypeError/ValueError inside the new code, and the broad except converts that implementation bug into the ordinary rejection error — so every well-formed input is silently reported as malformed instead of failing loudly.
Detection procedure
  1. From the task, note the function's contract: which inputs must succeed, and which single error type signals rejection.
  2. In the patch, list every symbol the new fast path relies on (patterns, group names, constants, helper maps) and confirm each is defined in the diff or verifiably present with exactly the assumed shape — e.g. capture-group names identical to the constructor's parameters, and none of them optional.
  3. Check the try/except scope: does it wrap only the input-dependent rejection, or also the construction/conversion step where a wrong assumption would surface? Widened clauses (except (ValueError, TypeError), except Exception) around the new logic are the red flag.
  4. Mentally run one canonical valid input from the workload through the new path, substituting the real (not assumed) helper definition, and see whether it reaches the success return.
Discriminator
A real violation is a fast path whose failure mode is indistinguishable from legitimate rejection, so a wrong assumption yields uniformly rejected valid input; a look-alike that is fine keeps the exception handler narrowly around genuinely input-driven errors (or has no handler at all around the construction step), letting any internal mistake propagate as a distinct error, and defines/validates every symbol it depends on.
Consequence
The workload may still run (and even look faster, since it now short-circuits), but the covering tests fail wholesale — valid values are rejected with the "not a valid value" error, so essentially every test touching this parsing path breaks rather than a single edge case.
id e63d8a80fd63 · mined from marshmallow-code/marshmallow marshmallow-code__marshmallow.perf_2
raw text (what the judge reads)
### Hand-rolled fast path that swallows its own bugs as "invalid input"
- **Applies when**: `task` -- the patch replaces a library/stdlib parsing or validation call with a hand-written fast path (regex + manual construction, manual arithmetic, table lookup) and wraps the new code in an `except` that re-raises the function's normal "bad value" error.
- **Pattern**: The rewrite depends on assumptions that are never verified (a pre-existing pattern/constant exists, its capture groups exactly match the constructor's keyword names, no group is optional/None, ranges are already validated). Any mismatch raises `TypeError`/`ValueError` inside the new code, and the broad `except` converts that implementation bug into the ordinary rejection error — so *every* well-formed input is silently reported as malformed instead of failing loudly.
- **Detection procedure**:
  1. From the task, note the function's contract: which inputs must succeed, and which single error type signals rejection.
  2. In the patch, list every symbol the new fast path relies on (patterns, group names, constants, helper maps) and confirm each is defined in the diff or verifiably present with exactly the assumed shape — e.g. capture-group names identical to the constructor's parameters, and none of them optional.
  3. Check the `try`/`except` scope: does it wrap only the input-dependent rejection, or also the construction/conversion step where a wrong assumption would surface? Widened clauses (`except (ValueError, TypeError)`, `except Exception`) around the new logic are the red flag.
  4. Mentally run one canonical *valid* input from the workload through the new path, substituting the real (not assumed) helper definition, and see whether it reaches the success return.
- **Discriminator**: A real violation is a fast path whose failure mode is indistinguishable from legitimate rejection, so a wrong assumption yields uniformly rejected valid input; a look-alike that is fine keeps the exception handler narrowly around genuinely input-driven errors (or has no handler at all around the construction step), letting any internal mistake propagate as a distinct error, and defines/validates every symbol it depends on.
- **Consequence**: The workload may still run (and even look faster, since it now short-circuits), but the covering tests fail wholesale — valid values are rejected with the "not a valid value" error, so essentially every test touching this parsing path breaks rather than a single edge case.
28Dropping the implicit type normalization performed by an element-wise accumulation looptaskrsalmei/alive-progress
Applies when
task -- the patch replaces a loop that appends/slices items one at a time into a fresh container with a single bulk concatenation/sum/+= over whole inputs, in order to remove per-element overhead.
Pattern
The original loop silently normalized heterogeneous inputs (strings, lists, generators, other sequences) into the accumulator's type because it only ever touched one element/slice at a time; the "equivalent" bulk version concatenates the raw inputs directly, so it only works when every caller already passes exactly the accumulator's type, and raises or returns a differently-typed result otherwise.
Detection procedure
  1. Read the removed loop and ask what type each element/slice contributed versus the type of the argument as a whole — if they differ (e.g., slicing a string yields items that concatenate into the result container), the loop was doing a conversion, not just iteration.
  2. Check the new expression's type requirements: bulk +/sum/extend demands all operands be the same type as the seed, with no per-item coercion.
  3. Grep every call site (and doc/tests) of the changed helper for arguments that are not already that type — strings, lists, tuples-of-mixed, iterators/generators (which the loop may consume once and the bulk version cannot re-index).
  4. If any caller, including ones outside the timing script's path, supplies another type, the patch is a behaviour change, not an optimization.
Discriminator
A real violation exists when at least one reachable caller passes an input whose whole-object type differs from the accumulator's (or is a one-shot iterable); it is a look-alike-but-fine change if the function is documented/typed to accept only that exact container type and all call sites provably pass it, so the loop was pure overhead.
Consequence
The workload may look fast (or crash immediately), but the repository's tests covering the helper fail en masse with TypeError: can only concatenate ... or with results of the wrong type/shape, i.e. speed bought by breaking supported input forms.
id 23389df31bd1 · mined from rsalmei/alive-progress rsalmei__alive-progress.perf_2
raw text (what the judge reads)
### Dropping the implicit type normalization performed by an element-wise accumulation loop
- **Applies when**: `task` -- the patch replaces a loop that appends/slices items one at a time into a fresh container with a single bulk concatenation/`sum`/`+=` over whole inputs, in order to remove per-element overhead.
- **Pattern**: The original loop silently normalized heterogeneous inputs (strings, lists, generators, other sequences) into the accumulator's type because it only ever touched one element/slice at a time; the "equivalent" bulk version concatenates the raw inputs directly, so it only works when every caller already passes exactly the accumulator's type, and raises or returns a differently-typed result otherwise.
- **Detection procedure**:
  1. Read the removed loop and ask what type each element/slice contributed versus the type of the argument as a whole — if they differ (e.g., slicing a string yields items that concatenate into the result container), the loop was doing a conversion, not just iteration.
  2. Check the new expression's type requirements: bulk `+`/`sum`/`extend` demands all operands be the same type as the seed, with no per-item coercion.
  3. Grep every call site (and doc/tests) of the changed helper for arguments that are not already that type — strings, lists, tuples-of-mixed, iterators/generators (which the loop may consume once and the bulk version cannot re-index).
  4. If any caller, including ones outside the timing script's path, supplies another type, the patch is a behaviour change, not an optimization.
- **Discriminator**: A real violation exists when at least one reachable caller passes an input whose whole-object type differs from the accumulator's (or is a one-shot iterable); it is a look-alike-but-fine change if the function is documented/typed to accept only that exact container type and all call sites provably pass it, so the loop was pure overhead.
- **Consequence**: The workload may look fast (or crash immediately), but the repository's tests covering the helper fail en masse with `TypeError: can only concatenate ...` or with results of the wrong type/shape, i.e. speed bought by breaking supported input forms.
29Reusing an assumed-existing helper/pattern to replace hand-written validation, without verifying it exists and is semantically identicaltaskpydicom/pydicom
Applies when
task -- a patch replaces an explicit, multi-step check loop (character-set scan, encoding check, length checks) with a single call to a named module-level constant, compiled pattern, or helper that the diff itself does not define or import.
Pattern
The optimization assumes a symbol from elsewhere in the codebase (or a stdlib module) is already in scope and already encodes exactly the same rules; it also silently deletes one or more of the original validation steps (e.g. an encoding/ASCII guard, an anchoring requirement, an empty-input case) that the replacement may not cover.
Detection procedure
  1. List every check the original code performs and the exact accept/reject set it implies, including degenerate inputs (empty string, empty sequence, non-ASCII, over-length).
  2. For each new symbol the patch references, confirm the diff itself defines/imports it, or point to the exact existing definition; check its literal text (anchors, character class, flags) and whether it is defined before the point of use.
  3. Map each original check onto the replacement; any check with no counterpart is a behaviour change unless you can argue the replacement subsumes it for all inputs, not just the workload's inputs.
  4. Check whether applying the replacement per-element instead of to the joined/whole input changes semantics for boundary cases.
Discriminator
A real violation is a patch whose new symbol is unresolved in that scope, or whose semantics provably differ (unanchored vs anchored, missing ASCII/length guard, different empty-input result). A look-alike that is fine defines or imports the symbol in the same diff and demonstrates the pattern accepts/rejects exactly the same inputs as the removed code, including edge cases.
Consequence
The covering tests fail wholesale (import/NameError at module load, or systematic accept/reject flips), so the timing script's speedup is meaningless — correctness is broken before performance is even measured.
id 8c1ea7dde461 · mined from pydicom/pydicom pydicom__pydicom.perf_1
raw text (what the judge reads)
### Reusing an assumed-existing helper/pattern to replace hand-written validation, without verifying it exists and is semantically identical
- **Applies when**: `task` -- a patch replaces an explicit, multi-step check loop (character-set scan, encoding check, length checks) with a single call to a named module-level constant, compiled pattern, or helper that the diff itself does not define or import.
- **Pattern**: The optimization assumes a symbol from elsewhere in the codebase (or a stdlib module) is already in scope and already encodes exactly the same rules; it also silently deletes one or more of the original validation steps (e.g. an encoding/ASCII guard, an anchoring requirement, an empty-input case) that the replacement may not cover.
- **Detection procedure**:
  1. List every check the original code performs and the exact accept/reject set it implies, including degenerate inputs (empty string, empty sequence, non-ASCII, over-length).
  2. For each new symbol the patch references, confirm the diff itself defines/imports it, or point to the exact existing definition; check its literal text (anchors, character class, flags) and whether it is defined before the point of use.
  3. Map each original check onto the replacement; any check with no counterpart is a behaviour change unless you can argue the replacement subsumes it for all inputs, not just the workload's inputs.
  4. Check whether applying the replacement per-element instead of to the joined/whole input changes semantics for boundary cases.
- **Discriminator**: A real violation is a patch whose new symbol is unresolved in that scope, or whose semantics provably differ (unanchored vs anchored, missing ASCII/length guard, different empty-input result). A look-alike that is fine defines or imports the symbol in the same diff and demonstrates the pattern accepts/rejects exactly the same inputs as the removed code, including edge cases.
- **Consequence**: The covering tests fail wholesale (import/NameError at module load, or systematic accept/reject flips), so the timing script's speedup is meaningless — correctness is broken before performance is even measured.
30Replacing a recomputed aggregate with an incremental counter without proving path-for-path equivalencetaskPyCQA/flake8
Applies when
task -- the hot spot is a quadratic "recompute the whole aggregate each iteration" pattern (e.g. len("".join(acc)), sum(acc), max(acc) inside a loop) and the patch introduces a running variable that is updated per item instead.
Pattern
The patch adds one accumulator update next to the one place it noticed, but does not verify that the accumulator's value is, at every point where the aggregate is read, identical to recomputing it from the live collection — because some paths append/extend/reset/mutate the collection, rewrite the item after measuring it, continue/break around the update, or the collection is also touched by other methods, initialization is misplaced relative to a reset, or the edit is anchored to stale context so it does not integrate with the surrounding code at all.
Detection procedure
  1. From the workload, identify the loop and the exact aggregate expression being eliminated, and note which value(s) callers/tests actually consume (here: the recorded offsets, not just the final total).
  2. Enumerate every statement in the enclosing scope (and any other method/branch) that creates, resets, appends to, extends, reorders, or edits elements of the accumulated collection, plus every early-exit path inside the loop.
  3. For each such site, check the patch keeps the counter in lockstep: initialized where the collection is created/cleared, updated with the final item value after any rewriting, and not skipped by continue/exception paths.
  4. Confirm the patch is a well-formed edit of the current code (context lines, indentation, variable actually defined before first use) so the module still imports — a mass failure of all covering tests, not a few, is the signature of an edit that does not integrate.
Discriminator
Fine if the collection is strictly append-only within one scope, the appended value is never modified afterwards, the counter is initialized exactly where the collection is, and every loop exit path either updates both or neither. A violation exists when at least one reachable path changes the collection (or the item) without a matching counter update, or when the counter is introduced in a way that breaks compilation/import of the module.
Consequence
The workload may well show the expected speedup (the loop is genuinely cheaper), while the covering tests fail — either with off-by-N/incorrect derived offsets and wrong diagnostics, or with a wholesale collection/import error that fails every test in the suite.
id 5ac882dc092b · mined from PyCQA/flake8 PyCQA__flake8.perf_3
raw text (what the judge reads)
### Replacing a recomputed aggregate with an incremental counter without proving path-for-path equivalence
- **Applies when**: `task` -- the hot spot is a quadratic "recompute the whole aggregate each iteration" pattern (e.g. `len("".join(acc))`, `sum(acc)`, `max(acc)` inside a loop) and the patch introduces a running variable that is updated per item instead.
- **Pattern**: The patch adds one accumulator update next to the one place it noticed, but does not verify that the accumulator's value is, at every point where the aggregate is read, identical to recomputing it from the live collection — because some paths append/extend/reset/mutate the collection, rewrite the item after measuring it, `continue`/`break` around the update, or the collection is also touched by other methods, initialization is misplaced relative to a reset, or the edit is anchored to stale context so it does not integrate with the surrounding code at all.
- **Detection procedure**:
  1. From the workload, identify the loop and the exact aggregate expression being eliminated, and note which value(s) callers/tests actually consume (here: the recorded offsets, not just the final total).
  2. Enumerate *every* statement in the enclosing scope (and any other method/branch) that creates, resets, appends to, extends, reorders, or edits elements of the accumulated collection, plus every early-exit path inside the loop.
  3. For each such site, check the patch keeps the counter in lockstep: initialized where the collection is created/cleared, updated with the *final* item value after any rewriting, and not skipped by `continue`/exception paths.
  4. Confirm the patch is a well-formed edit of the current code (context lines, indentation, variable actually defined before first use) so the module still imports — a mass failure of all covering tests, not a few, is the signature of an edit that does not integrate.
- **Discriminator**: Fine if the collection is strictly append-only within one scope, the appended value is never modified afterwards, the counter is initialized exactly where the collection is, and every loop exit path either updates both or neither. A violation exists when at least one reachable path changes the collection (or the item) without a matching counter update, or when the counter is introduced in a way that breaks compilation/import of the module.
- **Consequence**: The workload may well show the expected speedup (the loop is genuinely cheaper), while the covering tests fail — either with off-by-N/incorrect derived offsets and wrong diagnostics, or with a wholesale collection/import error that fails every test in the suite.
31Caching decorator added without verifying its prerequisites (import present, instance can hold the cache)taskPyCQA/flake8
Applies when
task -- the patch speeds up repeated lookups by swapping a recomputed property/method for a memoizing decorator (functools.cached_property, lru_cache, a custom cache helper) on a class.
Pattern
The diff changes only the decorator line and assumes the supporting machinery already exists — i.e. it never checks that the decorating module/helper is imported in that file, that the class permits per-instance attribute assignment (no __slots__, no frozen dataclass, no custom __setattr__), and that the cached name doesn't collide with an existing attribute. A missing import makes the name unresolvable at class-definition time, so the whole module fails to import.
Detection procedure
  1. In the workload, confirm the target really is the hot repeated lookup (here: many calls per constructed object) so the caching idea is on the hot path.
  2. Read the unchanged context of the patched file in the diff: does the required module/helper name (e.g. functools) actually appear in the file's imports, and does the patch add it if not?
  3. Check the class definition for anything that blocks setting a new instance attribute (__slots__, frozen/immutable dataclass, overridden __setattr__) and for an existing attribute/method of the same name.
  4. Check the cache lifetime/key: per-instance cache is fine only if the underlying data is immutable for the object's lifetime; a module- or class-level cache keyed on the object or on nothing at all is a red flag.
Discriminator
A real violation is a one-line decorator swap whose supporting name is not imported/valid in that file, or whose class cannot store the cached value — this is a hard failure, not a slowdown. A look-alike that is fine is the same swap where the import already exists in the visible file header (or is added by the diff) and the class is an ordinary mutable object whose backing state is set once in __init__.
Consequence
The module raises NameError/AttributeError at import or first access, so essentially every covering test errors out at collection (0 passed) and the timing script cannot even construct the object — a total correctness break rather than the expected speedup.
id e44c7db8a7d9 · mined from PyCQA/flake8 PyCQA__flake8.perf_0
raw text (what the judge reads)
### Caching decorator added without verifying its prerequisites (import present, instance can hold the cache)
- **Applies when**: `task` -- the patch speeds up repeated lookups by swapping a recomputed `property`/method for a memoizing decorator (`functools.cached_property`, `lru_cache`, a custom cache helper) on a class.
- **Pattern**: The diff changes only the decorator line and assumes the supporting machinery already exists — i.e. it never checks that the decorating module/helper is imported in that file, that the class permits per-instance attribute assignment (no `__slots__`, no frozen dataclass, no custom `__setattr__`), and that the cached name doesn't collide with an existing attribute. A missing import makes the name unresolvable at class-definition time, so the whole module fails to import.
- **Detection procedure**:
  1. In the workload, confirm the target really is the hot repeated lookup (here: many calls per constructed object) so the caching idea is on the hot path.
  2. Read the *unchanged* context of the patched file in the diff: does the required module/helper name (e.g. `functools`) actually appear in the file's imports, and does the patch add it if not?
  3. Check the class definition for anything that blocks setting a new instance attribute (`__slots__`, frozen/immutable dataclass, overridden `__setattr__`) and for an existing attribute/method of the same name.
  4. Check the cache lifetime/key: per-instance cache is fine only if the underlying data is immutable for the object's lifetime; a module- or class-level cache keyed on the object or on nothing at all is a red flag.
- **Discriminator**: A real violation is a one-line decorator swap whose supporting name is not imported/valid in that file, or whose class cannot store the cached value — this is a hard failure, not a slowdown. A look-alike that is fine is the same swap where the import already exists in the visible file header (or is added by the diff) and the class is an ordinary mutable object whose backing state is set once in `__init__`.
- **Consequence**: The module raises `NameError`/`AttributeError` at import or first access, so essentially every covering test errors out at collection (0 passed) and the timing script cannot even construct the object — a total correctness break rather than the expected speedup.
32Assumed drop-in equivalence of rewritten library calls / container semantics, with no output-equality checktaskdatamade/usaddress
Applies when
task -- the task demands "make it faster without changing behaviour" for a small pure function, and the patch rewrites its internals via supposedly equivalent idioms (pre-compiled regexes, dropped flags/keyword arguments, replacing an explicit scan loop with a in-container membership test, hoisting work to module import time).
Pattern
The patch swaps each construct for a "known-equivalent" faster form, relying on the reviewer's intuition that the semantics are identical, but never demonstrates that the function's full output is byte-identical on the workload's inputs and on edge inputs; details that carry semantics (regex flags/positional-arg order, whether the iterated object yields keys vs. values vs. normalized forms, empty/short strings, the exact truthiness of the value stored in a feature/result field) are silently altered or presumed harmless.
Detection procedure
  1. Read the workload to identify the single function whose outputs are consumed, and note that its return value feeds a downstream model/consumer, so every field's type and value, not just its truthiness, must be preserved.
  2. For each edit in the patch, name the semantic detail it depends on (compile-time flags, membership vs. iteration over a mapping/list and what that mapping's elements actually are, argument positions, side-effect ordering of new module-level statements) and check the repository definition of the constants/objects involved rather than assuming.
  3. Require the patch to include or reference an equivalence harness: run the old and new implementations on the exact workload tokens plus edge cases (empty string, punctuation-only, single character, non-ASCII) and assert full-output equality; absence of any such evidence is the flag.
  4. Confirm the rewrite also actually removes the dominant cost (e.g., a linear scan over a large container) rather than only shaving constant-factor regex overhead; if it does both, still demand step 3.
Discriminator
A genuine violation is a rewrite whose correctness rests on an unverified assumption about an external object's contents or an API's default behaviour, with no equality evidence. A look-alike that is fine is a rewrite where the equivalence is locally provable from the patch text alone (e.g., hoisting an identical pattern string with identical flags into a module constant, or a membership test against a container literally defined a few lines away as a plain set of the same normalized strings) and where field types/values are visibly unchanged.
Consequence
The workload may well get faster, but the covering tests fail wholesale — a subtly different feature/result value shifts every downstream prediction, so the failure looks catastrophic (near-0% pass) rather than like a narrow edge-case bug.
id 744d7c395f4a · mined from datamade/usaddress datamade__usaddress.perf_2
raw text (what the judge reads)
### Assumed drop-in equivalence of rewritten library calls / container semantics, with no output-equality check
- **Applies when**: `task` -- the task demands "make it faster without changing behaviour" for a small pure function, and the patch rewrites its internals via supposedly equivalent idioms (pre-compiled regexes, dropped flags/keyword arguments, replacing an explicit scan loop with a `in`-container membership test, hoisting work to module import time).
- **Pattern**: The patch swaps each construct for a "known-equivalent" faster form, relying on the reviewer's intuition that the semantics are identical, but never demonstrates that the function's full output is byte-identical on the workload's inputs and on edge inputs; details that carry semantics (regex flags/positional-arg order, whether the iterated object yields keys vs. values vs. normalized forms, empty/short strings, the exact truthiness of the value stored in a feature/result field) are silently altered or presumed harmless.
- **Detection procedure**:
  1. Read the workload to identify the single function whose outputs are consumed, and note that its return value feeds a downstream model/consumer, so every field's *type and value*, not just its truthiness, must be preserved.
  2. For each edit in the patch, name the semantic detail it depends on (compile-time flags, membership vs. iteration over a mapping/list and what that mapping's elements actually are, argument positions, side-effect ordering of new module-level statements) and check the repository definition of the constants/objects involved rather than assuming.
  3. Require the patch to include or reference an equivalence harness: run the old and new implementations on the exact workload tokens plus edge cases (empty string, punctuation-only, single character, non-ASCII) and assert full-output equality; absence of any such evidence is the flag.
  4. Confirm the rewrite also actually removes the dominant cost (e.g., a linear scan over a large container) rather than only shaving constant-factor regex overhead; if it does both, still demand step 3.
- **Discriminator**: A genuine violation is a rewrite whose correctness rests on an unverified assumption about an external object's contents or an API's default behaviour, with no equality evidence. A look-alike that is fine is a rewrite where the equivalence is locally provable from the patch text alone (e.g., hoisting an identical pattern string *with identical flags* into a module constant, or a membership test against a container literally defined a few lines away as a plain set of the same normalized strings) and where field types/values are visibly unchanged.
- **Consequence**: The workload may well get faster, but the covering tests fail wholesale — a subtly different feature/result value shifts every downstream prediction, so the failure looks catastrophic (near-0% pass) rather than like a narrow edge-case bug.
33Substituting a "faster equivalent" lookup table for a pattern scan without proving the table encodes the same matching semanticstasksunpy/sunpy
Applies when
task -- a hot loop matches a word/string against a list of patterns (regex, globs, suffix expressions) and the patch replaces that loop with membership tests against a precomputed dict/set/bucketed-by-length structure that already exists in the module.
Pattern
The patch assumes the precomputed structure is a drop-in index of the same pattern list, and swaps in literal in/slice lookups; in reality the structure is built from a different (usually smaller or differently normalized) source collection, or the original patterns are not plain literals (alternations, character classes, optional groups, anchors, case-folding), so some inputs that used to match no longer do, or vice versa.
Detection procedure
  1. From the workload, note which inputs traverse the replaced loop and what the loop's match outcome controls (early return, branch selection) — a flipped outcome here silently changes results, not just speed.
  2. Locate the definitions of both the original pattern list and the substituted lookup structure; confirm the latter is constructed from the former (same variable, same generation code) rather than merely similarly named.
  3. Check that every original pattern is a bare literal suffix/prefix: scan for regex metacharacters, alternation, optional or repeated groups, and check the case/normalization applied on both sides.
  4. Check boundary behaviour of the new indexing: inputs shorter than the bucket length, empty inputs, and patterns whose length differs from the bucket key — verify these produce the same answer as the pattern scan.
Discriminator
A genuine violation is when the two collections can disagree on at least one plausible input (differing membership, non-literal pattern, or different normalization) — enumerate the symmetric difference to confirm. A look-alike that is fine is a table provably generated from the same list at import time, containing only literal strings, with identical case handling; then the rewrite is a pure speedup.
Consequence
The workload may indeed get faster, but repository tests covering the pluralization/matching behaviour fail on the diverging inputs (e.g. one of the covering tests fails while the rest pass), because a class of words is now routed down the wrong branch and produces a different output string.
id c696c58d4031 · mined from sunpy/sunpy sunpy__sunpy.perf_1
raw text (what the judge reads)
### Substituting a "faster equivalent" lookup table for a pattern scan without proving the table encodes the same matching semantics
- **Applies when**: `task` -- a hot loop matches a word/string against a list of patterns (regex, globs, suffix expressions) and the patch replaces that loop with membership tests against a precomputed dict/set/bucketed-by-length structure that already exists in the module.
- **Pattern**: The patch assumes the precomputed structure is a drop-in index of the same pattern list, and swaps in literal `in`/slice lookups; in reality the structure is built from a *different* (usually smaller or differently normalized) source collection, or the original patterns are not plain literals (alternations, character classes, optional groups, anchors, case-folding), so some inputs that used to match no longer do, or vice versa.
- **Detection procedure**:
  1. From the workload, note which inputs traverse the replaced loop and what the loop's match outcome controls (early return, branch selection) — a flipped outcome here silently changes results, not just speed.
  2. Locate the definitions of both the original pattern list and the substituted lookup structure; confirm the latter is *constructed from* the former (same variable, same generation code) rather than merely similarly named.
  3. Check that every original pattern is a bare literal suffix/prefix: scan for regex metacharacters, alternation, optional or repeated groups, and check the case/normalization applied on both sides.
  4. Check boundary behaviour of the new indexing: inputs shorter than the bucket length, empty inputs, and patterns whose length differs from the bucket key — verify these produce the same answer as the pattern scan.
- **Discriminator**: A genuine violation is when the two collections can disagree on at least one plausible input (differing membership, non-literal pattern, or different normalization) — enumerate the symmetric difference to confirm. A look-alike that is fine is a table provably generated from the same list at import time, containing only literal strings, with identical case handling; then the rewrite is a pure speedup.
- **Consequence**: The workload may indeed get faster, but repository tests covering the pluralization/matching behaviour fail on the diverging inputs (e.g. one of the covering tests fails while the rest pass), because a class of words is now routed down the wrong branch and produces a different output string.
34Substituting a hot-path traversal with an assumed helper accessor that isn't actually definedtaskpython-openxml/python-docx
Applies when
task -- a patch speeds up a property/method by replacing explicit lookup/traversal code (XPath, dict walks, manual search) with a shorter call to a helper attribute, generated accessor, or delegated property on a related class.
Pattern
The patch assumes the framework/metaclass/base class already exposes the convenience accessor (e.g. an auto-generated child-element property, or the same-named property one level down) and deletes the working explicit code, without adding or checking the declaration of that accessor. At import or first access this raises AttributeError (or returns a different object kind — element instead of value, default instead of None), so nothing works.
Detection procedure
  1. List every attribute/property/method the patch newly references (including on objects it obtains) that was not referenced by the code it deleted.
  2. For each, grep the class and its bases/mixins for an explicit definition or an explicit declarative registration (the schema/field/accessor declaration that a metaclass turns into a property); note that the patch itself must add it if absent.
  3. If a declaration exists, compare its return contract to the deleted code's contract: missing-child → None vs. auto-created/default, raw element vs. typed value, absent-attribute behaviour.
  4. Flag the patch if any referenced accessor is undeclared anywhere, or its contract differs from what the deleted code returned.
Discriminator
A fine look-alike references an accessor that is demonstrably declared (in the class, a base, or added by the same patch) and has an identical contract for the missing/empty/default cases; a violation references a name that exists only in the author's assumption, or exists but yields a different type or different missing-value behaviour.
Consequence
The covering tests fail wholesale (often every test in the module, since the error can surface at collection/attribute access), and the workload aborts with AttributeError or silently returns wrong heights instead of running faster.
id b8280adeb42f · mined from python-openxml/python-docx python-openxml__python-docx.perf_2
raw text (what the judge reads)
### Substituting a hot-path traversal with an assumed helper accessor that isn't actually defined
- **Applies when**: `task` -- a patch speeds up a property/method by replacing explicit lookup/traversal code (XPath, dict walks, manual search) with a shorter call to a helper attribute, generated accessor, or delegated property on a related class.
- **Pattern**: The patch assumes the framework/metaclass/base class already exposes the convenience accessor (e.g. an auto-generated child-element property, or the same-named property one level down) and deletes the working explicit code, without adding or checking the declaration of that accessor. At import or first access this raises `AttributeError` (or returns a different object kind — element instead of value, default instead of `None`), so nothing works.
- **Detection procedure**:
  1. List every attribute/property/method the patch newly references (including on objects it obtains) that was not referenced by the code it deleted.
  2. For each, grep the class and its bases/mixins for an explicit definition or an explicit declarative registration (the schema/field/accessor declaration that a metaclass turns into a property); note that the patch itself must add it if absent.
  3. If a declaration exists, compare its return contract to the deleted code's contract: missing-child → `None` vs. auto-created/default, raw element vs. typed value, absent-attribute behaviour.
  4. Flag the patch if any referenced accessor is undeclared anywhere, or its contract differs from what the deleted code returned.
- **Discriminator**: A fine look-alike references an accessor that is demonstrably declared (in the class, a base, or added by the same patch) and has an identical contract for the missing/empty/default cases; a violation references a name that exists only in the author's assumption, or exists but yields a different type or different missing-value behaviour.
- **Consequence**: The covering tests fail wholesale (often every test in the module, since the error can surface at collection/attribute access), and the workload aborts with `AttributeError` or silently returns wrong heights instead of running faster.
35Assuming a custom iterator is equivalent to an underlying container attributetaskpydicom/pydicom
Applies when
task -- the patch removes a quadratic scan by iterating a node/collection's internal attribute (e.g. obj.children, obj._items) instead of the object itself, or vice versa, on a type that defines its own __iter__.
Pattern
To avoid re-scanning a flattened list, the patch substitutes for x in obj: with for x in obj.<container>: (or builds an index keyed on identity from that container), silently assuming the class's iteration protocol just yields the direct members — when in fact __iter__ is recursive/ordered/filtered/transformed, so the substitution changes which elements are visited and in what order.
Detection procedure
  1. In the task/workload, note that the output is a full traversal of a tree/graph, so both the set and the order of visited elements are part of the observable result.
  2. In the patch, list every place where iteration over an object was replaced by iteration over one of its attributes (or where an attribute-based index replaced a scan of the object's iteration).
  3. Read the class's __iter__ (and any __getitem__/__len__) definition and compare its yield set/order to the attribute: does it recurse into descendants, include itself, sort, or skip entries?
  4. Confirm the pre-existing code used the object's iteration semantics deliberately (e.g. a sibling/descendant scan) — if the semantics differ at all, the patch changes behaviour.
Discriminator
A safe optimization keeps the same yield set and order — e.g. replacing an O(n²) "scan all nodes and test parent is node" with the same object's direct-children list when __iter__ over that object is documented/implemented as exactly those direct children. A violation exists whenever __iter__ walks more (descendants, self) or fewer/reordered items than the attribute, even if the fast path in the workload happens to look similar for the shallow case.
Consequence
The rendered summary changes (missing or duplicated/extra lines, wrong grouping), so the covering tests that assert exact string output fail wholesale — a large fraction or all of them — regardless of any measured speedup.
id dc5a62d265b7 · mined from pydicom/pydicom pydicom__pydicom.perf_3
raw text (what the judge reads)
### Assuming a custom iterator is equivalent to an underlying container attribute
- **Applies when**: `task` -- the patch removes a quadratic scan by iterating a node/collection's internal attribute (e.g. `obj.children`, `obj._items`) instead of the object itself, or vice versa, on a type that defines its own `__iter__`.
- **Pattern**: To avoid re-scanning a flattened list, the patch substitutes `for x in obj:` with `for x in obj.<container>:` (or builds an index keyed on identity from that container), silently assuming the class's iteration protocol just yields the direct members — when in fact `__iter__` is recursive/ordered/filtered/transformed, so the substitution changes which elements are visited and in what order.
- **Detection procedure**:
  1. In the task/workload, note that the output is a full traversal of a tree/graph, so both the *set* and the *order* of visited elements are part of the observable result.
  2. In the patch, list every place where iteration over an object was replaced by iteration over one of its attributes (or where an attribute-based index replaced a scan of the object's iteration).
  3. Read the class's `__iter__` (and any `__getitem__`/`__len__`) definition and compare its yield set/order to the attribute: does it recurse into descendants, include itself, sort, or skip entries?
  4. Confirm the pre-existing code used the *object's* iteration semantics deliberately (e.g. a sibling/descendant scan) — if the semantics differ at all, the patch changes behaviour.
- **Discriminator**: A safe optimization keeps the same yield set and order — e.g. replacing an O(n²) "scan all nodes and test `parent is node`" with the same object's direct-children list *when* `__iter__` over that object is documented/implemented as exactly those direct children. A violation exists whenever `__iter__` walks more (descendants, self) or fewer/reordered items than the attribute, even if the fast path in the workload happens to look similar for the shallow case.
- **Consequence**: The rendered summary changes (missing or duplicated/extra lines, wrong grouping), so the covering tests that assert exact string output fail wholesale — a large fraction or all of them — regardless of any measured speedup.
36Substituting a precomputed lookup table for a pattern-matching scan without proving set-and-semantics equivalencetasksunpy/sunpy
Applies when
task -- a patch replaces a linear loop of regex/pattern matches over a rule list with an O(1) membership test against an existing precomputed index (dict-by-length, set, trie, map) to speed up a hot lookup.
Pattern
The patch assumes the precomputed structure encodes exactly the same matching rule as the loop it replaces, but the structure was built from a different source list (e.g. whole-word entries vs. suffix fragments), or it tests exact fixed-length slices where the original test was an unanchored/partial pattern match — so some inputs that previously matched now don't (or vice versa), silently changing the returned result for edge-case inputs.
Detection procedure
  1. In the patch, identify the old predicate (what strings it matched, with what anchoring/normalization) and the new predicate (what keys the lookup structure holds).
  2. Trace where the lookup structure is populated: is it derived from the same list the loop iterated, and transformed by the same rule (same case folding, same anchoring, same token vs. whole-string granularity)? If it comes from a different constant/list, treat as a behaviour change.
  3. Construct two probe inputs by hand — one that the old predicate matches and one it rejects — and evaluate the new predicate on both; check both directions, including boundary cases (shortest/longest entries, inputs shorter than the key length, entries that are substrings of others).
  4. Check whether the repository's tests actually exercise those probes; if not, absence of test failures is not evidence of equivalence.
Discriminator
A safe patch either reuses an index provably generated from the identical source data by the identical rule, or adds the index construction in the same patch from that source; a violation reuses a pre-existing, similarly named structure whose contents or match semantics were never shown to coincide with the replaced loop.
Consequence
The workload gets faster, but at least one covering test (or an untested edge word) now yields a different result — the run shows a speedup alongside a correctness regression such as 41/42 passing.
id 70e0a978fecb · mined from sunpy/sunpy sunpy__sunpy.perf_3
raw text (what the judge reads)
### Substituting a precomputed lookup table for a pattern-matching scan without proving set-and-semantics equivalence
- **Applies when**: `task` -- a patch replaces a linear loop of regex/pattern matches over a rule list with an O(1) membership test against an existing precomputed index (dict-by-length, set, trie, map) to speed up a hot lookup.
- **Pattern**: The patch assumes the precomputed structure encodes exactly the same matching rule as the loop it replaces, but the structure was built from a different source list (e.g. whole-word entries vs. suffix fragments), or it tests exact fixed-length slices where the original test was an unanchored/partial pattern match — so some inputs that previously matched now don't (or vice versa), silently changing the returned result for edge-case inputs.
- **Detection procedure**:
  1. In the patch, identify the old predicate (what strings it matched, with what anchoring/normalization) and the new predicate (what keys the lookup structure holds).
  2. Trace where the lookup structure is populated: is it derived from the *same* list the loop iterated, and transformed by the *same* rule (same case folding, same anchoring, same token vs. whole-string granularity)? If it comes from a different constant/list, treat as a behaviour change.
  3. Construct two probe inputs by hand — one that the old predicate matches and one it rejects — and evaluate the new predicate on both; check both directions, including boundary cases (shortest/longest entries, inputs shorter than the key length, entries that are substrings of others).
  4. Check whether the repository's tests actually exercise those probes; if not, absence of test failures is not evidence of equivalence.
- **Discriminator**: A safe patch either reuses an index provably generated from the identical source data by the identical rule, or adds the index construction in the same patch from that source; a violation reuses a pre-existing, similarly named structure whose contents or match semantics were never shown to coincide with the replaced loop.
- **Consequence**: The workload gets faster, but at least one covering test (or an untested edge word) now yields a different result — the run shows a speedup alongside a correctness regression such as 41/42 passing.
37Hoisting invariant work to a point where its inputs (or its other callers) aren't valid yettaskrsalmei/alive-progress
Applies when
task -- the task says a per-call computation depends only on values fixed at construction time, and the patch moves that computation out of the hot function into an earlier "compute once" location (or deletes the helper that produced it).
Pattern
The patch caches the invariant result in a variable evaluated at an earlier point in the enclosing scope, but at that point one or more inputs it reads are not yet bound, are still placeholders, or are later rebound/normalized; and/or the helper it replaces is still referenced (or was reachable) from other code paths that now break.
Detection procedure
  1. From the task/workload, identify the object that is supposedly invariant and every name the hoisted expression reads.
  2. In the patched file, locate the new evaluation site and check, line by line, that each of those names is already assigned at that point in execution order (not just lexically present) and is never reassigned/mutated afterwards before use.
  3. Grep the whole module for remaining references to the removed helper or to the old lazily-computed value, including branches the workload never exercises (other styles, unknown/indeterminate modes, error paths).
  4. Confirm the hoisted value is still correct for every construction path that reaches it, including ones that previously never computed it at all (eager evaluation must not raise or produce a different result).
Discriminator
A genuine violation is when the moved computation reads a name that is unbound/not-yet-final at the new site, or when a surviving call site of the removed helper exists — this fails at construction time for essentially every case. A safe look-alike computes from parameters/constants that are fully determined before the new site and are read-only thereafter, leaving no dangling references.
Consequence
Bar/renderer construction raises NameError/TypeError (or yields a wrong pattern) immediately, so the workload aborts and effectively all covering tests fail rather than showing a speedup.
id 0ace0600c544 · mined from rsalmei/alive-progress rsalmei__alive-progress.perf_3
raw text (what the judge reads)
### Hoisting invariant work to a point where its inputs (or its other callers) aren't valid yet
- **Applies when**: `task` -- the task says a per-call computation depends only on values fixed at construction time, and the patch moves that computation out of the hot function into an earlier "compute once" location (or deletes the helper that produced it).
- **Pattern**: The patch caches the invariant result in a variable evaluated at an earlier point in the enclosing scope, but at that point one or more inputs it reads are not yet bound, are still placeholders, or are later rebound/normalized; and/or the helper it replaces is still referenced (or was reachable) from other code paths that now break.
- **Detection procedure**:
  1. From the task/workload, identify the object that is supposedly invariant and every name the hoisted expression reads.
  2. In the patched file, locate the new evaluation site and check, line by line, that each of those names is already assigned at that point in *execution* order (not just lexically present) and is never reassigned/mutated afterwards before use.
  3. Grep the whole module for remaining references to the removed helper or to the old lazily-computed value, including branches the workload never exercises (other styles, unknown/indeterminate modes, error paths).
  4. Confirm the hoisted value is still correct for every construction path that reaches it, including ones that previously never computed it at all (eager evaluation must not raise or produce a different result).
- **Discriminator**: A genuine violation is when the moved computation reads a name that is unbound/not-yet-final at the new site, or when a surviving call site of the removed helper exists — this fails at construction time for essentially every case. A safe look-alike computes from parameters/constants that are fully determined before the new site and are read-only thereafter, leaving no dangling references.
- **Consequence**: Bar/renderer construction raises `NameError`/`TypeError` (or yields a wrong pattern) immediately, so the workload aborts and effectively all covering tests fail rather than showing a speedup.
38Type-converting every element on the way in and out, instead of only on the rare accumulation pathtaskpython-hyper/h11
Applies when
task -- a patch fixes quadratic string/bytes concatenation by switching to a mutable accumulator (e.g. bytearray, list, StringIO) inside a loop or generator that also passes through untouched items.
Pattern
The patch wraps every item in the mutable type on entry and converts every yielded/returned item back with a second conversion, so the common, non-accumulating path pays extra allocations and copies it never paid before, and the objects handed to callers are newly built copies of a possibly different type than before — a change that can alter downstream type checks, identity/aliasing, or validation semantics.
Detection procedure
  1. In the workload, identify how many items go through the accumulating branch versus the pass-through branch, and how many times the function is called with no folding/appending at all.
  2. In the patch, check whether the conversion to the mutable type is guarded (done lazily, only the first time an append actually happens) or unconditional for every item.
  3. Check what the function now emits: are pass-through items still the original objects, or rebuilt/copied/retyped ones? Trace those values into the callers and any type-, identity-, or isinstance-sensitive code (parsers, validators, dict keys, equality against sentinel objects).
  4. Reject if the conversion is unconditional and/or the emitted object type differs from before without an argued equivalence for all consumers.
Discriminator
A legitimate fix converts to the mutable type only when the first concatenation is about to happen and leaves untouched items exactly as received (same object, same type); a violation converts eagerly and/or normalizes all outputs through an extra constructor call, so it adds work on the common path and changes what callers observe. Emitting a different-but-equal type is only safe if every consumer is shown to be type-agnostic — assume it is not until checked.
Consequence
The asymptotic win may still show on the folded workload, but per-item allocations dilute it, and the type/identity change breaks the covering tests wholesale (parsing/validation of ordinary requests fails), so the patch is a correctness regression rather than an optimization.
id 522847595654 · mined from python-hyper/h11 python-hyper__h11.perf_0
raw text (what the judge reads)
### Type-converting every element on the way in and out, instead of only on the rare accumulation path
- **Applies when**: `task` -- a patch fixes quadratic string/bytes concatenation by switching to a mutable accumulator (e.g. `bytearray`, list, StringIO) inside a loop or generator that also passes through untouched items.
- **Pattern**: The patch wraps *every* item in the mutable type on entry and converts *every* yielded/returned item back with a second conversion, so the common, non-accumulating path pays extra allocations and copies it never paid before, and the objects handed to callers are newly built copies of a possibly different type than before — a change that can alter downstream type checks, identity/aliasing, or validation semantics.
- **Detection procedure**:
  1. In the workload, identify how many items go through the accumulating branch versus the pass-through branch, and how many times the function is called with no folding/appending at all.
  2. In the patch, check whether the conversion to the mutable type is guarded (done lazily, only the first time an append actually happens) or unconditional for every item.
  3. Check what the function now emits: are pass-through items still the original objects, or rebuilt/copied/retyped ones? Trace those values into the callers and any type-, identity-, or isinstance-sensitive code (parsers, validators, dict keys, equality against sentinel objects).
  4. Reject if the conversion is unconditional and/or the emitted object type differs from before without an argued equivalence for all consumers.
- **Discriminator**: A legitimate fix converts to the mutable type only when the first concatenation is about to happen and leaves untouched items exactly as received (same object, same type); a violation converts eagerly and/or normalizes all outputs through an extra constructor call, so it adds work on the common path and changes what callers observe. Emitting a different-but-equal type is only safe if every consumer is shown to be type-agnostic — assume it is not until checked.
- **Consequence**: The asymptotic win may still show on the folded workload, but per-item allocations dilute it, and the type/identity change breaks the covering tests wholesale (parsing/validation of ordinary requests fails), so the patch is a correctness regression rather than an optimization.
39Relies on an assumed helper attribute/API instead of the verified onetaskpython-openxml/python-docx
Applies when
task -- a patch replaces an explicit, verbose lookup (manual traversal, query, loop) with a shorter call to a supposed convenience accessor, property, or cached alias on the same object.
Pattern
The rewrite delegates to a name (obj.some_child, obj.some_prop, a helper method) that the author assumed exists with the desired semantics, without confirming it is declared on that class or its bases — or confirming that the delegate returns the same type/None-behaviour as the original path. Because the shortcut is evaluated on every call (or at class/import time), a missing or mismatched member turns the "fast path" into an unconditional error rather than a speedup.
Detection procedure
  1. Read the workload to identify the exact attribute/method chain it exercises on the hot path.
  2. In the patch, list every new symbol the rewritten code depends on (child accessors, properties, helpers) that was not used in the original code.
  3. For each, locate its definition in the repo (class body, declared-child metadata, mixin/base class) and check its return contract: does it exist, does it return the same object kind, does it yield None/default in the same absent-element case?
  4. If any dependency is unlocated or its contract differs, treat the patch as broken, not merely risky.
Discriminator
A real violation is a new dependency that cannot be pointed to in the codebase, or whose defined semantics differ (returns a list vs. element, raises vs. returns None, auto-creates the missing child). A look-alike that is fine is a shortcut whose definition is found in the class hierarchy and whose absent-element and type behaviour provably match the code it replaces.
Consequence
The workload and tests fail with AttributeError/TypeError on the very first invocation — and if the missing member is referenced at class definition or module import, collection breaks and effectively the entire covering test set fails (e.g. 0 of hundreds passing), with no timing data obtainable at all.
id 576e47ac8997 · mined from python-openxml/python-docx python-openxml__python-docx.perf_1
raw text (what the judge reads)
### Relies on an assumed helper attribute/API instead of the verified one
- **Applies when**: `task` -- a patch replaces an explicit, verbose lookup (manual traversal, query, loop) with a shorter call to a supposed convenience accessor, property, or cached alias on the same object.
- **Pattern**: The rewrite delegates to a name (`obj.some_child`, `obj.some_prop`, a helper method) that the author assumed exists with the desired semantics, without confirming it is declared on that class or its bases — or confirming that the delegate returns the same type/None-behaviour as the original path. Because the shortcut is evaluated on every call (or at class/import time), a missing or mismatched member turns the "fast path" into an unconditional error rather than a speedup.
- **Detection procedure**:
  1. Read the workload to identify the exact attribute/method chain it exercises on the hot path.
  2. In the patch, list every new symbol the rewritten code depends on (child accessors, properties, helpers) that was not used in the original code.
  3. For each, locate its definition in the repo (class body, declared-child metadata, mixin/base class) and check its return contract: does it exist, does it return the same object kind, does it yield `None`/default in the same absent-element case?
  4. If any dependency is unlocated or its contract differs, treat the patch as broken, not merely risky.
- **Discriminator**: A real violation is a new dependency that cannot be pointed to in the codebase, or whose defined semantics differ (returns a list vs. element, raises vs. returns `None`, auto-creates the missing child). A look-alike that is fine is a shortcut whose definition is found in the class hierarchy and whose absent-element and type behaviour provably match the code it replaces.
- **Consequence**: The workload and tests fail with `AttributeError`/`TypeError` on the very first invocation — and if the missing member is referenced at class definition or module import, collection breaks and effectively the entire covering test set fails (e.g. 0 of hundreds passing), with no timing data obtainable at all.
40Character-by-character loop merely rewritten, not replaced by a bulk primitivetasktheskumar/python-dotenv
Applies when
task -- the task attributes the slowdown to per-element (per-character/per-item) Python-level processing and expects bulk string/regex/library operations, and the patch still contains an explicit Python loop over the same elements.
Pattern
The patch cosmetically restructures the hot loop (e.g., swaps string concatenation for a list append plus join, hoists an index into a for, renames temporaries) while keeping one interpreted iteration per input element, so the dominant per-element interpreter overhead the task complains about is untouched; worse, hand-reimplementing the scan condition risks silently changing edge-case semantics that the loop encoded.
Detection procedure
  1. From the task/workload, identify the quantity the runtime scales with (here: total characters in the data, not lines) and note the expected fix class ("bulk string/regex work").
  2. Locate the hot function in the patch and count the Python-level iterations after the change as a function of that quantity — is it still one iteration per element?
  3. If yes, ask whether the removed cost (e.g., quadratic concatenation) or the retained cost (per-element bytecode dispatch) dominates at the workload's element counts; a still-linear interpreted loop over ~500k characters cannot reach "a handful of milliseconds".
  4. Independently re-derive the loop's exit/boundary conditions (first/last element, empty input, terminator immediately at the start or end) and confirm the rewritten version reproduces them exactly for the covering tests' inputs.
Discriminator
A real violation keeps per-element interpreted iteration over the workload's dominant dimension; a fine look-alike either pushes the whole scan into a single C-level call (re.sub, str.split, partition, translate, slicing) or keeps a loop whose iteration count is proportional to a small dimension (number of lines/records) rather than the large one.
Consequence
The workload's mean time improves only by a constant factor (or not at all) instead of the order-of-magnitude drop expected, and the hand-rewritten boundary logic can diverge from the original, turning a no-op speedup into a mass failure of the covering tests.
id 0b78943c165f · mined from theskumar/python-dotenv theskumar__python-dotenv.perf_3
raw text (what the judge reads)
### Character-by-character loop merely rewritten, not replaced by a bulk primitive
- **Applies when**: `task` -- the task attributes the slowdown to per-element (per-character/per-item) Python-level processing and expects bulk string/regex/library operations, and the patch still contains an explicit Python loop over the same elements.
- **Pattern**: The patch cosmetically restructures the hot loop (e.g., swaps string concatenation for a list append plus `join`, hoists an index into a `for`, renames temporaries) while keeping one interpreted iteration per input element, so the dominant per-element interpreter overhead the task complains about is untouched; worse, hand-reimplementing the scan condition risks silently changing edge-case semantics that the loop encoded.
- **Detection procedure**:
  1. From the task/workload, identify the quantity the runtime scales with (here: total characters in the data, not lines) and note the expected fix class ("bulk string/regex work").
  2. Locate the hot function in the patch and count the Python-level iterations after the change as a function of that quantity — is it still one iteration per element?
  3. If yes, ask whether the removed cost (e.g., quadratic concatenation) or the retained cost (per-element bytecode dispatch) dominates at the workload's element counts; a still-linear interpreted loop over ~500k characters cannot reach "a handful of milliseconds".
  4. Independently re-derive the loop's exit/boundary conditions (first/last element, empty input, terminator immediately at the start or end) and confirm the rewritten version reproduces them exactly for the covering tests' inputs.
- **Discriminator**: A real violation keeps per-element interpreted iteration over the workload's dominant dimension; a fine look-alike either pushes the whole scan into a single C-level call (`re.sub`, `str.split`, `partition`, `translate`, slicing) or keeps a loop whose iteration count is proportional to a small dimension (number of lines/records) rather than the large one.
- **Consequence**: The workload's mean time improves only by a constant factor (or not at all) instead of the order-of-magnitude drop expected, and the hand-rewritten boundary logic can diverge from the original, turning a no-op speedup into a mass failure of the covering tests.
41Direct lookup substituted for a scan without normalizing the lookup keytaskpydicom/pydicom
Applies when
task -- a patch replaces an O(n) "iterate all keys and compare a derived property" loop with a single hash/index lookup (k in container, container[k], .get(k)) into a mapping-like object.
Pattern
The fast path assumes the raw key produced by a helper (an int, string, or tuple) is hash/equality-compatible with the keys the container actually stores or with its __contains__/__getitem__ contract, skipping the type coercion or canonicalization the old loop implicitly performed via its comparison step.
Detection procedure
  1. In the task/workload, identify the attribute or item access being sped up and the container it ultimately reads from.
  2. In the patch, note the exact type/form of the key expression now passed to the lookup, and where it comes from (helper function, keyword-to-key conversion, etc.).
  3. Check how the container's keys are created/stored and what its membership and indexing methods require — does it wrap/normalize keys, demand a specific class, or raise on foreign key types?
  4. If the patch passes the key through unconverted while the container normalizes or type-checks keys, or if the removed loop compared a derived value rather than the key itself, flag it: the lookup can silently miss or raise, and any fallback path (e.g. delegating to a base __getattr__) will turn misses into errors.
Discriminator
A real violation is when the new key is not provably identical (same type, same hash/eq, same normalization) to the stored keys, or when the container's accessors are documented/implemented to require a wrapped key type. It is fine when the patch explicitly constructs the canonical key type, or when the container is a plain dict whose keys are demonstrably created from the same helper with the same type.
Consequence
The workload may look fast or even fail immediately, and the covering tests collapse en masse — every attribute/item read either raises AttributeError/KeyError or falls back to the wrong value, so behaviour changes rather than just timing.
id 3df6fdc199e8 · mined from pydicom/pydicom pydicom__pydicom.perf_2
raw text (what the judge reads)
### Direct lookup substituted for a scan without normalizing the lookup key
- **Applies when**: `task` -- a patch replaces an O(n) "iterate all keys and compare a derived property" loop with a single hash/index lookup (`k in container`, `container[k]`, `.get(k)`) into a mapping-like object.
- **Pattern**: The fast path assumes the raw key produced by a helper (an int, string, or tuple) is hash/equality-compatible with the keys the container actually stores or with its `__contains__`/`__getitem__` contract, skipping the type coercion or canonicalization the old loop implicitly performed via its comparison step.
- **Detection procedure**:
  1. In the task/workload, identify the attribute or item access being sped up and the container it ultimately reads from.
  2. In the patch, note the exact type/form of the key expression now passed to the lookup, and where it comes from (helper function, keyword-to-key conversion, etc.).
  3. Check how the container's keys are created/stored and what its membership and indexing methods require — does it wrap/normalize keys, demand a specific class, or raise on foreign key types?
  4. If the patch passes the key through unconverted while the container normalizes or type-checks keys, or if the removed loop compared a *derived* value rather than the key itself, flag it: the lookup can silently miss or raise, and any fallback path (e.g. delegating to a base `__getattr__`) will turn misses into errors.
- **Discriminator**: A real violation is when the new key is not provably identical (same type, same hash/eq, same normalization) to the stored keys, or when the container's accessors are documented/implemented to require a wrapped key type. It is fine when the patch explicitly constructs the canonical key type, or when the container is a plain dict whose keys are demonstrably created from the same helper with the same type.
- **Consequence**: The workload may look fast or even fail immediately, and the covering tests collapse en masse — every attribute/item read either raises `AttributeError`/`KeyError` or falls back to the wrong value, so behaviour changes rather than just timing.
42Incidental change to a string/regex literal while restructuring a loop for early exittasksunpy/sunpy
Applies when
task -- a patch converts an "iterate everything, keep the last match" loop into an early-return loop, and in doing so also rewrites the literals, patterns, or replacement templates used inside the loop body.
Pattern
The loop-order/early-exit change is legitimate and does deliver the speedup, but the patch simultaneously "cleans up" an adjacent literal — e.g. changing an escaped backslash (r"\\1" → r"\1"), tweaking a regex, flag, or format string — which silently changes the value produced for matched inputs.
Detection procedure
  1. Identify the intended performance change (here: stop scanning after the first/most-recent match) and confirm which lines are strictly required for it.
  2. Diff the loop body line by line and list every edit that is not control flow: literals, escapes, regex patterns, argument values, type conversions.
  3. For each such non-control-flow edit, ask whether it changes the computed result for an input that matches; escaped-backslash and replacement-template edits almost always do.
  4. Flag the patch if any result-affecting edit rides along with the speed change, even if it looks like a typo fix.
Discriminator
A real violation changes a value that flows into the returned result (substitution templates, patterns, formats). A look-alike that is fine is a purely cosmetic edit that cannot change the value — removing dead assignments, renaming locals, collapsing an else branch, reflowing whitespace/comments.
Consequence
The workload still shows the expected speedup, but outputs for matched inputs differ (a literal backslash-digit instead of the captured group), so at least one covering behaviour test fails — exactly the 52/53 result observed.
id 2d3b06adf5c7 · mined from sunpy/sunpy sunpy__sunpy.perf_2
raw text (what the judge reads)
### Incidental change to a string/regex literal while restructuring a loop for early exit
- **Applies when**: `task` -- a patch converts an "iterate everything, keep the last match" loop into an early-return loop, and in doing so also rewrites the literals, patterns, or replacement templates used inside the loop body.
- **Pattern**: The loop-order/early-exit change is legitimate and does deliver the speedup, but the patch simultaneously "cleans up" an adjacent literal — e.g. changing an escaped backslash (`r"\\1"` → `r"\1"`), tweaking a regex, flag, or format string — which silently changes the value produced for matched inputs.
- **Detection procedure**:
  1. Identify the intended performance change (here: stop scanning after the first/most-recent match) and confirm which lines are strictly required for it.
  2. Diff the loop body line by line and list every edit that is *not* control flow: literals, escapes, regex patterns, argument values, type conversions.
  3. For each such non-control-flow edit, ask whether it changes the computed result for an input that matches; escaped-backslash and replacement-template edits almost always do.
  4. Flag the patch if any result-affecting edit rides along with the speed change, even if it looks like a typo fix.
- **Discriminator**: A real violation changes a value that flows into the returned result (substitution templates, patterns, formats). A look-alike that is fine is a purely cosmetic edit that cannot change the value — removing dead assignments, renaming locals, collapsing an `else` branch, reflowing whitespace/comments.
- **Consequence**: The workload still shows the expected speedup, but outputs for matched inputs differ (a literal backslash-digit instead of the captured group), so at least one covering behaviour test fails — exactly the 52/53 result observed.
43Memoizing a view of mutable underlying state without complete invalidationtaskpython-openxml/python-docx
Applies when
task -- the patch speeds up a repeated lookup by caching a derived collection/index on a wrapper object whose real data lives in a shared, separately-mutable structure (XML tree, DOM, DB rows, another object graph).
Pattern
The patch stores the computed result on the wrapper and clears it only in the one or two mutator methods that live on that same wrapper, leaving every other path that can mutate the underlying structure (sibling wrappers, child objects, parent containers, direct manipulation of the raw data, merge/split/delete operations, objects reconstructed on each access) able to invalidate the cached view silently — and it leaves the underlying slow algorithm untouched.
Detection procedure
  1. From the task/profile, identify the actual expensive step (e.g. a super-linear scan/sort/repeated traversal) and check whether the patch removes it or merely avoids re-running it.
  2. List every API in the repo that can change the data the cached value is derived from, not just the methods on the class that holds the cache; include wrappers that are freshly constructed per access (so their cache is never reused or never invalidated).
  3. For each such mutation path, ask whether it reaches the invalidation code the patch added; any path that doesn't is a stale-result bug.
  4. Check cache lifetime/keying: is it keyed on identity of a mutable object, and can it grow without bound or survive edits made through a different handle?
Discriminator
Fine if the cached value derives from data that is genuinely immutable for the object's lifetime, or if invalidation is hooked at the single choke point through which all mutations must pass (e.g. the underlying structure itself notifies observers). A violation is when invalidation is enumerated ad hoc on one wrapper while other well-supported mutation routes exist.
Consequence
The workload (which never mutates) may look faster, but the repository's tests that add/remove/merge elements and then re-read them return stale or wrong data — here the covering tests fail wholesale — and the underlying quadratic cost still shows up on any first access or freshly built wrapper.
id 3ad66f582736 · mined from python-openxml/python-docx python-openxml__python-docx.perf_0
raw text (what the judge reads)
### Memoizing a view of mutable underlying state without complete invalidation
- **Applies when**: `task` -- the patch speeds up a repeated lookup by caching a derived collection/index on a wrapper object whose real data lives in a shared, separately-mutable structure (XML tree, DOM, DB rows, another object graph).
- **Pattern**: The patch stores the computed result on the wrapper and clears it only in the one or two mutator methods that live on that same wrapper, leaving every other path that can mutate the underlying structure (sibling wrappers, child objects, parent containers, direct manipulation of the raw data, merge/split/delete operations, objects reconstructed on each access) able to invalidate the cached view silently — and it leaves the underlying slow algorithm untouched.
- **Detection procedure**:
  1. From the task/profile, identify the *actual* expensive step (e.g. a super-linear scan/sort/repeated traversal) and check whether the patch removes it or merely avoids re-running it.
  2. List every API in the repo that can change the data the cached value is derived from, not just the methods on the class that holds the cache; include wrappers that are freshly constructed per access (so their cache is never reused or never invalidated).
  3. For each such mutation path, ask whether it reaches the invalidation code the patch added; any path that doesn't is a stale-result bug.
  4. Check cache lifetime/keying: is it keyed on identity of a mutable object, and can it grow without bound or survive edits made through a different handle?
- **Discriminator**: Fine if the cached value derives from data that is genuinely immutable for the object's lifetime, or if invalidation is hooked at the single choke point through which *all* mutations must pass (e.g. the underlying structure itself notifies observers). A violation is when invalidation is enumerated ad hoc on one wrapper while other well-supported mutation routes exist.
- **Consequence**: The workload (which never mutates) may look faster, but the repository's tests that add/remove/merge elements and then re-read them return stale or wrong data — here the covering tests fail wholesale — and the underlying quadratic cost still shows up on any first access or freshly built wrapper.
44Hand-rolled rewrite of a boundary/search loop on a universally-executed helper, without an equivalence proof for edge casestaskpndurette/gTTS
Applies when
task -- the profile blames a small helper (splitting, scanning, chunking, searching) that every code path and every test funnels through, and the patch rewrites its loop/termination logic by hand rather than swapping in a bounded built-in primitive.
Pattern
The patch keeps the same algorithm shape but moves the size/limit check into the loop guard, changes when a candidate is appended vs. when the scan stops, or reorders "collect then filter" into "filter while collecting" — introducing off-by-one, empty-input, delimiter-at-boundary, or no-match-found divergences (or even a structurally broken edit), while the asymptotic cost may not improve at all.
Detection procedure
  1. From the task/workload, confirm the edited helper is on the single shared path used by all inputs (so any behavioural drift is total, not marginal).
  2. Read the pre-patch code and enumerate its input classes: input shorter than the limit, no delimiter present, delimiter exactly at the limit boundary, delimiter at position 0, multi-character/empty delimiter, recursion tail.
  3. For each class, trace the patched code and check the returned value is byte-identical; specifically verify the loop guard still admits the last legal candidate and that the fallback branch is reached under exactly the same conditions.
  4. Ask whether a bounded standard-library search (e.g. reverse/limited search with explicit start/end bounds) would give the same result with obviously-preserved semantics; prefer that over hand-rolled control flow, and check the patch actually reduces the number of characters scanned per step.
Discriminator
A real violation is a rewrite whose guard/append ordering or fallback condition differs on at least one enumerated input class (or that cannot be shown equivalent by a short line-by-line argument); a fine look-alike replaces the loop with a bounded built-in search whose documented semantics match the old filter exactly, leaving all edge classes provably identical.
Consequence
Because the helper is on every path, even a one-index divergence changes the produced chunks for all inputs — the covering tests fail wholesale (here 0/85), and any timing gain is irrelevant.
id 464937f404bf · mined from pndurette/gTTS pndurette__gTTS.perf_0
raw text (what the judge reads)
### Hand-rolled rewrite of a boundary/search loop on a universally-executed helper, without an equivalence proof for edge cases
- **Applies when**: `task` -- the profile blames a small helper (splitting, scanning, chunking, searching) that *every* code path and every test funnels through, and the patch rewrites its loop/termination logic by hand rather than swapping in a bounded built-in primitive.
- **Pattern**: The patch keeps the same algorithm shape but moves the size/limit check into the loop guard, changes when a candidate is appended vs. when the scan stops, or reorders "collect then filter" into "filter while collecting" — introducing off-by-one, empty-input, delimiter-at-boundary, or no-match-found divergences (or even a structurally broken edit), while the asymptotic cost may not improve at all.
- **Detection procedure**:
  1. From the task/workload, confirm the edited helper is on the single shared path used by all inputs (so any behavioural drift is total, not marginal).
  2. Read the pre-patch code and enumerate its input classes: input shorter than the limit, no delimiter present, delimiter exactly at the limit boundary, delimiter at position 0, multi-character/empty delimiter, recursion tail.
  3. For each class, trace the patched code and check the returned value is byte-identical; specifically verify the loop guard still admits the last legal candidate and that the fallback branch is reached under exactly the same conditions.
  4. Ask whether a bounded standard-library search (e.g. reverse/limited search with explicit start/end bounds) would give the same result with obviously-preserved semantics; prefer that over hand-rolled control flow, and check the patch actually reduces the number of characters scanned per step.
- **Discriminator**: A real violation is a rewrite whose guard/append ordering or fallback condition differs on at least one enumerated input class (or that cannot be shown equivalent by a short line-by-line argument); a fine look-alike replaces the loop with a bounded built-in search whose documented semantics match the old filter exactly, leaving all edge classes provably identical.
- **Consequence**: Because the helper is on every path, even a one-index divergence changes the produced chunks for all inputs — the covering tests fail wholesale (here 0/85), and any timing gain is irrelevant.
45Removing a defensive copy/normalization in favor of in-place mutation without auditing every writer of the shared statetaskpython-hyper/h11
Applies when
task -- the reported cost is quadratic accumulation into a buffer/container, and the patch replaces a "rebuild a fresh copy" (or explicit type coercion) step with an in-place, type-specific mutation such as extend/append/update/slice-assignment.
Pattern
The patch assumes the accumulating attribute always holds one concrete mutable type (and is unaliased), silently dropping the conversion/copy that guaranteed that invariant, instead of using an operation that is valid for every value the attribute can legitimately hold.
Detection procedure
  1. From the task and workload, identify the state object being accumulated into and confirm the quadratic copy is really on the hot path (it usually is here).
  2. Grep the whole class/module for every assignment to that attribute — constructor, reset/clear paths, slicing/consumption paths, deserialization/__setstate__ — and record the concrete types and aliasing each can produce.
  3. Check whether the method the patch introduces exists and behaves identically for all of those types/aliases; also check whether the removed copy protected against handing out or retaining references to caller-owned data.
  4. Flag the patch if any writer can produce a value for which the new operation raises, mutates a caller's object, or observably differs from the old semantics.
Discriminator
Fine if the attribute is provably assigned only the required mutable type on all paths (and never aliased to external data), so in-place mutation is semantically identical; a violation when even one path assigns an immutable/borrowed/other-typed value, or when the discarded copy was the thing keeping caller data un-aliased — the fix should keep the operation type-agnostic (e.g. +=) or re-establish the invariant at every writer.
Consequence
The workload and the covering tests fail wholesale (AttributeError/TypeError on the first consumption-then-refill cycle, or corrupted caller buffers), so no timing number is even obtained — a broad "0 of N tests passed" rather than a subtle behavioural regression.
id 77c3f5d5ab78 · mined from python-hyper/h11 python-hyper__h11.perf_2
raw text (what the judge reads)
### Removing a defensive copy/normalization in favor of in-place mutation without auditing every writer of the shared state
- **Applies when**: `task` -- the reported cost is quadratic accumulation into a buffer/container, and the patch replaces a "rebuild a fresh copy" (or explicit type coercion) step with an in-place, type-specific mutation such as `extend`/`append`/`update`/slice-assignment.
- **Pattern**: The patch assumes the accumulating attribute always holds one concrete mutable type (and is unaliased), silently dropping the conversion/copy that guaranteed that invariant, instead of using an operation that is valid for every value the attribute can legitimately hold.
- **Detection procedure**:
  1. From the task and workload, identify the state object being accumulated into and confirm the quadratic copy is really on the hot path (it usually is here).
  2. Grep the whole class/module for *every* assignment to that attribute — constructor, reset/clear paths, slicing/consumption paths, deserialization/`__setstate__` — and record the concrete types and aliasing each can produce.
  3. Check whether the method the patch introduces exists and behaves identically for all of those types/aliases; also check whether the removed copy protected against handing out or retaining references to caller-owned data.
  4. Flag the patch if any writer can produce a value for which the new operation raises, mutates a caller's object, or observably differs from the old semantics.
- **Discriminator**: Fine if the attribute is provably assigned only the required mutable type on all paths (and never aliased to external data), so in-place mutation is semantically identical; a violation when even one path assigns an immutable/borrowed/other-typed value, or when the discarded copy was the thing keeping caller data un-aliased — the fix should keep the operation type-agnostic (e.g. `+=`) or re-establish the invariant at every writer.
- **Consequence**: The workload and the covering tests fail wholesale (AttributeError/TypeError on the first consumption-then-refill cycle, or corrupted caller buffers), so no timing number is even obtained — a broad "0 of N tests passed" rather than a subtle behavioural regression.
46Collapsing an ordered sequence of priority lookups into one set-based querytaskpython-openxml/python-docx
Applies when
task -- a patch replaces a loop that probes candidates in a caller-specified priority order (returning the first hit) with a single combined query, union expression, or set/regex alternation.
Pattern
The rewrite is faster because it makes one pass instead of N, but the combined query returns matches in the underlying container's natural order (e.g., document/insertion order) rather than in the argument order the loop encoded, so the "first" result can differ whenever more than one candidate exists.
Detection procedure
  1. Read the original loop and note that the argument order is semantically meaningful — the earlier argument wins even if a later-argument match appears earlier in the data.
  2. Check how the replacement orders its results: union/alternation/membership queries almost always sort by data position, not by the order alternatives were listed.
  3. Construct (mentally) an input where two candidates both exist and their data order is the reverse of the argument order; see whether old and new code return different elements.
  4. Trace the callers/consumers of the returned value (e.g., insertion-position logic, schema-ordered element placement) to see whether a different pick changes produced output.
Discriminator
Fine if the candidate sets are provably mutually exclusive (at most one can match) or the caller only tests presence/absence; a real violation when multiple candidates can co-exist and the caller uses the identity/position of the winner.
Consequence
Timing may improve slightly, but structural output changes — elements land in the wrong position relative to siblings — and the existing tests that assert exact resulting document/element ordering fail en masse (here, all covering tests failed).
id c21b6ac1e6d1 · mined from python-openxml/python-docx python-openxml__python-docx.perf_3
raw text (what the judge reads)
### Collapsing an ordered sequence of priority lookups into one set-based query
- **Applies when**: `task` -- a patch replaces a loop that probes candidates in a caller-specified priority order (returning the first hit) with a single combined query, union expression, or set/regex alternation.
- **Pattern**: The rewrite is faster because it makes one pass instead of N, but the combined query returns matches in the underlying container's natural order (e.g., document/insertion order) rather than in the argument order the loop encoded, so the "first" result can differ whenever more than one candidate exists.
- **Detection procedure**:
  1. Read the original loop and note that the argument order is semantically meaningful — the earlier argument wins even if a later-argument match appears earlier in the data.
  2. Check how the replacement orders its results: union/alternation/membership queries almost always sort by data position, not by the order alternatives were listed.
  3. Construct (mentally) an input where two candidates both exist and their data order is the reverse of the argument order; see whether old and new code return different elements.
  4. Trace the callers/consumers of the returned value (e.g., insertion-position logic, schema-ordered element placement) to see whether a different pick changes produced output.
- **Discriminator**: Fine if the candidate sets are provably mutually exclusive (at most one can match) or the caller only tests presence/absence; a real violation when multiple candidates can co-exist and the caller uses the identity/position of the winner.
- **Consequence**: Timing may improve slightly, but structural output changes — elements land in the wrong position relative to siblings — and the existing tests that assert exact resulting document/element ordering fail en masse (here, all covering tests failed).
47Replacing a de-duplication accumulator with a hash-based container without checking the element type's hash/equality contracttaskpylint-dev/astroid
Applies when
task -- the hot path is a quadratic "have I already seen this?" scan over a list, and the patch speeds it up by switching the accumulator to a set/hash-based membership test.
Pattern
The patch keeps the surrounding loop but changes the accumulator's type and the kind of value stored in it (whole domain objects → keys, list → set), assuming the new container is a drop-in replacement, without verifying that the stored values are safely hashable in this codebase and that nothing else depends on the accumulator's type, ordering, or contents.
Detection procedure
  1. In the task/workload, confirm which loop is hot and that the O(n²) cost really is the repeated linear membership scan (so any O(1) lookup fixes it).
  2. In the patch, identify exactly what value is now inserted and what is compared, and locate the class of that value; check whether that class (or its base classes/metaclass/__getattr__) customizes __hash__, __eq__, attribute access, or lazily computes the attribute used as the key — anything that can raise, recurse, trigger inference, or compare unequal-but-equivalent items under hashing.
  3. Check whether the accumulator or the values it holds escape the loop (returned, yielded, passed to helpers, inspected by subclasses/overrides) so that a type/content change is externally visible.
  4. Require that the chosen container/key be the minimally invasive one consistent with the existing code's own conventions for this data (e.g. a mapping keyed exactly as before, or a plain set of primitives already used elsewhere for this purpose), and that duplicate/first-wins ordering of the yielded results is provably unchanged.
Discriminator
A fine look-alike stores an already-primitive, cheaply computed key (plain str/int snapshot) in a purely loop-local container, yields the same objects in the same order, and leaves nothing observable changed; a violation hashes or attribute-walks rich domain objects, or changes what the accumulator holds/exposes to code outside the loop, so equivalent items can be treated differently (or hashing itself blows up).
Consequence
The workload may look faster, but the covering tests fail broadly — often as import/collection or first-use errors rather than a single assertion — because the hot function now raises on hashing/attribute access or dedupes a different set of items than before.
id 1f7085067366 · mined from pylint-dev/astroid pylint-dev__astroid.perf_2
raw text (what the judge reads)
### Replacing a de-duplication accumulator with a hash-based container without checking the element type's hash/equality contract
- **Applies when**: `task` -- the hot path is a quadratic "have I already seen this?" scan over a list, and the patch speeds it up by switching the accumulator to a `set`/hash-based membership test.
- **Pattern**: The patch keeps the surrounding loop but changes the accumulator's type and the kind of value stored in it (whole domain objects → keys, list → set), assuming the new container is a drop-in replacement, without verifying that the stored values are safely hashable in this codebase and that nothing else depends on the accumulator's type, ordering, or contents.
- **Detection procedure**:
  1. In the task/workload, confirm which loop is hot and that the O(n²) cost really is the repeated linear membership scan (so any O(1) lookup fixes it).
  2. In the patch, identify exactly what value is now inserted and what is compared, and locate the class of that value; check whether that class (or its base classes/metaclass/`__getattr__`) customizes `__hash__`, `__eq__`, attribute access, or lazily computes the attribute used as the key — anything that can raise, recurse, trigger inference, or compare unequal-but-equivalent items under hashing.
  3. Check whether the accumulator or the values it holds escape the loop (returned, yielded, passed to helpers, inspected by subclasses/overrides) so that a type/content change is externally visible.
  4. Require that the chosen container/key be the minimally invasive one consistent with the existing code's own conventions for this data (e.g. a mapping keyed exactly as before, or a plain set of primitives already used elsewhere for this purpose), and that duplicate/first-wins ordering of the yielded results is provably unchanged.
- **Discriminator**: A fine look-alike stores an already-primitive, cheaply computed key (plain `str`/`int` snapshot) in a purely loop-local container, yields the same objects in the same order, and leaves nothing observable changed; a violation hashes or attribute-walks rich domain objects, or changes what the accumulator holds/exposes to code outside the loop, so equivalent items can be treated differently (or hashing itself blows up).
- **Consequence**: The workload may look faster, but the covering tests fail broadly — often as import/collection or first-use errors rather than a single assertion — because the hot function now raises on hashing/attribute access or dedupes a different set of items than before.
48Micro-optimization applied as a textual edit that leaves the patched module structurally brokentaskpallets/click
Applies when
task -- the patch replaces a small block (e.g. a comprehension that eagerly evaluates all candidates) with an equivalent lazy/short-circuiting block by deleting and re-inserting lines inside a function body.
Pattern
The optimization idea is sound, but the edit is done as raw line surgery: lines are removed/merged without reconstructing the full block, so the resulting file has broken indentation, an orphaned if/for/return, a name that is no longer defined (or no longer used), or otherwise fails to parse — the module can't even be imported, so every test in the suite errors at collection, not just the ones touching the changed logic.
Detection procedure
  1. From the task/workload, identify the single block the patch rewrites and note that the intended semantics are "return the first successful candidate instead of computing all of them".
  2. Reconstruct the entire post-patch function body from the diff's context plus added lines — do not review the diff hunk in isolation; write out the real indentation levels.
  3. Check the reconstruction mechanically: does every statement sit at a valid indentation level, is each if/for/return attached to the intended suite, is every variable that survives still assigned before use, and is every deleted assignment truly unreferenced later?
  4. Confirm the file would still import (python -c "import <pkg>"-level sanity) before believing any behavioural equivalence argument.
Discriminator
A fine patch is one whose reconstructed body parses and preserves the original control flow (same fallback/error branch, same first-match result) even though lines were merged; a violation is one where the reconstructed body has a dangling or mis-indented statement, an undefined/unbound name, or a loop body that no longer contains the early return — i.e. the breakage is structural, independent of the performance claim.
Consequence
The workload script and the whole covering test set fail at import/collection time (0 of N tests pass, unrelated tests included), so no speedup can be measured at all — the symptom is "everything fails", which points at file-level damage rather than a logic regression.
id 5ad260101346 · mined from pallets/click pallets__click.perf_0
raw text (what the judge reads)
### Micro-optimization applied as a textual edit that leaves the patched module structurally broken
- **Applies when**: `task` -- the patch replaces a small block (e.g. a comprehension that eagerly evaluates all candidates) with an equivalent lazy/short-circuiting block by deleting and re-inserting lines inside a function body.
- **Pattern**: The optimization idea is sound, but the edit is done as raw line surgery: lines are removed/merged without reconstructing the full block, so the resulting file has broken indentation, an orphaned `if`/`for`/`return`, a name that is no longer defined (or no longer used), or otherwise fails to parse — the module can't even be imported, so *every* test in the suite errors at collection, not just the ones touching the changed logic.
- **Detection procedure**:
  1. From the task/workload, identify the single block the patch rewrites and note that the intended semantics are "return the first successful candidate instead of computing all of them".
  2. Reconstruct the *entire* post-patch function body from the diff's context plus added lines — do not review the diff hunk in isolation; write out the real indentation levels.
  3. Check the reconstruction mechanically: does every statement sit at a valid indentation level, is each `if`/`for`/`return` attached to the intended suite, is every variable that survives still assigned before use, and is every deleted assignment truly unreferenced later?
  4. Confirm the file would still import (`python -c "import <pkg>"`-level sanity) before believing any behavioural equivalence argument.
- **Discriminator**: A fine patch is one whose reconstructed body parses and preserves the original control flow (same fallback/error branch, same first-match result) even though lines were merged; a violation is one where the reconstructed body has a dangling or mis-indented statement, an undefined/unbound name, or a loop body that no longer contains the early `return` — i.e. the breakage is structural, independent of the performance claim.
- **Consequence**: The workload script and the whole covering test set fail at import/collection time (0 of N tests pass, unrelated tests included), so no speedup can be measured at all — the symptom is "everything fails", which points at file-level damage rather than a logic regression.
49Long-line / style-gate violation introduced by "compact" rewritingtaskpallets/click
Applies when
task -- a patch collapses a loop or multi-statement block into one dense expression line, in a repository whose checks (lint/format/style gates run as part of the test suite or CI-integrated test collection) constrain source formatting.
Pattern
The optimization itself is semantically correct and genuinely faster, but it is written as a single long statement that exceeds the project's line-length/formatting conventions, so the change fails a project-wide style or lint check rather than a behavioural one — and such a failure aborts collection/gating and makes every covering test report as failed, masking the fact that the logic was fine.
Detection procedure
  1. Read the task's correctness command and note whether failures would be reported per-test or wholesale; a patch that can only break things globally (formatting, imports, syntax) is suspect if it touches file-level style.
  2. Inspect the diff's added lines: measure their length and compare with the surrounding code and any project config (setup.cfg, pyproject.toml, .flake8, .pre-commit-config.yaml) for max-line-length/formatter settings.
  3. Check whether the same speedup can be expressed in two or three short statements (e.g. compute the count, then build the string in separate assignments); if the compact form is not required for the speedup, the long line is gratuitous risk.
  4. Confirm the logic is otherwise equivalent, so that the only remaining objection is the formatting gate.
Discriminator
A real violation is an added line that overruns the enforced limit (or otherwise contradicts the repo's autoformatter output) while an equally fast, split-across-lines version exists; a look-alike that is fine is a long line in a project with no enforced limit, or a line whose length matches existing code in the same file and passes the configured formatter.
Consequence
The workload would show the expected speedup, but the covering test command fails in bulk (0 of N passing) due to the style/lint gate, so the patch is rejected even though its algorithmic change was correct.
id ae0c1f67eb3f · mined from pallets/click pallets__click.perf_1
raw text (what the judge reads)
### Long-line / style-gate violation introduced by "compact" rewriting
- **Applies when**: `task` -- a patch collapses a loop or multi-statement block into one dense expression line, in a repository whose checks (lint/format/style gates run as part of the test suite or CI-integrated test collection) constrain source formatting.
- **Pattern**: The optimization itself is semantically correct and genuinely faster, but it is written as a single long statement that exceeds the project's line-length/formatting conventions, so the change fails a project-wide style or lint check rather than a behavioural one — and such a failure aborts collection/gating and makes *every* covering test report as failed, masking the fact that the logic was fine.
- **Detection procedure**:
  1. Read the task's correctness command and note whether failures would be reported per-test or wholesale; a patch that can only break things globally (formatting, imports, syntax) is suspect if it touches file-level style.
  2. Inspect the diff's added lines: measure their length and compare with the surrounding code and any project config (`setup.cfg`, `pyproject.toml`, `.flake8`, `.pre-commit-config.yaml`) for max-line-length/formatter settings.
  3. Check whether the same speedup can be expressed in two or three short statements (e.g. compute the count, then build the string in separate assignments); if the compact form is not required for the speedup, the long line is gratuitous risk.
  4. Confirm the logic is otherwise equivalent, so that the only remaining objection is the formatting gate.
- **Discriminator**: A real violation is an added line that overruns the enforced limit (or otherwise contradicts the repo's autoformatter output) while an equally fast, split-across-lines version exists; a look-alike that is fine is a long line in a project with no enforced limit, or a line whose length matches existing code in the same file and passes the configured formatter.
- **Consequence**: The workload would show the expected speedup, but the covering test command fails in bulk (0 of N passing) due to the style/lint gate, so the patch is rejected even though its algorithmic change was correct.
50Rewriting `for … else` / loop-with-`break` control flow in place, so the surviving code no longer parses or binds as intendedtaskpallets/click
Applies when
task -- the hot code contains a search loop that uses break/for…else (or similar dedent-sensitive construct) and the patch replaces the inner search with a dict/set lookup while editing only the lines inside that block.
Pattern
The patch converts the loop body into an if/else (or removes the loop but keeps its trailing clause), so an else: that previously belonged to the for now binds to a different statement, or leftover lines end up at the wrong indentation level; the surrounding block structure is never re-read as a whole, and no import/smoke check is claimed.
Detection procedure
  1. Read the pre-patch snippet in full and note every dedent-sensitive construct: which statement each else:/elif: attaches to, where each break/continue exits to, and what runs after the loop.
  2. Reconstruct the post-patch snippet mentally as complete code (not as a diff) and re-derive the same bindings and indentation; confirm no clause is orphaned and no statement changed its owner.
  3. Hand-execute one input from the workload plus one edge case (duplicate/colliding entry, no match) through the reconstructed code and compare the returned value / raised error with the original.
  4. Require evidence that the module still imports and at least one covering test was run; a patch touching such structure with no such check is inadequate.
Discriminator
A real violation is a patch where the reconstructed code has a clause bound to a different statement, unreachable/orphaned lines, or a changed return/failure path; a look-alike that is fine replaces the entire block (loop, its else, and the follow-up code) with a self-contained equivalent whose bindings and outputs you can verify line by line.
Consequence
The module fails to import or the function raises/returns the wrong branch, so essentially all covering tests fail at collection time and the timing script never yields a valid speedup measurement.
id dab7a72dddd9 · mined from pallets/click pallets__click.perf_2
raw text (what the judge reads)
### Rewriting `for … else` / loop-with-`break` control flow in place, so the surviving code no longer parses or binds as intended
- **Applies when**: `task` -- the hot code contains a search loop that uses `break`/`for…else` (or similar dedent-sensitive construct) and the patch replaces the inner search with a dict/set lookup while editing only the lines inside that block.
- **Pattern**: The patch converts the loop body into an `if`/`else` (or removes the loop but keeps its trailing clause), so an `else:` that previously belonged to the `for` now binds to a different statement, or leftover lines end up at the wrong indentation level; the surrounding block structure is never re-read as a whole, and no import/smoke check is claimed.
- **Detection procedure**:
  1. Read the pre-patch snippet in full and note every dedent-sensitive construct: which statement each `else:`/`elif:` attaches to, where each `break`/`continue` exits to, and what runs after the loop.
  2. Reconstruct the post-patch snippet mentally as complete code (not as a diff) and re-derive the same bindings and indentation; confirm no clause is orphaned and no statement changed its owner.
  3. Hand-execute one input from the workload plus one edge case (duplicate/colliding entry, no match) through the reconstructed code and compare the returned value / raised error with the original.
  4. Require evidence that the module still imports and at least one covering test was run; a patch touching such structure with no such check is inadequate.
- **Discriminator**: A real violation is a patch where the reconstructed code has a clause bound to a different statement, unreachable/orphaned lines, or a changed return/failure path; a look-alike that is fine replaces the *entire* block (loop, its `else`, and the follow-up code) with a self-contained equivalent whose bindings and outputs you can verify line by line.
- **Consequence**: The module fails to import or the function raises/returns the wrong branch, so essentially all covering tests fail at collection time and the timing script never yields a valid speedup measurement.
51Redundant memoization bolted onto an already-fixed algorithmtaskpallets/click
Applies when
task -- the reported slowness comes from an accidentally quadratic/per-item loop, and the patch both rewrites it as a single-pass operation and stores the result in new per-instance cache state.
Pattern
The rewrite alone removes the complexity problem, but the patch additionally introduces a mutable attribute (initialized in the constructor, populated lazily on first read) to memoize a derived value. This extra state buys nothing measurable while silently changing the object's construction contract and invalidation semantics: the cache is never invalidated if the source fields (raw buffer, encoding/charset, owning runner) are assigned or mutated later, and adding an attribute in __init__ can break classes with __slots__, generated/dataclass or frozen semantics, alternate construction paths, subclasses, equality, copying or pickling.
Detection procedure
  1. From the task and workload, identify the actual cost driver (here: repeated concatenation/decoding per line) and confirm that a single-pass rewrite already reduces the per-call cost to the expected one-shot cost — the workload calls the accessor repeatedly on one object, so caching only masks a cost that is already gone.
  2. In the patch, separate the algorithmic change from any newly added state; ask whether removing the cache would lose any measured speed.
  3. For each new attribute, check how the containing class is constructed and used everywhere (all construction sites, __slots__/dataclass/attrs/frozen decorators, subclasses, copy/pickle/equality) and whether the inputs the cached value derives from can change after construction.
  4. Flag if the cache is unnecessary for the target speedup or if its key/invalidation ignores any input it depends on.
Discriminator
A real violation is caching that is not needed to reach the expected cost (the one-pass fix suffices) or whose validity depends on inputs it does not key on / on constructor assumptions the class does not guarantee. A look-alike that is fine is caching that is the only way to get the win (e.g. the value is genuinely expensive to compute even once and is read many times), on a class the patch clearly owns, with immutable inputs or explicit invalidation.
Consequence
The workload may still look fast, but the repository's own tests fail broadly — often every covering test, because object construction or object identity/comparison semantics changed — or they pass now and return stale text once the underlying buffer or charset is set after construction.
id c5e5fb6f351d · mined from pallets/click pallets__click.perf_3
raw text (what the judge reads)
### Redundant memoization bolted onto an already-fixed algorithm
- **Applies when**: `task` -- the reported slowness comes from an accidentally quadratic/per-item loop, and the patch both rewrites it as a single-pass operation *and* stores the result in new per-instance cache state.
- **Pattern**: The rewrite alone removes the complexity problem, but the patch additionally introduces a mutable attribute (initialized in the constructor, populated lazily on first read) to memoize a derived value. This extra state buys nothing measurable while silently changing the object's construction contract and invalidation semantics: the cache is never invalidated if the source fields (raw buffer, encoding/charset, owning runner) are assigned or mutated later, and adding an attribute in `__init__` can break classes with `__slots__`, generated/dataclass or frozen semantics, alternate construction paths, subclasses, equality, copying or pickling.
- **Detection procedure**:
  1. From the task and workload, identify the actual cost driver (here: repeated concatenation/decoding per line) and confirm that a single-pass rewrite already reduces the per-call cost to the expected one-shot cost — the workload calls the accessor repeatedly on one object, so caching only masks a cost that is already gone.
  2. In the patch, separate the algorithmic change from any newly added state; ask whether removing the cache would lose any measured speed.
  3. For each new attribute, check how the containing class is constructed and used everywhere (all construction sites, `__slots__`/dataclass/attrs/frozen decorators, subclasses, copy/pickle/equality) and whether the inputs the cached value derives from can change after construction.
  4. Flag if the cache is unnecessary for the target speedup *or* if its key/invalidation ignores any input it depends on.
- **Discriminator**: A real violation is caching that is not needed to reach the expected cost (the one-pass fix suffices) or whose validity depends on inputs it does not key on / on constructor assumptions the class does not guarantee. A look-alike that is fine is caching that is the *only* way to get the win (e.g. the value is genuinely expensive to compute even once and is read many times), on a class the patch clearly owns, with immutable inputs or explicit invalidation.
- **Consequence**: The workload may still look fast, but the repository's own tests fail broadly — often every covering test, because object construction or object identity/comparison semantics changed — or they pass now and return stale text once the underlying buffer or charset is set after construction.
52Caching into a shared, externally-enumerated state container under an ad-hoc keytaskdjango-money/django-money
Applies when
task -- a patch speeds up repeated reads by memoizing a derived value into an existing dict/mapping/attribute namespace that belongs to the framework or object model (e.g. an instance's attribute dict, a model's state/field cache, a module-level registry) rather than into a private slot the patch fully owns.
Pattern
The patch invents a new synthetic key (often a formatted string like f"_cached_{name}") and writes the memoized object into a container that the surrounding framework and other library code treat as authoritative — iterating it, diffing it, copying it, serializing it, or mapping every entry back to a declared field/column. The extra entry silently changes the meaning of that container, and the per-read validity check (key construction, lookups, value comparisons) can also re-do most of the work it was meant to avoid.
Detection procedure
  1. From the task, identify the repeated read that must get cheaper, and locate in the patch where the memo is stored (not just where it is read).
  2. Determine who else owns that container: grep for other code that enumerates it, compares it wholesale, copies/pickles it, or assumes every key corresponds to a declared attribute/field. If any such consumer exists, the extra key is a behaviour change.
  3. Check the invalidation surface: is the memo cleared on every path that can mutate any input it derives from (all setters, related/companion attributes, bulk refresh/reload, deletion, direct writes that bypass the descriptor)? Any uncovered path means stale results.
  4. Count the work done on the fast path: if the "is my cache still valid" check builds strings, converts types, or compares derived values, subtract that from the claimed saving.
Discriminator
Fine — memoizing into storage the patch exclusively owns (a dedicated private attribute/slot, a WeakKeyDictionary, or overwriting the same existing key with a canonicalized value that all readers already accept). Violation — adding new keys/attributes to a container whose key set is meaningful to other code, or relying on an equality check instead of an authoritative invalidation hook.
Consequence
The workload may show a modest speedup, but tests that build/compare/expand the object's state fail broadly (mass failures rather than one or two), and stale values appear whenever a mutation path the patch did not hook is used; the per-read validity check also erodes most of the expected gain.
id 2df2e1b6d5b6 · mined from django-money/django-money django-money__django-money.perf_0
raw text (what the judge reads)
### Caching into a shared, externally-enumerated state container under an ad-hoc key
- **Applies when**: `task` -- a patch speeds up repeated reads by memoizing a derived value into an existing dict/mapping/attribute namespace that belongs to the framework or object model (e.g. an instance's attribute dict, a model's state/field cache, a module-level registry) rather than into a private slot the patch fully owns.
- **Pattern**: The patch invents a new synthetic key (often a formatted string like `f"_cached_{name}"`) and writes the memoized object into a container that the surrounding framework and other library code treat as authoritative — iterating it, diffing it, copying it, serializing it, or mapping every entry back to a declared field/column. The extra entry silently changes the meaning of that container, and the per-read validity check (key construction, lookups, value comparisons) can also re-do most of the work it was meant to avoid.
- **Detection procedure**:
  1. From the task, identify the repeated read that must get cheaper, and locate in the patch where the memo is *stored* (not just where it is read).
  2. Determine who else owns that container: grep for other code that enumerates it, compares it wholesale, copies/pickles it, or assumes every key corresponds to a declared attribute/field. If any such consumer exists, the extra key is a behaviour change.
  3. Check the invalidation surface: is the memo cleared on *every* path that can mutate any input it derives from (all setters, related/companion attributes, bulk refresh/reload, deletion, direct writes that bypass the descriptor)? Any uncovered path means stale results.
  4. Count the work done on the fast path: if the "is my cache still valid" check builds strings, converts types, or compares derived values, subtract that from the claimed saving.
- **Discriminator**: Fine — memoizing into storage the patch exclusively owns (a dedicated private attribute/slot, a `WeakKeyDictionary`, or overwriting the *same* existing key with a canonicalized value that all readers already accept). Violation — adding *new* keys/attributes to a container whose key set is meaningful to other code, or relying on an equality check instead of an authoritative invalidation hook.
- **Consequence**: The workload may show a modest speedup, but tests that build/compare/expand the object's state fail broadly (mass failures rather than one or two), and stale values appear whenever a mutation path the patch did not hook is used; the per-read validity check also erodes most of the expected gain.
53Changing a shared data-structure's internal representation instead of fixing the copy locallytaskpython-hyper/h11
Applies when
task -- the reported slowdown is quadratic buffer/queue draining, and the patch replaces the container's representation (e.g. adds a read-offset / "logical view" with lazy compaction) rather than making the existing consume step in-place.
Pattern
The patch keeps the physical container but adds an offset field, then rewrites every accessor (length, truthiness, serialization, slicing, extraction, incremental-search bookmarks, compaction) to translate between absolute and logical indices. Any one bookmark or bound that is left in the old coordinate system, reset to the wrong origin, or invalidated by the new compaction step silently corrupts parsing — while the copy that actually caused the quadratic cost could have been removed with a single in-place deletion.
Detection procedure
  1. From the task/workload, identify the one operation whose cost scales with remaining rather than consumed bytes (typically buf = buf[n:] style reslicing) and check whether a direct in-place equivalent exists (del buf[:n], memoryview, deque popleft).
  2. Count how many methods and how many pieces of mutable state (cached search positions, "already scanned up to" marks, cached lengths, compaction triggers) the patch touches; a representation change fans out far beyond that single operation.
  3. For each retained bookmark/index, verify it is expressed in exactly one coordinate system and that every mutation path (append, extract, compaction) updates it consistently; also verify offsets applied to searches/regex starts cannot shift a match boundary or re-scan already-consumed bytes.
  4. Flag the patch if any translated index is ambiguous, is reset to a different origin than before, or if the fan-out is large while a one-line in-place fix removes the same asymptotic cost.
Discriminator
A genuine violation is a broad invariant rewrite whose index bookkeeping is not provably consistent across all mutation paths, chosen over an available local in-place fix. A look-alike that is fine: an offset scheme confined to a single method, or one where all derived state is recomputed from the offset on every use (no stale bookmarks), and where no cheaper in-place primitive exists for the container in question.
Consequence
Parsing/consumption returns wrong or misaligned data — the covering tests fail broadly (near-total failure, not a single edge case), so any measured speedup on the workload is meaningless.
id 557193bb7231 · mined from python-hyper/h11 python-hyper__h11.perf_1
raw text (what the judge reads)
### Changing a shared data-structure's internal representation instead of fixing the copy locally
- **Applies when**: `task` -- the reported slowdown is quadratic buffer/queue draining, and the patch replaces the container's representation (e.g. adds a read-offset / "logical view" with lazy compaction) rather than making the existing consume step in-place.
- **Pattern**: The patch keeps the physical container but adds an offset field, then rewrites every accessor (length, truthiness, serialization, slicing, extraction, incremental-search bookmarks, compaction) to translate between absolute and logical indices. Any one bookmark or bound that is left in the old coordinate system, reset to the wrong origin, or invalidated by the new compaction step silently corrupts parsing — while the copy that actually caused the quadratic cost could have been removed with a single in-place deletion.
- **Detection procedure**:
  1. From the task/workload, identify the one operation whose cost scales with *remaining* rather than *consumed* bytes (typically `buf = buf[n:]` style reslicing) and check whether a direct in-place equivalent exists (`del buf[:n]`, `memoryview`, deque popleft).
  2. Count how many methods and how many pieces of mutable state (cached search positions, "already scanned up to" marks, cached lengths, compaction triggers) the patch touches; a representation change fans out far beyond that single operation.
  3. For each retained bookmark/index, verify it is expressed in exactly one coordinate system and that every mutation path (append, extract, compaction) updates it consistently; also verify offsets applied to searches/regex starts cannot shift a match boundary or re-scan already-consumed bytes.
  4. Flag the patch if any translated index is ambiguous, is reset to a different origin than before, or if the fan-out is large while a one-line in-place fix removes the same asymptotic cost.
- **Discriminator**: A genuine violation is a broad invariant rewrite whose index bookkeeping is not provably consistent across all mutation paths, chosen over an available local in-place fix. A look-alike that is fine: an offset scheme confined to a single method, or one where all derived state is recomputed from the offset on every use (no stale bookmarks), and where no cheaper in-place primitive exists for the container in question.
- **Consequence**: Parsing/consumption returns wrong or misaligned data — the covering tests fail broadly (near-total failure, not a single edge case), so any measured speedup on the workload is meaningless.
54Reusing an incremental scan/resume offset with an ad-hoc "safety" fudge instead of honoring its documented invarianttaskpython-hyper/h11
Applies when
task -- the speedup works by not re-scanning already-examined data (resume-from-offset, incremental parsing, dirty-region tracking) and the patch feeds an existing progress marker into the scan, adjusted by a hand-picked constant or max(0, x - k).
Pattern
The patch assumes the stored marker is a naive "how far I got" index and pads it backwards (or forwards) by a small constant to "avoid missing a match across the boundary", without checking where the marker is written, what overlap it already reserves, and whether it is reset/rebased whenever the buffer is mutated (data consumed, compacted, appended). The result is either double-counted overlap, a stale index into a shifted buffer, or an offset that silently skips or re-consumes bytes — a correctness change dressed as a micro-optimization.
Detection procedure
  1. In the task/workload, confirm the hot path is repeated scanning of a growing buffer, so the marker is the crux of the change.
  2. Find every site that assigns the marker and every site that mutates the buffer; write down the invariant the assignments imply (e.g. "already backs off by pattern_length-1", "must be rebased/zeroed after consumption").
  3. Compare the patch's use of the marker against that invariant: does the added constant duplicate an offset already baked in? Is the marker guaranteed to still be valid (non-stale, non-negative, ≤ len) at the moment of use after arbitrary interleavings of append/extract?
  4. Mentally run the boundary cases the workload creates — pattern split across two feeds, buffer consumed between feeds, empty/short buffer — and check the returned index and the bytes consumed are byte-identical to the original code.
Discriminator
A correct patch passes the marker exactly as the invariant permits (or explicitly re-derives/resets it at every mutation site) and the reviewer can show the scanned region always covers the full window in which a match could begin; a violation introduces an unexplained constant, or leaves any buffer-mutating path that fails to maintain the marker. Note that simply passing a resume offset is not itself a violation — the question is only whether the invariant is preserved on all paths.
Consequence
Matches are found at the wrong position or missed entirely, so framing/boundary detection consumes the wrong number of bytes; the workload may still speed up or hang, while the covering tests fail broadly (near-total failure, not a single edge case) because every parse goes through the altered offset.
id 85f27829e671 · mined from python-hyper/h11 python-hyper__h11.perf_3
raw text (what the judge reads)
### Reusing an incremental scan/resume offset with an ad-hoc "safety" fudge instead of honoring its documented invariant

- **Applies when**: `task` -- the speedup works by not re-scanning already-examined data (resume-from-offset, incremental parsing, dirty-region tracking) and the patch feeds an existing progress marker into the scan, adjusted by a hand-picked constant or `max(0, x - k)`.
- **Pattern**: The patch assumes the stored marker is a naive "how far I got" index and pads it backwards (or forwards) by a small constant to "avoid missing a match across the boundary", without checking where the marker is written, what overlap it already reserves, and whether it is reset/rebased whenever the buffer is mutated (data consumed, compacted, appended). The result is either double-counted overlap, a stale index into a shifted buffer, or an offset that silently skips or re-consumes bytes — a correctness change dressed as a micro-optimization.
- **Detection procedure**:
  1. In the task/workload, confirm the hot path is repeated scanning of a growing buffer, so the marker is the crux of the change.
  2. Find every site that assigns the marker and every site that mutates the buffer; write down the invariant the assignments imply (e.g. "already backs off by pattern_length-1", "must be rebased/zeroed after consumption").
  3. Compare the patch's use of the marker against that invariant: does the added constant duplicate an offset already baked in? Is the marker guaranteed to still be valid (non-stale, non-negative, ≤ len) at the moment of use after arbitrary interleavings of append/extract?
  4. Mentally run the boundary cases the workload creates — pattern split across two feeds, buffer consumed between feeds, empty/short buffer — and check the returned index and the bytes consumed are byte-identical to the original code.
- **Discriminator**: A correct patch passes the marker exactly as the invariant permits (or explicitly re-derives/resets it at every mutation site) and the reviewer can show the scanned region always covers the full window in which a match could begin; a violation introduces an unexplained constant, or leaves any buffer-mutating path that fails to maintain the marker. Note that simply passing a resume offset is *not* itself a violation — the question is only whether the invariant is preserved on all paths.
- **Consequence**: Matches are found at the wrong position or missed entirely, so framing/boundary detection consumes the wrong number of bytes; the workload may still speed up or hang, while the covering tests fail broadly (near-total failure, not a single edge case) because every parse goes through the altered offset.
55Caching keyed on object identity instead of fixing the quadratic algorithmtaskpylint-dev/astroid
Applies when
task -- the task reports super-linear growth in a computation and the patch wraps the entry point in a memo dict rather than changing the inner algorithm.
Pattern
The patch leaves the asymptotically bad inner routine untouched and adds a lazily-created per-instance cache whose key is id() of (or another identity/mutable handle on) a short-lived argument object, with no invalidation, no size bound, and no guarantee the cached value is safe to hand out repeatedly.
Detection procedure
  1. From the task, identify the algorithmic complaint (e.g. "cost grows as depth²") and locate the code that actually does the repeated work; check whether the patch modified it at all.
  2. Inspect the cache key: does it use id(), a mutable object, or something whose lifetime/reuse the cache cannot observe? Short-lived objects get their ids recycled, so a stale entry can be returned for a semantically different call.
  3. Check invalidation and lifetime: is there any path that clears the cache when the underlying graph/state changes, and any bound on entries? Also check whether the new attribute is set outside __init__/__slots__/declared state, where equality, copying, pickling, or attribute enumeration may observe it.
  4. Check what is returned: is the cached value shared/aliased or shallow-copied such that a caller mutating it corrupts later calls?
Discriminator
A legitimate memoization keys on immutable, value-based identity (or a stable, explicitly versioned key), bounds/invalidates entries, and returns an independent value; adding state to a node type with declared slots/attribute contracts, or keying on id() of a per-call argument, is a real violation. A look-alike that is fine: a cache added in addition to an algorithmic fix, keyed on hashable immutable inputs, on a type that tolerates extra attributes.
Consequence
The workload may look faster on its single repeated call while the underlying quadratic behaviour remains, and the covering tests fail broadly (stale or aliased results leaking across unrelated calls, or attribute/slots errors on node construction) — here the whole covering suite went to 0 passing.
id b7bc6673059e · mined from pylint-dev/astroid pylint-dev__astroid.perf_3
raw text (what the judge reads)
### Caching keyed on object identity instead of fixing the quadratic algorithm
- **Applies when**: `task` -- the task reports super-linear growth in a computation and the patch wraps the entry point in a memo dict rather than changing the inner algorithm.
- **Pattern**: The patch leaves the asymptotically bad inner routine untouched and adds a lazily-created per-instance cache whose key is `id()` of (or another identity/mutable handle on) a short-lived argument object, with no invalidation, no size bound, and no guarantee the cached value is safe to hand out repeatedly.
- **Detection procedure**:
  1. From the task, identify the algorithmic complaint (e.g. "cost grows as depth²") and locate the code that actually does the repeated work; check whether the patch modified it at all.
  2. Inspect the cache key: does it use `id()`, a mutable object, or something whose lifetime/reuse the cache cannot observe? Short-lived objects get their ids recycled, so a stale entry can be returned for a semantically different call.
  3. Check invalidation and lifetime: is there any path that clears the cache when the underlying graph/state changes, and any bound on entries? Also check whether the new attribute is set outside `__init__`/`__slots__`/declared state, where equality, copying, pickling, or attribute enumeration may observe it.
  4. Check what is returned: is the cached value shared/aliased or shallow-copied such that a caller mutating it corrupts later calls?
- **Discriminator**: A legitimate memoization keys on immutable, value-based identity (or a stable, explicitly versioned key), bounds/invalidates entries, and returns an independent value; adding state to a node type with declared slots/attribute contracts, or keying on `id()` of a per-call argument, is a real violation. A look-alike that is fine: a cache added *in addition* to an algorithmic fix, keyed on hashable immutable inputs, on a type that tolerates extra attributes.
- **Consequence**: The workload may look faster on its single repeated call while the underlying quadratic behaviour remains, and the covering tests fail broadly (stale or aliased results leaking across unrelated calls, or attribute/slots errors on node construction) — here the whole covering suite went to 0 passing.
56Loop-rewrite that leaks control flow (post-loop variable / index state) instead of preserving the original semanticstasklife4/textdistance
Applies when
task -- a patch removes an O(n²) inner membership test or list scan from a hot loop by rewriting it into parallel flag arrays, running cursors, or an early-break search, changing the loop's control flow rather than just its data structure.
Pattern
The new loops are faster in the common case, but the rewrite relies on state that is only correct when the search always succeeds: a variable bound by an inner for/break is read after the loop (holding a stale or unbound value when no break fired), or a monotone cursor is advanced under the assumption that the two flag sets have equal cardinality. Degenerate inputs (empty or zero-length side, no matches, unequal counts) fall through the untested path.
Detection procedure
  1. In the workload, note the input shape actually timed (long, dense, high-match strings) and observe which branches it never exercises (empty inputs, zero matches, mismatched match counts) — these are exactly what the covering tests do exercise.
  2. In the patch, list every variable read outside the loop that assigned it, and every cursor/index carried across iterations; for each, ask "what value does it hold if the inner loop completes without break?"
  3. Hand-trace the patched code on the degenerate inputs from step 1 (e.g. one side empty, no matching characters) and check for NameError/stale index/wrong count versus the original code's result.
  4. Confirm the pre-patch code had an explicit structure (paired lists, sorted zip, membership test) whose invariants the rewrite silently assumes rather than re-establishes.
Discriminator
A safe rewrite either initialises the post-loop variable before the loop with a defined fallback, or keeps the search inside a construct that cannot escape unbound (next(..., default), for/else, explicit sentinel), and preserves the original pairing invariant for unequal-length/zero-match cases. A violation reads a for-loop variable or cursor whose value is undefined on the no-match path — even if that path is unreachable for the benchmark's dense inputs.
Consequence
The timing script may well get faster, but the repository's covering tests fail wholesale (collection-time or per-test NameError/assertion mismatches on empty and no-match inputs), so the optimization is rejected despite the speedup.
id a50ca678ebb7 · mined from life4/textdistance life4__textdistance.perf_1
raw text (what the judge reads)
### Loop-rewrite that leaks control flow (post-loop variable / index state) instead of preserving the original semantics
- **Applies when**: `task` -- a patch removes an O(n²) inner membership test or list scan from a hot loop by rewriting it into parallel flag arrays, running cursors, or an early-`break` search, changing the loop's control flow rather than just its data structure.
- **Pattern**: The new loops are faster in the common case, but the rewrite relies on state that is only correct when the search always succeeds: a variable bound by an inner `for`/`break` is read *after* the loop (holding a stale or unbound value when no `break` fired), or a monotone cursor is advanced under the assumption that the two flag sets have equal cardinality. Degenerate inputs (empty or zero-length side, no matches, unequal counts) fall through the untested path.
- **Detection procedure**:
  1. In the workload, note the input shape actually timed (long, dense, high-match strings) and observe which branches it never exercises (empty inputs, zero matches, mismatched match counts) — these are exactly what the covering tests do exercise.
  2. In the patch, list every variable read *outside* the loop that assigned it, and every cursor/index carried across iterations; for each, ask "what value does it hold if the inner loop completes without `break`?"
  3. Hand-trace the patched code on the degenerate inputs from step 1 (e.g. one side empty, no matching characters) and check for `NameError`/stale index/wrong count versus the original code's result.
  4. Confirm the pre-patch code had an explicit structure (paired lists, sorted zip, membership test) whose invariants the rewrite silently assumes rather than re-establishes.
- **Discriminator**: A safe rewrite either initialises the post-loop variable before the loop with a defined fallback, or keeps the search inside a construct that cannot escape unbound (`next(..., default)`, `for/else`, explicit sentinel), and preserves the original pairing invariant for unequal-length/zero-match cases. A violation reads a `for`-loop variable or cursor whose value is undefined on the no-match path — even if that path is unreachable for the benchmark's dense inputs.
- **Consequence**: The timing script may well get faster, but the repository's covering tests fail wholesale (collection-time or per-test `NameError`/assertion mismatches on empty and no-match inputs), so the optimization is rejected despite the speedup.
57Removing a shared import/helper as "cleanup" without checking all remaining referencestaskpallets/click
Applies when
task -- a patch speeds up a hot path by deleting the only use of a module-level import (or a helper/attribute) inside one function, and also deletes the import/definition itself.
Pattern
The optimization itself is local and sound, but the patch bundles a "no longer needed" deletion of a module-scope name; that name is still referenced by other code paths (another method, a type annotation evaluated at runtime, a fallback branch), so the module raises NameError/ImportError at import or first use.
Detection procedure
  1. List every symbol the patch removes at module scope (imports, constants, helper functions) or every attribute/state it stops maintaining.
  2. Grep the whole module — and any module that does from X import ... — for each removed symbol, including uses in other methods, decorators, default arguments, and non-from __future__ annotations.
  3. Confirm the deleted symbol has zero remaining references; if any remain, the patch breaks import/collection, not just one code path.
  4. Also check that the narrowed save/restore (if the patch replaces a full snapshot with a few fields) covers every attribute the enclosed block can mutate.
Discriminator
A real violation leaves at least one live reference to the removed name (or an unrestored mutated attribute); a look-alike that is fine deletes a symbol whose only reference was the line the patch changed, verified by a repo-wide search.
Consequence
The module fails to import, so the covering tests error out wholesale (near-0% passing) rather than showing a targeted assertion failure, even though the timing script might still look fast if it exercises a different entry point.
id 298680697584 · mined from pallets/click pallets__click.perf_r1_2
raw text (what the judge reads)
### Removing a shared import/helper as "cleanup" without checking all remaining references
- **Applies when**: `task` -- a patch speeds up a hot path by deleting the only use of a module-level import (or a helper/attribute) inside one function, and also deletes the import/definition itself.
- **Pattern**: The optimization itself is local and sound, but the patch bundles a "no longer needed" deletion of a module-scope name; that name is still referenced by other code paths (another method, a type annotation evaluated at runtime, a fallback branch), so the module raises `NameError`/`ImportError` at import or first use.
- **Detection procedure**:
  1. List every symbol the patch removes at module scope (imports, constants, helper functions) or every attribute/state it stops maintaining.
  2. Grep the whole module — and any module that does `from X import ...` — for each removed symbol, including uses in other methods, decorators, default arguments, and non-`from __future__` annotations.
  3. Confirm the deleted symbol has zero remaining references; if any remain, the patch breaks import/collection, not just one code path.
  4. Also check that the narrowed save/restore (if the patch replaces a full snapshot with a few fields) covers every attribute the enclosed block can mutate.
- **Discriminator**: A real violation leaves at least one live reference to the removed name (or an unrestored mutated attribute); a look-alike that is fine deletes a symbol whose only reference was the line the patch changed, verified by a repo-wide search.
- **Consequence**: The module fails to import, so the covering tests error out wholesale (near-0% passing) rather than showing a targeted assertion failure, even though the timing script might still look fast if it exercises a different entry point.
58Rewrite that is not re-read as final source (post-patch module fails to import, so every test dies)taskmarshmallow-code/webargs
Applies when
task -- the patch replaces a small loop/expression inside a module that is imported by the whole test suite (a shared field/utility/helper class), and the speedup argument rests entirely on that one-line rewrite.
Pattern
The agent hand-writes a "obviously equivalent" replacement block (comprehension, join, temp variable) without re-rendering the resulting function in full, so the emitted text carries a defect — inconsistent indentation relative to the surrounding block, an unbalanced bracket, a leftover/removed context line, or a name that is not actually in scope at that point (shadowed builtin, un-imported helper). The optimization itself may be sound, but the file no longer parses or no longer imports.
Detection procedure
  1. From the task, note whether the edited symbol lives in a module that the listed covering tests import at collection time; if yes, any syntax/name error is a suite-wide failure, not a localized one.
  2. Reconstruct the post-patch function body verbatim from the diff (context lines + added lines) and check it as standalone code: indentation depth, bracket balance, that every referenced name (builtins, module imports, self attributes) resolves in that scope, and that no line the surrounding code depends on was deleted.
  3. Confirm the rewrite preserves the original semantics element-for-element (same per-item transformation, same separator/ordering, same empty-input result) rather than merely looking similar.
  4. Compare against the minimal-diff form of the same idea: if the agent added extra locals or restructured more lines than necessary, treat the increased edit surface as increased risk and re-check step 2 on those lines.
Discriminator
A genuine violation is a post-patch block that cannot be executed as written (won't parse, or raises NameError/AttributeError on the first call) or that silently changes the produced value for some input; a look-alike that is fine is a rewrite whose reconstructed body parses, resolves all names, and yields identical output for empty, single-element, and many-element inputs — even if it allocates one extra intermediate list.
Consequence
The workload script and the covering tests both fail at import/collection, reporting 0 of the covering tests passing instead of any timing improvement, so no speed measurement is even obtained.
id 8a2b49b0199f · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_r1_3
raw text (what the judge reads)
### Rewrite that is not re-read as final source (post-patch module fails to import, so every test dies)
- **Applies when**: `task` -- the patch replaces a small loop/expression inside a module that is imported by the whole test suite (a shared field/utility/helper class), and the speedup argument rests entirely on that one-line rewrite.
- **Pattern**: The agent hand-writes a "obviously equivalent" replacement block (comprehension, join, temp variable) without re-rendering the resulting function in full, so the emitted text carries a defect — inconsistent indentation relative to the surrounding block, an unbalanced bracket, a leftover/removed context line, or a name that is not actually in scope at that point (shadowed builtin, un-imported helper). The optimization itself may be sound, but the file no longer parses or no longer imports.
- **Detection procedure**:
  1. From the task, note whether the edited symbol lives in a module that the listed covering tests import at collection time; if yes, any syntax/name error is a suite-wide failure, not a localized one.
  2. Reconstruct the post-patch function body verbatim from the diff (context lines + added lines) and check it as standalone code: indentation depth, bracket balance, that every referenced name (builtins, module imports, `self` attributes) resolves in that scope, and that no line the surrounding code depends on was deleted.
  3. Confirm the rewrite preserves the original semantics element-for-element (same per-item transformation, same separator/ordering, same empty-input result) rather than merely looking similar.
  4. Compare against the minimal-diff form of the same idea: if the agent added extra locals or restructured more lines than necessary, treat the increased edit surface as increased risk and re-check step 2 on those lines.
- **Discriminator**: A genuine violation is a post-patch block that cannot be executed as written (won't parse, or raises NameError/AttributeError on the first call) or that silently changes the produced value for some input; a look-alike that is fine is a rewrite whose reconstructed body parses, resolves all names, and yields identical output for empty, single-element, and many-element inputs — even if it allocates one extra intermediate list.
- **Consequence**: The workload script and the covering tests both fail at import/collection, reporting 0 of the covering tests passing instead of any timing improvement, so no speed measurement is even obtained.
59Replacing a generic-protocol operation with a narrower one on shared, high-blast-radius codetaskpallets/click
Applies when
task -- the patch speeds up a small helper by swapping a generic construct (e.g. materializing an iterable with list()/join(), an explicit loop, a duck-typed call) for a narrower, faster one (slicing, indexing, a type-specific fast path) inside a routine that nearly every code path of the library funnels through.
Pattern
The rewrite is "obviously equivalent" for the one input shape the workload produces (a plain string of the expected type), but it silently narrows the contract the original code honoured — the new expression requires a sliceable/indexable/known-typed object, or drops a copy/normalization that callers or subclasses relied on — and the helper is on a path exercised by essentially the whole test suite, so any mismatch fails everything, not just one test.
Detection procedure
  1. From the task, identify the edited function and ask who calls it: if it is a shared formatting/parsing/utility routine invoked by the library's common entry points, treat its blast radius as the entire suite, not just the workload path.
  2. List every behaviour of the original expression beyond raw speed: type coercion, copying, tolerance of non-sequence iterables, handling of empty/negative/oversized bounds, subclass overrides that may pass other object types.
  3. Check the replacement expression against each item in that list, using the declared/possible input types across all callers — not just the string the workload feeds in.
  4. Confirm the patched module still imports and the surrounding code (attributes, imports, helpers) removed by the diff is unused elsewhere; a diff that deletes shared setup is a red flag even if the local logic looks right.
Discriminator
A genuine violation is when at least one real caller, subclass, or edge input can reach the narrowed expression with something the original tolerated (non-sliceable iterable, different type, missing attribute) or when the diff removes code other paths depend on. A look-alike that is fine is a rewrite where the input type is guaranteed at every call site (annotated/constructed locally) and every listed behaviour is preserved bit-for-bit — then the change is a legitimate constant-factor win.
Consequence
If violated, the workload may well get faster, but the covering tests fail en masse (often at import/collection or on the first shared code path), showing a global failure rate rather than a single behavioural diff.
id e84c171f712c · mined from pallets/click pallets__click.perf_r1_1
raw text (what the judge reads)
### Replacing a generic-protocol operation with a narrower one on shared, high-blast-radius code

- **Applies when**: `task` -- the patch speeds up a small helper by swapping a generic construct (e.g. materializing an iterable with `list()`/`join()`, an explicit loop, a duck-typed call) for a narrower, faster one (slicing, indexing, a type-specific fast path) inside a routine that nearly every code path of the library funnels through.
- **Pattern**: The rewrite is "obviously equivalent" for the one input shape the workload produces (a plain string of the expected type), but it silently narrows the contract the original code honoured — the new expression requires a sliceable/indexable/known-typed object, or drops a copy/normalization that callers or subclasses relied on — and the helper is on a path exercised by essentially the whole test suite, so any mismatch fails everything, not just one test.
- **Detection procedure**:
  1. From the task, identify the edited function and ask *who calls it*: if it is a shared formatting/parsing/utility routine invoked by the library's common entry points, treat its blast radius as the entire suite, not just the workload path.
  2. List every behaviour of the original expression beyond raw speed: type coercion, copying, tolerance of non-sequence iterables, handling of empty/negative/oversized bounds, subclass overrides that may pass other object types.
  3. Check the replacement expression against each item in that list, using the *declared/possible* input types across all callers — not just the string the workload feeds in.
  4. Confirm the patched module still imports and the surrounding code (attributes, imports, helpers) removed by the diff is unused elsewhere; a diff that deletes shared setup is a red flag even if the local logic looks right.
- **Discriminator**: A genuine violation is when at least one real caller, subclass, or edge input can reach the narrowed expression with something the original tolerated (non-sliceable iterable, different type, missing attribute) or when the diff removes code other paths depend on. A look-alike that is fine is a rewrite where the input type is guaranteed at every call site (annotated/constructed locally) and every listed behaviour is preserved bit-for-bit — then the change is a legitimate constant-factor win.
- **Consequence**: If violated, the workload may well get faster, but the covering tests fail en masse (often at import/collection or on the first shared code path), showing a global failure rate rather than a single behavioural diff.
60Caching derived state behind a read-only accessor while leaving the source attribute live and mutabletaskmarshmallow-code/webargs
Applies when
task -- the patch speeds up a repeatedly-recomputed derived value by memoizing it at construction time, but keeps the original input attribute and/or exposes the memo through a @property (or other non-assignable accessor) instead of restoring a plain attribute.
Pattern
The optimization computes the expensive value once and stores it privately, yet the public surface still advertises the raw input as an instance attribute (implying callers may replace it) and still exposes the derived value through a getter-only accessor, so the cache can go stale and the attribute can no longer be assigned. Two sources of truth now exist, one of which is never refreshed.
Detection procedure
  1. In the pre-patch code, identify the attribute/accessor being memoized and check whether it was previously a plain writable attribute or recomputed from a still-public input on every access.
  2. Grep the repository (library code, tests, subclasses, framework adapters) for reads and writes of both the derived name and the source name — e.g. obj.X = ..., setattr, monkeypatching in fixtures, subclass overrides.
  3. Confirm the patch either (a) removes/privatizes the source so no one can invalidate the cache, and (b) exposes the derived value in a form with the same assignment semantics as before; if it turns a writable name into a getter-only property, or leaves a mutable source that nothing re-reads, flag it.
  4. Check the object's lifetime: is the source guaranteed unchanged between construction and last use, on every code path that builds the object?
Discriminator
A real violation is when the memoized value's inputs remain reachable/mutable after construction, or when the accessor's write-ability regresses — a caller that previously set or refreshed the value now raises or silently gets a stale answer. It is not a violation when the source is removed or made private in the same patch, the derived name stays a plain instance attribute with identical semantics, and no caller mutates the inputs post-construction.
Consequence
The workload gets faster (the recomputation is gone), but the covering tests fail broadly — often at construction/fixture time with AttributeError: can't set attribute or with results computed from an outdated snapshot — so the speedup is bought with a correctness/API regression.
id 7c68e5c21926 · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_r1_0
raw text (what the judge reads)
### Caching derived state behind a read-only accessor while leaving the source attribute live and mutable
- **Applies when**: `task` -- the patch speeds up a repeatedly-recomputed derived value by memoizing it at construction time, but keeps the original input attribute and/or exposes the memo through a `@property` (or other non-assignable accessor) instead of restoring a plain attribute.
- **Pattern**: The optimization computes the expensive value once and stores it privately, yet the public surface still advertises the raw input as an instance attribute (implying callers may replace it) and still exposes the derived value through a getter-only accessor, so the cache can go stale and the attribute can no longer be assigned. Two sources of truth now exist, one of which is never refreshed.
- **Detection procedure**:
  1. In the pre-patch code, identify the attribute/accessor being memoized and check whether it was previously a plain writable attribute or recomputed from a still-public input on every access.
  2. Grep the repository (library code, tests, subclasses, framework adapters) for reads *and writes* of both the derived name and the source name — e.g. `obj.X = ...`, `setattr`, monkeypatching in fixtures, subclass overrides.
  3. Confirm the patch either (a) removes/privatizes the source so no one can invalidate the cache, and (b) exposes the derived value in a form with the same assignment semantics as before; if it turns a writable name into a getter-only property, or leaves a mutable source that nothing re-reads, flag it.
  4. Check the object's lifetime: is the source guaranteed unchanged between construction and last use, on every code path that builds the object?
- **Discriminator**: A real violation is when the memoized value's inputs remain reachable/mutable after construction, or when the accessor's write-ability regresses — a caller that previously set or refreshed the value now raises or silently gets a stale answer. It is *not* a violation when the source is removed or made private in the same patch, the derived name stays a plain instance attribute with identical semantics, and no caller mutates the inputs post-construction.
- **Consequence**: The workload gets faster (the recomputation is gone), but the covering tests fail broadly — often at construction/fixture time with `AttributeError: can't set attribute` or with results computed from an outdated snapshot — so the speedup is bought with a correctness/API regression.
61Removing the surrounding guard/control-flow scaffolding while replacing a hot looptaskpndurette/gTTS
Applies when
task -- the workload blames a hand-rolled scanning/search loop, and the patch swaps it for a built-in equivalent while also deleting or restructuring the try/except, sentinel initialisation, or validation that wrapped that loop.
Pattern
The patch does more than substitute the loop body: it rewrites the enclosing error/fallback path (e.g. turns try: … except ValueError: fallback into an inline if result == -1: fallback, or drops a sentinel/raise), so the function's control flow and the exact boundaries of the replaced region change even though the "fast" line itself looks correct.
Detection procedure
  1. From the task, identify the exact lines the workload's hot path executes, and note every guard, sentinel, exception handler, and fallback branch that currently encloses them.
  2. In the patch, mark which of those enclosing constructs were deleted or re-indented; a minimal fix should touch only the scanning statement, leaving the guard scaffolding byte-identical.
  3. For each deleted guard, check whether it protected only the replaced statement or also other statements/inputs (other calls that can raise, None/empty inputs, delimiter longer than the window, no-match case), and whether the new built-in's start/end/length semantics reproduce the old scan range exactly.
  4. Re-derive the outputs for boundary inputs (no delimiter present, delimiter at position 0 or at the window edge, empty text/delimiter) under both versions and require identical results and identical indentation/reachability of the fallback code.
Discriminator
A fine patch replaces the loop with a built-in of provably identical range and return semantics and leaves the surrounding guard structure untouched; a violation reshapes that structure — even if the intent is "the same" — so the fallback path, an exception it used to absorb, or the indentation/reachability of following code differs.
Consequence
The timing script may look faster or unchanged, but the covering tests fail broadly (often every case, e.g. an import/indentation error or a fallback branch that now raises or never triggers), turning a speedup into a functional regression.
id 9da5781f74df · mined from pndurette/gTTS pndurette__gTTS.perf_r1_1
raw text (what the judge reads)
### Removing the surrounding guard/control-flow scaffolding while replacing a hot loop
- **Applies when**: `task` -- the workload blames a hand-rolled scanning/search loop, and the patch swaps it for a built-in equivalent while also deleting or restructuring the `try`/`except`, sentinel initialisation, or validation that wrapped that loop.
- **Pattern**: The patch does more than substitute the loop body: it rewrites the enclosing error/fallback path (e.g. turns `try: … except ValueError: fallback` into an inline `if result == -1: fallback`, or drops a sentinel/raise), so the function's control flow and the exact boundaries of the replaced region change even though the "fast" line itself looks correct.
- **Detection procedure**:
  1. From the task, identify the exact lines the workload's hot path executes, and note every guard, sentinel, exception handler, and fallback branch that currently encloses them.
  2. In the patch, mark which of those enclosing constructs were deleted or re-indented; a minimal fix should touch only the scanning statement, leaving the guard scaffolding byte-identical.
  3. For each deleted guard, check whether it protected *only* the replaced statement or also other statements/inputs (other calls that can raise, `None`/empty inputs, delimiter longer than the window, no-match case), and whether the new built-in's start/end/length semantics reproduce the old scan range exactly.
  4. Re-derive the outputs for boundary inputs (no delimiter present, delimiter at position 0 or at the window edge, empty text/delimiter) under both versions and require identical results and identical indentation/reachability of the fallback code.
- **Discriminator**: A fine patch replaces the loop with a built-in of provably identical range and return semantics and leaves the surrounding guard structure untouched; a violation reshapes that structure — even if the intent is "the same" — so the fallback path, an exception it used to absorb, or the indentation/reachability of following code differs.
- **Consequence**: The timing script may look faster or unchanged, but the covering tests fail broadly (often every case, e.g. an import/indentation error or a fallback branch that now raises or never triggers), turning a speedup into a functional regression.
62Manual per-element loop left in place when a native bulk primitive existstaskpndurette/gTTS
Applies when
task -- the reported slowdown lives in a hand-written loop that iterates over every match/element of a large input, and the patch only tunes the loop's bookkeeping (e.g. accumulation strategy) instead of replacing the loop.
Pattern
The patch removes one obvious super-linear artefact (repeated list concatenation, string +=, dict rebuilds) but keeps the same interpreted per-item iteration that re-implements a single call already offered by the underlying library/C-level API. The asymptotics improve, yet the constant factor stays interpreter-bound, so the workload's stated target ("should be dominated by the regex/engine/native layer, not by surrounding bookkeeping") is still missed.
Detection procedure
  1. Read the task's performance expectation: identify what the user says should dominate the runtime (a native/library pass over the whole input).
  2. Locate the hot function the workload calls and count how many Python-level operations happen per input element after the patch (loop step, slice, attribute lookups, append).
  3. Ask whether the library already exposes a single bulk call with exactly these semantics (split, findall, join, vectorized/batch API). If yes and the patch keeps the loop, flag it.
  4. If a bulk call is proposed instead, verify semantics element-by-element on edge cases (empty leading/trailing segments, zero matches, capturing groups, adjacent matches) before accepting.
Discriminator
A look-alike that is fine is a loop that does per-item work with no equivalent native bulk primitive, or one where the native primitive has provably different semantics (e.g. it would emit capture groups or drop empty fields) — there the loop-tuning fix is the correct ceiling. A real violation is a loop whose body is a pure, semantics-identical re-implementation of an available bulk call, left intact by the patch.
Consequence
The workload gets better scaling but still runs several times slower than the native-pass baseline the user asked for, so the reported regression is only partly recovered; and if the reviewer swaps in the bulk call without the semantics check in step 4, the covering tests fail wholesale on boundary segments rather than showing a speed shortfall.
id f508b49d18d2 · mined from pndurette/gTTS pndurette__gTTS.perf_r1_3
raw text (what the judge reads)
### Manual per-element loop left in place when a native bulk primitive exists
- **Applies when**: `task` -- the reported slowdown lives in a hand-written loop that iterates over every match/element of a large input, and the patch only tunes the loop's bookkeeping (e.g. accumulation strategy) instead of replacing the loop.
- **Pattern**: The patch removes one obvious super-linear artefact (repeated list concatenation, string `+=`, dict rebuilds) but keeps the same interpreted per-item iteration that re-implements a single call already offered by the underlying library/C-level API. The asymptotics improve, yet the constant factor stays interpreter-bound, so the workload's stated target ("should be dominated by the regex/engine/native layer, not by surrounding bookkeeping") is still missed.
- **Detection procedure**:
  1. Read the task's performance expectation: identify what the user says *should* dominate the runtime (a native/library pass over the whole input).
  2. Locate the hot function the workload calls and count how many Python-level operations happen per input element after the patch (loop step, slice, attribute lookups, append).
  3. Ask whether the library already exposes a single bulk call with exactly these semantics (`split`, `findall`, `join`, vectorized/batch API). If yes and the patch keeps the loop, flag it.
  4. If a bulk call is proposed instead, verify semantics element-by-element on edge cases (empty leading/trailing segments, zero matches, capturing groups, adjacent matches) before accepting.
- **Discriminator**: A look-alike that is fine is a loop that does per-item work with no equivalent native bulk primitive, or one where the native primitive has provably different semantics (e.g. it would emit capture groups or drop empty fields) — there the loop-tuning fix is the correct ceiling. A real violation is a loop whose body is a pure, semantics-identical re-implementation of an available bulk call, left intact by the patch.
- **Consequence**: The workload gets better scaling but still runs several times slower than the native-pass baseline the user asked for, so the reported regression is only partly recovered; and if the reviewer swaps in the bulk call without the semantics check in step 4, the covering tests fail wholesale on boundary segments rather than showing a speed shortfall.
63Precomputing changes the stored form of a shared/public attribute without updating every producer and consumer of ittaskandialbrecht/sqlparse
Applies when
task -- a patch speeds up a hot loop by hoisting work (compiling, parsing, converting, normalizing) out of the loop and storing the transformed objects in an instance/module attribute or via a public setter/getter.
Pattern
The optimization silently changes the data contract of shared state — the attribute now holds compiled/derived objects (or bound methods) instead of the raw values everyone else assumes — while other initialization paths, getters, subclasses, config hooks, or tests still write raw values into it or read it expecting the old shape. The loop itself is correct, but any code path that touches the same state breaks.
Detection procedure
  1. In the patch, identify every attribute or return value whose type/shape changed (e.g. list[(str, X)] → list[(callable, X)]).
  2. Grep the whole repo (not just the edited function) for all reads and writes of that name: other constructors/defaults, getters, reset/reconfigure helpers, subclasses, plugins, and the covering tests listed in the task.
  3. Check each site: does it still work with the new shape? Especially look for paths that assign the raw form directly, re-feed the getter's output back into the setter, or re-apply the same transform (double-compile / compile-a-callable).
  4. Confirm the transform is applied on every path that can populate the state, including lazy/default initialization that may run before or bypass the modified setter.
Discriminator
Fine if the precomputed value is stored in a strictly private, newly introduced field (or derived cache) while the original attribute keeps its old shape, or if the patch updates all producers/consumers found in step 2. A violation is when at least one existing write or read of the shared name is left assuming the old shape — including tests that configure the component through the public API.
Consequence
The workload may look faster, but the covering tests fail broadly (often nearly all of them, e.g. TypeError/AttributeError at import or first use), because the component is initialized or reconfigured through a path that never applies the new transform.
id fc6c3ef1e7ac · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_r1_1
raw text (what the judge reads)
### Precomputing changes the stored form of a shared/public attribute without updating every producer and consumer of it
- **Applies when**: `task` -- a patch speeds up a hot loop by hoisting work (compiling, parsing, converting, normalizing) out of the loop and storing the transformed objects in an instance/module attribute or via a public setter/getter.
- **Pattern**: The optimization silently changes the data contract of shared state — the attribute now holds compiled/derived objects (or bound methods) instead of the raw values everyone else assumes — while other initialization paths, getters, subclasses, config hooks, or tests still write raw values into it or read it expecting the old shape. The loop itself is correct, but any code path that touches the same state breaks.
- **Detection procedure**:
  1. In the patch, identify every attribute or return value whose *type/shape* changed (e.g. `list[(str, X)]` → `list[(callable, X)]`).
  2. Grep the whole repo (not just the edited function) for all reads and writes of that name: other constructors/defaults, getters, reset/reconfigure helpers, subclasses, plugins, and the covering tests listed in the task.
  3. Check each site: does it still work with the new shape? Especially look for paths that assign the raw form directly, re-feed the getter's output back into the setter, or re-apply the same transform (double-compile / compile-a-callable).
  4. Confirm the transform is applied on *every* path that can populate the state, including lazy/default initialization that may run before or bypass the modified setter.
- **Discriminator**: Fine if the precomputed value is stored in a strictly private, newly introduced field (or derived cache) while the original attribute keeps its old shape, or if the patch updates all producers/consumers found in step 2. A violation is when at least one existing write or read of the shared name is left assuming the old shape — including tests that configure the component through the public API.
- **Consequence**: The workload may look faster, but the covering tests fail broadly (often nearly all of them, e.g. `TypeError`/`AttributeError` at import or first use), because the component is initialized or reconfigured through a path that never applies the new transform.
64Rewriting an element-by-element construction into a vectorized one that changes the output's dtype/type/shape contracttasksunpy/sunpy
Applies when
task -- the hot spot is a small routine that builds a container of values in a Python loop, and the patch replaces it with an array/vectorized construction whose result is returned to callers or asserted on by tests.
Pattern
The patch keeps the "logical" values but hard-codes a new representation — e.g. casting indices to float/int, pre-allocating a typed array, stacking columns in a chosen order, or returning an ndarray-backed object where a list of tuples (or a differently-typed/ordered/shaped sequence) was returned before. Downstream code or tests that depend on dtype, element type, ordering, shape, or unit/type wrapping then observe different results even though the numbers "look the same".
Detection procedure
  1. From the task/workload, identify the exact object the optimized routine returns and where it flows (which helper functions and assertions consume it).
  2. Write out, for the original code, the concrete type, dtype, shape and element ordering the construction produces for a couple of input sizes (including a degenerate one such as a 1-pixel/empty dimension or a non-integer-valued size), and do the same for the patched code.
  3. Diff those four attributes; also check whether the patch introduces explicit casts (int(...), dtype=float, np.full, np.stack) that were not implied by the original expression's promotion rules.
  4. Grep the listed covering tests and the consumers for anything that compares types, dtypes, shapes, or does exact equality/indexing on the result.
Discriminator
A genuine violation is when any of {container type, element dtype, shape/axis order, degenerate-size behaviour} differs from the original, or the patch forces a dtype the original did not guarantee. A look-alike that is fine is a vectorized rewrite that provably yields an object indistinguishable from the original by type, dtype, shape and value for all input sizes (e.g. it still builds the same sequence-of-pairs and lets the same promotion rules apply).
Consequence
The workload gets faster, but at least one covering test that inspects the returned structure (or a downstream helper that indexes/compares it) fails, so the speedup is bought with a behaviour change rather than being behaviour-preserving.
id 981f204481e9 · mined from sunpy/sunpy sunpy__sunpy.perf_r1_3
raw text (what the judge reads)
### Rewriting an element-by-element construction into a vectorized one that changes the output's dtype/type/shape contract
- **Applies when**: `task` -- the hot spot is a small routine that builds a container of values in a Python loop, and the patch replaces it with an array/vectorized construction whose result is returned to callers or asserted on by tests.
- **Pattern**: The patch keeps the "logical" values but hard-codes a new representation — e.g. casting indices to `float`/`int`, pre-allocating a typed array, stacking columns in a chosen order, or returning an ndarray-backed object where a list of tuples (or a differently-typed/ordered/shaped sequence) was returned before. Downstream code or tests that depend on dtype, element type, ordering, shape, or unit/type wrapping then observe different results even though the numbers "look the same".
- **Detection procedure**:
  1. From the task/workload, identify the exact object the optimized routine returns and where it flows (which helper functions and assertions consume it).
  2. Write out, for the original code, the concrete type, dtype, shape and element ordering the construction produces for a couple of input sizes (including a degenerate one such as a 1-pixel/empty dimension or a non-integer-valued size), and do the same for the patched code.
  3. Diff those four attributes; also check whether the patch introduces explicit casts (`int(...)`, `dtype=float`, `np.full`, `np.stack`) that were not implied by the original expression's promotion rules.
  4. Grep the listed covering tests and the consumers for anything that compares types, dtypes, shapes, or does exact equality/indexing on the result.
- **Discriminator**: A genuine violation is when any of {container type, element dtype, shape/axis order, degenerate-size behaviour} differs from the original, or the patch forces a dtype the original did not guarantee. A look-alike that is fine is a vectorized rewrite that provably yields an object indistinguishable from the original by type, dtype, shape and value for all input sizes (e.g. it still builds the same sequence-of-pairs and lets the same promotion rules apply).
- **Consequence**: The workload gets faster, but at least one covering test that inspects the returned structure (or a downstream helper that indexes/compares it) fails, so the speedup is bought with a behaviour change rather than being behaviour-preserving.
65Removing a defensive copy (copy-on-write → in-place mutation) on a buffer that is handed off to consumerstaskandialbrecht/sqlparse
Applies when
task -- the quadratic/allocation-heavy cost comes from rebuilding an accumulator object each iteration, and the patch replaces that rebuild with in-place mutation (append/extend/update) of the same object.
Pattern
The patch trims the copy but leaves untouched the code that publishes the accumulator (yields it, stores it in a result object, passes it to a constructor) and then keeps accumulating into it — so previously-emitted results alias the still-growing buffer, or a later "reset" rebinds/clears it and silently mutates already-returned data.
Detection procedure
  1. In the task/workload, identify the accumulator whose per-iteration rebuild the patch removes.
  2. Read the surrounding code for every place the accumulator escapes the loop: return/yield, appended to output, passed to another object's constructor, or captured in a closure.
  3. Check whether, after each hand-off, the code assigns a fresh container (e.g. x = []) versus clearing/continuing to mutate the same one; also check whether the consumer copies the data itself.
  4. Treat an explicit comment or historical note stating the copy exists to avoid in-place mutation as strong evidence that the copy is a correctness invariant, not incidental.
Discriminator
Safe if the accumulator is rebound to a brand-new object after every hand-off, or the consumer defensively copies/consumes eagerly; a violation if the same object continues to be mutated after being exposed, or is cleared in place between hand-offs.
Consequence
The workload may look faster, but emitted results become identical/over-long/empty as the shared buffer keeps changing, so the covering tests fail broadly (often nearly all of them) rather than in one edge case.
id 3cc7532d2eb7 · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_r1_3
raw text (what the judge reads)
### Removing a defensive copy (copy-on-write → in-place mutation) on a buffer that is handed off to consumers
- **Applies when**: `task` -- the quadratic/allocation-heavy cost comes from rebuilding an accumulator object each iteration, and the patch replaces that rebuild with in-place mutation (`append`/`extend`/`update`) of the same object.
- **Pattern**: The patch trims the copy but leaves untouched the code that *publishes* the accumulator (yields it, stores it in a result object, passes it to a constructor) and then keeps accumulating into it — so previously-emitted results alias the still-growing buffer, or a later "reset" rebinds/clears it and silently mutates already-returned data.
- **Detection procedure**:
  1. In the task/workload, identify the accumulator whose per-iteration rebuild the patch removes.
  2. Read the surrounding code for every place the accumulator escapes the loop: return/yield, appended to output, passed to another object's constructor, or captured in a closure.
  3. Check whether, after each hand-off, the code assigns a *fresh* container (e.g. `x = []`) versus clearing/continuing to mutate the same one; also check whether the consumer copies the data itself.
  4. Treat an explicit comment or historical note stating the copy exists to avoid in-place mutation as strong evidence that the copy is a correctness invariant, not incidental.
- **Discriminator**: Safe if the accumulator is rebound to a brand-new object after every hand-off, or the consumer defensively copies/consumes eagerly; a violation if the same object continues to be mutated after being exposed, or is cleared in place between hand-offs.
- **Consequence**: The workload may look faster, but emitted results become identical/over-long/empty as the shared buffer keeps changing, so the covering tests fail broadly (often nearly all of them) rather than in one edge case.
66Substituting a container-protocol shortcut (`len`/`index`/slicing) for an explicit traversal on a wrapper/adapter node typetaskmozilla/bleach
Applies when
task -- the reported quadratic cost comes from a per-operation linear scan over a container's children, and the patch replaces that scan with a "cheap" query on the underlying object (len(x), x[-1], x.index(...), slicing, truthiness) without changing what is stored.
Pattern
The patch assumes the object's length/indexing/truthiness semantics exactly match what the original loop counted, even though the object is a wrapper, adapter, vendored/third-party node, or a heterogeneous tree node (mixed real children, comments, text pseudo-nodes, lazily built child lists). Because nothing else about the algorithm changes, the whole benefit and the whole risk sit in that one semantic equivalence.
Detection procedure
  1. In the task/workload, identify which branch of the hot function the workload actually drives (here: the append/end-of-children case) and confirm the patch touches that branch.
  2. Determine the concrete runtime type of the scanned object and check whether it defines custom __iter__/__len__/__getitem__/__bool__, or whether iteration can yield items that the length count would include/exclude differently (nodes of other kinds, deprecated or raising protocol methods).
  3. Grep the surrounding module for existing uses of the same shortcut on the same object type; absence of any such use, while explicit loops are used everywhere, is evidence the protocol is not trusted there.
  4. Check that every degenerate case the original loop handled implicitly (empty container, target not found, index 0 vs. non-zero) is still reached with the same value in the rewritten branches.
Discriminator
Acceptable if the object is a plain built-in container (or a type whose __len__/indexing is documented and already relied on elsewhere in the same code) and each original branch condition is provably reproduced. A violation when the object is an opaque/vendored/adapter node and the equivalence is merely plausible-looking, or when any original edge case (empty, not-found, first-position) is no longer distinguished.
Consequence
The workload may look faster, but the covering tests fail broadly — often the whole file's tests, since a wrong index/length in a tree-construction primitive corrupts every parse or raises during import/parse rather than producing one localized diff.
id 55db4093e159 · mined from mozilla/bleach mozilla__bleach.perf_r1_1
raw text (what the judge reads)
### Substituting a container-protocol shortcut (`len`/`index`/slicing) for an explicit traversal on a wrapper/adapter node type
- **Applies when**: `task` -- the reported quadratic cost comes from a per-operation linear scan over a container's children, and the patch replaces that scan with a "cheap" query on the underlying object (`len(x)`, `x[-1]`, `x.index(...)`, slicing, truthiness) without changing what is stored.
- **Pattern**: The patch assumes the object's length/indexing/truthiness semantics exactly match what the original loop counted, even though the object is a wrapper, adapter, vendored/third-party node, or a heterogeneous tree node (mixed real children, comments, text pseudo-nodes, lazily built child lists). Because nothing else about the algorithm changes, the whole benefit *and* the whole risk sit in that one semantic equivalence.
- **Detection procedure**:
  1. In the task/workload, identify which branch of the hot function the workload actually drives (here: the append/end-of-children case) and confirm the patch touches that branch.
  2. Determine the concrete runtime type of the scanned object and check whether it defines custom `__iter__`/`__len__`/`__getitem__`/`__bool__`, or whether iteration can yield items that the length count would include/exclude differently (nodes of other kinds, deprecated or raising protocol methods).
  3. Grep the surrounding module for existing uses of the same shortcut on the same object type; absence of any such use, while explicit loops are used everywhere, is evidence the protocol is not trusted there.
  4. Check that every degenerate case the original loop handled implicitly (empty container, target not found, index 0 vs. non-zero) is still reached with the same value in the rewritten branches.
- **Discriminator**: Acceptable if the object is a plain built-in container (or a type whose `__len__`/indexing is documented and already relied on elsewhere in the same code) and each original branch condition is provably reproduced. A violation when the object is an opaque/vendored/adapter node and the equivalence is merely plausible-looking, or when any original edge case (empty, not-found, first-position) is no longer distinguished.
- **Consequence**: The workload may look faster, but the covering tests fail broadly — often the whole file's tests, since a wrong index/length in a tree-construction primitive corrupts every parse or raises during import/parse rather than producing one localized diff.
67Partial hoisting: batch call moved out of the loop, but per-element normalization left insidetasksunpy/sunpy
Applies when
task -- the task blames redundant per-iteration work (a conversion/transform re-run per segment/chunk), and the patch hoists that call out of the loop while keeping the follow-up post-processing (rounding, astype, unit stripping, reshaping) inside the loop on individual scalars.
Pattern
The expensive call is correctly batched once over the whole input, but the normalization step that originally operated on the array returned by that call is now re-executed element-by-element on scalars extracted from the batched result. This both leaves O(n) overhead in the loop and silently changes semantics, because array-level and scalar-level versions of the same operation can differ in dtype, precision, unit/masked handling, and edge cases (NaN/inf, empty or degenerate inputs, integer overflow), and because the number and order of conversions per value changes.
Detection procedure
  1. From the task/workload, identify the loop that dominates cost and the exact sequence of operations applied to each iteration's intermediate result in the original code.
  2. In the patch, check where the hoisted batch call ends and whether every downstream step that used to consume the whole per-iteration array now consumes the full batched array once, or instead individual elements inside the loop.
  3. For each step still inside the loop, ask whether it is type/shape-sensitive (rounding + integer cast, unit or Quantity handling, masking, normalization over a set of values) and whether applying it to a scalar can yield a different value or dtype than applying it to the array.
  4. Confirm the loop body is now genuinely cheap: count remaining allocations/conversions per iteration; if they scale with vertex/segment count, the reported "steep scaling with number of items" is only partly removed.
Discriminator
Fine if the only thing left in the loop is pure indexing plus the intrinsic per-segment computation, and every conversion/normalization is applied exactly once at array level as before. A violation is when a conversion or cast that the original code applied to an array is reissued per element (or reissued twice for shared boundary values), so results can differ or overhead persists.
Consequence
The workload shows only a partial speedup (per-iteration conversion overhead remains), and at least one covering test fails on a dtype/rounding/edge-case mismatch — i.e. behaviour changed while the promised "same results, faster" contract is broken.
id ee783e8dce23 · mined from sunpy/sunpy sunpy__sunpy.perf_r1_0
raw text (what the judge reads)
### Partial hoisting: batch call moved out of the loop, but per-element normalization left inside
- **Applies when**: `task` -- the task blames redundant per-iteration work (a conversion/transform re-run per segment/chunk), and the patch hoists that call out of the loop while keeping the follow-up post-processing (rounding, `astype`, unit stripping, reshaping) inside the loop on individual scalars.
- **Pattern**: The expensive call is correctly batched once over the whole input, but the normalization step that originally operated on the *array* returned by that call is now re-executed element-by-element on scalars extracted from the batched result. This both leaves O(n) overhead in the loop and silently changes semantics, because array-level and scalar-level versions of the same operation can differ in dtype, precision, unit/masked handling, and edge cases (NaN/inf, empty or degenerate inputs, integer overflow), and because the number and order of conversions per value changes.
- **Detection procedure**:
  1. From the task/workload, identify the loop that dominates cost and the exact sequence of operations applied to each iteration's intermediate result in the original code.
  2. In the patch, check where the hoisted batch call ends and whether every downstream step that used to consume the whole per-iteration array now consumes the full batched array once, or instead individual elements inside the loop.
  3. For each step still inside the loop, ask whether it is type/shape-sensitive (rounding + integer cast, unit or Quantity handling, masking, normalization over a set of values) and whether applying it to a scalar can yield a different value or dtype than applying it to the array.
  4. Confirm the loop body is now genuinely cheap: count remaining allocations/conversions per iteration; if they scale with vertex/segment count, the reported "steep scaling with number of items" is only partly removed.
- **Discriminator**: Fine if the only thing left in the loop is pure indexing plus the intrinsic per-segment computation, and every conversion/normalization is applied exactly once at array level as before. A violation is when a conversion or cast that the original code applied to an array is reissued per element (or reissued twice for shared boundary values), so results can differ or overhead persists.
- **Consequence**: The workload shows only a partial speedup (per-iteration conversion overhead remains), and at least one covering test fails on a dtype/rounding/edge-case mismatch — i.e. behaviour changed while the promised "same results, faster" contract is broken.
68Rewriting a snapshot-based loop over a mutating container into a hand-rolled live index cursortaskandialbrecht/sqlparse
Applies when
task -- the hot spot is an O(n²) loop that iterates a copy of a container while mutating the original (and pays a linear index lookup per element), and the patch replaces it with a while/index cursor walking the live, mutating container.
Pattern
Instead of keeping the snapshot traversal and replacing only the expensive position lookup with cheap incremental bookkeeping (e.g. a running offset that maps snapshot position → current position), the patch reinvents the traversal: manual idx += 1 scattered across every branch plus an ad-hoc reset after each mutation. The visited set, ordering, and the meaning of previously stored indices are silently redefined, so the loop is no longer provably equivalent to the original.
Detection procedure
  1. In the original code, identify why iteration is over a copy and what invariant the per-element position lookup guarantees (which elements are visited, and that stored indices from earlier iterations stay valid after mutations).
  2. In the patch, check whether the new cursor reproduces exactly that visited sequence: enumerate each branch's increment, the post-mutation reset value, and whether elements newly created/shifted by the mutation are re-visited or skipped.
  3. Check consistency of any state carried across iterations (stacks/lists of saved indices, offsets): after an in-place mutation shrinks or grows the container, are those saved values still correct under the new indexing scheme?
  4. Prefer/expect the minimal fix — same iteration source, one arithmetic offset updated at the mutation site — and treat a full restructuring of control flow as unproven until each branch is hand-simulated on a nested/unbalanced input.
Discriminator
Fine if the new cursor is accompanied by an explicit, checkable index-mapping argument and every branch (including early-continue and error paths) advances identically to the snapshot order, with saved indices rebased; a violation is when the mutation point resets or skips positions, or leaves earlier-saved indices interpreted under a different indexing basis — especially when the workload's flat input never exercises the nested/unbalanced paths that the restructuring affects.
Consequence
The workload may well get faster, but parsing/grouping results change or the loop mis-advances on nested or unbalanced input, so the covering tests fail wholesale (here 0/279 passed) — a speedup bought with incorrect output.
id 408e3756d67c · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_r1_2
raw text (what the judge reads)
### Rewriting a snapshot-based loop over a mutating container into a hand-rolled live index cursor

- **Applies when**: `task` -- the hot spot is an O(n²) loop that iterates a *copy* of a container while mutating the original (and pays a linear index lookup per element), and the patch replaces it with a `while`/index cursor walking the live, mutating container.
- **Pattern**: Instead of keeping the snapshot traversal and replacing only the expensive position lookup with cheap incremental bookkeeping (e.g. a running offset that maps snapshot position → current position), the patch reinvents the traversal: manual `idx += 1` scattered across every branch plus an ad-hoc reset after each mutation. The visited set, ordering, and the meaning of previously stored indices are silently redefined, so the loop is no longer provably equivalent to the original.
- **Detection procedure**:
  1. In the original code, identify why iteration is over a copy and what invariant the per-element position lookup guarantees (which elements are visited, and that stored indices from earlier iterations stay valid after mutations).
  2. In the patch, check whether the new cursor reproduces exactly that visited sequence: enumerate each branch's increment, the post-mutation reset value, and whether elements newly created/shifted by the mutation are re-visited or skipped.
  3. Check consistency of any state carried across iterations (stacks/lists of saved indices, offsets): after an in-place mutation shrinks or grows the container, are those saved values still correct under the new indexing scheme?
  4. Prefer/expect the minimal fix — same iteration source, one arithmetic offset updated at the mutation site — and treat a full restructuring of control flow as unproven until each branch is hand-simulated on a nested/unbalanced input.
- **Discriminator**: Fine if the new cursor is accompanied by an explicit, checkable index-mapping argument and every branch (including early-`continue` and error paths) advances identically to the snapshot order, with saved indices rebased; a violation is when the mutation point resets or skips positions, or leaves earlier-saved indices interpreted under a different indexing basis — especially when the workload's flat input never exercises the nested/unbalanced paths that the restructuring affects.
- **Consequence**: The workload may well get faster, but parsing/grouping results change or the loop mis-advances on nested or unbalanced input, so the covering tests fail wholesale (here 0/279 passed) — a speedup bought with incorrect output.