perf-rubrics-onpolicy-10

SWE-fficiency · 10 rubrics · first 10 of all
HF EdwardoSunny/perf-rubrics-onpolicy-10 · local data/libraries/perf-rubrics-onpolicy-10.json

0Patch is a mechanically unsound edit: the rewritten hot-path block no longer forms valid, importable codetaskalecthomas/voluptuous
Applies when
task -- the patch replaces or restructures a block inside a function of a module that the test suite (and the workload) imports, rather than adding a new isolated helper.
Pattern
The optimization idea is algorithmically right, but the applied hunk leaves the edited region structurally inconsistent — dangling/duplicated statements, wrong indentation relative to the surrounding if/else or loop, a return/branch moved out of its block, or references to names, annotations or imports that the same hunk removed. The module then fails at import/parse time, so nothing that depends on it runs at all.
Detection procedure
  1. From the workload, note which module(s) get imported to reach the hot path; any breakage there is fatal for every test, not just the changed function.
  2. Reconstruct the post-patch text of the whole edited function from the diff context (not just the + lines) and read it top to bottom as if it were a fresh file.
  3. Check mechanically: indentation of each surviving line matches its intended block; every branch still terminates with the value it used to return; every identifier, annotation, and helper referenced is still defined/imported; no leftover statement from the old algorithm remains half-removed.
  4. If the resulting function body is not something you could paste into the file and import cleanly, flag the patch regardless of how good the complexity argument is.
Discriminator
A real violation is a structural defect — the post-patch code cannot parse/import or a branch silently loses its result — which makes the entire suite fail uniformly. A look-alike that is fine is a patch that merely keeps redundant-looking scaffolding (e.g., an intermediate list, a now-unnecessary type annotation, an extra loop that only appends) while all names resolve and the block’s control flow and return values are intact; that is cosmetically clumsy but behaviour-preserving and still delivers the complexity win.
Consequence
The workload script and the covering tests both die during import/collection, showing 0 of N tests passing (an error, not assertion failures), so no timing measurement is even produced — as opposed to the correct version, which keeps all tests green and turns the quadratic combining step into an O(n log n) one.
id 5d3df984d519 · mined from alecthomas/voluptuous alecthomas__voluptuous.perf_2
raw text (what the judge reads)
### Patch is a mechanically unsound edit: the rewritten hot-path block no longer forms valid, importable code
- **Applies when**: `task` -- the patch replaces or restructures a block inside a function of a module that the test suite (and the workload) imports, rather than adding a new isolated helper.
- **Pattern**: The optimization idea is algorithmically right, but the applied hunk leaves the edited region structurally inconsistent — dangling/duplicated statements, wrong indentation relative to the surrounding `if/else` or loop, a `return`/branch moved out of its block, or references to names, annotations or imports that the same hunk removed. The module then fails at import/parse time, so nothing that depends on it runs at all.
- **Detection procedure**:
  1. From the workload, note which module(s) get imported to reach the hot path; any breakage there is fatal for every test, not just the changed function.
  2. Reconstruct the post-patch text of the whole edited function from the diff context (not just the `+` lines) and read it top to bottom as if it were a fresh file.
  3. Check mechanically: indentation of each surviving line matches its intended block; every branch still terminates with the value it used to return; every identifier, annotation, and helper referenced is still defined/imported; no leftover statement from the old algorithm remains half-removed.
  4. If the resulting function body is not something you could paste into the file and import cleanly, flag the patch regardless of how good the complexity argument is.
- **Discriminator**: A real violation is a *structural* defect — the post-patch code cannot parse/import or a branch silently loses its result — which makes the entire suite fail uniformly. A look-alike that is fine is a patch that merely keeps redundant-looking scaffolding (e.g., an intermediate list, a now-unnecessary type annotation, an extra loop that only appends) while all names resolve and the block’s control flow and return values are intact; that is cosmetically clumsy but behaviour-preserving and still delivers the complexity win.
- **Consequence**: The workload script and the covering tests both die during import/collection, showing 0 of N tests passing (an error, not assertion failures), so no timing measurement is even produced — as opposed to the correct version, which keeps all tests green and turns the quadratic combining step into an O(n log n) one.
1Caching a derived value behind an interface that callers still write to or expect to be livetaskmarshmallow-code/webargs
Applies when
task -- the patch removes repeated recomputation by memoizing a derived value (computed from constructor inputs) while leaving the original accessor — a property, method, or attribute — in place as the public entry point.
Pattern
The patch adds a private cached field populated once (usually in __init__) and makes the existing accessor return it, but does not check the accessor's full contract: it may still be read-only (a property with no setter), may hand out a shared mutable object where a fresh one was produced each call, and its inputs may be reassigned or mutated after construction. Callers, subclasses, or tests that assign to the accessor, mutate the returned object, or change the source input now get an exception or a stale value.
Detection procedure
  1. In the task/workload, identify the recomputed accessor named as the hot spot and note whether it is a property, plain attribute, or method, and whether it returns a mutable container.
  2. Grep the repository (library code, subclasses, and tests) for every read and write of that accessor and of the inputs it derives from — assignments, del, in-place mutation of the returned value, and post-construction mutation of the source.
  3. Check the patch: if it converts a writable-looking name into a read-only cached property, or freezes a value whose source is written to anywhere found in step 2, without adding a setter/invalidation or converting the name to a plain settable attribute, flag it.
  4. Confirm the cached value is computed from data that is genuinely fixed at construction time and not from per-call arguments.
Discriminator
A real violation is a cache whose accessor loses a capability it previously had (assignability, freshness after input change, per-call fresh object) with at least one existing caller relying on it. A look-alike that is fine: the value is derived only from immutable constructor state, the accessor is never assigned or its result never mutated anywhere in the codebase, and the patch either keeps it a plain attribute or supplies a setter/invalidation hook.
Consequence
The workload may show the expected speedup, but the covering tests fail broadly — typically AttributeError: can't set attribute at construction or in any test that overrides the value, or subtle wrong results from stale/shared state — turning a pure speed change into a behavioural regression.
id b0b6558f6fbc · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_0
raw text (what the judge reads)
### Caching a derived value behind an interface that callers still write to or expect to be live
- **Applies when**: `task` -- the patch removes repeated recomputation by memoizing a derived value (computed from constructor inputs) while leaving the original accessor — a property, method, or attribute — in place as the public entry point.
- **Pattern**: The patch adds a private cached field populated once (usually in `__init__`) and makes the existing accessor return it, but does not check the accessor's full contract: it may still be read-only (a property with no setter), may hand out a shared mutable object where a fresh one was produced each call, and its inputs may be reassigned or mutated after construction. Callers, subclasses, or tests that assign to the accessor, mutate the returned object, or change the source input now get an exception or a stale value.
- **Detection procedure**:
  1. In the task/workload, identify the recomputed accessor named as the hot spot and note whether it is a property, plain attribute, or method, and whether it returns a mutable container.
  2. Grep the repository (library code, subclasses, and tests) for every read *and write* of that accessor and of the inputs it derives from — assignments, `del`, in-place mutation of the returned value, and post-construction mutation of the source.
  3. Check the patch: if it converts a writable-looking name into a read-only cached property, or freezes a value whose source is written to anywhere found in step 2, without adding a setter/invalidation or converting the name to a plain settable attribute, flag it.
  4. Confirm the cached value is computed from data that is genuinely fixed at construction time and not from per-call arguments.
- **Discriminator**: A real violation is a cache whose accessor loses a capability it previously had (assignability, freshness after input change, per-call fresh object) with at least one existing caller relying on it. A look-alike that is fine: the value is derived only from immutable constructor state, the accessor is never assigned or its result never mutated anywhere in the codebase, and the patch either keeps it a plain attribute or supplies a setter/invalidation hook.
- **Consequence**: The workload may show the expected speedup, but the covering tests fail broadly — typically `AttributeError: can't set attribute` at construction or in any test that overrides the value, or subtle wrong results from stale/shared state — turning a pure speed change into a behavioural regression.
2Replacing a hand-rolled loop in a shared low-level helper with a library "fast constructor" whose accepted input types / edge-case semantics are narrowertaskpyca/pyopenssl
Applies when
task -- the profile points at per-byte Python work inside a small, widely-shared helper (buffer copying, encoding, conversion) and the patch swaps the explicit loop for a single built-in/FFI/library call that does the same thing in C.
Pattern
The patch assumes the one-shot constructor is a drop-in equivalent of the loop, but the loop tolerated inputs and corner cases the new call does not (bytearray/memoryview/other buffer-protocol objects, str, empty input, sentinel/terminator or length conventions, ownership/lifetime of the allocated memory). Because the helper sits under every entry point, a mismatch does not slow one path down — it makes the whole module unusable, which is why the entire covering suite can go from all-pass to zero-pass.
Detection procedure
  1. From the task and workload, identify the helper being edited and note how central it is: grep its call sites and ask "does every public operation, not just the timed one, flow through here?"
  2. Enumerate the input shapes those call sites (and the repo's tests/docs for them) actually pass: exact type(s), zero-length values, oversized values, and any implicit contract about terminators, reported length, or who owns/keeps alive the produced object.
  3. For each such shape, check the replacement call's documented behaviour — does it accept the type at all, allocate the same size, keep the buffer alive for the consumer's lifetime, and return the same length semantics? Any "probably fine" answer is a violation.
  4. Sanity-check that the edited module still imports and one non-timed operation still works, before trusting any timing number.
Discriminator
A real violation is when the new call's contract is provably narrower or different for at least one input shape reachable from a call site (e.g. only bytes accepted where callers may pass a mutable buffer, or the length/terminator convention shifts). It is not a violation when the surrounding code already normalizes/validates the input type before the helper is reached and the constructor's size, termination, and lifetime guarantees are documented to match — then the C-level replacement is exactly the intended fix.
Consequence
The workload may show the expected speedup (it exercises only the happy-path input), while the covering tests collapse wholesale — import or every-test errors rather than a few localized failures — signalling that a universally used helper's contract was broken for speed.
id 24e9ce529b79 · mined from pyca/pyopenssl pyca__pyopenssl.perf_1
raw text (what the judge reads)
### Replacing a hand-rolled loop in a shared low-level helper with a library "fast constructor" whose accepted input types / edge-case semantics are narrower

- **Applies when**: `task` -- the profile points at per-byte Python work inside a small, widely-shared helper (buffer copying, encoding, conversion) and the patch swaps the explicit loop for a single built-in/FFI/library call that does the same thing in C.
- **Pattern**: The patch assumes the one-shot constructor is a drop-in equivalent of the loop, but the loop tolerated inputs and corner cases the new call does not (bytearray/memoryview/other buffer-protocol objects, `str`, empty input, sentinel/terminator or length conventions, ownership/lifetime of the allocated memory). Because the helper sits under *every* entry point, a mismatch does not slow one path down — it makes the whole module unusable, which is why the entire covering suite can go from all-pass to zero-pass.
- **Detection procedure**:
  1. From the task and workload, identify the helper being edited and note how central it is: grep its call sites and ask "does every public operation, not just the timed one, flow through here?"
  2. Enumerate the input shapes those call sites (and the repo's tests/docs for them) actually pass: exact type(s), zero-length values, oversized values, and any implicit contract about terminators, reported length, or who owns/keeps alive the produced object.
  3. For each such shape, check the replacement call's documented behaviour — does it accept the type at all, allocate the same size, keep the buffer alive for the consumer's lifetime, and return the same length semantics? Any "probably fine" answer is a violation.
  4. Sanity-check that the edited module still imports and one non-timed operation still works, before trusting any timing number.
- **Discriminator**: A real violation is when the new call's contract is provably narrower or different for at least one input shape reachable from a call site (e.g. only `bytes` accepted where callers may pass a mutable buffer, or the length/terminator convention shifts). It is *not* a violation when the surrounding code already normalizes/validates the input type before the helper is reached and the constructor's size, termination, and lifetime guarantees are documented to match — then the C-level replacement is exactly the intended fix.
- **Consequence**: The workload may show the expected speedup (it exercises only the happy-path input), while the covering tests collapse wholesale — import or every-test errors rather than a few localized failures — signalling that a universally used helper's contract was broken for speed.
3Removing eager materialization of a lazy sequence, and early-exiting from it, without proving the full traversal wasn't load-bearingtaskandialbrecht/sqlparse
Applies when
task -- the patch speeds up a lookup by deleting a list(...)/tuple snapshot of a lazily produced sequence (generator, recursive tree walk, streaming iterator) and instead iterates it directly with an early return/break once the target is found.
Pattern
The rewrite is asymptotically better on paper (drops a quadratic prefix re-scan), but it silently changes when and how much of the producer runs: the producer is now advanced partially, entered re-entrantly, or interleaved with the caller's own work, and any side effect, shared cursor, lazy-initialization, or mutation-during-iteration that the eager snapshot used to isolate now leaks into unrelated behaviour.
Detection procedure
  1. From the task, note which shared/core API is being edited and how many tests are declared as "covering" it — a large covering set signals the function's data source is central shared machinery, not a leaf helper.
  2. Read the producer that was previously materialized: does it recurse over a structure that other code mutates, trigger lazy parsing/grouping/caching, keep per-object iteration state, or get consumed elsewhere while this call is in progress?
  3. Check the new control flow: the early return abandons the producer mid-stream, and the accumulator state (running offset, index) is now derived incrementally rather than recomputed — confirm both hold for every exit path, including "not found" and nested/re-entrant calls.
  4. If steps 2–3 leave any doubt, demand that the patch either keep a cheap materialization (still linear, e.g. one snapshot per call) or add evidence that partial consumption is safe; a patch that offers neither is inadequate before any timing is done.
Discriminator
Not a violation when the producer is a pure, side-effect-free read of an immutable/finished structure and no other code holds a live iterator over it — then partial consumption is genuinely equivalent and faster. It is a violation when the producer performs lazy construction, mutates or is mutated during traversal, or shares iteration state, so the old "consume everything into a list first" step was semantically load-bearing rather than merely wasteful.
Consequence
The workload may look faster, but the covering suite degrades far beyond the edited feature — collection/import-time or near-total failures across unrelated tests, because half-consumed lazy traversal leaves shared structures in an inconsistent state.
id be720c596805 · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_2
raw text (what the judge reads)
### Removing eager materialization of a lazy sequence, and early-exiting from it, without proving the full traversal wasn't load-bearing
- **Applies when**: `task` -- the patch speeds up a lookup by deleting a `list(...)`/tuple snapshot of a lazily produced sequence (generator, recursive tree walk, streaming iterator) and instead iterates it directly with an early `return`/`break` once the target is found.
- **Pattern**: The rewrite is asymptotically better on paper (drops a quadratic prefix re-scan), but it silently changes *when and how much* of the producer runs: the producer is now advanced partially, entered re-entrantly, or interleaved with the caller's own work, and any side effect, shared cursor, lazy-initialization, or mutation-during-iteration that the eager snapshot used to isolate now leaks into unrelated behaviour.
- **Detection procedure**:
  1. From the task, note which shared/core API is being edited and how many tests are declared as "covering" it — a large covering set signals the function's data source is central shared machinery, not a leaf helper.
  2. Read the producer that was previously materialized: does it recurse over a structure that other code mutates, trigger lazy parsing/grouping/caching, keep per-object iteration state, or get consumed elsewhere while this call is in progress?
  3. Check the new control flow: the early `return` abandons the producer mid-stream, and the accumulator state (running offset, index) is now derived incrementally rather than recomputed — confirm both hold for *every* exit path, including "not found" and nested/re-entrant calls.
  4. If steps 2–3 leave any doubt, demand that the patch either keep a cheap materialization (still linear, e.g. one snapshot per call) or add evidence that partial consumption is safe; a patch that offers neither is inadequate before any timing is done.
- **Discriminator**: Not a violation when the producer is a pure, side-effect-free read of an immutable/finished structure and no other code holds a live iterator over it — then partial consumption is genuinely equivalent and faster. It *is* a violation when the producer performs lazy construction, mutates or is mutated during traversal, or shares iteration state, so the old "consume everything into a list first" step was semantically load-bearing rather than merely wasteful.
- **Consequence**: The workload may look faster, but the covering suite degrades far beyond the edited feature — collection/import-time or near-total failures across unrelated tests, because half-consumed lazy traversal leaves shared structures in an inconsistent state.
4Unrelated collateral edits bundled with the hot-path change (and hand-rolled logic where a tested helper exists)taskandialbrecht/sqlparse
Applies when
task -- the patch is supposed to be a small, behaviour-preserving speedup of one lookup, but its diff touches lines that have nothing to do with the hot path (blank lines, indentation, decorators, imports, surrounding definitions) and/or re-implements inline a scan/search that the codebase already exposes as a shared helper.
Pattern
The agent makes the intended micro-optimization (e.g. replacing a full materialized scan with an early-exit loop) but also silently rewrites adjacent module structure/whitespace, so the file's top-level layout changes; and it inlines logic that an existing, already-tested primitive performs, duplicating semantics (skip rules, edge cases) instead of delegating.
Detection procedure
  1. From the task, identify the single function/expression that the workload times, and note the minimal edit needed there.
  2. Walk every hunk of the diff and classify each changed line as "required for the speedup" or "collateral" (formatting, blank-line count, class/def boundaries, imports, unrelated statements).
  3. Reject if any collateral line changes module-level structure or style that import-time checks, linters, or style tests in the suite could enforce — such edits can fail collection and take the entire suite down, independent of the optimization's merits.
  4. Check whether the rewritten logic duplicates an existing helper in the same module/class that already implements the same "first meaningful element" / skip semantics; if so, require delegation to it rather than a private copy.
Discriminator
A real violation is a diff whose non-essential lines alter file structure or duplicate semantics maintained elsewhere. A look-alike that is fine is a diff that only touches the timed expression (even if the new code is longer), or one that reformats lines it had to rewrite anyway, with no change to top-level blank-line/definition layout and no duplication of an existing helper's rules.
Consequence
The workload may well get faster, but the covering tests fail wholesale (0 passed — collection/style gate failure) or drift later as the duplicated skip logic diverges from the shared helper, so the change is rejected despite the speedup.
id 6eeca17c6203 · mined from andialbrecht/sqlparse andialbrecht__sqlparse.perf_3
raw text (what the judge reads)
### Unrelated collateral edits bundled with the hot-path change (and hand-rolled logic where a tested helper exists)
- **Applies when**: `task` -- the patch is supposed to be a small, behaviour-preserving speedup of one lookup, but its diff touches lines that have nothing to do with the hot path (blank lines, indentation, decorators, imports, surrounding definitions) and/or re-implements inline a scan/search that the codebase already exposes as a shared helper.
- **Pattern**: The agent makes the intended micro-optimization (e.g. replacing a full materialized scan with an early-exit loop) but also silently rewrites adjacent module structure/whitespace, so the file's top-level layout changes; and it inlines logic that an existing, already-tested primitive performs, duplicating semantics (skip rules, edge cases) instead of delegating.
- **Detection procedure**:
  1. From the task, identify the single function/expression that the workload times, and note the minimal edit needed there.
  2. Walk every hunk of the diff and classify each changed line as "required for the speedup" or "collateral" (formatting, blank-line count, class/def boundaries, imports, unrelated statements).
  3. Reject if any collateral line changes module-level structure or style that import-time checks, linters, or style tests in the suite could enforce — such edits can fail collection and take the entire suite down, independent of the optimization's merits.
  4. Check whether the rewritten logic duplicates an existing helper in the same module/class that already implements the same "first meaningful element" / skip semantics; if so, require delegation to it rather than a private copy.
- **Discriminator**: A real violation is a diff whose non-essential lines alter file structure or duplicate semantics maintained elsewhere. A look-alike that is fine is a diff that only touches the timed expression (even if the new code is longer), or one that reformats lines it had to rewrite anyway, with no change to top-level blank-line/definition layout and no duplication of an existing helper's rules.
- **Consequence**: The workload may well get faster, but the covering tests fail wholesale (0 passed — collection/style gate failure) or drift later as the duplicated skip logic diverges from the shared helper, so the change is rejected despite the speedup.
5New helper/module referenced without ensuring it is imported in the patched scopetaskInstagram/MonkeyType
Applies when
task -- a patch replaces a hand-rolled hot loop with a standard-library helper (e.g. a counter, cache, or itertools/heapq utility) or a new internal helper name.
Pattern
The rewrite is algorithmically correct, but the symbol it now calls is not bound in the edited module: the diff adds a use of somemodule.something / SomeHelper while touching only the function body and never adding (or verifying the pre-existing presence of) the corresponding import/from ... import. The result is a NameError/AttributeError at the first call, so every test exercising that code path — and often the whole test module — fails.
Detection procedure
  1. In the diff, list every name introduced on the added lines that is not a local variable, parameter, or builtin.
  2. For each such name, check whether the diff itself adds the import, or whether the unchanged header of the same file (context lines / the file as it exists) already imports it — do not assume it does because the name is "standard library".
  3. If the name's binding cannot be pointed to in either place, treat the patch as broken regardless of how plausible the optimization is.
  4. Also confirm the added call's signature/return shape matches how the result is consumed (e.g. iterating .items() on something that is actually a mapping).
Discriminator
A real violation is an unbound name — no import in the diff and none in the file. A look-alike that is fine is a name that is already imported at module top (or is a builtin such as len, dict, sorted, set), even though the diff does not mention the import; also fine is a fully-qualified use where the parent package is imported.
Consequence
The workload and the covering tests fail immediately at the first invocation with NameError: name '...' is not defined (typically all tests in the touched module report failure), so no speedup can be measured at all.
id a2c2565dd3aa · mined from Instagram/MonkeyType Instagram__MonkeyType.perf_3
raw text (what the judge reads)
### New helper/module referenced without ensuring it is imported in the patched scope
- **Applies when**: `task` -- a patch replaces a hand-rolled hot loop with a standard-library helper (e.g. a counter, cache, or itertools/heapq utility) or a new internal helper name.
- **Pattern**: The rewrite is algorithmically correct, but the symbol it now calls is not bound in the edited module: the diff adds a use of `somemodule.something` / `SomeHelper` while touching only the function body and never adding (or verifying the pre-existing presence of) the corresponding `import`/`from ... import`. The result is a `NameError`/`AttributeError` at the first call, so every test exercising that code path — and often the whole test module — fails.
- **Detection procedure**:
  1. In the diff, list every name introduced on the added lines that is not a local variable, parameter, or builtin.
  2. For each such name, check whether the diff itself adds the import, or whether the unchanged header of the same file (context lines / the file as it exists) already imports it — do not assume it does because the name is "standard library".
  3. If the name's binding cannot be pointed to in either place, treat the patch as broken regardless of how plausible the optimization is.
  4. Also confirm the added call's signature/return shape matches how the result is consumed (e.g. iterating `.items()` on something that is actually a mapping).
- **Discriminator**: A real violation is an unbound name — no import in the diff and none in the file. A look-alike that is fine is a name that *is* already imported at module top (or is a builtin such as `len`, `dict`, `sorted`, `set`), even though the diff does not mention the import; also fine is a fully-qualified use where the parent package is imported.
- **Consequence**: The workload and the covering tests fail immediately at the first invocation with `NameError: name '...' is not defined` (typically all tests in the touched module report failure), so no speedup can be measured at all.
6Hoisting per-request setup that silently drops request-dependent construction parameterstaskmarshmallow-code/webargs
Applies when
task -- the task asks to move repeated "build the schema/validator/parser object" work out of the per-call path into decoration/initialization time, and the patch constructs that object eagerly from a static spec.
Pattern
The patch pre-builds, at decoration/import time, an object that the runtime path previously built from a plain spec — but the runtime builder also injected contextual options (e.g. unknown-field policy derived from the data location, partial/many flags, per-call overrides, error handling config). The eager construction uses only default options, so the object now behaves differently, and the runtime path no longer recognizes the spec type well enough to re-apply the missing options.
Detection procedure
  1. In the unpatched code, find the runtime function that consumed the raw spec and note every argument it passed into the constructor, including values derived from the call site (location, request attributes, flags, defaults resolved lazily).
  2. Compare with the patch's eager construction site: list which of those arguments are unavailable or omitted there.
  3. Check whether the runtime path has a branch that treats an already-built object differently (skipping the option-injection step) — if so, the omitted options are now permanently lost.
  4. Look for a behavioural knob among the omitted options (unknown/extra-field handling, validation strictness, partial loading); any such omission is a correctness change, not a pure speedup.
Discriminator
A real violation omits or freezes an option that varies per call or per location, or that differs from the constructor's default; a benign look-alike hoists construction while explicitly threading through the same options (or defers construction until the first call, when the context is known) so the built object is provably identical to what the per-call path produced.
Consequence
The workload may indeed get faster (its single fixed configuration happens to match), but the covering tests fail broadly — typically every test through the decorated path errors or returns different validation results (e.g. extra/unknown inputs now rejected or silently kept), i.e. speed bought by changing semantics.
id 4b5dc68cb4d0 · mined from marshmallow-code/webargs marshmallow-code__webargs.perf_2
raw text (what the judge reads)
### Hoisting per-request setup that silently drops request-dependent construction parameters

- **Applies when**: `task` -- the task asks to move repeated "build the schema/validator/parser object" work out of the per-call path into decoration/initialization time, and the patch constructs that object eagerly from a static spec.
- **Pattern**: The patch pre-builds, at decoration/import time, an object that the runtime path previously built from a plain spec — but the runtime builder also injected contextual options (e.g. unknown-field policy derived from the data location, partial/many flags, per-call overrides, error handling config). The eager construction uses only default options, so the object now behaves differently, and the runtime path no longer recognizes the spec type well enough to re-apply the missing options.
- **Detection procedure**:
  1. In the unpatched code, find the runtime function that consumed the raw spec and note *every* argument it passed into the constructor, including values derived from the call site (location, request attributes, flags, defaults resolved lazily).
  2. Compare with the patch's eager construction site: list which of those arguments are unavailable or omitted there.
  3. Check whether the runtime path has a branch that treats an already-built object differently (skipping the option-injection step) — if so, the omitted options are now permanently lost.
  4. Look for a behavioural knob among the omitted options (unknown/extra-field handling, validation strictness, partial loading); any such omission is a correctness change, not a pure speedup.
- **Discriminator**: A real violation omits or freezes an option that varies per call or per location, or that differs from the constructor's default; a benign look-alike hoists construction while explicitly threading through the same options (or defers construction until the first call, when the context is known) so the built object is provably identical to what the per-call path produced.
- **Consequence**: The workload may indeed get faster (its single fixed configuration happens to match), but the covering tests fail broadly — typically every test through the decorated path errors or returns different validation results (e.g. extra/unknown inputs now rejected or silently kept), i.e. speed bought by changing semantics.
7Cache of per-class introspection stored where inheritance (or dynamic attributes) can leak the wrong entrytaskInstagram/MonkeyType
Applies when
task -- the hot path builds a lookup table by introspecting an object (dir, __dict__, attribute scans) and the patch memoizes that table on the class/type instead of recomputing it.
Pattern
The patch checks for the cache with an inheritance-visible lookup (hasattr(cls, ...) / getattr(cls, ...)) and stores it with setattr(cls, ...), so the first class in a hierarchy to populate the cache silently supplies its table to every base/subclass — and any attribute added after first use is never seen. The dispatch table is then wrong for exactly the objects whose behaviour depends on subclass overrides or extra hooks.
Detection procedure
  1. From the task/workload, identify which concrete object the hot path runs on and whether that type is a subclass (or has siblings) of the class whose introspection is being cached.
  2. In the patch, find where the cache is read and written; check whether the read can resolve to an attribute defined on a different class (base or sibling) than the one being written to, and whether a sentinel/cls.__dict__ check is used.
  3. Check whether the cached table can go stale: are the introspected attributes ever added/overridden per instance, per subclass, or at runtime after the first call?
  4. If the read is inheritance-visible or the table can go stale, treat the dispatch as potentially resolving to the wrong handler (or None) for some types.
Discriminator
Fine: cache keyed on cls.__dict__ presence or a module-level dict keyed by the exact type(self), populated from the full MRO, on a table that cannot change after class creation. Violation: hasattr(cls, name) / plain class attribute lookup as the presence check, or any scheme where one class's table can be served to another class in the hierarchy. Also fine is not caching at all if the lookup was replaced by a single direct attribute fetch (which is inherently correct under inheritance).
Consequence
Handlers resolve to the wrong (or missing) methods for subclasses, so the recursion falls through to fallback/error paths — the workload's rendering output changes or raises, and the covering tests fail broadly (near-total failure), even though the microbenchmark may look faster.
id bdbf6ff2027a · mined from Instagram/MonkeyType Instagram__MonkeyType.perf_2
raw text (what the judge reads)
### Cache of per-class introspection stored where inheritance (or dynamic attributes) can leak the wrong entry
- **Applies when**: `task` -- the hot path builds a lookup table by introspecting an object (`dir`, `__dict__`, attribute scans) and the patch memoizes that table on the class/type instead of recomputing it.
- **Pattern**: The patch checks for the cache with an inheritance-visible lookup (`hasattr(cls, ...)` / `getattr(cls, ...)`) and stores it with `setattr(cls, ...)`, so the first class in a hierarchy to populate the cache silently supplies its table to every base/subclass — and any attribute added after first use is never seen. The dispatch table is then wrong for exactly the objects whose behaviour depends on subclass overrides or extra hooks.
- **Detection procedure**:
  1. From the task/workload, identify which concrete object the hot path runs on and whether that type is a subclass (or has siblings) of the class whose introspection is being cached.
  2. In the patch, find where the cache is read and written; check whether the read can resolve to an attribute defined on a *different* class (base or sibling) than the one being written to, and whether a sentinel/`cls.__dict__` check is used.
  3. Check whether the cached table can go stale: are the introspected attributes ever added/overridden per instance, per subclass, or at runtime after the first call?
  4. If the read is inheritance-visible or the table can go stale, treat the dispatch as potentially resolving to the wrong handler (or `None`) for some types.
- **Discriminator**: Fine: cache keyed on `cls.__dict__` presence or a module-level dict keyed by the exact `type(self)`, populated from the full MRO, on a table that cannot change after class creation. Violation: `hasattr(cls, name)` / plain class attribute lookup as the presence check, or any scheme where one class's table can be served to another class in the hierarchy. Also fine is not caching at all if the lookup was replaced by a single direct attribute fetch (which is inherently correct under inheritance).
- **Consequence**: Handlers resolve to the wrong (or missing) methods for subclasses, so the recursion falls through to fallback/error paths — the workload's rendering output changes or raises, and the covering tests fail broadly (near-total failure), even though the microbenchmark may look faster.
8Opting into an internal compile/dispatch protocol instead of reusing already-precomputed worktaskalecthomas/voluptuous
Applies when
task -- the task asks to cut per-item overhead in a reusable validator/handler object, and the patch adds a hook method that makes the surrounding framework treat the object via a different (lower-level) calling convention, storing "compiled" state on the instance.
Pattern
The patch discovers that the object rebuilds work per call, but instead of hoisting that work into state already computed at construction time, it implements an undocumented framework extension hook (__*_compile__-style), returns a callback with a different signature (extra path/context args, different error/return contract), and caches the compiled callbacks on self — mixing two call conventions in one object and mutating shared instance state at schema-build time.
Detection procedure
  1. From the task, note that the expensive work is redundant construction that is already done and stored elsewhere in the object's constructor (or is trivially hoistable) — i.e., a one-line reuse would suffice.
  2. In the patch, check whether it instead opts the object into a framework-internal dispatch/compile protocol, changing the signature or error-wrapping semantics of the invoked callable relative to the plain path.
  3. Check whether the new path duplicates the validation/error semantics of the old path by hand (length checks, message substitution, exception re-wrapping, container type reconstruction, index/path bookkeeping) and whether any of those differ from the original.
  4. Check whether compiled state is stored on self and thus shared across every schema/thread that embeds the same instance, and whether the object can now be called both ways with different behaviour.
Discriminator
A real violation adds a new call path/protocol whose semantics must be re-derived and can diverge (or whose cached state leaks across users), when the same speedup was obtainable by reusing existing precomputed members. It is fine to hoist per-call construction into constructor-time state, or to memoize purely derived data, as long as the object's calling convention, return values, and raised-error types/messages are byte-for-byte unchanged and no state is shared across independent embeddings.
Consequence
Because the new protocol sits on a framework-wide dispatch path, a signature or error-contract mismatch fails far beyond the targeted case — the covering suite collapses almost entirely (here 0/146 passed), while the workload gain, if any, was achievable by a one-line reuse of existing precomputed validators.
id 622e8069d439 · mined from alecthomas/voluptuous alecthomas__voluptuous.perf_1
raw text (what the judge reads)
### Opting into an internal compile/dispatch protocol instead of reusing already-precomputed work
- **Applies when**: `task` -- the task asks to cut per-item overhead in a reusable validator/handler object, and the patch adds a hook method that makes the surrounding framework treat the object via a different (lower-level) calling convention, storing "compiled" state on the instance.
- **Pattern**: The patch discovers that the object rebuilds work per call, but instead of hoisting that work into state already computed at construction time, it implements an undocumented framework extension hook (`__*_compile__`-style), returns a callback with a different signature (extra path/context args, different error/return contract), and caches the compiled callbacks on `self` — mixing two call conventions in one object and mutating shared instance state at schema-build time.
- **Detection procedure**:
  1. From the task, note that the expensive work is redundant construction that is *already* done and stored elsewhere in the object's constructor (or is trivially hoistable) — i.e., a one-line reuse would suffice.
  2. In the patch, check whether it instead opts the object into a framework-internal dispatch/compile protocol, changing the signature or error-wrapping semantics of the invoked callable relative to the plain path.
  3. Check whether the new path duplicates the validation/error semantics of the old path by hand (length checks, message substitution, exception re-wrapping, container type reconstruction, index/path bookkeeping) and whether any of those differ from the original.
  4. Check whether compiled state is stored on `self` and thus shared across every schema/thread that embeds the same instance, and whether the object can now be called both ways with different behaviour.
- **Discriminator**: A real violation adds a *new* call path/protocol whose semantics must be re-derived and can diverge (or whose cached state leaks across users), when the same speedup was obtainable by reusing existing precomputed members. It is fine to hoist per-call construction into constructor-time state, or to memoize purely derived data, as long as the object's calling convention, return values, and raised-error types/messages are byte-for-byte unchanged and no state is shared across independent embeddings.
- **Consequence**: Because the new protocol sits on a framework-wide dispatch path, a signature or error-contract mismatch fails far beyond the targeted case — the covering suite collapses almost entirely (here 0/146 passed), while the workload gain, if any, was achievable by a one-line reuse of existing precomputed validators.
9Hoisting a statement out of a loop by re-indentation, without verifying the block structure still parses/behavestaskiterative/dvc
Applies when
task -- the speedup is achieved by moving an existing statement (a sort, flush, recompute, aggregation) from inside a loop to after it, i.e. the diff is essentially a change of indentation/scope rather than new logic.
Pattern
The patch relies on a purely mechanical scope change: one or more lines are de-indented (or re-nested) so they run once instead of per iteration. Nothing else is added, so the author assumes it is trivially safe and does no local syntax/semantic check — but the new indentation may not line up with the enclosing block (continuation lines, closing brackets, nested for/with/if), or the hoisted result may be needed inside the loop.
Detection procedure
  1. Read the task/workload to confirm the hot cost is the per-iteration repetition of that statement (quadratic growth is a strong hint), so hoisting is the right idea in principle.
  2. Reconstruct the full post-patch block from the diff, including unchanged context lines and bracket continuations, and check every indentation level: does the de-indented statement land at exactly the enclosing function/loop level, does it still follow a complete statement, and is it inside the right for/with/try?
  3. Check whether any code inside the loop reads the state the hoisted statement produced (ordering, flushed buffer, partial aggregate); if so, hoisting changes observable results, not just timing.
  4. Require evidence that the edited module still imports (a parse/import check or a single fast test) — a scope edit that cannot be shown to compile is not reviewable as "obviously equivalent".
Discriminator
A legitimate hoist has the moved statement at an indentation level that provably matches the enclosing block, with no in-loop consumer of its effect, so the final value is identical and only computed once. A violation is a hoist whose indentation cannot be reconciled with the surrounding lines/brackets (possible IndentationError/wrong nesting), or where the loop body depended on the per-iteration effect.
Consequence
If the block structure is broken, the module fails at import and every covering test errors out (0 passed) regardless of the workload; if an in-loop consumer existed, the workload gets faster but results/ordering change and the behaviour tests fail.
id 894d48b5e3ce · mined from iterative/dvc iterative__dvc.perf_2
raw text (what the judge reads)
### Hoisting a statement out of a loop by re-indentation, without verifying the block structure still parses/behaves
- **Applies when**: `task` -- the speedup is achieved by moving an existing statement (a sort, flush, recompute, aggregation) from inside a loop to after it, i.e. the diff is essentially a change of indentation/scope rather than new logic.
- **Pattern**: The patch relies on a purely mechanical scope change: one or more lines are de-indented (or re-nested) so they run once instead of per iteration. Nothing else is added, so the author assumes it is trivially safe and does no local syntax/semantic check — but the new indentation may not line up with the enclosing block (continuation lines, closing brackets, nested `for`/`with`/`if`), or the hoisted result may be needed inside the loop.
- **Detection procedure**:
  1. Read the task/workload to confirm the hot cost is the per-iteration repetition of that statement (quadratic growth is a strong hint), so hoisting is the right idea in principle.
  2. Reconstruct the *full* post-patch block from the diff, including unchanged context lines and bracket continuations, and check every indentation level: does the de-indented statement land at exactly the enclosing function/loop level, does it still follow a complete statement, and is it inside the right `for`/`with`/`try`?
  3. Check whether any code inside the loop reads the state the hoisted statement produced (ordering, flushed buffer, partial aggregate); if so, hoisting changes observable results, not just timing.
  4. Require evidence that the edited module still imports (a parse/`import` check or a single fast test) — a scope edit that cannot be shown to compile is not reviewable as "obviously equivalent".
- **Discriminator**: A legitimate hoist has the moved statement at an indentation level that provably matches the enclosing block, with no in-loop consumer of its effect, so the final value is identical and only computed once. A violation is a hoist whose indentation cannot be reconciled with the surrounding lines/brackets (possible `IndentationError`/wrong nesting), or where the loop body depended on the per-iteration effect.
- **Consequence**: If the block structure is broken, the module fails at import and *every* covering test errors out (0 passed) regardless of the workload; if an in-loop consumer existed, the workload gets faster but results/ordering change and the behaviour tests fail.