← main results | causal dashboard P/R comparison + traces sanity + distributions rubric library (183)

Conceptual rubric library & experiment metrics

183 rubrics · split rubric-gen-full75-v1 (15 train / 60 eval) · generated 2026-08-08 21:06:34 UTC · regenerate with python3 make_browser.py

Experiment metrics

truth source: annotation (evidence-grounded, per rubric-state) · 7188 blind decisions across 14 held-out competitions
familyprecision95% CIrecall95% CItp/fp/fn
conceptual0.139[0.0837, 0.1947]0.478[0.3514, 0.6164]32/198/35
control_generic0.000[0.0, 0.0]——0/52/0
control_shuffled0.000[0.0, 0.0]——0/3/0
control_conceptual0.000[0.0, 0.0]——0/55/0
Controls are the floor: vague criticism and domain-swapped rubrics score zero precision; whatever the real rubrics score above that is earned signal.
ground-truth annotator agreement: 97.8% per group of 4 (reference protocol reports 97%)

Utility: does giving the rubrics to a coding agent help?

paired statesguided betterrevise bettertiesmean advantagemedian advantage
8521-0.3%0.1%
guided = revision with retrieved rubrics; revise = same revision without them (the control). The rubrics' contribution is the difference, paired on the same starting program.

Causal: does removing the flagged pattern improve the score?

fires measuredimprovedcausal precisionmean improvementoutcomes
56200.3570.007266{"measured": 56, "no_patch": 63, "patched_run_failed": 112}
Patch out exactly what the rubric flagged, re-run, re-grade against the official answers. Failed edits are excluded from the denominator; they are facts about the editor.

Rubric library (conceptual / performance errors)

C11 · spec-noncompliance Hardcoded label-parsing stride copied from submission record width

P1 — Hardcoded label-parsing stride copied from submission record width

• Pattern: Detects training-annotation strings parsed by iterating a flat token list with a hardcoded per-record stride that matches the wider prediction/output record layout (including its extra leading score field) rather than the ground-truth record's actual field count, with no divisibility or schema validation on the label tokens.
• Detection procedure:
1. Locate a statement that splits a string value taken from a column of a training/label file (e.g., row["LabelString"].split() or equivalent) into a flat list of tokens.
2. Within the same function or script, locate iteration over that token list in fixed-size chunks using a hardcoded integer literal N (e.g., for i in range(0, len(tokens), N) or slices of the form tokens[i:i+N]).
3. Elsewhere in the same source file, locate serialization of prediction records into a delimited string with exactly N fields per record, where the first field is a confidence/score value (e.g., a format string whose first placeholder is filled from a variable named like conf, score, or a constant such as 1.0).
4. Confirm the label-parsing loop from step 2 indexes chunk positions assuming the same layout as step 3 (e.g., skipping or casting element 0 as a score, reading the class name from index N-1).
5. Confirm there is no check that len(tokens) is divisible by N and no validation that the parsed class-name token belongs to a known category set.
6. PRESENT when steps 1, 2, 3, 4, and 5 all hold: the training-label stride is copied from the output record width instead of derived from the label tokens themselves.
• Predicted impact:
• Add score for C11: Deviates from the specified method — 3
• weight: 3
• confidence: high
• Evidence: for i in range(0, len(tokens), 9) applied to ground-truth records that contain only 8 fields (no leading confidence); from the second record onward every field was shifted, class names were read from numeric positions, per-class box statistics silently fell back to hardcoded defaults, and the run scored the full contrastive gap (1.0 of the competition metric) below an otherwise identical variant that parsed the label schema correctly.
• Applies when: Ground-truth annotations are serialized as delimited strings inside a table column, and the required submission/output format differs from the label format by one or more extra fields (typically a leading confidence value).
• Example:
• Input:
    # label format per record: x y z w l h yaw name  (8 fields)
    dims_by_class = {}
    for _, row in train_df.iterrows():
        tokens = str(row["LabelString"]).split()
        for i in range(0, len(tokens), 9):   # 9 = submission stride (conf first)
            rec = tokens[i:i+9]
            name = rec[8]                     # reads a numeric token, not a class
            dims_by_class.setdefault(name, []).append(
                [float(rec[4]), float(rec[5]), float(rec[6])])
    # later: submission written as "conf x y z w l h yaw name" (9 fields)
• Consequence:
    Every record after the first is misaligned; `name` holds numeric strings, so
    dims_by_class contains no real class keys. Downstream lookups miss and fall
    back to default box sizes for all classes; training data contributes nothing
    and the competition metric drops by the full measured gap (1.0) versus
    correctly parsed labels.
• Counter-example:
• Input:
    LABEL_FIELDS = 8  # x y z w l h yaw name — no confidence in ground truth
    KNOWN = {"LABEL_A", "LABEL_B"}
    tokens = str(row["LabelString"]).split()
    assert len(tokens) % LABEL_FIELDS == 0
    for i in range(0, len(tokens), LABEL_FIELDS):
        rec = tokens[i:i+LABEL_FIELDS]
        name = rec[7]
        if name in KNOWN:
            dims_by_class.setdefault(name, []).append(
                [float(rec[3]), float(rec[4]), float(rec[5])])
• Why it does not fire: the stride matches the label record's own field count, is guarded by a divisibility check, and the parsed class name is validated against a known category set, so steps 4 and 5 fail.
C11 · spec-noncompliance Label parsing with submission-format stride

P1 — Label parsing with submission-format stride

• Pattern: Detects parsing of concatenated fixed-field ground-truth annotation strings using the per-record field stride of the prediction/submission serialization (which includes an extra confidence field) rather than the label format's own field count, so token offsets drift and dimension/class fields are read from the wrong positions.
• Detection procedure:
1. Locate code that reads annotation strings from a training-labels file (e.g., a column of train.csv) and tokenizes them with .split() or equivalent whitespace splitting into a flat token list.
2. Within the same function body, locate a loop stepping through that token list with a literal stride, e.g., range(0, len(t), S) or repeated slices t[i:i+S], where S is an integer literal or a constant assigned an integer literal.
3. Elsewhere in the same program, locate the format string or join expression that serializes prediction records for the output file; count its whitespace-separated fields, including any confidence/score field, call it F.
4. Check whether the parsing code contains, before accumulating parsed values, either (a) a divisibility guard on the token count against the stride (len(t) % S == 0 or equivalent) or (b) a check that the token at the class-name position is non-numeric.
5. PRESENT when S == F, the emitted format at step 3 includes a confidence field that the training labels do not contain, and neither guard from step 4 exists.
• Predicted impact:
• Add score for C11: 3
• weight: 3
• confidence: high
• Evidence: Training label records contained 8 fields per box but were stepped through with the 9-field submission stride (range(0, len(t), 9) where record 0 onward drifts by one field per record); the class-name position read a numeric token, per-class size statistics silently stayed empty, and the program fell back to hardcoded default box dimensions. Between the defective and corrected versions, the overlap-based competition metric differed by the full score range (1.0).
• Applies when: Ground-truth annotations are serialized as concatenated fixed-field record strings inside a single cell/field, and the required output format differs from the label format by one prepended field (typically a confidence score), and the program derives statistics from parsed labels.
• Example:
• Input:
    STRIDE = 9  # conf x y z w l h yaw class  (same layout as submission)
    dims = defaultdict(list)
    for s in train_df["gt_string"].fillna(""):
        t = s.split()
        for i in range(0, len(t) - STRIDE + 1, STRIDE):
            w, l, h = float(t[i+4]), float(t[i+5]), float(t[i+6])
            dims[t[i+8]].append((w, l, h))
    med = dims.get("LABEL_A") or [(1.9, 4.5, 1.7)]   # silent default
    box_w, box_l, box_h = np.median(np.array(med), axis=0)
• Consequence:
    Label records actually have 8 fields (no confidence), so offsets drift by
    one field per record: width/length/height are read from position/yaw
    tokens and the "class" slot holds a numeric token. dims["LABEL_A"] stays
    empty, hardcoded defaults replace learned sizes, predicted box overlap
    (IoU) drops, and the mAP-based score falls (measured gap: full metric
    range, 1.0, versus the correctly parsed version). Exit code is 0.
• Counter-example:
• Input:
    STRIDE = 8  # x y z w l h yaw class  (label layout, no confidence)
    dims = defaultdict(list)
    for s in train_df["gt_string"].fillna(""):
        t = s.split()
        assert len(t) % STRIDE == 0
        for i in range(0, len(t), STRIDE):
            cls = t[i+7]
            if cls.replace(".", "", 1).lstrip("-").isdigit():
                continue
            dims[cls].append((float(t[i+3]), float(t[i+4]), float(t[i+5])))
• Why it does not fire: the stride matches the label layout (one fewer field than the confidence-bearing submission format) and both a divisibility guard and a non-numeric class-token check protect the accumulation, so steps 4–5 fail.
C11 · spec-noncompliance Label strings parsed with prediction-format field stridegithub_occurrence

P1 — Label strings parsed with prediction-format field stride

• Pattern: Detects parsing of ground-truth concatenated-record label strings using the per-object field count of the prediction/submission format (which includes a leading confidence field), so every parsed field is shifted by one and the class token is read from a position holding a numeric value.
• Detection procedure:
1. Locate a statement that reads a ground-truth/training annotations file (e.g., pd.read_csv("train.csv") or equivalent) and, within the same function or script, calls .split() (or equivalent tokenization) on a string column whose value is a whitespace-joined sequence of fixed-length per-object records.
2. Find the fixed group size K used to partition those tokens: a slice of the form toks[i:i+K] inside a loop for i in range(0, len(toks), K), a reshape to (-1, K), or equivalent, where K is a literal integer.
3. Find where the same program formats its own prediction strings for output; count the fields written per object (call it M) and confirm the first written field is a confidence/score value formatted as a float.
4. Alternatively to step 3, check whether the label-parsing code converts the FIRST token of each K-sized group to a float treated as a confidence/score.
5. PRESENT when K equals M (or step 4 holds): the ground-truth strings are partitioned with a stride that assumes a leading confidence field, so the token used as the class name comes from a position that holds a numeric field in the true label layout.
• Predicted impact:
• Add score for C11 (Deviates from the specified method): 3
• weight: 3
• confidence: high
• Evidence: Label strings were partitioned with toks[i:i+9] although the ground-truth record has 8 fields (no confidence); the class key became a numeric string, so per-class size statistics were empty and the program fell back to hardcoded box dimensions. In a measured contrastive pair, the correctly parsed variant scored 1.0 higher on the competition metric than the shifted variant.
• Applies when: Ground-truth labels and predictions share a concatenated per-object string format that differs by exactly one leading field (a confidence present only in predictions), and the program learns statistics (class frequencies, per-class size priors) from the parsed labels.
• Example:
• Input:
    df = pd.read_csv("train.csv")
    dims_by_cls = {}
    for s in df["col_a"].fillna(""):
        toks = s.split()
        for i in range(0, len(toks), 9):      # 9 = prediction stride
            obj = toks[i:i+9]
            cls = obj[8]                       # actually yaw in label layout
            dims_by_cls.setdefault(cls, []).append(
                [float(obj[4]), float(obj[5]), float(obj[6])])
• Consequence:
    Every class key is a numeric string like "-1.5708"; lookups for real class
    names return nothing, so hardcoded default box sizes are used for all
    detections. Detection metric drops sharply versus correct parsing
    (observed gap: 1.0 on the competition metric between the two variants).
• Counter-example:
• Input:
    df = pd.read_csv("train.csv")
    dims_by_cls = {}
    for s in df["col_a"].fillna(""):
        toks = s.split()
        for i in range(0, len(toks), 8):      # 8 = ground-truth stride
            obj = toks[i:i+8]
            cls = obj[7]
            assert not cls.lstrip("-").replace(".", "").isdigit()
            dims_by_cls.setdefault(cls, []).append(
                [float(obj[3]), float(obj[4]), float(obj[5])])
• Why it does not fire: the label stride (8) differs from the prediction field count (9), no leading token is parsed as a confidence float, and the class token is validated as non-numeric.
C12 · silent-degradation Silent empty-prediction fallback on empty training setverified_trace · effect +0.0004

P1 — Silent empty-prediction fallback on empty training set

• Pattern: Detects a training/prediction stage wrapped in a conditional on the assembled training data being non-empty, whose else-branch fills the output rows with empty or constant placeholder predictions instead of raising, so a failed upstream data-assembly step still produces a complete output file.
• Detection procedure:
1. Locate a conditional whose test checks the size of a training data collection, e.g. if len(train_rows) > 0, if X_train.shape[0] > 0, if train_features: or equivalent, where the tested variable is populated by earlier assembly code (loops, joins, merges, or file parsing) in the same script.
2. Confirm the true branch of that conditional contains a model-fitting call and a prediction step (a call whose result contributes to the per-row output values).
3. Confirm the false branch (an else, or fall-through with a pre-initialized default) assigns an empty string, empty list, zero, or other constant as the prediction value for every output row, and contains no raise, assert, sys.exit with nonzero status, or logging-then-abort construct.
4. Confirm the script subsequently writes the per-row predictions to an output file (e.g. to_csv, json.dump, or equivalent) regardless of which branch executed.
5. PRESENT when all of steps 1–4 hold: the empty-training case is absorbed into a fully written placeholder output with no failure signal.
• Predicted impact:
• Add score for C12 (Failure absorbed into a worse result): 3
• weight: 3
• confidence: high
• Evidence: An else: branch appending {'Id': token, 'PredictionString': ''} for every test row when the assembled training list was empty; the run exited cleanly and wrote a structurally valid submission that scored 0 on the competition metric, while a comparable script that surfaced its assembly counts scored a full 1.0 higher.
• Applies when: A script assembles training data through a multi-step pipeline (parsing, joining, or filtering) and then writes a per-row prediction file consumed downstream, where an empty training set indicates an upstream bug rather than a legitimate state.
• Example:
• Input:
    train_rows = []
    for rec in records:
        if rec['token'] in label_map:
            train_rows.append(build_features(rec))

    if len(train_rows) > 0:
        model.fit(stack(train_rows), targets)
        preds = [format_pred(model, r) for r in test_rows]
    else:
        preds = ['' for _ in test_rows]

    pd.DataFrame({'Id': ids, 'Pred': preds}).to_csv('submission.csv', index=False)
• Consequence:
    A key mismatch in the assembly loop yields zero training rows; the run
    exits 0 and writes a valid-looking file of empty predictions, dropping
    the evaluation metric to its floor (0) instead of failing, and hiding
    the join bug from the operator.
• Counter-example:
• Input:
    train_rows = []
    for rec in records:
        if rec['token'] in label_map:
            train_rows.append(build_features(rec))

    print("assembled training rows:", len(train_rows))
    if len(train_rows) == 0:
        raise RuntimeError("no training rows assembled — check join keys")

    model.fit(stack(train_rows), targets)
    preds = [format_pred(model, r) for r in test_rows]
    pd.DataFrame({'Id': ids, 'Pred': preds}).to_csv('submission.csv', index=False)
• Why it does not fire: the empty-training branch raises with a diagnostic instead of writing placeholder predictions, so the upstream failure surfaces loudly rather than being absorbed into a blank output.
C12 · silent-degradation Silent empty-submission fallback on prediction failure

P1 — Silent empty-submission fallback on prediction failure

• Pattern: Detects a broad exception handler around the prediction pipeline whose recovery path writes an output file with the prediction column set to an empty string, so total prediction failure still produces a valid-looking but empty result.
• Detection procedure:
1. Find a try/except block whose try body invokes the main prediction routine (a call to a function that produces per-example predictions, or a loop that appends predictions per example).
2. Confirm the except clause catches Exception or a bare except, i.e., it is not restricted to a narrow, expected error type.
3. Within that except block (or a function it calls), find an assignment of an empty string, empty list, or equivalent empty value to the prediction column of an output frame (e.g., sub["col_a"] = "" or building rows with "" predictions), followed by a call that writes that frame to disk (to_csv or equivalent).
4. Confirm the except block does not re-raise, call sys.exit with a nonzero code, or otherwise abort after writing the file.
5. The pattern is PRESENT when steps 1–4 all hold: a broad handler catches prediction failure, substitutes empty predictions, writes the file, and lets the program terminate as if successful.
• Predicted impact:
• Add score for C12: Failure absorbed into a worse result: 3
• weight: 3
• confidence: high
• Evidence: A script wrapped main() in except Exception: and, on failure, executed sub["PredictionString"] = "" then wrote the file; the run exited cleanly in under 10 seconds and scored exactly the empty-submission baseline (0.0), while a version without the silent fallback scored 1.0 higher on the competition metric.
• Applies when: A script produces a per-example prediction file (free-text or structured prediction column) and wraps the prediction computation in exception handling.
• Example:
• Input:
    if __name__ == "__main__":
        try:
            main()
        except Exception as e:
            print("FATAL:", repr(e))
            sub = pd.read_csv("sample_submission.csv")
            sub["col_a"] = ""
            sub.to_csv("submission.csv", index=False)
            print("wrote fallback empty submission")
• Consequence:
    Any systematic bug in main() (e.g., a bad path or dtype) silently converts
    every prediction to an empty string; the run exits 0 in seconds and the
    metric collapses from its achievable value to the empty-submission
    baseline of 0.0, with no error surfaced.
• Counter-example:
• Input:
    if __name__ == "__main__":
        try:
            main()
        except Exception as e:
            print("FATAL:", repr(e))
            raise
        sub = pd.read_csv("submission.csv")
        n_empty = (sub["col_a"].fillna("") == "").sum()
        assert n_empty < 0.5 * len(sub), "too many empty predictions"
• Why it does not fire: the handler re-raises instead of writing empty predictions, and the script validates that most rows are non-empty before accepting the output.
C12 · silent-degradation Unverified template join with constant fillverified_trace · effect +0.1093

P1 — Unverified template join with constant fill

• Pattern: Detects a submission assembled by left-joining computed predictions onto a template id column and then filling unmatched rows with a constant, with no subsequent check that the join matched a non-trivial fraction of rows.
• Detection procedure:
1. Find a merge/join call (e.g. .merge(...), .join(...), or equivalent) whose left side is a frame derived from a template or sample output file and whose join mode is a left join (how="left" or default left-preserving join) on an id-like key column.
2. Within the same function body, find a subsequent call on the merged frame's prediction column that replaces missing values with a constant (e.g. .fillna(0), .fillna(-1), or equivalent), followed by writing the frame to an output file.
3. Between step 2's fill and the file write, check for any statement that inspects null counts or match coverage of the merged prediction column (e.g. assert, an if on .isna().sum(), .notna().mean(), a raised error, or equivalent).
4. PRESENT if steps 1 and 2 are found and no coverage-checking statement from step 3 exists between the merge and the file write.
• Predicted impact:
• Add score for C12 (Failure absorbed into a worse result): 2
• weight: 2
• confidence: high
• Evidence: s2 = sub[["Id"]].merge(out, on="Id", how="left") followed immediately by s2["Predicted"] = s2["Predicted"].fillna(0).astype(int) and a write to disk; an id dtype/value mismatch converted every row to the constant fallback, and the run scored the trivial baseline while the paired run without the blind fill scored 1.0 higher on the evaluation metric.
• Applies when: The program writes a prediction/output file whose row ids must match a provided template or sample file, and predictions are aligned to that template by a join.
• Example:
• Input:
    out = pd.DataFrame({"Id": test_ids, "Predicted": preds})
    # align to template ordering / ids
    s2 = template[["Id"]].merge(out, on="Id", how="left")
    s2["Predicted"] = s2["Predicted"].fillna(0).astype(int)
    s2.to_csv("submission.csv", index=False)
• Consequence:
    template "Id" is string dtype while test_ids are integers, so zero rows
    match; every "Predicted" value becomes the constant 0. The file is written
    successfully and the run exits 0, but the class-averaged evaluation score
    drops from the model's actual level to the trivial baseline (~0).
• Counter-example:
• Input:
    out = pd.DataFrame({"Id": test_ids, "Predicted": preds})
    s2 = template[["Id"]].merge(out, on="Id", how="left")
    matched = s2["Predicted"].notna().mean()
    if matched < 0.99:
        raise ValueError(f"join matched only {matched:.1%} of template rows")
    s2["Predicted"] = s2["Predicted"].fillna(0).astype(int)
    s2.to_csv("submission.csv", index=False)
• Why it does not fire: A coverage check on the merged prediction column exists between the merge and the file write, so a silent all-constant output cannot occur.
C12 · silent-degradation Defaulting label lookup fabricates classes for unmatched idstrace_observed

P1 — Defaulting label lookup fabricates classes for unmatched ids

• Pattern: Detects training labels assigned via a dictionary lookup with a constant fallback value keyed on a string-formatted identifier, so samples whose id fails to match a key are silently given a fabricated class label instead of being flagged as missing.
• Detection procedure:
1. Locate a dictionary or mapping built from a label file or label column (e.g. via dict(zip(...)), Series.to_dict(), a dict comprehension over label rows, or equivalent).
2. Within the same script, find a call to .get(key, default) on that mapping (or a mapping[key] if key in mapping else default expression) where default is a literal constant (such as 0, -1, or 0.5) and the result is appended to or assigned into the training target array or column used later in a fitting call.
3. Verify the key argument at step 2 is derived from a formatting or type-conversion operation on an identifier — any of str(...), .zfill(...), int(...), .split(...), .strip(...), slicing, or string concatenation — applied to a filename, directory name, or id column.
4. Confirm no subsequent statement before the fitting call checks membership of ids in the mapping, asserts full coverage, or raises/logs on missing keys.
5. PRESENT when steps 1–4 all hold: a constant-defaulted lookup on a formatted id feeds the training target with no coverage check before fitting.
• Predicted impact:
• Add score for C12 (Failure absorbed into a worse result): 3
• weight: 3
• confidence: high
• Evidence: Training labels attached with label_map.get(str(case_id), 0) keyed on string-formatted ids; an id-format mismatch mislabeled affected rows as class 0, and the run with this construct scored 0.154 worse on the evaluation metric than an equivalent pipeline that joined labels with an explicit merge and coverage handling.
• Applies when: Labels live in a separate file or table keyed by an identifier, and training rows are enumerated from files, directories, or another table whose ids must be reformatted to match the label keys before fitting a model.
• Example:
• Input:
    labels_df = pd.read_csv("train.csv")
    label_map = dict(zip(labels_df["id"].astype(str), labels_df["LABEL_A"]))

    X, y = [], []
    for case_dir in sorted(os.listdir(train_dir)):
        feats = extract_features(case_dir)
        X.append(feats)
        y.append(label_map.get(str(int(case_dir)), 0))

    model.fit(np.array(X), np.array(y))
• Consequence:
    Ids in the label file are zero-padded strings while the lookup key is
    de-padded via int(); every sample misses the map and receives label 0,
    so the fitted model predicts near-constant outputs and the ranking
    metric drops by ~0.15 versus a correctly joined label set, with exit
    code 0 and no warning emitted.
• Counter-example:
• Input:
    labels_df = pd.read_csv("train.csv")
    labels_df["id"] = labels_df["id"].astype(str).str.zfill(5)
    ids = pd.DataFrame({"id": [str(d).zfill(5) for d in os.listdir(train_dir)]})
    joined = ids.merge(labels_df, on="id", how="left")
    assert joined["LABEL_A"].notna().all(), "missing labels for some ids"
    X = np.stack([extract_features(i) for i in joined["id"]])
    model.fit(X, joined["LABEL_A"].values)
• Why it does not fire: labels are attached via an explicit merge followed by an assertion that every training id received a real label, so a format mismatch surfaces immediately instead of silently defaulting to a fabricated class.
C12 · silent-degradation Constant placeholder prediction string for every test row

P1 — Constant placeholder prediction string for every test row

• Pattern: Detects a submission-assembly loop that assigns an empty or constant literal to the prediction field on every branch, so the serialized output rows never incorporate any value computed from earlier fitting or feature work in the script.
• Detection procedure:
1. Locate the loop that iterates over test identifiers and appends rows (dict, tuple, or list) destined for the output file, where one field holds the prediction payload (e.g., a key like PredictionString, prediction, or the second column of the row).
2. Within that loop body, collect every assignment to the prediction payload variable. Check whether each assignment's right-hand side is a string, numeric, or empty-container literal (e.g., "", 0, []) with no reference to any variable defined outside the loop from a fitting step, statistic, or per-sample data lookup.
3. Confirm that no branch of the loop assigns the payload from a call, subscript, or expression involving a model, statistics dict, or feature variable computed earlier in the same file.
4. Confirm the rows built in this loop are written to the output artifact (e.g., via to_csv or equivalent serialization).
5. PRESENT if all assignments to the prediction payload inside the row-assembly loop are constant literals (steps 2–3) and the loop's rows are what get serialized (step 4).
• Predicted impact:
• Add score for C12 (Failure absorbed into a worse result): 3
• weight: 3
• confidence: high
• Evidence: A row-assembly loop assigned pred_string = "" on both branches of its conditional before appending {'Id': sample_id, 'PredictionString': pred_string}; the run completed cleanly but scored exactly 0.0 on the competition metric, while a sibling script that filled the payload from fitted per-class statistics scored 1.0 higher on the same metric.
• Applies when: The program produces a per-row prediction file (detection or structured-output style, e.g., an ID plus a prediction string or value column) and earlier code in the same script computes models, statistics, or features intended to drive those predictions.
• Example:
• Input:
    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        if sample_id in test_lookup:
            # avoid false positives by predicting nothing
            pred_string = ""
        else:
            pred_string = ""
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv(OUT_PATH, index=False)
• Consequence:
    Submission is entirely blank: metric collapses to the empty-baseline score
    (0.0 here), a full 1.0 below a script emitting crude prior-based guesses,
    despite all earlier fitting/statistics code running to completion.
• Counter-example:
• Input:
    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        if sample_id in test_lookup:
            pred_string = build_pred(stats, test_lookup[sample_id])
        else:
            pred_string = ""  # unmatched ids get no prediction
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv(OUT_PATH, index=False)
• Why it does not fire: One branch derives the payload from fitted statistics and per-sample data, so not all assignments to the prediction field are constant literals; the empty string is only a fallback for unmatched identifiers.
C12 · silent-degradation Silent empty-prediction fallback under broad exception handling

P1 — Silent empty-prediction fallback under broad exception handling

• Pattern: Detects a per-example inference loop (or its enclosing pipeline) whose broad exception handler substitutes an empty or trivial prediction and proceeds to write the output file, with no count or threshold on how many examples fell back, so a systematic bug yields a valid-looking but degenerate result.
• Detection procedure:
1. Locate a loop that iterates over evaluation or test identifiers and produces one prediction value per iteration that is later written to an output file (e.g., via to_csv, csv.writer, or equivalent).
2. Within that loop body, or in a try block wrapping the entire pipeline whose handler still writes the output file, find an except Exception clause or a bare except clause.
3. In that handler's body, verify the prediction for the affected sample(s) is set to an empty or constant default — a literal "", an empty list, or a column-wide assignment of an empty value — and the handler contains no raise statement.
4. Search the module for any counter incremented in that handler that is later compared against a threshold followed by raise, sys.exit with a nonzero code, or assert; confirm no such check exists before the output file is written.
5. PRESENT if steps 2–4 all hold: broad catch, empty/default substitution without re-raise, and no fallback-rate guard before writing the file.
• Predicted impact:
• Add score for C12: Failure absorbed into a worse result: 2
• weight: 2
• confidence: high
• Evidence: A per-example detection pipeline wrapped in except Exception: that substituted "" predictions, plus a top-level handler assigning sub["pred_col"] = "" and writing the file anyway; the run finished in ~10 seconds over thousands of samples with the metric at 0.0 versus a nonzero score for the working variant — a 100% relative gap with exit code 0.
• Applies when: The program performs per-example inference and writes a prediction file where empty or missing predictions are syntactically valid output.
• Example:
• Input:
    rows = []
    for sid in sample_ids:
        try:
            pts = load_cloud(sid)
            pred = detect(pts)          # a bug here raises on every sample
        except Exception:
            pred = ""                   # silently substitute empty prediction
        rows.append((sid, pred))
    pd.DataFrame(rows, columns=["Id", "pred_col"]).to_csv("out.csv", index=False)
• Consequence:
    Every iteration takes the fallback branch; the file contains only empty
    predictions and the metric collapses to the trivial baseline (0.0 vs the
    nonzero score of the working pipeline) while the run exits 0 in seconds.
• Counter-example:
• Input:
    rows, n_fail = [], 0
    for sid in sample_ids:
        try:
            pred = detect(load_cloud(sid))
        except ValueError:              # narrow, expected case only
            pred = ""
            n_fail += 1
        rows.append((sid, pred))
    if n_fail > 0.05 * len(sample_ids):
        raise RuntimeError(f"{n_fail} fallbacks; refusing to write output")
    pd.DataFrame(rows, columns=["Id", "pred_col"]).to_csv("out.csv", index=False)
• Why it does not fire: the handler catches only a narrow expected exception and a fallback-rate guard aborts before writing the file when failures are widespread, so a systematic bug cannot silently produce an all-empty submission.
C12 · silent-degradation Defaulted label lookup fabricates training labels on key mismatchtrace_observed

P1 — Defaulted label lookup fabricates training labels on key mismatch

• Pattern: Detects a supervised training-data assembly loop that retrieves each example's target label from a mapping keyed by identifier using a lookup that substitutes a hardcoded constant class value when the identifier is absent, so mismatched keys silently enter the training set with fabricated labels.
• Detection procedure:
1. Locate a dictionary (or Series/mapping) whose values are target labels, built from a label table such as train.csv (e.g. via dict(zip(...)), set_index(...).to_dict(), or equivalent).
2. Locate a loop that iterates over identifiers obtained from a source other than that label table (e.g. directory listing, file glob, separate index file) and builds feature rows for training.
3. Within that loop body, find a label retrieval from the mapping of step 1 using .get(key, <literal>) where the second argument is a literal constant, or a try/except KeyError block whose handler assigns a literal constant to the label variable.
4. Confirm the retrieved label value is appended to (or stored in) the collection later passed as the target argument y to a model-fitting call (e.g. model.fit(X, y) or equivalent).
5. PRESENT when all of steps 1–4 hold: labels for training rows can come from the hardcoded default rather than the label table, and no check skips or aborts on the missing key.
• Predicted impact:
• Add score for C12 (Failure absorbed into a worse result): 3
• weight: 2
• confidence: high
• Evidence: A training loop used labels.get(case_id, 0) to attach targets to feature rows built from a directory listing; identifier formatting mismatches caused rows to receive the default class, and the contrastive run that joined labels strictly scored 0.109 higher on the held-out metric with no error emitted by the weaker run.
• Applies when: Supervised training where target labels live in a separate table and are joined to feature rows by an identifier key, especially when identifiers come from filenames or directories and may differ in padding, case, or type from the label table's keys.
• Example:
• Input:
    labels = dict(zip(df['id'], df['LABEL_A']))
    X, y = [], []
    for case_dir in sorted(train_dir.iterdir()):
        case_id = case_dir.name
        feats = extract_features(case_dir)
        X.append(feats)
        y.append(labels.get(case_id, 0))  # missing key -> fabricated label
    model.fit(np.array(X), np.array(y))
• Consequence:
    Every identifier that fails to match (e.g. '00005' vs 5) is trained with
    label 0 regardless of its true class; the proportion of mislabeled rows
    grows with the mismatch rate, pushing held-out score toward chance —
    observed as a 0.109 drop in the evaluation metric versus a strict join,
    with the run completing cleanly and printing no warning.
• Counter-example:
• Input:
    labels = dict(zip(df['id'], df['LABEL_A']))
    X, y = [], []
    for case_dir in sorted(train_dir.iterdir()):
        case_id = case_dir.name
        if case_id not in labels:
            continue
        X.append(extract_features(case_dir))
        y.append(labels[case_id])
    assert len(y) == len(df), "label join lost rows"
    model.fit(np.array(X), np.array(y))
• Why it does not fire: the lookup is guarded by a membership check that skips unmatched identifiers and a count assertion, so no default label is ever fabricated into the training targets.
C12 · silent-degradation Silent empty-prediction fallback on broad exceptiongithub_occurrence

P1 — Silent empty-prediction fallback on broad exception

• Pattern: Detects a broad exception handler around the prediction-generation step whose recovery path fills the prediction field with an empty string and still writes the output file, with no subsequent check that a non-trivial fraction of predictions are non-empty.
• Detection procedure:
1. Locate a try block enclosing the code that computes prediction values (a per-example loop or a call to the main prediction routine), paired with a handler written as except Exception or a bare except (or a handler catching a comparably broad base class).
2. Inside that handler's body (or in a fallback branch it triggers), find an assignment or append where the prediction output — a column of a results frame, or the element added to the list of prediction strings — is set to the empty string literal "" or an equivalent trivial constant.
3. Confirm that after this handler executes, a call that persists the output file (to_csv, to_json, an opened-file write, or equivalent) is still reached on that code path.
4. Scan all statements between the handler and the persisting call: there is no assert, raise, conditional exit, or comparison that measures the count or fraction of empty predictions against a threshold.
5. PRESENT when steps 1–4 all hold: broad catch, empty-value substitution, output still written, and no emptiness check before writing.
• Predicted impact:
• Add score for C12: Failure absorbed into a worse result: 2
• weight: 2
• confidence: high
• Evidence: An except Exception handler around the full pipeline set sub["pred"] = "" for every row and wrote the file as a "fallback empty submission"; the run exited cleanly but scored the trivial-baseline value, a full 1.0 below a comparable version that emitted real predictions.
• Applies when: The program produces a per-example prediction output (file or frame) where an empty prediction value is syntactically valid, and prediction generation is wrapped in exception handling.
• Example:
• Input:
    if __name__ == "__main__":
        try:
            main()
        except Exception as e:
            print("FATAL:", repr(e))
            sub = pd.read_csv("sample_submission.csv")
            sub["pred"] = ""
            sub.to_csv("submission.csv", index=False)
            print("wrote fallback empty submission")
• Consequence:
    Any failure anywhere in main() silently yields a file of all-empty
    predictions; the evaluation metric collapses to the trivial baseline
    (score 0 vs. ~1.0 achievable) while the process exits with status 0.
• Counter-example:
• Input:
    n_fail = 0
    preds = []
    for ex in examples:
        try:
            preds.append(predict_one(ex))
        except ValueError:
            n_fail += 1
            preds.append("")
    assert n_fail / len(examples) < 0.05, "too many empty predictions"
    df["pred"] = preds
    df.to_csv("submission.csv", index=False)
• Why it does not fire: the handler catches only a specific expected exception and the write is gated by an assertion that the fraction of empty predictions stays below a threshold, so a degenerate all-empty output cannot be persisted silently.
C12 · silent-degradation Unprobed metadata-joined paths with silent per-file fallbackverified_trace · effect +2.3536

P1 — Unprobed metadata-joined paths with silent per-file fallback

• Pattern: Detects batch feature extraction over file paths assembled by joining a hardcoded directory prefix with per-record metadata fields, where each per-file read is wrapped in an error handler that substitutes a default value, and no existence check is performed on any constructed path before the batch loop begins.
• Detection procedure:
1. Find an expression that builds a file path by combining a string literal directory prefix with a field taken from a metadata structure (e.g., os.path.join(DIR, rec["file_name"]), string concatenation, or an f-string containing both a literal directory segment and a record field), and this expression's results are collected into a list or iterated for reading.
2. Within the loop or worker function that opens each constructed path (via an image/file open call or equivalent), confirm the open/read is inside a try/except (or equivalent handler) whose handler returns or appends a constant default such as a zero array, None-then-fill, or a fixed placeholder, rather than re-raising or aborting.
3. Search the source between the path-construction site and the start of the batch loop for any call to os.path.exists, os.path.isfile, Path.exists, glob, os.listdir on the prefix, or an assert/conditional that verifies at least one constructed path (or a fraction of them) resolves on disk; confirm no such probe exists.
4. The pattern is PRESENT when steps 1 and 2 hold and step 3 finds no probe.
• Predicted impact:
• Add score for C12 (Failure absorbed into a worse result): 2
• weight: 2
• confidence: high
• Evidence: Paths built as os.path.join(DATA, rec["file_name"]) were consumed by an extractor whose per-file except returned np.zeros(dim); the directory layout contained an extra intermediate folder, so every read failed silently, all feature vectors were zeros, every prediction collapsed to one class, and the class-averaged metric dropped from a positive score to ~0.
• Applies when: The program processes an image- or file-based dataset whose file locations are assembled from metadata records plus an assumed directory prefix, and features or predictions are computed per file in a batch loop.
• Example:
• Input:
    DATA = "/data/prepared/public"
    recs = json.load(open(os.path.join(DATA, "meta.json")))["images"]
    paths = [os.path.join(DATA, r["file_name"]) for r in recs]

    def feat(p):
        try:
            img = Image.open(p).convert("RGB").resize((32, 32))
            return np.asarray(img, dtype=np.float32).ravel()
        except Exception:
            return np.zeros(3072, dtype=np.float32)

    X = np.stack([feat(p) for p in paths])
• Consequence:
    When the on-disk layout includes an extra subdirectory not reflected in the
    prefix, every open fails and is replaced by a zero vector; all rows of X are
    identical, predictions collapse to a single class, and the class-averaged
    evaluation metric falls to ~0 while the run still exits with status 0.
• Counter-example:
• Input:
    DATA = "/data/prepared/public"
    recs = json.load(open(os.path.join(DATA, "meta.json")))["images"]
    paths = [os.path.join(DATA, r["file_name"]) for r in recs]
    if not os.path.exists(paths[0]):
        alt = os.path.join(DATA, "images")
        paths = [os.path.join(alt, r["file_name"]) for r in recs]
    assert sum(os.path.exists(p) for p in paths[:100]) > 90, "layout mismatch"

    def feat(p):
        img = Image.open(p).convert("RGB").resize((32, 32))
        return np.asarray(img, dtype=np.float32).ravel()

    X = np.stack([feat(p) for p in paths])
• Why it does not fire: A constructed path is probed with an existence check before the batch loop, an alternate layout is tried, and a minimum-fraction assertion aborts the run instead of silently absorbing failures into default vectors.
C12 · silent-degradation Blank-output fallback that masks fatal pipeline failuregithub_occurrence

P1 — Blank-output fallback that masks fatal pipeline failure

• Pattern: Detects a top-level exception handler around the script's main pipeline that, on any caught exception, writes a syntactically valid output artifact with empty or placeholder prediction values and then lets the process exit successfully.
• Detection procedure:
1. Locate the script entry point: an if __name__ == "__main__": block, or the top-level call that invokes the main pipeline function.
2. Check that the invocation at step 1 is enclosed in a try block whose handler is except Exception (or broader, including a bare except:).
3. Within that handler's body, find a statement that assigns an empty string, empty list, or constant placeholder to the prediction field(s) of a result table or record set (e.g., sub["col_a"] = "" or equivalent), followed by a write of that table to the designated output file (to_csv, to_json, a file write call, or equivalent).
4. Confirm that after the write at step 3 completes, the handler contains no raise statement and no process-exit call with a nonzero status on that success path (a nonzero exit only inside a nested handler for the fallback write itself does not count).
5. The pattern is PRESENT when steps 2, 3, and 4 all hold: any pipeline exception is absorbed, a blank-prediction artifact is written, and the process exits with status zero.
• Predicted impact:
• Add score for C12: Failure absorbed into a worse result: 2
• weight: 2
• confidence: high
• Evidence: An entry point wrapped in try: main() except Exception: whose handler executed sub["PredictionString"] = "" and wrote the file; a fatal error in an early pipeline stage produced a valid-looking submission scoring at the trivial zero baseline, with exit status 0 and no distinguishable failure signal, while an equivalent run without the absorber produced real predictions (metric difference of the full score range).
• Applies when: The program's purpose is to produce a scored output artifact (e.g., a predictions file) and it has a single script entry point invoking the pipeline.
• Example:
• Input:
    if __name__ == "__main__":
        try:
            main()
        except Exception as e:
            log("FATAL:", repr(e))
            sub = pd.read_csv("sample_submission.csv")
            sub["col_a"] = ""
            sub.to_csv("submission.csv", index=False)
            log("wrote fallback empty submission")
• Consequence:
    A failure in any pipeline stage still yields a valid submission file and
    exit code 0; the evaluation metric is pinned at the trivial all-blank
    baseline (near the minimum of the score range) instead of the score real
    predictions would earn, and monitoring cannot distinguish the failed run
    from a successful one.
• Counter-example:
• Input:
    if __name__ == "__main__":
        try:
            main()
        except Exception as e:
            log("FATAL:", repr(e))
            sub = pd.read_csv("sample_submission.csv")
            sub["col_a"] = ""
            sub.to_csv("submission_DEGRADED.csv", index=False)
            log("wrote degraded fallback artifact")
            sys.exit(1)
• Why it does not fire: the handler writes the fallback to an explicitly marked degraded artifact and then exits with a nonzero status, so the failure propagates as a visible error rather than masquerading as a completed run.
C12 · silent-degradation Half-precision reductions with artifact maskingtrace_observed

P1 — Half-precision reductions with artifact masking

• Pattern: Detects statistical reductions (mean, sum, std, max) computed directly on arrays kept in half precision after loading, with the resulting non-finite values later replaced by a finite-fill sanitizer instead of upcasting before the reduction.
• Detection procedure:
1. Find a call that loads array data from disk (np.load or equivalent) whose result is assigned to a variable, where between that assignment and the first reduction there is no cast widening the dtype (no .astype(np.float32), .astype(np.float64), .astype("float32"), or a dtype= argument specifying a 32- or 64-bit float on the reduction itself), and the data is known or declared to be 16-bit float (a comment, a dtype check, or an explicit float16 mention in the same file, or absence of any dtype handling in a pipeline whose loader stores float16).
2. Within the same function body or module scope, find at least one reduction over that variable: a call to .mean(), .sum(), .std(), .max(), or the equivalent module-level functions (np.mean(arr), etc.) applied to the loaded variable or a slice of it.
3. Find a subsequent call that replaces non-finite values with finite constants — np.nan_to_num(...) or equivalent (e.g. assignment through a mask of np.isinf/np.isnan to a constant) — applied to the feature matrix built from those reductions.
4. PRESENT when all three hold: half-precision arrays are reduced without a widening cast between load and reduction, and the downstream feature matrix is sanitized with a finite-fill replacement.
• Predicted impact:
• Add score for C12: Failure absorbed into a worse result: 3
• weight: 3
• confidence: high
• Evidence: Reductions such as arr.sum() and arr.std() on float16 arrays overflowed to inf, which np.nan_to_num(X, posinf=0.0, neginf=0.0) silently zeroed; the classifier trained on near-constant features and held-out AUC fell to roughly chance (0.492 vs 0.513 with an explicit float32 upcast before reducing).
• Applies when: Array data is stored or loaded in 16-bit floating point and handcrafted summary statistics over those arrays are used as model features.
• Example:
• Input:
    def load_and_process(ids, folder):
        feats = []
        for i in ids:
            arr = np.load(os.path.join(folder, i + ".npy"))  # stored float16
            feats.append([arr.mean(), arr.sum(), arr.std(), arr.max()])
        return np.array(feats)

    X = load_and_process(train_ids, train_dir)
    X = np.nan_to_num(X, nan=0.0, posinf=0.0, neginf=0.0)
    model.fit(scaler.fit_transform(X), y)
• Consequence:
    Sum/std reductions overflow float16's ~65504 max to inf, nan_to_num zeroes
    them, and whole feature columns become constant zeros; held-out AUC drops
    from ~0.513 to ~0.492 (at or below chance) versus the float32-upcast run.
• Counter-example:
• Input:
    def load_and_process(ids, folder):
        feats = []
        for i in ids:
            arr = np.load(os.path.join(folder, i + ".npy")).astype(np.float32)
            feats.append([arr.mean(), arr.sum(), arr.std(), arr.max()])
        return np.array(feats)

    X = load_and_process(train_ids, train_dir)
    X = np.nan_to_num(X, nan=0.0)  # only genuinely undefined values
    model.fit(scaler.fit_transform(X), y)
• Why it does not fire: The loaded array is widened to float32 before any reduction, so no overflow artifacts exist for the sanitizer to mask (step 1 fails).
C12 · silent-degradation Constant zero-vector substituted for failed record decoding

P1 — Constant zero-vector substituted for failed record decoding

• Pattern: Detects a broad per-record exception handler in a feature-extraction loop that appends a constant all-zero placeholder vector to the same feature collection as successfully decoded records, which is subsequently used for model fitting or prediction.
• Detection procedure:
1. Locate a loop (or per-record function called from a loop) that decodes or transforms individual records into numeric feature vectors and appends the result to a list or array (e.g., features.append(feat)).
2. Within the same loop or function body, find a try/except where the except clause is bare or catches Exception (or a similarly broad class).
3. Inside that except block, find an append (or assignment into the same collection) of a constant-valued vector — a call to np.zeros(...), [0] * n, np.full(..., 0), or equivalent constant construction — rather than a continue, pass-then-skip, or routing of the record identifier to a separate fallback path.
4. Confirm that the collection receiving both real features and the constant placeholder is later passed to a model fitting method or a prediction method (e.g., .fit(...), .predict(...), or equivalent).
5. PRESENT when all of steps 1–4 hold: broad handler, constant placeholder appended into the shared feature collection, and that collection feeds training or inference.
• Predicted impact:
• Add score for C12: Failure absorbed into a worse result: 2
• weight: 2
• confidence: high
• Evidence: A per-record decoder with except: features.append(np.zeros(N)) fed zero rows into both training and prediction; a contrastive variant that skipped undecodable records during training and assigned the modal class at inference scored 0.326 higher on the task metric.
• Applies when: A pipeline decodes individual records (images or other binary payloads) into numeric feature vectors inside a loop with per-record error handling, and the resulting feature matrix is used to fit a classifier or produce predictions.
• Example:
• Input:
    feats, labels = [], []
    for rec in records:
        try:
            img = Image.open(io.BytesIO(rec["payload"]))
            feats.append(extract(img))
        except Exception:
            feats.append(np.zeros(FEAT_DIM))
        labels.append(rec["label"])
    model.fit(np.array(feats), labels)
    preds = model.predict(np.array(test_feats))
• Consequence:
    Every undecodable record becomes an identical all-zero row: training rows
    at the origin pull the fitted decision boundary toward an arbitrary class,
    and failed test records all receive whatever class the model assigns to
    the zero vector. Classification accuracy drops measurably (observed gap:
    0.33 on the task metric versus skipping failed records).
• Counter-example:
• Input:
    feats, labels, failed_ids = [], [], []
    for rec in records:
        try:
            img = Image.open(io.BytesIO(rec["payload"]))
            feats.append(extract(img))
            labels.append(rec["label"])
        except Exception:
            failed_ids.append(rec["_id"])
            continue
    model.fit(np.array(feats), labels)
    fallback = Counter(labels).most_common(1)[0][0]
• Why it does not fire: the broad handler skips the failed record (continue) and records its identifier for an explicit fallback prediction, so no constant placeholder vector enters the feature collection used for fitting or prediction.
C1 · data-leakage Uncertainty scaled from in-sample per-entity fit residuals

P1 — Uncertainty scaled from in-sample per-entity fit residuals

• Pattern: Detects a submitted uncertainty/confidence value that is derived from residuals of a curve fitted on each entity's own complete target history, with no out-of-fold evaluation of the actual inference-time prediction path.
• Detection procedure:
1. Locate the construction of the output frame that is written to the submission file and identify the column holding an uncertainty, confidence, or sigma value (e.g., a column passed alongside the point prediction that the metric interprets as a scale).
2. Trace the value assigned to that column back to its source expression. Check whether it is computed from residuals — differences between observed target values and fitted values of a model fitted per entity on ALL of that entity's rows in the training frame (e.g., a per-group loop or groupby where a regression's own training-set predictions are subtracted from the same rows' targets, then aggregated with a standard deviation or mean absolute value).
3. Search the same script for any cross-validation split (e.g., GroupKFold, KFold, or equivalent manual fold loop) whose held-out predictions follow the same code path as the final test-time prediction (same baseline offset, same feature construction, same model call) and whose held-out errors feed into the uncertainty column identified in step 1.
4. PRESENT if the uncertainty column traces to in-sample per-entity fit residuals (step 2 holds) and no out-of-fold error distribution from the deployment prediction path contributes to it (step 3 finds nothing).
• Predicted impact:
• Add score for C1: Train/eval contamination: 3
• weight: 3
• confidence: high
• Evidence: Submission's uncertainty was set from np.std(y - lr.predict(X)) inside a per-entity loop over each entity's full history; the sibling script that instead calibrated the scale against grouped out-of-fold absolute errors of the deployed baseline-plus-slope path scored 0.053 better on the likelihood metric.
• Applies when: A longitudinal or forecasting task where the scored output includes both a point prediction and an uncertainty/scale value, and the metric penalizes cases where the true error exceeds the submitted scale.
• Example:
• Input:
    conf = {}
    for pid, g in train.groupby("entity_id"):
        lr = LinearRegression().fit(g[["week"]], g["target"])
        resid = g["target"] - lr.predict(g[["week"]])
        conf[pid] = max(resid.std(), 50.0)

    sub["pred"] = sub["entity_id"].map(slopes_pred_fn)
    sub["confidence"] = sub["entity_id"].map(conf).fillna(150.0)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Submitted sigma reflects smooth in-sample fit noise (~50-80 units) while
    actual forecast errors of the deployment path are several times larger;
    the negative-log-likelihood metric worsens on nearly every row where
    true error exceeds sigma, lowering the overall score by ~0.05.
• Counter-example:
• Input:
    gkf = GroupKFold(n_splits=5)
    oof_err = np.zeros(len(tr))
    for tr_idx, va_idx in gkf.split(tr, groups=tr["entity_id"]):
        m = Ridge().fit(X[tr_idx], y[tr_idx])
        pred = base[va_idx] + m.predict(X[va_idx]) * dw[va_idx]
        oof_err[va_idx] = np.abs(tr["target"].values[va_idx] - pred)

    c0, c1 = fit_scale(oof_err, np.abs(dw))
    sub["confidence"] = np.clip(c0 + c1 * np.abs(sub["dw"]), 70, 1000)
• Why it does not fire: the uncertainty scale is calibrated from held-out errors produced by the same baseline-plus-model prediction path used at inference, so step 3 finds an out-of-fold source feeding the confidence column.
C1 · data-leakage Uncertainty calibrated on in-sample residualsgithub_occurrence

P1 — Uncertainty calibrated on in-sample residuals

• Pattern: Detects an uncertainty or confidence estimate derived from residuals between labels and predictions where the predicting model was fitted on those same rows (e.g., per-entity fits on each entity's full label history), with no held-out or out-of-fold separation between the rows used to fit and the rows used to compute residuals.
• Detection procedure:
1. Locate a statement that computes residuals: a subtraction whose operands are a label column of a training frame (e.g., df["target"]) and the output of a prediction call (.predict(...) or equivalent) on rows of that same frame.
2. Trace the model object used in the prediction call at step 1 to its fitting call (.fit(...) or equivalent) within the same script; confirm the rows passed to the fitting call include the rows predicted at step 1 (same variable, same slice, or a per-entity group inside a loop or groupby over an entity-id column where both fit and predict receive the same group frame).
3. Confirm the residuals from step 1 flow into a spread or scale statistic (std, mean of absolute values, a quantile, or equivalent) that is later assigned to an uncertainty, confidence, or interval-width column of an output frame.
4. Confirm no cross-validation splitter, leave-one-group-out loop, or explicit holdout mask separates the rows used in the fitting call at step 2 from the rows used to compute residuals at step 1.
5. PRESENT if steps 1–4 all hold: an uncertainty value is calibrated from residuals of a model evaluated on its own training rows.
• Predicted impact:
• Add score for C1 (Train/eval contamination): 3
• weight: 3
• confidence: high
• Evidence: Per-entity linear fits on each entity's full label history produced residuals used to set the output Confidence value; because the fits saw every label they were scored against, the residual spread was optimistically small, and the overconfidence-penalizing metric was 0.037 worse than a variant that tuned the confidence scale against a proper validation score.
• Applies when: The script outputs a per-row uncertainty, confidence, or interval-width value, and that value is computed from residuals of some fitted model on longitudinal or per-entity training data.
• Example:
• Input:
    sigmas = []
    for pid, g in train.groupby("entity_id"):
        m = LinearRegression().fit(g[["week"]], g["value"])
        resid = g["value"] - m.predict(g[["week"]])
        sigmas.append(resid.std())
    conf = float(np.nanmean(sigmas))
    sub["Confidence"] = conf
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Predicted uncertainty is systematically smaller than the true out-of-sample
    error, so a likelihood-based metric that penalizes overconfidence degrades
    (observed gap ~0.037 on the competition metric versus a calibration that
    did not reuse fitted rows).
• Counter-example:
• Input:
    oof = cross_val_predict(
        model, X, y, cv=GroupKFold(5), groups=train["entity_id"]
    )
    resid = y - oof
    sub["Confidence"] = np.abs(resid).std()
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the residuals come from out-of-fold predictions where each row's prediction was made by a model that never saw that row's entity during fitting, so step 4's no-separation condition fails.
C1 · data-leakage Uncertainty parameter tuned on in-sample residuals

P1 — Uncertainty parameter tuned on in-sample residuals

• Pattern: Detects tuning of a per-prediction uncertainty parameter (a sigma/confidence value entering a likelihood-style score) by optimizing the metric over residuals computed from the model's predictions on the same rows it was fitted on, with no out-of-fold or held-out split.
• Detection procedure:
1. Locate a model-fitting call (e.g. model.fit(X, y) or equivalent) and note the array or frame of rows used for fitting.
2. Locate a prediction call (model.predict(...) or equivalent) whose input is the same variable (or a subscript/copy of it) as the fitting rows from step 1, and a residual computation of the form target minus that prediction (or its absolute value) within the same script.
3. Locate an optimization or search over an uncertainty parameter — a call to a scalar minimizer (minimize, minimize_scalar, or equivalent) or an explicit loop over candidate sigma values — whose objective function consumes the residuals from step 2 and includes the uncertainty parameter in a log/likelihood-style expression (e.g. terms containing both the residual and log of the parameter, or a clipped absolute-error-over-sigma term).
4. Confirm there is no resampling construct (cross-validation splitter, grouped fold iterator, or an explicit index split into disjoint fit/evaluate row sets) separating the rows used in step 1 from the rows whose residuals feed step 3.
5. The pattern is PRESENT when steps 1–3 all hold and step 4 finds no such separation.
• Predicted impact:
• Add score for C1 (Train/eval contamination): 3
• weight: 3
• confidence: high
• Evidence: A script computed resid = y_train - model.predict(X_train) on the fitting rows and then ran minimize(nll, sigma0) over those residuals to pick best_sigma, which was written as the constant confidence column; the likelihood-style leaderboard score was ~0.2 absolute (~3% relative) worse than a variant that derived the uncertainty from held-out-style residual structure.
• Applies when: The task requires submitting a per-prediction uncertainty (sigma/confidence) that is scored by a likelihood- or pinball-style metric, and the script tunes that uncertainty numerically from model residuals.
• Example:
• Input:
    model.fit(X_train, y_train)
    resid = y_train - model.predict(X_train)

    def nll(sigma):
        s = max(sigma, 70.0)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(nll, x0=[200.0]).x[0]
    sub["Confidence"] = best_sigma
• Consequence:
    In-sample residuals understate true error, so the tuned sigma is
    systematically too small; the likelihood-style score on unseen rows
    degrades by roughly 0.2 absolute (~3% relative) versus a sigma tuned
    on held-out residuals.
• Counter-example:
• Input:
    gkf = GroupKFold(n_splits=5)
    oof = np.zeros(len(y))
    for tr, va in gkf.split(X, y, groups=df["group_id"]):
        m = Ridge().fit(X[tr], y[tr])
        oof[va] = m.predict(X[va])
    resid = y - oof

    def nll(sigma):
        s = max(sigma, 70.0)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(nll, x0=[200.0]).x[0]
• Why it does not fire: The residuals feeding the sigma optimizer come from out-of-fold predictions produced by a grouped resampling split, so the rows being predicted were never in the corresponding fitting set (step 4 finds a separation).
C1 · data-leakage In-sample residuals used as submitted uncertainty scale

P1 — In-sample residuals used as submitted uncertainty scale

• Pattern: Detects an uncertainty or confidence value in the submitted output whose scale is derived from residuals computed on the same rows the model (or per-group fit) was trained on, with no out-of-fold or held-out separation.
• Detection procedure:
1. Identify a call that fits a model to a feature matrix and target (e.g. model.fit(X, y), np.polyfit(...), a per-group regression inside a groupby loop, or equivalent).
2. Within the same script, identify a prediction call whose input rows are the same variable (or a frame derived from it without row exclusion) as the fitting input at step 1, and a residual computed as the difference between the training target and that prediction.
3. Identify a scalar or per-row statistic of those residuals (e.g. .std(), np.percentile(...), mean absolute value, or a scale fitted to their magnitudes) that is assigned to an uncertainty/confidence column of the output written to the submission file.
4. Confirm there is no cross-validation split, cross_val_predict (or equivalent out-of-fold loop), or explicit held-out partition separating the rows used at step 1 from the rows predicted at step 2.
5. PRESENT when steps 1–4 all hold: the submitted uncertainty scale comes from predictions on rows the fit already saw.
• Predicted impact:
• Add score for C1 (Train/eval contamination): 2
• weight: 2
• confidence: high
• Evidence: A pipeline computing resid = y - model.predict(X) on the fitting rows and setting the submission's confidence column from resid.std() scored 0.056 worse on a likelihood-style metric than a pipeline that calibrated the same confidence from grouped out-of-fold residuals.
• Applies when: The evaluation metric consumes a predicted uncertainty/confidence value, and that value is calibrated from residuals of the model's own predictions.
• Example:
• Input:
    model.fit(X_tr, y_tr)
    resid = y_tr - model.predict(X_tr)
    sigma = float(np.std(resid))

    preds = model.predict(X_te)
    out = pd.DataFrame({
        "id": test_ids,
        "value": preds,
        "confidence": sigma,
    })
    out.to_csv("submission.csv", index=False)
• Consequence:
    In-sample residual spread understates held-out error, so the submitted
    confidence is too tight; the likelihood metric over-penalizes large
    errors and the final score drops (observed gap ~0.056 versus an
    out-of-fold calibration on the same features and model).
• Counter-example:
• Input:
    oof = cross_val_predict(model, X_tr, y_tr,
                            cv=GroupKFold(5), groups=grp)
    resid = y_tr - oof
    sigma = float(np.std(resid))

    model.fit(X_tr, y_tr)
    out = pd.DataFrame({"id": test_ids,
                        "value": model.predict(X_te),
                        "confidence": sigma})
    out.to_csv("submission.csv", index=False)
• Why it does not fire: the residuals feeding the uncertainty scale come from out-of-fold predictions, so each row's prediction was produced by a model that never trained on that row.
C1 · data-leakage Uncertainty scale tuned on in-sample residualsverified_trace · effect +0.2196

P1 — Uncertainty scale tuned on in-sample residuals

• Pattern: Detects a per-prediction uncertainty/confidence value whose scale parameter is optimized against residuals computed from predictions on the same rows the model was fitted on, with no held-out or out-of-fold split separating fitting from residual computation.
• Detection procedure:
1. Locate a model-fitting call (e.g., .fit(X, y) or equivalent) whose feature and target arrays are derived from a single training frame (e.g., train loaded from train.csv).
2. Locate, after step 1, a prediction call (e.g., model.predict(X) or equivalent) whose input array is the same variable, or is derived from the same training frame rows, as the fitting input in step 1 — with no row subsetting, fold indexing, or cross-validation loop separating fit rows from prediction rows.
3. Locate a residual computation of the form true-minus-predicted (any subtraction between the target column of that frame and the predictions from step 2), whose result feeds an optimization over a scalar (a call to a numerical minimizer, a grid/loop selecting an argmin/argmax over candidate values, or a closed-form quantile/std of the residuals).
4. The scalar produced in step 3 is assigned into an uncertainty/confidence output column of the final prediction frame (a column written to the output file alongside the point predictions).
5. PRESENT when all of steps 1–4 hold and there is no fold structure (e.g., KFold, GroupKFold, or equivalent manual index partition) between the fitting in step 1 and the prediction in step 2.
• Predicted impact:
• Add score for C1 (Train/eval contamination): 2
• weight: 2
• confidence: high
• Evidence: A scalar tuned via minimize(loss, sigma0) on residuals from model.predict(X_train) (the fit rows) and written as the constant Confidence column produced a likelihood-style score ~2.5% worse than a variant computing the confidence from held-out-aware residual structure; a second matched pair showed ~2.8% degradation.
• Applies when: Regression/forecasting task whose evaluation metric requires a per-prediction uncertainty or confidence value in addition to the point prediction, and the code tunes that value from residual statistics.
• Example:
• Input:
    model.fit(X_train, y_train)
    preds = model.predict(X_train)
    resid = y_train - preds

    def neg_ll(sigma):
        s = np.maximum(sigma, 70)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(neg_ll, x0=[200], method="Nelder-Mead").x[0]
    sub["Confidence"] = best_sigma
• Consequence:
    In-sample residuals are systematically smaller than held-out residuals, so the
    optimizer selects a sigma that is too small; on unseen rows the true errors exceed
    the optimistic scale and the log-likelihood metric degrades by ~2.5% relative
    versus tuning sigma on out-of-fold residuals.
• Counter-example:
• Input:
    oof = np.zeros(len(train))
    gkf = GroupKFold(n_splits=5)
    for tr, va in gkf.split(X, y, groups=train["subject_id"]):
        m = Ridge().fit(X[tr], y[tr])
        oof[va] = m.predict(X[va])
    resid = y - oof

    def neg_ll(sigma):
        s = np.maximum(sigma, 70)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(neg_ll, x0=[200], method="Nelder-Mead").x[0]
• Why it does not fire: The residuals feeding the sigma optimization come from out-of-fold predictions — each row is predicted by a model that never saw it — so step 5's no-fold-structure condition fails.
C1 · data-leakage Uncertainty from in-sample residuals of per-entity fitsgithub_occurrence

P1 — Uncertainty from in-sample residuals of per-entity fits

• Pattern: Detects a submitted uncertainty/confidence value computed from the spread of residuals between a fitted model's predictions and the very target rows that model was fitted on (e.g., per-entity curves scored against each entity's own training points), with no out-of-fold or held-out split separating fitting from residual measurement.
• Detection procedure:
1. Locate the construction of the output frame written to the submission file and identify a column representing uncertainty (a column whose name contains Confidence, sigma, std, or uncertainty, or a value passed to the metric's uncertainty slot).
2. Trace the value assigned to that column backwards within the script to a spread statistic — a call to .std(), np.std, np.percentile/np.quantile of absolute differences, or mean absolute difference — applied to a residual array.
3. Identify how the residual array is computed: it is y - model.predict(X) (or an equivalent subtraction of fitted values from targets, e.g. np.polyval(coeffs, x) - y), where the fitting call (fit, np.polyfit, curve_fit, or equivalent) received the same rows X, y (or the same per-group subset selected by an identical group key) as the prediction call producing the residuals.
4. Confirm there is no cross-validation splitter (KFold, GroupKFold, or equivalent), no explicit train/validation row partition, and no out-of-fold prediction array between the fitting call in step 3 and the residual computation in step 2.
5. PRESENT when the submitted uncertainty column (step 1) derives from a spread statistic (step 2) over residuals whose fitting and evaluation rows are identical (step 3) and no held-out split intervenes (step 4).
• Predicted impact:
• Add score for C1: Train/eval contamination: 2
• weight: 2
• confidence: high
• Evidence: Uncertainty column filled from np.std(y_group - fitted_curve(x_group)) where the curve was fit on the same group's rows; the resulting sigma understated true unseen-point error and the metric's error-over-sigma penalty inflated, costing ~0.05 metric units versus an out-of-fold residual estimate in a measured contrastive pair.
• Applies when: The evaluation metric scores a submitted uncertainty estimate alongside point predictions (e.g., a Gaussian log-likelihood or pinball-style penalty), and the data is grouped per entity with predictions made per group.
• Example:
• Input:
    sigmas = {}
    for pid, g in train.groupby("entity_id"):
        coeffs = np.polyfit(g["week"], g["col_a"], 1)
        resid = g["col_a"] - np.polyval(coeffs, g["week"])
        sigmas[pid] = resid.std()

    sub["pred"] = sub["entity_id"].map(slopes_pred)
    sub["Confidence"] = sub["entity_id"].map(sigmas).fillna(150.0)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Submitted sigma is systematically smaller than the true error on unseen
    weeks; the metric's |error|/sigma penalty term rises and the final score
    degrades by ~0.05 relative to an out-of-fold residual estimate.
• Counter-example:
• Input:
    oof = np.zeros(len(train))
    for tr, va in GroupKFold(5).split(X, y, groups=train["entity_id"]):
        m = Ridge().fit(X[tr], y[tr])
        oof[va] = m.predict(X[va])
    resid = y - oof
    base_sigma = np.std(resid)
    sub["Confidence"] = base_sigma + 0.5 * np.abs(sub["dw"])
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the residuals feeding the spread statistic come from out-of-fold predictions, so the rows used to fit each model are disjoint from the rows on which residuals are measured.
C1 · data-leakage Uncertainty parameter tuned on in-sample residualsverified_trace · effect +0.2196

P1 — Uncertainty parameter tuned on in-sample residuals

• Pattern: Detects a predictive-uncertainty parameter (e.g., a sigma or confidence width) that is optimized against residuals computed from a model's predictions on the same rows the model was fitted on, with no out-of-fold or held-out prediction step separating fitting from residual computation.
• Detection procedure:
1. Locate a call that fits a predictive model (e.g., .fit(...), a closed-form regression solve, np.polyfit, or equivalent) and record the variable(s) holding the rows used for fitting.
2. Locate a residual computation of the form y_true - y_pred (or its absolute/squared variant), where y_pred comes from applying the fitted model (via .predict(...) or equivalent evaluation of fitted coefficients) to rows that are a subset of, or identical to, the rows recorded in step 1.
3. Locate a subsequent optimization or search over a scalar uncertainty parameter — e.g., a call to scipy.optimize.minimize, minimize_scalar, a loop over candidate sigma values selecting the best score, or equivalent — whose objective function consumes the residuals from step 2.
4. Confirm there is no cross-validation split (KFold, GroupKFold, train_test_split, or equivalent) or held-out mask applied between the fit in step 1 and the residual computation in step 2 such that the residuals come from rows unseen by the fitted model.
5. PRESENT when steps 1–3 are all found and the condition in step 4 holds (residuals used for tuning the uncertainty parameter are in-sample).
• Predicted impact:
• Add score for C1 (Train/eval contamination): 2
• weight: 2
• confidence: high
• Evidence: A pipeline fit a regression on all training rows, computed resid = train["y"] - model_pred(train) on those same rows, and passed them to minimize to pick best_sigma used as the submitted uncertainty; a variant that derived the uncertainty width from grouped out-of-fold residuals scored 0.02456 higher on the likelihood-based competition metric.
• Applies when: The evaluation metric includes a predicted uncertainty/confidence term, and that uncertainty value is chosen by optimizing over residuals of a fitted model.
• Example:
• Input:
    from scipy.optimize import minimize
    coef = np.linalg.lstsq(X_train, y_train, rcond=None)[0]
    preds = X_train @ coef
    resid = y_train - preds

    def neg_score(sigma):
        s = np.maximum(sigma, 70)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(neg_score, x0=200.0).x[0]
    submission["Confidence"] = best_sigma
• Consequence:
    In-sample residuals understate the true error spread, so best_sigma is
    systematically too small; on unseen rows the likelihood metric applies a
    larger penalty per error unit, and the leaderboard score is measurably
    lower (~0.02 worse on the metric) than tuning on held-out residuals.
• Counter-example:
• Input:
    from sklearn.model_selection import GroupKFold
    oof = np.zeros(len(y))
    for tr, va in GroupKFold(5).split(X, y, groups=subj):
        m = Ridge().fit(X[tr], y[tr])
        oof[va] = m.predict(X[va])
    resid = y - oof

    def neg_score(sigma):
        s = max(sigma, 70)
        return np.mean(np.sqrt(2) * np.minimum(np.abs(resid), 1000) / s
                       + np.log(np.sqrt(2) * s))
    best_sigma = minimize(neg_score, x0=200.0).x[0]
• Why it does not fire: The residuals fed to the sigma optimizer are out-of-fold predictions from models that never saw the corresponding validation rows, so step 4's condition fails.
C1 · data-leakage Zero-offset baseline rows included in residual calibration fit

P1 — Zero-offset baseline rows included in residual calibration fit

• Pattern: Detects training data constructed by pairing each observation with a same-group baseline row, where rows whose offset from the baseline is zero (so the target trivially equals the baseline feature) are retained in the set used to fit the model and to compute residual-based uncertainty statistics.
• Detection procedure:
1. Find a construction that joins a table to itself or to a per-group baseline row (e.g., merge on a group id, a groupby-first/min selection, or equivalent), producing columns such as a baseline measurement (e.g., base_val) and a baseline position (e.g., base_week).
2. Find a derived offset column computed as the difference between the row's position column and the baseline position column (e.g., df["dw"] = df["Weeks"] - df["base_week"] or equivalent subtraction).
3. Confirm the target passed to a fitting call is either the raw measurement whose value equals the baseline feature when the offset is zero, or a delta that is exactly zero on those rows, and that a residual/error quantity computed on the fitted rows (e.g., y - pred) feeds an uncertainty or spread statistic (std, quantile, optimized sigma) used in the output.
4. Confirm that between the offset computation at step 2 and the fitting/residual computation at step 3, there is no row filter excluding zero-offset rows (no boolean mask, query, or slice comparing the offset column to 0 such as df[df["dw"] != 0]).
5. PRESENT when steps 1–3 hold and the filter in step 4 is absent.
• Predicted impact:
• Add score for C1: 2
• weight: 2
• confidence: high
• Evidence: A pipeline built training rows by merging each row with its group baseline, kept rows where dw == 0 (target identical to base_val) in the fit, and optimized a global sigma from the resulting residuals; the near-zero residuals on the trivial rows shrank the residual distribution and the calibrated uncertainty, and the likelihood-scored metric on held-out data was 0.029 worse than a variant that handled zero-offset rows separately.
• Applies when: Tabular regression where the target is derived from a per-group baseline row and residuals from the fit are used to calibrate an uncertainty or confidence output.
• Example:
• Input:
    base = df.sort_values("Weeks").groupby("Patient").first().reset_index()
    df = df.merge(base[["Patient", "Weeks", "val"]].rename(
        columns={"Weeks": "base_week", "val": "base_val"}), on="Patient")
    df["dw"] = df["Weeks"] - df["base_week"]
    X = df[["dw", "base_val", "col_a"]].values
    y = df["val"].values
    model.fit(X, y)
    resid = y - model.predict(X)
    sigma = np.std(resid)          # deflated by zero-offset rows
    out["Confidence"] = sigma
• Consequence:
    Zero-offset rows have target == base_val, so their residuals are near
    zero; the pooled residual std is biased low, the emitted confidence is
    systematically too small, and the held-out likelihood-based score
    degrades (observed ~0.029 worse on the competition metric).
• Counter-example:
• Input:
    base = df.sort_values("Weeks").groupby("Patient").first().reset_index()
    df = df.merge(base[["Patient", "Weeks", "val"]].rename(
        columns={"Weeks": "base_week", "val": "base_val"}), on="Patient")
    df["dw"] = df["Weeks"] - df["base_week"]
    fit_df = df[df["dw"] != 0]
    X = fit_df[["dw", "base_val", "col_a"]].values
    y = fit_df["val"].values
    model.fit(X, y)
    resid = y - model.predict(X)
    sigma = np.std(resid)
• Why it does not fire: the df[df["dw"] != 0] filter removes the degenerate zero-offset rows before both the fit and the residual computation, so the calibration statistic is not deflated.
C1 · data-leakage Uncertainty parameter tuned on in-sample residuals

P1 — Uncertainty parameter tuned on in-sample residuals

• Pattern: Detects a scalar uncertainty/confidence value for a likelihood-scored submission that is optimized against residuals computed from a model's predictions on the very rows the model was fitted on, with no held-out or out-of-fold partition separating fitting from residual computation.
• Detection procedure:
1. Locate a model-fitting call (e.g., .fit(X, y), closed-form coefficient solve, or equivalent) whose feature and target arrays are derived from the full training frame.
2. Locate a residual computation of the form y - preds (or absolute/squared variants) within the same script, where preds is produced by applying the model fitted at step 1 to the same feature array (or a frame containing the same rows) that was passed to the fitting call.
3. Locate a search over an uncertainty parameter — a call to a numeric optimizer, a loop over candidate values with argmin/argmax selection, or equivalent — whose objective function reads the residuals from step 2, and whose selected value is later assigned to the submission's uncertainty/confidence column.
4. Verify there is no train/validation split, cross-validation loop, or out-of-fold prediction array standing between the fitting call at step 1 and the residuals used at step 3 (e.g., no split function or fold iterator whose held-out indices feed the residual computation).
5. PRESENT when steps 1–3 all match and step 4 confirms the residuals feeding the uncertainty search are strictly in-sample.
• Predicted impact:
• Add score for C1: 2
• weight: 2
• confidence: high
• Evidence: A pipeline computed resid = train["target"] - predict(train) after fitting on all of train, then ran minimize(neg_loglik, sigma0) over those residuals and wrote the resulting best_sigma as a constant confidence; the held-out clipped log-likelihood metric was ~2.8% relatively worse than a counterpart that scaled uncertainty from out-of-fold–calibrated residual behavior.
• Applies when: The task is scored by a probabilistic/likelihood metric and the submission includes a per-row uncertainty or confidence value alongside point predictions, on tabular data.
• Example:
• Input:
    model.fit(X_train, y_train)
    resid = y_train - model.predict(X_train)

    def neg_loglik(sigma):
        s = np.maximum(sigma, 70.0)
        d = np.minimum(np.abs(resid), 1000.0)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(neg_loglik, x0=[200.0]).x[0]
    sub["Confidence"] = best_sigma
• Consequence:
    In-sample residuals are optimistically small, so the optimizer selects a
    sigma below the true predictive error; on unseen rows the clipped
    log-likelihood metric degrades (~2.8% relative worse than tuning the same
    sigma on out-of-fold residuals).
• Counter-example:
• Input:
    oof = np.zeros(len(y))
    for tr, va in GroupKFold(5).split(X, y, groups):
        m = Ridge().fit(X[tr], y[tr])
        oof[va] = m.predict(X[va])
    resid = y - oof
    best_sigma = minimize(neg_loglik_from(resid), x0=[200.0]).x[0]
    model = Ridge().fit(X, y)
    sub["Confidence"] = best_sigma
• Why it does not fire: the residuals feeding the sigma search come from out-of-fold predictions produced by fold models that never saw the rows they score, so step 4's separation condition holds.
C1 · data-leakage Uncertainty scale from per-subject in-sample residuals mismatched to deployed feature-only modelgithub_occurrence

P1 — Uncertainty scale from per-subject in-sample residuals mismatched to deployed feature-only model

• Pattern: Detects an uncertainty/confidence output whose scale is derived from residuals of per-group models fitted on each group's full outcome history, while the predictions actually submitted come from a different model that sees only baseline features, so the reported error scale describes a stronger model than the one producing the predictions.
• Detection procedure:
1. Find a loop or group-wise apply (e.g., for pid, g in df.groupby("subject_id") or equivalent) in which a model is fitted using the target column of that group's own rows (e.g., regressing g["target"] on g["time"]).
2. Within the same script, find residuals computed as the difference between that per-group fitted model's predictions and the same group's observed target values, aggregated into a scalar spread statistic (a call to .std(), a mean absolute residual, or equivalent) assigned to a variable, call it sigma.
3. Find a second, separately fitted model whose fitting call does not receive the per-group target history used in step 1 (fit only on baseline/feature columns), and whose predict call generates the values written to the prediction column of the output file.
4. Confirm that sigma from step 2 (possibly after arithmetic scaling) is assigned to the uncertainty/confidence column of the same output file.
5. PRESENT when steps 1–4 all hold: the output's prediction column comes from the feature-only model of step 3 while its uncertainty column comes from residuals of the per-group history-fitted models of steps 1–2.
• Predicted impact:
• Add score for C1: Train/eval contamination: 2
• weight: 2
• confidence: high
• Evidence: sigma = np.std(all_per_subject_residuals) computed from per-subject slope fits on each subject's full outcome trajectory was written as the confidence value alongside predictions from a global feature-only regressor; the measured likelihood-style competition metric was 0.028 worse than a variant whose error scale matched the deployed model.
• Applies when: the task has longitudinal per-subject data, the output requires both a point prediction and an uncertainty/confidence value per row, and the score penalizes the ratio of absolute error to the stated uncertainty.
• Example:
• Input:
    resids = []
    for pid, g in train.groupby("subject_id"):
        lr = LinearRegression().fit(g[["week"]], g["target"])
        resids.extend(g["target"] - lr.predict(g[["week"]]))
    sigma = np.std(resids)  # in-sample, per-subject history fits

    model = LinearRegression().fit(train[feature_cols], train["target"])
    sub["target"] = model.predict(sub[feature_cols])
    sub["Confidence"] = sigma
    sub.to_csv("submission.csv", index=False)
• Consequence:
    sigma reflects tight per-subject trajectory fits, not the deployed
    feature-only model's true error; the confidence value is systematically
    too small, so the |error|/sigma penalty term is inflated on every row
    and the likelihood-style score drops (~0.028 worse on the metric than
    when the scale matches the deployed model).
• Counter-example:
• Input:
    tr, va = train_test_split(train, test_size=0.2, random_state=0)
    model = LinearRegression().fit(tr[feature_cols], tr["target"])
    va_resid = va["target"] - model.predict(va[feature_cols])
    sigma = np.std(va_resid)  # held-out residuals of the deployed model

    model = LinearRegression().fit(train[feature_cols], train["target"])
    sub["target"] = model.predict(sub[feature_cols])
    sub["Confidence"] = sigma
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the uncertainty scale is computed from held-out residuals of the same feature-only model class that generates the submitted predictions, with no per-group fits on outcome history involved.
C2 · invalid-validation Uncertainty scale calibrated from in-sample residuals

P1 — Uncertainty scale calibrated from in-sample residuals

• Pattern: Detects a submitted uncertainty/confidence value derived from residuals of a model evaluated on the very observations it was fitted to, rather than from residuals collected on held-out points.
• Detection procedure:
1. Identify the assignment that produces the uncertainty output column of the submission frame (e.g. a column named like Confidence, sigma, uncertainty written alongside the point prediction before the final to_csv or equivalent write).
2. Trace the value assigned in step 1 back to a residual computation: a subtraction (or absolute/squared difference) between the observed target array and predictions from a fitted model (a call to .predict(...), np.polyval(...), or equivalent).
3. Check whether the rows passed to the fitting call (.fit(X, y), np.polyfit(...), or equivalent) and the rows passed to the prediction call in step 2 are the same object, the same slice, or the same per-entity group, with no split, mask, or fold index excluding the predicted rows from the fitting rows anywhere between the two calls (including inside a per-entity loop where each group's model is fit on that group's full history).
4. Check that the quantity flowing into step 1 is a scale statistic of those residuals (standard deviation, mean absolute value, quantile, or a constant multiple thereof).
5. PRESENT if steps 1–4 all hold: the submitted uncertainty is a scale statistic of residuals whose fitting rows and evaluation rows coincide.
• Predicted impact:
• Add score for C2: Validation scheme does not match the data structure: 2
• weight: 2
• confidence: high
• Evidence: Per-entity curve fits produced residuals on their own training points, and Confidence = residuals.std() was submitted as the uncertainty; the resulting scale was smaller than true extrapolation error and the likelihood-style score was 0.037 worse than a variant that calibrated its uncertainty scale on out-of-fold residuals.
• Applies when: The task's metric requires an uncertainty output alongside each point prediction, and predictions are extrapolated per entity (e.g. per-entity time series forecast) beyond the observed points.
• Example:
• Input:
    confs = {}
    for eid, g in df.groupby("entity_id"):
        coef = np.polyfit(g["week"], g["target"], 1)
        pred = np.polyval(coef, g["week"])
        resid = g["target"] - pred          # residuals on the fitting rows
        confs[eid] = max(70.0, resid.std())
    sub["Confidence"] = sub["entity_id"].map(confs)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Submitted uncertainty values understate the true forecast error at
    unseen weeks; the likelihood-style score degrades (e.g. ~0.04 worse)
    versus deriving the scale from held-out residuals.
• Counter-example:
• Input:
    oof_resid = []
    for tr_idx, va_idx in GroupKFold(5).split(X, y, groups=df["entity_id"]):
        model.fit(X[tr_idx], y[tr_idx])
        oof_resid.append(y[va_idx] - model.predict(X[va_idx]))
    scale = np.abs(np.concatenate(oof_resid)).mean()
    sub["Confidence"] = np.maximum(70.0, scale * (1 + 0.02 * sub["week_gap"]))
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the residuals feeding the uncertainty scale come from validation indices excluded from each fold's fitting rows, so the fitting and evaluation rows do not coincide (step 3 fails).
C2 · invalid-validation Uncertainty parameter tuned on in-sample residualsverified_trace · effect +0.0603

P1 — Uncertainty parameter tuned on in-sample residuals

• Pattern: Detects an uncertainty or confidence parameter for a likelihood-scored forecast being optimized against residuals computed by predicting on the same rows the model was fitted on, with no out-of-fold or held-out split before the residual computation.
• Detection procedure:
1. Locate a model-fitting call (e.g., .fit(X, y) or equivalent) where X and y are derived from the full training frame, with no preceding train/validation split (train_test_split, KFold, GroupKFold, cross_val_predict, or equivalent) applied to those rows within the same script.
2. Locate a subsequent prediction call (e.g., .predict(X) or equivalent) whose input array is the same variable, or derived from the same rows, as the fitting input at step 1, and a residual computation of the form true values minus those predictions (or an equivalent error term) on those rows.
3. Locate an optimization or search over a scalar spread/confidence parameter (e.g., a call to minimize, a grid loop over candidate sigma values, or equivalent) whose objective is computed from the residuals of step 2, and whose result is later written into the output's confidence/uncertainty column.
4. PRESENT if all of steps 1–3 hold: the residuals feeding the uncertainty-parameter optimization come from predictions on the fitting rows themselves, with no out-of-fold or group-held-out prediction step between fitting and residual computation.
• Predicted impact:
• Add score for C2 (Validation scheme does not match the data structure): 2
• weight: 2
• confidence: high
• Evidence: A pipeline fitted a regressor on all training rows, computed residuals = y - model.predict(X) on those same rows, then ran minimize over sigma on those residuals and wrote best_sigma into the Confidence column; a variant deriving the spread from held-out-style behavior scored ~0.029 better (≈2.9% relative) on the likelihood metric.
• Applies when: Tabular or longitudinal data where the evaluation metric includes an uncertainty/confidence term (e.g., a Laplace or Gaussian log-likelihood) and the script both fits a point-prediction model and tunes a spread parameter in the same run.
• Example:
• Input:
    model.fit(X_train, y_train)
    resid = y_train - model.predict(X_train)

    def neg_ll(sigma):
        s = max(sigma[0], 70.0)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))

    best_sigma = minimize(neg_ll, x0=[200.0]).x[0]
    sub_df["Confidence"] = best_sigma
• Consequence:
    In-sample residuals understate true error, so the optimized sigma is
    systematically too small; on held-out entities the log-likelihood metric
    degrades by roughly 0.2 absolute (~3% relative) versus tuning on
    out-of-fold residuals.
• Counter-example:
• Input:
    gkf = GroupKFold(n_splits=5)
    oof = np.zeros(len(y))
    for tr, va in gkf.split(X, y, groups=df["entity_id"]):
        m = Ridge().fit(X[tr], y[tr])
        oof[va] = m.predict(X[va])
    resid = y - oof
    best_sigma = minimize(neg_ll_from(resid), x0=[200.0]).x[0]
    sub_df["Confidence"] = best_sigma
• Why it does not fire: the residuals feeding the sigma optimization come from group-aware out-of-fold predictions, so no row is predicted by a model that was fitted on it.
C2 · invalid-validation Row-level random split across derived variants of the same source itemverified_trace · effect +0.1106

P1 — Row-level random split across derived variants of the same source item

• Pattern: Detects a feature matrix built by stacking multiple derived variants of each source item (several rows produced per underlying item) that is then partitioned into train and validation by a row-level random split with no grouping by source item, so variants of one item can land on both sides of the split.
• Detection procedure:
1. Locate code that constructs the training matrix by iterating over a single list of source identifiers (file names, item ids) more than once — e.g., one loop per variant subdirectory, transformation, or label class applied to the same identifier list — and appends one row per (identifier, variant) pair via append/vstack/concat or equivalent.
2. Confirm that the stacked rows share source identifiers across variants: the same identifier list or directory listing feeds every variant loop from step 1, or rows are generated as for variant in ...: for item in items: ....
3. Locate a call to train_test_split (or equivalent random row-wise splitter such as a shuffled index permutation followed by slicing, or KFold/ShuffleSplit without a groups argument) whose inputs are the stacked matrix and labels from step 1, within the same script.
4. Verify that no group-aware mechanism ties variants together at split time: no GroupKFold, GroupShuffleSplit, StratifiedGroupKFold (or equivalent), no groups= argument, and no manual split performed on the unique identifier list before row expansion.
5. The pattern is PRESENT when steps 1–4 all hold: multi-variant rows per source item exist and the train/validation partition is drawn at the row level with no per-item grouping.
• Predicted impact:
• Add score for C2 (Validation scheme does not match the data structure): 3
• weight: 3
• confidence: high
• Evidence: A pipeline stacked several transformed versions of each source item into one matrix and split it with train_test_split(X, y, test_size=0.2); the sibling run that split at the source-item level scored 0.288 higher on the final held-out metric, showing the row-level split had tuned early stopping and model selection to a leaked validation estimate.
• Applies when: The training set contains multiple rows derived from the same underlying source item (augmented copies, per-variant encodings, paired originals and modifications) and a holdout or cross-validation split is used to report a score, select a model, or drive early stopping.
• Example:
• Input:
    items = sorted(os.listdir(src_dir))
    rows, labels = [], []
    for variant in ["orig", "mod_a", "mod_b"]:
        for name in items:
            rows.append(extract_features(os.path.join(src_dir, variant, name)))
            labels.append(0 if variant == "orig" else 1)
    X, y = np.vstack(rows), np.array(labels)
    Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.2, random_state=0)
    model.fit(Xtr, ytr, eval_set=[(Xva, yva)], early_stopping_rounds=50)
    print("val score:", score(yva, model.predict(Xva)))
• Consequence:
    Reported validation score is optimistically inflated (observed gap ~0.29 on
    the final metric vs. a group-level split): near-duplicate variants of the
    same item appear in both partitions, early stopping halts at an iteration
    tuned to leaked rows, and the held-out score comes in well below the
    validation number used for model selection.
• Counter-example:
• Input:
    items = sorted(os.listdir(src_dir))
    tr_items, va_items = train_test_split(items, test_size=0.2, random_state=0)
    def build(subset):
        rows, labels = [], []
        for variant in ["orig", "mod_a", "mod_b"]:
            for name in subset:
                rows.append(extract_features(os.path.join(src_dir, variant, name)))
                labels.append(0 if variant == "orig" else 1)
        return np.vstack(rows), np.array(labels)
    Xtr, ytr = build(tr_items)
    Xva, yva = build(va_items)
    model.fit(Xtr, ytr, eval_set=[(Xva, yva)], early_stopping_rounds=50)
• Why it does not fire: the random split is performed on the unique source-item list before row expansion, so all variants of a given item fall entirely in train or entirely in validation (step 4 fails).
C3 · metric-mismatch Unbounded linear trend extrapolation over the prediction horizon

P1 — Unbounded linear trend extrapolation over the prediction horizon

• Pattern: Detects a per-entity fitted trend coefficient extrapolated across the full prediction horizon where neither the coefficient nor the resulting predicted values are constrained to a range derived from the observed training target distribution before the output file is written.
• Detection procedure:
1. Locate an assignment where a trend or slope value is obtained from a fitted linear estimator (e.g., the coef_ attribute of a linear model, the first return value of np.polyfit with degree 1, or an equivalent manual least-squares slope computation).
2. Locate a subsequent expression that computes predictions of the form base value plus that trend value multiplied by a time-offset variable (an expression containing the slope variable multiplied by a difference or offset of a time/step column).
3. Locate the statement that writes the predictions to the output file (e.g., a call to to_csv or equivalent on a frame containing the predicted column).
4. Check whether, between step 1 and step 3, the slope variable OR the predicted values are passed through a bounding operation (np.clip, .clip(...), min(...)/max(...) composition, or a conditional reassignment comparing against numeric bounds).
5. The pattern is PRESENT if steps 1–3 are found and no bounding operation from step 4 is applied to either the slope variable or the predicted values before the write in step 3.
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 2
• weight: 2
• confidence: high
• Evidence: A pipeline computing pred = base + slope * dw from per-entity fitted slopes with no np.clip on either slope or pred scored 0.05257 worse on the competition metric than an otherwise similar pipeline that clipped the predicted slope and final values to bounds taken from the training target range; outlier slopes produced physically impossible values at the horizon extremes, inflating absolute error on those rows.
• Applies when: The program forecasts a numeric target over a long fixed horizon by extrapolating a fitted linear or trend-based predictor, and the evaluation metric penalizes absolute or squared deviation on every horizon row.
• Example:
• Input:
    for pid, grp in train.groupby("entity_id"):
        slope = np.polyfit(grp["week"], grp["target"], 1)[0]
        slopes[pid] = slope

    for i, (pid, w) in enumerate(zip(entities, weeks)):
        base = baselines[pid]
        dw = w - base_week[pid]
        preds[i] = base + slopes[pid] * dw

    sub["target"] = preds
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Entities with two closely spaced noisy observations get extreme slope
    estimates; at horizon offsets of +/-100 steps the predictions leave the
    physically plausible target range, driving per-row absolute error to the
    metric's error cap on those rows and lowering the averaged score by ~0.05
    versus a bounded extrapolation.
• Counter-example:
• Input:
    for pid, grp in train.groupby("entity_id"):
        slope = np.polyfit(grp["week"], grp["target"], 1)[0]
        slopes[pid] = np.clip(slope, -12, 2)

    for i, (pid, w) in enumerate(zip(entities, weeks)):
        dw = w - base_week[pid]
        preds[i] = baselines[pid] + slopes[pid] * dw

    lo, hi = train["target"].min(), train["target"].max()
    sub["target"] = np.clip(preds, lo, hi)
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: Both the fitted slope and the final predictions are clipped to bounds derived from the training target distribution before the output is written, so step 5's conjunction fails.
C3 · metric-mismatch fixed per-class quota subsampling under accuracy scoringgithub_occurrence

P1 — fixed per-class quota subsampling under accuracy scoring

• Pattern: Detects construction of a training subsample by capping the number of records kept per class at a small fixed quota, on a class-imbalanced multiclass task evaluated by plain accuracy, which flattens the class prior the model learns relative to the skewed prior the metric averages over.
• Detection procedure:
1. In the training-data collection code, locate a per-class counter (e.g., a dict or Counter keyed by a label value, or a length check on a per-class list such as len(bucket[label])).
2. Confirm that within the same loop or comprehension, a record is skipped or the loop continues when that per-class count reaches a fixed constant threshold (a literal integer or a constant like PER_CLASS = 40), so every class contributes at most the same number of samples.
3. Confirm the task is multiclass classification whose final output is a single predicted label per row (submission or evaluation compares hard labels, i.e., plain accuracy or equivalent), and there is no later step that reweights predictions, class weights, or decision scores by the empirical class frequencies of the full dataset.
4. PRESENT if steps 1–3 all hold: per-class quota capping in training-sample construction, on a hard-label accuracy-style task, with no downstream prior correction.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A pipeline that filled per-class buckets with if len(bucket[cat]) < PER_CLASS: bucket[cat].append(feat) scored 0.69 lower on the accuracy-style metric than an otherwise similar pipeline that streamed training records at a uniform stride, preserving the natural skewed class frequencies; the quota version systematically under-predicted the frequent classes that dominate the test distribution.
• Applies when: Multiclass classification with a heavily skewed class distribution, the training set is subsampled for tractability, and the score is plain accuracy (or any metric that averages over the natural test distribution rather than per class).
• Example:
• Input:
    PER_CLASS = 40
    bucket = {}
    for rec in iter_records("train.csv"):
        y = rec["label"]
        lst = bucket.setdefault(y, [])
        if len(lst) < PER_CLASS:
            lst.append(featurize(rec))
    X = np.vstack([f for lst in bucket.values() for f in lst])
    labels = [y for y, lst in bucket.items() for _ in lst]
    model.fit(X, labels)
• Consequence:
    Test accuracy drops sharply (observed gap ~0.69 on one skewed task) versus
    prior-preserving subsampling: frequent classes, which dominate the test
    rows, are predicted far less often than their true rate, while rare
    classes are over-predicted, so the per-row accuracy average falls.
• Counter-example:
• Input:
    STRIDE = 25
    feats, labels = [], []
    for i, rec in enumerate(iter_records("train.csv")):
        if i % STRIDE != 0:
            continue
        feats.append(featurize(rec))
        labels.append(rec["label"])
    model.fit(np.vstack(feats), labels)
• Why it does not fire: the subsample is taken at a uniform stride across all records with no per-class counter or quota, so the sample's class frequencies match the natural skewed distribution the metric averages over.
C3 · metric-mismatch Constant confidence for AP-scored predictionsverified_trace · effect +0.0002

P1 — Constant confidence for AP-scored predictions

• Pattern: Detects that every prediction emitted for a task scored by average precision carries the same hard-coded constant confidence value, so no prediction can be ranked above any other.
• Detection procedure:
1. Confirm the program writes multiple predictions per sample into a submission or output file whose format includes a per-prediction confidence/score field (e.g. a prediction string beginning with a score, or a score/confidence column), and the surrounding task is evaluated by average precision or mean average precision.
2. Locate the expression that supplies the confidence field for each emitted prediction (a variable interpolated into the prediction string, or a value assigned to the score column).
3. Trace that expression: it is a numeric literal (e.g. 1.0, 0.5), or a variable assigned exactly once to a numeric literal before the emission loop and never reassigned or modified inside the loop.
4. Verify no per-prediction quantity (model output, cluster size, count, distance, probability, or equivalent) flows into the confidence expression anywhere in the file.
5. PRESENT when steps 1–4 all hold: the confidence written for every prediction is provably the same constant for all predictions.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A script defined confidence = 1.0 once and interpolated it identically into every emitted prediction string; a contrastive run that instead ranked the same candidate boxes by conf = max(0.01, min(1.0, cd["n"] / maxn)) (normalized cluster support) scored measurably higher on the AP-based metric, with the full observed gap attributable to the ranking.
• Applies when: The program produces multiple scored predictions per sample for a task evaluated by average precision, mean average precision, or another ranking-sensitive precision-recall metric.
• Example:
• Input:
    confidence = 1.0
    pred_strings = []
    for sample_id in sub_df["Id"]:
        parts = []
        for cand in candidates:
            parts.append(f"{confidence} {cand['x']:.3f} {cand['y']:.3f} "
                         f"{cand['w']:.3f} {cand['h']:.3f} LABEL_A")
        pred_strings.append(" ".join(parts))
    sub_df["PredictionString"] = pred_strings
    sub_df.to_csv("submission.csv", index=False)
• Consequence:
    All predictions tie at the same score, so the precision-recall sweep cannot
    place likely hits before unlikely ones; measured mAP is lower than the
    identical set of predictions ranked by empirical support (a full metric-unit
    gap was observed in a controlled contrastive run).
• Counter-example:
• Input:
    max_n = max(c["n"] for c in candidates) if candidates else 1
    pred_strings = []
    for sample_id in sub_df["Id"]:
        parts = []
        for cand in candidates:
            conf = max(0.01, min(1.0, cand["n"] / max_n))
            parts.append(f"{conf} {cand['x']:.3f} {cand['y']:.3f} "
                         f"{cand['w']:.3f} {cand['h']:.3f} LABEL_A")
        pred_strings.append(" ".join(parts))
    sub_df["PredictionString"] = pred_strings
• Why it does not fire: the confidence expression is recomputed inside the emission loop from a per-prediction quantity (cand["n"]), so emitted scores differ across predictions and induce a meaningful ranking.
C3 · metric-mismatch Unweighted probability blend without validated weightsgithub_occurrence

P1 — Unweighted probability blend without validated weights

• Pattern: Detects blending of class-probability outputs from multiple heterogeneous classifiers by fixed equal-weight arithmetic averaging, with no out-of-fold or held-out predictions used to select blend weights, when the evaluated metric scores probabilities directly.
• Detection procedure:
1. Identify two or more distinct estimator objects of different model families (e.g., a tree ensemble, a naive Bayes model, a linear model, or equivalents in any library) that are each fitted on the training features.
2. Find, for each of those estimators, a call to a probability-output method (predict_proba or equivalent) on the test/inference features, with results stored in separate variables.
3. Find an expression that combines those probability arrays by summing them and dividing by the number of models (e.g., (p1 + p2 + p3) / 3), or averaging with np.mean over a stacked axis, or multiplying each by the same literal constant, with all coefficients equal.
4. Search the whole script for any of: predictions generated on validation folds or a held-out split from the training data, a loop or optimizer that varies blend coefficients, or a metric call (e.g., a log-loss or equivalent probabilistic scoring function) evaluated on candidate blended predictions. Confirm none exist.
5. PRESENT when steps 1–3 hold and step 4 confirms no weight selection against held-out or out-of-fold predictions occurs anywhere before the equal-weight blend is written to output.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: ensemble_pred = (pred1 + pred2 + pred3) / 3 over three heterogeneous classifiers' probability outputs, with no out-of-fold evaluation or weight search; the equal-weight blend scored 0.418 worse on the probabilistic metric than a blend whose weights were grid-searched on out-of-fold predictions.
• Applies when: A multi-classifier ensemble produces class probabilities for a task scored by a probabilistic metric (e.g., log loss), and the members differ in model family or calibration quality.
• Example:
• Input:
    clf1.fit(X_tr, y)   # tree ensemble
    clf2.fit(X_tr, y)   # naive Bayes
    clf3.fit(X_tr, y)   # linear model
    p1 = clf1.predict_proba(X_te)
    p2 = clf2.predict_proba(X_te)
    p3 = clf3.predict_proba(X_te)
    blend = (p1 + p2 + p3) / 3
    pd.DataFrame(blend, columns=classes).to_csv("out.csv", index=False)
• Consequence:
    Log loss on the evaluation set is substantially higher (observed +0.42)
    than a blend with weights optimized on out-of-fold predictions, because
    poorly calibrated members drag the average toward miscalibrated
    probabilities instead of being down-weighted.
• Counter-example:
• Input:
    for name, m in models.items():
        oof[name], tst[name] = get_oof_preds(m, X_tr, y, X_te)
    best = min(
        ((log_loss(y, w*oof["a"] + (1-w)*oof["b"]), w)
         for w in np.arange(0, 1.01, 0.05))
    )
    _, w = best
    blend = w * tst["a"] + (1 - w) * tst["b"]
    pd.DataFrame(blend, columns=classes).to_csv("out.csv", index=False)
• Why it does not fire: the blend coefficients are selected by evaluating the probabilistic metric on out-of-fold predictions, so step 4's search for weight selection succeeds and the conjunction fails.
C3 · metric-mismatch Per-class binary classifiers with post-hoc sum-normalization for a mutually exclusive multiclass targetverified_trace · effect +0.0306

P1 — Per-class binary classifiers with post-hoc sum-normalization for a mutually exclusive multiclass target

• Pattern: Detects a mutually exclusive multiclass prediction task modeled by fitting one independent binary classifier per class and then dividing each classifier's positive-class output by the sum of all classifiers' outputs, instead of training a single jointly optimized multiclass model.
• Detection procedure:
1. Identify two or more classifier objects assigned to distinct variables, each fitted via a call to a fitting method (fit or equivalent) with a different binary target array, where the binary targets are derived from mutually exclusive one-hot columns of the same source frame (e.g., each target is one column of a group of indicator columns, or an equality comparison of a single label column against different values).
2. Identify that each fitted model's per-row score for the positive class (e.g., predict_proba(...)[:, 1] or equivalent) is stored in a separate variable.
3. Identify an expression that sums those per-model score variables element-wise (optionally with a small epsilon constant) and divides each individual score variable by that sum.
4. Identify that the divided (renormalized) values are written to the output columns of the final prediction artifact (e.g., assembled into a frame written with to_csv or equivalent).
5. The pattern is PRESENT when steps 1–4 all hold: multiple independently fitted binary classifiers over mutually exclusive targets, whose outputs are renormalized by their sum and used as the final per-class probabilities.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: Three independent binary models fitted per class with outputs combined as pred_x / (pred_a + pred_b + pred_c + 1e-10) scored measurably worse on multiclass log loss (difference ≈ 0.039) than a single model trained with a multinomial objective on the same features; the independently estimated probabilities are not jointly calibrated, so the renormalized distribution is systematically miscalibrated.
• Applies when: The task requires predicting a probability distribution over three or more mutually exclusive classes, and the evaluation metric is a proper scoring rule over the full distribution (e.g., multiclass log loss).
• Example:
• Input:
    m1 = LogisticRegression().fit(X_tr, y_is_class1)
    m2 = LogisticRegression().fit(X_tr, y_is_class2)
    m3 = LogisticRegression().fit(X_tr, y_is_class3)
    p1 = m1.predict_proba(X_te)[:, 1]
    p2 = m2.predict_proba(X_te)[:, 1]
    p3 = m3.predict_proba(X_te)[:, 1]
    total = p1 + p2 + p3 + 1e-10
    sub = pd.DataFrame({'id': ids, 'LABEL_A': p1/total,
                        'LABEL_B': p2/total, 'LABEL_C': p3/total})
    sub.to_csv('submission.csv', index=False)
• Consequence:
    Multiclass log loss is higher (worse) than a single softmax model on the
    same features — the independently fitted per-class probabilities are not
    jointly calibrated, and sum-normalization systematically distorts the
    predicted distribution (observed gap ≈ 0.04 log loss on the same data).
• Counter-example:
• Input:
    y = np.argmax(df[['LABEL_A', 'LABEL_B', 'LABEL_C']].values, axis=1)
    model = LogisticRegression(multi_class='multinomial', max_iter=1000)
    model.fit(X_tr, y)
    preds = model.predict_proba(X_te)
    sub = pd.DataFrame({'id': ids, 'LABEL_A': preds[:, 0],
                        'LABEL_B': preds[:, 1], 'LABEL_C': preds[:, 2]})
    sub.to_csv('submission.csv', index=False)
• Why it does not fire: The one-hot columns are collapsed into a single categorical label and one jointly trained multinomial model produces the full distribution directly, so there are no independently fitted per-class binaries and no post-hoc sum-normalization.
C3 · metric-mismatch Default-strength regularization for probabilistic linear classifier on high-dimensional handcrafted featuresverified_trace · effect +2.4295

P1 — Default-strength regularization for probabilistic linear classifier on high-dimensional handcrafted features

• Pattern: Detects a linear probabilistic classifier fitted on high-dimensional handcrafted feature vectors for a many-class task, instantiated without any explicit regularization-strength argument (library default) and with no penalty-strength search or probability-calibration step before its class-probability outputs are written as the final predictions.
• Detection procedure:
1. Locate a constructor call for a linear classifier that produces class probabilities (e.g. LogisticRegression, SGDClassifier(loss='log_loss'), or equivalent) whose result is later passed to a .fit(...) call.
2. Verify the constructor call contains no regularization-strength keyword (no C=, alpha=, or equivalent penalty-strength argument) or leaves it at the library default value.
3. Verify the feature matrix passed to .fit originates from handcrafted feature extraction (e.g. gradient histograms, color statistics, manual concatenations of per-sample descriptor arrays), not from a learned embedding, and its second dimension is a constant or expression of several hundred or more.
4. Verify that within the same script there is no loop or search over multiple regularization-strength values scored on a held-out split, and no post-hoc calibration wrapper (e.g. a calibration class or blending of the probability matrix toward a uniform or prior distribution) applied before the probabilities are written to the output file.
5. Verify the classifier's probability outputs (e.g. from predict_proba or equivalent) are written directly to the submission/output file. PRESENT if steps 1–5 all hold.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: LogisticRegression() fitted with default regularization strength on thousands-dimensional handcrafted descriptors for a many-class probability task produced sharply peaked, overconfident probabilities; a run identical except for a 100× stronger penalty (C=0.01) plus a small uniform blend scored ~0.62 better on the probability-quality metric.
• Applies when: The task is many-class classification scored by log loss or another metric that penalizes overconfident probability estimates, and the features are handcrafted/weak (not learned representations), making the linear boundary prone to overfitting noise directions.
• Example:
• Input:
    from sklearn.linear_model import LogisticRegression
    from sklearn.preprocessing import StandardScaler

    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)        # X: (n, 3000) handcrafted descriptors
    clf = LogisticRegression(max_iter=1000)   # default regularization strength
    clf.fit(X_scaled, y)                      # y has ~100 classes
    probs = clf.predict_proba(scaler.transform(X_test))
    pd.DataFrame(probs, columns=classes).to_csv("submission.csv", index=False)
• Consequence:
    The run completes normally, but the near-interpolating boundary emits
    probabilities close to 0/1 on many test rows; wrong confident rows incur
    large per-row log-loss penalties, and the overall log loss is substantially
    higher (worse) than the same pipeline with a 100x stronger penalty, which
    scored ~0.62 better on the probability-quality metric.
• Counter-example:
• Input:
    from sklearn.linear_model import LogisticRegression
    from sklearn.preprocessing import StandardScaler

    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    clf = LogisticRegression(C=0.01, max_iter=1000)  # strong penalty
    clf.fit(X_scaled, y)
    probs = clf.predict_proba(scaler.transform(X_test))
    probs = 0.9 * probs + 0.1 / probs.shape[1]       # blend toward uniform
    pd.DataFrame(probs, columns=classes).to_csv("submission.csv", index=False)
• Why it does not fire: the constructor explicitly sets a much stronger regularization strength and the probabilities are tempered toward a uniform prior before being written, so steps 2 and 4 both fail.
C3 · metric-mismatch Uniform probability averaging without validation-based weight selection

P1 — Uniform probability averaging without validation-based weight selection

• Pattern: Detects an ensemble that averages predicted class probabilities from multiple heterogeneous classifiers using fixed equal weights, with no out-of-fold or held-out evaluation of the scored metric used to choose blend weights or confirm each member helps.
• Detection procedure:
1. Identify two or more distinct classifier objects instantiated from different model classes (or with materially different configurations) within the same script, each fitted on the same training features.
2. Locate an expression that combines the probability outputs of these models (e.g. results of predict_proba or equivalent probability-emitting calls on the test/inference split) by summing them and dividing by their count, or by multiplying each by the same literal constant before summing.
3. Search the whole script for any computation of a validation metric (e.g. log_loss, roc_auc_score, or equivalent) applied to predictions on rows held out from fitting (cross-validation folds, a validation split, or out-of-fold arrays), whose result influences the blend — either by selecting weights via a loop/optimizer or by conditionally including/excluding a member.
4. PRESENT if step 2's equal-weight combination exists and no metric computation from step 3 influences the combination weights or membership anywhere in the script.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: ensemble_pred = (pred1 + pred2 + pred3) / 3 blending three heterogeneous classifiers with no held-out metric evaluation scored 0.41778 worse on the probabilistic competition metric than a variant that computed out-of-fold predictions per member and grid-searched blend weights on the metric before applying them to test predictions.
• Applies when: The script produces probabilistic predictions scored by a probability-sensitive metric (e.g. log loss) and combines outputs from two or more classifiers of different model families.
• Example:
• Input:
    clf1.fit(X_train, y_train)
    clf2.fit(X_train, y_train)
    clf3.fit(X_train, y_train)
    pred1 = clf1.predict_proba(X_test)
    pred2 = clf2.predict_proba(X_test)
    pred3 = clf3.predict_proba(X_test)
    ensemble_pred = (pred1 + pred2 + pred3) / 3
    pd.DataFrame(ensemble_pred).to_csv("submission.csv", index=False)
• Consequence:
    Log loss on the evaluation set is substantially higher (worse) than a
    validation-tuned blend — observed gap of ~0.42 — because a poorly
    calibrated member receives the same 1/3 weight as well-calibrated
    members instead of being down-weighted toward zero.
• Counter-example:
• Input:
    for name, model in models.items():
        oof[name], tst[name] = cross_val_probs(model, X_train, y, X_test)
    best = None
    for w1 in np.arange(0, 1.01, 0.05):
        for w2 in np.arange(0, 1.01 - w1 + 1e-9, 0.05):
            w3 = 1 - w1 - w2
            ll = log_loss(y, w1*oof["a"] + w2*oof["b"] + w3*oof["c"])
            if best is None or ll < best[0]:
                best = (ll, w1, w2, w3)
    _, w1, w2, w3 = best
    pred = w1*tst["a"] + w2*tst["b"] + w3*tst["c"]
• Why it does not fire: The blend weights are selected by minimizing a held-out metric on out-of-fold predictions, so the combination is not a fixed equal-weight average and step 3's condition is satisfied.
C3 · metric-mismatch Hardcoded uncertainty value in a metric that scores uncertaintygithub_occurrence

P1 — Hardcoded uncertainty value in a metric that scores uncertainty

• Pattern: Detects a program that fills a predicted-uncertainty output column with literal numeric constants (or constants plus a fixed heuristic term) without any code that evaluates the scoring metric's formula over candidate uncertainty values to select them.
• Detection procedure:
1. Identify the code that writes the final output file (e.g. a call to to_csv, to_parquet, or equivalent) and locate the output column representing an uncertainty, confidence, or standard-deviation value that the task's metric consumes alongside the point prediction.
2. Trace every assignment to that column within the program. Mark the pattern-relevant case if each such assignment is a numeric literal (e.g. 150.0), a numeric literal used in fillna, or an arithmetic expression whose coefficients are all numeric literals not assigned from any search loop.
3. Search the whole source for a loop or optimizer call that (a) iterates over candidate values for those coefficients and (b) inside its body computes a score using both residuals (difference between predictions and held-out true targets) and the candidate uncertainty value — e.g. an expression combining an absolute or clipped error term with a logarithm or division involving the uncertainty variable — and keeps the best-scoring candidate.
4. The pattern is PRESENT if step 2 finds only literal-derived assignments to the uncertainty column and step 3 finds no such metric-evaluating search anywhere in the source.
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 2
• weight: 2
• confidence: high
• Evidence: A submission writer using submission["uncertainty"].fillna(150.0) and constant per-row values scored measurably worse (~0.037 absolute, ~3.7% relative) on a likelihood-style metric than an otherwise similar pipeline that grid-searched an intercept-plus-slope uncertainty parameterization by directly evaluating the metric on held-out residuals.
• Applies when: The task's scoring metric jointly penalizes prediction error and a per-row predicted uncertainty value, and the required output includes an uncertainty/confidence field.
• Example:
• Input:
    preds = model.predict(X_test)
    sub = sample[["row_id"]].copy()
    sub["value"] = preds
    sub["uncertainty"] = 150.0  # fixed guess
    sub["value"] = sub["value"].fillna(train["value"].median())
    sub["uncertainty"] = sub["uncertainty"].fillna(150.0)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    The likelihood-style score is systematically lower (observed ~3.7% relative
    degradation) because the constant uncertainty is mis-sized relative to the
    actual residual distribution the metric penalizes: too small inflates the
    error term, too large inflates the log-penalty term.
• Counter-example:
• Input:
    best = (1e9, None, None)
    for c0 in range(60, 301, 10):
        for c1 in np.arange(0.0, 6.0, 0.5):
            sigma = np.clip(c0 + c1 * np.abs(dw_val), 70, None)
            err = np.clip(np.abs(y_val - pred_val), 0, 1000)
            score = -np.mean(-np.sqrt(2) * err / sigma - np.log(np.sqrt(2) * sigma))
            if score < best[0]:
                best = (score, c0, c1)
    _, C0, C1 = best
    sub["uncertainty"] = np.clip(C0 + C1 * np.abs(sub["dw"]), 70, None)
• Why it does not fire: the uncertainty coefficients are chosen by a search loop that evaluates the metric formula on held-out residuals, so step 3 finds a metric-evaluating search and step 4's conjunction fails.
C3 · metric-mismatch Class rebalancing before probability-scored output without recalibrationgithub_occurrence

P1 — Class rebalancing before probability-scored output without recalibration

• Pattern: Detects a classifier fitted with class rebalancing (balanced class weights or explicit minority resampling) whose raw predicted probabilities are then written as the final output for a task evaluated by log loss or another proper scoring rule, with no recalibration step in between.
• Detection procedure:
1. Locate any classifier construction that enables class rebalancing: a constructor argument class_weight='balanced' (or an explicit weight dictionary), or a resampling call such as an over/under-sampler applied to the training rows before the fit call (or equivalent in any library).
2. Locate a later call on that same fitted model producing per-class probabilities (e.g., predict_proba or equivalent), whose result flows — possibly through arithmetic normalization only — into a data frame or array written to the final output file.
3. Scan the code between the probability call and the output write for any recalibration construct: a calibration wrapper (e.g., CalibratedClassifierCV or equivalent), an isotonic/sigmoid fit on held-out predictions, or an explicit prior-correction formula that rescales probabilities using original class frequencies.
4. Confirm the evaluation target is a probability-scored metric: the output columns are probabilities (values compared against a log loss, Brier, or similar proper scoring rule) rather than hard class labels.
5. PRESENT when steps 1, 2, and 4 hold and step 3 finds no recalibration between fitting and writing the probabilities.
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 2
• weight: 2
• confidence: high
• Evidence: LogisticRegression(class_weight='balanced') fitted per class, then raw predict_proba outputs normalized and written directly to the submission; the rebalanced model scored measurably worse (≈0.024 higher loss) than an otherwise comparable model trained without reweighting, because probabilities were shifted away from the empirical class priors.
• Applies when: The program trains a classifier and its output is per-class probabilities scored by log loss or another proper scoring rule (not accuracy/F1 on hard labels).
• Example:
• Input:
    model = LogisticRegression(class_weight='balanced', max_iter=1000)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
    submission = pd.DataFrame(probs, columns=['LABEL_A', 'LABEL_B', 'LABEL_C'])
    submission.insert(0, 'id', test_df['id'])
    submission.to_csv('submission.csv', index=False)
• Consequence:
    Log loss increases (e.g., +0.02 or more versus an unweighted fit) because
    every row's predicted probabilities are systematically inflated for
    minority classes and deflated for the majority class, deviating from the
    true class priors that the proper scoring rule rewards.
• Counter-example:
• Input:
    base = LogisticRegression(class_weight='balanced', max_iter=1000)
    model = CalibratedClassifierCV(base, method='isotonic', cv=3)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
    submission = pd.DataFrame(probs, columns=['LABEL_A', 'LABEL_B', 'LABEL_C'])
    submission.insert(0, 'id', test_df['id'])
    submission.to_csv('submission.csv', index=False)
• Why it does not fire: although the base estimator uses balanced class weights, the probabilities are recalibrated by an isotonic calibration wrapper before being written, restoring alignment with the empirical class priors.
C3 · metric-mismatch Hardcoded uncertainty value where the metric scores calibrationgithub_occurrence

P1 — Hardcoded uncertainty value where the metric scores calibration

• Pattern: Detects a submission for a likelihood-style metric that scores both a point prediction and an uncertainty value, where the uncertainty output column is filled with hardcoded constants or ad-hoc heuristic expressions and the source never evaluates the stated metric formula on residuals nor optimizes the uncertainty parameter against it.
• Detection procedure:
1. Confirm the task's scoring formula takes two outputs per row: a point prediction and an uncertainty/confidence/scale value (e.g. a Laplace or Gaussian log-likelihood with a sigma term), and the output frame has a column for that uncertainty value.
2. In the source, locate every assignment to the uncertainty column of the output frame (e.g. submission["Confidence"] = ..., df["sigma"] = ... or equivalent).
3. Check whether each such assignment's right-hand side is a numeric literal, a literal passed through fillna, or an arithmetic expression combining literals with feature values (e.g. a base constant plus a distance-from-baseline term), with no dependence on prediction residuals.
4. Search the whole file for any computation of residuals (a difference between predicted and observed target values on training or held-out rows) that feeds into an expression implementing the task's scoring formula, and for any call to a numerical optimizer or grid search (e.g. scipy.optimize.minimize, a loop over candidate sigma values selecting the best score, or equivalent) whose result is assigned to the uncertainty column.
5. PRESENT if step 3 holds for all uncertainty-column assignments and step 4 finds no residual-based evaluation of the scoring formula anywhere in the file.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: A run assigning submission["Confidence"] = 150.0 (and fillna(150.0)) scored measurably worse than an otherwise comparable run that computed training residuals, implemented the metric formula, and set the uncertainty column from an optimized best_sigma — a ~0.029 gap on the competition metric for equivalent point predictions.
• Applies when: The task's evaluation metric has a closed-form formula, stated in the task description, that jointly scores a point prediction and a per-row uncertainty/scale value, and the program writes both to a submission file.
• Example:
• Input:
    model = LinearRegression().fit(X_train, y_train)
    submission["value"] = model.predict(X_test)
    # confidence: fixed floor plus distance-from-baseline heuristic
    submission["confidence"] = 150.0 + 2.0 * submission["weeks_from_base"].abs()
    submission["confidence"] = submission["confidence"].fillna(150.0)
    submission.to_csv("submission.csv", index=False)
• Consequence:
    The likelihood-based score is systematically worse (~2.9% relative in the
    measured pair) than the score achievable with the same point predictions,
    because the emitted uncertainty is miscalibrated versus the residual scale
    that the metric formula rewards.
• Counter-example:
• Input:
    resid = y_train - model.predict(X_train)
    def metric_loss(sigma):
        s = np.maximum(sigma, 70.0)
        d = np.minimum(np.abs(resid), 1000.0)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))
    best_sigma = minimize(metric_loss, x0=[200.0], method="Nelder-Mead").x[0]
    submission["value"] = model.predict(X_test)
    submission["confidence"] = best_sigma
    submission.to_csv("submission.csv", index=False)
• Why it does not fire: the uncertainty column is set from a value obtained by evaluating the task's scoring formula on training residuals and numerically optimizing it, so step 4 finds residual-based optimization and the pattern is absent.
C3 · metric-mismatch Constant global uncertainty ignoring error-driving covariate

P1 — Constant global uncertainty ignoring error-driving covariate

• Pattern: Detects a submission that pairs each point prediction with an uncertainty value assigned as a single scalar constant for all rows, never computed from any per-row covariate such as horizon or extrapolation distance.
• Detection procedure:
1. Identify the code that builds the output frame written by a call to to_csv (or equivalent), and locate the column holding a per-row uncertainty, spread, or confidence value alongside the point-prediction column.
2. Trace the assignment of that uncertainty column. Check whether the right-hand side is a scalar: a numeric literal, a variable bound earlier to a single float (e.g., the result of a scalar optimization or a global residual statistic), or a scalar broadcast into the column.
3. Within the same script, check whether that scalar is ever combined arithmetically with any per-row array or column (e.g., a time offset, distance, or other feature that varies across output rows) before being written to the uncertainty column.
4. PRESENT if the uncertainty column is assigned from a scalar (step 2) and no per-row covariate enters its computation (step 3), while the point-prediction column itself does vary per row.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: A submission assigned sub_merged["Confidence"] = best_sigma, a single fitted scalar, for every output row; a contrasting version computing conf = a + b * np.abs(dw) scaled uncertainty with the prediction offset and improved the averaged log-likelihood metric by ~0.028.
• Applies when: The output file is scored by a probabilistic metric that rewards per-row calibrated uncertainty, and the rows span varying prediction difficulty (e.g., differing forecast horizons or extrapolation distances).
• Example:
• Input:
    resid = train["y"] - train_pred
    best_sigma = float(np.std(resid))

    sub["y_pred"] = predict(sub_features)
    sub["uncertainty"] = best_sigma
    sub[["row_id", "y_pred", "uncertainty"]].to_csv("submission.csv", index=False)
• Consequence:
    Averaged log-likelihood score drops relative to horizon-scaled uncertainty:
    near-baseline rows (small true error) are over-penalized by an inflated sigma,
    while far-horizon rows (large true error) incur heavy miss penalties from an
    underestimated sigma. Observed gap on the same predictions: ~0.028 metric points.
• Counter-example:
• Input:
    resid = train["y"] - train_pred
    a, b = fit_sigma_params(resid, train["dw"])

    sub["y_pred"] = predict(sub_features)
    sub["uncertainty"] = np.maximum(a + b * np.abs(sub["dw"].values), 70.0)
    sub[["row_id", "y_pred", "uncertainty"]].to_csv("submission.csv", index=False)
• Why it does not fire: The uncertainty column is computed from a per-row covariate (sub["dw"]), so it varies with prediction difficulty even though scalar parameters a and b appear in the expression.
C3 · metric-mismatch Raw probabilities submitted without clipping or prior-blending under log lossverified_trace · effect +1.2715

P1 — Raw probabilities submitted without clipping or prior-blending under log loss

• Pattern: Detects a probability-scored multiclass submission where the model's predicted class probabilities are written to the output file with no clipping away from zero and no blending or shrinkage toward a uniform or prior distribution, leaving confident wrong predictions to incur near-unbounded per-row log-loss penalties.
• Detection procedure:
1. Locate a variable assigned from a call to a probability-prediction method on a fitted classifier (e.g. predict_proba, a softmax over logits, or equivalent) applied to the evaluation/test split.
2. Trace every transformation of that variable (including any DataFrame it is copied into) between the assignment at step 1 and a call that writes it to a submission file (to_csv or equivalent).
3. Check whether any of those transformations bounds the values away from zero: a call to np.clip/.clip with a positive lower bound, an elementwise maximum with a positive constant, or an affine blend of the form a * probs + b with b > 0 (uniform/prior mixing), followed or not by row renormalization. Row renormalization alone (dividing by row sums) does not count, since it cannot move a zero entry away from zero.
4. The pattern is PRESENT if the probabilities reach the submission write with no such bounding transformation applied anywhere on the path from step 1 to the write.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: probs = clf.predict_proba(X_test); ...; prob_df.div(row_sums, axis=0); submission.to_csv(...) with no clip or prior blend; the measured metric was 0.61 worse than an otherwise identical pipeline that applied 0.9 * probs + 0.1 / K before writing, with the raw-probability run scoring worse than a uniform-prediction baseline.
• Applies when: A multiclass classification task is scored by log loss (or another metric that heavily penalizes near-zero probability on the true class) and the program writes per-class probabilities to a submission file.
• Example:
• Input:
    probs = model.predict_proba(X_test_scaled)
    prob_df = pd.DataFrame(probs, columns=classes)
    prob_df = prob_df.reindex(columns=submission_cols, fill_value=0.0)
    row_sums = prob_df.sum(axis=1).replace(0, 1)
    prob_df = prob_df.div(row_sums, axis=0)
    submission = pd.concat([pd.DataFrame({"id": test_ids}), prob_df], axis=1)
    submission.to_csv("submission.csv", index=False)
• Consequence:
    Log loss inflates far above the uniform baseline (e.g. 11.9 vs ln(K)=4.79),
    because rows where the model assigns near-zero probability to the true
    class contribute penalties approaching -log(0); a clipped/blended variant
    of the same pipeline scores ~0.6 better on the competition metric.
• Counter-example:
• Input:
    probs = model.predict_proba(X_test_scaled)
    out = pd.DataFrame(0.0, index=range(len(test_ids)), columns=classes)
    for j, c in enumerate(model.classes_):
        out[c] = probs[:, j]
    K = len(classes)
    out = 0.9 * out.values + 0.1 * (1.0 / K)
    sub = pd.DataFrame(out, columns=classes)
    sub.insert(0, "id", test_ids)
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the affine blend 0.9 * probs + 0.1 / K adds a positive constant to every cell, bounding all probabilities away from zero before the submission is written.
C3 · metric-mismatch Unvalidated equal-weight probability blend

P1 — Unvalidated equal-weight probability blend

• Pattern: Detects averaging of predicted class probabilities from multiple heterogeneous classifiers with fixed uniform weights, where the blend is written to the output file without any out-of-fold or held-out computation of the scored metric to weight, select, or drop component models.
• Detection procedure:
1. Identify two or more distinct classifier objects of different model families assigned in the same script and each fitted on the same training features (e.g. multiple fit(...) calls, or equivalent, on the same feature matrix variable).
2. Identify variables assigned from probability-prediction calls (predict_proba(...) or equivalent) on the test/evaluation features, one per classifier from step 1.
3. Find an expression combining the variables from step 2 by elementwise addition divided by a literal integer count, or by a mean over them with no learned weight vector (e.g. (p1 + p2 + p3) / 3 or np.mean([...], axis=0)), whose result is written into the final output frame or file.
4. Within the same script, search for any computation of a validation score on held-out or out-of-fold predictions (e.g. a call to a loss/scoring function such as log_loss, cross_val_score, cross_val_predict, or an equivalent manual fold loop that scores predictions on rows not used for fitting).
5. The pattern is PRESENT if steps 1–3 all hold and step 4 finds no such validation computation influencing which models are blended or how they are weighted.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: ensemble_pred = (pred1 + pred2 + pred3) / 3 submitted directly with no out-of-fold scoring; a poorly calibrated tree ensemble in the uniform average produced roughly 2x worse log loss (a 0.4996 metric gap) than an out-of-fold-validated stacked blend of the same base models.
• Applies when: A script blends probability outputs from two or more classifiers and the task is scored on a probabilistic metric (e.g. log loss, Brier score, AUC on probabilities).
• Example:
• Input:
    clf1.fit(X_train, y_train)
    clf2.fit(X_train, y_train)
    clf3.fit(X_train, y_train)

    p1 = clf1.predict_proba(X_test)
    p2 = clf2.predict_proba(X_test)
    p3 = clf3.predict_proba(X_test)

    blend = (p1 + p2 + p3) / 3
    sub = pd.DataFrame(blend, columns=classes)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    The scored log loss is substantially worse (observed ~2x higher, a 0.50
    absolute gap) than a validated blend, because a miscalibrated component
    emitting near-0/1 probabilities drags the uniform average toward
    confident wrong predictions with no mechanism to detect or down-weight it.
• Counter-example:
• Input:
    oof = [cross_val_predict(m, X_train, y_train, cv=5,
                             method="predict_proba") for m in models]
    scores = [log_loss(y_train, o) for o in oof]
    keep = [m for m, s in zip(models, scores) if s < 1.5 * min(scores)]
    for m in keep:
        m.fit(X_train, y_train)
    preds = [m.predict_proba(X_test) for m in keep]
    blend = np.mean(preds, axis=0)
    sub = pd.DataFrame(blend, columns=classes)
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the script computes out-of-fold log loss per component and uses it to drop weak models before averaging, so the blend is validated on the scored metric (step 4 finds a validation computation influencing the blend).
C3 · metric-mismatch Hardcoded uncertainty constants never tuned against the metricgithub_occurrence

P1 — Hardcoded uncertainty constants never tuned against the metric

• Pattern: Detects a submission pipeline for a metric that jointly scores prediction error and a submitted uncertainty value, where the uncertainty column is filled from literal numeric constants (a fixed value, or a fixed base plus a fixed per-unit slope) with no code anywhere that evaluates the scoring formula on held-out or out-of-fold residuals to select those constants.
• Detection procedure:
1. Identify the assignment(s) that populate the uncertainty/confidence column of the output frame (e.g., a key such as "Confidence", "sigma", "uncertainty" in the DataFrame written by to_csv or equivalent).
2. Check whether every such assignment is built only from literal numeric constants, optionally combined linearly with a distance/offset term (e.g., 150.0, c0 + c1 * abs(dw) where c0 and c1 are assigned literal numbers and are never reassigned inside any loop).
3. Scan the whole file for a function or expression that implements the competition's uncertainty-aware scoring formula (a computation combining residuals, the uncertainty value, and a log or clip term, or equivalent) applied to validation or out-of-fold predictions.
4. Scan the whole file for a loop or grid enumeration over candidate uncertainty parameter values whose body computes such a score and keeps the argmax/argmin.
5. PRESENT if step 2 holds and neither step 3 nor step 4 finds a match.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: Submission code set the uncertainty column via submission["Confidence"].fillna(150.0) and per-row literal constants with no scoring-formula evaluation anywhere in the file; a variant that grid-searched the base and slope of the confidence against the metric on out-of-fold residuals scored about 0.053 better (≈5% relative) on the same data.
• Applies when: The evaluation metric jointly scores a point prediction and a submitted uncertainty/confidence value per row, and the submission file contains that uncertainty column (regression or tabular forecasting settings).
• Example:
• Input:
    c0, c1 = 200.0, 3.0
    for i, (p, w) in enumerate(zip(patients, weeks)):
        dw = w - base_week[p]
        fvc[i] = base_val[p] + slope[p] * dw
        conf[i] = c0 + c1 * abs(dw)
    out = pd.DataFrame({"Patient_Week": ids,
                        "FVC": fvc,
                        "Confidence": conf})
    out.to_csv("submission.csv", index=False)
• Consequence:
    The uncertainty-aware metric (Gaussian-NLL-style, higher is better) lands
    roughly 5% relative below a solution that tunes the confidence intercept
    and distance slope against the metric on out-of-fold residuals; the
    submitted confidences are systematically mis-calibrated in one direction.
• Counter-example:
• Input:
    best = (-1e9, None)
    for c0 in range(100, 401, 25):
        for c1 in [0.0, 1.0, 2.0, 4.0]:
            sig = np.clip(c0 + c1 * np.abs(oof_dw), 70, 1000)
            d = np.minimum(np.abs(oof_resid), 1000)
            score = np.mean(-np.sqrt(2) * d / sig - np.log(np.sqrt(2) * sig))
            if score > best[0]:
                best = (score, (c0, c1))
    c0, c1 = best[1]
    conf = np.clip(c0 + c1 * np.abs(test_dw), 70, 1000)
    out = pd.DataFrame({"Patient_Week": ids, "FVC": fvc, "Confidence": conf})
• Why it does not fire: the confidence parameters are chosen by enumerating candidates and evaluating the metric formula on out-of-fold residuals, so steps 3 and 4 both find a match and the final assignment is not built from untuned literals.
C3 · metric-mismatch Validation scores computed but ignored in ensemble combinationtrace_observed

P1 — Validation scores computed but ignored in ensemble combination

• Pattern: Detects an ensemble in which a per-candidate validation score is computed for each member model, yet the members' test predictions are combined with an unconditional uniform average that never uses those scores to weight, select, or drop any member.
• Detection procedure:
1. Find a loop (or repeated block) over two or more candidate models in which each model's out-of-sample validation score is computed via cross_val_score, cross_validate, a manual K-fold scoring loop, or equivalent, and the result is assigned to a variable or printed.
2. Within the same loop or block, confirm each model's predictions on the held-out/test inputs are appended to a shared collection (e.g., preds.append(...) or column-stacking of per-model outputs).
3. Find the statement that combines that collection into the final prediction, and confirm it is an unweighted aggregation — np.mean(preds, axis=0), sum(preds)/len(preds), an average with no weights= argument, or equivalent.
4. Confirm that between step 1 and step 3, no variable derived from the validation scores is used in any conditional, weighting expression, sorting, or filtering that affects which predictions enter the combination or with what coefficient.
5. PRESENT when steps 1–4 all hold: scores are computed for every member, predictions are pooled, the pool is combined with fixed equal weights, and the scores influence nothing downstream.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: medium
• Evidence: A pipeline computed sc = cross_val_score(m, Xtr, ytr, cv=cv, scoring="roc_auc") per model and only printed sc.mean(), then combined all members via p = np.mean(preds, axis=0); the measured ranking metric was 0.08114 lower than a comparable pipeline whose combination reflected member quality.
• Applies when: The program trains two or more distinct candidate models, evaluates each with a validation score, and the final output is a combination of their predictions scored on a rank-sensitive metric (e.g., AUC).
• Example:
• Input:
    models = {"a": model_a, "b": model_b}
    preds = []
    for name, m in models.items():
        sc = cross_val_score(m, X_tr, y_tr, cv=cv, scoring="roc_auc")
        print(name, "cv", sc.mean())
        m.fit(X_tr, y_tr)
        preds.append(m.predict_proba(X_te)[:, 1])
    p = np.mean(preds, axis=0)
    sub["LABEL_A"] = p
    sub.to_csv("submission.csv", index=False)
• Consequence:
    A member whose validation AUC is markedly worse than the best member
    contributes with the same weight, dragging the blended test AUC below
    what the best member or a score-weighted blend achieves (observed gap
    on the order of 0.08 in the ranking metric).
• Counter-example:
• Input:
    scores, preds = [], []
    for name, m in models.items():
        sc = cross_val_score(m, X_tr, y_tr, cv=cv, scoring="roc_auc").mean()
        scores.append(sc)
        m.fit(X_tr, y_tr)
        preds.append(m.predict_proba(X_te)[:, 1])
    w = np.array(scores) - 0.5
    w = np.clip(w, 0, None) / np.clip(w, 0, None).sum()
    p = np.average(preds, axis=0, weights=w)
• Why it does not fire: the validation scores flow into weights= of the combining call, so member contributions are conditioned on measured quality (step 4 fails).
C3 · metric-mismatch Uncalibrated probabilities written to log-loss submission

P1 — Uncalibrated probabilities written to log-loss submission

• Pattern: Detects predicted class probabilities from a fitted classifier written directly into a submission file scored by log loss, with no clipping, uniform-blend, or other bounding of extreme values between prediction and write.
• Detection procedure:
1. Locate a call producing per-class probability estimates from a fitted model (e.g. predict_proba, predict(..., softmax output), or equivalent) whose result is assigned to a variable.
2. Trace that variable (and any DataFrame constructed from it) forward within the same script to a call that writes a submission file (e.g. to_csv or equivalent) in a context where the task is scored by log loss or another metric penalizing confident errors unboundedly.
3. Confirm that on every path between step 1 and step 2 there is no operation that bounds the values away from 0 and 1: no clip with a small epsilon, no elementwise mix with a constant such as (1-eps)*p + eps/K, no temperature or calibration transform, no floor applied before renormalization (row renormalization alone does not count).
4. PRESENT when the probability variable from step 1 reaches the file write in step 2 with no bounding operation found in step 3.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: probs = clf.predict_proba(X_test_scaled) written to the submission with only row-sum renormalization; log loss measured at 11.63 versus 4.55 for an otherwise identical pipeline whose only difference was a 0.9/0.1 blend with the uniform distribution — worse than the ~4.79 all-uniform baseline.
• Applies when: the program produces a probability-per-class submission evaluated by log loss (or cross-entropy), especially with a weak or high-variance feature pipeline where confident wrong predictions are likely.
• Example:
• Input:
    clf.fit(X_train, y_train)
    probs = clf.predict_proba(X_test)
    prob_df = pd.DataFrame(probs, columns=classes)
    row_sums = prob_df.sum(axis=1).replace(0, 1)
    prob_df = prob_df.div(row_sums, axis=0)
    sub = pd.concat([pd.DataFrame({"id": test_ids}), prob_df], axis=1)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Log loss inflated to 11.63 vs 4.55 for the same pipeline with a
    uniform blend; overconfident wrong rows contribute near-unbounded
    -log(p) terms, pushing the score below the ~4.79 uniform baseline.
• Counter-example:
• Input:
    clf.fit(X_train, y_train)
    probs = clf.predict_proba(X_test)
    K = probs.shape[1]
    probs = 0.9 * probs + 0.1 / K
    prob_df = pd.DataFrame(probs, columns=classes)
    sub = pd.concat([pd.DataFrame({"id": test_ids}), prob_df], axis=1)
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the elementwise blend 0.9 * probs + 0.1 / K bounds every probability away from 0 and 1 before the write, so step 3's search for a bounding operation succeeds and the conjunction fails.
C3 · metric-mismatch Empty prediction strings for all detection instancesgithub_occurrence

P1 — Empty prediction strings for all detection instances

• Pattern: Detects a submission-building loop for a detection/retrieval task that assigns an empty or constant-empty prediction string to every evaluation instance, producing zero predicted detections overall.
• Detection procedure:
1. Locate the code that constructs the output rows written to the submission file (e.g., a loop over instance identifiers appending to a list later passed to a dataframe writer, or a direct column assignment on the sample-submission frame).
2. Within that construction, inspect every value bound to the prediction field (e.g., a column like PredictionString or the second element of each appended row dict/tuple): check whether all reachable branches assign the empty string "", None, or an unmodified copy of an empty placeholder from the sample file.
3. Confirm that no branch in the construction assigns a value derived from a model output, a training-statistics prior, or any per-instance computation that can yield a non-empty prediction.
4. PRESENT if steps 2 and 3 both hold: every evaluation instance receives an empty prediction and no code path can emit a non-empty one.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A submission loop set pred_string = "" on all branches (with a comment about avoiding false positives) for every evaluation instance; the resulting score was exactly 0, while a sibling program emitting a few constant prior boxes per instance transformed into each frame scored strictly higher (observed gap of 1.0 on the normalized metric).
• Applies when: The task is scored by a metric that rewards true positives (mAP, AP@IoU, F-beta, recall-weighted retrieval scores) and the program writes a per-instance prediction field in its output file.
• Example:
• Input:
    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        # avoid false positives which penalize the score
        pred_string = ""
        predictions.append({'Id': sample_id,
                            'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Consequence:
    Recall is 0 for every instance, so the precision/recall-based metric is
    pinned at 0.0 regardless of upstream computation; emitting even a fixed
    prior detection per instance yields a strictly positive score.
• Counter-example:
• Input:
    strings = []
    for tok in sub["Id"].values:
        pose = poses.get(tok)
        if pose is None:
            strings.append("")
            continue
        strings.append(format_prior_boxes(pose, priors))
    sub["PredictionString"] = strings
    sub.to_csv(OUT, index=False)
• Why it does not fire: only the missing-pose fallback branch is empty; the main branch emits non-empty predictions derived from per-instance data, so step 3 fails.
C3 · metric-mismatch Single default-configured fit with no validation-guided selection for a probability metric

P1 — Single default-configured fit with no validation-guided selection for a probability metric

• Pattern: Detects a pipeline scored by a probabilistic metric that fits exactly one classifier with default regularization on the full training set and writes its raw predicted probabilities to the output, with no held-out or cross-validated computation of the scored metric anywhere in the script to guide model, hyperparameter, or blend choices.
• Detection procedure:
1. Confirm the task is scored by a probabilistic metric: the output file contains per-class probability columns or a probability column (e.g., values written from a call producing class probabilities such as predict_proba or equivalent), and the task context specifies log loss, AUC, or Brier score.
2. Count calls to fitting methods (.fit(...) or equivalent) on classifier objects in the script; find exactly one, and its constructor call passes no regularization or complexity hyperparameters (e.g., no C=, alpha=, max_depth=, or equivalent tuning arguments beyond iteration/seed settings).
3. Search the script for any train/validation split or fold construct (train_test_split, KFold, StratifiedKFold, cross_val_score, cross_val_predict, GridSearchCV, or equivalent); find none, or find one whose resulting metric value is never compared against an alternative configuration (no loop over hyperparameter values, model constructors, or blend weights that selects by the metric).
4. Confirm the probabilities written to the output come directly from the single fitted model of step 2, with no averaging, weighting, or calibration step in between.
5. PRESENT when all of steps 1–4 hold: one default-configured fit, no metric-guided selection, raw single-model probabilities submitted.
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 2
• weight: 2
• confidence: high
• Evidence: A script calling model = LogisticRegression(max_iter=1000); model.fit(X_train_tfidf, y_train) and writing model.predict_proba(X_test_tfidf) directly to the submission scored 0.576 log loss, while a variant of the same feature pipeline that computed stratified out-of-fold predictions for three model families and grid-searched blend weights against the metric scored 0.383 — a 0.336 degradation from skipping validation-guided selection.
• Applies when: The program produces a prediction file for a task scored by log loss, AUC, Brier score, or another probability-sensitive metric, and the time budget permits fitting the model more than once (i.e., cross-validation is feasible).
• Example:
• Input:
    vec = TfidfVectorizer(ngram_range=(1, 2))
    Xtr = vec.fit_transform(train_df["text_col"])
    Xte = vec.transform(test_df["text_col"])
    model = LogisticRegression(max_iter=1000, random_state=42)
    model.fit(Xtr, train_df["target"])
    proba = model.predict_proba(Xte)
    sub = pd.DataFrame(proba, columns=model.classes_)
    sub.insert(0, "id", test_df["id"])
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Log loss on the leaderboard is ~0.58 versus ~0.38 for the same features
    with cross-validated hyperparameter and blend selection: a ~0.20 worse
    score from submitting unoptimized default-regularization probabilities.
• Counter-example:
• Input:
    best = None
    for C in [0.1, 0.5, 1.0, 4.0]:
        skf = StratifiedKFold(5, shuffle=True, random_state=42)
        oof = np.zeros((len(y), 3))
        for tr, va in skf.split(Xtr, y):
            m = LogisticRegression(C=C, max_iter=1000).fit(Xtr[tr], y[tr])
            oof[va] = m.predict_proba(Xtr[va])
        s = log_loss(y, oof)
        if best is None or s < best[0]:
            best = (s, C)
    final = LogisticRegression(C=best[1], max_iter=1000).fit(Xtr, y)
    sub = pd.DataFrame(final.predict_proba(Xte), columns=final.classes_)
• Why it does not fire: the script computes an out-of-fold estimate of the scored metric and uses it to select the regularization strength before the final fit, so step 3's "no metric-guided comparison" condition fails.
C3 · metric-mismatch Hardcoded uncertainty constants never tuned against the scoring formula

P1 — Hardcoded uncertainty constants never tuned against the scoring formula

• Pattern: Detects a submitted uncertainty/confidence value computed purely from hardcoded numeric literals (a fixed base, optionally plus a fixed coefficient times a distance term) with no search over those literals evaluated against the task's scoring formula on held-out prediction errors.
• Detection procedure:
1. Locate the statement that writes the output file (a call to to_csv, to_parquet, or equivalent) and identify the column, array, or field holding the uncertainty/confidence value alongside the point prediction.
2. Trace every assignment feeding that uncertainty value within the program; the pattern requires that each such assignment is either a numeric literal, or an arithmetic expression whose only non-literal operand is a distance/gap quantity (e.g., a literal plus a literal multiplied by an absolute difference of two values).
3. Search the entire source file for any loop or comprehension that iterates over candidate values for those literals (e.g., for c0 in, np.arange, itertools.product, or equivalent) whose body evaluates an expression combining held-out or out-of-fold residuals with the candidate uncertainty value (a clipped/penalized error or likelihood term) and keeps an argmax/argmin.
4. PRESENT if step 2 holds (uncertainty derives solely from hardcoded literals) and step 3 finds no such search anywhere in the source.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: conf[i] = c0 + c1 * abs(dw) with c0, c1 fixed constants and no tuning loop; a contrastive run that grid-searched the same two parameters against the exact scoring formula on out-of-fold errors scored about 5% better relative on the joint point-plus-uncertainty metric.
• Applies when: the task's evaluation metric jointly scores a point prediction and a submitted uncertainty/confidence value (a likelihood-style or penalized-error metric), and the program produces both outputs for a regression-style target.
• Example:
• Input:
    fvc = np.zeros(len(sub))
    conf = np.zeros(len(sub))
    for i, (p, w) in enumerate(zip(patients, weeks)):
        r = base.loc[p]
        dw = w - r["base_week"]
        fvc[i] = r["base_val"] + r["pred_slope"] * dw
        conf[i] = 200.0 + 5.0 * abs(dw)   # fixed heuristic constants
    out = pd.DataFrame({"id": sub["id"], "pred": fvc, "sigma": conf})
    out.to_csv("submission.csv", index=False)
• Consequence:
    Program completes and produces a valid submission, but the likelihood-style
    metric is systematically worse (~5% relative) than the same pipeline with
    the two uncertainty constants grid-searched against the exact scoring
    formula on held-out errors, because the fixed sigma is miscalibrated to
    the actual error distribution.
• Counter-example:
• Input:
    best = (None, -1e18)
    for c0 in np.arange(100, 400, 10):
        for c1 in np.arange(0, 10, 0.5):
            sig = np.clip(c0 + c1 * np.abs(oof_dw), 70, None)
            d = np.minimum(np.abs(oof_err), 1000)
            score = np.mean(-np.sqrt(2)*d/sig - np.log(np.sqrt(2)*sig))
            if score > best[1]:
                best = ((c0, c1), score)
    (c0, c1) = best[0]
    conf = c0 + c1 * np.abs(dw)
• Why it does not fire: the uncertainty still uses the base-plus-distance form, but the constants are chosen by a search that evaluates the exact scoring formula on out-of-fold errors, so step 3 finds the tuning loop.
C3 · metric-mismatch No held-out estimate of the scored metric before writing predictions

P1 — No held-out estimate of the scored metric before writing predictions

• Pattern: Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
• Detection procedure:
1. Locate every call to a model-fitting method (e.g. fit or equivalent) in the script; confirm each such call receives features and labels derived from the full training frame, with no row-subset indexing (no fold indices, no boolean masks, no slice of a shuffled index) applied to the labels before fitting.
2. Search the entire script for any call to a training/validation splitting utility (train_test_split, KFold, StratifiedKFold, cross_val_score, cross_val_predict, or equivalent) or any manual construction of disjoint train/validation index arrays; confirm none exists.
3. Search the entire script for any call to a metric function (log_loss, accuracy_score, roc_auc_score, a score method on the fitted model applied to held-back rows, or equivalent) whose arguments include true labels; confirm none exists.
4. Confirm the script calls a prediction method on features derived from a test file and passes the result to a file-writing call (to_csv or equivalent).
5. PRESENT when all of steps 1–4 hold: full-data fit, no split utility, no metric evaluation against true labels, and test predictions written to disk.
• Predicted impact:
• Add score for C3: 2
• weight: 2
• confidence: high
• Evidence: A script consisting only of model.fit(X_train_full, y) followed by model.predict_proba(X_test) and to_csv, with no split or metric call anywhere, scored 0.44 worse on the official probabilistic metric than a pipeline on the same data that computed out-of-fold metric values and used them to select among configurations.
• Applies when: A supervised-learning script produces a prediction file scored by an external metric and the training labels are available in the script, making a held-out estimate of that metric computable.
• Example:
• Input:
    import pandas as pd
    from sklearn.linear_model import LogisticRegression

    train = pd.read_csv("train.csv")
    test = pd.read_csv("test.csv")
    X, y = train[["col_a", "col_b"]], train["target"]

    model = LogisticRegression(C=1.0, max_iter=1000)
    model.fit(X, y)
    preds = model.predict_proba(test[["col_a", "col_b"]])
    pd.DataFrame(preds).to_csv("submission.csv", index=False)
• Consequence:
    The regularization strength, feature set, and model family are never
    checked against the scored metric; the submitted log loss lands
    substantially higher (e.g. +0.4) than a configuration selected via
    out-of-fold measurement on the same data.
• Counter-example:
• Input:
    import pandas as pd
    from sklearn.linear_model import LogisticRegression
    from sklearn.model_selection import cross_val_predict
    from sklearn.metrics import log_loss

    train = pd.read_csv("train.csv")
    test = pd.read_csv("test.csv")
    X, y = train[["col_a", "col_b"]], train["target"]

    model = LogisticRegression(C=1.0, max_iter=1000)
    oof = cross_val_predict(model, X, y, method="predict_proba", cv=5)
    print("cv log loss:", log_loss(y, oof))
    model.fit(X, y)
    pd.DataFrame(model.predict_proba(test[["col_a", "col_b"]])).to_csv("submission.csv", index=False)
• Why it does not fire: the script computes an out-of-fold estimate of the scored metric via a cross-validation utility and a metric call on true labels before writing the final predictions, so steps 2 and 3 fail.
C3 · metric-mismatch Per-class binary decomposition with prior-distorting reweighting under a probability-scored multiclass metricgithub_occurrence

P1 — Per-class binary decomposition with prior-distorting reweighting under a probability-scored multiclass metric

• Pattern: Detects a mutually exclusive multiclass target, scored by a metric evaluating predicted probabilities, being modeled as multiple independently fitted binary classifiers whose training reweights classes away from their empirical priors and whose per-class probabilities are renormalized to sum to one after prediction.
• Detection procedure:
1. Identify two or more classifier objects constructed within the same script, each fitted on a binary indicator target derived from mutually exclusive outcome columns or a one-vs-rest encoding of a single label (e.g. separate fit calls each receiving a different 0/1 label vector over the same rows).
2. Check that at least one of these classifiers is constructed with a class-rebalancing option such as class_weight='balanced', a per-class sample-weight vector inversely proportional to class frequency, or an equivalent reweighting argument.
3. Find, after the per-classifier probability predictions (via predict_proba or equivalent), an arithmetic renormalization step: the positive-class probabilities are summed across the classifiers and each is divided by that sum before being written to the output.
4. Confirm that the evaluation or objective mentioned in the script (a probability-based loss such as log loss / cross-entropy, or probability columns in the output artifact) scores the predicted probabilities themselves rather than hard labels.
5. The pattern is PRESENT when steps 1, 3, and 4 all hold and step 2 holds for at least one of the binary classifiers.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: Three binary classifiers each built with class_weight='balanced' were fitted on separate 0/1 indicators of a three-way outcome, their positive-class probabilities divided by their sum before writing the submission; the resulting probabilities were systematically shifted away from the true class priors, and a single jointly fitted multiclass model on the same features scored about 0.039 better on the probability-based competition metric.
• Applies when: The task has a single mutually exclusive multiclass outcome and the score is computed from predicted class probabilities (log loss / cross-entropy or similar), so calibration of the joint probability vector matters.
• Example:
• Input:
    m_a = LogisticRegression(class_weight='balanced')
    m_b = LogisticRegression(class_weight='balanced')
    m_c = LogisticRegression(class_weight='balanced')
    m_a.fit(X_tr, y_a); m_b.fit(X_tr, y_b); m_c.fit(X_tr, y_c)
    p_a = m_a.predict_proba(X_te)[:, 1]
    p_b = m_b.predict_proba(X_te)[:, 1]
    p_c = m_c.predict_proba(X_te)[:, 1]
    tot = p_a + p_b + p_c + 1e-10
    out = pd.DataFrame({'LABEL_A': p_a/tot, 'LABEL_B': p_b/tot,
                        'LABEL_C': p_c/tot})
• Consequence:
    Predicted probability vectors are pushed toward uniform class priors by the
    balanced reweighting, then distorted further by post-hoc renormalization;
    log loss on held-out data is roughly 0.04 worse than a single multiclass
    model fitted on identical features.
• Counter-example:
• Input:
    y = df['label'].map({'LABEL_A': 0, 'LABEL_B': 1, 'LABEL_C': 2})
    model = LogisticRegression(multi_class='multinomial', max_iter=1000)
    model.fit(X_tr, y)
    probs = model.predict_proba(X_te)
    out = pd.DataFrame({'LABEL_A': probs[:, 0], 'LABEL_B': probs[:, 1],
                        'LABEL_C': probs[:, 2]})
• Why it does not fire: A single multiclass model is fitted on one categorical label with no class reweighting and no post-hoc renormalization, so its probabilities are jointly calibrated by the softmax objective.
C3 · metric-mismatch Class rebalancing under a proper scoring rulegithub_occurrence

P1 — Class rebalancing under a proper scoring rule

• Pattern: Detects a classifier fitted with class rebalancing (weighted classes or a resampled training set) whose raw predicted probabilities are submitted directly for evaluation under a probability-scoring metric such as log loss, without any post-hoc recalibration to the original class priors.
• Detection procedure:
1. Locate every classifier constructor or fit call in the source; mark those that pass a class-weighting argument (e.g. class_weight='balanced', class_weight={...}, scale_pos_weight, is_unbalance=True, or equivalent) or whose training frame is built by oversampling/undersampling rows of one class (e.g. resample, SMOTE, repeated concat of a class subset, or equivalent).
2. Confirm the evaluation target is a proper scoring rule over probabilities: the script writes probability columns to the output file (values from predict_proba, predict returning per-class probabilities, or equivalent), or the metric is computed with log_loss, a multi_logloss/binary_logloss objective metric, Brier score, or equivalent.
3. Verify that between the probability prediction at step 2 and the final output there is no recalibration step: no calibration wrapper (e.g. CalibratedClassifierCV or equivalent), no prior-correction arithmetic on the probabilities (multiplying/dividing by class frequencies followed by renormalization), and no isotonic/Platt refit on held-out data.
4. PRESENT when a marked rebalanced model from step 1 produces the probabilities identified in step 2 and step 3 confirms those probabilities reach the output unrecalibrated.
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 2
• weight: 2
• confidence: high
• Evidence: Classifiers constructed with class_weight='balanced' and whose predict_proba outputs were normalized and written directly as the submission scored measurably worse (log-loss gap ≈ 0.039) than an unweighted fit on the same features, because rebalancing shifts predicted probabilities away from the true empirical class frequencies.
• Applies when: A classification task is scored by log loss, Brier score, or another proper scoring rule over predicted probabilities, and the class distribution is unequal (so rebalancing actually alters the effective priors).
• Example:
• Input:
    from sklearn.linear_model import LogisticRegression
    model = LogisticRegression(class_weight='balanced', max_iter=1000)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
    submission = pd.DataFrame(probs, columns=['LABEL_A', 'LABEL_B', 'LABEL_C'])
    submission.insert(0, 'id', test_df['id'])
    submission.to_csv('submission.csv', index=False)
• Consequence:
    Predicted probabilities are pulled toward a uniform class distribution
    rather than the true skewed priors; minority classes are systematically
    over-weighted, and log loss on the evaluation set rises (~0.03-0.04 worse)
    compared with an unweighted fit on identical features.
• Counter-example:
• Input:
    from sklearn.linear_model import LogisticRegression
    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
    submission = pd.DataFrame(probs, columns=['LABEL_A', 'LABEL_B', 'LABEL_C'])
    submission.insert(0, 'id', test_df['id'])
    submission.to_csv('submission.csv', index=False)
• Why it does not fire: The classifier is fitted without class weighting or resampling, so its probabilities reflect the empirical class priors and step 1 finds no rebalanced model.
C3 · metric-mismatch Blind probability submission without held-out metric evaluation

P1 — Blind probability submission without held-out metric evaluation

• Pattern: Detects a script that fits a probabilistic classifier on the full training data and writes predicted class probabilities to a submission file without ever computing a probability-based loss on held-out data or performing any validation-driven selection of model, hyperparameters, or ensemble weights.
• Detection procedure:
1. Locate a call that produces per-class probability predictions (e.g. predict_proba, predict(..., output_probabilities=True), or equivalent) whose result is written to an output file via to_csv, to_parquet, or equivalent.
2. Search the entire script for any import or call of a probability-scoring function (e.g. log_loss, roc_auc_score, brier_score_loss, or equivalent) applied to labels and predictions.
3. Search the entire script for any validation-split construct: a train/validation split call (e.g. train_test_split or equivalent), a k-fold iterator (e.g. KFold, StratifiedKFold, cross_val_score, cross_val_predict, or equivalent), or a hyperparameter search object (e.g. GridSearchCV, RandomizedSearchCV, or equivalent).
4. Check whether the fitting call at step 1 receives the complete labeled dataset (no index subsetting or split output) and the fitted model's probabilities feed the output file directly.
5. PRESENT when step 1 succeeds and both step 2 and step 3 find no occurrences anywhere in the script.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A script that ran model.fit(X_train_tfidf, y_train_encoded) on all labeled rows and wrote model.predict_proba(X_test_tfidf) straight to the submission, with no validation split or loss computation anywhere, scored 0.33507 worse on the probability-loss metric than a variant of the same model family that used stratified out-of-fold predictions to evaluate the loss and select ensemble weights.
• Applies when: The task is scored on a probability-quality metric (log loss or similar) and the script produces a probability submission from labeled training data.
• Example:
• Input:
    vec = TfidfVectorizer(max_features=20000)
    X = vec.fit_transform(train["text"])
    Xt = vec.transform(test["text"])
    model = LogisticRegression(max_iter=1000)
    model.fit(X, train["label"])
    proba = model.predict_proba(Xt)
    sub = pd.DataFrame(proba, columns=classes)
    sub.insert(0, "id", test["id"])
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Model choice, regularization strength, and calibration are never checked
    against the scored loss; submitted log loss is systematically higher
    (observed gap: 0.335 worse) than metric-guided selection over the same
    model family would achieve.
• Counter-example:
• Input:
    skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
    oof = np.zeros((len(train), n_classes))
    test_pred = np.zeros((len(test), n_classes))
    for tr, va in skf.split(X, y):
        m = LogisticRegression(max_iter=1000).fit(X[tr], y[tr])
        oof[va] = m.predict_proba(X[va])
        test_pred += m.predict_proba(Xt) / skf.n_splits
    print("oof log loss:", log_loss(y, oof))
    pd.DataFrame(test_pred, columns=classes).to_csv("submission.csv", index=False)
• Why it does not fire: The script computes out-of-fold probabilities via a k-fold iterator and evaluates log_loss on held-out predictions, so steps 2 and 3 both find occurrences.
C3 · metric-mismatch Unclipped probabilities under an unbounded log-based metric

P1 — Unclipped probabilities under an unbounded log-based metric

• Pattern: Detects probabilistic class predictions written directly to an output file scored by log loss (or a similar unbounded proper scoring rule) without any clipping away from 0 and 1, smoothing, or blending toward a uniform prior.
• Detection procedure:
1. Identify a variable assigned from a probability-producing prediction call on a fitted classifier (e.g., predict_proba, a softmax over model outputs, or equivalent) within the script.
2. Trace every subsequent expression through which that variable (or a DataFrame/array built from it) flows before it reaches a file-writing call such as to_csv, savetxt, or equivalent.
3. Check whether any of those expressions bounds the values away from the extremes: a call to clip/np.clip (or equivalent) with a positive lower bound, a convex combination with a constant (e.g., a * p + b where b > 0), addition of a positive constant followed by renormalization, or a temperature/power transform followed by renormalization.
4. Confirm the task is evaluated by log loss or another metric that is unbounded as a predicted probability for the true class approaches 0 (stated in comments, metric imports such as log_loss, or the surrounding task context).
5. PRESENT if the probability variable reaches the file-writing call with none of the bounding transformations from step 3 applied, and the condition in step 4 holds.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: probs = clf.predict_proba(X_test_scaled) written to the submission after only column reordering and row renormalization — no clipping or prior blending — produced a log loss of 11.92, versus 4.55 for an otherwise identical pipeline that blended predictions with a uniform prior; the unbounded version scored worse than a uniform-prediction baseline (~4.79).
• Applies when: A multiclass or binary probabilistic prediction task is scored by log loss or another proper scoring rule with unbounded penalty for confident wrong predictions, and the model can emit probabilities at or near 0 or 1.
• Example:
• Input:
    probs = clf.predict_proba(X_test)
    prob_df = pd.DataFrame(probs, columns=classes)
    prob_df = prob_df[submission_cols]
    row_sums = prob_df.sum(axis=1)
    prob_df = prob_df.div(row_sums, axis=0)
    sub = pd.concat([pd.DataFrame({"id": test_ids}), prob_df], axis=1)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Rows where the model assigns near-zero probability to the true class each
    contribute -log(p) ≈ 15-35 to the mean log loss; overall score degrades
    from ~4.5 (with smoothing) to ~11.9, worse than a uniform-prediction
    baseline of ~4.8.
• Counter-example:
• Input:
    probs = clf.predict_proba(X_test)
    K = probs.shape[1]
    probs = 0.9 * probs + 0.1 / K
    prob_df = pd.DataFrame(probs, columns=classes)
    sub = pd.concat([pd.DataFrame({"id": test_ids}), prob_df], axis=1)
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: the convex blend with a uniform prior (0.9 * probs + 0.1 / K) bounds every predicted probability away from 0, capping the per-row penalty, so step 3's condition is satisfied.
C3 · metric-mismatch Arithmetic mean of wrap-around anglesgithub_occurrence

P1 — Arithmetic mean of wrap-around angles

• Pattern: Detects aggregation of angular quantities (yaw, heading, orientation, phase) with a plain arithmetic mean or median over raw angle values instead of a circular mean of their sine/cosine components.
• Detection procedure:
1. Identify variables holding angles: any variable whose name contains yaw, heading, theta, angle, phase, or rot, or that is assigned from a call to atan2/arctan2 (or equivalent), or that is passed through a wrap-to-range operation (modulo 2*pi, or an add/subtract-2*pi normalization).
2. Find any call, within the same script, that aggregates one of these variables across multiple samples: np.mean, np.median, statistics.mean, a pandas .mean() on a column or groupby(...).mean() (or equivalent), or a manual sum(x)/len(x) where x is the angle collection.
3. Check whether, between step 1 and step 2, the angle values are transformed into sine and cosine components (calls to sin and cos on the variable) that are aggregated separately and recombined via atan2/arctan2, or the aggregation is performed by a dedicated circular-statistics routine (e.g. circmean or equivalent).
4. PRESENT if step 2 finds an aggregation of raw angle values and step 3 finds no sine/cosine decomposition or circular-statistics routine applied to that aggregation.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: A pipeline computed a representative orientation per group with np.mean(yaws) on raw wrapped angles; a comparable pipeline that aggregated orientations correctly scored 0.877 higher on the overlap-based competition metric, since angles clustered near the ±π boundary averaged to a value pointing roughly opposite the true orientation.
• Applies when: The program's output or an intermediate statistic includes an orientation, heading, or phase value that is computed by aggregating multiple angle observations across samples or group members.
• Example:
• Input:
    import numpy as np

    def summarize_group(rows):
        yaws = np.array([r["yaw"] for r in rows])  # radians in [-pi, pi]
        return {
            "cx": np.mean([r["x"] for r in rows]),
            "cy": np.mean([r["y"] for r in rows]),
            "yaw": np.mean(yaws),
        }
• Consequence:
    For groups whose yaws straddle the ±π wrap (e.g. 3.1 and -3.1 rad), the
    aggregated yaw comes out near 0 rad — roughly 180° off the true heading —
    so oriented-box overlap with ground truth drops sharply and the
    overlap-based evaluation score is systematically lower than with a
    circular mean.
• Counter-example:
• Input:
    import numpy as np

    def summarize_group(rows):
        yaws = np.array([r["yaw"] for r in rows])
        mean_yaw = np.arctan2(np.mean(np.sin(yaws)),
                              np.mean(np.cos(yaws)))
        return {
            "cx": np.mean([r["x"] for r in rows]),
            "cy": np.mean([r["y"] for r in rows]),
            "yaw": mean_yaw,
        }
• Why it does not fire: the angles are decomposed into sine and cosine components that are averaged separately and recombined with arctan2, which is the circular mean, so step 3 finds the required transformation and the pattern is absent.
C3 · metric-mismatch Empty prediction strings submitted under an average-precision metric

P1 — Empty prediction strings submitted under an average-precision metric

• Pattern: Detects a submission-writing routine for a detection task scored by average precision that assigns an empty or blank prediction string to every row, emitting no candidate detections at all.
• Detection procedure:
1. Locate the code that builds the per-sample prediction column of the output file (e.g., a loop or comprehension appending rows with a field such as PredictionString, or a direct column assignment on a dataframe before to_csv or equivalent).
2. Check whether the value assigned to that prediction field is the empty string "", None, np.nan, or a whitespace-only literal, on every control-flow path that reaches the row-append or column assignment (i.e., all branches of any if/else assign an empty/blank value, or there is a single unconditional empty assignment).
3. Confirm that no other code path between that assignment and the file write populates the prediction field with detection tokens (no concatenation, join of a non-empty list, or reassignment from a model output).
4. The pattern is PRESENT when the written prediction field is empty/blank for all rows (steps 2 and 3 both hold) and the task is evaluated by average precision or mean average precision.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: The offending program set pred_string = "" on every branch of its per-sample loop, with a comment claiming this "avoids false positives which heavily penalize the score"; the resulting submission scored 0.0 under the mAP-style metric, a full 1.0 below a sibling run that emitted heuristic detections with confidence scores, because average precision cannot exceed zero when recall is zero.
• Applies when: The program produces a submission or results file for a detection task whose evaluation metric is average precision, mean average precision, or another precision-recall-curve metric averaged over predicted detections.
• Example:
• Input:
    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        if sample_id in test_token_to_sample:
            # Empty predictions to avoid false positives penalizing the score
            pred_string = ""
        else:
            pred_string = ""
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Consequence:
    Recall is 0 for every class, so average precision is 0.0 for all classes and
    the final mAP score is pinned at the metric's minimum — strictly below any
    run that emits even weak candidate detections, which scored 1.0 higher on
    the same evaluation.
• Counter-example:
• Input:
    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        boxes = predict_boxes(sample_id)  # list of (conf, x, y, z, w, l, h, yaw, cls)
        if boxes:
            pred_string = " ".join(format_box(b) for b in boxes)
        else:
            pred_string = ""  # only when the model truly found nothing here
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Why it does not fire: An empty string is assigned only on the branch where the model returned no boxes for that individual sample, while another reachable path populates the prediction field from model output, so the field is not blank on every control-flow path.
C3 · metric-mismatch Constant confidence for ranked-detection scoringverified_trace · effect +0.0002

P1 — Constant confidence for ranked-detection scoring

• Pattern: Detects submission-writing code that assigns the same hard-coded constant as the confidence/score field of every emitted prediction when the output is evaluated by a confidence-ranked metric such as average precision, so no prediction can be ranked above any other.
• Detection procedure:
1. Locate the code that builds per-prediction output records (rows appended to a list, strings formatted into a prediction column, or dicts written to a results file) that include a confidence or score field alongside geometric/label fields.
2. Trace the value written into that confidence field back to its assignment: it is PRESENT-relevant if the value is a numeric literal (e.g. 1.0, 0.5) or a variable assigned a numeric literal once and never reassigned before use, and the assignment is not inside any branch that varies per prediction.
3. Confirm the confidence value does not depend on any per-prediction quantity (loop variable, model output, count, distance, frequency, or other statistic that differs across emitted predictions).
4. Confirm the surrounding task scores the output with a ranking-sensitive metric (average precision, mAP, precision-recall over ranked detections) — indicated by the output schema containing a confidence/score field per prediction.
5. PRESENT when all predictions receive an identical constant confidence (steps 2–3) and the output format includes a confidence field consumed by a ranked metric (step 4).
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 3
• weight: 3
• confidence: high
• Evidence: Every emitted detection used conf = 1.0 regardless of how the candidate was generated; the confidence-ranked metric fell roughly 8x (7e-05 vs 5.7e-04) compared with an otherwise similar solution that derived confidence from a per-candidate support count.
• Applies when: The program writes a prediction file whose per-prediction records include a confidence or score field, and the task is scored by a metric that ranks predictions by that confidence (average precision or similar precision-recall ranking).
• Example:
• Input:
    pred_tokens = []
    for cand in candidates:
        x, y, z = cand["x"], cand["y"], cand["z"]
        conf = 1.0
        pred_tokens.append(f"{conf} {x:.3f} {y:.3f} {z:.3f} LABEL_A")
    results.append({"Id": sample_id, "PredictionString": " ".join(pred_tokens)})
    pd.DataFrame(results).to_csv("submission.csv", index=False)
• Consequence:
    Average precision drops sharply (observed ~8x lower) because weak
    candidates tie with strong ones at identical rank, interleaving false
    positives among true positives in the precision-recall sweep.
• Counter-example:
• Input:
    pred_tokens = []
    max_n = max(c["n"] for c in candidates)
    for cand in candidates:
        x, y, z = cand["x"], cand["y"], cand["z"]
        conf = max(0.01, min(1.0, cand["n"] / max_n))
        pred_tokens.append(f"{conf} {x:.3f} {y:.3f} {z:.3f} LABEL_A")
    results.append({"Id": sample_id, "PredictionString": " ".join(pred_tokens)})
    pd.DataFrame(results).to_csv("submission.csv", index=False)
• Why it does not fire: The confidence is derived from a per-candidate support statistic (cand["n"]) and varies across predictions, so stronger candidates rank above weaker ones.
C3 · metric-mismatch Hard-coded uncertainty value in likelihood-scored submission

P1 — Hard-coded uncertainty value in likelihood-scored submission

• Pattern: Detects a submitted uncertainty/confidence column populated from literal numeric constants or hand-written arithmetic formulas, with no code path that evaluates or optimizes the uncertainty value against residuals from training or validation predictions.
• Detection procedure:
1. Identify the statement that writes the final output file (a call to to_csv or equivalent) and the frame it writes; confirm that frame contains a column whose name contains confidence, sigma, std, or uncertainty (case-insensitive).
2. Trace every assignment to that column within the script. Mark it CONSTANT-DERIVED if the assigned value is a numeric literal, a call combining literals such as max(x, <literal>) or <literal> + <literal> * expr, or a fill of missing values with a numeric literal.
3. Search the whole script for any computation of residuals — an expression subtracting predicted values from observed target values on training or held-out rows — that is subsequently used inside a loop, grid search, or call to a numerical optimizer (e.g. scipy.optimize.minimize or equivalent) whose objective references the uncertainty value.
4. PRESENT if step 2 finds only CONSTANT-DERIVED assignments and step 3 finds no residual-based evaluation or optimization feeding the uncertainty value.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: An uncertainty column set via constants like (pid, median_fvc, 150.0) and fillna(150.0) scored measurably worse (~0.029 absolute on the competition metric, ~3% relative) than an otherwise similar pipeline that chose the uncertainty by minimizing the negative mean of the exact metric over training residuals.
• Applies when: The task's scoring function jointly penalizes a point prediction and a submitted per-row uncertainty value (e.g. a Gaussian log-likelihood or pinball-style metric), and the submission file includes that uncertainty column.
• Example:
• Input:
    preds = model.predict(X_test)
    submission = pd.DataFrame({
        "id": test_ids,
        "value": preds,
        "Confidence": 150.0,   # fixed guess
    })
    submission["Confidence"] = submission["Confidence"].fillna(150.0)
    submission.to_csv("submission.csv", index=False)
• Consequence:
    The likelihood-based score is ~3% worse than the same predictions
    submitted with an uncertainty tuned on training residuals, because the
    fixed sigma is miscalibrated relative to the score-maximizing value.
• Counter-example:
• Input:
    resid = train["value"] - model.predict(X_train)
    def neg_metric(sigma):
        s = np.maximum(sigma, 70)
        d = np.minimum(np.abs(resid), 1000)
        return np.mean(np.sqrt(2) * d / s + np.log(np.sqrt(2) * s))
    best_sigma = minimize(neg_metric, x0=[200.0], method="Nelder-Mead").x[0]
    submission["Confidence"] = best_sigma
    submission.to_csv("submission.csv", index=False)
• Why it does not fire: The uncertainty value is chosen by numerically optimizing the actual scoring formula over residuals computed on training data, so step 3 finds a residual-based optimization feeding the submitted column.
C3 · metric-mismatch Single hardcoded prediction per multi-object samplegithub_occurrence

P1 — Single hardcoded prediction per multi-object sample

• Pattern: Detects a variable-cardinality set-prediction task (many objects of multiple classes per input) collapsed into exactly one predicted object per input with a single string-literal class label shared by all outputs.
• Detection procedure:
1. Locate the loop that appends one output row per test sample (e.g., appending a dict with an id field and a prediction-string field to a list later written with to_csv or equivalent).
2. Within that loop body, confirm the prediction-string field is built by a single formatting expression (one f-string, format call, or concatenation) that encodes exactly one object's values, with no inner loop over multiple objects, classes, or counts that accumulates additional tokens into the same field.
3. Confirm the class-label token inside that formatting expression is a string literal written directly in the source (not a variable derived from per-sample model output), so every sample receives the identical class.
4. Confirm the training annotations permit multiple objects per sample: the ground-truth table is grouped or filtered so that only one annotation row per sample id is retained (e.g., groupby(...).first(), taking index [0] of a per-sample annotation list, or a merge that drops duplicate sample ids) before fitting regressors of per-object coordinates against one shared per-sample feature vector.
5. PRESENT when steps 2, 3, and 4 all hold for the loop found in step 1.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A submission builder emitting pred_string = f"{conf} {x} {y} {z} {w} {l} {h} {yaw} LABEL_A" once per sample, with regressors fitted on one annotation per sample, scored 1.0 worse on the competition metric than a variant emitting multiple objects across classes per sample; recall was capped at one object per scene and every non-hardcoded class contributed zero.
• Applies when: The task's ground truth contains a variable number of labeled objects per input, spanning more than one class, and the evaluation metric is an mAP-style or set-matching score over all predicted objects.
• Example:
• Input:
    preds = []
    for s in test_samples:
        feats = extract_features(s)              # one vector per scene
        x = model_x.predict([feats])[0]
        y = model_y.predict([feats])[0]
        z = model_z.predict([feats])[0]
        pred = f"1.0 {x:.3f} {y:.3f} {z:.3f} 1.9 4.5 1.6 0.0 LABEL_A"
        preds.append({"Id": s["token"], "PredictionString": pred})
    pd.DataFrame(preds).to_csv("submission.csv", index=False)
• Consequence:
    Recall is capped at one object per scene even when scenes contain dozens
    of objects, and every class other than the hardcoded one scores exactly
    zero AP; the mAP-style metric drops to near or exactly 0 versus a
    baseline that emits multiple per-class predictions per scene.
• Counter-example:
• Input:
    preds = []
    for s in test_samples:
        tokens = []
        for cls, stats in class_stats.items():          # every class
            for _ in range(stats["avg_count"]):         # variable count
                x, y, z = sample_position(s, stats)
                tokens.append(f"1.0 {x:.3f} {y:.3f} {z:.3f} "
                              f"{stats['w']} {stats['l']} {stats['h']} 0.0 {cls}")
        preds.append({"Id": s["token"], "PredictionString": " ".join(tokens)})
    pd.DataFrame(preds).to_csv("submission.csv", index=False)
• Why it does not fire: the prediction field is accumulated over an inner loop across classes and per-class counts, so multiple objects with variable class labels are emitted per sample, failing steps 2 and 3.
C3 · metric-mismatch Empty prediction strings under a recall-bounded metricgithub_occurrence

P1 — Empty prediction strings under a recall-bounded metric

• Pattern: Detects a submission-writing loop that assigns an empty string (or empty list) as the prediction field for every evaluated sample when the scoring metric is precision-recall based, deliberately emitting zero positive predictions.
• Detection procedure:
1. Locate the code that builds the output rows written to the submission file (a loop or comprehension appending dictionaries/rows containing an identifier column and a prediction column, later passed to a CSV/JSON writer).
2. Within that loop or comprehension, check whether the prediction field is assigned a literal empty value — "", [], or a constant variable that was only ever assigned an empty literal — on every branch that reaches the append (i.e., no branch assigns a non-empty prediction derived from data or a model).
3. Confirm there is no other code path in the file that populates the prediction column with non-empty content before the file is written (no later assignment, fillna with non-empty content, merge, or update targeting that column).
4. The pattern is PRESENT when all rows written to the submission receive an empty prediction field and the task context (comments, metric name, per-object prediction-string format) indicates scoring by average precision or a similar precision-recall metric over predicted objects.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: The offending program set pred_string = "" on every branch of its submission loop, with a comment claiming this "avoids false positives which heavily penalize the score"; the measured score was exactly 0, while a sibling program emitting crude prior-based candidate objects scored 1.0 higher on the same metric.
• Applies when: The program produces a per-sample list of predicted objects (detections, spans, instances) scored by average precision, mAP, or another metric bounded above by recall, and the program controls how many predictions to emit.
• Example:
• Input:
    predictions = []
    for sample_id in test_ids:
        # Emit nothing to avoid false positives, which are penalized
        pred_string = ""
        predictions.append({"Id": sample_id, "PredictionString": pred_string})

    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Consequence:
    Average-precision score is exactly 0 because recall is 0 for every sample;
    even naive prior-based guesses (fixed-size boxes at expected locations)
    score strictly higher, so the submission underperforms any trivial baseline.
• Counter-example:
• Input:
    predictions = []
    for sample_id in test_ids:
        boxes = model_boxes.get(sample_id, [])
        if not boxes:
            pred_string = ""  # only when the model truly found nothing
        else:
            pred_string = " ".join(fmt(b) for b in boxes)
        predictions.append({"Id": sample_id, "PredictionString": pred_string})
    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Why it does not fire: A branch populates the prediction field with non-empty, data-derived content, so empty strings occur only for samples where the model genuinely produced no candidates, not universally by design.
C3 · metric-mismatch Uniform blend includes uncalibrated bagged-tree probabilities under a proper scoring rulegithub_occurrence

P1 — Uniform blend includes uncalibrated bagged-tree probabilities under a proper scoring rule

• Pattern: Detects a probability ensemble scored by log loss (or another proper scoring rule) that averages the class-probability output of a bagged tree ensemble fitted on high-dimensional sparse text features together with other models using equal, untuned weights and no calibration step for the tree member.
• Detection procedure:
1. Identify a model assigned from a bagging/random-forest classifier constructor (e.g. RandomForestClassifier, BaggingClassifier, ExtraTreesClassifier, or equivalent) that is fitted on a feature matrix produced by a sparse text vectorizer (e.g. TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent) within the same script.
2. Identify a call to that model's probability-output method (predict_proba or equivalent) whose result is combined with the probability outputs of at least one other model via an arithmetic expression with equal coefficients — e.g. (p1 + p2 + p3) / 3, np.mean([...], axis=0), or a sum multiplied by a constant equal to the reciprocal of the member count.
3. Confirm the tree model is never wrapped in or passed to a probability-calibration construct (e.g. CalibratedClassifierCV or equivalent) and that no weight in the blend expression is assigned from a held-out evaluation loop or search over weights within the script.
4. Confirm the blended probabilities are written as the final prediction output and the task's stated metric is log loss or another proper scoring rule (e.g. a log_loss computation, or a comment/metric reference to logarithmic loss).
5. PRESENT when steps 1–4 all hold: an uncalibrated bagged-tree member on sparse high-dimensional features contributes with equal, untuned weight to a probability blend evaluated by a proper scoring rule.
• Predicted impact:
• Add score for C3: Optimizing or reporting the wrong quantity: 2
• weight: 2
• confidence: high
• Evidence: ensemble_pred = (pred1 + pred2 + pred3) / 3 blending a random-forest member's vote-fraction probabilities on TF-IDF features with two calibrated linear/probabilistic models; the equally weighted blend scored ~0.40 worse on the log-loss metric than a tuned 0.7/0.3 blend of the two well-calibrated members alone.
• Applies when: the task is evaluated by log loss or another proper scoring rule on predicted class probabilities, the script builds a heterogeneous ensemble of classifiers, and features come from a sparse high-dimensional text vectorizer.
• Example:
• Input:
    vec = TfidfVectorizer(max_features=50000)
    X_tr = vec.fit_transform(train_text); X_te = vec.transform(test_text)
    clf1 = RandomForestClassifier(n_estimators=100)
    clf2 = MultinomialNB(alpha=0.1)
    clf3 = LogisticRegression(C=1.0)
    for c in (clf1, clf2, clf3):
        c.fit(X_tr, y)
    p1 = clf1.predict_proba(X_te)
    p2 = clf2.predict_proba(X_te)
    p3 = clf3.predict_proba(X_te)
    ensemble_pred = (p1 + p2 + p3) / 3
    pd.DataFrame(ensemble_pred, columns=classes).to_csv("submission.csv")
• Consequence:
    Log loss on the held-out evaluation increases (worsens) by roughly 0.4
    relative to a validation-weighted blend of the calibrated members only,
    because the forest's vote-fraction probabilities are miscalibrated on
    sparse text features and pull every prediction toward those values with
    full equal weight.
• Counter-example:
• Input:
    vec = TfidfVectorizer(max_features=50000)
    X_tr = vec.fit_transform(train_text); X_te = vec.transform(test_text)
    rf = CalibratedClassifierCV(RandomForestClassifier(n_estimators=100), cv=3)
    lr = LogisticRegression(C=1.0)
    rf.fit(X_tr, y); lr.fit(X_tr, y)
    best_w = min(weight_grid, key=lambda w: log_loss(
        y_val, w * rf.predict_proba(X_val) + (1 - w) * lr.predict_proba(X_val)))
    final = best_w * rf.predict_proba(X_te) + (1 - best_w) * lr.predict_proba(X_te)
    pd.DataFrame(final, columns=classes).to_csv("submission.csv")
• Why it does not fire: the tree member is wrapped in a calibration construct and the blend weight is chosen by held-out log-loss search rather than fixed equal coefficients, so steps 2 and 3 fail.
C3 · metric-mismatch Constant uncertainty under a calibration-scored metric

P1 — Constant uncertainty under a calibration-scored metric

• Pattern: Detects a submission whose per-prediction uncertainty column is filled with a single scalar constant for every row, even though the evaluation metric scores calibrated uncertainty and a difficulty covariate (such as forecast horizon or distance from a known baseline) is available per row.
• Detection procedure:
1. Identify the code that builds the output frame written to the submission file (e.g., a DataFrame later passed to .to_csv, or equivalent), and locate the column representing uncertainty/confidence/standard deviation.
2. Trace the value assigned to that column: check whether it is a scalar variable, a numeric literal, or the result of an optimization that returns one scalar (e.g., a single minimize result), broadcast to all rows via assignment like df["conf_col"] = sigma.
3. Confirm that no per-row quantity (a column of the prediction frame, an array of the same length as the predictions, or a function of a horizon/gap/distance column) enters the expression assigned to the uncertainty column anywhere before the file is written.
4. Confirm the prediction rows carry a covariate that varies with expected difficulty (e.g., a time-offset, horizon, or gap-from-baseline column exists in the prediction frame).
5. PRESENT when the uncertainty column is a broadcast scalar (steps 2–3) while a per-row difficulty covariate exists (step 4).
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: sub_df["Confidence"] = best_sigma broadcast one optimized scalar to every prediction row; a contrastive run assigning conf = a + b * np.abs(dw) scored 0.029 better on the same likelihood-based metric, because the constant over-states uncertainty on near-baseline rows and under-states it on far-horizon rows.
• Applies when: The evaluation metric is likelihood- or calibration-based (scores both a point prediction and a per-row uncertainty), and the prediction task involves forecasting or extrapolation where a per-row difficulty covariate is available.
• Example:
• Input:
    from scipy.optimize import minimize

    res = minimize(neg_metric, x0=[200.0], args=(y_val, pred_val))
    best_sigma = res.x[0]

    sub_df["value"] = model.predict(X_sub)
    sub_df["conf"] = best_sigma          # same constant for every row
    sub_df["conf"] = sub_df["conf"].round(1)
    sub_df[["row_id", "value", "conf"]].to_csv("submission.csv", index=False)
• Consequence:
    Average likelihood-based score drops (~0.03 worse than a horizon-scaled
    confidence on the same predictions): easy near-baseline rows are penalized
    by inflated sigma, and far-horizon rows are penalized by clipped errors
    exceeding the understated sigma.
• Counter-example:
• Input:
    dw = sub_df["target_week"] - sub_df["base_week"]
    conf = a + b * np.abs(dw.values.astype(float))
    conf = np.maximum(conf, 70.0)        # scalar floor, not a constant value
    sub_df["conf"] = conf
    sub_df[["row_id", "value", "conf"]].to_csv("submission.csv", index=False)
• Why it does not fire: the uncertainty column is a per-row function of the horizon covariate; the scalar 70.0 only appears as a lower clip, not as the value broadcast to all rows.
C3 · metric-mismatch Hardcoded uncertainty channel under a likelihood-based metricgithub_occurrence

P1 — Hardcoded uncertainty channel under a likelihood-based metric

• Pattern: Detects a regression submission scored by a metric that jointly evaluates the point prediction and a predicted uncertainty value, where the uncertainty column is filled from fixed numeric constants or a global residual statistic with hardcoded coefficients, with no search over uncertainty parameters evaluated against the scoring formula on held-out or out-of-fold errors.
• Detection procedure:
1. Identify the frame written to the output file (e.g. via to_csv or equivalent) and locate a column whose name denotes uncertainty or confidence (e.g. contains the substring Confidence, sigma, std, uncertainty, case-insensitive).
2. Trace every assignment to that column within the script. Mark it HEURISTIC if each assignment is (a) a numeric literal, (b) a literal-coefficient expression such as a + b * <term> where a and b are numeric literals or names assigned only numeric literals, or (c) a global standard deviation of residuals possibly scaled/offset by numeric literals — in all cases with no dependence on an evaluated score.
3. Search the whole script for any construct that evaluates candidate uncertainty values against the scoring formula on prediction errors: a loop or grid over candidate values (e.g. for s in ..., np.linspace, itertools.product, or equivalent) whose body computes a likelihood/score from residuals, or a call to a numeric optimizer (scipy.optimize.minimize or equivalent) whose objective involves both residuals and the candidate uncertainty.
4. PRESENT if step 2 marks the uncertainty column HEURISTIC and step 3 finds no such evaluation-driven selection anywhere in the script.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: A script assigned the uncertainty column via conf = a + b * np.abs(dw) with a, b as hardcoded constants and a literal fallback Confidence = 150.0, never evaluating any candidate sigma against the likelihood-based scoring formula; a comparable pipeline that tuned sigma on held-out errors scored 0.056 better on the same metric (≈5.6% relative).
• Applies when: The task is a regression whose score is a likelihood-style function of both the predicted value and a predicted uncertainty (e.g. Laplace or Gaussian log-likelihood with a per-row sigma), and the script produces that uncertainty column.
• Example:
• Input:
    delta = model.predict(Xte)
    fvc = pred_df["base_val"].values + delta
    a, b = 200.0, 3.0
    conf = a + b * np.abs(pred_df["dw"].values.astype(float))
    conf = np.maximum(conf, 70.0)
    out = pd.DataFrame({
        "row_id": pred_df["row_id"],
        "value": fvc,
        "Confidence": conf,
    })
    out.to_csv("submission.csv", index=False)
• Consequence:
    The predicted sigma is systematically miscalibrated relative to the
    likelihood metric; final leaderboard score is several percent worse
    (observed 5.6% relative) than the same point predictions paired with
    a sigma tuned on out-of-fold errors.
• Counter-example:
• Input:
    oof_err = np.abs(y_oof - oof_pred)
    dw = np.abs(oof_dw)
    best, best_ll = None, -np.inf
    for a in np.linspace(50, 400, 36):
        for b in np.linspace(0, 10, 21):
            s = np.maximum(a + b * dw, 70.0)
            ll = np.mean(-np.sqrt(2) * oof_err / s - np.log(np.sqrt(2) * s))
            if ll > best_ll:
                best_ll, best = ll, (a, b)
    a, b = best
    conf = np.maximum(a + b * np.abs(test_dw), 70.0)
• Why it does not fire: the uncertainty coefficients are selected by a grid search that evaluates the scoring formula on out-of-fold errors, so step 3 finds an evaluation-driven selection and the conjunction in step 4 fails.
C3 · metric-mismatch Universally empty prediction field in a ranked-detection submission

P1 — Universally empty prediction field in a ranked-detection submission

• Pattern: Detects a submission-writing routine for a detection or retrieval task that assigns an empty string to the prediction field on every control-flow path, so every instance receives zero predicted objects.
• Detection procedure:
1. Locate a loop (or vectorized assignment) that iterates over the identifiers of the evaluation set (e.g., rows of a sample-submission file or a list of test tokens) and builds a per-instance prediction value that is subsequently written to an output file via to_csv, json.dump, or equivalent.
2. Within that loop or assignment, enumerate every branch that sets the per-instance prediction value (e.g., a variable later placed in a PredictionString-style column, or appended to the list that becomes that column).
3. Check whether each such branch assigns the empty string literal "" (or an empty list/None later serialized as empty), with no branch constructing a non-empty prediction from model output, heuristics, or priors.
4. Confirm the output file is the artifact scored by a precision/recall-family metric (average precision, mean AP, or F-score), as indicated by the task setup, the output column naming, or comments referencing false-positive penalties.
5. PRESENT if all per-instance prediction assignments reachable in the loop are empty (steps 2–3) and that output is the scored artifact (step 4).
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: Every branch of the submission loop executed pred_string = "", justified by a comment that empty output "avoids false positives which heavily penalize the score"; recall was zero on every instance, so the AP-style metric evaluated to exactly 0.0, while an alternative program that emitted a small prior-derived candidate set with confidence values scored strictly higher (0.0 → 0.00057).
• Applies when: The program produces a per-instance list of predicted objects (boxes, spans, retrieved items) scored by average precision, mean AP, or an F-based metric that ranks or matches predictions against ground truth.
• Example:
• Input:
    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        if sample_id in test_index:
            # Empty output avoids false positives, which are penalized
            pred_string = ""
        else:
            pred_string = ""
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Consequence:
    Recall is 0 for every instance, so the average-precision-style score is
    exactly 0.0; any submission containing even one low-confidence prior-based
    candidate per instance scores strictly higher (e.g., 0.0 -> 0.00057),
    because AP ranks predictions by confidence and cannot fall below empty.
• Counter-example:
• Input:
    strings = []
    for tok in sub["Id"].values:
        pose = test_meta.get(tok)
        if pose is None or not cands:
            strings.append("")
            continue
        out = ["%.3f %.3f %.3f %.3f" % (c["conf"], c["x"], c["y"], c["z"])
               for c in cands]
        strings.append(" ".join(out))
    sub["PredictionString"] = strings
    sub.to_csv(OUT, index=False)
• Why it does not fire: The empty string is only a fallback for instances lacking metadata or candidates; the main path constructs non-empty, confidence-bearing predictions, so not all reachable assignments are empty.
C3 · metric-mismatch Single fixed-confidence detection per sample under an average-precision metricgithub_occurrence

P1 — Single fixed-confidence detection per sample under an average-precision metric

• Pattern: Detects prediction-writing code that emits at most one candidate object per sample, with the confidence field set to a hardcoded literal constant, for a task whose score is average precision over a variable number of ground-truth objects per sample.
• Detection procedure:
1. Locate the loop or comprehension that builds the per-sample prediction output (e.g. a string or row appended per sample token before writing a submission file with to_csv or equivalent).
2. Within that loop body, count how many candidate detections are added per sample: PRESENT-eligible if exactly one formatted detection (or an empty string) is produced per iteration, with no inner loop over multiple candidates for the same sample.
3. Inspect the confidence value in the formatted detection: PRESENT-eligible if it is a numeric literal (e.g. 1.0, 0.5) or a variable assigned once from a numeric literal, with no dependence on per-sample or per-candidate data.
4. Confirm the surrounding task is scored by average precision (mAP or equivalent) over samples that may contain zero, one, or many objects — indicated by a submission format that accepts a variable-length list of confidence x y z ...-style detections per row.
5. The pattern is PRESENT when steps 2, 3, and 4 all hold: one-or-zero detections per sample, constant confidence, AP-style scoring over variable object counts.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A pipeline emitting one detection per sample with a constant confidence, e.g. pred_string = f"1.0 {x} {y} {z} ..." appended once per sample, scored 1.0 worse on the competition metric than a variant that emitted multiple clustered candidates per sample with confidences derived from cluster support (conf = cd["n"] / maxn).
• Applies when: The program produces a detection-style submission (multiple objects allowed per row) for a task scored by average precision or mean average precision, and the true number of objects per sample varies.
• Example:
• Input:
    predictions = []
    for sample in test_samples:
        x, y, z = estimate_center(sample)
        w, l, h, yaw = 1.9, 4.5, 1.7, 0.0
        pred_string = f"1.0 {x:.3f} {y:.3f} {z:.3f} {w} {l} {h} {yaw} OBJ_A"
        predictions.append({'Id': sample['token'],
                            'PredictionString': pred_string})
    pd.DataFrame(predictions).to_csv('submission.csv', index=False)
• Consequence:
    Recall is capped at one object per sample while scenes contain many;
    the constant confidence provides no ranking signal for the AP sweep,
    so mAP collapses toward zero relative to multi-candidate outputs
    (observed metric gap of 1.0 versus a graded multi-candidate baseline).
• Counter-example:
• Input:
    strings = []
    for sample in test_samples:
        out = []
        for cd in candidates_for(sample):
            conf = max(0.01, min(1.0, cd['support'] / max_support))
            out.append(f"{conf:.4f} {cd['x']:.3f} {cd['y']:.3f} {cd['z']:.3f} "
                       f"{cd['w']} {cd['l']} {cd['h']} {cd['yaw']} OBJ_A")
        strings.append(" ".join(out))
    sub['PredictionString'] = strings
    sub.to_csv('submission.csv', index=False)
• Why it does not fire: the inner loop emits multiple candidates per sample and the confidence is computed from per-candidate data, so neither the single-detection condition nor the constant-confidence condition holds.
C3 · metric-mismatch Unvalidated equal-weight blend including tree-ensemble probabilities on sparse text under a probabilistic metric

P1 — Unvalidated equal-weight blend including tree-ensemble probabilities on sparse text under a probabilistic metric

• Pattern: Detects final class probabilities produced by averaging the probability outputs of several classifiers — at least one a bagged/randomized tree ensemble fitted on high-dimensional sparse text features — with equal or hardcoded weights and no held-out evaluation of a probabilistic loss to select members or weights.
• Detection procedure:
1. Identify a text featurization step producing a sparse high-dimensional matrix, e.g. a call to TfidfVectorizer or CountVectorizer (or equivalent bag-of-words/n-gram featurizer), whose fitted output is assigned to a variable later passed to model-fitting calls.
2. Within the same script, find two or more classifier objects each fitted (via .fit or equivalent) on that sparse matrix, where at least one is a bagged or randomized tree ensemble (e.g. RandomForestClassifier, ExtraTreesClassifier, or equivalent).
3. Find variables assigned from each classifier's class-probability method (predict_proba or equivalent) applied to the test/inference matrix.
4. Find an expression combining those probability variables using only literal numeric constants — e.g. summing them and dividing by a literal integer, or multiplying each by a literal float — with the result written into the final output (submission frame, saved file, or returned predictions).
5. Confirm the script contains no evaluation of a probabilistic loss (e.g. log_loss or equivalent, or cross-validated scoring with a log-loss/probabilistic scorer) whose result is computed before the combining expression and whose value could influence which members or weights appear in step 4.
6. PRESENT when steps 1–4 all hold and step 5 confirms the absence of any such evaluation.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: ensemble_pred = (pred1 + pred2 + pred3) / 3 blending a tree ensemble's predict_proba output on sparse text features with two other classifiers, with no held-out metric check; the measured contrastive pair showed the validated linear blend scoring 0.42 better on the probabilistic competition metric.
• Applies when: the task is scored by a probabilistic loss (log loss, Brier score, or similar), features are sparse high-dimensional text representations, and multiple classifiers' probability outputs are combined into the final prediction.
• Example:
• Input:
    vec = TfidfVectorizer()
    X_tr = vec.fit_transform(train['col_a']); X_te = vec.transform(test['col_a'])
    clf1 = LogisticRegression().fit(X_tr, y)
    clf2 = MultinomialNB().fit(X_tr, y)
    clf3 = RandomForestClassifier(n_estimators=100).fit(X_tr, y)
    p1 = clf1.predict_proba(X_te)
    p2 = clf2.predict_proba(X_te)
    p3 = clf3.predict_proba(X_te)
    final = (p1 + p2 + p3) / 3
    pd.DataFrame(final).to_csv('submission.csv', index=False)
• Consequence:
    Log loss on the scored set rises substantially (observed +0.42 versus a
    validated blend of well-calibrated models on the same features), because
    the tree ensemble's overconfident/flat probabilities on sparse text pull
    the averaged estimates away from calibrated values.
• Counter-example:
• Input:
    vec = TfidfVectorizer()
    X_tr = vec.fit_transform(train['col_a']); X_te = vec.transform(test['col_a'])
    oof_lr, oof_nb = np.zeros((len(y), 3)), np.zeros((len(y), 3))
    for tr_idx, va_idx in StratifiedKFold(5).split(X_tr, y):
        lr = LogisticRegression().fit(X_tr[tr_idx], y[tr_idx])
        nb = MultinomialNB().fit(X_tr[tr_idx], y[tr_idx])
        oof_lr[va_idx] = lr.predict_proba(X_tr[va_idx])
        oof_nb[va_idx] = nb.predict_proba(X_tr[va_idx])
    best_w = min(np.arange(0, 1.05, 0.05),
                 key=lambda w: log_loss(y, w*oof_lr + (1-w)*oof_nb))
    final = best_w * lr.predict_proba(X_te) + (1-best_w) * nb.predict_proba(X_te)
• Why it does not fire: the blend weight is chosen by minimizing log_loss on out-of-fold predictions (failing step 5), and no tree ensemble is included in the combination (failing step 2).
C3 · metric-mismatch Constant uncertainty ignoring forecast horizonverified_trace · effect +0.0603

P1 — Constant uncertainty ignoring forecast horizon

• Pattern: Detects a submission whose predicted-uncertainty column is assigned a single scalar constant for all rows, even though the prediction rows span multiple forecast horizons relative to a per-entity baseline time.
• Detection procedure:
1. Identify the output frame written to the submission file (e.g. via to_csv or equivalent) and locate the column holding the predicted uncertainty (a column separate from the point prediction, e.g. named like Confidence, sigma, or uncertainty).
2. Trace the assignment to that column within the same script: check whether the right-hand side is a scalar variable, a numeric literal, or an expression that does not reference any per-row time/horizon quantity (no column derived from a difference between a prediction time column and a baseline time column).
3. Confirm that elsewhere in the script a per-row time-gap quantity exists or is derivable (a prediction time column and a per-entity baseline time are both present, e.g. a computed difference like df["t"] - df["t_base"] or the raw columns to compute it), so horizons genuinely vary across rows.
4. PRESENT if the uncertainty column receives a horizon-independent scalar or constant expression (step 2) while the rows span varying forecast horizons (step 3).
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: sub_merged["Confidence"] = best_sigma assigned one global scalar to every row despite the script computing Weeks_since_base per row; a variant setting sigma as c0 + c1 * abs(dw) scored 0.02456 better on the same likelihood-based metric.
• Applies when: The task is a forecasting problem scored by a per-row likelihood or penalty that divides the prediction error by a model-supplied uncertainty, and prediction rows extend across multiple time offsets from an observed baseline.
• Example:
• Input:
    sub = sample_sub.merge(base, on="entity", how="left")
    sub["dt"] = sub["t"] - sub["t_base"]
    sub["value"] = predict_value(sub)
    best_sigma = 230.0
    sub["Confidence"] = best_sigma
    sub[["row_id", "value", "Confidence"]].to_csv("submission.csv", index=False)
• Consequence:
    Rows far from the baseline time have larger true errors but the same sigma,
    so their error/sigma penalty inflates; the averaged likelihood score drops
    (observed gap ~0.025 on the metric versus a horizon-scaled sigma).
• Counter-example:
• Input:
    sub = sample_sub.merge(base, on="entity", how="left")
    sub["dt"] = sub["t"] - sub["t_base"]
    sub["value"] = predict_value(sub)
    c0, c1 = 180.0, 4.5
    sub["Confidence"] = c0 + c1 * sub["dt"].abs()
    sub[["row_id", "value", "Confidence"]].to_csv("submission.csv", index=False)
• Why it does not fire: The uncertainty column is an expression of the per-row time gap dt, so it scales with the forecast horizon rather than being a global constant.
C3 · metric-mismatch Default regularization for probabilistic linear classifier on high-dimensional weak features

P1 — Default regularization for probabilistic linear classifier on high-dimensional weak features

• Pattern: Detects a linear probabilistic classifier fitted on high-dimensional hand-crafted features with the library-default regularization strength, whose probability outputs are written to the scored artifact without any held-out evaluation of the probabilistic metric to select that strength.
• Detection procedure:
1. Find a constructor call for a regularized linear classifier that produces class probabilities (e.g. LogisticRegression(...) or equivalent) where the regularization argument (C, alpha, or equivalent) is absent from the argument list, so the library default applies.
2. Confirm the fitted object's probability method (e.g. predict_proba(...) or equivalent) is called on a test/inference matrix, and its output flows (possibly via a dataframe) into a file-writing call such as .to_csv(...).
3. Confirm the feature matrix passed to .fit(...) is built from an extraction routine producing a fixed-length vector per sample whose declared or computed dimensionality is in the thousands or more (e.g. a feature-length constant or an array allocated with a large second dimension), i.e. features are engineered descriptors rather than a low-dimensional table.
4. Search the whole file for any of: a hyperparameter search construct (GridSearchCV, RandomizedSearchCV, or equivalent), a loop over candidate regularization values, or a call computing a probabilistic loss (log_loss or equivalent) on a held-out split.
5. PRESENT if step 1 finds a default-regularization fit AND steps 2–3 hold AND step 4 finds none of the listed constructs.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: LogisticRegression() fitted at default strength on thousands of gradient-histogram features and its predict_proba output submitted directly; a contrastive run identical except for strong regularization (C=0.01) plus a uniform-prior blend improved the probabilistic competition metric by 0.609, indicating the default fit produced overconfident probabilities that inflated log loss on misclassified rows.
• Applies when: The task is scored by a probabilistic metric such as multi-class log loss, the model is a maximum-likelihood linear classifier, and the feature dimensionality is large relative to the number of samples per class.
• Example:
• Input:
    from sklearn.linear_model import LogisticRegression
    from sklearn.preprocessing import StandardScaler

    # X: (n_samples, 8100) descriptor features, y: many-class labels
    scaler = StandardScaler()
    Xs = scaler.fit_transform(X)
    clf = LogisticRegression(max_iter=1000)
    clf.fit(Xs, y)
    probs = clf.predict_proba(scaler.transform(X_test))
    pd.DataFrame(probs, columns=classes).to_csv("submission.csv", index=False)
• Consequence:
    Under-regularized weights drive predicted probabilities toward 0/1;
    every misclassified row incurs a near-maximal per-row log loss, and the
    submitted log loss is more than 2x worse than a strongly regularized fit
    of the same features (observed gap ~0.61 on the probabilistic metric).
• Counter-example:
• Input:
    from sklearn.linear_model import LogisticRegression
    from sklearn.metrics import log_loss

    best_C, best_ll = None, float("inf")
    for C in [0.001, 0.01, 0.1, 1.0]:
        m = LogisticRegression(C=C, max_iter=1000).fit(X_tr, y_tr)
        ll = log_loss(y_va, m.predict_proba(X_va))
        if ll < best_ll:
            best_C, best_ll = C, ll
    clf = LogisticRegression(C=best_C, max_iter=1000).fit(X, y)
    probs = clf.predict_proba(X_test)
• Why it does not fire: the regularization strength is chosen by evaluating the probabilistic loss on a held-out split, so step 4 finds both a loop over candidate values and a log_loss evaluation.
C3 · metric-mismatch Constant confidence for all emitted detections under a ranking-sensitive metric

P1 — Constant confidence for all emitted detections under a ranking-sensitive metric

• Pattern: Detects a detection-output writer that assigns one fixed confidence value to every emitted prediction instead of a per-candidate score, when the evaluation metric averages precision over confidence-ordered predictions.
• Detection procedure:
1. Locate the code that builds per-sample prediction strings or rows for a detection submission, i.e., a loop that formats multiple candidate boxes/objects per sample and writes them to an output file (via to_csv, json.dump, write, or equivalent).
2. Identify the value emitted in the confidence position of each prediction (the field documented or formatted first/labeled as confidence or score).
3. Check whether that value is a literal constant (e.g., 1.0) or a variable assigned once outside all per-candidate loops and never reassigned per candidate, so every emitted prediction receives the identical confidence.
4. Check that no per-candidate quantity (cluster size, class frequency, model output, distance, count) contributes to the confidence value at any point before emission.
5. PRESENT if steps 3 and 4 both hold: the output contains multiple predictions per file, all carrying one identical constant confidence, with no per-candidate scoring anywhere in the pipeline.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 3
• weight: 3
• confidence: high
• Evidence: A submission writer set confidence = 1.0 once and emitted it for every candidate box, while a variant computing conf = max(0.01, min(1.0, cd["n"] / maxn)) from per-candidate support scored strictly higher on the same average-precision-style metric; the measured metric gap between the two runs was 1.0 in favor of the ranked-confidence version.
• Applies when: The task is a detection or multi-candidate prediction problem whose evaluation metric is average precision (or any metric computed over predictions sorted by confidence), and the program emits more than one candidate per sample or per file.
• Example:
• Input:
    confidence = 1.0
    rows = []
    for sample_id in sub_df["Id"]:
        parts = []
        for dx, dy in offsets:
            wx, wy = transform(dx, dy)
            parts.append(f"{confidence} {wx:.3f} {wy:.3f} {wz:.3f} "
                         f"{w:.3f} {l:.3f} {h:.3f} {yaw:.4f} LABEL_A")
        rows.append(" ".join(parts))
    sub_df["PredictionString"] = rows
    sub_df.to_csv("submission.csv", index=False)
• Consequence:
    Average precision drops relative to score-ordered output: because every
    prediction shares the same confidence, low-quality candidates are ranked
    tied with the best ones, so false positives appear early in the
    precision-recall sweep and depress precision at every recall level.
• Counter-example:
• Input:
    rows = []
    for sample_id in sub_df["Id"]:
        parts = []
        for cand in candidates:
            conf = max(0.01, min(1.0, cand["support"] / max_support))
            wx, wy = transform(cand["dx"], cand["dy"])
            parts.append(f"{conf} {wx:.3f} {wy:.3f} {wz:.3f} "
                          f"{cand['w']:.3f} {cand['l']:.3f} {cand['h']:.3f} {yaw:.4f} LABEL_A")
        rows.append(" ".join(parts))
    sub_df["PredictionString"] = rows
    sub_df.to_csv("submission.csv", index=False)
• Why it does not fire: The confidence is reassigned inside the per-candidate loop from a per-candidate quantity (cand["support"]), so predictions carry distinct plausibility-based scores and are ordered meaningfully.
C3 · metric-mismatch Hard-coded uncertainty term in a likelihood-scored submission

P1 — Hard-coded uncertainty term in a likelihood-scored submission

• Pattern: Detects a submitted uncertainty/confidence column built purely from hand-picked numeric constants (fixed literal, or literal base plus literal slope) with no search of those parameter values against the scored metric on held-out or out-of-fold predictions.
• Detection procedure:
1. Locate a statement writing a tabular output file (e.g. a call to to_csv or equivalent) whose frame contains both a point-prediction column and a second numeric column representing per-row uncertainty (names such as Confidence, sigma, std, or a column documented as the uncertainty term of the scoring metric).
2. Trace every assignment producing that uncertainty column within the same script; check whether each value expression is composed only of numeric literals, or numeric literals combined arithmetically with a feature (e.g. a + b * abs(dw) where a and b are assigned literal constants), rather than being an output of a fitted model (e.g. quantile spread, residual-based estimate) or of an optimization.
3. Search the whole script for any loop, grid, or optimizer call that iterates over candidate values of those constants and, inside the iteration, evaluates a function computing the task's likelihood/score on predictions from a validation split or out-of-fold predictions.
4. PRESENT if step 2 finds only literal-constant construction AND step 3 finds no such search anywhere before the output file is written.
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: Uncertainty column filled via submission["Confidence"].fillna(150.0) and conf = a + b * np.abs(dw) with hand-picked a, b; the equivalent pipeline that grid-searched these two parameters against the exact scored metric on out-of-fold predictions scored ~5.6% better on the official likelihood metric.
• Applies when: The task's score is an uncertainty-aware likelihood or interval metric in which a submitted per-row confidence/standard-deviation value directly enters the score, and the program produces such a column (regression/tabular setting).
• Example:
• Input:
    dw = np.abs(pred_df["weeks"] - pred_df["base_week"])
    conf = 150.0 + 3.0 * dw          # hand-picked base and slope
    conf = np.maximum(conf, 70.0)
    out = pd.DataFrame({
        "row_id": pred_df["row_id"],
        "value": preds,
        "Confidence": conf,
    })
    out.to_csv("submission.csv", index=False)
• Consequence:
    The sigma dimension of the likelihood metric is uncalibrated: on the
    contrastive pair, the hard-coded-confidence run scored ~5.6% worse on
    the official metric than the same pipeline with the base/slope
    grid-searched against the metric on out-of-fold predictions.
• Counter-example:
• Input:
    best = (None, -1e18)
    for a in range(60, 300, 10):
        for b in np.arange(0.0, 8.0, 0.5):
            sigma = np.maximum(a + b * oof_dw, 70.0)
            score = metric(oof_true, oof_pred, sigma)
            if score > best[1]:
                best = ((a, b), score)
    a, b = best[0]
    conf = np.maximum(a + b * dw, 70.0)
    out = pd.DataFrame({"row_id": ids, "value": preds, "Confidence": conf})
    out.to_csv("submission.csv", index=False)
• Why it does not fire: the same base-plus-slope parameterization appears, but the constants are chosen by an explicit grid search that evaluates the scored metric on out-of-fold predictions, so step 3 finds a search and step 4's conjunction fails.
C3 · metric-mismatch Constant per-row uncertainty despite horizon covariateverified_trace · effect +0.0603

P1 — Constant per-row uncertainty despite horizon covariate

• Pattern: Detects a submission whose per-row uncertainty/confidence column is filled with a single scalar constant for all rows, while the same output frame contains a horizon-like covariate (e.g., elapsed time or distance from a baseline observation) that is never used to vary the uncertainty.
• Detection procedure:
1. Locate an assignment to an output-frame column whose name contains Confidence, sigma, std, uncertainty, or an equivalent per-row uncertainty field required by the scoring metric.
2. Check that the right-hand side of that assignment is a scalar variable or numeric literal (broadcast to all rows), not an expression indexed by or computed from a per-row column of the frame.
3. Within the same script, check that a per-row covariate representing prediction horizon or offset exists — a column computed as the difference between a time/week/step column and a baseline value (e.g., df["dw"] = df["Weeks"] - df["base_week"] or equivalent), or the raw time column itself.
4. Verify that no other assignment to the uncertainty column anywhere in the script references that covariate (or any per-row column) in its right-hand side.
5. PRESENT when the uncertainty column is set only from a scalar (steps 1–2) and a horizon covariate is available but unused in any uncertainty expression (steps 3–4).
• Predicted impact:
• Add score for C3 (Optimizing or reporting the wrong quantity): 2
• weight: 2
• confidence: high
• Evidence: sub_df["Confidence"] = best_sigma assigns one optimized scalar to every row even though Weeks_since_base was computed in the same frame; the horizon-scaled variant conf = a + b * np.abs(pred_df["dw"]) improved the averaged likelihood metric by ~0.028–0.029 in two measured contrastive pairs.
• Applies when: The evaluation metric scores each prediction jointly with a per-row uncertainty value, and the prediction task spans varying horizons (rows at different distances from a baseline observation) so that expected error magnitude differs across rows.
• Example:
• Input:
    sub = sample_sub.merge(base_small, on="id_a", how="left")
    sub["weeks_since_base"] = sub["week"] - sub["week_base"]
    sub["pred"] = model.predict(make_features(sub))
    best_sigma = fit_global_sigma(train_residuals)
    sub["Confidence"] = best_sigma
    sub[["row_id", "pred", "Confidence"]].to_csv("submission.csv", index=False)
• Consequence:
    Averaged likelihood metric drops (~0.03 worse) versus horizon-scaled
    uncertainty: far-horizon rows with large true error are penalized by an
    under-sized sigma, while near-horizon rows waste an over-sized sigma.
• Counter-example:
• Input:
    sub = sample_sub.merge(base_small, on="id_a", how="left")
    sub["weeks_since_base"] = sub["week"] - sub["week_base"]
    sub["pred"] = model.predict(make_features(sub))
    conf = a + b * np.abs(sub["weeks_since_base"].astype(float))
    sub["Confidence"] = np.maximum(conf, min_sigma)
    sub[["row_id", "pred", "Confidence"]].to_csv("submission.csv", index=False)
• Why it does not fire: the uncertainty column is computed from the per-row horizon covariate, so its right-hand side is not a broadcast scalar and step 4 fails.
C4 · transform-skew Test-set imputation statistics diverge from training imputation

P1 — Test-set imputation statistics diverge from training imputation

• Pattern: Detects missing-value imputation applied to the held-out/test feature frame using statistics computed from that same test frame, while the training frame is imputed with statistics computed from itself.
• Detection procedure:
1. Identify two distinct dataframe variables: one used as the fitting input to a model-training call (e.g. model.fit(X, y) or equivalent), call it the training frame; and one passed to an inference call (e.g. model.predict(...) or equivalent), call it the test frame.
2. Locate a fill or imputation operation on the test frame of the form X_test.fillna(expr) (or an in-place equivalent, or an imputer's fit_transform called directly on the test frame) where expr is derived from the test frame itself — e.g. X_test.mean(), X_test.median(), X_test.mode() — rather than from a variable computed from the training frame.
3. Verify that no variable holding statistics computed from the training frame (e.g. X.mean() assigned earlier, or an imputer object fitted on the training frame) is used in the test-frame fill expression at step 2.
4. The pattern is PRESENT when the test frame's missing values are filled with values derived from the test frame (steps 2–3) and that frame is subsequently passed to the inference call from step 1.
• Predicted impact:
• Add score for C4 (Train and inference paths differ): 2
• weight: 2
• confidence: high
• Evidence: X_test.fillna(X_test.mean()) used to prepare inference inputs while the training frame was filled with its own means; in a measured contrastive pair the run reusing training statistics scored 0.023 higher on the competition metric.
• Applies when: Tabular pipeline where features may contain missing values and imputation is performed before fitting and predicting with separate train and test frames.
• Example:
• Input:
    X = train_df.drop(columns=["label"])
    y = train_df["label"]
    X_test = test_df.copy()

    X = X.fillna(X.mean())
    X_test = X_test.fillna(X_test.mean())

    model.fit(X, y)
    preds = model.predict(X_test)
• Consequence:
    Columns with missing values are filled with different constants at train
    and inference time; test rows are shifted relative to the distribution the
    model learned, lowering held-out accuracy (observed -0.023 on the
    evaluation metric versus reusing training-set statistics).
• Counter-example:
• Input:
    X = train_df.drop(columns=["label"])
    y = train_df["label"]
    X_test = test_df.copy()

    fill_values = X.mean()
    X = X.fillna(fill_values)
    X_test = X_test.fillna(fill_values)

    model.fit(X, y)
    preds = model.predict(X_test)
• Why it does not fire: the test frame's fill expression is a variable computed from the training frame, so both partitions are imputed with identical statistics.
C4 · transform-skew Row-identifier column used as a model featuretrace_observed

P1 — Row-identifier column used as a model feature

• Pattern: Detects a tabular training pipeline that fits a tree-based model on a feature matrix which still contains the row-identifier/primary-key column, so splits can be learned on identifier values whose range differs between the training and test sets.
• Detection procedure:
1. Identify the code that loads a training table (e.g. pd.read_csv('train.csv') or equivalent) and locate the column used purely as a row identifier: it is later written as the first column of the output/submission frame (e.g. pd.DataFrame({'Id': test_ids, ...})) or extracted with a name like id, Id, ID, row_id, or set/reset as an index.
2. Trace how the feature matrix passed to the model's fitting call (fit(X, y), lgb.train, or equivalent) is constructed from that table: record every drop(...), column-list selection (e.g. df[feats]), pop(...), or set_index(...) applied between loading and fitting.
3. Check whether the identifier column found in step 1 appears in any of the exclusion operations recorded in step 2, either by literal name or via a list comprehension that filters it out (e.g. [c for c in df.columns if c not in (id_col, target_col)]).
4. The pattern is PRESENT when the fitted estimator is a tree-based model and the identifier column from step 1 is absent from every exclusion operation in step 2, so it remains a column of the matrix passed to the fitting call.
• Predicted impact:
• Add score for C4: Train and inference paths differ: 2
• weight: 2
• confidence: high
• Evidence: A pipeline that built its feature matrix with X = train_df.drop('Cover_Type', axis=1) while leaving the Id column in place, then fitted a random-forest classifier, scored about 2.4% lower on the held-out metric than an otherwise comparable pipeline that constructed feats = [c for c in train.columns if c not in ("Id", target)] before fitting; splits learned on training-range identifiers routed all out-of-range test rows down uninformative branches.
• Applies when: Tabular supervised learning where the input table carries a monotonically assigned row-identifier column, the model is tree-based, and the test rows' identifier values lie outside (or in a disjoint range from) the training rows' identifiers.
• Example:
• Input:
    import pandas as pd
    from sklearn.ensemble import RandomForestClassifier
    train = pd.read_csv('train.csv')
    test = pd.read_csv('test.csv')
    y = train['target']
    X = train.drop('target', axis=1)      # 'Id' column still in X
    model = RandomForestClassifier(n_estimators=100, random_state=42)
    model.fit(X, y)
    preds = model.predict(test.drop(columns=[]))  # 'Id' still in features
    pd.DataFrame({'Id': test['Id'], 'target': preds}).to_csv('sub.csv', index=False)
• Consequence:
    Trees split on the identifier column; test rows have identifier values above
    the training range, so they all fall into the same branch of those splits.
    Held-out accuracy drops (observed ~2.4% relative deficit versus an identical
    pipeline that excludes the identifier), with predictions systematically
    biased toward the classes dominant in the highest-Id training rows.
• Counter-example:
• Input:
    import pandas as pd
    from sklearn.ensemble import RandomForestClassifier
    train = pd.read_csv('train.csv')
    test = pd.read_csv('test.csv')
    feats = [c for c in train.columns if c not in ('Id', 'target')]
    y = train['target']
    model = RandomForestClassifier(n_estimators=100, random_state=42)
    model.fit(train[feats], y)
    preds = model.predict(test[feats])
    pd.DataFrame({'Id': test['Id'], 'target': preds}).to_csv('sub.csv', index=False)
• Why it does not fire: the feature list explicitly filters out the identifier column before fitting, so the matrix passed to fit no longer contains it even though Id still appears when writing the output file.
C4 · transform-skew Test-frame-derived imputation statisticsverified_trace · effect +0.0202

P1 — Test-frame-derived imputation statistics

• Pattern: Detects missing-value imputation on the inference/test feature frame using statistics computed from that same test frame instead of statistics derived from the training data, so imputed values at inference come from a different distribution than the one the model was fit on.
• Detection procedure:
1. Identify a variable holding the test/inference feature frame, i.e. a frame loaded from a test file (e.g. pd.read_csv('test.csv') or equivalent) or derived from it, that is later passed to a prediction call (model.predict(...) or equivalent) but never passed to a fitting call.
2. Find a call on that variable to fillna(...) (or equivalent missing-value replacement) whose argument is an aggregate computed from the same variable — e.g. X_test.fillna(X_test.mean()), X_test.fillna(X_test.median()), or an imputer object on which fit or fit_transform is called with the test frame as input.
3. Confirm that within the same script the argument to that replacement is NOT an aggregate or fitted imputer derived from the training frame (no variable assigned from train_frame.mean()/train_frame.median() or an imputer fitted on the training frame is reused here).
4. PRESENT when steps 1–3 all hold: the test frame's missing values are filled from test-frame statistics with no reuse of training-derived statistics.
• Predicted impact:
• Add score for C4 (Train and inference paths differ): 2
• weight: 2
• confidence: high
• Evidence: X_test = X_test.fillna(X_test.mean(numeric_only=True)) while the training frame was filled with its own separate means; in two measured contrastive pairs the variant with test-frame-derived imputation scored 0.023 and 0.022 lower on the competition metric than otherwise comparable pipelines.
• Applies when: A tabular pipeline imputes missing values and produces predictions on a separate test/inference frame.
• Example:
• Input:
    train_df = pd.read_csv('train.csv')
    test_df = pd.read_csv('test.csv')
    X = train_df.drop('LABEL_A', axis=1).fillna(train_df.mean(numeric_only=True))
    y = train_df['LABEL_A']
    X_test = test_df.fillna(test_df.mean(numeric_only=True))
    model = RandomForestClassifier(n_estimators=100).fit(X, y)
    preds = model.predict(X_test[X.columns])
• Consequence:
    Rows in the test frame containing missing values receive imputed inputs
    shifted relative to the training-time fill values; held-out accuracy drops
    (~0.02 on the evaluation metric in measured runs), concentrated on rows
    with missing entries.
• Counter-example:
• Input:
    train_df = pd.read_csv('train.csv')
    test_df = pd.read_csv('test.csv')
    fill_values = train_df.drop('LABEL_A', axis=1).mean(numeric_only=True)
    X = train_df.drop('LABEL_A', axis=1).fillna(fill_values)
    y = train_df['LABEL_A']
    X_test = test_df.fillna(fill_values)
    model = RandomForestClassifier(n_estimators=100).fit(X, y)
    preds = model.predict(X_test[X.columns])
• Why it does not fire: the test frame's fillna argument is an aggregate computed once from the training frame and reused, so train and inference imputation statistics are identical.
C4 · transform-skew Zero-block stand-in for missing feature source before scalinggithub_occurrence

P1 — Zero-block stand-in for missing feature source before scaling

• Pattern: Detects per-sample feature construction that substitutes an all-zero block for an absent source or modality and concatenates it with real features that are subsequently standardized, with no fitted imputation step distinguishing absence from measured zeros.
• Detection procedure:
1. Locate a function or loop that builds a per-sample feature vector by concatenating blocks derived from more than one source (e.g., separate files, channels, or sub-directories per sample), where at least one branch handles the case that a source is unavailable (an if on file existence, empty read, len(...) == 0, or a try/except around loading).
2. In that unavailable-source branch, check that the appended block is composed of zeros: a call to np.zeros(...) or equivalent, a list literal containing only 0 or 0.0 repeated, or a multiplication of [0]/[0.0] by an integer.
3. Verify that the matrix assembled from these vectors is later passed to a fitted scaling or standardization transform (StandardScaler, RobustScaler, a manual (X - X.mean()) / X.std(), or equivalent) either directly or inside a pipeline.
4. Verify that no imputation transform (SimpleImputer, KNNImputer, a masked-fill using column statistics, or equivalent) is applied between feature assembly and model fitting, and that the zero blocks are not accompanied by a separate missingness-indicator column.
5. PRESENT if steps 1–4 all hold: zero blocks stand in for absent sources, the features are scaled, and no imputer or missingness indicator separates absence from real zeros.
• Predicted impact:
• Add score for C4 (Train and inference paths differ): 2
• weight: 2
• confidence: high
• Evidence: A per-sample extractor appended np.zeros(n_feats) whenever a source was unreadable, then fed the matrix through a fitted standard scaler; the sibling version filled absent blocks via a median imputer fitted on training rows. On the same task the imputer-based version scored 0.079 higher on the evaluation metric, because zero vectors were standardized into extreme pseudo-measurements that distorted both fitting and test-time predictions.
• Applies when: Per-sample features are concatenated from multiple sources or modalities, some sources can be absent for individual samples, and a scaling or standardization step is applied before model fitting or prediction.
• Example:
• Input:
    def build_features(sample_dir):
        blocks = []
        for src in ["src_a", "src_b", "src_c"]:
            path = sample_dir / f"{src}.npy"
            if path.exists():
                blocks.append(summarize(np.load(path)))   # 8 stats
            else:
                blocks.append(np.zeros(8))                # missing source
        return np.concatenate(blocks)

    X = np.stack([build_features(d) for d in sample_dirs])
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
• Consequence:
    Samples missing a source get standardized values several standard
    deviations from the column mean (zeros masquerading as real extreme
    measurements); the model fits these artifacts and the held-out
    classification metric drops by roughly 0.08 relative to imputing
    missing blocks from training-set statistics.
• Counter-example:
• Input:
    def build_features(sample_dir):
        blocks = []
        for src in ["src_a", "src_b", "src_c"]:
            path = sample_dir / f"{src}.npy"
            if path.exists():
                blocks.append(summarize(np.load(path)))
            else:
                blocks.append(np.full(8, np.nan))         # sentinel
        return np.concatenate(blocks)

    imp = SimpleImputer(strategy="median")
    X_train = imp.fit_transform(np.stack([build_features(d) for d in train_dirs]))
    X_test = imp.transform(np.stack([build_features(d) for d in test_dirs]))
• Why it does not fire: the absent-source branch appends NaN sentinels rather than zeros, and a fitted imputer replaces them with training-set column statistics before any scaling, so missingness is never conflated with a genuine zero measurement.
C4 · transform-skew Test-frame imputation uses test-derived statistics

P1 — Test-frame imputation uses test-derived statistics

• Pattern: Detects missing-value imputation where the test/inference frame is filled using statistics computed from the test frame itself while the training frame is filled with training-derived statistics, so the two paths use different fill values.
• Detection procedure:
1. Locate a statement in which a frame used for training features is filled by a missing-value replacement call (e.g. fillna, or an imputer's fit_transform, or equivalent) whose fill values are aggregate statistics (mean(), median(), mode(), or an imputer fitted on that same training frame).
2. Locate a separate statement in which the test/inference frame — a frame later passed to a model's prediction call rather than its fitting call — is filled by a missing-value replacement call whose fill values are aggregate statistics computed on that test frame itself (e.g. test_df.fillna(test_df.mean()), or an imputer's fit_transform invoked directly on the test frame), rather than reusing the values or fitted object from step 1.
3. PRESENT when both statements exist in the same script and the fill values applied to the test frame are derived from the test frame instead of the training frame.
• Predicted impact:
• Add score for C4 (Train and inference paths differ): 2
• weight: 2
• confidence: high
• Evidence: test_df.fillna(test_df.mean(numeric_only=True)) alongside train_df.fillna(train_df.mean(numeric_only=True)) produced a measured drop of about 0.021 in the held-out competition metric versus a variant that imputed both frames with training-derived statistics.
• Applies when: Tabular pipeline with separate train and test frames where missing values are imputed before model fitting and prediction.
• Example:
• Input:
    train_df = pd.read_csv('train.csv')
    test_df = pd.read_csv('test.csv')
    X_train = train_df.drop('label', axis=1).fillna(
        train_df.mean(numeric_only=True))
    y_train = train_df['label']
    X_test = test_df.fillna(test_df.mean(numeric_only=True))
    model.fit(X_train, y_train)
    preds = model.predict(X_test)
• Consequence:
    Test rows containing missing values are filled from the test frame's own
    column means, which differ from the training means the model was fitted
    against; affected feature values shift distribution and held-out accuracy
    on those rows drops (observed ~0.02 lower on the evaluation metric).
• Counter-example:
• Input:
    train_df = pd.read_csv('train.csv')
    test_df = pd.read_csv('test.csv')
    fill_values = train_df.mean(numeric_only=True)
    X_train = train_df.drop('label', axis=1).fillna(fill_values)
    y_train = train_df['label']
    X_test = test_df.fillna(fill_values)
    model.fit(X_train, y_train)
    preds = model.predict(X_test)
• Why it does not fire: The test frame is filled with the same training-derived statistics (fill_values) used for the training frame, so both paths impute identically.
C5 · budget-misuse One-vs-rest incremental linear classifier under extreme class countgithub_occurrence

P1 — One-vs-rest incremental linear classifier under extreme class count

• Pattern: Detects a linear one-vs-rest classifier trained incrementally on small minibatches while the full label vocabulary — numbering in the thousands — is registered upfront, so each per-class binary learner receives far fewer than one positive example per update.
• Detection procedure:
1. Locate an estimator whose multiclass strategy is one-vs-rest (e.g., SGDClassifier, PassiveAggressiveClassifier, Perceptron, or equivalent linear online learner) instantiated and then updated via an incremental fitting call such as partial_fit (or equivalent) inside a loop that iterates over minibatches of data.
2. Confirm the incremental fitting call receives (in its first invocation or via a classes= argument) an array of all distinct labels built from the complete label set — e.g., np.unique(labels), a precomputed class list, or a set accumulated over the whole training source — rather than only the classes present in the current batch.
3. Determine the number of distinct classes from a literal, a length check, a comment, or the construction at step 2; it must be at least 1000.
4. Determine the minibatch size from the literal constant or variable controlling batch accumulation in the loop at step 1; it must be numerically smaller than the class count found at step 3.
5. PRESENT when steps 1–4 all hold: an OvR-style linear learner is incrementally updated on batches smaller than a thousands-sized class vocabulary registered upfront.
• Predicted impact:
• Add score for C5: 3
• weight: 3
• confidence: high
• Evidence: clf.partial_fit(X_batch, y_batch, classes=all_classes) with thousands of classes and batches of a few hundred rows produced accuracy near a trivial majority baseline; an equivalent shared-softmax model on comparable features scored roughly 6x higher on the same metric.
• Applies when: A classification task has a very large label space (roughly 1000+ classes) and training is performed incrementally over minibatches with a linear classifier whose multiclass handling is one binary learner per class.
• Example:
• Input:
    from sklearn.linear_model import SGDClassifier
    import numpy as np

    all_classes = np.unique(labels)   # ~5000 distinct labels
    clf = SGDClassifier(loss="log_loss")
    BATCH = 512
    for X_batch, y_batch in iter_minibatches(feats, labels, BATCH):
        clf.partial_fit(X_batch, y_batch, classes=all_classes)
    preds = clf.predict(X_test)
• Consequence:
    Each of the ~5000 per-class binary learners sees on average ~0.1 positive
    examples per 512-row update, so weights barely move off the negative
    class; top-1 accuracy lands near the majority-class baseline, roughly
    6x lower than a single softmax model trained on the same features.
• Counter-example:
• Input:
    from sklearn.linear_model import SGDClassifier
    import numpy as np

    all_classes = np.unique(labels)   # 10 distinct labels
    clf = SGDClassifier(loss="log_loss")
    BATCH = 512
    for X_batch, y_batch in iter_minibatches(feats, labels, BATCH):
        clf.partial_fit(X_batch, y_batch, classes=all_classes)
    preds = clf.predict(X_test)
• Why it does not fire: the class count is far below 1000 and each 512-row batch supplies on average ~50 positives per class, so every binary learner receives ample positive signal per update.
C5 · budget-misuse Crippled boosted-ensemble capacity without validation-based selection

P1 — Crippled boosted-ensemble capacity without validation-based selection

• Pattern: Detects a boosted or bagged tree ensemble configured with a single-digit estimator count together with severely reduced tree size, trained once on the full data with no early stopping or validation-based model selection.
• Detection procedure:
1. Locate a constructor call for a gradient-boosting or tree-ensemble classifier/regressor (e.g. LGBMClassifier, XGBClassifier, GradientBoostingClassifier, RandomForestClassifier, or equivalent).
2. In that constructor's keyword arguments, find an estimator-count parameter (n_estimators, num_iterations, num_boost_round, or equivalent) assigned a literal integer less than or equal to 10.
3. In the same constructor, find at least one tree-size parameter assigned a literal value well below the library default: max_depth <= 3, num_leaves <= 8, or max_bin <= 32 (or equivalent parameters in another library).
4. Search the same script for any early-stopping mechanism (early_stopping_rounds, an early-stopping callback, or equivalent) or a validation split passed to the fitting call (eval_set or equivalent); confirm none is present.
5. Confirm the fitting call at step 1's model receives the full training frame (no hyperparameter search loop or comparison of multiple candidate models before the final fit).
6. The pattern is PRESENT when steps 2, 3, 4, and 5 all hold for the same model object.
• Predicted impact:
• Add score for C5: 3
• weight: 3
• confidence: high
• Evidence: A model configured as n_estimators=10, max_depth=3, num_leaves=7, max_bin=31 and fitted once with no validation scored ~0.10 worse on the held-out competition metric than the same library's model at default-scale capacity (n_estimators=100, num_leaves=31), a ~10% relative accuracy loss from underfitting.
• Applies when: Tabular supervised learning with a boosted-tree or tree-ensemble model on a training set large enough that default library capacity fits within the runtime budget.
• Example:
• Input:
    import lightgbm as lgb
    model = lgb.LGBMClassifier(
        n_estimators=10,
        learning_rate=0.5,
        max_depth=3,
        num_leaves=7,
        max_bin=31,
        random_state=42,
    )
    model.fit(X, y)
    preds = model.predict(X_test)
• Consequence:
    Held-out accuracy drops ~10% relative to the same model at default capacity
    (100 estimators, 31 leaves): the ensemble underfits, plateauing far below the
    achievable score while consuming only a small fraction of the time budget.
• Counter-example:
• Input:
    import lightgbm as lgb
    model = lgb.LGBMClassifier(
        n_estimators=1000,
        learning_rate=0.05,
        num_leaves=31,
        random_state=42,
    )
    model.fit(X_tr, y_tr, eval_set=[(X_val, y_val)],
              callbacks=[lgb.early_stopping(50)])
    preds = model.predict(X_test)
• Why it does not fire: The estimator budget is large and the number of trees actually used is chosen by early stopping on a validation split, so capacity is data-driven rather than fixed at a crippling value.
C5 · budget-misuse Blind fit with no validation-guided capacity selection

P1 — Blind fit with no validation-guided capacity selection

• Pattern: Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
• Detection procedure:
1. Locate a call to a model-fitting method (e.g. .fit(...), lgb.train(...), xgb.train(...), or equivalent) whose training inputs are the full training features and labels loaded from the training file, with no prior partitioning of those frames.
2. Confirm all hyperparameters passed to the model constructor or training call are literal constants (integers, floats, strings) rather than variables assigned from a search, loop, or selection routine.
3. Scan the entire script for any of: a call to a train/validation splitting function (e.g. train_test_split, KFold, or equivalent), a cross-validation utility, an early-stopping callback or early_stopping_rounds-style argument, or any metric function applied to predictions on rows held out from the fitting call at step 1. Record whether any is present.
4. Confirm the fitted model's prediction method is called on the test features and its output is written to the final output file without any comparison against an alternative model or configuration.
5. PRESENT when step 1 and step 2 hold, step 3 finds none of the listed constructs, and step 4 holds.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: model.fit(X, y) with all-literal constructor arguments followed immediately by model.predict(X_test) and file write, with no split, metric, or early stopping anywhere in the script; measured held-out score was 0.0215 lower (~2.2% relative accuracy) than a comparable-compute model that used a 10% validation split with early stopping to select its iteration count.
• Applies when: A supervised-learning script trains an iterative or ensemble model on a sizable tabular dataset and produces test-set predictions as its final artifact.
• Example:
• Input:
    train_df = pd.read_csv('train.csv')
    test_df = pd.read_csv('test.csv')
    X = train_df.drop(columns=['label'])
    y = train_df['label']
    model = RandomForestClassifier(n_estimators=100, max_depth=20,
                                   min_samples_leaf=2, random_state=42)
    model.fit(X, y)
    preds = model.predict(test_df)
    pd.DataFrame({'id': test_df['id'], 'label': preds}).to_csv('out.csv', index=False)
• Consequence:
    Final test accuracy is systematically below what the same compute achieves
    with validation-guided capacity selection: ~2% relative accuracy lost
    versus an early-stopped boosted model of comparable training cost, with
    no signal in the run indicating the model is under- or over-fitted.
• Counter-example:
• Input:
    X_tr, X_val, y_tr, y_val = train_test_split(X, y, test_size=0.1, random_state=42)
    train_set = lgb.Dataset(X_tr, label=y_tr)
    val_set = lgb.Dataset(X_val, label=y_val, reference=train_set)
    model = lgb.train(params, train_set, num_boost_round=1000,
                      valid_sets=[val_set],
                      callbacks=[lgb.early_stopping(stopping_rounds=30)])
    preds = model.predict(X_test, num_iteration=model.best_iteration)
• Why it does not fire: the script carves out a validation split and uses an early-stopping callback so the iteration count is selected by a monitored validation metric, failing step 3.
C5 · budget-misuse Default regularization on high-dimensional linear probability model

P1 — Default regularization on high-dimensional linear probability model

• Pattern: Detects a probabilistic linear classifier fitted on a high-dimensional handcrafted feature matrix using the library's default regularization strength, with no held-out or cross-validated search over regularization values.
• Detection procedure:
1. Locate a constructor call for a regularized linear classification model that outputs class probabilities (e.g., LogisticRegression or equivalent) whose fitted object is later used to produce per-class probability predictions (e.g., a call to predict_proba or equivalent).
2. Check the constructor's arguments: the regularization-strength parameter (e.g., C, alpha, or equivalent) is either absent or set to the library default value.
3. Verify the feature matrix passed to the model's fit call is built by concatenating or stacking handcrafted descriptors (histogram, gradient, statistical, or hand-coded pixel features) into a vector whose declared or computed dimension is in the thousands, or is comparable to or larger than the number of training rows visible in the script.
4. Search the same file for any hyperparameter-search construct over the regularization parameter: a loop over candidate values with evaluation on a held-out split, or a call to a cross-validated search utility (e.g., GridSearchCV, LogisticRegressionCV, or equivalent). None is present.
5. PRESENT when steps 1–4 all hold: default-strength regularization, high-dimensional handcrafted features, probability output, and no tuning of the regularization parameter anywhere in the file.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: clf = LogisticRegression(max_iter=1000) fitted on a multi-thousand-dimensional handcrafted descriptor matrix with many classes and few samples per class produced overconfident probabilities; a variant identical except for strong regularization (C=0.01) scored 4.55 on the log-loss-style metric versus 11.92 for the default, a 0.62 normalized-metric gap.
• Applies when: The task is multi-class probabilistic classification scored by a calibration-sensitive metric (e.g., log loss), the feature dimension is comparable to or exceeds the number of training samples, and a regularized linear model is the final estimator.
• Example:
• Input:
    X = np.array([extract_hog(p) for p in train_paths], dtype=np.float32)  # ~8000 dims
    scaler = StandardScaler()
    Xs = scaler.fit_transform(X)

    clf = LogisticRegression(max_iter=1000)   # default C=1.0, no tuning
    clf.fit(Xs, y)

    Xt = scaler.transform(np.array([extract_hog(p) for p in test_paths]))
    probs = clf.predict_proba(Xt)
    pd.DataFrame(probs, columns=clf.classes_).to_csv("submission.csv", index=False)
• Consequence:
    The model overfits the underdetermined feature space and emits sharply
    peaked, poorly calibrated probabilities; log loss on held-out data is
    roughly 2-3x worse than a strongly regularized fit of the same model
    (observed: 11.92 vs 4.55 on the same features).
• Counter-example:
• Input:
    X = np.array([extract_hog(p) for p in train_paths], dtype=np.float32)
    Xs = StandardScaler().fit_transform(X)

    best_c, best_loss = None, np.inf
    for c in [0.001, 0.01, 0.1, 1.0]:
        m = LogisticRegression(C=c, max_iter=1000)
        loss = -cross_val_score(m, Xs, y, scoring="neg_log_loss", cv=3).mean()
        if loss < best_loss:
            best_c, best_loss = c, loss
    clf = LogisticRegression(C=best_c, max_iter=1000).fit(Xs, y)
• Why it does not fire: The regularization strength is selected by a cross-validated grid over candidate values, so step 4's absence-of-tuning condition fails.
C5 · budget-misuse Hardcoded tiny training subset with sequential per-file feature extractionverified_trace · effect +0.1420

P1 — Hardcoded tiny training subset with sequential per-file feature extraction

• Pattern: Detects a training pipeline that truncates the list of available training files to a small hardcoded literal count and then extracts features from them one file at a time in a single-process loop, leaving most of the data and all available parallelism unused.
• Detection procedure:
1. Locate an assignment that builds a collection of training file names, paths, or ids (e.g., from os.listdir, glob.glob, Path.iterdir, or reading an index/metadata file, or equivalent).
2. Check that this collection is reduced using a hardcoded integer literal — via a slice such as [:900], a call like random.sample(files, 900) or .sample(n=900), or a loop break after a literal counter threshold — where the literal is written directly in source and is not computed from the collection's length, a measured time budget, or a command-line/config parameter.
3. Find a for loop (or list comprehension) over the reduced collection whose body calls a function that opens or reads each file and computes per-file features, executing in the main process.
4. Verify that this extraction is not dispatched through multiprocessing.Pool, concurrent.futures, joblib.Parallel, or an equivalent worker-pool mechanism anywhere in the same script.
5. The pattern is PRESENT when a hardcoded literal truncation of the training file list (step 2) feeds a sequential single-process per-file extraction loop (steps 3–4) that produces the training matrix.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A run that truncated the training file list with a hardcoded slice (selecting roughly 1% of ~70,000 available files, 900 ids) and extracted features in a sequential loop scored 0.58 on the held-out metric, while the same pipeline using a worker pool over ~9,000 ids in comparable wall-clock scored 0.82 — a 0.29 gap with no other change to the model.
• Applies when: The training set contains many individual files (thousands or more), features are computed per file, and the runtime budget permits substantially more processing than the hardcoded subset consumes.
• Example:
• Input:
    files = sorted(os.listdir(train_dir))[:900]
    feats = []
    labels = []
    for fname in files:
        f = extract_features(os.path.join(train_dir, fname))
        feats.append(f)
        labels.append(label_for(fname))
    X = np.vstack(feats)
    y = np.array(labels)
    model.fit(X, y)
• Consequence:
    Held-out metric drops from 0.82 to 0.58 (-0.29) versus the same pipeline
    trained on 10x more files, which fits within the same wall-clock budget
    when extraction is parallelized; the run exits cleanly either way.
• Counter-example:
• Input:
    files = sorted(os.listdir(train_dir))
    n = min(len(files), int(budget_seconds / est_sec_per_file * os.cpu_count()))
    subset = files[:n]
    paths = [os.path.join(train_dir, f) for f in subset]
    with Pool(os.cpu_count()) as pool:
        feats = pool.map(extract_features, paths)
    X = np.vstack(feats)
    model.fit(X, np.array([label_for(f) for f in subset]))
• Why it does not fire: the subset size is computed from the compute budget and available core count rather than a hardcoded literal, and extraction is dispatched through a worker pool instead of a sequential single-process loop.
C5 · budget-misuse Early-stopped model predicts test without refit on full training dataverified_trace · effect +0.1106

P1 — Early-stopped model predicts test without refit on full training data

• Pattern: Detects a gradient-boosted or iteratively trained model that is fit only on a train partition (with a held-out validation split used to select a stopping iteration) and is then used directly to predict the test set, with no subsequent refit on the combined train-plus-validation rows.
• Detection procedure:
1. Find a call that splits labeled data into two partitions (e.g., train_test_split or equivalent slicing/masking that assigns rows to a train subset and a validation subset within the same script).
2. Find a model fitting call that receives the train subset as training data and the validation subset as an evaluation set together with an early-stopping mechanism (e.g., early_stopping callback, early_stopping_rounds argument, or equivalent), and note the variable the fitted model is assigned to.
3. Find a prediction call on the test feature matrix (a frame or array derived from unlabeled/test files) and note which model variable it is invoked on.
4. Check whether, between step 2 and step 3, there is any second fitting call whose training data is the union of the train and validation subsets (or the full pre-split labeled matrix) and whose model output is the one used at step 3.
5. PRESENT if the model variable used for the test prediction at step 3 is the same variable assigned at step 2, and no refit on the full labeled data as described in step 4 occurs before the test prediction.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: preds = model.predict(X_test, num_iteration=model.best_iteration) where model was fit on only the post-split train partition; a contrastive run that refit on the full labeled matrix at approximately the selected iteration count scored 0.13705 higher on the evaluation metric.
• Applies when: Supervised training scripts that carve a validation fraction out of the labeled data to drive early stopping and then produce test-set predictions in the same run, especially when the labeled set is small enough that 10–20% of rows is a meaningful fraction of capacity.
• Example:
• Input:
    Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.15, random_state=42)
    dtr = lgb.Dataset(Xtr, label=ytr)
    dva = lgb.Dataset(Xva, label=yva, reference=dtr)
    model = lgb.train(
        params, dtr, num_boost_round=3000, valid_sets=[dva],
        callbacks=[lgb.early_stopping(100)],
    )
    preds = model.predict(X_test, num_iteration=model.best_iteration)
    sub = pd.DataFrame({"Id": ids, "Label": preds})
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Final model is fit on ~85% of the labeled rows; the test-set metric is
    measurably lower (observed gap ~0.14 on the evaluation metric) than a
    model refit on 100% of the labeled data at the selected iteration count.
• Counter-example:
• Input:
    Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.15, random_state=42)
    dtr = lgb.Dataset(Xtr, label=ytr)
    dva = lgb.Dataset(Xva, label=yva, reference=dtr)
    model = lgb.train(
        params, dtr, num_boost_round=3000, valid_sets=[dva],
        callbacks=[lgb.early_stopping(100)],
    )
    best_it = model.best_iteration or 500
    final = lgb.train(params, lgb.Dataset(X, label=y),
                      num_boost_round=max(50, int(best_it * 1.1)))
    preds = final.predict(X_test)
• Why it does not fire: the test prediction is invoked on final, a second model refit on the full pre-split labeled matrix X for approximately the iteration count selected by early stopping, satisfying the refit condition in step 4.
C5 · budget-misuse CV metric printed but never used for model selectiongithub_occurrence

P1 — CV metric printed but never used for model selection

• Pattern: Detects a script that computes an out-of-fold cross-validation metric solely to print it, while training exactly one hardcoded model configuration with no loop over hyperparameter values or alternative model families whose results feed a selection step.
• Detection procedure:
1. Find a k-fold cross-validation loop (e.g. iterating over splits from StratifiedKFold, KFold, or equivalent) that fits a model per fold and stores validation-split predictions into an out-of-fold array.
2. Find a call to a scoring function (e.g. roc_auc_score, log_loss, or equivalent) applied to that out-of-fold array, whose return value is either passed directly to print/logging or assigned to a variable that is never subsequently compared, ranked, or used in any conditional or selection expression.
3. Confirm the model constructor inside the fold loop is instantiated with literal constant hyperparameters (e.g. a fixed regularization value) and that this constructor call is not enclosed in any loop or comprehension over multiple hyperparameter values or multiple model classes.
4. Confirm there is no other code in the file that fits a second model family or a second hyperparameter setting and compares metrics between configurations.
5. PRESENT when all of steps 1–4 hold: the CV harness exists, its metric is print-only, and exactly one fixed configuration is ever trained.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A single fixed linear classifier LogisticRegression(C=0.01, max_iter=2000) was trained inside a 5-fold CV harness whose OOF AUC was only printed; the run finished in ~2 minutes of a multi-hour budget and scored at chance level (~0.50), roughly 0.13 AUC below what a tuned nonlinear learner achieved on the same features.
• Applies when: The script already contains a working k-fold CV harness with out-of-fold predictions, the task has a substantial runtime budget relative to observed runtime, and only one model configuration appears in the file.
• Example:
• Input:
    skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
    oof = np.zeros(len(y))
    test_pred = np.zeros(len(X_test))
    for tr, va in skf.split(X, y):
        m = make_pipeline(StandardScaler(),
                          LogisticRegression(C=0.01, max_iter=2000))
        m.fit(X[tr], y[tr])
        oof[va] = m.predict_proba(X[va])[:, 1]
        test_pred += m.predict_proba(X_test)[:, 1] / 5
    print("OOF AUC:", roc_auc_score(y, oof))
    sub["LABEL_A"] = test_pred
• Consequence:
    Held-out AUC lands near chance (~0.50) while a small grid over
    regularization strengths plus one boosted-tree alternative, selected
    by the same OOF metric, reaches ~0.58 — a ~13% relative deficit —
    with the run using ~2 minutes of a multi-hour compute budget.
• Counter-example:
• Input:
    best_score, best_cfg = -1, None
    for C in [0.01, 0.1, 1.0, 10.0]:
        oof = np.zeros(len(y))
        for tr, va in skf.split(X, y):
            m = make_pipeline(StandardScaler(), LogisticRegression(C=C))
            m.fit(X[tr], y[tr])
            oof[va] = m.predict_proba(X[va])[:, 1]
        score = roc_auc_score(y, oof)
        if score > best_score:
            best_score, best_cfg = score, C
    final = make_pipeline(StandardScaler(),
                          LogisticRegression(C=best_cfg)).fit(X, y)
• Why it does not fire: The OOF metric is compared across a grid of hyperparameter values and drives selection of the configuration used for the final fit, so the metric is not print-only and more than one configuration is trained.
C5 · budget-misuse Small hard vocabulary cap on sparse-text features for a linear model

P1 — Small hard vocabulary cap on sparse-text features for a linear model

• Pattern: Detects a text vectorizer configured with a small hard vocabulary-size cap (a literal integer of 10000 or less) whose output feeds a sparse linear or naive Bayes classifier, truncating discriminative terms without any stated memory or resource constraint.
• Detection procedure:
1. Locate a construction of a bag-of-words or TF-IDF text vectorizer (e.g. TfidfVectorizer, CountVectorizer, or equivalent) that includes a vocabulary-cap keyword argument (e.g. max_features=) set to a literal integer whose value is less than or equal to 10000.
2. Follow the variable holding that vectorizer's transformed output (via fit_transform or transform, or equivalent) to a .fit(...) call on a model object.
3. Check that the model object was constructed from a linear classifier or naive Bayes class (e.g. LogisticRegression, SGDClassifier, LinearSVC, MultinomialNB, or equivalent) — models whose training cost scales tractably with sparse feature width.
4. Check that no comment or configuration in the same file states a memory, latency, or deployment-size constraint motivating the cap.
5. The pattern is PRESENT when steps 1–4 all hold: a literal cap ≤ 10000 on the vectorizer, its output trains a sparse linear or naive Bayes model, and no resource constraint is stated.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 3
• weight: 3
• confidence: high
• Evidence: A pipeline using TfidfVectorizer(max_features=5000, ...) feeding a logistic regression scored ~0.58 log loss, while an otherwise comparable pipeline with a vocabulary two orders of magnitude larger (plus min_df pruning) scored ~0.35 on the same probabilistic metric — a 0.23 degradation caused by dropping rare but highly class-indicative n-grams.
• Applies when: Text classification with bag-of-words or TF-IDF features feeding a sparse linear or naive Bayes model, where the corpus fits in memory and no explicit resource budget is documented.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(max_features=5000, stop_words='english')
    X_tr = vec.fit_transform(df_train['col_a'])
    X_te = vec.transform(df_test['col_a'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    probs = model.predict_proba(X_te)
• Consequence:
    Log loss on the held-out set worsens from ~0.35 (uncapped vocabulary
    with min_df pruning) to ~0.58, because rare but strongly
    class-indicative n-grams are excluded from the 5000-term vocabulary.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(min_df=3, max_features=100000,
                          stop_words='english')
    X_tr = vec.fit_transform(df_train['col_a'])
    X_te = vec.transform(df_test['col_a'])

    model = LogisticRegression(C=4.0, max_iter=1000)
    model.fit(X_tr, y_train)
    probs = model.predict_proba(X_te)
• Why it does not fire: The vocabulary cap is 100000 (above the 10000 threshold) and dimensionality is controlled with min_df instead of an aggressive fixed feature count, so discriminative terms are retained.
C5 · budget-misuse Untuned single default-regularized classifier on a probability metric

P1 — Untuned single default-regularized classifier on a probability metric

• Pattern: Detects a probability-scored prediction pipeline that fits exactly one classifier using the library-default regularization strength, with no hyperparameter search, no validation-based tuning, and no combination of predictions from a second model family before writing the output probabilities.
• Detection procedure:
1. Locate the call that writes the final prediction file (e.g. to_csv or equivalent) and confirm the values written come from a per-class probability method (e.g. predict_proba or equivalent softmax output).
2. Trace the model object whose probability method feeds step 1 to its constructor; check whether the constructor call omits the regularization-strength argument (e.g. C= or alpha= or equivalent), leaving the library default.
3. Scan the whole script for any hyperparameter-search construct (e.g. GridSearchCV, RandomizedSearchCV, a loop over candidate parameter values scored on a held-out split, or equivalent); note whether none exists.
4. Scan the whole script for a second fitted estimator of a different family whose probability outputs are arithmetically combined (weighted sum, mean, or stacking) with the first model's outputs before step 1; note whether none exists.
5. PRESENT when the conditions in steps 2, 3, and 4 all hold: the written probabilities come from a single classifier at default regularization strength with no tuning loop and no cross-family blending anywhere in the script.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A script fitting one linear classifier via LogisticRegression(max_iter=1000, solver='lbfgs') (default regularization) and submitting its predict_proba output scored 0.37863 worse on the probability metric than an equally cheap run that set the regularization parameter and averaged in a second model family's probabilities, despite ample remaining time budget.
• Applies when: The task is scored on predicted probabilities (e.g. log loss), classical ML estimators are used, and the runtime budget comfortably allows fitting more than one cheap model or a small parameter sweep.
• Example:
• Input:
    vec = TfidfVectorizer(max_features=10000, stop_words='english')
    X_tr = vec.fit_transform(df['col_a'])
    X_te = vec.transform(test['col_a'])

    model = LogisticRegression(max_iter=1000, random_state=42)
    model.fit(X_tr, y)
    proba = model.predict_proba(X_te)

    sub = pd.DataFrame(proba, columns=classes)
    sub.insert(0, 'id', test['id'])
    sub.to_csv('submission.csv', index=False)
• Consequence:
    Log loss on the held-out evaluation is substantially higher (observed gap
    ~0.38) than an equally cheap run that tuned the regularization strength
    and blended a second model family, leaving most of the available
    training-time budget unused.
• Counter-example:
• Input:
    lr = LogisticRegression(C=4, max_iter=1000)
    lr.fit(X_tr, y)
    p_lr = lr.predict_proba(X_te)

    nb = MultinomialNB(alpha=0.05)
    nb.fit(X_tr, y)
    p_nb = nb.predict_proba(X_te)

    proba = 0.7 * p_lr + 0.3 * p_nb
    sub = pd.DataFrame(proba, columns=classes)
    sub.insert(0, 'id', test['id'])
    sub.to_csv('submission.csv', index=False)
• Why it does not fire: the regularization strength is explicitly set on both models (failing step 2) and the submitted probabilities are a weighted blend of two complementary model families (failing step 4).
C5 · budget-misuse Small hardcoded vocabulary cap on a large text corpusverified_trace · effect +0.2170

P1 — Small hardcoded vocabulary cap on a large text corpus

• Pattern: Detects a text-vectorization step whose vocabulary is capped by a hardcoded feature limit in the low thousands (5,000 or fewer) while feeding a fast-to-train sparse linear or naive-Bayes model, truncating rarer discriminative n-grams despite ample compute headroom.
• Detection procedure:
1. Locate any instantiation of a text vectorizer that produces a sparse term-document matrix (e.g. TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent) whose keyword argument max_features (or equivalent vocabulary-limiting parameter such as n_features) is set to a literal integer.
2. Check that the literal integer at step 1 is less than or equal to 5000.
3. Within the same script, find the matrix produced by the vectorizer at step 1 passed to a .fit(...) call (or equivalent training call) of a linear classifier or naive-Bayes classifier (e.g. LogisticRegression, SGDClassifier, MultinomialNB, LinearSVC, or equivalent).
4. Confirm no other vectorizer in the script produces features for the same model with a larger or absent vocabulary limit (e.g. no second vectorizer whose output is stacked with the first via hstack or a feature union with a higher cap).
5. The pattern is PRESENT when steps 1–4 all hold: a literal vocabulary cap of at most 5000 gates the only sparse text representation consumed by a cheap sparse-friendly classifier.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using TfidfVectorizer(max_features=5000) feeding linear and naive-Bayes classifiers scored 0.40 worse on the held-out probabilistic metric than an otherwise similar pipeline that used an order-of-magnitude larger word-plus-character n-gram vocabulary; both runs completed in seconds, so the truncated vocabulary discarded rare discriminative n-grams for no compute savings.
• Applies when: A text-classification script vectorizes a corpus of thousands of documents or more into sparse features for a linear or naive-Bayes model, and the runtime budget is not tightly constrained (single fast fit, no memory-limited environment indicated).
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(max_features=3000, ngram_range=(1, 2))
    X_train = vec.fit_transform(train["text"])
    X_test = vec.transform(test["text"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
• Consequence:
    Held-out log loss is systematically higher (e.g. ~0.4 worse) than the same
    pipeline with the cap removed or raised to tens of thousands of features,
    because rare but discriminative n-grams are truncated from the vocabulary;
    training time remains seconds either way, so the cap buys nothing.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression
    from scipy.sparse import hstack

    word_vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
    char_vec = TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5),
                               max_features=50000)
    X_train = hstack([word_vec.fit_transform(train["text"]),
                      char_vec.fit_transform(train["text"])])
    model = LogisticRegression(max_iter=1000).fit(X_train, y_train)
• Why it does not fire: The word vectorizer has no vocabulary cap and the character vectorizer's cap is 50,000, so no vectorizer feeding the model is limited to 5,000 or fewer features.
C5 · budget-misuse Unvalidated high-capacity boosted ensemble on small handcrafted-feature settrace_observed

P1 — Unvalidated high-capacity boosted ensemble on small handcrafted-feature set

• Pattern: Detects a boosted or deep tree ensemble classifier fitted exactly once with hard-coded, unvalidated capacity settings (100+ rounds and unbounded or deep trees, no minimum-leaf-size constraint) on a small handcrafted-feature matrix, with its probability outputs written directly as the ranked predictions.
• Detection procedure:
1. Find a constructor call for a tree-ensemble classifier (e.g. GradientBoostingClassifier, XGBClassifier, LGBMClassifier, or equivalent) whose arguments include a literal estimator/round count of 100 or greater, and which either omits a tree-depth argument entirely or passes a literal depth of 5 or more, and which passes no minimum-leaf-size argument (min_samples_leaf, min_child_weight, min_data_in_leaf, or equivalent).
2. Verify the training feature matrix passed to that model's fit call is assembled inside the same file by iterating over per-sample inputs and computing summary statistics or descriptors (appending per-sample vectors to a list, then converting to an array), rather than being loaded from a pre-built tabular training file with thousands of rows.
3. Confirm the file contains no call to cross_val_score, GridSearchCV, RandomizedSearchCV, cross_validate, or equivalent applied to that model class, and no second fit of the model class with different capacity arguments compared on a held-out split created by train_test_split or equivalent.
4. Confirm the fitted model's predict_proba (or equivalent probability method) output on the test features is written to the final output file without any post-hoc calibration or blending with another model.
5. PRESENT when steps 1–4 all hold: fixed high-capacity settings, small handcrafted feature matrix, no held-out or cross-validated capacity selection, and raw probabilities emitted.
• Predicted impact:
• Add score for C5: 2
• weight: 2
• confidence: high
• Evidence: A boosted-style ensemble instantiated with fixed literal capacity arguments and fit once on a few hundred rows of handcrafted image statistics (model.fit(X_scaled, y) with no validation loop) produced held-out ranking quality of ~0.49 — at or below chance — versus ~0.55 (a 10.9% relative gap) for a leaf-regularized ensemble on identical-style features.
• Applies when: the labeled training set is small (hundreds of rows), features are weak handcrafted descriptors extracted per sample, and the evaluation metric is a probability-ranking metric on a binary target.
• Example:
• Input:
    X, y = [], []
    for sample_id in train_ids:
        X.append(extract_stats(sample_id))   # handcrafted stats per sample
        y.append(labels[sample_id])
    X = scaler.fit_transform(np.array(X))    # ~400 rows
    model = GradientBoostingClassifier(n_estimators=200, max_depth=6)
    model.fit(X, y)
    proba = model.predict_proba(scaler.transform(X_test))[:, 1]
    pd.DataFrame({'id': test_ids, 'LABEL_A': proba}).to_csv('out.csv', index=False)
• Consequence:
    The ensemble memorizes feature noise on the tiny training set; held-out
    ranking score drops to ~0.49 (chance level), versus ~0.55 achievable with
    capacity chosen against a held-out estimate — a ~11% relative degradation.
• Counter-example:
• Input:
    X, y = [], []
    for sample_id in train_ids:
        X.append(extract_stats(sample_id))
        y.append(labels[sample_id])
    X = scaler.fit_transform(np.array(X))
    grid = {'max_depth': [2, 3], 'min_samples_leaf': [5, 20]}
    search = GridSearchCV(GradientBoostingClassifier(n_estimators=100),
                          grid, cv=5, scoring='roc_auc')
    search.fit(X, y)
    proba = search.best_estimator_.predict_proba(scaler.transform(X_test))[:, 1]
    pd.DataFrame({'id': test_ids, 'LABEL_A': proba}).to_csv('out.csv', index=False)
• Why it does not fire: depth and leaf-size are selected via cross-validation over a candidate grid, so the capacity settings are validated against a held-out estimate rather than hard-coded (step 3 fails).
C5 · budget-misuse Train-once fit with no validation estimate before submissionverified_trace · effect +0.0698

P1 — Train-once fit with no validation estimate before submission

• Pattern: Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
• Detection procedure:
1. Locate a call that fits a supervised model (e.g. .fit(X, y), lgb.train, xgb.train, or equivalent) inside the script that also writes the prediction/submission file.
2. Check whether, anywhere in the same script, the training rows are partitioned before that fit — e.g. a call to train_test_split, KFold, GroupKFold, StratifiedKFold, or equivalent manual index masking that separates rows into a fitted subset and a held-out subset.
3. Check whether any evaluation-metric function (e.g. roc_auc_score, accuracy_score, log_loss, a custom metric function, or equivalent) is called on predictions from held-out rows before the submission-writing call (e.g. .to_csv).
4. Check whether the fit call receives any validation set or early-stopping mechanism (e.g. valid_sets=, eval_set=, an early-stopping callback, or equivalent).
5. PRESENT if step 1 finds a fit whose inputs are all assembled training rows, AND steps 2, 3, and 4 all find nothing — no partition, no pre-submission metric on held-out data, and no validation/early-stopping argument to the fit.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A pipeline that ran clf.fit(X_scaled, y) on all rows with default hyperparameters and immediately wrote predict_proba outputs to the submission scored ~0.098 lower on the task metric than an otherwise comparable pipeline that held out a validation split, early-stopped on it, printed a validation metric, and refit at the selected iteration budget.
• Applies when: A supervised-learning script trains a tunable or iterative model and writes test-set predictions to a submission or output file in the same run.
• Example:
• Input:
    df = pd.read_csv("train.csv")
    X = df[feature_cols].values
    y = df["target"].values
    scaler = StandardScaler().fit(X)
    model = RandomForestClassifier(n_estimators=100, random_state=0)
    model.fit(scaler.transform(X), y)
    test = pd.read_csv("test.csv")
    preds = model.predict_proba(scaler.transform(test[feature_cols]))[:, 1]
    pd.DataFrame({"Id": test["Id"], "Label": preds}).to_csv("submission.csv", index=False)
• Consequence:
    Model capacity and iteration budget are never selected against held-out
    performance; the scored metric lands measurably below what the same
    pipeline achieves with a validated, early-stopped configuration
    (observed ~10% relative shortfall on the task metric).
• Counter-example:
• Input:
    Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.2, random_state=0)
    dtr = lgb.Dataset(Xtr, label=ytr)
    dva = lgb.Dataset(Xva, label=yva)
    model = lgb.train(params, dtr, num_boost_round=3000, valid_sets=[dva],
                      callbacks=[lgb.early_stopping(100)])
    print("val AUC:", roc_auc_score(yva, model.predict(Xva)))
    final = lgb.train(params, lgb.Dataset(X, label=y),
                      num_boost_round=model.best_iteration)
    pd.DataFrame({"Id": ids, "Label": final.predict(Xte)}).to_csv("submission.csv", index=False)
• Why it does not fire: The training rows are split, the fit uses a validation set with early stopping, and a metric on held-out rows is computed before the submission is written — the final full-data refit merely reuses the validated iteration budget.
C5 · budget-misuse Validation metric computed but never used for hyperparameter selectiongithub_occurrence

P1 — Validation metric computed but never used for hyperparameter selection

• Pattern: Detects a pipeline that holds out a validation split and scores a probabilistic classifier on it, then refits the final model with the same hardcoded hyperparameter values, using the validation estimate for nothing beyond printing or logging.
• Detection procedure:
1. Find a call that partitions labeled training data into two subsets (e.g., train_test_split or equivalent slicing that produces a train part and a held-out part within the same script).
2. Find a classifier constructed with literal hyperparameter values (e.g., a regularization strength, tree depth, or estimator count written as a constant), fitted on the train part, and evaluated on the held-out part via a scoring call (e.g., log_loss, accuracy_score, or equivalent) whose result is only passed to print, a logger, or an f-string — never compared, stored in a running best, or used in a conditional.
3. Find a second classifier of the same estimator type constructed later in the same script with hyperparameter values that are all literal constants (no variable set inside a loop or search, and no call to a search utility such as GridSearchCV or equivalent anywhere in the script), fitted on the full training data, and used to produce the final predictions.
4. PRESENT when steps 1–3 all hold: a held-out score exists, no hyperparameter in the final fit is derived from any comparison over multiple candidate values, and no automated search construct appears.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A script computed a held-out score with print(log_loss(y_val, p_val)) and then refit LogisticRegression(max_iter=2000, C=1.0) on the full data with the same hardcoded C; an otherwise identical pipeline that looped over four values of C, selected the best on the same held-out split, and refit achieved a log loss roughly 0.10 better on the evaluation metric.
• Applies when: The script trains a probabilistic classifier with at least one tunable complexity/regularization hyperparameter, creates a held-out validation split from the training data, and is evaluated on a continuous metric such as log loss.
• Example:
• Input:
    xa, xb, ya, yb = train_test_split(X, y, test_size=0.1, random_state=0)
    clf = LogisticRegression(max_iter=2000, C=1.0)
    clf.fit(xa, ya)
    p = clf.predict_proba(xb)[:, 1]
    print("val logloss:", log_loss(yb, p))

    clf_full = LogisticRegression(max_iter=2000, C=1.0)
    clf_full.fit(X, y)
    preds = clf_full.predict_proba(X_test)[:, 1]
• Consequence:
    Final test log loss is ~10% higher (worse) than the same pipeline that
    sweeps C over a small grid, picks the best value on the held-out split,
    and refits on all data; the validation score is computed but wasted.
• Counter-example:
• Input:
    xa, xb, ya, yb = train_test_split(X, y, test_size=0.1, random_state=0)
    best_c, best_s = None, float("inf")
    for C in [0.003, 0.01, 0.03, 0.1]:
        m = LogisticRegression(max_iter=1000, C=C)
        m.fit(xa, ya)
        s = log_loss(yb, m.predict_proba(xb)[:, 1])
        if s < best_s:
            best_s, best_c = s, C
    model = LogisticRegression(max_iter=2000, C=best_c)
    model.fit(X, y)
    preds = model.predict_proba(X_test)[:, 1]
• Why it does not fire: The held-out score participates in a comparison (s < best_s) that selects the hyperparameter used in the final refit, so the validation estimate is consumed by a selection loop rather than only printed.
C5 · budget-misuse Hardcoded small vocabulary cap on sparse text features feeding a linear modelverified_trace · effect +0.1475

P1 — Hardcoded small vocabulary cap on sparse text features feeding a linear model

• Pattern: Detects a text-to-sparse-matrix vectorizer constructed with a small hardcoded vocabulary-size cap whose output is consumed by a sparse-capable linear or naive-Bayes classifier, truncating the feature space the model could otherwise use.
• Detection procedure:
1. Locate a call constructing a bag-of-words or TF-IDF style vectorizer (TfidfVectorizer, CountVectorizer, or equivalent) whose keyword arguments include max_features set to a literal integer of 50000 or less.
2. Within the same script or module, find a variable assigned from that vectorizer's fit_transform (or fit followed by transform) applied to a text column of a training frame.
3. Find that variable (or a derived split of it) passed to the fit method of a linear classifier or naive-Bayes classifier (LogisticRegression, SGDClassifier, LinearSVC, MultinomialNB, ComplementNB, or equivalent) that accepts sparse input.
4. Confirm there is no other vectorizer in the script producing an uncapped feature set that is concatenated or stacked with the capped one before the fit at step 3.
5. PRESENT when all of steps 1–4 hold: a capped vocabulary (literal max_features ≤ 50000) is the sole sparse text representation fitted by a sparse-capable linear or naive-Bayes model.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A run using TfidfVectorizer(max_features=5000, ngram_range=(1, 2), min_df=2, max_df=0.8) feeding a logistic regression scored 0.54 multiclass log loss, while a comparable pipeline consuming the full vocabulary scored 0.33 — a ~38% relative degradation attributable to feature impoverishment, since discriminative rare terms were pruned from the vocabulary.
• Applies when: A text-classification script builds bag-of-words or TF-IDF features and trains a linear or naive-Bayes model on the resulting sparse matrix; not in scope when the downstream model requires dense input or memory is demonstrably constrained (e.g., the matrix is explicitly densified).
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    tfidf = TfidfVectorizer(max_features=5000, ngram_range=(1, 2), min_df=2)
    X_tr = tfidf.fit_transform(df["text_col"])
    X_te = tfidf.transform(test_df["text_col"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Consequence:
    Multiclass log loss rises from ~0.33 (full vocabulary) to ~0.54 with the
    5000-feature cap: rare but class-discriminative terms are dropped, so
    predicted probabilities are systematically less confident on the correct
    class. Direction: evaluation metric worsens by ~38% relative.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    tfidf = TfidfVectorizer(ngram_range=(1, 2), min_df=2, sublinear_tf=True)
    X_tr = tfidf.fit_transform(df["text_col"])
    X_te = tfidf.transform(test_df["text_col"])

    model = LogisticRegression(C=2.0, max_iter=1000)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Why it does not fire: The vectorizer omits max_features entirely, letting the full vocabulary reach the sparse-capable linear model, with noise controlled by min_df and regularization instead of a vocabulary cap.
C5 · budget-misuse Single default-configured probabilistic model with no ensembling or validationverified_trace · effect +0.0902

P1 — Single default-configured probabilistic model with no ensembling or validation

• Pattern: Detects a script that fits exactly one probabilistic classifier, writes its predicted class probabilities straight to the output file, and contains no cross-validation loop, no second model, and no out-of-fold evaluation, leaving obvious cheap ensembling capacity unused.
• Detection procedure:
1. Count the number of distinct estimator objects instantiated in the script that expose a fitting method (e.g. fit, or equivalent); the count of classifier instances is exactly one.
2. Confirm the script's output artifact (e.g. a call to to_csv or equivalent file write) receives columns derived directly from a single call to a probability-prediction method (e.g. predict_proba or equivalent) on that one classifier, with no arithmetic combining outputs from more than one model.
3. Confirm there is no cross-validation construct in the file: no fold-splitting object (e.g. StratifiedKFold, KFold, cross_val_predict, or equivalent), no loop over train/validation index pairs, and no computation of the scored loss on held-out rows.
4. Confirm the classifier's constructor call passes no hyperparameters that alter model capacity or regularization strength (arguments limited to bookkeeping such as random_state, max_iter, n_jobs, or none at all).
5. PRESENT when all of steps 1–4 hold: one default-capacity classifier, direct write of its probabilities, and no validation or blending anywhere in the script.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A script that ran model = LogisticRegression(max_iter=1000); model.fit(...) once and wrote model.predict_proba(...) directly scored 0.54 log loss in ~5 seconds, while a same-features blend of three cheap diverse classifiers with out-of-fold stacking scored 0.30 in ~64 seconds — a 0.44 metric gap with both runs far under the time budget.
• Applies when: The task is probabilistic classification scored by a proper loss (e.g. multiclass log loss), the feature matrix supports several fast linear/count-based learners, and observed or expected runtime is a small fraction of the allowed budget.
• Example:
• Input:
    tfidf = TfidfVectorizer(max_features=5000)
    X_tr = tfidf.fit_transform(train_df["col_a"])
    X_te = tfidf.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000, random_state=42)
    model.fit(X_tr, y_enc)
    proba = model.predict_proba(X_te)

    sub = pd.DataFrame(proba, columns=classes)
    sub.insert(0, "id", test_df["id"].values)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Multiclass log loss is substantially higher (worse) than a same-budget
    blend of several cheap diverse classifiers: 0.54 vs 0.30 on identical
    features, while using ~8% of the wall-clock the stronger run needed and
    a tiny fraction of the overall runtime budget.
• Counter-example:
• Input:
    models = [LogisticRegression(max_iter=2000),
              MultinomialNB(alpha=0.5)]
    oof = np.zeros((X_tr.shape[0], n_classes)); preds = []
    for m in models:
        for tr, va in skf.split(X_tr, y):
            m.fit(X_tr[tr], y[tr])
            oof[va] += m.predict_proba(X_tr[va]) / len(models)
        m.fit(X_tr, y)
        preds.append(m.predict_proba(X_te))
    print("oof loss:", log_loss(y, oof))
    pd.DataFrame(np.mean(preds, axis=0), columns=classes).to_csv("submission.csv", index=False)
• Why it does not fire: the script instantiates more than one classifier, uses a fold-splitting loop to compute out-of-fold loss, and the written probabilities are an average across models, failing steps 1, 2, and 3.
C5 · budget-misuse Validation set supplied to boosting fit without early stoppingtrace_observed

P1 — Validation set supplied to boosting fit without early stopping

• Pattern: Detects a gradient-boosted model fitted with a validation evaluation set and a large fixed iteration budget but no early-stopping mechanism, so the validation scores are computed yet never used to halt training at the best iteration.
• Detection procedure:
1. Find a call to a fitting method of a gradient-boosting estimator (e.g. LightGBM/XGBoost/CatBoost fit or train, or equivalent) whose arguments include an evaluation-set parameter such as eval_set= (or equivalent) referencing data distinct from the training arguments.
2. In the estimator's construction or the fit call, find an iteration-count parameter (n_estimators=, num_boost_round=, iterations=, or equivalent) set to a literal integer of 300 or more.
3. Inspect the same fit/train call for any early-stopping configuration: a callbacks= list containing an early-stopping callback, an early_stopping_rounds= argument, an early_stopping= constructor flag, or equivalent. Also check the estimator constructor for such an argument. callbacks=None or absence of all of these counts as no early stopping.
4. PRESENT when the fit call passes an evaluation set (step 1), the iteration budget is a large fixed literal (step 2), and no early-stopping configuration exists (step 3).
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A fit call with eval_set=[(X_va, y_va)], eval_metric="auc", callbacks=None and n_estimators=1000 trained every fold to the full 1000 trees; the identical pipeline with an early-stopping callback of patience 100 scored 0.20026 higher on the held-out competition metric.
• Applies when: A gradient-boosted tree model is trained (typically inside a cross-validation loop) with a validation partition passed to the fit call and a fixed, large number of boosting rounds.
• Example:
• Input:
    for tr_idx, va_idx in skf.split(X, y):
        model = LGBMClassifier(
            n_estimators=1000,
            learning_rate=0.03,
        )
        model.fit(
            X[tr_idx], y[tr_idx],
            eval_set=[(X[va_idx], y[va_idx])],
            eval_metric="auc",
            callbacks=None,
        )
        oof[va_idx] = model.predict_proba(X[va_idx])[:, 1]
• Consequence:
    Every fold trains all 1000 boosting rounds regardless of validation
    performance; the model overfits past the validation optimum on noisy
    features, lowering held-out AUC (~0.20 lower in a measured pair) and
    spending several times the necessary training compute per fold.
• Counter-example:
• Input:
    for tr_idx, va_idx in skf.split(X, y):
        model = LGBMClassifier(
            n_estimators=3000,
            learning_rate=0.03,
        )
        model.fit(
            X[tr_idx], y[tr_idx],
            eval_set=[(X[va_idx], y[va_idx])],
            eval_metric="auc",
            callbacks=[lgb.early_stopping(100, verbose=False)],
        )
        oof[va_idx] = model.predict_proba(X[va_idx])[:, 1]
• Why it does not fire: The fit call includes an early-stopping callback, so the large iteration count is only an upper bound and training halts at the validation-optimal round.
C5 · budget-misuse Low vocabulary cap starving a sparse linear text model

P1 — Low vocabulary cap starving a sparse linear text model

• Pattern: Detects a sparse text-vectorization step whose vocabulary is explicitly truncated to a few thousand features before being fed to a linear or naive-Bayes classifier that scales cheaply to far larger sparse inputs.
• Detection procedure:
1. Locate a call constructing a text vectorizer that produces sparse token-count or weighted-count features (e.g. TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent).
2. In that constructor call, find a keyword argument that caps the vocabulary or output dimensionality (max_features= or n_features= or equivalent) with a literal integer value less than or equal to 10000.
3. Trace the variable assigned from that vectorizer's fit/transform output (or a pipeline containing the vectorizer) to a .fit(...) call on a model constructed from a linear classifier or multinomial/Bernoulli naive-Bayes class (e.g. LogisticRegression, LinearSVC, SGDClassifier, MultinomialNB, or equivalent), all within the same script or module.
4. Confirm there is no surrounding code that justifies the cap by consuming a dense form of the matrix (e.g. a .toarray() / .todense() call, or a downstream model that requires dense input) between the vectorizer output and the classifier fit.
5. The pattern is PRESENT when steps 1–4 all hold: a literal cap ≤ 10000 on the vectorizer feeds a sparse-capable linear or naive-Bayes model with no dense conversion requiring it.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A run using a vectorizer with a small literal max_features cap into a logistic-regression classifier scored 0.379 worse on the competition metric (log loss) than an uncapped sparse setup, and the capped run was not faster; the truncation removed informative rare terms while buying no resource savings.
• Applies when: A text-classification pipeline vectorizes raw text into sparse token features and trains a linear or naive-Bayes model directly on those features, with no memory constraint forcing dimensionality reduction.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(max_features=5000, stop_words='english')
    X_tr = vec.fit_transform(train['col_a'])
    X_te = vec.transform(test['col_a'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    preds = model.predict_proba(X_te)
• Consequence:
    Log loss on the held-out evaluation rises (e.g. ~0.55 vs ~0.35 for the
    uncapped vocabulary) because the linear model is denied tens of thousands
    of discriminative rare terms; wall-clock time is not reduced, so the
    truncation trades accuracy for nothing.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(min_df=2, max_df=0.95, stop_words='english')
    X_tr = vec.fit_transform(train['col_a'])
    X_te = vec.transform(test['col_a'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    preds = model.predict_proba(X_te)
• Why it does not fire: The vectorizer prunes vocabulary only by document-frequency thresholds (min_df/max_df) with no hard integer cap on feature count, so the sparse linear model retains the full informative vocabulary.
C5 · budget-misuse Small capped vocabulary as sole text representation

P1 — Small capped vocabulary as sole text representation

• Pattern: Detects a text-to-sparse-features transformation whose vocabulary is capped at a small fixed size (a literal in the low thousands) and that serves as the only text representation feeding a sparse linear or naive-Bayes model, with no character-level n-gram view alongside it.
• Detection procedure:
1. Locate every construction of a bag-of-words or tf-idf vectorizer (TfidfVectorizer, CountVectorizer, or equivalent text-to-sparse-matrix transformer) in the source file.
2. For each such construction, check whether the keyword argument max_features (or the equivalent vocabulary-size cap in another library) is set to a literal integer less than or equal to 10000.
3. Check whether any vectorizer construction in the same file passes analyzer='char' or analyzer='char_wb' (or an equivalent character-level tokenization option), or whether the outputs of two or more distinct vectorizers are combined (e.g., via hstack, FeatureUnion, or equivalent concatenation of sparse blocks).
4. Check that the matrix produced by the capped vectorizer is passed to the fit method of a linear classifier or naive-Bayes model (a model whose training cost scales cheaply with sparse feature count).
5. The pattern is PRESENT when step 2 holds for a vectorizer, step 3 finds no character-level or combined representation, and step 4 holds for that vectorizer's output.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using a single word-level tf-idf transform with max_features=5000 feeding a logistic regression scored 0.336 worse on the held-out metric than a variant using uncapped word plus character n-gram blocks; the weaker run used 221s versus 581s for the richer variant, both well under budget, confirming the cap was unnecessary.
• Applies when: The task is text classification with sparse features feeding a linear or naive-Bayes model, and the available compute budget comfortably permits feature spaces of 10^5–10^6 columns.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vectorizer = TfidfVectorizer(max_features=5000, ngram_range=(1, 2))
    X_tr = vectorizer.fit_transform(train_df["col_a"])
    X_te = vectorizer.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    preds = model.predict_proba(X_te)
• Consequence:
    Held-out log loss is ~0.19 absolute worse than the same model trained on
    an uncapped word + character n-gram representation; runtime (221s) was
    far below the allowed budget (richer variant fits in 581s), so the
    vocabulary restriction sacrificed accuracy for headroom that was unused.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack
    from sklearn.linear_model import LogisticRegression

    word_vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
    char_vec = TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5))
    X_tr = hstack([word_vec.fit_transform(train_df["col_a"]),
                   char_vec.fit_transform(train_df["col_a"])]).tocsr()
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
• Why it does not fire: No small literal vocabulary cap is set and a character-level n-gram block is concatenated with the word-level block, so both step 2 and step 3 fail.
C5 · budget-misuse Small vocabulary cap on sparse text features with an untuned single linear model

P1 — Small vocabulary cap on sparse text features with an untuned single linear model

• Pattern: Detects a bag-of-words or TF-IDF text feature pipeline whose vocabulary is truncated by an explicit feature-count limit of a few thousand terms, with the resulting features fed to a single linear classifier fitted once with default regularization and no hyperparameter search or model combination.
• Detection procedure:
1. Locate a construction of a sparse text vectorizer (TfidfVectorizer, CountVectorizer, or equivalent bag-of-words/TF-IDF extractor) whose constructor arguments include a keyword named max_features (or an equivalent vocabulary-size limit) assigned a literal integer less than or equal to 10000.
2. Trace the matrix produced by that vectorizer's fit/transform call to the training input of a classifier fit; confirm the classifier is a linear model (LogisticRegression, SGDClassifier, LinearSVC, or equivalent) constructed without an explicit regularization-strength argument (no C= or alpha= keyword with a non-default literal).
3. Within the same module, confirm there is no hyperparameter search construct (GridSearchCV, RandomizedSearchCV, a loop over regularization values with validation scoring, or equivalent) and no combination of predictions from more than one fitted model.
4. PRESENT when all three hold: a literal vocabulary cap ≤ 10000 on the text vectorizer, a single default-regularized linear classifier consuming its output, and no tuning or ensembling anywhere in the module.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using a vectorizer with max_features=5000 and one default LogisticRegression() fit reached a held-out log loss of ~0.58 in 196s of a much larger time budget; an otherwise similar pipeline with the full vocabulary, character n-grams, and tuned/blended models reached ~0.38 in 581s — a 33% relative metric degradation from the capped, untuned version.
• Applies when: The task is text classification scored on a probabilistic or accuracy-like metric, features come from bag-of-words/TF-IDF vectorization, and the run completes well inside the available compute budget.
• Example:
• Input:
    vectorizer = TfidfVectorizer(
        max_features=5000,
        ngram_range=(1, 2),
        stop_words='english'
    )
    X_tr = vectorizer.fit_transform(train_text)
    X_te = vectorizer.transform(test_text)

    model = LogisticRegression(max_iter=1000, random_state=42)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Consequence:
    Held-out log loss ~0.58 versus ~0.38 achievable with the full vocabulary
    and tuned regularization on the same data — a ~33% relative degradation —
    while using roughly one third of the available wall-clock budget.
• Counter-example:
• Input:
    vectorizer = TfidfVectorizer(
        max_features=200000,
        ngram_range=(1, 2),
        sublinear_tf=True
    )
    X_tr = vectorizer.fit_transform(train_text)
    X_te = vectorizer.transform(test_text)

    grid = GridSearchCV(LogisticRegression(max_iter=1000),
                        {'C': [0.1, 1.0, 10.0]}, scoring='neg_log_loss', cv=5)
    grid.fit(X_tr, y_train)
    proba = grid.predict_proba(X_te)
• Why it does not fire: the vocabulary limit is far above the few-thousand threshold and the regularization strength is selected by a validation search, so neither the capacity cap nor the untuned-single-model condition holds.
C5 · budget-misuse Linear head on a projection far narrower than the label spacegithub_occurrence

P1 — Linear head on a projection far narrower than the label space

• Pattern: Detects a multiclass pipeline with thousands of distinct labels where features are linearly projected down to a component count far below the label cardinality and the projected features are then fitted by a purely linear classifier, capping achievable class separation.
• Detection procedure:
1. Find evidence of label cardinality in the thousands: a literal constant assigned to a name containing CLASS, NUM, LABEL, or CAT with value greater than 1000; a call counting unique label values whose result is compared to or documented as exceeding 1000; or a comment stating the class count.
2. Within the same script, find a dimensionality-reduction fit — a call to PCA, TruncatedSVD, a randomized projection, or an equivalent construction of a projection matrix applied by matrix multiplication — whose component count is a literal integer (or a variable assigned a literal integer) less than one fifth of the class count found in step 1.
3. Find a classifier fitted on the output of the reduction from step 2 that is linear-only: LogisticRegression, a linear SVM, a single softmax/linear layer with no hidden layer, a nearest-centroid rule on the projected space, or equivalent — with no nonlinear transformation between the projection and the decision layer.
4. PRESENT when all of steps 1–3 hold: class count > 1000, projection width < one fifth of the class count, and the model consuming the projection is purely linear.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A pipeline projecting image features with a ~100-component decomposition (n_components=100) and fitting a linear decision rule over several thousand classes scored ~0.69 lower on the task metric than a same-budget model operating on the unreduced feature representation.
• Applies when: Multiclass classification with class count in the thousands, a dimensionality-reduction step precedes the classifier, and the compute budget visibly permits fitting on the wider representation (the full features are already materialized in the script).
• Example:
• Input:
    NUM_CLASSES = 5270          # thousands of labels
    pca = PCA(n_components=100)
    Xtr = pca.fit_transform(X)  # X is 4096-dim features
    clf = LogisticRegression(max_iter=200)
    clf.fit(Xtr, y)
    preds = clf.predict(pca.transform(X_test))
• Consequence:
    Held-out accuracy is substantially lower (observed gap ~0.69 on the task
    metric) than a same-budget linear model fitted on the unreduced features:
    a 100-dimensional linear decision space cannot separate 5000+ classes,
    so many classes collapse onto the same decision regions.
• Counter-example:
• Input:
    NUM_CLASSES = 5270
    pca = PCA(n_components=256)
    Xtr = pca.fit_transform(X)
    clf = MLPClassifier(hidden_layer_sizes=(1024,))
    clf.fit(Xtr, y)
    preds = clf.predict(pca.transform(X_test))
• Why it does not fire: the model consuming the projection has a nonlinear hidden layer, so step 3's requirement of a purely linear classifier on the reduced features is not met.
C5 · budget-misuse Tiny boosting budget masked by inflated learning rate

P1 — Tiny boosting budget masked by inflated learning rate

• Pattern: Detects a gradient-boosted tree ensemble constructed with a very small number of boosting rounds combined with a learning rate raised several-fold above the library default, trading convergence for speed instead of allowing enough rounds at a moderate step size.
• Detection procedure:
1. Locate every constructor call or parameter dictionary that configures a gradient-boosted tree model (e.g., LGBMClassifier, XGBRegressor, GradientBoostingClassifier, CatBoostClassifier, or equivalent, including their train/params-dict APIs).
2. Within that call or dictionary, find the boosting-round parameter (n_estimators, num_boost_round, iterations, or equivalent) assigned a literal integer value less than or equal to 25.
3. Within the same call or dictionary, find the learning-rate parameter (learning_rate, eta, or equivalent) assigned a literal numeric value greater than or equal to 0.3.
4. Verify that no early-stopping mechanism (early_stopping_rounds, an early-stopping callback, or equivalent) is passed to the same model's fitting call within the same script.
5. The pattern is PRESENT when steps 2, 3, and 4 all hold for the same model configuration.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A model configured with n_estimators=10, learning_rate=0.5 scored 0.10421 lower on the evaluation metric than an otherwise comparable model using n_estimators=100, learning_rate=0.1 on the same data; the tiny-budget model's additive ensemble stops far before the loss curve flattens.
• Applies when: The script trains a gradient-boosted tree model with explicitly specified boosting-round and learning-rate hyperparameters, and total runtime is not so constrained that fewer than ~25 rounds is unavoidable.
• Example:
• Input:
    import lightgbm as lgb

    model = lgb.LGBMClassifier(
        n_estimators=10,
        learning_rate=0.5,
        max_depth=3,
        random_state=42,
    )
    model.fit(X, y)
    preds = model.predict(X_test)
• Consequence:
    Training loss plateaus after only a handful of coarse, high-step updates;
    held-out accuracy lands roughly 0.10 below what the same wall-clock budget
    achieves with learning_rate=0.1 and n_estimators=100.
• Counter-example:
• Input:
    import lightgbm as lgb

    model = lgb.LGBMClassifier(
        n_estimators=100,
        learning_rate=0.1,
        max_depth=3,
        random_state=42,
    )
    model.fit(X, y)
    preds = model.predict(X_test)
• Why it does not fire: The boosting-round count is well above 25 and the learning rate matches the library default, so neither threshold in steps 2 and 3 is met.
C5 · budget-misuse Small hardcoded vocabulary cap on sparse text features for a linear model

P1 — Small hardcoded vocabulary cap on sparse text features for a linear model

• Pattern: Detects a bag-of-words or TF-IDF text vectorizer constructed with a small hardcoded vocabulary-size cap (a literal of a few thousand or less) whose output feeds a sparse-capable linear classifier, discarding most of the n-gram feature space with no evidence the cap was tuned or validated.
• Detection procedure:
1. Locate a construction of a text vectorizer that produces sparse token-count or TF-IDF features (TfidfVectorizer, CountVectorizer, or equivalent) whose keyword arguments include a vocabulary-size cap (max_features= or equivalent) set to a literal integer less than or equal to 20000.
2. Confirm the matrix returned by that vectorizer's fit/transform call is, within the same script, passed to the fitting method of a linear or naive-Bayes-style classifier that accepts sparse input (LogisticRegression, SGDClassifier, LinearSVC, MultinomialNB, or equivalent).
3. Confirm the literal cap value is not chosen by any search or comparison: it does not appear inside a hyperparameter-search construct, a loop over candidate values, or a conditional comparing held-out metric scores.
4. PRESENT when all of steps 1–3 hold: a small literal vocabulary cap feeds a sparse linear model with no tuning or validation of the cap.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A vectorizer built with max_features=5000, ngram_range=(1, 2) feeding a linear classifier produced a scored probabilistic metric roughly 0.44 worse (log loss ~0.54 vs ~0.30) than a pipeline using the uncapped vocabulary, while runtime remained a small fraction of the available budget (~6s vs ~64s).
• Applies when: A text-classification script vectorizes raw text into sparse bag-of-words or TF-IDF features and trains a linear or naive-Bayes classifier on them, with no runtime or memory constraint that would justify truncating the vocabulary.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    tfidf = TfidfVectorizer(max_features=5000, ngram_range=(1, 2), min_df=2)
    X_tr = tfidf.fit_transform(train_df["col_a"])
    X_te = tfidf.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    probs = model.predict_proba(X_te)
• Consequence:
    Held-out log loss rises from ~0.30 (full vocabulary) to ~0.54 with the
    5000-feature cap; most discriminative rare n-grams are discarded while
    total runtime stays far below the available budget (~6s vs ~64s used
    by the uncapped pipeline).
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    tfidf = TfidfVectorizer(ngram_range=(1, 2), min_df=2, sublinear_tf=True)
    X_tr = tfidf.fit_transform(train_df["col_a"])
    X_te = tfidf.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    probs = model.predict_proba(X_te)
• Why it does not fire: The vectorizer sets no vocabulary-size cap, so the sparse linear model trains on the full n-gram feature space; only mild frequency pruning (min_df=2) is applied.
C5 · budget-misuse Validation score computed but never used to select anythinggithub_occurrence

P1 — Validation score computed but never used to select anything

• Pattern: Detects a pipeline that splits off a validation set and evaluates a single fixed-hyperparameter model on it, then refits and predicts with those same hardcoded hyperparameters, so the validation metric influences no hyperparameter choice.
• Detection procedure:
1. Find a call that partitions training data into a fit portion and a held-out portion (e.g. train_test_split or equivalent index slicing) within the script.
2. Find a model constructor whose hyperparameter arguments are all literal constants (numbers, strings) or defaults — no argument is a variable assigned inside a loop or from a comparison of scores.
3. Confirm the held-out portion is passed to a metric function (e.g. log_loss, accuracy_score, or equivalent) whose result is only printed, logged, or discarded — the returned value is never compared against another score, never used in a conditional, and never determines any argument of a later model constructor or fit call.
4. Confirm there is exactly one model-family constructor invocation per fit stage (no loop or list comprehension iterating over candidate hyperparameter values that each get fitted and scored).
5. Confirm a final model is fitted (on the full or fit portion of the data) with the same literal hyperparameters from step 2 and used to produce the output predictions.
6. PRESENT when steps 1–5 all hold: a validation split and metric exist, but exactly one hardcoded-hyperparameter configuration is ever fitted and the validation score selects nothing.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A pipeline computed a held-out score for a single linear classifier built as LogisticRegression(max_iter=2000, C=1.0) and then refit that same fixed configuration for the final predictions; an otherwise-identical solution that looped over a small grid of the regularization value on the same split achieved roughly a 6% relative improvement in the task metric.
• Applies when: The script creates a validation split and reports a metric on it, fits a supervised model whose key capacity/regularization hyperparameter is set by a literal or left at default, and produces final predictions from a single configuration.
• Example:
• Input:
    Xa, Xb, ya, yb = train_test_split(X, y, test_size=0.15, random_state=42)
    clf = LogisticRegression(max_iter=2000, C=1.0)
    clf.fit(Xa, ya)
    val_ll = log_loss(yb, clf.predict_proba(Xb)[:, 1])
    print("val logloss:", val_ll)

    clf_full = LogisticRegression(max_iter=2000, C=1.0)
    clf_full.fit(X, y)
    preds = clf_full.predict_proba(X_test)[:, 1]
    pd.DataFrame({"id": ids, "label": preds}).to_csv("out.csv", index=False)
• Consequence:
    Final held-out log loss is systematically higher than a run that tried a
    small grid of C values on the same validation split; observed gap ~6%
    relative log-loss deficit with the untuned default regularization.
• Counter-example:
• Input:
    Xa, Xb, ya, yb = train_test_split(X, y, test_size=0.15, random_state=42)
    best_C, best_ll = None, None
    for C in [0.003, 0.01, 0.03, 0.1, 1.0]:
        clf = LogisticRegression(max_iter=500, C=C)
        clf.fit(Xa, ya)
        ll = log_loss(yb, clf.predict_proba(Xb)[:, 1])
        if best_ll is None or ll < best_ll:
            best_ll, best_C = ll, C
    clf_full = LogisticRegression(max_iter=500, C=best_C)
    clf_full.fit(X, y)
    preds = clf_full.predict_proba(X_test)[:, 1]
• Why it does not fire: The validation metric is compared across candidate hyperparameter values inside a loop and the winning value (best_C, a variable, not a literal) parameterizes the final refit, so the score selects the configuration.
C5 · budget-misuse Small hard vocabulary cap with untuned classifier regularization

P1 — Small hard vocabulary cap with untuned classifier regularization

• Pattern: Detects a text-feature vectorizer constructed with a small hard vocabulary cap (a literal max_features of a few thousand or less) whose output feeds a linear probabilistic classifier left at default regularization strength, with no vocabulary-size justification such as frequency-based filtering or hyperparameter search.
• Detection procedure:
1. Locate a call that constructs a bag-of-words or TF-IDF style text vectorizer (e.g. TfidfVectorizer(...), CountVectorizer(...), or equivalent) whose arguments include max_features= followed by a literal integer less than or equal to 10000.
2. Confirm the object from step 1 has its fit_transform (or fit then transform, or equivalent) called on a text column of a training frame within the same script.
3. Locate a linear or naive-Bayes probabilistic classifier constructor (e.g. LogisticRegression(...), SGDClassifier(...), or equivalent) fitted on the matrix produced in step 2, whose argument list contains no regularization-strength keyword (C=, alpha=, or equivalent) and is not wrapped in any search object (GridSearchCV, RandomizedSearchCV, or equivalent) or cross-validated variant (e.g. a class name ending in CV).
4. Confirm the script contains no min_df= or max_df= argument on the vectorizer from step 1 that would substitute frequency-based filtering for the hard cap.
5. PRESENT when all of steps 1–4 hold: a literal cap ≤ 10000 on the vocabulary, features fitted from training text, and a downstream probabilistic classifier with default, unsearched regularization.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A pipeline using TfidfVectorizer(max_features=5000) into a default-regularization logistic model scored ~0.38 worse (relative ~38% higher log loss) on a probabilistic metric than an otherwise-similar pipeline that removed the cap, added character n-grams, and set a tuned regularization strength; both finished well within the runtime budget.
• Applies when: Text-classification code building sparse bag-of-words/TF-IDF features for a probability-scored metric, where the training corpus vocabulary plausibly exceeds the cap and runtime is not the binding constraint.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(max_features=5000, ngram_range=(1, 2))
    X_tr = vec.fit_transform(train_df["col_a"])
    X_te = vec.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000, random_state=42)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Consequence:
    Log loss on the held-out set is materially higher (~38% relative worse in the
    measured pair) than an identical pipeline with the vocabulary cap removed and
    regularization strength tuned; wall-clock time stays far under the available
    budget, so the discarded capacity buys nothing.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(min_df=3, sublinear_tf=True, ngram_range=(1, 2))
    X_tr = vec.fit_transform(train_df["col_a"])
    X_te = vec.transform(test_df["col_a"])

    model = LogisticRegression(C=4, max_iter=1000)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Why it does not fire: There is no literal max_features cap — vocabulary size is controlled by frequency filtering (min_df) — and the classifier's regularization strength is explicitly set rather than left at its default.
C5 · budget-misuse Starved boosting budget without validation-based sizing

P1 — Starved boosting budget without validation-based sizing

• Pattern: Detects a boosted tree ensemble fitted on a large training set with a tiny iteration count (at or below roughly 20) combined with severely capped tree complexity, and no early-stopping or validation-based selection of the iteration count.
• Detection procedure:
1. Locate a constructor call for a gradient-boosted tree model (LGBMClassifier, XGBClassifier, GradientBoostingClassifier, CatBoostClassifier, their regressor variants, or equivalent) whose result is later passed a call to a fitting method (fit, train, or equivalent) in the same script.
2. In that constructor's keyword arguments, check that the boosting-round parameter (n_estimators, num_boost_round, iterations, or equivalent) is a literal integer less than or equal to 20.
3. In the same constructor, check that tree complexity is also capped low: max_depth is a literal integer less than or equal to 4, or num_leaves (or equivalent leaf-count parameter) is a literal integer less than or equal to 15.
4. Confirm that neither the constructor nor the fitting call includes an early-stopping mechanism (early_stopping_rounds, an early-stopping callback, eval_set used for stopping, or equivalent), and that the script contains no loop or search that selects the round count from a held-out score.
5. The pattern is PRESENT when steps 1–4 all hold: a boosted ensemble is fitted with both a tiny fixed round budget and capped tree complexity, with no validation-driven selection of that budget.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: A model constructed with n_estimators=10, max_depth=3, num_leaves=7 and fitted directly on the full training frame scored ~9.6% relative worse on the held-out metric (0.854 vs 0.945) than a run using ~90 rounds with ~48 leaves on the same data, with runtime well within budget in both cases.
• Applies when: A supervised model is trained on a sizable tabular dataset with a gradient-boosted tree library, and the run is not otherwise shown to be constrained to seconds of compute.
• Example:
• Input:
    import lightgbm as lgb
    model = lgb.LGBMClassifier(
        n_estimators=10,
        learning_rate=0.5,
        max_depth=3,
        num_leaves=7,
        random_state=42,
    )
    model.fit(X, y)
    preds = model.predict(X_test)
• Consequence:
    Held-out accuracy drops materially (observed ~0.09 absolute, ~9.6%
    relative) compared with a moderately sized ensemble trained on the
    same data in a similar time budget; the model systematically
    underfits and predictions cluster on majority classes.
• Counter-example:
• Input:
    import lightgbm as lgb
    model = lgb.LGBMClassifier(
        n_estimators=500,
        learning_rate=0.1,
        num_leaves=48,
        random_state=42,
    )
    model.fit(X_tr, y_tr, eval_set=[(X_val, y_val)],
              callbacks=[lgb.early_stopping(30)])
    preds = model.predict(X_test)
• Why it does not fire: The round budget is not a tiny literal and an early-stopping callback on a validation split selects the effective iteration count, so steps 2 and 4 both fail.
C5 · budget-misuse Default regularization on high-dimensional probabilistic linear modelverified_trace · effect +2.4138

P1 — Default regularization on high-dimensional probabilistic linear model

• Pattern: Detects a probabilistic linear classifier fitted on high-dimensional handcrafted features with the library-default regularization strength, its probability outputs written directly to a submission scored by a metric that penalizes overconfident errors.
• Detection procedure:
1. Locate a constructor call for a linear probabilistic classifier (e.g. LogisticRegression or equivalent logistic/softmax linear model) whose argument list contains no regularization-strength argument (no C=, alpha=, or equivalent penalty-strength keyword).
2. Confirm the object from step 1 is fitted, within the same script, on a feature matrix produced by a handcrafted feature extractor (e.g. gradient-histogram, color-histogram, or other manual descriptor functions, or a matrix whose second dimension is set by a feature-length constant in the thousands), not by a learned embedding.
3. Confirm the fitted object's probability-output method (e.g. predict_proba or equivalent) is called and its result is written to an output file scored on probabilities (columns of per-class probabilities saved via to_csv or equivalent).
4. Confirm there is no hyperparameter search over regularization strength anywhere in the script (no loop or search utility varying C/alpha/penalty over multiple values with held-out evaluation).
5. PRESENT when steps 1–4 all hold: default-strength regularization, handcrafted high-dimensional features, probability outputs submitted, and no regularization tuning.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 3
• weight: 3
• confidence: high
• Evidence: clf = LogisticRegression(max_iter=1000) fitted on thousands of handcrafted descriptor features with default C=1.0, probabilities written to the submission; log loss was 11.63 versus 4.55 for the identical pipeline with C=0.01 plus mild uniform smoothing — a 0.61 gap on the competition metric from regularization strength alone.
• Applies when: The task is scored on predicted class probabilities (log loss or similar), the feature matrix is high-dimensional relative to samples per class, and the model is a linear probabilistic classifier trained in a single pass without hyperparameter selection.
• Example:
• Input:
    from sklearn.linear_model import LogisticRegression
    from sklearn.preprocessing import StandardScaler

    # X: (n_samples, 8000) handcrafted descriptor features
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    X_test_scaled = scaler.transform(X_test)

    clf = LogisticRegression(max_iter=1000)
    clf.fit(X_scaled, y)
    probs = clf.predict_proba(X_test_scaled)
    pd.DataFrame(probs, columns=classes).to_csv("submission.csv", index=False)
• Consequence:
    The near-interpolating fit emits sharply peaked probability vectors that are
    frequently wrong; log loss rises to ~11.6 versus ~4.6 for the same features
    with strong regularization — a 0.6+ degradation on the probability metric.
• Counter-example:
• Input:
    from sklearn.linear_model import LogisticRegression
    from sklearn.preprocessing import StandardScaler

    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    X_test_scaled = scaler.transform(X_test)

    clf = LogisticRegression(C=0.01, max_iter=1000)
    clf.fit(X_scaled, y)
    probs = clf.predict_proba(X_test_scaled)
    probs = 0.9 * probs + 0.1 / probs.shape[1]
    pd.DataFrame(probs, columns=classes).to_csv("submission.csv", index=False)
• Why it does not fire: the constructor explicitly sets a strong regularization strength (C=0.01), so step 1's condition of a missing penalty-strength argument fails.
C5 · budget-misuse Global-cap subsampling without per-class quota in extreme multiclassgithub_occurrence

P1 — Global-cap subsampling without per-class quota in extreme multiclass

• Pattern: Detects training-set subsampling that stops at a fixed total example count (or uniform stride) with no per-class counter or quota, in a classification task whose label space contains thousands of distinct classes.
• Detection procedure:
1. Locate a loop or slice that selects training examples and terminates or filters based on a single scalar budget: a comparison of one running counter against a literal cap (e.g. if n_kept >= MAX_TRAIN: break), a stride condition (e.g. if i % STRIDE != 0: continue), or a head slice such as df[:N] / itertools.islice(..., N) applied to the training source.
2. Confirm the task is extreme multiclass: the label column or class list feeds a classifier fit, and the code or comments indicate class cardinality in the thousands (e.g. n_classes derived from unique labels of a large catalog, or classes= passed with a collection built from thousands of unique values), or the cap-to-class ratio implied by literals in the file is below roughly 20 examples per class.
3. Verify that within the selection loop or slice there is no dictionary, Counter, defaultdict, or array indexed by the example's label that gates whether the example is kept (i.e. no keep/skip decision reads a per-label count), and no stratified splitting utility (train_test_split(..., stratify=...), StratifiedShuffleSplit, or equivalent) is applied to produce the subsample.
4. PRESENT when step 1's scalar-budget selection exists, step 2's extreme class cardinality holds, and step 3's absence of any per-class quota or stratified accumulation is confirmed.
• Predicted impact:
• Add score for C5 (Capacity or training budget left on the table): 2
• weight: 2
• confidence: high
• Evidence: A training loader using if kept >= CAP: break with a uniform stride over a streamed file, no per-label counter anywhere in the selection loop, left most of thousands of classes with single-digit example counts; the resulting model scored roughly half the competition metric of an otherwise identical pipeline that kept examples until each class reached a small quota (0.048 vs 0.093).
• Applies when: The label space has thousands of classes, the full training set is too large to use entirely (streamed or capped), and the code subsamples training examples before fitting.
• Example:
• Input:
    CAP = 60000
    X, y = [], []
    kept = 0
    for rec in stream_records(TRAIN_FILE):
        if kept >= CAP:
            break
        X.append(featurize(rec["img"]))
        y.append(rec["label"])       # ~5000 distinct labels overall
        kept += 1
    model.fit(np.array(X), np.array(y))
• Consequence:
    Average examples per class falls to ~12, and because the stream is not
    label-balanced most classes receive 0-3 examples; held-out accuracy is
    roughly half that of the same pipeline with a per-class quota
    (e.g. 0.048 vs 0.093 on the evaluation metric).
• Counter-example:
• Input:
    QUOTA, CAP = 40, 200000
    counts = {}
    X, y, kept = [], [], 0
    for rec in stream_records(TRAIN_FILE):
        if kept >= CAP:
            break
        lbl = rec["label"]
        if counts.get(lbl, 0) >= QUOTA:
            continue
        counts[lbl] = counts.get(lbl, 0) + 1
        X.append(featurize(rec["img"])); y.append(lbl); kept += 1
    model.fit(np.array(X), np.array(y))
• Why it does not fire: although a scalar cap is present, the keep/skip decision reads a per-label counter (counts) that enforces a quota per class, so step 3's condition fails.
C5 · budget-misuse Boosting fit with eval_set but no early stoppingtrace_observed

P1 — Boosting fit with eval_set but no early stopping

• Pattern: Detects a gradient-boosted ensemble fitted with a large fixed boosting-round count and a supplied validation evaluation set, but with no early-stopping mechanism registered, so training always runs to the full round count regardless of validation performance.
• Detection procedure:
1. Locate a call to a fitting method (e.g. fit) on a gradient-boosting model object (LGBMClassifier, XGBClassifier, CatBoostClassifier, or equivalent) whose constructor or params dict, within the same file, sets the boosting-round count (n_estimators, num_boost_round, iterations, or equivalent) to a literal integer of 300 or more.
2. Confirm the fitting call at step 1 receives an evaluation split via eval_set (or equivalent keyword accepting held-out data for monitoring).
3. Confirm the fitting call and the constructor/params from step 1 contain no early-stopping registration: no early_stopping_rounds keyword, no callback list containing an early-stopping callback (e.g. early_stopping(...) or equivalent), and no early_stopping constructor parameter; a callbacks=None argument or absent callbacks counts as no registration.
4. PRESENT when all of steps 1–3 hold for the same fitting call.
• Predicted impact:
• Add score for C5: Capacity or training budget left on the table: 2
• weight: 2
• confidence: high
• Evidence: model.fit(X_tr, y_tr, eval_set=[(X_va, y_va)], eval_metric="auc", callbacks=None) with n_estimators=1000 trained every fold to the full 1000 rounds; a comparable run that registered an early-stopping callback on the same validation metric scored 0.28 higher on the held-out ranking metric.
• Applies when: A gradient-boosting library model is trained on tabular or extracted features with a monitored validation split available, and the round count is a fixed literal rather than tuned per fold.
• Example:
• Input:
    model = LGBMClassifier(
        n_estimators=1000,
        learning_rate=0.03,
        num_leaves=31,
    )
    model.fit(
        X_tr, y_tr,
        eval_set=[(X_va, y_va)],
        eval_metric="auc",
        callbacks=None,
    )
• Consequence:
    Every fold boosts all 1000 rounds past the validation optimum; held-out
    AUC drops (observed ~0.28 lower) relative to stopping at the best
    iteration, and training wall-clock roughly doubles.
• Counter-example:
• Input:
    model = LGBMClassifier(
        n_estimators=1000,
        learning_rate=0.03,
        num_leaves=31,
    )
    model.fit(
        X_tr, y_tr,
        eval_set=[(X_va, y_va)],
        eval_metric="auc",
        callbacks=[lgb.early_stopping(60, verbose=False)],
    )
• Why it does not fire: An early-stopping callback keyed to the monitored validation metric is registered, so the large round count is only an upper bound and step 3 fails.
C6 · wasted-compute Per-row model inference inside a Python loopgithub_occurrence

P1 — Per-row model inference inside a Python loop

• Pattern: Detects a fitted model's prediction call (and any preprocessing transform feeding it) being invoked on single-row inputs inside a Python loop over evaluation samples, instead of one batched call on a stacked feature matrix.
• Detection procedure:
1. Locate a for loop whose iterable is a collection of test/evaluation items (e.g., a list of file names, rows of a frame, or range(len(...)) over the test set).
2. Inside that loop body, find a call to a fitted estimator's prediction method (predict, predict_proba, or equivalent) whose argument is constructed from a single item of the current iteration — e.g., a one-element list literal like [features], a reshape(1, -1), or a single-row slice.
3. Optionally, also inside the loop body, find a call to a fitted transformer's transform (or equivalent) on the same single-item input before the prediction call.
4. Confirm the loop appends or stores each per-item prediction into a growing collection (append, indexed assignment) and that no batched prediction call on the full stacked matrix exists for the same test set elsewhere.
5. PRESENT when steps 1, 2, and 4 all hold: prediction is executed once per sample inside the loop with no full-matrix batched alternative.
• Predicted impact:
• Add score for C6 (Avoidably expensive computation): 2
• weight: 2
• confidence: high
• Evidence: A per-sample loop calling scaler.transform([features]) then clf.predict_proba(features_scaled)[0][1] for each test item took ~2.2x longer overall than a run that stacked all features and predicted once, despite the loop-based run processing ~8x fewer items.
• Applies when: The code performs inference over a test or evaluation set whose features can be assembled into a single 2D matrix, and the estimator supports batched prediction.
• Example:
• Input:
    predictions = []
    for fname in test_files:
        features = extract_features(os.path.join(test_dir, fname))
        features_scaled = scaler.transform([features])
        pred = model.predict_proba(features_scaled)[0][1]
        predictions.append((fname, pred))
    sub = pd.DataFrame(predictions, columns=["Id", "Label"])
    sub.to_csv(output_path, index=False)
• Consequence:
    Inference wall-clock time grows by roughly an order of magnitude versus a
    single batched transform/predict on the stacked matrix, because per-call
    overhead (validation, dispatch, memory allocation) is paid once per sample
    instead of once per dataset; total runtime increases with no accuracy gain.
• Counter-example:
• Input:
    rows = []
    for fname in test_files:
        rows.append(extract_features(os.path.join(test_dir, fname)))
    X_test = np.vstack(rows)
    X_test = scaler.transform(X_test)
    preds = model.predict_proba(X_test)[:, 1]
    sub = pd.DataFrame({"Id": test_files, "Label": preds})
    sub.to_csv(output_path, index=False)
• Why it does not fire: The loop only accumulates raw features; the transform and prediction calls occur once outside the loop on the full stacked matrix, so no per-sample inference call exists.
C6 · wasted-compute Per-file serial extraction with per-sample inference calls

P1 — Per-file serial extraction with per-sample inference calls

• Pattern: Detects a loop over many independent input files that extracts features and invokes model inference on each single feature vector inside the loop body, instead of parallelizing extraction and predicting once on the stacked feature matrix.
• Detection procedure:
1. Locate a for loop (or list comprehension) that iterates over a collection of file names or file paths, e.g. produced by os.listdir, glob.glob, Path.iterdir, or equivalent.
2. Within the loop body, find a call to a user-defined or library function whose argument includes the loop variable (the file path/name) and whose result is used as a feature vector.
3. Within the same loop body, find a call to a model prediction method (predict, predict_proba, or equivalent) whose argument is that single feature vector, typically wrapped in a one-element list or a 2-D array with one row (e.g. [features] or features.reshape(1, -1)).
4. Confirm the loop is not dispatched through a process or thread pool (no ProcessPoolExecutor, multiprocessing.Pool, joblib.Parallel, or equivalent wrapping the per-file work).
5. PRESENT when steps 1–4 all hold: per-file feature extraction and per-sample inference both occur serially inside the same loop over independent files.
• Predicted impact:
• Add score for C6: Avoidably expensive computation: 3
• weight: 3
• confidence: high
• Evidence: A serial loop calling features = extract_features(path) followed by clf.predict_proba(scaler.transform([features])) per file took ~806s versus ~108s for a pooled-extraction, single-batched-predict variant while the slow version processed 2.3x fewer images, shrinking the training set achievable within the compute budget.
• Applies when: The program processes many (thousands or more) independent input files, each requiring a per-file feature extraction step before model inference.
• Example:
• Input:
    test_files = sorted(f for f in os.listdir(test_dir) if f.endswith('.jpg'))
    predictions = []
    for fname in test_files:
        features = extract_features(os.path.join(test_dir, fname))
        features_scaled = scaler.transform([features])
        pred = model.predict_proba(features_scaled)[0][1]
        predictions.append((fname, pred))
    sub = pd.DataFrame(predictions, columns=['id', 'score'])
    sub.to_csv('submission.csv', index=False)
• Consequence:
    Wall-clock time for extraction plus inference rises roughly 7x (e.g. 806s
    vs 108s on the same file set); within a fixed compute budget, ~2.3x fewer
    files can be processed, shrinking the usable training/evaluation set and
    degrading the final metric.
• Counter-example:
• Input:
    test_files = sorted(f for f in os.listdir(test_dir) if f.endswith('.jpg'))
    paths = [os.path.join(test_dir, f) for f in test_files]
    with ProcessPoolExecutor() as ex:
        feats = list(ex.map(extract_features, paths))
    X = scaler.transform(np.vstack(feats))
    preds = model.predict_proba(X)[:, 1]
    sub = pd.DataFrame({'id': test_files, 'score': preds})
    sub.to_csv('submission.csv', index=False)
• Why it does not fire: Feature extraction is mapped over files with a process pool and inference is a single call on the stacked matrix, so no per-file serial loop contains a per-sample prediction call.
C6 · wasted-compute Full-resolution array materialization before downscale in per-item image loopgithub_occurrence

P1 — Full-resolution array materialization before downscale in per-item image loop

• Pattern: Detects a loop over dataset records that decodes each image and converts it to a full-resolution numeric array before any downscaling, so every iteration materializes a large array whose contents are immediately reduced or discarded.
• Detection procedure:
1. Locate a for or while loop whose body decodes an image from bytes or a file per iteration (e.g., Image.open(...), cv2.imdecode(...), or equivalent), where the iterable is a dataset stream, file list, or record collection.
2. Within the same loop body, find an array-conversion call applied to the decoded image object (e.g., np.array(img), np.asarray(img), img_to_array(img), or equivalent) that occurs textually before any size-reduction call on the image object itself (e.g., .resize(...), .thumbnail(...), .draft(...), or equivalent).
3. Confirm that either (a) a size-reduction operation is applied to the resulting array afterwards (e.g., an array-based resize, slicing to a smaller shape, or equivalent), or (b) the loop stores only a reduced/derived version of the array, never the full-resolution array.
4. The pattern is PRESENT when steps 1–3 all hold: a per-item decode loop converts each image to a full-resolution array first, and only a downscaled derivative of that array is retained.
• Predicted impact:
• Add score for C6 (Avoidably expensive computation): 2
• weight: 2
• confidence: high
• Evidence: A per-record loop calling np.array(Image.open(io.BytesIO(...))) before resizing caused the run to take ~50% longer wall-clock while processing ~140x fewer training records than a variant that downscaled during decode, producing a large deficit on the evaluation metric.
• Applies when: The program iterates over a large collection of images (thousands to millions of records) and derives small fixed-size features or inputs from each image inside the loop.
• Example:
• Input:
    feats = []
    for rec in records:
        img = Image.open(io.BytesIO(rec["picture"]))
        arr = np.array(img)                    # full-resolution array
        small = skimage.transform.resize(arr, (32, 32))
        feats.append(small.flatten())
• Consequence:
    Per-item processing time rises severalfold because each iteration allocates
    and fills a full-resolution array that is immediately thrown away; within a
    fixed runtime budget the loop covers orders of magnitude fewer records
    (observed ~140x fewer training examples at ~50% greater wall-clock),
    lowering the final evaluation metric.
• Counter-example:
• Input:
    feats = []
    for rec in records:
        img = Image.open(io.BytesIO(rec["picture"]))
        img.draft("JPEG", (32, 32))
        img = img.resize((32, 32))
        arr = np.asarray(img)                  # only the small array exists
        feats.append(arr.flatten())
• Why it does not fire: The image object is downscaled (via draft-mode decode and resize) before array conversion, so no full-resolution array is ever materialized inside the loop.
C6 · wasted-compute Sequential per-file feature extraction over a large file corpusverified_trace · effect +0.0758

P1 — Sequential per-file feature extraction over a large file corpus

• Pattern: Detects a single-process sequential loop that reads and computes features from each file in a large collection of independent input files, with no process-pool or worker-based parallel map anywhere in the extraction stage.
• Detection procedure:
1. Locate a loop (a for statement, or a list/generator comprehension) that iterates over a collection of file identifiers or paths built by directory listing or glob (e.g. os.listdir, glob.glob, Path.iterdir, or equivalent), where the collection plausibly contains many entries (built from a directory scan rather than a small literal list).
2. Confirm that inside that loop body (or in a function called from it) each iteration loads one file (e.g. np.load, open, an image/audio read call, or equivalent) and computes per-file values that are appended or assigned into an accumulating structure, with no dependency between iterations (no iteration reads a value produced by a prior iteration other than the accumulator append).
3. Search the entire module for any construct that distributes this per-file work across processes or workers: ProcessPoolExecutor, multiprocessing.Pool, joblib.Parallel, a map submitted to an executor, or equivalent. None is applied to this extraction loop or its per-file function.
4. PRESENT if steps 1–2 hold for the extraction loop and step 3 finds no parallel-execution construct applied to it.
• Predicted impact:
• Add score for C6 (Avoidably expensive computation): 2
• weight: 2
• confidence: high
• Evidence: A sequential for file in sorted(os.listdir(subdir_path)) loop calling np.load and per-file statistics one file at a time took ~1437s for the extraction stage, versus ~66s total for an otherwise comparable pipeline that mapped the same per-file function with a process pool; the wasted wall-clock budget precluded further data processing and model search, coinciding with a 0.23 metric gap.
• Applies when: The program extracts features from many (thousands or more) independent files on a multi-core machine, and each file's computation does not depend on any other file's result.
• Example:
• Input:
    files = [f for f in os.listdir(data_dir) if f.endswith('.npy')]
    features = []
    for fname in files:
        arr = np.load(os.path.join(data_dir, fname))
        features.append([arr.mean(), arr.std(), arr.max(), arr.min()])
    X = np.array(features)
• Consequence:
    Feature extraction runs on a single core; wall-clock time for this stage
    grows ~Nx relative to an N-worker pool (e.g. ~1400s instead of ~90s on a
    16-core host), consuming the run budget that would otherwise fund richer
    features, more training rounds, or hyperparameter search.
• Counter-example:
• Input:
    def feats(path):
        arr = np.load(path)
        return [arr.mean(), arr.std(), arr.max(), arr.min()]

    files = sorted(glob.glob(os.path.join(data_dir, '*.npy')))
    with ProcessPoolExecutor(max_workers=os.cpu_count()) as ex:
        features = list(ex.map(feats, files, chunksize=64))
    X = np.array(features)
• Why it does not fire: the per-file feature function is mapped over the file list with a process pool, so step 3 finds a parallel-execution construct applied to the extraction work.
C6 · wasted-compute Raw-file re-parsing inside the epoch loopgithub_occurrence

P1 — Raw-file re-parsing inside the epoch loop

• Pattern: Detects a multi-epoch training loop whose body re-opens or re-iterates a large raw data file and re-runs per-record decoding/feature extraction on every epoch, instead of iterating over a feature array extracted once before the loop.
• Detection procedure:
1. Locate a loop whose iteration variable or range is derived from an epoch count (e.g., for epoch in range(N) or a variable whose name contains epoch, or a documented pass-over-data loop with a fixed pass count greater than 1).
2. Inside that loop's body (including functions called from it), find a call that opens or streams a data file — open(...), a generator that reads from a file path, or an equivalent record-iterating reader — where the file path argument refers to the training data file, not a cached artifact.
3. Within the same loop body, confirm that per-record transformation work occurs on the streamed records (e.g., byte decoding, image decoding, string parsing, or a feature-extraction function applied to each record) before those features feed an incremental fit/update call.
4. Confirm there is no statement before the epoch loop that materializes the full training feature set into an in-memory or on-disk cached array (e.g., an array built once by a single pass and then indexed inside the loop).
5. PRESENT when steps 1–4 all hold: the epoch loop re-streams and re-decodes the raw training file on every pass with no pre-extracted cached feature matrix.
• Predicted impact:
• Add score for C6: 3
• weight: 3
• confidence: high
• Evidence: A training script placed for doc in iter_records(TRAIN_PATH) with per-record extract_features(...) inside its epoch loop, re-decoding a multi-gigabyte file each pass; the run took 2.2x longer while processing 7.5x fewer training samples and completing far fewer effective epochs, contributing to a ~0.84 relative deficit on the evaluation metric versus a version that extracted features once and looped epochs over the cached array.
• Applies when: Training uses multiple passes (epochs) over a large sequential data file read via a streaming reader, and the extracted per-record features are small enough to cache as an array.
• Example:
• Input:
    for epoch in range(EPOCHS):
        for doc in iter_records(TRAIN_PATH):
            img = doc.get('imgs')[0].get('picture')
            feats = extract_features(img)
            batch_X.append(feats)
            batch_y.append(doc['label'])
            if len(batch_X) >= BATCH:
                model.partial_fit(np.array(batch_X), batch_y,
                                  classes=all_classes)
                batch_X, batch_y = [], []
• Consequence:
    Each epoch re-reads and re-decodes the entire multi-GB raw file, so
    per-epoch wall clock is dominated by I/O and decoding; total runtime
    scales as epochs x full-file decode time, forcing fewer effective
    epochs / samples within the time budget and lowering the final
    evaluation metric relative to a cached-feature run.
• Counter-example:
• Input:
    X_list, y_list = [], []
    for doc in iter_records(TRAIN_PATH):
        img = doc.get('imgs')[0].get('picture')
        X_list.append(extract_features(img))
        y_list.append(doc['label'])
    X = np.array(X_list, dtype=np.float32); y = np.array(y_list)
    for epoch in range(EPOCHS):
        order = np.random.permutation(len(X))
        for i in range(0, len(X), BATCH):
            idx = order[i:i+BATCH]
            model.partial_fit(X[idx], y[idx], classes=all_classes)
• Why it does not fire: the raw file is streamed and decoded exactly once before the epoch loop, and all epochs iterate over the cached in-memory array with in-memory shuffling, so no file read or per-record decoding occurs inside the loop.
C6 · wasted-compute Sequential per-file feature extraction over a large image collectionverified_trace · effect +0.0645

P1 — Sequential per-file feature extraction over a large image collection

• Pattern: Detects a single-process for loop that computes per-file features for every image in a directory listing one at a time, with no worker pool, thread pool, or batched/vectorized extraction path anywhere in the script.
• Detection procedure:
1. Locate a list of file names or paths built from a directory listing (e.g., a call to os.listdir, glob.glob, Path.iterdir, or equivalent), optionally filtered by an image extension such as .jpg, .png, or .jpeg.
2. Find a for loop whose iterable is that list (directly or after sorting/filtering) and whose body calls a function that opens or decodes the file (e.g., a function whose body calls Image.open, imread, or equivalent, or a wrapper function receiving the path) and appends the result to an accumulator list.
3. Search the entire source file for any occurrence of multiprocessing.Pool, Pool(, ProcessPoolExecutor, ThreadPoolExecutor, joblib.Parallel, Parallel(, map_async, imap, a data-loader class with a worker-count argument, or an equivalent parallel-map construct applied to the per-file work.
4. The pattern is PRESENT when steps 1 and 2 match and step 3 finds no parallel construct in the file, and the loop body contains no side effects other than appending to local accumulators (i.e., the per-item work is independent).
• Predicted impact:
• Add score for C6 (Avoidably expensive computation): 2
• weight: 2
• confidence: high
• Evidence: for fname in test_files: feat = extract_features(fpath); X_test_list.append(feat) executed sequentially over tens of thousands of images took 256s versus 77s (~3.3x slower) for the same feature workload when the pure per-file function was mapped over a process pool, consuming run budget that a parallel variant spent on a validation sweep over regularization strength, yielding a 0.055–0.061 improvement on the evaluation metric.
• Applies when: A script extracts hand-crafted features from a large number (thousands or more) of independent image files under a wall-clock or compute budget, and the per-file computation is pure (no shared mutable state).
• Example:
• Input:
    files = [f for f in os.listdir(IMG_DIR) if f.endswith(".jpg")]
    feats = []
    for fname in files:
        path = os.path.join(IMG_DIR, fname)
        feats.append(extract_features(path))  # decodes + computes descriptors
    X = np.vstack(feats)
• Consequence:
    Feature extraction wall-clock scales as 1x-core throughput: ~256s instead
    of ~77s on the same workload with a pool sized to available CPUs (~3.3x
    slower), leaving less budget for hyperparameter search or larger feature
    sets; the downstream metric was 0.056-0.061 worse in matched runs.
• Counter-example:
• Input:
    from multiprocessing import Pool
    files = sorted(f for f in os.listdir(IMG_DIR) if f.endswith(".jpg"))
    paths = [os.path.join(IMG_DIR, f) for f in files]
    with Pool(os.cpu_count()) as pool:
        feats = pool.map(extract_features, paths)
    X = np.vstack(feats)
• Why it does not fire: The per-file extraction is mapped over a process pool sized to the machine's CPUs, so step 3 finds a parallel construct applied to the per-file work.
C6 · wasted-compute Sequential per-file feature extraction on large file setsverified_trace · effect +0.0758

P1 — Sequential per-file feature extraction on large file sets

• Pattern: Detects feature extraction that iterates one-by-one over a large collection of independent data files in a plain sequential Python loop, loading and processing each file in the main process with no process pool, worker mapping, or parallel batching.
• Detection procedure:
1. Locate a function or loop body that loads a data file from disk (e.g. np.load, pd.read_csv, open followed by parsing, or equivalent) and computes derived features from its contents.
2. Confirm the load-and-process call is invoked inside a for loop (or list comprehension) that iterates over a collection of file identifiers or paths gathered from a directory listing or an id column of a metadata table, where each iteration depends only on its own file (no accumulator from iteration i is read at iteration i+1 other than appending results to a list).
3. Search the entire source file for any parallel-execution construct applied to that per-file function: ProcessPoolExecutor, multiprocessing.Pool, joblib.Parallel, concurrent.futures, dask, ray, or equivalent worker-mapping API.
4. PRESENT when step 2's loop exists, the per-file computations are mutually independent, and step 3 finds no parallel construct anywhere applied to that per-file work.
• Predicted impact:
• Add score for C6: Avoidably expensive computation: 2
• weight: 2
• confidence: high
• Evidence: A per-file loop for fid in ids: X.append(process(np.load(path(fid)))) run serially over tens of thousands of array files took ~1437s versus ~109s for a worker-pool version of the same extraction (>10x slower on a multi-core machine), consuming time budget that the faster pipeline spent on richer features and cross-validated modeling.
• Applies when: The program extracts features from many (hundreds or more) independent per-file inputs on a machine where multiple CPU cores are available, and the extraction loop dominates or substantially contributes to total runtime.
• Example:
• Input:
    import os, numpy as np

    def extract(path):
        arr = np.load(path)
        return [arr.mean(), arr.std(), np.abs(arr).max()]

    ids = sorted(f[:-4] for f in os.listdir("train") if f.endswith(".npy"))
    feats = []
    for fid in ids:
        feats.append(extract(os.path.join("train", fid + ".npy")))
    X = np.array(feats)
• Consequence:
    Feature extraction over ~50k files runs on one core: wall-clock time
    scales up by roughly the available-core factor (e.g. ~1437s vs ~109s,
    >10x slower), leaving less budget for feature engineering or model
    tuning within the same run and yielding a weaker final metric.
• Counter-example:
• Input:
    import os, numpy as np
    from concurrent.futures import ProcessPoolExecutor

    def extract(path):
        arr = np.load(path)
        return [arr.mean(), arr.std(), np.abs(arr).max()]

    ids = sorted(f[:-4] for f in os.listdir("train") if f.endswith(".npy"))
    paths = [os.path.join("train", fid + ".npy") for fid in ids]
    with ProcessPoolExecutor(max_workers=os.cpu_count()) as ex:
        X = np.array(list(ex.map(extract, paths, chunksize=64)))
• Why it does not fire: the same per-file extraction function is mapped over the paths through a process pool sized to available cores, so step 3 finds a parallel construct applied to the per-file work.
C6 · wasted-compute Sequential per-file feature extraction with per-file directory walk fallbacktrace_observed

P1 — Sequential per-file feature extraction with per-file directory walk fallback

• Pattern: Detects a single-process loop that extracts features from a large collection of files one at a time and, when a constructed path does not exist, performs a full recursive directory walk inside the loop body to locate each missing file, instead of building an id-to-path mapping once and distributing the extraction across workers.
• Detection procedure:
1. Locate a loop (a for statement or a sequential list comprehension) that iterates over a collection of file identifiers or rows read from a metadata table (e.g., a column of train.csv) and, in each iteration, opens and processes one file (via np.load, open, an image/audio read call, or equivalent).
2. Confirm the loop body is executed in the main process only: there is no use of ProcessPoolExecutor, multiprocessing.Pool, joblib.Parallel, or equivalent wrapping the per-file work within the same function body.
3. Within the same loop body, find a path-existence check (os.path.exists, Path.exists, a try/except around the file open, or equivalent) whose miss branch invokes a recursive search over the data directory — os.walk, glob.glob with recursive=True, Path.rglob, or equivalent — to find that single file.
4. Confirm no dictionary or table mapping file identifiers to paths is built once before the loop (e.g., a single glob/rglob pass whose results are stored in a dict keyed by filename) and consulted inside the loop.
5. The pattern is PRESENT when steps 1–4 all hold: sequential per-file processing, no worker pool, per-file recursive walk on path miss, and no precomputed path map.
• Predicted impact:
• Add score for C6: Avoidably expensive computation: 3
• weight: 3
• confidence: high
• Evidence: A sequential extraction loop with a per-miss os.walk fallback took roughly 1243s on a many-file dataset, versus 67s for the equivalent pipeline that globbed the tree once into an id-to-path dict and ran extraction in a ProcessPoolExecutor — about an 18x wall-clock difference, budget that funded a stronger model (metric gap 0.20026 in favor of the faster pipeline).
• Applies when: The dataset consists of many (thousands or more) independent files, each requiring a per-file feature extraction step before model training, and the files may live in nested subdirectories.
• Example:
• Input:
    feats = []
    for fid in df["id"]:
        path = os.path.join(DATA_DIR, f"{fid}.npy")
        if not os.path.exists(path):
            for root, _, files in os.walk(DATA_DIR):
                if f"{fid}.npy" in files:
                    path = os.path.join(root, f"{fid}.npy")
                    break
        arr = np.load(path)
        feats.append(extract_stats(arr))
    X = np.array(feats)
• Consequence:
    Feature extraction runs single-core and re-walks the entire directory tree
    for every file whose flat path misses; wall-clock is roughly an order of
    magnitude higher than a mapped/parallel equivalent (observed ~1243s vs
    ~67s), leaving no time budget for richer features or more model tuning.
• Counter-example:
• Input:
    path_map = {os.path.basename(p): p
                for p in glob.glob(os.path.join(DATA_DIR, "**", "*.npy"),
                                   recursive=True)}
    def load_one(fid):
        return extract_stats(np.load(path_map[f"{fid}.npy"]))

    with ProcessPoolExecutor() as ex:
        feats = list(ex.map(load_one, df["id"]))
    X = np.array(feats)
• Why it does not fire: the recursive glob runs exactly once before the loop to build an id-to-path map, and per-file extraction is distributed across a process pool, so steps 2–4 all fail.
C7 · nondeterminism Unordered set intersection determines feature column ordergithub_occurrence

P1 — Unordered set intersection determines feature column order

• Pattern: Detects a model feature list constructed by converting column collections through an unordered set operation (e.g. intersecting two sets and materializing the result as a list), so the order of features fed to a fitted estimator depends on hash randomization rather than a stable, deterministic ordering.
• Detection procedure:
1. Find an expression that applies set(...) to a dataframe's columns or a list of column names — e.g. set(df.columns), set(cols) — and combines it with another set using &, |, -, .intersection(...), .union(...), or .difference(...), or equivalent.
2. Check that the result of step 1 is materialized into an ordered container via list(...) (or unpacked/iterated directly) without a subsequent sorted(...) call or reordering against an original ordered column sequence (such as a comprehension filtering df.columns).
3. Check that the container from step 2 is used to subscript a dataframe (e.g. df[cols]) whose result is later passed as the feature matrix to a fitting call (.fit(...) or equivalent) or a prediction call of an estimator.
4. PRESENT when all of steps 1–3 hold: an unsorted set-derived column list determines the column order of the feature matrix used for fitting or prediction.
• Predicted impact:
• Add score for C7 (Result not reproducible): 2
• weight: 2
• confidence: high
• Evidence: common_cols = list(set(X.columns) & set(X_test.columns)); X = X[common_cols] fed to a seeded tree ensemble; a run with deterministic column ordering scored 0.02333 higher on the evaluation metric, and repeated runs of the set-based version produce different fitted models and predictions despite a fixed seed.
• Applies when: Tabular pipelines that align or select feature columns between a training frame and an evaluation frame before fitting an estimator whose result depends on feature order (e.g. tree ensembles, gradient boosting).
• Example:
• Input:
    common_cols = list(set(X.columns) & set(X_test.columns))
    X = X[common_cols]
    X_test = X_test[common_cols]

    model = SomeTreeEnsemble(n_estimators=100, random_state=42)
    model.fit(X, y)
    preds = model.predict(X_test)
    pd.DataFrame({'Id': ids, 'LABEL_A': preds}).to_csv('out.csv', index=False)
• Consequence:
    Feature column order changes between interpreter invocations due to hash
    randomization, so the fitted model and its predictions differ run to run
    despite random_state=42; the evaluation metric fluctuates across
    identical reruns (observed gap up to ~0.02 versus a deterministic
    ordering), making the reported score irreproducible.
• Counter-example:
• Input:
    common_cols = [c for c in X.columns if c in set(X_test.columns)]
    X = X[common_cols]
    X_test = X_test[common_cols]

    model = SomeTreeEnsemble(n_estimators=100, random_state=42)
    model.fit(X, y)
    preds = model.predict(X_test)
• Why it does not fire: The set is used only as a membership test inside a comprehension that iterates the original ordered X.columns, so the resulting column order is deterministic across runs.
C7 · nondeterminism Feature list from unordered set intersection

P1 — Feature list from unordered set intersection

• Pattern: Detects a model's feature column list built by converting a set operation over column names into a list, so the ordering of features fed to the estimator depends on per-process hash randomization rather than a deterministic sequence.
• Detection procedure:
1. Locate any expression that constructs a set (via set(...) or a set literal/comprehension) whose elements come from dataframe column collections (e.g. df.columns, a list of column names), combined with a set operator such as &, |, -, .intersection(...), .union(...), or .difference(...), or a bare set(...) of columns.
2. Check that the result is materialized into an ordered container by wrapping it in list(...) (or iterating it directly) WITHOUT an enclosing sorted(...) call and without re-deriving order from an ordered source (e.g. a list comprehension over the original column list filtered by set membership).
3. Check that the resulting name collection is used, within the same script, to subscript one or more dataframes (e.g. df[cols]) that are subsequently passed to a model-fitting or prediction call, or equivalent.
4. PRESENT if all of steps 1–3 hold: an unsorted, set-derived column list determines the feature ordering used for training or inference.
• Predicted impact:
• Add score for C7: Result not reproducible: 2
• weight: 2
• confidence: high
• Evidence: common_cols = list(set(X.columns) & set(X_test.columns)); X = X[common_cols] — feature order varied across otherwise identical runs because set iteration order depends on per-process string hashing; tree tie-breaks then differ, and the observed metric shifted by 0.02333 between runs despite a fixed model seed.
• Applies when: Tabular pipelines that align or select feature columns across dataframes before fitting a model.
• Example:
• Input:
    import pandas as pd
    from sklearn.ensemble import RandomForestClassifier

    X = train.drop(columns=["target"])
    X_test = test.copy()
    common_cols = list(set(X.columns) & set(X_test.columns))
    X = X[common_cols]
    X_test = X_test[common_cols]
    model = RandomForestClassifier(random_state=42)
    model.fit(X, train["target"])
    preds = model.predict(X_test)
• Consequence:
    Feature column order differs between reruns of the same script (set
    iteration follows per-process string hash randomization), changing
    tree-split tie-breaking; the reported metric drifts run-to-run
    (observed delta ~0.023) even though random_state is fixed, so the
    result cannot be reproduced.
• Counter-example:
• Input:
    import pandas as pd
    from sklearn.ensemble import RandomForestClassifier

    X = train.drop(columns=["target"])
    X_test = test.copy()
    common_cols = [c for c in X.columns if c in set(X_test.columns)]
    X = X[common_cols]
    X_test = X_test[common_cols]
    model = RandomForestClassifier(random_state=42)
    model.fit(X, train["target"])
    preds = model.predict(X_test)
• Why it does not fire: the set is used only for membership testing while the ordering is inherited deterministically from the original X.columns sequence, so feature order is identical on every run.
C8 · output-misalignment Label-feature alignment by prefix truncation after filtering

P1 — Label-feature alignment by prefix truncation after filtering

• Pattern: Detects supervised training where the label array is built by conditionally filtering out invalid records while the feature matrix is aligned to it merely by slicing its first len(labels) rows, so features and labels are paired by position after the row correspondence has been broken.
• Detection procedure:
1. Find an assignment where a label variable (e.g. y, y_train, labels) is built from a collection using a conditional filter — a list comprehension with an if clause, a call to filter(...), or boolean-mask indexing applied to the labels only.
2. Within the same function body or module scope, find a subsequent assignment where a feature variable is sliced with the length of that label variable, e.g. X[:len(y)], X[0:len(y_train)], or an equivalent .head(len(y)) / X[:y.shape[0]] construct.
3. Verify that no shared boolean mask or index array derived from the same validity condition is applied to the feature variable anywhere between steps 1 and 2.
4. Verify that both variables are later passed together to a model-fitting call (e.g. model.fit(X_sliced, y) or equivalent).
5. The pattern is PRESENT when the label filtering (step 1), the prefix slice of features by label length (step 2), the absence of a shared mask on features (step 3), and the joint use in fitting (step 4) all hold.
• Predicted impact:
• Add score for C8: Correct predictions, wrong arrangement: 3
• weight: 3
• confidence: high
• Evidence: y_train = le.fit_transform([c for c in cats if c is not None]) followed by X_train_filtered = X_train[:len(y_train)] and clf.fit(X_train_filtered, y_train); the run with this construct scored 0.326 lower on the evaluation metric than an implementation that appended feature/label pairs only when both were valid, because every feature row after the first dropped record was paired with the wrong label.
• Applies when: Supervised training code that loads records via streaming or iteration, skips some records based on a validity condition, and then assembles feature and label arrays separately before fitting.
• Example:
• Input:
    feats = []
    cats = []
    for rec in iterate_records("train.bin"):
        feats.append(extract(rec))
        cats.append(rec.get("label"))
    X = np.array(feats)
    y = enc.fit_transform([c for c in cats if c is not None])
    X = X[:len(y)]
    model.fit(X, y)
• Consequence:
    Any record with a missing label that is not at the tail shifts all
    subsequent labels relative to their feature rows; training accuracy
    and held-out metric drop toward chance level (observed -0.33 on the
    evaluation metric versus correctly paired data), with no error raised.
• Counter-example:
• Input:
    feats = []
    cats = []
    for rec in iterate_records("train.bin"):
        c = rec.get("label")
        if c is None:
            continue
        feats.append(extract(rec))
        cats.append(c)
    X = np.array(feats)
    y = enc.fit_transform(cats)
    model.fit(X, y)
• Why it does not fire: features and labels are appended together only when the record is valid, so no conditional filter is applied to labels alone and no prefix slice of the feature matrix by label length occurs.
C8 · output-misalignment Pose lookup by sample id without sensor/key-frame filterverified_trace · effect +0.0001

P1 — Pose lookup by sample id without sensor/key-frame filter

• Pattern: Detects per-sample pose or calibration lookup that selects the first record matching a sample identifier from a multi-sensor record table without restricting to the key-frame flag and the sensor modality the annotations are defined against, then uses that pose to place predicted geometry.
• Detection procedure:
1. Locate code that loads a table or list of per-sensor-record entries (e.g., a sample_data JSON/CSV or equivalent) where each entry carries both a sample identifier field and a pose or calibration reference field, and multiple entries can share the same sample identifier.
2. Within the same file, locate construction of a mapping from sample identifier to a pose/calibration record: a loop that assigns mapping[rec[<sample_id_field>]] = ... guarded only by a membership test such as not in mapping (or no guard at all), or a groupby(...).first() / drop_duplicates(<sample_id_field>) call, or equivalent first-match selection.
3. Confirm that between step 1 and step 2 there is no conditional or row filter testing BOTH (a) a key-frame boolean field (e.g., is_key_frame or equivalent) and (b) a sensor-channel/modality field, applied to the records before selection.
4. Confirm the mapped pose is later read (translation and rotation components) and applied arithmetically to predicted coordinates (e.g., rotating/offsetting box centers) that are written into an output file.
5. PRESENT when steps 2, 3, and 4 all hold: first-match pose selection lacking both filters feeds the spatial transform of emitted predictions.
• Predicted impact:
• Add score for C8: Correct predictions, wrong arrangement: 3
• weight: 3
• confidence: high
• Evidence: A mapping built with if rec["sample_token"] not in pose_map: pose_map[rec["sample_token"]] = rec over an unfiltered multi-sensor record table selected poses from arbitrary sensors/timestamps; the run scored the metric floor, while an otherwise similar script that filtered to key-frame records of the annotated sensor before the pose join scored the full normalized gap (1.0) higher.
• Applies when: A multi-sensor dataset where predictions are expressed in a world/global frame via a per-sample pose or calibration record, and annotations are tied to one specific sensor's key-frame timestamps.
• Example:
• Input:
    sd = json.load(open("sample_data.json"))
    poses = {p["token"]: p for p in json.load(open("ego_pose.json"))}
    sample_to_pose = {}
    for rec in sd:
        tok = rec["sample_token"]
        if tok not in sample_to_pose:
            sample_to_pose[tok] = poses[rec["ego_pose_token"]]
    for tok in df["Id"]:
        ego = sample_to_pose.get(tok)
        tx, ty, tz = ego["translation"]
        rows.append(f"{tx + dx:.3f} {ty + dy:.3f} {tz:.3f}")
• Consequence:
    Each sample's pose comes from whichever sensor record happens first in the
    table, often a non-key-frame camera sweep captured tens of milliseconds
    away from the annotated timestamp. Predicted boxes are systematically
    offset by the ego displacement over that interval (meters at driving
    speed), so overlap with ground truth collapses at tight IoU thresholds
    and the mean-AP metric drops sharply relative to a filtered pose join.
• Counter-example:
• Input:
    SENSOR_A = "sensor_a"
    sd = json.load(open("sample_data.json"))
    poses = {p["token"]: p for p in json.load(open("ego_pose.json"))}
    sample_to_pose = {}
    for rec in sd:
        if not rec["is_key_frame"] or rec["channel"] != SENSOR_A:
            continue
        sample_to_pose[rec["sample_token"]] = poses[rec["ego_pose_token"]]
    for tok in df["Id"]:
        ego = sample_to_pose.get(tok)
• Why it does not fire: the loop filters records on both the key-frame flag and the sensor-channel field before selecting the pose, so step 3's missing-filter condition fails.
C8 · output-misalignment Cross-table lookup keyed on wrong identifierverified_trace · effect +0.0002

P1 — Cross-table lookup keyed on wrong identifier

• Pattern: Detects a cross-table join that builds a lookup mapping keyed on each record's own primary identifier field and then probes it with a different table's identifier values, so the foreign-key field the probe values actually reference is never used and every membership test misses.
• Detection procedure:
1. Locate a mapping construction — a dict comprehension or a loop performing d[key] = value — where the key expression is a subscript of the iterated record with a literal string field name denoting the record's own identifier (for example 'token', 'id', or 'uid'), and the iterated collection was loaded from one file or table (via json.load, pd.read_json, or equivalent).
2. Locate a later probe of that same mapping — an in test, a .get(x) call, or a subscript d[x] — where x is a subscript of a record iterated from a different collection, loaded from a different file or table than the collection in step 1, and x uses the same literal field name as the key in step 1.
3. Within the same file, records of the step-1 collection are subscripted elsewhere (or their fields are visibly enumerated) with a literal field name that is a compound containing the step-2 collection's noun plus the identifier suffix (for example 'sample_token' where the step-1 key was 'token'), and that compound field name never appears as the key expression in the step-1 mapping construction.
4. PRESENT when steps 1–3 all hold and no assertion, exception, or conditional abort in the same file checks that the count of successful probes is greater than zero.
• Predicted impact:
• Add score for C8 (Correct predictions, wrong arrangement): 3
• weight: 3
• confidence: high
• Evidence: A pipeline built lookup = {a['token']: a for a in anns} and probed it with s['token'] taken from a different table whose identifiers live in a disjoint space; the child records' actual foreign-key field (a compound name referencing the parent table) was never used as the key. The join matched zero rows, the pipeline fell through to its empty-placeholder branch, and the run scored exactly the trivial baseline while an otherwise-comparable run keyed on the foreign key scored strictly higher (measured gap: full metric range).
• Applies when: The program loads two or more relational tables (JSON, CSV, or equivalent) that are linked by identifier fields, and assembles training or output rows by dictionary-based membership tests between them rather than a declarative merge.
• Example:
• Input:
    anns = json.load(open('table_b.json'))
    samples = json.load(open('table_a.json'))
    lookup = {a['token']: a for a in anns}   # keyed on child's own id
    rows = []
    for s in samples:
        if s['token'] in lookup:             # probes with parent's id
            rows.append((s['token'], lookup[s['token']]['value']))
    print(len(rows))
• Consequence:
    len(rows) is 0 on every run: the assembled training set is empty, the
    pipeline takes its no-data fallback and emits only blank placeholder
    outputs, and the evaluation metric lands at the trivial baseline (0.0)
    instead of the positive score a correctly keyed join produces.
• Counter-example:
• Input:
    anns = json.load(open('table_b.json'))
    samples = json.load(open('table_a.json'))
    lookup = {}
    for a in anns:
        lookup.setdefault(a['sample_token'], []).append(a)  # foreign key
    rows = []
    for s in samples:
        for a in lookup.get(s['token'], []):
            rows.append((s['token'], a['value']))
    assert len(rows) > 0, "join produced no matches"
• Why it does not fire: the mapping is keyed on the compound foreign-key field that the probing table's identifiers actually reference, and a nonzero-match assertion guards the join, so steps 3 and 4 both fail.
C9 · discarded-signal Stop-word removal or small vocabulary cap on a stylometric text task

P1 — Stop-word removal or small vocabulary cap on a stylometric text task

• Pattern: Detects a bag-of-words or TF-IDF text vectorizer for an authorship/style-attribution task configured to discard stop words or to cap the vocabulary at a small hardcoded feature count, throwing away the function-word frequencies and rare-token usage that carry the primary stylistic signal.
• Detection procedure:
1. Locate any constructor call that builds a token-frequency text feature extractor (e.g., TfidfVectorizer(...), CountVectorizer(...), or equivalent) whose output is later passed to a classifier fitting method within the same script.
2. Confirm the prediction target is an author or writing-style identity: the label column is fed to a classifier and the task setup (comments, label values used as output columns, or problem framing in the script) indicates each class corresponds to a distinct writer or style rather than a topic-independent category.
3. In the constructor call at step 1, check whether either of the following keyword arguments appears: (a) stop_words set to any value other than None (e.g., a language string or an explicit word list), or (b) max_features set to a literal integer less than or equal to 20000.
4. The pattern is PRESENT if steps 1–3 all hold, i.e., the vectorizer feeding the style classifier strips stop words, caps the vocabulary at a small literal size, or both.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A vectorizer built with stop_words='english', max_features=5000 on an author-attribution task produced a multi-class log loss roughly 0.44 worse than an otherwise comparable pipeline that retained stop words and the full vocabulary; function-word frequencies are the strongest per-author discriminator, so removing them directly degrades the metric.
• Applies when: A text-classification pipeline predicts author or style identity from raw text using token-frequency features (bag-of-words, TF-IDF, or equivalent).
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # predict which author wrote each sentence
    vec = TfidfVectorizer(stop_words='english', max_features=5000)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author_label'])
    probs = model.predict_proba(X_test)
• Consequence:
    Multi-class log loss substantially worse (~0.4 higher, tens of percent
    relative degradation) than the same pipeline with stop words retained and
    no vocabulary cap, because per-author function-word usage rates — the
    strongest stylistic discriminator — are absent from the feature matrix.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # predict which author wrote each sentence
    vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author_label'])
    probs = model.predict_proba(X_test)
• Why it does not fire: The vectorizer neither sets stop_words to a non-None value nor caps max_features, so function words and the full vocabulary remain available to the classifier.
C9 · discarded-signal Fixed central-fraction slice window in volumetric feature extractiontrace_observed

P1 — Fixed central-fraction slice window in volumetric feature extraction

• Pattern: Detects per-subject 2D-slice feature extraction from a 3D volume that restricts processing to a hardcoded fractional window of the slice index range (e.g. the middle 30%–70%), discarding all slices outside that window regardless of where the target structure lies in each subject.
• Detection procedure:
1. Locate a loop or function that processes one subject/case at a time and builds an ordered collection of 2D slices (a sorted list of per-slice image files, or a 3D array whose first or last axis is the slice axis).
2. Within the same function body, find an expression computing slice-index bounds by multiplying the slice count (e.g. len(files), arr.shape[0], or a variable assigned from either) by a literal fractional constant strictly between 0 and 1 (such as 0.3, 0.4, 0.25) or by integer division with a literal (such as n // 3), where both a lower bound greater than zero and an upper bound less than the count are produced.
3. Confirm those bounds are used to subscript the slice collection (e.g. files[lo:hi] or arr[lo:hi] or equivalent iteration limits), and that slices outside [lo, hi) are not read, aggregated, or otherwise used anywhere else in feature construction for that subject.
4. Confirm there is no content-based fallback or selection (no check of per-slice pixel statistics such as nonzero fraction or intensity sum used to choose slices) and no evenly spaced sampling over the full index range (no call like np.linspace(0, n - 1, k) or equivalent).
5. PRESENT when steps 2, 3, and 4 all hold: features are computed exclusively from a positionally fixed central fraction of each subject's volume.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline computing per-slice features only from files[int(n*0.3):int(n*0.7)] scored 0.17637 worse on the ranking metric than an otherwise comparable pipeline that drew features from across the full slice range; subjects whose relevant structure sat outside the fixed window contributed no discriminative signal.
• Applies when: The code extracts tabular or image features from 2D slices of per-subject 3D scan volumes (e.g. stacks of DICOM/NIfTI/PNG slices) where the axial position of the structure of interest varies between subjects.
• Example:
• Input:
    def subject_features(subject_dir):
        files = sorted(os.listdir(subject_dir))
        n = len(files)
        lo, hi = int(n * 0.3), int(n * 0.7)
        feats = []
        for f in files[lo:hi]:
            img = load_slice(os.path.join(subject_dir, f))
            feats.append([img.mean(), img.std(), (img > 0).mean()])
        return np.array(feats).mean(axis=0)
• Consequence:
    Subjects whose target structure lies in the first 30% or last 30% of the
    slice stack yield features indistinguishable from background; ranking
    metric (AUC) drops by ~0.17 versus sampling across the full slice range.
• Counter-example:
• Input:
    def subject_features(subject_dir):
        files = sorted(os.listdir(subject_dir))
        n = len(files)
        idx = np.linspace(0, n - 1, 16).astype(int)
        feats = []
        for i in idx:
            img = load_slice(os.path.join(subject_dir, files[i]))
            feats.append([img.mean(), img.std(), (img > 0).mean()])
        return np.array(feats).mean(axis=0)
• Why it does not fire: the slice indices are spread evenly across the entire index range via np.linspace(0, n - 1, k), so no positional window excludes any region of the volume.
C9 · discarded-signal Word-only text features where sub-word style carries signal

P1 — Word-only text features where sub-word style carries signal

• Pattern: Detects a stylistic text-classification pipeline that builds its feature matrix exclusively from word-level token n-grams, never encoding any character-level or sub-word features from the raw text.
• Detection procedure:
1. Identify every text-vectorization construction in the source: calls to TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent, whose output is later passed to a model's fitting method.
2. For each such construction, inspect its arguments: it is word-level if the analyzer argument is absent, or is the literal "word"; it is character-level if analyzer is "char" or "char_wb", or an equivalent sub-word tokenizer is used.
3. Check whether the feature matrix passed to any fitted model includes output from at least one character-level construction from step 2 (directly, via hstack, via a feature-union/pipeline combinator, or equivalent concatenation).
4. Confirm the prediction target is a per-text style-like label (e.g., which of several writers/sources produced the text) rather than a topical category, judged from how the label column is derived and used.
5. PRESENT if at least one word-level vectorizer feeds a fitted model, no character-level feature block is concatenated into any model's input anywhere in the program, and the condition in step 4 holds.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only TfidfVectorizer(ngram_range=(1, 2)) word features scored 0.39974 worse on the log-loss-based competition metric than an otherwise similar model that concatenated a analyzer="char_wb", ngram_range=(2, 4) block via sparse hstack; punctuation, morphology, and spelling signal in the raw text was never encoded by the word-only version.
• Applies when: Classical-ML text classification where the label reflects who or what style produced the text, features are bag-of-n-gram matrices, and the raw text field is available at feature-construction time.
• Example:
• Input:
    vec = TfidfVectorizer(ngram_range=(1, 2), min_df=2)
    X_tr = vec.fit_transform(train_df["text"])
    X_te = vec.transform(test_df["text"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, train_df["label"])
    proba = model.predict_proba(X_te)
• Consequence:
    Multiclass log loss is ~0.4 higher than the same classifier fed the
    word block concatenated with a character n-gram TF-IDF block; the
    orthographic/stylistic signal in the text is systematically discarded,
    so probability estimates are less confident and less accurate.
• Counter-example:
• Input:
    word_vec = TfidfVectorizer(ngram_range=(1, 2), min_df=2)
    char_vec = TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 4), min_df=3)
    Xw_tr = word_vec.fit_transform(train_df["text"])
    Xc_tr = char_vec.fit_transform(train_df["text"])
    X_tr = hstack([Xw_tr, Xc_tr]).tocsr()

    model = LogisticRegression(C=4.0, max_iter=1000)
    model.fit(X_tr, train_df["label"])
• Why it does not fire: a character-level vectorizer (analyzer="char_wb") is constructed and its output is concatenated into the fitted model's feature matrix, so step 3 finds a sub-word feature block.
C9 · discarded-signal Prefix slice of flattened image pixels as feature vectorgithub_occurrence

P1 — Prefix slice of flattened image pixels as feature vector

• Pattern: Detects a handcrafted image feature vector built by taking a fixed-length prefix slice of a flattened 2-D or 3-D pixel array, so the retained features cover only a contiguous band of rows at one edge of every image and the remaining spatial content is discarded.
• Detection procedure:
1. Locate a call that loads or resizes an image into an array with known spatial dimensions (e.g. img.resize((W, H)), cv2.resize(img, (W, H)), or equivalent), where W and H are literal integers.
2. Within the same function body or loop body, locate an expression that flattens that array into one dimension via .flatten(), .ravel(), .reshape(-1), or equivalent.
3. Check whether the flattened expression is immediately or subsequently subscripted with a slice whose stop bound is a literal integer N and whose start is absent or 0 (i.e. [:N]), with no step, and N is less than half of W * H (times the channel count if the array is 3-D).
4. Confirm the sliced result is appended to or assembled into the feature matrix later passed to a model-fitting call.
5. The pattern is PRESENT when steps 1–4 all hold: a flattened pixel array is prefix-sliced to a literal length far below the full pixel count and used as the model's features.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A pipeline built features with np.array(img).flatten()[:N] where N retained roughly 13% of the downsampled pixels (only the top rows of every image); the resulting classifier's log loss was about 20% worse than a pipeline using a spatially complete descriptor over the same images.
• Applies when: An image classification or regression task uses handcrafted (non-learned) features derived directly from pixel arrays fed to a classical model.
• Example:
• Input:
    features = []
    for path in image_paths:
        img = Image.open(path).convert('L').resize((64, 64))
        arr = np.array(img, dtype=np.float32) / 255.0
        features.append(arr.flatten()[:512])
    X = np.vstack(features)
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    model.fit(X_scaled, y)
• Consequence:
    Only 512 of 4096 pixels (the top 8 rows of each 64x64 image) reach the
    model; the bottom 87% of every image is discarded. Validation log loss
    is ~20% higher than the same model trained on a spatially complete
    feature representation of identical dimensionality.
• Counter-example:
• Input:
    features = []
    for path in image_paths:
        img = Image.open(path).convert('L').resize((16, 16))
        arr = np.array(img, dtype=np.float32) / 255.0
        features.append(arr.flatten())
    X = np.vstack(features)
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    model.fit(X_scaled, y)
• Why it does not fire: The dimensionality is reduced by resizing the whole image before flattening, so the full flattened vector (all 256 values) covers the entire spatial extent and no prefix slice is applied.
C9 · discarded-signal Global intensity statistics computed over zero-valued background pixelstrace_observed

P1 — Global intensity statistics computed over zero-valued background pixels

• Pattern: Detects handcrafted intensity features (mean, standard deviation, percentiles, skewness, or similar distributional statistics) computed over an entire scan-like image array — including its large zero-valued background — without first restricting the pixel set to nonzero or otherwise-masked foreground pixels.
• Detection procedure:
1. Locate a function that receives a 2D or 3D numeric array loaded from an image or scan file (e.g., via pydicom.dcmread, nib.load, Image.open, np.load, or equivalent) and returns a list, tuple, or vector of scalar features used as model input.
2. Within that function body, find calls to distributional-statistic operations applied to the array — np.mean, np.std, np.percentile, np.median, skew, kurtosis, or equivalent — where the argument is the full array (e.g., arr, arr.flatten(), arr.ravel()).
3. Check that between the array's creation and each such statistic call there is no boolean-mask subscript excluding zeros or background (no expression of the form arr[arr > 0], arr[mask], arr[arr != 0], or an equivalent call filtering to nonzero elements such as arr[np.nonzero(arr)]), and no cropping step that removes background rows/columns before the statistics.
4. PRESENT when at least one distributional statistic in the feature-extraction function is computed over the unmasked full array under the conditions of steps 1–3.
• Predicted impact:
• Add score for C9: 2
• weight: 2
• confidence: high
• Evidence: Feature extraction called statistics like np.mean(arr) and np.percentile(arr, ...) on full scan slices dominated by zero background; a contrastive run whose features isolated informative pixel regions scored ~0.10–0.15 higher on the classification metric (AUC), because the unmasked features tracked background proportion and framing rather than tissue intensity distribution.
• Applies when: A model is trained on handcrafted per-image scalar features derived from scan-like images (medical or scientific imaging) whose frames contain large regions of exactly-zero background, and the features include global intensity distribution statistics.
• Example:
• Input:
    def extract_features(path):
        arr = pydicom.dcmread(path).pixel_array.astype(np.float32)
        feats = [
            np.mean(arr),
            np.std(arr),
            np.percentile(arr, 25),
            np.percentile(arr, 75),
            skew(arr.flatten()),
        ]
        return feats
• Consequence:
    Classification AUC drops ~10% relative versus foreground-masked features:
    per-subject feature values vary with the fraction of zero background
    (framing/crop) instead of the foreground intensity distribution, so the
    classifier learns background area rather than the discriminative signal.
• Counter-example:
• Input:
    def extract_features(path):
        arr = pydicom.dcmread(path).pixel_array.astype(np.float32)
        fg = arr[arr > 0]
        if fg.size == 0:
            return [0.0] * 6
        feats = [
            np.mean(fg),
            np.std(fg),
            np.percentile(fg, 25),
            np.percentile(fg, 75),
            skew(fg),
            fg.size / arr.size,  # background fraction as its own feature
        ]
        return feats
• Why it does not fire: the statistics are computed on arr[arr > 0], a nonzero-foreground subset, so the mask-exclusion condition in step 3 is not met.
C9 · discarded-signal Single-mean-plus-noise spatial prior instead of multi-mode extractiongithub_occurrence

P1 — Single-mean-plus-noise spatial prior instead of multi-mode extraction

• Pattern: Detects a prediction routine that reduces the training set's empirical distribution of target locations to one per-group mean coordinate and then generates predicted positions by adding random noise around that mean, rather than deriving multiple representative locations from the same data.
• Detection procedure:
1. Locate an aggregation over training-set coordinate columns that computes a single central value per group, e.g. .mean(), np.mean, or a manually accumulated average producing one value per coordinate axis per group, with no clustering call (KMeans, DBSCAN, MeanShift, histogram-mode extraction, or equivalent) applied to those coordinates anywhere in the same script.
2. Within the loop or function that emits predictions, find that each predicted coordinate is computed as that per-group mean plus a random draw, e.g. mean_x + rng.normal(...), mean_x + np.random.randn(...) * std_x, or equivalent random-perturbation expressions.
3. Confirm the emitted predictions per input contain only points generated from this single mean (possibly repeated with different noise draws), and no other mechanism produces multiple distinct data-derived locations.
4. PRESENT when step 1's single-central-value aggregation exists without any multi-mode extraction, and step 2's noise-around-mean expression feeds the emitted predictions.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: Predictions built as stats["rel_x"] + rng.normal(0, stats["std_x"] * 0.5) around a single per-class mean location scored 0.87719 lower on the overlap-based competition metric than a variant that clustered the same training coordinates and emitted the cluster centers as candidates.
• Applies when: A script predicts spatial locations (2D/3D coordinates) for evaluation inputs using only aggregate statistics computed from training-set locations, with no learned model per input.
• Example:
• Input:
    stats = train.groupby("cls")[["x", "y", "z"]].agg(["mean", "std"])
    rng = np.random.default_rng(0)
    preds = []
    for sample_id in test_ids:
        for c in classes:
            mx, my, mz = stats.loc[c, [("x","mean"),("y","mean"),("z","mean")]]
            px = mx + rng.normal(0, stats.loc[c, ("x","std")] * 0.5)
            py = my + rng.normal(0, stats.loc[c, ("y","std")] * 0.5)
            preds.append((sample_id, c, px, py, mz))
• Consequence:
    Predicted locations scatter around a single low-density centroid between the
    true density modes; overlap-based precision drops toward zero (observed gap
    ~0.88 on the evaluation metric versus a mode-seeking baseline using the same
    training statistics).
• Counter-example:
• Input:
    from sklearn.cluster import KMeans
    km = KMeans(n_clusters=8, random_state=0).fit(train[["x", "y"]])
    cands = [(cx, cy, n) for (cx, cy), n in
             zip(km.cluster_centers_, np.bincount(km.labels_))]
    preds = []
    for sample_id in test_ids:
        for cx, cy, n in cands:
            preds.append((sample_id, cx, cy, n / len(train)))
• Why it does not fire: The training coordinates feed a clustering step that yields multiple deterministic density modes as candidate locations, so no single-mean-plus-random-noise construction is present.
C9 · discarded-signal Hardcoded spatial offsets ignoring available training coordinatesgithub_occurrence

P1 — Hardcoded spatial offsets ignoring available training coordinates

• Pattern: Detects a prediction-generation loop that places predicted object positions using literal numeric offset constants defined in the source, while training annotation files containing target coordinates are loaded or available but never aggregated into positional statistics.
• Detection procedure:
1. Locate a module-level or function-level assignment of a list or tuple of numeric literal pairs/triples (e.g., offsets = [(16.0, 0.0), (8.0, 3.0), ...]) whose elements are all numeric literals with no variable derived from data.
2. Confirm that inside a loop producing output rows (rows later written via to_csv, to_json, or equivalent), predicted coordinates are computed arithmetically from elements of the structure found in step 1 combined with a per-sample reference pose (translation/rotation), and from no other data-derived quantity.
3. Search the whole source for any statement that reads a training annotation source (a file or table containing ground-truth coordinates, e.g., loading a train.csv or annotation JSON with position fields) AND feeds those coordinates into an aggregation such as mean, std, median, a histogram, or a fitted distribution whose result is used in step 2's coordinate computation.
4. PRESENT if step 1 and step 2 hold and no such aggregation from step 3 flows into the predicted coordinates — i.e., positions come solely from hardcoded literals despite training coordinates being available in the parsed inputs.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)] used to place all predicted boxes relative to the reference pose, while parsed training annotations with per-object coordinates were never aggregated; the overlap-thresholded evaluation metric dropped by 1.0 (to zero) versus a variant that fitted per-class relative-position statistics from the same annotations.
• Applies when: A heuristic (non-learned) baseline generates spatial predictions and the program has access to training annotations containing target coordinates for the same object classes.
• Example:
• Input:
    offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)]
    rows = []
    for sid in sub_df["Id"]:
        ego = poses[sid]
        tx, ty, tz = ego["translation"]
        yaw = ego["yaw"]
        parts = []
        for dx, dy in offsets:
            wx = tx + dx * math.cos(yaw) - dy * math.sin(yaw)
            wy = ty + dx * math.sin(yaw) + dy * math.cos(yaw)
            parts.append(f"1.0 {wx:.3f} {wy:.3f} {tz:.3f} {W} {L} {H} {yaw:.4f} LABEL_A")
        rows.append(" ".join(parts))
• Consequence:
    Predicted positions cluster at fixed guessed offsets that rarely overlap the
    empirical object-location distribution; IoU with ground truth falls below
    the matching threshold for nearly all predictions, driving the overlap-
    thresholded metric to ~0 versus a data-fitted placement baseline.
• Counter-example:
• Input:
    ann = pd.read_csv("train.csv")
    stats = ann.groupby("label")[["rel_x", "rel_y", "rel_z"]].agg(["mean", "std"])
    rows = []
    for sid in sub_df["Id"]:
        ego = poses[sid]
        m = stats.loc["LABEL_A"]
        rx = m[("rel_x", "mean")] + rng.normal(0, m[("rel_x", "std")] * 0.5)
        ry = m[("rel_y", "mean")] + rng.normal(0, m[("rel_y", "std")] * 0.5)
        center = ego["trans"] + ego["R"].dot([rx, ry, m[("rel_z", "mean")]])
        rows.append(f"1.0 {center[0]:.3f} {center[1]:.3f} {center[2]:.3f}")
• Why it does not fire: predicted coordinates are derived from means and spreads aggregated out of the training annotations, so step 3's aggregation flows into the placement computation and no hardcoded literal offset list determines position.
C9 · discarded-signal Stop-word removal in an authorship/style text pipelineverified_trace · effect +0.0768

P1 — Stop-word removal in an authorship/style text pipeline

• Pattern: Detects a bag-of-words or TF-IDF text vectorizer configured to discard stop words when the prediction target is the author or writing style of the text rather than its topic, removing function-word frequencies that carry the discriminative stylistic signal.
• Detection procedure:
1. Identify a construction of a token-count or TF-IDF text vectorizer (e.g. TfidfVectorizer(...), CountVectorizer(...), or equivalent) whose keyword arguments include stop_words set to a language name string (such as 'english') or to a non-empty list/set of words.
2. Confirm the matrix produced by this vectorizer's fit_transform/transform (or equivalent) is passed, within the same script, to a classifier fitting call whose label vector is derived from a column identifying the author, writer, or style of each text sample (e.g. the label column is used as y and the task description, comments, or output columns indicate per-author probabilities or author identities).
3. The pattern is PRESENT when both hold: the vectorizer feeding the classifier removes stop words (step 1) and the classification target is authorship/style rather than topic (step 2).
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A pipeline using TfidfVectorizer(stop_words='english', ...) for multiclass author prediction scored ~0.40 worse on multiclass log loss than an otherwise comparable pipeline that retained stop words, because function words (the strongest per-author discriminators) were excluded from the feature space.
• Applies when: The script performs text classification where the label identifies the author or writing style of each document, and features are built with a bag-of-words or TF-IDF vectorizer.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # target: which author wrote each sentence
    vec = TfidfVectorizer(stop_words='english', max_features=20000)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author'])
    probs = model.predict_proba(X_test)
• Consequence:
    Multiclass log loss on the authorship task worsens substantially
    (observed ~0.40 absolute increase versus an identical pipeline that
    keeps stop words), because function-word frequencies — the strongest
    per-author discriminators — are removed before the model sees them.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # target: which author wrote each sentence
    vec = TfidfVectorizer(ngram_range=(1, 2), max_features=20000)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author'])
    probs = model.predict_proba(X_test)
• Why it does not fire: The vectorizer leaves stop_words at its default (no removal), so function-word frequencies remain in the feature space for the authorship classifier.
C9 · discarded-signal Capped word-only TF-IDF with no character n-gram featuresverified_trace · effect +0.1015

P1 — Capped word-only TF-IDF with no character n-gram features

• Pattern: Detects a text-classification pipeline whose only text representation is a word-level TF-IDF with a small hardcoded vocabulary cap (a few thousand features or fewer) and no character-level n-gram feature block, discarding sub-word stylistic signal present in the raw text.
• Detection procedure:
1. Locate every construction of a TF-IDF or count-based text vectorizer (e.g. TfidfVectorizer, CountVectorizer, or equivalent) in the source.
2. Check whether every such constructor either omits analyzer or sets analyzer='word' (or equivalent word-level tokenization); i.e., no constructor sets analyzer='char' or analyzer='char_wb' (or an equivalent character-level mode).
3. Check whether at least one such word-level vectorizer passes a max_features argument (or equivalent vocabulary-cap parameter) with a literal integer value less than or equal to 10000.
4. Check that the features fed to the model-fitting call come only from these vectorizers — no concatenation (e.g. hstack, FeatureUnion, ColumnTransformer, or equivalent) with any character-level text feature block appears between vectorization and fitting.
5. The pattern is PRESENT when steps 2, 3, and 4 all hold: the sole text representation is word-level, vocabulary-capped at a small literal value, and no character-level features are ever combined in.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using TfidfVectorizer(max_features=5000) word features alone scored ~0.60 multiclass log loss, while an otherwise comparable pipeline that concatenated word and character n-gram TF-IDF blocks via sparse hstack scored ~0.35 — roughly 42% relative worse for the capped word-only representation.
• Applies when: The task is classification over raw text documents (especially style- or author-sensitive tasks), features are built with TF-IDF or count vectorization, and the evaluation metric rewards well-calibrated fine-grained distinctions (e.g. log loss).
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(max_features=5000, stop_words='english')
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
• Consequence:
    Multiclass log loss is substantially higher than with an uncapped
    word + character n-gram representation (observed ~0.60 vs ~0.35,
    about 42% relative degradation) on an otherwise identical pipeline.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack
    from sklearn.linear_model import LogisticRegression

    word_vec = TfidfVectorizer(ngram_range=(1, 2), min_df=2)
    char_vec = TfidfVectorizer(analyzer='char_wb', ngram_range=(2, 5), max_features=50000)
    X_train = hstack([word_vec.fit_transform(train['text']),
                      char_vec.fit_transform(train['text'])]).tocsr()
    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)
• Why it does not fire: A character-level n-gram block is concatenated with the word features before fitting, and the word vectorizer has no small literal vocabulary cap, so steps 2–4 fail.
C9 · discarded-signal Word-token-only features for style-based text classification

P1 — Word-token-only features for style-based text classification

• Pattern: Detects a text-classification pipeline whose label depends on writing style (all classes share the same domain/topic) that builds every feature matrix from word-level tokenization only, with no character-level n-gram representation anywhere before model fitting.
• Detection procedure:
1. Confirm the script reads a text column and a categorical label column, fits a classifier on features derived from the text, and writes per-class predictions (a text-classification pipeline).
2. List every text-vectorization construct in the script, e.g. instantiations of TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent tokenizing/featurizing calls that transform the raw text column.
3. For each construct found in step 2, check whether it is configured for character-level analysis: an analyzer argument equal to 'char' or 'char_wb', or an equivalent character-n-gram option in another library.
4. Check whether any feature matrix passed to a model-fitting call is a concatenation (e.g. hstack, np.hstack, FeatureUnion, ColumnTransformer, or equivalent) that includes a character-level representation from step 3.
5. PRESENT if at least one vectorizer from step 2 exists, none of them satisfies step 3, and no character-level feature block reaches any fitting call per step 4.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only TfidfVectorizer(max_features=5000, ngram_range=(1, 2)) word features scored 0.44 worse on the log-loss-style metric than a pipeline that stacked word and character n-gram feature blocks before fitting; sub-word stylistic signal (punctuation habits, morphology, spelling) was never represented.
• Applies when: The task is text classification where classes are distinguished by authorial or stylistic properties rather than topic (e.g. same subject matter across all classes), and the metric rewards sharper class probabilities.
• Example:
• Input:
    tfidf = TfidfVectorizer(max_features=5000, ngram_range=(1, 2), min_df=2)
    X_tr = tfidf.fit_transform(train_df["text"])
    X_te = tfidf.transform(test_df["text"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
    pd.DataFrame(proba, columns=classes).to_csv("submission.csv", index=False)
• Consequence:
    Log loss on the held-out evaluation is ~0.44 higher (worse) than the same
    classifier fed word-plus-character-n-gram features; sub-word stylistic
    cues carried by punctuation and morphology never enter the model.
• Counter-example:
• Input:
    word_vec = TfidfVectorizer(ngram_range=(1, 2), min_df=2)
    char_vec = TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5))
    X_tr = hstack([word_vec.fit_transform(train_df["text"]),
                   char_vec.fit_transform(train_df["text"])])
    X_te = hstack([word_vec.transform(test_df["text"]),
                   char_vec.transform(test_df["text"])])
    model = LogisticRegression(max_iter=2000)
    model.fit(X_tr, y_train)
• Why it does not fire: A vectorizer with analyzer="char_wb" exists and its output is stacked into the feature matrix passed to the fitting call, satisfying steps 3 and 4.
C9 · discarded-signal Pixel-domain aggregates for compression-domain perturbation detection

P1 — Pixel-domain aggregates for compression-domain perturbation detection

• Pattern: Detects a pipeline for classifying subtle modifications hidden in a lossy-compressed image format that decodes files to spatial pixel arrays and derives only global aggregate statistics (means, standard deviations, histograms, gradient summaries), while never reading the compression-domain representation (quantized transform coefficients or quantization tables) where such modifications are embedded.
• Detection procedure:
1. Confirm the task context: the code reads image files with a lossy-compressed extension (e.g. files ending in .jpg or .jpeg) and produces per-image scores or labels for detecting subtle alterations (a binary or ranking target, not object recognition with many semantic classes).
2. Locate the feature-extraction function(s): any function that takes an image path or decoded array and returns a numeric feature vector used for model fitting or prediction.
3. Inside those functions, check that images are decoded to pixels via a spatial-domain loader such as Image.open(...) followed by np.array(...), or cv2.imread(...), or equivalent, and that all computed features are aggregates of pixel values or pixel differences: calls like np.mean, np.std, np.histogram, np.percentile, np.gradient, np.diff, or equivalent, reduced to scalars or short vectors.
4. Search the entire source for any access to the compression-domain representation: block transform calls such as dct, dctn, idct (or equivalent), or coefficient/table accessors such as quantization, coef_arrays, quant_tables, or a decoder invoked with a flag that returns raw coefficients.
5. The pattern is PRESENT if step 1 and step 3 hold and step 4 finds no compression-domain access anywhere in the code.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: Feature extraction of the form img = np.array(Image.open(path)); feats = [img.mean(), img.std(), *np.histogram(img, bins=16)[0]] fed to a classifier produced a weighted ranking metric of 0.582, versus 0.673 (~13% relative better) for an otherwise comparable pipeline that quantized block transform coefficients using the file's quantization table and built coefficient histograms; the perturbation lives in the quantized coefficients and is nearly invisible in global pixel aggregates.
• Applies when: The task is to detect subtle, embedded modifications in images stored in a lossy transform-coded format, and the modification is applied in (or is best observable in) the compression domain rather than as a visible spatial change.
• Example:
• Input:
    def extract_features(path):
        img = np.array(Image.open(path).convert('L'), dtype=np.float32)
        gx = np.abs(np.diff(img, axis=1)).mean()
        gy = np.abs(np.diff(img, axis=0)).mean()
        hist, _ = np.histogram(img, bins=16, range=(0, 255))
        return [img.mean(), img.std(), gx, gy, *hist]

    X = np.array([extract_features(p) for p in train_paths])
    clf = RandomForestClassifier(n_estimators=200).fit(X, y)
• Consequence:
    Held-out weighted AUC 0.582, barely above chance and ~0.09 absolute
    (~13% relative) below a baseline using quantized transform-coefficient
    histograms on the same data; ranking of positives is systematically
    poor because the embedding signal is absent from pixel aggregates.
• Counter-example:
• Input:
    def extract_features(path):
        im = Image.open(path)
        qt = np.array(im.quantization[0]).reshape(8, 8)
        y = np.array(im.convert('L'), dtype=np.float32) - 128
        blocks = y[:504, :504].reshape(63, 8, 63, 8).transpose(0, 2, 1, 3)
        coefs = np.round(dctn(blocks, axes=(2, 3), norm='ortho') / qt)
        hist, _ = np.histogram(coefs, bins=np.arange(-8, 9))
        return hist.tolist()
• Why it does not fire: although it also decodes pixels, it reads the quantization table and quantizes block transform coefficients, so step 4 finds compression-domain access and the conjunction in step 5 fails.
C9 · discarded-signal Contiguous-prefix training subsample of a sequential data fileverified_trace · effect +0.1361

P1 — Contiguous-prefix training subsample of a sequential data file

• Pattern: Detects training-subset construction that reads only the first N records of a large sequentially-stored data file — stopping iteration once a count reaches a cap — instead of striding or randomizing across the whole file, so records and classes appearing later in the file are never seen by the model.
• Detection procedure:
1. Locate a loop that iterates over records of a data file or streaming decoder (e.g., a BSON/CSV/JSON-lines iterator, for record in reader: or equivalent) and appends per-record features or labels to lists later used for fitting a model.
2. Within that loop body, find a counter variable incremented once per record and a conditional of the form if counter >= LIMIT: break (or a while counter < LIMIT loop condition, or slicing such as itertools.islice(reader, LIMIT) / nrows=LIMIT on the read call) where LIMIT is far smaller than the file's record count.
3. Confirm that no other statement in the same loop or in the record source skips records by a stride (e.g., if i % step != 0: continue), draws random offsets, or shuffles record order before the cap is applied.
4. PRESENT when the fitting data comes only from the first LIMIT records read in file order (steps 1–2) and no stride/shuffle/random selection intervenes (step 3).
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A training loop of the form if count >= max_items: break over a sequential record stream fitted a classifier only on the file's earliest records; a variant that strided across the entire file at the same sample size scored 0.97 higher on the evaluation metric because classes occurring only later in the file were otherwise unpredictable.
• Applies when: Training data is stored in one large sequential file with many classes, and the pipeline can only afford to fit on a subset of the records.
• Example:
• Input:
    X, y = [], []
    count = 0
    for record in iter_records("train_data.bin"):
        X.append(extract_features(record))
        y.append(record["label"])
        count += 1
        if count >= 50000:
            break
    model.fit(np.asarray(X), y)
• Consequence:
    The model is fitted only on classes present in the first 50k records of a
    multi-million-record file stored in grouped order; any class whose records
    all occur later has zero recall, and overall accuracy drops sharply versus
    a strided sample of identical size (observed metric gap ~0.97).
• Counter-example:
• Input:
    X, y = [], []
    step = max(1, total_records // 50000)
    for i, record in enumerate(iter_records("train_data.bin")):
        if i % step != 0:
            continue
        X.append(extract_features(record))
        y.append(record["label"])
        if len(X) >= 50000:
            break
    model.fit(np.asarray(X), y)
• Why it does not fire: the i % step != 0: continue stride spreads the 50k sampled records across the entire file, so the cap terminates only after the whole file has been spanned and late-file classes remain represented.
C9 · discarded-signal Axis-only reductions for faint slanted-signal detectiongithub_occurrence

P1 — Axis-only reductions for faint slanted-signal detection

• Pattern: Detects handcrafted feature extraction over 2D time-frequency panels that reduces each panel only to global scalar statistics and single-axis 1D profiles, with no accumulation of intensity along any family of slanted candidate trajectories, when the target is a faint drifting line-like signal.
• Detection procedure:
1. Locate a function or loop body that receives a 2D array (or a stack of 2D arrays) per sample and produces a fixed-length numeric feature vector later passed to the fit method of a tabular classifier or regressor.
2. Within that function or loop body, list every array-reducing operation applied to the 2D data: calls such as mean, std, min, max, median, percentile, sum, argmax, ptp, or equivalent, either over the whole array or with a single-axis argument (axis=0 or axis=1).
3. Check the same function or loop body for any construct that integrates along non-axis-aligned paths: a loop or vectorized sweep over a slope/shift/drift parameter combined with per-row or per-column index offsetting (e.g., np.roll with a row-dependent shift, fancy indexing with a linearly increasing offset, an affine/rotation warp, or a Radon/Hough transform or equivalent).
4. Confirm from surrounding context (array shape usage, per-sample 2D loading, binary target column) that samples are 2D arrays scored for presence of a weak structured signal rather than aggregate-level statistics.
5. PRESENT when step 2 finds only whole-array or single-axis reductions, step 3 finds no slanted-path accumulation anywhere in the feature pipeline, and step 4 holds.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: Feature builder computing only arr.mean(), arr.std(), np.percentile(arr, q), and arr.sum(axis=0)-style profiles per panel yielded out-of-fold AUC ≈ 0.51 (chance level); the same model class fed features that integrated intensity along a grid of drift slopes reached ≈ 0.71 AUC on identical folds — a 0.28 absolute (28% relative) ranking loss.
• Applies when: Per-sample inputs are 2D arrays (e.g., spectrogram-like panels) and the label marks presence of a faint, narrow, possibly drifting structure whose per-pixel amplitude is at or below the noise floor; features are handcrafted rather than learned by a spatial model.
• Example:
• Input:
    def make_features(path):
        arr = np.load(path).astype(np.float32)   # shape (T, F)
        feats = []
        for panel in arr:
            feats += [panel.mean(), panel.std(),
                      np.percentile(panel, 99),
                      panel.max(axis=0).mean(),
                      panel.sum(axis=1).max(),
                      float(np.argmax(panel.mean(axis=0)))]
        return np.array(feats)

    X = np.stack([make_features(p) for p in paths])
    model.fit(X, y)
• Consequence:
    Out-of-fold ranking metric sits near chance (AUC ~0.51) because a faint
    drifting line contributes negligibly to every global or axis-aligned
    statistic; trajectory-integrated features on the same classifier reach
    ~0.71 AUC, a ~0.28 absolute degradation with no error or warning emitted.
• Counter-example:
• Input:
    def make_features(path):
        arr = np.load(path).astype(np.float32)   # shape (T, F)
        feats = [arr.mean(), arr.std()]
        T = arr.shape[0]
        for slope in np.linspace(-4, 4, 17):
            shifted = np.stack([np.roll(arr[t], int(slope * t))
                                for t in range(T)])
            profile = shifted.sum(axis=0)         # line integrals per slope
            feats += [profile.max(), profile.max() / (profile.std() + 1e-6)]
        return np.array(feats)
• Why it does not fire: alongside the scalar summaries, the loop over slope with a row-dependent np.roll accumulates intensity along slanted candidate trajectories, satisfying step 3 and breaking the required conjunction.
C9 · discarded-signal Grayscale collapse discards chroma signal in subtle-signal image taskgithub_occurrence

P1 — Grayscale collapse discards chroma signal in subtle-signal image task

• Pattern: Detects a feature-extraction pipeline for detecting subtle or hidden signals in color images that converts each image to a single grayscale/luminance channel before computing features, discarding the chroma channels that may carry part of the target signal.
• Detection procedure:
1. Identify a function or loop that opens image files and returns numeric feature vectors later used to fit or apply a classifier (e.g. arrays passed to a gradient-boosting, linear, or tree-model fitting call, or equivalent).
2. Within that function or loop, find a color-to-single-channel conversion applied to every image before feature computation: a call to .convert('L') on a PIL image, cv2.cvtColor(..., cv2.COLOR_BGR2GRAY), cv2.imread(..., 0) / cv2.IMREAD_GRAYSCALE, indexing or averaging that reduces a 3-channel array to one channel (e.g. img[..., 0], img.mean(axis=2)), or equivalent.
3. Confirm that no other code path in the feature extraction computes features from more than one channel of the decoded image (no per-channel loop, no multi-channel color-space conversion such as to YCbCr with statistics taken on each plane).
4. Confirm the task context is detection of a weak or embedded signal in color inputs (labels distinguish altered vs. unaltered images, or the metric is an AUC-style discrimination score on color image files).
5. PRESENT when all images are reduced to one channel before feature extraction (steps 2–3) in such a detection context (step 4).
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: Feature extractor called Image.open(path).convert('L') and computed all statistics on the single grayscale plane; an otherwise equivalent pipeline extracting per-channel features from a luma/chroma decomposition scored ~0.024 higher on the competition's weighted-AUC metric.
• Applies when: Classical (non-deep) feature extraction over color image files for a discrimination task where the target signal is subtle and may be embedded across color channels.
• Example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('L')
        arr = np.asarray(img, dtype=np.float32)
        resid = arr - scipy.ndimage.median_filter(arr, size=3)
        return np.array([resid.mean(), resid.std(),
                         np.abs(resid).mean(), (resid**2).mean()])

    X = np.vstack([extract_features(p) for p in train_paths])
    model = lgb.train(params, lgb.Dataset(X, label=y))
• Consequence:
    Model trains and predicts cleanly, but the discrimination metric
    (weighted AUC) is ~1.4 points absolute lower than the same pipeline
    computing residual statistics on each channel of a YCbCr decomposition,
    because signal embedded in the chroma planes never reaches the model.
• Counter-example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('YCbCr')
        arr = np.asarray(img, dtype=np.float32)
        feats = []
        for c in range(3):
            plane = arr[..., c]
            resid = plane - scipy.ndimage.median_filter(plane, size=3)
            feats += [resid.mean(), resid.std(), np.abs(resid).mean()]
        return np.array(feats)
• Why it does not fire: The image is decoded to a multi-channel color space and features are computed per channel, so no chroma information is discarded before feature extraction.
C9 · discarded-signal CV scores computed but ensemble averaged with fixed weightsgithub_occurrence

P1 — CV scores computed but ensemble averaged with fixed weights

• Pattern: Detects code that evaluates each of several candidate models with cross-validation, then combines their held-out predictions with an unconditional mean or hard-coded constant weights that never reference the computed validation scores.
• Detection procedure:
1. Locate a loop or repeated block that iterates over two or more distinct fitted estimators, and within the same block, a call to a cross-validation scoring routine (e.g. cross_val_score, a manual K-fold scoring loop, or equivalent) whose result is assigned to a variable.
2. Verify that each such score variable is only passed to printing/logging calls (print, a logger method, string formatting) and is never assigned to any collection or variable that survives past its loop iteration.
3. Locate, after the loop, an expression that combines the per-model prediction arrays: a call to an elementwise mean over the collected predictions, or a sum/weighted sum whose coefficients are all numeric literals.
4. Confirm that no variable appearing in the combining expression of step 3 is derived (directly or through intermediate assignments) from any score variable found in step 1, and no conditional branch selects among models based on those scores.
5. PRESENT when steps 1–4 all hold: per-model validation scores exist in the program, and the final blended prediction is formed with weights that are literals or an unweighted mean.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline ran sc = cross_val_score(m, Xtr, ytr, ...) for each of two models, only printed sc.mean(), then blended with p = np.mean(preds, axis=0); a comparable pipeline that let validation performance drive the final predictions scored 0.17637 higher on the same competition metric.
• Applies when: The program trains two or more candidate models on the same task, computes a per-model cross-validation or hold-out score for each, and produces a single combined prediction for an evaluation set.
• Example:
• Input:
    models = {"a": model_a, "b": model_b}
    preds = []
    for name, m in models.items():
        sc = cross_val_score(m, Xtr, ytr, cv=5, scoring="roc_auc")
        print(name, "cv", sc.mean())
        m.fit(Xtr, ytr)
        preds.append(m.predict_proba(Xte)[:, 1])
    p = np.mean(preds, axis=0)
    sub["LABEL_A"] = p
    sub.to_csv("submission.csv", index=False)
• Consequence:
    When one member's CV score is materially lower than the other's (e.g. 0.55
    vs 0.72 AUC), the uniform average pulls the final test metric toward the
    weak member; observed drop of ~0.18 in the leaderboard metric versus a
    score-informed combination, with no error or warning emitted.
• Counter-example:
• Input:
    scores, preds = [], []
    for name, m in models.items():
        sc = cross_val_score(m, Xtr, ytr, cv=5, scoring="roc_auc").mean()
        scores.append(sc)
        m.fit(Xtr, ytr)
        preds.append(m.predict_proba(Xte)[:, 1])
    w = np.array(scores) - 0.5
    w = np.clip(w, 0, None); w = w / w.sum()
    p = np.average(preds, axis=0, weights=w)
• Why it does not fire: the combining expression's weights w are derived from the collected cross-validation scores, so step 4's requirement that scores never reach the blend is violated.
C9 · discarded-signal Hardcoded cap on rows sent through the fitted model at inference timegithub_occurrence

P1 — Hardcoded cap on rows sent through the fitted model at inference time

• Pattern: Detects an inference loop over test records that compares a running counter against a hardcoded numeric limit and, once the limit is reached, stops routing rows through the fitted model and instead writes a constant fallback label for every remaining row.
• Detection procedure:
1. Locate a loop that iterates over test/evaluation records and, inside its body, either calls a prediction method (predict, predict_proba, or equivalent) on features derived from the current record, or appends those features to a buffer that is later passed to a prediction method.
2. Within the same loop body, find a conditional whose test compares a counter variable (incremented inside the loop) against a name bound to a numeric literal, or against a numeric literal directly, where the prediction/buffering path from step 1 executes only when the counter is below the limit.
3. Confirm that when the condition in step 2 is false, the loop body still emits an output row for the record (e.g., a writerow call or append to a results list) using a value that does not depend on the current record's features — a constant, a precomputed majority label, or a single default.
4. Confirm the numeric limit is not derived from the total number of test records (e.g., it is not len(test) or a count computed from the test source), so records beyond the limit are expected to exist.
5. PRESENT when steps 1–4 all hold: prediction is gated by a hardcoded count and rows beyond the count receive a feature-independent constant output.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: An inference loop guarded by if processed_full < MAX_TEST_FULL_PROCESS predicted only the first slice of test records and wrote int(most_common_label) for all remaining rows; roughly 98% of test rows received the constant majority label, and the evaluation metric dropped from ~0.30 to ~0.013 relative to a version that predicted every row.
• Applies when: The program streams or iterates over a test set larger than the hardcoded limit and produces one prediction per record with a fitted model available in scope.
• Example:
• Input:
    MAX_ROWS = 50000
    processed = 0
    for rec in iter_records(test_path):
        rid = rec["_id"]
        if processed < MAX_ROWS:
            feat = extract_features(rec)
            if feat is not None:
                pred = model.predict([feat])[0]
                writer.writerow([rid, int(pred)])
                processed += 1
                continue
        writer.writerow([rid, int(most_common_label)])
• Consequence:
    Every test row beyond the first 50,000 receives the constant majority label;
    on a multi-million-row test set, accuracy collapses toward the
    majority-class baseline (observed metric fell from ~0.30 to ~0.013).
• Counter-example:
• Input:
    for rec in iter_records(test_path):
        rid = rec["_id"]
        feat = extract_features(rec)
        if feat is None:
            missing_ids.append(rid)
            continue
        buf_f.append(feat)
        buf_id.append(rid)
        if len(buf_f) >= 4096:
            flush_and_predict(buf_f, buf_id)
    flush_and_predict(buf_f, buf_id)
    for rid in missing_ids:
        writer.writerow([rid, fallback_cat])
• Why it does not fire: the 4096 literal only controls batch flushing — every decodable row still reaches the model, and the constant fallback is used solely for rows whose input genuinely failed to produce features, not rows beyond a count cap.
C9 · discarded-signal Handcrafted intensity features from unstandardized per-sample arraysverified_trace · effect +0.0887

P1 — Handcrafted intensity features from unstandardized per-sample arrays

• Pattern: Detects aggregate intensity statistics (mean, std, max, percentiles, channel contrasts) computed directly on each raw loaded sensor array to build a per-sample feature vector, with no per-sample or per-channel centering and scaling applied between loading and feature computation.
• Detection procedure:
1. Locate a loop or function that loads one array per sample identifier from files (a call to np.load, imread, read, or equivalent inside an iteration over sample ids).
2. Within that same loop or function body, identify calls that reduce the loaded array to scalar statistics — mean, std, max, min, percentile, median, sums of slices, or differences between slice aggregates — whose results are appended to a per-sample feature row.
3. Inspect all statements between the load call at step 1 and the first statistic call at step 2: check whether the loaded array is reassigned by an expression that both subtracts a per-array (or per-channel, via an axis= and keepdims= reduction on the same array) center statistic and divides by a per-array or per-channel spread statistic.
4. A scaler fitted on the assembled feature matrix after the loop (e.g., StandardScaler().fit_transform(X) or equivalent column-wise scaling) does NOT satisfy step 3, because it normalizes across samples per feature, not within each sample.
5. PRESENT if step 2 statistics are computed on the array from step 1 and no per-sample standardization from step 3 occurs between the load and the statistic calls.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A pipeline appending arr.mean(), arr.std(), arr.max() directly from each loaded .npy array, relying only on a dataset-level column scaler afterward, scored 0.31113 worse on the ranking metric than an otherwise comparable pipeline that standardized each sample before feature extraction; the raw-array features encoded per-sample gain and noise-floor variation instead of the target signal.
• Applies when: Building handcrafted statistical features from per-sample sensor/signal arrays on a task where each sample was individually normalized or has independently varying gain/noise floor (so absolute intensity levels carry no target information).
• Example:
• Input:
    def load_features(ids, folder):
        rows = []
        for i in ids:
            arr = np.load(os.path.join(folder, i + ".npy"))
            rows.append([arr.mean(), arr.std(), arr.max(),
                         np.percentile(arr, 99),
                         arr[0].mean() - arr[1].mean()])
        return np.array(rows)

    X = load_features(train_ids, train_dir)
    X = StandardScaler().fit_transform(X)  # column-wise only
    model.fit(X, y)
• Consequence:
    Validation ranking AUC collapses toward 0.5 (chance level); feature
    vectors are dominated by per-sample gain and noise-floor differences
    rather than the target, costing roughly 0.3 absolute on the ranking
    metric versus the same features computed after per-sample
    standardization. Runtime and exit status are unaffected.
• Counter-example:
• Input:
    def load_features(ids, folder):
        rows = []
        for i in ids:
            arr = np.load(os.path.join(folder, i + ".npy"))
            arr = arr - np.median(arr, axis=-1, keepdims=True)
            arr = arr / (arr.std(axis=-1, keepdims=True) + 1e-8)
            rows.append([arr.mean(), arr.std(), arr.max(),
                         np.percentile(arr, 99),
                         arr[0].mean() - arr[1].mean()])
        return np.array(rows)
• Why it does not fire: the loaded array is re-centered by a per-channel median and divided by a per-channel std between the load and every statistic call, satisfying the per-sample standardization check in step 3.
C9 · discarded-signal Single-scale detection statistics on spectrogram arraystrace_observed

P1 — Single-scale detection statistics on spectrogram arrays

• Pattern: Detects handcrafted signal-detection features (peak, line-integral, or column-aggregate statistics) computed from 2D spectrogram-like arrays only at native frequency resolution, with no additional copy of the same statistics computed on a smoothed or downsampled version of the array along the frequency axis.
• Detection procedure:
1. Locate a feature-extraction function or loop that loads 2D (or stacked 2D) numeric arrays and computes aggregate statistics along one axis — e.g. np.max, np.mean, np.sum, np.percentile with an axis= argument, or subtractions/ratios of such aggregates — and collects the results into a feature matrix later passed to a classifier or ranker fit call.
2. Within the same feature-extraction function (and any helper it calls), search for any smoothing or rescaling of the array applied before those aggregates: calls to uniform_filter1d, gaussian_filter, gaussian_filter1d, medfilt, convolve, correlate, a moving-average built from np.cumsum differences, pooling/downsampling via reshape-and-mean over blocks, or equivalent multi-scale constructs.
3. Check whether the aggregate statistics from step 1 are computed on more than one resolution of the array (native plus at least one filtered/pooled copy) and concatenated into the feature vector.
4. The pattern is PRESENT if step 1 holds, no construct from step 2 is applied to the array before the aggregation, and the feature vector contains statistics from only the single native-resolution array.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline extracting per-column peak statistics from raw 2D arrays (e.g. arr.max(axis=0) features only) and feeding them to a gradient-boosted classifier scored ~10% relative lower on the ranking metric than an otherwise-similar pipeline that additionally computed the same statistics on a uniform_filter1d-smoothed copy of each array; signals spanning several frequency bins had their matched-filter response diluted by adjacent noise bins at native resolution.
• Applies when: Handcrafted (non-deep-learning) feature extraction from 2D spectrogram-like or time–frequency arrays for a detection/classification task where target signal width in bins is unknown or variable.
• Example:
• Input:
    def features(path):
        arr = np.load(path).astype(np.float32)   # (T, F) spectrogram
        arr = (arr - arr.mean()) / (arr.std() + 1e-6)
        col = arr.mean(axis=0)                   # line integral per freq bin
        feats = [col.max(), col.mean(),
                 np.percentile(col, 99),
                 arr.max(), arr.max(axis=0).std()]
        return np.array(feats, dtype=np.float32)

    X = np.stack([features(p) for p in paths])
    model.fit(X, y)
• Consequence:
    Ranking metric (AUC) is ~10% relative lower than the same pipeline with
    duplicated statistics on a frequency-smoothed copy: signals wider than one
    frequency bin produce a peak response diluted by neighboring noise bins,
    so wide-signal positives score close to noise and are ranked below them.
• Counter-example:
• Input:
    from scipy.ndimage import uniform_filter1d

    def features(path):
        arr = np.load(path).astype(np.float32)
        arr = (arr - arr.mean()) / (arr.std() + 1e-6)
        feats = []
        for k in (1, 4, 16):
            a = arr if k == 1 else uniform_filter1d(arr, size=k, axis=1)
            col = a.mean(axis=0)
            feats += [col.max(), col.mean(), np.percentile(col, 99)]
        return np.array(feats, dtype=np.float32)
• Why it does not fire: the same detection statistics are computed at native resolution and on two frequency-smoothed copies of the array and concatenated, so step 3's multi-scale condition is satisfied.
C9 · discarded-signal Fitted results never reach the prediction path

P1 — Fitted results never reach the prediction path

• Pattern: Detects a script that computes training-derived statistics or fits a model into named identifiers, yet none of those identifiers are referenced in the code that constructs the values written to the prediction output, leaving every prediction a constant or empty placeholder.
• Detection procedure:
1. Identify identifiers assigned from operations on training data: calls to a .fit(...) method, aggregation calls such as .mean(), .std(), .median(), .groupby(...) (or equivalent), or dictionaries populated inside a loop over rows/records loaded from a training file (e.g. train.csv or a training annotation file).
2. Identify the code region that builds the rows or column values ultimately written to the output file (the loop or expression whose results are passed to to_csv, json.dump, or an equivalent write call producing the submission/predictions).
3. For each identifier from step 1, check whether it (or any identifier derived from it) is read anywhere inside the region from step 2, including via function calls made from that region.
4. Check whether the predicted value assigned in the region from step 2 is a literal constant (e.g. "", 0, a fixed string) or is computed only from the test-row identifier itself.
5. PRESENT if at least one identifier from step 1 exists, no identifier from step 1 is read in the region from step 2 (step 3 fails for all of them), and step 4 holds.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A script iterated over training annotations to build per-class statistics, then in the prediction loop assigned pred_string = "" for every test row; the fitted statistics were never referenced. The contrastive variant that injected those statistics into each prediction row scored the full metric difference (1.0) higher on the same task.
• Applies when: The script has a distinguishable training/statistics phase and a separate phase that writes per-row predictions to an output file.
• Example:
• Input:
    train = pd.read_csv("train.csv")
    stats = train.groupby("LABEL_A")["col_a"].mean().to_dict()

    rows = []
    for sample_id in sample_submission["Id"]:
        pred_string = ""  # placeholder, stats never consulted
        rows.append({"Id": sample_id, "PredictionString": pred_string})

    pd.DataFrame(rows).to_csv("submission.csv", index=False)
• Consequence:
    Every output row is empty; the submission scores at the trivial
    no-prediction baseline (metric near 0) despite the compute spent
    building `stats`, a large drop versus using the fitted statistics.
• Counter-example:
• Input:
    train = pd.read_csv("train.csv")
    stats = train.groupby("LABEL_A")["col_a"].mean().to_dict()

    rows = []
    for sample_id in sample_submission["Id"]:
        est = stats.get("LABEL_A", 0.0)
        pred_string = f"{est:.3f}"
        rows.append({"Id": sample_id, "PredictionString": pred_string})

    pd.DataFrame(rows).to_csv("submission.csv", index=False)
• Why it does not fire: the training-derived identifier stats is read inside the loop that builds the output rows, so the fitted results flow into the written predictions.
C9 · discarded-signal Prefix-only subsample of a sequentially-stored training fileverified_trace · effect +0.1361

P1 — Prefix-only subsample of a sequentially-stored training file

• Pattern: Detects a training subsample built by collecting only the first N records of a large sequentially-read data file (exiting the read loop once a fixed count is reached) with no stride, random skip, or reservoir replacement, so ordered or class-clustered storage leaves most of the label space unrepresented in the sample.
• Detection procedure:
1. Locate a loop that iterates over a streaming reader of a training data file — e.g., for row in csv.reader(f), for line in f, for doc in decode_file_iter(f), or equivalent record-by-record iteration — whose body appends features and/or labels to containers that are later passed to a model-fitting call (fit, partial_fit, or equivalent) within the same module.
2. Within that loop body, find a counter variable incremented once per collected record and a conditional break (or a loop condition) that terminates iteration when the counter reaches a fixed integer constant.
3. Within the same loop body, check for any of: a modulus-based skip such as if i % k != 0: continue, a per-record random acceptance test (a comparison against a call to a random-number function guarding the append), or an index-replacement assignment into an already-full container (reservoir sampling). None is present.
4. PRESENT when all hold: the loop feeds training containers (step 1), it exits early at a fixed record count (step 2), and no stride, random-skip, or reservoir mechanism exists in the loop (step 3).
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A training loop of the form if n_collected >= MAX_TRAIN: break over a streaming file iterator sampled only the file prefix; on the same record budget, a version that skipped records at a computed stride across the whole file scored 0.956 higher on the competition metric, because classes stored beyond the prefix were never seen at fit time and could never be predicted.
• Applies when: A large training file is read record-by-record and only a subset of records fits the training budget, particularly in multiclass problems where the file may be sorted or clustered by label or by any label-correlated key.
• Example:
• Input:
    X, y = [], []
    n = 0
    with open("train.bin", "rb") as f:
        for rec in stream_records(f):
            feat = featurize(rec["payload"])
            if feat is not None:
                X.append(feat)
                y.append(rec["label"])
                n += 1
            if n >= 100000:
                break
    model.fit(np.asarray(X), np.asarray(y))
• Consequence:
    The sample covers only the first ~2% of the file; labels stored later never
    appear in the training set, so the fitted model assigns them zero probability.
    Held-out accuracy is bounded by the label mass of the file prefix — measured
    0.956 lower on the task metric than a stride-sampled model with the same
    100k-record budget and identical features.
• Counter-example:
• Input:
    X, y = [], []
    n, i = 0, 0
    stride = est_total_records // 100000
    with open("train.bin", "rb") as f:
        for rec in stream_records(f):
            i += 1
            if i % stride != 0:
                continue
            feat = featurize(rec["payload"])
            if feat is not None:
                X.append(feat); y.append(rec["label"]); n += 1
    model.fit(np.asarray(X), np.asarray(y))
• Why it does not fire: the loop skips records at a computed stride (if i % stride != 0: continue) so the collected sample spans the entire file, and there is no fixed-count early break cutting off the tail of the label distribution.
C9 · discarded-signal Constant placeholder written for every test instance despite training computation

P1 — Constant placeholder written for every test instance despite training computation

• Pattern: Detects a prediction-writing loop or assignment where the value stored for every test instance is an unconditional constant literal (such as an empty string, zero, or fixed placeholder) that never references any variable derived from training data, leaving all upstream fitting or statistics computation as dead code.
• Detection procedure:
1. Locate the code that produces the per-instance output values written to the submission or results file (e.g., a loop appending to a list later assigned to an output column, or a direct column assignment before a call to to_csv or equivalent).
2. For each branch of that code that assigns the per-instance value, check whether the assigned expression is a constant literal ("", 0, 0.0, a fixed string) with no reference to any variable, model object, dictionary, or array defined elsewhere in the program.
3. Elsewhere in the same program, identify at least one computation over training data: a call to a fitting method, an aggregation over rows loaded from a training file, or the construction of statistics (means, counts, cluster centers, priors) from training inputs.
4. Verify that none of the artifacts produced in step 3 (fitted objects, learned tables, aggregated values) appear in any expression from step 2, directly or through intermediate variables.
5. PRESENT when every per-instance output value is an unconditional constant literal (step 2 holds for all branches) AND the program computes training-derived artifacts (step 3) that never flow into the output expressions (step 4).
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A submission loop assigned pred_string = "" in both branches of an if/else, so every row of the output received an empty prediction; the program's earlier loading and processing of training annotations was never referenced. The measured metric was exactly 0.0, versus a nonzero score for a comparable program that projected training-derived aggregate boxes into each test frame.
• Applies when: The program performs a prediction task, computes anything from training data (fits a model or aggregates statistics), and writes a per-instance output file such as a submission CSV.
• Example:
• Input:
    train_df = pd.read_csv("train.csv")
    priors = train_df.groupby("col_a").size() / len(train_df)  # computed, never used

    sub = pd.read_csv("sample_submission.csv")
    predictions = []
    for row_id in sub["Id"]:
        if row_id in known_ids:
            pred = ""   # placeholder
        else:
            pred = ""
        predictions.append(pred)
    sub["Prediction"] = predictions
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Evaluation metric collapses to the trivial-baseline floor (0.0), strictly
    below what even a crude training-derived prior would score, because no
    learned information reaches the output file.
• Counter-example:
• Input:
    train_df = pd.read_csv("train.csv")
    prior = train_df["col_a"].mean()

    sub = pd.read_csv("sample_submission.csv")
    predictions = []
    for row_id in sub["Id"]:
        if row_id in known_ids:
            predictions.append(prior)   # training-derived value
        else:
            predictions.append("")      # fallback only for unknown ids
    sub["Prediction"] = predictions
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: at least one branch of the per-instance assignment references a training-derived artifact (prior), so the constant appears only as a fallback rather than the unconditional value for every instance.
C9 · discarded-signal Per-group history collapsed to a single summary-statistic targetverified_trace · effect +0.0563

P1 — Per-group history collapsed to a single summary-statistic target

• Pattern: Detects longitudinal training pipelines that aggregate each subject's repeated measurements into one derived summary value (such as a per-group fitted slope or rate) and fit the model on one row per group, instead of regressing on every (baseline, timepoint) measurement row.
• Detection procedure:
1. Locate a training frame (e.g. read from train.csv) that contains a group identifier column and a time-like column, where multiple rows share the same group identifier value.
2. Find a group-wise aggregation over that identifier — a call to .groupby(...) followed by .apply(...), .agg(...), or a loop over groupby groups — whose body computes a single scalar per group from the time and measurement columns (e.g. a call to np.polyfit(...), a per-group least-squares fit, a difference of last and first values divided by elapsed time, or equivalent).
3. Trace the resulting one-row-per-group structure: its scalar column is passed as the target argument y to a model-fitting call (model.fit(X, y) or equivalent), and no model-fitting call in the script receives a target built from the original multi-row-per-group measurement column.
4. PRESENT when steps 1–3 all hold: the only fitted target is the per-group scalar summary and the model never trains on individual measurement rows.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline that computed one slope per subject via groupby(id).apply(lambda g: np.polyfit(g["week"], g["val"], 1)[0]) and regressed on those scalars scored 0.056 worse on the evaluation metric than a pipeline that trained on every (baseline, timepoint) pair with a time-delta feature and predicted the measurement delta directly.
• Applies when: The task is longitudinal forecasting with multiple timestamped measurements per group identifier, and a supervised model is trained to predict future measurement values.
• Example:
• Input:
    train = pd.read_csv("train.csv")  # cols: id, week, val, col_a

    def slope(g):
        return np.polyfit(g["week"], g["val"], 1)[0]

    per_id = train.groupby("id").apply(slope).rename("slope").reset_index()
    per_id = per_id.merge(train.groupby("id").first().reset_index(), on="id")
    X = per_id[["col_a"]]
    y = per_id["slope"]
    model = LinearRegression().fit(X, y)
    pred_val = base_val + model.predict(X_test) * weeks_ahead
• Consequence:
    Training set shrinks from ~N_measurements rows to ~N_subjects rows (a
    reduction by the average measurements-per-subject factor), and the model
    fits noisy per-subject slope estimates; forecast error on the evaluation
    metric is ~0.056 worse than regressing on all measurement rows directly.
• Counter-example:
• Input:
    train = pd.read_csv("train.csv")  # cols: id, week, val, col_a
    base = train.sort_values("week").groupby("id").first().reset_index()
    base = base.rename(columns={"week": "base_week", "val": "base_val"})
    pairs = train.merge(base[["id", "base_week", "base_val"]], on="id")
    pairs["dw"] = pairs["week"] - pairs["base_week"]
    X = pairs[["base_val", "dw", "col_a"]]
    y = pairs["val"] - pairs["base_val"]
    model = Ridge().fit(X, y)
• Why it does not fire: although a groupby(...).first() aggregation appears, it only extracts baseline features; the fitted target y is built from the full multi-row-per-group measurement column, so the model trains on one row per (baseline, timepoint) pair.
C9 · discarded-signal Word-tokens-only features for a stylistic-identity targetverified_trace · effect +0.2170

P1 — Word-tokens-only features for a stylistic-identity target

• Pattern: Detects a text-classification pipeline whose target is authorial or stylistic identity that builds its feature matrix exclusively from word-level token features, with no character-level or sub-word n-gram features extracted anywhere in the program.
• Detection procedure:
1. Confirm the task is stylistic/authorial identity: the label column being fitted identifies who wrote the text (e.g., an author/writer identifier), as evidenced by column names, comments, or the mapping from text rows to a small set of author labels.
2. Enumerate every feature-extraction construct applied to the raw text column: constructions of TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent tokenizing/vectorizing calls, including any manual n-gram builders.
3. For each construct found in step 2, inspect its analyzer argument (or equivalent granularity setting): record whether any of them uses 'char', 'char_wb', or otherwise iterates over characters/sub-word units rather than whole tokens.
4. Check whether any feature matrices are horizontally stacked (e.g., hstack or equivalent) with a matrix produced by a character-level extractor from step 3.
5. PRESENT if step 1 holds, at least one word-level extractor exists in step 2, and steps 3–4 find no character-level or sub-word feature source anywhere in the pipeline.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only a word-level TfidfVectorizer(...) before fitting linear and ensemble classifiers scored 0.39222 worse on the log-loss-style competition metric than an otherwise comparable pipeline that added character n-gram features and combined them via hstack with the word features; orthographic and morphological style cues were never available to the weaker model.
• Applies when: The program performs supervised text classification where the label is the author or stylistic source of each passage, and features are built with sparse bag-of-tokens vectorizers rather than a pretrained language model.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # target: which author wrote each sentence
    vec = TfidfVectorizer(ngram_range=(1, 2), max_features=50000)
    X_train = vec.fit_transform(train["text"])
    X_test = vec.transform(test["text"])

    model = LogisticRegression(C=5, max_iter=1000)
    model.fit(X_train, train["author_label"])
    preds = model.predict_proba(X_test)
• Consequence:
    Program runs cleanly, but multiclass log loss is materially higher
    (observed gap ~0.39 on the evaluation metric) than the same pipeline
    augmented with character n-gram features, because character-level
    style signal (spelling, punctuation habits, morphology) is never
    presented to the model.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack
    from sklearn.linear_model import LogisticRegression

    vec_w = TfidfVectorizer(ngram_range=(1, 2), max_features=50000)
    vec_c = TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 5))
    X_train = hstack([vec_w.fit_transform(train["text"]),
                      vec_c.fit_transform(train["text"])]).tocsr()
    X_test = hstack([vec_w.transform(test["text"]),
                     vec_c.transform(test["text"])]).tocsr()
    model = LogisticRegression(C=5, max_iter=1000).fit(X_train, train["author_label"])
• Why it does not fire: A character-level extractor (analyzer="char_wb") is present and its output is stacked with the word features, so step 5's conjunction fails.
C9 · discarded-signal Style signal discarded by stop-word removal plus capped word-only vocabulary

P1 — Style signal discarded by stop-word removal plus capped word-only vocabulary

• Pattern: Detects a text-classification pipeline where the discriminative signal is authorship or writing style, yet the only text featurization is a word-level bag/TF-IDF built with stop-word removal and a hard vocabulary size cap, with no character-level n-gram features anywhere in the pipeline.
• Detection procedure:
1. Identify the task as text classification over raw text where the labels correspond to authors, writers, or stylistic sources (e.g. the target column names or comments indicate author/style identification, or the model is fitted on a text column mapped to per-source labels).
2. Locate every text-vectorization constructor in the source (e.g. TfidfVectorizer, CountVectorizer, or equivalent). For each, record its keyword arguments.
3. Check that every such constructor uses word-level tokenization (no analyzer='char' or analyzer='char_wb' or equivalent character-level setting anywhere in the file).
4. Check that at least one such constructor passes both a stop-word removal argument (e.g. stop_words='english' or a non-empty stop-word list) and a hard vocabulary cap (e.g. max_features= with a literal integer).
5. Check that the outputs of these vectorizers are the only text-derived features passed to any model-fitting call (no concatenation with a second, character-level vectorizer via hstack, FeatureUnion, ColumnTransformer, or equivalent).
6. PRESENT when steps 1, 3, 4, and 5 all hold: a style-discrimination task is featurized exclusively by capped, stop-word-stripped word tokens with no character n-grams.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A pipeline using TfidfVectorizer(stop_words='english', max_features=10000) as its sole text featurization on an authorship-style task scored ~0.50 worse on the log-loss competition metric than a pipeline retaining the full function-word vocabulary and concatenating character n-gram features; stylistic markers live in function words and sub-word patterns that the weaker configuration deletes.
• Applies when: The program classifies raw text into classes distinguished by writing style or authorship, using bag-of-words / TF-IDF style featurization before a classical classifier.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # classify sentences by which writer produced them
    vec = TfidfVectorizer(stop_words='english', max_features=10000)
    X_tr = vec.fit_transform(train['text'])
    X_te = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, train['label'])
    pred = model.predict_proba(X_te)
• Consequence:
    Multiclass log loss roughly doubles (≈0.60 vs ≈0.30) relative to a pipeline
    that keeps function words and adds character n-gram features; probability
    estimates are systematically less confident on the correct class because
    the strongest stylistic markers were removed before fitting.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack
    from sklearn.linear_model import LogisticRegression

    word_vec = TfidfVectorizer(min_df=2, sublinear_tf=True)
    char_vec = TfidfVectorizer(analyzer='char_wb', ngram_range=(2, 5))
    X_tr = hstack([word_vec.fit_transform(train['text']),
                   char_vec.fit_transform(train['text'])])
    model = LogisticRegression(max_iter=2000)
    model.fit(X_tr, train['label'])
• Why it does not fire: the word vectorizer keeps the full min_df-filtered vocabulary with no stop-word removal or hard cap, and character n-gram features are concatenated in, so steps 3–5 all fail.
C9 · discarded-signal Text represented only by scalar surface statisticsgithub_occurrence

P1 — Text represented only by scalar surface statistics

• Pattern: Detects a supervised pipeline on a content-dependent text task where the model's feature matrix is built exclusively from scalar surface statistics of the text (character/word lengths, punctuation or marker counts) and no token-level lexical representation of the text is ever fitted or included.
• Detection procedure:
1. Identify columns of the loaded data that are passed to string-processing operations (e.g. .str.len(), .apply(len), .str.count(...), counting of characters such as ?, !, or newline markers) and whose values originate from free-text fields.
2. Collect every variable that is later passed as the feature argument to a model-fitting call (e.g. .fit(X, y), lgb.train, or equivalent).
3. Check whether any of those feature variables is produced, in whole or in part, by a text-vectorization or embedding operation applied to the text columns from step 1 — e.g. a call to fit_transform/transform on a bag-of-words, TF-IDF, hashing, or n-gram vectorizer, or an embedding/encoder model invoked on the raw text (or equivalent).
4. PRESENT if the feature variables from step 2 contain only scalar statistics derived from the text columns (step 1 constructs, aggregations of lengths/counts) and no output from step 3 is concatenated into or used as any feature matrix passed to a fitting call.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: Feature matrix built solely from constructs like df['col_a'].str.len() and df['col_a'].str.count('?') while an imported vectorizer was never applied; a contrastive run that added a shallow lexical representation improved the classification metric by ~0.039 (about 4% relative), showing nearly all discriminative content was discarded.
• Applies when: The task is supervised learning over inputs that include long free-text fields, and the target plausibly depends on the semantic content of the text rather than only on its length or formatting.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer  # imported, unused
    from sklearn.linear_model import LogisticRegression

    df = pd.read_csv('train.csv')
    df['len_a'] = df['col_a'].str.len()
    df['words_a'] = df['col_a'].str.split().str.len()
    df['excl_a'] = df['col_a'].str.count('!')
    X = df[['len_a', 'words_a', 'excl_a']].values

    model = LogisticRegression()
    model.fit(X, df['LABEL_A'])
    preds = model.predict_proba(pd.read_csv('test.csv').pipe(featurize))
• Consequence:
    Run completes and writes predictions, but held-out log loss is ~4%
    relative worse than the same model with even a shallow lexical
    representation added; predictions track only text length/style, so
    examples with similar lengths but opposite content are scored alike.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack

    df = pd.read_csv('train.csv')
    df['len_a'] = df['col_a'].str.len()
    df['excl_a'] = df['col_a'].str.count('!')

    vec = TfidfVectorizer(max_features=5000, ngram_range=(1, 2))
    X_text = vec.fit_transform(df['col_a'])
    X = hstack([X_text, df[['len_a', 'excl_a']].values])
    model.fit(X, df['LABEL_A'])
• Why it does not fire: The same scalar surface statistics are computed, but a fitted lexical vectorizer's output is concatenated into the feature matrix passed to the fitting call, so step 4's exclusivity condition fails.
C9 · discarded-signal World-frame targets regressed without pose transformgithub_occurrence

P1 — World-frame targets regressed without pose transform

• Pattern: Detects a model trained to predict object positions in absolute world coordinates directly from per-sample sensor or appearance features, while per-sample ego-pose/localization records (sensor-to-world translation and rotation) that are available in the input data are never used to transform targets or predictions.
• Detection procedure:
1. Find a model-fitting call (e.g., .fit(X, y) or equivalent) whose target array y is built from annotation position columns expressed in a global/world frame (columns holding absolute x, y, z object centers taken directly from the annotation table without subtraction of any per-sample pose values).
2. Find the feature matrix X for that same fitting call and confirm it is built only from per-sample sensor content (image pixels, point statistics, appearance descriptors) with no columns derived from an ego-pose or localization table.
3. Scan the whole script for any load of a pose/ego/localization table (a file or record set containing per-sample translation and rotation of the vehicle or sensor); if such a table exists in the provided data, check whether any of its numeric fields appear in arithmetic expressions (addition, subtraction, matrix multiply, trig-based rotation) applied to y before step 1 or to the model's predicted positions before they are written to the output file.
4. PRESENT if the target in step 1 is world-frame positions, the features in step 2 exclude pose-derived values, and no pose-field arithmetic from step 3 exists anywhere between annotation loading and output writing.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A regressor was fitted with model.fit(image_features, annotations[['x','y','z']]) where the targets were absolute world coordinates and the ego-pose table shipped with the data was never joined or applied; predicted locations were essentially uncorrelated with ground truth and the overlap-based precision metric stayed near zero, while a variant that expressed positions in the ego-relative frame and rotated/translated predictions back to world coordinates per sample scored 1.0 higher on the same metric.
• Applies when: The task requires predicting 3D object locations in a global/world coordinate frame, the ground-truth annotations are stored in that world frame, and the provided data includes per-sample ego-pose or localization records relating the sensor frame to the world frame.
• Example:
• Input:
    ann = pd.read_csv('train.csv')          # world-frame centers
    X = np.stack([hog_features(load_img(t)) for t in ann['token']])
    y = ann[['center_x', 'center_y', 'center_z']].values
    model = RandomForestRegressor().fit(X, y)

    test_X = np.stack([hog_features(load_img(t)) for t in test_tokens])
    preds = model.predict(test_X)            # written as-is, no pose used
    rows = [{'Id': t, 'Pred': fmt(p)} for t, p in zip(test_tokens, preds)]
    pd.DataFrame(rows).to_csv('submission.csv', index=False)
• Consequence:
    Predicted world positions cluster around the training-set mean location and
    are uncorrelated with true object positions, because vehicle pose — the
    dominant determinant of world coordinates — is absent from the features.
    Overlap-based average precision stays near 0.0 versus a pose-aware baseline.
• Counter-example:
• Input:
    ann = pd.read_csv('train.csv').merge(ego_pose, on='token')
    # express targets in ego-relative frame
    dx = ann['center_x'] - ann['ego_x']; dy = ann['center_y'] - ann['ego_y']
    c, s = np.cos(-ann['yaw']), np.sin(-ann['yaw'])
    y = np.stack([c*dx - s*dy, s*dx + c*dy, ann['center_z'] - ann['ego_z']], 1)
    model = RandomForestRegressor().fit(X, y)

    rel = model.predict(test_X)
    c, s = np.cos(test_yaw), np.sin(test_yaw)
    wx = test_ex + c*rel[:,0] - s*rel[:,1]
    wy = test_ey + s*rel[:,0] + c*rel[:,1]
• Why it does not fire: pose fields (ego_x, ego_y, yaw) are used arithmetically to transform targets into the ego frame before fitting and to rotate/translate predictions back to world coordinates, so step 3's condition fails.
C9 · discarded-signal Word-only features with stop words stripped for style attribution

P1 — Word-only features with stop words stripped for style attribution

• Pattern: Detects a style- or author-attribution text pipeline whose only text representation is a word-level vectorizer configured to remove stop words, with no character-level n-gram features constructed anywhere.
• Detection procedure:
1. Confirm the task is stylistic: the target column encodes an author, writer, or style identity rather than topical content (labels are drawn from a fixed small set of person/style identifiers and the input is raw text).
2. Locate every call that constructs a text feature extractor (e.g. TfidfVectorizer, CountVectorizer, or equivalent) whose output is later passed to a model-fitting call within the same script.
3. Check that at least one such constructor receives a stop-word removal argument (e.g. stop_words='english' or an explicit stop-word list) and that none of them receives a character-level analyzer argument (e.g. analyzer='char' or analyzer='char_wb', or an equivalent sub-word tokenization).
4. Check that no other feature source derived from character-level or sub-word units of the raw text (character n-gram counts, punctuation counts, character-level embeddings) is concatenated or stacked with the word features before model fitting.
5. PRESENT when steps 1–4 all hold: the model's only input is word tokens with function words removed and no character-level representation exists.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only TfidfVectorizer(stop_words='english', ngram_range=(1, 2)) scored 0.379 worse on the evaluation metric (log-loss) than an otherwise comparable pipeline that stacked a word vectorizer with a analyzer='char' n-gram vectorizer, because function-word usage, punctuation habits, and morphology are primary signals for distinguishing writing styles.
• Applies when: A classification task over short raw-text passages where the label reflects who wrote the text or its stylistic register, and features are built with bag-of-words style vectorizers.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(stop_words='english', ngram_range=(1, 2))
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author'])
    proba = model.predict_proba(X_test)
• Consequence:
    Multiclass log-loss is substantially higher (observed ~0.38 worse) than
    the same classifier fed word features stacked with character n-grams;
    function-word and punctuation signal that separates authors is discarded.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack
    from sklearn.linear_model import LogisticRegression

    word_vec = TfidfVectorizer(stop_words='english', ngram_range=(1, 2))
    char_vec = TfidfVectorizer(analyzer='char', ngram_range=(2, 5))
    X_train = hstack([word_vec.fit_transform(train['text']),
                      char_vec.fit_transform(train['text'])])
    X_test = hstack([word_vec.transform(test['text']),
                     char_vec.transform(test['text'])])
    LogisticRegression(max_iter=1000).fit(X_train, train['author'])
• Why it does not fire: Although the word vectorizer still removes stop words, a character-level n-gram vectorizer is constructed and stacked with the word features, so the sub-word and function-word stylistic signal is retained.
C9 · discarded-signal Contiguous middle-block slice sampling in stack feature extractiongithub_occurrence

P1 — Contiguous middle-block slice sampling in stack feature extraction

• Pattern: Detects per-stack feature extraction that reads slices only from a single contiguous window (e.g., a fixed-size block centered on the middle index) of an ordered slice list, so content near the ends of the stack never contributes to any feature.
• Detection procedure:
1. Locate a function that builds features for one sample by iterating over an ordered collection of slice/frame files or arrays (e.g., sorted file list from glob/os.listdir, or the first axis of a 3D array).
2. Within that function, find the expression that selects which slices are processed. Check whether it is a contiguous slice of the ordered list whose bounds are derived from the middle index or fixed offsets — e.g., files[mid - k : mid + k], files[n//2 - k : n//2 + k], or arr[c-k:c+k] — with no step argument greater than 1.
3. Confirm there is no other selection in the same function that spreads indices across the full extent, such as np.linspace(0, n - 1, k).astype(int), a slice with a computed stride like files[::n // k], or an explicit index loop covering both ends (or equivalent).
4. Confirm the per-slice values computed inside the loop are aggregated (mean, max, histogram, statistics) into the sample's feature vector.
5. PRESENT when the only slices contributing to the aggregated features come from the contiguous window found in step 2 and no full-extent selection from step 3 exists.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A per-volume feature extractor selected files[mid - 15 : mid + 15] from the sorted slice list before computing intensity statistics; a comparable pipeline sampling indices with an even spread across the whole stack scored ~0.10 higher on the evaluation metric, since target regions located outside the central block contributed nothing to the contiguous-window features.
• Applies when: The program aggregates hand-crafted per-slice features over an ordered stack of slices/frames (volumetric or sequential imagery) into one feature vector per sample.
• Example:
• Input:
    def extract_features(case_dir):
        files = sorted(os.listdir(case_dir))
        mid = len(files) // 2
        chosen = files[mid - 10 : mid + 10]
        stats = []
        for f in chosen:
            img = load_slice(os.path.join(case_dir, f))
            stats.append([img.mean(), img.std(), img.max()])
        return np.array(stats).mean(axis=0)
• Consequence:
    Features encode only the central ~20 slices; samples whose informative
    regions lie near the top or bottom of the stack produce feature vectors
    indistinguishable from background, lowering the classification metric by
    roughly 10% relative to full-extent strided sampling.
• Counter-example:
• Input:
    def extract_features(case_dir):
        files = sorted(os.listdir(case_dir))
        idx = np.linspace(0, len(files) - 1, 20).astype(int)
        stats = []
        for i in idx:
            img = load_slice(os.path.join(case_dir, files[i]))
            stats.append([img.mean(), img.std(), img.max()])
        return np.array(stats).mean(axis=0)
• Why it does not fire: The selected indices are spread evenly from the first to the last slice via np.linspace, so every region of the stack can contribute to the aggregated features.
C9 · discarded-signal Per-sample min-max rescaling before intensity-statistic featurestrace_observed

P1 — Per-sample min-max rescaling before intensity-statistic features

• Pattern: Detects per-sample rescaling of each image or slice by its own minimum and maximum before the feature extraction step computes only intensity statistics (means, standard deviations, percentiles, threshold fractions), erasing the cross-sample absolute-intensity differences those statistics exist to capture.
• Detection procedure:
1. Locate a function or loop body that processes one image, slice, or volume at a time (e.g., iterates over file paths or array entries and loads pixel data into an array such as img).
2. Within that same function or loop body, find an arithmetic expression that subtracts the array's own minimum and divides by its own range or maximum — e.g., (img - img.min()) / (img.max() - img.min()), img / img.max(), or equivalent — where both the minuend/dividend and the min/max come from the same per-sample array.
3. After that expression, in the same function or loop body, identify the feature values collected for the sample, and confirm they consist only of scalar intensity statistics of the rescaled array: calls such as .mean(), .std(), np.percentile(...), np.median(...), or fractions of pixels above/below a numeric threshold (e.g., (img > 0.5).mean()), with no spatial, textural, or shape descriptors computed from raw values.
4. Confirm these per-sample statistics are assembled into a feature matrix later passed to a model-fitting call.
5. The pattern is PRESENT when the per-sample min-max rescaling (step 2) occurs before the statistic extraction (step 3) and no intensity statistic is computed on the un-rescaled pixel values.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline computing img = (img - img.min()) / (img.max() - img.min() + 1e-6) per sample and then extracting only mean, std, and percentile features produced a held-out ranking metric near chance (~0.50 AUC), while the same modeling pipeline using intensity statistics on raw pixel values scored ~0.22 higher, because per-sample rescaling forces every sample's feature distribution onto the same [0,1] scale.
• Applies when: An image-classification or image-regression task builds handcrafted per-sample features from pixel intensities, and those features are limited to global intensity statistics rather than spatial or learned representations.
• Example:
• Input:
    feats = []
    for path in image_paths:
        img = load_pixels(path).astype(np.float32)
        img = (img - img.min()) / (img.max() - img.min() + 1e-6)
        feats.append([
            img.mean(),
            img.std(),
            np.percentile(img, 90),
            (img > 0.5).mean(),
        ])
    X = np.array(feats)
    model.fit(X, y)
• Consequence:
    Held-out ranking metric collapses toward chance (AUC ~0.50 vs ~0.65 for the
    identical pipeline computing the same statistics on raw pixel values):
    every sample is forced onto the same [0,1] intensity scale, so between-sample
    absolute-brightness differences — the only signal these global statistics
    carry — are removed before the model ever sees them.
• Counter-example:
• Input:
    feats = []
    for path in image_paths:
        img = load_pixels(path).astype(np.float32)
        feats.append([
            img.mean(),
            img.std(),
            np.percentile(img, 90),
            (img > img.mean()).mean(),
        ])
    X = np.array(feats)
    scaler = StandardScaler().fit(X[train_idx])
    model.fit(scaler.transform(X[train_idx]), y[train_idx])
• Why it does not fire: the intensity statistics are computed on raw pixel values with no per-sample min-max rescaling before extraction; normalization happens only afterward via a scaler fitted across samples, preserving cross-sample intensity differences.
C9 · discarded-signal Small fixed vocabulary cap on sparse text features fed to sparse-capable models

P1 — Small fixed vocabulary cap on sparse text features fed to sparse-capable models

• Pattern: Detects a sparse bag-of-words or TF-IDF text vectorizer constructed with a small fixed vocabulary-size cap (a few thousand or fewer features) whose output is consumed only by linear or probabilistic models that scale to high-dimensional sparse input, discarding the long tail of discriminative terms.
• Detection procedure:
1. Locate a call constructing a sparse text feature extractor (e.g., TfidfVectorizer, CountVectorizer, or equivalent bag-of-words/TF-IDF transformer).
2. Check that the constructor call includes a keyword argument max_features (or equivalent vocabulary-size cap) set to a literal integer less than or equal to 10000.
3. Trace the variable holding the extractor's fit_transform/transform output (or equivalent) and confirm every model fitted on it is a linear classifier/regressor or a naive Bayes variant (e.g., LogisticRegression, SGDClassifier, MultinomialNB, LinearSVC, or equivalent), with no downstream model that requires dense or low-dimensional input.
4. PRESENT when a literal cap ≤ 10000 from step 2 exists and all consumers in step 3 are sparse-capable linear/probabilistic models.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A text-classification pipeline using TfidfVectorizer(max_features=5000) feeding logistic regression and naive Bayes scored 0.42 worse on the log-loss-style metric than an otherwise similar pipeline with an uncapped vocabulary, because rare but discriminative n-grams were dropped from the feature space.
• Applies when: The program vectorizes free text into sparse features (bag-of-words or TF-IDF) and the resulting matrix is used only by linear or naive Bayes models, so high dimensionality carries negligible memory or fitting cost.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression
    from sklearn.naive_bayes import MultinomialNB

    vec = TfidfVectorizer(max_features=5000, ngram_range=(1, 2))
    X_train = vec.fit_transform(df_train['col_a'])
    X_test = vec.transform(df_test['col_a'])
    lr = LogisticRegression(max_iter=1000).fit(X_train, y)
    nb = MultinomialNB().fit(X_train, y)
    pred = (lr.predict_proba(X_test) + nb.predict_proba(X_test)) / 2
• Consequence:
    Validation log loss rises materially (observed ~0.4 absolute worsening,
    ~42% relative) versus the same pipeline with no vocabulary cap: the
    truncated vocabulary silently discards infrequent but highly
    discriminative n-grams the linear models could have exploited.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(ngram_range=(1, 2), min_df=2, max_df=0.9)
    X_train = vec.fit_transform(df_train['col_a'])
    X_test = vec.transform(df_test['col_a'])
    lr = LogisticRegression(max_iter=1000).fit(X_train, y)
    pred = lr.predict_proba(X_test)
• Why it does not fire: The vectorizer has no fixed max_features cap — it prunes only by document-frequency thresholds, keeping the full long tail of discriminative terms available to the linear model.
C9 · discarded-signal Axis-aligned-only features for a drifting-line detection targettrace_observed

P1 — Axis-aligned-only features for a drifting-line detection target

• Pattern: Detects a feature extractor for a 2D-array detection task, whose target is described as a line translating across one axis as a function of the other, that computes only fixed-axis aggregates and gradient statistics and never performs a shift-and-integrate over a range of candidate drift rates.
• Detection procedure:
1. Confirm the task context (comments, docstrings, task description loaded in the script, or variable names such as spectrogram, cadence, waterfall, drift) indicates each sample is a 2D array in which the target signal moves across one axis as the index of the other axis increases.
2. Locate the function(s) that convert each 2D array into a feature vector fed to a classifier. Record every aggregation call in those functions (np.mean, np.std, np.max, np.percentile, np.sum, np.gradient, or equivalent) and note whether each is applied with a constant axis= argument or over the flattened array.
3. Search the same function(s) — and any helper they call — for any construct that shifts rows (or columns) by an amount proportional to the row (or column) index before summing: a loop or vectorized expression over a set of candidate rates in which np.roll, fancy indexing with an index offset computed from the loop variable times the row index, scipy.ndimage.shift, an affine/rotation transform, a Hough or Radon transform, or equivalent is applied, followed by a reduction along the shifted axis.
4. PRESENT when step 1 holds, step 2 finds only constant-axis or whole-array aggregations, and step 3 finds no shift-proportional-to-index construct followed by a reduction anywhere in the feature-extraction path.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: Feature extraction consisting solely of calls like np.mean(arr, axis=0), np.std(arr, axis=1), and gradient-magnitude statistics fed to a tabular classifier scored 0.233 lower on the ranking metric than a pipeline that, for a grid of candidate slopes, index-shifted each row proportionally, summed along the shifted axis, and used contrast statistics of the dedrifted profiles as features.
• Applies when: A supervised detection task over per-sample 2D arrays (e.g., time–frequency or trajectory images) where the stated target is a faint line that drifts/translates across one axis, and features are hand-computed rather than learned by a 2D convolutional model.
• Example:
• Input:
    def extract_features(arr):  # arr: 2D array, target line drifts across columns over rows
        feats = []
        feats.append(arr.mean())
        feats.append(arr.std())
        feats.extend(arr.mean(axis=0))          # per-column average
        feats.extend(arr.max(axis=0))           # per-column peak
        gy, gx = np.gradient(arr)
        feats.append(np.abs(gx).mean())
        feats.append(np.abs(gy).mean())
        return np.array(feats)

    X = np.stack([extract_features(np.load(p)) for p in paths])
• Consequence:
    Energy of the drifting line is smeared across many columns instead of
    accumulated along its trajectory, so faint targets stay at noise level in
    every feature; the ranking metric (AUC) lands ~0.23 below an otherwise
    identical pipeline whose features integrate along candidate drift slopes.
• Counter-example:
• Input:
    def extract_features(arr, rates=range(-8, 9)):
        feats = [arr.mean(), arr.std()]
        n_rows = arr.shape[0]
        for r in rates:
            shifted = np.stack([np.roll(arr[i], -r * i) for i in range(n_rows)])
            profile = shifted.sum(axis=0)       # integrate along candidate slope
            feats.append(profile.max() / (profile.std() + 1e-9))
        return np.array(feats)
• Why it does not fire: the extractor shifts each row by an amount proportional to its index over a grid of candidate rates and reduces along the shifted axis, so step 3 finds a shift-and-integrate construct and the conjunction in step 4 fails.
C9 · discarded-signal Stop-word removal in authorship-style classification

P1 — Stop-word removal in authorship-style classification

• Pattern: Detects a text-feature extraction step that discards function words via a built-in or explicit stop-word list when the prediction target is the writer or style of the text rather than its topic.
• Detection procedure:
1. Identify the prediction target: locate the column or array passed as the label to a classifier's fitting method; confirm from surrounding code (label values used as output columns, comments, or variable names such as author, writer, style) that the task is to identify who wrote the text or its stylistic origin, not its subject matter.
2. Locate a bag-of-words or tf-idf vectorizer construction (e.g. TfidfVectorizer(...), CountVectorizer(...), or equivalent) whose transformed output is fed, directly or after concatenation, to the classifier from step 1.
3. Within that constructor call, find a stop-word argument set to a non-empty value: stop_words='english', stop_words= a list or set literal containing common function words, or an equivalent parameter in another library; alternatively, a preprocessing loop that removes tokens found in a stop-word collection before vectorization.
4. Confirm no second feature source in the same pipeline retains function words (e.g. an additional character-level or unfiltered word-level vectorizer whose output is stacked with the filtered one).
5. PRESENT when the label is an author/style identity (step 1), the features feeding the classifier come from the stop-word-filtered vectorizer (steps 2–3), and no parallel unfiltered feature source exists (step 4).
• Predicted impact:
• Add score for C9: 3
• weight: 3
• confidence: high
• Evidence: A vectorizer configured with stop_words='english' in an author-identification pipeline produced a multi-class log loss roughly 0.34–0.40 worse (absolute) than otherwise similar pipelines that retained all tokens; function-word frequencies are among the strongest per-author discriminators, and filtering them removes that signal entirely.
• Applies when: A text-classification pipeline where the label identifies the author, writer, or stylistic source of each document, and features are bag-of-words or tf-idf counts over the text.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vectorizer = TfidfVectorizer(max_features=50000,
                                 stop_words='english',
                                 ngram_range=(1, 2))
    X_tr = vectorizer.fit_transform(train['text'])
    X_te = vectorizer.transform(test['text'])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, train['author'])   # target: who wrote it
    proba = model.predict_proba(X_te)
• Consequence:
    Multi-class log loss increases substantially (observed ~0.34-0.40 absolute,
    ~34% relative worse) because high-frequency function words — the strongest
    per-author style markers — are excluded from the feature space, so the
    classifier's predicted probabilities are systematically less confident and
    less accurate on every fold.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vectorizer = TfidfVectorizer(max_features=50000,
                                 stop_words='english',
                                 ngram_range=(1, 2))
    X_tr = vectorizer.fit_transform(train['text'])
    X_te = vectorizer.transform(test['text'])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, train['topic'])   # target: subject category
    proba = model.predict_proba(X_te)
• Why it does not fire: the prediction target is the document's topic, not its author or style, so removing function words does not discard the task's primary signal and step 1 fails.
C9 · discarded-signal Word-only features for style attribution

P1 — Word-only features for style attribution

• Pattern: Detects an authorship or style text-classification pipeline that builds its feature matrix exclusively from word-level token vectorizers, with no character-level n-gram features extracted anywhere before model fitting.
• Detection procedure:
1. Confirm the task is stylistic or authorship classification of text: the label column identifies a writer, source, or style, and raw text strings are the model input.
2. Locate every text-vectorization construct in the source (e.g., TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent) that is fitted or applied to the text column.
3. For each such construct, check its analyzer setting: it is word-level if it has no analyzer argument, or has analyzer='word'; it is character-level if it has analyzer='char' or analyzer='char_wb' (or an equivalent sub-word tokenization option in another library).
4. Check whether any feature-combination step (e.g., hstack, FeatureUnion, ColumnTransformer, or equivalent concatenation of feature matrices) merges output from a character-level construct into the matrix passed to the model's fitting call.
5. PRESENT if step 1 holds, at least one word-level vectorizer feeds the model, and no character-level vectorizer output reaches any model fitting call in the script.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline vectorizing text with only a default word-level TfidfVectorizer before fitting classifiers scored 0.42 worse (relative log-loss gap) on the competition metric than an otherwise similar pipeline that used hstack to concatenate word-level and character-level n-gram features, because punctuation habits, morphology, and spelling signal were discarded.
• Applies when: The program performs supervised text classification where labels correspond to authors, sources, or writing styles, and features are derived from raw text via bag-of-words-style vectorization.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    vec = TfidfVectorizer(max_features=20000, ngram_range=(1, 2))
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
• Consequence:
    Held-out multiclass log loss is substantially higher (observed ~0.42
    worse) than the same pipeline augmented with character n-gram features,
    because sub-word stylistic cues (punctuation, spelling, morphology)
    never enter the feature matrix.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from scipy.sparse import hstack
    from sklearn.linear_model import LogisticRegression

    word_vec = TfidfVectorizer(ngram_range=(1, 2))
    char_vec = TfidfVectorizer(analyzer='char_wb', ngram_range=(2, 5))
    X_train = hstack([word_vec.fit_transform(train['text']),
                      char_vec.fit_transform(train['text'])]).tocsr()
    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)
• Why it does not fire: A character-level vectorizer (analyzer='char_wb') is fitted and its output is stacked into the matrix passed to the model's fitting call, so sub-word signal is used.
C9 · discarded-signal Global-scale division after per-channel baseline in peak-feature normalizationtrace_observed

P1 — Global-scale division after per-channel baseline in peak-feature normalization

• Pattern: Detects normalization of a 2D signal array that subtracts a per-channel location statistic computed along one axis but then divides by a single scalar dispersion statistic computed over the whole array, before extremum-based features (max, peak, top-k) are extracted from the result.
• Detection procedure:
1. In a function body, find an expression that subtracts from a 2D (or higher) array a statistic computed with an explicit axis argument — e.g. arr - arr.mean(axis=1, keepdims=True), arr - np.median(arr, axis=1, keepdims=True), or equivalent — producing a baseline-removed array.
2. In the same function body, find that this baseline-removed array (or the original array) is divided by a dispersion statistic — std, mad, interquartile range, or equivalent — computed WITHOUT an axis argument, or with the result reduced to a single scalar (e.g. arr.std(), np.std(arr), arr.std() + eps).
3. Downstream in the same function or module, the normalized array is passed to an extremum-style reduction — max, np.max, argmax, np.partition/np.sort followed by selection of the largest entries, top-k, or equivalent — whose output is stored as a feature or score.
4. PRESENT when all three hold: per-axis location subtraction (step 1), scalar/whole-array dispersion division (step 2), and extremum-based feature extraction from the result (step 3).
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline normalizing 2D signal panels with x = x - np.median(x, axis=1, keepdims=True) followed by x = x / (x.std() + 1e-6) before taking max/peak features scored ~0.10 lower on the ranking metric than an otherwise-matched pipeline that divided each channel by its own dispersion; high-variance interference channels dominated the extremum features in the weaker run.
• Applies when: 2D spectrogram-like or multi-channel time-series arrays are normalized as a feature-extraction step, and detection features are derived from maxima, peaks, or top-k values of the normalized array.
• Example:
• Input:
    def features(x):  # x: (channels, time)
        x = x - np.median(x, axis=1, keepdims=True)
        x = x / (x.std() + 1e-6)          # single scalar scale
        f1 = x.max()
        f2 = np.sort(x.ravel())[-5:].mean()
        return np.array([f1, f2])
• Consequence:
    Extremum features saturate on the few channels with intrinsically high
    noise variance instead of the target signal; ranking metric (AUC) drops
    by roughly 0.10 relative to per-channel scaling on the same downstream model.
• Counter-example:
• Input:
    def features(x):  # x: (channels, time)
        x = x - np.median(x, axis=1, keepdims=True)
        s = x.std(axis=1, keepdims=True) + 1e-6
        x = x / s                          # per-channel scale
        f1 = x.max()
        f2 = np.sort(x.ravel())[-5:].mean()
        return np.array([f1, f2])
• Why it does not fire: the dispersion divisor is computed with an explicit axis argument and broadcast per channel, so step 2's scalar-division condition is not met.
C9 · discarded-signal Predictions from label priors only; raw per-sample inputs never read

P1 — Predictions from label priors only; raw per-sample inputs never read

• Pattern: Flags a prediction-generating script that computes every output value from aggregate statistics of the training labels or scene metadata (means, modes, class frequencies, pose records) while never reading any of the provided per-sample raw input files (point clouds, images, signal arrays) that carry the target-determining signal.
• Detection procedure:
1. Confirm the script writes a predictions or submission artifact: a call to to_csv, to_parquet, json.dump, or an open(..., "w") write of per-sample predictions, or equivalent.
2. Enumerate every file-reading operation in the script: calls to read_csv, read_json, json.load, open(..., "r"), np.fromfile, np.load, image-loading calls, point-cloud loaders, or equivalent.
3. Check whether any operation from step 2 targets a per-sample raw input file — a file whose path is constructed per sample id or iterated from a directory of sensor/measurement files (binary arrays, images, per-sample records) — as opposed to a single table of labels, ids, or metadata.
4. Trace the values written per sample in step 1: check whether each one is derived only from (a) constants, (b) aggregates over a labels/metadata table (e.g., .mean(), .mode(), .value_counts(), most-common class), or (c) per-sample metadata fields such as pose or timestamp — with no dependence on any quantity loaded from a per-sample raw input file.
5. PRESENT when the artifact from step 1 is written, no per-sample raw input file is read (step 3 finds none), and all predicted values satisfy step 4.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A submission generator built each sample's prediction as fixed geometric offsets around ego pose with global mean box dimensions and the most common class (f"{confidence} {wx:.3f} {wy:.3f} ... {most_common_class}"), never opening any point-cloud or image file; its competition score was the full metric range below a comparable script that conditioned on per-sample statistics, matching a constant-prediction baseline.
• Applies when: The task provides per-sample raw sensor or feature files (point clouds, images, per-sample measurement arrays) that determine the target, and the script under review produces the final predictions for evaluation.
• Example:
• Input:
    import pandas as pd
    train = pd.read_csv("train_labels.csv")
    meta = pd.read_csv("sample_metadata.csv")
    mean_w = train["col_w"].mean()
    mean_h = train["col_h"].mean()
    top_class = train["label"].mode()[0]
    sub = pd.read_csv("sample_submission.csv")
    sub["Prediction"] = f"1.0 {mean_w:.3f} {mean_h:.3f} {top_class}"
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Evaluation score sits at the constant-prediction baseline, orders of
    magnitude below any model that reads the per-sample sensor files;
    every sample receives the same prior-derived output regardless of
    its actual content.
• Counter-example:
• Input:
    import numpy as np, pandas as pd
    train = pd.read_csv("train_labels.csv")
    mean_w = train["col_w"].mean()
    sub = pd.read_csv("sample_submission.csv")
    preds = []
    for sid in sub["Id"]:
        pts = np.fromfile(f"samples/{sid}.bin", dtype=np.float32).reshape(-1, 4)
        cx, cy = pts[:, 0].mean(), pts[:, 1].mean()
        preds.append(f"1.0 {cx:.3f} {cy:.3f} {mean_w:.3f}")
    sub["Prediction"] = preds
    sub.to_csv("submission.csv", index=False)
• Why it does not fire: The script loads a per-sample raw input file for each id (np.fromfile on a per-sample path) and the written values depend on quantities computed from that file, so step 3 finds a per-sample read and step 4 fails.
C9 · discarded-signal Prefix-truncated flattened pixel featuresgithub_occurrence

P1 — Prefix-truncated flattened pixel features

• Pattern: Detects an image feature extractor that flattens a downsampled pixel array and then keeps only a fixed-index prefix slice covering a fraction of the flattened values, so the feature vector describes only one spatial band of the image rather than the whole frame.
• Detection procedure:
1. Locate a call that produces a 1-D array from a 2-D or 3-D pixel array — .flatten(), .ravel(), or .reshape(-1) (or equivalent) — applied to an array derived from an image load/resize operation.
2. In the same expression or on the same variable within the same function body, locate a subscript slice with no start and a literal integer stop, e.g. [:N].
3. In the same function body, find the resize target dimensions (literal tuple such as (H, W) passed to a resize call, or equivalent), and compute the total flattened length T = H × W × (number of channels if the array is not converted to a single channel).
4. Confirm the sliced result is used as a per-sample feature vector: it is returned, appended to a list later stacked into a matrix, or passed to a model-fitting or transform call.
5. PRESENT when the literal stop N from step 2 is strictly less than T from step 3 and the condition in step 4 holds.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline building features via arr.flatten()[:N] with N a small fraction of the flattened resize output produced a classification log loss ~15% worse (0.654 vs 0.552) than a pipeline whose descriptor summarized the entire image.
• Applies when: Handcrafted image features are constructed from raw pixel arrays and fed to a classical classifier (no learned convolutional feature extractor).
• Example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('L').resize((32, 32))
        arr = np.array(img, dtype=np.float32) / 255.0
        return arr.flatten()[:256]  # 256 of 1024 values

    X = np.vstack([extract_features(p) for p in paths])
    model.fit(X, y)
• Consequence:
    Features encode only the top quarter of each image (row-major prefix),
    discarding 75% of the pixels; classification log loss rises ~15%
    (0.654 vs 0.552) compared to a spatially complete descriptor.
• Counter-example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('L').resize((16, 16))
        arr = np.array(img, dtype=np.float32) / 255.0
        return arr.flatten()[:256]  # 16*16 = 256, full array

    X = np.vstack([extract_features(p) for p in paths])
    model.fit(X, y)
• Why it does not fire: the literal stop equals the total flattened length (16 × 16 = 256), so the slice retains every pixel and the descriptor covers the whole image.
C9 · discarded-signal Paired-panel contrast reduced to global aggregates onlytrace_observed

P1 — Paired-panel contrast reduced to global aggregates only

• Pattern: Detects feature construction for paired signal-bearing/control panel data where the difference between the two groups is reduced solely to whole-array aggregate statistics, with no step that locates a peak in the difference and extracts per-panel values at that same location.
• Detection procedure:
1. In the source, find a feature-extraction function or loop that splits a multi-panel array into two groups by index (e.g., arr[0::2] and arr[1::2], or slicing by a group axis) and forms a difference such as on.mean(axis=0) - off.mean(axis=0) or an elementwise subtraction of the two group arrays.
2. Within the same function body, list every operation applied to that difference array before it is appended to the feature vector; check whether all of them are global reductions over the full array — calls such as .mean(), .std(), .max(), .min(), np.percentile(...) or equivalent with no retained index.
3. Within the same function body, check for any operation that computes a location from the difference or from a filtered version of it — np.argmax, np.unravel_index, index of a maximum along an axis, or equivalent — whose result is later used to subscript the individual panels or the per-group arrays.
4. PRESENT if step 2 finds only global reductions and step 3 finds no location-based extraction: the two groups are compared exclusively through whole-array aggregates of their difference.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: Feature code of the form d = on.mean(0) - off.mean(0) followed only by d.mean(), d.std(), d.max(), d.min() produced a ranking metric 0.28 lower than an otherwise similar pipeline that located the strongest response in the difference and emitted sorted per-panel values at that location; the localized contrast was averaged away across the full array.
• Applies when: The task provides paired observations (on-target panels interleaved or grouped with off-target/control panels) and the discriminative signal is expected to appear in the signal-bearing group only, concentrated at a small region of the array; features are hand-built scalars fed to a downstream classifier.
• Example:
• Input:
    def features(arr):                 # arr: (6, T, F); even panels on-target, odd off-target
        on = arr[0::2].mean(axis=0)
        off = arr[1::2].mean(axis=0)
        d = on - off
        return [d.mean(), d.std(), d.max(), d.min(),
                np.percentile(d, 99), np.percentile(d, 1)]

    X = np.stack([features(load_panels(i)) for i in ids])
    model.fit(X, y)
• Consequence:
    Out-of-fold ranking metric sits near chance (~0.55-0.60 AUC) instead of
    the ~0.80+ achievable from the same arrays: a narrow signal occupying a
    few cells of a T x F difference is diluted by a factor of roughly T*F in
    every global aggregate, so signal-bearing and control examples receive
    nearly identical feature vectors.
• Counter-example:
• Input:
    def features(arr):                 # arr: (6, T, F); even panels on-target, odd off-target
        on = arr[0::2].mean(axis=0)
        off = arr[1::2].mean(axis=0)
        d = uniform_filter1d(on - off, size=5, axis=0)
        r, c = np.unravel_index(np.argmax(d), d.shape)
        per_panel = np.sort(arr[:, r, c])
        return list(per_panel) + [d.max(), per_panel[-1] - per_panel[0]]
• Why it does not fire: step 3 finds a location computed from the difference (np.argmax + np.unravel_index) that is used to subscript the individual panels, so the on/off contrast is evaluated at the candidate signal's location rather than only through global aggregates.
C9 · discarded-signal Small hardcoded vocabulary cap on sparse text features for a linear modelverified_trace · effect +0.0795

P1 — Small hardcoded vocabulary cap on sparse text features for a linear model

• Pattern: Detects a text-vectorization step whose vocabulary is truncated by a small hardcoded feature limit (10,000 or fewer) before feeding a sparse linear or naive-Bayes classifier, discarding the long tail of rare discriminative terms.
• Detection procedure:
1. Locate a call constructing a bag-of-words or TF-IDF text vectorizer (e.g. TfidfVectorizer(...), CountVectorizer(...), or equivalent) whose output is later passed to a fit or fit_transform on raw text columns.
2. In that constructor's arguments, find a keyword argument named max_features (or the library's equivalent vocabulary-size cap) assigned a literal integer whose value is less than or equal to 10000.
3. Within the same script, find that the matrix produced by this vectorizer is passed to the fit method of a linear or naive-Bayes estimator (e.g. LogisticRegression, LinearSVC, SGDClassifier, MultinomialNB, or equivalent) — not a dense neural network or a model requiring dense low-dimensional input.
4. Confirm there is no subsequent step that re-expands the representation (no second uncapped vectorizer whose output is combined via hstack or equivalent with the capped one).
5. The pattern is PRESENT when steps 1–4 all hold: a literal cap ≤ 10000 truncates the vocabulary feeding a sparse-capable linear or naive-Bayes model with no compensating uncapped representation.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: TfidfVectorizer(max_features=5000, ngram_range=(1, 2), min_df=2, max_df=0.8) feeding a logistic regression on a multiclass text task; measured contrastive pairs showed the capped version scoring ~0.44 worse on the competition metric than an uncapped-vocabulary pipeline, because informative low-frequency n-grams were excluded from the representation.
• Applies when: A text classification or regression pipeline builds sparse bag-of-words/TF-IDF features and trains a linear or naive-Bayes model that handles high-dimensional sparse input cheaply.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    tfidf = TfidfVectorizer(max_features=5000, ngram_range=(1, 2), min_df=2)
    X_tr = tfidf.fit_transform(train_df["col_a"])
    X_te = tfidf.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    probs = model.predict_proba(X_te)
• Consequence:
    Multiclass log loss is substantially higher (~0.44 worse in measured
    runs) than with the full vocabulary, because rare but class-indicative
    n-grams are silently dropped from the feature space; the program still
    exits cleanly and writes predictions.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    tfidf = TfidfVectorizer(ngram_range=(1, 2), min_df=2, max_df=0.95)
    X_tr = tfidf.fit_transform(train_df["col_a"])
    X_te = tfidf.transform(test_df["col_a"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    probs = model.predict_proba(X_te)
• Why it does not fire: the vectorizer relies only on min_df/max_df pruning with no max_features cap, so the full discriminative vocabulary is retained for the sparse linear model.
C9 · discarded-signal World-frame position regression ignoring available ego posegithub_occurrence

P1 — World-frame position regression ignoring available ego pose

• Pattern: Detects a spatial-prediction pipeline that fits a regressor directly on absolute world-frame position targets and emits its raw outputs as world coordinates, while per-sample ego/agent pose metadata available in the input data is never loaded or applied to transform targets or predictions.
• Detection procedure:
1. Locate a call to a model-fitting method (e.g., .fit(X, y) or equivalent) where the target array y is built from columns or fields representing object center positions (names containing x, y, z, center, or translation) taken directly from the training annotation table with no arithmetic combining them with a per-sample translation vector or rotation matrix before fitting.
2. Locate the code that writes predictions to the output file, and confirm the model's predicted position values are formatted into the output with no intervening addition of a per-sample translation or multiplication by a per-sample rotation.
3. Scan the entire script for any read of a pose table (a file or record set whose name or fields include ego_pose, pose, translation+rotation, or equivalent per-sample agent-position metadata); confirm either no such read exists, or the values read are never used in the target construction of step 1 or the output path of step 2.
4. PRESENT if steps 1, 2, and 3 all hold: world-frame targets in, raw model outputs out, pose metadata unused.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: model.fit(features, train_df[['center_x','center_y','center_z']]) with predictions written verbatim to the submission while the provided per-sample pose table was never read; predicted volumes landed arbitrarily far from ground truth as the agent moved through the world, and the IoU-based mAP metric was exactly 0, versus 1.0 relative-frame gain for an otherwise similar pipeline that modeled agent-relative offsets and transformed back per sample.
• Applies when: The task requires predicting object positions or extents in a global/world coordinate frame, and the provided data includes per-sample agent or sensor pose metadata (translation plus rotation) relating the agent's local frame to the world frame.
• Example:
• Input:
    train = pd.read_csv('train.csv')
    X = extract_features(train)          # per-sample sensor features
    y = train[['cx', 'cy', 'cz']].values  # absolute world coordinates
    model = RandomForestRegressor().fit(X, y)

    X_test = extract_features(test)
    preds = model.predict(X_test)        # raw world-frame outputs
    for row, (px, py, pz) in zip(test.itertuples(), preds):
        out.append(f"1.0 {px:.3f} {py:.3f} {pz:.3f} {w} {l} {h} 0 LABEL_A")
• Consequence:
    Predicted boxes are anchored near the world-coordinate range seen in
    training; as the agent traverses new regions, predictions land tens to
    hundreds of meters from ground truth, IoU with every true box is 0, and
    the IoU-thresholded mAP score drops to exactly 0.
• Counter-example:
• Input:
    pose = load_pose_table('pose.json')   # per-sample translation + rotation
    rel = []
    for row in train.itertuples():
        p = pose[row.sample_id]
        rel.append(p['R'].T.dot(np.array([row.cx, row.cy, row.cz]) - p['t']))
    model = RandomForestRegressor().fit(X, np.array(rel))

    for row, r in zip(test.itertuples(), model.predict(X_test)):
        p = pose[row.sample_id]
        world = p['t'] + p['R'].dot(r)
        out.append(f"1.0 {world[0]:.3f} {world[1]:.3f} {world[2]:.3f} ...")
• Why it does not fire: The pose table is loaded and applied — targets are converted to the agent-relative frame before fitting (fails step 1) and predictions are transformed back to world coordinates per sample before output (fails steps 2 and 3).
C9 · discarded-signal Single-slice feature extraction from multi-file image seriestrace_observed

P1 — Single-slice feature extraction from multi-file image series

• Pattern: Detects per-case feature extraction that reads only the first file from each case's directory of slice files (after a plain listing or lexicographic sort), so features are computed from a single edge slice while all remaining slices in the series are discarded.
• Detection procedure:
1. Locate a function or loop that, for each case identifier, builds a list of image files in a per-case directory via os.listdir, glob, Path.iterdir, or equivalent, optionally followed by sorted(...) with no numeric key.
2. Within the same function or loop body, check whether the file list is reduced to exactly one element via a subscript with literal index 0 (e.g., files[0], sorted(files)[0]) or by breaking out of an iteration after the first file is read, before any image-loading call (pydicom.dcmread, imread, nib.load, or equivalent).
3. Confirm that the feature vector appended to the per-case feature matrix is derived only from that single loaded file, with no loop that reads additional files from the same list and no aggregation (mean, std, stack) across multiple slices.
4. The pattern is PRESENT when a per-case multi-file directory is reduced to its first file and the case's features are computed solely from that one file.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A pipeline computed handcrafted features from sorted(os.listdir(case_dir))[0] per case, using only the first slice of each series; a comparable pipeline sampling and aggregating multiple central slices scored roughly 0.08 higher on the ranking metric (AUC), i.e., the single-edge-slice version dropped toward chance.
• Applies when: Inputs are per-case folders containing many slice files forming a volumetric or sequential image series, and the pipeline builds classical/handcrafted feature vectors from the images.
• Example:
• Input:
    def extract_features(case_id):
        case_dir = os.path.join(DATA_DIR, case_id)
        files = sorted(os.listdir(case_dir))
        ds = pydicom.dcmread(os.path.join(case_dir, files[0]))
        img = ds.pixel_array.astype(np.float32)
        return np.array([img.mean(), img.std(), img.max(), img.min()])

    X = np.vstack([extract_features(c) for c in case_ids])
    model.fit(X, y)
• Consequence:
    Each case is represented by an edge slice that mostly contains background;
    the interior slices holding the discriminative signal never enter the
    feature matrix. Ranking metric (AUC) falls ~8% relative versus sampling
    multiple central slices, approaching chance-level performance.
• Counter-example:
• Input:
    def extract_features(case_id):
        case_dir = os.path.join(DATA_DIR, case_id)
        files = sorted(os.listdir(case_dir), key=lambda f: int(f.split('-')[-1].split('.')[0]))
        mid = len(files) // 2
        feats = []
        for f in files[mid - 3: mid + 4]:
            img = pydicom.dcmread(os.path.join(case_dir, f)).pixel_array.astype(np.float32)
            feats.append([img.mean(), img.std(), img.max(), img.min()])
        feats = np.array(feats)
        return np.concatenate([feats.mean(0), feats.std(0)])
• Why it does not fire: The file list is numerically sorted, multiple central slices are read in a loop, and per-slice features are aggregated across slices, so no reduction to a single first file occurs.
C9 · discarded-signal Constant empty prediction written for every test rowgithub_occurrence

P1 — Constant empty prediction written for every test row

• Pattern: Detects a submission-building loop that assigns a constant empty or hard-coded prediction value to every test row, so no statistic, model output, or data-derived quantity computed earlier ever reaches the written output file.
• Detection procedure:
1. Locate a loop that iterates over test identifiers (e.g. rows or a column of a sample-submission frame, a list of test ids) and appends or assigns one output record per identifier.
2. Within that loop body, find the variable stored into the prediction field of each record (the non-id field of the appended dict/row).
3. Check that on every control-flow branch inside the loop, this variable is assigned only a literal constant (empty string "", empty list, 0, or a fixed literal), with no assignment deriving from any variable computed outside the loop from training data, fitted models, or per-sample inputs.
4. Confirm the resulting collection is written to the submission/output file (via to_csv, to_json, a file write, or equivalent).
5. PRESENT when all branches assign only a literal constant prediction (step 3) and that constant-valued collection is what gets written (step 4), even though the script computes data-derived statistics or model outputs elsewhere.
• Predicted impact:
• Add score for C9: 3
• weight: 3
• confidence: high
• Evidence: A script computed per-class geometry statistics from training data, but its output loop set pred_string = "" on both branches of an if/else and wrote that for every test id; the run scored 0.0 on the competition metric with recall exactly zero, while a variant that emitted statistics-derived predictions per row scored measurably above zero.
• Applies when: The program assembles a per-test-id submission or output file for a structured-prediction, detection, or scoring task and elsewhere computes any data-derived quantity (statistics, fitted model, per-sample features).
• Example:
• Input:
    stats = compute_class_stats(train_df)  # data-derived, computed but unused below

    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        if sample_id in test_index:
            pred_string = ""
        else:
            pred_string = ""
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})

    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Consequence:
    Evaluation score collapses to the empty-submission baseline (metric = 0.0,
    recall exactly 0), versus a strictly positive score for any submission whose
    prediction strings are constructed from the computed statistics.
• Counter-example:
• Input:
    stats = compute_class_stats(train_df)

    predictions = []
    for sample_id in sample_submission['Id'].tolist():
        if sample_id in test_index:
            pred_string = build_prediction(stats, test_index[sample_id])
        else:
            pred_string = ""
        predictions.append({'Id': sample_id, 'PredictionString': pred_string})

    pd.DataFrame(predictions).to_csv(OUTPUT_PATH, index=False)
• Why it does not fire: One branch assigns the prediction from a call that consumes the training-derived stats and per-sample data, so not every control-flow path yields a literal constant, and computed signal flows into the written output.
C9 · discarded-signal Word-only text features for style-sensitive probabilistic classification

P1 — Word-only text features for style-sensitive probabilistic classification

• Pattern: Detects a text classification pipeline scored on a probabilistic metric that vectorizes documents exclusively with word-level token features, never constructing or combining any character n-gram representation, when the target reflects writing style or authorship rather than topic.
• Detection procedure:
1. Identify every call that constructs a text vectorizer (e.g. TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent) whose output is later passed to a model-fitting call within the same script.
2. For each such constructor call, inspect its keyword arguments: check whether analyzer is set to 'char' or 'char_wb' (or an equivalent option selecting character-level analysis in another library).
3. Check whether any feature matrices from multiple vectorizers are combined (e.g. via hstack, FeatureUnion, ColumnTransformer, or equivalent) before fitting, such that at least one combined source is character-level.
4. Check that the model's per-class probability outputs (e.g. from predict_proba or equivalent) are written to the prediction output, indicating a probabilistic scoring metric.
5. The pattern is PRESENT when every vectorizer found in step 1 uses word-level analysis (step 2 finds no character-level analyzer), no character-level feature source is combined in step 3, and step 4 confirms probabilistic outputs are produced.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only TfidfVectorizer(stop_words='english') word features with a linear classifier scored 0.399 worse on the log-loss-style metric than a pipeline that additionally built analyzer="char_wb", ngram_range=(2, 4) features and combined them via sparse hstack before fitting; character-level cues (affixes, punctuation habits) discriminating the classes were never exposed to the weaker model.
• Applies when: The task is text classification whose labels depend on fine-grained lexical or orthographic style (e.g. authorship, register, dialect) rather than topical content, and the pipeline fits a classical model on vectorized text and outputs class probabilities.
• Example:
• Input:
    vectorizer = TfidfVectorizer(max_df=0.95, lowercase=True,
                                 stop_words='english')
    X_tr = vectorizer.fit_transform(train["text"])
    X_te = vectorizer.transform(test["text"])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y)
    proba = model.predict_proba(X_te)
    sub = pd.DataFrame(proba, columns=classes)
    sub.insert(0, "id", test["id"].values)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Log loss on held-out data is substantially higher (worse) than a pipeline
    that also stacks character n-gram TF-IDF features; observed gap of ~0.40
    on the same data because sub-word stylistic signal (affixes, punctuation
    patterns) is never presented to the model.
• Counter-example:
• Input:
    word_vec = TfidfVectorizer(sublinear_tf=True)
    char_vec = TfidfVectorizer(analyzer="char_wb", ngram_range=(2, 4),
                               min_df=3)
    Xw_tr = word_vec.fit_transform(train["text"])
    Xc_tr = char_vec.fit_transform(train["text"])
    X_tr = hstack([Xw_tr, Xc_tr]).tocsr()
    model = LogisticRegression(C=4.0, max_iter=1000)
    model.fit(X_tr, y)
    proba = model.predict_proba(hstack([word_vec.transform(test["text"]),
                                        char_vec.transform(test["text"])]))
• Why it does not fire: A second vectorizer with analyzer="char_wb" is constructed and its features are combined via hstack before fitting, so step 2 finds a character-level analyzer and step 3 finds a combined character-level source.
C9 · discarded-signal Absolute-value regression ignores known per-subject baseline anchorverified_trace · effect +0.0603

P1 — Absolute-value regression ignores known per-subject baseline anchor

• Pattern: Detects a longitudinal forecasting pipeline where each test subject's exact baseline measurement is available at inference, yet a pooled regressor is fitted on the absolute target value and its raw outputs are written as the predictions, without reconstructing forecasts as the known baseline plus a predicted change.
• Detection procedure:
1. Identify a model-fitting call (e.g. .fit(X, y) or equivalent) where the target array y is assigned directly from a measurement column of the training frame (e.g. y = df["col_a"].values), not from a difference, ratio, or slope derived by subtracting or dividing by a per-subject baseline value.
2. Confirm the test-time data provides, for every subject, at least one exactly known measurement of that same column (e.g. the test frame contains one observed row per subject, or a baseline table is constructible from it).
3. Locate the prediction step (e.g. model.predict(X_test) or equivalent) and confirm its output is assigned to the submission/output column with at most clipping, rounding, or filling — no arithmetic that adds a per-subject baseline value or otherwise forces the prediction at the baseline time point to equal the known measurement.
4. PRESENT when all of: the fitted target is the absolute measurement (step 1), an exact per-subject baseline exists at inference (step 2), and predictions are emitted without being anchored to that baseline (step 3).
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: sub_df["FVC_pred"] = model.predict(X_sub) wrote pooled absolute-value regression outputs directly to the submission even though each test subject's exact baseline measurement was supplied; a variant that predicted a per-subject slope and reconstructed base_value + slope * (week - base_week) scored 0.02527 better on the competition metric.
• Applies when: the task asks for per-subject trajectories over a time axis, and the inference data includes one exactly observed value of the target for each subject (a baseline anchor).
• Example:
• Input:
    X = train[["age", "week", "col_a"]].values
    y = train["target"].values
    model = Ridge().fit(X, y)

    test["target_pred"] = model.predict(test[["age", "week", "col_a"]].values)
    sub = test[["row_id", "target_pred"]].rename(columns={"target_pred": "target"})
    sub["target"] = sub["target"].round(1)
    sub.to_csv("submission.csv", index=False)
• Consequence:
    Predictions at and near each subject's baseline week deviate from the
    exactly known measurement by the model's pooled residual error, adding
    avoidable error the metric penalizes; anchoring forecasts to the given
    baseline improved the competition metric by ~0.025.
• Counter-example:
• Input:
    base = train.groupby("subject").first()[["week", "target"]]
    train = train.join(base, on="subject", rsuffix="_base")
    train["slope"] = (train["target"] - train["target_base"]) / \
                     (train["week"] - train["week_base"]).replace(0, 1)
    model = Ridge().fit(train[["age", "col_a"]].values, train["slope"].values)

    tb = test.set_index("subject")
    dw = weeks - tb.loc[subjects, "week_base"].values
    pred = tb.loc[subjects, "target_base"].values + model.predict(Xt) * dw
• Why it does not fire: the regressor is fitted on a slope derived relative to the baseline (fails step 1), and predictions are reconstructed as the known baseline plus predicted change scaled by horizon (fails step 3), so zero-horizon forecasts equal the given measurement.
C9 · discarded-signal Global-only color histograms without spatially localized color statisticsverified_trace · effect +1.7687

P1 — Global-only color histograms without spatially localized color statistics

• Pattern: Detects handcrafted image-classification feature extraction that encodes color exclusively as whole-image per-channel histograms (concatenated with grayscale texture descriptors), while computing no spatially localized color statistic such as per-grid-cell channel means, so the location of colors within the frame is discarded.
• Detection procedure:
1. Locate a function or code block that converts image files into fixed-length numeric feature vectors later passed to a classifier's fitting method (e.g., fit on the resulting matrix, or equivalent).
2. Within that function or block, find at least one call computing a histogram over pixel channel values of the entire image (e.g., np.histogram on a full color channel array, or equivalent) whose output is concatenated into the feature vector.
3. Within the same function or block, check for any operation that partitions the image spatially before computing a color statistic: slicing or reshaping the color array into tiles/blocks followed by a per-tile reduction (mean, median, histogram) over color channels, or any resized small color image flattened into features.
4. Confirm that texture descriptors present (e.g., gradient-orientation histograms or equivalent) operate on a grayscale conversion of the image, so they carry no channel information.
5. The pattern is PRESENT when step 2 finds whole-image channel histograms, step 3 finds no spatially localized color statistic anywhere in the feature construction, and step 4 confirms the remaining descriptors are grayscale-only.
• Predicted impact:
• Add score for C9: 2
• weight: 2
• confidence: high
• Evidence: Feature vectors built as np.concatenate([hog_grayscale, hist_r, hist_g, hist_b]) with histograms taken over the full frame produced classification log loss roughly 4% relative worse than an otherwise identical pipeline that also appended coarse-grid per-channel means; a second variant of the same global-only design measured 0.095 worse on the same metric.
• Applies when: Image classification on color images using handcrafted (non-learned) features fed to a classical classifier, where the classes plausibly differ in where colors appear, not only in overall color distribution.
• Example:
• Input:
    def extract_features(path):
        img = np.asarray(Image.open(path).convert("RGB").resize((96, 96)))
        gray = rgb2gray(img)
        tex = hog(gray, pixels_per_cell=(12, 12), cells_per_block=(2, 2))
        hists = []
        for c in range(3):
            h, _ = np.histogram(img[:, :, c], bins=16, range=(0, 255), density=True)
            hists.append(h)
        return np.concatenate([tex] + hists).astype(np.float32)
• Consequence:
    Classification log loss is ~4% relative higher than an identical pipeline
    that additionally includes per-grid-cell channel means; ranking metric on
    held-out data degrades correspondingly because spatial color layout that
    discriminates the classes is never presented to the model.
• Counter-example:
• Input:
    def extract_features(path):
        img = np.asarray(Image.open(path).convert("RGB").resize((96, 96)))
        gray = rgb2gray(img)
        tex = hog(gray, pixels_per_cell=(12, 12), cells_per_block=(2, 2))
        hists = [np.histogram(img[:, :, c], bins=16, range=(0, 255), density=True)[0]
                 for c in range(3)]
        cells = img.reshape(4, 24, 4, 24, 3).mean(axis=(1, 3)).ravel() / 255.0
        return np.concatenate([tex] + hists + [cells]).astype(np.float32)
• Why it does not fire: the reshape(...).mean(...) computes per-grid-cell channel means, a spatially localized color statistic, so step 3 finds a spatial color feature and the conjunction fails.
C9 · discarded-signal Word-only text features for authorship/style prediction

P1 — Word-only text features for authorship/style prediction

• Pattern: Detects a text classification pipeline whose target represents a writer's identity or style, where the raw text is converted only into word-level token features and no character n-gram features are ever built or combined with them.
• Detection procedure:
1. Confirm the task setting: the label column being fitted takes values identifying who wrote each text (author/writer identifiers), and each training row pairs a raw text field with such a label.
2. Locate every text-vectorization construction in the source (e.g. TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent). For each, check the analyzer argument: it is absent (defaulting to word-level) or set to 'word'.
3. Search the entire source for any vectorizer constructed with analyzer='char' or analyzer='char_wb' (or an equivalent sub-word/character featurizer), and for any horizontal stacking of feature matrices (hstack or equivalent) that would combine such features with the word-level matrix.
4. PRESENT when step 1 holds, at least one word-level vectorizer from step 2 is fitted on the text and its output is passed to a model's fitting call, and step 3 finds no character-level vectorizer anywhere in the source.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A model fitted on only vectorizer = TfidfVectorizer(...) word features (with stop_words='english' further removing function words) scored multi-class log loss ~0.58, while an otherwise similar linear model additionally given analyzer='char' n-gram features via hstack scored ~0.36 — a ~38% relative degradation from discarding sub-word stylistic signal (punctuation habits, morphology, function-word patterns).
• Applies when: Text classification where the target is an author/writer identity or otherwise stylistic rather than topical, features are built by sparse bag-of-tokens vectorization, and a linear or similar sparse-input model is fitted on them.
• Example:
• Input:
    vec = TfidfVectorizer(max_features=20000, stop_words='english')
    X_tr = vec.fit_transform(train['text'])
    X_te = vec.transform(test['text'])

    le = LabelEncoder()
    y = le.fit_transform(train['author'])

    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y)
    proba = model.predict_proba(X_te)
• Consequence:
    Multi-class log loss ~0.58 instead of ~0.36 achievable with the same
    linear model when character n-gram features are stacked in — a ~38%
    relative worsening because punctuation, morphology, and function-word
    style cues are never presented to the model.
• Counter-example:
• Input:
    word_vec = TfidfVectorizer(ngram_range=(1, 2))
    char_vec = TfidfVectorizer(analyzer='char', ngram_range=(2, 4))

    X_tr = hstack([word_vec.fit_transform(train['text']),
                   char_vec.fit_transform(train['text'])]).tocsr()
    X_te = hstack([word_vec.transform(test['text']),
                   char_vec.transform(test['text'])]).tocsr()

    model = LogisticRegression(C=4, max_iter=1000)
    model.fit(X_tr, train['author_id'])
• Why it does not fire: A character-level vectorizer (analyzer='char') is constructed, fitted on the same text, and horizontally stacked with the word-level matrix before fitting, so step 3 finds sub-word features and the conjunction in step 4 fails.
C9 · discarded-signal Hardcoded spatial offsets instead of label-derived positions

P1 — Hardcoded spatial offsets instead of label-derived positions

• Pattern: Detects prediction of object locations by adding a fixed hardcoded list of numeric coordinate offsets to the ego/sensor pose, while the training annotations' position fields are never read or aggregated to determine where objects actually occur.
• Detection procedure:
1. Locate a literal collection (list or tuple of tuples/lists of numeric literals) of two or more coordinate offsets defined at module or function scope, e.g. offsets = [(16.0, 0.0), (8.0, 3.0), ...].
2. Within the loop that builds output predictions, confirm each element of that collection is combined arithmetically with the ego/sensor pose translation (addition, optionally after rotation by the pose yaw) to produce predicted box center coordinates.
3. Search the entire source file for any read of position/translation fields from the training annotation records (e.g. indexing a translation, center, or x/y/z field of loaded label data) that flows into an aggregation such as a mean, median, histogram, or clustering call whose result is used in the prediction loop.
4. PRESENT if steps 1 and 2 hold and step 3 finds no such flow — predicted centers come only from hardcoded literals plus the pose, even though annotation position data is available in the loaded label files.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A script defined offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)] and emitted boxes at tx + dx*cos(yaw) - dy*sin(yaw) around the ego pose, reading labels only for average box dimensions; the contrastive version that clustered annotation positions in the ego-relative frame scored the full metric difference (1.0) higher on the overlap-based evaluation metric.
• Applies when: The program produces spatial detection outputs (boxes or positions) scored by an overlap/IoU-based metric, and the provided training labels include object position coordinates.
• Example:
• Input:
    offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)]
    for sample_id in sub_df["Id"]:
        tx, ty, tz = ego_pose[sample_id]["translation"]
        yaw = pose_yaw(ego_pose[sample_id])
        parts = []
        for dx, dy in offsets:
            wx = tx + dx * math.cos(yaw) - dy * math.sin(yaw)
            wy = ty + dx * math.sin(yaw) + dy * math.cos(yaw)
            parts.append(f"1.0 {wx:.3f} {wy:.3f} {tz:.3f} {mw} {ml} {mh} {yaw} LABEL_A")
        preds.append(" ".join(parts))
• Consequence:
    Predicted box centers land at arbitrary fixed spots around the sensor that
    almost never overlap true objects at the required IoU thresholds; mean
    average precision drops to ~0.0 versus a baseline that places boxes at
    positions aggregated from the training annotations.
• Counter-example:
• Input:
    rel = []
    for ann in train_annotations:
        rx, ry, rz = to_ego_frame(ann["translation"], ego_pose[ann["sample"]])
        rel.append([rx, ry, rz])
    centers = KMeans(n_clusters=8).fit(np.array(rel)).cluster_centers_
    for sample_id in sub_df["Id"]:
        tx, ty, tz = ego_pose[sample_id]["translation"]
        yaw = pose_yaw(ego_pose[sample_id])
        for rx, ry, rz in centers:
            wx = tx + rx * math.cos(yaw) - ry * math.sin(yaw)
            wy = ty + rx * math.sin(yaw) + ry * math.cos(yaw)
            preds_add(wx, wy, tz + rz)
• Why it does not fire: the candidate positions combined with the ego pose are cluster centers computed from annotation translations, not a hardcoded literal offset list, so step 3's flow from label positions to the prediction loop exists.
C9 · discarded-signal Single positive variant selected per base item when multiple variant sources existverified_trace · effect +0.0338

P1 — Single positive variant selected per base item when multiple variant sources exist

• Pattern: Detects a training-set builder that, given several parallel directories of positive variants (one per generation method) for each base item, attaches only one variant per base chosen by index arithmetic or fixed selection, so the remaining variant files are never loaded into the training set.
• Detection procedure:
1. Locate where the script enumerates positive-class inputs: a list or collection of two or more directory names (or path prefixes) that each hold a variant of the same base items, e.g. variant_dirs = ["method_a", "method_b", "method_c"] or equivalent.
2. Within the loop that builds training examples over base items, check whether the variant directory used for base item i is chosen by an expression that yields exactly one directory per base — a modulo expression such as variant_dirs[i % len(variant_dirs)], a random single choice such as random.choice(variant_dirs), or a constant subscript such as variant_dirs[0] — rather than an inner loop over all directories.
3. Confirm that no other statement in the training-data construction phase iterates over the full variant-directory collection to add the non-selected variants of the same base item to the training features/labels.
4. PRESENT when steps 1–3 all hold: multiple variant sources exist, one is selected per base item, and the unselected variants are never appended to the training set.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: Training pairs built with alg = variant_dirs[i % 3] used one positive file per base item out of three provided; a comparable solution that loaded all three variants per base item scored ~0.17 higher on the weighted-AUC competition metric.
• Applies when: A binary detection task supplies, for each original item, multiple positive variants produced by different generation methods (one file per method in parallel directories), and the script constructs a tabular/feature training set from those files.
• Example:
• Input:
    variant_dirs = ["method_a", "method_b", "method_c"]
    X, y = [], []
    for i, name in enumerate(base_names):
        X.append(feats(os.path.join(DATA, "originals", name)))
        y.append(0)
        alg = variant_dirs[i % len(variant_dirs)]
        X.append(feats(os.path.join(DATA, alg, name)))
        y.append(1)
• Consequence:
    Only one third of the available positive files enter training; the model
    never sees all generation methods applied to the same content, and
    held-out weighted AUC drops by roughly 0.15-0.17 versus training on all
    provided variants.
• Counter-example:
• Input:
    variant_dirs = ["method_a", "method_b", "method_c"]
    X, y, group = [], [], []
    for i, name in enumerate(base_names):
        X.append(feats(os.path.join(DATA, "originals", name)))
        y.append(0); group.append(i)
        for alg in variant_dirs:
            X.append(feats(os.path.join(DATA, alg, name)))
            y.append(1); group.append(i)
• Why it does not fire: the inner loop iterates over the full variant-directory collection, so every provided positive variant of each base item is added to the training set (step 3 fails).
C9 · discarded-signal pixel-and-moment-only features for shape-dependent image classificationgithub_occurrence

P1 — pixel-and-moment-only features for shape-dependent image classification

• Pattern: Detects a classical (non-deep-learning) image classification pipeline whose feature extraction consists solely of raw resized pixel intensities and/or channel-wise moment statistics (mean, standard deviation, min, max, histograms of raw values), with no gradient-orientation, edge, or texture descriptor computed anywhere before the classifier is fitted.
• Detection procedure:
1. Confirm the script loads image files (e.g., via PIL.Image.open, cv2.imread, imageio.imread, or equivalent) and fits a classical classifier (e.g., RandomForestClassifier, LogisticRegression, SVC, gradient boosting, or equivalent) on features derived from those images, with no convolutional/deep model (no torch, tensorflow, keras, or pretrained-embedding import used for feature extraction).
2. Locate every function or code block whose return value or output array is later stacked/concatenated into the matrix passed to the classifier's fitting call. Within those blocks, list all operations applied to the image arrays.
3. Check whether any listed operation computes structural descriptors: calls such as hog, local_binary_pattern, canny, sobel, Sobel, Laplacian, HOGDescriptor, SIFT, ORB, daisy, graycomatrix, or equivalent, or any explicit finite-difference/gradient computation (e.g., np.gradient, np.diff on pixel axes) whose result feeds the feature matrix.
4. Check whether the only operations feeding the feature matrix are: resizing/flattening raw pixels, and/or per-channel aggregations of raw values (mean, std, min, max, median, percentile, histogram of intensities).
5. PRESENT if step 1 holds, step 3 finds no structural descriptor, and step 4 confirms features are exclusively raw intensities and/or channel-wise moment statistics.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline whose extractor produced only a downsampled pixel vector plus np.mean/np.std per channel, fed to a tree ensemble via clf.fit(features_scaled, y), scored ~0.16 worse on the log-loss metric than an otherwise comparable pipeline that added a gradient-orientation descriptor to the same feature set.
• Applies when: The task is image classification where class identity depends on object shape or texture, the solution uses classical machine learning (no deep or pretrained feature extractor), and features are hand-computed from image arrays.
• Example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('RGB').resize((32, 32))
        arr = np.asarray(img, dtype=np.float32) / 255.0
        pixels = arr.flatten()
        means = arr.mean(axis=(0, 1))
        stds = arr.std(axis=(0, 1))
        return np.concatenate([pixels, means, stds])

    X = np.vstack([extract_features(p) for p in train_paths])
    clf = RandomForestClassifier(n_estimators=300)
    clf.fit(scaler.fit_transform(X), y)
• Consequence:
    The classifier receives no shape or edge information, so held-out log loss
    is meaningfully higher (observed ~0.16 absolute, ~16% relative worse) than
    a same-cost pipeline that includes a gradient-orientation descriptor;
    accuracy on shape-distinguished classes drops correspondingly.
• Counter-example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('RGB').resize((64, 64))
        arr = np.asarray(img, dtype=np.float32) / 255.0
        gray = rgb2gray(arr)
        hog_feats = hog(gray, orientations=9, pixels_per_cell=(8, 8))
        means = arr.mean(axis=(0, 1))
        stds = arr.std(axis=(0, 1))
        return np.concatenate([hog_feats, means, stds])

    X = np.vstack([extract_features(p) for p in train_paths])
    clf = LogisticRegression(max_iter=2000)
    clf.fit(scaler.fit_transform(X), y)
• Why it does not fire: The feature block also computes a gradient-orientation descriptor (hog) that is concatenated into the matrix passed to the classifier, so step 3 finds a structural descriptor and the conjunction in step 5 fails.
C9 · discarded-signal Stop-word removal or tight vocabulary cap in stylometry featuresverified_trace · effect +0.1652

P1 — Stop-word removal or tight vocabulary cap in stylometry features

• Pattern: Detects a bag-of-words/TF-IDF vectorizer for a style- or author-attribution text classification task configured to remove stop words or to cap the vocabulary at a small fixed size, discarding the function-word and high-frequency-token signal that carries authorial style.
• Detection procedure:
1. Locate a construction of a text vectorizer that builds token-count or TF-IDF features (e.g. TfidfVectorizer(...) or CountVectorizer(...), or equivalent), whose output is later passed to a classifier fit call within the same script.
2. Confirm the prediction target is a stylistic/authorship-style label: the target column or class values identify who wrote the text or a per-writer/style category, rather than topical content (decidable from the target variable's source column, class-name constants, or submission column construction in the same script).
3. In the vectorizer's keyword arguments, check whether either holds: (a) a stop_words argument set to a non-None value (a language string or an explicit word list), or (b) a max_features argument set to a literal integer less than or equal to 20000.
4. The pattern is PRESENT if step 1 and step 2 hold and at least one condition in step 3 holds.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using a vectorizer configured with stop_words='english' and max_features=5000 for an author-attribution task scored ~0.60 log loss, versus ~0.35 for an otherwise comparable pipeline retaining stop words with an uncapped word+character n-gram representation — a ~42% relative degradation on the same data.
• Applies when: The task is text classification where the label reflects who wrote the text (or another stylistic category), and features are built with a bag-of-words or TF-IDF vectorizer.
• Example:
• Input:
    vec = TfidfVectorizer(stop_words='english', max_features=5000)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author_label'])
    pred = model.predict_proba(X_test)
    pd.DataFrame(pred, columns=classes).to_csv('submission.csv', index=False)
• Consequence:
    Multiclass log loss degrades substantially (observed ~0.60 vs ~0.35 for the
    same pipeline without stop-word removal and vocabulary capping): function
    words like "the", "upon", "which" — the strongest per-author discriminators —
    are stripped from the feature space, so predicted probabilities are less
    confident and less accurate on every fold.
• Counter-example:
• Input:
    vec = TfidfVectorizer(ngram_range=(1, 2), sublinear_tf=True)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, train['author_label'])
    pred = model.predict_proba(X_test)
    pd.DataFrame(pred, columns=classes).to_csv('submission.csv', index=False)
• Why it does not fire: The vectorizer passes no stop_words argument and no small max_features cap, so high-frequency function-word features that carry the stylistic signal are retained.
C9 · discarded-signal Interior-only slice window in volume subsamplingtrace_observed

P1 — Interior-only slice window in volume subsampling

• Pattern: Detects per-volume slice-index sampling whose lower and upper bounds are hardcoded fractional offsets into the stack (e.g., 25%–75% of the slice count), so the outer portions of every volume are excluded from feature extraction.
• Detection procedure:
1. Locate a loop or function that processes a collection of image files or 2D arrays belonging to one volume (e.g., a sorted list of per-slice files or a 3D array's first axis).
2. Within that scope, find the expression that produces the slice indices to read, such as a call to np.linspace, range, or an equivalent index generator, or a slicing expression on the sorted list.
3. Check whether the start bound is computed as the slice count multiplied by a literal fraction strictly greater than 0 (e.g., int(n * 0.25), len(files) // 4) AND the end bound is the slice count multiplied by a literal fraction strictly less than 1 (e.g., int(n * 0.75)), with no alternative branch that samples outside this window.
4. Confirm the resulting indices are used to load or select the slices from which features or statistics are computed, and no other code path in the same pipeline reads slices outside the window.
5. PRESENT when steps 3 and 4 both hold: features are computed only from a fixed interior fraction of every volume.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline selecting slices via np.linspace(int(n*0.25), int(n*0.75), k) per volume discarded roughly half of each subject's acquired slices; per-subject aggregate features built from the restricted window produced a ranking metric 0.22 lower than a comparable pipeline whose features drew on the full slice range.
• Applies when: The program extracts hand-crafted features or statistics from multi-slice / volumetric image stacks by subsampling slice indices per volume.
• Example:
• Input:
    def volume_features(slice_files):
        slice_files = sorted(slice_files)
        n = len(slice_files)
        lo, hi = int(n * 0.25), int(n * 0.75)
        idxs = np.linspace(lo, hi - 1, 10).astype(int)
        feats = []
        for i in idxs:
            img = load_slice(slice_files[i])
            feats.append([img.mean(), img.std(), (img > 0).mean()])
        return np.array(feats).mean(axis=0)
• Consequence:
    Half of every subject's slices never contribute to features; per-subject
    aggregate statistics carry less discriminative signal, and the held-out
    ranking metric drops by ~0.22 versus sampling across the full index range.
• Counter-example:
• Input:
    def volume_features(slice_files):
        slice_files = sorted(slice_files)
        n = len(slice_files)
        idxs = np.linspace(0, n - 1, 10).astype(int)
        feats = []
        for i in idxs:
            img = load_slice(slice_files[i])
            feats.append([img.mean(), img.std(), (img > 0).mean()])
        return np.array(feats).mean(axis=0)
• Why it does not fire: The sampled indices span the full range 0 to n - 1, so no portion of the volume is categorically excluded by hardcoded fractional bounds.
C9 · discarded-signal Color reduced to global per-channel scalar meansgithub_occurrence

P1 — Color reduced to global per-channel scalar means

• Pattern: Detects a handcrafted image feature pipeline that summarizes color solely as one global scalar statistic per channel, with no per-channel histograms or spatially pooled color statistics, while all shape/texture descriptors are computed from a grayscale conversion.
• Detection procedure:
1. Locate a function that takes an image path or image array and returns a numeric feature vector consumed by a classifier fit later in the file.
2. Within that function body, confirm the image is converted to grayscale (e.g., a call to rgb2gray, convert("L"), cv2.cvtColor(..., COLOR_BGR2GRAY), or equivalent) and a shape/texture descriptor (e.g., hog or equivalent) is computed on the grayscale result.
3. Within the same function body, identify all operations on the color (multi-channel) array: check whether every color-derived feature is a global reduction to one scalar per channel, such as arr.mean(axis=(0, 1)), np.mean(arr[:, :, c]), or an equivalent whole-image mean/std per channel.
4. Within the same function body, confirm there is NO call to a histogram routine on channel values (np.histogram, cv2.calcHist, np.bincount, or equivalent) and NO loop or reshape that partitions the image into spatial cells before reducing channel values (no slicing of the height/width axes into blocks followed by a per-block mean).
5. PRESENT when steps 2, 3, and 4 all hold: grayscale-only shape features plus color represented exclusively by global per-channel scalars, with no histogram and no spatial pooling of color anywhere in the extraction function.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: An extraction function computing hog(rgb2gray(img)) plus only img.mean(axis=(0, 1)) for color produced a competition-metric score 0.10023 worse than an otherwise similar pipeline whose feature vector included per-channel histograms and grid-pooled channel means.
• Applies when: Image classification solved with handcrafted (non-deep) features fed to a shallow classifier, where the source images are multi-channel color.
• Example:
• Input:
    def extract_features(path):
        img = np.asarray(Image.open(path).resize((64, 64))) / 255.0
        gray = rgb2gray(img)
        shape_feat = hog(gray, pixels_per_cell=(8, 8), cells_per_block=(2, 2))
        color_feat = img.mean(axis=(0, 1))          # 3 scalars, only color info
        return np.concatenate([shape_feat, color_feat])

    X = np.vstack([extract_features(p) for p in train_paths])
    clf = LogisticRegression(max_iter=2000).fit(scaler.fit_transform(X), y)
• Consequence:
    Validation log loss is roughly 0.10 higher than the same pipeline with
    per-channel histograms and grid-pooled channel means added; the classifier
    cannot separate classes that differ in color distribution or spatial color
    layout, so predicted probabilities are systematically less confident and
    less accurate.
• Counter-example:
• Input:
    def extract_features(path):
        img = np.asarray(Image.open(path).resize((64, 64))) / 255.0
        gray = rgb2gray(img)
        shape_feat = hog(gray, pixels_per_cell=(8, 8), cells_per_block=(2, 2))
        hists = [np.histogram(img[:, :, c], bins=16, range=(0, 1))[0]
                 for c in range(3)]
        cells = img.reshape(4, 16, 4, 16, 3).mean(axis=(1, 3)).ravel()
        return np.concatenate([shape_feat, *hists, cells])
• Why it does not fire: the extraction function still uses grayscale HOG, but color is represented with per-channel 16-bin histograms and a 4×4 spatial grid of per-cell channel means, so step 4's absence conditions fail.
C9 · discarded-signal Global per-channel mean as sole color featuregithub_occurrence

P1 — Global per-channel mean as sole color feature

• Pattern: Detects handcrafted image-feature pipelines that reduce all color information to a single global scalar per channel (whole-image per-channel mean) concatenated with grayscale texture descriptors, with no distribution-resolved color representation anywhere in the feature vector.
• Detection procedure:
1. Locate a function or loop body that loads image files and constructs a per-image feature vector for a classifier (evidence: a grayscale conversion such as rgb2gray or equivalent, plus a texture descriptor call such as hog, local binary patterns, or equivalent, feeding into np.concatenate, np.hstack, or list-based vector assembly).
2. Within the same function or loop body, find color-derived features computed as a global mean over each channel: expressions of the form arr[:, :, c].mean(), np.mean(arr, axis=(0, 1)), or equivalent, where the reduction covers all pixels of the image.
3. Within the same function or loop body, check for any distribution-resolved color feature on the color channels: a call to np.histogram or equivalent applied to channel data, per-channel standard deviation or percentile calls, or any pooling that partitions the image spatially (subscripting with computed row/column ranges) before reducing a color channel.
4. PRESENT if step 2 finds global per-channel means contributing to the feature vector AND step 3 finds no other color-derived feature, so the means are the only chromatic components alongside grayscale texture descriptors.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: Feature extractor built vectors as np.concatenate([hog_feat, [r.mean(), g.mean(), b.mean()]]); the otherwise-identical pipeline that appended normalized per-channel intensity histograms scored ~6% relatively better on the classification metric (lower log loss).
• Applies when: Image classification with handcrafted (non-learned) features fed to a shallow classifier, where the source images are multi-channel color images.
• Example:
• Input:
    def extract_features(path):
        img = np.asarray(Image.open(path).resize((64, 64)))
        gray = rgb2gray(img)
        hog_feat = hog(gray, pixels_per_cell=(8, 8))
        color_feat = [img[:, :, 0].mean(),
                      img[:, :, 1].mean(),
                      img[:, :, 2].mean()]
        return np.concatenate([hog_feat, color_feat])
• Consequence:
    Classification log loss rises by ~6% relative to the same pipeline with
    per-channel histograms: nearly all chromatic distributional signal is
    discarded, so classes separable by color composition are confused.
• Counter-example:
• Input:
    def extract_features(path):
        img = np.asarray(Image.open(path).resize((64, 64)))
        gray = rgb2gray(img)
        hog_feat = hog(gray, pixels_per_cell=(8, 8))
        color_feat = []
        for c in range(3):
            hist, _ = np.histogram(img[:, :, c], bins=32, range=(0, 255))
            color_feat.append(hist / hist.sum())
        return np.concatenate([hog_feat] + color_feat)
• Why it does not fire: the color channels contribute 32-bin normalized histograms, a distribution-resolved representation, so step 3 finds a chromatic feature beyond a global scalar mean.
C9 · discarded-signal Label space restricted to a training subsampleverified_trace · effect +0.0003

P1 — Label space restricted to a training subsample

• Pattern: Detects a multi-class classifier whose set of predictable classes is derived solely from the labels present in a subsampled training set, while an auxiliary file enumerating the complete class vocabulary is available but never used to define the model's output classes.
• Detection procedure:
1. Locate a training-data loading step that takes only a subset of the full training data — e.g. a call with an explicit row limit (nrows=, head(n), slicing like [:N] with a literal integer, an early break in a reading loop, or a stride/sampling parameter).
2. Locate a label-mapping construction fitted on the labels from step 1's subset — e.g. LabelEncoder().fit(...) or fit_transform(...) on the sampled label column, a dict or set built by iterating the sampled labels, or equivalent — with no classes= argument or class-list parameter supplied to the estimator from any other source.
3. Check whether the source elsewhere reads a file whose contents enumerate class identifiers (e.g. a categories/labels metadata file loaded via read_csv or equivalent); PRESENT if either (a) such a file exists in the task's provided inputs and is never read, or (b) it is read but its class column is never passed into the label mapping of step 2 or into the estimator's class-set parameter.
4. The pattern is PRESENT when steps 1 and 2 hold and the condition in step 3 holds, so the estimator's output space is limited to classes observed in the subsample.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: le = LabelEncoder(); y_train = le.fit_transform([c for c in train_categories if c is not None]) fitted on a small prefix subsample of a task with thousands of classes; every test item whose true class was absent from the subsample was guaranteed wrong, contributing to a ~5.4x accuracy gap (0.0089 vs 0.0477) against a run that initialized the model with the full class set.
• Applies when: Multi-class classification where the number of classes is large relative to the training rows actually loaded, the training data is subsampled, and the complete class vocabulary is available (in an auxiliary metadata file or derivable from the full label column).
• Example:
• Input:
    df = pd.read_csv("train.csv", nrows=5000)  # ~5000 of millions of rows
    X = build_features(df)

    le = LabelEncoder()
    y = le.fit_transform(df["label"])  # only labels seen in the subsample

    clf = RandomForestClassifier(n_estimators=50, n_jobs=-1)
    clf.fit(X, y)

    preds = le.inverse_transform(clf.predict(build_features(test_df)))
    pd.DataFrame({"id": test_df["id"], "label": preds}).to_csv("out.csv", index=False)
• Consequence:
    The model can only ever emit the ~800 classes present in the 5000-row
    subsample out of ~5000 total classes; every test item belonging to one of
    the ~4200 unseen classes is scored wrong with certainty, capping accuracy
    far below what the same features achieve with the full class set
    (observed ~5.4x lower metric on a clean, zero-exit run).
• Counter-example:
• Input:
    all_classes = pd.read_csv("category_names.csv")["category_id"].values

    df = pd.read_csv("train.csv", nrows=5000)
    X = build_features(df)

    clf = SGDClassifier(loss="log_loss")
    clf.partial_fit(X, df["label"].values, classes=all_classes)

    preds = clf.predict(build_features(test_df))
    pd.DataFrame({"id": test_df["id"], "label": preds}).to_csv("out.csv", index=False)
• Why it does not fire: although training still uses a subsample, the estimator's class set is initialized from the auxiliary file enumerating the full vocabulary via classes=all_classes, so every class remains predictable.
C9 · discarded-signal Single-slice read from multi-file volumetric stacktrace_observed

P1 — Single-slice read from multi-file volumetric stack

• Pattern: Detects per-sample feature extraction over directories containing many slice files where the code selects exactly one file by a literal index into the file listing (e.g., the first sorted filename) and derives all features for that sample from that lone slice.
• Detection procedure:
1. Find a loop that iterates over per-sample directories (e.g., iterating over subdirectories of a train or test folder, one directory per sample id).
2. Within the loop body (or a function called with the sample directory), find a file-listing expression such as sorted(os.listdir(...)), sorted(glob.glob(...)), sorted(path.iterdir()), or equivalent, whose result is immediately subscripted with a single literal integer index such as [0] or [-1].
3. Confirm that the only image/volume file opened or decoded per sample inside that loop body is the file obtained at step 2 — there is no inner loop, slicing (e.g., [::k]), index sampling (e.g., evenly spaced indices over the listing length), or aggregation over multiple entries of the listing.
4. Confirm the value returned from reading that single file feeds the sample's feature vector or model input.
5. PRESENT when steps 1–4 all hold: multi-file stacks exist per sample, exactly one file per sample is read via a literal index, and features come only from that file.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A feature extractor used files = sorted(case_dir.glob("*")); img = read(files[0]) per sample directory containing hundreds of slice files; the held-out competition metric was ~15% relative lower (near chance) than a variant that sampled slices evenly across each stack, because >99% of per-sample pixel data never influenced any feature.
• Applies when: Each sample is stored as a directory (or series) of many slice/frame files forming one volume or sequence, and features or model inputs are computed per sample from those files.
• Example:
• Input:
    def process_case(case_dir):
        files = sorted(case_dir.glob("*.dcm"))
        ds = pydicom.dcmread(str(files[0]))
        img = ds.pixel_array.astype(np.float32)
        return np.array([img.mean(), img.std(), img.max()])

    X = []
    for case_id in sorted(d.name for d in train_dir.iterdir() if d.is_dir()):
        X.append(process_case(train_dir / case_id))
• Consequence:
    Held-out score drops toward chance (~15% relative lower on the evaluation
    metric than a slice-sampled variant): features derive from one arbitrary
    edge slice, so >99% of each sample's pixel data never influences the model.
• Counter-example:
• Input:
    def process_case(case_dir, n_slices=16):
        files = sorted(case_dir.glob("*.dcm"))
        idxs = np.linspace(0, len(files) - 1, n_slices).astype(int)
        stats = []
        for i in idxs:
            img = pydicom.dcmread(str(files[i])).pixel_array.astype(np.float32)
            stats.append([img.mean(), img.std(), img.max()])
        return np.array(stats).mean(axis=0)
• Why it does not fire: the file listing is indexed at multiple evenly spaced positions and per-slice statistics are aggregated across the stack, so no single literal-index file is the sole source of features.
C9 · discarded-signal Global-statistic features on weak-signal 2D arrays without baseline subtraction or coherent integrationtrace_observed

P1 — Global-statistic features on weak-signal 2D arrays without baseline subtraction or coherent integration

• Pattern: Detects feature extraction for weak-signal detection in 2D intensity arrays that computes only whole-array or single-axis summary statistics (moments, percentiles, extremes) on the raw pixel values, with no per-row or per-column baseline subtraction and no summing of values along candidate signal trajectories before feature computation.
• Detection procedure:
1. Locate a function or loop that loads 2D (or stacked-2D) numeric arrays (e.g., np.load, image/array readers, or equivalent) and produces a fixed-length numeric feature vector per array that is later passed to a classifier's fitting method.
2. Within that function or loop, list every operation applied to the loaded array before values are appended to the feature vector: check whether all such operations are reductions of the form mean/std/min/max/percentile/skew/kurtosis/sum (or equivalent) applied either to the whole array or along a single axis of the raw values.
3. Check that no statement subtracts a per-row or per-column statistic from the array (e.g., no expression of the form arr - reduce(arr, axis=k, keepdims=True) or an equivalent broadcasted subtraction of a row/column baseline) before any reduction in step 2.
4. Check that no statement accumulates the array along shifted or offset index paths (no loop or vectorized construct that sums elements taken at indices offset per row/column, i.e., shift-and-sum along candidate trajectories) before any reduction in step 2.
5. PRESENT if step 2 holds (all features are raw global/axis summary statistics) AND steps 3 and 4 both hold (no baseline subtraction and no trajectory integration anywhere in the feature-extraction code).
• Predicted impact:
• Add score for C9: 3
• weight: 3
• confidence: high
• Evidence: Feature extraction of the form feats = [arr.mean(), arr.std(), np.percentile(arr, 99), arr.max(), arr.sum(axis=0).max()] on raw 2D arrays fed to a gradient-boosted classifier scored roughly 20% relative lower on the ranking metric (AUC) than a pipeline that subtracted a per-column median baseline, normalized per panel, and summed along a grid of candidate drift trajectories before computing on/off contrasts.
• Applies when: The task is detecting faint, extended, or drifting signals in 2D intensity arrays (spectrograms, images) whose per-pixel amplitude is at or below the noise floor, and features are handcrafted rather than learned by a convolutional model.
• Example:
• Input:
    def extract_features(path):
        arr = np.load(path).astype(np.float32)  # shape (T, F)
        feats = []
        feats += [arr.mean(), arr.std(), arr.max(), arr.min()]
        feats += list(np.percentile(arr, [50, 90, 99]))
        col_sum = arr.sum(axis=0)
        feats += [col_sum.max(), col_sum.std()]
        row_sum = arr.sum(axis=1)
        feats += [row_sum.max(), row_sum.std()]
        return feats
• Consequence:
    Signals below per-pixel noise contribute nothing distinguishable to any
    feature; the classifier's ranking metric (AUC) lands ~20% relative below
    a pipeline that baseline-subtracts per column and integrates along
    candidate trajectories, on identical data and model settings.
• Counter-example:
• Input:
    def extract_features(path):
        arr = np.load(path).astype(np.float32)  # shape (T, F)
        arr = arr - np.median(arr, axis=0, keepdims=True)  # per-column baseline
        feats = [arr.mean(), arr.std(), np.percentile(arr, 99)]
        best = -1e9
        for drift in range(-4, 5):  # shift-and-sum along candidate slopes
            shifted = np.stack([np.roll(arr[t], t * drift) for t in range(arr.shape[0])])
            best = max(best, shifted.sum(axis=0).max())
        feats.append(best)
        return feats
• Why it does not fire: The array is baseline-subtracted per column and integrated along shifted trajectories before reduction, so steps 3 and 4 fail and the conjunction in step 5 is not met.
C9 · discarded-signal Constant empty prediction string for every test rowgithub_occurrence

P1 — Constant empty prediction string for every test row

• Pattern: Detects a submission-assembly loop whose per-row prediction value is set only from constant literals (such as the empty string) on every control path, so no quantity derived from the training data ever reaches the output file.
• Detection procedure:
1. Locate the code that writes the output file consumed for scoring (a call to to_csv, csv.writer, or equivalent on a frame or rows containing an identifier column and a prediction column).
2. Within the loop or comprehension that builds the prediction column's values, list every assignment to the per-row prediction variable (or every value appended to the list that becomes that column).
3. For each such assignment, check whether the right-hand side is a constant literal (e.g. "", 0, a fixed string) with no reference to any variable computed from training files, fitted models, per-sample features, or aggregated statistics.
4. Confirm that elsewhere in the same script, training data is loaded or statistics/models are fitted from labeled data (a read of a training file or a call to a fitting/aggregation routine whose result is stored in a variable).
5. PRESENT when every control path in step 2 assigns only constant literals (step 3 holds for all of them) while step 4 finds at least one fitted or aggregated artifact that is never referenced inside the prediction-assembly loop.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A script loaded training annotations and enumerated test samples, but the output loop assigned pred_string = "" on both branches of its conditional; the resulting file scored exactly 0.0 on the evaluation metric, equal to the trivial empty-submission baseline, while a variant that emitted per-row strings from aggregated training statistics scored 1.0 points higher.
• Applies when: The script produces a per-row prediction output for scoring, training labels or annotations are available to the script, and the script exits cleanly after writing the output.
• Example:
• Input:
    train_df = pd.read_csv("train.csv")
    stats = train_df.groupby("col_a").mean()  # fitted artifact

    sub = pd.read_csv("sample_submission.csv")
    predictions = []
    for row_id in sub["Id"]:
        if row_id in known_ids:
            pred_string = ""   # avoid penalties
        else:
            pred_string = ""
        predictions.append({"Id": row_id, "PredictionString": pred_string})
    pd.DataFrame(predictions).to_csv("submission.csv", index=False)
• Consequence:
    Run completes with exit code 0 and a well-formed output file, but every
    prediction cell is empty; the evaluation metric equals the empty-submission
    baseline (0.0), strictly below any run whose output consumes the fitted
    statistics computed earlier in the same script.
• Counter-example:
• Input:
    train_df = pd.read_csv("train.csv")
    stats = train_df.groupby("col_a").mean()

    sub = pd.read_csv("sample_submission.csv")
    predictions = []
    for row_id in sub["Id"]:
        feat = features.get(row_id)
        if feat is None:
            pred_string = ""            # fallback for missing rows only
        else:
            pred_string = format_boxes(stats, feat)
        predictions.append({"Id": row_id, "PredictionString": pred_string})
    pd.DataFrame(predictions).to_csv("submission.csv", index=False)
• Why it does not fire: at least one control path assigns the prediction from format_boxes(stats, feat), which consumes the fitted artifact, so the constant empty string is only a fallback for missing rows rather than the value on every path.
C9 · discarded-signal Fixed-length prefix truncation of flattened raster featuresgithub_occurrence

P1 — Fixed-length prefix truncation of flattened raster features

• Pattern: Detects a flattened multi-dimensional image or grid array being truncated to a fixed-length prefix slice before use as model features, so only a contiguous top band of the raster contributes signal and the remainder is silently discarded.
• Detection procedure:
1. Locate a call that flattens a multi-dimensional array, e.g. .flatten(), .ravel(), or .reshape(-1) (or equivalent), applied to data loaded from an image file or a 2-D/3-D grid.
2. On the result of step 1 (either directly chained or via a variable assigned from it within the same function body), find a subscript of the form [:N] where N is a literal integer or an integer variable smaller than the flattened array's full length (full length inferable from a preceding resize/reshape with literal dimensions whose product exceeds N).
3. Confirm the truncated array is appended to, stacked into, or otherwise collected as a feature row that is later passed to a model-fitting or transform call.
4. The pattern is PRESENT when a flattened raster is prefix-sliced to fewer elements than its full length (steps 1–2) and the truncated result feeds model training or inference (step 3).
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: img_array.flatten()[:1024] on images resized to 64x64x3 (12288 values) kept only the first 1024 values — roughly the top few rows of the raster — and the resulting classifier's log loss was ~16% relative worse than a model using a whole-raster descriptor.
• Applies when: The program builds handcrafted feature vectors from images or other spatial grids (resize + flatten pipelines) for a classical model, rather than feeding full tensors to a convolutional model.
• Example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('RGB').resize((64, 64))
        arr = np.array(img, dtype=np.float32) / 255.0
        return arr.flatten()[:1024]  # keep first 1024 values

    X = np.vstack([extract_features(p) for p in train_paths])
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    model.fit(X_scaled, y)
• Consequence:
    Each 12288-value image representation is cut to its first 1024 values,
    i.e. only the top ~5 rows of pixels; the classifier sees a narrow spatial
    band, and hold-out log loss is ~16% relative worse than the same pipeline
    using the full flattened raster or a whole-image descriptor.
• Counter-example:
• Input:
    def extract_features(path):
        img = Image.open(path).convert('RGB').resize((16, 16))
        arr = np.array(img, dtype=np.float32) / 255.0
        return arr.flatten()  # all 768 values, whole image

    X = np.vstack([extract_features(p) for p in train_paths])
    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)
    model.fit(X_scaled, y)
• Why it does not fire: The dimensionality is reduced by resizing the whole image before flattening, so every spatial region contributes and no prefix slice discards part of the raster.
C9 · discarded-signal Hardcoded coordinate offsets instead of label-derived object placementgithub_occurrence

P1 — Hardcoded coordinate offsets instead of label-derived object placement

• Pattern: Detects spatial predictions whose positions are generated from a hardcoded literal list of coordinate offsets applied to every sample, while the training annotations containing target positions are loaded but never used to derive those positions.
• Detection procedure:
1. Identify a container (list or tuple) of numeric literal coordinate pairs or triples assigned to a variable in the module (e.g. offsets = [(16.0, 0.0), (8.0, 3.0), ...]), where every element is a numeric literal, not a computed value.
2. Confirm the elements of that container are iterated inside a loop over prediction/output rows and combined arithmetically with a per-sample reference pose (translation and/or rotation) to produce output coordinates written into the prediction column or file.
3. Identify a load of training annotations (e.g. pd.read_csv("train.csv"), json.load of a labels file, or equivalent) whose records include position/coordinate fields.
4. Verify that no position/coordinate field from the loaded annotations flows into the values written as predicted coordinates — the annotation frame may be read for other attributes (size averages, most frequent class), but no clustering, histogram, mean, or other aggregation of annotation positions is assigned to a variable used in step 2.
5. PRESENT when steps 1–2 hold (positions come from constant literals) and steps 3–4 hold (annotation positions are available but unused for placement).
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)] applied identically to every sample while training labels were loaded only to compute mean box dimensions and the most common class; the overlap-based detection metric (mAP over IoU thresholds) was 0.0, versus a nonzero score (0.00057) from a variant that clustered annotation positions in the reference frame and emitted the density peaks.
• Applies when: The task requires predicting object locations (boxes, points, or regions) and training annotations containing target positions are available to the program; scoring uses an overlap- or distance-based match against ground truth.
• Example:
• Input:
    ann = json.load(open("train_annotations.json"))  # has x, y, z per object
    mean_w = np.mean([a["w"] for a in ann])
    offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)]
    rows = []
    for tok in sub["Id"]:
        tx, ty, tz, yaw = poses[tok]
        parts = []
        for dx, dy in offsets:
            wx = tx + dx * math.cos(yaw) - dy * math.sin(yaw)
            wy = ty + dx * math.sin(yaw) + dy * math.cos(yaw)
            parts.append(f"1.0 {wx:.3f} {wy:.3f} {tz:.3f} {mean_w:.3f}")
        rows.append(" ".join(parts))
• Consequence:
    Predicted boxes sit at four fixed offsets that rarely coincide with real
    object locations; overlap-based mAP falls to ~0.0, while a baseline that
    places boxes at annotation-derived density peaks scores measurably above
    zero on the same metric.
• Counter-example:
• Input:
    ann = json.load(open("train_annotations.json"))
    rel = np.array([to_ref_frame(a, poses[a["tok"]]) for a in ann])  # x,y,z
    km = KMeans(n_clusters=4).fit(rel[:, :2])
    cands = km.cluster_centers_
    for tok in sub["Id"]:
        tx, ty, tz, yaw = poses[tok]
        parts = []
        for dx, dy in cands:
            wx = tx + dx * math.cos(yaw) - dy * math.sin(yaw)
            wy = ty + dx * math.sin(yaw) + dy * math.cos(yaw)
            parts.append(f"1.0 {wx:.3f} {wy:.3f} {tz:.3f}")
• Why it does not fire: the per-sample offsets are computed by clustering annotation positions transformed into the reference frame, so predicted coordinates flow from the training labels rather than from a literal constant list.
C9 · discarded-signal Global-scalar-only image features for many-class classification

P1 — Global-scalar-only image features for many-class classification

• Pattern: Detects an image-classification pipeline whose entire feature vector per image is a small fixed set of whole-image scalar reductions (mean, standard deviation, min, max, aggregate edge/gradient response), producing a feature dimensionality far below the number of target classes, with no downsampled pixel values, histograms, or local descriptors retained.
• Detection procedure:
1. Locate the function or code block that converts an image (array or decoded bytes) into the feature vector later passed to a classifier's fitting method. It qualifies if the returned vector is assembled as a list or array literal (e.g. np.array([...]), [...]) whose elements are all scalar reductions computed over the full image, such as .mean(), .std(), .min(), .max(), np.mean(img), np.sum(edges), or an aggregate of an edge/gradient map (e.g. a Canny or Sobel output reduced with .mean() or .sum(), or equivalent).
2. Confirm that within this function there is no operation that preserves spatial or distributional layout: no flattening of a resized image (e.g. resize(...) followed by .flatten()/.ravel() or equivalent), no histogram with multiple bins, and no local-descriptor extractor (HOG, LBP, SIFT-style, or equivalent) whose full output vector is kept.
3. Count the scalar elements in the returned vector; the pattern requires this count to be a fixed small number (16 or fewer).
4. Confirm the label set is large: the labels fitted by the classifier come from encoding a category column (via a label encoder or equivalent) whose distinct-value count is not a small hardcoded number — e.g. the task involves dozens or more classes, or the class count is data-driven with no cap below the feature count.
5. PRESENT when steps 1–4 all hold: every image is reduced to ≤16 global scalars, no spatial/color-layout features survive, and the classifier is trained to separate more classes than there are feature dimensions.
• Predicted impact:
• Add score for C9 (Available signal never used): 3
• weight: 3
• confidence: high
• Evidence: A pipeline extracting features as [img.mean(), img.std(), img.min(), img.max(), edges.mean()] and fitting a tree ensemble on a many-class label set scored 0.32597 lower on the competition metric than a pipeline that flattened downsampled RGB pixels into the feature vector, despite the weaker pipeline processing at least as many training samples.
• Applies when: The program performs multi-class image classification with hand-built (non-deep-learning) features, and the number of target classes is large relative to a handful of scalars.
• Example:
• Input:
    def extract_features(img_bytes):
        img = np.array(Image.open(io.BytesIO(img_bytes)))
        gray = rgb2gray(img)
        edges = feature.canny(gray)
        return np.array([img.mean(), img.std(), img.min(),
                         img.max(), edges.mean()])

    X = np.stack([extract_features(b) for b in image_blobs])
    le = LabelEncoder()
    y = le.fit_transform(categories)   # thousands of distinct classes
    model.fit(X, y)
• Consequence:
    Classes are inseparable in a 5-dimensional feature space spanning a
    class count orders of magnitude larger; held-out accuracy falls to
    near the majority-class baseline, ~33 points (relative) below a
    flattened-downsampled-pixel feature set on the same data.
• Counter-example:
• Input:
    def extract_features(img_bytes):
        img = Image.open(io.BytesIO(img_bytes)).resize((16, 16))
        arr = np.asarray(img, dtype=np.float32) / 255.0
        pixel_feats = arr.flatten()          # 16*16*3 = 768 dims
        return np.concatenate([pixel_feats,
                               [arr.mean(), arr.std()]])

    X = np.stack([extract_features(b) for b in image_blobs])
    model.fit(X, LabelEncoder().fit_transform(categories))
• Why it does not fire: although global mean and std scalars are appended, the vector also retains a flattened downsampled pixel array, so step 2 fails and the feature dimensionality is not far below the class count.
C9 · discarded-signal Word-only n-gram features for style-based text classificationverified_trace · effect +0.0979

P1 — Word-only n-gram features for style-based text classification

• Pattern: Detects a style- or authorship-attribution text pipeline whose only text featurization is word-level n-grams, with no character-level n-gram features anywhere in the feature set, discarding sub-word stylistic signal such as punctuation, morphology, and spelling habits.
• Detection procedure:
1. Determine the task is stylistic text classification: the target column identifies an author, writing style, or persona (not the document's topic), and the input features are raw text strings fed to one or more text vectorizers.
2. List every text-vectorization construct in the source (e.g., TfidfVectorizer, CountVectorizer, HashingVectorizer, or equivalent) whose output is used, directly or in a stacked/concatenated matrix, to fit any model whose predictions reach the final output.
3. For each such construct, check its analyzer setting: it either omits the analyzer argument (defaulting to word analysis) or sets it to 'word'; no construct sets analyzer to 'char' or 'char_wb' (or equivalent character-level tokenization), and no other character-level feature extraction (e.g., per-character counts, punctuation frequency features) is concatenated into the model input.
4. The pattern is PRESENT when steps 1–3 all hold: a stylistic classification task whose entire text feature matrix is built exclusively from word-level n-grams.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only TfidfVectorizer(ngram_range=(1, 2), analyzer='word') features for author attribution scored ~0.44 worse on the multiclass log-loss metric than an otherwise comparable pipeline that also included character-level n-gram features (~0.54 vs ~0.30 log loss).
• Applies when: The program performs supervised text classification where classes are distinguished by writing style or author identity rather than subject matter, and features are derived from raw text via bag-of-n-grams vectorization.
• Example:
• Input:
    train = pd.read_csv('train.csv')
    test = pd.read_csv('test.csv')
    y = LabelEncoder().fit_transform(train['author'])
    tfidf = TfidfVectorizer(max_features=5000, ngram_range=(1, 2),
                            min_df=2, analyzer='word')
    X_tr = tfidf.fit_transform(train['text'])
    X_te = tfidf.transform(test['text'])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y)
    pred = model.predict_proba(X_te)
    pd.DataFrame(pred).to_csv('submission.csv', index=False)
• Consequence:
    Multiclass log loss on held-out data is substantially higher (worse)
    than the same pipeline with character n-gram features added
    (~0.54 vs ~0.30 in a measured contrastive run); punctuation and
    sub-word style cues carrying class signal are never encoded.
• Counter-example:
• Input:
    word_vec = TfidfVectorizer(ngram_range=(1, 2), min_df=2)
    char_vec = TfidfVectorizer(analyzer='char_wb', ngram_range=(2, 5),
                               min_df=2)
    X_tr = hstack([word_vec.fit_transform(train['text']),
                   char_vec.fit_transform(train['text'])])
    X_te = hstack([word_vec.transform(test['text']),
                   char_vec.transform(test['text'])])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, LabelEncoder().fit_transform(train['author']))
    pred = model.predict_proba(X_te)
• Why it does not fire: A second vectorizer with analyzer='char_wb' is horizontally stacked into the model input, so character-level stylistic signal is present in the feature matrix.
C9 · discarded-signal Word-only features for style/authorship target

P1 — Word-only features for style/authorship target

• Pattern: Detects a text-classification pipeline whose target is authorship or writing style that builds all text features from word-level tokens, with no character n-gram representation anywhere in the feature set.
• Detection procedure:
1. Confirm the task predicts an author, writer identity, or stylistic category from raw text: the target column holds author/style labels (e.g., a label column joined to a text column, with class names denoting persons or styles) rather than topical categories.
2. Locate every text-vectorization construct in the source: instantiations of TfidfVectorizer, CountVectorizer, HashingVectorizer (or equivalent tokenizing feature extractors), or any custom tokenizer feeding a matrix used for model fitting.
3. For each such construct, inspect its analyzer configuration: it uses word tokens if it passes analyzer='word', omits the analyzer argument entirely (word is the default in common libraries), or tokenizes on whitespace/word boundaries in custom code.
4. Check whether ANY vectorizer in the pipeline uses analyzer='char' or analyzer='char_wb' (or equivalent character-level n-gram extraction) whose output is concatenated, stacked, or otherwise included in the features passed to a fitting call.
5. PRESENT if step 1 holds, at least one word-level vectorizer from step 3 exists, and no character-level representation from step 4 exists anywhere in the code.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline using only TfidfVectorizer(..., analyzer='word') before a linear classifier on an authorship task scored 0.44 worse on multiclass log loss than the same task solved with stacked word- and character-level features; punctuation, morphology, and spelling habits carried by sub-word patterns were never exploited.
• Applies when: The program performs supervised text classification and the labels identify who wrote the text or what stylistic register it belongs to, rather than what topic it covers.
• Example:
• Input:
    tfidf = TfidfVectorizer(
        ngram_range=(1, 2),
        max_df=0.95,
        analyzer='word'
    )
    X_tr = tfidf.fit_transform(train['text'])
    X_te = tfidf.transform(test['text'])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Consequence:
    Multiclass log loss is measurably higher (observed gap ~0.44) than the same
    classifier given word features stacked with character n-gram features;
    stylistic sub-word signal (punctuation habits, spelling, morphology)
    present in the text is never used, degrading probability estimates in the
    wrong direction on every class.
• Counter-example:
• Input:
    word_vec = TfidfVectorizer(ngram_range=(1, 2), analyzer='word')
    char_vec = TfidfVectorizer(analyzer='char_wb', ngram_range=(2, 5))
    X_tr = hstack([word_vec.fit_transform(train['text']),
                   char_vec.fit_transform(train['text'])])
    X_te = hstack([word_vec.transform(test['text']),
                   char_vec.transform(test['text'])])
    model = LogisticRegression(max_iter=1000)
    model.fit(X_tr, y_train)
    proba = model.predict_proba(X_te)
• Why it does not fire: a character n-gram vectorizer (analyzer='char_wb') is stacked with the word features before fitting, so step 4 finds a character-level representation and the conjunction in step 5 fails.
C9 · discarded-signal Single edge-slice read from a multi-file volumetric seriestrace_observed

P1 — Single edge-slice read from a multi-file volumetric series

• Pattern: Detects per-subject feature extraction over a directory of image-series files that loads only the first file in sorted listing order and never reads the remaining files, so each subject's volume is reduced to one arbitrary edge slice.
• Detection procedure:
1. Locate a function or block that builds a per-subject feature vector and, within it, an expression that lists the files of a subject's series directory (e.g. os.listdir, glob.glob, Path.iterdir, or equivalent), optionally wrapped in sorted(...).
2. Check that the result of step 1 is indexed with a literal [0] (or sliced [:1]) and that the selected single filename is passed to an image/series reader (e.g. pydicom.dcmread, nibabel.load, Image.open, cv2.imread, or equivalent).
3. Within the same function body, confirm there is no for loop, comprehension, or vectorized call that reads any other element of the file list from step 1, and no positional selection of a central index (e.g. len(files) // 2) or stride over the list.
4. PRESENT when steps 1–3 all hold: the subject's multi-file series is represented solely by the first file in directory order, with all other slices unread.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A pipeline extracting per-subject statistics from sorted(os.listdir(series_dir))[0] scored ≈0.49 AUC (near chance), while an otherwise comparable pipeline aggregating the same statistic family over central slices of the series scored ≈0.58 AUC — a 0.09 absolute drop attributable to representing each volume by one edge slice.
• Applies when: Each subject/sample is stored as a directory of multiple image files forming a 3D series (one file per slice), and features are handcrafted per subject rather than learned end-to-end on the full volume.
• Example:
• Input:
    def extract_features(case_dir):
        files = sorted(os.listdir(case_dir))
        img = pydicom.dcmread(os.path.join(case_dir, files[0])).pixel_array
        img = img.astype(np.float32)
        return np.array([img.mean(), img.std(), img.max(), img.min()])

    X = np.stack([extract_features(d) for d in case_dirs])
    model.fit(X, y)
• Consequence:
    Each subject is summarized by the first slice in filename order — typically
    a near-empty edge slice — so features carry almost no signal; held-out AUC
    falls to ~0.49 (chance level), ~0.09 absolute below the same statistics
    aggregated over central slices of each series.
• Counter-example:
• Input:
    def extract_features(case_dir):
        files = sorted(os.listdir(case_dir))
        mid = len(files) // 2
        feats = []
        for f in files[mid - 5 : mid + 5]:
            img = pydicom.dcmread(os.path.join(case_dir, f)).pixel_array.astype(np.float32)
            feats.append([img.mean(), img.std(), img.max(), img.min()])
        feats = np.array(feats)
        return np.concatenate([feats.mean(axis=0), feats.std(axis=0)])
• Why it does not fire: the file list is not truncated to element [0]; a window of central slices is read in a loop and per-slice features are aggregated into the subject vector, so step 2 and step 3 both fail.
C9 · discarded-signal Stop-word removal in stylometric text features

P1 — Stop-word removal in stylometric text features

• Pattern: Detects a bag-of-words or TF-IDF text vectorizer configured with a semantic stop-word list when the prediction target is the author or writing style of the text, discarding the function-word frequencies that carry the strongest per-author signal.
• Detection procedure:
1. Identify the prediction target: locate the column of the training frame passed as the label argument to a model-fitting call (e.g. the second positional argument of fit), and check from surrounding code or comments that it identifies the writer, author, or style of a text field rather than the topic or content.
2. Within the same script, find the construction of a token-count or TF-IDF text vectorizer (TfidfVectorizer, CountVectorizer, or equivalent) whose output is used, directly or via transformation, as model input features.
3. Check whether that constructor call includes a stop-word argument set to a non-null value — a named language list such as stop_words='english', or an explicit list/set of common function words (e.g. containing 'the', 'and', 'of').
4. PRESENT if the target from step 1 is an author/style identity and the vectorizer from step 2 has the stop-word argument from step 3 enabled.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A vectorizer built with stop_words='english' for an author-identification target produced multi-class log loss roughly 0.34–0.40 worse than otherwise similar pipelines that retained function words in the vocabulary, across two measured contrastive pairs.
• Applies when: Text classification code where the label is authorship or writing style and features are derived from token counts or TF-IDF over the raw text.
• Example:
• Input:
    train = pd.read_csv("train.csv")
    vectorizer = TfidfVectorizer(
        max_features=20000,
        stop_words='english',
        lowercase=True,
    )
    X = vectorizer.fit_transform(train["text_col"])
    model = LogisticRegression(max_iter=1000)
    model.fit(X, train["author_col"])
    proba = model.predict_proba(vectorizer.transform(test["text_col"]))
• Consequence:
    Multi-class log loss on the held-out evaluation rises materially (observed
    ~0.34–0.40 absolute worse, a ~40% relative degradation) because per-author
    function-word usage frequencies are stripped from the feature space before
    the model sees them.
• Counter-example:
• Input:
    train = pd.read_csv("train.csv")
    vectorizer = TfidfVectorizer(
        max_features=20000,
        min_df=3,
        max_df=0.95,
        lowercase=True,
    )
    X = vectorizer.fit_transform(train["text_col"])
    model = LogisticRegression(max_iter=1000)
    model.fit(X, train["author_col"])
• Why it does not fire: The vectorizer prunes only by document-frequency thresholds (min_df/max_df) and never enables a stop-word list, so common function words remain in the vocabulary.
C9 · discarded-signal Hardcoded placement offsets instead of label-derived statistics

P1 — Hardcoded placement offsets instead of label-derived statistics

• Pattern: Detects heuristic baseline prediction code that positions predicted objects using author-invented literal coordinate offsets or counts, while the provided training annotations containing real object locations and sizes are never loaded or aggregated.
• Detection procedure:
1. Locate a loop over evaluation samples that builds a prediction string or record containing spatial coordinates (e.g., formatted floats for x, y, z, width, length, height, or yaw) written to a submission file.
2. Within the same module, find the source of the coordinate values placed into those predictions: check whether they originate from a literal list or tuple of numeric constants (e.g., offsets = [(16.0, 0.0), (8.0, 3.0), ...]) or from single hardcoded scalar assignments, rather than from any variable computed by reading the training label file (train.csv or equivalent annotation source) and aggregating its coordinate columns (via mean, median, std, groupby, histogramming, or equivalent).
3. Confirm that no statement in the module reads per-object location columns from the training annotations and feeds any statistic of them into the placement of predicted objects (reading labels only for class names or counts of rows does not count as placement fitting).
4. The pattern is PRESENT when predictions with spatial coordinates are emitted for every evaluation sample AND all placement offsets/positions trace back to literal numeric constants AND training annotations containing object locations are available but never aggregated into those placements.
• Predicted impact:
• Add score for C9: 3
• weight: 3
• confidence: high
• Evidence: A baseline emitted boxes at a fixed literal set offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)] rotated into the world frame; the overlap-based detection metric scored exactly zero, while a sibling script that fitted per-class relative-position, size, and yaw statistics from the training annotations and sampled placements from them scored measurably above zero (contrastive gap equal to the full metric range).
• Applies when: The task is spatial prediction (object locations/boxes in 2D or 3D), training labels containing per-object coordinates are available on disk, and the code under review is a heuristic (non-learned) baseline that must still place objects.
• Example:
• Input:
    offsets = [(16.0, 0.0), (8.0, 3.0), (8.0, -3.0), (-8.0, 0.0)]
    mean_w, mean_l, mean_h = 2.0, 5.0, 1.8
    parts = []
    for sample_id in sub_df["Id"]:
        tx, ty, tz, yaw = ego_pose[sample_id]
        row = []
        for dx, dy in offsets:
            wx = tx + dx * math.cos(yaw) - dy * math.sin(yaw)
            wy = ty + dx * math.sin(yaw) + dy * math.cos(yaw)
            row.append(f"1.0 {wx:.3f} {wy:.3f} {tz:.3f} {mean_w} {mean_l} {mean_h} {yaw:.4f} LABEL_A")
        parts.append(" ".join(row))
    sub_df["PredictionString"] = parts
• Consequence:
    Overlap-based detection score drops to exactly 0.0: guessed placements
    almost never intersect ground-truth volumes, versus a nonzero score
    (full metric gap of 1.0 observed) when placements are sampled from
    location/size statistics fitted on the training annotations.
• Counter-example:
• Input:
    ann = pd.read_csv("train.csv")
    stats = ann.groupby("class")[["rel_x", "rel_y", "rel_z", "w", "l", "h"]].agg(["mean", "std"])
    parts = []
    for sample_id in sub_df["Id"]:
        tx, ty, tz, yaw = ego_pose[sample_id]
        row = []
        for c, s in stats.iterrows():
            wx = tx + rng.normal(s[("rel_x", "mean")], s[("rel_x", "std")])
            wy = ty + rng.normal(s[("rel_y", "mean")], s[("rel_y", "std")])
            row.append(f"1.0 {wx:.3f} {wy:.3f} {tz:.3f} {s[('w','mean')]:.3f} {s[('l','mean')]:.3f} {s[('h','mean')]:.3f} {yaw:.4f} {c}")
        parts.append(" ".join(row))
    sub_df["PredictionString"] = parts
• Why it does not fire: The placement coordinates and box sizes are derived from aggregated training-annotation statistics rather than literal numeric constants, so step 2's condition fails.
C9 · discarded-signal Image features reduced to scalar intensity statistics onlygithub_occurrence

P1 — Image features reduced to scalar intensity statistics only

• Pattern: Detects an image-classification feature pipeline that collapses each image or volume into only global scalar intensity summaries (means, standard deviations, percentiles, threshold fractions, gradient magnitudes), feeding the model a feature vector with no spatially resolved representation such as a flattened downscaled pixel grid.
• Detection procedure:
1. Identify a per-sample feature-extraction function or loop that reads image data (e.g. via pydicom, PIL.Image, cv2.imread, nibabel, or equivalent) and returns a numeric vector later stacked into a training matrix passed to a classical model's fit (e.g. LogisticRegression, RandomForestClassifier, gradient boosting, or equivalent).
2. Within that function or loop, list every value appended to the output vector; confirm each is produced by a global reduction over pixel values — calls such as mean, std, percentile, median, max, min, a comparison-then-mean fraction (e.g. (img > t).mean()), or a norm/sum of a gradient array — applied over the whole image or volume.
3. Confirm the output vector contains no term that preserves spatial layout: no flattened result of a resize/downscale call (resize(img, (k, k)).ravel() or equivalent), no per-region/per-patch statistics indexed by position, and no histogram-of-location features.
4. PRESENT when step 2 holds for every element of the feature vector and step 3 confirms no spatially resolved component reaches the model input.
• Predicted impact:
• Add score for C9: Available signal never used: 2
• weight: 2
• confidence: high
• Evidence: A pipeline whose per-volume features were only mean, percentile, threshold fractions, and gradient energy scored ~5.5% relative lower on the ranking metric (AUC) than an otherwise similar pipeline that additionally appended a flattened coarse downscaled slice per modality to the same statistics.
• Applies when: The task is classification from image or volumetric input, features are hand-built (no deep-learning framework in use), and the target signal may depend on where intensity patterns occur, not only on their global distribution.
• Example:
• Input:
    def extract(img):
        feats = [
            img.mean(),
            img.std(),
            np.percentile(img, 90),
            (img > img.mean()).mean(),
            np.abs(np.gradient(img.astype(float))[0]).mean(),
        ]
        return np.array(feats)

    Xtr = np.stack([extract(load(i)) for i in train_ids])
    model.fit(Xtr, ytr)
• Consequence:
    Ranking metric (AUC) lands ~5% relatively closer to chance than the same
    pipeline with a coarse spatial grid appended; all location-dependent
    signal in the images is discarded before the model ever sees it.
• Counter-example:
• Input:
    def extract(img):
        stats = [img.mean(), img.std(), np.percentile(img, 90)]
        grid = resize(img.astype(float), (8, 8), anti_aliasing=True)
        return np.concatenate([stats, grid.ravel()])

    Xtr = np.stack([extract(load(i)) for i in train_ids])
    model.fit(Xtr, ytr)
• Why it does not fire: The feature vector includes a flattened 8x8 downscaled grid, so a spatially resolved representation reaches the model alongside the scalar statistics, failing step 3.
C9 · discarded-signal Stop-word removal or tight vocabulary cap on style/authorship text features

P1 — Stop-word removal or tight vocabulary cap on style/authorship text features

• Pattern: Detects a bag-of-words or TF-IDF feature extractor for an authorship- or writing-style-prediction task that is configured to discard stop words or cap the vocabulary to a few thousand terms, removing the function-word frequencies that carry stylistic signal.
• Detection procedure:
1. Confirm the task is style- or author-oriented: the target column being fit represents an author, writer, or style identity rather than a topic (e.g., the label values are author identifiers and the script's comments, filenames, or class list indicate author/style prediction).
2. Locate the construction of a token-frequency vectorizer (TfidfVectorizer, CountVectorizer, or equivalent) whose output is later passed to a model-fitting call within the same script.
3. In that constructor call, check whether the keyword stop_words is set to any non-None value (such as 'english' or a list), OR the keyword max_features (or an equivalent vocabulary-size limit) is set to a literal integer less than or equal to 10000.
4. The pattern is PRESENT if the task is author/style prediction (step 1) and the vectorizer feeding the model satisfies either condition in step 3.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A vectorizer constructed with stop_words='english', max_features=5000 on an author-identification task, versus an otherwise similar pipeline keeping the full vocabulary with character n-grams, produced multiclass log loss of ~0.60 vs ~0.35 — roughly 42% relatively worse — because function words (the most discriminative authorship features) were excluded from the representation.
• Applies when: Text classification where the label is the author or writing style, using bag-of-words/TF-IDF style features fed to a classifier.
• Example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # Task: predict which author wrote each sentence
    vec = TfidfVectorizer(stop_words='english', max_features=5000)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])
    model = LogisticRegression()
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
• Consequence:
    Multiclass log loss degrades substantially (observed ~0.60 vs ~0.35 for the
    full-vocabulary variant, ~42% relatively worse) because function-word
    frequencies — the strongest authorship signal — are stripped from the
    feature space before the model ever sees them.
• Counter-example:
• Input:
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression

    # Task: predict which author wrote each sentence
    vec = TfidfVectorizer(sublinear_tf=True, ngram_range=(1, 2), min_df=2)
    X_train = vec.fit_transform(train['text'])
    X_test = vec.transform(test['text'])
    model = LogisticRegression()
    model.fit(X_train, y_train)
    probs = model.predict_proba(X_test)
• Why it does not fire: The vectorizer keeps stop words (no stop_words argument) and sets no small max_features cap; min_df=2 only prunes near-unique tokens and preserves the high-frequency function words carrying stylistic signal.
C9 · discarded-signal Global scalar image statistics as sole classifier featuresgithub_occurrence

P1 — Global scalar image statistics as sole classifier features

• Pattern: Detects an image classification pipeline over many classes where each image is reduced to a short vector of global scalar statistics (mean, standard deviation, min, max, aggregated edge or gradient magnitude), with no spatially- or color-resolved representation such as a downscaled pixel grid, per-region statistics, or histograms.
• Detection procedure:
1. Locate a function or code block that takes decoded image data (e.g. output of Image.open, imread, or equivalent) and produces the per-sample feature vector that is later stacked into the matrix passed to a classifier's fitting call (fit or equivalent).
2. Within that function or block, list every value appended to the feature vector; check whether each is a global scalar aggregate over the whole image — calls such as .mean(), .std(), .min(), .max(), np.sum(...), or a scalar aggregate of an edge/gradient map (e.g. mean of feature.canny(...), sobel(...), or equivalent) applied without any spatial partitioning.
3. Confirm the feature vector contains no flattened resized pixel array (e.g. img.resize(...) followed by .flatten() / np.asarray(...).ravel() or equivalent), no histogram call (np.histogram, .histogram(), or equivalent), and no per-tile/per-channel-region aggregation loop.
4. Confirm the classification target has many distinct classes — the label source is a column or list encoded via a label encoder or similar, and there is no evidence the problem is binary or few-class (e.g. no hard-coded two-class threshold).
5. PRESENT when the matrix given to the classifier's fitting call is built exclusively from the global scalar aggregates of step 2 (feature dimensionality on the order of ten or fewer per image) and the conditions of steps 3 and 4 hold.
• Predicted impact:
• Add score for C9: Available signal never used: 3
• weight: 3
• confidence: high
• Evidence: A pipeline built features as [img.mean(), img.std(), img.min(), img.max(), edges.mean()] per image and trained a tree ensemble on them for a many-class image task; it scored 0.009 on the competition metric while an otherwise comparable program using flattened downscaled pixels scored 0.30 on the same budget.
• Applies when: The program trains a classical (non-deep) classifier for image classification with many classes, and features are computed directly from decoded image arrays in the same script.
• Example:
• Input:
    def img_features(img):
        arr = np.asarray(img.convert('RGB'), dtype=np.float32)
        gray = rgb2gray(arr)
        edges = feature.canny(gray)
        return [arr.mean(), arr.std(), arr.min(), arr.max(), edges.mean()]

    X_train = np.array([img_features(im) for im in train_images])
    le = LabelEncoder()
    y_train = le.fit_transform(train_labels)
    clf = RandomForestClassifier(n_estimators=50, n_jobs=-1)
    clf.fit(X_train, y_train)
• Consequence:
    Held-out accuracy collapses to near-chance (0.009 vs 0.30 for the same
    pipeline using a flattened 16x16x3 downscaled pixel grid), because five
    global scalars cannot separate thousands of visually distinct classes.
• Counter-example:
• Input:
    def img_features(img):
        small = img.convert('RGB').resize((16, 16))
        px = np.asarray(small, dtype=np.float32).ravel() / 255.0
        arr = np.asarray(img.convert('RGB'), dtype=np.float32)
        return np.concatenate([px, [arr.mean(), arr.std()]])

    X_train = np.array([img_features(im) for im in train_images])
    clf = RandomForestClassifier(n_estimators=50, n_jobs=-1)
    clf.fit(X_train, LabelEncoder().fit_transform(train_labels))
• Why it does not fire: The feature vector includes a flattened downscaled pixel grid alongside the scalar statistics, so step 3's condition (no spatially-resolved representation) fails.
C9 · discarded-signal Uncalibrated block-transform features for perturbation detectionverified_trace · effect +0.0456

P1 — Uncalibrated block-transform features for perturbation detection

• Pattern: Detects a feature-extraction pipeline for detecting subtle perturbations in block-compressed images that computes block-transform statistics only on the aligned block grid and never contrasts them with the same statistics recomputed on a sub-block-shifted or cropped copy of the image, so no calibrated-difference features are produced.
• Detection procedure:
1. Locate a function that converts an image array into a numeric feature vector and, within its body, computes a block-wise frequency transform (e.g., a call to dctn, dct, or an equivalent block transform applied over fixed-size tiles such as 8x8) and aggregates statistics (histograms, means, variances, counts) of the transform coefficients.
2. Within the same function body (or any function it calls), check whether the input image is also sliced, cropped, or shifted by an offset that is not a multiple of the block size (e.g., a subscript like img[k:, k:] or img[a:b, c:d] where the start offsets are nonzero and not divisible by the tile size) before the same block transform is applied a second time.
3. Check whether any feature values are formed by subtracting or otherwise combining statistics from step 1 with statistics from a misaligned copy as in step 2.
4. PRESENT if step 1 holds and neither step 2 nor step 3 holds anywhere in the feature-extraction code path.
• Predicted impact:
• Add score for C9 (Available signal never used): 2
• weight: 2
• confidence: high
• Evidence: A pipeline computing only aligned-grid statistics such as dctn(block) histograms per 8x8 tile, with no shifted-crop reference, produced features dominated by image content; the variant that added differences between aligned and grid-misaligned statistics scored ~17% higher relative on the weighted-AUC metric.
• Applies when: The task is detecting subtle, embedded, or tampering-induced perturbations in block-compressed images, and hand-crafted block-transform statistics are used as model features.
• Example:
• Input:
    def extract_features(img):
        feats = []
        H, W = img.shape
        for i in range(0, H - 7, 8):
            for j in range(0, W - 7, 8):
                block = img[i:i+8, j:j+8].astype(np.float64)
                c = dctn(block, norm="ortho")
                feats.append(np.abs(c).ravel())
        F = np.vstack(feats)
        return np.concatenate([F.mean(0), F.var(0)])
• Consequence:
    Features vary mainly with image content rather than the perturbation;
    validation weighted AUC lands ~15-20% relative below a pipeline that
    also subtracts the same statistics from a few-pixel-shifted crop.
• Counter-example:
• Input:
    def extract_features(img):
        def stats(x):
            fs = []
            for i in range(0, x.shape[0] - 7, 8):
                for j in range(0, x.shape[1] - 7, 8):
                    c = dctn(x[i:i+8, j:j+8].astype(np.float64), norm="ortho")
                    fs.append(np.abs(c).ravel())
            F = np.vstack(fs)
            return np.concatenate([F.mean(0), F.var(0)])
        a = stats(img)
        b = stats(img[3:, 3:])          # grid-misaligned calibrated reference
        return np.concatenate([a, a - b])
• Why it does not fire: The same block-transform statistics are recomputed on a copy shifted by 3 pixels (not a multiple of the 8-pixel block size) and their difference is included as features, so the calibrated reference is present.