← main results | causal dashboard P/R comparison + traces sanity + distributions rubric library (183)

Conceptual rubrics — causal traces

pair standard: annotator marks rubrics explaining a weak-vs-strong gap · causal standard: patch the fire, re-run, re-grade · causal precision = fires improving above noise (4) / tested fires (56) · causal recall = states with ≥1 above-noise-improving fire (3) / tested states with above-noise gap to best (11) · generated 2026-08-08 21:02:16 UTC
improved +66.367% Raw probabilities submitted without clipping or prior-blending under log loss leaf-classification
rubric c3__metric-mismatch__6ab5ffe7927a
pattern Detects a probability-scored multiclass submission where the model's predicted class probabilities are written to the output file with no clipping away from zero and no blending or shrinkage toward a uniform or prior distribution, leaving confident wrong predictions to incur near-unbounded per-row log-loss penalties.
judge evidence for the fire probs = final_clf.predict_proba(X_test_scaled) written directly to CSV; no clipping or uniform blending applied.
proposed fix applied After predict_proba, blend with a uniform prior (e.g. p = 0.9*p + 0.1/K) or clip probabilities to [eps, 1-eps] and renormalize before writing the submission.
score 0.48765 → 0.16401 (improvement +0.323640, relative +66.367%, noise floor ±1.54%)
state 10f3a5199a87 · leaf-classification/init-004
improved +7.476% Capped word-only TF-IDF with no character n-gram features detecting-insults-in-social-commentary
rubric c9__discarded-signal__1885509badbd
pattern Detects a text-classification pipeline whose only text representation is a word-level TF-IDF with a small hardcoded vocabulary cap (a few thousand features or fewer) and no character-level n-gram feature block, discarding sub-word stylistic signal present in the raw text.
judge evidence for the fire Single TfidfVectorizer(max_features=5000, ngram_range=(1,2)) word-level only in Pipeline; no char analyzer or hstack of char features.
proposed fix applied Remove or greatly raise the vocabulary cap and concatenate a character n-gram TF-IDF block (e.g. char 2-5 grams via sparse hstack) alongside word n-grams.
score 0.84778 → 0.91116 (improvement +0.063380, relative +7.476%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
improved +7.328% Word-token-only features for style-based text classification detecting-insults-in-social-commentary
rubric c9__discarded-signal__1a1906afc4d2
pattern Detects a text-classification pipeline whose label depends on writing style (all classes share the same domain/topic) that builds every feature matrix from word-level tokenization only, with no character-level n-gram representation anywhere before model fitting.
judge evidence for the fire Only TfidfVectorizer(ngram_range=(1,2), stop_words='english') word features; no analyzer='char'/'char_wb' or stacked char block anywhere.
proposed fix applied
score 0.84644 → 0.90847 (improvement +0.062030, relative +7.328%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
improved +5.679% Text represented only by scalar surface statistics random-acts-of-pizza
rubric c9__discarded-signal__50a89888e872
pattern Detects a supervised pipeline on a content-dependent text task where the model's feature matrix is built exclusively from scalar surface statistics of the text (character/word lengths, punctuation or marker counts) and no token-level lexical representation of the text is ever fitted or included.
judge evidence for the fire TfidfVectorizer imported but never used; features only text_length, word_count, has_question/thanks/please/exclamation plus metadata passed to LogisticRegression.fit
proposed fix applied Vectorize the text fields (e.g. TF-IDF on prompt+response, optionally dimensionality-reduced) and concatenate with the scalar surface features before fitting the model.
score 0.61379 → 0.64865 (improvement +0.034860, relative +5.679%, noise floor ±1.54%)
state 1502f0c270f1 · random-acts-of-pizza/init-000
harmed -173.898% Unclipped probabilities under an unbounded log-based metric leaf-classification
rubric c3__metric-mismatch__9c1fd08f97b7
pattern Detects probabilistic class predictions written directly to an output file scored by log loss (or a similar unbounded proper scoring rule) without any clipping away from 0 and 1, smoothing, or blending toward a uniform prior.
judge evidence for the fire probs = final_clf.predict_proba(X_test_scaled) written directly via to_csv with no clipping or uniform blending; metric is multiclass log loss
proposed fix applied
score 0.05988 → 0.16401 (improvement -0.104130, relative -173.898%, noise floor ±1.54%)
state 10f3a5199a87 · leaf-classification/init-004
harmed -1.653% Uncalibrated probabilities written to log-loss submission leaf-classification
rubric c3__metric-mismatch__7e04f476b882
pattern Detects predicted class probabilities from a fitted classifier written directly into a submission file scored by log loss, with no clipping, uniform-blend, or other bounding of extreme values between prediction and write.
judge evidence for the fire probs = final_clf.predict_proba(X_test_scaled) goes straight into DataFrame and to_csv with no clipping or uniform blend
proposed fix applied
score 0.05988 → 0.06087 (improvement -0.000990, relative -1.653%, noise floor ±1.54%)
state 10f3a5199a87 · leaf-classification/init-004
within noise -1.421% Train-once fit with no validation estimate before submission random-acts-of-pizza
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire clf.fit(X_train, target) and gb.fit(...) on all rows; no train_test_split/CV, no roc_auc_score, no eval_set before to_csv
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.68664 → 0.67688 (improvement -0.009760, relative -1.421%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise -0.742% Train-once fit with no validation estimate before submission denoising-dirty-documents
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire model.fit(X_train, y_train) on all rows, fixed hyperparameters, no train_test_split/KFold, no metric computed, no eval_set; writes submission directly
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.02021 → 0.02036 (improvement -0.000150, relative -0.742%, noise floor ±1.54%)
state c6a549b1c22c · denoising-dirty-documents/init-004
within noise +0.727% No held-out estimate of the scored metric before writing predictions random-acts-of-pizza
rubric c3__metric-mismatch__972aa6849971
pattern Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
judge evidence for the fire clf.fit(X_train, target) and gb.fit(...) on full data; no train_test_split/KFold/cross_val, no roc_auc_score; predictions written to_csv
proposed fix applied
score 0.68664 → 0.69163 (improvement +0.004990, relative +0.727%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise -0.655% Blind fit with no validation-guided capacity selection nomad2018-predict-transparent-conductors
rubric c5__budget-misuse__2dd35f9c2787
pattern Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
judge evidence for the fire XGBRegressor(n_estimators=100, max_depth=6, learning_rate=0.1).fit(X_train_scaled, y) on full data; train_test_split imported but never called; no CV/early stopping/metric.
proposed fix applied
score 0.07942 → 0.07994 (improvement -0.000520, relative -0.655%, noise floor ±1.54%)
state fa8cd512c368 · nomad2018-predict-transparent-conductors/init-000
within noise +0.608% No held-out estimate of the scored metric before writing predictions random-acts-of-pizza
rubric c3__metric-mismatch__972aa6849971
pattern Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
judge evidence for the fire model.fit(X_train_scaled, y_train) on full data; no train_test_split/CV/metric call anywhere; predict_proba then to_csv
proposed fix applied
score 0.61379 → 0.61752 (improvement +0.003730, relative +0.608%, noise floor ±1.54%)
state 1502f0c270f1 · random-acts-of-pizza/init-000
within noise -0.593% Capped word-only TF-IDF with no character n-gram features random-acts-of-pizza
rubric c9__discarded-signal__1885509badbd
pattern Detects a text-classification pipeline whose only text representation is a word-level TF-IDF with a small hardcoded vocabulary cap (a few thousand features or fewer) and no character-level n-gram feature block, discarding sub-word stylistic signal present in the raw text.
judge evidence for the fire TfidfVectorizer(max_features=3000, ngram_range=(1,2)) word-level only; hstack adds only numeric features, no char n-gram block.
proposed fix applied Remove or greatly raise the vocabulary cap and concatenate a character n-gram TF-IDF block (e.g. char 2-5 grams via sparse hstack) alongside word n-grams.
score 0.68664 → 0.68257 (improvement -0.004070, relative -0.593%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise -0.578% Blind fit with no validation-guided capacity selection random-acts-of-pizza
rubric c5__budget-misuse__2dd35f9c2787
pattern Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
judge evidence for the fire LogisticRegression(C=1.0...) and GradientBoostingClassifier(n_estimators=200...) fit on full train, literal hyperparameters, no split/CV/metric/early stopping; blended preds written directly.
proposed fix applied
score 0.68664 → 0.68267 (improvement -0.003970, relative -0.578%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise +0.558% Small hardcoded vocabulary cap on sparse text features for a linear model random-acts-of-pizza
rubric c5__budget-misuse__e1da43ff7311
pattern Detects a bag-of-words or TF-IDF text vectorizer constructed with a small hardcoded vocabulary-size cap (a literal of a few thousand or less) whose output feeds a sparse-capable linear classifier, discarding most of the n-gram feature space with no evidence the cap was tuned or validated.
judge evidence for the fire TfidfVectorizer(max_features=3000, ngram_range=(1,2), min_df=3) feeds LogisticRegression via hstack; cap is a hardcoded literal, untuned.
proposed fix applied
score 0.68664 → 0.69047 (improvement +0.003830, relative +0.558%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise +0.558% Small fixed vocabulary cap on sparse text features fed to sparse-capable models random-acts-of-pizza
rubric c9__discarded-signal__5a2ed37eb9b5
pattern Detects a sparse bag-of-words or TF-IDF text vectorizer constructed with a small fixed vocabulary-size cap (a few thousand or fewer features) whose output is consumed only by linear or probabilistic models that scale to high-dimensional sparse input, discarding the long tail of discriminative terms.
judge evidence for the fire TfidfVectorizer(max_features=3000, ...) output hstacked and consumed only by LogisticRegression; GradientBoosting uses numeric features only.
proposed fix applied
score 0.68664 → 0.69047 (improvement +0.003830, relative +0.558%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise +0.558% Small hardcoded vocabulary cap on a large text corpus random-acts-of-pizza
rubric c5__budget-misuse__44a4ebec0fd4
pattern Detects a text-vectorization step whose vocabulary is capped by a hardcoded feature limit in the low thousands (5,000 or fewer) while feeding a fast-to-train sparse linear or naive-Bayes model, truncating rarer discriminative n-grams despite ample compute headroom.
judge evidence for the fire TfidfVectorizer(max_features=3000, ...) is the only text vectorizer; its hstacked matrix feeds LogisticRegression.fit(X_train, target)
proposed fix applied Raise or remove the feature cap (e.g. tens of thousands of word plus character n-grams) when models are sparse-linear and runtime allows.
score 0.68664 → 0.69047 (improvement +0.003830, relative +0.558%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise +0.558% Hardcoded small vocabulary cap on sparse text features feeding a linear model random-acts-of-pizza
rubric c5__budget-misuse__7d7614d9ca1d
pattern Detects a text-to-sparse-matrix vectorizer constructed with a small hardcoded vocabulary-size cap whose output is consumed by a sparse-capable linear or naive-Bayes classifier, truncating the feature space the model could otherwise use.
judge evidence for the fire TfidfVectorizer(max_features=3000,...) fit_transform → hstack with numeric → LogisticRegression.fit; only vectorizer, capped.
proposed fix applied Omit `max_features` (or set it far above the natural vocabulary size) when feeding sparse features into linear/NB models; control noise with `min_df` and regularization instead.
score 0.68664 → 0.69047 (improvement +0.003830, relative +0.558%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise +0.558% Small hardcoded vocabulary cap on sparse text features for a linear model random-acts-of-pizza
rubric c9__discarded-signal__76b1d2e2316e
pattern Detects a text-vectorization step whose vocabulary is truncated by a small hardcoded feature limit (10,000 or fewer) before feeding a sparse linear or naive-Bayes classifier, discarding the long tail of rare discriminative terms.
judge evidence for the fire TfidfVectorizer(max_features=3000, ...) hstacked only with numeric features and fed to LogisticRegression.fit; no uncapped vectorizer present
proposed fix applied Remove the vocabulary cap (or set it far above the corpus vocabulary size) and rely on min_df plus regularization; sparse linear models handle hundreds of thousands of features cheaply.
score 0.68664 → 0.69047 (improvement +0.003830, relative +0.558%, noise floor ±1.54%)
state 3094c3d76da8 · random-acts-of-pizza/init-001
within noise +0.463% Blind fit with no validation-guided capacity selection random-acts-of-pizza
rubric c5__budget-misuse__2dd35f9c2787
pattern Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
judge evidence for the fire LogisticRegression(max_iter=1000, random_state=42) fit on full X_train_scaled, predict_proba on test, no train_test_split/CV/early stopping anywhere
proposed fix applied
score 0.61379 → 0.61663 (improvement +0.002840, relative +0.463%, noise floor ±1.54%)
state 1502f0c270f1 · random-acts-of-pizza/init-000
within noise +0.463% Train-once fit with no validation estimate before submission random-acts-of-pizza
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire model.fit(X_train_scaled, y_train) on all rows; no train_test_split/CV, no roc_auc_score, no eval_set; predictions written directly to submission.csv
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.61379 → 0.61663 (improvement +0.002840, relative +0.463%, noise floor ±1.54%)
state 1502f0c270f1 · random-acts-of-pizza/init-000
within noise +0.448% Validation score computed but never used to select anything tabular-playground-series-may-2022
rubric c5__budget-misuse__e1e508029bc9
pattern Detects a pipeline that splits off a validation set and evaluates a single fixed-hyperparameter model on it, then refits and predicts with those same hardcoded hyperparameters, so the validation metric influences no hyperparameter choice.
judge evidence for the fire train_test_split + eval_set=[(X_val,y_val)] with eval_metric='logloss', verbose=False, no early_stopping_rounds; single XGBClassifier with all-literal hyperparameters used for predictions.
proposed fix applied Loop over a small grid of the key regularization/capacity hyperparameter, score each candidate on the validation split with the task metric, and refit the best value on the full training set.
score 0.93833 → 0.94253 (improvement +0.004200, relative +0.448%, noise floor ±1.54%)
state f5d804f5e008 · tabular-playground-series-may-2022/init-000
within noise +0.366% Small hardcoded vocabulary cap on sparse text features for a linear model detecting-insults-in-social-commentary
rubric c5__budget-misuse__e1da43ff7311
pattern Detects a bag-of-words or TF-IDF text vectorizer constructed with a small hardcoded vocabulary-size cap (a literal of a few thousand or less) whose output feeds a sparse-capable linear classifier, discarding most of the n-gram feature space with no evidence the cap was tuned or validated.
judge evidence for the fire TfidfVectorizer(max_features=5000, ngram_range=(1,2)) in Pipeline with LogisticRegression, no tuning or search over the cap
proposed fix applied
score 0.84644 → 0.84954 (improvement +0.003100, relative +0.366%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise +0.366% Small fixed vocabulary cap on sparse text features fed to sparse-capable models detecting-insults-in-social-commentary
rubric c9__discarded-signal__5a2ed37eb9b5
pattern Detects a sparse bag-of-words or TF-IDF text vectorizer constructed with a small fixed vocabulary-size cap (a few thousand or fewer features) whose output is consumed only by linear or probabilistic models that scale to high-dimensional sparse input, discarding the long tail of discriminative terms.
judge evidence for the fire TfidfVectorizer(max_features=5000, ...) in Pipeline feeding only LogisticRegression
proposed fix applied
score 0.84644 → 0.84954 (improvement +0.003100, relative +0.366%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise -0.227% Train-once fit with no validation estimate before submission nomad2018-predict-transparent-conductors
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire model_formation.fit(X_train_scaled, y_formation) on all rows; train_test_split imported but unused; no eval_set, no held-out metric before to_csv
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.07942 → 0.0796 (improvement -0.000180, relative -0.227%, noise floor ±1.54%)
state fa8cd512c368 · nomad2018-predict-transparent-conductors/init-000
within noise +0.208% Small hardcoded vocabulary cap on a large text corpus detecting-insults-in-social-commentary
rubric c5__budget-misuse__44a4ebec0fd4
pattern Detects a text-vectorization step whose vocabulary is capped by a hardcoded feature limit in the low thousands (5,000 or fewer) while feeding a fast-to-train sparse linear or naive-Bayes model, truncating rarer discriminative n-grams despite ample compute headroom.
judge evidence for the fire TfidfVectorizer(max_features=5000, ngram_range=(1,2)) is sole vectorizer in Pipeline feeding LogisticRegression
proposed fix applied Raise or remove the feature cap (e.g. tens of thousands of word plus character n-grams) when models are sparse-linear and runtime allows.
score 0.84778 → 0.84954 (improvement +0.001760, relative +0.208%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise +0.208% Hardcoded small vocabulary cap on sparse text features feeding a linear model detecting-insults-in-social-commentary
rubric c5__budget-misuse__7d7614d9ca1d
pattern Detects a text-to-sparse-matrix vectorizer constructed with a small hardcoded vocabulary-size cap whose output is consumed by a sparse-capable linear or naive-Bayes classifier, truncating the feature space the model could otherwise use.
judge evidence for the fire TfidfVectorizer(max_features=5000, ngram_range=(1,2)) in Pipeline feeding LogisticRegression; sole text representation, no uncapped vectorizer stacked.
proposed fix applied Omit `max_features` (or set it far above the natural vocabulary size) when feeding sparse features into linear/NB models; control noise with `min_df` and regularization instead.
score 0.84778 → 0.84954 (improvement +0.001760, relative +0.208%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise +0.208% Small hardcoded vocabulary cap on sparse text features for a linear model detecting-insults-in-social-commentary
rubric c9__discarded-signal__76b1d2e2316e
pattern Detects a text-vectorization step whose vocabulary is truncated by a small hardcoded feature limit (10,000 or fewer) before feeding a sparse linear or naive-Bayes classifier, discarding the long tail of rare discriminative terms.
judge evidence for the fire TfidfVectorizer(max_features=5000, ngram_range=(1,2)) in Pipeline feeding LogisticRegression; no second uncapped vectorizer combined.
proposed fix applied Remove the vocabulary cap (or set it far above the corpus vocabulary size) and rely on min_df plus regularization; sparse linear models handle hundreds of thousands of features cheaply.
score 0.84778 → 0.84954 (improvement +0.001760, relative +0.208%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise -0.158% Train-once fit with no validation estimate before submission detecting-insults-in-social-commentary
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire pipeline.fit(X_train_text, y_train) on all rows; no train_test_split/CV, no metric computed, no eval_set before to_csv.
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.84778 → 0.84644 (improvement -0.001340, relative -0.158%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise -0.151% Blind fit with no validation-guided capacity selection nomad2018-predict-transparent-conductors
rubric c5__budget-misuse__2dd35f9c2787
pattern Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
judge evidence for the fire XGBRegressor(n_estimators=100, max_depth=6, learning_rate=0.1) fit on full train, predict test, no split/CV/early stopping anywhere
proposed fix applied
score 0.18563 → 0.18591 (improvement -0.000280, relative -0.151%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise -0.151% Train-once fit with no validation estimate before submission nomad2018-predict-transparent-conductors
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire XGBRegressor fit on all rows with fixed params, no train_test_split/KFold, no eval_set, no metric before to_csv
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.18563 → 0.18591 (improvement -0.000280, relative -0.151%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise +0.109% Train-once fit with no validation estimate before submission nomad2018-predict-transparent-conductors
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire model_fe.fit(X_train, y_train_fe) on all rows; no train_test_split/KFold, no metric computed, no eval_set before to_csv
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.06409 → 0.06402 (improvement +0.000070, relative +0.109%, noise floor ±1.54%)
state ee8c04d68c9a · nomad2018-predict-transparent-conductors/init-004
within noise -0.085% Single default-configured probabilistic model with no ensembling or validation random-acts-of-pizza
rubric c5__budget-misuse__9cc60fcf16ee
pattern Detects a script that fits exactly one probabilistic classifier, writes its predicted class probabilities straight to the output file, and contains no cross-validation loop, no second model, and no out-of-fold evaluation, leaving obvious cheap ensembling capacity unused.
judge evidence for the fire Single LogisticRegression(max_iter=1000, random_state=42, n_jobs=-1, solver='lbfgs'); predict_proba written directly to submission; no KFold, no OOF, no second model.
proposed fix applied Generate out-of-fold predictions from several fast diverse models (e.g. logistic regression, naive Bayes, calibrated SVM) and combine them by averaging or a stacking meta-learner selected on out-of-fold loss.
score 0.61379 → 0.61327 (improvement -0.000520, relative -0.085%, noise floor ±1.54%)
state 1502f0c270f1 · random-acts-of-pizza/init-000
within noise -0.031% Train-once fit with no validation estimate before submission nomad2018-predict-transparent-conductors
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire rf_fe.fit(X, y_fe) / rf_bg.fit(X, y_bg) on all rows; no split, no CV, no metric, no early stopping before to_csv
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.06388 → 0.0639 (improvement -0.000020, relative -0.031%, noise floor ±1.54%)
state 4ba18d257d77 · nomad2018-predict-transparent-conductors/init-001
within noise -0.015% Single default-configured probabilistic model with no ensembling or validation detecting-insults-in-social-commentary
rubric c5__budget-misuse__9cc60fcf16ee
pattern Detects a script that fits exactly one probabilistic classifier, writes its predicted class probabilities straight to the output file, and contains no cross-validation loop, no second model, and no out-of-fold evaluation, leaving obvious cheap ensembling capacity unused.
judge evidence for the fire Single LogisticRegression(C=1.0 default) fit once; predict_proba written directly to CSV; no KFold/CV, no second model or blending.
proposed fix applied Generate out-of-fold predictions from several fast diverse models (e.g. logistic regression, naive Bayes, calibrated SVM) and combine them by averaging or a stacking meta-learner selected on out-of-fold loss.
score 0.84778 → 0.84765 (improvement -0.000130, relative -0.015%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise +0.000% No held-out estimate of the scored metric before writing predictions detecting-insults-in-social-commentary
rubric c3__metric-mismatch__972aa6849971
pattern Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
judge evidence for the fire pipeline.fit(X_train_text, y_train) on full data; no train_test_split/CV/metric call; predict_proba then to_csv.
proposed fix applied
score 0.84644 → 0.84644 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise +0.000% No held-out estimate of the scored metric before writing predictions aerial-cactus-identification
rubric c3__metric-mismatch__972aa6849971
pattern Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
judge evidence for the fire clf.fit(X_train_scaled, y_train) on full data; no train_test_split/KFold/metric call; predict_proba then to_csv.
proposed fix applied
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% No held-out estimate of the scored metric before writing predictions nomad2018-predict-transparent-conductors
rubric c3__metric-mismatch__972aa6849971
pattern Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
judge evidence for the fire Both XGBRegressors fit on full X_train_scaled; train_test_split imported but never called; no metric computed; predictions written to submission.csv.
proposed fix applied
score 0.07942 → 0.07942 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state fa8cd512c368 · nomad2018-predict-transparent-conductors/init-000
within noise +0.000% No held-out estimate of the scored metric before writing predictions nomad2018-predict-transparent-conductors
rubric c3__metric-mismatch__972aa6849971
pattern Detects a script that fits a single predictive model on the entire training set and writes its test-set predictions to an output file without ever computing a validation, cross-validated, or out-of-fold value of any evaluation metric, leaving every model, feature, and hyperparameter choice unmeasured against the objective.
judge evidence for the fire Fits XGBRegressor on full train_X_scaled for both targets; no train_test_split/KFold/cross_val, no metric call; predicts test and to_csv.
proposed fix applied
score 0.18563 → 0.18563 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise +0.000% Blind fit with no validation-guided capacity selection detecting-insults-in-social-commentary
rubric c5__budget-misuse__2dd35f9c2787
pattern Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
judge evidence for the fire pipeline.fit(X_train_text, y_train) with literal TF-IDF/LogisticRegression params, predict_proba on test, no train_test_split, CV, or validation metric anywhere
proposed fix applied
score 0.84644 → 0.84644 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state f26051bba78d · detecting-insults-in-social-commentary/init-003
within noise +0.000% Blind fit with no validation-guided capacity selection aerial-cactus-identification
rubric c5__budget-misuse__2dd35f9c2787
pattern Detects a supervised model fitted once on the entire training frame with fixed hand-picked hyperparameters and used to predict the test set directly, while the script contains no held-out validation split, no validation metric computation, and no early-stopping or iteration-selection mechanism anywhere.
judge evidence for the fire RandomForestClassifier(n_estimators=100, max_depth=20, random_state=42) fit on all training data; no split, CV, metric, or early stopping anywhere.
proposed fix applied
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% Per-file serial extraction with per-sample inference calls aerial-cactus-identification
rubric c6__wasted-compute__29375256d4c1
pattern Detects a loop over many independent input files that extracts features and invokes model inference on each single feature vector inside the loop body, instead of parallelizing extraction and predicting once on the stacked feature matrix.
judge evidence for the fire Loop over sample_submission['id'] calls extract_features then scaler.transform([features]) and clf.predict_proba per file, serially, no pool.
proposed fix applied
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% Unprobed metadata-joined paths with silent per-file fallback nomad2018-predict-transparent-conductors
rubric c12__silent-degradation__9251124b4217
pattern Detects batch feature extraction over file paths assembled by joining a hardcoded directory prefix with per-record metadata fields, where each per-file read is wrapped in an error handler that substitutes a default value, and no existence check is performed on any constructed path before the batch loop begins.
judge evidence for the fire Path os.path.join(BASE, folder, str(id_), 'geometry.xyz') read in try/except that silently returns empty → NaN/median fallback; no existence check before loop.
proposed fix applied Before batch extraction, probe one constructed path with an existence check and switch between candidate layouts (with/without the directory prefix), or assert a minimum fraction of paths exist and abort otherwise.
score 0.06409 → 0.06409 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state ee8c04d68c9a · nomad2018-predict-transparent-conductors/init-004
within noise +0.000% Unprobed metadata-joined paths with silent per-file fallback nomad2018-predict-transparent-conductors
rubric c12__silent-degradation__9251124b4217
pattern Detects batch feature extraction over file paths assembled by joining a hardcoded directory prefix with per-record metadata fields, where each per-file read is wrapped in an error handler that substitutes a default value, and no existence check is performed on any constructed path before the batch loop begins.
judge evidence for the fire f-string geom_path from literal prefix + row['id']; extract_geometry_features has bare `except: return None` replaced by {} then fillna(0); no exists/glob probe before loop
proposed fix applied Before batch extraction, probe one constructed path with an existence check and switch between candidate layouts (with/without the directory prefix), or assert a minimum fraction of paths exist and abort otherwise.
score 0.18563 → 0.18563 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise +0.000% Zero-block stand-in for missing feature source before scaling nomad2018-predict-transparent-conductors
rubric c4__transform-skew__b6484536490f
pattern Detects per-sample feature construction that substitutes an all-zero block for an absent source or modality and concatenates it with real features that are subsequently standardized, with no fitted imputation step distinguishing absence from measured zeros.
judge evidence for the fire except: return None → geom_features={} then pd.DataFrame(...).fillna(0), plus row.get('lv1',0) zeros; matrix then StandardScaler.fit_transform with no imputer/indicator
proposed fix applied Fill missing-source feature blocks with NaN and apply a `SimpleImputer` fit on training data to both train and test matrices before scaling.
score 0.18563 → 0.18563 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise +0.000% Train-once fit with no validation estimate before submission aerial-cactus-identification
rubric c5__budget-misuse__60f0ba653435
pattern Detects a supervised model fitted once with fixed hyperparameters on all assembled training rows and used to predict the test set directly, with no held-out validation split, no early stopping, and no computation of any evaluation metric before the submission file is written.
judge evidence for the fire clf.fit(X_train_scaled, y_train) on all rows; no train_test_split, no CV, no metric computed, predictions written directly to CSV
proposed fix applied Hold out a validation split (group-aware if samples are derived variants of a shared source), select iterations/hyperparameters via early stopping on it, then optionally refit on all data at the selected budget.
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% Validation set supplied to boosting fit without early stopping tabular-playground-series-may-2022
rubric c5__budget-misuse__a0167b4b8af9
pattern Detects a gradient-boosted model fitted with a validation evaluation set and a large fixed iteration budget but no early-stopping mechanism, so the validation scores are computed yet never used to halt training at the best iteration.
judge evidence for the fire fit(eval_set=[(X_val,y_val)]) with n_estimators=500 and no early_stopping_rounds/callbacks; comment claims early stopping but none configured
proposed fix applied Attach an early-stopping callback (e.g. `lgb.early_stopping(patience)`) whenever an eval_set is provided, and predict with the best iteration.
score 0.93833 → 0.93833 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state f5d804f5e008 · tabular-playground-series-may-2022/init-000
within noise +0.000% Boosting fit with eval_set but no early stopping tabular-playground-series-may-2022
rubric c5__budget-misuse__fcd923ce3aa1
pattern Detects a gradient-boosted ensemble fitted with a large fixed boosting-round count and a supplied validation evaluation set, but with no early-stopping mechanism registered, so training always runs to the full round count regardless of validation performance.
judge evidence for the fire XGBClassifier(n_estimators=500,...) fit with eval_set=[(X_val,y_val)] and no early_stopping_rounds/callbacks despite comment claiming early stopping
proposed fix applied Attach an early-stopping callback (or `early_stopping_rounds`) keyed to the validation metric whenever an eval_set is supplied to a boosting fit.
score 0.93833 → 0.93833 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state f5d804f5e008 · tabular-playground-series-may-2022/init-000
within noise +0.000% Per-row model inference inside a Python loop aerial-cactus-identification
rubric c6__wasted-compute__24fcb3d222d6
pattern Detects a fitted model's prediction call (and any preprocessing transform feeding it) being invoked on single-row inputs inside a Python loop over evaluation samples, instead of one batched call on a stacked feature matrix.
judge evidence for the fire for image_id in sample_submission['id']: scaler.transform([features]); clf.predict_proba(features_scaled)[0][1]; predictions.append(...)
proposed fix applied Accumulate test features into one 2D array, then call transform and predict_proba once on the full matrix.
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% Sequential per-file feature extraction over a large file corpus nomad2018-predict-transparent-conductors
rubric c6__wasted-compute__62bb3397ad68
pattern Detects a single-process sequential loop that reads and computes features from each file in a large collection of independent input files, with no process-pool or worker-based parallel map anywhere in the extraction stage.
judge evidence for the fire Sequential `for idx, row in train_df.iterrows()` loops open each geometry.xyz and compute pairwise-distance features; no ProcessPoolExecutor/joblib anywhere.
proposed fix applied Map the per-file feature function over the file list with a process pool sized to available cores, using chunked submission.
score 0.18563 → 0.18563 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise +0.000% Sequential per-file feature extraction over a large file corpus nomad2018-predict-transparent-conductors
rubric c6__wasted-compute__62bb3397ad68
pattern Detects a single-process sequential loop that reads and computes features from each file in a large collection of independent input files, with no process-pool or worker-based parallel map anywhere in the extraction stage.
judge evidence for the fire build_features loops `for idx, row in df.iterrows()` over 2400 ids, opening each geometry.xyz and computing O(N^2) Python pairwise distances; no ProcessPool/joblib anywhere.
proposed fix applied Map the per-file feature function over the file list with a process pool sized to available cores, using chunked submission.
score 0.07942 → 0.07942 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state fa8cd512c368 · nomad2018-predict-transparent-conductors/init-000
within noise +0.000% Sequential per-file feature extraction over a large image collection aerial-cactus-identification
rubric c6__wasted-compute__ac9b0b8b0bcc
pattern Detects a single-process `for` loop that computes per-file features for every image in a directory listing one at a time, with no worker pool, thread pool, or batched/vectorized extraction path anywhere in the script.
judge evidence for the fire Sequential loops `for idx, row in train_df.iterrows()` and `for image_id in sample_submission['id']` call extract_features per file; no Pool/Parallel anywhere.
proposed fix applied Map the pure per-file extraction function over the path list with a multiprocessing pool sized to available CPUs, falling back to the sequential loop on failure.
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% Sequential per-file feature extraction on large file sets aerial-cactus-identification
rubric c6__wasted-compute__daffeafdd94a
pattern Detects feature extraction that iterates one-by-one over a large collection of independent data files in a plain sequential Python loop, loading and processing each file in the main process with no process pool, worker mapping, or parallel batching.
judge evidence for the fire `for idx, row in train_df.iterrows(): ... extract_features(image_path)` and per-id test loop; no ProcessPoolExecutor/joblib anywhere
proposed fix applied Map the per-file feature function over paths with a ProcessPoolExecutor sized to available cores, with a serial fallback.
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000
within noise +0.000% Sequential per-file feature extraction on large file sets nomad2018-predict-transparent-conductors
rubric c6__wasted-compute__daffeafdd94a
pattern Detects feature extraction that iterates one-by-one over a large collection of independent data files in a plain sequential Python loop, loading and processing each file in the main process with no process pool, worker mapping, or parallel batching.
judge evidence for the fire `for idx, row in df.iterrows()` in build_features opens each geometry.xyz serially for 2160+240 files; no ProcessPoolExecutor/joblib anywhere.
proposed fix applied Map the per-file feature function over paths with a ProcessPoolExecutor sized to available cores, with a serial fallback.
score 0.07942 → 0.07942 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state fa8cd512c368 · nomad2018-predict-transparent-conductors/init-000
within noise +0.000% Sequential per-file feature extraction on large file sets nomad2018-predict-transparent-conductors
rubric c6__wasted-compute__daffeafdd94a
pattern Detects feature extraction that iterates one-by-one over a large collection of independent data files in a plain sequential Python loop, loading and processing each file in the main process with no process pool, worker mapping, or parallel batching.
judge evidence for the fire Sequential `for idx, row in train_df.iterrows()` (and test) calling extract_geometry_features(open/parse + O(n^2) distance loop) over 2400 files; no pool/joblib anywhere.
proposed fix applied Map the per-file feature function over paths with a ProcessPoolExecutor sized to available cores, with a serial fallback.
score 0.18563 → 0.18563 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 55c93d8bac27 · nomad2018-predict-transparent-conductors/init-003
within noise +0.000% Sequential per-file feature extraction on large file sets nomad2018-predict-transparent-conductors
rubric c6__wasted-compute__daffeafdd94a
pattern Detects feature extraction that iterates one-by-one over a large collection of independent data files in a plain sequential Python loop, loading and processing each file in the main process with no process pool, worker mapping, or parallel batching.
judge evidence for the fire compute_geom_features opens/parses geometry.xyz inside `for id_ in df['id']` for 2160+240 files; no ProcessPoolExecutor/joblib anywhere.
proposed fix applied Map the per-file feature function over paths with a ProcessPoolExecutor sized to available cores, with a serial fallback.
score 0.06409 → 0.06409 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state ee8c04d68c9a · nomad2018-predict-transparent-conductors/init-004
within noise +0.000% Grayscale collapse discards chroma signal in subtle-signal image task aerial-cactus-identification
rubric c9__discarded-signal__2dc268fdea17
pattern Detects a feature-extraction pipeline for detecting subtle or hidden signals in color images that converts each image to a single grayscale/luminance channel before computing features, discarding the chroma channels that may carry part of the target signal.
judge evidence for the fire extract_features uses Image.open(path).convert('L'); all HOG and stats computed on single grayscale plane, no per-channel features; AUC task on color aerial images.
proposed fix applied Decode to a multi-channel color space (e.g. YCbCr) and compute per-channel statistics/residual features rather than collapsing to one luminance channel.
score 0.5 → 0.5 (improvement +0.000000, relative +0.000%, noise floor ±1.54%)
state 6ff07afa1981 · aerial-cactus-identification/init-000