mle-rubrics-onpolicy-10

MLE-bench · 10 rubrics · first 10 of all
HF EdwardoSunny/mle-rubrics-onpolicy-10 · local data/libraries/mle-rubrics-onpolicy-10.json

0No held-out validation against the competition's stated metric before submittingtaskda-code
Applies when
task -- the task specifies an explicit evaluation metric (e.g., log loss, AUC, RMSE) for predictions on an unlabeled test set, and the script trains models and writes a submission file.
Pattern
The script fits one or more models on 100% of the labeled data, blends them with hand-picked weights, and writes predictions straight to disk — never computing the stated metric on a validation split, cross-validation folds, or even against a trivial baseline (e.g., class-prior probabilities). Hyperparameters, imputation choices, and ensemble weights are therefore unjustified, and there is no evidence the output is better than random or than a constant prediction.
Detection procedure
  1. Read the task and note the exact scoring metric and its inputs (probabilities vs. labels, per-row normalization, clipping).
  2. Scan the script for any train/validation split, cross_val_score/KFold, or explicit computation of that metric on labeled data; check whether printed diagnostics include a metric value at all.
  3. Check whether model/ensemble choices (weights, depth, learning rate) are tied to any measured score, or are literal constants written by the author.
  4. Inspect the answer/submission: is any score, baseline comparison, or sanity check (row count equal to test rows, ids matching test ids, probability ranges/sums, no degenerate constant column) reported?
Discriminator
A real violation is when no estimate of the stated metric exists anywhere for any candidate model, so the attempt cannot distinguish a good submission from a broken one. It is not a violation if the script measures the metric via CV/holdout (even briefly) and uses it to pick among options, nor if the metric is unmeasurable because labels genuinely do not exist for any subset.
Consequence
The submission may be systematically miscalibrated, mis-ordered, mislabeled by class column, or simply far worse than a simple baseline; the grader reports a failing/incorrect submission with no diagnostic trail, since the agent had no internal score to catch it.
id 83f41d76ebb5 · mined from da-code dacode-ml-competition-005
raw text (what the judge reads)
### No held-out validation against the competition's stated metric before submitting
- **Applies when**: `task` -- the task specifies an explicit evaluation metric (e.g., log loss, AUC, RMSE) for predictions on an unlabeled test set, and the script trains models and writes a submission file.
- **Pattern**: The script fits one or more models on 100% of the labeled data, blends them with hand-picked weights, and writes predictions straight to disk — never computing the stated metric on a validation split, cross-validation folds, or even against a trivial baseline (e.g., class-prior probabilities). Hyperparameters, imputation choices, and ensemble weights are therefore unjustified, and there is no evidence the output is better than random or than a constant prediction.
- **Detection procedure**:
  1. Read the task and note the exact scoring metric and its inputs (probabilities vs. labels, per-row normalization, clipping).
  2. Scan the script for any train/validation split, `cross_val_score`/`KFold`, or explicit computation of that metric on labeled data; check whether printed diagnostics include a metric value at all.
  3. Check whether model/ensemble choices (weights, depth, learning rate) are tied to any measured score, or are literal constants written by the author.
  4. Inspect the answer/submission: is any score, baseline comparison, or sanity check (row count equal to test rows, ids matching test ids, probability ranges/sums, no degenerate constant column) reported?
- **Discriminator**: A real violation is when *no* estimate of the stated metric exists anywhere for any candidate model, so the attempt cannot distinguish a good submission from a broken one. It is *not* a violation if the script measures the metric via CV/holdout (even briefly) and uses it to pick among options, nor if the metric is unmeasurable because labels genuinely do not exist for any subset.
- **Consequence**: The submission may be systematically miscalibrated, mis-ordered, mislabeled by class column, or simply far worse than a simple baseline; the grader reports a failing/incorrect submission with no diagnostic trail, since the agent had no internal score to catch it.
1Ships model predictions with no held-out validation and no distribution sanity checktaskda-code
Applies when
task -- a script trains a regressor/classifier on one file and writes predictions for another file as the deliverable, with no ground truth available for the predicted rows.
Pattern
The attempt fits a single model on 100% of the training rows, immediately predicts on the target rows, and writes the output without (a) any hold-out/CV error estimate, (b) a comparison of the predicted value distribution against the training target distribution, or (c) a check that the feature columns/dtypes/encodings used at prediction time are the same as those learned at fit time. Degenerate output (e.g. the vast majority of predictions collapsed to the same value, or a range far narrower than the target's) is accepted as-is.
Detection procedure
  1. Read the task: identify the deliverable (file, required column names/order, row alignment, rounding/units) and note that no labels exist for the predicted rows.
  2. Read the script: check whether a train/validation split or cross-validation with a printed error metric exists, and whether the same column names, imputation, and category encodings are applied consistently to both datasets (watch for near-identical but non-identical column spellings, or one file's schema being assumed for the other).
  3. Read the script's output stage: check whether the predicted values' summary statistics are compared to the training target's summary statistics, and whether any explicit guard rejects a degenerate/off-scale prediction vector.
  4. Inspect the submitted answer: compute the share of identical values and the min/max; if predictions are dominated by a single value or their spread is an order of magnitude off the target's documented spread, and no validation metric was reported, flag it.
Discriminator
A real violation is an unvalidated pipeline whose output is visibly degenerate or whose feature handling silently differs between fit and predict; it is not a violation if the script reports a hold-out/CV score and the prediction distribution plausibly matches the training target (a skewed target legitimately yields many small values, provided the reported validation error supports it).
Consequence
The written file passes format checks but its values are near-constant/mis-scaled, so the grader's accuracy or error tolerance against the reference targets fails, marking the deliverable WRONG.
id 4e7f8b275c9f · mined from da-code dacode-ml-regression-008
raw text (what the judge reads)
### Ships model predictions with no held-out validation and no distribution sanity check
- **Applies when**: `task` -- a script trains a regressor/classifier on one file and writes predictions for another file as the deliverable, with no ground truth available for the predicted rows.
- **Pattern**: The attempt fits a single model on 100% of the training rows, immediately predicts on the target rows, and writes the output without (a) any hold-out/CV error estimate, (b) a comparison of the predicted value distribution against the training target distribution, or (c) a check that the feature columns/dtypes/encodings used at prediction time are the same as those learned at fit time. Degenerate output (e.g. the vast majority of predictions collapsed to the same value, or a range far narrower than the target's) is accepted as-is.
- **Detection procedure**:
  1. Read the task: identify the deliverable (file, required column names/order, row alignment, rounding/units) and note that no labels exist for the predicted rows.
  2. Read the script: check whether a train/validation split or cross-validation with a printed error metric exists, and whether the same column names, imputation, and category encodings are applied consistently to both datasets (watch for near-identical but non-identical column spellings, or one file's schema being assumed for the other).
  3. Read the script's output stage: check whether the predicted values' summary statistics are compared to the training target's summary statistics, and whether any explicit guard rejects a degenerate/off-scale prediction vector.
  4. Inspect the submitted answer: compute the share of identical values and the min/max; if predictions are dominated by a single value or their spread is an order of magnitude off the target's documented spread, and no validation metric was reported, flag it.
- **Discriminator**: A real violation is an unvalidated pipeline whose output is visibly degenerate or whose feature handling silently differs between fit and predict; it is *not* a violation if the script reports a hold-out/CV score and the prediction distribution plausibly matches the training target (a skewed target legitimately yields many small values, provided the reported validation error supports it).
- **Consequence**: The written file passes format checks but its values are near-constant/mis-scaled, so the grader's accuracy or error tolerance against the reference targets fails, marking the deliverable WRONG.
2Hypothesis test run on the full table instead of the task-specified population/subsettaskda-code
Applies when
task -- the task asks for a statistical test, metric, or comparison whose population is qualified by filters (date range, category/tournament type, subgroup, one-sided direction) stated in the prompt, README, or problem context.
Pattern
The attempt loads the raw files and computes the statistic on every row of both groups, silently dropping the qualifying conditions (and/or defaulting to a two-sided test when a directional hypothesis was specified), so the reported p-value/metric describes a different population than the one asked about.
Detection procedure
  1. Read the task and any README, and list every explicit or implied restriction on the rows/columns to be analyzed (time window, subset of categories, groups compared) plus the alternative hypothesis direction and significance level.
  2. Read the scripts and check that each restriction appears as a concrete filter/dtype conversion (e.g., date parsing then range filter, category equality filter) before the statistic is computed, and that the test call matches the stated alternative and test type (paired vs independent, equal-variance assumption).
  3. Compare the row counts used in the test against the raw file row counts; if the script never prints or reduces counts, treat the population as unverified.
  4. Check the reported statistic's magnitude for plausibility given the intended (usually much smaller) subset — extreme p-values (e.g., 1e-100 or smaller) usually signal a far larger n than intended.
Discriminator
A real violation is a missing or incorrect filter/direction that changes which rows enter the computation; a look-alike that is fine is a script that applies the filters in a different but equivalent way (e.g., filtering at load time, using a query string) and can show the reduced counts/subset consistent with the task description.
Consequence
The reported p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample, so it will not match the expected value in the output file and the check fails even though the file format is correct.
id e6246376b83e · mined from da-code dacode-data-sa-001
raw text (what the judge reads)
### Hypothesis test run on the full table instead of the task-specified population/subset
- **Applies when**: `task` -- the task asks for a statistical test, metric, or comparison whose population is qualified by filters (date range, category/tournament type, subgroup, one-sided direction) stated in the prompt, README, or problem context.
- **Pattern**: The attempt loads the raw files and computes the statistic on every row of both groups, silently dropping the qualifying conditions (and/or defaulting to a two-sided test when a directional hypothesis was specified), so the reported p-value/metric describes a different population than the one asked about.
- **Detection procedure**:
  1. Read the task and any README, and list every explicit or implied restriction on the rows/columns to be analyzed (time window, subset of categories, groups compared) plus the alternative hypothesis direction and significance level.
  2. Read the scripts and check that each restriction appears as a concrete filter/dtype conversion (e.g., date parsing then range filter, category equality filter) before the statistic is computed, and that the test call matches the stated alternative and test type (paired vs independent, equal-variance assumption).
  3. Compare the row counts used in the test against the raw file row counts; if the script never prints or reduces counts, treat the population as unverified.
  4. Check the reported statistic's magnitude for plausibility given the intended (usually much smaller) subset — extreme p-values (e.g., 1e-100 or smaller) usually signal a far larger n than intended.
- **Discriminator**: A real violation is a missing or incorrect filter/direction that changes which rows enter the computation; a look-alike that is fine is a script that applies the filters in a different but equivalent way (e.g., filtering at load time, using a query string) and can show the reduced counts/subset consistent with the task description.
- **Consequence**: The reported p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample, so it will not match the expected value in the output file and the check fails even though the file format is correct.
3Output template file never actually inspectedtaskda-code
Applies when
task -- the task says to write results into an output file "following the exact structure/format of" a provided sample/template file, and the scripts build that output.
Pattern
The attempt never loads or prints the sample/template file; it hard-codes guessed column names, column order, row ordering, rounding/units and header text from the prose of the task, then asserts in the answer that the output "matches the required format."
Detection procedure
  1. In the task, note that a reference/sample output file is supplied and is the authority on schema and formatting.
  2. Search the scripts for any read/print/comparison of that sample file (e.g., loading it, checking its columns, dtypes, row count, decimal places, sort order).
  3. If absent, check whether the emitted frame's column names, column order, sort order and numeric formatting are instead invented in code or copied from the task wording.
  4. Check the answer for unverified claims of format compliance (no printed diff against the template).
Discriminator
A real violation is when no code path ever reads the template, so agreement with it is pure luck; it is fine if the script reads the template (or explicitly reindexes/renames/rounds/sorts to the template's columns and formatting) and prints a shape/column/dtype comparison, even if the final naming happens to be hard-coded afterwards.
Consequence
The graded file is judged WRONG/MISSING on a strict file comparison — mismatched header names or order, wrong row ordering, or unrounded/differently scaled values — even when the underlying aggregation logic is right.
id 375545aa1e68 · mined from da-code dacode-dm-csv-011
raw text (what the judge reads)
### Output template file never actually inspected
- **Applies when**: `task` -- the task says to write results into an output file "following the exact structure/format of" a provided sample/template file, and the scripts build that output.
- **Pattern**: The attempt never loads or prints the sample/template file; it hard-codes guessed column names, column order, row ordering, rounding/units and header text from the prose of the task, then asserts in the answer that the output "matches the required format."
- **Detection procedure**:
  1. In the task, note that a reference/sample output file is supplied and is the authority on schema and formatting.
  2. Search the scripts for any read/print/comparison of that sample file (e.g., loading it, checking its columns, dtypes, row count, decimal places, sort order).
  3. If absent, check whether the emitted frame's column names, column order, sort order and numeric formatting are instead invented in code or copied from the task wording.
  4. Check the answer for unverified claims of format compliance (no printed diff against the template).
- **Discriminator**: A real violation is when no code path ever reads the template, so agreement with it is pure luck; it is fine if the script reads the template (or explicitly reindexes/renames/rounds/sorts to the template's columns and formatting) and prints a shape/column/dtype comparison, even if the final naming happens to be hard-coded afterwards.
- **Consequence**: The graded file is judged WRONG/MISSING on a strict file comparison — mismatched header names or order, wrong row ordering, or unrounded/differently scaled values — even when the underlying aggregation logic is right.
4Unverified output artifact: schema/content of the saved file never checked against the requested spectaskda-code
Applies when
task -- the task requires writing results to a named file with an explicitly specified column naming scheme and one row per input record, and the scripts generate that file programmatically.
Pattern
The attempt builds a feature matrix after silently dropping/imputing/encoding columns, fits a model, writes the file, and then reports a prose summary (counts, chosen hyperparameter, feature legend) without ever re-reading the written file to confirm it exists at the expected path and that its header names, column count, row count, and label values match the requested format; degenerate or minimal-complexity results (e.g., the smallest possible number of groups) are accepted without a sanity check.
Detection procedure
  1. From the task, list the required artifact path, the exact column-naming convention, the expected number of rows (records) and the expected value semantics of the label column.
  2. In the scripts, locate the write call and check whether the DataFrame passed in has columns generated to match the convention exactly (no index column, no leftover original names, no extra/missing feature columns relative to the matrix actually clustered) and whether every input record survives preprocessing (no silent row drops from NaN handling or filtering).
  3. Check for a read-back/assert step after writing: does any code load the file and print/verify shape, header, row count, and label distribution? Compare that to what the final answer claims.
  4. Inspect the reported result for degeneracy or implausibility (a single dominant group, the minimum possible number of clusters, row count ≠ dataset size) that no validation step addressed.
Discriminator
A real violation is when the answer's claims about the file are asserted from in-memory variables or narrative only, with no post-write verification and no shape/format assertion — or when preprocessing changed the row/column set without reconciling it to the spec. It is not a violation if the script (or a follow-up run) reloads the artifact and asserts the header pattern, row count equal to the number of input records, and valid label values, even if the modeling choices are debatable.
Consequence
The grader reads the artifact and finds it missing, misnamed, mis-headered, or with the wrong number of rows/columns (or a degenerate labeling), scoring the file check WRONG/MISSING despite a confident-sounding summary.
id 2fad5c23094f · mined from da-code dacode-ml-cluster-014
raw text (what the judge reads)
### Unverified output artifact: schema/content of the saved file never checked against the requested spec
- **Applies when**: `task` -- the task requires writing results to a named file with an explicitly specified column naming scheme and one row per input record, and the scripts generate that file programmatically.
- **Pattern**: The attempt builds a feature matrix after silently dropping/imputing/encoding columns, fits a model, writes the file, and then reports a prose summary (counts, chosen hyperparameter, feature legend) without ever re-reading the written file to confirm it exists at the expected path and that its header names, column count, row count, and label values match the requested format; degenerate or minimal-complexity results (e.g., the smallest possible number of groups) are accepted without a sanity check.
- **Detection procedure**:
  1. From the task, list the required artifact path, the exact column-naming convention, the expected number of rows (records) and the expected value semantics of the label column.
  2. In the scripts, locate the write call and check whether the DataFrame passed in has columns generated to match the convention exactly (no index column, no leftover original names, no extra/missing feature columns relative to the matrix actually clustered) and whether every input record survives preprocessing (no silent row drops from NaN handling or filtering).
  3. Check for a read-back/assert step after writing: does any code load the file and print/verify shape, header, row count, and label distribution? Compare that to what the final answer claims.
  4. Inspect the reported result for degeneracy or implausibility (a single dominant group, the minimum possible number of clusters, row count ≠ dataset size) that no validation step addressed.
- **Discriminator**: A real violation is when the answer's claims about the file are asserted from in-memory variables or narrative only, with no post-write verification and no shape/format assertion — or when preprocessing changed the row/column set without reconciling it to the spec. It is *not* a violation if the script (or a follow-up run) reloads the artifact and asserts the header pattern, row count equal to the number of input records, and valid label values, even if the modeling choices are debatable.
- **Consequence**: The grader reads the artifact and finds it missing, misnamed, mis-headered, or with the wrong number of rows/columns (or a degenerate labeling), scoring the file check WRONG/MISSING despite a confident-sounding summary.
5Submission artifact never validated against the provided template (path, columns, ids, dtype)taskda-code
Applies when
task -- the task requires writing a prediction/result file whose format is defined by a sample/template file, and the scripts generate that file at a hard-coded location.
Pattern
The attempt builds the output frame from its own assumptions (its own column names, its own output directory, its own value dtype such as forcibly rounding/casting continuous predictions), prints a few rows as "verification", and never programmatically loads the template to confirm the file is written where the task expects, with the same column names/order, the same number and set of identifiers, and value types consistent with the evaluation metric.
Detection procedure
  1. In the task/README, note the required output filename, its expected location relative to the working directory, and the template file that defines its schema.
  2. In the scripts, find every write of the output file: check the path string, the constructed column names/order, and any post-processing of predictions (rounding, clipping, int casting, sorting).
  3. Check whether any script actually reads the template and asserts equality of columns, row count, and id set/order against the written file — printing head() or shapes without comparison does not count.
  4. In the answer, check whether it states the verified output location and schema match, or only narrates modeling choices and CV scores.
Discriminator
A real violation is when no code compares the produced file to the template/required path, or when values are transformed in a way the metric does not ask for (e.g., integer rounding of a continuous/log-scale target). It is fine if the script asserts column equality, id alignment and row counts against the template (even implicitly by copying the template and overwriting the prediction column) and writes to the location the task specifies.
Consequence
The grader reports the expected result file as missing or wrong (not found at the expected path, mismatched columns/ids, or degraded score from unnecessary value transformation), so the attempt scores 0 despite a plausible-looking model and CV metrics.
id 5c50253d9364 · mined from da-code dacode-ml-competition-009
raw text (what the judge reads)
### Submission artifact never validated against the provided template (path, columns, ids, dtype)
- **Applies when**: `task` -- the task requires writing a prediction/result file whose format is defined by a sample/template file, and the scripts generate that file at a hard-coded location.
- **Pattern**: The attempt builds the output frame from its own assumptions (its own column names, its own output directory, its own value dtype such as forcibly rounding/casting continuous predictions), prints a few rows as "verification", and never programmatically loads the template to confirm the file is written where the task expects, with the same column names/order, the same number and set of identifiers, and value types consistent with the evaluation metric.
- **Detection procedure**:
  1. In the task/README, note the required output filename, its expected location relative to the working directory, and the template file that defines its schema.
  2. In the scripts, find every write of the output file: check the path string, the constructed column names/order, and any post-processing of predictions (rounding, clipping, int casting, sorting).
  3. Check whether any script actually reads the template and asserts equality of columns, row count, and id set/order against the written file — printing `head()` or shapes without comparison does not count.
  4. In the answer, check whether it states the verified output location and schema match, or only narrates modeling choices and CV scores.
- **Discriminator**: A real violation is when no code compares the produced file to the template/required path, or when values are transformed in a way the metric does not ask for (e.g., integer rounding of a continuous/log-scale target). It is fine if the script asserts column equality, id alignment and row counts against the template (even implicitly by copying the template and overwriting the prediction column) and writes to the location the task specifies.
- **Consequence**: The grader reports the expected result file as missing or wrong (not found at the expected path, mismatched columns/ids, or degraded score from unnecessary value transformation), so the attempt scores 0 despite a plausible-looking model and CV metrics.
6Dropping rows with missing values instead of imputing, shrinking the required outputtaskda-code
Applies when
task -- the task asks for a per-record output file (labels, predictions, scores) covering a dataset that contains missing values or mixed dtypes.
Pattern
The script handles missing data with a blanket dropna() (or restricts to rows/columns that are fully populated), so the produced file contains far fewer records than the input dataset, and the answer reports the reduced count as if it were the full result.
Detection procedure
  1. Read the task for the required output granularity — does it implicitly require one row per input record, and does it state any filtering? If no filtering is stated, every input record must appear.
  2. In the script, look for dropna/subsetting before the model fit and for whether the output is written from the reduced frame.
  3. Compare the row count claimed in the answer (or in the written file) against the raw dataset's row count; also check the number of feature columns against the number of usable numeric columns.
  4. Flag if rows were silently lost and no imputation (mean/median/etc.) or justification was applied.
Discriminator
A real violation is unrequested row loss that changes the output's coverage; it is fine if the task explicitly asks to filter/subset, or if only a couple of records are dropped for a documented reason and the task does not require complete coverage.
Consequence
The output file has the wrong shape/row count and cluster labels that cannot be aligned to the expected per-record results, so the file comparison fails outright.
id e8f839fe2e4e · mined from da-code dacode-ml-cluster-009
raw text (what the judge reads)
### Dropping rows with missing values instead of imputing, shrinking the required output
- **Applies when**: `task` -- the task asks for a per-record output file (labels, predictions, scores) covering a dataset that contains missing values or mixed dtypes.
- **Pattern**: The script handles missing data with a blanket `dropna()` (or restricts to rows/columns that are fully populated), so the produced file contains far fewer records than the input dataset, and the answer reports the reduced count as if it were the full result.
- **Detection procedure**:
  1. Read the task for the required output granularity — does it implicitly require one row per input record, and does it state any filtering? If no filtering is stated, every input record must appear.
  2. In the script, look for `dropna`/subsetting before the model fit and for whether the output is written from the reduced frame.
  3. Compare the row count claimed in the answer (or in the written file) against the raw dataset's row count; also check the number of feature columns against the number of usable numeric columns.
  4. Flag if rows were silently lost and no imputation (mean/median/etc.) or justification was applied.
- **Discriminator**: A real violation is unrequested row loss that changes the output's coverage; it is fine if the task explicitly asks to filter/subset, or if only a couple of records are dropped for a documented reason **and** the task does not require complete coverage.
- **Consequence**: The output file has the wrong shape/row count and cluster labels that cannot be aligned to the expected per-record results, so the file comparison fails outright.
7Output schema deviation: extra/renamed columns in the required result filetaskda-code
Applies when
task -- the task specifies an exact output file with an exact set/naming of columns, and the script builds that file from an intermediate DataFrame.
Pattern
The script writes the file but adds identifier/bookkeeping columns, keeps original names, mis-indexes the required suffix numbering, or otherwise emits a superset/variant of the requested schema instead of exactly the stated columns (and does not assert the schema before saving).
Detection procedure
  1. From the task statement, write down the exact required column list (names, naming convention, and whether anything else is allowed) and any index/ordering requirement.
  2. In the script, trace the DataFrame that is passed to the write call: list every column added (insert, copy of source columns, reset_index) and every rename mapping, plus the index= argument.
  3. Compare that final column list to the required list; also check the numbering convention starts/increments as the task implies and that no ID/label leftovers survive.
  4. Check the answer/verification output: does it print the saved file's columns.tolist() and compare to the spec, or merely print a head and declare success? Also check the answer's described pipeline matches the script actually run.
Discriminator
A real violation is a saved file whose column set differs from the specification (extra column, unrenamed column, off-by-one naming, index written as a column). A look-alike that is fine: extra columns exist only in in-memory/intermediate frames or in separate diagnostic files, while the required file contains exactly the specified columns.
Consequence
The grader reads the result file, fails the schema/column check (or mis-aligns the feature vector), and marks the expected file WRONG/MISSING regardless of the clustering quality.
id de25d1ca3a10 · mined from da-code dacode-ml-cluster-016
raw text (what the judge reads)
### Output schema deviation: extra/renamed columns in the required result file
- **Applies when**: `task` -- the task specifies an exact output file with an exact set/naming of columns, and the script builds that file from an intermediate DataFrame.
- **Pattern**: The script writes the file but adds identifier/bookkeeping columns, keeps original names, mis-indexes the required suffix numbering, or otherwise emits a superset/variant of the requested schema instead of exactly the stated columns (and does not assert the schema before saving).
- **Detection procedure**:
  1. From the task statement, write down the exact required column list (names, naming convention, and whether anything else is allowed) and any index/ordering requirement.
  2. In the script, trace the DataFrame that is passed to the write call: list every column added (`insert`, `copy` of source columns, `reset_index`) and every rename mapping, plus the `index=` argument.
  3. Compare that final column list to the required list; also check the numbering convention starts/increments as the task implies and that no ID/label leftovers survive.
  4. Check the answer/verification output: does it print the saved file's `columns.tolist()` and compare to the spec, or merely print a head and declare success? Also check the answer's described pipeline matches the script actually run.
- **Discriminator**: A real violation is a saved file whose column set differs from the specification (extra column, unrenamed column, off-by-one naming, index written as a column). A look-alike that is fine: extra columns exist only in in-memory/intermediate frames or in separate diagnostic files, while the required file contains exactly the specified columns.
- **Consequence**: The grader reads the result file, fails the schema/column check (or mis-aligns the feature vector), and marks the expected file WRONG/MISSING regardless of the clustering quality.
8Validation split and feature set that don't mirror the actual prediction settingtaskda-code
Applies when
task -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation.
Pattern
The attempt builds features by blanket-excluding columns (dropping some that exist in both train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predictions against a simple baseline or against the target's own distribution on a held-out block that resembles the test rows.
Detection procedure
  1. From the task and test file schema, list the columns actually available at prediction time; then read the scripts' feature-selection code and note any available column that is excluded or silently dropped (e.g., via an exclude list, select_dtypes, or all-NaN filtering) despite being predictive of the target.
  2. Inspect the holdout logic: does it use train_test_split (shuffled) / random CV on rows that are sequential in time, rather than a chronological or block split matching the test period?
  3. Check whether the answer's reported metric is the only evidence of correctness — i.e., no baseline comparison (persistence/mean/available forecast column) and no sanity check that predicted distribution (mean, range, count, ordering) matches the training target and the expected output rows.
  4. Flag if (1) or (2) holds and (3) holds.
Discriminator
A real violation excludes usable, task-legitimate predictors and/or validates in a way that leaks temporally adjacent rows, so the quoted metric cannot be trusted; a look-alike that is fine excludes only columns genuinely absent/unusable in the test file (true leakage or all-missing), and validates on a chronological holdout that reproduces the test-time information set, with a baseline comparison reported.
Consequence
Validation metrics look strong (low MAE, high R²) while the submitted predictions are systematically off on the real test rows, so the graded file fails the accuracy/tolerance check despite the correct file name and row count.
id fa642c18c0f2 · mined from da-code dacode-ml-regression-002
raw text (what the judge reads)
### Validation split and feature set that don't mirror the actual prediction setting
- **Applies when**: `task` -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation.
- **Pattern**: The attempt builds features by blanket-excluding columns (dropping some that exist in *both* train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predictions against a simple baseline or against the target's own distribution on a held-out block that resembles the test rows.
- **Detection procedure**:
  1. From the task and test file schema, list the columns actually available at prediction time; then read the scripts' feature-selection code and note any available column that is excluded or silently dropped (e.g., via an `exclude` list, `select_dtypes`, or all-NaN filtering) despite being predictive of the target.
  2. Inspect the holdout logic: does it use `train_test_split` (shuffled) / random CV on rows that are sequential in time, rather than a chronological or block split matching the test period?
  3. Check whether the answer's reported metric is the only evidence of correctness — i.e., no baseline comparison (persistence/mean/available forecast column) and no sanity check that predicted distribution (mean, range, count, ordering) matches the training target and the expected output rows.
  4. Flag if (1) or (2) holds and (3) holds.
- **Discriminator**: A real violation excludes usable, task-legitimate predictors and/or validates in a way that leaks temporally adjacent rows, so the quoted metric cannot be trusted; a look-alike that is fine excludes only columns genuinely absent/unusable in the test file (true leakage or all-missing), and validates on a chronological holdout that reproduces the test-time information set, with a baseline comparison reported.
- **Consequence**: Validation metrics look strong (low MAE, high R²) while the submitted predictions are systematically off on the real test rows, so the graded file fails the accuracy/tolerance check despite the correct file name and row count.
9Incomplete deliverables for a spec-driven plotting tasktaskda-code
Applies when
task -- the task points to an external format spec (e.g. a YAML/JSON config) and grading depends on machine-readable artifacts (plot data/series arrays) in addition to a rendered image.
Pattern
The attempt produces only the image and narrates that it "follows the spec", without saving the accompanying data/spec artifacts the grader expects, and without keeping reproducible scripts; internal inconsistencies (e.g. a title naming one date range while the described data covers another) are left unresolved.
Detection procedure
  1. Read the task and any referenced spec file; enumerate every required output artifact (image, serialized plot spec, numeric array of plotted values) and every named property (figsize, color, title, axis labels, ticks, ordering, aggregation level).
  2. Check the working directory / scripts for each enumerated artifact actually being written, and check that a script exists that reproduces them from the raw data.
  3. Cross-check the answer's claimed properties against the spec verbatim (string equality of title/labels, tick list, series length) and against the described data range for contradictions.
  4. Verify the plotted series is derived with the stated granularity and ordering (e.g. correct time aggregation, sorted x-values, no dropped or duplicated periods) rather than asserted.
Discriminator
A real violation is missing required artifacts or a property that mismatches the spec text (including inconsistent date ranges/labels); it is fine if all required files are written and the only differences are cosmetic extras (markers, gridlines, legend) not constrained by the spec.
Consequence
Grader checks on the expected data artifacts report WRONG/MISSING and the attempt scores 0 even though an image exists.
id cd5c52df7363 · mined from da-code dacode-plot-line-015
raw text (what the judge reads)
### Incomplete deliverables for a spec-driven plotting task
- **Applies when**: `task` -- the task points to an external format spec (e.g. a YAML/JSON config) and grading depends on machine-readable artifacts (plot data/series arrays) in addition to a rendered image.
- **Pattern**: The attempt produces only the image and narrates that it "follows the spec", without saving the accompanying data/spec artifacts the grader expects, and without keeping reproducible scripts; internal inconsistencies (e.g. a title naming one date range while the described data covers another) are left unresolved.
- **Detection procedure**:
  1. Read the task and any referenced spec file; enumerate every required output artifact (image, serialized plot spec, numeric array of plotted values) and every named property (figsize, color, title, axis labels, ticks, ordering, aggregation level).
  2. Check the working directory / scripts for each enumerated artifact actually being written, and check that a script exists that reproduces them from the raw data.
  3. Cross-check the answer's claimed properties against the spec verbatim (string equality of title/labels, tick list, series length) and against the described data range for contradictions.
  4. Verify the plotted series is derived with the stated granularity and ordering (e.g. correct time aggregation, sorted x-values, no dropped or duplicated periods) rather than asserted.
- **Discriminator**: A real violation is missing required artifacts or a property that mismatches the spec text (including inconsistent date ranges/labels); it is fine if all required files are written and the only differences are cosmetic extras (markers, gridlines, legend) not constrained by the spec.
- **Consequence**: Grader checks on the expected data artifacts report WRONG/MISSING and the attempt scores 0 even though an image exists.