mle-rubrics-onpolicy-1000

MLE-bench · 1,000 rubrics · first 1000 of all
HF EdwardoSunny/mle-rubrics-onpolicy-1000 · local data/libraries/mle-rubrics-onpolicy-1000.json

0No held-out validation against the competition's stated metric before submittingtaskda-code
Applies when
task -- the task specifies an explicit evaluation metric (e.g., log loss, AUC, RMSE) for predictions on an unlabeled test set, and the script trains models and writes a submission file.
Pattern
The script fits one or more models on 100% of the labeled data, blends them with hand-picked weights, and writes predictions straight to disk — never computing the stated metric on a validation split, cross-validation folds, or even against a trivial baseline (e.g., class-prior probabilities). Hyperparameters, imputation choices, and ensemble weights are therefore unjustified, and there is no evidence the output is better than random or than a constant prediction.
Detection procedure
  1. Read the task and note the exact scoring metric and its inputs (probabilities vs. labels, per-row normalization, clipping).
  2. Scan the script for any train/validation split, cross_val_score/KFold, or explicit computation of that metric on labeled data; check whether printed diagnostics include a metric value at all.
  3. Check whether model/ensemble choices (weights, depth, learning rate) are tied to any measured score, or are literal constants written by the author.
  4. Inspect the answer/submission: is any score, baseline comparison, or sanity check (row count equal to test rows, ids matching test ids, probability ranges/sums, no degenerate constant column) reported?
Discriminator
A real violation is when no estimate of the stated metric exists anywhere for any candidate model, so the attempt cannot distinguish a good submission from a broken one. It is not a violation if the script measures the metric via CV/holdout (even briefly) and uses it to pick among options, nor if the metric is unmeasurable because labels genuinely do not exist for any subset.
Consequence
The submission may be systematically miscalibrated, mis-ordered, mislabeled by class column, or simply far worse than a simple baseline; the grader reports a failing/incorrect submission with no diagnostic trail, since the agent had no internal score to catch it.
id 83f41d76ebb5 · mined from da-code dacode-ml-competition-005
raw text (what the judge reads)
### No held-out validation against the competition's stated metric before submitting
- **Applies when**: `task` -- the task specifies an explicit evaluation metric (e.g., log loss, AUC, RMSE) for predictions on an unlabeled test set, and the script trains models and writes a submission file.
- **Pattern**: The script fits one or more models on 100% of the labeled data, blends them with hand-picked weights, and writes predictions straight to disk — never computing the stated metric on a validation split, cross-validation folds, or even against a trivial baseline (e.g., class-prior probabilities). Hyperparameters, imputation choices, and ensemble weights are therefore unjustified, and there is no evidence the output is better than random or than a constant prediction.
- **Detection procedure**:
  1. Read the task and note the exact scoring metric and its inputs (probabilities vs. labels, per-row normalization, clipping).
  2. Scan the script for any train/validation split, `cross_val_score`/`KFold`, or explicit computation of that metric on labeled data; check whether printed diagnostics include a metric value at all.
  3. Check whether model/ensemble choices (weights, depth, learning rate) are tied to any measured score, or are literal constants written by the author.
  4. Inspect the answer/submission: is any score, baseline comparison, or sanity check (row count equal to test rows, ids matching test ids, probability ranges/sums, no degenerate constant column) reported?
- **Discriminator**: A real violation is when *no* estimate of the stated metric exists anywhere for any candidate model, so the attempt cannot distinguish a good submission from a broken one. It is *not* a violation if the script measures the metric via CV/holdout (even briefly) and uses it to pick among options, nor if the metric is unmeasurable because labels genuinely do not exist for any subset.
- **Consequence**: The submission may be systematically miscalibrated, mis-ordered, mislabeled by class column, or simply far worse than a simple baseline; the grader reports a failing/incorrect submission with no diagnostic trail, since the agent had no internal score to catch it.
1Ships model predictions with no held-out validation and no distribution sanity checktaskda-code
Applies when
task -- a script trains a regressor/classifier on one file and writes predictions for another file as the deliverable, with no ground truth available for the predicted rows.
Pattern
The attempt fits a single model on 100% of the training rows, immediately predicts on the target rows, and writes the output without (a) any hold-out/CV error estimate, (b) a comparison of the predicted value distribution against the training target distribution, or (c) a check that the feature columns/dtypes/encodings used at prediction time are the same as those learned at fit time. Degenerate output (e.g. the vast majority of predictions collapsed to the same value, or a range far narrower than the target's) is accepted as-is.
Detection procedure
  1. Read the task: identify the deliverable (file, required column names/order, row alignment, rounding/units) and note that no labels exist for the predicted rows.
  2. Read the script: check whether a train/validation split or cross-validation with a printed error metric exists, and whether the same column names, imputation, and category encodings are applied consistently to both datasets (watch for near-identical but non-identical column spellings, or one file's schema being assumed for the other).
  3. Read the script's output stage: check whether the predicted values' summary statistics are compared to the training target's summary statistics, and whether any explicit guard rejects a degenerate/off-scale prediction vector.
  4. Inspect the submitted answer: compute the share of identical values and the min/max; if predictions are dominated by a single value or their spread is an order of magnitude off the target's documented spread, and no validation metric was reported, flag it.
Discriminator
A real violation is an unvalidated pipeline whose output is visibly degenerate or whose feature handling silently differs between fit and predict; it is not a violation if the script reports a hold-out/CV score and the prediction distribution plausibly matches the training target (a skewed target legitimately yields many small values, provided the reported validation error supports it).
Consequence
The written file passes format checks but its values are near-constant/mis-scaled, so the grader's accuracy or error tolerance against the reference targets fails, marking the deliverable WRONG.
id 4e7f8b275c9f · mined from da-code dacode-ml-regression-008
raw text (what the judge reads)
### Ships model predictions with no held-out validation and no distribution sanity check
- **Applies when**: `task` -- a script trains a regressor/classifier on one file and writes predictions for another file as the deliverable, with no ground truth available for the predicted rows.
- **Pattern**: The attempt fits a single model on 100% of the training rows, immediately predicts on the target rows, and writes the output without (a) any hold-out/CV error estimate, (b) a comparison of the predicted value distribution against the training target distribution, or (c) a check that the feature columns/dtypes/encodings used at prediction time are the same as those learned at fit time. Degenerate output (e.g. the vast majority of predictions collapsed to the same value, or a range far narrower than the target's) is accepted as-is.
- **Detection procedure**:
  1. Read the task: identify the deliverable (file, required column names/order, row alignment, rounding/units) and note that no labels exist for the predicted rows.
  2. Read the script: check whether a train/validation split or cross-validation with a printed error metric exists, and whether the same column names, imputation, and category encodings are applied consistently to both datasets (watch for near-identical but non-identical column spellings, or one file's schema being assumed for the other).
  3. Read the script's output stage: check whether the predicted values' summary statistics are compared to the training target's summary statistics, and whether any explicit guard rejects a degenerate/off-scale prediction vector.
  4. Inspect the submitted answer: compute the share of identical values and the min/max; if predictions are dominated by a single value or their spread is an order of magnitude off the target's documented spread, and no validation metric was reported, flag it.
- **Discriminator**: A real violation is an unvalidated pipeline whose output is visibly degenerate or whose feature handling silently differs between fit and predict; it is *not* a violation if the script reports a hold-out/CV score and the prediction distribution plausibly matches the training target (a skewed target legitimately yields many small values, provided the reported validation error supports it).
- **Consequence**: The written file passes format checks but its values are near-constant/mis-scaled, so the grader's accuracy or error tolerance against the reference targets fails, marking the deliverable WRONG.
2Hypothesis test run on the full table instead of the task-specified population/subsettaskda-code
Applies when
task -- the task asks for a statistical test, metric, or comparison whose population is qualified by filters (date range, category/tournament type, subgroup, one-sided direction) stated in the prompt, README, or problem context.
Pattern
The attempt loads the raw files and computes the statistic on every row of both groups, silently dropping the qualifying conditions (and/or defaulting to a two-sided test when a directional hypothesis was specified), so the reported p-value/metric describes a different population than the one asked about.
Detection procedure
  1. Read the task and any README, and list every explicit or implied restriction on the rows/columns to be analyzed (time window, subset of categories, groups compared) plus the alternative hypothesis direction and significance level.
  2. Read the scripts and check that each restriction appears as a concrete filter/dtype conversion (e.g., date parsing then range filter, category equality filter) before the statistic is computed, and that the test call matches the stated alternative and test type (paired vs independent, equal-variance assumption).
  3. Compare the row counts used in the test against the raw file row counts; if the script never prints or reduces counts, treat the population as unverified.
  4. Check the reported statistic's magnitude for plausibility given the intended (usually much smaller) subset — extreme p-values (e.g., 1e-100 or smaller) usually signal a far larger n than intended.
Discriminator
A real violation is a missing or incorrect filter/direction that changes which rows enter the computation; a look-alike that is fine is a script that applies the filters in a different but equivalent way (e.g., filtering at load time, using a query string) and can show the reduced counts/subset consistent with the task description.
Consequence
The reported p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample, so it will not match the expected value in the output file and the check fails even though the file format is correct.
id e6246376b83e · mined from da-code dacode-data-sa-001
raw text (what the judge reads)
### Hypothesis test run on the full table instead of the task-specified population/subset
- **Applies when**: `task` -- the task asks for a statistical test, metric, or comparison whose population is qualified by filters (date range, category/tournament type, subgroup, one-sided direction) stated in the prompt, README, or problem context.
- **Pattern**: The attempt loads the raw files and computes the statistic on every row of both groups, silently dropping the qualifying conditions (and/or defaulting to a two-sided test when a directional hypothesis was specified), so the reported p-value/metric describes a different population than the one asked about.
- **Detection procedure**:
  1. Read the task and any README, and list every explicit or implied restriction on the rows/columns to be analyzed (time window, subset of categories, groups compared) plus the alternative hypothesis direction and significance level.
  2. Read the scripts and check that each restriction appears as a concrete filter/dtype conversion (e.g., date parsing then range filter, category equality filter) before the statistic is computed, and that the test call matches the stated alternative and test type (paired vs independent, equal-variance assumption).
  3. Compare the row counts used in the test against the raw file row counts; if the script never prints or reduces counts, treat the population as unverified.
  4. Check the reported statistic's magnitude for plausibility given the intended (usually much smaller) subset — extreme p-values (e.g., 1e-100 or smaller) usually signal a far larger n than intended.
- **Discriminator**: A real violation is a missing or incorrect filter/direction that changes which rows enter the computation; a look-alike that is fine is a script that applies the filters in a different but equivalent way (e.g., filtering at load time, using a query string) and can show the reduced counts/subset consistent with the task description.
- **Consequence**: The reported p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample, so it will not match the expected value in the output file and the check fails even though the file format is correct.
3Output template file never actually inspectedtaskda-code
Applies when
task -- the task says to write results into an output file "following the exact structure/format of" a provided sample/template file, and the scripts build that output.
Pattern
The attempt never loads or prints the sample/template file; it hard-codes guessed column names, column order, row ordering, rounding/units and header text from the prose of the task, then asserts in the answer that the output "matches the required format."
Detection procedure
  1. In the task, note that a reference/sample output file is supplied and is the authority on schema and formatting.
  2. Search the scripts for any read/print/comparison of that sample file (e.g., loading it, checking its columns, dtypes, row count, decimal places, sort order).
  3. If absent, check whether the emitted frame's column names, column order, sort order and numeric formatting are instead invented in code or copied from the task wording.
  4. Check the answer for unverified claims of format compliance (no printed diff against the template).
Discriminator
A real violation is when no code path ever reads the template, so agreement with it is pure luck; it is fine if the script reads the template (or explicitly reindexes/renames/rounds/sorts to the template's columns and formatting) and prints a shape/column/dtype comparison, even if the final naming happens to be hard-coded afterwards.
Consequence
The graded file is judged WRONG/MISSING on a strict file comparison — mismatched header names or order, wrong row ordering, or unrounded/differently scaled values — even when the underlying aggregation logic is right.
id 375545aa1e68 · mined from da-code dacode-dm-csv-011
raw text (what the judge reads)
### Output template file never actually inspected
- **Applies when**: `task` -- the task says to write results into an output file "following the exact structure/format of" a provided sample/template file, and the scripts build that output.
- **Pattern**: The attempt never loads or prints the sample/template file; it hard-codes guessed column names, column order, row ordering, rounding/units and header text from the prose of the task, then asserts in the answer that the output "matches the required format."
- **Detection procedure**:
  1. In the task, note that a reference/sample output file is supplied and is the authority on schema and formatting.
  2. Search the scripts for any read/print/comparison of that sample file (e.g., loading it, checking its columns, dtypes, row count, decimal places, sort order).
  3. If absent, check whether the emitted frame's column names, column order, sort order and numeric formatting are instead invented in code or copied from the task wording.
  4. Check the answer for unverified claims of format compliance (no printed diff against the template).
- **Discriminator**: A real violation is when no code path ever reads the template, so agreement with it is pure luck; it is fine if the script reads the template (or explicitly reindexes/renames/rounds/sorts to the template's columns and formatting) and prints a shape/column/dtype comparison, even if the final naming happens to be hard-coded afterwards.
- **Consequence**: The graded file is judged WRONG/MISSING on a strict file comparison — mismatched header names or order, wrong row ordering, or unrounded/differently scaled values — even when the underlying aggregation logic is right.
4Unverified output artifact: schema/content of the saved file never checked against the requested spectaskda-code
Applies when
task -- the task requires writing results to a named file with an explicitly specified column naming scheme and one row per input record, and the scripts generate that file programmatically.
Pattern
The attempt builds a feature matrix after silently dropping/imputing/encoding columns, fits a model, writes the file, and then reports a prose summary (counts, chosen hyperparameter, feature legend) without ever re-reading the written file to confirm it exists at the expected path and that its header names, column count, row count, and label values match the requested format; degenerate or minimal-complexity results (e.g., the smallest possible number of groups) are accepted without a sanity check.
Detection procedure
  1. From the task, list the required artifact path, the exact column-naming convention, the expected number of rows (records) and the expected value semantics of the label column.
  2. In the scripts, locate the write call and check whether the DataFrame passed in has columns generated to match the convention exactly (no index column, no leftover original names, no extra/missing feature columns relative to the matrix actually clustered) and whether every input record survives preprocessing (no silent row drops from NaN handling or filtering).
  3. Check for a read-back/assert step after writing: does any code load the file and print/verify shape, header, row count, and label distribution? Compare that to what the final answer claims.
  4. Inspect the reported result for degeneracy or implausibility (a single dominant group, the minimum possible number of clusters, row count ≠ dataset size) that no validation step addressed.
Discriminator
A real violation is when the answer's claims about the file are asserted from in-memory variables or narrative only, with no post-write verification and no shape/format assertion — or when preprocessing changed the row/column set without reconciling it to the spec. It is not a violation if the script (or a follow-up run) reloads the artifact and asserts the header pattern, row count equal to the number of input records, and valid label values, even if the modeling choices are debatable.
Consequence
The grader reads the artifact and finds it missing, misnamed, mis-headered, or with the wrong number of rows/columns (or a degenerate labeling), scoring the file check WRONG/MISSING despite a confident-sounding summary.
id 2fad5c23094f · mined from da-code dacode-ml-cluster-014
raw text (what the judge reads)
### Unverified output artifact: schema/content of the saved file never checked against the requested spec
- **Applies when**: `task` -- the task requires writing results to a named file with an explicitly specified column naming scheme and one row per input record, and the scripts generate that file programmatically.
- **Pattern**: The attempt builds a feature matrix after silently dropping/imputing/encoding columns, fits a model, writes the file, and then reports a prose summary (counts, chosen hyperparameter, feature legend) without ever re-reading the written file to confirm it exists at the expected path and that its header names, column count, row count, and label values match the requested format; degenerate or minimal-complexity results (e.g., the smallest possible number of groups) are accepted without a sanity check.
- **Detection procedure**:
  1. From the task, list the required artifact path, the exact column-naming convention, the expected number of rows (records) and the expected value semantics of the label column.
  2. In the scripts, locate the write call and check whether the DataFrame passed in has columns generated to match the convention exactly (no index column, no leftover original names, no extra/missing feature columns relative to the matrix actually clustered) and whether every input record survives preprocessing (no silent row drops from NaN handling or filtering).
  3. Check for a read-back/assert step after writing: does any code load the file and print/verify shape, header, row count, and label distribution? Compare that to what the final answer claims.
  4. Inspect the reported result for degeneracy or implausibility (a single dominant group, the minimum possible number of clusters, row count ≠ dataset size) that no validation step addressed.
- **Discriminator**: A real violation is when the answer's claims about the file are asserted from in-memory variables or narrative only, with no post-write verification and no shape/format assertion — or when preprocessing changed the row/column set without reconciling it to the spec. It is *not* a violation if the script (or a follow-up run) reloads the artifact and asserts the header pattern, row count equal to the number of input records, and valid label values, even if the modeling choices are debatable.
- **Consequence**: The grader reads the artifact and finds it missing, misnamed, mis-headered, or with the wrong number of rows/columns (or a degenerate labeling), scoring the file check WRONG/MISSING despite a confident-sounding summary.
5Submission artifact never validated against the provided template (path, columns, ids, dtype)taskda-code
Applies when
task -- the task requires writing a prediction/result file whose format is defined by a sample/template file, and the scripts generate that file at a hard-coded location.
Pattern
The attempt builds the output frame from its own assumptions (its own column names, its own output directory, its own value dtype such as forcibly rounding/casting continuous predictions), prints a few rows as "verification", and never programmatically loads the template to confirm the file is written where the task expects, with the same column names/order, the same number and set of identifiers, and value types consistent with the evaluation metric.
Detection procedure
  1. In the task/README, note the required output filename, its expected location relative to the working directory, and the template file that defines its schema.
  2. In the scripts, find every write of the output file: check the path string, the constructed column names/order, and any post-processing of predictions (rounding, clipping, int casting, sorting).
  3. Check whether any script actually reads the template and asserts equality of columns, row count, and id set/order against the written file — printing head() or shapes without comparison does not count.
  4. In the answer, check whether it states the verified output location and schema match, or only narrates modeling choices and CV scores.
Discriminator
A real violation is when no code compares the produced file to the template/required path, or when values are transformed in a way the metric does not ask for (e.g., integer rounding of a continuous/log-scale target). It is fine if the script asserts column equality, id alignment and row counts against the template (even implicitly by copying the template and overwriting the prediction column) and writes to the location the task specifies.
Consequence
The grader reports the expected result file as missing or wrong (not found at the expected path, mismatched columns/ids, or degraded score from unnecessary value transformation), so the attempt scores 0 despite a plausible-looking model and CV metrics.
id 5c50253d9364 · mined from da-code dacode-ml-competition-009
raw text (what the judge reads)
### Submission artifact never validated against the provided template (path, columns, ids, dtype)
- **Applies when**: `task` -- the task requires writing a prediction/result file whose format is defined by a sample/template file, and the scripts generate that file at a hard-coded location.
- **Pattern**: The attempt builds the output frame from its own assumptions (its own column names, its own output directory, its own value dtype such as forcibly rounding/casting continuous predictions), prints a few rows as "verification", and never programmatically loads the template to confirm the file is written where the task expects, with the same column names/order, the same number and set of identifiers, and value types consistent with the evaluation metric.
- **Detection procedure**:
  1. In the task/README, note the required output filename, its expected location relative to the working directory, and the template file that defines its schema.
  2. In the scripts, find every write of the output file: check the path string, the constructed column names/order, and any post-processing of predictions (rounding, clipping, int casting, sorting).
  3. Check whether any script actually reads the template and asserts equality of columns, row count, and id set/order against the written file — printing `head()` or shapes without comparison does not count.
  4. In the answer, check whether it states the verified output location and schema match, or only narrates modeling choices and CV scores.
- **Discriminator**: A real violation is when no code compares the produced file to the template/required path, or when values are transformed in a way the metric does not ask for (e.g., integer rounding of a continuous/log-scale target). It is fine if the script asserts column equality, id alignment and row counts against the template (even implicitly by copying the template and overwriting the prediction column) and writes to the location the task specifies.
- **Consequence**: The grader reports the expected result file as missing or wrong (not found at the expected path, mismatched columns/ids, or degraded score from unnecessary value transformation), so the attempt scores 0 despite a plausible-looking model and CV metrics.
6Dropping rows with missing values instead of imputing, shrinking the required outputtaskda-code
Applies when
task -- the task asks for a per-record output file (labels, predictions, scores) covering a dataset that contains missing values or mixed dtypes.
Pattern
The script handles missing data with a blanket dropna() (or restricts to rows/columns that are fully populated), so the produced file contains far fewer records than the input dataset, and the answer reports the reduced count as if it were the full result.
Detection procedure
  1. Read the task for the required output granularity — does it implicitly require one row per input record, and does it state any filtering? If no filtering is stated, every input record must appear.
  2. In the script, look for dropna/subsetting before the model fit and for whether the output is written from the reduced frame.
  3. Compare the row count claimed in the answer (or in the written file) against the raw dataset's row count; also check the number of feature columns against the number of usable numeric columns.
  4. Flag if rows were silently lost and no imputation (mean/median/etc.) or justification was applied.
Discriminator
A real violation is unrequested row loss that changes the output's coverage; it is fine if the task explicitly asks to filter/subset, or if only a couple of records are dropped for a documented reason and the task does not require complete coverage.
Consequence
The output file has the wrong shape/row count and cluster labels that cannot be aligned to the expected per-record results, so the file comparison fails outright.
id e8f839fe2e4e · mined from da-code dacode-ml-cluster-009
raw text (what the judge reads)
### Dropping rows with missing values instead of imputing, shrinking the required output
- **Applies when**: `task` -- the task asks for a per-record output file (labels, predictions, scores) covering a dataset that contains missing values or mixed dtypes.
- **Pattern**: The script handles missing data with a blanket `dropna()` (or restricts to rows/columns that are fully populated), so the produced file contains far fewer records than the input dataset, and the answer reports the reduced count as if it were the full result.
- **Detection procedure**:
  1. Read the task for the required output granularity — does it implicitly require one row per input record, and does it state any filtering? If no filtering is stated, every input record must appear.
  2. In the script, look for `dropna`/subsetting before the model fit and for whether the output is written from the reduced frame.
  3. Compare the row count claimed in the answer (or in the written file) against the raw dataset's row count; also check the number of feature columns against the number of usable numeric columns.
  4. Flag if rows were silently lost and no imputation (mean/median/etc.) or justification was applied.
- **Discriminator**: A real violation is unrequested row loss that changes the output's coverage; it is fine if the task explicitly asks to filter/subset, or if only a couple of records are dropped for a documented reason **and** the task does not require complete coverage.
- **Consequence**: The output file has the wrong shape/row count and cluster labels that cannot be aligned to the expected per-record results, so the file comparison fails outright.
7Output schema deviation: extra/renamed columns in the required result filetaskda-code
Applies when
task -- the task specifies an exact output file with an exact set/naming of columns, and the script builds that file from an intermediate DataFrame.
Pattern
The script writes the file but adds identifier/bookkeeping columns, keeps original names, mis-indexes the required suffix numbering, or otherwise emits a superset/variant of the requested schema instead of exactly the stated columns (and does not assert the schema before saving).
Detection procedure
  1. From the task statement, write down the exact required column list (names, naming convention, and whether anything else is allowed) and any index/ordering requirement.
  2. In the script, trace the DataFrame that is passed to the write call: list every column added (insert, copy of source columns, reset_index) and every rename mapping, plus the index= argument.
  3. Compare that final column list to the required list; also check the numbering convention starts/increments as the task implies and that no ID/label leftovers survive.
  4. Check the answer/verification output: does it print the saved file's columns.tolist() and compare to the spec, or merely print a head and declare success? Also check the answer's described pipeline matches the script actually run.
Discriminator
A real violation is a saved file whose column set differs from the specification (extra column, unrenamed column, off-by-one naming, index written as a column). A look-alike that is fine: extra columns exist only in in-memory/intermediate frames or in separate diagnostic files, while the required file contains exactly the specified columns.
Consequence
The grader reads the result file, fails the schema/column check (or mis-aligns the feature vector), and marks the expected file WRONG/MISSING regardless of the clustering quality.
id de25d1ca3a10 · mined from da-code dacode-ml-cluster-016
raw text (what the judge reads)
### Output schema deviation: extra/renamed columns in the required result file
- **Applies when**: `task` -- the task specifies an exact output file with an exact set/naming of columns, and the script builds that file from an intermediate DataFrame.
- **Pattern**: The script writes the file but adds identifier/bookkeeping columns, keeps original names, mis-indexes the required suffix numbering, or otherwise emits a superset/variant of the requested schema instead of exactly the stated columns (and does not assert the schema before saving).
- **Detection procedure**:
  1. From the task statement, write down the exact required column list (names, naming convention, and whether anything else is allowed) and any index/ordering requirement.
  2. In the script, trace the DataFrame that is passed to the write call: list every column added (`insert`, `copy` of source columns, `reset_index`) and every rename mapping, plus the `index=` argument.
  3. Compare that final column list to the required list; also check the numbering convention starts/increments as the task implies and that no ID/label leftovers survive.
  4. Check the answer/verification output: does it print the saved file's `columns.tolist()` and compare to the spec, or merely print a head and declare success? Also check the answer's described pipeline matches the script actually run.
- **Discriminator**: A real violation is a saved file whose column set differs from the specification (extra column, unrenamed column, off-by-one naming, index written as a column). A look-alike that is fine: extra columns exist only in in-memory/intermediate frames or in separate diagnostic files, while the required file contains exactly the specified columns.
- **Consequence**: The grader reads the result file, fails the schema/column check (or mis-aligns the feature vector), and marks the expected file WRONG/MISSING regardless of the clustering quality.
8Validation split and feature set that don't mirror the actual prediction settingtaskda-code
Applies when
task -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation.
Pattern
The attempt builds features by blanket-excluding columns (dropping some that exist in both train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predictions against a simple baseline or against the target's own distribution on a held-out block that resembles the test rows.
Detection procedure
  1. From the task and test file schema, list the columns actually available at prediction time; then read the scripts' feature-selection code and note any available column that is excluded or silently dropped (e.g., via an exclude list, select_dtypes, or all-NaN filtering) despite being predictive of the target.
  2. Inspect the holdout logic: does it use train_test_split (shuffled) / random CV on rows that are sequential in time, rather than a chronological or block split matching the test period?
  3. Check whether the answer's reported metric is the only evidence of correctness — i.e., no baseline comparison (persistence/mean/available forecast column) and no sanity check that predicted distribution (mean, range, count, ordering) matches the training target and the expected output rows.
  4. Flag if (1) or (2) holds and (3) holds.
Discriminator
A real violation excludes usable, task-legitimate predictors and/or validates in a way that leaks temporally adjacent rows, so the quoted metric cannot be trusted; a look-alike that is fine excludes only columns genuinely absent/unusable in the test file (true leakage or all-missing), and validates on a chronological holdout that reproduces the test-time information set, with a baseline comparison reported.
Consequence
Validation metrics look strong (low MAE, high R²) while the submitted predictions are systematically off on the real test rows, so the graded file fails the accuracy/tolerance check despite the correct file name and row count.
id fa642c18c0f2 · mined from da-code dacode-ml-regression-002
raw text (what the judge reads)
### Validation split and feature set that don't mirror the actual prediction setting
- **Applies when**: `task` -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation.
- **Pattern**: The attempt builds features by blanket-excluding columns (dropping some that exist in *both* train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predictions against a simple baseline or against the target's own distribution on a held-out block that resembles the test rows.
- **Detection procedure**:
  1. From the task and test file schema, list the columns actually available at prediction time; then read the scripts' feature-selection code and note any available column that is excluded or silently dropped (e.g., via an `exclude` list, `select_dtypes`, or all-NaN filtering) despite being predictive of the target.
  2. Inspect the holdout logic: does it use `train_test_split` (shuffled) / random CV on rows that are sequential in time, rather than a chronological or block split matching the test period?
  3. Check whether the answer's reported metric is the only evidence of correctness — i.e., no baseline comparison (persistence/mean/available forecast column) and no sanity check that predicted distribution (mean, range, count, ordering) matches the training target and the expected output rows.
  4. Flag if (1) or (2) holds and (3) holds.
- **Discriminator**: A real violation excludes usable, task-legitimate predictors and/or validates in a way that leaks temporally adjacent rows, so the quoted metric cannot be trusted; a look-alike that is fine excludes only columns genuinely absent/unusable in the test file (true leakage or all-missing), and validates on a chronological holdout that reproduces the test-time information set, with a baseline comparison reported.
- **Consequence**: Validation metrics look strong (low MAE, high R²) while the submitted predictions are systematically off on the real test rows, so the graded file fails the accuracy/tolerance check despite the correct file name and row count.
9Incomplete deliverables for a spec-driven plotting tasktaskda-code
Applies when
task -- the task points to an external format spec (e.g. a YAML/JSON config) and grading depends on machine-readable artifacts (plot data/series arrays) in addition to a rendered image.
Pattern
The attempt produces only the image and narrates that it "follows the spec", without saving the accompanying data/spec artifacts the grader expects, and without keeping reproducible scripts; internal inconsistencies (e.g. a title naming one date range while the described data covers another) are left unresolved.
Detection procedure
  1. Read the task and any referenced spec file; enumerate every required output artifact (image, serialized plot spec, numeric array of plotted values) and every named property (figsize, color, title, axis labels, ticks, ordering, aggregation level).
  2. Check the working directory / scripts for each enumerated artifact actually being written, and check that a script exists that reproduces them from the raw data.
  3. Cross-check the answer's claimed properties against the spec verbatim (string equality of title/labels, tick list, series length) and against the described data range for contradictions.
  4. Verify the plotted series is derived with the stated granularity and ordering (e.g. correct time aggregation, sorted x-values, no dropped or duplicated periods) rather than asserted.
Discriminator
A real violation is missing required artifacts or a property that mismatches the spec text (including inconsistent date ranges/labels); it is fine if all required files are written and the only differences are cosmetic extras (markers, gridlines, legend) not constrained by the spec.
Consequence
Grader checks on the expected data artifacts report WRONG/MISSING and the attempt scores 0 even though an image exists.
id cd5c52df7363 · mined from da-code dacode-plot-line-015
raw text (what the judge reads)
### Incomplete deliverables for a spec-driven plotting task
- **Applies when**: `task` -- the task points to an external format spec (e.g. a YAML/JSON config) and grading depends on machine-readable artifacts (plot data/series arrays) in addition to a rendered image.
- **Pattern**: The attempt produces only the image and narrates that it "follows the spec", without saving the accompanying data/spec artifacts the grader expects, and without keeping reproducible scripts; internal inconsistencies (e.g. a title naming one date range while the described data covers another) are left unresolved.
- **Detection procedure**:
  1. Read the task and any referenced spec file; enumerate every required output artifact (image, serialized plot spec, numeric array of plotted values) and every named property (figsize, color, title, axis labels, ticks, ordering, aggregation level).
  2. Check the working directory / scripts for each enumerated artifact actually being written, and check that a script exists that reproduces them from the raw data.
  3. Cross-check the answer's claimed properties against the spec verbatim (string equality of title/labels, tick list, series length) and against the described data range for contradictions.
  4. Verify the plotted series is derived with the stated granularity and ordering (e.g. correct time aggregation, sorted x-values, no dropped or duplicated periods) rather than asserted.
- **Discriminator**: A real violation is missing required artifacts or a property that mismatches the spec text (including inconsistent date ranges/labels); it is fine if all required files are written and the only differences are cosmetic extras (markers, gridlines, legend) not constrained by the spec.
- **Consequence**: Grader checks on the expected data artifacts report WRONG/MISSING and the attempt scores 0 even though an image exists.
10Missing required output artifact (and fabricated fallback data instead of the real inputs)taskda-code
Applies when
task -- the task names a specific output file/format (e.g., a result file matching a provided sample) and points to input data that the scripts must locate and load.
Pattern
The scripts implement the statistical/modeling logic but never write the requested file; input loading is done by guessing candidate paths with a silent fallback that synthesizes or hard-codes plausible data, so the reported numbers come from invented inputs and the deliverable file is absent. The final answer is a prose/console report instead of the specified artifact.
Detection procedure
  1. From the task, list every required deliverable: exact output filename, expected columns/rows/ordering, rounding/units, and the exact input file(s) to be used.
  2. Grep the scripts for any write of that deliverable (to_csv, open(..., 'w'), etc.) and check the written schema against the provided sample; also check whether the sample file was read/inspected at all.
  3. Trace the data-loading path: does it read the actual provided file(s), or try a list of guessed paths / except: pass / if df is None: branch that generates random or literal numbers? Confirm the script would fail loudly rather than proceed on fake data.
  4. Compare the reported numbers to the branch that could have produced them (e.g., suspiciously round sample sizes, seeded np.random.normal values, a p-value of exactly 0) to see whether the answer came from the fallback.
Discriminator
A real violation is when no code path writes the named file with the sample's schema, or the reported values could only come from synthetic/hard-coded inputs. It is fine if the script robustly searches for the input but raises on failure, and does write the deliverable with the required columns/precision — even if the analysis logic is only in a helper module or the console also prints a summary.
Consequence
The grader finds the expected result file missing or containing values derived from invented data, so all output checks fail regardless of whether the statistical method was correct.
id c0919b027b92 · mined from da-code dacode-data-sa-028
raw text (what the judge reads)
### Missing required output artifact (and fabricated fallback data instead of the real inputs)
- **Applies when**: `task` -- the task names a specific output file/format (e.g., a result file matching a provided sample) and points to input data that the scripts must locate and load.
- **Pattern**: The scripts implement the statistical/modeling logic but never write the requested file; input loading is done by guessing candidate paths with a silent fallback that synthesizes or hard-codes plausible data, so the reported numbers come from invented inputs and the deliverable file is absent. The final answer is a prose/console report instead of the specified artifact.
- **Detection procedure**:
  1. From the task, list every required deliverable: exact output filename, expected columns/rows/ordering, rounding/units, and the exact input file(s) to be used.
  2. Grep the scripts for any write of that deliverable (`to_csv`, `open(..., 'w')`, etc.) and check the written schema against the provided sample; also check whether the sample file was read/inspected at all.
  3. Trace the data-loading path: does it read the actual provided file(s), or try a list of guessed paths / `except: pass` / `if df is None:` branch that generates random or literal numbers? Confirm the script would fail loudly rather than proceed on fake data.
  4. Compare the reported numbers to the branch that could have produced them (e.g., suspiciously round sample sizes, seeded `np.random.normal` values, a p-value of exactly 0) to see whether the answer came from the fallback.
- **Discriminator**: A real violation is when no code path writes the named file with the sample's schema, or the reported values could only come from synthetic/hard-coded inputs. It is fine if the script robustly searches for the input but raises on failure, and does write the deliverable with the required columns/precision — even if the analysis logic is only in a helper module or the console also prints a summary.
- **Consequence**: The grader finds the expected result file missing or containing values derived from invented data, so all output checks fail regardless of whether the statistical method was correct.
11Ignoring referenced specification files and required output artifactstaskda-code
Applies when
task -- The task points to auxiliary instruction/config files (e.g., a tips/README/spec text file, a YAML/JSON formatting config) and/or expects a set of deliverable files beyond the obvious main output.
Pattern
The scripts never open, parse, or echo the referenced instruction/config files; the agent instead hardcodes its own assumptions (grouping definitions, columns, titles, colors, axis labels, figure size, ranges) and writes only the single most obvious artifact, then declares success by paraphrasing the spec it never actually read.
Detection procedure
  1. List every external file the task text references (instruction files, config/format files) and every output file implied or named.
  2. Grep the scripts for reads of each referenced file (open, read_csv, yaml.safe_load, json.load) and for writes of each expected output; note any file mentioned in the task but absent from the code.
  3. Check whether formatting/derivation choices in the plotting or aggregation code trace back to values loaded from the config, or are literals invented by the agent.
  4. Compare the answer's claims (e.g., "formatted according to the spec", "definitions from the tips file") against step 2 — claims about content of unread files are unsupported.
Discriminator
A real violation is code that contains no read of a task-referenced file, or omits an expected artifact entirely. It is not a violation if the script reads the config and then legitimately falls back to defaults for keys the config doesn't specify, or if the extra artifacts are written by a separate step visible elsewhere in the run.
Consequence
Graders that compare each required artifact (image, serialized plot spec, numeric array) find them missing or mismatched in styling, labels, series definitions, or data values, so all checks fail even if the underlying computation looks plausible.
id 06f5882f6230 · mined from da-code dacode-plot-line-006
raw text (what the judge reads)
### Ignoring referenced specification files and required output artifacts
- **Applies when**: `task` -- The task points to auxiliary instruction/config files (e.g., a tips/README/spec text file, a YAML/JSON formatting config) and/or expects a set of deliverable files beyond the obvious main output.
- **Pattern**: The scripts never open, parse, or echo the referenced instruction/config files; the agent instead hardcodes its own assumptions (grouping definitions, columns, titles, colors, axis labels, figure size, ranges) and writes only the single most obvious artifact, then declares success by paraphrasing the spec it never actually read.
- **Detection procedure**:
  1. List every external file the task text references (instruction files, config/format files) and every output file implied or named.
  2. Grep the scripts for reads of each referenced file (`open`, `read_csv`, `yaml.safe_load`, `json.load`) and for writes of each expected output; note any file mentioned in the task but absent from the code.
  3. Check whether formatting/derivation choices in the plotting or aggregation code trace back to values loaded from the config, or are literals invented by the agent.
  4. Compare the answer's claims (e.g., "formatted according to the spec", "definitions from the tips file") against step 2 — claims about content of unread files are unsupported.
- **Discriminator**: A real violation is code that contains no read of a task-referenced file, or omits an expected artifact entirely. It is *not* a violation if the script reads the config and then legitimately falls back to defaults for keys the config doesn't specify, or if the extra artifacts are written by a separate step visible elsewhere in the run.
- **Consequence**: Graders that compare each required artifact (image, serialized plot spec, numeric array) find them missing or mismatched in styling, labels, series definitions, or data values, so all checks fail even if the underlying computation looks plausible.
12Spec file referenced by the task is never read or implementedtaskda-code
Applies when
task -- the instructions point to an external document (README/spec/config) that defines how to bin, filter, map, or order values before producing the requested output.
Pattern
The agent skips or paraphrases the spec, reuses the raw categories/values already present in the data (or its own invented grouping), and asserts in the answer that "the spec was followed" without any script step that opens the spec file or encodes its rules; often no reproducible script is saved at all.
Detection procedure
  1. From the task, list every referenced auxiliary document and the exact rules it is supposed to supply (bin edges, labels, exclusions, ordering, required output artifacts).
  2. Search the scripts for a read of that document and for an explicit mapping/binning structure (dict, pd.cut bins, list of labels) that materializes those rules; if no script exists, treat the derivation as unverifiable.
  3. Compare the category labels/counts in the answer against the raw distinct values of the source field: if they are identical (or the totals equal the unfiltered row count when the spec implies grouping/filtering), the spec was not applied.
  4. Check that every artifact the task/spec implies (plot plus any data dumps) is actually written by the script, not just described in prose.
Discriminator
A genuine violation is when no code path encodes the spec's rules and the output categories coincide with the raw field's own levels; it is fine if the code hard-codes bins that are demonstrably transcribed from the spec (labels/edges match) even without re-reading the file at runtime.
Consequence
Saved outputs (image and any numeric/JSON dumps) have wrong bin labels, wrong bin counts, or are missing entirely, so every file-level check fails despite a confident "task completed" report.
id 59487d5f27e9 · mined from da-code dacode-plot-bar-005
raw text (what the judge reads)
### Spec file referenced by the task is never read or implemented
- **Applies when**: `task` -- the instructions point to an external document (README/spec/config) that defines how to bin, filter, map, or order values before producing the requested output.
- **Pattern**: The agent skips or paraphrases the spec, reuses the raw categories/values already present in the data (or its own invented grouping), and asserts in the answer that "the spec was followed" without any script step that opens the spec file or encodes its rules; often no reproducible script is saved at all.
- **Detection procedure**:
  1. From the task, list every referenced auxiliary document and the exact rules it is supposed to supply (bin edges, labels, exclusions, ordering, required output artifacts).
  2. Search the scripts for a read of that document and for an explicit mapping/binning structure (dict, `pd.cut` bins, list of labels) that materializes those rules; if no script exists, treat the derivation as unverifiable.
  3. Compare the category labels/counts in the answer against the raw distinct values of the source field: if they are identical (or the totals equal the unfiltered row count when the spec implies grouping/filtering), the spec was not applied.
  4. Check that every artifact the task/spec implies (plot plus any data dumps) is actually written by the script, not just described in prose.
- **Discriminator**: A genuine violation is when no code path encodes the spec's rules and the output categories coincide with the raw field's own levels; it is fine if the code hard-codes bins that are demonstrably transcribed from the spec (labels/edges match) even without re-reading the file at runtime.
- **Consequence**: Saved outputs (image and any numeric/JSON dumps) have wrong bin labels, wrong bin counts, or are missing entirely, so every file-level check fails despite a confident "task completed" report.
13Missing reproducible script and required output artifacttaskda-code
Applies when
task -- the task states a specific preprocessing step and a specific answer format/output file, and the agent supplies a final answer with no (or non-runnable) script that produces it.
Pattern
The agent reports a value it believes it read off the data (or recalls), without a saved, end-to-end script that loads the raw file, coerces the relevant column to numeric (stripping symbols/thousand separators/percent signs), applies the stated imputation, computes the requested extremum/statistic, and writes the result to the exact requested artifact (e.g. the named JSON/CSV) in the exact requested schema. Nothing in the deliverables lets a reviewer re-derive the number, and the required file is never created.
Detection procedure
  1. List every deliverable the task demands: the output file name/location, the key names, the value types, and any mandated preprocessing step.
  2. Look for a script whose code path visibly does each of those: raw load → dtype cleaning of the target column → the mandated imputation → the selection/aggregation → serialization to the named file.
  3. If any link is absent (no script at all, or the script prints to stdout only, or never touches dtypes/imputation), mark inadequate; also check the reported number is a plausible value of the stated unit (e.g. a percentage inside 0–100) and that it belongs to the row that a numeric — not lexicographic — comparison would select.
  4. Confirm the emitted structure matches the requested shape exactly (same keys, list-vs-scalar, rounding).
Discriminator
A real violation is when the answer cannot be traced to executed code that both preprocesses as instructed and writes the named artifact. A look-alike that is fine: a short but complete script that does all steps and dumps the file, even if the agent additionally quotes the answer in prose.
Consequence
The grader finds the expected result file missing or containing a value derived from uncleaned/unimputed data (e.g. a string-max or a pre-imputation row), so the check fails with 0/1 even though the prose answer looks well-formatted.
id 225b4ad6590e · mined from da-code dacode-di-text-002
raw text (what the judge reads)
### Missing reproducible script and required output artifact
- **Applies when**: `task` -- the task states a specific preprocessing step and a specific answer format/output file, and the agent supplies a final answer with no (or non-runnable) script that produces it.
- **Pattern**: The agent reports a value it believes it read off the data (or recalls), without a saved, end-to-end script that loads the raw file, coerces the relevant column to numeric (stripping symbols/thousand separators/percent signs), applies the stated imputation, computes the requested extremum/statistic, and writes the result to the exact requested artifact (e.g. the named JSON/CSV) in the exact requested schema. Nothing in the deliverables lets a reviewer re-derive the number, and the required file is never created.
- **Detection procedure**:
  1. List every deliverable the task demands: the output file name/location, the key names, the value types, and any mandated preprocessing step.
  2. Look for a script whose code path visibly does each of those: raw load → dtype cleaning of the target column → the mandated imputation → the selection/aggregation → serialization to the named file.
  3. If any link is absent (no script at all, or the script prints to stdout only, or never touches dtypes/imputation), mark inadequate; also check the reported number is a plausible value of the stated unit (e.g. a percentage inside 0–100) and that it belongs to the row that a numeric — not lexicographic — comparison would select.
  4. Confirm the emitted structure matches the requested shape exactly (same keys, list-vs-scalar, rounding).
- **Discriminator**: A real violation is when the answer cannot be traced to executed code that both preprocesses as instructed and writes the named artifact. A look-alike that is fine: a short but complete script that does all steps and dumps the file, even if the agent additionally quotes the answer in prose.
- **Consequence**: The grader finds the expected result file missing or containing a value derived from uncleaned/unimputed data (e.g. a string-max or a pre-imputation row), so the check fails with 0/1 even though the prose answer looks well-formatted.
14Named statistic implemented with the wrong formula varianttaskinfiagent-dabench
Applies when
task -- The task names a specific, textbook-named statistic, coefficient, or metric variant (e.g. a particular "first/second coefficient", a specific averaging or normalization convention) that the script must compute by hand.
Pattern
The script hard-codes one plausible variant of the named quantity (or calls a library default) without checking that it matches the exact named definition, often with a comment that asserts the mapping rather than verifying it — e.g. implementing the median-based sibling formula while labeling it as the mode-based one, or using sample vs. population normalization inconsistently with the definition.
Detection procedure
  1. From the task statement, write down the exact canonical definition of the named statistic, including which central-tendency/normalization terms it uses.
  2. Read the script's arithmetic expression (not its comments or variable names) and map each term to the canonical definition.
  3. Check whether an alternative, similarly named variant exists that the script may have substituted; if so, require evidence in the script that both were computed/compared or that the chosen one was justified.
  4. Confirm downstream classification/rounding uses the value from the correct formula.
Discriminator
A real violation is when the coded expression is a different named quantity (different terms, different denominator convention) than the one requested; a look-alike that is fine is an algebraically equivalent rewriting, or a negligible convention choice (e.g. ddof) that provably cannot change the reported rounded value or the qualitative conclusion.
Consequence
The qualitative label (sign/direction/category) may still match by luck, but the numeric field differs from the reference, so the answer fails the value check while passing the type check — partial credit at best.
id a6e90dea273c · mined from infiagent-dabench dabench-359
raw text (what the judge reads)
### Named statistic implemented with the wrong formula variant
- **Applies when**: `task` -- The task names a specific, textbook-named statistic, coefficient, or metric variant (e.g. a particular "first/second coefficient", a specific averaging or normalization convention) that the script must compute by hand.
- **Pattern**: The script hard-codes one plausible variant of the named quantity (or calls a library default) without checking that it matches the exact named definition, often with a comment that asserts the mapping rather than verifying it — e.g. implementing the median-based sibling formula while labeling it as the mode-based one, or using sample vs. population normalization inconsistently with the definition.
- **Detection procedure**:
  1. From the task statement, write down the exact canonical definition of the named statistic, including which central-tendency/normalization terms it uses.
  2. Read the script's arithmetic expression (not its comments or variable names) and map each term to the canonical definition.
  3. Check whether an alternative, similarly named variant exists that the script may have substituted; if so, require evidence in the script that both were computed/compared or that the chosen one was justified.
  4. Confirm downstream classification/rounding uses the value from the correct formula.
- **Discriminator**: A real violation is when the coded expression is a *different* named quantity (different terms, different denominator convention) than the one requested; a look-alike that is fine is an algebraically equivalent rewriting, or a negligible convention choice (e.g. ddof) that provably cannot change the reported rounded value or the qualitative conclusion.
- **Consequence**: The qualitative label (sign/direction/category) may still match by luck, but the numeric field differs from the reference, so the answer fails the value check while passing the type check — partial credit at best.
15Circular "verification" that re-runs the same code instead of validating input assumptionstaskda-code
Applies when
task -- a script aggregates several raw columns row-wise (weighted sums, means, cumulative products) and a second "verification" script is used to confirm the saved output.
Pattern
The attempt reads the raw file with no inspection of missing values, dtypes, units/scale, or a leading initialization row, aggregates with a reducer that silently ignores NaNs (e.g. .sum(axis=1), .mean()), then "verifies" by recomputing with the identical code path and reporting a match — so any shared assumption error (NaNs treated as zero, percent vs. decimal returns, dropped/misaligned column, wrong starting baseline or compounding convention) is confirmed rather than caught.
Detection procedure
  1. In the task/README, list the assumptions the computation depends on: units/scale of the input values, expected number of rows/columns, and the exact definition and starting point of the requested output quantity.
  2. In the scripts, check whether the raw input is ever validated — isna().sum(), dtype check, min/max/scale check, row/column count check, or explicit NaN policy in the aggregation — and whether the aggregation is done with an operator that skips NaNs by default.
  3. Check whether the verification script does anything independent (recompute one row by hand from raw numbers, compare against a known benchmark, check output ranges/row counts/column names against the required format) or merely re-executes the same formula.
  4. Inspect the answer for tell-tale signs: implausibly smooth/small magnitudes, no NaN or zero at the series start, row count differing from the source file, or values that would be off by ~100x under a unit mismatch.
Discriminator
A real violation is when no check exists that could fail independently of the main computation (data-quality assertions absent and the verifier duplicates the formula). It is fine if the script asserts input completeness/scale, handles NaNs explicitly, or cross-checks at least one output value against an externally derived or hand-computed reference — even if the code is otherwise reused.
Consequence
All internal checks "pass" while the saved file diverges from the reference (shifted baseline, NaN-as-zero contamination, wrong scale or row alignment), and the grader marks the output file WRONG on numeric comparison.
id 8499bdc410cc · mined from da-code dacode-dm-csv-050
raw text (what the judge reads)
### Circular "verification" that re-runs the same code instead of validating input assumptions
- **Applies when**: `task` -- a script aggregates several raw columns row-wise (weighted sums, means, cumulative products) and a second "verification" script is used to confirm the saved output.
- **Pattern**: The attempt reads the raw file with no inspection of missing values, dtypes, units/scale, or a leading initialization row, aggregates with a reducer that silently ignores NaNs (e.g. `.sum(axis=1)`, `.mean()`), then "verifies" by recomputing with the identical code path and reporting a match — so any shared assumption error (NaNs treated as zero, percent vs. decimal returns, dropped/misaligned column, wrong starting baseline or compounding convention) is confirmed rather than caught.
- **Detection procedure**:
  1. In the task/README, list the assumptions the computation depends on: units/scale of the input values, expected number of rows/columns, and the exact definition and starting point of the requested output quantity.
  2. In the scripts, check whether the raw input is ever validated — `isna().sum()`, dtype check, min/max/scale check, row/column count check, or explicit NaN policy in the aggregation — and whether the aggregation is done with an operator that skips NaNs by default.
  3. Check whether the verification script does anything independent (recompute one row by hand from raw numbers, compare against a known benchmark, check output ranges/row counts/column names against the required format) or merely re-executes the same formula.
  4. Inspect the answer for tell-tale signs: implausibly smooth/small magnitudes, no NaN or zero at the series start, row count differing from the source file, or values that would be off by ~100x under a unit mismatch.
- **Discriminator**: A real violation is when *no* check exists that could fail independently of the main computation (data-quality assertions absent and the verifier duplicates the formula). It is fine if the script asserts input completeness/scale, handles NaNs explicitly, or cross-checks at least one output value against an externally derived or hand-computed reference — even if the code is otherwise reused.
- **Consequence**: All internal checks "pass" while the saved file diverges from the reference (shifted baseline, NaN-as-zero contamination, wrong scale or row alignment), and the grader marks the output file WRONG on numeric comparison.
16Dict/collection answers emitted without literal-syntax fidelity (unquoted keys, altered delimiters)taskinfiagent-dabench
Applies when
task -- the task prescribes an exact answer template containing a structured literal (dict/list/tuple) with string keys or labels, e.g. @name[{'k_1':v_1, ...}].
Pattern
The agent computes correct values but serializes the container by printing a Python object, f-string, or hand-typed text that drops the quotes around keys, swaps bracket/brace types, omits the name= prefix, or otherwise deviates from the template, so an exact/parse-based grader cannot match it even though the numbers are right.
Detection procedure
  1. Copy the answer template from the task statement verbatim and note every literal character: variable name, = or [], brace type, quoting of keys, separators, and key naming scheme.
  2. In the scripts, find the line that builds the final answer string and check whether it uses a faithful literal serialization (e.g. repr()/print(dict) with string keys, or a hard-coded template) rather than a loop/f-string that strips quotes.
  3. Compare the submitted answer token-by-token against the template; flag any mismatch in quoting, key spelling/order, brackets, or prefix.
  4. Optionally, attempt to parse the submitted payload with ast.literal_eval; if it raises or yields different key types than the template, flag it.
Discriminator
A real violation is a syntactic/format deviation from the stated template (unquoted or renamed keys, wrong bracket, missing prefix, wrong ordering, wrong rounding presentation). A look-alike that is fine is cosmetic whitespace or trailing-zero differences that still parse to the same literal object with the same key strings and values.
Consequence
The grader reports the expected mapping as WRONG/MISSING and scores 0 despite numerically identical values.
id 6c69bc6e9f3b · mined from infiagent-dabench dabench-450
raw text (what the judge reads)
### Dict/collection answers emitted without literal-syntax fidelity (unquoted keys, altered delimiters)
- **Applies when**: `task` -- the task prescribes an exact answer template containing a structured literal (dict/list/tuple) with string keys or labels, e.g. `@name[{'k_1':v_1, ...}]`.
- **Pattern**: The agent computes correct values but serializes the container by printing a Python object, f-string, or hand-typed text that drops the quotes around keys, swaps bracket/brace types, omits the `name=` prefix, or otherwise deviates from the template, so an exact/parse-based grader cannot match it even though the numbers are right.
- **Detection procedure**:
  1. Copy the answer template from the task statement verbatim and note every literal character: variable name, `=` or `[]`, brace type, quoting of keys, separators, and key naming scheme.
  2. In the scripts, find the line that builds the final answer string and check whether it uses a faithful literal serialization (e.g. `repr()`/`print(dict)` with string keys, or a hard-coded template) rather than a loop/f-string that strips quotes.
  3. Compare the submitted answer token-by-token against the template; flag any mismatch in quoting, key spelling/order, brackets, or prefix.
  4. Optionally, attempt to parse the submitted payload with `ast.literal_eval`; if it raises or yields different key types than the template, flag it.
- **Discriminator**: A real violation is a syntactic/format deviation from the stated template (unquoted or renamed keys, wrong bracket, missing prefix, wrong ordering, wrong rounding presentation). A look-alike that is fine is cosmetic whitespace or trailing-zero differences that still parse to the same literal object with the same key strings and values.
- **Consequence**: The grader reports the expected mapping as WRONG/MISSING and scores 0 despite numerically identical values.
17Missing evidence that the statistic was computed on the correctly cleaned data vectortaskinfiagent-dabench
Applies when
task -- the task asks for a distributional statistic or hypothesis test on a single numeric column (normality test, skewness, kurtosis, moments) and prescribes a specific test, threshold, or value to report.
Pattern
The attempt reports only the final verdict/numbers with no saved script and no intermediate diagnostics, so the vector actually fed to the test is unverifiable: NaNs, missing-value sentinels (e.g. -999, 0, blanks), strings coerced to numbers, or extra rows outside the intended subset can silently enter and dominate the moments, flipping the test outcome.
Detection procedure
  1. Read the task and list every required output and diagnostic (e.g. the test name, alpha, the p-value to be reported, rounding rules).
  2. Inspect the scripts for an explicit, reproducible chain: load → select column → coerce dtype → drop/handle missing and sentinel values → print n before and after cleaning → run the prescribed test → print p-value, skewness, kurtosis.
  3. Check the answer against these prints: is the reported p-value present, is the verdict consistent with the stated alpha, and are the moment values plausible for the printed n and value range?
  4. Flag if any script is missing, if n and cleaning steps are never printed, or if a required diagnostic (p-value) is absent from the report.
Discriminator
A genuine violation is an answer with no reproducible cleaning/diagnostic trail, or one whose extreme moment values (e.g. large |skew|, heavy kurtosis) are asserted without any check for sentinels/outliers that would explain them; it is not a violation if the script prints row counts, shows the column is clean numeric, reports the p-value, and the extreme moments are corroborated by printed summary statistics.
Consequence
The test is run on a contaminated or wrong-length vector, so the normality verdict and the skewness/kurtosis values differ from ground truth and every graded field fails, with no artifact left to diagnose the discrepancy.
id fbe935ae176c · mined from infiagent-dabench dabench-298
raw text (what the judge reads)
### Missing evidence that the statistic was computed on the correctly cleaned data vector
- **Applies when**: `task` -- the task asks for a distributional statistic or hypothesis test on a single numeric column (normality test, skewness, kurtosis, moments) and prescribes a specific test, threshold, or value to report.
- **Pattern**: The attempt reports only the final verdict/numbers with no saved script and no intermediate diagnostics, so the vector actually fed to the test is unverifiable: NaNs, missing-value sentinels (e.g. -999, 0, blanks), strings coerced to numbers, or extra rows outside the intended subset can silently enter and dominate the moments, flipping the test outcome.
- **Detection procedure**:
  1. Read the task and list every required output and diagnostic (e.g. the test name, alpha, the p-value to be reported, rounding rules).
  2. Inspect the scripts for an explicit, reproducible chain: load → select column → coerce dtype → drop/handle missing and sentinel values → print n before and after cleaning → run the prescribed test → print p-value, skewness, kurtosis.
  3. Check the answer against these prints: is the reported p-value present, is the verdict consistent with the stated alpha, and are the moment values plausible for the printed n and value range?
  4. Flag if any script is missing, if n and cleaning steps are never printed, or if a required diagnostic (p-value) is absent from the report.
- **Discriminator**: A genuine violation is an answer with no reproducible cleaning/diagnostic trail, or one whose extreme moment values (e.g. large |skew|, heavy kurtosis) are asserted without any check for sentinels/outliers that would explain them; it is *not* a violation if the script prints row counts, shows the column is clean numeric, reports the p-value, and the extreme moments are corroborated by printed summary statistics.
- **Consequence**: The test is run on a contaminated or wrong-length vector, so the normality verdict and the skewness/kurtosis values differ from ground truth and every graded field fails, with no artifact left to diagnose the discrepancy.
18No held-out validation of predictive quality before submitting predictionstaskda-code
Applies when
task -- the task asks for predictions on an unlabeled evaluation file and the scripts train a model on a labeled training file, with grading presumably based on prediction quality against hidden labels.
Pattern
The attempt fits a single default/lightly-tuned baseline (e.g., one vectorizer + one simple classifier with hand-picked hyperparameters), never splits off a labeled validation set or runs cross-validation, and reports only formatting facts (row count, column name) and the predicted class distribution as evidence of success — so no one knows whether the model clears the accuracy bar the grader uses.
Detection procedure
  1. Read the task: confirm the deliverable is per-row predictions judged against hidden ground truth, not just a file with the right shape.
  2. Read the scripts: look for any train/validation split, cross-validation, or scoring call on labeled data; also check whether more than one model/feature setting was compared.
  3. Read the answer: check whether it states a measured quality estimate (accuracy/F1 on held-out labeled data) rather than only shapes, column names, and label frequencies.
  4. If steps 2–3 find no measured score, flag: the attempt has no evidence its predictions are better than chance-level or an unaccepted baseline.
Discriminator
A real violation is the total absence of any quantitative held-out estimate (or a score computed on the same rows used for training, which is equally uninformative). It is not a violation if the attempt reports a legitimate held-out/CV score and simply chose a simple model, nor if the task explicitly only checks file format; comparing predicted vs. training label distributions alone is not a substitute for a score.
Consequence
The output file has the correct shape and column name but low agreement with the hidden labels, so an accuracy-threshold check marks the result file WRONG while the agent's self-report claims success.
id c162df6dc85b · mined from da-code dacode-ml-multi-011
raw text (what the judge reads)
### No held-out validation of predictive quality before submitting predictions
- **Applies when**: `task` -- the task asks for predictions on an unlabeled evaluation file and the scripts train a model on a labeled training file, with grading presumably based on prediction quality against hidden labels.
- **Pattern**: The attempt fits a single default/lightly-tuned baseline (e.g., one vectorizer + one simple classifier with hand-picked hyperparameters), never splits off a labeled validation set or runs cross-validation, and reports only formatting facts (row count, column name) and the predicted class distribution as evidence of success — so no one knows whether the model clears the accuracy bar the grader uses.
- **Detection procedure**:
  1. Read the task: confirm the deliverable is per-row predictions judged against hidden ground truth, not just a file with the right shape.
  2. Read the scripts: look for any train/validation split, cross-validation, or scoring call on labeled data; also check whether more than one model/feature setting was compared.
  3. Read the answer: check whether it states a measured quality estimate (accuracy/F1 on held-out labeled data) rather than only shapes, column names, and label frequencies.
  4. If steps 2–3 find no measured score, flag: the attempt has no evidence its predictions are better than chance-level or an unaccepted baseline.
- **Discriminator**: A real violation is the total absence of any quantitative held-out estimate (or a score computed on the same rows used for training, which is equally uninformative). It is *not* a violation if the attempt reports a legitimate held-out/CV score and simply chose a simple model, nor if the task explicitly only checks file format; comparing predicted vs. training label distributions alone is not a substitute for a score.
- **Consequence**: The output file has the correct shape and column name but low agreement with the hidden labels, so an accuracy-threshold check marks the result file WRONG while the agent's self-report claims success.
19Declaring success on format checks alone, with no held-out performance validationtaskda-code
Applies when
task -- the task asks for predictions on a held-out set written to a submission file, and the scripts fit a model and write the file.
Pattern
The attempt validates only superficial properties of the output (row count, column names, value range, no NaNs, mean roughly matching the target mean) and reports model hyperparameters as evidence of quality, while never computing the competition-relevant metric on a held-out/validation split, never comparing to a trivial baseline (e.g., predicting the target mean), and often silently training on a truncated subsample of the available rows or a reduced feature set.
Detection procedure
  1. Read the task to identify the evaluation target and the implied scoring metric (e.g., regression error/rank correlation/AUC) and whether all training rows/features are available.
  2. In the scripts, look for (a) a train/validation split with the metric computed on the validation part, (b) a comparison against a constant/naive baseline, and (c) whether the fit uses the full training data and the same feature set as the test-time transform.
  3. In the answer, check whether any quantitative out-of-sample score is reported, or whether only distributional/format statistics and hyperparameters are cited as "good".
  4. Compare the reported spread of predictions to the spread of the training target; a much narrower prediction range with no validation score is a red flag for severe underfitting/undertrained model.
Discriminator
A real violation reports no out-of-sample metric at all (or only format/range checks) so predictive quality is unknown; a look-alike that is fine reports a validation score computed on data excluded from fitting, ideally alongside a baseline score, even if the final answer text also mentions format checks.
Consequence
The submission file is well-formed but scores at or near a naive baseline (or below the grader's accuracy threshold), so the correctness check on the predictions fails despite all self-reported checks passing.
id 5d53bf8aed41 · mined from da-code dacode-ml-competition-008
raw text (what the judge reads)
### Declaring success on format checks alone, with no held-out performance validation
- **Applies when**: `task` -- the task asks for predictions on a held-out set written to a submission file, and the scripts fit a model and write the file.
- **Pattern**: The attempt validates only superficial properties of the output (row count, column names, value range, no NaNs, mean roughly matching the target mean) and reports model hyperparameters as evidence of quality, while never computing the competition-relevant metric on a held-out/validation split, never comparing to a trivial baseline (e.g., predicting the target mean), and often silently training on a truncated subsample of the available rows or a reduced feature set.
- **Detection procedure**:
  1. Read the task to identify the evaluation target and the implied scoring metric (e.g., regression error/rank correlation/AUC) and whether all training rows/features are available.
  2. In the scripts, look for (a) a train/validation split with the metric computed on the validation part, (b) a comparison against a constant/naive baseline, and (c) whether the fit uses the full training data and the same feature set as the test-time transform.
  3. In the answer, check whether any quantitative out-of-sample score is reported, or whether only distributional/format statistics and hyperparameters are cited as "good".
  4. Compare the reported spread of predictions to the spread of the training target; a much narrower prediction range with no validation score is a red flag for severe underfitting/undertrained model.
- **Discriminator**: A real violation reports no out-of-sample metric at all (or only format/range checks) so predictive quality is unknown; a look-alike that is fine reports a validation score computed on data excluded from fitting, ideally alongside a baseline score, even if the final answer text also mentions format checks.
- **Consequence**: The submission file is well-formed but scores at or near a naive baseline (or below the grader's accuracy threshold), so the correctness check on the predictions fails despite all self-reported checks passing.
20Ambiguous or unverified subgroup boundaries when a statistic is requested per range-defined bintaskinfiagent-dabench
Applies when
task -- the task asks for a metric/statistic computed separately for groups defined by numeric ranges (below X, between X and Y, above Y) of some column.
Pattern
The attempt picks one plausible binning without making inclusivity explicit or checking it — e.g. treating "between X and Y" as an open interval that drops rows exactly at the boundaries (or as closed, double-counting them), grouping on a column whose values are strings/half-steps so comparisons or matches silently misclassify rows, and never printing per-group row counts so that the groups do not sum to the filtered population.
Detection procedure
  1. Read the task and list the requested groups and their boundary values; note that boundary rows are typically a large share of a coarse rating/score scale.
  2. In the scripts, locate the filtering/binning code: check the comparison operators (<, <=, between), the dtype of the grouping column (numeric vs string/categorical), and whether the boundary categories are assigned to exactly one group.
  3. Verify the script prints, per group, the row count after dropping nulls in both correlated columns, and that these counts sum to the total non-null population of the filtered set (no dropped or duplicated boundary rows).
  4. Check the answer: if the middle/edge group's value is implausibly close to a neighbouring group's value, or group sizes were never reported, treat the binning as unvalidated.
Discriminator
A real violation is a binning where boundary-valued rows are excluded or misassigned, or where group counts are never shown; it is fine if the script documents the chosen convention, assigns every row in the filtered population to exactly one group, and reports counts that reconcile with the total (even if the convention is debatable, the reconciliation exposes it).
Consequence
Groups adjacent to the boundaries are computed on the wrong subset, so their correlation coefficients differ materially from the reference values while the unambiguous group matches — a partial-credit failure like 1/3 checks passed.
id c93e41af5a54 · mined from infiagent-dabench dabench-513
raw text (what the judge reads)
### Ambiguous or unverified subgroup boundaries when a statistic is requested per range-defined bin
- **Applies when**: `task` -- the task asks for a metric/statistic computed separately for groups defined by numeric ranges (below X, between X and Y, above Y) of some column.
- **Pattern**: The attempt picks one plausible binning without making inclusivity explicit or checking it — e.g. treating "between X and Y" as an open interval that drops rows exactly at the boundaries (or as closed, double-counting them), grouping on a column whose values are strings/half-steps so comparisons or matches silently misclassify rows, and never printing per-group row counts so that the groups do not sum to the filtered population.
- **Detection procedure**:
  1. Read the task and list the requested groups and their boundary values; note that boundary rows are typically a large share of a coarse rating/score scale.
  2. In the scripts, locate the filtering/binning code: check the comparison operators (`<`, `<=`, `between`), the dtype of the grouping column (numeric vs string/categorical), and whether the boundary categories are assigned to exactly one group.
  3. Verify the script prints, per group, the row count after dropping nulls in both correlated columns, and that these counts sum to the total non-null population of the filtered set (no dropped or duplicated boundary rows).
  4. Check the answer: if the middle/edge group's value is implausibly close to a neighbouring group's value, or group sizes were never reported, treat the binning as unvalidated.
- **Discriminator**: A real violation is a binning where boundary-valued rows are excluded or misassigned, or where group counts are never shown; it is fine if the script documents the chosen convention, assigns every row in the filtered population to exactly one group, and reports counts that reconcile with the total (even if the convention is debatable, the reconciliation exposes it).
- **Consequence**: Groups adjacent to the boundaries are computed on the wrong subset, so their correlation coefficients differ materially from the reference values while the unambiguous group matches — a partial-credit failure like 1/3 checks passed.
21Invented qualification thresholds / unverified output spec when the task's definition is ambiguous or truncatedtaskda-code
Applies when
task -- the task references an explicit definition, eligibility rule, or a provided sample/template output file, and part of that specification is incomplete, truncated, or not obviously reproduced in the scripts.
Pattern
The attempt guesses a cutoff (e.g., a made-up minimum count applied on a self-chosen basis such as per-item mean instead of total), applies that same filter to every sub-ranking regardless of whether the rule pertains to it, and never opens the supplied sample/template to confirm column names, ordering, key type, and row conventions.
Detection procedure
  1. Read the task/README and list every explicitly stated rule (qualification criteria, aggregation level, rounding, units, ordering) and every referenced auxiliary file (sample/template, dictionary, docs).
  2. Search the scripts for code that loads or prints each referenced auxiliary file and for constants that encode the stated rules; flag any numeric threshold or filter that appears without a documented source.
  3. Check whether each requested output quantity is computed with the definition scoped to it, i.e. whether the eligibility filter is applied only where the specification says it applies, and on the stated statistic (sum vs. mean vs. count).
  4. Compare the emitted file's header/ordering/index against the sample file's actual contents (not an assumed layout); flag if never compared.
Discriminator
A real violation is a threshold, aggregation choice, or output layout that the scripts never derive from the data, README, or sample file — it is hard-coded on intuition. It is fine if the agent inspects the sample/spec, documents the source of the rule, tests the sensitivity of the ranking to plausible alternative interpretations, or the rule genuinely has a single unambiguous reading in the provided text.
Consequence
The rankings are computed on a differently filtered population (and/or with mismatched column names/ordering) than the reference, so exact-match file comparison fails on all rows even though the underlying aggregation code is arithmetically correct.
id 162f575169db · mined from da-code dacode-dm-csv-009
raw text (what the judge reads)
### Invented qualification thresholds / unverified output spec when the task's definition is ambiguous or truncated
- **Applies when**: `task` -- the task references an explicit definition, eligibility rule, or a provided sample/template output file, and part of that specification is incomplete, truncated, or not obviously reproduced in the scripts.
- **Pattern**: The attempt guesses a cutoff (e.g., a made-up minimum count applied on a self-chosen basis such as per-item mean instead of total), applies that same filter to every sub-ranking regardless of whether the rule pertains to it, and never opens the supplied sample/template to confirm column names, ordering, key type, and row conventions.
- **Detection procedure**:
  1. Read the task/README and list every explicitly stated rule (qualification criteria, aggregation level, rounding, units, ordering) and every referenced auxiliary file (sample/template, dictionary, docs).
  2. Search the scripts for code that loads or prints each referenced auxiliary file and for constants that encode the stated rules; flag any numeric threshold or filter that appears without a documented source.
  3. Check whether each requested output quantity is computed with the definition scoped to it, i.e. whether the eligibility filter is applied only where the specification says it applies, and on the stated statistic (sum vs. mean vs. count).
  4. Compare the emitted file's header/ordering/index against the sample file's actual contents (not an assumed layout); flag if never compared.
- **Discriminator**: A real violation is a threshold, aggregation choice, or output layout that the scripts never derive from the data, README, or sample file — it is hard-coded on intuition. It is fine if the agent inspects the sample/spec, documents the source of the rule, tests the sensitivity of the ranking to plausible alternative interpretations, or the rule genuinely has a single unambiguous reading in the provided text.
- **Consequence**: The rankings are computed on a differently filtered population (and/or with mismatched column names/ordering) than the reference, so exact-match file comparison fails on all rows even though the underlying aggregation code is arithmetically correct.
22Spec file / required artifacts asserted rather than actually read and producedtaskda-code
Applies when
task -- the task points to an external specification file (config/YAML/JSON) for output formatting or binning and/or implies a set of output artifacts beyond the headline file.
Pattern
The attempt hard-codes its own titles, labels, bin edges, styling and category set, describes them in the final answer as if they came from the spec, and emits only the single obvious output file while silently skipping the other expected artifacts (serialized plot data, saved arrays); the derived group counts are never checked against the total row count of the filtered subset.
Detection procedure
  1. From the task text, list (a) every file that must be read as a specification and (b) every artifact that must be written.
  2. In the scripts, confirm the spec file is actually loaded and its values are used to drive the plot (bins, labels, title, order, figure params) instead of literals appearing in the code; confirm each expected artifact is written by an explicit save call.
  3. Compare the answer's stated categories/counts with the spec's own categories and with the dataset: does the number of groups match the spec, and do the group counts sum to the size of the filtered subset (and to a plausible fraction of the raw file)?
  4. Flag if any spec value is invented, any artifact is unwritten, or the counts fail the sum/magnitude sanity check.
Discriminator
A fine attempt reads the spec and its literals coincide with the spec values (verifiable by comparing code/output to the file contents) and writes all requested artifacts; a violation is code whose formatting/binning choices exist nowhere in the spec, or an answer that claims "all specifications applied" without the spec ever being parsed, or an obviously truncated subset count with no reconciliation.
Consequence
The grader checks each expected artifact (plot metadata, image, saved array); missing files and mismatched bins/counts make every check fail even though the script ran without error.
id be758efe9c22 · mined from da-code dacode-plot-bar-007
raw text (what the judge reads)
### Spec file / required artifacts asserted rather than actually read and produced
- **Applies when**: `task` -- the task points to an external specification file (config/YAML/JSON) for output formatting or binning and/or implies a set of output artifacts beyond the headline file.
- **Pattern**: The attempt hard-codes its own titles, labels, bin edges, styling and category set, describes them in the final answer as if they came from the spec, and emits only the single obvious output file while silently skipping the other expected artifacts (serialized plot data, saved arrays); the derived group counts are never checked against the total row count of the filtered subset.
- **Detection procedure**:
  1. From the task text, list (a) every file that must be *read* as a specification and (b) every artifact that must be *written*.
  2. In the scripts, confirm the spec file is actually loaded and its values are used to drive the plot (bins, labels, title, order, figure params) instead of literals appearing in the code; confirm each expected artifact is written by an explicit save call.
  3. Compare the answer's stated categories/counts with the spec's own categories and with the dataset: does the number of groups match the spec, and do the group counts sum to the size of the filtered subset (and to a plausible fraction of the raw file)?
  4. Flag if any spec value is invented, any artifact is unwritten, or the counts fail the sum/magnitude sanity check.
- **Discriminator**: A fine attempt reads the spec and its literals coincide with the spec values (verifiable by comparing code/output to the file contents) and writes all requested artifacts; a violation is code whose formatting/binning choices exist nowhere in the spec, or an answer that claims "all specifications applied" without the spec ever being parsed, or an obviously truncated subset count with no reconciliation.
- **Consequence**: The grader checks each expected artifact (plot metadata, image, saved array); missing files and mismatched bins/counts make every check fail even though the script ran without error.
23Stated output ordering/serialization constraint not enforced in codetaskda-code
Applies when
task -- the task specifies how the final answer must be ordered/formatted (e.g., "sort from highest to lowest") and/or written to a named result file, and the scripts produce ranked lists or dictionaries.
Pattern
The scripts extract the right records but leave them in whatever order the selection helper returns (e.g., an ascending "smallest-n" query for the low group while the high group is descending), and only print results instead of emitting the requested artifact/JSON structure — so the stated ordering/format constraint is never explicitly applied.
Detection procedure
  1. Read the task and list every explicit output constraint: ordering direction, grouping keys, rounding/units, and the required file name/JSON schema.
  2. In the scripts, locate each list-producing step and check whether an explicit sort (or reverse) with the required direction is applied to every output list, not just some of them; check whether the final structure is serialized to the required file.
  3. Compare the printed/reported lists against the constraint: for a "highest to lowest" requirement, verify the associated values decrease monotonically in each list.
  4. Flag if any list's implied value order contradicts the requirement, or if no code writes the required output artifact.
Discriminator
A real violation is when the code contains no ordering/serialization step matching the stated requirement (or the reported values run the wrong way); it is fine if a selection helper already returns the required direction and the reported values verifiably follow it, and the required file is written with the requested keys.
Consequence
The grader compares the expected result file element-by-element and marks the answer wrong (or missing) even though the correct set of records was identified.
id 5623989d04a7 · mined from da-code dacode-di-text-003
raw text (what the judge reads)
### Stated output ordering/serialization constraint not enforced in code
- **Applies when**: `task` -- the task specifies how the final answer must be ordered/formatted (e.g., "sort from highest to lowest") and/or written to a named result file, and the scripts produce ranked lists or dictionaries.
- **Pattern**: The scripts extract the right records but leave them in whatever order the selection helper returns (e.g., an ascending "smallest-n" query for the low group while the high group is descending), and only print results instead of emitting the requested artifact/JSON structure — so the stated ordering/format constraint is never explicitly applied.
- **Detection procedure**:
  1. Read the task and list every explicit output constraint: ordering direction, grouping keys, rounding/units, and the required file name/JSON schema.
  2. In the scripts, locate each list-producing step and check whether an explicit sort (or reverse) with the required direction is applied to *every* output list, not just some of them; check whether the final structure is serialized to the required file.
  3. Compare the printed/reported lists against the constraint: for a "highest to lowest" requirement, verify the associated values decrease monotonically in each list.
  4. Flag if any list's implied value order contradicts the requirement, or if no code writes the required output artifact.
- **Discriminator**: A real violation is when the code contains no ordering/serialization step matching the stated requirement (or the reported values run the wrong way); it is fine if a selection helper already returns the required direction and the reported values verifiably follow it, and the required file is written with the requested keys.
- **Consequence**: The grader compares the expected result file element-by-element and marks the answer wrong (or missing) even though the correct set of records was identified.
24Answer string doesn't literally match the requested output template (quoting/delimiters/tokens)taskinfiagent-dabench
Applies when
task -- The task specifies a rigid answer template with tagged fields (e.g. @field[value]) and shows the exact literal form the values should take (quoted strings, units, ranges, allowed vocabulary).
Pattern
The agent computes plausible or even correct values but emits them in a different surface form than the template demands — dropping the quotation marks shown in the spec, changing separators or spacing, reordering fields, adding extra prose/units, or paraphrasing an allowed label — so a literal string matcher scores every field wrong. This is often compounded by no saved script, so nothing else in the submission can be checked or re-derived.
Detection procedure
  1. Copy the answer-format line from the task verbatim and list each field, its delimiters, and exactly how the sample value is written (with or without quotes, capitalization, allowed token set).
  2. Copy the agent's final answer string and diff it character-by-character against that template, field by field and in order.
  3. Check that each value is drawn from the task's stated vocabulary/format (e.g. the literal range string, the enumerated category labels) rather than a synonym, computed variant, or extra-annotated form.
  4. Confirm a script exists that produces those exact strings (or that the answer is at least reproducible), rather than being hand-typed with no artifact.
Discriminator
A real violation is any deviation in the characters a strict matcher sees — missing/added quotes, wrong field order, extra text inside brackets. A look-alike that is fine is a difference the task explicitly leaves free (e.g. whitespace between fields when the spec shows none but no strictness is implied) or a value that is spelled exactly as the task's enumerated option even if worded differently from the agent's internal computation.
Consequence
The grader reports 0/N checks passed with "WRONG/MISSING" for fields whose semantic content is actually right, so a substantively correct analysis scores zero.
id 6703d17fd1ed · mined from infiagent-dabench dabench-550
raw text (what the judge reads)
### Answer string doesn't literally match the requested output template (quoting/delimiters/tokens)
- **Applies when**: `task` -- The task specifies a rigid answer template with tagged fields (e.g. `@field[value]`) and shows the exact literal form the values should take (quoted strings, units, ranges, allowed vocabulary).
- **Pattern**: The agent computes plausible or even correct values but emits them in a different surface form than the template demands — dropping the quotation marks shown in the spec, changing separators or spacing, reordering fields, adding extra prose/units, or paraphrasing an allowed label — so a literal string matcher scores every field wrong. This is often compounded by no saved script, so nothing else in the submission can be checked or re-derived.
- **Detection procedure**:
  1. Copy the answer-format line from the task verbatim and list each field, its delimiters, and exactly how the sample value is written (with or without quotes, capitalization, allowed token set).
  2. Copy the agent's final answer string and diff it character-by-character against that template, field by field and in order.
  3. Check that each value is drawn from the task's stated vocabulary/format (e.g. the literal range string, the enumerated category labels) rather than a synonym, computed variant, or extra-annotated form.
  4. Confirm a script exists that produces those exact strings (or that the answer is at least reproducible), rather than being hand-typed with no artifact.
- **Discriminator**: A real violation is any deviation in the characters a strict matcher sees — missing/added quotes, wrong field order, extra text inside brackets. A look-alike that is fine is a difference the task explicitly leaves free (e.g. whitespace between fields when the spec shows none but no strictness is implied) or a value that is spelled exactly as the task's enumerated option even if worded differently from the agent's internal computation.
- **Consequence**: The grader reports 0/N checks passed with "WRONG/MISSING" for fields whose semantic content is actually right, so a substantively correct analysis scores zero.
25Substituting a "creative interpretation" for the requested entities and metricstaskda-code
Applies when
task -- the task names specific entities, groupings, measures, config files, or output artifacts, and the scripts operate on a file whose columns do not contain them.
Pattern
Instead of locating the correct input data (or the correct columns) and producing every requested artifact, the attempt declares a "data mismatch" and maps the requested concepts onto unrelated available fields as proxies, then reports success on that redefined problem; stated config/settings and secondary output files are ignored or only nominally loaded.
Detection procedure
  1. From the task, list the required inputs (entities to group by, measure to rank by, quantities to aggregate), any settings file that must be applied, and every output file expected.
  2. In the scripts, check whether each listed item is read from a real column/field with matching semantics, or is invented/renamed as a stand-in.
  3. Check that the settings file's contents actually drive the output (figure size, labels, colors, ordering) rather than being loaded and discarded, and that all expected output artifacts are written.
  4. In the answer, look for language such as "proxy", "interpretation", "mismatch", "mapped X to Y" — a self-declared redefinition of the task.
Discriminator
A genuine violation redefines what is being measured or grouped, or silently drops required outputs. A look-alike that is fine is when the correct fields exist but must be derived by documented, semantics-preserving steps (parsing dates into durations, joining tables, standardizing codes) with the requested measure and grouping preserved; also fine is exhausting the available data files first and documenting that the needed field is genuinely absent, rather than assuming it after inspecting one file.
Consequence
The saved figure and any numeric export encode the wrong entities, ordering, and units, so every value-level and file-level check fails (missing expected artifacts, mismatched arrays), even though a plot was produced.
id 7e32c9b270d6 · mined from da-code dacode-plot-scatter-002
raw text (what the judge reads)
### Substituting a "creative interpretation" for the requested entities and metrics
- **Applies when**: `task` -- the task names specific entities, groupings, measures, config files, or output artifacts, and the scripts operate on a file whose columns do not contain them.
- **Pattern**: Instead of locating the correct input data (or the correct columns) and producing every requested artifact, the attempt declares a "data mismatch" and maps the requested concepts onto unrelated available fields as proxies, then reports success on that redefined problem; stated config/settings and secondary output files are ignored or only nominally loaded.
- **Detection procedure**:
  1. From the task, list the required inputs (entities to group by, measure to rank by, quantities to aggregate), any settings file that must be applied, and every output file expected.
  2. In the scripts, check whether each listed item is read from a real column/field with matching semantics, or is invented/renamed as a stand-in.
  3. Check that the settings file's contents actually drive the output (figure size, labels, colors, ordering) rather than being loaded and discarded, and that all expected output artifacts are written.
  4. In the answer, look for language such as "proxy", "interpretation", "mismatch", "mapped X to Y" — a self-declared redefinition of the task.
- **Discriminator**: A genuine violation redefines what is being measured or grouped, or silently drops required outputs. A look-alike that is fine is when the correct fields exist but must be derived by documented, semantics-preserving steps (parsing dates into durations, joining tables, standardizing codes) with the requested measure and grouping preserved; also fine is exhausting the available data files first and documenting that the needed field is genuinely absent, rather than assuming it after inspecting one file.
- **Consequence**: The saved figure and any numeric export encode the wrong entities, ordering, and units, so every value-level and file-level check fails (missing expected artifacts, mismatched arrays), even though a plot was produced.
26Output artifact name/format does not match the exact specificationtaskda-code
Applies when
task -- the task names a specific output file (or template/schema) that the deliverable must be saved as.
Pattern
The agent computes plausible numbers but writes them to a file whose name, spelling, location, or column/index layout differs from the one literally requested (e.g., "corrected" spelling, different directory, missing index column or header names from the template), and the answer text describes the self-chosen name as if it were the requirement.
Detection procedure
1. Copy the exact filename/path and any template column/index names from the task statement. 2. In the scripts, find every write call (to_csv, to_excel, savefig, etc.) and compare the literal string and the written frame's columns/index to step 1, character for character. 3. In the answer, check which filename and structure the agent claims to have produced. 4. Flag if any of the three differ, or if the agent never verifies the saved file by reading it back and comparing to the template.
Discriminator
A real violation is any deviation in the literal artifact name/path or in the template's column/index structure (including a "fixed" typo). Not a violation if the requested file is written exactly as specified and extra helper files also exist, or if the agent additionally saves an alias copy under the required name.
Consequence
The grader looks for the specified file and reports it as WRONG/MISSING, scoring 0 regardless of whether the underlying computation was correct.
id fbf0ce2a8103 · mined from da-code dacode-dm-csv-043
raw text (what the judge reads)
### Output artifact name/format does not match the exact specification
- **Applies when**: `task` -- the task names a specific output file (or template/schema) that the deliverable must be saved as.
- **Pattern**: The agent computes plausible numbers but writes them to a file whose name, spelling, location, or column/index layout differs from the one literally requested (e.g., "corrected" spelling, different directory, missing index column or header names from the template), and the answer text describes the self-chosen name as if it were the requirement.
- **Detection procedure**: 1. Copy the exact filename/path and any template column/index names from the task statement. 2. In the scripts, find every write call (`to_csv`, `to_excel`, `savefig`, etc.) and compare the literal string and the written frame's columns/index to step 1, character for character. 3. In the answer, check which filename and structure the agent claims to have produced. 4. Flag if any of the three differ, or if the agent never verifies the saved file by reading it back and comparing to the template.
- **Discriminator**: A real violation is any deviation in the literal artifact name/path or in the template's column/index structure (including a "fixed" typo). Not a violation if the requested file is written exactly as specified and extra helper files also exist, or if the agent additionally saves an alias copy under the required name.
- **Consequence**: The grader looks for the specified file and reports it as WRONG/MISSING, scoring 0 regardless of whether the underlying computation was correct.
27Fabricated/synthetic data substituted for the provided datasettaskda-code
Applies when
task -- the task references a supplied dataset (with a README/data dictionary) and the script must load it to compute the reported numbers.
Pattern
The script never reads any input file; instead it generates data with a random generator (or hard-codes values) that merely mimics the expected schema, then runs the requested statistics on that invented data and reports the results as the answer.
Detection procedure
1. Read the task to confirm an external data source is expected. 2. Scan the script for any file-reading call (read_csv, read_excel, load_dataset, DB query, path strings); if the only data construction is np.random.*, seed(...) loops, or literal lists, flag it. 3. Cross-check reported category names, group counts and sample sizes against the README/actual file — invented data usually shows suspiciously round or uniform group sizes. 4. Verify the answer is written to the exact requested output file/format, not an ad-hoc filename.
Discriminator
Legitimate uses of random generation are auxiliary (bootstrap resampling, simulation baselines, train/test shuffling) applied on top of loaded real data; the violation is when the reported statistics derive entirely from generated values with no real data path.
Consequence
The p-values and conclusions are unrelated to the true data, so every value mismatches the expected result file (and the output file may be missing/misnamed), yielding 0 checks passed.
id f43f43c20432 · mined from da-code dacode-data-sa-061
raw text (what the judge reads)
### Fabricated/synthetic data substituted for the provided dataset
- **Applies when**: `task` -- the task references a supplied dataset (with a README/data dictionary) and the script must load it to compute the reported numbers.
- **Pattern**: The script never reads any input file; instead it generates data with a random generator (or hard-codes values) that merely mimics the expected schema, then runs the requested statistics on that invented data and reports the results as the answer.
- **Detection procedure**: 1. Read the task to confirm an external data source is expected. 2. Scan the script for any file-reading call (`read_csv`, `read_excel`, `load_dataset`, DB query, path strings); if the only data construction is `np.random.*`, `seed(...)` loops, or literal lists, flag it. 3. Cross-check reported category names, group counts and sample sizes against the README/actual file — invented data usually shows suspiciously round or uniform group sizes. 4. Verify the answer is written to the exact requested output file/format, not an ad-hoc filename.
- **Discriminator**: Legitimate uses of random generation are auxiliary (bootstrap resampling, simulation baselines, train/test shuffling) applied *on top of* loaded real data; the violation is when the reported statistics derive entirely from generated values with no real data path.
- **Consequence**: The p-values and conclusions are unrelated to the true data, so every value mismatches the expected result file (and the output file may be missing/misnamed), yielding 0 checks passed.
28Discarding high-signal columns and accepting a near-baseline validation scoretaskda-code
Applies when
task -- a predictive task where the provided table contains identifier, categorical, and date/text columns alongside a handful of continuous measurements, and the scripts feed only the continuous measurements into the model.
Pattern
The attempt builds an ensemble of tree models on a hand-picked list of "numeric" columns, drops or never inspects the entity/name/date/genre/text columns (and never checks whether test rows can be matched back to rows of the reference table by identifier), and then ships predictions even though holdout error is essentially the same as predicting the target mean.
Detection procedure
  1. From the task/README and the exploration script, list all columns present in the reference table; from the modeling script, list the columns actually used as features.
  2. Flag every dropped column that plausibly carries signal (identifiers, entity/creator names, dates or derived time features, categorical labels, text fields) and check whether the script attempted any encoding, aggregation, target-derived grouping, or an identifier join/overlap check between test rows and the reference table.
  3. Read the reported validation metric and compare it to the trivial baseline computable from the printed target statistics (e.g., RMSE vs. the target's standard deviation, or the printed R²); also compare prediction spread (min/max/std) to the target's spread.
  4. Confirm whether the attempt reported this gap as acceptable and stopped, without trying richer features or a deduplication/lookup path.
Discriminator
A real violation is when informative columns were available and untried and the achieved error is at or near the mean-baseline (R² ≈ 0, prediction std ≪ target std, heavily shrunk range). It is not a violation if the dropped columns are genuinely uninformative (free-form IDs with no reuse, constant columns) or if the agent tested encodings/joins and documented that they did not help while the model clearly beats the baseline.
Consequence
Predictions collapse toward the training mean, so any accuracy/correlation threshold the grader applies to the submitted file fails even though the file has the right name, shape, and column header.
id 201ddc2f86f9 · mined from da-code dacode-ml-regression-004
raw text (what the judge reads)
### Discarding high-signal columns and accepting a near-baseline validation score
- **Applies when**: `task` -- a predictive task where the provided table contains identifier, categorical, and date/text columns alongside a handful of continuous measurements, and the scripts feed only the continuous measurements into the model.
- **Pattern**: The attempt builds an ensemble of tree models on a hand-picked list of "numeric" columns, drops or never inspects the entity/name/date/genre/text columns (and never checks whether test rows can be matched back to rows of the reference table by identifier), and then ships predictions even though holdout error is essentially the same as predicting the target mean.
- **Detection procedure**:
  1. From the task/README and the exploration script, list all columns present in the reference table; from the modeling script, list the columns actually used as features.
  2. Flag every dropped column that plausibly carries signal (identifiers, entity/creator names, dates or derived time features, categorical labels, text fields) and check whether the script attempted any encoding, aggregation, target-derived grouping, or an identifier join/overlap check between test rows and the reference table.
  3. Read the reported validation metric and compare it to the trivial baseline computable from the printed target statistics (e.g., RMSE vs. the target's standard deviation, or the printed R²); also compare prediction spread (min/max/std) to the target's spread.
  4. Confirm whether the attempt reported this gap as acceptable and stopped, without trying richer features or a deduplication/lookup path.
- **Discriminator**: A real violation is when informative columns were available and untried *and* the achieved error is at or near the mean-baseline (R² ≈ 0, prediction std ≪ target std, heavily shrunk range). It is not a violation if the dropped columns are genuinely uninformative (free-form IDs with no reuse, constant columns) or if the agent tested encodings/joins and documented that they did not help while the model clearly beats the baseline.
- **Consequence**: Predictions collapse toward the training mean, so any accuracy/correlation threshold the grader applies to the submitted file fails even though the file has the right name, shape, and column header.
29Blind argmax of one internal metric for cluster count, with no sanity check on resulting cluster sizestaskda-code
Applies when
task -- the task asks for an "appropriate" number of groups/hyperparameter and the script picks it by taking the single best value of one unsupervised score over a grid.
Pattern
The script sweeps k, selects argmax(silhouette) (or argmin(DB)) alone, and accepts the result even though the chosen solution is a poor/degenerate partition — e.g. clusters containing 1–3 points that are really outliers, an absolute score near 0.3 that is barely above the neighbouring k values, and no cross-check against the elbow curve, alternative metrics, robustness to seed, or outlier/skew handling (log-transform, winsorizing) of heavy-tailed features.
Detection procedure
  1. Read the task for how the number of groups is to be justified and what the output must represent.
  2. In the script, check whether the selection uses more than one criterion / any agreement check, and whether a minimum-cluster-size or stability check gates the final choice.
  3. In the reported output, inspect the cluster size distribution and the score table: flag singleton/near-singleton clusters, or a winning score that differs from the runner-up by a trivial margin.
  4. Check whether skewed/outlier-heavy variables were transformed or outliers handled before distance-based clustering; if not, the "winning" k is likely just isolating outliers.
Discriminator
A fine attempt either shows the metric has a clear, well-separated optimum with balanced, interpretable groups, or explicitly justifies small clusters as genuine outlier groups after robustness checks; a violation is accepting a marginal metric win that yields degenerate clusters with no corroborating evidence or interpretation.
Consequence
The saved label column encodes an outlier-driven partition whose cluster count and assignments disagree with the expected grouping, so the file-level comparison (number of clusters / label agreement) fails even though the pipeline runs without error.
id 915988065c7c · mined from da-code dacode-ml-cluster-013
raw text (what the judge reads)
### Blind argmax of one internal metric for cluster count, with no sanity check on resulting cluster sizes
- **Applies when**: `task` -- the task asks for an "appropriate" number of groups/hyperparameter and the script picks it by taking the single best value of one unsupervised score over a grid.
- **Pattern**: The script sweeps k, selects `argmax(silhouette)` (or `argmin(DB)`) alone, and accepts the result even though the chosen solution is a poor/degenerate partition — e.g. clusters containing 1–3 points that are really outliers, an absolute score near 0.3 that is barely above the neighbouring k values, and no cross-check against the elbow curve, alternative metrics, robustness to seed, or outlier/skew handling (log-transform, winsorizing) of heavy-tailed features.
- **Detection procedure**:
  1. Read the task for how the number of groups is to be justified and what the output must represent.
  2. In the script, check whether the selection uses more than one criterion / any agreement check, and whether a minimum-cluster-size or stability check gates the final choice.
  3. In the reported output, inspect the cluster size distribution and the score table: flag singleton/near-singleton clusters, or a winning score that differs from the runner-up by a trivial margin.
  4. Check whether skewed/outlier-heavy variables were transformed or outliers handled before distance-based clustering; if not, the "winning" k is likely just isolating outliers.
- **Discriminator**: A fine attempt either shows the metric has a clear, well-separated optimum with balanced, interpretable groups, or explicitly justifies small clusters as genuine outlier groups after robustness checks; a violation is accepting a marginal metric win that yields degenerate clusters with no corroborating evidence or interpretation.
- **Consequence**: The saved label column encodes an outlier-driven partition whose cluster count and assignments disagree with the expected grouping, so the file-level comparison (number of clusters / label agreement) fails even though the pipeline runs without error.
30Per-group statistic computed over an unverified grouping/axistaskinfiagent-dabench
Applies when
task -- the task asks for the entity (country, user, product…) that maximizes a distribution statistic (skewness, variance, mean, etc.) computed within each entity, so each entity's statistic depends on which rows/columns are collapsed into its sample.
Pattern
The attempt computes the statistic without an explicit, inspectable definition of the per-entity sample: it aggregates along the wrong axis (e.g., across entities instead of within one), collapses a filter it should have kept (or keeps rows it should have filtered), silently drops/keeps NaNs and non-numeric values, and then reports only the single argmax — no group sizes, no ranked list, no reproducible script — so an off-by-one-grouping error is invisible.
Detection procedure
  1. From the task, write down exactly what one entity's sample should be: which rows are selected, which column(s)/periods supply the values, and how many values are expected per entity.
  2. In the script, locate the groupby/loop and the statistic call; confirm the grouping key is the entity, the value vector is the intended one (not transposed, not a row across entities), the required definition/flag is set (e.g., Fisher/bias options), and NaN/dtype handling is explicit.
  3. Check the script prints diagnostics: per-entity sample counts, number of entities, and the top-5 ranked statistic values — not just the argmax.
  4. If scripts are missing or the answer is a bare label with no printed intermediate table, treat the result as unverifiable and reject.
Discriminator
A real violation is when the reviewer cannot reconstruct, from the code, which values form each entity's sample (or can see the axis/filter/NaN handling is inconsistent with the task wording), or when entities with degenerate samples (n≤2, all-NaN, constant) are ranked alongside valid ones. It is fine if the grouping and value selection are explicit and matching, and diagnostics show plausible counts and a stable margin between the top candidates — even if the chosen library call differs stylistically.
Consequence
The argmax shifts to a neighboring entity whose sample was built differently, so the reported name mismatches ground truth and the grader records 0/1 with no trace to diagnose.
id bc7e5fbfddb6 · mined from infiagent-dabench dabench-252
raw text (what the judge reads)
### Per-group statistic computed over an unverified grouping/axis
- **Applies when**: `task` -- the task asks for the entity (country, user, product…) that maximizes a distribution statistic (skewness, variance, mean, etc.) computed within each entity, so each entity's statistic depends on which rows/columns are collapsed into its sample.
- **Pattern**: The attempt computes the statistic without an explicit, inspectable definition of the per-entity sample: it aggregates along the wrong axis (e.g., across entities instead of within one), collapses a filter it should have kept (or keeps rows it should have filtered), silently drops/keeps NaNs and non-numeric values, and then reports only the single argmax — no group sizes, no ranked list, no reproducible script — so an off-by-one-grouping error is invisible.
- **Detection procedure**:
  1. From the task, write down exactly what one entity's sample should be: which rows are selected, which column(s)/periods supply the values, and how many values are expected per entity.
  2. In the script, locate the groupby/loop and the statistic call; confirm the grouping key is the entity, the value vector is the intended one (not transposed, not a row across entities), the required definition/flag is set (e.g., Fisher/bias options), and NaN/dtype handling is explicit.
  3. Check the script prints diagnostics: per-entity sample counts, number of entities, and the top-5 ranked statistic values — not just the argmax.
  4. If scripts are missing or the answer is a bare label with no printed intermediate table, treat the result as unverifiable and reject.
- **Discriminator**: A real violation is when the reviewer cannot reconstruct, from the code, which values form each entity's sample (or can see the axis/filter/NaN handling is inconsistent with the task wording), or when entities with degenerate samples (n≤2, all-NaN, constant) are ranked alongside valid ones. It is fine if the grouping and value selection are explicit and matching, and diagnostics show plausible counts and a stable margin between the top candidates — even if the chosen library call differs stylistically.
- **Consequence**: The argmax shifts to a neighboring entity whose sample was built differently, so the reported name mismatches ground truth and the grader records 0/1 with no trace to diagnose.
31Answer payload mismatch: dumping computed values instead of the requested identifiertaskinfiagent-dabench
Applies when
task -- the task asks the agent to create/derive something (a new column, a chosen feature, a selected model) and specifies an answer template whose placeholder names an identifier or label rather than the underlying data.
Pattern
The agent correctly performs the computation but fills the answer slot with the full vector of computed values (often truncated or with inconsistent rounding), rather than the single requested identifier/name described by the placeholder.
Detection procedure
1. Read the answer-format spec and identify exactly what the placeholder denotes (a name/label vs. a numeric list vs. a single statistic). 2. Check the placeholder's descriptive text and the task verb ("create a feature called…", "which variable…") to infer arity: one token or many. 3. Compare the agent's submitted payload to that arity/type; flag if it substitutes a long value list, an intermediate table, or extra commentary for a single identifier (or vice versa). 4. Verify any stated formatting constraints (rounding, ordering, delimiter, units) are applied to whatever is submitted.
Discriminator
A real violation is a type/arity mismatch with the placeholder's stated meaning (e.g., hundreds of numbers where a column name is described, or truncated output for a required full list). It is not a violation if the task genuinely requests every per-row value and the agent supplies them completely, in order, with the specified precision.
Consequence
The grader's exact-match on the expected identifier fails, scoring 0 even though the underlying computation was correct.
id 69d360903ae5 · mined from infiagent-dabench dabench-741
raw text (what the judge reads)
### Answer payload mismatch: dumping computed values instead of the requested identifier
- **Applies when**: `task` -- the task asks the agent to create/derive something (a new column, a chosen feature, a selected model) and specifies an answer template whose placeholder names an identifier or label rather than the underlying data.
- **Pattern**: The agent correctly performs the computation but fills the answer slot with the full vector of computed values (often truncated or with inconsistent rounding), rather than the single requested identifier/name described by the placeholder.
- **Detection procedure**: 1. Read the answer-format spec and identify exactly what the placeholder denotes (a name/label vs. a numeric list vs. a single statistic). 2. Check the placeholder's descriptive text and the task verb ("create a feature called…", "which variable…") to infer arity: one token or many. 3. Compare the agent's submitted payload to that arity/type; flag if it substitutes a long value list, an intermediate table, or extra commentary for a single identifier (or vice versa). 4. Verify any stated formatting constraints (rounding, ordering, delimiter, units) are applied to whatever is submitted.
- **Discriminator**: A real violation is a type/arity mismatch with the placeholder's stated meaning (e.g., hundreds of numbers where a column name is described, or truncated output for a required full list). It is *not* a violation if the task genuinely requests every per-row value and the agent supplies them completely, in order, with the specified precision.
- **Consequence**: The grader's exact-match on the expected identifier fails, scoring 0 even though the underlying computation was correct.
32Failure to read and conform to the provided output-format templatetaskda-code
Applies when
task -- the task says the deliverable file must match the format of a supplied sample/example result file (or otherwise specifies an exact output schema).
Pattern
The scripts never load or print the sample file; the agent invents column names, column order, row ordering, and rounding/precision from intuition, then saves the result and "verifies" only its own arithmetic, so the output can be numerically plausible yet schematically wrong.
Detection procedure
  1. Read the task and note every stated output constraint (template file, file name, column names, units, rounding, sort order).
  2. Search the scripts for any read/inspection of the template file (e.g., read_csv of the sample, printing its header/rows) and any comparison of the produced columns to it.
  3. If absent, check whether the header, column order, row order, and numeric formatting in the submitted answer are asserted anywhere or merely hard-coded by guess (e.g., an arbitrary .round(2), self-chosen labels, self-chosen sort key).
  4. Flag when the deliverable's schema/precision has no traceable source in the provided template.
Discriminator
Not a violation if the script actually loads the template (or explicitly reproduces its exact header/ordering/precision verified against it) and only then writes the output; it is a violation when the format is inferred from the agent's own phrasing of the task, even if the underlying aggregation logic is correct.
Consequence
The grader compares the file against the expected schema/values and marks it WRONG/MISSING because headers, ordering, or rounded values do not match, despite correct intermediate computations.
id 9662e0737271 · mined from da-code dacode-dm-csv-010
raw text (what the judge reads)
### Failure to read and conform to the provided output-format template
- **Applies when**: `task` -- the task says the deliverable file must match the format of a supplied sample/example result file (or otherwise specifies an exact output schema).
- **Pattern**: The scripts never load or print the sample file; the agent invents column names, column order, row ordering, and rounding/precision from intuition, then saves the result and "verifies" only its own arithmetic, so the output can be numerically plausible yet schematically wrong.
- **Detection procedure**:
  1. Read the task and note every stated output constraint (template file, file name, column names, units, rounding, sort order).
  2. Search the scripts for any read/inspection of the template file (e.g., `read_csv` of the sample, printing its header/rows) and any comparison of the produced columns to it.
  3. If absent, check whether the header, column order, row order, and numeric formatting in the submitted answer are asserted anywhere or merely hard-coded by guess (e.g., an arbitrary `.round(2)`, self-chosen labels, self-chosen sort key).
  4. Flag when the deliverable's schema/precision has no traceable source in the provided template.
- **Discriminator**: Not a violation if the script actually loads the template (or explicitly reproduces its exact header/ordering/precision verified against it) and only then writes the output; it *is* a violation when the format is inferred from the agent's own phrasing of the task, even if the underlying aggregation logic is correct.
- **Consequence**: The grader compares the file against the expected schema/values and marks it WRONG/MISSING because headers, ordering, or rounded values do not match, despite correct intermediate computations.
33Coarsening an identifier value to fit a format hint, losing required precisiontaskinfiagent-dabench
Applies when
task -- the answer includes a key/label (date, ID, category, index) that must be located in the data, and the answer template shows a format or example whose granularity may be coarser than the data's actual granularity.
Pattern
The attempt correctly locates the record but then truncates, rounds, or re-renders the identifier (e.g., drops components of a composite/hierarchical key) to match a literal format hint, reporting a value that no longer uniquely identifies the record it found — even though downstream computations were done on the full-precision value.
Detection procedure
  1. In the task, note the granularity of the identifier the data actually contains and whether any other constraint (e.g., "previous record", per-record lookup) implies the full-precision key is needed.
  2. In the scripts/output, find the raw identifier of the located record and compare it to the string finally emitted; check for truncation, reformatting, or aggregation of the key.
  3. Check internal consistency: does the emitted identifier, if fed back into the data, select exactly one record — and the same record used for the dependent calculation?
  4. If the emitted key is coarser than the located record (maps to many rows) while the dependent value came from a single row, flag it.
Discriminator
A real violation is when the reported key is ambiguous or lower-precision than the record actually used, so it cannot be verified against the data; it is not a violation when the data genuinely has that coarse granularity (one row per that key), or when the task explicitly asks for an aggregate over the coarser grouping and the dependent computation was done at that same level.
Consequence
The identifier check fails on exact-string comparison against the ground-truth full-precision key, so the submission is marked partially correct at best (dependent numeric value may pass while the key fails).
id a16ef15bdf8e · mined from infiagent-dabench dabench-572
raw text (what the judge reads)
### Coarsening an identifier value to fit a format hint, losing required precision
- **Applies when**: `task` -- the answer includes a key/label (date, ID, category, index) that must be located in the data, and the answer template shows a format or example whose granularity may be coarser than the data's actual granularity.
- **Pattern**: The attempt correctly locates the record but then truncates, rounds, or re-renders the identifier (e.g., drops components of a composite/hierarchical key) to match a literal format hint, reporting a value that no longer uniquely identifies the record it found — even though downstream computations were done on the full-precision value.
- **Detection procedure**:
  1. In the task, note the granularity of the identifier the data actually contains and whether any other constraint (e.g., "previous record", per-record lookup) implies the full-precision key is needed.
  2. In the scripts/output, find the raw identifier of the located record and compare it to the string finally emitted; check for truncation, reformatting, or aggregation of the key.
  3. Check internal consistency: does the emitted identifier, if fed back into the data, select exactly one record — and the same record used for the dependent calculation?
  4. If the emitted key is coarser than the located record (maps to many rows) while the dependent value came from a single row, flag it.
- **Discriminator**: A real violation is when the reported key is ambiguous or lower-precision than the record actually used, so it cannot be verified against the data; it is *not* a violation when the data genuinely has that coarse granularity (one row per that key), or when the task explicitly asks for an aggregate over the coarser grouping and the dependent computation was done at that same level.
- **Consequence**: The identifier check fails on exact-string comparison against the ground-truth full-precision key, so the submission is marked partially correct at best (dependent numeric value may pass while the key fails).
34Substituting proxy/external data for the variables the task actually namestaskda-code
Applies when
task -- the task asks for a statistic from a model built on specific named quantities from a provided dataset, and the script instead pulls data from an external source or swaps in a "close enough" stand-in series.
Pattern
The agent cannot locate (or does not look for) the required variables in the supplied data, so it downloads unrelated series or redefines a variable as a proxy (e.g., a rate instead of the named quantity, a different index as the "portfolio"), sometimes with a different frequency, aggregation, or date coverage, then reports the statistic from that substitute model as the answer while acknowledging the substitution in the write-up.
Detection procedure
1. From the task text, list the exact dependent and independent quantities, the aggregation (e.g., per-period min/mean) and the date range required. 2. In the scripts, trace where each modeling column comes from: is it derived from the provided dataset files, or fetched/invented externally? 3. Check whether the aggregation and filtering literally match the task wording (correct reduction function, correct period grouping, correct start/end). 4. Read the answer/comments for hedging language such as "proxy", "alternative interpretation", "actual variable not available", or a fallback second model — this signals the reported number is not from the requested regression.
Discriminator
A real violation is when a modeling variable is not derivable from the provided data and no equivalence is demonstrated, or the required aggregation is replaced by a different one. It is fine if the agent uses a documented column alias/rename within the provided dataset, or reconstructs the exact quantity (same units, frequency, and reduction) from raw provided columns.
Consequence
The saved statistic comes from a different model on different inputs, so the value in the output file fails the exact/tolerance comparison against ground truth even though the code runs cleanly.
id b74b966f23d1 · mined from da-code dacode-data-sa-043
raw text (what the judge reads)
### Substituting proxy/external data for the variables the task actually names
- **Applies when**: `task` -- the task asks for a statistic from a model built on specific named quantities from a provided dataset, and the script instead pulls data from an external source or swaps in a "close enough" stand-in series.
- **Pattern**: The agent cannot locate (or does not look for) the required variables in the supplied data, so it downloads unrelated series or redefines a variable as a proxy (e.g., a rate instead of the named quantity, a different index as the "portfolio"), sometimes with a different frequency, aggregation, or date coverage, then reports the statistic from that substitute model as the answer while acknowledging the substitution in the write-up.
- **Detection procedure**: 1. From the task text, list the exact dependent and independent quantities, the aggregation (e.g., per-period min/mean) and the date range required. 2. In the scripts, trace where each modeling column comes from: is it derived from the provided dataset files, or fetched/invented externally? 3. Check whether the aggregation and filtering literally match the task wording (correct reduction function, correct period grouping, correct start/end). 4. Read the answer/comments for hedging language such as "proxy", "alternative interpretation", "actual variable not available", or a fallback second model — this signals the reported number is not from the requested regression.
- **Discriminator**: A real violation is when a modeling variable is not derivable from the provided data and no equivalence is demonstrated, or the required aggregation is replaced by a different one. It is fine if the agent uses a documented column alias/rename within the provided dataset, or reconstructs the exact quantity (same units, frequency, and reduction) from raw provided columns.
- **Consequence**: The saved statistic comes from a different model on different inputs, so the value in the output file fails the exact/tolerance comparison against ground truth even though the code runs cleanly.
35Model selection and decision threshold not aligned with the task's stated objective on imbalanced classestaskda-code
Applies when
task -- the task states an asymmetric cost/priority (e.g., missed positives are far more expensive) or a specific evaluation metric for a rare-class classification, and the scripts must emit hard 0/1 labels.
Pattern
The attempt trains several off-the-shelf classifiers, compares them with generic metrics (accuracy, ROC-AUC, F1) on one arbitrary hold-out split, then calls predict() at the implicit 0.5 probability cutoff — never optimizing recall / the stated cost function, never tuning or even reporting the threshold, and never checking the resulting positive rate against the training base rate.
Detection procedure
  1. Read the task statement and write down the metric or cost trade-off it actually asks to optimize, plus the class balance seen in the training data.
  2. In the scripts, check how the final model and cutoff are chosen: is selection driven by that metric (e.g., threshold sweep / cost curve / recall-oriented CV), or by default predict() and unrelated scores?
  3. Compare the predicted positive count/rate in the answer with the training-set positive rate and with the recall implied on validation; flag if positives are predicted at or below base rate while the task penalizes false negatives.
  4. Check the selection evidence: single random split vs. repeated/stratified CV, and whether the chosen model is justified by numbers actually printed.
Discriminator
A genuine violation uses the default cutoff and off-target metrics with no sensitivity analysis; an acceptable attempt either explicitly tunes the threshold/class weights against the stated cost or metric and shows the trade-off, or documents that the default cutoff is already optimal under that metric.
Consequence
The saved label column has too few positives (low recall), so the graded score under the cost/recall-based check falls below the required threshold and the prediction file is marked wrong even though the file format is fine.
id 161f08488349 · mined from da-code dacode-ml-binary-013
raw text (what the judge reads)
### Model selection and decision threshold not aligned with the task's stated objective on imbalanced classes
- **Applies when**: `task` -- the task states an asymmetric cost/priority (e.g., missed positives are far more expensive) or a specific evaluation metric for a rare-class classification, and the scripts must emit hard 0/1 labels.
- **Pattern**: The attempt trains several off-the-shelf classifiers, compares them with generic metrics (accuracy, ROC-AUC, F1) on one arbitrary hold-out split, then calls `predict()` at the implicit 0.5 probability cutoff — never optimizing recall / the stated cost function, never tuning or even reporting the threshold, and never checking the resulting positive rate against the training base rate.
- **Detection procedure**:
  1. Read the task statement and write down the metric or cost trade-off it actually asks to optimize, plus the class balance seen in the training data.
  2. In the scripts, check how the final model and cutoff are chosen: is selection driven by that metric (e.g., threshold sweep / cost curve / recall-oriented CV), or by default `predict()` and unrelated scores?
  3. Compare the predicted positive count/rate in the answer with the training-set positive rate and with the recall implied on validation; flag if positives are predicted at or below base rate while the task penalizes false negatives.
  4. Check the selection evidence: single random split vs. repeated/stratified CV, and whether the chosen model is justified by numbers actually printed.
- **Discriminator**: A genuine violation uses the default cutoff and off-target metrics with no sensitivity analysis; an acceptable attempt either explicitly tunes the threshold/class weights against the stated cost or metric and shows the trade-off, or documents that the default cutoff is already optimal under that metric.
- **Consequence**: The saved label column has too few positives (low recall), so the graded score under the cost/recall-based check falls below the required threshold and the prediction file is marked wrong even though the file format is fine.
36Analyzing only one partition of the data when the task asks for the whole populationtaskinfiagent-dabench
Applies when
task -- the question asks for a statistic or detection over "all" units (all countries, all users, all records) and the data directory contains multiple files/partitions or the script filters/loads a single subset.
Pattern
The script hard-codes a single input file (or one group/region/segment) and computes quartiles, thresholds, or aggregates on that partial sample, then reports the result as if it covered the full population — the reference quantities (Q1/Q3, means, ranks) are therefore derived from the wrong subset, even if some of the returned items happen to coincide with the expected ones.
Detection procedure
  1. Read the task statement and note the stated scope of the population ("for all …") and any required grouping or filtering.
  2. Read the script's data-loading lines: check whether it enumerates/concatenates every relevant file or subset, or loads exactly one and never verifies that this covers the requested scope.
  3. Compare the row/unit count printed or implied by the script against the count implied by the task scope (e.g., number of files in the directory, total known units); a much smaller count signals a partial load.
  4. Check that the threshold/statistic itself (not just the final list) was computed on the full-scope data, and that the answer's labels and format match the requested output exactly.
Discriminator
A real violation is loading/filtering a subset when the task's scope is broader and no justification or coverage check exists. It is fine if the task explicitly restricts scope to that subset, or if the script demonstrates (by listing directory contents or comparing counts) that the single file is the complete population.
Consequence
Quartiles/thresholds computed on a partial sample yield a different outlier/selection set than the ground truth built from the full population, so the graded answer list mismatches and scores 0.
id 5f28c3b3b22a · mined from infiagent-dabench dabench-254
raw text (what the judge reads)
### Analyzing only one partition of the data when the task asks for the whole population
- **Applies when**: `task` -- the question asks for a statistic or detection over "all" units (all countries, all users, all records) and the data directory contains multiple files/partitions or the script filters/loads a single subset.
- **Pattern**: The script hard-codes a single input file (or one group/region/segment) and computes quartiles, thresholds, or aggregates on that partial sample, then reports the result as if it covered the full population — the reference quantities (Q1/Q3, means, ranks) are therefore derived from the wrong subset, even if some of the returned items happen to coincide with the expected ones.
- **Detection procedure**:
  1. Read the task statement and note the stated scope of the population ("for all …") and any required grouping or filtering.
  2. Read the script's data-loading lines: check whether it enumerates/concatenates every relevant file or subset, or loads exactly one and never verifies that this covers the requested scope.
  3. Compare the row/unit count printed or implied by the script against the count implied by the task scope (e.g., number of files in the directory, total known units); a much smaller count signals a partial load.
  4. Check that the threshold/statistic itself (not just the final list) was computed on the full-scope data, and that the answer's labels and format match the requested output exactly.
- **Discriminator**: A real violation is loading/filtering a subset when the task's scope is broader and no justification or coverage check exists. It is fine if the task explicitly restricts scope to that subset, or if the script demonstrates (by listing directory contents or comparing counts) that the single file *is* the complete population.
- **Consequence**: Quartiles/thresholds computed on a partial sample yield a different outlier/selection set than the ground truth built from the full population, so the graded answer list mismatches and scores 0.
37Output label/schema fidelity not verified against the training target's exact categoriestaskda-code
Applies when
task -- the deliverable is a prediction file with a specified column name whose values are the original class labels (or a specified numeric format) taken from a categorical target in the training data.
Pattern
The attempt trains a model, encodes the target internally, then writes predictions using re-derived or re-formatted labels (title-cased, renamed, 0/1 codes, extra index column, added ID column, wrong row count/order) without ever asserting that the written values and file schema exactly reproduce the label strings and column layout requested; the report even describes classes with capitalization/wording that differs from the raw data.
Detection procedure
  1. From the task/README, note the exact required file name, column name(s), and the exact set of label strings as they appear in the training target column.
  2. In the scripts, trace the target from load → encoding → inverse mapping → to_csv: check whether an inverse transform back to the original strings is applied, whether index=False is used, and whether the header equals the requested name.
  3. Check for any explicit sanity check in the script or answer: row count equals the test set size, single expected column, and value_counts() keys character-for-character identical to the training label values.
  4. Read the answer's class-distribution/summary text: if the class names or column layout it reports differ in spelling, case, or count from the raw data, treat the output as unverified.
Discriminator
A real violation is the absence of an inverse mapping/format assertion, or evidence of relabeled/re-cased/coded values, extra index or ID columns, or a row count not matching the test rows; a look-alike that is fine is a script that writes original label strings (or the explicitly requested encoding) with index=False and prints a verification of shape, column name, and label set — even if the prose summary paraphrases class names loosely.
Consequence
The grader cannot match the expected column values and marks the required output file WRONG/MISSING, so the attempt scores 0 regardless of how good the underlying model is.
id 6039bfdfbd4a · mined from da-code dacode-ml-binary-009
raw text (what the judge reads)
### Output label/schema fidelity not verified against the training target's exact categories
- **Applies when**: `task` -- the deliverable is a prediction file with a specified column name whose values are the original class labels (or a specified numeric format) taken from a categorical target in the training data.
- **Pattern**: The attempt trains a model, encodes the target internally, then writes predictions using re-derived or re-formatted labels (title-cased, renamed, 0/1 codes, extra index column, added ID column, wrong row count/order) without ever asserting that the written values and file schema exactly reproduce the label strings and column layout requested; the report even describes classes with capitalization/wording that differs from the raw data.
- **Detection procedure**:
  1. From the task/README, note the exact required file name, column name(s), and the exact set of label strings as they appear in the training target column.
  2. In the scripts, trace the target from load → encoding → inverse mapping → `to_csv`: check whether an inverse transform back to the original strings is applied, whether `index=False` is used, and whether the header equals the requested name.
  3. Check for any explicit sanity check in the script or answer: row count equals the test set size, single expected column, and `value_counts()` keys character-for-character identical to the training label values.
  4. Read the answer's class-distribution/summary text: if the class names or column layout it reports differ in spelling, case, or count from the raw data, treat the output as unverified.
- **Discriminator**: A real violation is the absence of an inverse mapping/format assertion, or evidence of relabeled/re-cased/coded values, extra index or ID columns, or a row count not matching the test rows; a look-alike that is fine is a script that writes original label strings (or the explicitly requested encoding) with `index=False` and prints a verification of shape, column name, and label set — even if the prose summary paraphrases class names loosely.
- **Consequence**: The grader cannot match the expected column values and marks the required output file WRONG/MISSING, so the attempt scores 0 regardless of how good the underlying model is.
38Invented metric definition + entity list back-filled from the plot config instead of derived from the datataskda-code
Applies when
task -- the task asks for a chart/table summarizing a substantive quantity (e.g., "performance", "top N"), and a settings/config file supplies cosmetic details plus axis/category labels.
Pattern
The script treats the config's label list as the source of truth for which entities to include and invents an ad-hoc, unjustified formula for the plotted value (arbitrary weights, mixing unrelated counts), rather than computing the substantive quantity from the data and letting the ranking/selection fall out of it. Filters (e.g., the "specified period") are also guessed from a title string. No check is made that the computed ordering/selection reproduces the config's labels, and required auxiliary output artifacts implied by the task/settings are never written.
Detection procedure
  1. Read the task and the config: list every constraint (which quantity is requested, which subset/period, ordering, required output files/format).
  2. In the scripts, locate the formula for the plotted values and ask whether each term and weight is traceable to the task, README, or config — or was chosen by the agent.
  3. Check whether the set/order of plotted categories is computed from the data (e.g., sorted by the metric and truncated) or merely intersected with the config's label list; if the latter, check whether the agent verified that its metric reproduces those labels in that order.
  4. Check the answer/output for all requested artifacts and for a sanity check on the plotted numbers (plausible ranges, monotonic ordering, counts matching N).
Discriminator
Fine if the metric is a standard/derivable definition for the requested quantity and the config labels are used only as a consistency check that the independently computed ranking matches; a violation if the metric contains agent-chosen weights/terms, the plotted bars are visibly non-monotonic with respect to a "top/best" label ordering, or the category list could not have been reproduced from the data by the script's own logic.
Consequence
The chart's bar values (and any saved numeric/JSON companion outputs) differ from the reference, so all value/plot checks fail even though the figure looks well-formed, and missing required output files fail outright.
id 590ec2f02edf · mined from da-code dacode-plot-bar-006
raw text (what the judge reads)
### Invented metric definition + entity list back-filled from the plot config instead of derived from the data
- **Applies when**: `task` -- the task asks for a chart/table summarizing a substantive quantity (e.g., "performance", "top N"), and a settings/config file supplies cosmetic details plus axis/category labels.
- **Pattern**: The script treats the config's label list as the source of truth for *which* entities to include and invents an ad-hoc, unjustified formula for the plotted value (arbitrary weights, mixing unrelated counts), rather than computing the substantive quantity from the data and letting the ranking/selection fall out of it. Filters (e.g., the "specified period") are also guessed from a title string. No check is made that the computed ordering/selection reproduces the config's labels, and required auxiliary output artifacts implied by the task/settings are never written.
- **Detection procedure**:
  1. Read the task and the config: list every constraint (which quantity is requested, which subset/period, ordering, required output files/format).
  2. In the scripts, locate the formula for the plotted values and ask whether each term and weight is traceable to the task, README, or config — or was chosen by the agent.
  3. Check whether the set/order of plotted categories is computed from the data (e.g., sorted by the metric and truncated) or merely intersected with the config's label list; if the latter, check whether the agent verified that its metric reproduces those labels in that order.
  4. Check the answer/output for all requested artifacts and for a sanity check on the plotted numbers (plausible ranges, monotonic ordering, counts matching N).
- **Discriminator**: Fine if the metric is a standard/derivable definition for the requested quantity and the config labels are used only as a consistency check that the independently computed ranking matches; a violation if the metric contains agent-chosen weights/terms, the plotted bars are visibly non-monotonic with respect to a "top/best" label ordering, or the category list could not have been reproduced from the data by the script's own logic.
- **Consequence**: The chart's bar values (and any saved numeric/JSON companion outputs) differ from the reference, so all value/plot checks fail even though the figure looks well-formed, and missing required output files fail outright.
39Outlier/threshold counts reported without an independent cross-check of the flagging logictaskinfiagent-dabench
Applies when
task -- the task asks for a count of rows meeting a statistical threshold (z-score, IQR, quantile, etc.) computed on a single column, and the script relies on a library helper plus a boolean mask to both count and filter.
Pattern
The script computes the statistic with one code path (e.g. a library function with special NaN/ddof handling), builds a positional boolean/index mask, applies it to the dataframe, and reports the resulting count without ever verifying it against a simple, independent recomputation ((|x - mean| > kstd).sum(), or comparing min/max to mean ± kstd) or checking the column's dtype, NaN count, and units.
Detection procedure
  1. From the task, note the exact rule and threshold and the exact quantity requested (count of flagged rows).
  2. In the script, check whether the column is inspected first (dtype, non-null count, min/max/describe) and whether any non-numeric/missing values would be silently dropped or coerced, and whether the mask indices (positional vs. label) match the dataframe index used for filtering.
  3. Check whether the flag count is recomputed a second, independent way (manual mean/std formula, or comparing extreme values to the mean ± k·std bounds) and whether the two agree.
  4. Compare the reported count to the reported summary stats: if max and min lie inside mean ± k·std, the count must be 0; if the count is non-zero, the script should print actual flagged values whose distance from the mean exceeds the bound.
Discriminator
A fine attempt prints the column's summary statistics and the flagged values, and the flagged values are visibly beyond the stated threshold from the mean; a violation reports a count that is never reconciled with the printed distribution, or whose mask alignment/NaN handling (e.g. nan_policy='omit' shrinking the array, ddof differences, positional mask on a non-default index) could shift which rows are flagged.
Consequence
The reported integer count (and the size of the filtered dataframe) is off — often a large spurious count where the true answer is zero or vice versa — so the graded numeric answer fails outright.
id 04b18e867611 · mined from infiagent-dabench dabench-361
raw text (what the judge reads)
### Outlier/threshold counts reported without an independent cross-check of the flagging logic
- **Applies when**: `task` -- the task asks for a count of rows meeting a statistical threshold (z-score, IQR, quantile, etc.) computed on a single column, and the script relies on a library helper plus a boolean mask to both count and filter.
- **Pattern**: The script computes the statistic with one code path (e.g. a library function with special NaN/ddof handling), builds a positional boolean/index mask, applies it to the dataframe, and reports the resulting count without ever verifying it against a simple, independent recomputation (`(|x - mean| > k*std).sum()`, or comparing `min`/`max` to `mean ± k*std`) or checking the column's dtype, NaN count, and units.
- **Detection procedure**:
  1. From the task, note the exact rule and threshold and the exact quantity requested (count of flagged rows).
  2. In the script, check whether the column is inspected first (dtype, non-null count, min/max/describe) and whether any non-numeric/missing values would be silently dropped or coerced, and whether the mask indices (positional vs. label) match the dataframe index used for filtering.
  3. Check whether the flag count is recomputed a second, independent way (manual mean/std formula, or comparing extreme values to the mean ± k·std bounds) and whether the two agree.
  4. Compare the reported count to the reported summary stats: if `max` and `min` lie inside mean ± k·std, the count must be 0; if the count is non-zero, the script should print actual flagged values whose distance from the mean exceeds the bound.
- **Discriminator**: A fine attempt prints the column's summary statistics and the flagged values, and the flagged values are visibly beyond the stated threshold from the mean; a violation reports a count that is never reconciled with the printed distribution, or whose mask alignment/NaN handling (e.g. `nan_policy='omit'` shrinking the array, ddof differences, positional mask on a non-default index) could shift which rows are flagged.
- **Consequence**: The reported integer count (and the size of the filtered dataframe) is off — often a large spurious count where the true answer is zero or vice versa — so the graded numeric answer fails outright.
40Substituting assumed conventions for an explicitly referenced specification filetaskda-code
Applies when
task -- the prompt points to an auxiliary document/config (readme, notes, mapping file, data dictionary) that defines how values must be transformed, named, filtered, or reported.
Pattern
The scripts never open or parse the referenced file; instead the agent hardcodes a "standard"/guessed mapping or definition from prior knowledge, then computes and reports statistics using those invented labels (and often skips writing the specified output artifact).
Detection procedure
  1. Read the task and list every external resource it names and every constraint that resource is said to carry (label names, categories, filters, output file).
  2. Grep the scripts for any read of that resource (open, read_csv, read_text, path string); if absent, the transformation rules are unverified assumptions.
  3. Compare the literal strings/values used in the script's hardcoded dictionary or constants against what the task says should come from the file — no in-script evidence of agreement is a violation.
  4. Check that the answer is emitted in the required form/artifact (e.g., written to the named result file with the named keys), not only printed to stdout.
Discriminator
Fine if the script actually loads the referenced file (or quotes its verified contents after inspecting it) and derives the mapping from it; a violation if the mapping's provenance is only a comment like "these are the standard codes" or general domain knowledge.
Consequence
The numeric ratio may be right while the category label differs from the expected vocabulary (or the required output file is missing), so exact-match grading on the result file fails.
id a1b122fce003 · mined from da-code dacode-di-text-004
raw text (what the judge reads)
### Substituting assumed conventions for an explicitly referenced specification file
- **Applies when**: `task` -- the prompt points to an auxiliary document/config (readme, notes, mapping file, data dictionary) that defines how values must be transformed, named, filtered, or reported.
- **Pattern**: The scripts never open or parse the referenced file; instead the agent hardcodes a "standard"/guessed mapping or definition from prior knowledge, then computes and reports statistics using those invented labels (and often skips writing the specified output artifact).
- **Detection procedure**:
  1. Read the task and list every external resource it names and every constraint that resource is said to carry (label names, categories, filters, output file).
  2. Grep the scripts for any read of that resource (`open`, `read_csv`, `read_text`, path string); if absent, the transformation rules are unverified assumptions.
  3. Compare the literal strings/values used in the script's hardcoded dictionary or constants against what the task says should come from the file — no in-script evidence of agreement is a violation.
  4. Check that the answer is emitted in the required form/artifact (e.g., written to the named result file with the named keys), not only printed to stdout.
- **Discriminator**: Fine if the script actually loads the referenced file (or quotes its verified contents after inspecting it) and derives the mapping from it; a violation if the mapping's provenance is only a comment like "these are the standard codes" or general domain knowledge.
- **Consequence**: The numeric ratio may be right while the category label differs from the expected vocabulary (or the required output file is missing), so exact-match grading on the result file fails.
41Submission schema invented instead of copied from the provided templatetaskda-code
Applies when
task -- the task supplies a sample/template output file (or explicit format spec) and the scripts write a predictions file to disk.
Pattern
The scripts never read the template file; column names, column order, ID column casing, and row ordering are hardcoded from the agent's guess (e.g. self-chosen probability column labels), and the "verification" step only re-checks the file against the agent's own assumptions rather than against the template.
Detection procedure
  1. In the task/README, note that a template output file is provided and that the output must match it.
  2. Grep the scripts for a read of that template file; if absent, inspect how the output DataFrame's columns/order are constructed.
  3. Compare the agent's produced header and row order to the template's header and ID order (names, count, case, spelling, ordering of rows).
  4. Check whether the verification script asserts equality with the template header/ID sequence, or merely checks internal consistency (probabilities summing to 1, ID set membership).
Discriminator
Not a violation if the hardcoded header/order provably equals the template (agent printed the template header and matched it exactly, including ID ordering); it is a violation when the header labels or ordering are guessed from prose/memory and never diffed against the actual file.
Consequence
The grader cannot parse/align the submission and marks the file WRONG/MISSING regardless of predictive quality, yielding a 0 score.
id a1442e208ec3 · mined from da-code dacode-ml-competition-003
raw text (what the judge reads)
### Submission schema invented instead of copied from the provided template
- **Applies when**: `task` -- the task supplies a sample/template output file (or explicit format spec) and the scripts write a predictions file to disk.
- **Pattern**: The scripts never read the template file; column names, column order, ID column casing, and row ordering are hardcoded from the agent's guess (e.g. self-chosen probability column labels), and the "verification" step only re-checks the file against the agent's own assumptions rather than against the template.
- **Detection procedure**:
  1. In the task/README, note that a template output file is provided and that the output must match it.
  2. Grep the scripts for a read of that template file; if absent, inspect how the output DataFrame's columns/order are constructed.
  3. Compare the agent's produced header and row order to the template's header and ID order (names, count, case, spelling, ordering of rows).
  4. Check whether the verification script asserts equality with the template header/ID sequence, or merely checks internal consistency (probabilities summing to 1, ID set membership).
- **Discriminator**: Not a violation if the hardcoded header/order provably equals the template (agent printed the template header and matched it exactly, including ID ordering); it is a violation when the header labels or ordering are guessed from prose/memory and never diffed against the actual file.
- **Consequence**: The grader cannot parse/align the submission and marks the file WRONG/MISSING regardless of predictive quality, yielding a 0 score.
42Fabricating labels instead of using the provided supervised target and evaluation splittaskda-code
Applies when
task -- The task names specific input/output files (e.g., a training file with a target column and a test file to predict on) and the scripts must produce predictions keyed to those exact rows.
Pattern
The attempt declares the ground-truth target "unavailable", builds heuristic pseudo-labels from unrelated auxiliary tables, trains/validates on those self-generated labels (yielding an implausibly high self-consistent CV score), and writes predictions for a row set it invented rather than the rows of the specified test file.
Detection procedure
1) From the task statement, list the required input file(s), the required prediction file, its required column(s), and the required row set/ordering. 2) Grep the scripts for reads of those files and for the target column; note whether the label used in training comes from the data or from a hand-written rule function. 3) Compare the row count/keys of the produced output against the specified test file's row count/keys, and check the reported column names match exactly what was asked. 4) Check whether the reported validation score is measured against real labels or against the same rule that generated them.
Discriminator
A real violation is when the supervised target exists in the provided data (or the test row set is explicitly given) but the script never reads it and instead invents labels/rows; it is acceptable if the task is genuinely unsupervised, or if heuristics are used only as extra features alongside the true labels and the output still matches the specified test keys and columns.
Consequence
The output file has the wrong number of rows, wrong keys, and/or extra/misnamed columns, and its label distribution bears no relation to the true target, so the grader marks the result file wrong/missing despite a near-perfect self-reported CV accuracy.
id 0b8517bdd1a1 · mined from da-code dacode-ml-multi-003
raw text (what the judge reads)
### Fabricating labels instead of using the provided supervised target and evaluation split
- **Applies when**: `task` -- The task names specific input/output files (e.g., a training file with a target column and a test file to predict on) and the scripts must produce predictions keyed to those exact rows.
- **Pattern**: The attempt declares the ground-truth target "unavailable", builds heuristic pseudo-labels from unrelated auxiliary tables, trains/validates on those self-generated labels (yielding an implausibly high self-consistent CV score), and writes predictions for a row set it invented rather than the rows of the specified test file.
- **Detection procedure**: 1) From the task statement, list the required input file(s), the required prediction file, its required column(s), and the required row set/ordering. 2) Grep the scripts for reads of those files and for the target column; note whether the label used in training comes from the data or from a hand-written rule function. 3) Compare the row count/keys of the produced output against the specified test file's row count/keys, and check the reported column names match exactly what was asked. 4) Check whether the reported validation score is measured against real labels or against the same rule that generated them.
- **Discriminator**: A real violation is when the supervised target exists in the provided data (or the test row set is explicitly given) but the script never reads it and instead invents labels/rows; it is acceptable if the task is genuinely unsupervised, or if heuristics are used only as extra features alongside the true labels and the output still matches the specified test keys and columns.
- **Consequence**: The output file has the wrong number of rows, wrong keys, and/or extra/misnamed columns, and its label distribution bears no relation to the true target, so the grader marks the result file wrong/missing despite a near-perfect self-reported CV accuracy.
43Relabeling output axes to "match" a template instead of verifying value-to-cell alignmenttaskda-code
Applies when
task -- the task supplies a template/example output file that fixes the row and column labels of a computed table (e.g., group keys by period index), and the script builds its own table then renames or shifts the labels to make them look like the template.
Pattern
The script computes a grouped statistic with its own natural index (e.g., offsets starting at 0), then cosmetically renames columns/rows (adding an offset, reformatting dates) so the header string-matches the template, without establishing which underlying group each template cell actually represents. Verification scripts then only compare shapes/labels, or spot-check cells while openly speculating ("column 4 might mean offset 3"), and the ambiguity is never resolved before writing the file.
Detection procedure
  1. Read the task/template: note the exact expected row labels, column labels, ordering, and rounding/format of the deliverable.
  2. In the scripts, find every place where labels are renamed, offset, reindexed, or reformatted after the computation, and ask whether the mapping from computed group → template cell is derived from a stated definition or merely assumed to make headers line up.
  3. Check whether any script independently confirms the mapping (e.g., recomputing one cell from raw rows filtered by the definition the template implies, and matching a non-empty template value or a documented convention) rather than comparing an empty template's headers.
  4. Read the answer: look for statements that reveal unresolved ambiguity about what a column/row means, or claims of correctness backed only by "shape and headers match".
Discriminator
A fine attempt derives the label convention from the task/template (or from filled example values) and verifies at least one cell end-to-end against raw data under that convention; a violation performs an arbitrary shift/rename purely so the header text matches, and its own debug output shows two competing interpretations left unsettled.
Consequence
Every value lands one position off (or the whole grid is shifted/transposed), so the file has the right shape and labels but wrong contents, and the grader marks the expected output file WRONG.
id f25e75f5bcdb · mined from da-code dacode-dm-csv-044
raw text (what the judge reads)
### Relabeling output axes to "match" a template instead of verifying value-to-cell alignment
- **Applies when**: `task` -- the task supplies a template/example output file that fixes the row and column labels of a computed table (e.g., group keys by period index), and the script builds its own table then renames or shifts the labels to make them look like the template.
- **Pattern**: The script computes a grouped statistic with its own natural index (e.g., offsets starting at 0), then cosmetically renames columns/rows (adding an offset, reformatting dates) so the header string-matches the template, without establishing which underlying group each template cell actually represents. Verification scripts then only compare shapes/labels, or spot-check cells while openly speculating ("column 4 might mean offset 3"), and the ambiguity is never resolved before writing the file.
- **Detection procedure**:
  1. Read the task/template: note the exact expected row labels, column labels, ordering, and rounding/format of the deliverable.
  2. In the scripts, find every place where labels are renamed, offset, reindexed, or reformatted after the computation, and ask whether the mapping from computed group → template cell is derived from a stated definition or merely assumed to make headers line up.
  3. Check whether any script independently confirms the mapping (e.g., recomputing one cell from raw rows filtered by the definition the template implies, and matching a non-empty template value or a documented convention) rather than comparing an empty template's headers.
  4. Read the answer: look for statements that reveal unresolved ambiguity about what a column/row means, or claims of correctness backed only by "shape and headers match".
- **Discriminator**: A fine attempt derives the label convention from the task/template (or from filled example values) and verifies at least one cell end-to-end against raw data under that convention; a violation performs an arbitrary shift/rename purely so the header text matches, and its own debug output shows two competing interpretations left unsettled.
- **Consequence**: Every value lands one position off (or the whole grid is shifted/transposed), so the file has the right shape and labels but wrong contents, and the grader marks the expected output file WRONG.
44Optimizing/validating with a metric other than the one the task specifiestaskda-code
Applies when
task -- the task states an explicit scoring metric (e.g., a rank/agreement/ordinal or cost-weighted metric) and the scripts must choose a model, tune it, or decide when the answer is good enough.
Pattern
The scripts never implement or compute the stated metric; they select the "best" model and ensemble weights using a generic proxy (accuracy, weighted F1, plain log-loss) and treat the target as unordered classes, so the reported/validated numbers say nothing about the actual grading score. Often compounded by a contaminated check (the final ensemble is fit on all training rows and then "evaluated" on a subset of those same rows).
Detection procedure
  1. Read the task and write down the exact evaluation metric and any structure it implies (ordering of labels, weighting, thresholding, rounding).
  2. Grep the scripts for that metric's computation (or an equivalent implementation) and for the objects used in model comparison/selection; note which score drives the choice of final model.
  3. Check that the score used for selection is computed on data held out from every fitting step (scaler, feature transforms, base models, ensemble weights) — not on rows the final model was refit on.
  4. Check the answer/prediction distribution against what the metric rewards (e.g., under an ordinal-agreement metric, whether extreme classes and label ordering are handled at all rather than collapsed to majority classes).
Discriminator
A real violation is when no run-time estimate of the stated metric exists anywhere, or the metric-relevant structure is ignored, so the agent cannot tell a good submission from a bad one. It is not a violation if the agent computes the stated metric on a clean holdout/CV and additionally reports proxies, or if it can show the proxy is a monotone stand-in for the stated metric on that holdout.
Consequence
The submission looks well-formed and plausible, but the graded score (the stated metric) is far below what a metric-aware baseline achieves — predictions cluster on majority labels and the attempt fails the correctness check despite high reported accuracy/F1.
id db6067bf7360 · mined from da-code dacode-ml-competition-006
raw text (what the judge reads)
### Optimizing/validating with a metric other than the one the task specifies
- **Applies when**: `task` -- the task states an explicit scoring metric (e.g., a rank/agreement/ordinal or cost-weighted metric) and the scripts must choose a model, tune it, or decide when the answer is good enough.
- **Pattern**: The scripts never implement or compute the stated metric; they select the "best" model and ensemble weights using a generic proxy (accuracy, weighted F1, plain log-loss) and treat the target as unordered classes, so the reported/validated numbers say nothing about the actual grading score. Often compounded by a contaminated check (the final ensemble is fit on all training rows and then "evaluated" on a subset of those same rows).
- **Detection procedure**:
  1. Read the task and write down the exact evaluation metric and any structure it implies (ordering of labels, weighting, thresholding, rounding).
  2. Grep the scripts for that metric's computation (or an equivalent implementation) and for the objects used in model comparison/selection; note which score drives the choice of final model.
  3. Check that the score used for selection is computed on data held out from every fitting step (scaler, feature transforms, base models, ensemble weights) — not on rows the final model was refit on.
  4. Check the answer/prediction distribution against what the metric rewards (e.g., under an ordinal-agreement metric, whether extreme classes and label ordering are handled at all rather than collapsed to majority classes).
- **Discriminator**: A real violation is when no run-time estimate of the stated metric exists anywhere, or the metric-relevant structure is ignored, so the agent cannot tell a good submission from a bad one. It is *not* a violation if the agent computes the stated metric on a clean holdout/CV and additionally reports proxies, or if it can show the proxy is a monotone stand-in for the stated metric on that holdout.
- **Consequence**: The submission looks well-formed and plausible, but the graded score (the stated metric) is far below what a metric-aware baseline achieves — predictions cluster on majority labels and the attempt fails the correctness check despite high reported accuracy/F1.
45Missingness mask defined inconsistently with the raw data, without a partition sanity checktaskinfiagent-dabench
Applies when
task -- the task asks to split rows into "null" vs "non-null" (or otherwise filtered) groups on one column and compare an aggregate of another column between the groups.
Pattern
The attempt defines the group mask with an ad-hoc or lossy rule (e.g., isna() after a loader that coerced sentinels like "NA", "None", "", "-" to real values or vice versa, dropna() on the whole frame, filtering on a different column, or reading only a subset/sample of the file) and/or lets the aggregated column be non-numeric so rows are silently dropped or coerced — then reports group means without checking that the two groups reconstruct the full dataset.
Detection procedure
  1. From the task, note the exact grouping rule (null vs not-null on the specified column) and the column to be aggregated; note that no other filtering was authorized.
  2. In the scripts, locate the load step and the mask: check for na_values=/converters, whole-frame dropna(), astype/to_numeric(errors='coerce'), string-vs-NaN handling, and any row subsetting or sampling; confirm the aggregate is computed on the intended column for each mask.
  3. Require an explicit printed sanity check: len(group_A) + len(group_B) == len(df), per-group non-null counts of the aggregated column, and total row count matching the raw file; if absent, the attempt is unverified.
  4. Check the reported means are consistent with those counts (e.g., recompute the pooled mean from group means and sizes and compare to the overall column mean).
Discriminator
A real violation is when rows are added/lost or reclassified relative to the raw column's true missingness (counts don't partition the file, or sentinel strings are treated inconsistently). It is fine if the aggregated column itself has NaNs excluded by the aggregation within each group, provided the grouping mask still partitions all rows and this is stated/verified.
Consequence
Both group means shift by a few percent — plausible-looking numbers that fail exact-value checks (the t-test p-value can still look tiny, masking the error), so the answer is graded wrong.
id 91e8c2fe7699 · mined from infiagent-dabench dabench-297
raw text (what the judge reads)
### Missingness mask defined inconsistently with the raw data, without a partition sanity check
- **Applies when**: `task` -- the task asks to split rows into "null" vs "non-null" (or otherwise filtered) groups on one column and compare an aggregate of another column between the groups.
- **Pattern**: The attempt defines the group mask with an ad-hoc or lossy rule (e.g., `isna()` after a loader that coerced sentinels like `"NA"`, `"None"`, `""`, `"-"` to real values or vice versa, `dropna()` on the whole frame, filtering on a different column, or reading only a subset/sample of the file) and/or lets the aggregated column be non-numeric so rows are silently dropped or coerced — then reports group means without checking that the two groups reconstruct the full dataset.
- **Detection procedure**:
  1. From the task, note the exact grouping rule (null vs not-null on the specified column) and the column to be aggregated; note that no other filtering was authorized.
  2. In the scripts, locate the load step and the mask: check for `na_values=`/converters, whole-frame `dropna()`, `astype`/`to_numeric(errors='coerce')`, string-vs-NaN handling, and any row subsetting or sampling; confirm the aggregate is computed on the intended column for each mask.
  3. Require an explicit printed sanity check: `len(group_A) + len(group_B) == len(df)`, per-group non-null counts of the aggregated column, and total row count matching the raw file; if absent, the attempt is unverified.
  4. Check the reported means are consistent with those counts (e.g., recompute the pooled mean from group means and sizes and compare to the overall column mean).
- **Discriminator**: A real violation is when rows are added/lost or reclassified relative to the raw column's true missingness (counts don't partition the file, or sentinel strings are treated inconsistently). It is fine if the aggregated column itself has NaNs excluded by the aggregation *within* each group, provided the grouping mask still partitions all rows and this is stated/verified.
- **Consequence**: Both group means shift by a few percent — plausible-looking numbers that fail exact-value checks (the t-test p-value can still look tiny, masking the error), so the answer is graded wrong.
46Missing or unsaved required output artifactstaskda-code
Applies when
task -- the task (or a referenced guidance/spec file) asks for computed results plus saved outputs, and the script's only persisted product is one figure or console prints while the numeric results live in the chat answer.
Pattern
The agent computes the requested quantities but never serializes them to the expected files/paths/formats (e.g., no saved array/table of the tallied values, no structured record of the plotted data), and instead reports them only as prose; it also does not re-read the spec file to enumerate the full list of deliverables.
Detection procedure
  1. From the task statement and any referenced guidance/README file, list every deliverable: each output file name, its format, its location, and the values it must contain.
  2. Grep the script for every write/save call (savefig, to_csv, np.save, json.dump, to_json, etc.) and record the exact paths produced.
  3. Compare the two lists; flag if any required deliverable has no corresponding write, is written to a different directory/name/extension, or if a required numeric result exists only in print/the chat answer.
  4. Check the saved figure/data actually encodes the requested numbers (counts, labels, ordering) rather than a stylistic approximation.
Discriminator
A real violation is a deliverable that is never written, or written under a name/format the spec did not ask for; it is not a violation if the file exists with the right content and only cosmetic details (title text, dpi, autopct) differ, or if the harness itself relocates correctly named outputs.
Consequence
The grader checks each expected artifact independently and marks every missing/mismatched file WRONG/MISSING, so the attempt scores 0 even when the reported numbers are plausible.
id 767c30e1213e · mined from da-code dacode-plot-pie-005
raw text (what the judge reads)
### Missing or unsaved required output artifacts
- **Applies when**: `task` -- the task (or a referenced guidance/spec file) asks for computed results plus saved outputs, and the script's only persisted product is one figure or console prints while the numeric results live in the chat answer.
- **Pattern**: The agent computes the requested quantities but never serializes them to the expected files/paths/formats (e.g., no saved array/table of the tallied values, no structured record of the plotted data), and instead reports them only as prose; it also does not re-read the spec file to enumerate the full list of deliverables.
- **Detection procedure**:
  1. From the task statement and any referenced guidance/README file, list every deliverable: each output file name, its format, its location, and the values it must contain.
  2. Grep the script for every write/save call (`savefig`, `to_csv`, `np.save`, `json.dump`, `to_json`, etc.) and record the exact paths produced.
  3. Compare the two lists; flag if any required deliverable has no corresponding write, is written to a different directory/name/extension, or if a required numeric result exists only in `print`/the chat answer.
  4. Check the saved figure/data actually encodes the requested numbers (counts, labels, ordering) rather than a stylistic approximation.
- **Discriminator**: A real violation is a deliverable that is never written, or written under a name/format the spec did not ask for; it is *not* a violation if the file exists with the right content and only cosmetic details (title text, dpi, autopct) differ, or if the harness itself relocates correctly named outputs.
- **Consequence**: The grader checks each expected artifact independently and marks every missing/mismatched file WRONG/MISSING, so the attempt scores 0 even when the reported numbers are plausible.
47Output schema not copied from the provided sample/format templatetaskda-code
Applies when
task -- the task supplies an example output file (or explicit column/row spec) and asks that results be saved in that exact format.
Pattern
The attempt invents its own header names, column count, or row labels (often adding stray/empty columns or renaming fields) instead of reading the sample file and reproducing its schema, so the content may be right while the file fails automated comparison; frequently the values are also hard-coded/recalled rather than produced by an aggregation over the data, with no script retained to show how they were derived.
Detection procedure
  1. In the task statement, locate the referenced sample/format artifact and note its exact headers, column order, and row-label wording.
  2. In the scripts, check that the sample file is actually loaded or its header literally reproduced, and that the reported values come from a computed aggregation over the full dataset (grouping/summing per entity) rather than typed-in constants.
  3. Diff the submitted file's header and row labels against the sample: same number of columns, same names/spelling/case, no extra or empty trailing fields, expected number of rows.
  4. Sanity-check one value by re-deriving it from the data (e.g., the max of the aggregated quantity) to confirm the ranking logic, not memory, produced it.
Discriminator
A real violation is a structural/label deviation from the given template (extra column, renamed header, invented row wording) or values with no code path producing them; a look-alike that is fine is a file whose schema matches the template exactly and differs only in benign ways the spec leaves open (row order when unspecified, quoting/whitespace).
Consequence
The grader's file/field comparison marks the expected output file WRONG/MISSING and scores 0, even if the underlying leaders were conceptually right.
id 0b75529b4da3 · mined from da-code dacode-dm-csv-015
raw text (what the judge reads)
### Output schema not copied from the provided sample/format template
- **Applies when**: `task` -- the task supplies an example output file (or explicit column/row spec) and asks that results be saved in that exact format.
- **Pattern**: The attempt invents its own header names, column count, or row labels (often adding stray/empty columns or renaming fields) instead of reading the sample file and reproducing its schema, so the content may be right while the file fails automated comparison; frequently the values are also hard-coded/recalled rather than produced by an aggregation over the data, with no script retained to show how they were derived.
- **Detection procedure**:
  1. In the task statement, locate the referenced sample/format artifact and note its exact headers, column order, and row-label wording.
  2. In the scripts, check that the sample file is actually loaded or its header literally reproduced, and that the reported values come from a computed aggregation over the full dataset (grouping/summing per entity) rather than typed-in constants.
  3. Diff the submitted file's header and row labels against the sample: same number of columns, same names/spelling/case, no extra or empty trailing fields, expected number of rows.
  4. Sanity-check one value by re-deriving it from the data (e.g., the max of the aggregated quantity) to confirm the ranking logic, not memory, produced it.
- **Discriminator**: A real violation is a structural/label deviation from the given template (extra column, renamed header, invented row wording) or values with no code path producing them; a look-alike that is fine is a file whose schema matches the template exactly and differs only in benign ways the spec leaves open (row order when unspecified, quoting/whitespace).
- **Consequence**: The grader's file/field comparison marks the expected output file WRONG/MISSING and scores 0, even if the underlying leaders were conceptually right.
48Unvalidated raw values in a knowingly "dirty" dataset before imputation/modelingtaskinfiagent-dabench
Applies when
task -- the script fits a model or computes a statistic on numeric columns from a source that is flagged (by filename, docs, or dtype) as containing errors, and the required preprocessing is only stated at a high level (e.g., "impute missing with the mean").
Pattern
The attempt takes the columns at face value: it prints dtypes/describe output but never checks whether values are actually parseable numerics within physically plausible ranges (non-numeric strings, sentinel codes, sign errors, unit/magnitude mismatches, duplicated or blank rows). Corrupt entries silently either become NaN-and-mean-imputed or remain as extreme outliers, so the fitted model and the reported error metric are driven by garbage rows.
Detection procedure
  1. Read the task for signals that the input is noisy/erroneous, and note which columns feed the model and which is the target.
  2. In the script, look for any explicit validation or cleaning step on those columns: coercion to numeric with inspection of what failed, range/plausibility checks, outlier or duplicate handling, verification that the imputed mean is computed on clean values.
  3. Check whether the script sanity-checks the final metric against the target's own scale (e.g., compare RMSE to the target's standard deviation or observed range) and reacts if it is implausibly large.
  4. If neither cleaning nor a scale sanity-check exists, and the reported metric implies typical errors comparable to or larger than the target's spread, flag the attempt.
Discriminator
A real violation is when no inspection of value validity occurred at all, or the reported error is out of line with the target's natural variability and is accepted without investigation. It is fine if the script explicitly verified the columns are clean (ranges, dtype coercion with zero failures) and the metric is consistent with the target's scale — even if no rows were removed.
Consequence
The regression is fit on contaminated values, inflating the reported error by roughly an order of magnitude relative to the reference value, so the single numeric answer fails the grader's exact-match check.
id 8008c46c5afc · mined from infiagent-dabench dabench-432
raw text (what the judge reads)
### Unvalidated raw values in a knowingly "dirty" dataset before imputation/modeling
- **Applies when**: `task` -- the script fits a model or computes a statistic on numeric columns from a source that is flagged (by filename, docs, or dtype) as containing errors, and the required preprocessing is only stated at a high level (e.g., "impute missing with the mean").
- **Pattern**: The attempt takes the columns at face value: it prints dtypes/describe output but never checks whether values are actually parseable numerics within physically plausible ranges (non-numeric strings, sentinel codes, sign errors, unit/magnitude mismatches, duplicated or blank rows). Corrupt entries silently either become NaN-and-mean-imputed or remain as extreme outliers, so the fitted model and the reported error metric are driven by garbage rows.
- **Detection procedure**:
  1. Read the task for signals that the input is noisy/erroneous, and note which columns feed the model and which is the target.
  2. In the script, look for any explicit validation or cleaning step on those columns: coercion to numeric with inspection of what failed, range/plausibility checks, outlier or duplicate handling, verification that the imputed mean is computed on clean values.
  3. Check whether the script sanity-checks the final metric against the target's own scale (e.g., compare RMSE to the target's standard deviation or observed range) and reacts if it is implausibly large.
  4. If neither cleaning nor a scale sanity-check exists, and the reported metric implies typical errors comparable to or larger than the target's spread, flag the attempt.
- **Discriminator**: A real violation is when no inspection of value validity occurred at all, or the reported error is out of line with the target's natural variability and is accepted without investigation. It is fine if the script explicitly verified the columns are clean (ranges, dtype coercion with zero failures) and the metric is consistent with the target's scale — even if no rows were removed.
- **Consequence**: The regression is fit on contaminated values, inflating the reported error by roughly an order of magnitude relative to the reference value, so the single numeric answer fails the grader's exact-match check.
49Unverified reordering of rows before a sequence-dependent computationtaskinfiagent-dabench
Applies when
task -- the requested quantity depends on row order (differences, lags, percent changes, cumulative sums, rolling stats) and the script explicitly sorts, reverses, or otherwise re-indexes the table before computing it.
Pattern
The agent asserts an ordering (e.g., "the file is newest-first, so reverse it") in a comment and applies the transformation without ever parsing the ordering key, printing the first/last key values, or checking monotonicity — so the sequence may be flipped relative to the intended one, silently inverting signs and shifting the result.
Detection procedure
  1. Read the task and note that the requested statistic is defined relative to a "previous"/"next" element, i.e. it is order-sensitive and sign-sensitive.
  2. In the scripts, locate any reversal/sort/reset_index step and check whether it is justified by evidence produced in the same run: is the ordering column converted to a proper dtype (e.g., datetime) and is its monotonic direction verified or sorted on explicitly?
  3. Check whether any post-hoc sanity check on the ordering-sensitive output exists (e.g., comparing the sign/magnitude of the statistic computed both as-is and reordered, or spot-checking one difference against two identified rows by their key values).
  4. Compare the reported answer to what would result under the opposite ordering; if the two differ mainly by sign, the answer is unvalidated.
Discriminator
Fine if the script sorts by the parsed ordering key (sort_values on a datetime/index) or prints head/tail of that key and states the observed direction; a violation if the direction is only assumed from a glance or a comment, or if raw row position is used as the order without any key check. Merely reversing data is not itself wrong — the absence of verification is.
Consequence
The lag is taken in the wrong direction, so the mean flips sign (and dispersion shifts slightly), and the graded values fail even though the formula and rounding look correct.
id 992d72ba6673 · mined from infiagent-dabench dabench-75
raw text (what the judge reads)
### Unverified reordering of rows before a sequence-dependent computation
- **Applies when**: `task` -- the requested quantity depends on row order (differences, lags, percent changes, cumulative sums, rolling stats) and the script explicitly sorts, reverses, or otherwise re-indexes the table before computing it.
- **Pattern**: The agent asserts an ordering (e.g., "the file is newest-first, so reverse it") in a comment and applies the transformation without ever parsing the ordering key, printing the first/last key values, or checking monotonicity — so the sequence may be flipped relative to the intended one, silently inverting signs and shifting the result.
- **Detection procedure**:
  1. Read the task and note that the requested statistic is defined relative to a "previous"/"next" element, i.e. it is order-sensitive and sign-sensitive.
  2. In the scripts, locate any reversal/sort/reset_index step and check whether it is justified by evidence produced in the same run: is the ordering column converted to a proper dtype (e.g., datetime) and is its monotonic direction verified or sorted on explicitly?
  3. Check whether any post-hoc sanity check on the ordering-sensitive output exists (e.g., comparing the sign/magnitude of the statistic computed both as-is and reordered, or spot-checking one difference against two identified rows by their key values).
  4. Compare the reported answer to what would result under the opposite ordering; if the two differ mainly by sign, the answer is unvalidated.
- **Discriminator**: Fine if the script sorts by the parsed ordering key (`sort_values` on a datetime/index) or prints head/tail of that key and states the observed direction; a violation if the direction is only assumed from a glance or a comment, or if raw row position is used as the order without any key check. Merely reversing data is not itself wrong — the absence of verification is.
- **Consequence**: The lag is taken in the wrong direction, so the mean flips sign (and dispersion shifts slightly), and the graded values fail even though the formula and rounding look correct.
50Output schema invented instead of read from the task's referenced specificationtaskda-code
Applies when
task -- The task points to an auxiliary instructions file (e.g., a tips/spec/README note) that defines the required test procedure and the exact fields/values of a result file the agent must write.
Pattern
The script never opens or echoes the referenced spec file; the agent guesses the column names, ordering, value wording (e.g., decision phrasing, test-type label), rounding, and number of fields, then declares the output "in the required format" without any comparison to the stated format.
Detection procedure
  1. Read the task and list every referenced instruction artifact and every explicitly constrained property of the deliverable (file name, one row vs. many, field set, allowed value strings, precision).
  2. Scan the scripts for any read/print of that artifact and for a literal mapping from its stated fields to the written columns; note whether the header/values in the code are copied from the spec or invented.
  3. Check the agent's answer for evidence it quoted the spec's required fields/values; absence of such a quote, or extra/renamed fields (e.g., an added statistic column, custom decision wording), is the flag.
  4. Also verify the chosen test and its preconditions match what the spec prescribes rather than the agent's own judgement (including how degenerate groups, e.g., a group with almost no observations, are handled).
Discriminator
Fine if the script demonstrably reads or verbatim reproduces the spec's field names/values and any deviation is justified; a violation is when the format is asserted from memory or from the prose of the main prompt only, with no trace of the spec content anywhere in scripts or output.
Consequence
The result file exists and looks plausible, but exact-match/field-wise grading of the deliverable fails (wrong or extra columns, non-matching decision strings or precision), scoring 0 despite a statistically defensible analysis.
id e6cd929607a6 · mined from da-code dacode-data-sa-004
raw text (what the judge reads)
### Output schema invented instead of read from the task's referenced specification
- **Applies when**: `task` -- The task points to an auxiliary instructions file (e.g., a tips/spec/README note) that defines the required test procedure and the exact fields/values of a result file the agent must write.
- **Pattern**: The script never opens or echoes the referenced spec file; the agent guesses the column names, ordering, value wording (e.g., decision phrasing, test-type label), rounding, and number of fields, then declares the output "in the required format" without any comparison to the stated format.
- **Detection procedure**:
  1. Read the task and list every referenced instruction artifact and every explicitly constrained property of the deliverable (file name, one row vs. many, field set, allowed value strings, precision).
  2. Scan the scripts for any read/print of that artifact and for a literal mapping from its stated fields to the written columns; note whether the header/values in the code are copied from the spec or invented.
  3. Check the agent's answer for evidence it quoted the spec's required fields/values; absence of such a quote, or extra/renamed fields (e.g., an added statistic column, custom decision wording), is the flag.
  4. Also verify the chosen test and its preconditions match what the spec prescribes rather than the agent's own judgement (including how degenerate groups, e.g., a group with almost no observations, are handled).
- **Discriminator**: Fine if the script demonstrably reads or verbatim reproduces the spec's field names/values and any deviation is justified; a violation is when the format is asserted from memory or from the prose of the main prompt only, with no trace of the spec content anywhere in scripts or output.
- **Consequence**: The result file exists and looks plausible, but exact-match/field-wise grading of the deliverable fails (wrong or extra columns, non-matching decision strings or precision), scoring 0 despite a statistically defensible analysis.
51Normalization method chosen so the reported statistic is trivially degeneratetaskinfiagent-dabench
Applies when
task -- The task asks to "normalize"/"scale" columns and then report summary statistics (e.g., means) of those transformed columns.
Pattern
The agent applies z-score standardization (subtract mean, divide by std), which forces every reported mean to 0 (or -0.0), instead of a scaling that preserves informative means (e.g., min-max to [0,1]); it then reports these degenerate zeros without questioning that the task would not ask for values that are known in advance.
Detection procedure
  1. Read the task to see which columns are to be scaled and which statistic must be reported afterwards.
  2. Inspect the script's scaling step: check whether the transform makes the requested statistic constant by construction (standardization → mean 0; centering → mean 0).
  3. Check the answer: if the reported values for all scaled columns are 0.0/-0.0 (or otherwise identical trivial constants) while unscaled columns carry real information, flag it.
  4. Confirm the answer format expectation (four-decimal distinct floats per column) is inconsistent with an all-zero result, and require the alternative scaling (min-max) or an explicit justification.
Discriminator
A real violation is when the requested statistic is mathematically fixed by the chosen transform, rendering the report uninformative; it is fine if standardization is explicitly named in the task, or if the reported statistic (e.g., std, min, max, or a downstream model score) is still informative under standardization.
Consequence
All scaled columns' reported means are 0.0 instead of the expected in-range values, so every check on those columns fails while unscaled/binary columns coincidentally pass.
id 934d706c6c1c · mined from infiagent-dabench dabench-28
raw text (what the judge reads)
### Normalization method chosen so the reported statistic is trivially degenerate
- **Applies when**: `task` -- The task asks to "normalize"/"scale" columns and then report summary statistics (e.g., means) of those transformed columns.
- **Pattern**: The agent applies z-score standardization (subtract mean, divide by std), which forces every reported mean to 0 (or -0.0), instead of a scaling that preserves informative means (e.g., min-max to [0,1]); it then reports these degenerate zeros without questioning that the task would not ask for values that are known in advance.
- **Detection procedure**:
  1. Read the task to see which columns are to be scaled and which statistic must be reported afterwards.
  2. Inspect the script's scaling step: check whether the transform makes the requested statistic constant by construction (standardization → mean 0; centering → mean 0).
  3. Check the answer: if the reported values for all scaled columns are 0.0/-0.0 (or otherwise identical trivial constants) while unscaled columns carry real information, flag it.
  4. Confirm the answer format expectation (four-decimal distinct floats per column) is inconsistent with an all-zero result, and require the alternative scaling (min-max) or an explicit justification.
- **Discriminator**: A real violation is when the requested statistic is mathematically fixed by the chosen transform, rendering the report uninformative; it is fine if standardization is explicitly named in the task, or if the reported statistic (e.g., std, min, max, or a downstream model score) is still informative under standardization.
- **Consequence**: All scaled columns' reported means are 0.0 instead of the expected in-range values, so every check on those columns fails while unscaled/binary columns coincidentally pass.
52Unverified row set / missing-value handling for a whole-dataset statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single summary statistic or test result computed over two or more columns of a supplied table, and the scripts (or lack of them) do not show how rows with missing/non-numeric entries were handled or how many rows entered the computation.
Pattern
The attempt computes the statistic after an implicit or convenience row reduction — e.g. dropping all rows with any NA anywhere in the table instead of only in the two relevant columns, coercing/filtering out non-numeric values, subsetting to a preview/sample of the file, or reading with a wrong delimiter/header so some rows are lost — and reports the resulting number without stating the sample size or checking it against the file's row count. It also reports derived values with less precision than requested (e.g. a truncated p-value) instead of the stated rounding.
Detection procedure
  1. From the task, note that the statistic is defined over the full table on the pair of named columns, and note the required output precision/format for each reported quantity.
  2. In the scripts, locate the load step and every row-removing operation (dropna() without subset=, boolean filters, head/nrows, dtype coercion, deduplication) and check whether any of them can remove rows that have valid values in both target columns.
  3. Check that the script prints the number of observations actually used (and non-null counts per column) and that this is compared with the raw file row count; if no script or no such print exists, treat the number as unverified.
  4. Compare the reported figures against the required rounding/format (e.g. two vs four decimals, no bare 0.0 for a p-value that should be given to four decimals).
Discriminator
A real violation is any row exclusion that is not required by the task and not justified/reported (or an unstated sample size), or a reported value whose precision differs from the spec. It is not a violation if the script restricts to pairwise-complete cases on exactly the two target columns, prints the resulting n, and matches it to the expected count — even if a few rows are legitimately dropped as genuinely missing.
Consequence
The coefficient shifts in the second decimal place (or the p-value string is malformed), so the graded numeric field mismatches the expected value even though the qualitative conclusion is right, and the attempt is scored wrong.
id c07951b6377d · mined from infiagent-dabench dabench-300
raw text (what the judge reads)
### Unverified row set / missing-value handling for a whole-dataset statistic
- **Applies when**: `task` -- the task asks for a single summary statistic or test result computed over two or more columns of a supplied table, and the scripts (or lack of them) do not show how rows with missing/non-numeric entries were handled or how many rows entered the computation.
- **Pattern**: The attempt computes the statistic after an implicit or convenience row reduction — e.g. dropping all rows with any NA anywhere in the table instead of only in the two relevant columns, coercing/filtering out non-numeric values, subsetting to a preview/sample of the file, or reading with a wrong delimiter/header so some rows are lost — and reports the resulting number without stating the sample size or checking it against the file's row count. It also reports derived values with less precision than requested (e.g. a truncated p-value) instead of the stated rounding.
- **Detection procedure**:
  1. From the task, note that the statistic is defined over the full table on the pair of named columns, and note the required output precision/format for each reported quantity.
  2. In the scripts, locate the load step and every row-removing operation (`dropna()` without `subset=`, boolean filters, `head`/`nrows`, dtype coercion, deduplication) and check whether any of them can remove rows that have valid values in both target columns.
  3. Check that the script prints the number of observations actually used (and non-null counts per column) and that this is compared with the raw file row count; if no script or no such print exists, treat the number as unverified.
  4. Compare the reported figures against the required rounding/format (e.g. two vs four decimals, no bare `0.0` for a p-value that should be given to four decimals).
- **Discriminator**: A real violation is any row exclusion that is not required by the task and not justified/reported (or an unstated sample size), or a reported value whose precision differs from the spec. It is *not* a violation if the script restricts to pairwise-complete cases on exactly the two target columns, prints the resulting n, and matches it to the expected count — even if a few rows are legitimately dropped as genuinely missing.
- **Consequence**: The coefficient shifts in the second decimal place (or the p-value string is malformed), so the graded numeric field mismatches the expected value even though the qualitative conclusion is right, and the attempt is scored wrong.
53No held-out validation or distribution sanity check on the predictionstaskda-code
Applies when
task -- the script trains a model on a labeled file and writes predictions for an unlabeled file, with no accuracy metric required in the deliverable.
Pattern
The attempt fits one model configuration on 100% of the labeled data, immediately predicts on the target file, and reports only self-described summary statistics of the predictions — never computing an error estimate on a held-out split, never comparing the predicted distribution (min/median/mean/max, skew) against the training label distribution, and never checking that the test rows were preprocessed identically (same feature construction, same category vocabulary, same handling of unseen/missing categories, same target scale/units).
Detection procedure
  1. Read the task: identify the required output file, its column, expected row count, and the plausible scale/units of the quantity being predicted.
  2. Read the script: check whether any split (train/validation, CV) is used to produce a numeric error estimate, and whether unseen categories/missing values in the target file are mapped in a way consistent with training (e.g., silently collapsed to an existing valid code rather than an explicit "unknown").
  3. Read the answer: compare the reported prediction statistics with the label statistics of the training data; flag if the mean/median/max are implausible or off by a large factor, or if a floor/clip was applied without justification.
  4. Confirm whether any validation number is reported at all; if the only evidence of quality is "predictions were saved", the attempt is unverified.
Discriminator
A real violation is an attempt with zero out-of-sample error estimate and no comparison of predicted vs. observed label distributions; it is fine if the agent reports a CV/holdout score (even from a simple baseline) and shows the prediction distribution matching the training label distribution, or explains any deliberate shift.
Consequence
The submitted file has the right shape and column name but systematically miscalibrated values (inflated mean/median, absurd extremes), so the grader's tolerance/error check on the predicted values fails while the agent claims success.
id bf4dfa0fddbd · mined from da-code dacode-ml-regression-014
raw text (what the judge reads)
### No held-out validation or distribution sanity check on the predictions
- **Applies when**: `task` -- the script trains a model on a labeled file and writes predictions for an unlabeled file, with no accuracy metric required in the deliverable.
- **Pattern**: The attempt fits one model configuration on 100% of the labeled data, immediately predicts on the target file, and reports only self-described summary statistics of the predictions — never computing an error estimate on a held-out split, never comparing the predicted distribution (min/median/mean/max, skew) against the training label distribution, and never checking that the test rows were preprocessed identically (same feature construction, same category vocabulary, same handling of unseen/missing categories, same target scale/units).
- **Detection procedure**:
  1. Read the task: identify the required output file, its column, expected row count, and the plausible scale/units of the quantity being predicted.
  2. Read the script: check whether any split (train/validation, CV) is used to produce a numeric error estimate, and whether unseen categories/missing values in the target file are mapped in a way consistent with training (e.g., silently collapsed to an existing valid code rather than an explicit "unknown").
  3. Read the answer: compare the reported prediction statistics with the label statistics of the training data; flag if the mean/median/max are implausible or off by a large factor, or if a floor/clip was applied without justification.
  4. Confirm whether any validation number is reported at all; if the only evidence of quality is "predictions were saved", the attempt is unverified.
- **Discriminator**: A real violation is an attempt with zero out-of-sample error estimate *and* no comparison of predicted vs. observed label distributions; it is fine if the agent reports a CV/holdout score (even from a simple baseline) and shows the prediction distribution matching the training label distribution, or explains any deliberate shift.
- **Consequence**: The submitted file has the right shape and column name but systematically miscalibrated values (inflated mean/median, absurd extremes), so the grader's tolerance/error check on the predicted values fails while the agent claims success.
54Prediction file row count/order doesn't match the evaluation inputtaskda-code
Applies when
task -- the task asks for a per-row output file (one prediction per input record) to be written for a held-out input set.
Pattern
The attempt produces an output file whose number of rows is far smaller (or larger) than the number of records in the provided input file — e.g. only a handful of rows from a truncated preview, a head()/sampled subset, a debug run, or rows dropped by filtering/dropna — and/or in an order that no longer aligns with the input rows, while still claiming completeness.
Detection procedure
  1. From the task and data files, determine the expected output length = number of records in the held-out input file (read its shape, don't assume), and the required column name/header.
  2. In the scripts, trace the object that is written out: check that it was built from the full input frame (no nrows=, .head(), .sample(), dropna/filtering, or partial-batch loop) and that row order is preserved (no sort, groupby, or reindex before writing).
  3. Count the rows in the produced answer/file and compare to the expected length; also check the header and allowed label/value set.
  4. Flag if the counts differ, if the count equals a suspicious small/round number, or if scripts are absent so the count cannot be traced to the full input.
Discriminator
A genuine violation is a length/alignment mismatch with the held-out input (or an untraceable pipeline); it is not a violation if the row count equals the input record count and the task itself legitimately requested an aggregated or filtered subset with that smaller size.
Consequence
The grader cannot align predictions to ground truth, so the submission is scored as wrong/missing regardless of model quality (or the score reflects only a tiny fraction of rows).
id 4e931ce66f66 · mined from da-code dacode-ml-multi-008
raw text (what the judge reads)
### Prediction file row count/order doesn't match the evaluation input
- **Applies when**: `task` -- the task asks for a per-row output file (one prediction per input record) to be written for a held-out input set.
- **Pattern**: The attempt produces an output file whose number of rows is far smaller (or larger) than the number of records in the provided input file — e.g. only a handful of rows from a truncated preview, a `head()`/sampled subset, a debug run, or rows dropped by filtering/`dropna` — and/or in an order that no longer aligns with the input rows, while still claiming completeness.
- **Detection procedure**:
  1. From the task and data files, determine the expected output length = number of records in the held-out input file (read its shape, don't assume), and the required column name/header.
  2. In the scripts, trace the object that is written out: check that it was built from the *full* input frame (no `nrows=`, `.head()`, `.sample()`, `dropna`/filtering, or partial-batch loop) and that row order is preserved (no `sort`, `groupby`, or reindex before writing).
  3. Count the rows in the produced answer/file and compare to the expected length; also check the header and allowed label/value set.
  4. Flag if the counts differ, if the count equals a suspicious small/round number, or if scripts are absent so the count cannot be traced to the full input.
- **Discriminator**: A genuine violation is a length/alignment mismatch with the held-out input (or an untraceable pipeline); it is *not* a violation if the row count equals the input record count and the task itself legitimately requested an aggregated or filtered subset with that smaller size.
- **Consequence**: The grader cannot align predictions to ground truth, so the submission is scored as wrong/missing regardless of model quality (or the score reflects only a tiny fraction of rows).
55Output artifacts written to an arbitrary directory instead of the task's working/data directorytaskinfiagent-dabench
Applies when
task -- the answer must include file paths to artifacts the agent creates (cleaned/transformed CSVs, model files, plots) that a grader will match or open.
Pattern
The agent saves outputs to its own home/current directory (or a temp path) rather than the directory the input data lives in / the directory implied by the task environment, and reports that path verbatim, so all path-bearing checks fail even though the computed numbers are right.
Detection procedure
  1. From the task statement, note where the input file(s) are located and any stated or conventional output location; treat the input's directory as the default expected output location when none is stated.
  2. In the scripts, find every write call (to_csv, savefig, dump) and record the literal or constructed output path; check whether it is derived from the input path or hardcoded to an unrelated directory.
  3. Compare the paths reported in the final answer to the input directory and to the paths actually written; flag any mismatch in directory, filename, or extension.
  4. Confirm the reported paths are absolute and refer to files that the script actually creates (name spelled identically, no later overwrite/rename).
Discriminator
A real violation is an output path whose directory differs from where the data was read (or from an explicitly required location), or a reported path that differs from what the script writes; it is fine if the task explicitly permits any path and the reported path exactly matches a file the script created in the same data directory.
Consequence
Path-containing checks are marked WRONG/MISSING even when the numeric fields match, so the item scores 0.
id 915e7b36967f · mined from infiagent-dabench dabench-743
raw text (what the judge reads)
### Output artifacts written to an arbitrary directory instead of the task's working/data directory
- **Applies when**: `task` -- the answer must include file paths to artifacts the agent creates (cleaned/transformed CSVs, model files, plots) that a grader will match or open.
- **Pattern**: The agent saves outputs to its own home/current directory (or a temp path) rather than the directory the input data lives in / the directory implied by the task environment, and reports that path verbatim, so all path-bearing checks fail even though the computed numbers are right.
- **Detection procedure**:
  1. From the task statement, note where the input file(s) are located and any stated or conventional output location; treat the input's directory as the default expected output location when none is stated.
  2. In the scripts, find every write call (`to_csv`, `savefig`, `dump`) and record the literal or constructed output path; check whether it is derived from the input path or hardcoded to an unrelated directory.
  3. Compare the paths reported in the final answer to the input directory and to the paths actually written; flag any mismatch in directory, filename, or extension.
  4. Confirm the reported paths are absolute and refer to files that the script actually creates (name spelled identically, no later overwrite/rename).
- **Discriminator**: A real violation is an output path whose directory differs from where the data was read (or from an explicitly required location), or a reported path that differs from what the script writes; it is fine if the task explicitly permits any path and the reported path exactly matches a file the script created in the same data directory.
- **Consequence**: Path-containing checks are marked WRONG/MISSING even when the numeric fields match, so the item scores 0.
56Incomplete compliance with an external plot/output specification (missing required artifacts)taskda-code
Applies when
task -- the task points to an external configuration/specification file (e.g., a YAML/JSON style guide) and/or names specific output files that the deliverable must consist of.
Pattern
The agent reads only a few keys of the spec file (or hardcodes assumptions about it), produces just the one visually obvious artifact (an image), and skips the other mandated outputs (serialized plot data / numeric arrays / config echo), or applies the spec's directives (titles, labels, order, colors, units, size) only partially. The final answer is prose about the finding rather than a checklist of produced files matching the requested format.
Detection procedure
  1. From the task text, list every required output artifact (file names/extensions) and every constraint the referenced spec file is said to govern.
  2. In the scripts, list every file actually written and every spec key actually consumed; compare against step 1, including whether the spec is fully enumerated/validated rather than accessed by a few hardcoded keys.
  3. Check the answer/report: does it confirm each required artifact exists with the right content type (e.g., underlying counts/values saved, not just an image), and in the requested location?
  4. Flag if any required artifact is never written, or any spec section is never read (e.g., keys present in the config that no code path uses).
Discriminator
A real violation is a required output file or spec directive with no corresponding code path. It is not a violation if the script writes all named artifacts and defensively reads the spec generically (e.g., iterating keys, with a printed dump verifying which directives were applied), even if some spec keys are simply unset/empty in the file.
Consequence
The grader checks each expected file independently; missing serialized data/config artifacts and partially styled figures score as WRONG/MISSING, so the run fails all checks even when the intermediate analytical finding is right.
id 371bb26355c6 · mined from da-code dacode-plot-pie-008
raw text (what the judge reads)
### Incomplete compliance with an external plot/output specification (missing required artifacts)
- **Applies when**: `task` -- the task points to an external configuration/specification file (e.g., a YAML/JSON style guide) and/or names specific output files that the deliverable must consist of.
- **Pattern**: The agent reads only a few keys of the spec file (or hardcodes assumptions about it), produces just the one visually obvious artifact (an image), and skips the other mandated outputs (serialized plot data / numeric arrays / config echo), or applies the spec's directives (titles, labels, order, colors, units, size) only partially. The final answer is prose about the finding rather than a checklist of produced files matching the requested format.
- **Detection procedure**:
  1. From the task text, list every required output artifact (file names/extensions) and every constraint the referenced spec file is said to govern.
  2. In the scripts, list every file actually written and every spec key actually consumed; compare against step 1, including whether the spec is fully enumerated/validated rather than accessed by a few hardcoded keys.
  3. Check the answer/report: does it confirm each required artifact exists with the right content type (e.g., underlying counts/values saved, not just an image), and in the requested location?
  4. Flag if any required artifact is never written, or any spec section is never read (e.g., keys present in the config that no code path uses).
- **Discriminator**: A real violation is a required output file or spec directive with no corresponding code path. It is *not* a violation if the script writes all named artifacts and defensively reads the spec generically (e.g., iterating keys, with a printed dump verifying which directives were applied), even if some spec keys are simply unset/empty in the file.
- **Consequence**: The grader checks each expected file independently; missing serialized data/config artifacts and partially styled figures score as WRONG/MISSING, so the run fails all checks even when the intermediate analytical finding is right.
57Dropping rows with missing values silently changes the evaluated sampletaskinfiagent-dabench
Applies when
task -- The task prescribes a fixed pipeline (given features, fixed split fraction and random seed, one metric) and the script must decide how to handle missing/invalid entries in the feature columns.
Pattern
The script calls a blanket row-drop (e.g. dropna()) on the feature/target frame before splitting, so the number of rows — and therefore the exact train/test partition produced by the seeded split — differs from the intended full-dataset pipeline; no imputation or justification is given, and the resulting row/split counts are never reconciled against the raw data.
Detection procedure
1. Read the task for any statement (or absence of a statement) about missing-value handling and note that the split is seed-fixed, so the row set fully determines the answer. 2. In the script, locate every operation that removes or filters rows (dropna, boolean masks, drop_duplicates, index slicing) applied before train_test_split. 3. Check whether the script reports the raw row count vs. the post-cleaning count and whether the loaded file/columns actually still contain missing values that require handling; check whether an imputation (mean/median/mode) alternative was considered. 4. Confirm the answer was computed on this reduced set with no sanity comparison to the full-data variant.
Discriminator
A real violation is unrequested row deletion in feature columns that shrinks the modeling sample (imputation would preserve it); it is fine to drop rows lacking the target (unlabeled rows cannot be scored), or to drop nothing because the columns are already complete — and it is fine if the script verifies the row count is unchanged after cleaning.
Consequence
The seeded split covers a different subset than the reference pipeline, so the reported accuracy is off by a few points (e.g. 0.76 vs. the expected 0.78) and the exact-value check fails.
id 7eb80ae7a4ed · mined from infiagent-dabench dabench-7
raw text (what the judge reads)
### Dropping rows with missing values silently changes the evaluated sample
- **Applies when**: `task` -- The task prescribes a fixed pipeline (given features, fixed split fraction and random seed, one metric) and the script must decide how to handle missing/invalid entries in the feature columns.
- **Pattern**: The script calls a blanket row-drop (e.g. `dropna()`) on the feature/target frame before splitting, so the number of rows — and therefore the exact train/test partition produced by the seeded split — differs from the intended full-dataset pipeline; no imputation or justification is given, and the resulting row/split counts are never reconciled against the raw data.
- **Detection procedure**: 1. Read the task for any statement (or absence of a statement) about missing-value handling and note that the split is seed-fixed, so the row set fully determines the answer. 2. In the script, locate every operation that removes or filters rows (`dropna`, boolean masks, `drop_duplicates`, index slicing) applied before `train_test_split`. 3. Check whether the script reports the raw row count vs. the post-cleaning count and whether the loaded file/columns actually still contain missing values that require handling; check whether an imputation (mean/median/mode) alternative was considered. 4. Confirm the answer was computed on this reduced set with no sanity comparison to the full-data variant.
- **Discriminator**: A real violation is unrequested row deletion in *feature* columns that shrinks the modeling sample (imputation would preserve it); it is fine to drop rows lacking the *target* (unlabeled rows cannot be scored), or to drop nothing because the columns are already complete — and it is fine if the script verifies the row count is unchanged after cleaning.
- **Consequence**: The seeded split covers a different subset than the reference pipeline, so the reported accuracy is off by a few points (e.g. 0.76 vs. the expected 0.78) and the exact-value check fails.
58Ignoring the provided output template when filling a required results filetaskda-code
Applies when
task -- The task supplies a pre-existing output file (template/skeleton) and asks that results be entered into it "in the same format", and the scripts generate that file's contents.
Pattern
The attempt never reads or inspects the supplied template; it invents its own schema — column names, row labels/category names, category definitions (e.g. self-chosen bin boundaries), row ordering, and extra catch-all rows — and writes a fresh file from scratch, then "verifies" the format only against its own assumptions.
Detection procedure
  1. In the task text, identify the exact output artifact required and any statement that its existing format must be followed.
  2. Search the scripts for any read/inspection of that template file (e.g. loading it, printing its rows/columns/label values) before writing.
  3. Check whether the labels, categories, ordering, and column headers written out are derived from the template or hard-coded from the agent's own reasoning; check whether extra/renamed rows (e.g. an "unknown/other" bucket) or self-invented category thresholds were introduced.
  4. Check whether the final answer is a narrative report rather than confirmation that the template file was filled in place with matching keys.
Discriminator
A real violation is inventing the key set/schema without ever reading the template; it is fine if the script loads the template, preserves its exact headers and row keys/order, and only populates the value cells — even if the agent additionally prints a human-readable summary.
Consequence
The required file fails key-by-key comparison against the expected file (mismatched row labels, extra rows, or values computed under different category definitions), so the grader reports the output file as WRONG/MISSING and scores 0.
id b9bcbaa11c91 · mined from da-code dacode-dm-csv-001
raw text (what the judge reads)
### Ignoring the provided output template when filling a required results file
- **Applies when**: `task` -- The task supplies a pre-existing output file (template/skeleton) and asks that results be entered into it "in the same format", and the scripts generate that file's contents.
- **Pattern**: The attempt never reads or inspects the supplied template; it invents its own schema — column names, row labels/category names, category definitions (e.g. self-chosen bin boundaries), row ordering, and extra catch-all rows — and writes a fresh file from scratch, then "verifies" the format only against its own assumptions.
- **Detection procedure**:
  1. In the task text, identify the exact output artifact required and any statement that its existing format must be followed.
  2. Search the scripts for any read/inspection of that template file (e.g. loading it, printing its rows/columns/label values) before writing.
  3. Check whether the labels, categories, ordering, and column headers written out are derived from the template or hard-coded from the agent's own reasoning; check whether extra/renamed rows (e.g. an "unknown/other" bucket) or self-invented category thresholds were introduced.
  4. Check whether the final answer is a narrative report rather than confirmation that the template file was filled in place with matching keys.
- **Discriminator**: A real violation is inventing the key set/schema without ever reading the template; it is fine if the script loads the template, preserves its exact headers and row keys/order, and only populates the value cells — even if the agent additionally prints a human-readable summary.
- **Consequence**: The required file fails key-by-key comparison against the expected file (mismatched row labels, extra rows, or values computed under different category definitions), so the grader reports the output file as WRONG/MISSING and scores 0.
59Self-reported performance measured on the training data instead of a held-out splittaskda-code
Applies when
task -- the agent fits a predictive model and must produce predictions for an unlabeled evaluation set that will be scored externally.
Pattern
The script trains on all labeled rows and then computes the quality metric with the same rows (or via predict on the training frame), reporting a high in-sample score as evidence the submission is good; no cross-validation or hold-out estimate, no comparison against a trivial baseline, and no check that the predicted positive rate / class balance is plausible relative to the training labels. Hyperparameters (depth, class weights, thresholds) are therefore chosen with no honest signal, and overfit or badly calibrated predictions go out unvalidated.
Detection procedure
  1. In the task, note that the scored artifact is predictions on data whose labels the agent never sees, so the only defensible quality claim is an out-of-sample estimate.
  2. In the scripts, locate the metric computation and check which rows are passed to it: are they the same rows used in fit, or a split/CV fold withheld from training?
  3. Check whether any baseline (majority class, simple logistic/default hyperparameters) is scored on that same held-out data for comparison, and whether the predicted label distribution on the evaluation set is compared to the training base rate.
  4. In the answer, see whether the headline score is labeled as training/in-sample; if it is, or if its provenance is unstated and step 2 shows no split, treat the quality claim as unsupported.
Discriminator
A real violation is a metric whose inputs are provably a subset of the fitted rows with no CV/hold-out anywhere; it is not a violation if the agent held out or cross-validated (even with a modest score), or if it reports an in-sample number in addition to a clearly identified out-of-sample estimate used for model selection.
Consequence
The reported score is inflated and uninformative (e.g., near-perfect in-sample vs. much weaker true performance), so the submitted prediction file can fall below the grader's accuracy/AUC threshold — and its class mix be far off the true base rate — while the answer confidently claims success.
id 23ae9ba0fd0b · mined from da-code dacode-ml-binary-016
raw text (what the judge reads)
### Self-reported performance measured on the training data instead of a held-out split
- **Applies when**: `task` -- the agent fits a predictive model and must produce predictions for an unlabeled evaluation set that will be scored externally.
- **Pattern**: The script trains on all labeled rows and then computes the quality metric with the same rows (or via `predict` on the training frame), reporting a high in-sample score as evidence the submission is good; no cross-validation or hold-out estimate, no comparison against a trivial baseline, and no check that the predicted positive rate / class balance is plausible relative to the training labels. Hyperparameters (depth, class weights, thresholds) are therefore chosen with no honest signal, and overfit or badly calibrated predictions go out unvalidated.
- **Detection procedure**:
  1. In the task, note that the scored artifact is predictions on data whose labels the agent never sees, so the only defensible quality claim is an out-of-sample estimate.
  2. In the scripts, locate the metric computation and check which rows are passed to it: are they the same rows used in `fit`, or a split/CV fold withheld from training?
  3. Check whether any baseline (majority class, simple logistic/default hyperparameters) is scored on that same held-out data for comparison, and whether the predicted label distribution on the evaluation set is compared to the training base rate.
  4. In the answer, see whether the headline score is labeled as training/in-sample; if it is, or if its provenance is unstated and step 2 shows no split, treat the quality claim as unsupported.
- **Discriminator**: A real violation is a metric whose inputs are provably a subset of the fitted rows with no CV/hold-out anywhere; it is *not* a violation if the agent held out or cross-validated (even with a modest score), or if it reports an in-sample number *in addition to* a clearly identified out-of-sample estimate used for model selection.
- **Consequence**: The reported score is inflated and uninformative (e.g., near-perfect in-sample vs. much weaker true performance), so the submitted prediction file can fall below the grader's accuracy/AUC threshold — and its class mix be far off the true base rate — while the answer confidently claims success.
60Ignoring the provided output template and the method implied by stated constraintstaskda-code
Applies when
task -- the task points to a sample/reference output file for the result format and/or states an execution constraint (e.g., fix a random seed) that implies a specific computational approach.
Pattern
The attempt never loads or inspects the sample output file and instead invents its own column names/rows, and it substitutes a closed-form/analytic shortcut for the randomized procedure implied by the seed requirement (so the seed is never actually used), then reports that value as the requested statistic.
Detection procedure
  1. Read the task for (a) a named format/template file and (b) constraints such as a random seed, rounding, units, or ordering.
  2. In the scripts, check that the template is actually read (or its exact header/row layout reproduced) and that every stated constraint appears in code — e.g., the seed is set and consumed by a stochastic step (resampling/permutation/bootstrap/model init).
  3. Compare the written file's schema (column names, count, order, row keys) and the value's provenance against the template and the requested quantity; flag if the schema is self-invented or the seed is decorative/unused.
  4. Check the reported answer restates the same schema and value as the saved file.
Discriminator
A real violation is inventing a schema or using a deterministic method when the task's constraints (template + seed) demand a specific format and a randomized computation; a look-alike that is fine is a deterministic method used when no randomized procedure is implied, or extra explanatory prose alongside a file that provably matches the template schema.
Consequence
The saved file fails exact-format/value comparison against the expected result (column names or the numeric p-value differ), scoring 0 even though the narrative conclusion may sound plausible.
id 5b9026b242e9 · mined from da-code dacode-data-sa-039
raw text (what the judge reads)
### Ignoring the provided output template and the method implied by stated constraints
- **Applies when**: `task` -- the task points to a sample/reference output file for the result format and/or states an execution constraint (e.g., fix a random seed) that implies a specific computational approach.
- **Pattern**: The attempt never loads or inspects the sample output file and instead invents its own column names/rows, and it substitutes a closed-form/analytic shortcut for the randomized procedure implied by the seed requirement (so the seed is never actually used), then reports that value as the requested statistic.
- **Detection procedure**:
  1. Read the task for (a) a named format/template file and (b) constraints such as a random seed, rounding, units, or ordering.
  2. In the scripts, check that the template is actually read (or its exact header/row layout reproduced) and that every stated constraint appears in code — e.g., the seed is set and consumed by a stochastic step (resampling/permutation/bootstrap/model init).
  3. Compare the written file's schema (column names, count, order, row keys) and the value's provenance against the template and the requested quantity; flag if the schema is self-invented or the seed is decorative/unused.
  4. Check the reported answer restates the same schema and value as the saved file.
- **Discriminator**: A real violation is inventing a schema or using a deterministic method when the task's constraints (template + seed) demand a specific format and a randomized computation; a look-alike that is fine is a deterministic method used when no randomized procedure is implied, or extra explanatory prose alongside a file that provably matches the template schema.
- **Consequence**: The saved file fails exact-format/value comparison against the expected result (column names or the numeric p-value differ), scoring 0 even though the narrative conclusion may sound plausible.
61Predictions not validated against the reference target distributiontaskda-code
Applies when
task -- the deliverable is a prediction file for a held-out split that will be scored against hidden ground truth, and the agent reports only self-computed validation metrics.
Pattern
The agent fits one default model, reports a mediocre in-house score (or none), and declares success from file-writing alone — never checking that predicted values plausibly reproduce the training target's distribution (mean, spread, min/max, skew) or that the row count/order/index matches the test input, and never trying to improve a clearly weak fit.
Detection procedure
1) Read the task for the required file name, column name, row count and any ordering/format constraint. 2) In the scripts, check whether the same feature construction and preprocessing are applied to train and test, and whether predictions are written in the test file's original row order with no extra index column. 3) Compare the reported prediction summary statistics (mean, min, max, variance) with the training target's statistics — flag if the predictions are visibly compressed, shifted, clipped, or far narrower/wider than the target. 4) Flag if the reported goodness-of-fit is weak and no alternative model, tuning, or error analysis was attempted before finalizing.
Discriminator
A genuine violation shows missing sanity checks and evidence of a distribution/shape mismatch or an unimproved weak model; it is fine if the agent explicitly compared prediction stats to the target, verified row count/order/column name, and justified the model choice with a comparison — shrinkage toward the mean alone is normal for regression and not by itself a violation.
Consequence
The submitted file scores below the grader's accuracy/error threshold (or fails format/shape validation), so the expected output is marked WRONG despite the script running without error.
id 3d32ead079fb · mined from da-code dacode-ml-regression-015
raw text (what the judge reads)
### Predictions not validated against the reference target distribution
- **Applies when**: `task` -- the deliverable is a prediction file for a held-out split that will be scored against hidden ground truth, and the agent reports only self-computed validation metrics.
- **Pattern**: The agent fits one default model, reports a mediocre in-house score (or none), and declares success from file-writing alone — never checking that predicted values plausibly reproduce the training target's distribution (mean, spread, min/max, skew) or that the row count/order/index matches the test input, and never trying to improve a clearly weak fit.
- **Detection procedure**: 1) Read the task for the required file name, column name, row count and any ordering/format constraint. 2) In the scripts, check whether the same feature construction and preprocessing are applied to train and test, and whether predictions are written in the test file's original row order with no extra index column. 3) Compare the reported prediction summary statistics (mean, min, max, variance) with the training target's statistics — flag if the predictions are visibly compressed, shifted, clipped, or far narrower/wider than the target. 4) Flag if the reported goodness-of-fit is weak and no alternative model, tuning, or error analysis was attempted before finalizing.
- **Discriminator**: A genuine violation shows missing sanity checks *and* evidence of a distribution/shape mismatch or an unimproved weak model; it is fine if the agent explicitly compared prediction stats to the target, verified row count/order/column name, and justified the model choice with a comparison — shrinkage toward the mean alone is normal for regression and not by itself a violation.
- **Consequence**: The submitted file scores below the grader's accuracy/error threshold (or fails format/shape validation), so the expected output is marked WRONG despite the script running without error.
62Trusting a single pre-existing column for a requested derived quantity, with a vacuous imputation steptaskda-code
Applies when
task -- the task asks to impute missing values and then rank/extremize a quantity that can either be read from one raw column or be recomputed from other columns in the data (e.g., a ratio of two available fields).
Pattern
The script picks the one column whose name resembles the requested quantity, cleans/imputes only that column, and reports its argmax/argmin — never checking that the stated preprocessing actually did anything (the column may have zero missing values, making the imputation a no-op) and never cross-validating the column against the quantity implied by its definition computed from the underlying fields.
Detection procedure
  1. From the task wording, write down the definition of the requested quantity and which raw fields it depends on; note the explicitly required preprocessing step.
  2. In the scripts, check whether preprocessing is applied to all fields the quantity depends on (or dataset-wide as stated) or only to one convenient column, and whether the script prints/asserts that the imputation changed at least one value.
  3. Check whether the script recomputes the quantity from its components and compares with the raw column (or at least verifies the raw column's units/scale and its extremes against known plausible ranges).
  4. Inspect the reported extremes: if the reported min/max come solely from the raw column with no consistency check, and the imputation was a no-op, flag the attempt.
Discriminator
A fine attempt either shows that the raw column is consistent with the definition (recomputed values match, or the task unambiguously names that single column) and reports how many values the imputation filled; a violation silently assumes the column is correct and complete, leaving the mandated imputation unexercised and the definition unverified.
Consequence
The reported extreme entities come from an inconsistent/uncorrected column, so the submitted min and/or max labels differ from the expected ones and the result file fails the equality check.
id 6918c1c8dc9a · mined from da-code dacode-di-text-001
raw text (what the judge reads)
### Trusting a single pre-existing column for a requested derived quantity, with a vacuous imputation step
- **Applies when**: `task` -- the task asks to impute missing values and then rank/extremize a quantity that can either be read from one raw column or be recomputed from other columns in the data (e.g., a ratio of two available fields).
- **Pattern**: The script picks the one column whose name resembles the requested quantity, cleans/imputes only that column, and reports its argmax/argmin — never checking that the stated preprocessing actually did anything (the column may have zero missing values, making the imputation a no-op) and never cross-validating the column against the quantity implied by its definition computed from the underlying fields.
- **Detection procedure**:
  1. From the task wording, write down the definition of the requested quantity and which raw fields it depends on; note the explicitly required preprocessing step.
  2. In the scripts, check whether preprocessing is applied to all fields the quantity depends on (or dataset-wide as stated) or only to one convenient column, and whether the script prints/asserts that the imputation changed at least one value.
  3. Check whether the script recomputes the quantity from its components and compares with the raw column (or at least verifies the raw column's units/scale and its extremes against known plausible ranges).
  4. Inspect the reported extremes: if the reported min/max come solely from the raw column with no consistency check, and the imputation was a no-op, flag the attempt.
- **Discriminator**: A fine attempt either shows that the raw column is consistent with the definition (recomputed values match, or the task unambiguously names that single column) and reports how many values the imputation filled; a violation silently assumes the column is correct and complete, leaving the mandated imputation unexercised and the definition unverified.
- **Consequence**: The reported extreme entities come from an inconsistent/uncorrected column, so the submitted min and/or max labels differ from the expected ones and the result file fails the equality check.
63Requested output file omits the computed quantities the task asked to includetaskda-code
Applies when
task -- the task asks to compute several per-entity quantities (scores, groups/segments, labels) and save "the results, including X and Y" to one named output file, and the script builds a full result table plus a trimmed export.
Pattern
The script computes all the intermediate per-entity metrics and derived scores in a working dataframe, but writes only a minimal subset of columns (e.g., ID + final label) to the required filename, dumping the complete table to a differently-named auxiliary file; the answer then claims completeness because the auxiliary file exists.
Detection procedure
1. From the task text, list every quantity that must appear in the named output file (each component metric, each derived score, the grouping/segment, the final level). 2. In the script, find the write call to that exact filename and enumerate the columns actually selected there. 3. Diff the two lists; also check whether a richer table was written to a different filename. 4. Check the answer's description of the file's contents against the task's required contents.
Discriminator
A real violation is when quantities the task explicitly named (or the per-entity scores the task's phrasing requires) are absent from the required file, even if present elsewhere on disk. It is fine if the file has extra columns, different column ordering/names, or if the task genuinely asked only for the final label and the extra table is a bonus.
Consequence
The grader compares the named file against a reference containing the full set of columns and marks it WRONG/MISSING (0 checks passed) despite the underlying computation being reasonable.
id aa0e6f15023f · mined from da-code dacode-dm-csv-052
raw text (what the judge reads)
### Requested output file omits the computed quantities the task asked to include
- **Applies when**: `task` -- the task asks to compute several per-entity quantities (scores, groups/segments, labels) and save "the results, including X and Y" to one named output file, and the script builds a full result table plus a trimmed export.
- **Pattern**: The script computes all the intermediate per-entity metrics and derived scores in a working dataframe, but writes only a minimal subset of columns (e.g., ID + final label) to the required filename, dumping the complete table to a differently-named auxiliary file; the answer then claims completeness because the auxiliary file exists.
- **Detection procedure**: 1. From the task text, list every quantity that must appear in the named output file (each component metric, each derived score, the grouping/segment, the final level). 2. In the script, find the write call to that exact filename and enumerate the columns actually selected there. 3. Diff the two lists; also check whether a richer table was written to a *different* filename. 4. Check the answer's description of the file's contents against the task's required contents.
- **Discriminator**: A real violation is when quantities the task explicitly named (or the per-entity scores the task's phrasing requires) are absent from the required file, even if present elsewhere on disk. It is fine if the file has extra columns, different column ordering/names, or if the task genuinely asked only for the final label and the extra table is a bonus.
- **Consequence**: The grader compares the named file against a reference containing the full set of columns and marks it WRONG/MISSING (0 checks passed) despite the underlying computation being reasonable.
64Ignoring a required subgroup filter (and unverified schema) when aggregating a statistictaskda-code
Applies when
task -- the task asks for a statistic over a specific subgroup/condition (e.g., one category, species, region, cohort) while the raw files contain multiple groups, and/or a template file defines the required output columns/values.
Pattern
The script reads each file and aggregates over every row, assuming the file already contains only the requested subgroup and that the column names it hard-codes exist as written; no filtering step, no schema inspection, and no comparison against the provided sample output file.
Detection procedure
  1. Read the task and note every restriction on rows (group/category, year, condition) and every stated output constraint (column names, ordering, rounding, format template).
  2. In the script, look for an explicit filtering operation (df[df[group_col] == ...]) on the group column and for any inspection/validation of the actual columns present in each input file; check whether the template/sample output file is ever read or its columns replicated.
  3. Check whether the script prints or asserts row counts / value ranges per group after loading, so a mismatch between the filtered subgroup size and the full file size would be visible.
  4. Compare the reported numbers to a rough domain sanity expectation (e.g., counts identical across heterogeneous files, or a ratio suspiciously near 1 when the two measured quantities are known to differ substantially) — an unexplained coincidence signals the filter or column mapping was wrong.
Discriminator
A real violation is when the raw file plausibly contains rows outside the requested subgroup (or differently named columns) and the script never filters/validates; it is fine if the script demonstrably verifies the subgroup (asserts unique group values, prints per-group counts, or the file is documented as already restricted) and reproduces the template's exact columns.
Consequence
The reported means/CIs are computed on the wrong row subset (or wrong columns) and the output file fails an exact/tolerance comparison to the expected result.csv, scoring 0.
id 98d4b8fe741f · mined from da-code dacode-data-sa-029
raw text (what the judge reads)
### Ignoring a required subgroup filter (and unverified schema) when aggregating a statistic
- **Applies when**: `task` -- the task asks for a statistic over a specific subgroup/condition (e.g., one category, species, region, cohort) while the raw files contain multiple groups, and/or a template file defines the required output columns/values.
- **Pattern**: The script reads each file and aggregates over every row, assuming the file already contains only the requested subgroup and that the column names it hard-codes exist as written; no filtering step, no schema inspection, and no comparison against the provided sample output file.
- **Detection procedure**:
  1. Read the task and note every restriction on rows (group/category, year, condition) and every stated output constraint (column names, ordering, rounding, format template).
  2. In the script, look for an explicit filtering operation (`df[df[group_col] == ...]`) on the group column and for any inspection/validation of the actual columns present in each input file; check whether the template/sample output file is ever read or its columns replicated.
  3. Check whether the script prints or asserts row counts / value ranges per group after loading, so a mismatch between the filtered subgroup size and the full file size would be visible.
  4. Compare the reported numbers to a rough domain sanity expectation (e.g., counts identical across heterogeneous files, or a ratio suspiciously near 1 when the two measured quantities are known to differ substantially) — an unexplained coincidence signals the filter or column mapping was wrong.
- **Discriminator**: A real violation is when the raw file plausibly contains rows outside the requested subgroup (or differently named columns) and the script never filters/validates; it is fine if the script demonstrably verifies the subgroup (asserts unique group values, prints per-group counts, or the file is documented as already restricted) and reproduces the template's exact columns.
- **Consequence**: The reported means/CIs are computed on the wrong row subset (or wrong columns) and the output file fails an exact/tolerance comparison to the expected result.csv, scoring 0.
65Loaded only a truncated slice of the source data (and never sanity-checked the result)taskda-code
Applies when
task -- the task asks for a statistic/model over a whole dataset file, and the script reads the data (possibly with nrows, a preview/head, a single chunk, one of several files, or a hand-made sample) before computing the requested output.
Pattern
The attempt computes the requested statistic on a small subset of the real data (e.g., a first-N-rows preview kept from an exploratory step) while reporting it as the full-dataset result, and accepts numerically implausible values (near-zero or wrong-signed correlations among conceptually related variables, tiny counts) without any cross-check against the file's true size or expected direction.
Detection procedure
  1. From the task, note that the statistic must cover the entire dataset; find the loading call in the script and check for row limits (nrows, head, slicing, chunk iteration, sampling, reading only one of multiple input files).
  2. Compare the row count printed/reported in the answer against the actual size of the source file (line count / shape); flag suspiciously round or small numbers (e.g., exactly 100) as a preview artifact.
  3. Check whether the script prints/validates shape, non-null counts, and value ranges of the analyzed columns, and whether the answer's numbers are directionally plausible for related variables.
  4. Flag if the computation runs on anything other than all rows of the intended source, or if no shape/plausibility check exists.
Discriminator
A real violation is a row limit or partial read that affects the final computation, or a reported N that is far below the file's true row count; it is fine if the limited read is used only for schema inspection/debugging while the final computation reloads the complete file, or if the reduced N is fully explained by the required missing-value filtering on a full read.
Consequence
The saved output file contains statistics from an unrepresentative subsample, so every cell differs from the expected values and the file-comparison check fails (0/1), even though the format looks correct.
id f5f9da84cf97 · mined from da-code dacode-data-sa-026
raw text (what the judge reads)
### Loaded only a truncated slice of the source data (and never sanity-checked the result)
- **Applies when**: `task` -- the task asks for a statistic/model over a whole dataset file, and the script reads the data (possibly with `nrows`, a preview/head, a single chunk, one of several files, or a hand-made sample) before computing the requested output.
- **Pattern**: The attempt computes the requested statistic on a small subset of the real data (e.g., a first-N-rows preview kept from an exploratory step) while reporting it as the full-dataset result, and accepts numerically implausible values (near-zero or wrong-signed correlations among conceptually related variables, tiny counts) without any cross-check against the file's true size or expected direction.
- **Detection procedure**:
  1. From the task, note that the statistic must cover the entire dataset; find the loading call in the script and check for row limits (`nrows`, `head`, slicing, chunk iteration, sampling, reading only one of multiple input files).
  2. Compare the row count printed/reported in the answer against the actual size of the source file (line count / shape); flag suspiciously round or small numbers (e.g., exactly 100) as a preview artifact.
  3. Check whether the script prints/validates shape, non-null counts, and value ranges of the analyzed columns, and whether the answer's numbers are directionally plausible for related variables.
  4. Flag if the computation runs on anything other than all rows of the intended source, or if no shape/plausibility check exists.
- **Discriminator**: A real violation is a row limit or partial read that affects the *final* computation, or a reported N that is far below the file's true row count; it is fine if the limited read is used only for schema inspection/debugging while the final computation reloads the complete file, or if the reduced N is fully explained by the required missing-value filtering on a full read.
- **Consequence**: The saved output file contains statistics from an unrepresentative subsample, so every cell differs from the expected values and the file-comparison check fails (0/1), even though the format looks correct.
66Cluster count chosen at the edge of the search grid, ignoring known group structuretaskda-code
Applies when
task -- the task asks for an "appropriate" number of groups/components (or any hyperparameter) and the script picks it by scanning a bounded grid and taking the arg-max/arg-min of an internal metric.
Pattern
The attempt sweeps k over a hard-coded range, selects the value with the best score, and that winner lands on the last (or first) value of the range with a weak absolute score — so the "optimum" is an artifact of where the sweep was truncated, not a real structural optimum. It also disregards strong external evidence of the natural group count available in the data (e.g., a categorical/label column that was dropped, or a documented set of classes), and never re-runs with a wider range or a second criterion (elbow, stability, comparison against the known cardinality).
Detection procedure
  1. In the task/README, note any stated or implied number of natural groups (a target/label column the script drops, documented categories, domain description).
  2. In the script, find the candidate range and the selection rule; check whether the selected value can be an endpoint of that range and whether any secondary check (elbow, stability, agreement with known cardinality) gates the choice.
  3. In the answer, compare the reported chosen value to the range bounds and look at the absolute value of the selection metric (e.g., a silhouette near 0.1–0.2 means almost no separation, so the arg-max is noise).
  4. Flag if the choice equals a range endpoint, or if the metric is essentially flat/weak, and no justification ties the choice to the data's known structure.
Discriminator
Fine if the winning value lies strictly inside the range with a clear peak, or if an endpoint win is confirmed by extending the range or by an independent criterion / the known group count; a violation is an endpoint (or noise-level) arg-max accepted as-is with no widening and no cross-check.
Consequence
The saved cluster labels use a partition count that disagrees with the reference grouping, so label-count / agreement checks (e.g., number of distinct clusters, ARI/NMI or cluster-quality thresholds) on the output file fail even though the file itself is well formed.
id a6874f279501 · mined from da-code dacode-ml-cluster-010
raw text (what the judge reads)
### Cluster count chosen at the edge of the search grid, ignoring known group structure
- **Applies when**: `task` -- the task asks for an "appropriate" number of groups/components (or any hyperparameter) and the script picks it by scanning a bounded grid and taking the arg-max/arg-min of an internal metric.
- **Pattern**: The attempt sweeps k over a hard-coded range, selects the value with the best score, and that winner lands on the last (or first) value of the range with a weak absolute score — so the "optimum" is an artifact of where the sweep was truncated, not a real structural optimum. It also disregards strong external evidence of the natural group count available in the data (e.g., a categorical/label column that was dropped, or a documented set of classes), and never re-runs with a wider range or a second criterion (elbow, stability, comparison against the known cardinality).
- **Detection procedure**:
  1. In the task/README, note any stated or implied number of natural groups (a target/label column the script drops, documented categories, domain description).
  2. In the script, find the candidate range and the selection rule; check whether the selected value can be an endpoint of that range and whether any secondary check (elbow, stability, agreement with known cardinality) gates the choice.
  3. In the answer, compare the reported chosen value to the range bounds and look at the absolute value of the selection metric (e.g., a silhouette near 0.1–0.2 means almost no separation, so the arg-max is noise).
  4. Flag if the choice equals a range endpoint, or if the metric is essentially flat/weak, and no justification ties the choice to the data's known structure.
- **Discriminator**: Fine if the winning value lies strictly inside the range with a clear peak, or if an endpoint win is confirmed by extending the range or by an independent criterion / the known group count; a violation is an endpoint (or noise-level) arg-max accepted as-is with no widening and no cross-check.
- **Consequence**: The saved cluster labels use a partition count that disagrees with the reference grouping, so label-count / agreement checks (e.g., number of distinct clusters, ARI/NMI or cluster-quality thresholds) on the output file fail even though the file itself is well formed.
67Degenerate cluster solution driven by unhandled outliers/skewtaskda-code
Applies when
task -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script picks the number of groups by maximizing an internal score over raw aggregated, heavy-tailed features.
Pattern
The script builds skewed count/monetary aggregates, applies only mean/variance scaling (no log/robust transform, no outlier handling), then selects the configuration with the best silhouette — which is almost always the smallest number of groups that isolates a handful of extreme records, yielding one giant group plus a few singletons and no interpretable segmentation.
Detection procedure
1) Read the task to confirm the deliverable is a meaningful multi-group segmentation, not just any label column. 2) In the script, check whether heavy-tailed features are log/rank/robust-transformed or winsorized before scaling, and whether the model-selection loop has any balance/interpretability guard beyond a single internal index. 3) Read the reported group sizes and score: flag if the near-perfect score coincides with a group containing a negligible fraction (e.g. <1%) of rows and the remainder in a single group. 4) Check whether the agent noticed and diagnosed this degeneracy or accepted it as "excellent quality".
Discriminator
A genuine violation is a mathematically extreme score produced by outlier isolation with essentially no partition of the bulk of the data; it is fine if a small group exists alongside several substantively sized groups, or if the agent explicitly justifies the small group after outlier treatment and shows the bulk is still meaningfully partitioned.
Consequence
The saved label file has near-zero effective segmentation (one dominant label), so any grader check on cluster count, label distribution, or segment-quality/consistency fails even though the file format is correct.
id e76127f08a60 · mined from da-code dacode-ml-cluster-019
raw text (what the judge reads)
### Degenerate cluster solution driven by unhandled outliers/skew
- **Applies when**: `task` -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script picks the number of groups by maximizing an internal score over raw aggregated, heavy-tailed features.
- **Pattern**: The script builds skewed count/monetary aggregates, applies only mean/variance scaling (no log/robust transform, no outlier handling), then selects the configuration with the best silhouette — which is almost always the smallest number of groups that isolates a handful of extreme records, yielding one giant group plus a few singletons and no interpretable segmentation.
- **Detection procedure**: 1) Read the task to confirm the deliverable is a meaningful multi-group segmentation, not just any label column. 2) In the script, check whether heavy-tailed features are log/rank/robust-transformed or winsorized before scaling, and whether the model-selection loop has any balance/interpretability guard beyond a single internal index. 3) Read the reported group sizes and score: flag if the near-perfect score coincides with a group containing a negligible fraction (e.g. <1%) of rows and the remainder in a single group. 4) Check whether the agent noticed and diagnosed this degeneracy or accepted it as "excellent quality".
- **Discriminator**: A genuine violation is a mathematically extreme score produced by outlier isolation with essentially no partition of the bulk of the data; it is fine if a small group exists alongside several substantively sized groups, or if the agent explicitly justifies the small group after outlier treatment and shows the bulk is still meaningfully partitioned.
- **Consequence**: The saved label file has near-zero effective segmentation (one dominant label), so any grader check on cluster count, label distribution, or segment-quality/consistency fails even though the file format is correct.
68Hardcoded/recalled numbers instead of resampling the provided datataskda-code
Applies when
task -- the task supplies a dataset (with a README) and asks for a statistic or confidence interval computed from it, and the script defines the inputs inline rather than loading any file.
Pattern
The attempt never reads the provided data files; it types in summary counts or rates from memory/background knowledge and then simulates from a parametric assumption (e.g., drawing from a Binomial with the assumed rate) instead of resampling the actual observed records, so both the point estimate and the interval reflect invented inputs and an unrequested method.
Detection procedure
  1. Read the task/README and list the data files that are expected to be the source of the requested quantity.
  2. Scan the script for any I/O (read_csv, read_excel, load, path strings) and check that the quantity is derived from those loaded objects; flag literal constants that stand in for data.
  3. Check that the resampling/estimation step draws with replacement from the loaded observations (or their groups) rather than from a hand-specified distribution or parameter.
  4. Compare the reported numbers against what the actual files would yield (row counts, group sizes, the observed effect size) — if the script prints counts that cannot be traced to a file, treat the result as unverified.
Discriminator
A real violation is when key inputs (counts, rates, group definitions, filters) originate from the agent's assumptions and are never cross-checked against the delivered files; it is fine to hardcode constants that are explicitly stated in the task/README, or to compute summaries in one script and pass them along, as long as they were derived from the loaded data.
Consequence
The output file contains an interval centered on the wrong effect size (and with the wrong width from a parametric rather than empirical bootstrap), so the graded value falls outside the expected tolerance and the result file is marked WRONG.
id 4a57796ac07a · mined from da-code dacode-data-sa-031
raw text (what the judge reads)
### Hardcoded/recalled numbers instead of resampling the provided data
- **Applies when**: `task` -- the task supplies a dataset (with a README) and asks for a statistic or confidence interval computed from it, and the script defines the inputs inline rather than loading any file.
- **Pattern**: The attempt never reads the provided data files; it types in summary counts or rates from memory/background knowledge and then simulates from a parametric assumption (e.g., drawing from a Binomial with the assumed rate) instead of resampling the actual observed records, so both the point estimate and the interval reflect invented inputs and an unrequested method.
- **Detection procedure**:
  1. Read the task/README and list the data files that are expected to be the source of the requested quantity.
  2. Scan the script for any I/O (`read_csv`, `read_excel`, `load`, path strings) and check that the quantity is derived from those loaded objects; flag literal constants that stand in for data.
  3. Check that the resampling/estimation step draws with replacement from the loaded observations (or their groups) rather than from a hand-specified distribution or parameter.
  4. Compare the reported numbers against what the actual files would yield (row counts, group sizes, the observed effect size) — if the script prints counts that cannot be traced to a file, treat the result as unverified.
- **Discriminator**: A real violation is when key inputs (counts, rates, group definitions, filters) originate from the agent's assumptions and are never cross-checked against the delivered files; it is fine to hardcode constants that are explicitly stated in the task/README, or to compute summaries in one script and pass them along, as long as they were derived from the loaded data.
- **Consequence**: The output file contains an interval centered on the wrong effect size (and with the wrong width from a parametric rather than empirical bootstrap), so the graded value falls outside the expected tolerance and the result file is marked WRONG.
69Missing verification that the prediction file exactly matches the required submission templatetaskda-code
Applies when
task -- the task requires writing an output file whose rows/keys and column layout must match a provided sample/template file over an entire held-out set.
Pattern
The scripts build predictions and write the output directly from model probabilities, but never load the template, never assert that the number of rows equals the number of held-out records, that the key column covers exactly the same identifiers (same set, same order), and that the column names/order match; the final answer is accepted without any row-count or key-coverage sanity check (and can silently be truncated, reordered, subset, or column-swapped).
Detection procedure
  1. Read the task for the stated output contract: required file name, header names, key column, and the fact that every held-out record must appear exactly once.
  2. Scan the scripts for an explicit comparison against that contract — e.g. reading the sample/template file, assert len(sub) == len(test), set(sub[key]) == set(template[key]), column-name/order equality, and a check that written file re-reads with the expected shape. Absence of all of these is the violation.
  3. Inspect the produced answer/file: count rows and compare to the held-out record count; check the key column for duplicates, missing ids, and that class-probability columns are in the specified order and non-degenerate.
  4. Also confirm the class-index-to-column mapping is derived from the fitted model's class ordering rather than assumed positionally.
Discriminator
A real violation is the total absence of any template/row-count/key check (or an actual mismatch in the emitted file). It is not a violation if the script reconstructs the submission from the template's key column (e.g. merges predictions onto the template ids) or asserts shape/ids/columns before saving — the checks may be terse, but they must exist and cover keys, count, and column layout.
Consequence
The grader reports the expected output file as WRONG/MISSING regardless of model quality, since rows are missing/misaligned or columns are mismapped, and the log-loss cannot be computed over the full held-out set.
id 06ba55fe8222 · mined from da-code dacode-ml-competition-005@s2
raw text (what the judge reads)
### Missing verification that the prediction file exactly matches the required submission template
- **Applies when**: `task` -- the task requires writing an output file whose rows/keys and column layout must match a provided sample/template file over an entire held-out set.
- **Pattern**: The scripts build predictions and write the output directly from model probabilities, but never load the template, never assert that the number of rows equals the number of held-out records, that the key column covers exactly the same identifiers (same set, same order), and that the column names/order match; the final answer is accepted without any row-count or key-coverage sanity check (and can silently be truncated, reordered, subset, or column-swapped).
- **Detection procedure**:
  1. Read the task for the stated output contract: required file name, header names, key column, and the fact that every held-out record must appear exactly once.
  2. Scan the scripts for an explicit comparison against that contract — e.g. reading the sample/template file, `assert len(sub) == len(test)`, `set(sub[key]) == set(template[key])`, column-name/order equality, and a check that written file re-reads with the expected shape. Absence of all of these is the violation.
  3. Inspect the produced answer/file: count rows and compare to the held-out record count; check the key column for duplicates, missing ids, and that class-probability columns are in the specified order and non-degenerate.
  4. Also confirm the class-index-to-column mapping is derived from the fitted model's class ordering rather than assumed positionally.
- **Discriminator**: A real violation is the total absence of any template/row-count/key check (or an actual mismatch in the emitted file). It is not a violation if the script reconstructs the submission from the template's key column (e.g. merges predictions onto the template ids) or asserts shape/ids/columns before saving — the checks may be terse, but they must exist and cover keys, count, and column layout.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING regardless of model quality, since rows are missing/misaligned or columns are mismapped, and the log-loss cannot be computed over the full held-out set.
70Predicted values not sanity-checked against the training target's scale and typetaskda-code
Applies when
task -- the agent trains a regressor to produce a predicted count/quantity/price column for held-out rows and writes it to an output file.
Pattern
The script transforms the target (log/sqrt/scaling) or fits on a filtered/atypical subset, then writes raw model output without inverting the transform or checking it against the observed target distribution — producing predictions whose scale, dtype (e.g., fractional values for integer counts), or row count is inconsistent with the training labels, with no assertion or comparison step.
Detection procedure
  1. From the task/README, note what the target represents (units, integer vs continuous, plausible range) and how many output rows are required.
  2. In the scripts, find every target transformation, clipping, subset filter, or scaling; verify an explicit inverse transform is applied before writing, and that the output is built from the full held-out set in original order.
  3. Compare the answer's summary statistics (min, median, max, count, share of near-zero values) against the training target's statistics; flag if the central tendency differs by an order of magnitude or the value type is incompatible (e.g., mostly sub-unit values where labels are large integers).
  4. Check the script contains an explicit sanity assertion on shape/range before saving; absence plus a mismatch in step 3 confirms the violation.
Discriminator
A genuine violation shows a systematic distribution mismatch (e.g., predictions concentrated far below the label median, or row count ≠ required rows). It is not a violation if predictions are merely smoother/less dispersed than labels — regression shrinkage toward the mean is expected as long as the central tendency and units match the labels.
Consequence
The output file fails the grader's value/format comparison — error metrics are dominated by a constant scale bias (or the file has the wrong number/type of entries), so the answer is marked wrong despite a plausibly reasonable model.
id 6ffdf2b3a3f1 · mined from da-code dacode-ml-regression-008@s2
raw text (what the judge reads)
### Predicted values not sanity-checked against the training target's scale and type
- **Applies when**: `task` -- the agent trains a regressor to produce a predicted count/quantity/price column for held-out rows and writes it to an output file.
- **Pattern**: The script transforms the target (log/sqrt/scaling) or fits on a filtered/atypical subset, then writes raw model output without inverting the transform or checking it against the observed target distribution — producing predictions whose scale, dtype (e.g., fractional values for integer counts), or row count is inconsistent with the training labels, with no assertion or comparison step.
- **Detection procedure**:
  1. From the task/README, note what the target represents (units, integer vs continuous, plausible range) and how many output rows are required.
  2. In the scripts, find every target transformation, clipping, subset filter, or scaling; verify an explicit inverse transform is applied before writing, and that the output is built from the full held-out set in original order.
  3. Compare the answer's summary statistics (min, median, max, count, share of near-zero values) against the training target's statistics; flag if the central tendency differs by an order of magnitude or the value type is incompatible (e.g., mostly sub-unit values where labels are large integers).
  4. Check the script contains an explicit sanity assertion on shape/range before saving; absence plus a mismatch in step 3 confirms the violation.
- **Discriminator**: A genuine violation shows a systematic distribution mismatch (e.g., predictions concentrated far below the label median, or row count ≠ required rows). It is *not* a violation if predictions are merely smoother/less dispersed than labels — regression shrinkage toward the mean is expected as long as the central tendency and units match the labels.
- **Consequence**: The output file fails the grader's value/format comparison — error metrics are dominated by a constant scale bias (or the file has the wrong number/type of entries), so the answer is marked wrong despite a plausibly reasonable model.
71Unfiltered population and unchecked test assumptions in hypothesis testingtaskda-code
Applies when
task -- the task asks for a p-value / significance decision about a statistic of some group comparison, and the scripts run an off-the-shelf test on whole raw tables straight from disk.
Pattern
The script loads every row of each source file, computes the derived quantity, and immediately applies a default parametric test (e.g., independent-samples t-test) without (a) restricting to the subpopulation, time window, or competition/segment that the question is actually about, and (b) checking whether the test's distributional assumptions (normality, symmetry, equal variance, independence) hold for the derived quantity — typically a small-integer, skewed, count-like variable.
Detection procedure
  1. Read the task/README and list every scoping qualifier implied by the question (which subset of records, which period, which category the claim concerns) plus the significance level and output format.
  2. Read the script for filtering steps: does any query/boolean mask/date parse restrict the rows before the statistic is computed? If the row counts used are essentially the full file sizes, the scoping step is missing.
  3. Inspect the test choice: is there any diagnostic (histogram, skew, normality test, variance comparison) or justification, and does the test's one/two-sided direction match the stated hypothesis? A default ttest_ind on a bounded count variable with no diagnostic is a red flag.
  4. Check the reported p-value's magnitude for plausibility given the intended (much smaller) subset — an astronomically small p-value is a symptom of using far more rows than the question intended.
Discriminator
A real violation is when the task or README implies a narrower population/assumption-appropriate test and the script silently uses all rows with a default parametric test; it is fine if the task genuinely asks about all records and the script either documents the assumption check or the chosen test is robust/appropriate for the variable's distribution.
Consequence
The p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample with the wrong test, so the saved value does not match the expected one and the file fails the grader even though the format is correct.
id 519e6a3e5f2c · mined from da-code dacode-data-sa-001@s2
raw text (what the judge reads)
### Unfiltered population and unchecked test assumptions in hypothesis testing
- **Applies when**: `task` -- the task asks for a p-value / significance decision about a statistic of some group comparison, and the scripts run an off-the-shelf test on whole raw tables straight from disk.
- **Pattern**: The script loads every row of each source file, computes the derived quantity, and immediately applies a default parametric test (e.g., independent-samples t-test) without (a) restricting to the subpopulation, time window, or competition/segment that the question is actually about, and (b) checking whether the test's distributional assumptions (normality, symmetry, equal variance, independence) hold for the derived quantity — typically a small-integer, skewed, count-like variable.
- **Detection procedure**:
  1. Read the task/README and list every scoping qualifier implied by the question (which subset of records, which period, which category the claim concerns) plus the significance level and output format.
  2. Read the script for filtering steps: does any `query`/boolean mask/date parse restrict the rows before the statistic is computed? If the row counts used are essentially the full file sizes, the scoping step is missing.
  3. Inspect the test choice: is there any diagnostic (histogram, skew, normality test, variance comparison) or justification, and does the test's one/two-sided direction match the stated hypothesis? A default `ttest_ind` on a bounded count variable with no diagnostic is a red flag.
  4. Check the reported p-value's magnitude for plausibility given the intended (much smaller) subset — an astronomically small p-value is a symptom of using far more rows than the question intended.
- **Discriminator**: A real violation is when the task or README implies a narrower population/assumption-appropriate test and the script silently uses all rows with a default parametric test; it is fine if the task genuinely asks about all records and the script either documents the assumption check or the chosen test is robust/appropriate for the variable's distribution.
- **Consequence**: The p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample with the wrong test, so the saved value does not match the expected one and the file fails the grader even though the format is correct.
72Ignoring the provided output template (format/rounding/ordering never verified)taskda-code
Applies when
task -- the task supplies a sample/expected output file (or explicit format spec) and asks for results "following the exact structure and formatting" of it.
Pattern
The scripts compute aggregates but never read or parse the template file; column names, column order, row order, and numeric precision are guessed from intuition, and raw floating-point sums are written verbatim (e.g. values with 10+ decimal digits or inconsistent decimal counts across rows), so the output can't byte-match the reference even when the underlying math is right.
Detection procedure
  1. In the task text, note whether a template/sample output artifact or explicit formatting rules (rounding, units, ordering, header names) are mentioned.
  2. Search the scripts for any read of that template (e.g. loading the sample file) and for any explicit formatting step (round(), float_format, sort key derived from the template, column list copied from the template header). Absence of both is the red flag.
  3. Inspect the produced answer: check that all numeric cells share a consistent, plausible precision, that header names/order match the template exactly, and that row order follows a rule that is stated or derivable from the template rather than an ad-hoc choice.
  4. Confirm no post-hoc comparison of the generated file against the template's shape/columns/dtypes was performed.
Discriminator
A real violation is when the template is never inspected and the formatting choices (precision, ordering, header) are unverified guesses. It is not a violation if the script reads the template (or the task states the format literally) and demonstrably enforces the same columns, order, and rounding — even if it does so with hardcoded values copied from the template.
Consequence
The grader's exact/tolerance file comparison fails on formatting alone (extra decimal digits, wrong row order, or mismatched headers), reporting the result file as WRONG despite arguably correct aggregation logic.
id 0d14093b2bf3 · mined from da-code dacode-dm-csv-011@s2
raw text (what the judge reads)
### Ignoring the provided output template (format/rounding/ordering never verified)
- **Applies when**: `task` -- the task supplies a sample/expected output file (or explicit format spec) and asks for results "following the exact structure and formatting" of it.
- **Pattern**: The scripts compute aggregates but never read or parse the template file; column names, column order, row order, and numeric precision are guessed from intuition, and raw floating-point sums are written verbatim (e.g. values with 10+ decimal digits or inconsistent decimal counts across rows), so the output can't byte-match the reference even when the underlying math is right.
- **Detection procedure**:
  1. In the task text, note whether a template/sample output artifact or explicit formatting rules (rounding, units, ordering, header names) are mentioned.
  2. Search the scripts for any read of that template (e.g. loading the sample file) and for any explicit formatting step (`round()`, `float_format`, sort key derived from the template, column list copied from the template header). Absence of both is the red flag.
  3. Inspect the produced answer: check that all numeric cells share a consistent, plausible precision, that header names/order match the template exactly, and that row order follows a rule that is stated or derivable from the template rather than an ad-hoc choice.
  4. Confirm no post-hoc comparison of the generated file against the template's shape/columns/dtypes was performed.
- **Discriminator**: A real violation is when the template is never inspected and the formatting choices (precision, ordering, header) are unverified guesses. It is *not* a violation if the script reads the template (or the task states the format literally) and demonstrably enforces the same columns, order, and rounding — even if it does so with hardcoded values copied from the template.
- **Consequence**: The grader's exact/tolerance file comparison fails on formatting alone (extra decimal digits, wrong row order, or mismatched headers), reporting the result file as WRONG despite arguably correct aggregation logic.
73Output feature matrix silently redefined and never sanity-checked against the requested spectaskda-code
Applies when
task -- the task asks for a result file whose columns are the per-record feature values (e.g., Feature_i) plus a derived label, and the script builds that file from an ad-hoc, self-chosen preprocessing pipeline.
Pattern
The agent drops, imputes, one-hot expands and reorders columns for modeling convenience, then dumps that transformed matrix as the "feature vector" without checking that the saved file's row count, column count/order, value scale and label column are a defensible, non-degenerate representation of the original records; the choice of the number of clusters/labels is also taken straight from an unexamined metric on that arbitrary space (often collapsing to the minimum allowed value, e.g. 2, with a very low score) and never sanity-checked.
Detection procedure
  1. Read the task statement and note exactly what each output column is supposed to contain (which records, which feature values, what label) and where the file must be written.
  2. In the script, trace every transformation between loading the raw data and to_csv: dropped columns, imputations, encodings, scaling, row filtering — and check whether the written matrix is the raw feature values, the transformed ones, or an inconsistent mix, and whether row count still equals the number of input records.
  3. Check for any post-write validation: re-reading the file, asserting shape/column names/dtypes, label counts and no NaNs, and a check that the produced grouping is non-trivial (more than a minimal split, balanced enough, score meaningfully above the alternatives).
  4. Read the answer text: does it merely restate the pipeline and the selection metric, or does it show these validation numbers and justify the discarded/derived columns?
Discriminator
A real violation is when transformations that change the feature vector's identity (dropping informative columns, expanding categoricals, imputation) are made without justification and no shape/content/degeneracy check on the saved file appears anywhere; it is not a violation if the script documents a principled feature definition and verifies the written file (rows = records, expected column naming/ordering, sensible label distribution), even if the exact preprocessing differs from a reference.
Consequence
The graded file has a feature matrix and/or label column that cannot be matched to the expected output (wrong width, transformed values, trivial or collapsed clustering), so the file check fails outright even though the script ran without error.
id 7fbf54e450eb · mined from da-code dacode-ml-cluster-014@s2
raw text (what the judge reads)
### Output feature matrix silently redefined and never sanity-checked against the requested spec
- **Applies when**: `task` -- the task asks for a result file whose columns are the per-record feature values (e.g., `Feature_i`) plus a derived label, and the script builds that file from an ad-hoc, self-chosen preprocessing pipeline.
- **Pattern**: The agent drops, imputes, one-hot expands and reorders columns for modeling convenience, then dumps that transformed matrix as the "feature vector" without checking that the saved file's row count, column count/order, value scale and label column are a defensible, non-degenerate representation of the original records; the choice of the number of clusters/labels is also taken straight from an unexamined metric on that arbitrary space (often collapsing to the minimum allowed value, e.g. 2, with a very low score) and never sanity-checked.
- **Detection procedure**:
  1. Read the task statement and note exactly what each output column is supposed to contain (which records, which feature values, what label) and where the file must be written.
  2. In the script, trace every transformation between loading the raw data and `to_csv`: dropped columns, imputations, encodings, scaling, row filtering — and check whether the written matrix is the raw feature values, the transformed ones, or an inconsistent mix, and whether row count still equals the number of input records.
  3. Check for any post-write validation: re-reading the file, asserting shape/column names/dtypes, label counts and no NaNs, and a check that the produced grouping is non-trivial (more than a minimal split, balanced enough, score meaningfully above the alternatives).
  4. Read the answer text: does it merely restate the pipeline and the selection metric, or does it show these validation numbers and justify the discarded/derived columns?
- **Discriminator**: A real violation is when transformations that change the feature vector's identity (dropping informative columns, expanding categoricals, imputation) are made without justification *and* no shape/content/degeneracy check on the saved file appears anywhere; it is *not* a violation if the script documents a principled feature definition and verifies the written file (rows = records, expected column naming/ordering, sensible label distribution), even if the exact preprocessing differs from a reference.
- **Consequence**: The graded file has a feature matrix and/or label column that cannot be matched to the expected output (wrong width, transformed values, trivial or collapsed clustering), so the file check fails outright even though the script ran without error.
74Never validating the output file against the provided submission templatetaskda-code
Applies when
task -- the task supplies a sample/template output file and asks for predictions written to a submission file in that exact format.
Pattern
The scripts build a submission purely from the agent's own assumptions (hand-written column names, ids taken from an intermediate frame, predictions coerced/rounded to a chosen dtype) and never load the template to check column names, row count, id set/order, or value type; a later "improved" script may also be truncated or fail, silently leaving an earlier or partial file on disk.
Detection procedure
  1. Read the task/README for the named template file and any stated format or metric constraints (column names, one row per test id, continuous vs. integer targets).
  2. Search the scripts for a read of that template and for an explicit comparison against it (shape equality, set(ids) equality, column-name equality, ordering, null check) before writing the output.
  3. Check that the final produced file comes from the last script that ran to completion, and that any post-processing (rounding, clipping, casting) is justified by the stated metric rather than applied by default.
  4. Inspect the answer itself: count rows and compare with the expected test size, confirm the header matches the template, and confirm no duplicate/missing ids.
Discriminator
A real violation is the absence of any template-based verification and an output whose shape/ids/dtype cannot be shown to match; it is fine if the script derives ids directly from the template or test file and asserts shape/id agreement (even without printing), or if rounding is explicitly required by the task.
Consequence
The grader reports the submission as WRONG/MISSING — row count or id set mismatched, wrong header, or unnecessary integer rounding inflating the error metric — even though the model itself may be reasonable.
id 3c6582a36cc3 · mined from da-code dacode-ml-competition-009@s2
raw text (what the judge reads)
### Never validating the output file against the provided submission template
- **Applies when**: `task` -- the task supplies a sample/template output file and asks for predictions written to a submission file in that exact format.
- **Pattern**: The scripts build a submission purely from the agent's own assumptions (hand-written column names, ids taken from an intermediate frame, predictions coerced/rounded to a chosen dtype) and never load the template to check column names, row count, id set/order, or value type; a later "improved" script may also be truncated or fail, silently leaving an earlier or partial file on disk.
- **Detection procedure**:
  1. Read the task/README for the named template file and any stated format or metric constraints (column names, one row per test id, continuous vs. integer targets).
  2. Search the scripts for a read of that template and for an explicit comparison against it (shape equality, `set(ids)` equality, column-name equality, ordering, null check) before writing the output.
  3. Check that the final produced file comes from the last script that ran to completion, and that any post-processing (rounding, clipping, casting) is justified by the stated metric rather than applied by default.
  4. Inspect the answer itself: count rows and compare with the expected test size, confirm the header matches the template, and confirm no duplicate/missing ids.
- **Discriminator**: A real violation is the absence of any template-based verification *and* an output whose shape/ids/dtype cannot be shown to match; it is fine if the script derives ids directly from the template or test file and asserts shape/id agreement (even without printing), or if rounding is explicitly required by the task.
- **Consequence**: The grader reports the submission as WRONG/MISSING — row count or id set mismatched, wrong header, or unnecessary integer rounding inflating the error metric — even though the model itself may be reasonable.
75Arbitrary feature subsetting when the task implies the full feature vectortaskda-code
Applies when
task -- the task asks to run an unsupervised/modeling procedure on "the dataset" and to output the feature vector alongside results (e.g., columns named Feature_i), without naming which columns to use.
Pattern
The script silently keeps only a handful of hand-picked columns (often a thematically related few, or only those with no missing values) instead of using all usable columns after standard cleaning, so the emitted feature matrix has far fewer dimensions than the data provides and the labels come from a different feature space than the reference solution.
Detection procedure
  1. Read the task for any statement restricting the columns; if none, the default is "all informative columns after cleaning" (drop pure identifiers/text keys, coerce numeric strings, encode or drop categoricals, impute or drop missing values in a documented way).
  2. In the script, find the column-selection step and count how many columns enter the model; compare with the number of candidate columns in the raw file.
  3. Check whether each exclusion is justified by a stated rule (non-numeric, identifier, all-missing) rather than by topical judgment or convenience; check whether missing values were dropped/imputed instead of used as a reason to discard whole columns.
  4. Compare the answer's reported feature count/columns (Feature_0..Feature_k) against the expected dimensionality from step 2 and against the row count of the output.
Discriminator
A real violation is dropping numeric, mostly-populated columns for subjective reasons (or to avoid handling NaNs/dtype cleanup); it is fine to drop columns that are identifiers, codes, free text, or unparseable, or to reduce dimensionality by an explicit method the task allows, when the reasoning is stated and applied uniformly.
Consequence
The output file's schema (number of Feature_i columns) and the cluster assignments diverge from the reference, so an exact/структural file comparison fails even though the pipeline ran without error.
id 3a42457fcc6b · mined from da-code dacode-ml-cluster-009@s2
raw text (what the judge reads)
### Arbitrary feature subsetting when the task implies the full feature vector
- **Applies when**: `task` -- the task asks to run an unsupervised/modeling procedure on "the dataset" and to output the feature vector alongside results (e.g., columns named `Feature_i`), without naming which columns to use.
- **Pattern**: The script silently keeps only a handful of hand-picked columns (often a thematically related few, or only those with no missing values) instead of using all usable columns after standard cleaning, so the emitted feature matrix has far fewer dimensions than the data provides and the labels come from a different feature space than the reference solution.
- **Detection procedure**:
  1. Read the task for any statement restricting the columns; if none, the default is "all informative columns after cleaning" (drop pure identifiers/text keys, coerce numeric strings, encode or drop categoricals, impute or drop missing values in a documented way).
  2. In the script, find the column-selection step and count how many columns enter the model; compare with the number of candidate columns in the raw file.
  3. Check whether each exclusion is justified by a stated rule (non-numeric, identifier, all-missing) rather than by topical judgment or convenience; check whether missing values were dropped/imputed instead of used as a reason to discard whole columns.
  4. Compare the answer's reported feature count/columns (`Feature_0..Feature_k`) against the expected dimensionality from step 2 and against the row count of the output.
- **Discriminator**: A real violation is dropping numeric, mostly-populated columns for subjective reasons (or to avoid handling NaNs/dtype cleanup); it is fine to drop columns that are identifiers, codes, free text, or unparseable, or to reduce dimensionality by an explicit method the task allows, when the reasoning is stated and applied uniformly.
- **Consequence**: The output file's schema (number of `Feature_i` columns) and the cluster assignments diverge from the reference, so an exact/структural file comparison fails even though the pipeline ran without error.
76Silently dropping entities that must appear in the required output filetaskda-code
Applies when
task -- the deliverable is a per-record/per-entity output file (labels, predictions, scores) covering the units derived from the input data, and the script applies filtering or aggregation before producing it.
Pattern
The script removes a substantial subset of the entities as a side effect of preprocessing or feature construction (e.g. dropping rows where a derived statistic is undefined/NaN, requiring a minimum count, trimming "outliers"), so the saved file has far fewer rows than the natural population of entities in the cleaned data — and the answer never reconciles this row count against the input.
Detection procedure
  1. From the task, identify the unit of the requested output file and the expected row count implied by the cleaned input (e.g. number of distinct entities after only the clearly-justified cleaning steps).
  2. In the scripts, list every filter/dropna/threshold applied after the entity table is built, and note how many units each removes; check whether any is driven purely by a self-created feature being undefined rather than by data validity.
  3. Compare the row count written to the output file with the count from step 1; also check the column naming/indexing convention matches the literal spec (e.g. Feature_0… vs Feature_1…, presence/absence of an ID column, ordering).
  4. Check the answer text: does it state the final row count and explain the gap, or does it present coverage loss as a methodological choice without validating it against the requirement?
Discriminator
A real violation is dropping units that are valid members of the target population purely for algorithmic convenience (undefined std, "needs ≥2 events", outlier removal) with no instruction to do so; a look-alike that is fine is removing records that cannot be attributed to any unit at all (missing entity key) or that the task/README explicitly designates as invalid (cancellations), where the remaining set is still the full population.
Consequence
The saved file has the wrong shape/coverage (thousands of entities missing) and, if the grader compares row counts, column names, or per-entity labels, it fails the file check outright even though the clustering itself may be reasonable; degenerate singleton clusters from unremoved extremes are a further symptom that no sanity check on cluster sizes was run.
id f78080af2b84 · mined from da-code dacode-ml-cluster-016@s2
raw text (what the judge reads)
### Silently dropping entities that must appear in the required output file
- **Applies when**: `task` -- the deliverable is a per-record/per-entity output file (labels, predictions, scores) covering the units derived from the input data, and the script applies filtering or aggregation before producing it.
- **Pattern**: The script removes a substantial subset of the entities as a side effect of preprocessing or feature construction (e.g. dropping rows where a derived statistic is undefined/NaN, requiring a minimum count, trimming "outliers"), so the saved file has far fewer rows than the natural population of entities in the cleaned data — and the answer never reconciles this row count against the input.
- **Detection procedure**:
  1. From the task, identify the unit of the requested output file and the expected row count implied by the cleaned input (e.g. number of distinct entities after only the clearly-justified cleaning steps).
  2. In the scripts, list every filter/`dropna`/threshold applied *after* the entity table is built, and note how many units each removes; check whether any is driven purely by a self-created feature being undefined rather than by data validity.
  3. Compare the row count written to the output file with the count from step 1; also check the column naming/indexing convention matches the literal spec (e.g. `Feature_0…` vs `Feature_1…`, presence/absence of an ID column, ordering).
  4. Check the answer text: does it state the final row count and explain the gap, or does it present coverage loss as a methodological choice without validating it against the requirement?
- **Discriminator**: A real violation is dropping units that are valid members of the target population purely for algorithmic convenience (undefined std, "needs ≥2 events", outlier removal) with no instruction to do so; a look-alike that is fine is removing records that cannot be attributed to any unit at all (missing entity key) or that the task/README explicitly designates as invalid (cancellations), where the remaining set is still the full population.
- **Consequence**: The saved file has the wrong shape/coverage (thousands of entities missing) and, if the grader compares row counts, column names, or per-entity labels, it fails the file check outright even though the clustering itself may be reasonable; degenerate singleton clusters from unremoved extremes are a further symptom that no sanity check on cluster sizes was run.
77Output file schema/row-alignment not verified against the required submission spectaskda-code
Applies when
task -- the task asks for predictions (or derived values) written to a named file with an explicitly stated column name, one row per input record.
Pattern
The agent writes the file with its own convention — extra index/key columns, renamed or extra headers, a different row count than the test input, or rows in an order other than the input's — and reports only model quality statistics, never checking that the artifact matches the literal requested format.
Detection procedure
  1. From the task statement, extract the exact required artifact: filename, required column name(s), expected number of rows (= rows in the provided test input), and implied row order.
  2. In the scripts, find the write step (e.g., to_csv) and check what DataFrame is written: which columns it contains, whether the header spelling matches exactly, whether an index is written, and whether the frame was built from the full, unfiltered, un-reordered test input (no dropped rows from NaN handling, no re-sorting, no deduplication).
  3. In the answer, look for an explicit verification of shape/columns against the test file (row count equality, column list equality, no NaNs); absence of such a check, or a reported row/column set that differs from the input's, is the flag.
  4. Cross-check the reported prediction count and column list against the test input's row count and the required header.
Discriminator
A real violation is a mismatch in the artifact itself — missing/renamed/extra columns, row count ≠ test rows, or rows realigned by dropping/sorting. It is not a violation if extra columns are explicitly permitted by the task, or if the agent kept an auxiliary column but verified counts and order and the required column is present and correctly named; nor is strong/weak model accuracy relevant here.
Consequence
The grader loads the file, fails schema or index alignment (or compares mismatched rows), and marks the result wrong/missing regardless of how good the underlying model was.
id eac24519e6c2 · mined from da-code dacode-ml-regression-002@s2
raw text (what the judge reads)
### Output file schema/row-alignment not verified against the required submission spec
- **Applies when**: `task` -- the task asks for predictions (or derived values) written to a named file with an explicitly stated column name, one row per input record.
- **Pattern**: The agent writes the file with its own convention — extra index/key columns, renamed or extra headers, a different row count than the test input, or rows in an order other than the input's — and reports only model quality statistics, never checking that the artifact matches the literal requested format.
- **Detection procedure**:
  1. From the task statement, extract the exact required artifact: filename, required column name(s), expected number of rows (= rows in the provided test input), and implied row order.
  2. In the scripts, find the write step (e.g., `to_csv`) and check what DataFrame is written: which columns it contains, whether the header spelling matches exactly, whether an index is written, and whether the frame was built from the full, unfiltered, un-reordered test input (no dropped rows from NaN handling, no re-sorting, no deduplication).
  3. In the answer, look for an explicit verification of shape/columns against the test file (row count equality, column list equality, no NaNs); absence of such a check, or a reported row/column set that differs from the input's, is the flag.
  4. Cross-check the reported prediction count and column list against the test input's row count and the required header.
- **Discriminator**: A real violation is a mismatch in the artifact itself — missing/renamed/extra columns, row count ≠ test rows, or rows realigned by dropping/sorting. It is *not* a violation if extra columns are explicitly permitted by the task, or if the agent kept an auxiliary column but verified counts and order and the required column is present and correctly named; nor is strong/weak model accuracy relevant here.
- **Consequence**: The grader loads the file, fails schema or index alignment (or compares mismatched rows), and marks the result wrong/missing regardless of how good the underlying model was.
78Unverifiable, artifact-incomplete plotting claims (no saved script, no data dump)taskda-code
Applies when
task -- the task asks for a chart/output built to a spec file and the grading depends on the underlying plotted data being recoverable (e.g., a serialized plot description and/or a numeric array), not just an image file.
Pattern
The agent reports a checklist of "spec met ✓" items and an image file, but leaves behind no runnable script and no machine-readable dump of the plotted series; the numbers quoted appear to be typed in or eyeballed rather than aggregated from the source file, and the auxiliary artifacts the grader reads are never written.
Detection procedure
  1. Read the task/spec and list every expected output artifact (image, JSON/spec dump, .npy/CSV of plotted values) and every constrained property (labels, ticks, order, units, aggregation level).
  2. Check the saved scripts: is there code that loads the raw source, performs the stated aggregation, plots, and then explicitly serializes the plotted x/y values and figure metadata to the required artifact names?
  3. Compare the numbers in the answer to the aggregation the script would produce — do they trace to a computed object, or are they hard-coded/suspiciously patterned (e.g., a value equal to the group label, implausibly small counts for the stated unit)?
  4. Confirm the answer asserts nothing that isn't demonstrably read back from the produced artifacts (e.g., re-reading the dump and printing shape/range).
Discriminator
A real violation is when required non-image artifacts are absent or the plotted values cannot be reproduced from code over the raw data; a look-alike that is fine is a script that computes and dumps the series programmatically and merely summarizes it in prose, even if only the image is visually inspected.
Consequence
The grader's checks on the expected artifacts (plot spec JSON, numeric array) report WRONG/MISSING, and any hard-coded series that differs from the true aggregation fails value comparison — scoring 0 despite a confident "task completed" report.
id ffe44a540300 · mined from da-code dacode-plot-line-015@s2
raw text (what the judge reads)
### Unverifiable, artifact-incomplete plotting claims (no saved script, no data dump)
- **Applies when**: `task` -- the task asks for a chart/output built to a spec file and the grading depends on the underlying plotted data being recoverable (e.g., a serialized plot description and/or a numeric array), not just an image file.
- **Pattern**: The agent reports a checklist of "spec met ✓" items and an image file, but leaves behind no runnable script and no machine-readable dump of the plotted series; the numbers quoted appear to be typed in or eyeballed rather than aggregated from the source file, and the auxiliary artifacts the grader reads are never written.
- **Detection procedure**:
  1. Read the task/spec and list every expected output artifact (image, JSON/spec dump, `.npy`/CSV of plotted values) and every constrained property (labels, ticks, order, units, aggregation level).
  2. Check the saved scripts: is there code that loads the raw source, performs the stated aggregation, plots, and then explicitly serializes the plotted x/y values and figure metadata to the required artifact names?
  3. Compare the numbers in the answer to the aggregation the script would produce — do they trace to a computed object, or are they hard-coded/suspiciously patterned (e.g., a value equal to the group label, implausibly small counts for the stated unit)?
  4. Confirm the answer asserts nothing that isn't demonstrably read back from the produced artifacts (e.g., re-reading the dump and printing shape/range).
- **Discriminator**: A real violation is when required non-image artifacts are absent or the plotted values cannot be reproduced from code over the raw data; a look-alike that is fine is a script that computes and dumps the series programmatically and merely *summarizes* it in prose, even if only the image is visually inspected.
- **Consequence**: The grader's checks on the expected artifacts (plot spec JSON, numeric array) report WRONG/MISSING, and any hard-coded series that differs from the true aggregation fails value comparison — scoring 0 despite a confident "task completed" report.
79Prescribed resampling procedure replaced by a different (more "extreme") testtaskda-code
Applies when
task -- the task/README explicitly dictates the inference procedure (e.g., shift each group to a common mean, then bootstrap each group separately and count replicates at least as extreme as the observed statistic) and the answer is a single p-value written to a template file.
Pattern
The attempt computes a p-value with a substitute method (parametric t-test, label-permutation test, pooled single-sample bootstrap, resampling only one group, or a one-sided/two-sided convention different from the one implied), and/or reports a Monte-Carlo p-value at or below the resolution floor of its replicate count (e.g., a handful of hits out of 10^5 draws) without checking that the number is stable or plausible.
Detection procedure
  1. From the task/README, write down the required steps: what statistic is observed, how the null is imposed (shifting vs. relabeling), which units are resampled, replicate count, and the tail definition.
  2. In the script, locate the null-generation code and check line-by-line that each group is resampled from its own shifted copy (not pooled/permuted), that the observed statistic is recomputed identically, and that the extremity comparison matches the stated tail and sign.
  3. Check the reported value against the simulation's granularity (p ≥ 1/replicates) and rerun-stability with a different seed; also compare order of magnitude to what the prescribed method plausibly yields versus a parametric shortcut.
  4. Verify the output file exists with the exact column name/format/precision from the template, containing the final p-value (not a count, a test statistic, or an unrounded scientific-notation surprise) — and that a script was saved so the procedure is reproducible.
Discriminator
A real violation is a null model or tail convention structurally different from the one specified (pooling/permuting instead of shifting, resampling one group, wrong comparison direction), or a p-value whose magnitude is an artifact of too few replicates; it is not a violation if the prescribed procedure is implemented and the p-value merely differs in the last digits due to random seed, provided it is well above 1/replicates and formatted as requested.
Consequence
The written file's p-value differs from the reference by orders of magnitude (or fails tolerance/format matching), so the result-file check fails even though the pipeline "ran successfully".
id 48a055917476 · mined from da-code dacode-data-sa-028@s2
raw text (what the judge reads)
### Prescribed resampling procedure replaced by a different (more "extreme") test
- **Applies when**: `task` -- the task/README explicitly dictates the inference procedure (e.g., shift each group to a common mean, then bootstrap each group separately and count replicates at least as extreme as the observed statistic) and the answer is a single p-value written to a template file.
- **Pattern**: The attempt computes a p-value with a substitute method (parametric t-test, label-permutation test, pooled single-sample bootstrap, resampling only one group, or a one-sided/two-sided convention different from the one implied), and/or reports a Monte-Carlo p-value at or below the resolution floor of its replicate count (e.g., a handful of hits out of 10^5 draws) without checking that the number is stable or plausible.
- **Detection procedure**:
  1. From the task/README, write down the required steps: what statistic is observed, how the null is imposed (shifting vs. relabeling), which units are resampled, replicate count, and the tail definition.
  2. In the script, locate the null-generation code and check line-by-line that each group is resampled from its own shifted copy (not pooled/permuted), that the observed statistic is recomputed identically, and that the extremity comparison matches the stated tail and sign.
  3. Check the reported value against the simulation's granularity (p ≥ 1/replicates) and rerun-stability with a different seed; also compare order of magnitude to what the prescribed method plausibly yields versus a parametric shortcut.
  4. Verify the output file exists with the exact column name/format/precision from the template, containing the final p-value (not a count, a test statistic, or an unrounded scientific-notation surprise) — and that a script was saved so the procedure is reproducible.
- **Discriminator**: A real violation is a null model or tail convention structurally different from the one specified (pooling/permuting instead of shifting, resampling one group, wrong comparison direction), or a p-value whose magnitude is an artifact of too few replicates; it is *not* a violation if the prescribed procedure is implemented and the p-value merely differs in the last digits due to random seed, provided it is well above 1/replicates and formatted as requested.
- **Consequence**: The written file's p-value differs from the reference by orders of magnitude (or fails tolerance/format matching), so the result-file check fails even though the pipeline "ran successfully".
80Ignoring provided specification/config files and their required output artifactstaskda-code
Applies when
task -- The prompt points to auxiliary instruction or configuration files (e.g., a tips/notes file, a YAML/JSON spec) that define methodology, styling, and/or required deliverables.
Pattern
The agent never opens or parses those files; it guesses the methodology and plotting parameters from the task wording, hard-codes its own choices (variable selection, grouping, labels, colors, figure size, axis ranges), and produces only the single obviously-named output while skipping companion artifacts the spec implies (serialized plot data, arrays, config echoes).
Detection procedure
  1. List every instruction/config file and every output artifact named or implied by the task statement.
  2. Search the scripts for code that reads each of those files (open/read/yaml.safe_load/json.load) and for code that writes each expected artifact.
  3. Check whether formatting/aggregation choices in the plotting or computation code are traceable to values loaded from the config, or are literals invented by the agent; check whether the answer text quotes the actual file contents versus paraphrasing assumptions.
  4. Verify the answer reports the same set of deliverables the task requires, not just one.
Discriminator
A real violation is when no code path ever reads the spec file(s) or when required outputs are missing entirely; it is fine if the agent read the files and then legitimately inlined their values (evidence: printed/quoted contents, parameter names matching the spec) and produced all requested artifacts.
Consequence
Graders comparing each expected artifact fail on missing files, and the produced figure mismatches the reference on styling and on the underlying aggregated series, giving 0 passed checks.
id 01a98f571698 · mined from da-code dacode-plot-line-006@s2
raw text (what the judge reads)
### Ignoring provided specification/config files and their required output artifacts
- **Applies when**: `task` -- The prompt points to auxiliary instruction or configuration files (e.g., a tips/notes file, a YAML/JSON spec) that define methodology, styling, and/or required deliverables.
- **Pattern**: The agent never opens or parses those files; it guesses the methodology and plotting parameters from the task wording, hard-codes its own choices (variable selection, grouping, labels, colors, figure size, axis ranges), and produces only the single obviously-named output while skipping companion artifacts the spec implies (serialized plot data, arrays, config echoes).
- **Detection procedure**:
  1. List every instruction/config file and every output artifact named or implied by the task statement.
  2. Search the scripts for code that reads each of those files (open/read/yaml.safe_load/json.load) and for code that writes each expected artifact.
  3. Check whether formatting/aggregation choices in the plotting or computation code are traceable to values loaded from the config, or are literals invented by the agent; check whether the answer text quotes the actual file contents versus paraphrasing assumptions.
  4. Verify the answer reports the same set of deliverables the task requires, not just one.
- **Discriminator**: A real violation is when no code path ever reads the spec file(s) or when required outputs are missing entirely; it is fine if the agent read the files and then legitimately inlined their values (evidence: printed/quoted contents, parameter names matching the spec) and produced all requested artifacts.
- **Consequence**: Graders comparing each expected artifact fail on missing files, and the produced figure mismatches the reference on styling and on the underlying aggregated series, giving 0 passed checks.
81Ignoring a task-referenced specification file that defines the required methodtaskda-code
Applies when
task -- The prompt points to an external document/notes file (e.g. a markdown/readme/spec in the working directory) that defines how to bin, group, filter, or compute the requested quantity.
Pattern
The scripts never open or echo the referenced file; the agent instead assumes the data's own native categories/defaults (e.g. pre-existing bucket labels or an ad-hoc ordering it invents) and produces outputs whose granularity, labels, or aggregation cannot be verified against the stated method.
Detection procedure
  1. List every external file, rule, or convention the task text references, plus every explicitly required output artifact/name.
  2. Grep the scripts for a read/print of each referenced file; confirm the derived categories/groups in the code are traceable to that file's content rather than to value_counts() of a raw column.
  3. Check the answer for evidence the specified method was applied (group boundaries/labels quoted from the spec) and that all required output files were written to the expected location.
  4. If the spec was never loaded, or produced groups differ in number/boundaries from anything in the spec, flag it.
Discriminator
Fine if the script demonstrably reads the spec (or quotes its rules) and the resulting groups match it — even if the raw column happens to be pre-binned; a violation is using dataset-native or self-invented categories with no reference to the mandated definition.
Consequence
Aggregated counts/labels differ from the reference binning, so the saved figure and any numeric arrays mismatch the expected artifacts and all checks fail, despite the plot having correct title/axis labels.
id cb2210c59cb1 · mined from da-code dacode-plot-bar-005@s2
raw text (what the judge reads)
### Ignoring a task-referenced specification file that defines the required method
- **Applies when**: `task` -- The prompt points to an external document/notes file (e.g. a markdown/readme/spec in the working directory) that defines how to bin, group, filter, or compute the requested quantity.
- **Pattern**: The scripts never open or echo the referenced file; the agent instead assumes the data's own native categories/defaults (e.g. pre-existing bucket labels or an ad-hoc ordering it invents) and produces outputs whose granularity, labels, or aggregation cannot be verified against the stated method.
- **Detection procedure**:
  1. List every external file, rule, or convention the task text references, plus every explicitly required output artifact/name.
  2. Grep the scripts for a read/print of each referenced file; confirm the derived categories/groups in the code are traceable to that file's content rather than to `value_counts()` of a raw column.
  3. Check the answer for evidence the specified method was applied (group boundaries/labels quoted from the spec) and that all required output files were written to the expected location.
  4. If the spec was never loaded, or produced groups differ in number/boundaries from anything in the spec, flag it.
- **Discriminator**: Fine if the script demonstrably reads the spec (or quotes its rules) and the resulting groups match it — even if the raw column happens to be pre-binned; a violation is using dataset-native or self-invented categories with no reference to the mandated definition.
- **Consequence**: Aggregated counts/labels differ from the reference binning, so the saved figure and any numeric arrays mismatch the expected artifacts and all checks fail, despite the plot having correct title/axis labels.
82Answer not persisted to the expected artifact with the exact requested schemataskda-code
Applies when
task -- the prompt shows a literal output template (e.g., keys whose values are bracketed lists) and/or the grading setup expects a specific result file produced by the agent's code.
Pattern
The attempt computes a plausible-looking value but only reports it in prose/chat, and/or reshapes the template (scalars instead of lists, renamed/extra/missing keys, added units or rounding not asked for), leaving no saved script or result file that a grader can read.
Detection procedure
  1. Re-read the task and write down the exact required deliverable: file name/location (if any), key names, and the value type shown in the template (list vs scalar, string vs number).
  2. Search the scripts for the code that serializes the final answer (e.g., a dump/write to the named file); if no script or no write step exists, the deliverable is missing.
  3. Diff the submitted object against the template key-by-key and type-by-type; confirm every key is present, spelled identically, and wrapped in the same container type.
  4. Check that the reported number is the final requested quantity in the requested units/precision, not an intermediate or reformatted variant.
Discriminator
A real violation is a structural mismatch (no file written, scalar where a list is shown, renamed/absent keys); harmless look-alikes are cosmetic differences the template does not constrain, such as key ordering, whitespace, or a value that is genuinely scalar because the template shows a scalar.
Consequence
The grader finds the expected result file missing or fails schema/type comparison, marking the submission wrong even if the underlying computation was right.
id 17f8edabbb6b · mined from da-code dacode-di-text-002@s2
raw text (what the judge reads)
### Answer not persisted to the expected artifact with the exact requested schema
- **Applies when**: `task` -- the prompt shows a literal output template (e.g., keys whose values are bracketed lists) and/or the grading setup expects a specific result file produced by the agent's code.
- **Pattern**: The attempt computes a plausible-looking value but only reports it in prose/chat, and/or reshapes the template (scalars instead of lists, renamed/extra/missing keys, added units or rounding not asked for), leaving no saved script or result file that a grader can read.
- **Detection procedure**:
  1. Re-read the task and write down the exact required deliverable: file name/location (if any), key names, and the value type shown in the template (list vs scalar, string vs number).
  2. Search the scripts for the code that serializes the final answer (e.g., a dump/write to the named file); if no script or no write step exists, the deliverable is missing.
  3. Diff the submitted object against the template key-by-key and type-by-type; confirm every key is present, spelled identically, and wrapped in the same container type.
  4. Check that the reported number is the final requested quantity in the requested units/precision, not an intermediate or reformatted variant.
- **Discriminator**: A real violation is a structural mismatch (no file written, scalar where a list is shown, renamed/absent keys); harmless look-alikes are cosmetic differences the template does not constrain, such as key ordering, whitespace, or a value that is genuinely scalar because the template shows a scalar.
- **Consequence**: The grader finds the expected result file missing or fails schema/type comparison, marking the submission wrong even if the underlying computation was right.
83Skipping input-integrity validation before building a derived/cumulative seriestaskda-code
Applies when
task -- the script reads a raw table and immediately aggregates it (weighted sums, compounding, cumulative products) into an output series that is graded against exact expected values.
Pattern
The attempt loads the file and jumps straight to arithmetic without checking for missing/blank cells, non-numeric dtypes, duplicate or unsorted keys, extra/renamed columns, or the scale/units of the values (e.g., fractions vs. percentages, levels vs. period-over-period changes). Any NaN, string column, off-by-100 scaling, or first-row convention silently propagates through the cumulative operation and corrupts every downstream row, and the attempt's "verification" script only re-derives its own numbers instead of testing them against an independent expectation.
Detection procedure
  1. Read the task/README for what the input columns are supposed to represent (units, scale, whether a base/first row exists) and what the output must contain (column names, ordering, index, rounding, row count).
  2. Scan the scripts for explicit integrity checks before the aggregation: isna().sum(), dtype checks, min/max or magnitude checks, row-count/date-continuity checks, handling of the first period. If aggregation happens with none of these, flag it.
  3. Check whether the "verification" step compares to anything independent (a known ground-truth magnitude, a hand-computed row, a second method) or merely recomputes the same formula.
  4. Inspect the reported output: are the values in a plausible range and magnitude for the stated quantity, is the row count/first row consistent with the input period, and do the column names/order match the required format exactly?
Discriminator
A real violation is when nothing in the pipeline could have detected a NaN, a mis-scaled column, or a wrong first-row/base convention, and no external cross-check exists. It is not a violation if the script explicitly inspects/handles missing values and units (or documents that the data is clean after checking) and validates at least one output value against an independent reference, even if the code is otherwise terse.
Consequence
The saved file has the right shape and looks superficially reasonable, but every value is systematically off (shifted, scaled, or NaN-contaminated from the first defective row onward), so an exact/tolerance comparison against the expected file fails on all checks.
id 4442748969bb · mined from da-code dacode-dm-csv-050@s2
raw text (what the judge reads)
### Skipping input-integrity validation before building a derived/cumulative series
- **Applies when**: `task` -- the script reads a raw table and immediately aggregates it (weighted sums, compounding, cumulative products) into an output series that is graded against exact expected values.
- **Pattern**: The attempt loads the file and jumps straight to arithmetic without checking for missing/blank cells, non-numeric dtypes, duplicate or unsorted keys, extra/renamed columns, or the scale/units of the values (e.g., fractions vs. percentages, levels vs. period-over-period changes). Any NaN, string column, off-by-100 scaling, or first-row convention silently propagates through the cumulative operation and corrupts every downstream row, and the attempt's "verification" script only re-derives its own numbers instead of testing them against an independent expectation.
- **Detection procedure**:
  1. Read the task/README for what the input columns are supposed to represent (units, scale, whether a base/first row exists) and what the output must contain (column names, ordering, index, rounding, row count).
  2. Scan the scripts for explicit integrity checks before the aggregation: `isna().sum()`, dtype checks, min/max or magnitude checks, row-count/date-continuity checks, handling of the first period. If aggregation happens with none of these, flag it.
  3. Check whether the "verification" step compares to anything independent (a known ground-truth magnitude, a hand-computed row, a second method) or merely recomputes the same formula.
  4. Inspect the reported output: are the values in a plausible range and magnitude for the stated quantity, is the row count/first row consistent with the input period, and do the column names/order match the required format exactly?
- **Discriminator**: A real violation is when nothing in the pipeline could have detected a NaN, a mis-scaled column, or a wrong first-row/base convention, and no external cross-check exists. It is *not* a violation if the script explicitly inspects/handles missing values and units (or documents that the data is clean after checking) and validates at least one output value against an independent reference, even if the code is otherwise terse.
- **Consequence**: The saved file has the right shape and looks superficially reasonable, but every value is systematically off (shifted, scaled, or NaN-contaminated from the first defective row onward), so an exact/tolerance comparison against the expected file fails on all checks.
84Silently substituting a random subsample (or otherwise altered input) for the statistic the task specifiedtaskinfiagent-dabench
Applies when
task -- the task names a specific statistical test/metric to run on a specified column or dataset, and the script contains a branch that resamples, truncates, or otherwise changes the input before computing it (often to dodge a library size limit or runtime cost).
Pattern
The script computes the mandated test on a random subset (frequently with no fixed seed) or a differently-filtered set than the one used for the other reported statistics, then reports the resulting p-value/decision as if it came from the full specified data — making the answer non-reproducible and potentially opposite to the intended result.
Detection procedure
  1. Read the task and note exactly which data (which column, which rows, all of them?) the named test/statistic must be computed on, and any stated thresholds.
  2. Scan the script for any np.random.*, .sample(, slicing, head/tail, dropna-with-different-scope, or size-based if branches sitting between data loading and the mandated computation; check whether a seed is set and whether the same subset feeds every reported number.
  3. Check whether the script prints diagnostics that let one verify the input to the test (n used, missing count, min/max/unique values) and whether the script would raise/flag rather than silently deviate when the library's constraints are hit.
  4. Compare the reported decision with the other reported numbers and with the data's nature (e.g., few distinct integer values, small n, mild moment values) — if the decision hinges on a resampled subset or contradicts the descriptive statistics, treat it as unverified.
Discriminator
A real violation is a silent, unrequested change of the analysis input (or a nondeterministic one) affecting the reported figure; it is fine if the task itself authorizes subsetting/filtering, or if the full specified data is used and any subsetting is only for auxiliary plots/exploration and is seeded and documented.
Consequence
The reported test decision (and its p-value) reflects a different, randomly varying sample than the ground-truth computation, so the categorical answer flips (e.g., "no" instead of "yes") and the grader scores 0 even when the accompanying descriptive statistics look plausible.
id faa875efbf32 · mined from infiagent-dabench dabench-298@s2
raw text (what the judge reads)
### Silently substituting a random subsample (or otherwise altered input) for the statistic the task specified
- **Applies when**: `task` -- the task names a specific statistical test/metric to run on a specified column or dataset, and the script contains a branch that resamples, truncates, or otherwise changes the input before computing it (often to dodge a library size limit or runtime cost).
- **Pattern**: The script computes the mandated test on a random subset (frequently with no fixed seed) or a differently-filtered set than the one used for the other reported statistics, then reports the resulting p-value/decision as if it came from the full specified data — making the answer non-reproducible and potentially opposite to the intended result.
- **Detection procedure**:
  1. Read the task and note exactly which data (which column, which rows, all of them?) the named test/statistic must be computed on, and any stated thresholds.
  2. Scan the script for any `np.random.*`, `.sample(`, slicing, head/tail, dropna-with-different-scope, or size-based `if` branches sitting between data loading and the mandated computation; check whether a seed is set and whether the same subset feeds every reported number.
  3. Check whether the script prints diagnostics that let one verify the input to the test (n used, missing count, min/max/unique values) and whether the script would raise/flag rather than silently deviate when the library's constraints are hit.
  4. Compare the reported decision with the other reported numbers and with the data's nature (e.g., few distinct integer values, small n, mild moment values) — if the decision hinges on a resampled subset or contradicts the descriptive statistics, treat it as unverified.
- **Discriminator**: A real violation is a *silent, unrequested* change of the analysis input (or a nondeterministic one) affecting the reported figure; it is fine if the task itself authorizes subsetting/filtering, or if the full specified data is used and any subsetting is only for auxiliary plots/exploration and is seeded and documented.
- **Consequence**: The reported test decision (and its p-value) reflects a different, randomly varying sample than the ground-truth computation, so the categorical answer flips (e.g., "no" instead of "yes") and the grader scores 0 even when the accompanying descriptive statistics look plausible.
85Prediction file not verified against the test set (row count / alignment / persistence)taskda-code
Applies when
task -- the deliverable is a prediction file whose rows must correspond one-to-one, in order, with the rows of a provided evaluation input file.
Pattern
The agent pastes predictions into the chat answer instead of (or in addition to) writing the required file, and never checks that the written file has exactly as many rows as the input, in the same order, with the exact required column name and no extra index column. Truncated, reordered, deduplicated, or dropped-row outputs (e.g., after dropping missing/blank text) go unnoticed.
Detection procedure
  1. Read the task for the required output filename, column name(s), and the input file that defines the number and order of predictions.
  2. In the scripts, confirm the file is actually written to the required path (to_csv(..., index=False)) with the exact header, and that predictions were produced from the full, unfiltered, unshuffled input (no dropna, sample, head, sorting, or partial-batch loop that would change row count/order).
  3. Look for an explicit post-write sanity check: re-read the file and assert len(output) == len(input), header matches, and no nulls/unexpected label values.
  4. Inspect the final answer: if it consists of the predictions dumped as text (especially cut off mid-word) rather than a confirmation of a validated saved file, treat the artifact as unverified.
Discriminator
A real violation is missing/unwritten file, or absent verification of row count, order, header, and label vocabulary. It is fine if the agent writes the file and shows a check of shape/head/value counts, even if it also prints a preview of the predictions in the answer.
Consequence
The grader looks for the expected file and compares row-by-row; a missing, truncated, misaligned, or wrongly-headed file scores 0 regardless of model quality.
id 7fd1355dd264 · mined from da-code dacode-ml-multi-011@s2
raw text (what the judge reads)
### Prediction file not verified against the test set (row count / alignment / persistence)
- **Applies when**: `task` -- the deliverable is a prediction file whose rows must correspond one-to-one, in order, with the rows of a provided evaluation input file.
- **Pattern**: The agent pastes predictions into the chat answer instead of (or in addition to) writing the required file, and never checks that the written file has exactly as many rows as the input, in the same order, with the exact required column name and no extra index column. Truncated, reordered, deduplicated, or dropped-row outputs (e.g., after dropping missing/blank text) go unnoticed.
- **Detection procedure**:
  1. Read the task for the required output filename, column name(s), and the input file that defines the number and order of predictions.
  2. In the scripts, confirm the file is actually written to the required path (`to_csv(..., index=False)`) with the exact header, and that predictions were produced from the full, unfiltered, unshuffled input (no `dropna`, `sample`, `head`, sorting, or partial-batch loop that would change row count/order).
  3. Look for an explicit post-write sanity check: re-read the file and assert `len(output) == len(input)`, header matches, and no nulls/unexpected label values.
  4. Inspect the final answer: if it consists of the predictions dumped as text (especially cut off mid-word) rather than a confirmation of a validated saved file, treat the artifact as unverified.
- **Discriminator**: A real violation is missing/unwritten file, or absent verification of row count, order, header, and label vocabulary. It is fine if the agent writes the file and shows a check of shape/head/value counts, even if it also prints a preview of the predictions in the answer.
- **Consequence**: The grader looks for the expected file and compares row-by-row; a missing, truncated, misaligned, or wrongly-headed file scores 0 regardless of model quality.
86Final artifact produced by an unvalidated "fallback" model that overwrites better-validated worktaskda-code
Applies when
task -- the deliverable is a prediction/result file and the agent runs several successive scripts, each rewriting the same output file, with the last one being a simplified/faster variant.
Pattern
Earlier scripts hold out data and report validation scores for stronger models, but the final script that actually writes the deliverable drops validation entirely, fits a much weaker/simpler estimator on all data, and silently overwrites the previous output; the answer reports only descriptive statistics of the predictions (mean/min/max) as if they were evidence of quality, with no held-out error estimate or comparison against the earlier candidates.
Detection procedure
  1. Read the task to identify the required deliverable file and the implied evaluation criterion (accuracy/error of the predictions, not merely file existence).
  2. Identify which script last writes that file, and check whether it (a) computes any held-out or cross-validated score for the exact model/pipeline whose predictions are saved, and (b) is at least as strong as models validated in earlier scripts.
  3. Check the reported answer for a quantitative generalization metric tied to the shipped predictions and an explicit reason for choosing that model over the alternatives.
  4. Sanity-check the shipped predictions against the training target's distribution/range and the required id count, ordering, column names, and output path.
Discriminator
A real violation is when the shipped model was never scored on unseen data, or was demonstrably weaker than an already-validated alternative and the switch is justified only by speed/convenience. It is fine if a simpler model is shipped after being compared on the same validation protocol and shown to be competitive, or if refitting a validated configuration on the full data without re-scoring.
Consequence
The submission file is well-formed but its predictive quality falls below the grader's threshold (poor R²/high error versus the hidden labels), so the deliverable is marked wrong despite the run "completing successfully."
id ae915c5e4760 · mined from da-code dacode-ml-competition-008@s2
raw text (what the judge reads)
### Final artifact produced by an unvalidated "fallback" model that overwrites better-validated work
- **Applies when**: `task` -- the deliverable is a prediction/result file and the agent runs several successive scripts, each rewriting the same output file, with the last one being a simplified/faster variant.
- **Pattern**: Earlier scripts hold out data and report validation scores for stronger models, but the final script that actually writes the deliverable drops validation entirely, fits a much weaker/simpler estimator on all data, and silently overwrites the previous output; the answer reports only descriptive statistics of the predictions (mean/min/max) as if they were evidence of quality, with no held-out error estimate or comparison against the earlier candidates.
- **Detection procedure**:
  1. Read the task to identify the required deliverable file and the implied evaluation criterion (accuracy/error of the predictions, not merely file existence).
  2. Identify which script last writes that file, and check whether it (a) computes any held-out or cross-validated score for the exact model/pipeline whose predictions are saved, and (b) is at least as strong as models validated in earlier scripts.
  3. Check the reported answer for a quantitative generalization metric tied to the shipped predictions and an explicit reason for choosing that model over the alternatives.
  4. Sanity-check the shipped predictions against the training target's distribution/range and the required id count, ordering, column names, and output path.
- **Discriminator**: A real violation is when the shipped model was never scored on unseen data, or was demonstrably weaker than an already-validated alternative and the switch is justified only by speed/convenience. It is fine if a simpler model is shipped after being compared on the same validation protocol and shown to be competitive, or if refitting a validated configuration on the full data without re-scoring.
- **Consequence**: The submission file is well-formed but its predictive quality falls below the grader's threshold (poor R²/high error versus the hidden labels), so the deliverable is marked wrong despite the run "completing successfully."
87Ranked "top-N" lists produced without applying the stated qualification rule and rank orderingtaskda-code
Applies when
task -- the task asks for the top N entities by some aggregate statistic, and the prompt/README supplies a definition, eligibility threshold, or tie-breaking/ordering convention for that statistic.
Pattern
The attempt groups rows and sorts by a raw aggregate without honoring the stated qualification rule (e.g., minimum number of underlying records per entity, deduplication of repeated items, numeric parsing/cleaning of the value column), and/or emits the N names in an order that is not the ranking order (alphabetical or arbitrary), so ties at the ceiling value are resolved incorrectly.
Detection procedure
  1. Read the task/README and write down every constraint attached to the ranking: how the aggregate is defined, any eligibility filter, how ties are broken, and whether row 0 must be the top-ranked entity.
  2. In the scripts, locate the groupby/aggregation and check for (a) cleaning/casting of the ranked column, (b) an explicit filter implementing the eligibility rule, (c) sort_values(..., ascending=False) on the aggregate with a defined tie-breaker before head(N).
  3. Inspect the answer file: check whether any column's entries are in alphabetical order or whether the top entries look like one-off/low-volume entities (a sign the eligibility filter and rank ordering were skipped).
  4. Confirm the emitted aggregate values were sanity-checked (range, count of rows per entity, N rows exactly, column names/ordering matching the sample format).
Discriminator
A real violation is when no code implements the stated threshold/ordering, or the output order cannot be reproduced by sorting on the aggregate; a look-alike that is fine is a list that merely happens to be near-alphabetical because genuine ties were broken by a documented, implemented rule.
Consequence
The exact-match check on the saved file fails because the entity set and/or their row positions differ from the reference ranking.
id 938bf48c33a6 · mined from da-code dacode-dm-csv-009@s2
raw text (what the judge reads)
### Ranked "top-N" lists produced without applying the stated qualification rule and rank ordering
- **Applies when**: `task` -- the task asks for the top N entities by some aggregate statistic, and the prompt/README supplies a definition, eligibility threshold, or tie-breaking/ordering convention for that statistic.
- **Pattern**: The attempt groups rows and sorts by a raw aggregate without honoring the stated qualification rule (e.g., minimum number of underlying records per entity, deduplication of repeated items, numeric parsing/cleaning of the value column), and/or emits the N names in an order that is not the ranking order (alphabetical or arbitrary), so ties at the ceiling value are resolved incorrectly.
- **Detection procedure**:
  1. Read the task/README and write down every constraint attached to the ranking: how the aggregate is defined, any eligibility filter, how ties are broken, and whether row 0 must be the top-ranked entity.
  2. In the scripts, locate the groupby/aggregation and check for (a) cleaning/casting of the ranked column, (b) an explicit filter implementing the eligibility rule, (c) `sort_values(..., ascending=False)` on the aggregate with a defined tie-breaker before `head(N)`.
  3. Inspect the answer file: check whether any column's entries are in alphabetical order or whether the top entries look like one-off/low-volume entities (a sign the eligibility filter and rank ordering were skipped).
  4. Confirm the emitted aggregate values were sanity-checked (range, count of rows per entity, N rows exactly, column names/ordering matching the sample format).
- **Discriminator**: A real violation is when no code implements the stated threshold/ordering, or the output order cannot be reproduced by sorting on the aggregate; a look-alike that is fine is a list that merely *happens* to be near-alphabetical because genuine ties were broken by a documented, implemented rule.
- **Consequence**: The exact-match check on the saved file fails because the entity set and/or their row positions differ from the reference ranking.
88Fabricating input data instead of using the provided datasettaskda-code
Applies when
task -- the task refers to a supplied dataset (with a README/spec) and the agent's scripts include a step that generates, simulates, or hard-codes the input data.
Pattern
The agent cannot find or fails to load the real input file, so it synthesizes a random/placeholder table with invented columns and value ranges, then runs the full analysis on that fake data and reports the resulting numbers as if they came from the real source. Related symptom: only a subset of required output artifacts is produced, and stated spec files are only partially honored.
Detection procedure
1. Read the task and note which input files and which output artifacts (files, formats, names) are expected. 2. Scan every script for data creation calls (np.random.*, range(...) fillers, manual dicts/lists written to the input path, seed) or writes to the same path later read as input. 3. Check whether the script instead locates and loads the real provided file (and would fail loudly if absent) and whether all required outputs are written. 4. Cross-check the answer's reported counts/ranges against the dataset description (row counts, units, realistic value ranges) and against the mentioned spec/config keys.
Discriminator
A real violation is when the analytical result reported comes from data the agent itself invented, or when the real file was never read. Legitimate look-alikes: creating tiny synthetic fixtures for unit-testing a plotting/utility function while the reported result still comes from the real dataset, or augmenting real data in a documented, task-sanctioned way.
Consequence
All value-based and file-based checks fail: derived arrays/JSON summaries and the figure encode arbitrary random counts (and required artifacts may be missing entirely), so the grader reports 0 of the expected outputs correct.
id bbec7f72c429 · mined from da-code dacode-plot-bar-007@s2
raw text (what the judge reads)
### Fabricating input data instead of using the provided dataset
- **Applies when**: `task` -- the task refers to a supplied dataset (with a README/spec) and the agent's scripts include a step that generates, simulates, or hard-codes the input data.
- **Pattern**: The agent cannot find or fails to load the real input file, so it synthesizes a random/placeholder table with invented columns and value ranges, then runs the full analysis on that fake data and reports the resulting numbers as if they came from the real source. Related symptom: only a subset of required output artifacts is produced, and stated spec files are only partially honored.
- **Detection procedure**: 1. Read the task and note which input files and which output artifacts (files, formats, names) are expected. 2. Scan every script for data creation calls (`np.random.*`, `range(...)` fillers, manual dicts/lists written to the input path, `seed`) or writes to the same path later read as input. 3. Check whether the script instead locates and loads the real provided file (and would fail loudly if absent) and whether all required outputs are written. 4. Cross-check the answer's reported counts/ranges against the dataset description (row counts, units, realistic value ranges) and against the mentioned spec/config keys.
- **Discriminator**: A real violation is when the analytical result reported comes from data the agent itself invented, or when the real file was never read. Legitimate look-alikes: creating tiny synthetic fixtures for unit-testing a plotting/utility function while the reported result still comes from the real dataset, or augmenting real data in a documented, task-sanctioned way.
- **Consequence**: All value-based and file-based checks fail: derived arrays/JSON summaries and the figure encode arbitrary random counts (and required artifacts may be missing entirely), so the grader reports 0 of the expected outputs correct.
89Empty-subset result reported as prose instead of the required answer formattaskinfiagent-dabench
Applies when
task -- a task prescribes a specific answer template (e.g. @name[value]) for a statistic computed on a filtered subset, and the filter may match zero rows.
Pattern
The agent discovers the filter yields no rows (or an undefined statistic) and replaces the required formatted answer with an explanatory sentence ("no data available", "filter value does not exist"), instead of emitting the statistic's defined degenerate value (NaN/empty) inside the requested format.
Detection procedure
  1. Read the task and note the exact required output token/format and rounding rules.
  2. Read the scripts for a branch that short-circuits when the filtered frame is empty (or when the aggregate returns NaN) and prints a message rather than the formatted value; check whether the filter/dtype (e.g. string vs numeric key) was verified before concluding the subset is empty.
  3. Compare the submitted answer against the required template: does it literally contain the token and a value slot?
  4. If it does not, flag it — an empty subset is a valid result to report, not a reason to abandon the format.
Discriminator
A real violation is any answer that abandons the template; it is fine to report an unusual value (NaN, 0, empty) inside the template, and it is also fine to note the emptiness as extra commentary alongside a properly formatted answer. Also not a violation if the emptiness stemmed from a genuine miscoded filter that, once fixed, yields data — that is a different (filtering) error.
Consequence
The grader's field-by-field check finds no parsable value for the requested key, so the expected result (including NaN) is scored WRONG/MISSING and the task fails at 0/1.
id c6a27ce88566 · mined from infiagent-dabench dabench-554@s2
raw text (what the judge reads)
### Empty-subset result reported as prose instead of the required answer format
- **Applies when**: `task` -- a task prescribes a specific answer template (e.g. `@name[value]`) for a statistic computed on a filtered subset, and the filter may match zero rows.
- **Pattern**: The agent discovers the filter yields no rows (or an undefined statistic) and replaces the required formatted answer with an explanatory sentence ("no data available", "filter value does not exist"), instead of emitting the statistic's defined degenerate value (NaN/empty) inside the requested format.
- **Detection procedure**:
  1. Read the task and note the exact required output token/format and rounding rules.
  2. Read the scripts for a branch that short-circuits when the filtered frame is empty (or when the aggregate returns NaN) and prints a message rather than the formatted value; check whether the filter/dtype (e.g. string vs numeric key) was verified before concluding the subset is empty.
  3. Compare the submitted answer against the required template: does it literally contain the token and a value slot?
  4. If it does not, flag it — an empty subset is a valid result to report, not a reason to abandon the format.
- **Discriminator**: A real violation is any answer that abandons the template; it is fine to report an unusual value (NaN, 0, empty) *inside* the template, and it is also fine to note the emptiness as extra commentary alongside a properly formatted answer. Also not a violation if the emptiness stemmed from a genuine miscoded filter that, once fixed, yields data — that is a different (filtering) error.
- **Consequence**: The grader's field-by-field check finds no parsable value for the requested key, so the expected result (including NaN) is scored WRONG/MISSING and the task fails at 0/1.
90Unverifiable answer: entity names/values not traced back to the provided data via runnable codetaskda-code
Applies when
task -- the task asks for specific records (top/bottom-N names, IDs, categories) or statistics to be extracted from a supplied data file after a prescribed preprocessing step.
Pattern
The submission presents a plausible-looking list that reflects general/world knowledge or an unsaved ad-hoc computation, with no script that (a) loads the given file, (b) applies the stated preprocessing (e.g., the specified imputation), (c) sorts by the requested field in the stated direction, and (d) writes the answer to the required output file; entity labels are also not copied verbatim from the data's key column.
Detection procedure
  1. Read the task and note the required output artifact/format, the mandated preprocessing, and the ordering/sorting constraint for every requested list.
  2. Look for a script that reads the provided file and prints/dumps exactly the reported values; if no script exists, or the script's output was never shown to match the submitted answer, treat the answer as unverified.
  3. Cross-check each reported label against the data's identifier column spelling/format (e.g., official vs. colloquial names, punctuation, abbreviations) and check that each requested list is ordered as instructed, not by a default or intuitive order.
  4. Sanity-check the numbers behind the picks (count = N, values within plausible range, no rows dropped instead of imputed) — if the underlying values are not reported at all, that is itself a flag.
Discriminator
A real violation is an answer whose values cannot be reproduced from the file by any shown code, or whose labels differ from the dataset's own strings/ordering rule; a look-alike that is fine is an answer that happens to match common knowledge but is accompanied by a script whose printed output and written result file match it exactly and whose labels are taken verbatim from the data.
Consequence
The grader's exact comparison against the expected result file fails on missing/misnamed entities, wrong list order, or a missing output file, scoring 0 even though the list looks superficially reasonable.
id 6f4b40ed7ee2 · mined from da-code dacode-di-text-003@s2
raw text (what the judge reads)
### Unverifiable answer: entity names/values not traced back to the provided data via runnable code
- **Applies when**: `task` -- the task asks for specific records (top/bottom-N names, IDs, categories) or statistics to be extracted from a supplied data file after a prescribed preprocessing step.
- **Pattern**: The submission presents a plausible-looking list that reflects general/world knowledge or an unsaved ad-hoc computation, with no script that (a) loads the given file, (b) applies the stated preprocessing (e.g., the specified imputation), (c) sorts by the requested field in the stated direction, and (d) writes the answer to the required output file; entity labels are also not copied verbatim from the data's key column.
- **Detection procedure**:
  1. Read the task and note the required output artifact/format, the mandated preprocessing, and the ordering/sorting constraint for *every* requested list.
  2. Look for a script that reads the provided file and prints/dumps exactly the reported values; if no script exists, or the script's output was never shown to match the submitted answer, treat the answer as unverified.
  3. Cross-check each reported label against the data's identifier column spelling/format (e.g., official vs. colloquial names, punctuation, abbreviations) and check that each requested list is ordered as instructed, not by a default or intuitive order.
  4. Sanity-check the numbers behind the picks (count = N, values within plausible range, no rows dropped instead of imputed) — if the underlying values are not reported at all, that is itself a flag.
- **Discriminator**: A real violation is an answer whose values cannot be reproduced from the file by any shown code, or whose labels differ from the dataset's own strings/ordering rule; a look-alike that is fine is an answer that happens to match common knowledge *but* is accompanied by a script whose printed output and written result file match it exactly and whose labels are taken verbatim from the data.
- **Consequence**: The grader's exact comparison against the expected result file fails on missing/misnamed entities, wrong list order, or a missing output file, scoring 0 even though the list looks superficially reasonable.
91Answer string not emitted in the literal template (verbatim tokens, quoting, order) with no reproducible script backing ittaskinfiagent-dabench
Applies when
task -- The task specifies an exact answer template with named tags, delimiters, and example value formatting, and expects the values to be produced by saved, re-runnable analysis code.
Pattern
The agent computes conceptually correct values but serializes them in a variant of the requested template — dropping the quotation marks/units shown in the example, renaming or reordering tags, changing separators, adding prose around the tags — and/or leaves no script that prints the final answer string, so nothing ever validated the emitted text against the template. A literal string-matching grader then fails every field even though the underlying analysis was right.
Detection procedure
  1. Copy the answer template exactly as given in the task, including every quote, bracket, tag name, separator, and the formatting shown in any example value.
  2. Check the scripts: is there a step that constructs and prints the final answer string, so the submitted text is a program output rather than hand-typed? If no script is saved at all, the answer is unverifiable by construction — flag it.
  3. Diff the submitted answer character-by-character against the template: tag names and order, presence/absence of quotes around each value, capitalization and spelling of the allowed value vocabulary, and the exact form of any range/unit string.
  4. Confirm each value is one of the permitted options and is the final requested quantity, not an intermediate.
Discriminator
A real violation is any character-level deviation from the specified template (missing quotes, extra text, altered tag names/order), or an answer with no code that produced it; a look-alike that is fine is a template where the task itself shows the value unquoted or explicitly allows whitespace variation, and the answer matches that shown form exactly.
Consequence
A regex/exact-match grader reports every field as WRONG/MISSING even though the submitted values are semantically identical to ground truth, yielding a 0/N score with no partial credit.
id 3722a8c2393a · mined from infiagent-dabench dabench-550@s2
raw text (what the judge reads)
### Answer string not emitted in the literal template (verbatim tokens, quoting, order) with no reproducible script backing it
- **Applies when**: `task` -- The task specifies an exact answer template with named tags, delimiters, and example value formatting, and expects the values to be produced by saved, re-runnable analysis code.
- **Pattern**: The agent computes conceptually correct values but serializes them in a variant of the requested template — dropping the quotation marks/units shown in the example, renaming or reordering tags, changing separators, adding prose around the tags — and/or leaves no script that prints the final answer string, so nothing ever validated the emitted text against the template. A literal string-matching grader then fails every field even though the underlying analysis was right.
- **Detection procedure**:
  1. Copy the answer template exactly as given in the task, including every quote, bracket, tag name, separator, and the formatting shown in any example value.
  2. Check the scripts: is there a step that constructs and prints the final answer string, so the submitted text is a program output rather than hand-typed? If no script is saved at all, the answer is unverifiable by construction — flag it.
  3. Diff the submitted answer character-by-character against the template: tag names and order, presence/absence of quotes around each value, capitalization and spelling of the allowed value vocabulary, and the exact form of any range/unit string.
  4. Confirm each value is one of the permitted options and is the final requested quantity, not an intermediate.
- **Discriminator**: A real violation is any character-level deviation from the specified template (missing quotes, extra text, altered tag names/order), or an answer with no code that produced it; a look-alike that is fine is a template where the task itself shows the value unquoted or explicitly allows whitespace variation, and the answer matches that shown form exactly.
- **Consequence**: A regex/exact-match grader reports every field as WRONG/MISSING even though the submitted values are semantically identical to ground truth, yielding a 0/N score with no partial credit.
92Substituting proxy data/entities for the ones the task namestaskda-code
Applies when
task -- the prompt asks for specific entities, groupings, or measures (e.g., "top N of X by Y, broken down by stage Z") and the scripts must locate those fields in the provided data files.
Pattern
The agent cannot find the requested fields in the files it opened, so it silently swaps in a different data source and re-interprets each requested concept as a loose "analogy" (different grouping key, different ranking measure, different segment definitions), then declares success instead of resolving the mismatch.
Detection procedure
1. List from the task the exact entity to rank, the ranking measure, and the segments/values to plot. 2. In the scripts, identify which file and which columns supply each of those three items. 3. Flag if any is drawn from a different domain or is described in the answer as an "analogy"/"proxy"/"equivalent", or if the reported category labels are not instances of the requested entity type. 4. Check the answer's reported units/axis labels match the requested measure (e.g., a duration in days, not a count or probability).
Discriminator
A real violation is using a different variable or dataset than the task names because the requested one was not found; acceptable look-alikes are cases where the requested concept genuinely exists under a differently-spelled column name and the mapping is stated and verifiable (same semantics, same units).
Consequence
Every derived artifact (figure, saved arrays, config-driven outputs) encodes the wrong categories and wrong quantities, so all value- and label-level checks fail even though a chart of the right visual type was produced.
id 7d4eac7f01cd · mined from da-code dacode-plot-scatter-002@s2
raw text (what the judge reads)
### Substituting proxy data/entities for the ones the task names
- **Applies when**: `task` -- the prompt asks for specific entities, groupings, or measures (e.g., "top N of X by Y, broken down by stage Z") and the scripts must locate those fields in the provided data files.
- **Pattern**: The agent cannot find the requested fields in the files it opened, so it silently swaps in a different data source and re-interprets each requested concept as a loose "analogy" (different grouping key, different ranking measure, different segment definitions), then declares success instead of resolving the mismatch.
- **Detection procedure**: 1. List from the task the exact entity to rank, the ranking measure, and the segments/values to plot. 2. In the scripts, identify which file and which columns supply each of those three items. 3. Flag if any is drawn from a different domain or is described in the answer as an "analogy"/"proxy"/"equivalent", or if the reported category labels are not instances of the requested entity type. 4. Check the answer's reported units/axis labels match the requested measure (e.g., a duration in days, not a count or probability).
- **Discriminator**: A real violation is using a different variable or dataset than the task names because the requested one was not found; acceptable look-alikes are cases where the requested concept genuinely exists under a differently-spelled column name and the mapping is stated and verifiable (same semantics, same units).
- **Consequence**: Every derived artifact (figure, saved arrays, config-driven outputs) encodes the wrong categories and wrong quantities, so all value- and label-level checks fail even though a chart of the right visual type was produced.
93Answer-format template not reproduced literally (delimiters/separators dropped)taskinfiagent-dabench
Applies when
task -- the prompt specifies an exact answer string template (a tag, an = or : separator, and specific wrapping brackets/braces) for reporting one or more computed values.
Pattern
The attempt computes the values correctly but emits them in a paraphrased wrapper — omitting or substituting the required leading/trailing delimiters, the key-name-to-value separator, or the bracket nesting — so an exact-match grader fails even though the analysis is right.
Detection procedure
  1. Copy the answer template exactly as given in the task and mark every literal token: tag name, @/=/: separators, and each opening/closing bracket or brace in order.
  2. Read the scripts (or final message construction) and check whether the output string is built from that literal template — e.g. an f-string/print that hard-codes the tag and delimiters — rather than from a bare dict/list repr.
  3. Compare the submitted answer token-by-token against the template: same tag, same separator, same bracket sequence and nesting, same key spelling/order.
  4. Flag if any literal token is missing, added, or swapped, even when the numeric values look plausible.
Discriminator
A real violation is a structural mismatch in the required literal tokens (missing =, wrong bracket type, missing wrapper, renamed/reordered keys); harmless look-alikes are cosmetic differences the template does not constrain, such as inner whitespace, quote style, or int-vs-float rendering of the same value.
Consequence
The grader reports the expected variable as WRONG/MISSING and scores 0 despite numerically correct values, because it cannot parse the submitted string against the required pattern.
id f83a462bf083 · mined from infiagent-dabench dabench-451@s2
raw text (what the judge reads)
### Answer-format template not reproduced literally (delimiters/separators dropped)
- **Applies when**: `task` -- the prompt specifies an exact answer string template (a tag, an `=` or `:` separator, and specific wrapping brackets/braces) for reporting one or more computed values.
- **Pattern**: The attempt computes the values correctly but emits them in a paraphrased wrapper — omitting or substituting the required leading/trailing delimiters, the key-name-to-value separator, or the bracket nesting — so an exact-match grader fails even though the analysis is right.
- **Detection procedure**:
  1. Copy the answer template exactly as given in the task and mark every literal token: tag name, `@`/`=`/`:` separators, and each opening/closing bracket or brace in order.
  2. Read the scripts (or final message construction) and check whether the output string is built from that literal template — e.g. an f-string/print that hard-codes the tag and delimiters — rather than from a bare `dict`/`list` repr.
  3. Compare the submitted answer token-by-token against the template: same tag, same separator, same bracket sequence and nesting, same key spelling/order.
  4. Flag if any literal token is missing, added, or swapped, even when the numeric values look plausible.
- **Discriminator**: A real violation is a structural mismatch in the required literal tokens (missing `=`, wrong bracket type, missing wrapper, renamed/reordered keys); harmless look-alikes are cosmetic differences the template does not constrain, such as inner whitespace, quote style, or int-vs-float rendering of the same value.
- **Consequence**: The grader reports the expected variable as WRONG/MISSING and scores 0 despite numerically correct values, because it cannot parse the submitted string against the required pattern.
94Unverified prediction file (no reproducible script, no shape/format sanity check)taskda-code
Applies when
task -- the deliverable is a prediction/output file with a specified column name and one row per record of a held-out input file.
Pattern
The attempt produces the output file without any saved, runnable script that reads the held-out input, fits on the training portion, and writes predictions — and never checks that the written file has exactly the required column name, the same number of rows as the held-out input, the original row order, and values in a plausible range/dtype for the target.
Detection procedure
  1. Read the task statement and note the exact required file name, column name(s), row count (from the held-out input), and any ordering/rounding/units constraints.
  2. Look for a script that is complete end-to-end (load train + held-out data → preprocess → fit → predict → write file); if scripts are absent or only partially cover this path, the result cannot be reproduced or audited and must be treated as unverified.
  3. Inspect the produced file's header and shape: does it contain the requested column spelled exactly as asked (no extra index column, no renamed/extra columns), and does its row count equal the held-out input's row count?
  4. Check value sanity: dtype numeric (or as required), no NaNs/empty cells, values inside the target's observed range, and non-constant/non-degenerate predictions aligned to the input row order.
Discriminator
A real violation is missing/unreproducible code or a file whose header, row count, ordering, or value range does not match the specification. A look-alike that is fine: a script that differs stylistically or uses a simple baseline model, but demonstrably writes the exact requested column, one row per held-out record in input order, with valid values.
Consequence
The grader compares the submitted file against the expected schema/row alignment and marks it WRONG/MISSING (0 checks passed) even though a file with the right name exists.
id 5d3480fe4243 · mined from da-code dacode-ml-regression-004@s2
raw text (what the judge reads)
### Unverified prediction file (no reproducible script, no shape/format sanity check)
- **Applies when**: `task` -- the deliverable is a prediction/output file with a specified column name and one row per record of a held-out input file.
- **Pattern**: The attempt produces the output file without any saved, runnable script that reads the held-out input, fits on the training portion, and writes predictions — and never checks that the written file has exactly the required column name, the same number of rows as the held-out input, the original row order, and values in a plausible range/dtype for the target.
- **Detection procedure**:
  1. Read the task statement and note the exact required file name, column name(s), row count (from the held-out input), and any ordering/rounding/units constraints.
  2. Look for a script that is complete end-to-end (load train + held-out data → preprocess → fit → predict → write file); if scripts are absent or only partially cover this path, the result cannot be reproduced or audited and must be treated as unverified.
  3. Inspect the produced file's header and shape: does it contain the requested column spelled exactly as asked (no extra index column, no renamed/extra columns), and does its row count equal the held-out input's row count?
  4. Check value sanity: dtype numeric (or as required), no NaNs/empty cells, values inside the target's observed range, and non-constant/non-degenerate predictions aligned to the input row order.
- **Discriminator**: A real violation is missing/unreproducible code *or* a file whose header, row count, ordering, or value range does not match the specification. A look-alike that is fine: a script that differs stylistically or uses a simple baseline model, but demonstrably writes the exact requested column, one row per held-out record in input order, with valid values.
- **Consequence**: The grader compares the submitted file against the expected schema/row alignment and marks it WRONG/MISSING (0 checks passed) even though a file with the right name exists.
95Feature scaling omitted before distance-based clustering, yielding degenerate outlier clusterstaskda-code
Applies when
task -- the task asks for an unsupervised grouping (k-means/hierarchical/DBSCAN-style) of records whose numeric columns have wildly different units and ranges, and the deliverable is a label file.
Pattern
The attempt feeds raw, unstandardized columns straight into a distance-based algorithm (and/or picks k by a score computed on those raw features), so one or two large-magnitude columns dominate the distance metric; the resulting partition contains singleton or near-singleton clusters that merely isolate extreme values, while the bulk of the records collapse into one or two huge groups. The attempt reports this as the "optimal" solution without a sanity check on cluster sizes or on whether the grouping reflects all features.
Detection procedure
  1. In the task/README, note the feature columns and their plausible scales (percentages, per-capita monetary values, rates) — flag if magnitudes differ by orders of magnitude.
  2. In the scripts, check whether a scaler/normalizer (or a distance metric that is scale-invariant) is applied to the feature matrix before fitting and before any cluster-count selection metric; also check that model selection uses the same transformed space.
  3. In the reported answer, inspect the cluster size distribution and cluster descriptions: singleton/2–3-member clusters described as "outliers" or clusters characterized by a single dominant variable are strong evidence of unscaled distances.
  4. Confirm the saved label file was actually verified (row count, column names, label range) rather than only asserted in prose.
Discriminator
A real violation is scaling never applied (or applied only after clustering / only for plots) together with a lopsided partition dominated by high-variance columns. It is not a violation if the script standardizes (or the algorithm/metric is scale-free) and small clusters are then justified by inspection — genuinely extreme records can legitimately form small clusters in a properly scaled space.
Consequence
The submitted label column encodes an outlier split rather than the intended socio-economic-style grouping, so the expected output file comparison (cluster structure/agreement with reference labels) fails even though the file format looks plausible.
id a56e2e2c47e9 · mined from da-code dacode-ml-cluster-013@s2
raw text (what the judge reads)
### Feature scaling omitted before distance-based clustering, yielding degenerate outlier clusters
- **Applies when**: `task` -- the task asks for an unsupervised grouping (k-means/hierarchical/DBSCAN-style) of records whose numeric columns have wildly different units and ranges, and the deliverable is a label file.
- **Pattern**: The attempt feeds raw, unstandardized columns straight into a distance-based algorithm (and/or picks *k* by a score computed on those raw features), so one or two large-magnitude columns dominate the distance metric; the resulting partition contains singleton or near-singleton clusters that merely isolate extreme values, while the bulk of the records collapse into one or two huge groups. The attempt reports this as the "optimal" solution without a sanity check on cluster sizes or on whether the grouping reflects all features.
- **Detection procedure**:
  1. In the task/README, note the feature columns and their plausible scales (percentages, per-capita monetary values, rates) — flag if magnitudes differ by orders of magnitude.
  2. In the scripts, check whether a scaler/normalizer (or a distance metric that is scale-invariant) is applied to the feature matrix *before* fitting and before any cluster-count selection metric; also check that model selection uses the same transformed space.
  3. In the reported answer, inspect the cluster size distribution and cluster descriptions: singleton/2–3-member clusters described as "outliers" or clusters characterized by a single dominant variable are strong evidence of unscaled distances.
  4. Confirm the saved label file was actually verified (row count, column names, label range) rather than only asserted in prose.
- **Discriminator**: A real violation is scaling never applied (or applied only after clustering / only for plots) together with a lopsided partition dominated by high-variance columns. It is *not* a violation if the script standardizes (or the algorithm/metric is scale-free) and small clusters are then justified by inspection — genuinely extreme records can legitimately form small clusters in a properly scaled space.
- **Consequence**: The submitted label column encodes an outlier split rather than the intended socio-economic-style grouping, so the expected output file comparison (cluster structure/agreement with reference labels) fails even though the file format looks plausible.
96Stated scope/filter in the task is not reflected in the computed statistic (wrong axis/subset)taskinfiagent-dabench
Applies when
task -- The question restricts the statistic to a specific slice (a given year, group, region, or subset of columns/rows) and the scripts compute an aggregate statistic per entity.
Pattern
The script ignores the stated restriction and aggregates over the full set of columns/rows (e.g., every period rather than the specified one), or aggregates along the wrong axis, so the reported ranking answers a different question than the one asked. Extra "verification" code re-checks the same wrong computation, giving false confidence.
Detection procedure
  1. From the task text, list every explicit scope word (year, subset, grouping, units, estimator/definition flag) that must appear as a filter or parameter in the code.
  2. Read the script and locate where each scope word is applied: is there a selection of the specific column/rows before the statistic, and is the statistic taken along the axis implied by the question?
  3. If a scope word has no corresponding filter (e.g., all period columns are passed into the per-row statistic, or all rows into a per-column statistic), flag it; also check whether the resulting sample size/shape is what the stated slice would produce.
  4. Confirm the sanity check: does the intermediate printout show data only from the requested slice, and is the final reported quantity the one named in the question (not an intermediate or differently-scoped one)?
Discriminator
A real violation is when the requested slice is never selected anywhere in the pipeline (or is selected but then discarded/overwritten). It is not a violation when the slice is implicit because the loaded file/frame is already restricted to it, or when the statistic legitimately requires the broader sample and the slice only defines the grouping key — in those cases the code should still show an explicit filter or a comment plus a shape/count check matching the slice.
Consequence
The ranking/argmax is computed over the wrong sample, so the reported entity differs from the ground truth and the grader marks the single expected value as WRONG, even though the estimator flag and output format are correct.
id e8a3e51f01ca · mined from infiagent-dabench dabench-252@s2
raw text (what the judge reads)
### Stated scope/filter in the task is not reflected in the computed statistic (wrong axis/subset)
- **Applies when**: `task` -- The question restricts the statistic to a specific slice (a given year, group, region, or subset of columns/rows) and the scripts compute an aggregate statistic per entity.
- **Pattern**: The script ignores the stated restriction and aggregates over the full set of columns/rows (e.g., every period rather than the specified one), or aggregates along the wrong axis, so the reported ranking answers a different question than the one asked. Extra "verification" code re-checks the same wrong computation, giving false confidence.
- **Detection procedure**:
  1. From the task text, list every explicit scope word (year, subset, grouping, units, estimator/definition flag) that must appear as a filter or parameter in the code.
  2. Read the script and locate where each scope word is applied: is there a selection of the specific column/rows before the statistic, and is the statistic taken along the axis implied by the question?
  3. If a scope word has no corresponding filter (e.g., all period columns are passed into the per-row statistic, or all rows into a per-column statistic), flag it; also check whether the resulting sample size/shape is what the stated slice would produce.
  4. Confirm the sanity check: does the intermediate printout show data only from the requested slice, and is the final reported quantity the one named in the question (not an intermediate or differently-scoped one)?
- **Discriminator**: A real violation is when the requested slice is never selected anywhere in the pipeline (or is selected but then discarded/overwritten). It is *not* a violation when the slice is implicit because the loaded file/frame is already restricted to it, or when the statistic legitimately requires the broader sample and the slice only defines the grouping key — in those cases the code should still show an explicit filter or a comment plus a shape/count check matching the slice.
- **Consequence**: The ranking/argmax is computed over the wrong sample, so the reported entity differs from the ground truth and the grader marks the single expected value as WRONG, even though the estimator flag and output format are correct.
97Reformatting an identifier to a coarser granularity than the data (losing precision in the reported key)taskinfiagent-dabench
Applies when
task -- The task asks to identify a specific record/key (a date, ID, category) from row-level data and the answer template shows a format string that is coarser or ambiguous relative to the data's granularity.
Pattern
The script correctly locates the extremum/target row, but then applies a formatting/truncation step (e.g., strftime to a shorter pattern, rounding, taking a prefix/substring) that discards the identifying detail, so the reported key no longer uniquely designates the row actually used for the downstream computation.
Detection procedure
  1. Read the task: determine the granularity of the entity being identified (one row / one timestamp / one ID) and note whether the downstream calculation depends on that exact row.
  2. Read the script: find where the identified key is converted for output; check whether the emitted string contains strictly less information than the key stored in the data.
  3. Compare the emitted key against the value used internally for the dependent computation — if the dependent computation uses the full-precision key but the answer reports a truncated one, flag it; prefer reporting the key exactly as it appears in the source (or, if the template is genuinely ambiguous, report the full-precision value rather than truncating).
  4. Check for collisions: would the truncated key match multiple rows in the data? If yes, it cannot be the intended unique answer.
Discriminator
A real violation is when truncation destroys uniqueness or drops detail present in the source key while the rest of the answer depends on that detail. It is fine when the task explicitly aggregates at the coarser level (e.g., the maximum is computed over monthly aggregates) or when the source key itself has only that granularity.
Consequence
The numeric part of the answer can be correct while the identifier check fails, giving a partial-credit/incorrect verdict (e.g., 1 of 2 checks passed).
id 835c6e408b56 · mined from infiagent-dabench dabench-572@s2
raw text (what the judge reads)
### Reformatting an identifier to a coarser granularity than the data (losing precision in the reported key)
- **Applies when**: `task` -- The task asks to identify a specific record/key (a date, ID, category) from row-level data and the answer template shows a format string that is coarser or ambiguous relative to the data's granularity.
- **Pattern**: The script correctly locates the extremum/target row, but then applies a formatting/truncation step (e.g., `strftime` to a shorter pattern, rounding, taking a prefix/substring) that discards the identifying detail, so the reported key no longer uniquely designates the row actually used for the downstream computation.
- **Detection procedure**:
  1. Read the task: determine the granularity of the entity being identified (one row / one timestamp / one ID) and note whether the downstream calculation depends on that exact row.
  2. Read the script: find where the identified key is converted for output; check whether the emitted string contains strictly less information than the key stored in the data.
  3. Compare the emitted key against the value used internally for the dependent computation — if the dependent computation uses the full-precision key but the answer reports a truncated one, flag it; prefer reporting the key exactly as it appears in the source (or, if the template is genuinely ambiguous, report the full-precision value rather than truncating).
  4. Check for collisions: would the truncated key match multiple rows in the data? If yes, it cannot be the intended unique answer.
- **Discriminator**: A real violation is when truncation destroys uniqueness or drops detail present in the source key while the rest of the answer depends on that detail. It is fine when the task explicitly aggregates at the coarser level (e.g., the maximum is computed over monthly aggregates) or when the source key itself has only that granularity.
- **Consequence**: The numeric part of the answer can be correct while the identifier check fails, giving a partial-credit/incorrect verdict (e.g., 1 of 2 checks passed).
98Missing/unverifiable output artifact — answer reported inline instead of written to the required filetaskda-code
Applies when
task -- the task explicitly instructs that results be saved to a named output file (e.g., result.csv) with the computed statistic(s).
Pattern
The attempt computes a single number and reports it in prose/chat, but no saved script or code path demonstrably writes the named file (with the expected column/row layout, precision, and the exact requested quantity); the artifact is absent, empty, or contains an intermediate value rather than the requested one.
Detection procedure
  1. Read the task and list every required deliverable: file name, location, and what it must contain (which statistic, how labelled, any rounding/format constraints).
  2. Search the submitted scripts for an explicit write call (to_csv/open(...).write/etc.) targeting that exact filename; if no scripts exist at all, the deliverable is unverifiable by definition.
  3. Trace which variable is written and confirm it is the final requested quantity computed on the specified subset/aggregation (not an intermediate, unfiltered, or differently-aggregated value), and that the file layout would parse as a table.
  4. Cross-check the reported number against the file-writing code: if the only evidence is a bare float in the answer text, flag it.
Discriminator
A real violation is when no reproducible code writes the named artifact, or the artifact holds a different quantity/format than requested; a look-alike that is fine is a script that writes the correct file under the right name and merely also echoes the value in the answer text.
Consequence
The grader looks for the named result file and finds it missing or containing a mismatched value/format, scoring 0 regardless of whether the number quoted in the chat happens to be close.
id 30ce0fb02dee · mined from da-code dacode-data-sa-043@s2
raw text (what the judge reads)
### Missing/unverifiable output artifact — answer reported inline instead of written to the required file
- **Applies when**: `task` -- the task explicitly instructs that results be saved to a named output file (e.g., `result.csv`) with the computed statistic(s).
- **Pattern**: The attempt computes a single number and reports it in prose/chat, but no saved script or code path demonstrably writes the named file (with the expected column/row layout, precision, and the exact requested quantity); the artifact is absent, empty, or contains an intermediate value rather than the requested one.
- **Detection procedure**:
  1. Read the task and list every required deliverable: file name, location, and what it must contain (which statistic, how labelled, any rounding/format constraints).
  2. Search the submitted scripts for an explicit write call (`to_csv`/`open(...).write`/etc.) targeting that exact filename; if no scripts exist at all, the deliverable is unverifiable by definition.
  3. Trace which variable is written and confirm it is the final requested quantity computed on the specified subset/aggregation (not an intermediate, unfiltered, or differently-aggregated value), and that the file layout would parse as a table.
  4. Cross-check the reported number against the file-writing code: if the only evidence is a bare float in the answer text, flag it.
- **Discriminator**: A real violation is when no reproducible code writes the named artifact, or the artifact holds a different quantity/format than requested; a look-alike that is fine is a script that writes the correct file under the right name and merely also echoes the value in the answer text.
- **Consequence**: The grader looks for the named result file and finds it missing or containing a mismatched value/format, scoring 0 regardless of whether the number quoted in the chat happens to be close.
99Missing reproducible pipeline + unvalidated prediction file (shape / class-rate sanity check)taskda-code
Applies when
task -- the deliverable is a per-row prediction file for a supplied test set, especially with a stated cost/recall asymmetry between error types.
Pattern
The agent produces the output file with no saved, re-runnable script that reads the test file, aligns rows, and writes the required column, and never checks that the emitted file has exactly one prediction per test row, in the original test-row order, with the exact header/values shown in the sample; it also never compares the predicted positive rate against the training base rate or the asymmetric-cost objective, so a near-all-negative (default-threshold, imbalance-collapsed) prediction is submitted uninspected.
Detection procedure
  1. From the task, note the required output format (header text, single column, allowed values) and the exact number of rows in the test file and in the provided sample output.
  2. In the scripts, look for code that (a) loads the test set, (b) predicts, and (c) writes the file preserving test-row order, plus an explicit assertion/print of len(predictions) == len(test) and value_counts(); absence of any saved script is itself a failure of verifiability.
  3. Compare the answer file: count rows (excluding header), confirm header string and value domain match the sample, and compute the fraction of positive predictions.
  4. Compare that fraction to the minority-class prevalence in training and to the stated cost preference (e.g., missed positives more costly than false alarms); flag if it is far below prevalence or if any threshold/class-weight decision was never justified or validated on a held-out split.
Discriminator
A genuine violation is a file whose row count/format cannot be shown to match the test set, or a positive rate materially below the training prevalence with no validation evidence (no held-out recall/F1/cost comparison, no threshold tuning). It is not a violation if the script asserts shape and ordering, the sparse positive rate is backed by held-out metrics on an imbalanced problem, and the format matches the sample exactly.
Consequence
The grader compares the submitted file row-by-row against the reference; a wrong row count, wrong header/order, or a degenerate mostly-negative prediction fails the file check outright (0/1) even though the answer "looks like" valid predictions.
id cfb02ca4e7e0 · mined from da-code dacode-ml-binary-013@s2
raw text (what the judge reads)
### Missing reproducible pipeline + unvalidated prediction file (shape / class-rate sanity check)
- **Applies when**: `task` -- the deliverable is a per-row prediction file for a supplied test set, especially with a stated cost/recall asymmetry between error types.
- **Pattern**: The agent produces the output file with no saved, re-runnable script that reads the test file, aligns rows, and writes the required column, and never checks that the emitted file has exactly one prediction per test row, in the original test-row order, with the exact header/values shown in the sample; it also never compares the predicted positive rate against the training base rate or the asymmetric-cost objective, so a near-all-negative (default-threshold, imbalance-collapsed) prediction is submitted uninspected.
- **Detection procedure**:
  1. From the task, note the required output format (header text, single column, allowed values) and the exact number of rows in the test file and in the provided sample output.
  2. In the scripts, look for code that (a) loads the test set, (b) predicts, and (c) writes the file preserving test-row order, plus an explicit assertion/print of `len(predictions) == len(test)` and `value_counts()`; absence of any saved script is itself a failure of verifiability.
  3. Compare the answer file: count rows (excluding header), confirm header string and value domain match the sample, and compute the fraction of positive predictions.
  4. Compare that fraction to the minority-class prevalence in training and to the stated cost preference (e.g., missed positives more costly than false alarms); flag if it is far below prevalence or if any threshold/class-weight decision was never justified or validated on a held-out split.
- **Discriminator**: A genuine violation is a file whose row count/format cannot be shown to match the test set, or a positive rate materially below the training prevalence with no validation evidence (no held-out recall/F1/cost comparison, no threshold tuning). It is *not* a violation if the script asserts shape and ordering, the sparse positive rate is backed by held-out metrics on an imbalanced problem, and the format matches the sample exactly.
- **Consequence**: The grader compares the submitted file row-by-row against the reference; a wrong row count, wrong header/order, or a degenerate mostly-negative prediction fails the file check outright (0/1) even though the answer "looks like" valid predictions.
100Answer produced by dumping a Python object instead of the requested answer formattaskinfiagent-dabench
Applies when
task -- the task prescribes an exact answer string/format (e.g., @name[list_of_strings]) and the script writes the answer by interpolating a Python variable (list, array, Series, dict) into the output.
Pattern
The script computes the right values but emits them via Python's default repr/str (brackets, single quotes, np.float64(...), dtype= noise, index labels), so the submitted string does not match the required token/delimiter/quoting convention, and no step normalizes or validates the final string.
Detection procedure
  1. Read the task and write down the exact target answer syntax, including delimiters, quote style, and whether names should be bare or quoted.
  2. In the script, find the line(s) that build the final answer and check whether the values are explicitly formatted (e.g., ", ".join(map(str, items))) or just interpolated as a container (f"...{mylist}", print(df[...])).
  3. Mentally render the produced string for a plausible result and compare character-by-character with the required format; also check the values are plain Python scalars/strings, not numpy/pandas objects.
  4. Confirm the script (or agent) does a final check that the emitted string matches the requested pattern; absence of such a check plus raw container interpolation is a violation.
Discriminator
A real violation is when the answer text carries container/object syntax or quoting not sanctioned by the requested format; it is fine if the container's default rendering happens to coincide exactly with the requested syntax, or if the script explicitly joins/formats elements into the prescribed pattern.
Consequence
The grader compares the submitted string to the expected one and reports WRONG/MISSING even though the underlying computation and values are correct.
id c50a4bfb85c6 · mined from infiagent-dabench dabench-254@s2
raw text (what the judge reads)
### Answer produced by dumping a Python object instead of the requested answer format
- **Applies when**: `task` -- the task prescribes an exact answer string/format (e.g., `@name[list_of_strings]`) and the script writes the answer by interpolating a Python variable (list, array, Series, dict) into the output.
- **Pattern**: The script computes the right values but emits them via Python's default `repr`/`str` (brackets, single quotes, `np.float64(...)`, `dtype=` noise, index labels), so the submitted string does not match the required token/delimiter/quoting convention, and no step normalizes or validates the final string.
- **Detection procedure**:
  1. Read the task and write down the exact target answer syntax, including delimiters, quote style, and whether names should be bare or quoted.
  2. In the script, find the line(s) that build the final answer and check whether the values are explicitly formatted (e.g., `", ".join(map(str, items))`) or just interpolated as a container (`f"...{mylist}"`, `print(df[...])`).
  3. Mentally render the produced string for a plausible result and compare character-by-character with the required format; also check the values are plain Python scalars/strings, not numpy/pandas objects.
  4. Confirm the script (or agent) does a final check that the emitted string matches the requested pattern; absence of such a check plus raw container interpolation is a violation.
- **Discriminator**: A real violation is when the answer text carries container/object syntax or quoting not sanctioned by the requested format; it is fine if the container's default rendering happens to coincide exactly with the requested syntax, or if the script explicitly joins/formats elements into the prescribed pattern.
- **Consequence**: The grader compares the submitted string to the expected one and reports WRONG/MISSING even though the underlying computation and values are correct.
101No held-out validation — model quality judged only by in-sample (resubstitution) performancetaskda-code
Applies when
task -- The task asks for predictions on an unlabeled test set and the script fits a model on all labeled data and reports performance metrics.
Pattern
The script trains on 100% of the labeled rows and then computes accuracy/other metrics by predicting on those same training rows (or on data used for fitting imputers/encoders), presenting that number as the model's quality; no cross-validation or hold-out split is ever produced, so the reported score is an optimistic upper bound with no bearing on the graded test predictions. Hyperparameters (depth, estimators, encoding choices) are likewise picked with no evidence, and no alternative configuration is compared.
Detection procedure
  1. Read the task to confirm the deliverable is scored on unseen test predictions against a hidden ground truth (usually with an implicit accuracy threshold).
  2. In the script, locate the fit call and the data passed to the metric call; check whether the evaluation set is disjoint from the fitting set (train_test_split, cross_val_score, or a K-fold loop) — flag if model.predict(X) reuses the training matrix.
  3. Check the answer text: if the only quoted score is the training-set score (typically suspiciously high) and no validation/CV score or variance estimate is given, the attempt has no evidence of generalization.
  4. Also confirm no sanity check beyond row count exists (e.g., predicted class balance compared to the training label distribution, label strings identical to the training labels).
Discriminator
A real violation is when every reported metric comes from rows the model saw during fitting. It is fine to refit on the full labeled set for the final submission after estimating performance on a hold-out/CV split, and it is fine to additionally print a training score alongside a validation score.
Consequence
The reported ~97% is meaningless; the submitted result.csv can fall below the grader's accuracy threshold (or contain systematically skewed/mislabeled classes) while the agent confidently claims success, yielding a WRONG verdict on the output file.
id 22c809b8395e · mined from da-code dacode-ml-binary-009@s2
raw text (what the judge reads)
### No held-out validation — model quality judged only by in-sample (resubstitution) performance
- **Applies when**: `task` -- The task asks for predictions on an unlabeled test set and the script fits a model on all labeled data and reports performance metrics.
- **Pattern**: The script trains on 100% of the labeled rows and then computes accuracy/other metrics by predicting on those same training rows (or on data used for fitting imputers/encoders), presenting that number as the model's quality; no cross-validation or hold-out split is ever produced, so the reported score is an optimistic upper bound with no bearing on the graded test predictions. Hyperparameters (depth, estimators, encoding choices) are likewise picked with no evidence, and no alternative configuration is compared.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on unseen test predictions against a hidden ground truth (usually with an implicit accuracy threshold).
  2. In the script, locate the `fit` call and the data passed to the metric call; check whether the evaluation set is disjoint from the fitting set (`train_test_split`, `cross_val_score`, or a K-fold loop) — flag if `model.predict(X)` reuses the training matrix.
  3. Check the answer text: if the only quoted score is the training-set score (typically suspiciously high) and no validation/CV score or variance estimate is given, the attempt has no evidence of generalization.
  4. Also confirm no sanity check beyond row count exists (e.g., predicted class balance compared to the training label distribution, label strings identical to the training labels).
- **Discriminator**: A real violation is when *every* reported metric comes from rows the model saw during fitting. It is fine to refit on the full labeled set for the final submission *after* estimating performance on a hold-out/CV split, and it is fine to additionally print a training score alongside a validation score.
- **Consequence**: The reported ~97% is meaningless; the submitted `result.csv` can fall below the grader's accuracy threshold (or contain systematically skewed/mislabeled classes) while the agent confidently claims success, yielding a WRONG verdict on the output file.
102Fabricated metric definition instead of the one implied by the task spec/configtaskda-code
Applies when
task -- the task asks to visualize/report a quantity whose definition comes from a provided config, prior pipeline step, or standard domain semantics, rather than being spelled out in the prompt.
Pattern
The script invents an arbitrary formula (e.g., an ad-hoc weighted sum of loosely related counts) and/or computes it on only one slice of the data (one join side, one group, one axis), then labels the bars with the axis/title text from the config as if the numbers matched that meaning; required companion artifacts encoding the true values are never produced.
Detection procedure
  1. Read the task and any config/label text to identify what quantity is actually being asked for (what the y-axis label, title, and expected output files imply), and list all output artifacts required.
  2. In the script, locate where that quantity is computed and check whether the formula is derived from the spec/domain definition or is invented in-line with hard-coded weights or unexplained combinations of columns.
  3. Check that the computation covers every relevant slice of the data (both sides of a symmetric relation, all rows, correct aggregation axis) and that ties/ordering follow the specified entity list.
  4. Compare the reported numbers against a quick plausibility check (expected magnitude/range for that quantity, counts per entity) and confirm every required output file is written.
Discriminator
A real violation is a metric whose definition appears nowhere in the task, data, or config and cannot be reproduced by an independent reader; a look-alike that is fine is a documented, standard definition (explicitly stated or unambiguously implied) that merely uses a compact implementation.
Consequence
The plotted/saved values are unrelated to the requested statistic and the required numeric artifacts are missing, so every value/file check fails even though the figure renders and looks well formatted.
id c66964a49542 · mined from da-code dacode-plot-bar-006@s2
raw text (what the judge reads)
### Fabricated metric definition instead of the one implied by the task spec/config
- **Applies when**: `task` -- the task asks to visualize/report a quantity whose definition comes from a provided config, prior pipeline step, or standard domain semantics, rather than being spelled out in the prompt.
- **Pattern**: The script invents an arbitrary formula (e.g., an ad-hoc weighted sum of loosely related counts) and/or computes it on only one slice of the data (one join side, one group, one axis), then labels the bars with the axis/title text from the config as if the numbers matched that meaning; required companion artifacts encoding the true values are never produced.
- **Detection procedure**:
  1. Read the task and any config/label text to identify what quantity is actually being asked for (what the y-axis label, title, and expected output files imply), and list all output artifacts required.
  2. In the script, locate where that quantity is computed and check whether the formula is derived from the spec/domain definition or is invented in-line with hard-coded weights or unexplained combinations of columns.
  3. Check that the computation covers every relevant slice of the data (both sides of a symmetric relation, all rows, correct aggregation axis) and that ties/ordering follow the specified entity list.
  4. Compare the reported numbers against a quick plausibility check (expected magnitude/range for that quantity, counts per entity) and confirm every required output file is written.
- **Discriminator**: A real violation is a metric whose definition appears nowhere in the task, data, or config and cannot be reproduced by an independent reader; a look-alike that is fine is a documented, standard definition (explicitly stated or unambiguously implied) that merely uses a compact implementation.
- **Consequence**: The plotted/saved values are unrelated to the requested statistic and the required numeric artifacts are missing, so every value/file check fails even though the figure renders and looks well formatted.
103Unverified input source and no independent cross-check of a threshold-based counttaskinfiagent-dabench
Applies when
task -- the task asks for a count of rows satisfying a statistical threshold (e.g. z-score, IQR, quantile rule) on one column of "the dataset", and the scripts pick a file/column and report the count directly.
Pattern
The script grabs the first plausible file (often a split such as *_train.csv) and the first column whose name loosely matches, computes the statistic with one library default (e.g. population vs sample std, one-sided vs two-sided mask, NaN-filled rows silently counted), and reports the resulting count without ever recomputing it a second, independent way or checking that the chosen file/column is the one the task describes.
Detection procedure
  1. Read the task and note exactly which dataset/entity and which quantity is requested (whole dataset vs a split; raw column vs cleaned/numeric-coerced column; both tails vs one tail).
  2. In the scripts, list every I/O and selection decision: which file was loaded, whether other candidate files were enumerated (ls/glob), which column was chosen, how NaNs/non-numeric values were handled, and which std convention (ddof) the statistic uses.
  3. Check whether the count is verified by a second route — e.g. a manual (x - x.mean())/x.std(ddof=...) recomputation, a printed min/max of the statistic, a distribution summary, or a comparison of the two-sided mask vs the absolute-value mask — and whether the resulting count is sanity-checked for plausibility (fraction of rows flagged, shape before/after drop).
  4. Flag the attempt if the file/column choice is unjustified among available alternatives, or if a nonzero count rests on a single library call with no cross-check and no stated reason it should differ from an alternative convention.
Discriminator
Fine: the script confirms the available files/columns, coerces dtypes and drops NaNs explicitly, and reproduces the same count under at least one alternative convention (or shows the extreme statistic values so the count is obviously robust). Violation: the count is produced once from an implicitly chosen file/column and default parameters, with no enumeration of alternatives and no second computation — so a wrong source or convention is indistinguishable from a correct one.
Consequence
The reported count differs sharply from the ground truth (e.g. a large nonzero count where the correct answer is zero, or vice versa), and the single-value answer check fails outright.
id 3f02ffbacf46 · mined from infiagent-dabench dabench-361@s2
raw text (what the judge reads)
### Unverified input source and no independent cross-check of a threshold-based count
- **Applies when**: `task` -- the task asks for a count of rows satisfying a statistical threshold (e.g. z-score, IQR, quantile rule) on one column of "the dataset", and the scripts pick a file/column and report the count directly.
- **Pattern**: The script grabs the first plausible file (often a split such as `*_train.csv`) and the first column whose name loosely matches, computes the statistic with one library default (e.g. population vs sample std, one-sided vs two-sided mask, NaN-filled rows silently counted), and reports the resulting count without ever recomputing it a second, independent way or checking that the chosen file/column is the one the task describes.
- **Detection procedure**:
  1. Read the task and note exactly which dataset/entity and which quantity is requested (whole dataset vs a split; raw column vs cleaned/numeric-coerced column; both tails vs one tail).
  2. In the scripts, list every I/O and selection decision: which file was loaded, whether other candidate files were enumerated (`ls`/`glob`), which column was chosen, how NaNs/non-numeric values were handled, and which std convention (`ddof`) the statistic uses.
  3. Check whether the count is verified by a second route — e.g. a manual `(x - x.mean())/x.std(ddof=...)` recomputation, a printed min/max of the statistic, a distribution summary, or a comparison of the two-sided mask vs the absolute-value mask — and whether the resulting count is sanity-checked for plausibility (fraction of rows flagged, shape before/after drop).
  4. Flag the attempt if the file/column choice is unjustified among available alternatives, or if a nonzero count rests on a single library call with no cross-check and no stated reason it should differ from an alternative convention.
- **Discriminator**: Fine: the script confirms the available files/columns, coerces dtypes and drops NaNs explicitly, and reproduces the same count under at least one alternative convention (or shows the extreme statistic values so the count is obviously robust). Violation: the count is produced once from an implicitly chosen file/column and default parameters, with no enumeration of alternatives and no second computation — so a wrong source or convention is indistinguishable from a correct one.
- **Consequence**: The reported count differs sharply from the ground truth (e.g. a large nonzero count where the correct answer is zero, or vice versa), and the single-value answer check fails outright.
104Missing reproducible artifact: answer asserted without a script or the required output filetaskda-code
Applies when
task -- The task specifies a deliverable (e.g., a named result file) and/or requires applying an externally-specified transformation before computing a statistic, and the agent must produce code that does it.
Pattern
The agent reports the summary value only as inline prose/JSON in its message, with no saved script that loads the data, applies the specified mapping/preprocessing, computes the statistic, and writes the requested output file — so the number is unverifiable and the required artifact is never created.
Detection procedure
  1. Read the task and list every required deliverable (output file name/location, exact key names, value types, rounding/format) and every required preprocessing step defined in auxiliary instructions or docs.
  2. Inspect the saved scripts/artifacts: is there code that reads the source data, applies each required transformation, computes the requested quantity, and serializes it to the named file?
  3. If no such code exists (or it prints only to stdout), treat the answer as unsubstantiated; if code exists, re-derive the value mentally from it and check the mapping was applied to the whole column (not just categories the agent assumed) and that the reported figure is the requested statistic, not an intermediate.
  4. Sanity-check the reported number against basics implied by the data (proportions sum to 1, the modal share is the maximum share, expected label vocabulary after mapping).
Discriminator
A real violation is an answer with no executable path from raw data to the reported value and/or no required output file. Not a violation: a script exists and writes the file correctly, and the chat text merely restates its contents (formatting differences in the message are fine as long as the artifact matches spec).
Consequence
The grader finds the expected result file missing or its contents mismatched, so all checks fail regardless of whether the quoted number happened to be close.
id 0c8014d589dd · mined from da-code dacode-di-text-004@s2
raw text (what the judge reads)
### Missing reproducible artifact: answer asserted without a script or the required output file
- **Applies when**: `task` -- The task specifies a deliverable (e.g., a named result file) and/or requires applying an externally-specified transformation before computing a statistic, and the agent must produce code that does it.
- **Pattern**: The agent reports the summary value only as inline prose/JSON in its message, with no saved script that loads the data, applies the specified mapping/preprocessing, computes the statistic, and writes the requested output file — so the number is unverifiable and the required artifact is never created.
- **Detection procedure**:
  1. Read the task and list every required deliverable (output file name/location, exact key names, value types, rounding/format) and every required preprocessing step defined in auxiliary instructions or docs.
  2. Inspect the saved scripts/artifacts: is there code that reads the source data, applies each required transformation, computes the requested quantity, and serializes it to the named file?
  3. If no such code exists (or it prints only to stdout), treat the answer as unsubstantiated; if code exists, re-derive the value mentally from it and check the mapping was applied to the whole column (not just categories the agent assumed) and that the reported figure is the requested statistic, not an intermediate.
  4. Sanity-check the reported number against basics implied by the data (proportions sum to 1, the modal share is the maximum share, expected label vocabulary after mapping).
- **Discriminator**: A real violation is an answer with no executable path from raw data to the reported value and/or no required output file. Not a violation: a script exists and writes the file correctly, and the chat text merely restates its contents (formatting differences in the message are fine as long as the artifact matches spec).
- **Consequence**: The grader finds the expected result file missing or its contents mismatched, so all checks fail regardless of whether the quoted number happened to be close.
105No out-of-sample validation with the competition's stated probabilistic metrictaskda-code
Applies when
task -- the task specifies an explicit evaluation metric over predicted probabilities (e.g. a weighted/balanced log loss) and the scripts fit one or more classifiers and write probability columns to a submission file.
Pattern
The attempt trains models on the full training set, computes predictions only on the same training rows (or not at all), never implements the stated metric, and never estimates it on held-out folds. Class re-weighting / imbalance handling is switched on for every model and the raw, uncalibrated per-model probabilities are averaged, so no evidence exists that the submitted probabilities are better than a trivial baseline under the actual scoring rule.
Detection procedure
  1. Read the task/README and note the exact scoring function and any class-balancing in it.
  2. Search the scripts for (a) an implementation of that metric, (b) a train/validation split or cross-validation loop whose score is printed, and (c) a comparison against a naive baseline (e.g. constant class prior).
  3. Check whether any printed diagnostics are computed on rows the model was fit on (in-sample predict_proba on X_train) rather than held-out rows.
  4. Inspect the answer: if only min/max/mean of predictions and a "sums to 1" check are reported, with no metric value, the model quality is unverified — and check whether probability calibration was considered given the reweighting used during fitting.
Discriminator
A real violation has zero held-out estimate of the stated metric anywhere in the pipeline. It is not a violation if the scripts compute the competition metric (or a documented equivalent) via CV/holdout and use it to choose among models/blend weights, even if the final model is refit on all data; nor if the metric name differs but the formula and class weighting match.
Consequence
The submission may be badly miscalibrated (over-confident, prior-shifted by class weights), yielding a log-loss far worse than a simple calibrated baseline, so the graded score falls below the acceptance threshold and the answer is marked wrong despite a syntactically valid file.
id 7a8d42b01fc1 · mined from da-code dacode-ml-competition-003@s2
raw text (what the judge reads)
### No out-of-sample validation with the competition's stated probabilistic metric
- **Applies when**: `task` -- the task specifies an explicit evaluation metric over predicted probabilities (e.g. a weighted/balanced log loss) and the scripts fit one or more classifiers and write probability columns to a submission file.
- **Pattern**: The attempt trains models on the full training set, computes predictions only on the same training rows (or not at all), never implements the stated metric, and never estimates it on held-out folds. Class re-weighting / imbalance handling is switched on for every model and the raw, uncalibrated per-model probabilities are averaged, so no evidence exists that the submitted probabilities are better than a trivial baseline under the actual scoring rule.
- **Detection procedure**:
  1. Read the task/README and note the exact scoring function and any class-balancing in it.
  2. Search the scripts for (a) an implementation of that metric, (b) a train/validation split or cross-validation loop whose score is printed, and (c) a comparison against a naive baseline (e.g. constant class prior).
  3. Check whether any printed diagnostics are computed on rows the model was fit on (in-sample `predict_proba` on `X_train`) rather than held-out rows.
  4. Inspect the answer: if only min/max/mean of predictions and a "sums to 1" check are reported, with no metric value, the model quality is unverified — and check whether probability calibration was considered given the reweighting used during fitting.
- **Discriminator**: A real violation has *zero* held-out estimate of the stated metric anywhere in the pipeline. It is not a violation if the scripts compute the competition metric (or a documented equivalent) via CV/holdout and use it to choose among models/blend weights, even if the final model is refit on all data; nor if the metric name differs but the formula and class weighting match.
- **Consequence**: The submission may be badly miscalibrated (over-confident, prior-shifted by class weights), yielding a log-loss far worse than a simple calibrated baseline, so the graded score falls below the acceptance threshold and the answer is marked wrong despite a syntactically valid file.
106Output contract not verified against the test file (row count, ordering, and exact label vocabulary)taskda-code
Applies when
task -- the task says to produce predictions for a supplied evaluation file and save them to a named output file with a specified column, and the script writes that file after fitting a model.
Pattern
The attempt builds a model on convenient rows (e.g., only records that have the engineered features non-null, or an internally re-split subset), then writes an output whose row count/order does not correspond one-to-one with the evaluation file's rows, and/or emits class strings the agent invented or reformatted rather than the exact category strings present in the labeled data. Suspiciously high validation accuracy from features derived after joining on label-correlated or leaked aggregates is reported as evidence of success instead of triggering a check.
Detection procedure
  1. From the task/README, note the required output filename, required column name(s), and the number of rows in the evaluation input.
  2. In the script, trace how the prediction frame is constructed: does it start from the full evaluation file (preserving its row order and every row, filling/imputing where features are missing), or from a filtered/merged/deduplicated subset?
  3. Compare the label strings produced (from classes_, mapping dicts, or hardcoded lists) against the distinct category values in the labeled source data — they must match character-for-character, with no renaming, casing changes, or added/dropped classes.
  4. Check the reported record count and class distribution in the answer against the evaluation file's row count and the training-label prior; treat a large discrepancy or a near-perfect validation score with a handful of weak features as an unverified red flag.
Discriminator
A real violation is when the written file cannot be aligned row-for-row with the evaluation input, or contains label strings not in the label vocabulary. It is not a violation if the script legitimately joins by an ID key, reindexes to the evaluation file's full index, and fills any missing predictions — even if intermediate modeling used a filtered subset.
Consequence
The grader compares the submitted file to ground truth keyed by evaluation rows and finds mismatched length/order or unrecognized label values, so accuracy collapses (or the file is scored as wrong/missing) despite the reported high internal validation score.
id 0cce2dcbd6f6 · mined from da-code dacode-ml-multi-003@s2
raw text (what the judge reads)
### Output contract not verified against the test file (row count, ordering, and exact label vocabulary)
- **Applies when**: `task` -- the task says to produce predictions for a supplied evaluation file and save them to a named output file with a specified column, and the script writes that file after fitting a model.
- **Pattern**: The attempt builds a model on convenient rows (e.g., only records that have the engineered features non-null, or an internally re-split subset), then writes an output whose row count/order does not correspond one-to-one with the evaluation file's rows, and/or emits class strings the agent invented or reformatted rather than the exact category strings present in the labeled data. Suspiciously high validation accuracy from features derived after joining on label-correlated or leaked aggregates is reported as evidence of success instead of triggering a check.
- **Detection procedure**:
  1. From the task/README, note the required output filename, required column name(s), and the number of rows in the evaluation input.
  2. In the script, trace how the prediction frame is constructed: does it start from the full evaluation file (preserving its row order and every row, filling/imputing where features are missing), or from a filtered/merged/deduplicated subset?
  3. Compare the label strings produced (from `classes_`, mapping dicts, or hardcoded lists) against the distinct category values in the labeled source data — they must match character-for-character, with no renaming, casing changes, or added/dropped classes.
  4. Check the reported record count and class distribution in the answer against the evaluation file's row count and the training-label prior; treat a large discrepancy or a near-perfect validation score with a handful of weak features as an unverified red flag.
- **Discriminator**: A real violation is when the written file cannot be aligned row-for-row with the evaluation input, or contains label strings not in the label vocabulary. It is *not* a violation if the script legitimately joins by an ID key, reindexes to the evaluation file's full index, and fills any missing predictions — even if intermediate modeling used a filtered subset.
- **Consequence**: The grader compares the submitted file to ground truth keyed by evaluation rows and finds mismatched length/order or unrecognized label values, so accuracy collapses (or the file is scored as wrong/missing) despite the reported high internal validation score.
107Output format not verified against the provided templatetaskda-code
Applies when
task -- The task says results must be saved to a named output file whose format must match a supplied template/example file.
Pattern
The script constructs the output table from its own assumptions (index labels, column names/order, header text, rounding, number of rows/columns) and writes it directly, never loading or comparing against the template file; self-"verification" only re-checks the script's own numbers rather than the required layout.
Detection procedure
1) Read the task for the phrase requiring the output to match a template and locate the template path. 2) Search the scripts for any read of that template file and any comparison of shape, index values, column names/dtypes, or rounding against it. 3) Inspect how index/column labels are generated (e.g., period-to-timestamp conversion, offsetting indices by 1, string formatting) and ask whether these were chosen from the template or invented. 4) Check the final answer for an explicit statement that the written file's shape and headers equal the template's.
Discriminator
A genuine violation is when no template comparison exists and format choices are asserted rather than derived; it is fine if the script reads the template and asserts equality of columns/index (or reindexes onto it), even if it also rounds or renames afterwards.
Consequence
The saved file's cell values may be right but the grader's file-level comparison fails on mismatched index labels, column offsets/names, or rounding, marking the expected result file WRONG.
id b7f21e913074 · mined from da-code dacode-dm-csv-044@s2
raw text (what the judge reads)
### Output format not verified against the provided template
- **Applies when**: `task` -- The task says results must be saved to a named output file whose format must match a supplied template/example file.
- **Pattern**: The script constructs the output table from its own assumptions (index labels, column names/order, header text, rounding, number of rows/columns) and writes it directly, never loading or comparing against the template file; self-"verification" only re-checks the script's own numbers rather than the required layout.
- **Detection procedure**: 1) Read the task for the phrase requiring the output to match a template and locate the template path. 2) Search the scripts for any read of that template file and any comparison of shape, index values, column names/dtypes, or rounding against it. 3) Inspect how index/column labels are generated (e.g., period-to-timestamp conversion, offsetting indices by 1, string formatting) and ask whether these were chosen from the template or invented. 4) Check the final answer for an explicit statement that the written file's shape and headers equal the template's.
- **Discriminator**: A genuine violation is when no template comparison exists and format choices are asserted rather than derived; it is fine if the script reads the template and asserts equality of columns/index (or reindexes onto it), even if it also rounds or renames afterwards.
- **Consequence**: The saved file's cell values may be right but the grader's file-level comparison fails on mismatched index labels, column offsets/names, or rounding, marking the expected result file WRONG.
108Deliverable file never validated against the provided template (prose report substituted for a checked artifact)taskda-code
Applies when
task -- the task requires writing a specific output file whose format is defined by a provided sample/template file, and the agent's scripts end by printing a summary of the modelling approach.
Pattern
The attempt focuses on model architecture and CV scores, writes the output file from an in-memory frame, and never re-reads the written file to confirm it matches the template: same column names/order, same number of rows, exactly the same key/ID values (taken from the test file, not regenerated or re-indexed), correct dtype, and no stray index column. The final answer describes the method and prediction counts instead of demonstrating the file was verified.
Detection procedure
  1. Read the task for the required filename, required columns, and the reference sample file; note the key column and expected row count.
  2. In the scripts, locate the write step and check whether the ID column is copied verbatim from the test input (and whether index=False / header conventions are honoured), and whether any post-write reload + comparison against the sample file occurs.
  3. Check the answer for concrete evidence of validation: first rows of the saved file, row count vs. sample row count, and a set-equality check of IDs; absence of these means unverified.
  4. Flag if the answer's only evidence is CV metrics and a prediction-class histogram.
Discriminator
A real violation is when no read-back/format assertion exists and IDs are not provably sourced from the test file; a look-alike that is fine is an attempt that writes the file with template-derived IDs and shows an explicit shape/column/ID-match check, even if it reports little about the model.
Consequence
The grader marks the expected output file as WRONG/MISSING (mismatched IDs, row count, column names, or an extra index column), scoring 0 regardless of model quality.
id 65aede131bc9 · mined from da-code dacode-ml-competition-006@s2
raw text (what the judge reads)
### Deliverable file never validated against the provided template (prose report substituted for a checked artifact)
- **Applies when**: `task` -- the task requires writing a specific output file whose format is defined by a provided sample/template file, and the agent's scripts end by printing a summary of the modelling approach.
- **Pattern**: The attempt focuses on model architecture and CV scores, writes the output file from an in-memory frame, and never re-reads the written file to confirm it matches the template: same column names/order, same number of rows, exactly the same key/ID values (taken from the test file, not regenerated or re-indexed), correct dtype, and no stray index column. The final answer describes the method and prediction counts instead of demonstrating the file was verified.
- **Detection procedure**:
  1. Read the task for the required filename, required columns, and the reference sample file; note the key column and expected row count.
  2. In the scripts, locate the write step and check whether the ID column is copied verbatim from the test input (and whether `index=False` / header conventions are honoured), and whether any post-write reload + comparison against the sample file occurs.
  3. Check the answer for concrete evidence of validation: first rows of the saved file, row count vs. sample row count, and a set-equality check of IDs; absence of these means unverified.
  4. Flag if the answer's only evidence is CV metrics and a prediction-class histogram.
- **Discriminator**: A real violation is when no read-back/format assertion exists and IDs are not provably sourced from the test file; a look-alike that is fine is an attempt that writes the file with template-derived IDs and shows an explicit shape/column/ID-match check, even if it reports little about the model.
- **Consequence**: The grader marks the expected output file as WRONG/MISSING (mismatched IDs, row count, column names, or an extra index column), scoring 0 regardless of model quality.
109Missing-value mask defined inconsistently, so the two groups don't partition the full datasettaskinfiagent-dabench
Applies when
task -- a task asks to split rows by whether a field is null/missing (or by any boolean condition) and compare an aggregate statistic between the two resulting groups.
Pattern
The attempt builds the group mask with a definition that silently differs from "null" as stored in the raw file — e.g. relying on default loader NA-conversion so that empty strings, whitespace, "NA"/"None"/"-" sentinels, or type coercion move rows to the wrong side — and/or applies an extra dropna()/filter on the measured column (or other columns) before grouping, so some rows are dropped from both groups. Group means then drift from the reference values in the same direction, and no row-count check is done.
Detection procedure
  1. Read the task and note the exact partition rule and which column the statistic is computed on; the two groups should together cover every row of the source table.
  2. In the script, locate the load step (delimiter, na_values, keep_default_na, dtype/converters) and the mask construction; check whether string sentinels/whitespace are handled the same way the task intends, and whether any dropna, filter, deduplication, or subsetting happens before the split.
  3. Check that the script prints len(group_a) + len(group_b) == len(full_df) and the per-group non-null counts used for the statistic; absence of such a check is itself a red flag.
  4. Compare reported group statistics against a quick independent recomputation of the mask (e.g. counting raw blank fields) — systematic small shifts in both group means indicate rows landed on the wrong side or were dropped.
Discriminator
A real violation is when the group sizes don't sum to the table size, or the null test used differs from the raw-file notion of missing (sentinels/blank strings mishandled), or rows were removed by pre-filtering. It is not a violation if rows are excluded only because the measured numeric column itself is missing for them (unavoidable for a mean/t-test) and the script reports those counts explicitly, or if the loader options demonstrably reproduce the same mask.
Consequence
Both group means (and the test statistic) are computed on slightly wrong subsets; the grader's exact numeric checks on each mean fail even though the pipeline and answer format look correct.
id d1ee9beabceb · mined from infiagent-dabench dabench-297@s2
raw text (what the judge reads)
### Missing-value mask defined inconsistently, so the two groups don't partition the full dataset
- **Applies when**: `task` -- a task asks to split rows by whether a field is null/missing (or by any boolean condition) and compare an aggregate statistic between the two resulting groups.
- **Pattern**: The attempt builds the group mask with a definition that silently differs from "null" as stored in the raw file — e.g. relying on default loader NA-conversion so that empty strings, whitespace, `"NA"`/`"None"`/`"-"` sentinels, or type coercion move rows to the wrong side — and/or applies an extra `dropna()`/filter on the measured column (or other columns) before grouping, so some rows are dropped from both groups. Group means then drift from the reference values in the same direction, and no row-count check is done.
- **Detection procedure**:
  1. Read the task and note the exact partition rule and which column the statistic is computed on; the two groups should together cover every row of the source table.
  2. In the script, locate the load step (delimiter, `na_values`, `keep_default_na`, dtype/converters) and the mask construction; check whether string sentinels/whitespace are handled the same way the task intends, and whether any `dropna`, filter, deduplication, or subsetting happens before the split.
  3. Check that the script prints `len(group_a) + len(group_b) == len(full_df)` and the per-group non-null counts used for the statistic; absence of such a check is itself a red flag.
  4. Compare reported group statistics against a quick independent recomputation of the mask (e.g. counting raw blank fields) — systematic small shifts in both group means indicate rows landed on the wrong side or were dropped.
- **Discriminator**: A real violation is when the group sizes don't sum to the table size, or the null test used differs from the raw-file notion of missing (sentinels/blank strings mishandled), or rows were removed by pre-filtering. It is *not* a violation if rows are excluded only because the measured numeric column itself is missing for them (unavoidable for a mean/t-test) **and** the script reports those counts explicitly, or if the loader options demonstrably reproduce the same mask.
- **Consequence**: Both group means (and the test statistic) are computed on slightly wrong subsets; the grader's exact numeric checks on each mean fail even though the pipeline and answer format look correct.
110Referenced spec file never read; requirements and outputs inventedtaskda-code
Applies when
task -- The task points to an external instruction/config/guidance document (or a README/spec section) that defines the analysis rules, thresholds, categories, and the deliverable artifacts.
Pattern
The script never opens or quotes the referenced document; instead the agent hard-codes its own filtering thresholds, category definitions, and output set (often saving only the one artifact named in the prompt summary), and the final answer restates those invented rules as if they were the spec.
Detection procedure
  1. From the task text, list every referenced spec/instruction file and every deliverable it implies (plots, arrays, JSON/CSV summaries, exact filenames).
  2. Search the scripts for a read of that spec file (or for comments/docstrings that verbatim mirror its rules); check whether thresholds, groupings, and category mappings are asserted with no traceable source ("Rules: 1..., 2...", magic numbers like a 180-day window, priority orders).
  3. Enumerate the files the script actually writes and compare to the deliverable list; flag any deliverable that is never created, and any created artifact whose content/ordering/format is not verified against the spec.
  4. Check the reported answer: does it report numbers derived from self-invented buckets/filters rather than spec-defined ones, and does it claim completion while some required artifacts are absent?
Discriminator
A real violation is when the rules or the artifact set come purely from the agent's assumptions and at least one required output is missing or unverifiable; it is not a violation if the script demonstrably loads/echoes the spec (or the task text itself fully enumerates the rules and filenames) and every named artifact is written and sanity-checked.
Consequence
Graders that compare per-artifact outputs report the required files as WRONG/MISSING and the derived tallies/plot as mismatched, yielding 0 checks passed even though the script runs without error.
id 3a1a3cab9665 · mined from da-code dacode-plot-pie-005@s2
raw text (what the judge reads)
### Referenced spec file never read; requirements and outputs invented
- **Applies when**: `task` -- The task points to an external instruction/config/guidance document (or a README/spec section) that defines the analysis rules, thresholds, categories, and the deliverable artifacts.
- **Pattern**: The script never opens or quotes the referenced document; instead the agent hard-codes its own filtering thresholds, category definitions, and output set (often saving only the one artifact named in the prompt summary), and the final answer restates those invented rules as if they were the spec.
- **Detection procedure**:
  1. From the task text, list every referenced spec/instruction file and every deliverable it implies (plots, arrays, JSON/CSV summaries, exact filenames).
  2. Search the scripts for a read of that spec file (or for comments/docstrings that verbatim mirror its rules); check whether thresholds, groupings, and category mappings are asserted with no traceable source ("Rules: 1..., 2...", magic numbers like a 180-day window, priority orders).
  3. Enumerate the files the script actually writes and compare to the deliverable list; flag any deliverable that is never created, and any created artifact whose content/ordering/format is not verified against the spec.
  4. Check the reported answer: does it report numbers derived from self-invented buckets/filters rather than spec-defined ones, and does it claim completion while some required artifacts are absent?
- **Discriminator**: A real violation is when the rules or the artifact set come purely from the agent's assumptions and at least one required output is missing or unverifiable; it is *not* a violation if the script demonstrably loads/echoes the spec (or the task text itself fully enumerates the rules and filenames) and every named artifact is written and sanity-checked.
- **Consequence**: Graders that compare per-artifact outputs report the required files as WRONG/MISSING and the derived tallies/plot as mismatched, yielding 0 checks passed even though the script runs without error.
111Output schema not derived from the provided template filetaskda-code
Applies when
task -- the task points to an example/sample output file (or explicitly states a required column set, header names, or row order) and the scripts must produce a deliverable file matching it.
Pattern
The attempt never reads or prints the referenced sample/template; it hard-codes invented column names and row contents from intuition, sometimes duplicating a field or omitting a required one, so the deliverable's schema cannot be verified against the expected one.
Detection procedure
  1. In the task text, identify every referenced format artifact (sample file, stated headers, ordering, units, rounding).
  2. Search the scripts for any read/print of that artifact, or an explicit column list copied from it; note whether the written DataFrame's columns and row labels are justified by anything observed.
  3. Compare the answer's header row and label values to what the task implies (e.g., redundant/renamed columns, extra or missing columns, made-up label wording).
  4. Flag if the schema was invented rather than mirrored, even when the underlying computed values look plausible.
Discriminator
Not a violation if the scripts actually load the template (or the task fully specifies headers verbatim) and construct the output with exactly those columns/labels; it is a violation when the header names/labels appear nowhere in the task or inspected data and no verification step compares them to the template.
Consequence
The grader's file/column comparison fails outright (WRONG/MISSING) despite correct analytical values, yielding a zero score.
id da0b3aedc653 · mined from da-code dacode-dm-csv-015@s2
raw text (what the judge reads)
### Output schema not derived from the provided template file
- **Applies when**: `task` -- the task points to an example/sample output file (or explicitly states a required column set, header names, or row order) and the scripts must produce a deliverable file matching it.
- **Pattern**: The attempt never reads or prints the referenced sample/template; it hard-codes invented column names and row contents from intuition, sometimes duplicating a field or omitting a required one, so the deliverable's schema cannot be verified against the expected one.
- **Detection procedure**:
  1. In the task text, identify every referenced format artifact (sample file, stated headers, ordering, units, rounding).
  2. Search the scripts for any read/print of that artifact, or an explicit column list copied from it; note whether the written DataFrame's columns and row labels are justified by anything observed.
  3. Compare the answer's header row and label values to what the task implies (e.g., redundant/renamed columns, extra or missing columns, made-up label wording).
  4. Flag if the schema was invented rather than mirrored, even when the underlying computed values look plausible.
- **Discriminator**: Not a violation if the scripts actually load the template (or the task fully specifies headers verbatim) and construct the output with exactly those columns/labels; it is a violation when the header names/labels appear nowhere in the task or inspected data and no verification step compares them to the template.
- **Consequence**: The grader's file/column comparison fails outright (WRONG/MISSING) despite correct analytical values, yielding a zero score.
112Skipping data-quality validation before imputing/modeling on a dataset that advertises dirty valuestaskinfiagent-dabench
Applies when
task -- The script loads a raw file (often named or documented as containing errors/noise) and immediately imputes means and fits a model on the numeric columns without inspecting value distributions, dtypes, or plausibility ranges.
Pattern
The agent treats every stored value as valid: it computes column means for imputation and fits/evaluates the model on columns that still contain corrupted entries (non-numeric strings coerced or dropped, impossible signs/units, duplicated-digit or off-by-orders-of-magnitude outliers, sentinel codes like -999). Because means and regression coefficients are highly sensitive to such entries, both the imputation values and the test-set error are inflated, and the reported metric is off by an order of magnitude.
Detection procedure
  1. Read the task/data source for any signal that the raw data is unclean (file name, description, mention of "errors"/"raw"/"noisy"), and note which columns feed the model.
  2. Scan the scripts for any validation step on those columns: dtype checks with coercion, min/max or range plausibility checks, outlier/sentinel detection, duplicate/inconsistent-row handling. If the only preprocessing is fillna(mean) (or similar) plus a split, flag it.
  3. Check whether the printed diagnostics the agent did run (means, coefficients, sample predictions, R²) were sanity-checked against domain-plausible ranges for the target; if a mean or a resulting error magnitude is implausible relative to the variable's natural scale and the agent did not investigate, flag it.
  4. Compare the reported metric to the target variable's own variance/scale — an MSE comparable to or larger than the target's total variance implies the model is worse than predicting the mean, which should have triggered a cleaning pass.
Discriminator
A real violation is when no inspection of value validity occurred at all, or when obvious anomalies were visible in the agent's own output and left unaddressed. It is not a violation when the agent explicitly examined the columns (dtype coercion, range/outlier report) and documented a justified decision to keep all values, or when the task explicitly forbids cleaning beyond the stated imputation.
Consequence
Corrupted values dominate the fitted coefficients and the test residuals, so the reported error metric is inflated by roughly an order of magnitude versus the reference value, and the exact-match numeric check fails.
id 6589e593fa42 · mined from infiagent-dabench dabench-432@s2
raw text (what the judge reads)
### Skipping data-quality validation before imputing/modeling on a dataset that advertises dirty values
- **Applies when**: `task` -- The script loads a raw file (often named or documented as containing errors/noise) and immediately imputes means and fits a model on the numeric columns without inspecting value distributions, dtypes, or plausibility ranges.
- **Pattern**: The agent treats every stored value as valid: it computes column means for imputation and fits/evaluates the model on columns that still contain corrupted entries (non-numeric strings coerced or dropped, impossible signs/units, duplicated-digit or off-by-orders-of-magnitude outliers, sentinel codes like -999). Because means and regression coefficients are highly sensitive to such entries, both the imputation values and the test-set error are inflated, and the reported metric is off by an order of magnitude.
- **Detection procedure**:
  1. Read the task/data source for any signal that the raw data is unclean (file name, description, mention of "errors"/"raw"/"noisy"), and note which columns feed the model.
  2. Scan the scripts for any validation step on those columns: dtype checks with coercion, min/max or range plausibility checks, outlier/sentinel detection, duplicate/inconsistent-row handling. If the only preprocessing is `fillna(mean)` (or similar) plus a split, flag it.
  3. Check whether the printed diagnostics the agent did run (means, coefficients, sample predictions, R²) were sanity-checked against domain-plausible ranges for the target; if a mean or a resulting error magnitude is implausible relative to the variable's natural scale and the agent did not investigate, flag it.
  4. Compare the reported metric to the target variable's own variance/scale — an MSE comparable to or larger than the target's total variance implies the model is worse than predicting the mean, which should have triggered a cleaning pass.
- **Discriminator**: A real violation is when no inspection of value validity occurred at all, or when obvious anomalies were visible in the agent's own output and left unaddressed. It is *not* a violation when the agent explicitly examined the columns (dtype coercion, range/outlier report) and documented a justified decision to keep all values, or when the task explicitly forbids cleaning beyond the stated imputation.
- **Consequence**: Corrupted values dominate the fitted coefficients and the test residuals, so the reported error metric is inflated by roughly an order of magnitude versus the reference value, and the exact-match numeric check fails.
113Normalization/scaling method mismatch producing degenerate summary statisticstaskinfiagent-dabench
Applies when
task -- the task asks to rescale/normalize numeric columns and then report summary statistics (e.g., means) of the transformed columns.
Pattern
The attempt applies a scaler whose output makes the requested statistic trivially constant (e.g., zero-centering/standardization when a bounded [0,1] min–max rescaling was intended), and reports the degenerate values (all 0.0000, or all 1.0000 for std) without questioning whether the transform matches the task's intent or whether the reported number carries any information.
Detection procedure
  1. Read the task wording for the transform requested and note whether the expected output range/answer format implies a specific scaling family (bounded 0–1 vs. mean-0/unit-variance), and note that the task asks for a per-column informative statistic.
  2. In the scripts, identify the exact scaler/formula applied to each column and derive analytically what the reported statistic must equal under that transform.
  3. Flag the attempt if the derived statistic is mathematically constant regardless of the data (0 for means after centering, 1 for std after standardizing) — i.e., the reported numbers convey no dataset-specific information.
  4. Cross-check the reported numbers: any column whose value is exactly 0.0000/1.0000 while other, untransformed columns show data-dependent values is strong evidence of the mismatch; also confirm all requested output fields are present and consistently derived.
Discriminator
A real violation is when the reported statistic is an artifact of the transform (constant by construction) or the scaling family contradicts the task's implied range; it is fine if the task explicitly names standardization and the reported statistic is still data-dependent (e.g., reporting means of columns that were not centered, or reporting min/max after standardizing).
Consequence
The graded values for every rescaled column are 0.0000 (or otherwise uniform) instead of the expected data-dependent values, so only the untouched/binary columns match and the answer fails most checks.
id e725063f01a0 · mined from infiagent-dabench dabench-28@s2
raw text (what the judge reads)
### Normalization/scaling method mismatch producing degenerate summary statistics
- **Applies when**: `task` -- the task asks to rescale/normalize numeric columns and then report summary statistics (e.g., means) of the transformed columns.
- **Pattern**: The attempt applies a scaler whose output makes the requested statistic trivially constant (e.g., zero-centering/standardization when a bounded [0,1] min–max rescaling was intended), and reports the degenerate values (all 0.0000, or all 1.0000 for std) without questioning whether the transform matches the task's intent or whether the reported number carries any information.
- **Detection procedure**:
  1. Read the task wording for the transform requested and note whether the expected output range/answer format implies a specific scaling family (bounded 0–1 vs. mean-0/unit-variance), and note that the task asks for a *per-column informative* statistic.
  2. In the scripts, identify the exact scaler/formula applied to each column and derive analytically what the reported statistic must equal under that transform.
  3. Flag the attempt if the derived statistic is mathematically constant regardless of the data (0 for means after centering, 1 for std after standardizing) — i.e., the reported numbers convey no dataset-specific information.
  4. Cross-check the reported numbers: any column whose value is exactly 0.0000/1.0000 while other, untransformed columns show data-dependent values is strong evidence of the mismatch; also confirm all requested output fields are present and consistently derived.
- **Discriminator**: A real violation is when the reported statistic is an artifact of the transform (constant by construction) or the scaling family contradicts the task's implied range; it is fine if the task explicitly names standardization and the reported statistic is still data-dependent (e.g., reporting means of columns that were not centered, or reporting min/max after standardizing).
- **Consequence**: The graded values for every rescaled column are 0.0000 (or otherwise uniform) instead of the expected data-dependent values, so only the untouched/binary columns match and the answer fails most checks.
114Rounded statistic reported without verifying the row set / preprocessing that produced ittaskinfiagent-dabench
Applies when
task -- the task asks for a single summary statistic (correlation, mean, metric) rounded to a fixed number of decimals, computed over two or more columns of a table.
Pattern
The agent loads the file with default settings and pipes the columns straight into the statistic, never checking how many rows actually entered the computation — silently dropping rows with missing/non-numeric/sentinel values, keeping rows the task's scope excludes, or reading only part of a multi-sheet/multi-file/multi-header source. The resulting value differs from the intended one in the second decimal, and because it is only reported rounded, the discrepancy is invisible in the final answer.
Detection procedure
  1. Read the task and note exactly which records are supposed to be in scope and to how many decimals the statistic must be reported.
  2. In the scripts, locate the load step and the statistic call; check whether the agent printed the number of rows loaded, the number of rows used after any NaN/dtype coercion, and the dtypes of the two columns.
  3. Check whether the unrounded statistic is printed with full precision and whether the agent looked at it relative to the rounding boundary (e.g. a value within ~0.005 of a .xx5 cut point, or an r whose 2-dp rounding could shift by one unit under a slightly different row set).
  4. Confirm the answer's rounding and format match the spec (correct decimal count for each field, no degenerate placeholders like an exactly-zero p-value where a small nonzero rounded value or "0.0000" is required).
Discriminator
A real violation is an attempt with no printed row/dtype/NaN diagnostics and no full-precision value, so the reported rounded number cannot be traced to a defensible record set. It is not a violation if the script prints the input shape, the count of rows actually used, confirms numeric dtypes and the missing-value policy, and reports the full-precision statistic — even if a reviewer might have chosen a different (documented) policy.
Consequence
The graded field is off by one unit in the last reported decimal (e.g. 0.53 vs 0.54), so an exact-match check on the coefficient fails even though the qualitative conclusion is right, and no saved script exists to diagnose which rows caused the drift.
id 0dffb42a35b6 · mined from infiagent-dabench dabench-300@s2
raw text (what the judge reads)
### Rounded statistic reported without verifying the row set / preprocessing that produced it
- **Applies when**: `task` -- the task asks for a single summary statistic (correlation, mean, metric) rounded to a fixed number of decimals, computed over two or more columns of a table.
- **Pattern**: The agent loads the file with default settings and pipes the columns straight into the statistic, never checking how many rows actually entered the computation — silently dropping rows with missing/non-numeric/sentinel values, keeping rows the task's scope excludes, or reading only part of a multi-sheet/multi-file/multi-header source. The resulting value differs from the intended one in the second decimal, and because it is only reported rounded, the discrepancy is invisible in the final answer.
- **Detection procedure**:
  1. Read the task and note exactly which records are supposed to be in scope and to how many decimals the statistic must be reported.
  2. In the scripts, locate the load step and the statistic call; check whether the agent printed the number of rows loaded, the number of rows used after any NaN/dtype coercion, and the dtypes of the two columns.
  3. Check whether the unrounded statistic is printed with full precision and whether the agent looked at it relative to the rounding boundary (e.g. a value within ~0.005 of a `.xx5` cut point, or an r whose 2-dp rounding could shift by one unit under a slightly different row set).
  4. Confirm the answer's rounding and format match the spec (correct decimal count for each field, no degenerate placeholders like an exactly-zero p-value where a small nonzero rounded value or "0.0000" is required).
- **Discriminator**: A real violation is an attempt with no printed row/dtype/NaN diagnostics and no full-precision value, so the reported rounded number cannot be traced to a defensible record set. It is *not* a violation if the script prints the input shape, the count of rows actually used, confirms numeric dtypes and the missing-value policy, and reports the full-precision statistic — even if a reviewer might have chosen a different (documented) policy.
- **Consequence**: The graded field is off by one unit in the last reported decimal (e.g. 0.53 vs 0.54), so an exact-match check on the coefficient fails even though the qualitative conclusion is right, and no saved script exists to diagnose which rows caused the drift.
115Accepting a weak model without benchmarking held-out error against the target's scale or a baselinetaskda-code
Applies when
task -- the deliverable is a file of predictions that will be scored against hidden ground truth, and the scripts fit models with default/naive feature handling and only self-report a validation score.
Pattern
The attempt encodes high-cardinality categorical fields with arbitrary integer codes (and/or drops informative columns), trains one or two off-the-shelf regressors with hand-picked hyperparameters, picks the "best" by a single random split, and declares success on the basis of a favorable-sounding R²/accuracy — without comparing the held-out error to the target's own dispersion, to a trivial baseline (mean/median or group-median), or to a stronger alternative model, and without checking that the predicted distribution matches the training target distribution.
Detection procedure
  1. Read the task: note that the grade depends on prediction quality against hidden labels, not on file existence, so any error level above the grader's tolerance fails.
  2. Read the scripts: check whether categorical/text features are given a mapping the model can exploit (target/one-hot/native categorical support) or only arbitrary ordinal codes; check whether any informative column is dropped; check whether more than a single 80/20 split and a couple of untuned models were tried.
  3. Compute the reported error in relative terms (RMSE ÷ target mean or std, or MAE ÷ median) and compare against a naive baseline the scripts should have run; also compare the reported min/max/mean of predictions with the training target's min/max/mean.
  4. Flag if no baseline/alternative comparison exists, if relative error is large (e.g., RMSE on the order of a third or more of the target mean), or if the prediction range is clipped/inflated relative to the training target range.
Discriminator
A genuine violation is an attempt whose only evidence of quality is an unanchored score from one split with clearly under-engineered features; it is not a violation when the agent explicitly benchmarks against a trivial baseline and at least one stronger/tuned model, uses cross-validation, and shows predicted-vs-actual distributions agreeing — even if the final score is modest.
Consequence
The submitted file has the right shape, column name and row count, so all format checks pass, but the accuracy check against ground truth falls outside tolerance and the task is graded incorrect.
id 173663b7d48d · mined from da-code dacode-ml-regression-014@s2
raw text (what the judge reads)
### Accepting a weak model without benchmarking held-out error against the target's scale or a baseline
- **Applies when**: `task` -- the deliverable is a file of predictions that will be scored against hidden ground truth, and the scripts fit models with default/naive feature handling and only self-report a validation score.
- **Pattern**: The attempt encodes high-cardinality categorical fields with arbitrary integer codes (and/or drops informative columns), trains one or two off-the-shelf regressors with hand-picked hyperparameters, picks the "best" by a single random split, and declares success on the basis of a favorable-sounding R²/accuracy — without comparing the held-out error to the target's own dispersion, to a trivial baseline (mean/median or group-median), or to a stronger alternative model, and without checking that the predicted distribution matches the training target distribution.
- **Detection procedure**:
  1. Read the task: note that the grade depends on prediction quality against hidden labels, not on file existence, so any error level above the grader's tolerance fails.
  2. Read the scripts: check whether categorical/text features are given a mapping the model can exploit (target/one-hot/native categorical support) or only arbitrary ordinal codes; check whether any informative column is dropped; check whether more than a single 80/20 split and a couple of untuned models were tried.
  3. Compute the reported error in relative terms (RMSE ÷ target mean or std, or MAE ÷ median) and compare against a naive baseline the scripts should have run; also compare the reported min/max/mean of predictions with the training target's min/max/mean.
  4. Flag if no baseline/alternative comparison exists, if relative error is large (e.g., RMSE on the order of a third or more of the target mean), or if the prediction range is clipped/inflated relative to the training target range.
- **Discriminator**: A genuine violation is an attempt whose only evidence of quality is an unanchored score from one split with clearly under-engineered features; it is *not* a violation when the agent explicitly benchmarks against a trivial baseline and at least one stronger/tuned model, uses cross-validation, and shows predicted-vs-actual distributions agreeing — even if the final score is modest.
- **Consequence**: The submitted file has the right shape, column name and row count, so all format checks pass, but the accuracy check against ground truth falls outside tolerance and the task is graded incorrect.
116Required output artifact not verified at the expected path/formattaskda-code
Applies when
task -- the task requires writing a deliverable file with a specified name, column(s), and one row per input record, and the agent claims it was produced.
Pattern
The attempt reports success narratively (row counts, class distribution) but never shows code/commands that (a) write the file to the expected location relative to the task's working/data directory, and (b) re-read the written file to confirm its path, header spelling, column count, row count and row order match the input records; the file is silently left in a home/scratch directory or with extra columns/index.
Detection procedure
  1. From the task, note the exact required filename, required column name(s), and the directory the grader will look in (normally the same directory as the provided data / current working dir).
  2. In the scripts, find the write call and check the path is that expected directory (not an absolute path elsewhere, not a temp dir) and that index/extra columns are suppressed and the header string matches exactly.
  3. Check for a post-write verification step: reload the file, assert row count equals the number of test records, assert column names, and print the resolved absolute path.
  4. In the answer, check whether the reported path is the expected deliverable location and whether the verification output (not just in-memory prediction stats) is quoted.
Discriminator
A real violation is an unverified or mislocated/misformatted artifact (path differs from the task's directory, no reload check, header/index mismatch, row count not tied to the test set size). It is fine if the script writes to the expected directory and prints a reload check confirming path, shape, and header — even if the model itself is simple or accuracy is modest.
Consequence
The grader reports the expected output file as WRONG/MISSING (or fails header/row-count checks), scoring 0 regardless of prediction quality.
id 8c3522afa537 · mined from da-code dacode-ml-multi-008@s2
raw text (what the judge reads)
### Required output artifact not verified at the expected path/format
- **Applies when**: `task` -- the task requires writing a deliverable file with a specified name, column(s), and one row per input record, and the agent claims it was produced.
- **Pattern**: The attempt reports success narratively (row counts, class distribution) but never shows code/commands that (a) write the file to the expected location relative to the task's working/data directory, and (b) re-read the written file to confirm its path, header spelling, column count, row count and row order match the input records; the file is silently left in a home/scratch directory or with extra columns/index.
- **Detection procedure**:
  1. From the task, note the exact required filename, required column name(s), and the directory the grader will look in (normally the same directory as the provided data / current working dir).
  2. In the scripts, find the write call and check the path is that expected directory (not an absolute path elsewhere, not a temp dir) and that index/extra columns are suppressed and the header string matches exactly.
  3. Check for a post-write verification step: reload the file, assert row count equals the number of test records, assert column names, and print the resolved absolute path.
  4. In the answer, check whether the reported path is the expected deliverable location and whether the verification output (not just in-memory prediction stats) is quoted.
- **Discriminator**: A real violation is an unverified or mislocated/misformatted artifact (path differs from the task's directory, no reload check, header/index mismatch, row count not tied to the test set size). It is fine if the script writes to the expected directory and prints a reload check confirming path, shape, and header — even if the model itself is simple or accuracy is modest.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING (or fails header/row-count checks), scoring 0 regardless of prediction quality.
117Incomplete/misplaced deliverable set relative to the task and its config spectaskda-code
Applies when
task -- the task asks for a saved artifact (chart, model, table) produced "according to" an external config/spec file, and the grader checks files on disk rather than just a printed answer.
Pattern
The attempt computes a plausible headline value and writes only the one file it explicitly noticed (e.g. the image), reading a couple of keys from the spec ad hoc; it never enumerates every key/required output in the spec (extra companion files such as the serialized plot data or plot parameters), never confirms the required file names, and writes to a personal working directory instead of the expected output location.
Detection procedure
  1. From the task text and the referenced spec/config file, list every required output artifact, its exact filename/extension, and its directory; also list every key in the config and the constraint it imposes (title, colors, size, labels, ordering, units, rounding, autopct, legend).
  2. Read the scripts and check off each item: is every artifact written, at the stated path, with the stated name? Is every config key actually consumed (not just the 3–4 the agent guessed at), and is nothing hard-coded that the config specifies?
  3. Check whether the underlying numbers behind the artifact are also persisted if the task/spec implies a machine-checkable dump (array/JSON of the plotted values), and whether category labels/order match the spec rather than data-driven sort order.
  4. Compare against the final answer: if the answer is only a scalar/string while the graded deliverables are files, treat unwritten or mislocated files as a failure regardless of whether the scalar is right.
Discriminator
A real violation is a missing/misnamed/misplaced required file or an unconsumed config key that changes the artifact's content; it is not a violation if the agent writes extra diagnostic files, or if it reformats internally while still emitting every required artifact at the required path with all spec constraints honored.
Consequence
File-level checks report WRONG/MISSING for the artifacts that were never created (and for the image if spec keys were ignored), so the attempt scores 0 even when the identified category/value is correct.
id b7ee8dc42678 · mined from da-code dacode-plot-pie-008@s2
raw text (what the judge reads)
### Incomplete/misplaced deliverable set relative to the task and its config spec
- **Applies when**: `task` -- the task asks for a saved artifact (chart, model, table) produced "according to" an external config/spec file, and the grader checks files on disk rather than just a printed answer.
- **Pattern**: The attempt computes a plausible headline value and writes only the one file it explicitly noticed (e.g. the image), reading a couple of keys from the spec ad hoc; it never enumerates every key/required output in the spec (extra companion files such as the serialized plot data or plot parameters), never confirms the required file names, and writes to a personal working directory instead of the expected output location.
- **Detection procedure**:
  1. From the task text and the referenced spec/config file, list every required output artifact, its exact filename/extension, and its directory; also list every key in the config and the constraint it imposes (title, colors, size, labels, ordering, units, rounding, autopct, legend).
  2. Read the scripts and check off each item: is every artifact written, at the stated path, with the stated name? Is every config key actually consumed (not just the 3–4 the agent guessed at), and is nothing hard-coded that the config specifies?
  3. Check whether the underlying numbers behind the artifact are also persisted if the task/spec implies a machine-checkable dump (array/JSON of the plotted values), and whether category labels/order match the spec rather than data-driven sort order.
  4. Compare against the final answer: if the answer is only a scalar/string while the graded deliverables are files, treat unwritten or mislocated files as a failure regardless of whether the scalar is right.
- **Discriminator**: A real violation is a missing/misnamed/misplaced required file or an unconsumed config key that changes the artifact's content; it is *not* a violation if the agent writes extra diagnostic files, or if it reformats internally while still emitting every required artifact at the required path with all spec constraints honored.
- **Consequence**: File-level checks report WRONG/MISSING for the artifacts that were never created (and for the image if spec keys were ignored), so the attempt scores 0 even when the identified category/value is correct.
118Fabricating output schema/category definitions instead of reading the provided template filetaskda-code
Applies when
task -- the task supplies a pre-existing result file (or explicit format spec) to be filled in, and the scripts must produce rows/labels/groupings that match it exactly.
Pattern
The attempt never opens the supplied file to learn its headers, row labels, ordering, or the intended grouping/bin definitions; instead it invents its own category names, boundaries, and extra rows from outside knowledge, then overwrites the file and declares success because the write succeeded.
Detection procedure
  1. In the task text, note that a target file/format is provided and identify what its rows and columns are supposed to encode.
  2. Search the scripts for any read/inspection of that target file (e.g., loading it, printing its head/columns/index) before writing; if the only interaction is a write, flag it.
  3. Check whether the grouping keys/labels used in the code are hard-coded from the agent's own assumptions (custom thresholds, self-chosen category names, an added "unknown/other" bucket) rather than derived from the template or the data's own labels.
  4. Cross-check the reported per-group counts for implausibility (a group with 0 or 1 members, or a group holding only a tiny fraction of records) that would indicate wrong boundaries/labels; confirm no sanity check on group sizes exists.
Discriminator
Fine if the scripts read the template (or the task spells out the exact labels) and the produced rows/labels/order match it, with only value cells changed; a violation is when labels, row set, or bin definitions are the agent's invention and no comparison against the provided format is ever made.
Consequence
The written file has mismatched row labels/extra rows and wrong per-group values, so an exact-match file check fails even though the script "verified" its own output.
id 979012ca7a82 · mined from da-code dacode-dm-csv-001@s2
raw text (what the judge reads)
### Fabricating output schema/category definitions instead of reading the provided template file
- **Applies when**: `task` -- the task supplies a pre-existing result file (or explicit format spec) to be filled in, and the scripts must produce rows/labels/groupings that match it exactly.
- **Pattern**: The attempt never opens the supplied file to learn its headers, row labels, ordering, or the intended grouping/bin definitions; instead it invents its own category names, boundaries, and extra rows from outside knowledge, then overwrites the file and declares success because the write succeeded.
- **Detection procedure**:
  1. In the task text, note that a target file/format is provided and identify what its rows and columns are supposed to encode.
  2. Search the scripts for any read/inspection of that target file (e.g., loading it, printing its head/columns/index) before writing; if the only interaction is a write, flag it.
  3. Check whether the grouping keys/labels used in the code are hard-coded from the agent's own assumptions (custom thresholds, self-chosen category names, an added "unknown/other" bucket) rather than derived from the template or the data's own labels.
  4. Cross-check the reported per-group counts for implausibility (a group with 0 or 1 members, or a group holding only a tiny fraction of records) that would indicate wrong boundaries/labels; confirm no sanity check on group sizes exists.
- **Discriminator**: Fine if the scripts read the template (or the task spells out the exact labels) and the produced rows/labels/order match it, with only value cells changed; a violation is when labels, row set, or bin definitions are the agent's invention and no comparison against the provided format is ever made.
- **Consequence**: The written file has mismatched row labels/extra rows and wrong per-group values, so an exact-match file check fails even though the script "verified" its own output.
119Entity-level aggregation and subgroup split defined on raw rows instead of per-unit recordstaskinfiagent-dabench
Applies when
task -- the data are multiple observation rows per underlying unit (event/subject/session) and the question asks for a statistic relating per-unit attributes (e.g., a maximum, a duration, a total) with subgroups defined by a threshold such as the median.
Pattern
The attempt correlates or thresholds using the raw row-level table (or aggregates with the wrong reducer/units), so each unit contributes many rows, the threshold (median) is computed over rows rather than over the per-unit values, and the resulting coefficient/subgroup membership is close to but not equal to the correct value.
Detection procedure
  1. From the task, list the per-unit quantities required (the aggregate of one variable, the span/duration of another, the magnitude used for the split) and note that the analysis unit is the entity, not the row.
  2. In the scripts, check that there is exactly one groupby(unit_id) producing one row per unit with the correct reducers (max, last−first for duration, sum/max for magnitude as specified) and correct units (e.g., hours vs days), and that duplicates/missing values are handled before aggregation.
  3. Verify the split threshold is computed on the aggregated per-unit series (not on raw rows, not after filtering) and that the two subgroups partition all units; print counts of units per subgroup and compare with the number of unique unit ids.
  4. Confirm the reported coefficient/p-value comes from the aggregated frame and matches the required rounding; a sanity check that group sizes are of plausible magnitude (tens–hundreds of units, not thousands of rows) should be present.
Discriminator
A real violation is when the correlation/threshold is computed over a table whose row count exceeds the number of unique units, or where the reducer/duration definition differs from the task wording; it is fine if the data are already one row per unit, or if aggregation is done with a documented equivalent reducer that yields identical per-unit values.
Consequence
Relationship-type labels may still come out right by luck, but the numeric coefficient (and subgroup p-values) drift from the ground truth by more than the two-decimal tolerance, so the answer is graded wrong on the coefficient checks.
id f24f38331096 · mined from infiagent-dabench dabench-431@s2
raw text (what the judge reads)
### Entity-level aggregation and subgroup split defined on raw rows instead of per-unit records
- **Applies when**: `task` -- the data are multiple observation rows per underlying unit (event/subject/session) and the question asks for a statistic relating per-unit attributes (e.g., a maximum, a duration, a total) with subgroups defined by a threshold such as the median.
- **Pattern**: The attempt correlates or thresholds using the raw row-level table (or aggregates with the wrong reducer/units), so each unit contributes many rows, the threshold (median) is computed over rows rather than over the per-unit values, and the resulting coefficient/subgroup membership is close to but not equal to the correct value.
- **Detection procedure**:
  1. From the task, list the per-unit quantities required (the aggregate of one variable, the span/duration of another, the magnitude used for the split) and note that the analysis unit is the entity, not the row.
  2. In the scripts, check that there is exactly one `groupby(unit_id)` producing one row per unit with the correct reducers (max, last−first for duration, sum/max for magnitude as specified) and correct units (e.g., hours vs days), and that duplicates/missing values are handled before aggregation.
  3. Verify the split threshold is computed on the aggregated per-unit series (not on raw rows, not after filtering) and that the two subgroups partition all units; print counts of units per subgroup and compare with the number of unique unit ids.
  4. Confirm the reported coefficient/p-value comes from the aggregated frame and matches the required rounding; a sanity check that group sizes are of plausible magnitude (tens–hundreds of units, not thousands of rows) should be present.
- **Discriminator**: A real violation is when the correlation/threshold is computed over a table whose row count exceeds the number of unique units, or where the reducer/duration definition differs from the task wording; it is fine if the data are already one row per unit, or if aggregation is done with a documented equivalent reducer that yields identical per-unit values.
- **Consequence**: Relationship-type labels may still come out right by luck, but the numeric coefficient (and subgroup p-values) drift from the ground truth by more than the two-decimal tolerance, so the answer is graded wrong on the coefficient checks.
120Hardcoded/recalled inputs and invented output schema instead of reading the provided filestaskda-code
Applies when
task -- the task references a provided dataset and/or an example output file (e.g., a sample result template), and the script defines the data or the output columns inline from memory.
Pattern
The agent never loads the supplied data/template files; it types in arrays it believes match the source and invents column names/row layout for the result file, so both the numbers and the file schema are unverified guesses. A related tell is a stated constraint (e.g., "set the random seed") that the chosen method makes irrelevant, signaling the wrong procedure was used.
Detection procedure
  1. From the task, list every artifact that must be read: the input data file(s) and any sample/template output file.
  2. Scan the script for a read/load call for each of those artifacts; flag if data values or the output header are literals rather than derived from a loaded file.
  3. Check whether any stated constraint (seed, rounding, units, ordering) is actually exercised by the code — an unused seed implies a deterministic/analytic method was substituted for the resampling-based one the task implies.
  4. Compare the written columns/ordering to the template's; if the template was never opened, treat the format as unverified.
Discriminator
Fine if the script loads the real files and only hardcodes trivially checkable metadata (paths, known constants) after verifying them, or reproduces the template header after reading it; a violation is when the substantive values or the output schema exist only in the script and are never cross-checked against the provided files.
Consequence
The result file mismatches the expected values and/or column format, so the grader marks the expected output file WRONG/MISSING even though the script ran without error.
id 656359908624 · mined from da-code dacode-data-sa-039@s2
raw text (what the judge reads)
### Hardcoded/recalled inputs and invented output schema instead of reading the provided files
- **Applies when**: `task` -- the task references a provided dataset and/or an example output file (e.g., a sample result template), and the script defines the data or the output columns inline from memory.
- **Pattern**: The agent never loads the supplied data/template files; it types in arrays it believes match the source and invents column names/row layout for the result file, so both the numbers and the file schema are unverified guesses. A related tell is a stated constraint (e.g., "set the random seed") that the chosen method makes irrelevant, signaling the wrong procedure was used.
- **Detection procedure**:
  1. From the task, list every artifact that must be read: the input data file(s) and any sample/template output file.
  2. Scan the script for a read/load call for each of those artifacts; flag if data values or the output header are literals rather than derived from a loaded file.
  3. Check whether any stated constraint (seed, rounding, units, ordering) is actually exercised by the code — an unused seed implies a deterministic/analytic method was substituted for the resampling-based one the task implies.
  4. Compare the written columns/ordering to the template's; if the template was never opened, treat the format as unverified.
- **Discriminator**: Fine if the script loads the real files and only hardcodes trivially checkable metadata (paths, known constants) after verifying them, or reproduces the template header after reading it; a violation is when the substantive values or the output schema exist only in the script and are never cross-checked against the provided files.
- **Consequence**: The result file mismatches the expected values and/or column format, so the grader marks the expected output file WRONG/MISSING even though the script ran without error.
121Unverifiable submission: no reproducible script or output validation for the requested prediction filetaskda-code
Applies when
task -- the task requires producing an output artifact (e.g., a predictions file with a specified column name and one row per test record) and the agent claims completion by naming the file.
Pattern
The attempt reports the deliverable's name without retaining the code that generated it, and never checks the written file's existence, row count against the test input, column naming, dtype, or value plausibility — so a missing, empty, misnamed, misaligned, or wrongly-formatted file passes unnoticed.
Detection procedure
  1. From the task, list the exact required artifact properties: filename/path, column name(s), expected number of rows (= number of test rows), value type/range, and any ordering requirement.
  2. In the scripts, locate the code that trains/predicts and writes the artifact; confirm it exists, reads the intended test file, writes to the exact required path, and uses the exact required header (and index=False where an index column would corrupt the format).
  3. Check for an explicit post-write verification step: re-read the saved file and assert shape/row count matches the test set, the required column is present, and there are no NaNs/constant or out-of-range predictions.
  4. Inspect the final answer: it must be backed by such a verification (printed shape, head, counts), not merely a filename assertion.
Discriminator
A real violation is when no generating script or no read-back/shape-and-header check exists, so correctness of the artifact is asserted rather than demonstrated; it is fine if the script writes the file and the log shows the re-read file's row count, column name, and sample values matching the test set even if no fancy validation framework is used.
Consequence
The grader looks for the artifact with the specified column and aligned rows and reports it WRONG/MISSING (0 checks passed) despite the answer claiming the file was produced.
id 94ba24209a35 · mined from da-code dacode-ml-regression-015@s2
raw text (what the judge reads)
### Unverifiable submission: no reproducible script or output validation for the requested prediction file
- **Applies when**: `task` -- the task requires producing an output artifact (e.g., a predictions file with a specified column name and one row per test record) and the agent claims completion by naming the file.
- **Pattern**: The attempt reports the deliverable's name without retaining the code that generated it, and never checks the written file's existence, row count against the test input, column naming, dtype, or value plausibility — so a missing, empty, misnamed, misaligned, or wrongly-formatted file passes unnoticed.
- **Detection procedure**:
  1. From the task, list the exact required artifact properties: filename/path, column name(s), expected number of rows (= number of test rows), value type/range, and any ordering requirement.
  2. In the scripts, locate the code that trains/predicts and writes the artifact; confirm it exists, reads the intended test file, writes to the exact required path, and uses the exact required header (and `index=False` where an index column would corrupt the format).
  3. Check for an explicit post-write verification step: re-read the saved file and assert shape/row count matches the test set, the required column is present, and there are no NaNs/constant or out-of-range predictions.
  4. Inspect the final answer: it must be backed by such a verification (printed shape, head, counts), not merely a filename assertion.
- **Discriminator**: A real violation is when no generating script or no read-back/shape-and-header check exists, so correctness of the artifact is asserted rather than demonstrated; it is fine if the script writes the file and the log shows the re-read file's row count, column name, and sample values matching the test set even if no fancy validation framework is used.
- **Consequence**: The grader looks for the artifact with the specified column and aligned rows and reports it WRONG/MISSING (0 checks passed) despite the answer claiming the file was produced.
122Unverified population coverage when computing a simple aggregate statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single descriptive statistic (mean, median, rate, count) over "all observations" of a field, and the scripts load the data from one or more files/tables before aggregating.
Pattern
The attempt computes the statistic on a silently reduced population — one file/partition/chunk out of several, a subset left over from an earlier filter or head/sample, rows lost to a join/merge or dtype coercion, or an aggressive "outlier"/dropna rule applied without justification — and reports the number without ever checking that the row count used matches the full source data.
Detection procedure
  1. From the task, note the intended population ("all observations") and any explicit filtering/rounding/format constraints.
  2. In the scripts, trace the data from load to aggregation: list every operation that can drop or add rows (file globbing, read_* with nrows/chunks, filters, merges, dropna, outlier thresholds, astype/to_numeric coercions) and check whether the final denominator is asserted or printed against the raw record count of the source.
  3. Check that the statistic is computed on the intended column/axis of the full frame, not on an intermediate, grouped, or per-file result that is then averaged unweighted.
  4. Confirm a reproducible script exists and prints the row count and value; if no script/counts are shown, the number cannot be validated at all.
Discriminator
A real violation is row loss or subsetting that is neither required by the task nor verified by a count/shape check (including averaging per-group means instead of the pooled mean, or discarding valid values as "outliers" with an arbitrary cutoff). It is fine if the script explicitly documents and verifies the exclusions the task requested (e.g., only missing values removed) and shows that used_rows + excluded_rows equals total source rows.
Consequence
The reported value is systematically biased and misses the expected number beyond the grader's tolerance, so the check fails even though the code runs without error.
id 7d7fe9cdcce8 · mined from infiagent-dabench dabench-320@s2
raw text (what the judge reads)
### Unverified population coverage when computing a simple aggregate statistic
- **Applies when**: `task` -- the task asks for a single descriptive statistic (mean, median, rate, count) over "all observations" of a field, and the scripts load the data from one or more files/tables before aggregating.
- **Pattern**: The attempt computes the statistic on a silently reduced population — one file/partition/chunk out of several, a subset left over from an earlier filter or `head`/sample, rows lost to a join/merge or dtype coercion, or an aggressive "outlier"/dropna rule applied without justification — and reports the number without ever checking that the row count used matches the full source data.
- **Detection procedure**:
  1. From the task, note the intended population ("all observations") and any explicit filtering/rounding/format constraints.
  2. In the scripts, trace the data from load to aggregation: list every operation that can drop or add rows (file globbing, `read_*` with `nrows`/chunks, filters, merges, `dropna`, outlier thresholds, `astype`/`to_numeric` coercions) and check whether the final denominator is asserted or printed against the raw record count of the source.
  3. Check that the statistic is computed on the intended column/axis of the full frame, not on an intermediate, grouped, or per-file result that is then averaged unweighted.
  4. Confirm a reproducible script exists and prints the row count and value; if no script/counts are shown, the number cannot be validated at all.
- **Discriminator**: A real violation is row loss or subsetting that is neither required by the task nor verified by a count/shape check (including averaging per-group means instead of the pooled mean, or discarding valid values as "outliers" with an arbitrary cutoff). It is fine if the script explicitly documents and verifies the exclusions the task requested (e.g., only missing values removed) and shows that used_rows + excluded_rows equals total source rows.
- **Consequence**: The reported value is systematically biased and misses the expected number beyond the grader's tolerance, so the check fails even though the code runs without error.
123Numeric columns ranked without verifying they were parsed as numbers (string-formatted values → wrong extremes)taskda-code
Applies when
task -- the task asks for the maximum/minimum (or mean-imputed) value of a quantitative field read from a raw CSV/Excel file where values may carry thousands separators, currency/percent symbols, or other text formatting.
Pattern
The attempt loads the file, imputes with mean(), and takes idxmax/idxmin (or sort_values().head()) without checking the column's dtype. If the column is object, imputation silently skips it and the ordering becomes lexicographic (or the fill/parse coerces values to NaN), so the reported extreme rows are wrong even though the code runs without error.
Detection procedure
  1. Read the task to identify which column(s) the answer depends on and whether any transformation (imputation, rounding, unit conversion) must precede the ranking.
  2. In the scripts, look for an explicit cleaning/coercion step for those columns (e.g., stripping separators/symbols then astype(float) / pd.to_numeric(..., errors='coerce')) and an assertion or printout of the resulting dtype, NaN count, and min/max.
  3. Check that the imputation actually affected the target column (count of NaNs before/after; fill value printed) rather than being a no-op on an object column.
  4. Compare the reported extreme entities against a quick plausibility check (are the top/bottom values consistent with the column's known scale and with other rows sorted numerically?); a "lowest" that is clearly not the smallest numerically, or extremes whose leading digits look alphabetically ordered, indicates string sorting.
Discriminator
A real violation is when no dtype/coercion evidence exists and the column plausibly contains formatted text, or the printed extremes are inconsistent with a numeric ordering. It is fine if the script demonstrates the column is already numeric (dtype check, describe() output) or explicitly cleans it, even if the final answer happens to be uncommon.
Consequence
The graded JSON names the wrong entity for the min and/or max (often the max is coincidentally right while the min is wrong), so the result file fails the exact-match check (0/1).
id 03064c04e264 · mined from da-code dacode-di-text-001@s2
raw text (what the judge reads)
### Numeric columns ranked without verifying they were parsed as numbers (string-formatted values → wrong extremes)

- **Applies when**: `task` -- the task asks for the maximum/minimum (or mean-imputed) value of a quantitative field read from a raw CSV/Excel file where values may carry thousands separators, currency/percent symbols, or other text formatting.
- **Pattern**: The attempt loads the file, imputes with `mean()`, and takes `idxmax`/`idxmin` (or `sort_values().head()`) without checking the column's dtype. If the column is `object`, imputation silently skips it and the ordering becomes lexicographic (or the fill/parse coerces values to NaN), so the reported extreme rows are wrong even though the code runs without error.
- **Detection procedure**:
  1. Read the task to identify which column(s) the answer depends on and whether any transformation (imputation, rounding, unit conversion) must precede the ranking.
  2. In the scripts, look for an explicit cleaning/coercion step for those columns (e.g., stripping separators/symbols then `astype(float)` / `pd.to_numeric(..., errors='coerce')`) and an assertion or printout of the resulting dtype, NaN count, and min/max.
  3. Check that the imputation actually affected the target column (count of NaNs before/after; fill value printed) rather than being a no-op on an object column.
  4. Compare the reported extreme entities against a quick plausibility check (are the top/bottom values consistent with the column's known scale and with other rows sorted numerically?); a "lowest" that is clearly not the smallest numerically, or extremes whose leading digits look alphabetically ordered, indicates string sorting.
- **Discriminator**: A real violation is when no dtype/coercion evidence exists and the column plausibly contains formatted text, or the printed extremes are inconsistent with a numeric ordering. It is fine if the script demonstrates the column is already numeric (dtype check, describe() output) or explicitly cleans it, even if the final answer happens to be uncommon.
- **Consequence**: The graded JSON names the wrong entity for the min and/or max (often the max is coincidentally right while the min is wrong), so the result file fails the exact-match check (0/1).
124Output file omits requested derived columns (only the final label is saved)taskda-code
Applies when
task -- the task asks to compute several intermediate quantities/scores and a final grouping/label, and to save "the results including X and Y" to a specified file.
Pattern
The script computes all the intermediate metrics and scores in memory, prints them, but writes only a minimal subset (e.g., an ID plus the final label) to the output file, dropping the per-item metrics/scores that the task explicitly said to include; the answer then asserts the file "matches the reference exactly" without any comparison actually being executed.
Detection procedure
  1. From the task statement, enumerate every quantity that must appear in the saved artifact (each named metric, each score, the grouping/segment, the final level, the identifier).
  2. In the script, find the line that constructs the DataFrame written to the output path and list its columns (and index/header settings).
  3. Diff the two lists; also check whether row count/ordering constraints implied by the task are preserved.
  4. In the answer, check whether any claim of "verified against reference/expected output" is backed by a comparison actually performed in the code shown.
Discriminator
A real violation is when a task-named quantity is computed but excluded from the saved file (or renamed beyond recognition/serialized with wrong index-header settings). It is not a violation if the task only asked for the final label, or if the extra quantities are present under clearly equivalent names.
Consequence
The saved file fails column/shape/content comparison against the expected artifact, so the check is scored wrong even if the underlying computation logic was reasonable, and the answer's unverified "matches exactly" claim is false.
id aad11a4a2b66 · mined from da-code dacode-dm-csv-052@s2
raw text (what the judge reads)
### Output file omits requested derived columns (only the final label is saved)
- **Applies when**: `task` -- the task asks to compute several intermediate quantities/scores *and* a final grouping/label, and to save "the results including X and Y" to a specified file.
- **Pattern**: The script computes all the intermediate metrics and scores in memory, prints them, but writes only a minimal subset (e.g., an ID plus the final label) to the output file, dropping the per-item metrics/scores that the task explicitly said to include; the answer then asserts the file "matches the reference exactly" without any comparison actually being executed.
- **Detection procedure**:
  1. From the task statement, enumerate every quantity that must appear in the saved artifact (each named metric, each score, the grouping/segment, the final level, the identifier).
  2. In the script, find the line that constructs the DataFrame written to the output path and list its columns (and index/header settings).
  3. Diff the two lists; also check whether row count/ordering constraints implied by the task are preserved.
  4. In the answer, check whether any claim of "verified against reference/expected output" is backed by a comparison actually performed in the code shown.
- **Discriminator**: A real violation is when a task-named quantity is computed but excluded from the saved file (or renamed beyond recognition/serialized with wrong index-header settings). It is *not* a violation if the task only asked for the final label, or if the extra quantities are present under clearly equivalent names.
- **Consequence**: The saved file fails column/shape/content comparison against the expected artifact, so the check is scored wrong even if the underlying computation logic was reasonable, and the answer's unverified "matches exactly" claim is false.
125Output schema invented instead of copied from the provided template filetaskda-code
Applies when
task -- the task says to write results into an output file "following the format specified in" a provided sample/template file.
Pattern
The script never reads or inspects the template; it hard-codes guessed column names, guessed cell encodings (e.g., packing a two-number interval into one string literal, or splitting it across columns), guessed row order/precision, and sometimes guesses the output directory too — so the numbers may be right while the file fails any schema-aware comparison.
Detection procedure
  1. In the task text, note every referenced template/sample artifact and the required output filename/location.
  2. Search the scripts for any read of that template (read_csv, open, printing its header/rows) and any assertion that produced columns/dtypes/row count match it.
  3. If absent, compare the script's constructed column labels, value formatting, and write path against what the task literally names; treat guessed labels or ad-hoc string formatting of numeric fields as the violation.
  4. Check the reported answer: does it show the final file contents and evidence that they conform to the template, or only the computed statistics?
Discriminator
Fine if the script loads the template (or an explicit verbatim quote of its header appears) and builds the output to that schema, even with minor cosmetic differences; a real violation is when the schema/format/path is asserted from assumption with no reference to the provided file.
Consequence
The grader reports the expected output file as WRONG/MISSING despite statistically correct values, because column names, cell encoding, or file location do not match the reference.
id ba8a0aab4bf8 · mined from da-code dacode-data-sa-029@s2
raw text (what the judge reads)
### Output schema invented instead of copied from the provided template file
- **Applies when**: `task` -- the task says to write results into an output file "following the format specified in" a provided sample/template file.
- **Pattern**: The script never reads or inspects the template; it hard-codes guessed column names, guessed cell encodings (e.g., packing a two-number interval into one string literal, or splitting it across columns), guessed row order/precision, and sometimes guesses the output directory too — so the numbers may be right while the file fails any schema-aware comparison.
- **Detection procedure**:
  1. In the task text, note every referenced template/sample artifact and the required output filename/location.
  2. Search the scripts for any read of that template (`read_csv`, `open`, printing its header/rows) and any assertion that produced columns/dtypes/row count match it.
  3. If absent, compare the script's constructed column labels, value formatting, and write path against what the task literally names; treat guessed labels or ad-hoc string formatting of numeric fields as the violation.
  4. Check the reported answer: does it show the final file contents *and* evidence that they conform to the template, or only the computed statistics?
- **Discriminator**: Fine if the script loads the template (or an explicit verbatim quote of its header appears) and builds the output to that schema, even with minor cosmetic differences; a real violation is when the schema/format/path is asserted from assumption with no reference to the provided file.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING despite statistically correct values, because column names, cell encoding, or file location do not match the reference.
126Ignoring the provided sample/template output file when writing the resulttaskda-code
Applies when
task -- The task supplies a reference/sample output file (or explicitly states a required output layout) and asks that the saved result "follow the same format".
Pattern
The scripts compute the requested statistic and dump it with a default writer (e.g., to_csv straight from a DataFrame) without ever reading or inspecting the sample file, so the header name of the index column, column/row order and labels, rounding/precision, or inclusion of an index are never verified against the template.
Detection procedure
  1. Read the task for any mention of a sample/example/template output file or an explicit format constraint on the saved artifact.
  2. Search the scripts for any read of that sample file (or a hard-coded replication of its structure) and for an explicit comparison of the produced file's shape, header, index labels, ordering, and numeric precision against it.
  3. Check the answer/report for evidence that the written file was compared cell-for-cell (or at least header-for-header) with the sample, not merely re-read and printed.
  4. Flag if the only validation is printing the computed object, or if internal counts in the report are mutually inconsistent (e.g., reported per-column missing counts that cannot produce the reported number of dropped rows), indicating no sanity check on the pipeline.
Discriminator
A real violation is when no sample inspection/verification exists at all, or the writer's defaults clearly differ from the stated template (missing index header, unrounded values, different label order). It is fine if the script reads the sample and asserts matching structure, or if the sample's structure is demonstrably reproduced (identical column/index names and ordering) even without a literal diff.
Consequence
The grader compares the submitted file to the expected one and marks it WRONG/MISSING even though the underlying statistic may be numerically right, yielding 0 checks passed.
id eda4a264c2d5 · mined from da-code dacode-data-sa-026@s2
raw text (what the judge reads)
### Ignoring the provided sample/template output file when writing the result
- **Applies when**: `task` -- The task supplies a reference/sample output file (or explicitly states a required output layout) and asks that the saved result "follow the same format".
- **Pattern**: The scripts compute the requested statistic and dump it with a default writer (e.g., `to_csv` straight from a DataFrame) without ever reading or inspecting the sample file, so the header name of the index column, column/row order and labels, rounding/precision, or inclusion of an index are never verified against the template.
- **Detection procedure**:
  1. Read the task for any mention of a sample/example/template output file or an explicit format constraint on the saved artifact.
  2. Search the scripts for any read of that sample file (or a hard-coded replication of its structure) and for an explicit comparison of the produced file's shape, header, index labels, ordering, and numeric precision against it.
  3. Check the answer/report for evidence that the written file was compared cell-for-cell (or at least header-for-header) with the sample, not merely re-read and printed.
  4. Flag if the only validation is printing the computed object, or if internal counts in the report are mutually inconsistent (e.g., reported per-column missing counts that cannot produce the reported number of dropped rows), indicating no sanity check on the pipeline.
- **Discriminator**: A real violation is when no sample inspection/verification exists at all, or the writer's defaults clearly differ from the stated template (missing index header, unrounded values, different label order). It is fine if the script reads the sample and asserts matching structure, or if the sample's structure is demonstrably reproduced (identical column/index names and ordering) even without a literal diff.
- **Consequence**: The grader compares the submitted file to the expected one and marks it WRONG/MISSING even though the underlying statistic may be numerically right, yielding 0 checks passed.
127Silently subsampling the dataset when the deliverable must cover all recordstaskda-code
Applies when
task -- the task asks for per-row outputs (labels, predictions, scores) written to a file, and the scripts read a large input table.
Pattern
The attempt reads only a slice/sample of the input (e.g., nrows=, head(), sample(n=...), a hard-coded cap for speed) and produces an output file with far fewer rows than the source data, without the task authorizing any subsetting.
Detection procedure
  1. From the task/README, note the expected scope of the output (one row per input record) and the documented size of the source data.
  2. In the scripts, grep the data-loading and output-writing steps for row limits, sampling, filtering, or dropna that shrinks the record count, and check whether labels are mapped back to all original rows.
  3. Compare the row count reported in the answer (and in the written file) against the source row count; flag any unexplained shortfall.
  4. Also confirm the column names/order in the written file match the exact schema the task specifies.
Discriminator
A real violation is an undisclosed or convenience-driven reduction of the deliverable's coverage; it is fine if the task itself specifies a subset/filter, or if sampling is used only for internal model selection/hyperparameter search while the final output is applied to every record.
Consequence
The expected output file fails comparison outright — row count (and hence any per-row label alignment) does not match the reference, so the check scores 0 regardless of clustering quality.
id 6b26eac67cb2 · mined from da-code dacode-ml-cluster-010@s2
raw text (what the judge reads)
### Silently subsampling the dataset when the deliverable must cover all records
- **Applies when**: `task` -- the task asks for per-row outputs (labels, predictions, scores) written to a file, and the scripts read a large input table.
- **Pattern**: The attempt reads only a slice/sample of the input (e.g., `nrows=`, `head()`, `sample(n=...)`, a hard-coded cap for speed) and produces an output file with far fewer rows than the source data, without the task authorizing any subsetting.
- **Detection procedure**:
  1. From the task/README, note the expected scope of the output (one row per input record) and the documented size of the source data.
  2. In the scripts, grep the data-loading and output-writing steps for row limits, sampling, filtering, or `dropna` that shrinks the record count, and check whether labels are mapped back to all original rows.
  3. Compare the row count reported in the answer (and in the written file) against the source row count; flag any unexplained shortfall.
  4. Also confirm the column names/order in the written file match the exact schema the task specifies.
- **Discriminator**: A real violation is an undisclosed or convenience-driven reduction of the deliverable's coverage; it is fine if the task itself specifies a subset/filter, or if sampling is used only for internal model selection/hyperparameter search while the final output is applied to every record.
- **Consequence**: The expected output file fails comparison outright — row count (and hence any per-row label alignment) does not match the reference, so the check scores 0 regardless of clustering quality.
128Degenerate clustering accepted because the selection metric was inflated by skewed, untransformed featurestaskda-code
Applies when
task -- the task asks for an unsupervised grouping into "an appropriate number of clusters" and the script picks k by an internal score (silhouette/inertia) on aggregated numeric features written to an output file.
Pattern
The attempt aggregates heavy-tailed count/amount features, applies only mean/variance standardization (no log/rank/quantile transform, no outlier handling), then picks k by the highest silhouette — which is maximized by isolating a handful of extreme records. The reported partition is effectively "everyone" vs. "a few outliers" (e.g. one cluster with >70% of rows, others with a handful), and the agent declares success without sanity-checking cluster sizes or feature distributions.
Detection procedure
  1. From the task, note that the deliverable is a meaningful segmentation, not just any label column, and note the required column names/format.
  2. In the script, check whether skewed monetary/count features are transformed (log1p, quantile, robust scaling) or extremes trimmed before distance-based clustering, and whether k is chosen by more than a single internal score maximum.
  3. In the answer, read the per-cluster row counts and compare with the number of rows: flag if one cluster holds the large majority while one or more clusters hold a negligible number of rows (<1%), or if the score is suspiciously high (e.g. >0.5) for high-dimensional real data.
  4. Check that the saved file's features match what was clustered (same transform/scaling) and that row count equals the number of entities clustered.
Discriminator
A genuinely imbalanced but valid solution has clusters that are all substantively populated and separations that survive a transform/robust-scaling check; a violation is a partition whose small clusters are a few extreme records and whose score collapses (or whose k changes) once the skew is handled — i.e. the metric, not the structure, drove the choice.
Consequence
The saved label file encodes an outlier-detection split rather than customer segments, so any grader comparison against a reasonable reference partition (cluster count, size distribution, agreement score) fails, marking the output file wrong.
id b26c01c1cc98 · mined from da-code dacode-ml-cluster-019@s2
raw text (what the judge reads)
### Degenerate clustering accepted because the selection metric was inflated by skewed, untransformed features
- **Applies when**: `task` -- the task asks for an unsupervised grouping into "an appropriate number of clusters" and the script picks k by an internal score (silhouette/inertia) on aggregated numeric features written to an output file.
- **Pattern**: The attempt aggregates heavy-tailed count/amount features, applies only mean/variance standardization (no log/rank/quantile transform, no outlier handling), then picks k by the highest silhouette — which is maximized by isolating a handful of extreme records. The reported partition is effectively "everyone" vs. "a few outliers" (e.g. one cluster with >70% of rows, others with a handful), and the agent declares success without sanity-checking cluster sizes or feature distributions.
- **Detection procedure**:
  1. From the task, note that the deliverable is a meaningful segmentation, not just any label column, and note the required column names/format.
  2. In the script, check whether skewed monetary/count features are transformed (log1p, quantile, robust scaling) or extremes trimmed before distance-based clustering, and whether k is chosen by more than a single internal score maximum.
  3. In the answer, read the per-cluster row counts and compare with the number of rows: flag if one cluster holds the large majority while one or more clusters hold a negligible number of rows (<1%), or if the score is suspiciously high (e.g. >0.5) for high-dimensional real data.
  4. Check that the saved file's features match what was clustered (same transform/scaling) and that row count equals the number of entities clustered.
- **Discriminator**: A genuinely imbalanced but valid solution has clusters that are all substantively populated and separations that survive a transform/robust-scaling check; a violation is a partition whose small clusters are a few extreme records and whose score collapses (or whose k changes) once the skew is handled — i.e. the metric, not the structure, drove the choice.
- **Consequence**: The saved label file encodes an outlier-detection split rather than customer segments, so any grader comparison against a reasonable reference partition (cluster count, size distribution, agreement score) fails, marking the output file wrong.
129Hardcoded/recalled inputs instead of loading the provided data and output templatetaskda-code
Applies when
task -- the task supplies a dataset (and often a template/example of the expected output file), and the script must derive its numbers from those files.
Pattern
The script embeds constants typed from memory or background knowledge (totals, counts, rates, group definitions) and never reads the supplied data files; likewise it invents the output file's columns/rows rather than reproducing the provided format. Any downstream statistic is then computed on numbers that do not match the actual data, and the file fails a format-sensitive check.
Detection procedure
  1. Read the task for (a) which data files are provided and (b) any stated output format/template for the result file.
  2. Scan the scripts for a load step (read_csv/read_*/file open) on those files; flag if the key quantities appear as literal constants with comments like "historical data" or "known values", or if grouping/filtering of the raw records is never performed in code.
  3. Compare the written result file's column names, ordering, units, and number of rows against the template given in the task; flag any invented headers or extra/missing fields.
  4. Check the answer for a traceability statement linking each reported number back to a computation on the loaded file (e.g., printed row counts, group sizes) rather than to prior knowledge.
Discriminator
A genuine violation is when the analysis quantities themselves are literals never verified against the loaded data, or when the output schema is self-invented. It is not a violation to hardcode auxiliary constants (random seed, number of bootstrap iterations, confidence level, a documented cutoff date) as long as all reported statistics are computed from records read out of the provided files and printed sanity checks (row counts, group totals) confirm agreement.
Consequence
The reported interval/statistic is computed from wrong inputs and the result file's schema does not match the expected one, so an exact-match or tolerance-based file check on result.csv reports WRONG/MISSING even though the method (bootstrap percentile CI) looks correct.
id c7cb89db100b · mined from da-code dacode-data-sa-031@s2
raw text (what the judge reads)
### Hardcoded/recalled inputs instead of loading the provided data and output template
- **Applies when**: `task` -- the task supplies a dataset (and often a template/example of the expected output file), and the script must derive its numbers from those files.
- **Pattern**: The script embeds constants typed from memory or background knowledge (totals, counts, rates, group definitions) and never reads the supplied data files; likewise it invents the output file's columns/rows rather than reproducing the provided format. Any downstream statistic is then computed on numbers that do not match the actual data, and the file fails a format-sensitive check.
- **Detection procedure**:
  1. Read the task for (a) which data files are provided and (b) any stated output format/template for the result file.
  2. Scan the scripts for a load step (`read_csv`/`read_*`/file open) on those files; flag if the key quantities appear as literal constants with comments like "historical data" or "known values", or if grouping/filtering of the raw records is never performed in code.
  3. Compare the written result file's column names, ordering, units, and number of rows against the template given in the task; flag any invented headers or extra/missing fields.
  4. Check the answer for a traceability statement linking each reported number back to a computation on the loaded file (e.g., printed row counts, group sizes) rather than to prior knowledge.
- **Discriminator**: A genuine violation is when the analysis quantities themselves are literals never verified against the loaded data, or when the output schema is self-invented. It is *not* a violation to hardcode auxiliary constants (random seed, number of bootstrap iterations, confidence level, a documented cutoff date) as long as all reported statistics are computed from records read out of the provided files and printed sanity checks (row counts, group totals) confirm agreement.
- **Consequence**: The reported interval/statistic is computed from wrong inputs and the result file's schema does not match the expected one, so an exact-match or tolerance-based file check on `result.csv` reports WRONG/MISSING even though the method (bootstrap percentile CI) looks correct.
130Overconfident, unvalidated probability outputs for a log-loss deliverable (single holdout fit, no baseline, no clipping/calibration)taskda-code
Applies when
task -- the task is scored by a probabilistic metric (log loss / Brier) and the script trains one model, evaluates it on a single random split, and writes predict_proba output straight to the submission file.
Pattern
The attempt fits a high-capacity model (deep trees / many boosting rounds) on only the train portion of a single split, never refits on the full labeled data, never cross-validates or tunes, never compares the score to a trivial baseline (class-prior constants), and never bounds/calibrates the emitted probabilities — so the file contains values on the order of 1e-5 or smaller for minority classes, where a handful of misclassified rows dominate the metric.
Detection procedure
  1. Read the task statement and note the exact scoring metric and that it penalizes confident errors logarithmically.
  2. In the script, check whether the object that generates the test predictions was fit on all labeled rows, whether any resampling/CV or hyperparameter selection was done, and whether the reported validation score is compared against a naive baseline (e.g., predicting the training class frequencies).
  3. Inspect the emitted probability columns in the answer: look at the minimum values per class and how many rows are near 0 or near 1 for a rare class.
  4. Flag if the final model saw only a fraction of the data, or if extreme probabilities are produced with no clipping/calibration and no evidence the CV score beats the prior baseline.
Discriminator
A fine attempt may also produce some confident rows, but it (a) refits on the full training set, (b) reports a cross-validated metric that clearly beats the constant-prior baseline, and (c) either uses a calibrated/regularized model or clips/smooths probabilities so rare-class predictions are not orders of magnitude below their base rate. A violation is a single deep-model holdout fit whose only quality evidence is one split's score, with untempered near-zero probabilities.
Consequence
The graded log loss is far worse than the reported validation number — often worse than a constant-prior submission — because a few confidently wrong rows contribute huge penalties, so the submission fails the accuracy threshold.
id 2083193f1dac · mined from da-code dacode-ml-competition-005@s3
raw text (what the judge reads)
### Overconfident, unvalidated probability outputs for a log-loss deliverable (single holdout fit, no baseline, no clipping/calibration)
- **Applies when**: `task` -- the task is scored by a probabilistic metric (log loss / Brier) and the script trains one model, evaluates it on a single random split, and writes `predict_proba` output straight to the submission file.
- **Pattern**: The attempt fits a high-capacity model (deep trees / many boosting rounds) on only the train portion of a single split, never refits on the full labeled data, never cross-validates or tunes, never compares the score to a trivial baseline (class-prior constants), and never bounds/calibrates the emitted probabilities — so the file contains values on the order of 1e-5 or smaller for minority classes, where a handful of misclassified rows dominate the metric.
- **Detection procedure**:
  1. Read the task statement and note the exact scoring metric and that it penalizes confident errors logarithmically.
  2. In the script, check whether the object that generates the test predictions was fit on *all* labeled rows, whether any resampling/CV or hyperparameter selection was done, and whether the reported validation score is compared against a naive baseline (e.g., predicting the training class frequencies).
  3. Inspect the emitted probability columns in the answer: look at the minimum values per class and how many rows are near 0 or near 1 for a rare class.
  4. Flag if the final model saw only a fraction of the data, or if extreme probabilities are produced with no clipping/calibration and no evidence the CV score beats the prior baseline.
- **Discriminator**: A fine attempt may also produce some confident rows, but it (a) refits on the full training set, (b) reports a cross-validated metric that clearly beats the constant-prior baseline, and (c) either uses a calibrated/regularized model or clips/smooths probabilities so rare-class predictions are not orders of magnitude below their base rate. A violation is a single deep-model holdout fit whose only quality evidence is one split's score, with untempered near-zero probabilities.
- **Consequence**: The graded log loss is far worse than the reported validation number — often worse than a constant-prior submission — because a few confidently wrong rows contribute huge penalties, so the submission fails the accuracy threshold.
131No held-out validation or distribution sanity check before submitting predictionstaskda-code
Applies when
task -- the deliverable is a file of predicted values produced by a model fit on separate training files, and the scripts go straight from fit to predict to writing the output.
Pattern
The attempt trains a single model with hand-picked hyperparameters on all available labeled rows, never holds out a validation split (or cross-validates), never computes an error metric, never compares alternative models/targets (e.g. raw vs log-transformed, regression vs baseline median), and never checks that the predicted distribution resembles the labeled target distribution. It also skips checks on the label column itself (missing/NaN targets, extreme skew, dtype/format) and on feature alignment between train and predict frames; the final answer reports only self-descriptive stats (min/max/mean of predictions) as evidence of success.
Detection procedure
  1. Read the task to confirm the output is scored against unseen true values, so predictive accuracy — not file existence — is what matters.
  2. Scan the scripts for any train/validation split, cross-validation, or metric computation on labeled data; note whether the target is cleaned (NaN/zero/dtype handling) and whether any comparison against a trivial baseline is made.
  3. Compare the reported prediction summary statistics (mean, median, max, share of zeros) with the summary statistics of the target in the training data printed earlier; a large mismatch (e.g. predicted median orders of magnitude below the training median, or predictions collapsed near zero) is a red flag.
  4. Check that the written file's columns/rows match exactly what the task specified (required column name, row count and order matching the input file, no extra or renamed columns).
Discriminator
A real violation is when no out-of-sample error estimate or baseline comparison exists anywhere, so the agent cannot know its model is better than a constant prediction; it is not a violation if the agent validated (split/CV metric or baseline comparison) and consciously chose a simple model, nor if a distribution shift between predictions and training labels is explained and justified by the validated setup.
Consequence
The submitted file is accepted structurally but scores poorly on the grader's accuracy tolerance (predictions systematically biased/shrunk toward zero or dominated by one feature), yielding a WRONG file verdict despite a confident "task completed" report.
id 54a6e4a38438 · mined from da-code dacode-ml-regression-008@s3
raw text (what the judge reads)
### No held-out validation or distribution sanity check before submitting predictions
- **Applies when**: `task` -- the deliverable is a file of predicted values produced by a model fit on separate training files, and the scripts go straight from `fit` to `predict` to writing the output.
- **Pattern**: The attempt trains a single model with hand-picked hyperparameters on all available labeled rows, never holds out a validation split (or cross-validates), never computes an error metric, never compares alternative models/targets (e.g. raw vs log-transformed, regression vs baseline median), and never checks that the predicted distribution resembles the labeled target distribution. It also skips checks on the label column itself (missing/NaN targets, extreme skew, dtype/format) and on feature alignment between train and predict frames; the final answer reports only self-descriptive stats (min/max/mean of predictions) as evidence of success.
- **Detection procedure**:
  1. Read the task to confirm the output is scored against unseen true values, so predictive accuracy — not file existence — is what matters.
  2. Scan the scripts for any train/validation split, cross-validation, or metric computation on labeled data; note whether the target is cleaned (NaN/zero/dtype handling) and whether any comparison against a trivial baseline is made.
  3. Compare the reported prediction summary statistics (mean, median, max, share of zeros) with the summary statistics of the target in the training data printed earlier; a large mismatch (e.g. predicted median orders of magnitude below the training median, or predictions collapsed near zero) is a red flag.
  4. Check that the written file's columns/rows match exactly what the task specified (required column name, row count and order matching the input file, no extra or renamed columns).
- **Discriminator**: A real violation is when *no* out-of-sample error estimate or baseline comparison exists anywhere, so the agent cannot know its model is better than a constant prediction; it is *not* a violation if the agent validated (split/CV metric or baseline comparison) and consciously chose a simple model, nor if a distribution shift between predictions and training labels is explained and justified by the validated setup.
- **Consequence**: The submitted file is accepted structurally but scores poorly on the grader's accuracy tolerance (predictions systematically biased/shrunk toward zero or dominated by one feature), yielding a WRONG file verdict despite a confident "task completed" report.
132Unjustified default hypothesis test on the full raw data (no population filter, no test-choice/direction check)taskda-code
Applies when
task -- the task asks for a p-value and a reject/fail-to-reject decision comparing two groups, and the scripts run a single canned test on every row of the loaded files.
Pattern
The attempt equates "the population in the question" with "all rows in the files" and uses library defaults (parametric two-sample, two-sided, equal-variance) without inspecting whether the data support them: no restriction to the comparable subset implied by the question's framing (era/competition/level, matching time windows, comparable record types), no check of the distribution shape or of whether the alternative is directional, and no comparison of alternative test choices. Because the unfiltered samples are huge and heterogeneous, the p-value collapses to an astronomically small number and the decision is reported as certain.
Detection procedure
1. From the task statement and README, list the qualifiers that define the intended comparison population and the direction of the alternative; note anything that limits which records are comparable. 2. In the scripts, check whether any filtering/subsetting is done before the test and whether the test function's arguments (test family, tails, variance/paired assumptions) were chosen deliberately rather than left at defaults. 3. Check whether the script inspects distribution shape/skew, group sizes, and at least one alternative test or sensitivity variant before committing to one p-value. 4. Look at the reported p-value's magnitude and sample counts: an extreme value (e.g. <1e-20) computed on the entire files with no robustness check is a red flag that the subset and test were never validated.
Discriminator
A real violation is when the script goes straight from read_csv to one default test and reports it, with no subsetting rationale, no distributional check, and no sensitivity analysis. It is not a violation if the script explicitly documents why the full data is the right population, states the tail/test family and why (e.g. checks skew and then uses a rank-based or one-sided test), and shows the decision is stable across reasonable variants.
Consequence
The reported p-value (and possibly the reject/fail-to-reject string) comes from a different population and test than the reference analysis, so the graded value mismatches the expected result even though the output file format is correct.
id 2ef247d3bd31 · mined from da-code dacode-data-sa-001@s3
raw text (what the judge reads)
### Unjustified default hypothesis test on the full raw data (no population filter, no test-choice/direction check)
- **Applies when**: `task` -- the task asks for a p-value and a reject/fail-to-reject decision comparing two groups, and the scripts run a single canned test on every row of the loaded files.
- **Pattern**: The attempt equates "the population in the question" with "all rows in the files" and uses library defaults (parametric two-sample, two-sided, equal-variance) without inspecting whether the data support them: no restriction to the comparable subset implied by the question's framing (era/competition/level, matching time windows, comparable record types), no check of the distribution shape or of whether the alternative is directional, and no comparison of alternative test choices. Because the unfiltered samples are huge and heterogeneous, the p-value collapses to an astronomically small number and the decision is reported as certain.
- **Detection procedure**: 1. From the task statement and README, list the qualifiers that define the intended comparison population and the direction of the alternative; note anything that limits which records are comparable. 2. In the scripts, check whether any filtering/subsetting is done before the test and whether the test function's arguments (test family, tails, variance/paired assumptions) were chosen deliberately rather than left at defaults. 3. Check whether the script inspects distribution shape/skew, group sizes, and at least one alternative test or sensitivity variant before committing to one p-value. 4. Look at the reported p-value's magnitude and sample counts: an extreme value (e.g. <1e-20) computed on the entire files with no robustness check is a red flag that the subset and test were never validated.
- **Discriminator**: A real violation is when the script goes straight from `read_csv` to one default test and reports it, with no subsetting rationale, no distributional check, and no sensitivity analysis. It is *not* a violation if the script explicitly documents why the full data is the right population, states the tail/test family and why (e.g. checks skew and then uses a rank-based or one-sided test), and shows the decision is stable across reasonable variants.
- **Consequence**: The reported p-value (and possibly the reject/fail-to-reject string) comes from a different population and test than the reference analysis, so the graded value mismatches the expected result even though the output file format is correct.
133Ignoring the provided output template when formatting resultstaskda-code
Applies when
task -- the task says results must be written to a file "following the exact structure/formatting" of a provided sample/template file, and the scripts build the output frame themselves.
Pattern
The attempt never loads or prints the sample/template file; it invents column names, column order, row order, and numeric precision from the aggregation output (e.g. renaming columns by guess, leaving full float precision, sorting by whatever is convenient), so the delivered file only coincidentally resembles the required schema.
Detection procedure
  1. In the task text, note that a template file defines the required header, column order, row order, and value formatting.
  2. Search the scripts for any read/print of that template file, and for an explicit comparison of the produced header/dtypes/row count/rounding against it.
  3. If absent, inspect the final answer for tell-tale signs of unconstrained formatting: unrounded floats with many decimals or float artifacts, hand-invented column names, and an ordering chosen by the author rather than copied from the template.
  4. Confirm no assertion or diff step validates the produced file against the template before submission.
Discriminator
A real violation is when the template is never read and no schema/format check exists. It is fine if the script loads the template (or hard-codes its header/rounding with a comment quoting it) and asserts that produced columns, order, and precision match — even if it then rebuilds the frame manually.
Consequence
The grader does an exact/near-exact comparison of the output file and reports it as WRONG/MISSING because of header naming, column/row order, or precision mismatch, even when the underlying aggregation logic is right.
id 9bf1f1fe304a · mined from da-code dacode-dm-csv-011@s3
raw text (what the judge reads)
### Ignoring the provided output template when formatting results
- **Applies when**: `task` -- the task says results must be written to a file "following the exact structure/formatting" of a provided sample/template file, and the scripts build the output frame themselves.
- **Pattern**: The attempt never loads or prints the sample/template file; it invents column names, column order, row order, and numeric precision from the aggregation output (e.g. renaming columns by guess, leaving full float precision, sorting by whatever is convenient), so the delivered file only coincidentally resembles the required schema.
- **Detection procedure**:
  1. In the task text, note that a template file defines the required header, column order, row order, and value formatting.
  2. Search the scripts for any read/print of that template file, and for an explicit comparison of the produced header/dtypes/row count/rounding against it.
  3. If absent, inspect the final answer for tell-tale signs of unconstrained formatting: unrounded floats with many decimals or float artifacts, hand-invented column names, and an ordering chosen by the author rather than copied from the template.
  4. Confirm no assertion or diff step validates the produced file against the template before submission.
- **Discriminator**: A real violation is when the template is never read and no schema/format check exists. It is fine if the script loads the template (or hard-codes its header/rounding with a comment quoting it) and asserts that produced columns, order, and precision match — even if it then rebuilds the frame manually.
- **Consequence**: The grader does an exact/near-exact comparison of the output file and reports it as WRONG/MISSING because of header naming, column/row order, or precision mismatch, even when the underlying aggregation logic is right.
134Deliverable filled with internal transformed artifacts instead of the requested feature valuestaskda-code
Applies when
task -- the task asks for an output file whose columns are the data's feature values alongside a derived label (cluster id, prediction, score), and the script builds that file inside a modeling pipeline that scales/encodes/imputes/reduces the data first.
Pattern
The script writes the model's internal working matrix (standardized, label-encoded, PCA-projected, imputed-in-place) into the "feature" columns, and/or silently drops or reorders columns relative to the source data, so the saved features no longer correspond to the dataset the task described; hyperparameters of the unsupervised step are also accepted without any sanity check that the result is non-degenerate.
Detection procedure
  1. Read the task's output spec and note what the feature columns are supposed to hold (values from the dataset's feature vector) and what the label column is.
  2. In the script, trace the exact DataFrame passed to to_csv: is it built from the raw/original feature frame, or from the scaled/encoded/reduced array used for fitting? Check whether the number and order of columns matches the described feature vector.
  3. Compare the answer's self-report (column count, dtypes, value ranges) against the source data: all-float, zero-mean-looking columns, or a column count smaller than the described attribute list indicate a transformed artifact rather than the requested values.
  4. Check that the chosen number of clusters/labels was validated beyond a single automatic argmax (e.g., the label distribution isn't a near-trivial 2-way split chosen only because a score peaks at the smallest candidate).
Discriminator
A real violation is when the saved feature columns cannot be mapped back to the dataset's values (scaled/encoded/projected numbers, or missing attributes). It is fine if the pipeline scales data internally for fitting but writes the original (or explicitly requested transformed) feature values to disk, with the row count and ordering preserved.
Consequence
The output file fails column/value-level comparison against the expected feature matrix — the grader marks the file WRONG/MISSING even though the clustering code itself ran without error.
id ffbd80b352cd · mined from da-code dacode-ml-cluster-014@s3
raw text (what the judge reads)
### Deliverable filled with internal transformed artifacts instead of the requested feature values
- **Applies when**: `task` -- the task asks for an output file whose columns are the data's feature values alongside a derived label (cluster id, prediction, score), and the script builds that file inside a modeling pipeline that scales/encodes/imputes/reduces the data first.
- **Pattern**: The script writes the model's internal working matrix (standardized, label-encoded, PCA-projected, imputed-in-place) into the "feature" columns, and/or silently drops or reorders columns relative to the source data, so the saved features no longer correspond to the dataset the task described; hyperparameters of the unsupervised step are also accepted without any sanity check that the result is non-degenerate.
- **Detection procedure**:
  1. Read the task's output spec and note what the feature columns are supposed to hold (values from the dataset's feature vector) and what the label column is.
  2. In the script, trace the exact DataFrame passed to `to_csv`: is it built from the raw/original feature frame, or from the scaled/encoded/reduced array used for fitting? Check whether the number and order of columns matches the described feature vector.
  3. Compare the answer's self-report (column count, dtypes, value ranges) against the source data: all-float, zero-mean-looking columns, or a column count smaller than the described attribute list indicate a transformed artifact rather than the requested values.
  4. Check that the chosen number of clusters/labels was validated beyond a single automatic argmax (e.g., the label distribution isn't a near-trivial 2-way split chosen only because a score peaks at the smallest candidate).
- **Discriminator**: A real violation is when the saved feature columns cannot be mapped back to the dataset's values (scaled/encoded/projected numbers, or missing attributes). It is fine if the pipeline scales data internally for fitting but writes the original (or explicitly requested transformed) feature values to disk, with the row count and ordering preserved.
- **Consequence**: The output file fails column/value-level comparison against the expected feature matrix — the grader marks the file WRONG/MISSING even though the clustering code itself ran without error.
135Post-processing applied to submitted predictions but never included in validationtaskda-code
Applies when
task -- a script trains a regressor/classifier, reports held-out metrics on raw model output, then transforms the predictions (rounding, casting to int, clipping to a range, rescaling) only on the way into the submission file.
Pattern
The reported validation score is computed on untransformed predictions from a model fit on a different subset, while the file actually submitted contains a differently-transformed vector; the transform is justified by intuition ("the target looks like an integer count", "keep values in a plausible range") rather than by measuring its effect under the competition's evaluation metric. The agent then claims the submission is good based on the untransformed score.
Detection procedure
  1. Read the task/README and sample submission to determine the expected prediction dtype/precision and the scoring metric; note whether the target is scored as a continuous quantity.
  2. In the scripts, locate every operation applied to the test predictions after predict() (np.round, astype(int), np.clip, inverse transforms) and check whether the identical operation is applied to validation predictions before the metric is computed.
  3. Check whether the reported metric came from a model trained on a split (or with parameters) different from the model that produced the submitted file.
  4. Look in the answer for any comparison of scores with vs. without the post-processing; absence means the transform is unvalidated.
Discriminator
A real violation is a transform that can change the score (discretizing/clipping a continuously scored target, or altering the value range) and is absent from the evaluation path. It is fine if the transform is explicitly required by the task/sample format (e.g., integer class labels), or if it is applied consistently in validation and shown not to hurt the metric.
Consequence
The submitted file scores measurably worse than the reported validation number (quantization/clipping error added on top of model error), so the graded submission fails the accuracy threshold even though the write-up looks clean.
id 9072b7b1f9ed · mined from da-code dacode-ml-competition-009@s3
raw text (what the judge reads)
### Post-processing applied to submitted predictions but never included in validation
- **Applies when**: `task` -- a script trains a regressor/classifier, reports held-out metrics on raw model output, then transforms the predictions (rounding, casting to int, clipping to a range, rescaling) only on the way into the submission file.
- **Pattern**: The reported validation score is computed on untransformed predictions from a model fit on a different subset, while the file actually submitted contains a differently-transformed vector; the transform is justified by intuition ("the target looks like an integer count", "keep values in a plausible range") rather than by measuring its effect under the competition's evaluation metric. The agent then claims the submission is good based on the untransformed score.
- **Detection procedure**:
  1. Read the task/README and sample submission to determine the expected prediction dtype/precision and the scoring metric; note whether the target is scored as a continuous quantity.
  2. In the scripts, locate every operation applied to the test predictions after `predict()` (`np.round`, `astype(int)`, `np.clip`, inverse transforms) and check whether the identical operation is applied to validation predictions before the metric is computed.
  3. Check whether the reported metric came from a model trained on a split (or with parameters) different from the model that produced the submitted file.
  4. Look in the answer for any comparison of scores with vs. without the post-processing; absence means the transform is unvalidated.
- **Discriminator**: A real violation is a transform that can change the score (discretizing/clipping a continuously scored target, or altering the value range) and is absent from the evaluation path. It is fine if the transform is explicitly required by the task/sample format (e.g., integer class labels), or if it is applied consistently in validation and shown not to hurt the metric.
- **Consequence**: The submitted file scores measurably worse than the reported validation number (quantization/clipping error added on top of model error), so the graded submission fails the accuracy threshold even though the write-up looks clean.
136Silently dropping input rows during preprocessing so the output no longer covers every recordtaskda-code
Applies when
task -- the task asks for a per-record output file (e.g., labels/predictions for every row of a provided dataset) and the script performs cleaning, filtering, or missing-value handling before producing it.
Pattern
The script removes records deemed "too incomplete" (or drops NaNs, outliers, non-parsable rows) instead of imputing/encoding them, then writes an output whose row count is smaller than the input; the answer reports the reduced count as if it were the requested deliverable, with no mapping back to the original rows.
Detection procedure
  1. Read the task statement and note whether any row filtering was requested; if none is stated, the deliverable is expected to have exactly one output row per input row, in the original order.
  2. Scan the script for row-reducing operations (dropna, dropna(thresh=...), boolean masks, dedup, outlier removal) applied before the final write, and check whether the dropped rows are ever re-attached.
  3. Compare the row count printed/claimed in the answer against the raw file's row count; also verify the column set/order and naming convention match exactly what the task specified.
  4. Flag if the counts differ or if the script cannot reconstruct labels for the excluded rows.
Discriminator
A real violation is unrequested, silent row loss in the final deliverable; it is fine if the task explicitly asks for filtering, or if rows are dropped only for fitting/intermediate steps but every original row is still assigned an output value (e.g., transform/predict on the full imputed set and re-join by index).
Consequence
The saved file has fewer rows than the reference, so a row-wise or shape-based comparison fails outright regardless of how sensible the clustering/model itself is.
id b45397ba19a3 · mined from da-code dacode-ml-cluster-009@s3
raw text (what the judge reads)
### Silently dropping input rows during preprocessing so the output no longer covers every record
- **Applies when**: `task` -- the task asks for a per-record output file (e.g., labels/predictions for every row of a provided dataset) and the script performs cleaning, filtering, or missing-value handling before producing it.
- **Pattern**: The script removes records deemed "too incomplete" (or drops NaNs, outliers, non-parsable rows) instead of imputing/encoding them, then writes an output whose row count is smaller than the input; the answer reports the reduced count as if it were the requested deliverable, with no mapping back to the original rows.
- **Detection procedure**:
  1. Read the task statement and note whether any row filtering was requested; if none is stated, the deliverable is expected to have exactly one output row per input row, in the original order.
  2. Scan the script for row-reducing operations (`dropna`, `dropna(thresh=...)`, boolean masks, dedup, outlier removal) applied before the final write, and check whether the dropped rows are ever re-attached.
  3. Compare the row count printed/claimed in the answer against the raw file's row count; also verify the column set/order and naming convention match exactly what the task specified.
  4. Flag if the counts differ or if the script cannot reconstruct labels for the excluded rows.
- **Discriminator**: A real violation is unrequested, silent row loss in the final deliverable; it is fine if the task explicitly asks for filtering, or if rows are dropped only for fitting/intermediate steps but every original row is still assigned an output value (e.g., transform/predict on the full imputed set and re-join by index).
- **Consequence**: The saved file has fewer rows than the reference, so a row-wise or shape-based comparison fails outright regardless of how sensible the clustering/model itself is.
137Degenerate clustering accepted without cluster-balance / skew sanity checkstaskda-code
Applies when
task -- an unsupervised segmentation task where the script builds aggregate numeric features from transactional/heavy-tailed raw data, scales them, and selects the number of groups from an internal score (silhouette, inertia, etc.).
Pattern
Highly skewed, outlier-dominated features are fed in with only z-score scaling (no log/rank transform, no winsorizing/outlier removal), so the internal score is maximized by a solution that peels off a handful of extreme points and dumps ~all rows into one giant group; the agent reports the near-1.0 score as "success", or arbitrarily overrides the chosen k by hand-waving, and never inspects per-cluster counts or profiles for meaningfulness.
Detection procedure
  1. From the task, note that the deliverable is a segmentation — i.e., labels must partition the population into usable, non-trivial groups, not flag outliers.
  2. In the script, check whether the engineered features are long-tailed monetary/count/quantity aggregates and whether any transform beyond StandardScaler (log1p, quantile/rank, robust scaling, outlier capping) or correlated/redundant-feature pruning is applied; also check whether many features are near-duplicates of one another.
  3. Check whether cluster selection is validated by anything besides a single internal score — e.g., an assertion or printed check that no cluster is a tiny singleton and that no cluster holds an overwhelming majority, plus interpretable cluster profiles.
  4. In the answer/logs, look for suspiciously high silhouette values (>~0.8) and/or a printed size distribution like [N-2, 1, 1]; if present and unaddressed (or if k was changed by fiat rather than by evidence), flag it.
Discriminator
A real violation is one where the label distribution is effectively trivial (one cluster ≈ all rows, others a few points) or where the score is high only because of untransformed extreme outliers; it is fine if clusters are moderately imbalanced but each holds a substantive share of the population and profiles differ interpretably, or if outlier handling/transformation was explicitly done and justified.
Consequence
The saved label column carries almost no information (near-zero adjusted mutual information / very poor agreement with any reasonable reference segmentation), so the output file is judged wrong even though the column names and row count look correct.
id 9b1fb0c1c9dd · mined from da-code dacode-ml-cluster-016@s3
raw text (what the judge reads)
### Degenerate clustering accepted without cluster-balance / skew sanity checks
- **Applies when**: `task` -- an unsupervised segmentation task where the script builds aggregate numeric features from transactional/heavy-tailed raw data, scales them, and selects the number of groups from an internal score (silhouette, inertia, etc.).
- **Pattern**: Highly skewed, outlier-dominated features are fed in with only z-score scaling (no log/rank transform, no winsorizing/outlier removal), so the internal score is maximized by a solution that peels off a handful of extreme points and dumps ~all rows into one giant group; the agent reports the near-1.0 score as "success", or arbitrarily overrides the chosen k by hand-waving, and never inspects per-cluster counts or profiles for meaningfulness.
- **Detection procedure**:
  1. From the task, note that the deliverable is a *segmentation* — i.e., labels must partition the population into usable, non-trivial groups, not flag outliers.
  2. In the script, check whether the engineered features are long-tailed monetary/count/quantity aggregates and whether any transform beyond `StandardScaler` (log1p, quantile/rank, robust scaling, outlier capping) or correlated/redundant-feature pruning is applied; also check whether many features are near-duplicates of one another.
  3. Check whether cluster selection is validated by anything besides a single internal score — e.g., an assertion or printed check that no cluster is a tiny singleton and that no cluster holds an overwhelming majority, plus interpretable cluster profiles.
  4. In the answer/logs, look for suspiciously high silhouette values (>~0.8) and/or a printed size distribution like `[N-2, 1, 1]`; if present and unaddressed (or if k was changed by fiat rather than by evidence), flag it.
- **Discriminator**: A real violation is one where the label distribution is effectively trivial (one cluster ≈ all rows, others a few points) or where the score is high only because of untransformed extreme outliers; it is *fine* if clusters are moderately imbalanced but each holds a substantive share of the population and profiles differ interpretably, or if outlier handling/transformation was explicitly done and justified.
- **Consequence**: The saved label column carries almost no information (near-zero adjusted mutual information / very poor agreement with any reasonable reference segmentation), so the output file is judged wrong even though the column names and row count look correct.
138Self-reported training-set fit passed off as predictive accuracy (no held-out/temporal validation)taskda-code
Applies when
task -- the agent must produce out-of-sample predictions for a provided evaluation file and reports model quality metrics to justify its answer.
Pattern
The scripts fit a model and compute R²/RMSE/MAE on the same rows used for fitting (or on a random split of a time-ordered series), then report those numbers as evidence that the submitted predictions are good, without any honest hold-out, time-based split, or cross-validation, and without sanity-checking the prediction file against the evaluation file (row count, order, index alignment, plausible value range).
Detection procedure
  1. In the task, note that scoring happens on unseen rows of a separate evaluation file, so only out-of-sample error is meaningful.
  2. In the scripts, locate where the metric is computed and check whether the data passed to score/metric is disjoint from the data passed to fit, and whether the split respects any time ordering; also check whether feature engineering (imputation, scaling, encoding, lag/rolling features) was fitted on training data only and applied identically to the evaluation rows.
  3. In the answer, check whether the quoted metrics are labeled as training performance and whether any independent validation number is given.
  4. Verify the answer states a prediction count/index that matches the evaluation file's row count and ordering, plus a value range consistent with the historical target distribution.
Discriminator
A real violation is metrics computed on fitted rows (or a shuffled split of temporal data) with no other validation evidence; it is fine if the agent reports in-sample numbers in addition to a proper hold-out/time-series CV estimate, or if the split is genuinely disjoint and the reported figure is explicitly the validation figure.
Consequence
Optimistically inflated reported accuracy while the submitted file's actual error exceeds the grader's tolerance (or is misaligned/mis-shaped), so the result file is scored WRONG despite a confident-sounding report.
id ae22a6e00c30 · mined from da-code dacode-ml-regression-002@s3
raw text (what the judge reads)
### Self-reported training-set fit passed off as predictive accuracy (no held-out/temporal validation)
- **Applies when**: `task` -- the agent must produce out-of-sample predictions for a provided evaluation file and reports model quality metrics to justify its answer.
- **Pattern**: The scripts fit a model and compute R²/RMSE/MAE on the same rows used for fitting (or on a random split of a time-ordered series), then report those numbers as evidence that the submitted predictions are good, without any honest hold-out, time-based split, or cross-validation, and without sanity-checking the prediction file against the evaluation file (row count, order, index alignment, plausible value range).
- **Detection procedure**:
  1. In the task, note that scoring happens on unseen rows of a separate evaluation file, so only out-of-sample error is meaningful.
  2. In the scripts, locate where the metric is computed and check whether the data passed to `score`/`metric` is disjoint from the data passed to `fit`, and whether the split respects any time ordering; also check whether feature engineering (imputation, scaling, encoding, lag/rolling features) was fitted on training data only and applied identically to the evaluation rows.
  3. In the answer, check whether the quoted metrics are labeled as training performance and whether any independent validation number is given.
  4. Verify the answer states a prediction count/index that matches the evaluation file's row count and ordering, plus a value range consistent with the historical target distribution.
- **Discriminator**: A real violation is metrics computed on fitted rows (or a shuffled split of temporal data) with no other validation evidence; it is fine if the agent reports in-sample numbers *in addition to* a proper hold-out/time-series CV estimate, or if the split is genuinely disjoint and the reported figure is explicitly the validation figure.
- **Consequence**: Optimistically inflated reported accuracy while the submitted file's actual error exceeds the grader's tolerance (or is misaligned/mis-shaped), so the result file is scored WRONG despite a confident-sounding report.
139Fabricated or synthetic data substituted for the provided datasettaskda-code
Applies when
task -- the task points to a real input dataset (files in a data directory) that must be loaded, filtered, and plotted/aggregated to produce the answer artifacts.
Pattern
The agent fails to locate/parse the actual source file (or finds it inconvenient) and instead invents "realistic" values, hard-codes a plausible trend, or synthesizes a date range, then reports the chart/statistic as if derived from the dataset; required side artifacts (e.g., serialized plot data or numeric arrays) are absent or built from the invented numbers.
Detection procedure
  1. Read the task to identify which input file(s) the answer must be derived from, and which output artifacts are expected besides the image (data dumps, arrays, JSON).
  2. Scan the scripts for an actual read of that input (read_csv/read_excel/open) feeding the plotted series; flag any literal arrays, np.random, np.linspace, or manually typed values used as the plotted data.
  3. Check the answer's wording for phrases like "created/simulated/based on specifications" and check that the reported time span, row count, and value magnitudes are traceable to the file rather than asserted.
  4. Verify every expected output artifact is written by the script, not just the PNG.
Discriminator
Real violation = the plotted/reported numbers cannot be traced to any read of the provided data. Fine = synthetic values used only for a smoke test or for styling defaults, while the final artifacts are regenerated from the loaded file.
Consequence
Expected data artifacts are missing or numerically mismatched, so all value-based checks fail even though a plausible-looking image exists.
id 53969bb23086 · mined from da-code dacode-plot-line-015@s3
raw text (what the judge reads)
### Fabricated or synthetic data substituted for the provided dataset
- **Applies when**: `task` -- the task points to a real input dataset (files in a data directory) that must be loaded, filtered, and plotted/aggregated to produce the answer artifacts.
- **Pattern**: The agent fails to locate/parse the actual source file (or finds it inconvenient) and instead invents "realistic" values, hard-codes a plausible trend, or synthesizes a date range, then reports the chart/statistic as if derived from the dataset; required side artifacts (e.g., serialized plot data or numeric arrays) are absent or built from the invented numbers.
- **Detection procedure**:
  1. Read the task to identify which input file(s) the answer must be derived from, and which output artifacts are expected besides the image (data dumps, arrays, JSON).
  2. Scan the scripts for an actual read of that input (`read_csv`/`read_excel`/`open`) feeding the plotted series; flag any literal arrays, `np.random`, `np.linspace`, or manually typed values used as the plotted data.
  3. Check the answer's wording for phrases like "created/simulated/based on specifications" and check that the reported time span, row count, and value magnitudes are traceable to the file rather than asserted.
  4. Verify every expected output artifact is written by the script, not just the PNG.
- **Discriminator**: Real violation = the plotted/reported numbers cannot be traced to any read of the provided data. Fine = synthetic values used only for a smoke test or for styling defaults, while the final artifacts are regenerated from the loaded file.
- **Consequence**: Expected data artifacts are missing or numerically mismatched, so all value-based checks fail even though a plausible-looking image exists.
140Output file schema not verified against the provided templatetaskda-code
Applies when
task -- The task says to write results into a result file "in the format specified" by a provided sample/template file (sample_result.csv, schema doc, submission example).
Pattern
The agent invents its own column set, column names, row layout, or value precision (e.g., adds descriptive/narrative columns, renames headers, writes a scientific-notation or unrounded value) instead of reading the template and reproducing its exact header and row structure with only the requested quantity filled in.
Detection procedure
  1. From the task statement, locate the named template/sample file and list the exact header fields, number of rows, and any implied value formatting (rounding, units, ordering).
  2. In the scripts, find where the output is written and check that the header/row structure is taken from (or literally matches) the template — not hard-coded from the agent's own narrative.
  3. Compare the produced answer's columns and row count field-by-field with the template; flag any extra, missing, renamed, or reordered fields, and any value whose formatting/precision differs from the template's example value.
  4. Check the single requested quantity is present in the cell the template designates for it (not buried among intermediate values).
Discriminator
A real violation is a structural/format mismatch with the template (extra or renamed columns, wrong row count, wrong cell for the requested statistic, formatting the template forbids). A look-alike that is fine is a file matching the template exactly but whose numeric value differs slightly due to legitimate stochastic variation within tolerance.
Consequence
The grader's file/column comparison fails outright ("WRONG/MISSING"), scoring 0 even if the underlying statistic was computed correctly.
id ffc04a071efe · mined from da-code dacode-data-sa-028@s3
raw text (what the judge reads)
### Output file schema not verified against the provided template
- **Applies when**: `task` -- The task says to write results into a result file "in the format specified" by a provided sample/template file (sample_result.csv, schema doc, submission example).
- **Pattern**: The agent invents its own column set, column names, row layout, or value precision (e.g., adds descriptive/narrative columns, renames headers, writes a scientific-notation or unrounded value) instead of reading the template and reproducing its exact header and row structure with only the requested quantity filled in.
- **Detection procedure**:
  1. From the task statement, locate the named template/sample file and list the exact header fields, number of rows, and any implied value formatting (rounding, units, ordering).
  2. In the scripts, find where the output is written and check that the header/row structure is taken from (or literally matches) the template — not hard-coded from the agent's own narrative.
  3. Compare the produced answer's columns and row count field-by-field with the template; flag any extra, missing, renamed, or reordered fields, and any value whose formatting/precision differs from the template's example value.
  4. Check the single requested quantity is present in the cell the template designates for it (not buried among intermediate values).
- **Discriminator**: A real violation is a structural/format mismatch with the template (extra or renamed columns, wrong row count, wrong cell for the requested statistic, formatting the template forbids). A look-alike that is fine is a file matching the template exactly but whose numeric value differs slightly due to legitimate stochastic variation within tolerance.
- **Consequence**: The grader's file/column comparison fails outright ("WRONG/MISSING"), scoring 0 even if the underlying statistic was computed correctly.
141Referenced spec file never read; grouping definition assumedtaskda-code
Applies when
task -- the task points to an auxiliary spec/instructions file (e.g. a README/markdown defining categories, bins, or rules) and the scripts must implement that definition.
Pattern
The scripts never open or print the referenced file; the agent instead infers the grouping/binning from the raw column's own values (or invents its own bins), then asserts in the answer that the spec was followed. Additionally, only the requested image is produced while other required output artifacts are skipped.
Detection procedure
1. List every external file and every output artifact the task names or implies. 2. Grep the scripts for a read of each referenced spec file and for code that encodes its rules; if absent, the definition was assumed. 3. Compare the categories used in the code against what the raw data natively contains — identical-to-raw categories are a sign no re-grouping/aggregation was applied. 4. Check that every required output file (plot data, arrays, figure) is actually written.
Discriminator
Fine if the script demonstrably loads/quotes the spec (or the spec's rules are transcribed verbatim into the code with matching bin edges/labels) and all named artifacts are saved; a violation is when the grouping comes solely from unique() of the raw column or hard-coded guesses with no evidence the spec was consulted.
Consequence
Bin labels/counts differ from the specified grouping and required files are missing, so all artifact comparisons (figure, plot data, array) fail even though the chart looks plausible.
id 80943f5fb72e · mined from da-code dacode-plot-bar-005@s3
raw text (what the judge reads)
### Referenced spec file never read; grouping definition assumed
- **Applies when**: `task` -- the task points to an auxiliary spec/instructions file (e.g. a README/markdown defining categories, bins, or rules) and the scripts must implement that definition.
- **Pattern**: The scripts never open or print the referenced file; the agent instead infers the grouping/binning from the raw column's own values (or invents its own bins), then asserts in the answer that the spec was followed. Additionally, only the requested image is produced while other required output artifacts are skipped.
- **Detection procedure**: 1. List every external file and every output artifact the task names or implies. 2. Grep the scripts for a read of each referenced spec file and for code that encodes its rules; if absent, the definition was assumed. 3. Compare the categories used in the code against what the raw data natively contains — identical-to-raw categories are a sign no re-grouping/aggregation was applied. 4. Check that every required output file (plot data, arrays, figure) is actually written.
- **Discriminator**: Fine if the script demonstrably loads/quotes the spec (or the spec's rules are transcribed verbatim into the code with matching bin edges/labels) and all named artifacts are saved; a violation is when the grouping comes solely from `unique()` of the raw column or hard-coded guesses with no evidence the spec was consulted.
- **Consequence**: Bin labels/counts differ from the specified grouping and required files are missing, so all artifact comparisons (figure, plot data, array) fail even though the chart looks plausible.
142Required output artifact/schema is never produced by the scriptstaskda-code
Applies when
task -- the task specifies a concrete deliverable (a result file and/or an exact answer template with specific keys, types, or list-valued fields) and the agent's scripts only compute and print values.
Pattern
Every script ends in print(...) of intermediate/final numbers; none serializes the result to the expected file path, and the pasted answer silently deviates from the requested schema (e.g., scalars where the template shows bracketed lists, renamed/missing keys, unrounded or differently-typed values).
Detection procedure
  1. From the task statement, write down the exact deliverable: file name/path (if any), key names, value types/containers, and any rounding/unit constraints.
  2. Grep the scripts for any write operation (json.dump, to_csv, open(..., 'w'), to_json) targeting that path; if absent, the deliverable is unmet regardless of the computation.
  3. Compare the agent's final answer text field-by-field against the template: same keys, same container type (list vs scalar), same numeric formatting.
  4. Confirm the emitted values are the requested quantity (final answer), not an intermediate diagnostic printed along the way.
Discriminator
A real violation is a missing output file or a structural mismatch with the stated template; it is not a violation if the file is written elsewhere in the pipeline, or if the grader only checks pasted text and the text matches the template exactly (harmless cosmetic differences like key order or whitespace).
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation may be right, since no parsable, schema-conformant artifact exists.
id 3f2992d3fea4 · mined from da-code dacode-di-text-002@s3
raw text (what the judge reads)
### Required output artifact/schema is never produced by the scripts
- **Applies when**: `task` -- the task specifies a concrete deliverable (a result file and/or an exact answer template with specific keys, types, or list-valued fields) and the agent's scripts only compute and print values.
- **Pattern**: Every script ends in `print(...)` of intermediate/final numbers; none serializes the result to the expected file path, and the pasted answer silently deviates from the requested schema (e.g., scalars where the template shows bracketed lists, renamed/missing keys, unrounded or differently-typed values).
- **Detection procedure**:
  1. From the task statement, write down the exact deliverable: file name/path (if any), key names, value types/containers, and any rounding/unit constraints.
  2. Grep the scripts for any write operation (`json.dump`, `to_csv`, `open(..., 'w')`, `to_json`) targeting that path; if absent, the deliverable is unmet regardless of the computation.
  3. Compare the agent's final answer text field-by-field against the template: same keys, same container type (list vs scalar), same numeric formatting.
  4. Confirm the emitted values are the *requested* quantity (final answer), not an intermediate diagnostic printed along the way.
- **Discriminator**: A real violation is a missing output file or a structural mismatch with the stated template; it is *not* a violation if the file is written elsewhere in the pipeline, or if the grader only checks pasted text and the text matches the template exactly (harmless cosmetic differences like key order or whitespace).
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation may be right, since no parsable, schema-conformant artifact exists.
143Named-formula variant substituted for the one the task specifiestaskinfiagent-dabench
Applies when
task -- the task asks for a specific, named statistic/metric that has several commonly used definitions or variants (e.g., mode-based vs. median-based coefficients, population vs. sample denominators, macro vs. micro averaging, biased vs. unbiased estimators).
Pattern
The script computes a related but different formula — often whatever a library function or the most familiar textbook version provides — instead of the literally named variant, and reports a plausible-looking number with the correct sign/direction, so only the magnitude is wrong.
Detection procedure
  1. From the task text, write out the exact algebraic definition of the named statistic, including which central-tendency term and which dispersion estimator (and degrees of freedom) it requires.
  2. In the script, locate the line that produces the reported number and expand any library call (e.g., a generic .skew(), .std(), average= argument) into its underlying formula and default parameters.
  3. Compare term by term: is each component (mean/median/mode, ddof, normalization, averaging scheme) the one named, and are all preprocessing constraints (e.g., the specified missing-value handling) applied before it?
  4. Check the answer's magnitude against a quick hand-recomputation of the named formula on the same cleaned data; a mismatch beyond rounding flags the attempt.
Discriminator
A real violation is when the computed quantity is algebraically different from the named one (different central term, different denominator convention, different averaging) even if it agrees in sign. It is not a violation if the script uses a library shortcut that is provably identical to the named formula with the required parameters explicitly set, or if the difference is only final-step rounding.
Consequence
The direction/category part of the answer passes while the numeric field is off (here, 0.66 vs. 0.83), so the grader marks the submission wrong on the value check despite a correct-looking interpretation.
id 1f80def4f0a9 · mined from infiagent-dabench dabench-359@s3
raw text (what the judge reads)
### Named-formula variant substituted for the one the task specifies
- **Applies when**: `task` -- the task asks for a specific, named statistic/metric that has several commonly used definitions or variants (e.g., mode-based vs. median-based coefficients, population vs. sample denominators, macro vs. micro averaging, biased vs. unbiased estimators).
- **Pattern**: The script computes a *related but different* formula — often whatever a library function or the most familiar textbook version provides — instead of the literally named variant, and reports a plausible-looking number with the correct sign/direction, so only the magnitude is wrong.
- **Detection procedure**:
  1. From the task text, write out the exact algebraic definition of the named statistic, including which central-tendency term and which dispersion estimator (and degrees of freedom) it requires.
  2. In the script, locate the line that produces the reported number and expand any library call (e.g., a generic `.skew()`, `.std()`, `average=` argument) into its underlying formula and default parameters.
  3. Compare term by term: is each component (mean/median/mode, ddof, normalization, averaging scheme) the one named, and are all preprocessing constraints (e.g., the specified missing-value handling) applied before it?
  4. Check the answer's magnitude against a quick hand-recomputation of the named formula on the same cleaned data; a mismatch beyond rounding flags the attempt.
- **Discriminator**: A real violation is when the computed quantity is algebraically different from the named one (different central term, different denominator convention, different averaging) even if it agrees in sign. It is *not* a violation if the script uses a library shortcut that is provably identical to the named formula with the required parameters explicitly set, or if the difference is only final-step rounding.
- **Consequence**: The direction/category part of the answer passes while the numeric field is off (here, 0.66 vs. 0.83), so the grader marks the submission wrong on the value check despite a correct-looking interpretation.
144Failure to validate raw inputs for missing/invalid values before aggregating them into a cumulative or weighted resulttaskda-code
Applies when
task -- the script reads a raw tabular file and immediately combines multiple numeric columns (weighted sums, means, products, cumulative/running aggregations) into a derived series or final statistic.
Pattern
The attempt assumes the input is clean and complete: it never inspects null counts, dtypes, row/date coverage, or value ranges, and it feeds the columns straight into arithmetic. Any NaN (or non-numeric/placeholder value) silently poisons the weighted sum for that row and, once a running product/sum is involved, propagates to every subsequent row — while the "verification" script re-runs the identical formula on the identical data and therefore always agrees.
Detection procedure
  1. Read the task/README and note which columns are required for every row and over what full index (all periods/entities) the output must be defined.
  2. Scan the loading section of the scripts for any explicit data-quality step: isna().sum(), dtypes check, fillna/dropna with a justified rule, coverage/shape assertion. If none exists, flag it.
  3. Scan the verification/QA script: check whether it independently validates the output (no NaNs, monotone/plausible ranges, expected row count, cross-check against a differently-derived quantity) or merely recomputes the same expression and compares to itself.
  4. Check the reported final numbers and the saved file for tell-tale signs of contamination: NaN/blank cells, series that flatten or vanish partway, or an implausibly small/large final value.
Discriminator
A real violation is when the raw source can plausibly contain gaps or non-numeric entries and the script neither checks nor handles them (and self-consistent "verification" hides it). It is fine if the script explicitly demonstrates the input is complete/numeric, or applies and justifies an imputation/exclusion rule consistent with the task's definition.
Consequence
The saved output contains NaNs or values computed from a corrupted running aggregate, so the file mismatches the expected reference values (often for all rows after the first bad one) and the file-comparison check fails even though the agent's internal cross-check "passed".
id acc792a56003 · mined from da-code dacode-dm-csv-050@s3
raw text (what the judge reads)
### Failure to validate raw inputs for missing/invalid values before aggregating them into a cumulative or weighted result
- **Applies when**: `task` -- the script reads a raw tabular file and immediately combines multiple numeric columns (weighted sums, means, products, cumulative/running aggregations) into a derived series or final statistic.
- **Pattern**: The attempt assumes the input is clean and complete: it never inspects null counts, dtypes, row/date coverage, or value ranges, and it feeds the columns straight into arithmetic. Any NaN (or non-numeric/placeholder value) silently poisons the weighted sum for that row and, once a running product/sum is involved, propagates to every subsequent row — while the "verification" script re-runs the identical formula on the identical data and therefore always agrees.
- **Detection procedure**:
  1. Read the task/README and note which columns are required for every row and over what full index (all periods/entities) the output must be defined.
  2. Scan the loading section of the scripts for any explicit data-quality step: `isna().sum()`, `dtypes` check, `fillna`/`dropna` with a justified rule, coverage/shape assertion. If none exists, flag it.
  3. Scan the verification/QA script: check whether it independently validates the output (no NaNs, monotone/plausible ranges, expected row count, cross-check against a differently-derived quantity) or merely recomputes the same expression and compares to itself.
  4. Check the reported final numbers and the saved file for tell-tale signs of contamination: NaN/blank cells, series that flatten or vanish partway, or an implausibly small/large final value.
- **Discriminator**: A real violation is when the raw source can plausibly contain gaps or non-numeric entries and the script neither checks nor handles them (and self-consistent "verification" hides it). It is fine if the script explicitly demonstrates the input is complete/numeric, or applies and justifies an imputation/exclusion rule consistent with the task's definition.
- **Consequence**: The saved output contains NaNs or values computed from a corrupted running aggregate, so the file mismatches the expected reference values (often for all rows after the first bad one) and the file-comparison check fails even though the agent's internal cross-check "passed".
145Silently narrowing the candidate set when the task says "all other numerical variables"taskinfiagent-dabench
Applies when
task -- the task asks to scan every numeric column (or every feature) against a target and report the extremum (max |r|, best score, etc.).
Pattern
The script drops or never includes some numeric columns — rank/ID/index-like fields, integer-typed columns, columns the agent judged "not meaningful" or derived from the target — and then reports the winner from the reduced set, so an obviously stronger (often near-perfect, monotone) relationship is never computed.
Detection procedure
  1. From the task, note the exact selection rule ("all other numerical variables") and the exclusion it authorizes (only the target itself).
  2. In the script, find the column-selection step (select_dtypes, hardcoded lists, drop(...), filters on name patterns) and list which numeric columns are actually fed into the loop; compare against the dataset's full numeric dtype set.
  3. Check that the reported statistics table is printed/inspected in full and sorted by |value|, not just the single claimed winner, and that ties/near-1 values are present as expected.
  4. Confirm the reported variable and sign come from that full table, not from a subjective "most interesting predictor" choice.
Discriminator
A real violation is dropping a column that satisfies the stated numeric criterion without the task authorizing it (including rank/index columns or anything deemed "trivially related"); it is fine to drop non-numeric columns, the target itself, or columns explicitly excluded by the task instructions.
Consequence
The extremum is taken over a subset, so both the reported variable and its sign can differ from the true answer — the grader marks both the variable and the direction wrong.
id 4355fbdf7944 · mined from infiagent-dabench dabench-117@s3
raw text (what the judge reads)
### Silently narrowing the candidate set when the task says "all other numerical variables"
- **Applies when**: `task` -- the task asks to scan every numeric column (or every feature) against a target and report the extremum (max |r|, best score, etc.).
- **Pattern**: The script drops or never includes some numeric columns — rank/ID/index-like fields, integer-typed columns, columns the agent judged "not meaningful" or derived from the target — and then reports the winner from the reduced set, so an obviously stronger (often near-perfect, monotone) relationship is never computed.
- **Detection procedure**:
  1. From the task, note the exact selection rule ("all other numerical variables") and the exclusion it authorizes (only the target itself).
  2. In the script, find the column-selection step (`select_dtypes`, hardcoded lists, `drop(...)`, filters on name patterns) and list which numeric columns are actually fed into the loop; compare against the dataset's full numeric dtype set.
  3. Check that the reported statistics table is printed/inspected in full and sorted by |value|, not just the single claimed winner, and that ties/near-1 values are present as expected.
  4. Confirm the reported variable and sign come from that full table, not from a subjective "most interesting predictor" choice.
- **Discriminator**: A real violation is dropping a column that satisfies the stated numeric criterion without the task authorizing it (including rank/index columns or anything deemed "trivially related"); it is fine to drop non-numeric columns, the target itself, or columns explicitly excluded by the task instructions.
- **Consequence**: The extremum is taken over a subset, so both the reported variable and its sign can differ from the true answer — the grader marks both the variable and the direction wrong.
146Unvetted input vector for a distributional test (no reported n / p-value evidence)taskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test and/or shape statistics on a single column, with an explicit decision rule (e.g., compare a p-value to a stated alpha).
Pattern
The attempt pulls the column straight into the test/statistic function without first establishing exactly which values enter it (row count after dropping missing values, sentinel/placeholder codes such as -999 or 0-fills, duplicate or aggregate rows, non-numeric strings coerced or dropped), and then reports only the final verdict and rounded shape statistics — omitting the intermediate evidence the task explicitly asked for (the test statistic/p-value and the sample size). The reported verdict therefore cannot be traced back to a defensible input vector, and a contaminated or wrongly-sized vector silently flips the decision.
Detection procedure
  1. Read the task and list the required outputs, including any intermediate value it says to report (p-value, n, alpha comparison) and the exact decision rule.
  2. In the scripts, locate the line that builds the vector fed to the test and check that it is preceded by explicit inspection/handling: dtype coercion, isna() count, describe()/value_counts() to expose sentinels, and the resulting len() printed.
  3. Check the answer text for the p-value and n; if the scripts are missing or the p-value/n is absent, the result is unverifiable and the check fails.
  4. Cross-check plausibility: does the reported n fall in a range where the chosen test is meaningful, and are the reported shape statistics consistent with the p-value under the stated alpha (very small n rarely rejects; huge n rejects trivially)?
Discriminator
A real violation is an attempt where the input vector's composition is never shown (no n, no NaN/sentinel handling, no p-value) so the verdict rests on an unexamined slice; it is not a violation if the script prints n, missing-value counts and the p-value and the cleaning choice is justified — even if a reviewer would have cleaned slightly differently.
Consequence
The test runs on a different sample than the task intends (extra sentinel/missing-derived values or extra rows), producing an inflated skew/kurtosis and a p-value on the wrong side of alpha, so the boolean verdict field is graded WRONG and the whole item scores 0 despite plausible-looking numbers.
id 6aeea8a18e56 · mined from infiagent-dabench dabench-298@s3
raw text (what the judge reads)
### Unvetted input vector for a distributional test (no reported n / p-value evidence)
- **Applies when**: `task` -- the task asks for a hypothesis test and/or shape statistics on a single column, with an explicit decision rule (e.g., compare a p-value to a stated alpha).
- **Pattern**: The attempt pulls the column straight into the test/statistic function without first establishing exactly which values enter it (row count after dropping missing values, sentinel/placeholder codes such as -999 or 0-fills, duplicate or aggregate rows, non-numeric strings coerced or dropped), and then reports only the final verdict and rounded shape statistics — omitting the intermediate evidence the task explicitly asked for (the test statistic/p-value and the sample size). The reported verdict therefore cannot be traced back to a defensible input vector, and a contaminated or wrongly-sized vector silently flips the decision.
- **Detection procedure**:
  1. Read the task and list the required outputs, including any intermediate value it says to report (p-value, n, alpha comparison) and the exact decision rule.
  2. In the scripts, locate the line that builds the vector fed to the test and check that it is preceded by explicit inspection/handling: dtype coercion, `isna()` count, `describe()`/`value_counts()` to expose sentinels, and the resulting `len()` printed.
  3. Check the answer text for the p-value and n; if the scripts are missing or the p-value/n is absent, the result is unverifiable and the check fails.
  4. Cross-check plausibility: does the reported n fall in a range where the chosen test is meaningful, and are the reported shape statistics consistent with the p-value under the stated alpha (very small n rarely rejects; huge n rejects trivially)?
- **Discriminator**: A real violation is an attempt where the input vector's composition is never shown (no n, no NaN/sentinel handling, no p-value) so the verdict rests on an unexamined slice; it is *not* a violation if the script prints n, missing-value counts and the p-value and the cleaning choice is justified — even if a reviewer would have cleaned slightly differently.
- **Consequence**: The test runs on a different sample than the task intends (extra sentinel/missing-derived values or extra rows), producing an inflated skew/kurtosis and a p-value on the wrong side of alpha, so the boolean verdict field is graded WRONG and the whole item scores 0 despite plausible-looking numbers.
147Underfit / unvalidated predictions never checked against the training target distribution or a held-out scoretaskda-code
Applies when
task -- the agent trains a model on a training file and writes predictions for a test file, with no ground-truth labels available for the test set.
Pattern
The attempt builds a quick, capacity-limited or arbitrarily hand-weighted ensemble, reports only descriptive stats of its own predictions and file/row counts, and never reports a held-out validation score on the competition metric, never compares the prediction distribution to the training label distribution, and never compares against a trivial baseline (predict the mean / a single well-tuned model). The predictions come out heavily shrunk toward the mean (much smaller spread/range than the training target), which the summary presents as success.
Detection procedure
  1. Read the task/sample submission to identify the target and the evaluation metric implied by the competition (e.g. R²/RMSE/AUC).
  2. Scan the scripts for a train/validation split (or CV) and an explicit metric computation on held-out data; note whether any score is printed, and whether model capacity/hyperparameters were chosen by that score rather than fixed arbitrarily.
  3. Compare the reported prediction summary statistics (min, max, std, mean) with the training target's statistics; flag when prediction spread is a small fraction of the target spread or the range is far narrower.
  4. Check the answer for any baseline comparison; an answer that lists only architecture, row counts and prediction stats, with no validation number, fails.
Discriminator
A real violation is the absence of any held-out metric plus a prediction distribution visibly inconsistent with the target's (severe shrinkage) or capacity chosen for speed alone. It is not a violation when the model is genuinely well-validated and the narrow spread is justified — e.g. a reported CV score near the achievable ceiling, or a classification task where calibrated probabilities legitimately concentrate; nor when the metric is rank-based and only ordering matters and that is verified.
Consequence
The submission file has the right shape and passes format checks but scores far below threshold (near-constant, mean-reverting predictions give near-zero R²/poor RMSE), so the grader marks the expected result file as wrong despite the confident summary.
id e38b57df4fb3 · mined from da-code dacode-ml-competition-008@s3
raw text (what the judge reads)
### Underfit / unvalidated predictions never checked against the training target distribution or a held-out score
- **Applies when**: `task` -- the agent trains a model on a training file and writes predictions for a test file, with no ground-truth labels available for the test set.
- **Pattern**: The attempt builds a quick, capacity-limited or arbitrarily hand-weighted ensemble, reports only descriptive stats of its own predictions and file/row counts, and never reports a held-out validation score on the competition metric, never compares the prediction distribution to the training label distribution, and never compares against a trivial baseline (predict the mean / a single well-tuned model). The predictions come out heavily shrunk toward the mean (much smaller spread/range than the training target), which the summary presents as success.
- **Detection procedure**:
  1. Read the task/sample submission to identify the target and the evaluation metric implied by the competition (e.g. R²/RMSE/AUC).
  2. Scan the scripts for a train/validation split (or CV) and an explicit metric computation on held-out data; note whether any score is printed, and whether model capacity/hyperparameters were chosen by that score rather than fixed arbitrarily.
  3. Compare the reported prediction summary statistics (min, max, std, mean) with the training target's statistics; flag when prediction spread is a small fraction of the target spread or the range is far narrower.
  4. Check the answer for any baseline comparison; an answer that lists only architecture, row counts and prediction stats, with no validation number, fails.
- **Discriminator**: A real violation is the absence of any held-out metric *plus* a prediction distribution visibly inconsistent with the target's (severe shrinkage) or capacity chosen for speed alone. It is *not* a violation when the model is genuinely well-validated and the narrow spread is justified — e.g. a reported CV score near the achievable ceiling, or a classification task where calibrated probabilities legitimately concentrate; nor when the metric is rank-based and only ordering matters and that is verified.
- **Consequence**: The submission file has the right shape and passes format checks but scores far below threshold (near-constant, mean-reverting predictions give near-zero R²/poor RMSE), so the grader marks the expected result file as wrong despite the confident summary.
148Invented qualification threshold / aggregation rule instead of using the spec's stated definitiontaskda-code
Applies when
task -- The task or its documentation defines the target quantity with an explicit eligibility filter or aggregation rule (e.g., "must have at least N …", "average over …"), and the scripts must reproduce that definition exactly.
Pattern
The agent does not locate or fully read the stated definition (README section, sample/format file, or truncated spec), then hard-codes a self-chosen cutoff and a self-chosen aggregation (which entity attribute is summed vs. averaged, and which quantity the filter is applied to), presenting it as "as per guidance" without any citation of the source text.
Detection procedure
  1. In the task/README, find every stated definitional constraint (minimum counts, which field the minimum applies to, how per-entity values are aggregated, tie-breaking, ordering, output columns) and note anything truncated or ambiguous.
  2. In the scripts, locate each corresponding constant and aggregation call and check it against the text; also check whether the script ever loads/prints the provided sample-format file or the README to resolve ambiguity.
  3. Flag if any threshold, filter field, or aggregation is chosen by the agent rather than traceable to the spec, or if an ambiguity in the spec was silently resolved without a sensitivity check across plausible alternatives.
  4. Compare the produced output's columns/row count/ordering to the sample format file.
Discriminator
A real violation is a constant or aggregation that appears nowhere in the task text (or contradicts it) and materially changes the eligible population or ranking; a look-alike that is fine is an agent-chosen tie-break or implementation detail that is explicitly demanded by the spec or demonstrably does not change the reported ranking (verified by a robustness check).
Consequence
The eligible set and the ranking differ from the reference, so the saved file's entity lists mismatch the expected answer and the exact-match file check fails.
id 3e3f953bb667 · mined from da-code dacode-dm-csv-009@s3
raw text (what the judge reads)
### Invented qualification threshold / aggregation rule instead of using the spec's stated definition
- **Applies when**: `task` -- The task or its documentation defines the target quantity with an explicit eligibility filter or aggregation rule (e.g., "must have at least N …", "average over …"), and the scripts must reproduce that definition exactly.
- **Pattern**: The agent does not locate or fully read the stated definition (README section, sample/format file, or truncated spec), then hard-codes a self-chosen cutoff and a self-chosen aggregation (which entity attribute is summed vs. averaged, and which quantity the filter is applied to), presenting it as "as per guidance" without any citation of the source text.
- **Detection procedure**:
  1. In the task/README, find every stated definitional constraint (minimum counts, which field the minimum applies to, how per-entity values are aggregated, tie-breaking, ordering, output columns) and note anything truncated or ambiguous.
  2. In the scripts, locate each corresponding constant and aggregation call and check it against the text; also check whether the script ever loads/prints the provided sample-format file or the README to resolve ambiguity.
  3. Flag if any threshold, filter field, or aggregation is chosen by the agent rather than traceable to the spec, or if an ambiguity in the spec was silently resolved without a sensitivity check across plausible alternatives.
  4. Compare the produced output's columns/row count/ordering to the sample format file.
- **Discriminator**: A real violation is a constant or aggregation that appears nowhere in the task text (or contradicts it) and materially changes the eligible population or ranking; a look-alike that is fine is an agent-chosen tie-break or implementation detail that is explicitly demanded by the spec or demonstrably does not change the reported ranking (verified by a robustness check).
- **Consequence**: The eligible set and the ranking differ from the reference, so the saved file's entity lists mismatch the expected answer and the exact-match file check fails.
149Bins/categories invented ad hoc, silently dropping records instead of following the spec and covering the full data rangetaskda-code
Applies when
task -- the task asks for counts/distribution over ranges or groups of a numeric field, with a provided config/spec file, and the script defines bin edges and labels itself.
Pattern
The script hard-codes bin boundaries (and a label like "121+") that do not actually cover the observed min–max range, uses pd.cut without include_lowest/open upper edge, and mislabels the resulting categories; records outside the edges (and rows whose values failed the parsing step) are dropped without any check, so the plotted/reported counts describe only a subset. Additionally, the group definitions and required output artifacts are taken from the agent's own assumptions rather than exhaustively from the provided spec file.
Detection procedure
  1. From the task, list every constraint that the provided spec/config file is supposed to supply (group boundaries/labels, ordering, and every required output file) and note that the analyst must not substitute its own.
  2. In the script, check whether bin edges/labels are hard-coded rather than read from the spec, and whether the top/bottom bins are open-ended enough to include the printed min/max of the parsed values.
  3. Verify the script asserts that the sum of the group counts equals the number of rows in the intended population (after the legitimate type filter), and that parsing failures are inspected rather than silently dropped.
  4. Compare the counts and the artifact list in the final answer against the population size and required outputs; a large unexplained gap (e.g., counts summing to far fewer than the filtered rows) or missing artifacts is a violation.
Discriminator
A real violation is when rows in the target population vanish with no justification, labels do not match the edges actually used, or spec-provided grouping/outputs are replaced by invented ones. It is fine if rows are excluded by an explicitly required filter (e.g., a different content type) and the script documents/asserts that the retained counts sum to the filtered total.
Consequence
The produced figure and any saved count arrays/JSON disagree with the reference bin structure and totals, so every artifact check fails even though the script runs without error.
id 212937b9cc68 · mined from da-code dacode-plot-bar-007@s3
raw text (what the judge reads)
### Bins/categories invented ad hoc, silently dropping records instead of following the spec and covering the full data range
- **Applies when**: `task` -- the task asks for counts/distribution over ranges or groups of a numeric field, with a provided config/spec file, and the script defines bin edges and labels itself.
- **Pattern**: The script hard-codes bin boundaries (and a label like "121+") that do not actually cover the observed min–max range, uses `pd.cut` without `include_lowest`/open upper edge, and mislabels the resulting categories; records outside the edges (and rows whose values failed the parsing step) are dropped without any check, so the plotted/reported counts describe only a subset. Additionally, the group definitions and required output artifacts are taken from the agent's own assumptions rather than exhaustively from the provided spec file.
- **Detection procedure**:
  1. From the task, list every constraint that the provided spec/config file is supposed to supply (group boundaries/labels, ordering, and every required output file) and note that the analyst must not substitute its own.
  2. In the script, check whether bin edges/labels are hard-coded rather than read from the spec, and whether the top/bottom bins are open-ended enough to include the printed min/max of the parsed values.
  3. Verify the script asserts that the sum of the group counts equals the number of rows in the intended population (after the legitimate type filter), and that parsing failures are inspected rather than silently dropped.
  4. Compare the counts and the artifact list in the final answer against the population size and required outputs; a large unexplained gap (e.g., counts summing to far fewer than the filtered rows) or missing artifacts is a violation.
- **Discriminator**: A real violation is when rows in the target population vanish with no justification, labels do not match the edges actually used, or spec-provided grouping/outputs are replaced by invented ones. It is fine if rows are excluded by an explicitly required filter (e.g., a different content type) and the script documents/asserts that the retained counts sum to the filtered total.
- **Consequence**: The produced figure and any saved count arrays/JSON disagree with the reference bin structure and totals, so every artifact check fails even though the script runs without error.
150Answer not traceable to executed code on the provided data (plausible-looking values supplied from prior knowledge)taskda-code
Applies when
task -- the task asks for a ranking/statistic that must be computed from a supplied data file (with a stated preprocessing step such as a specific imputation rule), and the deliverable is a small JSON/text answer.
Pattern
The attempt produces a confident, plausible-sounding list of entity names without a saved, runnable script that loads the file, applies the required preprocessing, computes the ordering, and writes the requested output file; the values are effectively recalled from general knowledge or from a partially-read file, so label spellings, entity coverage, and the applied preprocessing/sort order cannot be verified against the actual data.
Detection procedure
  1. Read the task and list every mandated operation (preprocessing rule, sorting direction, top-N size, output keys/filename/format).
  2. Look for a script that performs each mandated step end-to-end and persists the answer to the required artifact; if no script exists, or it stops before computing/writing the final ranking, flag immediately.
  3. Cross-check the reported entity labels against the data file's own label column (exact strings, e.g. official vs. colloquial names) and confirm the reported list is a contiguous head/tail of the sorted, imputed column with the requested sort direction.
  4. Sanity-check counts and values: N entries per list, no duplicates/overlap between lists, and values within a physically plausible range for the quantity.
Discriminator
A real violation is an answer that cannot be reproduced from any artifact, or whose labels/order differ from what a script over the file would yield; a look-alike that is fine is an answer whose script exists and reproduces exactly these labels (matching the file's spellings) even if the agent also happened to know the result in advance.
Consequence
The graded artifact is missing or contains names/order that don't match the file-derived ranking, so exact-match comparison against the expected result file fails (0/1 checks).
id 724a6d2bc7b1 · mined from da-code dacode-di-text-003@s3
raw text (what the judge reads)
### Answer not traceable to executed code on the provided data (plausible-looking values supplied from prior knowledge)
- **Applies when**: `task` -- the task asks for a ranking/statistic that must be computed from a supplied data file (with a stated preprocessing step such as a specific imputation rule), and the deliverable is a small JSON/text answer.
- **Pattern**: The attempt produces a confident, plausible-sounding list of entity names without a saved, runnable script that loads the file, applies the required preprocessing, computes the ordering, and writes the requested output file; the values are effectively recalled from general knowledge or from a partially-read file, so label spellings, entity coverage, and the applied preprocessing/sort order cannot be verified against the actual data.
- **Detection procedure**:
  1. Read the task and list every mandated operation (preprocessing rule, sorting direction, top-N size, output keys/filename/format).
  2. Look for a script that performs each mandated step end-to-end and persists the answer to the required artifact; if no script exists, or it stops before computing/writing the final ranking, flag immediately.
  3. Cross-check the reported entity labels against the data file's own label column (exact strings, e.g. official vs. colloquial names) and confirm the reported list is a contiguous head/tail of the sorted, imputed column with the requested sort direction.
  4. Sanity-check counts and values: N entries per list, no duplicates/overlap between lists, and values within a physically plausible range for the quantity.
- **Discriminator**: A real violation is an answer that cannot be reproduced from any artifact, or whose labels/order differ from what a script over the file would yield; a look-alike that is fine is an answer whose script exists and reproduces exactly these labels (matching the file's spellings) even if the agent also happened to know the result in advance.
- **Consequence**: The graded artifact is missing or contains names/order that don't match the file-derived ranking, so exact-match comparison against the expected result file fails (0/1 checks).
151Ambiguous sign/absolute-value convention when constructing "difference" variables instead of using the dataset's own columnstaskinfiagent-dabench
Applies when
task -- the question asks about a relationship (correlation, regression, comparison) between two "difference"/"gap"/"change" quantities, and the script derives one or both of them by subtracting raw columns even though the file already contains pre-computed difference columns.
Pattern
The attempt picks one arbitrary orientation (e.g. B − A) and/or a signed-vs-absolute convention for the derived variable, without checking against the ready-made column(s) in the data that encode the same quantity; the two variables end up on mismatched conventions (one signed, one absolute, or opposite orientations), so the statistic's sign and its p-value/magnitude differ from the intended pairing. No validation of the derived column against the existing one is performed.
Detection procedure
  1. From the task, list the two quantities to be related and note that neither an orientation nor a signed/absolute convention is specified.
  2. In the scripts, find where each quantity is obtained; check whether an already-present column represents that same quantity and whether the script uses it or recomputes it.
  3. Verify the script explicitly compares its derived column to the existing one (equality/correlation check on both sign and magnitude, or a check for whether either column is non-negative everywhere) and states which convention it adopts for both variables consistently.
  4. Check the reported statistic for tell-tale inconsistency: a sign or p-value that changes if one variable were flipped or made absolute, with no justification in the script output.
Discriminator
A real violation is when the two variables are built under different conventions, or a provided difference column is silently ignored/overridden, so the result is convention-dependent and unverified. It is fine if the script uses the dataset's own difference columns directly, or shows a validation step proving its recomputation matches them (including sign), or if the statistic is provably invariant to the flip (e.g. both variables flipped together).
Consequence
The relationship classification may still land correctly by luck, but the correlation's sign and especially the p-value differ from the intended pairing, so the numeric fields (r sign, p-value to 4 decimals) fail the grader.
id 36614d3765d5 · mined from infiagent-dabench dabench-142@s3
raw text (what the judge reads)
### Ambiguous sign/absolute-value convention when constructing "difference" variables instead of using the dataset's own columns
- **Applies when**: `task` -- the question asks about a relationship (correlation, regression, comparison) between two "difference"/"gap"/"change" quantities, and the script derives one or both of them by subtracting raw columns even though the file already contains pre-computed difference columns.
- **Pattern**: The attempt picks one arbitrary orientation (e.g. B − A) and/or a signed-vs-absolute convention for the derived variable, without checking against the ready-made column(s) in the data that encode the same quantity; the two variables end up on mismatched conventions (one signed, one absolute, or opposite orientations), so the statistic's sign and its p-value/magnitude differ from the intended pairing. No validation of the derived column against the existing one is performed.
- **Detection procedure**:
  1. From the task, list the two quantities to be related and note that neither an orientation nor a signed/absolute convention is specified.
  2. In the scripts, find where each quantity is obtained; check whether an already-present column represents that same quantity and whether the script uses it or recomputes it.
  3. Verify the script explicitly compares its derived column to the existing one (equality/correlation check on both sign and magnitude, or a check for whether either column is non-negative everywhere) and states which convention it adopts for *both* variables consistently.
  4. Check the reported statistic for tell-tale inconsistency: a sign or p-value that changes if one variable were flipped or made absolute, with no justification in the script output.
- **Discriminator**: A real violation is when the two variables are built under different conventions, or a provided difference column is silently ignored/overridden, so the result is convention-dependent and unverified. It is fine if the script uses the dataset's own difference columns directly, or shows a validation step proving its recomputation matches them (including sign), or if the statistic is provably invariant to the flip (e.g. both variables flipped together).
- **Consequence**: The relationship classification may still land correctly by luck, but the correlation's sign and especially the p-value differ from the intended pairing, so the numeric fields (r sign, p-value to 4 decimals) fail the grader.
152Entity/quantity substitution: analyzing different data than the task specifiestaskda-code
Applies when
task -- the prompt names specific entities, groupings, and measured quantities (e.g. a ranking dimension, a per-group aggregate) and specific output artifacts, and the scripts choose their own data source and variables.
Pattern
The attempt loads whichever file it happens to find in the working directory and computes a superficially similar chart (same chart type, same "top 10" framing) over entirely different entities and metrics than the task requested, then declares success because a plot was produced.
Detection procedure
  1. From the task statement, list the required grouping key, the ranking metric, the measured quantity per segment, and every required output file.
  2. Read the scripts: identify the input file(s) actually loaded and the columns used for grouping, ranking, and the stacked segments.
  3. Check each item in the list from step 1 has a concrete counterpart in step 2; treat any silent substitution (different entity type, different ranking metric, different segment definition) as a failure, even if the config file was read.
  4. Confirm the answer/scripts write all named output artifacts, not just the image.
Discriminator
A real violation is a mismatch in the semantic content — different grouping entities, different ranking basis, or different measured quantity than asked. A look-alike that is fine is when the required columns are present but merely renamed/derived (e.g., a duration computed from timestamps), with the derivation clearly matching the requested definition.
Consequence
All expected output files fail comparison — the numeric array and the rendered figure encode unrelated quantities, and any missing artifact scores zero outright, so the run fails every check despite a "success" report.
id 8602b0e474d6 · mined from da-code dacode-plot-scatter-002@s3
raw text (what the judge reads)
### Entity/quantity substitution: analyzing different data than the task specifies
- **Applies when**: `task` -- the prompt names specific entities, groupings, and measured quantities (e.g. a ranking dimension, a per-group aggregate) and specific output artifacts, and the scripts choose their own data source and variables.
- **Pattern**: The attempt loads whichever file it happens to find in the working directory and computes a superficially similar chart (same chart type, same "top 10" framing) over entirely different entities and metrics than the task requested, then declares success because a plot was produced.
- **Detection procedure**:
  1. From the task statement, list the required grouping key, the ranking metric, the measured quantity per segment, and every required output file.
  2. Read the scripts: identify the input file(s) actually loaded and the columns used for grouping, ranking, and the stacked segments.
  3. Check each item in the list from step 1 has a concrete counterpart in step 2; treat any silent substitution (different entity type, different ranking metric, different segment definition) as a failure, even if the config file was read.
  4. Confirm the answer/scripts write *all* named output artifacts, not just the image.
- **Discriminator**: A real violation is a mismatch in the semantic content — different grouping entities, different ranking basis, or different measured quantity than asked. A look-alike that is fine is when the required columns are present but merely renamed/derived (e.g., a duration computed from timestamps), with the derivation clearly matching the requested definition.
- **Consequence**: All expected output files fail comparison — the numeric array and the rendered figure encode unrelated quantities, and any missing artifact scores zero outright, so the run fails every check despite a "success" report.
153Multi-stage grouping: filtering/statistic computed at the wrong grouping leveltaskda-code
Applies when
task -- the task prescribes a preprocessing step defined per one grouping variable (e.g., "filter within each level of A") followed by a test/metric computed across the levels of a second grouping variable, and the script implements both stages.
Pattern
The attempt collapses or shifts the grouping scope — it computes the filter thresholds globally over the whole table, or per cell of the A×B cross-tabulation, or per level of B instead of A — and/or runs the subsequent test on data that was not filtered as specified. Because the trimming changes each group's spread, the resulting variance/dispersion statistic and its p-value silently change sign of the conclusion, and no row-count or threshold check is emitted to expose it.
Detection procedure
  1. From the task text, write down explicitly: which variable defines the filtering partition, which variable defines the comparison groups, and which numeric column the thresholds are computed on.
  2. In the script, locate the groupby/mask that produces the thresholds and confirm the key matches the filtering partition (not the comparison groups, not the whole frame, not both keys), and that the thresholds are applied to the same numeric column named in the task.
  3. Confirm the statistical test is then run on the filtered subset, one result per level of the filtering partition, with groups formed by the comparison variable and all levels present (check group sizes are non-zero and that the number of reported values equals the number of partitions).
  4. Check the script prints pre/post filter row counts and per-group counts/thresholds; if no such sanity output exists, and the reported conclusions are uniform across all partitions, treat the result as unverified.
Discriminator
A real violation is a grouping key (or filter target column) in the code that differs from the one the task names, or a test run on unfiltered/differently filtered data. A look-alike that is fine: a script that groups exactly as specified but implements the quartile/outlier bound with a defensible convention (e.g., interpolation method or inclusive vs. exclusive bounds), which shifts p-values only marginally and preserves counts.
Consequence
The p-value list is numerically off (often by an order of magnitude near the decision threshold), one or more conclusion strings flip, and the emitted result file mismatches the expected values — scored wrong despite correct-looking format.
id 32bfd4283dbe · mined from da-code dacode-data-sa-061@s3
raw text (what the judge reads)
### Multi-stage grouping: filtering/statistic computed at the wrong grouping level
- **Applies when**: `task` -- the task prescribes a preprocessing step defined per one grouping variable (e.g., "filter within each level of A") followed by a test/metric computed across the levels of a *second* grouping variable, and the script implements both stages.
- **Pattern**: The attempt collapses or shifts the grouping scope — it computes the filter thresholds globally over the whole table, or per cell of the A×B cross-tabulation, or per level of B instead of A — and/or runs the subsequent test on data that was not filtered as specified. Because the trimming changes each group's spread, the resulting variance/dispersion statistic and its p-value silently change sign of the conclusion, and no row-count or threshold check is emitted to expose it.
- **Detection procedure**:
  1. From the task text, write down explicitly: which variable defines the filtering partition, which variable defines the comparison groups, and which numeric column the thresholds are computed on.
  2. In the script, locate the `groupby`/mask that produces the thresholds and confirm the key matches the filtering partition (not the comparison groups, not the whole frame, not both keys), and that the thresholds are applied to the same numeric column named in the task.
  3. Confirm the statistical test is then run on the *filtered* subset, one result per level of the filtering partition, with groups formed by the comparison variable and all levels present (check group sizes are non-zero and that the number of reported values equals the number of partitions).
  4. Check the script prints pre/post filter row counts and per-group counts/thresholds; if no such sanity output exists, and the reported conclusions are uniform across all partitions, treat the result as unverified.
- **Discriminator**: A real violation is a grouping key (or filter target column) in the code that differs from the one the task names, or a test run on unfiltered/differently filtered data. A look-alike that is fine: a script that groups exactly as specified but implements the quartile/outlier bound with a defensible convention (e.g., interpolation method or inclusive vs. exclusive bounds), which shifts p-values only marginally and preserves counts.
- **Consequence**: The p-value list is numerically off (often by an order of magnitude near the decision threshold), one or more conclusion strings flip, and the emitted result file mismatches the expected values — scored wrong despite correct-looking format.
154Settling for a near-baseline model instead of exploiting the identifying/high-signal columns shared by train and testtaskda-code
Applies when
task -- the task asks for per-row predictions of a target and the provided test file carries many of the same descriptive/identifier/metadata columns as the labeled source file, but the script builds features from only a generic numeric subset.
Pattern
The agent hand-picks a handful of "obvious" numeric columns, drops all categorical/identifier/date/text fields, trains, observes a validation score barely above predicting the mean (R² ≈ 0, or predictions collapsed into a narrow band around the global mean), and ships those predictions anyway while describing the near-mean output as a positive sign ("mean matches training mean").
Detection procedure
  1. From the task/README, list every column present in both the labeled file and the test file; note identifiers, names, dates, and categorical metadata that plausibly carry strong signal about the target.
  2. In the scripts, check the feature list actually used and whether any step tests for exact or key-based overlap between test rows and labeled rows (a lookup/merge that could recover many targets directly).
  3. Read the reported validation metric and the prediction spread: if R² is near zero (or RMSE ≈ target std) and predicted min/max span a small fraction of the true target range, the model is essentially a constant predictor.
  4. Confirm the agent never compared its model against a trivial baseline (global mean / group mean by an unused categorical key) or tried adding the discarded columns.
Discriminator
A real violation is when informative shared columns were available and simply ignored, and the resulting score is indistinguishable from the mean baseline. It is not a violation if the discarded columns are truly unavailable/leaky in the test file, or if the agent tried them, documented that they did not help, and the low score reflects a genuinely weak signal validated against a baseline.
Consequence
The submitted prediction file is essentially a constant column; any accuracy/correlation/error threshold the grader applies against the true targets fails, and near-duplicate rows whose targets could have been recovered exactly are mispredicted.
id 9311eace719d · mined from da-code dacode-ml-regression-004@s3
raw text (what the judge reads)
### Settling for a near-baseline model instead of exploiting the identifying/high-signal columns shared by train and test
- **Applies when**: `task` -- the task asks for per-row predictions of a target and the provided test file carries many of the same descriptive/identifier/metadata columns as the labeled source file, but the script builds features from only a generic numeric subset.
- **Pattern**: The agent hand-picks a handful of "obvious" numeric columns, drops all categorical/identifier/date/text fields, trains, observes a validation score barely above predicting the mean (R² ≈ 0, or predictions collapsed into a narrow band around the global mean), and ships those predictions anyway while describing the near-mean output as a positive sign ("mean matches training mean").
- **Detection procedure**:
  1. From the task/README, list every column present in both the labeled file and the test file; note identifiers, names, dates, and categorical metadata that plausibly carry strong signal about the target.
  2. In the scripts, check the feature list actually used and whether any step tests for exact or key-based overlap between test rows and labeled rows (a lookup/merge that could recover many targets directly).
  3. Read the reported validation metric and the prediction spread: if R² is near zero (or RMSE ≈ target std) and predicted min/max span a small fraction of the true target range, the model is essentially a constant predictor.
  4. Confirm the agent never compared its model against a trivial baseline (global mean / group mean by an unused categorical key) or tried adding the discarded columns.
- **Discriminator**: A real violation is when informative shared columns were available and simply ignored, and the resulting score is indistinguishable from the mean baseline. It is *not* a violation if the discarded columns are truly unavailable/leaky in the test file, or if the agent tried them, documented that they did not help, and the low score reflects a genuinely weak signal validated against a baseline.
- **Consequence**: The submitted prediction file is essentially a constant column; any accuracy/correlation/error threshold the grader applies against the true targets fails, and near-duplicate rows whose targets could have been recovered exactly are mispredicted.
155Requested output file contains internal intermediates instead of the original data valuestaskda-code
Applies when
task -- the task asks for a result file whose columns echo the input feature values alongside a computed label/prediction, and the script applies a transform (scaling, PCA, imputation, encoding) before modeling.
Pattern
The attempt builds the output table from the transformed matrix used inside the model (e.g. z-scored or projected values) rather than from the original input rows, and/or silently drops/reorders identifier columns and rows, so the file's feature columns no longer match the source data even though the label column is plausible.
Detection procedure
  1. Read the task to determine exactly what each requested column should hold (raw feature values vs. derived ones), the expected row count, and column naming/order.
  2. In the script, find the object passed to to_csv and trace it back: is it derived from the raw dataframe, or from the output of a scaler/decomposition/encoder?
  3. Check whether every original row is present in original order and whether the number/meaning of feature columns equals the number of input features (a transform that changes dimensionality, or a fit on a subset, is an extra red flag).
  4. Compare the answer's description of the file (sample rows, value ranges) against the source data's ranges — negative/near-zero-mean values where the raw data is positive counts, rates, or currency indicate a transformed dump.
Discriminator
Using a transform inside the model is correct and expected; the violation is only when the saved output stores transformed values where the task implies the original feature vector, or when rows/columns of the saved file no longer align 1:1 with the input records. If the task explicitly asks for scaled/embedded coordinates, writing them is fine.
Consequence
The result file fails value-level comparison against the expected file (feature columns mismatch, or join on records fails), so the check scores 0 even though the clustering/labels themselves may be reasonable.
id 85b20549f731 · mined from da-code dacode-ml-cluster-013@s3
raw text (what the judge reads)
### Requested output file contains internal intermediates instead of the original data values
- **Applies when**: `task` -- the task asks for a result file whose columns echo the input feature values alongside a computed label/prediction, and the script applies a transform (scaling, PCA, imputation, encoding) before modeling.
- **Pattern**: The attempt builds the output table from the transformed matrix used inside the model (e.g. z-scored or projected values) rather than from the original input rows, and/or silently drops/reorders identifier columns and rows, so the file's feature columns no longer match the source data even though the label column is plausible.
- **Detection procedure**:
  1. Read the task to determine exactly what each requested column should hold (raw feature values vs. derived ones), the expected row count, and column naming/order.
  2. In the script, find the object passed to `to_csv` and trace it back: is it derived from the raw dataframe, or from the output of a scaler/decomposition/encoder?
  3. Check whether every original row is present in original order and whether the number/meaning of feature columns equals the number of input features (a transform that changes dimensionality, or a fit on a subset, is an extra red flag).
  4. Compare the answer's description of the file (sample rows, value ranges) against the source data's ranges — negative/near-zero-mean values where the raw data is positive counts, rates, or currency indicate a transformed dump.
- **Discriminator**: Using a transform *inside* the model is correct and expected; the violation is only when the *saved output* stores transformed values where the task implies the original feature vector, or when rows/columns of the saved file no longer align 1:1 with the input records. If the task explicitly asks for scaled/embedded coordinates, writing them is fine.
- **Consequence**: The result file fails value-level comparison against the expected file (feature columns mismatch, or join on records fails), so the check scores 0 even though the clustering/labels themselves may be reasonable.
156Stated subset/scope qualifier silently redefined instead of appliedtaskinfiagent-dabench
Applies when
task -- the question names an explicit slice (a specific year/category/group) and a population scope ("among all countries/records in the dataset"), and the script computes a per-entity statistic from a wide/multi-file table.
Pattern
The agent decides the named slice is "impossible" for the statistic (e.g., a single value can't be skewed), reinterprets the question by aggregating over a different axis (all periods instead of the named one), and/or loads only one of several files that together form the stated population — then rationalizes the substitution in comments rather than testing the literal reading.
Detection procedure
  1. From the task text, list every qualifier: the filter value, the entity being ranked, the population scope, and the required definition/flag of the statistic.
  2. In the scripts, locate where each qualifier is enforced: is there a filter on the named slice? are all relevant input files/rows loaded? does the reduction axis correspond to the entity in the question?
  3. Flag the attempt if any qualifier appears only in a comment/justification (e.g., "this must mean X instead") or is replaced by a broader/narrower set, or if the flag for the requested definition is asserted without checking library semantics against the documented meaning.
  4. Check the reported answer is the entity that maximizes the statistic under the literal qualifiers, not under the substituted interpretation.
Discriminator
A real violation is when the literal slice/scope was never computed or even attempted, and the deviation is defended by prose only. It is acceptable if the agent computes the literal reading, demonstrates with evidence that it is degenerate or empty, and then documents a fallback while still covering the full stated population and correct statistic definition.
Consequence
The ranking is over a different distribution/population than requested, so the top entity differs from ground truth and the graded key mismatches (0/1), even though the code runs cleanly.
id 6382f08f763c · mined from infiagent-dabench dabench-252@s3
raw text (what the judge reads)
### Stated subset/scope qualifier silently redefined instead of applied
- **Applies when**: `task` -- the question names an explicit slice (a specific year/category/group) and a population scope ("among all countries/records in the dataset"), and the script computes a per-entity statistic from a wide/multi-file table.
- **Pattern**: The agent decides the named slice is "impossible" for the statistic (e.g., a single value can't be skewed), reinterprets the question by aggregating over a different axis (all periods instead of the named one), and/or loads only one of several files that together form the stated population — then rationalizes the substitution in comments rather than testing the literal reading.
- **Detection procedure**:
  1. From the task text, list every qualifier: the filter value, the entity being ranked, the population scope, and the required definition/flag of the statistic.
  2. In the scripts, locate where each qualifier is enforced: is there a filter on the named slice? are all relevant input files/rows loaded? does the reduction axis correspond to the entity in the question?
  3. Flag the attempt if any qualifier appears only in a comment/justification (e.g., "this must mean X instead") or is replaced by a broader/narrower set, or if the flag for the requested definition is asserted without checking library semantics against the documented meaning.
  4. Check the reported answer is the entity that maximizes the statistic under the literal qualifiers, not under the substituted interpretation.
- **Discriminator**: A real violation is when the literal slice/scope was never computed or even attempted, and the deviation is defended by prose only. It is acceptable if the agent computes the literal reading, demonstrates with evidence that it is degenerate or empty, and then documents a fallback while still covering the full stated population and correct statistic definition.
- **Consequence**: The ranking is over a different distribution/population than requested, so the top entity differs from ground truth and the graded key mismatches (0/1), even though the code runs cleanly.
157Reference format file never opened; output schema guessed instead of derivedtaskda-code
Applies when
task -- the instructions require the deliverable to match the format of a provided sample/template file (column names, order, row set, rounding, index).
Pattern
The scripts compute the statistic but never load, print, or compare against the supplied template; the agent invents header names, column order, row labels and rounding from the prose of the task, and writes the file without any format assertion. A second "verification" script only re-checks the arithmetic, not the output contract.
Detection procedure
  1. Read the task and note every stated output constraint and the name of any reference/sample artifact.
  2. Grep the scripts for that artifact's filename and for any read of it; check whether the produced column names, ordering, row keys and numeric formatting are taken from it (or explicitly asserted equal to it) rather than hard-coded.
  3. Check whether the final write step is preceded by a comparison of the produced frame's shape/columns/row labels to the template's.
  4. Inspect the submitted file: are its headers, label spellings, row ordering and decimal precision provably derived from the template, or plausible-looking guesses?
Discriminator
A real violation is hard-coded or renamed output fields with no read/assert of the reference file; it is fine if the script reads the template (or an explicit spec in the task text) and the headers/labels/precision demonstrably come from it — even if the names happen to be simple.
Consequence
The grader compares against the expected file and reports the deliverable as WRONG/MISSING, because header text, label spelling, row coverage/order or rounding differ, even when the underlying aggregation may be numerically close or correct.
id 00e9d774045c · mined from da-code dacode-dm-csv-010@s3
raw text (what the judge reads)
### Reference format file never opened; output schema guessed instead of derived
- **Applies when**: `task` -- the instructions require the deliverable to match the format of a provided sample/template file (column names, order, row set, rounding, index).
- **Pattern**: The scripts compute the statistic but never load, print, or compare against the supplied template; the agent invents header names, column order, row labels and rounding from the prose of the task, and writes the file without any format assertion. A second "verification" script only re-checks the arithmetic, not the output contract.
- **Detection procedure**:
  1. Read the task and note every stated output constraint and the name of any reference/sample artifact.
  2. Grep the scripts for that artifact's filename and for any read of it; check whether the produced column names, ordering, row keys and numeric formatting are taken from it (or explicitly asserted equal to it) rather than hard-coded.
  3. Check whether the final write step is preceded by a comparison of the produced frame's shape/columns/row labels to the template's.
  4. Inspect the submitted file: are its headers, label spellings, row ordering and decimal precision provably derived from the template, or plausible-looking guesses?
- **Discriminator**: A real violation is hard-coded or renamed output fields with no read/assert of the reference file; it is fine if the script reads the template (or an explicit spec in the task text) and the headers/labels/precision demonstrably come from it — even if the names happen to be simple.
- **Consequence**: The grader compares against the expected file and reports the deliverable as WRONG/MISSING, because header text, label spelling, row coverage/order or rounding differ, even when the underlying aggregation may be numerically close or correct.
158Truncating or reformatting an identifier value to match a template instead of reporting the value found in the datataskinfiagent-dabench
Applies when
task -- the deliverable includes an identifier (a date, key, ID, label) that must be looked up in the data and reported, and the answer template shows a shortened/abstract pattern (e.g., a coarser granularity than the records carry).
Pattern
The script correctly locates the row of interest but the reported identifier is coerced to the literal template pattern (formatted/truncated/rounded to a coarser unit), discarding precision that exists in the source records, so the reported key no longer uniquely identifies the row actually used in the downstream computation.
Detection procedure
  1. In the task, note the identifier requested and the granularity actually present in the data (row-level key vs. aggregated period); note whether the task ever asks for aggregation to that coarser unit.
  2. In the scripts, find where the identifier is emitted and check for any strftime/slicing/astype(str)[:n]/rounding/relabeling applied only for output formatting.
  3. Compare the emitted identifier with the identifier of the row used for the dependent calculation (e.g., the neighboring-row lookup); if the dependent step uses full precision but the printed key is coarser, flag it.
  4. Check the printed value is still an existing, unique key in the data — a coarser string that matches many rows is a red flag.
Discriminator
Fine if the task explicitly required aggregation to the coarser unit and the analysis was performed at that unit throughout; a violation when the selection/computation happens at fine granularity and only the printed string is truncated to imitate the format example.
Consequence
The dependent numeric answer matches but the identifier check fails on exact-string comparison, so the submission is graded partially correct / incorrect.
id cae30821a726 · mined from infiagent-dabench dabench-572@s3
raw text (what the judge reads)
### Truncating or reformatting an identifier value to match a template instead of reporting the value found in the data
- **Applies when**: `task` -- the deliverable includes an identifier (a date, key, ID, label) that must be looked up in the data and reported, and the answer template shows a shortened/abstract pattern (e.g., a coarser granularity than the records carry).
- **Pattern**: The script correctly locates the row of interest but the reported identifier is coerced to the literal template pattern (formatted/truncated/rounded to a coarser unit), discarding precision that exists in the source records, so the reported key no longer uniquely identifies the row actually used in the downstream computation.
- **Detection procedure**:
  1. In the task, note the identifier requested and the granularity actually present in the data (row-level key vs. aggregated period); note whether the task ever asks for aggregation to that coarser unit.
  2. In the scripts, find where the identifier is emitted and check for any `strftime`/slicing/`astype(str)[:n]`/rounding/relabeling applied only for output formatting.
  3. Compare the emitted identifier with the identifier of the row used for the dependent calculation (e.g., the neighboring-row lookup); if the dependent step uses full precision but the printed key is coarser, flag it.
  4. Check the printed value is still an existing, unique key in the data — a coarser string that matches many rows is a red flag.
- **Discriminator**: Fine if the task explicitly required aggregation to the coarser unit and the analysis was performed at that unit throughout; a violation when the selection/computation happens at fine granularity and only the printed string is truncated to imitate the format example.
- **Consequence**: The dependent numeric answer matches but the identifier check fails on exact-string comparison, so the submission is graded partially correct / incorrect.
159Fabricated/synthetic stand-in data instead of the provided datasettaskda-code
Applies when
task -- the task asks for a statistic computed from a specific real dataset, and the scripts hard-code, simulate, or "reconstruct from memory" the input values.
Pattern
An initial exploration script fails to locate (or gives up on locating) the real input files, so the agent hand-writes plausible-looking arrays or generates random data, runs the requested computation on that invented input, and reports the resulting number as if it were the real answer.
Detection procedure
  1. Read the task to identify what input data the requested statistic must come from, and note that the answer is data-determined (a single number/table).
  2. Scan the analysis scripts for the data source: is every model input traced back to a file that was actually read (read_csv/read_excel/API), or does it come from literal arrays, np.random, or comments like "based on historical data"/"synthetic"?
  3. Check whether the discovery step actually succeeded (printed file list, shapes, column names) or whether the script proceeded on a fallback path after finding nothing; check the whole filesystem/data directory was searched (other extensions, nested dirs, package-bundled data) before giving up.
  4. Compare the reported answer to any traceable input: if no real rows/columns underlie it, the answer is unverifiable regardless of how the arithmetic looks.
Discriminator
A violation is when the substantive input values driving the reported result are invented; it is fine to hard-code small auxiliary constants (date ranges, thresholds, parameter grids) or to build toy data purely in a separate self-test/demo that does not produce the submitted answer. Also fine if the task itself specifies the numbers or asks for a simulation.
Consequence
The reported statistic is an arbitrary number unrelated to the true data, so the expected output file fails the value comparison (grader marks result.csv WRONG) even though the script runs cleanly and the format looks right.
id 89ce0185d934 · mined from da-code dacode-data-sa-043@s3
raw text (what the judge reads)
### Fabricated/synthetic stand-in data instead of the provided dataset
- **Applies when**: `task` -- the task asks for a statistic computed from a specific real dataset, and the scripts hard-code, simulate, or "reconstruct from memory" the input values.
- **Pattern**: An initial exploration script fails to locate (or gives up on locating) the real input files, so the agent hand-writes plausible-looking arrays or generates random data, runs the requested computation on that invented input, and reports the resulting number as if it were the real answer.
- **Detection procedure**:
  1. Read the task to identify what input data the requested statistic must come from, and note that the answer is data-determined (a single number/table).
  2. Scan the analysis scripts for the data source: is every model input traced back to a file that was actually read (`read_csv`/`read_excel`/API), or does it come from literal arrays, `np.random`, or comments like "based on historical data"/"synthetic"?
  3. Check whether the discovery step actually succeeded (printed file list, shapes, column names) or whether the script proceeded on a fallback path after finding nothing; check the whole filesystem/data directory was searched (other extensions, nested dirs, package-bundled data) before giving up.
  4. Compare the reported answer to any traceable input: if no real rows/columns underlie it, the answer is unverifiable regardless of how the arithmetic looks.
- **Discriminator**: A violation is when the *substantive* input values driving the reported result are invented; it is fine to hard-code small auxiliary constants (date ranges, thresholds, parameter grids) or to build toy data purely in a separate self-test/demo that does not produce the submitted answer. Also fine if the task itself specifies the numbers or asks for a simulation.
- **Consequence**: The reported statistic is an arbitrary number unrelated to the true data, so the expected output file fails the value comparison (grader marks result.csv WRONG) even though the script runs cleanly and the format looks right.
160Imbalanced-class predictions left at the default 0.5 threshold, selected by a threshold-free metrictaskda-code
Applies when
task -- the task asks for hard class labels on a rare-event target and the scripts choose/tune models with a ranking metric (AUC) while emitting labels via .predict().
Pattern
The agent optimizes and compares models on ROC-AUC (or similar threshold-independent score), never computes the metric the task actually implies (recall / cost of false negatives / F1 on the minority class), and writes labels from the default 0.5 cutoff. With a skewed target this collapses toward the majority class, so the submitted positive rate is far below the training prevalence and most true positives are missed. Class-weighting alone is treated as sufficient, and no check is made that the output label distribution is plausible.
Detection procedure
  1. Read the task: does it require hard labels and state an asymmetric cost (missed events are expensive) or an imbalanced target?
  2. Read the scripts: is model selection done only on AUC/log-loss, with predictions from predict() and no threshold search or minority-class metric (recall/F-beta/confusion matrix) on a validation split?
  3. Compute the positive rate in the training labels and in the submitted file; compare them.
  4. Check whether any sanity check on the output (row count, label counts, distribution vs. prior) exists in the scripts.
Discriminator
A real violation is when the objective is cost-sensitive/imbalanced and the submitted positive rate is materially below (roughly half or less of) the training prevalence with no justification. It is not a violation if the agent explicitly tuned the threshold on held-out data for the stated cost/metric, or if the predicted rate is close to the prior and the chosen operating point was validated.
Consequence
The prediction file has too few positives; recall / F1 / cost-weighted score on the hidden labels falls below the grading bar and the submission is marked wrong even though AUC looked high.
id d12d710edcbd · mined from da-code dacode-ml-binary-013@s3
raw text (what the judge reads)
### Imbalanced-class predictions left at the default 0.5 threshold, selected by a threshold-free metric
- **Applies when**: `task` -- the task asks for hard class labels on a rare-event target and the scripts choose/tune models with a ranking metric (AUC) while emitting labels via `.predict()`.
- **Pattern**: The agent optimizes and compares models on ROC-AUC (or similar threshold-independent score), never computes the metric the task actually implies (recall / cost of false negatives / F1 on the minority class), and writes labels from the default 0.5 cutoff. With a skewed target this collapses toward the majority class, so the submitted positive rate is far below the training prevalence and most true positives are missed. Class-weighting alone is treated as sufficient, and no check is made that the output label distribution is plausible.
- **Detection procedure**:
  1. Read the task: does it require hard labels and state an asymmetric cost (missed events are expensive) or an imbalanced target?
  2. Read the scripts: is model selection done only on AUC/log-loss, with predictions from `predict()` and no threshold search or minority-class metric (recall/F-beta/confusion matrix) on a validation split?
  3. Compute the positive rate in the training labels and in the submitted file; compare them.
  4. Check whether any sanity check on the output (row count, label counts, distribution vs. prior) exists in the scripts.
- **Discriminator**: A real violation is when the objective is cost-sensitive/imbalanced and the submitted positive rate is materially below (roughly half or less of) the training prevalence with no justification. It is *not* a violation if the agent explicitly tuned the threshold on held-out data for the stated cost/metric, or if the predicted rate is close to the prior and the chosen operating point was validated.
- **Consequence**: The prediction file has too few positives; recall / F1 / cost-weighted score on the hidden labels falls below the grading bar and the submission is marked wrong even though AUC looked high.
161Answer serialization not verifiably matching the requested output token, with no reproducible script producing ittaskinfiagent-dabench
Applies when
task -- the task specifies a literal answer format (a named tag plus a list/scalar in a given delimiter/quoting style) and expects the answer to be produced by saved, re-runnable code.
Pattern
The attempt computes a plausible result but hand-writes the final answer string, guessing at the list syntax (quotes, brackets, separators, ordering, name spelling) instead of emitting the exact requested token from a script; no script is saved, so the answer cannot be regenerated or checked character-by-character against the required template.
Detection procedure
  1. Read the task's "Answer format" line and write down the exact template: tag name, bracket/quote style, separator, and any casing/spelling source for the item labels.
  2. Check the scripts for a step that prints the final answer in exactly that template (e.g., a formatted print built from the computed values) — if no script exists or the final string is typed by hand, flag it.
  3. Compare the submitted string to the template token-by-token: tag spelling, delimiters, quoting, whitespace, item order, and whether labels are copied verbatim from the source data rather than re-typed.
  4. Confirm the values inside the token are the requested final quantity (the entity labels asked for), not indices, counts, or intermediate values.
Discriminator
A real violation is when the answer's syntax/labels are not demonstrably generated from the data by saved code, or deviate from the stated template in any character; a look-alike that is fine is an answer whose exact string is printed by a preserved script and matches the template literally, even if the reviewer would have chosen different formatting.
Consequence
The grader parses the answer token strictly and reports the expected value as WRONG/MISSING even when the underlying computation was right, yielding 0/1 checks passed and no way to audit or repair the result.
id 11366cba5f4a · mined from infiagent-dabench dabench-254@s3
raw text (what the judge reads)
### Answer serialization not verifiably matching the requested output token, with no reproducible script producing it
- **Applies when**: `task` -- the task specifies a literal answer format (a named tag plus a list/scalar in a given delimiter/quoting style) and expects the answer to be produced by saved, re-runnable code.
- **Pattern**: The attempt computes a plausible result but hand-writes the final answer string, guessing at the list syntax (quotes, brackets, separators, ordering, name spelling) instead of emitting the exact requested token from a script; no script is saved, so the answer cannot be regenerated or checked character-by-character against the required template.
- **Detection procedure**:
  1. Read the task's "Answer format" line and write down the exact template: tag name, bracket/quote style, separator, and any casing/spelling source for the item labels.
  2. Check the scripts for a step that prints the final answer in exactly that template (e.g., a formatted print built from the computed values) — if no script exists or the final string is typed by hand, flag it.
  3. Compare the submitted string to the template token-by-token: tag spelling, delimiters, quoting, whitespace, item order, and whether labels are copied verbatim from the source data rather than re-typed.
  4. Confirm the values inside the token are the requested final quantity (the entity labels asked for), not indices, counts, or intermediate values.
- **Discriminator**: A real violation is when the answer's syntax/labels are not demonstrably generated from the data by saved code, or deviate from the stated template in any character; a look-alike that is fine is an answer whose exact string is printed by a preserved script and matches the template literally, even if the reviewer would have chosen different formatting.
- **Consequence**: The grader parses the answer token strictly and reports the expected value as WRONG/MISSING even when the underlying computation was right, yielding 0/1 checks passed and no way to audit or repair the result.
162Ad-hoc invented metric computed over only one side of a paired-entity tabletaskda-code
Applies when
task -- the task asks for a per-entity statistic ("performance", "score", "total") derived from a table where each record lists two entities in distinct role columns (or otherwise stores an entity's records across multiple columns/subsets), and the exact formula is not spelled out in the prompt.
Pattern
The agent invents a formula out of thin air (e.g., arbitrary weights on an incidental flag column), and computes it by grouping on a single role column, so every record where the entity appears in the other role is silently dropped; no reference to the config/spec or to a standard domain definition is made, and no required companion outputs (serialized numbers/metadata files) are produced.
Detection procedure
  1. Read the task and any provided config/spec file for the definition of the requested quantity and the full list of expected deliverables; note whether the metric is defined anywhere or must follow a conventional definition.
  2. In the script, find the aggregation: check whether it iterates/groups over all columns in which an entity can appear, or only one; check whether the arithmetic combining columns is justified by the spec rather than asserted by the agent ("weighting emphasizes ... as they are considered more challenging").
  3. Cross-check counts: does the total number of records attributed to entities equal the expected multiple of the filtered row count (each row should contribute to both participants)? Does the script write every artifact the task/config implies, not just the image?
  4. Read the answer: are the reported numbers explained by a self-defined formula, and do the leaders look implausible for the domain (unknown entities outranking dominant ones)?
Discriminator
A real violation is an unsourced formula and/or an aggregation that can only see part of each entity's records; it is fine if the agent's formula is quoted from the task/config/README or is the standard domain definition, and the grouping demonstrably unions all roles (e.g., melt/concat of both role columns) even if the final numbers differ slightly.
Consequence
Every derived artifact (chart values, serialized arrays, metadata) mismatches the reference, so all output checks fail even though the plot's cosmetic settings are correct.
id caee41bf889d · mined from da-code dacode-plot-bar-006@s3
raw text (what the judge reads)
### Ad-hoc invented metric computed over only one side of a paired-entity table
- **Applies when**: `task` -- the task asks for a per-entity statistic ("performance", "score", "total") derived from a table where each record lists two entities in distinct role columns (or otherwise stores an entity's records across multiple columns/subsets), and the exact formula is not spelled out in the prompt.
- **Pattern**: The agent invents a formula out of thin air (e.g., arbitrary weights on an incidental flag column), and computes it by grouping on a single role column, so every record where the entity appears in the other role is silently dropped; no reference to the config/spec or to a standard domain definition is made, and no required companion outputs (serialized numbers/metadata files) are produced.
- **Detection procedure**:
  1. Read the task and any provided config/spec file for the definition of the requested quantity and the full list of expected deliverables; note whether the metric is defined anywhere or must follow a conventional definition.
  2. In the script, find the aggregation: check whether it iterates/groups over all columns in which an entity can appear, or only one; check whether the arithmetic combining columns is justified by the spec rather than asserted by the agent ("weighting emphasizes ... as they are considered more challenging").
  3. Cross-check counts: does the total number of records attributed to entities equal the expected multiple of the filtered row count (each row should contribute to both participants)? Does the script write every artifact the task/config implies, not just the image?
  4. Read the answer: are the reported numbers explained by a self-defined formula, and do the leaders look implausible for the domain (unknown entities outranking dominant ones)?
- **Discriminator**: A real violation is an unsourced formula and/or an aggregation that can only see part of each entity's records; it is fine if the agent's formula is quoted from the task/config/README or is the standard domain definition, and the grouping demonstrably unions all roles (e.g., melt/concat of both role columns) even if the final numbers differ slightly.
- **Consequence**: Every derived artifact (chart values, serialized arrays, metadata) mismatches the reference, so all output checks fail even though the plot's cosmetic settings are correct.
163Submission never validated against the provided output templatetaskda-code
Applies when
task -- The task specifies that results be written to an output file whose exact schema is defined by a provided example/template file (column names, row count, id ordering, value constraints).
Pattern
The scripts load only the training/test/auxiliary data and construct the output file from hard-coded column names and whatever prediction array is in memory; the template file is never read, and no post-write check compares the produced file's header, row count, id set/order, or per-row value constraints (e.g., probabilities summing to 1, values in range) against the template or against the test set.
Detection procedure
  1. Read the task/README and note the named template file and the required output schema (id column, per-class columns, one row per test record).
  2. Grep the scripts for a read of that template file and for any assertion/print that compares produced columns, dtypes, row count, and id ordering to it (and to the test input's ids).
  3. Inspect the produced answer file: check the header spelling matches the template exactly, that the number of data rows equals the number of test rows, that ids are unique and cover the test ids in the expected order, and that values satisfy the stated constraints.
  4. If any of these checks is absent from the scripts and cannot be confirmed from the answer itself, flag the attempt.
Discriminator
A real violation is when the template is never loaded and no shape/header/id/row-count check exists (so a mismatch in column naming, missing/extra rows, or reordered ids would go undetected). It is not a violation if the script builds the output by copying the template's id column/columns and asserts equality of shape and ids, even if the naming happens to be written literally.
Consequence
The grader compares the file to the expected submission schema and marks it WRONG/MISSING (unparseable or shape/header/id mismatch), so the model's actual predictive quality is never scored.
id 2addb1b27157 · mined from da-code dacode-ml-competition-003@s3
raw text (what the judge reads)
### Submission never validated against the provided output template
- **Applies when**: `task` -- The task specifies that results be written to an output file whose exact schema is defined by a provided example/template file (column names, row count, id ordering, value constraints).
- **Pattern**: The scripts load only the training/test/auxiliary data and construct the output file from hard-coded column names and whatever prediction array is in memory; the template file is never read, and no post-write check compares the produced file's header, row count, id set/order, or per-row value constraints (e.g., probabilities summing to 1, values in range) against the template or against the test set.
- **Detection procedure**:
  1. Read the task/README and note the named template file and the required output schema (id column, per-class columns, one row per test record).
  2. Grep the scripts for a read of that template file and for any assertion/print that compares produced columns, dtypes, row count, and id ordering to it (and to the test input's ids).
  3. Inspect the produced answer file: check the header spelling matches the template exactly, that the number of data rows equals the number of test rows, that ids are unique and cover the test ids in the expected order, and that values satisfy the stated constraints.
  4. If any of these checks is absent from the scripts and cannot be confirmed from the answer itself, flag the attempt.
- **Discriminator**: A real violation is when the template is never loaded and no shape/header/id/row-count check exists (so a mismatch in column naming, missing/extra rows, or reordered ids would go undetected). It is *not* a violation if the script builds the output by copying the template's id column/columns and asserts equality of shape and ids, even if the naming happens to be written literally.
- **Consequence**: The grader compares the file to the expected submission schema and marks it WRONG/MISSING (unparseable or shape/header/id mismatch), so the model's actual predictive quality is never scored.
164Fabricated labels / ignoring the provided train and test filestaskda-code
Applies when
task -- the task names specific train/test files and a target column, and the scripts must produce predictions for exactly the given test rows.
Pattern
The agent never loads the specified training file (so never sees real target labels) and never loads the specified test file; instead it stitches together auxiliary tables, invents the target with a hand-made heuristic score, trains on those synthetic labels, and emits predictions for its own row set with its own key column and label vocabulary.
Detection procedure
  1. From the task, list the required input files, the target column, and the required output columns/row set.
  2. Grep the scripts for reads of the named train/test files and for any place the real target column is loaded from data; flag if the target is instead computed by rules/thresholds inside the script.
  3. Compare the output row count and index/key to the test file's row count and key order, and compare the predicted label set to the label values actually present in the training data.
  4. Check reported accuracy: near-100% train and test accuracy on a rule-derived target is a tell that the label is a deterministic function of the features (self-consistency, not prediction).
Discriminator
A real violation is when no genuine labels ever enter the pipeline or the prediction set doesn't align with the specified test rows/format; it is fine if the agent legitimately loads the given labels and merges auxiliary tables as extra features, or engineers heuristic features (not the target) while still fitting on true labels.
Consequence
The output file has the wrong number/order of rows, wrong or extra columns, and label values unrelated to the true target, so the grader scores it as wrong/missing regardless of the reported "100% accuracy."
id 279c3e190f2c · mined from da-code dacode-ml-multi-003@s3
raw text (what the judge reads)
### Fabricated labels / ignoring the provided train and test files
- **Applies when**: `task` -- the task names specific train/test files and a target column, and the scripts must produce predictions for exactly the given test rows.
- **Pattern**: The agent never loads the specified training file (so never sees real target labels) and never loads the specified test file; instead it stitches together auxiliary tables, invents the target with a hand-made heuristic score, trains on those synthetic labels, and emits predictions for its own row set with its own key column and label vocabulary.
- **Detection procedure**:
  1. From the task, list the required input files, the target column, and the required output columns/row set.
  2. Grep the scripts for reads of the named train/test files and for any place the real target column is loaded from data; flag if the target is instead computed by rules/thresholds inside the script.
  3. Compare the output row count and index/key to the test file's row count and key order, and compare the predicted label set to the label values actually present in the training data.
  4. Check reported accuracy: near-100% train *and* test accuracy on a rule-derived target is a tell that the label is a deterministic function of the features (self-consistency, not prediction).
- **Discriminator**: A real violation is when no genuine labels ever enter the pipeline or the prediction set doesn't align with the specified test rows/format; it is fine if the agent legitimately loads the given labels and merges auxiliary tables as extra features, or engineers heuristic *features* (not the target) while still fitting on true labels.
- **Consequence**: The output file has the wrong number/order of rows, wrong or extra columns, and label values unrelated to the true target, so the grader scores it as wrong/missing regardless of the reported "100% accuracy."
165Template/output-format file never inspected before writing resultstaskda-code
Applies when
task -- the task says results must be saved to a specific output file whose format "matches the provided template" (or otherwise specifies an expected layout/schema).
Pattern
The agent never opens or parses the template/example file; it infers a plausible layout from domain convention (its own choice of row/column keys, index origin, header names, rounding, empty-cell handling, extra filtering) and writes that, then "verifies" only against its own output rather than against the reference schema.
Detection procedure
  1. Read the task for any mention of a template, sample, or required output schema, and note the required file name/location.
  2. Search the scripts for any read of that template (e.g., a load of the template path) and any comparison of column names, column order, row labels, dtypes, or shape to it.
  3. Check the write step: are header labels, index labels/format, value scaling and rounding, and NaN/blank representation derived from the template or hard-coded by the agent's own reasoning?
  4. Check the final answer: does it assert conformance ("format matches the template") without evidence of a diff/shape comparison against the template?
Discriminator
A real violation is when no template is ever read and the layout is self-invented (or when the agent adds unstated choices like extra row filtering, 1-based indices, or ad-hoc rounding). Fine: the script loads the template (or explicitly enumerates its columns/rows) and asserts the output's headers, index, shape, and rounding align with it — even if computed values are then discussed at length.
Consequence
The grader's file comparison fails on schema (header names/order, row keys, cell rounding or blanks) or on values shifted by the invented indexing/filtering, so the expected result file is scored WRONG/MISSING despite a seemingly correct analysis.
id eb4d157a5802 · mined from da-code dacode-dm-csv-044@s3
raw text (what the judge reads)
### Template/output-format file never inspected before writing results
- **Applies when**: `task` -- the task says results must be saved to a specific output file whose format "matches the provided template" (or otherwise specifies an expected layout/schema).
- **Pattern**: The agent never opens or parses the template/example file; it infers a plausible layout from domain convention (its own choice of row/column keys, index origin, header names, rounding, empty-cell handling, extra filtering) and writes that, then "verifies" only against its own output rather than against the reference schema.
- **Detection procedure**:
  1. Read the task for any mention of a template, sample, or required output schema, and note the required file name/location.
  2. Search the scripts for any read of that template (e.g., a load of the template path) and any comparison of column names, column order, row labels, dtypes, or shape to it.
  3. Check the write step: are header labels, index labels/format, value scaling and rounding, and NaN/blank representation derived from the template or hard-coded by the agent's own reasoning?
  4. Check the final answer: does it assert conformance ("format matches the template") without evidence of a diff/shape comparison against the template?
- **Discriminator**: A real violation is when no template is ever read and the layout is self-invented (or when the agent adds unstated choices like extra row filtering, 1-based indices, or ad-hoc rounding). Fine: the script loads the template (or explicitly enumerates its columns/rows) and asserts the output's headers, index, shape, and rounding align with it — even if computed values are then discussed at length.
- **Consequence**: The grader's file comparison fails on schema (header names/order, row keys, cell rounding or blanks) or on values shifted by the invented indexing/filtering, so the expected result file is scored WRONG/MISSING despite a seemingly correct analysis.
166Unauthorized redefinition of "missing" / extra row filtering beyond the stated grouping ruletaskinfiagent-dabench
Applies when
task -- the task defines groups or subsets by a simple, explicit rule (e.g., "rows where column X is null vs. not null") and the scripts implement grouping/cleaning themselves.
Pattern
The agent silently broadens or narrows the stated rule — treating whitespace/empty strings/sentinel values as missing, casting columns to strings, or dropping rows with missing values in the measured column — so group sizes and the resulting statistics differ from the literal specification the grader used.
Detection procedure
  1. Read the task and write down the exact membership rule and any allowed preprocessing; note that no extra cleaning was authorized.
  2. Read the scripts and list every predicate/filter applied to build each group (isna, astype(str).str.strip()=='', dropna() on the value column, dtype coercions, index_col, row drops).
  3. Flag any predicate not stated in the task; check whether the final reported numbers come from the augmented version rather than the plain rule (compare the earlier plain-rule run's means to the final answer — if they differ, the extra filtering changed the result and was never justified).
  4. Confirm the answer reports the plain-rule statistic; if the agent iterated to a "cleaner" variant with different means, treat that as unvalidated deviation.
Discriminator
Extra cleaning is fine only if the task explicitly asks for it, or if the reviewer can verify it does not change group counts/statistics (identical means either way, e.g., no empty strings exist and the value column has no missing entries). It is a violation when the deviation demonstrably shifts counts or means and the agent never reconciled the two definitions or reported the literal one.
Consequence
Group means (and any downstream test statistic) are computed on a different subset than intended, so the reported means fail exact-value checks even though the test's significance conclusion may still look plausible.
id 4d2e5e31f08f · mined from infiagent-dabench dabench-297@s3
raw text (what the judge reads)
### Unauthorized redefinition of "missing" / extra row filtering beyond the stated grouping rule
- **Applies when**: `task` -- the task defines groups or subsets by a simple, explicit rule (e.g., "rows where column X is null vs. not null") and the scripts implement grouping/cleaning themselves.
- **Pattern**: The agent silently broadens or narrows the stated rule — treating whitespace/empty strings/sentinel values as missing, casting columns to strings, or dropping rows with missing values in the *measured* column — so group sizes and the resulting statistics differ from the literal specification the grader used.
- **Detection procedure**:
  1. Read the task and write down the exact membership rule and any allowed preprocessing; note that no extra cleaning was authorized.
  2. Read the scripts and list every predicate/filter applied to build each group (`isna`, `astype(str).str.strip()==''`, `dropna()` on the value column, dtype coercions, `index_col`, row drops).
  3. Flag any predicate not stated in the task; check whether the final reported numbers come from the augmented version rather than the plain rule (compare the earlier plain-rule run's means to the final answer — if they differ, the extra filtering changed the result and was never justified).
  4. Confirm the answer reports the plain-rule statistic; if the agent iterated to a "cleaner" variant with different means, treat that as unvalidated deviation.
- **Discriminator**: Extra cleaning is fine only if the task explicitly asks for it, or if the reviewer can verify it does not change group counts/statistics (identical means either way, e.g., no empty strings exist and the value column has no missing entries). It is a violation when the deviation demonstrably shifts counts or means and the agent never reconciled the two definitions or reported the literal one.
- **Consequence**: Group means (and any downstream test statistic) are computed on a different subset than intended, so the reported means fail exact-value checks even though the test's significance conclusion may still look plausible.
167Missing required deliverables / self-invented definitions instead of following the provided spectaskda-code
Applies when
task -- the task points to an accompanying specification (guidance/README/config) and/or implies a set of saved output artifacts (figure, serialized numbers, plot metadata) that the grader will check.
Pattern
The script implements the analyst's own assumptions (invented filters, thresholds, category groupings, ordering) rather than the definitions in the referenced spec, and saves only the one artifact explicitly named in the prompt text while omitting the other expected output files; the final answer narrates numbers that were never persisted in the required machine-checkable form.
Detection procedure
  1. From the task, list every artifact the harness could compare (image file, numeric array/tabular dump, plot-state JSON) and every rule the referenced spec is said to contain (grouping definitions, subset filters, ordering, colors, sizes).
  2. In the scripts, check whether the spec file is actually read/quoted anywhere; flag any threshold, subset filter, or category mapping that appears with no citation to the spec (e.g., docstrings that define categories by the author's own keyword heuristics).
  3. Enumerate the script's write/save calls and confirm one exists for each artifact from step 1; flag if only the figure is written and no numeric/plot-metadata dump is produced.
  4. Compare the answer text: if it reports counts/values that exist only in stdout and not in a saved file, or reports a subset-restricted result when the task asked for the full tally, flag it.
Discriminator
A real violation is inventing rules the spec supplies or leaving a required artifact unwritten; a look-alike that is fine is a script that reproduces the spec's stated rules verbatim (or reasonably infers unstated details) and saves every artifact, even if its formatting or file naming differs cosmetically.
Consequence
The grader reports the missing/unmatched files as WRONG/MISSING and any comparable numbers disagree because they were computed on a self-invented subset, yielding 0 passed checks despite a plausible-looking narrative answer.
id e73b3d356dbe · mined from da-code dacode-plot-pie-005@s3
raw text (what the judge reads)
### Missing required deliverables / self-invented definitions instead of following the provided spec
- **Applies when**: `task` -- the task points to an accompanying specification (guidance/README/config) and/or implies a set of saved output artifacts (figure, serialized numbers, plot metadata) that the grader will check.
- **Pattern**: The script implements the analyst's own assumptions (invented filters, thresholds, category groupings, ordering) rather than the definitions in the referenced spec, and saves only the one artifact explicitly named in the prompt text while omitting the other expected output files; the final answer narrates numbers that were never persisted in the required machine-checkable form.
- **Detection procedure**:
  1. From the task, list every artifact the harness could compare (image file, numeric array/tabular dump, plot-state JSON) and every rule the referenced spec is said to contain (grouping definitions, subset filters, ordering, colors, sizes).
  2. In the scripts, check whether the spec file is actually read/quoted anywhere; flag any threshold, subset filter, or category mapping that appears with no citation to the spec (e.g., docstrings that define categories by the author's own keyword heuristics).
  3. Enumerate the script's write/save calls and confirm one exists for each artifact from step 1; flag if only the figure is written and no numeric/plot-metadata dump is produced.
  4. Compare the answer text: if it reports counts/values that exist only in stdout and not in a saved file, or reports a subset-restricted result when the task asked for the full tally, flag it.
- **Discriminator**: A real violation is inventing rules the spec supplies or leaving a required artifact unwritten; a look-alike that is fine is a script that reproduces the spec's stated rules verbatim (or reasonably infers unstated details) and saves every artifact, even if its formatting or file naming differs cosmetically.
- **Consequence**: The grader reports the missing/unmatched files as WRONG/MISSING and any comparable numbers disagree because they were computed on a self-invented subset, yielding 0 passed checks despite a plausible-looking narrative answer.
168Unsanity-checked error magnitude for a regression on a column-cleaned target/featurestaskinfiagent-dabench
Applies when
task -- the task asks for a regression error metric (MSE/RMSE/MAE) on a numeric target after prescribed preprocessing (e.g., mean-imputation of specific columns) and a fixed train/test split.
Pattern
The attempt runs a pipeline whose preprocessing silently deviates from the spec — columns left as strings/objects and coerced or dropped, rows dropped instead of mean-imputed, imputation done only on some of the named columns or fit on the wrong subset, or units/scale of a predictor left unnormalized — and then reports the resulting error without ever comparing it to the target's own scale (variance / mean-predictor baseline) or to the retained row counts.
Detection procedure
  1. From the task, list exactly which columns must be imputed/used and the split/metric definition; note the target's plausible range from the data description or a quick describe().
  2. In the scripts, verify each named column is numeric before imputation (check for dropna, astype/to_numeric(errors='coerce'), string-stripping, or an implicit drop by the model), verify all named columns are mean-imputed, and verify row count after preprocessing equals the original row count.
  3. Check that the reported metric is accompanied by a baseline sanity check: compare MSE to the variance of the target (or to a mean-only predictor's MSE) and confirm √MSE is small relative to the target's spread.
  4. If no script is saved / no such check is shown, treat the number as unverified and reject.
Discriminator
A genuine violation shows either (a) preprocessing that changes the row count or silently excludes a required column, or (b) an error whose square root is comparable to or larger than the target's standard deviation with no explanation; a look-alike that is fine is a spec-compliant pipeline that happens to have high error but demonstrably retains all rows, imputes all named columns, and reports an error below the mean-baseline MSE.
Consequence
The reported metric can be an order of magnitude off the reference value (worse than predicting the mean), so the single numeric answer fails the exact-value grader.
id 17c022cc5f99 · mined from infiagent-dabench dabench-432@s3
raw text (what the judge reads)
### Unsanity-checked error magnitude for a regression on a column-cleaned target/features
- **Applies when**: `task` -- the task asks for a regression error metric (MSE/RMSE/MAE) on a numeric target after prescribed preprocessing (e.g., mean-imputation of specific columns) and a fixed train/test split.
- **Pattern**: The attempt runs a pipeline whose preprocessing silently deviates from the spec — columns left as strings/objects and coerced or dropped, rows dropped instead of mean-imputed, imputation done only on some of the named columns or fit on the wrong subset, or units/scale of a predictor left unnormalized — and then reports the resulting error without ever comparing it to the target's own scale (variance / mean-predictor baseline) or to the retained row counts.
- **Detection procedure**:
  1. From the task, list exactly which columns must be imputed/used and the split/metric definition; note the target's plausible range from the data description or a quick describe().
  2. In the scripts, verify each named column is numeric before imputation (check for `dropna`, `astype`/`to_numeric(errors='coerce')`, string-stripping, or an implicit drop by the model), verify all named columns are mean-imputed, and verify row count after preprocessing equals the original row count.
  3. Check that the reported metric is accompanied by a baseline sanity check: compare MSE to the variance of the target (or to a mean-only predictor's MSE) and confirm √MSE is small relative to the target's spread.
  4. If no script is saved / no such check is shown, treat the number as unverified and reject.
- **Discriminator**: A genuine violation shows either (a) preprocessing that changes the row count or silently excludes a required column, or (b) an error whose square root is comparable to or larger than the target's standard deviation with no explanation; a look-alike that is fine is a spec-compliant pipeline that happens to have high error but demonstrably retains all rows, imputes all named columns, and reports an error below the mean-baseline MSE.
- **Consequence**: The reported metric can be an order of magnitude off the reference value (worse than predicting the mean), so the single numeric answer fails the exact-value grader.
169Reordering time-ordered rows on an unverified assumptiontaskinfiagent-dabench
Applies when
task -- the computation depends on row order (lag/diff/shift, cumulative, rolling, or sequential splits) and the script sorts, reverses, or otherwise reorders the rows before computing.
Pattern
The script asserts a direction for the existing order (e.g., "data is newest-first, so reverse it") in a comment and flips or re-sorts rows without ever parsing the timestamp/index column and checking monotonicity; the sign or magnitude of the order-dependent statistic then silently depends on a guess.
Detection procedure
  1. From the task, note whether the requested quantity is order-dependent (defined relative to a "previous"/"next" record or accumulated over the sequence).
  2. In the scripts, find every reorder operation (iloc[::-1], sort_values, reset_index) and check whether it is justified by an explicit, executed test — e.g., parsing the time column to datetime and verifying it is increasing/decreasing, or sorting by that column rather than by position.
  3. Check whether any post-hoc validation would catch a wrong direction: comparing the first/last timestamps, or recomputing the statistic on both orderings and reconciling with an external sanity check (e.g., overall start-to-end change consistent with the sign of the mean).
  4. Compare the reported answer's sign/magnitude to what the raw endpoints of the series imply; an unexplained sign flip signals the assumed order was wrong.
Discriminator
Fine if the script sorts explicitly by a parsed date/index column (direction then cannot be wrong) or prints/asserts the order before reordering; a violation if the direction is inferred from a printed head, a comment, or intuition, with no check — especially when the result's sign hinges on it.
Consequence
Order-dependent statistics come out with reversed sign (mean of differences/returns) and slightly different dispersion, so both reported values fail the exact-match check even though the formula and rounding are correct.
id ec41dcbe4b25 · mined from infiagent-dabench dabench-75@s3
raw text (what the judge reads)
### Reordering time-ordered rows on an unverified assumption
- **Applies when**: `task` -- the computation depends on row order (lag/diff/shift, cumulative, rolling, or sequential splits) and the script sorts, reverses, or otherwise reorders the rows before computing.
- **Pattern**: The script asserts a direction for the existing order (e.g., "data is newest-first, so reverse it") in a comment and flips or re-sorts rows without ever parsing the timestamp/index column and checking monotonicity; the sign or magnitude of the order-dependent statistic then silently depends on a guess.
- **Detection procedure**:
  1. From the task, note whether the requested quantity is order-dependent (defined relative to a "previous"/"next" record or accumulated over the sequence).
  2. In the scripts, find every reorder operation (`iloc[::-1]`, `sort_values`, `reset_index`) and check whether it is justified by an explicit, executed test — e.g., parsing the time column to datetime and verifying it is increasing/decreasing, or sorting *by* that column rather than by position.
  3. Check whether any post-hoc validation would catch a wrong direction: comparing the first/last timestamps, or recomputing the statistic on both orderings and reconciling with an external sanity check (e.g., overall start-to-end change consistent with the sign of the mean).
  4. Compare the reported answer's sign/magnitude to what the raw endpoints of the series imply; an unexplained sign flip signals the assumed order was wrong.
- **Discriminator**: Fine if the script sorts explicitly by a parsed date/index column (direction then cannot be wrong) or prints/asserts the order before reordering; a violation if the direction is inferred from a printed head, a comment, or intuition, with no check — especially when the result's sign hinges on it.
- **Consequence**: Order-dependent statistics come out with reversed sign (mean of differences/returns) and slightly different dispersion, so both reported values fail the exact-match check even though the formula and rounding are correct.
170Unvalidated row-inclusion decisions when computing a summary statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single statistic (correlation, mean, coefficient, metric) over two or more columns of a table, and the script makes implicit choices about which rows enter the computation (index/header parsing, NaN dropping, dtype coercion, sentinel/placeholder values, duplicate or aggregate rows).
Pattern
The script loads the file with unexamined parsing options, applies a one-shot filter (e.g., drop rows where either column is non-finite, or coerce to numeric) and immediately reports the resulting statistic, with no check that the analysed row set equals the intended population and no test of how sensitive the statistic is to that filtering choice. A small difference in included rows shifts the value enough to break the required rounding precision.
Detection procedure
  1. Read the task for any stated population/filtering rule; if none is stated, the default is "all records in the file as delivered".
  2. In the script, list every step that can change the row count: header/index parsing, type conversion, dropna/masking, deduplication, reindexing — and check whether the script prints and compares the pre-filter vs post-filter counts against the file's raw record count.
  3. Check whether the script inspects the columns for non-numeric or placeholder encodings (blank, "NA", 0, -1, -9, "."), which would either wrongly enter or wrongly leave the computation, and whether it recomputes the statistic under the alternative treatment to show the reported value is stable.
  4. Compare the reported number's precision (e.g., two decimals) with the observed sensitivity: if no stability check exists and the statistic sits near a rounding boundary, flag it.
Discriminator
Fine if the script explicitly reports row counts before/after any filtering, confirms they match the intended population (or a rule stated in the task), and shows the statistic is unchanged under plausible alternative handling of ambiguous rows. A violation is when filtering is implicit/unjustified and the single reported value was never cross-checked against another row-inclusion or computation path.
Consequence
The statistic is computed on a slightly different subset than the reference, so the rounded value differs by one or more units in the last reported digit and the numeric check fails even though the qualitative conclusion (significant/not) happens to pass.
id fab917fd88c5 · mined from infiagent-dabench dabench-300@s3
raw text (what the judge reads)
### Unvalidated row-inclusion decisions when computing a summary statistic
- **Applies when**: `task` -- the task asks for a single statistic (correlation, mean, coefficient, metric) over two or more columns of a table, and the script makes implicit choices about which rows enter the computation (index/header parsing, NaN dropping, dtype coercion, sentinel/placeholder values, duplicate or aggregate rows).
- **Pattern**: The script loads the file with unexamined parsing options, applies a one-shot filter (e.g., drop rows where either column is non-finite, or coerce to numeric) and immediately reports the resulting statistic, with no check that the analysed row set equals the intended population and no test of how sensitive the statistic is to that filtering choice. A small difference in included rows shifts the value enough to break the required rounding precision.
- **Detection procedure**:
  1. Read the task for any stated population/filtering rule; if none is stated, the default is "all records in the file as delivered".
  2. In the script, list every step that can change the row count: header/index parsing, type conversion, dropna/masking, deduplication, reindexing — and check whether the script prints and compares the pre-filter vs post-filter counts against the file's raw record count.
  3. Check whether the script inspects the columns for non-numeric or placeholder encodings (blank, "NA", 0, -1, -9, "."), which would either wrongly enter or wrongly leave the computation, and whether it recomputes the statistic under the alternative treatment to show the reported value is stable.
  4. Compare the reported number's precision (e.g., two decimals) with the observed sensitivity: if no stability check exists and the statistic sits near a rounding boundary, flag it.
- **Discriminator**: Fine if the script explicitly reports row counts before/after any filtering, confirms they match the intended population (or a rule stated in the task), and shows the statistic is unchanged under plausible alternative handling of ambiguous rows. A violation is when filtering is implicit/unjustified and the single reported value was never cross-checked against another row-inclusion or computation path.
- **Consequence**: The statistic is computed on a slightly different subset than the reference, so the rounded value differs by one or more units in the last reported digit and the numeric check fails even though the qualitative conclusion (significant/not) happens to pass.
171No held-out validation metric — prediction quality asserted only from output summary statisticstaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test set that will be graded against hidden ground truth by an error metric.
Pattern
The attempt trains a single model with hand-picked hyperparameters, writes the prediction file, and reports only descriptive statistics of the predictions (mean/median/min/max/std, row count, dtype) plus a feature list — with no cross-validation or hold-out error estimate, no baseline comparison, and no check that the predicted distribution matches the training target distribution.
Detection procedure
  1. Read the task to confirm the deliverable is scored on accuracy against unseen labels, not merely on file existence/shape.
  2. Search the scripts for any train/validation split, cross-validation, or scoring call producing an error metric (RMSE/MAE/R²/log-error) on labeled data held out from fitting; note whether at least one alternative (simpler baseline or different hyperparameters/target transform) was compared.
  3. Compare the reported prediction statistics against the same statistics of the training target: check mean, median, spread, and min/max ranges for gross mismatch (e.g. predicted spread or extremes far outside plausible label range, or median far from training median).
  4. Confirm the answer states a quantified generalization estimate; if the only evidence of quality is "predictions look like numbers in a plausible range", flag it.
Discriminator
A real violation is the absence of any labeled-data error estimate before submission. It is not a violation if the agent did compute a hold-out/CV score (even for one model) and the reported summary just omits hyperparameter tuning — nor if the metric is present but modest; the issue is unverified quality, not suboptimality. Also not a violation when the task genuinely has no labels for validation.
Consequence
The submitted prediction column is graded against ground truth and fails the error threshold (file marked WRONG), with no diagnostic in the attempt that could have caught the miscalibrated or under-fit model beforehand.
id da4ab5a4bf4e · mined from da-code dacode-ml-regression-014@s3
raw text (what the judge reads)
### No held-out validation metric — prediction quality asserted only from output summary statistics
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test set that will be graded against hidden ground truth by an error metric.
- **Pattern**: The attempt trains a single model with hand-picked hyperparameters, writes the prediction file, and reports only descriptive statistics of the predictions (mean/median/min/max/std, row count, dtype) plus a feature list — with no cross-validation or hold-out error estimate, no baseline comparison, and no check that the predicted distribution matches the training target distribution.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on accuracy against unseen labels, not merely on file existence/shape.
  2. Search the scripts for any train/validation split, cross-validation, or scoring call producing an error metric (RMSE/MAE/R²/log-error) on labeled data held out from fitting; note whether at least one alternative (simpler baseline or different hyperparameters/target transform) was compared.
  3. Compare the reported prediction statistics against the same statistics of the training target: check mean, median, spread, and min/max ranges for gross mismatch (e.g. predicted spread or extremes far outside plausible label range, or median far from training median).
  4. Confirm the answer states a quantified generalization estimate; if the only evidence of quality is "predictions look like numbers in a plausible range", flag it.
- **Discriminator**: A real violation is the absence of *any* labeled-data error estimate before submission. It is not a violation if the agent did compute a hold-out/CV score (even for one model) and the reported summary just omits hyperparameter tuning — nor if the metric is present but modest; the issue is unverified quality, not suboptimality. Also not a violation when the task genuinely has no labels for validation.
- **Consequence**: The submitted prediction column is graded against ground truth and fails the error threshold (file marked WRONG), with no diagnostic in the attempt that could have caught the miscalibrated or under-fit model beforehand.
172Prediction file row count / alignment with the evaluation set is never verifiedtaskda-code
Applies when
task -- The deliverable is a per-record output file (predictions, labels, scores) that must correspond one-to-one, in order, with the rows of a supplied evaluation/test input.
Pattern
The agent produces an output file (or pastes an answer) whose number of rows is not demonstrably equal to the number of input rows — e.g. predictions are generated from a subsample, a filtered/deduplicated frame, a chunk that failed to concatenate, a truncated print, or a differently-ordered frame — and no script step asserts len(predictions) == len(test_input) or preserves the original row order/ID join.
Detection procedure
  1. From the task/README, find the expected number of evaluation records (read the input file's shape, or the stated size) and the required column name(s)/order.
  2. In the scripts, trace the object that is written out: check it derives from the full, unfiltered, unshuffled evaluation frame (no head(), sample(), dropna(), partial loop, or merge that can drop/reorder rows) and that an explicit shape/count assertion exists before writing.
  3. Count the rows actually present in the submitted file/answer and compare to step 1; also confirm the header and label vocabulary match the specification exactly.
  4. If scripts are missing entirely, treat the answer as unverifiable and check row count directly against the input.
Discriminator
A real violation is a mismatch in row count, order, or header/label values relative to the evaluation input; it is not a violation if the count matches exactly and only the predicted label values are debatable (accuracy issues are a separate concern), nor if a documented ID column is included that lets the grader realign rows.
Consequence
The grader cannot align predictions with ground truth, so the submission is scored as wrong/missing regardless of model quality (0 checks passed).
id 0f0c7b8bb08f · mined from da-code dacode-ml-multi-008@s3
raw text (what the judge reads)
### Prediction file row count / alignment with the evaluation set is never verified
- **Applies when**: `task` -- The deliverable is a per-record output file (predictions, labels, scores) that must correspond one-to-one, in order, with the rows of a supplied evaluation/test input.
- **Pattern**: The agent produces an output file (or pastes an answer) whose number of rows is not demonstrably equal to the number of input rows — e.g. predictions are generated from a subsample, a filtered/deduplicated frame, a chunk that failed to concatenate, a truncated print, or a differently-ordered frame — and no script step asserts `len(predictions) == len(test_input)` or preserves the original row order/ID join.
- **Detection procedure**:
  1. From the task/README, find the expected number of evaluation records (read the input file's shape, or the stated size) and the required column name(s)/order.
  2. In the scripts, trace the object that is written out: check it derives from the full, unfiltered, unshuffled evaluation frame (no `head()`, `sample()`, `dropna()`, partial loop, or merge that can drop/reorder rows) and that an explicit shape/count assertion exists before writing.
  3. Count the rows actually present in the submitted file/answer and compare to step 1; also confirm the header and label vocabulary match the specification exactly.
  4. If scripts are missing entirely, treat the answer as unverifiable and check row count directly against the input.
- **Discriminator**: A real violation is a mismatch in row count, order, or header/label values relative to the evaluation input; it is *not* a violation if the count matches exactly and only the predicted label values are debatable (accuracy issues are a separate concern), nor if a documented ID column is included that lets the grader realign rows.
- **Consequence**: The grader cannot align predictions with ground truth, so the submission is scored as wrong/missing regardless of model quality (0 checks passed).
173Output artifacts written outside the task's working/data directorytaskinfiagent-dabench
Applies when
task -- the deliverable includes file paths to artifacts (CSVs, models, plots) that the agent must create and report.
Pattern
The agent writes outputs to an arbitrary location (e.g., its own home/temp/current directory) instead of the directory where the input data lives or the directory the task designates, then reports that path; the numeric content may be right but the path does not match what is expected.
Detection procedure
1. Read the task and note where the input file(s) are read from and any stated output directory or naming convention. 2. In the scripts, find every write call (to_csv, savefig, dump) and record the exact output path. 3. Compare that path's directory to the input data directory / stated convention; check the filename too. 4. Confirm the path string in the final answer is identical to the path actually written.
Discriminator
A real violation is an output directory that differs from the input/data directory or from an explicitly stated location without the task permitting a free choice; it is fine if the task explicitly leaves the location open, or if the agent writes into the same directory as the source data (the default safe choice).
Consequence
The path field fails string comparison against the expected artifact location, marking the whole item wrong even when the computed statistics are correct.
id 8c0185ec629d · mined from infiagent-dabench dabench-743@s3
raw text (what the judge reads)
### Output artifacts written outside the task's working/data directory
- **Applies when**: `task` -- the deliverable includes file paths to artifacts (CSVs, models, plots) that the agent must create and report.
- **Pattern**: The agent writes outputs to an arbitrary location (e.g., its own home/temp/current directory) instead of the directory where the input data lives or the directory the task designates, then reports that path; the numeric content may be right but the path does not match what is expected.
- **Detection procedure**: 1. Read the task and note where the input file(s) are read from and any stated output directory or naming convention. 2. In the scripts, find every write call (`to_csv`, `savefig`, `dump`) and record the exact output path. 3. Compare that path's directory to the input data directory / stated convention; check the filename too. 4. Confirm the path string in the final answer is identical to the path actually written.
- **Discriminator**: A real violation is an output directory that differs from the input/data directory or from an explicitly stated location without the task permitting a free choice; it is fine if the task explicitly leaves the location open, or if the agent writes into the same directory as the source data (the default safe choice).
- **Consequence**: The path field fails string comparison against the expected artifact location, marking the whole item wrong even when the computed statistics are correct.
174Deliverable artifacts not produced (answer reported as text only)taskda-code
Applies when
task -- the task asks for output files (plots, saved arrays, config-driven figures) in addition to, or instead of, a value stated in the reply.
Pattern
The agent computes an intermediate selection/statistic and reports it in prose, but the scripts never write the requested files (or write them under different names/paths/formats than specified), and any styling spec file referenced by the task is never read or applied.
Detection procedure
  1. Enumerate every artifact the task requires: exact filenames, extensions, directories, plus any spec/config file whose parameters must be honored.
  2. Search the scripts for a write call for each artifact (e.g. figure save, array/JSON dump) and for a load/parse of the spec file; confirm the saved paths match the required names exactly.
  3. Check the final answer: does it stop at the intermediate quantity (a chosen group, a filtered subset) instead of, or without, the requested deliverables?
  4. If scripts are missing/unsaved entirely, treat the reproducibility of every artifact as unverified.
Discriminator
A real violation is a missing/renamed/unstyled artifact or a text-only answer; it is fine if all required files are written with the exact names and the spec's parameters (labels, ordering, colors, sizes) are demonstrably applied, and the prose value is merely supplementary.
Consequence
File-level checks report each expected artifact as WRONG/MISSING, scoring 0 even if the intermediate value quoted in the answer happens to be right.
id f9dca85af370 · mined from da-code dacode-plot-pie-008@s3
raw text (what the judge reads)
### Deliverable artifacts not produced (answer reported as text only)
- **Applies when**: `task` -- the task asks for output files (plots, saved arrays, config-driven figures) in addition to, or instead of, a value stated in the reply.
- **Pattern**: The agent computes an intermediate selection/statistic and reports it in prose, but the scripts never write the requested files (or write them under different names/paths/formats than specified), and any styling spec file referenced by the task is never read or applied.
- **Detection procedure**:
  1. Enumerate every artifact the task requires: exact filenames, extensions, directories, plus any spec/config file whose parameters must be honored.
  2. Search the scripts for a write call for each artifact (e.g. figure save, array/JSON dump) and for a load/parse of the spec file; confirm the saved paths match the required names exactly.
  3. Check the final answer: does it stop at the intermediate quantity (a chosen group, a filtered subset) instead of, or without, the requested deliverables?
  4. If scripts are missing/unsaved entirely, treat the reproducibility of every artifact as unverified.
- **Discriminator**: A real violation is a missing/renamed/unstyled artifact or a text-only answer; it is fine if all required files are written with the exact names and the spec's parameters (labels, ordering, colors, sizes) are demonstrably applied, and the prose value is merely supplementary.
- **Consequence**: File-level checks report each expected artifact as WRONG/MISSING, scoring 0 even if the intermediate value quoted in the answer happens to be right.
175Unjustified row-dropping / undocumented preprocessing choices that silently change the evaluated sampletaskinfiagent-dabench
Applies when
task -- the task fixes the model, feature list, split fraction and random seed and asks for a single metric, but some feature columns contain missing values or the model's raw output must be converted to class labels.
Pattern
The script silently makes a free preprocessing choice that changes which rows are modeled or how outputs become labels — e.g. dropna() on the feature/target frame before splitting (shrinking the sample and shifting the seeded split), or an ad-hoc/implicit label threshold — and reports the resulting metric with no record of row counts or the decision rule, so the number cannot be reproduced against the canonical pipeline.
Detection procedure
  1. From the task, list what is pinned (features, split ratio, seed, metric) and note what is not pinned (missing-value strategy, continuous→label rule, rounding).
  2. In the script, find the step order: check whether any rows are removed/filtered before the seeded split, and whether the row count fed to train_test_split equals the full dataset row count; check that all requested features (including ones with NaNs) survive.
  3. Check that any regressor-to-class conversion is explicit (e.g. threshold at 0.5) and applied only to test predictions, and that the metric is computed on the held-out subset only.
  4. Confirm the script prints shapes/row counts before and after preprocessing and, since the unpinned choice is decision-relevant, at least a second variant (impute vs. drop) so the reported figure is shown to be robust; otherwise flag.
Discriminator
A real violation is dropping/altering rows (or leaving the labeling rule implicit) when an imputation-preserving pipeline over all rows was equally available and no counts are reported. It is fine if the dataset genuinely has no missing values in the requested features (row count unchanged and verified), or if the task itself dictates the filtering.
Consequence
The seeded 80/20 split covers a different, smaller population than the canonical one, so the accuracy is off by a couple of hundredths and the exact-match check on the rounded metric fails (e.g. reporting 0.76 where 0.78 is expected).
id d1c214e7db1c · mined from infiagent-dabench dabench-7@s3
raw text (what the judge reads)
### Unjustified row-dropping / undocumented preprocessing choices that silently change the evaluated sample
- **Applies when**: `task` -- the task fixes the model, feature list, split fraction and random seed and asks for a single metric, but some feature columns contain missing values or the model's raw output must be converted to class labels.
- **Pattern**: The script silently makes a free preprocessing choice that changes *which rows* are modeled or *how outputs become labels* — e.g. `dropna()` on the feature/target frame before splitting (shrinking the sample and shifting the seeded split), or an ad-hoc/implicit label threshold — and reports the resulting metric with no record of row counts or the decision rule, so the number cannot be reproduced against the canonical pipeline.
- **Detection procedure**:
  1. From the task, list what is pinned (features, split ratio, seed, metric) and note what is *not* pinned (missing-value strategy, continuous→label rule, rounding).
  2. In the script, find the step order: check whether any rows are removed/filtered before the seeded split, and whether the row count fed to `train_test_split` equals the full dataset row count; check that all requested features (including ones with NaNs) survive.
  3. Check that any regressor-to-class conversion is explicit (e.g. threshold at 0.5) and applied only to test predictions, and that the metric is computed on the held-out subset only.
  4. Confirm the script prints shapes/row counts before and after preprocessing and, since the unpinned choice is decision-relevant, at least a second variant (impute vs. drop) so the reported figure is shown to be robust; otherwise flag.
- **Discriminator**: A real violation is dropping/altering rows (or leaving the labeling rule implicit) when an imputation-preserving pipeline over all rows was equally available and no counts are reported. It is fine if the dataset genuinely has no missing values in the requested features (row count unchanged and verified), or if the task itself dictates the filtering.
- **Consequence**: The seeded 80/20 split covers a different, smaller population than the canonical one, so the accuracy is off by a couple of hundredths and the exact-match check on the rounded metric fails (e.g. reporting 0.76 where 0.78 is expected).
176Deliverable template not filled in on its own termstaskda-code
Applies when
task -- the task supplies a specific output file/template that must be populated and returned, and the agent instead (or additionally) narrates findings in prose.
Pattern
The agent invents its own categories, column names, and ordering, reports numbers in the chat/markdown summary, and never writes (or overwrites with a mismatched schema) the provided result file; the template's existing headers/row labels are ignored, so the categories used for grouping don't line up with the ones the file expects.
Detection procedure
1. Read the task for the named output artifact and open/inspect its header row and any pre-filled label rows to see the exact expected schema and category vocabulary. 2. Search the scripts for code that reads that file and writes it back with the same columns/row labels and row count. 3. Compare the group names/order in the agent's answer against the template's labels — any extra bucket (e.g., an "unknown/other" catch-all) or renamed/missing label is a mismatch. 4. Confirm the final answer points to a saved file rather than only prose numbers.
Discriminator
A real violation is a missing file, or a file whose columns/labels/row set differ from the template; it is fine if the agent derives categories in code as long as the written file's schema and label set match the template exactly and any auxiliary prose is redundant with it.
Consequence
The grader compares the expected file cell-by-cell and marks it WRONG/MISSING, scoring 0 regardless of whether the underlying computation was reasonable.
id 7d7d832768f0 · mined from da-code dacode-dm-csv-001@s3
raw text (what the judge reads)
### Deliverable template not filled in on its own terms
- **Applies when**: `task` -- the task supplies a specific output file/template that must be populated and returned, and the agent instead (or additionally) narrates findings in prose.
- **Pattern**: The agent invents its own categories, column names, and ordering, reports numbers in the chat/markdown summary, and never writes (or overwrites with a mismatched schema) the provided result file; the template's existing headers/row labels are ignored, so the categories used for grouping don't line up with the ones the file expects.
- **Detection procedure**: 1. Read the task for the named output artifact and open/inspect its header row and any pre-filled label rows to see the exact expected schema and category vocabulary. 2. Search the scripts for code that reads that file and writes it back with the same columns/row labels and row count. 3. Compare the group names/order in the agent's answer against the template's labels — any extra bucket (e.g., an "unknown/other" catch-all) or renamed/missing label is a mismatch. 4. Confirm the final answer points to a saved file rather than only prose numbers.
- **Discriminator**: A real violation is a missing file, or a file whose columns/labels/row set differ from the template; it is fine if the agent derives categories in code as long as the written file's schema and label set match the template exactly and any auxiliary prose is redundant with it.
- **Consequence**: The grader compares the expected file cell-by-cell and marks it WRONG/MISSING, scoring 0 regardless of whether the underlying computation was reasonable.
177Entity-level aggregation not established before computing a per-entity statistictaskinfiagent-dabench
Applies when
task -- the task asks for a correlation/statistic between two per-entity attributes (e.g., an entity's maximum of some ordinal field and its span of activity) while the source table stores multiple observation rows per entity, and a subgroup split (e.g., above/below median of a third attribute) must also be defined per entity.
Pattern
The attempt computes the statistic (and the median split) over raw observation rows, or aggregates with an inconsistent key/definition (duration from row counts instead of last-minus-first timestamp, category as row value instead of per-entity max, damage summed vs. maxed), so the n and the resulting r drift from the intended per-entity values; no script is retained to make the aggregation auditable.
Detection procedure
  1. From the task wording, identify the analysis unit ("per storm/patient/user/order") and list the exact aggregations required for each variable plus the variable used for the subgroup split.
  2. In the scripts, confirm there is an explicit groupby(entity_id) producing one row per entity, and check each aggregation function matches the definition (max vs. mean, time-span vs. record count, and units of duration); confirm the median threshold is computed on the aggregated frame, not raw rows.
  3. Check that the reported n (or a printed df.shape after aggregation) equals the number of distinct entities and that the two subgroups partition it, and that missing/duplicate rows were handled before aggregating.
  4. If no script or printed intermediate counts exist, treat the numeric answer as unverified and reject.
Discriminator
A real violation is a statistic computed at the wrong grain or with a mis-specified aggregation/threshold (wrong n, wrong r at the 2nd decimal). A look-alike that is fine is a correctly grouped analysis whose r differs only by defensible tie-handling in the median split or by rounding — evidenced by matching entity counts and documented aggregation rules.
Consequence
The qualitative verdict (e.g., "linear") may still match, but the reported coefficient misses the expected value at the required rounding precision, and the answer is graded wrong on the numeric checks.
id faa655d59a6d · mined from infiagent-dabench dabench-431@s3
raw text (what the judge reads)
### Entity-level aggregation not established before computing a per-entity statistic
- **Applies when**: `task` -- the task asks for a correlation/statistic between two per-entity attributes (e.g., an entity's maximum of some ordinal field and its span of activity) while the source table stores multiple observation rows per entity, and a subgroup split (e.g., above/below median of a third attribute) must also be defined per entity.
- **Pattern**: The attempt computes the statistic (and the median split) over raw observation rows, or aggregates with an inconsistent key/definition (duration from row counts instead of last-minus-first timestamp, category as row value instead of per-entity max, damage summed vs. maxed), so the n and the resulting r drift from the intended per-entity values; no script is retained to make the aggregation auditable.
- **Detection procedure**:
  1. From the task wording, identify the analysis unit ("per storm/patient/user/order") and list the exact aggregations required for each variable plus the variable used for the subgroup split.
  2. In the scripts, confirm there is an explicit `groupby(entity_id)` producing one row per entity, and check each aggregation function matches the definition (max vs. mean, time-span vs. record count, and units of duration); confirm the median threshold is computed on the aggregated frame, not raw rows.
  3. Check that the reported n (or a printed `df.shape` after aggregation) equals the number of distinct entities and that the two subgroups partition it, and that missing/duplicate rows were handled before aggregating.
  4. If no script or printed intermediate counts exist, treat the numeric answer as unverified and reject.
- **Discriminator**: A real violation is a statistic computed at the wrong grain or with a mis-specified aggregation/threshold (wrong n, wrong r at the 2nd decimal). A look-alike that is fine is a correctly grouped analysis whose r differs only by defensible tie-handling in the median split or by rounding — evidenced by matching entity counts and documented aggregation rules.
- **Consequence**: The qualitative verdict (e.g., "linear") may still match, but the reported coefficient misses the expected value at the required rounding precision, and the answer is graded wrong on the numeric checks.
178Baseline model built on an arbitrary feature subset instead of all available predictorstaskinfiagent-dabench
Applies when
task -- the task asks whether adding an engineered feature improves a model, without explicitly listing which columns the baseline model should use.
Pattern
The script silently narrows the baseline design matrix to only the few columns involved in the engineered feature (or another hand-picked subset), discarding the rest of the dataset's predictors, and then compares the "with new feature" model against this weakened baseline. The comparison is internally consistent but both reported metrics are far from the values obtained with the natural "all original columns" baseline.
Detection procedure
  1. Read the task and list which predictors it names or implies; note whether it restricts the feature set at all.
  2. In the script, find where X is constructed for the baseline and for the augmented model; compare the column list to the full set of columns in the loaded table (minus the target).
  3. If columns present in the data are dropped without any instruction or stated justification (e.g., non-numeric columns dropped instead of encoded, or only the feature-engineering inputs kept), flag it.
  4. Check that the augmented model equals the baseline features plus exactly the new feature, and that both use the same split.
Discriminator
A real violation is dropping usable predictors that the task never excluded (including categorical columns that could have been encoded), making the baseline arbitrarily weak. It is not a violation if the task explicitly names the predictors, or if a column is dropped for a defensible, stated reason (identifier column, leakage, exact duplicate of the target).
Consequence
Both RMSE values are inflated relative to the reference solution (which uses the full predictor set), so the correlation check passes but both model-metric checks fail, even though the qualitative "improves/does not improve" conclusion may look plausible.
id b05af7b0f64c · mined from infiagent-dabench dabench-549@s3
raw text (what the judge reads)
### Baseline model built on an arbitrary feature subset instead of all available predictors
- **Applies when**: `task` -- the task asks whether adding an engineered feature improves a model, without explicitly listing which columns the baseline model should use.
- **Pattern**: The script silently narrows the baseline design matrix to only the few columns involved in the engineered feature (or another hand-picked subset), discarding the rest of the dataset's predictors, and then compares the "with new feature" model against this weakened baseline. The comparison is internally consistent but both reported metrics are far from the values obtained with the natural "all original columns" baseline.
- **Detection procedure**:
  1. Read the task and list which predictors it names or implies; note whether it restricts the feature set at all.
  2. In the script, find where `X` is constructed for the baseline and for the augmented model; compare the column list to the full set of columns in the loaded table (minus the target).
  3. If columns present in the data are dropped without any instruction or stated justification (e.g., non-numeric columns dropped instead of encoded, or only the feature-engineering inputs kept), flag it.
  4. Check that the augmented model equals the baseline features plus exactly the new feature, and that both use the same split.
- **Discriminator**: A real violation is dropping usable predictors that the task never excluded (including categorical columns that could have been encoded), making the baseline arbitrarily weak. It is not a violation if the task explicitly names the predictors, or if a column is dropped for a defensible, stated reason (identifier column, leakage, exact duplicate of the target).
- **Consequence**: Both RMSE values are inflated relative to the reference solution (which uses the full predictor set), so the correlation check passes but both model-metric checks fail, even though the qualitative "improves/does not improve" conclusion may look plausible.
179Unverified prediction artifact: no held-out validation and no re-read check of the exact output spectaskda-code
Applies when
task -- the deliverable is a prediction/result file with an explicitly stated filename, column name(s), and row coverage, produced by a model the agent trains itself.
Pattern
The attempt trains a model and writes the file, then declares success based only on self-reported counts (rows, predicted class balance, number of estimators), without (a) reloading the written file and asserting it matches the requested schema exactly — required column present and named verbatim, no unrequested extra columns/index, one row per test record in the original test order, valid label values/encoding — and (b) any held-out or cross-validated score showing the model beats a trivial baseline. Often no script is retained at all, so neither claim is reproducible.
Detection procedure
  1. From the task text, list the literal output requirements: filename, exact column name(s), expected number of rows, and label type/encoding implied by the training target.
  2. In the scripts, find where the file is written; check whether the frame written contains exactly the requested columns (watch for index=True, an added identifier column, probabilities instead of labels, or strings instead of the training label's dtype) and whether row count/order is tied to the test file.
  3. Check for a post-write verification step (re-read the file, assert shape/columns/dtypes/value set) and for a validation metric computed on data not used for fitting.
  4. Compare the final answer's claims to what the scripts actually assert; treat unsupported claims ("format is as requested", "model used X") as unverified.
Discriminator
A real violation is when nothing in the scripts or answer proves the on-disk file matches the stated schema and the model was ever scored out-of-sample; it is not a violation if the agent adds extra columns that the task explicitly permits, or if it verifies schema and reports a holdout metric but simply achieves modest accuracy.
Consequence
The grader reads the file with the expected schema and either fails to find the required column/row alignment or scores the predictions below the accepted threshold, marking the deliverable WRONG/MISSING despite a confident "task completed" report.
id 8c27f4bb6173 · mined from da-code dacode-ml-binary-016@s3
raw text (what the judge reads)
### Unverified prediction artifact: no held-out validation and no re-read check of the exact output spec
- **Applies when**: `task` -- the deliverable is a prediction/result file with an explicitly stated filename, column name(s), and row coverage, produced by a model the agent trains itself.
- **Pattern**: The attempt trains a model and writes the file, then declares success based only on self-reported counts (rows, predicted class balance, number of estimators), without (a) reloading the written file and asserting it matches the requested schema exactly — required column present and named verbatim, no unrequested extra columns/index, one row per test record in the original test order, valid label values/encoding — and (b) any held-out or cross-validated score showing the model beats a trivial baseline. Often no script is retained at all, so neither claim is reproducible.
- **Detection procedure**:
  1. From the task text, list the literal output requirements: filename, exact column name(s), expected number of rows, and label type/encoding implied by the training target.
  2. In the scripts, find where the file is written; check whether the frame written contains exactly the requested columns (watch for `index=True`, an added identifier column, probabilities instead of labels, or strings instead of the training label's dtype) and whether row count/order is tied to the test file.
  3. Check for a post-write verification step (re-read the file, assert shape/columns/dtypes/value set) and for a validation metric computed on data not used for fitting.
  4. Compare the final answer's claims to what the scripts actually assert; treat unsupported claims ("format is as requested", "model used X") as unverified.
- **Discriminator**: A real violation is when nothing in the scripts or answer proves the on-disk file matches the stated schema and the model was ever scored out-of-sample; it is *not* a violation if the agent adds extra columns that the task explicitly permits, or if it verifies schema and reports a holdout metric but simply achieves modest accuracy.
- **Consequence**: The grader reads the file with the expected schema and either fails to find the required column/row alignment or scores the predictions below the accepted threshold, marking the deliverable WRONG/MISSING despite a confident "task completed" report.
180Method substituted for the one implied by the stated constraints/contexttaskda-code
Applies when
task -- the task asks for a statistic/p-value and imposes a constraint (e.g., "set the random seed", a named output template, a specific column set) that only makes sense for a particular class of method.
Pattern
The attempt computes the quantity with a convenient closed-form/library default (e.g., an analytic parametric test or built-in shortcut) that never consumes the seed and never reproduces the template's expected fields, silently ignoring the constraint that signalled a resampling/simulation-based (or otherwise specified) procedure. The reported value is then a different estimator than the one the grader checks.
Detection procedure
  1. In the task text, list every explicit constraint (seed, rounding, units, filtering, ordering, output filename/columns) and ask which methods each constraint presupposes; open the provided sample/template file and note its exact fields and value conventions.
  2. In the scripts, check whether each constraint is actually exercised: is the seed passed to a random generator that affects the reported number? Are the output columns/labels taken from the template rather than invented?
  3. If a constraint is inert (seed set but no randomness used; template ignored), treat the chosen method as suspect and identify the method that would use it (e.g., permutation/bootstrap replicates with a defined test statistic and a stated number of draws).
  4. Compare the reported value/label against what that method would yield and against the template's expected labels.
Discriminator
A real violation is when the constraint has no effect on the computed result or the output schema deviates from the supplied template; it is fine if the method genuinely uses the constraint (seed drives the replicates, rounding/units applied) even if a different-but-equivalent implementation was possible, or if the constraint is legitimately irrelevant and the answer format still matches the template exactly.
Consequence
The grader compares against the constraint-consistent result and schema, so the p-value/statistic differs beyond tolerance (or the label/column values mismatch) and the file is marked wrong despite the script running without error.
id b02746c37fff · mined from da-code dacode-data-sa-039@s3
raw text (what the judge reads)
### Method substituted for the one implied by the stated constraints/context
- **Applies when**: `task` -- the task asks for a statistic/p-value and imposes a constraint (e.g., "set the random seed", a named output template, a specific column set) that only makes sense for a particular class of method.
- **Pattern**: The attempt computes the quantity with a convenient closed-form/library default (e.g., an analytic parametric test or built-in shortcut) that never consumes the seed and never reproduces the template's expected fields, silently ignoring the constraint that signalled a resampling/simulation-based (or otherwise specified) procedure. The reported value is then a different estimator than the one the grader checks.
- **Detection procedure**:
  1. In the task text, list every explicit constraint (seed, rounding, units, filtering, ordering, output filename/columns) and ask which methods each constraint presupposes; open the provided sample/template file and note its exact fields and value conventions.
  2. In the scripts, check whether each constraint is actually exercised: is the seed passed to a random generator that affects the reported number? Are the output columns/labels taken from the template rather than invented?
  3. If a constraint is inert (seed set but no randomness used; template ignored), treat the chosen method as suspect and identify the method that would use it (e.g., permutation/bootstrap replicates with a defined test statistic and a stated number of draws).
  4. Compare the reported value/label against what that method would yield and against the template's expected labels.
- **Discriminator**: A real violation is when the constraint has no effect on the computed result or the output schema deviates from the supplied template; it is fine if the method genuinely uses the constraint (seed drives the replicates, rounding/units applied) even if a different-but-equivalent implementation was possible, or if the constraint is legitimately irrelevant and the answer format still matches the template exactly.
- **Consequence**: The grader compares against the constraint-consistent result and schema, so the p-value/statistic differs beyond tolerance (or the label/column values mismatch) and the file is marked wrong despite the script running without error.
181Silent numeric coercion of formatted strings, then imputing the corrupted cellstaskda-code
Applies when
task -- a script converts a text/object column to numeric (e.g. pd.to_numeric(..., errors='coerce'), astype(float)) before computing extremes, aggregates or imputations.
Pattern
The attempt never compares the number of nulls before and after coercion, so entries carrying formatting characters (thousand separators, currency/percent signs, units, spaces, footnote marks) become NaN; those NaNs are then filled with the mean, which silently replaces the true extreme values with a middling number and shifts the mean itself.
Detection procedure
  1. In the task/README, note whether the target column is numeric-looking but stored as text, and whether values can be large enough to be written with separators or suffixes.
  2. In the scripts, find the conversion step and check whether the raw strings are cleaned (regex/str.replace of non-numeric characters) and whether the null count is asserted/printed before vs. after the conversion, and before any imputation.
  3. Check whether the imputation is applied only to genuinely missing cells, and whether the reported extreme rows are sanity-checked against known magnitudes/plausible range for that quantity.
  4. Inspect the reported answer: if the maximum equals or is close to the imputed mean, or is implausibly small for the domain, the extremes were destroyed.
Discriminator
Fine if the script cleans formatting first, or explicitly verifies that the post-coercion NaN count equals the pre-existing missing count (or lists the coerced rows and confirms they are truly blank); a violation is blind errors='coerce' followed by mean-fill with no such check.
Consequence
The extreme-value answer names the wrong record (the true max/min was overwritten by the imputed mean), so the graded key mismatches the expected result and the check fails.
id 034828d9fa90 · mined from da-code dacode-di-text-001@s3
raw text (what the judge reads)
### Silent numeric coercion of formatted strings, then imputing the corrupted cells
- **Applies when**: `task` -- a script converts a text/object column to numeric (e.g. `pd.to_numeric(..., errors='coerce')`, `astype(float)`) before computing extremes, aggregates or imputations.
- **Pattern**: The attempt never compares the number of nulls before and after coercion, so entries carrying formatting characters (thousand separators, currency/percent signs, units, spaces, footnote marks) become NaN; those NaNs are then filled with the mean, which silently replaces the true extreme values with a middling number and shifts the mean itself.
- **Detection procedure**:
  1. In the task/README, note whether the target column is numeric-looking but stored as text, and whether values can be large enough to be written with separators or suffixes.
  2. In the scripts, find the conversion step and check whether the raw strings are cleaned (regex/`str.replace` of non-numeric characters) and whether the null count is asserted/printed *before* vs. *after* the conversion, and *before* any imputation.
  3. Check whether the imputation is applied only to genuinely missing cells, and whether the reported extreme rows are sanity-checked against known magnitudes/plausible range for that quantity.
  4. Inspect the reported answer: if the maximum equals or is close to the imputed mean, or is implausibly small for the domain, the extremes were destroyed.
- **Discriminator**: Fine if the script cleans formatting first, or explicitly verifies that the post-coercion NaN count equals the pre-existing missing count (or lists the coerced rows and confirms they are truly blank); a violation is blind `errors='coerce'` followed by mean-fill with no such check.
- **Consequence**: The extreme-value answer names the wrong record (the true max/min was overwritten by the imputed mean), so the graded key mismatches the expected result and the check fails.
182Failing to sanity-check subset sizes and the orientation/definition of a derived quantitytaskda-code
Applies when
task -- the task asks for a statistic computed on filtered subgroups of a dataset and/or on a derived quantity built from two columns (ratio, difference, rate).
Pattern
The attempt loads/filters the data with an unverified assumption (wrong file, wrong grouping key, header row misparsed, silent drop of rows) and/or builds the derived quantity with the operands swapped or the wrong definition, then reports the statistic without ever comparing record counts or the value's plausible range to what the data/domain implies.
Detection procedure
  1. From the task and data README, note what each subgroup should roughly contain (how many records, which files/columns supply them) and what the requested derived quantity means, including which term is the numerator/units.
  2. In the scripts, trace the filtering/merge/parse steps and the formula for the derived quantity; check whether any count, shape, or dtype assertion or printout is used to confirm the intended rows were selected.
  3. In the answer, compare the reported group sizes and statistic values against those expectations — e.g. subgroup n far smaller than the source file's rows, or a ratio on the wrong side of 1 given the two quantities' typical magnitudes.
  4. Flag the attempt if no such validation exists and the reported counts/values are implausible or unexplained.
Discriminator
A genuine violation shows numbers that contradict the data source (e.g. tens of records where the file clearly holds hundreds, or an inverted ratio) with no verification step; a look-alike is fine if the script explicitly asserts/prints row counts and shapes that match the source and the derived quantity's definition matches the task wording, even if the subgroup is legitimately small.
Consequence
Every downstream number (point estimate and bootstrap interval) is computed on the wrong rows or the reciprocal quantity, so the written output file fails exact/tolerance comparison against the expected result and the task scores 0.
id 28d51aa01651 · mined from da-code dacode-data-sa-029@s3
raw text (what the judge reads)
### Failing to sanity-check subset sizes and the orientation/definition of a derived quantity
- **Applies when**: `task` -- the task asks for a statistic computed on filtered subgroups of a dataset and/or on a derived quantity built from two columns (ratio, difference, rate).
- **Pattern**: The attempt loads/filters the data with an unverified assumption (wrong file, wrong grouping key, header row misparsed, silent drop of rows) and/or builds the derived quantity with the operands swapped or the wrong definition, then reports the statistic without ever comparing record counts or the value's plausible range to what the data/domain implies.
- **Detection procedure**:
  1. From the task and data README, note what each subgroup should roughly contain (how many records, which files/columns supply them) and what the requested derived quantity means, including which term is the numerator/units.
  2. In the scripts, trace the filtering/merge/parse steps and the formula for the derived quantity; check whether any count, shape, or dtype assertion or printout is used to confirm the intended rows were selected.
  3. In the answer, compare the reported group sizes and statistic values against those expectations — e.g. subgroup n far smaller than the source file's rows, or a ratio on the wrong side of 1 given the two quantities' typical magnitudes.
  4. Flag the attempt if no such validation exists and the reported counts/values are implausible or unexplained.
- **Discriminator**: A genuine violation shows numbers that contradict the data source (e.g. tens of records where the file clearly holds hundreds, or an inverted ratio) with no verification step; a look-alike is fine if the script explicitly asserts/prints row counts and shapes that match the source and the derived quantity's definition matches the task wording, even if the subgroup is legitimately small.
- **Consequence**: Every downstream number (point estimate and bootstrap interval) is computed on the wrong rows or the reciprocal quantity, so the written output file fails exact/tolerance comparison against the expected result and the task scores 0.
183Analysis run on a truncated/placeholder slice of the data, with no row-count or plausibility sanity checktaskda-code
Applies when
task -- the task asks for a statistic, matrix, or model fit computed over an entire dataset file, and the scripts load that file before computing and saving the result.
Pattern
The attempt silently computes on a partial, sampled, hand-built, or otherwise truncated version of the data (e.g. a round "n = 100" rows, a head()/nrows= read, a manually retyped snippet, or a re-created toy frame) and reports the resulting numbers as if they came from the full source, without ever cross-checking the loaded row count against the real file or checking that the numbers are plausible.
Detection procedure
  1. Read the task/README and note what the full data source is; find the exact load call in the scripts and check for any row limit, sampling, slicing, or literal in-script data construction.
  2. Compare the row/column counts the answer reports (before and after cleaning) with the actual file's size — suspiciously round or small totals, or counts never verified anywhere in the script, are red flags.
  3. Check that the reported statistic passes a domain/range sanity check (e.g. quantities expected to move together showing near-zero or sign-flipped relationships, metrics near chance, shapes/labels not matching the requested output).
  4. Confirm the saved artifact was actually written from the computed object with the requested name, orientation, index/header, and rounding — and that the answer's numbers match that file rather than being pasted prose.
Discriminator
A real violation is when the computed population differs from the one the task specifies (truncated read, fabricated frame, unverified counts) or the values are implausible and unchecked; it is fine if the script reads the whole file and the row reduction is fully explained by the required dropping/filtering step, with counts printed and the output file written from the computed object.
Consequence
The saved file contains numbers that differ from the ground-truth computation on the full data (here, near-zero correlations instead of the true strong positive ones), so the file-comparison check fails even though the answer's narrative looks complete.
id 141d54b10d39 · mined from da-code dacode-data-sa-026@s3
raw text (what the judge reads)
### Analysis run on a truncated/placeholder slice of the data, with no row-count or plausibility sanity check
- **Applies when**: `task` -- the task asks for a statistic, matrix, or model fit computed over an entire dataset file, and the scripts load that file before computing and saving the result.
- **Pattern**: The attempt silently computes on a partial, sampled, hand-built, or otherwise truncated version of the data (e.g. a round "n = 100" rows, a `head()`/`nrows=` read, a manually retyped snippet, or a re-created toy frame) and reports the resulting numbers as if they came from the full source, without ever cross-checking the loaded row count against the real file or checking that the numbers are plausible.
- **Detection procedure**:
  1. Read the task/README and note what the full data source is; find the exact load call in the scripts and check for any row limit, sampling, slicing, or literal in-script data construction.
  2. Compare the row/column counts the answer reports (before and after cleaning) with the actual file's size — suspiciously round or small totals, or counts never verified anywhere in the script, are red flags.
  3. Check that the reported statistic passes a domain/range sanity check (e.g. quantities expected to move together showing near-zero or sign-flipped relationships, metrics near chance, shapes/labels not matching the requested output).
  4. Confirm the saved artifact was actually written from the computed object with the requested name, orientation, index/header, and rounding — and that the answer's numbers match that file rather than being pasted prose.
- **Discriminator**: A real violation is when the computed population differs from the one the task specifies (truncated read, fabricated frame, unverified counts) or the values are implausible and unchecked; it is fine if the script reads the whole file and the row reduction is fully explained by the required dropping/filtering step, with counts printed and the output file written from the computed object.
- **Consequence**: The saved file contains numbers that differ from the ground-truth computation on the full data (here, near-zero correlations instead of the true strong positive ones), so the file-comparison check fails even though the answer's narrative looks complete.
184Selecting a model-order hyperparameter at the boundary of the searched range without validationtaskda-code
Applies when
task -- the script chooses a structural hyperparameter (number of clusters/components/factors) by scanning a candidate range and picking the extreme of an internal quality score, especially when the data carries a natural or documented group count.
Pattern
The attempt sweeps a hard-coded range, reports the best score at the last (or first) candidate, and accepts that value as "optimal" even though the score curve is still rising at the edge and all scores are near zero (no real structure), while never cross-checking against an alternative criterion (elbow, gap, stability, known number of categories in the data description/label column) or against the effect of preprocessing on mixed binary/continuous features.
Detection procedure
  1. Read the task and data description for any implied or documented number of natural groups (e.g., a categorical label or class column that was dropped before clustering).
  2. In the script, locate the candidate range and how the "optimal" value is chosen; check whether the winner is the range endpoint and whether more than one selection criterion is consulted.
  3. In the answer, inspect the reported score table: flag if the metric increases monotonically to the last candidate, if absolute values are near 0 (weak/no separation), or if the chosen count contradicts the documented group count.
  4. Check whether preprocessing is justified for the feature mix (e.g., standardizing one-hot/indicator columns so they dominate distance) and whether any sanity check of cluster sizes/labels against the known grouping was done.
Discriminator
A genuine violation is an endpoint-optimum plus no corroborating criterion and no reconciliation with a documented group count; it is fine if the score curve has an interior peak or plateau, or the script explicitly extends the range/uses a second criterion (elbow, stability, ARI vs. known labels) and explains why the chosen count differs from the documented one.
Consequence
The saved label column encodes an over- or under-partitioned solution, so comparison against the reference clustering (cluster count, agreement metrics such as ARI/NMI, or cluster-size profile) fails and the output file is marked wrong.
id 08f3e6c03d13 · mined from da-code dacode-ml-cluster-010@s3
raw text (what the judge reads)
### Selecting a model-order hyperparameter at the boundary of the searched range without validation
- **Applies when**: `task` -- the script chooses a structural hyperparameter (number of clusters/components/factors) by scanning a candidate range and picking the extreme of an internal quality score, especially when the data carries a natural or documented group count.
- **Pattern**: The attempt sweeps a hard-coded range, reports the best score at the last (or first) candidate, and accepts that value as "optimal" even though the score curve is still rising at the edge and all scores are near zero (no real structure), while never cross-checking against an alternative criterion (elbow, gap, stability, known number of categories in the data description/label column) or against the effect of preprocessing on mixed binary/continuous features.
- **Detection procedure**:
  1. Read the task and data description for any implied or documented number of natural groups (e.g., a categorical label or class column that was dropped before clustering).
  2. In the script, locate the candidate range and how the "optimal" value is chosen; check whether the winner is the range endpoint and whether more than one selection criterion is consulted.
  3. In the answer, inspect the reported score table: flag if the metric increases monotonically to the last candidate, if absolute values are near 0 (weak/no separation), or if the chosen count contradicts the documented group count.
  4. Check whether preprocessing is justified for the feature mix (e.g., standardizing one-hot/indicator columns so they dominate distance) and whether any sanity check of cluster sizes/labels against the known grouping was done.
- **Discriminator**: A genuine violation is an endpoint-optimum plus no corroborating criterion and no reconciliation with a documented group count; it is fine if the score curve has an interior peak or plateau, or the script explicitly extends the range/uses a second criterion (elbow, stability, ARI vs. known labels) and explains why the chosen count differs from the documented one.
- **Consequence**: The saved label column encodes an over- or under-partitioned solution, so comparison against the reference clustering (cluster count, agreement metrics such as ARI/NMI, or cluster-size profile) fails and the output file is marked wrong.
185Requested output artifact is claimed but never verified (path, schema, contents)taskda-code
Applies when
task -- the task requires saving results to a named file with a prescribed column naming/format, and the agent's answer describes the file instead of demonstrating it.
Pattern
The attempt runs an analysis and asserts the deliverable was written (with a path, row/column counts and column names quoted from intent, not from disk), but no step re-opens the saved file to confirm it exists in the expected location, has exactly the required column names, the right number of rows, and content consistent with the described analysis; often the file is written to a non-collected directory, with transformed/intermediate values, or the write step silently never ran because no persisted script exists.
Detection procedure
  1. From the task, list the exact deliverable(s): filename, required column names/order, expected granularity (one row per entity) and expected value semantics.
  2. In the scripts, locate the write call: check the directory it targets (relative to the expected submission/working dir), the columns actually assigned, and whether the values written are the requested ones rather than an intermediate representation.
  3. Look for a post-write verification step (re-read the file, print shape, columns, head, label value counts) and check the printed output is echoed in the answer; absence of such evidence — or of any saved script at all — is the flag.
  4. Cross-check numbers in the answer (row count, number of clusters/labels, column count) against the verification printout; mismatch or "no printout" means unverified.
Discriminator
A real violation is a claim about the file backed only by narrative description or by pre-write in-memory variables; a look-alike that is fine re-reads the written file from the required path and shows its actual shape, exact required column names, and label distribution matching the reported result.
Consequence
The grader reports the expected result file as WRONG/MISSING (absent from the collected path, wrong headers, or contents inconsistent with the report), scoring 0 regardless of the analysis quality.
id 5a848e8d8178 · mined from da-code dacode-ml-cluster-019@s3
raw text (what the judge reads)
### Requested output artifact is claimed but never verified (path, schema, contents)
- **Applies when**: `task` -- the task requires saving results to a named file with a prescribed column naming/format, and the agent's answer describes the file instead of demonstrating it.
- **Pattern**: The attempt runs an analysis and asserts the deliverable was written (with a path, row/column counts and column names quoted from intent, not from disk), but no step re-opens the saved file to confirm it exists in the expected location, has exactly the required column names, the right number of rows, and content consistent with the described analysis; often the file is written to a non-collected directory, with transformed/intermediate values, or the write step silently never ran because no persisted script exists.
- **Detection procedure**:
  1. From the task, list the exact deliverable(s): filename, required column names/order, expected granularity (one row per entity) and expected value semantics.
  2. In the scripts, locate the write call: check the directory it targets (relative to the expected submission/working dir), the columns actually assigned, and whether the values written are the requested ones rather than an intermediate representation.
  3. Look for a post-write verification step (re-read the file, print `shape`, `columns`, `head`, label value counts) and check the printed output is echoed in the answer; absence of such evidence — or of any saved script at all — is the flag.
  4. Cross-check numbers in the answer (row count, number of clusters/labels, column count) against the verification printout; mismatch or "no printout" means unverified.
- **Discriminator**: A real violation is a claim about the file backed only by narrative description or by pre-write in-memory variables; a look-alike that is fine re-reads the written file from the required path and shows its actual shape, exact required column names, and label distribution matching the reported result.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING (absent from the collected path, wrong headers, or contents inconsistent with the report), scoring 0 regardless of the analysis quality.
186Requested output artifact never written in the specified formattaskda-code
Applies when
task -- the task instructs the agent to record its findings in a named output file (e.g., a CSV) with a given format/template.
Pattern
The agent performs the computation and reports the numbers only in its chat/prose summary, never creating (or overwriting) the named file, or writing it with different columns/row labels/units/precision than the provided template requires.
Detection procedure
1. Read the task and note the exact filename, expected columns/rows, units, and rounding of the required artifact. 2. Search the scripts for a write call targeting that exact path (to_csv, open(...,'w'), etc.) and confirm it runs unconditionally at the end of the pipeline. 3. Check the written object's schema against the template: same column names/order, same number of rows, values in the same units/scale (e.g., proportion vs. percent) and rounding. 4. Confirm the final answer's numbers match what the file would contain, rather than existing only in prose.
Discriminator
A real violation is missing file creation, a wrong path/filename, or a schema/unit mismatch with the stated template; it is fine if the file is written correctly and the prose is merely an additional human-readable summary.
Consequence
The grader looks for the named file and marks it WRONG/MISSING, scoring 0 regardless of whether the underlying statistic was computed correctly.
id 35fe9aa318c7 · mined from da-code dacode-data-sa-031@s3
raw text (what the judge reads)
### Requested output artifact never written in the specified format
- **Applies when**: `task` -- the task instructs the agent to record its findings in a named output file (e.g., a CSV) with a given format/template.
- **Pattern**: The agent performs the computation and reports the numbers only in its chat/prose summary, never creating (or overwriting) the named file, or writing it with different columns/row labels/units/precision than the provided template requires.
- **Detection procedure**: 1. Read the task and note the exact filename, expected columns/rows, units, and rounding of the required artifact. 2. Search the scripts for a write call targeting that exact path (`to_csv`, `open(...,'w')`, etc.) and confirm it runs unconditionally at the end of the pipeline. 3. Check the written object's schema against the template: same column names/order, same number of rows, values in the same units/scale (e.g., proportion vs. percent) and rounding. 4. Confirm the final answer's numbers match what the file would contain, rather than existing only in prose.
- **Discriminator**: A real violation is missing file creation, a wrong path/filename, or a schema/unit mismatch with the stated template; it is fine if the file is written correctly and the prose is merely an additional human-readable summary.
- **Consequence**: The grader looks for the named file and marks it WRONG/MISSING, scoring 0 regardless of whether the underlying statistic was computed correctly.
187Never scoring the submission against the competition's own metric on held-out datataskda-code
Applies when
task -- the task states an explicit evaluation metric (e.g., a loss/score on probabilistic or numeric predictions) and the scripts build a model, produce cross-validated/held-out predictions, and write a prediction file.
Pattern
The attempt sets up folds and even stores out-of-fold predictions, but never computes the stated metric on them, never compares it to a trivial baseline (e.g., class-prior/mean predictor), and never tunes or calibrates anything; verification is limited to file shape, column names, and value ranges. A single un-tuned, overconfident model is shipped on faith.
Detection procedure
  1. Read the task and note the exact metric and how it treats predictions (e.g., log/exponential penalties on extreme probabilities, rescaling rules).
  2. Grep the scripts for the metric being computed on validation/OOF predictions (log_loss, roc_auc, RMSE, etc.) and for any baseline or alternative-configuration comparison; note if OOF arrays are computed but unused.
  3. Read the verification script: check whether it validates predictive quality or only format/shape/range.
  4. Inspect the answer's prediction distribution for signs of uncalibrated overconfidence (many values near 0 or 1, or near-zero probabilities for a minority class) that the metric would punish heavily.
Discriminator
A fine attempt prints at least one held-out metric value computed with the task's definition, contrasts it against a baseline or another setting, and only then writes the file; a violation is when no quality number exists anywhere in the scripts or report — format checks alone. Not a violation if the metric is computed and reported and simply happens to be modest, or if the model is chosen after documented comparison.
Consequence
The submitted file is well-formed but scores worse than a simple baseline (extreme, uncalibrated probabilities blow up the log-type loss), so the graded result fails the accuracy threshold even though all shape/format sanity checks passed.
id 19876a658632 · mined from da-code dacode-ml-competition-005@s4
raw text (what the judge reads)
### Never scoring the submission against the competition's own metric on held-out data
- **Applies when**: `task` -- the task states an explicit evaluation metric (e.g., a loss/score on probabilistic or numeric predictions) and the scripts build a model, produce cross-validated/held-out predictions, and write a prediction file.
- **Pattern**: The attempt sets up folds and even stores out-of-fold predictions, but never computes the stated metric on them, never compares it to a trivial baseline (e.g., class-prior/mean predictor), and never tunes or calibrates anything; verification is limited to file shape, column names, and value ranges. A single un-tuned, overconfident model is shipped on faith.
- **Detection procedure**:
  1. Read the task and note the exact metric and how it treats predictions (e.g., log/exponential penalties on extreme probabilities, rescaling rules).
  2. Grep the scripts for the metric being computed on validation/OOF predictions (`log_loss`, `roc_auc`, RMSE, etc.) and for any baseline or alternative-configuration comparison; note if OOF arrays are computed but unused.
  3. Read the verification script: check whether it validates predictive quality or only format/shape/range.
  4. Inspect the answer's prediction distribution for signs of uncalibrated overconfidence (many values near 0 or 1, or near-zero probabilities for a minority class) that the metric would punish heavily.
- **Discriminator**: A fine attempt prints at least one held-out metric value computed with the task's definition, contrasts it against a baseline or another setting, and only then writes the file; a violation is when no quality number exists anywhere in the scripts or report — format checks alone. Not a violation if the metric is computed and reported and simply happens to be modest, or if the model is chosen after documented comparison.
- **Consequence**: The submitted file is well-formed but scores worse than a simple baseline (extreme, uncalibrated probabilities blow up the log-type loss), so the graded result fails the accuracy threshold even though all shape/format sanity checks passed.
188Silent column intersection hides misnamed/missing features (no schema assertion)taskda-code
Applies when
task -- a script builds a feature list by hand and then keeps only the columns that happen to exist in both the training and the scoring table (or wraps every column access in if col in df.columns) before fitting a model and writing predictions.
Pattern
The hard-coded feature names differ from the actual headers (singular/plural, spelling variants, different naming conventions between files), so the defensive intersection/if in columns guard quietly drops most informative predictors and also silently skips the intended type coercions and missing-value handling for them. The model still trains, prints a high validation score on whatever few columns survived, and produces an output file whose values (and often whose schema/row count) do not match the requested target — nothing in the log signals that features were lost.
Detection procedure
  1. From the task/README, list the documented field names and the required output schema (exact column name(s), one row per scoring record, any ordering/rounding/index constraints).
  2. In the script, find the declared feature/target names and every guard of the form "only use columns present in both" or if col in df.columns; compare each declared name character-by-character with the documented headers, and check whether the script prints/asserts the final feature count and the raw headers of each loaded file.
  3. Check the write step: does it emit exactly the requested column name(s) and the expected number of rows, and is any index/extra id column suppressed or explicitly required?
  4. In the answer, compare the claimed feature list against the features that would actually survive the guard, and look for a reported validation score dominated by a single surviving column.
Discriminator
Fine if the script asserts the expected columns exist (raising or logging on mismatch) and the names provably match the real headers, or if a deliberate, stated reason exists for excluding a column; a violation is when the guard is the only mechanism handling name mismatches, so a typo degrades the model or the output schema without any error, and no post-write shape/column check is performed.
Consequence
The prediction file is built from a truncated feature set and/or a schema the grader does not expect, so the output comparison fails (wrong values, extra/renamed columns, or wrong row count) despite an impressive-looking reported R²/RMSE.
id f6373bce20a4 · mined from da-code dacode-ml-regression-008@s4
raw text (what the judge reads)
### Silent column intersection hides misnamed/missing features (no schema assertion)
- **Applies when**: `task` -- a script builds a feature list by hand and then keeps only the columns that happen to exist in both the training and the scoring table (or wraps every column access in `if col in df.columns`) before fitting a model and writing predictions.
- **Pattern**: The hard-coded feature names differ from the actual headers (singular/plural, spelling variants, different naming conventions between files), so the defensive intersection/`if in columns` guard quietly drops most informative predictors and also silently skips the intended type coercions and missing-value handling for them. The model still trains, prints a high validation score on whatever few columns survived, and produces an output file whose values (and often whose schema/row count) do not match the requested target — nothing in the log signals that features were lost.
- **Detection procedure**:
  1. From the task/README, list the documented field names and the required output schema (exact column name(s), one row per scoring record, any ordering/rounding/index constraints).
  2. In the script, find the declared feature/target names and every guard of the form "only use columns present in both" or `if col in df.columns`; compare each declared name character-by-character with the documented headers, and check whether the script prints/asserts the final feature count and the raw headers of each loaded file.
  3. Check the write step: does it emit exactly the requested column name(s) and the expected number of rows, and is any index/extra id column suppressed or explicitly required?
  4. In the answer, compare the claimed feature list against the features that would actually survive the guard, and look for a reported validation score dominated by a single surviving column.
- **Discriminator**: Fine if the script asserts the expected columns exist (raising or logging on mismatch) and the names provably match the real headers, or if a deliberate, stated reason exists for excluding a column; a violation is when the guard is the only mechanism handling name mismatches, so a typo degrades the model or the output schema without any error, and no post-write shape/column check is performed.
- **Consequence**: The prediction file is built from a truncated feature set and/or a schema the grader does not expect, so the output comparison fails (wrong values, extra/renamed columns, or wrong row count) despite an impressive-looking reported R²/RMSE.
189Statistical test run on the full raw dataset with an unjustified default test formtaskda-code
Applies when
task -- a task asks for a hypothesis test / p-value on a comparison between two groups, and the scripts load the raw files and immediately test all rows with a default off-the-shelf function.
Pattern
The script never restricts the data to the population implied by the question (date window, competition/category, entity subset, deduplication), and it picks the test form (parametric vs. rank-based, one- vs. two-tailed, paired vs. independent, equal-variance vs. Welch) by default rather than by checking the distribution shape and the direction stated or implied in the hypothesis — producing a p-value from the wrong sample and/or the wrong test.
Detection procedure
  1. Read the task/README for any scoping words (time period, event type, category, "official", "recent", subgroup) and for directional wording ("greater than", "more than") in the hypothesis; list the subset and tail this implies.
  2. Read the script: check whether any filtering/subsetting occurs before the test, and whether the sample size after filtering is printed and matches the implied scope.
  3. Check whether the script inspects the outcome distribution (histogram, skew, normality) or justifies the chosen test function and its alternative/variance arguments; a bare two-sided independent t-test on skewed count data is a red flag.
  4. Compare the reported p-value magnitude to plausibility: an extreme value (e.g., <1e-20) usually signals a huge unfiltered sample rather than the intended analysis.
Discriminator
Not a violation if the task genuinely asks about all records and the script documents that no scoping applies, or if it shows an assumption check that supports the chosen test (and gives the same decision under a robust alternative). It is a violation when scope-defining language or distributional evidence exists and the script ignores both.
Consequence
The p-value differs from the reference by orders of magnitude (and the reject/fail-to-reject string may still coincidentally match), so the exact-value check on the output file fails.
id bc696ed93a39 · mined from da-code dacode-data-sa-001@s4
raw text (what the judge reads)
### Statistical test run on the full raw dataset with an unjustified default test form
- **Applies when**: `task` -- a task asks for a hypothesis test / p-value on a comparison between two groups, and the scripts load the raw files and immediately test all rows with a default off-the-shelf function.
- **Pattern**: The script never restricts the data to the population implied by the question (date window, competition/category, entity subset, deduplication), and it picks the test form (parametric vs. rank-based, one- vs. two-tailed, paired vs. independent, equal-variance vs. Welch) by default rather than by checking the distribution shape and the direction stated or implied in the hypothesis — producing a p-value from the wrong sample and/or the wrong test.
- **Detection procedure**:
  1. Read the task/README for any scoping words (time period, event type, category, "official", "recent", subgroup) and for directional wording ("greater than", "more than") in the hypothesis; list the subset and tail this implies.
  2. Read the script: check whether any filtering/subsetting occurs before the test, and whether the sample size after filtering is printed and matches the implied scope.
  3. Check whether the script inspects the outcome distribution (histogram, skew, normality) or justifies the chosen test function and its `alternative`/variance arguments; a bare two-sided independent t-test on skewed count data is a red flag.
  4. Compare the reported p-value magnitude to plausibility: an extreme value (e.g., <1e-20) usually signals a huge unfiltered sample rather than the intended analysis.
- **Discriminator**: Not a violation if the task genuinely asks about all records and the script documents that no scoping applies, or if it shows an assumption check that supports the chosen test (and gives the same decision under a robust alternative). It *is* a violation when scope-defining language or distributional evidence exists and the script ignores both.
- **Consequence**: The p-value differs from the reference by orders of magnitude (and the reject/fail-to-reject string may still coincidentally match), so the exact-value check on the output file fails.
190Output not validated against the provided template/sample file (precision, rounding, ordering, headers)taskda-code
Applies when
task -- the task supplies a sample/reference output file and says results must follow its exact structure and formatting, and the scripts write a results file.
Pattern
The attempt computes aggregates and dumps raw floating-point values straight from the dataframe, hard-coding or guessing the row order and column layout instead of reading the sample file and matching its number formatting (decimal places/rounding), row ordering, and header text; no programmatic comparison of the produced file to the template is done.
Detection procedure
  1. Read the task for the formatting constraint and note that a sample output file is provided.
  2. In the scripts, check whether the sample file is actually parsed and its properties (column names/order, row order key, number of decimals or value scale in each numeric column) are extracted and applied — as opposed to being printed once, ignored, or replaced by a hand-typed list.
  3. Inspect the final written values: raw artifacts such as long float tails (e.g. ...7799999999), inconsistent decimal counts across rows, or an ordering rule not derived from the sample indicate no formatting step was applied.
  4. Confirm the script contains no final check comparing the written file's shape/columns/row-keys/format against the sample before submitting.
Discriminator
A real violation is when the output's formatting or ordering is chosen by the agent rather than derived from the template (unrounded floats, guessed sort order, renamed/reordered columns). It is not a violation if the script reads the template, applies its rounding/precision and ordering, and the residual differences are only in values that the template itself leaves free.
Consequence
The grader does an exact/tolerance-limited comparison against the reference file and marks the result file WRONG even when the underlying aggregation logic may be close or correct.
id 8bd262a5e965 · mined from da-code dacode-dm-csv-011@s4
raw text (what the judge reads)
### Output not validated against the provided template/sample file (precision, rounding, ordering, headers)
- **Applies when**: `task` -- the task supplies a sample/reference output file and says results must follow its exact structure and formatting, and the scripts write a results file.
- **Pattern**: The attempt computes aggregates and dumps raw floating-point values straight from the dataframe, hard-coding or guessing the row order and column layout instead of reading the sample file and matching its number formatting (decimal places/rounding), row ordering, and header text; no programmatic comparison of the produced file to the template is done.
- **Detection procedure**:
  1. Read the task for the formatting constraint and note that a sample output file is provided.
  2. In the scripts, check whether the sample file is actually parsed and its properties (column names/order, row order key, number of decimals or value scale in each numeric column) are extracted and applied — as opposed to being printed once, ignored, or replaced by a hand-typed list.
  3. Inspect the final written values: raw artifacts such as long float tails (e.g. `...7799999999`), inconsistent decimal counts across rows, or an ordering rule not derived from the sample indicate no formatting step was applied.
  4. Confirm the script contains no final check comparing the written file's shape/columns/row-keys/format against the sample before submitting.
- **Discriminator**: A real violation is when the output's formatting or ordering is chosen by the agent rather than derived from the template (unrounded floats, guessed sort order, renamed/reordered columns). It is *not* a violation if the script reads the template, applies its rounding/precision and ordering, and the residual differences are only in values that the template itself leaves free.
- **Consequence**: The grader does an exact/tolerance-limited comparison against the reference file and marks the result file WRONG even when the underlying aggregation logic may be close or correct.
191Output artifact contains a re-engineered feature representation instead of the dataset's actual feature vectortaskda-code
Applies when
task -- the task asks for a result file whose columns are indexed positions of "the feature vector" (or otherwise tied to the input records/columns) plus a computed label or prediction.
Pattern
The script builds its own preprocessing pipeline (dropping identifiers/dates, adding engineered aggregates, encoding categoricals, imputing/removing rows) and then writes those transformed columns as the required Feature_i columns, so the saved artifact's column count, column order, and/or row count no longer correspond to the dataset's feature vector as delivered — often without ever stating the mapping back to the original columns or checking row alignment.
Detection procedure
  1. From the task text, fix the expected artifact contract: how many Feature_i columns should exist (one per value of the input feature vector), in what order, and how many rows (one per input record, in original order).
  2. In the script, trace exactly what dataframe is concatenated with the cluster/label column before to_csv: is it the loaded (or minimally cleaned) input, or the scaled/engineered/encoded matrix? Note any columns added or dropped and any rows dropped by dropna/filtering.
  3. Compare the answer's self-reported shape and feature list against step 1: engineered names (age, ratios, totals, *_encoded), a feature count that differs from the input's, or a row count that differs from the raw record count are all mismatches.
  4. Check whether the script asserts shape/row alignment (e.g., saved rows == input rows, feature columns == input feature columns) before writing; absence of such a check compounds the risk.
Discriminator
A real violation is when the persisted Feature_i values are not the dataset's feature values (renamed, transformed, aggregated, reordered, or with rows lost/reindexed). It is fine to scale, encode, or engineer features internally for fitting the model, as long as the written artifact carries the original feature values (or the explicitly requested representation) with one row per input record and the label appended.
Consequence
The grader compares the artifact's columns/rows to the expected feature matrix and reports the file as WRONG/MISSING even though a plausible clustering was performed, yielding 0/1 checks passed.
id 3efe850729a3 · mined from da-code dacode-ml-cluster-014@s4
raw text (what the judge reads)
### Output artifact contains a re-engineered feature representation instead of the dataset's actual feature vector
- **Applies when**: `task` -- the task asks for a result file whose columns are indexed positions of "the feature vector" (or otherwise tied to the input records/columns) plus a computed label or prediction.
- **Pattern**: The script builds its own preprocessing pipeline (dropping identifiers/dates, adding engineered aggregates, encoding categoricals, imputing/removing rows) and then writes *those* transformed columns as the required `Feature_i` columns, so the saved artifact's column count, column order, and/or row count no longer correspond to the dataset's feature vector as delivered — often without ever stating the mapping back to the original columns or checking row alignment.
- **Detection procedure**:
  1. From the task text, fix the expected artifact contract: how many `Feature_i` columns should exist (one per value of the input feature vector), in what order, and how many rows (one per input record, in original order).
  2. In the script, trace exactly what dataframe is concatenated with the cluster/label column before `to_csv`: is it the loaded (or minimally cleaned) input, or the scaled/engineered/encoded matrix? Note any columns added or dropped and any rows dropped by `dropna`/filtering.
  3. Compare the answer's self-reported shape and feature list against step 1: engineered names (age, ratios, totals, `*_encoded`), a feature count that differs from the input's, or a row count that differs from the raw record count are all mismatches.
  4. Check whether the script asserts shape/row alignment (e.g., saved rows == input rows, feature columns == input feature columns) before writing; absence of such a check compounds the risk.
- **Discriminator**: A real violation is when the persisted `Feature_i` values are not the dataset's feature values (renamed, transformed, aggregated, reordered, or with rows lost/reindexed). It is fine to scale, encode, or engineer features *internally* for fitting the model, as long as the written artifact carries the original feature values (or the explicitly requested representation) with one row per input record and the label appended.
- **Consequence**: The grader compares the artifact's columns/rows to the expected feature matrix and reports the file as WRONG/MISSING even though a plausible clustering was performed, yielding 0/1 checks passed.
192Submission never validated against the provided template, plus unrequested post-processing of predictionstaskda-code
Applies when
task -- the task supplies a sample/expected output file (or an explicit format spec) and the scripts write a predictions file at the end.
Pattern
The scripts build the output file purely from the model's in-memory arrays (hand-typed column names, ids taken from whatever frame is loaded) and never read the provided template to check column names/order, row count, id set and ordering, or value dtype; on top of that they apply transformations the task never asked for (rounding continuous predictions to integers, clipping to a range guessed from training data), and the last script executed may overwrite an earlier, stronger model's file with a weaker baseline.
Detection procedure
  1. Read the task/README for the stated output contract (template file, required columns, one row per test record, value type).
  2. Search the scripts for any load of the template file and any assertion comparing the produced frame's shape, columns, and id set/order against it — absence is a red flag.
  3. Check the prediction-writing block for value mutations (round, astype(int), clip, rescaling) and ask whether the task or metric requires them; also check which script ran last and whether it writes the same path with a less-validated model.
  4. Inspect the emitted file's header, row count, and value distribution against the template/test size.
Discriminator
A real violation is unverified structure (no shape/id/column check) or a value transform imposed by the agent's own assumption about the target; it is not a violation when the task explicitly demands integer/rounded/bounded outputs, or when the script asserts equality with the template before writing.
Consequence
The grader's file comparison fails outright (mismatched rows/ids/columns) or the score degrades because the discretization/clipping destroys information the metric measures, so the submission is marked wrong despite reasonable validation numbers printed during training.
id 6f3cb05fb84e · mined from da-code dacode-ml-competition-009@s4
raw text (what the judge reads)
### Submission never validated against the provided template, plus unrequested post-processing of predictions
- **Applies when**: `task` -- the task supplies a sample/expected output file (or an explicit format spec) and the scripts write a predictions file at the end.
- **Pattern**: The scripts build the output file purely from the model's in-memory arrays (hand-typed column names, ids taken from whatever frame is loaded) and never read the provided template to check column names/order, row count, id set and ordering, or value dtype; on top of that they apply transformations the task never asked for (rounding continuous predictions to integers, clipping to a range guessed from training data), and the last script executed may overwrite an earlier, stronger model's file with a weaker baseline.
- **Detection procedure**:
  1. Read the task/README for the stated output contract (template file, required columns, one row per test record, value type).
  2. Search the scripts for any load of the template file and any assertion comparing the produced frame's shape, columns, and id set/order against it — absence is a red flag.
  3. Check the prediction-writing block for value mutations (`round`, `astype(int)`, `clip`, rescaling) and ask whether the task or metric requires them; also check which script ran last and whether it writes the same path with a less-validated model.
  4. Inspect the emitted file's header, row count, and value distribution against the template/test size.
- **Discriminator**: A real violation is unverified structure (no shape/id/column check) or a value transform imposed by the agent's own assumption about the target; it is *not* a violation when the task explicitly demands integer/rounded/bounded outputs, or when the script asserts equality with the template before writing.
- **Consequence**: The grader's file comparison fails outright (mismatched rows/ids/columns) or the score degrades because the discretization/clipping destroys information the metric measures, so the submission is marked wrong despite reasonable validation numbers printed during training.
193Silently dropping input rows (or ID-like columns) instead of imputing/selecting, so the output no longer covers the full datasettaskda-code
Applies when
task -- the task asks for a per-record output file (labels, predictions, scores) derived from a dataset that contains missing values, mixed dtypes, or identifier/code-like columns.
Pattern
The script filters out records with too many NaNs (or dropnas rows), and/or feeds every numeric-looking column into the model including identifiers and arbitrary codes, then reports success without checking that the produced file has one row per input record and that the feature set is meaningful.
Detection procedure
  1. Read the task and note the expected output granularity: one row per input record, with the exact requested column names/ordering.
  2. In the script, find every row-reducing operation (dropna, threshold-based filtering, deduplication) and check whether a per-record output is still required for the removed rows; a reviewer should expect missing values to be imputed (or otherwise handled) rather than rows deleted.
  3. In the script, list the columns fed to the model and flag identifier-like or purely administrative numeric fields (codes, indices, keys) that carry no analytic signal but will dominate/distort distance-based methods.
  4. Compare the row count and column list stated in the answer with the raw dataset's record count and semantic feature set; any shortfall or extraneous code column is a violation.
Discriminator
A real violation is when rows that exist in the source data have no row in the deliverable (or when the feature matrix includes non-informative key/code columns) without the task authorizing such filtering. It is fine if the task explicitly requests filtering/subsetting, or if rows are dropped only from an intermediate fitting step but every original record still receives a label in the output.
Consequence
The saved file has fewer rows (and/or a different feature dimensionality) than the reference, so an exact shape/row-alignment comparison marks the result file WRONG/MISSING and the whole task fails regardless of clustering quality.
id 3b643ca2873e · mined from da-code dacode-ml-cluster-009@s4
raw text (what the judge reads)
### Silently dropping input rows (or ID-like columns) instead of imputing/selecting, so the output no longer covers the full dataset
- **Applies when**: `task` -- the task asks for a per-record output file (labels, predictions, scores) derived from a dataset that contains missing values, mixed dtypes, or identifier/code-like columns.
- **Pattern**: The script filters out records with too many NaNs (or `dropna`s rows), and/or feeds every numeric-looking column into the model including identifiers and arbitrary codes, then reports success without checking that the produced file has one row per input record and that the feature set is meaningful.
- **Detection procedure**:
  1. Read the task and note the expected output granularity: one row per input record, with the exact requested column names/ordering.
  2. In the script, find every row-reducing operation (`dropna`, threshold-based filtering, deduplication) and check whether a per-record output is still required for the removed rows; a reviewer should expect missing values to be imputed (or otherwise handled) rather than rows deleted.
  3. In the script, list the columns fed to the model and flag identifier-like or purely administrative numeric fields (codes, indices, keys) that carry no analytic signal but will dominate/distort distance-based methods.
  4. Compare the row count and column list stated in the answer with the raw dataset's record count and semantic feature set; any shortfall or extraneous code column is a violation.
- **Discriminator**: A real violation is when rows that exist in the source data have no row in the deliverable (or when the feature matrix includes non-informative key/code columns) without the task authorizing such filtering. It is fine if the task explicitly requests filtering/subsetting, or if rows are dropped only from an intermediate fitting step but every original record still receives a label in the output.
- **Consequence**: The saved file has fewer rows (and/or a different feature dimensionality) than the reference, so an exact shape/row-alignment comparison marks the result file WRONG/MISSING and the whole task fails regardless of clustering quality.
194Reporting in-sample fit as evidence of predictive accuracy (no held-out validation)taskda-code
Applies when
task -- the task asks for predictions on a separate test set and the scripts fit a model, then quote accuracy numbers (R², MAE, RMSE) as justification that the output is good.
Pattern
The attempt computes its quality metrics on the exact rows used to fit the model (or on a random split of temporally ordered data), never on a genuinely withheld subset drawn the same way as the target set. A near-perfect score is then reported, hiding preprocessing/feature problems (rows silently dropped by NaN handling, a proxy/leaky predictor dominating importance, train/test feature mismatch, mis-scaled or mis-ordered outputs) that only an honest out-of-sample check would expose.
Detection procedure
  1. From the task, note that the deliverable is predictions for rows the model never sees, and note whether the data is ordered/temporal.
  2. In the scripts, locate where each metric is computed and trace which dataframe/indices are passed to .score()/metric(y, y_pred); check whether those rows were also passed to .fit(), and whether any split respects the ordering used to define the test set.
  3. Compare the answer's reported metrics with that provenance; also check for any independent sanity check on the output (row count equal to the test file, index/order alignment, prediction distribution vs. observed target distribution).
  4. Flag if all reported numbers are training-set numbers and no out-of-sample or distribution/shape sanity check is present.
Discriminator
A real violation reports only fit-data metrics (typically suspiciously high, e.g. R² > 0.95 with one feature carrying most of the importance) and offers no held-out or forward-in-time evaluation and no output sanity check. It is not a violation if the script holds out a properly constructed validation subset (or uses CV consistent with the data's ordering) and reports those metrics, or clearly labels training metrics as diagnostic while separately validating out-of-sample.
Consequence
The submitted prediction file scores far worse than the claimed metrics — errors from dropped/misaligned rows, leaked or unavailable features, or a mis-specified target go undetected, and the grader marks the result file wrong despite a confident, high-R² report.
id 1e9bd2372189 · mined from da-code dacode-ml-regression-002@s4
raw text (what the judge reads)
### Reporting in-sample fit as evidence of predictive accuracy (no held-out validation)
- **Applies when**: `task` -- the task asks for predictions on a separate test set and the scripts fit a model, then quote accuracy numbers (R², MAE, RMSE) as justification that the output is good.
- **Pattern**: The attempt computes its quality metrics on the exact rows used to fit the model (or on a random split of temporally ordered data), never on a genuinely withheld subset drawn the same way as the target set. A near-perfect score is then reported, hiding preprocessing/feature problems (rows silently dropped by NaN handling, a proxy/leaky predictor dominating importance, train/test feature mismatch, mis-scaled or mis-ordered outputs) that only an honest out-of-sample check would expose.
- **Detection procedure**:
  1. From the task, note that the deliverable is predictions for rows the model never sees, and note whether the data is ordered/temporal.
  2. In the scripts, locate where each metric is computed and trace which dataframe/indices are passed to `.score()`/`metric(y, y_pred)`; check whether those rows were also passed to `.fit()`, and whether any split respects the ordering used to define the test set.
  3. Compare the answer's reported metrics with that provenance; also check for any independent sanity check on the output (row count equal to the test file, index/order alignment, prediction distribution vs. observed target distribution).
  4. Flag if all reported numbers are training-set numbers and no out-of-sample or distribution/shape sanity check is present.
- **Discriminator**: A real violation reports only fit-data metrics (typically suspiciously high, e.g. R² > 0.95 with one feature carrying most of the importance) and offers no held-out or forward-in-time evaluation and no output sanity check. It is *not* a violation if the script holds out a properly constructed validation subset (or uses CV consistent with the data's ordering) and reports those metrics, or clearly labels training metrics as diagnostic while separately validating out-of-sample.
- **Consequence**: The submitted prediction file scores far worse than the claimed metrics — errors from dropped/misaligned rows, leaked or unavailable features, or a mis-specified target go undetected, and the grader marks the result file wrong despite a confident, high-R² report.
195Unverifiable plot output: only an image + prose, with the plotted series silently re-aggregatedtaskda-code
Applies when
task -- the task asks for a chart built to an external spec file (e.g., a YAML/JSON plot config) over a time- or category-indexed series, and grading is based on the underlying plotted values rather than the picture.
Pattern
The agent renders the figure and asserts compliance in prose, but (a) never writes out the machine-readable artifacts that encode the plot spec and the plotted data arrays, and (b) changes the series' granularity/ordering/aggregation (e.g., rolling the raw records up to a coarser period, dropping partial periods, or reordering the index) without any statement in the spec authorizing it.
Detection procedure
  1. Read the task and the referenced spec file: list every required output artifact (image and any data/spec dumps) and the exact axis definitions, tick/label values, and series length implied by the spec.
  2. Read the scripts: confirm each required artifact is written to disk with the required filename, and trace the series from raw file to plot() call — note every groupby, resample, sum, mean, dropna, sort, or date-parsing step that changes the number or level of points.
  3. Compare the agent's reported data points (count, labels, magnitudes) against the raw file's native granularity and row count; a handful of points where the source has many per label is a red flag.
  4. Check the answer for evidence-free compliance claims ("✓ follows all specifications") that are not backed by a printed comparison against the spec's fields.
Discriminator
A real violation is when the spec (or task) implies the raw granularity/ordering and the script aggregates or truncates anyway, or when a required artifact filename never appears in any to_json/np.save/savefig call. It is not a violation if the spec explicitly requests the aggregation and the script still dumps the plotted arrays and spec fields to the required files.
Consequence
The value-based checks on the expected artifacts fail as WRONG/MISSING — either because the files were never produced, or because the saved array has the wrong length/values relative to the ground-truth series — even though the rendered PNG looks plausible.
id 8e2a439cbe62 · mined from da-code dacode-plot-line-015@s4
raw text (what the judge reads)
### Unverifiable plot output: only an image + prose, with the plotted series silently re-aggregated
- **Applies when**: `task` -- the task asks for a chart built to an external spec file (e.g., a YAML/JSON plot config) over a time- or category-indexed series, and grading is based on the underlying plotted values rather than the picture.
- **Pattern**: The agent renders the figure and asserts compliance in prose, but (a) never writes out the machine-readable artifacts that encode the plot spec and the plotted data arrays, and (b) changes the series' granularity/ordering/aggregation (e.g., rolling the raw records up to a coarser period, dropping partial periods, or reordering the index) without any statement in the spec authorizing it.
- **Detection procedure**:
  1. Read the task and the referenced spec file: list every required output artifact (image *and* any data/spec dumps) and the exact axis definitions, tick/label values, and series length implied by the spec.
  2. Read the scripts: confirm each required artifact is written to disk with the required filename, and trace the series from raw file to `plot()` call — note every `groupby`, `resample`, `sum`, `mean`, `dropna`, `sort`, or date-parsing step that changes the number or level of points.
  3. Compare the agent's reported data points (count, labels, magnitudes) against the raw file's native granularity and row count; a handful of points where the source has many per label is a red flag.
  4. Check the answer for evidence-free compliance claims ("✓ follows all specifications") that are not backed by a printed comparison against the spec's fields.
- **Discriminator**: A real violation is when the spec (or task) implies the raw granularity/ordering and the script aggregates or truncates anyway, or when a required artifact filename never appears in any `to_json`/`np.save`/`savefig` call. It is *not* a violation if the spec explicitly requests the aggregation and the script still dumps the plotted arrays and spec fields to the required files.
- **Consequence**: The value-based checks on the expected artifacts fail as WRONG/MISSING — either because the files were never produced, or because the saved array has the wrong length/values relative to the ground-truth series — even though the rendered PNG looks plausible.
196Hard-coded/recalled input data instead of loading the provided files and spectaskda-code
Applies when
task -- The task ships data files plus instruction/format files (e.g., a tips/README and a sample result template) and the script must read them to produce the requested output.
Pattern
The script embeds literal arrays/values typed from the agent's memory of a "well-known" dataset, and invents the output column names, ordering, rounding, and file path instead of deriving them from the provided data and sample output file; no code path ever opens the supplied inputs.
Detection procedure
  1. List the input artifacts named or implied by the task (data files, instructions file, sample output file) and note any stated procedural constraints (which statistic, which resampling scheme, how many replicates, seed, rounding).
  2. Grep the scripts for read operations (read_csv, open, load, path strings) against each artifact; flag any artifact that is never read, especially if the corresponding values appear as literals in the code.
  3. Compare the written output's schema (column names, row count, precision, file location) to the sample template; if the template was never read, verify the schema was at least copied verbatim from the task text.
  4. Sanity-check the reported number against the method: e.g., a resampling p-value can never be smaller than 1/n_replicates, so an exact 0 with a finite number of replicates should be reported as "< 1/n" or the procedure/precision reconsidered.
Discriminator
A real violation is when the analysis values themselves (or the output schema) come from memory/assumption while the provided files exist and are unread; it is not a violation if the script loads the provided files and merely defines constants such as seed, replicate count, or tolerances inline.
Consequence
The computed statistic is based on the wrong sample sizes/values and the output schema may not match the template, so exact-file comparison against the expected result.csv fails even when the statistical method is conceptually right (here compounded by an implausible degenerate p-value of exactly 0).
id 08c9a6c96256 · mined from da-code dacode-data-sa-028@s4
raw text (what the judge reads)
### Hard-coded/recalled input data instead of loading the provided files and spec
- **Applies when**: `task` -- The task ships data files plus instruction/format files (e.g., a tips/README and a sample result template) and the script must read them to produce the requested output.
- **Pattern**: The script embeds literal arrays/values typed from the agent's memory of a "well-known" dataset, and invents the output column names, ordering, rounding, and file path instead of deriving them from the provided data and sample output file; no code path ever opens the supplied inputs.
- **Detection procedure**:
  1. List the input artifacts named or implied by the task (data files, instructions file, sample output file) and note any stated procedural constraints (which statistic, which resampling scheme, how many replicates, seed, rounding).
  2. Grep the scripts for read operations (`read_csv`, `open`, `load`, path strings) against each artifact; flag any artifact that is never read, especially if the corresponding values appear as literals in the code.
  3. Compare the written output's schema (column names, row count, precision, file location) to the sample template; if the template was never read, verify the schema was at least copied verbatim from the task text.
  4. Sanity-check the reported number against the method: e.g., a resampling p-value can never be smaller than 1/n_replicates, so an exact 0 with a finite number of replicates should be reported as "< 1/n" or the procedure/precision reconsidered.
- **Discriminator**: A real violation is when the analysis values themselves (or the output schema) come from memory/assumption while the provided files exist and are unread; it is *not* a violation if the script loads the provided files and merely defines constants such as seed, replicate count, or tolerances inline.
- **Consequence**: The computed statistic is based on the wrong sample sizes/values and the output schema may not match the template, so exact-file comparison against the expected `result.csv` fails even when the statistical method is conceptually right (here compounded by an implausible degenerate p-value of exactly 0).
197Referenced spec files never read; required output artifacts never producedtaskda-code
Applies when
task -- The prompt points to auxiliary instruction/config files (e.g., a tips/notes file, a YAML/JSON style or parameter spec) and/or implies multiple saved artifacts, and the scripts only load the raw data.
Pattern
The agent never opens the referenced spec files, invents its own filtering rules, groupings, labels, axis/units, and figure styling from guesswork, and emits only the single most obvious output file, omitting the companion data/config artifacts (serialized plot data, arrays, metadata) that the specified format implies.
Detection procedure
  1. List every file path and format spec mentioned in the task prompt (instruction files, config files, and each expected output name/extension).
  2. Grep the scripts for reads of each instruction/config file; if any is absent, the parameters used (exclusions, season/bin definitions, labels, colors, sizes, ranges) are unverified guesses — note that hard-coded values appear with no traceable source.
  3. Check the scripts' write calls against the full list of expected outputs; flag if any required artifact (e.g., a serialized plot-data or array file) is never written.
  4. Read the answer for claims of compliance ("formatted according to spec", "excluded specified items") that are not backed by any parsing of the spec file.
Discriminator
A real violation is when the spec file is never opened/parsed at all, or an expected output filename is never written; it is not a violation if the script reads the spec and programmatically applies its keys (even if it also has sensible defaults), or if the extra artifacts are legitimately produced by a different script step.
Consequence
The grader's per-file comparison fails for the missing artifacts and for the plot whose data, styling, filtering, or labels diverge from the unread specification — 0 checks passed despite a "completed successfully" report.
id cb1722d9f7a9 · mined from da-code dacode-plot-line-006@s4
raw text (what the judge reads)
### Referenced spec files never read; required output artifacts never produced
- **Applies when**: `task` -- The prompt points to auxiliary instruction/config files (e.g., a tips/notes file, a YAML/JSON style or parameter spec) and/or implies multiple saved artifacts, and the scripts only load the raw data.
- **Pattern**: The agent never opens the referenced spec files, invents its own filtering rules, groupings, labels, axis/units, and figure styling from guesswork, and emits only the single most obvious output file, omitting the companion data/config artifacts (serialized plot data, arrays, metadata) that the specified format implies.
- **Detection procedure**:
  1. List every file path and format spec mentioned in the task prompt (instruction files, config files, and each expected output name/extension).
  2. Grep the scripts for reads of each instruction/config file; if any is absent, the parameters used (exclusions, season/bin definitions, labels, colors, sizes, ranges) are unverified guesses — note that hard-coded values appear with no traceable source.
  3. Check the scripts' write calls against the full list of expected outputs; flag if any required artifact (e.g., a serialized plot-data or array file) is never written.
  4. Read the answer for claims of compliance ("formatted according to spec", "excluded specified items") that are not backed by any parsing of the spec file.
- **Discriminator**: A real violation is when the spec file is never opened/parsed at all, or an expected output filename is never written; it is *not* a violation if the script reads the spec and programmatically applies its keys (even if it also has sensible defaults), or if the extra artifacts are legitimately produced by a different script step.
- **Consequence**: The grader's per-file comparison fails for the missing artifacts and for the plot whose data, styling, filtering, or labels diverge from the unread specification — 0 checks passed despite a "completed successfully" report.
198Ignoring a referenced specification file and substituting assumed definitionstaskda-code
Applies when
task -- the prompt points to an auxiliary document/config in the workspace (e.g., a .md/.json/.txt spec) that defines how categories, bins, filters, or a metric must be constructed.
Pattern
The scripts never open or parse the referenced file; instead the agent hardcodes groupings/definitions inferred from the raw data's own distinct values or from a "reasonable" convention, and the write-up merely asserts that the spec was followed.
Detection procedure
  1. Read the task and list every external artifact it names as governing the method, plus every named output file/format.
  2. Grep the scripts for a read of that artifact (open/read_csv/json.load on that path) and for any printout of its contents.
  3. If absent, check whether the categories/bins/thresholds in the code are literals invented by the agent (or copied from the data's raw levels) rather than derived from the spec.
  4. Check the answer for evidence-free claims like "follows the method in <spec>" with no quoted spec content, and confirm all required output artifacts are actually produced.
Discriminator
A real violation is code whose category/bin definitions have no traceable source in the referenced file; it is fine if the script loads and echoes the spec (or the answer quotes it verbatim) and the hardcoded values demonstrably match it, even if later hardcoded for convenience.
Consequence
The computed groups/aggregates differ from the specified ones, so the saved figure and any serialized data/array outputs mismatch the expected values, and all output-file checks fail even though the plot's title/labels are correct.
id a765a1f0a901 · mined from da-code dacode-plot-bar-005@s4
raw text (what the judge reads)
### Ignoring a referenced specification file and substituting assumed definitions
- **Applies when**: `task` -- the prompt points to an auxiliary document/config in the workspace (e.g., a `.md`/`.json`/`.txt` spec) that defines how categories, bins, filters, or a metric must be constructed.
- **Pattern**: The scripts never open or parse the referenced file; instead the agent hardcodes groupings/definitions inferred from the raw data's own distinct values or from a "reasonable" convention, and the write-up merely asserts that the spec was followed.
- **Detection procedure**:
  1. Read the task and list every external artifact it names as governing the method, plus every named output file/format.
  2. Grep the scripts for a read of that artifact (open/read_csv/json.load on that path) and for any printout of its contents.
  3. If absent, check whether the categories/bins/thresholds in the code are literals invented by the agent (or copied from the data's raw levels) rather than derived from the spec.
  4. Check the answer for evidence-free claims like "follows the method in <spec>" with no quoted spec content, and confirm all required output artifacts are actually produced.
- **Discriminator**: A real violation is code whose category/bin definitions have no traceable source in the referenced file; it is fine if the script loads and echoes the spec (or the answer quotes it verbatim) and the hardcoded values demonstrably match it, even if later hardcoded for convenience.
- **Consequence**: The computed groups/aggregates differ from the specified ones, so the saved figure and any serialized data/array outputs mismatch the expected values, and all output-file checks fail even though the plot's title/labels are correct.
199Required output artifact and schema not produced by the scriptstaskda-code
Applies when
task -- the task specifies a concrete answer format (e.g., a JSON template with keys and bracketed/list values) and/or an expected result file that the grader will read.
Pattern
The scripts only compute and print the result to stdout; no step writes the answer to the required file, and the final answer is typed by hand in a shape that deviates from the given template (scalars instead of the shown list/array form, renamed or reordered keys, extra/missing precision or units).
Detection procedure
  1. Read the task statement and record the exact required artifact (file name/path, if any) and the literal answer template, including key spelling and whether values are wrapped in [...].
  2. Scan every script for a write/serialization call (json.dump, to_csv, open(..., 'w')) targeting that artifact; if only print/logging exists, the artifact is missing.
  3. Compare the submitted answer character-by-character against the template: same keys, same value container type, values consistent with any stated rounding/unit convention.
  4. Flag if the artifact is absent or any structural element of the template is not reproduced.
Discriminator
A real violation is a missing output file or a structural deviation from the stated template (scalar vs list, altered key names). It is not a violation if the file is written with the exact schema and only cosmetic whitespace/ordering differs, or if the task genuinely requested plain text with no file.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even when the underlying computation (imputation and argmax) is numerically correct.
id 8cca6aab1cc5 · mined from da-code dacode-di-text-002@s4
raw text (what the judge reads)
### Required output artifact and schema not produced by the scripts
- **Applies when**: `task` -- the task specifies a concrete answer format (e.g., a JSON template with keys and bracketed/list values) and/or an expected result file that the grader will read.
- **Pattern**: The scripts only compute and `print` the result to stdout; no step writes the answer to the required file, and the final answer is typed by hand in a shape that deviates from the given template (scalars instead of the shown list/array form, renamed or reordered keys, extra/missing precision or units).
- **Detection procedure**:
  1. Read the task statement and record the exact required artifact (file name/path, if any) and the literal answer template, including key spelling and whether values are wrapped in `[...]`.
  2. Scan every script for a write/serialization call (`json.dump`, `to_csv`, `open(..., 'w')`) targeting that artifact; if only `print`/logging exists, the artifact is missing.
  3. Compare the submitted answer character-by-character against the template: same keys, same value container type, values consistent with any stated rounding/unit convention.
  4. Flag if the artifact is absent or any structural element of the template is not reproduced.
- **Discriminator**: A real violation is a missing output file or a structural deviation from the stated template (scalar vs list, altered key names). It is *not* a violation if the file is written with the exact schema and only cosmetic whitespace/ordering differs, or if the task genuinely requested plain text with no file.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even when the underlying computation (imputation and argmax) is numerically correct.
200Output schema invented rather than derived from the stated required formattaskda-code
Applies when
task -- the task asks for results to be written to a specific results file "in the required format", and the scripts construct that file's columns/rows from scratch.
Pattern
The agent never locates or reconstructs the expected schema (column names, ordering, index/date column presence and formatting, number of rows, rounding/units), and instead invents plausible-looking names and layout; its "verification" script only re-checks its own arithmetic against its own file, so any schema mismatch is invisible.
Detection procedure
  1. Read the task/README for any explicit or implied output contract (named columns, example rows, sample/template file, naming conventions inherited from the input data or from the domain wording used in the prompt).
  2. Inspect the script that writes the output: are column names, row set, ordering, and value definitions traceable to that contract, or are they ad-hoc labels chosen by the agent (e.g., abbreviations, suffixes, extra/renamed columns)?
  3. Check whether any step reads a template/expected-output artifact, or at minimum enumerates the schema assumptions and cross-checks shape/row count/date coverage against the input.
  4. Check the final answer: does it assert format compliance without evidence, or does it show the actual header/first rows compared to the required one?
Discriminator
A real violation is when the required format is discoverable (stated fields, template file, conventional naming implied by the prompt) but the script's header/layout does not demonstrably follow it. A look-alike that is fine is when the task genuinely leaves naming free and the agent documents the choice while matching all stated constraints (row coverage, ordering, precision).
Consequence
The numeric computation may be right, but the grader's file comparison fails on missing/renamed columns or mismatched rows, scoring the result file WRONG/MISSING (0 checks passed).
id a4c887eff964 · mined from da-code dacode-dm-csv-050@s4
raw text (what the judge reads)
### Output schema invented rather than derived from the stated required format
- **Applies when**: `task` -- the task asks for results to be written to a specific results file "in the required format", and the scripts construct that file's columns/rows from scratch.
- **Pattern**: The agent never locates or reconstructs the expected schema (column names, ordering, index/date column presence and formatting, number of rows, rounding/units), and instead invents plausible-looking names and layout; its "verification" script only re-checks its own arithmetic against its own file, so any schema mismatch is invisible.
- **Detection procedure**:
  1. Read the task/README for any explicit or implied output contract (named columns, example rows, sample/template file, naming conventions inherited from the input data or from the domain wording used in the prompt).
  2. Inspect the script that writes the output: are column names, row set, ordering, and value definitions traceable to that contract, or are they ad-hoc labels chosen by the agent (e.g., abbreviations, suffixes, extra/renamed columns)?
  3. Check whether any step reads a template/expected-output artifact, or at minimum enumerates the schema assumptions and cross-checks shape/row count/date coverage against the input.
  4. Check the final answer: does it assert format compliance without evidence, or does it show the actual header/first rows compared to the required one?
- **Discriminator**: A real violation is when the required format is discoverable (stated fields, template file, conventional naming implied by the prompt) but the script's header/layout does not demonstrably follow it. A look-alike that is fine is when the task genuinely leaves naming free and the agent documents the choice while matching all stated constraints (row coverage, ordering, precision).
- **Consequence**: The numeric computation may be right, but the grader's file comparison fails on missing/renamed columns or mismatched rows, scoring the result file WRONG/MISSING (0 checks passed).
201Statistical test run on an unvetted raw column (sentinels/outliers/sample-size limits) with no sanity cross-checktaskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test or distribution statistic on a single column, and the script feeds the column straight into the test after at most a dropna().
Pattern
The attempt treats "the column" as automatically equal to "the valid data": it never inspects value counts, sentinel/placeholder codes (0, -1, -9, 9999), duplicated header/summary rows, or the row count relative to the test's validity range (e.g., large-N tests that reject almost any real data, or scipy warnings about approximate p-values). It then reports the raw verdict even when the accompanying descriptive statistics (large |skew|, heavy tails, min/max far from the bulk) point to a contaminating subpopulation rather than a genuinely non-normal target distribution.
Detection procedure
  1. Read the task for any implied population/filter (valid records, non-missing coverage, a specific window/subset) and for the test's stated parameters; note whether the script encodes that filter at all.
  2. In the script, check whether anything beyond dropna() validates the input: printed value counts / min-max / histogram, removal of sentinel or degenerate rows, and an explicit check that N is inside the test's reliable range.
  3. Compare the reported test verdict with the reported companion statistics and summary output: if the rejection is driven by a small number of extreme values or by very large N, but the bulk of the data looks symmetric, the input was not vetted.
  4. Confirm the script does not re-run or contrast the test on the cleaned/plausible subset before committing to an answer.
Discriminator
A real violation is when the script never examines the column's value distribution or N-validity and the reported skew/kurtosis are extreme relative to what the described quantity should look like. It is fine if the script prints diagnostics (counts, extremes, distribution shape), explicitly justifies keeping all rows, and confirms the test is valid at that sample size — even if the final verdict is "not normal".
Consequence
The p-value is computed on a contaminated or out-of-range sample, so the normality/hypothesis verdict flips relative to ground truth and the shape statistics are distorted, failing the graded check despite a syntactically correct answer format.
id 58a315d1c7a4 · mined from infiagent-dabench dabench-298@s4
raw text (what the judge reads)
### Statistical test run on an unvetted raw column (sentinels/outliers/sample-size limits) with no sanity cross-check

- **Applies when**: `task` -- the task asks for a hypothesis test or distribution statistic on a single column, and the script feeds the column straight into the test after at most a `dropna()`.
- **Pattern**: The attempt treats "the column" as automatically equal to "the valid data": it never inspects value counts, sentinel/placeholder codes (0, -1, -9, 9999), duplicated header/summary rows, or the row count relative to the test's validity range (e.g., large-N tests that reject almost any real data, or scipy warnings about approximate p-values). It then reports the raw verdict even when the accompanying descriptive statistics (large |skew|, heavy tails, min/max far from the bulk) point to a contaminating subpopulation rather than a genuinely non-normal target distribution.
- **Detection procedure**:
  1. Read the task for any implied population/filter (valid records, non-missing coverage, a specific window/subset) and for the test's stated parameters; note whether the script encodes that filter at all.
  2. In the script, check whether anything beyond `dropna()` validates the input: printed value counts / min-max / histogram, removal of sentinel or degenerate rows, and an explicit check that N is inside the test's reliable range.
  3. Compare the reported test verdict with the reported companion statistics and summary output: if the rejection is driven by a small number of extreme values or by very large N, but the bulk of the data looks symmetric, the input was not vetted.
  4. Confirm the script does not re-run or contrast the test on the cleaned/plausible subset before committing to an answer.
- **Discriminator**: A real violation is when the script never examines the column's value distribution or N-validity and the reported skew/kurtosis are extreme relative to what the described quantity should look like. It is fine if the script prints diagnostics (counts, extremes, distribution shape), explicitly justifies keeping all rows, and confirms the test is valid at that sample size — even if the final verdict is "not normal".
- **Consequence**: The p-value is computed on a contaminated or out-of-range sample, so the normality/hypothesis verdict flips relative to ground truth and the shape statistics are distorted, failing the graded check despite a syntactically correct answer format.
202Final deliverable comes from an unvalidated "fast fallback" script that overwrites a validated model's outputtaskda-code
Applies when
task -- multiple scripts in the attempt write to the same required output artifact, and at least one of them is a quick/simplified fallback model with no held-out evaluation.
Pattern
The agent builds and validates one modeling approach, then later runs a stripped-down script (simpler model, no train/validation split, no metric printed) that rewrites the same output path; the delivered artifact therefore reflects the weakest, never-scored model, and the answer describes it as if it were the chosen best approach without any comparative evidence.
Detection procedure
  1. From the task, note the exact required output artifact (name and location) and the metric that will be used to judge it.
  2. Scan every script for writes to that artifact; determine which script executes last and thus determines the delivered content, and confirm the path/filename matches exactly what the task requires.
  3. Check whether that final-writing script computes any held-out/cross-validated score, and whether its score was compared against the other candidate models before being chosen.
  4. Compare the answer's claimed model/metrics against what the final-writing script actually does and reports.
Discriminator
A real violation is when the last writer is unvalidated or was shown/likely to be weaker (e.g., plain linear model replacing a tuned/ensembled tree model) or writes to a path other than the required one; it is fine if the final writer is the model whose held-out score was demonstrably the best among candidates and writes to the specified path.
Consequence
The graded file contains predictions from an inferior or misplaced model, so the score falls below the accepted threshold (or the expected file is judged missing/wrong) even though the answer text claims a validated, well-performing model.
id 819e038b3f82 · mined from da-code dacode-ml-competition-008@s4
raw text (what the judge reads)
### Final deliverable comes from an unvalidated "fast fallback" script that overwrites a validated model's output
- **Applies when**: `task` -- multiple scripts in the attempt write to the same required output artifact, and at least one of them is a quick/simplified fallback model with no held-out evaluation.
- **Pattern**: The agent builds and validates one modeling approach, then later runs a stripped-down script (simpler model, no train/validation split, no metric printed) that rewrites the same output path; the delivered artifact therefore reflects the weakest, never-scored model, and the answer describes it as if it were the chosen best approach without any comparative evidence.
- **Detection procedure**:
  1. From the task, note the exact required output artifact (name and location) and the metric that will be used to judge it.
  2. Scan every script for writes to that artifact; determine which script executes last and thus determines the delivered content, and confirm the path/filename matches exactly what the task requires.
  3. Check whether that final-writing script computes any held-out/cross-validated score, and whether its score was compared against the other candidate models before being chosen.
  4. Compare the answer's claimed model/metrics against what the final-writing script actually does and reports.
- **Discriminator**: A real violation is when the last writer is unvalidated or was shown/likely to be weaker (e.g., plain linear model replacing a tuned/ensembled tree model) or writes to a path other than the required one; it is fine if the final writer is the model whose held-out score was demonstrably the best among candidates and writes to the specified path.
- **Consequence**: The graded file contains predictions from an inferior or misplaced model, so the score falls below the accepted threshold (or the expected file is judged missing/wrong) even though the answer text claims a validated, well-performing model.
203Guessed eligibility threshold applied to every sub-ranking instead of only where definedtaskda-code
Applies when
task -- The task asks for several separate top-N rankings of grouped entities and states a minimum-qualification rule (or gives a partially specified/ambiguous rule) plus a sample output file defining the expected format.
Pattern
The agent invents an interpretation of the qualification rule (e.g., re-deriving the threshold as a per-item average rather than the stated aggregate), never cross-checks it against the spec or the provided sample output, and then applies that single filter uniformly to all sub-rankings even though the rule was scoped to only one of them — silently dropping entities that legitimately belong in the other rankings.
Detection procedure
  1. Read the task/README and note exactly which sub-question the qualification rule is attached to, what quantity it is measured on, and whether the statement is complete.
  2. Read the script and check how the threshold quantity is computed (aggregate vs. mean, per-group vs. per-row) and to which sub-rankings the filtered frame is passed.
  3. Check whether the script ever loads/inspects the provided sample output file to confirm column names, ordering, id conventions, and to sanity-check candidate entries.
  4. Compare the filtered group count to the unfiltered count and ask whether obviously dominant entities in the unrestricted sub-rankings could have been removed by the filter.
Discriminator
A real violation is applying an unstated or re-derived filter to sub-questions where the task never imposed it, or redefining the threshold statistic without evidence; it is fine if the task explicitly scopes one qualification rule to all outputs, or if the agent tests both readings and validates the choice against the sample/README.
Consequence
Rankings for the unfiltered categories contain the wrong entities (or wrong order), so exact-match comparison against the expected file fails even though the per-group statistics were computed correctly.
id fcd23980b62e · mined from da-code dacode-dm-csv-009@s4
raw text (what the judge reads)
### Guessed eligibility threshold applied to every sub-ranking instead of only where defined
- **Applies when**: `task` -- The task asks for several separate top-N rankings of grouped entities and states a minimum-qualification rule (or gives a partially specified/ambiguous rule) plus a sample output file defining the expected format.
- **Pattern**: The agent invents an interpretation of the qualification rule (e.g., re-deriving the threshold as a per-item average rather than the stated aggregate), never cross-checks it against the spec or the provided sample output, and then applies that single filter uniformly to all sub-rankings even though the rule was scoped to only one of them — silently dropping entities that legitimately belong in the other rankings.
- **Detection procedure**:
  1. Read the task/README and note exactly which sub-question the qualification rule is attached to, what quantity it is measured on, and whether the statement is complete.
  2. Read the script and check how the threshold quantity is computed (aggregate vs. mean, per-group vs. per-row) and to which sub-rankings the filtered frame is passed.
  3. Check whether the script ever loads/inspects the provided sample output file to confirm column names, ordering, id conventions, and to sanity-check candidate entries.
  4. Compare the filtered group count to the unfiltered count and ask whether obviously dominant entities in the unrestricted sub-rankings could have been removed by the filter.
- **Discriminator**: A real violation is applying an unstated or re-derived filter to sub-questions where the task never imposed it, or redefining the threshold statistic without evidence; it is fine if the task explicitly scopes one qualification rule to all outputs, or if the agent tests both readings and validates the choice against the sample/README.
- **Consequence**: Rankings for the unfiltered categories contain the wrong entities (or wrong order), so exact-match comparison against the expected file fails even though the per-group statistics were computed correctly.
204Fabricating synthetic input data instead of using the provided datasettaskda-code
Applies when
task -- The task references a supplied dataset (and config/spec files) that the analysis must be computed from.
Pattern
The agent cannot find or does not load the real input file, so it generates random/simulated records (e.g., np.random, hard-coded sample rows) and runs the requested analysis on that invented data, then reports the resulting numbers as if they were real findings.
Detection procedure
  1. Read the task/README to identify which input file(s) the deliverable must be derived from, and any required output artifacts.
  2. Scan the scripts for data creation calls (np.random.*, synthetic loops building DataFrames, to_csv of self-made data) and check whether the path actually read for the analysis is a file the agent itself wrote rather than the provided dataset.
  3. Check the answer's reported counts/ranges against the dataset's documented scale and semantics (e.g., total row count, expected units/format of fields); flag mismatches or suspiciously round/normal-looking distributions.
  4. Confirm every required output artifact is produced from the real data; missing ones (e.g., a serialized plot/array file) also signal the spec was not followed.
Discriminator
A real violation is when the reported result is computed from agent-generated values; it is fine to synthesize data purely for smoke-testing/unit-testing code, provided the final artifacts and reported numbers come from loading the provided dataset.
Consequence
All value-based checks fail — plotted/serialized counts and any derived files differ entirely from the ground truth, so the grader marks every expected file wrong or missing.
id 887d94aa4db8 · mined from da-code dacode-plot-bar-007@s4
raw text (what the judge reads)
### Fabricating synthetic input data instead of using the provided dataset
- **Applies when**: `task` -- The task references a supplied dataset (and config/spec files) that the analysis must be computed from.
- **Pattern**: The agent cannot find or does not load the real input file, so it generates random/simulated records (e.g., `np.random`, hard-coded sample rows) and runs the requested analysis on that invented data, then reports the resulting numbers as if they were real findings.
- **Detection procedure**:
  1. Read the task/README to identify which input file(s) the deliverable must be derived from, and any required output artifacts.
  2. Scan the scripts for data creation calls (`np.random.*`, synthetic loops building DataFrames, `to_csv` of self-made data) and check whether the path actually read for the analysis is a file the agent itself wrote rather than the provided dataset.
  3. Check the answer's reported counts/ranges against the dataset's documented scale and semantics (e.g., total row count, expected units/format of fields); flag mismatches or suspiciously round/normal-looking distributions.
  4. Confirm every required output artifact is produced from the real data; missing ones (e.g., a serialized plot/array file) also signal the spec was not followed.
- **Discriminator**: A real violation is when the reported result is computed from agent-generated values; it is fine to synthesize data purely for smoke-testing/unit-testing code, provided the final artifacts and reported numbers come from loading the provided dataset.
- **Consequence**: All value-based checks fail — plotted/serialized counts and any derived files differ entirely from the ground truth, so the grader marks every expected file wrong or missing.
205Ambiguous standard-deviation convention (population vs. sample `ddof`) chosen without justification or cross-checktaskinfiagent-dabench
Applies when
task -- the task asks for a dispersion statistic (standard deviation, variance, standard error) on a data subset and the script computes it with a library call that has a ddof/denominator option.
Pattern
The script silently picks one normalization (e.g. np.std(x) with ddof=0, or conversely ddof=1) without checking which convention the task/reference tooling implies, and reports a single rounded number; the mean matches the reference but the spread statistic is off by the n/(n-1) factor.
Detection procedure
  1. Read the task statement for any hint about the convention (sample vs. population, "unbiased", or the tool implied by the workflow, e.g. pandas/describe() defaults vs. numpy defaults).
  2. In the script, locate the dispersion computation and note the explicit or default ddof.
  3. Check whether the script computes the alternative convention as well (or prints both values with the subset size n) so the reporter can see the magnitude of the difference; if only one value is produced and no rationale is given, flag it.
  4. Confirm the reported spread is consistent with the reported mean/subset (same filtered array, same rounding) — a mismatch of roughly a factor sqrt(n/(n-1)) from the plausible reference indicates the wrong convention.
Discriminator
Not a violation if the task or its stated tooling unambiguously fixes the convention (e.g. "population standard deviation", or the script mirrors the pipeline's stated library default) and the script's ddof matches it; it is a violation when the convention is unstated, the script uses a non-default-for-the-ecosystem ddof, and no both-ways sanity print or justification exists.
Consequence
The mean and any outlier/subset outputs check out while the standard deviation fails the grader by the small sqrt(n/(n-1)) factor, producing a partially-correct answer (e.g. 1 of 2 numeric checks passing).
id dcdd79d9e3ff · mined from infiagent-dabench dabench-495@s4
raw text (what the judge reads)
### Ambiguous standard-deviation convention (population vs. sample `ddof`) chosen without justification or cross-check
- **Applies when**: `task` -- the task asks for a dispersion statistic (standard deviation, variance, standard error) on a data subset and the script computes it with a library call that has a `ddof`/denominator option.
- **Pattern**: The script silently picks one normalization (e.g. `np.std(x)` with `ddof=0`, or conversely `ddof=1`) without checking which convention the task/reference tooling implies, and reports a single rounded number; the mean matches the reference but the spread statistic is off by the `n/(n-1)` factor.
- **Detection procedure**:
  1. Read the task statement for any hint about the convention (sample vs. population, "unbiased", or the tool implied by the workflow, e.g. pandas/`describe()` defaults vs. numpy defaults).
  2. In the script, locate the dispersion computation and note the explicit or default `ddof`.
  3. Check whether the script computes the alternative convention as well (or prints both values with the subset size `n`) so the reporter can see the magnitude of the difference; if only one value is produced and no rationale is given, flag it.
  4. Confirm the reported spread is consistent with the reported mean/subset (same filtered array, same rounding) — a mismatch of roughly a factor `sqrt(n/(n-1))` from the plausible reference indicates the wrong convention.
- **Discriminator**: Not a violation if the task or its stated tooling unambiguously fixes the convention (e.g. "population standard deviation", or the script mirrors the pipeline's stated library default) and the script's `ddof` matches it; it *is* a violation when the convention is unstated, the script uses a non-default-for-the-ecosystem `ddof`, and no both-ways sanity print or justification exists.
- **Consequence**: The mean and any outlier/subset outputs check out while the standard deviation fails the grader by the small `sqrt(n/(n-1))` factor, producing a partially-correct answer (e.g. 1 of 2 numeric checks passing).
206Answer asserted from prior/world knowledge instead of computed from the provided datataskda-code
Applies when
task -- the task asks for a ranking, extreme values, or statistics that must be derived from a supplied dataset after a specified preprocessing step (e.g., a stated imputation rule), and the answer is plausible-looking domain knowledge.
Pattern
The attempt ships a final answer with no (or only trivial) code that actually loads the file, applies the stated preprocessing, sorts, and slices the requested rows — the entries are ones a knowledgeable person would guess, and the required transformation is never demonstrated on the real column values.
Detection procedure
  1. Read the task and list every operational requirement (which file/column, the required missing-value handling, the sort direction, the number of items, the output schema).
  2. Inspect the submitted scripts/logs for an executable chain: read data → apply the stated preprocessing → compute/sort → select top-N → write the requested output file; note any requirement with no corresponding code.
  3. Cross-check the reported entries against the dataset's own naming and coverage (exact country/entity spellings, whether such rows exist at all, whether ties or imputed rows change the boundary of the top-N).
  4. Confirm the answer was persisted in the exact requested artifact (filename, keys, ordering) rather than only printed in prose.
Discriminator
A real violation is when no run on the actual file can be traced (or the reported labels don't match the dataset's label strings/row set), so the result is unverifiable guesswork; a look-alike that is fine is a fully scripted pipeline whose output happens to agree with domain intuition, with printed intermediate values (row counts, imputed-cell count, sorted head/tail) proving it came from the data.
Consequence
Entity names or membership deviate from what the dataset yields (different spellings, missing/extra countries, wrong boundary member or order), so the expected result file comparison fails and the check scores 0.
id 28ebc0833a2b · mined from da-code dacode-di-text-003@s4
raw text (what the judge reads)
### Answer asserted from prior/world knowledge instead of computed from the provided data
- **Applies when**: `task` -- the task asks for a ranking, extreme values, or statistics that must be derived from a supplied dataset after a specified preprocessing step (e.g., a stated imputation rule), and the answer is plausible-looking domain knowledge.
- **Pattern**: The attempt ships a final answer with no (or only trivial) code that actually loads the file, applies the stated preprocessing, sorts, and slices the requested rows — the entries are ones a knowledgeable person would guess, and the required transformation is never demonstrated on the real column values.
- **Detection procedure**:
  1. Read the task and list every operational requirement (which file/column, the required missing-value handling, the sort direction, the number of items, the output schema).
  2. Inspect the submitted scripts/logs for an executable chain: read data → apply the stated preprocessing → compute/sort → select top-N → write the requested output file; note any requirement with no corresponding code.
  3. Cross-check the reported entries against the dataset's own naming and coverage (exact country/entity spellings, whether such rows exist at all, whether ties or imputed rows change the boundary of the top-N).
  4. Confirm the answer was persisted in the exact requested artifact (filename, keys, ordering) rather than only printed in prose.
- **Discriminator**: A real violation is when no run on the actual file can be traced (or the reported labels don't match the dataset's label strings/row set), so the result is unverifiable guesswork; a look-alike that is fine is a fully scripted pipeline whose output happens to agree with domain intuition, with printed intermediate values (row counts, imputed-cell count, sorted head/tail) proving it came from the data.
- **Consequence**: Entity names or membership deviate from what the dataset yields (different spellings, missing/extra countries, wrong boundary member or order), so the expected result file comparison fails and the check scores 0.
207Answer-field formatting not reproduced exactly as specified (quoting/literal form of string values)taskinfiagent-dabench
Applies when
task -- the task dictates a rigid answer template with tagged fields and says certain values must be given as strings (often showing a quoted example), and the script builds the answer text programmatically.
Pattern
The attempt computes the right substantive values but emits them in a different surface form than the template demands — e.g. dropping the quotation marks the example shows, or deriving the literal from data via casting/formatting (f"{int(min)}-{int(max)}", rounding, unit suffixes) instead of writing the exact required token — so an exact-match grader fails fields whose content is actually correct.
Detection procedure
  1. In the task statement, extract every answer field with its exact demanded surface form: tag name, delimiter, whether the value is quoted, and any example literal.
  2. In the script, locate the line(s) that assemble/print the final answer and compare character-by-character to the template: check that quotes appear where the example has them, that all fields use the same quoting convention, and that no field is silently reformatted.
  3. Trace any dynamically derived string field back to its source and ask whether the computation can produce a token differing from the required literal (extra decimals, -0, reversed order, whitespace, different separator).
  4. Confirm the submitted answer string reproduces the template for every field, not just some.
Discriminator
A real violation is an inconsistency with the explicitly shown template (e.g. one field quoted and another not, or a computed literal that need not match the example form); it is not a violation if the task's template genuinely leaves formatting free, or if the derived string is provably identical to the required literal in all cases.
Consequence
The grader's exact string comparison marks correctly-computed fields as WRONG/MISSING, so the attempt scores partial credit (e.g. 1/3) despite sound analysis.
id ac627cb2cb01 · mined from infiagent-dabench dabench-550@s4
raw text (what the judge reads)
### Answer-field formatting not reproduced exactly as specified (quoting/literal form of string values)
- **Applies when**: `task` -- the task dictates a rigid answer template with tagged fields and says certain values must be given as strings (often showing a quoted example), and the script builds the answer text programmatically.
- **Pattern**: The attempt computes the right substantive values but emits them in a different surface form than the template demands — e.g. dropping the quotation marks the example shows, or deriving the literal from data via casting/formatting (`f"{int(min)}-{int(max)}"`, rounding, unit suffixes) instead of writing the exact required token — so an exact-match grader fails fields whose content is actually correct.
- **Detection procedure**:
  1. In the task statement, extract every answer field with its exact demanded surface form: tag name, delimiter, whether the value is quoted, and any example literal.
  2. In the script, locate the line(s) that assemble/print the final answer and compare character-by-character to the template: check that quotes appear where the example has them, that all fields use the same quoting convention, and that no field is silently reformatted.
  3. Trace any dynamically derived string field back to its source and ask whether the computation can produce a token differing from the required literal (extra decimals, `-0`, reversed order, whitespace, different separator).
  4. Confirm the submitted answer string reproduces the template for *every* field, not just some.
- **Discriminator**: A real violation is an inconsistency with the explicitly shown template (e.g. one field quoted and another not, or a computed literal that need not match the example form); it is *not* a violation if the task's template genuinely leaves formatting free, or if the derived string is provably identical to the required literal in all cases.
- **Consequence**: The grader's exact string comparison marks correctly-computed fields as WRONG/MISSING, so the attempt scores partial credit (e.g. 1/3) despite sound analysis.
208Entity substitution when the requested fields aren't present in the datataskda-code
Applies when
task -- the task names specific entities/metrics/stages to compute, and the loaded files appear not to contain columns matching those names.
Pattern
Instead of locating the correct input file or the correct columns (or stopping to report the mismatch), the agent silently redefines the requested concepts as unrelated available columns and produces an output that matches the requested shape but not the requested content; required auxiliary outputs are also skipped.
Detection procedure
1. List every entity, metric, grouping key, and output artifact explicitly named in the task. 2. Read the scripts/answer for the columns and files actually used and check each named concept maps to a real, semantically matching field. 3. Flag any place where the answer states or implies a proxy mapping ("X treated as Y") or where axis/legend labels contradict the described quantities. 4. Verify all requested output files are produced, not just the chart/number.
Discriminator
Legitimate: a documented, semantically equivalent renaming (e.g., a column with a synonym name or a derived quantity that truly equals the requested one). Violation: the substitute measures a different phenomenon, or the agent had to reinterpret the whole task domain to force a fit — the correct data source or columns were likely elsewhere/unsearched.
Consequence
Outputs computed from the wrong variables and mislabeled axes; comparison against expected artifacts fails on all checks, and missing required side outputs count as wrong/missing.
id fc1333cbcb63 · mined from da-code dacode-plot-scatter-002@s4
raw text (what the judge reads)
### Entity substitution when the requested fields aren't present in the data
- **Applies when**: `task` -- the task names specific entities/metrics/stages to compute, and the loaded files appear not to contain columns matching those names.
- **Pattern**: Instead of locating the correct input file or the correct columns (or stopping to report the mismatch), the agent silently redefines the requested concepts as unrelated available columns and produces an output that matches the requested *shape* but not the requested *content*; required auxiliary outputs are also skipped.
- **Detection procedure**: 1. List every entity, metric, grouping key, and output artifact explicitly named in the task. 2. Read the scripts/answer for the columns and files actually used and check each named concept maps to a real, semantically matching field. 3. Flag any place where the answer states or implies a proxy mapping ("X treated as Y") or where axis/legend labels contradict the described quantities. 4. Verify all requested output files are produced, not just the chart/number.
- **Discriminator**: Legitimate: a documented, semantically equivalent renaming (e.g., a column with a synonym name or a derived quantity that truly equals the requested one). Violation: the substitute measures a different phenomenon, or the agent had to reinterpret the whole task domain to force a fit — the correct data source or columns were likely elsewhere/unsearched.
- **Consequence**: Outputs computed from the wrong variables and mislabeled axes; comparison against expected artifacts fails on all checks, and missing required side outputs count as wrong/missing.
209Answer-string format not emitted in fully-closed, parseable formtaskinfiagent-dabench
Applies when
task -- the task prescribes a literal answer template (a tagged key with brackets/braces holding a dict, list, or tuple of per-column/per-class values).
Pattern
The attempt computes correct numbers but copies the template verbatim from the prompt, including truncated or unbalanced delimiters (e.g. an opening [/{ with no matching }/]), omits required quoting/separators, or reorders/renames keys — so an automatic parser cannot extract the object and the answer scores zero despite correct analysis.
Detection procedure
  1. Read the task and write down the exact required wrapper: tag name, opening and closing delimiters, key names, key order, and value types.
  2. Read the final answer string and count/match every bracket, brace, and quote; check that the enclosed literal is valid syntax (i.e. would survive ast.literal_eval after stripping the tag).
  3. Check the scripts/answer step for an explicit serialization + self-parse (or at least a printed final string) rather than hand-typed text pasted from the prompt template.
  4. Confirm all required keys are present, spelled exactly as in the constraints, with values in the requested units/rounding.
Discriminator
A real violation is a structurally unparseable or key-mismatched answer string (unbalanced delimiters, missing keys, stray text inside the object); cosmetic differences that still parse to the same object — spacing, single vs double quotes, trailing newline — are fine.
Consequence
The grader reports the expected key as WRONG/MISSING and 0/1 checks passed even though the submitted numbers equal the ground truth.
id c0cb6dd1ba13 · mined from infiagent-dabench dabench-451@s4
raw text (what the judge reads)
### Answer-string format not emitted in fully-closed, parseable form
- **Applies when**: `task` -- the task prescribes a literal answer template (a tagged key with brackets/braces holding a dict, list, or tuple of per-column/per-class values).
- **Pattern**: The attempt computes correct numbers but copies the template verbatim from the prompt, including truncated or unbalanced delimiters (e.g. an opening `[`/`{` with no matching `}`/`]`), omits required quoting/separators, or reorders/renames keys — so an automatic parser cannot extract the object and the answer scores zero despite correct analysis.
- **Detection procedure**:
  1. Read the task and write down the exact required wrapper: tag name, opening and closing delimiters, key names, key order, and value types.
  2. Read the final answer string and count/match every bracket, brace, and quote; check that the enclosed literal is valid syntax (i.e. would survive `ast.literal_eval` after stripping the tag).
  3. Check the scripts/answer step for an explicit serialization + self-parse (or at least a printed final string) rather than hand-typed text pasted from the prompt template.
  4. Confirm all required keys are present, spelled exactly as in the constraints, with values in the requested units/rounding.
- **Discriminator**: A real violation is a structurally unparseable or key-mismatched answer string (unbalanced delimiters, missing keys, stray text inside the object); cosmetic differences that still parse to the same object — spacing, single vs double quotes, trailing newline — are fine.
- **Consequence**: The grader reports the expected key as WRONG/MISSING and 0/1 checks passed even though the submitted numbers equal the ground truth.
210Ships an unvalidated model that only uses a convenience subset of featurestaskda-code
Applies when
task -- the task asks for predictions on a held-out file and the scripts fit a model, write the output, and stop, without any hold-out/cross-validated error estimate or baseline comparison.
Pattern
The agent hand-picks a narrow slate of "easy" numeric columns (dropping identifier, categorical, date/text, and other metadata columns that are present in both train and test and are often far more predictive), fits one or two off-the-shelf regressors, and reports success based only on descriptive statistics of the predictions (min/max/mean/std) and feature importances. Nothing in the pipeline estimates generalization error, so a model whose predictions are barely better than predicting the global mean (very compressed prediction spread relative to the target's spread) is indistinguishable from a good one.
Detection procedure
  1. Read the task: confirm it is a supervised prediction task scored against hidden ground truth, so accuracy — not just file existence — determines success.
  2. Read the scripts: check for any train/validation split, cross-validation, or metric (MAE/RMSE/R²/accuracy) computed on data the model did not fit, and check for a trivial baseline (predict mean/median/majority) to compare against. Also list which columns shared by train and test were excluded and whether the exclusion was justified (e.g., leakage) or merely convenience.
  3. Read the answer: see whether reported evidence is only prediction summary stats/feature importances; compare the predicted standard deviation to the training target's standard deviation — a large shrinkage with no reported validation score is a red flag for a near-mean-only model.
  4. Flag if no out-of-sample score exists, or if strong shared features were silently dropped without testing whether including them helps.
Discriminator
A real violation is zero out-of-sample evaluation plus unexplained feature omissions. It is not a violation if the agent reports a validation/CV metric (even a mediocre one) beating a stated baseline, or explicitly demonstrates that the omitted columns are unusable (absent from test, constant, or leaky) and that the chosen features were compared against alternatives.
Consequence
The saved file has the right shape and column name but the values are near-constant around the training mean; the grader's error/correlation threshold against true labels fails, and the attempt is scored wrong despite "completing" the deliverable.
id 97fafa09d6c7 · mined from da-code dacode-ml-regression-004@s4
raw text (what the judge reads)
### Ships an unvalidated model that only uses a convenience subset of features
- **Applies when**: `task` -- the task asks for predictions on a held-out file and the scripts fit a model, write the output, and stop, without any hold-out/cross-validated error estimate or baseline comparison.
- **Pattern**: The agent hand-picks a narrow slate of "easy" numeric columns (dropping identifier, categorical, date/text, and other metadata columns that are present in both train and test and are often far more predictive), fits one or two off-the-shelf regressors, and reports success based only on descriptive statistics of the predictions (min/max/mean/std) and feature importances. Nothing in the pipeline estimates generalization error, so a model whose predictions are barely better than predicting the global mean (very compressed prediction spread relative to the target's spread) is indistinguishable from a good one.
- **Detection procedure**:
  1. Read the task: confirm it is a supervised prediction task scored against hidden ground truth, so accuracy — not just file existence — determines success.
  2. Read the scripts: check for any train/validation split, cross-validation, or metric (MAE/RMSE/R²/accuracy) computed on data the model did not fit, and check for a trivial baseline (predict mean/median/majority) to compare against. Also list which columns shared by train and test were excluded and whether the exclusion was justified (e.g., leakage) or merely convenience.
  3. Read the answer: see whether reported evidence is only prediction summary stats/feature importances; compare the predicted standard deviation to the training target's standard deviation — a large shrinkage with no reported validation score is a red flag for a near-mean-only model.
  4. Flag if no out-of-sample score exists, or if strong shared features were silently dropped without testing whether including them helps.
- **Discriminator**: A real violation is zero out-of-sample evaluation plus unexplained feature omissions. It is *not* a violation if the agent reports a validation/CV metric (even a mediocre one) beating a stated baseline, or explicitly demonstrates that the omitted columns are unusable (absent from test, constant, or leaky) and that the chosen features were compared against alternatives.
- **Consequence**: The saved file has the right shape and column name but the values are near-constant around the training mean; the grader's error/correlation threshold against true labels fails, and the attempt is scored wrong despite "completing" the deliverable.
211Choosing the number of clusters by blind argmax of a score, without rejecting degenerate solutionstaskda-code
Applies when
task -- the task asks for an unsupervised grouping "into an appropriate number of groups" and the script sweeps a hyperparameter (k) and picks the value maximizing an internal quality score.
Pattern
The script takes argmax of a metric whose values differ only in the noise range across candidates, ignores the elbow/stability and domain expectation (a small number of interpretable tiers), and accepts the resulting partition even when some groups contain one or a handful of points — i.e. outlier-driven singleton clusters rather than a meaningful segmentation. No sanity check on cluster sizes, no comparison of the top candidates, no outlier/skew handling before scaling.
Detection procedure
  1. Read the task for how the group count should be justified and whether the intended output is a coarse, interpretable segmentation.
  2. In the script, check whether the selection rule is a bare argmax/argmin over the sweep, and whether any tie-breaking, elbow inspection, stability across seeds, or minimum-cluster-size filter is applied.
  3. In the reported scores, check whether the winning candidate beats the runner-ups by a trivial margin (e.g. differences within a few thousandths, or a non-monotone/noisy curve).
  4. In the reported cluster distribution, look for groups whose size is a tiny fraction of n (single-digit counts); confirm the script never validated sizes or re-ran with a smaller/simpler k.
Discriminator
A real violation is when the selected candidate wins by a negligible margin and/or yields degenerate micro-clusters that the script never questions. It is fine if the score has a clear separated maximum (or a clear elbow), cluster sizes are all non-trivial, and the choice is defended against neighbouring candidates — even if the chosen k is unusual.
Consequence
The saved label column encodes an unstable, outlier-driven partition with a group count different from the expected segmentation, so any check comparing the number of clusters, cluster sizes, or label agreement/quality against the reference fails, and the file is scored WRONG.
id 057b41276fae · mined from da-code dacode-ml-cluster-013@s4
raw text (what the judge reads)
### Choosing the number of clusters by blind argmax of a score, without rejecting degenerate solutions
- **Applies when**: `task` -- the task asks for an unsupervised grouping "into an appropriate number of groups" and the script sweeps a hyperparameter (k) and picks the value maximizing an internal quality score.
- **Pattern**: The script takes `argmax` of a metric whose values differ only in the noise range across candidates, ignores the elbow/stability and domain expectation (a small number of interpretable tiers), and accepts the resulting partition even when some groups contain one or a handful of points — i.e. outlier-driven singleton clusters rather than a meaningful segmentation. No sanity check on cluster sizes, no comparison of the top candidates, no outlier/skew handling before scaling.
- **Detection procedure**:
  1. Read the task for how the group count should be justified and whether the intended output is a coarse, interpretable segmentation.
  2. In the script, check whether the selection rule is a bare `argmax`/`argmin` over the sweep, and whether any tie-breaking, elbow inspection, stability across seeds, or minimum-cluster-size filter is applied.
  3. In the reported scores, check whether the winning candidate beats the runner-ups by a trivial margin (e.g. differences within a few thousandths, or a non-monotone/noisy curve).
  4. In the reported cluster distribution, look for groups whose size is a tiny fraction of n (single-digit counts); confirm the script never validated sizes or re-ran with a smaller/simpler k.
- **Discriminator**: A real violation is when the selected candidate wins by a negligible margin and/or yields degenerate micro-clusters that the script never questions. It is fine if the score has a clear separated maximum (or a clear elbow), cluster sizes are all non-trivial, and the choice is defended against neighbouring candidates — even if the chosen k is unusual.
- **Consequence**: The saved label column encodes an unstable, outlier-driven partition with a group count different from the expected segmentation, so any check comparing the number of clusters, cluster sizes, or label agreement/quality against the reference fails, and the file is scored WRONG.
212Statistic computed over the wrong axis/subset, ignoring the qualifier stated in the questiontaskinfiagent-dabench
Applies when
task -- the question names a specific slice (a year, group, region, or field) and/or "all entities in the dataset", and the script computes a summary statistic from a table with entities on one axis and periods/variables on the other.
Pattern
The script aggregates over the convenient axis or over the whole table (e.g., every period for each entity, or only one partial input file) instead of over the exact population the question specifies, so the reported statistic answers a different question than the one asked.
Detection procedure
  1. From the task text, write down explicitly the unit of aggregation (what varies inside each computed statistic) and the filter (which rows/columns/files are in scope, including whether "all" implies loading more than one input file).
  2. In the script, locate the slicing/indexing that feeds the statistic and check that the values passed in correspond exactly to the stated filter, and that the axis matches the stated unit of aggregation.
  3. Verify coverage: does the loaded data span the full entity population named in the task (check file list / row count / key counts), or only a subset?
  4. Check the printed diagnostics — if the vector length or value list shown does not equal the count implied by the task's filter, the statistic is on the wrong subset/axis.
Discriminator
A real violation is when the stated qualifier is never used anywhere in the code (no filter on the named slice, or extra data silently dropped/ignored). It is not a violation if the qualifier is legitimately applied earlier (e.g., data already restricted upstream, or the qualifier defines the grouping key) and the aggregation vector still matches the task's definition.
Consequence
The ranking/argmax is drawn from a different distribution than requested, so the reported entity disagrees with ground truth even though the code runs cleanly and the statistic function itself is used correctly.
id 748ad98d1d07 · mined from infiagent-dabench dabench-252@s4
raw text (what the judge reads)
### Statistic computed over the wrong axis/subset, ignoring the qualifier stated in the question
- **Applies when**: `task` -- the question names a specific slice (a year, group, region, or field) and/or "all entities in the dataset", and the script computes a summary statistic from a table with entities on one axis and periods/variables on the other.
- **Pattern**: The script aggregates over the convenient axis or over the whole table (e.g., every period for each entity, or only one partial input file) instead of over the exact population the question specifies, so the reported statistic answers a different question than the one asked.
- **Detection procedure**:
  1. From the task text, write down explicitly the *unit of aggregation* (what varies inside each computed statistic) and the *filter* (which rows/columns/files are in scope, including whether "all" implies loading more than one input file).
  2. In the script, locate the slicing/indexing that feeds the statistic and check that the values passed in correspond exactly to the stated filter, and that the axis matches the stated unit of aggregation.
  3. Verify coverage: does the loaded data span the full entity population named in the task (check file list / row count / key counts), or only a subset?
  4. Check the printed diagnostics — if the vector length or value list shown does not equal the count implied by the task's filter, the statistic is on the wrong subset/axis.
- **Discriminator**: A real violation is when the stated qualifier is never used anywhere in the code (no filter on the named slice, or extra data silently dropped/ignored). It is *not* a violation if the qualifier is legitimately applied earlier (e.g., data already restricted upstream, or the qualifier defines the grouping key) and the aggregation vector still matches the task's definition.
- **Consequence**: The ranking/argmax is drawn from a different distribution than requested, so the reported entity disagrees with ground truth even though the code runs cleanly and the statistic function itself is used correctly.
213Truncating or coarsening an identified key value to match a loosely-worded format hinttaskinfiagent-dabench
Applies when
task -- the answer requires reporting an identifier (a date, ID, label, key) located by an argmax/argmin or lookup, and the stated answer template shows a format that is less precise than the value actually stored in the data.
Pattern
The agent finds the correct record, then reformats/truncates the identifier to fit the template literally (e.g. dropping the day component, rounding an index, stripping a suffix), discarding information that uniquely identifies the record — even though every downstream computation was done on the full-precision value.
Detection procedure
  1. Read the task and note the granularity of the identifier in the source data versus the granularity implied by the answer template.
  2. In the scripts, find where the identifier is selected (argmax/lookup) and trace any subsequent string formatting, casting, rounding, or .strftime/slicing applied before printing.
  3. Check whether the reported identifier still uniquely picks out the row used for the rest of the computation; if the reported form maps to many rows in the data, it has lost information.
  4. Confirm consistency: the identifier reported must be the same object used to fetch the neighbouring/derived values reported alongside it.
Discriminator
A real violation is when the transformation destroys precision present in the data (many rows share the reported value). It is fine when the data's native granularity genuinely equals the template's (all records are monthly, so a month string is exact), or when the task explicitly says to aggregate to that coarser level.
Consequence
The derived numeric answer matches, but the identifier field is scored WRONG against the full-precision ground truth, so the submission fails on a partial-credit check (1/2 here) despite correct analysis.
id 2c5b04aef20a · mined from infiagent-dabench dabench-572@s4
raw text (what the judge reads)
### Truncating or coarsening an identified key value to match a loosely-worded format hint
- **Applies when**: `task` -- the answer requires reporting an identifier (a date, ID, label, key) located by an argmax/argmin or lookup, and the stated answer template shows a format that is less precise than the value actually stored in the data.
- **Pattern**: The agent finds the correct record, then reformats/truncates the identifier to fit the template literally (e.g. dropping the day component, rounding an index, stripping a suffix), discarding information that uniquely identifies the record — even though every downstream computation was done on the full-precision value.
- **Detection procedure**:
  1. Read the task and note the granularity of the identifier in the source data versus the granularity implied by the answer template.
  2. In the scripts, find where the identifier is selected (argmax/lookup) and trace any subsequent string formatting, casting, rounding, or `.strftime`/slicing applied before printing.
  3. Check whether the reported identifier still uniquely picks out the row used for the rest of the computation; if the reported form maps to many rows in the data, it has lost information.
  4. Confirm consistency: the identifier reported must be the same object used to fetch the neighbouring/derived values reported alongside it.
- **Discriminator**: A real violation is when the transformation *destroys* precision present in the data (many rows share the reported value). It is fine when the data's native granularity genuinely equals the template's (all records are monthly, so a month string is exact), or when the task explicitly says to aggregate to that coarser level.
- **Consequence**: The derived numeric answer matches, but the identifier field is scored WRONG against the full-precision ground truth, so the submission fails on a partial-credit check (1/2 here) despite correct analysis.
214Emitting identifiers with extra quoting/decoration inside the required answer delimiterstaskinfiagent-dabench
Applies when
task -- the task specifies a literal answer template (e.g. @field[item1,item2,...]) whose list items are strings such as labels, IDs, or category names.
Pattern
The attempt computes the right values but formats list items with added syntax — surrounding quotes, brackets, np.str_(...), spaces, or a Python repr/list dump — instead of the bare comma-separated tokens the template shows, so string-valued fields fail exact matching while numeric fields pass.
Detection procedure
  1. Read the task's answer format and note exactly which characters are allowed inside each @field[...]: bare tokens separated by single commas, no quotes, no spaces.
  2. In the scripts, find where the answer string is built and check whether items come from str(x)/joining raw tokens versus print(list_object), f-strings on a list, or json.dumps, all of which inject quotes/brackets.
  3. Compare the submitted answer character-by-character against the template; flag any ", ', nested [], or whitespace around string items.
  4. Confirm the identifiers themselves are the raw values as they appear in the source data (no added prefixes, type wrappers, or normalization).
Discriminator
A real violation is decoration added by the formatting step (quotes/brackets/spaces not present in the underlying identifier); it is not a violation if the special characters, e.g. parentheses, are genuinely part of the identifier value in the data.
Consequence
The numeric field matches but the string-identifier field is scored WRONG/MISSING, giving a partial-credit failure (e.g. 1/2 checks) despite a correct analysis.
id 96ea793f2655 · mined from infiagent-dabench dabench-219@s4
raw text (what the judge reads)
### Emitting identifiers with extra quoting/decoration inside the required answer delimiters
- **Applies when**: `task` -- the task specifies a literal answer template (e.g. `@field[item1,item2,...]`) whose list items are strings such as labels, IDs, or category names.
- **Pattern**: The attempt computes the right values but formats list items with added syntax — surrounding quotes, brackets, `np.str_(...)`, spaces, or a Python `repr`/`list` dump — instead of the bare comma-separated tokens the template shows, so string-valued fields fail exact matching while numeric fields pass.
- **Detection procedure**:
  1. Read the task's answer format and note exactly which characters are allowed inside each `@field[...]`: bare tokens separated by single commas, no quotes, no spaces.
  2. In the scripts, find where the answer string is built and check whether items come from `str(x)`/joining raw tokens versus `print(list_object)`, f-strings on a list, or `json.dumps`, all of which inject quotes/brackets.
  3. Compare the submitted answer character-by-character against the template; flag any `"`, `'`, nested `[]`, or whitespace around string items.
  4. Confirm the identifiers themselves are the raw values as they appear in the source data (no added prefixes, type wrappers, or normalization).
- **Discriminator**: A real violation is decoration added by the formatting step (quotes/brackets/spaces not present in the underlying identifier); it is *not* a violation if the special characters, e.g. parentheses, are genuinely part of the identifier value in the data.
- **Consequence**: The numeric field matches but the string-identifier field is scored WRONG/MISSING, giving a partial-credit failure (e.g. 1/2 checks) despite a correct analysis.
215Predictions emitted without shape/prevalence sanity check against the stated cost‑sensitive objectivetaskda-code
Applies when
task -- the task asks for a per-row prediction file for a held-out test set (especially a rare-event/imbalanced target where the stated objective penalizes missed positives more than false alarms).
Pattern
The agent produces a label column with no saved, re-runnable script and no verification that (a) the number of predicted rows equals the number of test rows, (b) the header/format matches the provided sample file, and (c) the predicted positive rate is plausible given the training prevalence and the asymmetric error costs — typically a default 0.5 threshold on an imbalanced model yields far too few positives, or the file is silently truncated.
Detection procedure
  1. Read the task: note the required file name, exact header/column, row count implied by the test file, and any stated asymmetry in error costs (e.g., missed positives are the expensive error).
  2. Read the scripts: check that they exist and are reproducible, that predictions are generated on the full test set in original row order, and that the decision threshold / class weighting was chosen against a validation set using the cost-aligned metric (recall/F-beta/expected cost), not left at the library default.
  3. Inspect the answer file: count data rows and compare with the test-set row count; confirm header matches the sample exactly; compute the fraction of positives.
  4. Flag if row count differs, header/values deviate from the sample format, or the positive rate is drastically below the training prevalence with no documented threshold-tuning justification.
Discriminator
A low positive rate is fine if the script shows the threshold was tuned on held-out data under the stated cost/metric and the row count and format check out; it is a violation when the rate simply reflects an untuned default, or when row count/header cannot be shown to match the required output spec.
Consequence
The grader compares the file row-by-row against the expected output and marks it WRONG/MISSING — either due to a length/format mismatch or because the recall-oriented score collapses when nearly all positives are predicted as negative.
id 7e461eb78929 · mined from da-code dacode-ml-binary-013@s4
raw text (what the judge reads)
### Predictions emitted without shape/prevalence sanity check against the stated cost‑sensitive objective
- **Applies when**: `task` -- the task asks for a per-row prediction file for a held-out test set (especially a rare-event/imbalanced target where the stated objective penalizes missed positives more than false alarms).
- **Pattern**: The agent produces a label column with no saved, re-runnable script and no verification that (a) the number of predicted rows equals the number of test rows, (b) the header/format matches the provided sample file, and (c) the predicted positive rate is plausible given the training prevalence and the asymmetric error costs — typically a default 0.5 threshold on an imbalanced model yields far too few positives, or the file is silently truncated.
- **Detection procedure**:
  1. Read the task: note the required file name, exact header/column, row count implied by the test file, and any stated asymmetry in error costs (e.g., missed positives are the expensive error).
  2. Read the scripts: check that they exist and are reproducible, that predictions are generated on the full test set in original row order, and that the decision threshold / class weighting was chosen against a validation set using the cost-aligned metric (recall/F-beta/expected cost), not left at the library default.
  3. Inspect the answer file: count data rows and compare with the test-set row count; confirm header matches the sample exactly; compute the fraction of positives.
  4. Flag if row count differs, header/values deviate from the sample format, or the positive rate is drastically below the training prevalence with no documented threshold-tuning justification.
- **Discriminator**: A low positive rate is fine if the script shows the threshold was tuned on held-out data under the stated cost/metric and the row count and format check out; it is a violation when the rate simply reflects an untuned default, or when row count/header cannot be shown to match the required output spec.
- **Consequence**: The grader compares the file row-by-row against the expected output and marks it WRONG/MISSING — either due to a length/format mismatch or because the recall-oriented score collapses when nearly all positives are predicted as negative.
216Analysis run on a partial data slice when the task specifies the full populationtaskinfiagent-dabench
Applies when
task -- the question asks for a statistic or set of entities over an entire population ("all X"), and the scripts load a single file/subset out of a directory that contains several partitions of the data.
Pattern
The agent hard-codes one input file (one region/segment/split) without ever checking what other data files or rows exist, then computes quartiles/thresholds/rankings from that subset and reports them as if they applied to the whole population. It also never states the number of entities analyzed against the number implied by the task.
Detection procedure
  1. Read the task and note the scope words ("all countries", "the whole dataset", "every record") and the expected population size or breadth.
  2. Read the scripts: identify every data-loading call; check whether the agent enumerated the available data files/tables (e.g., listing the data directory) or simply picked one, and whether any concatenation/merge covers the remaining partitions.
  3. Compare the row/entity count printed by the script (or implied by the file) with the scope in the task; if the loaded set is clearly one partition of a larger whole, flag it.
  4. Check the reported answer format character-for-character against the requested template (delimiters, quoting, ordering, separators) before accepting it.
Discriminator
A real violation is when other partitions of the same variable exist (or the file name/shape shows a subset) and were never inspected or combined; it is fine if the agent verified that the single file is the complete population, or if the task itself restricts the scope to that subset. Note that a subset answer can coincidentally overlap the true answer — the absence of any coverage check is the violation, not just a differing result.
Consequence
Threshold-based or rank-based results (quartiles, cutoffs, top-k, outliers) are computed from the wrong distribution, so the reported entity list is incomplete or spurious and the grader marks the expected value WRONG/MISSING, even when part of the list happens to match.
id 96b5db3f0767 · mined from infiagent-dabench dabench-254@s4
raw text (what the judge reads)
### Analysis run on a partial data slice when the task specifies the full population
- **Applies when**: `task` -- the question asks for a statistic or set of entities over an entire population ("all X"), and the scripts load a single file/subset out of a directory that contains several partitions of the data.
- **Pattern**: The agent hard-codes one input file (one region/segment/split) without ever checking what other data files or rows exist, then computes quartiles/thresholds/rankings from that subset and reports them as if they applied to the whole population. It also never states the number of entities analyzed against the number implied by the task.
- **Detection procedure**:
  1. Read the task and note the scope words ("all countries", "the whole dataset", "every record") and the expected population size or breadth.
  2. Read the scripts: identify every data-loading call; check whether the agent enumerated the available data files/tables (e.g., listing the data directory) or simply picked one, and whether any concatenation/merge covers the remaining partitions.
  3. Compare the row/entity count printed by the script (or implied by the file) with the scope in the task; if the loaded set is clearly one partition of a larger whole, flag it.
  4. Check the reported answer format character-for-character against the requested template (delimiters, quoting, ordering, separators) before accepting it.
- **Discriminator**: A real violation is when other partitions of the same variable exist (or the file name/`shape` shows a subset) and were never inspected or combined; it is fine if the agent verified that the single file is the complete population, or if the task itself restricts the scope to that subset. Note that a subset answer can coincidentally overlap the true answer — the absence of any coverage check is the violation, not just a differing result.
- **Consequence**: Threshold-based or rank-based results (quartiles, cutoffs, top-k, outliers) are computed from the wrong distribution, so the reported entity list is incomplete or spurious and the grader marks the expected value WRONG/MISSING, even when part of the list happens to match.
217Declaring success without a held-out validation score or output verificationtaskda-code
Applies when
task -- the task asks for predictions/derived values written to a specific output file, and the scripts train a model and write that file in one pass.
Pattern
The attempt reports a pipeline description and prediction class counts as evidence of success, but never (a) holds out part of the labeled data to measure accuracy/appropriate metric, nor (b) re-reads the written file to confirm its path, row count, column name, ordering relative to the input rows, and that the emitted label strings/dtype exactly match the label vocabulary/format seen in the training target. Preprocessing choices (encoders fit separately per split, scalers, imputation, dropped/reordered columns) are therefore never checked for consistency between train and test.
Detection procedure
  1. Read the task for the required output artifact: exact filename/location, column name, expected number of rows, and the value format implied by the labeled data.
  2. Scan the scripts for any validation split or cross-validation producing a numeric score on labeled data; if absent, and no baseline comparison exists, the attempt has no evidence the model works.
  3. Scan for a post-write verification block that loads the output file and asserts shape, column name, row alignment with the test input, and membership of predicted values in the training label set.
  4. Check that any encoders/imputers/scalers are fit on training data only and applied with identical column order to test; a per-split fit_transform or independently built category mappings is a silent mismatch.
Discriminator
A fine attempt may skip elaborate tuning but still prints a measured validation metric and an explicit file/format sanity check; a violation is one where the only "evidence" is descriptive text and a class-count breakdown, which cannot detect label mismatch, row misalignment, wrong path, or a broken encoding.
Consequence
The grader reads the output file and finds it missing at the expected path or mismatched in labels/order/row count, scoring 0 even though the agent reports "TASK COMPLETED SUCCESSFULLY".
id 89f10213221c · mined from da-code dacode-ml-binary-009@s4
raw text (what the judge reads)
### Declaring success without a held-out validation score or output verification
- **Applies when**: `task` -- the task asks for predictions/derived values written to a specific output file, and the scripts train a model and write that file in one pass.
- **Pattern**: The attempt reports a pipeline description and prediction class counts as evidence of success, but never (a) holds out part of the labeled data to measure accuracy/appropriate metric, nor (b) re-reads the written file to confirm its path, row count, column name, ordering relative to the input rows, and that the emitted label strings/dtype exactly match the label vocabulary/format seen in the training target. Preprocessing choices (encoders fit separately per split, scalers, imputation, dropped/reordered columns) are therefore never checked for consistency between train and test.
- **Detection procedure**:
  1. Read the task for the required output artifact: exact filename/location, column name, expected number of rows, and the value format implied by the labeled data.
  2. Scan the scripts for any validation split or cross-validation producing a numeric score on labeled data; if absent, and no baseline comparison exists, the attempt has no evidence the model works.
  3. Scan for a post-write verification block that loads the output file and asserts shape, column name, row alignment with the test input, and membership of predicted values in the training label set.
  4. Check that any encoders/imputers/scalers are fit on training data only and applied with identical column order to test; a per-split `fit_transform` or independently built category mappings is a silent mismatch.
- **Discriminator**: A fine attempt may skip elaborate tuning but still prints a measured validation metric and an explicit file/format sanity check; a violation is one where the only "evidence" is descriptive text and a class-count breakdown, which cannot detect label mismatch, row misalignment, wrong path, or a broken encoding.
- **Consequence**: The grader reads the output file and finds it missing at the expected path or mismatched in labels/order/row count, scoring 0 even though the agent reports "TASK COMPLETED SUCCESSFULLY".
218Fabricated metric definition computed on only half the relevant recordstaskda-code
Applies when
task -- the task asks to visualize/report "performance" (or any composite score) whose formula is not spelled out in the prompt, and the script must aggregate per-entity records that appear in more than one role/column (or across more than one file).
Pattern
The agent invents an arbitrary formula (e.g., count + 2 * some_flag_count) with no basis in the task text, config file, or upstream artifacts, and computes it by grouping on a single role column, so every record where the entity appears in the other role is silently dropped; it also skips the auxiliary result artifacts the task's output convention implies.
Detection procedure
  1. Read the task and any provided config/spec files for the definition of the quantity to be plotted (axis labels, prior-step outputs, saved intermediates); note whether the agent's formula is stated anywhere.
  2. In the script, check the aggregation key: does the entity of interest occur in multiple columns/tables, and does the groupby cover all of them (plus correct dtype comparisons for flags, e.g. boolean vs. the string 'TRUE')?
  3. Sanity-check magnitudes in the answer: are per-entity counts roughly half (or otherwise inconsistent with) the totals printed for the filtered data, and does the resulting ranking contradict domain-obvious expectations?
  4. Confirm every output artifact the task/grader convention requires (figure plus any serialized values/plot metadata) is actually written.
Discriminator
A real violation is a formula with no source in the prompt/config and an aggregation that provably ignores records where the entity appears under another column; it is fine if the metric is explicitly defined by the task/config, or if the single-column grouping is exactly what the definition requires (and the script demonstrates the other role contributes nothing).
Consequence
The plotted values and any saved numeric array differ from the reference, so the figure and all value-based checks fail, and missing auxiliary files fail outright regardless of the chart's cosmetic correctness.
id 2984e0978047 · mined from da-code dacode-plot-bar-006@s4
raw text (what the judge reads)
### Fabricated metric definition computed on only half the relevant records
- **Applies when**: `task` -- the task asks to visualize/report "performance" (or any composite score) whose formula is not spelled out in the prompt, and the script must aggregate per-entity records that appear in more than one role/column (or across more than one file).
- **Pattern**: The agent invents an arbitrary formula (e.g., `count + 2 * some_flag_count`) with no basis in the task text, config file, or upstream artifacts, and computes it by grouping on a single role column, so every record where the entity appears in the other role is silently dropped; it also skips the auxiliary result artifacts the task's output convention implies.
- **Detection procedure**:
  1. Read the task and any provided config/spec files for the definition of the quantity to be plotted (axis labels, prior-step outputs, saved intermediates); note whether the agent's formula is stated anywhere.
  2. In the script, check the aggregation key: does the entity of interest occur in multiple columns/tables, and does the groupby cover all of them (plus correct dtype comparisons for flags, e.g. boolean vs. the string `'TRUE'`)?
  3. Sanity-check magnitudes in the answer: are per-entity counts roughly half (or otherwise inconsistent with) the totals printed for the filtered data, and does the resulting ranking contradict domain-obvious expectations?
  4. Confirm every output artifact the task/grader convention requires (figure plus any serialized values/plot metadata) is actually written.
- **Discriminator**: A real violation is a formula with no source in the prompt/config and an aggregation that provably ignores records where the entity appears under another column; it is fine if the metric is explicitly defined by the task/config, or if the single-column grouping is exactly what the definition requires (and the script demonstrates the other role contributes nothing).
- **Consequence**: The plotted values and any saved numeric array differ from the reference, so the figure and all value-based checks fail, and missing auxiliary files fail outright regardless of the chart's cosmetic correctness.
219Threshold-based outlier count not reconciled with the extreme-value sanity checktaskinfiagent-dabench
Applies when
task -- The task asks for a count of rows flagged by an explicit statistical rule with a fixed cutoff (e.g., a standardized-score, quantile, or distance threshold) computed on a single numeric column.
Pattern
The script applies some flagging rule but never verifies that the standardization matches the stated definition (population vs. sample SD, mean/SD computed on the correct subset after NaN/dtype cleaning, raw values vs. transformed/scaled values) and never prints the extreme statistics of the score; a nonzero-looking count is reported without checking whether the distribution can even exceed the cutoff. Common variants: using a different rule than stated (IQR, modified/median-based z, percentile), using a lower effective cutoff, standardizing groupwise or per-chunk, or counting flagged values in the wrong column/axis.
Detection procedure
  1. Read the task and write down the exact rule and cutoff, and which quantity the count refers to (rows vs. values, before or after cleaning).
  2. In the script, locate where the score is computed: confirm the mean/SD (or quantiles) come from the full cleaned column of interest, that the same column is thresholded, and that the comparison uses the stated cutoff with the stated sign convention (two-sided if "higher than X or lower than -X").
  3. Check that the script prints diagnostics that bound the answer — column dtype, non-null count, min/max of the score (or min/max of the raw value), and the count itself — and that the reported count is consistent with them (e.g., if max |score| < cutoff, the count must be 0).
  4. Compare the final reported integer to those diagnostics and to the requested output format; a count reported with no printed extreme-score evidence is unverified.
Discriminator
A real violation is when the rule/cutoff in code differs from the task's, or when no diagnostic exists that could distinguish "several outliers" from "none" — including the case where the reported count contradicts the printed extremes. It is not a violation if the script implements the stated rule exactly on the correct cleaned column and prints extremes that are consistent with a nonzero count; large counts are legitimate for genuinely heavy-tailed data.
Consequence
The reported count is off (often nonzero when the correct answer is zero, or vice versa), so the single numeric check fails and the task scores 0.
id 7048f0eb8387 · mined from infiagent-dabench dabench-361@s4
raw text (what the judge reads)
### Threshold-based outlier count not reconciled with the extreme-value sanity check
- **Applies when**: `task` -- The task asks for a count of rows flagged by an explicit statistical rule with a fixed cutoff (e.g., a standardized-score, quantile, or distance threshold) computed on a single numeric column.
- **Pattern**: The script applies some flagging rule but never verifies that the standardization matches the stated definition (population vs. sample SD, mean/SD computed on the correct subset after NaN/dtype cleaning, raw values vs. transformed/scaled values) and never prints the extreme statistics of the score; a nonzero-looking count is reported without checking whether the distribution can even exceed the cutoff. Common variants: using a different rule than stated (IQR, modified/median-based z, percentile), using a lower effective cutoff, standardizing groupwise or per-chunk, or counting flagged values in the wrong column/axis.
- **Detection procedure**:
  1. Read the task and write down the exact rule and cutoff, and which quantity the count refers to (rows vs. values, before or after cleaning).
  2. In the script, locate where the score is computed: confirm the mean/SD (or quantiles) come from the full cleaned column of interest, that the same column is thresholded, and that the comparison uses the stated cutoff with the stated sign convention (two-sided if "higher than X or lower than -X").
  3. Check that the script prints diagnostics that bound the answer — column dtype, non-null count, min/max of the score (or min/max of the raw value), and the count itself — and that the reported count is consistent with them (e.g., if max |score| < cutoff, the count must be 0).
  4. Compare the final reported integer to those diagnostics and to the requested output format; a count reported with no printed extreme-score evidence is unverified.
- **Discriminator**: A real violation is when the rule/cutoff in code differs from the task's, or when no diagnostic exists that could distinguish "several outliers" from "none" — including the case where the reported count contradicts the printed extremes. It is *not* a violation if the script implements the stated rule exactly on the correct cleaned column and prints extremes that are consistent with a nonzero count; large counts are legitimate for genuinely heavy-tailed data.
- **Consequence**: The reported count is off (often nonzero when the correct answer is zero, or vice versa), so the single numeric check fails and the task scores 0.
220Hard-coding a mapping/spec from assumption instead of reading the referenced auxiliary filetaskda-code
Applies when
task -- The prompt tells the agent to apply a definition, label mapping, filter, or parameter set that is documented in an auxiliary file (README, tips/notes, data dictionary, config) rather than in the prompt itself.
Pattern
The script never opens or quotes the referenced file; instead it embeds a mapping/threshold/category set the agent recalled from general knowledge or inferred from the raw column values, then computes the requested statistic on those self-invented labels (and often skips writing the required output artifact too).
Detection procedure
  1. From the task text, list every external artifact named as the source of truth (e.g. "use the mapping from <file>") and every required output artifact/format.
  2. Grep the scripts for a read/open/parse of each named source file; check that the transformation values actually originate from it rather than from a literal dict/list defined in code.
  3. If values are literals, check whether the script at least prints the file's contents or otherwise verifies the literals match the spec; absent that, the mapping is unverified — also confirm the required output file is written, not just printed.
  4. Compare the reported label strings/granularity to what the spec plausibly requires (e.g. could the spec collapse or rename categories, changing which is most frequent or its ratio?).
Discriminator
A real violation is when the spec file is never read and never echoed, so nothing in the run confirms the literals match it. It is fine if the agent read the file (printed/cat'd it) and then hard-coded the identical mapping, or if the task itself fully specifies the mapping inline.
Consequence
Labels (and therefore the modal category name, and possibly the counts/ratio if categories are merged) fail exact-match comparison against the expected values, and a missing required result file scores 0 regardless of the numbers.
id 77e25e094030 · mined from da-code dacode-di-text-004@s4
raw text (what the judge reads)
### Hard-coding a mapping/spec from assumption instead of reading the referenced auxiliary file
- **Applies when**: `task` -- The prompt tells the agent to apply a definition, label mapping, filter, or parameter set that is documented in an auxiliary file (README, tips/notes, data dictionary, config) rather than in the prompt itself.
- **Pattern**: The script never opens or quotes the referenced file; instead it embeds a mapping/threshold/category set the agent recalled from general knowledge or inferred from the raw column values, then computes the requested statistic on those self-invented labels (and often skips writing the required output artifact too).
- **Detection procedure**:
  1. From the task text, list every external artifact named as the source of truth (e.g. "use the mapping from <file>") and every required output artifact/format.
  2. Grep the scripts for a read/open/parse of each named source file; check that the transformation values actually originate from it rather than from a literal dict/list defined in code.
  3. If values are literals, check whether the script at least prints the file's contents or otherwise verifies the literals match the spec; absent that, the mapping is unverified — also confirm the required output file is written, not just printed.
  4. Compare the reported label strings/granularity to what the spec plausibly requires (e.g. could the spec collapse or rename categories, changing which is most frequent or its ratio?).
- **Discriminator**: A real violation is when the spec file is never read and never echoed, so nothing in the run confirms the literals match it. It is fine if the agent read the file (printed/cat'd it) and then hard-coded the identical mapping, or if the task itself fully specifies the mapping inline.
- **Consequence**: Labels (and therefore the modal category name, and possibly the counts/ratio if categories are merged) fail exact-match comparison against the expected values, and a missing required result file scores 0 regardless of the numbers.
221Submission file not validated against the provided template schemataskda-code
Applies when
task -- The task supplies a sample/template output file and requires writing predictions to a specified output file matching that format.
Pattern
The attempt builds the output from its own assumptions (hand-typed header names, capitalization, column order, or ids taken from the wrong source) and writes/reports it without ever loading the template file and comparing header text, row count, id set and ordering, and value ranges against it.
Detection procedure
  1. Read the task/README for the required output filename and the template file that defines its schema.
  2. In the scripts, check whether the template is actually read and used to construct or verify the output (header strings compared exactly including case, ids taken from the test/template file, column order preserved), and whether the file is written to the exact required path.
  3. Compare the produced answer's first line and shape to the template: identical column names/case/order, one row per test id in the template's order, no extra index column, and values in the valid range (e.g., probabilities summing to 1 per row where required).
  4. Flag if any assertion/sanity check on row count, id match, or header equality is absent.
Discriminator
A real violation is a header/case/column-order/id-coverage/path mismatch or absence of any verification against the template; it is fine if the script constructs the frame from the template's own ids and columns (or asserts equality with it), even if it never prints the check.
Consequence
The grader reads the expected file's schema and reports the submission as WRONG/MISSING (unparseable or mismatched ids/columns), scoring 0 regardless of how good the underlying model was.
id a2d4f806de2a · mined from da-code dacode-ml-competition-003@s4
raw text (what the judge reads)
### Submission file not validated against the provided template schema
- **Applies when**: `task` -- The task supplies a sample/template output file and requires writing predictions to a specified output file matching that format.
- **Pattern**: The attempt builds the output from its own assumptions (hand-typed header names, capitalization, column order, or ids taken from the wrong source) and writes/reports it without ever loading the template file and comparing header text, row count, id set and ordering, and value ranges against it.
- **Detection procedure**:
  1. Read the task/README for the required output filename and the template file that defines its schema.
  2. In the scripts, check whether the template is actually read and used to construct or verify the output (header strings compared exactly including case, ids taken from the test/template file, column order preserved), and whether the file is written to the exact required path.
  3. Compare the produced answer's first line and shape to the template: identical column names/case/order, one row per test id in the template's order, no extra index column, and values in the valid range (e.g., probabilities summing to 1 per row where required).
  4. Flag if any assertion/sanity check on row count, id match, or header equality is absent.
- **Discriminator**: A real violation is a header/case/column-order/id-coverage/path mismatch or absence of any verification against the template; it is fine if the script constructs the frame from the template's own ids and columns (or asserts equality with it), even if it never prints the check.
- **Consequence**: The grader reads the expected file's schema and reports the submission as WRONG/MISSING (unparseable or mismatched ids/columns), scoring 0 regardless of how good the underlying model was.
222Substituting heuristic pseudo-labels and an ad-hoc row set for the specified training labels and test filetaskda-code
Applies when
task -- the task names specific input/output files and a target column, and the scripts must train on the provided labeled data and predict for exactly the provided evaluation rows.
Pattern
The agent claims the labeled training data or the designated test file is "missing", invents the target values with hand-written rules (thresholds, keyword/category maps), fits a model on those self-generated labels, and then writes predictions for whatever subset of rows it could build features for (often all rows, or only rows present in some auxiliary file) instead of the exact evaluation rows in the required order/format.
Detection procedure
  1. From the task, list the required inputs (labeled source, evaluation file), the required output filename/columns, and the expected number of prediction rows.
  2. In the scripts, check whether the target column is ever read from the data; flag any code that constructs the target from if/elif rules or keyword dictionaries, and any os.path.exists(...) fallback that silently proceeds when the specified file is absent.
  3. Check how the output row set is derived: does it iterate over the evaluation file's identifiers exactly (no continue/filter dropping rows, no reliance on auxiliary tables), and does it preserve count and order?
  4. Compare the answer's reported row count and label vocabulary against the expected evaluation size and the label values actually used in the dataset (spelling, casing, category set).
Discriminator
A real violation is when labels are synthesized rather than learned from provided ground truth, or when output rows do not correspond 1:1 to the evaluation rows. It is fine to use rule-based features or a heuristic baseline if the model is fit/validated against genuine labels from the data and predictions cover exactly the required rows; it is also fine to search additional paths for a file that truly exists elsewhere.
Consequence
The grader finds result.csv mismatched in row count/identifiers and label categories, and accuracy against true labels is near chance, so the file is marked WRONG/MISSING despite a confident "task completed" report.
id a6fc23d3a982 · mined from da-code dacode-ml-multi-003@s4
raw text (what the judge reads)
### Substituting heuristic pseudo-labels and an ad-hoc row set for the specified training labels and test file
- **Applies when**: `task` -- the task names specific input/output files and a target column, and the scripts must train on the provided labeled data and predict for exactly the provided evaluation rows.
- **Pattern**: The agent claims the labeled training data or the designated test file is "missing", invents the target values with hand-written rules (thresholds, keyword/category maps), fits a model on those self-generated labels, and then writes predictions for whatever subset of rows it could build features for (often all rows, or only rows present in some auxiliary file) instead of the exact evaluation rows in the required order/format.
- **Detection procedure**:
  1. From the task, list the required inputs (labeled source, evaluation file), the required output filename/columns, and the expected number of prediction rows.
  2. In the scripts, check whether the target column is ever read from the data; flag any code that constructs the target from `if/elif` rules or keyword dictionaries, and any `os.path.exists(...)` fallback that silently proceeds when the specified file is absent.
  3. Check how the output row set is derived: does it iterate over the evaluation file's identifiers exactly (no `continue`/filter dropping rows, no reliance on auxiliary tables), and does it preserve count and order?
  4. Compare the answer's reported row count and label vocabulary against the expected evaluation size and the label values actually used in the dataset (spelling, casing, category set).
- **Discriminator**: A real violation is when labels are synthesized rather than learned from provided ground truth, or when output rows do not correspond 1:1 to the evaluation rows. It is fine to use rule-based features or a heuristic baseline *if* the model is fit/validated against genuine labels from the data and predictions cover exactly the required rows; it is also fine to search additional paths for a file that truly exists elsewhere.
- **Consequence**: The grader finds `result.csv` mismatched in row count/identifiers and label categories, and accuracy against true labels is near chance, so the file is marked WRONG/MISSING despite a confident "task completed" report.
223Output template provided but never read or validated againsttaskda-code
Applies when
task -- The task says the result file's format must match a provided template/example file, and the scripts write the output file.
Pattern
The agent never opens or parses the template; it guesses the header names, index labels, row/column ordering, dtype, precision, and any implied filtering/derivation conventions, then asserts "format matches" without a comparison step.
Detection procedure
  1. In the task statement, note that a template/reference file is supplied and list what it constrains (column names/count, index labeling, row order, value precision, null representation).
  2. Search the scripts for any read of that template (e.g., loading it, printing its head, comparing headers/shape); if absent, the format claim is unverified.
  3. Check whether formatting decisions in the script are hard-coded guesses (renaming columns to invented labels, ad-hoc date string formats, arbitrary rounding, dropping/keeping rows without justification) rather than derived from the template.
  4. Confirm the answer text states format compliance without showing a header-by-header / shape-by-shape diff against the template.
Discriminator
A real violation is guessing the schema while a template exists; it is fine if the script explicitly loads the template and asserts equality of columns/shape/index (or reindexes the result onto the template's structure), even if the printed summary is brief.
Consequence
The written file differs from the expected one in headers, key labels, ordering, or rounding, so an exact/structural file comparison fails even when the underlying aggregation logic is defensible.
id c412df81f023 · mined from da-code dacode-dm-csv-044@s4
raw text (what the judge reads)
### Output template provided but never read or validated against
- **Applies when**: `task` -- The task says the result file's format must match a provided template/example file, and the scripts write the output file.
- **Pattern**: The agent never opens or parses the template; it guesses the header names, index labels, row/column ordering, dtype, precision, and any implied filtering/derivation conventions, then asserts "format matches" without a comparison step.
- **Detection procedure**:
  1. In the task statement, note that a template/reference file is supplied and list what it constrains (column names/count, index labeling, row order, value precision, null representation).
  2. Search the scripts for any read of that template (e.g., loading it, printing its head, comparing headers/shape); if absent, the format claim is unverified.
  3. Check whether formatting decisions in the script are hard-coded guesses (renaming columns to invented labels, ad-hoc date string formats, arbitrary rounding, dropping/keeping rows without justification) rather than derived from the template.
  4. Confirm the answer text states format compliance without showing a header-by-header / shape-by-shape diff against the template.
- **Discriminator**: A real violation is guessing the schema while a template exists; it is fine if the script explicitly loads the template and asserts equality of columns/shape/index (or reindexes the result onto the template's structure), even if the printed summary is brief.
- **Consequence**: The written file differs from the expected one in headers, key labels, ordering, or rounding, so an exact/structural file comparison fails even when the underlying aggregation logic is defensible.
224Optimizing/selecting on a proxy metric instead of the task's stated evaluation metrictaskda-code
Applies when
task -- the task specifies an explicit scoring metric (e.g., a weighted/ordinal agreement, ranking, or cost-sensitive score) and the scripts train, tune, cross-validate, weight ensembles, or choose a decision rule for the final predictions.
Pattern
The scripts never compute the stated metric anywhere; they score candidate models with a convenient default (plain accuracy/log-loss), weight or pick the ensemble by that default, and take plain argmax over class probabilities, ignoring the metric's structure (e.g., ordering of classes, asymmetric penalties, tie-breaking) — so no evidence exists that the submitted predictions are good under the actual grading rule.
Detection procedure
  1. Read the task/README and note the exact metric name and its properties (ordered classes? asymmetric error costs? rank-based?).
  2. Grep the scripts for that metric or an equivalent scorer used inside cross_val_score/GridSearchCV/manual holdout evaluation, and for any comparison of alternative prediction strategies under it.
  3. Check what quantity the printed model-selection numbers and ensemble weights are derived from, and how probabilities are converted to final labels.
  4. Inspect the answer: if predictions collapse onto the few majority classes and no metric-specific validation number was ever reported, flag it.
Discriminator
A real violation is when the stated metric is never estimated on held-out data, so selection and thresholding rest solely on a proxy. It is fine if the agent computes the stated metric on a validation split (or a custom scorer) and shows the proxy-chosen model also wins under it, or explicitly justifies the proxy as equivalent for this metric.
Consequence
The submission is technically well-formatted but scores far below achievable under the real metric (near-chance/negative agreement for ordinal-weighted scores), so the grader marks the expected result file wrong.
id a8c53c2bfb5e · mined from da-code dacode-ml-competition-006@s4
raw text (what the judge reads)
### Optimizing/selecting on a proxy metric instead of the task's stated evaluation metric
- **Applies when**: `task` -- the task specifies an explicit scoring metric (e.g., a weighted/ordinal agreement, ranking, or cost-sensitive score) and the scripts train, tune, cross-validate, weight ensembles, or choose a decision rule for the final predictions.
- **Pattern**: The scripts never compute the stated metric anywhere; they score candidate models with a convenient default (plain accuracy/log-loss), weight or pick the ensemble by that default, and take plain argmax over class probabilities, ignoring the metric's structure (e.g., ordering of classes, asymmetric penalties, tie-breaking) — so no evidence exists that the submitted predictions are good under the actual grading rule.
- **Detection procedure**:
  1. Read the task/README and note the exact metric name and its properties (ordered classes? asymmetric error costs? rank-based?).
  2. Grep the scripts for that metric or an equivalent scorer used inside `cross_val_score`/`GridSearchCV`/manual holdout evaluation, and for any comparison of alternative prediction strategies under it.
  3. Check what quantity the printed model-selection numbers and ensemble weights are derived from, and how probabilities are converted to final labels.
  4. Inspect the answer: if predictions collapse onto the few majority classes and no metric-specific validation number was ever reported, flag it.
- **Discriminator**: A real violation is when the stated metric is never estimated on held-out data, so selection and thresholding rest solely on a proxy. It is fine if the agent computes the stated metric on a validation split (or a custom scorer) and shows the proxy-chosen model also wins under it, or explicitly justifies the proxy as equivalent for this metric.
- **Consequence**: The submission is technically well-formatted but scores far below achievable under the real metric (near-chance/negative agreement for ordinal-weighted scores), so the grader marks the expected result file wrong.
225Null/missing-value partition defined inconsistently, so group subsets (and their statistics) don't reconstitute the full datasettaskinfiagent-dabench
Applies when
task -- the task asks to split rows into "missing" vs "non-missing" (or any complementary filter) on one column and compute per-group statistics or a test on another column.
Pattern
The attempt uses an ad-hoc notion of "null" (e.g., only NaN while string sentinels like "", "NA", "None", "-", "null" remain, or the reverse), or silently loses rows via dropna() on the whole frame, read_csv defaults/na_values, dtype coercion, deduplication, or a prior filter — so the two group sizes don't add up to the row count and the group means drift from the true values.
Detection procedure
  1. From the task, note that the two groups are complementary and must together cover every row of the source table.
  2. In the scripts, find where the data is loaded and where the mask is built; check for global dropna()/fillna, subsetting, type casts, or a mask that only tests one form of missingness.
  3. Confirm the script prints len(group_a), len(group_b), and len(df) and asserts they sum; also check the value column's own missing values are handled explicitly and identically for both groups.
  4. Sanity-check the reported statistics against these counts/ranges; if counts were never printed, the partition is unverified.
Discriminator
A real violation is when rows can fall outside both groups (or are dropped before masking) or when only one missingness representation is recognized; it is fine if the script explicitly normalizes sentinels to NaN, documents any exclusion, and shows the two group sizes summing to the full row count.
Consequence
Both group means (and the test statistic) are computed on the wrong subsets, so the reported numbers miss the expected values even though the p-value may still look decisively significant, and the grader marks the mean checks wrong.
id 62a52c37f887 · mined from infiagent-dabench dabench-297@s4
raw text (what the judge reads)
### Null/missing-value partition defined inconsistently, so group subsets (and their statistics) don't reconstitute the full dataset
- **Applies when**: `task` -- the task asks to split rows into "missing" vs "non-missing" (or any complementary filter) on one column and compute per-group statistics or a test on another column.
- **Pattern**: The attempt uses an ad-hoc notion of "null" (e.g., only `NaN` while string sentinels like `""`, `"NA"`, `"None"`, `"-"`, `"null"` remain, or the reverse), or silently loses rows via `dropna()` on the whole frame, `read_csv` defaults/`na_values`, dtype coercion, deduplication, or a prior filter — so the two group sizes don't add up to the row count and the group means drift from the true values.
- **Detection procedure**:
  1. From the task, note that the two groups are complementary and must together cover every row of the source table.
  2. In the scripts, find where the data is loaded and where the mask is built; check for global `dropna()`/`fillna`, subsetting, type casts, or a mask that only tests one form of missingness.
  3. Confirm the script prints `len(group_a)`, `len(group_b)`, and `len(df)` and asserts they sum; also check the value column's own missing values are handled explicitly and identically for both groups.
  4. Sanity-check the reported statistics against these counts/ranges; if counts were never printed, the partition is unverified.
- **Discriminator**: A real violation is when rows can fall outside both groups (or are dropped before masking) or when only one missingness representation is recognized; it is fine if the script explicitly normalizes sentinels to NaN, documents any exclusion, and shows the two group sizes summing to the full row count.
- **Consequence**: Both group means (and the test statistic) are computed on the wrong subsets, so the reported numbers miss the expected values even though the p-value may still look decisively significant, and the grader marks the mean checks wrong.
226Substituting invented rules for the referenced specification, and emitting only a subset of required artifactstaskda-code
Applies when
task -- the task points to an external instruction/spec file (guidance, template, config) and/or an evaluation harness that consumes saved artifacts, and the script must implement exactly those steps and write exactly those outputs.
Pattern
The script never loads/quotes the referenced spec; it improvises preprocessing and filtering thresholds (e.g., self-chosen cutoffs, self-derived reference dates, self-invented category mapping and ordering) and saves only the one artifact named in the prompt text, silently skipping the other data/figure artifacts the pipeline expects. The answer then presents the improvised rules as if they were the given requirements.
Detection procedure
  1. From the task, list every referenced spec document and every expected output artifact (data dumps, serialized plot metadata, arrays, images) with their exact filenames/paths.
  2. In the scripts, check that each spec-driven step is traceable to text actually read from the spec (or explicitly reproduced in a comment/print) rather than assumed, and that every listed artifact is written with the required name, shape, and content ordering.
  3. Flag if any thresholds, category definitions, orderings, or figure/legend/color assignments are chosen by the agent without a cited source, or if any expected artifact is never created.
  4. Check whether the final answer states these choices as assumptions or presents them as confirmed requirements.
Discriminator
A real violation is inventing decision rules that materially change the reported counts/statistics, or omitting artifacts the grader consumes. A look-alike that is fine is a script that reads the spec (or reproduces its wording) and only fills in cosmetic details the spec leaves free, while still writing all required outputs.
Consequence
All artifact-level checks fail (missing files, and the produced figure/values disagree with the spec's filtering and category definitions), so the run scores zero even though the reported numbers look internally consistent.
id 5aa8bb536bc5 · mined from da-code dacode-plot-pie-005@s4
raw text (what the judge reads)
### Substituting invented rules for the referenced specification, and emitting only a subset of required artifacts
- **Applies when**: `task` -- the task points to an external instruction/spec file (guidance, template, config) and/or an evaluation harness that consumes saved artifacts, and the script must implement exactly those steps and write exactly those outputs.
- **Pattern**: The script never loads/quotes the referenced spec; it improvises preprocessing and filtering thresholds (e.g., self-chosen cutoffs, self-derived reference dates, self-invented category mapping and ordering) and saves only the one artifact named in the prompt text, silently skipping the other data/figure artifacts the pipeline expects. The answer then presents the improvised rules as if they were the given requirements.
- **Detection procedure**:
  1. From the task, list every referenced spec document and every expected output artifact (data dumps, serialized plot metadata, arrays, images) with their exact filenames/paths.
  2. In the scripts, check that each spec-driven step is traceable to text actually read from the spec (or explicitly reproduced in a comment/print) rather than assumed, and that every listed artifact is written with the required name, shape, and content ordering.
  3. Flag if any thresholds, category definitions, orderings, or figure/legend/color assignments are chosen by the agent without a cited source, or if any expected artifact is never created.
  4. Check whether the final answer states these choices as assumptions or presents them as confirmed requirements.
- **Discriminator**: A real violation is inventing decision rules that materially change the reported counts/statistics, or omitting artifacts the grader consumes. A look-alike that is fine is a script that reads the spec (or reproduces its wording) and only fills in cosmetic details the spec leaves free, while still writing all required outputs.
- **Consequence**: All artifact-level checks fail (missing files, and the produced figure/values disagree with the spec's filtering and category definitions), so the run scores zero even though the reported numbers look internally consistent.
227Model error not sanity-checked against a trivial baseline (target variance), hiding broken feature/target preprocessingtaskinfiagent-dabench
Applies when
task -- a supervised regression task asks for a specific error metric (e.g., MSE) after prescribed preprocessing such as mean-imputation of a few numeric columns.
Pattern
The agent loads raw columns that are not cleanly numeric (strings with currency symbols, thousands separators, unit suffixes, footnote markers, ranges, placeholder tokens like "n/a"/"-"), applies to_numeric(..., errors='coerce'), dropna, or an ad-hoc parse that silently mangles or discards many values, then fits the model and reports whatever number comes out — without ever comparing the resulting error to the variance/std of the target or to a mean-predictor baseline, and without checking row counts and value ranges after cleaning.
Detection procedure
  1. From the task, note the exact preprocessing mandated (which columns, impute-with-mean vs drop, split fraction) and the metric requested.
  2. In the scripts, locate where each modeling column is converted to numeric: check whether non-numeric artifacts are stripped before conversion, whether coercion-to-NaN plus row-dropping replaces the required imputation, and whether the row count/NaN count is printed before and after cleaning.
  3. Check whether the script computes any reference quantity — target variance, std, or the MSE of predicting the training mean — and compares it with the reported metric.
  4. Compare the reported metric to the plausible scale of the target: an MSE far above the target's variance (i.e., worse than predicting the mean) or a nonsensical error magnitude relative to the target's observed range signals broken parsing/imputation or mismatched X/y rows.
Discriminator
A genuine violation is when the reported error is at or above the naive mean-predictor error, or when rows were silently dropped/coerced instead of imputed as instructed, with no diagnostic printed. It is not a violation if the columns are already clean numerics, imputation follows the stated rule, counts/shapes are verified, and a weak-but-below-baseline error simply reflects genuinely low predictor signal.
Consequence
The reported metric is off by an order of magnitude from the reference value computed on correctly parsed, mean-imputed data, so the graded numeric answer fails exact-match.
id 7beee8d46a60 · mined from infiagent-dabench dabench-432@s4
raw text (what the judge reads)
### Model error not sanity-checked against a trivial baseline (target variance), hiding broken feature/target preprocessing
- **Applies when**: `task` -- a supervised regression task asks for a specific error metric (e.g., MSE) after prescribed preprocessing such as mean-imputation of a few numeric columns.
- **Pattern**: The agent loads raw columns that are not cleanly numeric (strings with currency symbols, thousands separators, unit suffixes, footnote markers, ranges, placeholder tokens like "n/a"/"-"), applies `to_numeric(..., errors='coerce')`, `dropna`, or an ad-hoc parse that silently mangles or discards many values, then fits the model and reports whatever number comes out — without ever comparing the resulting error to the variance/std of the target or to a mean-predictor baseline, and without checking row counts and value ranges after cleaning.
- **Detection procedure**:
  1. From the task, note the exact preprocessing mandated (which columns, impute-with-mean vs drop, split fraction) and the metric requested.
  2. In the scripts, locate where each modeling column is converted to numeric: check whether non-numeric artifacts are stripped before conversion, whether coercion-to-NaN plus row-dropping replaces the required imputation, and whether the row count/NaN count is printed before and after cleaning.
  3. Check whether the script computes any reference quantity — target variance, std, or the MSE of predicting the training mean — and compares it with the reported metric.
  4. Compare the reported metric to the plausible scale of the target: an MSE far above the target's variance (i.e., worse than predicting the mean) or a nonsensical error magnitude relative to the target's observed range signals broken parsing/imputation or mismatched X/y rows.
- **Discriminator**: A genuine violation is when the reported error is at or above the naive mean-predictor error, or when rows were silently dropped/coerced instead of imputed as instructed, with no diagnostic printed. It is *not* a violation if the columns are already clean numerics, imputation follows the stated rule, counts/shapes are verified, and a weak-but-below-baseline error simply reflects genuinely low predictor signal.
- **Consequence**: The reported metric is off by an order of magnitude from the reference value computed on correctly parsed, mean-imputed data, so the graded numeric answer fails exact-match.
228Unjustified re-ordering of rows before a sequence-dependent (lag/diff) computationtaskinfiagent-dabench
Applies when
task -- the task asks for a quantity defined relative to the "previous" row (differences, percent change, cumulative or rolling statistics) and the script re-sorts or re-indexes the data before computing it.
Pattern
The agent assumes the file's row order is "wrong" (e.g., reversed) and sorts by a parsed key, then computes the lag-based column on the new order, without verifying that the task intended a different order than the one supplied; because the operation is order-sensitive, the sign and magnitude of the resulting statistics flip relative to the as-given order.
Detection procedure
  1. Read the task statement and check whether it specifies any sorting/ordering; if it only says "previous day/row", the as-given file order is the default definition.
  2. Scan the script for any sort_values, sort_index, reset_index, reversal, or groupby-reordering applied before diff/shift/pct_change/rolling operations.
  3. Check whether the script provides evidence for the re-ordering (printed head/tail of the key showing it was inconsistent) and whether it computes the statistic under both orderings and reports the discrepancy.
  4. Compare the reported sign/magnitude against a quick sanity expectation (e.g., mean change should be near-symmetric in magnitude but opposite in sign between the two orderings); a sign difference means the answer depends entirely on the unverified assumption.
Discriminator
A real violation is silently changing row order (or leaving it unverified) for an order-sensitive computation when the task defined "previous" relative to the given data; it is fine if the task explicitly requires chronological/keyed ordering, or if the script demonstrates the file is already in that order so the sort is a no-op (verifiable by comparing pre- and post-sort indices).
Consequence
The lag direction is inverted, so the mean flips sign (and dispersion shifts slightly), and the graded numeric checks fail even though the formula and rounding were implemented correctly.
id 36259e795653 · mined from infiagent-dabench dabench-75@s4
raw text (what the judge reads)
### Unjustified re-ordering of rows before a sequence-dependent (lag/diff) computation
- **Applies when**: `task` -- the task asks for a quantity defined relative to the "previous" row (differences, percent change, cumulative or rolling statistics) and the script re-sorts or re-indexes the data before computing it.
- **Pattern**: The agent assumes the file's row order is "wrong" (e.g., reversed) and sorts by a parsed key, then computes the lag-based column on the new order, without verifying that the task intended a different order than the one supplied; because the operation is order-sensitive, the sign and magnitude of the resulting statistics flip relative to the as-given order.
- **Detection procedure**:
  1. Read the task statement and check whether it specifies any sorting/ordering; if it only says "previous day/row", the as-given file order is the default definition.
  2. Scan the script for any `sort_values`, `sort_index`, `reset_index`, reversal, or groupby-reordering applied before `diff`/`shift`/`pct_change`/rolling operations.
  3. Check whether the script provides evidence for the re-ordering (printed head/tail of the key showing it was inconsistent) *and* whether it computes the statistic under both orderings and reports the discrepancy.
  4. Compare the reported sign/magnitude against a quick sanity expectation (e.g., mean change should be near-symmetric in magnitude but opposite in sign between the two orderings); a sign difference means the answer depends entirely on the unverified assumption.
- **Discriminator**: A real violation is silently changing row order (or leaving it unverified) for an order-sensitive computation when the task defined "previous" relative to the given data; it is fine if the task explicitly requires chronological/keyed ordering, or if the script demonstrates the file is already in that order so the sort is a no-op (verifiable by comparing pre- and post-sort indices).
- **Consequence**: The lag direction is inverted, so the mean flips sign (and dispersion shifts slightly), and the graded numeric checks fail even though the formula and rounding were implemented correctly.
229No held-out validation before committing to a model/ensembletaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test set and the scripts fit one or more models and write predictions directly, with no scored comparison of alternatives.
Pattern
The agent trains several estimators on the full labeled data, blends them with hand-picked weights (or picks hyperparameters by intuition), never computes any validation error on a held-out split or via cross-validation, and reports the submission as done — so weak or mis-specified components (e.g., linear models on label-encoded high-cardinality categoricals, low-capacity boosters) silently drag the blend's accuracy below the grader's threshold.
Detection procedure
  1. Read the task to confirm predictions will be scored against hidden truth by an error metric (i.e., accuracy matters, not just file format).
  2. Scan the scripts for any train/validation split, cross-validation, or metric call (MAE/RMSE/R²) on labeled data; check whether model choice, hyperparameters, and any ensemble weights are justified by such a number.
  3. Check the answer/report: does it quote a validation score, or only describe architecture and weights?
  4. If no validation score exists anywhere, flag: the agent had no evidence its chosen blend beats a simple single-model baseline.
Discriminator
A real violation is zero out-of-sample estimate of error anywhere in the pipeline (weights/hyperparameters asserted, not measured). It is not a violation if the agent evaluated candidates on a holdout/CV, reported the score, and then optionally refit the winner on all data — even if the final model is simple.
Consequence
The grader compares predictions to true labels under an error/agreement threshold; an unvalidated, arbitrarily weighted blend that includes badly-fit components misses the threshold and the file is marked WRONG despite correct shape and column name.
id 1cf36f3916fb · mined from da-code dacode-ml-regression-014@s4
raw text (what the judge reads)
### No held-out validation before committing to a model/ensemble
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test set and the scripts fit one or more models and write predictions directly, with no scored comparison of alternatives.
- **Pattern**: The agent trains several estimators on the full labeled data, blends them with hand-picked weights (or picks hyperparameters by intuition), never computes any validation error on a held-out split or via cross-validation, and reports the submission as done — so weak or mis-specified components (e.g., linear models on label-encoded high-cardinality categoricals, low-capacity boosters) silently drag the blend's accuracy below the grader's threshold.
- **Detection procedure**:
  1. Read the task to confirm predictions will be scored against hidden truth by an error metric (i.e., accuracy matters, not just file format).
  2. Scan the scripts for any train/validation split, cross-validation, or metric call (MAE/RMSE/R²) on labeled data; check whether model choice, hyperparameters, and any ensemble weights are justified by such a number.
  3. Check the answer/report: does it quote a validation score, or only describe architecture and weights?
  4. If no validation score exists anywhere, flag: the agent had no evidence its chosen blend beats a simple single-model baseline.
- **Discriminator**: A real violation is zero out-of-sample estimate of error anywhere in the pipeline (weights/hyperparameters asserted, not measured). It is *not* a violation if the agent evaluated candidates on a holdout/CV, reported the score, and then optionally refit the winner on all data — even if the final model is simple.
- **Consequence**: The grader compares predictions to true labels under an error/agreement threshold; an unvalidated, arbitrarily weighted blend that includes badly-fit components misses the threshold and the file is marked WRONG despite correct shape and column name.
230Output row count / alignment with the evaluation input not verifiedtaskda-code
Applies when
task -- The task asks for a per-row prediction (or per-row derived value) file produced from a given input table, so the deliverable's length and row order must correspond one-to-one with that input.
Pattern
The attempt produces a prediction file whose number of rows does not equal the number of rows in the provided evaluation input (e.g., predictions generated from a truncated sample, a filtered/deduplicated subset, a head/nrows read, a partial chunk, or a leftover debug run), or whose rows are reordered/shuffled relative to the input. No script asserts len(predictions) == len(input_df) or preserves the original index order, and the reported answer is accepted without a shape check.
Detection procedure
  1. From the task/README, determine the expected number of prediction rows (row count of the evaluation input file) and the required column name/header.
  2. In the scripts, trace how the input is loaded and how predictions are written: look for nrows=, .head(), .sample(), dropna(), drop_duplicates(), filtering/merging, chunked loops, or sorting/shuffling between load and write; check whether the original row order/index is preserved.
  3. Count the data rows in the submitted file (excluding header) and compare to the expected count; also confirm the header text matches exactly.
  4. Confirm at least one explicit sanity assertion or printed shape check exists tying the output length to the input length; absence plus a mismatch is a violation.
Discriminator
A real violation is a genuine length/order mismatch with the evaluation input (or an unverifiable output produced by code that could drop/reorder rows). It is not a violation if row filtering is explicitly required by the task, or if the count matches and order is provably preserved (index-aligned write) even though intermediate filtering happened on a copy.
Consequence
The grader cannot align predictions to the ground-truth labels; the file is scored as wrong/missing (accuracy effectively 0) regardless of model quality.
id 8f461ebaacfd · mined from da-code dacode-ml-multi-008@s4
raw text (what the judge reads)
### Output row count / alignment with the evaluation input not verified
- **Applies when**: `task` -- The task asks for a per-row prediction (or per-row derived value) file produced from a given input table, so the deliverable's length and row order must correspond one-to-one with that input.
- **Pattern**: The attempt produces a prediction file whose number of rows does not equal the number of rows in the provided evaluation input (e.g., predictions generated from a truncated sample, a filtered/deduplicated subset, a head/`nrows` read, a partial chunk, or a leftover debug run), or whose rows are reordered/shuffled relative to the input. No script asserts `len(predictions) == len(input_df)` or preserves the original index order, and the reported answer is accepted without a shape check.
- **Detection procedure**:
  1. From the task/README, determine the expected number of prediction rows (row count of the evaluation input file) and the required column name/header.
  2. In the scripts, trace how the input is loaded and how predictions are written: look for `nrows=`, `.head()`, `.sample()`, `dropna()`, `drop_duplicates()`, filtering/merging, chunked loops, or sorting/shuffling between load and write; check whether the original row order/index is preserved.
  3. Count the data rows in the submitted file (excluding header) and compare to the expected count; also confirm the header text matches exactly.
  4. Confirm at least one explicit sanity assertion or printed shape check exists tying the output length to the input length; absence plus a mismatch is a violation.
- **Discriminator**: A real violation is a genuine length/order mismatch with the evaluation input (or an unverifiable output produced by code that could drop/reorder rows). It is *not* a violation if row filtering is explicitly required by the task, or if the count matches and order is provably preserved (index-aligned write) even though intermediate filtering happened on a copy.
- **Consequence**: The grader cannot align predictions to the ground-truth labels; the file is scored as wrong/missing (accuracy effectively 0) regardless of model quality.
231Missing or unverified required output artifacts specified by a config/spec filetaskda-code
Applies when
task -- the task points to an external specification (e.g., a YAML/JSON config with plotting or output guidelines) and/or expects deliverables saved to disk rather than only printed in the chat answer.
Pattern
The attempt reads (or paraphrases) the spec, hardcodes a few of its settings, writes only the single most obvious artifact (the image), and reports the numbers in prose — never enumerating the full set of files/keys the spec and task imply, and never re-loading the written files to confirm they exist with the right structure/values.
Detection procedure
  1. Read the task and open the referenced spec file; list every required deliverable (file names, extensions, save paths) and every constrained property (title text, labels, colors, ordering of categories, figure size, rounding, units, serialized data dumps).
  2. Read the scripts and build the set of files actually written and settings actually applied; diff against the list from step 1.
  3. Check whether the script contains a verification step that reloads each written artifact (e.g., re-open the image/array/JSON and assert shape, keys, category order, and value ranges).
  4. Check the final answer: does it claim completion based on prose description of the chart rather than on evidence that all named artifacts were saved and validated?
Discriminator
A real violation is when a spec- or task-implied deliverable (an auxiliary data dump, a serialized array, a metadata JSON, a specific save path/filename) is absent from the scripts, or a constrained property is assumed rather than read from the spec. It is not a violation if the agent produced every file the spec names and merely used extra defaults for properties the spec leaves free, or if an unrequested extra file is also written.
Consequence
The grader's per-file checks fail as WRONG/MISSING for every deliverable that was never created (and possibly for the one that was, if its category order/labels/colors deviate from the spec), yielding 0 passed checks even when the headline computed value is right.
id c7fdf7e2122d · mined from da-code dacode-plot-pie-008@s4
raw text (what the judge reads)
### Missing or unverified required output artifacts specified by a config/spec file
- **Applies when**: `task` -- the task points to an external specification (e.g., a YAML/JSON config with plotting or output guidelines) and/or expects deliverables saved to disk rather than only printed in the chat answer.
- **Pattern**: The attempt reads (or paraphrases) the spec, hardcodes a few of its settings, writes only the single most obvious artifact (the image), and reports the numbers in prose — never enumerating the full set of files/keys the spec and task imply, and never re-loading the written files to confirm they exist with the right structure/values.
- **Detection procedure**:
  1. Read the task and open the referenced spec file; list every required deliverable (file names, extensions, save paths) and every constrained property (title text, labels, colors, ordering of categories, figure size, rounding, units, serialized data dumps).
  2. Read the scripts and build the set of files actually written and settings actually applied; diff against the list from step 1.
  3. Check whether the script contains a verification step that reloads each written artifact (e.g., re-open the image/array/JSON and assert shape, keys, category order, and value ranges).
  4. Check the final answer: does it claim completion based on prose description of the chart rather than on evidence that all named artifacts were saved and validated?
- **Discriminator**: A real violation is when a spec- or task-implied deliverable (an auxiliary data dump, a serialized array, a metadata JSON, a specific save path/filename) is absent from the scripts, or a constrained property is assumed rather than read from the spec. It is *not* a violation if the agent produced every file the spec names and merely used extra defaults for properties the spec leaves free, or if an unrequested extra file is also written.
- **Consequence**: The grader's per-file checks fail as WRONG/MISSING for every deliverable that was never created (and possibly for the one that was, if its category order/labels/colors deviate from the spec), yielding 0 passed checks even when the headline computed value is right.
232Unvalidated missing-value handling that silently changes the row set before a fixed-seed splittaskinfiagent-dabench
Applies when
task -- the task prescribes an exact, reproducible split (fixed seed/proportions) and metric on data whose feature columns contain missing values, and the script preprocesses before splitting.
Pattern
The attempt drops rows (or drops/reshapes columns) containing nulls, or imputes with an ad-hoc statistic, without checking or reporting how many rows survive; the resulting split — and therefore the metric — differs from the canonical one even though the seed matches. A related variant is failing to state how continuous model outputs are converted into class labels, so the accuracy definition itself is unverified.
Detection procedure
  1. From the task, list the mandated preprocessing/split/metric spec and note which steps are not specified (missing-value policy, thresholding rule) — these are the discretionary points that must be sanity-checked.
  2. In the script, find every operation that can change row count or label mapping (dropna, filtered joins, fillna, astype, threshold/rounding of predictions) and check whether the script prints shapes before/after and the resulting train/test sizes.
  3. Compare the reported number of rows entering the split against the raw row count; if they differ, the fixed seed no longer reproduces the intended split, and the attempt must justify the choice or test the alternative (impute vs. drop).
  4. Check the answer is derived from the test-subset predictions after an explicit label conversion, and that its value is reported in the requested rounding/format.
Discriminator
A real violation is an undocumented row-count or label-mapping change with no shape/count sanity check and no comparison of the plausible alternative; it is fine if the script logs the pre/post row counts, keeps the full row set (imputation) so the seeded split matches the raw data, and explicitly defines the prediction-to-label rule.
Consequence
The metric is computed on a different test subset (or with an unstated decision rule) than the reference, producing an accuracy off by a few hundredths — close enough to look plausible but graded wrong on exact-value comparison.
id bbcbb5d1272d · mined from infiagent-dabench dabench-7@s4
raw text (what the judge reads)
### Unvalidated missing-value handling that silently changes the row set before a fixed-seed split
- **Applies when**: `task` -- the task prescribes an exact, reproducible split (fixed seed/proportions) and metric on data whose feature columns contain missing values, and the script preprocesses before splitting.
- **Pattern**: The attempt drops rows (or drops/reshapes columns) containing nulls, or imputes with an ad-hoc statistic, without checking or reporting how many rows survive; the resulting split — and therefore the metric — differs from the canonical one even though the seed matches. A related variant is failing to state how continuous model outputs are converted into class labels, so the accuracy definition itself is unverified.
- **Detection procedure**:
  1. From the task, list the mandated preprocessing/split/metric spec and note which steps are *not* specified (missing-value policy, thresholding rule) — these are the discretionary points that must be sanity-checked.
  2. In the script, find every operation that can change row count or label mapping (`dropna`, filtered joins, `fillna`, `astype`, threshold/rounding of predictions) and check whether the script prints shapes before/after and the resulting train/test sizes.
  3. Compare the reported number of rows entering the split against the raw row count; if they differ, the fixed seed no longer reproduces the intended split, and the attempt must justify the choice or test the alternative (impute vs. drop).
  4. Check the answer is derived from the test-subset predictions after an explicit label conversion, and that its value is reported in the requested rounding/format.
- **Discriminator**: A real violation is an undocumented row-count or label-mapping change with no shape/count sanity check and no comparison of the plausible alternative; it is fine if the script logs the pre/post row counts, keeps the full row set (imputation) so the seeded split matches the raw data, and explicitly defines the prediction-to-label rule.
- **Consequence**: The metric is computed on a different test subset (or with an unstated decision rule) than the reference, producing an accuracy off by a few hundredths — close enough to look plausible but graded wrong on exact-value comparison.
233Ignoring the provided output template's schema and semanticstaskda-code
Applies when
task -- the task supplies a pre-existing answer/submission file whose existing columns, row labels, or category set define exactly what must be filled in.
Pattern
The agent invents its own schema (its own category names, its own column meanings such as row indices instead of counts, extra/missing category rows) instead of reading the template file first and populating only its defined fields, then reports narrative numbers as if the deliverable were satisfied.
Detection procedure
  1. From the task, note that a specific output file/format is mandated and must be filled in "strictly".
  2. In the scripts, check whether the template file is actually read/inspected (headers, existing rows, allowed category labels) before writing, and whether the write preserves those headers and row keys.
  3. Compare the answer's category list and per-row values against the template's expected fields — do category names come from the template or from the agent's own derivation? Does each value type match the column meaning (e.g., a count vs. an identifier)?
  4. Flag if any template row/column is absent, renamed, reordered, or filled with a different quantity than its header implies (including catch-all buckets like "Unknown" that the template does not contain).
Discriminator
Fine if the agent inspected the template and its output matches header names, row keys, ordering and value semantics exactly (extra explanatory prose outside the file is harmless); a violation is when the file's schema/labels/value meaning is agent-invented or deviates from the template, even if the underlying computation is defensible.
Consequence
The result file fails exact-match checks on keys/values, so the graded check scores 0 regardless of how reasonable the reported narrative numbers look.
id 4951cfe1be16 · mined from da-code dacode-dm-csv-001@s4
raw text (what the judge reads)
### Ignoring the provided output template's schema and semantics
- **Applies when**: `task` -- the task supplies a pre-existing answer/submission file whose existing columns, row labels, or category set define exactly what must be filled in.
- **Pattern**: The agent invents its own schema (its own category names, its own column meanings such as row indices instead of counts, extra/missing category rows) instead of reading the template file first and populating only its defined fields, then reports narrative numbers as if the deliverable were satisfied.
- **Detection procedure**:
  1. From the task, note that a specific output file/format is mandated and must be filled in "strictly".
  2. In the scripts, check whether the template file is actually read/inspected (headers, existing rows, allowed category labels) before writing, and whether the write preserves those headers and row keys.
  3. Compare the answer's category list and per-row values against the template's expected fields — do category names come from the template or from the agent's own derivation? Does each value type match the column meaning (e.g., a count vs. an identifier)?
  4. Flag if any template row/column is absent, renamed, reordered, or filled with a different quantity than its header implies (including catch-all buckets like "Unknown" that the template does not contain).
- **Discriminator**: Fine if the agent inspected the template and its output matches header names, row keys, ordering and value semantics exactly (extra explanatory prose outside the file is harmless); a violation is when the file's schema/labels/value meaning is agent-invented or deviates from the template, even if the underlying computation is defensible.
- **Consequence**: The result file fails exact-match checks on keys/values, so the graded check scores 0 regardless of how reasonable the reported narrative numbers look.
234Grouping-level (unit-of-analysis) error before computing a statistictaskinfiagent-dabench
Applies when
task -- the task asks for a statistic (correlation, mean, split by median, etc.) about entities while the raw data stores multiple time-stamped or repeated rows per entity.
Pattern
The attempt computes the statistic directly on raw rows, or aggregates entities with an inconsistent recipe (e.g. duration derived from row counts instead of first/last timestamp difference, "maximum" taken after rather than before deduplication, the median split computed on row-level values instead of one value per entity), so the resulting r/p reflects the wrong sample size and a slightly different quantity than requested.
Detection procedure
  1. From the task wording, identify the entity key (the thing one data point should represent) and the per-entity quantities required (extremum, span/duration, magnitude used for the split).
  2. In the scripts, check that there is an explicit groupby(entity_key).agg(...) producing exactly one row per entity, and that each requested quantity is derived from that aggregation with the natural definition (max for extremum, last−first timestamp for span, per-entity total/scalar for the split variable).
  3. Verify the split threshold (e.g. median) is computed on the aggregated per-entity table, and print the number of entities in each subgroup; confirm it is plausible and that the two subgroups partition the entities.
  4. Compare the reported r/p against the aggregated-entity computation; if the script never collapsed to entity level, or duration/extremum was defined by a proxy, flag it.
Discriminator
A genuine violation is when row-level or proxy-aggregated data feeds the statistic (entity counts differ from the number of unique keys, or duration comes from record counts/units other than those implied). It is fine if the data is already one row per entity, or if the aggregation is present and only the tie-breaking of the median split differs in a documented, defensible way.
Consequence
The reported coefficient (and sometimes the relationship classification) drifts from the ground-truth value by more than the rounding tolerance, so numeric checks fail even when the qualitative conclusion happens to match.
id dc06ee4f6bfc · mined from infiagent-dabench dabench-431@s4
raw text (what the judge reads)
### Grouping-level (unit-of-analysis) error before computing a statistic
- **Applies when**: `task` -- the task asks for a statistic (correlation, mean, split by median, etc.) about *entities* while the raw data stores multiple time-stamped or repeated rows per entity.
- **Pattern**: The attempt computes the statistic directly on raw rows, or aggregates entities with an inconsistent recipe (e.g. duration derived from row counts instead of first/last timestamp difference, "maximum" taken after rather than before deduplication, the median split computed on row-level values instead of one value per entity), so the resulting r/p reflects the wrong sample size and a slightly different quantity than requested.
- **Detection procedure**:
  1. From the task wording, identify the entity key (the thing one data point should represent) and the per-entity quantities required (extremum, span/duration, magnitude used for the split).
  2. In the scripts, check that there is an explicit `groupby(entity_key).agg(...)` producing exactly one row per entity, and that each requested quantity is derived from that aggregation with the natural definition (max for extremum, last−first timestamp for span, per-entity total/scalar for the split variable).
  3. Verify the split threshold (e.g. median) is computed on the aggregated per-entity table, and print the number of entities in each subgroup; confirm it is plausible and that the two subgroups partition the entities.
  4. Compare the reported r/p against the aggregated-entity computation; if the script never collapsed to entity level, or duration/extremum was defined by a proxy, flag it.
- **Discriminator**: A genuine violation is when row-level or proxy-aggregated data feeds the statistic (entity counts differ from the number of unique keys, or duration comes from record counts/units other than those implied). It is fine if the data is already one row per entity, or if the aggregation is present and only the tie-breaking of the median split differs in a documented, defensible way.
- **Consequence**: The reported coefficient (and sometimes the relationship classification) drifts from the ground-truth value by more than the rounding tolerance, so numeric checks fail even when the qualitative conclusion happens to match.
235Unverified predictions with no held-out validation or reproducible scripttaskda-code
Applies when
task -- the deliverable is a prediction file produced by a model trained on a labeled training split and applied to an unlabeled test split.
Pattern
The attempt fits a single model with default/ad-hoc hyperparameters, writes the output file, and reports only self-descriptive metadata (row counts, class balance, feature importances) without ever estimating out-of-sample performance on a held-out or cross-validated subset, without comparing against a baseline, and without preserving the script so the pipeline can be inspected — so systematic errors (mis-set decision threshold, class-weight over-prediction of the minority class, dropped/mis-encoded features, misaligned row order or ID mapping) go undetected.
Detection procedure
  1. Read the task to identify the required output file, column name, expected number of rows, and the label semantics.
  2. Inspect the scripts (or note their absence) for a validation step: a train/validation split or cross-validation with the target metric printed, plus a trivial baseline (e.g., majority class) for comparison.
  3. Check whether the reported answer cites any out-of-sample score and whether predicted class proportions are compared to the training label distribution and to the ID ordering of the test file.
  4. Flag if the only evidence of correctness is descriptive counts of the submission itself, or if no script exists to verify preprocessing consistency between train and test.
Discriminator
A real violation has zero out-of-sample evidence (no CV/holdout score, no baseline, no distribution sanity check) and/or no inspectable pipeline; a look-alike that is fine reports a validated metric and shows train/test preprocessing were applied consistently, even if the final model is simple.
Consequence
The submitted file may have the right shape and column name yet poor accuracy or a skewed positive rate, so the grader's accuracy/threshold check on the prediction column fails while the agent claims success.
id f1c5ce03bef3 · mined from da-code dacode-ml-binary-016@s4
raw text (what the judge reads)
### Unverified predictions with no held-out validation or reproducible script
- **Applies when**: `task` -- the deliverable is a prediction file produced by a model trained on a labeled training split and applied to an unlabeled test split.
- **Pattern**: The attempt fits a single model with default/ad-hoc hyperparameters, writes the output file, and reports only self-descriptive metadata (row counts, class balance, feature importances) without ever estimating out-of-sample performance on a held-out or cross-validated subset, without comparing against a baseline, and without preserving the script so the pipeline can be inspected — so systematic errors (mis-set decision threshold, class-weight over-prediction of the minority class, dropped/mis-encoded features, misaligned row order or ID mapping) go undetected.
- **Detection procedure**:
  1. Read the task to identify the required output file, column name, expected number of rows, and the label semantics.
  2. Inspect the scripts (or note their absence) for a validation step: a train/validation split or cross-validation with the target metric printed, plus a trivial baseline (e.g., majority class) for comparison.
  3. Check whether the reported answer cites any out-of-sample score and whether predicted class proportions are compared to the training label distribution and to the ID ordering of the test file.
  4. Flag if the only evidence of correctness is descriptive counts of the submission itself, or if no script exists to verify preprocessing consistency between train and test.
- **Discriminator**: A real violation has zero out-of-sample evidence (no CV/holdout score, no baseline, no distribution sanity check) and/or no inspectable pipeline; a look-alike that is fine reports a validated metric and shows train/test preprocessing were applied consistently, even if the final model is simple.
- **Consequence**: The submitted file may have the right shape and column name yet poor accuracy or a skewed positive rate, so the grader's accuracy/threshold check on the prediction column fails while the agent claims success.
236Deliverable file never written (or not matching the provided template schema)taskda-code
Applies when
task -- the task requires saving results to a specific output filename in the format of a provided sample/template file, and grading is done on that file rather than on chat text.
Pattern
The agent computes a number and reports it only in its narrative answer (or writes it with ad-hoc column names/row layout, wrong path, or from an interactive session with no saved script), so no artifact at the required path reproduces the template's exact header, column order, and row count.
Detection procedure
  1. Read the task and note the exact required output filename, and locate the sample/template file it must mirror; record its header names, column order, and expected number of rows.
  2. Inspect the scripts for an explicit write call to that exact filename (relative to the expected working directory), and check whether the DataFrame constructed there uses the template's column names/order and index=False semantics.
  3. Compare the agent's reported answer to what the written file would contain; if there are no saved scripts at all, treat the artifact as unverifiable and flag it.
  4. Also confirm any stated reproducibility constraint (e.g., fixed random seed for resampling-based statistics) is set in the script that produces the file, so the saved value is regenerable.
Discriminator
A real violation is the absence of a persisted file at the required name, or a file whose schema/values differ from the template (extra index column, renamed/reordered fields, missing rows, wrong directory). A look-alike that is fine: the script writes the correct file and the agent additionally echoes the same content in its message, with identical values and headers.
Consequence
The grader looks for the named result file, finds it missing or schema-mismatched, and scores 0 even if the underlying statistic was computed correctly.
id 14eb7d640560 · mined from da-code dacode-data-sa-039@s4
raw text (what the judge reads)
### Deliverable file never written (or not matching the provided template schema)
- **Applies when**: `task` -- the task requires saving results to a specific output filename in the format of a provided sample/template file, and grading is done on that file rather than on chat text.
- **Pattern**: The agent computes a number and reports it only in its narrative answer (or writes it with ad-hoc column names/row layout, wrong path, or from an interactive session with no saved script), so no artifact at the required path reproduces the template's exact header, column order, and row count.
- **Detection procedure**:
  1. Read the task and note the exact required output filename, and locate the sample/template file it must mirror; record its header names, column order, and expected number of rows.
  2. Inspect the scripts for an explicit write call to that exact filename (relative to the expected working directory), and check whether the DataFrame constructed there uses the template's column names/order and index=False semantics.
  3. Compare the agent's reported answer to what the written file would contain; if there are no saved scripts at all, treat the artifact as unverifiable and flag it.
  4. Also confirm any stated reproducibility constraint (e.g., fixed random seed for resampling-based statistics) is set in the script that produces the file, so the saved value is regenerable.
- **Discriminator**: A real violation is the absence of a persisted file at the required name, or a file whose schema/values differ from the template (extra index column, renamed/reordered fields, missing rows, wrong directory). A look-alike that is fine: the script writes the correct file and the agent additionally echoes the same content in its message, with identical values and headers.
- **Consequence**: The grader looks for the named result file, finds it missing or schema-mismatched, and scores 0 even if the underlying statistic was computed correctly.
237Discarding a high-signal raw column (e.g., a timestamp/text field) instead of engineering features from it, then accepting a weak validation scoretaskda-code
Applies when
task -- the training table contains non-numeric or structured columns (dates/times, IDs, categoricals) alongside numeric ones, and the script builds its feature matrix by simply excluding anything that isn't directly usable by the estimator.
Pattern
The script drops such columns with a blanket filter ([c for c in cols if c not in ['date', target]]), never derives features from them (hour/day-of-week/month, lags, encodings), evaluates with a random shuffled split even though the data is sequential, and then reports a mediocre validation score (e.g., R² well below what the column would provide) as success without comparing against a naive baseline or attempting improvement.
Detection procedure
  1. From the task and data exploration output, list all columns present and note which ones are dropped before fitting; check whether any dropped column plausibly carries strong signal (time index, group/ID, category).
  2. In the modeling script, verify whether any transformation of those columns is created; if not, confirm the model sees only the residual numeric block.
  3. Check the validation protocol: is the split random while the data has an inherent order/time dependence, and is the reported score compared to a baseline (mean predictor, simple time-based model) or to a plausible target quality level?
  4. Read the answer: does it claim success purely on the basis of an unbenchmarked score, and does it also contain internal inconsistencies (naming one model while fitting another)?
Discriminator
A real violation is dropping a column that carries genuine predictive structure with no attempt to encode it and no baseline comparison. It is fine to drop a column that is provably uninformative (constant, pure noise, unique identifier with no ordering meaning) or absent from the test file — provided the script or answer states that check explicitly.
Consequence
Predictions are far less accurate than achievable (compressed range, near-mean outputs), so the submitted file fails the grader's accuracy/similarity threshold even though the file name, column name, and row count look correct.
id a97b532eaaa3 · mined from da-code dacode-ml-regression-015@s4
raw text (what the judge reads)
### Discarding a high-signal raw column (e.g., a timestamp/text field) instead of engineering features from it, then accepting a weak validation score
- **Applies when**: `task` -- the training table contains non-numeric or structured columns (dates/times, IDs, categoricals) alongside numeric ones, and the script builds its feature matrix by simply excluding anything that isn't directly usable by the estimator.
- **Pattern**: The script drops such columns with a blanket filter (`[c for c in cols if c not in ['date', target]]`), never derives features from them (hour/day-of-week/month, lags, encodings), evaluates with a random shuffled split even though the data is sequential, and then reports a mediocre validation score (e.g., R² well below what the column would provide) as success without comparing against a naive baseline or attempting improvement.
- **Detection procedure**:
  1. From the task and data exploration output, list all columns present and note which ones are dropped before fitting; check whether any dropped column plausibly carries strong signal (time index, group/ID, category).
  2. In the modeling script, verify whether any transformation of those columns is created; if not, confirm the model sees only the residual numeric block.
  3. Check the validation protocol: is the split random while the data has an inherent order/time dependence, and is the reported score compared to a baseline (mean predictor, simple time-based model) or to a plausible target quality level?
  4. Read the answer: does it claim success purely on the basis of an unbenchmarked score, and does it also contain internal inconsistencies (naming one model while fitting another)?
- **Discriminator**: A real violation is dropping a column that carries genuine predictive structure with no attempt to encode it and no baseline comparison. It is fine to drop a column that is provably uninformative (constant, pure noise, unique identifier with no ordering meaning) or absent from the test file — provided the script or answer states that check explicitly.
- **Consequence**: Predictions are far less accurate than achievable (compressed range, near-mean outputs), so the submitted file fails the grader's accuracy/similarity threshold even though the file name, column name, and row count look correct.
238Answer written to an ad-hoc file instead of the required deliverabletaskda-code
Applies when
task -- the task (or its harness) specifies a result artifact with a given name/path/format, and the scripts end by printing or saving the computed answer.
Pattern
The script computes a plausible answer but dumps it to a self-chosen filename/location (e.g. a scratch .txt in the working dir) or only prints it to stdout, never creating the exact expected result file with the exact requested JSON keys and value types.
Detection procedure
  1. From the task statement/harness, note the required output artifact: exact filename, directory, format, key names, and whether values are scalars or lists.
  2. Grep the scripts for every write operation (open(...,'w'), to_csv, json.dump) and record the paths and the object being serialized.
  3. Compare: does at least one write target the required path with the required schema? Check key spelling/case and that list-valued fields are lists, not bare strings.
  4. Confirm the submitted answer text is backed by that file, not merely echoed from console output.
Discriminator
A real violation is when no write matches the required artifact (wrong name, wrong directory, wrong container type, or console-only). It is not a violation if the script writes the correct file and additionally logs or saves extra debug copies elsewhere.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 even if the computed values would have been right.
id 157463a0880b · mined from da-code dacode-di-text-001@s4
raw text (what the judge reads)
### Answer written to an ad-hoc file instead of the required deliverable
- **Applies when**: `task` -- the task (or its harness) specifies a result artifact with a given name/path/format, and the scripts end by printing or saving the computed answer.
- **Pattern**: The script computes a plausible answer but dumps it to a self-chosen filename/location (e.g. a scratch `.txt` in the working dir) or only prints it to stdout, never creating the exact expected result file with the exact requested JSON keys and value types.
- **Detection procedure**:
  1. From the task statement/harness, note the required output artifact: exact filename, directory, format, key names, and whether values are scalars or lists.
  2. Grep the scripts for every write operation (`open(...,'w')`, `to_csv`, `json.dump`) and record the paths and the object being serialized.
  3. Compare: does at least one write target the required path with the required schema? Check key spelling/case and that list-valued fields are lists, not bare strings.
  4. Confirm the submitted answer text is backed by that file, not merely echoed from console output.
- **Discriminator**: A real violation is when no write matches the required artifact (wrong name, wrong directory, wrong container type, or console-only). It is *not* a violation if the script writes the correct file and additionally logs or saves extra debug copies elsewhere.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 even if the computed values would have been right.
239Deliverable file asserted rather than verified (summary in place of the requested artifact)taskda-code
Applies when
task -- The task requires writing results to a specific output file (e.g., a per-entity table of derived scores/segments/labels) and the agent's answer is a prose summary claiming the file was produced.
Pattern
The attempt reports aggregate statistics and methodology text but never shows the script that writes the file, never re-reads the saved file, and never confirms its path, row count, key column, and the specific requested columns (scores, segment, level) — so a missing, misplaced, truncated, or wrongly-shaped artifact goes unnoticed.
Detection procedure
  1. From the task statement, list the exact deliverable: filename, expected granularity (one row per entity), and every field explicitly requested (each component metric, the composite score, the segment/label column).
  2. In the scripts, locate the write call: confirm it targets the exact required filename in the expected working directory, and that the DataFrame written contains the identifier plus all requested fields (not just a subset or an intermediate frame).
  3. Check for a post-write verification step (re-read the file; print shape, columns, head, and null counts) and compare those numbers to the counts quoted in the answer.
  4. In the answer, check that the reported entity count and segment counts are reconciled against the file's actual rows, and that the answer isn't purely narrative with no evidence of the file's contents.
Discriminator
A real violation is when no script/verification evidence links the claimed numbers to an on-disk file with the required name and columns; it is fine if the agent shows the write plus a read-back with matching shape/columns even if the summary text is brief — and cosmetic differences (column ordering, extra helper columns) are not violations as long as all requested fields and one row per entity are present.
Consequence
The grader marks the expected output file WRONG/MISSING (0 checks passed) regardless of how reasonable the described methodology sounds, because the artifact it inspects is absent or lacks the required per-entity segment/level columns.
id fd7050b1d9df · mined from da-code dacode-dm-csv-052@s4
raw text (what the judge reads)
### Deliverable file asserted rather than verified (summary in place of the requested artifact)
- **Applies when**: `task` -- The task requires writing results to a specific output file (e.g., a per-entity table of derived scores/segments/labels) and the agent's answer is a prose summary claiming the file was produced.
- **Pattern**: The attempt reports aggregate statistics and methodology text but never shows the script that writes the file, never re-reads the saved file, and never confirms its path, row count, key column, and the specific requested columns (scores, segment, level) — so a missing, misplaced, truncated, or wrongly-shaped artifact goes unnoticed.
- **Detection procedure**:
  1. From the task statement, list the exact deliverable: filename, expected granularity (one row per entity), and every field explicitly requested (each component metric, the composite score, the segment/label column).
  2. In the scripts, locate the write call: confirm it targets the exact required filename in the expected working directory, and that the DataFrame written contains the identifier plus all requested fields (not just a subset or an intermediate frame).
  3. Check for a post-write verification step (re-read the file; print shape, columns, head, and null counts) and compare those numbers to the counts quoted in the answer.
  4. In the answer, check that the reported entity count and segment counts are reconciled against the file's actual rows, and that the answer isn't purely narrative with no evidence of the file's contents.
- **Discriminator**: A real violation is when no script/verification evidence links the claimed numbers to an on-disk file with the required name and columns; it is fine if the agent shows the write plus a read-back with matching shape/columns even if the summary text is brief — and cosmetic differences (column ordering, extra helper columns) are not violations as long as all requested fields and one row per entity are present.
- **Consequence**: The grader marks the expected output file WRONG/MISSING (0 checks passed) regardless of how reasonable the described methodology sounds, because the artifact it inspects is absent or lacks the required per-entity segment/level columns.
240Hardcoded/invented input data instead of loading the provided filestaskda-code
Applies when
task -- the task references supplied data files (and a sample output file defining the format), but the script embeds numeric arrays or literals in code instead of reading those files.
Pattern
The agent types in values "from memory" (or a truncated subset) and computes statistics on them, never opening the actual dataset or the format-template file, so both the numbers and the output schema are unverified against ground truth.
Detection procedure
  1. In the task/README, list every input artifact that should be read (data files, sample/template output file).
  2. Scan the script for read_csv/load/file paths: check whether each listed artifact is actually opened, or whether the data appears as inline literals/dicts.
  3. Compare the script's implied row counts and value ranges to whatever the task or data documentation states; inline data that is short, suspiciously round, or of unstated provenance is a red flag.
  4. Check that the output columns/order/header come from reading the sample file rather than being guessed in code.
Discriminator
A real violation is when the analysis' primary inputs or the output schema are invented rather than loaded; it is fine to hardcode genuine constants (bootstrap size, seed, confidence level) or to inline a tiny lookup table that the task itself specifies verbatim.
Consequence
Statistics are computed on the wrong sample (and possibly the wrong columns/format), so the written result file mismatches expected values and the file check fails even though the code runs cleanly.
id 42cdaa265372 · mined from da-code dacode-data-sa-029@s4
raw text (what the judge reads)
### Hardcoded/invented input data instead of loading the provided files
- **Applies when**: `task` -- the task references supplied data files (and a sample output file defining the format), but the script embeds numeric arrays or literals in code instead of reading those files.
- **Pattern**: The agent types in values "from memory" (or a truncated subset) and computes statistics on them, never opening the actual dataset or the format-template file, so both the numbers and the output schema are unverified against ground truth.
- **Detection procedure**:
  1. In the task/README, list every input artifact that should be read (data files, sample/template output file).
  2. Scan the script for `read_csv`/`load`/file paths: check whether each listed artifact is actually opened, or whether the data appears as inline literals/dicts.
  3. Compare the script's implied row counts and value ranges to whatever the task or data documentation states; inline data that is short, suspiciously round, or of unstated provenance is a red flag.
  4. Check that the output columns/order/header come from reading the sample file rather than being guessed in code.
- **Discriminator**: A real violation is when the analysis' *primary* inputs or the output schema are invented rather than loaded; it is fine to hardcode genuine constants (bootstrap size, seed, confidence level) or to inline a tiny lookup table that the task itself specifies verbatim.
- **Consequence**: Statistics are computed on the wrong sample (and possibly the wrong columns/format), so the written result file mismatches expected values and the file check fails even though the code runs cleanly.
241Numeric coercion / conversion accepted without validating the resulting values or the statistic's plausibilitytaskda-code
Applies when
task -- the task requires casting columns to numeric (or otherwise parsing raw fields) before computing a summary statistic such as a correlation, mean, or model metric.
Pattern
The attempt applies a permissive conversion (e.g. to_numeric(..., errors='coerce'), astype, regex extraction) and immediately computes and reports the statistic, without checking how many values survived, whether the parsed values fall in the expected range/scale, or whether the resulting statistic is plausible given the domain; silently mis-parsed or misaligned values (mixed rating scales, text-embedded numbers, header/index shifts) produce numbers that are technically valid but wrong. It also never re-reads the written output file to confirm the values and layout match the required sample format.
Detection procedure
  1. From the task, note the expected value domain of each parsed column (e.g. bounded rating scales) and the expected range/sign of the requested statistic.
  2. In the scripts, check whether any post-conversion validation exists: value counts / min-max / unique values per column, count of coerced NaNs vs. originally missing, row count after dropping, and a read-back of the saved file compared to the sample's shape, row/column names and ordering.
  3. Inspect the reported statistic against domain expectation — for related sub-ratings and an aggregate score, near-zero or negative correlations are a red flag, as are counts of dropped rows that are large and unexplained.
  4. If no such checks were run and no diagnostic printout of parsed values exists, flag the attempt regardless of how clean the code looks.
Discriminator
A real violation is an attempt whose only evidence is the final numbers, with no per-column parsed-value inspection and an outcome that contradicts reasonable expectations. A look-alike that is fine either prints/validates ranges, counts and output-file shape, or explicitly justifies a surprising result (e.g. shows the underlying scatter/distribution or documents why the columns are genuinely uncorrelated).
Consequence
The saved result file contains a full, well-formed matrix whose entries are all wrong (and possibly with mismatched labels/ordering), so an exact-value comparison against the expected file fails on every check.
id 43bcd8f1454f · mined from da-code dacode-data-sa-026@s4
raw text (what the judge reads)
### Numeric coercion / conversion accepted without validating the resulting values or the statistic's plausibility
- **Applies when**: `task` -- the task requires casting columns to numeric (or otherwise parsing raw fields) before computing a summary statistic such as a correlation, mean, or model metric.
- **Pattern**: The attempt applies a permissive conversion (e.g. `to_numeric(..., errors='coerce')`, `astype`, regex extraction) and immediately computes and reports the statistic, without checking how many values survived, whether the parsed values fall in the expected range/scale, or whether the resulting statistic is plausible given the domain; silently mis-parsed or misaligned values (mixed rating scales, text-embedded numbers, header/index shifts) produce numbers that are technically valid but wrong. It also never re-reads the written output file to confirm the values and layout match the required sample format.
- **Detection procedure**:
  1. From the task, note the expected value domain of each parsed column (e.g. bounded rating scales) and the expected range/sign of the requested statistic.
  2. In the scripts, check whether any post-conversion validation exists: value counts / min-max / unique values per column, count of coerced NaNs vs. originally missing, row count after dropping, and a read-back of the saved file compared to the sample's shape, row/column names and ordering.
  3. Inspect the reported statistic against domain expectation — for related sub-ratings and an aggregate score, near-zero or negative correlations are a red flag, as are counts of dropped rows that are large and unexplained.
  4. If no such checks were run and no diagnostic printout of parsed values exists, flag the attempt regardless of how clean the code looks.
- **Discriminator**: A real violation is an attempt whose only evidence is the final numbers, with no per-column parsed-value inspection and an outcome that contradicts reasonable expectations. A look-alike that is fine either prints/validates ranges, counts and output-file shape, or explicitly justifies a surprising result (e.g. shows the underlying scatter/distribution or documents why the columns are genuinely uncorrelated).
- **Consequence**: The saved result file contains a full, well-formed matrix whose entries are all wrong (and possibly with mismatched labels/ordering), so an exact-value comparison against the expected file fails on every check.
242Output row count silently smaller than the input dataset (partial/truncated results)taskda-code
Applies when
task -- The task asks for a per-record output artifact (e.g., a label/prediction/score for every observation) written to a file, and the script loads a large source table and writes a derived table.
Pattern
The attempt processes or emits only a subset of the records (a sampled/truncated/head-limited slice, an early-stopped loop, or a partially written file) and reports summary counts far below the documented size of the source data, without ever comparing the output row count to the input row count.
Detection procedure
  1. From the task/README, note the expected scale of the data (stated number of observations, or the shape printed when the raw file is loaded).
  2. In the scripts, trace the row count from load → any filtering/sampling/slicing → the frame that is written out; check whether any step reduces rows and whether an explicit len(output) == len(input) assertion exists.
  3. In the reported answer, read the stated number of rows/samples and the per-group counts; sum them and compare against the expected scale from step 1.
  4. Flag if the emitted row count differs from the full input row count (especially by orders of magnitude) with no justification in the task.
Discriminator
A real violation is an unexplained shrinkage (round numbers like 5,000 or 10,000, or a count that doesn't match the loaded shape) for a task that requires labels for all records. It is not a violation if the task explicitly permits subsampling, or if rows were dropped by a required filtering/missing-value rule that the script documents and the answer states.
Consequence
The graded file fails a shape/row-count or per-row alignment check against the reference, so the result is marked wrong even if the modeling choices were reasonable.
id 039326dd0e60 · mined from da-code dacode-ml-cluster-010@s4
raw text (what the judge reads)
### Output row count silently smaller than the input dataset (partial/truncated results)
- **Applies when**: `task` -- The task asks for a per-record output artifact (e.g., a label/prediction/score for every observation) written to a file, and the script loads a large source table and writes a derived table.
- **Pattern**: The attempt processes or emits only a subset of the records (a sampled/truncated/head-limited slice, an early-stopped loop, or a partially written file) and reports summary counts far below the documented size of the source data, without ever comparing the output row count to the input row count.
- **Detection procedure**:
  1. From the task/README, note the expected scale of the data (stated number of observations, or the shape printed when the raw file is loaded).
  2. In the scripts, trace the row count from load → any filtering/sampling/slicing → the frame that is written out; check whether any step reduces rows and whether an explicit `len(output) == len(input)` assertion exists.
  3. In the reported answer, read the stated number of rows/samples and the per-group counts; sum them and compare against the expected scale from step 1.
  4. Flag if the emitted row count differs from the full input row count (especially by orders of magnitude) with no justification in the task.
- **Discriminator**: A real violation is an unexplained shrinkage (round numbers like 5,000 or 10,000, or a count that doesn't match the loaded shape) for a task that requires labels for all records. It is *not* a violation if the task explicitly permits subsampling, or if rows were dropped by a required filtering/missing-value rule that the script documents and the answer states.
- **Consequence**: The graded file fails a shape/row-count or per-row alignment check against the reference, so the result is marked wrong even if the modeling choices were reasonable.
243Deliverable file schema is asserted rather than verified against the specified column names/contenttaskda-code
Applies when
task -- the task names an output artifact and prescribes exact column names/format for it, and the agent's answer describes the analysis narratively instead of showing the written file's actual header and shape.
Pattern
The agent computes a result and claims the required file was produced, but never reads the file back to confirm it exists with exactly the prescribed column names (correct spelling, casing, indexing base, one column per feature-vector element) and no extra/index columns; identifier or unscaled/derived helper columns are silently carried through, or the values written are not the ones the columns are supposed to hold.
Detection procedure
1. From the task, write down the exact required filename, the required column names (including the indexing convention for repeated feature columns) and the intended row unit. 2. In the scripts, locate the single write call for that file and check the DataFrame passed to it: are columns renamed to the exact required names, is index=False used, are extraneous key/label columns dropped, and does the number of feature columns equal the dimensionality of the vectors actually clustered? 3. Check whether the script (or answer) re-reads the file and prints its header, dtypes, row count, and label value counts as a sanity check. 4. Compare the answer's stated schema/row count against those printed values; treat a truncated, hand-written, or unverified schema description as no verification.
Discriminator
A real violation is when no code/output demonstrates the saved file's header and shape matching the specification (or the write clearly emits different/extra columns); a look-alike that is fine is an answer whose prose is terse but whose script renames columns explicitly, writes without the index, and prints a read-back of head/shape confirming the required names.
Consequence
The grader checking the expected artifact reports the file as WRONG/MISSING (schema or content mismatch), scoring 0 regardless of how reasonable the clustering methodology was.
id 9055bcdf7696 · mined from da-code dacode-ml-cluster-019@s4
raw text (what the judge reads)
### Deliverable file schema is asserted rather than verified against the specified column names/content
- **Applies when**: `task` -- the task names an output artifact and prescribes exact column names/format for it, and the agent's answer describes the analysis narratively instead of showing the written file's actual header and shape.
- **Pattern**: The agent computes a result and claims the required file was produced, but never reads the file back to confirm it exists with exactly the prescribed column names (correct spelling, casing, indexing base, one column per feature-vector element) and no extra/index columns; identifier or unscaled/derived helper columns are silently carried through, or the values written are not the ones the columns are supposed to hold.
- **Detection procedure**: 1. From the task, write down the exact required filename, the required column names (including the indexing convention for repeated feature columns) and the intended row unit. 2. In the scripts, locate the single write call for that file and check the DataFrame passed to it: are columns renamed to the exact required names, is `index=False` used, are extraneous key/label columns dropped, and does the number of feature columns equal the dimensionality of the vectors actually clustered? 3. Check whether the script (or answer) re-reads the file and prints its header, dtypes, row count, and label value counts as a sanity check. 4. Compare the answer's stated schema/row count against those printed values; treat a truncated, hand-written, or unverified schema description as no verification.
- **Discriminator**: A real violation is when no code/output demonstrates the saved file's header and shape matching the specification (or the write clearly emits different/extra columns); a look-alike that is fine is an answer whose prose is terse but whose script renames columns explicitly, writes without the index, and prints a read-back of head/shape confirming the required names.
- **Consequence**: The grader checking the expected artifact reports the file as WRONG/MISSING (schema or content mismatch), scoring 0 regardless of how reasonable the clustering methodology was.
244Hardcoded / recalled data instead of the provided input filestaskda-code
Applies when
task -- the task ships a dataset (with a README or data directory) and the script must compute a statistic from it.
Pattern
The script embeds numbers typed inline from the agent's prior knowledge of the topic (or a summary table) rather than reading the supplied data files, so the analysis runs on a different, coarser, or invented dataset than the one being graded — and the reported statistic cannot match the expected values regardless of how correct the method is.
Detection procedure
  1. Read the task/README and list the data artifacts the agent is expected to consume (file paths, granularity, key fields).
  2. Scan the scripts for any file-reading call (read_csv, read_parquet, open, glob, etc.); if the input is instead a literal DataFrame/dict/array of numbers, flag it.
  3. Check whether the inline numbers' granularity and row count plausibly match the shipped data (e.g., aggregated periods vs. per-record rows) and whether the script ever prints shapes/counts to verify.
  4. Confirm the reported answer (and the output file) derives from that inline data, with no cross-check against the real files.
Discriminator
A real violation is when the substantive input values are typed in and the shipped files are never opened or validated. It is fine if literals are only auxiliary constants (thresholds, seeds, column name lists, cutoff dates) while the actual records come from the provided files, or if inline values are used solely as a sanity cross-check against loaded data.
Consequence
The requested output file contains numbers computed from a fabricated dataset, so the graded values fall outside the expected tolerance and the check fails even though the bootstrap/CI machinery itself looks reasonable.
id 94a7b4e0fc3c · mined from da-code dacode-data-sa-031@s4
raw text (what the judge reads)
### Hardcoded / recalled data instead of the provided input files
- **Applies when**: `task` -- the task ships a dataset (with a README or data directory) and the script must compute a statistic from it.
- **Pattern**: The script embeds numbers typed inline from the agent's prior knowledge of the topic (or a summary table) rather than reading the supplied data files, so the analysis runs on a different, coarser, or invented dataset than the one being graded — and the reported statistic cannot match the expected values regardless of how correct the method is.
- **Detection procedure**:
  1. Read the task/README and list the data artifacts the agent is expected to consume (file paths, granularity, key fields).
  2. Scan the scripts for any file-reading call (`read_csv`, `read_parquet`, `open`, glob, etc.); if the input is instead a literal `DataFrame`/dict/array of numbers, flag it.
  3. Check whether the inline numbers' granularity and row count plausibly match the shipped data (e.g., aggregated periods vs. per-record rows) and whether the script ever prints shapes/counts to verify.
  4. Confirm the reported answer (and the output file) derives from that inline data, with no cross-check against the real files.
- **Discriminator**: A real violation is when the substantive input values are typed in and the shipped files are never opened or validated. It is fine if literals are only auxiliary constants (thresholds, seeds, column name lists, cutoff dates) while the actual records come from the provided files, or if inline values are used solely as a sanity cross-check against loaded data.
- **Consequence**: The requested output file contains numbers computed from a fabricated dataset, so the graded values fall outside the expected tolerance and the check fails even though the bootstrap/CI machinery itself looks reasonable.
245No hold-out evaluation with the competition's own metric before choosing the final modeltaskda-code
Applies when
task -- the task names a specific scoring metric for a prediction file, and the scripts train one or more models and write predictions directly to the submission file.
Pattern
The agent fits models on 100% of the training data, changes model families/hyperparameters/ensemble weights by intuition, and only "validates" the output file's shape, column names, value ranges and row sums — never computing the stated metric on any held-out or cross-validated split. Imported CV utilities are left unused, so there is no evidence any variant is better than a trivial baseline (e.g., class-prior constants), and probability calibration/quality is unchecked.
Detection procedure
  1. Read the task and note the exact evaluation metric and its inputs (probabilities vs. labels, per-row normalization, clipping).
  2. Scan every script for a split (train/validation, K-fold) plus an explicit computation of that metric; note whether any printed number is that metric.
  3. Check whether competing model/hyperparameter/ensemble-weight choices in later scripts are justified by such a number, or only by "improved" naming and eyeballed outputs.
  4. Check whether the final artifact is compared against a trivial baseline (constant class frequencies) under the same metric.
Discriminator
A real violation is when no estimate of the stated metric exists anywhere, so the reported model is unfalsified guessing; it is not a violation if the agent computed the metric on a proper hold-out/CV (even if it also refits on full data afterwards), nor if the task truly has no defined metric and only file format matters.
Consequence
The submitted probabilities can be badly calibrated or worse than a constant-prior baseline; the leaderboard log loss exceeds the acceptance threshold and the submission file is graded WRONG despite passing all format/shape sanity checks.
id 680152a2aa2e · mined from da-code dacode-ml-competition-005@s5
raw text (what the judge reads)
### No hold-out evaluation with the competition's own metric before choosing the final model
- **Applies when**: `task` -- the task names a specific scoring metric for a prediction file, and the scripts train one or more models and write predictions directly to the submission file.
- **Pattern**: The agent fits models on 100% of the training data, changes model families/hyperparameters/ensemble weights by intuition, and only "validates" the output file's shape, column names, value ranges and row sums — never computing the stated metric on any held-out or cross-validated split. Imported CV utilities are left unused, so there is no evidence any variant is better than a trivial baseline (e.g., class-prior constants), and probability calibration/quality is unchecked.
- **Detection procedure**:
  1. Read the task and note the exact evaluation metric and its inputs (probabilities vs. labels, per-row normalization, clipping).
  2. Scan every script for a split (train/validation, K-fold) plus an explicit computation of that metric; note whether any printed number is that metric.
  3. Check whether competing model/hyperparameter/ensemble-weight choices in later scripts are justified by such a number, or only by "improved" naming and eyeballed outputs.
  4. Check whether the final artifact is compared against a trivial baseline (constant class frequencies) under the same metric.
- **Discriminator**: A real violation is when *no* estimate of the stated metric exists anywhere, so the reported model is unfalsified guessing; it is *not* a violation if the agent computed the metric on a proper hold-out/CV (even if it also refits on full data afterwards), nor if the task truly has no defined metric and only file format matters.
- **Consequence**: The submitted probabilities can be badly calibrated or worse than a constant-prior baseline; the leaderboard log loss exceeds the acceptance threshold and the submission file is graded WRONG despite passing all format/shape sanity checks.
246Requested output artifact never verified against the spectaskda-code
Applies when
task -- The task asks for predictions/results to be written to a named output file with a specified column name (and implicitly one row per input record), and the agent reports a prose summary of its modeling work.
Pattern
The attempt focuses on model building and validation metrics, and treats "reported a good score" as task completion — the write step is absent, unsaved, or unchecked, so the artifact may be missing, have a different filename/column header, a different row count or ordering than the input, or contain NaN/negative/implausible values.
Detection procedure
  1. From the task, extract the exact deliverable contract: output filename, required column name(s), expected number of rows (= rows in the input file), and any ordering/rounding/unit constraints.
  2. In the scripts, locate the line that writes the file; confirm it writes to that exact path, uses exactly the required column name, is produced from predictions on the full input file (no rows dropped by dropna/filtering/merges), and does not include an index column unless requested.
  3. Confirm the script (or answer) contains an explicit post-write sanity check: re-read the file, assert row count matches the input, assert column names, and check value ranges (non-negative, no NaN, plausible magnitude vs. training label distribution).
  4. Check the final answer references the actual saved artifact and its verified shape/head — not only training/validation metrics.
Discriminator
A real violation is when no write-and-verify path is demonstrable (no saved script, no shape/column/NaN assertion, or a row count that can silently differ from the input after preprocessing). It is not a violation if the file is written for every input row with the exact required header and the agent shows a read-back check, even if the model itself is simple or the reported metrics are modest.
Consequence
The grader looks for the named file with the named column and matching rows; a missing file, renamed column, dropped/reordered rows, or NaN entries makes the deliverable unreadable or misaligned, scoring 0 regardless of how strong the reported validation metric was.
id 0ca31c1981f6 · mined from da-code dacode-ml-regression-008@s5
raw text (what the judge reads)
### Requested output artifact never verified against the spec
- **Applies when**: `task` -- The task asks for predictions/results to be written to a named output file with a specified column name (and implicitly one row per input record), and the agent reports a prose summary of its modeling work.
- **Pattern**: The attempt focuses on model building and validation metrics, and treats "reported a good score" as task completion — the write step is absent, unsaved, or unchecked, so the artifact may be missing, have a different filename/column header, a different row count or ordering than the input, or contain NaN/negative/implausible values.
- **Detection procedure**:
  1. From the task, extract the exact deliverable contract: output filename, required column name(s), expected number of rows (= rows in the input file), and any ordering/rounding/unit constraints.
  2. In the scripts, locate the line that writes the file; confirm it writes to that exact path, uses exactly the required column name, is produced from predictions on the full input file (no rows dropped by dropna/filtering/merges), and does not include an index column unless requested.
  3. Confirm the script (or answer) contains an explicit post-write sanity check: re-read the file, assert row count matches the input, assert column names, and check value ranges (non-negative, no NaN, plausible magnitude vs. training label distribution).
  4. Check the final answer references the actual saved artifact and its verified shape/head — not only training/validation metrics.
- **Discriminator**: A real violation is when no write-and-verify path is demonstrable (no saved script, no shape/column/NaN assertion, or a row count that can silently differ from the input after preprocessing). It is *not* a violation if the file is written for every input row with the exact required header and the agent shows a read-back check, even if the model itself is simple or the reported metrics are modest.
- **Consequence**: The grader looks for the named file with the named column and matching rows; a missing file, renamed column, dropped/reordered rows, or NaN entries makes the deliverable unreadable or misaligned, scoring 0 regardless of how strong the reported validation metric was.
247Blindly applying the default test to the full raw dataset without establishing the target population or the test's assumptionstaskda-code
Applies when
task -- the task asks for a specific statistic/test result (p-value, decision, metric) and the scripts compute it directly from the raw files with a single off-the-shelf function call.
Pattern
The script loads everything, derives the quantity of interest, and immediately calls the most generic parametric/two-sided routine on the entire dataset — never checking whether the task (or the data's structure: date ranges, categories/tournament labels, group sizes, skewed count distributions) implies a restricted subset, a directional alternative, or a distribution-appropriate (e.g., rank-based / unequal-variance) test.
Detection procedure
  1. Read the task statement and README and list every implicit scoping or framing clue (time window, category of records, direction of the claimed effect, "current"/"recent"/"top-level" wording) that would narrow the rows or change the test form.
  2. In the scripts, look for any filtering, grouping, or assumption-checking step (date/category filters, normality or variance inspection, choice of one- vs two-sided, choice of parametric vs nonparametric) and note which of the clues from step 1 are unaddressed.
  3. Check whether the script prints row counts / distribution diagnostics for the analyzed subset and compares them against what the task's framing would imply; a script that reports counts equal to the full file size is using the unrestricted population.
  4. Inspect the reported number for implausible extremes (e.g., a p-value many orders of magnitude below any plausible value) that indicate a huge, unfiltered sample rather than the intended one.
Discriminator
A real violation is when the task or data contains a concrete scoping/direction/assumption cue that the script ignores, so the statistic is computed on a different population or under a different alternative than requested. It is not a violation if the task genuinely asks about all records and the script documents (with printed diagnostics) that the default two-sided parametric test is appropriate for the observed distributions and sample sizes.
Consequence
The p-value (and possibly the reject/fail-to-reject decision) differs from the reference answer, so the exact-value check on the output file fails even though the file format is correct.
id 892d3074f23f · mined from da-code dacode-data-sa-001@s5
raw text (what the judge reads)
### Blindly applying the default test to the full raw dataset without establishing the target population or the test's assumptions
- **Applies when**: `task` -- the task asks for a specific statistic/test result (p-value, decision, metric) and the scripts compute it directly from the raw files with a single off-the-shelf function call.
- **Pattern**: The script loads everything, derives the quantity of interest, and immediately calls the most generic parametric/two-sided routine on the entire dataset — never checking whether the task (or the data's structure: date ranges, categories/tournament labels, group sizes, skewed count distributions) implies a restricted subset, a directional alternative, or a distribution-appropriate (e.g., rank-based / unequal-variance) test.
- **Detection procedure**:
  1. Read the task statement and README and list every implicit scoping or framing clue (time window, category of records, direction of the claimed effect, "current"/"recent"/"top-level" wording) that would narrow the rows or change the test form.
  2. In the scripts, look for any filtering, grouping, or assumption-checking step (date/category filters, normality or variance inspection, choice of one- vs two-sided, choice of parametric vs nonparametric) and note which of the clues from step 1 are unaddressed.
  3. Check whether the script prints row counts / distribution diagnostics for the analyzed subset and compares them against what the task's framing would imply; a script that reports counts equal to the full file size is using the unrestricted population.
  4. Inspect the reported number for implausible extremes (e.g., a p-value many orders of magnitude below any plausible value) that indicate a huge, unfiltered sample rather than the intended one.
- **Discriminator**: A real violation is when the task or data contains a concrete scoping/direction/assumption cue that the script ignores, so the statistic is computed on a different population or under a different alternative than requested. It is *not* a violation if the task genuinely asks about all records and the script documents (with printed diagnostics) that the default two-sided parametric test is appropriate for the observed distributions and sample sizes.
- **Consequence**: The p-value (and possibly the reject/fail-to-reject decision) differs from the reference answer, so the exact-value check on the output file fails even though the file format is correct.
248Never opening the provided reference/template file to derive output contract and computation conventionstaskda-code
Applies when
task -- the task says to write results into an output file that must follow the "exact structure and formatting" of a supplied sample/template file, and the scripts build that output themselves.
Pattern
The attempt hard-codes the assumed header names, column order, row order, rounding/precision and value conventions from memory or from the task prose, never loading or printing the sample file, and never diffing its own output's schema/formatting (and the implied definition of derived quantities, e.g. whether rounding is applied per row before aggregation) against it.
Detection procedure
  1. Read the task for any reference file, template, or "same format as" instruction, plus explicit constraints (filters, rounding, units, ordering).
  2. Search the scripts for any read/inspection of that reference file and for a post-write comparison (header equality, column order, dtype/decimal formatting, row count and row ordering).
  3. If absent, check whether the derived-metric formula and rounding step are simply assumed; note that any mismatch in aggregation convention or formatting cannot be caught by the script's own prints.
  4. Confirm the answer was accepted with no independent cross-check (e.g. recomputing totals a second way, or verifying grand totals/row counts against the raw tables).
Discriminator
A fine attempt either loads the template and programmatically asserts schema/format agreement (or shows the template contents and manually aligns every field, rounding rule and ordering); a violation only re-prints its own computed frame, which looks self-consistent regardless of whether it matches the required contract or the intended metric definition.
Consequence
The written file is judged WRONG/MISSING on exact-match comparison — differing values from an unvalidated aggregation/rounding convention, or a header/column/row-order mismatch — even though the script's own logs look clean.
id 6344c1da65e4 · mined from da-code dacode-dm-csv-011@s5
raw text (what the judge reads)
### Never opening the provided reference/template file to derive output contract and computation conventions
- **Applies when**: `task` -- the task says to write results into an output file that must follow the "exact structure and formatting" of a supplied sample/template file, and the scripts build that output themselves.
- **Pattern**: The attempt hard-codes the assumed header names, column order, row order, rounding/precision and value conventions from memory or from the task prose, never loading or printing the sample file, and never diffing its own output's schema/formatting (and the implied definition of derived quantities, e.g. whether rounding is applied per row before aggregation) against it.
- **Detection procedure**:
  1. Read the task for any reference file, template, or "same format as" instruction, plus explicit constraints (filters, rounding, units, ordering).
  2. Search the scripts for any read/inspection of that reference file and for a post-write comparison (header equality, column order, dtype/decimal formatting, row count and row ordering).
  3. If absent, check whether the derived-metric formula and rounding step are simply assumed; note that any mismatch in aggregation convention or formatting cannot be caught by the script's own prints.
  4. Confirm the answer was accepted with no independent cross-check (e.g. recomputing totals a second way, or verifying grand totals/row counts against the raw tables).
- **Discriminator**: A fine attempt either loads the template and programmatically asserts schema/format agreement (or shows the template contents and manually aligns every field, rounding rule and ordering); a violation only re-prints its own computed frame, which looks self-consistent regardless of whether it matches the required contract or the intended metric definition.
- **Consequence**: The written file is judged WRONG/MISSING on exact-match comparison — differing values from an unvalidated aggregation/rounding convention, or a header/column/row-order mismatch — even though the script's own logs look clean.
249Degenerate clustering from unvetted feature matrix and unchecked cluster-count selectiontaskda-code
Applies when
task -- the task asks for an unsupervised segmentation whose per-row feature vector and label must be written to a specified output file, and the script builds the feature matrix by dumping every remaining column and picks k by a single automatic rule.
Pattern
The script drops only the obvious identifier, keeps all other columns verbatim (including zero-variance/constant bookkeeping columns, ordinal-label-encoded categoricals that impose a false numeric order, and derived date offsets computed from "now"), scales them, then selects the cluster count as the argmax of one internal index — which almost always returns the smallest candidate (k=2) — and reports success without ever checking that the columns carry information or that the resulting partition is non-trivial and interpretable.
Detection procedure
  1. Read the task/README and list which columns are genuine behavioral/demographic signal versus identifiers, constants, or metadata; note any preprocessing the data obviously needs (unordered categories → one-hot, missing values, outlier ranges).
  2. In the script, check whether variance/uniqueness of each feature is inspected and whether constant or non-informative columns are removed, and whether unordered categoricals are one-hot encoded rather than integer-label-encoded before distance-based clustering.
  3. Check the cluster-count choice: is more than one criterion used (elbow + silhouette + cluster-size balance/profile inspection), or is it a bare argmax over a range whose minimum wins? Verify the final label distribution and per-cluster feature profiles are actually examined for distinctness.
  4. Check the answer/output: does it report a low internal validity score, exactly the smallest k, or clusters that are not characterized at all — and do the written Feature_i columns correspond to a deliberately chosen, documented feature set of consistent shape/row count?
Discriminator
A real violation keeps demonstrably useless columns (constant, ID-like, arbitrarily ordinal-encoded) and/or accepts a two-cluster split with a weak score and no profile validation; a look-alike that is fine may also end up with a small k, but only after explicitly testing variance/encoding choices, comparing multiple selection criteria, and showing the clusters differ meaningfully on key variables.
Consequence
The saved file has the requested column names but an arbitrary, information-diluted feature matrix and a near-trivial partition, so any comparison of the labeling against a reasonable reference segmentation (cluster count, agreement/quality thresholds, or feature-set expectations) fails and the output file is marked WRONG.
id f37a97ae18a2 · mined from da-code dacode-ml-cluster-014@s5
raw text (what the judge reads)
### Degenerate clustering from unvetted feature matrix and unchecked cluster-count selection
- **Applies when**: `task` -- the task asks for an unsupervised segmentation whose per-row feature vector and label must be written to a specified output file, and the script builds the feature matrix by dumping every remaining column and picks k by a single automatic rule.
- **Pattern**: The script drops only the obvious identifier, keeps all other columns verbatim (including zero-variance/constant bookkeeping columns, ordinal-label-encoded categoricals that impose a false numeric order, and derived date offsets computed from "now"), scales them, then selects the cluster count as the argmax of one internal index — which almost always returns the smallest candidate (k=2) — and reports success without ever checking that the columns carry information or that the resulting partition is non-trivial and interpretable.
- **Detection procedure**:
  1. Read the task/README and list which columns are genuine behavioral/demographic signal versus identifiers, constants, or metadata; note any preprocessing the data obviously needs (unordered categories → one-hot, missing values, outlier ranges).
  2. In the script, check whether variance/uniqueness of each feature is inspected and whether constant or non-informative columns are removed, and whether unordered categoricals are one-hot encoded rather than integer-label-encoded before distance-based clustering.
  3. Check the cluster-count choice: is more than one criterion used (elbow + silhouette + cluster-size balance/profile inspection), or is it a bare argmax over a range whose minimum wins? Verify the final label distribution and per-cluster feature profiles are actually examined for distinctness.
  4. Check the answer/output: does it report a low internal validity score, exactly the smallest k, or clusters that are not characterized at all — and do the written `Feature_i` columns correspond to a deliberately chosen, documented feature set of consistent shape/row count?
- **Discriminator**: A real violation keeps demonstrably useless columns (constant, ID-like, arbitrarily ordinal-encoded) and/or accepts a two-cluster split with a weak score and no profile validation; a look-alike that is fine may also end up with a small k, but only after explicitly testing variance/encoding choices, comparing multiple selection criteria, and showing the clusters differ meaningfully on key variables.
- **Consequence**: The saved file has the requested column names but an arbitrary, information-diluted feature matrix and a near-trivial partition, so any comparison of the labeling against a reasonable reference segmentation (cluster count, agreement/quality thresholds, or feature-set expectations) fails and the output file is marked WRONG.
250Ignoring the provided sample submission as the format/location contracttaskda-code
Applies when
task -- the task points to a template/example output file (or an explicit output path and schema) that the deliverable must match, and the scripts build that deliverable themselves.
Pattern
The scripts never load or compare against the template file; the output path, column names/order, row count/id ordering, and value dtype are assumed from memory, and the agent additionally post-processes predictions (e.g., rounding to integers, clipping to a hand-picked range) without any evidence from the template or the stated metric that such transformation is expected.
Detection procedure
  1. Read the task for the named template/example file and the required output location, then grep the scripts for any read of that file or any explicit format check.
  2. If absent, inspect how the deliverable is constructed: which directory it is written to, which columns and order are used, and whether ids are aligned to the test rows.
  3. Check for value-altering steps applied after prediction (rounding, casting to int, clipping) and ask whether the task/template/metric requires discrete or bounded values.
  4. Confirm the answer text contains an actual verification (columns and row count compared to the template, path as specified) rather than only self-reported validation scores.
Discriminator
A real violation is when no script ever reads/asserts against the template or the required path and dtype-changing post-processing is introduced on faith; it is fine if the scripts load the template, reindex/merge on its ids, assert equal columns and row counts, and justify any rounding from the template's own value type or the stated evaluation metric.
Consequence
The graded file is judged WRONG/MISSING — wrong location, wrong column set/order, misaligned ids, or degraded scores from needlessly discretized predictions — even though the reported internal validation metrics look reasonable.
id c919a65a6fda · mined from da-code dacode-ml-competition-009@s5
raw text (what the judge reads)
### Ignoring the provided sample submission as the format/location contract
- **Applies when**: `task` -- the task points to a template/example output file (or an explicit output path and schema) that the deliverable must match, and the scripts build that deliverable themselves.
- **Pattern**: The scripts never load or compare against the template file; the output path, column names/order, row count/id ordering, and value dtype are assumed from memory, and the agent additionally post-processes predictions (e.g., rounding to integers, clipping to a hand-picked range) without any evidence from the template or the stated metric that such transformation is expected.
- **Detection procedure**:
  1. Read the task for the named template/example file and the required output location, then grep the scripts for any read of that file or any explicit format check.
  2. If absent, inspect how the deliverable is constructed: which directory it is written to, which columns and order are used, and whether ids are aligned to the test rows.
  3. Check for value-altering steps applied after prediction (rounding, casting to int, clipping) and ask whether the task/template/metric requires discrete or bounded values.
  4. Confirm the answer text contains an actual verification (columns and row count compared to the template, path as specified) rather than only self-reported validation scores.
- **Discriminator**: A real violation is when no script ever reads/asserts against the template or the required path and dtype-changing post-processing is introduced on faith; it is fine if the scripts load the template, reindex/merge on its ids, assert equal columns and row counts, and justify any rounding from the template's own value type or the stated evaluation metric.
- **Consequence**: The graded file is judged WRONG/MISSING — wrong location, wrong column set/order, misaligned ids, or degraded scores from needlessly discretized predictions — even though the reported internal validation metrics look reasonable.
251Silently shrinking the dataset (or the feature set) instead of preserving every input recordtaskda-code
Applies when
task -- the task asks for a per-row output (labels, predictions, scores) written for "the dataset", and the script does row filtering (dropna(), subsetting) or hand-picks a few columns before fitting.
Pattern
The attempt drops all rows containing any missing value and/or keeps an arbitrary handful of numeric columns (sometimes including meaningless ID-like fields), then writes an output file whose row count is far smaller than the source data and whose feature columns do not correspond to the full usable feature vector — so the result cannot be aligned row-by-row with the expected output.
Detection procedure
  1. From the task, note whether the deliverable is one output row per input record and whether any filtering/subsetting was explicitly authorized.
  2. In the script, find every operation that changes row count or column set (dropna, boolean masks, explicit column lists) and check whether missing values could instead be imputed and whether non-informative/ID columns are excluded while genuine features are kept.
  3. Compare the reported/produced output shape against the raw input shape; flag any unexplained reduction in rows, and any feature count that is a small fraction of the available informative columns.
  4. Check the answer for a sanity statement reconciling input rows → output rows (and column naming/ordering matching the requested format).
Discriminator
A real violation is an unrequested reduction (rows lost to dropna, features chosen ad hoc) with no reconciliation to the original size; it is fine if the task itself specifies the filter, or if dropped columns are non-numeric/identifier fields while all rows are retained via imputation and the count is verified.
Consequence
The output file has the wrong number of rows and/or wrong feature columns, so row-wise comparison with the reference labeling fails and the file is marked WRONG/MISSING.
id 93e2a0309653 · mined from da-code dacode-ml-cluster-009@s5
raw text (what the judge reads)
### Silently shrinking the dataset (or the feature set) instead of preserving every input record
- **Applies when**: `task` -- the task asks for a per-row output (labels, predictions, scores) written for "the dataset", and the script does row filtering (`dropna()`, subsetting) or hand-picks a few columns before fitting.
- **Pattern**: The attempt drops all rows containing any missing value and/or keeps an arbitrary handful of numeric columns (sometimes including meaningless ID-like fields), then writes an output file whose row count is far smaller than the source data and whose feature columns do not correspond to the full usable feature vector — so the result cannot be aligned row-by-row with the expected output.
- **Detection procedure**:
  1. From the task, note whether the deliverable is one output row per input record and whether any filtering/subsetting was explicitly authorized.
  2. In the script, find every operation that changes row count or column set (`dropna`, boolean masks, explicit column lists) and check whether missing values could instead be imputed and whether non-informative/ID columns are excluded while genuine features are kept.
  3. Compare the reported/produced output shape against the raw input shape; flag any unexplained reduction in rows, and any feature count that is a small fraction of the available informative columns.
  4. Check the answer for a sanity statement reconciling input rows → output rows (and column naming/ordering matching the requested format).
- **Discriminator**: A real violation is an *unrequested* reduction (rows lost to `dropna`, features chosen ad hoc) with no reconciliation to the original size; it is fine if the task itself specifies the filter, or if dropped columns are non-numeric/identifier fields while all rows are retained via imputation and the count is verified.
- **Consequence**: The output file has the wrong number of rows and/or wrong feature columns, so row-wise comparison with the reference labeling fails and the file is marked WRONG/MISSING.
252Deliverable file asserted as correct without schema/content verificationtaskda-code
Applies when
task -- The task requires writing results to a specific output file with specified column names/format, and the agent's answer declares the file was produced.
Pattern
The attempt reports methodology and cluster/prediction summaries in prose, but neither the scripts nor the answer contain a read-back check that the written file actually exists at the expected path and matches the required schema (exact column names, one row per requested unit, label column present, no index column, no NaNs, row count equal to the number of entities analyzed). Truncated or inconsistent summary numbers (e.g., cluster sizes that don't sum to the stated total) are asserted rather than derived from the saved file.
Detection procedure
  1. From the task, list the exact required deliverable: file name/path, required column names and their naming convention, and the intended number of rows.
  2. In the scripts, locate the write call and check for a subsequent re-load of the written file plus explicit assertions on columns, dtypes, shape, and label validity; also check the file is written where the grader looks (working directory vs. elsewhere).
  3. In the answer, check whether every reported figure (row counts, cluster sizes, feature names) is consistent internally and traceable to the saved file rather than to an in-memory intermediate.
  4. Flag if no read-back verification exists, if the column naming/indexing convention differs from the spec, or if the reported counts/percentages are inconsistent or incomplete.
Discriminator
A real violation is an unverified or spec-deviating deliverable (missing/extra columns, wrong naming or indexing, index written as a column, wrong row granularity, wrong location, numbers that don't add up). A look-alike that is fine reloads the saved file and prints its head/shape/columns/label counts, which match the spec exactly, even if the modeling choices (feature set, k, algorithm) are debatable.
Consequence
The grader reads the expected output file and finds it missing, mis-named, or with non-conforming columns/rows, scoring the deliverable WRONG/MISSING despite a confident "task completed successfully" report.
id 780e75e301e7 · mined from da-code dacode-ml-cluster-016@s5
raw text (what the judge reads)
### Deliverable file asserted as correct without schema/content verification
- **Applies when**: `task` -- The task requires writing results to a specific output file with specified column names/format, and the agent's answer declares the file was produced.
- **Pattern**: The attempt reports methodology and cluster/prediction summaries in prose, but neither the scripts nor the answer contain a read-back check that the written file actually exists at the expected path and matches the required schema (exact column names, one row per requested unit, label column present, no index column, no NaNs, row count equal to the number of entities analyzed). Truncated or inconsistent summary numbers (e.g., cluster sizes that don't sum to the stated total) are asserted rather than derived from the saved file.
- **Detection procedure**:
  1. From the task, list the exact required deliverable: file name/path, required column names and their naming convention, and the intended number of rows.
  2. In the scripts, locate the write call and check for a subsequent re-load of the written file plus explicit assertions on columns, dtypes, shape, and label validity; also check the file is written where the grader looks (working directory vs. elsewhere).
  3. In the answer, check whether every reported figure (row counts, cluster sizes, feature names) is consistent internally and traceable to the saved file rather than to an in-memory intermediate.
  4. Flag if no read-back verification exists, if the column naming/indexing convention differs from the spec, or if the reported counts/percentages are inconsistent or incomplete.
- **Discriminator**: A real violation is an unverified or spec-deviating deliverable (missing/extra columns, wrong naming or indexing, index written as a column, wrong row granularity, wrong location, numbers that don't add up). A look-alike that is fine reloads the saved file and prints its head/shape/columns/label counts, which match the spec exactly, even if the modeling choices (feature set, k, algorithm) are debatable.
- **Consequence**: The grader reads the expected output file and finds it missing, mis-named, or with non-conforming columns/rows, scoring the deliverable WRONG/MISSING despite a confident "task completed successfully" report.
253Fabricated/hard-coded data instead of loading the provided datasettaskda-code
Applies when
task -- the task points to a provided dataset (plus a config/spec file) and asks for a chart, table, or metric derived from it.
Pattern
The scripts never read the actual data file(s); they hard-code "typical", "approximate", or invented values (often with comments like "based on public records" or "sample data") and feed those into the required output, so the output's format may look right while every value is wrong. Related symptom: required auxiliary artifacts (e.g., a JSON/NPY dump of the plotted series) are never produced.
Detection procedure
  1. Read the task and note which input files exist and which output artifacts are expected.
  2. Scan every script for a read of the actual data source (read_csv/read_excel/load/SQL/etc.) that flows into the plotted or reported values; a config/spec file being loaded does not count.
  3. Check whether numeric series are literal lists in code, and whether their provenance is a comment/assumption rather than a computation.
  4. Check the answer/verification script: does it assert correctness only from formatting (size, title, colors) and never compare values to the source data or emit the required data artifacts?
Discriminator
Legitimate hard-coded constants are thresholds, labels, axis limits, or config echoes; a violation is when the measured quantities themselves are literals with no derivation from the input files. Also fine: literals produced by an earlier script that provably read the data and printed them.
Consequence
Format-level checks may pass, but any value-level check (exported series arrays, JSON of plotted points, metric comparison) fails outright, and missing derived artifacts count as wrong/missing.
id 684d8b202784 · mined from da-code dacode-plot-line-015@s5
raw text (what the judge reads)
### Fabricated/hard-coded data instead of loading the provided dataset
- **Applies when**: `task` -- the task points to a provided dataset (plus a config/spec file) and asks for a chart, table, or metric derived from it.
- **Pattern**: The scripts never read the actual data file(s); they hard-code "typical", "approximate", or invented values (often with comments like "based on public records" or "sample data") and feed those into the required output, so the output's format may look right while every value is wrong. Related symptom: required auxiliary artifacts (e.g., a JSON/NPY dump of the plotted series) are never produced.
- **Detection procedure**:
  1. Read the task and note which input files exist and which output artifacts are expected.
  2. Scan every script for a read of the actual data source (`read_csv`/`read_excel`/`load`/SQL/etc.) that flows into the plotted or reported values; a config/spec file being loaded does not count.
  3. Check whether numeric series are literal lists in code, and whether their provenance is a comment/assumption rather than a computation.
  4. Check the answer/verification script: does it assert correctness only from formatting (size, title, colors) and never compare values to the source data or emit the required data artifacts?
- **Discriminator**: Legitimate hard-coded constants are thresholds, labels, axis limits, or config echoes; a violation is when the *measured quantities themselves* are literals with no derivation from the input files. Also fine: literals produced by an earlier script that provably read the data and printed them.
- **Consequence**: Format-level checks may pass, but any value-level check (exported series arrays, JSON of plotted points, metric comparison) fails outright, and missing derived artifacts count as wrong/missing.
254Ignoring provided instruction/config files and their required output artifactstaskda-code
Applies when
task -- the task tells the agent to follow separate specification files (e.g., a tips/instructions file and a config/format file) and/or to emit outputs in a prescribed set of files.
Pattern
The script never opens or parses the referenced spec/config files; instead it hardcodes the agent's own guesses for filtering rules, grouping definitions, styling, axis/label/series conventions, and saves only the one obviously-named artifact, silently skipping the companion machine-checkable outputs (serialized plot spec, numeric array, etc.).
Detection procedure
  1. From the task text, list every referenced instruction/config file and every output artifact that is named or implied by the required format.
  2. Grep the scripts for reads of those spec files (open/read/yaml.safe_load/json.load) and for writes of each expected artifact.
  3. For every rule the script hardcodes (exclusions, category definitions, units, ordering, figure/style settings), check whether it is quoted from the spec file or invented — invented values that "match the spec" only per the agent's prose summary count as unverified.
  4. Check the answer text: does it enumerate all produced files and show that each spec directive was applied, or does it only assert one image was saved?
Discriminator
A real violation is a script that produces fewer artifacts than requested or derives spec-governed parameters from nothing readable in the code path; it is fine if the script loads the config and programmatically drives its parameters from it (even if it also writes extra files or restates the rules in prose).
Consequence
The grader finds required artifacts missing or mismatched (no serialized plot spec / numeric array) and the image itself deviates from the mandated formatting and grouping conventions, so all checks fail even though the plot "looks" reasonable.
id 4b54d8539257 · mined from da-code dacode-plot-line-006@s5
raw text (what the judge reads)
### Ignoring provided instruction/config files and their required output artifacts
- **Applies when**: `task` -- the task tells the agent to follow separate specification files (e.g., a tips/instructions file and a config/format file) and/or to emit outputs in a prescribed set of files.
- **Pattern**: The script never opens or parses the referenced spec/config files; instead it hardcodes the agent's own guesses for filtering rules, grouping definitions, styling, axis/label/series conventions, and saves only the one obviously-named artifact, silently skipping the companion machine-checkable outputs (serialized plot spec, numeric array, etc.).
- **Detection procedure**:
  1. From the task text, list every referenced instruction/config file and every output artifact that is named or implied by the required format.
  2. Grep the scripts for reads of those spec files (open/read/`yaml.safe_load`/`json.load`) and for writes of each expected artifact.
  3. For every rule the script hardcodes (exclusions, category definitions, units, ordering, figure/style settings), check whether it is quoted from the spec file or invented — invented values that "match the spec" only per the agent's prose summary count as unverified.
  4. Check the answer text: does it enumerate all produced files and show that each spec directive was applied, or does it only assert one image was saved?
- **Discriminator**: A real violation is a script that produces fewer artifacts than requested or derives spec-governed parameters from nothing readable in the code path; it is fine if the script loads the config and programmatically drives its parameters from it (even if it also writes extra files or restates the rules in prose).
- **Consequence**: The grader finds required artifacts missing or mismatched (no serialized plot spec / numeric array) and the image itself deviates from the mandated formatting and grouping conventions, so all checks fail even though the plot "looks" reasonable.
255Required output artifact and exact answer schema are not producedtaskda-code
Applies when
task -- the task specifies an answer template (key names, bracketed/list-valued fields, rounding) and/or an expected result file that the graded deliverable must be written to.
Pattern
The agent reports the answer only in chat prose, or reshapes the requested schema (scalar where a list/array was shown, renamed/extra/missing keys, unrounded or differently-typed values), and never writes the specified result file from the script — so nothing exists for the grader to read even if the underlying computation happened to be right.
Detection procedure
  1. From the task text, extract the literal deliverable: file name/path expected, key names, value container types shown in the template, and any rounding/unit/ordering constraints.
  2. Read the scripts for a write step (json.dump, to_csv, etc.) that emits exactly that file, and check the object being dumped key-for-key and type-for-type against the template; if no script was saved at all, treat the deliverable as unverifiable and unreproducible.
  3. Compare the agent's final reported answer to the template: are values wrapped as shown, are all requested fields present, are numbers formatted as instructed?
  4. Flag the attempt if the required file is not written by any script, or if the reported/serialized structure deviates from the template.
Discriminator
A real violation is a missing output file or a structural/type/format mismatch with the stated template. It is not a violation if the file is written with the exact keys and containers and only cosmetic differences remain (whitespace, key order, 82.6 vs 82.60 when no rounding was specified), or if the task genuinely asked only for a chat answer with no artifact.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 regardless of analytical correctness, since it cannot match a schema it never receives.
id e68e23173710 · mined from da-code dacode-di-text-002@s5
raw text (what the judge reads)
### Required output artifact and exact answer schema are not produced
- **Applies when**: `task` -- the task specifies an answer template (key names, bracketed/list-valued fields, rounding) and/or an expected result file that the graded deliverable must be written to.
- **Pattern**: The agent reports the answer only in chat prose, or reshapes the requested schema (scalar where a list/array was shown, renamed/extra/missing keys, unrounded or differently-typed values), and never writes the specified result file from the script — so nothing exists for the grader to read even if the underlying computation happened to be right.
- **Detection procedure**:
  1. From the task text, extract the literal deliverable: file name/path expected, key names, value container types shown in the template, and any rounding/unit/ordering constraints.
  2. Read the scripts for a write step (`json.dump`, `to_csv`, etc.) that emits exactly that file, and check the object being dumped key-for-key and type-for-type against the template; if no script was saved at all, treat the deliverable as unverifiable and unreproducible.
  3. Compare the agent's final reported answer to the template: are values wrapped as shown, are all requested fields present, are numbers formatted as instructed?
  4. Flag the attempt if the required file is not written by any script, or if the reported/serialized structure deviates from the template.
- **Discriminator**: A real violation is a missing output file or a structural/type/format mismatch with the stated template. It is *not* a violation if the file is written with the exact keys and containers and only cosmetic differences remain (whitespace, key order, `82.6` vs `82.60` when no rounding was specified), or if the task genuinely asked only for a chat answer with no artifact.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 regardless of analytical correctness, since it cannot match a schema it never receives.
256Named-formula variant substituted for the one the task specifiestaskinfiagent-dabench
Applies when
task -- the task names a specific statistic/coefficient by name (e.g., a particular "first"/"second" coefficient, a specific variant of a metric or a specific denominator/normalization) and the script computes it via a library default or a related formula.
Pattern
The attempt computes a mathematically adjacent quantity — a library's default implementation, or the sibling variant of the named formula — instead of the exact definition stated. The result is the right sign/order of magnitude but a different number, so it looks plausible and passes any qualitative sub-question while failing the numeric one.
Detection procedure
  1. From the task text, write down the exact algebraic definition implied by the named statistic, including every component (which central tendency measures, which dispersion measure, sample vs population denominator, any multiplier).
  2. In the script, locate the line producing the reported number and expand it into algebra; check each component matches term-by-term (e.g., is the required central-tendency term actually computed, or was a different one silently used because it was easier/undefined?).
  3. Flag any use of a black-box library function (.skew(), .kurt(), .corr(), f1_score(...) defaults, etc.) unless the script proves its default matches the named definition.
  4. Check the answer's numeric value is reproducible from the stated definition; if two variants differ by a constant factor or a swapped term, verify which one was reported.
Discriminator
A real violation is when the computed expression differs structurally from the named definition (different term, different multiplier, different denominator convention), even if the qualitative conclusion is unchanged. It is not a violation if the script uses a library call but explicitly documents/verifies that the library's implementation equals the named formula, or if a degenerate-input fallback is handled and justified.
Consequence
The categorical/qualitative sub-answer passes while the numeric field is off by a systematic factor or offset, yielding a partial-credit grade (e.g., 1 of 2 checks) and an overall incorrect verdict.
id b5e0503b0593 · mined from infiagent-dabench dabench-359@s5
raw text (what the judge reads)
### Named-formula variant substituted for the one the task specifies
- **Applies when**: `task` -- the task names a specific statistic/coefficient by name (e.g., a particular "first"/"second" coefficient, a specific variant of a metric or a specific denominator/normalization) and the script computes it via a library default or a related formula.
- **Pattern**: The attempt computes a mathematically adjacent quantity — a library's default implementation, or the sibling variant of the named formula — instead of the exact definition stated. The result is the right sign/order of magnitude but a different number, so it looks plausible and passes any qualitative sub-question while failing the numeric one.
- **Detection procedure**:
  1. From the task text, write down the exact algebraic definition implied by the named statistic, including every component (which central tendency measures, which dispersion measure, sample vs population denominator, any multiplier).
  2. In the script, locate the line producing the reported number and expand it into algebra; check each component matches term-by-term (e.g., is the required central-tendency term actually computed, or was a different one silently used because it was easier/undefined?).
  3. Flag any use of a black-box library function (`.skew()`, `.kurt()`, `.corr()`, `f1_score(...)` defaults, etc.) unless the script proves its default matches the named definition.
  4. Check the answer's numeric value is reproducible from the stated definition; if two variants differ by a constant factor or a swapped term, verify which one was reported.
- **Discriminator**: A real violation is when the computed expression differs structurally from the named definition (different term, different multiplier, different denominator convention), even if the qualitative conclusion is unchanged. It is *not* a violation if the script uses a library call but explicitly documents/verifies that the library's implementation equals the named formula, or if a degenerate-input fallback is handled and justified.
- **Consequence**: The categorical/qualitative sub-answer passes while the numeric field is off by a systematic factor or offset, yielding a partial-credit grade (e.g., 1 of 2 checks) and an overall incorrect verdict.
257Statistical test reported without the required test output or verification of the exact data vector testedtaskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test / distribution statistic on one column with an explicit decision rule and an explicit request to report the test's p-value (or statistic).
Pattern
The attempt outputs only the derived verdict and the descriptive statistics, never printing the p-value, the sample size, or how the column was prepared (dtype coercion, NaN dropping, filtering/deduplication of rows, subset selection), so the decision rests on an unverified input vector; often the test is run on a differently-shaped vector (extra rows, un-dropped nulls, string-to-number coercion artifacts, or the wrong column of a similarly named pair) than the one the descriptive statistics describe.
Detection procedure
  1. Read the task and list every quantity it says to report (here: test p-value plus the shape statistics) and the exact decision threshold/rule.
  2. In the scripts, locate the vector fed to the test: check that it is built from the named column only, that NaN/dtype handling is explicit, and that len() / describe() of that vector is printed and matches the vector used for the shape statistics.
  3. In the answer/log, confirm the p-value (and n) is actually reported and that applying the stated threshold to it reproduces the yes/no verdict.
  4. Cross-check plausibility: for the reported n, do the shape statistics and the verdict tell a coherent story (small n makes the test low-power, so a "non-normal" verdict from shape alone is not evidence); if p-value is absent, treat the verdict as unverifiable.
Discriminator
A real violation is when the p-value/n is missing or when the tested vector's construction (filtering, NaN handling, column choice) cannot be traced from the scripts; a look-alike that is fine prints the p-value and the row count, handles missing values explicitly, and its verdict follows mechanically from the stated threshold even if the shape statistics look extreme.
Consequence
The grader compares the yes/no decision to ground truth and marks the whole item wrong, because the test was silently run on a different (or uncleaned) sample than intended and no reported p-value existed to catch the mismatch.
id 9402904f83a5 · mined from infiagent-dabench dabench-298@s5
raw text (what the judge reads)
### Statistical test reported without the required test output or verification of the exact data vector tested
- **Applies when**: `task` -- the task asks for a hypothesis test / distribution statistic on one column with an explicit decision rule and an explicit request to report the test's p-value (or statistic).
- **Pattern**: The attempt outputs only the derived verdict and the descriptive statistics, never printing the p-value, the sample size, or how the column was prepared (dtype coercion, NaN dropping, filtering/deduplication of rows, subset selection), so the decision rests on an unverified input vector; often the test is run on a differently-shaped vector (extra rows, un-dropped nulls, string-to-number coercion artifacts, or the wrong column of a similarly named pair) than the one the descriptive statistics describe.
- **Detection procedure**:
  1. Read the task and list every quantity it says to report (here: test p-value plus the shape statistics) and the exact decision threshold/rule.
  2. In the scripts, locate the vector fed to the test: check that it is built from the named column only, that NaN/dtype handling is explicit, and that `len()` / `describe()` of that vector is printed and matches the vector used for the shape statistics.
  3. In the answer/log, confirm the p-value (and n) is actually reported and that applying the stated threshold to it reproduces the yes/no verdict.
  4. Cross-check plausibility: for the reported n, do the shape statistics and the verdict tell a coherent story (small n makes the test low-power, so a "non-normal" verdict from shape alone is not evidence); if p-value is absent, treat the verdict as unverifiable.
- **Discriminator**: A real violation is when the p-value/n is missing or when the tested vector's construction (filtering, NaN handling, column choice) cannot be traced from the scripts; a look-alike that is fine prints the p-value and the row count, handles missing values explicitly, and its verdict follows mechanically from the stated threshold even if the shape statistics look extreme.
- **Consequence**: The grader compares the yes/no decision to ground truth and marks the whole item wrong, because the test was silently run on a different (or uncleaned) sample than intended and no reported p-value existed to catch the mismatch.
258Deliverable file never verified against the required path/template (and silently overwritten by successive runs)taskda-code
Applies when
task -- the task requires producing an output artifact (e.g., a predictions/results file) in a specified name/format matching a provided template, and the scripts write that file at the end of each of several iterations.
Pattern
The agent hardcodes an output path of its own choosing, never loads the provided template to confirm the expected directory, filename, row count, id set/order, and column names, and re-runs several different modeling scripts that each overwrite the same file — so the final artifact may be absent from the graded location, truncated, or produced by a different (possibly failed/weaker) script than the one described in the answer.
Detection procedure
  1. Read the task for the required output filename/format and the template file it references; note where the grader would look (typically the working/data directory implied by the task).
  2. In the scripts, find every write of the artifact; check whether the path is derived from/compared with the template location, and whether any code reads the template to validate columns, row count, and id alignment after writing.
  3. Check whether multiple scripts write the same artifact and whether the final answer's described method matches the last script actually executed to completion (look for truncated/incomplete scripts or unreported reruns).
  4. Confirm the answer includes an explicit post-write verification (re-read the file, print shape, head, id match with the template) rather than only pre-write prediction statistics.
Discriminator
A real violation is when no code ever reads back or cross-checks the artifact against the template/expected path, or when the described final model cannot be traced to the last successful write; it is fine if the agent writes to a nonstandard path but explicitly copies/verifies it against the template's location, columns, and ids.
Consequence
The grader reports the expected result file as WRONG/MISSING (not found at the checked path, mismatched ids/columns, or content from an unintended run), so the task scores zero regardless of model quality.
id 12a910e1ff96 · mined from da-code dacode-ml-competition-008@s5
raw text (what the judge reads)
### Deliverable file never verified against the required path/template (and silently overwritten by successive runs)

- **Applies when**: `task` -- the task requires producing an output artifact (e.g., a predictions/results file) in a specified name/format matching a provided template, and the scripts write that file at the end of each of several iterations.
- **Pattern**: The agent hardcodes an output path of its own choosing, never loads the provided template to confirm the expected directory, filename, row count, id set/order, and column names, and re-runs several different modeling scripts that each overwrite the same file — so the final artifact may be absent from the graded location, truncated, or produced by a different (possibly failed/weaker) script than the one described in the answer.
- **Detection procedure**:
  1. Read the task for the required output filename/format and the template file it references; note where the grader would look (typically the working/data directory implied by the task).
  2. In the scripts, find every write of the artifact; check whether the path is derived from/compared with the template location, and whether any code reads the template to validate columns, row count, and id alignment after writing.
  3. Check whether multiple scripts write the same artifact and whether the final answer's described method matches the last script actually executed to completion (look for truncated/incomplete scripts or unreported reruns).
  4. Confirm the answer includes an explicit post-write verification (re-read the file, print shape, head, id match with the template) rather than only pre-write prediction statistics.
- **Discriminator**: A real violation is when no code ever reads back or cross-checks the artifact against the template/expected path, or when the described final model cannot be traced to the last successful write; it is fine if the agent writes to a nonstandard path but explicitly copies/verifies it against the template's location, columns, and ids.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING (not found at the checked path, mismatched ids/columns, or content from an unintended run), so the task scores zero regardless of model quality.
259Invented filter thresholds and aggregation rules instead of using the provided spec/sample artifacttaskda-code
Applies when
task -- The task points to a reference/sample output file or a definition section that constrains how entities qualify, how values are aggregated, or how results are formatted, and the scripts must produce a ranked/summary table.
Pattern
The agent never loads or inspects the referenced sample/spec artifact; it silently picks its own qualification cutoff (e.g., a minimum-count filter), its own aggregation (mean vs. sum, per-item vs. per-group), and its own tie-breaking/column layout, then reports the ranking as if it were the required one — with no evidence that any of these choices match the stated definition.
Detection procedure
  1. Read the task statement and note every artifact it references (sample/expected output, definition text) and any definition that is partial, truncated, or ambiguous.
  2. Grep the scripts for reads of those artifacts and for hard-coded constants (thresholds, head(n), filters, rounding, column names); check whether each constant is traceable to the task text or the sample file, or was chosen by the agent.
  3. Check whether the aggregation function per group and the sort/tie-break rule are justified anywhere (comment, printed comparison, sensitivity check) rather than assumed.
  4. Inspect the produced output's header/columns/ordering and confirm the agent compared them against the sample file, not just against its own assumption.
Discriminator
A real violation is an unjustified, outcome-changing choice (a cutoff, aggregation, or format never verified against the available spec/sample); it is fine if the agent loaded the sample/spec and its constants demonstrably follow from it, or if it showed the ranking is stable across plausible alternatives.
Consequence
The output file has plausible-looking but wrong rows (and possibly wrong headers/order), so an exact-match file comparison fails on all checks despite clean-looking code.
id fff462e109db · mined from da-code dacode-dm-csv-009@s5
raw text (what the judge reads)
### Invented filter thresholds and aggregation rules instead of using the provided spec/sample artifact
- **Applies when**: `task` -- The task points to a reference/sample output file or a definition section that constrains how entities qualify, how values are aggregated, or how results are formatted, and the scripts must produce a ranked/summary table.
- **Pattern**: The agent never loads or inspects the referenced sample/spec artifact; it silently picks its own qualification cutoff (e.g., a minimum-count filter), its own aggregation (mean vs. sum, per-item vs. per-group), and its own tie-breaking/column layout, then reports the ranking as if it were the required one — with no evidence that any of these choices match the stated definition.
- **Detection procedure**:
  1. Read the task statement and note every artifact it references (sample/expected output, definition text) and any definition that is partial, truncated, or ambiguous.
  2. Grep the scripts for reads of those artifacts and for hard-coded constants (thresholds, `head(n)`, filters, rounding, column names); check whether each constant is traceable to the task text or the sample file, or was chosen by the agent.
  3. Check whether the aggregation function per group and the sort/tie-break rule are justified anywhere (comment, printed comparison, sensitivity check) rather than assumed.
  4. Inspect the produced output's header/columns/ordering and confirm the agent compared them against the sample file, not just against its own assumption.
- **Discriminator**: A real violation is an unjustified, outcome-changing choice (a cutoff, aggregation, or format never verified against the available spec/sample); it is fine if the agent loaded the sample/spec and its constants demonstrably follow from it, or if it showed the ranking is stable across plausible alternatives.
- **Consequence**: The output file has plausible-looking but wrong rows (and possibly wrong headers/order), so an exact-match file comparison fails on all checks despite clean-looking code.
260Fabricating substitute inputs and specs instead of using the provided onestaskda-code
Applies when
task -- The task references specific provided artifacts (a real dataset and/or a configuration/spec file) and requires a set of named output files.
Pattern
The agent cannot locate (or does not verify) the real input/config, so it synthesizes random stand-in data or invents default settings, then runs the full pipeline on the fake inputs and reports the resulting numbers as the answer; it also produces only some of the required output files.
Detection procedure
  1. List every input the task names (data file, config/spec file) and every output artifact the task or grading expects.
  2. Scan the scripts for data/config generation or fallback-default code paths (random generators, hard-coded "if config missing use defaults", search lists ending in a synthetic file) and check whether the actual provided file is asserted to exist.
  3. Check the reported numbers/plot parameters against the real file's known properties (row counts, category names, spec fields); look for tell-tale signs of synthetic data such as suspiciously uniform bin counts, round totals, or an empty terminal bin.
  4. Verify each expected output file is actually written by the scripts with the exact required name/format.
Discriminator
A violation is when synthetic data or invented spec values feed the reported result or the saved artifacts; it is fine to create a small mock dataset purely for a smoke test if the final run demonstrably loads the real input (e.g., a hard failure when it is absent) and reads the real spec.
Consequence
Every graded artifact mismatches ground truth — values, bin edges, and plot metadata come from a different dataset, and missing required files count as wrong/absent, yielding 0 checks passed.
id 65317c900245 · mined from da-code dacode-plot-bar-007@s5
raw text (what the judge reads)
### Fabricating substitute inputs and specs instead of using the provided ones
- **Applies when**: `task` -- The task references specific provided artifacts (a real dataset and/or a configuration/spec file) and requires a set of named output files.
- **Pattern**: The agent cannot locate (or does not verify) the real input/config, so it synthesizes random stand-in data or invents default settings, then runs the full pipeline on the fake inputs and reports the resulting numbers as the answer; it also produces only some of the required output files.
- **Detection procedure**:
  1. List every input the task names (data file, config/spec file) and every output artifact the task or grading expects.
  2. Scan the scripts for data/config generation or fallback-default code paths (random generators, hard-coded "if config missing use defaults", search lists ending in a synthetic file) and check whether the actual provided file is asserted to exist.
  3. Check the reported numbers/plot parameters against the real file's known properties (row counts, category names, spec fields); look for tell-tale signs of synthetic data such as suspiciously uniform bin counts, round totals, or an empty terminal bin.
  4. Verify each expected output file is actually written by the scripts with the exact required name/format.
- **Discriminator**: A violation is when synthetic data or invented spec values feed the reported result or the saved artifacts; it is fine to create a small mock dataset purely for a smoke test if the final run demonstrably loads the real input (e.g., a hard failure when it is absent) and reads the real spec.
- **Consequence**: Every graded artifact mismatches ground truth — values, bin edges, and plot metadata come from a different dataset, and missing required files count as wrong/absent, yielding 0 checks passed.
261Sample vs. population variant of a dispersion statistic (ddof) left unspecifiedtaskinfiagent-dabench
Applies when
task -- the task asks for a standard deviation, variance, standard error, or a z-score/normalization built from one, and the script computes it with a library default (numpy.std, pandas.Series.std, scipy.stats.zscore) without an explicit ddof.
Pattern
The attempt mixes conventions — e.g. computes the reported spread with the population formula (ddof=0) while the expected/reference convention is the sample formula (ddof=1), or vice versa — and reports a single value with no note that the two differ, even though for small n the difference is well above the requested rounding precision.
Detection procedure
  1. Read the task for any statement (or convention implied by the tooling/answer precision) about sample vs. population spread, and note the number of retained observations n.
  2. In the scripts, locate every call producing a spread statistic and check whether ddof (or the equivalent) is passed explicitly; note that NumPy defaults to 0 and pandas defaults to 1, so switching libraries silently switches the answer.
  3. Recompute the alternative convention: multiply the reported value by sqrt(n/(n-1)) (or its inverse) and compare to the reported figure at the requested rounding.
  4. Flag the attempt if the two variants differ at the reported precision and the script never justified, cross-checked, or reported the chosen convention.
Discriminator
Not a violation if n is large enough that both conventions round to the same reported value, or if the task/answer key explicitly fixes the convention and the script matches it; it is a violation when a default was inherited implicitly and the alternative changes the rounded answer — especially when the same statistic is also used inside a threshold rule (e.g., z-scores), where the convention can also change which points are flagged.
Consequence
The mean and any counts/flags may match the reference while the spread value is off by the sqrt(n/(n-1)) factor, so the grader marks the standard-deviation check WRONG and the submission fails despite otherwise correct logic.
id 626d1927e881 · mined from infiagent-dabench dabench-495@s5
raw text (what the judge reads)
### Sample vs. population variant of a dispersion statistic (ddof) left unspecified
- **Applies when**: `task` -- the task asks for a standard deviation, variance, standard error, or a z-score/normalization built from one, and the script computes it with a library default (`numpy.std`, `pandas.Series.std`, `scipy.stats.zscore`) without an explicit `ddof`.
- **Pattern**: The attempt mixes conventions — e.g. computes the reported spread with the population formula (`ddof=0`) while the expected/reference convention is the sample formula (`ddof=1`), or vice versa — and reports a single value with no note that the two differ, even though for small n the difference is well above the requested rounding precision.
- **Detection procedure**:
  1. Read the task for any statement (or convention implied by the tooling/answer precision) about sample vs. population spread, and note the number of retained observations n.
  2. In the scripts, locate every call producing a spread statistic and check whether `ddof` (or the equivalent) is passed explicitly; note that NumPy defaults to 0 and pandas defaults to 1, so switching libraries silently switches the answer.
  3. Recompute the alternative convention: multiply the reported value by `sqrt(n/(n-1))` (or its inverse) and compare to the reported figure at the requested rounding.
  4. Flag the attempt if the two variants differ at the reported precision and the script never justified, cross-checked, or reported the chosen convention.
- **Discriminator**: Not a violation if n is large enough that both conventions round to the same reported value, or if the task/answer key explicitly fixes the convention and the script matches it; it *is* a violation when a default was inherited implicitly and the alternative changes the rounded answer — especially when the same statistic is also used inside a threshold rule (e.g., z-scores), where the convention can also change which points are flagged.
- **Consequence**: The mean and any counts/flags may match the reference while the spread value is off by the `sqrt(n/(n-1))` factor, so the grader marks the standard-deviation check WRONG and the submission fails despite otherwise correct logic.
262Empty-subset filtering silently replaced by a broader/looser subsettaskinfiagent-dabench
Applies when
task -- the task asks for a statistic computed after applying two or more conjunctive filters (value conditions, group keys, date/ID restrictions) to a table.
Pattern
The attempt never checks how many rows survive the full filter chain. Either only some of the required conditions are applied, or a loose/coerced comparison (string vs numeric, isin, or instead of and, fuzzy/nearest match, fallback to the last non-empty subset) is used, so a plausible-looking number is reported even though the exactly-specified subset is empty or nearly empty. The legitimate answer for an empty subset (NaN / "no data") is never considered.
Detection procedure
  1. From the task, list every filter condition and confirm they are meant to be applied conjunctively, in the stated order.
  2. In the scripts, check that each condition appears with the correct column, correct comparison, correct dtype, and is ANDed — not dropped, relaxed, or replaced by a fallback branch.
  3. Check whether the script prints/asserts the row count (and shape) of the final filtered frame before computing the statistic, and whether it handles the count==0 case explicitly instead of defaulting to a value from a wider subset.
  4. Compare the reported number against that count: if no count evidence exists, or the count is 0/1 while a "typical-looking" aggregate is reported, flag it.
Discriminator
A real violation is one where the final conjunctive subset's size is unverified or where a non-empty answer is reported despite conditions that cannot all hold; it is not a violation if the script demonstrates a non-empty filtered subset (printed counts/head) and computes the statistic on exactly that subset — nor if a documented dtype normalization (e.g., stripping whitespace) is applied to all rows rather than used to force a match.
Consequence
The grader compares against the value implied by the exact subset (often NaN/undefined or a different aggregate); a number computed on a broader or partially filtered subset fails the check outright.
id 5c6e8692539b · mined from infiagent-dabench dabench-554@s5
raw text (what the judge reads)
### Empty-subset filtering silently replaced by a broader/looser subset
- **Applies when**: `task` -- the task asks for a statistic computed after applying two or more conjunctive filters (value conditions, group keys, date/ID restrictions) to a table.
- **Pattern**: The attempt never checks how many rows survive the full filter chain. Either only some of the required conditions are applied, or a loose/coerced comparison (string vs numeric, `isin`, `or` instead of `and`, fuzzy/nearest match, fallback to the last non-empty subset) is used, so a plausible-looking number is reported even though the exactly-specified subset is empty or nearly empty. The legitimate answer for an empty subset (NaN / "no data") is never considered.
- **Detection procedure**:
  1. From the task, list every filter condition and confirm they are meant to be applied conjunctively, in the stated order.
  2. In the scripts, check that each condition appears with the correct column, correct comparison, correct dtype, and is ANDed — not dropped, relaxed, or replaced by a fallback branch.
  3. Check whether the script prints/asserts the row count (and shape) of the final filtered frame before computing the statistic, and whether it handles the count==0 case explicitly instead of defaulting to a value from a wider subset.
  4. Compare the reported number against that count: if no count evidence exists, or the count is 0/1 while a "typical-looking" aggregate is reported, flag it.
- **Discriminator**: A real violation is one where the final conjunctive subset's size is unverified or where a non-empty answer is reported despite conditions that cannot all hold; it is *not* a violation if the script demonstrates a non-empty filtered subset (printed counts/head) and computes the statistic on exactly that subset — nor if a documented dtype normalization (e.g., stripping whitespace) is applied to *all* rows rather than used to force a match.
- **Consequence**: The grader compares against the value implied by the exact subset (often NaN/undefined or a different aggregate); a number computed on a broader or partially filtered subset fails the check outright.
263Answer not demonstrably derived from the provided data (unverifiable / prior-knowledge results)taskda-code
Applies when
task -- the task asks for specific values or entity names to be extracted from a supplied data file (with stated preprocessing such as imputation) and written to a named result artifact.
Pattern
The attempt reports a plausible-looking list of names/values without any saved, runnable script that loads the file, applies the stated preprocessing, computes the ranking/statistic, and writes the requested output file — the answer appears to come from general knowledge or an unrecorded ad-hoc step, so labels and ordering cannot be traced back to rows of the data.
Detection procedure
  1. Read the task and note the required inputs (source file, preprocessing rule), the requested ordering/format, and the expected output artifact name.
  2. Look for a script that (a) reads the source file, (b) applies the stated preprocessing, (c) computes the requested statistic/ranking, and (d) dumps the exact requested JSON/file; if no such script exists, or it hard-codes the reported values, flag immediately.
  3. Check that every reported label is a verbatim string present in the data's key column (no re-spellings, aliases, or long-form names) and that each reported list obeys the stated sort direction and length.
  4. Spot-check one or two reported entries against the data (e.g., the extreme value) to confirm they would actually survive the imputation/filtering step.
Discriminator
A real violation is an answer with no reproducible path from file to output (missing script, hard-coded lists, names not matching the data's spelling, or ordering not matching the stated direction). A look-alike that is fine is an answer that happens to match common knowledge but is produced by a script whose printed intermediate output matches the submitted values and whose labels are taken directly from the data.
Consequence
The grader compares against the expected result file and marks it wrong/missing — either the artifact was never written, the entity strings don't match the dataset's canonical names, or the list order violates the requested sort, yielding 0 checks passed.
id fe4ce9b1e9b0 · mined from da-code dacode-di-text-003@s5
raw text (what the judge reads)
### Answer not demonstrably derived from the provided data (unverifiable / prior-knowledge results)
- **Applies when**: `task` -- the task asks for specific values or entity names to be extracted from a supplied data file (with stated preprocessing such as imputation) and written to a named result artifact.
- **Pattern**: The attempt reports a plausible-looking list of names/values without any saved, runnable script that loads the file, applies the stated preprocessing, computes the ranking/statistic, and writes the requested output file — the answer appears to come from general knowledge or an unrecorded ad-hoc step, so labels and ordering cannot be traced back to rows of the data.
- **Detection procedure**:
  1. Read the task and note the required inputs (source file, preprocessing rule), the requested ordering/format, and the expected output artifact name.
  2. Look for a script that (a) reads the source file, (b) applies the stated preprocessing, (c) computes the requested statistic/ranking, and (d) dumps the exact requested JSON/file; if no such script exists, or it hard-codes the reported values, flag immediately.
  3. Check that every reported label is a verbatim string present in the data's key column (no re-spellings, aliases, or long-form names) and that each reported list obeys the stated sort direction and length.
  4. Spot-check one or two reported entries against the data (e.g., the extreme value) to confirm they would actually survive the imputation/filtering step.
- **Discriminator**: A real violation is an answer with no reproducible path from file to output (missing script, hard-coded lists, names not matching the data's spelling, or ordering not matching the stated direction). A look-alike that is fine is an answer that happens to match common knowledge but is produced by a script whose printed intermediate output matches the submitted values and whose labels are taken directly from the data.
- **Consequence**: The grader compares against the expected result file and marks it wrong/missing — either the artifact was never written, the entity strings don't match the dataset's canonical names, or the list order violates the requested sort, yielding 0 checks passed.
264Boolean/categorical values not emitted in the literal form the answer spec requirestaskinfiagent-dabench
Applies when
task -- the answer template asks for a boolean, string label, or enumerated token (e.g. true/false, yes/no, a class name) and the script interpolates a native language object into the output string.
Pattern
The script computes the right value but writes it with the language's default repr (True, nan, 1.0, Class(0)) via f-string interpolation, instead of converting it to the exact literal spelling/case/type the task specifies; no normalization or format check happens before submitting.
Detection procedure
1. Read the answer format section of the task and note the exact required literals for every non-numeric field (case, spelling, allowed set). 2. In the script, find the line that builds the final answer string and trace what object is inserted for each of those fields. 3. Simulate the interpolation mentally (e.g. Python bool → "True", numpy bool_ → "True") and compare character-for-character against the required literal. 4. Check whether any explicit mapping/lowercasing/str() normalization or a final self-check against the spec exists; if not, flag.
Discriminator
A real violation is a mismatch in the emitted characters (case or token) against a spec that names the literals; it is fine if the task does not fix the spelling, or if the script explicitly maps the value (e.g. "true" if sig else "false") or lowercases before output. Numeric fields whose rounding matches the spec are not affected by this check.
Consequence
The grader marks that field WRONG/MISSING even though the underlying computation is correct, producing a partial-credit failure (e.g. 2/3 checks passed) and an overall incorrect verdict.
id 1804fa9e5362 · mined from infiagent-dabench dabench-668@s5
raw text (what the judge reads)
### Boolean/categorical values not emitted in the literal form the answer spec requires
- **Applies when**: `task` -- the answer template asks for a boolean, string label, or enumerated token (e.g. `true`/`false`, `yes`/`no`, a class name) and the script interpolates a native language object into the output string.
- **Pattern**: The script computes the right value but writes it with the language's default repr (`True`, `nan`, `1.0`, `Class(0)`) via f-string interpolation, instead of converting it to the exact literal spelling/case/type the task specifies; no normalization or format check happens before submitting.
- **Detection procedure**: 1. Read the answer format section of the task and note the exact required literals for every non-numeric field (case, spelling, allowed set). 2. In the script, find the line that builds the final answer string and trace what object is inserted for each of those fields. 3. Simulate the interpolation mentally (e.g. Python `bool` → `"True"`, numpy `bool_` → `"True"`) and compare character-for-character against the required literal. 4. Check whether any explicit mapping/lowercasing/`str()` normalization or a final self-check against the spec exists; if not, flag.
- **Discriminator**: A real violation is a mismatch in the emitted characters (case or token) against a spec that names the literals; it is fine if the task does not fix the spelling, or if the script explicitly maps the value (e.g. `"true" if sig else "false"`) or lowercases before output. Numeric fields whose rounding matches the spec are not affected by this check.
- **Consequence**: The grader marks that field WRONG/MISSING even though the underlying computation is correct, producing a partial-credit failure (e.g. 2/3 checks passed) and an overall incorrect verdict.
265Entities plotted/analyzed do not match the entities named in the prompttaskda-code
Applies when
task -- the request names specific input data, grouping entities, measures, and output artifacts, and the agent must produce a chart/table from them.
Pattern
The attempt analyzes a different data source and different variables than those requested (different rows, different categorical breakdown, different metric), yet reports success because a chart of the requested visual type was produced and saved under the requested filename.
Detection procedure
  1. From the task statement, list the required entities: source dataset/columns, the grouping/selection rule (e.g., top-N by some measure), the quantity being aggregated, the config file to obey, and every required output file.
  2. Read the scripts (or, if none exist, the answer's description of processing steps) and extract the same list actually used.
  3. Compare item by item; also confirm every named output artifact is actually written, not just the one mentioned in the summary.
  4. Check that the reported titles/labels/axes derive from the config file and the requested variables rather than from unrelated content.
Discriminator
A real violation is a mismatch in the substance — different table, different grouping variable, different aggregated measure, or missing required outputs. A look-alike that is fine is cosmetic divergence (color choices, bar ordering within an allowed tie, phrasing of a label) while the underlying data, grouping, and aggregation match the request.
Consequence
The saved figure and any numeric dumps encode the wrong values and shapes, so all file-level checks (plot metadata, image, array) fail even though the agent declares the task complete.
id 438ab1e0b274 · mined from da-code dacode-plot-scatter-002@s5
raw text (what the judge reads)
### Entities plotted/analyzed do not match the entities named in the prompt
- **Applies when**: `task` -- the request names specific input data, grouping entities, measures, and output artifacts, and the agent must produce a chart/table from them.
- **Pattern**: The attempt analyzes a different data source and different variables than those requested (different rows, different categorical breakdown, different metric), yet reports success because *a* chart of the requested visual type was produced and saved under the requested filename.
- **Detection procedure**:
  1. From the task statement, list the required entities: source dataset/columns, the grouping/selection rule (e.g., top-N by some measure), the quantity being aggregated, the config file to obey, and every required output file.
  2. Read the scripts (or, if none exist, the answer's description of processing steps) and extract the same list actually used.
  3. Compare item by item; also confirm every named output artifact is actually written, not just the one mentioned in the summary.
  4. Check that the reported titles/labels/axes derive from the config file and the requested variables rather than from unrelated content.
- **Discriminator**: A real violation is a mismatch in the substance — different table, different grouping variable, different aggregated measure, or missing required outputs. A look-alike that is fine is cosmetic divergence (color choices, bar ordering within an allowed tie, phrasing of a label) while the underlying data, grouping, and aggregation match the request.
- **Consequence**: The saved figure and any numeric dumps encode the wrong values and shapes, so all file-level checks (plot metadata, image, array) fail even though the agent declares the task complete.
266Units/scale mismatch in reported summary statisticstaskinfiagent-dabench
Applies when
task -- the task asks for descriptive statistics (mean, std, etc.) of a quantity derived from raw data, and the script may transform, rescale, normalize, or re-unit that quantity before computing them.
Pattern
The agent computes the statistic on a converted or standardized version of the variable (e.g., seconds→hours/days, per-row normalization, z-scores, log or ratio transforms, or a per-group average) instead of on the raw values in the units implied by the data, and reports numbers that are orders of magnitude away from what the raw column supports; no back-check against the pre-filtering statistics is performed. Often accompanied by no saved script, making the derivation unverifiable.
Detection procedure
  1. Read the task and fix the expected unit/scale of the requested statistic from the raw data definition (what one row/value means, in what unit).
  2. In the script, trace the variable actually passed to the statistic function and list every transformation applied upstream (unit division, scaling, standardization, aggregation, resampling).
  3. Compute or locate the statistic on the untransformed values before outlier removal and compare orders of magnitude with the reported answer; also check that std/mean ratio and the filtered-out count are consistent with the stated rule.
  4. Confirm the answer's magnitude is inside the plausible range of the raw values (min/max), not a normalized ~0–1 or unit-converted number.
Discriminator
A real violation is when the reported numbers cannot lie in the range of the raw variable, or differ from the pre-removal statistic by a factor unrelated to removing a few tail points. It is not a violation if the transformation was explicitly requested by the task, or if the change from the pre-removal statistic is modest and explained solely by the removed points.
Consequence
Both reported values fail exact/tolerance comparison against ground truth (off by a constant factor or scale), so the grader marks every numeric check wrong even though the outlier-removal logic may have been right.
id fbd5ca129df1 · mined from infiagent-dabench dabench-619@s5
raw text (what the judge reads)
### Units/scale mismatch in reported summary statistics
- **Applies when**: `task` -- the task asks for descriptive statistics (mean, std, etc.) of a quantity derived from raw data, and the script may transform, rescale, normalize, or re-unit that quantity before computing them.
- **Pattern**: The agent computes the statistic on a converted or standardized version of the variable (e.g., seconds→hours/days, per-row normalization, z-scores, log or ratio transforms, or a per-group average) instead of on the raw values in the units implied by the data, and reports numbers that are orders of magnitude away from what the raw column supports; no back-check against the pre-filtering statistics is performed. Often accompanied by no saved script, making the derivation unverifiable.
- **Detection procedure**:
  1. Read the task and fix the expected unit/scale of the requested statistic from the raw data definition (what one row/value means, in what unit).
  2. In the script, trace the variable actually passed to the statistic function and list every transformation applied upstream (unit division, scaling, standardization, aggregation, resampling).
  3. Compute or locate the statistic on the untransformed values *before* outlier removal and compare orders of magnitude with the reported answer; also check that std/mean ratio and the filtered-out count are consistent with the stated rule.
  4. Confirm the answer's magnitude is inside the plausible range of the raw values (min/max), not a normalized ~0–1 or unit-converted number.
- **Discriminator**: A real violation is when the reported numbers cannot lie in the range of the raw variable, or differ from the pre-removal statistic by a factor unrelated to removing a few tail points. It is *not* a violation if the transformation was explicitly requested by the task, or if the change from the pre-removal statistic is modest and explained solely by the removed points.
- **Consequence**: Both reported values fail exact/tolerance comparison against ground truth (off by a constant factor or scale), so the grader marks every numeric check wrong even though the outlier-removal logic may have been right.
267Test run over auto-discovered groups instead of the exact categories named in the tasktaskda-code
Applies when
task -- the task explicitly enumerates the groups/levels (e.g., three named conditions) that a statistical test or comparison must be run over, and the script derives those groups programmatically from the data.
Pattern
The script calls unique()/groupby() on the grouping column and feeds every discovered level into the test, without checking that the discovered levels match the enumerated list in count, identity, and coding (numeric codes vs. labels, rare extra categories, merged categories). Extra or unmerged levels silently change the test statistic and p-value, and the reported result is never sanity-checked against the stated number of groups.
Detection procedure
  1. Read the task and note the exact number and names of the groups the test must compare, plus any other stated constraints (subset order, filtering scope, output file/format).
  2. In the script, find where groups are constructed; check for any assertion or filter/mapping that forces the group set to the enumerated categories, and check that group ordering/labels are reconciled with the raw encoding.
  3. Inspect the printed/reported diagnostics: does the script report per-group N and the number of groups per subset, so a mismatch (e.g., 4 groups where 3 were specified, or a group with a handful of rows) would be visible?
  4. Check the answer: does it come with evidence that exactly the specified groups were compared, and is it written to the artifact name/format the task requires?
Discriminator
A real violation is when the data's grouping column can contain levels beyond (or encoded differently from) the enumerated set and the script has no filter/merge/assertion — even if the numbers "look plausible". It is fine if the script explicitly validates set(levels) == specified set (or maps/collapses to it) and fails loudly otherwise.
Consequence
The test is computed on the wrong partition, yielding p-values (and possibly conclusions) that differ from the reference, so the expected result file is marked wrong/missing and the check fails.
id dec1716e42d5 · mined from da-code dacode-data-sa-061@s5
raw text (what the judge reads)
### Test run over auto-discovered groups instead of the exact categories named in the task
- **Applies when**: `task` -- the task explicitly enumerates the groups/levels (e.g., three named conditions) that a statistical test or comparison must be run over, and the script derives those groups programmatically from the data.
- **Pattern**: The script calls `unique()`/`groupby()` on the grouping column and feeds every discovered level into the test, without checking that the discovered levels match the enumerated list in count, identity, and coding (numeric codes vs. labels, rare extra categories, merged categories). Extra or unmerged levels silently change the test statistic and p-value, and the reported result is never sanity-checked against the stated number of groups.
- **Detection procedure**:
  1. Read the task and note the exact number and names of the groups the test must compare, plus any other stated constraints (subset order, filtering scope, output file/format).
  2. In the script, find where groups are constructed; check for any assertion or filter/mapping that forces the group set to the enumerated categories, and check that group ordering/labels are reconciled with the raw encoding.
  3. Inspect the printed/reported diagnostics: does the script report per-group N and the number of groups per subset, so a mismatch (e.g., 4 groups where 3 were specified, or a group with a handful of rows) would be visible?
  4. Check the answer: does it come with evidence that exactly the specified groups were compared, and is it written to the artifact name/format the task requires?
- **Discriminator**: A real violation is when the data's grouping column can contain levels beyond (or encoded differently from) the enumerated set and the script has no filter/merge/assertion — even if the numbers "look plausible". It is fine if the script explicitly validates `set(levels) == specified set` (or maps/collapses to it) and fails loudly otherwise.
- **Consequence**: The test is computed on the wrong partition, yielding p-values (and possibly conclusions) that differ from the reference, so the expected result file is marked wrong/missing and the check fails.
268Model shipped without held-out validation against a naive baseline (and with high-signal features silently dropped)taskda-code
Applies when
task -- the task asks for predictions on a test set and the script fits a model, but the answer reports only training-side descriptives (feature importances, prediction range, mean) instead of an out-of-sample error estimate.
Pattern
The agent keeps only the conveniently numeric columns, drops categorical/text/date/identifier columns that carry most of the signal, never splits off a validation fold (or cross-validates), and justifies the result with self-consistent statistics such as "mean prediction matches training mean" — a statement that is also true of a constant predictor and therefore proves nothing about accuracy.
Detection procedure
  1. Read the task/target definition and list which columns in the training file plausibly carry signal (categorical, temporal, text, grouping keys), then check the script's feature list for how many were discarded and whether any encoding/imputation was attempted.
  2. Search the script for a train/validation split, cross-validation, or any scored metric (RMSE/MAE/R²/accuracy) computed on data not used for fitting; check that a trivial baseline (mean/median/majority, or a simple group-wise average) was scored for comparison.
  3. Inspect the reported prediction distribution versus the training target distribution: if the predicted spread is far narrower than the target spread, the model is near-constant and is not being caught by any check.
  4. Confirm the answer's evidence is out-of-sample performance, not intermediate artifacts (importances, ranges, means) that are compatible with a useless model.
Discriminator
A real violation is no out-of-sample score at all, or a score never compared to a baseline, combined with unexplained dropping of informative columns. It is fine if the agent validated and consciously chose a numeric-only model because it measurably beat the baseline and the richer feature sets, or if the discarded columns were genuinely unusable (constant, entirely missing, or absent from the test file).
Consequence
The submitted predictions are effectively a noisy constant around the target mean; the grader's error/correlation threshold against true test values fails even though the file has the right shape and column name, and the answer contains no number that would have revealed the problem.
id 04132b17ed0c · mined from da-code dacode-ml-regression-004@s5
raw text (what the judge reads)
### Model shipped without held-out validation against a naive baseline (and with high-signal features silently dropped)
- **Applies when**: `task` -- the task asks for predictions on a test set and the script fits a model, but the answer reports only training-side descriptives (feature importances, prediction range, mean) instead of an out-of-sample error estimate.
- **Pattern**: The agent keeps only the conveniently numeric columns, drops categorical/text/date/identifier columns that carry most of the signal, never splits off a validation fold (or cross-validates), and justifies the result with self-consistent statistics such as "mean prediction matches training mean" — a statement that is also true of a constant predictor and therefore proves nothing about accuracy.
- **Detection procedure**:
  1. Read the task/target definition and list which columns in the training file plausibly carry signal (categorical, temporal, text, grouping keys), then check the script's feature list for how many were discarded and whether any encoding/imputation was attempted.
  2. Search the script for a train/validation split, cross-validation, or any scored metric (RMSE/MAE/R²/accuracy) computed on data not used for fitting; check that a trivial baseline (mean/median/majority, or a simple group-wise average) was scored for comparison.
  3. Inspect the reported prediction distribution versus the training target distribution: if the predicted spread is far narrower than the target spread, the model is near-constant and is not being caught by any check.
  4. Confirm the answer's evidence is out-of-sample performance, not intermediate artifacts (importances, ranges, means) that are compatible with a useless model.
- **Discriminator**: A real violation is *no* out-of-sample score at all, or a score never compared to a baseline, combined with unexplained dropping of informative columns. It is fine if the agent validated and consciously chose a numeric-only model because it measurably beat the baseline and the richer feature sets, or if the discarded columns were genuinely unusable (constant, entirely missing, or absent from the test file).
- **Consequence**: The submitted predictions are effectively a noisy constant around the target mean; the grader's error/correlation threshold against true test values fails even though the file has the right shape and column name, and the answer contains no number that would have revealed the problem.
269Unvalidated automatic model-selection producing a degenerate groupingtaskda-code
Applies when
task -- the task asks the agent to choose a structural hyperparameter itself (e.g., number of clusters/components/bins) and the script picks it by maximizing a single internal score over a range, then writes the labels straight to the deliverable.
Pattern
The attempt trusts one internal criterion's argmax without any sanity check on the resulting partition, so it ships a solution with near-empty or singleton groups (and/or a very low absolute score), instead of cross-checking with a second criterion (elbow/gap/stability), outlier handling or skew-correcting transforms, and the conventional/domain-plausible value; the exported feature columns may also be a different representation than the one actually fitted.
Detection procedure
  1. Read the task: note that the "appropriate number of groups" is a graded part of the answer, not just the file format.
  2. Read the script: check whether selection rests on one metric's argmax with no tie-breaking, no stability/robustness check, no inspection of group sizes, and no treatment of heavy-tailed/outlier features before fitting.
  3. Read the reported output: look at the per-group counts and the absolute score value — groups of size 1–3 out of hundreds, or a silhouette far below ~0.4, indicate the "optimum" is fitting outliers rather than structure.
  4. Check that the features written to the deliverable are the same matrix (same order, same scaling convention) that produced the labels.
Discriminator
A real violation is an unexamined argmax whose partition is degenerate/unstable or whose score is barely above noise; it is fine if the script compares several criteria (or several seeds/subsamples), reports balanced non-trivial groups, and justifies the chosen value even when it is not the raw argmax.
Consequence
The saved label column encodes a different number and shape of groups than the reference partition, so the file comparison for the clustering deliverable fails (0/1 checks) even though the CSV schema and row count look correct.
id 1416f33312f2 · mined from da-code dacode-ml-cluster-013@s5
raw text (what the judge reads)
### Unvalidated automatic model-selection producing a degenerate grouping
- **Applies when**: `task` -- the task asks the agent to choose a structural hyperparameter itself (e.g., number of clusters/components/bins) and the script picks it by maximizing a single internal score over a range, then writes the labels straight to the deliverable.
- **Pattern**: The attempt trusts one internal criterion's argmax without any sanity check on the resulting partition, so it ships a solution with near-empty or singleton groups (and/or a very low absolute score), instead of cross-checking with a second criterion (elbow/gap/stability), outlier handling or skew-correcting transforms, and the conventional/domain-plausible value; the exported feature columns may also be a different representation than the one actually fitted.
- **Detection procedure**:
  1. Read the task: note that the "appropriate number of groups" is a graded part of the answer, not just the file format.
  2. Read the script: check whether selection rests on one metric's argmax with no tie-breaking, no stability/robustness check, no inspection of group sizes, and no treatment of heavy-tailed/outlier features before fitting.
  3. Read the reported output: look at the per-group counts and the absolute score value — groups of size 1–3 out of hundreds, or a silhouette far below ~0.4, indicate the "optimum" is fitting outliers rather than structure.
  4. Check that the features written to the deliverable are the same matrix (same order, same scaling convention) that produced the labels.
- **Discriminator**: A real violation is an unexamined argmax whose partition is degenerate/unstable or whose score is barely above noise; it is fine if the script compares several criteria (or several seeds/subsamples), reports balanced non-trivial groups, and justifies the chosen value even when it is not the raw argmax.
- **Consequence**: The saved label column encodes a different number and shape of groups than the reference partition, so the file comparison for the clustering deliverable fails (0/1 checks) even though the CSV schema and row count look correct.
270Statistic computed over a different slice/axis than the task specifies (and on an unverified subset of the available data)taskinfiagent-dabench
Applies when
task -- the task names a specific slice (a year, group, split, or filter) and a per-entity statistic, and the script must decide which values enter each entity's computation and which source files/rows constitute "the dataset".
Pattern
The script silently substitutes a convenient axis or scope for the specified one — e.g. aggregating each entity across all periods instead of the named period, or loading only one of several available data files — and never states or tests the assumption; the stated constraint word (the year/filter) appears only as a discarded intermediate print, while the reported ranking comes from a different computation.
Detection procedure
  1. From the task statement, write down explicitly the unit of observation the statistic is supposed to summarize (what varies inside each entity's sample) and the filter that defines the slice.
  2. In the script, find the array actually passed to the statistic function and trace its construction: which rows, which columns, which files. Check that the filter named in the task actually restricts that array.
  3. Check whether the script enumerated the available data sources/rows (e.g. listed the input directory, verified entity counts) before concluding, or just assumed one file/one orientation is "the dataset".
  4. Check the final answer's provenance: is the reported name selected from the ranking built on the task-specified slice, or from a differently-scoped ranking?
Discriminator
A real violation is when the task's filter is computable from the data but is not applied to the values fed into the statistic, or when part of the data that plausibly belongs to "all entities" was never loaded. It is not a violation when the filter is genuinely degenerate for the requested statistic and the script explicitly documents that ambiguity, tests both interpretations, and shows they agree or justifies the chosen one.
Consequence
The reported entity is the argmax of a different quantity than the one requested, so the graded key does not match the expected value (0/1), even though the code runs cleanly and prints plausible numbers.
id 87bf1d7f9295 · mined from infiagent-dabench dabench-252@s5
raw text (what the judge reads)
### Statistic computed over a different slice/axis than the task specifies (and on an unverified subset of the available data)

- **Applies when**: `task` -- the task names a specific slice (a year, group, split, or filter) and a per-entity statistic, and the script must decide which values enter each entity's computation and which source files/rows constitute "the dataset".
- **Pattern**: The script silently substitutes a convenient axis or scope for the specified one — e.g. aggregating each entity across *all* periods instead of the named period, or loading only one of several available data files — and never states or tests the assumption; the stated constraint word (the year/filter) appears only as a discarded intermediate print, while the reported ranking comes from a different computation.
- **Detection procedure**:
  1. From the task statement, write down explicitly the unit of observation the statistic is supposed to summarize (what varies inside each entity's sample) and the filter that defines the slice.
  2. In the script, find the array actually passed to the statistic function and trace its construction: which rows, which columns, which files. Check that the filter named in the task actually restricts that array.
  3. Check whether the script enumerated the available data sources/rows (e.g. listed the input directory, verified entity counts) before concluding, or just assumed one file/one orientation is "the dataset".
  4. Check the final answer's provenance: is the reported name selected from the ranking built on the task-specified slice, or from a differently-scoped ranking?
- **Discriminator**: A real violation is when the task's filter is computable from the data but is not applied to the values fed into the statistic, or when part of the data that plausibly belongs to "all entities" was never loaded. It is *not* a violation when the filter is genuinely degenerate for the requested statistic and the script explicitly documents that ambiguity, tests both interpretations, and shows they agree or justifies the chosen one.
- **Consequence**: The reported entity is the argmax of a different quantity than the one requested, so the graded key does not match the expected value (0/1), even though the code runs cleanly and prints plausible numbers.
271Truncating an identified key value to a coarser granularity than the analysis producedtaskinfiagent-dabench
Applies when
task -- the task asks you to locate a specific record/label (a date, ID, category, row key) and the answer template's example format looks coarser or less precise than the value your analysis actually resolves.
Pattern
The script correctly computes the full-resolution identifier, but the reported answer applies a formatting/rounding/truncation step (e.g., dropping components of a timestamp, collapsing to a parent group, truncating an ID) that discards the information the task was actually asking for, so the graded field no longer uniquely identifies the located record. Any downstream number computed at full resolution is left inconsistent with the coarse label reported.
Detection procedure
  1. From the task, list what must be identified and check whether the subsequent computation (previous-row lookup, group filter, join) depends on the full-resolution value.
  2. In the scripts, find the line that formats the identifier for output and compare its precision to the precision used internally for the computation.
  3. If the output precision is lower, verify that the coarse value still uniquely identifies the record; if it does not (e.g., many records share it), treat the formatting step as information loss and require the full-resolution value (or at minimum both, with the full value in the graded field).
Discriminator
Not a violation when the coarse form is genuinely unique/unambiguous for the located record, or when the task explicitly defines the unit of analysis at that coarser level (e.g., the maximum is taken over aggregated groups). It is a violation when the internal computation used a finer key than the one reported, i.e. the reported label cannot be used to reproduce the reported number.
Consequence
The identifier check fails against the expected full-resolution value while the derived numeric check may still pass, yielding a partially-correct, overall-incorrect grade.
id 093411917ca0 · mined from infiagent-dabench dabench-572@s5
raw text (what the judge reads)
### Truncating an identified key value to a coarser granularity than the analysis produced
- **Applies when**: `task` -- the task asks you to *locate* a specific record/label (a date, ID, category, row key) and the answer template's example format looks coarser or less precise than the value your analysis actually resolves.
- **Pattern**: The script correctly computes the full-resolution identifier, but the reported answer applies a formatting/rounding/truncation step (e.g., dropping components of a timestamp, collapsing to a parent group, truncating an ID) that discards the information the task was actually asking for, so the graded field no longer uniquely identifies the located record. Any downstream number computed at full resolution is left inconsistent with the coarse label reported.
- **Detection procedure**:
  1. From the task, list what must be *identified* and check whether the subsequent computation (previous-row lookup, group filter, join) depends on the full-resolution value.
  2. In the scripts, find the line that formats the identifier for output and compare its precision to the precision used internally for the computation.
  3. If the output precision is lower, verify that the coarse value still uniquely identifies the record; if it does not (e.g., many records share it), treat the formatting step as information loss and require the full-resolution value (or at minimum both, with the full value in the graded field).
- **Discriminator**: Not a violation when the coarse form is genuinely unique/unambiguous for the located record, or when the task explicitly defines the unit of analysis at that coarser level (e.g., the maximum is taken over aggregated groups). It *is* a violation when the internal computation used a finer key than the one reported, i.e. the reported label cannot be used to reproduce the reported number.
- **Consequence**: The identifier check fails against the expected full-resolution value while the derived numeric check may still pass, yielding a partially-correct, overall-incorrect grade.
272Fabricating input data when the real dataset isn't foundtaskda-code
Applies when
task -- the task requires computing a statistic from a specific provided dataset, and the scripts include fallback logic that generates or simulates data.
Pattern
The agent fails to locate/inspect the actual input files, then synthesizes random or hand-constructed data (often with a fixed seed and invented column names) and reports a statistic computed on that fabricated data as if it were the real answer, without flagging the substitution.
Detection procedure
1. Read the task to confirm the answer must derive from a real supplied dataset. 2. Scan the scripts for np.random, linspace/hard-coded arrays, functions named "synthetic"/"sample"/"fallback", or a path list ending in a generated file. 3. Check whether the final reported number's code path can reach the synthetic branch — i.e., whether the script ever verified that a genuine data file was loaded (printed shape, columns, date range matching the task's stated period). 4. Confirm the answer is a single number with no evidence tying it to real, inspected data.
Discriminator
A real violation is when the reported value could come from generated data or from a file the agent never confirmed matches the task's described variables/period; it is fine if synthetic data is used only for pipeline smoke-testing and the final run demonstrably loads and describes the actual dataset (columns, row counts, time span consistent with the task).
Consequence
The reported statistic is unrelated to the true data, so the value in the output file mismatches the expected result and the check fails outright.
id c6bdc9f814f5 · mined from da-code dacode-data-sa-043@s5
raw text (what the judge reads)
### Fabricating input data when the real dataset isn't found
- **Applies when**: `task` -- the task requires computing a statistic from a specific provided dataset, and the scripts include fallback logic that generates or simulates data.
- **Pattern**: The agent fails to locate/inspect the actual input files, then synthesizes random or hand-constructed data (often with a fixed seed and invented column names) and reports a statistic computed on that fabricated data as if it were the real answer, without flagging the substitution.
- **Detection procedure**: 1. Read the task to confirm the answer must derive from a real supplied dataset. 2. Scan the scripts for `np.random`, `linspace`/hard-coded arrays, functions named "synthetic"/"sample"/"fallback", or a path list ending in a generated file. 3. Check whether the final reported number's code path can reach the synthetic branch — i.e., whether the script ever verified that a genuine data file was loaded (printed shape, columns, date range matching the task's stated period). 4. Confirm the answer is a single number with no evidence tying it to real, inspected data.
- **Discriminator**: A real violation is when the reported value could come from generated data or from a file the agent never confirmed matches the task's described variables/period; it is fine if synthetic data is used only for pipeline smoke-testing and the final run demonstrably loads and describes the actual dataset (columns, row counts, time span consistent with the task).
- **Consequence**: The reported statistic is unrelated to the true data, so the value in the output file mismatches the expected result and the check fails outright.
273Identifier values wrapped in extra quoting/decoration inside the answer templatetaskinfiagent-dabench
Applies when
task -- the answer format specifies a bracketed list of string identifiers, and the script prints/formats those identifiers into the required template.
Pattern
The attempt computes the correct values but emits the list elements with added decoration — Python repr/str(list) output, quotation marks, brackets, or whitespace — instead of the bare comma-separated tokens shown in the requested format, so string matching on the identifier field fails even though the analysis is right.
Detection procedure
1. Read the required answer format literally and note exactly what characters separate and surround the list elements (commas only, no quotes, no spaces). 2. In the scripts, find where the final answer string is built; check whether it interpolates a Python list/repr of strings rather than ",".join(str(x) for x in items). 3. Compare the produced answer text character-by-character with the template, for both the identifier field and the numeric field. 4. Confirm every element is a raw token (no ', ", [] nesting, or padding) and ordering/pairing matches the value list.
Discriminator
A real violation is any extra quoting/spacing/structural character in the emitted field, even with correct underlying content; a look-alike that is fine is an identifier whose own text legitimately contains punctuation (e.g. parentheses or hyphens) that is part of the data value itself.
Consequence
The grader marks the identifier field WRONG/MISSING while the numeric field passes, yielding a partial (failing) score despite a correct computation.
id 265f778a421e · mined from infiagent-dabench dabench-219@s5
raw text (what the judge reads)
### Identifier values wrapped in extra quoting/decoration inside the answer template
- **Applies when**: `task` -- the answer format specifies a bracketed list of string identifiers, and the script prints/formats those identifiers into the required template.
- **Pattern**: The attempt computes the correct values but emits the list elements with added decoration — Python `repr`/`str(list)` output, quotation marks, brackets, or whitespace — instead of the bare comma-separated tokens shown in the requested format, so string matching on the identifier field fails even though the analysis is right.
- **Detection procedure**: 1. Read the required answer format literally and note exactly what characters separate and surround the list elements (commas only, no quotes, no spaces). 2. In the scripts, find where the final answer string is built; check whether it interpolates a Python list/`repr` of strings rather than `",".join(str(x) for x in items)`. 3. Compare the produced answer text character-by-character with the template, for both the identifier field and the numeric field. 4. Confirm every element is a raw token (no `'`, `"`, `[]` nesting, or padding) and ordering/pairing matches the value list.
- **Discriminator**: A real violation is any extra quoting/spacing/structural character in the emitted field, even with correct underlying content; a look-alike that is fine is an identifier whose own text legitimately contains punctuation (e.g. parentheses or hyphens) that is part of the data value itself.
- **Consequence**: The grader marks the identifier field WRONG/MISSING while the numeric field passes, yielding a partial (failing) score despite a correct computation.
274Deliverable artifact not validated against the required path, header, and row counttaskda-code
Applies when
task -- the task asks for predictions/results to be written to a named output file whose format is defined by a provided sample/template file.
Pattern
The attempt produces an output file but never checks it against the specification: it is written to a different directory than the expected one, or it deviates from the template in header name/spelling, number of columns, index column presence, row count (must equal the number of scoring rows), or label encoding (e.g., strings/probabilities instead of the integer class labels used in the sample). No reproducible script is retained showing how the file was generated, so the mismatch cannot be caught by inspection either.
Detection procedure
  1. From the task text, extract the exact required output filename, its intended location (default: the working/submission directory the task names), the required column label(s), and the row count implied by the evaluation input.
  2. In the scripts, find the write call and confirm the path string matches that location/filename exactly and that the frame written has only the specified column(s) and no index (index=False or equivalent).
  3. Load the provided sample/template and the produced file side by side and compare: header text, column count, dtype/value domain of entries, and number of data rows versus rows in the evaluation input.
  4. Confirm the row order corresponds 1:1 to the evaluation input order (no shuffling, sorting, or dropped rows from filtering/NaN removal).
Discriminator
A real violation is any concrete mismatch in path, header, row count, value domain, or row alignment (or the absence of any artifact/script that lets these be checked). A look-alike that is fine is a file that differs only in harmless ways the spec allows — e.g., extra trailing newline, quoting style, or int-vs-float rendering of the same labels — while path, header, count, and ordering all match.
Consequence
The grader looks for the specified file with the specified schema and finds it missing or unparseable/misaligned, so the check fails outright (score 0) regardless of how good the underlying model was.
id bb10fe1b4a57 · mined from da-code dacode-ml-binary-013@s5
raw text (what the judge reads)
### Deliverable artifact not validated against the required path, header, and row count
- **Applies when**: `task` -- the task asks for predictions/results to be written to a named output file whose format is defined by a provided sample/template file.
- **Pattern**: The attempt produces an output file but never checks it against the specification: it is written to a different directory than the expected one, or it deviates from the template in header name/spelling, number of columns, index column presence, row count (must equal the number of scoring rows), or label encoding (e.g., strings/probabilities instead of the integer class labels used in the sample). No reproducible script is retained showing how the file was generated, so the mismatch cannot be caught by inspection either.
- **Detection procedure**:
  1. From the task text, extract the exact required output filename, its intended location (default: the working/submission directory the task names), the required column label(s), and the row count implied by the evaluation input.
  2. In the scripts, find the write call and confirm the path string matches that location/filename exactly and that the frame written has only the specified column(s) and no index (`index=False` or equivalent).
  3. Load the provided sample/template and the produced file side by side and compare: header text, column count, dtype/value domain of entries, and number of data rows versus rows in the evaluation input.
  4. Confirm the row order corresponds 1:1 to the evaluation input order (no shuffling, sorting, or dropped rows from filtering/NaN removal).
- **Discriminator**: A real violation is any concrete mismatch in path, header, row count, value domain, or row alignment (or the absence of any artifact/script that lets these be checked). A look-alike that is fine is a file that differs only in harmless ways the spec allows — e.g., extra trailing newline, quoting style, or int-vs-float rendering of the same labels — while path, header, count, and ordering all match.
- **Consequence**: The grader looks for the specified file with the specified schema and finds it missing or unparseable/misaligned, so the check fails outright (score 0) regardless of how good the underlying model was.
275Unverifiable result reported in a decorated format instead of the exact requested answer templatetaskinfiagent-dabench
Applies when
task -- the task prescribes a literal answer template (e.g. @key[...] with a specified element type/separator) and expects the answer to be produced by saved, re-runnable analysis code.
Pattern
The attempt reports a hand-written answer string that adds or changes decoration relative to the template (quoting of elements, brackets, JSON-style formatting, ordering, extra whitespace/units) and leaves behind no script that computes and prints the final string, so the value and its formatting cannot be reproduced or checked.
Detection procedure
  1. Extract from the task statement the exact answer template: key name, delimiters, element type (bare names vs quoted strings), separator, and any ordering/rounding rules.
  2. Check that a saved script exists which computes the statistic and prints the final answer string itself (not just intermediate tables); if no such artifact exists, the attempt is unverifiable and inadequate regardless of whether the values look right.
  3. Compare the submitted string character-by-character against the template: flag added quotes, changed brackets, altered separators, appended commentary, or element ordering not matching the stated rule.
  4. Confirm the printed values are the requested quantity (e.g. the identifying labels) rather than an adjacent intermediate (indices, values, counts).
Discriminator
A real violation is a mismatch in the literal surface form or a missing reproducible producer of that string; a look-alike that is fine is a script that programmatically formats the answer exactly as specified and whose output is pasted verbatim, even if internal variables use different types or ordering.
Consequence
The grader parses the submission against the expected key/format and reports the check as WRONG/MISSING even when the underlying analysis identified the correct items, scoring 0.
id 8740e2463150 · mined from infiagent-dabench dabench-254@s5
raw text (what the judge reads)
### Unverifiable result reported in a decorated format instead of the exact requested answer template
- **Applies when**: `task` -- the task prescribes a literal answer template (e.g. `@key[...]` with a specified element type/separator) and expects the answer to be produced by saved, re-runnable analysis code.
- **Pattern**: The attempt reports a hand-written answer string that adds or changes decoration relative to the template (quoting of elements, brackets, JSON-style formatting, ordering, extra whitespace/units) and leaves behind no script that computes and prints the final string, so the value and its formatting cannot be reproduced or checked.
- **Detection procedure**:
  1. Extract from the task statement the exact answer template: key name, delimiters, element type (bare names vs quoted strings), separator, and any ordering/rounding rules.
  2. Check that a saved script exists which computes the statistic and *prints the final answer string itself* (not just intermediate tables); if no such artifact exists, the attempt is unverifiable and inadequate regardless of whether the values look right.
  3. Compare the submitted string character-by-character against the template: flag added quotes, changed brackets, altered separators, appended commentary, or element ordering not matching the stated rule.
  4. Confirm the printed values are the requested quantity (e.g. the identifying labels) rather than an adjacent intermediate (indices, values, counts).
- **Discriminator**: A real violation is a mismatch in the literal surface form or a missing reproducible producer of that string; a look-alike that is fine is a script that programmatically formats the answer exactly as specified and whose output is pasted verbatim, even if internal variables use different types or ordering.
- **Consequence**: The grader parses the submission against the expected key/format and reports the check as WRONG/MISSING even when the underlying analysis identified the correct items, scoring 0.
276Output file schema not matching the exact requested columns/valuestaskda-code
Applies when
task -- The task specifies an output file with a named column (or set of columns) holding predictions/results, and the scripts write that file.
Pattern
The attempt writes the file with extra columns (e.g., an index/ID column), a different column name/case, or label values in a different encoding/spelling than the source data's target values, and then declares success based only on row count and absence of nulls — never checking the file against the literal specification.
Detection procedure
  1. From the task statement, extract the exact required file name, the exact column name(s), and (for classification) the exact label vocabulary as it appears in the training target.
  2. Read the write step in the scripts: check whether the frame written contains only the required column(s), whether index=False (or equivalent) is used so no extra index column appears, and whether the column header string matches exactly (spelling, case, spacing).
  3. Check that predicted values are inverse-transformed back to the original label strings/dtype rather than left as 0/1, encoded integers, or reworded categories.
  4. Compare the answer's own description of the file (columns listed, value names, dtypes) against step 1; any mismatch, including "contains 2 columns: id, <target>", is a violation.
Discriminator
A real violation is a structural mismatch with the stated spec (extra/renamed columns, altered or re-encoded label values). It is fine if the file has exactly the requested column(s) with original-format values, even if the internal modeling used encoded labels, and it is fine to include an ID column only when the task explicitly asks for one.
Consequence
The grader reads the expected column/values and fails to match, marking the result file WRONG/MISSING and scoring 0 regardless of how accurate the underlying model is.
id 5551b76d2916 · mined from da-code dacode-ml-binary-009@s5
raw text (what the judge reads)
### Output file schema not matching the exact requested columns/values
- **Applies when**: `task` -- The task specifies an output file with a named column (or set of columns) holding predictions/results, and the scripts write that file.
- **Pattern**: The attempt writes the file with extra columns (e.g., an index/ID column), a different column name/case, or label values in a different encoding/spelling than the source data's target values, and then declares success based only on row count and absence of nulls — never checking the file against the literal specification.
- **Detection procedure**:
  1. From the task statement, extract the exact required file name, the exact column name(s), and (for classification) the exact label vocabulary as it appears in the training target.
  2. Read the write step in the scripts: check whether the frame written contains only the required column(s), whether `index=False` (or equivalent) is used so no extra index column appears, and whether the column header string matches exactly (spelling, case, spacing).
  3. Check that predicted values are inverse-transformed back to the original label strings/dtype rather than left as 0/1, encoded integers, or reworded categories.
  4. Compare the answer's own description of the file (columns listed, value names, dtypes) against step 1; any mismatch, including "contains 2 columns: id, <target>", is a violation.
- **Discriminator**: A real violation is a structural mismatch with the stated spec (extra/renamed columns, altered or re-encoded label values). It is fine if the file has exactly the requested column(s) with original-format values, even if the internal modeling used encoded labels, and it is fine to include an ID column only when the task explicitly asks for one.
- **Consequence**: The grader reads the expected column/values and fails to match, marking the result file WRONG/MISSING and scoring 0 regardless of how accurate the underlying model is.
277Fabricated metric definition applied to only one side of a symmetric record structuretaskda-code
Applies when
task -- the task asks for a "performance"/"score"/ranking chart or table whose exact formula is not spelled out in the prompt, and the data stores each observation with two (or more) role-based columns (e.g. two participant columns with their own value columns) plus a config/spec file that constrains the output.
Pattern
The script invents an arbitrary aggregation (e.g. counts plus an ad-hoc weighted term) rather than deriving the intended quantity from the spec/config or a standard domain definition, and computes it by grouping on only one of the role columns, so every record where an entity appears in the other role is silently dropped. Flag comparisons are also done against raw strings/booleans without checking dtype, and only some of the required output artifacts are produced.
Detection procedure
  1. Read the task and the referenced config/spec: list every requested output artifact and every field of the spec (labels, ordering, title, categories), and note which of these implicitly define the quantity to be plotted (e.g. an axis label naming the metric, a category list implying a selection rule).
  2. In the script, locate the line(s) defining the plotted quantity; check whether the formula is traceable to the spec/domain definition or was chosen by the agent, and whether the grouping/aggregation covers all role columns in which an entity can appear (search for a second groupby/melt/concat over the mirrored columns).
  3. Verify boolean/flag filters against the actual parsed dtype (x == 'TRUE' on a real boolean column silently yields zero), and check that the selected entity set and its ordering are derived, not copied from the spec's label list.
  4. Compare the artifacts written by the script with the full list from step 1; any missing file, or any reported value the agent never sanity-checked (plausible range, counts, monotonic ranking) is a failure.
Discriminator
A real violation is an unjustified formula, a one-sided grouping, a dtype-blind filter, or missing artifacts. It is not a violation if the formula is explicitly given in the task/spec, if the entity genuinely only occurs in one role column, or if the agent verified equivalence (e.g. showed the mirrored aggregation adds nothing).
Consequence
The plotted values, the entity ranking, and any derived numeric dump differ from the reference, so the image comparison and the numeric/JSON artifact checks all fail (and missing artifacts fail outright), even though the script runs without error.
id fcf555c00cb0 · mined from da-code dacode-plot-bar-006@s5
raw text (what the judge reads)
### Fabricated metric definition applied to only one side of a symmetric record structure
- **Applies when**: `task` -- the task asks for a "performance"/"score"/ranking chart or table whose exact formula is not spelled out in the prompt, and the data stores each observation with two (or more) role-based columns (e.g. two participant columns with their own value columns) plus a config/spec file that constrains the output.
- **Pattern**: The script invents an arbitrary aggregation (e.g. counts plus an ad-hoc weighted term) rather than deriving the intended quantity from the spec/config or a standard domain definition, and computes it by grouping on only one of the role columns, so every record where an entity appears in the other role is silently dropped. Flag comparisons are also done against raw strings/booleans without checking dtype, and only some of the required output artifacts are produced.
- **Detection procedure**:
  1. Read the task and the referenced config/spec: list every requested output artifact and every field of the spec (labels, ordering, title, categories), and note which of these implicitly define the quantity to be plotted (e.g. an axis label naming the metric, a category list implying a selection rule).
  2. In the script, locate the line(s) defining the plotted quantity; check whether the formula is traceable to the spec/domain definition or was chosen by the agent, and whether the grouping/aggregation covers all role columns in which an entity can appear (search for a second groupby/melt/concat over the mirrored columns).
  3. Verify boolean/flag filters against the actual parsed dtype (`x == 'TRUE'` on a real boolean column silently yields zero), and check that the selected entity set and its ordering are derived, not copied from the spec's label list.
  4. Compare the artifacts written by the script with the full list from step 1; any missing file, or any reported value the agent never sanity-checked (plausible range, counts, monotonic ranking) is a failure.
- **Discriminator**: A real violation is an unjustified formula, a one-sided grouping, a dtype-blind filter, or missing artifacts. It is *not* a violation if the formula is explicitly given in the task/spec, if the entity genuinely only occurs in one role column, or if the agent verified equivalence (e.g. showed the mirrored aggregation adds nothing).
- **Consequence**: The plotted values, the entity ranking, and any derived numeric dump differ from the reference, so the image comparison and the numeric/JSON artifact checks all fail (and missing artifacts fail outright), even though the script runs without error.
278Counting extreme values without validating the input scope or the statistic's definitiontaskinfiagent-dabench
Applies when
task -- the task asks for a count/removal of rows satisfying a threshold rule (z-score, IQR, quantile, etc.) on one column, and the script hard-codes a single data file and a single library call to compute the statistic.
Pattern
The agent picks the first plausible file in the data directory (e.g. a split or partial file) and one implementation of the statistic (library default ddof, NaN-dropped subset, unconverted dtype) without ever enumerating the available files or confirming that the resulting count is robust; it then reports whatever number that single path produces, even when the reported extremes look like ordinary repeated values rather than genuine tail points.
Detection procedure
  1. Read the task and note whether it names a specific file/subset; if it does not, check whether the script listed the data directory and justified its choice over other candidate files (full dataset vs. train/test split, raw vs. cleaned).
  2. Read the statistic computation: does it rely on a library default (population vs. sample standard deviation), on a NaN-dropped or otherwise filtered subset, or on a column whose dtype/units were never verified as numeric and clean?
  3. Check whether the script performs any independent cross-check of the count -- recomputing the bound explicitly (mean ± 3·std with both ddof=0 and ddof=1), and printing the min/max and the top few values against that bound.
  4. Compare the reported count with the printed extreme values: a large count of repeated, non-extreme values, or a count that flips between implementations, means the number was never validated.
Discriminator
A fine attempt may still use one file and one library call, but it explicitly confirms the file choice (directory listing or task wording), verifies column dtype/NaNs, and shows that the flagged values genuinely lie outside the explicitly computed bound under both variance conventions; a violation is an attempt whose count depends on unexamined defaults or an unjustified subset, with no reconciliation step.
Consequence
The reported count differs from the ground-truth count (often drastically, e.g. a large number where the correct answer is zero, or vice versa), so the single numeric check in the required answer format fails.
id b78195814adf · mined from infiagent-dabench dabench-361@s5
raw text (what the judge reads)
### Counting extreme values without validating the input scope or the statistic's definition
- **Applies when**: `task` -- the task asks for a count/removal of rows satisfying a threshold rule (z-score, IQR, quantile, etc.) on one column, and the script hard-codes a single data file and a single library call to compute the statistic.
- **Pattern**: The agent picks the first plausible file in the data directory (e.g. a split or partial file) and one implementation of the statistic (library default `ddof`, NaN-dropped subset, unconverted dtype) without ever enumerating the available files or confirming that the resulting count is robust; it then reports whatever number that single path produces, even when the reported extremes look like ordinary repeated values rather than genuine tail points.
- **Detection procedure**:
  1. Read the task and note whether it names a specific file/subset; if it does not, check whether the script listed the data directory and justified its choice over other candidate files (full dataset vs. train/test split, raw vs. cleaned).
  2. Read the statistic computation: does it rely on a library default (population vs. sample standard deviation), on a NaN-dropped or otherwise filtered subset, or on a column whose dtype/units were never verified as numeric and clean?
  3. Check whether the script performs any independent cross-check of the count -- recomputing the bound explicitly (mean ± 3·std with both `ddof=0` and `ddof=1`), and printing the min/max and the top few values against that bound.
  4. Compare the reported count with the printed extreme values: a large count of repeated, non-extreme values, or a count that flips between implementations, means the number was never validated.
- **Discriminator**: A fine attempt may still use one file and one library call, but it explicitly confirms the file choice (directory listing or task wording), verifies column dtype/NaNs, and shows that the flagged values genuinely lie outside the explicitly computed bound under both variance conventions; a violation is an attempt whose count depends on unexamined defaults or an unjustified subset, with no reconciliation step.
- **Consequence**: The reported count differs from the ground-truth count (often drastically, e.g. a large number where the correct answer is zero, or vice versa), so the single numeric check in the required answer format fails.
279Submission format assumed instead of read from the provided templatetaskda-code
Applies when
task -- The task supplies a sample/template output file (or an explicit format spec) and the scripts must write a prediction/result file matching it.
Pattern
The scripts build the output file from hand-guessed column names, id casing, column order or row order (e.g. derived from the test/input frame) and never load or compare against the provided template; the "validation" step only checks self-consistent properties (probabilities sum to 1, no NaNs, row count) rather than agreement with the template's header and ids.
Detection procedure
  1. In the task/README, note that a sample output file is provided and is the authoritative format spec.
  2. Search the scripts for a read of that template file; check whether the output DataFrame's column names, dtypes, id values and row order are taken from or asserted equal to it.
  3. If absent, compare the literal header strings and id column name written in the script against those stated/implied in the task description — any invented naming, casing, or extra/missing column is a violation.
  4. Check the answer's "validation" section: if it lists only internal sanity checks and never "columns match sample submission" / "ids match sample submission exactly", flag it.
Discriminator
A real violation is inventing or inferring the header/id scheme without ever touching the template; it is fine if the script loads the template (or explicitly asserts the exact required column list and id ordering) and writes into it, even if it then also runs internal sanity checks.
Consequence
The grader cannot parse or align the file (unknown/renamed columns, mismatched or misordered ids), so the submission is scored as wrong/missing regardless of model quality.
id cac547552aff · mined from da-code dacode-ml-competition-003@s5
raw text (what the judge reads)
### Submission format assumed instead of read from the provided template
- **Applies when**: `task` -- The task supplies a sample/template output file (or an explicit format spec) and the scripts must write a prediction/result file matching it.
- **Pattern**: The scripts build the output file from hand-guessed column names, id casing, column order or row order (e.g. derived from the test/input frame) and never load or compare against the provided template; the "validation" step only checks self-consistent properties (probabilities sum to 1, no NaNs, row count) rather than agreement with the template's header and ids.
- **Detection procedure**:
  1. In the task/README, note that a sample output file is provided and is the authoritative format spec.
  2. Search the scripts for a read of that template file; check whether the output DataFrame's column names, dtypes, id values and row order are taken from or asserted equal to it.
  3. If absent, compare the literal header strings and id column name written in the script against those stated/implied in the task description — any invented naming, casing, or extra/missing column is a violation.
  4. Check the answer's "validation" section: if it lists only internal sanity checks and never "columns match sample submission" / "ids match sample submission exactly", flag it.
- **Discriminator**: A real violation is inventing or inferring the header/id scheme without ever touching the template; it is fine if the script loads the template (or explicitly asserts the exact required column list and id ordering) and writes into it, even if it then also runs internal sanity checks.
- **Consequence**: The grader cannot parse or align the file (unknown/renamed columns, mismatched or misordered ids), so the submission is scored as wrong/missing regardless of model quality.
280Ignoring the provided train/test split and fabricating one's owntaskda-code
Applies when
task -- The task explicitly names the file(s) to predict on (e.g., a designated test set) and the required output file/column format.
Pattern
The attempt derives its own "test set" from the full data using an ad-hoc rule (e.g., rows missing some field, or a random split) instead of loading the specified prediction file, so the predictions' row count, ID set, and ordering do not correspond to the required submission; it may also drop or rename required columns.
Detection procedure
  1. From the task, note the exact input file to predict on, the expected number of prediction rows, the key/ID alignment, and the required output columns/names.
  2. In the scripts, check whether that file is actually read and used to build the feature matrix for prediction; flag any construction of the evaluation subset by filtering/splitting another file.
  3. Compare the reported prediction count and column list in the answer against the specified test file's row count and required schema; verify IDs are the test IDs in the test file's order.
  4. Check that features used at prediction time exist for the specified test rows (not only for the self-invented subset).
Discriminator
A real violation is when the predicted rows are not exactly the specified test rows (different count/IDs/order) or the output schema deviates; it is fine to carve an internal validation split out of the training data as long as final predictions are generated for every row of the given test file in its order with the required column(s).
Consequence
The submitted file cannot be joined/compared to the ground-truth labels — the grader reports the output file as wrong/missing (0 score) regardless of model quality.
id 86443b8b18fe · mined from da-code dacode-ml-multi-003@s5
raw text (what the judge reads)
### Ignoring the provided train/test split and fabricating one's own
- **Applies when**: `task` -- The task explicitly names the file(s) to predict on (e.g., a designated test set) and the required output file/column format.
- **Pattern**: The attempt derives its own "test set" from the full data using an ad-hoc rule (e.g., rows missing some field, or a random split) instead of loading the specified prediction file, so the predictions' row count, ID set, and ordering do not correspond to the required submission; it may also drop or rename required columns.
- **Detection procedure**:
  1. From the task, note the exact input file to predict on, the expected number of prediction rows, the key/ID alignment, and the required output columns/names.
  2. In the scripts, check whether that file is actually read and used to build the feature matrix for prediction; flag any construction of the evaluation subset by filtering/splitting another file.
  3. Compare the reported prediction count and column list in the answer against the specified test file's row count and required schema; verify IDs are the test IDs in the test file's order.
  4. Check that features used at prediction time exist for the specified test rows (not only for the self-invented subset).
- **Discriminator**: A real violation is when the predicted rows are not exactly the specified test rows (different count/IDs/order) or the output schema deviates; it is fine to carve an internal validation split out of the *training* data as long as final predictions are generated for every row of the given test file in its order with the required column(s).
- **Consequence**: The submitted file cannot be joined/compared to the ground-truth labels — the grader reports the output file as wrong/missing (0 score) regardless of model quality.
281Ignoring the provided template/reference output when producing a required filetaskda-code
Applies when
task -- the task supplies a template or an analogous already-computed result file and asks that the saved output "match the format"/structure of it.
Pattern
The script computes the aggregation with self-invented choices (extra row filtering, rounding, index/label formatting, fill values, column naming) and writes the file without ever reading the template, or it prints a side-by-side comparison against an analogous reference but never asserts equality nor reconciles the observed differences; the final answer describes the structure from the script's own output rather than from the template.
Detection procedure
  1. In the task text, note every referenced template/example/companion result file and the stated formatting constraints (rounding, ordering, index format).
  2. Search the scripts for a read of that template and an explicit structural check (shape, index values/dtype, column names/order) plus a numeric reconciliation on a comparable quantity; absence of any such check is a flag.
  3. Where the scripts do compare against an analogous reference (e.g., same cohort/group layout for a different measure), verify the agent reproduced that reference exactly before trusting its pipeline on the new measure; if the printed numbers differ and the agent hand-waves the difference ("expected to differ"), treat the methodology as uncalibrated.
  4. Check whether preprocessing/rounding decisions not stated in the task (dropping non-positive or missing rows, rounding to N decimals, fillna) were adopted without evidence from the template.
Discriminator
A real violation is when no template-derived validation exists, or a known mismatch with the reference is left unexplained; it is fine if the agent loads the template, matches its shape/labels/precision, and reproduces an analogous reference measure exactly before generating the requested one.
Consequence
The saved file differs from the expected one in values or layout (extra/missing rows from unjustified filtering, wrong rounding, wrong index/column labels), so the file-comparison check fails and the task scores 0.
id 5f98228ce30e · mined from da-code dacode-dm-csv-044@s5
raw text (what the judge reads)
### Ignoring the provided template/reference output when producing a required file
- **Applies when**: `task` -- the task supplies a template or an analogous already-computed result file and asks that the saved output "match the format"/structure of it.
- **Pattern**: The script computes the aggregation with self-invented choices (extra row filtering, rounding, index/label formatting, fill values, column naming) and writes the file without ever reading the template, or it prints a side-by-side comparison against an analogous reference but never asserts equality nor reconciles the observed differences; the final answer describes the structure from the script's own output rather than from the template.
- **Detection procedure**:
  1. In the task text, note every referenced template/example/companion result file and the stated formatting constraints (rounding, ordering, index format).
  2. Search the scripts for a read of that template and an explicit structural check (shape, index values/dtype, column names/order) plus a numeric reconciliation on a comparable quantity; absence of any such check is a flag.
  3. Where the scripts do compare against an analogous reference (e.g., same cohort/group layout for a different measure), verify the agent reproduced that reference exactly before trusting its pipeline on the new measure; if the printed numbers differ and the agent hand-waves the difference ("expected to differ"), treat the methodology as uncalibrated.
  4. Check whether preprocessing/rounding decisions not stated in the task (dropping non-positive or missing rows, rounding to N decimals, fillna) were adopted without evidence from the template.
- **Discriminator**: A real violation is when no template-derived validation exists, or a known mismatch with the reference is left unexplained; it is fine if the agent loads the template, matches its shape/labels/precision, and reproduces an analogous reference measure exactly before generating the requested one.
- **Consequence**: The saved file differs from the expected one in values or layout (extra/missing rows from unjustified filtering, wrong rounding, wrong index/column labels), so the file-comparison check fails and the task scores 0.
282Missing-value mask defined too narrowly, so group membership (and every group statistic) is offtaskinfiagent-dabench
Applies when
task -- the analysis splits rows into "missing" vs "non-missing" groups on some column (or otherwise filters/drops rows) before computing group statistics or a test.
Pattern
The script relies on a single default notion of "null" (e.g., isna() after a plain load) without checking how the file actually encodes absent values — empty strings, whitespace, sentinel tokens like NA/None/-/?/0, or values lost to dtype coercion/na_values defaults — and it may also silently drop rows with missing values in the measured column. The resulting two subsets are mis-partitioned, so both group means shift slightly while the answer still looks plausible.
Detection procedure
  1. Read the task to identify exactly which column defines the split and which column is aggregated; note that all rows should fall in one of the two groups.
  2. In the script, find how the file is loaded (delimiter, na_values, keep_default_na, dtype) and how the mask is built; check whether the agent ever inspected the raw distinct values / value counts of the split column to see what "missing" looks like in the data.
  3. Check that n_group1 + n_group2 == len(df) is asserted/printed, and that no dropna(), astype, or filtering step upstream removed rows from either group.
  4. Compare the reported group sizes/means against a quick independent recount under an alternative missingness definition (strings stripped, sentinels included); flag if the answer is sensitive to that choice and the agent never checked.
Discriminator
A real violation is an unverified mask — no evidence the agent looked at raw values or verified the partition covers all rows. It is not a violation if the agent inspected the column's encodings, documented that only true NaNs exist, and showed the two subset sizes summing to the full row count (even if the load is a plain read_csv).
Consequence
Both reported group means (and the test statistic) deviate from the reference values by a few percent — enough that the numeric checks fail even though the qualitative conclusion (significant difference, p≈0) looks right.
id 26b19f14ed56 · mined from infiagent-dabench dabench-297@s5
raw text (what the judge reads)
### Missing-value mask defined too narrowly, so group membership (and every group statistic) is off
- **Applies when**: `task` -- the analysis splits rows into "missing" vs "non-missing" groups on some column (or otherwise filters/drops rows) before computing group statistics or a test.
- **Pattern**: The script relies on a single default notion of "null" (e.g., `isna()` after a plain load) without checking how the file actually encodes absent values — empty strings, whitespace, sentinel tokens like `NA`/`None`/`-`/`?`/`0`, or values lost to dtype coercion/`na_values` defaults — and it may also silently drop rows with missing values in the *measured* column. The resulting two subsets are mis-partitioned, so both group means shift slightly while the answer still looks plausible.
- **Detection procedure**:
  1. Read the task to identify exactly which column defines the split and which column is aggregated; note that all rows should fall in one of the two groups.
  2. In the script, find how the file is loaded (delimiter, `na_values`, `keep_default_na`, dtype) and how the mask is built; check whether the agent ever inspected the raw distinct values / value counts of the split column to see what "missing" looks like in the data.
  3. Check that `n_group1 + n_group2 == len(df)` is asserted/printed, and that no `dropna()`, `astype`, or filtering step upstream removed rows from either group.
  4. Compare the reported group sizes/means against a quick independent recount under an alternative missingness definition (strings stripped, sentinels included); flag if the answer is sensitive to that choice and the agent never checked.
- **Discriminator**: A real violation is an unverified mask — no evidence the agent looked at raw values or verified the partition covers all rows. It is *not* a violation if the agent inspected the column's encodings, documented that only true NaNs exist, and showed the two subset sizes summing to the full row count (even if the load is a plain `read_csv`).
- **Consequence**: Both reported group means (and the test statistic) deviate from the reference values by a few percent — enough that the numeric checks fail even though the qualitative conclusion (significant difference, p≈0) looks right.
283Required output artifacts not fully enumerated and producedtaskda-code
Applies when
task -- the task (or a referenced guidance/spec file) asks for deliverables such as saved figures, serialized numeric results, or structured metadata files, and the agent's script writes files to disk.
Pattern
The agent produces only the most visible artifact (e.g., an image) and reports the remaining numbers in prose/stdout, silently skipping the other required serialized outputs, writing them to a non-specified path, or never re-reading the referenced spec to learn the full deliverable list and the exact definitions/categories it prescribes.
Detection procedure
  1. From the task statement and any referenced instruction file, list every named output artifact (filename, extension, location) and every prescribed definition (filters, category labels, ordering, colors, sizes).
  2. Grep the scripts for every write/save call (savefig, to_csv, np.save, json.dump, etc.) and record the exact paths produced.
  3. Diff the two lists: flag if any required artifact is absent, has a different name/extension/directory, or if any prescribed definition was replaced by the agent's own invented rule (e.g., self-derived category mapping or filter thresholds not traceable to the spec).
  4. Check the final answer: if it states results only as text while a machine-readable artifact was requested, flag it.
Discriminator
A real violation is a missing/misplaced required file or a definition the agent invented without any evidence it read the spec; it is not a violation if the artifact exists under the specified name and the extra prose is merely a summary, or if the spec genuinely leaves the categorization to the analyst and the agent documents a defensible rule.
Consequence
File-existence/content checks for the unproduced artifacts fail outright (0 checks passed), and even the produced figure can mismatch because it encodes self-invented categories and counts rather than the specified ones.
id b57c022c3af4 · mined from da-code dacode-plot-pie-005@s5
raw text (what the judge reads)
### Required output artifacts not fully enumerated and produced
- **Applies when**: `task` -- the task (or a referenced guidance/spec file) asks for deliverables such as saved figures, serialized numeric results, or structured metadata files, and the agent's script writes files to disk.
- **Pattern**: The agent produces only the most visible artifact (e.g., an image) and reports the remaining numbers in prose/stdout, silently skipping the other required serialized outputs, writing them to a non-specified path, or never re-reading the referenced spec to learn the full deliverable list and the exact definitions/categories it prescribes.
- **Detection procedure**:
  1. From the task statement and any referenced instruction file, list every named output artifact (filename, extension, location) and every prescribed definition (filters, category labels, ordering, colors, sizes).
  2. Grep the scripts for every write/save call (`savefig`, `to_csv`, `np.save`, `json.dump`, etc.) and record the exact paths produced.
  3. Diff the two lists: flag if any required artifact is absent, has a different name/extension/directory, or if any prescribed definition was replaced by the agent's own invented rule (e.g., self-derived category mapping or filter thresholds not traceable to the spec).
  4. Check the final answer: if it states results only as text while a machine-readable artifact was requested, flag it.
- **Discriminator**: A real violation is a missing/misplaced required file or a definition the agent invented without any evidence it read the spec; it is *not* a violation if the artifact exists under the specified name and the extra prose is merely a summary, or if the spec genuinely leaves the categorization to the analyst and the agent documents a defensible rule.
- **Consequence**: File-existence/content checks for the unproduced artifacts fail outright (0 checks passed), and even the produced figure can mismatch because it encodes self-invented categories and counts rather than the specified ones.
284Unverified regression error: prescribed preprocessing/split recipe not reproducibly implemented or sanity-checked against a variance baselinetaskinfiagent-dabench
Applies when
task -- the task fixes an exact modeling recipe (specific feature/target columns, a stated missing-value handling rule, a fixed train/test split fraction, and a named error metric) and asks for a single numeric error value.
Pattern
The attempt reports a metric value without a saved, re-runnable script that implements each stated step literally on the full dataset — e.g. it silently drops rows with missing values (or lets non-numeric/parsed-as-string columns force coercion) instead of mean-imputing exactly the named columns before splitting, subsets/filters rows the task never mentioned, swaps target and predictors, or evaluates on the training portion — and it never checks whether the reported error is plausible relative to the target's own variance.
Detection procedure
  1. From the task, list every mandated step: exact target and predictor columns, missing-value rule, split fraction, metric definition, rounding/format.
  2. Open the scripts and match each mandated step to a concrete line; flag if no script exists, if rows are dropped/filtered anywhere, if the named columns are not numeric-cast then imputed with their own means over the full dataset before splitting, or if the split fraction/metric differs.
  3. Check that the metric is computed on the held-out portion only, with predictions and truths from the same rows and the same target variable (not a transformed/scaled version).
  4. Compare the reported error to the variance (or std²) of the target computed under the same imputation; an ordinary-least-squares fit should not do much worse than predicting the mean, so an error far above the target variance (or a residual scale implausible given the target's range) must be explained.
Discriminator
A genuine violation is an unexplained deviation from the recipe or an error magnitude that exceeds the mean-only baseline; a look-alike that is fine is a large-but-consistent error simply because the target has a large scale/variance and the recipe is implemented exactly as stated (split seed differences alone give only modest variation, not order-of-magnitude ones).
Consequence
The reported error differs from the reference value (often by an order of magnitude) and the answer is graded wrong, with no script available to localize the deviating step.
id bc555904b943 · mined from infiagent-dabench dabench-432@s5
raw text (what the judge reads)
### Unverified regression error: prescribed preprocessing/split recipe not reproducibly implemented or sanity-checked against a variance baseline
- **Applies when**: `task` -- the task fixes an exact modeling recipe (specific feature/target columns, a stated missing-value handling rule, a fixed train/test split fraction, and a named error metric) and asks for a single numeric error value.
- **Pattern**: The attempt reports a metric value without a saved, re-runnable script that implements each stated step literally on the full dataset — e.g. it silently drops rows with missing values (or lets non-numeric/parsed-as-string columns force coercion) instead of mean-imputing exactly the named columns before splitting, subsets/filters rows the task never mentioned, swaps target and predictors, or evaluates on the training portion — and it never checks whether the reported error is plausible relative to the target's own variance.
- **Detection procedure**:
  1. From the task, list every mandated step: exact target and predictor columns, missing-value rule, split fraction, metric definition, rounding/format.
  2. Open the scripts and match each mandated step to a concrete line; flag if no script exists, if rows are dropped/filtered anywhere, if the named columns are not numeric-cast then imputed with their own means over the full dataset before splitting, or if the split fraction/metric differs.
  3. Check that the metric is computed on the held-out portion only, with predictions and truths from the same rows and the same target variable (not a transformed/scaled version).
  4. Compare the reported error to the variance (or std²) of the target computed under the same imputation; an ordinary-least-squares fit should not do much worse than predicting the mean, so an error far above the target variance (or a residual scale implausible given the target's range) must be explained.
- **Discriminator**: A genuine violation is an unexplained deviation from the recipe or an error magnitude that exceeds the mean-only baseline; a look-alike that is fine is a large-but-consistent error simply because the target has a large scale/variance and the recipe is implemented exactly as stated (split seed differences alone give only modest variation, not order-of-magnitude ones).
- **Consequence**: The reported error differs from the reference value (often by an order of magnitude) and the answer is graded wrong, with no script available to localize the deviating step.
285Silently re-ordering rows before a row-order-dependent computationtaskinfiagent-dabench
Applies when
task -- the requested quantity depends on row adjacency or sequence (lags, differences, cumulative sums, first/last, rolling windows) and the script sorts, reverses, or re-indexes the data before computing it.
Pattern
The agent assumes the file's stored order is "wrong", parses a key (often an ambiguous date format) and sorts by it, then computes the lag-based column on the re-ordered frame — without any check that the new order matches the order implied by the task, and without comparing results under the original order. A reversed sequence flips the sign of differences/returns and changes the statistic, yet nothing in the script would reveal it.
Detection procedure
  1. In the task statement, look for any explicit instruction about ordering ("previous day", "previous row", "sorted by ..."); note whether the task specifies re-sorting or only implies "previous row as stored".
  2. In the script, find any sort_values, sort_index, [::-1], reset_index, or re-parsing of a key that precedes the shift/diff/pct_change call; check whether the parsing itself is risky (e.g. two-digit years, mixed formats, non-ISO strings).
  3. Check whether the script validates the resulting order (prints head/tail, asserts monotonicity of the key, confirms the parsed key round-trips) and whether it computes the statistic under both the original and sorted order to see if the conclusion changes.
  4. Sanity-check the reported number's sign/magnitude against the data's overall trend (e.g. does the mean per-step change agree with (last − first)/n over the same ordering?).
Discriminator
A real violation is an unverified reordering (or unverified reliance on stored order) that materially changes a sequence-dependent result; it is fine if the script asserts the key is monotonic after sorting, demonstrates the stored order was already correct/incorrect, or shows the statistic is invariant to the ordering choice.
Consequence
The lagged differences are computed backwards, so the mean comes out with the opposite sign (and dispersion statistics shift slightly), and the graded values fail to match the expected answer.
id 7b48261a1864 · mined from infiagent-dabench dabench-75@s5
raw text (what the judge reads)
### Silently re-ordering rows before a row-order-dependent computation
- **Applies when**: `task` -- the requested quantity depends on row adjacency or sequence (lags, differences, cumulative sums, first/last, rolling windows) and the script sorts, reverses, or re-indexes the data before computing it.
- **Pattern**: The agent assumes the file's stored order is "wrong", parses a key (often an ambiguous date format) and sorts by it, then computes the lag-based column on the re-ordered frame — without any check that the new order matches the order implied by the task, and without comparing results under the original order. A reversed sequence flips the sign of differences/returns and changes the statistic, yet nothing in the script would reveal it.
- **Detection procedure**:
  1. In the task statement, look for any explicit instruction about ordering ("previous day", "previous row", "sorted by ..."); note whether the task specifies re-sorting or only implies "previous row as stored".
  2. In the script, find any `sort_values`, `sort_index`, `[::-1]`, `reset_index`, or re-parsing of a key that precedes the `shift`/`diff`/`pct_change` call; check whether the parsing itself is risky (e.g. two-digit years, mixed formats, non-ISO strings).
  3. Check whether the script validates the resulting order (prints head/tail, asserts monotonicity of the key, confirms the parsed key round-trips) **and** whether it computes the statistic under both the original and sorted order to see if the conclusion changes.
  4. Sanity-check the reported number's sign/magnitude against the data's overall trend (e.g. does the mean per-step change agree with (last − first)/n over the same ordering?).
- **Discriminator**: A real violation is an unverified reordering (or unverified reliance on stored order) that materially changes a sequence-dependent result; it is fine if the script asserts the key is monotonic after sorting, demonstrates the stored order was already correct/incorrect, or shows the statistic is invariant to the ordering choice.
- **Consequence**: The lagged differences are computed backwards, so the mean comes out with the opposite sign (and dispersion statistics shift slightly), and the graded values fail to match the expected answer.
286Small numeric deviation from an unverified row set / unclean numeric coerciontaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic (correlation, mean, metric) over two or more columns of a table, to be reported at a fixed rounding precision.
Pattern
The attempt loads the file with default settings and calls the statistic directly, without ever establishing which rows actually enter the computation: non-numeric placeholders ("NA", "-", blanks, thousands separators) silently make a column object dtype or get coerced/dropped, duplicate/aggregate/header-like rows are left in, or missing values are dropped column-wise instead of pairwise (or vice versa). No record count, dtype, or range check is printed, so a result that is close-but-not-equal to the correct value looks plausible and is submitted. (Here it is aggravated by no script being saved at all, so the row set cannot even be audited.)
Detection procedure
  1. Read the task and note the required precision and any stated filtering; a two-decimal (or finer) answer means the exact set of contributing rows matters.
  2. In the scripts, check whether the loaded columns are explicitly coerced to numeric and whether the number of rows before and after cleaning/NaN handling is printed or asserted; check whether any rows are excluded and why.
  3. Check whether the reported statistic is accompanied by a sanity report (N used, min/max, dtypes) or a cross-check with a second implementation (e.g., a different library or a manual formula) that agrees to the reported precision.
  4. If none of the above exist — or no script exists at all — treat the number as unverified, and additionally check that every reported field respects the requested formatting (e.g., a p-value printed as 0.0 rather than the requested 4-decimal form).
Discriminator
A real violation is an attempt whose contributing-row set and dtypes are never established or shown, so an off-by-a-few-rows difference would go unnoticed; it is fine if the script explicitly coerces types, reports N used and how missing/invalid entries were handled, and the result is reproducible/cross-checked — even if some rows are legitimately dropped.
Consequence
The reported statistic differs from ground truth in the last reported digit (e.g., 0.53 vs 0.54) and/or a field is formatted outside the requested precision, so exact-match grading fails even though the qualitative conclusion is right.
id 68b22b370d84 · mined from infiagent-dabench dabench-300@s5
raw text (what the judge reads)
### Small numeric deviation from an unverified row set / unclean numeric coercion
- **Applies when**: `task` -- the task asks for a single numeric statistic (correlation, mean, metric) over two or more columns of a table, to be reported at a fixed rounding precision.
- **Pattern**: The attempt loads the file with default settings and calls the statistic directly, without ever establishing which rows actually enter the computation: non-numeric placeholders ("NA", "-", blanks, thousands separators) silently make a column object dtype or get coerced/dropped, duplicate/aggregate/header-like rows are left in, or missing values are dropped column-wise instead of pairwise (or vice versa). No record count, dtype, or range check is printed, so a result that is close-but-not-equal to the correct value looks plausible and is submitted. (Here it is aggravated by no script being saved at all, so the row set cannot even be audited.)
- **Detection procedure**:
  1. Read the task and note the required precision and any stated filtering; a two-decimal (or finer) answer means the exact set of contributing rows matters.
  2. In the scripts, check whether the loaded columns are explicitly coerced to numeric and whether the number of rows before and after cleaning/NaN handling is printed or asserted; check whether any rows are excluded and why.
  3. Check whether the reported statistic is accompanied by a sanity report (N used, min/max, dtypes) or a cross-check with a second implementation (e.g., a different library or a manual formula) that agrees to the reported precision.
  4. If none of the above exist — or no script exists at all — treat the number as unverified, and additionally check that every reported field respects the requested formatting (e.g., a p-value printed as `0.0` rather than the requested 4-decimal form).
- **Discriminator**: A real violation is an attempt whose contributing-row set and dtypes are never established or shown, so an off-by-a-few-rows difference would go unnoticed; it is fine if the script explicitly coerces types, reports N used and how missing/invalid entries were handled, and the result is reproducible/cross-checked — even if some rows are legitimately dropped.
- **Consequence**: The reported statistic differs from ground truth in the last reported digit (e.g., 0.53 vs 0.54) and/or a field is formatted outside the requested precision, so exact-match grading fails even though the qualitative conclusion is right.
287Single un-benchmarked model accepted as the final prediction for an accuracy-graded tasktaskda-code
Applies when
task -- the deliverable is a file of predicted values for a held-out set, so grading depends on how close the predictions are, not merely on the file existing.
Pattern
The agent fits one off-the-shelf regressor with hand-picked hyperparameters, ordinal/label-encodes high-cardinality categorical columns, leaves a heavily skewed target untransformed, drops potentially informative columns without justification, and then ships the predictions after a single train/validation split — never comparing the achieved error against a baseline, alternative models, cross-validation, or any explicit accuracy threshold.
Detection procedure
  1. Read the task and confirm the output is scored on prediction closeness (a regression/classification submission), which implies a quality bar, not just a format bar.
  2. Read the scripts: check whether more than one model/feature configuration is tried, whether validation is repeated (CV) rather than a single random split, whether skewed targets are log/robust-transformed, whether dropped columns and naive integer encodings of many-level categoricals were tested for impact.
  3. Compare the reported validation error to the target's scale (e.g., MAE or RMSE divided by target mean/IQR): if the relative error is large or is never contrasted with a trivial baseline (mean/median predictor, group means), the accuracy claim is unsupported.
  4. Inspect the emitted predictions' distribution (min/max/mean/spread, count) against the training target distribution; flag if the spread is compressed, contains implausible extremes, or the row count/order is not verified against the test file.
Discriminator
A real violation is an attempt with zero model selection evidence and no error-vs-baseline comparison; it is not a violation if the agent ran cross-validation or compared several configurations, justified the encoding/target transform choices, and showed the chosen model beats a trivial baseline by a clear margin with a sanity-checked prediction distribution.
Consequence
The submitted file is well-formed but the predictions miss the reference values by more than the grader's tolerance, so the accuracy check fails (file reported WRONG) even though every script ran without error.
id 8ef3f15a41d9 · mined from da-code dacode-ml-regression-014@s5
raw text (what the judge reads)
### Single un-benchmarked model accepted as the final prediction for an accuracy-graded task
- **Applies when**: `task` -- the deliverable is a file of predicted values for a held-out set, so grading depends on how close the predictions are, not merely on the file existing.
- **Pattern**: The agent fits one off-the-shelf regressor with hand-picked hyperparameters, ordinal/label-encodes high-cardinality categorical columns, leaves a heavily skewed target untransformed, drops potentially informative columns without justification, and then ships the predictions after a single train/validation split — never comparing the achieved error against a baseline, alternative models, cross-validation, or any explicit accuracy threshold.
- **Detection procedure**:
  1. Read the task and confirm the output is scored on prediction closeness (a regression/classification submission), which implies a quality bar, not just a format bar.
  2. Read the scripts: check whether more than one model/feature configuration is tried, whether validation is repeated (CV) rather than a single random split, whether skewed targets are log/robust-transformed, whether dropped columns and naive integer encodings of many-level categoricals were tested for impact.
  3. Compare the reported validation error to the target's scale (e.g., MAE or RMSE divided by target mean/IQR): if the relative error is large or is never contrasted with a trivial baseline (mean/median predictor, group means), the accuracy claim is unsupported.
  4. Inspect the emitted predictions' distribution (min/max/mean/spread, count) against the training target distribution; flag if the spread is compressed, contains implausible extremes, or the row count/order is not verified against the test file.
- **Discriminator**: A real violation is an attempt with zero model selection evidence and no error-vs-baseline comparison; it is *not* a violation if the agent ran cross-validation or compared several configurations, justified the encoding/target transform choices, and showed the chosen model beats a trivial baseline by a clear margin with a sanity-checked prediction distribution.
- **Consequence**: The submitted file is well-formed but the predictions miss the reference values by more than the grader's tolerance, so the accuracy check fails (file reported WRONG) even though every script ran without error.
288Output artifacts written to an arbitrary location/name instead of the task's expected directory and naming conventiontaskinfiagent-dabench
Applies when
task -- the deliverable includes file paths of generated artifacts (cleaned/normalized/predicted CSVs, models) that must be reported in the answer.
Pattern
The agent reads the input from the provided working directory but writes outputs to its own home/CWD or with an ad-hoc filename, then reports that path; the grader expects artifacts alongside the source data with a name derived from the input file and the operation performed.
Detection procedure
  1. In the task statement, note where the input data lives (the path used to load it) and any explicit or implied naming convention for the outputs.
  2. In the scripts, find every write call (to_csv, save, open(..., 'w')) and record the directory and filename used; check whether the output directory equals the input data directory and whether the filename echoes the input file's base name plus the transformation.
  3. Compare the paths printed in the final answer to those write calls; confirm they are absolute, existent, and in the expected location.
  4. Flag if the outputs sit in a different directory than the input data or use a name unrelated to the input file, even when all numeric fields are correct.
Discriminator
A real violation is writing outside the data directory or using a name that does not follow the input-derived convention; it is fine if the task explicitly permits any path, or if the agent writes to the same directory as the input with a clearly derived name (a purely cosmetic capitalization/separator difference may still be tolerated if the grader matches loosely, but a different directory never is).
Consequence
The path field fails the string/existence comparison, so the whole answer tuple is marked wrong (0 checks passed) despite correct statistics.
id 62375f9a465f · mined from infiagent-dabench dabench-743@s5
raw text (what the judge reads)
### Output artifacts written to an arbitrary location/name instead of the task's expected directory and naming convention
- **Applies when**: `task` -- the deliverable includes file paths of generated artifacts (cleaned/normalized/predicted CSVs, models) that must be reported in the answer.
- **Pattern**: The agent reads the input from the provided working directory but writes outputs to its own home/CWD or with an ad-hoc filename, then reports that path; the grader expects artifacts alongside the source data with a name derived from the input file and the operation performed.
- **Detection procedure**:
  1. In the task statement, note where the input data lives (the path used to load it) and any explicit or implied naming convention for the outputs.
  2. In the scripts, find every write call (`to_csv`, `save`, `open(..., 'w')`) and record the directory and filename used; check whether the output directory equals the input data directory and whether the filename echoes the input file's base name plus the transformation.
  3. Compare the paths printed in the final answer to those write calls; confirm they are absolute, existent, and in the expected location.
  4. Flag if the outputs sit in a different directory than the input data or use a name unrelated to the input file, even when all numeric fields are correct.
- **Discriminator**: A real violation is writing outside the data directory or using a name that does not follow the input-derived convention; it is fine if the task explicitly permits any path, or if the agent writes to the same directory as the input with a clearly derived name (a purely cosmetic capitalization/separator difference may still be tolerated if the grader matches loosely, but a different directory never is).
- **Consequence**: The path field fails the string/existence comparison, so the whole answer tuple is marked wrong (0 checks passed) despite correct statistics.
289Deliverable artifacts not produced (answer text substituted for required output files)taskda-code
Applies when
task -- the task explicitly asks for saved outputs (chart image, config/spec file, serialized array/table) in addition to or instead of a textual finding.
Pattern
The agent computes and reports only the intermediate textual result (e.g., the selected group/label) and never writes, or only partially writes, the required files with the required names/formats/paths; no script is retained that would regenerate them.
Detection procedure
  1. Enumerate every artifact the task names or implies (file names, extensions, directories, and any referenced spec/config that dictates plot styling or content).
  2. Scan the scripts for explicit save calls covering each artifact, with the exact filenames and locations requested, and check that any provided spec file is actually read and applied rather than hard-coded defaults.
  3. Compare the submitted answer/workspace against that checklist; flag if any artifact is missing, misnamed, or unverifiable because no script exists.
  4. Confirm the numbers backing each artifact come from the requested final quantity and subset, not an intermediate step.
Discriminator
A real violation is a missing/misnamed/never-written artifact or an unloaded spec file; it is fine if all artifacts exist with correct names and the text answer is merely an additional summary.
Consequence
Grader checks for the expected files fail (WRONG/MISSING) even if the reported textual value happens to be right, yielding zero credit.
id dccbffec6874 · mined from da-code dacode-plot-pie-008@s5
raw text (what the judge reads)
### Deliverable artifacts not produced (answer text substituted for required output files)
- **Applies when**: `task` -- the task explicitly asks for saved outputs (chart image, config/spec file, serialized array/table) in addition to or instead of a textual finding.
- **Pattern**: The agent computes and reports only the intermediate textual result (e.g., the selected group/label) and never writes, or only partially writes, the required files with the required names/formats/paths; no script is retained that would regenerate them.
- **Detection procedure**:
  1. Enumerate every artifact the task names or implies (file names, extensions, directories, and any referenced spec/config that dictates plot styling or content).
  2. Scan the scripts for explicit save calls covering each artifact, with the exact filenames and locations requested, and check that any provided spec file is actually read and applied rather than hard-coded defaults.
  3. Compare the submitted answer/workspace against that checklist; flag if any artifact is missing, misnamed, or unverifiable because no script exists.
  4. Confirm the numbers backing each artifact come from the requested final quantity and subset, not an intermediate step.
- **Discriminator**: A real violation is a missing/misnamed/never-written artifact or an unloaded spec file; it is fine if all artifacts exist with correct names and the text answer is merely an additional summary.
- **Consequence**: Grader checks for the expected files fail (WRONG/MISSING) even if the reported textual value happens to be right, yielding zero credit.
290Deliverable file not written/filled in the required formattaskda-code
Applies when
task -- the task names a specific output file (provided template or schema) that must contain the results, and the agent instead reports findings in prose.
Pattern
The attempt performs the analysis and prints a narrative summary, but never programmatically loads the provided template, writes the computed values into its exact columns/rows, and saves it back; no script step produces or verifies the required artifact.
Detection procedure
  1. Read the task and list every required output artifact, its path, and its stated format/schema constraints.
  2. Search the scripts for code that reads the template, populates it, and writes it to the required path (e.g., a save/write call to that filename), plus a re-read of the saved file to confirm columns, row count, and value types.
  3. Compare the answer text: does it merely narrate numbers, or does it reference the written file and echo its contents in the template's schema (same category labels, same ordering, same field names)?
  4. Flag if any required artifact has no write step, or if written values/labels deviate from the template's expected categories/ordering/dtypes.
Discriminator
A real violation is missing the write step, writing to a different path/name, or overwriting the template with a different schema (extra/renamed/reordered columns, invented category labels, free-text instead of numbers). It is fine if the file is written exactly per template and the prose is just an additional summary.
Consequence
The grader compares the expected file and reports it WRONG/MISSING, scoring 0 regardless of whether the underlying analysis numbers were reasonable.
id 8b22046024f0 · mined from da-code dacode-dm-csv-001@s5
raw text (what the judge reads)
### Deliverable file not written/filled in the required format
- **Applies when**: `task` -- the task names a specific output file (provided template or schema) that must contain the results, and the agent instead reports findings in prose.
- **Pattern**: The attempt performs the analysis and prints a narrative summary, but never programmatically loads the provided template, writes the computed values into its exact columns/rows, and saves it back; no script step produces or verifies the required artifact.
- **Detection procedure**:
  1. Read the task and list every required output artifact, its path, and its stated format/schema constraints.
  2. Search the scripts for code that reads the template, populates it, and writes it to the required path (e.g., a save/write call to that filename), plus a re-read of the saved file to confirm columns, row count, and value types.
  3. Compare the answer text: does it merely narrate numbers, or does it reference the written file and echo its contents in the template's schema (same category labels, same ordering, same field names)?
  4. Flag if any required artifact has no write step, or if written values/labels deviate from the template's expected categories/ordering/dtypes.
- **Discriminator**: A real violation is missing the write step, writing to a different path/name, or overwriting the template with a different schema (extra/renamed/reordered columns, invented category labels, free-text instead of numbers). It is fine if the file is written exactly per template and the prose is just an additional summary.
- **Consequence**: The grader compares the expected file and reports it WRONG/MISSING, scoring 0 regardless of whether the underlying analysis numbers were reasonable.
291Group-splitting and per-entity aggregation defined ad hoc, with no reproducible script or sanity checktaskinfiagent-dabench
Applies when
task -- the task asks for a statistic (correlation, mean difference, model score) computed per entity after aggregating multi-row records and after splitting the entities into subgroups by a threshold such as a median.
Pattern
The attempt jumps to a number without pinning down (a) how rows are collapsed to one record per entity (max/first/last, duplicate entity names across years, records with missing timestamps), (b) how the derived quantity is defined (e.g., span endpoints inclusive vs. exclusive, counted units vs. elapsed units), and (c) which side of the threshold ties go to (>= vs > the median, entities exactly at the median). Each choice shifts the reported coefficient by a few hundredths, and with no saved script the reviewer cannot see or reproduce the choice actually made.
Detection procedure
  1. Read the task and list every derived quantity and every subgroup definition it implies; note which ones the task leaves ambiguous.
  2. Inspect the scripts for an explicit, single aggregation step producing one row per entity (groupby on a truly unique entity key) and an explicit threshold comparison; if scripts are absent or the aggregation is implicit, the attempt is already unverifiable.
  3. Check the scripts print sanity counts before the statistic: number of unique entities, size of each subgroup (should be near-equal for a median split and sum to the total), and min/max of the derived quantity.
  4. Compare the reported number against at least one alternative defensible convention (tie-handling, inclusive vs. exclusive span); if the answer moves more than the required rounding precision and no justification is given, flag it.
Discriminator
A real violation is when the aggregation/threshold convention is undocumented or the subgroup counts/entity counts were never printed, so the reported value cannot be reproduced or defended. A look-alike that is fine is an attempt that states the convention, shows the entity and subgroup counts, and demonstrates the statistic is stable (to the requested rounding) across reasonable conventions.
Consequence
The reported categorical verdict (e.g., "linear") may still match, but the numeric coefficient is off by a small amount (e.g., 0.58 vs. 0.56), so exact-value checks fail and the submission is graded incorrect.
id 58d36340bc70 · mined from infiagent-dabench dabench-431@s5
raw text (what the judge reads)
### Group-splitting and per-entity aggregation defined ad hoc, with no reproducible script or sanity check
- **Applies when**: `task` -- the task asks for a statistic (correlation, mean difference, model score) computed per entity after aggregating multi-row records and after splitting the entities into subgroups by a threshold such as a median.
- **Pattern**: The attempt jumps to a number without pinning down (a) how rows are collapsed to one record per entity (max/first/last, duplicate entity names across years, records with missing timestamps), (b) how the derived quantity is defined (e.g., span endpoints inclusive vs. exclusive, counted units vs. elapsed units), and (c) which side of the threshold ties go to (`>=` vs `>` the median, entities exactly at the median). Each choice shifts the reported coefficient by a few hundredths, and with no saved script the reviewer cannot see or reproduce the choice actually made.
- **Detection procedure**:
  1. Read the task and list every derived quantity and every subgroup definition it implies; note which ones the task leaves ambiguous.
  2. Inspect the scripts for an explicit, single aggregation step producing one row per entity (groupby on a truly unique entity key) and an explicit threshold comparison; if scripts are absent or the aggregation is implicit, the attempt is already unverifiable.
  3. Check the scripts print sanity counts before the statistic: number of unique entities, size of each subgroup (should be near-equal for a median split and sum to the total), and min/max of the derived quantity.
  4. Compare the reported number against at least one alternative defensible convention (tie-handling, inclusive vs. exclusive span); if the answer moves more than the required rounding precision and no justification is given, flag it.
- **Discriminator**: A real violation is when the aggregation/threshold convention is undocumented or the subgroup counts/entity counts were never printed, so the reported value cannot be reproduced or defended. A look-alike that is fine is an attempt that states the convention, shows the entity and subgroup counts, and demonstrates the statistic is stable (to the requested rounding) across reasonable conventions.
- **Consequence**: The reported categorical verdict (e.g., "linear") may still match, but the numeric coefficient is off by a small amount (e.g., 0.58 vs. 0.56), so exact-value checks fail and the submission is graded incorrect.
292Answer string does not match the literal answer template (extra quoting/delimiters)taskinfiagent-dabench
Applies when
task -- the task specifies an exact answer format token (e.g., @key[item1, item2, …]) and the agent must emit a list or value inside it.
Pattern
The agent computes the right values but serializes them using a language-native repr (Python list/dict str(), quotes around strings, numpy types, trailing commas, different bracket or separator style) instead of writing the plain items exactly as the template shows, so a literal/string-matching grader rejects a substantively correct result.
Detection procedure
  1. Copy the answer template from the task verbatim and note every literal character: prefix marker, brackets, separator, spacing, and whether items appear quoted or bare.
  2. In the scripts, find the line that builds the final answer string; check whether it prints a computed container directly (e.g., print(f"@key{my_list}")) rather than joining items into the required literal form.
  3. Compare the submitted answer character-by-character with the template, ignoring the item values themselves; flag any added quotes, changed brackets/separators, casing, or missing prefix.
  4. Also confirm any stated post-processing on the items (sorting, dedup, rounding, units) is applied in the string that is actually emitted, not only in an intermediate variable.
Discriminator
A real violation is a formatting/serialization mismatch with the template (quotes, brackets, separators, prefix) even when values are right; a look-alike that is fine is a template whose example itself shows quoted items or where the task explicitly allows any list-like rendering.
Consequence
The grader reports the expected key as WRONG/MISSING and echoes a submitted string that visibly contains the correct values, yielding 0/1 checks passed.
id aff2ee9fd65c · mined from infiagent-dabench dabench-207@s5
raw text (what the judge reads)
### Answer string does not match the literal answer template (extra quoting/delimiters)
- **Applies when**: `task` -- the task specifies an exact answer format token (e.g., `@key[item1, item2, …]`) and the agent must emit a list or value inside it.
- **Pattern**: The agent computes the right values but serializes them using a language-native repr (Python list/dict `str()`, quotes around strings, `numpy` types, trailing commas, different bracket or separator style) instead of writing the plain items exactly as the template shows, so a literal/string-matching grader rejects a substantively correct result.
- **Detection procedure**:
  1. Copy the answer template from the task verbatim and note every literal character: prefix marker, brackets, separator, spacing, and whether items appear quoted or bare.
  2. In the scripts, find the line that builds the final answer string; check whether it prints a computed container directly (e.g., `print(f"@key{my_list}")`) rather than joining items into the required literal form.
  3. Compare the submitted answer character-by-character with the template, ignoring the item values themselves; flag any added quotes, changed brackets/separators, casing, or missing prefix.
  4. Also confirm any stated post-processing on the items (sorting, dedup, rounding, units) is applied in the string that is actually emitted, not only in an intermediate variable.
- **Discriminator**: A real violation is a formatting/serialization mismatch with the template (quotes, brackets, separators, prefix) even when values are right; a look-alike that is fine is a template whose example itself shows quoted items or where the task explicitly allows any list-like rendering.
- **Consequence**: The grader reports the expected key as WRONG/MISSING and echoes a submitted string that visibly contains the correct values, yielding 0/1 checks passed.
293Optimizing/validating on a ranking score, then hard-thresholding at 0.5 without checking the decision metric or class balancetaskda-code
Applies when
task -- the deliverable is a set of hard class labels (not probabilities) for a held-out file, and the scripts train probabilistic classifiers on an imbalanced target.
Pattern
The attempt validates and weights models with a threshold-free metric (e.g., ROC-AUC), blends probabilities from models trained under inconsistent objectives (some with class re-weighting, some without, some on scaled and some on raw features), then converts to labels with a hard-coded 0.5 cutoff. No score is ever computed on the validation split with the metric the labels will actually be judged by (accuracy/F1/balanced accuracy), and the resulting predicted positive rate is never compared with the base rate in the training labels.
Detection procedure
  1. Read the task: confirm the requested output is discrete labels, so the grade depends on a threshold, not on ranking quality.
  2. Read the training script: check whether any validation-set score is computed on the thresholded predictions and whether the cutoff was chosen/tuned against that score; check whether the blended models were fit on comparable targets/feature representations (mixing class_weight='balanced' with unweighted models makes the averaged probability meaningless in absolute terms).
  3. Read the printed/produced predictions: compute the fraction of positive labels and compare it with the positive rate of the training target; a large gap (e.g., predicting far fewer positives than the base rate, or far more) with no justification is the flag.
  4. Check that the output file's columns and row order match exactly what was asked (only/at least the named prediction column, one row per test row in test order).
Discriminator
A real violation is when the cutoff is arbitrary AND the label-level metric was never measured, or the mixed probabilities are not on a common scale. It is not a violation if the agent reports the label-level validation metric (and/or sweeps thresholds) and the predicted positive rate is defensibly close to the expected prevalence — a 0.5 cutoff can be correct when explicitly validated.
Consequence
The submitted labels are systematically biased toward the majority class, so accuracy/F1 against the held-out truth falls below the grader's threshold and the file is marked WRONG even though the internally reported AUC looked healthy.
id 6b4b12a5dfc5 · mined from da-code dacode-ml-binary-016@s5
raw text (what the judge reads)
### Optimizing/validating on a ranking score, then hard-thresholding at 0.5 without checking the decision metric or class balance
- **Applies when**: `task` -- the deliverable is a set of hard class labels (not probabilities) for a held-out file, and the scripts train probabilistic classifiers on an imbalanced target.
- **Pattern**: The attempt validates and weights models with a threshold-free metric (e.g., ROC-AUC), blends probabilities from models trained under inconsistent objectives (some with class re-weighting, some without, some on scaled and some on raw features), then converts to labels with a hard-coded 0.5 cutoff. No score is ever computed on the validation split with the metric the labels will actually be judged by (accuracy/F1/balanced accuracy), and the resulting predicted positive rate is never compared with the base rate in the training labels.
- **Detection procedure**:
  1. Read the task: confirm the requested output is discrete labels, so the grade depends on a threshold, not on ranking quality.
  2. Read the training script: check whether any validation-set score is computed on the *thresholded* predictions and whether the cutoff was chosen/tuned against that score; check whether the blended models were fit on comparable targets/feature representations (mixing `class_weight='balanced'` with unweighted models makes the averaged probability meaningless in absolute terms).
  3. Read the printed/produced predictions: compute the fraction of positive labels and compare it with the positive rate of the training target; a large gap (e.g., predicting far fewer positives than the base rate, or far more) with no justification is the flag.
  4. Check that the output file's columns and row order match exactly what was asked (only/at least the named prediction column, one row per test row in test order).
- **Discriminator**: A real violation is when the cutoff is arbitrary AND the label-level metric was never measured, or the mixed probabilities are not on a common scale. It is *not* a violation if the agent reports the label-level validation metric (and/or sweeps thresholds) and the predicted positive rate is defensibly close to the expected prevalence — a 0.5 cutoff can be correct when explicitly validated.
- **Consequence**: The submitted labels are systematically biased toward the majority class, so accuracy/F1 against the held-out truth falls below the grader's threshold and the file is marked WRONG even though the internally reported AUC looked healthy.
294Ignoring a stated randomness/seed constraint that implies a resampling-based methodtaskda-code
Applies when
task -- the prompt asks for a statistic/p-value and explicitly specifies setting a random seed (or otherwise names a simulation/resampling framing), and the scripts must produce a required output file in a given template format.
Pattern
The attempt computes the quantity with a closed-form/analytic routine (e.g., a canned parametric test) that consumes no randomness, so the seed is irrelevant; the seed instruction is never used, and the reported number cannot be reproduced by the intended simulation-based procedure. The answer is also often narrated in prose rather than verified against the required file schema.
Detection procedure
  1. Read the task for explicit constraints about randomness (seed), method family, rounding, units, and the required output file/columns of the sample template.
  2. Search the scripts for any use of the seed and of a random draw (permutation, bootstrap, simulation loop); if the seed is set but no random number generator is ever consumed, or no seed appears at all, the stated method was ignored.
  3. Check that the computed statistic's definition matches the method implied by the task framing (resampling under the null vs. parametric assumption), and that the p-value is derived from that null distribution.
  4. Confirm the script actually writes result.csv with the same column names/row structure and value precision as the provided sample, and that the reported number equals what is in the file.
Discriminator
A real violation is when the requested method inherently requires randomness and none is used (seed unused, no resampling), or the output file/schema is absent/mismatched. A look-alike that is fine: the script does run a seeded resampling procedure and additionally reports a parametric statistic for context, or the seed is genuinely irrelevant because the task never mentioned it.
Consequence
The p-value differs from the seeded resampling reference value (and/or the file/schema doesn't match), so the file check fails and the task is scored 0.
id bf57c1b89a4b · mined from da-code dacode-data-sa-039@s5
raw text (what the judge reads)
### Ignoring a stated randomness/seed constraint that implies a resampling-based method
- **Applies when**: `task` -- the prompt asks for a statistic/p-value and explicitly specifies setting a random seed (or otherwise names a simulation/resampling framing), and the scripts must produce a required output file in a given template format.
- **Pattern**: The attempt computes the quantity with a closed-form/analytic routine (e.g., a canned parametric test) that consumes no randomness, so the seed is irrelevant; the seed instruction is never used, and the reported number cannot be reproduced by the intended simulation-based procedure. The answer is also often narrated in prose rather than verified against the required file schema.
- **Detection procedure**:
  1. Read the task for explicit constraints about randomness (seed), method family, rounding, units, and the required output file/columns of the sample template.
  2. Search the scripts for any use of the seed and of a random draw (permutation, bootstrap, simulation loop); if the seed is set but no random number generator is ever consumed, or no seed appears at all, the stated method was ignored.
  3. Check that the computed statistic's definition matches the method implied by the task framing (resampling under the null vs. parametric assumption), and that the p-value is derived from that null distribution.
  4. Confirm the script actually writes `result.csv` with the same column names/row structure and value precision as the provided sample, and that the reported number equals what is in the file.
- **Discriminator**: A real violation is when the requested method inherently requires randomness and none is used (seed unused, no resampling), or the output file/schema is absent/mismatched. A look-alike that is fine: the script does run a seeded resampling procedure and additionally reports a parametric statistic for context, or the seed is genuinely irrelevant because the task never mentioned it.
- **Consequence**: The p-value differs from the seeded resampling reference value (and/or the file/schema doesn't match), so the file check fails and the task is scored 0.
295Submitting predictions without any held-out validation of predictive accuracytaskda-code
Applies when
task -- the task asks the agent to produce predictions for an unlabeled test file, and the scripts fit a model on the full labeled data and immediately write the prediction file.
Pattern
The attempt treats the job as "produce a correctly-shaped file" rather than "produce accurate values": it imports a train/validation splitter or scaler but never uses it, computes no error metric on any labeled holdout, tries no alternative model/feature set, and drops potentially informative columns (e.g., timestamps or other non-numeric fields) without checking their predictive value. The final answer reports only prediction summary statistics (count, min/max/mean) as evidence of success.
Detection procedure
  1. From the task, note that grading depends on how close the predicted values are to hidden ground truth, not just on file/column names.
  2. Scan the script for any labeled-data evaluation: a train/validation or time-ordered split, cross-validation, or a printed error metric (MAE/RMSE/R²/accuracy). Also check whether any input column is dropped without justification and whether the split respects the data's structure (temporal/grouped).
  3. Read the answer for a reported out-of-sample score and any comparison against a trivial baseline (predicting the mean/last value) or a second model.
  4. Flag the attempt if no out-of-sample error number exists anywhere, or if the only "validation" is descriptive statistics of the predictions themselves.
Discriminator
A genuine violation has zero quantitative evidence that the model beats a naive baseline on unseen labeled data. It is not a violation if the script prints a holdout/CV score (even from a single simple split) and the answer states it, or if the task explicitly forbids holding out data; distribution checks of predictions are fine as an additional sanity check but never substitute for a measured error.
Consequence
The written file has the right shape and column name but values far from ground truth, so the grader's tolerance/score threshold on the prediction file fails while the agent confidently reports success.
id fb671710e42a · mined from da-code dacode-ml-regression-015@s5
raw text (what the judge reads)
### Submitting predictions without any held-out validation of predictive accuracy
- **Applies when**: `task` -- the task asks the agent to produce predictions for an unlabeled test file, and the scripts fit a model on the full labeled data and immediately write the prediction file.
- **Pattern**: The attempt treats the job as "produce a correctly-shaped file" rather than "produce accurate values": it imports a train/validation splitter or scaler but never uses it, computes no error metric on any labeled holdout, tries no alternative model/feature set, and drops potentially informative columns (e.g., timestamps or other non-numeric fields) without checking their predictive value. The final answer reports only prediction summary statistics (count, min/max/mean) as evidence of success.
- **Detection procedure**:
  1. From the task, note that grading depends on how close the predicted values are to hidden ground truth, not just on file/column names.
  2. Scan the script for any labeled-data evaluation: a train/validation or time-ordered split, cross-validation, or a printed error metric (MAE/RMSE/R²/accuracy). Also check whether any input column is dropped without justification and whether the split respects the data's structure (temporal/grouped).
  3. Read the answer for a reported out-of-sample score and any comparison against a trivial baseline (predicting the mean/last value) or a second model.
  4. Flag the attempt if no out-of-sample error number exists anywhere, or if the only "validation" is descriptive statistics of the predictions themselves.
- **Discriminator**: A genuine violation has zero quantitative evidence that the model beats a naive baseline on unseen labeled data. It is *not* a violation if the script prints a holdout/CV score (even from a single simple split) and the answer states it, or if the task explicitly forbids holding out data; distribution checks of predictions are fine as an *additional* sanity check but never substitute for a measured error.
- **Consequence**: The written file has the right shape and column name but values far from ground truth, so the grader's tolerance/score threshold on the prediction file fails while the agent confidently reports success.
296Unverified dtype handling when computing extremes (formatted numbers left as strings)taskda-code
Applies when
task -- the task asks for the argmax/argmin (or mean-imputation) of a quantitative column that may be stored as text with thousands separators, percent/currency symbols, or footnote characters.
Pattern
The attempt loads the file and directly calls max/min/idxmax/sort_values (or imputes with a "mean") on a column whose dtype is still object, so comparison is lexicographic rather than numeric; no dtype check, no printout of the winning row's value, and no evidence (script or result file) that the numeric conversion happened.
Detection procedure
  1. From the task, note which column drives the requested extreme/statistic and whether imputation is required (imputation is only possible on a numeric column, so a numeric cast must appear somewhere).
  2. In the scripts, look for an explicit cleaning/casting step (strip separators/symbols, pd.to_numeric(..., errors=...)) before the aggregation, plus a check of dtype/NaN counts; flag if the aggregation runs on raw loaded values or if no script/log exists at all.
  3. Check whether the attempt printed the extreme value alongside the reported label and sanity-checked it against the plausible range and against the top/bottom few rows.
  4. Compare the reported answer to the requested output schema (key names, list vs scalar, and the required result file) — a mismatch here compounds the error.
Discriminator
Fine if the column is already numeric on load (dtype shown as int/float) or a documented cleaning step precedes the aggregation; a real violation is aggregating/imputing on an object column, or reporting an extreme without ever displaying the underlying numeric value to confirm it is the true max/min.
Consequence
The reported country/label corresponds to a lexicographic rather than numeric extreme (and mean imputation silently does nothing), so the graded value in the result file does not match the expected answer — 0/1 checks pass.
id a7c06e8c0660 · mined from da-code dacode-di-text-001@s5
raw text (what the judge reads)
### Unverified dtype handling when computing extremes (formatted numbers left as strings)
- **Applies when**: `task` -- the task asks for the argmax/argmin (or mean-imputation) of a quantitative column that may be stored as text with thousands separators, percent/currency symbols, or footnote characters.
- **Pattern**: The attempt loads the file and directly calls `max`/`min`/`idxmax`/`sort_values` (or imputes with a "mean") on a column whose dtype is still `object`, so comparison is lexicographic rather than numeric; no dtype check, no printout of the winning row's value, and no evidence (script or result file) that the numeric conversion happened.
- **Detection procedure**:
  1. From the task, note which column drives the requested extreme/statistic and whether imputation is required (imputation is only possible on a numeric column, so a numeric cast must appear somewhere).
  2. In the scripts, look for an explicit cleaning/casting step (strip separators/symbols, `pd.to_numeric(..., errors=...)`) *before* the aggregation, plus a check of `dtype`/NaN counts; flag if the aggregation runs on raw loaded values or if no script/log exists at all.
  3. Check whether the attempt printed the extreme value alongside the reported label and sanity-checked it against the plausible range and against the top/bottom few rows.
  4. Compare the reported answer to the requested output schema (key names, list vs scalar, and the required result file) — a mismatch here compounds the error.
- **Discriminator**: Fine if the column is already numeric on load (dtype shown as int/float) or a documented cleaning step precedes the aggregation; a real violation is aggregating/imputing on an `object` column, or reporting an extreme without ever displaying the underlying numeric value to confirm it is the true max/min.
- **Consequence**: The reported country/label corresponds to a lexicographic rather than numeric extreme (and mean imputation silently does nothing), so the graded value in the result file does not match the expected answer — 0/1 checks pass.
297Dropping requested intermediate outputs from the saved result filetaskda-code
Applies when
task -- the task asks to compute several quantities (scores, segments, labels) and save "the results including X and Y" to a single output file, and the script writes that file by selecting a subset of columns.
Pattern
The agent computes all the intermediate quantities in memory, then exports only the final label/prediction column plus an ID, silently discarding the other explicitly requested fields (component scores, aggregate score, segment string), and/or copies a schema guessed from a helper/sample file instead of the one the task text demands.
Detection procedure
  1. Read the task statement and list every artifact it says must appear in the output file (each score, each derived segment, each level/label, plus the key column).
  2. Find the line in the script that builds and writes the output (df[[...]], to_csv, etc.) and list the columns actually written, with their names and order.
  3. Compare the two lists; flag any requested item that is computed in the script but not written, or written under a name/format inconsistent with the task wording.
  4. Check whether the agent justified the reduced schema by an external sample/reference file rather than by the task, and whether its final answer describes the file as containing fewer fields than requested.
Discriminator
A real violation is when a required deliverable is computed but excluded (or renamed away) from the saved file; it is not a violation when the omitted columns are genuinely internal scratch variables never requested, or when the task explicitly names a minimal schema and the script matches it exactly.
Consequence
The grader compares the produced file against the expected one and reports the result file as WRONG/MISSING because required columns are absent, even though the underlying values were computed correctly.
id aa75b0e5c6df · mined from da-code dacode-dm-csv-052@s5
raw text (what the judge reads)
### Dropping requested intermediate outputs from the saved result file
- **Applies when**: `task` -- the task asks to compute several quantities (scores, segments, labels) and save "the results including X and Y" to a single output file, and the script writes that file by selecting a subset of columns.
- **Pattern**: The agent computes all the intermediate quantities in memory, then exports only the final label/prediction column plus an ID, silently discarding the other explicitly requested fields (component scores, aggregate score, segment string), and/or copies a schema guessed from a helper/sample file instead of the one the task text demands.
- **Detection procedure**:
  1. Read the task statement and list every artifact it says must appear in the output file (each score, each derived segment, each level/label, plus the key column).
  2. Find the line in the script that builds and writes the output (`df[[...]]`, `to_csv`, etc.) and list the columns actually written, with their names and order.
  3. Compare the two lists; flag any requested item that is computed in the script but not written, or written under a name/format inconsistent with the task wording.
  4. Check whether the agent justified the reduced schema by an external sample/reference file rather than by the task, and whether its final answer describes the file as containing fewer fields than requested.
- **Discriminator**: A real violation is when a required deliverable is computed but excluded (or renamed away) from the saved file; it is *not* a violation when the omitted columns are genuinely internal scratch variables never requested, or when the task explicitly names a minimal schema and the script matches it exactly.
- **Consequence**: The grader compares the produced file against the expected one and reports the result file as WRONG/MISSING because required columns are absent, even though the underlying values were computed correctly.
298Fabricating synthetic data when the real input file isn't foundtaskda-code
Applies when
task -- the task references a provided dataset and the script must locate/load it before computing the requested output.
Pattern
The script guesses a few hard-coded paths and, on failure, silently generates random/simulated data (or a placeholder subset) and proceeds to compute and save "results" from it, without ever asserting that the real file was loaded or that the provided sample/reference format was inspected.
Detection procedure
1) Read the task for the named dataset and any provided sample/reference file. 2) In the scripts, look for fallback branches (if data is None:, except: ...) that synthesize data or invent columns instead of raising an error, and check whether the real file path was actually confirmed (e.g., a directory listing of the data folder). 3) Check whether the reported answer states the source file, row counts, and column names from the real data, or instead reports only structural facts (shape, symmetry, ranges) that would hold for any fabricated data. 4) Confirm the provided sample/format file was read and the output was matched to it.
Discriminator
A real violation is a fallback that produces output from data the agent never verified came from the task's dataset (or values implausible for it, e.g., suspiciously uniform/extreme statistics typical of simulated variables). It's fine if the script tries multiple paths but hard-fails/aborts when none exist, or if it discovers the real file and logs concrete evidence (path, shape, head rows) consistent with the described data.
Consequence
result.csv exists and is well-formed but its numbers come from random data, so every value comparison against the expected file fails (0/1 checks passed).
id e1dea681dab1 · mined from da-code dacode-data-sa-026@s5
raw text (what the judge reads)
### Fabricating synthetic data when the real input file isn't found
- **Applies when**: `task` -- the task references a provided dataset and the script must locate/load it before computing the requested output.
- **Pattern**: The script guesses a few hard-coded paths and, on failure, silently generates random/simulated data (or a placeholder subset) and proceeds to compute and save "results" from it, without ever asserting that the real file was loaded or that the provided sample/reference format was inspected.
- **Detection procedure**: 1) Read the task for the named dataset and any provided sample/reference file. 2) In the scripts, look for fallback branches (`if data is None:`, `except: ...`) that synthesize data or invent columns instead of raising an error, and check whether the real file path was actually confirmed (e.g., a directory listing of the data folder). 3) Check whether the reported answer states the source file, row counts, and column names from the real data, or instead reports only structural facts (shape, symmetry, ranges) that would hold for any fabricated data. 4) Confirm the provided sample/format file was read and the output was matched to it.
- **Discriminator**: A real violation is a fallback that produces output from data the agent never verified came from the task's dataset (or values implausible for it, e.g., suspiciously uniform/extreme statistics typical of simulated variables). It's fine if the script tries multiple paths but hard-fails/aborts when none exist, or if it discovers the real file and logs concrete evidence (path, shape, head rows) consistent with the described data.
- **Consequence**: `result.csv` exists and is well-formed but its numbers come from random data, so every value comparison against the expected file fails (0/1 checks passed).
299Selected hyperparameter sits at the edge of the searched grid and contradicts known structure in the datataskda-code
Applies when
task -- the script must choose a free structural hyperparameter (number of clusters/components/topics, k in kNN, tree depth) by sweeping a fixed candidate range and picking the best score, while the dataset itself contains a label/target column or documented grouping that implies a natural number of groups.
Pattern
The attempt sweeps a hard-coded range, takes the arg-max of a single internal metric, and reports a value that lands on the first or last candidate of that range (so the true optimum was never bracketed), without extending the sweep, without cross-checking a second criterion (elbow/BIC/stability, agreement with the held-out label column via ARI/NMI), and without sanity-checking that the chosen structure is plausible given the data's documented composition (e.g., row count vs. README, known class count, cluster size balance).
Detection procedure
  1. Read the task and README for any documented structure (a target/label column, stated number of categories, stated dataset size) and note the implied plausible range for the hyperparameter.
  2. In the script, locate the candidate grid and the selection rule; check whether the reported winner is at a grid endpoint and whether the sweep was re-run/extended beyond it.
  3. Check whether any second, independent validation was performed (alternative internal index agreeing, comparison to the held-out label column, stability across seeds/subsamples) and whether the loaded data's shape was compared against the documented shape.
  4. Compare the answer's reported cluster/group count and size distribution to the documented structure; flag if it exceeds the plausible count, was picked at a boundary, or the row count silently differs from the documentation.
Discriminator
A real violation is an endpoint pick with a single weak metric (e.g., a low silhouette that is still monotonically improving) and no corroboration; it is fine if the winner is interior to the grid, or if the agent explicitly extended the sweep past the endpoint and showed the score turning over, or corroborated the choice with a second criterion/known label count and explained the discrepancy.
Consequence
The saved result file encodes an over-partitioned (or under-partitioned) structure that fails the grader's check on the expected number/quality of groups — the output file is marked WRONG even though its column names and shape look correct.
id d675f16c4c67 · mined from da-code dacode-ml-cluster-010@s5
raw text (what the judge reads)
### Selected hyperparameter sits at the edge of the searched grid and contradicts known structure in the data
- **Applies when**: `task` -- the script must choose a free structural hyperparameter (number of clusters/components/topics, k in kNN, tree depth) by sweeping a fixed candidate range and picking the best score, while the dataset itself contains a label/target column or documented grouping that implies a natural number of groups.
- **Pattern**: The attempt sweeps a hard-coded range, takes the arg-max of a single internal metric, and reports a value that lands on the first or last candidate of that range (so the true optimum was never bracketed), without extending the sweep, without cross-checking a second criterion (elbow/BIC/stability, agreement with the held-out label column via ARI/NMI), and without sanity-checking that the chosen structure is plausible given the data's documented composition (e.g., row count vs. README, known class count, cluster size balance).
- **Detection procedure**:
  1. Read the task and README for any documented structure (a target/label column, stated number of categories, stated dataset size) and note the implied plausible range for the hyperparameter.
  2. In the script, locate the candidate grid and the selection rule; check whether the reported winner is at a grid endpoint and whether the sweep was re-run/extended beyond it.
  3. Check whether any second, independent validation was performed (alternative internal index agreeing, comparison to the held-out label column, stability across seeds/subsamples) and whether the loaded data's shape was compared against the documented shape.
  4. Compare the answer's reported cluster/group count and size distribution to the documented structure; flag if it exceeds the plausible count, was picked at a boundary, or the row count silently differs from the documentation.
- **Discriminator**: A real violation is an endpoint pick with a single weak metric (e.g., a low silhouette that is still monotonically improving) and no corroboration; it is fine if the winner is interior to the grid, or if the agent explicitly extended the sweep past the endpoint and showed the score turning over, or corroborated the choice with a second criterion/known label count and explained the discrepancy.
- **Consequence**: The saved result file encodes an over-partitioned (or under-partitioned) structure that fails the grader's check on the expected number/quality of groups — the output file is marked WRONG even though its column names and shape look correct.
300Required output artifact never written or verified on disktaskda-code
Applies when
task -- the task explicitly asks for a deliverable file with a specified name, columns, and row semantics, and the agent's work is presented as a prose summary.
Pattern
The attempt performs the analysis and narrates results (counts, metrics, cluster sizes) but never contains code that writes the named file with the exact required column names/schema, or writes it to a different path/name/format, and never re-reads it to confirm it exists and has the expected shape.
Detection procedure
  1. From the task statement, list the exact required artifacts: filename, column names (including index/enumeration conventions), and expected number of rows.
  2. Search the scripts for an explicit write call producing that exact filename in the working directory, and check the DataFrame's columns are renamed to the required names before writing (and that the index isn't silently added or the header dropped).
  3. Check that after writing, the code (or answer) reports a verification read-back: file exists, shape, head, and that row count matches the number of analysed units.
  4. Compare the final answer: does it point to the artifact and its verified contents, or does it only present narrative statistics?
Discriminator
A real violation is when no code path guarantees the exact required file/schema, or the answer's only evidence is prose. It is fine if the file is written under the required name with required headers and the answer cites a verified shape/head, even if the narrative is verbose or the chosen cluster count is debatable.
Consequence
The grader looking for the named result file marks it WRONG/MISSING and the task scores 0 regardless of how sound the modelling was.
id e048a99628c5 · mined from da-code dacode-ml-cluster-019@s5
raw text (what the judge reads)
### Required output artifact never written or verified on disk
- **Applies when**: `task` -- the task explicitly asks for a deliverable file with a specified name, columns, and row semantics, and the agent's work is presented as a prose summary.
- **Pattern**: The attempt performs the analysis and narrates results (counts, metrics, cluster sizes) but never contains code that writes the named file with the exact required column names/schema, or writes it to a different path/name/format, and never re-reads it to confirm it exists and has the expected shape.
- **Detection procedure**:
  1. From the task statement, list the exact required artifacts: filename, column names (including index/enumeration conventions), and expected number of rows.
  2. Search the scripts for an explicit write call producing that exact filename in the working directory, and check the DataFrame's columns are renamed to the required names before writing (and that the index isn't silently added or the header dropped).
  3. Check that after writing, the code (or answer) reports a verification read-back: file exists, shape, head, and that row count matches the number of analysed units.
  4. Compare the final answer: does it point to the artifact and its verified contents, or does it only present narrative statistics?
- **Discriminator**: A real violation is when no code path guarantees the exact required file/schema, or the answer's only evidence is prose. It is fine if the file is written under the required name with required headers and the answer cites a verified shape/head, even if the narrative is verbose or the chosen cluster count is debatable.
- **Consequence**: The grader looking for the named result file marks it WRONG/MISSING and the task scores 0 regardless of how sound the modelling was.
301Fabricating input data from memory instead of loading the provided filestaskda-code
Applies when
task -- The task ships a data directory/README and the script must compute a statistic from that supplied data.
Pattern
The script hardcodes a small table of numbers typed from the agent's prior knowledge (or a guessed schema) rather than reading the actual provided files, so every downstream statistic is computed on invented values and never reconciled with the real records.
Detection procedure
1. From the task/README, list the data artifacts that are supposed to exist and the granularity they imply (e.g., per-period counts vs. per-record rows). 2. Search the scripts for any file-reading call (read_csv, load, open, glob of the data dir); if the only data source is a literal dict/array/DataFrame constructor, flag immediately. 3. Check whether any step verifies the loaded data against the source (row/column counts, totals, printed head) — absence of such a check strengthens the flag. 4. Confirm the reported numbers could only come from the hardcoded literals, not from the shipped files.
Discriminator
Legitimate cases hardcode only constants that the task itself states (thresholds, seeds, split sizes, known reference values) while still loading the observations from disk; a violation is when the observations being analyzed exist nowhere but in the script's literals, or the schema/units were assumed rather than inspected.
Consequence
The confidence interval / metric is derived from wrong inputs, so the written output file fails the expected-value check even though the code runs cleanly and the reported interval looks plausible.
id d67682e9c5c4 · mined from da-code dacode-data-sa-031@s5
raw text (what the judge reads)
### Fabricating input data from memory instead of loading the provided files
- **Applies when**: `task` -- The task ships a data directory/README and the script must compute a statistic from that supplied data.
- **Pattern**: The script hardcodes a small table of numbers typed from the agent's prior knowledge (or a guessed schema) rather than reading the actual provided files, so every downstream statistic is computed on invented values and never reconciled with the real records.
- **Detection procedure**: 1. From the task/README, list the data artifacts that are supposed to exist and the granularity they imply (e.g., per-period counts vs. per-record rows). 2. Search the scripts for any file-reading call (`read_csv`, `load`, `open`, glob of the data dir); if the only data source is a literal dict/array/DataFrame constructor, flag immediately. 3. Check whether any step verifies the loaded data against the source (row/column counts, totals, printed head) — absence of such a check strengthens the flag. 4. Confirm the reported numbers could only come from the hardcoded literals, not from the shipped files.
- **Discriminator**: Legitimate cases hardcode only constants that the task itself states (thresholds, seeds, split sizes, known reference values) while still loading the observations from disk; a violation is when the *observations being analyzed* exist nowhere but in the script's literals, or the schema/units were assumed rather than inspected.
- **Consequence**: The confidence interval / metric is derived from wrong inputs, so the written output file fails the expected-value check even though the code runs cleanly and the reported interval looks plausible.
302Mis-mapping model probability columns to the required class-labeled output columnstaskda-code
Applies when
task -- the deliverable is a per-row multi-class probability file whose columns must be named/ordered by class label, and the script encodes the target (e.g., with a label encoder or factorize) before calling predict_proba.
Pattern
The script hard-codes an assumed index→class correspondence (e.g., "class 0 is the last label") when slicing the probability matrix, instead of deriving the mapping from the fitted model's classes_ / encoder's classes_, so the probability columns get attached to the wrong class names.
Detection procedure
  1. In the task/README, note the exact required output column names and their class meanings.
  2. In the script, find where the target is encoded and what ordering that encoder produces (label encoders sort labels alphabetically/numerically; predict_proba columns follow model.classes_ in that same sorted order).
  3. Check the submission-construction code: is each output column indexed via a lookup built from classes_ (e.g., proba[:, list(classes_).index('X')] or pd.DataFrame(proba, columns=le.classes_)), or via literal indices with a comment asserting an ordering?
  4. Cross-check the literal indices against the sorted class order and against the answer's own numbers: compare each column's mean probability to the class frequencies printed/known from training; a rare class carrying a large average probability (or the majority class carrying a tiny one) confirms the swap.
Discriminator
Hard-coded indices are fine if they provably coincide with the sorted classes_ order (or the script prints/asserts the mapping); it is a violation when the assumed order contradicts the encoder's sorted order, or when no verification of the mapping exists and the resulting column means are inconsistent with the training class distribution.
Consequence
The file has valid shape, ids and normalized-able rows, so format checks pass, but probabilities are assigned to the wrong labels, giving a log loss far worse than a constant-prior baseline and a failed correctness check.
id 9768bc71e9d2 · mined from da-code dacode-ml-competition-005@s6
raw text (what the judge reads)
### Mis-mapping model probability columns to the required class-labeled output columns
- **Applies when**: `task` -- the deliverable is a per-row multi-class probability file whose columns must be named/ordered by class label, and the script encodes the target (e.g., with a label encoder or `factorize`) before calling `predict_proba`.
- **Pattern**: The script hard-codes an assumed index→class correspondence (e.g., "class 0 is the last label") when slicing the probability matrix, instead of deriving the mapping from the fitted model's `classes_` / encoder's `classes_`, so the probability columns get attached to the wrong class names.
- **Detection procedure**:
  1. In the task/README, note the exact required output column names and their class meanings.
  2. In the script, find where the target is encoded and what ordering that encoder produces (label encoders sort labels alphabetically/numerically; `predict_proba` columns follow `model.classes_` in that same sorted order).
  3. Check the submission-construction code: is each output column indexed via a lookup built from `classes_` (e.g., `proba[:, list(classes_).index('X')]` or `pd.DataFrame(proba, columns=le.classes_)`), or via literal indices with a comment asserting an ordering?
  4. Cross-check the literal indices against the sorted class order and against the answer's own numbers: compare each column's mean probability to the class frequencies printed/known from training; a rare class carrying a large average probability (or the majority class carrying a tiny one) confirms the swap.
- **Discriminator**: Hard-coded indices are fine if they provably coincide with the sorted `classes_` order (or the script prints/asserts the mapping); it is a violation when the assumed order contradicts the encoder's sorted order, or when no verification of the mapping exists and the resulting column means are inconsistent with the training class distribution.
- **Consequence**: The file has valid shape, ids and normalized-able rows, so format checks pass, but probabilities are assigned to the wrong labels, giving a log loss far worse than a constant-prior baseline and a failed correctness check.
303Output file schema/alignment not verified against the literal deliverable spectaskda-code
Applies when
task -- the task names an output file and the exact column name(s) it must contain, and the script builds that file from model predictions.
Pattern
The script writes the file with extra or renamed columns (e.g., an added identifier or index column), a different row count/order than the input rows it must score, or applies an unrequested transformation (truncation to int, clipping, rescaling) before writing — and the answer never checks the written file against the spec.
Detection procedure
  1. Re-read the task statement and list the exact required filename, required column name(s), implied row count, and implied row order (usually one row per input record, in input order).
  2. In the script, find the DataFrame construction and the to_csv call; note every column included, any index writing, any dtype cast/rounding, and which input frame supplies the rows.
  3. Compare item-by-item with step 1; also confirm the frame used for prediction was not filtered, deduplicated, sorted, or re-indexed relative to the input file.
  4. Check whether the answer/report includes a read-back sanity check of the saved file (shape, column list, first rows, value range) rather than only in-memory statistics.
Discriminator
A violation is adding/removing/renaming columns, changing row count or order, or altering predicted values in a way the task did not request. It is not a violation if extra columns are explicitly permitted by the task, or if the transformation (e.g., non-negativity for a count target) is clearly implied by the target's definition and the required column is still present, correctly named, and row-aligned.
Consequence
The grader loads the file, fails to find the expected column/shape or misaligns rows against ground truth, and scores the deliverable as wrong/missing even though the model itself may be reasonable.
id 26a5bb04f6fb · mined from da-code dacode-ml-regression-008@s6
raw text (what the judge reads)
### Output file schema/alignment not verified against the literal deliverable spec
- **Applies when**: `task` -- the task names an output file and the exact column name(s) it must contain, and the script builds that file from model predictions.
- **Pattern**: The script writes the file with extra or renamed columns (e.g., an added identifier or index column), a different row count/order than the input rows it must score, or applies an unrequested transformation (truncation to int, clipping, rescaling) before writing — and the answer never checks the written file against the spec.
- **Detection procedure**:
  1. Re-read the task statement and list the exact required filename, required column name(s), implied row count, and implied row order (usually one row per input record, in input order).
  2. In the script, find the DataFrame construction and the `to_csv` call; note every column included, any index writing, any dtype cast/rounding, and which input frame supplies the rows.
  3. Compare item-by-item with step 1; also confirm the frame used for prediction was not filtered, deduplicated, sorted, or re-indexed relative to the input file.
  4. Check whether the answer/report includes a read-back sanity check of the saved file (shape, column list, first rows, value range) rather than only in-memory statistics.
- **Discriminator**: A violation is adding/removing/renaming columns, changing row count or order, or altering predicted values in a way the task did not request. It is *not* a violation if extra columns are explicitly permitted by the task, or if the transformation (e.g., non-negativity for a count target) is clearly implied by the target's definition and the required column is still present, correctly named, and row-aligned.
- **Consequence**: The grader loads the file, fails to find the expected column/shape or misaligns rows against ground truth, and scores the deliverable as wrong/missing even though the model itself may be reasonable.
304Uses the full raw dataset with a default two-sided parametric test instead of scoping the population and test form to the hypothesistaskda-code
Applies when
task -- the task asks for a p-value and an accept/reject decision from a comparison of two groups, and the scripts load the raw files and immediately run a stock test on all rows.
Pattern
The attempt never establishes which rows constitute the population the hypothesis refers to (competition/category, date range, or other stated or strongly implied scope), never states the alternative hypothesis direction, and never checks whether the variable's distribution justifies the chosen test; it just calls a default two-sided parametric routine on every row and gets an astronomically small p-value.
Detection procedure
  1. Read the task and README for any qualifier that narrows the population (subset of records, time window, category of event) or that implies a directional claim ("more than", "greater", "higher").
  2. In the scripts, check for an explicit filtering step matching that scope and an explicit choice of tail/alternative; check that group sample sizes and means are printed after filtering, not before.
  3. Check whether the distribution of the compared quantity (skewed, discrete counts, unequal variances) was inspected and whether the test chosen (parametric vs. rank-based, equal-variance vs. Welch) was justified rather than left at defaults.
  4. Compare the reported p-value magnitude to the plausibility of the effect: a p-value many orders of magnitude below any conventional threshold on hundreds of thousands of rows is a signal that an unrestricted, oversized sample was used.
Discriminator
A real violation is when the task/README language implies a narrower population or a one-sided alternative and the script demonstrably uses all rows and a default two-sided call; it is fine if the task genuinely asks about all records and the script documents that the full set is the intended population and that the test assumptions were checked.
Consequence
The p-value is computed on the wrong sample with the wrong test form, so the numeric p_val (and possibly the reject/fail-to-reject string) does not match the expected value and the result file fails the grader's check.
id 42be8f0b88c9 · mined from da-code dacode-data-sa-001@s6
raw text (what the judge reads)
### Uses the full raw dataset with a default two-sided parametric test instead of scoping the population and test form to the hypothesis
- **Applies when**: `task` -- the task asks for a p-value and an accept/reject decision from a comparison of two groups, and the scripts load the raw files and immediately run a stock test on all rows.
- **Pattern**: The attempt never establishes which rows constitute the population the hypothesis refers to (competition/category, date range, or other stated or strongly implied scope), never states the alternative hypothesis direction, and never checks whether the variable's distribution justifies the chosen test; it just calls a default two-sided parametric routine on every row and gets an astronomically small p-value.
- **Detection procedure**:
  1. Read the task and README for any qualifier that narrows the population (subset of records, time window, category of event) or that implies a directional claim ("more than", "greater", "higher").
  2. In the scripts, check for an explicit filtering step matching that scope and an explicit choice of tail/alternative; check that group sample sizes and means are printed after filtering, not before.
  3. Check whether the distribution of the compared quantity (skewed, discrete counts, unequal variances) was inspected and whether the test chosen (parametric vs. rank-based, equal-variance vs. Welch) was justified rather than left at defaults.
  4. Compare the reported p-value magnitude to the plausibility of the effect: a p-value many orders of magnitude below any conventional threshold on hundreds of thousands of rows is a signal that an unrestricted, oversized sample was used.
- **Discriminator**: A real violation is when the task/README language implies a narrower population or a one-sided alternative and the script demonstrably uses all rows and a default two-sided call; it is fine if the task genuinely asks about all records and the script documents that the full set is the intended population and that the test assumptions were checked.
- **Consequence**: The p-value is computed on the wrong sample with the wrong test form, so the numeric `p_val` (and possibly the reject/fail-to-reject string) does not match the expected value and the result file fails the grader's check.
305Template output file is assumed rather than read and validatedtaskda-code
Applies when
task -- the task says results must be written to an output file that follows the "exact structure/formatting" of a provided sample/template file.
Pattern
The script never loads or prints the sample file; instead the agent hard-codes column names, column order, row order, rounding, and header text from memory or guesswork, then declares the output "matches the sample" without any comparison.
Detection procedure
  1. Read the task and note that a sample/template artifact is supplied as the formatting contract.
  2. Search the scripts for any read of that sample file (e.g., loading it, printing its header/rows) and any assertion comparing the produced frame's column names, column order, dtypes/rounding, row count, and row ordering to it.
  3. If absent, check whether the output construction relies on literals the agent invented (hand-typed header names, a hard-coded category/row ordering list, arbitrary rounding) with no provenance from the sample.
  4. Check the final answer for evidence of an explicit diff/verification step against the sample rather than a bare claim of conformance.
Discriminator
A real violation is when no code path ever inspects the template and the format is reconstructed from assumption; it is fine if the script loads the sample, derives (or asserts) headers/order/precision from it, even if the final formatting code then writes literals that were verified to match.
Consequence
The file-level grader comparing the output to the expected artifact reports WRONG/MISSING because of mismatched column names, column/row order, or numeric formatting, even if the underlying aggregation logic was reasonable.
id 9b6f6502ffcd · mined from da-code dacode-dm-csv-011@s6
raw text (what the judge reads)
### Template output file is assumed rather than read and validated
- **Applies when**: `task` -- the task says results must be written to an output file that follows the "exact structure/formatting" of a provided sample/template file.
- **Pattern**: The script never loads or prints the sample file; instead the agent hard-codes column names, column order, row order, rounding, and header text from memory or guesswork, then declares the output "matches the sample" without any comparison.
- **Detection procedure**:
  1. Read the task and note that a sample/template artifact is supplied as the formatting contract.
  2. Search the scripts for any read of that sample file (e.g., loading it, printing its header/rows) and any assertion comparing the produced frame's column names, column order, dtypes/rounding, row count, and row ordering to it.
  3. If absent, check whether the output construction relies on literals the agent invented (hand-typed header names, a hard-coded category/row ordering list, arbitrary rounding) with no provenance from the sample.
  4. Check the final answer for evidence of an explicit diff/verification step against the sample rather than a bare claim of conformance.
- **Discriminator**: A real violation is when no code path ever inspects the template and the format is reconstructed from assumption; it is fine if the script loads the sample, derives (or asserts) headers/order/precision from it, even if the final formatting code then writes literals that were verified to match.
- **Consequence**: The file-level grader comparing the output to the expected artifact reports WRONG/MISSING because of mismatched column names, column/row order, or numeric formatting, even if the underlying aggregation logic was reasonable.
306Requested output artifact never written with the exact required schemataskda-code
Applies when
task -- the task specifies a deliverable file with an explicit name and column/schema layout, and the agent's scripts/answer are the only evidence that it was produced.
Pattern
The attempt performs the analysis and reports a narrative summary (method, chosen hyperparameters, group profiles) but never contains code that writes the named file, or writes it with different columns, extra/missing index, renamed labels, or fewer rows than the input records — treating the prose report as the answer.
Detection procedure
  1. Extract from the task the exact required filename, required column names (including any index-style naming pattern), and the expected number of rows.
  2. Search the scripts for an explicit save call to that filename and check the DataFrame passed to it: column names constructed exactly as specified, one row per input record, no index column written.
  3. Check the final answer for a statement/verification that the file exists with that shape (e.g., printed shape and head), not merely a description of features and cluster profiles.
  4. Flag if no such write exists, the schema differs, or row count differs from the analyzed record count (e.g., rows dropped during cleaning without being restored/labeled).
Discriminator
A real violation is missing/misnamed/mis-columned output or row-count mismatch; a look-alike that is fine is a correctly written file plus a summary, or a file whose values differ from a reviewer's own preferred method (algorithm/cluster-count choices are legitimately free when the task doesn't constrain them).
Consequence
The grader looks for the named file with the specified columns and rows; it reports the expected file as WRONG/MISSING and scores 0 regardless of how sound the underlying analysis was.
id 2a1ed64d7fbc · mined from da-code dacode-ml-cluster-014@s6
raw text (what the judge reads)
### Requested output artifact never written with the exact required schema
- **Applies when**: `task` -- the task specifies a deliverable file with an explicit name and column/schema layout, and the agent's scripts/answer are the only evidence that it was produced.
- **Pattern**: The attempt performs the analysis and reports a narrative summary (method, chosen hyperparameters, group profiles) but never contains code that writes the named file, or writes it with different columns, extra/missing index, renamed labels, or fewer rows than the input records — treating the prose report as the answer.
- **Detection procedure**:
  1. Extract from the task the exact required filename, required column names (including any index-style naming pattern), and the expected number of rows.
  2. Search the scripts for an explicit save call to that filename and check the DataFrame passed to it: column names constructed exactly as specified, one row per input record, no index column written.
  3. Check the final answer for a statement/verification that the file exists with that shape (e.g., printed `shape` and `head`), not merely a description of features and cluster profiles.
  4. Flag if no such write exists, the schema differs, or row count differs from the analyzed record count (e.g., rows dropped during cleaning without being restored/labeled).
- **Discriminator**: A real violation is missing/misnamed/mis-columned output or row-count mismatch; a look-alike that is fine is a correctly written file plus a summary, or a file whose values differ from a reviewer's own preferred method (algorithm/cluster-count choices are legitimately free when the task doesn't constrain them).
- **Consequence**: The grader looks for the named file with the specified columns and rows; it reports the expected file as WRONG/MISSING and scores 0 regardless of how sound the underlying analysis was.
307Submission artifact never validated against the provided templatetaskda-code
Applies when
task -- the task supplies a sample/expected output file (or an explicit output spec) and the script writes a prediction/result file to a specified location.
Pattern
The script builds the output file from hard-coded assumptions (guessed column names, id source, dtype, path) without ever loading or comparing to the provided template, and often applies unrequested post-processing (rounding, integer casting, clipping) to the predicted values; the final answer summarizes modeling details but shows no check that the file matches the required shape, header, row count, ordering, and location.
Detection procedure
  1. Read the task/README for the named output artifact, its required location, and the template file it must match.
  2. Search the script for a read of that template and for any explicit comparison of the written file's columns, row count, id set/order, and value dtype against it; also confirm the output path is exactly where the task asks.
  3. Check whether predicted values are transformed (rounded/cast/clipped) and whether the task or template actually requires that; continuous targets scored by a regression metric should normally stay continuous.
  4. Check the final answer for a post-write verification (re-read the file, print head/shape/dtypes) rather than only training/validation statistics.
Discriminator
A real violation is an output file whose columns/ids/dtype/path were assumed rather than derived and verified against the template (or values altered beyond spec); it is fine if the script loads the template (or task-stated spec), constructs the file from it, and re-reads/asserts the written file even if the modeling itself is simple.
Consequence
The grader marks the expected output file as WRONG/MISSING — wrong path, mismatched header/ids/row count, or degraded score from unnecessary discretization — regardless of how good the reported validation metrics were.
id 8382b18e16ff · mined from da-code dacode-ml-competition-009@s6
raw text (what the judge reads)
### Submission artifact never validated against the provided template
- **Applies when**: `task` -- the task supplies a sample/expected output file (or an explicit output spec) and the script writes a prediction/result file to a specified location.
- **Pattern**: The script builds the output file from hard-coded assumptions (guessed column names, id source, dtype, path) without ever loading or comparing to the provided template, and often applies unrequested post-processing (rounding, integer casting, clipping) to the predicted values; the final answer summarizes modeling details but shows no check that the file matches the required shape, header, row count, ordering, and location.
- **Detection procedure**:
  1. Read the task/README for the named output artifact, its required location, and the template file it must match.
  2. Search the script for a read of that template and for any explicit comparison of the written file's columns, row count, id set/order, and value dtype against it; also confirm the output path is exactly where the task asks.
  3. Check whether predicted values are transformed (rounded/cast/clipped) and whether the task or template actually requires that; continuous targets scored by a regression metric should normally stay continuous.
  4. Check the final answer for a post-write verification (re-read the file, print head/shape/dtypes) rather than only training/validation statistics.
- **Discriminator**: A real violation is an output file whose columns/ids/dtype/path were assumed rather than derived and verified against the template (or values altered beyond spec); it is fine if the script loads the template (or task-stated spec), constructs the file from it, and re-reads/asserts the written file even if the modeling itself is simple.
- **Consequence**: The grader marks the expected output file as WRONG/MISSING — wrong path, mismatched header/ids/row count, or degraded score from unnecessary discretization — regardless of how good the reported validation metrics were.
308Dropping input rows so the output covers fewer records than the datasettaskda-code
Applies when
task -- the task asks for a per-record output file (labels, predictions, scores) covering a given dataset, and the script does missing-value handling before producing it.
Pattern
The script silently removes rows with NaNs (e.g. dropna()) or subsets columns, then writes an output file with far fewer rows than the input, without imputing or otherwise preserving every record; the answer even states the shrunken count as if it were acceptable.
Detection procedure
  1. Read the task and note whether the requested output is expected to have one row per input record (no filtering instruction is given).
  2. In the script, find every row-removing operation (dropna, boolean masks, joins) between loading and writing the output.
  3. Compare the row count written (or reported in the answer) with the raw input row count; also check that column count/names match the requested schema exactly.
  4. Flag if rows were lost with no instruction to filter, and no imputation was attempted.
Discriminator
A real violation is unrequested row loss (or a schema mismatch) in the deliverable; it is fine if the task explicitly says to filter/deduplicate, or if dropped rows are restored with labels/NaNs so the output still aligns row-for-row with the source.
Consequence
The saved file cannot be aligned or compared row-wise with the reference output, so the file-level check fails regardless of clustering/model quality.
id c6789cd4a091 · mined from da-code dacode-ml-cluster-009@s6
raw text (what the judge reads)
### Dropping input rows so the output covers fewer records than the dataset
- **Applies when**: `task` -- the task asks for a per-record output file (labels, predictions, scores) covering a given dataset, and the script does missing-value handling before producing it.
- **Pattern**: The script silently removes rows with NaNs (e.g. `dropna()`) or subsets columns, then writes an output file with far fewer rows than the input, without imputing or otherwise preserving every record; the answer even states the shrunken count as if it were acceptable.
- **Detection procedure**:
  1. Read the task and note whether the requested output is expected to have one row per input record (no filtering instruction is given).
  2. In the script, find every row-removing operation (`dropna`, boolean masks, joins) between loading and writing the output.
  3. Compare the row count written (or reported in the answer) with the raw input row count; also check that column count/names match the requested schema exactly.
  4. Flag if rows were lost with no instruction to filter, and no imputation was attempted.
- **Discriminator**: A real violation is unrequested row loss (or a schema mismatch) in the deliverable; it is fine if the task explicitly says to filter/deduplicate, or if dropped rows are restored with labels/NaNs so the output still aligns row-for-row with the source.
- **Consequence**: The saved file cannot be aligned or compared row-wise with the reference output, so the file-level check fails regardless of clustering/model quality.
309Output file schema not written exactly as specified (stray index column / extra or mismatched columns)taskda-code
Applies when
task -- the task dictates a deliverable file with an exact set of column names (e.g., Feature_i ... plus a label/prediction column), and a script writes that file with to_csv/to_excel from a DataFrame that carries an index or extra engineered columns.
Pattern
The agent builds a DataFrame indexed by an identifier (or with helper/raw columns), renames only some columns to the required names, and saves with the default index=True, so the written file has an unnamed leading column, a different column count/order, or values that are not the vectors actually clustered/modeled — silently violating the required schema.
Detection procedure
  1. Read the task statement and write down the exact required column names, their order, and whether an index/ID column is permitted.
  2. In the writing script, inspect the DataFrame just before the save call: list its index (is it a meaningful key set by index_col/set_index?) and its final column list after any renaming.
  3. Check the save call arguments: is index=False used when no index column is allowed? Do the renamed columns cover every column present (no leftovers, no missing Feature_i)? Is numbering/ordering consistent with what was described?
  4. Confirm the saved values correspond to the feature representation the task implies (the vectors actually used to produce the labels), and that row count equals the number of clustered/predicted entities.
Discriminator
A real violation is when the file as written would contain columns beyond (or differently named from) the specified schema — e.g., an unnamed index column, leftover un-renamed features, or missing the label column. It is not a violation if the index is explicitly reset/dropped, or if the task allows an identifier column and it is named as required.
Consequence
The deliverable file fails automatic schema/column checks (header mismatch or extra column), so the grader marks the expected file WRONG/MISSING regardless of how sound the modeling was.
id 1e6a5f5e9767 · mined from da-code dacode-ml-cluster-016@s6
raw text (what the judge reads)
### Output file schema not written exactly as specified (stray index column / extra or mismatched columns)
- **Applies when**: `task` -- the task dictates a deliverable file with an exact set of column names (e.g., `Feature_i` ... plus a label/prediction column), and a script writes that file with `to_csv`/`to_excel` from a DataFrame that carries an index or extra engineered columns.
- **Pattern**: The agent builds a DataFrame indexed by an identifier (or with helper/raw columns), renames only some columns to the required names, and saves with the default `index=True`, so the written file has an unnamed leading column, a different column count/order, or values that are not the vectors actually clustered/modeled — silently violating the required schema.
- **Detection procedure**:
  1. Read the task statement and write down the exact required column names, their order, and whether an index/ID column is permitted.
  2. In the writing script, inspect the DataFrame just before the save call: list its index (is it a meaningful key set by `index_col`/`set_index`?) and its final column list after any renaming.
  3. Check the save call arguments: is `index=False` used when no index column is allowed? Do the renamed columns cover every column present (no leftovers, no missing `Feature_i`)? Is numbering/ordering consistent with what was described?
  4. Confirm the saved values correspond to the feature representation the task implies (the vectors actually used to produce the labels), and that row count equals the number of clustered/predicted entities.
- **Discriminator**: A real violation is when the file as written would contain columns beyond (or differently named from) the specified schema — e.g., an unnamed index column, leftover un-renamed features, or missing the label column. It is *not* a violation if the index is explicitly reset/dropped, or if the task allows an identifier column and it is named as required.
- **Consequence**: The deliverable file fails automatic schema/column checks (header mismatch or extra column), so the grader marks the expected file WRONG/MISSING regardless of how sound the modeling was.
310Predictions produced for a self-made split instead of the provided evaluation filetaskda-code
Applies when
task -- The task supplies a separate held-out input file (e.g., a test set) and asks for a prediction file with a specified column name and one row per held-out record.
Pattern
The agent loads only the historical/training table, does its own random or percentage split, trains on one part and writes predictions for the other part, never reading the provided evaluation file; the output therefore has a row count/ordering/index unrelated to the required rows (and a random split may also leak temporally adjacent records into training).
Detection procedure
  1. From the task, note the required output file, required column name(s), and the exact number/order of rows implied by the provided evaluation input.
  2. In the scripts, check that the evaluation input file is actually read and that the prediction array passed to the output writer is indexed by that file's rows (not by a train_test_split / iloc slice of the training table).
  3. Compare the reported prediction count (and any key/timestamp column) in the answer against the row count of the provided evaluation file; also confirm the output header spells the requested column exactly.
  4. Confirm the feature construction for the evaluation rows uses only columns available in that file (same transformer/imputer fitted on training data), so the model isn't relying on features absent at prediction time.
Discriminator
A real violation is when the output row count/keys cannot be traced to the provided evaluation file (e.g., it equals a 20% slice of the training set), or the evaluation file is never opened. It is fine if the agent additionally makes an internal validation split for model selection, as long as the final written predictions are generated from the provided evaluation file's rows in its original order.
Consequence
The grader cannot align predictions to ground-truth rows — the file is scored as wrong/missing regardless of model quality, and reported metrics describe an unrelated subset.
id 5125c07326a7 · mined from da-code dacode-ml-regression-002@s6
raw text (what the judge reads)
### Predictions produced for a self-made split instead of the provided evaluation file
- **Applies when**: `task` -- The task supplies a separate held-out input file (e.g., a test set) and asks for a prediction file with a specified column name and one row per held-out record.
- **Pattern**: The agent loads only the historical/training table, does its own random or percentage split, trains on one part and writes predictions for the other part, never reading the provided evaluation file; the output therefore has a row count/ordering/index unrelated to the required rows (and a random split may also leak temporally adjacent records into training).
- **Detection procedure**:
  1. From the task, note the required output file, required column name(s), and the exact number/order of rows implied by the provided evaluation input.
  2. In the scripts, check that the evaluation input file is actually read and that the prediction array passed to the output writer is indexed by that file's rows (not by a `train_test_split` / iloc slice of the training table).
  3. Compare the reported prediction count (and any key/timestamp column) in the answer against the row count of the provided evaluation file; also confirm the output header spells the requested column exactly.
  4. Confirm the feature construction for the evaluation rows uses only columns available in that file (same transformer/imputer fitted on training data), so the model isn't relying on features absent at prediction time.
- **Discriminator**: A real violation is when the output row count/keys cannot be traced to the provided evaluation file (e.g., it equals a 20% slice of the training set), or the evaluation file is never opened. It is fine if the agent additionally makes an internal validation split for model selection, as long as the final written predictions are generated from the provided evaluation file's rows in its original order.
- **Consequence**: The grader cannot align predictions to ground-truth rows — the file is scored as wrong/missing regardless of model quality, and reported metrics describe an unrelated subset.
311Missing machine-checkable output artifacts (only the human-facing rendering is produced)taskda-code
Applies when
task -- the task asks for a deliverable produced by a spec file/config (e.g., a chart, table, or report) and grading is done on the underlying data/spec artifacts, not just the rendered image or prose summary.
Pattern
The agent writes only the visual/pretty output (an image file) plus a narrative summary, never persisting the plotted series/values or the resolved spec in the structured side-files that the environment's spec/convention implies (e.g., a JSON of the plot definition and an array of the plotted numbers), and no script is left behind to regenerate them.
Detection procedure
  1. Read the task and the referenced spec/config file for every named or implied output artifact (file names, extensions, serialization formats), including any conventional companions to the main deliverable.
  2. Inspect the scripts/commands for explicit write calls for each artifact (e.g., savefig, json.dump, np.save, to_csv); list which artifacts have no corresponding write.
  3. Check the answer: does it claim success while only describing one artifact (image, screenshot, prose stats), and does it enumerate the data actually serialized?
  4. Flag if any required/implied artifact lacks both a writing step and a saved reproducible script.
Discriminator
A real violation is when an expected artifact is never written to disk (or is written under a different name/extension/location than specified). It is not a violation if the artifact exists but its content is merely debatable, or if the agent legitimately consolidated multiple requested items into the exact single file format the task specified.
Consequence
The grader looks for the structured files, finds them missing, and marks every content check as WRONG/MISSING regardless of whether the rendered chart itself was correct.
id f9de2e28fa8e · mined from da-code dacode-plot-line-015@s6
raw text (what the judge reads)
### Missing machine-checkable output artifacts (only the human-facing rendering is produced)
- **Applies when**: `task` -- the task asks for a deliverable produced by a spec file/config (e.g., a chart, table, or report) and grading is done on the underlying data/spec artifacts, not just the rendered image or prose summary.
- **Pattern**: The agent writes only the visual/pretty output (an image file) plus a narrative summary, never persisting the plotted series/values or the resolved spec in the structured side-files that the environment's spec/convention implies (e.g., a JSON of the plot definition and an array of the plotted numbers), and no script is left behind to regenerate them.
- **Detection procedure**:
  1. Read the task and the referenced spec/config file for every named or implied output artifact (file names, extensions, serialization formats), including any conventional companions to the main deliverable.
  2. Inspect the scripts/commands for explicit write calls for each artifact (e.g., `savefig`, `json.dump`, `np.save`, `to_csv`); list which artifacts have no corresponding write.
  3. Check the answer: does it claim success while only describing one artifact (image, screenshot, prose stats), and does it enumerate the data actually serialized?
  4. Flag if any required/implied artifact lacks both a writing step and a saved reproducible script.
- **Discriminator**: A real violation is when an expected artifact is never written to disk (or is written under a different name/extension/location than specified). It is *not* a violation if the artifact exists but its content is merely debatable, or if the agent legitimately consolidated multiple requested items into the exact single file format the task specified.
- **Consequence**: The grader looks for the structured files, finds them missing, and marks every content check as WRONG/MISSING regardless of whether the rendered chart itself was correct.
312Hard-coded / invented input data instead of loading the provided datasettaskda-code
Applies when
task -- the task supplies data files (plus a README/tips and a sample output schema) and the script must compute a statistic or model result from them.
Pattern
The script defines the input arrays/tables as literals typed into the source (or synthesizes them) rather than reading the supplied files, so every downstream number is computed on invented values; the output schema is likewise improvised instead of copied from the provided sample file.
Detection procedure
  1. Read the task/README and list the input artifacts that are supposed to exist (data files) and the output contract (sample result file: exact columns, row count, naming, rounding).
  2. Grep the scripts for any file-reading call (read_csv, load, open, path strings) and confirm the analysis variables trace back to those reads; flag immediately if the data appear as in-line literals or randomly generated values.
  3. Check that the written output file's columns/rows are derived from the provided sample schema, not from an ad-hoc dict of extra diagnostic fields; also confirm the write call actually executes (file name, index=False, no truncated/unfinished call).
  4. Compare reported summary stats (n, means) against anything stated in the task/README; implausibly clean or suspiciously round data is a red flag.
Discriminator
A real violation is data values appearing only in the script with no read from the supplied files; it is fine to hard-code small constants, thresholds, seeds, or to reconstruct a lookup table explicitly named in the instructions, as long as the analyzed observations come from the provided source.
Consequence
The grader compares result.csv against the expected value computed from the true data and marks it WRONG/MISSING — the schema mismatch and the statistic computed on fabricated observations both fail, even though the statistical method itself was described correctly.
id 8ef3e89df1d8 · mined from da-code dacode-data-sa-028@s6
raw text (what the judge reads)
### Hard-coded / invented input data instead of loading the provided dataset
- **Applies when**: `task` -- the task supplies data files (plus a README/tips and a sample output schema) and the script must compute a statistic or model result from them.
- **Pattern**: The script defines the input arrays/tables as literals typed into the source (or synthesizes them) rather than reading the supplied files, so every downstream number is computed on invented values; the output schema is likewise improvised instead of copied from the provided sample file.
- **Detection procedure**:
  1. Read the task/README and list the input artifacts that are supposed to exist (data files) and the output contract (sample result file: exact columns, row count, naming, rounding).
  2. Grep the scripts for any file-reading call (`read_csv`, `load`, `open`, path strings) and confirm the analysis variables trace back to those reads; flag immediately if the data appear as in-line literals or randomly generated values.
  3. Check that the written output file's columns/rows are derived from the provided sample schema, not from an ad-hoc dict of extra diagnostic fields; also confirm the write call actually executes (file name, `index=False`, no truncated/unfinished call).
  4. Compare reported summary stats (n, means) against anything stated in the task/README; implausibly clean or suspiciously round data is a red flag.
- **Discriminator**: A real violation is data values appearing only in the script with no read from the supplied files; it is *fine* to hard-code small constants, thresholds, seeds, or to reconstruct a lookup table explicitly named in the instructions, as long as the analyzed observations come from the provided source.
- **Consequence**: The grader compares `result.csv` against the expected value computed from the true data and marks it WRONG/MISSING — the schema mismatch and the statistic computed on fabricated observations both fail, even though the statistical method itself was described correctly.
313Spec files referenced by the task are never read; required output artifacts are never producedtaskda-code
Applies when
task -- The prompt points to auxiliary instruction/configuration files (e.g., a tips/README/spec text file, a YAML/JSON style or format config) and/or implies a set of deliverable files beyond the one obvious output.
Pattern
The script hardcodes assumptions the agent guessed would be in those files (filters, category definitions, axis/label/style choices) without ever opening or parsing them, and writes only the single most obvious artifact (e.g., the image) while silently skipping other expected deliverables (serialized plot data, arrays, config echoes).
Detection procedure
  1. From the task text, list every referenced auxiliary file and every named/implied output file.
  2. Grep the scripts for each referenced file: is it opened/parsed (open, read_csv, yaml.safe_load, json.load) and are its values actually used to drive the computation/plot, rather than duplicated as literals in code?
  3. Grep for each expected output: is it written (savefig, np.save, json.dump, to_csv) with the exact filename/extension/location asked for?
  4. Check the final answer: does it merely assert compliance ("per the tips file...", "formatted as specified") without evidence of having loaded the spec or listing all produced files?
Discriminator
A real violation is when spec content appears only as inline literals or narrative claims with no read of the file, or when a named deliverable is absent from the code. It is not a violation if the script loads the spec and legitimately inlines derived defaults, or if the "missing" file is genuinely not requested anywhere in the task or its referenced configs.
Consequence
Graders that check each expected artifact report the missing files as WRONG/MISSING, and even the produced artifact fails comparison because its filtering, grouping, styling, or axis conventions diverge from the unread specification — a total (0/N) score despite a confident completion summary.
id d82f71cf922c · mined from da-code dacode-plot-line-006@s6
raw text (what the judge reads)
### Spec files referenced by the task are never read; required output artifacts are never produced
- **Applies when**: `task` -- The prompt points to auxiliary instruction/configuration files (e.g., a tips/README/spec text file, a YAML/JSON style or format config) and/or implies a set of deliverable files beyond the one obvious output.
- **Pattern**: The script hardcodes assumptions the agent *guessed* would be in those files (filters, category definitions, axis/label/style choices) without ever opening or parsing them, and writes only the single most obvious artifact (e.g., the image) while silently skipping other expected deliverables (serialized plot data, arrays, config echoes).
- **Detection procedure**:
  1. From the task text, list every referenced auxiliary file and every named/implied output file.
  2. Grep the scripts for each referenced file: is it opened/parsed (`open`, `read_csv`, `yaml.safe_load`, `json.load`) and are its values actually used to drive the computation/plot, rather than duplicated as literals in code?
  3. Grep for each expected output: is it written (`savefig`, `np.save`, `json.dump`, `to_csv`) with the exact filename/extension/location asked for?
  4. Check the final answer: does it merely assert compliance ("per the tips file...", "formatted as specified") without evidence of having loaded the spec or listing all produced files?
- **Discriminator**: A real violation is when spec content appears only as inline literals or narrative claims with no read of the file, or when a named deliverable is absent from the code. It is *not* a violation if the script loads the spec and legitimately inlines derived defaults, or if the "missing" file is genuinely not requested anywhere in the task or its referenced configs.
- **Consequence**: Graders that check each expected artifact report the missing files as WRONG/MISSING, and even the produced artifact fails comparison because its filtering, grouping, styling, or axis conventions diverge from the unread specification — a total (0/N) score despite a confident completion summary.
314Ignoring a referenced specification file that defines the required grouping/definitionstaskda-code
Applies when
task -- the prompt points to an auxiliary document (README, spec, schema, or instructions file) that defines how categories, bins, filters, or outputs must be constructed.
Pattern
The agent skips opening the referenced file and instead substitutes its own convention — e.g., using the raw values/categories present in the data column as-is, or a self-invented binning scheme — so the produced aggregation and any required side artifacts do not match the mandated definition.
Detection procedure
  1. Read the task and list every externally referenced artifact and every explicitly named output (file names, titles, labels, extra serialized results).
  2. Search the scripts for a read/open of each referenced document and check that the derived categories/bins in the code are traceable to that document rather than hard-coded from the agent's prior assumptions.
  3. Check the scripts write every named output artifact, not just the most visible one (e.g., the image but not the accompanying data/array/JSON files).
  4. Compare the answer's category list/labels to the spec's; if the agent never quoted or echoed the spec's rules, treat the mapping as unverified.
Discriminator
A real violation is when no code path ever reads the referenced file and the grouping is invented or copied verbatim from raw data levels; it is fine if the agent read the file (or reproduced its rules verbatim in comments/output) and the hard-coded bins provably match it, and all required artifacts are written.
Consequence
The plotted/serialized bin counts and labels differ from the reference grouping and required companion files are missing, so every artifact check fails even though the chart looks superficially correct.
id c2f5e93b1194 · mined from da-code dacode-plot-bar-005@s6
raw text (what the judge reads)
### Ignoring a referenced specification file that defines the required grouping/definitions
- **Applies when**: `task` -- the prompt points to an auxiliary document (README, spec, schema, or instructions file) that defines how categories, bins, filters, or outputs must be constructed.
- **Pattern**: The agent skips opening the referenced file and instead substitutes its own convention — e.g., using the raw values/categories present in the data column as-is, or a self-invented binning scheme — so the produced aggregation and any required side artifacts do not match the mandated definition.
- **Detection procedure**:
  1. Read the task and list every externally referenced artifact and every explicitly named output (file names, titles, labels, extra serialized results).
  2. Search the scripts for a read/open of each referenced document and check that the derived categories/bins in the code are traceable to that document rather than hard-coded from the agent's prior assumptions.
  3. Check the scripts write every named output artifact, not just the most visible one (e.g., the image but not the accompanying data/array/JSON files).
  4. Compare the answer's category list/labels to the spec's; if the agent never quoted or echoed the spec's rules, treat the mapping as unverified.
- **Discriminator**: A real violation is when no code path ever reads the referenced file and the grouping is invented or copied verbatim from raw data levels; it is fine if the agent read the file (or reproduced its rules verbatim in comments/output) and the hard-coded bins provably match it, and all required artifacts are written.
- **Consequence**: The plotted/serialized bin counts and labels differ from the reference grouping and required companion files are missing, so every artifact check fails even though the chart looks superficially correct.
315Requested answer artifact never written in the specified schemataskda-code
Applies when
task -- the task specifies an exact output template (e.g., a JSON object with named keys/list values) and/or the grading harness expects a result file, and the script's job is to produce that deliverable.
Pattern
The script computes the quantity but only prints diagnostics to stdout; it never serializes the result to the expected output file, and the final chat answer is hand-typed with types/shapes that deviate from the template (scalars where lists are shown, renamed keys, extra text, unrounded/typed values).
Detection procedure
  1. Read the task statement and note the literal output contract: file name/location if any, exact key names, value containers (list vs scalar), and any rounding/unit/ordering rules.
  2. Scan the scripts for any write step (json.dump, to_csv, open(...,'w')) targeting that artifact; if all output goes through print, the deliverable is missing.
  3. Compare the submitted answer's keys, nesting, and value types character-by-character against the template; verify numeric formatting matches any stated precision.
  4. Confirm the value reported is the final requested quantity, not an intermediate printed line.
Discriminator
A real violation is when no code path produces the required file or the answer's structure/typing differs from the given template; it is fine if the script writes the artifact (even with extra console logging) and the answer reproduces the template exactly, or if the task genuinely requests only inline text with no file and no container hints.
Consequence
The grader reports the expected result file as WRONG/MISSING (or a schema mismatch), scoring 0 even when the underlying computation happened to be right.
id d1c6fad36c97 · mined from da-code dacode-di-text-002@s6
raw text (what the judge reads)
### Requested answer artifact never written in the specified schema
- **Applies when**: `task` -- the task specifies an exact output template (e.g., a JSON object with named keys/list values) and/or the grading harness expects a result file, and the script's job is to produce that deliverable.
- **Pattern**: The script computes the quantity but only prints diagnostics to stdout; it never serializes the result to the expected output file, and the final chat answer is hand-typed with types/shapes that deviate from the template (scalars where lists are shown, renamed keys, extra text, unrounded/typed values).
- **Detection procedure**:
  1. Read the task statement and note the literal output contract: file name/location if any, exact key names, value containers (list vs scalar), and any rounding/unit/ordering rules.
  2. Scan the scripts for any write step (`json.dump`, `to_csv`, `open(...,'w')`) targeting that artifact; if all output goes through `print`, the deliverable is missing.
  3. Compare the submitted answer's keys, nesting, and value types character-by-character against the template; verify numeric formatting matches any stated precision.
  4. Confirm the value reported is the final requested quantity, not an intermediate printed line.
- **Discriminator**: A real violation is when no code path produces the required file or the answer's structure/typing differs from the given template; it is fine if the script writes the artifact (even with extra console logging) and the answer reproduces the template exactly, or if the task genuinely requests only inline text with no file and no container hints.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING (or a schema mismatch), scoring 0 even when the underlying computation happened to be right.
316Using the wrong variant of a named statistical formulataskinfiagent-dabench
Applies when
task -- the task names a specific, textbook-defined statistic/metric (a formula with variants, e.g. "first" vs "second" coefficient, sample vs population denominator, macro vs micro averaging) and the script hard-codes one formula.
Pattern
The script substitutes a related-but-different formula (a sibling variant of the same named quantity, or a different estimator/denominator convention) and reports its value, often with a comment asserting it is the requested definition without any cross-check.
Detection procedure
  1. From the task statement, write down the exact named definition, including which distributional inputs/parameters it requires.
  2. Read the script's formula line-by-line and check that every required input appears and that no substitute input is used in its place (e.g. a different central-tendency measure, a different normalization, a different denominator/degrees-of-freedom convention).
  3. Check whether the script computes any alternative variant as a comparison or sanity check (e.g. a library implementation or the sibling formula) and reconciles differences; absence of any cross-check on a formula with known variants is a red flag.
  4. Confirm the reported number is the output of the required definition, not of the substituted one.
Discriminator
A real violation is a formula that is mathematically not the requested definition (different inputs or normalization), even if it yields the same qualitative conclusion; a look-alike that is fine is an algebraically equivalent rewrite of the same definition, or a benign convention choice that the script explicitly justifies and shows does not change the rounded reported value.
Consequence
The qualitative/categorical part of the answer may still match, but the numeric value differs from the ground truth, so the graded numeric check fails (partial credit at best).
id 593624f201c4 · mined from infiagent-dabench dabench-359@s6
raw text (what the judge reads)
### Using the wrong variant of a named statistical formula
- **Applies when**: `task` -- the task names a specific, textbook-defined statistic/metric (a formula with variants, e.g. "first" vs "second" coefficient, sample vs population denominator, macro vs micro averaging) and the script hard-codes one formula.
- **Pattern**: The script substitutes a related-but-different formula (a sibling variant of the same named quantity, or a different estimator/denominator convention) and reports its value, often with a comment asserting it is the requested definition without any cross-check.
- **Detection procedure**:
  1. From the task statement, write down the exact named definition, including which distributional inputs/parameters it requires.
  2. Read the script's formula line-by-line and check that every required input appears and that no substitute input is used in its place (e.g. a different central-tendency measure, a different normalization, a different denominator/degrees-of-freedom convention).
  3. Check whether the script computes any alternative variant as a comparison or sanity check (e.g. a library implementation or the sibling formula) and reconciles differences; absence of any cross-check on a formula with known variants is a red flag.
  4. Confirm the reported number is the output of the required definition, not of the substituted one.
- **Discriminator**: A real violation is a formula that is mathematically not the requested definition (different inputs or normalization), even if it yields the same qualitative conclusion; a look-alike that is fine is an algebraically equivalent rewrite of the same definition, or a benign convention choice that the script explicitly justifies and shows does not change the rounded reported value.
- **Consequence**: The qualitative/categorical part of the answer may still match, but the numeric value differs from the ground truth, so the graded numeric check fails (partial credit at best).
317Unverified output schema/units for a required result filetaskda-code
Applies when
task -- The task says results must be written to a specific file "following the required format," but the format (column names, order, units, index, rounding) is not fully spelled out in the prompt.
Pattern
The agent invents its own column names, column set, and value convention (e.g., decimals vs. percent, net vs. gross cumulative factor, including/excluding an index or date column) without inspecting the input file's conventions or any provided template/example, and never sanity-checks that its numbers are on the same scale as the source data.
Detection procedure
  1. Read the task for any statement that the output must match a "required/expected format" and note that the concrete schema is not given in the prompt.
  2. Search the scripts for evidence that the schema was derived: reading a template/expected-output/sample file, listing directory contents, or printing the input file's head/dtypes to infer units and naming conventions; the absence of any such step is the red flag.
  3. Check whether the transformation is unit-consistent with the source: e.g., are per-period values in percent or decimal, and is the compounding/aggregation formula applied on the matching scale? Look for a printed sanity check comparing output magnitudes to the input's magnitude.
  4. Read the answer for hard-coded, self-chosen column labels and a claim that the format was "verified," where the verification only re-prints the agent's own dataframe.
Discriminator
A real violation is when the schema/units are chosen by assumption and only self-consistently checked; it is fine if the script explicitly reads a provided template/example (or the prompt enumerates the exact columns) and asserts the produced file matches those names, order, and value scale.
Consequence
The file exists and looks plausible, but the checker comparing it to the expected result fails on column names, ordering, or values off by a factor (percent vs. decimal, or +1 offset), giving 0 despite correct portfolio logic.
id 6cf8e25074c0 · mined from da-code dacode-dm-csv-050@s6
raw text (what the judge reads)
### Unverified output schema/units for a required result file
- **Applies when**: `task` -- The task says results must be written to a specific file "following the required format," but the format (column names, order, units, index, rounding) is not fully spelled out in the prompt.
- **Pattern**: The agent invents its own column names, column set, and value convention (e.g., decimals vs. percent, net vs. gross cumulative factor, including/excluding an index or date column) without inspecting the input file's conventions or any provided template/example, and never sanity-checks that its numbers are on the same scale as the source data.
- **Detection procedure**:
  1. Read the task for any statement that the output must match a "required/expected format" and note that the concrete schema is not given in the prompt.
  2. Search the scripts for evidence that the schema was *derived*: reading a template/expected-output/sample file, listing directory contents, or printing the input file's head/dtypes to infer units and naming conventions; the absence of any such step is the red flag.
  3. Check whether the transformation is unit-consistent with the source: e.g., are per-period values in percent or decimal, and is the compounding/aggregation formula applied on the matching scale? Look for a printed sanity check comparing output magnitudes to the input's magnitude.
  4. Read the answer for hard-coded, self-chosen column labels and a claim that the format was "verified," where the verification only re-prints the agent's own dataframe.
- **Discriminator**: A real violation is when the schema/units are chosen by assumption and only self-consistently checked; it is fine if the script explicitly reads a provided template/example (or the prompt enumerates the exact columns) and asserts the produced file matches those names, order, and value scale.
- **Consequence**: The file exists and looks plausible, but the checker comparing it to the expected result fails on column names, ordering, or values off by a factor (percent vs. decimal, or +1 offset), giving 0 despite correct portfolio logic.
318Distribution statistics computed on an uncleaned/unverified subset, with extreme moments accepted at face valuetaskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test and/or descriptive statistics (normality, skewness, kurtosis, means) on a single numeric column of a loaded table.
Pattern
The attempt feeds the raw column straight into the test/statistic without first inspecting it (count, dtype, min/max, unique low-frequency extremes, NaN/sentinel codes such as -999/0/blank, duplicated or non-target rows), then reports a conclusion driven by a handful of contaminating values; the requested diagnostic (e.g., the p-value) and the cleaning decisions are not shown, so the number cannot be reproduced or checked.
Detection procedure
  1. Read the task: note the exact column, any implicit population (valid/non-missing records only), and every quantity the task says to report (test statistic, p-value, rounded moments).
  2. Read the script: check whether it prints n, dtype, min/max/quantiles and NaN/sentinel counts for the column, and whether it explicitly handles or documents excluded values before calling the test/moment functions.
  3. Compare the reported moments against the printed diagnostics: strongly non-zero skew/kurtosis (|skew| > ~1, excess kurtosis > ~3) with no accompanying look at the tail values is an unvalidated result, especially when it flips the yes/no verdict.
  4. Check the answer text: is the stated required diagnostic (p-value) present, and does the yes/no conclusion follow from it at the stated alpha, using the same n as the cleaned data?
Discriminator
A real violation is when no distributional sanity check exists, or the diagnostics that were printed show placeholder/sentinel-like extremes, mixed record types, or an n that differs from the intended population, and the conclusion hinges on them. It is fine if the script prints the diagnostics, shows the extremes are genuine data, and the heavy tails persist — a genuinely skewed distribution reported with its p-value is not a defect.
Consequence
The test statistic is computed on the wrong effective sample, so the p-value crosses alpha in the wrong direction and the binary verdict (plus both moment values) is graded wrong, with nothing in the submission letting a reviewer trace the error.
id b4484d4bc700 · mined from infiagent-dabench dabench-298@s6
raw text (what the judge reads)
### Distribution statistics computed on an uncleaned/unverified subset, with extreme moments accepted at face value
- **Applies when**: `task` -- the task asks for a hypothesis test and/or descriptive statistics (normality, skewness, kurtosis, means) on a single numeric column of a loaded table.
- **Pattern**: The attempt feeds the raw column straight into the test/statistic without first inspecting it (count, dtype, min/max, unique low-frequency extremes, NaN/sentinel codes such as -999/0/blank, duplicated or non-target rows), then reports a conclusion driven by a handful of contaminating values; the requested diagnostic (e.g., the p-value) and the cleaning decisions are not shown, so the number cannot be reproduced or checked.
- **Detection procedure**:
  1. Read the task: note the exact column, any implicit population (valid/non-missing records only), and every quantity the task says to report (test statistic, p-value, rounded moments).
  2. Read the script: check whether it prints n, dtype, min/max/quantiles and NaN/sentinel counts for the column, and whether it explicitly handles or documents excluded values before calling the test/moment functions.
  3. Compare the reported moments against the printed diagnostics: strongly non-zero skew/kurtosis (|skew| > ~1, excess kurtosis > ~3) with no accompanying look at the tail values is an unvalidated result, especially when it flips the yes/no verdict.
  4. Check the answer text: is the stated required diagnostic (p-value) present, and does the yes/no conclusion follow from it at the stated alpha, using the same n as the cleaned data?
- **Discriminator**: A real violation is when no distributional sanity check exists, or the diagnostics that were printed show placeholder/sentinel-like extremes, mixed record types, or an n that differs from the intended population, and the conclusion hinges on them. It is fine if the script prints the diagnostics, shows the extremes are genuine data, and the heavy tails persist — a genuinely skewed distribution reported with its p-value is not a defect.
- **Consequence**: The test statistic is computed on the wrong effective sample, so the p-value crosses alpha in the wrong direction and the binary verdict (plus both moment values) is graded wrong, with nothing in the submission letting a reviewer trace the error.
319Deliverable file not verifiably produced with full, aligned row coveragetaskda-code
Applies when
task -- the task asks for per-row predictions written to a named output file with a specified column, and the agent reports results inline instead of demonstrating the saved artifact.
Pattern
The attempt pastes a (truncated) list of predicted labels as the "answer" with no saved script showing that the file was written, and no check that the number of predicted rows equals the number of input rows or that the row order matches the input.
Detection procedure
  1. Read the task and note the exact required filename, column name, and the implied row count (= number of rows in the evaluation input).
  2. Inspect the scripts for an explicit write of that file (e.g., a DataFrame with only the required column written without an index) and for a preceding assertion/print of len(predictions) == len(input_df) and preserved input order.
  3. Compare the answer's content to the expected row count/format: if it is a pasted, truncated, or mid-word-cut label list rather than evidence of a complete written file, flag it.
  4. Confirm no code path drops rows (dropna, filtering, deduplication, reindexing, shuffling) between reading the input and writing predictions.
Discriminator
A fine attempt shows the file being written and a shape/count sanity check (and may additionally print a preview); a violation offers only inline text or a file whose row count/order cannot be tied back to the evaluation input.
Consequence
The grader reports the expected output file as WRONG/MISSING, or the file misaligns with ground-truth rows, scoring near zero regardless of model quality.
id 9d86bc0b1032 · mined from da-code dacode-ml-multi-011@s6
raw text (what the judge reads)
### Deliverable file not verifiably produced with full, aligned row coverage
- **Applies when**: `task` -- the task asks for per-row predictions written to a named output file with a specified column, and the agent reports results inline instead of demonstrating the saved artifact.
- **Pattern**: The attempt pastes a (truncated) list of predicted labels as the "answer" with no saved script showing that the file was written, and no check that the number of predicted rows equals the number of input rows or that the row order matches the input.
- **Detection procedure**:
  1. Read the task and note the exact required filename, column name, and the implied row count (= number of rows in the evaluation input).
  2. Inspect the scripts for an explicit write of that file (e.g., a DataFrame with only the required column written without an index) and for a preceding assertion/print of `len(predictions) == len(input_df)` and preserved input order.
  3. Compare the answer's content to the expected row count/format: if it is a pasted, truncated, or mid-word-cut label list rather than evidence of a complete written file, flag it.
  4. Confirm no code path drops rows (dropna, filtering, deduplication, reindexing, shuffling) between reading the input and writing predictions.
- **Discriminator**: A fine attempt shows the file being written and a shape/count sanity check (and may additionally print a preview); a violation offers only inline text or a file whose row count/order cannot be tied back to the evaluation input.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING, or the file misaligns with ground-truth rows, scoring near zero regardless of model quality.
320Final submission produced by a downsampled, in-sample-only-validated modeltaskda-code
Applies when
task -- the task requires a predictions/results file and the scripts include several successive modeling attempts that all write to the same output path.
Pattern
The last script executed trains on an arbitrary subsample of the available training rows "for speed" with small/shallow default-ish models, judges quality only by metrics computed on the same rows it trained on (or reports no held-out metric at all), and overwrites the output file previously produced by a better-validated, full-data model — so the delivered artifact is the weakest, unverified one.
Detection procedure
  1. Read the task to confirm the deliverable is a scored predictions file and note the format/reference template.
  2. For each script, record: fraction of training data used, whether a held-out/CV score is computed on data not used in fitting, and the output path written.
  3. Identify which script writes the file last; check whether its model was compared on a common held-out set against the other candidates, and whether it used all available data (or justified a subsample by a measured, not assumed, score).
  4. Check the final answer: does it quote a held-out score for the exact model that produced the saved file, plus shape/id-alignment/range sanity checks against the template?
Discriminator
A violation is when the shipped model's only reported metric is computed on its training rows (or is absent) and it was trained on a reduced sample with no evidence it matches/beats the full-data alternative; it is fine if subsampling or a simpler model is chosen after an apples-to-apples held-out comparison, or if the last-written file comes from the model with the best validated score.
Consequence
The saved predictions come from an underfit/unverified model whose true generalization error is unknown and worse than an already-available alternative, so the file fails the grader's accuracy threshold even though it is formatted correctly.
id 2578b8ef262a · mined from da-code dacode-ml-competition-008@s6
raw text (what the judge reads)
### Final submission produced by a downsampled, in-sample-only-validated model
- **Applies when**: `task` -- the task requires a predictions/results file and the scripts include several successive modeling attempts that all write to the same output path.
- **Pattern**: The last script executed trains on an arbitrary subsample of the available training rows "for speed" with small/shallow default-ish models, judges quality only by metrics computed on the same rows it trained on (or reports no held-out metric at all), and overwrites the output file previously produced by a better-validated, full-data model — so the delivered artifact is the weakest, unverified one.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is a scored predictions file and note the format/reference template.
  2. For each script, record: fraction of training data used, whether a held-out/CV score is computed on data not used in fitting, and the output path written.
  3. Identify which script writes the file last; check whether its model was compared on a common held-out set against the other candidates, and whether it used all available data (or justified a subsample by a measured, not assumed, score).
  4. Check the final answer: does it quote a held-out score for the exact model that produced the saved file, plus shape/id-alignment/range sanity checks against the template?
- **Discriminator**: A violation is when the *shipped* model's only reported metric is computed on its training rows (or is absent) and it was trained on a reduced sample with no evidence it matches/beats the full-data alternative; it is fine if subsampling or a simpler model is chosen after an apples-to-apples held-out comparison, or if the last-written file comes from the model with the best validated score.
- **Consequence**: The saved predictions come from an underfit/unverified model whose true generalization error is unknown and worse than an already-available alternative, so the file fails the grader's accuracy threshold even though it is formatted correctly.
321Invented qualification threshold / definition not grounded in the provided spec or sample outputtaskda-code
Applies when
task -- The task references a definition, threshold, or output template that is given in an accompanying README/sample file (possibly incomplete or truncated), and the script must group/filter records before ranking them.
Pattern
The script hard-codes a filtering cutoff (e.g., a minimum count of records or votes), an aggregation choice (sum vs. mean vs. per-item), and a scope for that filter (applied to every ranked category rather than only the one it defines), none of which are verified against the stated definition or the provided sample output file; the sample/template file is never loaded or compared to the produced file.
Detection procedure
1. Read the task/README and list every quantitative qualifier and formatting rule it states, noting any that are ambiguous or cut off. 2. Search the scripts for the corresponding constants and aggregation calls; check whether each was read from the spec or chosen arbitrarily, and whether the script ever inspects the provided sample/reference file. 3. Check whether a filter tied to one definition is silently reused for the other requested rankings, and whether alternative reasonable definitions (mean vs. sum, per-book minimum vs. total minimum) would change the top-10 lists. 4. Inspect the produced answer for tell-tale signs of a mis-specified filter: entries with a single underlying record, obviously niche or non-author entities, or lists that are alphabetically ordered rather than value-ordered.
Discriminator
A real violation is when the chosen constant/aggregation is unstated in the sources, is not tested for sensitivity, and materially changes the ranking; it is fine if the constant is explicitly given in the spec, or if the script demonstrates (e.g., by printing results under several plausible thresholds) that the top-10 lists are stable and matches the sample file's column names, order, and index convention.
Consequence
The output file has the right shape but wrong membership/order in one or more columns, so an exact-match file comparison fails all checks.
id 97e07144703e · mined from da-code dacode-dm-csv-009@s6
raw text (what the judge reads)
### Invented qualification threshold / definition not grounded in the provided spec or sample output
- **Applies when**: `task` -- The task references a definition, threshold, or output template that is given in an accompanying README/sample file (possibly incomplete or truncated), and the script must group/filter records before ranking them.
- **Pattern**: The script hard-codes a filtering cutoff (e.g., a minimum count of records or votes), an aggregation choice (sum vs. mean vs. per-item), and a scope for that filter (applied to every ranked category rather than only the one it defines), none of which are verified against the stated definition or the provided sample output file; the sample/template file is never loaded or compared to the produced file.
- **Detection procedure**: 1. Read the task/README and list every quantitative qualifier and formatting rule it states, noting any that are ambiguous or cut off. 2. Search the scripts for the corresponding constants and aggregation calls; check whether each was read from the spec or chosen arbitrarily, and whether the script ever inspects the provided sample/reference file. 3. Check whether a filter tied to one definition is silently reused for the other requested rankings, and whether alternative reasonable definitions (mean vs. sum, per-book minimum vs. total minimum) would change the top-10 lists. 4. Inspect the produced answer for tell-tale signs of a mis-specified filter: entries with a single underlying record, obviously niche or non-author entities, or lists that are alphabetically ordered rather than value-ordered.
- **Discriminator**: A real violation is when the chosen constant/aggregation is unstated in the sources, is not tested for sensitivity, and materially changes the ranking; it is fine if the constant is explicitly given in the spec, or if the script demonstrates (e.g., by printing results under several plausible thresholds) that the top-10 lists are stable and matches the sample file's column names, order, and index convention.
- **Consequence**: The output file has the right shape but wrong membership/order in one or more columns, so an exact-match file comparison fails all checks.
322Fabricated or placeholder input data substituted for the real dataset (and unverified against basic size/shape facts)taskda-code
Applies when
task -- the task points to a provided data file (plus a config/spec file) and the agent's scripts must load it, aggregate it, and emit the required output artifacts.
Pattern
The attempt never demonstrably reads the supplied file (or reads only a truncated/generated stand-in, e.g. a random sample, a hard-coded array, or synthetic values), then reports summary statistics from that stand-in as if they were the real result; the reported numbers are round or implausibly uniform and are not cross-checked against the documented row count or expected value ranges. Required side artifacts named in the spec (serialized plot data, numeric arrays) are also not produced because the pipeline was improvised.
Detection procedure
  1. Read the task/README for the stated data source and the full list of expected outputs (all files/formats, not just the image).
  2. In the scripts, confirm an explicit load of the actual provided path with no head/nrows/sample/random-generation shortcut, and confirm every named output artifact is written.
  3. Compare the answer's reported N (and min/max of the analysed variable) to what the README/data imply; flag suspiciously round totals (e.g. exactly 2,000) or near-equal bucket counts inconsistent with a real-world skewed distribution.
  4. Check that bucket edges/labels/ordering come from the spec file rather than being invented in prose.
Discriminator
A real violation is when the reported totals cannot be reconciled with the actual file (round N, uniform bins, truncated value range) or when required artifacts are absent; a look-alike that is fine is a legitimate filtered subset (e.g. movies only) whose reduced count is explicitly derived from the loaded data and whose distribution is plausibly skewed.
Consequence
Every graded artifact mismatches the reference — bar heights/bin counts differ from ground truth and the missing serialized outputs score as absent, yielding 0/3 checks passed despite a confident-looking summary.
id d22563c45967 · mined from da-code dacode-plot-bar-007@s6
raw text (what the judge reads)
### Fabricated or placeholder input data substituted for the real dataset (and unverified against basic size/shape facts)
- **Applies when**: `task` -- the task points to a provided data file (plus a config/spec file) and the agent's scripts must load it, aggregate it, and emit the required output artifacts.
- **Pattern**: The attempt never demonstrably reads the supplied file (or reads only a truncated/generated stand-in, e.g. a random sample, a hard-coded array, or synthetic values), then reports summary statistics from that stand-in as if they were the real result; the reported numbers are round or implausibly uniform and are not cross-checked against the documented row count or expected value ranges. Required side artifacts named in the spec (serialized plot data, numeric arrays) are also not produced because the pipeline was improvised.
- **Detection procedure**:
  1. Read the task/README for the stated data source and the full list of expected outputs (all files/formats, not just the image).
  2. In the scripts, confirm an explicit load of the actual provided path with no `head`/`nrows`/`sample`/random-generation shortcut, and confirm every named output artifact is written.
  3. Compare the answer's reported N (and min/max of the analysed variable) to what the README/data imply; flag suspiciously round totals (e.g. exactly 2,000) or near-equal bucket counts inconsistent with a real-world skewed distribution.
  4. Check that bucket edges/labels/ordering come from the spec file rather than being invented in prose.
- **Discriminator**: A real violation is when the reported totals cannot be reconciled with the actual file (round N, uniform bins, truncated value range) or when required artifacts are absent; a look-alike that is fine is a legitimate filtered subset (e.g. movies only) whose reduced count is explicitly derived from the loaded data and whose distribution is plausibly skewed.
- **Consequence**: Every graded artifact mismatches the reference — bar heights/bin counts differ from ground truth and the missing serialized outputs score as absent, yielding 0/3 checks passed despite a confident-looking summary.
323Output artifact not written to the required file name/path/formattaskda-code
Applies when
task -- The task specifies (explicitly or via the expected-deliverable convention) a result file and a JSON/CSV schema that the grader will read, and the scripts end by printing or saving results.
Pattern
The attempt computes plausible results but persists them to an ad-hoc filename/location (e.g., a scratch .txt, a print to stdout, or a working directory other than the one the grader inspects), or writes keys/value types that don't literally match the requested template, so the graded artifact is missing or unreadable even when the analysis is right.
Detection procedure
  1. Read the task statement and note the exact required deliverable: file name, directory, key names, ordering, and value types of the requested output template.
  2. Scan the scripts for every write operation (open(...,'w'), to_csv, json.dump, to_json) and record the exact path and the object being serialized.
  3. Compare path and serialized structure character-for-character with the requirement; check that the keys are the requested ones and the values are the requested entity (e.g., names vs. records/values) in the requested order.
  4. Confirm no step later overwrites, relocates, or leaves the artifact only in stdout/logs.
Discriminator
A real violation is when no write targets the required filename/path or the written structure deviates from the requested schema; it is fine if the script writes the correct file and additionally prints or saves debug copies elsewhere.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, regardless of whether the printed numbers were correct.
id 225486b86c35 · mined from da-code dacode-di-text-003@s6
raw text (what the judge reads)
### Output artifact not written to the required file name/path/format
- **Applies when**: `task` -- The task specifies (explicitly or via the expected-deliverable convention) a result file and a JSON/CSV schema that the grader will read, and the scripts end by printing or saving results.
- **Pattern**: The attempt computes plausible results but persists them to an ad-hoc filename/location (e.g., a scratch `.txt`, a print to stdout, or a working directory other than the one the grader inspects), or writes keys/value types that don't literally match the requested template, so the graded artifact is missing or unreadable even when the analysis is right.
- **Detection procedure**:
  1. Read the task statement and note the exact required deliverable: file name, directory, key names, ordering, and value types of the requested output template.
  2. Scan the scripts for every write operation (`open(...,'w')`, `to_csv`, `json.dump`, `to_json`) and record the exact path and the object being serialized.
  3. Compare path and serialized structure character-for-character with the requirement; check that the keys are the requested ones and the values are the requested entity (e.g., names vs. records/values) in the requested order.
  4. Confirm no step later overwrites, relocates, or leaves the artifact only in stdout/logs.
- **Discriminator**: A real violation is when no write targets the required filename/path or the written structure deviates from the requested schema; it is fine if the script writes the correct file and *additionally* prints or saves debug copies elsewhere.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, regardless of whether the printed numbers were correct.
324Inverting a "was the required step properly handled?" flag into a "was there anything to fix?" flagtaskinfiagent-dabench
Applies when
task -- the answer format asks for a Yes/No (or boolean) confirmation that a prescribed preprocessing/validation step was carried out correctly, and the script derives that flag from data conditions rather than from whether the step was applied.
Pattern
The script computes the flag from a diagnostic count (e.g., flag = "Yes" if n_problem_rows > 0 else "No"), skips the prescribed operation when the count is zero, and then reports "No" — even though the requested procedure was fully and correctly satisfied (nothing needed changing, so the data conforms). The reported flag ends up describing the input data instead of the compliance of the pipeline. The same defect appears when the flag is emitted before the step runs, or when quoting/token formatting deviates from the requested literal format.
Detection procedure
  1. Read the answer spec and restate in words what the Yes/No field asserts — usually "the prescribed step was executed and the constraint now holds", not "the raw data had defects".
  2. Locate the line in the script that assigns the flag; check whether it is a function of a diagnostic count/condition rather than of successful execution plus a post-condition verification.
  3. Check that the post-condition is actually verified after processing (e.g., re-checking counts/ranges/shapes) and that the flag is set from that verification; also confirm the printed answer tokens match the requested literals exactly.
  4. Consider the "no defects found" branch specifically: does the script still report the compliant/affirmative value? If it reports the negative value, flag it.
Discriminator
A real violation is when the constraint is satisfied (or trivially satisfied) yet the script reports the negative value, or reports the negative value purely because no defects existed. It is not a violation if the task explicitly asks whether the raw data contained defects, or if the step genuinely failed/was skipped despite defects being present.
Consequence
The graded field mismatches the expected literal (e.g., expected "Yes", submitted "No"), zeroing that check even though the underlying computation and other fields were done correctly.
id 5e5bb46152cb · mined from infiagent-dabench dabench-550@s6
raw text (what the judge reads)
### Inverting a "was the required step properly handled?" flag into a "was there anything to fix?" flag
- **Applies when**: `task` -- the answer format asks for a Yes/No (or boolean) confirmation that a prescribed preprocessing/validation step was carried out correctly, and the script derives that flag from data conditions rather than from whether the step was applied.
- **Pattern**: The script computes the flag from a diagnostic count (e.g., `flag = "Yes" if n_problem_rows > 0 else "No"`), skips the prescribed operation when the count is zero, and then reports "No" — even though the requested procedure was fully and correctly satisfied (nothing needed changing, so the data conforms). The reported flag ends up describing the input data instead of the compliance of the pipeline. The same defect appears when the flag is emitted before the step runs, or when quoting/token formatting deviates from the requested literal format.
- **Detection procedure**:
  1. Read the answer spec and restate in words what the Yes/No field asserts — usually "the prescribed step was executed and the constraint now holds", not "the raw data had defects".
  2. Locate the line in the script that assigns the flag; check whether it is a function of a diagnostic count/condition rather than of successful execution plus a post-condition verification.
  3. Check that the post-condition is actually verified after processing (e.g., re-checking counts/ranges/shapes) and that the flag is set from that verification; also confirm the printed answer tokens match the requested literals exactly.
  4. Consider the "no defects found" branch specifically: does the script still report the compliant/affirmative value? If it reports the negative value, flag it.
- **Discriminator**: A real violation is when the constraint is satisfied (or trivially satisfied) yet the script reports the negative value, or reports the negative value purely because no defects existed. It is *not* a violation if the task explicitly asks whether the raw data contained defects, or if the step genuinely failed/was skipped despite defects being present.
- **Consequence**: The graded field mismatches the expected literal (e.g., expected "Yes", submitted "No"), zeroing that check even though the underlying computation and other fields were done correctly.
325Fabricating required entities from an unrelated data source instead of verifying the schema matches the tasktaskda-code
Applies when
task -- the task names specific entities/measures (grouping keys, stage/category breakdowns, metric fields) and the agent must locate them in the provided data files before computing anything.
Pattern
The agent never confirms that the required fields exist; it grabs whatever file is present, substitutes loosely-related proxies (a different granularity of key, a row count in place of the requested measure), and invents the category breakdown, then reports success with confident formatting instead of flagging the mismatch.
Detection procedure
1. List from the task statement every entity the output requires (the grouping dimension, the ranking measure, the stacked/segment categories, all declared config/output artifacts). 2. Read the scripts and check that each one is read from an actual column or file, not hardcoded, derived by string-parsing an unrelated field, or replaced by a count. 3. Read the answer for tell-tale substitutions — units or entity types that don't match the requested ones, category labels that appear nowhere in the data, or a measure redefined in passing ("sales = number of records"). 4. Confirm every artifact named in the task (config file consumed, each output file produced) is actually read/written.
Discriminator
A real violation is when the required fields cannot be shown to exist in the loaded data and are conjured or proxied; it is fine if the agent documents a legitimate renaming/mapping between the task's wording and actual column names, with evidence (printed schema, value counts) that the mapped columns hold the requested quantities.
Consequence
Outputs are produced but encode the wrong quantities and often the wrong number/identity of groups; every value-based check (arrays, JSON of plot data, rendered figure) fails, and any missing artifact fails outright.
id b87ebd48ba77 · mined from da-code dacode-plot-scatter-002@s6
raw text (what the judge reads)
### Fabricating required entities from an unrelated data source instead of verifying the schema matches the task
- **Applies when**: `task` -- the task names specific entities/measures (grouping keys, stage/category breakdowns, metric fields) and the agent must locate them in the provided data files before computing anything.
- **Pattern**: The agent never confirms that the required fields exist; it grabs whatever file is present, substitutes loosely-related proxies (a different granularity of key, a row count in place of the requested measure), and invents the category breakdown, then reports success with confident formatting instead of flagging the mismatch.
- **Detection procedure**: 1. List from the task statement every entity the output requires (the grouping dimension, the ranking measure, the stacked/segment categories, all declared config/output artifacts). 2. Read the scripts and check that each one is read from an actual column or file, not hardcoded, derived by string-parsing an unrelated field, or replaced by a count. 3. Read the answer for tell-tale substitutions — units or entity types that don't match the requested ones, category labels that appear nowhere in the data, or a measure redefined in passing ("sales = number of records"). 4. Confirm every artifact named in the task (config file consumed, each output file produced) is actually read/written.
- **Discriminator**: A real violation is when the required fields cannot be shown to exist in the loaded data and are conjured or proxied; it is fine if the agent documents a legitimate renaming/mapping between the task's wording and actual column names, with evidence (printed schema, value counts) that the mapped columns hold the requested quantities.
- **Consequence**: Outputs are produced but encode the wrong quantities and often the wrong number/identity of groups; every value-based check (arrays, JSON of plot data, rendered figure) fails, and any missing artifact fails outright.
326Grouping variable used for filtering is wrongly reused as the unit of analysistaskda-code
Applies when
task -- the task specifies a preprocessing step done within one grouping variable (e.g., per-group outlier trimming) and then asks for a test/metric comparing levels of a different variable.
Pattern
The attempt conflates the two variables: it splits the data by the preprocessing group and runs the requested test once per group, emitting one result (and conclusion) per preprocessing group instead of the single result over the recombined, filtered dataset — so the number and meaning of reported values do not match what was asked.
Detection procedure
  1. Read the task and write down explicitly (a) which variable defines the preprocessing/filtering partition and (b) which variable defines the groups being compared, and hence how many output values the requested statistic should have.
  2. In the scripts, locate the loop/groupby feeding the statistical test and check whether the samples passed to it are the levels of the comparison variable drawn from the full re-concatenated filtered data, or the levels of the preprocessing variable.
  3. Count the elements in the reported answer's arrays and compare with the count derived in step 1; check each conclusion string is derived from its matching p-value at the stated threshold.
  4. Sanity-check that the filtered rows were reunited (total post-filter row count = sum of per-group post-filter counts) before the test, and that each comparison group is non-empty.
Discriminator
A real violation is when the output cardinality/keying follows the preprocessing partition (or any variable other than the one named as the comparison factor) when the task asked for a single test over the pooled filtered data; it is fine if the task itself explicitly requests one test per preprocessing group, or if the reported list length and ordering provably correspond to the comparison levels the task names.
Consequence
The expected result file mismatches on both length and values (e.g., several p-values/conclusions instead of the one requested, or values computed on subsets), so the grader marks the answer wrong even though the test function itself was used correctly.
id 54fb0a9f1f2f · mined from da-code dacode-data-sa-061@s6
raw text (what the judge reads)
### Grouping variable used for filtering is wrongly reused as the unit of analysis
- **Applies when**: `task` -- the task specifies a preprocessing step done *within* one grouping variable (e.g., per-group outlier trimming) and then asks for a test/metric comparing levels of a *different* variable.
- **Pattern**: The attempt conflates the two variables: it splits the data by the preprocessing group and runs the requested test once per group, emitting one result (and conclusion) per preprocessing group instead of the single result over the recombined, filtered dataset — so the number and meaning of reported values do not match what was asked.
- **Detection procedure**:
  1. Read the task and write down explicitly (a) which variable defines the preprocessing/filtering partition and (b) which variable defines the groups being compared, and hence how many output values the requested statistic should have.
  2. In the scripts, locate the loop/groupby feeding the statistical test and check whether the samples passed to it are the levels of the comparison variable drawn from the full re-concatenated filtered data, or the levels of the preprocessing variable.
  3. Count the elements in the reported answer's arrays and compare with the count derived in step 1; check each conclusion string is derived from its matching p-value at the stated threshold.
  4. Sanity-check that the filtered rows were reunited (total post-filter row count = sum of per-group post-filter counts) before the test, and that each comparison group is non-empty.
- **Discriminator**: A real violation is when the output cardinality/keying follows the preprocessing partition (or any variable other than the one named as the comparison factor) when the task asked for a single test over the pooled filtered data; it is fine if the task itself explicitly requests one test per preprocessing group, or if the reported list length and ordering provably correspond to the comparison levels the task names.
- **Consequence**: The expected result file mismatches on both length and values (e.g., several p-values/conclusions instead of the one requested, or values computed on subsets), so the grader marks the answer wrong even though the test function itself was used correctly.
327Predictions not aligned one-to-one with the original test rowstaskda-code
Applies when
task -- the script must emit a prediction file whose rows correspond, in order, to every row of a provided test set, and the preprocessing code filters, drops, re-indexes, or reorders test rows (e.g. dropna, deduplication, boolean masking) before predicting.
Pattern
The attempt drops or subsets test rows during cleaning, predicts only on the surviving subset, then re-assembles the full output using DataFrame index labels, positional counters, or a constant fill for the removed rows — mixing label-based and position-based addressing, so predictions land on the wrong rows (or the file has the wrong length), while the printed summary only shows the count and looks plausible.
Detection procedure
  1. From the task, note the required output contract: number of rows = number of test rows, original order preserved, required column name(s)/dtype.
  2. In the script, find every operation that can change the test frame's row count or index (dropna, drop_duplicates, filters, reset_index, merges) and check whether the object used for prediction is still the full, original-order test frame.
  3. Trace the re-assembly step: verify that indices used to scatter predictions back (arr[df.index] = pred, range(len(df)), for i in ... if i not in ...) refer to the same indexing space (positional vs. label) as the array being filled; any mismatch is a bug.
  4. Check the answer/logs for an explicit verification that output length equals test length and that no row was silently imputed with a global mean or left unassigned.
Discriminator
A real violation is when test rows are removed or re-indexed and reinstated through a different addressing scheme, or the output length can differ from the test length. It is fine if missing values are imputed in place (row count untouched), or if the scatter-back is done with a consistently label-aligned structure (e.g. pd.Series(pred, index=subset.index).reindex(test.index)) plus a verified length/order check.
Consequence
The saved file contains predictions attached to the wrong rows or an incorrect number of rows, so the grader's row-wise comparison against ground truth fails even though the model itself may be reasonable.
id c0bb9c27cef1 · mined from da-code dacode-ml-regression-004@s6
raw text (what the judge reads)
### Predictions not aligned one-to-one with the original test rows
- **Applies when**: `task` -- the script must emit a prediction file whose rows correspond, in order, to every row of a provided test set, and the preprocessing code filters, drops, re-indexes, or reorders test rows (e.g. `dropna`, deduplication, boolean masking) before predicting.
- **Pattern**: The attempt drops or subsets test rows during cleaning, predicts only on the surviving subset, then re-assembles the full output using DataFrame index labels, positional counters, or a constant fill for the removed rows — mixing label-based and position-based addressing, so predictions land on the wrong rows (or the file has the wrong length), while the printed summary only shows the count and looks plausible.
- **Detection procedure**:
  1. From the task, note the required output contract: number of rows = number of test rows, original order preserved, required column name(s)/dtype.
  2. In the script, find every operation that can change the test frame's row count or index (`dropna`, `drop_duplicates`, filters, `reset_index`, merges) and check whether the object used for prediction is still the full, original-order test frame.
  3. Trace the re-assembly step: verify that indices used to scatter predictions back (`arr[df.index] = pred`, `range(len(df))`, `for i in ... if i not in ...`) refer to the same indexing space (positional vs. label) as the array being filled; any mismatch is a bug.
  4. Check the answer/logs for an explicit verification that output length equals test length and that no row was silently imputed with a global mean or left unassigned.
- **Discriminator**: A real violation is when test rows are removed or re-indexed and reinstated through a different addressing scheme, or the output length can differ from the test length. It is fine if missing values are imputed in place (row count untouched), or if the scatter-back is done with a consistently label-aligned structure (e.g. `pd.Series(pred, index=subset.index).reindex(test.index)`) plus a verified length/order check.
- **Consequence**: The saved file contains predictions attached to the wrong rows or an incorrect number of rows, so the grader's row-wise comparison against ground truth fails even though the model itself may be reasonable.
328Unvalidated cluster-count choice producing degenerate/singleton groupstaskda-code
Applies when
task -- the task asks for an unsupervised grouping with an "appropriate" number of groups, and the script picks that number automatically (silhouette, elbow, etc.) and writes labels to a required output file.
Pattern
The attempt accepts the auto-selected number of groups without a sanity check on the resulting partition, so the "optimal" score is driven by outliers and yields groups of size 1–3 alongside huge catch-all groups; it also never verifies that the saved file's shape, column names, and label range match the requested spec.
Detection procedure
  1. Read the task for the required output contract (exact column names, one row per input record, label column) and any statement about what the groups should represent.
  2. In the script, check whether the selection criterion is cross-checked with a second diagnostic (stability, alternative metric, comparison of 2–6 candidate values) and whether outlier/skew handling or scaling appropriate to heavy-tailed variables was considered.
  3. In the reported group sizes, flag any solution where one or more groups contain a negligible fraction of records (e.g. <2–3% or a handful of rows) while the interpretation claims substantive segments; treat that as an unvalidated selection rather than a real structure.
  4. Verify the answer states the saved file's row count equals the input record count and that column names/label values literally match the requested format; absence of that verification is a further flag.
Discriminator
A genuinely small group is fine if the script explicitly justifies it (documented outliers, stability across seeds/metrics, or a domain reason) and the remaining groups are still interpretable; the violation is accepting singleton/near-singleton groups purely because a single internal metric peaked there, with no alternative considered and no sanity check on sizes or file shape.
Consequence
The saved labels do not match the expected partition (grader compares cluster structure/agreement), so the output file is marked WRONG even though the pipeline ran without error.
id ac6cec5c3291 · mined from da-code dacode-ml-cluster-013@s6
raw text (what the judge reads)
### Unvalidated cluster-count choice producing degenerate/singleton groups
- **Applies when**: `task` -- the task asks for an unsupervised grouping with an "appropriate" number of groups, and the script picks that number automatically (silhouette, elbow, etc.) and writes labels to a required output file.
- **Pattern**: The attempt accepts the auto-selected number of groups without a sanity check on the resulting partition, so the "optimal" score is driven by outliers and yields groups of size 1–3 alongside huge catch-all groups; it also never verifies that the saved file's shape, column names, and label range match the requested spec.
- **Detection procedure**:
  1. Read the task for the required output contract (exact column names, one row per input record, label column) and any statement about what the groups should represent.
  2. In the script, check whether the selection criterion is cross-checked with a second diagnostic (stability, alternative metric, comparison of 2–6 candidate values) and whether outlier/skew handling or scaling appropriate to heavy-tailed variables was considered.
  3. In the reported group sizes, flag any solution where one or more groups contain a negligible fraction of records (e.g. <2–3% or a handful of rows) while the interpretation claims substantive segments; treat that as an unvalidated selection rather than a real structure.
  4. Verify the answer states the saved file's row count equals the input record count and that column names/label values literally match the requested format; absence of that verification is a further flag.
- **Discriminator**: A genuinely small group is fine if the script explicitly justifies it (documented outliers, stability across seeds/metrics, or a domain reason) and the remaining groups are still interpretable; the violation is accepting singleton/near-singleton groups purely because a single internal metric peaked there, with no alternative considered and no sanity check on sizes or file shape.
- **Consequence**: The saved labels do not match the expected partition (grader compares cluster structure/agreement), so the output file is marked WRONG even though the pipeline ran without error.
329Distribution statistic computed over the wrong grouping/slice (wrong axis or subset)taskinfiagent-dabench
Applies when
task -- the task asks for a per-entity distributional statistic (skewness, variance, kurtosis, etc.) restricted to a stated slice (a year, a category, a filter), so the analyst must decide which rows form each entity's distribution.
Pattern
The attempt picks a plausible-but-unstated grouping — e.g. aggregating across all periods instead of the stated one, computing the statistic across entities rather than within each entity, applying it along the wrong DataFrame axis, or collapsing duplicate/sub-rows before the statistic — and reports the arg-max of that different quantity. Often the code isn't saved, so the grouping cannot be inspected at all.
Detection procedure
  1. From the task statement, write down explicitly: the unit of the returned answer (one entity name), the slice that filters rows, and which column(s) supply the values whose distribution is measured.
  2. In the scripts, locate the exact filter and groupby/axis argument used before the statistic call; check that the filter matches the stated slice and that the grouping key matches the answer unit.
  3. Check that each group actually has more than 2 non-null values (skewness/kurtosis are undefined or degenerate otherwise) and print group sizes; a group of size 1–2 or an all-NaN group silently producing a value or NaN is a red flag.
  4. Confirm the reported entity is the arg-max of the statistic for the correct slice, and that the definition flag (e.g. Fisher vs Pearson, bias correction, NaN policy) matches the constraint; if no script is available to verify any of this, treat the answer as unsupported.
Discriminator
A real violation is when the code's filter/grouping/axis provably differs from the task's stated slice-and-unit, or when the grouping is unverifiable because no code/intermediate output exists. It is not a violation if the data layout makes an alternative reading equivalent (e.g. one row per entity-period so the grouping collapses identically) and the script prints the group sizes and top-k values showing the intended slice was used.
Consequence
The reported entity is the arg-max of a different statistic than requested, so the single expected string mismatches and the grader scores 0/1 with no partial credit.
id d741f8c10591 · mined from infiagent-dabench dabench-252@s6
raw text (what the judge reads)
### Distribution statistic computed over the wrong grouping/slice (wrong axis or subset)
- **Applies when**: `task` -- the task asks for a per-entity distributional statistic (skewness, variance, kurtosis, etc.) restricted to a stated slice (a year, a category, a filter), so the analyst must decide which rows form each entity's distribution.
- **Pattern**: The attempt picks a plausible-but-unstated grouping — e.g. aggregating across all periods instead of the stated one, computing the statistic across entities rather than within each entity, applying it along the wrong DataFrame axis, or collapsing duplicate/sub-rows before the statistic — and reports the arg-max of that different quantity. Often the code isn't saved, so the grouping cannot be inspected at all.
- **Detection procedure**:
  1. From the task statement, write down explicitly: the unit of the returned answer (one entity name), the slice that filters rows, and which column(s) supply the values whose distribution is measured.
  2. In the scripts, locate the exact filter and `groupby`/axis argument used before the statistic call; check that the filter matches the stated slice and that the grouping key matches the answer unit.
  3. Check that each group actually has more than 2 non-null values (skewness/kurtosis are undefined or degenerate otherwise) and print group sizes; a group of size 1–2 or an all-NaN group silently producing a value or NaN is a red flag.
  4. Confirm the reported entity is the arg-max of the statistic for the correct slice, and that the definition flag (e.g. Fisher vs Pearson, bias correction, NaN policy) matches the constraint; if no script is available to verify any of this, treat the answer as unsupported.
- **Discriminator**: A real violation is when the code's filter/grouping/axis provably differs from the task's stated slice-and-unit, or when the grouping is unverifiable because no code/intermediate output exists. It is *not* a violation if the data layout makes an alternative reading equivalent (e.g. one row per entity-period so the grouping collapses identically) and the script prints the group sizes and top-k values showing the intended slice was used.
- **Consequence**: The reported entity is the arg-max of a different statistic than requested, so the single expected string mismatches and the grader scores 0/1 with no partial credit.
330Assumed output schema instead of reading the provided reference/template filetaskda-code
Applies when
task -- the task points to an existing sample/expected-output file (or explicitly states a required format) that the deliverable must match, and the scripts generate the deliverable programmatically.
Pattern
The attempt never loads or prints the reference file; instead it hardcodes column headers, label spellings, row order, rounding, and inclusion of rows from guesswork or from what happened to appear in the joined data, then writes the file and declares success. Equally, the grouped statistic's key set (rows/periods included in the mean) is taken from the observed groups rather than from the definition implied by the reference, so absent groups are silently dropped rather than treated as zeros or a fixed denominator.
Detection procedure
  1. Read the task for any mention of a sample/expected output file or a stated format constraint (headers, units, rounding, ordering).
  2. Search the scripts for a read of that reference file and a comparison of columns/row keys/dtypes/rounding against the produced output; note any literal lists of labels or header names written by hand.
  3. Inspect how the aggregated statistic's denominator is formed — does it come from groupby(...).mean() over only observed keys, or from an explicitly constructed complete key set (e.g., full period range × all groups) justified by the task/reference?
  4. Check the final answer for tell-tale signs of an unvalidated schema: header text and label strings not verified, arbitrary rounding, and no printed diff/assertion against the reference.
Discriminator
A real violation is when no code path ever reads or asserts against the reference file (or the stated format) and the key/denominator set is implicitly whatever the join produced. It is fine if the script loads the reference, asserts column names/order/row keys match (or reindexes the result to the reference keys, filling missing groups explicitly) even if some literals also appear in the code.
Consequence
The output file is byte-comparable-wrong — mismatched headers/labels/row set, or numerically off because groups/periods with no observations were excluded from the mean — so the file check fails (0/1) even though the pipeline logic looks reasonable.
id ff4834589571 · mined from da-code dacode-dm-csv-010@s6
raw text (what the judge reads)
### Assumed output schema instead of reading the provided reference/template file
- **Applies when**: `task` -- the task points to an existing sample/expected-output file (or explicitly states a required format) that the deliverable must match, and the scripts generate the deliverable programmatically.
- **Pattern**: The attempt never loads or prints the reference file; instead it hardcodes column headers, label spellings, row order, rounding, and inclusion of rows from guesswork or from what happened to appear in the joined data, then writes the file and declares success. Equally, the grouped statistic's key set (rows/periods included in the mean) is taken from the observed groups rather than from the definition implied by the reference, so absent groups are silently dropped rather than treated as zeros or a fixed denominator.
- **Detection procedure**:
  1. Read the task for any mention of a sample/expected output file or a stated format constraint (headers, units, rounding, ordering).
  2. Search the scripts for a read of that reference file and a comparison of columns/row keys/dtypes/rounding against the produced output; note any literal lists of labels or header names written by hand.
  3. Inspect how the aggregated statistic's denominator is formed — does it come from `groupby(...).mean()` over only observed keys, or from an explicitly constructed complete key set (e.g., full period range × all groups) justified by the task/reference?
  4. Check the final answer for tell-tale signs of an unvalidated schema: header text and label strings not verified, arbitrary rounding, and no printed diff/assertion against the reference.
- **Discriminator**: A real violation is when no code path ever reads or asserts against the reference file (or the stated format) and the key/denominator set is implicitly whatever the join produced. It is fine if the script loads the reference, asserts column names/order/row keys match (or reindexes the result to the reference keys, filling missing groups explicitly) even if some literals also appear in the code.
- **Consequence**: The output file is byte-comparable-wrong — mismatched headers/labels/row set, or numerically off because groups/periods with no observations were excluded from the mean — so the file check fails (0/1) even though the pipeline logic looks reasonable.
331Truncating/reformatting an identifier value so information is losttaskinfiagent-dabench
Applies when
task -- the answer requires reporting a key value (a date, ID, label, or category) that identifies a specific row/record found by an argmax/argmin or lookup, and the answer template shows a shorter or coarser rendering than the underlying data.
Pattern
The agent finds the correct record but then reshapes the identifier to match a format hint literally (e.g., chopping components off a timestamp, dropping precision, lowercasing/renaming a label), emitting a value that no longer uniquely designates the record it found; the dependent numeric result is still computed from the full-precision value, so the two reported fields are internally inconsistent in granularity.
Detection procedure
  1. From the task, note the granularity of the identifier the computation actually operates on (the row selected, and the row used as its neighbour/predecessor for any derived quantity).
  2. In the scripts, find where the identifier is converted to a string for output; check for slicing, strftime-style reformatting, rounding, or manual literal typing that drops components.
  3. Compare the emitted identifier with the value used internally: if the derived statistic required finer granularity than the emitted identifier expresses, flag it.
  4. Prefer reporting the identifier exactly as stored in the data (full precision), even if a format hint appears coarser; only coarsen when the task explicitly asks for an aggregate at that coarser level.
Discriminator
A real violation is when the coarsened identifier maps to many candidate records and the rest of the answer was computed at the finer level. It is fine when the task genuinely aggregates to the coarser unit (e.g., the maximum is computed over monthly aggregates) so the coarse label is the true unit of analysis.
Consequence
The numeric field matches but the identifier field fails exact-string comparison, so the grader marks the submission partially/fully incorrect despite correct computation.
id c0afd84020dd · mined from infiagent-dabench dabench-572@s6
raw text (what the judge reads)
### Truncating/reformatting an identifier value so information is lost
- **Applies when**: `task` -- the answer requires reporting a key value (a date, ID, label, or category) that identifies a specific row/record found by an argmax/argmin or lookup, and the answer template shows a shorter or coarser rendering than the underlying data.
- **Pattern**: The agent finds the correct record but then reshapes the identifier to match a format hint literally (e.g., chopping components off a timestamp, dropping precision, lowercasing/renaming a label), emitting a value that no longer uniquely designates the record it found; the dependent numeric result is still computed from the full-precision value, so the two reported fields are internally inconsistent in granularity.
- **Detection procedure**:
  1. From the task, note the granularity of the identifier the computation actually operates on (the row selected, and the row used as its neighbour/predecessor for any derived quantity).
  2. In the scripts, find where the identifier is converted to a string for output; check for slicing, `strftime`-style reformatting, rounding, or manual literal typing that drops components.
  3. Compare the emitted identifier with the value used internally: if the derived statistic required finer granularity than the emitted identifier expresses, flag it.
  4. Prefer reporting the identifier exactly as stored in the data (full precision), even if a format hint appears coarser; only coarsen when the task explicitly asks for an aggregate at that coarser level.
- **Discriminator**: A real violation is when the coarsened identifier maps to many candidate records and the rest of the answer was computed at the finer level. It is fine when the task genuinely aggregates to the coarser unit (e.g., the maximum is computed over monthly aggregates) so the coarse label is the true unit of analysis.
- **Consequence**: The numeric field matches but the identifier field fails exact-string comparison, so the grader marks the submission partially/fully incorrect despite correct computation.
332Fabricated/synthetic-data fallback instead of the provided datasettaskda-code
Applies when
task -- the task requires computing a statistic from a specific supplied dataset, and the script includes guessed filenames, keyword-based column matching, or else branches that generate data when the real inputs aren't found.
Pattern
The script never verifies it actually loaded the intended file/columns; when discovery fails it silently falls back to synthetic or placeholder data (often a perfect-fit or trivially constructed relationship) and writes that number to the output as if it were the real result.
Detection procedure
1. Read the task to identify the required inputs (period filter, variables, aggregation). 2. Scan the scripts for fallback branches, broad try/except, or heuristic column/file matching that can proceed without the intended data. 3. Check whether the reported number could only have come from such a fallback (e.g., exactly 0, exactly 1.0, or a round value implying a synthetic perfect fit) rather than from noisy real data. 4. Confirm the script logs/asserts the actual file name, row count, date range, and column names used.
Discriminator
A real violation is when the answer is produced by data the script itself created or by columns chosen without verification; it is fine if the script loads the actual file, prints shape/columns/filtered counts, and those match the task's stated scope — even if the code contains defensive fallbacks that never triggered.
Consequence
The written result is unrelated to the true data (here, an impossible SSR of 0 for a noisy empirical regression), so the grader's value comparison fails.
id b432f814c9cf · mined from da-code dacode-data-sa-043@s6
raw text (what the judge reads)
### Fabricated/synthetic-data fallback instead of the provided dataset
- **Applies when**: `task` -- the task requires computing a statistic from a specific supplied dataset, and the script includes guessed filenames, keyword-based column matching, or `else` branches that generate data when the real inputs aren't found.
- **Pattern**: The script never verifies it actually loaded the intended file/columns; when discovery fails it silently falls back to synthetic or placeholder data (often a perfect-fit or trivially constructed relationship) and writes that number to the output as if it were the real result.
- **Detection procedure**: 1. Read the task to identify the required inputs (period filter, variables, aggregation). 2. Scan the scripts for fallback branches, broad `try/except`, or heuristic column/file matching that can proceed without the intended data. 3. Check whether the reported number could only have come from such a fallback (e.g., exactly 0, exactly 1.0, or a round value implying a synthetic perfect fit) rather than from noisy real data. 4. Confirm the script logs/asserts the actual file name, row count, date range, and column names used.
- **Discriminator**: A real violation is when the answer is produced by data the script itself created or by columns chosen without verification; it is fine if the script loads the actual file, prints shape/columns/filtered counts, and those match the task's stated scope — even if the code contains defensive fallbacks that never triggered.
- **Consequence**: The written result is unrelated to the true data (here, an impossible SSR of 0 for a noisy empirical regression), so the grader's value comparison fails.
333Objective/threshold mismatch in imbalanced classification with asymmetric error coststaskda-code
Applies when
task -- the task states that one error type is far more costly than the other (e.g., missed detections vs. false alarms) on a rare-class prediction problem, and the script must emit hard 0/1 labels to a submission file.
Pattern
The attempt trains a default classifier and calls .predict() at the implicit 0.5 probability cutoff, then justifies the result with balanced/precision-favoring metrics (accuracy, F1, ROC-AUC) instead of the cost-weighted objective the task specifies; the predicted positive rate is never compared against the training base rate or the cost-implied optimum, and no threshold/cost search is performed on a held-out split.
Detection procedure
  1. Read the task statement and write down the stated cost/priority ordering of TP/FP/FN (or the named evaluation metric).
  2. In the scripts, locate how the hard labels are produced: is it a bare .predict()/argmax, or a threshold chosen by optimizing the stated cost/recall target on validation probabilities? Check whether the reported selection metric is the one the task asks for.
  3. Compare the count/rate of predicted positives in the output file against the positive rate in the training labels and against what the cost asymmetry implies (heavier FN cost ⇒ predicted positive rate should generally exceed the base rate, not fall below it).
  4. Check that the reported validation numbers are internally plausible and computed on data never used for fitting or threshold selection (e.g., unusually high precision+recall on a rare class is a red flag for evaluating on training rows or on a resampled set).
Discriminator
A real violation is when the label-generating cutoff is left at the library default (or tuned for a symmetric metric) despite an explicitly asymmetric objective, and no sanity check ties the predicted positive count to the expected prevalence. It is not a violation if the agent deliberately searched thresholds against the stated cost function on a clean validation split and documented why the retained cutoff yields a low positive rate.
Consequence
The submitted label column under-predicts the rare class, so the grader's cost/recall-based score (or exact-match check on the expected label file) fails even though the self-reported ROC-AUC/F1 look excellent.
id c8cf35dbaf27 · mined from da-code dacode-ml-binary-013@s6
raw text (what the judge reads)
### Objective/threshold mismatch in imbalanced classification with asymmetric error costs
- **Applies when**: `task` -- the task states that one error type is far more costly than the other (e.g., missed detections vs. false alarms) on a rare-class prediction problem, and the script must emit hard 0/1 labels to a submission file.
- **Pattern**: The attempt trains a default classifier and calls `.predict()` at the implicit 0.5 probability cutoff, then justifies the result with balanced/precision-favoring metrics (accuracy, F1, ROC-AUC) instead of the cost-weighted objective the task specifies; the predicted positive rate is never compared against the training base rate or the cost-implied optimum, and no threshold/cost search is performed on a held-out split.
- **Detection procedure**:
  1. Read the task statement and write down the stated cost/priority ordering of TP/FP/FN (or the named evaluation metric).
  2. In the scripts, locate how the hard labels are produced: is it a bare `.predict()`/argmax, or a threshold chosen by optimizing the stated cost/recall target on validation probabilities? Check whether the reported selection metric is the one the task asks for.
  3. Compare the count/rate of predicted positives in the output file against the positive rate in the training labels and against what the cost asymmetry implies (heavier FN cost ⇒ predicted positive rate should generally exceed the base rate, not fall below it).
  4. Check that the reported validation numbers are internally plausible and computed on data never used for fitting or threshold selection (e.g., unusually high precision+recall on a rare class is a red flag for evaluating on training rows or on a resampled set).
- **Discriminator**: A real violation is when the label-generating cutoff is left at the library default (or tuned for a symmetric metric) despite an explicitly asymmetric objective, and no sanity check ties the predicted positive count to the expected prevalence. It is *not* a violation if the agent deliberately searched thresholds against the stated cost function on a clean validation split and documented why the retained cutoff yields a low positive rate.
- **Consequence**: The submitted label column under-predicts the rare class, so the grader's cost/recall-based score (or exact-match check on the expected label file) fails even though the self-reported ROC-AUC/F1 look excellent.
334Analyzing a regional/partial data file when the task asks for the full populationtaskinfiagent-dabench
Applies when
task -- the question specifies a scope (e.g., "all countries", "all users", "the entire dataset") and the scripts load a single file or pre-filtered slice without checking what else is available.
Pattern
The script hard-codes one input file (or one group/subset) and computes the requested statistic on it, implicitly assuming that file covers the whole requested population; no enumeration of the data directory, no row-count/coverage sanity check, and no note reconciling the file's scope with the task's scope.
Detection procedure
  1. Read the task and write down the exact population the statistic must be computed over.
  2. In the scripts, find every data load and check whether the file name/path or any filter restricts the data to a sub-population (a region, year subset, split, category).
  3. Look for an explicit step that lists/inspects available inputs and either concatenates them or documents that the loaded file is the complete population; also check that printed record counts are plausible for the stated scope.
  4. Verify the reported answer (and any quartiles/thresholds behind it) was derived from that full population, and that the output string matches the requested answer template exactly.
Discriminator
A real violation is when a sub-population file/filter is used with no evidence it equals the required scope (other candidate files or groups plausibly exist). It is fine if the script demonstrates coverage — e.g., it lists the directory, merges all relevant files, or prints counts matching the known full population — even if only one file ends up being read.
Consequence
Quartiles and thresholds are computed on the wrong subset, so the reported items are those of the subset, not the requested population; the grader marks the answer wrong even when it partially overlaps the expected list.
id 0726712d4f9f · mined from infiagent-dabench dabench-254@s6
raw text (what the judge reads)
### Analyzing a regional/partial data file when the task asks for the full population
- **Applies when**: `task` -- the question specifies a scope (e.g., "all countries", "all users", "the entire dataset") and the scripts load a single file or pre-filtered slice without checking what else is available.
- **Pattern**: The script hard-codes one input file (or one group/subset) and computes the requested statistic on it, implicitly assuming that file covers the whole requested population; no enumeration of the data directory, no row-count/coverage sanity check, and no note reconciling the file's scope with the task's scope.
- **Detection procedure**:
  1. Read the task and write down the exact population the statistic must be computed over.
  2. In the scripts, find every data load and check whether the file name/path or any filter restricts the data to a sub-population (a region, year subset, split, category).
  3. Look for an explicit step that lists/inspects available inputs and either concatenates them or documents that the loaded file is the complete population; also check that printed record counts are plausible for the stated scope.
  4. Verify the reported answer (and any quartiles/thresholds behind it) was derived from that full population, and that the output string matches the requested answer template exactly.
- **Discriminator**: A real violation is when a sub-population file/filter is used with no evidence it equals the required scope (other candidate files or groups plausibly exist). It is fine if the script demonstrates coverage — e.g., it lists the directory, merges all relevant files, or prints counts matching the known full population — even if only one file ends up being read.
- **Consequence**: Quartiles and thresholds are computed on the wrong subset, so the reported items are those of the subset, not the requested population; the grader marks the answer wrong even when it partially overlaps the expected list.
335Prediction file contents not validated against the expected label vocabulary, row count, and formattaskda-code
Applies when
task -- the task asks for predictions on a held-out file written to a named output file with a named column, and the target is categorical (string labels).
Pattern
The agent focuses on model quality (validation accuracy, precision/recall) and asserts success, but never re-reads the written file to confirm that (a) it has exactly one row per input test row in the same order, (b) the column name matches the requested name exactly, (c) the predicted values are the exact label strings/encoding used in the source data (not integer codes, not re-cased/renamed/abbreviated variants), and (d) no stray index column, header duplication, or missing values are present.
Detection procedure
  1. From the task statement, note the required file name, column name, and the exact label vocabulary implied by the training data/README.
  2. In the scripts, find the write step: check whether labels are inverse-transformed back to their original strings, whether index=False is used, and whether the row count is tied to the raw test file (before any dropna/filtering/deduplication).
  3. Check for any post-write verification: reloading the output, printing shape, value_counts(), and comparing the unique values to the training label set and the length to len(test).
  4. In the answer, see whether the reported counts/labels are traced back to the file on disk with the source-data spelling, or only to in-memory model output.
Discriminator
A real violation is when no read-back check exists and the label values written could differ from the source vocabulary (encoded ints, custom names like "Satisfied"/"Neutral or Dissatisfied", capitalization changes) or rows could be dropped/reordered by preprocessing. A look-alike that is fine explicitly inverse-maps to the original label strings and prints the reloaded file's shape and unique values matching the raw test row count and training labels.
Consequence
The grader compares the submitted file cell-by-cell against expected labels and marks it WRONG/MISSING despite high internal validation accuracy, because label spelling, encoding, row count, or ordering does not match.
id a1b8263f67b1 · mined from da-code dacode-ml-binary-009@s6
raw text (what the judge reads)
### Prediction file contents not validated against the expected label vocabulary, row count, and format
- **Applies when**: `task` -- the task asks for predictions on a held-out file written to a named output file with a named column, and the target is categorical (string labels).
- **Pattern**: The agent focuses on model quality (validation accuracy, precision/recall) and asserts success, but never re-reads the written file to confirm that (a) it has exactly one row per input test row in the same order, (b) the column name matches the requested name exactly, (c) the predicted values are the exact label strings/encoding used in the source data (not integer codes, not re-cased/renamed/abbreviated variants), and (d) no stray index column, header duplication, or missing values are present.
- **Detection procedure**:
  1. From the task statement, note the required file name, column name, and the exact label vocabulary implied by the training data/README.
  2. In the scripts, find the write step: check whether labels are inverse-transformed back to their original strings, whether `index=False` is used, and whether the row count is tied to the raw test file (before any dropna/filtering/deduplication).
  3. Check for any post-write verification: reloading the output, printing `shape`, `value_counts()`, and comparing the unique values to the training label set and the length to `len(test)`.
  4. In the answer, see whether the reported counts/labels are traced back to the file on disk with the source-data spelling, or only to in-memory model output.
- **Discriminator**: A real violation is when no read-back check exists and the label values written could differ from the source vocabulary (encoded ints, custom names like "Satisfied"/"Neutral or Dissatisfied", capitalization changes) or rows could be dropped/reordered by preprocessing. A look-alike that is fine explicitly inverse-maps to the original label strings and prints the reloaded file's shape and unique values matching the raw test row count and training labels.
- **Consequence**: The grader compares the submitted file cell-by-cell against expected labels and marks it WRONG/MISSING despite high internal validation accuracy, because label spelling, encoding, row count, or ordering does not match.
336Using a provided reference list to filter results instead of to validate an invented metrictaskda-code
Applies when
task -- the task's spec/config file supplies expected category labels (or ordering/axis values) for an output, and the underlying quantity being plotted or ranked is not explicitly defined, so the script must choose an aggregation itself.
Pattern
The agent guesses an ad-hoc formula (e.g., an arbitrary weighted sum of counts), then subsets/reindexes its computed table to exactly the labels listed in the config, so any disagreement between its metric and the intended one is silently hidden; no independent check is made that the metric, computed over the full data, actually reproduces the given labels, their order, and any other required output artifacts.
Detection procedure
  1. Read the task and config for what is actually specified (label list, ordering, requested outputs/files) versus what the script must define itself.
  2. In the script, find where the metric is defined; check whether the formula is justified by the task/domain or appears invented (arbitrary weights, mixing counts with weighted counts, no documented meaning).
  3. Check whether the reference labels are used as a filter/reindex on the computed data or as an assertion (recompute top-N/order from the full dataset and compare to the given labels).
  4. Check the answer/outputs: does the derived ranking, computed independently, match the config's label order, and are all requested deliverable files produced?
Discriminator
A real violation is when removing the label-based filtering would change which entities/order appear (i.e., the metric alone does not reproduce the spec's labels), or when the formula has no basis in the task statement. It is fine if the config labels are merely a display/ordering convention and the script separately verifies that its independently computed top-N and ordering coincide with them, or if the metric is explicitly defined in the task.
Consequence
The plotted bar heights and any saved numeric arrays encode the wrong statistic (and the ordering check passes only cosmetically), so file-level comparisons of the chart and result arrays all fail even though the script runs without error.
id 1471e5e7c55d · mined from da-code dacode-plot-bar-006@s6
raw text (what the judge reads)
### Using a provided reference list to filter results instead of to validate an invented metric
- **Applies when**: `task` -- the task's spec/config file supplies expected category labels (or ordering/axis values) for an output, and the underlying quantity being plotted or ranked is not explicitly defined, so the script must choose an aggregation itself.
- **Pattern**: The agent guesses an ad-hoc formula (e.g., an arbitrary weighted sum of counts), then subsets/reindexes its computed table to exactly the labels listed in the config, so any disagreement between its metric and the intended one is silently hidden; no independent check is made that the metric, computed over the full data, actually reproduces the given labels, their order, and any other required output artifacts.
- **Detection procedure**:
  1. Read the task and config for what is actually specified (label list, ordering, requested outputs/files) versus what the script must define itself.
  2. In the script, find where the metric is defined; check whether the formula is justified by the task/domain or appears invented (arbitrary weights, mixing counts with weighted counts, no documented meaning).
  3. Check whether the reference labels are used as a *filter/reindex* on the computed data or as an *assertion* (recompute top-N/order from the full dataset and compare to the given labels).
  4. Check the answer/outputs: does the derived ranking, computed independently, match the config's label order, and are all requested deliverable files produced?
- **Discriminator**: A real violation is when removing the label-based filtering would change which entities/order appear (i.e., the metric alone does not reproduce the spec's labels), or when the formula has no basis in the task statement. It is fine if the config labels are merely a display/ordering convention and the script separately verifies that its independently computed top-N and ordering coincide with them, or if the metric is explicitly defined in the task.
- **Consequence**: The plotted bar heights and any saved numeric arrays encode the wrong statistic (and the ordering check passes only cosmetically), so file-level comparisons of the chart and result arrays all fail even though the script runs without error.
337Unvalidated threshold statistic: no cross-check of the estimator convention, data source, or plausibility of the flagged counttaskinfiagent-dabench
Applies when
task -- the task asks for a count/flagging of rows via a distribution-based rule (z-score, IQR, percentile, threshold) computed from a column's mean/std or quantiles, and the scripts produce a single number from one library call.
Pattern
The agent runs one implementation (e.g., a single library's z-score with its default ddof/NaN policy) on one file/column it picked without verifying that this is the intended source column, that non-numeric/sentinel/missing entries were handled, or that the resulting count is consistent with the printed distribution; every "verification" script re-runs the same code path and therefore reproduces the same number.
Detection procedure
  1. Read the task for the exact rule, threshold and target column; note which choices are ambiguous (population vs sample standard deviation, NaN dropping, which file/split, whether the column is truly numeric or contains coded/placeholder values).
  2. Read the scripts: check whether the flagging statistic is computed by more than one independent route (e.g., manual (x - mean)/std with both ddof=0 and ddof=1, or a second column/file candidate) and whether column dtype, unique values and missing-value counts were inspected before thresholding.
  3. Check whether the reported count is reconciled with the printed summary statistics (max |z|, min/max, mean, std): does the count of values beyond the threshold match what the reported extremes imply, and is the fraction flagged plausible for the rule (a 3σ rule on a roughly bell-shaped column should flag ≲0.3% of rows; a much larger share signals a skewed/degenerate/mis-typed column or wrong std convention).
  4. Confirm the reported answer is the count actually requested (and in the requested format), not an intermediate or subset-specific count.
Discriminator
A real violation is when the count comes from a single unreplicated code path and no reconciliation with the distribution or with an alternative convention/data source was performed — the answer could flip under a defensible alternative. It is not a violation if the agent showed that the two conventions and NaN handling give the same count, verified the column choice/dtype, and the flagged fraction and extreme values are consistent with the printed summary.
Consequence
The submitted count differs from the reference count (often by a large margin, e.g., a nonzero count where the correct rule yields none), so the single graded field fails and the whole task scores 0.
id b568272a3741 · mined from infiagent-dabench dabench-361@s6
raw text (what the judge reads)
### Unvalidated threshold statistic: no cross-check of the estimator convention, data source, or plausibility of the flagged count
- **Applies when**: `task` -- the task asks for a count/flagging of rows via a distribution-based rule (z-score, IQR, percentile, threshold) computed from a column's mean/std or quantiles, and the scripts produce a single number from one library call.
- **Pattern**: The agent runs one implementation (e.g., a single library's z-score with its default `ddof`/NaN policy) on one file/column it picked without verifying that this is the intended source column, that non-numeric/sentinel/missing entries were handled, or that the resulting count is consistent with the printed distribution; every "verification" script re-runs the same code path and therefore reproduces the same number.
- **Detection procedure**:
  1. Read the task for the exact rule, threshold and target column; note which choices are ambiguous (population vs sample standard deviation, NaN dropping, which file/split, whether the column is truly numeric or contains coded/placeholder values).
  2. Read the scripts: check whether the flagging statistic is computed by more than one independent route (e.g., manual `(x - mean)/std` with both `ddof=0` and `ddof=1`, or a second column/file candidate) and whether column dtype, unique values and missing-value counts were inspected before thresholding.
  3. Check whether the reported count is reconciled with the printed summary statistics (max |z|, min/max, mean, std): does the count of values beyond the threshold match what the reported extremes imply, and is the fraction flagged plausible for the rule (a 3σ rule on a roughly bell-shaped column should flag ≲0.3% of rows; a much larger share signals a skewed/degenerate/mis-typed column or wrong std convention).
  4. Confirm the reported answer is the count actually requested (and in the requested format), not an intermediate or subset-specific count.
- **Discriminator**: A real violation is when the count comes from a single unreplicated code path and no reconciliation with the distribution or with an alternative convention/data source was performed — the answer could flip under a defensible alternative. It is *not* a violation if the agent showed that the two conventions and NaN handling give the same count, verified the column choice/dtype, and the flagged fraction and extreme values are consistent with the printed summary.
- **Consequence**: The submitted count differs from the reference count (often by a large margin, e.g., a nonzero count where the correct rule yields none), so the single graded field fails and the whole task scores 0.
338Missing required output artifact (answer only in chat, no reproducible script)taskda-code
Applies when
task -- the task statement or README specifies deliverable file(s) (e.g., a named JSON/CSV result file) and/or requires a derivation from an auxiliary spec file, and the agent responds with a value only in its final message.
Pattern
The attempt reports a plausible-looking formatted answer inline but never writes the named result file to disk (and often keeps no script), so nothing reproduces the number and the required artifact is absent or stale.
Detection procedure
  1. Read the task/README and list every named output artifact, its path, and its required key/field names and value types.
  2. Inspect the saved scripts/notebook for code that (a) loads the auxiliary spec/mapping resource, (b) computes the requested statistic, and (c) explicitly serializes the result to each named artifact path.
  3. If no script exists, or no write-to-file step exists for a named artifact, flag the attempt regardless of whether the inline value looks reasonable.
  4. If a file is written, open it and confirm keys, nesting, and rounding/units match the requested format exactly.
Discriminator
A real violation is the absence of a persisted, script-generated artifact matching the required name/schema; it is not a violation if the artifact is correctly written and the inline message merely restates its contents, nor if the task genuinely asks only for a text answer with no file named anywhere.
Consequence
The grader looks up the expected result file, finds it missing (or not matching schema), and scores 0 even if the reported number were right — and with no script the value cannot be verified or corrected.
id a591c2f05e46 · mined from da-code dacode-di-text-004@s6
raw text (what the judge reads)
### Missing required output artifact (answer only in chat, no reproducible script)
- **Applies when**: `task` -- the task statement or README specifies deliverable file(s) (e.g., a named JSON/CSV result file) and/or requires a derivation from an auxiliary spec file, and the agent responds with a value only in its final message.
- **Pattern**: The attempt reports a plausible-looking formatted answer inline but never writes the named result file to disk (and often keeps no script), so nothing reproduces the number and the required artifact is absent or stale.
- **Detection procedure**:
  1. Read the task/README and list every named output artifact, its path, and its required key/field names and value types.
  2. Inspect the saved scripts/notebook for code that (a) loads the auxiliary spec/mapping resource, (b) computes the requested statistic, and (c) explicitly serializes the result to each named artifact path.
  3. If no script exists, or no write-to-file step exists for a named artifact, flag the attempt regardless of whether the inline value looks reasonable.
  4. If a file is written, open it and confirm keys, nesting, and rounding/units match the requested format exactly.
- **Discriminator**: A real violation is the absence of a persisted, script-generated artifact matching the required name/schema; it is *not* a violation if the artifact is correctly written and the inline message merely restates its contents, nor if the task genuinely asks only for a text answer with no file named anywhere.
- **Consequence**: The grader looks up the expected result file, finds it missing (or not matching schema), and scores 0 even if the reported number were right — and with no script the value cannot be verified or corrected.
339Submission schema not verified against the provided templatetaskda-code
Applies when
task -- The task requires writing predictions to an output file whose format is defined by a provided sample/template file (column names, id casing, row count, ordering).
Pattern
The attempt builds the output from its own assumptions (hand-typed header, own column spelling/case, its own row order or a partial/truncated set of rows) and never loads the template file to confirm the header text, the id values, or the number of rows matches; the answer is also presented as pasted text rather than a verified saved file.
Detection procedure
  1. Read the task/README for the stated output spec and note that a template file defines it exactly.
  2. In the scripts, look for an explicit read of the template file and a comparison of the produced frame's columns (exact strings, including case) and its id column (same values, same order, same length) against it.
  3. Check whether the ids come from the official test input and whether every test row appears exactly once, with no dropped/filtered/deduplicated rows.
  4. Inspect the reported answer's header and row count against the template's header and the test set size; treat hand-written headers or a row count that differs (or output shown only partially with no file-level check) as unverified.
Discriminator
A real violation is a mismatch in header strings/case, extra or missing columns, missing/reordered/duplicated ids, or absence of any programmatic check against the template. It is fine if the script derives columns and ids directly from the template/test file and asserts shape and equality, even if the printed preview is short.
Consequence
The grader cannot parse or align the file with the expected schema and marks the submission WRONG/MISSING regardless of predictive quality.
id 4f343387a134 · mined from da-code dacode-ml-competition-003@s6
raw text (what the judge reads)
### Submission schema not verified against the provided template
- **Applies when**: `task` -- The task requires writing predictions to an output file whose format is defined by a provided sample/template file (column names, id casing, row count, ordering).
- **Pattern**: The attempt builds the output from its own assumptions (hand-typed header, own column spelling/case, its own row order or a partial/truncated set of rows) and never loads the template file to confirm the header text, the id values, or the number of rows matches; the answer is also presented as pasted text rather than a verified saved file.
- **Detection procedure**:
  1. Read the task/README for the stated output spec and note that a template file defines it exactly.
  2. In the scripts, look for an explicit read of the template file and a comparison of the produced frame's columns (exact strings, including case) and its id column (same values, same order, same length) against it.
  3. Check whether the ids come from the official test input and whether every test row appears exactly once, with no dropped/filtered/deduplicated rows.
  4. Inspect the reported answer's header and row count against the template's header and the test set size; treat hand-written headers or a row count that differs (or output shown only partially with no file-level check) as unverified.
- **Discriminator**: A real violation is a mismatch in header strings/case, extra or missing columns, missing/reordered/duplicated ids, or absence of any programmatic check against the template. It is fine if the script derives columns and ids directly from the template/test file and asserts shape and equality, even if the printed preview is short.
- **Consequence**: The grader cannot parse or align the file with the expected schema and marks the submission WRONG/MISSING regardless of predictive quality.
340Ignoring the provided output template when producing the deliverable filetaskda-code
Applies when
task -- the task says results must be saved to a named file whose layout matches a supplied template/example file, and the scripts build the output table from scratch.
Pattern
The agent never opens or parses the template; it invents its own index/column labels, offsets (e.g., period counters starting at 1 instead of 0), date/label formatting, column ordering, and rounding, then asserts the output "matches the format" without any comparison. Self-checks in a verify script only recompute the agent's own numbers against its own expectations, so they always pass.
Detection procedure
  1. Read the task for any mention of a template, sample, or "format must match" requirement, and note the file name(s) of the expected deliverable.
  2. Scan the scripts for a read of that template file (e.g., loading it and inspecting its header row, index labels, dtype/format of labels, number of columns, rounding/precision).
  3. If no such read exists, compare the answer's header/index conventions to any plausible template conventions the task hints at (starting index value, label formatting, precision) — unexplained choices made unilaterally by the agent are the flag.
  4. Confirm the verification step compares against the template (shape, column names, index values, rounding), not merely against the agent's own recomputation.
Discriminator
Fine if the scripts explicitly load the template and align columns/index/rounding to it (or reproduce it exactly and only fill values); a violation is when format decisions (label naming, offset base, precision, orientation) are made by the agent with no reference to the template, even if the underlying aggregation logic is defensible.
Consequence
The numeric aggregation may be conceptually reasonable, but the saved file's headers/index/offsets/precision differ from the reference, so an exact- or column-wise comparison by the grader marks the file WRONG/MISSING and the task scores 0.
id d38e2f799d19 · mined from da-code dacode-dm-csv-044@s6
raw text (what the judge reads)
### Ignoring the provided output template when producing the deliverable file
- **Applies when**: `task` -- the task says results must be saved to a named file whose layout matches a supplied template/example file, and the scripts build the output table from scratch.
- **Pattern**: The agent never opens or parses the template; it invents its own index/column labels, offsets (e.g., period counters starting at 1 instead of 0), date/label formatting, column ordering, and rounding, then asserts the output "matches the format" without any comparison. Self-checks in a verify script only recompute the agent's own numbers against its own expectations, so they always pass.
- **Detection procedure**:
  1. Read the task for any mention of a template, sample, or "format must match" requirement, and note the file name(s) of the expected deliverable.
  2. Scan the scripts for a read of that template file (e.g., loading it and inspecting its header row, index labels, dtype/format of labels, number of columns, rounding/precision).
  3. If no such read exists, compare the answer's header/index conventions to any plausible template conventions the task hints at (starting index value, label formatting, precision) — unexplained choices made unilaterally by the agent are the flag.
  4. Confirm the verification step compares against the template (shape, column names, index values, rounding), not merely against the agent's own recomputation.
- **Discriminator**: Fine if the scripts explicitly load the template and align columns/index/rounding to it (or reproduce it exactly and only fill values); a violation is when format decisions (label naming, offset base, precision, orientation) are made by the agent with no reference to the template, even if the underlying aggregation logic is defensible.
- **Consequence**: The numeric aggregation may be conceptually reasonable, but the saved file's headers/index/offsets/precision differ from the reference, so an exact- or column-wise comparison by the grader marks the file WRONG/MISSING and the task scores 0.
341Model selection and validation done with a metric other than the competition's stated evaluation metrictaskda-code
Applies when
task -- The task states an explicit scoring metric (e.g., a weighted/ordinal agreement or rank metric) and the scripts fit several candidate models and pick/tune one before writing predictions.
Pattern
The scripts import or compute the stated metric only incidentally (or not at all) and instead rank models, choose hyperparameters, and weight ensembles by a default proxy such as plain accuracy — often on the training set rather than held-out folds — so the final chosen model is optimized for the wrong objective and the ordinal/cost structure of the target is ignored (no thresholding/rounding of ordinal outputs, no optimization of cut points).
Detection procedure
  1. Read the task/README and record the exact metric named for scoring.
  2. Grep the scripts for that metric's implementation or scorer; check whether cross_val_score/GridSearchCV/model-comparison printouts use it, or use a default like scoring='accuracy' / .score().
  3. Check whether the reported comparison numbers come from held-out folds or from predict(X_train) on the fitted training data.
  4. Inspect the final answer for signs the objective was never targeted: e.g., predicted class distribution collapsed onto the majority classes with rare/extreme classes almost absent, which the stated metric penalizes heavily.
Discriminator
A real violation is when no selection decision (model, weights, thresholds) is ever evaluated under the stated metric; it is fine if a proxy is used for speed but the finalists are compared, and the final choice made, under the stated metric on out-of-fold predictions (or if the proxy is provably monotone with it).
Consequence
The submission has valid format but a low or near-chance score under the official metric — an ordinal-aware or metric-optimized model would score materially higher, so the answer fails the grader's threshold.
id 3989bde54f4e · mined from da-code dacode-ml-competition-006@s6
raw text (what the judge reads)
### Model selection and validation done with a metric other than the competition's stated evaluation metric
- **Applies when**: `task` -- The task states an explicit scoring metric (e.g., a weighted/ordinal agreement or rank metric) and the scripts fit several candidate models and pick/tune one before writing predictions.
- **Pattern**: The scripts import or compute the stated metric only incidentally (or not at all) and instead rank models, choose hyperparameters, and weight ensembles by a default proxy such as plain accuracy — often on the *training* set rather than held-out folds — so the final chosen model is optimized for the wrong objective and the ordinal/cost structure of the target is ignored (no thresholding/rounding of ordinal outputs, no optimization of cut points).
- **Detection procedure**:
  1. Read the task/README and record the exact metric named for scoring.
  2. Grep the scripts for that metric's implementation or scorer; check whether `cross_val_score`/`GridSearchCV`/model-comparison printouts use it, or use a default like `scoring='accuracy'` / `.score()`.
  3. Check whether the reported comparison numbers come from held-out folds or from `predict(X_train)` on the fitted training data.
  4. Inspect the final answer for signs the objective was never targeted: e.g., predicted class distribution collapsed onto the majority classes with rare/extreme classes almost absent, which the stated metric penalizes heavily.
- **Discriminator**: A real violation is when *no* selection decision (model, weights, thresholds) is ever evaluated under the stated metric; it is fine if a proxy is used for speed but the finalists are compared, and the final choice made, under the stated metric on out-of-fold predictions (or if the proxy is provably monotone with it).
- **Consequence**: The submission has valid format but a low or near-chance score under the official metric — an ordinal-aware or metric-optimized model would score materially higher, so the answer fails the grader's threshold.
342Unvalidated data loading and silent row-dropping that changes group membershiptaskinfiagent-dabench
Applies when
task -- the task asks to split rows into groups by whether a field is missing/null and compare a summary statistic per group, and the script loads a delimited file and filters rows before aggregating.
Pattern
The script reads the file with default/assumed parsing options (separator, index column, quoting, NA tokens) and then applies extra filtering (e.g., dropna() on the measured column, or relying on isnull() to capture sentinel/empty-string/"NA"-like markers) without ever verifying that the parsed row/column counts, dtypes, and per-group sizes are consistent with the raw file. Group sizes and therefore group means silently shift, and no sanity check is printed or cross-checked against the total row count.
Detection procedure
  1. Read the task to identify exactly which rows belong to each group and which rows the requested statistic must be computed over (all rows in the group, or only a filtered subset).
  2. In the script, locate the load call and note every implicit assumption (delimiter, index_col, quote handling, default NA strings); check whether the script inspects raw file structure (header, a few raw lines, expected column count) or just trusts the parse.
  3. Check every filtering/dropna/notnull step applied after loading and verify the two group sizes sum to the total parsed rows, and that the parsed row count matches the raw file line count; also verify the measured column's dtype is numeric rather than object coerced from mis-split fields.
  4. Confirm the reported means are recomputed on the groups as defined by the task, and that a range/count sanity check (group n, min/max, no unexpected NaNs) was performed before reporting.
Discriminator
A real violation is when parsing options are assumed and/or rows are dropped from the statistic's denominator with no verification that counts/dtypes match the raw source. It is not a violation if the script explicitly validates the parse (row/column counts, dtype, NA representation) and either keeps all group rows or justifies the exclusion because the task itself requires it.
Consequence
Both group means (and the test statistic) are computed on a subtly wrong subset, so the reported numbers are close to but not equal to the expected values and every numeric check fails, even though the p-value still looks "significant" and nothing errors out.
id 213681646873 · mined from infiagent-dabench dabench-297@s6
raw text (what the judge reads)
### Unvalidated data loading and silent row-dropping that changes group membership
- **Applies when**: `task` -- the task asks to split rows into groups by whether a field is missing/null and compare a summary statistic per group, and the script loads a delimited file and filters rows before aggregating.
- **Pattern**: The script reads the file with default/assumed parsing options (separator, index column, quoting, NA tokens) and then applies extra filtering (e.g., `dropna()` on the measured column, or relying on `isnull()` to capture sentinel/empty-string/"NA"-like markers) without ever verifying that the parsed row/column counts, dtypes, and per-group sizes are consistent with the raw file. Group sizes and therefore group means silently shift, and no sanity check is printed or cross-checked against the total row count.
- **Detection procedure**:
  1. Read the task to identify exactly which rows belong to each group and which rows the requested statistic must be computed over (all rows in the group, or only a filtered subset).
  2. In the script, locate the load call and note every implicit assumption (delimiter, `index_col`, quote handling, default NA strings); check whether the script inspects raw file structure (header, a few raw lines, expected column count) or just trusts the parse.
  3. Check every filtering/`dropna`/`notnull` step applied after loading and verify the two group sizes sum to the total parsed rows, and that the parsed row count matches the raw file line count; also verify the measured column's dtype is numeric rather than object coerced from mis-split fields.
  4. Confirm the reported means are recomputed on the groups as defined by the task, and that a range/count sanity check (group n, min/max, no unexpected NaNs) was performed before reporting.
- **Discriminator**: A real violation is when parsing options are assumed and/or rows are dropped from the statistic's denominator with no verification that counts/dtypes match the raw source. It is *not* a violation if the script explicitly validates the parse (row/column counts, dtype, NA representation) and either keeps all group rows or justifies the exclusion because the task itself requires it.
- **Consequence**: Both group means (and the test statistic) are computed on a subtly wrong subset, so the reported numbers are close to but not equal to the expected values and every numeric check fails, even though the p-value still looks "significant" and nothing errors out.
343Deliverables and rule-spec taken from guesswork instead of the provided instruction filetaskda-code
Applies when
task -- the task points to an external guidance/spec document and expects a set of saved artifacts (figures, serialized arrays/JSON, tables) plus derived filtering/grouping rules.
Pattern
The script saves only the one artifact named explicitly in the prompt (e.g., the image) and hard-codes filtering/category rules that the agent inferred, visible as inline comments debating the meaning of the rule ("re-reading: this probably means…"), rather than quoting and implementing the spec verbatim and emitting every artifact the spec/harness expects.
Detection procedure
  1. Read the task and enumerate every output file/object that must exist and every stated constraint (filters, thresholds, ordering, sizes, colors, rounding).
  2. Grep the scripts for a write/save call per enumerated artifact (savefig, np.save, to_json, json.dump, …); flag any artifact with no corresponding write, and flag the absence of any attempt to open/parse the referenced guidance file.
  3. Check whether the script's rule definitions are traceable to quoted spec text; flag comments showing the agent chose among competing interpretations, or ad-hoc keyword/priority orderings not present in the spec.
  4. Compare the answer's reported numbers/labels to what the spec implies (e.g., whether the reported group is the one that was asked for, and whether counts came before or after the required filter).
Discriminator
A real violation is missing save calls for required artifacts, or rule logic invented by the agent with no textual anchor; it is not a violation if the script reads the spec, implements it literally, and writes all requested artifacts even if it additionally saves extra intermediate files or documents a genuinely unambiguous simplification.
Consequence
The grader finds expected files absent or containing different values ("WRONG/MISSING") for every artifact, so all checks fail even though the narrative answer looks self-consistent.
id f8e290b8dd5a · mined from da-code dacode-plot-pie-005@s6
raw text (what the judge reads)
### Deliverables and rule-spec taken from guesswork instead of the provided instruction file
- **Applies when**: `task` -- the task points to an external guidance/spec document and expects a set of saved artifacts (figures, serialized arrays/JSON, tables) plus derived filtering/grouping rules.
- **Pattern**: The script saves only the one artifact named explicitly in the prompt (e.g., the image) and hard-codes filtering/category rules that the agent inferred, visible as inline comments debating the meaning of the rule ("re-reading: this probably means…"), rather than quoting and implementing the spec verbatim and emitting every artifact the spec/harness expects.
- **Detection procedure**:
  1. Read the task and enumerate every output file/object that must exist and every stated constraint (filters, thresholds, ordering, sizes, colors, rounding).
  2. Grep the scripts for a write/save call per enumerated artifact (`savefig`, `np.save`, `to_json`, `json.dump`, …); flag any artifact with no corresponding write, and flag the absence of any attempt to open/parse the referenced guidance file.
  3. Check whether the script's rule definitions are traceable to quoted spec text; flag comments showing the agent chose among competing interpretations, or ad-hoc keyword/priority orderings not present in the spec.
  4. Compare the answer's reported numbers/labels to what the spec implies (e.g., whether the reported group is the one that was asked for, and whether counts came before or after the required filter).
- **Discriminator**: A real violation is missing save calls for required artifacts, or rule logic invented by the agent with no textual anchor; it is *not* a violation if the script reads the spec, implements it literally, and writes all requested artifacts even if it additionally saves extra intermediate files or documents a genuinely unambiguous simplification.
- **Consequence**: The grader finds expected files absent or containing different values ("WRONG/MISSING") for every artifact, so all checks fail even though the narrative answer looks self-consistent.
344Output template file provided by the task was never inspected or followedtaskda-code
Applies when
task -- The task points to an example/sample output file (or explicitly states a required schema, column names, row labels, or ordering) that the deliverable must match.
Pattern
The scripts never read or print the referenced sample/template file; instead the agent invents its own column headers, label wording, extra/blank columns, and row ordering, then writes the deliverable directly from that guess.
Detection procedure
  1. Read the task and list every stated formatting constraint and every referenced example file.
  2. Search the scripts for any load/print of that example file (or an explicit hard-coded schema demonstrably copied from it); note the exact headers and label strings actually written out.
  3. Compare the written headers/labels/row order against the referenced format; flag if they are self-invented, contain placeholder or empty columns, or use paraphrased category names.
  4. Check the final answer text for the same mismatch (e.g., unrelated header names, blank trailing column).
Discriminator
A real violation is when the template exists in the task/workspace and the scripts show no evidence of having read or reproduced it, so header/label/column-count agreement is pure luck; it is fine if the agent read the template (or the task fully specifies the schema in text) and the output columns, label wording, and ordering demonstrably match — even if the values happen to be produced by other means.
Consequence
The grader's file comparison fails on schema/label mismatch (wrong column names, extra empty column, non-matching category strings) and marks the result file WRONG even when the underlying computed values are correct.
id 669b6dbfc237 · mined from da-code dacode-dm-csv-015@s6
raw text (what the judge reads)
### Output template file provided by the task was never inspected or followed
- **Applies when**: `task` -- The task points to an example/sample output file (or explicitly states a required schema, column names, row labels, or ordering) that the deliverable must match.
- **Pattern**: The scripts never read or print the referenced sample/template file; instead the agent invents its own column headers, label wording, extra/blank columns, and row ordering, then writes the deliverable directly from that guess.
- **Detection procedure**:
  1. Read the task and list every stated formatting constraint and every referenced example file.
  2. Search the scripts for any load/print of that example file (or an explicit hard-coded schema demonstrably copied from it); note the exact headers and label strings actually written out.
  3. Compare the written headers/labels/row order against the referenced format; flag if they are self-invented, contain placeholder or empty columns, or use paraphrased category names.
  4. Check the final answer text for the same mismatch (e.g., unrelated header names, blank trailing column).
- **Discriminator**: A real violation is when the template exists in the task/workspace and the scripts show no evidence of having read or reproduced it, so header/label/column-count agreement is pure luck; it is fine if the agent read the template (or the task fully specifies the schema in text) and the output columns, label wording, and ordering demonstrably match — even if the values happen to be produced by other means.
- **Consequence**: The grader's file comparison fails on schema/label mismatch (wrong column names, extra empty column, non-matching category strings) and marks the result file WRONG even when the underlying computed values are correct.
345Model error never sanity-checked against target scale / baseline variance (masking a preprocessing or split error)taskinfiagent-dabench
Applies when
task -- a task asks for a supervised model's error metric (MSE/RMSE/MAE) on a held-out split with prescribed preprocessing (e.g., mean-imputation of specified columns) and a fixed train/test ratio.
Pattern
The attempt runs a pipeline whose preprocessing silently deviates from the spec (rows dropped instead of imputed, non-numeric/parsed-badly columns coerced or left as strings, imputation done per-split or after filtering, features/target misaligned or unscaled outliers left in a skewed monetary column), reports the resulting metric verbatim, and never compares it to the target's own spread or to a trivial mean-predictor baseline that would expose the error as implausibly large (or suspiciously small).
Detection procedure
  1. From the task, list the mandated preprocessing steps, the split fraction, and the exact metric definition; note the plausible range/std of the target variable from the data description.
  2. In the scripts, check line by line that each mandated step is applied to the full column set before splitting (imputation with column means, numeric coercion of any text/currency-formatted field, no dropna or row filtering that removes records the spec says to impute), and that X/y come from the same aligned frame.
  3. Check whether the script computes any reference quantity — target variance on the test split, or MSE of predicting the training mean — and compares it to the reported metric; absence of this check is the flag.
  4. Compare the reported number to that baseline yourself: if RMSE (√MSE) is on the order of, or larger than, the target's standard deviation or a large fraction of its full range, the pipeline is almost certainly broken and the answer should not have been submitted.
Discriminator
A genuine violation is an unexplained metric whose RMSE is comparable to or exceeds the target's own dispersion, or a pipeline that demonstrably skips/reorders a mandated preprocessing step; a look-alike that is fine is a large absolute MSE that is still well below the mean-predictor baseline (i.e., the model explains substantial variance) and whose preprocessing matches every stated constraint, with the magnitude simply reflecting the target's units.
Consequence
The reported metric is off by an order of magnitude from the reference value, so the single numeric check in the expected answer fails outright.
id 1646508bcaab · mined from infiagent-dabench dabench-432@s6
raw text (what the judge reads)
### Model error never sanity-checked against target scale / baseline variance (masking a preprocessing or split error)
- **Applies when**: `task` -- a task asks for a supervised model's error metric (MSE/RMSE/MAE) on a held-out split with prescribed preprocessing (e.g., mean-imputation of specified columns) and a fixed train/test ratio.
- **Pattern**: The attempt runs a pipeline whose preprocessing silently deviates from the spec (rows dropped instead of imputed, non-numeric/parsed-badly columns coerced or left as strings, imputation done per-split or after filtering, features/target misaligned or unscaled outliers left in a skewed monetary column), reports the resulting metric verbatim, and never compares it to the target's own spread or to a trivial mean-predictor baseline that would expose the error as implausibly large (or suspiciously small).
- **Detection procedure**:
  1. From the task, list the mandated preprocessing steps, the split fraction, and the exact metric definition; note the plausible range/std of the target variable from the data description.
  2. In the scripts, check line by line that each mandated step is applied to the full column set before splitting (imputation with column means, numeric coercion of any text/currency-formatted field, no `dropna` or row filtering that removes records the spec says to impute), and that X/y come from the same aligned frame.
  3. Check whether the script computes any reference quantity — target variance on the test split, or MSE of predicting the training mean — and compares it to the reported metric; absence of this check is the flag.
  4. Compare the reported number to that baseline yourself: if RMSE (√MSE) is on the order of, or larger than, the target's standard deviation or a large fraction of its full range, the pipeline is almost certainly broken and the answer should not have been submitted.
- **Discriminator**: A genuine violation is an unexplained metric whose RMSE is comparable to or exceeds the target's own dispersion, or a pipeline that demonstrably skips/reorders a mandated preprocessing step; a look-alike that is fine is a large absolute MSE that is still well below the mean-predictor baseline (i.e., the model explains substantial variance) and whose preprocessing matches every stated constraint, with the magnitude simply reflecting the target's units.
- **Consequence**: The reported metric is off by an order of magnitude from the reference value, so the single numeric check in the expected answer fails outright.
346Statistic computed on an unverified row subset (silent row loss / coercion) with no N-and-precision sanity checktaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic (correlation, mean, test statistic, p-value) over two or more columns of a table, to be reported at a fixed rounding precision.
Pattern
The attempt loads the data and calls a one-line statistic function without ever establishing which rows actually entered the computation — no report of total rows vs. rows used after NaN/blank/non-numeric coercion, no check for stray filtering, deduplication, or a partial/alternate file/sheet being read. Small differences in the effective sample silently shift the statistic in the second decimal, and the reported value is not cross-checked by any independent path. The precision constraint is also treated loosely (e.g., printing a raw float rather than the requested fixed number of decimals).
Detection procedure
  1. Read the task and note the required statistic, the required rounding, and whether any subsetting is authorized; absent explicit instructions, the default is all rows with valid paired values.
  2. Read the script: locate the data load and every operation between load and the statistic call (dropna, astype, boolean masks, head, merges, groupby, reading only some rows/sheet). Check whether non-numeric values would be coerced or would silently drop rows, and whether NaN handling is pairwise-complete across exactly the columns involved.
  3. Check whether the script prints the row count actually used (and ideally min/max/dtype of each column) alongside the statistic, and whether the statistic is corroborated by a second method (e.g., a different library or a manual formula).
  4. Check the final answer: each reported number must be rounded exactly as specified (correct decimal places, including trailing zeros) and must match the script's printed value.
Discriminator
A real violation is when the effective N is never printed/justified, or rows are dropped/coerced by side effect of the pipeline rather than by an explicit, task-justified rule — so an off-by-a-few-rows sample cannot be ruled out. It is not a violation if the script explicitly reports total vs. used rows, the dropping rule is the natural pairwise-complete one (or one the task demands), and the printed value is formatted to the requested precision.
Consequence
The reported coefficient differs from the reference in the last required decimal (e.g., 0.53 vs 0.54) and/or the precision-formatted field is malformed, so exact-match checks on the numeric fields fail even though the qualitative conclusion is right.
id 099a0e9d1286 · mined from infiagent-dabench dabench-300@s6
raw text (what the judge reads)
### Statistic computed on an unverified row subset (silent row loss / coercion) with no N-and-precision sanity check
- **Applies when**: `task` -- the task asks for a single numeric statistic (correlation, mean, test statistic, p-value) over two or more columns of a table, to be reported at a fixed rounding precision.
- **Pattern**: The attempt loads the data and calls a one-line statistic function without ever establishing which rows actually entered the computation — no report of total rows vs. rows used after NaN/blank/non-numeric coercion, no check for stray filtering, deduplication, or a partial/alternate file/sheet being read. Small differences in the effective sample silently shift the statistic in the second decimal, and the reported value is not cross-checked by any independent path. The precision constraint is also treated loosely (e.g., printing a raw float rather than the requested fixed number of decimals).
- **Detection procedure**:
  1. Read the task and note the required statistic, the required rounding, and whether any subsetting is authorized; absent explicit instructions, the default is all rows with valid paired values.
  2. Read the script: locate the data load and every operation between load and the statistic call (`dropna`, `astype`, boolean masks, `head`, merges, groupby, reading only some rows/sheet). Check whether non-numeric values would be coerced or would silently drop rows, and whether NaN handling is pairwise-complete across exactly the columns involved.
  3. Check whether the script prints the row count actually used (and ideally min/max/dtype of each column) alongside the statistic, and whether the statistic is corroborated by a second method (e.g., a different library or a manual formula).
  4. Check the final answer: each reported number must be rounded exactly as specified (correct decimal places, including trailing zeros) and must match the script's printed value.
- **Discriminator**: A real violation is when the effective N is never printed/justified, or rows are dropped/coerced by side effect of the pipeline rather than by an explicit, task-justified rule — so an off-by-a-few-rows sample cannot be ruled out. It is *not* a violation if the script explicitly reports total vs. used rows, the dropping rule is the natural pairwise-complete one (or one the task demands), and the printed value is formatted to the requested precision.
- **Consequence**: The reported coefficient differs from the reference in the last required decimal (e.g., 0.53 vs 0.54) and/or the precision-formatted field is malformed, so exact-match checks on the numeric fields fail even though the qualitative conclusion is right.
347Prediction distribution never sanity-checked against the training targettaskda-code
Applies when
task -- The task asks for a file of model predictions on a held-out set, and the scripts fit a model and write predictions without any held-out evaluation or comparison of predicted vs. observed target statistics.
Pattern
The agent reports only summary numbers of its own output (range, mean, median) and treats "file written with right column name" as success, even though those summaries are implausible relative to the training target (e.g. a floor at exactly 0, an extreme maximum, or a mean/median far from the training distribution) — a symptom of broken encoding, mis-joined features, wrong target scale/transform, or rows written in the wrong order.
Detection procedure
  1. From the task/README, note the expected scale, sign, and shape of the target, and the required output shape (one row per test row, in test order, with the named column).
  2. In the scripts, look for (a) a train/validation split with a reported error metric, and (b) an explicit comparison of predicted-vs-training target summary statistics or a check on row count/order; note whether categorical encodings and missing values are fit on train and applied consistently to test.
  3. In the answer, compare the reported prediction statistics with the training target statistics; flag degenerate boundaries (many exact 0s/negatives), maxima far outside the training range, or means/medians shifted by a large factor.
  4. Flag the attempt if no held-out metric is reported and no distribution/shape sanity check is done, or if the reported prediction summary is clearly inconsistent with the target.
Discriminator
A real violation shows implausible or unverified prediction statistics with no validation score and no distribution check; a look-alike that is fine reports a held-out error metric plus prediction summaries that match the training target's scale and support (skew and a few high-value outliers alone are not a violation if the training target is similarly skewed).
Consequence
The saved prediction file scores far worse than the grader's error/correlation threshold (or fails row-alignment), so the file is marked WRONG despite having the correct name and column.
id dafb26014913 · mined from da-code dacode-ml-regression-014@s6
raw text (what the judge reads)
### Prediction distribution never sanity-checked against the training target
- **Applies when**: `task` -- The task asks for a file of model predictions on a held-out set, and the scripts fit a model and write predictions without any held-out evaluation or comparison of predicted vs. observed target statistics.
- **Pattern**: The agent reports only summary numbers of its own output (range, mean, median) and treats "file written with right column name" as success, even though those summaries are implausible relative to the training target (e.g. a floor at exactly 0, an extreme maximum, or a mean/median far from the training distribution) — a symptom of broken encoding, mis-joined features, wrong target scale/transform, or rows written in the wrong order.
- **Detection procedure**:
  1. From the task/README, note the expected scale, sign, and shape of the target, and the required output shape (one row per test row, in test order, with the named column).
  2. In the scripts, look for (a) a train/validation split with a reported error metric, and (b) an explicit comparison of predicted-vs-training target summary statistics or a check on row count/order; note whether categorical encodings and missing values are fit on train and applied consistently to test.
  3. In the answer, compare the reported prediction statistics with the training target statistics; flag degenerate boundaries (many exact 0s/negatives), maxima far outside the training range, or means/medians shifted by a large factor.
  4. Flag the attempt if no held-out metric is reported *and* no distribution/shape sanity check is done, or if the reported prediction summary is clearly inconsistent with the target.
- **Discriminator**: A real violation shows implausible or unverified prediction statistics with no validation score and no distribution check; a look-alike that is fine reports a held-out error metric plus prediction summaries that match the training target's scale and support (skew and a few high-value outliers alone are not a violation if the training target is similarly skewed).
- **Consequence**: The saved prediction file scores far worse than the grader's error/correlation threshold (or fails row-alignment), so the file is marked WRONG despite having the correct name and column.
348Incomplete deliverables: only the visible artifact is produced, and reported numbers aren't self-consistenttaskda-code
Applies when
task -- the task points to an external spec/config file (e.g. a YAML/JSON of plotting or output guidelines) and/or implies saved artifacts beyond the one file explicitly named in the prompt.
Pattern
The agent reads only the parts of the spec it finds convenient (colors, size, title), writes the single named image/output file, and skips the companion machine-readable artifacts (serialized plot parameters/data, numeric arrays of the computed values) that the spec or workspace convention requires; the summary text also quotes percentages/counts that do not reconcile with each other, showing no post-hoc check.
Detection procedure
  1. Open the referenced spec/config and enumerate every key it defines (figure params, labels, ordering, and any output/serialization directives); list every artifact the task or spec implies must exist on disk.
  2. Read the scripts (or note their absence — unsaved scripts are itself a red flag) and check that each enumerated artifact is written with the exact required filename/format, not just the one file named in the prompt.
  3. Recompute the reported derived quantities from the reported raw counts (e.g. share = count / total) and compare with the stated values and with the category ordering/labels required by the spec.
  4. Flag if any required artifact is unaccounted for, or if step 3 disagrees with the answer text.
Discriminator
A real violation is a missing/renamed/mis-formatted required artifact or arithmetic that does not reconcile; it is not a violation if the extra artifacts are genuinely absent from the spec and every reported number reproduces from the stated counts (harmless extra files or rounding at the last digit are fine).
Consequence
The grader checks each expected output file independently and marks the missing/derived-inconsistent ones WRONG/MISSING, so the run scores 0 even when the headline category identified is plausible.
id a26cefb4b386 · mined from da-code dacode-plot-pie-008@s6
raw text (what the judge reads)
### Incomplete deliverables: only the visible artifact is produced, and reported numbers aren't self-consistent
- **Applies when**: `task` -- the task points to an external spec/config file (e.g. a YAML/JSON of plotting or output guidelines) and/or implies saved artifacts beyond the one file explicitly named in the prompt.
- **Pattern**: The agent reads only the parts of the spec it finds convenient (colors, size, title), writes the single named image/output file, and skips the companion machine-readable artifacts (serialized plot parameters/data, numeric arrays of the computed values) that the spec or workspace convention requires; the summary text also quotes percentages/counts that do not reconcile with each other, showing no post-hoc check.
- **Detection procedure**:
  1. Open the referenced spec/config and enumerate every key it defines (figure params, labels, ordering, *and* any output/serialization directives); list every artifact the task or spec implies must exist on disk.
  2. Read the scripts (or note their absence — unsaved scripts are itself a red flag) and check that each enumerated artifact is written with the exact required filename/format, not just the one file named in the prompt.
  3. Recompute the reported derived quantities from the reported raw counts (e.g. share = count / total) and compare with the stated values and with the category ordering/labels required by the spec.
  4. Flag if any required artifact is unaccounted for, or if step 3 disagrees with the answer text.
- **Discriminator**: A real violation is a missing/renamed/mis-formatted required artifact or arithmetic that does not reconcile; it is *not* a violation if the extra artifacts are genuinely absent from the spec and every reported number reproduces from the stated counts (harmless extra files or rounding at the last digit are fine).
- **Consequence**: The grader checks each expected output file independently and marks the missing/derived-inconsistent ones WRONG/MISSING, so the run scores 0 even when the headline category identified is plausible.
349Undocumented row-dropping / ad-hoc preprocessing choices that change the evaluated sampletaskinfiagent-dabench
Applies when
task -- a task specifies an exact modeling recipe (fixed features, fixed split seed, fixed metric) on a dataset that has missing values or a target/prediction type mismatch, so preprocessing choices silently change which rows are scored and how predictions are thresholded.
Pattern
the attempt makes a convenient, unstated preprocessing decision — e.g. dropping all rows with any missing value instead of imputing (or vice versa), dropping/keeping extra columns, encoding differently than instructed, or converting continuous model output to class labels with an arbitrary or missing rounding/threshold rule — so the row count, feature matrix, and split contents differ from the canonical recipe and the metric lands a few points off.
Detection procedure
  1. From the task, list every fixed element of the recipe (feature set, encoding type, split fraction and seed, model class, metric, rounding) and note that any change in row count or label-derivation rule shifts the metric.
  2. In the scripts, find each preprocessing step and record: rows before vs after (shape prints), how missing values are handled, whether the split is applied to the full processed table, and exactly how continuous predictions are turned into 0/1 labels.
  3. Flag if rows are silently removed, if no missing-value strategy is stated/justified, or if the label-derivation threshold/rounding is arbitrary and untested against the alternative (e.g. no comparison of >=0.5 vs round, or drop-rows vs impute).
  4. Check the final number is the metric on the held-out split of the intended full sample and that a sanity check (row counts, split sizes, accuracy plausibly above the majority-class baseline) was printed.
Discriminator
a real violation is a preprocessing/label-derivation choice that changes the evaluated rows or labels without being required by the task and without any check of the alternative; it is fine if the task explicitly prescribes the handling, or if the script reports the row counts and shows the metric is stable across the reasonable alternatives.
Consequence
the reported accuracy is close to but not equal to the reference value (e.g. off by 0.02 after rounding), so an exact-match grader marks the answer wrong even though the pipeline "looks" correct.
id 45d61098a955 · mined from infiagent-dabench dabench-7@s6
raw text (what the judge reads)
### Undocumented row-dropping / ad-hoc preprocessing choices that change the evaluated sample
- **Applies when**: `task` -- a task specifies an exact modeling recipe (fixed features, fixed split seed, fixed metric) on a dataset that has missing values or a target/prediction type mismatch, so preprocessing choices silently change which rows are scored and how predictions are thresholded.
- **Pattern**: the attempt makes a convenient, unstated preprocessing decision — e.g. dropping all rows with any missing value instead of imputing (or vice versa), dropping/keeping extra columns, encoding differently than instructed, or converting continuous model output to class labels with an arbitrary or missing rounding/threshold rule — so the row count, feature matrix, and split contents differ from the canonical recipe and the metric lands a few points off.
- **Detection procedure**:
  1. From the task, list every fixed element of the recipe (feature set, encoding type, split fraction and seed, model class, metric, rounding) and note that any change in row count or label-derivation rule shifts the metric.
  2. In the scripts, find each preprocessing step and record: rows before vs after (`shape` prints), how missing values are handled, whether the split is applied to the full processed table, and exactly how continuous predictions are turned into 0/1 labels.
  3. Flag if rows are silently removed, if no missing-value strategy is stated/justified, or if the label-derivation threshold/rounding is arbitrary and untested against the alternative (e.g. no comparison of `>=0.5` vs `round`, or drop-rows vs impute).
  4. Check the final number is the metric on the held-out split of the intended full sample and that a sanity check (row counts, split sizes, accuracy plausibly above the majority-class baseline) was printed.
- **Discriminator**: a real violation is a preprocessing/label-derivation choice that changes the evaluated rows or labels without being required by the task and without any check of the alternative; it is fine if the task explicitly prescribes the handling, or if the script reports the row counts and shows the metric is stable across the reasonable alternatives.
- **Consequence**: the reported accuracy is close to but not equal to the reference value (e.g. off by 0.02 after rounding), so an exact-match grader marks the answer wrong even though the pipeline "looks" correct.
350Inventing the output schema/categories instead of deriving them from the provided template filetaskda-code
Applies when
task -- The task says results must be written into a provided pre-existing output file "adhering strictly to its format", and the scripts define groupings/labels/rows themselves.
Pattern
The agent never reads the supplied template (its header, row labels, ordering, expected number of rows), invents its own category boundaries and names (including catch-all buckets like "Unknown"/"N/A" for missing values), and writes a brand-new file whose rows/labels/columns only coincidentally resemble the required schema.
Detection procedure
  1. In the task, note that an output artifact with a fixed format is provided and identify what it constrains (column names, row keys, order, row count).
  2. Search the scripts for any read of that template file (e.g., loading it to get its header/index values); if absent, or if labels/bins are hard-coded from the agent's own assumptions, flag it.
  3. Compare the answer's written rows against the template's expected keys: check for extra rows (e.g., a bucket for unclassifiable/missing records), missing rows, renamed labels, or a different ordering/sorting than the template.
  4. Check that the guessed grouping rule is justified by the data or task (documented thresholds), not chosen arbitrarily, and that records with missing grouping values are handled as the template implies rather than dumped into a new category.
Discriminator
A real violation is when the required categories/columns are knowable from the provided file (or task text) and the script never consults them, producing extra/renamed/reordered rows; it is fine if the script loads the template, preserves its exact header and key set/order, and only fills in the values.
Consequence
The output file fails exact-match comparison (unexpected or missing rows such as an "unknown" category, mismatched labels, wrong ordering or counts), so the file check is marked WRONG even though the underlying counting logic may be arithmetically self-consistent.
id 2dbe279caf87 · mined from da-code dacode-dm-csv-001@s6
raw text (what the judge reads)
### Inventing the output schema/categories instead of deriving them from the provided template file
- **Applies when**: `task` -- The task says results must be written into a provided pre-existing output file "adhering strictly to its format", and the scripts define groupings/labels/rows themselves.
- **Pattern**: The agent never reads the supplied template (its header, row labels, ordering, expected number of rows), invents its own category boundaries and names (including catch-all buckets like "Unknown"/"N/A" for missing values), and writes a brand-new file whose rows/labels/columns only coincidentally resemble the required schema.
- **Detection procedure**:
  1. In the task, note that an output artifact with a fixed format is provided and identify what it constrains (column names, row keys, order, row count).
  2. Search the scripts for any read of that template file (e.g., loading it to get its header/index values); if absent, or if labels/bins are hard-coded from the agent's own assumptions, flag it.
  3. Compare the answer's written rows against the template's expected keys: check for extra rows (e.g., a bucket for unclassifiable/missing records), missing rows, renamed labels, or a different ordering/sorting than the template.
  4. Check that the guessed grouping rule is justified by the data or task (documented thresholds), not chosen arbitrarily, and that records with missing grouping values are handled as the template implies rather than dumped into a new category.
- **Discriminator**: A real violation is when the required categories/columns are knowable from the provided file (or task text) and the script never consults them, producing extra/renamed/reordered rows; it is fine if the script loads the template, preserves its exact header and key set/order, and only fills in the values.
- **Consequence**: The output file fails exact-match comparison (unexpected or missing rows such as an "unknown" category, mismatched labels, wrong ordering or counts), so the file check is marked WRONG even though the underlying counting logic may be arithmetically self-consistent.
351Unvalidated prediction file (class balance and column schema never checked against the task spec and training base rate)taskda-code
Applies when
task -- the task asks for a prediction file with a named target column, and the scripts train a classifier and write out predicted labels.
Pattern
The agent trains a model, writes the output file, and reports only self-generated summary statistics (model hyperparameters, counts) without (a) comparing the predicted positive rate to the label prevalence observed in the training data, (b) confirming the file's columns/row count/dtypes/label encoding exactly match what was requested, or (c) reporting a held-out validation score that justifies the chosen decision threshold. A default 0.5 threshold on an imbalanced target typically yields far fewer positives than the true rate, and extra or renamed columns silently break the expected format.
Detection procedure
  1. Read the task statement and note the exact required file name, required column name(s), and expected number of rows (from the test input).
  2. Read the scripts: find where the output is written; check whether the written frame contains exactly the requested column(s) with the requested name and label encoding, and whether any explicit assertion on row count/columns exists.
  3. Compute or locate the positive-class prevalence in the training labels and compare it with the predicted positive rate stated in the answer; flag a large discrepancy (e.g., predicted rate less than ~⅔ or more than ~1.5× the training rate) with no threshold tuning or class-weighting justification.
  4. Check whether the answer reports any held-out validation metric at all; an answer consisting only of hyperparameters and predicted counts is unverified.
Discriminator
A real violation is an unjustified mismatch — no assertion on file schema, and/or a predicted positive rate far off the training prevalence with no stated threshold/metric rationale. It is not a violation if the agent deliberately optimizes a threshold for a stated metric (e.g., maximizing F1 on a validation split) and reports that score, or if the schema is explicitly asserted/printed and matches the request even when extra identifier columns are permitted by the task.
Consequence
The saved file either fails a format/column check outright or scores far below the achievable accuracy/F1 because most true positives are predicted negative, so the grader marks the expected output file WRONG while the agent's summary looks internally consistent.
id e5bb910e89ac · mined from da-code dacode-ml-binary-016@s6
raw text (what the judge reads)
### Unvalidated prediction file (class balance and column schema never checked against the task spec and training base rate)
- **Applies when**: `task` -- the task asks for a prediction file with a named target column, and the scripts train a classifier and write out predicted labels.
- **Pattern**: The agent trains a model, writes the output file, and reports only self-generated summary statistics (model hyperparameters, counts) without (a) comparing the predicted positive rate to the label prevalence observed in the training data, (b) confirming the file's columns/row count/dtypes/label encoding exactly match what was requested, or (c) reporting a held-out validation score that justifies the chosen decision threshold. A default 0.5 threshold on an imbalanced target typically yields far fewer positives than the true rate, and extra or renamed columns silently break the expected format.
- **Detection procedure**:
  1. Read the task statement and note the exact required file name, required column name(s), and expected number of rows (from the test input).
  2. Read the scripts: find where the output is written; check whether the written frame contains exactly the requested column(s) with the requested name and label encoding, and whether any explicit assertion on row count/columns exists.
  3. Compute or locate the positive-class prevalence in the training labels and compare it with the predicted positive rate stated in the answer; flag a large discrepancy (e.g., predicted rate less than ~⅔ or more than ~1.5× the training rate) with no threshold tuning or class-weighting justification.
  4. Check whether the answer reports any held-out validation metric at all; an answer consisting only of hyperparameters and predicted counts is unverified.
- **Discriminator**: A real violation is an unjustified mismatch — no assertion on file schema, and/or a predicted positive rate far off the training prevalence with no stated threshold/metric rationale. It is *not* a violation if the agent deliberately optimizes a threshold for a stated metric (e.g., maximizing F1 on a validation split) and reports that score, or if the schema is explicitly asserted/printed and matches the request even when extra identifier columns are permitted by the task.
- **Consequence**: The saved file either fails a format/column check outright or scores far below the achievable accuracy/F1 because most true positives are predicted negative, so the grader marks the expected output file WRONG while the agent's summary looks internally consistent.
352Ignoring a stated randomness/seed constraint (and the provided output template) that implies a resampling-based proceduretaskda-code
Applies when
task -- the prompt fixes a random seed and/or points to a sample output file, and the agent must produce a statistic (p-value, CI, score) written to a specified file format.
Pattern
The agent computes the quantity with a deterministic closed-form/library shortcut (e.g., an analytic parametric test or default estimator) in which the seed plays no role, then "sets the seed" cosmetically and writes a file whose column names/shape are invented rather than copied from the supplied sample; the number reported answers a related but methodologically different question than the one the course/task framing implies.
Detection procedure
  1. Read the task for explicit constraints: seed, sample/template output file, units, rounding, column names. Note that a required seed only matters if the method involves simulation/resampling/randomized splitting.
  2. Inspect the script: does any step actually consume randomness (permutation, bootstrap, shuffling, random init)? If not, the seed constraint has been silently dropped and the intended method was likely replaced.
  3. Check whether the script reads/echoes the provided sample output file to derive column names, ordering, and value formatting, instead of hard-coding a guessed schema.
  4. Compare the reported value's definition to the requested one (which test/statistic, one- vs two-sided, which groups/subset) and confirm the written file contains exactly that value in the template's shape.
Discriminator
A genuine violation is when no randomness is used anywhere despite a mandated seed, or the output schema was never validated against the sample; it is not a violation if the method is legitimately deterministic and the agent verified the sample template and matched it exactly (a seed set defensively is fine when the requested statistic is truly analytic and the template matches).
Consequence
The graded file mismatches the expected value and/or header, so the file-level check fails (0/1) even though the reported number looks plausible and the narrative conclusion sounds reasonable.
id cb21fa4f6963 · mined from da-code dacode-data-sa-039@s6
raw text (what the judge reads)
### Ignoring a stated randomness/seed constraint (and the provided output template) that implies a resampling-based procedure
- **Applies when**: `task` -- the prompt fixes a random seed and/or points to a sample output file, and the agent must produce a statistic (p-value, CI, score) written to a specified file format.
- **Pattern**: The agent computes the quantity with a deterministic closed-form/library shortcut (e.g., an analytic parametric test or default estimator) in which the seed plays no role, then "sets the seed" cosmetically and writes a file whose column names/shape are invented rather than copied from the supplied sample; the number reported answers a related but methodologically different question than the one the course/task framing implies.
- **Detection procedure**:
  1. Read the task for explicit constraints: seed, sample/template output file, units, rounding, column names. Note that a required seed only matters if the method involves simulation/resampling/randomized splitting.
  2. Inspect the script: does any step actually consume randomness (permutation, bootstrap, shuffling, random init)? If not, the seed constraint has been silently dropped and the intended method was likely replaced.
  3. Check whether the script reads/echoes the provided sample output file to derive column names, ordering, and value formatting, instead of hard-coding a guessed schema.
  4. Compare the reported value's definition to the requested one (which test/statistic, one- vs two-sided, which groups/subset) and confirm the written file contains exactly that value in the template's shape.
- **Discriminator**: A genuine violation is when no randomness is used anywhere despite a mandated seed, or the output schema was never validated against the sample; it is *not* a violation if the method is legitimately deterministic **and** the agent verified the sample template and matched it exactly (a seed set defensively is fine when the requested statistic is truly analytic and the template matches).
- **Consequence**: The graded file mismatches the expected value and/or header, so the file-level check fails (0/1) even though the reported number looks plausible and the narrative conclusion sounds reasonable.
353Predictions submitted with no held-out accuracy estimatetaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test file that will be graded against hidden ground truth, and the script trains a model and writes the output in a single pass.
Pattern
The script fits one default-configured model on all labeled rows, immediately predicts on the test file, and reports only descriptive statistics of the predictions (min/max/mean/std) as if they were "model performance metrics" — there is no train/validation split, cross-validation, error metric (RMSE/MAE/R²), or baseline comparison, so nothing establishes that the predictions are accurate enough to pass. It often also silently drops potentially informative columns (e.g., timestamps/IDs) or leaves feature handling unexamined, with no experiment showing this choice helps.
Detection procedure
  1. Read the task to confirm the deliverable is scored on predictive accuracy against unseen labels, not just file existence.
  2. Scan the script for any holdout split, cross-validation, or scoring call on labeled data; check whether any error metric is computed at all.
  3. Check whether columns dropped or transformations applied before fitting are justified by a measured comparison, and whether train and test feature sets are verified to align (same columns, same order, same dtypes).
  4. Read the answer: are the quoted "performance" numbers actually validation errors, or merely summary statistics of the predicted vector?
Discriminator
A real violation is when no labeled-data error estimate exists anywhere, so accuracy is unknown; it is not a violation if the agent reports a validation/CV score (even from a simple model) and the score is reasonable, nor if the task explicitly scores only format/existence.
Consequence
The grader compares predictions to true labels with an error/correlation threshold; an unvalidated, untuned model with mishandled features scores worse than the threshold and the file is marked WRONG, while the agent's reported "metrics" give no warning because they describe only the prediction distribution.
id 2acb570515de · mined from da-code dacode-ml-regression-015@s6
raw text (what the judge reads)
### Predictions submitted with no held-out accuracy estimate
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test file that will be graded against hidden ground truth, and the script trains a model and writes the output in a single pass.
- **Pattern**: The script fits one default-configured model on all labeled rows, immediately predicts on the test file, and reports only descriptive statistics of the predictions (min/max/mean/std) as if they were "model performance metrics" — there is no train/validation split, cross-validation, error metric (RMSE/MAE/R²), or baseline comparison, so nothing establishes that the predictions are accurate enough to pass. It often also silently drops potentially informative columns (e.g., timestamps/IDs) or leaves feature handling unexamined, with no experiment showing this choice helps.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on predictive accuracy against unseen labels, not just file existence.
  2. Scan the script for any holdout split, cross-validation, or scoring call on labeled data; check whether any error metric is computed at all.
  3. Check whether columns dropped or transformations applied before fitting are justified by a measured comparison, and whether train and test feature sets are verified to align (same columns, same order, same dtypes).
  4. Read the answer: are the quoted "performance" numbers actually validation errors, or merely summary statistics of the predicted vector?
- **Discriminator**: A real violation is when *no* labeled-data error estimate exists anywhere, so accuracy is unknown; it is not a violation if the agent reports a validation/CV score (even from a simple model) and the score is reasonable, nor if the task explicitly scores only format/existence.
- **Consequence**: The grader compares predictions to true labels with an error/correlation threshold; an unvalidated, untuned model with mishandled features scores worse than the threshold and the file is marked WRONG, while the agent's reported "metrics" give no warning because they describe only the prediction distribution.
354Unverified row filtering / partial data shrinking the population used for a summary statistictaskinfiagent-dabench
Applies when
task -- The task asks for a single aggregate statistic (mean, rate, count, correlation) over "all observations" of a field, and the script loads data and applies drops/filters (nulls, "outliers", dedup, one file/chunk of many) before aggregating.
Pattern
The attempt silently reduces the population before aggregating — e.g. applies an invented outlier rule (z-score/IQR/percentile clipping) to a categorical or coded integer field, drops rows on unrelated columns' nulls, reads only one shard/sheet/subset of the source, or aggregates after a merge that lost rows — and reports the resulting number without ever comparing it to the unfiltered value or to the source row count.
Detection procedure
  1. From the task, note the exact population ("all observations in column X") and which filters are actually mandated versus merely mentioned in passing.
  2. In the scripts, list every operation between load and aggregation that can change row count (dropna, boolean masks, quantile/std clipping, drop_duplicates, single-file/nrows/head reads, joins, groupby-then-mean) and check whether each is required by the task and defined by an explicit rule.
  3. Check whether the script prints/asserts the row count and value range used for the aggregate against the raw source count, and whether it reports the unfiltered statistic alongside the filtered one for comparison.
  4. Confirm the reported answer corresponds to the full mandated population (and to the requested statistic, in the requested format) — with saved, re-runnable code showing that count.
Discriminator
A real violation is filtering that is undefined, discretionary, or applied to a field where it is semantically meaningless (coded/categorical values have no "outliers"), or a data load that covers only part of the source, with no count check. Not a violation: dropping only true missing values in the target column, or an explicitly specified filter, when the script logs before/after counts and the aggregate is stable/justified.
Consequence
The reported value is biased away from the true population statistic (here a low mean), failing an exact/tolerance check even though the code runs without error.
id d0effff71ff6 · mined from infiagent-dabench dabench-320@s6
raw text (what the judge reads)
### Unverified row filtering / partial data shrinking the population used for a summary statistic
- **Applies when**: `task` -- The task asks for a single aggregate statistic (mean, rate, count, correlation) over "all observations" of a field, and the script loads data and applies drops/filters (nulls, "outliers", dedup, one file/chunk of many) before aggregating.
- **Pattern**: The attempt silently reduces the population before aggregating — e.g. applies an invented outlier rule (z-score/IQR/percentile clipping) to a categorical or coded integer field, drops rows on unrelated columns' nulls, reads only one shard/sheet/subset of the source, or aggregates after a merge that lost rows — and reports the resulting number without ever comparing it to the unfiltered value or to the source row count.
- **Detection procedure**:
  1. From the task, note the exact population ("all observations in column X") and which filters are actually mandated versus merely mentioned in passing.
  2. In the scripts, list every operation between load and aggregation that can change row count (`dropna`, boolean masks, quantile/std clipping, `drop_duplicates`, single-file/`nrows`/`head` reads, joins, groupby-then-mean) and check whether each is required by the task and defined by an explicit rule.
  3. Check whether the script prints/asserts the row count and value range used for the aggregate against the raw source count, and whether it reports the unfiltered statistic alongside the filtered one for comparison.
  4. Confirm the reported answer corresponds to the full mandated population (and to the requested statistic, in the requested format) — with saved, re-runnable code showing that count.
- **Discriminator**: A real violation is filtering that is undefined, discretionary, or applied to a field where it is semantically meaningless (coded/categorical values have no "outliers"), or a data load that covers only part of the source, with no count check. Not a violation: dropping only true missing values in the target column, or an explicitly specified filter, when the script logs before/after counts and the aggregate is stable/justified.
- **Consequence**: The reported value is biased away from the true population statistic (here a low mean), failing an exact/tolerance check even though the code runs without error.
355Ranking extremes on a column that was never verified to be numeric (and never sanity-checked)taskda-code
Applies when
task -- the task asks for the row(s) achieving a maximum/minimum (or any order statistic) of a field that, as delivered in the raw file, may be stored as text (thousand separators, %, currency/unit suffixes, blanks) or that must first be imputed/cleaned.
Pattern
The attempt loads the file with defaults, fills or ranks directly, and calls idxmax/idxmin/sort_values on a column whose dtype is object, so the comparison is lexicographic (or silently skips coerced-NaN rows) — producing an extreme that is not the true numeric extreme, with no printed dtype, value, or plausibility check.
Detection procedure
  1. From the task, note which field the extreme is taken over and any required preprocessing step (imputation, unit conversion, filtering).
  2. In the scripts, check for an explicit cleaning/coercion step (strip separators/symbols, astype(float) or pd.to_numeric(..., errors='coerce')) applied before imputation and ranking, and check that the mean used for filling is computed on the numeric version.
  3. Check that the script prints the winning rows with their values (and the top/bottom few), so the magnitudes can be inspected; confirm the answer is written to the required output file in the exact requested structure (e.g. list-valued fields if the template shows lists).
  4. Compare the reported extreme values/entities against common-sense magnitude expectations for the quantity; a max that is far from the largest plausible value, or an entity that would not plausibly hold the record, signals lexicographic or NaN-driven ranking.
Discriminator
A real violation is when no dtype/coercion evidence exists and the reported values are absent or implausible; it is fine if the column is genuinely numeric on load (or is explicitly coerced) and the script prints the extreme values, even if the resulting entity is surprising but consistent with the printed numbers.
Consequence
The grader compares the named entity/entities in the result file against the true argmax/argmin and marks the answer wrong (or missing/mis-shaped if the JSON structure differs from the template), scoring 0.
id 812d6dd0ec9a · mined from da-code dacode-di-text-001@s6
raw text (what the judge reads)
### Ranking extremes on a column that was never verified to be numeric (and never sanity-checked)
- **Applies when**: `task` -- the task asks for the row(s) achieving a maximum/minimum (or any order statistic) of a field that, as delivered in the raw file, may be stored as text (thousand separators, %, currency/unit suffixes, blanks) or that must first be imputed/cleaned.
- **Pattern**: The attempt loads the file with defaults, fills or ranks directly, and calls `idxmax`/`idxmin`/`sort_values` on a column whose dtype is `object`, so the comparison is lexicographic (or silently skips coerced-NaN rows) — producing an extreme that is not the true numeric extreme, with no printed dtype, value, or plausibility check.
- **Detection procedure**:
  1. From the task, note which field the extreme is taken over and any required preprocessing step (imputation, unit conversion, filtering).
  2. In the scripts, check for an explicit cleaning/coercion step (strip separators/symbols, `astype(float)` or `pd.to_numeric(..., errors='coerce')`) applied *before* imputation and ranking, and check that the mean used for filling is computed on the numeric version.
  3. Check that the script prints the winning rows *with their values* (and the top/bottom few), so the magnitudes can be inspected; confirm the answer is written to the required output file in the exact requested structure (e.g. list-valued fields if the template shows lists).
  4. Compare the reported extreme values/entities against common-sense magnitude expectations for the quantity; a max that is far from the largest plausible value, or an entity that would not plausibly hold the record, signals lexicographic or NaN-driven ranking.
- **Discriminator**: A real violation is when no dtype/coercion evidence exists **and** the reported values are absent or implausible; it is fine if the column is genuinely numeric on load (or is explicitly coerced) and the script prints the extreme values, even if the resulting entity is surprising but consistent with the printed numbers.
- **Consequence**: The grader compares the named entity/entities in the result file against the true argmax/argmin and marks the answer wrong (or missing/mis-shaped if the JSON structure differs from the template), scoring 0.
356Output file omits requested fields / doesn't match the reference schemataskda-code
Applies when
task -- the task asks for a saved results file that must contain several derived quantities (e.g., component scores, a segment label, and a level/class), and a sample/reference output file or column spec is available in the data directory.
Pattern
The script computes all the intermediate quantities in memory but writes only a minimal subset of columns (typically the ID plus the single final label), silently dropping other explicitly requested outputs; the agent then declares a "100% match" against the reference without actually comparing column sets, row counts, or ordering.
Detection procedure
  1. Read the task statement and list every quantity the deliverable file is required to contain (each noun in "save X, Y and Z" is a column).
  2. Find the line in the script that subsets/writes the dataframe and compare the written column list against that required list, and against the header of any provided sample/expected output.
  3. Check that the verification code actually asserts schema equality (same columns, same number of rows, same key set/order) rather than merging on a key and comparing one label column only.
  4. Check that the reported claim ("verified", "matches") is backed by printed evidence covering the whole file, not one column.
Discriminator
A real violation is when a required or reference-present column is computed but excluded from the saved file, or the saved row count/key set differs from the reference. It is not a violation if the task genuinely asks for a single label and the reference file has exactly those columns, or if extra columns are present beyond the required ones and the grader tolerates supersets.
Consequence
The graded file fails schema/content comparison against the expected artifact (missing columns or mismatched shape), so the check scores 0 even though the underlying computation may have been right.
id 7a7acab58c14 · mined from da-code dacode-dm-csv-052@s6
raw text (what the judge reads)
### Output file omits requested fields / doesn't match the reference schema
- **Applies when**: `task` -- the task asks for a saved results file that must contain several derived quantities (e.g., component scores, a segment label, and a level/class), and a sample/reference output file or column spec is available in the data directory.
- **Pattern**: The script computes all the intermediate quantities in memory but writes only a minimal subset of columns (typically the ID plus the single final label), silently dropping other explicitly requested outputs; the agent then declares a "100% match" against the reference without actually comparing column sets, row counts, or ordering.
- **Detection procedure**:
  1. Read the task statement and list every quantity the deliverable file is required to contain (each noun in "save X, Y and Z" is a column).
  2. Find the line in the script that subsets/writes the dataframe and compare the written column list against that required list, and against the header of any provided sample/expected output.
  3. Check that the verification code actually asserts schema equality (same columns, same number of rows, same key set/order) rather than merging on a key and comparing one label column only.
  4. Check that the reported claim ("verified", "matches") is backed by printed evidence covering the whole file, not one column.
- **Discriminator**: A real violation is when a required or reference-present column is computed but excluded from the saved file, or the saved row count/key set differs from the reference. It is *not* a violation if the task genuinely asks for a single label and the reference file has exactly those columns, or if extra columns are present beyond the required ones and the grader tolerates supersets.
- **Consequence**: The graded file fails schema/content comparison against the expected artifact (missing columns or mismatched shape), so the check scores 0 even though the underlying computation may have been right.
357Fabricating or regenerating the input data instead of using the provided datasettaskda-code
Applies when
task -- The task refers to a specific provided dataset (and possibly sample output files) that the scripts must read from disk.
Pattern
The agent cannot find/parse the real input (or doesn't check carefully), so it synthesizes a "realistic" substitute with random values, or writes to the expected data path, and then runs the requested analysis on that fake data, reporting the numbers as if they were real.
Detection procedure
  1. Read the task to identify the expected input file(s)/paths and any reference/sample output format.
  2. Scan every script for data-creation code (np.random.*, hardcoded literals, make_classification, manual DataFrame construction, or any write to the input data directory) rather than a pure read of the provided file.
  3. Check whether the analysis script consumes that generated file; also check whether the sample/reference format file was ever read and matched.
  4. Sanity-check the reported numbers against domain expectations (e.g., near-zero correlations among conceptually related rating columns, or suspiciously round row counts) — a sign of independently generated random data.
Discriminator
Creating synthetic data purely for unit-testing a function, while the final reported result still comes from reading the provided file, is fine; the violation is when the reported/saved output derives from data the agent invented, or when the agent overwrites/substitutes the real input path.
Consequence
The saved output file contains values unrelated to the true data (here, ~0 correlations from independent random draws), so any exact/tolerance comparison against the expected result fails outright.
id c30369b70b81 · mined from da-code dacode-data-sa-026@s6
raw text (what the judge reads)
### Fabricating or regenerating the input data instead of using the provided dataset
- **Applies when**: `task` -- The task refers to a specific provided dataset (and possibly sample output files) that the scripts must read from disk.
- **Pattern**: The agent cannot find/parse the real input (or doesn't check carefully), so it synthesizes a "realistic" substitute with random values, or writes to the expected data path, and then runs the requested analysis on that fake data, reporting the numbers as if they were real.
- **Detection procedure**:
  1. Read the task to identify the expected input file(s)/paths and any reference/sample output format.
  2. Scan every script for data-creation code (`np.random.*`, hardcoded literals, `make_classification`, manual DataFrame construction, or any write to the input data directory) rather than a pure read of the provided file.
  3. Check whether the analysis script consumes that generated file; also check whether the sample/reference format file was ever read and matched.
  4. Sanity-check the reported numbers against domain expectations (e.g., near-zero correlations among conceptually related rating columns, or suspiciously round row counts) — a sign of independently generated random data.
- **Discriminator**: Creating synthetic data purely for unit-testing a function, while the final reported result still comes from reading the provided file, is fine; the violation is when the reported/saved output derives from data the agent invented, or when the agent overwrites/substitutes the real input path.
- **Consequence**: The saved output file contains values unrelated to the true data (here, ~0 correlations from independent random draws), so any exact/tolerance comparison against the expected result fails outright.
358Optimal-hyperparameter chosen at the edge of the search grid (and against known structure)taskda-code
Applies when
task -- the script must choose a model/structure hyperparameter (e.g., number of clusters/components/topics) by scanning a candidate range and scoring each value with an internal criterion.
Pattern
The script picks the argmax/argmin over a fixed, arbitrarily truncated grid, and the winner lands on the last (or first) candidate in that grid, with a weak/flat score curve; the agent accepts it without extending the range, without a stability/elbow cross-check, and without comparing to a group count that the data or documentation strongly implies (e.g., the cardinality of an available label/category column that was dropped before modeling).
Detection procedure
  1. Read the task and data description for any explicit or implicit indication of the natural number of groups (a label column, documented categories, a stated count) and note it.
  2. In the script, find the candidate grid and the selection rule; check whether the selected value is at a grid endpoint and whether any secondary criterion (elbow, stability, gap statistic, agreement with known groups) is used.
  3. In the reported answer, compare the chosen value's score against neighbours: if the score is monotonically increasing toward the endpoint and/or the absolute score is very low (near-degenerate separation), flag it.
  4. Confirm no re-run with a wider grid or justification for the truncation appears anywhere.
Discriminator
A real violation is an endpoint selection with a monotone/flat score curve, no wider-grid check, and a plausible alternative count implied by the data that was never tested against. It is not a violation if the optimum lies strictly inside the grid with a clear peak, or if the endpoint choice is explicitly justified (wider range tested, or the criterion is known to saturate and other evidence supports the value).
Consequence
The saved output has a partition with the wrong number/granularity of clusters, so any grader comparing cluster count or label agreement (ARI/NMI vs. the natural grouping) against the expected file fails, even though the file format and row count look fine.
id 008d79af2659 · mined from da-code dacode-ml-cluster-010@s6
raw text (what the judge reads)
### Optimal-hyperparameter chosen at the edge of the search grid (and against known structure)
- **Applies when**: `task` -- the script must choose a model/structure hyperparameter (e.g., number of clusters/components/topics) by scanning a candidate range and scoring each value with an internal criterion.
- **Pattern**: The script picks the argmax/argmin over a fixed, arbitrarily truncated grid, and the winner lands on the last (or first) candidate in that grid, with a weak/flat score curve; the agent accepts it without extending the range, without a stability/elbow cross-check, and without comparing to a group count that the data or documentation strongly implies (e.g., the cardinality of an available label/category column that was dropped before modeling).
- **Detection procedure**:
  1. Read the task and data description for any explicit or implicit indication of the natural number of groups (a label column, documented categories, a stated count) and note it.
  2. In the script, find the candidate grid and the selection rule; check whether the selected value is at a grid endpoint and whether any secondary criterion (elbow, stability, gap statistic, agreement with known groups) is used.
  3. In the reported answer, compare the chosen value's score against neighbours: if the score is monotonically increasing toward the endpoint and/or the absolute score is very low (near-degenerate separation), flag it.
  4. Confirm no re-run with a wider grid or justification for the truncation appears anywhere.
- **Discriminator**: A real violation is an endpoint selection with a monotone/flat score curve, no wider-grid check, and a plausible alternative count implied by the data that was never tested against. It is *not* a violation if the optimum lies strictly inside the grid with a clear peak, or if the endpoint choice is explicitly justified (wider range tested, or the criterion is known to saturate and other evidence supports the value).
- **Consequence**: The saved output has a partition with the wrong number/granularity of clusters, so any grader comparing cluster count or label agreement (ARI/NMI vs. the natural grouping) against the expected file fails, even though the file format and row count look fine.
359Fabricated/hard-coded input data instead of loading the provided datasettaskda-code
Applies when
task -- The task supplies a data directory/README and asks for a statistic, interval, or model result derived from that data.
Pattern
The script contains literal arrays/constants typed into the source ("from historical records", "based on the README") and never reads any file from the provided data path, so every downstream number — however correctly the bootstrap/metric code is written — describes invented inputs rather than the real ones.
Detection procedure
  1. Read the task/README and list the data files the agent is expected to load.
  2. Grep the scripts for any file I/O (read_csv, open, load, glob of the data dir); if the only inputs are in-line literals, flag immediately.
  3. If literals exist, check whether their shape/counts/values can be traced to the real files (row counts, date ranges, group sizes); unverifiable "approximate historical values" count as fabricated.
  4. Cross-check the reported summary statistics against what the actual file would give (e.g., means, group sizes) before accepting the answer.
Discriminator
A real violation is when the analysis inputs are hard-coded and unverified; it is fine to hard-code constants that are genuinely specified by the task (thresholds, seeds, split dates, column names) while the observations themselves are read from the provided files.
Consequence
The output file parses and looks plausible, but the interval/metric values do not match the ground truth computed from the actual data, so the result-file check fails.
id 71a186d4b6bf · mined from da-code dacode-data-sa-031@s6
raw text (what the judge reads)
### Fabricated/hard-coded input data instead of loading the provided dataset
- **Applies when**: `task` -- The task supplies a data directory/README and asks for a statistic, interval, or model result derived from that data.
- **Pattern**: The script contains literal arrays/constants typed into the source ("from historical records", "based on the README") and never reads any file from the provided data path, so every downstream number — however correctly the bootstrap/metric code is written — describes invented inputs rather than the real ones.
- **Detection procedure**:
  1. Read the task/README and list the data files the agent is expected to load.
  2. Grep the scripts for any file I/O (`read_csv`, `open`, `load`, glob of the data dir); if the only inputs are in-line literals, flag immediately.
  3. If literals exist, check whether their shape/counts/values can be traced to the real files (row counts, date ranges, group sizes); unverifiable "approximate historical values" count as fabricated.
  4. Cross-check the reported summary statistics against what the actual file would give (e.g., means, group sizes) before accepting the answer.
- **Discriminator**: A real violation is when the *analysis inputs* are hard-coded and unverified; it is fine to hard-code constants that are genuinely specified by the task (thresholds, seeds, split dates, column names) while the observations themselves are read from the provided files.
- **Consequence**: The output file parses and looks plausible, but the interval/metric values do not match the ground truth computed from the actual data, so the result-file check fails.
360Unverified, unreproducible prediction artifacttaskda-code
Applies when
task -- the task requires producing an output file of per-row predictions (or derived values) for a given input set, and the agent's deliverable is that file.
Pattern
The attempt reports the filename as the answer without any retained/inspectable script that reads the input set, fits/applies a model, and writes the file — and without any check that the written file has the required column name, one row per input record, matching order, and plausible value ranges.
Detection procedure
  1. Read the task and list the explicit output contract: file name, exact column header(s), number of rows expected (= size of the provided input set), row ordering, and any dtype/rounding/unit constraints.
  2. Look for a saved script that performs the full path input-file → features → model → written output; if no script exists, or it cannot be traced end to end, the attempt is unreproducible and inadequate.
  3. In the script (or answer), locate the explicit verification of the artifact: re-read the written file and assert shape, header spelling/case, absence of an unexpected index column or NaNs, and that predicted values lie in a plausible range relative to training targets.
  4. Confirm the number of predicted rows equals the input row count and that no rows were dropped by preprocessing (e.g., dropna/filtering on the prediction set).
Discriminator
A real violation is the absence of any traceable generation step or any post-write assertion on the artifact's shape/header/values; a look-alike that is fine is a script that generates the file and prints/asserts its shape, columns, and summary statistics, even if the model itself is simple.
Consequence
The grader checking the output file finds it missing, mis-named, mis-headed, or with the wrong row count/degenerate values, so the file-existence/format check fails and the task scores 0.
id b0286a502307 · mined from da-code dacode-ml-regression-015@s3
raw text (what the judge reads)
### Unverified, unreproducible prediction artifact
- **Applies when**: `task` -- the task requires producing an output file of per-row predictions (or derived values) for a given input set, and the agent's deliverable is that file.
- **Pattern**: The attempt reports the filename as the answer without any retained/inspectable script that reads the input set, fits/applies a model, and writes the file — and without any check that the written file has the required column name, one row per input record, matching order, and plausible value ranges.
- **Detection procedure**:
  1. Read the task and list the explicit output contract: file name, exact column header(s), number of rows expected (= size of the provided input set), row ordering, and any dtype/rounding/unit constraints.
  2. Look for a saved script that performs the full path input-file → features → model → written output; if no script exists, or it cannot be traced end to end, the attempt is unreproducible and inadequate.
  3. In the script (or answer), locate the explicit verification of the artifact: re-read the written file and assert shape, header spelling/case, absence of an unexpected index column or NaNs, and that predicted values lie in a plausible range relative to training targets.
  4. Confirm the number of predicted rows equals the input row count and that no rows were dropped by preprocessing (e.g., dropna/filtering on the prediction set).
- **Discriminator**: A real violation is the absence of any traceable generation step or any post-write assertion on the artifact's shape/header/values; a look-alike that is fine is a script that generates the file and prints/asserts its shape, columns, and summary statistics, even if the model itself is simple.
- **Consequence**: The grader checking the output file finds it missing, mis-named, mis-headed, or with the wrong row count/degenerate values, so the file-existence/format check fails and the task scores 0.
361Deliverable file mismatch — final artifact never re-validated at the required pathtaskda-code
Applies when
task -- the task names a specific output file (name/format/row set) that must contain the model's predictions, and the scripts write, copy, or re-read prediction artifacts across several files.
Pattern
The modeling script writes results to one path, but later "verification" scripts (and the reported answer) read or produce a different file, so the officially required file is never confirmed to exist at the expected location with the complete, correct content; the checks pass on a file the grader does not look at, and/or the reported content is a partial/truncated copy.
Detection procedure
  1. From the task statement, record the exact required output filename, directory convention, header, and the expected number of rows (= number of scoring rows in the provided input file).
  2. Scan every script for file writes and reads: list each to_csv/open(...,'w') target and each read_csv source; check that the final validation script reads the same path that the required deliverable must occupy.
  3. Confirm at least one check compares the deliverable's row count and id set directly against the scoring input file (not against a sample file or an intermediate copy), and that the reported answer contains all those rows, not a preview.
  4. Flag if any path in step 2 diverges, if the deliverable path is never re-opened after being written, or if the answer's row count is smaller than the expected count.
Discriminator
A real violation is a genuine divergence in paths/names/row coverage (e.g., validation and the reported answer target a differently named or partially written file, or the answer holds fewer rows than the scoring set). Not a violation if the extra file is a deliberate duplicate written from the same in-memory frame and the required path is independently re-read and row/id-validated; display truncation in logs is fine as long as the on-disk file is verified complete.
Consequence
The grader reports the expected output file as missing or wrong even though the model ran, because the scored path is absent, stale, or contains only a fraction of the required rows — the metric cannot be computed and the attempt scores zero regardless of model quality.
id d03bd7acee59 · mined from da-code dacode-ml-competition-005@s7
raw text (what the judge reads)
### Deliverable file mismatch — final artifact never re-validated at the required path
- **Applies when**: `task` -- the task names a specific output file (name/format/row set) that must contain the model's predictions, and the scripts write, copy, or re-read prediction artifacts across several files.
- **Pattern**: The modeling script writes results to one path, but later "verification" scripts (and the reported answer) read or produce a *different* file, so the officially required file is never confirmed to exist at the expected location with the complete, correct content; the checks pass on a file the grader does not look at, and/or the reported content is a partial/truncated copy.
- **Detection procedure**:
  1. From the task statement, record the exact required output filename, directory convention, header, and the expected number of rows (= number of scoring rows in the provided input file).
  2. Scan every script for file writes and reads: list each `to_csv`/`open(...,'w')` target and each `read_csv` source; check that the final validation script reads *the same path* that the required deliverable must occupy.
  3. Confirm at least one check compares the deliverable's row count and id set directly against the scoring input file (not against a sample file or an intermediate copy), and that the reported answer contains all those rows, not a preview.
  4. Flag if any path in step 2 diverges, if the deliverable path is never re-opened after being written, or if the answer's row count is smaller than the expected count.
- **Discriminator**: A real violation is a genuine divergence in paths/names/row coverage (e.g., validation and the reported answer target a differently named or partially written file, or the answer holds fewer rows than the scoring set). Not a violation if the extra file is a deliberate duplicate written from the same in-memory frame *and* the required path is independently re-read and row/id-validated; display truncation in logs is fine as long as the on-disk file is verified complete.
- **Consequence**: The grader reports the expected output file as missing or wrong even though the model ran, because the scored path is absent, stale, or contains only a fraction of the required rows — the metric cannot be computed and the attempt scores zero regardless of model quality.
362Unvalidated near-perfect validation score with no prediction-distribution sanity checktaskda-code
Applies when
task -- a script trains a model on auxiliary/train files and writes predictions for a held-out file, and the answer cites a validation metric as evidence of quality.
Pattern
The attempt reports an implausibly strong holdout metric (e.g., R² ≈ 0.98, error orders of magnitude below the target's spread) for a noisy real-world count/price target, and accepts it without checking whether a feature encodes the target (leakage, duplicated rows across train/val, target-derived engineered feature) or whether the emitted predictions have a distribution comparable to the observed target; predicted central tendency and spread are never compared against the training target's.
Detection procedure
  1. From the task/README, note the target's nature (heavy-tailed counts, noisy behavioural signal) and form a prior on achievable accuracy from the available features.
  2. In the scripts, list every feature fed to the model and check for any that is computed from the target, aggregated over the full data including validation rows, or duplicated across splits; also confirm the split is done before any fitting/encoding.
  3. Compare the reported validation error to the target's own standard deviation/median in the training data — flag if the error is implausibly small relative to it.
  4. In the answer, compare reported prediction summary stats (min/median/mean/max, share of near-zero values) against the training target's summary stats and against the required output shape/column/location; flag any large mismatch that is asserted rather than verified.
Discriminator
A genuinely high score is fine when the features are legitimately predictive, the split is clean, and the prediction distribution plausibly matches the target's (similar median/quantiles); a violation is a high score paired with either a target-derived/leaked feature or predictions whose central tendency is far below/above the historical target distribution with no diagnostic run.
Consequence
The submitted prediction file scores poorly against the true labels (grader marks the output file WRONG) even though the reported internal metric looked excellent.
id 1d51d5694d4c · mined from da-code dacode-ml-regression-008@s7
raw text (what the judge reads)
### Unvalidated near-perfect validation score with no prediction-distribution sanity check
- **Applies when**: `task` -- a script trains a model on auxiliary/train files and writes predictions for a held-out file, and the answer cites a validation metric as evidence of quality.
- **Pattern**: The attempt reports an implausibly strong holdout metric (e.g., R² ≈ 0.98, error orders of magnitude below the target's spread) for a noisy real-world count/price target, and accepts it without checking whether a feature encodes the target (leakage, duplicated rows across train/val, target-derived engineered feature) or whether the emitted predictions have a distribution comparable to the observed target; predicted central tendency and spread are never compared against the training target's.
- **Detection procedure**:
  1. From the task/README, note the target's nature (heavy-tailed counts, noisy behavioural signal) and form a prior on achievable accuracy from the available features.
  2. In the scripts, list every feature fed to the model and check for any that is computed from the target, aggregated over the full data including validation rows, or duplicated across splits; also confirm the split is done before any fitting/encoding.
  3. Compare the reported validation error to the target's own standard deviation/median in the training data — flag if the error is implausibly small relative to it.
  4. In the answer, compare reported prediction summary stats (min/median/mean/max, share of near-zero values) against the training target's summary stats and against the required output shape/column/location; flag any large mismatch that is asserted rather than verified.
- **Discriminator**: A genuinely high score is fine when the features are legitimately predictive, the split is clean, and the prediction distribution plausibly matches the target's (similar median/quantiles); a violation is a high score paired with either a target-derived/leaked feature or predictions whose central tendency is far below/above the historical target distribution with no diagnostic run.
- **Consequence**: The submitted prediction file scores poorly against the true labels (grader marks the output file WRONG) even though the reported internal metric looked excellent.
363Untargeted population and default test choice for a hypothesis testtaskda-code
Applies when
task -- the task asks for a p-value and a reject/fail-to-reject decision comparing two groups, and the scripts pick both the analysis sample and the test procedure implicitly.
Pattern
The script loads every available row of both sources, computes the aggregate quantity on the entire history, and applies a default two-sided parametric test (e.g., ttest_ind) without (a) restricting to the comparable subpopulation/time window the question is really about, (b) checking whether the alternative should be one-sided, or (c) verifying the distributional assumptions (normality, equal variance, sample-size imbalance) that make the chosen test valid. The resulting p-value is astronomically small, which is accepted at face value.
Detection procedure
  1. Read the task/README for any scoping language (comparable competition tier, era, category, "official"/"same conditions") and for directional wording in the hypothesis ("greater than", "more than"); note whether the scripts encode these as filters or as a one-sided alternative.
  2. Inspect the scripts for any exploration of the outcome distribution (histogram, skew, normality test) and any comparison of group sizes/variances before the test is selected; absence means the test was chosen by default, not by evidence.
  3. Check whether an alternative test (rank-based/nonparametric) or a filtered subset would plausibly change the p-value by orders of magnitude; if the reported p-value is extreme (e.g., <1e-20) with heterogeneous, decades-spanning pooled data, treat it as a red flag rather than strong evidence.
  4. Confirm the reported p-value comes from the test the task's hypothesis implies (same tail, same sample) and not from a broader convenience sample.
Discriminator
A real violation is when the task or data description implies a narrower comparable sample and/or a directional/nonparametric procedure and the script silently pools everything with a default two-sided t-test. It is not a violation if the script explicitly justifies the pooled sample and the test with checks (distribution inspection, sensitivity to subsetting) and the task truly specifies no scope or direction.
Consequence
The p-value differs from the reference by many orders of magnitude (and can flip the reject/fail-to-reject decision), so the numeric column fails the grader's value check even when the file format and the decision string are well-formed.
id fcfdf599832b · mined from da-code dacode-data-sa-001@s7
raw text (what the judge reads)
### Untargeted population and default test choice for a hypothesis test
- **Applies when**: `task` -- the task asks for a p-value and a reject/fail-to-reject decision comparing two groups, and the scripts pick both the analysis sample and the test procedure implicitly.
- **Pattern**: The script loads every available row of both sources, computes the aggregate quantity on the *entire* history, and applies a default two-sided parametric test (e.g., `ttest_ind`) without (a) restricting to the comparable subpopulation/time window the question is really about, (b) checking whether the alternative should be one-sided, or (c) verifying the distributional assumptions (normality, equal variance, sample-size imbalance) that make the chosen test valid. The resulting p-value is astronomically small, which is accepted at face value.
- **Detection procedure**:
  1. Read the task/README for any scoping language (comparable competition tier, era, category, "official"/"same conditions") and for directional wording in the hypothesis ("greater than", "more than"); note whether the scripts encode these as filters or as a one-sided alternative.
  2. Inspect the scripts for any exploration of the outcome distribution (histogram, skew, normality test) and any comparison of group sizes/variances before the test is selected; absence means the test was chosen by default, not by evidence.
  3. Check whether an alternative test (rank-based/nonparametric) or a filtered subset would plausibly change the p-value by orders of magnitude; if the reported p-value is extreme (e.g., <1e-20) with heterogeneous, decades-spanning pooled data, treat it as a red flag rather than strong evidence.
  4. Confirm the reported p-value comes from the test the task's hypothesis implies (same tail, same sample) and not from a broader convenience sample.
- **Discriminator**: A real violation is when the task or data description implies a narrower comparable sample and/or a directional/nonparametric procedure and the script silently pools everything with a default two-sided t-test. It is *not* a violation if the script explicitly justifies the pooled sample and the test with checks (distribution inspection, sensitivity to subsetting) and the task truly specifies no scope or direction.
- **Consequence**: The p-value differs from the reference by many orders of magnitude (and can flip the reject/fail-to-reject decision), so the numeric column fails the grader's value check even when the file format and the decision string are well-formed.
364Output file stores internally transformed values instead of the requested feature vectortaskda-code
Applies when
task -- the task asks for a result file whose columns echo the input feature vector alongside a computed label/prediction, and the script applies scaling/encoding/imputation/dimensionality reduction before modeling.
Pattern
The agent writes the model-internal matrix (standardized, encoded, imputed, or projected values) into the deliverable's feature columns, and/or silently drops or reorders a subset of the original columns, so the saved rows no longer match the source records that a grader would join or compare against; the summary script only re-reads the agent's own file, so the mismatch is never noticed.
Detection procedure
  1. Read the task/README to determine what the feature columns of the deliverable are supposed to contain (raw dataset values, in dataset order, one row per record) and how many features/rows that implies.
  2. In the script, trace which array is passed to the DataFrame that gets written: is it X (the selected original values) or a transformed variant (X_scaled, encoded, PCA output)? Check whether the column subset and their order match the dataset's feature vector, including any columns silently excluded.
  3. Compare row count and column count of the written file against the source table; verify that a spot value in Feature_0 is a plausible raw value (e.g., an age/count/amount) and not a z-score near 0.
  4. Check that any verification step re-loads the source data and cross-checks against the output, rather than only re-printing the output file.
Discriminator
A real violation is when the deliverable's feature values cannot be matched back to the source records (mean≈0/std≈1 columns, missing columns, label-encoded integers where raw categories were expected). It is not a violation if the task explicitly asks for the preprocessed/embedded representation, or if a transformation is applied only inside the model while the raw values are what get written out.
Consequence
The grader's file-level comparison against expected records fails (values/shape don't align), so the submission is scored wrong even if the clustering itself is reasonable.
id e375ac4e3469 · mined from da-code dacode-ml-cluster-014@s7
raw text (what the judge reads)
### Output file stores internally transformed values instead of the requested feature vector
- **Applies when**: `task` -- the task asks for a result file whose columns echo the input feature vector alongside a computed label/prediction, and the script applies scaling/encoding/imputation/dimensionality reduction before modeling.
- **Pattern**: The agent writes the *model-internal* matrix (standardized, encoded, imputed, or projected values) into the deliverable's feature columns, and/or silently drops or reorders a subset of the original columns, so the saved rows no longer match the source records that a grader would join or compare against; the summary script only re-reads the agent's own file, so the mismatch is never noticed.
- **Detection procedure**:
  1. Read the task/README to determine what the feature columns of the deliverable are supposed to contain (raw dataset values, in dataset order, one row per record) and how many features/rows that implies.
  2. In the script, trace which array is passed to the DataFrame that gets written: is it `X` (the selected original values) or a transformed variant (`X_scaled`, encoded, PCA output)? Check whether the column subset and their order match the dataset's feature vector, including any columns silently excluded.
  3. Compare row count and column count of the written file against the source table; verify that a spot value in `Feature_0` is a plausible raw value (e.g., an age/count/amount) and not a z-score near 0.
  4. Check that any verification step re-loads the *source* data and cross-checks against the output, rather than only re-printing the output file.
- **Discriminator**: A real violation is when the deliverable's feature values cannot be matched back to the source records (mean≈0/std≈1 columns, missing columns, label-encoded integers where raw categories were expected). It is *not* a violation if the task explicitly asks for the preprocessed/embedded representation, or if a transformation is applied only inside the model while the raw values are what get written out.
- **Consequence**: The grader's file-level comparison against expected records fails (values/shape don't align), so the submission is scored wrong even if the clustering itself is reasonable.
365Discretizing / clipping continuous regression outputs without a metric-based justification (and fitting the final model only on a split)taskda-code
Applies when
task -- the target is numeric and the submission is scored by a continuous regression metric (or the metric is unstated), yet the script post-processes model outputs with rounding, casting to int, or clipping to observed bounds before writing the submission file.
Pattern
The agent trains regressors on a train/validation split, evaluates them on continuous predictions, then applies np.round(...).astype(int) and np.clip(...) to the test predictions — a transformation never validated against the same metric — and never refits the chosen model/ensemble on the full training data. The submitted values therefore differ systematically from the values whose quality was measured, and their dtype/precision may not match the reference submission file.
Detection procedure
  1. Read the task/README and the sample submission to determine the expected value type and precision; note whether any instruction requires integer or bounded outputs.
  2. In the scripts, locate the line that produces the submitted column and check for rounding, integer casting, clipping, or other post-processing applied after the last evaluation step.
  3. Verify whether the identical post-processing was scored on held-out data (i.e., the reported validation metric was computed on the transformed predictions); also check whether the final model was refit on all labeled rows before predicting.
  4. Inspect the submitted file: if values are all integers while the sample/target distribution or metric implies continuous predictions, flag it.
Discriminator
Rounding is legitimate only when the task/metric explicitly demands discrete labels (e.g., classification, an accuracy/exact-match metric, or a stated formatting rule) or when the agent demonstrated on validation data that the rounded predictions score at least as well as the raw ones. A violation is post-processing introduced purely by assumption ("target looks like integers"), plus evaluation performed only on the un-transformed predictions.
Consequence
The submitted predictions incur extra quantization error (and truncated tails from clipping) relative to what validation reported, and the model is weaker than it could be because 20% of labeled data was discarded; the grader's metric comparison against the reference solution falls outside tolerance and the file is marked wrong.
id 7c7ff6997671 · mined from da-code dacode-ml-competition-009@s7
raw text (what the judge reads)
### Discretizing / clipping continuous regression outputs without a metric-based justification (and fitting the final model only on a split)
- **Applies when**: `task` -- the target is numeric and the submission is scored by a continuous regression metric (or the metric is unstated), yet the script post-processes model outputs with rounding, casting to int, or clipping to observed bounds before writing the submission file.
- **Pattern**: The agent trains regressors on a train/validation split, evaluates them on continuous predictions, then applies `np.round(...).astype(int)` and `np.clip(...)` to the test predictions — a transformation never validated against the same metric — and never refits the chosen model/ensemble on the full training data. The submitted values therefore differ systematically from the values whose quality was measured, and their dtype/precision may not match the reference submission file.
- **Detection procedure**:
  1. Read the task/README and the sample submission to determine the expected value type and precision; note whether any instruction requires integer or bounded outputs.
  2. In the scripts, locate the line that produces the submitted column and check for rounding, integer casting, clipping, or other post-processing applied after the last evaluation step.
  3. Verify whether the identical post-processing was scored on held-out data (i.e., the reported validation metric was computed on the transformed predictions); also check whether the final model was refit on all labeled rows before predicting.
  4. Inspect the submitted file: if values are all integers while the sample/target distribution or metric implies continuous predictions, flag it.
- **Discriminator**: Rounding is legitimate only when the task/metric explicitly demands discrete labels (e.g., classification, an accuracy/exact-match metric, or a stated formatting rule) **or** when the agent demonstrated on validation data that the rounded predictions score at least as well as the raw ones. A violation is post-processing introduced purely by assumption ("target looks like integers"), plus evaluation performed only on the un-transformed predictions.
- **Consequence**: The submitted predictions incur extra quantization error (and truncated tails from clipping) relative to what validation reported, and the model is weaker than it could be because 20% of labeled data was discarded; the grader's metric comparison against the reference solution falls outside tolerance and the file is marked wrong.
366Arbitrary/undocumented feature-subset selection when the task implies the full feature vectortaskda-code
Applies when
task -- the task asks to run an unsupervised/derived-output procedure on "the dataset" and to save per-row feature values plus a result, without explicitly naming which columns to use.
Pattern
The script silently hand-picks a small, ad-hoc subset of columns (often mixing genuine measurements with identifier-like or non-informative fields such as codes, IDs, or coordinates) instead of using all usable columns after a principled, stated cleaning rule; the emitted Feature_i columns therefore correspond to a feature vector the grader cannot reproduce, and no justification or sensitivity check is given.
Detection procedure
  1. Read the task/README and count how many columns are plausibly usable (e.g., all numeric/parsable-numeric fields) and whether the task names a specific subset.
  2. In the script, find the column-selection step and compare its length and membership to that set; note any dropped informative columns and any retained identifier/administrative columns.
  3. Check whether the script states and applies a reproducible rule (e.g., "all columns coercible to numeric after stripping symbols, minus pure identifiers") versus a hardcoded list.
  4. Check the answer/output header: does the number of Feature_i columns match the rule-derived feature count, and are values in the expected scale/units (raw vs. standardized) for the requested output?
Discriminator
A real violation is a hardcoded list that omits many usable columns or includes obvious non-features, with no stated criterion; it is fine if the task itself names the columns, or if the script applies an explicit, documented, reproducible filter (missingness threshold, dtype coercion, dropping IDs) that a reviewer could re-derive from the data.
Consequence
The saved file has the wrong number/content of Feature_i columns and cluster assignments that don't align with the reference partition, so the file-level check fails outright even though the pipeline "ran successfully."
id a65095077b34 · mined from da-code dacode-ml-cluster-009@s7
raw text (what the judge reads)
### Arbitrary/undocumented feature-subset selection when the task implies the full feature vector
- **Applies when**: `task` -- the task asks to run an unsupervised/derived-output procedure on "the dataset" and to save per-row feature values plus a result, without explicitly naming which columns to use.
- **Pattern**: The script silently hand-picks a small, ad-hoc subset of columns (often mixing genuine measurements with identifier-like or non-informative fields such as codes, IDs, or coordinates) instead of using all usable columns after a principled, stated cleaning rule; the emitted `Feature_i` columns therefore correspond to a feature vector the grader cannot reproduce, and no justification or sensitivity check is given.
- **Detection procedure**:
  1. Read the task/README and count how many columns are plausibly usable (e.g., all numeric/parsable-numeric fields) and whether the task names a specific subset.
  2. In the script, find the column-selection step and compare its length and membership to that set; note any dropped informative columns and any retained identifier/administrative columns.
  3. Check whether the script states and applies a reproducible rule (e.g., "all columns coercible to numeric after stripping symbols, minus pure identifiers") versus a hardcoded list.
  4. Check the answer/output header: does the number of `Feature_i` columns match the rule-derived feature count, and are values in the expected scale/units (raw vs. standardized) for the requested output?
- **Discriminator**: A real violation is a hardcoded list that omits many usable columns or includes obvious non-features, with no stated criterion; it is fine if the task itself names the columns, or if the script applies an explicit, documented, reproducible filter (missingness threshold, dtype coercion, dropping IDs) that a reviewer could re-derive from the data.
- **Consequence**: The saved file has the wrong number/content of `Feature_i` columns and cluster assignments that don't align with the reference partition, so the file-level check fails outright even though the pipeline "ran successfully."
367Selecting a clustering solution by metric alone while ignoring degenerate, outlier-driven clusters on untransformed heavy-tailed featurestaskda-code
Applies when
task -- an unsupervised segmentation task where the script builds aggregate numeric features from raw transactional/skewed data, standardizes them, sweeps k, and picks the "best" k by an internal index (silhouette/DB/CH).
Pattern
The attempt only z-scores highly right-skewed features (no log/rank transform, no outlier handling, no removal of invalid/negative/cancelled records beyond a crude filter), so a few extreme points dominate distances; the internal metric is maximized by a solution that isolates a handful of outliers into micro-clusters, and the agent reports it as "optimal" without any sanity check on cluster sizes, entity counts, or feature distributions.
Detection procedure
  1. Read the task/README for records that are semantically invalid (returns, cancellations, non-positive quantities/prices, duplicate or missing keys) and check whether the script removes or at least inspects them before aggregating.
  2. In the script, check whether the aggregated features are skewness-corrected/outlier-robust before scaling; StandardScaler alone on unbounded monetary/count features is a red flag.
  3. In the reported output, inspect the cluster size distribution and the number of rows written: if any cluster holds a negligible fraction (e.g. <1% or a literal handful of members) while one cluster holds the majority, or if the row count was never reconciled against the expected number of entities, flag it.
  4. Check that the reported "optimal k" was justified by anything besides the raw index value (stability, elbow agreement, interpretability of segment sizes).
Discriminator
A genuinely rare-but-real segment is fine if the script explicitly examined and justified it (outliers inspected, transform tried and rejected, sizes reconciled with domain expectation); the violation is when micro-clusters appear as an unexamined by-product of an unsanitized, unstransformed feature space and the metric maximum is accepted mechanically.
Consequence
The written result file has a cluster labeling (and often a row count / feature representation) that does not match the reference partition — cluster balance and agreement scores (ARI/NMI or size checks) fail, so the file is graded WRONG despite the script running without error.
id 08d686652cd3 · mined from da-code dacode-ml-cluster-016@s7
raw text (what the judge reads)
### Selecting a clustering solution by metric alone while ignoring degenerate, outlier-driven clusters on untransformed heavy-tailed features
- **Applies when**: `task` -- an unsupervised segmentation task where the script builds aggregate numeric features from raw transactional/skewed data, standardizes them, sweeps k, and picks the "best" k by an internal index (silhouette/DB/CH).
- **Pattern**: The attempt only z-scores highly right-skewed features (no log/rank transform, no outlier handling, no removal of invalid/negative/cancelled records beyond a crude filter), so a few extreme points dominate distances; the internal metric is maximized by a solution that isolates a handful of outliers into micro-clusters, and the agent reports it as "optimal" without any sanity check on cluster sizes, entity counts, or feature distributions.
- **Detection procedure**:
  1. Read the task/README for records that are semantically invalid (returns, cancellations, non-positive quantities/prices, duplicate or missing keys) and check whether the script removes or at least inspects them before aggregating.
  2. In the script, check whether the aggregated features are skewness-corrected/outlier-robust before scaling; `StandardScaler` alone on unbounded monetary/count features is a red flag.
  3. In the reported output, inspect the cluster size distribution and the number of rows written: if any cluster holds a negligible fraction (e.g. <1% or a literal handful of members) while one cluster holds the majority, or if the row count was never reconciled against the expected number of entities, flag it.
  4. Check that the reported "optimal k" was justified by anything besides the raw index value (stability, elbow agreement, interpretability of segment sizes).
- **Discriminator**: A genuinely rare-but-real segment is fine if the script explicitly examined and justified it (outliers inspected, transform tried and rejected, sizes reconciled with domain expectation); the violation is when micro-clusters appear as an unexamined by-product of an unsanitized, unstransformed feature space and the metric maximum is accepted mechanically.
- **Consequence**: The written result file has a cluster labeling (and often a row count / feature representation) that does not match the reference partition — cluster balance and agreement scores (ARI/NMI or size checks) fail, so the file is graded WRONG despite the script running without error.
368Prediction file row count/order not validated against the test inputtaskda-code
Applies when
task -- the task asks for per-row predictions on a provided test file saved to a named output file with a specified column name.
Pattern
The agent produces an output file whose number of rows (and/or row order) does not correspond one-to-one with the test input rows — e.g. predictions only for a truncated subset, a resampled/aggregated version, or rows sorted/filtered differently — and never asserts the correspondence. Symptoms include a suspiciously small or round number of predictions, an obviously placeholder/constant final value, and no code that reads the test file's length and compares it to the output length.
Detection procedure
  1. From the task/data, determine the expected number of prediction rows (length of the test file) and the required column name/ordering.
  2. Search the scripts for the point where predictions are written: confirm they are generated from a feature matrix built by transforming all test rows in original order, and that an explicit check (e.g. len(pred) == len(test), index alignment, no dropna/groupby/filter applied to test) exists.
  3. Count the rows in the submitted output and compare to the expected count; inspect the first/last values for truncation artifacts, duplicates, or hand-typed placeholders.
  4. Flag if counts differ, ordering cannot be traced back to the test file, or no alignment check is present.
Discriminator
A genuine violation is a mismatch in row count or an untraceable ordering/index; it is not a violation if the count and order match the test file exactly and the values merely look unusual, nor if the task itself explicitly requests aggregated/fewer rows.
Consequence
The grader cannot align predictions with ground truth, so the file is scored as wrong/missing regardless of model quality (shape mismatch → 0 checks passed).
id a0f120834a57 · mined from da-code dacode-ml-regression-002@s7
raw text (what the judge reads)
### Prediction file row count/order not validated against the test input
- **Applies when**: `task` -- the task asks for per-row predictions on a provided test file saved to a named output file with a specified column name.
- **Pattern**: The agent produces an output file whose number of rows (and/or row order) does not correspond one-to-one with the test input rows — e.g. predictions only for a truncated subset, a resampled/aggregated version, or rows sorted/filtered differently — and never asserts the correspondence. Symptoms include a suspiciously small or round number of predictions, an obviously placeholder/constant final value, and no code that reads the test file's length and compares it to the output length.
- **Detection procedure**:
  1. From the task/data, determine the expected number of prediction rows (length of the test file) and the required column name/ordering.
  2. Search the scripts for the point where predictions are written: confirm they are generated from a feature matrix built by transforming *all* test rows in original order, and that an explicit check (e.g. `len(pred) == len(test)`, index alignment, no dropna/groupby/filter applied to test) exists.
  3. Count the rows in the submitted output and compare to the expected count; inspect the first/last values for truncation artifacts, duplicates, or hand-typed placeholders.
  4. Flag if counts differ, ordering cannot be traced back to the test file, or no alignment check is present.
- **Discriminator**: A genuine violation is a mismatch in row count or an untraceable ordering/index; it is *not* a violation if the count and order match the test file exactly and the values merely look unusual, nor if the task itself explicitly requests aggregated/fewer rows.
- **Consequence**: The grader cannot align predictions with ground truth, so the file is scored as wrong/missing regardless of model quality (shape mismatch → 0 checks passed).
369Plot/analysis output asserted to follow a spec without emitting or checking the machine-readable artifacts (and with silently changed aggregation)taskda-code
Applies when
task -- the task points to an external format/config file (e.g., a YAML/JSON spec) and expects saved outputs (image plus serialized plot/data artifacts) derived from a time series or grouped series in the data.
Pattern
The attempt renders a picture, aggregates or re-bins the source records to a coarser/different granularity than the data provides (or relabels axes/titles by guess), then declares "follows the spec: Yes" in prose — without reading the spec field-by-field, without dumping the plotted series to the required serialized artifacts, and without cross-checking the plotted values against the raw records.
Detection procedure
  1. From the task statement, list every required artifact and every spec source (config file, stated title/labels/ordering/units); confirm the scripts actually open and parse the spec file rather than hardcoding assumed values.
  2. In the scripts, check that each required artifact is written (not just the image) and that the array/series written is the exact series plotted.
  3. Compare the granularity, count, and order of the plotted points against the raw data's native records for the entity in question (number of rows/periods after filtering); flag any aggregation, resampling, or truncation of the range that the task did not request.
  4. Check the answer for self-certification language ("follows spec: Yes", "created successfully") that is not backed by a printed comparison of spec fields and data checks (shape, min/max, first/last labels).
Discriminator
A real violation is when the spec was never parsed, a required output file is absent, or the plotted series' length/granularity differs from the source records without the task asking for aggregation. It is fine if the script reads the spec and applies each field, writes all requested artifacts, and any aggregation is explicitly demanded by the task or spec.
Consequence
Automated checks on the expected serialized artifacts fail as WRONG/MISSING (missing file, or array of the wrong shape/values), so the task scores 0 even though the rendered image looks plausible.
id 43455db225d0 · mined from da-code dacode-plot-line-015@s7
raw text (what the judge reads)
### Plot/analysis output asserted to follow a spec without emitting or checking the machine-readable artifacts (and with silently changed aggregation)
- **Applies when**: `task` -- the task points to an external format/config file (e.g., a YAML/JSON spec) and expects saved outputs (image plus serialized plot/data artifacts) derived from a time series or grouped series in the data.
- **Pattern**: The attempt renders a picture, aggregates or re-bins the source records to a coarser/different granularity than the data provides (or relabels axes/titles by guess), then declares "follows the spec: Yes" in prose — without reading the spec field-by-field, without dumping the plotted series to the required serialized artifacts, and without cross-checking the plotted values against the raw records.
- **Detection procedure**:
  1. From the task statement, list every required artifact and every spec source (config file, stated title/labels/ordering/units); confirm the scripts actually open and parse the spec file rather than hardcoding assumed values.
  2. In the scripts, check that each required artifact is written (not just the image) and that the array/series written is the exact series plotted.
  3. Compare the granularity, count, and order of the plotted points against the raw data's native records for the entity in question (number of rows/periods after filtering); flag any aggregation, resampling, or truncation of the range that the task did not request.
  4. Check the answer for self-certification language ("follows spec: Yes", "created successfully") that is not backed by a printed comparison of spec fields and data checks (shape, min/max, first/last labels).
- **Discriminator**: A real violation is when the spec was never parsed, a required output file is absent, or the plotted series' length/granularity differs from the source records without the task asking for aggregation. It is fine if the script reads the spec and applies each field, writes all requested artifacts, and any aggregation is explicitly demanded by the task or spec.
- **Consequence**: Automated checks on the expected serialized artifacts fail as WRONG/MISSING (missing file, or array of the wrong shape/values), so the task scores 0 even though the rendered image looks plausible.
370Output file schema invented instead of copied from the provided templatetaskda-code
Applies when
task -- the task supplies a sample/example result file (or an explicit column/row spec) and asks for the computed value(s) to be written into a result file in that exact format.
Pattern
The attempt computes a statistic and writes a self-designed CSV with extra descriptive columns, renamed/reordered headers, different rounding, or a different number of rows, without ever reading the sample file to confirm the expected schema; the narrative answer is prose-heavy while the deliverable file silently deviates.
Detection procedure
  1. In the task/instructions, locate the referenced template or format spec and note exactly which columns, header names, row count, units and rounding are required.
  2. In the scripts, find where the result file is written and check whether the template was loaded/inspected (or its headers hard-copied) versus a hand-written header list; also check that the value written is the final requested quantity, not an intermediate.
  3. Compare the agent's produced file header/rows field-by-field against the template; flag any added, missing, renamed or reordered fields, or differing precision/row structure.
  4. Cross-check the input coverage that fed the number (record counts per group, filters applied) against the source data size, since a template mismatch often accompanies analyzing only a slice of the data.
Discriminator
A real violation is any structural deviation from the mandated file layout (extra columns, different names/order/row count) or a value derived from an unintended subset; harmless look-alikes are additional prose in the chat/log, or trailing whitespace/quoting differences, while the file's columns, order, and values match the template.
Consequence
The grader reads the expected result file key-by-key, finds the required field missing or the value not matching, and scores the file as WRONG/MISSING even when the underlying method looks reasonable.
id 9c4cabeff736 · mined from da-code dacode-data-sa-028@s7
raw text (what the judge reads)
### Output file schema invented instead of copied from the provided template
- **Applies when**: `task` -- the task supplies a sample/example result file (or an explicit column/row spec) and asks for the computed value(s) to be written into a result file in that exact format.
- **Pattern**: The attempt computes a statistic and writes a self-designed CSV with extra descriptive columns, renamed/reordered headers, different rounding, or a different number of rows, without ever reading the sample file to confirm the expected schema; the narrative answer is prose-heavy while the deliverable file silently deviates.
- **Detection procedure**:
  1. In the task/instructions, locate the referenced template or format spec and note exactly which columns, header names, row count, units and rounding are required.
  2. In the scripts, find where the result file is written and check whether the template was loaded/inspected (or its headers hard-copied) versus a hand-written header list; also check that the value written is the final requested quantity, not an intermediate.
  3. Compare the agent's produced file header/rows field-by-field against the template; flag any added, missing, renamed or reordered fields, or differing precision/row structure.
  4. Cross-check the input coverage that fed the number (record counts per group, filters applied) against the source data size, since a template mismatch often accompanies analyzing only a slice of the data.
- **Discriminator**: A real violation is any structural deviation from the mandated file layout (extra columns, different names/order/row count) or a value derived from an unintended subset; harmless look-alikes are additional prose in the chat/log, or trailing whitespace/quoting differences, while the file's columns, order, and values match the template.
- **Consequence**: The grader reads the expected result file key-by-key, finds the required field missing or the value not matching, and scores the file as WRONG/MISSING even when the underlying method looks reasonable.
371Ignoring a referenced specification file that defines the required categorization/methodtaskda-code
Applies when
task -- the task instructions point to an auxiliary document (README, spec, config, notes file) that defines how to bin, filter, label, or otherwise transform the data before producing the requested output.
Pattern
The agent never opens or parses the referenced document, and instead invents its own grouping/labels (e.g., reusing the raw categories already present in the source column, or a "reasonable" default binning), then produces the artifact with those self-chosen categories and asserts success.
Detection procedure
  1. Read the task and list every external file or rule it says the method must follow, plus every output artifact it implies.
  2. Scan the scripts for any read/open/import of that referenced document, or a comment/derivation showing its rules were applied; check whether category labels/boundaries are hard-coded from the agent's own assumption rather than traceable to the spec.
  3. Check the answer for evidence the spec was consulted (quoted rules, matching bin edges/counts, and the same number of categories) and that all required output files exist, not just the visually obvious one.
  4. Flag if the categories come straight from the raw data's existing levels or an unjustified default, or if the spec was never opened.
Discriminator
A real violation is when the spec is never read and the grouping originates solely from the data's own distinct values or the agent's intuition; it is fine if the script loads/quotes the spec and the resulting bins coincidentally match raw levels, or if it explicitly verifies the spec's rules map onto those levels.
Consequence
Category counts and their labels/ordering differ from the reference, so the saved figure and any accompanying serialized data (array/JSON of the plotted values) mismatch the expected outputs and every file-level check fails.
id 849262513268 · mined from da-code dacode-plot-bar-005@s7
raw text (what the judge reads)
### Ignoring a referenced specification file that defines the required categorization/method
- **Applies when**: `task` -- the task instructions point to an auxiliary document (README, spec, config, notes file) that defines how to bin, filter, label, or otherwise transform the data before producing the requested output.
- **Pattern**: The agent never opens or parses the referenced document, and instead invents its own grouping/labels (e.g., reusing the raw categories already present in the source column, or a "reasonable" default binning), then produces the artifact with those self-chosen categories and asserts success.
- **Detection procedure**:
  1. Read the task and list every external file or rule it says the method must follow, plus every output artifact it implies.
  2. Scan the scripts for any read/open/import of that referenced document, or a comment/derivation showing its rules were applied; check whether category labels/boundaries are hard-coded from the agent's own assumption rather than traceable to the spec.
  3. Check the answer for evidence the spec was consulted (quoted rules, matching bin edges/counts, and the same number of categories) and that all required output files exist, not just the visually obvious one.
  4. Flag if the categories come straight from the raw data's existing levels or an unjustified default, or if the spec was never opened.
- **Discriminator**: A real violation is when the spec is never read and the grouping originates solely from the data's own distinct values or the agent's intuition; it is fine if the script loads/quotes the spec and the resulting bins coincidentally match raw levels, or if it explicitly verifies the spec's rules map onto those levels.
- **Consequence**: Category counts and their labels/ordering differ from the reference, so the saved figure and any accompanying serialized data (array/JSON of the plotted values) mismatch the expected outputs and every file-level check fails.
372Required output artifact and literal value formatting not honoredtaskda-code
Applies when
task -- the prompt specifies a deliverable file and/or an exact answer template (key names, list-valued fields, numeric type, units, rounding), and the agent only prints or pastes an answer in chat.
Pattern
The attempt reports a plausible-looking result inline but never writes the expected result file, and/or reshapes the template — scalars instead of the shown list containers, numbers wrapped as strings with appended unit/percent symbols, extra rounding or renamed keys — so an automated checker cannot parse or match it.
Detection procedure
  1. From the task statement, list every hard output requirement: file name/path to be produced, exact keys, container types shown in the template, numeric type, units, and rounding/precision.
  2. Scan the scripts for a write step (json.dump/to_csv/etc.) to that exact filename, and check the object being serialized is built from the template, not ad-hoc.
  3. Compare the submitted answer element-by-element against the template: same keys, same nesting (list vs scalar), same value type (number vs string), no added symbols/suffixes not present in the source data.
  4. If any requirement is unmet, flag — regardless of whether the underlying computation looks correct.
Discriminator
A real violation is a missing artifact or a structural/type deviation from the shown template (scalar where a list is shown, "82.6%" where the source stores a bare number). Harmless look-alikes: cosmetic whitespace/key ordering, or a value formatted exactly as the source data stores it while still matching the template's container and type.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even when the computed quantity would otherwise have been right, making the substantive analysis unverifiable.
id c9eeb99ff900 · mined from da-code dacode-di-text-002@s7
raw text (what the judge reads)
### Required output artifact and literal value formatting not honored
- **Applies when**: `task` -- the prompt specifies a deliverable file and/or an exact answer template (key names, list-valued fields, numeric type, units, rounding), and the agent only prints or pastes an answer in chat.
- **Pattern**: The attempt reports a plausible-looking result inline but never writes the expected result file, and/or reshapes the template — scalars instead of the shown list containers, numbers wrapped as strings with appended unit/percent symbols, extra rounding or renamed keys — so an automated checker cannot parse or match it.
- **Detection procedure**:
  1. From the task statement, list every hard output requirement: file name/path to be produced, exact keys, container types shown in the template, numeric type, units, and rounding/precision.
  2. Scan the scripts for a write step (`json.dump`/`to_csv`/etc.) to that exact filename, and check the object being serialized is built from the template, not ad-hoc.
  3. Compare the submitted answer element-by-element against the template: same keys, same nesting (list vs scalar), same value type (number vs string), no added symbols/suffixes not present in the source data.
  4. If any requirement is unmet, flag — regardless of whether the underlying computation looks correct.
- **Discriminator**: A real violation is a missing artifact or a structural/type deviation from the shown template (scalar where a list is shown, `"82.6%"` where the source stores a bare number). Harmless look-alikes: cosmetic whitespace/key ordering, or a value formatted exactly as the source data stores it while still matching the template's container and type.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even when the computed quantity would otherwise have been right, making the substantive analysis unverifiable.
373Uses a look-alike formula instead of the specifically named statistictaskinfiagent-dabench
Applies when
task -- the task names a specific, named variant of a statistic/metric (e.g., a particular coefficient, "first"/"second" version, macro vs micro, population vs sample) and the script hard-codes one formula.
Pattern
The agent implements a different but related formula that shares the same family name (or a more familiar textbook version), gets a plausible number of the right sign/order of magnitude, and never checks the implemented expression against the exact named definition. Often a comment in the script asserts the formula equals the requested one without justification.
Detection procedure
  1. Extract from the task the exact name/qualifier of the requested statistic and write down its canonical formula (including which central-tendency terms, denominators, or ddof it uses).
  2. Read the script's arithmetic line and list its components; check term-by-term against the canonical formula (e.g., does it use the mode where the definition requires the mode, the correct multiplier, the correct dispersion measure and its ddof).
  3. Check whether the script computes any alternative/cross-check version and whether the agent reconciled discrepancies rather than just printing them "for reference".
  4. Confirm the reported answer comes from the correctly named variant, not the substitute.
Discriminator
A real violation is a formula whose components differ from the named definition (different central-tendency term, missing/extra factor, wrong ddof), even if the sign matches. Not a violation: an algebraically equivalent rearrangement, or a library call documented to implement exactly the named variant, or a case where the task defines the formula explicitly and the script matches that definition.
Consequence
The numeric field is wrong (here, a value off by a large margin) while the derived categorical/qualitative field may still pass, yielding a partial-credit failure.
id 845a6eb148e7 · mined from infiagent-dabench dabench-359@s7
raw text (what the judge reads)
### Uses a look-alike formula instead of the specifically named statistic
- **Applies when**: `task` -- the task names a specific, named variant of a statistic/metric (e.g., a particular coefficient, "first"/"second" version, macro vs micro, population vs sample) and the script hard-codes one formula.
- **Pattern**: The agent implements a different but related formula that shares the same family name (or a more familiar textbook version), gets a plausible number of the right sign/order of magnitude, and never checks the implemented expression against the exact named definition. Often a comment in the script asserts the formula equals the requested one without justification.
- **Detection procedure**:
  1. Extract from the task the exact name/qualifier of the requested statistic and write down its canonical formula (including which central-tendency terms, denominators, or ddof it uses).
  2. Read the script's arithmetic line and list its components; check term-by-term against the canonical formula (e.g., does it use the mode where the definition requires the mode, the correct multiplier, the correct dispersion measure and its ddof).
  3. Check whether the script computes any alternative/cross-check version and whether the agent reconciled discrepancies rather than just printing them "for reference".
  4. Confirm the reported answer comes from the correctly named variant, not the substitute.
- **Discriminator**: A real violation is a formula whose components differ from the named definition (different central-tendency term, missing/extra factor, wrong ddof), even if the sign matches. Not a violation: an algebraically equivalent rearrangement, or a library call documented to implement exactly the named variant, or a case where the task defines the formula explicitly and the script matches that definition.
- **Consequence**: The numeric field is wrong (here, a value off by a large margin) while the derived categorical/qualitative field may still pass, yielding a partial-credit failure.
374Unverified output schema/convention when the task demands a "required format"taskda-code
Applies when
task -- the task asks to save results to a specific file "in the required format" but the script hard-codes column names, ordering, and the metric's exact definition without deriving them from the input data, README, or any provided template/reference.
Pattern
The agent invents its own output layout (self-chosen column labels, dropped or added identifier columns, own convention for a compound/aggregate quantity such as subtracting a baseline, normalizing, or rescaling) and never cross-checks it against the source file's existing columns/conventions or any example of the expected output; it also does not keep the original per-asset/intermediate columns that the source schema implies should be preserved.
Detection procedure
  1. Read the task/README for any statement about the output file, its columns, or naming conventions; note whether the format is only implicitly defined (e.g., by the input file's own columns or by a package/tutorial convention).
  2. Read the script's write step: list the exact columns written, their names, order, index handling, rounding, and the formula used for the requested quantity (e.g., cumulative product vs cumulative product minus one).
  3. Check whether the script anywhere inspects the input file's full column set / any provided template and aligns names and definition to it, or at least prints a comparison; absence of such a check is the flag.
  4. Read the answer: if it reports headline numbers without stating that the file's schema was validated against the required/reference format, treat the format as unverified.
Discriminator
A real violation is inventing labels/definitions with no evidence tying them to the task's stated or implied format. It is fine if the task explicitly specifies the columns and the script matches them, or if the script programmatically derives names from the input/template and verifies row count, column set, and value convention.
Consequence
The values may be internally consistent yet the file fails an exact-schema/value comparison — the grader marks the expected result file WRONG/MISSING even though the analysis logic looks plausible.
id 7c72c032d6c1 · mined from da-code dacode-dm-csv-050@s7
raw text (what the judge reads)
### Unverified output schema/convention when the task demands a "required format"
- **Applies when**: `task` -- the task asks to save results to a specific file "in the required format" but the script hard-codes column names, ordering, and the metric's exact definition without deriving them from the input data, README, or any provided template/reference.
- **Pattern**: The agent invents its own output layout (self-chosen column labels, dropped or added identifier columns, own convention for a compound/aggregate quantity such as subtracting a baseline, normalizing, or rescaling) and never cross-checks it against the source file's existing columns/conventions or any example of the expected output; it also does not keep the original per-asset/intermediate columns that the source schema implies should be preserved.
- **Detection procedure**:
  1. Read the task/README for any statement about the output file, its columns, or naming conventions; note whether the format is only implicitly defined (e.g., by the input file's own columns or by a package/tutorial convention).
  2. Read the script's write step: list the exact columns written, their names, order, index handling, rounding, and the formula used for the requested quantity (e.g., cumulative product vs cumulative product minus one).
  3. Check whether the script anywhere inspects the input file's full column set / any provided template and aligns names and definition to it, or at least prints a comparison; absence of such a check is the flag.
  4. Read the answer: if it reports headline numbers without stating that the file's schema was validated against the required/reference format, treat the format as unverified.
- **Discriminator**: A real violation is inventing labels/definitions with no evidence tying them to the task's stated or implied format. It is fine if the task explicitly specifies the columns and the script matches them, or if the script programmatically derives names from the input/template and verifies row count, column set, and value convention.
- **Consequence**: The values may be internally consistent yet the file fails an exact-schema/value comparison — the grader marks the expected result file WRONG/MISSING even though the analysis logic looks plausible.
375Answer wrapper/format not reproduced literally as specifiedtaskinfiagent-dabench
Applies when
task -- the prompt dictates an exact answer template (a named variable, delimiters, quoting style, key names, ordering, rounding) and the agent must emit a final answer string.
Pattern
The agent computes correct numbers but re-serializes them in its own style — swapping the required assignment/delimiter syntax, using JSON double quotes instead of Python literals, adding or dropping brackets/braces, renaming or reordering keys, or omitting the variable-name prefix — so an exact/parsed match against the expected answer object fails even though the analysis was right.
Detection procedure
  1. Copy the answer template from the task verbatim and list its structural elements: variable name, assignment or bracket syntax, quote style, key naming pattern, element order, decimal places.
  2. Read the script's final print/emit statement (or the reported answer if no script) and tokenize it against that list, element by element.
  3. Flag any mismatch in wrapper syntax, quoting, key spelling/order, or number formatting, even if all values agree with the computation.
  4. Also check that a script exists and deterministically prints this exact string, rather than the answer being hand-typed.
Discriminator
A real violation is a structural/serialization deviation from the stated template (missing name=, wrong bracket/quote characters, reordered or renamed keys, unrounded floats); a look-alike that is fine is a purely cosmetic difference the task explicitly leaves free (e.g., surrounding whitespace, or trailing zeros when the task only fixes precision).
Consequence
The grader parses or string-matches the answer, finds the expected named result "WRONG/MISSING", and scores 0 despite numerically correct values.
id 3b3337053b9c · mined from infiagent-dabench dabench-450@s7
raw text (what the judge reads)
### Answer wrapper/format not reproduced literally as specified
- **Applies when**: `task` -- the prompt dictates an exact answer template (a named variable, delimiters, quoting style, key names, ordering, rounding) and the agent must emit a final answer string.
- **Pattern**: The agent computes correct numbers but re-serializes them in its own style — swapping the required assignment/delimiter syntax, using JSON double quotes instead of Python literals, adding or dropping brackets/braces, renaming or reordering keys, or omitting the variable-name prefix — so an exact/parsed match against the expected answer object fails even though the analysis was right.
- **Detection procedure**:
  1. Copy the answer template from the task verbatim and list its structural elements: variable name, assignment or bracket syntax, quote style, key naming pattern, element order, decimal places.
  2. Read the script's final print/emit statement (or the reported answer if no script) and tokenize it against that list, element by element.
  3. Flag any mismatch in wrapper syntax, quoting, key spelling/order, or number formatting, even if all values agree with the computation.
  4. Also check that a script exists and deterministically prints this exact string, rather than the answer being hand-typed.
- **Discriminator**: A real violation is a structural/serialization deviation from the stated template (missing `name=`, wrong bracket/quote characters, reordered or renamed keys, unrounded floats); a look-alike that is fine is a purely cosmetic difference the task explicitly leaves free (e.g., surrounding whitespace, or trailing zeros when the task only fixes precision).
- **Consequence**: The grader parses or string-matches the answer, finds the expected named result "WRONG/MISSING", and scores 0 despite numerically correct values.
376Statistic/test computed on an unverified data subset (and required intermediate not reported)taskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test or summary statistic on a specific column/subset, with a stated decision rule and a required reported quantity (e.g. the test p-value).
Pattern
The attempt loads the file, applies the test/statistic to whatever rows survive default loading (all rows, duplicates included, NaNs silently dropped or coerced, wrong sheet/subset/grouping level), never prints the sample size or the intermediate quantity the task explicitly asked for, and reports only the final verdict — so a subset or dtype mistake that flips the verdict is invisible.
Detection procedure
  1. From the task, list (a) the exact population the statistic must describe (which column, which rows, deduplicated or not, missing-value policy) and (b) every quantity that must be reported (p-value, rounded values, format tokens).
  2. In the scripts, locate the line that selects the data passed to the test/statistic; check it prints/records len(x), x.dtype, NaN count, and any filtering or aggregation done beforehand — and check whether that selection matches (a).
  3. Check the answer/log actually contains every quantity from (b); a missing p-value (or missing n) means the verdict cannot be audited and the check fails.
  4. Cross-check internal consistency: does the reported test decision agree with the reported shape statistics and with the sample size (e.g. extreme skew/kurtosis vs. "normal", or a near-zero p-value that is an artifact of testing far more rows than the intended units)? Flag any unexplained mismatch.
Discriminator
A real violation is when the analyzed vector's size/composition is never established or demonstrably differs from the task's stated population, or a mandated reported number is absent. It is not a violation if the script prints n, the NaN/dedup handling, and the required intermediate values, and those match the task specification — even if the final verdict happens to be surprising.
Consequence
The verdict (normal/not, class label, threshold decision) is derived from the wrong rows or an unaudited pipeline and comes out opposite to ground truth; the grader marks the key field WRONG with no way to trace the error back, and format/completeness checks for the required intermediate also fail.
id c10ca94bc4a5 · mined from infiagent-dabench dabench-298@s7
raw text (what the judge reads)
### Statistic/test computed on an unverified data subset (and required intermediate not reported)
- **Applies when**: `task` -- the task asks for a hypothesis test or summary statistic on a specific column/subset, with a stated decision rule and a required reported quantity (e.g. the test p-value).
- **Pattern**: The attempt loads the file, applies the test/statistic to whatever rows survive default loading (all rows, duplicates included, NaNs silently dropped or coerced, wrong sheet/subset/grouping level), never prints the sample size or the intermediate quantity the task explicitly asked for, and reports only the final verdict — so a subset or dtype mistake that flips the verdict is invisible.
- **Detection procedure**:
  1. From the task, list (a) the exact population the statistic must describe (which column, which rows, deduplicated or not, missing-value policy) and (b) every quantity that must be reported (p-value, rounded values, format tokens).
  2. In the scripts, locate the line that selects the data passed to the test/statistic; check it prints/records `len(x)`, `x.dtype`, NaN count, and any filtering or aggregation done beforehand — and check whether that selection matches (a).
  3. Check the answer/log actually contains every quantity from (b); a missing p-value (or missing n) means the verdict cannot be audited and the check fails.
  4. Cross-check internal consistency: does the reported test decision agree with the reported shape statistics and with the sample size (e.g. extreme skew/kurtosis vs. "normal", or a near-zero p-value that is an artifact of testing far more rows than the intended units)? Flag any unexplained mismatch.
- **Discriminator**: A real violation is when the analyzed vector's size/composition is never established or demonstrably differs from the task's stated population, or a mandated reported number is absent. It is *not* a violation if the script prints n, the NaN/dedup handling, and the required intermediate values, and those match the task specification — even if the final verdict happens to be surprising.
- **Consequence**: The verdict (normal/not, class label, threshold decision) is derived from the wrong rows or an unaudited pipeline and comes out opposite to ground truth; the grader marks the key field WRONG with no way to trace the error back, and format/completeness checks for the required intermediate also fail.
377Final model fit on a fragment of the training data (subsample/split confusion) and never validated against the target distributiontaskda-code
Applies when
task -- a script must train a model on a provided training file and emit predictions for a test file in a required submission format.
Pattern
The script draws an arbitrary subsample "for speed", then splits it and calls fit() on only one part (often with mixed-up variable names such as fitting on the object named "val"), so the deployed model sees a small fraction of the labeled rows; weak learners are then blended with hand-picked weights, and no check compares the prediction distribution to the training target distribution or verifies the submission is written to the required filename/location.
Detection procedure
  1. From the task, note the required output file (name, path, columns, row count) and that the full training set is available for fitting.
  2. In the script, trace exactly which array is passed to fit() for the model whose predictions are written out; compute the fraction of the full training rows that array represents, and check whether a final refit on all data occurs.
  3. Check whether the reported validation metric is compared against a trivial baseline (e.g. linear model or mean predictor) and whether predicted min/max/std are compared to the training target's min/max/std.
  4. Check that the output is saved to the exact requested path/filename with the exact expected id count and column names.
Discriminator
A deliberate, documented subsample used only for hyperparameter search followed by a full-data refit (or a sound k-fold CV with metrics near/above a stated baseline) is fine; the violation is when the model producing the submitted predictions was fit on a small slice (or the validation slice) and its predictions' spread/accuracy is never sanity-checked against the target's, or the file is not written where the task asked.
Consequence
Predictions are systematically over-smoothed (much smaller variance and range than the true targets) and the leaderboard/metric score falls below the acceptance threshold — or the expected submission file is judged missing/wrong — so the attempt fails despite "strong" self-reported validation numbers.
id 5fcd3da2e3cc · mined from da-code dacode-ml-competition-008@s7
raw text (what the judge reads)
### Final model fit on a fragment of the training data (subsample/split confusion) and never validated against the target distribution
- **Applies when**: `task` -- a script must train a model on a provided training file and emit predictions for a test file in a required submission format.
- **Pattern**: The script draws an arbitrary subsample "for speed", then splits it and calls `fit()` on only one part (often with mixed-up variable names such as fitting on the object named "val"), so the deployed model sees a small fraction of the labeled rows; weak learners are then blended with hand-picked weights, and no check compares the prediction distribution to the training target distribution or verifies the submission is written to the required filename/location.
- **Detection procedure**:
  1. From the task, note the required output file (name, path, columns, row count) and that the full training set is available for fitting.
  2. In the script, trace exactly which array is passed to `fit()` for the model whose predictions are written out; compute the fraction of the full training rows that array represents, and check whether a final refit on all data occurs.
  3. Check whether the reported validation metric is compared against a trivial baseline (e.g. linear model or mean predictor) and whether predicted min/max/std are compared to the training target's min/max/std.
  4. Check that the output is saved to the exact requested path/filename with the exact expected id count and column names.
- **Discriminator**: A deliberate, documented subsample used only for hyperparameter search followed by a full-data refit (or a sound k-fold CV with metrics near/above a stated baseline) is fine; the violation is when the model producing the submitted predictions was fit on a small slice (or the validation slice) and its predictions' spread/accuracy is never sanity-checked against the target's, or the file is not written where the task asked.
- **Consequence**: Predictions are systematically over-smoothed (much smaller variance and range than the true targets) and the leaderboard/metric score falls below the acceptance threshold — or the expected submission file is judged missing/wrong — so the attempt fails despite "strong" self-reported validation numbers.
378Ambiguous subgroup boundaries used without explicit, verifiable binning ruletaskinfiagent-dabench
Applies when
task -- the task asks for a statistic reported separately for value ranges of a grouping variable whose cut points are described in words (e.g., "below X", "between X and Y", "above Y"), and the scripts must translate that into filters.
Pattern
The attempt picks one boundary convention silently (or leaves no script at all), so borderline values are dropped or lumped into the wrong bin, and group sizes/subsets differ from the intended ones; the reported per-group statistics are then computed on the wrong rows, sometimes with no record of how each group was defined.
Detection procedure
  1. From the task, list the requested groups and note every boundary value and whether the wording leaves inclusivity/exclusivity or gap coverage undetermined (also note any required filter, e.g., only rows where the grouping variable and both correlated variables are non-null).
  2. In the scripts, find the exact filter expressions for each group; check that the union of filters covers all in-scope rows exactly once (no dropped values at cut points, no overlap) and that the same non-null/deduplication filtering is applied identically to each group.
  3. Check that the script prints per-group row counts (and value ranges of the grouping variable) alongside each statistic, so the grouping can be sanity-checked; if no script or no such diagnostics exist, the numbers are unverifiable.
  4. Compare the reported values against these diagnostics: groups with plausible sizes but a statistic near zero/one boundary-sensitive value should prompt re-running with the alternate convention to see whether the answer changes.
Discriminator
A real violation is when boundary handling is unstated/inconsistent, group counts are never checked, or an alternate defensible convention would change the reported figure; it is fine if the script explicitly encodes one convention, covers all rows without gaps/overlap, prints group sizes, and shows the statistic is stable (or the task's wording is genuinely unambiguous).
Consequence
Per-group statistics computed on shifted or truncated subsets — one group may coincidentally match the reference value while the others are off by a large margin, so the grader marks most of the requested values wrong.
id 8c7fbbab1f48 · mined from infiagent-dabench dabench-513@s7
raw text (what the judge reads)
### Ambiguous subgroup boundaries used without explicit, verifiable binning rule
- **Applies when**: `task` -- the task asks for a statistic reported separately for value ranges of a grouping variable whose cut points are described in words (e.g., "below X", "between X and Y", "above Y"), and the scripts must translate that into filters.
- **Pattern**: The attempt picks one boundary convention silently (or leaves no script at all), so borderline values are dropped or lumped into the wrong bin, and group sizes/subsets differ from the intended ones; the reported per-group statistics are then computed on the wrong rows, sometimes with no record of how each group was defined.
- **Detection procedure**:
  1. From the task, list the requested groups and note every boundary value and whether the wording leaves inclusivity/exclusivity or gap coverage undetermined (also note any required filter, e.g., only rows where the grouping variable and both correlated variables are non-null).
  2. In the scripts, find the exact filter expressions for each group; check that the union of filters covers all in-scope rows exactly once (no dropped values at cut points, no overlap) and that the same non-null/deduplication filtering is applied identically to each group.
  3. Check that the script prints per-group row counts (and value ranges of the grouping variable) alongside each statistic, so the grouping can be sanity-checked; if no script or no such diagnostics exist, the numbers are unverifiable.
  4. Compare the reported values against these diagnostics: groups with plausible sizes but a statistic near zero/one boundary-sensitive value should prompt re-running with the alternate convention to see whether the answer changes.
- **Discriminator**: A real violation is when boundary handling is unstated/inconsistent, group counts are never checked, or an alternate defensible convention would change the reported figure; it is fine if the script explicitly encodes one convention, covers all rows without gaps/overlap, prints group sizes, and shows the statistic is stable (or the task's wording is genuinely unambiguous).
- **Consequence**: Per-group statistics computed on shifted or truncated subsets — one group may coincidentally match the reference value while the others are off by a large margin, so the grader marks most of the requested values wrong.
379Invented qualification thresholds / definitions instead of the ones specified in the task materialstaskda-code
Applies when
task -- the task (or an accompanying README/definition section/sample output file) states how a derived quantity must be computed or which records qualify for a ranking, and the script must implement that rule.
Pattern
The script silently substitutes the agent's own interpretation — an arbitrary cutoff value, a different aggregation (e.g. sum where a count/mean is required, or vice-versa), or a filter applied to all sub-results when it was defined for only one — and never re-reads the provided definition text or the sample output file to confirm the rule, column names, ordering, and value formatting.
Detection procedure
  1. Read the task statement and any referenced spec/README/sample files and list every explicitly stated rule: eligibility filters and their numeric thresholds, the aggregation function, rounding/units, ordering, and the exact output schema.
  2. Scan the script for each rule: is the threshold hard-coded from the spec or chosen by the agent ("threshold = X" with no citation)? Is the aggregation the one named in the spec? Is the filter scoped to the sub-results the spec attaches it to?
  3. Check whether the script ever loads/prints the provided sample output to validate column names, row count, id convention and value formatting, rather than assuming them.
  4. Flag if any rule in step 1 has no corresponding, matching implementation in the script, or if a truncated/unread portion of the spec was never inspected.
Discriminator
A real violation is a rule that the task materials actually pin down (a stated minimum count, a named statistic, a sample header) but the code implements differently or invents. It is not a violation if the materials genuinely leave the choice open and the agent documents a reasonable choice, or if the agent's value provably reproduces the specified rule (e.g. sum equals count on a per-record basis here).
Consequence
The ranked sets contain the wrong members/order (and possibly wrong headers or row count), so an exact-match file comparison against the reference output fails on all checks even though the pipeline runs without error.
id 95397242923b · mined from da-code dacode-dm-csv-009@s7
raw text (what the judge reads)
### Invented qualification thresholds / definitions instead of the ones specified in the task materials
- **Applies when**: `task` -- the task (or an accompanying README/definition section/sample output file) states how a derived quantity must be computed or which records qualify for a ranking, and the script must implement that rule.
- **Pattern**: The script silently substitutes the agent's own interpretation — an arbitrary cutoff value, a different aggregation (e.g. sum where a count/mean is required, or vice-versa), or a filter applied to all sub-results when it was defined for only one — and never re-reads the provided definition text or the sample output file to confirm the rule, column names, ordering, and value formatting.
- **Detection procedure**:
  1. Read the task statement and any referenced spec/README/sample files and list every explicitly stated rule: eligibility filters and their numeric thresholds, the aggregation function, rounding/units, ordering, and the exact output schema.
  2. Scan the script for each rule: is the threshold hard-coded from the spec or chosen by the agent ("threshold = X" with no citation)? Is the aggregation the one named in the spec? Is the filter scoped to the sub-results the spec attaches it to?
  3. Check whether the script ever loads/prints the provided sample output to validate column names, row count, id convention and value formatting, rather than assuming them.
  4. Flag if any rule in step 1 has no corresponding, matching implementation in the script, or if a truncated/unread portion of the spec was never inspected.
- **Discriminator**: A real violation is a rule that the task materials actually pin down (a stated minimum count, a named statistic, a sample header) but the code implements differently or invents. It is not a violation if the materials genuinely leave the choice open and the agent documents a reasonable choice, or if the agent's value provably reproduces the specified rule (e.g. sum equals count on a per-record basis here).
- **Consequence**: The ranked sets contain the wrong members/order (and possibly wrong headers or row count), so an exact-match file comparison against the reference output fails on all checks even though the pipeline runs without error.
380Answer asserted from prior knowledge instead of computed from the provided datataskda-code
Applies when
task -- the task asks for specific values/rankings to be derived from a supplied dataset with prescribed preprocessing (e.g., impute missing values, then rank/aggregate), and the answer must be written to a result file.
Pattern
The attempt reports a plausible-looking list of entities/values that reflect general world knowledge or an unstated external source, with no saved script that loads the given file, applies the stated preprocessing, computes the statistic, and serializes the requested output; no intermediate check (row counts, imputed-column dtype, sorted head/tail) is shown.
Detection procedure
  1. From the task, list the mandatory computation steps and output artifacts (input file, preprocessing rule, sort order, JSON keys, result file path).
  2. Inspect the submitted scripts/logs: is there code that reads the provided data, performs the stated imputation, and prints/writes the ranked result? If no script or log output exists, the answer is unverifiable.
  3. Cross-check the reported entities/order against what the data could produce: are the values numeric-cleaned (strings with symbols/commas coerced), is the "lowest" list sorted as instructed, and do the names match the dataset's own naming/spelling conventions?
  4. Flag if any reported item cannot be traced to a printed intermediate from the data, or if the required output file was not produced.
Discriminator
A real violation is an answer with no reproducible data-derived trace (missing script, no printed sorted output, names/format not matching the dataset's conventions). A look-alike that is fine is an attempt whose script does read and process the data and whose printed output matches the reported answer, even if the answer happens to coincide with common knowledge.
Consequence
The graded result file is missing or its contents disagree with the dataset-derived ranking (wrong members, wrong order, or wrong naming), so all checks fail.
id 41d7c054d9a1 · mined from da-code dacode-di-text-003@s7
raw text (what the judge reads)
### Answer asserted from prior knowledge instead of computed from the provided data
- **Applies when**: `task` -- the task asks for specific values/rankings to be derived from a supplied dataset with prescribed preprocessing (e.g., impute missing values, then rank/aggregate), and the answer must be written to a result file.
- **Pattern**: The attempt reports a plausible-looking list of entities/values that reflect general world knowledge or an unstated external source, with no saved script that loads the given file, applies the stated preprocessing, computes the statistic, and serializes the requested output; no intermediate check (row counts, imputed-column dtype, sorted head/tail) is shown.
- **Detection procedure**:
  1. From the task, list the mandatory computation steps and output artifacts (input file, preprocessing rule, sort order, JSON keys, result file path).
  2. Inspect the submitted scripts/logs: is there code that reads the provided data, performs the stated imputation, and prints/writes the ranked result? If no script or log output exists, the answer is unverifiable.
  3. Cross-check the reported entities/order against what the data could produce: are the values numeric-cleaned (strings with symbols/commas coerced), is the "lowest" list sorted as instructed, and do the names match the dataset's own naming/spelling conventions?
  4. Flag if any reported item cannot be traced to a printed intermediate from the data, or if the required output file was not produced.
- **Discriminator**: A real violation is an answer with no reproducible data-derived trace (missing script, no printed sorted output, names/format not matching the dataset's conventions). A look-alike that is fine is an attempt whose script does read and process the data and whose printed output matches the reported answer, even if the answer happens to coincide with common knowledge.
- **Consequence**: The graded result file is missing or its contents disagree with the dataset-derived ranking (wrong members, wrong order, or wrong naming), so all checks fail.
381Over-aggressive row filtering on columns not used by the modeltaskinfiagent-dabench
Applies when
task -- The task names a specific feature column and target column, and the script cleans the table (dropna / to_numeric coercion / de-duplication) before splitting into train and test.
Pattern
The attempt applies dropna() or numeric coercion across the whole dataframe or across a set of columns wider than the ones actually modeled, so rows that are perfectly valid for the stated feature/target are discarded; the surviving row count (and hence the seeded 70/30 split composition) differs from the intended one, shifting the metric.
Detection procedure
  1. From the task, list exactly which columns are needed (feature(s) + target) and any filtering the task explicitly states.
  2. In the script, find every row-removing or coercing operation and note which columns it is evaluated over; check whether it touches columns outside that list, or whether unrelated aggregate/total/subgroup rows are silently kept or dropped.
  3. Compare the row count after cleaning against the raw row count and the count of rows with valid values in only the needed columns — if the script never prints/compares these, the filtering is unverified.
  4. Confirm the answer was produced from the split of the minimally-filtered subset; an unexplained shrinkage of the modeling set invalidates the seeded split.
Discriminator
A real violation is filtering keyed on columns (or non-numeric placeholder values) irrelevant to the model, or filtering the task never requested; it is fine to drop rows that are missing/non-numeric in the feature or target themselves, or to apply a filter the task explicitly states.
Consequence
The train/test partition under the fixed seed contains a different set of rows than the reference, so the reported MSE (or any test-set metric) is close to but not equal to the expected value and is graded wrong.
id 0a13a210cd28 · mined from infiagent-dabench dabench-23@s7
raw text (what the judge reads)
### Over-aggressive row filtering on columns not used by the model
- **Applies when**: `task` -- The task names a specific feature column and target column, and the script cleans the table (dropna / to_numeric coercion / de-duplication) before splitting into train and test.
- **Pattern**: The attempt applies `dropna()` or numeric coercion across the whole dataframe or across a set of columns wider than the ones actually modeled, so rows that are perfectly valid for the stated feature/target are discarded; the surviving row count (and hence the seeded 70/30 split composition) differs from the intended one, shifting the metric.
- **Detection procedure**:
  1. From the task, list exactly which columns are needed (feature(s) + target) and any filtering the task explicitly states.
  2. In the script, find every row-removing or coercing operation and note which columns it is evaluated over; check whether it touches columns outside that list, or whether unrelated aggregate/total/subgroup rows are silently kept or dropped.
  3. Compare the row count after cleaning against the raw row count and the count of rows with valid values in only the needed columns — if the script never prints/compares these, the filtering is unverified.
  4. Confirm the answer was produced from the split of the minimally-filtered subset; an unexplained shrinkage of the modeling set invalidates the seeded split.
- **Discriminator**: A real violation is filtering keyed on columns (or non-numeric placeholder values) irrelevant to the model, or filtering the task never requested; it is fine to drop rows that are missing/non-numeric in the feature or target themselves, or to apply a filter the task explicitly states.
- **Consequence**: The train/test partition under the fixed seed contains a different set of rows than the reference, so the reported MSE (or any test-set metric) is close to but not equal to the expected value and is graded wrong.
382Answer-template fidelity: dropping the required literal quoting/formatting on some fieldstaskinfiagent-dabench
Applies when
task -- the task specifies an exact answer template with tagged fields and gives example literals (e.g. quoted strings, units, or a fixed token form) for each field.
Pattern
The attempt computes the right values but emits them inconsistently with the template — quotes present on one field and omitted on others, or extra/missing delimiters, casing, or unit suffixes — so a string-matching grader fails those fields even though the analysis is correct.
Detection procedure
  1. From the task, write out the answer template verbatim and note, per field, the exact literal form implied (quoted string vs bare number, allowed vocabulary, units, separators).
  2. In the scripts, find the line(s) that build the final answer string and expand it mentally into the literal output.
  3. Compare that literal output field-by-field against the template; flag any field where quoting, spacing, case, or token spelling differs, and especially flag internal inconsistency (some fields quoted, others not).
  4. Confirm the submitted answer text (not just the reasoning) matches; if a field is described as "a string" in the task, require it to appear quoted as in the task's example.
Discriminator
A real violation is a deviation from the explicitly shown literal form (missing quotes on a field the task calls a string, changed token like "NonNormal" vs "Non-Normal"). A look-alike that is fine is harmless whitespace outside the tags, or a field the task shows unquoted being emitted unquoted.
Consequence
The grader reports those specific fields as WRONG/MISSING despite correct underlying values, producing a partial score (e.g. 1/3 checks passed) and an overall incorrect verdict.
id 99f72ba88caf · mined from infiagent-dabench dabench-550@s7
raw text (what the judge reads)
### Answer-template fidelity: dropping the required literal quoting/formatting on some fields
- **Applies when**: `task` -- the task specifies an exact answer template with tagged fields and gives example literals (e.g. quoted strings, units, or a fixed token form) for each field.
- **Pattern**: The attempt computes the right values but emits them inconsistently with the template — quotes present on one field and omitted on others, or extra/missing delimiters, casing, or unit suffixes — so a string-matching grader fails those fields even though the analysis is correct.
- **Detection procedure**:
  1. From the task, write out the answer template verbatim and note, per field, the exact literal form implied (quoted string vs bare number, allowed vocabulary, units, separators).
  2. In the scripts, find the line(s) that build the final answer string and expand it mentally into the literal output.
  3. Compare that literal output field-by-field against the template; flag any field where quoting, spacing, case, or token spelling differs, and especially flag internal inconsistency (some fields quoted, others not).
  4. Confirm the submitted answer text (not just the reasoning) matches; if a field is described as "a string" in the task, require it to appear quoted as in the task's example.
- **Discriminator**: A real violation is a deviation from the explicitly shown literal form (missing quotes on a field the task calls a string, changed token like "NonNormal" vs "Non-Normal"). A look-alike that is fine is harmless whitespace outside the tags, or a field the task shows unquoted being emitted unquoted.
- **Consequence**: The grader reports those specific fields as WRONG/MISSING despite correct underlying values, producing a partial score (e.g. 1/3 checks passed) and an overall incorrect verdict.
383Analysis performed on the wrong entities/quantities than the task specifiestaskda-code
Applies when
task -- the prompt names specific grouping keys, ranking metric, and measured quantity (plus required output files), and the agent must locate them in the provided data.
Pattern
The agent substitutes whatever columns/files it finds convenient, silently redefining the group entity, the ranking criterion, and the plotted measure (e.g. counts of categories instead of average durations per stage, ranked by frequency instead of the stated metric), and reports success without ever reconciling its variables against the task wording. It also produces only some of the requested artifacts.
Detection procedure
1. From the task, list the required grouping entity, the ranking metric for selecting the top-N, the measured statistic per stacked segment, and every required output file. 2. In the scripts/answer, identify which dataset, columns, and aggregation were actually used, and which files are written. 3. Check each item in the list from step 1 has an exact counterpart in step 2 (same entity, same ranking metric, same statistic and units, same file names). 4. Flag if any element is replaced by a proxy, if the ranking metric differs from the stated one, or if any required artifact is missing.
Discriminator
A genuine violation changes the semantics — different entity type, different ranking metric, or a count/sum where an average of a duration was asked, or missing output files. Acceptable look-alikes: the correct quantities computed via differently named columns after documented mapping, or extra outputs beyond those required.
Consequence
All content checks fail — the saved figure and numeric export encode unrelated values, required auxiliary files are absent, and the reported top-N list matches nothing in the expected result.
id 111337f35873 · mined from da-code dacode-plot-scatter-002@s7
raw text (what the judge reads)
### Analysis performed on the wrong entities/quantities than the task specifies
- **Applies when**: `task` -- the prompt names specific grouping keys, ranking metric, and measured quantity (plus required output files), and the agent must locate them in the provided data.
- **Pattern**: The agent substitutes whatever columns/files it finds convenient, silently redefining the group entity, the ranking criterion, and the plotted measure (e.g. counts of categories instead of average durations per stage, ranked by frequency instead of the stated metric), and reports success without ever reconciling its variables against the task wording. It also produces only some of the requested artifacts.
- **Detection procedure**: 1. From the task, list the required grouping entity, the ranking metric for selecting the top-N, the measured statistic per stacked segment, and every required output file. 2. In the scripts/answer, identify which dataset, columns, and aggregation were actually used, and which files are written. 3. Check each item in the list from step 1 has an exact counterpart in step 2 (same entity, same ranking metric, same statistic and units, same file names). 4. Flag if any element is replaced by a proxy, if the ranking metric differs from the stated one, or if any required artifact is missing.
- **Discriminator**: A genuine violation changes the semantics — different entity type, different ranking metric, or a count/sum where an average of a duration was asked, or missing output files. Acceptable look-alikes: the correct quantities computed via differently named columns after documented mapping, or extra outputs beyond those required.
- **Consequence**: All content checks fail — the saved figure and numeric export encode unrelated values, required auxiliary files are absent, and the reported top-N list matches nothing in the expected result.
384Answer string does not reproduce the required output template exactly (and no script exists to generate it)taskinfiagent-dabench
Applies when
task -- the task prescribes a literal answer template (a tag, brackets, key names, separators, ordering, rounding) and the agent must emit a dict/list-shaped result.
Pattern
The agent computes plausible numbers but hand-writes the final answer string instead of having a script emit it verbatim from the template, so it silently deviates from the prescribed syntax (extra/missing brackets or braces, altered spacing/quoting, renamed or reordered keys, values as floats vs ints) — the numbers can be right while the string fails an exact-match grader. The absence of any saved script also means the reviewer cannot re-derive the values.
Detection procedure
  1. Extract the answer template from the task character by character: tag name, opening/closing delimiters, key spelling and order, value types.
  2. Check whether any saved script constructs and prints the final answer string directly from computed values (rather than the agent typing it into prose); if no script/artifact exists, flag immediately as unverifiable.
  3. Diff the submitted answer against the template token-by-token — count brackets/braces, compare key names and order, check value formatting/rounding against stated constraints.
  4. Flag if any character-level deviation from the template exists or if the answer cannot be traced to script output.
Discriminator
A real violation is a deviation in the answer's literal syntax/keys/typing from the stated template, or an answer with no generating code; a look-alike that is fine is an answer whose delimiters, key set, order and value formatting match the template exactly and is printed by a script the reviewer can read.
Consequence
The grader's exact/structured match fails and the item scores 0 even though the underlying computed statistics were correct, with no script available to diagnose or re-submit.
id 1c9239bd1050 · mined from infiagent-dabench dabench-451@s7
raw text (what the judge reads)
### Answer string does not reproduce the required output template exactly (and no script exists to generate it)
- **Applies when**: `task` -- the task prescribes a literal answer template (a tag, brackets, key names, separators, ordering, rounding) and the agent must emit a dict/list-shaped result.
- **Pattern**: The agent computes plausible numbers but hand-writes the final answer string instead of having a script emit it verbatim from the template, so it silently deviates from the prescribed syntax (extra/missing brackets or braces, altered spacing/quoting, renamed or reordered keys, values as floats vs ints) — the numbers can be right while the string fails an exact-match grader. The absence of any saved script also means the reviewer cannot re-derive the values.
- **Detection procedure**:
  1. Extract the answer template from the task character by character: tag name, opening/closing delimiters, key spelling and order, value types.
  2. Check whether any saved script constructs and prints the final answer string directly from computed values (rather than the agent typing it into prose); if no script/artifact exists, flag immediately as unverifiable.
  3. Diff the submitted answer against the template token-by-token — count brackets/braces, compare key names and order, check value formatting/rounding against stated constraints.
  4. Flag if any character-level deviation from the template exists or if the answer cannot be traced to script output.
- **Discriminator**: A real violation is a deviation in the answer's literal syntax/keys/typing from the stated template, or an answer with no generating code; a look-alike that is fine is an answer whose delimiters, key set, order and value formatting match the template exactly and is printed by a script the reviewer can read.
- **Consequence**: The grader's exact/structured match fails and the item scores 0 even though the underlying computed statistics were correct, with no script available to diagnose or re-submit.
385Ignoring a stated per-group split: collapsing a grouped analysis into one aggregate resulttaskda-code
Applies when
task -- the task specifies preprocessing or a statistical test to be performed "for each" level of some grouping variable, and the requested output format is a list/array of values.
Pattern
The attempt applies the filtering step group-wise (or not at all), then pools all groups back together and computes a single statistic, reporting one number and one conclusion instead of one entry per group; equivalently, it treats a plural output container as a one-element list without checking how many entries the grouping implies.
Detection procedure
  1. Read the task and list every grouping/segmentation phrase ("for each X", "by X", "within each X") and count the distinct levels the data would yield for each.
  2. Read the script and check whether the loop/groupby over those levels wraps both the preprocessing and the reported statistic, or only the preprocessing before a concat/pooled call.
  3. Compare the length of each list in the answer against the expected number of levels (and check that conclusions are per-entry, in a stated or natural order such as the data's category order).
  4. Confirm the reported quantity is the requested final statistic per group, not an aggregate summary of them.
Discriminator
A genuine violation shows fewer output entries than the task's grouping implies (e.g., one p-value where the grouping has several levels) with no justification; it is fine if the task's grouping applies only to a preprocessing step and explicitly asks for a single pooled test, or if the grouping variable truly has one level in the filtered data (verified by a printed count).
Consequence
The graded file mismatches the reference on both length and values — the single pooled p-value/conclusion differs from the per-group ones, so all checks fail even if the test itself was coded correctly.
id 38386420bb7b · mined from da-code dacode-data-sa-061@s7
raw text (what the judge reads)
### Ignoring a stated per-group split: collapsing a grouped analysis into one aggregate result
- **Applies when**: `task` -- the task specifies preprocessing or a statistical test to be performed "for each" level of some grouping variable, and the requested output format is a list/array of values.
- **Pattern**: The attempt applies the filtering step group-wise (or not at all), then pools all groups back together and computes a single statistic, reporting one number and one conclusion instead of one entry per group; equivalently, it treats a plural output container as a one-element list without checking how many entries the grouping implies.
- **Detection procedure**:
  1. Read the task and list every grouping/segmentation phrase ("for each X", "by X", "within each X") and count the distinct levels the data would yield for each.
  2. Read the script and check whether the loop/`groupby` over those levels wraps *both* the preprocessing and the reported statistic, or only the preprocessing before a `concat`/pooled call.
  3. Compare the length of each list in the answer against the expected number of levels (and check that conclusions are per-entry, in a stated or natural order such as the data's category order).
  4. Confirm the reported quantity is the requested final statistic per group, not an aggregate summary of them.
- **Discriminator**: A genuine violation shows fewer output entries than the task's grouping implies (e.g., one p-value where the grouping has several levels) with no justification; it is fine if the task's grouping applies only to a preprocessing step and explicitly asks for a single pooled test, or if the grouping variable truly has one level in the filtered data (verified by a printed count).
- **Consequence**: The graded file mismatches the reference on both length and values — the single pooled p-value/conclusion differs from the per-group ones, so all checks fail even if the test itself was coded correctly.
386Unvalidated model: no held-out error estimate and silently discarded predictive columnstaskda-code
Applies when
task -- the task asks for predicted values on a test set to be scored against hidden ground truth, and the script trains a model and writes predictions directly without measuring accuracy.
Pattern
The attempt keeps only a convenient subset of columns (e.g., the numeric/"obvious" ones), drops all categorical, identifier, date/text metadata without justification, never holds out a validation split to compute an error metric, and reports only descriptive statistics of the predictions (mean, range, std) as evidence of success. The narrow, near-constant prediction spread compared with the training target distribution goes unremarked.
Detection procedure
  1. Read the task and the data columns: list which available fields plausibly carry signal about the target.
  2. Read the script's feature-selection step: check whether excluded columns were dropped for a stated reason (leakage, all-null, identifier) or just by dtype convenience; check that categorical/text/date fields were encoded rather than silently ignored.
  3. Search the script/answer for any held-out or cross-validated error metric (RMSE/MAE/R² on data not used for fitting). Absence of any such number is the red flag.
  4. Compare the reported prediction distribution (std, min, max) with the training target's distribution; a much narrower spread with no accompanying validation score indicates an underfit, unverified model.
Discriminator
A real violation is an attempt with no out-of-sample metric at all, or one whose feature exclusion is unexplained and removes fields that obviously relate to the target. It is not a violation if the agent reports a validation score and consciously justifies shrinkage/feature drops (e.g., leakage removal), even if the final score is modest; nor if the target genuinely has low variance.
Consequence
The saved prediction file is graded against ground truth and fails the accuracy threshold — predictions cluster near the mean and the error metric is no better (or worse) than a trivial baseline, while the answer text claims success based on self-referential statistics.
id fa5eb2dd202f · mined from da-code dacode-ml-regression-004@s7
raw text (what the judge reads)
### Unvalidated model: no held-out error estimate and silently discarded predictive columns
- **Applies when**: `task` -- the task asks for predicted values on a test set to be scored against hidden ground truth, and the script trains a model and writes predictions directly without measuring accuracy.
- **Pattern**: The attempt keeps only a convenient subset of columns (e.g., the numeric/"obvious" ones), drops all categorical, identifier, date/text metadata without justification, never holds out a validation split to compute an error metric, and reports only descriptive statistics of the predictions (mean, range, std) as evidence of success. The narrow, near-constant prediction spread compared with the training target distribution goes unremarked.
- **Detection procedure**:
  1. Read the task and the data columns: list which available fields plausibly carry signal about the target.
  2. Read the script's feature-selection step: check whether excluded columns were dropped for a stated reason (leakage, all-null, identifier) or just by dtype convenience; check that categorical/text/date fields were encoded rather than silently ignored.
  3. Search the script/answer for any held-out or cross-validated error metric (RMSE/MAE/R² on data not used for fitting). Absence of any such number is the red flag.
  4. Compare the reported prediction distribution (std, min, max) with the training target's distribution; a much narrower spread with no accompanying validation score indicates an underfit, unverified model.
- **Discriminator**: A real violation is an attempt with **no** out-of-sample metric at all, or one whose feature exclusion is unexplained and removes fields that obviously relate to the target. It is *not* a violation if the agent reports a validation score and consciously justifies shrinkage/feature drops (e.g., leakage removal), even if the final score is modest; nor if the target genuinely has low variance.
- **Consequence**: The saved prediction file is graded against ground truth and fails the accuracy threshold — predictions cluster near the mean and the error metric is no better (or worse) than a trivial baseline, while the answer text claims success based on self-referential statistics.
387Degenerate cluster solution accepted from a single automatic model-selection scoretaskda-code
Applies when
task -- the task asks for an unsupervised grouping with an "appropriate" number of groups and the script picks that number automatically from one internal score (silhouette, inertia elbow, etc.) on raw standardized features.
Pattern
The script scans a range of candidate group counts, takes the arg-max of a single index, and writes the labels out with no check on the resulting partition — so heavy-tailed/outlier-driven features can yield near-singleton groups (clusters of size 1–3) and a low absolute score, while the domain context (a small number of interpretable tiers, e.g. "needs aid / intermediate / developed") implies a small, balanced grouping.
Detection procedure
  1. Read the task for cues on what the grouping is for; note whether an interpretable, small number of tiers is implied and whether any distributional preprocessing (skew/outlier handling, log transform, PCA) is warranted for the feature types described.
  2. In the script, check whether the chosen count is justified by more than one criterion (elbow + silhouette + cluster-size/stability check) and whether any post-hoc validation of the partition exists.
  3. In the printed output/answer, inspect the per-cluster counts and the absolute score: flag if any cluster has ≲2% of the rows, if the score is low (e.g. <0.35) and barely separated from neighbouring k, or if the count exceeds the number of tiers the task narrative implies.
  4. Confirm the agent did not re-examine or defend the solution after seeing such counts.
Discriminator
A real violation is an unvalidated arg-max that produces obviously degenerate or uninterpretable groups; it is fine if the agent inspected cluster sizes/centroids, showed the small cluster is a genuine, stable subpopulation, or corroborated the count with a second criterion and the task narrative.
Consequence
The saved label column encodes the wrong number/assignment of groups, so a comparison against the reference partition (cluster count, size profile, or label agreement such as ARI) fails, marking the output file wrong even though its column names and row count are correct.
id 2d996b7cf5eb · mined from da-code dacode-ml-cluster-013@s7
raw text (what the judge reads)
### Degenerate cluster solution accepted from a single automatic model-selection score
- **Applies when**: `task` -- the task asks for an unsupervised grouping with an "appropriate" number of groups and the script picks that number automatically from one internal score (silhouette, inertia elbow, etc.) on raw standardized features.
- **Pattern**: The script scans a range of candidate group counts, takes the arg-max of a single index, and writes the labels out with no check on the resulting partition — so heavy-tailed/outlier-driven features can yield near-singleton groups (clusters of size 1–3) and a low absolute score, while the domain context (a small number of interpretable tiers, e.g. "needs aid / intermediate / developed") implies a small, balanced grouping.
- **Detection procedure**:
  1. Read the task for cues on what the grouping is *for*; note whether an interpretable, small number of tiers is implied and whether any distributional preprocessing (skew/outlier handling, log transform, PCA) is warranted for the feature types described.
  2. In the script, check whether the chosen count is justified by more than one criterion (elbow + silhouette + cluster-size/stability check) and whether any post-hoc validation of the partition exists.
  3. In the printed output/answer, inspect the per-cluster counts and the absolute score: flag if any cluster has ≲2% of the rows, if the score is low (e.g. <0.35) and barely separated from neighbouring k, or if the count exceeds the number of tiers the task narrative implies.
  4. Confirm the agent did not re-examine or defend the solution after seeing such counts.
- **Discriminator**: A real violation is an unvalidated arg-max that produces obviously degenerate or uninterpretable groups; it is fine if the agent inspected cluster sizes/centroids, showed the small cluster is a genuine, stable subpopulation, or corroborated the count with a second criterion and the task narrative.
- **Consequence**: The saved label column encodes the wrong number/assignment of groups, so a comparison against the reference partition (cluster count, size profile, or label agreement such as ARI) fails, marking the output file wrong even though its column names and row count are correct.
388Per-group distributional statistic computed over the wrong grouping/axis (and not reproducible)taskinfiagent-dabench
Applies when
task -- the task asks for a distribution-shaped statistic (skewness, kurtosis, variance, etc.) computed per entity and then compared across entities, using a specific library/definition.
Pattern
The attempt never pins down what set of rows constitutes each entity's distribution: it filters to the requested slice and then computes the statistic over a collapsed/aggregated series (e.g., one value per entity, or across entities instead of within them), or silently drops/imputes missing values, so the per-entity statistic is meaningless or reflects a different axis than requested. Often no script is saved, so the grouping cannot be inspected.
Detection procedure
  1. Read the task and write down explicitly: the grouping key, the slice/filter, and which column's multiple observations form each distribution; confirm the requested definition/flag of the statistic (e.g., Fisher vs. Pearson, bias correction).
  2. In the scripts, locate the groupby/filter chain and check that each group yields a vector of length > 2 of the target variable after the slice, and that the statistic is applied within groups (axis over rows of the group), not over the vector of group aggregates.
  3. Verify NaN handling and dtype coercion are explicit (dropna vs. nan_policy) and that the library/flags named in the task are actually used, not a hand-rolled or pandas default equivalent.
  4. Check the reported winner against a printed sanity table: number of groups, per-group sample sizes, and the top few statistic values; a single reproducible script must exist that regenerates this table.
Discriminator
A real violation is when the per-group vector length is 1 (or the statistic is computed across groups/columns), when the definition flag differs from the one stated, or when no script/intermediate table exists to verify the grouping. It is not a violation if the script groups correctly, keeps multiple observations per group, uses the specified function/flag, and prints group counts — even if the ranking is close between top candidates.
Consequence
The reported entity is the argmax of a differently-defined quantity, so the single-answer check fails outright (0/1) and the error is invisible without the grouping/count sanity table.
id cf898b9c3126 · mined from infiagent-dabench dabench-252@s7
raw text (what the judge reads)
### Per-group distributional statistic computed over the wrong grouping/axis (and not reproducible)

- **Applies when**: `task` -- the task asks for a distribution-shaped statistic (skewness, kurtosis, variance, etc.) computed *per entity* and then compared across entities, using a specific library/definition.
- **Pattern**: The attempt never pins down what set of rows constitutes each entity's distribution: it filters to the requested slice and then computes the statistic over a collapsed/aggregated series (e.g., one value per entity, or across entities instead of within them), or silently drops/imputes missing values, so the per-entity statistic is meaningless or reflects a different axis than requested. Often no script is saved, so the grouping cannot be inspected.
- **Detection procedure**:
  1. Read the task and write down explicitly: the grouping key, the slice/filter, and which column's *multiple observations* form each distribution; confirm the requested definition/flag of the statistic (e.g., Fisher vs. Pearson, bias correction).
  2. In the scripts, locate the groupby/filter chain and check that each group yields a vector of length > 2 of the target variable after the slice, and that the statistic is applied within groups (axis over rows of the group), not over the vector of group aggregates.
  3. Verify NaN handling and dtype coercion are explicit (dropna vs. nan_policy) and that the library/flags named in the task are actually used, not a hand-rolled or pandas default equivalent.
  4. Check the reported winner against a printed sanity table: number of groups, per-group sample sizes, and the top few statistic values; a single reproducible script must exist that regenerates this table.
- **Discriminator**: A real violation is when the per-group vector length is 1 (or the statistic is computed across groups/columns), when the definition flag differs from the one stated, or when no script/intermediate table exists to verify the grouping. It is *not* a violation if the script groups correctly, keeps multiple observations per group, uses the specified function/flag, and prints group counts — even if the ranking is close between top candidates.
- **Consequence**: The reported entity is the argmax of a differently-defined quantity, so the single-answer check fails outright (0/1) and the error is invisible without the grouping/count sanity table.
389Extra decoration/quoting of list items breaks the required literal answer formattaskinfiagent-dabench
Applies when
task -- The task specifies an exact answer template with bracketed, comma-separated lists (e.g. @keys[a,b] @values[x,y]) and the script builds that string programmatically from DataFrame columns.
Pattern
The script serializes list elements with added characters that the template never showed — quotation marks, brackets, spaces, str() of a Python list, units, or trailing zeros — and/or reorders one list independently of its paired list, so the emitted string is not a literal match for the requested tokens even though the underlying computation is right.
Detection procedure
  1. Copy the answer template from the task statement and note exactly which characters separate and delimit items (no quotes, no spaces, paired lists aligned element-by-element).
  2. Read the string-building code in the script (join, f-strings, print) and mentally render the output for a sample of items, especially string-valued identifiers.
  3. Compare the rendered tokens character-by-character with the template; check that any sorting is applied jointly to paired lists and that rounding/format rules from the task are applied to each value.
  4. Flag if any element carries characters not present in the template, or if the two lists could be ordered inconsistently with each other.
Discriminator
A real violation is a formatting difference in the submitted tokens themselves (quotes, brackets, spaces, wrong pairing). It is not a violation if the script merely prints extra diagnostic output around the final answer line, or if quoting appears only in intermediate debug prints while the final answer string is clean.
Consequence
String/exact-match grading of that field fails even when the numeric computation is correct, giving partial credit (e.g. values matched, identifiers "WRONG/MISSING") and an overall incorrect verdict.
id 01b673371a21 · mined from infiagent-dabench dabench-219@s7
raw text (what the judge reads)
### Extra decoration/quoting of list items breaks the required literal answer format
- **Applies when**: `task` -- The task specifies an exact answer template with bracketed, comma-separated lists (e.g. `@keys[a,b] @values[x,y]`) and the script builds that string programmatically from DataFrame columns.
- **Pattern**: The script serializes list elements with added characters that the template never showed — quotation marks, brackets, spaces, `str()` of a Python list, units, or trailing zeros — and/or reorders one list independently of its paired list, so the emitted string is not a literal match for the requested tokens even though the underlying computation is right.
- **Detection procedure**:
  1. Copy the answer template from the task statement and note exactly which characters separate and delimit items (no quotes, no spaces, paired lists aligned element-by-element).
  2. Read the string-building code in the script (`join`, f-strings, `print`) and mentally render the output for a sample of items, especially string-valued identifiers.
  3. Compare the rendered tokens character-by-character with the template; check that any sorting is applied jointly to paired lists and that rounding/format rules from the task are applied to each value.
  4. Flag if any element carries characters not present in the template, or if the two lists could be ordered inconsistently with each other.
- **Discriminator**: A real violation is a formatting difference in the submitted tokens themselves (quotes, brackets, spaces, wrong pairing). It is *not* a violation if the script merely prints extra diagnostic output around the final answer line, or if quoting appears only in intermediate debug prints while the final answer string is clean.
- **Consequence**: String/exact-match grading of that field fails even when the numeric computation is correct, giving partial credit (e.g. values matched, identifiers "WRONG/MISSING") and an overall incorrect verdict.
390Prediction file shape/coverage not validated against the test settaskda-code
Applies when
task -- The task asks for a per-row prediction (or per-row output) file for every observation in a supplied evaluation/test file, in a format shown by a sample submission.
Pattern
The agent produces an output file whose row count, ordering, header, or value domain doesn't match the test input exactly — e.g. predictions for only a subset of rows (truncated, filtered, deduplicated, or rows dropped by dropna/merge), a re-indexed order, or an extra index column — and never asserts len(predictions) == len(test) before saving.
Detection procedure
  1. From the task/README, note the exact expected number of output rows and the required header/column set from the sample file.
  2. In the scripts, trace the test dataframe from load to prediction: check for any row-dropping operation (NA handling, filtering, sampling, chunked loops, head()), any re-sorting, and whether the write step includes index=False and the required column name; confirm an explicit shape/row-count assertion exists.
  3. Count the rows in the produced answer file (excluding the header) and compare to the expected count; also check values fall in the allowed label set and the positive rate is plausible relative to training prevalence.
  4. If no script is available, treat the file itself as the evidence: row count and header must match the specification exactly.
Discriminator
A genuine violation is a mismatch in row count/order/header/value domain versus the test file specification. A look-alike that is fine is a file with the correct number of rows and header whose predicted class balance merely differs from expectation (that is a model-quality issue, not a format failure), or truncation that occurs only in the displayed excerpt while the saved file is complete.
Consequence
The grader cannot align predictions with ground-truth rows, so the file is scored as wrong/missing and the check fails regardless of model accuracy.
id 687bf20eb762 · mined from da-code dacode-ml-binary-013@s7
raw text (what the judge reads)
### Prediction file shape/coverage not validated against the test set
- **Applies when**: `task` -- The task asks for a per-row prediction (or per-row output) file for every observation in a supplied evaluation/test file, in a format shown by a sample submission.
- **Pattern**: The agent produces an output file whose row count, ordering, header, or value domain doesn't match the test input exactly — e.g. predictions for only a subset of rows (truncated, filtered, deduplicated, or rows dropped by `dropna`/merge), a re-indexed order, or an extra index column — and never asserts `len(predictions) == len(test)` before saving.
- **Detection procedure**:
  1. From the task/README, note the exact expected number of output rows and the required header/column set from the sample file.
  2. In the scripts, trace the test dataframe from load to prediction: check for any row-dropping operation (NA handling, filtering, sampling, chunked loops, `head()`), any re-sorting, and whether the write step includes `index=False` and the required column name; confirm an explicit shape/row-count assertion exists.
  3. Count the rows in the produced answer file (excluding the header) and compare to the expected count; also check values fall in the allowed label set and the positive rate is plausible relative to training prevalence.
  4. If no script is available, treat the file itself as the evidence: row count and header must match the specification exactly.
- **Discriminator**: A genuine violation is a mismatch in row count/order/header/value domain versus the test file specification. A look-alike that is fine is a file with the correct number of rows and header whose predicted class balance merely differs from expectation (that is a model-quality issue, not a format failure), or truncation that occurs only in the displayed excerpt while the saved file is complete.
- **Consequence**: The grader cannot align predictions with ground-truth rows, so the file is scored as wrong/missing and the check fails regardless of model accuracy.
391Scope of the analyzed population not verified against the task's stated universetaskinfiagent-dabench
Applies when
task -- the task asks for a statistic or detection over an entire population ("all countries", "all users", "the whole dataset") and the working directory contains several partitioned/regional/sharded data files or a filterable grouping column.
Pattern
The script loads a single convenient file (or silently keeps one subset/group) and computes the requested statistic on that partial slice, never enumerating the available data sources or checking that the row count matches the full universe implied by the task; the answer is then reported as if it covered everything.
Detection procedure
  1. Read the task and write down the intended population and its expected size/scope (e.g., all entities, not one region or one year-subset).
  2. In the script, find every data-loading call and every filter/subset; check whether all relevant files/partitions are read and concatenated, or whether only one was chosen.
  3. Look for an explicit sanity check in the script output (row/entity count, list of unique groups) that reconciles the loaded data with the expected population size; absence of such a check is a red flag.
  4. Confirm the reported answer (and any quartiles/bounds it depends on) was computed after the full population was assembled, and that the final string matches the requested answer format exactly.
Discriminator
A real violation is loading/keeping a proper subset when the task's universe is broader (or never verifying coverage); it is fine if the single file provably is the full population (documented row/entity count matching the task) or if the task itself restricts the analysis to that subset.
Consequence
Quartiles, thresholds, and hence the set of flagged entities are derived from the wrong distribution, so the reported list is incomplete or spurious and the grader's exact-set comparison fails even when some items happen to overlap.
id c37087e9a1ae · mined from infiagent-dabench dabench-254@s7
raw text (what the judge reads)
### Scope of the analyzed population not verified against the task's stated universe
- **Applies when**: `task` -- the task asks for a statistic or detection over an entire population ("all countries", "all users", "the whole dataset") and the working directory contains several partitioned/regional/sharded data files or a filterable grouping column.
- **Pattern**: The script loads a single convenient file (or silently keeps one subset/group) and computes the requested statistic on that partial slice, never enumerating the available data sources or checking that the row count matches the full universe implied by the task; the answer is then reported as if it covered everything.
- **Detection procedure**:
  1. Read the task and write down the intended population and its expected size/scope (e.g., all entities, not one region or one year-subset).
  2. In the script, find every data-loading call and every filter/subset; check whether all relevant files/partitions are read and concatenated, or whether only one was chosen.
  3. Look for an explicit sanity check in the script output (row/entity count, list of unique groups) that reconciles the loaded data with the expected population size; absence of such a check is a red flag.
  4. Confirm the reported answer (and any quartiles/bounds it depends on) was computed after the full population was assembled, and that the final string matches the requested answer format exactly.
- **Discriminator**: A real violation is loading/keeping a proper subset when the task's universe is broader (or never verifying coverage); it is fine if the single file provably *is* the full population (documented row/entity count matching the task) or if the task itself restricts the analysis to that subset.
- **Consequence**: Quartiles, thresholds, and hence the set of flagged entities are derived from the wrong distribution, so the reported list is incomplete or spurious and the grader's exact-set comparison fails even when some items happen to overlap.
392Unverified prediction file: no held-out score and no check that output values/order match the required formattaskda-code
Applies when
task -- the task asks the agent to train on a labeled split, predict on an unlabeled split, and write predictions to a specified file/column.
Pattern
The attempt fits a model on all training rows and dumps model.predict(test) straight to the output file, without (a) any held-out/CV estimate of accuracy, (b) confirming the written label values are the exact strings/encoding used in the source labels, and (c) confirming row count and row order correspond one-to-one to the test file as read (no shuffling, no dropped rows from NA handling, no index column written). The report cites only class counts and the algorithm name as evidence of success.
Detection procedure
  1. Read the task for the required output artifact: file name, column name, expected number of rows, and the label vocabulary implied by the training labels.
  2. In the scripts, look for a train/validation split (or CV) producing a reported score on data with known labels; if absent, the attempt has no evidence the model or the label mapping is right.
  3. Trace the prediction path: are rows ever dropped/filtered/reindexed after reading the test file, is any label-encoding inverted back to the original strings, and is the file written with index=False and only the requested column?
  4. Check the final answer for concrete verification (validation score, len(result) == len(test), printed head, set of unique output values equal to the training label set) rather than just counts and a model name.
Discriminator
A real violation has zero labeled-data score and no post-write shape/value/order sanity check; an attempt is fine if it reports a validation/CV metric and demonstrates the output file's length, column name, and unique values match the training label vocabulary and test row count — even if the model is simple.
Consequence
The grader compares result.csv to ground truth row-by-row and marks it WRONG/MISSING because labels are mis-encoded (e.g. 0/1 or renamed classes), rows are misaligned/shortened, or an extra index column shifts the schema — while the agent's self-report still claims success.
id ed6cef0da722 · mined from da-code dacode-ml-binary-009@s7
raw text (what the judge reads)
### Unverified prediction file: no held-out score and no check that output values/order match the required format
- **Applies when**: `task` -- the task asks the agent to train on a labeled split, predict on an unlabeled split, and write predictions to a specified file/column.
- **Pattern**: The attempt fits a model on all training rows and dumps `model.predict(test)` straight to the output file, without (a) any held-out/CV estimate of accuracy, (b) confirming the written label values are the exact strings/encoding used in the source labels, and (c) confirming row count and row order correspond one-to-one to the test file as read (no shuffling, no dropped rows from NA handling, no index column written). The report cites only class counts and the algorithm name as evidence of success.
- **Detection procedure**:
  1. Read the task for the required output artifact: file name, column name, expected number of rows, and the label vocabulary implied by the training labels.
  2. In the scripts, look for a train/validation split (or CV) producing a reported score on data with known labels; if absent, the attempt has no evidence the model or the label mapping is right.
  3. Trace the prediction path: are rows ever dropped/filtered/reindexed after reading the test file, is any label-encoding inverted back to the original strings, and is the file written with `index=False` and only the requested column?
  4. Check the final answer for concrete verification (validation score, `len(result) == len(test)`, printed head, set of unique output values equal to the training label set) rather than just counts and a model name.
- **Discriminator**: A real violation has zero labeled-data score *and* no post-write shape/value/order sanity check; an attempt is fine if it reports a validation/CV metric and demonstrates the output file's length, column name, and unique values match the training label vocabulary and test row count — even if the model is simple.
- **Consequence**: The grader compares `result.csv` to ground truth row-by-row and marks it WRONG/MISSING because labels are mis-encoded (e.g. 0/1 or renamed classes), rows are misaligned/shortened, or an extra index column shifts the schema — while the agent's self-report still claims success.
393Fabricated metric definition forced onto a provided plot spec, with no consistency checktaskda-code
Applies when
task -- the task says to visualize/report "performance" (or similar aggregate) using a supplied config/spec file, and the script must decide how to compute the underlying quantity and which records to aggregate over.
Pattern
The attempt invents an ad-hoc formula (e.g., an arbitrary weighted sum of counts) and aggregates over only one side/subset of the relevant entities (e.g., grouping by one role column while ignoring rows where the entity appears in the other role), then plots those numbers against the entity list/order taken from the config without ever checking that the computed values are consistent with the spec (implied ordering, axis range, ticks, chart orientation) or that all requested output artifacts are produced.
Detection procedure
  1. Read the task and the config/spec file: list every constraint it fixes (entity list and order, axis labels/limits, chart type/orientation, figure params) and every output file the task implies.
  2. In the script, locate the definition of the plotted quantity: is it derived from a stated definition in the task/README/spec, or invented? Check whether the aggregation covers all rows in which each entity participates and whether the time/subset filter comes from the task rather than a guess.
  3. Check whether the script validates its numbers against the spec: does the computed ranking/magnitude match the spec's implied order, tick range, or title semantics ("best", "top N")? Does it emit every required artifact (image plus any serialized values/config)?
  4. Read the answer: if it reports numbers whose ordering contradicts the given entity order or whose magnitudes are implausible for the stated concept, and no cross-check was run, flag it.
Discriminator
A real violation is a metric whose definition has no basis in the task, README, or spec and which was never reconciled with the spec's fixed elements (order/limits) or the full set of required outputs. A look-alike that is fine is a metric that is explicitly defined (or unambiguously standard for the domain, e.g., points/wins over all matches of a team) and whose computed values are shown to reproduce the spec's ordering/axis constraints.
Consequence
The plotted values, and any serialized numeric result, differ from the reference; the image and the accompanying value/config files all fail comparison, yielding 0 of the expected checks even though the script "ran successfully".
id 6bb8fb001b53 · mined from da-code dacode-plot-bar-006@s7
raw text (what the judge reads)
### Fabricated metric definition forced onto a provided plot spec, with no consistency check
- **Applies when**: `task` -- the task says to visualize/report "performance" (or similar aggregate) using a supplied config/spec file, and the script must decide how to compute the underlying quantity and which records to aggregate over.
- **Pattern**: The attempt invents an ad-hoc formula (e.g., an arbitrary weighted sum of counts) and aggregates over only one side/subset of the relevant entities (e.g., grouping by one role column while ignoring rows where the entity appears in the other role), then plots those numbers against the entity list/order taken from the config without ever checking that the computed values are consistent with the spec (implied ordering, axis range, ticks, chart orientation) or that all requested output artifacts are produced.
- **Detection procedure**:
  1. Read the task and the config/spec file: list every constraint it fixes (entity list and order, axis labels/limits, chart type/orientation, figure params) and every output file the task implies.
  2. In the script, locate the definition of the plotted quantity: is it derived from a stated definition in the task/README/spec, or invented? Check whether the aggregation covers all rows in which each entity participates and whether the time/subset filter comes from the task rather than a guess.
  3. Check whether the script validates its numbers against the spec: does the computed ranking/magnitude match the spec's implied order, tick range, or title semantics ("best", "top N")? Does it emit every required artifact (image plus any serialized values/config)?
  4. Read the answer: if it reports numbers whose ordering contradicts the given entity order or whose magnitudes are implausible for the stated concept, and no cross-check was run, flag it.
- **Discriminator**: A real violation is a metric whose definition has no basis in the task, README, or spec and which was never reconciled with the spec's fixed elements (order/limits) or the full set of required outputs. A look-alike that is fine is a metric that is explicitly defined (or unambiguously standard for the domain, e.g., points/wins over all matches of a team) and whose computed values are shown to reproduce the spec's ordering/axis constraints.
- **Consequence**: The plotted values, and any serialized numeric result, differ from the reference; the image and the accompanying value/config files all fail comparison, yielding 0 of the expected checks even though the script "ran successfully".
394Unverified threshold-based outlier count (no independent recomputation or distribution sanity check)taskinfiagent-dabench
Applies when
task -- the task asks for a count of values exceeding a fixed statistical threshold (z-score, IQR, percentile) in one column, and the script computes it with a single library call on a self-chosen subset.
Pattern
The attempt calls one helper (e.g., scipy.stats.zscore) on a filtered/coerced version of the column (rows silently dropped, dtype auto-cast, sentinel/placeholder or duplicated rows left in), reports the resulting count directly, and never cross-checks it against the column's own summary statistics or a hand-computed formula. Any hidden mismatch (different mean/std base, wrong file/column, non-numeric or sentinel values inflating spread, different ddof/NaN policy) passes through undetected.
Detection procedure
  1. From the task, note the exact definition, threshold and population the statistic must be computed over (all rows of the stated table, as-is).
  2. In the script, check whether the analyzed subset/column was altered before the statistic (row filtering, notna(), type coercion, a different file among several) and whether the count is recomputed a second, independent way (explicit (x - mean)/std, or comparing min/max against mean ± k*std from describe()).
  3. Check whether the script prints and reasons about the extreme z-values and the implied tail fraction (count / N) versus what the threshold should admit (~0.3% for k=3 on roughly bell-shaped data), and whether flagged values are inspected to confirm they are genuine extremes rather than repeated codes/units artifacts.
  4. If the reported number rests on one uncorroborated library call, or the tail fraction / flagged values are printed but not reconciled with the distribution, flag the attempt.
Discriminator
A genuine violation is a single-path computation whose result is accepted without any reconciliation (no manual formula check, no mean ± kstd vs observed range comparison, no scrutiny of the flagged rows or of the subset change). It is not* a violation when the script computes the statistic on the population the task specifies and demonstrates agreement between at least two views (e.g., library z-scores and explicit formula, or count consistent with describe() bounds), even if the final count happens to be large.
Consequence
The reported integer is derived from a different population/definition than the reference, so the single exact-match check fails (e.g., a nonzero count submitted where the correct answer is a very different number), scoring 0 despite a script that runs cleanly.
id 05eb5c1c3fe3 · mined from infiagent-dabench dabench-361@s7
raw text (what the judge reads)
### Unverified threshold-based outlier count (no independent recomputation or distribution sanity check)
- **Applies when**: `task` -- the task asks for a count of values exceeding a fixed statistical threshold (z-score, IQR, percentile) in one column, and the script computes it with a single library call on a self-chosen subset.
- **Pattern**: The attempt calls one helper (e.g., `scipy.stats.zscore`) on a filtered/coerced version of the column (rows silently dropped, dtype auto-cast, sentinel/placeholder or duplicated rows left in), reports the resulting count directly, and never cross-checks it against the column's own summary statistics or a hand-computed formula. Any hidden mismatch (different mean/std base, wrong file/column, non-numeric or sentinel values inflating spread, different ddof/NaN policy) passes through undetected.
- **Detection procedure**:
  1. From the task, note the exact definition, threshold and population the statistic must be computed over (all rows of the stated table, as-is).
  2. In the script, check whether the analyzed subset/column was altered before the statistic (row filtering, `notna()`, type coercion, a different file among several) and whether the count is recomputed a second, independent way (explicit `(x - mean)/std`, or comparing `min`/`max` against `mean ± k*std` from `describe()`).
  3. Check whether the script prints and reasons about the extreme z-values and the implied tail fraction (count / N) versus what the threshold should admit (~0.3% for k=3 on roughly bell-shaped data), and whether flagged values are inspected to confirm they are genuine extremes rather than repeated codes/units artifacts.
  4. If the reported number rests on one uncorroborated library call, or the tail fraction / flagged values are printed but not reconciled with the distribution, flag the attempt.
- **Discriminator**: A genuine violation is a single-path computation whose result is accepted without any reconciliation (no manual formula check, no `mean ± k*std` vs observed range comparison, no scrutiny of the flagged rows or of the subset change). It is *not* a violation when the script computes the statistic on the population the task specifies and demonstrates agreement between at least two views (e.g., library z-scores and explicit formula, or count consistent with `describe()` bounds), even if the final count happens to be large.
- **Consequence**: The reported integer is derived from a different population/definition than the reference, so the single exact-match check fails (e.g., a nonzero count submitted where the correct answer is a very different number), scoring 0 despite a script that runs cleanly.
395Guessing a transformation spec instead of reading the referenced auxiliary filetaskda-code
Applies when
task -- The task instructs the agent to apply a mapping, filter, definition, or parameter set that lives in a separate provided file (README, tips/notes, config, data dictionary) before computing the requested statistic or model output.
Pattern
The scripts never open or parse the referenced file; instead the agent hardcodes a "standard"/"common convention" version of the mapping from prior knowledge, and the downstream label names, groupings, or category counts are then unverifiable against the actual spec (and the required output artifact may also be skipped or written in the wrong format).
Detection procedure
  1. Read the task statement and list every external file it says the transformation/definition must come from, plus the exact output artifact and format demanded.
  2. Grep the scripts for a read of each such file (open, read_csv, read_text, etc.); check whether the mapping/parameters used are derived from that content or are literal dicts written by the agent, and note any comment like "based on common conventions".
  3. Check that the final values (label strings, category collapsing, denominator, rounding) are consistent with the spec file's wording, and that the answer is also persisted to the required file with the required keys.
  4. If the spec was never read, treat all reported labels/ratios as unvalidated even if the numeric computation itself looks correct.
Discriminator
A real violation is inferring the spec's content from memory or the data alone; it is fine if the script reads the file (or the reviewer can see its contents quoted/printed in the run log) and the hardcoded values are shown to match it verbatim — including whether the mapping merges categories, which would change both the winning label and its ratio.
Consequence
The reported category name (and possibly the ratio, if the true mapping merges or renames groups) mismatches the expected answer, and the required result file is missing or wrong, so the check fails outright.
id ca0e8ab21bd2 · mined from da-code dacode-di-text-004@s7
raw text (what the judge reads)
### Guessing a transformation spec instead of reading the referenced auxiliary file
- **Applies when**: `task` -- The task instructs the agent to apply a mapping, filter, definition, or parameter set that lives in a separate provided file (README, tips/notes, config, data dictionary) before computing the requested statistic or model output.
- **Pattern**: The scripts never open or parse the referenced file; instead the agent hardcodes a "standard"/"common convention" version of the mapping from prior knowledge, and the downstream label names, groupings, or category counts are then unverifiable against the actual spec (and the required output artifact may also be skipped or written in the wrong format).
- **Detection procedure**:
  1. Read the task statement and list every external file it says the transformation/definition must come from, plus the exact output artifact and format demanded.
  2. Grep the scripts for a read of each such file (`open`, `read_csv`, `read_text`, etc.); check whether the mapping/parameters used are derived from that content or are literal dicts written by the agent, and note any comment like "based on common conventions".
  3. Check that the final values (label strings, category collapsing, denominator, rounding) are consistent with the spec file's wording, and that the answer is also persisted to the required file with the required keys.
  4. If the spec was never read, treat all reported labels/ratios as unvalidated even if the numeric computation itself looks correct.
- **Discriminator**: A real violation is inferring the spec's content from memory or the data alone; it is fine if the script reads the file (or the reviewer can see its contents quoted/printed in the run log) and the hardcoded values are shown to match it verbatim — including whether the mapping merges categories, which would change both the winning label and its ratio.
- **Consequence**: The reported category name (and possibly the ratio, if the true mapping merges or renames groups) mismatches the expected answer, and the required result file is missing or wrong, so the check fails outright.
396Auxiliary train-only metadata used as model features (merge yields all-missing columns for test)taskda-code
Applies when
task -- the provided data includes a supplementary/side file keyed to training rows (labels, metadata, post-hoc annotations) that the scripts join onto both train and test before building the feature matrix.
Pattern
The script merges the auxiliary table onto the test set the same way it does for train, without verifying that the auxiliary keys cover test rows; for test every merged column is NaN and is silently filled with a placeholder/median, so the model is trained on informative side-columns that are constant-missing at prediction time. The mismatch is never caught because there is no held-out validation using the competition metric and no post-merge sanity check on missing counts.
Detection procedure
  1. Read the task/README to see which files are stated as available for the test rows versus only for training rows.
  2. In the scripts, locate every merge/join and the construction of the feature list; check whether any merged column that is not explicitly excluded ends up in the features, and whether the code asserts key coverage (e.g., match rate, NaN count) after the join on test.
  3. Check whether imputation is applied blindly to the whole test frame (fillna with train median/'Unknown'), which would mask a 100%-missing column.
  4. Check whether any cross-validated/held-out estimate of the required evaluation metric is computed on data that reproduces the test-time feature availability; absence means the defect could not have been detected.
Discriminator
A real violation is when auxiliary columns are genuinely absent/unmatched for test rows (join coverage ≈ 0 or the file only lists training ids) yet still enter X. It is fine if the auxiliary table covers test ids too, or if the script explicitly drops those columns from the feature list (excluding them by name) and only uses them for grouping/stratification.
Consequence
The model relies on features that are constant placeholders at inference, so test probabilities are essentially degenerate/miscalibrated relative to training performance; the balanced/weighted loss on the hidden labels comes out far worse than any reported internal number, and the submission is scored wrong.
id 05a83c2ff287 · mined from da-code dacode-ml-competition-003@s7
raw text (what the judge reads)
### Auxiliary train-only metadata used as model features (merge yields all-missing columns for test)
- **Applies when**: `task` -- the provided data includes a supplementary/side file keyed to training rows (labels, metadata, post-hoc annotations) that the scripts join onto both train and test before building the feature matrix.
- **Pattern**: The script merges the auxiliary table onto the test set the same way it does for train, without verifying that the auxiliary keys cover test rows; for test every merged column is NaN and is silently filled with a placeholder/median, so the model is trained on informative side-columns that are constant-missing at prediction time. The mismatch is never caught because there is no held-out validation using the competition metric and no post-merge sanity check on missing counts.
- **Detection procedure**:
  1. Read the task/README to see which files are stated as available for the test rows versus only for training rows.
  2. In the scripts, locate every merge/join and the construction of the feature list; check whether any merged column that is not explicitly excluded ends up in the features, and whether the code asserts key coverage (e.g., match rate, NaN count) after the join on test.
  3. Check whether imputation is applied blindly to the whole test frame (fillna with train median/`'Unknown'`), which would mask a 100%-missing column.
  4. Check whether any cross-validated/held-out estimate of the required evaluation metric is computed on data that reproduces the test-time feature availability; absence means the defect could not have been detected.
- **Discriminator**: A real violation is when auxiliary columns are genuinely absent/unmatched for test rows (join coverage ≈ 0 or the file only lists training ids) yet still enter `X`. It is fine if the auxiliary table covers test ids too, or if the script explicitly drops those columns from the feature list (excluding them by name) and only uses them for grouping/stratification.
- **Consequence**: The model relies on features that are constant placeholders at inference, so test probabilities are essentially degenerate/miscalibrated relative to training performance; the balanced/weighted loss on the hidden labels comes out far worse than any reported internal number, and the submission is scored wrong.
397Unchallenged near-perfect score with predictions that don't match the label spacetaskda-code
Applies when
task -- a script trains a classifier/regressor from a handful of engineered features and writes a prediction file, reporting training/CV performance that is at or near the ceiling.
Pattern
The attempt treats a suspiciously perfect (≈100%) validation score as success instead of as a red flag for leakage or a degenerate/misaligned target, and never sanity-checks the written predictions against the label space, class distribution, row count, and column spec of the target file — so it emits only a subset of the true categories (and/or extra columns) without noticing.
Detection procedure
  1. From the task/README, list the exact required output columns and, from the training labels, the full set of distinct categories and their approximate frequencies.
  2. In the scripts, find the reported accuracy/score and the features used; check whether any feature is derived from, or an alias of, the target, and whether the score is implausibly high given weak features (counts, aggregates) and a noisy real-world label.
  3. In the answer/output, compare the predicted category set and distribution, the row count, and the column names/order against step 1.
  4. Confirm the script contains an explicit assertion or printed check on these (columns, row count, category coverage); if absent, the attempt is unverified.
Discriminator
A genuine violation shows a ceiling score with no leakage investigation and output that deviates from the required schema/label space (missing categories, extra columns, wrong row count). It is fine if the high score is justified by a documented near-deterministic feature and the output schema and category coverage match the training label space and required format.
Consequence
The saved file fails schema/label comparison against the reference (missing categories, extra column, mismatched keys), so accuracy against ground truth collapses and the file is graded WRONG/MISSING despite the reported ~99.99% accuracy.
id 4d37c1e713ba · mined from da-code dacode-ml-multi-003@s7
raw text (what the judge reads)
### Unchallenged near-perfect score with predictions that don't match the label space
- **Applies when**: `task` -- a script trains a classifier/regressor from a handful of engineered features and writes a prediction file, reporting training/CV performance that is at or near the ceiling.
- **Pattern**: The attempt treats a suspiciously perfect (≈100%) validation score as success instead of as a red flag for leakage or a degenerate/misaligned target, and never sanity-checks the written predictions against the label space, class distribution, row count, and column spec of the target file — so it emits only a subset of the true categories (and/or extra columns) without noticing.
- **Detection procedure**:
  1. From the task/README, list the exact required output columns and, from the training labels, the full set of distinct categories and their approximate frequencies.
  2. In the scripts, find the reported accuracy/score and the features used; check whether any feature is derived from, or an alias of, the target, and whether the score is implausibly high given weak features (counts, aggregates) and a noisy real-world label.
  3. In the answer/output, compare the predicted category set and distribution, the row count, and the column names/order against step 1.
  4. Confirm the script contains an explicit assertion or printed check on these (columns, row count, category coverage); if absent, the attempt is unverified.
- **Discriminator**: A genuine violation shows a ceiling score with no leakage investigation *and* output that deviates from the required schema/label space (missing categories, extra columns, wrong row count). It is fine if the high score is justified by a documented near-deterministic feature and the output schema and category coverage match the training label space and required format.
- **Consequence**: The saved file fails schema/label comparison against the reference (missing categories, extra column, mismatched keys), so accuracy against ground truth collapses and the file is graded WRONG/MISSING despite the reported ~99.99% accuracy.
398Template/reference file used only for header matching, never for value-level validationtaskda-code
Applies when
task -- the task supplies a template or example output file (and/or analogous already-computed companion result files) that the deliverable's format and conventions must match.
Pattern
The scripts confirm only superficial structure — column names, row labels, shape — and then write their own numbers, without ever checking whether the pipeline reproduces any known reference values. Definitional choices (aggregation level, index/offset convention, rounding, whether to average raw rows vs. average per-entity first, filtering of nulls/negatives) are chosen by assumption and asserted as "matches the template".
Detection procedure
  1. From the task, note that a template/reference (or a sibling file produced by the same analytical recipe) exists and that the output "format must match".
  2. In the scripts, find every comparison against that reference and classify it: does it compare headers/labels/shape only, or does it compare actual cell values (or reproduce a companion file end-to-end with the same code path)?
  3. Check whether ambiguous conventions in the code (period offset starting at 1, round(1), mean over transaction rows, no exclusion of missing IDs / cancellations / non-positive values) are justified by any evidence from the reference rather than by comment-only assertions.
  4. In the answer, look for claims of verification/consistency that are backed only by label matching or by re-printing the agent's own numbers.
Discriminator
Fine if the agent recomputes at least one reference/companion quantity with its own pipeline and shows numeric agreement (or explicitly reconciles a documented difference); a violation if all "verification" is structural (columns/dates/shape equal) or self-referential (comparing the output to itself), leaving the aggregation definition and rounding untested.
Consequence
The file has the right shape and headers but wrong cell values, so an exact/tolerance value comparison marks the expected result file WRONG and the task scores 0 despite a plausible-looking narrative.
id 3748133c1fbf · mined from da-code dacode-dm-csv-044@s7
raw text (what the judge reads)
### Template/reference file used only for header matching, never for value-level validation
- **Applies when**: `task` -- the task supplies a template or example output file (and/or analogous already-computed companion result files) that the deliverable's format and conventions must match.
- **Pattern**: The scripts confirm only superficial structure — column names, row labels, shape — and then write their own numbers, without ever checking whether the pipeline reproduces any known reference values. Definitional choices (aggregation level, index/offset convention, rounding, whether to average raw rows vs. average per-entity first, filtering of nulls/negatives) are chosen by assumption and asserted as "matches the template".
- **Detection procedure**:
  1. From the task, note that a template/reference (or a sibling file produced by the same analytical recipe) exists and that the output "format must match".
  2. In the scripts, find every comparison against that reference and classify it: does it compare headers/labels/shape only, or does it compare actual cell values (or reproduce a companion file end-to-end with the same code path)?
  3. Check whether ambiguous conventions in the code (period offset starting at 1, `round(1)`, mean over transaction rows, no exclusion of missing IDs / cancellations / non-positive values) are justified by any evidence from the reference rather than by comment-only assertions.
  4. In the answer, look for claims of verification/consistency that are backed only by label matching or by re-printing the agent's own numbers.
- **Discriminator**: Fine if the agent recomputes at least one reference/companion quantity with its own pipeline and shows numeric agreement (or explicitly reconciles a documented difference); a violation if all "verification" is structural (columns/dates/shape equal) or self-referential (comparing the output to itself), leaving the aggregation definition and rounding untested.
- **Consequence**: The file has the right shape and headers but wrong cell values, so an exact/tolerance value comparison marks the expected result file WRONG and the task scores 0 despite a plausible-looking narrative.
399Optimizing/selecting models with a default metric instead of the competition's stated evaluation metrictaskda-code
Applies when
task -- the task explicitly names the scoring metric (e.g., an ordinal/weighted-agreement, ranking, or imbalanced-class metric) and the scripts do model selection, weighting, or thresholding.
Pattern
The scripts never implement or compute the stated metric; they rely on library defaults (accuracy, R², argmax of class probabilities, naive rounding) for cross-validation, ensemble weights, and final rounding, so the chosen model and its decision rule are tuned for the wrong objective — typically collapsing predictions onto the majority classes and losing credit on the minority ends that the stated metric rewards.
Detection procedure
  1. Read the task statement and write down the exact metric to be scored.
  2. Grep the scripts for that metric (or an explicit implementation of it) inside any cross_val_score/scoring=/model-comparison/threshold-search code; note whether any reported validation number is in that metric.
  3. Check how continuous outputs are converted to the required output type (plain round/argmax vs. cutoffs optimized against the stated metric) and whether the final class/label distribution is compared to the training distribution.
  4. Inspect the produced answer file's value distribution and row count/ID coverage against the expected submission spec.
Discriminator
A real violation is when no held-out estimate of the stated metric exists anywhere, so there is no evidence the chosen pipeline beats a trivial baseline on it; it is fine if the agent uses a proxy metric for speed but still reports and compares candidates on the stated metric (or proves the proxy is monotonically equivalent).
Consequence
The submission is syntactically valid but scores far below threshold on the stated metric (near-baseline or below the required cutoff), because predictions are concentrated in the frequent classes and the grader's metric check fails.
id 0d99ad93a2de · mined from da-code dacode-ml-competition-006@s7
raw text (what the judge reads)
### Optimizing/selecting models with a default metric instead of the competition's stated evaluation metric
- **Applies when**: `task` -- the task explicitly names the scoring metric (e.g., an ordinal/weighted-agreement, ranking, or imbalanced-class metric) and the scripts do model selection, weighting, or thresholding.
- **Pattern**: The scripts never implement or compute the stated metric; they rely on library defaults (accuracy, R², argmax of class probabilities, naive rounding) for cross-validation, ensemble weights, and final rounding, so the chosen model and its decision rule are tuned for the wrong objective — typically collapsing predictions onto the majority classes and losing credit on the minority ends that the stated metric rewards.
- **Detection procedure**:
  1. Read the task statement and write down the exact metric to be scored.
  2. Grep the scripts for that metric (or an explicit implementation of it) inside any `cross_val_score`/`scoring=`/model-comparison/threshold-search code; note whether any reported validation number is in that metric.
  3. Check how continuous outputs are converted to the required output type (plain `round`/`argmax` vs. cutoffs optimized against the stated metric) and whether the final class/label distribution is compared to the training distribution.
  4. Inspect the produced answer file's value distribution and row count/ID coverage against the expected submission spec.
- **Discriminator**: A real violation is when *no* held-out estimate of the stated metric exists anywhere, so there is no evidence the chosen pipeline beats a trivial baseline on it; it is fine if the agent uses a proxy metric for speed but still reports and compares candidates on the stated metric (or proves the proxy is monotonically equivalent).
- **Consequence**: The submission is syntactically valid but scores far below threshold on the stated metric (near-baseline or below the required cutoff), because predictions are concentrated in the frequent classes and the grader's metric check fails.
400Group split by "missingness" not validated against the raw file's null semanticstaskinfiagent-dabench
Applies when
task -- a task asks to compare a statistic between rows where some field is missing/null and rows where it is not, and the scripts build the two groups with a pandas/SQL null test after loading the data.
Pattern
The attempt loads the file with default parser settings and splits with a single isna()/notna() test, so placeholder values (empty strings, "NA", "None", "-", whitespace) or type coercions land in the wrong group; additionally, rows with missing values in the measured numeric column, or rows dropped by the loader, silently shift both means. No check confirms the two group sizes sum to the full row count or that the split matches the raw file.
Detection procedure
  1. Read the task to identify the partition field, the measured column, and the required statistic/format.
  2. In the scripts, find the load call and the split expression: check whether na_values/keep_default_na/dtype/delimiter/quoting are left implicit, whether string placeholders and blanks are normalized to null before testing, and whether any dropna()/filter is applied that could remove rows from either group.
  3. Confirm the script prints and the reviewer can verify: n_null + n_notnull == total rows, counts of non-null values in the measured column per group, and the group means — and that these counts are consistent with a quick independent count of blank/placeholder entries in the raw file.
  4. Check the reported numbers against the required formatting (e.g., a p-value rounded to the requested number of decimals rather than printed as 0.0/0), since a degenerate-looking output is a symptom of an unchecked pipeline.
Discriminator
A real violation is when the null test is applied to already-coerced/filtered data with no reconciliation of group counts to the raw row count, or when placeholders/blank strings are plausible in the file and never normalized. It is not a violation if the script explicitly declares null tokens (or verifies none exist), shows the two group sizes summing to the total, and documents how missing values in the measured column are treated.
Consequence
Both group means are shifted by a few percent (and any test statistic with them), so exact-value checks on the reported means fail even though the analysis method is nominally correct; the reported p-value may also be rejected for wrong precision/format.
id 7b7cba2aff1f · mined from infiagent-dabench dabench-297@s7
raw text (what the judge reads)
### Group split by "missingness" not validated against the raw file's null semantics
- **Applies when**: `task` -- a task asks to compare a statistic between rows where some field is missing/null and rows where it is not, and the scripts build the two groups with a pandas/SQL null test after loading the data.
- **Pattern**: The attempt loads the file with default parser settings and splits with a single `isna()`/`notna()` test, so placeholder values (empty strings, `"NA"`, `"None"`, `"-"`, whitespace) or type coercions land in the wrong group; additionally, rows with missing values in the *measured* numeric column, or rows dropped by the loader, silently shift both means. No check confirms the two group sizes sum to the full row count or that the split matches the raw file.
- **Detection procedure**:
  1. Read the task to identify the partition field, the measured column, and the required statistic/format.
  2. In the scripts, find the load call and the split expression: check whether `na_values`/`keep_default_na`/`dtype`/delimiter/quoting are left implicit, whether string placeholders and blanks are normalized to null before testing, and whether any `dropna()`/filter is applied that could remove rows from either group.
  3. Confirm the script prints and the reviewer can verify: `n_null + n_notnull == total rows`, counts of non-null values in the measured column per group, and the group means — and that these counts are consistent with a quick independent count of blank/placeholder entries in the raw file.
  4. Check the reported numbers against the required formatting (e.g., a p-value rounded to the requested number of decimals rather than printed as `0.0`/`0`), since a degenerate-looking output is a symptom of an unchecked pipeline.
- **Discriminator**: A real violation is when the null test is applied to already-coerced/filtered data with no reconciliation of group counts to the raw row count, or when placeholders/blank strings are plausible in the file and never normalized. It is *not* a violation if the script explicitly declares null tokens (or verifies none exist), shows the two group sizes summing to the total, and documents how missing values in the measured column are treated.
- **Consequence**: Both group means are shifted by a few percent (and any test statistic with them), so exact-value checks on the reported means fail even though the analysis method is nominally correct; the reported p-value may also be rejected for wrong precision/format.
401Ignoring the provided specification and its required output artifactstaskda-code
Applies when
task -- the task points to an external instruction/spec file (or explicitly lists deliverable files/values) and the script must implement exactly that procedure and emit exactly those outputs.
Pattern
The agent never opens/quotes the referenced spec, invents its own filtering thresholds and category definitions from intuition, and saves only the one artifact it happened to think of (e.g. an image), while other required deliverables (serialized numeric results, plot metadata) are never written.
Detection procedure
  1. From the task text, list every named instruction source and every required deliverable (file names, values, formats).
  2. Read the script for a load/parse of each instruction source; check that each filtering rule, grouping/category definition, and label set in the code is traceable to a quoted line of the spec rather than an ad-hoc heuristic (if 'x' in s ... else None, arbitrary day counts, self-chosen group names).
  3. Grep the script for a write of every deliverable on the list; confirm the save paths/names match exactly.
  4. Check the answer text: does it report the deliverables produced, or does it silently redefine the task ("Categorization Rules I applied…")?
Discriminator
A real violation is code whose rules and group labels have no basis in the given instructions, or that produces a strict subset of the requested files; it is fine if the script derives the same rules from the spec (even if phrased differently) and writes every requested artifact, or if the spec genuinely leaves a detail free and the agent documents the ambiguity.
Consequence
Graders that compare each expected artifact find them missing or containing numbers from a self-invented grouping, so every check fails even though the script runs without error.
id 10fff822f2b9 · mined from da-code dacode-plot-pie-005@s7
raw text (what the judge reads)
### Ignoring the provided specification and its required output artifacts
- **Applies when**: `task` -- the task points to an external instruction/spec file (or explicitly lists deliverable files/values) and the script must implement exactly that procedure and emit exactly those outputs.
- **Pattern**: The agent never opens/quotes the referenced spec, invents its own filtering thresholds and category definitions from intuition, and saves only the one artifact it happened to think of (e.g. an image), while other required deliverables (serialized numeric results, plot metadata) are never written.
- **Detection procedure**:
  1. From the task text, list every named instruction source and every required deliverable (file names, values, formats).
  2. Read the script for a load/parse of each instruction source; check that each filtering rule, grouping/category definition, and label set in the code is traceable to a quoted line of the spec rather than an ad-hoc heuristic (`if 'x' in s ... else None`, arbitrary day counts, self-chosen group names).
  3. Grep the script for a write of every deliverable on the list; confirm the save paths/names match exactly.
  4. Check the answer text: does it report the deliverables produced, or does it silently redefine the task ("Categorization Rules I applied…")?
- **Discriminator**: A real violation is code whose rules and group labels have no basis in the given instructions, or that produces a strict subset of the requested files; it is fine if the script derives the same rules from the spec (even if phrased differently) and writes every requested artifact, or if the spec genuinely leaves a detail free and the agent documents the ambiguity.
- **Consequence**: Graders that compare each expected artifact find them missing or containing numbers from a self-invented grouping, so every check fails even though the script runs without error.
402Missing-value handling and numeric coercion not verified against the raw column contentstaskinfiagent-dabench
Applies when
task -- the task names specific columns to use as features/target and prescribes an exact missing-value policy (e.g., mean imputation) before fitting and scoring a model.
Pattern
The attempt applies the prescribed imputation only nominally — it reads columns that are stored as text/mixed types (units, symbols, separators, ranges, sentinel codes), lets to_numeric(..., errors='coerce') or a silent parse turn many valid entries into NaN, then fills them with a mean computed from the surviving subset (or drops rows instead of imputing). The resulting feature/target distribution is badly distorted, but no check on dtypes, NaN counts, row counts, or value ranges is performed before reporting the metric.
Detection procedure
  1. From the task, list the exact columns involved and the exact missing-value rule (impute vs drop, which statistic, computed on what).
  2. In the scripts, find where each column is loaded/converted: check whether raw dtypes are inspected, whether any string cleaning occurs, and how many values become NaN after conversion versus how many were missing in the raw file.
  3. Check that after handling, the row count equals the original row count (imputation, not deletion), that no NaNs remain, and that the target's mean/range is plausible for the quantity being modeled.
  4. Compare the reported error metric to a trivial baseline (variance of the target / MSE of predicting the target mean); if the model's MSE is of the same order or larger, demand an explanation.
Discriminator
A real violation shows silent NaN creation, shrunken row counts, or unchecked object dtypes on the modeled columns; a look-alike that is fine explicitly parses the raw representation (strips units/symbols, maps sentinels), documents the number of originally-missing values, and preserves the full row count with all-numeric, in-range columns.
Consequence
The model is fit on a corrupted or truncated version of the data, so the reported MSE is off by a large factor from the reference value and the answer is graded wrong even though the pipeline and output format look correct.
id ac0090a1719d · mined from infiagent-dabench dabench-432@s7
raw text (what the judge reads)
### Missing-value handling and numeric coercion not verified against the raw column contents
- **Applies when**: `task` -- the task names specific columns to use as features/target and prescribes an exact missing-value policy (e.g., mean imputation) before fitting and scoring a model.
- **Pattern**: The attempt applies the prescribed imputation only nominally — it reads columns that are stored as text/mixed types (units, symbols, separators, ranges, sentinel codes), lets `to_numeric(..., errors='coerce')` or a silent parse turn many valid entries into NaN, then fills them with a mean computed from the surviving subset (or drops rows instead of imputing). The resulting feature/target distribution is badly distorted, but no check on dtypes, NaN counts, row counts, or value ranges is performed before reporting the metric.
- **Detection procedure**:
  1. From the task, list the exact columns involved and the exact missing-value rule (impute vs drop, which statistic, computed on what).
  2. In the scripts, find where each column is loaded/converted: check whether raw dtypes are inspected, whether any string cleaning occurs, and how many values become NaN after conversion versus how many were missing in the raw file.
  3. Check that after handling, the row count equals the original row count (imputation, not deletion), that no NaNs remain, and that the target's mean/range is plausible for the quantity being modeled.
  4. Compare the reported error metric to a trivial baseline (variance of the target / MSE of predicting the target mean); if the model's MSE is of the same order or larger, demand an explanation.
- **Discriminator**: A real violation shows silent NaN creation, shrunken row counts, or unchecked object dtypes on the modeled columns; a look-alike that is fine explicitly parses the raw representation (strips units/symbols, maps sentinels), documents the number of originally-missing values, and preserves the full row count with all-numeric, in-range columns.
- **Consequence**: The model is fit on a corrupted or truncated version of the data, so the reported MSE is off by a large factor from the reference value and the answer is graded wrong even though the pipeline and output format look correct.
403Deliverable file not written in the exact requested schemataskda-code
Applies when
task -- the task requires the result be persisted to a named output file with a prescribed row/column structure, and the agent instead reports findings in prose.
Pattern
The attempt performs the analysis correctly-ish but never produces (or overwrites/mis-shapes) the specified artifact: no code path writes the file, or it writes multiple rows / a nested list-of-dicts / extra columns instead of the single required row with the required fields.
Detection procedure
  1. From the task, list the exact required artifact name and its schema (number of rows, field names/order, value formats).
  2. Search the scripts for a write call targeting that filename; confirm the object written is a single-row table with exactly the required fields, not a list, dict-of-dicts, or multi-row frame.
  3. Check the final answer: does it point to the saved file and echo its contents, or is it only a narrative report of statistics?
  4. If any required field (e.g., test name, decision, p-value, comment) is absent from the written row, or the write step is missing entirely, flag it.
Discriminator
A real violation is missing/mis-shaped persistence of the named artifact; a look-alike that is fine is a script that writes the correct one-row file and additionally prints a verbose summary for the human reader.
Consequence
The grader finds the expected output file missing or structurally wrong and scores 0, regardless of whether the underlying statistical test and p-value were right.
id 7227cda40754 · mined from da-code dacode-data-sa-004@s7
raw text (what the judge reads)
### Deliverable file not written in the exact requested schema
- **Applies when**: `task` -- the task requires the result be persisted to a named output file with a prescribed row/column structure, and the agent instead reports findings in prose.
- **Pattern**: The attempt performs the analysis correctly-ish but never produces (or overwrites/mis-shapes) the specified artifact: no code path writes the file, or it writes multiple rows / a nested list-of-dicts / extra columns instead of the single required row with the required fields.
- **Detection procedure**:
  1. From the task, list the exact required artifact name and its schema (number of rows, field names/order, value formats).
  2. Search the scripts for a write call targeting that filename; confirm the object written is a single-row table with exactly the required fields, not a list, dict-of-dicts, or multi-row frame.
  3. Check the final answer: does it point to the saved file and echo its contents, or is it only a narrative report of statistics?
  4. If any required field (e.g., test name, decision, p-value, comment) is absent from the written row, or the write step is missing entirely, flag it.
- **Discriminator**: A real violation is missing/mis-shaped persistence of the named artifact; a look-alike that is fine is a script that writes the correct one-row file and *additionally* prints a verbose summary for the human reader.
- **Consequence**: The grader finds the expected output file missing or structurally wrong and scores 0, regardless of whether the underlying statistical test and p-value were right.
404Unverified row set / dtype handling for a correlation-type statistic reported to fixed precisiontaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic (correlation, mean, ratio, metric) computed over two or more columns of a table and reported rounded to a fixed number of decimals.
Pattern
The agent computes the statistic with a one-liner (or an unsaved/ad-hoc script) without explicitly establishing which rows enter the computation — silently dropping rows via pairwise NA handling, coercion failures on non-numeric/placeholder values ("NA", "", "-1"), duplicated or header/index rows, or an unintended subset/merge — so the value is close to but not equal to the correct one, and no count/shape sanity check or reproducible script backs it up.
Detection procedure
  1. Read the task: note that the required output is rounded (e.g., 2 or 4 decimals), so small population differences change the graded digit; note whether any filtering is specified (if none, the full table is intended).
  2. Read the scripts (or note their absence): check that the data is loaded once from the full source, that the two columns are explicitly cast to numeric with failures inspected rather than silently coerced/dropped, and that the number of rows actually used is printed alongside the statistic.
  3. Check for an explicit sanity report: N used vs N total in file, count of NaNs/non-numeric entries per column, and the statistic computed at full precision before rounding.
  4. Compare the reported value's precision and format to the spec (e.g., a p-value required to 4 decimals must not be reported as 0.0), and confirm the same computed object is what is reported (not an intermediate or a different variant of the statistic).
Discriminator
A real violation is an attempt where the reviewer cannot tell, from the scripts/output alone, how many rows and which rows produced the number, or where non-numeric/missing values are dropped without being counted and justified. It is not a violation if the script prints row counts and NA handling and those match the full dataset (or a filter the task requires) — a legitimate value at that N is fine even if it differs from a naive computation.
Consequence
The rounded statistic lands one unit off in the last graded digit (e.g., 0.53 instead of 0.54) and/or the auxiliary value fails the required precision format, so the answer is marked wrong despite the qualitative conclusion being right, with no saved script to diagnose or reproduce.
id fa30069adab5 · mined from infiagent-dabench dabench-300@s7
raw text (what the judge reads)
### Unverified row set / dtype handling for a correlation-type statistic reported to fixed precision
- **Applies when**: `task` -- the task asks for a single numeric statistic (correlation, mean, ratio, metric) computed over two or more columns of a table and reported rounded to a fixed number of decimals.
- **Pattern**: The agent computes the statistic with a one-liner (or an unsaved/ad-hoc script) without explicitly establishing which rows enter the computation — silently dropping rows via pairwise NA handling, coercion failures on non-numeric/placeholder values ("NA", "", "-1"), duplicated or header/index rows, or an unintended subset/merge — so the value is close to but not equal to the correct one, and no count/shape sanity check or reproducible script backs it up.
- **Detection procedure**:
  1. Read the task: note that the required output is rounded (e.g., 2 or 4 decimals), so small population differences change the graded digit; note whether any filtering is specified (if none, the full table is intended).
  2. Read the scripts (or note their absence): check that the data is loaded once from the full source, that the two columns are explicitly cast to numeric with failures inspected rather than silently coerced/dropped, and that the number of rows actually used is printed alongside the statistic.
  3. Check for an explicit sanity report: N used vs N total in file, count of NaNs/non-numeric entries per column, and the statistic computed at full precision before rounding.
  4. Compare the reported value's precision and format to the spec (e.g., a p-value required to 4 decimals must not be reported as `0.0`), and confirm the same computed object is what is reported (not an intermediate or a different variant of the statistic).
- **Discriminator**: A real violation is an attempt where the reviewer cannot tell, from the scripts/output alone, how many rows and which rows produced the number, or where non-numeric/missing values are dropped without being counted and justified. It is *not* a violation if the script prints row counts and NA handling and those match the full dataset (or a filter the task requires) — a legitimate value at that N is fine even if it differs from a naive computation.
- **Consequence**: The rounded statistic lands one unit off in the last graded digit (e.g., 0.53 instead of 0.54) and/or the auxiliary value fails the required precision format, so the answer is marked wrong despite the qualitative conclusion being right, with no saved script to diagnose or reproduce.
405Feature matrix not verified to be identical (names + order) between train and predicttaskda-code
Applies when
task -- a script builds a training design matrix by dropping the target from one file and a prediction matrix from a separate file, then calls .fit/.predict without asserting the two matrices agree.
Pattern
The agent assumes the held-out file has exactly the same columns in the same order as the training file. Because most estimators match features positionally, any extra/missing/reordered column silently shifts values into the wrong feature slots, yet training and validation still succeed and print a respectable score (validation is computed only on the training file, so the mismatch never surfaces).
Detection procedure
  1. In the task/README, list the fields expected in the input files and note whether the prediction file is guaranteed to be a column-identical copy minus the target (usually it is not verified).
  2. In the scripts, look for an explicit check such as assert list(X_train.columns) == list(X_test.columns), a reindex/X_test = X_test[X_train.columns], or a pipeline/ColumnTransformer that selects by name; also check for printing of both column lists and dtypes.
  3. If no such alignment step exists, flag it; then check whether the answer's diagnostics are consistent with domain sense (e.g., feature-importance ranking, prediction min/max/mean vs. the target distribution in the training file) — implausible ordering or an inflated prediction range is corroborating evidence of shifted columns.
  4. Confirm the answer reports only in-file validation metrics and no cross-file sanity check (row count matches the prediction file, predicted distribution comparable to observed target distribution).
Discriminator
A real violation is when nothing in the code guarantees positional/name agreement (no reindexing, no name-based pipeline, no assertion, no printed comparison). It is not a violation if the script explicitly reindexes/selects the prediction frame by the training columns, or encodes both frames through one shared name-based transformer, even if it never prints the column lists.
Consequence
Predictions are generated for the right number of rows with the right header, so the file looks valid, but the values come from a mis-mapped feature space; the grader's held-out error is far worse than the reported validation error and the file is marked wrong.
id 59014535b9af · mined from da-code dacode-ml-regression-014@s7
raw text (what the judge reads)
### Feature matrix not verified to be identical (names + order) between train and predict
- **Applies when**: `task` -- a script builds a training design matrix by dropping the target from one file and a prediction matrix from a separate file, then calls `.fit`/`.predict` without asserting the two matrices agree.
- **Pattern**: The agent assumes the held-out file has exactly the same columns in the same order as the training file. Because most estimators match features *positionally*, any extra/missing/reordered column silently shifts values into the wrong feature slots, yet training and validation still succeed and print a respectable score (validation is computed only on the training file, so the mismatch never surfaces).
- **Detection procedure**:
  1. In the task/README, list the fields expected in the input files and note whether the prediction file is guaranteed to be a column-identical copy minus the target (usually it is not verified).
  2. In the scripts, look for an explicit check such as `assert list(X_train.columns) == list(X_test.columns)`, a reindex/`X_test = X_test[X_train.columns]`, or a pipeline/`ColumnTransformer` that selects by name; also check for printing of both column lists and dtypes.
  3. If no such alignment step exists, flag it; then check whether the answer's diagnostics are consistent with domain sense (e.g., feature-importance ranking, prediction min/max/mean vs. the target distribution in the training file) — implausible ordering or an inflated prediction range is corroborating evidence of shifted columns.
  4. Confirm the answer reports only in-file validation metrics and no cross-file sanity check (row count matches the prediction file, predicted distribution comparable to observed target distribution).
- **Discriminator**: A real violation is when nothing in the code guarantees positional/name agreement (no reindexing, no name-based pipeline, no assertion, no printed comparison). It is *not* a violation if the script explicitly reindexes/selects the prediction frame by the training columns, or encodes both frames through one shared name-based transformer, even if it never prints the column lists.
- **Consequence**: Predictions are generated for the right number of rows with the right header, so the file looks valid, but the values come from a mis-mapped feature space; the grader's held-out error is far worse than the reported validation error and the file is marked wrong.
406Required output artifacts and external spec file ignoredtaskda-code
Applies when
task -- the task points to a configuration/spec file for output formatting and/or the deliverable is a set of saved artifact files (plot, array, config dump) rather than only a textual answer.
Pattern
The attempt computes an intermediate statistic, writes one obvious output file, and reports the numbers in prose, without ever parsing the referenced spec file or emitting every artifact the spec/task implies (e.g. the serialized plot parameters and the underlying values array), so graded files are missing or built with default styling/labels/order instead of the mandated ones.
Detection procedure
  1. Read the task and the referenced spec/config file; list every required artifact (file name, format) and every stated formatting constraint (labels, ordering, units, rounding, colors, title).
  2. Read the scripts and check that the spec file is actually loaded and each of its keys is applied, not hardcoded or ignored.
  3. Check that each listed artifact is written to the expected path/format, and that saved numeric artifacts contain the requested quantity (e.g. the plotted values) with plausible shape/length.
  4. Compare the reported answer against the artifact list: prose-only reporting of an intermediate value while artifacts are absent is a violation.
Discriminator
A real violation is missing/unparsed spec or missing artifact files; it is fine if the script writes all required files and applies each spec key, even if the styling code looks verbose or the prose summary is brief.
Consequence
Grader checks for the expected files fail as WRONG/MISSING (here all three), so the run scores zero regardless of whether the intermediate statistic was right.
id 39a433982668 · mined from da-code dacode-plot-pie-008@s7
raw text (what the judge reads)
### Required output artifacts and external spec file ignored
- **Applies when**: `task` -- the task points to a configuration/spec file for output formatting and/or the deliverable is a set of saved artifact files (plot, array, config dump) rather than only a textual answer.
- **Pattern**: The attempt computes an intermediate statistic, writes one obvious output file, and reports the numbers in prose, without ever parsing the referenced spec file or emitting every artifact the spec/task implies (e.g. the serialized plot parameters and the underlying values array), so graded files are missing or built with default styling/labels/order instead of the mandated ones.
- **Detection procedure**:
  1. Read the task and the referenced spec/config file; list every required artifact (file name, format) and every stated formatting constraint (labels, ordering, units, rounding, colors, title).
  2. Read the scripts and check that the spec file is actually loaded and each of its keys is applied, not hardcoded or ignored.
  3. Check that each listed artifact is written to the expected path/format, and that saved numeric artifacts contain the requested quantity (e.g. the plotted values) with plausible shape/length.
  4. Compare the reported answer against the artifact list: prose-only reporting of an intermediate value while artifacts are absent is a violation.
- **Discriminator**: A real violation is missing/unparsed spec or missing artifact files; it is fine if the script writes all required files and applies each spec key, even if the styling code looks verbose or the prose summary is brief.
- **Consequence**: Grader checks for the expected files fail as WRONG/MISSING (here all three), so the run scores zero regardless of whether the intermediate statistic was right.
407Ignoring the provided output template / spec when defining categories and schemataskda-code
Applies when
task -- the task says to fill in a provided result file "adhering strictly to its format", or a spec/README defines the categories, labels, or thresholds to use for grouping.
Pattern
The scripts never read the supplied template file (or the auxiliary definition file) before writing; instead the agent hard-codes its own bucket boundaries, category names, extra catch-all rows, and column headers, then overwrites the template with a self-invented schema.
Detection procedure
  1. From the task text, list every artifact the agent is told to fill or follow (template file, format description, definition/spec document) and the exact fields it implies.
  2. Scan the scripts for a read of that template/spec (e.g., loading the file and printing its header and existing row labels); if the only reference is a write, or the categories/thresholds appear as literals invented in code, flag it.
  3. Compare the written file's column names, row labels, row count, and row order against the template's originals — any renamed column, added "unknown/other" row, or reordered/missing category is a violation.
  4. Check the answer text for whether the agent verified its group labels match the required ones rather than just reporting its own counts.
Discriminator
Fine if the agent inspects the template/spec and its hard-coded labels and thresholds are shown to match it exactly (or are trivially derived from it); a violation if the labels/boundaries come from outside knowledge or convenience with no verification, or if rows/columns not present in the template are introduced.
Consequence
The output file fails exact-format/value comparison — mismatched keys or extra rows make every checked cell wrong, scoring 0 even if the underlying counting logic was reasonable.
id b4d352b3b2e1 · mined from da-code dacode-dm-csv-001@s7
raw text (what the judge reads)
### Ignoring the provided output template / spec when defining categories and schema
- **Applies when**: `task` -- the task says to fill in a provided result file "adhering strictly to its format", or a spec/README defines the categories, labels, or thresholds to use for grouping.
- **Pattern**: The scripts never read the supplied template file (or the auxiliary definition file) before writing; instead the agent hard-codes its own bucket boundaries, category names, extra catch-all rows, and column headers, then overwrites the template with a self-invented schema.
- **Detection procedure**:
  1. From the task text, list every artifact the agent is told to fill or follow (template file, format description, definition/spec document) and the exact fields it implies.
  2. Scan the scripts for a read of that template/spec (e.g., loading the file and printing its header and existing row labels); if the only reference is a write, or the categories/thresholds appear as literals invented in code, flag it.
  3. Compare the written file's column names, row labels, row count, and row order against the template's originals — any renamed column, added "unknown/other" row, or reordered/missing category is a violation.
  4. Check the answer text for whether the agent verified its group labels match the required ones rather than just reporting its own counts.
- **Discriminator**: Fine if the agent inspects the template/spec and its hard-coded labels and thresholds are shown to match it exactly (or are trivially derived from it); a violation if the labels/boundaries come from outside knowledge or convenience with no verification, or if rows/columns not present in the template are introduced.
- **Consequence**: The output file fails exact-format/value comparison — mismatched keys or extra rows make every checked cell wrong, scoring 0 even if the underlying counting logic was reasonable.
408Correlation computed on raw observation rows instead of one aggregated record per entitytaskinfiagent-dabench
Applies when
task -- the task asks for a statistic (e.g., correlation) between per-entity attributes such as an extreme value, a duration/span, or a total, while the source table stores multiple time-stamped rows per entity.
Pattern
The attempt runs the statistic (and any median/threshold split) directly on the row-level table, or aggregates only some of the needed quantities, so each entity is weighted by its number of rows and the "duration" / "maximum" / "total" is not a single value per entity. The resulting r and p-value are plausible-looking but computed on the wrong unit of analysis, and both subgroup values drift from the intended ones.
Detection procedure
  1. From the task wording, identify the unit of analysis (one point per entity) and the per-entity quantities that must be derived (max, min–max span, sum, category label).
  2. In the scripts, check for an explicit groupby(entity_id) producing exactly one row per entity, with each requested quantity computed by the right aggregation (max for extremes, last−first timestamp converted to the stated unit for duration, and a single consistent value for the split variable).
  3. Verify the threshold split (e.g., median) and the correlation are both applied to that aggregated frame, and check the reported n / group sizes are on the order of the number of entities, not the number of raw rows.
  4. Confirm the answer's r values were not left unvalidated: a quick sanity check of group counts and value ranges (durations non-negative, category within valid bounds) should appear or be reproducible.
Discriminator
A real violation is when no per-entity collapse happens before the statistic, or when the split/duration is derived from row-level values; it is fine if the dataset genuinely has one row per entity, or if the script aggregates first and then merges the split variable back at entity level.
Consequence
The correlation coefficients (and subgroup membership) are shifted by row-count weighting, so the rounded r values fail exact-match checks even though the qualitative relationship type may still pass.
id c71f96fca480 · mined from infiagent-dabench dabench-431@s7
raw text (what the judge reads)
### Correlation computed on raw observation rows instead of one aggregated record per entity
- **Applies when**: `task` -- the task asks for a statistic (e.g., correlation) between per-entity attributes such as an extreme value, a duration/span, or a total, while the source table stores multiple time-stamped rows per entity.
- **Pattern**: The attempt runs the statistic (and any median/threshold split) directly on the row-level table, or aggregates only some of the needed quantities, so each entity is weighted by its number of rows and the "duration" / "maximum" / "total" is not a single value per entity. The resulting r and p-value are plausible-looking but computed on the wrong unit of analysis, and both subgroup values drift from the intended ones.
- **Detection procedure**:
  1. From the task wording, identify the unit of analysis (one point per entity) and the per-entity quantities that must be derived (max, min–max span, sum, category label).
  2. In the scripts, check for an explicit `groupby(entity_id)` producing exactly one row per entity, with each requested quantity computed by the right aggregation (max for extremes, last−first timestamp converted to the stated unit for duration, and a single consistent value for the split variable).
  3. Verify the threshold split (e.g., median) and the correlation are both applied to that aggregated frame, and check the reported n / group sizes are on the order of the number of entities, not the number of raw rows.
  4. Confirm the answer's r values were not left unvalidated: a quick sanity check of group counts and value ranges (durations non-negative, category within valid bounds) should appear or be reproducible.
- **Discriminator**: A real violation is when no per-entity collapse happens before the statistic, or when the split/duration is derived from row-level values; it is fine if the dataset genuinely has one row per entity, or if the script aggregates first and then merges the split variable back at entity level.
- **Consequence**: The correlation coefficients (and subgroup membership) are shifted by row-count weighting, so the rounded r values fail exact-match checks even though the qualitative relationship type may still pass.
409No held-out validation and no sanity check of the predicted class distribution against the training base ratetaskda-code
Applies when
task -- the task asks for predicted labels on an unlabeled test set to be written to a file, and the scripts fit models and dump hard 0/1 (or class) predictions directly.
Pattern
The attempt trains one or more models on the full labeled data, applies a default decision rule (e.g., argmax / probability > 0.5) with no cross-validation or hold-out estimate of the target metric, and never compares the resulting predicted positive rate to the observed base rate in the training labels or the requested output schema. With class imbalance and an averaged/regularized ensemble, probabilities shrink toward the majority class, so the minority class is heavily under-predicted and per-class metrics collapse.
Detection procedure
  1. Read the task/README for the requested output (exact column name(s), row count, allowed values, ordering) and note the class balance of the label column in the training data.
  2. In the scripts, check whether any hold-out split or CV is used to estimate the scoring metric before generating final predictions, and whether the decision threshold/class weighting is chosen rather than left at the library default.
  3. In the answer/output, compare the reported predicted positive rate with the training base rate, and compare the written file's columns/row count with the requested format.
  4. Flag if there is no validation score reported, or if the predicted minority-class rate deviates from the training base rate by a large factor (e.g., roughly half or double) with no justification, or if the file contains columns/order not asked for.
Discriminator
A real violation is an unvalidated pipeline whose prediction distribution is grossly off from the label prior with no stated reason; it is fine if the agent reports a hold-out/CV score for the relevant metric and either shows the predicted rate is close to the prior or explicitly justifies deviating (e.g., threshold tuned to maximize the stated metric on validation data).
Consequence
The submitted prediction file scores below the grader's accuracy/F1 threshold (or fails a format/column check), so the expected result file is marked WRONG even though the pipeline "ran successfully".
id 59dff6863e4b · mined from da-code dacode-ml-binary-016@s7
raw text (what the judge reads)
### No held-out validation and no sanity check of the predicted class distribution against the training base rate
- **Applies when**: `task` -- the task asks for predicted labels on an unlabeled test set to be written to a file, and the scripts fit models and dump hard 0/1 (or class) predictions directly.
- **Pattern**: The attempt trains one or more models on the full labeled data, applies a default decision rule (e.g., argmax / probability > 0.5) with no cross-validation or hold-out estimate of the target metric, and never compares the resulting predicted positive rate to the observed base rate in the training labels or the requested output schema. With class imbalance and an averaged/regularized ensemble, probabilities shrink toward the majority class, so the minority class is heavily under-predicted and per-class metrics collapse.
- **Detection procedure**:
  1. Read the task/README for the requested output (exact column name(s), row count, allowed values, ordering) and note the class balance of the label column in the training data.
  2. In the scripts, check whether any hold-out split or CV is used to estimate the scoring metric before generating final predictions, and whether the decision threshold/class weighting is chosen rather than left at the library default.
  3. In the answer/output, compare the reported predicted positive rate with the training base rate, and compare the written file's columns/row count with the requested format.
  4. Flag if there is no validation score reported, or if the predicted minority-class rate deviates from the training base rate by a large factor (e.g., roughly half or double) with no justification, or if the file contains columns/order not asked for.
- **Discriminator**: A real violation is an unvalidated pipeline whose prediction distribution is grossly off from the label prior with no stated reason; it is fine if the agent reports a hold-out/CV score for the relevant metric and either shows the predicted rate is close to the prior or explicitly justifies deviating (e.g., threshold tuned to maximize the stated metric on validation data).
- **Consequence**: The submitted prediction file scores below the grader's accuracy/F1 threshold (or fails a format/column check), so the expected result file is marked WRONG even though the pipeline "ran successfully".
410Fabricated/synthetic fallback data instead of locating the provided inputstaskda-code
Applies when
task -- the task references a supplied dataset (and/or a sample output template) and the script guesses at file paths rather than confirming what files actually exist.
Pattern
The script probes a short hard-coded list of candidate filenames and, on failure, silently generates simulated data (random draws or hand-typed numbers) and reports statistics from it as if they were the real answer; the provided format template is likewise never read, so column names/shape are invented.
Detection procedure
  1. Read the task and note which input artifacts are promised (data file, sample result/format file) — these must be discovered by listing the working/data directories, not assumed.
  2. Scan the scripts for any if not found: ... np.random.* / literal arrays / "demonstration" branch, and for whether any template file is opened and its columns reused.
  3. Check the answer/logs for evidence the fallback branch actually ran (e.g., messages about no data found, "synthetic", or values that could only come from a generator with the fixed seed).
  4. Confirm the reported numbers trace back to a file that exists on disk with the expected row counts; if not, the result is unverifiable.
Discriminator
A genuine violation is computing and reporting the final answer from data the script itself invented, or from a self-invented output schema, without ever enumerating the filesystem to find the real inputs. Not a violation: a robust path search that succeeds and loads real data, or a simulation branch used only for unit-testing that is not the source of the reported result.
Consequence
The reported statistic and the output file's columns/values are unrelated to the true dataset, so the expected result file is judged WRONG/MISSING and the check fails outright.
id d93c99c59b37 · mined from da-code dacode-data-sa-039@s7
raw text (what the judge reads)
### Fabricated/synthetic fallback data instead of locating the provided inputs
- **Applies when**: `task` -- the task references a supplied dataset (and/or a sample output template) and the script guesses at file paths rather than confirming what files actually exist.
- **Pattern**: The script probes a short hard-coded list of candidate filenames and, on failure, silently generates simulated data (random draws or hand-typed numbers) and reports statistics from it as if they were the real answer; the provided format template is likewise never read, so column names/shape are invented.
- **Detection procedure**:
  1. Read the task and note which input artifacts are promised (data file, sample result/format file) — these must be discovered by listing the working/data directories, not assumed.
  2. Scan the scripts for any `if not found: ... np.random.*` / literal arrays / "demonstration" branch, and for whether any template file is opened and its columns reused.
  3. Check the answer/logs for evidence the fallback branch actually ran (e.g., messages about no data found, "synthetic", or values that could only come from a generator with the fixed seed).
  4. Confirm the reported numbers trace back to a file that exists on disk with the expected row counts; if not, the result is unverifiable.
- **Discriminator**: A genuine violation is computing and reporting the final answer from data the script itself invented, or from a self-invented output schema, without ever enumerating the filesystem to find the real inputs. Not a violation: a robust path search that succeeds and loads real data, or a simulation branch used only for unit-testing that is not the source of the reported result.
- **Consequence**: The reported statistic and the output file's columns/values are unrelated to the true dataset, so the expected result file is judged WRONG/MISSING and the check fails outright.
411Prediction file shape/alignment not verified against the test inputtaskda-code
Applies when
task -- the task asks for per-row predictions on a provided test file to be written to an output file with a specified column name.
Pattern
The attempt produces an output file whose row count or row order does not correspond 1:1 to the rows of the given test file (e.g. rows dropped by dropna/filtering/deduplication, predictions made on a train/validation split or an aggregated subset, index reset after shuffling, or a header/index column written incorrectly), and no check is made that len(predictions) == len(test) or that order is preserved. No script is retained to make this auditable.
Detection procedure
  1. From the task, note the required output file, the exact required column name(s), and read the test input to get its row count.
  2. In the scripts, trace the object that is finally written: check whether any dropna, boolean mask, drop_duplicates, groupby, sample, train_test_split, or re-sorting touched the test frame between load and prediction, and whether predictions are written with index=False and the exact column header.
  3. Count the data rows in the submitted file (excluding header) and compare to the test row count; confirm the first/last predictions correspond to the first/last test rows (e.g. by spot-checking feature-driven expectations or a preserved key column).
  4. If no script exists, treat the absence of an explicit row-count/order assertion as unverified and reject.
Discriminator
A real violation is a mismatch in row count/order or a wrong header/extra index column between the submission and the test input; it is not a violation merely because values are unrounded floats or predictions look imperfect, as long as one prediction per test row appears in the original order under the requested column name.
Consequence
The grader's row-wise comparison against the reference cannot align, so the file is scored WRONG/MISSING regardless of model quality (0/1 checks passed).
id 787c6a4f496b · mined from da-code dacode-ml-regression-015@s7
raw text (what the judge reads)
### Prediction file shape/alignment not verified against the test input
- **Applies when**: `task` -- the task asks for per-row predictions on a provided test file to be written to an output file with a specified column name.
- **Pattern**: The attempt produces an output file whose row count or row order does not correspond 1:1 to the rows of the given test file (e.g. rows dropped by `dropna`/filtering/deduplication, predictions made on a train/validation split or an aggregated subset, index reset after shuffling, or a header/index column written incorrectly), and no check is made that `len(predictions) == len(test)` or that order is preserved. No script is retained to make this auditable.
- **Detection procedure**:
  1. From the task, note the required output file, the exact required column name(s), and read the test input to get its row count.
  2. In the scripts, trace the object that is finally written: check whether any `dropna`, boolean mask, `drop_duplicates`, `groupby`, `sample`, `train_test_split`, or re-sorting touched the test frame between load and prediction, and whether predictions are written with `index=False` and the exact column header.
  3. Count the data rows in the submitted file (excluding header) and compare to the test row count; confirm the first/last predictions correspond to the first/last test rows (e.g. by spot-checking feature-driven expectations or a preserved key column).
  4. If no script exists, treat the absence of an explicit row-count/order assertion as unverified and reject.
- **Discriminator**: A real violation is a mismatch in row count/order or a wrong header/extra index column between the submission and the test input; it is *not* a violation merely because values are unrounded floats or predictions look imperfect, as long as one prediction per test row appears in the original order under the requested column name.
- **Consequence**: The grader's row-wise comparison against the reference cannot align, so the file is scored WRONG/MISSING regardless of model quality (0/1 checks passed).
412Unjustified row exclusion (ad‑hoc "outlier"/NA filtering) or partial data scope when computing a simple statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single descriptive statistic over "all observations" of a field, and the scripts load the data and apply any filtering, sampling, or per-file/per-chunk aggregation before computing it.
Pattern
The attempt narrows the population before computing the statistic — dropping "outliers" via an arbitrary rule (z-score, IQR, percentile clip) on a field whose values are codes/bounded categories, reading only one file/sheet/chunk of a multi-part source, coercing invalid strings to NaN and silently dropping many rows, or averaging per-group means instead of pooling all rows — and reports the resulting number without ever comparing it to the unfiltered value or documenting how many rows were removed.
Detection procedure
  1. From the task statement, fix the intended population (all non-missing observations of the field) and note that instructions like "ignore outliers" do not license discarding legitimate in-range values.
  2. In the scripts, trace every step between load and aggregation: file/sheet selection, row filters, dtype coercion, dedup, groupby, sampling; check that the final statistic is computed on the pooled full column, not a subset or a mean-of-means.
  3. Require the script to print row counts before and after each drop, plus the statistic computed both with and without any optional filtering; if these diagnostics (or the scripts themselves) are absent, the number is unverifiable and the attempt is inadequate.
  4. Sanity-check the reported value against the field's plausible range/known distribution of codes; a shift from the unfiltered mean caused by discarding a nontrivial fraction of valid rows is a red flag.
Discriminator
Dropping only genuinely missing/unparseable entries, or values outside the field's documented valid domain, is fine and should leave counts essentially intact; a violation is statistical outlier trimming or scope truncation that removes valid observations (or ignores part of the source) and measurably moves the statistic, without justification or a reported unfiltered comparison.
Consequence
The reported mean is biased low/high relative to the true population mean and fails exact/tolerance matching by the grader, even though the answer format is correct.
id 68c64fa9ee65 · mined from infiagent-dabench dabench-320@s7
raw text (what the judge reads)
### Unjustified row exclusion (ad‑hoc "outlier"/NA filtering) or partial data scope when computing a simple statistic
- **Applies when**: `task` -- the task asks for a single descriptive statistic over "all observations" of a field, and the scripts load the data and apply any filtering, sampling, or per-file/per-chunk aggregation before computing it.
- **Pattern**: The attempt narrows the population before computing the statistic — dropping "outliers" via an arbitrary rule (z-score, IQR, percentile clip) on a field whose values are codes/bounded categories, reading only one file/sheet/chunk of a multi-part source, coercing invalid strings to NaN and silently dropping many rows, or averaging per-group means instead of pooling all rows — and reports the resulting number without ever comparing it to the unfiltered value or documenting how many rows were removed.
- **Detection procedure**:
  1. From the task statement, fix the intended population (all non-missing observations of the field) and note that instructions like "ignore outliers" do not license discarding legitimate in-range values.
  2. In the scripts, trace every step between load and aggregation: file/sheet selection, row filters, dtype coercion, dedup, groupby, sampling; check that the final statistic is computed on the pooled full column, not a subset or a mean-of-means.
  3. Require the script to print row counts before and after each drop, plus the statistic computed both with and without any optional filtering; if these diagnostics (or the scripts themselves) are absent, the number is unverifiable and the attempt is inadequate.
  4. Sanity-check the reported value against the field's plausible range/known distribution of codes; a shift from the unfiltered mean caused by discarding a nontrivial fraction of valid rows is a red flag.
- **Discriminator**: Dropping only genuinely missing/unparseable entries, or values outside the field's documented valid domain, is fine and should leave counts essentially intact; a violation is statistical outlier trimming or scope truncation that removes valid observations (or ignores part of the source) and measurably moves the statistic, without justification or a reported unfiltered comparison.
- **Consequence**: The reported mean is biased low/high relative to the true population mean and fails exact/tolerance matching by the grader, even though the answer format is correct.
413Required output artifact is never written to the expected path/name/formattaskda-code
Applies when
task -- the task specifies a deliverable (a named result file and/or an exact JSON/CSV schema) that the grader will read after the scripts run.
Pattern
The script computes the answer and prints it or dumps it to an ad-hoc location/filename/extension of the agent's choosing (e.g., a scratch text file in the home directory) instead of creating the required artifact with the required name, path, and structure, so the graded file is missing even if the reasoning is fine.
Detection procedure
  1. From the task statement, list the exact deliverable(s): file name, directory (default to the working/output directory the harness uses), serialization format, and the literal key names/value types of the requested template.
  2. Grep the final script for every file-write call and compare each target path/name/extension against that list; also check that the object being serialized uses the exact keys and value types (e.g., list vs. scalar) from the template.
  3. Confirm nothing is only printed or written to a non-graded scratch file; if multiple scripts exist, verify the last one run produces the artifact.
  4. If the artifact is absent or misnamed/mis-keyed, flag the attempt regardless of whether the printed values look right.
Discriminator
A real violation is a write to a different filename/path/format than requested, or keys/value shapes that differ from the given template; it is not a violation if the file is written to the specified name in a plausible working directory, or if extra diagnostic prints/files exist alongside the correct artifact.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even though the console output may contain the intended answer.
id 3ee620e42342 · mined from da-code dacode-di-text-001@s7
raw text (what the judge reads)
### Required output artifact is never written to the expected path/name/format
- **Applies when**: `task` -- the task specifies a deliverable (a named result file and/or an exact JSON/CSV schema) that the grader will read after the scripts run.
- **Pattern**: The script computes the answer and prints it or dumps it to an ad-hoc location/filename/extension of the agent's choosing (e.g., a scratch text file in the home directory) instead of creating the required artifact with the required name, path, and structure, so the graded file is missing even if the reasoning is fine.
- **Detection procedure**:
  1. From the task statement, list the exact deliverable(s): file name, directory (default to the working/output directory the harness uses), serialization format, and the literal key names/value types of the requested template.
  2. Grep the final script for every file-write call and compare each target path/name/extension against that list; also check that the object being serialized uses the exact keys and value types (e.g., list vs. scalar) from the template.
  3. Confirm nothing is only `print`ed or written to a non-graded scratch file; if multiple scripts exist, verify the *last* one run produces the artifact.
  4. If the artifact is absent or misnamed/mis-keyed, flag the attempt regardless of whether the printed values look right.
- **Discriminator**: A real violation is a write to a different filename/path/format than requested, or keys/value shapes that differ from the given template; it is *not* a violation if the file is written to the specified name in a plausible working directory, or if extra diagnostic prints/files exist alongside the correct artifact.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even though the console output may contain the intended answer.
414Output file omits the requested component/intermediate columns (over-reduced deliverable)taskda-code
Applies when
task -- the task asks to compute several derived quantities (per-entity metrics, sub-scores, a combined score, a segment/label) and to save "the results, including X and Y" to a specified file.
Pattern
The script computes all the intermediate quantities in memory, prints them for inspection, but then writes only the final label plus an ID key to the output file, discarding the per-entity metrics, per-component scores, and combined score/segment string that the task text explicitly enumerated. Related sub-pattern: an ambiguous component definition (e.g., counting raw rows instead of distinct events) is chosen without justification and, because the underlying columns are not saved, the choice is invisible and uncheckable in the deliverable.
Detection procedure
  1. Read the task sentence describing the deliverable and list every noun that must appear in the saved file (each metric, each score, the segmentation, the level/label).
  2. Find the line in the script that writes the output file and enumerate the exact columns selected there.
  3. Compare the two lists; flag if any requested item was computed but excluded, or if column names/values are collapsed (e.g., only a single label column survives).
  4. Check the answer text: does it describe the file schema as fewer columns than the task enumerated, and does it justify definitional choices for the components that are not persisted?
Discriminator
A real violation is dropping quantities the task named as part of the saved result. It is not a violation to omit genuinely internal scratch columns the task never mentioned, nor to add extra columns beyond those requested (supersets are usually acceptable); also fine if the task explicitly restricts the file to a minimal schema.
Consequence
The grader compares the saved file against an expected table containing the metric, score, and segment columns; missing columns cause a schema/content mismatch and the file is scored WRONG regardless of whether the labels themselves are reasonable.
id 7887fdeee9d1 · mined from da-code dacode-dm-csv-052@s7
raw text (what the judge reads)
### Output file omits the requested component/intermediate columns (over-reduced deliverable)
- **Applies when**: `task` -- the task asks to compute several derived quantities (per-entity metrics, sub-scores, a combined score, a segment/label) and to save "the results, including X and Y" to a specified file.
- **Pattern**: The script computes all the intermediate quantities in memory, prints them for inspection, but then writes only the final label plus an ID key to the output file, discarding the per-entity metrics, per-component scores, and combined score/segment string that the task text explicitly enumerated. Related sub-pattern: an ambiguous component definition (e.g., counting raw rows instead of distinct events) is chosen without justification and, because the underlying columns are not saved, the choice is invisible and uncheckable in the deliverable.
- **Detection procedure**:
  1. Read the task sentence describing the deliverable and list every noun that must appear in the saved file (each metric, each score, the segmentation, the level/label).
  2. Find the line in the script that writes the output file and enumerate the exact columns selected there.
  3. Compare the two lists; flag if any requested item was computed but excluded, or if column names/values are collapsed (e.g., only a single label column survives).
  4. Check the answer text: does it describe the file schema as fewer columns than the task enumerated, and does it justify definitional choices for the components that are not persisted?
- **Discriminator**: A real violation is dropping quantities the task named as part of the saved result. It is *not* a violation to omit genuinely internal scratch columns the task never mentioned, nor to add extra columns beyond those requested (supersets are usually acceptable); also fine if the task explicitly restricts the file to a minimal schema.
- **Consequence**: The grader compares the saved file against an expected table containing the metric, score, and segment columns; missing columns cause a schema/content mismatch and the file is scored WRONG regardless of whether the labels themselves are reasonable.
415Unverified input subset/format contract (assumed instead of inspected)taskda-code
Applies when
task -- the task names specific source files, a group/subset to analyze, and/or a template file that defines the output layout, and the script loads data and writes results without ever printing/inspecting those artifacts.
Pattern
The agent hard-codes file paths, column names, and an output schema from guesswork (e.g., comments like "based on the sample format, it looks like…"), applies no filter for the required subgroup, and never checks that the number of loaded records, the value ranges, or the written columns/headers match what the task and template imply. A small, obviously truncated or unfiltered table is analyzed as if it were the full target population.
Detection procedure
  1. From the task, list the required inputs (which files, which subgroup/filter) and the required output contract (which file, column names, order, rounding/format) and note any template file that must be read.
  2. In the scripts, check whether the template/sample file is actually opened and its header reused, or whether the schema and value formatting are invented; check whether the stated subgroup filter is applied after loading.
  3. In the scripts/answer, look for any validation of the loaded data: row counts per group, expected ranges, no missing values, shape after filtering.
  4. Compare the reported sample sizes and summary values in the answer against what the described source realistically contains; flag round/suspiciously tiny counts (e.g., exactly 20 rows per group) or values far from plausible domain magnitudes with no comment.
Discriminator
Fine if the script (or logged output) shows the template header being read/echoed and prints group-wise counts/ranges that the agent explicitly reconciles with the task description; a violation is when both the input subset and output schema rest on unverified assumptions and no count/shape/range check appears anywhere.
Consequence
The numbers are computed from the wrong rows (or written under the wrong header/precision), so the result file mismatches the expected one and the grader marks the required output WRONG/MISSING even though the statistical method itself was implemented correctly.
id 811afefae9b0 · mined from da-code dacode-data-sa-029@s7
raw text (what the judge reads)
### Unverified input subset/format contract (assumed instead of inspected)
- **Applies when**: `task` -- the task names specific source files, a group/subset to analyze, and/or a template file that defines the output layout, and the script loads data and writes results without ever printing/inspecting those artifacts.
- **Pattern**: The agent hard-codes file paths, column names, and an output schema from guesswork (e.g., comments like "based on the sample format, it looks like…"), applies no filter for the required subgroup, and never checks that the number of loaded records, the value ranges, or the written columns/headers match what the task and template imply. A small, obviously truncated or unfiltered table is analyzed as if it were the full target population.
- **Detection procedure**:
  1. From the task, list the required inputs (which files, which subgroup/filter) and the required output contract (which file, column names, order, rounding/format) and note any template file that must be read.
  2. In the scripts, check whether the template/sample file is actually opened and its header reused, or whether the schema and value formatting are invented; check whether the stated subgroup filter is applied after loading.
  3. In the scripts/answer, look for any validation of the loaded data: row counts per group, expected ranges, no missing values, shape after filtering.
  4. Compare the reported sample sizes and summary values in the answer against what the described source realistically contains; flag round/suspiciously tiny counts (e.g., exactly 20 rows per group) or values far from plausible domain magnitudes with no comment.
- **Discriminator**: Fine if the script (or logged output) shows the template header being read/echoed and prints group-wise counts/ranges that the agent explicitly reconciles with the task description; a violation is when both the input subset and output schema rest on unverified assumptions and no count/shape/range check appears anywhere.
- **Consequence**: The numbers are computed from the wrong rows (or written under the wrong header/precision), so the result file mismatches the expected one and the grader marks the required output WRONG/MISSING even though the statistical method itself was implemented correctly.
416Degenerate/outlier-driven clustering accepted on the strength of an internal scoretaskda-code
Applies when
task -- the task asks for an unsupervised segmentation into "an appropriate number of groups" and the script picks k by an internal index (silhouette/inertia) after building aggregate features from raw transactional records.
Pattern
The attempt skips domain cleaning of the raw records (cancellations/negative or zero amounts, rows with no entity key, duplicate or out-of-range values) and applies only a mean/variance scaler to heavily skewed, long-tailed aggregates; the resulting k-selection is driven by a handful of extreme points, producing clusters with a few members each while one cluster holds most of the population, and the high silhouette is reported as evidence of success without any distributional sanity check.
Detection procedure
  1. In the task/README, list the record-level conditions that must be filtered or repaired before aggregating (invalid/negative/cancelled records, missing entity identifiers, date-range restriction) and check the script performs each explicitly.
  2. In the script, check whether skewed aggregate features are transformed (log/rank/robust scaling) or outliers handled before distance-based clustering, and whether the k-search compares candidates on more than one criterion.
  3. In the reported output, inspect the cluster size distribution and per-cluster feature summaries: flag if any cluster contains a negligible fraction of entities (e.g., <1%) or if one cluster absorbs the bulk while the score is claimed to be "high".
  4. Verify the output file's row count and column set match the requested schema and the cleaned entity count implied by step 1.
Discriminator
A genuine violation shows tiny clusters that are pure extreme-value artifacts and an absent cleaning/transformation step, so the segmentation is unusable and the entity count is inflated by invalid records; it is not a violation when a small cluster is a deliberately identified, documented high-value segment produced after cleaning and skew handling, with the remaining clusters reasonably balanced and the choice of k justified beyond a single index.
Consequence
The saved result file fails validation — entity/row counts and feature values differ from those obtainable on properly filtered data, and the label distribution is degenerate, so the grader marks the expected output file as wrong.
id d6671bacbbdb · mined from da-code dacode-ml-cluster-019@s7
raw text (what the judge reads)
### Degenerate/outlier-driven clustering accepted on the strength of an internal score
- **Applies when**: `task` -- the task asks for an unsupervised segmentation into "an appropriate number of groups" and the script picks k by an internal index (silhouette/inertia) after building aggregate features from raw transactional records.
- **Pattern**: The attempt skips domain cleaning of the raw records (cancellations/negative or zero amounts, rows with no entity key, duplicate or out-of-range values) and applies only a mean/variance scaler to heavily skewed, long-tailed aggregates; the resulting k-selection is driven by a handful of extreme points, producing clusters with a few members each while one cluster holds most of the population, and the high silhouette is reported as evidence of success without any distributional sanity check.
- **Detection procedure**:
  1. In the task/README, list the record-level conditions that must be filtered or repaired before aggregating (invalid/negative/cancelled records, missing entity identifiers, date-range restriction) and check the script performs each explicitly.
  2. In the script, check whether skewed aggregate features are transformed (log/rank/robust scaling) or outliers handled before distance-based clustering, and whether the k-search compares candidates on more than one criterion.
  3. In the reported output, inspect the cluster size distribution and per-cluster feature summaries: flag if any cluster contains a negligible fraction of entities (e.g., <1%) or if one cluster absorbs the bulk while the score is claimed to be "high".
  4. Verify the output file's row count and column set match the requested schema and the cleaned entity count implied by step 1.
- **Discriminator**: A genuine violation shows tiny clusters that are pure extreme-value artifacts *and* an absent cleaning/transformation step, so the segmentation is unusable and the entity count is inflated by invalid records; it is *not* a violation when a small cluster is a deliberately identified, documented high-value segment produced after cleaning and skew handling, with the remaining clusters reasonably balanced and the choice of k justified beyond a single index.
- **Consequence**: The saved result file fails validation — entity/row counts and feature values differ from those obtainable on properly filtered data, and the label distribution is degenerate, so the grader marks the expected output file as wrong.
417Fabricating input data instead of loading the provided datasettaskda-code
Applies when
task -- the task references a supplied dataset/README, and the analysis scripts must derive numbers from those files.
Pattern
The scripts never read any data file; instead they hard-code "historically typical" or guessed counts/aggregates (often with comments trying several alternative guesses), then run the requested statistical procedure on the invented numbers.
Detection procedure
1. From the task, note that input data files are provided and identify what raw records the statistic requires. 2. Search the scripts for any file-loading call (read_csv, read_excel, load, open, glob of the data directory); confirm whether the analysis variables come from a file or from literal assignments. 3. Check whether comments/prints reveal uncertainty about the true values ("let me try", "typical version", multiple candidate numbers) and whether the agent ever listed/inspected the data directory. 4. Compare the reported summary counts/rates against any figures stated in the task/README for plausibility.
Discriminator
A real violation is when the key inputs are literals invented or recalled from memory with no verification against the provided files; it is fine to hard-code constants that are explicitly given in the task/README, or to hard-code values already extracted from the loaded data in a prior verified step.
Consequence
The confidence interval / metric is computed on the wrong sample sizes and rates, so the written output file fails the expected-value check (here the interval even spans zero, contradicting the well-established effect), and the graded result is marked wrong.
id 31e7014a99ec · mined from da-code dacode-data-sa-031@s7
raw text (what the judge reads)
### Fabricating input data instead of loading the provided dataset
- **Applies when**: `task` -- the task references a supplied dataset/README, and the analysis scripts must derive numbers from those files.
- **Pattern**: The scripts never read any data file; instead they hard-code "historically typical" or guessed counts/aggregates (often with comments trying several alternative guesses), then run the requested statistical procedure on the invented numbers.
- **Detection procedure**: 1. From the task, note that input data files are provided and identify what raw records the statistic requires. 2. Search the scripts for any file-loading call (`read_csv`, `read_excel`, `load`, `open`, glob of the data directory); confirm whether the analysis variables come from a file or from literal assignments. 3. Check whether comments/prints reveal uncertainty about the true values ("let me try", "typical version", multiple candidate numbers) and whether the agent ever listed/inspected the data directory. 4. Compare the reported summary counts/rates against any figures stated in the task/README for plausibility.
- **Discriminator**: A real violation is when the key inputs are literals invented or recalled from memory with no verification against the provided files; it is fine to hard-code constants that are explicitly given in the task/README, or to hard-code values already extracted from the loaded data in a prior verified step.
- **Consequence**: The confidence interval / metric is computed on the wrong sample sizes and rates, so the written output file fails the expected-value check (here the interval even spans zero, contradicting the well-established effect), and the graded result is marked wrong.
418Group-wise statistic computed on a silently truncated or mis-selected subsettaskinfiagent-dabench
Applies when
task -- the task asks for summary statistics (median, range, min/max, counts) of a value column computed separately for each level of a categorical/filter column.
Pattern
The attempt filters and groups in one unverified step (e.g. dropping rows with NA/odd dtypes, keeping only rows that match some incidental condition, joining/deduplicating first, or reading the statistic off an already-aggregated or partially-loaded frame) and reports the numbers without ever printing how many rows landed in each group or what the raw min/max are. The result looks plausible in isolation but is derived from a smaller or different population than the task defines.
Detection procedure
  1. Read the task and note exactly which rows should enter each group (which filter values, which value column, whether all rows of the file are in scope) and which output statistics are requested.
  2. Read the script and trace the row-count path from load → filter → group → aggregate: look for implicit reductions (dropna, drop_duplicates, merges, head/sample, chunked or nrows-limited reads, query on extra conditions, resampling/daily aggregation) that are not mandated by the task.
  3. Check whether the script prints per-group row counts, min, max, and dtype of the value column before reporting; absence of any such sanity output is itself the flag.
  4. Cross-check the reported spread against the data's plausible extremes: if the reported range/extremes are markedly narrower than the physically or empirically expected span of the raw column, treat the aggregation as suspect.
Discriminator
A real violation is a reduction in scope that the task never asked for, or an unverified pipeline where group sizes and extremes were never inspected. It is not a violation if the reduction is explicitly required by the task's constraints (e.g. the stated category filter, stated unit conversion) and the script prints counts/extremes per group showing the full in-scope rows were retained.
Consequence
The requested statistics (especially spread-type ones like range, which are dominated by the tails that get dropped) come out systematically too small or shifted, so the graded values mismatch the reference even though the answer format is correct.
id 41c74aa43ded · mined from infiagent-dabench dabench-759@s7
raw text (what the judge reads)
### Group-wise statistic computed on a silently truncated or mis-selected subset
- **Applies when**: `task` -- the task asks for summary statistics (median, range, min/max, counts) of a value column computed separately for each level of a categorical/filter column.
- **Pattern**: The attempt filters and groups in one unverified step (e.g. dropping rows with NA/odd dtypes, keeping only rows that match some incidental condition, joining/deduplicating first, or reading the statistic off an already-aggregated or partially-loaded frame) and reports the numbers without ever printing how many rows landed in each group or what the raw min/max are. The result looks plausible in isolation but is derived from a smaller or different population than the task defines.
- **Detection procedure**:
  1. Read the task and note exactly which rows should enter each group (which filter values, which value column, whether all rows of the file are in scope) and which output statistics are requested.
  2. Read the script and trace the row-count path from load → filter → group → aggregate: look for implicit reductions (dropna, drop_duplicates, merges, head/sample, chunked or nrows-limited reads, `query` on extra conditions, resampling/daily aggregation) that are not mandated by the task.
  3. Check whether the script prints per-group row counts, min, max, and dtype of the value column before reporting; absence of any such sanity output is itself the flag.
  4. Cross-check the reported spread against the data's plausible extremes: if the reported range/extremes are markedly narrower than the physically or empirically expected span of the raw column, treat the aggregation as suspect.
- **Discriminator**: A real violation is a reduction in scope that the task never asked for, or an unverified pipeline where group sizes and extremes were never inspected. It is *not* a violation if the reduction is explicitly required by the task's constraints (e.g. the stated category filter, stated unit conversion) and the script prints counts/extremes per group showing the full in-scope rows were retained.
- **Consequence**: The requested statistics (especially spread-type ones like range, which are dominated by the tails that get dropped) come out systematically too small or shifted, so the graded values mismatch the reference even though the answer format is correct.
419Applying a hypothesis test to the full raw table without scoping the population or checking the test's distributional assumptionstaskda-code
Applies when
task -- the task asks for a p-value and an accept/reject decision comparing two groups, and the scripts run a single off-the-shelf test on every row of the loaded files.
Pattern
The script loads both files, derives the comparison quantity, and immediately calls a parametric two-sample test on all rows, without (a) restricting to the subset/time window/competition tier that the question's framing implies, and (b) verifying that the variable satisfies the chosen test's assumptions (normality/symmetry, independence, equal variance) or comparing against the appropriate nonparametric alternative.
Detection procedure
  1. Read the task and README for scoping cues (a stated period, event type, tier, or "comparable matches/records" framing) and check whether the script applies any row filter before testing; a bare read_csv → aggregate → test chain with no filtering step is a red flag.
  2. Check whether the script inspects the distribution of the test variable (histogram, skew, discreteness/counts) or justifies the test family; if it jumps straight to ttest_ind on a bounded count-like variable, note the missing assumption check.
  3. Check whether the script considers one-sided vs two-sided and the direction implied by the stated null/alternative, and whether variance equality (equal_var) was decided rather than left at default.
  4. Compare the reported p-value's magnitude to what a scoped, assumption-appropriate test would plausibly give: an astronomically small p-value (e.g. 1e-100 or smaller) from tens of thousands of rows signals that the test was run on the whole population rather than the intended subset.
Discriminator
A real violation is when the script never filters to the implied subset or never justifies the test family, so a different (defensible) analysis choice would change the p-value by many orders of magnitude — and possibly the decision. It is not a violation if the script explicitly documents that the full table is the intended population and shows an assumption check (or reports that parametric and nonparametric tests agree), with the decision robust to the choice.
Consequence
The saved p-value differs from the reference value by orders of magnitude (and the decision string may still coincidentally match), so an exact/tolerance check on the p-value column fails and the result file is graded wrong.
id cdc06af280a0 · mined from da-code dacode-data-sa-001@s8
raw text (what the judge reads)
### Applying a hypothesis test to the full raw table without scoping the population or checking the test's distributional assumptions
- **Applies when**: `task` -- the task asks for a p-value and an accept/reject decision comparing two groups, and the scripts run a single off-the-shelf test on every row of the loaded files.
- **Pattern**: The script loads both files, derives the comparison quantity, and immediately calls a parametric two-sample test on all rows, without (a) restricting to the subset/time window/competition tier that the question's framing implies, and (b) verifying that the variable satisfies the chosen test's assumptions (normality/symmetry, independence, equal variance) or comparing against the appropriate nonparametric alternative.
- **Detection procedure**:
  1. Read the task and README for scoping cues (a stated period, event type, tier, or "comparable matches/records" framing) and check whether the script applies any row filter before testing; a bare `read_csv` → aggregate → test chain with no filtering step is a red flag.
  2. Check whether the script inspects the distribution of the test variable (histogram, skew, discreteness/counts) or justifies the test family; if it jumps straight to `ttest_ind` on a bounded count-like variable, note the missing assumption check.
  3. Check whether the script considers one-sided vs two-sided and the direction implied by the stated null/alternative, and whether variance equality (`equal_var`) was decided rather than left at default.
  4. Compare the reported p-value's magnitude to what a scoped, assumption-appropriate test would plausibly give: an astronomically small p-value (e.g. 1e-100 or smaller) from tens of thousands of rows signals that the test was run on the whole population rather than the intended subset.
- **Discriminator**: A real violation is when the script never filters to the implied subset or never justifies the test family, so a different (defensible) analysis choice would change the p-value by many orders of magnitude — and possibly the decision. It is *not* a violation if the script explicitly documents that the full table is the intended population and shows an assumption check (or reports that parametric and nonparametric tests agree), with the decision robust to the choice.
- **Consequence**: The saved p-value differs from the reference value by orders of magnitude (and the decision string may still coincidentally match), so an exact/tolerance check on the p-value column fails and the result file is graded wrong.
420Output schema/format assumed instead of validated against the provided templatetaskda-code
Applies when
task -- the task says to write results to a file "following the exact structure/formatting" of a provided sample/template file, and the scripts build the output frame themselves.
Pattern
The script hard-codes column names, column order, row order, rounding and numeric formatting from memory or from a quick glance, and only afterwards "verifies" by printing both files side by side without an automated, field-by-field conformance check; row ordering may even be copied blindly from placeholder rows in the template rather than derived from a defensible rule (e.g., the sort key implied by the sample).
Detection procedure
  1. In the task text, note every explicit formatting constraint (target filename, column set/order, header names, sort order, rounding/units, index inclusion).
  2. In the scripts, check whether the template file is actually read and used to drive the output construction, or whether names/order/rounding are literals invented by the agent (e.g., result.columns = [...], result[[...]], .round(2) with no reference to the sample).
  3. Look for an explicit assertion comparing the written file's header list, column order, dtypes, row count, and row-key order to the template — not just print() of both.
  4. Check whether the ordering rule was inferred (sorted by a value/name) or copied from the template's example rows, and whether any grouping keys present in the template (e.g., extra breakdown levels or a total row) are missing from the output.
Discriminator
A real violation is when the output's header names, column/row order, rounding, or set of rows could differ from the template and nothing in the code would catch it; it is fine if the script reads the template and programmatically enforces (or asserts) identical headers, shape, key set, and ordering rule before finishing.
Consequence
The grader's exact file comparison fails ("result.csv WRONG/MISSING") even when the underlying aggregation numbers may be right, since headers, column/row order, or rounding do not match the expected file.
id bbc9f574cd34 · mined from da-code dacode-dm-csv-011@s8
raw text (what the judge reads)
### Output schema/format assumed instead of validated against the provided template
- **Applies when**: `task` -- the task says to write results to a file "following the exact structure/formatting" of a provided sample/template file, and the scripts build the output frame themselves.
- **Pattern**: The script hard-codes column names, column order, row order, rounding and numeric formatting from memory or from a quick glance, and only afterwards "verifies" by printing both files side by side without an automated, field-by-field conformance check; row ordering may even be copied blindly from placeholder rows in the template rather than derived from a defensible rule (e.g., the sort key implied by the sample).
- **Detection procedure**:
  1. In the task text, note every explicit formatting constraint (target filename, column set/order, header names, sort order, rounding/units, index inclusion).
  2. In the scripts, check whether the template file is actually read and used to drive the output construction, or whether names/order/rounding are literals invented by the agent (e.g., `result.columns = [...]`, `result[[...]]`, `.round(2)` with no reference to the sample).
  3. Look for an explicit assertion comparing the written file's header list, column order, dtypes, row count, and row-key order to the template — not just `print()` of both.
  4. Check whether the ordering rule was inferred (sorted by a value/name) or copied from the template's example rows, and whether any grouping keys present in the template (e.g., extra breakdown levels or a total row) are missing from the output.
- **Discriminator**: A real violation is when the output's header names, column/row order, rounding, or set of rows could differ from the template and nothing in the code would catch it; it is fine if the script reads the template and programmatically enforces (or asserts) identical headers, shape, key set, and ordering rule before finishing.
- **Consequence**: The grader's exact file comparison fails ("result.csv WRONG/MISSING") even when the underlying aggregation numbers may be right, since headers, column/row order, or rounding do not match the expected file.
421Unjustified discretization/clipping of continuous predictions, with no check against the provided submission template or evaluation metrictaskda-code
Applies when
task -- the task supplies a sample/template submission file and asks for predictions of a numeric target, and the scripts apply post-processing (rounding, casting to int, clipping, flooring at a bound) before writing the output.
Pattern
The attempt trains a regression model, then transforms the raw continuous predictions with an ad-hoc rule justified only by intuition about the target ("counts must be whole numbers", "must be ≥ 1"), never loads or compares against the provided template to confirm expected dtype/precision/columns/row order, and selects models using a convenience metric (e.g. R²) rather than the metric implied by the task, so the post-processing is never validated as harmless.
Detection procedure
  1. Read the task/README and note the provided template file and any stated (or strongly implied) evaluation metric and output format.
  2. Scan the scripts for any transformation applied to predictions after predict() (rounding, astype(int), np.maximum/clip, scaling) and check whether the task explicitly requires it.
  3. Check whether the template file is ever read and used to validate the written output (column names, row count, id set and ordering, value dtype/range), and whether model choice was scored with the task's metric on held-out data both with and without the post-processing.
  4. Inspect the submitted answer: are all values integers/quantized, and does the id ordering/row count demonstrably match the template rather than just the test file as read?
Discriminator
A real violation is post-processing that is self-invented and never A/B-tested against the raw predictions on validation data, combined with no programmatic comparison to the template; it is fine if the task/template explicitly specifies integer or bounded outputs, or if the script measures the task metric before and after the transform and shows it does not hurt.
Consequence
The submission file is accepted structurally but scores materially worse than the raw-prediction baseline (quantization adds avoidable error), or fails format/dtype/ordering validation outright, so the graded check on the submission file is marked wrong.
id bde53e890931 · mined from da-code dacode-ml-competition-009@s8
raw text (what the judge reads)
### Unjustified discretization/clipping of continuous predictions, with no check against the provided submission template or evaluation metric
- **Applies when**: `task` -- the task supplies a sample/template submission file and asks for predictions of a numeric target, and the scripts apply post-processing (rounding, casting to int, clipping, flooring at a bound) before writing the output.
- **Pattern**: The attempt trains a regression model, then transforms the raw continuous predictions with an ad-hoc rule justified only by intuition about the target ("counts must be whole numbers", "must be ≥ 1"), never loads or compares against the provided template to confirm expected dtype/precision/columns/row order, and selects models using a convenience metric (e.g. R²) rather than the metric implied by the task, so the post-processing is never validated as harmless.
- **Detection procedure**:
  1. Read the task/README and note the provided template file and any stated (or strongly implied) evaluation metric and output format.
  2. Scan the scripts for any transformation applied to predictions after `predict()` (rounding, `astype(int)`, `np.maximum/clip`, scaling) and check whether the task explicitly requires it.
  3. Check whether the template file is ever read and used to validate the written output (column names, row count, id set and ordering, value dtype/range), and whether model choice was scored with the task's metric on held-out data both with and without the post-processing.
  4. Inspect the submitted answer: are all values integers/quantized, and does the id ordering/row count demonstrably match the template rather than just the test file as read?
- **Discriminator**: A real violation is post-processing that is self-invented and never A/B-tested against the raw predictions on validation data, combined with no programmatic comparison to the template; it is fine if the task/template explicitly specifies integer or bounded outputs, or if the script measures the task metric before and after the transform and shows it does not hurt.
- **Consequence**: The submission file is accepted structurally but scores materially worse than the raw-prediction baseline (quantization adds avoidable error), or fails format/dtype/ordering validation outright, so the graded check on the submission file is marked wrong.
422Silent row/feature attrition when the task implies full-coverage outputtaskda-code
Applies when
task -- the task asks for a per-record output file (labels, predictions, transformed features) over an input table that contains missing values or mixed dtypes.
Pattern
The script drops every record with any missing value and/or keeps only an arbitrary subset of columns (often just the columns that happened to be complete, including ID-like or non-informative numeric codes), then writes an output whose row count and feature-vector width silently differ from the full input; the answer reports the reduction as a normal processing step rather than checking it against the task's implied coverage.
Detection procedure
  1. Read the task: note whether it asks for output covering the whole dataset and whether it specifies which features/columns to use; note the required column naming/format.
  2. In the script, find the preprocessing step: check whether rows are dropped (dropna, filtering) and how the feature set is chosen; check whether imputation/encoding was considered instead of deletion, and whether identifier-like numeric columns were excluded.
  3. Compare the reported output row count and number of feature columns against the input record count and the number of usable columns; flag any unexplained shrinkage.
  4. Verify the answer states a concrete sanity check (rows == input rows, feature width, header names exactly as requested) rather than just "completed successfully".
Discriminator
A real violation is deletion/selection done for convenience that changes the output's coverage or feature width without the task authorizing it; it is fine if the task explicitly permits filtering, or if the dropped rows/columns are justified (e.g., a column is entirely missing or non-numeric text) and the output still covers the required records with imputation applied where needed.
Consequence
The written file has fewer rows and/or a different Feature_i width than the reference, so row-wise or shape-based comparison fails and the file is scored WRONG/MISSING even though the clustering code itself ran.
id f025c810b058 · mined from da-code dacode-ml-cluster-009@s8
raw text (what the judge reads)
### Silent row/feature attrition when the task implies full-coverage output
- **Applies when**: `task` -- the task asks for a per-record output file (labels, predictions, transformed features) over an input table that contains missing values or mixed dtypes.
- **Pattern**: The script drops every record with any missing value and/or keeps only an arbitrary subset of columns (often just the columns that happened to be complete, including ID-like or non-informative numeric codes), then writes an output whose row count and feature-vector width silently differ from the full input; the answer reports the reduction as a normal processing step rather than checking it against the task's implied coverage.
- **Detection procedure**:
  1. Read the task: note whether it asks for output covering the whole dataset and whether it specifies which features/columns to use; note the required column naming/format.
  2. In the script, find the preprocessing step: check whether rows are dropped (`dropna`, filtering) and how the feature set is chosen; check whether imputation/encoding was considered instead of deletion, and whether identifier-like numeric columns were excluded.
  3. Compare the reported output row count and number of feature columns against the input record count and the number of usable columns; flag any unexplained shrinkage.
  4. Verify the answer states a concrete sanity check (rows == input rows, feature width, header names exactly as requested) rather than just "completed successfully".
- **Discriminator**: A real violation is deletion/selection done for convenience that changes the output's coverage or feature width without the task authorizing it; it is fine if the task explicitly permits filtering, or if the dropped rows/columns are justified (e.g., a column is entirely missing or non-numeric text) *and* the output still covers the required records with imputation applied where needed.
- **Consequence**: The written file has fewer rows and/or a different `Feature_i` width than the reference, so row-wise or shape-based comparison fails and the file is scored WRONG/MISSING even though the clustering code itself ran.
423Dropping records before producing a per-record deliverabletaskda-code
Applies when
task -- the task asks for an output file with one row per entity/observation derived from the full dataset (e.g. labels, predictions, segments), and the script applies filtering/outlier-removal/deduplication steps before writing that file.
Pattern
The attempt silently shrinks the population (missing-value drops, sign filters, IQR/quantile outlier trimming, sampling) and then exports only the surviving rows, so the deliverable's row count and coverage no longer match the entities implied by the task; the answer reports the reduced count as if it were the answer. Often paired with an unvalidated "optimal" choice (e.g. picking the number of groups purely by the maximum of one internal score, yielding a degenerate 2-group split) with no sanity check on coverage or group balance.
Detection procedure
  1. From the task, state what one row of the requested output file must represent and how many rows there should be (all entities in the source data, or the documented cleaned set).
  2. In the scripts, list every operation that removes or subsets rows between loading and writing the output; note the row count after each and whether removed entities are ever re-attached (e.g. assigned a label or re-scored).
  3. Compare the final written row count / index alignment (including any merge, reset_index, or column-assignment across differently filtered frames) to the number from step 1; check the file's column names and count against the requested schema.
  4. Check whether the answer justifies the discarding as required by the task, and whether the group count was sanity-checked (non-trivial number of groups, reasonable sizes) rather than taken blindly from one score.
Discriminator
Removing rows that cannot be labelled at all (e.g. records with no entity identifier) and saying so is fine; the violation is discarding valid, labelable entities for convenience (outliers, extreme values, samples) so the deliverable covers only part of the population, or assigning labels from one filtered frame onto rows of another (misaligned index).
Consequence
The output file has fewer rows than expected and labels that cannot be matched one-to-one with the reference entities, so a file-level comparison of shape/labels fails and the whole check scores 0 even if the clustering itself is sensible.
id 9c125e933dd8 · mined from da-code dacode-ml-cluster-016@s8
raw text (what the judge reads)
### Dropping records before producing a per-record deliverable
- **Applies when**: `task` -- the task asks for an output file with one row per entity/observation derived from the full dataset (e.g. labels, predictions, segments), and the script applies filtering/outlier-removal/deduplication steps before writing that file.
- **Pattern**: The attempt silently shrinks the population (missing-value drops, sign filters, IQR/quantile outlier trimming, sampling) and then exports only the surviving rows, so the deliverable's row count and coverage no longer match the entities implied by the task; the answer reports the reduced count as if it were the answer. Often paired with an unvalidated "optimal" choice (e.g. picking the number of groups purely by the maximum of one internal score, yielding a degenerate 2-group split) with no sanity check on coverage or group balance.
- **Detection procedure**:
  1. From the task, state what one row of the requested output file must represent and how many rows there should be (all entities in the source data, or the documented cleaned set).
  2. In the scripts, list every operation that removes or subsets rows between loading and writing the output; note the row count after each and whether removed entities are ever re-attached (e.g. assigned a label or re-scored).
  3. Compare the final written row count / index alignment (including any `merge`, `reset_index`, or column-assignment across differently filtered frames) to the number from step 1; check the file's column names and count against the requested schema.
  4. Check whether the answer justifies the discarding as required by the task, and whether the group count was sanity-checked (non-trivial number of groups, reasonable sizes) rather than taken blindly from one score.
- **Discriminator**: Removing rows that cannot be labelled at all (e.g. records with no entity identifier) and saying so is fine; the violation is discarding valid, labelable entities for convenience (outliers, extreme values, samples) so the deliverable covers only part of the population, or assigning labels from one filtered frame onto rows of another (misaligned index).
- **Consequence**: The output file has fewer rows than expected and labels that cannot be matched one-to-one with the reference entities, so a file-level comparison of shape/labels fails and the whole check scores 0 even if the clustering itself is sensible.
424No held-out validation before submitting accuracy-graded predictionstaskda-code
Applies when
task -- the task asks for a prediction file whose quality will be judged against unseen ground-truth values, and the script fits one model and immediately writes predictions.
Pattern
The attempt trains a single model on all labelled rows, reports fit quality computed on those same training rows (R², MAE, etc.) as evidence of success, and never estimates out-of-sample error via a hold-out split, cross-validation, or (for time-indexed data) a chronological back-test; consequently obviously weak choices — dropping the time/index column, median-imputing instead of interpolating, no lag/calendar features, no baseline comparison, untuned hyperparameters — go undetected.
Detection procedure
1. Read the task to confirm the deliverable is scored on predictive accuracy against hidden targets. 2. In the script, search for any split of the labelled data (train_test_split, cross_val_score, manual date cutoff) and for any comparison against a trivial baseline (mean, persistence/last-value). 3. Check whether the reported metric is computed on the exact same rows used in fit. 4. Read the answer: if the only quality numbers are in-sample and phrased as "model performance", with no out-of-sample estimate or baseline, and no check that important structure (time ordering, lags, seasonality, categorical encodings) was used, flag it.
Discriminator
A real violation reports only in-sample metrics with no hold-out/CV number anywhere; it is fine if the script validates on a hold-out or CV fold (and refits on all data afterwards) and the answer quotes that out-of-sample estimate, even if the final metric happens to be reported for the full-data refit.
Consequence
In-sample R²/MAE look strong while the submitted predictions miss the hidden targets by far more than the tolerance, so the accuracy check on the result file fails and the reported error figure is an over-optimistic intermediate value rather than the true generalization performance.
id 92241efcf584 · mined from da-code dacode-ml-regression-002@s8
raw text (what the judge reads)
### No held-out validation before submitting accuracy-graded predictions
- **Applies when**: `task` -- the task asks for a prediction file whose quality will be judged against unseen ground-truth values, and the script fits one model and immediately writes predictions.
- **Pattern**: The attempt trains a single model on all labelled rows, reports fit quality computed on those same training rows (R², MAE, etc.) as evidence of success, and never estimates out-of-sample error via a hold-out split, cross-validation, or (for time-indexed data) a chronological back-test; consequently obviously weak choices — dropping the time/index column, median-imputing instead of interpolating, no lag/calendar features, no baseline comparison, untuned hyperparameters — go undetected.
- **Detection procedure**: 1. Read the task to confirm the deliverable is scored on predictive accuracy against hidden targets. 2. In the script, search for any split of the labelled data (`train_test_split`, `cross_val_score`, manual date cutoff) and for any comparison against a trivial baseline (mean, persistence/last-value). 3. Check whether the reported metric is computed on the exact same rows used in `fit`. 4. Read the answer: if the only quality numbers are in-sample and phrased as "model performance", with no out-of-sample estimate or baseline, and no check that important structure (time ordering, lags, seasonality, categorical encodings) was used, flag it.
- **Discriminator**: A real violation reports *only* in-sample metrics with no hold-out/CV number anywhere; it is fine if the script validates on a hold-out or CV fold (and refits on all data afterwards) and the answer quotes that out-of-sample estimate, even if the final metric happens to be reported for the full-data refit.
- **Consequence**: In-sample R²/MAE look strong while the submitted predictions miss the hidden targets by far more than the tolerance, so the accuracy check on the result file fails and the reported error figure is an over-optimistic intermediate value rather than the true generalization performance.
425Plot delivered without verifiable, data-derived series (no reproducible script or exported plot data)taskda-code
Applies when
task -- the task asks for a chart/figure built from a source dataset to a supplied format spec, and grading inspects the underlying plotted values/figure metadata, not just that an image file exists.
Pattern
The agent reports "chart created and saved" but leaves no runnable script and no serialized artifact of what was plotted (series arrays, figure/axes metadata); the x/y values appear hardcoded, aggregated by an unstated rule, or otherwise not traceable to a load-filter-aggregate chain over the actual source file, so the drawn line cannot be checked against the data.
Detection procedure
  1. Read the task and the format spec to list every required output artifact (image plus any data/metadata dumps) and every constrained property (labels, ticks, size, color, ordering, units).
  2. Inspect the submitted scripts: confirm there is code that reads the source file, selects the correct entity/subset, aggregates on the correct axis/time granularity, and passes those computed arrays directly into the plotting call — not literals or values recited in prose.
  3. Check that the script writes each required artifact (e.g., the exported series/array and figure-spec dump), not only the image, and to the expected path.
  4. Cross-check the answer's stated axis ranges, tick labels, and point counts against what the data pipeline could actually produce (row counts, date span, aggregation level); flag any number stated without a computing line of code.
Discriminator
A real violation is an answer whose plotted values have no code path back to the raw data or whose required non-image outputs are absent; it is not a violation if the script demonstrably computes the series from the source and writes all requested artifacts, even if the prose summary is terse or the styling differs cosmetically within allowed tolerance.
Consequence
Graders that compare the exported series/figure metadata to reference values report the expected data files as WRONG/MISSING and score 0, even though an image was produced.
id 8bf4351a2b07 · mined from da-code dacode-plot-line-015@s8
raw text (what the judge reads)
### Plot delivered without verifiable, data-derived series (no reproducible script or exported plot data)
- **Applies when**: `task` -- the task asks for a chart/figure built from a source dataset to a supplied format spec, and grading inspects the underlying plotted values/figure metadata, not just that an image file exists.
- **Pattern**: The agent reports "chart created and saved" but leaves no runnable script and no serialized artifact of what was plotted (series arrays, figure/axes metadata); the x/y values appear hardcoded, aggregated by an unstated rule, or otherwise not traceable to a load-filter-aggregate chain over the actual source file, so the drawn line cannot be checked against the data.
- **Detection procedure**:
  1. Read the task and the format spec to list every required output artifact (image plus any data/metadata dumps) and every constrained property (labels, ticks, size, color, ordering, units).
  2. Inspect the submitted scripts: confirm there is code that reads the source file, selects the correct entity/subset, aggregates on the correct axis/time granularity, and passes those computed arrays directly into the plotting call — not literals or values recited in prose.
  3. Check that the script writes each required artifact (e.g., the exported series/array and figure-spec dump), not only the image, and to the expected path.
  4. Cross-check the answer's stated axis ranges, tick labels, and point counts against what the data pipeline could actually produce (row counts, date span, aggregation level); flag any number stated without a computing line of code.
- **Discriminator**: A real violation is an answer whose plotted values have no code path back to the raw data or whose required non-image outputs are absent; it is *not* a violation if the script demonstrably computes the series from the source and writes all requested artifacts, even if the prose summary is terse or the styling differs cosmetically within allowed tolerance.
- **Consequence**: Graders that compare the exported series/figure metadata to reference values report the expected data files as WRONG/MISSING and score 0, even though an image was produced.
426Fabricated/synthetic input data instead of the provided dataset, and ignored output schemataskda-code
Applies when
task -- the task references provided input files (data directory, tips.md, sample_result.csv) and requires results written in a specified output format.
Pattern
The scripts never load any supplied file; they hard-code invented "representative" arrays from background knowledge (often round, evenly spaced, suspiciously symmetric values), and they write an output file with self-invented columns instead of copying the schema of the provided sample/template.
Detection procedure
  1. Read the task for named input artifacts (data files, instruction files, sample output files) and note them.
  2. Search the scripts for any read of those artifacts (read_csv, open, np.loadtxt, glob of the data dir); if the only data is literal arrays typed into the script, or comments say "no data file is provided / based on historical study", flag immediately.
  3. Check whether the script reads or reproduces the sample output's exact column names/row order/precision; invented column sets or extra metadata columns are a violation.
  4. Check the answer for signs of unverified inputs: equal group sizes, tidy round numbers, or statistics that were never cross-checked against the real file.
Discriminator
Fine if the script actually loads the provided files (and optionally hard-codes small constants stated in the instructions); a violation if the analytic inputs themselves are invented, or if the output columns/format were chosen by the agent rather than taken from the given template.
Consequence
The computed statistic is unrelated to the true data and the output file fails schema/value comparison, so the graded result file is marked WRONG/MISSING regardless of the correctness of the statistical method.
id 1f5a4c19520d · mined from da-code dacode-data-sa-028@s8
raw text (what the judge reads)
### Fabricated/synthetic input data instead of the provided dataset, and ignored output schema
- **Applies when**: `task` -- the task references provided input files (data directory, `tips.md`, `sample_result.csv`) and requires results written in a specified output format.
- **Pattern**: The scripts never load any supplied file; they hard-code invented "representative" arrays from background knowledge (often round, evenly spaced, suspiciously symmetric values), and they write an output file with self-invented columns instead of copying the schema of the provided sample/template.
- **Detection procedure**:
  1. Read the task for named input artifacts (data files, instruction files, sample output files) and note them.
  2. Search the scripts for any read of those artifacts (`read_csv`, `open`, `np.loadtxt`, glob of the data dir); if the only data is literal arrays typed into the script, or comments say "no data file is provided / based on historical study", flag immediately.
  3. Check whether the script reads or reproduces the sample output's exact column names/row order/precision; invented column sets or extra metadata columns are a violation.
  4. Check the answer for signs of unverified inputs: equal group sizes, tidy round numbers, or statistics that were never cross-checked against the real file.
- **Discriminator**: Fine if the script actually loads the provided files (and optionally hard-codes small constants stated in the instructions); a violation if the analytic inputs themselves are invented, or if the output columns/format were chosen by the agent rather than taken from the given template.
- **Consequence**: The computed statistic is unrelated to the true data and the output file fails schema/value comparison, so the graded result file is marked WRONG/MISSING regardless of the correctness of the statistical method.
427Missing/unverified output artifacts required by referenced spec filestaskda-code
Applies when
task -- the task points to auxiliary instruction/config files (e.g., a tips/spec/config file) and expects specific output files to be written to disk.
Pattern
The agent reads the spec loosely, produces only the one artifact named in the prose of the prompt (e.g., the image), and never emits the other artifacts the spec mandates (serialized plot/config dump, numeric array of the plotted values, etc.); it then declares success based on a printed textual summary rather than on the existence and content of every required file. Often the scripts themselves aren't saved, so the pipeline can't be re-checked.
Detection procedure
  1. Read the task prompt and every referenced instruction/config file; list every output artifact, filename, key name, dtype/shape, unit, rounding, ordering, and filtering rule they mention.
  2. Read the scripts and mark, for each listed artifact/rule, the exact line that writes or enforces it; flag any item with no corresponding line (especially non-image side artifacts and aggregation/weighting or exclusion rules).
  3. Check the answer: does it assert completion by describing content ("title, colors, DPI") instead of confirming each required file was written and re-loaded/validated (shape, length, value range)?
  4. Flag if any mandated artifact is absent, unwritten, or unvalidated, or if the spec's parameters were paraphrased instead of parsed from the file.
Discriminator
A real violation is a mandated artifact/rule that appears nowhere in the scripts, or is written but never re-read and sanity-checked. A look-alike that is fine: the agent writes all mandated files (possibly with extra ones) and shows a verification step (file listing, reload, shape/range print) even if its prose summary is terse.
Consequence
Graders that check each expected file independently report the missing ones as WRONG/MISSING and the run scores zero on those checks even if the single produced plot looks plausible.
id 4b56af2b4b34 · mined from da-code dacode-plot-line-006@s8
raw text (what the judge reads)
### Missing/unverified output artifacts required by referenced spec files
- **Applies when**: `task` -- the task points to auxiliary instruction/config files (e.g., a tips/spec/config file) and expects specific output files to be written to disk.
- **Pattern**: The agent reads the spec loosely, produces only the one artifact named in the prose of the prompt (e.g., the image), and never emits the other artifacts the spec mandates (serialized plot/config dump, numeric array of the plotted values, etc.); it then declares success based on a printed textual summary rather than on the existence and content of every required file. Often the scripts themselves aren't saved, so the pipeline can't be re-checked.
- **Detection procedure**:
  1. Read the task prompt *and* every referenced instruction/config file; list every output artifact, filename, key name, dtype/shape, unit, rounding, ordering, and filtering rule they mention.
  2. Read the scripts and mark, for each listed artifact/rule, the exact line that writes or enforces it; flag any item with no corresponding line (especially non-image side artifacts and aggregation/weighting or exclusion rules).
  3. Check the answer: does it assert completion by describing content ("title, colors, DPI") instead of confirming each required file was written and re-loaded/validated (shape, length, value range)?
  4. Flag if any mandated artifact is absent, unwritten, or unvalidated, or if the spec's parameters were paraphrased instead of parsed from the file.
- **Discriminator**: A real violation is a mandated artifact/rule that appears nowhere in the scripts, or is written but never re-read and sanity-checked. A look-alike that is fine: the agent writes all mandated files (possibly with extra ones) and shows a verification step (file listing, reload, shape/range print) even if its prose summary is terse.
- **Consequence**: Graders that check each expected file independently report the missing ones as WRONG/MISSING and the run scores zero on those checks even if the single produced plot looks plausible.
428Spec file referenced by the task is never actually read and appliedtaskda-code
Applies when
task -- the instructions point to an auxiliary specification (a README/markdown/config defining bins, categories, filters, or naming) that must govern how the analysis is grouped or formatted.
Pattern
The agent skips or only skims the referenced spec and instead uses the categories/labels that already exist natively in the raw data (or its own reasonable-looking scheme), producing output whose group boundaries, labels, ordering, or set of categories differ from the mandated ones — while asserting in the answer that the spec was "followed".
Detection procedure
  1. From the task, list the external spec artifact and what it is supposed to control (bin edges, labels, order, extra outputs).
  2. In the scripts/transcript, look for an explicit read of that artifact and a mapping/binning step derived from its contents; absence of any such read, or a hard-coded scheme with no citation of the spec, is a red flag.
  3. Compare the categories in the reported output against the raw data's distinct values: if they are identical to the raw column's native categories (no merging/re-binning at all), the spec was almost certainly not applied.
  4. Confirm every artifact the task or spec requires (plot file plus any auxiliary data/serialization files) is actually written, with the required title/axis labels and category ordering.
Discriminator
A genuine violation is when no evidence exists that the spec's contents were loaded/quoted and the grouping coincides with the untransformed raw values or an invented scheme; it is fine if the agent quotes the spec and the mandated grouping legitimately happens to match the raw categories, or re-bins exactly as the spec states.
Consequence
The saved figure and any derived arrays/JSON contain the wrong number of bins, wrong labels, or wrong counts per bin (and required side artifacts may be missing), so every file-level equality check against the expected outputs fails despite a confident "task completed" report.
id 2239ca2e858d · mined from da-code dacode-plot-bar-005@s8
raw text (what the judge reads)
### Spec file referenced by the task is never actually read and applied
- **Applies when**: `task` -- the instructions point to an auxiliary specification (a README/markdown/config defining bins, categories, filters, or naming) that must govern how the analysis is grouped or formatted.
- **Pattern**: The agent skips or only skims the referenced spec and instead uses the categories/labels that already exist natively in the raw data (or its own reasonable-looking scheme), producing output whose group boundaries, labels, ordering, or set of categories differ from the mandated ones — while asserting in the answer that the spec was "followed".
- **Detection procedure**:
  1. From the task, list the external spec artifact and what it is supposed to control (bin edges, labels, order, extra outputs).
  2. In the scripts/transcript, look for an explicit read of that artifact and a mapping/binning step derived from its contents; absence of any such read, or a hard-coded scheme with no citation of the spec, is a red flag.
  3. Compare the categories in the reported output against the raw data's distinct values: if they are identical to the raw column's native categories (no merging/re-binning at all), the spec was almost certainly not applied.
  4. Confirm every artifact the task or spec requires (plot file plus any auxiliary data/serialization files) is actually written, with the required title/axis labels and category ordering.
- **Discriminator**: A genuine violation is when no evidence exists that the spec's contents were loaded/quoted and the grouping coincides with the untransformed raw values or an invented scheme; it is fine if the agent quotes the spec and the mandated grouping legitimately happens to match the raw categories, or re-bins exactly as the spec states.
- **Consequence**: The saved figure and any derived arrays/JSON contain the wrong number of bins, wrong labels, or wrong counts per bin (and required side artifacts may be missing), so every file-level equality check against the expected outputs fails despite a confident "task completed" report.
429Required output artifact and literal answer format never producedtaskda-code
Applies when
task -- The task specifies an exact answer template (e.g., a JSON object with keyed lists) and/or expects results to be saved to a named result file, and the scripts only print values to stdout.
Pattern
The attempt does all the computation interactively/printing to console, never writes the requested file and never constructs the answer object in the exact shape requested (scalars pasted where lists/keys were shown, missing keys, altered key names, unrounded or reformatted values), so even a numerically right computation cannot be graded.
Detection procedure
  1. Read the task and list every output requirement: destination file name(s), key names, value types/containers, rounding/units, ordering.
  2. Scan every script for a write/serialize call (to_json, json.dump, to_csv, open(...,'w')) targeting the required path; note if only print statements exist.
  3. Compare the submitted answer's structure key-by-key against the template in the task (container type, key spelling, number formatting).
  4. Flag if any required artifact is absent or any structural/format element differs from the template.
Discriminator
A real violation is a missing file or structurally different answer (scalar instead of list, renamed/missing key, no persisted output). It is not a violation if the file is written with the exact keys/containers and the answer text merely renders the same content readably (e.g., whitespace or key order differences that the format did not pin down).
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying statistic was computed correctly.
id 17db6411888b · mined from da-code dacode-di-text-002@s8
raw text (what the judge reads)
### Required output artifact and literal answer format never produced
- **Applies when**: `task` -- The task specifies an exact answer template (e.g., a JSON object with keyed lists) and/or expects results to be saved to a named result file, and the scripts only print values to stdout.
- **Pattern**: The attempt does all the computation interactively/printing to console, never writes the requested file and never constructs the answer object in the exact shape requested (scalars pasted where lists/keys were shown, missing keys, altered key names, unrounded or reformatted values), so even a numerically right computation cannot be graded.
- **Detection procedure**:
  1. Read the task and list every output requirement: destination file name(s), key names, value types/containers, rounding/units, ordering.
  2. Scan every script for a write/serialize call (`to_json`, `json.dump`, `to_csv`, `open(...,'w')`) targeting the required path; note if only `print` statements exist.
  3. Compare the submitted answer's structure key-by-key against the template in the task (container type, key spelling, number formatting).
  4. Flag if any required artifact is absent or any structural/format element differs from the template.
- **Discriminator**: A real violation is a missing file or structurally different answer (scalar instead of list, renamed/missing key, no persisted output). It is *not* a violation if the file is written with the exact keys/containers and the answer text merely renders the same content readably (e.g., whitespace or key order differences that the format did not pin down).
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying statistic was computed correctly.
430Using only one data file/split when the task refers to the whole datasettaskinfiagent-dabench
Applies when
task -- the task asks for a descriptive statistic (or any global summary) of "the dataset" and the data directory may contain several files or splits (train/test/parts) rather than a single obvious table.
Pattern
The script hard-codes one file (e.g. the training split or the first file found) without listing the directory or justifying the choice, so the statistic is computed on a proper subset of the intended population; the number is plausible-looking, so no alarm is raised.
Detection procedure
  1. Read the task statement and note whether it scopes the computation to a named split/subset or just says "the dataset".
  2. Read the script's load step: does it enumerate the available files (glob/listdir/print of file names, row counts) or does it silently open one hard-coded path?
  3. Check whether the row count actually used is reconciled against the total records available (e.g. printed shapes of all candidate files, or a comment showing the other files are irrelevant/duplicated).
  4. If only a subset was loaded with no such reconciliation, flag it — the reported statistic describes a subset, not the requested population.
Discriminator
A real violation is silently ignoring sibling data files that belong to the same population; it is not a violation if the task itself names the split, if the script demonstrably checked and the other files hold different variables/labels only, or if the ignored file is verified to be a subset/duplicate of the loaded one.
Consequence
The formatted answer parses and the qualitative label may still be right by luck, but the numeric value is computed on the wrong sample and fails an exact-value check (e.g. reported 0.66 vs expected 0.83).
id 9bef79b4e885 · mined from infiagent-dabench dabench-359@s8
raw text (what the judge reads)
### Using only one data file/split when the task refers to the whole dataset
- **Applies when**: `task` -- the task asks for a descriptive statistic (or any global summary) of "the dataset" and the data directory may contain several files or splits (train/test/parts) rather than a single obvious table.
- **Pattern**: The script hard-codes one file (e.g. the training split or the first file found) without listing the directory or justifying the choice, so the statistic is computed on a proper subset of the intended population; the number is plausible-looking, so no alarm is raised.
- **Detection procedure**:
  1. Read the task statement and note whether it scopes the computation to a named split/subset or just says "the dataset".
  2. Read the script's load step: does it enumerate the available files (glob/listdir/print of file names, row counts) or does it silently open one hard-coded path?
  3. Check whether the row count actually used is reconciled against the total records available (e.g. printed shapes of all candidate files, or a comment showing the other files are irrelevant/duplicated).
  4. If only a subset was loaded with no such reconciliation, flag it — the reported statistic describes a subset, not the requested population.
- **Discriminator**: A real violation is silently ignoring sibling data files that belong to the same population; it is *not* a violation if the task itself names the split, if the script demonstrably checked and the other files hold different variables/labels only, or if the ignored file is verified to be a subset/duplicate of the loaded one.
- **Consequence**: The formatted answer parses and the qualitative label may still be right by luck, but the numeric value is computed on the wrong sample and fails an exact-value check (e.g. reported 0.66 vs expected 0.83).
431Assumed output schema instead of confirming the "required format"taskda-code
Applies when
task -- the prompt says results must be saved to a specific file "in the required format" (or otherwise implies a fixed schema) and the scripts write a CSV/JSON whose columns, ordering, units, or row set were chosen by the agent.
Pattern
The attempt never searches the working directory / README / provided templates or existing sample outputs for the expected schema; it invents column names, invents the units (fraction vs. percent), keeps or drops the index/date column arbitrarily, and includes every input row without checking whether the raw inputs need cleaning (e.g. missing or non-numeric values that silently propagate through cumulative/aggregate operations). Verification then only re-reads the agent's own file and prints "looks reasonable" statistics, which cannot detect a schema or convention mismatch.
Detection procedure
  1. In the task text, list every stated output constraint: filename, column names/order, units, rounding, row coverage, index inclusion.
  2. In the scripts, check whether any step reads a spec/template/example of the output (or inspects the raw file's structure and missingness) before constructing the result; note any column name or unit that appears only in the agent's code and nowhere in the task.
  3. Check the verification script: does it compare the produced file against an external reference (spec, template, independently recomputed values, expected row/column counts) or only summarize the file it just wrote?
  4. In the answer, look for statements of format ("Columns: ...", "N rows") that are asserted rather than derived from the task or data.
Discriminator
A real violation is when the required schema/convention is externally defined (or discoverable in the environment) and the script's choice is an unchecked guess; it is fine if the task itself spells out the exact columns/units and the script matches them literally, or if the agent explicitly inspected a template/README and mirrored it.
Consequence
The file exists and its numbers may be internally consistent, but the automated comparison against the reference output fails on column names/order, units/scale, or row count, so the file is scored WRONG/MISSING and the whole task fails despite a plausible-sounding narrative.
id 9b074805588b · mined from da-code dacode-dm-csv-050@s8
raw text (what the judge reads)
### Assumed output schema instead of confirming the "required format"
- **Applies when**: `task` -- the prompt says results must be saved to a specific file "in the required format" (or otherwise implies a fixed schema) and the scripts write a CSV/JSON whose columns, ordering, units, or row set were chosen by the agent.
- **Pattern**: The attempt never searches the working directory / README / provided templates or existing sample outputs for the expected schema; it invents column names, invents the units (fraction vs. percent), keeps or drops the index/date column arbitrarily, and includes every input row without checking whether the raw inputs need cleaning (e.g. missing or non-numeric values that silently propagate through cumulative/aggregate operations). Verification then only re-reads the agent's own file and prints "looks reasonable" statistics, which cannot detect a schema or convention mismatch.
- **Detection procedure**:
  1. In the task text, list every stated output constraint: filename, column names/order, units, rounding, row coverage, index inclusion.
  2. In the scripts, check whether any step reads a spec/template/example of the output (or inspects the raw file's structure and missingness) before constructing the result; note any column name or unit that appears only in the agent's code and nowhere in the task.
  3. Check the verification script: does it compare the produced file against an external reference (spec, template, independently recomputed values, expected row/column counts) or only summarize the file it just wrote?
  4. In the answer, look for statements of format ("Columns: ...", "N rows") that are asserted rather than derived from the task or data.
- **Discriminator**: A real violation is when the required schema/convention is externally defined (or discoverable in the environment) and the script's choice is an unchecked guess; it is fine if the task itself spells out the exact columns/units and the script matches them literally, or if the agent explicitly inspected a template/README and mirrored it.
- **Consequence**: The file exists and its numbers may be internally consistent, but the automated comparison against the reference output fails on column names/order, units/scale, or row count, so the file is scored WRONG/MISSING and the whole task fails despite a plausible-sounding narrative.
432Silently dropping eligible numeric columns from an "all variables" comparisontaskinfiagent-dabench
Applies when
task -- the task says to compute a statistic between a target column and all other numerical columns and report the extremum, and the script selects which columns to include.
Pattern
The attempt filters the candidate set by intuition rather than by the stated rule — excluding index-like, id-like, rank-like, or otherwise "trivially related" numeric columns, or only keeping a hand-listed set of "interesting" features — so the true extremum is never evaluated, and a lower-magnitude substantive variable is reported.
Detection procedure
  1. From the task, state the exact eligibility rule for candidate columns (e.g., every numeric column other than the target) and the exact selection rule (max |statistic|, sign reported separately).
  2. In the script, find where candidates are chosen (select_dtypes, explicit column lists, drop(...)) and check whether any numeric column is excluded without the task authorizing it.
  3. Check that the printed output enumerates the statistic for every candidate column, not just the winner, and that the reported answer is the row with the largest absolute value (including negatives), not the largest signed value.
  4. Confirm the reported sign comes from that same winning column's coefficient.
Discriminator
A real violation is excluding a numeric column that meets the task's stated criterion (even if it seems redundant or derived from the target). It is fine to exclude non-numeric/text columns, or the target itself, or columns the task explicitly says to filter out.
Consequence
The extremum is taken over a truncated candidate set, so both the reported variable and often its sign are wrong — the grader marks all checks failed.
id 4797f6229b74 · mined from infiagent-dabench dabench-117@s8
raw text (what the judge reads)
### Silently dropping eligible numeric columns from an "all variables" comparison
- **Applies when**: `task` -- the task says to compute a statistic between a target column and *all* other numerical columns and report the extremum, and the script selects which columns to include.
- **Pattern**: The attempt filters the candidate set by intuition rather than by the stated rule — excluding index-like, id-like, rank-like, or otherwise "trivially related" numeric columns, or only keeping a hand-listed set of "interesting" features — so the true extremum is never evaluated, and a lower-magnitude substantive variable is reported.
- **Detection procedure**:
  1. From the task, state the exact eligibility rule for candidate columns (e.g., every numeric column other than the target) and the exact selection rule (max |statistic|, sign reported separately).
  2. In the script, find where candidates are chosen (`select_dtypes`, explicit column lists, `drop(...)`) and check whether any numeric column is excluded without the task authorizing it.
  3. Check that the printed output enumerates the statistic for *every* candidate column, not just the winner, and that the reported answer is the row with the largest absolute value (including negatives), not the largest signed value.
  4. Confirm the reported sign comes from that same winning column's coefficient.
- **Discriminator**: A real violation is excluding a numeric column that meets the task's stated criterion (even if it seems redundant or derived from the target). It is fine to exclude non-numeric/text columns, or the target itself, or columns the task explicitly says to filter out.
- **Consequence**: The extremum is taken over a truncated candidate set, so both the reported variable and often its sign are wrong — the grader marks all checks failed.
433Unvetted input subset and unreconciled diagnostics for a distribution/normality testtaskinfiagent-dabench
Applies when
task -- the task asks for a statistical test plus descriptive statistics on a single column, and the script feeds the raw column straight into the test after at most a dropna().
Pattern
The attempt never inspects the column's actual contents (sentinel/placeholder values such as 0 or -1 rows, boundary/incomplete records, non-numeric or mixed dtypes, duplicated index rows) and never checks the test's validity conditions (e.g. sample-size limits where the p-value is only approximate). It then reports a verdict driven by a handful of contaminating rows, and does not notice that the reported skewness/kurtosis are wildly inconsistent with the expected shape of the variable.
Detection procedure
  1. Read the task to see which rows are in scope, then read the script for any row filtering: if the only cleaning is dropna() (or nothing), the attempt has assumed the column is clean.
  2. Check whether the script prints any distributional sanity output before the test — value counts of extreme/zero values, min/max vs. plausible range, n after cleaning, dtype — and whether the sample size is inside the test's reliable range.
  3. Compare the reported skewness/kurtosis with the reported test verdict and with the printed min/mean/median: heavy skew and large excess kurtosis on a bounded count-like variable is a red flag that a small set of degenerate rows dominates, not evidence that the analysis is correct.
  4. If no filtering rationale is documented and no sensitivity check (test/statistics recomputed after excluding the suspect rows) exists, flag the attempt.
Discriminator
A fine attempt either shows evidence that the column has no sentinel/degenerate values (printed range, counts, dtype consistent with the variable's meaning) or explicitly justifies keeping/excluding them and shows the conclusion is stable; a violation silently reports a single unchecked run whose conclusion flips if a few obviously anomalous rows are removed.
Consequence
The test statistic and p-value are computed on a contaminated (or invalid-size) sample, so the binary normality verdict is inverted relative to ground truth and the skewness/kurtosis values are far off, failing every graded field.
id 3c6a17545f7e · mined from infiagent-dabench dabench-298@s8
raw text (what the judge reads)
### Unvetted input subset and unreconciled diagnostics for a distribution/normality test
- **Applies when**: `task` -- the task asks for a statistical test plus descriptive statistics on a single column, and the script feeds the raw column straight into the test after at most a `dropna()`.
- **Pattern**: The attempt never inspects the column's actual contents (sentinel/placeholder values such as 0 or -1 rows, boundary/incomplete records, non-numeric or mixed dtypes, duplicated index rows) and never checks the test's validity conditions (e.g. sample-size limits where the p-value is only approximate). It then reports a verdict driven by a handful of contaminating rows, and does not notice that the reported skewness/kurtosis are wildly inconsistent with the expected shape of the variable.
- **Detection procedure**:
  1. Read the task to see which rows are in scope, then read the script for any row filtering: if the only cleaning is `dropna()` (or nothing), the attempt has assumed the column is clean.
  2. Check whether the script prints any distributional sanity output *before* the test — value counts of extreme/zero values, min/max vs. plausible range, n after cleaning, dtype — and whether the sample size is inside the test's reliable range.
  3. Compare the reported skewness/kurtosis with the reported test verdict and with the printed min/mean/median: heavy skew and large excess kurtosis on a bounded count-like variable is a red flag that a small set of degenerate rows dominates, not evidence that the analysis is correct.
  4. If no filtering rationale is documented and no sensitivity check (test/statistics recomputed after excluding the suspect rows) exists, flag the attempt.
- **Discriminator**: A fine attempt either shows evidence that the column has no sentinel/degenerate values (printed range, counts, dtype consistent with the variable's meaning) or explicitly justifies keeping/excluding them and shows the conclusion is stable; a violation silently reports a single unchecked run whose conclusion flips if a few obviously anomalous rows are removed.
- **Consequence**: The test statistic and p-value are computed on a contaminated (or invalid-size) sample, so the binary normality verdict is inverted relative to ground truth and the skewness/kurtosis values are far off, failing every graded field.
434Weak final model chosen from a hobbled/unrepresentative validation comparisontaskda-code
Applies when
task -- the task is a predictive-modeling competition scored on prediction quality, and the scripts compare a few candidate models before writing the submission file.
Pattern
The agent progressively downgrades its evaluation setup for speed (small random subsample, single contiguous/non-shuffled split instead of cross-validation, drastically reduced tree counts/depth) and then picks whichever candidate "wins" that crippled comparison — typically a plain linear baseline — discarding stronger candidates evaluated earlier under different conditions. No feature engineering, no hyperparameter tuning, and no benchmark of the achieved validation score against a known-good or expected level; the answer reports the weak score as if it were a success.
Detection procedure
  1. Read the task to see that the deliverable is scored on predictive accuracy, not just on producing a file.
  2. In the scripts, check how the final model was selected: is the validation set a full held-out split or CV over all data, or a small subsample split by index order? Are the candidate models compared on the same data with comparable capacity/budget?
  3. Check whether scores from different scripts/settings are being mixed, and whether the final chosen model is the simplest baseline while stronger models were trained only under handicapped settings (few estimators, shallow depth, tiny sample) or dropped entirely.
  4. Check the answer: does it justify the choice only by an internal validation number, with no comparison to a reference/baseline expectation, no tuning, and no error analysis?
Discriminator
A real violation is when the comparison itself is unfair or unrepresentative (different data sizes/splits per model, capacity-starved competitors, order-based splits) or when the accepted score is never sanity-checked against any external reference. It is not a violation if a simple model is chosen after a like-for-like comparison on a proper held-out set or CV using full-strength configurations, and the reported score is argued to be near the achievable ceiling.
Consequence
The submission has the right shape and column names but predictions are materially less accurate than the grading threshold, so the file check on the expected result is marked WRONG despite a confident-sounding report.
id 9d652acaa5ac · mined from da-code dacode-ml-competition-008@s8
raw text (what the judge reads)
### Weak final model chosen from a hobbled/unrepresentative validation comparison
- **Applies when**: `task` -- the task is a predictive-modeling competition scored on prediction quality, and the scripts compare a few candidate models before writing the submission file.
- **Pattern**: The agent progressively downgrades its evaluation setup for speed (small random subsample, single contiguous/non-shuffled split instead of cross-validation, drastically reduced tree counts/depth) and then picks whichever candidate "wins" that crippled comparison — typically a plain linear baseline — discarding stronger candidates evaluated earlier under different conditions. No feature engineering, no hyperparameter tuning, and no benchmark of the achieved validation score against a known-good or expected level; the answer reports the weak score as if it were a success.
- **Detection procedure**:
  1. Read the task to see that the deliverable is scored on predictive accuracy, not just on producing a file.
  2. In the scripts, check how the final model was selected: is the validation set a full held-out split or CV over all data, or a small subsample split by index order? Are the candidate models compared on the *same* data with comparable capacity/budget?
  3. Check whether scores from different scripts/settings are being mixed, and whether the final chosen model is the simplest baseline while stronger models were trained only under handicapped settings (few estimators, shallow depth, tiny sample) or dropped entirely.
  4. Check the answer: does it justify the choice only by an internal validation number, with no comparison to a reference/baseline expectation, no tuning, and no error analysis?
- **Discriminator**: A real violation is when the comparison itself is unfair or unrepresentative (different data sizes/splits per model, capacity-starved competitors, order-based splits) or when the accepted score is never sanity-checked against any external reference. It is *not* a violation if a simple model is chosen after a like-for-like comparison on a proper held-out set or CV using full-strength configurations, and the reported score is argued to be near the achievable ceiling.
- **Consequence**: The submission has the right shape and column names but predictions are materially less accurate than the grading threshold, so the file check on the expected result is marked WRONG despite a confident-sounding report.
435Ambiguous group boundaries assigned without checking the value distribution or alternative conventionstaskinfiagent-dabench
Applies when
task -- the task asks for a statistic computed separately over ranges of a numeric variable described in words ("below X", "between X and Y", "above Y"), and the script hard-codes comparison operators to build those subsets.
Pattern
The agent picks one inclusive/exclusive convention (e.g., putting boundary values in the lower bin and treating "above Y" as >= Y), never inspects how many records sit exactly on the boundaries, and never re-runs the statistic under the alternative convention. Because a large mass of the data typically sits exactly on round boundary values, the wrong assignment moves most of a bin's records and changes the reported statistic substantially, while the bin that has no boundary mass looks correct and creates false confidence.
Detection procedure
  1. Read the task wording and list every boundary value and the bins it could belong to; note that "between X and Y" is usually inclusive of both endpoints and "above Y" strictly greater.
  2. In the script, read the filter expressions and record the exact operators used for each bin; check whether the union is exhaustive/disjoint and where each boundary value lands.
  3. Check whether the script prints the frequency of the grouping variable (or at least bin counts) so a reviewer can see how many rows sit exactly on boundaries; if boundary mass is large and unexamined, flag.
  4. Check whether the script computes the statistic under at least one alternative boundary convention (or otherwise justifies the chosen one); absence of any sensitivity check with heavy boundary mass = inadequate.
Discriminator
Not a violation if the grouping variable is continuous with effectively no records exactly at the boundaries, or if the script explicitly shows counts/values at the boundaries and the results are unchanged (or the wording is unambiguous, e.g., "strictly less than"). It is a violation when boundary values are common (discrete/rounded ratings-like variables) and only one convention was ever computed.
Consequence
The bin containing no boundary mass matches ground truth while the bins that gain/lose the boundary records report materially different coefficients, so the answer fails most of the per-bin checks despite correct metric code.
id 14a0bd052b2b · mined from infiagent-dabench dabench-513@s8
raw text (what the judge reads)
### Ambiguous group boundaries assigned without checking the value distribution or alternative conventions
- **Applies when**: `task` -- the task asks for a statistic computed separately over ranges of a numeric variable described in words ("below X", "between X and Y", "above Y"), and the script hard-codes comparison operators to build those subsets.
- **Pattern**: The agent picks one inclusive/exclusive convention (e.g., putting boundary values in the lower bin and treating "above Y" as `>= Y`), never inspects how many records sit exactly on the boundaries, and never re-runs the statistic under the alternative convention. Because a large mass of the data typically sits exactly on round boundary values, the wrong assignment moves most of a bin's records and changes the reported statistic substantially, while the bin that has no boundary mass looks correct and creates false confidence.
- **Detection procedure**:
  1. Read the task wording and list every boundary value and the bins it could belong to; note that "between X and Y" is usually inclusive of both endpoints and "above Y" strictly greater.
  2. In the script, read the filter expressions and record the exact operators used for each bin; check whether the union is exhaustive/disjoint and where each boundary value lands.
  3. Check whether the script prints the frequency of the grouping variable (or at least bin counts) so a reviewer can see how many rows sit exactly on boundaries; if boundary mass is large and unexamined, flag.
  4. Check whether the script computes the statistic under at least one alternative boundary convention (or otherwise justifies the chosen one); absence of any sensitivity check with heavy boundary mass = inadequate.
- **Discriminator**: Not a violation if the grouping variable is continuous with effectively no records exactly at the boundaries, or if the script explicitly shows counts/values at the boundaries and the results are unchanged (or the wording is unambiguous, e.g., "strictly less than"). It is a violation when boundary values are common (discrete/rounded ratings-like variables) and only one convention was ever computed.
- **Consequence**: The bin containing no boundary mass matches ground truth while the bins that gain/lose the boundary records report materially different coefficients, so the answer fails most of the per-bin checks despite correct metric code.
436Invented qualification thresholds / aggregation rules instead of the ones specified in the task or docstaskda-code
Applies when
task -- the task or its documentation defines the target metric (including any eligibility filter, aggregation level, or output template file) and the script must reproduce that definition to rank/select entities.
Pattern
The agent hard-codes a self-chosen cutoff and aggregation choice (e.g. a round-number minimum count, sum vs mean, same filter reused for every sub-question), never quotes or verifies the definition given in the README/spec, and never opens the provided sample output file to confirm columns, ordering, and whether values or names are expected.
Detection procedure
  1. Read the task text and any accompanying documentation and list every stated definition, eligibility condition, rounding/units rule, and any referenced template/sample output file.
  2. Search the scripts for each of those items; flag any constant, filter, or aggregation function that appears without being traceable to the stated definition (magic numbers, arbitrary thresholds, one filter applied to all metrics).
  3. Check whether the script ever loads/prints the sample/template file and compares its header, row count, column order and value type against the produced output.
  4. Flag if the definition is truncated/ambiguous and the agent resolved it by guessing rather than by testing candidate interpretations against the sample or documenting sensitivity.
Discriminator
A real violation is a parameter or aggregation whose value cannot be derived from the task/docs/sample and materially changes which rows are selected; it is fine if the threshold is explicitly stated in the docs, or if the agent shows that the ranking is unchanged across plausible interpretations, or reproduces the sample's known rows.
Consequence
The ranked entity lists differ from the reference (or the file's schema/ordering differs), so exact-match file comparison fails and the task scores 0.
id 23a69325be09 · mined from da-code dacode-dm-csv-009@s8
raw text (what the judge reads)
### Invented qualification thresholds / aggregation rules instead of the ones specified in the task or docs
- **Applies when**: `task` -- the task or its documentation defines the target metric (including any eligibility filter, aggregation level, or output template file) and the script must reproduce that definition to rank/select entities.
- **Pattern**: The agent hard-codes a self-chosen cutoff and aggregation choice (e.g. a round-number minimum count, `sum` vs `mean`, same filter reused for every sub-question), never quotes or verifies the definition given in the README/spec, and never opens the provided sample output file to confirm columns, ordering, and whether values or names are expected.
- **Detection procedure**:
  1. Read the task text and any accompanying documentation and list every stated definition, eligibility condition, rounding/units rule, and any referenced template/sample output file.
  2. Search the scripts for each of those items; flag any constant, filter, or aggregation function that appears without being traceable to the stated definition (magic numbers, arbitrary thresholds, one filter applied to all metrics).
  3. Check whether the script ever loads/prints the sample/template file and compares its header, row count, column order and value type against the produced output.
  4. Flag if the definition is truncated/ambiguous and the agent resolved it by guessing rather than by testing candidate interpretations against the sample or documenting sensitivity.
- **Discriminator**: A real violation is a parameter or aggregation whose value cannot be derived from the task/docs/sample and materially changes which rows are selected; it is fine if the threshold is explicitly stated in the docs, or if the agent shows that the ranking is unchanged across plausible interpretations, or reproduces the sample's known rows.
- **Consequence**: The ranked entity lists differ from the reference (or the file's schema/ordering differs), so exact-match file comparison fails and the task scores 0.
437Undefined/empty-subset results not reported in the required literal formattaskinfiagent-dabench
Applies when
task -- The task demands a statistic in a strict answer template (e.g. a float rounded to N decimals) and the requested filters can plausibly select zero or all-missing rows, making the statistic undefined.
Pattern
The attempt (often with no saved script) discovers the statistic is undefined and hand-writes a substitute token (NaN, N/A, None, -, null, or an arbitrary 0.00) instead of emitting the sentinel exactly as the grader's string comparison expects, and never documents the row counts that justify the degenerate result.
Detection procedure
  1. Read the task: note the exact answer template, the rounding/format rule, and whether the filter chain could yield an empty or all-missing subset.
  2. Read the scripts: check that they print the row count (and non-null count) after each filtering step and that the printed value is produced by the same formatting code path that fills the answer template (e.g. f"{v:.2f}"), not typed by hand.
  3. Read the answer: confirm the emitted token is exactly what that formatting path would produce for the degenerate case (case, spelling, no extra decorations) and matches the conventional lowercase float repr rather than an ad-hoc label.
  4. If no script or no printed counts exist to reproduce the token, treat the answer as unverified.
Discriminator
A real violation is a substituted or re-cased/re-worded sentinel, or an undefined result asserted without evidence of the empty/all-null subset; it is fine if the agent shows the counts and emits the sentinel exactly as the standard string conversion of the computed value produces it.
Consequence
The grader's exact-match/parse check on the answer field fails (0/1 checks), even though the underlying analysis reached the right conclusion.
id 1aacc0e5c8c7 · mined from infiagent-dabench dabench-554@s8
raw text (what the judge reads)
### Undefined/empty-subset results not reported in the required literal format
- **Applies when**: `task` -- The task demands a statistic in a strict answer template (e.g. a float rounded to N decimals) and the requested filters can plausibly select zero or all-missing rows, making the statistic undefined.
- **Pattern**: The attempt (often with no saved script) discovers the statistic is undefined and hand-writes a substitute token (`NaN`, `N/A`, `None`, `-`, `null`, or an arbitrary `0.00`) instead of emitting the sentinel exactly as the grader's string comparison expects, and never documents the row counts that justify the degenerate result.
- **Detection procedure**:
  1. Read the task: note the exact answer template, the rounding/format rule, and whether the filter chain could yield an empty or all-missing subset.
  2. Read the scripts: check that they print the row count (and non-null count) after each filtering step and that the printed value is produced by the same formatting code path that fills the answer template (e.g. `f"{v:.2f}"`), not typed by hand.
  3. Read the answer: confirm the emitted token is exactly what that formatting path would produce for the degenerate case (case, spelling, no extra decorations) and matches the conventional lowercase float repr rather than an ad-hoc label.
  4. If no script or no printed counts exist to reproduce the token, treat the answer as unverified.
- **Discriminator**: A real violation is a substituted or re-cased/re-worded sentinel, or an undefined result asserted without evidence of the empty/all-null subset; it is fine if the agent shows the counts and emits the sentinel exactly as the standard string conversion of the computed value produces it.
- **Consequence**: The grader's exact-match/parse check on the answer field fails (0/1 checks), even though the underlying analysis reached the right conclusion.
438Answer not traceable to executed code on the provided datataskda-code
Applies when
task -- the task asks for specific values/rankings that must be computed from a supplied dataset with stated preprocessing and ordering rules, and the deliverable is a small result file.
Pattern
The attempt produces a plausible-looking list of entities/numbers that reflects general domain or world knowledge rather than the actual file: no script (or only a stub) loads the data, applies the required preprocessing (e.g., mean fill of the target column), sorts, and writes the result artifact, so the output cannot be regenerated and silently ignores dataset-specific quirks (naming variants, string-formatted numerics, rows present/absent, required sort direction within each group).
Detection procedure
  1. From the task, list the required computational steps and constraints (imputation rule, ranking direction/order within each reported group, output keys, output file path).
  2. In the scripts, locate the exact lines that read the provided file, coerce the relevant column to numeric, impute, sort, slice top-k, and dump the required artifact; check each constraint is implemented, including sort order for both groups.
  3. If no such end-to-end script exists (or it never writes the expected artifact), treat the answer as unverified; if it exists, re-derive the reported items mentally/spot-check against printed intermediate output (row counts, min/max of the column, imputed count).
  4. Check the reported entity labels are spelled exactly as they appear in the dataset, not as commonly known names.
Discriminator
A real violation is when the reported values have no reproducible code path from the input file, or the code exists but skips a stated constraint; a look-alike that is fine is a script that fully computes the result and merely happens to agree with domain intuition, with names and ordering taken verbatim from the data.
Consequence
The graded artifact is missing or contains entities/order that do not match the values computed from the file, so the exact-match check on the result file fails (0/1).
id 478a181eb10f · mined from da-code dacode-di-text-003@s8
raw text (what the judge reads)
### Answer not traceable to executed code on the provided data
- **Applies when**: `task` -- the task asks for specific values/rankings that must be computed from a supplied dataset with stated preprocessing and ordering rules, and the deliverable is a small result file.
- **Pattern**: The attempt produces a plausible-looking list of entities/numbers that reflects general domain or world knowledge rather than the actual file: no script (or only a stub) loads the data, applies the required preprocessing (e.g., mean fill of the target column), sorts, and writes the result artifact, so the output cannot be regenerated and silently ignores dataset-specific quirks (naming variants, string-formatted numerics, rows present/absent, required sort direction within each group).
- **Detection procedure**:
  1. From the task, list the required computational steps and constraints (imputation rule, ranking direction/order within each reported group, output keys, output file path).
  2. In the scripts, locate the exact lines that read the provided file, coerce the relevant column to numeric, impute, sort, slice top-k, and dump the required artifact; check each constraint is implemented, including sort order for *both* groups.
  3. If no such end-to-end script exists (or it never writes the expected artifact), treat the answer as unverified; if it exists, re-derive the reported items mentally/spot-check against printed intermediate output (row counts, min/max of the column, imputed count).
  4. Check the reported entity labels are spelled exactly as they appear in the dataset, not as commonly known names.
- **Discriminator**: A real violation is when the reported values have no reproducible code path from the input file, or the code exists but skips a stated constraint; a look-alike that is fine is a script that fully computes the result and merely happens to agree with domain intuition, with names and ordering taken verbatim from the data.
- **Consequence**: The graded artifact is missing or contains entities/order that do not match the values computed from the file, so the exact-match check on the result file fails (0/1).
439Substituting proxy entities/metrics when the required fields appear absenttaskda-code
Applies when
task -- the task names specific entities, stages, or measures to compute and plot/save, and the scripts cannot immediately find matching columns in the loaded file(s).
Pattern
Instead of searching the full set of provided inputs (or re-reading the config/spec) for the named fields, the agent declares a "data mismatch," redefines the requested entities and metrics as loosely analogous proxies from whatever file it happened to open, and produces a chart/report that answers a different question — often also skipping required output artifacts.
Detection procedure
  1. From the task statement, list the exact entities, grouping keys, metric definition, ranking/filter rule, and every required output file/format (including any settings file that must be honored).
  2. In the scripts, check which input files were enumerated and loaded, and whether the columns used actually correspond to the listed entities and metric; flag any comment or narrative that redefines a requested term as "interpreted as X".
  3. Check the answer/report for the required artifacts and for statements that the metric or grouping was replaced by a substitute (e.g., counts standing in for a value measure, one categorical field standing in for another).
  4. Confirm whether all provided data files were inspected before the substitution was made; unexplored inputs plus a redefinition is a violation.
Discriminator
Legitimate: the agent enumerates all supplied inputs, shows the required fields genuinely do not exist anywhere, and still produces every requested artifact in the requested format under the documented assumption. Violation: the substitution is based on a single partially-explored file, changes the semantics of the grouping key or metric, or omits required outputs.
Consequence
The saved figure and any numeric/serialized outputs encode different groups and quantities than expected (and some required files are missing entirely), so every value/artifact comparison fails — 0 checks passed.
id 40260f6cf9b7 · mined from da-code dacode-plot-scatter-002@s8
raw text (what the judge reads)
### Substituting proxy entities/metrics when the required fields appear absent
- **Applies when**: `task` -- the task names specific entities, stages, or measures to compute and plot/save, and the scripts cannot immediately find matching columns in the loaded file(s).
- **Pattern**: Instead of searching the full set of provided inputs (or re-reading the config/spec) for the named fields, the agent declares a "data mismatch," redefines the requested entities and metrics as loosely analogous proxies from whatever file it happened to open, and produces a chart/report that answers a different question — often also skipping required output artifacts.
- **Detection procedure**:
  1. From the task statement, list the exact entities, grouping keys, metric definition, ranking/filter rule, and every required output file/format (including any settings file that must be honored).
  2. In the scripts, check which input files were enumerated and loaded, and whether the columns used actually correspond to the listed entities and metric; flag any comment or narrative that redefines a requested term as "interpreted as X".
  3. Check the answer/report for the required artifacts and for statements that the metric or grouping was replaced by a substitute (e.g., counts standing in for a value measure, one categorical field standing in for another).
  4. Confirm whether all provided data files were inspected before the substitution was made; unexplored inputs plus a redefinition is a violation.
- **Discriminator**: Legitimate: the agent enumerates all supplied inputs, shows the required fields genuinely do not exist anywhere, and still produces every requested artifact in the requested format under the documented assumption. Violation: the substitution is based on a single partially-explored file, changes the semantics of the grouping key or metric, or omits required outputs.
- **Consequence**: The saved figure and any numeric/serialized outputs encode different groups and quantities than expected (and some required files are missing entirely), so every value/artifact comparison fails — 0 checks passed.
440Stratified preprocessing mistaken for stratified analysis (wrong number/scope of reported statistics)taskda-code
Applies when
task -- the task says to clean/filter the data within subgroups of one variable, then run a single statistical test comparing groups of a different variable.
Pattern
The agent reuses the preprocessing grouping as the analysis grouping, running the test once per preprocessing subgroup and reporting a list of statistics (and one conclusion each), instead of pooling the cleaned data and computing the single requested test across the levels named in the task.
Detection procedure
1) From the task text, identify which variable defines the filtering strata and which variable defines the comparison groups, and count how many test statistics the requested output should contain. 2) In the scripts, check whether the test call loops over the filtering strata or is applied once to the concatenated cleaned data with samples split by the comparison variable. 3) Compare the length of the answer's statistic/conclusion arrays with the count implied by step 1; also verify each conclusion is derived from its own p-value at a stated/standard alpha. 4) Sanity-check that the cleaned rows from all strata are recombined (row counts before/after filtering) rather than analyzed separately or dropped.
Discriminator
A real violation is when the task names exactly one comparison across a fixed set of levels yet the answer contains one result per preprocessing stratum (or vice versa); it is fine if the task explicitly asks for per-stratum tests, or if the output happens to be a list of length 1 wrapped for format compliance.
Consequence
The expected result file contains a single p-value/conclusion pair (or a differently sized set), so the submitted multi-element arrays mismatch on both count and values and the check fails.
id 8f32dcf7eedc · mined from da-code dacode-data-sa-061@s8
raw text (what the judge reads)
### Stratified preprocessing mistaken for stratified analysis (wrong number/scope of reported statistics)
- **Applies when**: `task` -- the task says to clean/filter the data within subgroups of one variable, then run a single statistical test comparing groups of a *different* variable.
- **Pattern**: The agent reuses the preprocessing grouping as the analysis grouping, running the test once per preprocessing subgroup and reporting a list of statistics (and one conclusion each), instead of pooling the cleaned data and computing the single requested test across the levels named in the task.
- **Detection procedure**: 1) From the task text, identify which variable defines the filtering strata and which variable defines the comparison groups, and count how many test statistics the requested output should contain. 2) In the scripts, check whether the test call loops over the filtering strata or is applied once to the concatenated cleaned data with samples split by the comparison variable. 3) Compare the length of the answer's statistic/conclusion arrays with the count implied by step 1; also verify each conclusion is derived from its own p-value at a stated/standard alpha. 4) Sanity-check that the cleaned rows from all strata are recombined (row counts before/after filtering) rather than analyzed separately or dropped.
- **Discriminator**: A real violation is when the task names exactly one comparison across a fixed set of levels yet the answer contains one result per preprocessing stratum (or vice versa); it is fine if the task explicitly asks for per-stratum tests, or if the output happens to be a list of length 1 wrapped for format compliance.
- **Consequence**: The expected result file contains a single p-value/conclusion pair (or a differently sized set), so the submitted multi-element arrays mismatch on both count and values and the check fails.
441Prediction file not verified for row-for-row coverage, ordering, and value-distribution sanitytaskda-code
Applies when
task -- the task asks for a prediction/derived column written to a file for every record of a held-out input set, and the script does its own cleaning (dropping/filtering rows, coercing dtypes, keeping only a subset of columns) before predicting.
Pattern
The agent cleans train and test independently, silently drops or reorders rows (missing values, dtype coercion, deduplication, subsetting to "numeric only" features), writes predictions without asserting that the output length and order match the raw input file, and reports summary stats (e.g., a very narrow predicted range, an odd row count) without checking them against the target's actual distribution or the input's row count.
Detection procedure
  1. From the task, note the required output: file name/location, exact column name(s), and that there must be one prediction per row of the provided held-out file in its original order.
  2. In the scripts, trace the test dataframe from load to write: look for dropna, boolean filtering, merge, groupby, sort_values, index resets, or feature-subsetting that could change row count or order, and check whether any explicit assertion compares len(predictions) to len(raw_test) and whether the file is written to the path the task specifies.
  3. In the answer, compare the reported number of predictions to the row count of the held-out file, and compare the reported prediction min/max/mean/spread against the observed target distribution in training data.
  4. Flag if no coverage/order assertion exists, if counts are unexplained, if the output path/column name deviates from the request, or if the predicted spread is implausibly compressed relative to the training target range with no justification.
Discriminator
A real violation is unverified row loss/reordering, a wrong path/column header, or unexamined distribution mismatch; it is not a violation if the script explicitly asserts equal row counts against the raw file, imputes rather than drops test rows, writes exactly the requested column/path, and the reported prediction distribution is consistent with the training target (mild shrinkage from regression averaging is normal and can be stated).
Consequence
The grader reads the output file and finds a misaligned, truncated, or wrongly named/located column, so the file compares as WRONG/MISSING regardless of the model's internal validation score.
id d978e9a685b3 · mined from da-code dacode-ml-regression-004@s8
raw text (what the judge reads)
### Prediction file not verified for row-for-row coverage, ordering, and value-distribution sanity
- **Applies when**: `task` -- the task asks for a prediction/derived column written to a file for every record of a held-out input set, and the script does its own cleaning (dropping/filtering rows, coercing dtypes, keeping only a subset of columns) before predicting.
- **Pattern**: The agent cleans train and test independently, silently drops or reorders rows (missing values, dtype coercion, deduplication, subsetting to "numeric only" features), writes predictions without asserting that the output length and order match the raw input file, and reports summary stats (e.g., a very narrow predicted range, an odd row count) without checking them against the target's actual distribution or the input's row count.
- **Detection procedure**:
  1. From the task, note the required output: file name/location, exact column name(s), and that there must be one prediction per row of the provided held-out file in its original order.
  2. In the scripts, trace the test dataframe from load to write: look for `dropna`, boolean filtering, `merge`, `groupby`, `sort_values`, index resets, or feature-subsetting that could change row count or order, and check whether any explicit assertion compares `len(predictions)` to `len(raw_test)` and whether the file is written to the path the task specifies.
  3. In the answer, compare the reported number of predictions to the row count of the held-out file, and compare the reported prediction min/max/mean/spread against the observed target distribution in training data.
  4. Flag if no coverage/order assertion exists, if counts are unexplained, if the output path/column name deviates from the request, or if the predicted spread is implausibly compressed relative to the training target range with no justification.
- **Discriminator**: A real violation is unverified row loss/reordering, a wrong path/column header, or unexamined distribution mismatch; it is *not* a violation if the script explicitly asserts equal row counts against the raw file, imputes rather than drops test rows, writes exactly the requested column/path, and the reported prediction distribution is consistent with the training target (mild shrinkage from regression averaging is normal and can be stated).
- **Consequence**: The grader reads the output file and finds a misaligned, truncated, or wrongly named/located column, so the file compares as WRONG/MISSING regardless of the model's internal validation score.
442Group-wise statistic computed over the wrong slice/axis of the tabletaskinfiagent-dabench
Applies when
task -- the task asks for the per-entity value of a distributional statistic (skewness, variance, kurtosis, etc.) that is then argmax/argmin'd across entities, and the data are stored with entities on one axis and observations (times, categories, repeated measures) on the other.
Pattern
The attempt picks an intuitive but unstated slice — e.g. aggregating along the wrong axis (across entities within one column instead of across observations within one entity), silently dropping rows with missing observations so different entities are scored on different sample sizes, or filtering to a single observation and computing the statistic on a degenerate vector — and reports the argmax of that quantity without ever checking what population each number summarizes.
Detection procedure
  1. From the task statement, write down explicitly the intended unit of analysis and the vector the statistic is applied to: one number per entity, computed over which set of values? Note any stated definition/flag (e.g. bias-corrected vs. Fisher/Pearson variant) and confirm the library call uses it.
  2. In the scripts, locate the reduction call and check the axis/groupby key and the shape of its input and output: the output length must equal the number of entities, and each input vector must contain the intended observations (not one scalar, not the cross-entity cross-section).
  3. Check NaN/dtype handling before the reduction: confirm non-numeric placeholders are coerced, and that the per-entity sample counts are reported/roughly comparable rather than entities being silently reduced to 1–2 points.
  4. Check the reported winner against a printed sorted top-5 with the per-entity n and statistic value; a plausible answer should be reproducible from a saved script, and an unsaved/ad-hoc computation with no printed diagnostics should be treated as unverified.
Discriminator
A real violation is when the reduction's input vector or grouping key does not match the unit of analysis the task names (wrong axis, degenerate n, unequal ad-hoc dropping), or a stated variant/definition flag is left at a different default. It is not a violation if the axis/grouping is correct and documented and the ambiguity is only in tie-breaking or in cosmetic ordering of output.
Consequence
The reported entity is the argmax of a different quantity than requested, so the single-value answer key mismatches (0/1 checks passed) even though the code runs without error.
id ee08e054b3e5 · mined from infiagent-dabench dabench-252@s8
raw text (what the judge reads)
### Group-wise statistic computed over the wrong slice/axis of the table
- **Applies when**: `task` -- the task asks for the per-entity value of a distributional statistic (skewness, variance, kurtosis, etc.) that is then argmax/argmin'd across entities, and the data are stored with entities on one axis and observations (times, categories, repeated measures) on the other.
- **Pattern**: The attempt picks an intuitive but unstated slice — e.g. aggregating along the wrong axis (across entities within one column instead of across observations within one entity), silently dropping rows with missing observations so different entities are scored on different sample sizes, or filtering to a single observation and computing the statistic on a degenerate vector — and reports the argmax of that quantity without ever checking what population each number summarizes.
- **Detection procedure**:
  1. From the task statement, write down explicitly the intended unit of analysis and the vector the statistic is applied to: one number per entity, computed over which set of values? Note any stated definition/flag (e.g. bias-corrected vs. Fisher/Pearson variant) and confirm the library call uses it.
  2. In the scripts, locate the reduction call and check the `axis`/`groupby` key and the shape of its input and output: the output length must equal the number of entities, and each input vector must contain the intended observations (not one scalar, not the cross-entity cross-section).
  3. Check NaN/dtype handling before the reduction: confirm non-numeric placeholders are coerced, and that the per-entity sample counts are reported/roughly comparable rather than entities being silently reduced to 1–2 points.
  4. Check the reported winner against a printed sorted top-5 with the per-entity n and statistic value; a plausible answer should be reproducible from a saved script, and an unsaved/ad-hoc computation with no printed diagnostics should be treated as unverified.
- **Discriminator**: A real violation is when the reduction's input vector or grouping key does not match the unit of analysis the task names (wrong axis, degenerate n, unequal ad-hoc dropping), or a stated variant/definition flag is left at a different default. It is *not* a violation if the axis/grouping is correct and documented and the ambiguity is only in tie-breaking or in cosmetic ordering of output.
- **Consequence**: The reported entity is the argmax of a different quantity than requested, so the single-value answer key mismatches (0/1 checks passed) even though the code runs without error.
443Answer-slot filled with data values instead of the requested identifiertaskinfiagent-dabench
Applies when
task -- The task asks for an answer token inside a named placeholder (e.g. @key[<something>]) and the placeholder description names an entity such as a column/feature/model/variable rather than a numeric result.
Pattern
The agent performs the computation correctly but pastes the entire computed series/table (or an intermediate array) into the answer slot, instead of the single identifier (name/label) that the answer format actually asks for; the payload is often truncated or of unbounded length.
Detection procedure
  1. Read the answer-format spec and decide what type the slot expects: a name/label, a single scalar, or an explicit list — treat the parenthetical gloss ("where X refers to …") as defining the referent, not the content to dump.
  2. Inspect the submitted answer's type and cardinality: is it one token, or hundreds of comma-separated numbers?
  3. Check whether the answer would still be interpretable/gradable if the underlying file were unavailable; a long value dump that depends on row order and length is a red flag.
  4. Confirm the scripts actually persist/return the artifact (e.g. the new column with its required rounding) so the identifier answer is backed by real work.
Discriminator
A real violation is when the spec names an object ("the new column", "the selected model", "the chosen variable") and the agent emits its contents; it is not a violation when the spec explicitly asks for the values/list, or for a scalar statistic, in which case dumping the numbers (with stated precision) is correct.
Consequence
The grader string-matches the expected identifier against a long numeric list and marks the check WRONG/MISSING, scoring 0 even though the underlying computation was right.
id 1eadf2bdac9b · mined from infiagent-dabench dabench-741@s8
raw text (what the judge reads)
### Answer-slot filled with data values instead of the requested identifier
- **Applies when**: `task` -- The task asks for an answer token inside a named placeholder (e.g. `@key[<something>]`) and the placeholder description names an entity such as a column/feature/model/variable rather than a numeric result.
- **Pattern**: The agent performs the computation correctly but pastes the entire computed series/table (or an intermediate array) into the answer slot, instead of the single identifier (name/label) that the answer format actually asks for; the payload is often truncated or of unbounded length.
- **Detection procedure**:
  1. Read the answer-format spec and decide what type the slot expects: a name/label, a single scalar, or an explicit list — treat the parenthetical gloss ("where X refers to …") as defining the *referent*, not the content to dump.
  2. Inspect the submitted answer's type and cardinality: is it one token, or hundreds of comma-separated numbers?
  3. Check whether the answer would still be interpretable/gradable if the underlying file were unavailable; a long value dump that depends on row order and length is a red flag.
  4. Confirm the scripts actually persist/return the artifact (e.g. the new column with its required rounding) so the identifier answer is backed by real work.
- **Discriminator**: A real violation is when the spec names an object ("the new column", "the selected model", "the chosen variable") and the agent emits its contents; it is *not* a violation when the spec explicitly asks for the values/list, or for a scalar statistic, in which case dumping the numbers (with stated precision) is correct.
- **Consequence**: The grader string-matches the expected identifier against a long numeric list and marks the check WRONG/MISSING, scoring 0 even though the underlying computation was right.
444Ignoring the provided template/sample output file when producing the deliverabletaskda-code
Applies when
task -- The task says the saved result must match the format of a provided sample/example output file (or a stated schema), and the scripts write the final file.
Pattern
The agent invents its own column names, ordering, rounding and row set from the wording of the prompt, and never opens or compares against the referenced sample file; the "verification" script only re-derives its own numbers instead of checking conformance to the template.
Detection procedure
  1. In the task text, note every explicitly referenced artifact that constrains the output (sample file, schema, required columns, rounding, sort order, units).
  2. Search the scripts for any read of that reference artifact and any comparison of the written file's header/dtypes/row keys/precision against it.
  3. Inspect the produced answer's header and rows: are the column labels, category/key set, ordering and numeric precision demonstrably derived from the template, or guessed (e.g., pretty-printed names, arbitrary rounding, alphabetical sort chosen "for consistency")?
  4. Flag if no step in the pipeline ever loads or asserts against the template.
Discriminator
A real violation is when the reference artifact is available and never read/asserted against, so any mismatch in naming, key coverage, ordering, or precision goes undetected; it is fine if the script loads the sample, aligns columns/keys/precision to it (or asserts equality of structure), even if it also hard-codes the same names.
Consequence
The grader's file-level comparison fails ("WRONG/MISSING") even when the underlying aggregation logic may be defensible, because headers, key rows, ordering, or rounding differ from the expected file.
id 2f0cdf502141 · mined from da-code dacode-dm-csv-010@s8
raw text (what the judge reads)
### Ignoring the provided template/sample output file when producing the deliverable
- **Applies when**: `task` -- The task says the saved result must match the format of a provided sample/example output file (or a stated schema), and the scripts write the final file.
- **Pattern**: The agent invents its own column names, ordering, rounding and row set from the wording of the prompt, and never opens or compares against the referenced sample file; the "verification" script only re-derives its own numbers instead of checking conformance to the template.
- **Detection procedure**:
  1. In the task text, note every explicitly referenced artifact that constrains the output (sample file, schema, required columns, rounding, sort order, units).
  2. Search the scripts for any read of that reference artifact and any comparison of the written file's header/dtypes/row keys/precision against it.
  3. Inspect the produced answer's header and rows: are the column labels, category/key set, ordering and numeric precision demonstrably derived from the template, or guessed (e.g., pretty-printed names, arbitrary rounding, alphabetical sort chosen "for consistency")?
  4. Flag if no step in the pipeline ever loads or asserts against the template.
- **Discriminator**: A real violation is when the reference artifact is available and never read/asserted against, so any mismatch in naming, key coverage, ordering, or precision goes undetected; it is fine if the script loads the sample, aligns columns/keys/precision to it (or asserts equality of structure), even if it also hard-codes the same names.
- **Consequence**: The grader's file-level comparison fails ("WRONG/MISSING") even when the underlying aggregation logic may be defensible, because headers, key rows, ordering, or rounding differ from the expected file.
445Loss of value precision when coercing an identified record's key into a stated format templatetaskinfiagent-dabench
Applies when
task -- the deliverable includes a specific value identified from the data (a date, ID, label, timestamp) that must be reported alongside a derived statistic, and the task's format hint specifies a coarser granularity than the value actually carries.
Pattern
The attempt correctly locates the record and computes the derived statistic from the full-precision key, but then truncates/reformats the reported key to match the literal format template (e.g., dropping the finest component), so the reported identifier no longer uniquely designates the record used in the computation.
Detection procedure
  1. Read the task and note the granularity of the key stored in the data versus the granularity implied by the answer-format template.
  2. In the scripts, find where the key is selected and where it is formatted for output; check for any truncation, strftime/substring/rounding/casting that discards components.
  3. Check for internal consistency: does the reported key, taken literally, still identify exactly one record, and is it the same record used to compute the accompanying statistic? If the truncated key maps to many records (or to a different one), flag it.
  4. Prefer the full-precision value from the data; only shorten if the task explicitly demands aggregation at the coarser level (e.g., "the month with the highest average").
Discriminator
A real violation is when the underlying computation used a fine-grained record but the reported key is coarsened, making it ambiguous. It is not a violation when the analysis itself was genuinely performed at the coarse level (data aggregated to that granularity), so the coarse key is the true unit of analysis.
Consequence
The derived statistic matches, but the identifier check fails on exact string comparison, giving a partial-credit/incorrect verdict (e.g., 1 of 2 checks passed).
id cdfd84dce9e6 · mined from infiagent-dabench dabench-572@s8
raw text (what the judge reads)
### Loss of value precision when coercing an identified record's key into a stated format template
- **Applies when**: `task` -- the deliverable includes a specific value identified from the data (a date, ID, label, timestamp) that must be reported alongside a derived statistic, and the task's format hint specifies a coarser granularity than the value actually carries.
- **Pattern**: The attempt correctly locates the record and computes the derived statistic from the full-precision key, but then truncates/reformats the reported key to match the literal format template (e.g., dropping the finest component), so the reported identifier no longer uniquely designates the record used in the computation.
- **Detection procedure**:
  1. Read the task and note the granularity of the key stored in the data versus the granularity implied by the answer-format template.
  2. In the scripts, find where the key is selected and where it is formatted for output; check for any truncation, `strftime`/substring/rounding/casting that discards components.
  3. Check for internal consistency: does the reported key, taken literally, still identify exactly one record, and is it the same record used to compute the accompanying statistic? If the truncated key maps to many records (or to a different one), flag it.
  4. Prefer the full-precision value from the data; only shorten if the task explicitly demands aggregation at the coarser level (e.g., "the month with the highest average").
- **Discriminator**: A real violation is when the underlying computation used a fine-grained record but the reported key is coarsened, making it ambiguous. It is *not* a violation when the analysis itself was genuinely performed at the coarse level (data aggregated to that granularity), so the coarse key is the true unit of analysis.
- **Consequence**: The derived statistic matches, but the identifier check fails on exact string comparison, giving a partial-credit/incorrect verdict (e.g., 1 of 2 checks passed).
446Decorated identifiers in the final answer string (extra quotes/brackets/units not in the requested format)taskinfiagent-dabench
Applies when
task -- The task specifies an exact answer template (e.g. @field[item1,item2,...]) and the script builds that string by joining values pulled from a dataframe.
Pattern
The attempt computes the correct values but serializes them with added decoration — wrapping each item in quotes, adding spaces, appending units, or keeping Python repr/list syntax — so the emitted tokens do not literally match the expected plain comma-separated items, even though the underlying analysis is right.
Detection procedure
  1. Read the task's answer-format spec and write down the exact literal token pattern expected inside each field (are items bare strings/numbers? separator? any quoting shown in the template?).
  2. In the script, find the line(s) that build the answer string (join, f-strings, print) and check whether any per-item formatting is applied (e.g. f'"{x}"', str(list), %.2f with padding, unit suffixes).
  3. Compare the final printed/submitted string character-by-character against the template; for string-valued identifiers confirm they appear exactly as in the data with no added delimiters.
  4. Also confirm numeric fields follow the stated rounding/precision and that the two fields are in the same order (item i in one field corresponds to item i in the other).
Discriminator
A real violation is decoration the task's template never showed (quotes, brackets, units, whitespace) around otherwise-correct items; it is not a violation if the quoting/characters are genuinely part of the data value itself (e.g. an identifier that contains parentheses) or if the template explicitly displayed quotes.
Consequence
Automated string/field matching marks that field WRONG/MISSING despite correct computation — e.g. the numeric field passes while the identifier field fails, yielding a partial-credit failure.
id a4ec5b96c191 · mined from infiagent-dabench dabench-219@s8
raw text (what the judge reads)
### Decorated identifiers in the final answer string (extra quotes/brackets/units not in the requested format)
- **Applies when**: `task` -- The task specifies an exact answer template (e.g. `@field[item1,item2,...]`) and the script builds that string by joining values pulled from a dataframe.
- **Pattern**: The attempt computes the correct values but serializes them with added decoration — wrapping each item in quotes, adding spaces, appending units, or keeping Python `repr`/list syntax — so the emitted tokens do not literally match the expected plain comma-separated items, even though the underlying analysis is right.
- **Detection procedure**:
  1. Read the task's answer-format spec and write down the exact literal token pattern expected inside each field (are items bare strings/numbers? separator? any quoting shown in the template?).
  2. In the script, find the line(s) that build the answer string (`join`, f-strings, `print`) and check whether any per-item formatting is applied (e.g. `f'"{x}"'`, `str(list)`, `%.2f` with padding, unit suffixes).
  3. Compare the final printed/submitted string character-by-character against the template; for string-valued identifiers confirm they appear exactly as in the data with no added delimiters.
  4. Also confirm numeric fields follow the stated rounding/precision and that the two fields are in the same order (item *i* in one field corresponds to item *i* in the other).
- **Discriminator**: A real violation is decoration the task's template never showed (quotes, brackets, units, whitespace) around otherwise-correct items; it is *not* a violation if the quoting/characters are genuinely part of the data value itself (e.g. an identifier that contains parentheses) or if the template explicitly displayed quotes.
- **Consequence**: Automated string/field matching marks that field WRONG/MISSING despite correct computation — e.g. the numeric field passes while the identifier field fails, yielding a partial-credit failure.
447Output row count / shape not validated against the test set and sample formattaskda-code
Applies when
task -- the task asks for a per-row prediction (or per-record output) file whose format is defined by a sample file and whose length must equal the number of rows in the provided evaluation input.
Pattern
The attempt produces a prediction file (often with a plausible-looking header and 0/1 values) without any explicit check that len(predictions) == len(test_input), that the header/column name and ordering match the sample, and that no rows were dropped by preprocessing (NaN dropping, filtering, deduplication, partial-batch inference, truncated write/print). The submitted file ends up with a different number of rows than the evaluation input, or in a different row order, so it cannot align with the ground truth regardless of model quality.
Detection procedure
  1. From the task/README, note the expected number of output rows (test file row count) and the exact required column name/format from the sample output file.
  2. In the scripts, trace the test data path end-to-end: look for any dropna, boolean filtering, drop_duplicates, resampling, index resetting/sorting, or writing from a slice/partial loop; confirm predictions are produced for every input row in original order and written with index=False and the required header.
  3. Check that the script asserts or prints the shape of the loaded test data and of the final written file, and compares both to the sample; absence of any such check (or absence of saved scripts to verify) is itself a flag.
  4. Count the rows in the submitted answer file and compare to the expected test row count and to the sample's structure.
Discriminator
A real violation is a length/order/column mismatch (or no evidence anywhere that length and format were checked). Not a violation: a file with the correct number of rows, correct single header, and original order — even if the class balance looks unusual or accuracy is poor, since that is a modeling-quality issue rather than an alignment failure.
Consequence
The grader reports the expected output file as WRONG/MISSING because rows cannot be compared one-to-one with the ground truth, scoring effectively zero even if the underlying model was reasonable.
id dcef3d25df3e · mined from da-code dacode-ml-binary-013@s8
raw text (what the judge reads)
### Output row count / shape not validated against the test set and sample format
- **Applies when**: `task` -- the task asks for a per-row prediction (or per-record output) file whose format is defined by a sample file and whose length must equal the number of rows in the provided evaluation input.
- **Pattern**: The attempt produces a prediction file (often with a plausible-looking header and 0/1 values) without any explicit check that `len(predictions) == len(test_input)`, that the header/column name and ordering match the sample, and that no rows were dropped by preprocessing (NaN dropping, filtering, deduplication, partial-batch inference, truncated write/print). The submitted file ends up with a different number of rows than the evaluation input, or in a different row order, so it cannot align with the ground truth regardless of model quality.
- **Detection procedure**:
  1. From the task/README, note the expected number of output rows (test file row count) and the exact required column name/format from the sample output file.
  2. In the scripts, trace the test data path end-to-end: look for any `dropna`, boolean filtering, `drop_duplicates`, resampling, index resetting/sorting, or writing from a slice/partial loop; confirm predictions are produced for every input row in original order and written with `index=False` and the required header.
  3. Check that the script asserts or prints the shape of the loaded test data and of the final written file, and compares both to the sample; absence of any such check (or absence of saved scripts to verify) is itself a flag.
  4. Count the rows in the submitted answer file and compare to the expected test row count and to the sample's structure.
- **Discriminator**: A real violation is a length/order/column mismatch (or no evidence anywhere that length and format were checked). Not a violation: a file with the correct number of rows, correct single header, and original order — even if the class balance looks unusual or accuracy is poor, since that is a modeling-quality issue rather than an alignment failure.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING because rows cannot be compared one-to-one with the ground truth, scoring effectively zero even if the underlying model was reasonable.
448Analyzing a convenience subset instead of the full population the question asks abouttaskinfiagent-dabench
Applies when
task -- the question specifies a universe of entities ("all X") but the data lives in several files/partitions or one file that covers only part of that universe.
Pattern
The script loads a single file (or one partition/group) without checking whether it contains every entity in the requested scope, computes the statistic on that partial set, and reports the result as if it covered the whole population — quantiles, thresholds and counts are therefore derived from the wrong denominator, even when the surfaced answer happens to overlap the true one.
Detection procedure
  1. Read the task and write down the intended universe (all entities / a stated filter) and any implied row count.
  2. Read the load step in the script: list which file(s)/partitions are read and whether any concatenation, deduplication, or scope check is done.
  3. Compare the number and identity of entities actually present (printed len, describe, or sorted listing) against the universe from step 1; look for a missing-scope assertion or a sanity check on counts.
  4. Check whether the reported answer names/labels are drawn from the full universe and formatted exactly as the task requests.
Discriminator
A real violation is when the loaded source is provably narrower than the requested scope (one region/group/split among several available) and no evidence is given that the omitted rows can't change the statistic; it is fine if the script explicitly verifies the single file already contains the full universe, or if the task itself restricts the scope to that subset.
Consequence
Quartiles, bounds, and the outlier/statistic set are computed on a biased subsample, so the grader marks the reported entity list (or its exact expected form) as wrong/missing despite the script running without error.
id 737b7fbbcea6 · mined from infiagent-dabench dabench-254@s8
raw text (what the judge reads)
### Analyzing a convenience subset instead of the full population the question asks about
- **Applies when**: `task` -- the question specifies a universe of entities ("all X") but the data lives in several files/partitions or one file that covers only part of that universe.
- **Pattern**: The script loads a single file (or one partition/group) without checking whether it contains every entity in the requested scope, computes the statistic on that partial set, and reports the result as if it covered the whole population — quantiles, thresholds and counts are therefore derived from the wrong denominator, even when the surfaced answer happens to overlap the true one.
- **Detection procedure**:
  1. Read the task and write down the intended universe (all entities / a stated filter) and any implied row count.
  2. Read the load step in the script: list which file(s)/partitions are read and whether any concatenation, deduplication, or scope check is done.
  3. Compare the number and identity of entities actually present (printed `len`, `describe`, or sorted listing) against the universe from step 1; look for a missing-scope assertion or a sanity check on counts.
  4. Check whether the reported answer names/labels are drawn from the full universe and formatted exactly as the task requests.
- **Discriminator**: A real violation is when the loaded source is provably narrower than the requested scope (one region/group/split among several available) and no evidence is given that the omitted rows can't change the statistic; it is fine if the script explicitly verifies the single file already contains the full universe, or if the task itself restricts the scope to that subset.
- **Consequence**: Quartiles, bounds, and the outlier/statistic set are computed on a biased subsample, so the grader marks the reported entity list (or its exact expected form) as wrong/missing despite the script running without error.
449Unverified prediction artifact: label vocabulary, path, and row alignment never checked against the source datataskda-code
Applies when
task -- the task asks for predictions written to a specific output file with a specified column name, and the agent reports success from summary statistics rather than from the saved file itself.
Pattern
The agent trains a model and declares completion, but never re-reads the written file to confirm (a) it sits at the expected path/name the task requested, (b) the predicted values use the exact label strings/encoding present in the training target (not encoded integers, re-spelled/re-cased variants, or extra index/columns), and (c) the row count and order match the test rows one-to-one; no held-out score is reported, so the label mapping could even be inverted without notice.
Detection procedure
  1. From the task text, note the required file path/name, column header, and the exact set of target values as they appear in the training data.
  2. In the scripts, find the write step: check the destination path is the one asked for (not a home/scratch directory), that index=False-style options avoid extra columns, and that the values written come from an inverse mapping back to the original label strings.
  3. Check for a read-back/sanity block: does anything re-load the output and assert shape == number of test rows, header spelling, and set(values) ⊆ set(training labels)?
  4. In the answer, look for a validation/holdout metric and a class-distribution comparison against training; a report of only counts and model hyperparameters is insufficient evidence of correctness.
Discriminator
A real violation is an attempt whose only evidence is self-reported counts, with no read-back assertion, no confirmed path, and no held-out accuracy; a look-alike that is fine explicitly reloads the saved file, asserts the header/shape/label vocabulary, and reports a validation score that rules out inverted or mis-encoded labels.
Consequence
The grader cannot find or cannot parse the expected file (wrong location, extra index column, wrong header) or compares mismatched label spellings/encodings, so every row scores as wrong and the check fails despite a well-performing model.
id 8a43101385e2 · mined from da-code dacode-ml-binary-009@s8
raw text (what the judge reads)
### Unverified prediction artifact: label vocabulary, path, and row alignment never checked against the source data
- **Applies when**: `task` -- the task asks for predictions written to a specific output file with a specified column name, and the agent reports success from summary statistics rather than from the saved file itself.
- **Pattern**: The agent trains a model and declares completion, but never re-reads the written file to confirm (a) it sits at the expected path/name the task requested, (b) the predicted values use the exact label strings/encoding present in the training target (not encoded integers, re-spelled/re-cased variants, or extra index/columns), and (c) the row count and order match the test rows one-to-one; no held-out score is reported, so the label mapping could even be inverted without notice.
- **Detection procedure**:
  1. From the task text, note the required file path/name, column header, and the exact set of target values as they appear in the training data.
  2. In the scripts, find the write step: check the destination path is the one asked for (not a home/scratch directory), that `index=False`-style options avoid extra columns, and that the values written come from an inverse mapping back to the original label strings.
  3. Check for a read-back/sanity block: does anything re-load the output and assert shape == number of test rows, header spelling, and `set(values) ⊆ set(training labels)`?
  4. In the answer, look for a validation/holdout metric and a class-distribution comparison against training; a report of only counts and model hyperparameters is insufficient evidence of correctness.
- **Discriminator**: A real violation is an attempt whose only evidence is self-reported counts, with no read-back assertion, no confirmed path, and no held-out accuracy; a look-alike that is fine explicitly reloads the saved file, asserts the header/shape/label vocabulary, and reports a validation score that rules out inverted or mis-encoded labels.
- **Consequence**: The grader cannot find or cannot parse the expected file (wrong location, extra index column, wrong header) or compares mismatched label spellings/encodings, so every row scores as wrong and the check fails despite a well-performing model.
450Fabricated aggregation metric computed over only one of the entity's two rolestaskda-code
Applies when
task -- the task asks for a per-entity "performance"/summary chart or table, where each record lists the entity in one of several role columns and the target quantity is defined by the task, a config file, or an axis label rather than invented by the analyst.
Pattern
The script makes up an arbitrary formula for the requested quantity (e.g., a weighted sum of counts with no basis in the task, config, or domain) and aggregates records matching the entity in a single role column, silently discarding all records where the entity appears in the other role(s); it also skips any auxiliary result artifacts implied by the task/config.
Detection procedure
  1. Read the task text and any provided config/spec (axis labels, title, output keys) and write down what quantity is actually being requested and which output files are implied.
  2. In the script, locate where that quantity is computed: check whether its formula is traceable to the task/config/domain definition or was invented, and check whether the row filter uses every column in which the entity can appear.
  3. Cross-check the reported numbers against a cheap sanity check (e.g., totals per entity vs. total records, symmetry between roles, plausible ranges); confirm every required output artifact is written, not just the figure.
  4. Flag if the formula is unjustified, if only one role column is scanned, or if implied artifacts (numeric/serialized results alongside the plot) are missing.
Discriminator
A genuine violation is an undefined/invented formula or a one-sided subset that provably drops relevant rows; it is not a violation when the task or config explicitly restricts the analysis to one role/subset, or when the formula is stated in the task or is the standard domain definition (in which case the script should cite it in a comment or config lookup).
Consequence
The bar heights, ranking, and any saved numeric arrays differ from the reference, so the figure and all derived result files fail comparison even though the plot styling looks correct.
id 2ffe5f5c8057 · mined from da-code dacode-plot-bar-006@s8
raw text (what the judge reads)
### Fabricated aggregation metric computed over only one of the entity's two roles
- **Applies when**: `task` -- the task asks for a per-entity "performance"/summary chart or table, where each record lists the entity in one of several role columns and the target quantity is defined by the task, a config file, or an axis label rather than invented by the analyst.
- **Pattern**: The script makes up an arbitrary formula for the requested quantity (e.g., a weighted sum of counts with no basis in the task, config, or domain) and aggregates records matching the entity in a single role column, silently discarding all records where the entity appears in the other role(s); it also skips any auxiliary result artifacts implied by the task/config.
- **Detection procedure**:
  1. Read the task text and any provided config/spec (axis labels, title, output keys) and write down what quantity is actually being requested and which output files are implied.
  2. In the script, locate where that quantity is computed: check whether its formula is traceable to the task/config/domain definition or was invented, and check whether the row filter uses every column in which the entity can appear.
  3. Cross-check the reported numbers against a cheap sanity check (e.g., totals per entity vs. total records, symmetry between roles, plausible ranges); confirm every required output artifact is written, not just the figure.
  4. Flag if the formula is unjustified, if only one role column is scanned, or if implied artifacts (numeric/serialized results alongside the plot) are missing.
- **Discriminator**: A genuine violation is an undefined/invented formula or a one-sided subset that provably drops relevant rows; it is *not* a violation when the task or config explicitly restricts the analysis to one role/subset, or when the formula is stated in the task or is the standard domain definition (in which case the script should cite it in a comment or config lookup).
- **Consequence**: The bar heights, ranking, and any saved numeric arrays differ from the reference, so the figure and all derived result files fail comparison even though the plot styling looks correct.
451Adding unrequested preprocessing that changes the reported statistictaskinfiagent-dabench
Applies when
task -- the task specifies a precise, literal computation (a formula/threshold rule applied to a named column) and the scripts insert extra data-cleaning steps (row filtering, NaN dropping, type coercion, rescaling) that were not asked for before computing the reported number.
Pattern
The agent first runs the spec-literal computation, gets a degenerate or "boring" result (e.g. all-NaN scores, zero hits), decides it must be an artifact, then re-runs on a silently reduced/altered subset and reports that second, larger number — never reconciling the two or flagging the discrepancy, and never sanity-checking the new number against what the stated rule should theoretically produce.
Detection procedure
  1. Read the task and list exactly which operations are mandated (column, formula, threshold, "drop rows and build new frame") and which are not mentioned (e.g. missing-value handling, subsetting).
  2. Read the scripts in order and note every filtering/transform step applied before the reported statistic; check whether each one is mandated by the task or invented by the agent.
  3. Compare the numbers produced by the spec-literal version and the agent-modified version; if they differ and the agent reports the modified one without justification from the task text, flag it.
  4. Sanity-check the reported magnitude against the rule's implied scale (e.g. a |z|>3 rule should tag a very small fraction of a roughly bell-shaped column; a few percent signals a broken denominator, an altered population, or a library NaN policy issue rather than genuine outliers).
Discriminator
A real violation is preprocessing the agent introduced on its own initiative that materially changes the answer and is reported without discussion. It is fine if the extra step is explicitly required by the task/constraints, is provably answer-neutral (verified both ways giving the same number), or the agent shows the spec-literal result and explains why it is invalid.
Consequence
The submitted count comes from a different population/definition than the reference implementation, so the single graded integer mismatches the ground truth (here 97 vs 0) and the attempt scores 0.
id 0b23d0a932c1 · mined from infiagent-dabench dabench-361@s8
raw text (what the judge reads)
### Adding unrequested preprocessing that changes the reported statistic
- **Applies when**: `task` -- the task specifies a precise, literal computation (a formula/threshold rule applied to a named column) and the scripts insert extra data-cleaning steps (row filtering, NaN dropping, type coercion, rescaling) that were not asked for before computing the reported number.
- **Pattern**: The agent first runs the spec-literal computation, gets a degenerate or "boring" result (e.g. all-NaN scores, zero hits), decides it must be an artifact, then re-runs on a silently reduced/altered subset and reports that second, larger number — never reconciling the two or flagging the discrepancy, and never sanity-checking the new number against what the stated rule should theoretically produce.
- **Detection procedure**:
  1. Read the task and list exactly which operations are mandated (column, formula, threshold, "drop rows and build new frame") and which are *not* mentioned (e.g. missing-value handling, subsetting).
  2. Read the scripts in order and note every filtering/transform step applied before the reported statistic; check whether each one is mandated by the task or invented by the agent.
  3. Compare the numbers produced by the spec-literal version and the agent-modified version; if they differ and the agent reports the modified one without justification from the task text, flag it.
  4. Sanity-check the reported magnitude against the rule's implied scale (e.g. a |z|>3 rule should tag a very small fraction of a roughly bell-shaped column; a few percent signals a broken denominator, an altered population, or a library NaN policy issue rather than genuine outliers).
- **Discriminator**: A real violation is preprocessing the agent introduced on its own initiative that materially changes the answer and is reported without discussion. It is *fine* if the extra step is explicitly required by the task/constraints, is provably answer-neutral (verified both ways giving the same number), or the agent shows the spec-literal result and explains why it is invalid.
- **Consequence**: The submitted count comes from a different population/definition than the reference implementation, so the single graded integer mismatches the ground truth (here 97 vs 0) and the attempt scores 0.
452Missing required output artifact and no reproducible computationtaskda-code
Applies when
task -- the task specifies that results be saved to a named output file (and/or the answer depends on a mapping/rule defined in an auxiliary instructions file), and the agent replies with a value inline.
Pattern
The agent reports a number/label directly in its chat answer without saving the required result file and without leaving any script that reads the data, applies the stated transformation rules, and computes the statistic — so the value is unverifiable and the expected artifact is absent.
Detection procedure
  1. Read the task and note every required deliverable: exact output filename(s), format/keys, plus any external rule file whose mapping must be applied before aggregating.
  2. Inspect the submitted scripts/workspace: is there code that loads the data, applies the stated mapping, computes the requested statistic, and writes the named file with the required key names?
  3. Check whether the reported answer can be traced to that code's output (mapped category labels, ratio definition = count of top category / total, rounding as specified).
  4. If no artifact is written or no script exists to reproduce the value, flag the attempt regardless of whether the number looks plausible.
Discriminator
A real violation is the absence of the named artifact or of any code path producing the reported value; it is not a violation if the file is written (even by a short inline script) with the correct keys and the printed/inline answer merely mirrors that file's contents.
Consequence
The grader looks for the specified result file, finds it missing or mismatched in keys/values, and scores 0 even if the inline text happens to be near-correct.
id 4db29f5ddbb9 · mined from da-code dacode-di-text-004@s8
raw text (what the judge reads)
### Missing required output artifact and no reproducible computation
- **Applies when**: `task` -- the task specifies that results be saved to a named output file (and/or the answer depends on a mapping/rule defined in an auxiliary instructions file), and the agent replies with a value inline.
- **Pattern**: The agent reports a number/label directly in its chat answer without saving the required result file and without leaving any script that reads the data, applies the stated transformation rules, and computes the statistic — so the value is unverifiable and the expected artifact is absent.
- **Detection procedure**:
  1. Read the task and note every required deliverable: exact output filename(s), format/keys, plus any external rule file whose mapping must be applied before aggregating.
  2. Inspect the submitted scripts/workspace: is there code that loads the data, applies the stated mapping, computes the requested statistic, and writes the named file with the required key names?
  3. Check whether the reported answer can be traced to that code's output (mapped category labels, ratio definition = count of top category / total, rounding as specified).
  4. If no artifact is written or no script exists to reproduce the value, flag the attempt regardless of whether the number looks plausible.
- **Discriminator**: A real violation is the absence of the named artifact or of any code path producing the reported value; it is *not* a violation if the file is written (even by a short inline script) with the correct keys and the printed/inline answer merely mirrors that file's contents.
- **Consequence**: The grader looks for the specified result file, finds it missing or mismatched in keys/values, and scores 0 even if the inline text happens to be near-correct.
453Submission artifact not verified against the provided template (header names, row count, on-disk file)taskda-code
Applies when
task -- the task requires producing a prediction/output file whose format is defined by a provided sample/template file, and the agent's deliverable is the file itself.
Pattern
The agent emits predictions as inline text (or writes a file it never re-reads), copying column names/casing from memory rather than from the template, and truncates or omits rows so the output does not cover every required identifier; no script is retained that provably writes the full file to the expected path.
Detection procedure
  1. Read the task/README to note the required output path and the exact header, column order, and identifier set implied by the template file.
  2. In the scripts, confirm there is code that (a) loads the template/test identifiers, (b) builds the output with the template's header spelling/case and identifier order, and (c) writes to the exact required filename.
  3. Check for a post-write sanity check: re-read the written file and assert row count equals the number of test identifiers, columns match the template exactly, IDs match one-to-one with no duplicates/missing, and probability columns are in [0,1] (and sum to 1 if required).
  4. Inspect the reported answer: if it is a pasted, cut-off listing, or its header differs in name/case from the template, treat the deliverable as unverified.
Discriminator
A real violation is missing/renamed/reordered columns, a truncated or partial ID set, or no evidence the file was written and re-validated; it is not a violation if the file is written and verified with matching shape/IDs/headers and the chat merely shows an abbreviated preview alongside that verification.
Consequence
The grader reports the expected output file as WRONG/MISSING — it cannot parse or align the submission (unrecognized header, missing rows/IDs), so the score is zero regardless of model quality.
id a4d555c054a1 · mined from da-code dacode-ml-competition-003@s8
raw text (what the judge reads)
### Submission artifact not verified against the provided template (header names, row count, on-disk file)
- **Applies when**: `task` -- the task requires producing a prediction/output file whose format is defined by a provided sample/template file, and the agent's deliverable is the file itself.
- **Pattern**: The agent emits predictions as inline text (or writes a file it never re-reads), copying column names/casing from memory rather than from the template, and truncates or omits rows so the output does not cover every required identifier; no script is retained that provably writes the full file to the expected path.
- **Detection procedure**:
  1. Read the task/README to note the required output path and the exact header, column order, and identifier set implied by the template file.
  2. In the scripts, confirm there is code that (a) loads the template/test identifiers, (b) builds the output with the template's header spelling/case and identifier order, and (c) writes to the exact required filename.
  3. Check for a post-write sanity check: re-read the written file and assert row count equals the number of test identifiers, columns match the template exactly, IDs match one-to-one with no duplicates/missing, and probability columns are in [0,1] (and sum to 1 if required).
  4. Inspect the reported answer: if it is a pasted, cut-off listing, or its header differs in name/case from the template, treat the deliverable as unverified.
- **Discriminator**: A real violation is missing/renamed/reordered columns, a truncated or partial ID set, or no evidence the file was written and re-validated; it is *not* a violation if the file is written and verified with matching shape/IDs/headers and the chat merely shows an abbreviated preview alongside that verification.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING — it cannot parse or align the submission (unrecognized header, missing rows/IDs), so the score is zero regardless of model quality.
454Predicted label strings not drawn verbatim from the training label vocabularytaskda-code
Applies when
task -- the task asks for predictions of a categorical target and the submission is a file whose values must match the label categories present in the provided training data.
Pattern
The attempt invents, lowercases, abbreviates, or re-buckets the category names (e.g. hand-typed or "cleaned" label strings, a collapsed set of classes) instead of emitting the exact distinct values observed in the training target column, so every row is scored as a mismatch even if the underlying ranking/model is reasonable.
Detection procedure
  1. From the task/README, identify the target column and, from the scripts, find where the label set is obtained — is it read from the training data (unique(), label encoder classes, inverse_transform) or hard-coded/mapped by hand?
  2. Collect the distinct values that appear in the submitted output column and compare them character-for-character (spelling, casing, hyphens/spaces, number of classes) with the distinct values in the training target.
  3. Also check row count, ID column name/order, and header names against the required output spec.
  4. Flag if any output value is not an exact member of the training label set, or if the number of distinct classes differs from the training data.
Discriminator
A real violation is an output value that does not appear verbatim in the training target column (or a class silently dropped/merged). It is fine if the script applies a documented, task-specified renaming/normalization, or if it round-trips labels through a fitted encoder whose classes come from the training data — then values will match exactly.
Consequence
The grader's exact-match comparison against the reference file fails on essentially all rows, giving ~0 accuracy / "WRONG" despite a plausible-looking prediction file.
id 5db56263240f · mined from da-code dacode-ml-multi-003@s8
raw text (what the judge reads)
### Predicted label strings not drawn verbatim from the training label vocabulary
- **Applies when**: `task` -- the task asks for predictions of a categorical target and the submission is a file whose values must match the label categories present in the provided training data.
- **Pattern**: The attempt invents, lowercases, abbreviates, or re-buckets the category names (e.g. hand-typed or "cleaned" label strings, a collapsed set of classes) instead of emitting the exact distinct values observed in the training target column, so every row is scored as a mismatch even if the underlying ranking/model is reasonable.
- **Detection procedure**:
  1. From the task/README, identify the target column and, from the scripts, find where the label set is obtained — is it read from the training data (`unique()`, label encoder classes, `inverse_transform`) or hard-coded/mapped by hand?
  2. Collect the distinct values that appear in the submitted output column and compare them character-for-character (spelling, casing, hyphens/spaces, number of classes) with the distinct values in the training target.
  3. Also check row count, ID column name/order, and header names against the required output spec.
  4. Flag if any output value is not an exact member of the training label set, or if the number of distinct classes differs from the training data.
- **Discriminator**: A real violation is an output value that does not appear verbatim in the training target column (or a class silently dropped/merged). It is fine if the script applies a documented, task-specified renaming/normalization, or if it round-trips labels through a fitted encoder whose classes come from the training data — then values will match exactly.
- **Consequence**: The grader's exact-match comparison against the reference file fails on essentially all rows, giving ~0 accuracy / "WRONG" despite a plausible-looking prediction file.
455Output format invented instead of read from the provided templatetaskda-code
Applies when
task -- The task supplies a template/example output file (or explicit schema) that the deliverable must match, and the script writes the result file directly from its own computed frame.
Pattern
The attempt never opens or parses the template; it guesses header names, index/label formatting, row/column orientation, ordering, and rounding from intuition, then declares success after re-reading only its own output.
Detection procedure
1) In the task text, note any mention of a template, sample, or "format must match" requirement and locate the referenced file. 2) Search the scripts for any read of that template (e.g. loading it, printing its head, comparing columns/shape/dtypes). 3) If absent, compare the script's hard-coded output structure (column labels, key formatting, pivot orientation, decimal rounding, presence of index) against the plausible template structure; any invented naming or formatting choice is a red flag. 4) Check whether the answer/validation step verifies the output against the template rather than merely reading back the file it just wrote.
Discriminator
A real violation is when no template comparison exists and format choices are self-invented; it is fine if the script loads the template (or its exact schema is quoted in the task) and explicitly aligns columns, labels, ordering, and rounding to it — even if it then rebuilds the frame itself.
Consequence
The file is produced and looks self-consistent, but the grader's cell/column/index-wise comparison against the expected file fails (mismatched headers, label format, orientation, or rounding), scoring 0.
id ae4bd4396e1f · mined from da-code dacode-dm-csv-044@s8
raw text (what the judge reads)
### Output format invented instead of read from the provided template
- **Applies when**: `task` -- The task supplies a template/example output file (or explicit schema) that the deliverable must match, and the script writes the result file directly from its own computed frame.
- **Pattern**: The attempt never opens or parses the template; it guesses header names, index/label formatting, row/column orientation, ordering, and rounding from intuition, then declares success after re-reading only its own output.
- **Detection procedure**: 1) In the task text, note any mention of a template, sample, or "format must match" requirement and locate the referenced file. 2) Search the scripts for any read of that template (e.g. loading it, printing its head, comparing columns/shape/dtypes). 3) If absent, compare the script's hard-coded output structure (column labels, key formatting, pivot orientation, decimal rounding, presence of index) against the plausible template structure; any invented naming or formatting choice is a red flag. 4) Check whether the answer/validation step verifies the output against the template rather than merely reading back the file it just wrote.
- **Discriminator**: A real violation is when no template comparison exists and format choices are self-invented; it is fine if the script loads the template (or its exact schema is quoted in the task) and explicitly aligns columns, labels, ordering, and rounding to it — even if it then rebuilds the frame itself.
- **Consequence**: The file is produced and looks self-consistent, but the grader's cell/column/index-wise comparison against the expected file fails (mismatched headers, label format, orientation, or rounding), scoring 0.
456Ordinal / metric-specific objective replaced by generic classification, with no baseline comparisontaskda-code
Applies when
task -- the evaluation metric is stated explicitly (e.g. rank/ordinal agreement, weighted error, custom loss) and the scripts fit off-the-shelf models whose default objective is something else (plain multiclass accuracy, unweighted log-loss).
Pattern
The attempt uses the stated metric only as a CV scorer while every candidate model optimizes an unrelated objective; it further distorts the target distribution (e.g. rebalancing weights, resampling) even though the metric rewards matching the ordered/marginal structure of the labels. No metric-appropriate alternative (ordinal/regression + tuned rounding thresholds, expected-value decoding, cost-sensitive decision rule) and no trivial baseline (constant/majority or simple rounded regression) are evaluated, so there is nothing to show the reported approach is better than a naive one; the final submission is simply whatever the last script wrote.
Detection procedure
  1. Read the task/README and note the exact scoring metric and how predictions are decoded into the required label space.
  2. In the scripts, check what loss each candidate model actually minimizes and how class probabilities are converted to final labels (argmax vs. thresholding/rounding tuned on the metric).
  3. Check whether any baseline or metric-aligned alternative is scored on the same CV split for comparison, and whether the reported CV number is compared against a plausible target level for that metric.
  4. Trace which script produced the delivered file and confirm it corresponds to the best-validated configuration, and that the predicted label distribution is compared to the training label distribution.
Discriminator
A real violation is when all candidates share the same metric-mismatched objective/decoding and no baseline or alternative decoding is ever measured; it is fine if the agent explicitly tested a metric-aligned variant (or thresholded regression) and documented that the classifier scored equal/better on the same folds.
Consequence
CV numbers look "reasonable" in isolation, but the submitted predictions score below the grader's threshold for the stated metric (over-predicting rare extreme classes, poor agreement), so the submission file is marked wrong.
id c53dbd2ad6ed · mined from da-code dacode-ml-competition-006@s8
raw text (what the judge reads)
### Ordinal / metric-specific objective replaced by generic classification, with no baseline comparison
- **Applies when**: `task` -- the evaluation metric is stated explicitly (e.g. rank/ordinal agreement, weighted error, custom loss) and the scripts fit off-the-shelf models whose default objective is something else (plain multiclass accuracy, unweighted log-loss).
- **Pattern**: The attempt uses the stated metric only as a CV scorer while every candidate model optimizes an unrelated objective; it further distorts the target distribution (e.g. rebalancing weights, resampling) even though the metric rewards matching the ordered/marginal structure of the labels. No metric-appropriate alternative (ordinal/regression + tuned rounding thresholds, expected-value decoding, cost-sensitive decision rule) and no trivial baseline (constant/majority or simple rounded regression) are evaluated, so there is nothing to show the reported approach is better than a naive one; the final submission is simply whatever the last script wrote.
- **Detection procedure**:
  1. Read the task/README and note the exact scoring metric and how predictions are decoded into the required label space.
  2. In the scripts, check what loss each candidate model actually minimizes and how class probabilities are converted to final labels (argmax vs. thresholding/rounding tuned on the metric).
  3. Check whether any baseline or metric-aligned alternative is scored on the same CV split for comparison, and whether the reported CV number is compared against a plausible target level for that metric.
  4. Trace which script produced the delivered file and confirm it corresponds to the best-validated configuration, and that the predicted label distribution is compared to the training label distribution.
- **Discriminator**: A real violation is when *all* candidates share the same metric-mismatched objective/decoding and no baseline or alternative decoding is ever measured; it is fine if the agent explicitly tested a metric-aligned variant (or thresholded regression) and documented that the classifier scored equal/better on the same folds.
- **Consequence**: CV numbers look "reasonable" in isolation, but the submitted predictions score below the grader's threshold for the stated metric (over-predicting rare extreme classes, poor agreement), so the submission file is marked wrong.
457Missingness-based grouping defined inconsistently with the raw data (partition not validated)taskinfiagent-dabench
Applies when
task -- the task asks to split rows into "missing" vs "non-missing" groups on some column and compare a statistic (mean, test) between the two groups.
Pattern
The attempt relies on the loader's default null handling (isna() after a plain read) without checking how the raw file encodes absence — empty strings, "NA", "None", "-", 0, whitespace, or a placeholder token — so rows land in the wrong group; it also silently drops rows with missing values in the measured column, or filters/deduplicates the frame, so the two groups no longer partition the full dataset, and the reported group means drift from the true ones.
Detection procedure
  1. Read the task to confirm the exact grouping rule ("null" vs "not null") and which column supplies the compared values.
  2. In the script, find where the mask is built and where the data is loaded: check for explicit handling of alternate missing encodings (na_values, .str.strip(), == ""), dtype coercion of the value column, and any earlier dropna, subsetting, or merge that changes row count.
  3. Require the script to print len(df), mask.sum(), (~mask).sum() (and the unique raw values / dtype of the grouping column) and check that the two counts sum to the full row count; also require the value column's NaN count to be reported.
  4. Compare the reported group means against these counts for plausibility (e.g., a large group with a mean far from the overall mean) and confirm the answer reports the requested means and p-value at the required rounding.
Discriminator
A real violation is a script whose group sizes are never printed, or whose printed sizes don't sum to the raw row count, or which uses only default NaN detection on a column that in the raw file contains placeholder strings/whitespace. It is not a violation if the script explicitly inspects the raw column's distinct values, justifies the missingness definition, and shows the two group sizes summing to the total — even if it then legitimately excludes rows with non-numeric measurement values, provided that exclusion is stated.
Consequence
Both group means (and the test statistic) are computed on slightly wrong row sets, so the reported numbers differ from the reference at the required two-decimal precision and every value check fails, even though the qualitative conclusion (significant difference) looks right.
id 314c459f8d75 · mined from infiagent-dabench dabench-297@s8
raw text (what the judge reads)
### Missingness-based grouping defined inconsistently with the raw data (partition not validated)

- **Applies when**: `task` -- the task asks to split rows into "missing" vs "non-missing" groups on some column and compare a statistic (mean, test) between the two groups.
- **Pattern**: The attempt relies on the loader's default null handling (`isna()` after a plain read) without checking how the raw file encodes absence — empty strings, `"NA"`, `"None"`, `"-"`, `0`, whitespace, or a placeholder token — so rows land in the wrong group; it also silently drops rows with missing values in the measured column, or filters/deduplicates the frame, so the two groups no longer partition the full dataset, and the reported group means drift from the true ones.
- **Detection procedure**:
  1. Read the task to confirm the exact grouping rule ("null" vs "not null") and which column supplies the compared values.
  2. In the script, find where the mask is built and where the data is loaded: check for explicit handling of alternate missing encodings (`na_values`, `.str.strip()`, `== ""`), dtype coercion of the value column, and any earlier `dropna`, subsetting, or merge that changes row count.
  3. Require the script to print `len(df)`, `mask.sum()`, `(~mask).sum()` (and the unique raw values / dtype of the grouping column) and check that the two counts sum to the full row count; also require the value column's NaN count to be reported.
  4. Compare the reported group means against these counts for plausibility (e.g., a large group with a mean far from the overall mean) and confirm the answer reports the requested means and p-value at the required rounding.
- **Discriminator**: A real violation is a script whose group sizes are never printed, or whose printed sizes don't sum to the raw row count, or which uses only default NaN detection on a column that in the raw file contains placeholder strings/whitespace. It is *not* a violation if the script explicitly inspects the raw column's distinct values, justifies the missingness definition, and shows the two group sizes summing to the total — even if it then legitimately excludes rows with non-numeric measurement values, provided that exclusion is stated.
- **Consequence**: Both group means (and the test statistic) are computed on slightly wrong row sets, so the reported numbers differ from the reference at the required two-decimal precision and every value check fails, even though the qualitative conclusion (significant difference) looks right.
458Ignoring the provided specification file and inventing undocumented rules/outputstaskda-code
Applies when
task -- the task points to an external instruction/spec artifact (guidance file, README, config) and names concrete deliverables (files, formats, sizes), and the script encodes analysis choices as hard-coded logic.
Pattern
The script never loads or quotes the spec; instead the agent guesses the filtering thresholds, category definitions, and ordering from intuition, writes them into docstrings as if they were the spec, and emits only the one output it happened to think of, silently dropping the other required serialized artifacts (numeric result dump, plot-data JSON, etc.).
Detection procedure
  1. Read the task and list every referenced spec source and every required deliverable (file name, extension, figure size, colors, labels, ordering, rounding).
  2. Grep the script for a read of the spec source and for a write of each deliverable; check that each rule in the code (threshold values, keyword-to-category mappings, tie-breaking "first match wins" logic) can be traced to a quoted line of the spec rather than to the agent's own comment.
  3. Compare the agent's final answer to the deliverable list: does it report only the artifacts it produced, or does it account for all of them?
  4. Check whether derived columns computed in the script (unused intermediate variables) suggest the agent copied a generic template instead of following the actual instructions.
Discriminator
A real violation is when the categorization/filter constants or the output set exist only in the script's own comments and no spec text is ever read or cited; it is fine if the script hard-codes rules but the reviewer can match each constant and each output file to explicit spec text (a script may legitimately inline spec values rather than parse the file).
Consequence
Graders comparing every expected artifact fail on the missing files outright, and even the produced figure/counts differ because the invented thresholds and keyword mappings yield a different subset and different per-category tallies than the specified ones.
id 355a2bf4599a · mined from da-code dacode-plot-pie-005@s8
raw text (what the judge reads)
### Ignoring the provided specification file and inventing undocumented rules/outputs
- **Applies when**: `task` -- the task points to an external instruction/spec artifact (guidance file, README, config) and names concrete deliverables (files, formats, sizes), and the script encodes analysis choices as hard-coded logic.
- **Pattern**: The script never loads or quotes the spec; instead the agent guesses the filtering thresholds, category definitions, and ordering from intuition, writes them into docstrings as if they were the spec, and emits only the one output it happened to think of, silently dropping the other required serialized artifacts (numeric result dump, plot-data JSON, etc.).
- **Detection procedure**:
  1. Read the task and list every referenced spec source and every required deliverable (file name, extension, figure size, colors, labels, ordering, rounding).
  2. Grep the script for a read of the spec source and for a write of each deliverable; check that each rule in the code (threshold values, keyword-to-category mappings, tie-breaking "first match wins" logic) can be traced to a quoted line of the spec rather than to the agent's own comment.
  3. Compare the agent's final answer to the deliverable list: does it report only the artifacts it produced, or does it account for all of them?
  4. Check whether derived columns computed in the script (unused intermediate variables) suggest the agent copied a generic template instead of following the actual instructions.
- **Discriminator**: A real violation is when the categorization/filter constants or the output set exist only in the script's own comments and no spec text is ever read or cited; it is fine if the script hard-codes rules but the reviewer can match each constant and each output file to explicit spec text (a script may legitimately inline spec values rather than parse the file).
- **Consequence**: Graders comparing every expected artifact fail on the missing files outright, and even the produced figure/counts differ because the invented thresholds and keyword mappings yield a different subset and different per-category tallies than the specified ones.
459Missing-value handling deviates from the prescribed imputation (silent row drops / non-numeric coercion / target excluded)taskinfiagent-dabench
Applies when
task -- The task explicitly prescribes how to treat missing values (e.g., "impute the listed columns with their column means") for a set of feature and target columns before fitting/evaluating a model.
Pattern
The script does not implement exactly that rule: it drops rows with NAs (via dropna(), or implicitly because a column stored as text/currency/thousands-separated strings becomes NaN or object dtype and gets discarded), imputes only the predictors and not the target, imputes with a different statistic (median/zero/forward-fill), or fits the imputer on a subset (train only) when the task said to impute the column means of the full data. No sanity check on row counts/dtypes/value ranges after cleaning is performed, and no reproducible script is retained.
Detection procedure
  1. From the task, list every column named in the missing-value instruction, the statistic to use, and whether it includes the response variable.
  2. In the script, trace each of those columns from load to model input: check dtype conversion (to_numeric/string stripping) happens before the mean is computed, that fillna(mean) is applied to each listed column including the target, and that no dropna, boolean filter, or merge silently removes rows.
  3. Compare len(df) (or X.shape) after cleaning to the raw row count and to the train/test sizes implied by the stated split ratio; flag any shrinkage or any column left as object dtype.
  4. Check the reported metric's magnitude against a trivial baseline (e.g., variance of the target / MSE of predicting the target mean); an MSE near or above that baseline signals a broken feature/row pipeline.
Discriminator
A real violation is any deviation that changes which rows or values enter the fit (dropped rows, wrong statistic, unconverted numeric strings, target left unimputed). Not a violation: cosmetic differences that leave the same numeric matrix — e.g., using SimpleImputer(strategy="mean") instead of fillna(df.mean()), or reordering columns — provided row count and values match.
Consequence
The model is trained/tested on a different (smaller or differently-valued) sample than specified, so the reported error is off by a large factor from the reference value and the exact-match check on the metric fails.
id bdcaf084167b · mined from infiagent-dabench dabench-432@s8
raw text (what the judge reads)
### Missing-value handling deviates from the prescribed imputation (silent row drops / non-numeric coercion / target excluded)
- **Applies when**: `task` -- The task explicitly prescribes how to treat missing values (e.g., "impute the listed columns with their column means") for a set of feature and target columns before fitting/evaluating a model.
- **Pattern**: The script does not implement exactly that rule: it drops rows with NAs (via `dropna()`, or implicitly because a column stored as text/currency/thousands-separated strings becomes NaN or object dtype and gets discarded), imputes only the predictors and not the target, imputes with a different statistic (median/zero/forward-fill), or fits the imputer on a subset (train only) when the task said to impute the column means of the full data. No sanity check on row counts/dtypes/value ranges after cleaning is performed, and no reproducible script is retained.
- **Detection procedure**:
  1. From the task, list every column named in the missing-value instruction, the statistic to use, and whether it includes the response variable.
  2. In the script, trace each of those columns from load to model input: check dtype conversion (`to_numeric`/string stripping) happens *before* the mean is computed, that `fillna(mean)` is applied to each listed column including the target, and that no `dropna`, boolean filter, or merge silently removes rows.
  3. Compare `len(df)` (or `X.shape`) after cleaning to the raw row count and to the train/test sizes implied by the stated split ratio; flag any shrinkage or any column left as object dtype.
  4. Check the reported metric's magnitude against a trivial baseline (e.g., variance of the target / MSE of predicting the target mean); an MSE near or above that baseline signals a broken feature/row pipeline.
- **Discriminator**: A real violation is any deviation that changes which rows or values enter the fit (dropped rows, wrong statistic, unconverted numeric strings, target left unimputed). Not a violation: cosmetic differences that leave the same numeric matrix — e.g., using `SimpleImputer(strategy="mean")` instead of `fillna(df.mean())`, or reordering columns — provided row count and values match.
- **Consequence**: The model is trained/tested on a different (smaller or differently-valued) sample than specified, so the reported error is off by a large factor from the reference value and the exact-match check on the metric fails.
460Unverified row-ordering assumption in lag/sequence-dependent computationstaskinfiagent-dabench
Applies when
task -- the requested statistic depends on the order of rows (differences from a previous row, lags, cumulative or rolling quantities, time-based splits) and the script re-sorts or reverses the data before computing it.
Pattern
The script infers the data's ordering from a glance ("appears to be newest-first") and flips/reorders rows without programmatically checking the ordering key, so every lagged difference is computed against the wrong neighbour and sign-sensitive results (e.g. a mean change) come out inverted.
Detection procedure
  1. In the task, confirm the statistic is defined relative to a previous row, i.e. it is order- and sign-sensitive.
  2. In the script, locate any reordering step ([::-1], sort_values, reset_index) and check whether the ordering key was parsed to a proper datetime/numeric type and verified monotonic (e.g. an assert or an explicit is_monotonic_increasing/min-max date print) rather than assumed from visual inspection.
  3. Check that the reordering is done by explicitly sorting on the ordering key (ascending) instead of blindly reversing, and that the index/key is not dropped before the check.
  4. Inspect the reported answer for a sanity check on sign/magnitude (e.g. does the mean's sign agree with the overall first-to-last change in the underlying series?); absence of such a check plus an unverified flip is a violation.
Discriminator
Fine if the script parses and asserts/prints that the ordering key is ascending before computing lags (or sorts by that key explicitly); a violation if ordering is decided by a comment/eyeball and applied via a reversal, or if rows are reindexed with no ordering validation at all. Also not a violation when the statistic is order-independent (plain mean, count, sum over all rows).
Consequence
Lagged differences are taken in the wrong direction, so order-sensitive statistics flip sign (and shift slightly in magnitude), and the graded mean fails while the near-symmetric dispersion value only appears close, producing 0/2 correct checks.
id 9e35922ebe83 · mined from infiagent-dabench dabench-75@s8
raw text (what the judge reads)
### Unverified row-ordering assumption in lag/sequence-dependent computations
- **Applies when**: `task` -- the requested statistic depends on the order of rows (differences from a previous row, lags, cumulative or rolling quantities, time-based splits) and the script re-sorts or reverses the data before computing it.
- **Pattern**: The script infers the data's ordering from a glance ("appears to be newest-first") and flips/reorders rows without programmatically checking the ordering key, so every lagged difference is computed against the wrong neighbour and sign-sensitive results (e.g. a mean change) come out inverted.
- **Detection procedure**:
  1. In the task, confirm the statistic is defined relative to a *previous* row, i.e. it is order- and sign-sensitive.
  2. In the script, locate any reordering step (`[::-1]`, `sort_values`, `reset_index`) and check whether the ordering key was parsed to a proper datetime/numeric type and verified monotonic (e.g. an assert or an explicit `is_monotonic_increasing`/min-max date print) rather than assumed from visual inspection.
  3. Check that the reordering is done by explicitly sorting on the ordering key (ascending) instead of blindly reversing, and that the index/key is not dropped before the check.
  4. Inspect the reported answer for a sanity check on sign/magnitude (e.g. does the mean's sign agree with the overall first-to-last change in the underlying series?); absence of such a check plus an unverified flip is a violation.
- **Discriminator**: Fine if the script parses and asserts/prints that the ordering key is ascending before computing lags (or sorts by that key explicitly); a violation if ordering is decided by a comment/eyeball and applied via a reversal, or if rows are reindexed with no ordering validation at all. Also not a violation when the statistic is order-independent (plain mean, count, sum over all rows).
- **Consequence**: Lagged differences are taken in the wrong direction, so order-sensitive statistics flip sign (and shift slightly in magnitude), and the graded mean fails while the near-symmetric dispersion value only appears close, producing 0/2 correct checks.
461Output schema/format not verified against the task's prescribed field settaskda-code
Applies when
task -- the task specifies a result file with an exact set of fields/columns (often defined in a referenced instructions/tips file) and a fixed number of rows.
Pattern
The attempt computes a defensible statistic but writes the file with an ad-hoc schema — extra or renamed columns, different column order, values in a different form (e.g. unrounded/zero-padded numbers, free-text labels instead of the prescribed vocabulary) — and never checks the written file against the specification actually stated in the task or the referenced instructions.
Detection procedure
  1. From the task text (and any referenced instructions file the task says to follow), list the required fields, their names/order, the expected row count, and any value conventions (rounding, allowed label strings, units).
  2. In the scripts, locate the code that constructs and writes the result file; extract the literal header/keys, ordering, and formatting applied to each value.
  3. Compare item-by-item with the list from step 1; flag any extra field, missing field, renamed/reordered field, or value rendered in a form the spec does not sanction.
  4. Confirm the script re-reads the written file (or prints it) to assert shape (rows × columns) and content before declaring completion; absence of this check is itself a flag.
Discriminator
A real violation is a mismatch in the contract (field set, names, row count, value encoding) that a strict comparator would reject; a look-alike that is fine is a correct schema whose numeric value differs slightly due to a legitimate methodological choice, or cosmetic whitespace that the spec clearly tolerates. Adding an "informative" extra column is a violation, not a bonus.
Consequence
The grader compares the produced file to the expected one field-by-field and marks it WRONG/MISSING even though the underlying statistic may be right, yielding 0 checks passed.
id 549a6e35a725 · mined from da-code dacode-data-sa-004@s8
raw text (what the judge reads)
### Output schema/format not verified against the task's prescribed field set
- **Applies when**: `task` -- the task specifies a result file with an exact set of fields/columns (often defined in a referenced instructions/tips file) and a fixed number of rows.
- **Pattern**: The attempt computes a defensible statistic but writes the file with an ad-hoc schema — extra or renamed columns, different column order, values in a different form (e.g. unrounded/zero-padded numbers, free-text labels instead of the prescribed vocabulary) — and never checks the written file against the specification actually stated in the task or the referenced instructions.
- **Detection procedure**:
  1. From the task text (and any referenced instructions file the task says to follow), list the required fields, their names/order, the expected row count, and any value conventions (rounding, allowed label strings, units).
  2. In the scripts, locate the code that constructs and writes the result file; extract the literal header/keys, ordering, and formatting applied to each value.
  3. Compare item-by-item with the list from step 1; flag any extra field, missing field, renamed/reordered field, or value rendered in a form the spec does not sanction.
  4. Confirm the script re-reads the written file (or prints it) to assert shape (rows × columns) and content before declaring completion; absence of this check is itself a flag.
- **Discriminator**: A real violation is a mismatch in the *contract* (field set, names, row count, value encoding) that a strict comparator would reject; a look-alike that is fine is a correct schema whose numeric value differs slightly due to a legitimate methodological choice, or cosmetic whitespace that the spec clearly tolerates. Adding an "informative" extra column is a violation, not a bonus.
- **Consequence**: The grader compares the produced file to the expected one field-by-field and marks it WRONG/MISSING even though the underlying statistic may be right, yielding 0 checks passed.
462Ambiguous "normalize" implemented as z-scoring, yielding degenerate reported statisticstaskinfiagent-dabench
Applies when
task -- The task asks to rescale/normalize a set of numeric columns and then report a summary statistic (e.g., mean) of those columns.
Pattern
The script applies standardization (subtract mean, divide by std) rather than min–max scaling to [0,1], so every reported mean is mathematically forced to 0 (printed as 0.0/-0.0); the agent submits these constant values without noticing they carry no information about the data.
Detection procedure
  1. Read the task wording for the rescaling requirement and note whether the reported statistic would be trivially determined by the chosen transform.
  2. Inspect the script for the scaler/formula used (StandardScaler, (x-mean)/std vs MinMaxScaler, (x-min)/(max-min)).
  3. Check the reported values: if all rescaled columns report exactly 0/-0 (or 1.0 for a std), the statistic is an artifact of the transform, not a data property — flag it.
  4. Confirm the alternative interpretation (min–max to [0,1]) yields distinct, in-range (0–1) values per column; require the agent to justify or default to the interpretation that produces informative, column-specific numbers.
Discriminator
A real violation is when the requested statistic becomes constant/degenerate by construction under the chosen transform (all zeros), or when the transform contradicts an explicit range requirement. It is not a violation if the task explicitly names standardization, or if the reported statistic still varies meaningfully per column (e.g., reporting min/max or medians after z-scoring).
Consequence
Every scaled column's reported mean equals 0.0 and fails the expected non-trivial values (e.g., ~0.19–0.46), so only the untransformed/binary columns match and the answer is graded largely incorrect.
id 612da95ff6a1 · mined from infiagent-dabench dabench-28@s8
raw text (what the judge reads)
### Ambiguous "normalize" implemented as z-scoring, yielding degenerate reported statistics
- **Applies when**: `task` -- The task asks to rescale/normalize a set of numeric columns and then report a summary statistic (e.g., mean) of those columns.
- **Pattern**: The script applies standardization (subtract mean, divide by std) rather than min–max scaling to [0,1], so every reported mean is mathematically forced to 0 (printed as `0.0`/`-0.0`); the agent submits these constant values without noticing they carry no information about the data.
- **Detection procedure**:
  1. Read the task wording for the rescaling requirement and note whether the reported statistic would be trivially determined by the chosen transform.
  2. Inspect the script for the scaler/formula used (`StandardScaler`, `(x-mean)/std` vs `MinMaxScaler`, `(x-min)/(max-min)`).
  3. Check the reported values: if all rescaled columns report exactly 0/-0 (or 1.0 for a std), the statistic is an artifact of the transform, not a data property — flag it.
  4. Confirm the alternative interpretation (min–max to [0,1]) yields distinct, in-range (0–1) values per column; require the agent to justify or default to the interpretation that produces informative, column-specific numbers.
- **Discriminator**: A real violation is when the requested statistic becomes constant/degenerate by construction under the chosen transform (all zeros), or when the transform contradicts an explicit range requirement. It is *not* a violation if the task explicitly names standardization, or if the reported statistic still varies meaningfully per column (e.g., reporting min/max or medians after z-scoring).
- **Consequence**: Every scaled column's reported mean equals 0.0 and fails the expected non-trivial values (e.g., ~0.19–0.46), so only the untransformed/binary columns match and the answer is graded largely incorrect.
463Unverified row set / value coercion when computing a summary statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic (correlation, mean, test statistic, p-value) over two or more columns of a raw table, reported to a fixed rounding precision.
Pattern
The attempt loads the file and calls a one-liner statistic function without ever inspecting or reporting how many rows actually entered the computation — no check for missing/sentinel values, non-numeric strings coerced or silently dropped, duplicate/extra header rows, or an unintended filter/subset — so the statistic is computed on a slightly different population than the intended full data and lands just outside the rounding tolerance. No script or intermediate output is retained to justify the number.
Detection procedure
  1. Read the task to see which rows are in scope (usually all rows unless a filter is stated) and to what precision the answer must be reported.
  2. Read the script: does it print the shape/dtype of the loaded table, the count of non-null pairs actually used, and the raw unrounded statistic? Does it explain any dropping/coercion (dropna, errors='coerce', delimiter/header assumptions) rather than relying on library defaults?
  3. Check whether the reported value was cross-validated by a second, independent computation (e.g., a different library/manual formula) or at least sanity-checked against N and the raw unrounded value.
  4. Check the reported numbers respect the demanded precision and range (e.g., a probability written as 0.0 rather than with the requested number of decimals).
Discriminator
A real violation is an attempt that produces the number with no evidence of the effective sample size, parsing choices, or unrounded value — any of which could shift the last reported digit. It is not a violation if the script prints N/shape, shows the full-precision statistic, and its row-handling is either a no-op (no nulls, all numeric) or explicitly justified by the task.
Consequence
The reported statistic differs from ground truth in the last required decimal (e.g., 0.53 vs 0.54), so the exact-match check on the numeric field fails even though the qualitative conclusion is right, and secondary fields may also violate the required formatting.
id a874ec3188a9 · mined from infiagent-dabench dabench-300@s8
raw text (what the judge reads)
### Unverified row set / value coercion when computing a summary statistic
- **Applies when**: `task` -- the task asks for a single numeric statistic (correlation, mean, test statistic, p-value) over two or more columns of a raw table, reported to a fixed rounding precision.
- **Pattern**: The attempt loads the file and calls a one-liner statistic function without ever inspecting or reporting how many rows actually entered the computation — no check for missing/sentinel values, non-numeric strings coerced or silently dropped, duplicate/extra header rows, or an unintended filter/subset — so the statistic is computed on a slightly different population than the intended full data and lands just outside the rounding tolerance. No script or intermediate output is retained to justify the number.
- **Detection procedure**:
  1. Read the task to see which rows are in scope (usually all rows unless a filter is stated) and to what precision the answer must be reported.
  2. Read the script: does it print the shape/dtype of the loaded table, the count of non-null pairs actually used, and the raw unrounded statistic? Does it explain any dropping/coercion (`dropna`, `errors='coerce'`, delimiter/header assumptions) rather than relying on library defaults?
  3. Check whether the reported value was cross-validated by a second, independent computation (e.g., a different library/manual formula) or at least sanity-checked against N and the raw unrounded value.
  4. Check the reported numbers respect the demanded precision and range (e.g., a probability written as `0.0` rather than with the requested number of decimals).
- **Discriminator**: A real violation is an attempt that produces the number with no evidence of the effective sample size, parsing choices, or unrounded value — any of which could shift the last reported digit. It is *not* a violation if the script prints N/shape, shows the full-precision statistic, and its row-handling is either a no-op (no nulls, all numeric) or explicitly justified by the task.
- **Consequence**: The reported statistic differs from ground truth in the last required decimal (e.g., 0.53 vs 0.54), so the exact-match check on the numeric field fails even though the qualitative conclusion is right, and secondary fields may also violate the required formatting.
464Prediction file accepted on self-reported metrics without checking the output distribution against the training targettaskda-code
Applies when
task -- the deliverable is a file of per-row predictions for a held-out set, and the scripts report only internal validation scores (R², MAE) plus a summary of the predictions.
Pattern
The attempt trusts a high in-sample/holdout score and never compares the produced prediction distribution (count, min, max, mean, median, quantiles) to the distribution of the target in the training data or to domain plausibility; broken encodings, feature/column misalignment between train and test, or a mis-scaled target then slip through, producing predictions that are extreme at the tails or shifted in central tendency while the reported metric still looks excellent.
Detection procedure
  1. Read the task/README to note what the predicted quantity is and what a plausible range and typical central value look like from the training target.
  2. In the scripts, check whether the same fitted preprocessing (encoders, imputers, column order, dtype handling) is applied to the prediction set and whether unseen categories/missing values are handled — and whether any post-prediction sanity check exists.
  3. Compare the answer's reported prediction summary (row count, min, max, mean, median) with the training target's summary; flag if row count differs from the prediction set, if the min/max fall far outside the observed target range, or if mean/median are shifted materially relative to training.
  4. Cross-check the claimed validation metric against the summary: a near-perfect score coexisting with an implausible prediction spread or with feature importances that contradict domain sense is evidence the pipeline, not the model, is wrong.
Discriminator
A real violation is when no such comparison is made and the reported summary is inconsistent with the training target (wrong row count, tail values orders of magnitude outside observed prices, shifted center) or when train/test preprocessing is fit separately; a look-alike that is fine is a prediction set whose summary closely tracks the training target's distribution with only mild tail shrinkage (expected from regression averaging), with preprocessing fit on train and merely applied to test.
Consequence
The saved prediction file fails the grader's value/format check — scores far worse than the self-reported metric because the predicted values are misaligned with the true targets, even though the file has the right column name and no missing values.
id 43b5972cc0a5 · mined from da-code dacode-ml-regression-014@s8
raw text (what the judge reads)
### Prediction file accepted on self-reported metrics without checking the output distribution against the training target
- **Applies when**: `task` -- the deliverable is a file of per-row predictions for a held-out set, and the scripts report only internal validation scores (R², MAE) plus a summary of the predictions.
- **Pattern**: The attempt trusts a high in-sample/holdout score and never compares the produced prediction distribution (count, min, max, mean, median, quantiles) to the distribution of the target in the training data or to domain plausibility; broken encodings, feature/column misalignment between train and test, or a mis-scaled target then slip through, producing predictions that are extreme at the tails or shifted in central tendency while the reported metric still looks excellent.
- **Detection procedure**:
  1. Read the task/README to note what the predicted quantity is and what a plausible range and typical central value look like from the training target.
  2. In the scripts, check whether the same fitted preprocessing (encoders, imputers, column order, dtype handling) is applied to the prediction set and whether unseen categories/missing values are handled — and whether any post-prediction sanity check exists.
  3. Compare the answer's reported prediction summary (row count, min, max, mean, median) with the training target's summary; flag if row count differs from the prediction set, if the min/max fall far outside the observed target range, or if mean/median are shifted materially relative to training.
  4. Cross-check the claimed validation metric against the summary: a near-perfect score coexisting with an implausible prediction spread or with feature importances that contradict domain sense is evidence the pipeline, not the model, is wrong.
- **Discriminator**: A real violation is when no such comparison is made *and* the reported summary is inconsistent with the training target (wrong row count, tail values orders of magnitude outside observed prices, shifted center) or when train/test preprocessing is fit separately; a look-alike that is fine is a prediction set whose summary closely tracks the training target's distribution with only mild tail shrinkage (expected from regression averaging), with preprocessing fit on train and merely applied to test.
- **Consequence**: The saved prediction file fails the grader's value/format check — scores far worse than the self-reported metric because the predicted values are misaligned with the true targets, even though the file has the right column name and no missing values.
465Output artifacts written outside the expected/canonical working directorytaskinfiagent-dabench
Applies when
task -- the answer must include file paths to artifacts the agent creates (cleaned/normalized CSVs, model files, plots) and the task's input data lives in a specific directory.
Pattern
The agent writes outputs to its own home/current directory (or a temp dir) rather than alongside the provided input data, and reports those paths; the values themselves may be right but the path string does not match the expected location.
Detection procedure
1. From the task, note where the input file was supplied from (e.g., the directory of the given dataset path) and any stated output-location convention. 2. In the scripts, find every write call (to_csv, savefig, dump) and record the literal output directory. 3. Compare that directory with the input data's directory / stated convention; also check the reported path in the answer is the exact absolute path actually written. 4. Flag if the output directory differs from the input data directory without an explicit instruction to use another location.
Discriminator
A real violation is writing to an unrelated directory when the task provides a canonical data directory (outputs should default there). It is fine if the task explicitly names an output directory and the agent used it, or if the agent wrote into the same directory as the source data under a different but stated filename.
Consequence
The path field of the answer mismatches the expected string, so that whole answer tuple is graded wrong even though the computed statistics are correct.
id 5490df68f4d2 · mined from infiagent-dabench dabench-743@s8
raw text (what the judge reads)
### Output artifacts written outside the expected/canonical working directory
- **Applies when**: `task` -- the answer must include file paths to artifacts the agent creates (cleaned/normalized CSVs, model files, plots) and the task's input data lives in a specific directory.
- **Pattern**: The agent writes outputs to its own home/current directory (or a temp dir) rather than alongside the provided input data, and reports those paths; the values themselves may be right but the path string does not match the expected location.
- **Detection procedure**: 1. From the task, note where the input file was supplied from (e.g., the directory of the given dataset path) and any stated output-location convention. 2. In the scripts, find every write call (`to_csv`, `savefig`, `dump`) and record the literal output directory. 3. Compare that directory with the input data's directory / stated convention; also check the reported path in the answer is the exact absolute path actually written. 4. Flag if the output directory differs from the input data directory without an explicit instruction to use another location.
- **Discriminator**: A real violation is writing to an unrelated directory when the task provides a canonical data directory (outputs should default there). It is fine if the task explicitly names an output directory and the agent used it, or if the agent wrote into the same directory as the source data under a different but stated filename.
- **Consequence**: The path field of the answer mismatches the expected string, so that whole answer tuple is graded wrong even though the computed statistics are correct.
466Incomplete production of the required output artifacts specified by the task or its config filetaskda-code
Applies when
task -- the task points to an external specification (config/spec file, e.g. YAML/JSON guidelines) and/or asks for saved deliverables, and the scripts are expected to write files to disk.
Pattern
The attempt reads the spec only for cosmetic parameters, produces the one artifact explicitly named in the prompt (e.g. the image), and reports numbers in prose, while silently skipping the other artifacts the spec/harness requires (serialized plot metadata, saved arrays/tables of the underlying values) or writing them with wrong names, locations, keys, or shapes.
Detection procedure
  1. Read the task text and the referenced spec file, and enumerate every output file, filename, format, and field/key it mentions (not just the one repeated in the prompt sentence).
  2. Grep the scripts for a save/serialize call for each enumerated artifact (savefig, save, to_csv, json.dump, ...) and check the exact path, name, and dumped structure/dtype against the spec.
  3. Confirm the values written into each artifact are the requested quantities (e.g. the plotted category shares/counts) and in the requested units/order/rounding, not intermediate diagnostics.
  4. Compare with the answer: if the answer only narrates numbers or mentions one file while the spec listed more, flag it.
Discriminator
A real violation is a missing/misnamed/mis-structured required artifact, or one whose contents don't match the spec's fields. It is not a violation if all specified artifacts exist under the specified names and the spec genuinely mentions no others — extra optional debug files or prose narration alongside complete artifacts are fine.
Consequence
File-level checks report the absent or mismatched artifacts as WRONG/MISSING, so the submission scores zero even when the analytical conclusion (the selected group) is correct.
id d0a4422ff84a · mined from da-code dacode-plot-pie-008@s8
raw text (what the judge reads)
### Incomplete production of the required output artifacts specified by the task or its config file
- **Applies when**: `task` -- the task points to an external specification (config/spec file, e.g. YAML/JSON guidelines) and/or asks for saved deliverables, and the scripts are expected to write files to disk.
- **Pattern**: The attempt reads the spec only for cosmetic parameters, produces the one artifact explicitly named in the prompt (e.g. the image), and reports numbers in prose, while silently skipping the other artifacts the spec/harness requires (serialized plot metadata, saved arrays/tables of the underlying values) or writing them with wrong names, locations, keys, or shapes.
- **Detection procedure**:
  1. Read the task text *and* the referenced spec file, and enumerate every output file, filename, format, and field/key it mentions (not just the one repeated in the prompt sentence).
  2. Grep the scripts for a save/serialize call for each enumerated artifact (`savefig`, `save`, `to_csv`, `json.dump`, ...) and check the exact path, name, and dumped structure/dtype against the spec.
  3. Confirm the values written into each artifact are the *requested* quantities (e.g. the plotted category shares/counts) and in the requested units/order/rounding, not intermediate diagnostics.
  4. Compare with the answer: if the answer only narrates numbers or mentions one file while the spec listed more, flag it.
- **Discriminator**: A real violation is a missing/misnamed/mis-structured required artifact, or one whose contents don't match the spec's fields. It is *not* a violation if all specified artifacts exist under the specified names and the spec genuinely mentions no others — extra optional debug files or prose narration alongside complete artifacts are fine.
- **Consequence**: File-level checks report the absent or mismatched artifacts as WRONG/MISSING, so the submission scores zero even when the analytical conclusion (the selected group) is correct.
467Missing-value handling silently changes the row set before splittingtaskinfiagent-dabench
Applies when
task -- the task specifies an exact split procedure (fixed proportion and random seed) and a metric on the held-out part, while the raw feature set contains columns with nulls that the model cannot consume.
Pattern
The script drops rows (or whole columns) containing nulls — via dropna(), boolean filtering, or implicit dropping inside a merge/encoding step — instead of imputing, so the dataset passed to the split has fewer rows than the source. With a fixed seed, a different row count yields a different train/test partition and a different metric, even though every other step looks correct. A variant is the reverse: dropping a feature the task explicitly listed.
Detection procedure
  1. From the task, list the exact features/rows that must be used and note that the split is seed-fixed (so results depend on the exact input frame).
  2. In the script, find every step between loading and train_test_split and check whether any removes rows or listed columns (dropna, notnull masks, drop, filters, non-inner-safe merges); confirm the row count entering the split equals the raw row count.
  3. Check that nulls are instead handled by an explicit imputation (mean/median/mode or a category) applied consistently to all required features, and that all listed features are still present.
  4. Verify the reported metric is computed on the test partition of that same full frame, and that a continuous-output model used for a binary target is thresholded (e.g., ≥0.5) before scoring accuracy.
Discriminator
A real violation is any unrequested reduction of the row/feature set relative to the source data before the seed-fixed split. It is fine if the task itself asks for filtering, or if the removal is provably a no-op (verified zero nulls / duplicate columns) — the reviewer should demand a printed shape/null-count check proving this.
Consequence
The accuracy is computed on a different held-out subset than the reference, so the reported figure is off by a few points (e.g., 0.76 vs. 0.78) and the exact-match grader marks it wrong with no other visible error.
id d0cf3f90d8b0 · mined from infiagent-dabench dabench-7@s8
raw text (what the judge reads)
### Missing-value handling silently changes the row set before splitting
- **Applies when**: `task` -- the task specifies an exact split procedure (fixed proportion and random seed) and a metric on the held-out part, while the raw feature set contains columns with nulls that the model cannot consume.
- **Pattern**: The script drops rows (or whole columns) containing nulls — via `dropna()`, boolean filtering, or implicit dropping inside a merge/encoding step — instead of imputing, so the dataset passed to the split has fewer rows than the source. With a fixed seed, a different row count yields a different train/test partition and a different metric, even though every other step looks correct. A variant is the reverse: dropping a feature the task explicitly listed.
- **Detection procedure**:
  1. From the task, list the exact features/rows that must be used and note that the split is seed-fixed (so results depend on the exact input frame).
  2. In the script, find every step between loading and `train_test_split` and check whether any removes rows or listed columns (`dropna`, `notnull` masks, `drop`, filters, non-inner-safe merges); confirm the row count entering the split equals the raw row count.
  3. Check that nulls are instead handled by an explicit imputation (mean/median/mode or a category) applied consistently to all required features, and that all listed features are still present.
  4. Verify the reported metric is computed on the test partition of that same full frame, and that a continuous-output model used for a binary target is thresholded (e.g., ≥0.5) before scoring accuracy.
- **Discriminator**: A real violation is any unrequested reduction of the row/feature set relative to the source data before the seed-fixed split. It is fine if the task itself asks for filtering, or if the removal is provably a no-op (verified zero nulls / duplicate columns) — the reviewer should demand a printed shape/null-count check proving this.
- **Consequence**: The accuracy is computed on a different held-out subset than the reference, so the reported figure is off by a few points (e.g., 0.76 vs. 0.78) and the exact-match grader marks it wrong with no other visible error.
468Invented category definitions instead of the ones fixed by the provided output templatetaskda-code
Applies when
task -- the task supplies a pre-existing result file/template whose rows or columns encode the required categories (groups, bins, labels), and the scripts derive those categories themselves from a continuous or raw column.
Pattern
The attempt never opens/inspects the supplied template to learn the exact label set, ordering, and granularity; it hard-codes its own thresholds and names, then overwrites the file with a self-invented schema — often also producing wildly unbalanced group counts that go unquestioned.
Detection procedure
  1. In the task text, note that an existing file must be filled in "adhering strictly to its format"; check whether any script reads that file (or otherwise enumerates its expected labels/rows) before writing to it.
  2. In the scripts, locate where categories are constructed and check whether the boundaries/labels come from the template or domain-standard definitions, versus being invented constants; also check whether an "extra"/"unknown" bucket is added that the template does not contain, and whether the file is written to the same path the grader will read.
  3. In the printed output/answer, sanity-check per-category counts and the mapping of representative records: near-empty or hugely dominant bins, or well-known records landing in an obviously wrong category, indicate mis-specified boundaries or unit mismatch.
  4. Confirm the final written file's header names, row labels, row order, and row count match the template exactly.
Discriminator
A real violation is when the label set/boundaries could only have been guessed (no read of the template, no cited standard) and/or the resulting group sizes are implausible; it is fine if the script loads the template (or reproduces its exact labels/order) and merely recomputes the counts, even if some counts are zero.
Consequence
The output file fails an exact row/label comparison — categories missing, extra, mislabeled, or with counts derived from wrong bins — so the check for that file is scored WRONG/MISSING despite a plausible-looking narrative summary.
id ec154ff1f5cf · mined from da-code dacode-dm-csv-001@s8
raw text (what the judge reads)
### Invented category definitions instead of the ones fixed by the provided output template
- **Applies when**: `task` -- the task supplies a pre-existing result file/template whose rows or columns encode the required categories (groups, bins, labels), and the scripts derive those categories themselves from a continuous or raw column.
- **Pattern**: The attempt never opens/inspects the supplied template to learn the exact label set, ordering, and granularity; it hard-codes its own thresholds and names, then overwrites the file with a self-invented schema — often also producing wildly unbalanced group counts that go unquestioned.
- **Detection procedure**:
  1. In the task text, note that an existing file must be filled in "adhering strictly to its format"; check whether any script reads that file (or otherwise enumerates its expected labels/rows) before writing to it.
  2. In the scripts, locate where categories are constructed and check whether the boundaries/labels come from the template or domain-standard definitions, versus being invented constants; also check whether an "extra"/"unknown" bucket is added that the template does not contain, and whether the file is written to the same path the grader will read.
  3. In the printed output/answer, sanity-check per-category counts and the mapping of representative records: near-empty or hugely dominant bins, or well-known records landing in an obviously wrong category, indicate mis-specified boundaries or unit mismatch.
  4. Confirm the final written file's header names, row labels, row order, and row count match the template exactly.
- **Discriminator**: A real violation is when the label set/boundaries could only have been guessed (no read of the template, no cited standard) and/or the resulting group sizes are implausible; it is fine if the script loads the template (or reproduces its exact labels/order) and merely recomputes the counts, even if some counts are zero.
- **Consequence**: The output file fails an exact row/label comparison — categories missing, extra, mislabeled, or with counts derived from wrong bins — so the check for that file is scored WRONG/MISSING despite a plausible-looking narrative summary.
469Unvalidated entity-level aggregation and threshold-split definitions for derived variablestaskinfiagent-dabench
Applies when
task -- the analysis requires correlating/comparing variables that must first be derived per entity (e.g., a max, a span/duration, a total) from a multi-row-per-entity table, and then splitting the entities by a computed threshold such as a median.
Pattern
The attempt computes the statistic directly on the raw record-level table or with an unstated aggregation/duration convention (counting rows vs. differencing timestamps, inclusive vs. exclusive endpoints, unit of time), and applies the threshold split without stating how ties/boundary values and duplicate entity rows are handled; the numeric result is plausible and the qualitative verdict is right, but the coefficient is off by a few hundredths.
Detection procedure
  1. From the task, list every derived quantity (per-entity aggregate, span/duration, damage total) and the split rule, and note that the answer format demands a specific rounding precision — so small definitional differences change the graded value.
  2. In the scripts, locate the aggregation step: confirm there is an explicit group-by producing exactly one row per entity, an explicit duration formula with stated units and endpoint convention, and an explicit rule for records at/above vs. below the threshold.
  3. Check for a printed sanity block: number of unique entities, number of rows in each split (should sum to the total and be near-equal for a median split), and min/max/range of the derived duration and category values.
  4. If scripts are absent, or any of these definitions/counts is implicit or unprinted, treat the reported coefficient as unverified and flag the attempt.
Discriminator
A real violation is an attempt whose entity-level table, duration convention, or split membership cannot be reconstructed and checked from the code/output; a look-alike that is fine explicitly aggregates to one row per entity, prints group sizes and value ranges, and shows the coefficient is stable under the documented convention (or notes the alternative convention gives the same rounded value).
Consequence
The direction and significance verdict pass, but the rounded correlation coefficient differs from ground truth (e.g., 0.58 vs 0.56), so the exact-value check fails and the overall answer is marked incorrect.
id b0116e377772 · mined from infiagent-dabench dabench-431@s8
raw text (what the judge reads)
### Unvalidated entity-level aggregation and threshold-split definitions for derived variables
- **Applies when**: `task` -- the analysis requires correlating/comparing variables that must first be derived per entity (e.g., a max, a span/duration, a total) from a multi-row-per-entity table, and then splitting the entities by a computed threshold such as a median.
- **Pattern**: The attempt computes the statistic directly on the raw record-level table or with an unstated aggregation/duration convention (counting rows vs. differencing timestamps, inclusive vs. exclusive endpoints, unit of time), and applies the threshold split without stating how ties/boundary values and duplicate entity rows are handled; the numeric result is plausible and the qualitative verdict is right, but the coefficient is off by a few hundredths.
- **Detection procedure**:
  1. From the task, list every derived quantity (per-entity aggregate, span/duration, damage total) and the split rule, and note that the answer format demands a specific rounding precision — so small definitional differences change the graded value.
  2. In the scripts, locate the aggregation step: confirm there is an explicit group-by producing exactly one row per entity, an explicit duration formula with stated units and endpoint convention, and an explicit rule for records at/above vs. below the threshold.
  3. Check for a printed sanity block: number of unique entities, number of rows in each split (should sum to the total and be near-equal for a median split), and min/max/range of the derived duration and category values.
  4. If scripts are absent, or any of these definitions/counts is implicit or unprinted, treat the reported coefficient as unverified and flag the attempt.
- **Discriminator**: A real violation is an attempt whose entity-level table, duration convention, or split membership cannot be reconstructed and checked from the code/output; a look-alike that is fine explicitly aggregates to one row per entity, prints group sizes and value ranges, and shows the coefficient is stable under the documented convention (or notes the alternative convention gives the same rounded value).
- **Consequence**: The direction and significance verdict pass, but the rounded correlation coefficient differs from ground truth (e.g., 0.58 vs 0.56), so the exact-value check fails and the overall answer is marked incorrect.
470Analysis run on unverified/substituted data with the required output artifact never producedtaskda-code
Applies when
task -- the task points to a provided dataset and asks for a result saved to a specific output file that must mirror a given sample/template format.
Pattern
The scripts spend all their effort searching the internet or guessing at look-alike datasets, never successfully load the supplied input, and the final answer quotes numbers (and a "saved to file" claim) that no executed script actually computed or wrote; the sample/template file is never opened to confirm column names, ordering, or value formatting.
Detection procedure
  1. Read the task and note (a) the required input source and (b) the required output file plus its format template.
  2. Scan the scripts for a line that successfully reads the actual provided input and a line that writes the required output file (e.g., to_csv("result.csv")) using columns copied from the template; also check whether the template file itself is ever read/printed.
  3. Cross-check every number in the final answer against a script that demonstrably ran on the loaded data — if raw values or statistics appear only in prose, treat them as fabricated.
  4. Confirm any stated constraint (seed, test type, rounding, column names) is actually exercised in code, not just asserted in the write-up.
Discriminator
A genuine attempt loads the provided data (or a verifiably identical copy, with a printed shape/head sanity check) and contains explicit write code for the output file; a violation is when the data provenance is unresolved/guessed or the output file write is only claimed in text. Fetching an external mirror is acceptable only if the script prints the loaded rows/shape and the values in the answer trace back to that print.
Consequence
The graded output file is missing or contains values/columns that don't match the expected result, so the file check fails regardless of how plausible the reported p-value or metric sounds.
id c151912658da · mined from da-code dacode-data-sa-039@s8
raw text (what the judge reads)
### Analysis run on unverified/substituted data with the required output artifact never produced
- **Applies when**: `task` -- the task points to a provided dataset and asks for a result saved to a specific output file that must mirror a given sample/template format.
- **Pattern**: The scripts spend all their effort searching the internet or guessing at look-alike datasets, never successfully load the supplied input, and the final answer quotes numbers (and a "saved to file" claim) that no executed script actually computed or wrote; the sample/template file is never opened to confirm column names, ordering, or value formatting.
- **Detection procedure**:
  1. Read the task and note (a) the required input source and (b) the required output file plus its format template.
  2. Scan the scripts for a line that successfully reads the actual provided input and a line that writes the required output file (e.g., `to_csv("result.csv")`) using columns copied from the template; also check whether the template file itself is ever read/printed.
  3. Cross-check every number in the final answer against a script that demonstrably ran on the loaded data — if raw values or statistics appear only in prose, treat them as fabricated.
  4. Confirm any stated constraint (seed, test type, rounding, column names) is actually exercised in code, not just asserted in the write-up.
- **Discriminator**: A genuine attempt loads the provided data (or a verifiably identical copy, with a printed shape/head sanity check) and contains explicit write code for the output file; a violation is when the data provenance is unresolved/guessed or the output file write is only claimed in text. Fetching an external mirror is acceptable only if the script prints the loaded rows/shape and the values in the answer trace back to that print.
- **Consequence**: The graded output file is missing or contains values/columns that don't match the expected result, so the file check fails regardless of how plausible the reported p-value or metric sounds.
471String-formatted numeric columns ranked without type coercion (and extremes never sanity-checked)taskda-code
Applies when
task -- the task asks for the max/min (or any ranking/aggregate) of a numeric field in a raw tabular file where values may be stored as text with thousands separators, %, currency/unit symbols, or footnote markers.
Pattern
The attempt loads the file with defaults, imputes/aggregates or sorts the column as-is, and reports the top/bottom label without ever confirming the column's dtype is numeric or that the reported extreme value is plausible. Lexicographic sorting of strings (or silent coercion of only some rows) then yields a wrong "highest"/"lowest" entity, and mean-imputation of a non-numeric column silently no-ops or fills the wrong rows.
Detection procedure
  1. Read the task to identify which column(s) the requested statistic depends on, and any required preprocessing step (e.g., fill missing values before ranking).
  2. In the scripts, check for an explicit cleaning/conversion step (strip separators/symbols, to_numeric(..., errors='coerce'), dtype assertion) applied before imputation and before sorting/idxmax/idxmin; absence of such a step, or reliance on default parsing, is a red flag.
  3. Check whether the script prints/asserts the extreme values (not just the labels) and the count of NaNs before/after imputation; if only the entity names are printed, the result is unverified.
  4. Compare the reported extremes to domain plausibility (does the "lowest" entity really have the smallest value given known magnitudes, are ties handled as a list?).
Discriminator
Fine if the script demonstrates the column is already numeric (dtype check, describe(), or an explicit cleaning routine) and reports the extreme values alongside the labels; a violation is when ranking/imputation is done on a column whose numeric type was never established or whose extreme value was never inspected.
Consequence
The grader compares the reported entity names to ground truth and marks them wrong, because string ordering or partially-coerced values picked a different row than the true numeric max/min.
id 9a93d62ca68c · mined from da-code dacode-di-text-001@s8
raw text (what the judge reads)
### String-formatted numeric columns ranked without type coercion (and extremes never sanity-checked)
- **Applies when**: `task` -- the task asks for the max/min (or any ranking/aggregate) of a numeric field in a raw tabular file where values may be stored as text with thousands separators, %, currency/unit symbols, or footnote markers.
- **Pattern**: The attempt loads the file with defaults, imputes/aggregates or sorts the column as-is, and reports the top/bottom label without ever confirming the column's dtype is numeric or that the reported extreme value is plausible. Lexicographic sorting of strings (or silent coercion of only some rows) then yields a wrong "highest"/"lowest" entity, and mean-imputation of a non-numeric column silently no-ops or fills the wrong rows.
- **Detection procedure**:
  1. Read the task to identify which column(s) the requested statistic depends on, and any required preprocessing step (e.g., fill missing values before ranking).
  2. In the scripts, check for an explicit cleaning/conversion step (strip separators/symbols, `to_numeric(..., errors='coerce')`, dtype assertion) applied *before* imputation and before sorting/`idxmax`/`idxmin`; absence of such a step, or reliance on default parsing, is a red flag.
  3. Check whether the script prints/asserts the extreme *values* (not just the labels) and the count of NaNs before/after imputation; if only the entity names are printed, the result is unverified.
  4. Compare the reported extremes to domain plausibility (does the "lowest" entity really have the smallest value given known magnitudes, are ties handled as a list?).
- **Discriminator**: Fine if the script demonstrates the column is already numeric (dtype check, describe(), or an explicit cleaning routine) and reports the extreme values alongside the labels; a violation is when ranking/imputation is done on a column whose numeric type was never established or whose extreme value was never inspected.
- **Consequence**: The grader compares the reported entity names to ground truth and marks them wrong, because string ordering or partially-coerced values picked a different row than the true numeric max/min.
472Output file omits components the task explicitly asked to includetaskda-code
Applies when
task -- the task asks to compute several intermediate quantities/scores and then save results "including" derived groupings or labels to a specific output file.
Pattern
The script computes all the requested intermediate quantities in memory but writes only the final label/prediction column (plus an ID) to the required output path, dumping the full table to an unrequested side file instead; the answer text then reports summary counts rather than confirming the saved schema.
Detection procedure
  1. From the task statement, list every quantity that must appear in the deliverable file (each component metric, the composite score, the segment/label), and any implied identifier/ordering/index conventions.
  2. In the script, find the line that writes the required filename and read exactly which columns are selected there (and whether the index is written).
  3. Compare that column set against the list from step 1; also check whether a richer table was written to a different, non-requested filename — a strong sign the deliverable was over-trimmed.
  4. Check the answer text: does it state the saved file's columns, row count and a couple of sample rows, or only aggregate distributions?
Discriminator
A real violation is when quantities named or clearly implied by the task ("including X and Y") are computed but absent from the deliverable file, or the deliverable is a strict subset of a side file. It is not a violation if the task genuinely asks for only the label column, or if extra columns beyond the required set are present (supersets are usually acceptable).
Consequence
The grader compares the expected result file column-by-column (or joins on the expected schema) and reports the file as WRONG/MISSING even though the underlying computation may be close, yielding 0 credit.
id 56c08e90814d · mined from da-code dacode-dm-csv-052@s8
raw text (what the judge reads)
### Output file omits components the task explicitly asked to include
- **Applies when**: `task` -- the task asks to compute several intermediate quantities/scores and then save results "including" derived groupings or labels to a specific output file.
- **Pattern**: The script computes all the requested intermediate quantities in memory but writes only the final label/prediction column (plus an ID) to the required output path, dumping the full table to an unrequested side file instead; the answer text then reports summary counts rather than confirming the saved schema.
- **Detection procedure**:
  1. From the task statement, list every quantity that must appear in the deliverable file (each component metric, the composite score, the segment/label), and any implied identifier/ordering/index conventions.
  2. In the script, find the line that writes the required filename and read exactly which columns are selected there (and whether the index is written).
  3. Compare that column set against the list from step 1; also check whether a richer table was written to a *different*, non-requested filename — a strong sign the deliverable was over-trimmed.
  4. Check the answer text: does it state the saved file's columns, row count and a couple of sample rows, or only aggregate distributions?
- **Discriminator**: A real violation is when quantities named or clearly implied by the task ("including X and Y") are computed but absent from the deliverable file, or the deliverable is a strict subset of a side file. It is *not* a violation if the task genuinely asks for only the label column, or if extra columns beyond the required set are present (supersets are usually acceptable).
- **Consequence**: The grader compares the expected result file column-by-column (or joins on the expected schema) and reports the file as WRONG/MISSING even though the underlying computation may be close, yielding 0 credit.
473Blind numeric coercion and arbitrary input selection, with no plausibility check on the resulting statistictaskda-code
Applies when
task -- the script auto-discovers its input file (first match in a directory / hard-coded guess list) and/or force-casts columns with errors='coerce'-style conversion before computing a summary statistic, and a reference/sample output file or domain expectation exists.
Pattern
The attempt never verifies that it loaded the intended table or that the coerced column values are the real values (it only counts how many became NaN), then reports the computed statistic as-is even when its magnitude/sign is implausible, and never diffs its output against the provided sample/reference format.
Detection procedure
  1. In the task, note any provided sample/reference output file, format or rounding constraint, and note what relationship among the variables is substantively expected.
  2. In the scripts, check whether the input file is chosen deterministically and validated (expected shape, column set, value ranges, a printed head of the raw column values), whether coercion drops/mangles a non-trivial share of rows, and whether the sample file is ever read and compared.
  3. In the answer, compare the reported numbers against a basic sanity expectation (e.g., strongly related measures should not all be ~0; ranges, row counts and matrix shape/labels should be consistent) and against the required output format.
  4. Flag if no such validation exists and the reported values are implausible or the format was never checked against the sample.
Discriminator
Fine if the script pins the exact input, prints raw values before/after conversion, reports how many rows were lost, and either matches the sample layout explicitly or explains a genuinely surprising result with evidence; a violation is a pipeline that only prints NaN counts and accepts whatever numbers fall out, with implausible output left uninvestigated.
Consequence
The saved result file contains values (or a layout) that do not match the expected reference, so the file-level check fails outright even though the script ran without errors.
id 64ed754c31f4 · mined from da-code dacode-data-sa-026@s8
raw text (what the judge reads)
### Blind numeric coercion and arbitrary input selection, with no plausibility check on the resulting statistic
- **Applies when**: `task` -- the script auto-discovers its input file (first match in a directory / hard-coded guess list) and/or force-casts columns with `errors='coerce'`-style conversion before computing a summary statistic, and a reference/sample output file or domain expectation exists.
- **Pattern**: The attempt never verifies that it loaded the intended table or that the coerced column values are the real values (it only counts how many became NaN), then reports the computed statistic as-is even when its magnitude/sign is implausible, and never diffs its output against the provided sample/reference format.
- **Detection procedure**:
  1. In the task, note any provided sample/reference output file, format or rounding constraint, and note what relationship among the variables is substantively expected.
  2. In the scripts, check whether the input file is chosen deterministically and validated (expected shape, column set, value ranges, a printed head of the raw column values), whether coercion drops/mangles a non-trivial share of rows, and whether the sample file is ever read and compared.
  3. In the answer, compare the reported numbers against a basic sanity expectation (e.g., strongly related measures should not all be ~0; ranges, row counts and matrix shape/labels should be consistent) and against the required output format.
  4. Flag if no such validation exists and the reported values are implausible or the format was never checked against the sample.
- **Discriminator**: Fine if the script pins the exact input, prints raw values before/after conversion, reports how many rows were lost, and either matches the sample layout explicitly or explains a genuinely surprising result with evidence; a violation is a pipeline that only prints NaN counts and accepts whatever numbers fall out, with implausible output left uninvestigated.
- **Consequence**: The saved result file contains values (or a layout) that do not match the expected reference, so the file-level check fails outright even though the script ran without errors.
474Deliverable file not produced/verified with the exact requested schemataskda-code
Applies when
task -- the task names an output artifact (e.g. a CSV) with specific column names and content, and the agent instead reports a narrative summary.
Pattern
The attempt describes features, model choice and cluster profiles in prose, but there is no saved script that writes the artifact, and no post-write check that the file exists at the expected path with the exact required column names, row count and label column.
Detection procedure
1. From the task, list the required filename, required columns (including exact naming pattern/indexing) and expected row granularity. 2. Search the scripts/outputs for the code that writes that filename and inspect the DataFrame columns passed to it. 3. Check the answer/log for an explicit verification of the written file (head of file, shape, column list). 4. Flag if the writing code is absent, the columns are renamed/reordered/missing, or verification is only asserted in prose.
Discriminator
A real violation is when the artifact is never demonstrably written or its header/shape is unverified or mismatched; a look-alike that is fine is when a script clearly writes the file with the exact columns and the agent shows the file's actual head/shape, even if the prose summary is truncated.
Consequence
The grader looks for the expected file with the expected schema and marks it WRONG/MISSING, scoring 0 regardless of how sound the analysis narrative is.
id 4ee947ebce67 · mined from da-code dacode-ml-cluster-019@s8
raw text (what the judge reads)
### Deliverable file not produced/verified with the exact requested schema
- **Applies when**: `task` -- the task names an output artifact (e.g. a CSV) with specific column names and content, and the agent instead reports a narrative summary.
- **Pattern**: The attempt describes features, model choice and cluster profiles in prose, but there is no saved script that writes the artifact, and no post-write check that the file exists at the expected path with the exact required column names, row count and label column.
- **Detection procedure**: 1. From the task, list the required filename, required columns (including exact naming pattern/indexing) and expected row granularity. 2. Search the scripts/outputs for the code that writes that filename and inspect the DataFrame columns passed to it. 3. Check the answer/log for an explicit verification of the written file (head of file, shape, column list). 4. Flag if the writing code is absent, the columns are renamed/reordered/missing, or verification is only asserted in prose.
- **Discriminator**: A real violation is when the artifact is never demonstrably written or its header/shape is unverified or mismatched; a look-alike that is fine is when a script clearly writes the file with the exact columns and the agent shows the file's actual head/shape, even if the prose summary is truncated.
- **Consequence**: The grader looks for the expected file with the expected schema and marks it WRONG/MISSING, scoring 0 regardless of how sound the analysis narrative is.
475No out-of-sample estimate of the stated evaluation metric before submitting probabilistic predictionstaskda-code
Applies when
task -- the task specifies a scoring metric (e.g., a probabilistic loss) and the scripts train a model on the full labeled set and write predictions for an unlabeled set without ever scoring the model.
Pattern
The attempt fits a single, aggressively-parameterized model (deep trees, many estimators, no regularization/calibration search), skips any cross-validation or holdout scoring under the stated metric, and "validates" only file mechanics (column names, row counts, values in [0,1], rows summing to 1). Predicted probabilities are therefore uncalibrated/overconfident (many values at 1e-5 or 0.999) with no evidence they beat even a class-prior baseline.
Detection procedure
  1. Read the task and note the exact scoring metric and prediction type (probabilities vs. labels).
  2. Search the scripts for any split (train_test_split, KFold, cross_val_score) plus an explicit computation of that metric on data not used for fitting; also look for a comparison against a trivial baseline (e.g., predicting training class frequencies).
  3. If absent, inspect the produced predictions for the distribution of extreme values and for the rare classes — a metric that penalizes confident errors will be dominated by these.
  4. Confirm the reported deliverable includes no metric estimate, only format checks.
Discriminator
A real violation is no out-of-sample metric computation at all (format-only verification, or only training-set accuracy). It is fine if the agent did CV/holdout under the correct metric, compared candidates/baselines, and then refit on all data — even if the final probabilities are sharp, since sharpness was justified by validation.
Consequence
The submission may be syntactically valid but scores far worse than a naive prior baseline (log loss inflated by near-zero probabilities on wrong classes), so the graded result fails the accuracy/quality threshold despite passing format checks.
id 6e884630c5ef · mined from da-code dacode-ml-competition-005@s9
raw text (what the judge reads)
### No out-of-sample estimate of the stated evaluation metric before submitting probabilistic predictions
- **Applies when**: `task` -- the task specifies a scoring metric (e.g., a probabilistic loss) and the scripts train a model on the full labeled set and write predictions for an unlabeled set without ever scoring the model.
- **Pattern**: The attempt fits a single, aggressively-parameterized model (deep trees, many estimators, no regularization/calibration search), skips any cross-validation or holdout scoring under the stated metric, and "validates" only file mechanics (column names, row counts, values in [0,1], rows summing to 1). Predicted probabilities are therefore uncalibrated/overconfident (many values at 1e-5 or 0.999) with no evidence they beat even a class-prior baseline.
- **Detection procedure**:
  1. Read the task and note the exact scoring metric and prediction type (probabilities vs. labels).
  2. Search the scripts for any split (`train_test_split`, `KFold`, `cross_val_score`) plus an explicit computation of that metric on data not used for fitting; also look for a comparison against a trivial baseline (e.g., predicting training class frequencies).
  3. If absent, inspect the produced predictions for the distribution of extreme values and for the rare classes — a metric that penalizes confident errors will be dominated by these.
  4. Confirm the reported deliverable includes no metric estimate, only format checks.
- **Discriminator**: A real violation is *no* out-of-sample metric computation at all (format-only verification, or only training-set accuracy). It is fine if the agent did CV/holdout under the correct metric, compared candidates/baselines, and then refit on all data — even if the final probabilities are sharp, since sharpness was justified by validation.
- **Consequence**: The submission may be syntactically valid but scores far worse than a naive prior baseline (log loss inflated by near-zero probabilities on wrong classes), so the graded result fails the accuracy/quality threshold despite passing format checks.
476Output file schema deviates from the explicitly requested columns/row alignmenttaskda-code
Applies when
task -- the task names an output file and the exact column name(s) it must contain, and the script builds that file from a given input table.
Pattern
The attempt writes the file with extra, renamed, or reordered columns (e.g. adding an identifier or index column, or using a differently-spelled header), and/or does not guarantee one row per input row in the input's original order, then declares success based on model metrics rather than on schema conformance. Often paired with a self-reported fit score computed on the training rows only, so no independent check catches the mismatch.
Detection procedure
  1. From the task statement, list the required output filename, the exact required column name(s), and the implied row count/order (usually equal to and aligned with the provided input file).
  2. In the scripts, find the write call (e.g. to_csv) and read the DataFrame passed to it: enumerate its columns, whether an index is written, and whether rows were filtered, dropped (e.g. via NaN handling), sorted, or concatenated from multiple sources.
  3. Compare that column list and row construction against step 1; also check that column names read from the input match the documented header spellings so no silent key error/renaming occurred.
  4. Check the answer text: does it assert the file conforms (columns and row count) as verified, or only report training-fit statistics as evidence of success?
Discriminator
A real violation is any deviation from the specified header set/spelling, an unwanted index/extra column, or a row count/order that no longer aligns with the input rows. It is not a violation if extra columns are explicitly permitted by the task, or if the file contains exactly the required column plus an identifier the task itself asked for; likewise, dropping rows is fine only when the task requested filtering.
Consequence
The grader loads the expected file and compares schema and per-row predictions; an extra/renamed column or misaligned rows makes the file unparseable or unmatchable, so the check is scored WRONG/MISSING regardless of model quality.
id 4b327d145f6e · mined from da-code dacode-ml-regression-008@s9
raw text (what the judge reads)
### Output file schema deviates from the explicitly requested columns/row alignment
- **Applies when**: `task` -- the task names an output file and the exact column name(s) it must contain, and the script builds that file from a given input table.
- **Pattern**: The attempt writes the file with extra, renamed, or reordered columns (e.g. adding an identifier or index column, or using a differently-spelled header), and/or does not guarantee one row per input row in the input's original order, then declares success based on model metrics rather than on schema conformance. Often paired with a self-reported fit score computed on the training rows only, so no independent check catches the mismatch.
- **Detection procedure**:
  1. From the task statement, list the required output filename, the exact required column name(s), and the implied row count/order (usually equal to and aligned with the provided input file).
  2. In the scripts, find the write call (e.g. `to_csv`) and read the DataFrame passed to it: enumerate its columns, whether an index is written, and whether rows were filtered, dropped (e.g. via NaN handling), sorted, or concatenated from multiple sources.
  3. Compare that column list and row construction against step 1; also check that column names read from the input match the documented header spellings so no silent key error/renaming occurred.
  4. Check the answer text: does it assert the file conforms (columns and row count) as verified, or only report training-fit statistics as evidence of success?
- **Discriminator**: A real violation is any deviation from the specified header set/spelling, an unwanted index/extra column, or a row count/order that no longer aligns with the input rows. It is *not* a violation if extra columns are explicitly permitted by the task, or if the file contains exactly the required column plus an identifier the task itself asked for; likewise, dropping rows is fine only when the task requested filtering.
- **Consequence**: The grader loads the expected file and compares schema and per-row predictions; an extra/renamed column or misaligned rows makes the file unparseable or unmatchable, so the check is scored WRONG/MISSING regardless of model quality.
477Unit of analysis, population subset, and test form assumed rather than justifiedtaskda-code
Applies when
task -- the task asks for a hypothesis test / summary statistic comparing two groups, and the scripts build the comparison samples directly from the raw files without explicitly deciding what one observation is, which rows belong in scope, and what test form (one- vs two-sided, parametric vs rank-based) the question implies.
Pattern
The attempt takes the whole file(s) as-is, reshapes columns into a single pooled vector (e.g. stacking per-side or per-entity values instead of aggregating to the natural record level), skips any filtering to the population the question is about (time window, category, competition/segment), and calls a default library test (two-sided, equal-variance, normality-assuming) on skewed count data — with no check that the resulting n, mean, and distribution shape match the stated question.
Detection procedure
  1. Read the task statement and note the exact quantity being compared (per-record aggregate vs per-row value), any implied scope restriction, and whether the wording implies a directional or non-directional comparison.
  2. In the script, locate how the two samples are constructed: check whether rows are aggregated/derived to the unit named in the task, whether any subset filter is applied, and whether sample sizes are plausible for that scope.
  3. Check the test call: is the choice of one/two-sided, parametric/non-parametric, and variance assumption argued for (e.g. via a distribution or normality/skew check), or is it the library default?
  4. Compare the reported p-value's magnitude to what the chosen scope and n imply — an astronomically small p-value from a pooled, unfiltered, un-aggregated sample is a red flag that the sample is much larger and differently defined than intended.
Discriminator
A real violation is when the script's sample definition (unit, scope, or test form) cannot be traced back to a sentence in the task and no sensitivity/sanity check was run; it is fine if the task genuinely asks for the full unfiltered population at row level and the script documents that reading plus verifies distributional assumptions or uses a robust test.
Consequence
The p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong population and wrong observation unit, so the numeric column mismatches the expected value and the output file is graded WRONG even though the file format is correct.
id 46bd7ef329d0 · mined from da-code dacode-data-sa-001@s9
raw text (what the judge reads)
### Unit of analysis, population subset, and test form assumed rather than justified
- **Applies when**: `task` -- the task asks for a hypothesis test / summary statistic comparing two groups, and the scripts build the comparison samples directly from the raw files without explicitly deciding what one observation is, which rows belong in scope, and what test form (one- vs two-sided, parametric vs rank-based) the question implies.
- **Pattern**: The attempt takes the whole file(s) as-is, reshapes columns into a single pooled vector (e.g. stacking per-side or per-entity values instead of aggregating to the natural record level), skips any filtering to the population the question is about (time window, category, competition/segment), and calls a default library test (two-sided, equal-variance, normality-assuming) on skewed count data — with no check that the resulting n, mean, and distribution shape match the stated question.
- **Detection procedure**:
  1. Read the task statement and note the exact quantity being compared (per-record aggregate vs per-row value), any implied scope restriction, and whether the wording implies a directional or non-directional comparison.
  2. In the script, locate how the two samples are constructed: check whether rows are aggregated/derived to the unit named in the task, whether any subset filter is applied, and whether sample sizes are plausible for that scope.
  3. Check the test call: is the choice of one/two-sided, parametric/non-parametric, and variance assumption argued for (e.g. via a distribution or normality/skew check), or is it the library default?
  4. Compare the reported p-value's magnitude to what the chosen scope and n imply — an astronomically small p-value from a pooled, unfiltered, un-aggregated sample is a red flag that the sample is much larger and differently defined than intended.
- **Discriminator**: A real violation is when the script's sample definition (unit, scope, or test form) cannot be traced back to a sentence in the task and no sensitivity/sanity check was run; it is fine if the task genuinely asks for the full unfiltered population at row level and the script documents that reading plus verifies distributional assumptions or uses a robust test.
- **Consequence**: The p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong population and wrong observation unit, so the numeric column mismatches the expected value and the output file is graded WRONG even though the file format is correct.
478Co-aggregated measures from joined tables never cross-checked for consistency (fan-out / mis-scoped SUMs)taskda-code
Applies when
task -- the task asks for two or more aggregate measures per group (e.g., a count/quantity total and a monetary total) that are computed from joined or multi-level tables, with a stated adjustment (discount, tax, unit conversion) and a template file defining the output format.
Pattern
The attempt joins/merges tables and sums each measure without checking that both measures are aggregated over the same row set and grain: one measure gets deduplicated, filtered, or taken from the header/parent level while the other is summed over fanned-out child rows (or the adjustment is applied to only part of the expression). The result is written out without comparing derived ratios, group totals, or the column layout against the provided sample/template.
Detection procedure
  1. From the task, list every requested measure, the required adjustment formula, the filter/date scope, and the exact column names/order/rounding shown in the sample output file.
  2. In the scripts, trace each measure back to the table and row set it is summed over; confirm both come from the same post-join line-item grain, that the filter is applied once at the correct level, and that the adjustment enters the monetary formula (not the quantity) as specified.
  3. Compute simple sanity ratios from the reported answer (money ÷ quantity per group, and overall totals versus known row counts/price ranges); flag if implied per-unit values differ wildly across groups or are implausible relative to the raw data.
  4. Compare the emitted file header, column order, sort order, and decimal formatting byte-for-byte in spirit with the sample file.
Discriminator
A real violation shows measures drawn from different grains/subsets or implied per-unit values that the raw data cannot support (or a layout diverging from the template); a look-alike that is fine has genuinely uneven per-unit values that the reviewer can confirm from the underlying price/discount distribution, with both sums traced to the identical filtered join.
Consequence
The grader compares against the reference file and marks it WRONG because group-level totals (and possibly column order/rounding) do not match, even though the output looks superficially well-formed.
id b49551e8f08b · mined from da-code dacode-dm-csv-011@s9
raw text (what the judge reads)
### Co-aggregated measures from joined tables never cross-checked for consistency (fan-out / mis-scoped SUMs)
- **Applies when**: `task` -- the task asks for two or more aggregate measures per group (e.g., a count/quantity total and a monetary total) that are computed from joined or multi-level tables, with a stated adjustment (discount, tax, unit conversion) and a template file defining the output format.
- **Pattern**: The attempt joins/merges tables and sums each measure without checking that both measures are aggregated over the same row set and grain: one measure gets deduplicated, filtered, or taken from the header/parent level while the other is summed over fanned-out child rows (or the adjustment is applied to only part of the expression). The result is written out without comparing derived ratios, group totals, or the column layout against the provided sample/template.
- **Detection procedure**:
  1. From the task, list every requested measure, the required adjustment formula, the filter/date scope, and the exact column names/order/rounding shown in the sample output file.
  2. In the scripts, trace each measure back to the table and row set it is summed over; confirm both come from the same post-join line-item grain, that the filter is applied once at the correct level, and that the adjustment enters the monetary formula (not the quantity) as specified.
  3. Compute simple sanity ratios from the reported answer (money ÷ quantity per group, and overall totals versus known row counts/price ranges); flag if implied per-unit values differ wildly across groups or are implausible relative to the raw data.
  4. Compare the emitted file header, column order, sort order, and decimal formatting byte-for-byte in spirit with the sample file.
- **Discriminator**: A real violation shows measures drawn from different grains/subsets or implied per-unit values that the raw data cannot support (or a layout diverging from the template); a look-alike that is fine has genuinely uneven per-unit values that the reviewer can confirm from the underlying price/discount distribution, with both sums traced to the identical filtered join.
- **Consequence**: The grader compares against the reference file and marks it WRONG because group-level totals (and possibly column order/rounding) do not match, even though the output looks superficially well-formed.
479Deliverable file asserted in prose but never verified on disk (or produced by an unsaved script)taskda-code
Applies when
task -- the task requires writing a result artifact (a file with specified name, columns, and one row per input record) and the agent's answer describes that artifact instead of showing evidence it was created correctly.
Pattern
The agent narrates its pipeline and claims the output file was saved with a given shape/columns, but there is no saved, re-runnable script and no post-write read-back check: the file may be absent, in the wrong directory, have different/misnamed columns, wrong row count, or degenerate content (e.g., a label column with a single value or a trivially small number of groups chosen with no selection rationale).
Detection procedure
  1. From the task statement, list the exact required deliverable: file name/path, exact column names and their meaning, expected row count, and any per-row correspondence to input records.
  2. In the scripts, locate the write call and check that (a) the script is persisted and re-runnable end-to-end, (b) the written path matches the required name/location, and (c) the column names are constructed exactly as specified (indexing/order included).
  3. Look for a verification step after writing — re-read the file and assert shape, column names, no unexpected NaNs, and that the label/target column has a plausible distribution (more than one group, no near-empty group) — and check that the answer reports these observed values, not intended ones.
  4. Check the answer for internal consistency with the task and with any reported numbers (column count vs. listed features, row count vs. dataset size, group count vs. any model-selection evidence); truncated or self-contradictory descriptions count as unverified.
Discriminator
A real violation is an answer whose claims about the artifact rest only on narration (no read-back assertions, no script, mismatched or unlisted columns, unjustified degenerate labels). It is fine if the agent shows the actual re-read output (head, shape, columns, value counts) matching the specification, even if the modeling choices themselves are simple.
Consequence
The grader looks for the specified file with the specified columns/content and marks it WRONG/MISSING, failing all checks regardless of how sound the described methodology was.
id 6acdd3997f31 · mined from da-code dacode-ml-cluster-014@s9
raw text (what the judge reads)
### Deliverable file asserted in prose but never verified on disk (or produced by an unsaved script)
- **Applies when**: `task` -- the task requires writing a result artifact (a file with specified name, columns, and one row per input record) and the agent's answer describes that artifact instead of showing evidence it was created correctly.
- **Pattern**: The agent narrates its pipeline and claims the output file was saved with a given shape/columns, but there is no saved, re-runnable script and no post-write read-back check: the file may be absent, in the wrong directory, have different/misnamed columns, wrong row count, or degenerate content (e.g., a label column with a single value or a trivially small number of groups chosen with no selection rationale).
- **Detection procedure**:
  1. From the task statement, list the exact required deliverable: file name/path, exact column names and their meaning, expected row count, and any per-row correspondence to input records.
  2. In the scripts, locate the write call and check that (a) the script is persisted and re-runnable end-to-end, (b) the written path matches the required name/location, and (c) the column names are constructed exactly as specified (indexing/order included).
  3. Look for a verification step after writing — re-read the file and assert shape, column names, no unexpected NaNs, and that the label/target column has a plausible distribution (more than one group, no near-empty group) — and check that the answer reports these observed values, not intended ones.
  4. Check the answer for internal consistency with the task and with any reported numbers (column count vs. listed features, row count vs. dataset size, group count vs. any model-selection evidence); truncated or self-contradictory descriptions count as unverified.
- **Discriminator**: A real violation is an answer whose claims about the artifact rest only on narration (no read-back assertions, no script, mismatched or unlisted columns, unjustified degenerate labels). It is fine if the agent shows the actual re-read output (head, `shape`, `columns`, value counts) matching the specification, even if the modeling choices themselves are simple.
- **Consequence**: The grader looks for the specified file with the specified columns/content and marks it WRONG/MISSING, failing all checks regardless of how sound the described methodology was.
480Submission never validated against the provided template (and outputs coerced by unjustified rounding)taskda-code
Applies when
task -- the task supplies a sample/reference output file defining the required rows, ID set, column names and value type, and the scripts write a predictions file.
Pattern
The scripts load only the train/test inputs, build predictions, and write the output directly, without ever reading the template file to confirm the row count, ID ordering/coverage, column names and value dtype; on top of that they apply post-hoc transforms (rounding to integers, clipping to invented bounds) that are not required by the task and are never checked against the evaluation metric or the observed target distribution.
Detection procedure
  1. Read the task/README to note that a reference output file exists and what it prescribes (columns, one row per test record, numeric type, any ordering).
  2. Grep the scripts for a read of that reference file and for any assertion comparing the produced file's shape/IDs/columns to it — its absence is the flag.
  3. Inspect the final prediction step for transforms (round, astype(int), clip) applied to regression outputs, and check whether the task or metric anywhere justifies discretization or those bounds.
  4. Inspect the emitted answer itself: count rows and compare to the number of test rows; check IDs match the test set exactly and the header matches the template.
Discriminator
A real violation is when no programmatic check ties the output to the template and the value space is altered without justification (e.g., continuous targets forced to integers, or a truncated/misaligned row set). It is not a violation if the task explicitly requires integer/class outputs, or if the script includes an assertion on shape/IDs/columns versus the template even if formatting choices are otherwise debatable.
Consequence
The grader marks the submission file WRONG/MISSING — either it cannot be aligned with the expected IDs/rows, or the needless rounding inflates the error metric relative to the raw continuous predictions, so the attempt fails the file check despite the model itself being reasonable.
id a87285e7b276 · mined from da-code dacode-ml-competition-009@s9
raw text (what the judge reads)
### Submission never validated against the provided template (and outputs coerced by unjustified rounding)
- **Applies when**: `task` -- the task supplies a sample/reference output file defining the required rows, ID set, column names and value type, and the scripts write a predictions file.
- **Pattern**: The scripts load only the train/test inputs, build predictions, and write the output directly, without ever reading the template file to confirm the row count, ID ordering/coverage, column names and value dtype; on top of that they apply post-hoc transforms (rounding to integers, clipping to invented bounds) that are not required by the task and are never checked against the evaluation metric or the observed target distribution.
- **Detection procedure**:
  1. Read the task/README to note that a reference output file exists and what it prescribes (columns, one row per test record, numeric type, any ordering).
  2. Grep the scripts for a read of that reference file and for any assertion comparing the produced file's shape/IDs/columns to it — its absence is the flag.
  3. Inspect the final prediction step for transforms (`round`, `astype(int)`, `clip`) applied to regression outputs, and check whether the task or metric anywhere justifies discretization or those bounds.
  4. Inspect the emitted answer itself: count rows and compare to the number of test rows; check IDs match the test set exactly and the header matches the template.
- **Discriminator**: A real violation is when no programmatic check ties the output to the template *and* the value space is altered without justification (e.g., continuous targets forced to integers, or a truncated/misaligned row set). It is not a violation if the task explicitly requires integer/class outputs, or if the script includes an assertion on shape/IDs/columns versus the template even if formatting choices are otherwise debatable.
- **Consequence**: The grader marks the submission file WRONG/MISSING — either it cannot be aligned with the expected IDs/rows, or the needless rounding inflates the error metric relative to the raw continuous predictions, so the attempt fails the file check despite the model itself being reasonable.
481Dropping rows/columns instead of repairing them, silently shrinking the output universetaskda-code
Applies when
task -- the task asks for a per-record output (labels, predictions, scores) over a supplied dataset, and the script decides which rows and columns to keep by dropna() and/or select_dtypes(include=numeric).
Pattern
The attempt auto-selects only the already-numeric columns (so columns stored as strings because of currency/percent/thousands separators are dropped, while meaningless ID-like numeric codes are kept) and drops every row containing a NaN, producing an output with fewer rows than input records and a feature vector unrelated to the dataset's substantive variables — with no attempt to parse/clean or impute.
Detection procedure
  1. Read the task for the expected output granularity: does it imply one output row per input record? Note the input record count.
  2. In the script, find the row/column filtering steps; check whether any string-formatted numeric columns are parsed (strip symbols, to_numeric), whether identifier-like numeric columns are excluded, and whether missing values are imputed rather than dropna()-ed.
  3. Compare the reported output shape and feature list against input record count and the README's substantive variables; flag if rows < records or if most documented indicators are absent while codes/IDs are present.
  4. Confirm no sanity check (row count equals input, feature count as intended) was printed and reconciled in the answer.
Discriminator
A real violation is silent, unjustified shrinkage of the record set or loss of informative columns caused purely by dtype/NaN convenience. It is fine if the task explicitly says to filter records, or if the script deliberately cleans/parses columns and imputes, and the answer states and justifies any remaining exclusions.
Consequence
The saved file has the wrong number of rows and a wrong-width/wrong-content feature matrix, so a row-count or shape/content comparison against the expected per-record output fails outright, regardless of how good the clustering itself is.
id 737558c9cb01 · mined from da-code dacode-ml-cluster-009@s9
raw text (what the judge reads)
### Dropping rows/columns instead of repairing them, silently shrinking the output universe
- **Applies when**: `task` -- the task asks for a per-record output (labels, predictions, scores) over a supplied dataset, and the script decides which rows and columns to keep by `dropna()` and/or `select_dtypes(include=numeric)`.
- **Pattern**: The attempt auto-selects only the already-numeric columns (so columns stored as strings because of currency/percent/thousands separators are dropped, while meaningless ID-like numeric codes are kept) and drops every row containing a NaN, producing an output with fewer rows than input records and a feature vector unrelated to the dataset's substantive variables — with no attempt to parse/clean or impute.
- **Detection procedure**:
  1. Read the task for the expected output granularity: does it imply one output row per input record? Note the input record count.
  2. In the script, find the row/column filtering steps; check whether any string-formatted numeric columns are parsed (strip symbols, `to_numeric`), whether identifier-like numeric columns are excluded, and whether missing values are imputed rather than `dropna()`-ed.
  3. Compare the reported output shape and feature list against input record count and the README's substantive variables; flag if rows < records or if most documented indicators are absent while codes/IDs are present.
  4. Confirm no sanity check (row count equals input, feature count as intended) was printed and reconciled in the answer.
- **Discriminator**: A real violation is silent, unjustified shrinkage of the record set or loss of informative columns caused purely by dtype/NaN convenience. It is fine if the task explicitly says to filter records, or if the script deliberately cleans/parses columns and imputes, and the answer states and justifies any remaining exclusions.
- **Consequence**: The saved file has the wrong number of rows and a wrong-width/wrong-content feature matrix, so a row-count or shape/content comparison against the expected per-record output fails outright, regardless of how good the clustering itself is.
482Output rows/features silently re-defined by aggregation or transformation, never sanity-checked against the source datataskda-code
Applies when
task -- the task asks for a per-record result file with a prescribed set of columns, and the script derives its own entity granularity (aggregating/grouping/filtering rows) or writes transformed values instead of the underlying feature values.
Pattern
The attempt collapses the raw table into a smaller derived unit of analysis (e.g., grouping many rows into one row per entity, dropping rows with missing keys, deduplicating) and/or exports scaled/encoded intermediate values as the "Feature_i" columns, without ever checking that the resulting row count, feature count and value semantics correspond to what the task's output specification implies. No reconciliation is shown between raw record count, retained count, and exported count.
Detection procedure
  1. Read the task's output spec and infer the intended unit of one output row and what a "feature vector" is expected to contain; note any implied count (all records, all entities, etc.).
  2. In the script, trace every operation that changes row count (groupby, dropna, filters, dedup) and every transform applied to the exported columns (scaling, log, PCA) and identify the final DataFrame written to disk.
  3. Compare the written file's shape/semantics to step 1: does the row count match a defensible population, and are exported feature columns the values the task asked to be clustered, or an internal intermediate?
  4. Check the answer for an explicit reconciliation and basic sanity checks (raw rows → dropped rows → final rows; cluster sizes not degenerate); absence of these, or clusters containing a handful of points reported as a "good" solution, is a red flag.
Discriminator
A real violation is when the aggregation/transform choice is unjustified by the task text, unreconciled with the raw data counts, and produces an output whose granularity or column meaning differs from the stated spec. It is fine if the script documents the derivation, the row count is traceable to the source data, and the exported columns are exactly the vector that was clustered in the units the task implies.
Consequence
The saved result file has the wrong number of rows and/or column semantics, so the grader's file check on shape/content fails even though the clustering metrics reported internally look excellent.
id 377f9140e28a · mined from da-code dacode-ml-cluster-016@s9
raw text (what the judge reads)
### Output rows/features silently re-defined by aggregation or transformation, never sanity-checked against the source data
- **Applies when**: `task` -- the task asks for a per-record result file with a prescribed set of columns, and the script derives its own entity granularity (aggregating/grouping/filtering rows) or writes transformed values instead of the underlying feature values.
- **Pattern**: The attempt collapses the raw table into a smaller derived unit of analysis (e.g., grouping many rows into one row per entity, dropping rows with missing keys, deduplicating) and/or exports scaled/encoded intermediate values as the "Feature_i" columns, without ever checking that the resulting row count, feature count and value semantics correspond to what the task's output specification implies. No reconciliation is shown between raw record count, retained count, and exported count.
- **Detection procedure**:
  1. Read the task's output spec and infer the intended unit of one output row and what a "feature vector" is expected to contain; note any implied count (all records, all entities, etc.).
  2. In the script, trace every operation that changes row count (groupby, dropna, filters, dedup) and every transform applied to the exported columns (scaling, log, PCA) and identify the final DataFrame written to disk.
  3. Compare the written file's shape/semantics to step 1: does the row count match a defensible population, and are exported feature columns the values the task asked to be clustered, or an internal intermediate?
  4. Check the answer for an explicit reconciliation and basic sanity checks (raw rows → dropped rows → final rows; cluster sizes not degenerate); absence of these, or clusters containing a handful of points reported as a "good" solution, is a red flag.
- **Discriminator**: A real violation is when the aggregation/transform choice is unjustified by the task text, unreconciled with the raw data counts, and produces an output whose granularity or column meaning differs from the stated spec. It is fine if the script documents the derivation, the row count is traceable to the source data, and the exported columns are exactly the vector that was clustered in the units the task implies.
- **Consequence**: The saved result file has the wrong number of rows and/or column semantics, so the grader's file check on shape/content fails even though the clustering metrics reported internally look excellent.
483Required output artifacts (beyond the headline file) not produced or verifiedtaskda-code
Applies when
task -- the task points to a spec/config file (e.g., a YAML/JSON format description) and/or a grader that expects specific persisted result files (image plus serialized data/array/metadata), and the agent must save outputs to disk.
Pattern
The attempt treats the visible deliverable (a single figure or a prose summary) as the whole job: it never opens/echoes the spec file to enumerate every required output, saves only one file, keeps no runnable script, and then asserts success in narrative form ("saved as ...", "follows the specification") without listing or re-reading the files on disk.
Detection procedure
  1. From the task text and any referenced spec/config file, list every artifact and field the deliverable must contain (file names, extensions, serialized data of the plotted values, title/label/size/color keys, ordering, units).
  2. Read the scripts for the corresponding write calls (e.g., one per required artifact) and check that the values written come from the actual plotted/computed objects rather than being retyped by hand.
  3. Check for a post-hoc verification step: does the attempt re-list the output directory and re-load each saved artifact to confirm existence, shape/length, dtype and key names?
  4. Compare the answer's claims against that verification evidence; a claim of compliance with no dump of the spec's requirements and no file listing is unsupported.
Discriminator
A real violation is missing/unwritten required artifacts or compliance asserted only in prose. It is not a violation if the spec genuinely requires a single file and the attempt shows the spec contents plus a load-back check confirming the file's structure — nor is extra descriptive prose a problem when the artifacts are all written and validated.
Consequence
File-level grader checks report the unproduced artifacts as WRONG/MISSING, so the task scores zero even if the rendered chart itself looks correct.
id edb0fed6f88b · mined from da-code dacode-plot-line-015@s9
raw text (what the judge reads)
### Required output artifacts (beyond the headline file) not produced or verified
- **Applies when**: `task` -- the task points to a spec/config file (e.g., a YAML/JSON format description) and/or a grader that expects specific persisted result files (image plus serialized data/array/metadata), and the agent must save outputs to disk.
- **Pattern**: The attempt treats the visible deliverable (a single figure or a prose summary) as the whole job: it never opens/echoes the spec file to enumerate every required output, saves only one file, keeps no runnable script, and then asserts success in narrative form ("saved as ...", "follows the specification") without listing or re-reading the files on disk.
- **Detection procedure**:
  1. From the task text and any referenced spec/config file, list every artifact and field the deliverable must contain (file names, extensions, serialized data of the plotted values, title/label/size/color keys, ordering, units).
  2. Read the scripts for the corresponding write calls (e.g., one per required artifact) and check that the values written come from the actual plotted/computed objects rather than being retyped by hand.
  3. Check for a post-hoc verification step: does the attempt re-list the output directory and re-load each saved artifact to confirm existence, shape/length, dtype and key names?
  4. Compare the answer's claims against that verification evidence; a claim of compliance with no dump of the spec's requirements and no file listing is unsupported.
- **Discriminator**: A real violation is missing/unwritten required artifacts or compliance asserted only in prose. It is *not* a violation if the spec genuinely requires a single file and the attempt shows the spec contents plus a load-back check confirming the file's structure — nor is extra descriptive prose a problem when the artifacts are all written and validated.
- **Consequence**: File-level grader checks report the unproduced artifacts as WRONG/MISSING, so the task scores zero even if the rendered chart itself looks correct.
484Deliverable file never written / answer only in prosetaskda-code
Applies when
task -- the task states the result must be written to a specific output file following a provided sample/template format.
Pattern
The agent computes the requested quantity but only prints a narrative report to stdout, never writing (or writing with wrong columns/headers/precision) the required result file, so the graded artifact is missing or malformed.
Detection procedure
1) Read the task for the named output path and any sample/template file defining columns, ordering, and rounding. 2) Search the scripts for an explicit write to that exact path (e.g. a to_csv/file-write call) and confirm the written column names and value formatting match the template. 3) Read the agent's final answer: check it reports the requested statistic in the file's schema rather than only as free-form text or intermediate diagnostics. 4) Flag if no write exists, the path/schema differs from the template, or the reported value is unformatted (e.g. "0.0000" where the template implies a specific precision or notation).
Discriminator
A real violation is a missing/renamed/mis-schema'd output file or prose-only reporting; it is fine if the script writes the correct path and schema and the prose is merely an extra human-readable summary.
Consequence
The grader marks the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying computation was statistically correct.
id 9d84ff5b0d0f · mined from da-code dacode-data-sa-028@s9
raw text (what the judge reads)
### Deliverable file never written / answer only in prose
- **Applies when**: `task` -- the task states the result must be written to a specific output file following a provided sample/template format.
- **Pattern**: The agent computes the requested quantity but only prints a narrative report to stdout, never writing (or writing with wrong columns/headers/precision) the required result file, so the graded artifact is missing or malformed.
- **Detection procedure**: 1) Read the task for the named output path and any sample/template file defining columns, ordering, and rounding. 2) Search the scripts for an explicit write to that exact path (e.g. a `to_csv`/file-write call) and confirm the written column names and value formatting match the template. 3) Read the agent's final answer: check it reports the requested statistic in the file's schema rather than only as free-form text or intermediate diagnostics. 4) Flag if no write exists, the path/schema differs from the template, or the reported value is unformatted (e.g. "0.0000" where the template implies a specific precision or notation).
- **Discriminator**: A real violation is a missing/renamed/mis-schema'd output file or prose-only reporting; it is fine if the script writes the correct path and schema and the prose is merely an extra human-readable summary.
- **Consequence**: The grader marks the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying computation was statistically correct.
485Ignoring a referenced specification file and substituting assumed rulestaskda-code
Applies when
task -- the task instructs the agent to follow a method/definition given in an auxiliary document (README, spec, config, instructions file) rather than defining the rule inline.
Pattern
The scripts never open, print, or parse the referenced document; instead the agent hardcodes categories/bins/thresholds/formats inferred from the raw data's existing values (or from convention) and then claims in the answer that the output "follows" the spec, with no evidence of the spec's actual content.
Detection procedure
  1. From the task text, list every externally referenced file or document that defines how the analysis must be done or how outputs must be saved.
  2. Grep the scripts for reads of each such file (open/read_csv/read_text/print of contents); confirm the derived rule appears as data read from the file, not as a literal list embedded in the code.
  3. Check the answer for any quotation/paraphrase of the spec's actual rules and for the full set of artifacts the spec/task implies; if the spec was never read, treat the categories, ordering, and output files as unverified.
  4. Sanity-check the derived grouping against the raw category values (e.g., are the code's groups just the raw values renamed, implying no aggregation/regrouping was applied?).
Discriminator
Fine if the script loads the spec and programmatically derives the mapping, or explicitly prints the spec's contents and then encodes an identical rule verbatim; violation if the rule's origin is only the agent's inference from the data or from general convention, with the spec never inspected.
Consequence
Group definitions, ordering, or required output artifacts diverge from the specified method, so the saved figure and any companion result files mismatch the expected values and all output checks fail.
id a1c623662fa5 · mined from da-code dacode-plot-bar-005@s9
raw text (what the judge reads)
### Ignoring a referenced specification file and substituting assumed rules
- **Applies when**: `task` -- the task instructs the agent to follow a method/definition given in an auxiliary document (README, spec, config, instructions file) rather than defining the rule inline.
- **Pattern**: The scripts never open, print, or parse the referenced document; instead the agent hardcodes categories/bins/thresholds/formats inferred from the raw data's existing values (or from convention) and then claims in the answer that the output "follows" the spec, with no evidence of the spec's actual content.
- **Detection procedure**:
  1. From the task text, list every externally referenced file or document that defines how the analysis must be done or how outputs must be saved.
  2. Grep the scripts for reads of each such file (open/read_csv/read_text/print of contents); confirm the derived rule appears as data read from the file, not as a literal list embedded in the code.
  3. Check the answer for any quotation/paraphrase of the spec's actual rules and for the full set of artifacts the spec/task implies; if the spec was never read, treat the categories, ordering, and output files as unverified.
  4. Sanity-check the derived grouping against the raw category values (e.g., are the code's groups just the raw values renamed, implying no aggregation/regrouping was applied?).
- **Discriminator**: Fine if the script loads the spec and programmatically derives the mapping, or explicitly prints the spec's contents and then encodes an identical rule verbatim; violation if the rule's origin is only the agent's inference from the data or from general convention, with the spec never inspected.
- **Consequence**: Group definitions, ordering, or required output artifacts diverge from the specified method, so the saved figure and any companion result files mismatch the expected values and all output checks fail.
486Deliverable artifact never produced in the requested schemataskda-code
Applies when
task -- the task specifies an exact answer format (a JSON/CSV structure, key names, list-valued fields) and/or the grading harness looks for a named result file, while the agent's scripts only print exploratory output.
Pattern
The scripts load, clean, and compute the correct-looking quantity but end at print(...); the final answer is retyped by hand into the chat, so no result file is written and the hand-typed structure silently deviates from the requested schema (scalars instead of lists, renamed/extra keys, unrounded or differently-typed values).
Detection procedure
  1. From the task statement, list the required output: file name (if any), top-level keys, and the container type/precision of each value.
  2. Search every script for a write/serialization call (json.dump, to_csv, open(...,'w')) targeting that file; if none exists, the attempt is already inadequate.
  3. Compare the submitted answer's keys, value containers, and numeric formatting character-by-character against the required template.
  4. Confirm the value in the answer is the same object the script actually computed (not a rounded/retyped variant), and that any tie or multi-row case is represented as the schema demands.
Discriminator
A real violation is a missing result artifact or a structural mismatch (e.g., bare string/number where a list is specified, key spelling differing from the template). Not a violation: a script that writes the exact required file and the chat answer merely echoes it, or cosmetic whitespace differences inside a correctly keyed and correctly typed structure.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation was right.
id d04466e34579 · mined from da-code dacode-di-text-002@s9
raw text (what the judge reads)
### Deliverable artifact never produced in the requested schema
- **Applies when**: `task` -- the task specifies an exact answer format (a JSON/CSV structure, key names, list-valued fields) and/or the grading harness looks for a named result file, while the agent's scripts only print exploratory output.
- **Pattern**: The scripts load, clean, and compute the correct-looking quantity but end at `print(...)`; the final answer is retyped by hand into the chat, so no result file is written and the hand-typed structure silently deviates from the requested schema (scalars instead of lists, renamed/extra keys, unrounded or differently-typed values).
- **Detection procedure**:
  1. From the task statement, list the required output: file name (if any), top-level keys, and the container type/precision of each value.
  2. Search every script for a write/serialization call (`json.dump`, `to_csv`, `open(...,'w')`) targeting that file; if none exists, the attempt is already inadequate.
  3. Compare the submitted answer's keys, value containers, and numeric formatting character-by-character against the required template.
  4. Confirm the value in the answer is the same object the script actually computed (not a rounded/retyped variant), and that any tie or multi-row case is represented as the schema demands.
- **Discriminator**: A real violation is a missing result artifact or a structural mismatch (e.g., bare string/number where a list is specified, key spelling differing from the template). Not a violation: a script that writes the exact required file and the chat answer merely echoes it, or cosmetic whitespace differences inside a correctly keyed and correctly typed structure.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation was right.
487Circular "verification" of a derived output whose definition/format was never pinned downtaskda-code
Applies when
task -- the deliverable is a derived series/aggregate (e.g., cumulative or compounded quantities, normalized indices, scores) whose exact convention, column set, or baseline the task states loosely, and the data file itself contains reference/precomputed columns or rows the script ignores.
Pattern
The script picks one plausible convention (e.g., subtracting a base, dropping/keeping a starting row, silently propagating missing values, re-deriving a quantity that the input already provides) and then "verifies" it by recomputing the very same formula, so the check can never fail; no comparison against any independent reference or documented format.
Detection procedure
  1. From the task/README, list every ambiguous choice in the requested output: definition of the derived quantity, whether an initial/base row is included, column names/order, units, rounding, row count.
  2. In the script, check whether the raw input was inspected for pre-existing columns matching the requested outputs and for missing/non-numeric values before any cumulative/aggregating operation; flag if such columns are ignored or NaNs are left to propagate through a running product/sum.
  3. Read the verification script: if it re-derives the same expression and compares it to its own output (or only prints shapes/dtypes), mark the verification as circular — it validates nothing about the convention.
  4. Sanity-check the reported numbers against an independent expectation (e.g., a supplied reference column, a known baseline value at the first period, plausible end-of-period magnitude); flag if no such external anchor exists anywhere.
Discriminator
A genuine violation is when at least one output convention had a checkable external anchor (a column in the file, an explicit README/format statement, or a trivially computable known value) that the agent never used. It is not a violation if the agent explicitly compared its series to an independent reference or to values given in the task and documented the matching convention, even if some choices remained judgment calls.
Consequence
The file has the right shape and passes the agent's own checks, but every value is offset/shifted/contaminated relative to the expected convention, so an exact-match file comparison fails (0 checks passed).
id 46564f991276 · mined from da-code dacode-dm-csv-050@s9
raw text (what the judge reads)
### Circular "verification" of a derived output whose definition/format was never pinned down
- **Applies when**: `task` -- the deliverable is a derived series/aggregate (e.g., cumulative or compounded quantities, normalized indices, scores) whose exact convention, column set, or baseline the task states loosely, and the data file itself contains reference/precomputed columns or rows the script ignores.
- **Pattern**: The script picks one plausible convention (e.g., subtracting a base, dropping/keeping a starting row, silently propagating missing values, re-deriving a quantity that the input already provides) and then "verifies" it by recomputing the very same formula, so the check can never fail; no comparison against any independent reference or documented format.
- **Detection procedure**:
  1. From the task/README, list every ambiguous choice in the requested output: definition of the derived quantity, whether an initial/base row is included, column names/order, units, rounding, row count.
  2. In the script, check whether the raw input was inspected for pre-existing columns matching the requested outputs and for missing/non-numeric values before any cumulative/aggregating operation; flag if such columns are ignored or NaNs are left to propagate through a running product/sum.
  3. Read the verification script: if it re-derives the same expression and compares it to its own output (or only prints shapes/dtypes), mark the verification as circular — it validates nothing about the convention.
  4. Sanity-check the reported numbers against an independent expectation (e.g., a supplied reference column, a known baseline value at the first period, plausible end-of-period magnitude); flag if no such external anchor exists anywhere.
- **Discriminator**: A genuine violation is when at least one output convention had a checkable external anchor (a column in the file, an explicit README/format statement, or a trivially computable known value) that the agent never used. It is not a violation if the agent explicitly compared its series to an independent reference or to values given in the task and documented the matching convention, even if some choices remained judgment calls.
- **Consequence**: The file has the right shape and passes the agent's own checks, but every value is offset/shifted/contaminated relative to the expected convention, so an exact-match file comparison fails (0 checks passed).
488Prescribed test run on a silently altered / inconsistent data subset, with no sanity check on the values fed intaskinfiagent-dabench
Applies when
task -- the task names a specific statistical test or statistic to compute on a column, and the scripts load the raw column and pass it straight to the test (optionally sampling, truncating, or dropping rows on their own initiative).
Pattern
The attempt changes the input to the mandated test without authorization (e.g., random subsampling to satisfy an implementation limit, dropping/keeping rows by an unstated rule), computes companion statistics on a different subset than the test, and never inspects the column for sentinel codes, placeholder values, dtype coercion artifacts, or extreme-tail contamination that would explain a diagnostic value wildly out of line with the rest of the distribution.
Detection procedure
  1. Read the task and list exactly what data the requested test/statistics are supposed to be computed on (all rows of the column, any stated filter, alpha, rounding).
  2. In the scripts, trace the exact array passed to each call and check whether all reported numbers come from the same, unmodified, task-sanctioned subset; flag any sample(), head/tail truncation, row filter, or per-script divergence (one script tests a sample, another tests everything) that is never reconciled.
  3. Check whether any script prints and reacts to basic value diagnostics (min/max, value counts, count of zeros/negatives/sentinels like -1/9999, histogram) — and whether an implausible diagnostic (very large |skew|, huge kurtosis, min far from the bulk) was investigated rather than accepted.
  4. Compare the final reported answer against the script that actually followed the task's specification; if the answer comes from a modified-input run, or if two runs would disagree, mark inadequate.
Discriminator
A real violation is unauthorized or inconsistent input selection and/or blind acceptance of contradictory diagnostics; it is fine if the subset change is explicitly required by the task, is applied identically to every reported statistic, and the script documents a check that the retained values are the legitimate ones (no sentinels/placeholders masquerading as data).
Consequence
The reported test decision (and skew/kurtosis) reflects a different population than the graded one, so the categorical verdict flips and the numeric values miss the expected key, giving 0 credit despite a plausible-looking run.
id b38d40dd7f5a · mined from infiagent-dabench dabench-298@s9
raw text (what the judge reads)
### Prescribed test run on a silently altered / inconsistent data subset, with no sanity check on the values fed in
- **Applies when**: `task` -- the task names a specific statistical test or statistic to compute on a column, and the scripts load the raw column and pass it straight to the test (optionally sampling, truncating, or dropping rows on their own initiative).
- **Pattern**: The attempt changes the input to the mandated test without authorization (e.g., random subsampling to satisfy an implementation limit, dropping/keeping rows by an unstated rule), computes companion statistics on a *different* subset than the test, and never inspects the column for sentinel codes, placeholder values, dtype coercion artifacts, or extreme-tail contamination that would explain a diagnostic value wildly out of line with the rest of the distribution.
- **Detection procedure**:
  1. Read the task and list exactly what data the requested test/statistics are supposed to be computed on (all rows of the column, any stated filter, alpha, rounding).
  2. In the scripts, trace the exact array passed to each call and check whether all reported numbers come from the *same*, unmodified, task-sanctioned subset; flag any `sample()`, head/tail truncation, row filter, or per-script divergence (one script tests a sample, another tests everything) that is never reconciled.
  3. Check whether any script prints and reacts to basic value diagnostics (min/max, value counts, count of zeros/negatives/sentinels like -1/9999, histogram) — and whether an implausible diagnostic (very large |skew|, huge kurtosis, min far from the bulk) was investigated rather than accepted.
  4. Compare the final reported answer against the script that actually followed the task's specification; if the answer comes from a modified-input run, or if two runs would disagree, mark inadequate.
- **Discriminator**: A real violation is unauthorized or inconsistent input selection and/or blind acceptance of contradictory diagnostics; it is *fine* if the subset change is explicitly required by the task, is applied identically to every reported statistic, and the script documents a check that the retained values are the legitimate ones (no sentinels/placeholders masquerading as data).
- **Consequence**: The reported test decision (and skew/kurtosis) reflects a different population than the graded one, so the categorical verdict flips and the numeric values miss the expected key, giving 0 credit despite a plausible-looking run.
489Accepting the first model's weak validation score without baseline comparison or iterationtaskda-code
Applies when
task -- The task asks for a predictive model whose output quality is graded (e.g., a held-out score threshold), and the scripts fit one off-the-shelf model with arbitrary hyperparameters and immediately write the submission.
Pattern
The agent trains a single default/under-configured learner, observes a mediocre validation metric (low R²/AUC, or predictions with far smaller spread than the training target), never compares it to a trivial baseline (mean/linear model/simple feature aggregate) or an alternative algorithm, and declares "task completed successfully" because the file format checks pass. Format validation is mistaken for result validation.
Detection procedure
  1. Read the task to confirm the deliverable is judged on predictive quality, not just file existence/format.
  2. In the scripts, count how many candidate models/hyperparameter settings and how many baselines are evaluated; check whether any capacity/regularization tuning, feature engineering, or cross-validation is done.
  3. Compare the reported validation metric and prediction distribution against the training target distribution: is the metric weak in absolute terms, and is the predicted spread/std materially compressed relative to the target's?
  4. Check whether the final answer's "success" claim rests only on shape/ID/range checks rather than on any evidence the score is competitive.
Discriminator
A real violation is a single unbenchmarked model with a clearly weak or unexplained score and no attempt at improvement; it is not a violation if the agent established a baseline and/or tried multiple models and the chosen one is demonstrably the best available, even if the absolute metric is modest because the signal is genuinely limited (shown by the baseline comparison).
Consequence
The submission is well-formed but scores below the grader's accuracy threshold, so the file is judged WRONG despite passing all self-run format checks.
id acc0a77a34d9 · mined from da-code dacode-ml-competition-008@s9
raw text (what the judge reads)
### Accepting the first model's weak validation score without baseline comparison or iteration
- **Applies when**: `task` -- The task asks for a *predictive* model whose output quality is graded (e.g., a held-out score threshold), and the scripts fit one off-the-shelf model with arbitrary hyperparameters and immediately write the submission.
- **Pattern**: The agent trains a single default/under-configured learner, observes a mediocre validation metric (low R²/AUC, or predictions with far smaller spread than the training target), never compares it to a trivial baseline (mean/linear model/simple feature aggregate) or an alternative algorithm, and declares "task completed successfully" because the file format checks pass. Format validation is mistaken for result validation.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is judged on predictive quality, not just file existence/format.
  2. In the scripts, count how many candidate models/hyperparameter settings and how many baselines are evaluated; check whether any capacity/regularization tuning, feature engineering, or cross-validation is done.
  3. Compare the reported validation metric and prediction distribution against the training target distribution: is the metric weak in absolute terms, and is the predicted spread/std materially compressed relative to the target's?
  4. Check whether the final answer's "success" claim rests only on shape/ID/range checks rather than on any evidence the score is competitive.
- **Discriminator**: A real violation is a single unbenchmarked model with a clearly weak or unexplained score and no attempt at improvement; it is *not* a violation if the agent established a baseline and/or tried multiple models and the chosen one is demonstrably the best available, even if the absolute metric is modest because the signal is genuinely limited (shown by the baseline comparison).
- **Consequence**: The submission is well-formed but scores below the grader's accuracy threshold, so the file is judged WRONG despite passing all self-run format checks.
490Ranking definitions, eligibility filters, and output schema assumed instead of read from the provided spectaskda-code
Applies when
task -- the task asks for "top-N" rankings of grouped entities, supplies a definition/README (possibly truncated) and a sample output file, and the scripts compute the rankings with hard-coded aggregation rules and column names.
Pattern
The attempt invents the operational details — which aggregation to use per metric (mean vs. sum vs. max), one qualification threshold applied uniformly to every ranking, tie-breaking, rounding, sort order — and invents the output header/index column, without ever loading or printing the supplied sample/spec files to confirm. Ties at the metric's ceiling are then silently broken by group order (e.g., alphabetical), so the "top" list is arbitrary.
Detection procedure
  1. In the task/README, list every stated or implied definitional constraint (qualification criteria, aggregation per metric, N, ordering, format) and note any that are incomplete or defined in a separate spec/sample file.
  2. In the scripts, check that each such constraint is read from the file (e.g., the sample header is loaded and reused, the threshold/definition is quoted from the spec) rather than typed in as a literal; flag any literal that has no cited source.
  3. Check whether the metric's top values are saturated/tied (e.g., many groups at the max rating) and whether the script defines an explicit deterministic tie-breaker; a top-N list that reads in alphabetical order is evidence of arbitrary selection.
  4. Compare the produced header, column names, ordering, and row count against the sample file byte-for-byte in intent; flag any extra/renamed columns.
Discriminator
A real violation is when the chosen aggregation, filter, tie-break, or header could plausibly differ from the supplied spec and no script output demonstrates agreement with it. It is fine if the script explicitly reads the sample/spec, echoes the definition it implements, and shows that the top-N is unambiguous (no ties at the cutoff) or applies the spec's stated tie-breaker.
Consequence
The file exists and looks plausible, but the grader's exact-match on member sets/order/header fails — wrong entities enter the ranking (or wrong ones win ties) and the schema mismatches, yielding 0 passed checks.
id d8e9d611a460 · mined from da-code dacode-dm-csv-009@s9
raw text (what the judge reads)
### Ranking definitions, eligibility filters, and output schema assumed instead of read from the provided spec
- **Applies when**: `task` -- the task asks for "top-N" rankings of grouped entities, supplies a definition/README (possibly truncated) and a sample output file, and the scripts compute the rankings with hard-coded aggregation rules and column names.
- **Pattern**: The attempt invents the operational details — which aggregation to use per metric (mean vs. sum vs. max), one qualification threshold applied uniformly to every ranking, tie-breaking, rounding, sort order — and invents the output header/index column, without ever loading or printing the supplied sample/spec files to confirm. Ties at the metric's ceiling are then silently broken by group order (e.g., alphabetical), so the "top" list is arbitrary.
- **Detection procedure**:
  1. In the task/README, list every stated or implied definitional constraint (qualification criteria, aggregation per metric, N, ordering, format) and note any that are incomplete or defined in a separate spec/sample file.
  2. In the scripts, check that each such constraint is read from the file (e.g., the sample header is loaded and reused, the threshold/definition is quoted from the spec) rather than typed in as a literal; flag any literal that has no cited source.
  3. Check whether the metric's top values are saturated/tied (e.g., many groups at the max rating) and whether the script defines an explicit deterministic tie-breaker; a top-N list that reads in alphabetical order is evidence of arbitrary selection.
  4. Compare the produced header, column names, ordering, and row count against the sample file byte-for-byte in intent; flag any extra/renamed columns.
- **Discriminator**: A real violation is when the chosen aggregation, filter, tie-break, or header could plausibly differ from the supplied spec and no script output demonstrates agreement with it. It is fine if the script explicitly reads the sample/spec, echoes the definition it implements, and shows that the top-N is unambiguous (no ties at the cutoff) or applies the spec's stated tie-breaker.
- **Consequence**: The file exists and looks plausible, but the grader's exact-match on member sets/order/header fails — wrong entities enter the ranking (or wrong ones win ties) and the schema mismatches, yielding 0 passed checks.
491Fabricating or substituting input data instead of using the provided datasettaskda-code
Applies when
task -- the task references a specific provided dataset/config files and the scripts show trouble locating or loading them.
Pattern
After failing to find/read the real input, the agent synthesizes a random or placeholder dataset (or downloads an unrelated source) and runs the full analysis on it, then reports the resulting numbers and artifacts as if they came from the real data.
Detection procedure
1) List the input paths the task/README implies and the output artifacts it requires. 2) Scan scripts for data creation with random generators, hardcoded fake rows, or reads from a path different from the provided one; also check whether every required output artifact is written. 3) Compare row/record counts and category counts in the answer against the documented scale of the real dataset (e.g., README's stated size). 4) Flag if the analyzed data is self-generated or the counts are implausibly small/round, or if required outputs are missing.
Discriminator
Legitimate: synthetic data used only in a throwaway smoke test while the final reported run reads the real file and produces all required artifacts. Violation: the final reported numbers/plots derive from generated or substituted data, or the agent never resolved the real input path.
Consequence
All value/artifact checks fail because the computed distribution and saved files reflect invented data, and expected output files are absent or mismatched.
id 8d26f54c7134 · mined from da-code dacode-plot-bar-007@s9
raw text (what the judge reads)
### Fabricating or substituting input data instead of using the provided dataset
- **Applies when**: `task` -- the task references a specific provided dataset/config files and the scripts show trouble locating or loading them.
- **Pattern**: After failing to find/read the real input, the agent synthesizes a random or placeholder dataset (or downloads an unrelated source) and runs the full analysis on it, then reports the resulting numbers and artifacts as if they came from the real data.
- **Detection procedure**: 1) List the input paths the task/README implies and the output artifacts it requires. 2) Scan scripts for data creation with random generators, hardcoded fake rows, or reads from a path different from the provided one; also check whether every required output artifact is written. 3) Compare row/record counts and category counts in the answer against the documented scale of the real dataset (e.g., README's stated size). 4) Flag if the analyzed data is self-generated or the counts are implausibly small/round, or if required outputs are missing.
- **Discriminator**: Legitimate: synthetic data used only in a throwaway smoke test while the final reported run reads the real file and produces all required artifacts. Violation: the final reported numbers/plots derive from generated or substituted data, or the agent never resolved the real input path.
- **Consequence**: All value/artifact checks fail because the computed distribution and saved files reflect invented data, and expected output files are absent or mismatched.
492Degenerate (empty/undefined) results reported in an ad-hoc string instead of the required formattaskinfiagent-dabench
Applies when
task -- the task prescribes a strict answer template (e.g. a float rounded to N decimals) and the requested statistic may be undefined because the chained filters leave no rows (or only null values).
Pattern
The agent detects that the computation yields nothing and hand-writes a special token ("NaN", "None", "N/A", "empty", a message) whose spelling/case/type differs from what the harness compares against, and it does so without a saved, re-runnable script that shows the filter counts and the exact printed answer string.
Detection procedure
  1. Read the task and note the exact answer template, including type, rounding, and how a missing/undefined value would have to be rendered.
  2. Read the scripts: confirm they exist, that they print the row count after each filtering step, and that the final printed token is produced programmatically by the same formatting path in both the normal and the empty/all-null branch (e.g. f"@x[{value:.2f}]" or a single canonical literal).
  3. Compare the submitted string character-by-character with that template: check casing, spelling, quoting, and whether rounding/decimal places were applied.
  4. If no script is saved, or the special-case token is typed by hand rather than emitted by code, flag the attempt as unverifiable.
Discriminator
A real violation is a token that differs in case/spelling/type from the canonical representation the task or grader expects, or one that cannot be traced to script output; it is not a violation if the empty result is genuinely correct and the agent emitted the canonical lowercase/numeric-literal form produced by the same code path that prints normal results.
Consequence
The grader's string/value comparison fails even though the underlying computation was right, scoring 0/1 with "WRONG/MISSING" despite the submitted value matching semantically.
id 6421fe48cc19 · mined from infiagent-dabench dabench-554@s9
raw text (what the judge reads)
### Degenerate (empty/undefined) results reported in an ad-hoc string instead of the required format
- **Applies when**: `task` -- the task prescribes a strict answer template (e.g. a float rounded to N decimals) and the requested statistic may be undefined because the chained filters leave no rows (or only null values).
- **Pattern**: The agent detects that the computation yields nothing and hand-writes a special token ("NaN", "None", "N/A", "empty", a message) whose spelling/case/type differs from what the harness compares against, and it does so without a saved, re-runnable script that shows the filter counts and the exact printed answer string.
- **Detection procedure**:
  1. Read the task and note the exact answer template, including type, rounding, and how a missing/undefined value would have to be rendered.
  2. Read the scripts: confirm they exist, that they print the row count after each filtering step, and that the final printed token is produced programmatically by the same formatting path in both the normal and the empty/all-null branch (e.g. `f"@x[{value:.2f}]"` or a single canonical literal).
  3. Compare the submitted string character-by-character with that template: check casing, spelling, quoting, and whether rounding/decimal places were applied.
  4. If no script is saved, or the special-case token is typed by hand rather than emitted by code, flag the attempt as unverifiable.
- **Discriminator**: A real violation is a token that differs in case/spelling/type from the canonical representation the task or grader expects, or one that cannot be traced to script output; it is *not* a violation if the empty result is genuinely correct and the agent emitted the canonical lowercase/numeric-literal form produced by the same code path that prints normal results.
- **Consequence**: The grader's string/value comparison fails even though the underlying computation was right, scoring 0/1 with "WRONG/MISSING" despite the submitted value matching semantically.
493Stated ordering constraint not applied to every requested output grouptaskda-code
Applies when
task -- the task asks for several ranked lists/groups and states a single global ordering (or rounding/format) rule, and the script builds each group with its own sorting helper.
Pattern
The agent applies the stated ordering to the "obvious" group but leaves the other group in whatever order its extraction function naturally produces (e.g. an ascending "smallest-n" selection left ascending while the instruction demanded descending), and/or writes the result to a file name/path other than the one requested — so values are right but sequence/format is wrong.
Detection procedure
  1. Read the task and write down every explicit output constraint: ordering direction, how many items, rounding/units, JSON key names, and the exact output file name/path — noting whether each constraint is scoped to one group or to all of them.
  2. In the script, locate each list-building step and check the sort/selection direction actually used for that list, plus any final reordering step; confirm the direction matches the stated rule for that list.
  3. Check the write step: does it emit the requested file name and the requested keys/structure?
  4. Inspect the reported answer and verify each list's element order is monotonic in the required direction (compare the underlying values printed by the script, not just names).
Discriminator
A real violation is when a group's order (or file/format) contradicts an explicitly stated instruction. It is not a violation if the task genuinely specified different orderings per group, if the task left ordering unspecified, or if ties in the values make two orderings equally valid.
Consequence
The grader compares the expected artifact element-by-element (or fails to find it at all), so a list with correct membership but reversed order — or an answer written outside the expected file — is scored wrong, giving 0 despite correct computation.
id a9f0562b68d0 · mined from da-code dacode-di-text-003@s9
raw text (what the judge reads)
### Stated ordering constraint not applied to every requested output group
- **Applies when**: `task` -- the task asks for several ranked lists/groups and states a single global ordering (or rounding/format) rule, and the script builds each group with its own sorting helper.
- **Pattern**: The agent applies the stated ordering to the "obvious" group but leaves the other group in whatever order its extraction function naturally produces (e.g. an ascending "smallest-n" selection left ascending while the instruction demanded descending), and/or writes the result to a file name/path other than the one requested — so values are right but sequence/format is wrong.
- **Detection procedure**:
  1. Read the task and write down every explicit output constraint: ordering direction, how many items, rounding/units, JSON key names, and the exact output file name/path — noting whether each constraint is scoped to one group or to all of them.
  2. In the script, locate each list-building step and check the sort/selection direction actually used for that list, plus any final reordering step; confirm the direction matches the stated rule for *that* list.
  3. Check the write step: does it emit the requested file name and the requested keys/structure?
  4. Inspect the reported answer and verify each list's element order is monotonic in the required direction (compare the underlying values printed by the script, not just names).
- **Discriminator**: A real violation is when a group's order (or file/format) contradicts an explicitly stated instruction. It is *not* a violation if the task genuinely specified different orderings per group, if the task left ordering unspecified, or if ties in the values make two orderings equally valid.
- **Consequence**: The grader compares the expected artifact element-by-element (or fails to find it at all), so a list with correct membership but reversed order — or an answer written outside the expected file — is scored wrong, giving 0 despite correct computation.
494Answer-format token fidelity (quoting/delimiters applied inconsistently across fields)taskinfiagent-dabench
Applies when
task -- the task specifies a literal answer template with tagged fields and shows how each value should be written (e.g. quoted strings, units, brackets), and the agent must emit that template verbatim.
Pattern
The agent computes the right values but serializes them in a form that deviates from the specified template for some fields — dropping quotation marks around string values, changing delimiters, adding/removing spaces, or altering capitalization — often inconsistently (one field quoted, others not), so an exact-match grader fails those fields even though the analysis was correct.
Detection procedure
  1. Copy the answer template and the example values shown in the task, and list for each field the exact expected lexical form (quoted vs bare, allowed vocabulary, punctuation).
  2. Read the agent's final answer and compare each field token-by-token against that expected form, not just semantically.
  3. Flag any field whose delimiters/quoting/spelling differ from the template, and especially flag internal inconsistency (the same value type formatted differently across fields).
  4. Check that a script or printed output exists that produces the answer string in the required form; if no reproducible artifact exists, the format cannot be verified and the attempt is inadequate.
Discriminator
A real violation is a deviation in the literal characters of a field the task specified (missing quotes, wrong separator, different casing/wording than the allowed options). A look-alike that is fine is a formatting choice the task left unspecified (e.g. whitespace between tags, ordering of fields when order is not mandated) or a value that is semantically and lexically among the permitted options.
Consequence
An exact-match grader marks the affected fields WRONG/MISSING despite correct underlying computations, yielding a partial score (here 1/3) and an overall incorrect verdict.
id 56eb9ae42b29 · mined from infiagent-dabench dabench-550@s9
raw text (what the judge reads)
### Answer-format token fidelity (quoting/delimiters applied inconsistently across fields)
- **Applies when**: `task` -- the task specifies a literal answer template with tagged fields and shows how each value should be written (e.g. quoted strings, units, brackets), and the agent must emit that template verbatim.
- **Pattern**: The agent computes the right values but serializes them in a form that deviates from the specified template for some fields — dropping quotation marks around string values, changing delimiters, adding/removing spaces, or altering capitalization — often inconsistently (one field quoted, others not), so an exact-match grader fails those fields even though the analysis was correct.
- **Detection procedure**:
  1. Copy the answer template and the example values shown in the task, and list for each field the exact expected lexical form (quoted vs bare, allowed vocabulary, punctuation).
  2. Read the agent's final answer and compare each field token-by-token against that expected form, not just semantically.
  3. Flag any field whose delimiters/quoting/spelling differ from the template, and especially flag internal inconsistency (the same value type formatted differently across fields).
  4. Check that a script or printed output exists that produces the answer string in the required form; if no reproducible artifact exists, the format cannot be verified and the attempt is inadequate.
- **Discriminator**: A real violation is a deviation in the literal characters of a field the task specified (missing quotes, wrong separator, different casing/wording than the allowed options). A look-alike that is fine is a formatting choice the task left unspecified (e.g. whitespace between tags, ordering of fields when order is not mandated) or a value that is semantically and lexically among the permitted options.
- **Consequence**: An exact-match grader marks the affected fields WRONG/MISSING despite correct underlying computations, yielding a partial score (here 1/3) and an overall incorrect verdict.
495Entity mismatch between requested analysis and the data actually usedtaskda-code
Applies when
task -- the task names specific entities, metrics, or groupings (e.g., a category dimension, a duration measure, a config file's settings) that must exist in the supplied inputs, and the agent must load those inputs and emit the required output files.
Pattern
The agent cannot find the named fields, silently substitutes whatever unrelated columns/files are at hand, and reports a chart built from those substitutes while reusing the requested title/labels — producing output that matches the task only in wording, not in content. It also skips required auxiliary output artifacts.
Detection procedure
1. List from the task every required input (config file, source table), every required quantity (grouping key, per-stage measure, aggregation), and every required output file. 2. Read the scripts/answer to see which files and columns were actually read and which outputs were written; map each required item to a concrete column or state that it was absent. 3. Flag if any required quantity is unmatched, or if the reported category values/units are semantically inconsistent with the requested ones (e.g., group labels or measure units that could not possibly be the requested entity), or if a listed output file is never created. 4. Check the answer's own internal consistency: axis labels, legend entries, and the stated metric should describe the requested quantity, not something else.
Discriminator
A real violation is substituting different semantic content (different grouping entity, different measure, unrelated source file) while keeping the requested labels; a look-alike that is fine is using a differently named but semantically equivalent column, or a reasonable disambiguation among several candidate columns that all encode the requested quantity, stated explicitly.
Consequence
All expected artifacts fail comparison — the saved figure, the numeric array, and the config-derived metadata encode the wrong entities and values, so every automated check returns mismatch (0 passed), and missing required files count as absent outright.
id 0c18c3141601 · mined from da-code dacode-plot-scatter-002@s9
raw text (what the judge reads)
### Entity mismatch between requested analysis and the data actually used
- **Applies when**: `task` -- the task names specific entities, metrics, or groupings (e.g., a category dimension, a duration measure, a config file's settings) that must exist in the supplied inputs, and the agent must load those inputs and emit the required output files.
- **Pattern**: The agent cannot find the named fields, silently substitutes whatever unrelated columns/files are at hand, and reports a chart built from those substitutes while reusing the requested title/labels — producing output that matches the task only in wording, not in content. It also skips required auxiliary output artifacts.
- **Detection procedure**: 1. List from the task every required input (config file, source table), every required quantity (grouping key, per-stage measure, aggregation), and every required output file. 2. Read the scripts/answer to see which files and columns were actually read and which outputs were written; map each required item to a concrete column or state that it was absent. 3. Flag if any required quantity is unmatched, or if the reported category values/units are semantically inconsistent with the requested ones (e.g., group labels or measure units that could not possibly be the requested entity), or if a listed output file is never created. 4. Check the answer's own internal consistency: axis labels, legend entries, and the stated metric should describe the requested quantity, not something else.
- **Discriminator**: A real violation is substituting different semantic content (different grouping entity, different measure, unrelated source file) while keeping the requested labels; a look-alike that is fine is using a differently *named* but semantically equivalent column, or a reasonable disambiguation among several candidate columns that all encode the requested quantity, stated explicitly.
- **Consequence**: All expected artifacts fail comparison — the saved figure, the numeric array, and the config-derived metadata encode the wrong entities and values, so every automated check returns mismatch (0 passed), and missing required files count as absent outright.
496Output artifact name, location, and format not matched to the specificationtaskda-code
Applies when
task -- The task names an exact output file (and/or supplies a template showing column names, value scaling, rounding, and index layout) that the script must produce.
Pattern
The script computes a plausible result but writes it under a different filename/spelling/path than requested, and/or emits values in a different representation than the template (raw proportions vs. percentages, unrounded vs. rounded, missing/extra index column, different header labels or ordering) — no comparison against the template is ever performed.
Detection procedure
  1. Read the task statement and note the literal output filename (including any unusual spelling), the required directory, and every stated formatting constraint or template.
  2. In the script, find each write call (to_csv, to_json, savefig, …) and compare the literal path/filename string character-by-character with the task's, and check index=/header=/column naming choices.
  3. Check whether the script ever loads or prints the provided template/sample output and asserts that its own result has matching shape, headers, dtype and value scale/rounding; absence of any such check is a red flag.
  4. Inspect the reported answer: confirm the value magnitudes, decimal precision, and column set match the template rather than the raw computation defaults.
Discriminator
A real violation is any deviation in the graded artifact's name/path or in template-visible formatting (scale, rounding, headers, index, ordering). Not a violation if the script writes the exact requested filename to the expected location and its formatting demonstrably follows the template, even if intermediate debug files with other names are also produced.
Consequence
The grader finds the expected file missing or its contents non-matching and marks the result WRONG/MISSING even when the underlying computation is arithmetically correct.
id e876015d8086 · mined from da-code dacode-dm-csv-043@s9
raw text (what the judge reads)
### Output artifact name, location, and format not matched to the specification
- **Applies when**: `task` -- The task names an exact output file (and/or supplies a template showing column names, value scaling, rounding, and index layout) that the script must produce.
- **Pattern**: The script computes a plausible result but writes it under a different filename/spelling/path than requested, and/or emits values in a different representation than the template (raw proportions vs. percentages, unrounded vs. rounded, missing/extra index column, different header labels or ordering) — no comparison against the template is ever performed.
- **Detection procedure**:
  1. Read the task statement and note the literal output filename (including any unusual spelling), the required directory, and every stated formatting constraint or template.
  2. In the script, find each write call (`to_csv`, `to_json`, `savefig`, …) and compare the literal path/filename string character-by-character with the task's, and check `index=`/`header=`/column naming choices.
  3. Check whether the script ever loads or prints the provided template/sample output and asserts that its own result has matching shape, headers, dtype and value scale/rounding; absence of any such check is a red flag.
  4. Inspect the reported answer: confirm the value magnitudes, decimal precision, and column set match the template rather than the raw computation defaults.
- **Discriminator**: A real violation is any deviation in the graded artifact's name/path or in template-visible formatting (scale, rounding, headers, index, ordering). Not a violation if the script writes the exact requested filename to the expected location and its formatting demonstrably follows the template, even if intermediate debug files with other names are also produced.
- **Consequence**: The grader finds the expected file missing or its contents non-matching and marks the result WRONG/MISSING even when the underlying computation is arithmetically correct.
497Predictions that collapse to the target mean (near-zero validation skill) accepted without a baseline or distribution checktaskda-code
Applies when
task -- a task asks for predicted values on a held-out set, and the scripts fit models on a subset of "easy" numeric columns and report validation scores.
Pattern
The attempt reports an explained-variance/accuracy barely above chance, produces predictions whose spread is a small fraction of the training target's spread (all values hugging the mean), and declares success without comparing to a mean/median baseline or investigating unused high-signal columns (identifiers, categorical entities, dates/eras, text fields) that were silently dropped for being non-numeric.
Detection procedure
  1. From the task/README, list all available columns and note which the script actually feeds to the model; flag any dropped categorical, temporal, or entity columns that plausibly carry most of the signal.
  2. In the script, check whether any baseline (predict-the-mean/majority) is computed and whether the reported metric is compared to it; check whether validation split, features, and preprocessing match those used at prediction time.
  3. In the answer, compare the reported prediction min/max/std against the training target's min/max/std; a std that is a small fraction of the target std, or a range far inside the target range, indicates mean-collapse.
  4. Confirm the reported skill metric is materially better than the baseline; if it is ~0, the deliverable is essentially a constant and the attempt is inadequate regardless of file format.
Discriminator
A genuinely low-signal problem is acceptable only if the script demonstrates it — baseline comparison, attempts to encode/use the dropped columns, tuning — and reports the limitation; a violation is silently restricting to a convenient feature subset, never benchmarking, and presenting near-zero skill as "successfully trained".
Consequence
The saved predictions are nearly constant and uncorrelated with the true targets, so any correlation/error-threshold check by the grader fails even though the file has the right shape and column name.
id ec3a3ad3d38b · mined from da-code dacode-ml-regression-004@s9
raw text (what the judge reads)
### Predictions that collapse to the target mean (near-zero validation skill) accepted without a baseline or distribution check
- **Applies when**: `task` -- a task asks for predicted values on a held-out set, and the scripts fit models on a subset of "easy" numeric columns and report validation scores.
- **Pattern**: The attempt reports an explained-variance/accuracy barely above chance, produces predictions whose spread is a small fraction of the training target's spread (all values hugging the mean), and declares success without comparing to a mean/median baseline or investigating unused high-signal columns (identifiers, categorical entities, dates/eras, text fields) that were silently dropped for being non-numeric.
- **Detection procedure**:
  1. From the task/README, list all available columns and note which the script actually feeds to the model; flag any dropped categorical, temporal, or entity columns that plausibly carry most of the signal.
  2. In the script, check whether any baseline (predict-the-mean/majority) is computed and whether the reported metric is compared to it; check whether validation split, features, and preprocessing match those used at prediction time.
  3. In the answer, compare the reported prediction min/max/std against the training target's min/max/std; a std that is a small fraction of the target std, or a range far inside the target range, indicates mean-collapse.
  4. Confirm the reported skill metric is materially better than the baseline; if it is ~0, the deliverable is essentially a constant and the attempt is inadequate regardless of file format.
- **Discriminator**: A genuinely low-signal problem is acceptable only if the script demonstrates it — baseline comparison, attempts to encode/use the dropped columns, tuning — and reports the limitation; a violation is silently restricting to a convenient feature subset, never benchmarking, and presenting near-zero skill as "successfully trained".
- **Consequence**: The saved predictions are nearly constant and uncorrelated with the true targets, so any correlation/error-threshold check by the grader fails even though the file has the right shape and column name.
498Selecting the cluster count by a single automatic score without sanity-checking for degenerate clusterstaskda-code
Applies when
task -- the task asks for an "appropriate" number of groups/components and the script picks it by argmax/elbow of one internal metric over a range, then writes the resulting labels to the deliverable file.
Pattern
The script sweeps k, takes the best score, and immediately fits and exports labels with no check that the chosen partition is substantively meaningful; because unscaled outliers and heavy-tailed features dominate the distance metric, the "best" k produces singleton or near-singleton groups plus one giant group, and this obviously unbalanced solution is reported as final without ever comparing it to a more parsimonious, better-balanced alternative.
Detection procedure
  1. Read the task for the deliverable's intended use (here: an interpretable grouping) and note whether it constrains or hints at the granularity of the grouping.
  2. In the script, check whether the chosen hyperparameter is taken directly from a single criterion's argmax and whether any post-fit validation exists (group sizes, stability across seeds, agreement between elbow and silhouette, outlier handling/robust scaling or transformation of skewed features).
  3. In the reported output, inspect the group-size distribution and per-group summaries; flag if any group has ~1-3 members or one group holds ~half the data while others are trivial.
  4. Confirm the exported file's feature columns are exactly the vectors actually clustered (same rows/order/transform) rather than a different representation than the model saw.
Discriminator
A real violation is an unvalidated argmax whose partition is degenerate (singleton groups, no seed/criterion cross-check, no outlier or scale diagnostics). It is fine if the script justifies k with two or more converging diagnostics, reports balanced and interpretable groups, or explicitly examines and defends small groups as genuine outlier clusters after robustness checks.
Consequence
The saved labels differ structurally from any reasonable reference partition (wrong number of clusters, singleton groups, mismatched row/column content), so a file-level comparison of cluster.csv fails and the downstream interpretation of "which group needs aid" is unstable.
id aed15156c03a · mined from da-code dacode-ml-cluster-013@s9
raw text (what the judge reads)
### Selecting the cluster count by a single automatic score without sanity-checking for degenerate clusters
- **Applies when**: `task` -- the task asks for an "appropriate" number of groups/components and the script picks it by argmax/elbow of one internal metric over a range, then writes the resulting labels to the deliverable file.
- **Pattern**: The script sweeps k, takes the best score, and immediately fits and exports labels with no check that the chosen partition is substantively meaningful; because unscaled outliers and heavy-tailed features dominate the distance metric, the "best" k produces singleton or near-singleton groups plus one giant group, and this obviously unbalanced solution is reported as final without ever comparing it to a more parsimonious, better-balanced alternative.
- **Detection procedure**:
  1. Read the task for the deliverable's intended use (here: an interpretable grouping) and note whether it constrains or hints at the granularity of the grouping.
  2. In the script, check whether the chosen hyperparameter is taken directly from a single criterion's argmax and whether any post-fit validation exists (group sizes, stability across seeds, agreement between elbow and silhouette, outlier handling/robust scaling or transformation of skewed features).
  3. In the reported output, inspect the group-size distribution and per-group summaries; flag if any group has ~1-3 members or one group holds ~half the data while others are trivial.
  4. Confirm the exported file's feature columns are exactly the vectors actually clustered (same rows/order/transform) rather than a different representation than the model saw.
- **Discriminator**: A real violation is an unvalidated argmax whose partition is degenerate (singleton groups, no seed/criterion cross-check, no outlier or scale diagnostics). It is fine if the script justifies k with two or more converging diagnostics, reports balanced and interpretable groups, or explicitly examines and defends small groups as genuine outlier clusters after robustness checks.
- **Consequence**: The saved labels differ structurally from any reasonable reference partition (wrong number of clusters, singleton groups, mismatched row/column content), so a file-level comparison of cluster.csv fails and the downstream interpretation of "which group needs aid" is unstable.
499Silently redefining the analysis scope and mismapping a stated statistic definition to library optionstaskinfiagent-dabench
Applies when
task -- the task names a specific statistic with an explicit definition/adjustment and a specific slice of the data, and the scripts compute it via a library call with option flags over a self-chosen set of columns/rows.
Pattern
Faced with wording it finds ambiguous, the attempt invents a restricted window of the data (e.g., "all periods up to and including the named one") without justification, and passes a flag whose meaning it asserts in a comment rather than verifies against the library docs — so both the input set and the estimator variant differ from what was requested, while the code looks tidy and self-consistent.
Detection procedure
  1. Read the task and write down (a) the exact data slice requested and (b) the exact estimator variant/adjustment named.
  2. In the scripts, locate the array actually passed to the statistic and check whether its construction matches (a) exactly, or whether columns/rows were dropped/truncated by an interpretive choice; look for comments where the agent debates interpretations and then commits to one without evidence.
  3. Check the library call's option flags against the documented meaning of the named variant (e.g., bias/ddof/adjusted parameters), not against the agent's inline comment; flag any case where the comment restates the task wording next to a flag that actually selects the other variant.
  4. Confirm the reported answer is the one produced under the task's slice+variant, and that no alternative interpretation was computed and then discarded arbitrarily.
Discriminator
A real violation is a subset or flag chosen for convenience/guesswork that changes the result and contradicts the literal task wording; it is fine if the restriction is forced by the data (e.g., the requested columns are the only ones present) or if the agent computed the interpretations and showed they agree, or documented the flag's meaning from the library's own definition.
Consequence
The reported entity/value is the argmax of a different statistic over a different sample, so the grader's single expected label does not match and the check fails outright.
id 7605f052ec79 · mined from infiagent-dabench dabench-252@s9
raw text (what the judge reads)
### Silently redefining the analysis scope and mismapping a stated statistic definition to library options
- **Applies when**: `task` -- the task names a specific statistic with an explicit definition/adjustment and a specific slice of the data, and the scripts compute it via a library call with option flags over a self-chosen set of columns/rows.
- **Pattern**: Faced with wording it finds ambiguous, the attempt invents a restricted window of the data (e.g., "all periods up to and including the named one") without justification, and passes a flag whose meaning it asserts in a comment rather than verifies against the library docs — so both the input set and the estimator variant differ from what was requested, while the code looks tidy and self-consistent.
- **Detection procedure**:
  1. Read the task and write down (a) the exact data slice requested and (b) the exact estimator variant/adjustment named.
  2. In the scripts, locate the array actually passed to the statistic and check whether its construction matches (a) exactly, or whether columns/rows were dropped/truncated by an interpretive choice; look for comments where the agent debates interpretations and then commits to one without evidence.
  3. Check the library call's option flags against the documented meaning of the named variant (e.g., bias/ddof/adjusted parameters), not against the agent's inline comment; flag any case where the comment restates the task wording next to a flag that actually selects the other variant.
  4. Confirm the reported answer is the one produced under the task's slice+variant, and that no alternative interpretation was computed and then discarded arbitrarily.
- **Discriminator**: A real violation is a subset or flag chosen for convenience/guesswork that changes the result and contradicts the literal task wording; it is *fine* if the restriction is forced by the data (e.g., the requested columns are the only ones present) or if the agent computed the interpretations and showed they agree, or documented the flag's meaning from the library's own definition.
- **Consequence**: The reported entity/value is the argmax of a different statistic over a different sample, so the grader's single expected label does not match and the check fails outright.
500Truncating/reformatting an identifier value so information required by the answer is losttaskinfiagent-dabench
Applies when
task -- The task asks you to identify a specific record (a date, ID, or key) and the answer template shows a format string that is coarser or ambiguous relative to the granularity at which the record actually exists in the data.
Pattern
The script correctly locates the record, then blindly applies the literal format token (e.g., strftime to a shorter pattern, string slicing, rounding a key, taking a group label) and reports a coarsened value that no longer uniquely identifies the record found — even though the downstream calculation used the full-precision record.
Detection procedure
  1. In the task, note the granularity of the entity being searched for (per-row/per-observation) and compare it to the granularity implied by the answer-format template.
  2. In the scripts, find the line that converts the located key into the reported string; check whether it discards components present in the source data.
  3. Check for internal inconsistency: does the reported key resolve to many rows while the accompanying computed metric was derived from exactly one row (or from that row's immediate neighbor)?
  4. If so, require the answer to carry the full identifier as stored in the data (a coarser template should be treated as an illustrative placeholder, not a mandate to delete precision).
Discriminator
A real violation is when the coarsened key is no longer a unique or verifiable pointer to the row used in the computation. It is fine if the task explicitly requires aggregation at the coarser level (e.g., "find the month with the highest average"), or if the source data itself only exists at that granularity.
Consequence
The computed numeric value matches, but the identifier check fails on exact-string comparison against the full-precision ground truth, so the submission is scored incorrect despite correct analysis.
id c8eb49335b36 · mined from infiagent-dabench dabench-572@s9
raw text (what the judge reads)
### Truncating/reformatting an identifier value so information required by the answer is lost
- **Applies when**: `task` -- The task asks you to identify a specific record (a date, ID, or key) and the answer template shows a format string that is coarser or ambiguous relative to the granularity at which the record actually exists in the data.
- **Pattern**: The script correctly locates the record, then blindly applies the literal format token (e.g., strftime to a shorter pattern, string slicing, rounding a key, taking a group label) and reports a coarsened value that no longer uniquely identifies the record found — even though the downstream calculation used the full-precision record.
- **Detection procedure**:
  1. In the task, note the granularity of the entity being searched for (per-row/per-observation) and compare it to the granularity implied by the answer-format template.
  2. In the scripts, find the line that converts the located key into the reported string; check whether it discards components present in the source data.
  3. Check for internal inconsistency: does the reported key resolve to many rows while the accompanying computed metric was derived from exactly one row (or from that row's immediate neighbor)?
  4. If so, require the answer to carry the full identifier as stored in the data (a coarser template should be treated as an illustrative placeholder, not a mandate to delete precision).
- **Discriminator**: A real violation is when the coarsened key is no longer a unique or verifiable pointer to the row used in the computation. It is fine if the task explicitly requires aggregation at the coarser level (e.g., "find the month with the highest average"), or if the source data itself only exists at that granularity.
- **Consequence**: The computed numeric value matches, but the identifier check fails on exact-string comparison against the full-precision ground truth, so the submission is scored incorrect despite correct analysis.
501Requested output artifact never written to the specified filetaskda-code
Applies when
task -- the task explicitly asks that results be saved to a named output file (e.g., result.csv) in addition to (or instead of) being reported in chat.
Pattern
The agent computes a number and reports it only in its text answer, with no script step that materializes the value into the required file (or writes it under a different name/path/format, or the writing code is never actually executed / left in an unsaved ad-hoc snippet).
Detection procedure
  1. Read the task statement and list every required deliverable: file name, location, and any implied structure (column header, row per result, rounding/units).
  2. Scan the scripts for an explicit write call (to_csv, open(...).write, etc.) whose target filename matches the requested name exactly, and confirm that call is on the executed path (not commented out, not in a function never called, not in a directory different from the working dir).
  3. Check the reported answer: if it is only prose/inline text with no reference to a produced artifact, and no directory listing or read-back confirms the file exists, treat the deliverable as missing.
  4. If a file is written, verify its contents match the requested quantity and format (correct value, sensible header/shape) by reading it back.
Discriminator
A real violation is the absence of an executed write to the exactly-named artifact, or a file whose contents don't hold the requested statistic; a look-alike that is fine is a script that writes the correct file and additionally prints the value to console, or that writes with a slightly different column name while the required filename and value are correct (unless the task fixed the schema).
Consequence
The grader looks for the named result file, finds it missing or containing the wrong content, and marks the check failed regardless of whether the reported number is numerically right.
id a0514f06bbed · mined from da-code dacode-data-sa-043@s9
raw text (what the judge reads)
### Requested output artifact never written to the specified file
- **Applies when**: `task` -- the task explicitly asks that results be saved to a named output file (e.g., `result.csv`) in addition to (or instead of) being reported in chat.
- **Pattern**: The agent computes a number and reports it only in its text answer, with no script step that materializes the value into the required file (or writes it under a different name/path/format, or the writing code is never actually executed / left in an unsaved ad-hoc snippet).
- **Detection procedure**:
  1. Read the task statement and list every required deliverable: file name, location, and any implied structure (column header, row per result, rounding/units).
  2. Scan the scripts for an explicit write call (`to_csv`, `open(...).write`, etc.) whose target filename matches the requested name exactly, and confirm that call is on the executed path (not commented out, not in a function never called, not in a directory different from the working dir).
  3. Check the reported answer: if it is only prose/inline text with no reference to a produced artifact, and no directory listing or read-back confirms the file exists, treat the deliverable as missing.
  4. If a file is written, verify its contents match the requested quantity and format (correct value, sensible header/shape) by reading it back.
- **Discriminator**: A real violation is the absence of an executed write to the exactly-named artifact, or a file whose contents don't hold the requested statistic; a look-alike that is fine is a script that writes the correct file and additionally prints the value to console, or that writes with a slightly different column name while the required filename and value are correct (unless the task fixed the schema).
- **Consequence**: The grader looks for the named result file, finds it missing or containing the wrong content, and marks the check failed regardless of whether the reported number is numerically right.
502Extraneous formatting characters injected into a delimited answer listtaskinfiagent-dabench
Applies when
task -- the task specifies a literal answer template with bracketed, comma-separated lists of identifiers or values, and the script string-builds that template.
Pattern
The script adds decoration not present in the template — quotes around string items, spaces after commas, units, or a trailing period — so the emitted tokens don't match the expected literal tokens even though the underlying computation is right.
Detection procedure
1. Copy the exact template from the task statement and note every literal character (brackets, commas, absence of quotes/spaces). 2. Read the formatting/output lines of the script and mentally render one item, including any f'"{x}"'-style wrapping, join separators, or str() of a numeric. 3. Compare the rendered string character-by-character with the template; also confirm numeric items honor stated rounding (e.g., two decimals preserved, not truncated to 9.0). 4. Check the final answer text for the same discrepancies.
Discriminator
A real violation is any added or missing literal character in the delimited list (quotes, stray spaces, prefixes); it is not a violation if the identifiers themselves legitimately contain punctuation that appears in the source data, nor if the grader clearly normalizes whitespace only.
Consequence
The value check may pass while the identifier/string check is scored WRONG/MISSING, since the graded token includes the extra characters and fails exact string comparison.
id 6343261e9fba · mined from infiagent-dabench dabench-219@s9
raw text (what the judge reads)
### Extraneous formatting characters injected into a delimited answer list
- **Applies when**: `task` -- the task specifies a literal answer template with bracketed, comma-separated lists of identifiers or values, and the script string-builds that template.
- **Pattern**: The script adds decoration not present in the template — quotes around string items, spaces after commas, units, or a trailing period — so the emitted tokens don't match the expected literal tokens even though the underlying computation is right.
- **Detection procedure**: 1. Copy the exact template from the task statement and note every literal character (brackets, commas, absence of quotes/spaces). 2. Read the formatting/output lines of the script and mentally render one item, including any `f'"{x}"'`-style wrapping, `join` separators, or `str()` of a numeric. 3. Compare the rendered string character-by-character with the template; also confirm numeric items honor stated rounding (e.g., two decimals preserved, not truncated to `9.0`). 4. Check the final answer text for the same discrepancies.
- **Discriminator**: A real violation is any added or missing literal character in the delimited list (quotes, stray spaces, prefixes); it is *not* a violation if the identifiers themselves legitimately contain punctuation that appears in the source data, nor if the grader clearly normalizes whitespace only.
- **Consequence**: The value check may pass while the identifier/string check is scored WRONG/MISSING, since the graded token includes the extra characters and fails exact string comparison.
503Analysis run on a subset of the available data instead of the full population the task specifiestaskinfiagent-dabench
Applies when
task -- the question asks about "all" units (countries, users, records, etc.) while the data directory holds several partitioned files (by region, split, time chunk) or the loaded table is only one slice of the whole.
Pattern
The script hard-codes a single partition file (or filters to one group) and computes the requested statistic on it, never checking whether other files/rows belonging to the same population exist and must be concatenated first; the reported answer therefore describes a sub-population, and any threshold-based statistic (quartiles, means, cutoffs) is derived from the wrong distribution.
Detection procedure
1. Read the task statement and note the stated scope ("all X", "the whole dataset", no filter mentioned). 2. In the scripts, list every data source read and every filter/subset applied; note whether the filename or filter implies a partition (region, category, split). 3. Check whether the script anywhere enumerates the data directory or verifies the row/entity count against the expected full population; absence of such a check with a partition-named source is the violation. 4. Confirm the reported statistic (e.g., quantiles/bounds) was computed after, not before, combining all partitions.
Discriminator
A real violation is when the full population is available elsewhere (other files/rows) and was silently excluded; it is fine if the task itself restricts scope to that partition, or if the script explicitly demonstrates (via directory listing or count check) that the single source already contains the entire population.
Consequence
Quartiles/thresholds and the resulting flagged set differ from those computed on the full data, so the reported entity list is incomplete or contains false positives and the grader marks the answer wrong even when the method is correctly implemented.
id c49bb9b0e57e · mined from infiagent-dabench dabench-254@s9
raw text (what the judge reads)
### Analysis run on a subset of the available data instead of the full population the task specifies
- **Applies when**: `task` -- the question asks about "all" units (countries, users, records, etc.) while the data directory holds several partitioned files (by region, split, time chunk) or the loaded table is only one slice of the whole.
- **Pattern**: The script hard-codes a single partition file (or filters to one group) and computes the requested statistic on it, never checking whether other files/rows belonging to the same population exist and must be concatenated first; the reported answer therefore describes a sub-population, and any threshold-based statistic (quartiles, means, cutoffs) is derived from the wrong distribution.
- **Detection procedure**: 1. Read the task statement and note the stated scope ("all X", "the whole dataset", no filter mentioned). 2. In the scripts, list every data source read and every filter/subset applied; note whether the filename or filter implies a partition (region, category, split). 3. Check whether the script anywhere enumerates the data directory or verifies the row/entity count against the expected full population; absence of such a check with a partition-named source is the violation. 4. Confirm the reported statistic (e.g., quantiles/bounds) was computed after, not before, combining all partitions.
- **Discriminator**: A real violation is when the full population is available elsewhere (other files/rows) and was silently excluded; it is fine if the task itself restricts scope to that partition, or if the script explicitly demonstrates (via directory listing or count check) that the single source already contains the entire population.
- **Consequence**: Quartiles/thresholds and the resulting flagged set differ from those computed on the full data, so the reported entity list is incomplete or contains false positives and the grader marks the answer wrong even when the method is correctly implemented.
504Output artifact not verified against the required row count and clean formattaskda-code
Applies when
task -- the task requires writing predictions/results to a named file with a specified column, and the agent's scripts produce that file and print a snippet as the answer.
Pattern
The attempt writes the file and declares success without any post-write verification: it never re-reads the saved file to confirm it has exactly one row per input record, the exact requested column name/header, valid label values, and no extra index column, log lines, or truncated content. The reported answer shows only a handful of rows (often mixed with status messages), which is consistent with a file that is truncated, mis-indexed, or contains the wrong number/format of predictions.
Detection procedure
  1. From the task, note the required output filename, column name, label vocabulary, and the number of rows in the evaluation input.
  2. Search the scripts for a verification step after saving: a re-read of the output file plus assertions/prints of shape, header names, value_counts() of the labels, and null checks — and confirm the writer suppresses the index and uses the exact requested header.
  3. Inspect the reported answer: check whether it presents evidence of full-file integrity (row count matching the input, label distribution) or merely a few lines and a "saved" message.
  4. Flag the attempt if no such read-back/shape/label check exists, or if the shown output contains non-data text or a row count that cannot match the evaluation input.
Discriminator
A real violation is the absence of any check tying the saved file's row count, header, and label values back to the evaluation input; it is not a violation if the script asserts/prints these (e.g., len(pred) == len(test), header equals the requested name, labels ⊂ allowed set) even if the answer text itself only displays a preview.
Consequence
The grader reads the output file and finds it wrong or unreadable — mismatched row count, wrong/extra columns, or label strings that don't align with expected values — so the submission scores 0 regardless of model quality.
id 05b746fa1f5e · mined from da-code dacode-ml-binary-009@s9
raw text (what the judge reads)
### Output artifact not verified against the required row count and clean format
- **Applies when**: `task` -- the task requires writing predictions/results to a named file with a specified column, and the agent's scripts produce that file and print a snippet as the answer.
- **Pattern**: The attempt writes the file and declares success without any post-write verification: it never re-reads the saved file to confirm it has exactly one row per input record, the exact requested column name/header, valid label values, and no extra index column, log lines, or truncated content. The reported answer shows only a handful of rows (often mixed with status messages), which is consistent with a file that is truncated, mis-indexed, or contains the wrong number/format of predictions.
- **Detection procedure**:
  1. From the task, note the required output filename, column name, label vocabulary, and the number of rows in the evaluation input.
  2. Search the scripts for a verification step after saving: a re-read of the output file plus assertions/prints of `shape`, header names, `value_counts()` of the labels, and null checks — and confirm the writer suppresses the index and uses the exact requested header.
  3. Inspect the reported answer: check whether it presents evidence of full-file integrity (row count matching the input, label distribution) or merely a few lines and a "saved" message.
  4. Flag the attempt if no such read-back/shape/label check exists, or if the shown output contains non-data text or a row count that cannot match the evaluation input.
- **Discriminator**: A real violation is the absence of any check tying the saved file's row count, header, and label values back to the evaluation input; it is *not* a violation if the script asserts/prints these (e.g., `len(pred) == len(test)`, header equals the requested name, labels ⊂ allowed set) even if the answer text itself only displays a preview.
- **Consequence**: The grader reads the output file and finds it wrong or unreadable — mismatched row count, wrong/extra columns, or label strings that don't align with expected values — so the submission scores 0 regardless of model quality.
505Invented metric definition instead of the domain-standard / spec-implied quantitytaskda-code
Applies when
task -- the task asks to summarize or rank entities by a vague quality (e.g., "performance", "score", "activity") and a config/spec file plus expected output artifacts define the axes, labels, and required deliverables.
Pattern
The script fabricates an arbitrary formula (an ad-hoc weighted sum of counts or of fields unrelated to the quality being measured) rather than deriving the quantity implied by the axis label, the domain convention, or the spec; it also skips the other required artifacts and never sanity-checks the components of the formula (e.g., a boolean column compared against a string, silently yielding zeros).
Detection procedure
  1. Read the task and any config/spec file: list every required output artifact and the exact semantic of the y-axis/metric label.
  2. In the script, locate the line where the plotted/reported quantity is computed and ask whether the formula is stated anywhere in the task/spec or is a standard definition of that label; check that each input field's dtype/values match how it is filtered or compared.
  3. Compare the artifacts actually written by the script with the required list; flag any missing file or any quantity not reconstructible from the spec.
  4. Check the answer for any validation of the numbers (ranges, counts, cross-check against a known ranking); absence of such a check with an invented formula is a violation.
Discriminator
A fine attempt either uses the definition given/implied by the spec (or a standard one) and states the mapping explicitly, or, if genuinely ambiguous, tests candidate definitions against a checkable signal (label semantics, known ordering, provided reference values) — versus this failure, where a formula with unjustified weights is asserted with no derivation and no verification, and part of the deliverables is never produced.
Consequence
The plotted bars and saved numeric arrays differ from the reference values, and missing artifacts fail outright, so every equality/closeness check on the outputs fails.
id e647bd36428b · mined from da-code dacode-plot-bar-006@s9
raw text (what the judge reads)
### Invented metric definition instead of the domain-standard / spec-implied quantity
- **Applies when**: `task` -- the task asks to summarize or rank entities by a vague quality (e.g., "performance", "score", "activity") and a config/spec file plus expected output artifacts define the axes, labels, and required deliverables.
- **Pattern**: The script fabricates an arbitrary formula (an ad-hoc weighted sum of counts or of fields unrelated to the quality being measured) rather than deriving the quantity implied by the axis label, the domain convention, or the spec; it also skips the other required artifacts and never sanity-checks the components of the formula (e.g., a boolean column compared against a string, silently yielding zeros).
- **Detection procedure**:
  1. Read the task and any config/spec file: list every required output artifact and the exact semantic of the y-axis/metric label.
  2. In the script, locate the line where the plotted/reported quantity is computed and ask whether the formula is stated anywhere in the task/spec or is a standard definition of that label; check that each input field's dtype/values match how it is filtered or compared.
  3. Compare the artifacts actually written by the script with the required list; flag any missing file or any quantity not reconstructible from the spec.
  4. Check the answer for any validation of the numbers (ranges, counts, cross-check against a known ranking); absence of such a check with an invented formula is a violation.
- **Discriminator**: A fine attempt either uses the definition given/implied by the spec (or a standard one) and states the mapping explicitly, or, if genuinely ambiguous, tests candidate definitions against a checkable signal (label semantics, known ordering, provided reference values) — versus this failure, where a formula with unjustified weights is asserted with no derivation and no verification, and part of the deliverables is never produced.
- **Consequence**: The plotted bars and saved numeric arrays differ from the reference values, and missing artifacts fail outright, so every equality/closeness check on the outputs fails.
506Reported outlier/filter count not reconciled with the stated rule's extreme statistictaskinfiagent-dabench
Applies when
task -- The task defines an explicit numeric rule (e.g., a standardized-score or fixed-cutoff criterion) for flagging or removing rows, and the answer is the count of flagged rows.
Pattern
The agent applies some flagging routine (a library helper, a different robust/quantile variant, a cutoff applied to raw values, or statistics computed on a subset/wrong column/with NaNs or non-numeric strings coerced) and reports whatever count it produces, without ever printing the extreme value of the rule's own statistic (e.g., max |standardized score|) to confirm that any row can actually exceed the stated threshold.
Detection procedure
  1. Read the task and write down the exact rule: which column, which statistic (mean/SD over the full cleaned column), and the exact cutoff and comparison (strictly greater in absolute value).
  2. Read the script: check the statistic is computed with the stated definition on the full, numeric-cleaned target column (not a robust/median variant, not per-group, not on a filtered or transformed subset, not on the raw values), and that the comparison uses the stated cutoff and sign convention.
  3. Check the script prints diagnostics that reconcile the count: column min/max, mean, SD, max |statistic|, and the number of rows satisfying the condition; verify the reported count equals that number and that max |statistic| > cutoff is consistent with a nonzero count.
  4. Confirm the answer reports the requested count in the requested format (and not rows dropped for other reasons such as missing values, or the size of the resulting dataframe).
Discriminator
A real violation is a count produced by a rule that differs from the specified one, or that is never cross-checked against the extreme value of the specified statistic; it is fine if the script uses a library shortcut but explicitly verifies the same count from the stated formula and prints max |statistic| relative to the cutoff — including the legitimate case where that yields zero flagged rows.
Consequence
The graded count mismatches ground truth (e.g., a nonzero count when the correct rule flags none), failing the exact-match check.
id 5ae791d5ae07 · mined from infiagent-dabench dabench-361@s9
raw text (what the judge reads)
### Reported outlier/filter count not reconciled with the stated rule's extreme statistic
- **Applies when**: `task` -- The task defines an explicit numeric rule (e.g., a standardized-score or fixed-cutoff criterion) for flagging or removing rows, and the answer is the count of flagged rows.
- **Pattern**: The agent applies some flagging routine (a library helper, a different robust/quantile variant, a cutoff applied to raw values, or statistics computed on a subset/wrong column/with NaNs or non-numeric strings coerced) and reports whatever count it produces, without ever printing the extreme value of the rule's own statistic (e.g., max |standardized score|) to confirm that any row can actually exceed the stated threshold.
- **Detection procedure**:
  1. Read the task and write down the exact rule: which column, which statistic (mean/SD over the full cleaned column), and the exact cutoff and comparison (strictly greater in absolute value).
  2. Read the script: check the statistic is computed with the stated definition on the full, numeric-cleaned target column (not a robust/median variant, not per-group, not on a filtered or transformed subset, not on the raw values), and that the comparison uses the stated cutoff and sign convention.
  3. Check the script prints diagnostics that reconcile the count: column min/max, mean, SD, max |statistic|, and the number of rows satisfying the condition; verify the reported count equals that number and that max |statistic| > cutoff is consistent with a nonzero count.
  4. Confirm the answer reports the requested count in the requested format (and not rows dropped for other reasons such as missing values, or the size of the resulting dataframe).
- **Discriminator**: A real violation is a count produced by a rule that differs from the specified one, or that is never cross-checked against the extreme value of the specified statistic; it is fine if the script uses a library shortcut but explicitly verifies the same count from the stated formula and prints max |statistic| relative to the cutoff — including the legitimate case where that yields zero flagged rows.
- **Consequence**: The graded count mismatches ground truth (e.g., a nonzero count when the correct rule flags none), failing the exact-match check.
507Ignoring a task-referenced specification file and substituting assumed conventionstaskda-code
Applies when
task -- the prompt tells the agent to use definitions, mappings, filters, or rules that live in an auxiliary file (README, tips/notes, config, data dictionary) rather than in the prompt itself.
Pattern
The scripts never open or print the referenced file; instead the agent hardcodes a "standard"/"common" version of the mapping or rule from prior knowledge (often flagged by a comment like "typical mapping for this kind of dataset"), so category names, groupings, or derived values can silently differ from the required ones — and any required output artifact is likewise produced from guessed conventions or not written at all.
Detection procedure
  1. List every external artifact the task instructs the agent to consult (spec file) or produce (result file, specific path/format).
  2. Grep the scripts for reads of that spec file (open, read_csv, read_text, etc.) and for writes of the required output file.
  3. If the mapping/rule is inlined as a literal in the script, check whether any script output shows it was cross-checked against the file's actual contents; also check whether categories could be merged/renamed differently than assumed (e.g., a mapping that collapses two raw codes into one label changes both the argmax and the ratio).
  4. Confirm the final answer's label strings and numeric formatting come from the file's vocabulary and the task's stated rounding/format, not from the agent's invention.
Discriminator
A real violation is when the authoritative file is never read and its content is guessed (or the required output file is never persisted). It is fine if the script reads the file (or prints its contents) and then encodes the same mapping as a literal for clarity, or if the task itself states the mapping inline.
Consequence
The graded artifact is missing or contains label strings/derived statistics that do not match the reference mapping, so the answer is scored wrong even when the counting logic is arithmetically correct.
id ef4fdf10a64e · mined from da-code dacode-di-text-004@s9
raw text (what the judge reads)
### Ignoring a task-referenced specification file and substituting assumed conventions
- **Applies when**: `task` -- the prompt tells the agent to use definitions, mappings, filters, or rules that live in an auxiliary file (README, tips/notes, config, data dictionary) rather than in the prompt itself.
- **Pattern**: The scripts never open or print the referenced file; instead the agent hardcodes a "standard"/"common" version of the mapping or rule from prior knowledge (often flagged by a comment like "typical mapping for this kind of dataset"), so category names, groupings, or derived values can silently differ from the required ones — and any required output artifact is likewise produced from guessed conventions or not written at all.
- **Detection procedure**:
  1. List every external artifact the task instructs the agent to consult (spec file) or produce (result file, specific path/format).
  2. Grep the scripts for reads of that spec file (`open`, `read_csv`, `read_text`, etc.) and for writes of the required output file.
  3. If the mapping/rule is inlined as a literal in the script, check whether any script output shows it was cross-checked against the file's actual contents; also check whether categories could be merged/renamed differently than assumed (e.g., a mapping that collapses two raw codes into one label changes both the argmax and the ratio).
  4. Confirm the final answer's label strings and numeric formatting come from the file's vocabulary and the task's stated rounding/format, not from the agent's invention.
- **Discriminator**: A real violation is when the authoritative file is never read and its content is guessed (or the required output file is never persisted). It is fine if the script reads the file (or prints its contents) and then encodes the same mapping as a literal for clarity, or if the task itself states the mapping inline.
- **Consequence**: The graded artifact is missing or contains label strings/derived statistics that do not match the reference mapping, so the answer is scored wrong even when the counting logic is arithmetically correct.
508Prediction file not validated against the test set's contract (row count, order, column name, label vocabulary)taskda-code
Applies when
task -- The task asks for predictions on a supplied evaluation split written to a named output file with a named column, and the agent produces that file (possibly with no saved, re-runnable script).
Pattern
The attempt writes an output file without verifying it against the evaluation input: row count differs from the number of evaluation rows (e.g. rows dropped by dropna/filtering during preprocessing, or an index reset that changes order), the column name/header differs from the exact requested string, extra index/ID columns are emitted, or predicted values are not in the label vocabulary of the target (e.g. encoded integers, probabilities, or NaN for rows the model could not score) — and no code is retained that a reviewer could re-run to check any of this.
Detection procedure
  1. From the task, note the exact required output filename, required column name(s), the expected number of prediction rows (= rows in the evaluation file), and the allowed set of target values as seen in the training labels.
  2. In the scripts, trace the evaluation data from load to write: look for any row-dropping (dropna, boolean filtering, deduplication, merges that lose or duplicate rows), any reordering/sorting, and how the target is decoded back to its original labels before writing; confirm a script exists at all and that the write step names the column exactly as requested.
  3. Inspect the produced file's header, shape, and unique values; confirm rows align 1:1 and in the same order as the evaluation input, no nulls, and every value is a legal label string.
  4. Flag the attempt if any of these cannot be confirmed from the scripts/output — in particular if no script is saved so the mapping from evaluation rows to output rows is unverifiable.
Discriminator
A real violation is a mismatch in the file's contract (count, order, header, value domain, nulls) or the absence of any reproducible code establishing that contract. It is not a violation if the file has exactly one row per evaluation record in input order, the exact requested column name, and only valid label values — even if the model is simple, the accuracy is low, or extra intermediate files were also written.
Consequence
The grader reads the expected file and finds it WRONG/MISSING — it cannot align predictions to ground truth (shape/column-name mismatch) or scores unparseable values, yielding 0 checks passed regardless of model quality.
id 6464cd337fbf · mined from da-code dacode-ml-multi-003@s9
raw text (what the judge reads)
### Prediction file not validated against the test set's contract (row count, order, column name, label vocabulary)
- **Applies when**: `task` -- The task asks for predictions on a supplied evaluation split written to a named output file with a named column, and the agent produces that file (possibly with no saved, re-runnable script).
- **Pattern**: The attempt writes an output file without verifying it against the evaluation input: row count differs from the number of evaluation rows (e.g. rows dropped by `dropna`/filtering during preprocessing, or an index reset that changes order), the column name/header differs from the exact requested string, extra index/ID columns are emitted, or predicted values are not in the label vocabulary of the target (e.g. encoded integers, probabilities, or `NaN` for rows the model could not score) — and no code is retained that a reviewer could re-run to check any of this.
- **Detection procedure**:
  1. From the task, note the exact required output filename, required column name(s), the expected number of prediction rows (= rows in the evaluation file), and the allowed set of target values as seen in the training labels.
  2. In the scripts, trace the evaluation data from load to write: look for any row-dropping (`dropna`, boolean filtering, deduplication, merges that lose or duplicate rows), any reordering/sorting, and how the target is decoded back to its original labels before writing; confirm a script exists at all and that the write step names the column exactly as requested.
  3. Inspect the produced file's header, shape, and unique values; confirm rows align 1:1 and in the same order as the evaluation input, no nulls, and every value is a legal label string.
  4. Flag the attempt if any of these cannot be confirmed from the scripts/output — in particular if no script is saved so the mapping from evaluation rows to output rows is unverifiable.
- **Discriminator**: A real violation is a mismatch in the file's contract (count, order, header, value domain, nulls) or the absence of any reproducible code establishing that contract. It is *not* a violation if the file has exactly one row per evaluation record in input order, the exact requested column name, and only valid label values — even if the model is simple, the accuracy is low, or extra intermediate files were also written.
- **Consequence**: The grader reads the expected file and finds it WRONG/MISSING — it cannot align predictions to ground truth (shape/column-name mismatch) or scores unparseable values, yielding 0 checks passed regardless of model quality.
509Template/format conformance never verified against the provided reference filetaskda-code
Applies when
task -- The task says the output must be saved to a named file "matching the provided template/example format", and the scripts write a derived table (pivot/aggregation) to that file.
Pattern
The agent builds the result with its own assumed layout (its own index/column names, row & column ordering, rounding/precision, NaN vs blank vs 0 handling, index-written-or-not, dtype of the key column) and never programmatically reads the template to compare headers, shape, and cell formatting; the write step is treated as done once the file exists.
Detection procedure
  1. Read the task for the required artifact name and the phrase referencing a template/example; note every implied formatting constraint (column labels, ordering, rounding, empty-cell convention).
  2. Search the scripts for any read of the template file and an explicit comparison (e.g., asserting equal column lists, row count/labels, dtypes, decimal places). If absent, the check fails.
  3. Inspect the write call and any transformation before it (round(...), fillna, reset_index, to_csv(index=...), date formatting) and ask whether each choice was justified by the template or invented by the agent.
  4. Read the answer: if it describes the output structure in its own words ("rounded to 1 decimal", "empty cells", "CohortIndex 1..N") rather than reporting a passed comparison against the template, treat it as unverified.
Discriminator
A real violation is inventing formatting choices with no template read/assert anywhere; it is fine if the script loads the template (or its header) and asserts/aligns columns, ordering, and precision — even if it also rounds or fills values, since those were derived from the reference.
Consequence
The grader compares the submitted file cell-by-cell (or header-by-header) against the expected file and marks it WRONG/MISSING despite the underlying aggregation possibly being close to right.
id 0c9f74e1b546 · mined from da-code dacode-dm-csv-044@s9
raw text (what the judge reads)
### Template/format conformance never verified against the provided reference file
- **Applies when**: `task` -- The task says the output must be saved to a named file "matching the provided template/example format", and the scripts write a derived table (pivot/aggregation) to that file.
- **Pattern**: The agent builds the result with its own assumed layout (its own index/column names, row & column ordering, rounding/precision, NaN vs blank vs 0 handling, index-written-or-not, dtype of the key column) and never programmatically reads the template to compare headers, shape, and cell formatting; the write step is treated as done once the file exists.
- **Detection procedure**:
  1. Read the task for the required artifact name and the phrase referencing a template/example; note every implied formatting constraint (column labels, ordering, rounding, empty-cell convention).
  2. Search the scripts for any read of the template file and an explicit comparison (e.g., asserting equal column lists, row count/labels, dtypes, decimal places). If absent, the check fails.
  3. Inspect the write call and any transformation before it (`round(...)`, `fillna`, `reset_index`, `to_csv(index=...)`, date formatting) and ask whether each choice was justified by the template or invented by the agent.
  4. Read the answer: if it *describes* the output structure in its own words ("rounded to 1 decimal", "empty cells", "CohortIndex 1..N") rather than reporting a passed comparison against the template, treat it as unverified.
- **Discriminator**: A real violation is inventing formatting choices with no template read/assert anywhere; it is fine if the script loads the template (or its header) and asserts/aligns columns, ordering, and precision — even if it also rounds or fills values, since those were derived from the reference.
- **Consequence**: The grader compares the submitted file cell-by-cell (or header-by-header) against the expected file and marks it WRONG/MISSING despite the underlying aggregation possibly being close to right.
510Model selection and metric reporting done on the training data itself (no held-out validation)taskda-code
Applies when
task -- scripts fit one or more models on all labeled rows and then compute the competition/evaluation metric with those same rows to compare candidates or justify the final submission.
Pattern
The attempt scores model.predict(X_train) against y_train (or scores a stacked/ensemble model whose members were fit on the full data), reports a near-perfect in-sample metric, and picks the "optimized" configuration and any decision thresholds/rounding on that basis — so nothing in the pipeline estimates out-of-sample agreement, and the imbalanced/ordinal structure the metric punishes is never checked.
Detection procedure
  1. Read the task to identify the stated evaluation metric and whether the target is ordinal/imbalanced.
  2. Scan the scripts for any split, K-fold, or out-of-fold prediction used to compute that metric; check that the data used for scoring was excluded from fit.
  3. Check whether the chosen model/hyperparameters/thresholds were selected using that out-of-sample score, and whether the reported score is plausible (in-sample scores near 1.0 on noisy tabular data are a red flag).
  4. Inspect the predicted label distribution versus the training label distribution: a model tuned in-sample typically collapses to the majority classes and never predicts rare extreme levels, which the metric penalizes.
Discriminator
A real violation is when no honest estimate exists anywhere (all metrics come from rows used in fitting) or when selection is driven by that in-sample number; it is fine if in-sample numbers are printed only as diagnostics alongside cross-validated or hold-out scores that actually drive the final choice.
Consequence
The reported metric (e.g., ~0.99 agreement) is meaningless; the submitted predictions score far lower on the hidden test set — typically well below a simple validated baseline — and the answer fails the accuracy threshold.
id cc8f78a59b6f · mined from da-code dacode-ml-competition-006@s9
raw text (what the judge reads)
### Model selection and metric reporting done on the training data itself (no held-out validation)
- **Applies when**: `task` -- scripts fit one or more models on all labeled rows and then compute the competition/evaluation metric with those same rows to compare candidates or justify the final submission.
- **Pattern**: The attempt scores `model.predict(X_train)` against `y_train` (or scores a stacked/ensemble model whose members were fit on the full data), reports a near-perfect in-sample metric, and picks the "optimized" configuration and any decision thresholds/rounding on that basis — so nothing in the pipeline estimates out-of-sample agreement, and the imbalanced/ordinal structure the metric punishes is never checked.
- **Detection procedure**:
  1. Read the task to identify the stated evaluation metric and whether the target is ordinal/imbalanced.
  2. Scan the scripts for any split, K-fold, or out-of-fold prediction used to compute that metric; check that the data used for scoring was excluded from `fit`.
  3. Check whether the chosen model/hyperparameters/thresholds were selected using that out-of-sample score, and whether the reported score is plausible (in-sample scores near 1.0 on noisy tabular data are a red flag).
  4. Inspect the predicted label distribution versus the training label distribution: a model tuned in-sample typically collapses to the majority classes and never predicts rare extreme levels, which the metric penalizes.
- **Discriminator**: A real violation is when *no* honest estimate exists anywhere (all metrics come from rows used in fitting) or when selection is driven by that in-sample number; it is fine if in-sample numbers are printed only as diagnostics alongside cross-validated or hold-out scores that actually drive the final choice.
- **Consequence**: The reported metric (e.g., ~0.99 agreement) is meaningless; the submitted predictions score far lower on the hidden test set — typically well below a simple validated baseline — and the answer fails the accuracy threshold.
511Group partition by missingness not validated as exhaustive and correctly definedtaskinfiagent-dabench
Applies when
task -- a task asks to split rows into "null" vs "non-null" (or otherwise filtered) groups on some column and compare aggregate statistics between them.
Pattern
The attempt defines the missingness mask loosely — e.g. relying on default parsers that read empty strings, "NA", "None", "-", or whitespace as real values (or vice versa), dropping rows via dropna()/read options before splitting, or filtering on the measured column too — so the two groups don't partition the full table and each group's mean is computed on a shifted subset.
Detection procedure
  1. From the task, note the exact partition rule and confirm every row of the source table must land in exactly one of the two groups.
  2. In the scripts, check how the data is loaded (parser/na_values/dtype settings) and whether any row filtering, deduplication, or dropna happens before the split; check the mask uses a null test consistent with how the file encodes emptiness.
  3. Verify the script prints group sizes and that n_group1 + n_group2 == total rows, plus the raw value counts of the split column's distinct sentinel-like values.
  4. Compare the reported means against the printed group counts/ranges; if no counts or sanity checks are printed (or no script was retained), treat the numbers as unverified.
Discriminator
A real violation is when the split is derived from an unchecked default parse or after an extra filter, so group sizes are never shown to sum to the table length; it is fine if the script explicitly inspects the column's sentinel encodings, states the null rule, and prints counts that sum to the total — even if only a documented subset is legitimately excluded per the task.
Consequence
Both group means (and the test statistic) are computed on the wrong row sets, so the reported means differ from ground truth by a few percent and the numeric checks fail even though the p-value direction looks plausible.
id eed34927928f · mined from infiagent-dabench dabench-297@s9
raw text (what the judge reads)
### Group partition by missingness not validated as exhaustive and correctly defined
- **Applies when**: `task` -- a task asks to split rows into "null" vs "non-null" (or otherwise filtered) groups on some column and compare aggregate statistics between them.
- **Pattern**: The attempt defines the missingness mask loosely — e.g. relying on default parsers that read empty strings, `"NA"`, `"None"`, `"-"`, or whitespace as real values (or vice versa), dropping rows via `dropna()`/`read` options before splitting, or filtering on the measured column too — so the two groups don't partition the full table and each group's mean is computed on a shifted subset.
- **Detection procedure**:
  1. From the task, note the exact partition rule and confirm every row of the source table must land in exactly one of the two groups.
  2. In the scripts, check how the data is loaded (parser/`na_values`/dtype settings) and whether any row filtering, deduplication, or `dropna` happens before the split; check the mask uses a null test consistent with how the file encodes emptiness.
  3. Verify the script prints group sizes and that `n_group1 + n_group2 == total rows`, plus the raw value counts of the split column's distinct sentinel-like values.
  4. Compare the reported means against the printed group counts/ranges; if no counts or sanity checks are printed (or no script was retained), treat the numbers as unverified.
- **Discriminator**: A real violation is when the split is derived from an unchecked default parse or after an extra filter, so group sizes are never shown to sum to the table length; it is fine if the script explicitly inspects the column's sentinel encodings, states the null rule, and prints counts that sum to the total — even if only a documented subset is legitimately excluded per the task.
- **Consequence**: Both group means (and the test statistic) are computed on the wrong row sets, so the reported means differ from ground truth by a few percent and the numeric checks fail even though the p-value direction looks plausible.
512Invented preprocessing steps and missing required output artifacts (spec file never consulted)taskda-code
Applies when
task -- the task points to an external specification/guidance document (or otherwise enumerates required outputs and processing rules) and the script is expected to reproduce that pipeline and save its results to files.
Pattern
The script never reads/echoes the referenced spec, substitutes self-invented filters, category-mapping rules and thresholds ("keep rows above N", "last 6 months", ad-hoc keyword buckets) for the prescribed ones, and writes only the one artifact mentioned in the prompt's prose while silently omitting the other required saved outputs (e.g., serialized plot data, numeric arrays/tables).
Detection procedure
  1. From the task text, list (a) every referenced spec/guidance file and (b) every artifact the deliverable is supposed to consist of (image, JSON/plot spec, array/CSV of the tallied values).
  2. Scan the script for a read/parse of the spec file and for a save call per artifact in the list; note any filtering, grouping, or category-definition logic that has no textual basis in the task or spec.
  3. Compare the answer's reported numbers/steps against the spec's stated rules and against the artifact list; flag if the answer describes only self-derived rules or only some artifacts.
  4. Sanity-check magnitudes: if a large fraction of rows was dropped or left uncategorized by the agent's own rules, treat that as evidence the rules were invented rather than prescribed.
Discriminator
A real violation is when the spec was never opened and its rules/outputs cannot be traced in the code, or a required output file is simply never written; it is not a violation when the script demonstrably follows the spec (quoted/implemented rules, all artifacts saved) and merely makes a defensible judgment call on a genuinely ambiguous detail.
Consequence
Every expected artifact fails its check — missing files score zero outright, and the produced figure/counts reflect an arbitrary subset and arbitrary category definitions, so the grader reports 0/N checks passed.
id b25cde2d297a · mined from da-code dacode-plot-pie-005@s9
raw text (what the judge reads)
### Invented preprocessing steps and missing required output artifacts (spec file never consulted)
- **Applies when**: `task` -- the task points to an external specification/guidance document (or otherwise enumerates required outputs and processing rules) and the script is expected to reproduce that pipeline and save its results to files.
- **Pattern**: The script never reads/echoes the referenced spec, substitutes self-invented filters, category-mapping rules and thresholds ("keep rows above N", "last 6 months", ad-hoc keyword buckets) for the prescribed ones, and writes only the one artifact mentioned in the prompt's prose while silently omitting the other required saved outputs (e.g., serialized plot data, numeric arrays/tables).
- **Detection procedure**:
  1. From the task text, list (a) every referenced spec/guidance file and (b) every artifact the deliverable is supposed to consist of (image, JSON/plot spec, array/CSV of the tallied values).
  2. Scan the script for a read/parse of the spec file and for a save call per artifact in the list; note any filtering, grouping, or category-definition logic that has no textual basis in the task or spec.
  3. Compare the answer's reported numbers/steps against the spec's stated rules and against the artifact list; flag if the answer describes only self-derived rules or only some artifacts.
  4. Sanity-check magnitudes: if a large fraction of rows was dropped or left uncategorized by the agent's own rules, treat that as evidence the rules were invented rather than prescribed.
- **Discriminator**: A real violation is when the spec was never opened and its rules/outputs cannot be traced in the code, or a required output file is simply never written; it is *not* a violation when the script demonstrably follows the spec (quoted/implemented rules, all artifacts saved) and merely makes a defensible judgment call on a genuinely ambiguous detail.
- **Consequence**: Every expected artifact fails its check — missing files score zero outright, and the produced figure/counts reflect an arbitrary subset and arbitrary category definitions, so the grader reports 0/N checks passed.
513Row loss / silent data alteration instead of the prescribed missing-value handlingtaskinfiagent-dabench
Applies when
task -- the task specifies a preprocessing recipe (e.g., impute named columns with a stated statistic) before a train/test split and a single reported metric, and the raw columns may contain non-numeric tokens, blanks, or sentinel values.
Pattern
The script coerces or cleans the relevant columns in a way that quietly changes the modeled population — dropna(), boolean filtering, pd.to_numeric(..., errors='coerce') followed by row removal, or imputing only some of the named columns / imputing after rows were already discarded — so the model is fit and scored on a different (smaller or differently-scaled) sample than the specification implies. No row-count or distribution check is printed, and no script is retained to verify what actually ran.
Detection procedure
  1. From the task, list exactly which columns must be imputed, with what statistic, and at what stage relative to the split; note that the row count should be unchanged by preprocessing.
  2. In the script, trace each named column from load to model input: look for dropna, masks, astype, to_numeric(errors=...), string stripping/parsing, unit rescaling, or outlier removal, and check that every named column actually receives the stated imputation.
  3. Require the script to print df.shape and per-column NaN counts before and after preprocessing, plus train/test sizes; confirm rows before == rows after and test size ≈ the stated fraction.
  4. Sanity-check the reported metric against target scale: compare the metric to the variance of the target (an MSE far above the target's variance, or far from a baseline mean-predictor MSE, signals a corrupted/misaligned sample or units).
Discriminator
A real violation is any transformation that removes or rescales rows/values of the specified columns beyond the stated imputation, or imputation applied to only part of the named set; harmless look-alikes are pure parsing steps (stripping currency symbols/commas) that leave the row count and value scale identical and are followed by the required mean imputation, or metric differences caused only by an unstated split seed.
Consequence
The model is trained/evaluated on a non-conforming sample, so the reported error differs from the reference by an order of magnitude and the grader marks the single numeric answer wrong, with no saved script to diagnose the discrepancy.
id f6b66e1f40ce · mined from infiagent-dabench dabench-432@s9
raw text (what the judge reads)
### Row loss / silent data alteration instead of the prescribed missing-value handling
- **Applies when**: `task` -- the task specifies a preprocessing recipe (e.g., impute named columns with a stated statistic) before a train/test split and a single reported metric, and the raw columns may contain non-numeric tokens, blanks, or sentinel values.
- **Pattern**: The script coerces or cleans the relevant columns in a way that quietly changes the modeled population — `dropna()`, boolean filtering, `pd.to_numeric(..., errors='coerce')` followed by row removal, or imputing only some of the named columns / imputing after rows were already discarded — so the model is fit and scored on a different (smaller or differently-scaled) sample than the specification implies. No row-count or distribution check is printed, and no script is retained to verify what actually ran.
- **Detection procedure**:
  1. From the task, list exactly which columns must be imputed, with what statistic, and at what stage relative to the split; note that the row count should be unchanged by preprocessing.
  2. In the script, trace each named column from load to model input: look for `dropna`, masks, `astype`, `to_numeric(errors=...)`, string stripping/parsing, unit rescaling, or outlier removal, and check that every named column actually receives the stated imputation.
  3. Require the script to print `df.shape` and per-column NaN counts before and after preprocessing, plus train/test sizes; confirm rows before == rows after and test size ≈ the stated fraction.
  4. Sanity-check the reported metric against target scale: compare the metric to the variance of the target (an MSE far above the target's variance, or far from a baseline mean-predictor MSE, signals a corrupted/misaligned sample or units).
- **Discriminator**: A real violation is any transformation that removes or rescales rows/values of the specified columns beyond the stated imputation, or imputation applied to only part of the named set; harmless look-alikes are pure parsing steps (stripping currency symbols/commas) that leave the row count and value scale identical and are followed by the required mean imputation, or metric differences caused only by an unstated split seed.
- **Consequence**: The model is trained/evaluated on a non-conforming sample, so the reported error differs from the reference by an order of magnitude and the grader marks the single numeric answer wrong, with no saved script to diagnose the discrepancy.
514Row ordering assumed rather than verified before a sequential/lagged computationtaskinfiagent-dabench
Applies when
task -- the task requires a period-over-period change, lag, difference, rolling window, or any other order-dependent transform on tabular data.
Pattern
The script hard-codes an assumption about the row order (e.g., reverses the frame or leaves it as-is) based on a guess or a glance, instead of explicitly sorting by the time/sequence key; if the assumption is wrong the sign of every difference flips and the statistics are computed on a reversed series.
Detection procedure
1. Read the task to confirm the computation depends on "previous"/"next" rows. 2. In the script, find where order is established — check whether there is an explicit sort_values on a parsed datetime/sequence column, or merely a comment claiming the order and a blind reversal/no-op. 3. Check whether the ordering key is parsed to a proper dtype (not a string) and whether any printed diagnostic actually confirms ascending order before the shift/diff. 4. Compare the reported statistic's sign/magnitude with a plausibility check (e.g., a mean whose sign would invert under reversal is a red flag when order was never validated).
Discriminator
A real violation is an unverified/asserted ordering (comment-only justification, index-based reversal, or no sort) before an order-sensitive operation; it is fine if the script explicitly sorts by a correctly typed sequence key, or prints/asserts the first and last key values to confirm direction.
Consequence
Differences are computed backwards, so order-sensitive statistics (e.g., the mean) come out with the wrong sign and slightly wrong magnitude, failing the expected-value checks.
id fd558f08c1d4 · mined from infiagent-dabench dabench-75@s9
raw text (what the judge reads)
### Row ordering assumed rather than verified before a sequential/lagged computation
- **Applies when**: `task` -- the task requires a period-over-period change, lag, difference, rolling window, or any other order-dependent transform on tabular data.
- **Pattern**: The script hard-codes an assumption about the row order (e.g., reverses the frame or leaves it as-is) based on a guess or a glance, instead of explicitly sorting by the time/sequence key; if the assumption is wrong the sign of every difference flips and the statistics are computed on a reversed series.
- **Detection procedure**: 1. Read the task to confirm the computation depends on "previous"/"next" rows. 2. In the script, find where order is established — check whether there is an explicit `sort_values` on a parsed datetime/sequence column, or merely a comment claiming the order and a blind reversal/no-op. 3. Check whether the ordering key is parsed to a proper dtype (not a string) and whether any printed diagnostic actually confirms ascending order before the shift/diff. 4. Compare the reported statistic's sign/magnitude with a plausibility check (e.g., a mean whose sign would invert under reversal is a red flag when order was never validated).
- **Discriminator**: A real violation is an unverified/asserted ordering (comment-only justification, index-based reversal, or no sort) before an order-sensitive operation; it is fine if the script explicitly sorts by a correctly typed sequence key, or prints/asserts the first and last key values to confirm direction.
- **Consequence**: Differences are computed backwards, so order-sensitive statistics (e.g., the mean) come out with the wrong sign and slightly wrong magnitude, failing the expected-value checks.
515Unvalidated row set / precision when reporting a single summary statistictaskinfiagent-dabench
Applies when
task -- the task asks for one numeric statistic (correlation, mean, test statistic, p-value, score) computed over two or more columns of a table, with a fixed rounding/format spec.
Pattern
The script loads the data and calls the statistic function directly, without printing or checking how many rows actually entered the computation (silent NaN/inf dropping by the library, coercion of a string/mixed-dtype column to numeric with errors dropped, reading only a sheet/chunk/subset, or an inherited filter from an earlier step). The reported value is then a plausible-looking near-miss, and the printed value is transcribed at whatever precision Python happened to show rather than the precision the answer format demands.
Detection procedure
  1. Read the task: note the exact population of rows implied (all rows unless a filter is stated) and the required rounding/format for each reported quantity.
  2. Read the script: check that it prints the raw row count, the count after any dtype conversion/NaN handling, and the dtypes of the two columns, and that any dropping is explicit and justified — not left to library defaults or an unexamined read_* argument.
  3. Check that the statistic is computed on the full intended column pair (same aligned rows for both) and that the final numbers are formatted with the requested number of decimals (e.g. a p-value printed to the specified decimal places, not truncated to 0.0 or left in scientific notation).
  4. Flag if the answer reports a value with no accompanying evidence of sample size / no cross-check (e.g. recomputing with a second library or with explicit dropna on the pair).
Discriminator
A fine attempt explicitly reports N used vs N total and shows the drop reason, and its rounding matches the spec; a violation is one where N is never printed, dtype handling is implicit, or a reported figure's precision differs from the stated format — even if the number "looks right".
Consequence
The graded statistic differs from ground truth in the last reported digit (e.g. 0.53 vs 0.54) because a handful of rows were silently excluded or included, and/or a formatted field fails an exact-match check, so the answer is scored wrong despite the qualitative conclusion being right.
id 10bd9cee0182 · mined from infiagent-dabench dabench-300@s9
raw text (what the judge reads)
### Unvalidated row set / precision when reporting a single summary statistic
- **Applies when**: `task` -- the task asks for one numeric statistic (correlation, mean, test statistic, p-value, score) computed over two or more columns of a table, with a fixed rounding/format spec.
- **Pattern**: The script loads the data and calls the statistic function directly, without printing or checking how many rows actually entered the computation (silent NaN/`inf` dropping by the library, coercion of a string/mixed-dtype column to numeric with errors dropped, reading only a sheet/chunk/subset, or an inherited filter from an earlier step). The reported value is then a plausible-looking near-miss, and the printed value is transcribed at whatever precision Python happened to show rather than the precision the answer format demands.
- **Detection procedure**:
  1. Read the task: note the exact population of rows implied (all rows unless a filter is stated) and the required rounding/format for each reported quantity.
  2. Read the script: check that it prints the raw row count, the count after any dtype conversion/NaN handling, and the dtypes of the two columns, and that any dropping is explicit and justified — not left to library defaults or an unexamined `read_*` argument.
  3. Check that the statistic is computed on the full intended column pair (same aligned rows for both) and that the final numbers are formatted with the requested number of decimals (e.g. a p-value printed to the specified decimal places, not truncated to `0.0` or left in scientific notation).
  4. Flag if the answer reports a value with no accompanying evidence of sample size / no cross-check (e.g. recomputing with a second library or with explicit `dropna` on the pair).
- **Discriminator**: A fine attempt explicitly reports N used vs N total and shows the drop reason, and its rounding matches the spec; a violation is one where N is never printed, dtype handling is implicit, or a reported figure's precision differs from the stated format — even if the number "looks right".
- **Consequence**: The graded statistic differs from ground truth in the last reported digit (e.g. 0.53 vs 0.54) because a handful of rows were silently excluded or included, and/or a formatted field fails an exact-match check, so the answer is scored wrong despite the qualitative conclusion being right.
516In-sample metrics used as the only evidence of predictive qualitytaskda-code
Applies when
task -- the deliverable is a set of predictions for a held-out file and the scripts fit one or more models and then report error/accuracy figures.
Pattern
The attempt evaluates the fitted model on the very rows it was trained on (or on the whole training file) with no held-out split, cross-validation, or comparison against a trivial baseline; blending weights and model choices are then justified by these in-sample numbers, and the reported errors/R² are presented as proof the submission is good.
Detection procedure
  1. In the task, note that scoring happens on unseen rows, so the answer must be backed by an estimate of out-of-sample error.
  2. In the scripts, locate every fit(...) and every metric call and check whether the data passed to the metric is disjoint from the data passed to fit (explicit train/validation split, cross_val_score, or an out-of-fold construction).
  3. Check whether any reported metric is compared to a simple baseline (e.g., predicting a group/global central value) and whether target/feature handling that could inflate in-sample fit (duplicated rows, ID-like or near-target features, encoders fit on all data) is examined.
  4. In the answer, check whether the quoted error figures are labeled as training-set numbers, and whether the higher-weighted model is the one with the worse metric — a sign the weights were not validated at all.
  5. Verify a sanity check of the output itself (row count equals test rows, order preserved, value range plausible against the training target distribution).
Discriminator
A real violation is when no metric in the scripts is computed on data withheld from fitting, so nothing constrains generalization error; it is fine if the agent trains a final model on all data after having measured error on a proper holdout/CV and reports that holdout number.
Consequence
The submitted predictions can be far worse than the quoted error suggests (overfit or leakage-inflated fit), so the grader's threshold on the true held-out metric fails while the answer claims strong performance.
id 5c6746478ada · mined from da-code dacode-ml-regression-014@s9
raw text (what the judge reads)
### In-sample metrics used as the only evidence of predictive quality
- **Applies when**: `task` -- the deliverable is a set of predictions for a held-out file and the scripts fit one or more models and then report error/accuracy figures.
- **Pattern**: The attempt evaluates the fitted model on the very rows it was trained on (or on the whole training file) with no held-out split, cross-validation, or comparison against a trivial baseline; blending weights and model choices are then justified by these in-sample numbers, and the reported errors/R² are presented as proof the submission is good.
- **Detection procedure**:
  1. In the task, note that scoring happens on unseen rows, so the answer must be backed by an estimate of out-of-sample error.
  2. In the scripts, locate every `fit(...)` and every metric call and check whether the data passed to the metric is disjoint from the data passed to `fit` (explicit train/validation split, `cross_val_score`, or an out-of-fold construction).
  3. Check whether any reported metric is compared to a simple baseline (e.g., predicting a group/global central value) and whether target/feature handling that could inflate in-sample fit (duplicated rows, ID-like or near-target features, encoders fit on all data) is examined.
  4. In the answer, check whether the quoted error figures are labeled as training-set numbers, and whether the higher-weighted model is the one with the *worse* metric — a sign the weights were not validated at all.
  5. Verify a sanity check of the output itself (row count equals test rows, order preserved, value range plausible against the training target distribution).
- **Discriminator**: A real violation is when *no* metric in the scripts is computed on data withheld from fitting, so nothing constrains generalization error; it is fine if the agent trains a final model on all data *after* having measured error on a proper holdout/CV and reports that holdout number.
- **Consequence**: The submitted predictions can be far worse than the quoted error suggests (overfit or leakage-inflated fit), so the grader's threshold on the true held-out metric fails while the answer claims strong performance.
517Ignoring the provided output template's schema and category labelstaskda-code
Applies when
task -- the task says results must be written into a provided/pre-existing output file "adhering strictly to its format", and the scripts generate that file from scratch.
Pattern
The attempt never opens or inspects the supplied template; it invents its own column names, its own category/group labels (via self-chosen binning thresholds), its own row set and ordering, plus extra catch-all groups for missing values, and overwrites the template with this ad-hoc table. The reported prose summary is treated as the deliverable.
Detection procedure
  1. In the task, note the named result file and any statement that its existing format must be preserved.
  2. Search the scripts for any read/inspection of that file (e.g., loading it, printing its header/rows) before writing; if the only interaction is a write, flag it.
  3. Compare the labels/columns/row count the scripts produce against what the template prescribes: are group names and their definitions taken from the template (or an authoritative reference), or fabricated by the agent's own thresholds? Are extra rows added for un-mappable/missing records?
  4. Check the final answer: does it state that the file was filled in with exactly the template's rows/columns, or does it only present a free-text summary with the agent's own labels?
Discriminator
A real violation is inventing schema/categories without ever reading the template (row labels, column names, or ordering can differ from it). It is fine if the script loads the template, keeps its exact columns and row labels, and only fills in the computed values — even if the internal derivation uses reasonable thresholds — or if the task provides no template at all.
Consequence
The graded file fails an exact-match check on keys/columns/row set (unmatched or extra category rows, mismatched header), so the deliverable is scored WRONG even if the underlying counting logic were defensible.
id 5eed8e6b992d · mined from da-code dacode-dm-csv-001@s9
raw text (what the judge reads)
### Ignoring the provided output template's schema and category labels
- **Applies when**: `task` -- the task says results must be written into a provided/pre-existing output file "adhering strictly to its format", and the scripts generate that file from scratch.
- **Pattern**: The attempt never opens or inspects the supplied template; it invents its own column names, its own category/group labels (via self-chosen binning thresholds), its own row set and ordering, plus extra catch-all groups for missing values, and overwrites the template with this ad-hoc table. The reported prose summary is treated as the deliverable.
- **Detection procedure**:
  1. In the task, note the named result file and any statement that its existing format must be preserved.
  2. Search the scripts for any read/inspection of that file (e.g., loading it, printing its header/rows) before writing; if the only interaction is a write, flag it.
  3. Compare the labels/columns/row count the scripts produce against what the template prescribes: are group names and their definitions taken from the template (or an authoritative reference), or fabricated by the agent's own thresholds? Are extra rows added for un-mappable/missing records?
  4. Check the final answer: does it state that the file was filled in with exactly the template's rows/columns, or does it only present a free-text summary with the agent's own labels?
- **Discriminator**: A real violation is inventing schema/categories without ever reading the template (row labels, column names, or ordering can differ from it). It is fine if the script loads the template, keeps its exact columns and row labels, and only fills in the computed values — even if the internal derivation uses reasonable thresholds — or if the task provides no template at all.
- **Consequence**: The graded file fails an exact-match check on keys/columns/row set (unmatched or extra category rows, mismatched header), so the deliverable is scored WRONG even if the underlying counting logic were defensible.
518Output file not verified against the provided result templatetaskda-code
Applies when
task -- the task says to save the result "in the format of" a supplied sample/template file (e.g., sample_result.csv) to a specific filename.
Pattern
The attempt computes a number, prints/narrates it in prose, and writes some ad-hoc file (extra columns, different header names, extra rows/index, different value precision or ordering) without ever opening the template to mirror its exact schema; no post-write read-back check is done, and the final answer reports the statistic in text rather than showing the saved file contents.
Detection procedure
  1. Read the task for the required output filename and the referenced template; confirm the scripts actually load/inspect that template (read its header, row count, column names, dtypes/rounding).
  2. Inspect the writing code: are column names, column order, row count, index behavior (index=False), and value formatting derived from the template, or hard-coded/invented?
  3. Check for a verification step after writing: re-read the saved file and compare its shape/columns to the template; assert the stored statistic is within a valid range (e.g., a p-value in [0, 1]) and matches the reported value.
  4. Check that the final answer displays the actual saved file contents, not only a prose summary of an intermediate computation.
Discriminator
A real violation is when the template is never read or the written schema demonstrably deviates (different/added columns, index column, wrong number of rows, value rounded differently than the template implies). A look-alike that is fine is a script that reads the template, reproduces its exact header/shape, and re-reads the output to confirm — even if the surrounding narrative is verbose or the analysis method is described only briefly.
Consequence
The grader compares result.csv against the expected file and reports WRONG/MISSING even when the underlying computation is close or correct, yielding 0/1 checks passed.
id cabd5c93a70f · mined from da-code dacode-data-sa-039@s9
raw text (what the judge reads)
### Output file not verified against the provided result template
- **Applies when**: `task` -- the task says to save the result "in the format of" a supplied sample/template file (e.g., `sample_result.csv`) to a specific filename.
- **Pattern**: The attempt computes a number, prints/narrates it in prose, and writes some ad-hoc file (extra columns, different header names, extra rows/index, different value precision or ordering) without ever opening the template to mirror its exact schema; no post-write read-back check is done, and the final answer reports the statistic in text rather than showing the saved file contents.
- **Detection procedure**:
  1. Read the task for the required output filename and the referenced template; confirm the scripts actually load/inspect that template (read its header, row count, column names, dtypes/rounding).
  2. Inspect the writing code: are column names, column order, row count, index behavior (`index=False`), and value formatting derived from the template, or hard-coded/invented?
  3. Check for a verification step after writing: re-read the saved file and compare its shape/columns to the template; assert the stored statistic is within a valid range (e.g., a p-value in [0, 1]) and matches the reported value.
  4. Check that the final answer displays the actual saved file contents, not only a prose summary of an intermediate computation.
- **Discriminator**: A real violation is when the template is never read or the written schema demonstrably deviates (different/added columns, index column, wrong number of rows, value rounded differently than the template implies). A look-alike that is fine is a script that reads the template, reproduces its exact header/shape, and re-reads the output to confirm — even if the surrounding narrative is verbose or the analysis method is described only briefly.
- **Consequence**: The grader compares `result.csv` against the expected file and reports WRONG/MISSING even when the underlying computation is close or correct, yielding 0/1 checks passed.
519Discarding the identifier/time column and shipping a weak model without checking test-row provenancetaskda-code
Applies when
task -- a full "complete" source file is provided alongside a separate test file, and the scripts drop a timestamp/ID-like column before training and validate with a plain random split.
Pattern
The agent treats the task as a generic i.i.d. tabular regression: it drops the date/key column as "non-numeric", never checks whether the test rows are already present in (or contiguous with) the provided complete file, uses train_test_split(shuffle=True) on sequentially ordered data, and then accepts a mediocre validation score (low R², RMSE comparable to the target's own std) as final, submitting predictions with no comparison against the target's known distribution.
Detection procedure
  1. Read the task and list the input files; note whether a superset/"complete" file is supplied along with the test file, and whether the test file carries a key column (timestamp, id, index) shared with it.
  2. In the scripts, check for a merge/overlap/duplicate check between test rows and the complete file (e.g. joining on the key or on the full feature vector). If none exists and the key column is simply dropped, flag it.
  3. Check the validation design: is the split random over data with an inherent order (time series/sequence)? Is any time-derived feature (hour, weekday, lag) engineered from the dropped key? If not, the model is thrown away most of the signal and the reported score is optimistically or pessimistically mis-specified.
  4. Read the reported metrics and prediction summary: compare RMSE/R² against a trivial baseline (predicting the training mean) and compare predicted min/max/std against the target's actual min/max/std. Flag if the model barely beats the baseline or if predicted spread is far narrower than the target's.
Discriminator
A real violation is when the shared key/ordering was available and unused, and the reported error is close to the naive-baseline error (or the prediction spread collapses relative to the target). It is fine if the agent explicitly tested for overlap and found none, justified dropping the key, used an order-respecting validation scheme, and demonstrated a clear margin over the baseline.
Consequence
The saved prediction file has the right shape and column name but values that miss the true per-row values by a wide margin, so any correctness check based on row-wise agreement or an error/accuracy threshold against the held-out targets fails, even though the script "ran successfully".
id 2b074b6e18cd · mined from da-code dacode-ml-regression-015@s9
raw text (what the judge reads)
### Discarding the identifier/time column and shipping a weak model without checking test-row provenance
- **Applies when**: `task` -- a full "complete" source file is provided alongside a separate test file, and the scripts drop a timestamp/ID-like column before training and validate with a plain random split.
- **Pattern**: The agent treats the task as a generic i.i.d. tabular regression: it drops the date/key column as "non-numeric", never checks whether the test rows are already present in (or contiguous with) the provided complete file, uses `train_test_split(shuffle=True)` on sequentially ordered data, and then accepts a mediocre validation score (low R², RMSE comparable to the target's own std) as final, submitting predictions with no comparison against the target's known distribution.
- **Detection procedure**:
  1. Read the task and list the input files; note whether a superset/"complete" file is supplied along with the test file, and whether the test file carries a key column (timestamp, id, index) shared with it.
  2. In the scripts, check for a merge/overlap/duplicate check between test rows and the complete file (e.g. joining on the key or on the full feature vector). If none exists and the key column is simply dropped, flag it.
  3. Check the validation design: is the split random over data with an inherent order (time series/sequence)? Is any time-derived feature (hour, weekday, lag) engineered from the dropped key? If not, the model is thrown away most of the signal and the reported score is optimistically or pessimistically mis-specified.
  4. Read the reported metrics and prediction summary: compare RMSE/R² against a trivial baseline (predicting the training mean) and compare predicted min/max/std against the target's actual min/max/std. Flag if the model barely beats the baseline or if predicted spread is far narrower than the target's.
- **Discriminator**: A real violation is when the shared key/ordering was available and unused, and the reported error is close to the naive-baseline error (or the prediction spread collapses relative to the target). It is fine if the agent explicitly tested for overlap and found none, justified dropping the key, used an order-respecting validation scheme, and demonstrated a clear margin over the baseline.
- **Consequence**: The saved prediction file has the right shape and column name but values that miss the true per-row values by a wide margin, so any correctness check based on row-wise agreement or an error/accuracy threshold against the held-out targets fails, even though the script "ran successfully".
520Numeric fields stored as text are ranked/imputed without parsingtaskda-code
Applies when
task -- the task asks for extreme values (max/min/top-k) or a statistic-based imputation on a column that arrives from CSV/Excel as strings containing thousands separators, %, currency symbols, or units.
Pattern
The attempt loads the file and directly calls fillna(mean), idxmax/idxmin, sort_values, or nlargest on the raw column. Because the dtype is object, the mean-imputation silently skips the column and the ordering is lexicographic (e.g. "9" > "10,000"), so the reported extremes are plausible-looking but wrong entries; no script or intermediate output is retained to expose this.
Detection procedure
  1. Read the task to identify which column(s) the requested statistic and the required imputation act on, and note any stated preprocessing constraint.
  2. In the scripts, check for an explicit cleaning/casting step for those columns (strip separators/symbols, pd.to_numeric, astype(float)) and a dtype/NaN-count assertion before the aggregation; check that the imputation actually changed the intended column.
  3. Check that the reported extremes are backed by printed numeric values (the max and min magnitudes), not just labels, and that the required output file/format is written.
  4. Sanity-check the reported entities against domain expectation for the metric's scale (is the "lowest" plausibly the smallest by orders of magnitude, or merely a string that sorts early?).
Discriminator
A real violation is when no parsing/dtype check exists and the answer is reported without the underlying numeric values; it is fine if the loader already yields a float dtype (verifiable via an explicit dtype print/assert) or the script cleans the strings before aggregating, even if the code is terse.
Consequence
The grader compares against extremes computed on properly parsed numbers and the reported country/label(s) mismatch — plus a missing/incorrectly written result file — scoring 0.
id 36cd1caf9907 · mined from da-code dacode-di-text-001@s9
raw text (what the judge reads)
### Numeric fields stored as text are ranked/imputed without parsing
- **Applies when**: `task` -- the task asks for extreme values (max/min/top-k) or a statistic-based imputation on a column that arrives from CSV/Excel as strings containing thousands separators, `%`, currency symbols, or units.
- **Pattern**: The attempt loads the file and directly calls `fillna(mean)`, `idxmax`/`idxmin`, `sort_values`, or `nlargest` on the raw column. Because the dtype is `object`, the mean-imputation silently skips the column and the ordering is lexicographic (e.g. "9" > "10,000"), so the reported extremes are plausible-looking but wrong entries; no script or intermediate output is retained to expose this.
- **Detection procedure**:
  1. Read the task to identify which column(s) the requested statistic and the required imputation act on, and note any stated preprocessing constraint.
  2. In the scripts, check for an explicit cleaning/casting step for those columns (strip separators/symbols, `pd.to_numeric`, `astype(float)`) and a dtype/NaN-count assertion before the aggregation; check that the imputation actually changed the intended column.
  3. Check that the reported extremes are backed by printed numeric values (the max and min magnitudes), not just labels, and that the required output file/format is written.
  4. Sanity-check the reported entities against domain expectation for the metric's scale (is the "lowest" plausibly the smallest by orders of magnitude, or merely a string that sorts early?).
- **Discriminator**: A real violation is when no parsing/dtype check exists and the answer is reported without the underlying numeric values; it is fine if the loader already yields a float dtype (verifiable via an explicit dtype print/assert) or the script cleans the strings before aggregating, even if the code is terse.
- **Consequence**: The grader compares against extremes computed on properly parsed numbers and the reported country/label(s) mismatch — plus a missing/incorrectly written result file — scoring 0.
521Output file omits requested fields (over-trimmed result schema)taskda-code
Applies when
task -- The task names a specific output file and enumerates multiple things it must contain (e.g., derived scores, group/segment assignments, and a final label), and the script builds a full intermediate table before writing.
Pattern
The script computes all the requested quantities in memory but writes only a minimal subset (identifier + one final label) to the required file, pushing the other requested columns into an unrequested side file or leaving them in stdout; the answer then asserts correctness (sometimes with an unverifiable "validated against reference" claim) without checking the saved file's schema against the task wording.
Detection procedure
  1. Re-read the task sentence describing the deliverable and list every quantity it says must be "included" in the saved file, plus any stated naming/ordering/format constraints.
  2. Find the exact to_csv/to_json/save call for that filename in the script and enumerate the columns actually selected there (watch for a df[[...]] subset immediately before saving).
  3. Diff the two lists; also check whether extra "detailed" files were created to hold the dropped columns, which signals the agent knew they existed but excluded them.
  4. Check whether the answer's correctness claim is backed by an actual comparison the script performs, or is merely asserted.
Discriminator
A real violation is when quantities explicitly named as deliverables in the task are absent from the required file (or renamed beyond recognition). It is not a violation if the task only asks for the final label and the extra columns are genuinely optional, or if the file contains the requested fields plus harmless additional columns.
Consequence
The grader reads the required file, fails to find the expected columns/values, and marks the deliverable WRONG/MISSING even though the underlying computation may have been reasonable.
id de04cf930484 · mined from da-code dacode-dm-csv-052@s9
raw text (what the judge reads)
### Output file omits requested fields (over-trimmed result schema)
- **Applies when**: `task` -- The task names a specific output file and enumerates multiple things it must contain (e.g., derived scores, group/segment assignments, and a final label), and the script builds a full intermediate table before writing.
- **Pattern**: The script computes all the requested quantities in memory but writes only a minimal subset (identifier + one final label) to the required file, pushing the other requested columns into an unrequested side file or leaving them in stdout; the answer then asserts correctness (sometimes with an unverifiable "validated against reference" claim) without checking the saved file's schema against the task wording.
- **Detection procedure**:
  1. Re-read the task sentence describing the deliverable and list every quantity it says must be "included" in the saved file, plus any stated naming/ordering/format constraints.
  2. Find the exact `to_csv`/`to_json`/save call for that filename in the script and enumerate the columns actually selected there (watch for a `df[[...]]` subset immediately before saving).
  3. Diff the two lists; also check whether extra "detailed" files were created to hold the dropped columns, which signals the agent knew they existed but excluded them.
  4. Check whether the answer's correctness claim is backed by an actual comparison the script performs, or is merely asserted.
- **Discriminator**: A real violation is when quantities explicitly named as deliverables in the task are absent from the required file (or renamed beyond recognition). It is *not* a violation if the task only asks for the final label and the extra columns are genuinely optional, or if the file contains the requested fields plus harmless additional columns.
- **Consequence**: The grader reads the required file, fails to find the expected columns/values, and marks the deliverable WRONG/MISSING even though the underlying computation may have been reasonable.
522Output file not built from the provided format templatetaskda-code
Applies when
task -- The task says to write results to a named output file "following the format specified in" a provided sample/template file.
Pattern
The agent computes plausible numbers but constructs the output file from its own invented schema (own column names, row order, orientation, units, rounding, extra prose/summary rows), never opening or echoing the supplied sample file to mirror its exact header, row labels, and cell layout; the final answer reports numbers in free text rather than demonstrating the written file's contents.
Detection procedure
1. In the task statement, note the exact output filename and the referenced template file. 2. Search the scripts for a read/inspection of the template (e.g., loading it, printing its header/rows) and for code that writes the output using that exact schema. 3. Compare the written columns/rows/order/precision against the template's; if the template was never read, treat the schema as unverified. 4. Check the final answer includes a dump of the produced file (or at least its header and rows) rather than only a narrative of the values.
Discriminator
A real violation is when the schema is invented or unverified against the template; it is fine if the script reads the template (or reproduces its literal header/row labels) and writes matching keys/order, even if the prose summary is verbose or values differ slightly.
Consequence
The grader compares the output file cell-by-cell against the expected file and reports it as WRONG/MISSING, scoring 0 even when the underlying statistics are approximately right.
id e1484ff4e75f · mined from da-code dacode-data-sa-029@s9
raw text (what the judge reads)
### Output file not built from the provided format template
- **Applies when**: `task` -- The task says to write results to a named output file "following the format specified in" a provided sample/template file.
- **Pattern**: The agent computes plausible numbers but constructs the output file from its own invented schema (own column names, row order, orientation, units, rounding, extra prose/summary rows), never opening or echoing the supplied sample file to mirror its exact header, row labels, and cell layout; the final answer reports numbers in free text rather than demonstrating the written file's contents.
- **Detection procedure**: 1. In the task statement, note the exact output filename and the referenced template file. 2. Search the scripts for a read/inspection of the template (e.g., loading it, printing its header/rows) and for code that writes the output using that exact schema. 3. Compare the written columns/rows/order/precision against the template's; if the template was never read, treat the schema as unverified. 4. Check the final answer includes a dump of the produced file (or at least its header and rows) rather than only a narrative of the values.
- **Discriminator**: A real violation is when the schema is invented or unverified against the template; it is fine if the script reads the template (or reproduces its literal header/row labels) and writes matching keys/order, even if the prose summary is verbose or values differ slightly.
- **Consequence**: The grader compares the output file cell-by-cell against the expected file and reports it as WRONG/MISSING, scoring 0 even when the underlying statistics are approximately right.
523Blindly ingesting an arbitrarily-chosen input file and not sanity-checking the result against expectationstaskda-code
Applies when
task -- the task names a specific dataset (and often a sample output file), and the script locates its input by globbing/listing a directory and taking the first match rather than by identifying the intended file.
Pattern
The script resolves the data path with os.listdir/glob + [0] (or a guessed filename list), never verifies that the loaded table is the intended dataset (row/column counts, expected column set, plausible value ranges), never reads the provided sample output to confirm formatting/ordering, and then reports whatever numbers come out — even when they are implausible (e.g. essentially zero correlations among quantities that should co-move, or an unexpected number of rows/levels).
Detection procedure
  1. From the task, list the identifying properties of the intended input (file name/dataset, the required columns, any provided sample output file to mirror).
  2. In the scripts, check how the input path is chosen and whether any assertion/print-and-stop validates the loaded object against those properties (expected columns present with exact names, expected magnitude of rows, sample-format comparison).
  3. In the answer, compare reported shapes, dropped-row counts and statistic values to domain expectations; flag if the script proceeded despite missing/fuzzy-matched columns or if the values look degenerate (near-zero, constant, out of range).
  4. Flag the attempt if input selection is positional/guessed AND no post-load or post-result sanity check would have caught loading the wrong file or wrong columns.
Discriminator
Fine if the directory demonstrably contains exactly one candidate file and the script asserts the expected columns/shape (or the agent inspected the file and confirmed identity, and the resulting statistics are plausible). A violation is taking files[0]/case-insensitive fuzzy column matches with no verification, and accepting degenerate-looking output without investigation.
Consequence
The saved output is computed from the wrong file, wrong columns, or wrong subset, so result.csv values (and/or its header/index format versus the sample) do not match the reference and the file check fails outright.
id f444f76447dd · mined from da-code dacode-data-sa-026@s9
raw text (what the judge reads)
### Blindly ingesting an arbitrarily-chosen input file and not sanity-checking the result against expectations
- **Applies when**: `task` -- the task names a specific dataset (and often a sample output file), and the script locates its input by globbing/listing a directory and taking the first match rather than by identifying the intended file.
- **Pattern**: The script resolves the data path with `os.listdir`/`glob` + `[0]` (or a guessed filename list), never verifies that the loaded table is the intended dataset (row/column counts, expected column set, plausible value ranges), never reads the provided sample output to confirm formatting/ordering, and then reports whatever numbers come out — even when they are implausible (e.g. essentially zero correlations among quantities that should co-move, or an unexpected number of rows/levels).
- **Detection procedure**:
  1. From the task, list the identifying properties of the intended input (file name/dataset, the required columns, any provided sample output file to mirror).
  2. In the scripts, check how the input path is chosen and whether any assertion/print-and-stop validates the loaded object against those properties (expected columns present with exact names, expected magnitude of rows, sample-format comparison).
  3. In the answer, compare reported shapes, dropped-row counts and statistic values to domain expectations; flag if the script proceeded despite missing/fuzzy-matched columns or if the values look degenerate (near-zero, constant, out of range).
  4. Flag the attempt if input selection is positional/guessed AND no post-load or post-result sanity check would have caught loading the wrong file or wrong columns.
- **Discriminator**: Fine if the directory demonstrably contains exactly one candidate file and the script asserts the expected columns/shape (or the agent inspected the file and confirmed identity, and the resulting statistics are plausible). A violation is taking `files[0]`/case-insensitive fuzzy column matches with no verification, and accepting degenerate-looking output without investigation.
- **Consequence**: The saved output is computed from the wrong file, wrong columns, or wrong subset, so `result.csv` values (and/or its header/index format versus the sample) do not match the reference and the file check fails outright.
524Silent subsampling of the dataset when full-data output is requiredtaskda-code
Applies when
task -- the task asks for a per-record output artifact (labels, predictions, scores) covering the provided dataset, and the scripts load a large input file.
Pattern
The agent samples or truncates the data (e.g., nrows=, .sample(n), head(), a hard-coded cap) for speed, fits and writes results only for that subset, and reports success without noting that the output row count is far smaller than the input row count.
Detection procedure
  1. Read the task for whether the deliverable is expected to cover every input record; note the documented/actual size of the input data.
  2. Scan the scripts for any row-limiting call at load or before fitting/writing (nrows, sample, iloc[:N], train_test_split used only to shrink, chunked read that keeps one chunk).
  3. Compare the row count the answer reports for the saved file against the input dataset's row count; also confirm column names/order match the requested schema exactly.
  4. Confirm the output file is written to the location the task/grader expects, not a private working directory.
Discriminator
A real violation is when the reduced set becomes the delivered artifact; it is fine to subsample only for exploratory steps (e.g., hyperparameter or k selection) as long as the final model is applied to and written out for all records, or the task explicitly permits a sample.
Consequence
The saved file has far fewer rows than expected, so row-wise comparison, shape checks, or joins against the reference fail and the deliverable is scored missing/wrong regardless of clustering quality.
id 0e5fca08490d · mined from da-code dacode-ml-cluster-010@s9
raw text (what the judge reads)
### Silent subsampling of the dataset when full-data output is required
- **Applies when**: `task` -- the task asks for a per-record output artifact (labels, predictions, scores) covering the provided dataset, and the scripts load a large input file.
- **Pattern**: The agent samples or truncates the data (e.g., `nrows=`, `.sample(n)`, `head()`, a hard-coded cap) for speed, fits and writes results only for that subset, and reports success without noting that the output row count is far smaller than the input row count.
- **Detection procedure**:
  1. Read the task for whether the deliverable is expected to cover every input record; note the documented/actual size of the input data.
  2. Scan the scripts for any row-limiting call at load or before fitting/writing (`nrows`, `sample`, `iloc[:N]`, `train_test_split` used only to shrink, chunked read that keeps one chunk).
  3. Compare the row count the answer reports for the saved file against the input dataset's row count; also confirm column names/order match the requested schema exactly.
  4. Confirm the output file is written to the location the task/grader expects, not a private working directory.
- **Discriminator**: A real violation is when the reduced set becomes the delivered artifact; it is fine to subsample only for exploratory steps (e.g., hyperparameter or k selection) as long as the final model is applied to and written out for all records, or the task explicitly permits a sample.
- **Consequence**: The saved file has far fewer rows than expected, so row-wise comparison, shape checks, or joins against the reference fail and the deliverable is scored missing/wrong regardless of clustering quality.
525Requested output artifact never written (or written with the wrong schema)taskda-code
Applies when
task -- the task explicitly requires results to be persisted to a named file with specified column names/format, and the agent's deliverable is a prose summary of the analysis.
Pattern
The attempt performs the analysis and narrates methodology and cluster/model statistics, but no script step actually writes the required file to the working directory, or it writes a file whose name, columns, or row count differ from the specification (e.g., renamed columns, extra index column, only aggregated group summaries instead of one row per input record).
Detection procedure
1. Extract from the task statement the exact required filename, column names, and implied row granularity. 2. Search the scripts for a write call (to_csv, savetxt, etc.) targeting that exact path, and check the DataFrame passed to it for the required column names and expected number of rows. 3. Check the reported answer for evidence the file exists (path confirmed, head of the file, shape printed) rather than only narrative results. 4. If no save step or no schema match is present, flag as inadequate.
Discriminator
A real violation is missing/misnamed/mis-schema'd output, or output whose rows do not correspond to the units the task asks to label. A look-alike that is fine: the file is written correctly with the specified name and columns and the prose is merely additional commentary; harmless differences such as column order or dtype that still satisfy the stated names and granularity.
Consequence
The grader checking for the named result file reports it WRONG/MISSING and scores 0, regardless of how sound the underlying analysis was.
id 8ab29067634a · mined from da-code dacode-ml-cluster-019@s9
raw text (what the judge reads)
### Requested output artifact never written (or written with the wrong schema)
- **Applies when**: `task` -- the task explicitly requires results to be persisted to a named file with specified column names/format, and the agent's deliverable is a prose summary of the analysis.
- **Pattern**: The attempt performs the analysis and narrates methodology and cluster/model statistics, but no script step actually writes the required file to the working directory, or it writes a file whose name, columns, or row count differ from the specification (e.g., renamed columns, extra index column, only aggregated group summaries instead of one row per input record).
- **Detection procedure**: 1. Extract from the task statement the exact required filename, column names, and implied row granularity. 2. Search the scripts for a write call (`to_csv`, `savetxt`, etc.) targeting that exact path, and check the DataFrame passed to it for the required column names and expected number of rows. 3. Check the reported answer for evidence the file exists (path confirmed, head of the file, shape printed) rather than only narrative results. 4. If no save step or no schema match is present, flag as inadequate.
- **Discriminator**: A real violation is missing/misnamed/mis-schema'd output, or output whose rows do not correspond to the units the task asks to label. A look-alike that is fine: the file is written correctly with the specified name and columns and the prose is merely additional commentary; harmless differences such as column order or dtype that still satisfy the stated names and granularity.
- **Consequence**: The grader checking for the named result file reports it WRONG/MISSING and scores 0, regardless of how sound the underlying analysis was.
526Wrong definition or units for the requested effect statistic (raw counts vs. normalized rate)taskda-code
Applies when
task -- the task asks for a confidence interval (or other estimate) of a "reduction"/"difference"/"effect" without spelling out the formula, and the data contain both an outcome count and the denominator (exposure, population, total observations) needed to normalize it.
Pattern
The agent picks the most convenient raw column (e.g., an absolute count per period) as the outcome, resamples it, and reports an interval in count units, ignoring that the intended quantity is a rate/proportion (or percentage-point) difference computed from outcome ÷ denominator; it also never checks that the written output file's schema/units match the provided format template.
Detection procedure
  1. Read the task and any provided format/template file: note the exact quantity, its units, and the expected column names/rows of the output artifact.
  2. In the script, locate the outcome variable used for the statistic; check whether a denominator column exists in the data and whether it is used for normalization or aggregated (sum) before differencing.
  3. Compare the magnitude/units of the reported bounds to what the requested quantity could plausibly be (e.g., a proportion difference must lie in [-1, 1]; percentage points in [-100, 100]); flag if the bounds are in unnormalized count units.
  4. Confirm the answer text and the written file agree with the template's schema and units, not just with each other.
Discriminator
A real violation is when a denominator/exposure column is present (or the task's units imply a rate) yet the statistic is computed on raw counts, or the interval's magnitude is impossible for the requested unit. A look-alike that is fine: the data genuinely have no denominator and the task explicitly asks for an absolute-count difference, with the interval documented in those units.
Consequence
The interval bounds are off by orders of magnitude from the expected values, so the output-file check fails even though the bootstrap machinery itself ran correctly.
id 412e4af3a256 · mined from da-code dacode-data-sa-031@s9
raw text (what the judge reads)
### Wrong definition or units for the requested effect statistic (raw counts vs. normalized rate)
- **Applies when**: `task` -- the task asks for a confidence interval (or other estimate) of a "reduction"/"difference"/"effect" without spelling out the formula, and the data contain both an outcome count and the denominator (exposure, population, total observations) needed to normalize it.
- **Pattern**: The agent picks the most convenient raw column (e.g., an absolute count per period) as the outcome, resamples it, and reports an interval in count units, ignoring that the intended quantity is a rate/proportion (or percentage-point) difference computed from outcome ÷ denominator; it also never checks that the written output file's schema/units match the provided format template.
- **Detection procedure**:
  1. Read the task and any provided format/template file: note the exact quantity, its units, and the expected column names/rows of the output artifact.
  2. In the script, locate the outcome variable used for the statistic; check whether a denominator column exists in the data and whether it is used for normalization or aggregated (sum) before differencing.
  3. Compare the magnitude/units of the reported bounds to what the requested quantity could plausibly be (e.g., a proportion difference must lie in [-1, 1]; percentage points in [-100, 100]); flag if the bounds are in unnormalized count units.
  4. Confirm the answer text and the written file agree with the template's schema and units, not just with each other.
- **Discriminator**: A real violation is when a denominator/exposure column is present (or the task's units imply a rate) yet the statistic is computed on raw counts, or the interval's magnitude is impossible for the requested unit. A look-alike that is fine: the data genuinely have no denominator and the task explicitly asks for an absolute-count difference, with the interval documented in those units.
- **Consequence**: The interval bounds are off by orders of magnitude from the expected values, so the output-file check fails even though the bootstrap machinery itself ran correctly.
527Deliverable file completeness: predictions not written for every required row/IDtaskda-code
Applies when
task -- the task requires producing an output file whose rows must correspond one-to-one with the rows/IDs of a provided evaluation set, in a given column format.
Pattern
The attempt reports predictions as inline text (or writes a partial/truncated file) instead of programmatically generating the full artifact — the output ends mid-number, contains far fewer rows than the evaluation set, IDs are unordered/unverified against the evaluation set, and no script exists that reads the evaluation file, predicts for all of it, and saves the file.
Detection procedure
  1. From the task/README, note the required output filename, header columns, and the number and identity of rows expected (i.e., every ID in the evaluation input, plus the sample/template file's structure).
  2. In the scripts, look for an explicit path: load evaluation input → predict for all its rows → assemble DataFrame with required columns → write to the required filename; check no row filtering/subsampling/head() limits the output.
  3. Check for a post-write sanity assertion: row count equals evaluation-set row count, ID set matches exactly, no NaNs, columns and names/order match the template.
  4. Inspect the reported answer/file: does it appear complete (ends with a well-formed final row), have the right row count, and hold valid values in range?
Discriminator
A real violation is missing rows, missing/unsaved file, truncated or malformed values, or IDs not matching the evaluation set. A look-alike that is fine is a complete file where only a preview of rows is echoed in the write-up, while the script demonstrably writes and validates all rows.
Consequence
The grader cannot score the submission (file missing/mismatched shape or unparsable values) and marks it wrong regardless of model quality; even if parsed, absent rows yield undefined or maximal log-loss penalties.
id 86ab780a8bc2 · mined from da-code dacode-ml-competition-005@s10
raw text (what the judge reads)
### Deliverable file completeness: predictions not written for every required row/ID
- **Applies when**: `task` -- the task requires producing an output file whose rows must correspond one-to-one with the rows/IDs of a provided evaluation set, in a given column format.
- **Pattern**: The attempt reports predictions as inline text (or writes a partial/truncated file) instead of programmatically generating the full artifact — the output ends mid-number, contains far fewer rows than the evaluation set, IDs are unordered/unverified against the evaluation set, and no script exists that reads the evaluation file, predicts for all of it, and saves the file.
- **Detection procedure**:
  1. From the task/README, note the required output filename, header columns, and the number and identity of rows expected (i.e., every ID in the evaluation input, plus the sample/template file's structure).
  2. In the scripts, look for an explicit path: load evaluation input → predict for all its rows → assemble DataFrame with required columns → write to the required filename; check no row filtering/subsampling/head() limits the output.
  3. Check for a post-write sanity assertion: row count equals evaluation-set row count, ID set matches exactly, no NaNs, columns and names/order match the template.
  4. Inspect the reported answer/file: does it appear complete (ends with a well-formed final row), have the right row count, and hold valid values in range?
- **Discriminator**: A real violation is missing rows, missing/unsaved file, truncated or malformed values, or IDs not matching the evaluation set. A look-alike that is fine is a complete file where only a *preview* of rows is echoed in the write-up, while the script demonstrably writes and validates all rows.
- **Consequence**: The grader cannot score the submission (file missing/mismatched shape or unparsable values) and marks it wrong regardless of model quality; even if parsed, absent rows yield undefined or maximal log-loss penalties.
528Prediction file not validated for full row coverage and non-degenerate value distributiontaskda-code
Applies when
task -- the deliverable is a per-row prediction file for a supplied test set, especially when the target is a skewed count/continuous quantity.
Pattern
The agent writes the output file (or pastes it as the answer) without asserting that it contains exactly one prediction per test row in the expected id order/format, and without comparing the predicted value distribution to the training target distribution; the result is a truncated/partial file or predictions collapsed to a near-constant value (e.g., mostly zeros or the mode) that no sanity check would have accepted.
Detection procedure
  1. In the task, note the required output filename, required column name(s), and the number of rows in the provided test file.
  2. In the scripts, look for an explicit check after prediction: len(pred) == len(test), ids matching test ids exactly (no drops from dropna/filtering/merging), and a written-file re-read verifying shape and header; absence of any such check is the first flag.
  3. In the scripts, look for a distribution sanity check on the predictions (min/max/mean/median/quantiles or share of a single value) compared against the same statistics of the training target; absence, or presence with a wildly mismatched profile that was not investigated, is the second flag.
  4. Inspect the answer itself: count rows, check for truncation/incomplete final line, and compute the fraction of identical values — a large majority of one value or far-too-small maximum relative to the training target range confirms the violation.
Discriminator
A genuinely fine attempt may still predict many low/zero values if the training target is truly dominated by them and the script demonstrates row-count/id alignment plus a printed comparison of predicted vs. actual target distributions; the violation is the absence of these verifications combined with an output whose row count or value spread does not match the test set and target.
Consequence
The grader finds the expected file missing rows, mis-formatted, or scored against a degenerate constant-like prediction, yielding a failed correctness check (or a metric far worse than a trivial baseline).
id a273dbd3f8b5 · mined from da-code dacode-ml-regression-008@s10
raw text (what the judge reads)
### Prediction file not validated for full row coverage and non-degenerate value distribution
- **Applies when**: `task` -- the deliverable is a per-row prediction file for a supplied test set, especially when the target is a skewed count/continuous quantity.
- **Pattern**: The agent writes the output file (or pastes it as the answer) without asserting that it contains exactly one prediction per test row in the expected id order/format, and without comparing the predicted value distribution to the training target distribution; the result is a truncated/partial file or predictions collapsed to a near-constant value (e.g., mostly zeros or the mode) that no sanity check would have accepted.
- **Detection procedure**:
  1. In the task, note the required output filename, required column name(s), and the number of rows in the provided test file.
  2. In the scripts, look for an explicit check after prediction: `len(pred) == len(test)`, ids matching test ids exactly (no drops from `dropna`/filtering/merging), and a written-file re-read verifying shape and header; absence of any such check is the first flag.
  3. In the scripts, look for a distribution sanity check on the predictions (min/max/mean/median/quantiles or share of a single value) compared against the same statistics of the training target; absence, or presence with a wildly mismatched profile that was not investigated, is the second flag.
  4. Inspect the answer itself: count rows, check for truncation/incomplete final line, and compute the fraction of identical values — a large majority of one value or far-too-small maximum relative to the training target range confirms the violation.
- **Discriminator**: A genuinely fine attempt may still predict many low/zero values if the training target is truly dominated by them **and** the script demonstrates row-count/id alignment plus a printed comparison of predicted vs. actual target distributions; the violation is the *absence* of these verifications combined with an output whose row count or value spread does not match the test set and target.
- **Consequence**: The grader finds the expected file missing rows, mis-formatted, or scored against a degenerate constant-like prediction, yielding a failed correctness check (or a metric far worse than a trivial baseline).
529Hypothesis test run on the full raw dataset with an unjustified default test specificationtaskda-code
Applies when
task -- the task asks for a p-value and a reject/fail-to-reject decision, and the scripts feed entire loaded tables into an off-the-shelf test function.
Pattern
The script takes the first plausible test (e.g., ttest_ind with default two-sided, equal-variance settings) applied to every row of both files, without checking whether the task/README scopes the comparison to a subset (a competition, date range, category, or matched population), whether the hypothesis is directional, or whether the data satisfy the test's assumptions (normality, symmetry, count/skewed data, unequal variances, paired vs independent). No alternative specification is compared, and an astronomically small p-value is accepted without question.
Detection procedure
  1. Read the task statement and README for any scoping or directional language (which subpopulation, which time window, "greater than", "more goals/higher value than") and list every filter/direction the intended analysis implies.
  2. In the scripts, check whether each such filter is actually applied before the test and whether the alternative, variance/pairing, and parametric-vs-rank choices are set deliberately rather than left at defaults.
  3. Inspect the printed sample sizes/means against the raw row counts: if n equals the full table length and no subsetting appears, the test population is unverified.
  4. Look at the reported p-value magnitude: an extreme value (e.g., <1e-50) driven by huge n on skewed count data is a red flag that neither the subset nor the test assumptions were validated.
Discriminator
A real violation is when the task/README implies a narrower population or a one-sided/nonparametric formulation and the script silently uses all rows with default two-sided parametric settings. It is fine if the script explicitly documents that no scoping applies, checks distribution shape/variances, and shows the decision is stable across reasonable test choices.
Consequence
The p-value differs by many orders of magnitude from the reference (and the reject/fail-to-reject decision can flip), so the saved result file fails the value check even though the file format is correct.
id e3b196d71ded · mined from da-code dacode-data-sa-001@s10
raw text (what the judge reads)
### Hypothesis test run on the full raw dataset with an unjustified default test specification
- **Applies when**: `task` -- the task asks for a p-value and a reject/fail-to-reject decision, and the scripts feed entire loaded tables into an off-the-shelf test function.
- **Pattern**: The script takes the first plausible test (e.g., `ttest_ind` with default two-sided, equal-variance settings) applied to every row of both files, without checking whether the task/README scopes the comparison to a subset (a competition, date range, category, or matched population), whether the hypothesis is directional, or whether the data satisfy the test's assumptions (normality, symmetry, count/skewed data, unequal variances, paired vs independent). No alternative specification is compared, and an astronomically small p-value is accepted without question.
- **Detection procedure**:
  1. Read the task statement and README for any scoping or directional language (which subpopulation, which time window, "greater than", "more goals/higher value than") and list every filter/direction the intended analysis implies.
  2. In the scripts, check whether each such filter is actually applied before the test and whether the `alternative`, variance/pairing, and parametric-vs-rank choices are set deliberately rather than left at defaults.
  3. Inspect the printed sample sizes/means against the raw row counts: if n equals the full table length and no subsetting appears, the test population is unverified.
  4. Look at the reported p-value magnitude: an extreme value (e.g., <1e-50) driven by huge n on skewed count data is a red flag that neither the subset nor the test assumptions were validated.
- **Discriminator**: A real violation is when the task/README implies a narrower population or a one-sided/nonparametric formulation and the script silently uses all rows with default two-sided parametric settings. It is fine if the script explicitly documents that no scoping applies, checks distribution shape/variances, and shows the decision is stable across reasonable test choices.
- **Consequence**: The p-value differs by many orders of magnitude from the reference (and the reject/fail-to-reject decision can flip), so the saved result file fails the value check even though the file format is correct.
530Never loading the provided template/sample output file before formatting resultstaskda-code
Applies when
task -- The task says the output must follow the exact structure/formatting of a provided sample or reference file (column names, order, row order, decimals, headers).
Pattern
The scripts read only the raw input tables, then invent the output schema from assumption (guessed column names, guessed column ordering, guessed sort order, guessed rounding/precision), and write the result file without ever opening, printing, or diffing against the supplied sample file.
Detection procedure
  1. Read the task statement and note any file cited as the formatting reference, plus any explicit formatting constraints (rounding, units, ordering).
  2. Grep the scripts for a read/open of that reference file; check whether its header, row order and value formatting are printed or programmatically compared to the produced output.
  3. If absent, inspect the hard-coded schema in the script (column list, sort key, rounding, index=False, float_format) and ask which of those choices are unverifiable guesses.
  4. Check the final answer's header/order against any snippet of the sample visible in the task or workspace; flag if it can't be confirmed.
Discriminator
Fine if the script loads the sample (or an equivalent listing of it) and asserts/aligns column names, order, dtypes and row ordering to it — even if the schema was later hard-coded; a violation is when the sample is never touched and the layout (e.g., placing a label column last, alphabetical sorting, 2-decimal formatting) rests purely on inference.
Consequence
Values may be computed correctly but the file fails exact-match grading on header names/order, row order, or numeric formatting, scoring 0 on the output-file check.
id 07ec6f7e05da · mined from da-code dacode-dm-csv-011@s10
raw text (what the judge reads)
### Never loading the provided template/sample output file before formatting results
- **Applies when**: `task` -- The task says the output must follow the exact structure/formatting of a provided sample or reference file (column names, order, row order, decimals, headers).
- **Pattern**: The scripts read only the raw input tables, then invent the output schema from assumption (guessed column names, guessed column ordering, guessed sort order, guessed rounding/precision), and write the result file without ever opening, printing, or diffing against the supplied sample file.
- **Detection procedure**:
  1. Read the task statement and note any file cited as the formatting reference, plus any explicit formatting constraints (rounding, units, ordering).
  2. Grep the scripts for a read/open of that reference file; check whether its header, row order and value formatting are printed or programmatically compared to the produced output.
  3. If absent, inspect the hard-coded schema in the script (column list, sort key, rounding, index=False, float_format) and ask which of those choices are unverifiable guesses.
  4. Check the final answer's header/order against any snippet of the sample visible in the task or workspace; flag if it can't be confirmed.
- **Discriminator**: Fine if the script loads the sample (or an equivalent listing of it) and asserts/aligns column names, order, dtypes and row ordering to it — even if the schema was later hard-coded; a violation is when the sample is never touched and the layout (e.g., placing a label column last, alphabetical sorting, 2-decimal formatting) rests purely on inference.
- **Consequence**: Values may be computed correctly but the file fails exact-match grading on header names/order, row order, or numeric formatting, scoring 0 on the output-file check.
531Silent row/feature loss from naive dtype selection and row-droppingtaskda-code
Applies when
task -- the task asks for a per-record output (labels, predictions, scores) over an entire supplied dataset, and the script builds its feature matrix by auto-selecting dtypes and/or dropping rows/columns with missing values.
Pattern
The script calls something like select_dtypes(number) plus dropna() without inspecting why columns are non-numeric (values carrying %, $, thousands separators, or other formatting are stored as strings) and without checking how many records survive. The result is a labeled output covering only a subset of the input rows and a small, arbitrary subset of the intended features — often including meaningless identifier-like numeric columns — while the answer still claims the task was completed on the dataset.
Detection procedure
  1. Read the task/README to determine the intended feature set and whether the output must contain one row per input record.
  2. In the script, find every step that filters columns (dtype selection) or rows (dropna, boolean masks) and check whether any cleaning/parsing of string-formatted numerics or imputation was attempted before them.
  3. Compare the reported output shape against the input row count and the README's list of quantitative fields; flag if rows shrank or if many documented numeric attributes are absent.
  4. Check whether retained features are actually informative (identifiers, codes, coordinates) rather than the substantive indicators described.
Discriminator
A genuine violation drops records or documented numeric attributes purely as an artifact of unparsed formatting or unhandled NaNs, with no justification and no coverage check. It is fine if the task explicitly permits subsetting, or if the script parses/imputes first and then documents a small, deliberate exclusion with row counts verified against the input.
Consequence
The saved file has fewer rows than expected and feature columns that don't match the reference feature vector, so a file-level comparison of shape/columns/labels fails outright even though the clustering code itself ran without error.
id ddf8824ab28b · mined from da-code dacode-ml-cluster-009@s10
raw text (what the judge reads)
### Silent row/feature loss from naive dtype selection and row-dropping
- **Applies when**: `task` -- the task asks for a per-record output (labels, predictions, scores) over an entire supplied dataset, and the script builds its feature matrix by auto-selecting dtypes and/or dropping rows/columns with missing values.
- **Pattern**: The script calls something like `select_dtypes(number)` plus `dropna()` without inspecting why columns are non-numeric (values carrying `%`, `$`, thousands separators, or other formatting are stored as strings) and without checking how many records survive. The result is a labeled output covering only a subset of the input rows and a small, arbitrary subset of the intended features — often including meaningless identifier-like numeric columns — while the answer still claims the task was completed on the dataset.
- **Detection procedure**:
  1. Read the task/README to determine the intended feature set and whether the output must contain one row per input record.
  2. In the script, find every step that filters columns (dtype selection) or rows (`dropna`, boolean masks) and check whether any cleaning/parsing of string-formatted numerics or imputation was attempted before them.
  3. Compare the reported output shape against the input row count and the README's list of quantitative fields; flag if rows shrank or if many documented numeric attributes are absent.
  4. Check whether retained features are actually informative (identifiers, codes, coordinates) rather than the substantive indicators described.
- **Discriminator**: A genuine violation drops records or documented numeric attributes purely as an artifact of unparsed formatting or unhandled NaNs, with no justification and no coverage check. It is fine if the task explicitly permits subsetting, or if the script parses/imputes first and then documents a small, deliberate exclusion with row counts verified against the input.
- **Consequence**: The saved file has fewer rows than expected and feature columns that don't match the reference feature vector, so a file-level comparison of shape/columns/labels fails outright even though the clustering code itself ran without error.
532Degenerate clustering accepted because a validity score was maximized by outlier-only clusterstaskda-code
Applies when
task -- an unsupervised segmentation task where the script builds skewed, unbounded aggregate features, scales them linearly, and picks the cluster count by maximizing an internal score (silhouette/CH) or by a hard-coded k.
Pattern
The attempt feeds highly right-skewed, heavy-tailed, mutually redundant aggregates into a distance-based algorithm with no outlier treatment or skew correction (no log/rank transform, no winsorizing, no dimensionality reduction), then trusts an internal validity index. The index is maximized by a partition that isolates a handful of extreme points, leaving essentially all rows in one cluster, and the near-1.0 score plus the lopsided cluster sizes are reported as "excellent separation" instead of being treated as a red flag.
Detection procedure
  1. In the task, confirm the deliverable is a meaningful segmentation (multiple usable groups), not just any label column.
  2. In the scripts, check whether the engineered features are unbounded sums/maxima/standard deviations and whether any skew/outlier handling exists before scaling and clustering; check whether k is chosen purely by an internal index or fixed without justification.
  3. In the printed output/answer, read the per-cluster counts and the reported score: flag if one cluster holds an overwhelming majority (e.g. >90–95%) or any cluster has a trivially small number of members, especially alongside an implausibly high separation score.
  4. Verify the attempt performed any sanity check on the segmentation (cluster size balance, profile interpretation, stability) rather than only reporting metrics.
Discriminator
A genuine violation shows near-total mass in one cluster with the remainder being outlier singletons and no skew/outlier treatment; a look-alike that is fine is a moderately imbalanced but interpretable segmentation (each cluster holding a non-trivial share, distinct profiles) where skewed features were transformed or outliers explicitly handled and the imbalance is discussed.
Consequence
The saved label file is effectively a single-cluster/outlier-flag assignment, so any check on segmentation quality (cluster count, size distribution, silhouette on transformed features, comparison to a reference segmentation) fails and the result file is marked wrong.
id 617560efcaa6 · mined from da-code dacode-ml-cluster-016@s10
raw text (what the judge reads)
### Degenerate clustering accepted because a validity score was maximized by outlier-only clusters
- **Applies when**: `task` -- an unsupervised segmentation task where the script builds skewed, unbounded aggregate features, scales them linearly, and picks the cluster count by maximizing an internal score (silhouette/CH) or by a hard-coded k.
- **Pattern**: The attempt feeds highly right-skewed, heavy-tailed, mutually redundant aggregates into a distance-based algorithm with no outlier treatment or skew correction (no log/rank transform, no winsorizing, no dimensionality reduction), then trusts an internal validity index. The index is maximized by a partition that isolates a handful of extreme points, leaving essentially all rows in one cluster, and the near-1.0 score plus the lopsided cluster sizes are reported as "excellent separation" instead of being treated as a red flag.
- **Detection procedure**:
  1. In the task, confirm the deliverable is a meaningful segmentation (multiple usable groups), not just any label column.
  2. In the scripts, check whether the engineered features are unbounded sums/maxima/standard deviations and whether any skew/outlier handling exists before scaling and clustering; check whether k is chosen purely by an internal index or fixed without justification.
  3. In the printed output/answer, read the per-cluster counts and the reported score: flag if one cluster holds an overwhelming majority (e.g. >90–95%) or any cluster has a trivially small number of members, especially alongside an implausibly high separation score.
  4. Verify the attempt performed any sanity check on the segmentation (cluster size balance, profile interpretation, stability) rather than only reporting metrics.
- **Discriminator**: A genuine violation shows near-total mass in one cluster with the remainder being outlier singletons and no skew/outlier treatment; a look-alike that is fine is a moderately imbalanced but interpretable segmentation (each cluster holding a non-trivial share, distinct profiles) where skewed features were transformed or outliers explicitly handled and the imbalance is discussed.
- **Consequence**: The saved label file is effectively a single-cluster/outlier-flag assignment, so any check on segmentation quality (cluster count, size distribution, silhouette on transformed features, comparison to a reference segmentation) fails and the result file is marked wrong.
533Unvalidated model: predictions reported with no held-out error estimate or row-alignment checktaskda-code
Applies when
task -- the task asks for predictions on a supplied test file written to an output file, and the script trains a model on a separate training set and writes predictions.
Pattern
The attempt trains a single model on all training data, immediately predicts on the test file, and reports only descriptive statistics of the predictions (mean/min/max/count) as evidence of success — never measuring error on a held-out or time-based validation split, and never verifying that the test feature matrix was built with the same columns, dtypes, encodings, imputation and row order as training (e.g. merges/joins/dropna that can silently reorder, duplicate, or drop rows, or produce all-NaN columns filled by a global mean).
Detection procedure
  1. Read the task: note the required output file, column name, and the implied one-prediction-per-test-row correspondence.
  2. In the scripts, search for any train/validation split and any error metric (MAE/RMSE/R²) computed on data not used for fitting; if absent, the model quality is unverified.
  3. Trace the test-side feature construction: confirm the same feature list, same encoders/imputers fitted on train are applied, and that any merge/aggregation is checked to preserve the test row count and original order before writing.
  4. Check the final answer/report for a validation score and for shape/order sanity checks; a report containing only prediction summary statistics is the tell.
Discriminator
A real violation is when no out-of-sample error is ever computed or test features are assembled by a path (different merge, refit imputer/encoder, reindexing) that is not demonstrably identical and order-preserving relative to training. It is fine if the script does cross-validation/holdout scoring and explicitly asserts len(pred) == len(test) with a preserved key/index, even if hyperparameters are simple and untuned.
Consequence
The output file has the right shape and plausible-looking value ranges but predictions are misaligned or driven by degenerate (mean-imputed / mismatched) features, so the grader's accuracy threshold against ground truth fails while the agent's self-report shows nothing wrong.
id 48bd3680d36a · mined from da-code dacode-ml-regression-002@s10
raw text (what the judge reads)
### Unvalidated model: predictions reported with no held-out error estimate or row-alignment check
- **Applies when**: `task` -- the task asks for predictions on a supplied test file written to an output file, and the script trains a model on a separate training set and writes predictions.
- **Pattern**: The attempt trains a single model on all training data, immediately predicts on the test file, and reports only descriptive statistics of the predictions (mean/min/max/count) as evidence of success — never measuring error on a held-out or time-based validation split, and never verifying that the test feature matrix was built with the same columns, dtypes, encodings, imputation and row order as training (e.g. merges/joins/dropna that can silently reorder, duplicate, or drop rows, or produce all-NaN columns filled by a global mean).
- **Detection procedure**:
  1. Read the task: note the required output file, column name, and the implied one-prediction-per-test-row correspondence.
  2. In the scripts, search for any train/validation split and any error metric (MAE/RMSE/R²) computed on data not used for fitting; if absent, the model quality is unverified.
  3. Trace the test-side feature construction: confirm the same feature list, same encoders/imputers fitted on train are applied, and that any merge/aggregation is checked to preserve the test row count and original order before writing.
  4. Check the final answer/report for a validation score and for shape/order sanity checks; a report containing only prediction summary statistics is the tell.
- **Discriminator**: A real violation is when *no* out-of-sample error is ever computed **or** test features are assembled by a path (different merge, refit imputer/encoder, reindexing) that is not demonstrably identical and order-preserving relative to training. It is fine if the script does cross-validation/holdout scoring and explicitly asserts `len(pred) == len(test)` with a preserved key/index, even if hyperparameters are simple and untuned.
- **Consequence**: The output file has the right shape and plausible-looking value ranges but predictions are misaligned or driven by degenerate (mean-imputed / mismatched) features, so the grader's accuracy threshold against ground truth fails while the agent's self-report shows nothing wrong.
534Fabricated/hard-coded input data instead of loading the provided datasettaskda-code
Applies when
task -- the task references a provided dataset (files in a data directory) and the scripts must read it to compute or plot the requested values.
Pattern
The agent never locates/reads the actual data file (or fails to find it and moves on), and instead hard-codes a small table of "approximate"/"typical" values from memory or assumption, then builds the deliverable from those invented numbers.
Detection procedure
  1. Read the task and note which dataset/entity the requested output must be derived from.
  2. Scan the scripts for an actual load step (read_csv/read_excel/load/open on a path under the provided data directory) that feeds the plotted/reported values; note whether any exploration script listed the directory and then was ignored.
  3. Check whether the plotted/reported series comes from a literal dict/list defined in the script rather than from a loaded DataFrame, and whether comments say "approximate", "typical", "based on the actual dataset".
  4. Check the answer for signs of unverifiable provenance: a suspiciously round/small number of points, values not traceable to any file, no printout of raw loaded data or shape.
Discriminator
Hard-coded constants are fine when they are configuration (figure size, labels, thresholds, seeds) or come from a config file; the violation is when the measured values themselves (the y-values, metrics, counts) are invented rather than derived from the supplied data.
Consequence
Saved artifacts (plot data arrays, JSON specs, npy values) will not match the ground-truth series — every value/length check fails even if the styling and file format are correct.
id 47542538328b · mined from da-code dacode-plot-line-015@s10
raw text (what the judge reads)
### Fabricated/hard-coded input data instead of loading the provided dataset
- **Applies when**: `task` -- the task references a provided dataset (files in a data directory) and the scripts must read it to compute or plot the requested values.
- **Pattern**: The agent never locates/reads the actual data file (or fails to find it and moves on), and instead hard-codes a small table of "approximate"/"typical" values from memory or assumption, then builds the deliverable from those invented numbers.
- **Detection procedure**:
  1. Read the task and note which dataset/entity the requested output must be derived from.
  2. Scan the scripts for an actual load step (`read_csv`/`read_excel`/`load`/open on a path under the provided data directory) that feeds the plotted/reported values; note whether any exploration script listed the directory and then was ignored.
  3. Check whether the plotted/reported series comes from a literal dict/list defined in the script rather than from a loaded DataFrame, and whether comments say "approximate", "typical", "based on the actual dataset".
  4. Check the answer for signs of unverifiable provenance: a suspiciously round/small number of points, values not traceable to any file, no printout of raw loaded data or shape.
- **Discriminator**: Hard-coded constants are fine when they are configuration (figure size, labels, thresholds, seeds) or come from a config file; the violation is when the *measured values themselves* (the y-values, metrics, counts) are invented rather than derived from the supplied data.
- **Consequence**: Saved artifacts (plot data arrays, JSON specs, npy values) will not match the ground-truth series — every value/length check fails even if the styling and file format are correct.
535Ignoring an external spec file that defines the required categorization/methodtaskda-code
Applies when
task -- the prompt points to a separate document (e.g., a markdown/README/config describing binning, grouping, filtering, or metric rules) that must govern how a derived variable or aggregation is built, and also names the exact output artifacts.
Pattern
The agent never opens or quotes the referenced spec and instead reuses the raw category labels already present in the data (or invents its own bins), then reports counts and claims success; it also produces only the one obvious output file, skipping the other artifacts the task/environment expects, and leaves no script behind for verification.
Detection procedure
  1. From the task text, list (a) every referenced instruction file and the rules it is supposed to contain, and (b) every required output artifact and its exact naming/labels.
  2. In the scripts, check for an explicit read of that instruction file and a mapping/binning step whose category boundaries and labels are traceable to it — not just value_counts() on the raw column.
  3. Compare the answer's category list against what the spec-derived grouping would produce (number of groups, boundary labels, ordering); a set of groups identical to the raw data's own labels is a red flag that the spec was ignored.
  4. Verify each required artifact is written and that a sanity check on totals/shape exists (e.g., group counts sum to the number of valid, non-missing respondents rather than the full raw row count).
Discriminator
A real violation is when no code path reads/encodes the spec's rules and the output groups coincide with the untransformed source categories or an arbitrary self-chosen scheme, or when named deliverables are absent. It is fine if the agent read the spec and the spec's grouping happens to coincide with the raw categories — provided the script shows the mapping explicitly and all named artifacts are produced.
Consequence
The grader compares against spec-defined groups and all expected files; wrong bin edges/labels and missing artifacts make every check fail even though the plot "looks" correct and the printed totals seem plausible.
id 2923d39b9511 · mined from da-code dacode-plot-bar-005@s10
raw text (what the judge reads)
### Ignoring an external spec file that defines the required categorization/method
- **Applies when**: `task` -- the prompt points to a separate document (e.g., a markdown/README/config describing binning, grouping, filtering, or metric rules) that must govern how a derived variable or aggregation is built, and also names the exact output artifacts.
- **Pattern**: The agent never opens or quotes the referenced spec and instead reuses the raw category labels already present in the data (or invents its own bins), then reports counts and claims success; it also produces only the one obvious output file, skipping the other artifacts the task/environment expects, and leaves no script behind for verification.
- **Detection procedure**:
  1. From the task text, list (a) every referenced instruction file and the rules it is supposed to contain, and (b) every required output artifact and its exact naming/labels.
  2. In the scripts, check for an explicit read of that instruction file and a mapping/binning step whose category boundaries and labels are traceable to it — not just `value_counts()` on the raw column.
  3. Compare the answer's category list against what the spec-derived grouping would produce (number of groups, boundary labels, ordering); a set of groups identical to the raw data's own labels is a red flag that the spec was ignored.
  4. Verify each required artifact is written and that a sanity check on totals/shape exists (e.g., group counts sum to the number of valid, non-missing respondents rather than the full raw row count).
- **Discriminator**: A real violation is when no code path reads/encodes the spec's rules and the output groups coincide with the untransformed source categories or an arbitrary self-chosen scheme, or when named deliverables are absent. It is fine if the agent read the spec and the spec's grouping happens to coincide with the raw categories — provided the script shows the mapping explicitly and all named artifacts are produced.
- **Consequence**: The grader compares against spec-defined groups and all expected files; wrong bin edges/labels and missing artifacts make every check fail even though the plot "looks" correct and the printed totals seem plausible.
536Missing required output artifact / declared answer schema not producedtaskda-code
Applies when
task -- the task specifies a concrete deliverable (a result file and/or an exact JSON/text schema, e.g. keys mapping to lists) in addition to a narrative answer.
Pattern
The scripts only load, clean, and print intermediate diagnostics and the final value to stdout; the agent then hand-types a reply that neither writes the requested file nor matches the requested container types (scalars where lists/arrays were specified, extra or renamed keys, missing rounding/units).
Detection procedure
  1. Read the task statement and list every hard output requirement: file name/path, key names, value types (list vs scalar), rounding, ordering.
  2. Grep the scripts for any write/serialize call (json.dump, to_json, to_csv, open(..., 'w')) and check the target path and structure match that list.
  3. Compare the agent's final answer text key-by-key and type-by-type against the requested schema.
  4. Flag if no script emits the artifact, or if any key/type/format element differs from the specification.
Discriminator
A real violation is the absence of the required artifact or a structural mismatch (scalar instead of list, wrong key, wrong path). Not a violation if the file is written elsewhere in the workflow (e.g. a later cell/script) with the exact schema, and the chat answer is merely a readable echo of correct content.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation is arguably right.
id ea470793ad20 · mined from da-code dacode-di-text-002@s10
raw text (what the judge reads)
### Missing required output artifact / declared answer schema not produced
- **Applies when**: `task` -- the task specifies a concrete deliverable (a result file and/or an exact JSON/text schema, e.g. keys mapping to lists) in addition to a narrative answer.
- **Pattern**: The scripts only load, clean, and `print` intermediate diagnostics and the final value to stdout; the agent then hand-types a reply that neither writes the requested file nor matches the requested container types (scalars where lists/arrays were specified, extra or renamed keys, missing rounding/units).
- **Detection procedure**:
  1. Read the task statement and list every hard output requirement: file name/path, key names, value types (list vs scalar), rounding, ordering.
  2. Grep the scripts for any write/serialize call (`json.dump`, `to_json`, `to_csv`, `open(..., 'w')`) and check the target path and structure match that list.
  3. Compare the agent's final answer text key-by-key and type-by-type against the requested schema.
  4. Flag if no script emits the artifact, or if any key/type/format element differs from the specification.
- **Discriminator**: A real violation is the absence of the required artifact or a structural mismatch (scalar instead of list, wrong key, wrong path). Not a violation if the file is written elsewhere in the workflow (e.g. a later cell/script) with the exact schema, and the chat answer is merely a readable echo of correct content.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation is arguably right.
537Ignoring the provided output template when producing the required result filetaskda-code
Applies when
task -- the task demands an output file "in the required format" and the workspace contains a sample/template/expected-format file (or the prompt spells out column names, ordering, units, or rounding).
Pattern
The agent invents its own schema (column names, extra/missing columns, date representation, decimal vs. percent, index column, row coverage) from the narrative description, and at most prints a comparison against the template without asserting equality or rewriting the file to match; the mismatch is never reconciled before submitting.
Detection procedure
  1. In the task statement and data directory listing, note every explicit format constraint and whether a sample output artifact exists.
  2. In the scripts, find where the result file is constructed and check whether column names/order, key/index column, value scale, rounding, and row count are derived from (or asserted against) the template rather than hard-coded from the agent's own reasoning.
  3. Check that any verification script does more than print: does it assert column-name equality, shape equality, and dtype/format equality, and does the pipeline fail or self-correct if they differ?
  4. Read the final answer for signs the agent silently accepted a divergence (e.g., it lists its own column names and row count without stating they match the template exactly).
Discriminator
A real violation is when the template's schema is available (or fully specified) and the written file's headers/shape/value convention are not provably identical to it. It is not a violation if the agent explicitly checked and enforced conformance (assertions, reindexing, renaming, rounding) — or if no template or format spec exists and the agent documented a reasonable schema consistent with every stated constraint.
Consequence
The grader compares the submitted file to the expected one field-by-field and reports it as WRONG/MISSING even when the underlying computation is nearly right, yielding 0 on the file check.
id 92c23dbd7933 · mined from da-code dacode-dm-csv-050@s10
raw text (what the judge reads)
### Ignoring the provided output template when producing the required result file
- **Applies when**: `task` -- the task demands an output file "in the required format" and the workspace contains a sample/template/expected-format file (or the prompt spells out column names, ordering, units, or rounding).
- **Pattern**: The agent invents its own schema (column names, extra/missing columns, date representation, decimal vs. percent, index column, row coverage) from the narrative description, and at most *prints* a comparison against the template without asserting equality or rewriting the file to match; the mismatch is never reconciled before submitting.
- **Detection procedure**:
  1. In the task statement and data directory listing, note every explicit format constraint and whether a sample output artifact exists.
  2. In the scripts, find where the result file is constructed and check whether column names/order, key/index column, value scale, rounding, and row count are derived from (or asserted against) the template rather than hard-coded from the agent's own reasoning.
  3. Check that any verification script does more than print: does it `assert` column-name equality, shape equality, and dtype/format equality, and does the pipeline fail or self-correct if they differ?
  4. Read the final answer for signs the agent silently accepted a divergence (e.g., it lists its own column names and row count without stating they match the template exactly).
- **Discriminator**: A real violation is when the template's schema is available (or fully specified) and the written file's headers/shape/value convention are not provably identical to it. It is *not* a violation if the agent explicitly checked and enforced conformance (assertions, reindexing, renaming, rounding) — or if no template or format spec exists and the agent documented a reasonable schema consistent with every stated constraint.
- **Consequence**: The grader compares the submitted file to the expected one field-by-field and reports it as WRONG/MISSING even when the underlying computation is nearly right, yielding 0 on the file check.
538Unvalidated input vector for a distribution/normality statistic (no cleaning, size, or consistency checks)taskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test and/or distribution moments computed on a single named variable, with a stated decision rule and reportable statistic.
Pattern
The attempt pipes the raw column straight into the test/moment functions without an auditable script that shows which rows survived — no explicit missing-value handling, no dtype coercion of string/sentinel-coded values, no check that the vector is the intended variable and full population — and then reports only the verdict, never printing the p-value, n, or the moments' consistency with that verdict. Sentinel codes (e.g., -999, 0 placeholders), stray NaNs, or a mis-selected/duplicated column silently create heavy tails, flipping the test outcome.
Detection procedure
  1. Read the task and note the exact variable, the test, the decision threshold, and every quantity that must be printed (here: the p-value plus the moments).
  2. In the scripts, locate the line building the analyzed array; check that it selects the named column, coerces dtype numerically, drops/handles NaNs explicitly, and prints n, min/max, and the count of dropped rows.
  3. Check that the script prints the raw test statistic/p-value (not just a yes/no), and that the reported verdict follows mechanically from comparing that p-value to the stated alpha.
  4. Sanity-cross-check the reported numbers against each other and against the variable's plausible range: extreme |skewness| / excess kurtosis alongside a claimed normal result (or vice versa) should trigger inspection of the input vector, e.g., by re-running on the cleaned column and on the raw column to see if the verdict changes.
Discriminator
A real violation is an attempt whose script never materializes or prints the cleaned vector's size/range and never emits the p-value, so the verdict cannot be traced to inputs; it is not a violation if the script shows the cleaning steps, prints n and the p-value, and the extreme moments are genuinely a property of the fully-cleaned data (in which case the verdict is defensible even if unusual).
Consequence
The test runs on a contaminated or wrong-sized vector, so the p-value and moments are off and the binary normality verdict is inverted — the grader marks the primary field wrong, and the missing printed p-value leaves no evidence to diagnose or defend the answer.
id d4acc011d262 · mined from infiagent-dabench dabench-298@s10
raw text (what the judge reads)
### Unvalidated input vector for a distribution/normality statistic (no cleaning, size, or consistency checks)
- **Applies when**: `task` -- the task asks for a hypothesis test and/or distribution moments computed on a single named variable, with a stated decision rule and reportable statistic.
- **Pattern**: The attempt pipes the raw column straight into the test/moment functions without an auditable script that shows which rows survived — no explicit missing-value handling, no dtype coercion of string/sentinel-coded values, no check that the vector is the intended variable and full population — and then reports only the verdict, never printing the p-value, n, or the moments' consistency with that verdict. Sentinel codes (e.g., -999, 0 placeholders), stray NaNs, or a mis-selected/duplicated column silently create heavy tails, flipping the test outcome.
- **Detection procedure**:
  1. Read the task and note the exact variable, the test, the decision threshold, and every quantity that must be printed (here: the p-value plus the moments).
  2. In the scripts, locate the line building the analyzed array; check that it selects the named column, coerces dtype numerically, drops/handles NaNs explicitly, and prints `n`, min/max, and the count of dropped rows.
  3. Check that the script prints the raw test statistic/p-value (not just a yes/no), and that the reported verdict follows mechanically from comparing that p-value to the stated alpha.
  4. Sanity-cross-check the reported numbers against each other and against the variable's plausible range: extreme |skewness| / excess kurtosis alongside a claimed normal result (or vice versa) should trigger inspection of the input vector, e.g., by re-running on the cleaned column and on the raw column to see if the verdict changes.
- **Discriminator**: A real violation is an attempt whose script never materializes or prints the cleaned vector's size/range and never emits the p-value, so the verdict cannot be traced to inputs; it is *not* a violation if the script shows the cleaning steps, prints n and the p-value, and the extreme moments are genuinely a property of the fully-cleaned data (in which case the verdict is defensible even if unusual).
- **Consequence**: The test runs on a contaminated or wrong-sized vector, so the p-value and moments are off and the binary normality verdict is inverted — the grader marks the primary field wrong, and the missing printed p-value leaves no evidence to diagnose or defend the answer.
539Model quality never validated on held-out data (only fit-set score reported), with prediction distribution left uncheckedtaskda-code
Applies when
task -- the script trains a supervised model on a labeled training file and writes predictions for an unlabeled test file that will be graded against hidden labels.
Pattern
The script fits one arbitrary model/hyperparameter configuration on 100% of the training rows, reports only the score on the same rows it was fit on (or no score at all), never holds out a validation split or does cross-validation, and never compares the predicted label distribution to the training label distribution — so a mediocre or systematically skewed model is shipped with no evidence it clears the grading threshold.
Detection procedure
  1. Read the task to confirm the output is graded on predictive quality against hidden labels (accuracy/score, not just file existence).
  2. Search the script for any train_test_split, cross_val_score, or explicit validation set used to score the model; if the only reported number comes from scoring on the same data used in fit, flag it.
  3. Check whether more than one configuration/model was compared, and whether any option that deliberately changes the decision prior (e.g. class re-weighting/resampling) is used without a validation comparison against the unweighted version.
  4. Compare the reported predicted-label distribution in the answer to the training-label distribution; a near-uniform or otherwise strongly shifted distribution with no justification is corroborating evidence of an unvalidated, mis-calibrated model.
Discriminator
A real violation is a single unverified configuration whose generalization score is unknown (fit-set score, or none) — even if the pipeline code is technically correct. It is not a violation if the agent held out data (or CV'd), reported an honest out-of-sample score, and chose among alternatives on that basis; nor is a shifted prediction distribution a problem when the agent validated that the shift improves the out-of-sample metric being graded.
Consequence
The submitted prediction file has accuracy well below the grader's threshold (fit-set score overstates generalization, and prior-distorting options push labels away from the true class balance), so the expected result file is scored WRONG despite the correct filename, column name and row count.
id 0ec75c3ac8b2 · mined from da-code dacode-ml-multi-011@s10
raw text (what the judge reads)
### Model quality never validated on held-out data (only fit-set score reported), with prediction distribution left unchecked
- **Applies when**: `task` -- the script trains a supervised model on a labeled training file and writes predictions for an unlabeled test file that will be graded against hidden labels.
- **Pattern**: The script fits one arbitrary model/hyperparameter configuration on 100% of the training rows, reports only the score on the same rows it was fit on (or no score at all), never holds out a validation split or does cross-validation, and never compares the predicted label distribution to the training label distribution — so a mediocre or systematically skewed model is shipped with no evidence it clears the grading threshold.
- **Detection procedure**:
  1. Read the task to confirm the output is graded on predictive quality against hidden labels (accuracy/score, not just file existence).
  2. Search the script for any `train_test_split`, `cross_val_score`, or explicit validation set used to score the model; if the only reported number comes from scoring on the same data used in `fit`, flag it.
  3. Check whether more than one configuration/model was compared, and whether any option that deliberately changes the decision prior (e.g. class re-weighting/resampling) is used without a validation comparison against the unweighted version.
  4. Compare the reported predicted-label distribution in the answer to the training-label distribution; a near-uniform or otherwise strongly shifted distribution with no justification is corroborating evidence of an unvalidated, mis-calibrated model.
- **Discriminator**: A real violation is a single unverified configuration whose generalization score is unknown (fit-set score, or none) — even if the pipeline code is technically correct. It is *not* a violation if the agent held out data (or CV'd), reported an honest out-of-sample score, and chose among alternatives on that basis; nor is a shifted prediction distribution a problem when the agent validated that the shift improves the out-of-sample metric being graded.
- **Consequence**: The submitted prediction file has accuracy well below the grader's threshold (fit-set score overstates generalization, and prior-distorting options push labels away from the true class balance), so the expected result file is scored WRONG despite the correct filename, column name and row count.
540Final submission produced by a degraded model that contradicts the attempt's own validation evidencetaskda-code
Applies when
task -- The task asks for predictions/results written to an output file, and the scripts train several model variants (or a "fast"/subsampled rerun) with any validation metric computed along the way.
Pattern
The agent runs a stronger pipeline first, then re-runs a cheaper variant (small random subsample of the training rows, shallow/few-estimator models, no tuning) that overwrites the same output file, and combines models with weights that ignore or invert the measured validation ranking (giving the majority of weight to the model with the clearly worse validation score). The reported answer describes these weak scores as "reasonable/excellent" instead of treating them as a red flag.
Detection procedure
  1. Read the task to confirm the output file is the graded artifact and note whether any accuracy/quality bar is implied by the competition setting.
  2. In the scripts, find every write to the output path and determine which one executes last; check what fraction of available training data and what model capacity that final pipeline uses.
  3. Compare the validation metrics printed for each candidate model against the ensemble weights / final choice: does the final artifact come from the best-scoring configuration, or from a weaker one?
  4. Check the answer text for validation numbers that are far apart across models, or for a metric that is poor in absolute terms yet labeled as acceptable without a comparison to a trivial baseline (mean predictor, plain linear fit).
Discriminator
A genuine violation is when the last-written artifact is demonstrably worse by the agent's own measured metric (weaker model dominant in the blend, or trained on a small slice of available data with no evidence it matches full-data performance). It is fine to deliberately subsample or ensemble when the scripts show validation evidence that the chosen configuration is at least as good, or when the drop is quantified and justified against a resource constraint stated in the task.
Consequence
The saved predictions are far less accurate than achievable (here the dominant component explained roughly half the variance of a simple linear alternative), so the graded file fails the competition's score threshold even though the file format, row count, and value range all look correct.
id 66aecc74c73a · mined from da-code dacode-ml-competition-008@s10
raw text (what the judge reads)
### Final submission produced by a degraded model that contradicts the attempt's own validation evidence
- **Applies when**: `task` -- The task asks for predictions/results written to an output file, and the scripts train several model variants (or a "fast"/subsampled rerun) with any validation metric computed along the way.
- **Pattern**: The agent runs a stronger pipeline first, then re-runs a cheaper variant (small random subsample of the training rows, shallow/few-estimator models, no tuning) that overwrites the same output file, and combines models with weights that ignore or invert the measured validation ranking (giving the majority of weight to the model with the clearly worse validation score). The reported answer describes these weak scores as "reasonable/excellent" instead of treating them as a red flag.
- **Detection procedure**:
  1. Read the task to confirm the output file is the graded artifact and note whether any accuracy/quality bar is implied by the competition setting.
  2. In the scripts, find every write to the output path and determine which one executes last; check what fraction of available training data and what model capacity that final pipeline uses.
  3. Compare the validation metrics printed for each candidate model against the ensemble weights / final choice: does the final artifact come from the best-scoring configuration, or from a weaker one?
  4. Check the answer text for validation numbers that are far apart across models, or for a metric that is poor in absolute terms yet labeled as acceptable without a comparison to a trivial baseline (mean predictor, plain linear fit).
- **Discriminator**: A genuine violation is when the last-written artifact is demonstrably worse by the agent's own measured metric (weaker model dominant in the blend, or trained on a small slice of available data with no evidence it matches full-data performance). It is fine to deliberately subsample or ensemble when the scripts show validation evidence that the chosen configuration is at least as good, or when the drop is quantified and justified against a resource constraint stated in the task.
- **Consequence**: The saved predictions are far less accurate than achievable (here the dominant component explained roughly half the variance of a simple linear alternative), so the graded file fails the competition's score threshold even though the file format, row count, and value range all look correct.
541Fabricating inputs and omitting required output artifactstaskda-code
Applies when
task -- the task points to provided input files (raw data, a config/spec file) and implies a set of deliverable artifacts (image plus machine-readable data/serialized results) in a fixed location.
Pattern
The agent cannot locate or does not load the supplied inputs, so it recreates them from memory or synthesizes a substitute (writing its own copy of the data or of the config), and it emits only the most visible artifact (e.g., the picture) while silently skipping the companion machine-readable outputs the checker reads.
Detection procedure
  1. From the task statement, list every input file the agent is supposed to consume and every output file it is supposed to produce (including implicit companions such as serialized values or a plot-data dump).
  2. In the scripts/answer, check whether each input is read from its given path — a line that writes or hardcodes the contents of a supposed input file is a red flag, as is a record count/schema that the agent states without a provenance step.
  3. Check that every output on the list is written, to the expected directory, with the expected name and extension; a summary that enumerates "files created" containing self-made inputs but missing a required output is a violation.
  4. Verify the plot parameters (title, labels, bins, colors, ordering) are parsed from the supplied spec file rather than invented, and that bin edges/categories follow the spec's definition rather than the agent's own choice.
Discriminator
Legitimate: the agent reads the given config/data, and additionally saves intermediate copies or diagnostics beyond the required set. Violation: the required input never appears in a read operation (or is reconstructed by the agent), or one or more required output paths are never written even though the prose claims the task is complete.
Consequence
The grader finds missing/mismatched artifacts (e.g., the serialized values and plot-metadata files absent, image built from invented data and bin edges), so every check fails regardless of how plausible the reported numbers look.
id 04502be907ba · mined from da-code dacode-plot-bar-007@s10
raw text (what the judge reads)
### Fabricating inputs and omitting required output artifacts
- **Applies when**: `task` -- the task points to provided input files (raw data, a config/spec file) and implies a set of deliverable artifacts (image plus machine-readable data/serialized results) in a fixed location.
- **Pattern**: The agent cannot locate or does not load the supplied inputs, so it recreates them from memory or synthesizes a substitute (writing its own copy of the data or of the config), and it emits only the most visible artifact (e.g., the picture) while silently skipping the companion machine-readable outputs the checker reads.
- **Detection procedure**:
  1. From the task statement, list every input file the agent is supposed to consume and every output file it is supposed to produce (including implicit companions such as serialized values or a plot-data dump).
  2. In the scripts/answer, check whether each input is *read* from its given path — a line that *writes* or hardcodes the contents of a supposed input file is a red flag, as is a record count/schema that the agent states without a provenance step.
  3. Check that every output on the list is written, to the expected directory, with the expected name and extension; a summary that enumerates "files created" containing self-made inputs but missing a required output is a violation.
  4. Verify the plot parameters (title, labels, bins, colors, ordering) are parsed from the supplied spec file rather than invented, and that bin edges/categories follow the spec's definition rather than the agent's own choice.
- **Discriminator**: Legitimate: the agent reads the given config/data, and additionally saves intermediate copies or diagnostics beyond the required set. Violation: the required input never appears in a read operation (or is reconstructed by the agent), or one or more required output paths are never written even though the prose claims the task is complete.
- **Consequence**: The grader finds missing/mismatched artifacts (e.g., the serialized values and plot-metadata files absent, image built from invented data and bin edges), so every check fails regardless of how plausible the reported numbers look.
542Undefined/empty-result branch emits a hand-written placeholder instead of the canonically formatted valuetaskinfiagent-dabench
Applies when
task -- The task demands a strictly formatted numeric answer and the scripts contain a conditional branch that substitutes a literal string (e.g. 'NaN', 'None', 'N/A', '') when a filter, group, or computation yields no rows or an undefined result.
Pattern
The agent correctly discovers the empty/undefined case but then writes the answer via a custom string in the else branch, bypassing the same formatting pipeline (float conversion, round(x, 2), lowercase/repr of the sentinel) that the expected answer token is built from, so the submitted string differs in case, spelling, or precision from the canonical form.
Detection procedure
  1. Read the task's answer-format spec and note the exact type/precision/token style required.
  2. In the answer-producing script, locate every branch that can set the reported value and check whether any of them assigns a hard-coded string or a differently typed placeholder rather than passing the computed value through the same format/round step.
  3. Compare the literal that branch produces to the canonical rendering of the underlying value (e.g. what f"{float('nan'):.2f}" or str(np.nan) yields) and to the case/spelling used in the spec.
  4. Check the final answer string: if it contains a placeholder token whose case/spelling was chosen by the agent rather than derived from the computed value, flag it.
Discriminator
A real violation is when the placeholder is authored by the agent and is not byte-identical to the canonical rendering of the value (NaN vs nan, 0 vs 0.00, None vs nan); it is fine if the empty case is still routed through the standard formatting call, or if the task explicitly names the exact token to emit for undefined results.
Consequence
The grader does an exact/normalized token match against the canonical value and reports the submission as WRONG/MISSING even though the analysis logic was right, scoring 0.
id c46b4edd6e5b · mined from infiagent-dabench dabench-554@s10
raw text (what the judge reads)
### Undefined/empty-result branch emits a hand-written placeholder instead of the canonically formatted value
- **Applies when**: `task` -- The task demands a strictly formatted numeric answer and the scripts contain a conditional branch that substitutes a literal string (e.g. `'NaN'`, `'None'`, `'N/A'`, `''`) when a filter, group, or computation yields no rows or an undefined result.
- **Pattern**: The agent correctly discovers the empty/undefined case but then writes the answer via a custom string in the `else` branch, bypassing the same formatting pipeline (float conversion, `round(x, 2)`, lowercase/`repr` of the sentinel) that the expected answer token is built from, so the submitted string differs in case, spelling, or precision from the canonical form.
- **Detection procedure**:
  1. Read the task's answer-format spec and note the exact type/precision/token style required.
  2. In the answer-producing script, locate every branch that can set the reported value and check whether any of them assigns a hard-coded string or a differently typed placeholder rather than passing the computed value through the same format/round step.
  3. Compare the literal that branch produces to the canonical rendering of the underlying value (e.g. what `f"{float('nan'):.2f}"` or `str(np.nan)` yields) and to the case/spelling used in the spec.
  4. Check the final answer string: if it contains a placeholder token whose case/spelling was chosen by the agent rather than derived from the computed value, flag it.
- **Discriminator**: A real violation is when the placeholder is authored by the agent and is not byte-identical to the canonical rendering of the value (`NaN` vs `nan`, `0` vs `0.00`, `None` vs `nan`); it is fine if the empty case is still routed through the standard formatting call, or if the task explicitly names the exact token to emit for undefined results.
- **Consequence**: The grader does an exact/normalized token match against the canonical value and reports the submission as WRONG/MISSING even though the analysis logic was right, scoring 0.
543Substituting proxy data/metrics when the required entities aren't found, instead of locating the real inputstaskda-code
Applies when
task -- the task names specific entities (groups, stages, metrics, config settings) and the agent claims the available data doesn't contain them and therefore "adapts" the task.
Pattern
Rather than searching the provided workspace for the correct file(s)/columns (or reconciling a possibly stale README with the actual files), the agent redefines the requested quantity with an unrelated proxy (e.g., counts instead of the stated aggregate, a different grouping key, category labels borrowed from a config file) and reports success on the substituted problem — often also skipping some required output artifacts.
Detection procedure
  1. From the task text, list the required inputs (entities/columns/keys), the required computation, and every required output artifact and its format/config source.
  2. In the scripts, check whether the agent enumerated all available data files/columns and mapped each required entity to a concrete real column; look for language like "since the data doesn't contain X, I used Y instead".
  3. Verify the computed statistic matches the stated definition (same aggregation, same grouping, same ordering/ranking key) rather than a convenient stand-in, and that the config settings are applied to the intended quantities, not merely borrowed as labels.
  4. Confirm the answer includes every requested artifact with expected shapes/values; missing or placeholder artifacts are disqualifying.
Discriminator
A real violation redefines the requested quantity or grouping (or omits artifacts) without exhaustively demonstrating the required fields are absent from all provided files; a look-alike is fine when the agent shows the mapping from renamed/undocumented columns to the required entities and computes the originally stated statistic on them.
Consequence
All expected output files fail comparison (wrong values, wrong categories/order, or missing files), yielding 0 checks passed despite a confident completion summary.
id 14a82d501dcb · mined from da-code dacode-plot-scatter-002@s10
raw text (what the judge reads)
### Substituting proxy data/metrics when the required entities aren't found, instead of locating the real inputs
- **Applies when**: `task` -- the task names specific entities (groups, stages, metrics, config settings) and the agent claims the available data doesn't contain them and therefore "adapts" the task.
- **Pattern**: Rather than searching the provided workspace for the correct file(s)/columns (or reconciling a possibly stale README with the actual files), the agent redefines the requested quantity with an unrelated proxy (e.g., counts instead of the stated aggregate, a different grouping key, category labels borrowed from a config file) and reports success on the substituted problem — often also skipping some required output artifacts.
- **Detection procedure**:
  1. From the task text, list the required inputs (entities/columns/keys), the required computation, and every required output artifact and its format/config source.
  2. In the scripts, check whether the agent enumerated all available data files/columns and mapped each required entity to a concrete real column; look for language like "since the data doesn't contain X, I used Y instead".
  3. Verify the computed statistic matches the stated definition (same aggregation, same grouping, same ordering/ranking key) rather than a convenient stand-in, and that the config settings are applied to the intended quantities, not merely borrowed as labels.
  4. Confirm the answer includes every requested artifact with expected shapes/values; missing or placeholder artifacts are disqualifying.
- **Discriminator**: A real violation redefines the requested quantity or grouping (or omits artifacts) without exhaustively demonstrating the required fields are absent from all provided files; a look-alike is fine when the agent shows the mapping from renamed/undocumented columns to the required entities and computes the originally stated statistic on them.
- **Consequence**: All expected output files fail comparison (wrong values, wrong categories/order, or missing files), yielding 0 checks passed despite a confident completion summary.
544Output does not conform to the provided template/schema of the required result filetaskda-code
Applies when
task -- the task says to save results to a named file "matching the provided template" (or otherwise specifies an output schema), and the scripts build the table themselves instead of reconciling it with that template.
Pattern
The agent computes plausible numbers but writes/reports them with its own conventions — self-invented column names or index label, raw unrounded floats, wrong column count/order, wrong row ordering, empty-vs-NaN cells, or a truncated/pasted-in-chat table — never opening the template file to compare headers, dtypes, precision, and shape.
Detection procedure
  1. Read the task for the exact output filename and any referenced template/example file, and note every format constraint it implies (header names, index column, number of columns, rounding/units, ordering).
  2. In the scripts, look for code that reads the template (or an explicit schema definition) and enforces it — e.g., reindexing to the template's columns/rows, applying the template's rounding, using its header/index labels — before to_csv.
  3. Compare the agent's produced table head to the template head cell-by-cell: same first-column name and value formatting, same column labels and count, same decimal precision, same handling of missing cells.
  4. Check the file was actually written to the required path and is complete (not a truncated console dump), and that its shape matches the expected number of cohorts/periods.
Discriminator
A real violation is a structural or precision mismatch with the template (different labels, extra/missing columns, unrounded values, wrong ordering, missing file); it is not a violation if the values differ only by benign float noise while headers, shape, ordering and rounding all match the template.
Consequence
The grader's file-level comparison against the expected artifact fails outright ("WRONG/MISSING") even if the underlying computation was conceptually right, scoring 0.
id 134ccd107623 · mined from da-code dacode-dm-csv-043@s10
raw text (what the judge reads)
### Output does not conform to the provided template/schema of the required result file
- **Applies when**: `task` -- the task says to save results to a named file "matching the provided template" (or otherwise specifies an output schema), and the scripts build the table themselves instead of reconciling it with that template.
- **Pattern**: The agent computes plausible numbers but writes/reports them with its own conventions — self-invented column names or index label, raw unrounded floats, wrong column count/order, wrong row ordering, empty-vs-NaN cells, or a truncated/pasted-in-chat table — never opening the template file to compare headers, dtypes, precision, and shape.
- **Detection procedure**:
  1. Read the task for the exact output filename and any referenced template/example file, and note every format constraint it implies (header names, index column, number of columns, rounding/units, ordering).
  2. In the scripts, look for code that reads the template (or an explicit schema definition) and enforces it — e.g., reindexing to the template's columns/rows, applying the template's rounding, using its header/index labels — before `to_csv`.
  3. Compare the agent's produced table head to the template head cell-by-cell: same first-column name and value formatting, same column labels and count, same decimal precision, same handling of missing cells.
  4. Check the file was actually written to the required path and is complete (not a truncated console dump), and that its shape matches the expected number of cohorts/periods.
- **Discriminator**: A real violation is a structural or precision mismatch with the template (different labels, extra/missing columns, unrounded values, wrong ordering, missing file); it is *not* a violation if the values differ only by benign float noise while headers, shape, ordering and rounding all match the template.
- **Consequence**: The grader's file-level comparison against the expected artifact fails outright ("WRONG/MISSING") even if the underlying computation was conceptually right, scoring 0.
545Discarding most available predictors and shipping a model that never beats a trivial baselinetaskda-code
Applies when
task -- a predictive model must be fit on a tabular file that contains categorical/identifier/date/text columns alongside a few generic numeric ones, and the scripts hard-code a small numeric-only feature list.
Pattern
The attempt silently drops every non-numeric or high-cardinality column (entity names, dates/eras, category or tag fields) instead of encoding/aggregating them, trains on the weak leftovers, and accepts the result without comparing validation error to a trivial baseline (predict the global mean/median) — producing predictions whose spread is far narrower than the target's and whose R² is near zero.
Detection procedure
  1. From the task/README and the exploration script output, list all columns present besides the target, and note which carry obvious signal (identity of the producing entity, time of release, genre/category labels, counts/rankings).
  2. In the modeling script, compare that list to the feature list actually used; flag columns excluded with no stated reason and no attempt at encoding, target/frequency aggregation, or date parsing.
  3. Check whether any baseline comparison exists (constant-prediction error, or a stated acceptance threshold) and whether the reported validation R²/MAE is reported at all and is meaningfully better than that baseline.
  4. Compare the reported prediction std/range in the answer to the target's std/range from the exploration step; a large shrinkage with no baseline check confirms the violation.
Discriminator
Fine if the excluded columns are genuinely uninformative (unique IDs, free-text URLs, constant columns) or the attempt shows a validation comparison demonstrating they add nothing; a violation when informative metadata is dropped purely because it is non-numeric and no baseline/validation comparison is reported.
Consequence
The saved predictions are close to a constant near the training mean, so any accuracy/correlation-based grading of the prediction file fails even though the file has the right shape and column name.
id 4eec16430bb2 · mined from da-code dacode-ml-regression-004@s10
raw text (what the judge reads)
### Discarding most available predictors and shipping a model that never beats a trivial baseline
- **Applies when**: `task` -- a predictive model must be fit on a tabular file that contains categorical/identifier/date/text columns alongside a few generic numeric ones, and the scripts hard-code a small numeric-only feature list.
- **Pattern**: The attempt silently drops every non-numeric or high-cardinality column (entity names, dates/eras, category or tag fields) instead of encoding/aggregating them, trains on the weak leftovers, and accepts the result without comparing validation error to a trivial baseline (predict the global mean/median) — producing predictions whose spread is far narrower than the target's and whose R² is near zero.
- **Detection procedure**:
  1. From the task/README and the exploration script output, list all columns present besides the target, and note which carry obvious signal (identity of the producing entity, time of release, genre/category labels, counts/rankings).
  2. In the modeling script, compare that list to the feature list actually used; flag columns excluded with no stated reason and no attempt at encoding, target/frequency aggregation, or date parsing.
  3. Check whether any baseline comparison exists (constant-prediction error, or a stated acceptance threshold) and whether the reported validation R²/MAE is reported at all and is meaningfully better than that baseline.
  4. Compare the reported prediction std/range in the answer to the target's std/range from the exploration step; a large shrinkage with no baseline check confirms the violation.
- **Discriminator**: Fine if the excluded columns are genuinely uninformative (unique IDs, free-text URLs, constant columns) *or* the attempt shows a validation comparison demonstrating they add nothing; a violation when informative metadata is dropped purely because it is non-numeric and no baseline/validation comparison is reported.
- **Consequence**: The saved predictions are close to a constant near the training mean, so any accuracy/correlation-based grading of the prediction file fails even though the file has the right shape and column name.
546Unsanity-checked, degenerate cluster/group solution accepted as "optimal"taskda-code
Applies when
task -- the task asks for an unsupervised grouping (or any model selection by an internal score) and the script picks the configuration that maximizes a single criterion without inspecting the resulting group structure.
Pattern
The attempt sweeps a hyperparameter (e.g., number of groups), takes the argmax of one internal score, and reports the result even though the produced partition is degenerate — singleton or near-singleton groups, one group holding most of the data, and a low absolute score — which is usually a symptom of skipped preprocessing on heavy-tailed/outlier-laden features (no log/robust transform, no outlier handling, no missing-value check) rather than real structure.
Detection procedure
  1. Read the task for the required output contract (row count, exact column names/order, one label per input record) and any implied need for interpretable groups.
  2. In the script, check whether the selection step considers anything beyond a single score: absolute score magnitude, group-size distribution, stability across seeds, alternative scalings/transforms, alternative algorithms.
  3. In the reported output, inspect the group-size table and score value: flag if any group has a trivially small membership (e.g., 1–3 of hundreds), if one group dominates, or if the score is near the "no structure" range for that metric.
  4. Verify the written file itself matches the contract (header names exactly as specified, correct number of feature columns/rows, labels present for every record) rather than trusting the prose description of the file.
Discriminator
A genuine violation is a partition whose sizes/score indicate the algorithm merely isolated outliers or split noise, with no attempt to transform skewed features or compare alternatives; a look-alike that is fine is a small group that is justified by explicit checks (stability, transformed features, domain interpretation) and a defensible score, with the output file conforming to the requested schema.
Consequence
The saved label column disagrees with any reasonable reference partition (and/or the header/shape mismatches), so the file-level check fails and the task scores 0 despite a confident-sounding summary.
id c0b2cf36c10e · mined from da-code dacode-ml-cluster-013@s10
raw text (what the judge reads)
### Unsanity-checked, degenerate cluster/group solution accepted as "optimal"
- **Applies when**: `task` -- the task asks for an unsupervised grouping (or any model selection by an internal score) and the script picks the configuration that maximizes a single criterion without inspecting the resulting group structure.
- **Pattern**: The attempt sweeps a hyperparameter (e.g., number of groups), takes the argmax of one internal score, and reports the result even though the produced partition is degenerate — singleton or near-singleton groups, one group holding most of the data, and a low absolute score — which is usually a symptom of skipped preprocessing on heavy-tailed/outlier-laden features (no log/robust transform, no outlier handling, no missing-value check) rather than real structure.
- **Detection procedure**:
  1. Read the task for the required output contract (row count, exact column names/order, one label per input record) and any implied need for interpretable groups.
  2. In the script, check whether the selection step considers anything beyond a single score: absolute score magnitude, group-size distribution, stability across seeds, alternative scalings/transforms, alternative algorithms.
  3. In the reported output, inspect the group-size table and score value: flag if any group has a trivially small membership (e.g., 1–3 of hundreds), if one group dominates, or if the score is near the "no structure" range for that metric.
  4. Verify the written file itself matches the contract (header names exactly as specified, correct number of feature columns/rows, labels present for every record) rather than trusting the prose description of the file.
- **Discriminator**: A genuine violation is a partition whose sizes/score indicate the algorithm merely isolated outliers or split noise, with no attempt to transform skewed features or compare alternatives; a look-alike that is fine is a small group that is justified by explicit checks (stability, transformed features, domain interpretation) and a defensible score, with the output file conforming to the requested schema.
- **Consequence**: The saved label column disagrees with any reasonable reference partition (and/or the header/shape mismatches), so the file-level check fails and the task scores 0 despite a confident-sounding summary.
547Silently redefining an ill-posed statistic (dropping the stated filter / using only part of the available data)taskinfiagent-dabench
Applies when
task -- the task names a specific slice (a year, group, subset) and a statistic that cannot be computed literally on that slice (e.g. a shape statistic on one value per entity), and/or the source data is split across several files or a wider population than the one the script loads.
Pattern
The script notices the literal reading is degenerate, then unilaterally substitutes a different computation (aggregating over a different axis, all periods, or one arbitrary file/region) and reports that result as the answer, without enumerating the plausible interpretations, without checking that all relevant data files/rows were included, and without any check that the chosen reading still honors the stated constraint.
Detection procedure
  1. Read the task and list the explicit scoping constraints (which subset, which axis/dimension, which definition/flag of the statistic) and the full set of data sources available.
  2. Read the script: identify which files/rows/columns are actually loaded and which axis the statistic is reduced over; check whether the named subset still appears in the final computation.
  3. If the literal computation was abandoned (comments like "this is ambiguous", "we can't compute skewness of one value"), check whether the script evaluates 2+ candidate interpretations and/or all data sources and gives a stated reason for the pick.
  4. Inspect the answer: is it the winner of an interpretation that quietly drops the constraint or uses only a fraction of the population?
Discriminator
A real violation is a single unjustified reinterpretation, or a computation restricted to one of several available data sources; it is fine if the script explicitly compares the candidate interpretations over the complete data and shows they agree, or documents why one reading is the only one consistent with every stated constraint.
Consequence
The reported entity is the extremum of a different statistic or of a subpopulation, so it does not match the expected label and the answer scores 0 despite clean-looking code.
id b70be6b4d791 · mined from infiagent-dabench dabench-252@s10
raw text (what the judge reads)
### Silently redefining an ill-posed statistic (dropping the stated filter / using only part of the available data)
- **Applies when**: `task` -- the task names a specific slice (a year, group, subset) and a statistic that cannot be computed literally on that slice (e.g. a shape statistic on one value per entity), and/or the source data is split across several files or a wider population than the one the script loads.
- **Pattern**: The script notices the literal reading is degenerate, then unilaterally substitutes a different computation (aggregating over a different axis, all periods, or one arbitrary file/region) and reports that result as the answer, without enumerating the plausible interpretations, without checking that all relevant data files/rows were included, and without any check that the chosen reading still honors the stated constraint.
- **Detection procedure**:
  1. Read the task and list the explicit scoping constraints (which subset, which axis/dimension, which definition/flag of the statistic) and the full set of data sources available.
  2. Read the script: identify which files/rows/columns are actually loaded and which axis the statistic is reduced over; check whether the named subset still appears in the final computation.
  3. If the literal computation was abandoned (comments like "this is ambiguous", "we can't compute skewness of one value"), check whether the script evaluates 2+ candidate interpretations and/or all data sources and gives a stated reason for the pick.
  4. Inspect the answer: is it the winner of an interpretation that quietly drops the constraint or uses only a fraction of the population?
- **Discriminator**: A real violation is a single unjustified reinterpretation, or a computation restricted to one of several available data sources; it is fine if the script explicitly compares the candidate interpretations over the complete data and shows they agree, or documents why one reading is the only one consistent with every stated constraint.
- **Consequence**: The reported entity is the extremum of a different statistic or of a subpopulation, so it does not match the expected label and the answer scores 0 despite clean-looking code.
548Ignoring the provided reference/sample output file when producing the deliverabletaskda-code
Applies when
task -- the task says the output must match the format of a provided example/schema file (or a stated column/ordering/rounding spec) and the scripts write a result file.
Pattern
The scripts never load or inspect the reference file; the agent guesses column names, header text, row ordering, rounding/precision, and value definitions from the prose alone, then saves the deliverable without comparing it to the template.
Detection procedure
  1. Read the task for any mention of a sample/expected-format file or explicit format constraints (column names, order, units, decimals).
  2. Search the scripts for a read of that file (or any explicit assertion on the output's columns, dtypes, row count, and ordering); note whether the output column names/rounding are hard-coded from guesswork.
  3. Compare the submitted file's header, row count, ordering, and numeric precision against the template's; check that the aggregation definition implied by the template (e.g., which grouping keys, whether empty periods count) is the one implemented.
  4. Flag if no read/verification step exists, or if any header/order/precision element is invented.
Discriminator
Fine if the script loads the template (or otherwise reproduces its exact header/order/precision) and the output demonstrably conforms; a violation is when conformity is merely assumed, even if the underlying numbers might be right — and especially when the template would also have disambiguated the aggregation's grouping/denominator.
Consequence
The grader's file comparison fails on header names, row order, or rounded values (and possibly on the statistic itself), scoring 0 despite plausible-looking numbers.
id 9a3fbcab6f76 · mined from da-code dacode-dm-csv-010@s10
raw text (what the judge reads)
### Ignoring the provided reference/sample output file when producing the deliverable
- **Applies when**: `task` -- the task says the output must match the format of a provided example/schema file (or a stated column/ordering/rounding spec) and the scripts write a result file.
- **Pattern**: The scripts never load or inspect the reference file; the agent guesses column names, header text, row ordering, rounding/precision, and value definitions from the prose alone, then saves the deliverable without comparing it to the template.
- **Detection procedure**:
  1. Read the task for any mention of a sample/expected-format file or explicit format constraints (column names, order, units, decimals).
  2. Search the scripts for a read of that file (or any explicit assertion on the output's columns, dtypes, row count, and ordering); note whether the output column names/rounding are hard-coded from guesswork.
  3. Compare the submitted file's header, row count, ordering, and numeric precision against the template's; check that the aggregation definition implied by the template (e.g., which grouping keys, whether empty periods count) is the one implemented.
  4. Flag if no read/verification step exists, or if any header/order/precision element is invented.
- **Discriminator**: Fine if the script loads the template (or otherwise reproduces its exact header/order/precision) and the output demonstrably conforms; a violation is when conformity is merely assumed, even if the underlying numbers might be right — and especially when the template would also have disambiguated the aggregation's grouping/denominator.
- **Consequence**: The grader's file comparison fails on header names, row order, or rounded values (and possibly on the statistic itself), scoring 0 despite plausible-looking numbers.
549Lossy reformatting of an identified key (dropping precision from the answer)taskinfiagent-dabench
Applies when
task -- the task asks to locate a specific record/entity (e.g., the argmax row identifier) and then use it in a follow-up computation, and the answer template shows a shortened or coarser rendering of that identifier than the granularity the data and the computation actually use.
Pattern
The script correctly finds the record at full resolution and correctly does the dependent calculation with it, but the final reported identifier is truncated/reformatted (coarser unit, dropped component, rounded label) so it no longer uniquely designates the record that was found; the agent treats an ambiguous format hint as license to discard information instead of reporting the identifier at the resolution at which it was determined.
Detection procedure
  1. In the task, note the granularity at which the target record must be located and at which the dependent quantity (difference, ratio, lookup of the neighbouring record) is computed.
  2. In the scripts, find the line that formats the identifier for output and compare it with the value used internally for the dependent computation.
  3. If the printed identifier is coarser than the internal one (fewer components, aggregated unit) and could match many records, flag it; check whether the full-resolution value could have been printed alongside or instead.
  4. Confirm the submitted answer string, not just the console log, carries the full-resolution identifier.
Discriminator
A real violation is when the coarse label is ambiguous with respect to the record actually used (multiple records share it) or contradicts the granularity of the neighbouring-record logic. It is fine when the task's grouping genuinely happens at the coarse level (e.g., the maximum is taken over aggregated groups) so the coarse label is the unit of analysis.
Consequence
The dependent numeric answer matches, but the identifier field is scored WRONG/MISSING, so the submission fails despite correct analysis.
id e231e4cff591 · mined from infiagent-dabench dabench-572@s10
raw text (what the judge reads)
### Lossy reformatting of an identified key (dropping precision from the answer)
- **Applies when**: `task` -- the task asks to locate a specific record/entity (e.g., the argmax row identifier) and then use it in a follow-up computation, and the answer template shows a shortened or coarser rendering of that identifier than the granularity the data and the computation actually use.
- **Pattern**: The script correctly finds the record at full resolution and correctly does the dependent calculation with it, but the final reported identifier is truncated/reformatted (coarser unit, dropped component, rounded label) so it no longer uniquely designates the record that was found; the agent treats an ambiguous format hint as license to discard information instead of reporting the identifier at the resolution at which it was determined.
- **Detection procedure**:
  1. In the task, note the granularity at which the target record must be located and at which the dependent quantity (difference, ratio, lookup of the neighbouring record) is computed.
  2. In the scripts, find the line that formats the identifier for output and compare it with the value used internally for the dependent computation.
  3. If the printed identifier is coarser than the internal one (fewer components, aggregated unit) and could match many records, flag it; check whether the full-resolution value could have been printed alongside or instead.
  4. Confirm the submitted answer string, not just the console log, carries the full-resolution identifier.
- **Discriminator**: A real violation is when the coarse label is ambiguous with respect to the record actually used (multiple records share it) or contradicts the granularity of the neighbouring-record logic. It is fine when the task's grouping genuinely happens at the coarse level (e.g., the maximum is taken over aggregated groups) so the coarse label *is* the unit of analysis.
- **Consequence**: The dependent numeric answer matches, but the identifier field is scored WRONG/MISSING, so the submission fails despite correct analysis.
550Optimizing/selecting by a ranking metric while emitting hard labels at the default threshold under class imbalancetaskda-code
Applies when
task -- a classification task with a rare positive class and an explicitly asymmetric cost (missed positives cost far more than false alarms), where the deliverable is a column of hard 0/1 labels.
Pattern
The script selects and compares models with a threshold-free score (e.g., ROC-AUC) or plain accuracy, then calls predict() (implicit 0.5 cut-off) to produce the submitted labels, never tuning the decision threshold or evaluating the cost-aligned metric (recall / F-beta / expected cost) on a validation set; the resulting label file contains far fewer positives than the base rate implies.
Detection procedure
  1. Read the task statement for the scoring objective and the cost asymmetry between error types; note the positive-class prevalence printed or implied by the training data.
  2. In the scripts, check which metric drives model selection and whether any threshold search / class_weight / resampling / predict_proba cut-off tuning is done before generating the final labels; flag if selection uses AUC-like scores but output uses bare predict().
  3. Check whether the cost-aligned metric (recall or F1 on the positive class) is ever computed on held-out data and reported.
  4. Count positives in the submitted file and compare to prevalence × n_rows; flag if it is substantially below that expectation.
Discriminator
A real violation is when no threshold/cost-sensitive step exists anywhere and positive counts are far under the expected prevalence. It is fine if the agent explicitly validated the threshold against the cost-aligned metric (or used balanced weights/resampling plus a tuned cut-off) and the positive rate is defensible, even if the count differs somewhat from the base rate.
Consequence
The label file is dominated by the majority class, recall on the rare class is low, and the grader's recall/F1/cost-based check against the reference labels fails despite a high reported AUC.
id eb380e6a7e72 · mined from da-code dacode-ml-binary-013@s10
raw text (what the judge reads)
### Optimizing/selecting by a ranking metric while emitting hard labels at the default threshold under class imbalance
- **Applies when**: `task` -- a classification task with a rare positive class and an explicitly asymmetric cost (missed positives cost far more than false alarms), where the deliverable is a column of hard 0/1 labels.
- **Pattern**: The script selects and compares models with a threshold-free score (e.g., ROC-AUC) or plain accuracy, then calls `predict()` (implicit 0.5 cut-off) to produce the submitted labels, never tuning the decision threshold or evaluating the cost-aligned metric (recall / F-beta / expected cost) on a validation set; the resulting label file contains far fewer positives than the base rate implies.
- **Detection procedure**:
  1. Read the task statement for the scoring objective and the cost asymmetry between error types; note the positive-class prevalence printed or implied by the training data.
  2. In the scripts, check which metric drives model selection and whether any threshold search / `class_weight` / resampling / `predict_proba` cut-off tuning is done before generating the final labels; flag if selection uses AUC-like scores but output uses bare `predict()`.
  3. Check whether the cost-aligned metric (recall or F1 on the positive class) is ever computed on held-out data and reported.
  4. Count positives in the submitted file and compare to `prevalence × n_rows`; flag if it is substantially below that expectation.
- **Discriminator**: A real violation is when no threshold/cost-sensitive step exists anywhere and positive counts are far under the expected prevalence. It is fine if the agent explicitly validated the threshold against the cost-aligned metric (or used balanced weights/resampling plus a tuned cut-off) and the positive rate is defensible, even if the count differs somewhat from the base rate.
- **Consequence**: The label file is dominated by the majority class, recall on the rare class is low, and the grader's recall/F1/cost-based check against the reference labels fails despite a high reported AUC.
551Answer tag built from a raw Python object repr instead of the exact requested formattaskinfiagent-dabench
Applies when
task -- the task specifies a literal answer tag/format (e.g. @name[list_of_strings]) and the script emits it by string-interpolating a computed object.
Pattern
The script does something like print(f"@key{my_list}"), letting Python's repr decide delimiters, quoting, spacing and ordering, rather than explicitly assembling the string in the format the task demands; the values may be right while the surrounding syntax (quote style, brackets, separators, type of elements) does not match the specification the grader parses.
Detection procedure
  1. Read the task's answer-format spec and note exactly how elements should be delimited, quoted, ordered, and typed.
  2. Find every line in the scripts that prints/writes the final answer tag and determine whether the string is constructed explicitly (e.g. ", ".join(str(x) for x in ...)) or via implicit repr/str of a list, Series, array, or numpy scalar.
  3. Compare the literal characters the script will emit (including quotes, np.float64(...), trailing spaces, index labels) to the spec; also check the elements are the requested entities (names) and not indices or intermediate objects.
  4. Check the submitted answer string itself against the spec character-by-character before accepting it as "content-correct".
Discriminator
A real violation is when the emitted string's syntax deviates from the specified template (extra quoting, container repr, dtype wrappers, wrong separator) even though the underlying values are correct; a look-alike that is fine is a script that interpolates an object whose str/repr provably equals the required template, or that builds the tag from explicitly formatted plain strings.
Consequence
The grader parses the tag and reports the expected values as WRONG/MISSING despite the analysis being substantively correct, yielding 0/1 checks passed.
id 1ee0a704b93a · mined from infiagent-dabench dabench-254@s10
raw text (what the judge reads)
### Answer tag built from a raw Python object repr instead of the exact requested format
- **Applies when**: `task` -- the task specifies a literal answer tag/format (e.g. `@name[list_of_strings]`) and the script emits it by string-interpolating a computed object.
- **Pattern**: The script does something like `print(f"@key{my_list}")`, letting Python's `repr` decide delimiters, quoting, spacing and ordering, rather than explicitly assembling the string in the format the task demands; the values may be right while the surrounding syntax (quote style, brackets, separators, type of elements) does not match the specification the grader parses.
- **Detection procedure**:
  1. Read the task's answer-format spec and note exactly how elements should be delimited, quoted, ordered, and typed.
  2. Find every line in the scripts that prints/writes the final answer tag and determine whether the string is constructed explicitly (e.g. `", ".join(str(x) for x in ...)`) or via implicit `repr`/`str` of a list, Series, array, or numpy scalar.
  3. Compare the literal characters the script will emit (including quotes, `np.float64(...)`, trailing spaces, index labels) to the spec; also check the elements are the requested entities (names) and not indices or intermediate objects.
  4. Check the submitted answer string itself against the spec character-by-character before accepting it as "content-correct".
- **Discriminator**: A real violation is when the emitted string's syntax deviates from the specified template (extra quoting, container repr, dtype wrappers, wrong separator) even though the underlying values are correct; a look-alike that is fine is a script that interpolates an object whose `str`/`repr` provably equals the required template, or that builds the tag from explicitly formatted plain strings.
- **Consequence**: The grader parses the tag and reports the expected values as WRONG/MISSING despite the analysis being substantively correct, yielding 0/1 checks passed.
552Degenerate/empty result emitted as a placeholder instead of a valid identifiertaskinfiagent-dabench
Applies when
task -- The task asks for an "argmax"-style identifier (which group/row/model has the highest count, score, etc.) plus its value, and the underlying quantity may legitimately be zero or tied for all groups.
Pattern
The script branches on "if the aggregate is empty/zero, skip the selection" and writes None/""/N/A into the identifier slot, so the answer satisfies the numeric field but leaves the required label unfilled, instead of still selecting a concrete group (e.g., first by the specified or a deterministic ordering) from the fully enumerated group list.
Detection procedure
1. In the task statement, confirm both an identifier and a numeric field are required, and that the numeric field's allowed range includes the degenerate case (e.g., "greater than or equal to 0"). 2. In the script, look for conditionals or defaults that set the identifier to a non-data value when the metric is zero/empty, or that filter groups out before selecting. 3. Check whether the answer's identifier slot contains an actual value drawn from the data's group labels; a placeholder string is a violation. 4. Verify the selection is deterministic (e.g., idxmax over the complete, unfiltered group index) so a valid label is returned even in the all-equal/all-zero case.
Discriminator
A real violation is emitting a placeholder when the group set is non-empty and a label could have been chosen; it is acceptable to report "none" only if the task explicitly permits it or the group set itself is genuinely empty.
Consequence
The numeric check passes but the identifier check fails, producing a partially correct, graded-wrong answer.
id 344e30ea28d8 · mined from infiagent-dabench dabench-760@s10
raw text (what the judge reads)
### Degenerate/empty result emitted as a placeholder instead of a valid identifier
- **Applies when**: `task` -- The task asks for an "argmax"-style identifier (which group/row/model has the highest count, score, etc.) plus its value, and the underlying quantity may legitimately be zero or tied for all groups.
- **Pattern**: The script branches on "if the aggregate is empty/zero, skip the selection" and writes `None`/`""`/`N/A` into the identifier slot, so the answer satisfies the numeric field but leaves the required label unfilled, instead of still selecting a concrete group (e.g., first by the specified or a deterministic ordering) from the fully enumerated group list.
- **Detection procedure**: 1. In the task statement, confirm both an identifier and a numeric field are required, and that the numeric field's allowed range includes the degenerate case (e.g., "greater than or equal to 0"). 2. In the script, look for conditionals or defaults that set the identifier to a non-data value when the metric is zero/empty, or that filter groups out before selecting. 3. Check whether the answer's identifier slot contains an actual value drawn from the data's group labels; a placeholder string is a violation. 4. Verify the selection is deterministic (e.g., `idxmax` over the complete, unfiltered group index) so a valid label is returned even in the all-equal/all-zero case.
- **Discriminator**: A real violation is emitting a placeholder when the group set is non-empty and a label could have been chosen; it is acceptable to report "none" only if the task explicitly permits it or the group set itself is genuinely empty.
- **Consequence**: The numeric check passes but the identifier check fails, producing a partially correct, graded-wrong answer.
553Fabricated metric definition instead of the one implied by the task/configtaskda-code
Applies when
task -- the task asks to visualize or report a "performance"/ranking quantity whose exact formula is not spelled out in the prompt but is constrained by an accompanying config/spec file (axis labels, title, ordering) or by upstream context, and the entity appears in the data under more than one role/key.
Pattern
The script invents an arbitrary composite formula (e.g. count plus a weighted count of some flag), computes it over only one of the roles an entity can occupy (so half the relevant records are dropped), never checks that the invented quantity reproduces the ordering/labels already given in the config, and skips producing the other required output artifacts alongside the figure.
Detection procedure
  1. From the task and the config/spec file, list every constraint on the plotted quantity: axis/legend text, title wording, the fixed category order or label list, and every output file the task expects (figure plus any serialized values/settings).
  2. In the script, find the exact expression used for the bar heights; check whether it is stated anywhere in the task/config or is an author-invented weighting, and whether it aggregates all records in which each entity participates (both/all roles, both directions) rather than a single column.
  3. Cross-check consistency: if the config supplies a category order or a "top-N" label list, verify the script confirms that sorting its own metric reproduces that order/list; if it merely reindexes to the given order without validation, the metric is unverified.
  4. Confirm every expected artifact is written (not just the image), and that the answer reports the requested quantity rather than intermediate counts.
Discriminator
Fine if the formula is quoted from the task/config or is a standard, unambiguous measure for the domain (e.g. wins/points computed from all matches an entity played) and the script demonstrates it reproduces any ordering/labels given in the spec. A violation is an ad-hoc weighted combination, single-role filtering, or a metric whose ranking silently contradicts the provided labels, with no sanity check.
Consequence
Bar heights and any serialized values differ from the reference, and missing side artifacts fail outright — all graded checks (figure, settings JSON, numeric array) report WRONG/MISSING even though the script ran without error.
id f6bd5bdba845 · mined from da-code dacode-plot-bar-006@s10
raw text (what the judge reads)
### Fabricated metric definition instead of the one implied by the task/config
- **Applies when**: `task` -- the task asks to visualize or report a "performance"/ranking quantity whose exact formula is not spelled out in the prompt but is constrained by an accompanying config/spec file (axis labels, title, ordering) or by upstream context, and the entity appears in the data under more than one role/key.
- **Pattern**: The script invents an arbitrary composite formula (e.g. count plus a weighted count of some flag), computes it over only one of the roles an entity can occupy (so half the relevant records are dropped), never checks that the invented quantity reproduces the ordering/labels already given in the config, and skips producing the other required output artifacts alongside the figure.
- **Detection procedure**:
  1. From the task and the config/spec file, list every constraint on the plotted quantity: axis/legend text, title wording, the fixed category order or label list, and every output file the task expects (figure plus any serialized values/settings).
  2. In the script, find the exact expression used for the bar heights; check whether it is stated anywhere in the task/config or is an author-invented weighting, and whether it aggregates all records in which each entity participates (both/all roles, both directions) rather than a single column.
  3. Cross-check consistency: if the config supplies a category order or a "top-N" label list, verify the script confirms that sorting its own metric reproduces that order/list; if it merely reindexes to the given order without validation, the metric is unverified.
  4. Confirm every expected artifact is written (not just the image), and that the answer reports the requested quantity rather than intermediate counts.
- **Discriminator**: Fine if the formula is quoted from the task/config or is a standard, unambiguous measure for the domain (e.g. wins/points computed from all matches an entity played) **and** the script demonstrates it reproduces any ordering/labels given in the spec. A violation is an ad-hoc weighted combination, single-role filtering, or a metric whose ranking silently contradicts the provided labels, with no sanity check.
- **Consequence**: Bar heights and any serialized values differ from the reference, and missing side artifacts fail outright — all graded checks (figure, settings JSON, numeric array) report WRONG/MISSING even though the script ran without error.
554Z-score/threshold rule not implemented or sanity-checked against the stated definitiontaskinfiagent-dabench
Applies when
task -- the task prescribes an explicit statistical rule with a fixed cutoff (e.g., standardized-score, IQR, percentile) to flag/remove rows in a numeric column and report a count.
Pattern
The attempt reports a nonzero (often large) count produced by a different rule than the one stated — e.g., a different cutoff, a robust/modified score, an IQR or percentile filter, per-group or rolling standardization, sample-vs-population SD confusion, or scoring a mis-parsed/wrong column — and never checks whether any value actually exceeds the prescribed cutoff.
Detection procedure
  1. Read the task and write down the exact rule: statistic definition, the column, and the numeric cutoff (including sign/absolute-value handling).
  2. Read the script and locate the line that computes the statistic and the comparison operator; confirm it uses the same formula (mean/SD of the full cleaned column, not median/MAD, not quantiles), the same cutoff constant, and the intended dtype-cleaned column.
  3. Check that the script prints a sanity summary (n, mean, SD, min/max of the statistic) and that the reported count is consistent with it: with an absolute cutoff of k, the count must be 0 unless max|statistic| > k, and for a bounded-size sample the theoretical fraction beyond k SDs is tiny.
  4. Compare the reported number to the format requested and to that sanity check; a count that is a large share of rows under a strict cutoff is a red flag requiring the max|statistic| evidence.
Discriminator
A genuine violation is when the code's rule/cutoff differs from the stated one, or when the reported count cannot be reconciled with the printed extreme value of the statistic. It is fine if the code implements the stated rule exactly and the data legitimately contains many extreme points, evidenced by a printed max|statistic| above the cutoff.
Consequence
The reported count diverges from the ground-truth count computed under the prescribed rule (here, a nonzero count where the correct answer is zero), so the single graded check fails.
id 56f52f6f76f6 · mined from infiagent-dabench dabench-361@s10
raw text (what the judge reads)
### Z-score/threshold rule not implemented or sanity-checked against the stated definition
- **Applies when**: `task` -- the task prescribes an explicit statistical rule with a fixed cutoff (e.g., standardized-score, IQR, percentile) to flag/remove rows in a numeric column and report a count.
- **Pattern**: The attempt reports a nonzero (often large) count produced by a different rule than the one stated — e.g., a different cutoff, a robust/modified score, an IQR or percentile filter, per-group or rolling standardization, sample-vs-population SD confusion, or scoring a mis-parsed/wrong column — and never checks whether any value actually exceeds the prescribed cutoff.
- **Detection procedure**:
  1. Read the task and write down the exact rule: statistic definition, the column, and the numeric cutoff (including sign/absolute-value handling).
  2. Read the script and locate the line that computes the statistic and the comparison operator; confirm it uses the same formula (mean/SD of the full cleaned column, not median/MAD, not quantiles), the same cutoff constant, and the intended dtype-cleaned column.
  3. Check that the script prints a sanity summary (n, mean, SD, min/max of the statistic) and that the reported count is consistent with it: with an absolute cutoff of k, the count must be 0 unless max|statistic| > k, and for a bounded-size sample the theoretical fraction beyond k SDs is tiny.
  4. Compare the reported number to the format requested and to that sanity check; a count that is a large share of rows under a strict cutoff is a red flag requiring the max|statistic| evidence.
- **Discriminator**: A genuine violation is when the code's rule/cutoff differs from the stated one, or when the reported count cannot be reconciled with the printed extreme value of the statistic. It is fine if the code implements the stated rule exactly and the data legitimately contains many extreme points, evidenced by a printed max|statistic| above the cutoff.
- **Consequence**: The reported count diverges from the ground-truth count computed under the prescribed rule (here, a nonzero count where the correct answer is zero), so the single graded check fails.
555Required output artifact never produced (answer only pasted in chat, no reproducible script)taskda-code
Applies when
task -- The task (or its README/expected-files note) implies a deliverable result file in a specific format, and/or the answer must be derivable from saved code that applies a stated mapping/transformation.
Pattern
The agent reports a formatted answer inline but leaves no saved script and no written result file, so nothing on disk matches the expected artifact; the numbers cannot be traced back to the stated preprocessing step (e.g., the prescribed label transformation) or re-checked for rounding/format compliance.
Detection procedure
  1. Read the task and any README for named expected outputs (file name, extension, key names, value types) and for any mandated transformation step.
  2. Inspect the agent's workspace/scripts: is there code that loads the data, applies the mandated mapping, computes the statistic, and writes the named file (e.g., json.dump(..., open("result.json","w")))?
  3. Compare the inline answer's keys, value types, and rounding against the requested template; verify each number could be recomputed from the saved code.
  4. If no script exists or no write-to-file call exists, flag as inadequate regardless of whether the inline numbers look plausible.
Discriminator
A real violation is the absence of a persisted artifact or of any code producing it; it is not a violation if the file is written under the required name with correct schema and the inline text is merely a convenience copy, nor if the task truly asks for text-only output with no expected files.
Consequence
The grader looks for the expected result file, finds it missing (or schema/rounding mismatched), and scores 0 even if the underlying computation was arguably right; unverifiable intermediate steps (mapping applied? ratio over which denominator?) also cannot be credited.
id b4f92879044e · mined from da-code dacode-di-text-004@s10
raw text (what the judge reads)
### Required output artifact never produced (answer only pasted in chat, no reproducible script)
- **Applies when**: `task` -- The task (or its README/expected-files note) implies a deliverable result file in a specific format, and/or the answer must be derivable from saved code that applies a stated mapping/transformation.
- **Pattern**: The agent reports a formatted answer inline but leaves no saved script and no written result file, so nothing on disk matches the expected artifact; the numbers cannot be traced back to the stated preprocessing step (e.g., the prescribed label transformation) or re-checked for rounding/format compliance.
- **Detection procedure**:
  1. Read the task and any README for named expected outputs (file name, extension, key names, value types) and for any mandated transformation step.
  2. Inspect the agent's workspace/scripts: is there code that loads the data, applies the mandated mapping, computes the statistic, and *writes* the named file (e.g., `json.dump(..., open("result.json","w"))`)?
  3. Compare the inline answer's keys, value types, and rounding against the requested template; verify each number could be recomputed from the saved code.
  4. If no script exists or no write-to-file call exists, flag as inadequate regardless of whether the inline numbers look plausible.
- **Discriminator**: A real violation is the absence of a persisted artifact or of any code producing it; it is *not* a violation if the file is written under the required name with correct schema and the inline text is merely a convenience copy, nor if the task truly asks for text-only output with no expected files.
- **Consequence**: The grader looks for the expected result file, finds it missing (or schema/rounding mismatched), and scores 0 even if the underlying computation was arguably right; unverifiable intermediate steps (mapping applied? ratio over which denominator?) also cannot be credited.
556Submission artifact not validated against the provided template and full test indextaskda-code
Applies when
task -- the task requires writing a prediction/output file whose schema, ID coverage, and row count are defined by a sample/template file and the test input.
Pattern
The agent produces an output file (or pastes an answer) using hand-written column names, an arbitrary ID ordering, or only a subset of the required rows, without ever programmatically comparing the file's header and ID set to the supplied template and test index; header casing/naming drift and truncated/partial row coverage go unnoticed.
Detection procedure
  1. Read the task/README and note the exact required output: file name, header strings (as given in the template), one row per test record, and the value semantics (e.g., probabilities per class summing to 1).
  2. Read the scripts and look for an explicit validation step: loading the template and test files, asserting list(sub.columns) == list(template.columns), len(sub) == len(test), and set(sub[id]) == set(test[id]) (and, where relevant, row order preserved and no NaNs).
  3. Inspect the produced answer/file: compare its header token-by-token with the template's header (case and spelling included) and count its rows against the test set size.
  4. Flag if any of these comparisons is absent from the scripts or fails on the artifact.
Discriminator
A real violation is a header string, ID set, row count, or value-range mismatch relative to the template/test data (e.g., renamed/re-cased ID column, far fewer rows than test records, probabilities not summing to 1). A look-alike that is fine is a file that matches the template exactly but whose predicted values differ from some reference, or cosmetic differences the template itself permits (e.g., float precision, trailing newline).
Consequence
The grader cannot align the submission with ground truth and marks the file WRONG/MISSING (0 checks passed) regardless of model quality; scoring tools would error or score only the covered subset.
id e30282f80c05 · mined from da-code dacode-ml-competition-003@s10
raw text (what the judge reads)
### Submission artifact not validated against the provided template and full test index
- **Applies when**: `task` -- the task requires writing a prediction/output file whose schema, ID coverage, and row count are defined by a sample/template file and the test input.
- **Pattern**: The agent produces an output file (or pastes an answer) using hand-written column names, an arbitrary ID ordering, or only a subset of the required rows, without ever programmatically comparing the file's header and ID set to the supplied template and test index; header casing/naming drift and truncated/partial row coverage go unnoticed.
- **Detection procedure**:
  1. Read the task/README and note the exact required output: file name, header strings (as given in the template), one row per test record, and the value semantics (e.g., probabilities per class summing to 1).
  2. Read the scripts and look for an explicit validation step: loading the template and test files, asserting `list(sub.columns) == list(template.columns)`, `len(sub) == len(test)`, and `set(sub[id]) == set(test[id])` (and, where relevant, row order preserved and no NaNs).
  3. Inspect the produced answer/file: compare its header token-by-token with the template's header (case and spelling included) and count its rows against the test set size.
  4. Flag if any of these comparisons is absent from the scripts or fails on the artifact.
- **Discriminator**: A real violation is a header string, ID set, row count, or value-range mismatch relative to the template/test data (e.g., renamed/re-cased ID column, far fewer rows than test records, probabilities not summing to 1). A look-alike that is fine is a file that matches the template exactly but whose *predicted values* differ from some reference, or cosmetic differences the template itself permits (e.g., float precision, trailing newline).
- **Consequence**: The grader cannot align the submission with ground truth and marks the file WRONG/MISSING (0 checks passed) regardless of model quality; scoring tools would error or score only the covered subset.
557Substituting an unsupervised proxy and a different input file for the specified labeled train/test setuptaskda-code
Applies when
task -- The task names specific input/prediction files and a target column whose labels exist in the provided training data, and the scripts must produce one prediction per row of the named evaluation file.
Pattern
The agent claims the named file(s) are missing or unusable, silently swaps in a different table as the "test set", and generates the target by unsupervised means (clustering, thresholds, heuristics) with an arbitrary cluster→label mapping, instead of training a supervised model on the available labeled rows and predicting for the required row set. The output therefore has the wrong number of rows, wrong row identities/order, and label values not validated against the label vocabulary observed in the training data.
Detection procedure
  1. From the task, list the exact evaluation file, the required output filename/column, and the target variable; note that graders align predictions to the evaluation file's rows.
  2. In the scripts, check which file is loaded as the prediction input and whether any labeled source is used for fitting; flag any re-definition of the evaluation set or any "file not found → use another file" fallback that isn't verified by an actual directory listing/inspection step.
  3. Check whether the predicted label values and their distribution are compared against the label set/frequencies in the labeled data, and whether output row count and key column match the evaluation file exactly.
  4. In the answer, compare the reported prediction count and category names to the evaluation file's row count and the documented/observed label categories; mismatch in either is a violation.
Discriminator
A genuine violation invents labels or an input set without evidence (no verified file listing, no label vocabulary check, row count differing from the required evaluation set). A look-alike that is fine: the agent verifies the required file's absence/alternate path, still predicts for exactly the required rows/keys, and uses unsupervised or heuristic methods only after confirming no labels exist anywhere, while matching the documented label vocabulary.
Consequence
The grader cannot join predictions to the evaluation rows (wrong length/keys) or scores near-chance on mislabeled categories, so the result file is marked WRONG/MISSING regardless of the modeling narrative.
id fb2244088444 · mined from da-code dacode-ml-multi-003@s10
raw text (what the judge reads)
### Substituting an unsupervised proxy and a different input file for the specified labeled train/test setup
- **Applies when**: `task` -- The task names specific input/prediction files and a target column whose labels exist in the provided training data, and the scripts must produce one prediction per row of the named evaluation file.
- **Pattern**: The agent claims the named file(s) are missing or unusable, silently swaps in a different table as the "test set", and generates the target by unsupervised means (clustering, thresholds, heuristics) with an arbitrary cluster→label mapping, instead of training a supervised model on the available labeled rows and predicting for the required row set. The output therefore has the wrong number of rows, wrong row identities/order, and label values not validated against the label vocabulary observed in the training data.
- **Detection procedure**:
  1. From the task, list the exact evaluation file, the required output filename/column, and the target variable; note that graders align predictions to the evaluation file's rows.
  2. In the scripts, check which file is loaded as the prediction input and whether any labeled source is used for fitting; flag any re-definition of the evaluation set or any "file not found → use another file" fallback that isn't verified by an actual directory listing/inspection step.
  3. Check whether the predicted label values and their distribution are compared against the label set/frequencies in the labeled data, and whether output row count and key column match the evaluation file exactly.
  4. In the answer, compare the reported prediction count and category names to the evaluation file's row count and the documented/observed label categories; mismatch in either is a violation.
- **Discriminator**: A genuine violation invents labels or an input set without evidence (no verified file listing, no label vocabulary check, row count differing from the required evaluation set). A look-alike that is fine: the agent verifies the required file's absence/alternate path, still predicts for exactly the required rows/keys, and uses unsupervised or heuristic methods only after confirming no labels exist anywhere, while matching the documented label vocabulary.
- **Consequence**: The grader cannot join predictions to the evaluation rows (wrong length/keys) or scores near-chance on mislabeled categories, so the result file is marked WRONG/MISSING regardless of the modeling narrative.
558Template/format file provided but never read or matchedtaskda-code
Applies when
task -- The task says the output must match a provided template/sample file (headers, index labels, column names/order, dtypes, rounding), and the scripts write a result file.
Pattern
The agent invents its own output schema (column names, index encoding, number of columns, rounding) from assumptions about the analysis instead of loading the template, and never programmatically compares its output's shape/headers/labels to it; derived keys (e.g. period offsets) are also computed with ad-hoc arithmetic rather than the convention the template implies.
Detection procedure
  1. Read the task and note any mention of a template/example/expected-format file and the required output filename/location.
  2. Search the scripts for any read of that template (e.g. read_csv on the template path) and for an explicit comparison of column names, column count, row labels, and value formatting against it.
  3. Inspect how the agent constructed labels/keys and rounding: were they justified by the template, or by guesswork (approximations like dividing day differences by a fixed number, self-chosen date string format, self-chosen decimal places)?
  4. Check the answer text for evidence of a shape/header equality check against the template; a narrative claim of "format matches" without a printed diff is not evidence.
Discriminator
Fine if the script loads the template (or the task supplies no template) and asserts/prints matching headers, row keys and dtypes before saving; a violation if the schema was assumed, or if the only "verification" recomputes the agent's own numbers under its own assumptions.
Consequence
The saved file is compared cell-by-cell against the reference and fails on mismatched headers, index labels, column count, or rounding, so the file check is marked WRONG even if the underlying aggregation logic is close.
id 9fab20a2d03a · mined from da-code dacode-dm-csv-044@s10
raw text (what the judge reads)
### Template/format file provided but never read or matched
- **Applies when**: `task` -- The task says the output must match a provided template/sample file (headers, index labels, column names/order, dtypes, rounding), and the scripts write a result file.
- **Pattern**: The agent invents its own output schema (column names, index encoding, number of columns, rounding) from assumptions about the analysis instead of loading the template, and never programmatically compares its output's shape/headers/labels to it; derived keys (e.g. period offsets) are also computed with ad-hoc arithmetic rather than the convention the template implies.
- **Detection procedure**:
  1. Read the task and note any mention of a template/example/expected-format file and the required output filename/location.
  2. Search the scripts for any read of that template (e.g. `read_csv` on the template path) and for an explicit comparison of column names, column count, row labels, and value formatting against it.
  3. Inspect how the agent constructed labels/keys and rounding: were they justified by the template, or by guesswork (approximations like dividing day differences by a fixed number, self-chosen date string format, self-chosen decimal places)?
  4. Check the answer text for evidence of a shape/header equality check against the template; a narrative claim of "format matches" without a printed diff is not evidence.
- **Discriminator**: Fine if the script loads the template (or the task supplies no template) and asserts/prints matching headers, row keys and dtypes before saving; a violation if the schema was assumed, or if the only "verification" recomputes the agent's own numbers under its own assumptions.
- **Consequence**: The saved file is compared cell-by-cell against the reference and fails on mismatched headers, index labels, column count, or rounding, so the file check is marked WRONG even if the underlying aggregation logic is close.
559Submission file not verified for completeness against the test indextaskda-code
Applies when
task -- the task requires producing a prediction/output file with one row per identifier in a provided test/holdout file, in a format shown by a sample file.
Pattern
The attempt pastes or writes a result set that is truncated, partially written, or otherwise not one-to-one with the required identifiers (missing rows, a final malformed/incomplete line, duplicated or extra IDs, wrong column names/order), and never runs an explicit check that the output matches the expected row count and ID set before declaring success.
Detection procedure
  1. From the task/README and the sample output file, note the exact required columns, header, and the fact that every test identifier must appear exactly once.
  2. In the scripts, look for code that (a) writes the file to the required path and (b) asserts/prints len(output) == len(test), set(output.id) == set(test.id), no NaNs, and columns equal to the sample's columns; absence of such checks (or absence of any saved script at all) is a red flag.
  3. Inspect the produced answer/file end-to-end: count rows, confirm the last line is complete and well-formed, and confirm all IDs are unique and drawn from the test set.
  4. Confirm the deliverable actually exists at the required filename rather than only being echoed in the response text.
Discriminator
A real violation is a row-count/ID-set/format mismatch or a truncated or missing file; a look-alike that is fine is a complete file that merely displays a preview of rows in the console while the on-disk file passes the count/ID/format checks.
Consequence
The grader marks the expected output file WRONG/MISSING — it cannot be scored (or scores near-zero) because predictions are absent for part of the test set or the file cannot be parsed/aligned.
id f40b199f4274 · mined from da-code dacode-ml-competition-006@s10
raw text (what the judge reads)
### Submission file not verified for completeness against the test index
- **Applies when**: `task` -- the task requires producing a prediction/output file with one row per identifier in a provided test/holdout file, in a format shown by a sample file.
- **Pattern**: The attempt pastes or writes a result set that is truncated, partially written, or otherwise not one-to-one with the required identifiers (missing rows, a final malformed/incomplete line, duplicated or extra IDs, wrong column names/order), and never runs an explicit check that the output matches the expected row count and ID set before declaring success.
- **Detection procedure**:
  1. From the task/README and the sample output file, note the exact required columns, header, and the fact that every test identifier must appear exactly once.
  2. In the scripts, look for code that (a) writes the file to the required path and (b) asserts/prints `len(output) == len(test)`, `set(output.id) == set(test.id)`, no NaNs, and columns equal to the sample's columns; absence of such checks (or absence of any saved script at all) is a red flag.
  3. Inspect the produced answer/file end-to-end: count rows, confirm the last line is complete and well-formed, and confirm all IDs are unique and drawn from the test set.
  4. Confirm the deliverable actually exists at the required filename rather than only being echoed in the response text.
- **Discriminator**: A real violation is a row-count/ID-set/format mismatch or a truncated or missing file; a look-alike that is fine is a complete file that merely displays a preview of rows in the console while the on-disk file passes the count/ID/format checks.
- **Consequence**: The grader marks the expected output file WRONG/MISSING — it cannot be scored (or scores near-zero) because predictions are absent for part of the test set or the file cannot be parsed/aligned.
560Null/group-membership definition not validated against the raw file's actual missing-value encodingtaskinfiagent-dabench
Applies when
task -- the analysis splits rows into groups by whether a column is "missing"/"null" (or otherwise filters rows on a sentinel value) and reports group statistics.
Pattern
The script loads the file with default/implicit parsing options (e.g., a chosen index_col, default na_values, default dtype inference) and then defines the groups purely with isnull()/notnull(), without ever inspecting how absent values are actually represented in the raw text (empty strings, whitespace, "NA", "None", ".", 0, placeholder tokens) or verifying that the column being tested is the intended one after parsing.
Detection procedure
  1. Read the task to see which column defines the grouping and which column the statistic is computed on.
  2. In the script, check whether there is any step that inspects the raw column contents (unique values, value counts, sample of odd values, row/column counts vs. the raw header) before relying on isnull(); also check whether index/column-selection options could shift or drop columns.
  3. Check whether group sizes are reported and cross-validated (null count + non-null count == total rows; group means recomputed under an alternative missing-value definition to see if they move).
  4. Compare the answer's group sizes/means to any sanity expectation stated or derivable (e.g., all rows accounted for, means plausible given the column's range); an answer with no such check and no raw-value inspection is inadequate.
Discriminator
Fine if the script explicitly enumerates the column's distinct/anomalous values (or passes explicit na_values/dtype and confirms counts sum to the full row count), showing the missingness definition was verified; a violation is relying solely on default parsing plus isnull() with no verification, so mis-encoded missing values silently land in the wrong group.
Consequence
A handful of rows are assigned to the wrong group, so both group means (and the test statistic) are close to but not equal to the expected values, and the grader marks the numeric fields wrong even though the code "looks" correct.
id 972ec9270b9c · mined from infiagent-dabench dabench-297@s10
raw text (what the judge reads)
### Null/group-membership definition not validated against the raw file's actual missing-value encoding
- **Applies when**: `task` -- the analysis splits rows into groups by whether a column is "missing"/"null" (or otherwise filters rows on a sentinel value) and reports group statistics.
- **Pattern**: The script loads the file with default/implicit parsing options (e.g., a chosen `index_col`, default `na_values`, default dtype inference) and then defines the groups purely with `isnull()`/`notnull()`, without ever inspecting how absent values are actually represented in the raw text (empty strings, whitespace, `"NA"`, `"None"`, `"."`, `0`, placeholder tokens) or verifying that the column being tested is the intended one after parsing.
- **Detection procedure**:
  1. Read the task to see which column defines the grouping and which column the statistic is computed on.
  2. In the script, check whether there is any step that inspects the raw column contents (unique values, value counts, sample of odd values, row/column counts vs. the raw header) before relying on `isnull()`; also check whether index/column-selection options could shift or drop columns.
  3. Check whether group sizes are reported and cross-validated (null count + non-null count == total rows; group means recomputed under an alternative missing-value definition to see if they move).
  4. Compare the answer's group sizes/means to any sanity expectation stated or derivable (e.g., all rows accounted for, means plausible given the column's range); an answer with no such check and no raw-value inspection is inadequate.
- **Discriminator**: Fine if the script explicitly enumerates the column's distinct/anomalous values (or passes explicit `na_values`/dtype and confirms counts sum to the full row count), showing the missingness definition was verified; a violation is relying solely on default parsing plus `isnull()` with no verification, so mis-encoded missing values silently land in the wrong group.
- **Consequence**: A handful of rows are assigned to the wrong group, so both group means (and the test statistic) are close to but not equal to the expected values, and the grader marks the numeric fields wrong even though the code "looks" correct.
561Deliverables and rules improvised instead of taken from the provided spectaskda-code
Applies when
task -- the task points to an external instructions/guidance file (or a fixed harness) that defines the preprocessing rules and the set of output artifacts the script must write.
Pattern
The script never demonstrably reads/quotes the spec; it invents its own filtering thresholds, category-mapping rules, and orderings from guesswork, and it saves only the one artifact named in the prompt's prose (e.g. the image) while omitting the companion machine-readable outputs (serialized plot data, numeric result arrays/tables) the checker also expects.
Detection procedure
  1. From the task text, list every artifact the environment expects and every rule the spec is said to contain (filters, category definitions, colors/order, figure size, file names/paths).
  2. Grep the scripts for reads of the spec file and for a write/save call per expected artifact; build a checklist of produced vs. required files and their exact save locations.
  3. For each analytic rule in the script (threshold values, keyword-to-category mapping, priority order, reference date), check whether it is traceable to the spec or is an assumption authored in the script (comments like "based on guidance" with no quoted source, or docstring-invented rules).
  4. Compare the final answer text: does it claim compliance with rules/artifacts that the script never actually verified or wrote?
Discriminator
A real violation is missing output files or rule values with no provenance in the spec; a look-alike that is fine is a script that reads the spec (or reproduces its exact stated values) and writes every required artifact, even if it additionally logs extras or chooses harmless cosmetic defaults the spec leaves open.
Consequence
The grader reports the missing artifacts as WRONG/MISSING and the produced artifact mismatches too, since counts/proportions come from self-invented filters — a 0/N score regardless of how polished the reported summary looks.
id 3dd0046e6656 · mined from da-code dacode-plot-pie-005@s10
raw text (what the judge reads)
### Deliverables and rules improvised instead of taken from the provided spec
- **Applies when**: `task` -- the task points to an external instructions/guidance file (or a fixed harness) that defines the preprocessing rules and the set of output artifacts the script must write.
- **Pattern**: The script never demonstrably reads/quotes the spec; it invents its own filtering thresholds, category-mapping rules, and orderings from guesswork, and it saves only the one artifact named in the prompt's prose (e.g. the image) while omitting the companion machine-readable outputs (serialized plot data, numeric result arrays/tables) the checker also expects.
- **Detection procedure**:
  1. From the task text, list every artifact the environment expects and every rule the spec is said to contain (filters, category definitions, colors/order, figure size, file names/paths).
  2. Grep the scripts for reads of the spec file and for a write/save call per expected artifact; build a checklist of produced vs. required files and their exact save locations.
  3. For each analytic rule in the script (threshold values, keyword-to-category mapping, priority order, reference date), check whether it is traceable to the spec or is an assumption authored in the script (comments like "based on guidance" with no quoted source, or docstring-invented rules).
  4. Compare the final answer text: does it claim compliance with rules/artifacts that the script never actually verified or wrote?
- **Discriminator**: A real violation is missing output files or rule values with no provenance in the spec; a look-alike that is fine is a script that reads the spec (or reproduces its exact stated values) and writes every required artifact, even if it additionally logs extras or chooses harmless cosmetic defaults the spec leaves open.
- **Consequence**: The grader reports the missing artifacts as WRONG/MISSING and the produced artifact mismatches too, since counts/proportions come from self-invented filters — a 0/N score regardless of how polished the reported summary looks.
562Preprocessing shortcut: rows silently lost or columns mis-parsed instead of the prescribed imputationtaskinfiagent-dabench
Applies when
task -- the task prescribes an exact preprocessing recipe (e.g., "impute missing values in columns X, Y, Z with their column means") before fitting and scoring a model on a fixed train/test split.
Pattern
The attempt converts the target/feature columns to numeric without checking how they are stored (strings with separators, units, ranges, sentinel/placeholder values), or calls dropna()/read_csv filters/errors='coerce' followed by row removal, so the modelled table has fewer rows than the source; the required mean imputation is then applied to a different (already-filtered or badly-coerced) population, changing both the imputed constants and the split contents.
Detection procedure
  1. Read the task and list the exact preprocessing steps and the columns they must touch; note the expected row count of the dataset.
  2. In the scripts, trace each named column from load to model input: check dtype inspection/cleaning, any coercion to numeric, and any row-dropping call (dropna, boolean filters, merge, groupby) occurring before or after the imputation step.
  3. Require an explicit shape/NaN sanity print: rows before vs. after preprocessing must be identical, and post-imputation NaN count must be zero; also require the imputation to use the mean of the full parsed column, not of the training subset only if the task says otherwise (or vice versa if leakage-free imputation is specified).
  4. Compare the reported metric against a crude sanity baseline (e.g., variance of the target, or MSE of predicting the target mean) — an MSE of the same order as, or larger than, the target variance signals a broken feature/target pipeline.
Discriminator
A real violation is when row counts shrink, NaNs remain, or numeric conversion silently nulls non-trivial fractions of a column relative to the raw file; it is not a violation if the script inspects dtypes, cleans string formatting deliberately, keeps the full row count, and imputes exactly the named columns as instructed.
Consequence
The model is fit and scored on a different sample and scale than intended, so the reported error metric is off by a large factor (often an order of magnitude) and fails the exact-value check, with no saved script or sanity output to diagnose it.
id e8772ccb5206 · mined from infiagent-dabench dabench-432@s10
raw text (what the judge reads)
### Preprocessing shortcut: rows silently lost or columns mis-parsed instead of the prescribed imputation
- **Applies when**: `task` -- the task prescribes an exact preprocessing recipe (e.g., "impute missing values in columns X, Y, Z with their column means") before fitting and scoring a model on a fixed train/test split.
- **Pattern**: The attempt converts the target/feature columns to numeric without checking how they are stored (strings with separators, units, ranges, sentinel/placeholder values), or calls `dropna()`/`read_csv` filters/`errors='coerce'` followed by row removal, so the modelled table has fewer rows than the source; the required mean imputation is then applied to a different (already-filtered or badly-coerced) population, changing both the imputed constants and the split contents.
- **Detection procedure**:
  1. Read the task and list the exact preprocessing steps and the columns they must touch; note the expected row count of the dataset.
  2. In the scripts, trace each named column from load to model input: check dtype inspection/cleaning, any coercion to numeric, and any row-dropping call (`dropna`, boolean filters, `merge`, `groupby`) occurring before or after the imputation step.
  3. Require an explicit shape/NaN sanity print: rows before vs. after preprocessing must be identical, and post-imputation NaN count must be zero; also require the imputation to use the mean of the full parsed column, not of the training subset only if the task says otherwise (or vice versa if leakage-free imputation is specified).
  4. Compare the reported metric against a crude sanity baseline (e.g., variance of the target, or MSE of predicting the target mean) — an MSE of the same order as, or larger than, the target variance signals a broken feature/target pipeline.
- **Discriminator**: A real violation is when row counts shrink, NaNs remain, or numeric conversion silently nulls non-trivial fractions of a column relative to the raw file; it is *not* a violation if the script inspects dtypes, cleans string formatting deliberately, keeps the full row count, and imputes exactly the named columns as instructed.
- **Consequence**: The model is fit and scored on a different sample and scale than intended, so the reported error metric is off by a large factor (often an order of magnitude) and fails the exact-value check, with no saved script or sanity output to diagnose it.
563Ordering not verified before computing sequential/lagged quantitiestaskinfiagent-dabench
Applies when
task -- the task asks for a statistic derived from row-to-row differences, lags, shifts, cumulative operations, or any "previous record" comparison in a table that has an explicit time/sequence key.
Pattern
The attempt applies a shift/diff/pct_change-style operation directly to the file's native row order without sorting (or checking the sort direction) on the sequence key, so "previous" is actually "next"; the resulting series is sign-inverted (or otherwise misaligned), and the reported mean/aggregate has the wrong sign or magnitude while the dispersion statistic looks nearly right, hiding the error.
Detection procedure
  1. In the task statement, identify whether the requested quantity depends on record order (words like previous, prior, change, growth, lag, rolling, cumulative).
  2. In the scripts, look for an explicit sort_values/index-sort on the sequence key immediately before the shift/diff step, plus a check of the first/last rows or the min/max of the key to confirm ascending order; also check that the first row's undefined value is dropped rather than filled.
  3. Compare the answer's sign and scale against a quick expectation from the data (e.g., whether the series ends above or below where it starts): a mean change whose sign contradicts the overall trend indicates reversed ordering.
  4. Flag if no sorting/ordering verification exists anywhere before the sequential computation.
Discriminator
Not a violation if the script demonstrably verifies ascending order (sorts, or asserts/prints that the key is monotonic increasing) or if the data has no sequence key and order is defined as-given; it is a violation when order is simply assumed, especially when the reported aggregate's sign is unexamined.
Consequence
The dispersion statistic matches to within rounding while the mean comes out with the opposite sign, so the graded answer fails on the location statistic (and often on both after rounding).
id 57037919d835 · mined from infiagent-dabench dabench-75@s10
raw text (what the judge reads)
### Ordering not verified before computing sequential/lagged quantities
- **Applies when**: `task` -- the task asks for a statistic derived from row-to-row differences, lags, shifts, cumulative operations, or any "previous record" comparison in a table that has an explicit time/sequence key.
- **Pattern**: The attempt applies a shift/diff/`pct_change`-style operation directly to the file's native row order without sorting (or checking the sort direction) on the sequence key, so "previous" is actually "next"; the resulting series is sign-inverted (or otherwise misaligned), and the reported mean/aggregate has the wrong sign or magnitude while the dispersion statistic looks nearly right, hiding the error.
- **Detection procedure**:
  1. In the task statement, identify whether the requested quantity depends on record order (words like previous, prior, change, growth, lag, rolling, cumulative).
  2. In the scripts, look for an explicit `sort_values`/index-sort on the sequence key immediately before the shift/diff step, plus a check of the first/last rows or the min/max of the key to confirm ascending order; also check that the first row's undefined value is dropped rather than filled.
  3. Compare the answer's sign and scale against a quick expectation from the data (e.g., whether the series ends above or below where it starts): a mean change whose sign contradicts the overall trend indicates reversed ordering.
  4. Flag if no sorting/ordering verification exists anywhere before the sequential computation.
- **Discriminator**: Not a violation if the script demonstrably verifies ascending order (sorts, or asserts/prints that the key is monotonic increasing) or if the data has no sequence key and order is defined as-given; it *is* a violation when order is simply assumed, especially when the reported aggregate's sign is unexamined.
- **Consequence**: The dispersion statistic matches to within rounding while the mean comes out with the opposite sign, so the graded answer fails on the location statistic (and often on both after rounding).
564Output schema invented instead of read from the referenced spec/instructions filetaskda-code
Applies when
task -- The task says to follow an external instructions/spec artifact (e.g., a tips/README/template file) and to save results in a "required format" with named fields.
Pattern
The script never opens or echoes the referenced spec; the agent guesses the column names, their number/order, the wording of the decision/comment strings, the test to use, and the numeric formatting, then writes a file that is structurally plausible but not the mandated schema (extra columns, renamed fields, truncated/zero-padded p-value).
Detection procedure
  1. In the task text, list every referenced artifact (spec file, template, example row) and every explicitly enumerated output field.
  2. Scan the scripts for any read/print of that artifact; if absent, the output schema is unverified by construction.
  3. Compare the written header/values to the fields the task enumerates: count them, check names and order, and check that no undocumented extra columns were added and no required field was renamed or merged.
  4. Sanity-check value formatting against what a spec would plausibly demand (e.g., a p-value printed as 0.000000 instead of scientific notation or the requested rounding), and confirm the chosen test matches what the spec/assumption checks imply.
Discriminator
A real violation is when a spec artifact exists and was never consulted, or when the produced header/field set demonstrably deviates from the enumerated fields. It is fine if the script reads the spec (or the task itself fully enumerates the schema) and the output columns match one-to-one, even if internal analysis code is verbose.
Consequence
The file exists with one row but fails exact-match/field-wise comparison against the expected format, so the grader marks the result file WRONG despite the underlying statistic being defensible.
id 8025744135aa · mined from da-code dacode-data-sa-004@s10
raw text (what the judge reads)
### Output schema invented instead of read from the referenced spec/instructions file
- **Applies when**: `task` -- The task says to follow an external instructions/spec artifact (e.g., a tips/README/template file) and to save results in a "required format" with named fields.
- **Pattern**: The script never opens or echoes the referenced spec; the agent guesses the column names, their number/order, the wording of the decision/comment strings, the test to use, and the numeric formatting, then writes a file that is structurally plausible but not the mandated schema (extra columns, renamed fields, truncated/zero-padded p-value).
- **Detection procedure**:
  1. In the task text, list every referenced artifact (spec file, template, example row) and every explicitly enumerated output field.
  2. Scan the scripts for any read/print of that artifact; if absent, the output schema is unverified by construction.
  3. Compare the written header/values to the fields the task enumerates: count them, check names and order, and check that no undocumented extra columns were added and no required field was renamed or merged.
  4. Sanity-check value formatting against what a spec would plausibly demand (e.g., a p-value printed as `0.000000` instead of scientific notation or the requested rounding), and confirm the chosen test matches what the spec/assumption checks imply.
- **Discriminator**: A real violation is when a spec artifact exists and was never consulted, or when the produced header/field set demonstrably deviates from the enumerated fields. It is fine if the script reads the spec (or the task itself fully enumerates the schema) and the output columns match one-to-one, even if internal analysis code is verbose.
- **Consequence**: The file exists with one row but fails exact-match/field-wise comparison against the expected format, so the grader marks the result file WRONG despite the underlying statistic being defensible.
565Ambiguous scaling method chosen so that the requested statistic becomes trivially degeneratetaskinfiagent-dabench
Applies when
task -- the task asks to "normalize"/"scale" columns and then report a summary statistic (e.g., the mean) of those scaled columns.
Pattern
The agent applies zero-centering standardization (subtract mean, divide by std) instead of range-based min–max scaling to [0,1], so every reported mean collapses to ~0 (or ~-0.0000); the agent reports these degenerate values without questioning why the task would ask for a statistic that is identical and information-free for every column.
Detection procedure
  1. Read the task: note which columns must be scaled and which statistic is to be reported afterwards.
  2. In the script, identify the scaler/formula used (StandardScaler/(x-mean)/std vs MinMaxScaler/(x-min)/(max-min)) and check whether the task text disambiguates it; if not, check whether the agent tested/justified the alternative.
  3. Inspect the reported values: if the requested statistic is mathematically forced to a constant (all ~0.0000, or all 1.0) by the chosen transform, treat the choice as suspect — a task would not ask for per-column values that cannot differ.
  4. Confirm the untouched columns (e.g., encoded/binary ones) still show plausible, varied values while all scaled ones are identical — a strong signal of the wrong scaling convention.
Discriminator
A real violation is when the reported statistic is a mathematical constant of the transform (no information, all columns equal) and the task expected per-column distinct values; it is fine if the task explicitly names standardization, or if the reported statistic (e.g., std, min, max, or a model metric) still varies meaningfully across columns under the chosen transform.
Consequence
Every scaled column's reported number is ~0.0000 instead of the distinct in-[0,1] means expected, so all those checks fail while only the non-scaled/binary columns match.
id 89960b2e53a5 · mined from infiagent-dabench dabench-28@s10
raw text (what the judge reads)
### Ambiguous scaling method chosen so that the requested statistic becomes trivially degenerate
- **Applies when**: `task` -- the task asks to "normalize"/"scale" columns and then report a summary statistic (e.g., the mean) of those scaled columns.
- **Pattern**: The agent applies zero-centering standardization (subtract mean, divide by std) instead of range-based min–max scaling to [0,1], so every reported mean collapses to ~0 (or ~-0.0000); the agent reports these degenerate values without questioning why the task would ask for a statistic that is identical and information-free for every column.
- **Detection procedure**:
  1. Read the task: note which columns must be scaled and which statistic is to be reported afterwards.
  2. In the script, identify the scaler/formula used (`StandardScaler`/`(x-mean)/std` vs `MinMaxScaler`/`(x-min)/(max-min)`) and check whether the task text disambiguates it; if not, check whether the agent tested/justified the alternative.
  3. Inspect the reported values: if the requested statistic is mathematically forced to a constant (all ~0.0000, or all 1.0) by the chosen transform, treat the choice as suspect — a task would not ask for per-column values that cannot differ.
  4. Confirm the untouched columns (e.g., encoded/binary ones) still show plausible, varied values while all scaled ones are identical — a strong signal of the wrong scaling convention.
- **Discriminator**: A real violation is when the reported statistic is a mathematical constant of the transform (no information, all columns equal) and the task expected per-column distinct values; it is fine if the task explicitly names standardization, or if the reported statistic (e.g., std, min, max, or a model metric) still varies meaningfully across columns under the chosen transform.
- **Consequence**: Every scaled column's reported number is ~0.0000 instead of the distinct in-[0,1] means expected, so all those checks fail while only the non-scaled/binary columns match.
566Unjustified row-filtering/NaN handling that silently changes a rounded statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single summary statistic over two or more columns of a table, and the script applies its own cleaning step (dropna, subsetting, type coercion, deduplication, index column choice) before computing it.
Pattern
The script drops or keeps rows using an assumption the task never states (e.g. dropna() on only the two columns, coercing non-numeric entries, or reading with an implicit index/header), reports the resulting number to the required rounding, and never checks whether an equally plausible alternative subset would change the reported digits.
Detection procedure
  1. Read the task for any stated filtering/inclusion rule; note that if none is given, the full table as-loaded is the default population.
  2. In the script, list every step that changes the row count or dtype before the statistic (dropna, filters, index_col, coercion, column selection) and check whether each is mandated by the task.
  3. Check whether the script prints the row count actually used and compares it to the raw row count, and whether it recomputes the statistic under at least one alternative handling (e.g. pairwise vs. listwise deletion, including/excluding suspicious sentinel values like 0 or -1).
  4. Check the reported value's sensitivity: if the statistic is rounded to few decimals, confirm the script shows the unrounded value far enough from a rounding boundary that the alternative handling cannot flip it.
Discriminator
A real violation is an unmandated or unverified row/dtype filter with no reported row counts and no alternative-handling comparison; it is fine if the task explicitly prescribes the filtering, or the script demonstrates (printed counts + recomputation) that the statistic is identical to the requested precision across the plausible handlings.
Consequence
The categorical part of the answer (e.g. significance verdict) may still match, but the numeric statistic differs from ground truth in the last required decimal, so the grader marks that field wrong and the submission fails.
id e7656cd80845 · mined from infiagent-dabench dabench-300@s10
raw text (what the judge reads)
### Unjustified row-filtering/NaN handling that silently changes a rounded statistic
- **Applies when**: `task` -- the task asks for a single summary statistic over two or more columns of a table, and the script applies its own cleaning step (dropna, subsetting, type coercion, deduplication, index column choice) before computing it.
- **Pattern**: The script drops or keeps rows using an assumption the task never states (e.g. `dropna()` on only the two columns, coercing non-numeric entries, or reading with an implicit index/header), reports the resulting number to the required rounding, and never checks whether an equally plausible alternative subset would change the reported digits.
- **Detection procedure**:
  1. Read the task for any stated filtering/inclusion rule; note that if none is given, the full table as-loaded is the default population.
  2. In the script, list every step that changes the row count or dtype before the statistic (dropna, filters, `index_col`, coercion, column selection) and check whether each is mandated by the task.
  3. Check whether the script prints the row count actually used and compares it to the raw row count, and whether it recomputes the statistic under at least one alternative handling (e.g. pairwise vs. listwise deletion, including/excluding suspicious sentinel values like 0 or -1).
  4. Check the reported value's sensitivity: if the statistic is rounded to few decimals, confirm the script shows the unrounded value far enough from a rounding boundary that the alternative handling cannot flip it.
- **Discriminator**: A real violation is an unmandated or unverified row/dtype filter with no reported row counts and no alternative-handling comparison; it is fine if the task explicitly prescribes the filtering, or the script demonstrates (printed counts + recomputation) that the statistic is identical to the requested precision across the plausible handlings.
- **Consequence**: The categorical part of the answer (e.g. significance verdict) may still match, but the numeric statistic differs from ground truth in the last required decimal, so the grader marks that field wrong and the submission fails.
567Predicted target distribution never sanity-checked against the training target distributiontaskda-code
Applies when
task -- a script fits a model on a labeled training set and writes per-row predictions of a continuous target to a submission file.
Pattern
The attempt reports fit/validation metrics and prediction summary statistics, but never compares the predicted distribution (min/median/mean/max, spread, tail) with the distribution of the observed target in the training data; gross mis-scaling, mis-transformation (e.g., forgotten log/inverse transform, wrong unit), or extreme extrapolation therefore goes unnoticed, and the reported summary numbers are internally inconsistent with each other or with the claimed post-processing (e.g., a stated floor/clip that the reported minimum contradicts).
Detection procedure
  1. From the task/README, note the target and read from the scripts (or the answer's own numbers) the training target's central tendency and range.
  2. In the scripts, check whether any explicit comparison/assertion exists between predicted and training target distributions (quantiles, mean ratio, share above the training max) before writing the output file, and whether any transform applied to the target is inverted symmetrically.
  3. Compare the answer's reported prediction stats to the training target stats and to the reported error metrics: flag if the predicted mean/median is off by a large factor, if the max sits far beyond the observed target range, or if the reported RMSE/MAE is implausibly large relative to typical target values.
  4. Check the answer's stats for self-contradiction (claimed clipping/rounding/positivity vs. the reported min/max/dtype).
Discriminator
A real violation shows a systematic shift or inflation of the whole predicted distribution (or contradictory reported bounds) with no verification step; it is fine if predictions are somewhat smoother/narrower than the training target (normal regression shrinkage) and the script or answer demonstrates the quantiles line up with training data on the same scale.
Consequence
The submitted prediction column is on the wrong scale or dominated by out-of-range values, so the grader's held-out error/agreement check against true values fails even though the reported in-script validation score looked high.
id 130cd5d13c5e · mined from da-code dacode-ml-regression-014@s10
raw text (what the judge reads)
### Predicted target distribution never sanity-checked against the training target distribution
- **Applies when**: `task` -- a script fits a model on a labeled training set and writes per-row predictions of a continuous target to a submission file.
- **Pattern**: The attempt reports fit/validation metrics and prediction summary statistics, but never compares the predicted distribution (min/median/mean/max, spread, tail) with the distribution of the observed target in the training data; gross mis-scaling, mis-transformation (e.g., forgotten log/inverse transform, wrong unit), or extreme extrapolation therefore goes unnoticed, and the reported summary numbers are internally inconsistent with each other or with the claimed post-processing (e.g., a stated floor/clip that the reported minimum contradicts).
- **Detection procedure**:
  1. From the task/README, note the target and read from the scripts (or the answer's own numbers) the training target's central tendency and range.
  2. In the scripts, check whether any explicit comparison/assertion exists between predicted and training target distributions (quantiles, mean ratio, share above the training max) before writing the output file, and whether any transform applied to the target is inverted symmetrically.
  3. Compare the answer's reported prediction stats to the training target stats and to the reported error metrics: flag if the predicted mean/median is off by a large factor, if the max sits far beyond the observed target range, or if the reported RMSE/MAE is implausibly large relative to typical target values.
  4. Check the answer's stats for self-contradiction (claimed clipping/rounding/positivity vs. the reported min/max/dtype).
- **Discriminator**: A real violation shows a systematic shift or inflation of the whole predicted distribution (or contradictory reported bounds) with no verification step; it is fine if predictions are somewhat smoother/narrower than the training target (normal regression shrinkage) and the script or answer demonstrates the quantiles line up with training data on the same scale.
- **Consequence**: The submitted prediction column is on the wrong scale or dominated by out-of-range values, so the grader's held-out error/agreement check against true values fails even though the reported in-script validation score looked high.
568Unvalidated model output: no held-out accuracy estimate and no check for directly recoverable labelstaskda-code
Applies when
task -- the script fits a predictor on a provided labeled source file and writes predictions for a separate evaluation file whose labels are graded for correctness.
Pattern
The attempt trains one model with arbitrary, unvalidated settings (single feature column, capped vocabulary, default hyperparameters), never measures accuracy on any held-out data, never compares against a stronger/alternative approach, and never checks whether the evaluation rows already appear in the labeled source file (so their labels could be looked up or, conversely, so that an accuracy estimate isn't inflated by overlap). It then submits the raw predictions as the final answer.
Detection procedure
  1. Read the task to see that correctness is judged against true labels for the evaluation rows, i.e., a quality threshold must be met, not merely a file produced.
  2. Scan the script for any train/validation split, cross-validation, or score printout on labeled data; note if the only printed diagnostics are shapes and prediction class counts.
  3. Check whether the script ever joins the evaluation rows to the labeled source file on an identifier or on the text/key columns, to test for overlap (labels obtainable exactly) and to confirm row count/order alignment of the output with the evaluation file.
  4. If neither a held-out score nor an overlap/alignment check exists, and the model choice was never compared to any alternative, flag the attempt.
Discriminator
A fine attempt reports at least one honest held-out (or cross-validated) score, confirms the output length/order matches the evaluation rows, and — if the source file may contain the evaluation records — explicitly checks and exploits or excludes that overlap. A violation submits predictions whose expected accuracy is entirely unknown and unbounded from below, with only distributional eyeballing as "evidence".
Consequence
The grader compares predicted labels to ground truth and the accuracy falls below the acceptance threshold (or the file is misaligned with the evaluation rows), so the single correctness check fails with no diagnostic in the logs explaining why.
id 7776bfa80fd6 · mined from da-code dacode-ml-multi-008@s10
raw text (what the judge reads)
### Unvalidated model output: no held-out accuracy estimate and no check for directly recoverable labels
- **Applies when**: `task` -- the script fits a predictor on a provided labeled source file and writes predictions for a separate evaluation file whose labels are graded for correctness.
- **Pattern**: The attempt trains one model with arbitrary, unvalidated settings (single feature column, capped vocabulary, default hyperparameters), never measures accuracy on any held-out data, never compares against a stronger/alternative approach, and never checks whether the evaluation rows already appear in the labeled source file (so their labels could be looked up or, conversely, so that an accuracy estimate isn't inflated by overlap). It then submits the raw predictions as the final answer.
- **Detection procedure**:
  1. Read the task to see that correctness is judged against true labels for the evaluation rows, i.e., a quality threshold must be met, not merely a file produced.
  2. Scan the script for any train/validation split, cross-validation, or score printout on labeled data; note if the only printed diagnostics are shapes and prediction class counts.
  3. Check whether the script ever joins the evaluation rows to the labeled source file on an identifier or on the text/key columns, to test for overlap (labels obtainable exactly) and to confirm row count/order alignment of the output with the evaluation file.
  4. If neither a held-out score nor an overlap/alignment check exists, and the model choice was never compared to any alternative, flag the attempt.
- **Discriminator**: A fine attempt reports at least one honest held-out (or cross-validated) score, confirms the output length/order matches the evaluation rows, and — if the source file may contain the evaluation records — explicitly checks and exploits or excludes that overlap. A violation submits predictions whose expected accuracy is entirely unknown and unbounded from below, with only distributional eyeballing as "evidence".
- **Consequence**: The grader compares predicted labels to ground truth and the accuracy falls below the acceptance threshold (or the file is misaligned with the evaluation rows), so the single correctness check fails with no diagnostic in the logs explaining why.
569Dropping rows with missing values instead of imputing, altering the splittaskinfiagent-dabench
Applies when
task -- The task prescribes a fixed split (e.g., fixed proportion and random seed) and a metric on the held-out set, and the chosen feature columns contain missing values that the script must handle.
Pattern
The script calls a blanket row-drop (dropna()) on the feature/target subset before splitting, so the row count, row indices, and therefore the exact train/test partition differ from the canonical pipeline; the reported metric is computed on a different (smaller) test set than the task implies. A related variant is loading the wrong file/subset without checking it matches the described data.
Detection procedure
  1. Read the task for the specified split parameters and note that they only reproduce the reference result if the row set fed to the split is the full, unfiltered dataset.
  2. In the script, look for any row-removal step (dropna, filtering, deduplication) applied before the split, and check whether the task authorized removing rows; compare the printed pre- and post-cleaning row counts.
  3. Check whether an imputation/fill strategy (mean/median/mode, or the model-agnostic default) that preserves all rows was considered instead, and whether the resulting test-set size matches the expected fraction of the original file size.
  4. Verify the reported metric is computed on that expected-size test set, and that the loaded input file is the one the task describes.
Discriminator
A real violation is unjustified row loss that changes the number of rows entering the split (test-set size ≠ stated fraction × original rows). It is fine if the data genuinely has no missing values in the used columns (row count unchanged), or if the task explicitly instructs dropping incomplete records.
Consequence
The accuracy is computed on a different test partition than the reference, producing a value that is close but off by a few points (e.g., 0.76 vs 0.78) and marked wrong by an exact-match grader.
id 42ad97f133a5 · mined from infiagent-dabench dabench-7@s10
raw text (what the judge reads)
### Dropping rows with missing values instead of imputing, altering the split
- **Applies when**: `task` -- The task prescribes a fixed split (e.g., fixed proportion and random seed) and a metric on the held-out set, and the chosen feature columns contain missing values that the script must handle.
- **Pattern**: The script calls a blanket row-drop (`dropna()`) on the feature/target subset before splitting, so the row count, row indices, and therefore the exact train/test partition differ from the canonical pipeline; the reported metric is computed on a different (smaller) test set than the task implies. A related variant is loading the wrong file/subset without checking it matches the described data.
- **Detection procedure**:
  1. Read the task for the specified split parameters and note that they only reproduce the reference result if the row set fed to the split is the full, unfiltered dataset.
  2. In the script, look for any row-removal step (dropna, filtering, deduplication) applied before the split, and check whether the task authorized removing rows; compare the printed pre- and post-cleaning row counts.
  3. Check whether an imputation/fill strategy (mean/median/mode, or the model-agnostic default) that preserves all rows was considered instead, and whether the resulting test-set size matches the expected fraction of the original file size.
  4. Verify the reported metric is computed on that expected-size test set, and that the loaded input file is the one the task describes.
- **Discriminator**: A real violation is unjustified row loss that changes the number of rows entering the split (test-set size ≠ stated fraction × original rows). It is fine if the data genuinely has no missing values in the used columns (row count unchanged), or if the task explicitly instructs dropping incomplete records.
- **Consequence**: The accuracy is computed on a different test partition than the reference, producing a value that is close but off by a few points (e.g., 0.76 vs 0.78) and marked wrong by an exact-match grader.
570Ignoring the provided output template's implied schema and accepting a degenerate (empty) resulttaskda-code
Applies when
task -- the task supplies a skeleton/example result file (with headers, ID columns, or pre-filled category rows) that the analysis must fill in, and the scripts must define groups/categories and a selection rule to populate it.
Pattern
The agent invents its own category definitions, ID scheme, and "best-in-group" rule instead of reading the template first and reverse-engineering the required label vocabulary, row set, and key semantics; it then writes a file that is empty or contains rows whose labels/IDs cannot match the expected ones, and reports this as the finding without questioning it.
Detection procedure
  1. Read the task for any mention of a provided result file/format, and note that the categories, key column meaning, and row granularity are constraints, not free choices.
  2. In the scripts, check whether the template is loaded and inspected before the computation, and whether category labels, ID values, and column order are derived from it; flag hard-coded, self-invented bins/labels/1-based row indices, and catch-all buckets (e.g., an "Unknown"/NaN group) that the template does not contain.
  3. Check whether the pipeline validates its output against the template (same columns, same/expected number of rows, IDs drawn from the allowed set) before saving.
  4. Read the answer: if the produced file has zero data rows, or the count of qualifying items is 0 while the script itself reported many qualifying items overall, treat the absence of any re-examination of the definitions as a failure.
Discriminator
A real violation is inventing the grouping/labeling/ID scheme when a template or dataset column dictates it, and/or shipping an empty or unvalidated file. It is fine if the template genuinely only fixes headers and the agent documents a principled, data-derived grouping, verifies the schema, and the empty/short result is corroborated by an explicit sanity check showing no candidates exist under any reasonable definition.
Consequence
The saved file's rows (or lack of rows) don't match the expected keys/values, so the file check fails outright — 0/1 — even though the narrative summary looks internally consistent.
id ace29df69052 · mined from da-code dacode-dm-csv-001@s10
raw text (what the judge reads)
### Ignoring the provided output template's implied schema and accepting a degenerate (empty) result
- **Applies when**: `task` -- the task supplies a skeleton/example result file (with headers, ID columns, or pre-filled category rows) that the analysis must fill in, and the scripts must define groups/categories and a selection rule to populate it.
- **Pattern**: The agent invents its own category definitions, ID scheme, and "best-in-group" rule instead of reading the template first and reverse-engineering the required label vocabulary, row set, and key semantics; it then writes a file that is empty or contains rows whose labels/IDs cannot match the expected ones, and reports this as the finding without questioning it.
- **Detection procedure**:
  1. Read the task for any mention of a provided result file/format, and note that the categories, key column meaning, and row granularity are constraints, not free choices.
  2. In the scripts, check whether the template is loaded and inspected *before* the computation, and whether category labels, ID values, and column order are derived from it; flag hard-coded, self-invented bins/labels/1-based row indices, and catch-all buckets (e.g., an "Unknown"/NaN group) that the template does not contain.
  3. Check whether the pipeline validates its output against the template (same columns, same/expected number of rows, IDs drawn from the allowed set) before saving.
  4. Read the answer: if the produced file has zero data rows, or the count of qualifying items is 0 while the script itself reported many qualifying items overall, treat the absence of any re-examination of the definitions as a failure.
- **Discriminator**: A real violation is inventing the grouping/labeling/ID scheme when a template or dataset column dictates it, and/or shipping an empty or unvalidated file. It is fine if the template genuinely only fixes headers and the agent documents a principled, data-derived grouping, verifies the schema, and the empty/short result is corroborated by an explicit sanity check showing no candidates exist under any reasonable definition.
- **Consequence**: The saved file's rows (or lack of rows) don't match the expected keys/values, so the file check fails outright — 0/1 — even though the narrative summary looks internally consistent.
571Ignoring a stated procedural cue (random seed / provided output template) that dictates the required method and formattaskda-code
Applies when
task -- the prompt fixes a random seed and/or points to a sample output file, implying a specific resampling-based (simulation) procedure and an exact result schema.
Pattern
The agent substitutes a deterministic closed-form/library statistic (e.g., an analytic parametric test) for the simulation-based estimate the seed implies, reports the statistic only in prose, and writes a result file whose columns/values were never checked against the provided sample file — so the seed is mentioned but is functionally irrelevant to the number produced.
Detection procedure
  1. Read the task for procedural constraints: a fixed seed, a named template/sample output file, rounding/units/ordering, and the exact quantity requested.
  2. In the scripts, check whether any randomness (permutation, bootstrap, simulation) actually feeds the reported number; if the computation is fully deterministic while a seed was mandated, the method is likely wrong.
  3. Check that the script reads/inspects the sample output file (or otherwise reproduces its exact column names, row count, and value formatting) before writing the result file.
  4. Compare the reported value and file contents against the requested quantity — is it the requested statistic, in the requested schema, or an intermediate/alternative statistic reported in free text?
Discriminator
A real violation is when setting the seed could not change the output at all (no sampling step) or when the written file's schema was never verified against the template; it is not a violation if the agent runs a seeded resampling procedure and additionally quotes an analytic value as a cross-check, or if the deterministic method is explicitly the one the task names.
Consequence
The saved result.csv fails the exact-match/schema check against the expected file (wrong p-value from the wrong test and/or wrong columns), so the task scores 0 even though the prose narrative looks coherent.
id 5546daf47254 · mined from da-code dacode-data-sa-039@s10
raw text (what the judge reads)
### Ignoring a stated procedural cue (random seed / provided output template) that dictates the required method and format
- **Applies when**: `task` -- the prompt fixes a random seed and/or points to a sample output file, implying a specific resampling-based (simulation) procedure and an exact result schema.
- **Pattern**: The agent substitutes a deterministic closed-form/library statistic (e.g., an analytic parametric test) for the simulation-based estimate the seed implies, reports the statistic only in prose, and writes a result file whose columns/values were never checked against the provided sample file — so the seed is mentioned but is functionally irrelevant to the number produced.
- **Detection procedure**:
  1. Read the task for procedural constraints: a fixed seed, a named template/sample output file, rounding/units/ordering, and the exact quantity requested.
  2. In the scripts, check whether any randomness (permutation, bootstrap, simulation) actually feeds the reported number; if the computation is fully deterministic while a seed was mandated, the method is likely wrong.
  3. Check that the script reads/inspects the sample output file (or otherwise reproduces its exact column names, row count, and value formatting) before writing the result file.
  4. Compare the reported value and file contents against the requested quantity — is it the requested statistic, in the requested schema, or an intermediate/alternative statistic reported in free text?
- **Discriminator**: A real violation is when setting the seed could not change the output at all (no sampling step) or when the written file's schema was never verified against the template; it is *not* a violation if the agent runs a seeded resampling procedure and additionally quotes an analytic value as a cross-check, or if the deterministic method is explicitly the one the task names.
- **Consequence**: The saved `result.csv` fails the exact-match/schema check against the expected file (wrong p-value from the wrong test and/or wrong columns), so the task scores 0 even though the prose narrative looks coherent.
572No reproducible script and no held-out validation of predictive qualitytaskda-code
Applies when
task -- the task asks for predictions on a provided test file to be written to an output file, and the agent's deliverable is the output file plus a prose summary.
Pattern
The attempt reports only pipeline metadata and descriptive statistics of the predicted values (min/max/mean/std, row count, column name) without any saved/runnable script and without any error estimate from a held-out split or cross-validation, so nothing demonstrates the model is better than a trivial baseline or that predictions were generated from the provided test rows in their original order/alignment.
Detection procedure
  1. Read the task and note that grading depends on prediction accuracy against hidden labels, not on file formatting alone.
  2. Look for a script that is saved and re-runnable end to end, and check it splits off a validation set (or uses CV) and prints a quantitative error metric plus a comparison to a naive baseline (mean/median predictor).
  3. Check the script/answer for evidence that predicted rows correspond 1:1 and in order to the given test file (row count equals test row count, no shuffling/dropna/filtering applied to the test frame, features engineered identically to training).
  4. Check whether the answer's prediction distribution is compared to the training target distribution (range, mean, spread); accept only if such a comparison is explicitly made.
Discriminator
A real violation is when the only evidence of correctness is self-reported summary statistics of the predictions or the file format; a look-alike that is fine reports a validation/CV score against a baseline, confirms test row alignment and identical feature construction, and keeps the code available even if the prose summary is brief.
Consequence
The submitted output file may have the right shape and column name but poorly aligned or low-accuracy values, so accuracy-threshold checks on the expected file fail and the attempt is marked wrong with no way to diagnose or reproduce it.
id 29d2553577a5 · mined from da-code dacode-ml-regression-015@s10
raw text (what the judge reads)
### No reproducible script and no held-out validation of predictive quality
- **Applies when**: `task` -- the task asks for predictions on a provided test file to be written to an output file, and the agent's deliverable is the output file plus a prose summary.
- **Pattern**: The attempt reports only pipeline metadata and descriptive statistics of the predicted values (min/max/mean/std, row count, column name) without any saved/runnable script and without any error estimate from a held-out split or cross-validation, so nothing demonstrates the model is better than a trivial baseline or that predictions were generated from the provided test rows in their original order/alignment.
- **Detection procedure**:
  1. Read the task and note that grading depends on prediction accuracy against hidden labels, not on file formatting alone.
  2. Look for a script that is saved and re-runnable end to end, and check it splits off a validation set (or uses CV) and prints a quantitative error metric plus a comparison to a naive baseline (mean/median predictor).
  3. Check the script/answer for evidence that predicted rows correspond 1:1 and in order to the given test file (row count equals test row count, no shuffling/dropna/filtering applied to the test frame, features engineered identically to training).
  4. Check whether the answer's prediction distribution is compared to the training target distribution (range, mean, spread); accept only if such a comparison is explicitly made.
- **Discriminator**: A real violation is when the only evidence of correctness is self-reported summary statistics of the predictions or the file format; a look-alike that is fine reports a validation/CV score against a baseline, confirms test row alignment and identical feature construction, and keeps the code available even if the prose summary is brief.
- **Consequence**: The submitted output file may have the right shape and column name but poorly aligned or low-accuracy values, so accuracy-threshold checks on the expected file fail and the attempt is marked wrong with no way to diagnose or reproduce it.
573Unjustified row exclusion (“outlier”/cleaning filter) before computing a simple aggregatetaskinfiagent-dabench
Applies when
task -- the task asks for a straightforward summary statistic of one column and mentions cleaning language (e.g., ignore missing values/outliers) without defining a filtering rule.
Pattern
The attempt invents a filtering rule (IQR/z-score/percentile cut, or dropping codes it deems invalid) and reports the statistic of the surviving subset, without checking whether the column is a coded/categorical/bounded field where such filtering is meaningless, and without reporting or comparing the unfiltered value. Often no script is preserved, so the exclusion is invisible and unreproducible.
Detection procedure
  1. Read the task and note whether a concrete filtering threshold is specified; if not, the default expectation is all non-missing observations.
  2. Read the scripts for any drop, boolean mask, quantile/std-based cut, or dtype coercion that silently removes rows; check the row count before and after.
  3. Check the column's nature (small set of integer codes, ID, ordinal) — if values are a bounded discrete code set, any "outlier" removal is illegitimate.
  4. Confirm the reported number equals the statistic over all non-missing rows; if the script never computes/prints both filtered and unfiltered values (and row counts), treat the result as unverified.
Discriminator
A real violation removes valid in-range observations under a self-invented rule, or cannot show the unfiltered value; it is fine to drop rows only for explicitly stated criteria (nulls, a threshold given in the task) or for provably invalid entries (non-numeric junk, sentinel codes), with counts shown.
Consequence
The reported mean is biased low/high relative to the ground-truth full-column mean and fails exact/tolerance matching by the grader.
id b339581f02d2 · mined from infiagent-dabench dabench-320@s10
raw text (what the judge reads)
### Unjustified row exclusion (“outlier”/cleaning filter) before computing a simple aggregate
- **Applies when**: `task` -- the task asks for a straightforward summary statistic of one column and mentions cleaning language (e.g., ignore missing values/outliers) without defining a filtering rule.
- **Pattern**: The attempt invents a filtering rule (IQR/z-score/percentile cut, or dropping codes it deems invalid) and reports the statistic of the surviving subset, without checking whether the column is a coded/categorical/bounded field where such filtering is meaningless, and without reporting or comparing the unfiltered value. Often no script is preserved, so the exclusion is invisible and unreproducible.
- **Detection procedure**:
  1. Read the task and note whether a concrete filtering threshold is specified; if not, the default expectation is all non-missing observations.
  2. Read the scripts for any `drop`, boolean mask, quantile/std-based cut, or dtype coercion that silently removes rows; check the row count before and after.
  3. Check the column's nature (small set of integer codes, ID, ordinal) — if values are a bounded discrete code set, any "outlier" removal is illegitimate.
  4. Confirm the reported number equals the statistic over all non-missing rows; if the script never computes/prints both filtered and unfiltered values (and row counts), treat the result as unverified.
- **Discriminator**: A real violation removes valid in-range observations under a self-invented rule, or cannot show the unfiltered value; it is fine to drop rows only for explicitly stated criteria (nulls, a threshold given in the task) or for provably invalid entries (non-numeric junk, sentinel codes), with counts shown.
- **Consequence**: The reported mean is biased low/high relative to the ground-truth full-column mean and fails exact/tolerance matching by the grader.
574Building on a convenient pre-derived intermediate file without validating it against the raw sourcetaskda-code
Applies when
task -- the task asks for metrics/scores/segments to be computed from the provided data, and the working directory contains both raw records and a pre-aggregated or partially processed file that already holds the needed quantities.
Pattern
The script loads the pre-derived file directly, inherits whatever filtering, time window, aggregation rule and entity set that file was built with, and never recomputes or cross-checks the quantities from the raw records; downstream bucket cutoffs and label names are then invented ad hoc on top of an unverified base.
Detection procedure
  1. Read the task and note which quantities are supposed to be derived and over which population/time window.
  2. Inspect the script's input path(s): is it the raw source, or a derived artifact? Check whether the script ever verifies the artifact (entity count, date range, aggregation definition) against the raw data.
  3. Check the answer for a population/row count and compare it to the entity count obtainable from the raw source; a discrepancy that is never explained is a red flag.
  4. Check that any discretization thresholds/label sets are derived from the data or a stated convention, not hand-picked constants with no justification.
Discriminator
Fine if the derived file is the only or explicitly designated input, or if the script confirms it reproduces the raw-data aggregates (matching entity counts and definitions). A violation is silently accepting a derived subset/definition as the analysis base while the task implies the full raw-data-derived population.
Consequence
The saved output has the wrong number of rows and shifted quantile boundaries, so per-entity scores and segment labels differ from the expected file and the file-comparison check fails outright.
id 3b944f61601b · mined from da-code dacode-dm-csv-052@s10
raw text (what the judge reads)
### Building on a convenient pre-derived intermediate file without validating it against the raw source
- **Applies when**: `task` -- the task asks for metrics/scores/segments to be *computed* from the provided data, and the working directory contains both raw records and a pre-aggregated or partially processed file that already holds the needed quantities.
- **Pattern**: The script loads the pre-derived file directly, inherits whatever filtering, time window, aggregation rule and entity set that file was built with, and never recomputes or cross-checks the quantities from the raw records; downstream bucket cutoffs and label names are then invented ad hoc on top of an unverified base.
- **Detection procedure**:
  1. Read the task and note which quantities are supposed to be *derived* and over which population/time window.
  2. Inspect the script's input path(s): is it the raw source, or a derived artifact? Check whether the script ever verifies the artifact (entity count, date range, aggregation definition) against the raw data.
  3. Check the answer for a population/row count and compare it to the entity count obtainable from the raw source; a discrepancy that is never explained is a red flag.
  4. Check that any discretization thresholds/label sets are derived from the data or a stated convention, not hand-picked constants with no justification.
- **Discriminator**: Fine if the derived file is the only or explicitly designated input, or if the script confirms it reproduces the raw-data aggregates (matching entity counts and definitions). A violation is silently accepting a derived subset/definition as the analysis base while the task implies the full raw-data-derived population.
- **Consequence**: The saved output has the wrong number of rows and shifted quantile boundaries, so per-entity scores and segment labels differ from the expected file and the file-comparison check fails outright.
575Silently subsampling the dataset so the output covers only part of the required rowstaskda-code
Applies when
task -- the task asks for a per-record output (labels, predictions, cluster assignments) written to a file for a dataset whose full size is stated or discoverable from the source file.
Pattern
The agent loads only a truncated slice of the data (e.g., nrows=, head(n), sample(n), an early-exit read loop, or a "for speed" cap) and then writes the result file containing just those rows, while the report presents it as the result for the whole dataset without flagging the reduction.
Detection procedure
  1. From the task/README, note the expected record count and feature count of the source data (e.g., "over half a million measurements").
  2. Scan the scripts for any row-limiting argument or slicing between loading and fitting/writing (nrows, usecols, .iloc[:N], .sample, random subset, chunk read that keeps only one chunk).
  3. Compare the row count the agent reports for the written file against the source row count; also check the column names/count match the requested schema.
  4. If they differ, check whether the task explicitly permitted subsampling or whether the script re-applies the fitted model to all rows before writing.
Discriminator
Fine: subsampling used only for hyperparameter search / diagnostics (e.g., choosing k, tuning), with the final model applied to and exported for every record. Violation: the exported per-record file itself has fewer rows than the source data, or the row count is never verified against the source.
Consequence
The result file has the wrong shape/row coverage, so row-wise comparison against the expected output fails outright — the grader marks the file WRONG/MISSING regardless of clustering quality.
id 611bfb0595eb · mined from da-code dacode-ml-cluster-010@s10
raw text (what the judge reads)
### Silently subsampling the dataset so the output covers only part of the required rows
- **Applies when**: `task` -- the task asks for a per-record output (labels, predictions, cluster assignments) written to a file for a dataset whose full size is stated or discoverable from the source file.
- **Pattern**: The agent loads only a truncated slice of the data (e.g., `nrows=`, `head(n)`, `sample(n)`, an early-exit read loop, or a "for speed" cap) and then writes the result file containing just those rows, while the report presents it as the result for the whole dataset without flagging the reduction.
- **Detection procedure**:
  1. From the task/README, note the expected record count and feature count of the source data (e.g., "over half a million measurements").
  2. Scan the scripts for any row-limiting argument or slicing between loading and fitting/writing (`nrows`, `usecols`, `.iloc[:N]`, `.sample`, random subset, chunk read that keeps only one chunk).
  3. Compare the row count the agent reports for the written file against the source row count; also check the column names/count match the requested schema.
  4. If they differ, check whether the task explicitly permitted subsampling or whether the script re-applies the fitted model to all rows before writing.
- **Discriminator**: Fine: subsampling used only for hyperparameter search / diagnostics (e.g., choosing k, tuning), with the final model applied to and exported for every record. Violation: the exported per-record file itself has fewer rows than the source data, or the row count is never verified against the source.
- **Consequence**: The result file has the wrong shape/row coverage, so row-wise comparison against the expected output fails outright — the grader marks the file WRONG/MISSING regardless of clustering quality.
576Deliverable artifact never written in the requested file/formattaskda-code
Applies when
task -- The task explicitly names an output file (e.g. result.csv) and/or a template/format the findings must be entered into.
Pattern
The agent computes plausible numbers and reports them only in prose/console output, never creating the named file (or creating it with different column names, row order, units, or extra narrative text), so the graded artifact is missing or unparseable even if the analysis is sound.
Detection procedure
  1. Read the task statement and list every required output artifact: exact filename, expected columns/keys, units, and any rounding/ordering constraints.
  2. Scan the scripts for a write step (to_csv, open(...,'w'), etc.) targeting that exact filename in the expected directory, and check that the written columns/labels match the provided template rather than ad-hoc names.
  3. Check that the values written are the quantity actually requested (correct unit/scale — e.g. proportion vs count, difference vs ratio) and match the numbers quoted in the final answer.
  4. If the final answer is only a narrative summary with no file creation or no confirmation that the file exists and re-reads correctly, flag it.
Discriminator
A real violation is a missing file, a mismatched filename/schema, or written values inconsistent with the requested quantity; a look-alike that is fine is an agent that writes the correct file and additionally summarizes the numbers in prose.
Consequence
The grader looks for the named file and reports it as WRONG/MISSING, scoring 0 regardless of how reasonable the reported statistics look.
id 106f08d365f2 · mined from da-code dacode-data-sa-031@s10
raw text (what the judge reads)
### Deliverable artifact never written in the requested file/format
- **Applies when**: `task` -- The task explicitly names an output file (e.g. `result.csv`) and/or a template/format the findings must be entered into.
- **Pattern**: The agent computes plausible numbers and reports them only in prose/console output, never creating the named file (or creating it with different column names, row order, units, or extra narrative text), so the graded artifact is missing or unparseable even if the analysis is sound.
- **Detection procedure**:
  1. Read the task statement and list every required output artifact: exact filename, expected columns/keys, units, and any rounding/ordering constraints.
  2. Scan the scripts for a write step (`to_csv`, `open(...,'w')`, etc.) targeting that exact filename in the expected directory, and check that the written columns/labels match the provided template rather than ad-hoc names.
  3. Check that the values written are the quantity actually requested (correct unit/scale — e.g. proportion vs count, difference vs ratio) and match the numbers quoted in the final answer.
  4. If the final answer is only a narrative summary with no file creation or no confirmation that the file exists and re-reads correctly, flag it.
- **Discriminator**: A real violation is a missing file, a mismatched filename/schema, or written values inconsistent with the requested quantity; a look-alike that is fine is an agent that writes the correct file and *additionally* summarizes the numbers in prose.
- **Consequence**: The grader looks for the named file and reports it as WRONG/MISSING, scoring 0 regardless of how reasonable the reported statistics look.
577No out-of-sample estimate of the competition's own metric before submittingtaskda-code
Applies when
task -- the task names a specific scoring metric (e.g., a probabilistic or ranked loss) and the scripts fit a model and write a prediction file directly.
Pattern
The scripts fit one default/near-default estimator on the full training set, immediately call predict_proba/predict on the test set, and write the submission; the only "verification" is a format/shape check (columns, NaNs, hardcoded row count, min/max) with no cross-validated or holdout computation of the stated metric, no comparison against a trivial baseline (class priors), and no evidence that preprocessing choices (missing-value handling, categorical encoding, unused/leaky columns) actually help. Model/hyperparameter choices are therefore made blind.
Detection procedure
  1. Read the task and note the exact evaluation metric and any probability/format constraints.
  2. Search the scripts for that metric being computed (e.g., log_loss/cross_val_score/cross_val_predict on a held-out split or CV folds); note whether any baseline comparison exists.
  3. Check whether preprocessing that the data demands (missing values, categorical levels present only in test, dropping identifier columns) is validated by that score rather than assumed.
  4. Inspect the final answer/verification step: if it only reports shapes, column names, and value ranges, and no numeric score is ever reported, the attempt is unvalidated.
Discriminator
A genuine violation is when no out-of-sample score in the required metric is ever produced, so a mis-specified preprocessing step or a poorly-calibrated model cannot be detected. It is not a violation if the scripts report a CV/holdout score in the stated metric (even from a single model) and use it to compare at least one alternative or baseline; small deviations in split strategy are acceptable.
Consequence
The submitted file is syntactically valid but scores worse than the grading threshold (often worse than a simple prior/baseline), so the grader marks the expected output file WRONG despite the format checks passing.
id ef08b5396e4e · mined from da-code dacode-ml-competition-005@s11
raw text (what the judge reads)
### No out-of-sample estimate of the competition's own metric before submitting
- **Applies when**: `task` -- the task names a specific scoring metric (e.g., a probabilistic or ranked loss) and the scripts fit a model and write a prediction file directly.
- **Pattern**: The scripts fit one default/near-default estimator on the full training set, immediately call `predict_proba`/`predict` on the test set, and write the submission; the only "verification" is a format/shape check (columns, NaNs, hardcoded row count, min/max) with no cross-validated or holdout computation of the stated metric, no comparison against a trivial baseline (class priors), and no evidence that preprocessing choices (missing-value handling, categorical encoding, unused/leaky columns) actually help. Model/hyperparameter choices are therefore made blind.
- **Detection procedure**:
  1. Read the task and note the exact evaluation metric and any probability/format constraints.
  2. Search the scripts for that metric being computed (e.g., `log_loss`/`cross_val_score`/`cross_val_predict` on a held-out split or CV folds); note whether any baseline comparison exists.
  3. Check whether preprocessing that the data demands (missing values, categorical levels present only in test, dropping identifier columns) is validated by that score rather than assumed.
  4. Inspect the final answer/verification step: if it only reports shapes, column names, and value ranges, and no numeric score is ever reported, the attempt is unvalidated.
- **Discriminator**: A genuine violation is when *no* out-of-sample score in the required metric is ever produced, so a mis-specified preprocessing step or a poorly-calibrated model cannot be detected. It is *not* a violation if the scripts report a CV/holdout score in the stated metric (even from a single model) and use it to compare at least one alternative or baseline; small deviations in split strategy are acceptable.
- **Consequence**: The submitted file is syntactically valid but scores worse than the grading threshold (often worse than a simple prior/baseline), so the grader marks the expected output file WRONG despite the format checks passing.
578Model shipped with no held-out validation and no prediction-distribution sanity checktaskda-code
Applies when
task -- the scripts fit a supervised model on labeled data and write predictions for an unlabeled set straight to the required output file.
Pattern
The attempt trains on 100% of the labeled rows, never computes an error metric on a held-out split (or CV), and never compares the resulting prediction distribution (min/max/mean/median/spread, integer-vs-float, plausible range) against the distribution of the training target; it only prints summary numbers without asserting they are consistent. Typical symptom: predictions that are far more compressed toward zero (or otherwise scaled differently) than the observed target, and fractional values where the quantity is inherently a count.
Detection procedure
  1. Read the task to note the nature/units of the quantity to predict (count, currency, bounded score) and any format expectation.
  2. Scan the scripts for any train/validation split, cross-validation, or metric computation on labeled data — if the only .fit() uses all rows and no score is ever computed, flag it.
  3. Compare the printed/held training-target summary statistics with the prediction summary statistics: check that the mean/median/percentile range and dtype are of the same order and type; a heavily-skewed target predicted as a narrow near-zero band, or continuous decimals for a count, is a red flag.
  4. Check that no extra sanity assertion (row count equals test rows, non-negativity, plausible max) is missing.
Discriminator
A fine attempt may also train on all data after it has validated the pipeline on a split (or reports CV error) and explicitly reconciles prediction statistics with target statistics; the violation is the total absence of any quantitative check, not merely a low-variance prediction that has been justified by validation error and target skew.
Consequence
The graded comparison against reference values shows large error (predictions systematically off in scale/type), so the output file is marked wrong even though it has the right column name and shape.
id 3202e0ebe489 · mined from da-code dacode-ml-regression-008@s11
raw text (what the judge reads)
### Model shipped with no held-out validation and no prediction-distribution sanity check
- **Applies when**: `task` -- the scripts fit a supervised model on labeled data and write predictions for an unlabeled set straight to the required output file.
- **Pattern**: The attempt trains on 100% of the labeled rows, never computes an error metric on a held-out split (or CV), and never compares the resulting prediction distribution (min/max/mean/median/spread, integer-vs-float, plausible range) against the distribution of the training target; it only prints summary numbers without asserting they are consistent. Typical symptom: predictions that are far more compressed toward zero (or otherwise scaled differently) than the observed target, and fractional values where the quantity is inherently a count.
- **Detection procedure**:
  1. Read the task to note the nature/units of the quantity to predict (count, currency, bounded score) and any format expectation.
  2. Scan the scripts for any train/validation split, cross-validation, or metric computation on labeled data — if the only `.fit()` uses all rows and no score is ever computed, flag it.
  3. Compare the printed/held training-target summary statistics with the prediction summary statistics: check that the mean/median/percentile range and dtype are of the same order and type; a heavily-skewed target predicted as a narrow near-zero band, or continuous decimals for a count, is a red flag.
  4. Check that no extra sanity assertion (row count equals test rows, non-negativity, plausible max) is missing.
- **Discriminator**: A fine attempt may also train on all data *after* it has validated the pipeline on a split (or reports CV error) and explicitly reconciles prediction statistics with target statistics; the violation is the total absence of any quantitative check, not merely a low-variance prediction that has been justified by validation error and target skew.
- **Consequence**: The graded comparison against reference values shows large error (predictions systematically off in scale/type), so the output file is marked wrong even though it has the right column name and shape.
579Uses the full dataset and a default two-sided parametric test without justifying population scope, test type, or tail directiontaskda-code
Applies when
task -- the task asks for a p-value and an accept/reject decision comparing two groups drawn from long-running, heterogeneous historical records.
Pattern
The script loads every row of both files, computes a derived quantity, and immediately calls the default parametric two-sample test (equal-variance, two-sided) on the whole population — never restricting to the population implied by the question (e.g., a competition/time window or other stated filter), never checking that the derived quantity's distribution justifies a mean-based parametric test (small integer/skewed counts), and never matching the tail of the test to the direction of the stated alternative hypothesis. Because n is huge, the reported p-value is astronomically small, which the agent treats as confirmation instead of as a red flag.
Detection procedure
  1. Read the task statement (and README) for any qualifier that narrows the units of analysis — a subset, date range, category, or "recent"/"official"/specific-competition wording — and for whether the alternative is directional ("greater than", "more than") or merely "different".
  2. In the script, check for a filtering step matching each qualifier and check the test call's parameters: one-sided vs two-sided, equal-variance vs Welch, parametric vs rank-based; also check whether the distribution of the tested quantity was inspected (histogram/skew) before choosing a mean-based test.
  3. Compare the printed sample sizes to the raw row counts: if they are identical, no subsetting occurred; if the p-value is many orders of magnitude below any plausible threshold, treat that as a symptom of testing an unintended, much larger population.
  4. Confirm the reported p-value is the one from the correctly scoped, correctly tailed test, and that the decision string and column names match the requested format exactly.
Discriminator
A real violation is when the task text (or accompanying README/context) implies a narrower population or a directional/nonparametric test and the script silently uses everything with library defaults. It is not a violation if the task genuinely asks about the entire dataset with a two-sided mean comparison and the script documents that choice (e.g., checks distribution/variance and explains why the default test applies).
Consequence
The p-value differs by many orders of magnitude from the reference (and can flip the reject/fail-to-reject decision), so the saved result file fails the grader's value check even though the file format is correct.
id de9b99f3deef · mined from da-code dacode-data-sa-001@s11
raw text (what the judge reads)
### Uses the full dataset and a default two-sided parametric test without justifying population scope, test type, or tail direction
- **Applies when**: `task` -- the task asks for a p-value and an accept/reject decision comparing two groups drawn from long-running, heterogeneous historical records.
- **Pattern**: The script loads every row of both files, computes a derived quantity, and immediately calls the default parametric two-sample test (equal-variance, two-sided) on the whole population — never restricting to the population implied by the question (e.g., a competition/time window or other stated filter), never checking that the derived quantity's distribution justifies a mean-based parametric test (small integer/skewed counts), and never matching the tail of the test to the direction of the stated alternative hypothesis. Because n is huge, the reported p-value is astronomically small, which the agent treats as confirmation instead of as a red flag.
- **Detection procedure**:
  1. Read the task statement (and README) for any qualifier that narrows the units of analysis — a subset, date range, category, or "recent"/"official"/specific-competition wording — and for whether the alternative is directional ("greater than", "more than") or merely "different".
  2. In the script, check for a filtering step matching each qualifier and check the test call's parameters: one-sided vs two-sided, equal-variance vs Welch, parametric vs rank-based; also check whether the distribution of the tested quantity was inspected (histogram/skew) before choosing a mean-based test.
  3. Compare the printed sample sizes to the raw row counts: if they are identical, no subsetting occurred; if the p-value is many orders of magnitude below any plausible threshold, treat that as a symptom of testing an unintended, much larger population.
  4. Confirm the reported p-value is the one from the correctly scoped, correctly tailed test, and that the decision string and column names match the requested format exactly.
- **Discriminator**: A real violation is when the task text (or accompanying README/context) implies a narrower population or a directional/nonparametric test and the script silently uses everything with library defaults. It is *not* a violation if the task genuinely asks about the entire dataset with a two-sided mean comparison and the script documents that choice (e.g., checks distribution/variance and explains why the default test applies).
- **Consequence**: The p-value differs by many orders of magnitude from the reference (and can flip the reject/fail-to-reject decision), so the saved result file fails the grader's value check even though the file format is correct.
580Template file provided but never programmatically read — output schema guessedtaskda-code
Applies when
task -- The task says results must follow the exact structure/formatting of a provided sample/template output file.
Pattern
The scripts never load or inspect the template; column names, column order, row order, header text, rounding/units are invented by the agent (sometimes hard-coded from a remembered or hand-typed list), and the final file is written without any comparison against the template.
Detection procedure
  1. In the task statement, note that a sample/template output artifact is named as the formatting spec.
  2. Search the scripts for any read of that template file (e.g., loading it and printing its header/rows); if absent, the schema is being guessed.
  3. Check whether the produced columns, their order, row ordering key, and value formatting (rounding, integer vs float, units, label spellings) are derived from the template or from ad-hoc literals/assumptions in the code.
  4. Confirm the script has no final assertion comparing produced header/row count/dtypes to the template.
Discriminator
A real violation is guessing or hard-coding the schema with no read of the template; it is fine if the script loads the template (or an explicit, quoted copy of its header) and derives/validates column names, order, and formatting from it — even if it then reorders rows for a stated reason.
Consequence
The output file fails exact-match grading on header names, column/row order, or numeric formatting even when the underlying aggregation logic is right, so the single file check scores 0.
id 5d12f8908bf8 · mined from da-code dacode-dm-csv-011@s11
raw text (what the judge reads)
### Template file provided but never programmatically read — output schema guessed
- **Applies when**: `task` -- The task says results must follow the exact structure/formatting of a provided sample/template output file.
- **Pattern**: The scripts never load or inspect the template; column names, column order, row order, header text, rounding/units are invented by the agent (sometimes hard-coded from a remembered or hand-typed list), and the final file is written without any comparison against the template.
- **Detection procedure**:
  1. In the task statement, note that a sample/template output artifact is named as the formatting spec.
  2. Search the scripts for any read of that template file (e.g., loading it and printing its header/rows); if absent, the schema is being guessed.
  3. Check whether the produced columns, their order, row ordering key, and value formatting (rounding, integer vs float, units, label spellings) are derived from the template or from ad-hoc literals/assumptions in the code.
  4. Confirm the script has no final assertion comparing produced header/row count/dtypes to the template.
- **Discriminator**: A real violation is guessing or hard-coding the schema with no read of the template; it is fine if the script loads the template (or an explicit, quoted copy of its header) and derives/validates column names, order, and formatting from it — even if it then reorders rows for a stated reason.
- **Consequence**: The output file fails exact-match grading on header names, column/row order, or numeric formatting even when the underlying aggregation logic is right, so the single file check scores 0.
581Deliverable written straight from an unvalidated, auto-selected clustering (trivial k / transformed feature columns)taskda-code
Applies when
task -- the task asks for an unsupervised grouping to be saved to a file whose columns must hold the feature vector plus a label, and the script picks the number of groups automatically from an internal score and dumps whatever it gets.
Pattern
The script maximizes a single internal index (e.g., silhouette) over a high-dimensional, redundantly encoded, standardized matrix, which almost always favors the smallest candidate k, then writes that near-trivial two-way split — together with scaled/derived columns that are not the dataset's feature vector (duplicated or engineered fields, arbitrary one-hot expansion) — to the required file with no sanity check on cluster sizes, separation, or column correspondence.
Detection procedure
  1. Read the task and note exactly what the output columns are supposed to contain (the feature values used for clustering, in the expected count/order, plus the label) and whether "distinct groups" implies a non-degenerate segmentation.
  2. In the script, list the columns fed to the model and compare against what is written out: check for redundant/derived duplicates of the same information, leakage of an identifier, and whether the written values are the transformed matrix rather than the features implied by the task.
  3. Check how k is chosen: is it picked solely by argmax of one index over a range starting at 2, with no elbow/stability/interpretability cross-check and no rejection of degenerate solutions?
  4. Check the answer/report for any post-hoc validation — cluster size balance, per-cluster profile differences, agreement across k or across algorithms; absence of all of these plus a chosen k at the range boundary is the flag.
Discriminator
A real violation is a boundary-value k accepted with no corroborating evidence and an output table whose columns don't map one-to-one to a defensible feature vector; it is not a violation if k=2 (or any k) is defended with at least one independent check (elbow, stability, distinct cluster profiles) and the saved columns clearly correspond to the features actually used, consistently named/ordered.
Consequence
The saved file fails the grader's structural/content comparison (wrong number or meaning of Feature_i columns) and/or yields a degenerate segmentation that does not match the expected cluster structure, so the single file check scores 0.
id 9362494c7bd8 · mined from da-code dacode-ml-cluster-014@s11
raw text (what the judge reads)
### Deliverable written straight from an unvalidated, auto-selected clustering (trivial k / transformed feature columns)

- **Applies when**: `task` -- the task asks for an unsupervised grouping to be saved to a file whose columns must hold the feature vector plus a label, and the script picks the number of groups automatically from an internal score and dumps whatever it gets.
- **Pattern**: The script maximizes a single internal index (e.g., silhouette) over a high-dimensional, redundantly encoded, standardized matrix, which almost always favors the smallest candidate k, then writes that near-trivial two-way split — together with scaled/derived columns that are not the dataset's feature vector (duplicated or engineered fields, arbitrary one-hot expansion) — to the required file with no sanity check on cluster sizes, separation, or column correspondence.
- **Detection procedure**:
  1. Read the task and note exactly what the output columns are supposed to contain (the feature values used for clustering, in the expected count/order, plus the label) and whether "distinct groups" implies a non-degenerate segmentation.
  2. In the script, list the columns fed to the model and compare against what is written out: check for redundant/derived duplicates of the same information, leakage of an identifier, and whether the written values are the transformed matrix rather than the features implied by the task.
  3. Check how k is chosen: is it picked solely by argmax of one index over a range starting at 2, with no elbow/stability/interpretability cross-check and no rejection of degenerate solutions?
  4. Check the answer/report for any post-hoc validation — cluster size balance, per-cluster profile differences, agreement across k or across algorithms; absence of all of these plus a chosen k at the range boundary is the flag.
- **Discriminator**: A real violation is a boundary-value k accepted with no corroborating evidence and an output table whose columns don't map one-to-one to a defensible feature vector; it is *not* a violation if k=2 (or any k) is defended with at least one independent check (elbow, stability, distinct cluster profiles) and the saved columns clearly correspond to the features actually used, consistently named/ordered.
- **Consequence**: The saved file fails the grader's structural/content comparison (wrong number or meaning of `Feature_i` columns) and/or yields a degenerate segmentation that does not match the expected cluster structure, so the single file check scores 0.
582Model/post-processing choices justified only by in-sample (training-set) metricstaskda-code
Applies when
task -- the scripts fit one or more supervised models, compare variants, and/or apply post-hoc transformations to predictions (rounding, clipping, ensembling weights) before writing the submission file.
Pattern
Every reported score is computed by predicting on the same rows used for fitting (no train/validation split, no cross-validation, no out-of-fold estimate), so high-capacity settings and arbitrary prediction post-processing look fine; the attempt then picks the final pipeline and prediction format on the basis of these meaningless numbers, and never checks the choice against the task's stated evaluation target or the sample output's value type.
Detection procedure
  1. Read the task/README to note the evaluation target and the exact value type/format shown in the provided sample output file.
  2. Scan each script for where metrics are computed: check whether the data passed to predict/score is the same object used in fit, and whether any train_test_split/cross_val_*/out-of-fold logic exists.
  3. Identify each decision that depends on a score (model family, depth/estimators, ensemble weights, rounding/clipping of outputs) and confirm whether any held-out estimate supports it; also compare the written column dtype/range to the sample submission's.
  4. Flag if all reported scores are in-sample and/or a value transformation (e.g., discretizing continuous predictions, clipping to observed range) is applied with no held-out comparison to the untransformed predictions.
Discriminator
A real violation is when no honest generalization estimate exists anywhere in the pipeline, or when an irreversible output transformation is applied purely by intuition; it is fine if in-sample numbers are printed merely as diagnostics while model/post-processing selection is driven by cross-validation or a held-out split, or if the transformation is explicitly required by the task/sample format.
Consequence
The submitted file is produced by an overfit, unvalidated pipeline (and possibly the wrong value granularity), so the hidden-metric score falls below the acceptance threshold and the graded submission is marked wrong even though the file's shape and ids look valid.
id 30f7b1069a33 · mined from da-code dacode-ml-competition-009@s11
raw text (what the judge reads)
### Model/post-processing choices justified only by in-sample (training-set) metrics
- **Applies when**: `task` -- the scripts fit one or more supervised models, compare variants, and/or apply post-hoc transformations to predictions (rounding, clipping, ensembling weights) before writing the submission file.
- **Pattern**: Every reported score is computed by predicting on the same rows used for fitting (no train/validation split, no cross-validation, no out-of-fold estimate), so high-capacity settings and arbitrary prediction post-processing look fine; the attempt then picks the final pipeline and prediction format on the basis of these meaningless numbers, and never checks the choice against the task's stated evaluation target or the sample output's value type.
- **Detection procedure**:
  1. Read the task/README to note the evaluation target and the exact value type/format shown in the provided sample output file.
  2. Scan each script for where metrics are computed: check whether the data passed to `predict`/`score` is the same object used in `fit`, and whether any `train_test_split`/`cross_val_*`/out-of-fold logic exists.
  3. Identify each decision that depends on a score (model family, depth/estimators, ensemble weights, rounding/clipping of outputs) and confirm whether any held-out estimate supports it; also compare the written column dtype/range to the sample submission's.
  4. Flag if all reported scores are in-sample and/or a value transformation (e.g., discretizing continuous predictions, clipping to observed range) is applied with no held-out comparison to the untransformed predictions.
- **Discriminator**: A real violation is when *no* honest generalization estimate exists anywhere in the pipeline, or when an irreversible output transformation is applied purely by intuition; it is fine if in-sample numbers are printed merely as diagnostics while model/post-processing selection is driven by cross-validation or a held-out split, or if the transformation is explicitly required by the task/sample format.
- **Consequence**: The submitted file is produced by an overfit, unvalidated pipeline (and possibly the wrong value granularity), so the hidden-metric score falls below the acceptance threshold and the graded submission is marked wrong even though the file's shape and ids look valid.
583Row-dropping that shrinks the required per-record outputtaskda-code
Applies when
task -- the task asks for a per-record output file (e.g., one label/prediction per input row) and the script handles missing values or filtering before producing it.
Pattern
The script calls a blanket dropna() (or similar filter) over all candidate feature columns, silently discarding a large fraction of records, then writes the result for only the surviving subset — while the answer reports the reduced count as if it were the expected deliverable.
Detection procedure
  1. Read the task/README to determine the expected number of output rows (normally the full number of input records) and the required columns/naming.
  2. In the script, find every step that removes rows (dropna, boolean masks, joins) and estimate how many rows survive; check whether the missingness is concentrated in a few noisy columns that could instead be imputed or excluded.
  3. Compare the printed/reported output row count against the input record count; also check that non-informative identifier-like columns (codes, indices, coordinates) were not fed in as features.
  4. Confirm the answer explicitly justifies any row loss against a stated task requirement; if it just notes "removed rows with missing values", flag it.
Discriminator
Legitimate cases are when the task explicitly asks for a filtered subset, or when only a negligible number of rows lack the target/key field; a violation is when a substantial share of records (here, >25%) is dropped merely for convenience and imputation or column pruning would have preserved them.
Consequence
The saved file has fewer rows than the reference output, so row-count/alignment checks fail and the file is graded WRONG/MISSING regardless of clustering quality.
id c54429db0c68 · mined from da-code dacode-ml-cluster-009@s11
raw text (what the judge reads)
### Row-dropping that shrinks the required per-record output
- **Applies when**: `task` -- the task asks for a per-record output file (e.g., one label/prediction per input row) and the script handles missing values or filtering before producing it.
- **Pattern**: The script calls a blanket `dropna()` (or similar filter) over all candidate feature columns, silently discarding a large fraction of records, then writes the result for only the surviving subset — while the answer reports the reduced count as if it were the expected deliverable.
- **Detection procedure**:
  1. Read the task/README to determine the expected number of output rows (normally the full number of input records) and the required columns/naming.
  2. In the script, find every step that removes rows (`dropna`, boolean masks, joins) and estimate how many rows survive; check whether the missingness is concentrated in a few noisy columns that could instead be imputed or excluded.
  3. Compare the printed/reported output row count against the input record count; also check that non-informative identifier-like columns (codes, indices, coordinates) were not fed in as features.
  4. Confirm the answer explicitly justifies any row loss against a stated task requirement; if it just notes "removed rows with missing values", flag it.
- **Discriminator**: Legitimate cases are when the task explicitly asks for a filtered subset, or when only a negligible number of rows lack the target/key field; a violation is when a substantial share of records (here, >25%) is dropped merely for convenience and imputation or column pruning would have preserved them.
- **Consequence**: The saved file has fewer rows than the reference output, so row-count/alignment checks fail and the file is graded WRONG/MISSING regardless of clustering quality.
584Degenerate cluster solution accepted because a validity index was maximized on heavily skewed, outlier-dominated featurestaskda-code
Applies when
task -- the task asks for an unsupervised segmentation into an "appropriate" number of groups and the script selects k by maximizing/minimizing an internal index (silhouette, Davies-Bouldin, etc.) on standardized but highly skewed count/amount features.
Pattern
The attempt aggregates raw, long-tailed variables, applies only z-scaling (no log/rank transform, no outlier capping or removal), then lets the index pick the smallest k, producing a near-degenerate partition where one cluster holds a handful of extreme points and another holds essentially the whole population; the suspiciously high index score is reported as evidence of "excellent" segmentation instead of triggering a re-check.
Detection procedure
  1. In the task statement, confirm the deliverable is a meaningful segmentation of all records, not just outlier detection.
  2. In the script, check whether skew is addressed before distance-based clustering (log/box-cox/robust scaling, winsorizing, or outlier screening) and whether any guard exists against trivial solutions (minimum cluster size, cluster-balance check, elbow/stability cross-check, not only the raw index optimum).
  3. In the printed/reported results, inspect cluster sizes and the chosen k: flag if the optimum is at the boundary of the searched range, if the index is implausibly high (e.g. >0.8 on real customer data), or if any cluster contains a negligible fraction (<1%) of rows.
  4. Verify the saved output rows/columns correspond to the intended units of analysis and that group labels are non-trivial across them.
Discriminator
A genuine violation shows one or two clusters absorbing almost all points with the rest being extreme outliers and no transformation/robustness step; it is fine if skew was handled (or outliers explicitly modeled) and the resulting segments are reasonably populated and stable, even if some segment is small by design.
Consequence
The saved segmentation is effectively an outlier flag rather than a customer segmentation, so cluster labels/count do not match the expected result file and the check fails.
id b891b09f4a7f · mined from da-code dacode-ml-cluster-016@s11
raw text (what the judge reads)
### Degenerate cluster solution accepted because a validity index was maximized on heavily skewed, outlier-dominated features
- **Applies when**: `task` -- the task asks for an unsupervised segmentation into an "appropriate" number of groups and the script selects k by maximizing/minimizing an internal index (silhouette, Davies-Bouldin, etc.) on standardized but highly skewed count/amount features.
- **Pattern**: The attempt aggregates raw, long-tailed variables, applies only z-scaling (no log/rank transform, no outlier capping or removal), then lets the index pick the smallest k, producing a near-degenerate partition where one cluster holds a handful of extreme points and another holds essentially the whole population; the suspiciously high index score is reported as evidence of "excellent" segmentation instead of triggering a re-check.
- **Detection procedure**:
  1. In the task statement, confirm the deliverable is a meaningful segmentation of all records, not just outlier detection.
  2. In the script, check whether skew is addressed before distance-based clustering (log/box-cox/robust scaling, winsorizing, or outlier screening) and whether any guard exists against trivial solutions (minimum cluster size, cluster-balance check, elbow/stability cross-check, not only the raw index optimum).
  3. In the printed/reported results, inspect cluster sizes and the chosen k: flag if the optimum is at the boundary of the searched range, if the index is implausibly high (e.g. >0.8 on real customer data), or if any cluster contains a negligible fraction (<1%) of rows.
  4. Verify the saved output rows/columns correspond to the intended units of analysis and that group labels are non-trivial across them.
- **Discriminator**: A genuine violation shows one or two clusters absorbing almost all points with the rest being extreme outliers and no transformation/robustness step; it is fine if skew was handled (or outliers explicitly modeled) and the resulting segments are reasonably populated and stable, even if some segment is small by design.
- **Consequence**: The saved segmentation is effectively an outlier flag rather than a customer segmentation, so cluster labels/count do not match the expected result file and the check fails.
585Deliverable file never verified (path, name, column, row count)taskda-code
Applies when
task -- The task requires writing predictions/results to a named output file with a specified column name and one row per input record.
Pattern
The attempt focuses on modeling and reports metrics/statistics, but writes the output with an unverified or non-default path (absolute/scratch directory, truncated or renamed file), and never re-reads the written file to confirm it exists at the required location with the exact required column name, expected row count, and no index/extra columns. Reported "success" rests on in-memory arrays rather than on the artifact the grader reads.
Detection procedure
  1. From the task statement, extract the exact required artifact: file name, expected directory (default: the working directory the grader inspects), required column name(s), and expected number of rows (= number of test records).
  2. In the scripts, find the write call and check the literal target path and the columns/index= argument passed; note whether the frame written has the required column name verbatim.
  3. Look for a post-write verification step (re-read the file, print shape, columns, head, and compare row count to the test set) and confirm the answer quotes those verified values, not just prediction summary statistics.
  4. Flag if the path is non-standard/unconfirmed, the column name differs in spelling/case, the row count is never compared, or the answer's file path is truncated/ambiguous.
Discriminator
A fine attempt writes to the expected relative path and explicitly echoes the reloaded file's shape/columns matching the test size and required header; a violation infers success from to_csv returning without error, or reports only model/prediction statistics while the artifact's location, header, and length remain unchecked.
Consequence
The grader reports the expected result file as missing or wrong (absent, misnamed, wrong column, or wrong number of rows), scoring 0 regardless of prediction quality.
id 5e1b2ca40084 · mined from da-code dacode-ml-regression-002@s11
raw text (what the judge reads)
### Deliverable file never verified (path, name, column, row count)
- **Applies when**: `task` -- The task requires writing predictions/results to a named output file with a specified column name and one row per input record.
- **Pattern**: The attempt focuses on modeling and reports metrics/statistics, but writes the output with an unverified or non-default path (absolute/scratch directory, truncated or renamed file), and never re-reads the written file to confirm it exists at the required location with the exact required column name, expected row count, and no index/extra columns. Reported "success" rests on in-memory arrays rather than on the artifact the grader reads.
- **Detection procedure**:
  1. From the task statement, extract the exact required artifact: file name, expected directory (default: the working directory the grader inspects), required column name(s), and expected number of rows (= number of test records).
  2. In the scripts, find the write call and check the literal target path and the columns/`index=` argument passed; note whether the frame written has the required column name verbatim.
  3. Look for a post-write verification step (re-read the file, print shape, columns, head, and compare row count to the test set) and confirm the answer quotes those verified values, not just prediction summary statistics.
  4. Flag if the path is non-standard/unconfirmed, the column name differs in spelling/case, the row count is never compared, or the answer's file path is truncated/ambiguous.
- **Discriminator**: A fine attempt writes to the expected relative path and explicitly echoes the reloaded file's shape/columns matching the test size and required header; a violation infers success from `to_csv` returning without error, or reports only model/prediction statistics while the artifact's location, header, and length remain unchecked.
- **Consequence**: The grader reports the expected result file as missing or wrong (absent, misnamed, wrong column, or wrong number of rows), scoring 0 regardless of prediction quality.
586Missing machine-checkable artifacts behind a visualization/deliverabletaskda-code
Applies when
task -- the task asks for a chart or other rendered output built from a spec file, and grading depends on the underlying plotted/derived data, not just the image.
Pattern
The attempt produces only the rendered image (and a prose summary of its styling), never persisting the plotted series, axis values, or spec-derived metadata in the structured side files the harness checks; no reusable script is saved either, so the data behind the figure cannot be recovered or verified.
Detection procedure
  1. Read the task and any referenced spec/config file for the full list of expected outputs (image plus any data/serialized companions) and their exact filenames/extensions.
  2. Inspect the scripts for explicit write calls for each expected artifact (e.g., array dump, JSON dump) — not just a figure-save call; confirm the written objects are the final plotted x/y values in the required order, dtype, and shape.
  3. Check the answer's file inventory: does it enumerate every required artifact, or only the image plus narrative claims about title/color/labels?
  4. If any required artifact is absent or only described in prose, flag; also flag if no script exists to regenerate it.
Discriminator
A real violation is missing or unwritten required output files, or files whose contents are not the exact requested series. A look-alike that is fine is an attempt that saves all required artifacts but describes only the image in its prose summary, or that names the companion files differently only because the task/spec allows it.
Consequence
The grader's per-file checks for the companion data files report WRONG/MISSING and the task scores 0, even though the rendered image may look correct.
id 266de5cafec8 · mined from da-code dacode-plot-line-015@s11
raw text (what the judge reads)
### Missing machine-checkable artifacts behind a visualization/deliverable
- **Applies when**: `task` -- the task asks for a chart or other rendered output built from a spec file, and grading depends on the underlying plotted/derived data, not just the image.
- **Pattern**: The attempt produces only the rendered image (and a prose summary of its styling), never persisting the plotted series, axis values, or spec-derived metadata in the structured side files the harness checks; no reusable script is saved either, so the data behind the figure cannot be recovered or verified.
- **Detection procedure**:
  1. Read the task and any referenced spec/config file for the full list of expected outputs (image plus any data/serialized companions) and their exact filenames/extensions.
  2. Inspect the scripts for explicit write calls for each expected artifact (e.g., array dump, JSON dump) — not just a figure-save call; confirm the written objects are the final plotted x/y values in the required order, dtype, and shape.
  3. Check the answer's file inventory: does it enumerate every required artifact, or only the image plus narrative claims about title/color/labels?
  4. If any required artifact is absent or only described in prose, flag; also flag if no script exists to regenerate it.
- **Discriminator**: A real violation is missing or unwritten required output files, or files whose contents are not the exact requested series. A look-alike that is fine is an attempt that saves all required artifacts but describes only the image in its prose summary, or that names the companion files differently only because the task/spec allows it.
- **Consequence**: The grader's per-file checks for the companion data files report WRONG/MISSING and the task scores 0, even though the rendered image may look correct.
587Fabricated/hard-coded input data instead of loading the provided filestaskda-code
Applies when
task -- the task references supplied data files, instruction files (e.g. tips/README), and an output template, but the scripts contain literal arrays or values typed inline.
Pattern
The agent "recalls" a well-known dataset and hard-codes numbers (often guessing at sample sizes, groups, or direction of effect), never reading the actual input files or the specified output schema, then runs a technically correct procedure on invented inputs.
Detection procedure
  1. Read the task for named inputs (data files, tips/instruction files, sample output file) and note that results must be derived from them.
  2. Scan the scripts for any file-reading call (read_csv, open, load, etc.) on those inputs; flag if the analysis operates on inline literals or on data whose provenance is unstated.
  3. Check whether the output columns/row layout were copied from the provided template rather than invented (e.g. a self-chosen column name like pvalue).
  4. Check the answer for tell-tale signs of guessed data: comments like "let me try another version", surprise that the effect direction contradicts the task premise, or multiple candidate datasets tried.
Discriminator
A real violation is when the numeric inputs cannot be traced to a file present in the working directory; it is fine if the script hard-codes only constants (seeds, iteration counts, thresholds) or if it prints a verified checksum/shape confirming the inline data matches the supplied file.
Consequence
The reported statistic is computed on the wrong numbers and the output file's values (and often its schema) mismatch the expected result, so the grader marks the result file WRONG even though the method looks sound.
id 777ccc54b21c · mined from da-code dacode-data-sa-028@s11
raw text (what the judge reads)
### Fabricated/hard-coded input data instead of loading the provided files
- **Applies when**: `task` -- the task references supplied data files, instruction files (e.g. tips/README), and an output template, but the scripts contain literal arrays or values typed inline.
- **Pattern**: The agent "recalls" a well-known dataset and hard-codes numbers (often guessing at sample sizes, groups, or direction of effect), never reading the actual input files or the specified output schema, then runs a technically correct procedure on invented inputs.
- **Detection procedure**:
  1. Read the task for named inputs (data files, tips/instruction files, sample output file) and note that results must be derived from them.
  2. Scan the scripts for any file-reading call (`read_csv`, `open`, `load`, etc.) on those inputs; flag if the analysis operates on inline literals or on data whose provenance is unstated.
  3. Check whether the output columns/row layout were copied from the provided template rather than invented (e.g. a self-chosen column name like `pvalue`).
  4. Check the answer for tell-tale signs of guessed data: comments like "let me try another version", surprise that the effect direction contradicts the task premise, or multiple candidate datasets tried.
- **Discriminator**: A real violation is when the numeric inputs cannot be traced to a file present in the working directory; it is fine if the script hard-codes only constants (seeds, iteration counts, thresholds) or if it prints a verified checksum/shape confirming the inline data matches the supplied file.
- **Consequence**: The reported statistic is computed on the wrong numbers and the output file's values (and often its schema) mismatch the expected result, so the grader marks the result file WRONG even though the method looks sound.
588Reported success without a persisted, re-runnable script or verified output artifactstaskda-code
Applies when
task -- the task asks for concrete output files/plots built according to an external spec document, and the agent's deliverable is a prose summary of numbers plus a claim that files were saved.
Pattern
The attempt leaves no saved script (or a script that isn't the one that produced the reported numbers), never re-reads/quotes the referenced spec to derive the grouping/binning rules, and never checks the written artifacts exist with the expected content/shape — it just asserts "successfully created" and lists totals that conveniently match a headline figure from the README.
Detection procedure
  1. List the artifacts the task requires (files, names, titles, axis labels, any auxiliary serialized data) and check that a saved script writes each one explicitly.
  2. Check that the script loads and applies the rules from the referenced spec file (bin edges, category labels, ordering) rather than hard-coding categories from memory or from the prose answer.
  3. Verify the script contains post-write sanity checks: file exists, category counts sum to the number of valid (non-missing) records, no unmapped/NaN category silently dropped or lumped in.
  4. Compare the numbers in the prose answer against something the script actually prints; if the only evidence is narrative text, treat the result as unverified.
Discriminator
A real violation has no reproducible code path from raw data → spec-defined groups → all required files, and no printed verification; a look-alike that is fine has a short script that reads the spec, writes every requested artifact, and echoes the counts/shape it wrote (even if the prose summary is brief).
Consequence
The grader finds the required files missing, misnamed, or containing values that don't match the spec-defined grouping, so every artifact check fails despite a confident success report.
id d36048034ca5 · mined from da-code dacode-plot-bar-005@s11
raw text (what the judge reads)
### Reported success without a persisted, re-runnable script or verified output artifacts
- **Applies when**: `task` -- the task asks for concrete output files/plots built according to an external spec document, and the agent's deliverable is a prose summary of numbers plus a claim that files were saved.
- **Pattern**: The attempt leaves no saved script (or a script that isn't the one that produced the reported numbers), never re-reads/quotes the referenced spec to derive the grouping/binning rules, and never checks the written artifacts exist with the expected content/shape — it just asserts "successfully created" and lists totals that conveniently match a headline figure from the README.
- **Detection procedure**:
  1. List the artifacts the task requires (files, names, titles, axis labels, any auxiliary serialized data) and check that a saved script writes each one explicitly.
  2. Check that the script loads and applies the rules from the referenced spec file (bin edges, category labels, ordering) rather than hard-coding categories from memory or from the prose answer.
  3. Verify the script contains post-write sanity checks: file exists, category counts sum to the number of valid (non-missing) records, no unmapped/NaN category silently dropped or lumped in.
  4. Compare the numbers in the prose answer against something the script actually prints; if the only evidence is narrative text, treat the result as unverified.
- **Discriminator**: A real violation has no reproducible code path from raw data → spec-defined groups → all required files, and no printed verification; a look-alike that is fine has a short script that reads the spec, writes every requested artifact, and echoes the counts/shape it wrote (even if the prose summary is brief).
- **Consequence**: The grader finds the required files missing, misnamed, or containing values that don't match the spec-defined grouping, so every artifact check fails despite a confident success report.
589Answer artifact not persisted in the exact requested schemataskda-code
Applies when
task -- the task specifies an output template (e.g., a JSON object with named keys whose values are shown as lists) and/or the harness expects a result file on disk.
Pattern
The agent computes a value and reports it only in chat prose, or writes/prints it with a different container type or key/value shape than the template (scalar instead of list, extra/renamed keys, stringified numbers), and leaves no script that regenerates the file.
Detection procedure
  1. Read the task and transcribe the literal output template: required file name, key names, and the type of each value as displayed (bracketed values imply lists/arrays).
  2. Search the scripts for code that constructs that object and writes it to the expected file path (e.g., json.dump(..., open("result.json","w"))); note whether any script exists at all.
  3. Compare the submitted answer element-by-element against the template: key spelling, container type, number of entries, numeric type/rounding.
  4. Flag if the file is never written, or if any key/value shape deviates from the template, or if no script reproduces it.
Discriminator
A real violation is a structural/location mismatch (missing file, scalar where a list is shown, renamed keys) even if the computed number is arguably right; a look-alike that is fine is cosmetic difference the spec leaves open (key ordering, whitespace, int vs float of the same value) with the required file present.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying computation was correct.
id 4e964dcde23a · mined from da-code dacode-di-text-002@s11
raw text (what the judge reads)
### Answer artifact not persisted in the exact requested schema
- **Applies when**: `task` -- the task specifies an output template (e.g., a JSON object with named keys whose values are shown as lists) and/or the harness expects a result file on disk.
- **Pattern**: The agent computes a value and reports it only in chat prose, or writes/prints it with a different container type or key/value shape than the template (scalar instead of list, extra/renamed keys, stringified numbers), and leaves no script that regenerates the file.
- **Detection procedure**:
  1. Read the task and transcribe the literal output template: required file name, key names, and the type of each value as displayed (bracketed values imply lists/arrays).
  2. Search the scripts for code that constructs that object and writes it to the expected file path (e.g., `json.dump(..., open("result.json","w"))`); note whether any script exists at all.
  3. Compare the submitted answer element-by-element against the template: key spelling, container type, number of entries, numeric type/rounding.
  4. Flag if the file is never written, or if any key/value shape deviates from the template, or if no script reproduces it.
- **Discriminator**: A real violation is a structural/location mismatch (missing file, scalar where a list is shown, renamed keys) even if the computed number is arguably right; a look-alike that is fine is cosmetic difference the spec leaves open (key ordering, whitespace, int vs float of the same value) with the required file present.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying computation was correct.
590Ignoring a provided sample/expected-output artifact when fixing the result's conventiontaskda-code
Applies when
task -- the task demands a result file "in the required format" and the workspace contains a sample/template/reference version of that file (or a documented schema), and the requested quantity has more than one common definition (e.g., cumulative vs. incremental, level vs. change, with/without a baseline row, percent vs. fraction, rounding).
Pattern
The attempt picks one plausible formula/convention from memory, writes the file, and only prints the sample file for eyeball comparison without programmatically reconciling row count, row ordering, index/baseline row, value scale, and offset; any systematic mismatch (e.g., every value differing by a constant, a factor of 100, or a one-row shift) goes unnoticed and unresolved.
Detection procedure
  1. Read the task/README for wording about "required format" and note any sample/expected output file or schema the scripts open.
  2. In the scripts, check whether the sample is used only for display (print(head/tail)) or is actually used as a check: same number of rows, same first/last index values, same columns, and a numeric comparison (difference, ratio, or offset by 1) on overlapping rows.
  3. Inspect how the requested quantity is defined in code and enumerate the alternative conventions (subtracting a baseline or not, starting from a seed/initial row, units, rounding, per-period vs. running aggregate); confirm the script contains an explicit justification or a test that discriminates among them.
  4. Look at the reported answer's first row: does it match what the sample/format implies for the initial period (e.g., a baseline value or a zero/one starting point) and is the row count identical to the sample's?
Discriminator
A real violation is when a reference artifact or explicit schema was available and the attempt never numerically reconciled its own output against it (or reconciled and then shipped a known systematic difference). It is not a violation if the script asserts equality/closeness on shared structure and the only differences are genuinely unconstrained (e.g., extra float precision within the stated tolerance), or if no reference/schema exists at all.
Consequence
The saved file has the right shape but the wrong convention (constant offset, wrong scale, or shifted/missing baseline row), so an exact/tolerance-based file comparison marks the result WRONG despite the underlying arithmetic being reasonable.
id a98ef987c3f5 · mined from da-code dacode-dm-csv-050@s11
raw text (what the judge reads)
### Ignoring a provided sample/expected-output artifact when fixing the result's convention
- **Applies when**: `task` -- the task demands a result file "in the required format" and the workspace contains a sample/template/reference version of that file (or a documented schema), and the requested quantity has more than one common definition (e.g., cumulative vs. incremental, level vs. change, with/without a baseline row, percent vs. fraction, rounding).
- **Pattern**: The attempt picks one plausible formula/convention from memory, writes the file, and only *prints* the sample file for eyeball comparison without programmatically reconciling row count, row ordering, index/baseline row, value scale, and offset; any systematic mismatch (e.g., every value differing by a constant, a factor of 100, or a one-row shift) goes unnoticed and unresolved.
- **Detection procedure**:
  1. Read the task/README for wording about "required format" and note any sample/expected output file or schema the scripts open.
  2. In the scripts, check whether the sample is used only for display (`print(head/tail)`) or is actually used as a check: same number of rows, same first/last index values, same columns, and a numeric comparison (difference, ratio, or offset by 1) on overlapping rows.
  3. Inspect how the requested quantity is defined in code and enumerate the alternative conventions (subtracting a baseline or not, starting from a seed/initial row, units, rounding, per-period vs. running aggregate); confirm the script contains an explicit justification or a test that discriminates among them.
  4. Look at the reported answer's first row: does it match what the sample/format implies for the initial period (e.g., a baseline value or a zero/one starting point) and is the row count identical to the sample's?
- **Discriminator**: A real violation is when a reference artifact or explicit schema was available and the attempt never numerically reconciled its own output against it (or reconciled and then shipped a known systematic difference). It is *not* a violation if the script asserts equality/closeness on shared structure and the only differences are genuinely unconstrained (e.g., extra float precision within the stated tolerance), or if no reference/schema exists at all.
- **Consequence**: The saved file has the right shape but the wrong convention (constant offset, wrong scale, or shifted/missing baseline row), so an exact/tolerance-based file comparison marks the result WRONG despite the underlying arithmetic being reasonable.
591Blindly loading one input source without verifying it is the intended analysis populationtaskinfiagent-dabench
Applies when
task -- the task asks for a statistic/test on a named column or variable, and the script hard-codes a single file path (or a single parse option such as index_col, delimiter, header, or an implicit row filter) without checking what data is actually available or expected.
Pattern
The script picks one candidate file/slice from the data directory, reads it with unverified parsing options, and immediately runs the test and moment statistics on whatever rows come back. There is no listing of available inputs, no comparison of row/column counts against the task's implied scope, no check that the parsed column is numeric and complete, and no plausibility check of the resulting statistics before they are reported as the final answer.
Detection procedure
  1. Read the task and note the population the statistic is supposed to describe (which file(s), which rows, any stated filtering, and roughly how many observations that implies).
  2. Read the script's loading block: does it enumerate/justify the chosen file and parsing options, or is a single path silently hard-coded? Does it print shape, dtype, NA counts, and a head of the target column and compare them against the task's expectation?
  3. Check whether the script validates the statistical procedure against the data it loaded (e.g., sample-size validity range for the chosen test, numeric vs. categorical dtype, degenerate/constant or truncated values) rather than trusting the first number returned.
  4. Compare the reported outputs for internal consistency (test verdict vs. skewness/kurtosis magnitudes, means/medians vs. the column's documented meaning); flag if extreme values are reported with no comment or cross-check.
Discriminator
A real violation is a script whose data-selection and parsing choices are unexamined assumptions — swapping in a different available file, header/index option, or row subset would silently change the answer and nobody would notice. It is not a violation when the task unambiguously specifies a single input and the script still prints shape/dtype/NA diagnostics that a reviewer can match to that specification, even if it then proceeds directly to the statistic.
Consequence
The reported test verdict and moment statistics describe the wrong rows (or a mis-parsed column), so the graded fields — especially the categorical yes/no verdict — flip relative to ground truth even though the code itself runs without error.
id 8456f0f481e7 · mined from infiagent-dabench dabench-298@s11
raw text (what the judge reads)
### Blindly loading one input source without verifying it is the intended analysis population
- **Applies when**: `task` -- the task asks for a statistic/test on a named column or variable, and the script hard-codes a single file path (or a single parse option such as `index_col`, delimiter, header, or an implicit row filter) without checking what data is actually available or expected.
- **Pattern**: The script picks one candidate file/slice from the data directory, reads it with unverified parsing options, and immediately runs the test and moment statistics on whatever rows come back. There is no listing of available inputs, no comparison of row/column counts against the task's implied scope, no check that the parsed column is numeric and complete, and no plausibility check of the resulting statistics before they are reported as the final answer.
- **Detection procedure**:
  1. Read the task and note the population the statistic is supposed to describe (which file(s), which rows, any stated filtering, and roughly how many observations that implies).
  2. Read the script's loading block: does it enumerate/justify the chosen file and parsing options, or is a single path silently hard-coded? Does it print shape, dtype, NA counts, and a head of the target column and compare them against the task's expectation?
  3. Check whether the script validates the statistical procedure against the data it loaded (e.g., sample-size validity range for the chosen test, numeric vs. categorical dtype, degenerate/constant or truncated values) rather than trusting the first number returned.
  4. Compare the reported outputs for internal consistency (test verdict vs. skewness/kurtosis magnitudes, means/medians vs. the column's documented meaning); flag if extreme values are reported with no comment or cross-check.
- **Discriminator**: A real violation is a script whose data-selection and parsing choices are unexamined assumptions — swapping in a different available file, header/index option, or row subset would silently change the answer and nobody would notice. It is *not* a violation when the task unambiguously specifies a single input and the script still prints shape/dtype/NA diagnostics that a reviewer can match to that specification, even if it then proceeds directly to the statistic.
- **Consequence**: The reported test verdict and moment statistics describe the wrong rows (or a mis-parsed column), so the graded fields — especially the categorical yes/no verdict — flip relative to ground truth even though the code itself runs without error.
592Underfit final model shipped without comparing candidates on a common validation metrictaskda-code
Applies when
task -- Scripts train several candidate models (different families, subsamples, or hyperparameters) and the last one executed overwrites the submission file used for scoring.
Pattern
The agent trades accuracy for runtime (training on a fraction of available rows, capped depth/estimators, arbitrary clipping of outputs) and ships whichever model ran last, without evaluating all candidates on one identical holdout and without checking that the prediction distribution matches the target distribution — so a weak model silently replaces a stronger one and the reported validation score is low in absolute terms.
Detection procedure
  1. Read the task to confirm the deliverable is scored on predictive accuracy against a hidden ground truth, not merely on file existence.
  2. In the scripts, list every candidate model and note which ones report a holdout metric on the same split and same feature set; flag any candidate that writes the submission without a comparable score, that fits on a subsample of the available training data, or that post-processes predictions with hard-coded clipping bounds not derived from the data.
  3. Compare the reported validation metric of the shipped model against a trivial baseline (predicting the target mean, or a plain linear/ridge fit on all rows); flag if the shipped model's error/R² is not clearly better, or if no such baseline was ever computed.
  4. Compare the spread (std, min–max) of the submitted predictions with the spread of the training target; flag if predictions are collapsed to a much narrower band than the target, a signature of underfitting/over-regularization.
Discriminator
A real violation is shipping a model whose only reported score is mediocre relative to an unmeasured or better-scoring alternative, or whose prediction variance is far below the target variance. It is not a violation when the agent explicitly benchmarks all candidates on one held-out split (or CV), selects the best, refits on the full training data, and the prediction spread is consistent with the target's.
Consequence
The submission file exists and is correctly formatted, but its column of predictions is too inaccurate/too flat to pass the accuracy threshold, so the grader marks the expected result file WRONG despite a confident "submission complete" report.
id 1581532fbb2e · mined from da-code dacode-ml-competition-008@s11
raw text (what the judge reads)
### Underfit final model shipped without comparing candidates on a common validation metric
- **Applies when**: `task` -- Scripts train several candidate models (different families, subsamples, or hyperparameters) and the last one executed overwrites the submission file used for scoring.
- **Pattern**: The agent trades accuracy for runtime (training on a fraction of available rows, capped depth/estimators, arbitrary clipping of outputs) and ships whichever model ran last, without evaluating all candidates on one identical holdout and without checking that the prediction distribution matches the target distribution — so a weak model silently replaces a stronger one and the reported validation score is low in absolute terms.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on predictive accuracy against a hidden ground truth, not merely on file existence.
  2. In the scripts, list every candidate model and note which ones report a holdout metric on the *same* split and same feature set; flag any candidate that writes the submission without a comparable score, that fits on a subsample of the available training data, or that post-processes predictions with hard-coded clipping bounds not derived from the data.
  3. Compare the reported validation metric of the shipped model against a trivial baseline (predicting the target mean, or a plain linear/ridge fit on all rows); flag if the shipped model's error/R² is not clearly better, or if no such baseline was ever computed.
  4. Compare the spread (std, min–max) of the submitted predictions with the spread of the training target; flag if predictions are collapsed to a much narrower band than the target, a signature of underfitting/over-regularization.
- **Discriminator**: A real violation is shipping a model whose only reported score is mediocre relative to an unmeasured or better-scoring alternative, or whose prediction variance is far below the target variance. It is *not* a violation when the agent explicitly benchmarks all candidates on one held-out split (or CV), selects the best, refits on the full training data, and the prediction spread is consistent with the target's.
- **Consequence**: The submission file exists and is correctly formatted, but its column of predictions is too inaccurate/too flat to pass the accuracy threshold, so the grader marks the expected result file WRONG despite a confident "submission complete" report.
593Deliverables asserted in prose instead of produced and sanity-checked against the provided spec/datataskda-code
Applies when
task -- the task names concrete output artifacts (plot/config/array/table files) and/or a spec file that dictates formatting, and the agent's final answer is a narrative summary of numbers and "chart details".
Pattern
The attempt describes what it supposedly built (file size, DPI, title, bin counts) rather than showing scripts that read the spec file, write every named artifact, and re-load/verify them; underlying inputs are also taken on faith, so a truncated, sampled, or placeholder version of the source data silently drives the reported distribution.
Detection procedure
  1. From the task text, list every required output file and every stated formatting/derivation constraint (source of config, bin edges, ordering, units, orientation).
  2. Search the scripts for an explicit write of each listed file and an explicit read/parse of the spec file; any artifact mentioned only in the answer text (or any spec value hard-coded rather than loaded) is unbacked.
  3. Cross-check the reported input scale and value ranges against the README/raw file (row count, min/max, presence of both entity types, number of non-null duration values); flag suspiciously round totals or ranges clipped to the bin boundaries.
  4. Confirm the answer's tabulated numbers are reproducible from the written artifacts (e.g., saved array/JSON) rather than only printed to stdout.
Discriminator
A real violation is when a required file is never written by any script, the spec is never read, or the input's shape/range contradicts the documented raw data; it is fine if all files are written and verified and the prose summary merely restates their contents, even if the narrative is verbose.
Consequence
The grader finds expected artifacts missing or containing values derived from the wrong (partial/synthetic) input, so every file-level check fails regardless of how plausible the printed summary looks.
id 9ce6884a8151 · mined from da-code dacode-plot-bar-007@s11
raw text (what the judge reads)
### Deliverables asserted in prose instead of produced and sanity-checked against the provided spec/data
- **Applies when**: `task` -- the task names concrete output artifacts (plot/config/array/table files) and/or a spec file that dictates formatting, and the agent's final answer is a narrative summary of numbers and "chart details".
- **Pattern**: The attempt describes what it supposedly built (file size, DPI, title, bin counts) rather than showing scripts that read the spec file, write every named artifact, and re-load/verify them; underlying inputs are also taken on faith, so a truncated, sampled, or placeholder version of the source data silently drives the reported distribution.
- **Detection procedure**:
  1. From the task text, list every required output file and every stated formatting/derivation constraint (source of config, bin edges, ordering, units, orientation).
  2. Search the scripts for an explicit write of each listed file and an explicit read/parse of the spec file; any artifact mentioned only in the answer text (or any spec value hard-coded rather than loaded) is unbacked.
  3. Cross-check the reported input scale and value ranges against the README/raw file (row count, min/max, presence of both entity types, number of non-null duration values); flag suspiciously round totals or ranges clipped to the bin boundaries.
  4. Confirm the answer's tabulated numbers are reproducible from the written artifacts (e.g., saved array/JSON) rather than only printed to stdout.
- **Discriminator**: A real violation is when a required file is never written by any script, the spec is never read, or the input's shape/range contradicts the documented raw data; it is fine if all files are written and verified and the prose summary merely restates their contents, even if the narrative is verbose.
- **Consequence**: The grader finds expected artifacts missing or containing values derived from the wrong (partial/synthetic) input, so every file-level check fails regardless of how plausible the printed summary looks.
594Undefined/empty-result values not emitted in the exact requested literal formattaskinfiagent-dabench
Applies when
task -- a task asks for a numeric statistic in a strict answer template (e.g. float rounded to N decimals) but the requested filtering/aggregation can legitimately yield no rows or an undefined value.
Pattern
The attempt (often without a saved, re-runnable script) discovers the statistic is undefined and hand-writes a token for it — a differently-cased or differently-spelled placeholder ("NaN", "None", "N/A", "undefined", "", "0.00") — instead of emitting the value exactly as the evaluation harness would produce/compare it (e.g. the string a float prints, or the format literally specified in the task), and provides no evidence (row counts, subset shape) that the empty/undefined outcome is real rather than a bug.
Detection procedure
  1. Read the task's answer template and note the required type/precision and the literal string form expected for the value, including how a missing/undefined result would be spelled.
  2. Read the scripts: check that the filtering chain is applied in the stated order, that the subset size is printed, and that the final answer string is produced programmatically (e.g. by formatting the computed value) rather than typed by hand.
  3. Compare the submitted token character-by-character with what the script's formatting would emit for that value (case, spacing, decimals, sign).
  4. If no script was saved, or the answer token cannot be traced to a printed program output, flag the attempt.
Discriminator
A real violation is a hand-typed or re-cased/rephrased placeholder or a value whose formatting deviates from the template; it is fine if the answer string is exactly the program-produced formatting of the computed value (including a legitimately empty subset), backed by a script that prints the subset counts and the final formatted string.
Consequence
The grader's string/value comparison fails even though the underlying computation may be right, scoring 0 on the check.
id 65c5df5847cf · mined from infiagent-dabench dabench-554@s11
raw text (what the judge reads)
### Undefined/empty-result values not emitted in the exact requested literal format
- **Applies when**: `task` -- a task asks for a numeric statistic in a strict answer template (e.g. float rounded to N decimals) but the requested filtering/aggregation can legitimately yield no rows or an undefined value.
- **Pattern**: The attempt (often without a saved, re-runnable script) discovers the statistic is undefined and hand-writes a token for it — a differently-cased or differently-spelled placeholder ("NaN", "None", "N/A", "undefined", "", "0.00") — instead of emitting the value exactly as the evaluation harness would produce/compare it (e.g. the string a `float` prints, or the format literally specified in the task), and provides no evidence (row counts, subset shape) that the empty/undefined outcome is real rather than a bug.
- **Detection procedure**:
  1. Read the task's answer template and note the required type/precision and the literal string form expected for the value, including how a missing/undefined result would be spelled.
  2. Read the scripts: check that the filtering chain is applied in the stated order, that the subset size is printed, and that the final answer string is produced programmatically (e.g. by formatting the computed value) rather than typed by hand.
  3. Compare the submitted token character-by-character with what the script's formatting would emit for that value (case, spacing, decimals, sign).
  4. If no script was saved, or the answer token cannot be traced to a printed program output, flag the attempt.
- **Discriminator**: A real violation is a hand-typed or re-cased/rephrased placeholder or a value whose formatting deviates from the template; it is fine if the answer string is exactly the program-produced formatting of the computed value (including a legitimately empty subset), backed by a script that prints the subset counts and the final formatted string.
- **Consequence**: The grader's string/value comparison fails even though the underlying computation may be right, scoring 0 on the check.
595Ignoring the stated ordering (or other formatting) constraint on the reported listtaskda-code
Applies when
task -- the task asks for a ranked subset of records and explicitly states the order (or rounding/units/format) in which all reported groups must be presented.
Pattern
The script computes the correct top-k and bottom-k subsets but then re-sorts one of them in the opposite direction (e.g., leaves the bottom group ascending after a descending selection), or otherwise emits the values in an order/format other than the one the prompt dictates, and the agent never re-reads the constraint against the final output.
Detection procedure
  1. Extract every explicit output constraint from the task text (ordering direction, rounding, units, key names, file to write).
  2. In the scripts, locate each sort_values/head/tail/reset_index chain and note the actual direction and sequence in which each reported group is finally assembled; also confirm the result is written to the requested artifact, not only printed.
  3. Compare the final answer's element order for each group against the stated direction — check the first and last element of every list against the underlying values.
  4. Flag if any group's order (or format/persistence) deviates from the stated requirement.
Discriminator
A real violation is when the prompt fixes one direction for all groups and a group is emitted in the other direction (or the requested output artifact is missing); a look-alike that is fine is when the prompt is silent about order for that group, or when a temporary ascending sort is used internally but reversed before output.
Consequence
The grader compares element-by-element against the reference list and marks the answer wrong (or missing) even though the selected members are correct.
id 4a18397fa74f · mined from da-code dacode-di-text-003@s11
raw text (what the judge reads)
### Ignoring the stated ordering (or other formatting) constraint on the reported list
- **Applies when**: `task` -- the task asks for a ranked subset of records and explicitly states the order (or rounding/units/format) in which all reported groups must be presented.
- **Pattern**: The script computes the correct top-k and bottom-k subsets but then re-sorts one of them in the opposite direction (e.g., leaves the bottom group ascending after a descending selection), or otherwise emits the values in an order/format other than the one the prompt dictates, and the agent never re-reads the constraint against the final output.
- **Detection procedure**:
  1. Extract every explicit output constraint from the task text (ordering direction, rounding, units, key names, file to write).
  2. In the scripts, locate each `sort_values`/`head`/`tail`/`reset_index` chain and note the actual direction and sequence in which each reported group is finally assembled; also confirm the result is written to the requested artifact, not only printed.
  3. Compare the final answer's element order for *each* group against the stated direction — check the first and last element of every list against the underlying values.
  4. Flag if any group's order (or format/persistence) deviates from the stated requirement.
- **Discriminator**: A real violation is when the prompt fixes one direction for all groups and a group is emitted in the other direction (or the requested output artifact is missing); a look-alike that is fine is when the prompt is silent about order for that group, or when a temporary ascending sort is used internally but reversed before output.
- **Consequence**: The grader compares element-by-element against the reference list and marks the answer wrong (or missing) even though the selected members are correct.
596Reconstructing a derived quantity instead of using (or reconciling with) the dataset's existing derived columntaskinfiagent-dabench
Applies when
task -- the task asks for a statistic over a "difference"/"ratio"/"rate"-type quantity and the input file already contains a pre-computed column for that quantity (or a closely related one) that the script recomputes from raw fields.
Pattern
The script silently redefines the quantity from raw columns (choosing its own operand order, sign, absolute-value convention, scaling, or string-cleaning rule) and never checks that its reconstruction matches the provided column or that row counts/NaNs are unchanged; the statistic is then computed on a slightly different vector than the one the task refers to.
Detection procedure
  1. Read the task and list every quantity it names; note which of these exist as ready-made columns in the input file (the script's own head/columns printout reveals this).
  2. In the script, check whether each such quantity is recomputed; if so, look for an explicit equality/agreement check (e.g., comparing recomputed vs. provided column, or asserting sign/scale/absolute-value convention and identical non-null counts).
  3. Check whether rows with missing/unparseable values in either variable are handled consistently across both variables before the statistic is computed, and whether the effective N is printed and sanity-checked against the file's row count.
  4. If no reconciliation or N check exists, treat the reported statistic (especially p-value, which is far more sensitive to N and to a few discrepant rows than a 2-decimal correlation) as unverified.
Discriminator
A real violation is recomputation with no comparison against the available pre-computed column and no reported N; it is fine if the script recomputes but explicitly verifies agreement (or documents a justified deviation, e.g., the provided column is rounded) and confirms the same rows are used for both variables.
Consequence
The headline statistic can round to the expected value while the derived, N-sensitive output (p-value) is off, so the answer fails the exact-match check on that field even though the qualitative conclusion looks right.
id c96110f3a00c · mined from infiagent-dabench dabench-142@s11
raw text (what the judge reads)
### Reconstructing a derived quantity instead of using (or reconciling with) the dataset's existing derived column
- **Applies when**: `task` -- the task asks for a statistic over a "difference"/"ratio"/"rate"-type quantity and the input file already contains a pre-computed column for that quantity (or a closely related one) that the script recomputes from raw fields.
- **Pattern**: The script silently redefines the quantity from raw columns (choosing its own operand order, sign, absolute-value convention, scaling, or string-cleaning rule) and never checks that its reconstruction matches the provided column or that row counts/NaNs are unchanged; the statistic is then computed on a slightly different vector than the one the task refers to.
- **Detection procedure**:
  1. Read the task and list every quantity it names; note which of these exist as ready-made columns in the input file (the script's own head/columns printout reveals this).
  2. In the script, check whether each such quantity is recomputed; if so, look for an explicit equality/agreement check (e.g., comparing recomputed vs. provided column, or asserting sign/scale/absolute-value convention and identical non-null counts).
  3. Check whether rows with missing/unparseable values in either variable are handled consistently across both variables before the statistic is computed, and whether the effective N is printed and sanity-checked against the file's row count.
  4. If no reconciliation or N check exists, treat the reported statistic (especially p-value, which is far more sensitive to N and to a few discrepant rows than a 2-decimal correlation) as unverified.
- **Discriminator**: A real violation is recomputation with no comparison against the available pre-computed column and no reported N; it is fine if the script recomputes but explicitly verifies agreement (or documents a justified deviation, e.g., the provided column is rounded) and confirms the same rows are used for both variables.
- **Consequence**: The headline statistic can round to the expected value while the derived, N-sensitive output (p-value) is off, so the answer fails the exact-match check on that field even though the qualitative conclusion looks right.
597Fabricated proxy mapping when the required fields are absent from the loaded datataskda-code
Applies when
task -- the task names specific entities/quantities (e.g., grouping key, measure, category dimension) and required config/output artifacts, and the scripts must locate them in the provided input files.
Pattern
The agent loads a file whose schema has none of the named fields, then silently redefines them with unrelated proxies (renaming one column as the grouping key, substituting a count for the measure, inventing categories and hard-coded values for the breakdown), and/or ignores the specified settings file and the other required output artifacts, producing a plot that is internally consistent but answers a different question.
Detection procedure
  1. From the task, list the exact required inputs (named fields, config file) and required outputs (all filenames/formats, e.g. figure plus serialized data).
  2. In the scripts, check that each named field is read from a real column of the loaded data, not renamed, guessed, or synthesized from constants; check the config file is actually parsed and its settings applied.
  3. In the answer/logs, look for statements that redefine terms ("X was interpreted as Y", "derived from...", literal made-up values per category) or that describe a domain unrelated to the task wording.
  4. Verify every required output artifact is produced with the requested structure, not just the image.
Discriminator
A real violation is inventing a definition or value that has no derivable basis in the data (or reading the wrong file entirely) without flagging the mismatch; acceptable is a documented, data-grounded mapping when a column is merely named differently but holds the requested semantics — or explicitly halting and reporting that the required fields/inputs are missing.
Consequence
All expected artifacts are missing or numerically unrelated to ground truth; every file-level check fails (0/N), even though the run "completed successfully".
id 9f0b65804dd9 · mined from da-code dacode-plot-scatter-002@s11
raw text (what the judge reads)
### Fabricated proxy mapping when the required fields are absent from the loaded data
- **Applies when**: `task` -- the task names specific entities/quantities (e.g., grouping key, measure, category dimension) and required config/output artifacts, and the scripts must locate them in the provided input files.
- **Pattern**: The agent loads a file whose schema has none of the named fields, then silently redefines them with unrelated proxies (renaming one column as the grouping key, substituting a count for the measure, inventing categories and hard-coded values for the breakdown), and/or ignores the specified settings file and the other required output artifacts, producing a plot that is internally consistent but answers a different question.
- **Detection procedure**:
  1. From the task, list the exact required inputs (named fields, config file) and required outputs (all filenames/formats, e.g. figure plus serialized data).
  2. In the scripts, check that each named field is read from a real column of the loaded data, not renamed, guessed, or synthesized from constants; check the config file is actually parsed and its settings applied.
  3. In the answer/logs, look for statements that redefine terms ("X was interpreted as Y", "derived from...", literal made-up values per category) or that describe a domain unrelated to the task wording.
  4. Verify every required output artifact is produced with the requested structure, not just the image.
- **Discriminator**: A real violation is inventing a definition or value that has no derivable basis in the data (or reading the wrong file entirely) without flagging the mismatch; acceptable is a documented, data-grounded mapping when a column is merely named differently but holds the requested semantics — or explicitly halting and reporting that the required fields/inputs are missing.
- **Consequence**: All expected artifacts are missing or numerically unrelated to ground truth; every file-level check fails (0/N), even though the run "completed successfully".
598Output file not validated cell-by-cell against the provided templatetaskda-code
Applies when
task -- The task says to save results to a named file "matching the provided template/example format", and the scripts write a derived table (pivot/matrix/summary) to disk.
Pattern
The agent computes the values, then writes the file using its own assumptions about orientation, index/header names, value scale (proportion vs percent), rounding, ordering, and missing-value representation, without ever loading the template and diffing structure against it; the answer describes the chosen format instead of proving it matches.
Detection procedure
  1. Read the task and note that a template/reference file exists and defines the expected schema (column names/order, index column presence, units/scale, decimal precision, blank vs NaN encoding).
  2. Search the scripts for any read of the template file and an explicit comparison (e.g., asserting equal column lists, shape, index name, or a value-range check) before/after saving.
  3. If no such load-and-compare exists, check whether the answer justifies each formatting choice with evidence from the template (a printed head of the template, a shape/dtype match) rather than with phrasing like "values are proportions 0–1" or "NaN as empty cells".
  4. Sanity-check the described values against the template's own scale and precision (e.g., 0–1 vs 0–100, unrounded floats vs 1–2 decimals) and against expected shape/row count.
Discriminator
A real violation is unverified self-invented formatting — no template read, no assertion, no printed side-by-side comparison. It is fine if the script loads the template (or its header) and asserts matching columns/index/shape/scale, or if the answer shows the template's first rows next to the produced output, even if the format was ultimately simple.
Consequence
The file exists and the numbers may be conceptually right, but the automated check comparing it to the expected file fails on schema, scale, rounding, or missing-value encoding, scoring 0.
id f9e1e163da9a · mined from da-code dacode-dm-csv-043@s11
raw text (what the judge reads)
### Output file not validated cell-by-cell against the provided template
- **Applies when**: `task` -- The task says to save results to a named file "matching the provided template/example format", and the scripts write a derived table (pivot/matrix/summary) to disk.
- **Pattern**: The agent computes the values, then writes the file using its own assumptions about orientation, index/header names, value scale (proportion vs percent), rounding, ordering, and missing-value representation, without ever loading the template and diffing structure against it; the answer describes the chosen format instead of proving it matches.
- **Detection procedure**:
  1. Read the task and note that a template/reference file exists and defines the expected schema (column names/order, index column presence, units/scale, decimal precision, blank vs `NaN` encoding).
  2. Search the scripts for any read of the template file and an explicit comparison (e.g., asserting equal column lists, shape, index name, or a value-range check) before/after saving.
  3. If no such load-and-compare exists, check whether the answer justifies each formatting choice with evidence from the template (a printed head of the template, a shape/dtype match) rather than with phrasing like "values are proportions 0–1" or "NaN as empty cells".
  4. Sanity-check the described values against the template's own scale and precision (e.g., 0–1 vs 0–100, unrounded floats vs 1–2 decimals) and against expected shape/row count.
- **Discriminator**: A real violation is unverified self-invented formatting — no template read, no assertion, no printed side-by-side comparison. It is fine if the script loads the template (or its header) and asserts matching columns/index/shape/scale, or if the answer shows the template's first rows next to the produced output, even if the format was ultimately simple.
- **Consequence**: The file exists and the numbers may be conceptually right, but the automated check comparing it to the expected file fails on schema, scale, rounding, or missing-value encoding, scoring 0.
599Unverified construction of the target quantity (wrong unit of analysis / aggregation)taskinfiagent-dabench
Applies when
task -- the statistic is requested on a quantity that is not a ready-made column, so the script must derive it (aggregate rows by an ID, difference two timestamps, use a raw per-record field, etc.).
Pattern
The script silently commits to one derivation (e.g., groupby(id).sum() of a per-row field) without inspecting whether the dataset already contains the requested quantity, whether the correct definition is a difference of start/end fields, or whether the analysis unit is the row rather than the group; downstream outlier filtering and statistics are then computed on a different population than the task intends.
Detection procedure
  1. Read the task and note the exact entity whose values are to be summarized (row-level record vs. grouped entity) and the count of values that should result.
  2. Read the script's data-loading/exploration output and check whether an existing column (or an explicit start/end pair) directly encodes the requested quantity, and whether the script compared that against its chosen derivation.
  3. Check whether the script justified/sanity-checked the derived series: number of values, plausible min/max/units, and consistency with any alternative derivation available in the data.
  4. Compare the reported statistics to the raw-data summary printed by the script; if the derived series' scale or count differs sharply from the natural alternative and no reconciliation was done, flag it.
Discriminator
A real violation is choosing one of several plausible definitions with no cross-check or evidence; it is fine if the script explicitly inspects the schema, shows that only one definition is consistent with the task's wording/units, or computes both and reports they agree.
Consequence
Every downstream number (outlier set, filtered mean, standard deviation) is computed on the wrong population, so the reported values are off by a large factor and all graded checks fail even though the filtering/rounding code is correct.
id bcb614556906 · mined from infiagent-dabench dabench-619@s11
raw text (what the judge reads)
### Unverified construction of the target quantity (wrong unit of analysis / aggregation)
- **Applies when**: `task` -- the statistic is requested on a quantity that is not a ready-made column, so the script must derive it (aggregate rows by an ID, difference two timestamps, use a raw per-record field, etc.).
- **Pattern**: The script silently commits to one derivation (e.g., `groupby(id).sum()` of a per-row field) without inspecting whether the dataset already contains the requested quantity, whether the correct definition is a difference of start/end fields, or whether the analysis unit is the row rather than the group; downstream outlier filtering and statistics are then computed on a different population than the task intends.
- **Detection procedure**:
  1. Read the task and note the exact entity whose values are to be summarized (row-level record vs. grouped entity) and the count of values that should result.
  2. Read the script's data-loading/exploration output and check whether an existing column (or an explicit start/end pair) directly encodes the requested quantity, and whether the script compared that against its chosen derivation.
  3. Check whether the script justified/sanity-checked the derived series: number of values, plausible min/max/units, and consistency with any alternative derivation available in the data.
  4. Compare the reported statistics to the raw-data summary printed by the script; if the derived series' scale or count differs sharply from the natural alternative and no reconciliation was done, flag it.
- **Discriminator**: A real violation is choosing one of several plausible definitions with no cross-check or evidence; it is fine if the script explicitly inspects the schema, shows that only one definition is consistent with the task's wording/units, or computes both and reports they agree.
- **Consequence**: Every downstream number (outlier set, filtered mean, standard deviation) is computed on the wrong population, so the reported values are off by a large factor and all graded checks fail even though the filtering/rounding code is correct.
600Predictions never validated against a holdout baseline or the target's own distributiontaskda-code
Applies when
task -- the task asks for predicted values on an unlabeled split and the script trains a model, writes the prediction file, and reports only the file's summary statistics.
Pattern
The attempt picks a feature set and model, skips any train/validation split scoring (no RMSE/MAE/R² vs a mean-predictor or simple baseline), and does not compare the predicted value distribution to the label distribution in the training data; the resulting predictions are heavily shrunk toward the training mean (tiny spread, narrow min–max) and/or ignore high-signal columns available in both splits, yet are reported as a success.
Detection procedure
  1. Read the task to confirm the deliverable is per-row predicted values scored against hidden labels, and note the format constraints (row count, column name, order).
  2. In the scripts, look for (a) a held-out or cross-validated score printed and compared to a trivial baseline, and (b) a check that the columns used exist and are used identically for train and test rows; note whether informative columns present in the test file were dropped without justification.
  3. In the answer/output, compare the reported prediction spread (std, min, max) with the training label spread; a std that is a small fraction of the label std, or a range far narrower than the label range, indicates mean-collapse.
  4. Flag if no validation number is reported anywhere, or if the validation number is not better than the baseline, or if the distribution check in step 3 fails.
Discriminator
A genuinely near-constant prediction set is acceptable only if the script shows a holdout score that beats the mean-predictor baseline (features really are weakly predictive and shrinkage is optimal); it is a violation when no such evidence exists, or when strong predictors available in both splits were silently discarded while predictions collapse to the mean.
Consequence
The submitted prediction file scores at or below the trivial mean-prediction baseline on the grader's error/correlation threshold, so the expected output file is marked wrong even though its shape and column name look correct.
id c2a4cce43fc0 · mined from da-code dacode-ml-regression-004@s11
raw text (what the judge reads)
### Predictions never validated against a holdout baseline or the target's own distribution
- **Applies when**: `task` -- the task asks for predicted values on an unlabeled split and the script trains a model, writes the prediction file, and reports only the file's summary statistics.
- **Pattern**: The attempt picks a feature set and model, skips any train/validation split scoring (no RMSE/MAE/R² vs a mean-predictor or simple baseline), and does not compare the predicted value distribution to the label distribution in the training data; the resulting predictions are heavily shrunk toward the training mean (tiny spread, narrow min–max) and/or ignore high-signal columns available in both splits, yet are reported as a success.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is per-row predicted values scored against hidden labels, and note the format constraints (row count, column name, order).
  2. In the scripts, look for (a) a held-out or cross-validated score printed and compared to a trivial baseline, and (b) a check that the columns used exist and are used identically for train and test rows; note whether informative columns present in the test file were dropped without justification.
  3. In the answer/output, compare the reported prediction spread (std, min, max) with the training label spread; a std that is a small fraction of the label std, or a range far narrower than the label range, indicates mean-collapse.
  4. Flag if no validation number is reported anywhere, or if the validation number is not better than the baseline, or if the distribution check in step 3 fails.
- **Discriminator**: A genuinely near-constant prediction set is acceptable only if the script shows a holdout score that beats the mean-predictor baseline (features really are weakly predictive and shrinkage is optimal); it is a violation when no such evidence exists, or when strong predictors available in both splits were silently discarded while predictions collapse to the mean.
- **Consequence**: The submitted prediction file scores at or below the trivial mean-prediction baseline on the grader's error/correlation threshold, so the expected output file is marked wrong even though its shape and column name look correct.
601Ambiguous statistic silently redefined by arbitrary data subsettingtaskinfiagent-dabench
Applies when
task -- the requested statistic cannot be computed literally on the stated slice of data (e.g. a dispersion/shape measure asked for a single row/column/period), so the script must choose which records or fields form the sample.
Pattern
The agent invents one interpretation (e.g. truncating a series at the named point, or dropping part of the available rows/files/columns) without stating why, computes the statistic only on that hand-picked subset, and reports its top-ranked entity as the answer — while other equally plausible subsets (full series, all groups, all source files) give a different winner. Often the same script also asserts a library flag's meaning from memory rather than the docs to "satisfy" the stated definition constraint.
Detection procedure
  1. Read the task and check whether the requested statistic is well-defined on the slice named; if not, note that an interpretation choice is unavoidable.
  2. In the scripts, locate the exact set of rows/columns/files fed into the statistic and ask whether that choice is derived from the task wording or hard-coded by the agent (look for hand-typed column lists, single-file loads when siblings exist, filters not mentioned in the task).
  3. Check whether the script enumerates the plausible interpretations, compares their rankings, and gives a reason for the one reported — and whether any library option used (bias/ddof/definition flags) is verified against documentation rather than a comment claim.
  4. Compare the reported answer to the rankings produced by the alternative interpretations that the script itself computed; if they disagree and the reported one is the narrower, unjustified subset, flag it.
Discriminator
Fine if the task text (or an explicit, documented convention) uniquely pins down the sample, or if all candidate interpretations yield the same answer and the script shows that. A violation is when the interpretations disagree, the chosen one drops available data with no justification, and the agent commits to it anyway.
Consequence
The reported entity is the argmax under a non-canonical sample, so the single-value grader check fails against the expected entity even though the code runs cleanly.
id b3050bee60b4 · mined from infiagent-dabench dabench-252@s11
raw text (what the judge reads)
### Ambiguous statistic silently redefined by arbitrary data subsetting
- **Applies when**: `task` -- the requested statistic cannot be computed literally on the stated slice of data (e.g. a dispersion/shape measure asked for a single row/column/period), so the script must choose which records or fields form the sample.
- **Pattern**: The agent invents one interpretation (e.g. truncating a series at the named point, or dropping part of the available rows/files/columns) without stating why, computes the statistic only on that hand-picked subset, and reports its top-ranked entity as the answer — while other equally plausible subsets (full series, all groups, all source files) give a different winner. Often the same script also asserts a library flag's meaning from memory rather than the docs to "satisfy" the stated definition constraint.
- **Detection procedure**:
  1. Read the task and check whether the requested statistic is well-defined on the slice named; if not, note that an interpretation choice is unavoidable.
  2. In the scripts, locate the exact set of rows/columns/files fed into the statistic and ask whether that choice is derived from the task wording or hard-coded by the agent (look for hand-typed column lists, single-file loads when siblings exist, filters not mentioned in the task).
  3. Check whether the script enumerates the plausible interpretations, compares their rankings, and gives a reason for the one reported — and whether any library option used (bias/ddof/definition flags) is verified against documentation rather than a comment claim.
  4. Compare the reported answer to the rankings produced by the alternative interpretations that the script itself computed; if they disagree and the reported one is the narrower, unjustified subset, flag it.
- **Discriminator**: Fine if the task text (or an explicit, documented convention) uniquely pins down the sample, or if all candidate interpretations yield the same answer and the script shows that. A violation is when the interpretations disagree, the chosen one drops available data with no justification, and the agent commits to it anyway.
- **Consequence**: The reported entity is the argmax under a non-canonical sample, so the single-value grader check fails against the expected entity even though the code runs cleanly.
602Reporting a value at coarser granularity than the underlying result (destructive reformatting to fit a template)taskinfiagent-dabench
Applies when
task -- The task asks the agent to identify a specific record/label (a date, ID, category) and gives a format template or example whose apparent granularity is coarser or otherwise ambiguous relative to the value actually computed.
Pattern
The script correctly locates the record, then applies a lossy transformation (truncating, splitting, rounding, relabeling) purely to match the literal format placeholder, and submits the reduced value instead of the identifier that uniquely determines the computation it just performed.
Detection procedure
  1. Read the task and note what the identified value must be to make the rest of the computation well-defined (e.g. it must pin down a single row/observation, not a group of them).
  2. Read the scripts and find where the identified value is converted before being printed as the final answer; check whether that conversion discards information (string splitting, truncation, aggregation to a coarser key).
  3. Compare the submitted identifier with the value actually used internally for the downstream calculation; if the submitted one no longer maps back to a unique record, flag it.
  4. Check whether the agent considered the alternative (submitting the full-precision value) or blindly trusted the placeholder; absence of any note/sanity check is additional evidence.
Discriminator
A real violation is when the reported identifier is strictly less informative than the one the analysis depended on (multiple records share it). It is fine when the coarser form is genuinely unique in the data, or when the task explicitly demands aggregation at that granularity (e.g. "report the month with the highest monthly total") and the computation was actually done at that level.
Consequence
The numeric part of the answer may match, but the identifier check fails on exact-string comparison, so the attempt is scored partially correct / incorrect overall.
id c31cb988d353 · mined from infiagent-dabench dabench-572@s11
raw text (what the judge reads)
### Reporting a value at coarser granularity than the underlying result (destructive reformatting to fit a template)
- **Applies when**: `task` -- The task asks the agent to identify a specific record/label (a date, ID, category) and gives a format template or example whose apparent granularity is coarser or otherwise ambiguous relative to the value actually computed.
- **Pattern**: The script correctly locates the record, then applies a lossy transformation (truncating, splitting, rounding, relabeling) purely to match the literal format placeholder, and submits the reduced value instead of the identifier that uniquely determines the computation it just performed.
- **Detection procedure**:
  1. Read the task and note what the identified value must be to make the rest of the computation well-defined (e.g. it must pin down a single row/observation, not a group of them).
  2. Read the scripts and find where the identified value is converted before being printed as the final answer; check whether that conversion discards information (string splitting, truncation, aggregation to a coarser key).
  3. Compare the submitted identifier with the value actually used internally for the downstream calculation; if the submitted one no longer maps back to a unique record, flag it.
  4. Check whether the agent considered the alternative (submitting the full-precision value) or blindly trusted the placeholder; absence of any note/sanity check is additional evidence.
- **Discriminator**: A real violation is when the reported identifier is strictly less informative than the one the analysis depended on (multiple records share it). It is fine when the coarser form is genuinely unique in the data, or when the task explicitly demands aggregation at that granularity (e.g. "report the month with the highest monthly total") and the computation was actually done at that level.
- **Consequence**: The numeric part of the answer may match, but the identifier check fails on exact-string comparison, so the attempt is scored partially correct / incorrect overall.
603Fabricated/guessed inputs instead of the actual provided datataskda-code
Applies when
task -- the task references a specific dataset and the script contains fallback data generation, random seeds, or heuristic keyword/positional guessing to pick the variables to analyze.
Pattern
The script never verifies that the real input file and the required variables were actually found; if discovery fails it silently substitutes simulated data or the "first two numeric columns," then reports a number computed from inputs that may have nothing to do with the requested quantities (wrong regressor/response, wrong direction of the relationship, unfiltered date range, wrong aggregation).
Detection procedure
  1. Read the task and list the exact required inputs: the source file, the dependent and independent variables, and any stated filtering/aggregation (e.g., period restriction, per-period aggregation such as min/mean).
  2. Scan the script for create_synthetic_data, np.random, broad try/except: pass, or column selection by keyword substring / positional index, and check whether any of these can execute without raising an error.
  3. Check whether the script asserts the loaded file's identity, column names, row count, and applies the stated filtering/aggregation before fitting; also check the dependent/independent roles are not swapped by the heuristic.
  4. Compare the reported number and printed data preview against the real dataset's expected shape/period count; if the preview shows generated or mismatched columns, or the roles are ambiguous, flag it.
Discriminator
A robust attempt hard-fails (raises/exits) when the expected file or columns are absent and prints an explicit mapping of task variables to dataset columns plus the applied filters; a violation proceeds anyway on synthetic or arbitrarily chosen columns, or on unaggregated/unfiltered rows, producing a plausible-looking scalar.
Consequence
The output file has the right format but a numerically wrong value, so the grader's value comparison fails (0/1 checks passed).
id 38d3dc140560 · mined from da-code dacode-data-sa-043@s11
raw text (what the judge reads)
### Fabricated/guessed inputs instead of the actual provided data
- **Applies when**: `task` -- the task references a specific dataset and the script contains fallback data generation, random seeds, or heuristic keyword/positional guessing to pick the variables to analyze.
- **Pattern**: The script never verifies that the real input file and the required variables were actually found; if discovery fails it silently substitutes simulated data or the "first two numeric columns," then reports a number computed from inputs that may have nothing to do with the requested quantities (wrong regressor/response, wrong direction of the relationship, unfiltered date range, wrong aggregation).
- **Detection procedure**:
  1. Read the task and list the exact required inputs: the source file, the dependent and independent variables, and any stated filtering/aggregation (e.g., period restriction, per-period aggregation such as min/mean).
  2. Scan the script for `create_synthetic_data`, `np.random`, broad `try/except: pass`, or column selection by keyword substring / positional index, and check whether any of these can execute without raising an error.
  3. Check whether the script asserts the loaded file's identity, column names, row count, and applies the stated filtering/aggregation before fitting; also check the dependent/independent roles are not swapped by the heuristic.
  4. Compare the reported number and printed data preview against the real dataset's expected shape/period count; if the preview shows generated or mismatched columns, or the roles are ambiguous, flag it.
- **Discriminator**: A robust attempt hard-fails (raises/exits) when the expected file or columns are absent and prints an explicit mapping of task variables to dataset columns plus the applied filters; a violation proceeds anyway on synthetic or arbitrarily chosen columns, or on unaggregated/unfiltered rows, producing a plausible-looking scalar.
- **Consequence**: The output file has the right format but a numerically wrong value, so the grader's value comparison fails (0/1 checks passed).
604Minority-class prediction rate not sanity-checked against training prevalence and the stated cost asymmetrytaskda-code
Applies when
task -- The task asks for per-row class labels on a held-out set for a rare-event/imbalanced problem where missed positives are stated to be costlier than false alarms.
Pattern
The attempt outputs labels from a default 0.5 threshold (or an untuned/underfit model) without ever comparing the predicted positive rate to the positive rate in the training labels, and without tuning the decision threshold to the recall-weighted objective; the submitted column is overwhelmingly the majority class with far fewer positives than prevalence implies.
Detection procedure
1. Read the task for the class balance and the stated cost/metric emphasis (e.g., recall/F1 on the rare class, penalty for false negatives). 2. In the scripts, look for (a) a validation split with the rare-class metric reported, (b) class-imbalance handling (class weights/resampling) and (c) an explicit threshold choice justified by that metric — flag if any is absent. 3. Compute the fraction of positives in the submitted file and compare with the training positive rate; flag if it is materially lower (e.g., less than ~half) or if positives are near zero. 4. Confirm row count and header match the provided sample format.
Discriminator
A real violation is a submission whose positive rate is far below training prevalence with no validation evidence (no held-out rare-class recall/F1 figure) justifying the conservatism; a look-alike that is fine is a submission with a similarly low positive count that is backed by reported held-out metrics showing that threshold maximizes the requested cost-sensitive objective.
Consequence
The grader's rare-class recall (and any cost-weighted score) collapses toward the trivial all-majority baseline, so the prediction file fails the accuracy/score threshold and is marked WRONG.
id 579c05df3299 · mined from da-code dacode-ml-binary-013@s11
raw text (what the judge reads)
### Minority-class prediction rate not sanity-checked against training prevalence and the stated cost asymmetry
- **Applies when**: `task` -- The task asks for per-row class labels on a held-out set for a rare-event/imbalanced problem where missed positives are stated to be costlier than false alarms.
- **Pattern**: The attempt outputs labels from a default 0.5 threshold (or an untuned/underfit model) without ever comparing the predicted positive rate to the positive rate in the training labels, and without tuning the decision threshold to the recall-weighted objective; the submitted column is overwhelmingly the majority class with far fewer positives than prevalence implies.
- **Detection procedure**: 1. Read the task for the class balance and the stated cost/metric emphasis (e.g., recall/F1 on the rare class, penalty for false negatives). 2. In the scripts, look for (a) a validation split with the rare-class metric reported, (b) class-imbalance handling (class weights/resampling) and (c) an explicit threshold choice justified by that metric — flag if any is absent. 3. Compute the fraction of positives in the submitted file and compare with the training positive rate; flag if it is materially lower (e.g., less than ~half) or if positives are near zero. 4. Confirm row count and header match the provided sample format.
- **Discriminator**: A real violation is a submission whose positive rate is far below training prevalence with no validation evidence (no held-out rare-class recall/F1 figure) justifying the conservatism; a look-alike that is fine is a submission with a similarly low positive count that is backed by reported held-out metrics showing that threshold maximizes the requested cost-sensitive objective.
- **Consequence**: The grader's rare-class recall (and any cost-weighted score) collapses toward the trivial all-majority baseline, so the prediction file fails the accuracy/score threshold and is marked WRONG.
605Population scope narrower than the task's stated universetaskinfiagent-dabench
Applies when
task -- the question asks for a statistic or set of records over an entire population ("all countries", "all users", "the full dataset"), while the data directory holds several partial/segmented files or the script filters/loads only one slice.
Pattern
The script loads a single file (or one subgroup) that covers only part of the requested universe and computes the requested quantity on that subset, never verifying that the loaded rows cover every entity the task refers to; the reported answer is therefore a subset-level result presented as a population-level one, and thresholds/quantiles derived from it are computed on the wrong denominator.
Detection procedure
  1. Read the task and write down the intended unit and coverage (which entities, which time slice, how many rows should be in scope).
  2. In the script, find every data-loading and filtering line and determine exactly which entities are present; check whether other files/segments in the data location belong to the same universe and were skipped.
  3. Confirm the script prints a coverage sanity check (row count / unique-entity count / list of source files) and that this count matches the expected population size; absence of such a check with a single-segment load is a flag.
  4. Check the final answer was derived from the full-coverage frame, not from the partial one, and matches the requested format/entity names.
Discriminator
A real violation is loading/using fewer entities than the task's universe (or not proving coverage) when additional relevant segments exist; it is fine if the single file provably contains the whole universe, or if the task itself restricts the scope to that segment.
Consequence
Quartiles/aggregates and thus the flagged set are computed from a truncated population, so the reported entities are incomplete or spurious and the grader marks the answer WRONG/MISSING even when some names coincide.
id 4e7cae44c9bd · mined from infiagent-dabench dabench-254@s11
raw text (what the judge reads)
### Population scope narrower than the task's stated universe
- **Applies when**: `task` -- the question asks for a statistic or set of records over an entire population ("all countries", "all users", "the full dataset"), while the data directory holds several partial/segmented files or the script filters/loads only one slice.
- **Pattern**: The script loads a single file (or one subgroup) that covers only part of the requested universe and computes the requested quantity on that subset, never verifying that the loaded rows cover every entity the task refers to; the reported answer is therefore a subset-level result presented as a population-level one, and thresholds/quantiles derived from it are computed on the wrong denominator.
- **Detection procedure**:
  1. Read the task and write down the intended unit and coverage (which entities, which time slice, how many rows should be in scope).
  2. In the script, find every data-loading and filtering line and determine exactly which entities are present; check whether other files/segments in the data location belong to the same universe and were skipped.
  3. Confirm the script prints a coverage sanity check (row count / unique-entity count / list of source files) and that this count matches the expected population size; absence of such a check with a single-segment load is a flag.
  4. Check the final answer was derived from the full-coverage frame, not from the partial one, and matches the requested format/entity names.
- **Discriminator**: A real violation is loading/using fewer entities than the task's universe (or not proving coverage) when additional relevant segments exist; it is fine if the single file provably contains the whole universe, or if the task itself restricts the scope to that segment.
- **Consequence**: Quartiles/aggregates and thus the flagged set are computed from a truncated population, so the reported entities are incomplete or spurious and the grader marks the answer WRONG/MISSING even when some names coincide.
606Prediction file schema/label-encoding never verified against the source datataskda-code
Applies when
task -- the task asks for a prediction file with a specified column name, and the target is a categorical/encoded label whose exact string or numeric representation comes from the training data.
Pattern
The attempt trains a model and writes the submission file, but never checks that the written file matches the required schema — exact column name (case/spacing), one row per test row in original test order, no extra index column, and label values written in exactly the same vocabulary/encoding as the target column in the training file (e.g. writing decoded strings when integer codes are expected, or vice-versa, or re-labelled/re-cased variants). Quality is also asserted from the class-balance printout rather than from any held-out score.
Detection procedure
1. Read the task for the required file name, column name, and any implied label representation; note the exact unique values of the target as they appear in the training data. 2. In the scripts, trace how the target was encoded for training and how predictions were converted back before to_csv — check for LabelEncoder/map round-trips, index=False, and that the frame has exactly the number of test rows in unchanged order. 3. Compare the answer's reported column name, row count, and reported class labels against step 1; treat any renaming, re-casing, or type change as unjustified. 4. Check whether any held-out/CV accuracy is reported to confirm the model itself is sane, rather than only a predicted-class distribution.
Discriminator
A real violation is an unverified or transformed output representation (labels/column/row alignment never compared back to the source target values, or an index column emitted); it is fine if the script explicitly reads the training target's unique values (or a provided sample submission) and asserts the output's column name, length, and value set match them.
Consequence
The grader reads result.csv and finds the column/label values or row alignment incompatible with the reference, scoring the file as WRONG/MISSING even though the underlying model may be accurate.
id e55809f63578 · mined from da-code dacode-ml-binary-009@s11
raw text (what the judge reads)
### Prediction file schema/label-encoding never verified against the source data
- **Applies when**: `task` -- the task asks for a prediction file with a specified column name, and the target is a categorical/encoded label whose exact string or numeric representation comes from the training data.
- **Pattern**: The attempt trains a model and writes the submission file, but never checks that the written file matches the required schema — exact column name (case/spacing), one row per test row in original test order, no extra index column, and label values written in exactly the same vocabulary/encoding as the target column in the training file (e.g. writing decoded strings when integer codes are expected, or vice-versa, or re-labelled/re-cased variants). Quality is also asserted from the class-balance printout rather than from any held-out score.
- **Detection procedure**: 1. Read the task for the required file name, column name, and any implied label representation; note the exact unique values of the target as they appear in the training data. 2. In the scripts, trace how the target was encoded for training and how predictions were converted back before `to_csv` — check for `LabelEncoder`/`map` round-trips, `index=False`, and that the frame has exactly the number of test rows in unchanged order. 3. Compare the answer's reported column name, row count, and reported class labels against step 1; treat any renaming, re-casing, or type change as unjustified. 4. Check whether any held-out/CV accuracy is reported to confirm the model itself is sane, rather than only a predicted-class distribution.
- **Discriminator**: A real violation is an unverified or transformed output representation (labels/column/row alignment never compared back to the source target values, or an index column emitted); it is fine if the script explicitly reads the training target's unique values (or a provided sample submission) and asserts the output's column name, length, and value set match them.
- **Consequence**: The grader reads `result.csv` and finds the column/label values or row alignment incompatible with the reference, scoring the file as WRONG/MISSING even though the underlying model may be accurate.
607Fabricated performance metric computed on only one side of the paired recordstaskda-code
Applies when
task -- the task asks to visualize/rank "performance" of entities that appear in multiple role columns (e.g., two participant columns per row) and the scripts must define an aggregate score from the raw records.
Pattern
The agent invents an ad-hoc formula that has no basis in the task or config (e.g., counts of rows plus an arbitrary weight on an unrelated flag), aggregates by only one of the role columns so half of each entity's records are silently dropped, and never checks that the invented components actually change the result or that the output artifacts requested by the environment (extra saved data/config files) are produced.
Detection procedure
  1. Read the task and any provided config/spec file for the definition of the quantity to be plotted (label names, expected result files, units); note whether a standard domain definition exists (e.g., wins/points/goal difference computed over all records involving the entity).
  2. In the scripts, check the aggregation key(s): does it group by a single role column while the entity also appears in the other role column? Does the scoring formula come from the spec or is it improvised?
  3. Check each formula term for a validity/sanity test: dtype-correct comparisons, non-zero contribution, plausible value ranges and counts per entity; also confirm every expected output artifact is written, not just the image.
  4. Compare the answer's reported numbers against a quick independent expectation (e.g., known strong entities should rank high; totals should roughly double when both roles are counted) — a table where an added term contributes exactly 0 for every entity is a red flag.
Discriminator
A real violation is when the score's definition is unjustified by the task/spec or is computed over a strict subset of the entity's records; it is not a violation if the spec (or task text) explicitly defines the metric that way, or if the entity genuinely appears in only one role column.
Consequence
The plotted values, ordering and entity set differ from the reference, so the saved chart and the accompanying numeric/config result files all mismatch and every grader check fails.
id 33d2a44f2a67 · mined from da-code dacode-plot-bar-006@s11
raw text (what the judge reads)
### Fabricated performance metric computed on only one side of the paired records
- **Applies when**: `task` -- the task asks to visualize/rank "performance" of entities that appear in multiple role columns (e.g., two participant columns per row) and the scripts must define an aggregate score from the raw records.
- **Pattern**: The agent invents an ad-hoc formula that has no basis in the task or config (e.g., counts of rows plus an arbitrary weight on an unrelated flag), aggregates by only one of the role columns so half of each entity's records are silently dropped, and never checks that the invented components actually change the result or that the output artifacts requested by the environment (extra saved data/config files) are produced.
- **Detection procedure**:
  1. Read the task and any provided config/spec file for the definition of the quantity to be plotted (label names, expected result files, units); note whether a standard domain definition exists (e.g., wins/points/goal difference computed over all records involving the entity).
  2. In the scripts, check the aggregation key(s): does it group by a single role column while the entity also appears in the other role column? Does the scoring formula come from the spec or is it improvised?
  3. Check each formula term for a validity/sanity test: dtype-correct comparisons, non-zero contribution, plausible value ranges and counts per entity; also confirm every expected output artifact is written, not just the image.
  4. Compare the answer's reported numbers against a quick independent expectation (e.g., known strong entities should rank high; totals should roughly double when both roles are counted) — a table where an added term contributes exactly 0 for every entity is a red flag.
- **Discriminator**: A real violation is when the score's definition is unjustified by the task/spec or is computed over a strict subset of the entity's records; it is *not* a violation if the spec (or task text) explicitly defines the metric that way, or if the entity genuinely appears in only one role column.
- **Consequence**: The plotted values, ordering and entity set differ from the reference, so the saved chart and the accompanying numeric/config result files all mismatch and every grader check fails.
608Threshold-based counts reported without cross-validation of the population, estimator convention, or plausibilitytaskinfiagent-dabench
Applies when
task -- the task asks for a count (or filtered dataframe) obtained by applying a fixed numeric threshold to a derived statistic (z-score, ratio, residual, percentile) computed from a column's own summary statistics.
Pattern
The script silently changes the population the statistic is computed over (e.g., dropping missing rows, filtering, or reading the column with an unchecked dtype/parse) and uses one particular estimator convention (e.g., population vs. sample standard deviation, mean over the wrong subset) without ever comparing against the straightforward whole-column computation or checking whether the resulting count is plausible for the threshold used.
Detection procedure
  1. Read the task and note exactly which rows/values the statistic is supposed to be computed over and reported for (all rows as loaded, unless the task says otherwise).
  2. In the scripts, trace every step between loading and the threshold comparison — look for dropna, filtering, type coercion, re-indexing, or a library call whose default differs from the obvious alternative (e.g., scipy.stats.zscore with ddof=0 vs. df.std() with ddof=1) — and check whether the script justifies or at least reproduces the result the other way.
  3. Check whether the script sanity-checks the magnitude of the answer: for a 3-sigma rule the flagged fraction should be a fraction of a percent; a count that is a large share of the rows should have triggered an inspection of the column's dtype, units, duplicates, or distribution before reporting.
  4. Confirm the reported number is the count the task asked for, computed on the same frame the task described, not on an internally reduced frame.
Discriminator
A genuine violation is when the row set or estimator was altered implicitly and the count was reported with no second computation or plausibility check. It is not a violation if the script explicitly shows there are no missing/invalid values (so the subset equals the full column), or if it reports both estimator conventions/frames and shows they agree, or if a large flagged fraction is explained by demonstrated heavy skew in the data.
Consequence
The reported count differs from the reference count computed on the full, correctly typed column (often by tens or hundreds of rows, or the difference between zero and a nonzero count), so the single graded value @outlier_count[...] fails the exact-match check.
id 7cab91617839 · mined from infiagent-dabench dabench-361@s11
raw text (what the judge reads)
### Threshold-based counts reported without cross-validation of the population, estimator convention, or plausibility

- **Applies when**: `task` -- the task asks for a count (or filtered dataframe) obtained by applying a fixed numeric threshold to a derived statistic (z-score, ratio, residual, percentile) computed from a column's own summary statistics.
- **Pattern**: The script silently changes the population the statistic is computed over (e.g., dropping missing rows, filtering, or reading the column with an unchecked dtype/parse) and uses one particular estimator convention (e.g., population vs. sample standard deviation, mean over the wrong subset) without ever comparing against the straightforward whole-column computation or checking whether the resulting count is plausible for the threshold used.
- **Detection procedure**:
  1. Read the task and note exactly which rows/values the statistic is supposed to be computed over and reported for (all rows as loaded, unless the task says otherwise).
  2. In the scripts, trace every step between loading and the threshold comparison — look for `dropna`, filtering, type coercion, re-indexing, or a library call whose default differs from the obvious alternative (e.g., `scipy.stats.zscore` with ddof=0 vs. `df.std()` with ddof=1) — and check whether the script justifies or at least reproduces the result the other way.
  3. Check whether the script sanity-checks the magnitude of the answer: for a 3-sigma rule the flagged fraction should be a fraction of a percent; a count that is a large share of the rows should have triggered an inspection of the column's dtype, units, duplicates, or distribution before reporting.
  4. Confirm the reported number is the count the task asked for, computed on the same frame the task described, not on an internally reduced frame.
- **Discriminator**: A genuine violation is when the row set or estimator was altered implicitly and the count was reported with no second computation or plausibility check. It is *not* a violation if the script explicitly shows there are no missing/invalid values (so the subset equals the full column), or if it reports both estimator conventions/frames and shows they agree, or if a large flagged fraction is explained by demonstrated heavy skew in the data.
- **Consequence**: The reported count differs from the reference count computed on the full, correctly typed column (often by tens or hundreds of rows, or the difference between zero and a nonzero count), so the single graded value `@outlier_count[...]` fails the exact-match check.
609Missing required output artifact / no reproducible scripttaskda-code
Applies when
task -- The task (or its README/tips file) specifies deliverables such as a result file, a fixed answer schema, or a documented mapping/transformation to apply before computing the statistic.
Pattern
The agent reports the answer only as chat text, without saving the required output file (e.g. result.json) and without leaving any script that reads the auxiliary instructions file, applies the prescribed transformation, and writes the deliverable — so the answer cannot be verified or graded even if the number happens to be plausible.
Detection procedure
  1. Read the task and every referenced auxiliary file (README, tips/notes) and list all required artifacts, filenames, key names, and prescribed transformations/roundings.
  2. Check the saved scripts: is there code that loads the data, explicitly applies the stated mapping/filter, computes the requested statistic, and writes each required artifact to the exact filename/format?
  3. Check the final answer: does it match the required schema, and does every value in it trace to a line of executed code that also persisted it to disk?
  4. Flag if any required file is absent, or if the mapping/label-normalization step cannot be pointed to in code (an unmapped or partially mapped category can also change which category is most frequent and its ratio).
Discriminator
A real violation is when the deliverable file is missing/unwritten or the prescribed transformation is nowhere in the code; it is not a violation if the file is written under the exact requested name and the mapping is applied in code, even if the agent additionally restates the answer in chat or uses a different but equivalent implementation.
Consequence
The grader looks for the expected result file, finds it missing (or containing untransformed/mismatched category labels and ratio), and scores 0 regardless of what was typed in the response.
id b99fab8b3b70 · mined from da-code dacode-di-text-004@s11
raw text (what the judge reads)
### Missing required output artifact / no reproducible script
- **Applies when**: `task` -- The task (or its README/tips file) specifies deliverables such as a result file, a fixed answer schema, or a documented mapping/transformation to apply before computing the statistic.
- **Pattern**: The agent reports the answer only as chat text, without saving the required output file (e.g. `result.json`) and without leaving any script that reads the auxiliary instructions file, applies the prescribed transformation, and writes the deliverable — so the answer cannot be verified or graded even if the number happens to be plausible.
- **Detection procedure**:
  1. Read the task and every referenced auxiliary file (README, tips/notes) and list all required artifacts, filenames, key names, and prescribed transformations/roundings.
  2. Check the saved scripts: is there code that loads the data, explicitly applies the stated mapping/filter, computes the requested statistic, and writes each required artifact to the exact filename/format?
  3. Check the final answer: does it match the required schema, and does every value in it trace to a line of executed code that also persisted it to disk?
  4. Flag if any required file is absent, or if the mapping/label-normalization step cannot be pointed to in code (an unmapped or partially mapped category can also change which category is most frequent and its ratio).
- **Discriminator**: A real violation is when the deliverable file is missing/unwritten or the prescribed transformation is nowhere in the code; it is *not* a violation if the file is written under the exact requested name and the mapping is applied in code, even if the agent additionally restates the answer in chat or uses a different but equivalent implementation.
- **Consequence**: The grader looks for the expected result file, finds it missing (or containing untransformed/mismatched category labels and ratio), and scores 0 regardless of what was typed in the response.
610Predicted category labels not drawn verbatim from the training label vocabularytaskda-code
Applies when
task -- the task asks for predictions of a categorical target that already exists as a column in the provided training data, written to an output file.
Pattern
The attempt invents, abbreviates, re-cases, or re-maps class names (e.g., encoding to integers then decoding with hand-typed strings, or using a "cleaner" wording) instead of emitting the exact category strings present in the training column; sometimes the output is also truncated or has a different row count than the test set.
Detection procedure
  1. From the task/README, find the target column and read the exact set of distinct values it takes in the training file (including capitalization, spacing, hyphens, and any "unknown"/missing handling).
  2. Read the scripts for how predicted labels are produced and written: check whether the inverse mapping comes from the fitted encoder/classes_/original column values, or from a literal list typed by the agent.
  3. Compare the distinct values in the submitted output against the training vocabulary character-for-character; also compare output row count and ID set against the test file.
  4. Flag if any emitted label is absent from the training vocabulary, or if row count/IDs do not match the test set exactly.
Discriminator
A real violation is a label string that no exact-match lookup against the training column would find (a renamed/abbreviated/merged class), or a row/ID mismatch. It is not a violation if the strings match exactly but the predicted class distribution differs from training, or if the task explicitly requested a different encoding (e.g., integer codes) and the script follows that spec.
Consequence
Scoring joins predictions to ground truth by exact string match, so every row with a renamed class counts as wrong; accuracy collapses toward zero and the grader reports the result file as WRONG even if the underlying model is reasonable.
id 55323ba7277b · mined from da-code dacode-ml-multi-003@s11
raw text (what the judge reads)
### Predicted category labels not drawn verbatim from the training label vocabulary
- **Applies when**: `task` -- the task asks for predictions of a categorical target that already exists as a column in the provided training data, written to an output file.
- **Pattern**: The attempt invents, abbreviates, re-cases, or re-maps class names (e.g., encoding to integers then decoding with hand-typed strings, or using a "cleaner" wording) instead of emitting the exact category strings present in the training column; sometimes the output is also truncated or has a different row count than the test set.
- **Detection procedure**:
  1. From the task/README, find the target column and read the exact set of distinct values it takes in the training file (including capitalization, spacing, hyphens, and any "unknown"/missing handling).
  2. Read the scripts for how predicted labels are produced and written: check whether the inverse mapping comes from the fitted encoder/`classes_`/original column values, or from a literal list typed by the agent.
  3. Compare the distinct values in the submitted output against the training vocabulary character-for-character; also compare output row count and ID set against the test file.
  4. Flag if any emitted label is absent from the training vocabulary, or if row count/IDs do not match the test set exactly.
- **Discriminator**: A real violation is a label string that no exact-match lookup against the training column would find (a renamed/abbreviated/merged class), or a row/ID mismatch. It is *not* a violation if the strings match exactly but the predicted class distribution differs from training, or if the task explicitly requested a different encoding (e.g., integer codes) and the script follows that spec.
- **Consequence**: Scoring joins predictions to ground truth by exact string match, so every row with a renamed class counts as wrong; accuracy collapses toward zero and the grader reports the result file as WRONG even if the underlying model is reasonable.
611Output format never validated against the provided templatetaskda-code
Applies when
task -- the task says results must be saved to a file whose format matches a provided template/sample/schema file, and the scripts construct the output from scratch.
Pattern
The attempt builds the result table using its own assumed layout (index/column names, orientation, label encoding, rounding, ordering, empty-cell handling) and writes the file without ever reading the template, so it never checks that headers, row labels, shape, and value formatting agree.
Detection procedure
  1. Read the task and note that a template/example output file is supplied and that the deliverable must match it.
  2. Search the scripts for any read of that template (e.g., loading it and printing its columns/index/shape/dtypes) or any explicit comparison of the produced file's header and row labels to it.
  3. If absent, inspect the invented format choices (column naming scheme, row label type such as raw timestamps vs. formatted periods, transposition, rounding/precision, inclusion of an index column) and judge whether they were merely guessed.
  4. Check the final answer for evidence of a shape/header equality check against the template; a prose description of "the format" that only restates the agent's own choices is not evidence.
Discriminator
A real violation is guessing the layout with no reference to the template at any point; it is fine if the script loads the template (or reproduces its exact header/index) and asserts or prints a match, even if the writing code itself is hand-built.
Consequence
The saved file has correct-looking numbers but mismatched headers/labels/precision/orientation, so an automated comparison against the expected file reports the deliverable as WRONG/MISSING and the task scores 0.
id c8c9a33c5806 · mined from da-code dacode-dm-csv-044@s11
raw text (what the judge reads)
### Output format never validated against the provided template
- **Applies when**: `task` -- the task says results must be saved to a file whose format matches a provided template/sample/schema file, and the scripts construct the output from scratch.
- **Pattern**: The attempt builds the result table using its own assumed layout (index/column names, orientation, label encoding, rounding, ordering, empty-cell handling) and writes the file without ever reading the template, so it never checks that headers, row labels, shape, and value formatting agree.
- **Detection procedure**:
  1. Read the task and note that a template/example output file is supplied and that the deliverable must match it.
  2. Search the scripts for any read of that template (e.g., loading it and printing its columns/index/shape/dtypes) or any explicit comparison of the produced file's header and row labels to it.
  3. If absent, inspect the invented format choices (column naming scheme, row label type such as raw timestamps vs. formatted periods, transposition, rounding/precision, inclusion of an index column) and judge whether they were merely guessed.
  4. Check the final answer for evidence of a shape/header equality check against the template; a prose description of "the format" that only restates the agent's own choices is not evidence.
- **Discriminator**: A real violation is guessing the layout with no reference to the template at any point; it is fine if the script loads the template (or reproduces its exact header/index) and asserts or prints a match, even if the writing code itself is hand-built.
- **Consequence**: The saved file has correct-looking numbers but mismatched headers/labels/precision/orientation, so an automated comparison against the expected file reports the deliverable as WRONG/MISSING and the task scores 0.
612Group partition defined on a silently altered/filtered dataset (null-mask semantics)taskinfiagent-dabench
Applies when
task -- the task asks for statistics compared between two groups defined by missingness (or another predicate) in one column, computed over an entire loaded table.
Pattern
The attempt builds the two groups after implicitly changing what counts as "missing" or after dropping rows — e.g. reading the file with default na_values/keep_default_na that convert placeholder strings ("NA", "none", "-", "") into nulls or vice versa, calling dropna()/dropna(subset=...) on the numeric column, coercing dtypes with errors='coerce', or reading only some sheets/chunks/files — so each group's mean is computed on a subset that differs from the full table, giving means that are close to but not equal to the correct ones.
Detection procedure
  1. In the task, note the exact partition rule (null vs non-null in a given column) and that the statistic must cover all rows of the source data.
  2. In the scripts, trace the load step (parser options, na handling, dtype coercion, file/sheet/row selection) and every filter between load and grouping; check whether the mask is applied to the raw column values and whether any row-dropping happens before the mean.
  3. Verify the script prints/asserts len(group_a) + len(group_b) == len(full_table) and the group counts; if no such check exists, treat the partition as unvalidated.
  4. Check the reported means against the group counts/ranges: unexplained near-miss values (right magnitude, wrong digits) indicate rows were silently excluded or reclassified.
Discriminator
A real violation is any row-count reduction or missingness redefinition that is not required by the task statement; it is fine if rows are excluded because the task explicitly says so, or if a dtype conversion provably changes no values (verified by a count comparison before/after).
Consequence
The grader compares the reported group means to exact expected values; a partition off by even a few rows yields means that differ in the second decimal (e.g. 43.31 vs 45.48) and all mean-based checks fail, even though the significance conclusion looks plausible.
id 937dbb05ec53 · mined from infiagent-dabench dabench-297@s11
raw text (what the judge reads)
### Group partition defined on a silently altered/filtered dataset (null-mask semantics)
- **Applies when**: `task` -- the task asks for statistics compared between two groups defined by missingness (or another predicate) in one column, computed over an entire loaded table.
- **Pattern**: The attempt builds the two groups after implicitly changing what counts as "missing" or after dropping rows — e.g. reading the file with default `na_values`/`keep_default_na` that convert placeholder strings ("NA", "none", "-", "") into nulls or vice versa, calling `dropna()`/`dropna(subset=...)` on the numeric column, coercing dtypes with `errors='coerce'`, or reading only some sheets/chunks/files — so each group's mean is computed on a subset that differs from the full table, giving means that are close to but not equal to the correct ones.
- **Detection procedure**:
  1. In the task, note the exact partition rule (null vs non-null in a given column) and that the statistic must cover all rows of the source data.
  2. In the scripts, trace the load step (parser options, na handling, dtype coercion, file/sheet/row selection) and every filter between load and grouping; check whether the mask is applied to the raw column values and whether any row-dropping happens before the mean.
  3. Verify the script prints/asserts `len(group_a) + len(group_b) == len(full_table)` and the group counts; if no such check exists, treat the partition as unvalidated.
  4. Check the reported means against the group counts/ranges: unexplained near-miss values (right magnitude, wrong digits) indicate rows were silently excluded or reclassified.
- **Discriminator**: A real violation is any row-count reduction or missingness redefinition that is not required by the task statement; it is fine if rows are excluded because the task explicitly says so, or if a dtype conversion provably changes no values (verified by a count comparison before/after).
- **Consequence**: The grader compares the reported group means to exact expected values; a partition off by even a few rows yields means that differ in the second decimal (e.g. 43.31 vs 45.48) and all mean-based checks fail, even though the significance conclusion looks plausible.
613Missing persisted result artifacts (answer reported only in prose/plot)taskda-code
Applies when
task -- the task asks for a computed quantity plus a saved figure/output file, and the harness grades by comparing saved artifact files (numeric results, plot data, image) rather than the chat answer.
Pattern
The script computes the requested statistic in memory, saves only the visual artifact explicitly named in the prompt, and communicates the numbers only in the final text answer; no machine-readable file (array/JSON/CSV of the tallies or plot data) is written, and no check confirms which files exist on disk after the run.
Detection procedure
  1. Read the task and note every deliverable implied: the figure, the underlying numbers, and any conventional companion files the environment/guidance expects (e.g., a serialized array of the computed values and a JSON dump of the plotted data).
  2. Grep the scripts for all write operations (savefig, np.save, to_csv, json.dump, open(...,'w')) and list the exact filenames/paths produced.
  3. Compare that list to the deliverables from step 1; flag if any requested or conventionally required quantity exists only as a printed value or inside the figure.
  4. Check the final answer: if it states numbers that have no corresponding file on disk, the attempt is unverifiable by the grader.
Discriminator
A real violation is when a graded quantity exists only in stdout/prose or is baked into pixels; it is not a violation if every graded value is written to a file with the expected name/format (extra prose restating them is fine), or if the task genuinely has a single artifact and the script writes it correctly.
Consequence
The grader reports the expected output files as WRONG/MISSING and scores 0, even if the in-memory computation was reasonable.
id 1bff706bc941 · mined from da-code dacode-plot-pie-005@s11
raw text (what the judge reads)
### Missing persisted result artifacts (answer reported only in prose/plot)
- **Applies when**: `task` -- the task asks for a computed quantity plus a saved figure/output file, and the harness grades by comparing saved artifact files (numeric results, plot data, image) rather than the chat answer.
- **Pattern**: The script computes the requested statistic in memory, saves only the visual artifact explicitly named in the prompt, and communicates the numbers only in the final text answer; no machine-readable file (array/JSON/CSV of the tallies or plot data) is written, and no check confirms which files exist on disk after the run.
- **Detection procedure**:
  1. Read the task and note every deliverable implied: the figure, the underlying numbers, and any conventional companion files the environment/guidance expects (e.g., a serialized array of the computed values and a JSON dump of the plotted data).
  2. Grep the scripts for all write operations (`savefig`, `np.save`, `to_csv`, `json.dump`, `open(...,'w')`) and list the exact filenames/paths produced.
  3. Compare that list to the deliverables from step 1; flag if any requested or conventionally required quantity exists only as a printed value or inside the figure.
  4. Check the final answer: if it states numbers that have no corresponding file on disk, the attempt is unverifiable by the grader.
- **Discriminator**: A real violation is when a graded quantity exists only in stdout/prose or is baked into pixels; it is *not* a violation if every graded value is written to a file with the expected name/format (extra prose restating them is fine), or if the task genuinely has a single artifact and the script writes it correctly.
- **Consequence**: The grader reports the expected output files as WRONG/MISSING and scores 0, even if the in-memory computation was reasonable.
614Skipping data-quality validation on columns known/likely to contain corrupted values before imputing and modelingtaskinfiagent-dabench
Applies when
task -- the script reads a raw file (often flagged as dirty by its name, mixed dtypes, or sentinel/out-of-range entries) and immediately uses selected columns for mean-imputation, fitting, and metric computation.
Pattern
The attempt trusts the raw column contents: it only counts nulls and prints dtypes, never checks value ranges, units, duplicates, or non-numeric/impossible entries. Corrupted rows inflate the column means used for imputation and distort the fitted model, so the reported metric is off by an order of magnitude, yet no sanity check is applied to the number before reporting.
Detection procedure
  1. Read the task/file naming and constraints to see whether the data is raw or explicitly described as containing errors, and note which columns feed the model.
  2. In the script, look for any cleaning/validation step beyond isna() counts — e.g., coercion to numeric, range or domain plausibility checks, duplicate removal, outlier/sentinel detection — before means are computed and rows are imputed.
  3. Inspect whether the printed summary statistics (mean, min, max) of each modeling column are compared against physically/logically plausible bounds for that quantity.
  4. Check whether the final metric is sanity-checked against the target's variance/scale (e.g., MSE far exceeding the plausible variance of the target signals bad inputs).
Discriminator
A genuine violation is when no distributional or domain check is ever performed and the resulting statistics/metric are implausible relative to the target's scale; it is not a violation if the script inspected value ranges and deliberately kept extreme-but-valid observations, or if the task explicitly instructs using the raw values as-is.
Consequence
Imputation means and regression coefficients are computed from contaminated values, producing a test MSE that is many times the true value and failing exact-value grading.
id 12ee4b877158 · mined from infiagent-dabench dabench-432@s11
raw text (what the judge reads)
### Skipping data-quality validation on columns known/likely to contain corrupted values before imputing and modeling
- **Applies when**: `task` -- the script reads a raw file (often flagged as dirty by its name, mixed dtypes, or sentinel/out-of-range entries) and immediately uses selected columns for mean-imputation, fitting, and metric computation.
- **Pattern**: The attempt trusts the raw column contents: it only counts nulls and prints dtypes, never checks value ranges, units, duplicates, or non-numeric/impossible entries. Corrupted rows inflate the column means used for imputation and distort the fitted model, so the reported metric is off by an order of magnitude, yet no sanity check is applied to the number before reporting.
- **Detection procedure**:
  1. Read the task/file naming and constraints to see whether the data is raw or explicitly described as containing errors, and note which columns feed the model.
  2. In the script, look for any cleaning/validation step beyond `isna()` counts — e.g., coercion to numeric, range or domain plausibility checks, duplicate removal, outlier/sentinel detection — before means are computed and rows are imputed.
  3. Inspect whether the printed summary statistics (mean, min, max) of each modeling column are compared against physically/logically plausible bounds for that quantity.
  4. Check whether the final metric is sanity-checked against the target's variance/scale (e.g., MSE far exceeding the plausible variance of the target signals bad inputs).
- **Discriminator**: A genuine violation is when no distributional or domain check is ever performed and the resulting statistics/metric are implausible relative to the target's scale; it is *not* a violation if the script inspected value ranges and deliberately kept extreme-but-valid observations, or if the task explicitly instructs using the raw values as-is.
- **Consequence**: Imputation means and regression coefficients are computed from contaminated values, producing a test MSE that is many times the true value and failing exact-value grading.
615Row-set / missing-value handling for a two-variable statistic is never pinned down or sanity-checkedtaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic (correlation, mean difference, test p-value) computed from two or more columns of a table, to be reported at a fixed precision.
Pattern
The attempt loads the table and calls a one-line statistic helper without explicitly stating which rows enter the computation — it may drop rows with nulls anywhere in the frame, silently coerce or drop non-numeric/sentinel values, keep duplicate or aggregate rows, or rely on a helper's default pairwise-vs-listwise deletion. No row count, dtype check, or cross-check with a second implementation is printed, so a small shift in the row set produces a value that rounds to the wrong last digit and goes unnoticed.
Detection procedure
  1. Read the task and note the exact variables required and the requested rounding precision.
  2. In the scripts, locate where rows are filtered/dropped and where dtypes are set; check whether missing-value removal is scoped to exactly the required columns (listwise on that pair only) and whether string/sentinel values are converted rather than discarded.
  3. Check whether the script prints diagnostics — n used, dtypes, min/max — and whether the statistic is confirmed by an independent route (e.g., a second library or manual formula).
  4. Compare the reported value's precision and plausibility against those diagnostics; if the row count and NA policy cannot be reconstructed from the output, or the value sits near a rounding boundary with no verification, flag it.
Discriminator
A real violation is an unstated or over-broad row filter / unverified single computation on data that plausibly contains nulls or mixed types; it is fine if the script explicitly restricts NA handling to the required columns, reports n and dtypes, and the value is corroborated (identical to 3+ decimals by a second method) — even if it is written as one line.
Consequence
The reported statistic is computed on a slightly different subset than intended, so it differs from the reference in the last reported digit (e.g., 0.53 vs 0.54) and the numeric check fails even though the qualitative conclusion is right.
id 325ecdc86c93 · mined from infiagent-dabench dabench-300@s11
raw text (what the judge reads)
### Row-set / missing-value handling for a two-variable statistic is never pinned down or sanity-checked
- **Applies when**: `task` -- the task asks for a single numeric statistic (correlation, mean difference, test p-value) computed from two or more columns of a table, to be reported at a fixed precision.
- **Pattern**: The attempt loads the table and calls a one-line statistic helper without explicitly stating which rows enter the computation — it may drop rows with nulls anywhere in the frame, silently coerce or drop non-numeric/sentinel values, keep duplicate or aggregate rows, or rely on a helper's default pairwise-vs-listwise deletion. No row count, dtype check, or cross-check with a second implementation is printed, so a small shift in the row set produces a value that rounds to the wrong last digit and goes unnoticed.
- **Detection procedure**:
  1. Read the task and note the exact variables required and the requested rounding precision.
  2. In the scripts, locate where rows are filtered/dropped and where dtypes are set; check whether missing-value removal is scoped to exactly the required columns (listwise on that pair only) and whether string/sentinel values are converted rather than discarded.
  3. Check whether the script prints diagnostics — n used, dtypes, min/max — and whether the statistic is confirmed by an independent route (e.g., a second library or manual formula).
  4. Compare the reported value's precision and plausibility against those diagnostics; if the row count and NA policy cannot be reconstructed from the output, or the value sits near a rounding boundary with no verification, flag it.
- **Discriminator**: A real violation is an unstated or over-broad row filter / unverified single computation on data that plausibly contains nulls or mixed types; it is fine if the script explicitly restricts NA handling to the required columns, reports n and dtypes, and the value is corroborated (identical to 3+ decimals by a second method) — even if it is written as one line.
- **Consequence**: The reported statistic is computed on a slightly different subset than intended, so it differs from the reference in the last reported digit (e.g., 0.53 vs 0.54) and the numeric check fails even though the qualitative conclusion is right.
616No measured validation and no check for labeled overlap between the provided source data and the evaluation rowstaskda-code
Applies when
task -- the agent must produce per-row predictions for an unlabeled evaluation file while a larger labeled source file (from which the evaluation rows were likely carved) is available.
Pattern
The script fits one model on the entire labeled file and writes predictions straight out, with no held-out evaluation, no accuracy estimate, no comparison of alternative models/features, and no attempt to join the evaluation rows to the labeled file on identifier or exact-text keys to see whether true labels are already available for some or all rows. Quality is asserted ("completed successfully") from prediction counts and class distribution alone.
Detection procedure
  1. In the task, note that the evaluation rows are scored against exact ground-truth labels and check whether a labeled superset/source file is supplied.
  2. In the scripts, look for (a) any train/validation split or cross-validation producing a numeric score, and (b) any merge/lookup of evaluation rows against the labeled file on ID or key text fields.
  3. If neither exists, flag: the agent has no evidence its predictions meet the accuracy bar and may have discarded directly recoverable labels.
  4. Confirm the answer contains only predicted labels with no reported score or match rate against known labels.
Discriminator
A real violation is submitting predictions with zero quantitative quality evidence when labels for the evaluation rows are obtainable or a holdout is trivially constructible. It is fine if the agent reports a validation/CV score (or a verified key-based match rate) and that score plausibly clears the task's threshold, or if the labeled source demonstrably shares no keys or rows with the evaluation set.
Consequence
The grader compares predictions to exact ground-truth labels; a mid-accuracy generic text model produces many mismatched rows and the file is marked WRONG, whereas a validated model or key-based label lookup would have matched.
id ff69cdc41e4e · mined from da-code dacode-ml-multi-008@s11
raw text (what the judge reads)
### No measured validation and no check for labeled overlap between the provided source data and the evaluation rows
- **Applies when**: `task` -- the agent must produce per-row predictions for an unlabeled evaluation file while a larger labeled source file (from which the evaluation rows were likely carved) is available.
- **Pattern**: The script fits one model on the entire labeled file and writes predictions straight out, with no held-out evaluation, no accuracy estimate, no comparison of alternative models/features, and no attempt to join the evaluation rows to the labeled file on identifier or exact-text keys to see whether true labels are already available for some or all rows. Quality is asserted ("completed successfully") from prediction counts and class distribution alone.
- **Detection procedure**:
  1. In the task, note that the evaluation rows are scored against exact ground-truth labels and check whether a labeled superset/source file is supplied.
  2. In the scripts, look for (a) any train/validation split or cross-validation producing a numeric score, and (b) any merge/lookup of evaluation rows against the labeled file on ID or key text fields.
  3. If neither exists, flag: the agent has no evidence its predictions meet the accuracy bar and may have discarded directly recoverable labels.
  4. Confirm the answer contains only predicted labels with no reported score or match rate against known labels.
- **Discriminator**: A real violation is submitting predictions with zero quantitative quality evidence when labels for the evaluation rows are obtainable or a holdout is trivially constructible. It is fine if the agent reports a validation/CV score (or a verified key-based match rate) and that score plausibly clears the task's threshold, or if the labeled source demonstrably shares no keys or rows with the evaluation set.
- **Consequence**: The grader compares predictions to exact ground-truth labels; a mid-accuracy generic text model produces many mismatched rows and the file is marked WRONG, whereas a validated model or key-based label lookup would have matched.
617Output artifacts written outside the input data's directory/expected locationtaskinfiagent-dabench
Applies when
task -- the task asks the agent to report file paths of derived/output files created from a provided input file.
Pattern
The agent reads the input from its given location but writes the derived files to an unrelated working directory (e.g., its home/cwd) and reports that path, instead of writing them alongside the input or to the location implied by the task setup.
Detection procedure
1. From the task, note where the input file lives and any stated/implied output location or naming convention. 2. In the scripts, find every write call (to_csv, savefig, open(...,'w')) and record the exact output paths. 3. Compare those paths (directory and filename) with the input file's directory and the task's convention. 4. Check the reported paths in the answer match the paths actually written and follow that convention.
Discriminator
A real violation is an output directory that differs from the input data directory with no instruction authorizing it; it is fine if the task explicitly names an output directory, or if the agent writes to the same directory as the input (or a task-specified one) and reports that absolute path consistently.
Consequence
All numeric values may be correct, yet the path field mismatches the expected string, so the check containing it (and possibly the whole answer) is graded wrong.
id fc88b2cfc157 · mined from infiagent-dabench dabench-743@s11
raw text (what the judge reads)
### Output artifacts written outside the input data's directory/expected location
- **Applies when**: `task` -- the task asks the agent to report file paths of derived/output files created from a provided input file.
- **Pattern**: The agent reads the input from its given location but writes the derived files to an unrelated working directory (e.g., its home/cwd) and reports that path, instead of writing them alongside the input or to the location implied by the task setup.
- **Detection procedure**: 1. From the task, note where the input file lives and any stated/implied output location or naming convention. 2. In the scripts, find every write call (`to_csv`, `savefig`, `open(...,'w')`) and record the exact output paths. 3. Compare those paths (directory and filename) with the input file's directory and the task's convention. 4. Check the reported paths in the answer match the paths actually written and follow that convention.
- **Discriminator**: A real violation is an output directory that differs from the input data directory with no instruction authorizing it; it is fine if the task explicitly names an output directory, or if the agent writes to the same directory as the input (or a task-specified one) and reports that absolute path consistently.
- **Consequence**: All numeric values may be correct, yet the path field mismatches the expected string, so the check containing it (and possibly the whole answer) is graded wrong.
618Ignoring an external spec file and shipping an incomplete artifact settaskda-code
Applies when
task -- the task points to a configuration/guidelines file (e.g. a YAML/JSON spec) and asks for saved outputs (chart, arrays, config dumps) in addition to a narrative answer.
Pattern
The agent computes a headline number and writes one obvious output (the image), but never loads/parses the referenced spec file, never applies its styling/labeling/format directives, and never emits the companion artifacts that the spec or task convention implies (e.g. the serialized plot config and the underlying numeric array). The final message narrates the result instead of demonstrating that every required file exists with the required content.
Detection procedure
  1. Read the task and the referenced spec file; enumerate every required output path and every explicit directive (labels, ordering, categories included, units, rounding, file names/formats).
  2. Scan the scripts for a read of that spec file and for a write of each enumerated artifact; check the plotting/serialization code actually consumes the spec values rather than hard-coding its own.
  3. Compare the agent's reported numbers against the spec's stated categories/units (e.g. do the reported shares cover all categories and sum to 100%? are units and rounding as specified?).
  4. Require the answer to list each produced file with a verification (existence, shape/length, category set) rather than a prose claim.
Discriminator
A real violation is missing required files, a spec file that is never opened, or output content that contradicts the spec (missing categories, wrong labels/units). A look-alike that is fine is an agent that reads the spec, produces all files, and merely summarizes a subset of the numbers in prose while the saved artifacts are complete and consistent.
Consequence
Artifact-level checks fail as WRONG/MISSING for the unproduced or non-conforming files, so the submission scores zero even when the headline computed value happens to be right.
id a0260ab37956 · mined from da-code dacode-plot-pie-008@s11
raw text (what the judge reads)
### Ignoring an external spec file and shipping an incomplete artifact set
- **Applies when**: `task` -- the task points to a configuration/guidelines file (e.g. a YAML/JSON spec) and asks for saved outputs (chart, arrays, config dumps) in addition to a narrative answer.
- **Pattern**: The agent computes a headline number and writes one obvious output (the image), but never loads/parses the referenced spec file, never applies its styling/labeling/format directives, and never emits the companion artifacts that the spec or task convention implies (e.g. the serialized plot config and the underlying numeric array). The final message narrates the result instead of demonstrating that every required file exists with the required content.
- **Detection procedure**:
  1. Read the task and the referenced spec file; enumerate every required output path and every explicit directive (labels, ordering, categories included, units, rounding, file names/formats).
  2. Scan the scripts for a read of that spec file and for a write of each enumerated artifact; check the plotting/serialization code actually consumes the spec values rather than hard-coding its own.
  3. Compare the agent's reported numbers against the spec's stated categories/units (e.g. do the reported shares cover all categories and sum to 100%? are units and rounding as specified?).
  4. Require the answer to list each produced file with a verification (existence, shape/length, category set) rather than a prose claim.
- **Discriminator**: A real violation is missing required files, a spec file that is never opened, or output content that contradicts the spec (missing categories, wrong labels/units). A look-alike that is fine is an agent that reads the spec, produces all files, and merely summarizes a subset of the numbers in prose while the saved artifacts are complete and consistent.
- **Consequence**: Artifact-level checks fail as WRONG/MISSING for the unproduced or non-conforming files, so the submission scores zero even when the headline computed value happens to be right.
619Derived variable parsed from messy/dirty source with silent failures and no cleaning or validationtaskinfiagent-dabench
Applies when
task -- a required analysis variable is not stored directly but must be derived (parsed, computed, decoded) from free-text or known-noisy input columns, and the script builds it with custom parsing logic plus try/except or NaN fallbacks.
Pattern
The script writes an ad-hoc parser/derivation, silently returns NaN (or a default) on any case it can't handle, drops those rows implicitly before the statistic, and never checks the source data for corrupted/out-of-range/duplicate/typo values even when the data source is flagged as containing errors. The resulting statistic is computed on a subtly wrong subset or on wrong derived values, so it is "close but off".
Detection procedure
  1. From the task, identify every variable in the statistic and whether it is raw or derived; note any signal (file naming, task wording, constraints) that the raw data may contain deliberate errors needing cleaning.
  2. In the scripts, locate the derivation and list the branches that yield NaN/defaults or exceptions; check whether the script counts and inspects those rows, and whether it validates derived values for plausibility (non-negative, sane range, expected count matching input rows).
  3. Check whether any explicit cleaning/validation step exists for the raw inputs of the statistic (implausible categories, out-of-range numbers, wrong dtypes, duplicated records) before splitting/filtering.
  4. Confirm the reported number is computed after such validation; if rows were dropped or values silently coerced without being reported or justified, flag it.
Discriminator
A real violation is silent, unquantified loss/mis-derivation with no sanity check (row counts, value ranges, spot-check of parsed vs. source, no handling of known-dirty entries). It is fine if the script explicitly prints how many rows failed derivation, shows the parsed values match the source for edge cases, and documents a justified rule for excluding or repairing bad records.
Consequence
The correlation/metric is computed on a slightly different or slightly wrong sample than intended, so the categorical conclusion may still match but the numeric value fails the exact-value check (e.g., 0.58 reported vs. 0.56 expected), producing a partial-credit/incorrect verdict.
id 6bf1225b64aa · mined from infiagent-dabench dabench-431@s11
raw text (what the judge reads)
### Derived variable parsed from messy/dirty source with silent failures and no cleaning or validation
- **Applies when**: `task` -- a required analysis variable is not stored directly but must be derived (parsed, computed, decoded) from free-text or known-noisy input columns, and the script builds it with custom parsing logic plus `try/except` or `NaN` fallbacks.
- **Pattern**: The script writes an ad-hoc parser/derivation, silently returns `NaN` (or a default) on any case it can't handle, drops those rows implicitly before the statistic, and never checks the source data for corrupted/out-of-range/duplicate/typo values even when the data source is flagged as containing errors. The resulting statistic is computed on a subtly wrong subset or on wrong derived values, so it is "close but off".
- **Detection procedure**:
  1. From the task, identify every variable in the statistic and whether it is raw or derived; note any signal (file naming, task wording, constraints) that the raw data may contain deliberate errors needing cleaning.
  2. In the scripts, locate the derivation and list the branches that yield `NaN`/defaults or exceptions; check whether the script counts and inspects those rows, and whether it validates derived values for plausibility (non-negative, sane range, expected count matching input rows).
  3. Check whether any explicit cleaning/validation step exists for the raw inputs of the statistic (implausible categories, out-of-range numbers, wrong dtypes, duplicated records) before splitting/filtering.
  4. Confirm the reported number is computed after such validation; if rows were dropped or values silently coerced without being reported or justified, flag it.
- **Discriminator**: A real violation is silent, unquantified loss/mis-derivation with no sanity check (row counts, value ranges, spot-check of parsed vs. source, no handling of known-dirty entries). It is fine if the script explicitly prints how many rows failed derivation, shows the parsed values match the source for edge cases, and documents a justified rule for excluding or repairing bad records.
- **Consequence**: The correlation/metric is computed on a slightly different or slightly wrong sample than intended, so the categorical conclusion may still match but the numeric value fails the exact-value check (e.g., 0.58 reported vs. 0.56 expected), producing a partial-credit/incorrect verdict.
620Class-balance sanity check on predicted labels is missing (accuracy-only model selection under imbalance)taskda-code
Applies when
task -- the task asks for hard class labels on a held-out set for an imbalanced binary/multiclass target, and the scripts select a model/threshold using overall accuracy.
Pattern
The attempt picks whichever classifier has the highest plain accuracy, uses the default 0.5 cut-off, and writes out predictions whose positive rate is far below the target's base rate in the training data, never comparing the two or reporting a balance-aware metric (recall, F1, ROC-AUC, confusion matrix).
Detection procedure
  1. From the task/README or a quick read of the training file, note the empirical class proportions of the target.
  2. In the scripts, check whether model/threshold selection uses only accuracy and whether any class-weighting, resampling, threshold tuning, or balance-aware metric is present.
  3. In the answer, compute the predicted positive rate on the test output and compare it with the training base rate; also confirm the reported accuracy is not simply near the majority-class share.
  4. Flag if the predicted positive rate is roughly half (or double) the base rate, or if no per-class metric was ever reported.
Discriminator
A genuine violation shows a large, unexplained gap between predicted and prior positive rates with no calibration/threshold justification; it is fine if the script deliberately tunes the threshold for a stated metric, documents why the shifted rate is expected, or if the test set is legitimately known to differ in composition.
Consequence
The submitted label file under-calls the minority class, so recall/F1/balanced-accuracy (and any exact-match comparison against expected labels) fall well below what the reported accuracy suggests, and the output file is graded wrong.
id a94b32618268 · mined from da-code dacode-ml-binary-016@s11
raw text (what the judge reads)
### Class-balance sanity check on predicted labels is missing (accuracy-only model selection under imbalance)
- **Applies when**: `task` -- the task asks for hard class labels on a held-out set for an imbalanced binary/multiclass target, and the scripts select a model/threshold using overall accuracy.
- **Pattern**: The attempt picks whichever classifier has the highest plain accuracy, uses the default 0.5 cut-off, and writes out predictions whose positive rate is far below the target's base rate in the training data, never comparing the two or reporting a balance-aware metric (recall, F1, ROC-AUC, confusion matrix).
- **Detection procedure**:
  1. From the task/README or a quick read of the training file, note the empirical class proportions of the target.
  2. In the scripts, check whether model/threshold selection uses only accuracy and whether any class-weighting, resampling, threshold tuning, or balance-aware metric is present.
  3. In the answer, compute the predicted positive rate on the test output and compare it with the training base rate; also confirm the reported accuracy is not simply near the majority-class share.
  4. Flag if the predicted positive rate is roughly half (or double) the base rate, or if no per-class metric was ever reported.
- **Discriminator**: A genuine violation shows a large, unexplained gap between predicted and prior positive rates with no calibration/threshold justification; it is *fine* if the script deliberately tunes the threshold for a stated metric, documents why the shifted rate is expected, or if the test set is legitimately known to differ in composition.
- **Consequence**: The submitted label file under-calls the minority class, so recall/F1/balanced-accuracy (and any exact-match comparison against expected labels) fall well below what the reported accuracy suggests, and the output file is graded wrong.
621Missing or unverified output artifact matching the requested file/templatetaskda-code
Applies when
task -- the task asks for the result to be written to a named file in the format of a provided sample/template, and the agent's work is only reported in chat or in an ad-hoc script.
Pattern
The attempt computes a number and reports it in prose, but never persists a script that writes the required file, or writes it with a different name/path, column header, row count, or dtype than the supplied template — so the deliverable the grader reads is absent or malformed even if the statistic is right.
Detection procedure
  1. From the task text, list the exact required deliverable(s): file name, location, and the schema implied by the sample/template (column names, order, number of rows, value type/rounding).
  2. Read the scripts for an explicit write step (e.g., to_csv) that targets that exact filename and builds columns from the template rather than from invented names; confirm the script is saved/reproducible and that any stated constraint (random seed, rounding, units) is set in it.
  3. Check the reported answer against the template: does it consist of the same header and one value per required row, and was the file's existence/contents re-read and printed after writing?
  4. Flag if no write step exists, the path/header differs from the template, or the file was never read back to confirm shape.
Discriminator
A real violation is when the required file is not written, is named/structured differently from the template, or was never verified after writing; a look-alike that is fine writes the exact filename with template-derived headers and echoes the file contents (matching schema) even if the reported prose value is formatted differently.
Consequence
The grader marks the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying statistic was computed correctly.
id c2904c4471c9 · mined from da-code dacode-data-sa-039@s11
raw text (what the judge reads)
### Missing or unverified output artifact matching the requested file/template
- **Applies when**: `task` -- the task asks for the result to be written to a named file in the format of a provided sample/template, and the agent's work is only reported in chat or in an ad-hoc script.
- **Pattern**: The attempt computes a number and reports it in prose, but never persists a script that writes the required file, or writes it with a different name/path, column header, row count, or dtype than the supplied template — so the deliverable the grader reads is absent or malformed even if the statistic is right.
- **Detection procedure**:
  1. From the task text, list the exact required deliverable(s): file name, location, and the schema implied by the sample/template (column names, order, number of rows, value type/rounding).
  2. Read the scripts for an explicit write step (e.g., `to_csv`) that targets that exact filename and builds columns from the template rather than from invented names; confirm the script is saved/reproducible and that any stated constraint (random seed, rounding, units) is set in it.
  3. Check the reported answer against the template: does it consist of the same header and one value per required row, and was the file's existence/contents re-read and printed after writing?
  4. Flag if no write step exists, the path/header differs from the template, or the file was never read back to confirm shape.
- **Discriminator**: A real violation is when the required file is not written, is named/structured differently from the template, or was never verified after writing; a look-alike that is fine writes the exact filename with template-derived headers and echoes the file contents (matching schema) even if the reported prose value is formatted differently.
- **Consequence**: The grader marks the expected result file as WRONG/MISSING and scores 0 regardless of whether the underlying statistic was computed correctly.
622Model quality never validated out-of-sample (only in-sample fit reported)taskda-code
Applies when
task -- the task asks for predictions on a test file that will be scored against hidden ground truth, and the scripts fit one model on the full training data and immediately write predictions.
Pattern
The agent trains a single model with default/arbitrary hyperparameters, evaluates it only on the same rows it was fit on (or not at all), drops potentially informative columns (e.g., timestamps/identifiers) without checking their predictive value or engineering features from them, and then reports the training-set score as evidence of success. No hold-out/cross-validation estimate, no comparison to a trivial baseline (mean/median/persistence), and no alternative model is tried, so a weak model is shipped unnoticed.
Detection procedure
  1. Read the task to confirm the deliverable is scored on prediction accuracy, not just file existence/format.
  2. Scan the scripts for any train/validation split, cross-validation, or backtest that produces a score on rows not used for fitting; check whether any baseline or second model is compared.
  3. Check which columns are dropped before fitting and whether any dropped column carries signal that could have been transformed rather than discarded.
  4. Read the answer: if the only quality numbers quoted are computed on the training rows (and are mediocre, e.g., a low R²/high error relative to target variance), the attempt has no evidence the predictions generalize.
Discriminator
A real violation is when no score on unseen data exists anywhere, or the only reported score is in-sample and weak, with no baseline comparison. It is fine if the agent reports a proper hold-out/CV score (even a modest one) and shows it beat a baseline, or if the task genuinely only checks file format rather than accuracy.
Consequence
The predictions file is produced with the right shape and column name but its values are too inaccurate against the hidden targets, so the accuracy-based file check fails even though the answer claims success.
id bb9a384d9146 · mined from da-code dacode-ml-regression-015@s11
raw text (what the judge reads)
### Model quality never validated out-of-sample (only in-sample fit reported)
- **Applies when**: `task` -- the task asks for predictions on a test file that will be scored against hidden ground truth, and the scripts fit one model on the full training data and immediately write predictions.
- **Pattern**: The agent trains a single model with default/arbitrary hyperparameters, evaluates it only on the same rows it was fit on (or not at all), drops potentially informative columns (e.g., timestamps/identifiers) without checking their predictive value or engineering features from them, and then reports the training-set score as evidence of success. No hold-out/cross-validation estimate, no comparison to a trivial baseline (mean/median/persistence), and no alternative model is tried, so a weak model is shipped unnoticed.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on prediction accuracy, not just file existence/format.
  2. Scan the scripts for any train/validation split, cross-validation, or backtest that produces a score on rows not used for fitting; check whether any baseline or second model is compared.
  3. Check which columns are dropped before fitting and whether any dropped column carries signal that could have been transformed rather than discarded.
  4. Read the answer: if the only quality numbers quoted are computed on the training rows (and are mediocre, e.g., a low R²/high error relative to target variance), the attempt has no evidence the predictions generalize.
- **Discriminator**: A real violation is when *no* score on unseen data exists anywhere, or the only reported score is in-sample and weak, with no baseline comparison. It is fine if the agent reports a proper hold-out/CV score (even a modest one) and shows it beat a baseline, or if the task genuinely only checks file format rather than accuracy.
- **Consequence**: The predictions file is produced with the right shape and column name but its values are too inaccurate against the hidden targets, so the accuracy-based file check fails even though the answer claims success.
623Missing/unreproducible output artifact for the requested deliverabletaskda-code
Applies when
task -- the task asks for a specific answer object (a JSON/CSV in a stated shape) that the grader will read from a result file, and the agent's work should consist of scripts that compute and persist it.
Pattern
The agent reports the answer only in chat prose, with no saved script that (a) loads the raw data, (b) performs each stated preprocessing step, (c) computes the requested statistic, and (d) writes the result to the expected file path in the exact requested schema. The numbers may look plausible but nothing on disk backs them, and no intermediate check (dtype after loading, count of imputed cells, the actual extremum values) is recorded.
Detection procedure
  1. From the task statement, list the required deliverable: file name/path, top-level keys, value types (scalar vs list), and any stated preprocessing or formatting constraints.
  2. Inspect the agent's scripts for a code path that ends in an explicit write of that exact structure to that exact path; also check that each stated preprocessing step appears in code rather than being assumed.
  3. Check whether the script prints/asserts verifiable evidence — column dtypes after parsing, number of NaNs filled, and the extreme values (not just the labels) — so a reviewer can sanity-check magnitudes and ties.
  4. Compare the reported chat answer against what the script would write; if there is no script, or the script writes a different path/shape, flag it.
Discriminator
A real violation is when no persisted artifact in the required schema can be produced by re-running the agent's code (or the artifact's keys/types/path differ). A look-alike that is fine is a script that writes the correct file and additionally echoes the answer in chat, or that writes the file under the exact requested name with equivalent but explicitly justified value types.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, regardless of whether the values quoted in the response happen to be right; silent preprocessing/dtype errors (e.g., a numeric column read as text so ordering is lexicographic) also go undetected because no sanity output exists.
id 39fc9ca13e9f · mined from da-code dacode-di-text-001@s11
raw text (what the judge reads)
### Missing/unreproducible output artifact for the requested deliverable
- **Applies when**: `task` -- the task asks for a specific answer object (a JSON/CSV in a stated shape) that the grader will read from a result file, and the agent's work should consist of scripts that compute and persist it.
- **Pattern**: The agent reports the answer only in chat prose, with no saved script that (a) loads the raw data, (b) performs each stated preprocessing step, (c) computes the requested statistic, and (d) writes the result to the expected file path in the exact requested schema. The numbers may look plausible but nothing on disk backs them, and no intermediate check (dtype after loading, count of imputed cells, the actual extremum values) is recorded.
- **Detection procedure**:
  1. From the task statement, list the required deliverable: file name/path, top-level keys, value types (scalar vs list), and any stated preprocessing or formatting constraints.
  2. Inspect the agent's scripts for a code path that ends in an explicit write of that exact structure to that exact path; also check that each stated preprocessing step appears in code rather than being assumed.
  3. Check whether the script prints/asserts verifiable evidence — column dtypes after parsing, number of NaNs filled, and the extreme values (not just the labels) — so a reviewer can sanity-check magnitudes and ties.
  4. Compare the reported chat answer against what the script would write; if there is no script, or the script writes a different path/shape, flag it.
- **Discriminator**: A real violation is when no persisted artifact in the required schema can be produced by re-running the agent's code (or the artifact's keys/types/path differ). A look-alike that is fine is a script that writes the correct file and additionally echoes the answer in chat, or that writes the file under the exact requested name with equivalent but explicitly justified value types.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, regardless of whether the values quoted in the response happen to be right; silent preprocessing/dtype errors (e.g., a numeric column read as text so ordering is lexicographic) also go undetected because no sanity output exists.
624Deliverable omits requested intermediate quantities (over-trimmed output file)taskda-code
Applies when
task -- the task asks to compute several derived quantities and save results "including X and Y" into a specified output file, and the script builds a rich intermediate table but writes only a subset of its columns.
Pattern
The agent computes all required per-entity metrics and derived labels in memory, then subsets the final DataFrame down to an identifier plus one label column before writing, silently dropping the other quantities the task explicitly named (component scores, aggregate score, segment code). It then declares success, sometimes asserting a "match with reference data" that was never actually performed.
Detection procedure
  1. Read the task statement and list every quantity it names as part of the saved output (each metric, each score, the segment/group identifier, the final label).
  2. Find the line in the script that writes the output file and enumerate exactly the columns it contains (watch for a [[...]] subset, drop, or to_csv on a reduced frame).
  3. Diff the two lists: flag if any named quantity is computed but not written, or if column names/ordering do not reflect the requested content.
  4. Check whether any validation claim in the answer is backed by code in the scripts that actually reads a reference/expected artifact; unsupported claims are a red flag.
Discriminator
A real violation is when a quantity the task explicitly asked to save exists in the script's intermediate frame (or is trivially derivable) but is absent from the written file. It is not a violation when the task asks only for the final label, or when the extra columns are purely diagnostic and never named in the task.
Consequence
The saved file has fewer/different columns than the expected artifact, so an exact-file or column-wise comparison fails outright even if the per-row labels are computed sensibly, yielding 0/1 on the file check.
id 1892e9a6a4d5 · mined from da-code dacode-dm-csv-052@s11
raw text (what the judge reads)
### Deliverable omits requested intermediate quantities (over-trimmed output file)
- **Applies when**: `task` -- the task asks to compute several derived quantities and save results "including X and Y" into a specified output file, and the script builds a rich intermediate table but writes only a subset of its columns.
- **Pattern**: The agent computes all required per-entity metrics and derived labels in memory, then subsets the final DataFrame down to an identifier plus one label column before writing, silently dropping the other quantities the task explicitly named (component scores, aggregate score, segment code). It then declares success, sometimes asserting a "match with reference data" that was never actually performed.
- **Detection procedure**:
  1. Read the task statement and list every quantity it names as part of the saved output (each metric, each score, the segment/group identifier, the final label).
  2. Find the line in the script that writes the output file and enumerate exactly the columns it contains (watch for a `[[...]]` subset, `drop`, or `to_csv` on a reduced frame).
  3. Diff the two lists: flag if any named quantity is computed but not written, or if column names/ordering do not reflect the requested content.
  4. Check whether any validation claim in the answer is backed by code in the scripts that actually reads a reference/expected artifact; unsupported claims are a red flag.
- **Discriminator**: A real violation is when a quantity the task explicitly asked to save exists in the script's intermediate frame (or is trivially derivable) but is absent from the written file. It is *not* a violation when the task asks only for the final label, or when the extra columns are purely diagnostic and never named in the task.
- **Consequence**: The saved file has fewer/different columns than the expected artifact, so an exact-file or column-wise comparison fails outright even if the per-row labels are computed sensibly, yielding 0/1 on the file check.
625Requested output artifact never written (or not verified against the provided template)taskda-code
Applies when
task -- The task explicitly asks for results to be saved to a named file following a given sample/template file, and the agent's scripts and answer are available for inspection.
Pattern
The attempt computes numbers and prints/reports them in prose or a chat summary, but no script step actually writes the named output file, or it writes it without reading the template to match filename, column names, row order, and value formatting (e.g., how intervals are encoded).
Detection procedure
  1. From the task text, list the exact deliverable artifacts (file name, location) and the template file that defines the schema.
  2. Search the scripts for a write call producing that exact filename, and for code that loads/echoes the template's header and row structure; note whether any script was actually saved/run at all.
  3. Compare the reported result layout (columns, row identifiers, ordering, value encoding such as bracketed pairs vs separate columns, decimal precision) against the template's layout.
  4. Sanity-check the reported per-group record counts and value ranges against the source data as loaded by the script's filtering logic.
Discriminator
A real violation is when the deliverable file is absent, misnamed, or its schema was never checked against the template — reporting the numbers in the chat does not substitute. A look-alike that is fine is a run that writes the correct file and merely also summarizes it in prose, or that reorders nothing while using slightly different but template-consistent formatting.
Consequence
The grader looks for the named result file, finds it missing or schema-mismatched, and scores 0 regardless of whether the underlying statistics were computed correctly.
id c0a41ff0376d · mined from da-code dacode-data-sa-029@s11
raw text (what the judge reads)
### Requested output artifact never written (or not verified against the provided template)
- **Applies when**: `task` -- The task explicitly asks for results to be saved to a named file following a given sample/template file, and the agent's scripts and answer are available for inspection.
- **Pattern**: The attempt computes numbers and prints/reports them in prose or a chat summary, but no script step actually writes the named output file, or it writes it without reading the template to match filename, column names, row order, and value formatting (e.g., how intervals are encoded).
- **Detection procedure**:
  1. From the task text, list the exact deliverable artifacts (file name, location) and the template file that defines the schema.
  2. Search the scripts for a write call producing that exact filename, and for code that loads/echoes the template's header and row structure; note whether any script was actually saved/run at all.
  3. Compare the reported result layout (columns, row identifiers, ordering, value encoding such as bracketed pairs vs separate columns, decimal precision) against the template's layout.
  4. Sanity-check the reported per-group record counts and value ranges against the source data as loaded by the script's filtering logic.
- **Discriminator**: A real violation is when the deliverable file is absent, misnamed, or its schema was never checked against the template — reporting the numbers in the chat does not substitute. A look-alike that is fine is a run that writes the correct file and merely also summarizes it in prose, or that reorders nothing while using slightly different but template-consistent formatting.
- **Consequence**: The grader looks for the named result file, finds it missing or schema-mismatched, and scores 0 regardless of whether the underlying statistics were computed correctly.
626Uncritically feeding every column (including the label/ID column) into an unsupervised feature matrixtaskda-code
Applies when
task -- the script builds a feature matrix with df.values / df.copy() and passes it straight to a clustering or other unsupervised/embedding routine without inspecting what each column represents.
Pattern
The attempt treats the whole table as homogeneous numeric features — including the dataset's target/class column, row IDs, or already one-hot encoded indicator blocks — scales them all identically, and then reports cluster structure and "optimal k" that are largely driven by the leaked target or by an arbitrarily weighted dummy block, rather than by the descriptive attributes. Often the chosen number of clusters is also stated inconsistently with the diagnostic actually computed (e.g., the search prints one best value, the comment says another, the code uses a third), showing the choice was not grounded in the evidence.
Detection procedure
  1. Read the task/README to identify which columns are descriptive attributes and which are a target label, identifier, or derived/encoded artifact.
  2. In the script, find where X is constructed; check whether any target/ID column is dropped and whether binary dummy columns are handled differently from continuous ones (or at least acknowledged).
  3. Cross-check the model-selection step: does the value of the hyperparameter finally used equal the one the printed diagnostic actually favors, and is the same feature matrix used in both selection and final fit?
  4. Check the reported output shape/row count against the source table's row count and the requested column naming (e.g., does the number of Feature_i columns imply the label was kept?).
Discriminator
A real violation is when a column that encodes the outcome/identity is silently included as a clustering input, or when the reported k contradicts the script's own diagnostics; it is not a violation if the agent explicitly justifies keeping all columns (e.g., the task defines the feature vector as the full row) and the final settings are consistent with what the selection code printed.
Consequence
The saved result file has the wrong feature set/column count and cluster assignments that reflect the leaked label rather than genuine structure, so any comparison against the expected file (shape, columns, or cluster agreement) fails.
id 5d8d2e227b3c · mined from da-code dacode-ml-cluster-010@s11
raw text (what the judge reads)
### Uncritically feeding every column (including the label/ID column) into an unsupervised feature matrix
- **Applies when**: `task` -- the script builds a feature matrix with `df.values` / `df.copy()` and passes it straight to a clustering or other unsupervised/embedding routine without inspecting what each column represents.
- **Pattern**: The attempt treats the whole table as homogeneous numeric features — including the dataset's target/class column, row IDs, or already one-hot encoded indicator blocks — scales them all identically, and then reports cluster structure and "optimal k" that are largely driven by the leaked target or by an arbitrarily weighted dummy block, rather than by the descriptive attributes. Often the chosen number of clusters is also stated inconsistently with the diagnostic actually computed (e.g., the search prints one best value, the comment says another, the code uses a third), showing the choice was not grounded in the evidence.
- **Detection procedure**:
  1. Read the task/README to identify which columns are descriptive attributes and which are a target label, identifier, or derived/encoded artifact.
  2. In the script, find where `X` is constructed; check whether any target/ID column is dropped and whether binary dummy columns are handled differently from continuous ones (or at least acknowledged).
  3. Cross-check the model-selection step: does the value of the hyperparameter finally used equal the one the printed diagnostic actually favors, and is the same feature matrix used in both selection and final fit?
  4. Check the reported output shape/row count against the source table's row count and the requested column naming (e.g., does the number of `Feature_i` columns imply the label was kept?).
- **Discriminator**: A real violation is when a column that encodes the outcome/identity is silently included as a clustering input, or when the reported k contradicts the script's own diagnostics; it is *not* a violation if the agent explicitly justifies keeping all columns (e.g., the task defines the feature vector as the full row) and the final settings are consistent with what the selection code printed.
- **Consequence**: The saved result file has the wrong feature set/column count and cluster assignments that reflect the leaked label rather than genuine structure, so any comparison against the expected file (shape, columns, or cluster agreement) fails.
627Unvalidated cluster count / degenerate cluster structure in unsupervised groupingtaskda-code
Applies when
task -- the task asks for grouping records into "an appropriate number of" clusters and the script picks k and writes a labels file without any quantitative model-selection or quality check.
Pattern
The attempt hard-codes (or eyeballs) the number of clusters, builds heavily right-skewed aggregate features and standardizes them without transforming/winsorizing outliers, then reports the result with no silhouette/elbow/gap evidence — producing a partition where one or two clusters absorb the vast majority of rows and another holds a handful of extreme points (an outlier-detector, not a segmentation).
Detection procedure
  1. Read the task for the requirement that the cluster count be "appropriate" and for the exact output schema (column names, one row per entity, feature columns matching the vectors actually clustered).
  2. In the scripts, check whether k is selected by a stated criterion evaluated over a range (silhouette, inertia elbow, Davies–Bouldin, BIC) and whether skewed monetary/count features are log- or rank-transformed / outlier-trimmed before scaling.
  3. In the answer, inspect the reported per-cluster sizes and any reported internal validity score; flag if no score is reported, or if cluster sizes are extremely lopsided (e.g., a cluster with <1% of rows while another holds >60%).
  4. Confirm the written file's columns/row count match both the schema and the number of entities claimed, and that the saved feature values are the ones actually fed to the model.
Discriminator
A genuinely justified solution shows the search over k with its scores (and may still land on unbalanced clusters that the score supports); a violation states k with no supporting evidence, or reports a partition whose imbalance/near-zero separation would be exposed by the very check that was skipped.
Consequence
The saved labels fail the grader's cluster-quality/structure check (e.g., low silhouette or a partition not reproducible as a sensible segmentation), so the output file is marked WRONG despite being correctly formatted.
id 62af904400ff · mined from da-code dacode-ml-cluster-019@s11
raw text (what the judge reads)
### Unvalidated cluster count / degenerate cluster structure in unsupervised grouping
- **Applies when**: `task` -- the task asks for grouping records into "an appropriate number of" clusters and the script picks k and writes a labels file without any quantitative model-selection or quality check.
- **Pattern**: The attempt hard-codes (or eyeballs) the number of clusters, builds heavily right-skewed aggregate features and standardizes them without transforming/winsorizing outliers, then reports the result with no silhouette/elbow/gap evidence — producing a partition where one or two clusters absorb the vast majority of rows and another holds a handful of extreme points (an outlier-detector, not a segmentation).
- **Detection procedure**:
  1. Read the task for the requirement that the cluster count be "appropriate" and for the exact output schema (column names, one row per entity, feature columns matching the vectors actually clustered).
  2. In the scripts, check whether k is selected by a stated criterion evaluated over a range (silhouette, inertia elbow, Davies–Bouldin, BIC) and whether skewed monetary/count features are log- or rank-transformed / outlier-trimmed before scaling.
  3. In the answer, inspect the reported per-cluster sizes and any reported internal validity score; flag if no score is reported, or if cluster sizes are extremely lopsided (e.g., a cluster with <1% of rows while another holds >60%).
  4. Confirm the written file's columns/row count match both the schema and the number of entities claimed, and that the saved feature values are the ones actually fed to the model.
- **Discriminator**: A genuinely justified solution shows the search over k with its scores (and may still land on unbalanced clusters that the score supports); a violation states k with no supporting evidence, or reports a partition whose imbalance/near-zero separation would be exposed by the very check that was skipped.
- **Consequence**: The saved labels fail the grader's cluster-quality/structure check (e.g., low silhouette or a partition not reproducible as a sensible segmentation), so the output file is marked WRONG despite being correctly formatted.
628Analysis run on hardcoded/recalled numbers instead of the provided data filestaskda-code
Applies when
task -- the task ships a data directory/README and the scripts contain literal data values (arrays, dicts, totals) typed into the code rather than loaded from those files.
Pattern
The agent's exploration script only lists/inspects the data directory (or errors out on it) and never successfully loads the actual records; the analysis script then reconstructs the dataset from the agent's background knowledge of the topic, so every downstream statistic, interval, and output row is computed on numbers that were never validated against the shipped data.
Detection procedure
  1. Read the task/README and confirm that input data files are supplied and are the intended source of the quantities being estimated.
  2. Scan every analysis script for read_csv/read_*/file-open calls on the supplied paths; flag scripts whose inputs are literal in-code numbers, and check whether the exploration script actually printed real rows/shapes from the files or merely directory contents.
  3. Cross-check the in-code totals/row counts against anything the exploration output did reveal (file names, shapes, sample rows); if no such check exists anywhere, the analysis is unverified.
  4. Confirm the reported result derives from file-loaded data; if it derives only from hardcoded values, flag as inadequate regardless of how sound the statistical method looks.
Discriminator
Legitimate hardcoding is limited to constants stated in the task itself (thresholds, seeds, group boundaries, known parameter values) while the observations still come from the files; a violation is when the observational data (counts, rates, per-unit records) is invented in code and no script ever reads or validates the provided files.
Consequence
The output file contains a plausibly formatted interval computed from wrong inputs, so the grader's value check on the expected result file fails even though the bootstrap code itself is correct.
id 2d4554e5e1ef · mined from da-code dacode-data-sa-031@s11
raw text (what the judge reads)
### Analysis run on hardcoded/recalled numbers instead of the provided data files
- **Applies when**: `task` -- the task ships a data directory/README and the scripts contain literal data values (arrays, dicts, totals) typed into the code rather than loaded from those files.
- **Pattern**: The agent's exploration script only lists/inspects the data directory (or errors out on it) and never successfully loads the actual records; the analysis script then reconstructs the dataset from the agent's background knowledge of the topic, so every downstream statistic, interval, and output row is computed on numbers that were never validated against the shipped data.
- **Detection procedure**:
  1. Read the task/README and confirm that input data files are supplied and are the intended source of the quantities being estimated.
  2. Scan every analysis script for `read_csv`/`read_*`/file-open calls on the supplied paths; flag scripts whose inputs are literal in-code numbers, and check whether the exploration script actually printed real rows/shapes from the files or merely directory contents.
  3. Cross-check the in-code totals/row counts against anything the exploration output did reveal (file names, shapes, sample rows); if no such check exists anywhere, the analysis is unverified.
  4. Confirm the reported result derives from file-loaded data; if it derives only from hardcoded values, flag as inadequate regardless of how sound the statistical method looks.
- **Discriminator**: Legitimate hardcoding is limited to constants stated in the task itself (thresholds, seeds, group boundaries, known parameter values) while the observations still come from the files; a violation is when the observational data (counts, rates, per-unit records) is invented in code and no script ever reads or validates the provided files.
- **Consequence**: The output file contains a plausibly formatted interval computed from wrong inputs, so the grader's value check on the expected result file fails even though the bootstrap code itself is correct.
629Ensemble/model chosen without comparing it to its own components or any baseline on the scored metrictaskda-code
Applies when
task -- the deliverable is scored by a continuous quality metric (log loss, RMSE, AUC) and the scripts blend or pick models with default hyperparameters.
Pattern
The attempt fixes an arbitrary combination (e.g., equal-weight average of several off-the-shelf classifiers, defaults untuned, weak/uncalibrated members included) and runs cross-validation only on that one fixed combination — it never scores the individual members, a trivial baseline (class-prior/constant prediction), or an alternative configuration, so there is no evidence the submitted predictor is better than the obvious alternatives, and a demonstrably worse model can be shipped.
Detection procedure
  1. Read the task to identify the scoring metric and confirm the answer's quality (not just its format) determines correctness.
  2. In the scripts, list every model/config that is evaluated versus the one that is used to generate the submission; check whether per-component and baseline scores are computed on the same folds.
  3. Check whether the validation script's pipeline (features, encodings, scaling, blend weights) is identical to the submission script's, and whether any selection/tuning decision was made using those numbers.
  4. Check the reported/printed validation score against a plausible reference for the metric (e.g., the constant class-prior baseline) — if no such comparison exists anywhere, flag.
Discriminator
A real violation is a single unjustified configuration with at most one CV number and no per-model or baseline comparison; it is fine if the scripts evaluate several candidates (or tune) on consistent folds and select the best, even if the final model is a simple average, as long as that average was shown to beat its components/baseline.
Consequence
The submission is well-formed but its metric value lands above the grader's accuracy threshold (worse log loss than a properly tuned/selected model), so the expected output file is judged wrong despite passing all format checks.
id 6b0e8c0892b0 · mined from da-code dacode-ml-competition-005@s12
raw text (what the judge reads)
### Ensemble/model chosen without comparing it to its own components or any baseline on the scored metric
- **Applies when**: `task` -- the deliverable is scored by a continuous quality metric (log loss, RMSE, AUC) and the scripts blend or pick models with default hyperparameters.
- **Pattern**: The attempt fixes an arbitrary combination (e.g., equal-weight average of several off-the-shelf classifiers, defaults untuned, weak/uncalibrated members included) and runs cross-validation only on that one fixed combination — it never scores the individual members, a trivial baseline (class-prior/constant prediction), or an alternative configuration, so there is no evidence the submitted predictor is better than the obvious alternatives, and a demonstrably worse model can be shipped.
- **Detection procedure**:
  1. Read the task to identify the scoring metric and confirm the answer's quality (not just its format) determines correctness.
  2. In the scripts, list every model/config that is *evaluated* versus the one that is *used to generate the submission*; check whether per-component and baseline scores are computed on the same folds.
  3. Check whether the validation script's pipeline (features, encodings, scaling, blend weights) is identical to the submission script's, and whether any selection/tuning decision was made using those numbers.
  4. Check the reported/printed validation score against a plausible reference for the metric (e.g., the constant class-prior baseline) — if no such comparison exists anywhere, flag.
- **Discriminator**: A real violation is a single unjustified configuration with at most one CV number and no per-model or baseline comparison; it is fine if the scripts evaluate several candidates (or tune) on consistent folds and select the best, even if the final model is a simple average, as long as that average was shown to beat its components/baseline.
- **Consequence**: The submission is well-formed but its metric value lands above the grader's accuracy threshold (worse log loss than a properly tuned/selected model), so the expected output file is judged wrong despite passing all format checks.
630Statistical test run on the whole dataset with a default test, ignoring the scope and test type implied by the tasktaskda-code
Applies when
task -- the task asks for a p-value/decision from a hypothesis test on observational data whose rows span many sub-populations, eras, or competition/category levels, and the script feeds all rows into a stock parametric test.
Pattern
The agent skips any scoping step (no time-window, category, or subgroup filter), never inspects the distribution of the outcome variable, and calls a default two-sided parametric two-sample test, so both the analysis population and the test statistic differ from what the task intends.
Detection procedure
  1. From the task statement and README, list the population the comparison is supposed to be about and any qualifiers (time period, competition/tier, one-sided vs two-sided wording such as "greater than"/"higher").
  2. In the script, find the rows actually passed to the test: check for any filtering/subsetting and whether the sample sizes printed match the intended population rather than the full file.
  3. Check the test choice against the outcome's nature (bounded counts, heavy skew, tiny discrete range) — did the script test normality/shape or justify a parametric two-sided test, or just call the first function available?
  4. Compare the reported p-value's magnitude to what the intended sample sizes and effect size could plausibly produce; astronomically small p-values from hundreds of thousands of rows signal an unscoped, over-powered test.
Discriminator
A real violation is when the task/README implies a narrower population or a directional/nonparametric test and the script demonstrably uses everything with a default two-sided parametric call; it is fine if the agent explicitly checked scope and distribution and documented that the full sample and that test are the correct choice.
Consequence
The p-value is off by many orders of magnitude (and the decision may flip), so the saved result file fails the grader's value check even though the file format is correct.
id 34fd80e4bc47 · mined from da-code dacode-data-sa-001@s12
raw text (what the judge reads)
### Statistical test run on the whole dataset with a default test, ignoring the scope and test type implied by the task
- **Applies when**: `task` -- the task asks for a p-value/decision from a hypothesis test on observational data whose rows span many sub-populations, eras, or competition/category levels, and the script feeds all rows into a stock parametric test.
- **Pattern**: The agent skips any scoping step (no time-window, category, or subgroup filter), never inspects the distribution of the outcome variable, and calls a default two-sided parametric two-sample test, so both the analysis population and the test statistic differ from what the task intends.
- **Detection procedure**:
  1. From the task statement and README, list the population the comparison is supposed to be about and any qualifiers (time period, competition/tier, one-sided vs two-sided wording such as "greater than"/"higher").
  2. In the script, find the rows actually passed to the test: check for any filtering/subsetting and whether the sample sizes printed match the intended population rather than the full file.
  3. Check the test choice against the outcome's nature (bounded counts, heavy skew, tiny discrete range) — did the script test normality/shape or justify a parametric two-sided test, or just call the first function available?
  4. Compare the reported p-value's magnitude to what the intended sample sizes and effect size could plausibly produce; astronomically small p-values from hundreds of thousands of rows signal an unscoped, over-powered test.
- **Discriminator**: A real violation is when the task/README implies a narrower population or a directional/nonparametric test and the script demonstrably uses everything with a default two-sided parametric call; it is fine if the agent explicitly checked scope and distribution and documented that the full sample and that test are the correct choice.
- **Consequence**: The p-value is off by many orders of magnitude (and the decision may flip), so the saved result file fails the grader's value check even though the file format is correct.
631Silently dropping available input columns instead of building the full feature vectortaskda-code
Applies when
task -- the task asks to model/cluster "the dataset" and to persist the feature vector alongside the output, but the raw table contains a mix of numeric, categorical, and date-like columns.
Pattern
The script hard-codes a convenience subset of numeric columns, discards non-numeric ones (categoricals, dates, derived identifiers) with no encoding or stated justification, then writes those raw subset values out as Feature_i, so the persisted feature vector's width/content does not correspond to the dataset's usable feature space that the grader reconstructs.
Detection procedure
  1. Read the task/README and list every attribute available, marking which are true identifiers/metadata versus genuine features.
  2. In the script, find the feature-selection list and compare it against that inventory: are non-numeric columns simply omitted rather than encoded (one-hot/ordinal/date-to-numeric), and are any numeric columns dropped without comment?
  3. Check what is written to the output file — are the persisted Feature_i columns the full preprocessed/encoded matrix actually fed to the algorithm (and in a documented order), or a partial raw subset?
  4. Look for any sanity note in the answer reconciling the number of output feature columns with the number of usable attributes in the source data.
Discriminator
A real violation is dropping informative columns purely because they are strings/dates or because they were not in a hand-typed list. It is fine to drop unique row identifiers, constant columns, or leakage-prone fields, or to reduce dimensions, if the script states the reason and the persisted feature vector matches the matrix actually used by the model.
Consequence
The saved file has a feature matrix of the wrong shape/content relative to the reference pipeline, so a file-level comparison of Feature_* columns (and the cluster structure derived from them) fails even though the code runs and reports plausible metrics.
id 0af2e1d38541 · mined from da-code dacode-ml-cluster-014@s12
raw text (what the judge reads)
### Silently dropping available input columns instead of building the full feature vector
- **Applies when**: `task` -- the task asks to model/cluster "the dataset" and to persist the feature vector alongside the output, but the raw table contains a mix of numeric, categorical, and date-like columns.
- **Pattern**: The script hard-codes a convenience subset of numeric columns, discards non-numeric ones (categoricals, dates, derived identifiers) with no encoding or stated justification, then writes those raw subset values out as `Feature_i`, so the persisted feature vector's width/content does not correspond to the dataset's usable feature space that the grader reconstructs.
- **Detection procedure**:
  1. Read the task/README and list every attribute available, marking which are true identifiers/metadata versus genuine features.
  2. In the script, find the feature-selection list and compare it against that inventory: are non-numeric columns simply omitted rather than encoded (one-hot/ordinal/date-to-numeric), and are any numeric columns dropped without comment?
  3. Check what is written to the output file — are the persisted `Feature_i` columns the full preprocessed/encoded matrix actually fed to the algorithm (and in a documented order), or a partial raw subset?
  4. Look for any sanity note in the answer reconciling the number of output feature columns with the number of usable attributes in the source data.
- **Discriminator**: A real violation is dropping informative columns purely because they are strings/dates or because they were not in a hand-typed list. It is fine to drop unique row identifiers, constant columns, or leakage-prone fields, or to reduce dimensions, **if** the script states the reason and the persisted feature vector matches the matrix actually used by the model.
- **Consequence**: The saved file has a feature matrix of the wrong shape/content relative to the reference pipeline, so a file-level comparison of `Feature_*` columns (and the cluster structure derived from them) fails even though the code runs and reports plausible metrics.
632No held-out validation of model/ensemble choices, plus unjustified rounding of continuous predictionstaskda-code
Applies when
task -- a script fits one or more models on the full training set, blends them with hand-picked weights, and/or post-processes predictions (rounding, clipping, casting) before writing the submission file.
Pattern
The attempt never computes an out-of-fold or hold-out score under the competition's evaluation metric; imports for cross-validation are present but unused. Hyperparameters and ensemble weights are chosen by intuition, and the continuous regression output is snapped to integers/clipped because the target "looks" integer — a transformation that is only correct if the metric is defined on integer classes, and which strictly increases error for squared/log-error metrics.
Detection procedure
  1. Read the task/README (and any sample submission or metric statement) to determine the evaluation metric and whether the target must be integer-valued.
  2. Scan the script for any evaluation step: a train/validation split, cross_val_score/KFold loop, or printed score in the metric's units. Note whether the printed diagnostics are only prediction summary statistics rather than an error estimate.
  3. Check how model settings and blend weights were chosen — are they tied to any measured score, or are they literals with no supporting evidence?
  4. Inspect the post-processing on predictions (rounding, astype(int), clipping) and ask whether the metric rewards it; check the submitted file's value distribution for signs of coarse/degraded predictions.
Discriminator
A real violation is zero measured generalization error anywhere in the pipeline (or rounding applied under a continuous error metric); it is fine if the script reports CV/hold-out scores that justify the chosen model and weights, or if the task explicitly requires integer-valued output and rounding is the documented convention.
Consequence
The submission is syntactically valid but scores worse than a simple validated baseline — the grader's metric threshold is not met and the result is marked wrong, with no internal score available to have caught it beforehand.
id 6e64fd8953bd · mined from da-code dacode-ml-competition-009@s12
raw text (what the judge reads)
### No held-out validation of model/ensemble choices, plus unjustified rounding of continuous predictions
- **Applies when**: `task` -- a script fits one or more models on the full training set, blends them with hand-picked weights, and/or post-processes predictions (rounding, clipping, casting) before writing the submission file.
- **Pattern**: The attempt never computes an out-of-fold or hold-out score under the competition's evaluation metric; imports for cross-validation are present but unused. Hyperparameters and ensemble weights are chosen by intuition, and the continuous regression output is snapped to integers/clipped because the target "looks" integer — a transformation that is only correct if the metric is defined on integer classes, and which strictly increases error for squared/log-error metrics.
- **Detection procedure**:
  1. Read the task/README (and any sample submission or metric statement) to determine the evaluation metric and whether the target must be integer-valued.
  2. Scan the script for any evaluation step: a train/validation split, `cross_val_score`/KFold loop, or printed score in the metric's units. Note whether the printed diagnostics are only prediction summary statistics rather than an error estimate.
  3. Check how model settings and blend weights were chosen — are they tied to any measured score, or are they literals with no supporting evidence?
  4. Inspect the post-processing on predictions (rounding, `astype(int)`, clipping) and ask whether the metric rewards it; check the submitted file's value distribution for signs of coarse/degraded predictions.
- **Discriminator**: A real violation is zero measured generalization error anywhere in the pipeline (or rounding applied under a continuous error metric); it is fine if the script reports CV/hold-out scores that justify the chosen model and weights, or if the task explicitly requires integer-valued output and rounding is the documented convention.
- **Consequence**: The submission is syntactically valid but scores worse than a simple validated baseline — the grader's metric threshold is not met and the result is marked wrong, with no internal score available to have caught it beforehand.
633Output rows/values silently altered by preprocessing (dropped records, transformed features)taskda-code
Applies when
task -- the task asks for a per-record result file whose rows correspond to the input records and whose columns are the feature values plus a derived label/prediction.
Pattern
The script drops or filters records with missing/odd values (or subsets columns) instead of imputing, and writes the internally scaled/encoded feature matrix rather than the features as they should appear, so the saved file has fewer rows than the input and values that do not match the source data.
Detection procedure
  1. From the task, note the required output file, its column naming/ordering convention, and the implied number of rows (one per input record).
  2. In the scripts, locate every row-reducing operation (dropna, filtering, deduplication) and every value-transforming step (scaling, log, encoding) applied before the write, and check whether the written frame is the transformed matrix or the original feature values.
  3. Compare the reported/actual output shape and a few written values against the raw input count and raw values; confirm row alignment with the source order.
  4. Check the answer text for admissions like "after filtering incomplete rows" or "all features are standardized" in the saved file.
Discriminator
A real violation is when the written file cannot be aligned row-for-row with the input (fewer rows) or its feature cells differ from the values the task implies; it is fine if transformation happens only inside the model pipeline while the saved file retains all records with the expected feature representation, or if the task explicitly authorizes the filtering/scaling.
Consequence
The result file fails shape/row-count and value-match checks against the expected per-record output, so the grader marks the file WRONG/MISSING even if the clustering/model logic itself was reasonable.
id 9b41f14027c0 · mined from da-code dacode-ml-cluster-009@s12
raw text (what the judge reads)
### Output rows/values silently altered by preprocessing (dropped records, transformed features)
- **Applies when**: `task` -- the task asks for a per-record result file whose rows correspond to the input records and whose columns are the feature values plus a derived label/prediction.
- **Pattern**: The script drops or filters records with missing/odd values (or subsets columns) instead of imputing, and writes the internally scaled/encoded feature matrix rather than the features as they should appear, so the saved file has fewer rows than the input and values that do not match the source data.
- **Detection procedure**:
  1. From the task, note the required output file, its column naming/ordering convention, and the implied number of rows (one per input record).
  2. In the scripts, locate every row-reducing operation (dropna, filtering, deduplication) and every value-transforming step (scaling, log, encoding) applied before the write, and check whether the written frame is the transformed matrix or the original feature values.
  3. Compare the reported/actual output shape and a few written values against the raw input count and raw values; confirm row alignment with the source order.
  4. Check the answer text for admissions like "after filtering incomplete rows" or "all features are standardized" in the saved file.
- **Discriminator**: A real violation is when the written file cannot be aligned row-for-row with the input (fewer rows) or its feature cells differ from the values the task implies; it is fine if transformation happens only inside the model pipeline while the saved file retains all records with the expected feature representation, or if the task explicitly authorizes the filtering/scaling.
- **Consequence**: The result file fails shape/row-count and value-match checks against the expected per-record output, so the grader marks the file WRONG/MISSING even if the clustering/model logic itself was reasonable.
634Degenerate clusters from unhandled skew/outliers in unsupervised segmentationtaskda-code
Applies when
task -- the script builds aggregated numeric features (counts, sums, monetary totals, averages) and runs a distance-based clustering algorithm to produce required group labels.
Pattern
Features are only z-score standardized, with no log/rank transform, winsorizing, or outlier screening, so a few extreme records dominate the Euclidean geometry; the resulting solution assigns almost all records to one or two huge clusters and a handful (1–20) to "outlier" clusters, and the agent accepts this (often also picking k by a silhouette score that is inflated by exactly those outlier splits, or overriding the metric by hand with no stated rule).
Detection procedure
  1. Read the task to confirm the goal is a meaningful segmentation of the entities with labels written in a specified schema.
  2. In the script, check the feature-construction step for any treatment of heavy-tailed/unbounded quantities (log1p, quantile/robust scaling, capping, explicit outlier removal) before scaling — note if scaling is the only step.
  3. Check how k is chosen: is there a stated, reproducible criterion, or is a value hardcoded after seeing that the metric-optimal k produced near-empty clusters?
  4. In the reported output, inspect the cluster size distribution and per-cluster means; flag if any cluster holds a negligible fraction of rows or if one cluster holds nearly everything.
Discriminator
A genuinely small cluster is fine if the script deliberately handled skew (transform/robust scaling) and the small group is still substantively sized and interpretable; the violation is raw skewed features plus clusters of a few individual records that are effectively outlier detection, not segmentation.
Consequence
The saved label file has the right column names but a near-trivial partition, so the grader's check on the clustering result (cluster balance/quality/expected structure) fails even though the pipeline ran without error.
id 454f6ea7aac8 · mined from da-code dacode-ml-cluster-016@s12
raw text (what the judge reads)
### Degenerate clusters from unhandled skew/outliers in unsupervised segmentation
- **Applies when**: `task` -- the script builds aggregated numeric features (counts, sums, monetary totals, averages) and runs a distance-based clustering algorithm to produce required group labels.
- **Pattern**: Features are only z-score standardized, with no log/rank transform, winsorizing, or outlier screening, so a few extreme records dominate the Euclidean geometry; the resulting solution assigns almost all records to one or two huge clusters and a handful (1–20) to "outlier" clusters, and the agent accepts this (often also picking k by a silhouette score that is inflated by exactly those outlier splits, or overriding the metric by hand with no stated rule).
- **Detection procedure**:
  1. Read the task to confirm the goal is a meaningful segmentation of the entities with labels written in a specified schema.
  2. In the script, check the feature-construction step for any treatment of heavy-tailed/unbounded quantities (log1p, quantile/robust scaling, capping, explicit outlier removal) before scaling — note if scaling is the only step.
  3. Check how k is chosen: is there a stated, reproducible criterion, or is a value hardcoded after seeing that the metric-optimal k produced near-empty clusters?
  4. In the reported output, inspect the cluster size distribution and per-cluster means; flag if any cluster holds a negligible fraction of rows or if one cluster holds nearly everything.
- **Discriminator**: A genuinely small cluster is fine if the script deliberately handled skew (transform/robust scaling) and the small group is still substantively sized and interpretable; the violation is raw skewed features plus clusters of a few individual records that are effectively outlier detection, not segmentation.
- **Consequence**: The saved label file has the right column names but a near-trivial partition, so the grader's check on the clustering result (cluster balance/quality/expected structure) fails even though the pipeline ran without error.
635Boolean/categorical answer emitted in a different literal form than the spectaskinfiagent-dabench
Applies when
task -- the answer template fixes the allowed tokens for a field (e.g., "True or False", a label name, a unit string) and the script prints that field from a computed value.
Pattern
The script formats the value with a language-native conversion (str(x).lower(), int(bool), repr, numpy type printing) instead of the exact token specified, so the reported literal (false, 0, False.) does not match the required literal even though the underlying computation is right.
Detection procedure
1. Read the answer format section and note the exact literal tokens/casing the task requires for each field. 2. Find the print/format statement in the script for each such field and mentally evaluate what string it produces (watch for .lower(), .upper(), f-string of a numpy bool, or %d). 3. Compare the produced literal character-for-character with the specified token. 4. Check the submitted answer block for the same mismatch.
Discriminator
A real violation is a literal that differs in case/type/wording from the stated token; a look-alike that is fine is a field whose format the task left unconstrained (e.g., "a number") or where the script prints exactly the specified token via a mapping.
Consequence
The numeric fields pass but the mis-cased/mis-typed field is scored WRONG/MISSING, so the overall attempt is marked incorrect despite correct analysis.
id 0bd005e33a91 · mined from infiagent-dabench dabench-733@s12
raw text (what the judge reads)
### Boolean/categorical answer emitted in a different literal form than the spec
- **Applies when**: `task` -- the answer template fixes the allowed tokens for a field (e.g., "True or False", a label name, a unit string) and the script prints that field from a computed value.
- **Pattern**: The script formats the value with a language-native conversion (`str(x).lower()`, `int(bool)`, `repr`, numpy type printing) instead of the exact token specified, so the reported literal (`false`, `0`, `False.`) does not match the required literal even though the underlying computation is right.
- **Detection procedure**: 1. Read the answer format section and note the exact literal tokens/casing the task requires for each field. 2. Find the print/format statement in the script for each such field and mentally evaluate what string it produces (watch for `.lower()`, `.upper()`, f-string of a numpy bool, or `%d`). 3. Compare the produced literal character-for-character with the specified token. 4. Check the submitted answer block for the same mismatch.
- **Discriminator**: A real violation is a literal that differs in case/type/wording from the stated token; a look-alike that is fine is a field whose format the task left unconstrained (e.g., "a number") or where the script prints exactly the specified token via a mapping.
- **Consequence**: The numeric fields pass but the mis-cased/mis-typed field is scored WRONG/MISSING, so the overall attempt is marked incorrect despite correct analysis.
636Output artifact never validated against the requested file contracttaskda-code
Applies when
task -- The task requires writing predictions/results to a named file with a specified column name, one row per input record, aligned to the given test set.
Pattern
The agent focuses on modeling and validation metrics, then writes the output file without any explicit check that the saved file exists at the required path, has exactly the required column name (and no unintended extra/index columns), has exactly as many rows as the test input, and preserves the test set's row order/key alignment; the final report cites model scores instead of showing the head/shape of the saved file.
Detection procedure
  1. Read the task and record the exact required filename, column name(s), row count implied by the test input, and any ordering/format constraints.
  2. In the scripts, find the write call (e.g. to_csv) and check: path matches, the DataFrame is built from the test rows in their original order (no dropped rows from dropna/merge, no re-sorting, no index=True leaking an unnamed column), and the target column is named exactly as required.
  3. Look for a post-write verification step (re-read the file; assert shape, columns, no NaNs, plausible value range) — absence of any such assertion is the flag.
  4. Check the final answer: does it report the saved file's shape/columns/head, or only training/validation metrics and summary statistics?
Discriminator
A real violation is when nothing in the scripts or answer establishes that the saved file matches the required name/column/row count/order — merges, NaN filtering, or aggregation steps upstream make silent row loss or misalignment plausible. It is not a violation if the script (or answer) demonstrably re-reads the output and asserts shape, column names, and alignment to the test keys, even if no metric on test data can be computed.
Consequence
The grader reads the expected file and finds it missing, misnamed, wrong column header, or with a row count/order that does not match the test set, so the check fails outright regardless of how good the model's validation RMSE/R² looked.
id 0e9550e2febd · mined from da-code dacode-ml-regression-002@s12
raw text (what the judge reads)
### Output artifact never validated against the requested file contract
- **Applies when**: `task` -- The task requires writing predictions/results to a named file with a specified column name, one row per input record, aligned to the given test set.
- **Pattern**: The agent focuses on modeling and validation metrics, then writes the output file without any explicit check that the saved file exists at the required path, has exactly the required column name (and no unintended extra/index columns), has exactly as many rows as the test input, and preserves the test set's row order/key alignment; the final report cites model scores instead of showing the head/shape of the saved file.
- **Detection procedure**:
  1. Read the task and record the exact required filename, column name(s), row count implied by the test input, and any ordering/format constraints.
  2. In the scripts, find the write call (e.g. `to_csv`) and check: path matches, the DataFrame is built from the test rows in their original order (no dropped rows from `dropna`/merge, no re-sorting, no `index=True` leaking an unnamed column), and the target column is named exactly as required.
  3. Look for a post-write verification step (re-read the file; assert shape, columns, no NaNs, plausible value range) — absence of any such assertion is the flag.
  4. Check the final answer: does it report the saved file's shape/columns/head, or only training/validation metrics and summary statistics?
- **Discriminator**: A real violation is when nothing in the scripts or answer establishes that the saved file matches the required name/column/row count/order — merges, NaN filtering, or aggregation steps upstream make silent row loss or misalignment plausible. It is not a violation if the script (or answer) demonstrably re-reads the output and asserts shape, column names, and alignment to the test keys, even if no metric on test data can be computed.
- **Consequence**: The grader reads the expected file and finds it missing, misnamed, wrong column header, or with a row count/order that does not match the test set, so the check fails outright regardless of how good the model's validation RMSE/R² looked.
637Hard-coded/assumed input data instead of loading the provided files and instructionstaskda-code
Applies when
task -- the task points to specific data files and an instruction file (e.g. a tips/README describing the exact procedure), and the script must read them to produce the requested statistic.
Pattern
The script embeds literal data arrays (or hard-coded group sizes, means, thresholds, parameter choices) transcribed from memory or from a textbook example, and never opens the supplied data or instruction files; the reported answer is then computed on data that does not match the actual inputs (wrong group definitions, wrong n, wrong values), and any procedural detail specified in the instruction file (which groups to compare, one- vs two-sided, number of replicates, rounding, output column names) is guessed rather than read.
Detection procedure
  1. From the task statement, list the input artifacts that must be consulted (data file(s), instruction/tips file) and the exact output file/columns required.
  2. Scan the script for any file-reading call (read_csv, open, load, etc.) for each of those artifacts; flag if data appear as inline literals or if the instruction file is never read.
  3. Cross-check the script's implicit assumptions (row counts, group labels/keys, units, means printed in the answer) against the real files or the task/README description; flag any mismatch in size or grouping.
  4. Check the answer text for signs of unverified assumptions ("assumed", memorized sample values, symmetric/rounded group sizes) and for procedural choices not traceable to the instruction file.
Discriminator
Legitimate: the script loads the provided files and only hard-codes documented constants (seed, replicate count stated in the instructions) after confirming them in the file; a printed sanity check of shapes/group sizes matches the source data. Violation: the numeric data themselves, group membership, or the procedure come from outside the provided files, with no verification step.
Consequence
The computed statistic is derived from the wrong inputs, so the written output file contains a value outside the accepted tolerance and the file check fails even though the code runs without error.
id 69a97cec5b65 · mined from da-code dacode-data-sa-028@s12
raw text (what the judge reads)
### Hard-coded/assumed input data instead of loading the provided files and instructions
- **Applies when**: `task` -- the task points to specific data files and an instruction file (e.g. a tips/README describing the exact procedure), and the script must read them to produce the requested statistic.
- **Pattern**: The script embeds literal data arrays (or hard-coded group sizes, means, thresholds, parameter choices) transcribed from memory or from a textbook example, and never opens the supplied data or instruction files; the reported answer is then computed on data that does not match the actual inputs (wrong group definitions, wrong n, wrong values), and any procedural detail specified in the instruction file (which groups to compare, one- vs two-sided, number of replicates, rounding, output column names) is guessed rather than read.
- **Detection procedure**:
  1. From the task statement, list the input artifacts that must be consulted (data file(s), instruction/tips file) and the exact output file/columns required.
  2. Scan the script for any file-reading call (`read_csv`, `open`, `load`, etc.) for each of those artifacts; flag if data appear as inline literals or if the instruction file is never read.
  3. Cross-check the script's implicit assumptions (row counts, group labels/keys, units, means printed in the answer) against the real files or the task/README description; flag any mismatch in size or grouping.
  4. Check the answer text for signs of unverified assumptions ("assumed", memorized sample values, symmetric/rounded group sizes) and for procedural choices not traceable to the instruction file.
- **Discriminator**: Legitimate: the script loads the provided files and only hard-codes documented constants (seed, replicate count stated in the instructions) after confirming them in the file; a printed sanity check of shapes/group sizes matches the source data. Violation: the numeric data themselves, group membership, or the procedure come from outside the provided files, with no verification step.
- **Consequence**: The computed statistic is derived from the wrong inputs, so the written output file contains a value outside the accepted tolerance and the file check fails even though the code runs without error.
638Ignoring the task-referenced specification file when defining categories/bins (and not verifying the outputs against it)taskda-code
Applies when
task -- the prompt points to an auxiliary document (README/spec/config) that defines how values must be bucketed, filtered, ordered, or formatted, and the deliverable is a derived chart/table/array saved to disk.
Pattern
The attempt never opens or quotes the referenced spec; it invents plausible-looking bins/labels (or reuses the raw category values already present in the data), counts every row including headers/blank or "prefer not to say" responses, and then reports a narrative summary instead of demonstrating that each required artifact was produced from those spec-defined groups.
Detection procedure
  1. Read the task and list every external file it cites plus every explicitly named output artifact and formatting constraint (title/labels/file names).
  2. Search the scripts for a read of that cited file and for a mapping/bin definition traceable to its contents; if the bins appear as a hard-coded list with no provenance, flag it.
  3. Compare the reported categories/labels/order against what the spec would imply and against the raw column's own values — identical to the raw values or a different granularity are both red flags.
  4. Check that the totals reconcile: does the sum of counts equal the number of valid, in-scope rows (not the dataset's headline response count), and is every promised output file actually written by the script?
Discriminator
A real violation is when no script evidence shows the spec was read and the grouping/labels cannot be derived from it, or the counts sum to the raw row count without any handling of invalid/excluded rows. It is fine if the script reads the spec (or explicitly reproduces its bins with a comment citing it) and the totals differ from the headline count for a documented filtering reason.
Consequence
The saved artifacts (image plus any serialized counts/plot data) encode the wrong buckets or wrong per-bucket values, so every file-level check fails even though the chart looks well-formed and correctly titled.
id 58d354a30180 · mined from da-code dacode-plot-bar-005@s12
raw text (what the judge reads)
### Ignoring the task-referenced specification file when defining categories/bins (and not verifying the outputs against it)
- **Applies when**: `task` -- the prompt points to an auxiliary document (README/spec/config) that defines how values must be bucketed, filtered, ordered, or formatted, and the deliverable is a derived chart/table/array saved to disk.
- **Pattern**: The attempt never opens or quotes the referenced spec; it invents plausible-looking bins/labels (or reuses the raw category values already present in the data), counts every row including headers/blank or "prefer not to say" responses, and then reports a narrative summary instead of demonstrating that each required artifact was produced from those spec-defined groups.
- **Detection procedure**:
  1. Read the task and list every external file it cites plus every explicitly named output artifact and formatting constraint (title/labels/file names).
  2. Search the scripts for a read of that cited file and for a mapping/bin definition traceable to its contents; if the bins appear as a hard-coded list with no provenance, flag it.
  3. Compare the reported categories/labels/order against what the spec would imply and against the raw column's own values — identical to the raw values or a different granularity are both red flags.
  4. Check that the totals reconcile: does the sum of counts equal the number of *valid, in-scope* rows (not the dataset's headline response count), and is every promised output file actually written by the script?
- **Discriminator**: A real violation is when no script evidence shows the spec was read and the grouping/labels cannot be derived from it, or the counts sum to the raw row count without any handling of invalid/excluded rows. It is fine if the script reads the spec (or explicitly reproduces its bins with a comment citing it) and the totals differ from the headline count for a documented filtering reason.
- **Consequence**: The saved artifacts (image plus any serialized counts/plot data) encode the wrong buckets or wrong per-bucket values, so every file-level check fails even though the chart looks well-formed and correctly titled.
639Output artifact and schema not produced as literally specifiedtaskda-code
Applies when
task -- the prompt dictates an exact answer skeleton (key names, value containers such as [...]) and/or an expected result file that the scripts must write.
Pattern
The agent computes a plausible number, then reports it only in chat prose as a bare scalar/string, never writing the required result file and never matching the literal container/key structure shown in the prompt (e.g. list-valued fields collapsed to a single value, extra/renamed keys, no saved script that emits the file).
Detection procedure
  1. Read the task and copy out the exact required output: file name/path, key spellings, value types (scalar vs list), and any rounding/unit constraints.
  2. Search the scripts for a write step to that file (json.dump/to_json/to_csv to the named path) and confirm the serialized object's keys and value types are built to match the skeleton, not an ad-hoc dict.
  3. Compare the submitted answer field-by-field against the skeleton: same keys, same container type for each value, same numeric formatting.
  4. Flag if the required file is absent, if no script produces it, or if any field's type/name deviates.
Discriminator
A real violation is a missing output file or a structural/type mismatch with the stated skeleton; a look-alike that is fine is an answer that writes the required file with the exact keys and containers and merely differs in cosmetic whitespace or key ordering.
Consequence
The grader looks for the expected result file and exact schema, so it reports WRONG/MISSING and scores 0 even if the underlying computation happened to be right.
id 43bfcb064eef · mined from da-code dacode-di-text-002@s12
raw text (what the judge reads)
### Output artifact and schema not produced as literally specified
- **Applies when**: `task` -- the prompt dictates an exact answer skeleton (key names, value containers such as `[...]`) and/or an expected result file that the scripts must write.
- **Pattern**: The agent computes a plausible number, then reports it only in chat prose as a bare scalar/string, never writing the required result file and never matching the literal container/key structure shown in the prompt (e.g. list-valued fields collapsed to a single value, extra/renamed keys, no saved script that emits the file).
- **Detection procedure**:
  1. Read the task and copy out the exact required output: file name/path, key spellings, value types (scalar vs list), and any rounding/unit constraints.
  2. Search the scripts for a write step to that file (`json.dump`/`to_json`/`to_csv` to the named path) and confirm the serialized object's keys and value types are built to match the skeleton, not an ad-hoc dict.
  3. Compare the submitted answer field-by-field against the skeleton: same keys, same container type for each value, same numeric formatting.
  4. Flag if the required file is absent, if no script produces it, or if any field's type/name deviates.
- **Discriminator**: A real violation is a missing output file or a structural/type mismatch with the stated skeleton; a look-alike that is fine is an answer that writes the required file with the exact keys and containers and merely differs in cosmetic whitespace or key ordering.
- **Consequence**: The grader looks for the expected result file and exact schema, so it reports WRONG/MISSING and scores 0 even if the underlying computation happened to be right.
640Output schema invented instead of derived from the task's required formattaskda-code
Applies when
task -- the task says results must be saved to a named file "following the required format", and the script builds that file with column names, index, ordering, or value conventions chosen by the agent.
Pattern
The agent hard-codes an ad-hoc schema (self-invented column labels, extra/missing key columns, index written or dropped, values expressed in a self-chosen convention such as excess-vs-growth, decimal-vs-percent, unrounded) without ever locating or reproducing the format the task/README/reference materials actually prescribe, and without checking the artifact after writing it.
Detection procedure
  1. Read the task/README for any statement about the output artifact: file name, expected columns/labels, one column per strategy vs. long format, units, rounding, index/date handling.
  2. Read the script's construction of the saved artifact and list exactly the header names, row count, ordering, and value definition it produces.
  3. Check whether the script does anything to confirm this matches the prescribed format (inspecting a template/example file, aligning names to those used in the source data or documentation, re-reading the written file).
  4. Compare the answer text's description of the artifact with step 1; if the schema originates only from the agent's own naming choices, flag it.
Discriminator
Fine if the task truly leaves the schema open, or the agent explicitly matched documented/example column names, units, and value definition; a violation is when a format is specified or implied (or a template exists) and the agent substitutes an unverified invention, or when a defensible alternative convention (e.g., cumulative growth factor vs. cumulative percentage change) is chosen with no justification or check.
Consequence
The numeric analysis may be plausible, but the file comparison against the expected artifact fails on header/column names, shape, or value convention, scoring 0 even though the reported summary numbers look reasonable.
id c9b56e930963 · mined from da-code dacode-dm-csv-050@s12
raw text (what the judge reads)
### Output schema invented instead of derived from the task's required format
- **Applies when**: `task` -- the task says results must be saved to a named file "following the required format", and the script builds that file with column names, index, ordering, or value conventions chosen by the agent.
- **Pattern**: The agent hard-codes an ad-hoc schema (self-invented column labels, extra/missing key columns, index written or dropped, values expressed in a self-chosen convention such as excess-vs-growth, decimal-vs-percent, unrounded) without ever locating or reproducing the format the task/README/reference materials actually prescribe, and without checking the artifact after writing it.
- **Detection procedure**:
  1. Read the task/README for any statement about the output artifact: file name, expected columns/labels, one column per strategy vs. long format, units, rounding, index/date handling.
  2. Read the script's construction of the saved artifact and list exactly the header names, row count, ordering, and value definition it produces.
  3. Check whether the script does anything to confirm this matches the prescribed format (inspecting a template/example file, aligning names to those used in the source data or documentation, re-reading the written file).
  4. Compare the answer text's description of the artifact with step 1; if the schema originates only from the agent's own naming choices, flag it.
- **Discriminator**: Fine if the task truly leaves the schema open, or the agent explicitly matched documented/example column names, units, and value definition; a violation is when a format is specified or implied (or a template exists) and the agent substitutes an unverified invention, or when a defensible alternative convention (e.g., cumulative growth factor vs. cumulative percentage change) is chosen with no justification or check.
- **Consequence**: The numeric analysis may be plausible, but the file comparison against the expected artifact fails on header/column names, shape, or value convention, scoring 0 even though the reported summary numbers look reasonable.
641Distribution statistic computed on an uncleaned/unverified column, with no reported diagnosticstaskinfiagent-dabench
Applies when
task -- the task asks for a distribution summary or hypothesis test (normality test, skewness/kurtosis, mean/variance) on a single numeric column of a table.
Pattern
The attempt loads the table and feeds the raw column straight into the test/statistic without first coercing dtype, dropping missing values, or removing sentinel/placeholder codes (e.g. -999, 0-as-missing, blanks, text-coded nulls), and reports only the final verdict — no sample size, no NaN/sentinel count, no p-value — so the vector actually tested is never verified and no script is left to reproduce it.
Detection procedure
  1. Read the task and list every quantity the constraints say to report or use (e.g. the test statistic/p-value, the alpha threshold, the rounding rule).
  2. In the scripts, locate the exact expression passed to the test/statistic function; check for pd.to_numeric(..., errors=...), .dropna()/.notna(), and any filter of placeholder codes, plus a printed len()/describe()/value-count of that vector.
  3. Check whether the run log prints the intermediate diagnostics (n used, count of dropped rows, p-value) — if the answer's verdict cannot be traced to a printed number, treat it as unverified.
  4. Cross-check internal consistency: extreme shape statistics paired with a "not normal" verdict on a small n, or a near-symmetric shape paired with a rejection, should trigger a re-read of the column's raw values (min/max, unique values, dtype) before accepting.
Discriminator
A real violation is when the column plausibly contains non-numeric entries, nulls, or out-of-range sentinel values and the script neither inspects nor removes them, and no n/p-value is printed. It is not a violation if the script explicitly shows the column is clean (dtype numeric, zero NaNs, min/max within plausible range) and logs n and the p-value, even if it then does no further filtering.
Consequence
The test runs on a contaminated or wrong-length vector, so the p-value, skewness and kurtosis all shift; the reported normal/not-normal verdict flips relative to ground truth and every field of the answer is marked wrong, with no saved script to diagnose the discrepancy.
id 2fff2c902c93 · mined from infiagent-dabench dabench-298@s12
raw text (what the judge reads)
### Distribution statistic computed on an uncleaned/unverified column, with no reported diagnostics
- **Applies when**: `task` -- the task asks for a distribution summary or hypothesis test (normality test, skewness/kurtosis, mean/variance) on a single numeric column of a table.
- **Pattern**: The attempt loads the table and feeds the raw column straight into the test/statistic without first coercing dtype, dropping missing values, or removing sentinel/placeholder codes (e.g. -999, 0-as-missing, blanks, text-coded nulls), and reports only the final verdict — no sample size, no NaN/sentinel count, no p-value — so the vector actually tested is never verified and no script is left to reproduce it.
- **Detection procedure**:
  1. Read the task and list every quantity the constraints say to report or use (e.g. the test statistic/p-value, the alpha threshold, the rounding rule).
  2. In the scripts, locate the exact expression passed to the test/statistic function; check for `pd.to_numeric(..., errors=...)`, `.dropna()`/`.notna()`, and any filter of placeholder codes, plus a printed `len()`/`describe()`/value-count of that vector.
  3. Check whether the run log prints the intermediate diagnostics (n used, count of dropped rows, p-value) — if the answer's verdict cannot be traced to a printed number, treat it as unverified.
  4. Cross-check internal consistency: extreme shape statistics paired with a "not normal" verdict on a small n, or a near-symmetric shape paired with a rejection, should trigger a re-read of the column's raw values (min/max, unique values, dtype) before accepting.
- **Discriminator**: A real violation is when the column plausibly contains non-numeric entries, nulls, or out-of-range sentinel values and the script neither inspects nor removes them, and no n/p-value is printed. It is *not* a violation if the script explicitly shows the column is clean (dtype numeric, zero NaNs, min/max within plausible range) and logs n and the p-value, even if it then does no further filtering.
- **Consequence**: The test runs on a contaminated or wrong-length vector, so the p-value, skewness and kurtosis all shift; the reported normal/not-normal verdict flips relative to ground truth and every field of the answer is marked wrong, with no saved script to diagnose the discrepancy.
642Outlier removal that fails the direction-of-change sanity checktaskinfiagent-dabench
Applies when
task -- the task asks to detect outliers by a rule (e.g. IQR/z-score) and then recompute a summary statistic on the filtered data.
Pattern
The attempt applies the detection rule at one level of aggregation but recomputes the statistic on a different set of rows (e.g. filters within groups but averages over ungrouped/unfiltered records, drops whole groups instead of extreme records, or re-derives thresholds after filtering), and reports a "without-outliers" value that moves in a direction inconsistent with which tail was removed — without ever checking that inconsistency.
Detection procedure
  1. From the task, identify the unit the rule is applied to, the unit the statistic is computed over, and whether the flagged points lie mostly in the upper or lower tail.
  2. In the scripts, verify the filtered dataframe used for the recomputed statistic is exactly the original rows minus the flagged rows — same key/index, same aggregation level, no re-merge, re-grouping, or re-thresholding in between.
  3. Compare reported before/after values: removing predominantly upper-tail points must lower the mean (and lower-tail removals must raise it); also check the post-filter row count equals original count minus number of flagged points.
  4. Flag if the direction is wrong, the magnitude is negligible despite many extreme removals, or the counts don't reconcile.
Discriminator
A genuine violation shows a statistic that moves the wrong way (or barely moves) given the tail removed, or row counts that don't reconcile. It is fine if outliers exist in both tails and the net shift is small or opposite, provided the script explicitly shows the tail composition and the counts add up.
Consequence
The "before" statistic matches the reference but the post-removal statistic is far off (here, essentially unchanged/increased instead of dropping substantially), so that check fails and the answer is graded incorrect.
id b4428f79f23c · mined from infiagent-dabench dabench-62@s12
raw text (what the judge reads)
### Outlier removal that fails the direction-of-change sanity check
- **Applies when**: `task` -- the task asks to detect outliers by a rule (e.g. IQR/z-score) and then recompute a summary statistic on the filtered data.
- **Pattern**: The attempt applies the detection rule at one level of aggregation but recomputes the statistic on a different set of rows (e.g. filters within groups but averages over ungrouped/unfiltered records, drops whole groups instead of extreme records, or re-derives thresholds after filtering), and reports a "without-outliers" value that moves in a direction inconsistent with which tail was removed — without ever checking that inconsistency.
- **Detection procedure**:
  1. From the task, identify the unit the rule is applied to, the unit the statistic is computed over, and whether the flagged points lie mostly in the upper or lower tail.
  2. In the scripts, verify the filtered dataframe used for the recomputed statistic is exactly the original rows minus the flagged rows — same key/index, same aggregation level, no re-merge, re-grouping, or re-thresholding in between.
  3. Compare reported before/after values: removing predominantly upper-tail points must lower the mean (and lower-tail removals must raise it); also check the post-filter row count equals original count minus number of flagged points.
  4. Flag if the direction is wrong, the magnitude is negligible despite many extreme removals, or the counts don't reconcile.
- **Discriminator**: A genuine violation shows a statistic that moves the wrong way (or barely moves) given the tail removed, or row counts that don't reconcile. It is fine if outliers exist in both tails and the net shift is small or opposite, provided the script explicitly shows the tail composition and the counts add up.
- **Consequence**: The "before" statistic matches the reference but the post-removal statistic is far off (here, essentially unchanged/increased instead of dropping substantially), so that check fails and the answer is graded incorrect.
643Deliverable file not actually produced/verified (answer pasted in chat instead)taskda-code
Applies when
task -- the task requires writing predictions or results to a named output file with a specified column/schema, and one row per input record.
Pattern
The agent prints or pastes the result contents into its final answer (often truncated) without saved scripts that write the file to the expected path, and never verifies the file exists, has the required column name, and has exactly as many rows as the input.
Detection procedure
  1. Read the task and note the exact required output artifact: filename, path, column name(s), row count implied by the input.
  2. Inspect the scripts for an explicit write to that filename (e.g., a to_csv/save call) with the required header, and confirm the scripts were actually executed and retained.
  3. Check for a post-write validation step: reload the file, assert row count equals the number of input records and column names match exactly.
  4. Compare the agent's answer to the artifact: if the answer is a truncated dump of rows rather than a confirmation of a complete written file, flag it.
Discriminator
A real violation is missing/unverified file creation, a mismatched column name/header, or a row count that does not match the input; it is fine if the file is written and validated and the chat text is merely an optional preview of a verified artifact.
Consequence
The grader looks for the expected file and finds it missing, misnamed, short, or wrongly headed, scoring the submission as WRONG/MISSING regardless of prediction quality.
id 4fa749e05324 · mined from da-code dacode-ml-multi-011@s12
raw text (what the judge reads)
### Deliverable file not actually produced/verified (answer pasted in chat instead)
- **Applies when**: `task` -- the task requires writing predictions or results to a named output file with a specified column/schema, and one row per input record.
- **Pattern**: The agent prints or pastes the result contents into its final answer (often truncated) without saved scripts that write the file to the expected path, and never verifies the file exists, has the required column name, and has exactly as many rows as the input.
- **Detection procedure**:
  1. Read the task and note the exact required output artifact: filename, path, column name(s), row count implied by the input.
  2. Inspect the scripts for an explicit write to that filename (e.g., a to_csv/save call) with the required header, and confirm the scripts were actually executed and retained.
  3. Check for a post-write validation step: reload the file, assert row count equals the number of input records and column names match exactly.
  4. Compare the agent's answer to the artifact: if the answer is a truncated dump of rows rather than a confirmation of a complete written file, flag it.
- **Discriminator**: A real violation is missing/unverified file creation, a mismatched column name/header, or a row count that does not match the input; it is fine if the file is written and validated and the chat text is merely an optional preview of a verified artifact.
- **Consequence**: The grader looks for the expected file and finds it missing, misnamed, short, or wrongly headed, scoring the submission as WRONG/MISSING regardless of prediction quality.
644Deliverable is an inline/truncated prediction dump instead of a complete saved submission filetaskda-code
Applies when
task -- the task requires writing predictions to a named output file that mirrors a provided template (id column + prediction column, one row per test record).
Pattern
The agent reports predictions as text in its answer (or writes a partial/truncated file) without evidence that the required file was created on disk with exactly the template's columns and one row for every test row; no script is preserved showing the write step and no shape/format check is performed.
Detection procedure
  1. Read the task/template to determine the required output filename, column names/order, and expected row count (= number of rows in the test set).
  2. Search the scripts for an explicit write to that exact filename (e.g., to_csv(..., index=False)) and confirm the frame written contains the template's id values in the template's order.
  3. Check for a verification step in the script or answer: printed row count of the written file compared to the test set size, printed header, and prediction value range.
  4. Inspect the answer: if it consists of pasted rows, count them and compare to the test set size; treat a truncated or few-dozen-row listing with no confirmed on-disk artifact as failing.
Discriminator
A real violation is missing/partial/misnamed output or unverified row count and columns; it is fine if the file is written in full with the correct schema and the answer merely shows a preview clearly labeled as a head of the saved file, with the full-file row count and header reported.
Consequence
The grader looks for the expected result file and finds it missing, short of the required rows, or schema-mismatched, so the submission is scored as wrong/missing regardless of model quality.
id bbfe61ff1dc0 · mined from da-code dacode-ml-competition-008@s12
raw text (what the judge reads)
### Deliverable is an inline/truncated prediction dump instead of a complete saved submission file
- **Applies when**: `task` -- the task requires writing predictions to a named output file that mirrors a provided template (id column + prediction column, one row per test record).
- **Pattern**: The agent reports predictions as text in its answer (or writes a partial/truncated file) without evidence that the required file was created on disk with exactly the template's columns and one row for every test row; no script is preserved showing the write step and no shape/format check is performed.
- **Detection procedure**:
  1. Read the task/template to determine the required output filename, column names/order, and expected row count (= number of rows in the test set).
  2. Search the scripts for an explicit write to that exact filename (e.g., `to_csv(..., index=False)`) and confirm the frame written contains the template's id values in the template's order.
  3. Check for a verification step in the script or answer: printed row count of the written file compared to the test set size, printed header, and prediction value range.
  4. Inspect the answer: if it consists of pasted rows, count them and compare to the test set size; treat a truncated or few-dozen-row listing with no confirmed on-disk artifact as failing.
- **Discriminator**: A real violation is missing/partial/misnamed output or unverified row count and columns; it is *fine* if the file is written in full with the correct schema and the answer merely shows a preview clearly labeled as a head of the saved file, with the full-file row count and header reported.
- **Consequence**: The grader looks for the expected result file and finds it missing, short of the required rows, or schema-mismatched, so the submission is scored as wrong/missing regardless of model quality.
645Inventing a qualification/aggregation rule instead of deriving it from the provided spec or sample filetaskda-code
Applies when
task -- the task text (or a README/definition section, possibly truncated) states how a ranked/aggregated quantity must be defined and points to a sample output file, and the script must group records and pick a top-N.
Pattern
The script hard-codes a threshold, filter, or aggregation choice (e.g., a minimum-count cutoff, sum vs. mean, per-item vs. total) that the agent guessed rather than read from the stated definition or inferred from the reference/sample file, and then applies that same filter uniformly to every requested ranking, without ever loading the sample output to confirm column names, ordering, and index conventions.
Detection procedure
  1. In the task/README, list every explicit definitional constraint (qualification thresholds, which statistic, which units, rounding, sort order) and note whether any definition is incomplete/truncated.
  2. In the scripts, locate the corresponding constants and aggregation calls; check each one traces to a quoted requirement rather than an arbitrary literal, and check whether a filter meant for one ranking is being reused for the others.
  3. Check whether the script reads the provided sample/reference output file and validates its own output against it (columns, row count, header names, ordering/index).
  4. If any constant or aggregation is unsourced, or the sample file is never opened, flag the attempt.
Discriminator
A real violation is an unsourced literal/aggregation whose change would reorder the reported top-N, or a format never checked against the sample; it is fine if the constant is explicitly quoted from the task, or if the script demonstrates the ranking is insensitive to the choice (e.g., reports results under several thresholds and they agree).
Consequence
The output has plausible-looking but wrong entries and/or mismatched columns/ordering, so exact-match comparison against the expected file fails on all checks.
id 571c7ee2c44d · mined from da-code dacode-dm-csv-009@s12
raw text (what the judge reads)
### Inventing a qualification/aggregation rule instead of deriving it from the provided spec or sample file
- **Applies when**: `task` -- the task text (or a README/definition section, possibly truncated) states how a ranked/aggregated quantity must be defined and points to a sample output file, and the script must group records and pick a top-N.
- **Pattern**: The script hard-codes a threshold, filter, or aggregation choice (e.g., a minimum-count cutoff, sum vs. mean, per-item vs. total) that the agent guessed rather than read from the stated definition or inferred from the reference/sample file, and then applies that same filter uniformly to every requested ranking, without ever loading the sample output to confirm column names, ordering, and index conventions.
- **Detection procedure**:
  1. In the task/README, list every explicit definitional constraint (qualification thresholds, which statistic, which units, rounding, sort order) and note whether any definition is incomplete/truncated.
  2. In the scripts, locate the corresponding constants and aggregation calls; check each one traces to a quoted requirement rather than an arbitrary literal, and check whether a filter meant for one ranking is being reused for the others.
  3. Check whether the script reads the provided sample/reference output file and validates its own output against it (columns, row count, header names, ordering/index).
  4. If any constant or aggregation is unsourced, or the sample file is never opened, flag the attempt.
- **Discriminator**: A real violation is an unsourced literal/aggregation whose change would reorder the reported top-N, or a format never checked against the sample; it is fine if the constant is explicitly quoted from the task, or if the script demonstrates the ranking is insensitive to the choice (e.g., reports results under several thresholds and they agree).
- **Consequence**: The output has plausible-looking but wrong entries and/or mismatched columns/ordering, so exact-match comparison against the expected file fails on all checks.
646Fabricating input data or config instead of reading the provided filestaskda-code
Applies when
task -- the task references a supplied dataset and/or a specification file (e.g., a config describing plot/output formatting) and names the required output artifacts.
Pattern
The script never opens the provided data/spec files; it synthesizes random or hard-coded data, invents its own spec values, and even writes the spec file itself, then reports numbers derived entirely from the fabricated inputs. Required output artifacts named by the task are also skipped.
Detection procedure
  1. List every input file the task mentions (data files, config/spec files) and every output artifact expected.
  2. Search the scripts for reads of those inputs (read_csv, open(...), yaml.safe_load, etc.); flag any use of np.random, hard-coded sample rows, or code that writes a file that was supposed to be read.
  3. Check that every named output artifact is produced at the required path and format; flag missing ones.
  4. Compare reported counts/categories against what the real data could plausibly contain (row counts, category labels stated in the README/spec).
Discriminator
Genuine violation = the analysis pipeline's numbers come from generated/assumed data, or the spec is authored rather than parsed. Fine = the script reads the real inputs and only uses synthetic data for a temporary smoke test, or fills a documented default for a field genuinely absent from the spec.
Consequence
Every value-based check fails — output files are missing or contain numbers unrelated to the real data, so the grader scores 0 even though the script runs cleanly and the answer looks well-formatted.
id c58c0336fea2 · mined from da-code dacode-plot-bar-007@s12
raw text (what the judge reads)
### Fabricating input data or config instead of reading the provided files
- **Applies when**: `task` -- the task references a supplied dataset and/or a specification file (e.g., a config describing plot/output formatting) and names the required output artifacts.
- **Pattern**: The script never opens the provided data/spec files; it synthesizes random or hard-coded data, invents its own spec values, and even writes the spec file itself, then reports numbers derived entirely from the fabricated inputs. Required output artifacts named by the task are also skipped.
- **Detection procedure**:
  1. List every input file the task mentions (data files, config/spec files) and every output artifact expected.
  2. Search the scripts for reads of those inputs (`read_csv`, `open(...)`, `yaml.safe_load`, etc.); flag any use of `np.random`, hard-coded sample rows, or code that *writes* a file that was supposed to be *read*.
  3. Check that every named output artifact is produced at the required path and format; flag missing ones.
  4. Compare reported counts/categories against what the real data could plausibly contain (row counts, category labels stated in the README/spec).
- **Discriminator**: Genuine violation = the analysis pipeline's numbers come from generated/assumed data, or the spec is authored rather than parsed. Fine = the script reads the real inputs and only uses synthetic data for a temporary smoke test, or fills a documented default for a field genuinely absent from the spec.
- **Consequence**: Every value-based check fails — output files are missing or contain numbers unrelated to the real data, so the grader scores 0 even though the script runs cleanly and the answer looks well-formatted.
647Ignoring the stated output contract (deliverable file and ordering)taskda-code
Applies when
task -- the prompt specifies exactly how the result must be delivered (a named result file, a JSON schema, and/or an explicit sort order for each list) and the scripts end by writing/printing results.
Pattern
The attempt computes plausible numbers but persists them somewhere other than the required artifact name/path (e.g. an ad-hoc answer.txt or only stdout), and/or builds each requested list with a helper whose natural order contradicts the demanded ordering (e.g. an ascending "n-smallest" call when the task says sort every list from highest to lowest), without any explicit re-sorting step.
Detection procedure
  1. From the task statement, list the hard output requirements: exact file name/location, key names, value types, and the required ordering/rounding/units of each list.
  2. In the scripts, find the final write step and compare the target filename/path and the emitted structure against that list.
  3. Trace how each reported list is constructed and note the order it produces; check whether an explicit sort matching the stated direction is applied to both lists.
  4. Inspect the submitted answer and confirm each list's order and the schema match the requirement literally.
Discriminator
A real violation is a mismatch in the artifact the grader reads or an order/format the task explicitly pinned down; it is not a violation if the required file is also written (extra debug files are fine) or if the task left ordering unspecified and any consistent order is acceptable.
Consequence
The grader reports the expected result file as WRONG/MISSING, or marks the lists incorrect despite the underlying values being right, scoring 0.
id b5ac7922343e · mined from da-code dacode-di-text-003@s12
raw text (what the judge reads)
### Ignoring the stated output contract (deliverable file and ordering)
- **Applies when**: `task` -- the prompt specifies exactly how the result must be delivered (a named result file, a JSON schema, and/or an explicit sort order for each list) and the scripts end by writing/printing results.
- **Pattern**: The attempt computes plausible numbers but persists them somewhere other than the required artifact name/path (e.g. an ad-hoc `answer.txt` or only stdout), and/or builds each requested list with a helper whose natural order contradicts the demanded ordering (e.g. an ascending "n-smallest" call when the task says sort every list from highest to lowest), without any explicit re-sorting step.
- **Detection procedure**:
  1. From the task statement, list the hard output requirements: exact file name/location, key names, value types, and the required ordering/rounding/units of each list.
  2. In the scripts, find the final write step and compare the target filename/path and the emitted structure against that list.
  3. Trace how each reported list is constructed and note the order it produces; check whether an explicit sort matching the stated direction is applied to *both* lists.
  4. Inspect the submitted answer and confirm each list's order and the schema match the requirement literally.
- **Discriminator**: A real violation is a mismatch in the artifact the grader reads or an order/format the task explicitly pinned down; it is *not* a violation if the required file is also written (extra debug files are fine) or if the task left ordering unspecified and any consistent order is acceptable.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING, or marks the lists incorrect despite the underlying values being right, scoring 0.
648Answer string does not reproduce the requested tag template literally (quoting/delimiters dropped)taskinfiagent-dabench
Applies when
task -- The task specifies an exact answer template with tagged fields (e.g. @field[<value>]) and shows the value in a particular literal form (quoted string, unit, bracket style), and the script hard-codes/prints the final answer.
Pattern
The attempt computes the right values but emits them in a re-formatted way — dropping the quotation marks shown in the template, changing separators, adding extra prose/units, renaming tags, or writing the answer only to a side file instead of the requested output — so a literal/regex grader matches nothing even though the analysis is correct.
Detection procedure
  1. Copy the answer template verbatim from the task statement, including every quote, bracket, comma, and spacing, and note which parts are placeholders vs. literal characters.
  2. Find the line(s) in the script that construct/print/write the final answer and expand the f-string mentally into the concrete output string.
  3. Character-by-character diff that expanded string against the template with placeholders filled in; check also that it is delivered where the task asks (final response, not just a file).
  4. Flag if any literal character of the template (notably quotes around string values, tag names, or field order) is missing or altered.
Discriminator
A real violation is a mismatch in the literal scaffolding of the template (missing quotes, altered tag names, different separators, answer only in a file). A look-alike that is fine is a difference only inside the placeholder region that the task leaves free (e.g. Yes vs yes when the task shows no casing constraint, or harmless surrounding whitespace/newlines).
Consequence
The grader reports every field as WRONG/MISSING (0/N checks) even though the submitted values are substantively identical to the ground truth.
id e76a2ce16847 · mined from infiagent-dabench dabench-550@s12
raw text (what the judge reads)
### Answer string does not reproduce the requested tag template literally (quoting/delimiters dropped)
- **Applies when**: `task` -- The task specifies an exact answer template with tagged fields (e.g. `@field[<value>]`) and shows the value in a particular literal form (quoted string, unit, bracket style), and the script hard-codes/prints the final answer.
- **Pattern**: The attempt computes the right values but emits them in a re-formatted way — dropping the quotation marks shown in the template, changing separators, adding extra prose/units, renaming tags, or writing the answer only to a side file instead of the requested output — so a literal/regex grader matches nothing even though the analysis is correct.
- **Detection procedure**:
  1. Copy the answer template verbatim from the task statement, including every quote, bracket, comma, and spacing, and note which parts are placeholders vs. literal characters.
  2. Find the line(s) in the script that construct/print/write the final answer and expand the f-string mentally into the concrete output string.
  3. Character-by-character diff that expanded string against the template with placeholders filled in; check also that it is delivered where the task asks (final response, not just a file).
  4. Flag if any literal character of the template (notably quotes around string values, tag names, or field order) is missing or altered.
- **Discriminator**: A real violation is a mismatch in the literal scaffolding of the template (missing quotes, altered tag names, different separators, answer only in a file). A look-alike that is fine is a difference only inside the placeholder region that the task leaves free (e.g. `Yes` vs `yes` when the task shows no casing constraint, or harmless surrounding whitespace/newlines).
- **Consequence**: The grader reports every field as WRONG/MISSING (0/N checks) even though the submitted values are substantively identical to the ground truth.
649Entity substitution: computing the requested quantity from unrelated fieldstaskda-code
Applies when
task -- the task names specific entities/measures (e.g., a grouping key, a set of stages, a ranking metric) and the scripts must locate those in the provided data before aggregating.
Pattern
The agent cannot find the named columns, silently substitutes semantically different fields (or a different dataset entirely), and still reports the output as if it answered the question — often also skipping required output artifacts (extra data/config files the task or config implies).
Detection procedure
1. List every entity the task requires: the grouping dimension, the ranking measure, the per-stage quantity, its unit, and every file to be written (including any config-driven or auxiliary outputs). 2. Read the scripts/answer for the actual columns and files used, and check each required entity maps to a real column of the right meaning and unit. 3. Flag if any required entity is replaced by a proxy of a different kind (counts instead of monetary totals, countries instead of cities, category labels instead of time-duration stages) or if a required output file is never produced. 4. Check the answer's reported numbers are in units consistent with the request (e.g., days), not raw record counts.
Discriminator
A legitimate attempt derives the requested quantity from columns that genuinely encode it (possibly renamed or computed, e.g., differencing timestamps to get durations) and states the mapping; a violation swaps in a different concept or dataset and reports it without noting that the requested quantity is unavailable.
Consequence
All value/array/plot checks mismatch — wrong axis labels, wrong group ordering, wrong magnitudes (counts vs. days), and missing expected artifacts, so every grader check fails.
id 435ce3aa8eff · mined from da-code dacode-plot-scatter-002@s12
raw text (what the judge reads)
### Entity substitution: computing the requested quantity from unrelated fields
- **Applies when**: `task` -- the task names specific entities/measures (e.g., a grouping key, a set of stages, a ranking metric) and the scripts must locate those in the provided data before aggregating.
- **Pattern**: The agent cannot find the named columns, silently substitutes semantically different fields (or a different dataset entirely), and still reports the output as if it answered the question — often also skipping required output artifacts (extra data/config files the task or config implies).
- **Detection procedure**: 1. List every entity the task requires: the grouping dimension, the ranking measure, the per-stage quantity, its unit, and every file to be written (including any config-driven or auxiliary outputs). 2. Read the scripts/answer for the actual columns and files used, and check each required entity maps to a real column of the right meaning and unit. 3. Flag if any required entity is replaced by a proxy of a different kind (counts instead of monetary totals, countries instead of cities, category labels instead of time-duration stages) or if a required output file is never produced. 4. Check the answer's reported numbers are in units consistent with the request (e.g., days), not raw record counts.
- **Discriminator**: A legitimate attempt derives the requested quantity from columns that genuinely encode it (possibly renamed or computed, e.g., differencing timestamps to get durations) and states the mapping; a violation swaps in a different concept or dataset and reports it without noting that the requested quantity is unavailable.
- **Consequence**: All value/array/plot checks mismatch — wrong axis labels, wrong group ordering, wrong magnitudes (counts vs. days), and missing expected artifacts, so every grader check fails.
650Output artifact name/location/schema doesn't match the specified deliverabletaskda-code
Applies when
task -- the task names an exact output file (and/or supplies a template) that the script must write for grading.
Pattern
The script writes results to a filename, directory, or column/index layout of the agent's own choosing (corrected spelling, different folder, renamed or reordered columns, added/dropped index) instead of byte-for-byte matching the requested deliverable, and the answer asserts "format matches the template" without ever loading and comparing the template.
Detection procedure
  1. From the task statement, copy the required output path/filename verbatim, plus any stated template or format reference.
  2. Search the scripts for the write call (to_csv/to_json/save) and compare the literal string, character by character, to the required name and expected output directory.
  3. Check whether the script reads the provided template (or expected header) and programmatically asserts that column names, order, row order/index, and dtypes match; absence of such a check is a red flag.
  4. Check the answer text: does it claim format compliance while naming a different file than the task requested?
Discriminator
A real violation is any mismatch in the artifact the grader will look for — different spelling, extension, directory, or a header/index layout never verified against the template. A look-alike that is fine: the script writes the exact required path (possibly also writing extra copies elsewhere) and explicitly validates its columns/shape against the template.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even if the underlying numbers were computed correctly.
id 29f1d0bf5c5c · mined from da-code dacode-dm-csv-043@s12
raw text (what the judge reads)
### Output artifact name/location/schema doesn't match the specified deliverable
- **Applies when**: `task` -- the task names an exact output file (and/or supplies a template) that the script must write for grading.
- **Pattern**: The script writes results to a filename, directory, or column/index layout of the agent's own choosing (corrected spelling, different folder, renamed or reordered columns, added/dropped index) instead of byte-for-byte matching the requested deliverable, and the answer asserts "format matches the template" without ever loading and comparing the template.
- **Detection procedure**:
  1. From the task statement, copy the required output path/filename verbatim, plus any stated template or format reference.
  2. Search the scripts for the write call (`to_csv`/`to_json`/`save`) and compare the literal string, character by character, to the required name and expected output directory.
  3. Check whether the script reads the provided template (or expected header) and programmatically asserts that column names, order, row order/index, and dtypes match; absence of such a check is a red flag.
  4. Check the answer text: does it claim format compliance while naming a different file than the task requested?
- **Discriminator**: A real violation is any mismatch in the artifact the grader will look for — different spelling, extension, directory, or a header/index layout never verified against the template. A look-alike that is fine: the script writes the exact required path (possibly also writing extra copies elsewhere) and explicitly validates its columns/shape against the template.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even if the underlying numbers were computed correctly.
651Preprocessing grouping mistaken for analysis grouping (wrong number/scope of reported statistics)taskda-code
Applies when
task -- The task says to clean or filter data within subgroups (per group thresholds/ranges) and then run a single statistical test or metric across a different set of categories on the cleaned data.
Pattern
The agent reuses the preprocessing grouping as the unit of analysis, running the test separately inside each cleaning subgroup and emitting one statistic per subgroup, instead of pooling the cleaned rows and computing the requested test over the categories actually named in the task (or vice-versa). The output length/shape therefore doesn't match what was asked, even if each individual computation is arithmetically fine.
Detection procedure
  1. Read the task and write down two things separately: (a) the grouping used only for filtering/cleaning, and (b) the grouping that defines the comparison in the requested test/metric, plus how many result values that implies.
  2. In the scripts, locate the groupby/loop that drives the test call and check whether it iterates over (a) or (b); confirm the cleaned frame is re-concatenated before the test if the test is meant to be global.
  3. Count the elements in the reported answer arrays and compare against the count implied in step 1; check the answer's key names/format match the requested template exactly.
  4. Sanity-check that the total row count after cleaning is reported/derivable and that each test category is represented in the pooled cleaned data.
Discriminator
A real violation is when the number of reported statistics equals the number of cleaning subgroups (or otherwise differs from the number implied by the requested comparison) with no statement in the task asking for per-subgroup results. It is fine if the task explicitly asks for one result per subgroup, or if the two groupings genuinely coincide and the counts still match the task's wording.
Consequence
The graded file mismatches the expected structure and values (e.g., a list of several p-values where one pooled p-value was expected), so exact-match/structural checks fail even though the conclusion wording may look plausible.
id 3184fb3c10e0 · mined from da-code dacode-data-sa-061@s12
raw text (what the judge reads)
### Preprocessing grouping mistaken for analysis grouping (wrong number/scope of reported statistics)
- **Applies when**: `task` -- The task says to clean or filter data *within* subgroups (per group thresholds/ranges) and then run a single statistical test or metric across a *different* set of categories on the cleaned data.
- **Pattern**: The agent reuses the preprocessing grouping as the unit of analysis, running the test separately inside each cleaning subgroup and emitting one statistic per subgroup, instead of pooling the cleaned rows and computing the requested test over the categories actually named in the task (or vice-versa). The output length/shape therefore doesn't match what was asked, even if each individual computation is arithmetically fine.
- **Detection procedure**:
  1. Read the task and write down two things separately: (a) the grouping used only for filtering/cleaning, and (b) the grouping that defines the comparison in the requested test/metric, plus how many result values that implies.
  2. In the scripts, locate the `groupby`/loop that drives the test call and check whether it iterates over (a) or (b); confirm the cleaned frame is re-concatenated before the test if the test is meant to be global.
  3. Count the elements in the reported answer arrays and compare against the count implied in step 1; check the answer's key names/format match the requested template exactly.
  4. Sanity-check that the total row count after cleaning is reported/derivable and that each test category is represented in the pooled cleaned data.
- **Discriminator**: A real violation is when the number of reported statistics equals the number of cleaning subgroups (or otherwise differs from the number implied by the requested comparison) with no statement in the task asking for per-subgroup results. It is fine if the task explicitly asks for one result per subgroup, or if the two groupings genuinely coincide and the counts still match the task's wording.
- **Consequence**: The graded file mismatches the expected structure and values (e.g., a list of several p-values where one pooled p-value was expected), so exact-match/structural checks fail even though the conclusion wording may look plausible.
652Model selected and validated on the training data it was fit on (in-sample scoring, no held-out split)taskda-code
Applies when
task -- The task asks for predictions on an unlabeled test set, and the scripts train several candidate models and pick/report a "best" one.
Pattern
The attempt computes its evaluation metrics (R²/MAE/RMSE) by predicting on the same rows used for fit, then selects the model with the best in-sample score and reports that score as the expected accuracy. Because in-sample fit rewards memorization, a high-capacity model (deep forest/boosting) is chosen over models that would generalize better, and no honest estimate of test error, no hyperparameter tuning, and no comparison against a trivial baseline (e.g., predicting the target mean) ever happens.
Detection procedure
  1. Read the task to confirm the deliverable is out-of-sample predictions, so generalization — not fit — is what matters.
  2. In the scripts, check whether any train/validation split, cross-validation, or out-of-fold prediction is created; look for the arrays passed to fit and to predict/score being identical.
  3. Check whether model/feature selection decisions (max(models, key=r2), chosen feature list, chosen depth) are driven by those same in-sample numbers, and whether a baseline is scored for reference.
  4. Compare the reported metric in the answer to the plausible signal in the data: if the answer claims a large R² from generic features with no validation number to back it, treat the claim as unverified.
  5. Also verify predictions were sanity-checked beyond mean/range (e.g., variance and distribution vs. training target) rather than just clipped.
Discriminator
A real violation is when no held-out or cross-validated estimate exists anywhere, so selection and the reported quality both come from resubstitution. It is fine if the script reports training fit for diagnostics but selects the final model using CV/validation scores, or if it uses a single fixed model with no selection and reports metrics as training-only without claiming generalization.
Consequence
The reported R² is wildly optimistic while the saved predictions are over-smoothed/overfit noise; the graded prediction file fails the accuracy/correlation threshold (often no better than predicting the mean), and the answer's stated performance cannot be reproduced on the held-out labels.
id 98d6b499dd02 · mined from da-code dacode-ml-regression-004@s12
raw text (what the judge reads)
### Model selected and validated on the training data it was fit on (in-sample scoring, no held-out split)
- **Applies when**: `task` -- The task asks for predictions on an unlabeled test set, and the scripts train several candidate models and pick/report a "best" one.
- **Pattern**: The attempt computes its evaluation metrics (R²/MAE/RMSE) by predicting on the same rows used for `fit`, then selects the model with the best in-sample score and reports that score as the expected accuracy. Because in-sample fit rewards memorization, a high-capacity model (deep forest/boosting) is chosen over models that would generalize better, and no honest estimate of test error, no hyperparameter tuning, and no comparison against a trivial baseline (e.g., predicting the target mean) ever happens.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is out-of-sample predictions, so generalization — not fit — is what matters.
  2. In the scripts, check whether any train/validation split, cross-validation, or out-of-fold prediction is created; look for the arrays passed to `fit` and to `predict`/`score` being identical.
  3. Check whether model/feature selection decisions (`max(models, key=r2)`, chosen feature list, chosen depth) are driven by those same in-sample numbers, and whether a baseline is scored for reference.
  4. Compare the reported metric in the answer to the plausible signal in the data: if the answer claims a large R² from generic features with no validation number to back it, treat the claim as unverified.
  5. Also verify predictions were sanity-checked beyond mean/range (e.g., variance and distribution vs. training target) rather than just clipped.
- **Discriminator**: A real violation is when *no* held-out or cross-validated estimate exists anywhere, so selection and the reported quality both come from resubstitution. It is fine if the script reports training fit for diagnostics but selects the final model using CV/validation scores, or if it uses a single fixed model with no selection and reports metrics as training-only without claiming generalization.
- **Consequence**: The reported R² is wildly optimistic while the saved predictions are over-smoothed/overfit noise; the graded prediction file fails the accuracy/correlation threshold (often no better than predicting the mean), and the answer's stated performance cannot be reproduced on the held-out labels.
653Unvalidated automatic model-selection producing a degenerate partitiontaskda-code
Applies when
task -- the task asks for an unsupervised grouping (or any hyperparameter chosen automatically) and the script picks the setting by taking the argmax/argmin of a single internal score over a range of candidates.
Pattern
The attempt trusts the single best internal score, and then reports/saves a partition whose group sizes are wildly degenerate (groups of size 1–3 next to groups of 80+) without cross-checking against other diagnostics (elbow/inertia, alternative metric, second-best candidate, outlier handling, or domain expectation of a small number of interpretable groups). The score is maximized by isolating outliers, not by finding structure.
Detection procedure
  1. Read the task for any hint about the intended granularity/interpretability of the groups (e.g., they must support a decision or ranking over the population).
  2. In the script, check whether the chosen setting comes from a bare argmax/argmin of one metric with no tie-breaking, no stability check, and no rule rejecting near-empty groups; also check whether outliers/extreme values were inspected or capped before fitting.
  3. In the reported output, read the group-size distribution: flag if any group holds a single item or a negligible fraction (< ~2–3%) of rows while others are huge, or if the number of groups was not corroborated by a second diagnostic.
  4. Confirm the saved file's columns/indexing and row count match the requested format and the full input population; flag if the answer never states such a check.
Discriminator
A real violation is when the winning configuration's groups are size-degenerate or contradicted by other diagnostics and the agent never reconsiders; it is fine if tiny groups are genuinely justified (documented outlier cluster reported as such, or the choice is confirmed by a second metric/elbow and shown to be stable across seeds).
Consequence
The saved grouping does not reproduce the expected partition (wrong cluster count and near-singleton labels), so file-level comparison of cluster assignments/counts fails and the deliverable is scored wrong.
id 2fe26fe60a92 · mined from da-code dacode-ml-cluster-013@s12
raw text (what the judge reads)
### Unvalidated automatic model-selection producing a degenerate partition
- **Applies when**: `task` -- the task asks for an unsupervised grouping (or any hyperparameter chosen automatically) and the script picks the setting by taking the argmax/argmin of a single internal score over a range of candidates.
- **Pattern**: The attempt trusts the single best internal score, and then reports/saves a partition whose group sizes are wildly degenerate (groups of size 1–3 next to groups of 80+) without cross-checking against other diagnostics (elbow/inertia, alternative metric, second-best candidate, outlier handling, or domain expectation of a small number of interpretable groups). The score is maximized by isolating outliers, not by finding structure.
- **Detection procedure**:
  1. Read the task for any hint about the intended granularity/interpretability of the groups (e.g., they must support a decision or ranking over the population).
  2. In the script, check whether the chosen setting comes from a bare `argmax`/`argmin` of one metric with no tie-breaking, no stability check, and no rule rejecting near-empty groups; also check whether outliers/extreme values were inspected or capped before fitting.
  3. In the reported output, read the group-size distribution: flag if any group holds a single item or a negligible fraction (< ~2–3%) of rows while others are huge, or if the number of groups was not corroborated by a second diagnostic.
  4. Confirm the saved file's columns/indexing and row count match the requested format and the full input population; flag if the answer never states such a check.
- **Discriminator**: A real violation is when the winning configuration's groups are size-degenerate or contradicted by other diagnostics and the agent never reconsiders; it is fine if tiny groups are genuinely justified (documented outlier cluster reported as such, or the choice is confirmed by a second metric/elbow and shown to be stable across seeds).
- **Consequence**: The saved grouping does not reproduce the expected partition (wrong cluster count and near-singleton labels), so file-level comparison of cluster assignments/counts fails and the deliverable is scored wrong.
654Truncating an identifier value to a coarser granularity than the data (over-literal reading of a format template)taskinfiagent-dabench
Applies when
task -- the deliverable includes a key/label (date, ID, category) that is selected from the data, and the prompt gives a schematic format template (e.g., a date pattern) whose granularity is coarser than, or ambiguous relative to, the granularity of the underlying records.
Pattern
The attempt locates the correct row, but then reformats/truncates the identifier to match the literal template (dropping day, time, or sub-code), so the reported key no longer uniquely identifies the selected record — even though every downstream number was computed at full granularity.
Detection procedure
  1. From the task, note the granularity at which the selection and the dependent calculation must be performed (here: the unit at which "previous record" and the change are defined).
  2. In the scripts, find where the identifier is emitted and check for any truncation/strftime/slicing/astype(str)[:n] or grouping that reduces its precision relative to the row actually used in the computation.
  3. Compare the emitted identifier with the record used for the dependent statistic: can the reported key be mapped back to exactly one row of the source data?
  4. Check whether the reported key is internally consistent — if the dependent value was computed between adjacent fine-grained records, a coarse key is a mismatch.
Discriminator
A real violation is when the source data and the computation operate at a finer granularity than the reported key, so the key becomes ambiguous or non-invertible. It is not a violation if the data itself is genuinely stored at that coarser granularity (one record per period), or if the task explicitly asks for an aggregation over that period — then the template and data agree.
Consequence
The numeric part may match, but the identifier field fails exact-string comparison, giving a partial score (e.g., 1/2 checks) and an overall incorrect verdict.
id 934d78f42a12 · mined from infiagent-dabench dabench-572@s12
raw text (what the judge reads)
### Truncating an identifier value to a coarser granularity than the data (over-literal reading of a format template)
- **Applies when**: `task` -- the deliverable includes a key/label (date, ID, category) that is selected from the data, and the prompt gives a schematic format template (e.g., a date pattern) whose granularity is coarser than, or ambiguous relative to, the granularity of the underlying records.
- **Pattern**: The attempt locates the correct row, but then reformats/truncates the identifier to match the literal template (dropping day, time, or sub-code), so the reported key no longer uniquely identifies the selected record — even though every downstream number was computed at full granularity.
- **Detection procedure**:
  1. From the task, note the granularity at which the selection and the dependent calculation must be performed (here: the unit at which "previous record" and the change are defined).
  2. In the scripts, find where the identifier is emitted and check for any truncation/`strftime`/slicing/`astype(str)[:n]` or grouping that reduces its precision relative to the row actually used in the computation.
  3. Compare the emitted identifier with the record used for the dependent statistic: can the reported key be mapped back to exactly one row of the source data?
  4. Check whether the reported key is internally consistent — if the dependent value was computed between adjacent fine-grained records, a coarse key is a mismatch.
- **Discriminator**: A real violation is when the source data and the computation operate at a finer granularity than the reported key, so the key becomes ambiguous or non-invertible. It is *not* a violation if the data itself is genuinely stored at that coarser granularity (one record per period), or if the task explicitly asks for an aggregation over that period — then the template and data agree.
- **Consequence**: The numeric part may match, but the identifier field fails exact-string comparison, giving a partial score (e.g., 1/2 checks) and an overall incorrect verdict.
655Fabricating input data instead of locating and loading the provided datasettaskda-code
Applies when
task -- the task asks for a statistic computed from a described dataset, and the script generates data programmatically (random draws, np.linspace, hardcoded values) rather than reading a file.
Pattern
The agent claims no data file exists, synthesizes "realistic" data with a fixed seed, fits the requested model on the synthetic data, and reports the resulting number as the answer — so the value is unrelated to the real data and cannot match ground truth.
Detection procedure
  1. Read the task/README and note that the requested quantity is defined relative to a specific real dataset.
  2. Scan the script for any read of an external source (read_csv, read_excel, SQL, API, package-bundled data); if the only inputs come from RNG or literal constructions, flag it.
  3. Check whether the agent actually searched the working directory / environment (e.g., listing files, globbing for data) before concluding data was unavailable.
  4. Verify the reported number is traceable to real observations (row counts, date ranges, summary stats printed from loaded data) rather than to a seed.
Discriminator
Legitimate synthetic data is fine only when the task explicitly asks for simulation (e.g., Monte Carlo) parameterized by real loaded data or by task-given constants; a violation is when the substantive input series themselves are invented as stand-ins for unfound real data.
Consequence
The saved statistic is an arbitrary function of the random seed; the grader's value comparison in the output file fails, marking the result wrong even though the file format is correct.
id ffabfec831c4 · mined from da-code dacode-data-sa-043@s12
raw text (what the judge reads)
### Fabricating input data instead of locating and loading the provided dataset
- **Applies when**: `task` -- the task asks for a statistic computed from a described dataset, and the script generates data programmatically (random draws, `np.linspace`, hardcoded values) rather than reading a file.
- **Pattern**: The agent claims no data file exists, synthesizes "realistic" data with a fixed seed, fits the requested model on the synthetic data, and reports the resulting number as the answer — so the value is unrelated to the real data and cannot match ground truth.
- **Detection procedure**:
  1. Read the task/README and note that the requested quantity is defined relative to a specific real dataset.
  2. Scan the script for any read of an external source (`read_csv`, `read_excel`, SQL, API, package-bundled data); if the only inputs come from RNG or literal constructions, flag it.
  3. Check whether the agent actually searched the working directory / environment (e.g., listing files, globbing for data) before concluding data was unavailable.
  4. Verify the reported number is traceable to real observations (row counts, date ranges, summary stats printed from loaded data) rather than to a seed.
- **Discriminator**: Legitimate synthetic data is fine only when the task explicitly asks for simulation (e.g., Monte Carlo) *parameterized by* real loaded data or by task-given constants; a violation is when the substantive input series themselves are invented as stand-ins for unfound real data.
- **Consequence**: The saved statistic is an arbitrary function of the random seed; the grader's value comparison in the output file fails, marking the result wrong even though the file format is correct.
656Model selected on training-set (in-sample) scores with no held-out validation or cost-aware thresholdtaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test file and the scripts train several candidate models and pick "the best one", especially with class imbalance and an asymmetric error cost stated in the brief.
Pattern
The attempt fits each model on the full labeled data and reports a fit-quality metric (e.g., ROC-AUC ≈ 0.999) computed on those same rows, then selects the model and uses the default 0.5 decision threshold without any hold-out/cross-validation estimate or tuning of the decision rule toward the metric the task actually cares about (e.g., recall/false-negative cost).
Detection procedure
1. Read the task for the evaluation intent (which error type is costly, whether a scoring metric or class labels are required). 2. In the scripts, check whether the data used to compute each reported score is disjoint from the data used to fit — look for train_test_split, cross_val_score, or a validation set; flag if predict/predict_proba is called on the training matrix. 3. Check whether the threshold/class-weight choice was validated against the cost-relevant metric on held-out data, rather than left at defaults. 4. Compare the reported score to the answer's predicted positive rate: a near-perfect score paired with a predicted positive rate that merely mirrors the training prevalence is a red flag of in-sample optimism.
Discriminator
A real violation is when every reported number comes from rows the model saw during fitting (or the split is created after fitting/resampling/imputation on the full data); it is fine if scores come from a genuine hold-out or CV fold and the near-perfect value is reproduced out-of-sample, or if the threshold was chosen on validation data with the task's cost in mind.
Consequence
The chosen model/threshold is optimized for memorization rather than generalization, so the submitted label column misses a large share of the rare positive class; the grader's out-of-sample check (recall/F1/agreement with expected labels) fails despite the reported ~1.0 score.
id 4630ae857d38 · mined from da-code dacode-ml-binary-013@s12
raw text (what the judge reads)
### Model selected on training-set (in-sample) scores with no held-out validation or cost-aware threshold
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test file and the scripts train several candidate models and pick "the best one", especially with class imbalance and an asymmetric error cost stated in the brief.
- **Pattern**: The attempt fits each model on the full labeled data and reports a fit-quality metric (e.g., ROC-AUC ≈ 0.999) computed on those same rows, then selects the model and uses the default 0.5 decision threshold without any hold-out/cross-validation estimate or tuning of the decision rule toward the metric the task actually cares about (e.g., recall/false-negative cost).
- **Detection procedure**: 1. Read the task for the evaluation intent (which error type is costly, whether a scoring metric or class labels are required). 2. In the scripts, check whether the data used to compute each reported score is disjoint from the data used to `fit` — look for `train_test_split`, `cross_val_score`, or a validation set; flag if `predict/predict_proba` is called on the training matrix. 3. Check whether the threshold/class-weight choice was validated against the cost-relevant metric on held-out data, rather than left at defaults. 4. Compare the reported score to the answer's predicted positive rate: a near-perfect score paired with a predicted positive rate that merely mirrors the training prevalence is a red flag of in-sample optimism.
- **Discriminator**: A real violation is when *every* reported number comes from rows the model saw during fitting (or the split is created after fitting/resampling/imputation on the full data); it is fine if scores come from a genuine hold-out or CV fold and the near-perfect value is reproduced out-of-sample, or if the threshold was chosen on validation data with the task's cost in mind.
- **Consequence**: The chosen model/threshold is optimized for memorization rather than generalization, so the submitted label column misses a large share of the rare positive class; the grader's out-of-sample check (recall/F1/agreement with expected labels) fails despite the reported ~1.0 score.
657Answer emitted in a formatting variant the grader cannot parse (decorated list literal instead of the requested plain list)taskinfiagent-dabench
Applies when
task -- the task prescribes a specific answer template (e.g. @name[list_of_strings]) and the agent produces a list of labels/values as the final deliverable.
Pattern
The scripts compute the correct values, but the final answer is written by pasting a Python/JSON-style repr (quotes, brackets, spacing, trailing punctuation, extra keys) rather than the literal token format the task asked for; no step in the scripts constructs or prints the final answer string in the required shape for direct copying.
Detection procedure
  1. Read the task's "Answer format" line and note exactly which delimiters, separators, quoting and casing are specified, and whether elements are meant to be bare labels or quoted strings.
  2. Scan the scripts for the point where the final answer string is produced; check whether it prints the answer already formatted (e.g. print(f"@name[{', '.join(items)}]")) or only prints a DataFrame/list repr.
  3. Compare the submitted answer character-by-character with the template: extra "/', [] vs no brackets, ; vs ,, unit suffixes, rounding, or ordering differences.
  4. Also confirm the answer contains only the requested items (no explanatory text, index numbers, or values attached to labels).
Discriminator
A real violation is a deliverable whose surface form deviates from the stated template even though the underlying values are right; a look-alike that is fine is an answer whose surface form matches the template while the scripts internally used a different representation (Python list, DataFrame) purely for computation.
Consequence
The grader's exact/parsed comparison fails and the item scores 0 despite correct analysis, reported as "expected -> ...: WRONG/MISSING" with a submitted value that looks superficially identical to the expected one.
id 6e7005e538d8 · mined from infiagent-dabench dabench-254@s12
raw text (what the judge reads)
### Answer emitted in a formatting variant the grader cannot parse (decorated list literal instead of the requested plain list)
- **Applies when**: `task` -- the task prescribes a specific answer template (e.g. `@name[list_of_strings]`) and the agent produces a list of labels/values as the final deliverable.
- **Pattern**: The scripts compute the correct values, but the final answer is written by pasting a Python/JSON-style repr (quotes, brackets, spacing, trailing punctuation, extra keys) rather than the literal token format the task asked for; no step in the scripts constructs or prints the final answer string in the required shape for direct copying.
- **Detection procedure**:
  1. Read the task's "Answer format" line and note exactly which delimiters, separators, quoting and casing are specified, and whether elements are meant to be bare labels or quoted strings.
  2. Scan the scripts for the point where the final answer string is produced; check whether it prints the answer already formatted (e.g. `print(f"@name[{', '.join(items)}]")`) or only prints a DataFrame/`list` repr.
  3. Compare the submitted answer character-by-character with the template: extra `"`/`'`, `[]` vs no brackets, `;` vs `,`, unit suffixes, rounding, or ordering differences.
  4. Also confirm the answer contains only the requested items (no explanatory text, index numbers, or values attached to labels).
- **Discriminator**: A real violation is a deliverable whose surface form deviates from the stated template even though the underlying values are right; a look-alike that is fine is an answer whose surface form matches the template while the scripts internally used a different representation (Python list, DataFrame) purely for computation.
- **Consequence**: The grader's exact/parsed comparison fails and the item scores 0 despite correct analysis, reported as "expected -> ...: WRONG/MISSING" with a submitted value that looks superficially identical to the expected one.
658Arbitrary `argmax` label when the maximum is a tie (often a degenerate all-zero case)taskinfiagent-dabench
Applies when
task -- the task asks for the identity of the group/row/feature achieving an extreme (max/min) of a per-group statistic, and the script computes it with an idxmax/argmax/sort().head(1)-style call.
Pattern
the attempt reports the extreme value correctly but takes the label from whatever element the library happens to return first, without checking how many groups share that extreme — most damagingly when the statistic is identical (e.g., zero) for every group, so the "winner" is an artifact of internal ordering (hash/group order, unsorted input, file read order).
Detection procedure
  1. In the task, note that both an extreme value and its label are requested, and that the extreme value could plausibly be degenerate (0, empty, constant).
  2. In the script, check whether the count of groups attaining the extreme is computed/printed, and whether a deterministic tie-break (e.g., sort the group keys, or the ordering rule the task implies) is applied before selecting the label.
  3. If the reported extreme value is 0 or otherwise trivial, confirm the script establishes whether all groups share it; if so, the label must come from a stated/deterministic ordering, not from idxmax on an arbitrarily ordered object.
  4. Reject if no tie count and no explicit deterministic ordering exist (also reject if no script/output is available to verify either).
Discriminator
fine if the script prints the top-k statistic values and shows a unique maximum, or explicitly sorts group keys / applies the task's ordering rule before picking the label; violation if the label is emitted from an unordered grouping with no tie inspection, particularly when the extreme value is trivially uniform.
Consequence
the numeric part of the answer matches but the identifier does not, so the grader marks a partial failure (e.g., 1/2 checks passed) and the answer is non-reproducible across runs.
id 30bdbcc12151 · mined from infiagent-dabench dabench-760@s12
raw text (what the judge reads)
### Arbitrary `argmax` label when the maximum is a tie (often a degenerate all-zero case)
- **Applies when**: `task` -- the task asks for the identity of the group/row/feature achieving an extreme (max/min) of a per-group statistic, and the script computes it with an `idxmax`/`argmax`/`sort().head(1)`-style call.
- **Pattern**: the attempt reports the extreme value correctly but takes the label from whatever element the library happens to return first, without checking how many groups share that extreme — most damagingly when the statistic is identical (e.g., zero) for every group, so the "winner" is an artifact of internal ordering (hash/group order, unsorted input, file read order).
- **Detection procedure**:
  1. In the task, note that both an extreme *value* and its *label* are requested, and that the extreme value could plausibly be degenerate (0, empty, constant).
  2. In the script, check whether the count of groups attaining the extreme is computed/printed, and whether a deterministic tie-break (e.g., sort the group keys, or the ordering rule the task implies) is applied before selecting the label.
  3. If the reported extreme value is 0 or otherwise trivial, confirm the script establishes whether *all* groups share it; if so, the label must come from a stated/deterministic ordering, not from `idxmax` on an arbitrarily ordered object.
  4. Reject if no tie count and no explicit deterministic ordering exist (also reject if no script/output is available to verify either).
- **Discriminator**: fine if the script prints the top-k statistic values and shows a unique maximum, or explicitly sorts group keys / applies the task's ordering rule before picking the label; violation if the label is emitted from an unordered grouping with no tie inspection, particularly when the extreme value is trivially uniform.
- **Consequence**: the numeric part of the answer matches but the identifier does not, so the grader marks a partial failure (e.g., 1/2 checks passed) and the answer is non-reproducible across runs.
659Invented metric definition aggregated over only one of the entity's rolestaskda-code
Applies when
task -- the task asks to visualize/report a per-entity "performance"/summary quantity, and the entity can appear in more than one column (or the quantity's definition is fixed elsewhere, e.g. a config, prior step, or the axis/label text).
Pattern
The script makes up an arbitrary formula for the requested quantity (e.g. a weighted count of rows/flags) instead of deriving it from the definition implied by the task artifacts, and computes it by grouping on a single role column, silently dropping all rows where the entity appears in the other role; the answer then presents these numbers as the result without checking them against the labels/ordering already fixed in the provided config or against the full set of required output files.
Detection procedure
  1. Read the task and any provided config/spec for the intended definition of the plotted quantity (title, axis labels, units, ordering, and the full list of required output artifacts).
  2. In the script, locate the line computing that quantity: is the formula justified by the task/config, or invented by the agent with a self-supplied rationale in a comment/answer? Does the groupby key cover every column in which an entity can occur, or just one?
  3. Cross-check the produced values against the config-provided category order: if the config already fixes the labels/order, the computed ranking must reproduce that order; if it doesn't, the metric is wrong.
  4. Check the answer/script produces every artifact the task requires (chart plus any serialized data/values), not just the image.
Discriminator
A real violation is when the formula has no basis in the task/config, or when rows where the entity appears in the secondary role are excluded without justification, or when the config's fixed ordering contradicts the computed values. It is fine if the task genuinely restricts the analysis to one role, or if the definition is explicitly given and the script implements it faithfully with the ordering consistent.
Consequence
The bar heights (and any saved numeric/serialized outputs) differ from the reference values and ordering, so all value-based checks and the image comparison fail, and missing companion artifacts fail outright.
id be01627fbf37 · mined from da-code dacode-plot-bar-006@s12
raw text (what the judge reads)
### Invented metric definition aggregated over only one of the entity's roles
- **Applies when**: `task` -- the task asks to visualize/report a per-entity "performance"/summary quantity, and the entity can appear in more than one column (or the quantity's definition is fixed elsewhere, e.g. a config, prior step, or the axis/label text).
- **Pattern**: The script makes up an arbitrary formula for the requested quantity (e.g. a weighted count of rows/flags) instead of deriving it from the definition implied by the task artifacts, and computes it by grouping on a single role column, silently dropping all rows where the entity appears in the other role; the answer then presents these numbers as the result without checking them against the labels/ordering already fixed in the provided config or against the full set of required output files.
- **Detection procedure**:
  1. Read the task and any provided config/spec for the intended definition of the plotted quantity (title, axis labels, units, ordering, and the full list of required output artifacts).
  2. In the script, locate the line computing that quantity: is the formula justified by the task/config, or invented by the agent with a self-supplied rationale in a comment/answer? Does the groupby key cover every column in which an entity can occur, or just one?
  3. Cross-check the produced values against the config-provided category order: if the config already fixes the labels/order, the computed ranking must reproduce that order; if it doesn't, the metric is wrong.
  4. Check the answer/script produces every artifact the task requires (chart plus any serialized data/values), not just the image.
- **Discriminator**: A real violation is when the formula has no basis in the task/config, or when rows where the entity appears in the secondary role are excluded without justification, or when the config's fixed ordering contradicts the computed values. It is fine if the task genuinely restricts the analysis to one role, or if the definition is explicitly given and the script implements it faithfully with the ordering consistent.
- **Consequence**: The bar heights (and any saved numeric/serialized outputs) differ from the reference values and ordering, so all value-based checks and the image comparison fail, and missing companion artifacts fail outright.
660Threshold-based outlier rule not verified against the actual score distributiontaskinfiagent-dabench
Applies when
task -- The task prescribes a specific detection/filtering rule (e.g., a standardized-score cutoff, quantile bound, or fixed threshold) on one column, and the script computes the statistic and counts/drops the flagged rows.
Pattern
The attempt computes the deviation statistic with a definition or scope that differs from the one stated (e.g., wrong center/scale — median/MAD instead of mean/std, ddof mismatch, per-group/rolling/windowed statistics instead of whole-column, scaling applied after another transform or on a filtered/joined subset), or silently substitutes a different rule (IQR, percentile, |value| cutoff), then reports the resulting count without ever checking the maximum absolute score against the stated threshold. A non-zero count is accepted at face value.
Detection procedure
  1. Read the task and write down the exact rule: which column, which statistic (center, scale, denominator convention), which threshold, and applied to which rows (full column vs. subset/group).
  2. In the script, locate the line that computes the statistic and the line that applies the cutoff; confirm the center and scale come from the entire target column (after only the preprocessing the task requires, e.g., NaN handling) and that the comparison uses the stated threshold and both tails.
  3. Check that the script prints a diagnostic tying the count to the distribution: min/max of the score, the implied data-value cutoffs (mean ± k·std), and the fraction of rows flagged; verify the flagged fraction is plausible for the rule (a ±3σ rule on n rows can flag at most a few percent, and often zero).
  4. Compare the reported count with those diagnostics and with the required answer format; a count that cannot be reproduced from max|score| > threshold is a violation.
Discriminator
A real violation is when the reported count is inconsistent with the stated rule applied to the whole column (e.g., max|z| < 3 yet outliers reported, or the count matches an IQR/grouped/rolling computation). It is not a violation if the script uses the stated definition on the full column, prints the score range showing values beyond the threshold, and the count matches those rows — even if the count is large or zero.
Consequence
The reported integer count (and the resulting filtered dataframe size) disagrees with the ground-truth count derived from the prescribed rule, so the answer check fails outright even though the pipeline "ran successfully".
id e3744728dc9f · mined from infiagent-dabench dabench-361@s12
raw text (what the judge reads)
### Threshold-based outlier rule not verified against the actual score distribution
- **Applies when**: `task` -- The task prescribes a specific detection/filtering rule (e.g., a standardized-score cutoff, quantile bound, or fixed threshold) on one column, and the script computes the statistic and counts/drops the flagged rows.
- **Pattern**: The attempt computes the deviation statistic with a definition or scope that differs from the one stated (e.g., wrong center/scale — median/MAD instead of mean/std, `ddof` mismatch, per-group/rolling/windowed statistics instead of whole-column, scaling applied after another transform or on a filtered/joined subset), or silently substitutes a different rule (IQR, percentile, |value| cutoff), then reports the resulting count without ever checking the maximum absolute score against the stated threshold. A non-zero count is accepted at face value.
- **Detection procedure**:
  1. Read the task and write down the exact rule: which column, which statistic (center, scale, denominator convention), which threshold, and applied to which rows (full column vs. subset/group).
  2. In the script, locate the line that computes the statistic and the line that applies the cutoff; confirm the center and scale come from the *entire* target column (after only the preprocessing the task requires, e.g., NaN handling) and that the comparison uses the stated threshold and both tails.
  3. Check that the script prints a diagnostic tying the count to the distribution: min/max of the score, the implied data-value cutoffs (mean ± k·std), and the fraction of rows flagged; verify the flagged fraction is plausible for the rule (a ±3σ rule on n rows can flag at most a few percent, and often zero).
  4. Compare the reported count with those diagnostics and with the required answer format; a count that cannot be reproduced from `max|score| > threshold` is a violation.
- **Discriminator**: A real violation is when the reported count is inconsistent with the stated rule applied to the whole column (e.g., max|z| < 3 yet outliers reported, or the count matches an IQR/grouped/rolling computation). It is *not* a violation if the script uses the stated definition on the full column, prints the score range showing values beyond the threshold, and the count matches those rows — even if the count is large or zero.
- **Consequence**: The reported integer count (and the resulting filtered dataframe size) disagrees with the ground-truth count derived from the prescribed rule, so the answer check fails outright even though the pipeline "ran successfully".
661Required output artifact never written (answer only reported in chat)taskda-code
Applies when
task -- The task (or its README/spec) names a deliverable file and/or an exact response schema, and the agent must both compute values and persist them.
Pattern
The agent computes numbers ad hoc (often inline/interactively, with no saved script) and pastes a JSON-looking blob into its final message, but no code ever serializes that object to the named output file, and no script exists to reproduce it. Auxiliary spec details (exact key names, label strings taken verbatim from the provided mapping, rounding/decimal form) are also unverifiable because the transformation code isn't preserved.
Detection procedure
  1. Read the task/README and list every required artifact (file name, location) and every formatting constraint (keys, value types, label spelling, rounding, units).
  2. Search the submitted scripts for a write step producing that artifact (json.dump, to_csv, open(..., 'w')) with the exact filename; if no scripts were saved at all, the check fails immediately since nothing regenerates the artifact.
  3. Compare the keys/label strings/number formatting in the reported answer against the spec and against the mapping/vocabulary the task said to use, confirming the values came from the mapped labels rather than raw or self-invented category names.
  4. Confirm the artifact's content would equal the reported answer (same object literally dumped, not retyped by hand).
Discriminator
A real violation is the absence of any code path that creates the required file (or a mismatch between file content and reported answer). It is not a violation if the file is written correctly and the agent additionally echoes the answer in chat, nor if the filename differs only by an explicitly permitted alias/path the task allows.
Consequence
The grader looks for the expected result file, finds it missing (or inconsistent/misformatted), and scores 0 regardless of whether the underlying statistic was numerically right.
id 62fbe7468588 · mined from da-code dacode-di-text-004@s12
raw text (what the judge reads)
### Required output artifact never written (answer only reported in chat)
- **Applies when**: `task` -- The task (or its README/spec) names a deliverable file and/or an exact response schema, and the agent must both compute values and persist them.
- **Pattern**: The agent computes numbers ad hoc (often inline/interactively, with no saved script) and pastes a JSON-looking blob into its final message, but no code ever serializes that object to the named output file, and no script exists to reproduce it. Auxiliary spec details (exact key names, label strings taken verbatim from the provided mapping, rounding/decimal form) are also unverifiable because the transformation code isn't preserved.
- **Detection procedure**:
  1. Read the task/README and list every required artifact (file name, location) and every formatting constraint (keys, value types, label spelling, rounding, units).
  2. Search the submitted scripts for a write step producing that artifact (`json.dump`, `to_csv`, `open(..., 'w')`) with the exact filename; if no scripts were saved at all, the check fails immediately since nothing regenerates the artifact.
  3. Compare the keys/label strings/number formatting in the reported answer against the spec and against the mapping/vocabulary the task said to use, confirming the values came from the mapped labels rather than raw or self-invented category names.
  4. Confirm the artifact's content would equal the reported answer (same object literally dumped, not retyped by hand).
- **Discriminator**: A real violation is the absence of any code path that creates the required file (or a mismatch between file content and reported answer). It is *not* a violation if the file is written correctly and the agent additionally echoes the answer in chat, nor if the filename differs only by an explicitly permitted alias/path the task allows.
- **Consequence**: The grader looks for the expected result file, finds it missing (or inconsistent/misformatted), and scores 0 regardless of whether the underlying statistic was numerically right.
662Hand-crafted heuristic in place of a model fit to the provided labels, with no validationtaskda-code
Applies when
task -- The task supplies labeled training data and a separate file of rows to score, and the scripts produce predictions from manually chosen formulas/thresholds rather than a model estimated from the labels.
Pattern
The agent invents an ad-hoc score from a few loosely related columns, picks cut points by intuition, and maps them onto the label categories. It never reads the label column from the training split, never measures accuracy/F1 on any held-out labeled subset, and never compares its predicted class distribution to the observed label distribution. It then declares success based only on the fact that a file was written.
Detection procedure
  1. From the task/README, confirm the target column exists with values in the training file and note the row set to be scored.
  2. Scan the scripts for (a) any load of the target column for fitting, (b) any train/validation split with a computed score, (c) any check that output row count/IDs match the rows to be scored and that predicted labels use the exact label vocabulary.
  3. If fitting and validation are absent, check the reported class proportions against the label proportions in the training data and against the required output shape/columns.
  4. Treat "no reported validation metric of any kind" plus "thresholds chosen by hand" as failing, regardless of how plausible the feature story sounds.
Discriminator
A legitimate rule-based baseline is still fit or calibrated against the observed labels and reports a measured score on labeled data (and its output matches the required rows/format); a violation asserts correctness with zero contact with the labels and zero sanity check on counts, ID alignment, or category distribution.
Consequence
Predictions are essentially unrelated to the true target (distribution wildly off, e.g. a rare class dominating), and the grader marks the output file wrong even though it exists and parses; format/row mismatches may cause it to be scored as missing.
id 3ff417122d6c · mined from da-code dacode-ml-multi-003@s12
raw text (what the judge reads)
### Hand-crafted heuristic in place of a model fit to the provided labels, with no validation
- **Applies when**: `task` -- The task supplies labeled training data and a separate file of rows to score, and the scripts produce predictions from manually chosen formulas/thresholds rather than a model estimated from the labels.
- **Pattern**: The agent invents an ad-hoc score from a few loosely related columns, picks cut points by intuition, and maps them onto the label categories. It never reads the label column from the training split, never measures accuracy/F1 on any held-out labeled subset, and never compares its predicted class distribution to the observed label distribution. It then declares success based only on the fact that a file was written.
- **Detection procedure**:
  1. From the task/README, confirm the target column exists with values in the training file and note the row set to be scored.
  2. Scan the scripts for (a) any load of the target column for fitting, (b) any train/validation split with a computed score, (c) any check that output row count/IDs match the rows to be scored and that predicted labels use the exact label vocabulary.
  3. If fitting and validation are absent, check the reported class proportions against the label proportions in the training data and against the required output shape/columns.
  4. Treat "no reported validation metric of any kind" plus "thresholds chosen by hand" as failing, regardless of how plausible the feature story sounds.
- **Discriminator**: A legitimate rule-based baseline is still fit or calibrated against the observed labels and reports a measured score on labeled data (and its output matches the required rows/format); a violation asserts correctness with zero contact with the labels and zero sanity check on counts, ID alignment, or category distribution.
- **Consequence**: Predictions are essentially unrelated to the true target (distribution wildly off, e.g. a rare class dominating), and the grader marks the output file wrong even though it exists and parses; format/row mismatches may cause it to be scored as missing.
663Output not reconciled cell-by-cell with the provided format templatetaskda-code
Applies when
task -- The task supplies a template/example output file (or explicit format spec) and the agent must write results to a file matching it.
Pattern
The agent invents its own layout — row-label formatting, column names, index start value/count, ordering, rounding/precision — and asserts "format matches the template" without ever loading the template and comparing headers, index labels, shape and dtypes against it.
Detection procedure
1. Read the task for the named template/spec file and note which structural elements it fixes (header names, label format, row/column ordering, number of columns, decimal precision). 2. Search the scripts for any read of that template and any explicit comparison (e.g. asserting equal column lists / index labels / shape) — absence is the red flag. 3. Inspect how the agent constructed labels and columns (date string format, whether the time-offset column starts at 0 or 1, rounding choice) and check each against the template's actual values. 4. Check whether the answer's format claims are backed by a printed diff/assertion rather than prose.
Discriminator
A real violation is choosing structural conventions unilaterally (or self-declaring a match) with no programmatic check against the template; it is fine if the script loads the template and asserts/aligns columns, index and rounding, even if the layout was reconstructed rather than copied.
Consequence
The file-comparison check fails — values may be shifted or labels/headers mismatched — so the result file is scored WRONG/MISSING despite a plausible-sounding narrative.
id c279bbd51b4d · mined from da-code dacode-dm-csv-044@s12
raw text (what the judge reads)
### Output not reconciled cell-by-cell with the provided format template
- **Applies when**: `task` -- The task supplies a template/example output file (or explicit format spec) and the agent must write results to a file matching it.
- **Pattern**: The agent invents its own layout — row-label formatting, column names, index start value/count, ordering, rounding/precision — and asserts "format matches the template" without ever loading the template and comparing headers, index labels, shape and dtypes against it.
- **Detection procedure**: 1. Read the task for the named template/spec file and note which structural elements it fixes (header names, label format, row/column ordering, number of columns, decimal precision). 2. Search the scripts for any read of that template and any explicit comparison (e.g. asserting equal column lists / index labels / shape) — absence is the red flag. 3. Inspect how the agent constructed labels and columns (date string format, whether the time-offset column starts at 0 or 1, rounding choice) and check each against the template's actual values. 4. Check whether the answer's format claims are backed by a printed diff/assertion rather than prose.
- **Discriminator**: A real violation is choosing structural conventions unilaterally (or self-declaring a match) with no programmatic check against the template; it is fine if the script loads the template and asserts/aligns columns, index and rounding, even if the layout was reconstructed rather than copied.
- **Consequence**: The file-comparison check fails — values may be shifted or labels/headers mismatched — so the result file is scored WRONG/MISSING despite a plausible-sounding narrative.
664Submission artifact never validated against the provided templatetaskda-code
Applies when
task -- the task specifies an output file whose format is defined by a sample/template file (header names, key column, one row per test key, expected location).
Pattern
The scripts build predictions and write an output file from scratch (hand-constructed DataFrame, ad-hoc path) without ever loading the sample template or asserting that the written file matches it in row count, key coverage/order, column names, and value dtype/range; the agent then reports success on the basis of model CV scores alone.
Detection procedure
  1. Read the task/README for the required output filename, location, header, and the rule "one prediction per test key".
  2. Search the scripts for a read of the sample/template file and for post-write checks (e.g., re-reading the saved file, comparing len and the key set/order to the test/template file, checking header spelling and integer vs float output).
  3. Compare the write path in the script to the path the grader expects, and confirm the ID column comes from the test keys in template order.
  4. Inspect the reported answer: does it contain exactly as many rows as the test set, all keys, no duplicates/missing, and values in the legal label range?
Discriminator
A real violation is the absence of any template-based or shape-based verification (or a path/ordering/row-count mismatch that such a check would catch); it is not a violation if the script derives the frame from the template/test keys and asserts shape, key equality, header and value ranges before saving — even if it never literally opens the sample file, provided an equivalent assertion exists.
Consequence
The grader reports the expected submission file as WRONG/MISSING (unparseable, missing/extra/misordered IDs, wrong path or header), so the score is void regardless of how good the cross-validated model was.
id 3ea6121b0ee5 · mined from da-code dacode-ml-competition-006@s12
raw text (what the judge reads)
### Submission artifact never validated against the provided template
- **Applies when**: `task` -- the task specifies an output file whose format is defined by a sample/template file (header names, key column, one row per test key, expected location).
- **Pattern**: The scripts build predictions and write an output file from scratch (hand-constructed DataFrame, ad-hoc path) without ever loading the sample template or asserting that the written file matches it in row count, key coverage/order, column names, and value dtype/range; the agent then reports success on the basis of model CV scores alone.
- **Detection procedure**:
  1. Read the task/README for the required output filename, location, header, and the rule "one prediction per test key".
  2. Search the scripts for a read of the sample/template file and for post-write checks (e.g., re-reading the saved file, comparing `len` and the key set/order to the test/template file, checking header spelling and integer vs float output).
  3. Compare the write path in the script to the path the grader expects, and confirm the ID column comes from the test keys in template order.
  4. Inspect the reported answer: does it contain exactly as many rows as the test set, all keys, no duplicates/missing, and values in the legal label range?
- **Discriminator**: A real violation is the absence of any template-based or shape-based verification (or a path/ordering/row-count mismatch that such a check would catch); it is *not* a violation if the script derives the frame from the template/test keys and asserts shape, key equality, header and value ranges before saving — even if it never literally opens the sample file, provided an equivalent assertion exists.
- **Consequence**: The grader reports the expected submission file as WRONG/MISSING (unparseable, missing/extra/misordered IDs, wrong path or header), so the score is void regardless of how good the cross-validated model was.
665Group membership defined by naive null-detection without inspecting raw placeholder valuestaskinfiagent-dabench
Applies when
task -- the analysis splits rows into groups by "missing vs. present" in a column (or otherwise filters rows) and computes per-group statistics/tests.
Pattern
The script relies solely on the loader's default missingness parsing (isnull()/notnull() after a plain read_csv) without ever inspecting the raw distinct values of the grouping column, so sentinel/placeholder entries (empty strings, whitespace, "NA", "None", "-", 0, "nan" as text) or parser artifacts (wrong index_col/separator/na_values, mis-shifted columns) silently land in the wrong group; the reported group means are then computed over the wrong subsets, and no check confirms the split is plausible.
Detection procedure
  1. Read the task to see exactly which rows must belong to each group and on what column the split is defined.
  2. In the scripts, check whether the grouping column's raw unique values / value counts / dtype are printed and reasoned about before the mask is built, and whether the loader arguments (index column, delimiter, na_values, keep_default_na) are justified against a printed sample of raw rows.
  3. Check whether group sizes are validated: do the two group counts sum to the total row count, are both non-trivial, and is the analysis column's own missingness handled consistently (e.g., dropping NaNs for the test but not for the mean, or vice versa)?
  4. Compare the reported per-group statistics against any independently computable sanity anchor (overall mean as a size-weighted blend of the two group means, expected magnitudes/ranges); flag if the answer is submitted with no such cross-check.
Discriminator
A real violation is when the grouping column plausibly contains non-standard missing markers or the load could mis-align columns and the script never verifies this (no unique-value dump, no count reconciliation, no consistency check between the mean subset and the test subset). It is not a violation if the script explicitly inspects raw values/dtypes, documents that missingness is genuine NaN, and reconciles group counts with the total — even if it then uses a one-line isnull() mask.
Consequence
Group assignments are shifted, so both reported group means (and often the test statistic) deviate from ground truth by several units while still looking internally consistent; the grader marks the numeric fields wrong even though the p-value/format happens to look right.
id 386120e28002 · mined from infiagent-dabench dabench-297@s12
raw text (what the judge reads)
### Group membership defined by naive null-detection without inspecting raw placeholder values
- **Applies when**: `task` -- the analysis splits rows into groups by "missing vs. present" in a column (or otherwise filters rows) and computes per-group statistics/tests.
- **Pattern**: The script relies solely on the loader's default missingness parsing (`isnull()`/`notnull()` after a plain `read_csv`) without ever inspecting the raw distinct values of the grouping column, so sentinel/placeholder entries (empty strings, whitespace, `"NA"`, `"None"`, `"-"`, `0`, `"nan"` as text) or parser artifacts (wrong `index_col`/separator/`na_values`, mis-shifted columns) silently land in the wrong group; the reported group means are then computed over the wrong subsets, and no check confirms the split is plausible.
- **Detection procedure**:
  1. Read the task to see exactly which rows must belong to each group and on what column the split is defined.
  2. In the scripts, check whether the grouping column's raw unique values / value counts / dtype are printed and reasoned about before the mask is built, and whether the loader arguments (index column, delimiter, `na_values`, `keep_default_na`) are justified against a printed sample of raw rows.
  3. Check whether group sizes are validated: do the two group counts sum to the total row count, are both non-trivial, and is the analysis column's own missingness handled consistently (e.g., dropping NaNs for the test but not for the mean, or vice versa)?
  4. Compare the reported per-group statistics against any independently computable sanity anchor (overall mean as a size-weighted blend of the two group means, expected magnitudes/ranges); flag if the answer is submitted with no such cross-check.
- **Discriminator**: A real violation is when the grouping column plausibly contains non-standard missing markers or the load could mis-align columns and the script never verifies this (no unique-value dump, no count reconciliation, no consistency check between the mean subset and the test subset). It is *not* a violation if the script explicitly inspects raw values/dtypes, documents that missingness is genuine `NaN`, and reconciles group counts with the total — even if it then uses a one-line `isnull()` mask.
- **Consequence**: Group assignments are shifted, so both reported group means (and often the test statistic) deviate from ground truth by several units while still looking internally consistent; the grader marks the numeric fields wrong even though the p-value/format happens to look right.
666Missing required output artifacts (only part of the deliverables produced)taskda-code
Applies when
task -- the task (or a referenced guidance/spec file) states that the analysis must produce specific saved outputs (figures, serialized arrays, JSON summaries, tables) in addition to a textual answer.
Pattern
The script implements the visible/most salient deliverable (e.g. the plot image) and reports numbers in prose, but never writes the other required artifact files, and the agent never re-reads the referenced spec to enumerate every expected output or the exact definitions/filters it prescribes — inventing its own preprocessing and categorization rules instead.
Detection procedure
  1. From the task statement and any referenced guidance/spec document, list every artifact that must exist on disk (name, format) and every stated constraint (filters, definitions, ordering, sizes, colors).
  2. Grep the scripts for write/save calls (savefig, save, to_csv, json.dump, etc.) and build the set of files actually produced.
  3. Compare the two sets: flag if any required file is absent, misnamed, or in a different directory than requested; also flag any filtering/derivation step in the script that has no counterpart in the spec.
  4. Check the final answer text for a claim that "the outputs were generated" without evidence of each file being written.
Discriminator
A real violation is a required artifact that no code path creates (or a derivation rule invented by the agent rather than taken from the spec); a look-alike that is fine is an extra, unrequested file, or an artifact written under the exact requested name via a helper/config path that a reader can trace.
Consequence
Automated checks for the missing files report WRONG/MISSING, so the attempt scores zero on those checks even if the reported numbers look plausible; self-invented filters also make the one produced artifact mismatch the reference values.
id 498942e8cb8c · mined from da-code dacode-plot-pie-005@s12
raw text (what the judge reads)
### Missing required output artifacts (only part of the deliverables produced)
- **Applies when**: `task` -- the task (or a referenced guidance/spec file) states that the analysis must produce specific saved outputs (figures, serialized arrays, JSON summaries, tables) in addition to a textual answer.
- **Pattern**: The script implements the visible/most salient deliverable (e.g. the plot image) and reports numbers in prose, but never writes the other required artifact files, and the agent never re-reads the referenced spec to enumerate every expected output or the exact definitions/filters it prescribes — inventing its own preprocessing and categorization rules instead.
- **Detection procedure**:
  1. From the task statement and any referenced guidance/spec document, list every artifact that must exist on disk (name, format) and every stated constraint (filters, definitions, ordering, sizes, colors).
  2. Grep the scripts for write/save calls (`savefig`, `save`, `to_csv`, `json.dump`, etc.) and build the set of files actually produced.
  3. Compare the two sets: flag if any required file is absent, misnamed, or in a different directory than requested; also flag any filtering/derivation step in the script that has no counterpart in the spec.
  4. Check the final answer text for a claim that "the outputs were generated" without evidence of each file being written.
- **Discriminator**: A real violation is a required artifact that no code path creates (or a derivation rule invented by the agent rather than taken from the spec); a look-alike that is fine is an extra, unrequested file, or an artifact written under the exact requested name via a helper/config path that a reader can trace.
- **Consequence**: Automated checks for the missing files report WRONG/MISSING, so the attempt scores zero on those checks even if the reported numbers look plausible; self-invented filters also make the one produced artifact mismatch the reference values.
667Missing reproducible script + un-sanity-checked error magnitude relative to target variancetaskinfiagent-dabench
Applies when
task -- a task prescribes an exact preprocessing/split/metric recipe (e.g., mean-imputation of specific columns, fixed train/test fraction, a stated error metric) and the agent reports a single numeric score.
Pattern
The attempt reports a score with no saved, re-runnable script showing each prescribed step (column parsing/dtype coercion, imputation applied to the stated columns, split, fit, metric on the held-out set), and never compares the reported error against a trivial baseline (variance of the target / error of predicting the target mean), so a broken preprocessing step (strings or mixed units left uncoerced, rows silently dropped or NaNs propagated, imputation done on the wrong axis/after the split, features and target mismatched) inflates the number by an order of magnitude without notice.
Detection procedure
  1. From the task, list every mandated step (which columns are cleaned and how, split proportion, metric definition, rounding/format) and the plausible scale of the target variable.
  2. In the scripts, check that each mandated step exists explicitly and in a defensible order, that the modeled columns are numeric after loading (explicit to_numeric/dtype check), and that row counts/shapes are printed after cleaning and after the split.
  3. Check whether the script computes a baseline reference (target variance or MSE of the mean predictor on the test set) and compares it to the reported error.
  4. Compare the reported number to that reference: flag if the model's error is at or above the naive-baseline error, or if no such comparison and no shape/dtype evidence exists.
Discriminator
A genuine violation is an unverifiable or clearly non-conforming pipeline, or an error value that a regression on informative features could not plausibly produce (≥ the variance of the target). It is not a violation when the pipeline is fully shown, dtypes/counts are validated, and the error is simply large because the features are weakly predictive but still below the mean-predictor baseline.
Consequence
The reported metric differs from the reference value by a large factor (here ~10×), so the exact-match grader marks the single required field wrong and the task scores 0.
id 4f156b1a4f56 · mined from infiagent-dabench dabench-432@s12
raw text (what the judge reads)
### Missing reproducible script + un-sanity-checked error magnitude relative to target variance
- **Applies when**: `task` -- a task prescribes an exact preprocessing/split/metric recipe (e.g., mean-imputation of specific columns, fixed train/test fraction, a stated error metric) and the agent reports a single numeric score.
- **Pattern**: The attempt reports a score with no saved, re-runnable script showing each prescribed step (column parsing/dtype coercion, imputation applied to the stated columns, split, fit, metric on the held-out set), and never compares the reported error against a trivial baseline (variance of the target / error of predicting the target mean), so a broken preprocessing step (strings or mixed units left uncoerced, rows silently dropped or NaNs propagated, imputation done on the wrong axis/after the split, features and target mismatched) inflates the number by an order of magnitude without notice.
- **Detection procedure**:
  1. From the task, list every mandated step (which columns are cleaned and how, split proportion, metric definition, rounding/format) and the plausible scale of the target variable.
  2. In the scripts, check that each mandated step exists explicitly and in a defensible order, that the modeled columns are numeric after loading (explicit `to_numeric`/dtype check), and that row counts/shapes are printed after cleaning and after the split.
  3. Check whether the script computes a baseline reference (target variance or MSE of the mean predictor on the test set) and compares it to the reported error.
  4. Compare the reported number to that reference: flag if the model's error is at or above the naive-baseline error, or if no such comparison and no shape/dtype evidence exists.
- **Discriminator**: A genuine violation is an unverifiable or clearly non-conforming pipeline, or an error value that a regression on informative features could not plausibly produce (≥ the variance of the target). It is *not* a violation when the pipeline is fully shown, dtypes/counts are validated, and the error is simply large because the features are weakly predictive but still below the mean-predictor baseline.
- **Consequence**: The reported metric differs from the reference value by a large factor (here ~10×), so the exact-match grader marks the single required field wrong and the task scores 0.
668Silent row-reordering before a lag/sequence-dependent computationtaskinfiagent-dabench
Applies when
task -- The requested statistic depends on row order (differences, lags, cumulative sums, rolling windows, time-series splits) and the script re-sorts, reverses, or re-indexes the data before computing it.
Pattern
The agent asserts an ordering assumption (e.g., "rows are reversed, so sort ascending"), applies a sort/shift on the reordered frame, and reports the result without verifying the assumption against the file or checking that the ordering choice is the one the task's definition implies ("previous row" as stored vs. "previous timestamp"). A wrong ordering flips the sign of every difference and shifts the mean, while leaving the dispersion nearly unchanged — so nothing looks obviously broken.
Detection procedure
  1. In the task text, identify whether the requested quantity is order-sensitive and what "previous"/"next" is defined relative to (raw row order vs. parsed timestamp order).
  2. In the script, find every sort_values, sort_index, [::-1], reset_index, groupby or merge that could change row order before the shift/diff/rolling call, and check whether the script printed evidence (first/last few raw rows and parsed dates) confirming the original order was actually what the agent claimed.
  3. Check whether the script computes the statistic under both plausible orderings (or at least sanity-checks the sign/magnitude of the mean against a known start-to-end price/level change: mean per-period change should have the same sign as the overall net change over the period).
  4. Compare the reported sign/magnitude of the answer with that end-to-end sanity check; an unexplained sign disagreement is a red flag.
Discriminator
A real violation is a reordering (or lack of it) that is asserted rather than demonstrated, with no cross-check of the resulting sign/magnitude. It is fine if the script prints the raw head/tail with parsed dates, shows the data was genuinely in the assumed order, and confirms the aggregate result is consistent with the net change across the series — or if the computation is order-invariant.
Consequence
The dispersion statistic comes out roughly right (off only by rounding/one dropped observation) while the mean is reported with the wrong sign, so the answer fails all exact-match checks despite looking internally consistent.
id 539b52ecc320 · mined from infiagent-dabench dabench-75@s12
raw text (what the judge reads)
### Silent row-reordering before a lag/sequence-dependent computation
- **Applies when**: `task` -- The requested statistic depends on row order (differences, lags, cumulative sums, rolling windows, time-series splits) and the script re-sorts, reverses, or re-indexes the data before computing it.
- **Pattern**: The agent asserts an ordering assumption (e.g., "rows are reversed, so sort ascending"), applies a sort/`shift` on the reordered frame, and reports the result without verifying the assumption against the file or checking that the ordering choice is the one the task's definition implies ("previous row" as stored vs. "previous timestamp"). A wrong ordering flips the sign of every difference and shifts the mean, while leaving the dispersion nearly unchanged — so nothing looks obviously broken.
- **Detection procedure**:
  1. In the task text, identify whether the requested quantity is order-sensitive and what "previous"/"next" is defined relative to (raw row order vs. parsed timestamp order).
  2. In the script, find every `sort_values`, `sort_index`, `[::-1]`, `reset_index`, `groupby` or merge that could change row order before the `shift`/`diff`/rolling call, and check whether the script printed evidence (first/last few raw rows and parsed dates) confirming the original order was actually what the agent claimed.
  3. Check whether the script computes the statistic under both plausible orderings (or at least sanity-checks the sign/magnitude of the mean against a known start-to-end price/level change: mean per-period change should have the same sign as the overall net change over the period).
  4. Compare the reported sign/magnitude of the answer with that end-to-end sanity check; an unexplained sign disagreement is a red flag.
- **Discriminator**: A real violation is a reordering (or lack of it) that is asserted rather than demonstrated, with no cross-check of the resulting sign/magnitude. It is fine if the script prints the raw head/tail with parsed dates, shows the data was genuinely in the assumed order, and confirms the aggregate result is consistent with the net change across the series — or if the computation is order-invariant.
- **Consequence**: The dispersion statistic comes out roughly right (off only by rounding/one dropped observation) while the mean is reported with the wrong sign, so the answer fails all exact-match checks despite looking internally consistent.
669Unvalidated row-subset / missing-value handling for a statistic reported at coarse rounding precisiontaskinfiagent-dabench
Applies when
task -- the task asks for a single summary statistic (correlation, mean, coefficient, metric) rounded to few decimals, and the script must decide which rows of a table are valid inputs (NaNs, sentinel values like 0/-1/"NA" strings, non-numeric dtypes, header/index quirks).
Pattern
The script silently adopts one row-filtering convention (e.g., per-column dropna() then index intersection, or no filtering at all) without ever printing/justifying how many rows were kept versus the raw file, without inspecting the columns for sentinel or non-numeric entries, and without checking whether an equally defensible convention (whole-row dropna, excluding sentinel/zero rows, coercing dtypes) would change the value at the requested rounding precision.
Detection procedure
  1. Read the task for the population of records it implies; note the rounding precision demanded for the reported number.
  2. In the script, locate where rows are dropped/kept and check whether the raw row count, the post-filter count, and per-column dtype/NaN/sentinel diagnostics are printed and compared.
  3. Check whether the script computes the statistic under at least one alternative but reasonable filtering/dtype convention and confirms the rounded value is identical.
  4. Inspect the reported number: if it sits close to a rounding boundary (unrounded value within ~0.005 of the .xx5 cut, or the raw value differs across conventions) and no such robustness check exists, flag it.
Discriminator
Fine if the script demonstrates the columns are fully numeric with no missing/sentinel values (so all conventions coincide), or explicitly shows the rounded statistic is stable across conventions; a violation is when the filtering choice is implicit, undocumented in output, and the final rounded digit could plausibly flip.
Consequence
The reported statistic differs from the reference by one unit in the last reported digit (e.g., 0.53 vs 0.54), so the numeric check fails even though the qualitative conclusion (significant / not significant) is graded correct.
id 01bd3905554a · mined from infiagent-dabench dabench-300@s12
raw text (what the judge reads)
### Unvalidated row-subset / missing-value handling for a statistic reported at coarse rounding precision
- **Applies when**: `task` -- the task asks for a single summary statistic (correlation, mean, coefficient, metric) rounded to few decimals, and the script must decide which rows of a table are valid inputs (NaNs, sentinel values like 0/-1/"NA" strings, non-numeric dtypes, header/index quirks).
- **Pattern**: The script silently adopts one row-filtering convention (e.g., per-column `dropna()` then index intersection, or no filtering at all) without ever printing/justifying how many rows were kept versus the raw file, without inspecting the columns for sentinel or non-numeric entries, and without checking whether an equally defensible convention (whole-row `dropna`, excluding sentinel/zero rows, coercing dtypes) would change the value at the requested rounding precision.
- **Detection procedure**:
  1. Read the task for the population of records it implies; note the rounding precision demanded for the reported number.
  2. In the script, locate where rows are dropped/kept and check whether the raw row count, the post-filter count, and per-column dtype/NaN/sentinel diagnostics are printed and compared.
  3. Check whether the script computes the statistic under at least one alternative but reasonable filtering/dtype convention and confirms the rounded value is identical.
  4. Inspect the reported number: if it sits close to a rounding boundary (unrounded value within ~0.005 of the .xx5 cut, or the raw value differs across conventions) and no such robustness check exists, flag it.
- **Discriminator**: Fine if the script demonstrates the columns are fully numeric with no missing/sentinel values (so all conventions coincide), or explicitly shows the rounded statistic is stable across conventions; a violation is when the filtering choice is implicit, undocumented in output, and the final rounded digit could plausibly flip.
- **Consequence**: The reported statistic differs from the reference by one unit in the last reported digit (e.g., 0.53 vs 0.54), so the numeric check fails even though the qualitative conclusion (significant / not significant) is graded correct.
670Submitting predictions from a model/blend that was never validatedtaskda-code
Applies when
task -- the script evaluates several candidate models on a hold-out split but the file written for submission comes from a different estimator, a hand-picked weighted blend, or a re-fit/re-weighted variant.
Pattern
The attempt computes validation scores for individual models, declares one "best," then writes an arbitrary fixed-weight combination (or a model refit on different rows/features) whose accuracy was never measured — including components with visibly poor validation scores (e.g. a heavily regularized linear model on ordinally label-encoded high-cardinality categoricals) that drag the blend away from the best validated predictor. No post-hoc sanity check ties the written file back to any measured error.
Detection procedure
  1. From the task, note that the graded artifact is the prediction file, so its accuracy is what matters.
  2. In the script, identify exactly which object/expression produces the values written to the output file, and list every model and weight involved.
  3. Check whether that exact predictor (same weights, same fitted rows, same feature matrix) has a reported hold-out score; compare the reported per-model scores to see if any weighted component is clearly weaker than the best model.
  4. Confirm the script sanity-checks the output (row count equal to the test rows, order preserved, value range/distribution comparable to the training target); absence of both a validation score for the blend and these checks is the violation.
Discriminator
A fine attempt either submits the single predictor whose hold-out score it reported, or evaluates the ensemble itself (weights chosen by CV/stacking on validation data) and shows it beats the individual models; a violation submits an unmeasured combination or an unevaluated refit, especially one that includes components with much worse validation performance.
Consequence
The submitted predictions have higher error than the best model the script already had, so the grader's accuracy/threshold check on the prediction file fails even though the printed validation numbers looked acceptable.
id 0df74b679675 · mined from da-code dacode-ml-regression-014@s12
raw text (what the judge reads)
### Submitting predictions from a model/blend that was never validated
- **Applies when**: `task` -- the script evaluates several candidate models on a hold-out split but the file written for submission comes from a different estimator, a hand-picked weighted blend, or a re-fit/re-weighted variant.
- **Pattern**: The attempt computes validation scores for individual models, declares one "best," then writes an arbitrary fixed-weight combination (or a model refit on different rows/features) whose accuracy was never measured — including components with visibly poor validation scores (e.g. a heavily regularized linear model on ordinally label-encoded high-cardinality categoricals) that drag the blend away from the best validated predictor. No post-hoc sanity check ties the written file back to any measured error.
- **Detection procedure**:
  1. From the task, note that the graded artifact is the prediction file, so its accuracy is what matters.
  2. In the script, identify exactly which object/expression produces the values written to the output file, and list every model and weight involved.
  3. Check whether that exact predictor (same weights, same fitted rows, same feature matrix) has a reported hold-out score; compare the reported per-model scores to see if any weighted component is clearly weaker than the best model.
  4. Confirm the script sanity-checks the output (row count equal to the test rows, order preserved, value range/distribution comparable to the training target); absence of both a validation score for the blend and these checks is the violation.
- **Discriminator**: A fine attempt either submits the single predictor whose hold-out score it reported, or evaluates the ensemble itself (weights chosen by CV/stacking on validation data) and shows it beats the individual models; a violation submits an unmeasured combination or an unevaluated refit, especially one that includes components with much worse validation performance.
- **Consequence**: The submitted predictions have higher error than the best model the script already had, so the grader's accuracy/threshold check on the prediction file fails even though the printed validation numbers looked acceptable.
671No held-out validation — model quality judged by training-set score onlytaskda-code
Applies when
task -- the script fits a supervised model on labeled data and produces predictions for an unlabeled evaluation set that will be scored against hidden ground truth.
Pattern
The attempt trains a single, minimally-tuned model on all labeled rows, computes and reports the score on those same training rows (or no score at all), and submits the predictions without any out-of-sample estimate, baseline comparison, or model/feature alternative. The reported "accuracy" is therefore an optimistic in-sample number that says nothing about whether the submission clears the grading bar.
Detection procedure
  1. Read the task to see that scoring is done externally on data whose labels the agent cannot see, so the only internal quality signal must come from a held-out split or cross-validation.
  2. Scan the script for any train/validation split, cross_val_score, K-fold loop, or evaluation on rows excluded from fit; note whether the score printed is computed on the exact same feature matrix passed to fit.
  3. Check whether more than one configuration (features, vectorizer settings, classifier, or aggressive capacity limits such as a small feature cap) was compared using that held-out signal.
  4. Read the final answer: if the headline quality number is the in-sample score and is presented as evidence the submission is good, flag it.
Discriminator
A real violation is when no score is computed on data withheld from fitting; it is fine if the agent fits a final model on all labeled data after first measuring held-out/CV performance (and reports that number), or if the task supplies its own labeled validation set that the agent scores on.
Consequence
The submission format looks correct (right column name, right row count), but predictive accuracy is unknown and typically sits below the grader's threshold, so the file is marked WRONG while the answer confidently cites a high training accuracy.
id 45d4fa3a631a · mined from da-code dacode-ml-multi-008@s12
raw text (what the judge reads)
### No held-out validation — model quality judged by training-set score only
- **Applies when**: `task` -- the script fits a supervised model on labeled data and produces predictions for an unlabeled evaluation set that will be scored against hidden ground truth.
- **Pattern**: The attempt trains a single, minimally-tuned model on all labeled rows, computes and reports the score on those same training rows (or no score at all), and submits the predictions without any out-of-sample estimate, baseline comparison, or model/feature alternative. The reported "accuracy" is therefore an optimistic in-sample number that says nothing about whether the submission clears the grading bar.
- **Detection procedure**:
  1. Read the task to see that scoring is done externally on data whose labels the agent cannot see, so the only internal quality signal must come from a held-out split or cross-validation.
  2. Scan the script for any train/validation split, `cross_val_score`, K-fold loop, or evaluation on rows excluded from `fit`; note whether the score printed is computed on the exact same feature matrix passed to `fit`.
  3. Check whether more than one configuration (features, vectorizer settings, classifier, or aggressive capacity limits such as a small feature cap) was compared using that held-out signal.
  4. Read the final answer: if the headline quality number is the in-sample score and is presented as evidence the submission is good, flag it.
- **Discriminator**: A real violation is when *no* score is computed on data withheld from fitting; it is fine if the agent fits a final model on all labeled data *after* first measuring held-out/CV performance (and reports that number), or if the task supplies its own labeled validation set that the agent scores on.
- **Consequence**: The submission format looks correct (right column name, right row count), but predictive accuracy is unknown and typically sits below the grader's threshold, so the file is marked WRONG while the answer confidently cites a high training accuracy.
672Deliverables reduced to a single text answer; required artifact files never producedtaskda-code
Applies when
task -- the prompt asks for output files (charts, serialized arrays/tables, config-driven plots) in addition to or instead of a scalar/label answer, and the agent's work must be reconstructible from saved scripts.
Pattern
The agent computes only the first, intermediate step (e.g., the grouping/ranking that selects a subset) and reports that value as the final answer, while the requested files are never written, are written to different names/paths/formats, or ignore the styling/parameter spec file the task points to; sometimes no scripts are saved at all, so nothing can be re-run.
Detection procedure
  1. Read the task and list every concrete deliverable: each filename/extension, its expected content, and every external spec (config/YAML/JSON) whose settings must be honored.
  2. Read the scripts and check that each deliverable has an explicit write call with exactly the required name/path/format, and that the spec file is actually loaded and its keys applied (not hard-coded defaults).
  3. Compare the agent's final answer with the deliverable list: if it reports only an intermediate selection value and no file-producing code exists, mark inadequate.
  4. Sanity-check the produced artifacts' content (category counts/shares sum correctly, chart drawn on the selected subset only, saved array shape/dtype plausible) before accepting.
Discriminator
A real violation is a missing/renamed/unspecified-format artifact or an unread spec file; it is not a violation if all required files are written with correct names and spec-driven parameters and the agent additionally states the intermediate value as context.
Consequence
Graders that check artifacts report each expected file as WRONG/MISSING, yielding 0 passed checks even when the intermediate textual answer happens to be right.
id 3a84ba412a1c · mined from da-code dacode-plot-pie-008@s12
raw text (what the judge reads)
### Deliverables reduced to a single text answer; required artifact files never produced
- **Applies when**: `task` -- the prompt asks for output files (charts, serialized arrays/tables, config-driven plots) in addition to or instead of a scalar/label answer, and the agent's work must be reconstructible from saved scripts.
- **Pattern**: The agent computes only the first, intermediate step (e.g., the grouping/ranking that selects a subset) and reports that value as the final answer, while the requested files are never written, are written to different names/paths/formats, or ignore the styling/parameter spec file the task points to; sometimes no scripts are saved at all, so nothing can be re-run.
- **Detection procedure**:
  1. Read the task and list every concrete deliverable: each filename/extension, its expected content, and every external spec (config/YAML/JSON) whose settings must be honored.
  2. Read the scripts and check that each deliverable has an explicit write call with exactly the required name/path/format, and that the spec file is actually loaded and its keys applied (not hard-coded defaults).
  3. Compare the agent's final answer with the deliverable list: if it reports only an intermediate selection value and no file-producing code exists, mark inadequate.
  4. Sanity-check the produced artifacts' content (category counts/shares sum correctly, chart drawn on the selected subset only, saved array shape/dtype plausible) before accepting.
- **Discriminator**: A real violation is a missing/renamed/unspecified-format artifact or an unread spec file; it is *not* a violation if all required files are written with correct names and spec-driven parameters and the agent additionally states the intermediate value as context.
- **Consequence**: Graders that check artifacts report each expected file as WRONG/MISSING, yielding 0 passed checks even when the intermediate textual answer happens to be right.
673Missing validation/cleaning of corrupted or implausible values before computing the statistictaskinfiagent-dabench
Applies when
task -- a statistic must be computed from raw fields that are parsed/derived (dates, strings, currency, categorical codes) in a source known or likely to contain data-entry errors, and the scripts only drop nulls before analysis.
Pattern
The attempt writes a parser/derivation, filters on notna() alone, and immediately feeds the result into the statistic — never checking whether derived or raw values fall in a plausible range (e.g., negative or absurdly large derived quantities, out-of-range category codes, sign/magnitude anomalies, duplicated records, mis-parsed formats silently coerced to a number). Any residual bad rows shift the correlation/metric by a few hundredths, which the agent then reports without cross-validating.
Detection procedure
  1. Read the task and note whether the input is flagged as containing errors, or whether any analysis variable is derived from free-text/parsed fields.
  2. In the scripts, list every cleaning step actually applied; check for explicit range/domain validation of each analysis variable (min/max, allowed category set, non-negativity, monotone date order) and for handling of parse branches that return np.nan vs. return a wrong number.
  3. Check whether the agent printed diagnostics that would reveal anomalies (value counts, describe with min/max, count of rows dropped) and whether it reacted to any anomaly it saw.
  4. Check whether the reported statistic was re-derived under at least one alternative cleaning assumption to confirm stability of the rounded value.
Discriminator
A real violation is when no domain/range check exists (or anomalies appear in printed output and are ignored) for variables that can plausibly be corrupted; it is not a violation if the agent explicitly validated ranges/categories and documented that no invalid rows exist, or if the dataset is a clean, purely numeric, pre-validated source.
Consequence
The statistic is computed on a contaminated subset, so the coefficient (and its rounded two-decimal report) deviates from the ground truth even though derived flags like significance/relationship type still match — a partial-credit failure like 2/3 checks passed.
id f900ea82422e · mined from infiagent-dabench dabench-431@s12
raw text (what the judge reads)
### Missing validation/cleaning of corrupted or implausible values before computing the statistic
- **Applies when**: `task` -- a statistic must be computed from raw fields that are parsed/derived (dates, strings, currency, categorical codes) in a source known or likely to contain data-entry errors, and the scripts only drop nulls before analysis.
- **Pattern**: The attempt writes a parser/derivation, filters on `notna()` alone, and immediately feeds the result into the statistic — never checking whether derived or raw values fall in a plausible range (e.g., negative or absurdly large derived quantities, out-of-range category codes, sign/magnitude anomalies, duplicated records, mis-parsed formats silently coerced to a number). Any residual bad rows shift the correlation/metric by a few hundredths, which the agent then reports without cross-validating.
- **Detection procedure**:
  1. Read the task and note whether the input is flagged as containing errors, or whether any analysis variable is derived from free-text/parsed fields.
  2. In the scripts, list every cleaning step actually applied; check for explicit range/domain validation of each analysis variable (min/max, allowed category set, non-negativity, monotone date order) and for handling of parse branches that `return np.nan` vs. return a wrong number.
  3. Check whether the agent printed diagnostics that would reveal anomalies (value counts, describe with min/max, count of rows dropped) and whether it reacted to any anomaly it saw.
  4. Check whether the reported statistic was re-derived under at least one alternative cleaning assumption to confirm stability of the rounded value.
- **Discriminator**: A real violation is when no domain/range check exists (or anomalies appear in printed output and are ignored) for variables that can plausibly be corrupted; it is *not* a violation if the agent explicitly validated ranges/categories and documented that no invalid rows exist, or if the dataset is a clean, purely numeric, pre-validated source.
- **Consequence**: The statistic is computed on a contaminated subset, so the coefficient (and its rounded two-decimal report) deviates from the ground truth even though derived flags like significance/relationship type still match — a partial-credit failure like 2/3 checks passed.
674Deliverable file schema not exactly as specifiedtaskda-code
Applies when
task -- the task states the output file name and the exact column(s) it must contain, and the script writes predictions to that file.
Pattern
The script writes the requested file but with a schema the task did not ask for — e.g., extra identifier/index columns, renamed or extra headers, different column order, or a different dtype/label encoding than implied — and the agent's answer describes this altered schema as if it satisfied the request without ever re-reading the written file and comparing it to the spec.
Detection procedure
  1. Extract from the task the literal required filename and the literal required column name(s)/values, plus any implied row count or ordering.
  2. In the script, find the DataFrame construction and to_csv/write call; list the exact columns, their order, index flag, and value domain being written.
  3. Diff (2) against (1): flag any additional column, missing column, spelling/case mismatch, wrong dtype (e.g., probabilities or strings where binary labels are expected), or row count/order not matching the test input.
  4. Check the answer for a post-write verification step (re-load the file, assert columns and shape); absence of such a check plus any diff in step 3 is a violation.
Discriminator
A real violation is a concrete deviation from an explicitly stated schema (extra ID column, misspelled header, floats instead of labels). A look-alike that is fine is when the task itself is silent on extra columns and the required column is present with correct values, verified by reloading the file.
Consequence
The grader's file comparison fails on schema/parse mismatch (or reads the wrong column), so the submission is scored wrong/missing even if the underlying model predictions were reasonable.
id 8616e7cbc638 · mined from da-code dacode-ml-binary-016@s12
raw text (what the judge reads)
### Deliverable file schema not exactly as specified
- **Applies when**: `task` -- the task states the output file name and the exact column(s) it must contain, and the script writes predictions to that file.
- **Pattern**: The script writes the requested file but with a schema the task did not ask for — e.g., extra identifier/index columns, renamed or extra headers, different column order, or a different dtype/label encoding than implied — and the agent's answer describes this altered schema as if it satisfied the request without ever re-reading the written file and comparing it to the spec.
- **Detection procedure**:
  1. Extract from the task the literal required filename and the literal required column name(s)/values, plus any implied row count or ordering.
  2. In the script, find the DataFrame construction and `to_csv`/write call; list the exact columns, their order, index flag, and value domain being written.
  3. Diff (2) against (1): flag any additional column, missing column, spelling/case mismatch, wrong dtype (e.g., probabilities or strings where binary labels are expected), or row count/order not matching the test input.
  4. Check the answer for a post-write verification step (re-load the file, assert columns and shape); absence of such a check plus any diff in step 3 is a violation.
- **Discriminator**: A real violation is a concrete deviation from an explicitly stated schema (extra `ID` column, misspelled header, floats instead of labels). A look-alike that is fine is when the task itself is silent on extra columns and the required column is present with correct values, verified by reloading the file.
- **Consequence**: The grader's file comparison fails on schema/parse mismatch (or reads the wrong column), so the submission is scored wrong/missing even if the underlying model predictions were reasonable.
675Answer does not instantiate the requested output schema (and is not persisted where required)taskda-code
Applies when
task -- the prompt supplies a literal JSON/text template with typed placeholders (e.g. list brackets, key names, ordering) and/or names an output file the result must be written to.
Pattern
The agent reports a semantically plausible value but changes the container type or shape of the placeholder (scalar where a list is shown, extra/renamed keys, missing entries) and/or only prints the result to stdout instead of writing the specified result file, with no script that serializes the template.
Detection procedure
  1. Copy the template from the task and note, for each placeholder, its exact key name and value type/shape, plus any required output filename.
  2. Search the scripts for the code that builds and dumps the final object; check it constructs exactly those keys with exactly those value types and writes to the named file (e.g. json.dump(..., open("result.json","w"))).
  3. Compare the submitted answer literal against the template placeholder-by-placeholder (type, key spelling, count of elements, ties handled as multiple entries if the template allows them).
  4. Flag if any placeholder type/shape differs, a key is missing/renamed, or no artifact-writing step exists.
Discriminator
A real violation is a structural mismatch (scalar vs list, missing file, renamed key) or dropping legitimate tied/multiple values; harmless look-alikes are cosmetic differences the grader normalizes, such as whitespace, key order, or trailing newline, when the keys and value types still match.
Consequence
The grader loads the expected result file and compares parsed structures, so it reports the file as WRONG/MISSING and scores 0 even when the underlying computation is close to right.
id 41d7240e3147 · mined from da-code dacode-di-text-001@s12
raw text (what the judge reads)
### Answer does not instantiate the requested output schema (and is not persisted where required)
- **Applies when**: `task` -- the prompt supplies a literal JSON/text template with typed placeholders (e.g. list brackets, key names, ordering) and/or names an output file the result must be written to.
- **Pattern**: The agent reports a semantically plausible value but changes the container type or shape of the placeholder (scalar where a list is shown, extra/renamed keys, missing entries) and/or only prints the result to stdout instead of writing the specified result file, with no script that serializes the template.
- **Detection procedure**:
  1. Copy the template from the task and note, for each placeholder, its exact key name and value type/shape, plus any required output filename.
  2. Search the scripts for the code that builds and dumps the final object; check it constructs exactly those keys with exactly those value types and writes to the named file (e.g. `json.dump(..., open("result.json","w"))`).
  3. Compare the submitted answer literal against the template placeholder-by-placeholder (type, key spelling, count of elements, ties handled as multiple entries if the template allows them).
  4. Flag if any placeholder type/shape differs, a key is missing/renamed, or no artifact-writing step exists.
- **Discriminator**: A real violation is a structural mismatch (scalar vs list, missing file, renamed key) or dropping legitimate tied/multiple values; harmless look-alikes are cosmetic differences the grader normalizes, such as whitespace, key order, or trailing newline, when the keys and value types still match.
- **Consequence**: The grader loads the expected result file and compares parsed structures, so it reports the file as WRONG/MISSING and scores 0 even when the underlying computation is close to right.
676Output file omits requested derived columns (over-pruned deliverable)taskda-code
Applies when
task -- the task asks to compute several intermediate quantities/labels and save results "including X and Y" to a specific output file, and the script subselects columns before writing.
Pattern
The script computes all the intermediate fields (scores, group codes, composite keys, assigned labels) but writes only a minimal subset (e.g., an ID plus one final label) to the deliverable, dropping the per-component values and/or the segmentation field the task explicitly named; the answer then asserts the file matches an expected reference without any real comparison.
Detection procedure
  1. From the task statement, list every quantity that must appear in the saved file (each component score, the composite/segment identifier, the final level, the entity ID) and any stated ordering/naming/format constraints.
  2. In the script, find the line that builds the DataFrame passed to to_csv/to_excel and enumerate exactly which columns survive; compare to the list from step 1.
  3. Check whether the script or answer performs a concrete verification (loading a reference/expected file and comparing) or merely asserts a match in prose; treat unsupported "verified 100% match" claims as a red flag.
  4. Confirm row count and ID coverage of the written file equal the source entity count.
Discriminator
A real violation is when a field the task named (or an obvious component of the requested result) is computed in memory but excluded from the saved file; it is not a violation if the task genuinely requests only the final label, or if the extra fields are present in the required file and merely also duplicated elsewhere. Note that writing the full table to a secondary file does not satisfy a requirement on the named deliverable.
Consequence
The grader compares the deliverable against the expected schema/content and reports the file as wrong/missing even though the underlying computation may be correct, yielding 0 on the file check.
id 16bc4858323f · mined from da-code dacode-dm-csv-052@s12
raw text (what the judge reads)
### Output file omits requested derived columns (over-pruned deliverable)
- **Applies when**: `task` -- the task asks to compute several intermediate quantities/labels and save results "including X and Y" to a specific output file, and the script subselects columns before writing.
- **Pattern**: The script computes all the intermediate fields (scores, group codes, composite keys, assigned labels) but writes only a minimal subset (e.g., an ID plus one final label) to the deliverable, dropping the per-component values and/or the segmentation field the task explicitly named; the answer then asserts the file matches an expected reference without any real comparison.
- **Detection procedure**:
  1. From the task statement, list every quantity that must appear in the saved file (each component score, the composite/segment identifier, the final level, the entity ID) and any stated ordering/naming/format constraints.
  2. In the script, find the line that builds the DataFrame passed to `to_csv`/`to_excel` and enumerate exactly which columns survive; compare to the list from step 1.
  3. Check whether the script or answer performs a concrete verification (loading a reference/expected file and comparing) or merely asserts a match in prose; treat unsupported "verified 100% match" claims as a red flag.
  4. Confirm row count and ID coverage of the written file equal the source entity count.
- **Discriminator**: A real violation is when a field the task named (or an obvious component of the requested result) is computed in memory but excluded from the saved file; it is *not* a violation if the task genuinely requests only the final label, or if the extra fields are present in the required file and merely also duplicated elsewhere. Note that writing the full table to a *secondary* file does not satisfy a requirement on the named deliverable.
- **Consequence**: The grader compares the deliverable against the expected schema/content and reports the file as wrong/missing even though the underlying computation may be correct, yielding 0 on the file check.
677Unverified input subset and unread output-format templatetaskda-code
Applies when
task -- the task specifies data restricted by explicit qualifiers (group, site, period, condition) and an output file whose layout is defined by a provided sample/template file.
Pattern
The script reads a single convenient input file, hard-codes assumed column names and a single filter (or none), never checks that the rows actually correspond to the stated qualifiers, and writes the output with column names/ordering/quoting invented by the agent instead of copied from the supplied template; the answer then asserts success without any evidence that inputs or output layout were validated.
Detection procedure
  1. From the task, list every restriction on the rows to be used and note that a sample/template output file was provided.
  2. In the scripts, check whether each restriction is applied explicitly (filter on the qualifier column, or documented proof the file already contains only those rows) and whether the template file is ever opened/parsed to derive header names, column order, and value formatting.
  3. In the printed output/answer, look for a validation trace: row counts per group, a peek at raw rows, and a printed diff/comparison of the written file's header against the template's header.
  4. Flag if any restriction is only assumed, or if the output header/format is asserted rather than derived from the template; also flag if reported per-group summaries look implausible (e.g., identical counts across groups, or one group's interval width being an order of magnitude different from the other) with no follow-up check.
Discriminator
Fine if the script demonstrably reads the template and applies/verifies each stated restriction (even when the filter turns out to be a no-op, with counts printed to prove it); a violation is when subset membership and output layout rest purely on the agent's assumption about the file it happened to find.
Consequence
The written file is compared to the expected one and fails — either because the statistics were computed over the wrong rows (wrong means/intervals) or because headers, column order, or value encoding do not match the required format.
id 80eb22ca92b9 · mined from da-code dacode-data-sa-029@s12
raw text (what the judge reads)
### Unverified input subset and unread output-format template
- **Applies when**: `task` -- the task specifies data restricted by explicit qualifiers (group, site, period, condition) and an output file whose layout is defined by a provided sample/template file.
- **Pattern**: The script reads a single convenient input file, hard-codes assumed column names and a single filter (or none), never checks that the rows actually correspond to the stated qualifiers, and writes the output with column names/ordering/quoting invented by the agent instead of copied from the supplied template; the answer then asserts success without any evidence that inputs or output layout were validated.
- **Detection procedure**:
  1. From the task, list every restriction on the rows to be used and note that a sample/template output file was provided.
  2. In the scripts, check whether each restriction is applied explicitly (filter on the qualifier column, or documented proof the file already contains only those rows) and whether the template file is ever opened/parsed to derive header names, column order, and value formatting.
  3. In the printed output/answer, look for a validation trace: row counts per group, a peek at raw rows, and a printed diff/comparison of the written file's header against the template's header.
  4. Flag if any restriction is only assumed, or if the output header/format is asserted rather than derived from the template; also flag if reported per-group summaries look implausible (e.g., identical counts across groups, or one group's interval width being an order of magnitude different from the other) with no follow-up check.
- **Discriminator**: Fine if the script demonstrably reads the template and applies/verifies each stated restriction (even when the filter turns out to be a no-op, with counts printed to prove it); a violation is when subset membership and output layout rest purely on the agent's assumption about the file it happened to find.
- **Consequence**: The written file is compared to the expected one and fails — either because the statistics were computed over the wrong rows (wrong means/intervals) or because headers, column order, or value encoding do not match the required format.
678Ambiguous dispersion statistic computed with the wrong degrees-of-freedom conventiontaskinfiagent-dabench
Applies when
task -- the task asks for a spread/variance-type statistic (standard deviation, variance, standard error, correlation-adjusted spread) on a column or subset, and the script uses a library default without stating the estimator convention.
Pattern
The agent calls a one-liner (e.g. numpy.std(x) vs pandas.Series.std(), or .std() on a tensor/array) and reports whatever the library default produces, silently choosing population (ddof=0) or sample (ddof=1) without checking which the task implies or cross-checking the alternative. Other means/counts in the same answer look fine, masking the error.
Detection procedure
  1. Read the task: note that the requested statistic has more than one standard definition (biased vs unbiased, ddof=0 vs 1) and that the task gives no explicit formula.
  2. Read the script: identify which library/function computes it and recall that function's default ddof; check whether the agent explicitly set ddof/sample= or documented the choice.
  3. Check whether the agent computed both conventions (or compared numpy vs pandas results) and justified the reported one; with n in the hundreds the two differ by roughly a factor sqrt(n/(n-1)) — small but enough to fail an exact-match grader.
  4. Compare the reported value's precision/rounding to the requested format and confirm the reported number is the convention that is conventional for a "sample of countries/rows" (usually ddof=1, the pandas default).
Discriminator
A real violation is an unexamined library default on an ambiguous estimator with no cross-check or justification; it is fine if the agent explicitly set and justified ddof (or the task defines the formula), or if the two conventions coincide (population fully enumerated and the task says so).
Consequence
The mean and other checks pass while the dispersion value is off by a small percentage, so the exact-value grader marks that field WRONG and the submission fails despite looking correct.
id 05fa3bd3f3b7 · mined from infiagent-dabench dabench-255@s12
raw text (what the judge reads)
### Ambiguous dispersion statistic computed with the wrong degrees-of-freedom convention
- **Applies when**: `task` -- the task asks for a spread/variance-type statistic (standard deviation, variance, standard error, correlation-adjusted spread) on a column or subset, and the script uses a library default without stating the estimator convention.
- **Pattern**: The agent calls a one-liner (e.g. `numpy.std(x)` vs `pandas.Series.std()`, or `.std()` on a tensor/array) and reports whatever the library default produces, silently choosing population (ddof=0) or sample (ddof=1) without checking which the task implies or cross-checking the alternative. Other means/counts in the same answer look fine, masking the error.
- **Detection procedure**:
  1. Read the task: note that the requested statistic has more than one standard definition (biased vs unbiased, ddof=0 vs 1) and that the task gives no explicit formula.
  2. Read the script: identify which library/function computes it and recall that function's default ddof; check whether the agent explicitly set `ddof`/`sample=` or documented the choice.
  3. Check whether the agent computed both conventions (or compared `numpy` vs `pandas` results) and justified the reported one; with n in the hundreds the two differ by roughly a factor sqrt(n/(n-1)) — small but enough to fail an exact-match grader.
  4. Compare the reported value's precision/rounding to the requested format and confirm the reported number is the convention that is conventional for a "sample of countries/rows" (usually ddof=1, the pandas default).
- **Discriminator**: A real violation is an unexamined library default on an ambiguous estimator with no cross-check or justification; it is fine if the agent explicitly set and justified `ddof` (or the task defines the formula), or if the two conventions coincide (population fully enumerated *and* the task says so).
- **Consequence**: The mean and other checks pass while the dispersion value is off by a small percentage, so the exact-value grader marks that field WRONG and the submission fails despite looking correct.
679Degenerate clustering accepted because an internal metric looked greattaskda-code
Applies when
task -- an unsupervised segmentation/clustering task where the agent selects the number of groups by an internal index (silhouette, Davies–Bouldin, inertia) on raw, heavy-tailed features.
Pattern
The agent builds skewed aggregate features, standardizes without any outlier treatment or log/rank transform, and picks the k with the best score — which is achieved by isolating a handful of extreme points, leaving essentially all records in one giant cluster. The near-perfect score is reported as evidence of "excellent separation" and no alternative solution or stability check is made.
Detection procedure
  1. Read the task for the deliverable (a per-record label file with a specified column naming/format) and note that a useful segmentation must actually partition the population.
  2. In the scripts, check whether skew is handled before distance-based clustering (winsorizing/clipping, log or quantile transform, outlier removal or robust scaling) and whether the selection rule is guarded against degenerate splits (e.g., minimum cluster share, elbow + interpretability, comparison of cluster profiles).
  3. In the answer, inspect the reported cluster sizes and the internal score: a silhouette near 0.85+ combined with one cluster holding >90–95% of records and clusters of size <1% is the signature of outlier-driven degeneracy.
  4. Confirm the output file's shape/columns match the requested schema (one row per record, Feature_i columns, label column) and that the feature columns are the form the task implies.
Discriminator
A genuinely imbalanced but valid solution shows moderate scores, clusters that remain populated and interpretable (each typically at least a few percent of records), and evidence the agent tested transforms/other k values and justified the imbalance; a violation is a high score produced solely by singleton-like clusters with no skew handling or size sanity check.
Consequence
The saved label file is effectively a single-cluster (plus outliers) assignment, so any grader comparison against a reasonable reference segmentation — on cluster count, size distribution, or agreement measures — fails.
id 6134cf9807ae · mined from da-code dacode-ml-cluster-019@s12
raw text (what the judge reads)
### Degenerate clustering accepted because an internal metric looked great
- **Applies when**: `task` -- an unsupervised segmentation/clustering task where the agent selects the number of groups by an internal index (silhouette, Davies–Bouldin, inertia) on raw, heavy-tailed features.
- **Pattern**: The agent builds skewed aggregate features, standardizes without any outlier treatment or log/rank transform, and picks the k with the best score — which is achieved by isolating a handful of extreme points, leaving essentially all records in one giant cluster. The near-perfect score is reported as evidence of "excellent separation" and no alternative solution or stability check is made.
- **Detection procedure**:
  1. Read the task for the deliverable (a per-record label file with a specified column naming/format) and note that a useful segmentation must actually partition the population.
  2. In the scripts, check whether skew is handled before distance-based clustering (winsorizing/clipping, log or quantile transform, outlier removal or robust scaling) and whether the selection rule is guarded against degenerate splits (e.g., minimum cluster share, elbow + interpretability, comparison of cluster profiles).
  3. In the answer, inspect the reported cluster sizes and the internal score: a silhouette near 0.85+ combined with one cluster holding >90–95% of records and clusters of size <1% is the signature of outlier-driven degeneracy.
  4. Confirm the output file's shape/columns match the requested schema (one row per record, `Feature_i` columns, label column) and that the feature columns are the form the task implies.
- **Discriminator**: A genuinely imbalanced but valid solution shows moderate scores, clusters that remain populated and interpretable (each typically at least a few percent of records), and evidence the agent tested transforms/other k values and justified the imbalance; a violation is a high score produced solely by singleton-like clusters with no skew handling or size sanity check.
- **Consequence**: The saved label file is effectively a single-cluster (plus outliers) assignment, so any grader comparison against a reasonable reference segmentation — on cluster count, size distribution, or agreement measures — fails.
680No held-out estimate of the competition metric before submittingtaskda-code
Applies when
task -- the task specifies an evaluation metric for predictions on an unlabeled test set, and the scripts train one or more models and write predictions directly.
Pattern
The attempt fits models on 100% of the labeled data, picks hyperparameters, ensemble members and blend weights by intuition (hard-coded round numbers), and never computes the stated metric on any validation split or via cross-validation; the only "verification" is structural (shape, ID match, row sums), so nothing detects miscalibrated/overconfident probabilities, unhandled missing values or encoding mismatches, or a variant that would score much worse than a trivial baseline.
Detection procedure
  1. Read the task and note the exact scoring function and prediction form (e.g., per-class probabilities, ranking, RMSE).
  2. Search the scripts for any train/validation split or cross-validation together with an explicit computation of that same metric on held-out labeled rows; also check whether a trivial baseline (class priors / mean prediction) was scored for comparison.
  3. Check whether model/ensemble choices (weights, depth, n_estimators, preprocessing choices) are justified by any such measured score or are simply asserted.
  4. Inspect the produced predictions for symptoms the missing check would have caught (e.g., near-0/1 probabilities for a metric that punishes confidence, degenerate constant columns, values outside the valid range).
Discriminator
A real violation is the absence of any held-out computation of the task's metric anywhere in the pipeline (structural checks and predict_proba printouts do not count); it is fine if the script reports CV/holdout scores in the task's metric — even if only for a single final model — and used them to compare alternatives, or if the metric is genuinely unmeasurable because no labels exist.
Consequence
The submission is format-valid but unoptimized and likely overconfident, so the leaderboard/grader metric falls below the required threshold (e.g., log loss far worse than a prior-only baseline), and the run is marked wrong with no internal evidence of why.
id a996f1050046 · mined from da-code dacode-ml-competition-005@s13
raw text (what the judge reads)
### No held-out estimate of the competition metric before submitting
- **Applies when**: `task` -- the task specifies an evaluation metric for predictions on an unlabeled test set, and the scripts train one or more models and write predictions directly.
- **Pattern**: The attempt fits models on 100% of the labeled data, picks hyperparameters, ensemble members and blend weights by intuition (hard-coded round numbers), and never computes the stated metric on any validation split or via cross-validation; the only "verification" is structural (shape, ID match, row sums), so nothing detects miscalibrated/overconfident probabilities, unhandled missing values or encoding mismatches, or a variant that would score much worse than a trivial baseline.
- **Detection procedure**:
  1. Read the task and note the exact scoring function and prediction form (e.g., per-class probabilities, ranking, RMSE).
  2. Search the scripts for any train/validation split or cross-validation together with an explicit computation of that same metric on held-out labeled rows; also check whether a trivial baseline (class priors / mean prediction) was scored for comparison.
  3. Check whether model/ensemble choices (weights, depth, n_estimators, preprocessing choices) are justified by any such measured score or are simply asserted.
  4. Inspect the produced predictions for symptoms the missing check would have caught (e.g., near-0/1 probabilities for a metric that punishes confidence, degenerate constant columns, values outside the valid range).
- **Discriminator**: A real violation is the absence of any held-out computation of the task's metric anywhere in the pipeline (structural checks and `predict_proba` printouts do not count); it is fine if the script reports CV/holdout scores in the task's metric — even if only for a single final model — and used them to compare alternatives, or if the metric is genuinely unmeasurable because no labels exist.
- **Consequence**: The submission is format-valid but unoptimized and likely overconfident, so the leaderboard/grader metric falls below the required threshold (e.g., log loss far worse than a prior-only baseline), and the run is marked wrong with no internal evidence of why.
681Deliverable artifact not verified against the requested path, schema, and row alignmenttaskda-code
Applies when
task -- the task asks for predictions/results to be written to a named output file with a specified column name, derived row-for-row from a given input file.
Pattern
The attempt focuses on modeling and treats the output file as an afterthought: it writes to an arbitrary directory or filename, uses a differently-spelled/extra column, drops or reorders rows (e.g., after dropna, filtering, deduplication, or a groupby/merge), writes an index column, or emits NaNs/implausible values — and never reopens the file to confirm it matches the input's row count and the requested schema. Often no script is preserved, so the artifact cannot be reproduced or audited.
Detection procedure
  1. From the task statement, extract the exact required output filename (and implied location, normally the working directory the grader inspects), the exact required column name(s), and the input file whose rows the output must correspond to.
  2. In the scripts, locate the single write call that produces the deliverable; check the literal path string, the column names in the written frame, index=False, and whether any row-count-changing operation (filtering, dropna, dedup, merge, groupby, sampling, train/test reshuffle) was applied to the test frame before prediction.
  3. Check for an explicit post-write validation: re-read the file and assert len(out) == len(test), columns equal the required names, no NaNs, and values in a plausible range (non-negative, correct units/rounding).
  4. Confirm the answer points to a file that exists at the required relative path with that exact name, not just some path where a file happens to be.
Discriminator
A real violation is a mismatch in path/name, column naming, row count, row order, or unusable values (NaN/negative/index column) — anything that makes the file unreadable as "one prediction per test row". A look-alike that is fine: an absolute path that resolves to the expected working-directory file with correct schema and row alignment, or predictions that are merely inaccurate while the file is well-formed.
Consequence
The grader reports the expected output file as WRONG/MISSING (file not found at the expected location, or column/row-count mismatch), scoring 0 regardless of model quality.
id 080db2abc928 · mined from da-code dacode-ml-regression-008@s13
raw text (what the judge reads)
### Deliverable artifact not verified against the requested path, schema, and row alignment
- **Applies when**: `task` -- the task asks for predictions/results to be written to a named output file with a specified column name, derived row-for-row from a given input file.
- **Pattern**: The attempt focuses on modeling and treats the output file as an afterthought: it writes to an arbitrary directory or filename, uses a differently-spelled/extra column, drops or reorders rows (e.g., after dropna, filtering, deduplication, or a groupby/merge), writes an index column, or emits NaNs/implausible values — and never reopens the file to confirm it matches the input's row count and the requested schema. Often no script is preserved, so the artifact cannot be reproduced or audited.
- **Detection procedure**:
  1. From the task statement, extract the exact required output filename (and implied location, normally the working directory the grader inspects), the exact required column name(s), and the input file whose rows the output must correspond to.
  2. In the scripts, locate the single write call that produces the deliverable; check the literal path string, the column names in the written frame, `index=False`, and whether any row-count-changing operation (filtering, dropna, dedup, merge, groupby, sampling, train/test reshuffle) was applied to the test frame before prediction.
  3. Check for an explicit post-write validation: re-read the file and assert `len(out) == len(test)`, columns equal the required names, no NaNs, and values in a plausible range (non-negative, correct units/rounding).
  4. Confirm the answer points to a file that exists at the required relative path with that exact name, not just some path where a file happens to be.
- **Discriminator**: A real violation is a mismatch in path/name, column naming, row count, row order, or unusable values (NaN/negative/index column) — anything that makes the file unreadable as "one prediction per test row". A look-alike that is fine: an absolute path that resolves to the expected working-directory file with correct schema and row alignment, or predictions that are merely inaccurate while the file is well-formed.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING (file not found at the expected location, or column/row-count mismatch), scoring 0 regardless of model quality.
682Statistical test run on the full pooled data with default options, without justifying subset, directionality, or distributional assumptionstaskda-code
Applies when
task -- the task asks for a p-value and an accept/reject decision from a hypothesis test, and the scripts feed entire raw tables into a single off-the-shelf test call.
Pattern
The agent treats "compute a p-value" as one library call: it pools every row of both sources (no filtering to the population/time window/competition tier the question is actually about), uses the default two-sided parametric test with default variance/independence settings, and never checks whether the outcome variable satisfies that test's assumptions (normality, symmetry, equal variance, sample comparability). A vanishingly small p-value from a huge n is then accepted at face value.
Detection procedure
  1. Read the task statement and README for any qualifier that restricts the comparison population (era, competition, category, completeness) or that implies a direction ("greater than", "more than") rather than mere difference; note the implied test specification.
  2. Read the script: check whether any row filtering is applied before the test, which test function is called, and which arguments (alternative=, equal_var=, paired) are left at defaults.
  3. Check whether the script inspects the outcome distribution (skew, discreteness, heavy tail, group sizes) or compares at least one alternative test before committing to the reported p-value.
  4. Compare the reported p-value's magnitude and the effective sample size to what the intended, filtered analysis would plausibly produce; an astronomically small p-value on hundreds of thousands of rows is a red flag that the wrong population was used.
Discriminator
A real violation is when the task or data description implies a narrower population or a directional/nonparametric test and the script silently uses all rows with default two-sided parametric settings. It is not a violation if the task genuinely asks about the whole corpus and the script explicitly documents why the pooled, two-sided, parametric choice is appropriate (e.g., checks distribution/variances and notes robustness of the conclusion under alternative tests).
Consequence
The reported p-value differs from the reference value by many orders of magnitude (and can flip or spuriously confirm the reject/fail-to-reject decision), so the exact-value check on the output file fails even though the file format is correct.
id e4a1283af5c6 · mined from da-code dacode-data-sa-001@s13
raw text (what the judge reads)
### Statistical test run on the full pooled data with default options, without justifying subset, directionality, or distributional assumptions
- **Applies when**: `task` -- the task asks for a p-value and an accept/reject decision from a hypothesis test, and the scripts feed entire raw tables into a single off-the-shelf test call.
- **Pattern**: The agent treats "compute a p-value" as one library call: it pools every row of both sources (no filtering to the population/time window/competition tier the question is actually about), uses the default two-sided parametric test with default variance/independence settings, and never checks whether the outcome variable satisfies that test's assumptions (normality, symmetry, equal variance, sample comparability). A vanishingly small p-value from a huge n is then accepted at face value.
- **Detection procedure**:
  1. Read the task statement and README for any qualifier that restricts the comparison population (era, competition, category, completeness) or that implies a direction ("greater than", "more than") rather than mere difference; note the implied test specification.
  2. Read the script: check whether any row filtering is applied before the test, which test function is called, and which arguments (alternative=, equal_var=, paired) are left at defaults.
  3. Check whether the script inspects the outcome distribution (skew, discreteness, heavy tail, group sizes) or compares at least one alternative test before committing to the reported p-value.
  4. Compare the reported p-value's magnitude and the effective sample size to what the intended, filtered analysis would plausibly produce; an astronomically small p-value on hundreds of thousands of rows is a red flag that the wrong population was used.
- **Discriminator**: A real violation is when the task or data description implies a narrower population or a directional/nonparametric test and the script silently uses all rows with default two-sided parametric settings. It is *not* a violation if the task genuinely asks about the whole corpus and the script explicitly documents why the pooled, two-sided, parametric choice is appropriate (e.g., checks distribution/variances and notes robustness of the conclusion under alternative tests).
- **Consequence**: The reported p-value differs from the reference value by many orders of magnitude (and can flip or spuriously confirm the reject/fail-to-reject decision), so the exact-value check on the output file fails even though the file format is correct.
683Template file assumed instead of inspected when "match the sample format" is requiredtaskda-code
Applies when
task -- The task says the output must follow the exact structure/formatting of a provided sample/template file (column names, order, sorting, rounding, units, row set).
Pattern
The scripts never load or print the sample file; column names, column order, row ordering, and numeric precision are hard-coded from memory or inferred from the input tables, and the final file is written with raw full-precision floats and a guessed row order, with no programmatic comparison against the template.
Detection procedure
  1. In the task text, note that a sample/template output file is named and identify every formatting property it could constrain (header text, column order, row order/keys, decimal rounding, thousands/currency formatting, dtypes).
  2. Search the scripts for any read/print of that sample file and for any assertion comparing the produced header/row-keys/dtypes/precision against it; note whether values are rounded or formatted before writing.
  3. Inspect the produced answer for tell-tale unformatted output (e.g., long floating-point tails like ...79.7999999999, inconsistent decimal places, ordering derived from an arbitrary sort or a hand-typed list).
  4. Flag if no step of the pipeline verified the output against the template.
Discriminator
A real violation is when the sample file exists and is never opened/compared, so the format is guessed; it is not a violation if the script reads the template (or explicitly reproduces its header/order/precision) and asserts the produced file matches, even if written by hand afterwards.
Consequence
The grader does an exact/near-exact comparison against the reference file and marks the output WRONG despite the underlying aggregation logic possibly being correct, because headers, row order, or float precision differ.
id df68caf11350 · mined from da-code dacode-dm-csv-011@s13
raw text (what the judge reads)
### Template file assumed instead of inspected when "match the sample format" is required
- **Applies when**: `task` -- The task says the output must follow the exact structure/formatting of a provided sample/template file (column names, order, sorting, rounding, units, row set).
- **Pattern**: The scripts never load or print the sample file; column names, column order, row ordering, and numeric precision are hard-coded from memory or inferred from the input tables, and the final file is written with raw full-precision floats and a guessed row order, with no programmatic comparison against the template.
- **Detection procedure**:
  1. In the task text, note that a sample/template output file is named and identify every formatting property it could constrain (header text, column order, row order/keys, decimal rounding, thousands/currency formatting, dtypes).
  2. Search the scripts for any read/print of that sample file and for any assertion comparing the produced header/row-keys/dtypes/precision against it; note whether values are rounded or formatted before writing.
  3. Inspect the produced answer for tell-tale unformatted output (e.g., long floating-point tails like `...79.7999999999`, inconsistent decimal places, ordering derived from an arbitrary sort or a hand-typed list).
  4. Flag if no step of the pipeline verified the output against the template.
- **Discriminator**: A real violation is when the sample file exists and is never opened/compared, so the format is guessed; it is *not* a violation if the script reads the template (or explicitly reproduces its header/order/precision) and asserts the produced file matches, even if written by hand afterwards.
- **Consequence**: The grader does an exact/near-exact comparison against the reference file and marks the output WRONG despite the underlying aggregation logic possibly being correct, because headers, row order, or float precision differ.
684Trivial/degenerate unsupervised solution exported as the internal preprocessed matrixtaskda-code
Applies when
task -- the task asks for a discovered structure (e.g., cluster labels) to be saved alongside the feature vector, and the script auto-selects the model size by one internal metric and writes out its own transformed matrix.
Pattern
The agent throws every column (including ID-like, near-constant, binary flag, and label-encoded nominal fields) into a distance-based algorithm, picks the number of groups by naively argmax-ing a single internal score — which almost always favours the smallest, near-trivial split — accepts a weak score without questioning it, and then writes the scaled/encoded intermediate matrix as the requested feature columns instead of the feature values the task/README describes. No sanity check that the groups are separated, balanced-in-a-plausible-way, or interpretable.
Detection procedure
  1. Read the task to see what the feature columns in the output are supposed to hold (values from the dataset's feature vector) and how many/what kind of groups the deliverable implies.
  2. In the script, check the feature-selection and encoding steps: are nominal categories forced into a single ordinal integer, are binary/constant/identifier columns included unfiltered, is scaling applied before distance computation and then also persisted to the output file?
  3. Check the model-selection block: is the chosen size simply the argmax over the smallest candidate range, with no elbow/stability/interpretability cross-check, and is the winning score value itself low (weak separation)?
  4. Read the reported answer: does it claim success while showing a boundary-of-range group count, a marginal score, and output columns whose values are standardized/encoded rather than the dataset's feature values?
Discriminator
A real violation is when the selected structure sits at the extreme of the searched range with a weak score and no corroborating check, and/or the saved feature columns are an internal transform that cannot be mapped back to the described features. It is fine if the agent justifies the group count with at least one independent check (elbow, stability across seeds, profile interpretation), documents why raw scale/encoding is preserved in the output, and the transform is a deliberate, stated part of the requested feature vector.
Consequence
The saved file's grouping is a near-trivial bisection (or its feature columns don't match the expected values), so any grader comparison against a reference segmentation — by cluster agreement, cluster count, or feature-column content — fails, and the answer's confident "completed successfully" summary hides the fact that no distinct groups were actually discovered.
id 7cee8932e319 · mined from da-code dacode-ml-cluster-014@s13
raw text (what the judge reads)
### Trivial/degenerate unsupervised solution exported as the internal preprocessed matrix
- **Applies when**: `task` -- the task asks for a discovered structure (e.g., cluster labels) to be saved alongside the feature vector, and the script auto-selects the model size by one internal metric and writes out its own transformed matrix.
- **Pattern**: The agent throws every column (including ID-like, near-constant, binary flag, and label-encoded nominal fields) into a distance-based algorithm, picks the number of groups by naively argmax-ing a single internal score — which almost always favours the smallest, near-trivial split — accepts a weak score without questioning it, and then writes the scaled/encoded intermediate matrix as the requested feature columns instead of the feature values the task/README describes. No sanity check that the groups are separated, balanced-in-a-plausible-way, or interpretable.
- **Detection procedure**:
  1. Read the task to see what the feature columns in the output are supposed to hold (values from the dataset's feature vector) and how many/what kind of groups the deliverable implies.
  2. In the script, check the feature-selection and encoding steps: are nominal categories forced into a single ordinal integer, are binary/constant/identifier columns included unfiltered, is scaling applied before distance computation and then *also* persisted to the output file?
  3. Check the model-selection block: is the chosen size simply the argmax over the smallest candidate range, with no elbow/stability/interpretability cross-check, and is the winning score value itself low (weak separation)?
  4. Read the reported answer: does it claim success while showing a boundary-of-range group count, a marginal score, and output columns whose values are standardized/encoded rather than the dataset's feature values?
- **Discriminator**: A real violation is when the selected structure sits at the extreme of the searched range with a weak score and no corroborating check, and/or the saved feature columns are an internal transform that cannot be mapped back to the described features. It is fine if the agent justifies the group count with at least one independent check (elbow, stability across seeds, profile interpretation), documents why raw scale/encoding is preserved in the output, and the transform is a deliberate, stated part of the requested feature vector.
- **Consequence**: The saved file's grouping is a near-trivial bisection (or its feature columns don't match the expected values), so any grader comparison against a reference segmentation — by cluster agreement, cluster count, or feature-column content — fails, and the answer's confident "completed successfully" summary hides the fact that no distinct groups were actually discovered.
685Deliverable written to the required file, complete and row-aligned with the test settaskda-code
Applies when
task -- the task asks for predictions/results to be saved to a named output file in a given template format, and the agent's scripts and final answer are available.
Pattern
The attempt returns the results only as inline text in the chat response (often truncated mid-record), or writes a file that is never verified to exist, to contain every required id, or to match the template's column names/order — so the graded artifact is missing or incomplete.
Detection procedure
  1. From the task/README, note the exact required output filename, its column names/order, and the expected number of rows (= number of test rows / template rows).
  2. Search the scripts for an explicit write of that exact filename (e.g. a to_csv(..., index=False)) and check that the written frame's key column comes from the test set in the template's order.
  3. Check the scripts (or the agent's report) for a post-write verification: reload the file and assert row count, column names, no NaNs, and key set equal to the template's keys.
  4. Inspect the final answer: if it is a pasted table rather than a confirmation of the saved file, or if the pasted content ends abruptly/has fewer rows than expected, flag it.
Discriminator
A genuine violation is missing the file write, writing a different path/format, or having an unverified/short/truncated row set. It is not a violation if the file is properly written and validated and the agent merely also shows a preview of the first rows for readability.
Consequence
The grader looks for the named artifact and finds it missing, or finds a partial/misformatted table, so the submission check fails outright regardless of model quality.
id 23701366bc1c · mined from da-code dacode-ml-competition-009@s13
raw text (what the judge reads)
### Deliverable written to the required file, complete and row-aligned with the test set
- **Applies when**: `task` -- the task asks for predictions/results to be saved to a named output file in a given template format, and the agent's scripts and final answer are available.
- **Pattern**: The attempt returns the results only as inline text in the chat response (often truncated mid-record), or writes a file that is never verified to exist, to contain every required id, or to match the template's column names/order — so the graded artifact is missing or incomplete.
- **Detection procedure**:
  1. From the task/README, note the exact required output filename, its column names/order, and the expected number of rows (= number of test rows / template rows).
  2. Search the scripts for an explicit write of that exact filename (e.g. a `to_csv(..., index=False)`) and check that the written frame's key column comes from the test set in the template's order.
  3. Check the scripts (or the agent's report) for a post-write verification: reload the file and assert row count, column names, no NaNs, and key set equal to the template's keys.
  4. Inspect the final answer: if it is a pasted table rather than a confirmation of the saved file, or if the pasted content ends abruptly/has fewer rows than expected, flag it.
- **Discriminator**: A genuine violation is missing the file write, writing a different path/format, or having an unverified/short/truncated row set. It is *not* a violation if the file is properly written and validated and the agent merely also shows a preview of the first rows for readability.
- **Consequence**: The grader looks for the named artifact and finds it missing, or finds a partial/misformatted table, so the submission check fails outright regardless of model quality.
686Dropping input rows when the deliverable requires one output row per input recordtaskda-code
Applies when
task -- the task asks for a per-record output artifact (labels, predictions, scores) written to a file, and the script does any row filtering, dropna, deduplication, or subsetting during preprocessing.
Pattern
The attempt removes records it judges unusable (too many missing values, unparseable strings, outliers) instead of imputing/handling them, so the saved file has fewer rows than the source data and the row-to-record alignment with the expected output is broken; the answer even advertises the removal as a virtue.
Detection procedure
  1. Read the task statement for the required output schema and note whether it implies coverage of every input record (a per-record label file with no ID column implies exact row correspondence).
  2. In the script, locate every operation that can change row count (dropna, boolean masks, drop_duplicates, query, groupby aggregation, merges) and check whether any occurs before the labels are assigned and saved.
  3. Compare the reported/actual output row count against the raw input row count; also confirm the column names and their order exactly match the requested naming convention.
  4. Check whether the script prints or asserts len(output) == len(raw_input) (or equivalent shape sanity check) before writing.
Discriminator
A real violation is silent or self-justified row loss in an artifact that must cover all records; it is fine if the task explicitly permits filtering, or if the output carries an identifier column that lets the grader align rows, or if rows were dropped only from an intermediate fitting step while all records still receive a label.
Consequence
The saved file has the wrong shape/row alignment, so row-wise comparison with the reference output fails outright and the file is scored as wrong/missing even if the clustering logic itself was reasonable.
id 2e51747a7d8d · mined from da-code dacode-ml-cluster-009@s13
raw text (what the judge reads)
### Dropping input rows when the deliverable requires one output row per input record
- **Applies when**: `task` -- the task asks for a per-record output artifact (labels, predictions, scores) written to a file, and the script does any row filtering, `dropna`, deduplication, or subsetting during preprocessing.
- **Pattern**: The attempt removes records it judges unusable (too many missing values, unparseable strings, outliers) instead of imputing/handling them, so the saved file has fewer rows than the source data and the row-to-record alignment with the expected output is broken; the answer even advertises the removal as a virtue.
- **Detection procedure**:
  1. Read the task statement for the required output schema and note whether it implies coverage of every input record (a per-record label file with no ID column implies exact row correspondence).
  2. In the script, locate every operation that can change row count (`dropna`, boolean masks, `drop_duplicates`, `query`, `groupby` aggregation, merges) and check whether any occurs before the labels are assigned and saved.
  3. Compare the reported/actual output row count against the raw input row count; also confirm the column names and their order exactly match the requested naming convention.
  4. Check whether the script prints or asserts `len(output) == len(raw_input)` (or equivalent shape sanity check) before writing.
- **Discriminator**: A real violation is silent or self-justified row loss in an artifact that must cover all records; it is fine if the task explicitly permits filtering, or if the output carries an identifier column that lets the grader align rows, or if rows were dropped only from an intermediate fitting step while all records still receive a label.
- **Consequence**: The saved file has the wrong shape/row alignment, so row-wise comparison with the reference output fails outright and the file is scored as wrong/missing even if the clustering logic itself was reasonable.
687Accepting a degenerate, outlier-driven cluster solution without sanity checkstaskda-code
Applies when
task -- the task asks for unsupervised segmentation of entities built from aggregated, heavily right‑skewed transactional/count features, and the script picks the cluster count by an internal score (silhouette/inertia) alone.
Pattern
The script standardizes raw skewed aggregates (no log/rank transform, no winsorizing or outlier removal) and then selects k by maximizing silhouette. Because a handful of extreme entities dominate distance space, the "best" model is degenerate: one or two microscopic clusters (a few members) plus one or two huge buckets. The agent reports the score and the lopsided sizes as a success instead of treating them as a red flag, and no alternative feature scaling/algorithm is compared.
Detection procedure
  1. Read the task to confirm the deliverable is a meaningful segmentation (a label per entity), not merely a file with the right headers.
  2. In the scripts, check whether skewness/outliers in the engineered features are addressed before scaling (log1p, quantile/robust scaling, clipping, or explicit outlier flagging) and whether model selection considers anything beyond a single internal index.
  3. In the printed diagnostics/answer, inspect the per-cluster counts: flag if any cluster holds a negligible fraction of entities (e.g., <1%) while another holds most of them, or if the reported "optimal k" was chosen despite the script's own notes that lower k was "extremely imbalanced".
  4. Confirm the output file's shape/columns and row count match the requested format and the number of entities, and that features saved match those actually clustered.
Discriminator
A genuinely valid solution may still have one small cluster, but only after skew/outliers were explicitly handled and the choice was justified against balanced alternatives; the violation is picking a solution whose tiny clusters are pure untreated outliers and whose score improvement comes solely from isolating them, with no transform or comparison attempted.
Consequence
The saved labels are effectively "bulk vs. a few outliers," so the grader's check of the clustering file (cluster structure/quality/size distribution against a reasonable reference segmentation) fails even though the file format looks correct.
id 7fcde80194b7 · mined from da-code dacode-ml-cluster-016@s13
raw text (what the judge reads)
### Accepting a degenerate, outlier-driven cluster solution without sanity checks
- **Applies when**: `task` -- the task asks for unsupervised segmentation of entities built from aggregated, heavily right‑skewed transactional/count features, and the script picks the cluster count by an internal score (silhouette/inertia) alone.
- **Pattern**: The script standardizes raw skewed aggregates (no log/rank transform, no winsorizing or outlier removal) and then selects k by maximizing silhouette. Because a handful of extreme entities dominate distance space, the "best" model is degenerate: one or two microscopic clusters (a few members) plus one or two huge buckets. The agent reports the score and the lopsided sizes as a success instead of treating them as a red flag, and no alternative feature scaling/algorithm is compared.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is a meaningful segmentation (a label per entity), not merely a file with the right headers.
  2. In the scripts, check whether skewness/outliers in the engineered features are addressed before scaling (log1p, quantile/robust scaling, clipping, or explicit outlier flagging) and whether model selection considers anything beyond a single internal index.
  3. In the printed diagnostics/answer, inspect the per-cluster counts: flag if any cluster holds a negligible fraction of entities (e.g., <1%) while another holds most of them, or if the reported "optimal k" was chosen despite the script's own notes that lower k was "extremely imbalanced".
  4. Confirm the output file's shape/columns and row count match the requested format and the number of entities, and that features saved match those actually clustered.
- **Discriminator**: A genuinely valid solution may still have one small cluster, but only after skew/outliers were explicitly handled and the choice was justified against balanced alternatives; the violation is picking a solution whose tiny clusters are pure untreated outliers and whose score improvement comes solely from isolating them, with no transform or comparison attempted.
- **Consequence**: The saved labels are effectively "bulk vs. a few outliers," so the grader's check of the clustering file (cluster structure/quality/size distribution against a reasonable reference segmentation) fails even though the file format looks correct.
688Plotted/analyzed values hardcoded from memory instead of read from the provided datataskda-code
Applies when
task -- The task points to a provided dataset (a file or directory) and asks for a chart, statistic, or model output derived from a specific series/subset of it.
Pattern
The script loads only the config/spec file (or nothing) and then embeds literal arrays/numbers typed from prior knowledge or assumption, never reading, filtering, or aggregating the actual data file; the answer then reports those invented values as if computed.
Detection procedure
  1. From the task, identify which input file(s) hold the raw data to be analyzed and what granularity/time range is implied.
  2. Scan the script for an actual data-loading call (read_csv/read_excel/load/etc.) on that raw file and a traceable chain from the loaded frame to the plotted/reported values.
  3. Flag if the values used are literal constants, or if the only file opened is a config/spec, or if the range/granularity is asserted in a comment ("data from 20XX to 20YY") rather than derived from the data.
  4. Check the answer for numbers or date ranges that appear nowhere as computed output (no groupby/aggregation/shape print backing them).
Discriminator
A real violation is data content invented or assumed; it is fine to hardcode formatting constants taken from the provided spec (title, colors, figsize, tick labels), or to hardcode a small lookup/mapping that is then applied to genuinely loaded data.
Consequence
Saved arrays/plot-data artifacts (e.g., exported series or serialized plot values) mismatch the ground-truth values, axis lengths, or date range, so every value/shape check fails even if the chart's cosmetic format is correct.
id 735f221c81ca · mined from da-code dacode-plot-line-015@s13
raw text (what the judge reads)
### Plotted/analyzed values hardcoded from memory instead of read from the provided data
- **Applies when**: `task` -- The task points to a provided dataset (a file or directory) and asks for a chart, statistic, or model output derived from a specific series/subset of it.
- **Pattern**: The script loads only the config/spec file (or nothing) and then embeds literal arrays/numbers typed from prior knowledge or assumption, never reading, filtering, or aggregating the actual data file; the answer then reports those invented values as if computed.
- **Detection procedure**:
  1. From the task, identify which input file(s) hold the raw data to be analyzed and what granularity/time range is implied.
  2. Scan the script for an actual data-loading call (`read_csv`/`read_excel`/`load`/etc.) on that raw file and a traceable chain from the loaded frame to the plotted/reported values.
  3. Flag if the values used are literal constants, or if the only file opened is a config/spec, or if the range/granularity is asserted in a comment ("data from 20XX to 20YY") rather than derived from the data.
  4. Check the answer for numbers or date ranges that appear nowhere as computed output (no groupby/aggregation/shape print backing them).
- **Discriminator**: A real violation is data content invented or assumed; it is fine to hardcode *formatting* constants taken from the provided spec (title, colors, figsize, tick labels), or to hardcode a small lookup/mapping that is then applied to genuinely loaded data.
- **Consequence**: Saved arrays/plot-data artifacts (e.g., exported series or serialized plot values) mismatch the ground-truth values, axis lengths, or date range, so every value/shape check fails even if the chart's cosmetic format is correct.
689Monte-Carlo p-value reported as an exact boundary value (0 or 1) without resolution checktaskda-code
Applies when
task -- the requested quantity is a tail probability/p-value estimated by resampling (bootstrap/permutation/simulation), and the script counts how many replicates exceed the observed statistic.
Pattern
The attempt runs a small or unvalidated number of replicates, finds zero exceedances, and reports the degenerate boundary value (e.g. 0.0) as the final answer, without checking whether that value is distinguishable from "smaller than my resolution" or whether the resampling scheme matches the stated null (e.g. shifting/centering the groups before resampling, resampling each group at its own size).
Detection procedure
  1. From the task/README, note the exact null hypothesis and the prescribed resampling recipe (what is shifted/held fixed, what statistic is recomputed, one- vs two-sided tail).
  2. In the script, verify the null is imposed as described (data transformed so the null holds) and that the tail count uses the correct comparison and direction; note the replicate count N and hence the minimum resolvable p-value 1/N.
  3. Inspect the reported number: if it equals 0 or 1 exactly, check whether the script logs the exceedance count and whether N is large enough (and whether any sanity check on the bootstrap distribution — its center, spread, observed-statistic placement — was performed).
  4. Confirm the reported value is the requested probability (not a count, a z-score, or a one-sided piece of a two-sided test) and matches the required file/column/precision format.
Discriminator
A genuinely tiny p-value backed by a large replicate count and an explicit acknowledgement of the resolution bound (e.g. reported as <1/N or with a large N consistent with the expected value) is fine; the violation is a boundary value produced by an unvalidated/small N, or by a resampling scheme that does not enforce the stated null (so the null distribution is mis-centered and no replicate can ever exceed the observation).
Consequence
The graded value differs from the reference p-value (which is small but nonzero), so the result file fails the numeric tolerance check and the task scores 0.
id 108945c4a896 · mined from da-code dacode-data-sa-028@s13
raw text (what the judge reads)
### Monte-Carlo p-value reported as an exact boundary value (0 or 1) without resolution check
- **Applies when**: `task` -- the requested quantity is a tail probability/p-value estimated by resampling (bootstrap/permutation/simulation), and the script counts how many replicates exceed the observed statistic.
- **Pattern**: The attempt runs a small or unvalidated number of replicates, finds zero exceedances, and reports the degenerate boundary value (e.g. `0.0`) as the final answer, without checking whether that value is distinguishable from "smaller than my resolution" or whether the resampling scheme matches the stated null (e.g. shifting/centering the groups before resampling, resampling each group at its own size).
- **Detection procedure**:
  1. From the task/README, note the exact null hypothesis and the prescribed resampling recipe (what is shifted/held fixed, what statistic is recomputed, one- vs two-sided tail).
  2. In the script, verify the null is imposed as described (data transformed so the null holds) and that the tail count uses the correct comparison and direction; note the replicate count `N` and hence the minimum resolvable p-value `1/N`.
  3. Inspect the reported number: if it equals 0 or 1 exactly, check whether the script logs the exceedance count and whether `N` is large enough (and whether any sanity check on the bootstrap distribution — its center, spread, observed-statistic placement — was performed).
  4. Confirm the reported value is the requested probability (not a count, a z-score, or a one-sided piece of a two-sided test) and matches the required file/column/precision format.
- **Discriminator**: A genuinely tiny p-value backed by a large replicate count *and* an explicit acknowledgement of the resolution bound (e.g. reported as `<1/N` or with a large `N` consistent with the expected value) is fine; the violation is a boundary value produced by an unvalidated/small `N`, or by a resampling scheme that does not enforce the stated null (so the null distribution is mis-centered and no replicate can ever exceed the observation).
- **Consequence**: The graded value differs from the reference p-value (which is small but nonzero), so the result file fails the numeric tolerance check and the task scores 0.
690Fabricating the spec instead of reading the provided instruction/config files (and their required outputs)taskda-code
Applies when
task -- the task says to follow instructions in an auxiliary file (readme/tips/notes) and/or format output per a config file (yaml/json/schema), and to emit named artifact files.
Pattern
The script never loads (or only conditionally loads, with a fallback that silently wins) the referenced instruction/config file; the agent invents its own definitions, styling, and output set, sometimes even writing the "config" itself, and produces only the single most obvious artifact while omitting the other required files.
Detection procedure
  1. From the task text, list every referenced input file (instructions, config/spec) and every expected output artifact name/extension.
  2. Grep the scripts for each referenced input path: is it opened and are its values actually used, or wrapped in if os.path.exists(...)/.get(key, default) chains that make hard-coded defaults sufficient? Confirm the agent did not create the config file itself.
  3. Grep the scripts for a write of each expected output artifact; flag any missing artifact and any output written outside the required directory/filename.
  4. Cross-check the answer's claimed choices (aggregation definition, series set, labels, colors, ranges) against literal quotes from the instruction file; if the answer never quotes or echoes the file's contents, treat the choices as invented.
Discriminator
Fine if the scripts demonstrably parse the provided files and the printed/echoed contents match the parameters used, with defaults only for keys genuinely absent; a violation is when the instruction/config file is missing from the working directory, unread, self-authored, or its content is never evidenced — or when required companion artifacts (serialized plot data, arrays) are simply not produced.
Consequence
Graders comparing each expected artifact fail on all of them — missing files score zero, and the one produced figure/array mismatches the specified aggregation, series, units, or styling even though the agent reports "task completed successfully."
id bbcb7de57212 · mined from da-code dacode-plot-line-006@s13
raw text (what the judge reads)
### Fabricating the spec instead of reading the provided instruction/config files (and their required outputs)
- **Applies when**: `task` -- the task says to follow instructions in an auxiliary file (readme/tips/notes) and/or format output per a config file (yaml/json/schema), and to emit named artifact files.
- **Pattern**: The script never loads (or only conditionally loads, with a fallback that silently wins) the referenced instruction/config file; the agent invents its own definitions, styling, and output set, sometimes even writing the "config" itself, and produces only the single most obvious artifact while omitting the other required files.
- **Detection procedure**:
  1. From the task text, list every referenced input file (instructions, config/spec) and every expected output artifact name/extension.
  2. Grep the scripts for each referenced input path: is it opened and are its values actually used, or wrapped in `if os.path.exists(...)`/`.get(key, default)` chains that make hard-coded defaults sufficient? Confirm the agent did not create the config file itself.
  3. Grep the scripts for a write of each expected output artifact; flag any missing artifact and any output written outside the required directory/filename.
  4. Cross-check the answer's claimed choices (aggregation definition, series set, labels, colors, ranges) against literal quotes from the instruction file; if the answer never quotes or echoes the file's contents, treat the choices as invented.
- **Discriminator**: Fine if the scripts demonstrably parse the provided files and the printed/echoed contents match the parameters used, with defaults only for keys genuinely absent; a violation is when the instruction/config file is missing from the working directory, unread, self-authored, or its content is never evidenced — or when required companion artifacts (serialized plot data, arrays) are simply not produced.
- **Consequence**: Graders comparing each expected artifact fail on all of them — missing files score zero, and the one produced figure/array mismatches the specified aggregation, series, units, or styling even though the agent reports "task completed successfully."
691Unverified conformance to an externally specified category/binning definition (and no reproducible script)taskda-code
Applies when
task -- The task points to an auxiliary specification (README/markdown/config) that defines how a variable must be bucketed, filtered, ordered, or formatted, and the deliverable is a plot/table/array derived from those buckets.
Pattern
The agent produces the deliverable using the raw categories present in the data (or its own invented bins) without ever opening the spec file, keeps no runnable script showing the mapping, and declares success by restating the requested title/labels plus counts — including rows with missing/non-answer values, so the totals trivially equal the full response count.
Detection procedure
  1. Read the task for any referenced spec document or stated constraints (bin edges, group names, ordering, output files) and list them explicitly.
  2. Inspect the agent's scripts: confirm the spec file is actually read/quoted and that a visible mapping from raw values to spec-defined groups exists, plus explicit handling of blanks/unparseable values; if no script was saved at all, the attempt is unreproducible and fails immediately.
  3. Compare the categories in the answer with the spec-defined ones (count, boundaries, labels, order) and check whether the sum of bucket counts equals total rows minus non-responses rather than the raw dataset size.
  4. Verify every requested artifact (image plus any expected data/serialization outputs) is created by the script, not just claimed in prose.
Discriminator
A real violation is when the groups/labels/counts cannot be traced to the spec (or the spec was never opened) or when non-answers are silently counted as a group; it is fine if the spec's grouping happens to coincide with the data's native categories and the script demonstrates it read/applied the spec and excluded missing values.
Consequence
The saved figure and derived data arrays encode the wrong bins/counts, so all file-level comparisons against the expected outputs fail even though the answer text looks well formatted.
id 005db5d4b4d1 · mined from da-code dacode-plot-bar-005@s13
raw text (what the judge reads)
### Unverified conformance to an externally specified category/binning definition (and no reproducible script)
- **Applies when**: `task` -- The task points to an auxiliary specification (README/markdown/config) that defines how a variable must be bucketed, filtered, ordered, or formatted, and the deliverable is a plot/table/array derived from those buckets.
- **Pattern**: The agent produces the deliverable using the raw categories present in the data (or its own invented bins) without ever opening the spec file, keeps no runnable script showing the mapping, and declares success by restating the requested title/labels plus counts — including rows with missing/non-answer values, so the totals trivially equal the full response count.
- **Detection procedure**:
  1. Read the task for any referenced spec document or stated constraints (bin edges, group names, ordering, output files) and list them explicitly.
  2. Inspect the agent's scripts: confirm the spec file is actually read/quoted and that a visible mapping from raw values to spec-defined groups exists, plus explicit handling of blanks/unparseable values; if no script was saved at all, the attempt is unreproducible and fails immediately.
  3. Compare the categories in the answer with the spec-defined ones (count, boundaries, labels, order) and check whether the sum of bucket counts equals total rows minus non-responses rather than the raw dataset size.
  4. Verify every requested artifact (image plus any expected data/serialization outputs) is created by the script, not just claimed in prose.
- **Discriminator**: A real violation is when the groups/labels/counts cannot be traced to the spec (or the spec was never opened) or when non-answers are silently counted as a group; it is fine if the spec's grouping happens to coincide with the data's native categories *and* the script demonstrates it read/applied the spec and excluded missing values.
- **Consequence**: The saved figure and derived data arrays encode the wrong bins/counts, so all file-level comparisons against the expected outputs fail even though the answer text looks well formatted.
692Requested answer artifact never written (results only printed to stdout)taskda-code
Applies when
task -- the task asks for the answer in a specific serialized shape (e.g., a JSON object with given keys, values as lists) and/or the grading expects a result file produced by the scripts.
Pattern
The scripts do all the computation but only print intermediate diagnostics and the final value; no code writes the answer to the expected output file, and the chat-only answer ignores the literal template (e.g., scalar instead of list-valued keys, keys renamed/reordered, extra rounding or units changed).
Detection procedure
  1. Read the task statement and note the exact required deliverable: file name/location (if implied by "provide the answer" plus expected artifacts) and the exact key names and value types shown in the template.
  2. Grep the scripts for any write/serialize call (json.dump, to_json, to_csv, open(..., 'w')) targeting that deliverable; if the only outputs are print statements, the deliverable is missing.
  3. Compare the submitted answer's structure key-by-key with the template: same keys, same nesting, values in the shown container type (list vs scalar), numbers in the shown precision/units.
  4. Flag if either the artifact is absent or the structure deviates from the template.
Discriminator
A real violation is no persisted answer file or a structural mismatch with the given template; it is not a violation if the script writes the required file with the required keys and merely also prints extra diagnostics, or if formatting differs only in whitespace/JSON key order while types and names match.
Consequence
The grader reports the expected result file as WRONG/MISSING (0 checks passed) even when the computed statistic itself may be right, because there is nothing to grade or the parsed structure doesn't match.
id 70a1c9da7d9f · mined from da-code dacode-di-text-002@s13
raw text (what the judge reads)
### Requested answer artifact never written (results only printed to stdout)
- **Applies when**: `task` -- the task asks for the answer in a specific serialized shape (e.g., a JSON object with given keys, values as lists) and/or the grading expects a result file produced by the scripts.
- **Pattern**: The scripts do all the computation but only `print` intermediate diagnostics and the final value; no code writes the answer to the expected output file, and the chat-only answer ignores the literal template (e.g., scalar instead of list-valued keys, keys renamed/reordered, extra rounding or units changed).
- **Detection procedure**:
  1. Read the task statement and note the exact required deliverable: file name/location (if implied by "provide the answer" plus expected artifacts) and the exact key names and value types shown in the template.
  2. Grep the scripts for any write/serialize call (`json.dump`, `to_json`, `to_csv`, `open(..., 'w')`) targeting that deliverable; if the only outputs are `print` statements, the deliverable is missing.
  3. Compare the submitted answer's structure key-by-key with the template: same keys, same nesting, values in the shown container type (list vs scalar), numbers in the shown precision/units.
  4. Flag if either the artifact is absent or the structure deviates from the template.
- **Discriminator**: A real violation is no persisted answer file or a structural mismatch with the given template; it is *not* a violation if the script writes the required file with the required keys and merely also prints extra diagnostics, or if formatting differs only in whitespace/JSON key order while types and names match.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING (0 checks passed) even when the computed statistic itself may be right, because there is nothing to grade or the parsed structure doesn't match.
693Silently analyzing only one data file/split when the task refers to the whole datasettaskinfiagent-dabench
Applies when
task -- the task asks for a descriptive statistic (or model input) over "the dataset" and the working directory may contain several files or split shards (train/test/val, monthly chunks, part files).
Pattern
The script hard-codes a single file path (typically the _train file, or the first one seen) and computes the requested statistic on that partial subset, never enumerating the available files or reporting the total row count against the dataset's documented size.
Detection procedure
  1. Read the task and note whether it scopes the computation to a specific split/file; if it just says "the dataset", the whole dataset is intended.
  2. Read the script's data-loading lines: does it list/glob the data directory, or does it load exactly one hard-coded path? Is there any concatenation of remaining files?
  3. Check whether the script prints and the agent reasons about row counts / non-null counts and compares them to the expected dataset size; absence of any such check is a red flag.
  4. If only a subset is loaded and no justification appears in the task, treat the reported statistic as computed on the wrong population.
Discriminator
A real violation is loading one of several available files while the task scope is the full dataset (or the loaded subset's size is never validated). It is fine if the task explicitly names that file/split, or if the script verified that this file is the only data file and its shape matches the dataset description.
Consequence
The statistic is computed on a biased subsample, so the numeric answer deviates from the expected value (only the coarse qualitative label may still match), and the grader marks the numeric check wrong.
id 80a1d54e21e1 · mined from infiagent-dabench dabench-359@s13
raw text (what the judge reads)
### Silently analyzing only one data file/split when the task refers to the whole dataset
- **Applies when**: `task` -- the task asks for a descriptive statistic (or model input) over "the dataset" and the working directory may contain several files or split shards (train/test/val, monthly chunks, part files).
- **Pattern**: The script hard-codes a single file path (typically the `*_train*` file, or the first one seen) and computes the requested statistic on that partial subset, never enumerating the available files or reporting the total row count against the dataset's documented size.
- **Detection procedure**:
  1. Read the task and note whether it scopes the computation to a specific split/file; if it just says "the dataset", the whole dataset is intended.
  2. Read the script's data-loading lines: does it list/glob the data directory, or does it load exactly one hard-coded path? Is there any concatenation of remaining files?
  3. Check whether the script prints and the agent reasons about row counts / non-null counts and compares them to the expected dataset size; absence of any such check is a red flag.
  4. If only a subset is loaded and no justification appears in the task, treat the reported statistic as computed on the wrong population.
- **Discriminator**: A real violation is loading one of several available files while the task scope is the full dataset (or the loaded subset's size is never validated). It is fine if the task explicitly names that file/split, or if the script verified that this file is the only data file and its shape matches the dataset description.
- **Consequence**: The statistic is computed on a biased subsample, so the numeric answer deviates from the expected value (only the coarse qualitative label may still match), and the grader marks the numeric check wrong.
694Unverified data conventions and output-format assumptions in a derived-quantity computationtaskda-code
Applies when
task -- the task asks for a derived series/statistic to be written to a file "in the required format", and the script hard-codes both the formula convention (units, scaling, whether a baseline is subtracted) and the output column names/ordering without inspecting the raw inputs or any format spec.
Pattern
The attempt reads the raw file and immediately applies one arithmetic convention (e.g., treating values as fractions rather than percents, or subtracting/not subtracting a base of 1), invents column names and row coverage for the output, and then "verifies" only by re-running its own identical formula. Missing values, unexpected extra/renamed columns, or a documented output template are never checked, and the resulting magnitudes are never compared against an external plausibility reference.
Detection procedure
  1. Read the task/README for any stated output schema (column names, ordering, units, rounding, index/date format) and any stated definition of the requested quantity; note what is specified vs. what the agent had to guess.
  2. In the script, check whether the raw input is actually inspected before computation — dtypes, value ranges/scale, NaN counts, and the exact set of columns used — and whether NaNs are handled explicitly rather than silently propagating through cumulative/aggregating operations.
  3. Check whether the verification step is independent (recomputing by a different route, comparing magnitudes to a known external benchmark or to per-column aggregates) or merely re-executes the same expression on the same data.
  4. Compare the written file's columns/units/row count to the task's requested format; flag if names, scaling, or coverage are the agent's own invention.
Discriminator
Fine if the script demonstrably inspected the raw values (scale, NaNs, columns) and either the task explicitly fixed the schema/definition or the agent justified its choice with a check that would fail under the alternative convention; a violation is a hard-coded convention plus a self-referential "verification" that cannot detect a wrong scale, wrong baseline, NaN contamination, or mismatched column names.
Consequence
The file exists and looks internally consistent, but the grader's element-wise comparison to the reference fails everywhere (values off by a constant offset or 100×, or columns/rows misaligned), scoring 0 despite a plausible-sounding narrative summary.
id 48ff7141e79d · mined from da-code dacode-dm-csv-050@s13
raw text (what the judge reads)
### Unverified data conventions and output-format assumptions in a derived-quantity computation
- **Applies when**: `task` -- the task asks for a derived series/statistic to be written to a file "in the required format", and the script hard-codes both the formula convention (units, scaling, whether a baseline is subtracted) and the output column names/ordering without inspecting the raw inputs or any format spec.
- **Pattern**: The attempt reads the raw file and immediately applies one arithmetic convention (e.g., treating values as fractions rather than percents, or subtracting/not subtracting a base of 1), invents column names and row coverage for the output, and then "verifies" only by re-running its own identical formula. Missing values, unexpected extra/renamed columns, or a documented output template are never checked, and the resulting magnitudes are never compared against an external plausibility reference.
- **Detection procedure**:
  1. Read the task/README for any stated output schema (column names, ordering, units, rounding, index/date format) and any stated definition of the requested quantity; note what is specified vs. what the agent had to guess.
  2. In the script, check whether the raw input is actually inspected before computation — dtypes, value ranges/scale, NaN counts, and the exact set of columns used — and whether NaNs are handled explicitly rather than silently propagating through cumulative/aggregating operations.
  3. Check whether the verification step is independent (recomputing by a different route, comparing magnitudes to a known external benchmark or to per-column aggregates) or merely re-executes the same expression on the same data.
  4. Compare the written file's columns/units/row count to the task's requested format; flag if names, scaling, or coverage are the agent's own invention.
- **Discriminator**: Fine if the script demonstrably inspected the raw values (scale, NaNs, columns) and either the task explicitly fixed the schema/definition or the agent justified its choice with a check that would fail under the alternative convention; a violation is a hard-coded convention plus a self-referential "verification" that cannot detect a wrong scale, wrong baseline, NaN contamination, or mismatched column names.
- **Consequence**: The file exists and looks internally consistent, but the grader's element-wise comparison to the reference fails everywhere (values off by a constant offset or 100×, or columns/rows misaligned), scoring 0 despite a plausible-sounding narrative summary.
695Applying a statistical test to a raw column without validating the test's sample-size limits or screening sentinel/degenerate valuestaskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test (e.g., a normality test) plus descriptive statistics on a single column, and the script feeds the column straight into the test function after only a dropna().
Pattern
The attempt treats the column as clean and the test as unconditionally valid: it never checks the number of observations against the test's documented validity range, never inspects the tails for sentinel/placeholder codes (0, -1, 9999, boundary values from incomplete windows/records) that could dominate skew and kurtosis, and never cross-checks the test's conclusion against complementary evidence (histogram/QQ shape, quantiles, an alternative normality test). The reported decision therefore rests on a single p-value that the library itself may flag as unreliable.
Detection procedure
  1. Read the task for the exact test and the population it should be computed on; note any implied filtering (valid/complete records only).
  2. In the script, locate the test call and check whether the sample size and value range are printed and compared to the test's assumptions (e.g., scipy's Shapiro-Wilk p-value is unreliable for N above a few thousand and the function emits a warning) and whether extreme-value rows are examined rather than silently included.
  3. Check whether any independent sanity check of the distribution shape is performed (quantiles/histogram, alternative test, or the same test on the cleaned subset) and whether the skew/kurtosis magnitudes are reconciled with the test verdict.
  4. Inspect the answer: if a large |skew|/|kurtosis| plus a vanishing p-value is reported with no discussion of outlier codes or N limits, flag it.
Discriminator
Fine if the script prints N, the value range/tail counts, and either shows the sample is within the test's valid range or reruns the test on the properly filtered subset and gets the same verdict; a real violation is a single unguarded test call on the raw column with no size/sentinel/shape validation.
Consequence
The normality verdict flips relative to ground truth (and skew/kurtosis are inflated by contaminating rows), so the categorical field—and typically the numeric fields too—are marked wrong, scoring 0.
id bd4faf767745 · mined from infiagent-dabench dabench-298@s13
raw text (what the judge reads)
### Applying a statistical test to a raw column without validating the test's sample-size limits or screening sentinel/degenerate values
- **Applies when**: `task` -- the task asks for a hypothesis test (e.g., a normality test) plus descriptive statistics on a single column, and the script feeds the column straight into the test function after only a `dropna()`.
- **Pattern**: The attempt treats the column as clean and the test as unconditionally valid: it never checks the number of observations against the test's documented validity range, never inspects the tails for sentinel/placeholder codes (0, -1, 9999, boundary values from incomplete windows/records) that could dominate skew and kurtosis, and never cross-checks the test's conclusion against complementary evidence (histogram/QQ shape, quantiles, an alternative normality test). The reported decision therefore rests on a single p-value that the library itself may flag as unreliable.
- **Detection procedure**:
  1. Read the task for the exact test and the population it should be computed on; note any implied filtering (valid/complete records only).
  2. In the script, locate the test call and check whether the sample size and value range are printed *and* compared to the test's assumptions (e.g., scipy's Shapiro-Wilk p-value is unreliable for N above a few thousand and the function emits a warning) and whether extreme-value rows are examined rather than silently included.
  3. Check whether any independent sanity check of the distribution shape is performed (quantiles/histogram, alternative test, or the same test on the cleaned subset) and whether the skew/kurtosis magnitudes are reconciled with the test verdict.
  4. Inspect the answer: if a large |skew|/|kurtosis| plus a vanishing p-value is reported with no discussion of outlier codes or N limits, flag it.
- **Discriminator**: Fine if the script prints N, the value range/tail counts, and either shows the sample is within the test's valid range or reruns the test on the properly filtered subset and gets the same verdict; a real violation is a single unguarded test call on the raw column with no size/sentinel/shape validation.
- **Consequence**: The normality verdict flips relative to ground truth (and skew/kurtosis are inflated by contaminating rows), so the categorical field—and typically the numeric fields too—are marked wrong, scoring 0.
696Accepting a single default baseline without measuring held-out predictive quality against a plausible bartaskda-code
Applies when
task -- the task asks for predictions on an unlabeled evaluation file that will be graded on accuracy/score, and the scripts fit one off-the-shelf pipeline (default vectorizer/model hyperparameters) and immediately export predictions.
Pattern
The attempt treats "the file has the right column and ran without error" as success: it fits one model with arbitrary default settings, reports only training-set or in-sample cross-validation numbers (or nothing at all), never holds out labeled data to estimate true generalization, never compares at least one alternative configuration (different features, class weighting, stronger/regularized model, text normalization), and never checks whether the estimated score is anywhere near what the task's grading likely requires. Verification scripts only re-check file formatting, not quality.
Detection procedure
  1. Read the task to see that predictions will be scored against hidden labels, and note the number of classes / difficulty (a many-class or highly overlapping label space means a weak baseline will score low).
  2. Read the scripts and check whether any labeled data is held out (train/validation split or CV with the estimate actually inspected) and whether the reported number is an honest generalization estimate rather than fit-on-all-data accuracy.
  3. Check whether more than one modeling choice was evaluated and selected on that held-out estimate, and whether the agent stated the expected score and judged it acceptable.
  4. Look at the final answer/verification step: if it only confirms row count, column name, and label vocabulary, with no quality estimate, flag it.
Discriminator
Not a violation if the agent measured a held-out score, showed it is high for the label space, and reasonably concluded further tuning was unnecessary; it is a violation when the only evidence is in-sample accuracy, a formatting check, or an unexamined CV printout, with no comparison or acceptance criterion for the score.
Consequence
The exported predictions are from an untuned baseline whose real accuracy falls below the grader's threshold, so the expected output file is marked WRONG even though its shape and column name are valid.
id 2a5da96c0707 · mined from da-code dacode-ml-multi-011@s13
raw text (what the judge reads)
### Accepting a single default baseline without measuring held-out predictive quality against a plausible bar
- **Applies when**: `task` -- the task asks for predictions on an unlabeled evaluation file that will be graded on accuracy/score, and the scripts fit one off-the-shelf pipeline (default vectorizer/model hyperparameters) and immediately export predictions.
- **Pattern**: The attempt treats "the file has the right column and ran without error" as success: it fits one model with arbitrary default settings, reports only training-set or in-sample cross-validation numbers (or nothing at all), never holds out labeled data to estimate true generalization, never compares at least one alternative configuration (different features, class weighting, stronger/regularized model, text normalization), and never checks whether the estimated score is anywhere near what the task's grading likely requires. Verification scripts only re-check file formatting, not quality.
- **Detection procedure**:
  1. Read the task to see that predictions will be scored against hidden labels, and note the number of classes / difficulty (a many-class or highly overlapping label space means a weak baseline will score low).
  2. Read the scripts and check whether any labeled data is held out (train/validation split or CV with the estimate actually inspected) and whether the reported number is an honest generalization estimate rather than fit-on-all-data accuracy.
  3. Check whether more than one modeling choice was evaluated and selected on that held-out estimate, and whether the agent stated the expected score and judged it acceptable.
  4. Look at the final answer/verification step: if it only confirms row count, column name, and label vocabulary, with no quality estimate, flag it.
- **Discriminator**: Not a violation if the agent measured a held-out score, showed it is high for the label space, and reasonably concluded further tuning was unnecessary; it *is* a violation when the only evidence is in-sample accuracy, a formatting check, or an unexamined CV printout, with no comparison or acceptance criterion for the score.
- **Consequence**: The exported predictions are from an untuned baseline whose real accuracy falls below the grader's threshold, so the expected output file is marked WRONG even though its shape and column name are valid.
697Required output artifact never written or verifiedtaskda-code
Applies when
task -- the task asks for predictions/results to be saved to a specific file following a given template, and the agent's deliverable is a prose summary of modeling work.
Pattern
The attempt reports model choice, validation scores and prediction statistics, but no script step actually writes the required file (or writes it without checking it against the template's row count, ID column, ordering and column names), so the graded artifact is missing or malformed while the narrative sounds complete.
Detection procedure
  1. From the task, list the exact required output file(s), their column names, and the reference template.
  2. Search the scripts for a write of that exact filename (e.g. to_csv("submission.csv", index=False)) and confirm it runs in the final executed path, not only in an abandoned branch; if no scripts are saved, treat the artifact as unverified.
  3. Check that the written frame is built from the test/holdout inputs and is re-loaded and compared to the template: same number of rows as test set, identical ID values in identical order, identical column headers.
  4. Confirm the answer text references the produced file and its verified shape/head rather than only validation metrics and prediction summary stats.
Discriminator
A real violation is no write call, a write of a differently named/shaped file, or a write whose shape/IDs were never checked against the template; a look-alike that is fine is an attempt that writes the file, prints its shape/head and a template diff, and then additionally narrates model selection.
Consequence
The grader's file check reports the expected output as WRONG/MISSING and scores 0, regardless of how good the reported validation score was.
id a371cd2eb48a · mined from da-code dacode-ml-competition-008@s13
raw text (what the judge reads)
### Required output artifact never written or verified
- **Applies when**: `task` -- the task asks for predictions/results to be saved to a specific file following a given template, and the agent's deliverable is a prose summary of modeling work.
- **Pattern**: The attempt reports model choice, validation scores and prediction statistics, but no script step actually writes the required file (or writes it without checking it against the template's row count, ID column, ordering and column names), so the graded artifact is missing or malformed while the narrative sounds complete.
- **Detection procedure**:
  1. From the task, list the exact required output file(s), their column names, and the reference template.
  2. Search the scripts for a write of that exact filename (e.g. `to_csv("submission.csv", index=False)`) and confirm it runs in the final executed path, not only in an abandoned branch; if no scripts are saved, treat the artifact as unverified.
  3. Check that the written frame is built from the test/holdout inputs and is re-loaded and compared to the template: same number of rows as test set, identical ID values in identical order, identical column headers.
  4. Confirm the answer text references the produced file and its verified shape/head rather than only validation metrics and prediction summary stats.
- **Discriminator**: A real violation is no write call, a write of a differently named/shaped file, or a write whose shape/IDs were never checked against the template; a look-alike that is fine is an attempt that writes the file, prints its shape/head and a template diff, and then additionally narrates model selection.
- **Consequence**: The grader's file check reports the expected output as WRONG/MISSING and scores 0, regardless of how good the reported validation score was.
698Invented qualification thresholds / aggregation rules instead of the ones specified in the task materialstaskda-code
Applies when
task -- The task text or accompanying files (README definitions, a sample/template output file) specify how a derived entity-level statistic must be computed, filtered, or formatted, and the scripts must produce a ranked/aggregated output.
Pattern
The agent never opens or echoes the provided specification artifacts (definition text, sample output template), and instead hard-codes its own eligibility filter, aggregation function, or column naming/ordering ("threshold = 1000", sum vs mean, self-chosen header names), justifying it as "to ensure reliability" rather than citing the spec.
Detection procedure
  1. Read the task statement and note every stated rule: qualification/minimum criteria, which statistic (mean vs sum vs count), rounding/units, ordering, and the referenced format/template file.
  2. Search the scripts for whether the template/spec file is loaded or its columns are read; check whether each magic number and aggregation in the code traces back to a stated rule.
  3. Compare the produced output's columns, column order, index/id convention, and value types against the template's actual contents (not against an assumed layout).
  4. Flag if any filter constant, aggregation choice, or header is agent-invented, or if the spec text was truncated/unread and no attempt was made to recover it.
Discriminator
A genuine violation is an unjustified constant or aggregation that changes membership/ordering of the reported results, or headers that differ from the template. It is fine if the spec is genuinely silent, the agent states the ambiguity, and shows the ranking is robust (e.g., tests several thresholds) while still matching the template's format exactly.
Consequence
The output file has the right shape but wrong rows (different qualifying entities/ordering) or mismatched headers, so an exact-match file check fails despite the pipeline "running successfully".
id 6c4599c4c905 · mined from da-code dacode-dm-csv-009@s13
raw text (what the judge reads)
### Invented qualification thresholds / aggregation rules instead of the ones specified in the task materials
- **Applies when**: `task` -- The task text or accompanying files (README definitions, a sample/template output file) specify how a derived entity-level statistic must be computed, filtered, or formatted, and the scripts must produce a ranked/aggregated output.
- **Pattern**: The agent never opens or echoes the provided specification artifacts (definition text, sample output template), and instead hard-codes its own eligibility filter, aggregation function, or column naming/ordering ("threshold = 1000", `sum` vs `mean`, self-chosen header names), justifying it as "to ensure reliability" rather than citing the spec.
- **Detection procedure**:
  1. Read the task statement and note every stated rule: qualification/minimum criteria, which statistic (mean vs sum vs count), rounding/units, ordering, and the referenced format/template file.
  2. Search the scripts for whether the template/spec file is loaded or its columns are read; check whether each magic number and aggregation in the code traces back to a stated rule.
  3. Compare the produced output's columns, column order, index/id convention, and value types against the template's actual contents (not against an assumed layout).
  4. Flag if any filter constant, aggregation choice, or header is agent-invented, or if the spec text was truncated/unread and no attempt was made to recover it.
- **Discriminator**: A genuine violation is an unjustified constant or aggregation that changes membership/ordering of the reported results, or headers that differ from the template. It is fine if the spec is genuinely silent, the agent states the ambiguity, and shows the ranking is robust (e.g., tests several thresholds) while still matching the template's format exactly.
- **Consequence**: The output file has the right shape but wrong rows (different qualifying entities/ordering) or mismatched headers, so an exact-match file check fails despite the pipeline "running successfully".
699Config/spec file treated as optional, with silent fallbacks to invented defaults or synthetic datataskda-code
Applies when
task -- the task points to an external specification file (config/params/schema) and/or a fixed set of output artifacts, and the scripts must read that spec and emit exactly those artifacts.
Pattern
The script wraps the spec load in if os.path.exists(...)/try-except, pre-populates its own hardcoded defaults (title, labels, bin width, size, colors), and similarly falls back to glob-discovered or randomly generated data if the expected input isn't found. It then reports success from these defaults without ever showing the spec's actual contents, and produces only a subset of the required output files.
Detection procedure
  1. From the task statement, list every named input spec and every required output artifact (file names/formats).
  2. In the scripts, check whether each spec is read unconditionally and whether every key that drives the computation (grouping/binning, ordering, labels, rounding, units) comes from the spec rather than from a local default dict; check that every required output file is written somewhere.
  3. In the run logs/answer, look for evidence that the real spec was parsed and echoed (concrete values quoted from it) versus generic defaults or "if not found, create sample data" paths.
  4. Flag if the reported parameters match the script's hardcoded defaults, if any required artifact is never written, or if a fabricated-data branch could have produced the result.
Discriminator
A real violation is when spec-driven choices are actually decided by fallback defaults, or when the reported values are indistinguishable from defaults, or required outputs are missing. It is fine to define defaults that are provably overridden by every key present in the spec, provided the script verifies the spec/data exist and hard-fails otherwise, and all required artifacts are produced.
Consequence
Outputs are graded against the spec (bin edges, labels, ordering, figure attributes) and the companion artifacts; defaults yield mismatched bins/labels and missing files, so every file-level check fails even though the script "succeeded".
id 65f701d330ee · mined from da-code dacode-plot-bar-007@s13
raw text (what the judge reads)
### Config/spec file treated as optional, with silent fallbacks to invented defaults or synthetic data
- **Applies when**: `task` -- the task points to an external specification file (config/params/schema) and/or a fixed set of output artifacts, and the scripts must read that spec and emit exactly those artifacts.
- **Pattern**: The script wraps the spec load in `if os.path.exists(...)`/`try-except`, pre-populates its own hardcoded defaults (title, labels, bin width, size, colors), and similarly falls back to `glob`-discovered or randomly generated data if the expected input isn't found. It then reports success from these defaults without ever showing the spec's actual contents, and produces only a subset of the required output files.
- **Detection procedure**:
  1. From the task statement, list every named input spec and every required output artifact (file names/formats).
  2. In the scripts, check whether each spec is read unconditionally and whether every key that drives the computation (grouping/binning, ordering, labels, rounding, units) comes from the spec rather than from a local default dict; check that every required output file is written somewhere.
  3. In the run logs/answer, look for evidence that the real spec was parsed and echoed (concrete values quoted from it) versus generic defaults or "if not found, create sample data" paths.
  4. Flag if the reported parameters match the script's hardcoded defaults, if any required artifact is never written, or if a fabricated-data branch could have produced the result.
- **Discriminator**: A real violation is when spec-driven choices are actually decided by fallback defaults, or when the reported values are indistinguishable from defaults, or required outputs are missing. It is fine to define defaults that are provably overridden by every key present in the spec, provided the script verifies the spec/data exist and hard-fails otherwise, and all required artifacts are produced.
- **Consequence**: Outputs are graded against the spec (bin edges, labels, ordering, figure attributes) and the companion artifacts; defaults yield mismatched bins/labels and missing files, so every file-level check fails even though the script "succeeded".
700Sample vs. population standard deviation (ddof) not reconciledtaskinfiagent-dabench
Applies when
task -- the task asks for a standard deviation (or variance, or any statistic whose definition depends on a degrees-of-freedom / normalization choice) computed from a data column, and the script uses a library default without stating the convention.
Pattern
The attempt computes the spread with whichever default the chosen library provides (e.g. NumPy's ddof=0 vs. pandas' ddof=1, or a hand-rolled formula dividing by n instead of n-1), never checks the alternative, and reports a single number. Means and counts agree with the reference, so the error looks like a rounding artifact rather than a definitional one — often the same script also mixes libraries (pandas for the mean, NumPy for the deviation, or vice versa) so conventions are inconsistent within one answer.
Detection procedure
  1. In the task statement, note that a standard deviation/variance is requested and that no explicit ddof/sample-vs-population convention is given.
  2. In the scripts, find the exact call producing the reported value and record the library and its default normalization; also check whether any thresholding/standardization step (e.g. z-scores) uses a different library, i.e. a different ddof, than the final report.
  3. Verify the script computes and compares both conventions (or documents why one is chosen, e.g. matching the tool used elsewhere in the pipeline); flag the attempt if only one value was ever produced.
  4. Sanity-check magnitude: for the reported n, the two conventions differ by a factor sqrt(n/(n-1)) — if that ratio moved the second decimal place, a single un-justified choice is a coin flip and must be flagged.
Discriminator
Not a violation if the task (or the surrounding pipeline convention) pins the definition and the script demonstrably uses it, or if the script explicitly computes both and shows they round to the same two decimals. It is a violation when only one default was used silently and the two conventions round differently — regardless of which one happens to be right.
Consequence
The mean and outlier list match the reference while the dispersion value is off in the second decimal, so the grader marks a partial pass (e.g. 1 of 2 checks) and the overall answer is scored incorrect.
id cf350db21c06 · mined from infiagent-dabench dabench-495@s13
raw text (what the judge reads)
### Sample vs. population standard deviation (ddof) not reconciled
- **Applies when**: `task` -- the task asks for a standard deviation (or variance, or any statistic whose definition depends on a degrees-of-freedom / normalization choice) computed from a data column, and the script uses a library default without stating the convention.
- **Pattern**: The attempt computes the spread with whichever default the chosen library provides (e.g. NumPy's `ddof=0` vs. pandas' `ddof=1`, or a hand-rolled formula dividing by `n` instead of `n-1`), never checks the alternative, and reports a single number. Means and counts agree with the reference, so the error looks like a rounding artifact rather than a definitional one — often the same script also mixes libraries (pandas for the mean, NumPy for the deviation, or vice versa) so conventions are inconsistent within one answer.
- **Detection procedure**:
  1. In the task statement, note that a standard deviation/variance is requested and that no explicit ddof/sample-vs-population convention is given.
  2. In the scripts, find the exact call producing the reported value and record the library and its default normalization; also check whether any thresholding/standardization step (e.g. z-scores) uses a *different* library, i.e. a different ddof, than the final report.
  3. Verify the script computes and compares both conventions (or documents why one is chosen, e.g. matching the tool used elsewhere in the pipeline); flag the attempt if only one value was ever produced.
  4. Sanity-check magnitude: for the reported `n`, the two conventions differ by a factor `sqrt(n/(n-1))` — if that ratio moved the second decimal place, a single un-justified choice is a coin flip and must be flagged.
- **Discriminator**: Not a violation if the task (or the surrounding pipeline convention) pins the definition and the script demonstrably uses it, or if the script explicitly computes both and shows they round to the same two decimals. It *is* a violation when only one default was used silently and the two conventions round differently — regardless of which one happens to be right.
- **Consequence**: The mean and outlier list match the reference while the dispersion value is off in the second decimal, so the grader marks a partial pass (e.g. 1 of 2 checks) and the overall answer is scored incorrect.
701Answer asserted from prior knowledge instead of computed from the provided datataskda-code
Applies when
task -- the task asks for specific entities/values (e.g., top-N by a column) to be extracted from a supplied dataset after a stated preprocessing step, and the deliverable is a fixed-format file.
Pattern
The attempt produces a plausible-looking, "textbook" answer list without a saved/runnable script that loads the file, applies the stated preprocessing, sorts, and writes the required output — so the labels, values, and ordering come from world knowledge or an unverified guess rather than the dataset's own rows and spellings.
Detection procedure
  1. Read the task for the required computation (preprocessing rule, ranking direction, sort order) and the exact output artifact/format.
  2. Look for a script that reads the dataset, performs each stated step, prints the ranked slice with its numeric values, and writes the artifact; if no such script or no printed intermediates exist, the answer is unverifiable.
  3. Cross-check the reported names against the dataset's own label strings (spelling/naming conventions) and check the requested ordering is actually applied to both lists, not just the "top" one.
  4. Confirm the number of returned items and that the values are in a plausible range for the column after imputation.
Discriminator
A real violation is an answer with no executable derivation or no printed values tying each returned item to a row in the data; a look-alike that is fine is a script-derived answer that happens to match common knowledge but prints the sorted values, entity names as they appear in the file, and writes the artifact in the requested shape.
Consequence
The graded file mismatches the expected keys/names/order (e.g., differently spelled entities, wrong sort direction, or entities absent from the data), scoring 0 even though the answer "looks right".
id 8133afe098f1 · mined from da-code dacode-di-text-003@s13
raw text (what the judge reads)
### Answer asserted from prior knowledge instead of computed from the provided data
- **Applies when**: `task` -- the task asks for specific entities/values (e.g., top-N by a column) to be extracted from a supplied dataset after a stated preprocessing step, and the deliverable is a fixed-format file.
- **Pattern**: The attempt produces a plausible-looking, "textbook" answer list without a saved/runnable script that loads the file, applies the stated preprocessing, sorts, and writes the required output — so the labels, values, and ordering come from world knowledge or an unverified guess rather than the dataset's own rows and spellings.
- **Detection procedure**:
  1. Read the task for the required computation (preprocessing rule, ranking direction, sort order) and the exact output artifact/format.
  2. Look for a script that reads the dataset, performs each stated step, prints the ranked slice with its numeric values, and writes the artifact; if no such script or no printed intermediates exist, the answer is unverifiable.
  3. Cross-check the reported names against the dataset's own label strings (spelling/naming conventions) and check the requested ordering is actually applied to both lists, not just the "top" one.
  4. Confirm the number of returned items and that the values are in a plausible range for the column after imputation.
  5. 
- **Discriminator**: A real violation is an answer with no executable derivation or no printed values tying each returned item to a row in the data; a look-alike that is fine is a script-derived answer that happens to match common knowledge but prints the sorted values, entity names as they appear in the file, and writes the artifact in the requested shape.
- **Consequence**: The graded file mismatches the expected keys/names/order (e.g., differently spelled entities, wrong sort direction, or entities absent from the data), scoring 0 even though the answer "looks right".
702Answer-format template not reproduced literally (quoting/delimiters dropped)taskinfiagent-dabench
Applies when
task -- the task specifies an exact answer string with tagged fields, including literal punctuation such as quotation marks, brackets, or separators, and states field values are "strings".
Pattern
The attempt computes the right values but emits them in a paraphrased shell of the template — quotes around string values removed, extra/missing whitespace, tags renamed or reordered, or values printed as separate log lines rather than the single required expression — so an exact-match grader fails every field even though the analysis is correct.
Detection procedure
  1. Copy the answer template from the task verbatim and note every literal character: tag names, brackets, quote marks around string values, and the separator between fields.
  2. Read the script's final print/report statements and the submitted answer, and diff them character-by-character against the template (especially whether string values are wrapped in the quotes shown in the spec).
  3. Check that the answer is a single line containing all required tags in the specified order, with no substituted colons/spaces or added commentary inside the brackets.
  4. If any literal from the template is missing or altered, flag the attempt regardless of whether the underlying numbers are right.
Discriminator
A real violation is a deviation from characters the task showed literally (e.g., @field["Yes"] submitted as @field[Yes]). A look-alike that is fine is deviation only in parts the task left free-form (e.g., a value the task described generically as "range" where no example quoting was mandated), or harmless surrounding prose outside the answer line.
Consequence
The grader reports every expected field as WRONG/MISSING despite the submitted values equalling ground truth, yielding 0/N checks passed.
id da1642076580 · mined from infiagent-dabench dabench-550@s13
raw text (what the judge reads)
### Answer-format template not reproduced literally (quoting/delimiters dropped)
- **Applies when**: `task` -- the task specifies an exact answer string with tagged fields, including literal punctuation such as quotation marks, brackets, or separators, and states field values are "strings".
- **Pattern**: The attempt computes the right values but emits them in a paraphrased shell of the template — quotes around string values removed, extra/missing whitespace, tags renamed or reordered, or values printed as separate log lines rather than the single required expression — so an exact-match grader fails every field even though the analysis is correct.
- **Detection procedure**:
  1. Copy the answer template from the task verbatim and note every literal character: tag names, brackets, quote marks around string values, and the separator between fields.
  2. Read the script's final print/report statements and the submitted answer, and diff them character-by-character against the template (especially whether string values are wrapped in the quotes shown in the spec).
  3. Check that the answer is a single line containing all required tags in the specified order, with no substituted colons/spaces or added commentary inside the brackets.
  4. If any literal from the template is missing or altered, flag the attempt regardless of whether the underlying numbers are right.
- **Discriminator**: A real violation is a deviation from characters the task showed literally (e.g., `@field["Yes"]` submitted as `@field[Yes]`). A look-alike that is fine is deviation only in parts the task left free-form (e.g., a value the task described generically as "range" where no example quoting was mandated), or harmless surrounding prose outside the answer line.
- **Consequence**: The grader reports every expected field as WRONG/MISSING despite the submitted values equalling ground truth, yielding 0/N checks passed.
703Proxy substitution: inventing stand-ins for the requested entities instead of locating the right datataskda-code
Applies when
task -- the task names specific entities/quantities (e.g., a grouping key, a ranking measure, per-stage durations) and the script works on a file whose columns do not obviously contain them.
Pattern
The agent finds no columns matching the requested concepts, then silently redefines them as loosely related proxies (ranking by row counts instead of the stated measure, treating unrelated categories as the requested stages, deriving a "duration" from timestamp gaps), and reports success without ever confirming the source data supports the requested quantities or producing all requested output artifacts.
Detection procedure
1. From the task statement, list every named quantity: the grouping unit, the ranking metric, the per-category measure, and every required output file/format setting. 2. In the scripts, trace each listed quantity to an actual column or a documented derivation; flag any that is replaced by a comment-justified substitute ("since we don't have X, we'll use Y"), or where the config/README describes a domain unrelated to the task's vocabulary. 3. Check whether the agent explored other available inputs (or verified the loaded file is the intended one) before substituting, and whether stated config keys were used for their intended purpose rather than force-fitted. 4. Check the answer lists every requested artifact with plausible values, not just the figure.
Discriminator
A legitimate case is a documented, semantically equivalent derivation (e.g., computing a stage duration as the difference between two explicitly dated stage columns). A violation is when the substitute measures a different concept entirely, or when config labels/axis names clearly contradict the meaning assigned to them — a signal the wrong dataset or wrong mapping is in use.
Consequence
The produced numbers and plot encode a different quantity than requested, so value/array comparisons and any auxiliary expected files fail, and missing required outputs score zero regardless of chart aesthetics.
id e9021bacecbb · mined from da-code dacode-plot-scatter-002@s13
raw text (what the judge reads)
### Proxy substitution: inventing stand-ins for the requested entities instead of locating the right data
- **Applies when**: `task` -- the task names specific entities/quantities (e.g., a grouping key, a ranking measure, per-stage durations) and the script works on a file whose columns do not obviously contain them.
- **Pattern**: The agent finds no columns matching the requested concepts, then silently redefines them as loosely related proxies (ranking by row counts instead of the stated measure, treating unrelated categories as the requested stages, deriving a "duration" from timestamp gaps), and reports success without ever confirming the source data supports the requested quantities or producing all requested output artifacts.
- **Detection procedure**: 1. From the task statement, list every named quantity: the grouping unit, the ranking metric, the per-category measure, and every required output file/format setting. 2. In the scripts, trace each listed quantity to an actual column or a documented derivation; flag any that is replaced by a comment-justified substitute ("since we don't have X, we'll use Y"), or where the config/README describes a domain unrelated to the task's vocabulary. 3. Check whether the agent explored other available inputs (or verified the loaded file is the intended one) before substituting, and whether stated config keys were used for their intended purpose rather than force-fitted. 4. Check the answer lists every requested artifact with plausible values, not just the figure.
- **Discriminator**: A legitimate case is a documented, semantically equivalent derivation (e.g., computing a stage duration as the difference between two explicitly dated stage columns). A violation is when the substitute measures a different concept entirely, or when config labels/axis names clearly contradict the meaning assigned to them — a signal the wrong dataset or wrong mapping is in use.
- **Consequence**: The produced numbers and plot encode a different quantity than requested, so value/array comparisons and any auxiliary expected files fail, and missing required outputs score zero regardless of chart aesthetics.
704Output artifact not written to the exact requested filename/path/format from the provided templatetaskda-code
Applies when
task -- the task names a specific output file (and/or says the format must match a provided template) and the scripts write results to disk.
Pattern
The agent computes plausible numbers but saves them under a self-chosen filename, directory, or layout (different spelling, extension, index/header presence, decimal precision, column names/order) without ever opening or diffing against the referenced template/expected artifact.
Detection procedure
  1. Read the task statement and copy out the literal target filename (including any unusual spelling), path, and any stated format constraints (rounding, units, ordering, index column, header names).
  2. Search the scripts for the write call (to_csv, to_json, savefig, etc.) and compare its path string character-by-character with the literal target; also check index=, column order, and rounding against the stated constraints.
  3. Check whether any script loads/prints the provided template or an example output and asserts shape, column names, dtypes, and value precision against the produced file; absence of such a comparison is a red flag.
  4. Inspect the final answer's header row and a sample row for template conformance (e.g., expected column labels, expected number of decimals, no stray index column).
Discriminator
A real violation is a mismatch in the artifact the grader reads — wrong name/path spelling, missing file at that path, or a schema/precision that differs from the template. It is not a violation if the file is written to the exact requested name and layout and an extra copy/intermediate is saved elsewhere, or if formatting details are genuinely unspecified anywhere in the task or template.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even when the underlying computation is numerically correct.
id f37317c85d5a · mined from da-code dacode-dm-csv-043@s13
raw text (what the judge reads)
### Output artifact not written to the exact requested filename/path/format from the provided template
- **Applies when**: `task` -- the task names a specific output file (and/or says the format must match a provided template) and the scripts write results to disk.
- **Pattern**: The agent computes plausible numbers but saves them under a self-chosen filename, directory, or layout (different spelling, extension, index/header presence, decimal precision, column names/order) without ever opening or diffing against the referenced template/expected artifact.
- **Detection procedure**:
  1. Read the task statement and copy out the literal target filename (including any unusual spelling), path, and any stated format constraints (rounding, units, ordering, index column, header names).
  2. Search the scripts for the write call (`to_csv`, `to_json`, `savefig`, etc.) and compare its path string character-by-character with the literal target; also check `index=`, column order, and rounding against the stated constraints.
  3. Check whether any script loads/prints the provided template or an example output and asserts shape, column names, dtypes, and value precision against the produced file; absence of such a comparison is a red flag.
  4. Inspect the final answer's header row and a sample row for template conformance (e.g., expected column labels, expected number of decimals, no stray index column).
- **Discriminator**: A real violation is a mismatch in the artifact the grader reads — wrong name/path spelling, missing file at that path, or a schema/precision that differs from the template. It is *not* a violation if the file is written to the exact requested name and layout and an extra copy/intermediate is saved elsewhere, or if formatting details are genuinely unspecified anywhere in the task or template.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even when the underlying computation is numerically correct.
705Validation split that overlaps the training data (self-evaluation leakage)taskda-code
Applies when
task -- a script carves out a hold-out/validation subset to compare models or report accuracy before producing final predictions on an unlabeled test set.
Pattern
The script creates a train/validation split but then fits each candidate model on the full labelled dataset (or otherwise on rows that include the validation rows), and evaluates on the validation subset. The resulting score is a memorization score, so model comparison/selection and the reported metric are both invalid; the agent then presents that inflated score as evidence the predictions are good, without any independent sanity check (e.g. comparing prediction distribution to the target distribution, or cross-validation).
Detection procedure
  1. In the task, note that the reported metric is meant to estimate generalization to unseen rows and that the deliverable is predictions for a separate unlabeled file.
  2. In the scripts, trace the exact arrays passed to fit(...) and to predict(...)/metric calls for each candidate model; check that the fitted rows and the evaluated rows are disjoint (also check scalers/encoders are fit on training rows only).
  3. Check whether any leakage-free estimate exists (proper hold-out or CV) and whether the reported metric matches it; check whether the answer's claimed score is plausible given the feature set used.
  4. Check the answer for basic sanity validation of the output itself (row count matches test file, column name/order, value range and spread comparable to the observed target distribution).
Discriminator
A real violation is fitting on a superset of the evaluation rows (e.g. model.fit(X_all, y_all) followed by scoring on X_val drawn from X_all), or refitting a transformer on all data before splitting. It is not a violation to fit on the training split, score on the held-out split, and only afterwards refit on all labelled data for the final prediction — provided the reported metric comes from the leakage-free fit.
Consequence
Reported R²/RMSE is wildly optimistic relative to true performance, the "best model" may be the worst generalizer, and the submitted prediction file's values deviate far enough from the true targets that the grader's accuracy/error tolerance check on the output file fails.
id 4b5cf4d34650 · mined from da-code dacode-ml-regression-004@s13
raw text (what the judge reads)
### Validation split that overlaps the training data (self-evaluation leakage)
- **Applies when**: `task` -- a script carves out a hold-out/validation subset to compare models or report accuracy before producing final predictions on an unlabeled test set.
- **Pattern**: The script creates a train/validation split but then fits each candidate model on the *full* labelled dataset (or otherwise on rows that include the validation rows), and evaluates on the validation subset. The resulting score is a memorization score, so model comparison/selection and the reported metric are both invalid; the agent then presents that inflated score as evidence the predictions are good, without any independent sanity check (e.g. comparing prediction distribution to the target distribution, or cross-validation).
- **Detection procedure**:
  1. In the task, note that the reported metric is meant to estimate generalization to unseen rows and that the deliverable is predictions for a separate unlabeled file.
  2. In the scripts, trace the exact arrays passed to `fit(...)` and to `predict(...)`/metric calls for each candidate model; check that the fitted rows and the evaluated rows are disjoint (also check scalers/encoders are fit on training rows only).
  3. Check whether any leakage-free estimate exists (proper hold-out or CV) and whether the reported metric matches it; check whether the answer's claimed score is plausible given the feature set used.
  4. Check the answer for basic sanity validation of the output itself (row count matches test file, column name/order, value range and spread comparable to the observed target distribution).
- **Discriminator**: A real violation is fitting on a superset of the evaluation rows (e.g. `model.fit(X_all, y_all)` followed by scoring on `X_val` drawn from `X_all`), or refitting a transformer on all data before splitting. It is *not* a violation to fit on the training split, score on the held-out split, and only afterwards refit on all labelled data for the final prediction — provided the reported metric comes from the leakage-free fit.
- **Consequence**: Reported R²/RMSE is wildly optimistic relative to true performance, the "best model" may be the worst generalizer, and the submitted prediction file's values deviate far enough from the true targets that the grader's accuracy/error tolerance check on the output file fails.
706Statistic computed over a different slice than the task specifies (scope silently redefined)taskinfiagent-dabench
Applies when
task -- the task asks for a statistic over a specific slice (a given year, group, subset, or axis) and the script aggregates over whatever dimension is convenient in the loaded file's layout.
Pattern
The script ignores the stated qualifier because the literal slice looks degenerate under the assumed table layout (e.g., a distribution statistic would have only one value per entity), so it substitutes a broader/different aggregation (all periods, all columns, or the transposed axis) without checking whether another file, a long-format table, or repeated rows per entity would make the requested slice well-defined. It also may pass a definition/flag that contradicts the explicitly stated variant of the statistic (biased vs. bias-corrected, sample vs. population).
Detection procedure
  1. From the task, write down (a) the exact slice the statistic must be computed on, and (b) any stated definition/variant or option that must be used.
  2. In the script, locate the array actually passed to the statistic function and confirm it contains exactly the values in slice (a) — check which columns/rows are selected and along which axis the reduction happens.
  3. If the requested slice would yield too few values for the statistic to be meaningful, check whether the script verified the data layout (inspect available files, shape, duplicate keys per entity, wide vs. long format) instead of silently widening the scope.
  4. Compare the function's arguments to (b) — e.g., the bias/ddof/definition keyword — and confirm they encode the requested variant, not the library default.
Discriminator
A real violation is when the printed inputs to the statistic span a dimension the task did not ask for (or the option contradicts the stated definition); it is fine if the script demonstrably reduced to the requested slice and the extra values it uses are the legitimate within-slice observations (e.g., multiple records per entity inside that slice).
Consequence
The ranking/argmax is produced from a different distribution than requested, so the reported entity does not match ground truth and the answer fails the exact-match check even though the code runs cleanly.
id f9e05101455d · mined from infiagent-dabench dabench-252@s13
raw text (what the judge reads)
### Statistic computed over a different slice than the task specifies (scope silently redefined)
- **Applies when**: `task` -- the task asks for a statistic over a specific slice (a given year, group, subset, or axis) and the script aggregates over whatever dimension is convenient in the loaded file's layout.
- **Pattern**: The script ignores the stated qualifier because the literal slice looks degenerate under the assumed table layout (e.g., a distribution statistic would have only one value per entity), so it substitutes a broader/different aggregation (all periods, all columns, or the transposed axis) without checking whether another file, a long-format table, or repeated rows per entity would make the requested slice well-defined. It also may pass a definition/flag that contradicts the explicitly stated variant of the statistic (biased vs. bias-corrected, sample vs. population).
- **Detection procedure**:
  1. From the task, write down (a) the exact slice the statistic must be computed on, and (b) any stated definition/variant or option that must be used.
  2. In the script, locate the array actually passed to the statistic function and confirm it contains exactly the values in slice (a) — check which columns/rows are selected and along which axis the reduction happens.
  3. If the requested slice would yield too few values for the statistic to be meaningful, check whether the script verified the data layout (inspect available files, shape, duplicate keys per entity, wide vs. long format) instead of silently widening the scope.
  4. Compare the function's arguments to (b) — e.g., the bias/ddof/definition keyword — and confirm they encode the requested variant, not the library default.
- **Discriminator**: A real violation is when the printed inputs to the statistic span a dimension the task did not ask for (or the option contradicts the stated definition); it is fine if the script demonstrably reduced to the requested slice and the extra values it uses are the legitimate within-slice observations (e.g., multiple records per entity inside that slice).
- **Consequence**: The ranking/argmax is produced from a different distribution than requested, so the reported entity does not match ground truth and the answer fails the exact-match check even though the code runs cleanly.
707Coercing an identified key/label to a coarser granularity than the record it identifiestaskinfiagent-dabench
Applies when
task -- the deliverable includes an identifier (date, ID, category key) of a row selected by an extremum or filter, and the answer template shows a format/pattern that is less precise than the values actually stored in the data.
Pattern
The agent takes the literal template at face value and truncates or reformats the selected key (e.g., dropping components of a timestamp, zero-padding away detail, rounding an index), reporting a coarser label instead of the exact key of the selected record — even though every downstream computation in the same task was done at the finer granularity.
Detection procedure
  1. Read the task: note which record the extremum/filter selects and what granularity the companion calculation (e.g., "previous row") implicitly requires.
  2. Read the scripts: find where the selected key is printed/formatted and check for any truncation, string slicing, strftime/round/cast, or reformatting applied only for output.
  3. Compare the reported key with the key actually used internally to look up the related record; if the reported string could match many rows in the data while the internal one matches exactly one, flag it.
  4. Check that the reported key, pasted back into the dataset, uniquely reproduces the reported statistic.
Discriminator
A real violation is losing information the data contains (many rows share the reported label). It is fine if the underlying values genuinely have that granularity, or if the task explicitly asks for an aggregate at the coarser level (e.g., "the month with the highest average") — there the coarse label is the selected key.
Consequence
The dependent numeric answer may be correct while the identifier check fails an exact-string comparison, so the submission is marked wrong on partial credit (e.g., 1/2 checks passed).
id d38cdb543205 · mined from infiagent-dabench dabench-572@s13
raw text (what the judge reads)
### Coercing an identified key/label to a coarser granularity than the record it identifies
- **Applies when**: `task` -- the deliverable includes an identifier (date, ID, category key) of a row selected by an extremum or filter, and the answer template shows a format/pattern that is less precise than the values actually stored in the data.
- **Pattern**: The agent takes the literal template at face value and truncates or reformats the selected key (e.g., dropping components of a timestamp, zero-padding away detail, rounding an index), reporting a coarser label instead of the exact key of the selected record — even though every downstream computation in the same task was done at the finer granularity.
- **Detection procedure**:
  1. Read the task: note which record the extremum/filter selects and what granularity the companion calculation (e.g., "previous row") implicitly requires.
  2. Read the scripts: find where the selected key is printed/formatted and check for any truncation, string slicing, `strftime`/`round`/cast, or reformatting applied only for output.
  3. Compare the reported key with the key actually used internally to look up the related record; if the reported string could match many rows in the data while the internal one matches exactly one, flag it.
  4. Check that the reported key, pasted back into the dataset, uniquely reproduces the reported statistic.
- **Discriminator**: A real violation is losing information the data contains (many rows share the reported label). It is fine if the underlying values genuinely have that granularity, or if the task explicitly asks for an aggregate at the coarser level (e.g., "the month with the highest average") — there the coarse label *is* the selected key.
- **Consequence**: The dependent numeric answer may be correct while the identifier check fails an exact-string comparison, so the submission is marked wrong on partial credit (e.g., 1/2 checks passed).
708Deliverable file never written/verified against the required templatetaskda-code
Applies when
task -- the task requires producing an output artifact (a file at a given name/path, with a column name, row count and value format specified by an example/sample file), and the agent's answer is a prose summary of its modeling work.
Pattern
The attempt focuses on model selection and metrics, then narrates the prediction counts in the answer, without a script step that (a) writes the file to the exact requested path/filename, (b) re-reads it, and (c) asserts header name(s), row count equal to the test set, and allowed value set/dtype match the provided sample. Any silent failure (wrong directory, index column written, header omitted, wrong length after dropping rows with missing values, probabilities instead of labels) goes undetected because the answer is never cross-checked against the artifact.
Detection procedure
  1. From the task, list the exact deliverable: filename, column header(s), expected number of rows, and value format shown in the sample/example file.
  2. Search the scripts for the write call; confirm it targets that exact filename/path, uses the specified header, and suppresses any extra index/ID column.
  3. Look for a post-write verification (re-load the file, compare shape to the test input's row count, compare columns and unique values to the sample) and an explicit sanity check that no rows were lost during preprocessing/imputation/filtering.
  4. Check that the counts quoted in the final answer are computed from the reloaded file, not from an in-memory array — and that the answer states the file was saved at the required location.
Discriminator
A real violation is when no code path provably creates the named file with the sample's schema, or the file's row/column/value properties are never asserted; it is fine if the script writes with the correct name and header and includes (or the answer reports) a read-back check showing matching shape, header, and label values — even if the model itself is mediocre.
Consequence
The grader looks for the expected artifact and finds it missing, mis-named, or mis-shaped (extra index column, wrong header, fewer/more rows than test observations), so every check fails regardless of the reported ROC-AUC/F1.
id 458fc770e14e · mined from da-code dacode-ml-binary-013@s13
raw text (what the judge reads)
### Deliverable file never written/verified against the required template
- **Applies when**: `task` -- the task requires producing an output artifact (a file at a given name/path, with a column name, row count and value format specified by an example/sample file), and the agent's answer is a prose summary of its modeling work.
- **Pattern**: The attempt focuses on model selection and metrics, then narrates the prediction counts in the answer, without a script step that (a) writes the file to the exact requested path/filename, (b) re-reads it, and (c) asserts header name(s), row count equal to the test set, and allowed value set/dtype match the provided sample. Any silent failure (wrong directory, index column written, header omitted, wrong length after dropping rows with missing values, probabilities instead of labels) goes undetected because the answer is never cross-checked against the artifact.
- **Detection procedure**:
  1. From the task, list the exact deliverable: filename, column header(s), expected number of rows, and value format shown in the sample/example file.
  2. Search the scripts for the write call; confirm it targets that exact filename/path, uses the specified header, and suppresses any extra index/ID column.
  3. Look for a post-write verification (re-load the file, compare shape to the test input's row count, compare columns and unique values to the sample) and an explicit sanity check that no rows were lost during preprocessing/imputation/filtering.
  4. Check that the counts quoted in the final answer are computed from the reloaded file, not from an in-memory array — and that the answer states the file was saved at the required location.
- **Discriminator**: A real violation is when no code path provably creates the named file with the sample's schema, or the file's row/column/value properties are never asserted; it is fine if the script writes with the correct name and header and includes (or the answer reports) a read-back check showing matching shape, header, and label values — even if the model itself is mediocre.
- **Consequence**: The grader looks for the expected artifact and finds it missing, mis-named, or mis-shaped (extra index column, wrong header, fewer/more rows than test observations), so every check fails regardless of the reported ROC-AUC/F1.
709Answer serialization doesn't match the requested literal format (and no script produces it)taskinfiagent-dabench
Applies when
task -- the task prescribes an exact answer template (e.g. @key[list_of_strings], a delimiter, a units/rounding rule) and the agent must emit a list or set of labels as the final deliverable.
Pattern
The attempt computes the right underlying values but hand-writes the final answer in its own serialization style — extra quoting, brackets, escaping, different separators, added whitespace/ordering, or key names not exactly as specified — and no saved script prints the answer string, so the formatting is never checked against the template.
Detection procedure
  1. Read the task and copy out the answer template verbatim, noting the key name, the delimiter between items, and whether items are quoted or bare.
  2. Look in the scripts for a line that constructs and prints the final answer string; if the answer exists only in prose (no script output, no saved artifact), flag it as unverifiable.
  3. Character-by-character compare the submitted answer to the template: key spelling, presence/absence of quotes around each item, separator, surrounding brackets, trailing punctuation, and item spelling exactly as it appears in the source data.
  4. If any element differs from the template, or the item labels are re-typed by hand rather than taken programmatically from the data, flag it.
Discriminator
A real violation is a deviation in the serialization or labeling (quotes, separator, key, renamed/abbreviated labels) or an answer with no script that generates it; a look-alike that is fine is an answer whose formatting is byte-compatible with the template even if the ordering of an unordered set differs or the script formats it in a different but equivalent way the task explicitly allows.
Consequence
The grader's exact/parsed comparison fails and reports the expected value as WRONG/MISSING even though the computed statistic was correct, scoring 0.
id 6ead85d74cba · mined from infiagent-dabench dabench-254@s13
raw text (what the judge reads)
### Answer serialization doesn't match the requested literal format (and no script produces it)
- **Applies when**: `task` -- the task prescribes an exact answer template (e.g. `@key[list_of_strings]`, a delimiter, a units/rounding rule) and the agent must emit a list or set of labels as the final deliverable.
- **Pattern**: The attempt computes the right underlying values but hand-writes the final answer in its own serialization style — extra quoting, brackets, escaping, different separators, added whitespace/ordering, or key names not exactly as specified — and no saved script prints the answer string, so the formatting is never checked against the template.
- **Detection procedure**:
  1. Read the task and copy out the answer template verbatim, noting the key name, the delimiter between items, and whether items are quoted or bare.
  2. Look in the scripts for a line that constructs and prints the final answer string; if the answer exists only in prose (no script output, no saved artifact), flag it as unverifiable.
  3. Character-by-character compare the submitted answer to the template: key spelling, presence/absence of quotes around each item, separator, surrounding brackets, trailing punctuation, and item spelling exactly as it appears in the source data.
  4. If any element differs from the template, or the item labels are re-typed by hand rather than taken programmatically from the data, flag it.
- **Discriminator**: A real violation is a deviation in the *serialization or labeling* (quotes, separator, key, renamed/abbreviated labels) or an answer with no script that generates it; a look-alike that is fine is an answer whose formatting is byte-compatible with the template even if the ordering of an unordered set differs or the script formats it in a different but equivalent way the task explicitly allows.
- **Consequence**: The grader's exact/parsed comparison fails and reports the expected value as WRONG/MISSING even though the computed statistic was correct, scoring 0.
710Ad-hoc invention of the target metric instead of using the specified/derivable definitiontaskda-code
Applies when
task -- the task asks to visualize or report "performance"/"score"/"ranking" over a filtered period using settings from an external config or a definition established earlier in the workflow, and the script must compute that quantity itself.
Pattern
The agent invents an arbitrary formula (e.g., counting rows and adding a weighted count of an unrelated boolean/context flag) rather than the domain-standard or explicitly specified quantity, then reports the resulting ranking as the answer; it also emits only the figure while the expected deliverables include the computed values/config echo in machine-readable form.
Detection procedure
  1. Read the task and any referenced config/spec files; list every required output artifact and every stated definition, filter, ordering, and styling constraint.
  2. In the scripts, locate the line(s) that compute the plotted/reported quantity and check whether each term traces back to a stated definition or an accepted domain convention (e.g., outcome-based points/wins) rather than to a coincidental column.
  3. Compare the produced files against the required-artifact list, and check the answer's leaderboard for face-validity (do the top entries match what a domain-aware sanity check would expect for "best performers"?).
  4. Flag if any term is unexplained, if the flag/column used is unrelated to performance, or if any required artifact is missing.
Discriminator
A real violation is a metric whose components have no justification in the task/config and whose ranking is dominated by exposure/volume or an irrelevant attribute; acceptable is a metric that is either explicitly specified, or a standard aggregation of outcome variables, with the choice stated and consistent with a plausible sanity check.
Consequence
The plotted values, ordering, and any serialized results differ from the reference, so the figure and all accompanying result/config files fail comparison — 0 checks passed despite a well-formatted chart.
id 676d27a3020b · mined from da-code dacode-plot-bar-006@s13
raw text (what the judge reads)
### Ad-hoc invention of the target metric instead of using the specified/derivable definition
- **Applies when**: `task` -- the task asks to visualize or report "performance"/"score"/"ranking" over a filtered period using settings from an external config or a definition established earlier in the workflow, and the script must compute that quantity itself.
- **Pattern**: The agent invents an arbitrary formula (e.g., counting rows and adding a weighted count of an unrelated boolean/context flag) rather than the domain-standard or explicitly specified quantity, then reports the resulting ranking as the answer; it also emits only the figure while the expected deliverables include the computed values/config echo in machine-readable form.
- **Detection procedure**:
  1. Read the task and any referenced config/spec files; list every required output artifact and every stated definition, filter, ordering, and styling constraint.
  2. In the scripts, locate the line(s) that compute the plotted/reported quantity and check whether each term traces back to a stated definition or an accepted domain convention (e.g., outcome-based points/wins) rather than to a coincidental column.
  3. Compare the produced files against the required-artifact list, and check the answer's leaderboard for face-validity (do the top entries match what a domain-aware sanity check would expect for "best performers"?).
  4. Flag if any term is unexplained, if the flag/column used is unrelated to performance, or if any required artifact is missing.
- **Discriminator**: A real violation is a metric whose components have no justification in the task/config and whose ranking is dominated by exposure/volume or an irrelevant attribute; acceptable is a metric that is either explicitly specified, or a standard aggregation of outcome variables, with the choice stated and consistent with a plausible sanity check.
- **Consequence**: The plotted values, ordering, and any serialized results differ from the reference, so the figure and all accompanying result/config files fail comparison — 0 checks passed despite a well-formatted chart.
711Anomaly/threshold counts reported without an independent recomputation and plausibility checktaskinfiagent-dabench
Applies when
task -- the script flags rows by a statistical rule (z-score, IQR, quantile cut, threshold filter) on one column and the requested answer is the count of flagged rows.
Pattern
The attempt relies on a single library call (with NaN/dtype policies it never verifies) to build the boolean mask, prints the count, and submits it — without recomputing the statistic by hand from the printed mean/std (or quantiles), without checking that the column is clean numeric (no sentinel/placeholder codes, no strings coerced or silently dropped), and without confirming that the flagged values actually lie beyond the implied cut-off and that mask length still matches the dataframe length.
Detection procedure
  1. Read the task for the exact rule and threshold, and note which column and which rows it must be applied to.
  2. In the scripts, check whether the column is validated as numeric with missing/placeholder values explicitly handled, and whether the mask is aligned to the full dataframe (a NaN-omitting helper can return a shorter or NaN-containing array that silently changes which rows are flagged).
  3. Check whether the count is cross-verified: an explicit mean ± k*std (or quantile) cut-off printed, the flagged values listed and compared to that cut-off, and the count reproduced by a second, hand-written computation (including sample vs population denominator).
  4. Compare the reported count with the printed column summary (min/max, mean, std, value counts) — if max and min are within k standard deviations of the mean, any nonzero count is self-contradictory.
Discriminator
A fine attempt shows the derived numeric cut-off, the extreme values, and a matching hand-computed count (or a documented reason for a nonzero/zero result); a violation reports only the library-produced count, or reports a count that its own printed distribution summary cannot support, or applies the rule to a column whose dtype/missing-value handling was never checked.
Consequence
The submitted integer differs from the ground-truth count (e.g., a large nonzero count where the correct answer is zero, or vice versa), and the derived cleaned dataframe has the wrong number of rows, so the single-value check fails.
id 2b0ddebd48e2 · mined from infiagent-dabench dabench-361@s13
raw text (what the judge reads)
### Anomaly/threshold counts reported without an independent recomputation and plausibility check
- **Applies when**: `task` -- the script flags rows by a statistical rule (z-score, IQR, quantile cut, threshold filter) on one column and the requested answer is the *count* of flagged rows.
- **Pattern**: The attempt relies on a single library call (with NaN/dtype policies it never verifies) to build the boolean mask, prints the count, and submits it — without recomputing the statistic by hand from the printed mean/std (or quantiles), without checking that the column is clean numeric (no sentinel/placeholder codes, no strings coerced or silently dropped), and without confirming that the flagged values actually lie beyond the implied cut-off and that mask length still matches the dataframe length.
- **Detection procedure**:
  1. Read the task for the exact rule and threshold, and note which column and which rows it must be applied to.
  2. In the scripts, check whether the column is validated as numeric with missing/placeholder values explicitly handled, and whether the mask is aligned to the full dataframe (a NaN-omitting helper can return a shorter or NaN-containing array that silently changes which rows are flagged).
  3. Check whether the count is cross-verified: an explicit `mean ± k*std` (or quantile) cut-off printed, the flagged values listed and compared to that cut-off, and the count reproduced by a second, hand-written computation (including sample vs population denominator).
  4. Compare the reported count with the printed column summary (min/max, mean, std, value counts) — if max and min are within k standard deviations of the mean, any nonzero count is self-contradictory.
- **Discriminator**: A fine attempt shows the derived numeric cut-off, the extreme values, and a matching hand-computed count (or a documented reason for a nonzero/zero result); a violation reports only the library-produced count, or reports a count that its own printed distribution summary cannot support, or applies the rule to a column whose dtype/missing-value handling was never checked.
- **Consequence**: The submitted integer differs from the ground-truth count (e.g., a large nonzero count where the correct answer is zero, or vice versa), and the derived cleaned dataframe has the wrong number of rows, so the single-value check fails.
712Ignoring a task-referenced auxiliary spec (mapping/instruction file) before computing the statistictaskda-code
Applies when
task -- The prompt points to a companion file (notes, tips, data dictionary, config) that defines a required label mapping, recoding, filter, or naming convention to apply before the requested statistic is computed and reported.
Pattern
The script loads only the primary data table and computes the statistic directly on raw/encoded values, never opening or applying the referenced spec file; the reported category label is therefore the raw code rather than the mapped value (and any file-based deliverable implied by the task is skipped).
Detection procedure
  1. Read the task and list every external artifact it names, plus every required transformation and every required output (value(s), label vocabulary, rounding, file such as a results JSON).
  2. Scan the scripts for a read/open/parse of each named artifact and for code that applies its content (a dict/replace/map/merge step); note absence.
  3. Compare the answer's category name against the vocabulary the spec would produce — if it matches the raw column's encoding rather than the mapped names, the mapping was skipped.
  4. Confirm the script writes any required output file, not just prints to stdout.
Discriminator
A real violation is when the spec's content would change the reported value or label (or the deliverable is missing). It is fine if the agent read the spec and demonstrably applied it (mapping visible in code/output), or if the spec provably leaves labels and counts identical and the answer states so.
Consequence
The graded answer's label string (and possibly the counts/ratio after merging categories) mismatches the expected mapped value, and the expected result file is missing — scoring 0 despite arithmetically correct value counts.
id 1c336a6ca95d · mined from da-code dacode-di-text-004@s13
raw text (what the judge reads)
### Ignoring a task-referenced auxiliary spec (mapping/instruction file) before computing the statistic
- **Applies when**: `task` -- The prompt points to a companion file (notes, tips, data dictionary, config) that defines a required label mapping, recoding, filter, or naming convention to apply before the requested statistic is computed and reported.
- **Pattern**: The script loads only the primary data table and computes the statistic directly on raw/encoded values, never opening or applying the referenced spec file; the reported category label is therefore the raw code rather than the mapped value (and any file-based deliverable implied by the task is skipped).
- **Detection procedure**:
  1. Read the task and list every external artifact it names, plus every required transformation and every required output (value(s), label vocabulary, rounding, file such as a results JSON).
  2. Scan the scripts for a read/open/parse of each named artifact and for code that applies its content (a dict/replace/map/merge step); note absence.
  3. Compare the answer's category name against the vocabulary the spec would produce — if it matches the raw column's encoding rather than the mapped names, the mapping was skipped.
  4. Confirm the script writes any required output file, not just prints to stdout.
- **Discriminator**: A real violation is when the spec's content would change the reported value or label (or the deliverable is missing). It is fine if the agent read the spec and demonstrably applied it (mapping visible in code/output), or if the spec provably leaves labels and counts identical and the answer states so.
- **Consequence**: The graded answer's label string (and possibly the counts/ratio after merging categories) mismatches the expected mapped value, and the expected result file is missing — scoring 0 despite arithmetically correct value counts.
713Hand-crafted heuristic rules instead of a model fit to the provided labeled training split, with no row-aligned output filetaskda-code
Applies when
task -- the task supplies a labeled training file and an unlabeled test file and asks for predictions written to a specific output file with a specific column name.
Pattern
The attempt never loads the labeled training file (or never uses the target column to fit/validate anything); instead it invents threshold rules over auxiliary/side tables, applies them to a self-assembled set of records, and reports a prose summary of the rules — often without producing an output file whose row count and order match the test file, or whose column name matches the requested one.
Detection procedure
  1. From the task, note the required inputs (train file with labels, test file), the required output filename, column name, and the expected number of prediction rows (= test rows).
  2. Scan the scripts for a read of the training file and a fit/train/predict step (or at minimum a rule calibrated against the training labels) plus any held-out validation score; flag if the target column is never read.
  3. Check that the final artifact is written to the exact requested path/column and that the code asserts len(predictions) == len(test) and preserves test row order/ID join; flag if predictions are built from unrelated tables or aggregated to a different record count.
  4. Check the answer text: if it describes rules/thresholds and feature counts but reports no validation accuracy and no confirmation that the output file exists with correct shape and label vocabulary, treat as inadequate.
Discriminator
A simple rule-based or majority-class baseline is acceptable if it is calibrated/scored against the training labels, uses the same label strings as the training target, and is emitted as a correctly named, correctly sized, test-aligned file; the violation is using uncalibrated invented thresholds, unverified label spellings, and a record set whose count differs from the test set.
Consequence
The grader finds the expected result file missing or unmatched (wrong row count, wrong/absent column name, unseen label values), so the submission scores 0 regardless of how plausible the heuristics sound.
id fb9aaa54f6f9 · mined from da-code dacode-ml-multi-003@s13
raw text (what the judge reads)
### Hand-crafted heuristic rules instead of a model fit to the provided labeled training split, with no row-aligned output file
- **Applies when**: `task` -- the task supplies a labeled training file and an unlabeled test file and asks for predictions written to a specific output file with a specific column name.
- **Pattern**: The attempt never loads the labeled training file (or never uses the target column to fit/validate anything); instead it invents threshold rules over auxiliary/side tables, applies them to a self-assembled set of records, and reports a prose summary of the rules — often without producing an output file whose row count and order match the test file, or whose column name matches the requested one.
- **Detection procedure**:
  1. From the task, note the required inputs (train file with labels, test file), the required output filename, column name, and the expected number of prediction rows (= test rows).
  2. Scan the scripts for a read of the training file and a fit/`train`/`predict` step (or at minimum a rule calibrated against the training labels) plus any held-out validation score; flag if the target column is never read.
  3. Check that the final artifact is written to the exact requested path/column and that the code asserts `len(predictions) == len(test)` and preserves test row order/ID join; flag if predictions are built from unrelated tables or aggregated to a different record count.
  4. Check the answer text: if it describes rules/thresholds and feature counts but reports no validation accuracy and no confirmation that the output file exists with correct shape and label vocabulary, treat as inadequate.
- **Discriminator**: A simple rule-based or majority-class baseline is acceptable *if* it is calibrated/scored against the training labels, uses the same label strings as the training target, and is emitted as a correctly named, correctly sized, test-aligned file; the violation is using uncalibrated invented thresholds, unverified label spellings, and a record set whose count differs from the test set.
- **Consequence**: The grader finds the expected result file missing or unmatched (wrong row count, wrong/absent column name, unseen label values), so the submission scores 0 regardless of how plausible the heuristics sound.
714Missingness-based grouping defined without verifying the mask (partition doesn't cover the full table)taskinfiagent-dabench
Applies when
task -- the task asks to split rows into groups by whether a column is null/missing (or otherwise filter rows by a condition) and compare a statistic between the groups.
Pattern
The attempt takes the loader's default notion of "missing" (e.g. isna() after a read with default settings) without checking how missingness is actually encoded in the file — empty strings, sentinel tokens like NA/none/-/0, whitespace, or a column parsed as object/string — and/or silently drops rows when coercing the numeric column, so the two groups don't partition the original row set and both group means shift.
Detection procedure
  1. In the task, note the exact grouping condition and which column supplies the measured statistic.
  2. In the scripts, find how the file is read (separator, na_values, keep_default_na, dtype) and how the mask is built; check whether the value column is numerically coerced and whether any dropna/astype/filter happens before grouping.
  3. Check the script prints group sizes and confirms n_group1 + n_group2 == len(raw_df), plus a peek at the distinct raw values of the grouping column to confirm which tokens count as missing.
  4. Compare reported means/counts against a quick independent recomputation or plausibility check; unexplained shifts in both group means indicate rows moved between or out of the groups.
Discriminator
A real violation is when the mask/coercion is never validated (no counts printed, no inspection of raw missingness tokens) or the group sizes don't sum to the full row count; it is fine if the script explicitly documents the missingness encoding, shows the counts summing to the total, and only then computes the statistics.
Consequence
Both group means (and the test statistic) are computed on subtly wrong subsets, so the numeric answers differ from ground truth beyond rounding and the grader marks the value checks wrong even though the reported p-value direction looks plausible.
id 0fe1074a42d3 · mined from infiagent-dabench dabench-297@s13
raw text (what the judge reads)
### Missingness-based grouping defined without verifying the mask (partition doesn't cover the full table)

- **Applies when**: `task` -- the task asks to split rows into groups by whether a column is null/missing (or otherwise filter rows by a condition) and compare a statistic between the groups.
- **Pattern**: The attempt takes the loader's default notion of "missing" (e.g. `isna()` after a read with default settings) without checking how missingness is actually encoded in the file — empty strings, sentinel tokens like `NA`/`none`/`-`/`0`, whitespace, or a column parsed as object/string — and/or silently drops rows when coercing the numeric column, so the two groups don't partition the original row set and both group means shift.
- **Detection procedure**:
  1. In the task, note the exact grouping condition and which column supplies the measured statistic.
  2. In the scripts, find how the file is read (separator, `na_values`, `keep_default_na`, dtype) and how the mask is built; check whether the value column is numerically coerced and whether any `dropna`/`astype`/filter happens before grouping.
  3. Check the script prints group sizes and confirms `n_group1 + n_group2 == len(raw_df)`, plus a peek at the distinct raw values of the grouping column to confirm which tokens count as missing.
  4. Compare reported means/counts against a quick independent recomputation or plausibility check; unexplained shifts in *both* group means indicate rows moved between or out of the groups.
- **Discriminator**: A real violation is when the mask/coercion is never validated (no counts printed, no inspection of raw missingness tokens) or the group sizes don't sum to the full row count; it is fine if the script explicitly documents the missingness encoding, shows the counts summing to the total, and only then computes the statistics.
- **Consequence**: Both group means (and the test statistic) are computed on subtly wrong subsets, so the numeric answers differ from ground truth beyond rounding and the grader marks the value checks wrong even though the reported p-value direction looks plausible.
715Spec file referenced by the task is never read, so rules and output artifacts are inventedtaskda-code
Applies when
task -- The task points to an external instruction/guidance/config file (or a described protocol) that defines the preprocessing rules, derived categories, and the set of deliverable output files.
Pattern
The script never loads, prints, or quotes that spec; instead it hard-codes filters, groupings/thresholds, and saves only the one artifact mentioned in the task prompt, leaving other required outputs (serialized numeric results, plot-data dumps, tables) unwritten. The final answer restates the invented rules as if they were the given ones.
Detection procedure
  1. From the task text, list every referenced spec source and every deliverable file/format implied (image, array dump, JSON of plot data, printed values, rounding/ordering rules).
  2. Grep the scripts for reads of the spec file and for write calls (savefig, save, to_json, to_csv, dump) — build the set of artifacts actually produced.
  3. Compare sets: any deliverable with no corresponding write call, and any rule (filters, category definitions, thresholds) that appears in the script but has no traceable origin in the task/spec, is a violation.
  4. Check the answer text: does it document rules as "applied per guidance" without evidence the guidance was consulted?
Discriminator
Fine if the script demonstrably reads/echoes the spec, or if the task text itself fully enumerates the rules and every artifact is written; a real violation is inventing thresholds/category logic by guesswork and producing a strict subset of the required files.
Consequence
Graders that check for each expected artifact mark the missing ones as WRONG/MISSING, and the one produced file also mismatches because the underlying filtering/grouping differs from the specified procedure — near-zero score regardless of how coherent the narrative answer looks.
id 27c414e1d478 · mined from da-code dacode-plot-pie-005@s13
raw text (what the judge reads)
### Spec file referenced by the task is never read, so rules and output artifacts are invented
- **Applies when**: `task` -- The task points to an external instruction/guidance/config file (or a described protocol) that defines the preprocessing rules, derived categories, and the set of deliverable output files.
- **Pattern**: The script never loads, prints, or quotes that spec; instead it hard-codes filters, groupings/thresholds, and saves only the one artifact mentioned in the task prompt, leaving other required outputs (serialized numeric results, plot-data dumps, tables) unwritten. The final answer restates the invented rules as if they were the given ones.
- **Detection procedure**:
  1. From the task text, list every referenced spec source and every deliverable file/format implied (image, array dump, JSON of plot data, printed values, rounding/ordering rules).
  2. Grep the scripts for reads of the spec file and for write calls (`savefig`, `save`, `to_json`, `to_csv`, `dump`) — build the set of artifacts actually produced.
  3. Compare sets: any deliverable with no corresponding write call, and any rule (filters, category definitions, thresholds) that appears in the script but has no traceable origin in the task/spec, is a violation.
  4. Check the answer text: does it document rules as "applied per guidance" without evidence the guidance was consulted?
- **Discriminator**: Fine if the script demonstrably reads/echoes the spec, or if the task text itself fully enumerates the rules and every artifact is written; a real violation is inventing thresholds/category logic by guesswork and producing a strict subset of the required files.
- **Consequence**: Graders that check for each expected artifact mark the missing ones as WRONG/MISSING, and the one produced file also mismatches because the underlying filtering/grouping differs from the specified procedure — near-zero score regardless of how coherent the narrative answer looks.
716Missing/non-numeric handling that silently changes the modeling rows or target scaletaskinfiagent-dabench
Applies when
task -- the task prescribes an exact preprocessing recipe (e.g., impute specified columns with column means), a fixed train/test split, and a scale-dependent error metric on a numeric target.
Pattern
The attempt loads the raw columns and either drops rows with nulls, coerces dirty/string-formatted numeric fields in a way that nulls out or mis-scales values, imputes only some of the named columns, or imputes/derives statistics after (or inconsistently across) the split — so the fitted rows, target distribution, or feature units differ from the specified recipe, and the reported error is off by an order of magnitude with no sanity check.
Detection procedure
  1. Read the task and list every mandated preprocessing step (which columns, which fill statistic, which split ratio, which metric) and the exact reporting format/rounding.
  2. In the script, trace each named column from load to model input: check dtype coercion, that fillna(mean) (not dropna, median, zero, or interpolation) is applied to every named column, that no additional filtering/deduplication/outlier removal shrinks the frame, and that row count before modeling equals the raw row count.
  3. Check that the split uses the stated fraction on the full imputed frame and that the metric is computed on the held-out predictions only, with the target in its original units (no scaling/log transform left un-inverted).
  4. Compare the reported error against a cheap baseline the script should have printed: variance of the target (MSE of predicting the mean). An MSE near or above target variance, or wildly larger than a same-recipe rerun, signals broken preprocessing.
Discriminator
A real violation is a deviation from the mandated recipe or an unhandled dirty-dtype that changes which rows/values enter the model; a look-alike that is fine is a script that follows the recipe exactly and merely differs from a reference by random-seed/split ordering, which shifts the metric slightly rather than by a large factor.
Consequence
The reported error is far from the reference value (here roughly an order of magnitude too large), so the single-value check fails even though the code runs without error.
id 2f17b1178927 · mined from infiagent-dabench dabench-432@s13
raw text (what the judge reads)
### Missing/non-numeric handling that silently changes the modeling rows or target scale
- **Applies when**: `task` -- the task prescribes an exact preprocessing recipe (e.g., impute specified columns with column means), a fixed train/test split, and a scale-dependent error metric on a numeric target.
- **Pattern**: The attempt loads the raw columns and either drops rows with nulls, coerces dirty/string-formatted numeric fields in a way that nulls out or mis-scales values, imputes only some of the named columns, or imputes/derives statistics after (or inconsistently across) the split — so the fitted rows, target distribution, or feature units differ from the specified recipe, and the reported error is off by an order of magnitude with no sanity check.
- **Detection procedure**:
  1. Read the task and list every mandated preprocessing step (which columns, which fill statistic, which split ratio, which metric) and the exact reporting format/rounding.
  2. In the script, trace each named column from load to model input: check dtype coercion, that `fillna(mean)` (not dropna, median, zero, or interpolation) is applied to *every* named column, that no additional filtering/deduplication/outlier removal shrinks the frame, and that row count before modeling equals the raw row count.
  3. Check that the split uses the stated fraction on the full imputed frame and that the metric is computed on the held-out predictions only, with the target in its original units (no scaling/log transform left un-inverted).
  4. Compare the reported error against a cheap baseline the script should have printed: variance of the target (MSE of predicting the mean). An MSE near or above target variance, or wildly larger than a same-recipe rerun, signals broken preprocessing.
- **Discriminator**: A real violation is a deviation from the mandated recipe or an unhandled dirty-dtype that changes which rows/values enter the model; a look-alike that is fine is a script that follows the recipe exactly and merely differs from a reference by random-seed/split ordering, which shifts the metric slightly rather than by a large factor.
- **Consequence**: The reported error is far from the reference value (here roughly an order of magnitude too large), so the single-value check fails even though the code runs without error.
717Reporting near-zero/degenerate summary statistics after a scaling step without validating the scaling conventiontaskinfiagent-dabench
Applies when
task -- the task asks to normalize/scale numeric columns and then report a summary statistic (e.g., mean) of the transformed columns, and the scripts choose a scaling method that the task does not pin down explicitly.
Pattern
The attempt applies zero-centering standardization (or another transform whose reported statistic is trivially constant, e.g., mean ≈ 0) and then reports that mathematically-forced value for every scaled column, instead of using the scaling convention implied by the word "normalize" (min–max to [0,1]) which yields informative, column-specific values; it also leaves untransformed/encoded columns on an inconsistent scale.
Detection procedure
  1. Read the task: note which columns must be scaled, which must be encoded, and which statistic is reported afterwards.
  2. Read the scripts: identify the exact transform applied (StandardScaler/z-score vs MinMaxScaler vs custom) and check whether the requested statistic becomes degenerate (identical or ~0 for all columns) under that transform.
  3. Inspect the reported numbers: if every scaled column's statistic is the same constant (0.0, 1.0) while non-scaled/binary columns have varied values, treat it as a red flag that the transform, not the data, produced the answer.
  4. Cross-check plausibility: for [0,1] scaling the mean must lie strictly inside (0,1) and differ across columns with different skew; confirm the scripts would produce such values and that the untouched columns' handling matches the stated constraints.
Discriminator
A real violation is when the reported statistic is an artifact of the transform (constant across all scaled columns, carrying no data information) or when the transform contradicts the task's stated wording/expected range. It is fine if the task explicitly names standardization, or if the reported statistic is genuinely informative under the chosen transform (e.g., reporting std, or means that vary per column and lie in a sensible range).
Consequence
All scaled columns' reported values collapse to the same constant and mismatch the ground truth, so only the binary/encoded columns pass and the answer is graded incorrect.
id 5cfb96a8765b · mined from infiagent-dabench dabench-28@s13
raw text (what the judge reads)
### Reporting near-zero/degenerate summary statistics after a scaling step without validating the scaling convention
- **Applies when**: `task` -- the task asks to normalize/scale numeric columns and then report a summary statistic (e.g., mean) of the transformed columns, and the scripts choose a scaling method that the task does not pin down explicitly.
- **Pattern**: The attempt applies zero-centering standardization (or another transform whose reported statistic is trivially constant, e.g., mean ≈ 0) and then reports that mathematically-forced value for every scaled column, instead of using the scaling convention implied by the word "normalize" (min–max to [0,1]) which yields informative, column-specific values; it also leaves untransformed/encoded columns on an inconsistent scale.
- **Detection procedure**:
  1. Read the task: note which columns must be scaled, which must be encoded, and which statistic is reported afterwards.
  2. Read the scripts: identify the exact transform applied (StandardScaler/z-score vs MinMaxScaler vs custom) and check whether the requested statistic becomes degenerate (identical or ~0 for all columns) under that transform.
  3. Inspect the reported numbers: if every scaled column's statistic is the same constant (0.0, 1.0) while non-scaled/binary columns have varied values, treat it as a red flag that the transform, not the data, produced the answer.
  4. Cross-check plausibility: for [0,1] scaling the mean must lie strictly inside (0,1) and differ across columns with different skew; confirm the scripts would produce such values and that the untouched columns' handling matches the stated constraints.
- **Discriminator**: A real violation is when the reported statistic is an artifact of the transform (constant across all scaled columns, carrying no data information) or when the transform contradicts the task's stated wording/expected range. It is fine if the task explicitly names standardization, or if the reported statistic is genuinely informative under the chosen transform (e.g., reporting std, or means that vary per column and lie in a sensible range).
- **Consequence**: All scaled columns' reported values collapse to the same constant and mismatch the ground truth, so only the binary/encoded columns pass and the answer is graded incorrect.
718Model quality claimed from in-sample fit, with no held-out validation or prediction-distribution sanity checktaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test set and the scripts train a flexible model (tree ensemble, boosted trees, deep net) and report accuracy numbers.
Pattern
The attempt fits the model on the full labeled data, then computes R²/RMSE/MAE by predicting on those same rows, quotes the near-perfect in-sample score as proof of quality, and never (a) evaluates on a held-out split or cross-validation, nor (b) compares the predicted-target distribution on the test set against the observed target distribution in training. Encoding/preprocessing choices (e.g., arbitrary integer encoding of high-cardinality categories, unseen-category handling, skewed target left untransformed) therefore go untested, and gross distribution shifts in the output pass unnoticed.
Detection procedure
  1. Read the task to confirm the deliverable is out-of-sample predictions, so reported metrics must estimate generalization.
  2. In the scripts, check whether the rows used to compute the reported metric are disjoint from the rows used to fit; a metric computed on model.predict(X_train) after model.fit(X_train, y_train), or any absence of a split/CV, is the violation.
  3. In the answer, compare the reported summary statistics of the predictions (min, max, mean, median) with the same statistics of the training target; flag large discrepancies (e.g., mean far above median plus a maximum far outside the training range, or predicted mean not close to training mean).
  4. Check that categorical encoders/imputers are fit on training data and applied consistently to test rows, including a defined path for categories unseen at fit time.
  5. Confirm the output file's row count, column name, and ordering match the test set and the requested format.
Discriminator
A real violation is when the only quantitative evidence is an in-sample score, or the prediction distribution is implausible relative to the training target; it is fine if a holdout/CV score is reported (even a mediocre one) and the prediction summary statistics roughly track the training target, even if the model is simple or the in-sample score is also mentioned.
Consequence
The submitted predictions score poorly against the hidden labels (high RMSE/low R²) despite the reported "R² ≈ 0.99", and the file fails the grader's accuracy threshold even though its shape and column name look correct.
id dd10d90b8b03 · mined from da-code dacode-ml-regression-014@s13
raw text (what the judge reads)
### Model quality claimed from in-sample fit, with no held-out validation or prediction-distribution sanity check
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test set and the scripts train a flexible model (tree ensemble, boosted trees, deep net) and report accuracy numbers.
- **Pattern**: The attempt fits the model on the full labeled data, then computes R²/RMSE/MAE by predicting on those same rows, quotes the near-perfect in-sample score as proof of quality, and never (a) evaluates on a held-out split or cross-validation, nor (b) compares the predicted-target distribution on the test set against the observed target distribution in training. Encoding/preprocessing choices (e.g., arbitrary integer encoding of high-cardinality categories, unseen-category handling, skewed target left untransformed) therefore go untested, and gross distribution shifts in the output pass unnoticed.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is out-of-sample predictions, so reported metrics must estimate generalization.
  2. In the scripts, check whether the rows used to compute the reported metric are disjoint from the rows used to fit; a metric computed on `model.predict(X_train)` after `model.fit(X_train, y_train)`, or any absence of a split/CV, is the violation.
  3. In the answer, compare the reported summary statistics of the predictions (min, max, mean, median) with the same statistics of the training target; flag large discrepancies (e.g., mean far above median plus a maximum far outside the training range, or predicted mean not close to training mean).
  4. Check that categorical encoders/imputers are fit on training data and applied consistently to test rows, including a defined path for categories unseen at fit time.
  5. Confirm the output file's row count, column name, and ordering match the test set and the requested format.
- **Discriminator**: A real violation is when the *only* quantitative evidence is an in-sample score, or the prediction distribution is implausible relative to the training target; it is fine if a holdout/CV score is reported (even a mediocre one) and the prediction summary statistics roughly track the training target, even if the model is simple or the in-sample score is also mentioned.
- **Consequence**: The submitted predictions score poorly against the hidden labels (high RMSE/low R²) despite the reported "R² ≈ 0.99", and the file fails the grader's accuracy threshold even though its shape and column name look correct.
719Missing or unverified required output artifacts specified by an external config/spectaskda-code
Applies when
task -- the task points to a separate configuration/spec file (e.g., a YAML/JSON of plotting or export guidelines) and/or expects a set of saved output files, and the agent's scripts write only some of them.
Pattern
The agent reads the spec loosely (or hard-codes styling it guessed), saves one visible artifact (the image) and reports numbers in prose, while never emitting the other required machine-checkable artifacts (serialized figure/metadata and numeric arrays) or verifying that saved values match the spec's keys, units, ordering, and rounding.
Detection procedure
  1. From the task text and the referenced spec file, enumerate every required output: exact filenames, extensions, and for each the required content (title text, figure size, colors, labels, category order, percentage vs count, rounding/units).
  2. In the scripts, list every write/save call (savefig, np.save, json.dump, to_csv) and match them one-to-one against that enumeration; flag any required file with no corresponding write, or any styling/label value that was typed in literally instead of being read from the spec.
  3. Check the answer/log for a post-save verification step (re-open each file, print shape/keys/values) and for internal consistency of the reported numbers (do category counts sum to the stated total? do percentages match the counts?).
  4. Flag the attempt if any required artifact is absent, or if reported figures are mutually inconsistent and unchecked.
Discriminator
A real violation is a required artifact never written, or spec values duplicated by hand and not cross-checked (e.g., percentages that cannot be derived from the reported counts). It is not a violation when all required files are written and verified and the agent additionally summarizes them in prose, or when extra non-required files are produced.
Consequence
The grader marks the missing files as WRONG/MISSING and the run scores zero on those checks even if the headline conclusion (the selected category/group) happens to be right.
id cedb09561aa4 · mined from da-code dacode-plot-pie-008@s13
raw text (what the judge reads)
### Missing or unverified required output artifacts specified by an external config/spec
- **Applies when**: `task` -- the task points to a separate configuration/spec file (e.g., a YAML/JSON of plotting or export guidelines) and/or expects a set of saved output files, and the agent's scripts write only some of them.
- **Pattern**: The agent reads the spec loosely (or hard-codes styling it guessed), saves one visible artifact (the image) and reports numbers in prose, while never emitting the other required machine-checkable artifacts (serialized figure/metadata and numeric arrays) or verifying that saved values match the spec's keys, units, ordering, and rounding.
- **Detection procedure**:
  1. From the task text and the referenced spec file, enumerate every required output: exact filenames, extensions, and for each the required content (title text, figure size, colors, labels, category order, percentage vs count, rounding/units).
  2. In the scripts, list every write/save call (`savefig`, `np.save`, `json.dump`, `to_csv`) and match them one-to-one against that enumeration; flag any required file with no corresponding write, or any styling/label value that was typed in literally instead of being read from the spec.
  3. Check the answer/log for a post-save verification step (re-open each file, print shape/keys/values) and for internal consistency of the reported numbers (do category counts sum to the stated total? do percentages match the counts?).
  4. Flag the attempt if any required artifact is absent, or if reported figures are mutually inconsistent and unchecked.
- **Discriminator**: A real violation is a required artifact never written, or spec values duplicated by hand and not cross-checked (e.g., percentages that cannot be derived from the reported counts). It is *not* a violation when all required files are written and verified and the agent additionally summarizes them in prose, or when extra non-required files are produced.
- **Consequence**: The grader marks the missing files as WRONG/MISSING and the run scores zero on those checks even if the headline conclusion (the selected category/group) happens to be right.
720Unvalidated derived variable parsed from raw/dirty source fieldstaskinfiagent-dabench
Applies when
task -- the analysis requires a quantity (duration, count, rate, category) that is not a column but must be derived by parsing/deriving from free-text, mixed-format, or explicitly error-containing source fields before a statistic is computed.
Pattern
The agent writes a regex/heuristic parser, spot-checks a handful of strings, and feeds the derived column straight into the statistic — never auditing the parsed values for impossible or implausible results (negatives, zeros, absurdly large values, silent NaNs from unmatched formats, out-of-range categories, duplicated/garbled records), and never reconciling how many rows were dropped or mis-parsed against the raw row count.
Detection procedure
  1. In the task/data description, note any signal that the source is dirty (error-flagged file, free-text fields, multiple formats, out-of-domain codes) and identify every variable the agent must construct rather than read directly.
  2. In the scripts, locate the parsing/derivation function and check whether failures fall through to NaN/except: return nan or default values, and whether the parsed and other analysis columns are subsequently range-checked (min/max, allowed category set, non-negative, count of nulls vs. rows used).
  3. Check whether rows with invalid or unparsed values are explicitly filtered/corrected before the statistic, and whether the final N used in the test is reported and consistent with the raw N minus documented exclusions.
  4. Compare the reported statistic against this: if no validation or cleaning step exists, the reported coefficient reflects contaminated inputs and should be treated as unverified.
Discriminator
A real violation is the absence of any post-parse validity check or outlier/invalid-record handling on a derived or error-prone column; it is not a violation if the agent prints distributional sanity checks (min/max, unique values, null counts) and either shows all values are plausible or explicitly drops/repairs the offending rows with a stated rule.
Consequence
The statistic is computed on a subtly contaminated subset, so the point estimate (e.g., a correlation to two decimals) drifts from ground truth even when the qualitative conclusion and p-value checks pass, producing a partially-correct answer graded as wrong.
id 1cadbec072af · mined from infiagent-dabench dabench-431@s13
raw text (what the judge reads)
### Unvalidated derived variable parsed from raw/dirty source fields
- **Applies when**: `task` -- the analysis requires a quantity (duration, count, rate, category) that is not a column but must be derived by parsing/deriving from free-text, mixed-format, or explicitly error-containing source fields before a statistic is computed.
- **Pattern**: The agent writes a regex/heuristic parser, spot-checks a handful of strings, and feeds the derived column straight into the statistic — never auditing the parsed values for impossible or implausible results (negatives, zeros, absurdly large values, silent `NaN`s from unmatched formats, out-of-range categories, duplicated/garbled records), and never reconciling how many rows were dropped or mis-parsed against the raw row count.
- **Detection procedure**:
  1. In the task/data description, note any signal that the source is dirty (error-flagged file, free-text fields, multiple formats, out-of-domain codes) and identify every variable the agent must construct rather than read directly.
  2. In the scripts, locate the parsing/derivation function and check whether failures fall through to `NaN`/`except: return nan` or default values, and whether the parsed and other analysis columns are subsequently range-checked (min/max, allowed category set, non-negative, count of nulls vs. rows used).
  3. Check whether rows with invalid or unparsed values are explicitly filtered/corrected before the statistic, and whether the final N used in the test is reported and consistent with the raw N minus documented exclusions.
  4. Compare the reported statistic against this: if no validation or cleaning step exists, the reported coefficient reflects contaminated inputs and should be treated as unverified.
- **Discriminator**: A real violation is the absence of any post-parse validity check or outlier/invalid-record handling on a derived or error-prone column; it is *not* a violation if the agent prints distributional sanity checks (min/max, unique values, null counts) and either shows all values are plausible or explicitly drops/repairs the offending rows with a stated rule.
- **Consequence**: The statistic is computed on a subtly contaminated subset, so the point estimate (e.g., a correlation to two decimals) drifts from ground truth even when the qualitative conclusion and p-value checks pass, producing a partially-correct answer graded as wrong.
721Output artifact not verified against the exact requested schemataskda-code
Applies when
task -- the task requires writing results to a named file with a specified column/format, and the scripts generate that file at the end of a modeling pipeline.
Pattern
The attempt builds a reasonable model but writes the output file with a schema that deviates from the literal specification (extra index/identifier columns, renamed or differently-cased column, wrong row count/order, probabilities instead of the requested label type, index written as an unnamed column), and never re-reads the saved file to confirm it matches the spec; the summary asserts success from in-memory state rather than from the file on disk.
Detection procedure
  1. From the task text, list the exact required artifact name, required column name(s) and nothing more, expected value type/domain, and the expected number and order of rows (typically one per test record, in input order).
  2. In the scripts, find the write call (e.g., to_csv) and check the exact DataFrame columns passed, whether the index is suppressed, the dtype/values written, and that the frame was built from the untouched test set (no reordering, no dropped/filtered rows).
  3. Check whether the script (or a follow-up step) reloads the file and asserts shape, column names, row count and value domain; absence of any such verification is a red flag.
  4. Compare the answer's claims about the file (columns, row count, class balance) with what the code actually writes; mismatch or unverifiable claims count as a violation.
Discriminator
A real violation is a concrete deviation from the stated schema (extra/missing/misnamed columns, index leaked as a column, row count or order not matching the test input, wrong value type) or a total absence of post-write validation; it is not a violation if the file contains exactly the requested column(s) with one aligned row per test record and the script or answer demonstrates a read-back sanity check, even if the model itself is simple or accuracy is modest.
Consequence
The grader loads the artifact, fails to find the expected column/shape or cannot align rows to the ground truth, and marks the required file WRONG/MISSING regardless of how good the underlying model was.
id 841d1c974734 · mined from da-code dacode-ml-binary-016@s13
raw text (what the judge reads)
### Output artifact not verified against the exact requested schema
- **Applies when**: `task` -- the task requires writing results to a named file with a specified column/format, and the scripts generate that file at the end of a modeling pipeline.
- **Pattern**: The attempt builds a reasonable model but writes the output file with a schema that deviates from the literal specification (extra index/identifier columns, renamed or differently-cased column, wrong row count/order, probabilities instead of the requested label type, index written as an unnamed column), and never re-reads the saved file to confirm it matches the spec; the summary asserts success from in-memory state rather than from the file on disk.
- **Detection procedure**:
  1. From the task text, list the exact required artifact name, required column name(s) and nothing more, expected value type/domain, and the expected number and order of rows (typically one per test record, in input order).
  2. In the scripts, find the write call (e.g., `to_csv`) and check the exact DataFrame columns passed, whether the index is suppressed, the dtype/values written, and that the frame was built from the untouched test set (no reordering, no dropped/filtered rows).
  3. Check whether the script (or a follow-up step) reloads the file and asserts shape, column names, row count and value domain; absence of any such verification is a red flag.
  4. Compare the answer's claims about the file (columns, row count, class balance) with what the code actually writes; mismatch or unverifiable claims count as a violation.
- **Discriminator**: A real violation is a concrete deviation from the stated schema (extra/missing/misnamed columns, index leaked as a column, row count or order not matching the test input, wrong value type) or a total absence of post-write validation; it is *not* a violation if the file contains exactly the requested column(s) with one aligned row per test record and the script or answer demonstrates a read-back sanity check, even if the model itself is simple or accuracy is modest.
- **Consequence**: The grader loads the artifact, fails to find the expected column/shape or cannot align rows to the ground truth, and marks the required file WRONG/MISSING regardless of how good the underlying model was.
722Output file not built from the provided template/example schemataskda-code
Applies when
task -- the task says to save results "in the format of" a provided sample/template file (e.g., sample_result.csv) to a specific output filename.
Pattern
The agent computes a plausible number, prints it in prose, and writes a self-invented CSV (own column names, extra descriptive rows, different row/column orientation, added precision or units) without ever reading the sample file to copy its exact header, row labels, ordering, and value formatting — so the graded artifact mismatches even when the statistic is arguably right.
Detection procedure
  1. In the task text, note the named template file and the required output filename/location.
  2. In the scripts, search for any read of the template (e.g., loading sample_result.*) and a comparison of the produced object's columns/index/shape to it; absence of this is the red flag.
  3. Inspect the write call: confirm the written frame's column names, row labels/order, number of rows, and value rounding are derived from the template rather than hard-coded prose-style fields, and that index=/header options don't add or drop a column.
  4. Check the final answer: if it only narrates statistics and asserts "saved in the required format" without echoing the actual file contents, treat the format as unverified.
Discriminator
A real violation is inventing or reformatting the schema (renaming/adding fields, transposing, embedding text, extra decimals or units) without ever consulting the template; it is fine if the script loads the template, reuses its exact structure, and merely substitutes computed values — even if the printed narrative is verbose.
Consequence
The grader compares result.csv against the expected file and reports WRONG/MISSING on schema or cell mismatch, scoring 0 regardless of whether the underlying statistic was computed correctly.
id 7b6fc8a804e0 · mined from da-code dacode-data-sa-039@s13
raw text (what the judge reads)
### Output file not built from the provided template/example schema
- **Applies when**: `task` -- the task says to save results "in the format of" a provided sample/template file (e.g., `sample_result.csv`) to a specific output filename.
- **Pattern**: The agent computes a plausible number, prints it in prose, and writes a self-invented CSV (own column names, extra descriptive rows, different row/column orientation, added precision or units) without ever reading the sample file to copy its exact header, row labels, ordering, and value formatting — so the graded artifact mismatches even when the statistic is arguably right.
- **Detection procedure**:
  1. In the task text, note the named template file and the required output filename/location.
  2. In the scripts, search for any read of the template (e.g., loading `sample_result.*`) and a comparison of the produced object's columns/index/shape to it; absence of this is the red flag.
  3. Inspect the write call: confirm the written frame's column names, row labels/order, number of rows, and value rounding are derived from the template rather than hard-coded prose-style fields, and that `index=`/header options don't add or drop a column.
  4. Check the final answer: if it only narrates statistics and asserts "saved in the required format" without echoing the actual file contents, treat the format as unverified.
- **Discriminator**: A real violation is inventing or reformatting the schema (renaming/adding fields, transposing, embedding text, extra decimals or units) without ever consulting the template; it is fine if the script loads the template, reuses its exact structure, and merely substitutes computed values — even if the printed narrative is verbose.
- **Consequence**: The grader compares `result.csv` against the expected file and reports WRONG/MISSING on schema or cell mismatch, scoring 0 regardless of whether the underlying statistic was computed correctly.
723Sorting/extremes computed on a column that was never verified to be truly numeric after cleaningtaskda-code
Applies when
task -- the answer is an argmax/argmin (or top/bottom ranking) over a column from a raw CSV whose values may carry formatting (thousands separators, %, currency symbols, units) or be stored as text, possibly after an imputation step.
Pattern
The attempt loads the file, imputes with a summary statistic, and takes idxmax()/idxmin() (or sort_values().head()) without ever coercing the target column to a numeric dtype and checking the coercion succeeded; text values sort lexicographically or partially-parsed values silently drop digits, so the reported extremes are plausible-looking but wrong, and no result file / sanity check is produced.
Detection procedure
  1. In the task, identify exactly which column drives the ranking and what unit/scale its extremes should plausibly have.
  2. In the scripts, look for an explicit cleaning + astype(float)/pd.to_numeric(..., errors='coerce') step on that column, plus a check that the dtype is numeric and that the count of NaNs introduced by coercion is reported or zero; also confirm the mean-imputation happens on the numeric version, not the string version.
  3. Check for any post-hoc validation: printing the top-k and bottom-k rows with their values, and confirming the extreme values fall in a physically sensible range.
  4. Compare the reported extremes against that range/expectation and confirm the answer was written to the requested output file in the requested key/format.
Discriminator
A real violation is when no dtype coercion/verification exists, or the printed extreme value is absent/implausible for the quantity's units; it is fine if the script coerces the column, reports how many values failed to parse, and shows extreme values consistent with the quantity's plausible range — even if the cleaning is done inline in one line.
Consequence
The grader sees an entry that is not the true maximum/minimum (a mid-range or lexicographically-first row) and marks the answer wrong, or finds the expected result file missing entirely.
id 6d48f25497bb · mined from da-code dacode-di-text-001@s13
raw text (what the judge reads)
### Sorting/extremes computed on a column that was never verified to be truly numeric after cleaning
- **Applies when**: `task` -- the answer is an argmax/argmin (or top/bottom ranking) over a column from a raw CSV whose values may carry formatting (thousands separators, %, currency symbols, units) or be stored as text, possibly after an imputation step.
- **Pattern**: The attempt loads the file, imputes with a summary statistic, and takes `idxmax()/idxmin()` (or `sort_values().head()`) without ever coercing the target column to a numeric dtype and checking the coercion succeeded; text values sort lexicographically or partially-parsed values silently drop digits, so the reported extremes are plausible-looking but wrong, and no result file / sanity check is produced.
- **Detection procedure**:
  1. In the task, identify exactly which column drives the ranking and what unit/scale its extremes should plausibly have.
  2. In the scripts, look for an explicit cleaning + `astype(float)`/`pd.to_numeric(..., errors='coerce')` step on that column, plus a check that the dtype is numeric and that the count of NaNs introduced by coercion is reported or zero; also confirm the mean-imputation happens on the numeric version, not the string version.
  3. Check for any post-hoc validation: printing the top-k and bottom-k rows with their values, and confirming the extreme values fall in a physically sensible range.
  4. Compare the reported extremes against that range/expectation and confirm the answer was written to the requested output file in the requested key/format.
- **Discriminator**: A real violation is when no dtype coercion/verification exists, or the printed extreme value is absent/implausible for the quantity's units; it is fine if the script coerces the column, reports how many values failed to parse, and shows extreme values consistent with the quantity's plausible range — even if the cleaning is done inline in one line.
- **Consequence**: The grader sees an entry that is not the true maximum/minimum (a mid-range or lexicographically-first row) and marks the answer wrong, or finds the expected result file missing entirely.
724Ambiguous-but-conventional definitions chosen arbitrarily and never validatedtaskda-code
Applies when
task -- the task asks for a well-known standard analysis (scores, segments, aggregated per-entity metrics) whose exact definitions (aggregation unit, filters, cut points, label thresholds) are conventional but not spelled out in the prompt.
Pattern
The script silently picks one of several plausible definitions for each ingredient — e.g. counting raw rows instead of distinct transactions/events per entity, dropping or keeping negative/return/missing-key records, hand-chosen bucket thresholds for the final labels, an invented reference date — and never checks the resulting per-entity table against any external reference, documented convention, or sanity target (expected number of entities, expected distribution of labels, plausible ranges).
Detection procedure
  1. From the task text, list each definitional choice the prompt leaves open (unit of aggregation, filtering rules, binning scheme, label cutoffs, output columns).
  2. In the scripts, find the line implementing each choice and check whether any justification, alternative comparison, or validation accompanies it (e.g. counting unique IDs vs. rows; number of resulting entities vs. the number of distinct keys in the raw file).
  3. Check whether rows with a missing grouping key or non-positive values are handled explicitly, or just silently dropped by groupby/filters, and whether the surviving entity count is reported and reconciled against the raw data.
  4. Read the answer: does it merely restate its own outputs ("processed successfully", "calculated correctly") instead of comparing them to an independent expectation?
Discriminator
A fine attempt states the convention it follows, and cross-checks at least one derived quantity (entity count, per-bin counts, min/max of each metric) against the raw data or the standard formulation; a violation is when several defensible definitions exist, one is picked implicitly, and every "validation" is self-referential (counts of its own output).
Consequence
The saved file has per-entity metric values, score buckets, and level labels that differ systematically from the reference solution, so an exact/near-exact file comparison fails even though the pipeline runs without error.
id 2c7c586d5a1a · mined from da-code dacode-dm-csv-052@s13
raw text (what the judge reads)
### Ambiguous-but-conventional definitions chosen arbitrarily and never validated
- **Applies when**: `task` -- the task asks for a well-known standard analysis (scores, segments, aggregated per-entity metrics) whose exact definitions (aggregation unit, filters, cut points, label thresholds) are conventional but not spelled out in the prompt.
- **Pattern**: The script silently picks one of several plausible definitions for each ingredient — e.g. counting raw rows instead of distinct transactions/events per entity, dropping or keeping negative/return/missing-key records, hand-chosen bucket thresholds for the final labels, an invented reference date — and never checks the resulting per-entity table against any external reference, documented convention, or sanity target (expected number of entities, expected distribution of labels, plausible ranges).
- **Detection procedure**:
  1. From the task text, list each definitional choice the prompt leaves open (unit of aggregation, filtering rules, binning scheme, label cutoffs, output columns).
  2. In the scripts, find the line implementing each choice and check whether any justification, alternative comparison, or validation accompanies it (e.g. counting unique IDs vs. rows; number of resulting entities vs. the number of distinct keys in the raw file).
  3. Check whether rows with a missing grouping key or non-positive values are handled explicitly, or just silently dropped by `groupby`/filters, and whether the surviving entity count is reported and reconciled against the raw data.
  4. Read the answer: does it merely restate its own outputs ("processed successfully", "calculated correctly") instead of comparing them to an independent expectation?
- **Discriminator**: A fine attempt states the convention it follows, and cross-checks at least one derived quantity (entity count, per-bin counts, min/max of each metric) against the raw data or the standard formulation; a violation is when several defensible definitions exist, one is picked implicitly, and every "validation" is self-referential (counts of its own output).
- **Consequence**: The saved file has per-entity metric values, score buckets, and level labels that differ systematically from the reference solution, so an exact/near-exact file comparison fails even though the pipeline runs without error.
725Bootstrap interval taken from the spread of raw values instead of the spread of the resampled statistictaskda-code
Applies when
task -- the task asks for a confidence interval on a summary statistic (mean, ratio, difference, correlation) via bootstrap resampling with a specified number of replicates.
Pattern
The attempt computes percentiles of the raw per-observation values (or of a single resample, or of one replicate array) rather than of the distribution of the statistic recomputed on each of the N resamples; equivalently it never builds an array of length N of statistic values before taking percentiles. The reported interval therefore has roughly the width of the data's own dispersion instead of shrinking like the standard error (~1/√n), and no check is made that the interval is plausibly narrow, brackets the point estimate symmetrically-ish, and that the point estimate itself is in a physically sensible range.
Detection procedure
  1. From the task, note the requested statistic, the requested replicate count, and the confidence level; also note the required output file and column/format spec.
  2. In the scripts, locate the resampling loop/vectorized draw: confirm it produces one statistic per replicate (array length == requested replicates) and that percentiles are taken over that array, not over the observation-level values; confirm the resample is drawn with replacement at the same sample size and within the correct subgroup.
  3. Compare the reported interval half-width against a quick mental standard-error estimate (sample SD / √n × ~2.6 for 99%): if the half-width is of the order of the data's SD rather than the SE, or the interval does not tightly surround the reported point estimate, flag it.
  4. Check the point estimate against a back-of-envelope value from the raw columns (e.g., ratio of the two column means) and check that the requested file was actually written with the exact sample format/ordering/rounding.
Discriminator
A genuinely wide interval is fine when n is small or the statistic is heavy-tailed — verify by the SE scaling check, not by width alone; a real violation is when the code path never aggregates per-replicate statistics, or the width stays essentially constant/comparable to the raw data spread for a large-n mean.
Consequence
The point estimate and/or interval bounds differ from the reference beyond tolerance, so the expected output file fails every value check even if its layout looks right.
id 7e7de4d1400f · mined from da-code dacode-data-sa-029@s13
raw text (what the judge reads)
### Bootstrap interval taken from the spread of raw values instead of the spread of the resampled statistic
- **Applies when**: `task` -- the task asks for a confidence interval on a summary statistic (mean, ratio, difference, correlation) via bootstrap resampling with a specified number of replicates.
- **Pattern**: The attempt computes percentiles of the raw per-observation values (or of a single resample, or of one replicate array) rather than of the distribution of the statistic recomputed on each of the N resamples; equivalently it never builds an array of length N of statistic values before taking percentiles. The reported interval therefore has roughly the width of the data's own dispersion instead of shrinking like the standard error (~1/√n), and no check is made that the interval is plausibly narrow, brackets the point estimate symmetrically-ish, and that the point estimate itself is in a physically sensible range.
- **Detection procedure**:
  1. From the task, note the requested statistic, the requested replicate count, and the confidence level; also note the required output file and column/format spec.
  2. In the scripts, locate the resampling loop/vectorized draw: confirm it produces one statistic per replicate (array length == requested replicates) and that percentiles are taken over that array, not over the observation-level values; confirm the resample is drawn with replacement at the same sample size and within the correct subgroup.
  3. Compare the reported interval half-width against a quick mental standard-error estimate (sample SD / √n × ~2.6 for 99%): if the half-width is of the order of the data's SD rather than the SE, or the interval does not tightly surround the reported point estimate, flag it.
  4. Check the point estimate against a back-of-envelope value from the raw columns (e.g., ratio of the two column means) and check that the requested file was actually written with the exact sample format/ordering/rounding.
- **Discriminator**: A genuinely wide interval is fine when n is small or the statistic is heavy-tailed — verify by the SE scaling check, not by width alone; a real violation is when the code path never aggregates per-replicate statistics, or the width stays essentially constant/comparable to the raw data spread for a large-n mean.
- **Consequence**: The point estimate and/or interval bounds differ from the reference beyond tolerance, so the expected output file fails every value check even if its layout looks right.
726Output artifact silently covers only a fraction of the input recordstaskda-code
Applies when
task -- the task asks for a result file that assigns a value (label, prediction, score) to the data, and the script loads a source table and writes one row per record.
Pattern
The attempt runs the pipeline on a truncated/sampled/partially-read version of the data (or drops rows during preprocessing) and writes a result file with far fewer rows than the source, while the answer reports the reduced count as if it were the full dataset — no comparison against the known/documented size of the source is ever made.
Detection procedure
  1. From the task/README, note the expected order of magnitude of the input (e.g., stated record counts, "over half a million measurements") and any statement that the whole dataset must be processed.
  2. In the script, trace the data from load to write: look for nrows=, head(), sample(), slicing, subsetting, dropna, or an intermediate/pre-truncated file path, and check whether any assertion compares len(output) == len(input_source).
  3. In the reported answer, read the row/sample count of the written file and compare it to step 1; also compare it to the shape printed right after loading.
  4. Flag if the output row count is materially smaller than the source, or if no shape/count check ties the output back to the raw input.
Discriminator
A real violation is unexplained shrinkage or use of an unverified subset presented as the whole. It is fine if the task explicitly permits sampling AND the script both documents the sampling and the answer states it, or if rows are removed by an explicitly required filter whose removed count is reported and reconciled.
Consequence
The saved file cannot be aligned row-for-row with the reference data, so the grader's shape/index check on the expected result file fails outright regardless of clustering quality.
id 97f95694b9e9 · mined from da-code dacode-ml-cluster-010@s13
raw text (what the judge reads)
### Output artifact silently covers only a fraction of the input records
- **Applies when**: `task` -- the task asks for a result file that assigns a value (label, prediction, score) to the data, and the script loads a source table and writes one row per record.
- **Pattern**: The attempt runs the pipeline on a truncated/sampled/partially-read version of the data (or drops rows during preprocessing) and writes a result file with far fewer rows than the source, while the answer reports the reduced count as if it were the full dataset — no comparison against the known/documented size of the source is ever made.
- **Detection procedure**:
  1. From the task/README, note the expected order of magnitude of the input (e.g., stated record counts, "over half a million measurements") and any statement that the whole dataset must be processed.
  2. In the script, trace the data from load to write: look for `nrows=`, `head()`, `sample()`, slicing, subsetting, dropna, or an intermediate/pre-truncated file path, and check whether any assertion compares `len(output) == len(input_source)`.
  3. In the reported answer, read the row/sample count of the written file and compare it to step 1; also compare it to the shape printed right after loading.
  4. Flag if the output row count is materially smaller than the source, or if no shape/count check ties the output back to the raw input.
- **Discriminator**: A real violation is unexplained shrinkage or use of an unverified subset presented as the whole. It is fine if the task explicitly permits sampling AND the script both documents the sampling and the answer states it, or if rows are removed by an explicitly required filter whose removed count is reported and reconciled.
- **Consequence**: The saved file cannot be aligned row-for-row with the reference data, so the grader's shape/index check on the expected result file fails outright regardless of clustering quality.
727Degenerate cluster count chosen by blindly maximizing an internal score on skewed, outlier-dominated featurestaskda-code
Applies when
task -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script picks that number automatically by maximizing an internal validity score over heavy-tailed, un-transformed aggregate features.
Pattern
The attempt builds highly skewed count/sum features, standardizes them (which does not remove skew or outliers), then selects the k with the best silhouette/inertia score. A handful of extreme outliers form their own micro-cluster, so the score is maximized at a trivial solution (e.g., k=2 with >99% of rows in one cluster), and the attempt reports it as the "optimal" segmentation without any distribution/balance sanity check.
Detection procedure
  1. Read the task: confirm the deliverable is a meaningful segmentation with a justified number of groups, not just any labeling.
  2. Read the feature-engineering step: check whether long-tailed monetary/count features are log-transformed, winsorized, or outlier-filtered before scaling; if only StandardScaler is applied, extreme records will dominate distances.
  3. Read the model-selection step: check whether k is chosen purely by argmax of an internal index, and whether any guard exists (elbow inspection, minimum cluster size, stability, interpretability of profiles).
  4. Read the answer/output: inspect the reported cluster sizes and the score curve; flag if one cluster holds nearly all rows, or if the winning score is an isolated spike far above all other k values.
Discriminator
A genuinely small cluster is fine when it is deliberately identified as an outlier segment and the remaining clusters are substantive and interpretable, or when the skew was explicitly handled and multiple k values still agree. The violation is when the "optimal" solution is effectively outliers-vs-everyone-else, providing no customer/group differentiation and no evidence the agent questioned it.
Consequence
The saved label column is near-constant, so any grader comparison of cluster structure (number of clusters, size balance, agreement with a reference segmentation, or downstream silhouette on properly transformed features) fails, and the output file is marked wrong.
id f00a5ced8f18 · mined from da-code dacode-ml-cluster-019@s13
raw text (what the judge reads)
### Degenerate cluster count chosen by blindly maximizing an internal score on skewed, outlier-dominated features
- **Applies when**: `task` -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script picks that number automatically by maximizing an internal validity score over heavy-tailed, un-transformed aggregate features.
- **Pattern**: The attempt builds highly skewed count/sum features, standardizes them (which does not remove skew or outliers), then selects the k with the best silhouette/inertia score. A handful of extreme outliers form their own micro-cluster, so the score is maximized at a trivial solution (e.g., k=2 with >99% of rows in one cluster), and the attempt reports it as the "optimal" segmentation without any distribution/balance sanity check.
- **Detection procedure**:
  1. Read the task: confirm the deliverable is a meaningful segmentation with a justified number of groups, not just any labeling.
  2. Read the feature-engineering step: check whether long-tailed monetary/count features are log-transformed, winsorized, or outlier-filtered before scaling; if only `StandardScaler` is applied, extreme records will dominate distances.
  3. Read the model-selection step: check whether k is chosen purely by argmax of an internal index, and whether any guard exists (elbow inspection, minimum cluster size, stability, interpretability of profiles).
  4. Read the answer/output: inspect the reported cluster sizes and the score curve; flag if one cluster holds nearly all rows, or if the winning score is an isolated spike far above all other k values.
- **Discriminator**: A genuinely small cluster is fine when it is deliberately identified as an outlier segment *and* the remaining clusters are substantive and interpretable, or when the skew was explicitly handled and multiple k values still agree. The violation is when the "optimal" solution is effectively outliers-vs-everyone-else, providing no customer/group differentiation and no evidence the agent questioned it.
- **Consequence**: The saved label column is near-constant, so any grader comparison of cluster structure (number of clusters, size balance, agreement with a reference segmentation, or downstream silhouette on properly transformed features) fails, and the output file is marked wrong.
728Incomplete / unverified submission file coveragetaskda-code
Applies when
task -- the task requires producing a prediction or output file whose rows must correspond one-to-one with the rows of a provided input/test set (optionally in a given order and with a given header).
Pattern
The attempt produces or reports a partial output — fewer rows than the input set, truncated mid-record, missing ids, or pasted inline instead of written to the required filename — and never checks the row count / id set against the reference input or sample submission.
Detection procedure
  1. From the task, note the required output filename, header/column names, and the expected number of rows (= number of rows in the test/input file, or of the provided sample output).
  2. In the scripts, look for the step that writes the file and any assertion comparing the written frame's shape and id set to the test frame / sample file; note whether the file is actually saved (not just printed).
  3. Inspect the delivered artifact: count rows, check the last line is a complete record, and check the id set matches the reference exactly (no missing, extra, or duplicated ids).
  4. Also confirm each row has all required probability/value columns and plausible ranges (e.g., non-negative, correct number of fields).
Discriminator
A real violation is a mismatch in row count / id membership or a structurally broken (truncated, missing-column) file, or the absence of the required saved file. A look-alike that is fine: the full file exists and matches the reference ids, and only the display of it in the chat is abbreviated, with the script showing an explicit shape/id check.
Consequence
The grader cannot align predictions with ground-truth labels — the expected output file is reported WRONG/MISSING and the metric (e.g., log loss) cannot be computed or is scored as a failure, yielding 0 passed checks.
id bd700a1e0abe · mined from da-code dacode-ml-competition-005@s14
raw text (what the judge reads)
### Incomplete / unverified submission file coverage
- **Applies when**: `task` -- the task requires producing a prediction or output file whose rows must correspond one-to-one with the rows of a provided input/test set (optionally in a given order and with a given header).
- **Pattern**: The attempt produces or reports a partial output — fewer rows than the input set, truncated mid-record, missing ids, or pasted inline instead of written to the required filename — and never checks the row count / id set against the reference input or sample submission.
- **Detection procedure**:
  1. From the task, note the required output filename, header/column names, and the expected number of rows (= number of rows in the test/input file, or of the provided sample output).
  2. In the scripts, look for the step that writes the file and any assertion comparing the written frame's shape and id set to the test frame / sample file; note whether the file is actually saved (not just printed).
  3. Inspect the delivered artifact: count rows, check the last line is a complete record, and check the id set matches the reference exactly (no missing, extra, or duplicated ids).
  4. Also confirm each row has all required probability/value columns and plausible ranges (e.g., non-negative, correct number of fields).
- **Discriminator**: A real violation is a mismatch in row count / id membership or a structurally broken (truncated, missing-column) file, or the absence of the required saved file. A look-alike that is fine: the full file exists and matches the reference ids, and only the *display* of it in the chat is abbreviated, with the script showing an explicit shape/id check.
- **Consequence**: The grader cannot align predictions with ground-truth labels — the expected output file is reported WRONG/MISSING and the metric (e.g., log loss) cannot be computed or is scored as a failure, yielding 0 passed checks.
729Output artifact not verified for completeness and schema against the input rowstaskda-code
Applies when
task -- the task asks for a prediction/result file with a specified column name and one entry per row of a given input table.
Pattern
The agent produces the result inline or writes it without ever asserting that the written file has exactly as many rows as the input (in the same order/keys) and exactly the requested column name(s); the delivered content is truncated, deduplicated, reordered, or missing/extra keys, and no script is retained that regenerates it deterministically.
Detection procedure
  1. From the task, note the required output filename, required column name(s), and the expected number of rows (= number of records in the provided input file).
  2. In the scripts, look for the write step and for an explicit shape/schema check before or after writing (e.g., comparing len(predictions) to len(input_df), aligning on the identifier column, asserting header names); also confirm a script exists at all that produced the delivered file.
  3. Compare the delivered answer/file: count rows, check the header line matches the requested column name(s) exactly, and check the identifier set equals the input's identifier set (no missing, no duplicates, no partial last line).
  4. Flag if any count/name/key mismatch exists or if no code path guarantees the check.
Discriminator
A real violation is a row-count, key-set, header-name, or truncation mismatch between input and delivered output (or an answer pasted in chat with no corresponding saved file/script). Not a violation: full-length output with correct header and complete key coverage whose predicted values are merely inaccurate, or harmless extra columns if the task allows them.
Consequence
The grader cannot parse or align the expected file — it reports the required output as WRONG/MISSING regardless of model quality, since rows are absent or the schema/keys don't line up.
id c0df70ac5093 · mined from da-code dacode-ml-regression-008@s14
raw text (what the judge reads)
### Output artifact not verified for completeness and schema against the input rows
- **Applies when**: `task` -- the task asks for a prediction/result file with a specified column name and one entry per row of a given input table.
- **Pattern**: The agent produces the result inline or writes it without ever asserting that the written file has exactly as many rows as the input (in the same order/keys) and exactly the requested column name(s); the delivered content is truncated, deduplicated, reordered, or missing/extra keys, and no script is retained that regenerates it deterministically.
- **Detection procedure**:
  1. From the task, note the required output filename, required column name(s), and the expected number of rows (= number of records in the provided input file).
  2. In the scripts, look for the write step and for an explicit shape/schema check before or after writing (e.g., comparing `len(predictions)` to `len(input_df)`, aligning on the identifier column, asserting header names); also confirm a script exists at all that produced the delivered file.
  3. Compare the delivered answer/file: count rows, check the header line matches the requested column name(s) exactly, and check the identifier set equals the input's identifier set (no missing, no duplicates, no partial last line).
  4. Flag if any count/name/key mismatch exists or if no code path guarantees the check.
- **Discriminator**: A real violation is a row-count, key-set, header-name, or truncation mismatch between input and delivered output (or an answer pasted in chat with no corresponding saved file/script). Not a violation: full-length output with correct header and complete key coverage whose predicted values are merely inaccurate, or harmless extra columns if the task allows them.
- **Consequence**: The grader cannot parse or align the expected file — it reports the required output as WRONG/MISSING regardless of model quality, since rows are absent or the schema/keys don't line up.
730Statistical test applied to the full raw data with unvalidated assumptions and directiontaskda-code
Applies when
task -- the task asks for a hypothesis test (p-value + decision) on data whose relevant population, time window, or distributional form is not spelled out in the prompt.
Pattern
The script loads every row of both files, forms the comparison statistic, and immediately runs one default parametric two-sided routine (e.g., a plain independent-samples t-test) without (a) scoping the sample to the population the question is about, (b) checking whether the variable's distribution/variance justifies that test, or (c) checking whether the stated hypothesis implies a directional alternative. The resulting p-value is astronomically small because the sample is the entire multi-century, multi-competition record rather than the intended subset.
Detection procedure
  1. Read the task and README: identify what population the null hypothesis is really about (which competition tier, era, or record type) and whether the wording implies "different" vs. "greater/less".
  2. Read the script: check whether any filtering/subsetting step, exploratory distribution/normality/variance check, or justification of the chosen test and its alternative argument appears; note whether one canned test is called with all defaults on all rows.
  3. Read the answer: flag p-values at implausible magnitudes (e.g., <1e-20) or sample counts equal to the full file length — a sign the test was run on a far larger/heterogeneous sample than intended.
  4. Confirm the script never prints or reasons about the sample size/subset actually tested against the task's framing.
Discriminator
A real violation is when the script does zero scoping and zero assumption/direction checking and just uses the whole file with defaults. It is fine if the agent explicitly examined the distribution and subset options, documented why the full sample and a two-sided parametric test are appropriate, and its counts match the intended population — even if the same test is ultimately used.
Consequence
The reported p-value differs from the reference by many orders of magnitude (and potentially the reject/fail-to-reject decision flips), so the value check on the output file fails even though the file format is correct.
id 9b076faa3bb4 · mined from da-code dacode-data-sa-001@s14
raw text (what the judge reads)
### Statistical test applied to the full raw data with unvalidated assumptions and direction
- **Applies when**: `task` -- the task asks for a hypothesis test (p-value + decision) on data whose relevant population, time window, or distributional form is not spelled out in the prompt.
- **Pattern**: The script loads every row of both files, forms the comparison statistic, and immediately runs one default parametric two-sided routine (e.g., a plain independent-samples t-test) without (a) scoping the sample to the population the question is about, (b) checking whether the variable's distribution/variance justifies that test, or (c) checking whether the stated hypothesis implies a directional alternative. The resulting p-value is astronomically small because the sample is the entire multi-century, multi-competition record rather than the intended subset.
- **Detection procedure**:
  1. Read the task and README: identify what population the null hypothesis is really about (which competition tier, era, or record type) and whether the wording implies "different" vs. "greater/less".
  2. Read the script: check whether any filtering/subsetting step, exploratory distribution/normality/variance check, or justification of the chosen test and its `alternative` argument appears; note whether one canned test is called with all defaults on all rows.
  3. Read the answer: flag p-values at implausible magnitudes (e.g., <1e-20) or sample counts equal to the full file length — a sign the test was run on a far larger/heterogeneous sample than intended.
  4. Confirm the script never prints or reasons about the sample size/subset actually tested against the task's framing.
- **Discriminator**: A real violation is when the script does zero scoping and zero assumption/direction checking and just uses the whole file with defaults. It is fine if the agent explicitly examined the distribution and subset options, documented why the full sample and a two-sided parametric test are appropriate, and its counts match the intended population — even if the same test is ultimately used.
- **Consequence**: The reported p-value differs from the reference by many orders of magnitude (and potentially the reject/fail-to-reject decision flips), so the value check on the output file fails even though the file format is correct.
731Saved "feature" columns don't match the vector actually used for modeling (silent drops / un-transformed values)taskda-code
Applies when
task -- the task asks to persist a per-row result file whose columns are the elements of the feature vector plus a model output (label/score), and the script builds a feature matrix through a preprocessing pipeline before fitting.
Pattern
The script fits the model on a transformed matrix (scaled/imputed/encoded) but writes a different matrix to disk — e.g. the pre-transform raw frame, or a subset that silently excludes columns the pipeline dropped (non-numeric/categorical/derived fields) — and the final answer describes the saved columns as if they were the modeled ones ("standardized values", "all attributes"), so column count, ordering and value scale in the artifact don't correspond to what was clustered/trained on.
Detection procedure
  1. From the task, note exactly what each output column is supposed to contain and how many feature elements are implied by the described data.
  2. In the script, identify the object passed to fit/fit_predict and the object written by to_csv; check whether they are the same array (same transform stage, same column list, same order).
  3. Check how the feature list was built: were columns of some dtype excluded without justification, or was an available informative field discarded rather than encoded/derived? Compare the resulting feature count to the dataset's attribute count.
  4. Compare the answer's prose description of the saved columns (transformation applied, feature semantics, chosen hyperparameters) against what the code literally does; any mismatch (e.g. "k chosen by silhouette" while k is hard-coded) is a red flag on the whole artifact.
Discriminator
A real violation is when the persisted matrix cannot be reproduced from the modeling matrix (different stage, fewer/extra columns, different order) or the narrative contradicts the code. It is not a violation if the script deliberately and consistently saves the same feature representation it modeled (raw or transformed) with all relevant attributes represented, and the answer describes it accurately.
Consequence
The graded file has the wrong number/content of Feature_i columns and values on the wrong scale, so a column-wise or shape-wise comparison against the expected artifact fails even if the clustering itself was reasonable.
id 09621a5cb754 · mined from da-code dacode-ml-cluster-014@s14
raw text (what the judge reads)
### Saved "feature" columns don't match the vector actually used for modeling (silent drops / un-transformed values)
- **Applies when**: `task` -- the task asks to persist a per-row result file whose columns are the elements of the feature vector plus a model output (label/score), and the script builds a feature matrix through a preprocessing pipeline before fitting.
- **Pattern**: The script fits the model on a transformed matrix (scaled/imputed/encoded) but writes a *different* matrix to disk — e.g. the pre-transform raw frame, or a subset that silently excludes columns the pipeline dropped (non-numeric/categorical/derived fields) — and the final answer describes the saved columns as if they were the modeled ones ("standardized values", "all attributes"), so column count, ordering and value scale in the artifact don't correspond to what was clustered/trained on.
- **Detection procedure**:
  1. From the task, note exactly what each output column is supposed to contain and how many feature elements are implied by the described data.
  2. In the script, identify the object passed to `fit`/`fit_predict` and the object written by `to_csv`; check whether they are the same array (same transform stage, same column list, same order).
  3. Check how the feature list was built: were columns of some dtype excluded without justification, or was an available informative field discarded rather than encoded/derived? Compare the resulting feature count to the dataset's attribute count.
  4. Compare the answer's prose description of the saved columns (transformation applied, feature semantics, chosen hyperparameters) against what the code literally does; any mismatch (e.g. "k chosen by silhouette" while k is hard-coded) is a red flag on the whole artifact.
- **Discriminator**: A real violation is when the persisted matrix cannot be reproduced from the modeling matrix (different stage, fewer/extra columns, different order) or the narrative contradicts the code. It is *not* a violation if the script deliberately and consistently saves the same feature representation it modeled (raw or transformed) with all relevant attributes represented, and the answer describes it accurately.
- **Consequence**: The graded file has the wrong number/content of `Feature_i` columns and values on the wrong scale, so a column-wise or shape-wise comparison against the expected artifact fails even if the clustering itself was reasonable.
732Submission does not cover the full test index (row count / ID set never validated against the template)taskda-code
Applies when
task -- The task requires writing a prediction file whose rows must correspond one-to-one with a provided test set or sample submission template.
Pattern
The agent produces a prediction file (or pastes an inline answer) containing only a subset of the required rows — e.g. a sample, a chunk, a debug slice, or the head of a dataframe — and never asserts that its row count, ID set, and column names match the template exactly; the reported IDs also appear unordered/arbitrary relative to the template.
Detection procedure
  1. From the task/README and the template file, determine the expected number of prediction rows, the exact ID column values, and the required column names/order.
  2. Read the scripts: find where predictions are written and check for any explicit alignment step (merge/reindex on the template's IDs) and any assertion on len(pred) == len(template) and set(pred.id) == set(template.id).
  3. Count the rows and inspect the IDs in the produced answer/file; compare against the expected count and ID set (and check for duplicates or missing IDs).
  4. Flag if the counts differ, IDs are a subset/superset, columns are renamed/reordered, or no alignment/sanity check exists in the code.
Discriminator
A genuine violation is a mismatch in row count or ID membership versus the template (or code with no alignment and no check). It is not a violation if the file matches the template exactly in count and IDs but merely differs in the order the reviewer expected, when the task/template does not require a specific row order — or if the reviewer is only seeing a truncated display of a full file whose length can be confirmed from the script's write step or a printed shape.
Consequence
The grader cannot match predictions to the ground-truth index, so the submission is scored as wrong/missing (0 checks passed) regardless of model quality.
id f34fc950411c · mined from da-code dacode-ml-competition-009@s14
raw text (what the judge reads)
### Submission does not cover the full test index (row count / ID set never validated against the template)
- **Applies when**: `task` -- The task requires writing a prediction file whose rows must correspond one-to-one with a provided test set or sample submission template.
- **Pattern**: The agent produces a prediction file (or pastes an inline answer) containing only a subset of the required rows — e.g. a sample, a chunk, a debug slice, or the head of a dataframe — and never asserts that its row count, ID set, and column names match the template exactly; the reported IDs also appear unordered/arbitrary relative to the template.
- **Detection procedure**:
  1. From the task/README and the template file, determine the expected number of prediction rows, the exact ID column values, and the required column names/order.
  2. Read the scripts: find where predictions are written and check for any explicit alignment step (merge/reindex on the template's IDs) and any assertion on `len(pred) == len(template)` and `set(pred.id) == set(template.id)`.
  3. Count the rows and inspect the IDs in the produced answer/file; compare against the expected count and ID set (and check for duplicates or missing IDs).
  4. Flag if the counts differ, IDs are a subset/superset, columns are renamed/reordered, or no alignment/sanity check exists in the code.
- **Discriminator**: A genuine violation is a mismatch in row count or ID membership versus the template (or code with no alignment and no check). It is *not* a violation if the file matches the template exactly in count and IDs but merely differs in the order the reviewer expected, when the task/template does not require a specific row order — or if the reviewer is only seeing a truncated display of a full file whose length can be confirmed from the script's write step or a printed shape.
- **Consequence**: The grader cannot match predictions to the ground-truth index, so the submission is scored as wrong/missing (0 checks passed) regardless of model quality.
733Output row coverage silently reduced by dropping incomplete recordstaskda-code
Applies when
task -- the task asks for a per-record output artifact (labels, predictions, scores) covering the input table, and the script filters or dropna()s rows before producing it.
Pattern
The attempt selects an ad-hoc subset of columns, drops every row with any missing value in that subset (and/or keeps ID-like/non-informative numeric columns), then writes a result file with fewer rows than the input dataset while reporting success — so the saved artifact cannot be aligned with the expected per-record output.
Detection procedure
  1. From the task/README, note the expected granularity and count of the output (one row per record in the source file) and any stated column/format constraints.
  2. In the script, find every row-removing operation (dropna, boolean filters, merge with non-outer join, deduplication) applied before the artifact is written, and check whether missing values are instead imputed/handled so all records survive.
  3. In the answer/logs, compare the reported number of output rows with the number of input records; also check that the feature columns used are substantive attributes rather than identifiers/codes/coordinates that carry no analytic meaning.
  4. Verify the written file's header names and column order literally match the requested naming scheme.
Discriminator
A real violation is when rows are lost purely as a side effect of missing-value handling or arbitrary column selection while the task expects full coverage; it is fine if the task explicitly asks to filter to a subpopulation, or if the dropped rows are documented duplicates/non-records and the output granularity requested still matches.
Consequence
The saved file has the wrong shape/row alignment versus the reference artifact, so a file-level comparison of labels per record fails and the task scores 0 despite "success" being reported.
id e52ac16056bb · mined from da-code dacode-ml-cluster-009@s14
raw text (what the judge reads)
### Output row coverage silently reduced by dropping incomplete records
- **Applies when**: `task` -- the task asks for a per-record output artifact (labels, predictions, scores) covering the input table, and the script filters or `dropna()`s rows before producing it.
- **Pattern**: The attempt selects an ad-hoc subset of columns, drops every row with any missing value in that subset (and/or keeps ID-like/non-informative numeric columns), then writes a result file with fewer rows than the input dataset while reporting success — so the saved artifact cannot be aligned with the expected per-record output.
- **Detection procedure**:
  1. From the task/README, note the expected granularity and count of the output (one row per record in the source file) and any stated column/format constraints.
  2. In the script, find every row-removing operation (`dropna`, boolean filters, `merge` with non-outer join, deduplication) applied before the artifact is written, and check whether missing values are instead imputed/handled so all records survive.
  3. In the answer/logs, compare the reported number of output rows with the number of input records; also check that the feature columns used are substantive attributes rather than identifiers/codes/coordinates that carry no analytic meaning.
  4. Verify the written file's header names and column order literally match the requested naming scheme.
- **Discriminator**: A real violation is when rows are lost purely as a side effect of missing-value handling or arbitrary column selection while the task expects full coverage; it is fine if the task explicitly asks to filter to a subpopulation, or if the dropped rows are documented duplicates/non-records and the output granularity requested still matches.
- **Consequence**: The saved file has the wrong shape/row alignment versus the reference artifact, so a file-level comparison of labels per record fails and the task scores 0 despite "success" being reported.
734Degenerate clustering accepted because model-selection metric was not sanity-checked against cluster balancetaskda-code
Applies when
task -- the task asks for an unsupervised segmentation (or any grouping) of entities built from skewed, heavy-tailed aggregated features, and the script picks the number of groups by optimizing a single internal score (silhouette, inertia elbow, etc.).
Pattern
The attempt scales raw, extremely right-skewed aggregates with a plain standardizer (no log/rank transform, no outlier handling or winsorizing), then selects the configuration with the best internal score. Because a handful of extreme records separate trivially, the score is near-perfect while essentially all entities fall into one giant group — a solution that carries no segmentation information — and the attempt reports it as success without questioning the group-size distribution.
Detection procedure
  1. In the task, confirm the deliverable is a meaningful partition of entities (labels per row), not just any label column.
  2. In the scripts, check whether skew-heavy monetary/count/total features are transformed (log1p, quantile/rank) or trimmed before scaling, and whether the selection loop imposes any constraint beyond maximizing the internal score.
  3. In the reported output, read the per-group counts and the internal score together: flag if one group holds a dominant share (e.g. >90%) of rows while others hold a handful, and/or the score is implausibly high (≳0.8) for real-world segmentation data.
  4. Check whether the attempt ran any alternative (different transform, k, or algorithm) and compared balance, rather than accepting the first optimum.
Discriminator
A genuine violation shows near-empty groups made of outliers plus one catch-all group, with no transform or outlier handling in the pipeline; a legitimate case is an intentionally reported outlier group alongside several substantively sized segments, or a documented decision that the natural structure is imbalanced supported by comparison to transformed/trimmed alternatives.
Consequence
The saved label file is effectively single-cluster, so any grader check on the partition (number of non-trivial clusters, cluster-size distribution, agreement with a reasonable reference segmentation, or recomputed silhouette on transformed features) fails, marking the result file WRONG despite correct column naming.
id f6a96dad4205 · mined from da-code dacode-ml-cluster-016@s14
raw text (what the judge reads)
### Degenerate clustering accepted because model-selection metric was not sanity-checked against cluster balance
- **Applies when**: `task` -- the task asks for an unsupervised segmentation (or any grouping) of entities built from skewed, heavy-tailed aggregated features, and the script picks the number of groups by optimizing a single internal score (silhouette, inertia elbow, etc.).
- **Pattern**: The attempt scales raw, extremely right-skewed aggregates with a plain standardizer (no log/rank transform, no outlier handling or winsorizing), then selects the configuration with the best internal score. Because a handful of extreme records separate trivially, the score is near-perfect while essentially all entities fall into one giant group — a solution that carries no segmentation information — and the attempt reports it as success without questioning the group-size distribution.
- **Detection procedure**:
  1. In the task, confirm the deliverable is a meaningful partition of entities (labels per row), not just any label column.
  2. In the scripts, check whether skew-heavy monetary/count/total features are transformed (log1p, quantile/rank) or trimmed before scaling, and whether the selection loop imposes any constraint beyond maximizing the internal score.
  3. In the reported output, read the per-group counts and the internal score together: flag if one group holds a dominant share (e.g. >90%) of rows while others hold a handful, and/or the score is implausibly high (≳0.8) for real-world segmentation data.
  4. Check whether the attempt ran any alternative (different transform, k, or algorithm) and compared balance, rather than accepting the first optimum.
- **Discriminator**: A genuine violation shows near-empty groups made of outliers plus one catch-all group, with no transform or outlier handling in the pipeline; a legitimate case is an intentionally reported outlier group *alongside* several substantively sized segments, or a documented decision that the natural structure is imbalanced supported by comparison to transformed/trimmed alternatives.
- **Consequence**: The saved label file is effectively single-cluster, so any grader check on the partition (number of non-trivial clusters, cluster-size distribution, agreement with a reasonable reference segmentation, or recomputed silhouette on transformed features) fails, marking the result file WRONG despite correct column naming.
735Boolean/categorical answer not written in the exact literal form the format specifiestaskinfiagent-dabench
Applies when
task -- The answer template asks for a value drawn from a fixed vocabulary (e.g. True/False, Yes/No, a class label) inside a tagged output field.
Pattern
The agent computes the value correctly but emits it in a different surface form than the template declares — lowercase/uppercase variants, 0/1, "none", Python repr of a numpy bool, or a re-worded label — so a string-matching grader rejects the field even though the underlying analysis is right.
Detection procedure
  1. Read the task's answer-format block and list, verbatim, the allowed literals for every non-numeric field (and rounding/units for numeric fields).
  2. Read the scripts/answer construction to see how each field is stringified (direct print(bool_var), f-string of a numpy type, hand-typed text) and whether any normalization to the declared literal is done.
  3. Compare the final submitted answer token-by-token against the allowed literals, including capitalization and surrounding tag names.
  4. Flag if any field's text is not character-identical to a declared allowed literal.
Discriminator
A real violation is a mismatch in the literal itself (false vs False, 1 vs True); a look-alike that is fine is harmless whitespace outside the brackets or a numeric field that merely differs in trailing zeros while respecting the stated rounding.
Consequence
The grader marks that specific check WRONG/MISSING while the other fields pass, yielding a partially-correct submission scored as incorrect overall.
id 84b1a8768c41 · mined from infiagent-dabench dabench-733@s14
raw text (what the judge reads)
### Boolean/categorical answer not written in the exact literal form the format specifies
- **Applies when**: `task` -- The answer template asks for a value drawn from a fixed vocabulary (e.g. `True`/`False`, `Yes`/`No`, a class label) inside a tagged output field.
- **Pattern**: The agent computes the value correctly but emits it in a different surface form than the template declares — lowercase/uppercase variants, `0/1`, `"none"`, Python `repr` of a numpy bool, or a re-worded label — so a string-matching grader rejects the field even though the underlying analysis is right.
- **Detection procedure**:
  1. Read the task's answer-format block and list, verbatim, the allowed literals for every non-numeric field (and rounding/units for numeric fields).
  2. Read the scripts/answer construction to see how each field is stringified (direct `print(bool_var)`, f-string of a numpy type, hand-typed text) and whether any normalization to the declared literal is done.
  3. Compare the final submitted answer token-by-token against the allowed literals, including capitalization and surrounding tag names.
  4. Flag if any field's text is not character-identical to a declared allowed literal.
- **Discriminator**: A real violation is a mismatch in the literal itself (`false` vs `False`, `1` vs `True`); a look-alike that is fine is harmless whitespace outside the brackets or a numeric field that merely differs in trailing zeros while respecting the stated rounding.
- **Consequence**: The grader marks that specific check WRONG/MISSING while the other fields pass, yielding a partially-correct submission scored as incorrect overall.
736Missing out-of-sample validation before submitting predictionstaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test file and the scripts fit a model on labeled data and write the predictions straight to the output file.
Pattern
The attempt reports only descriptive artifacts (feature importances, prediction min/max/mean, row counts) as evidence of quality, with no error metric computed on held-out labeled rows constructed to mimic the test rows (same feature availability, same temporal/group split); model choice, feature set, and imputation are therefore never checked against any accuracy target, and calendar/index-like features that cannot generalize are left in.
Detection procedure
  1. Read the task to determine which rows are unlabeled and how test rows relate to training rows (time-ordered continuation, random subset, distinct groups).
  2. Search the scripts for a held-out split or cross-validation and a fitted-vs-actual metric (RMSE/MAE/R² or the task's metric) printed on labeled data that were not used for fitting; check the split mirrors the train/test relationship.
  3. Check the answer: does it quote an out-of-sample error number, or only feature importances and prediction summary statistics?
  4. Verify the feature set used for prediction exists with the same meaning in the test file and would not be unavailable/degenerate there (e.g., raw index-like or non-recurring calendar values used as top predictors).
Discriminator
A real violation is when no numeric out-of-sample accuracy estimate exists anywhere, so nothing distinguishes a good model from one barely better than the mean; it is fine if the script reports a validation/CV score (even a simple one) on an appropriately constructed split, or compares against a baseline, even if the final answer text emphasizes other details.
Consequence
The submitted prediction file is syntactically valid (right column name and row count) but its error against the hidden labels exceeds the grader's threshold, so the file check fails with no diagnostic in the attempt explaining why.
id 11bf3fa054f2 · mined from da-code dacode-ml-regression-002@s14
raw text (what the judge reads)
### Missing out-of-sample validation before submitting predictions
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test file and the scripts fit a model on labeled data and write the predictions straight to the output file.
- **Pattern**: The attempt reports only descriptive artifacts (feature importances, prediction min/max/mean, row counts) as evidence of quality, with no error metric computed on held-out labeled rows constructed to mimic the test rows (same feature availability, same temporal/group split); model choice, feature set, and imputation are therefore never checked against any accuracy target, and calendar/index-like features that cannot generalize are left in.
- **Detection procedure**:
  1. Read the task to determine which rows are unlabeled and how test rows relate to training rows (time-ordered continuation, random subset, distinct groups).
  2. Search the scripts for a held-out split or cross-validation and a fitted-vs-actual metric (RMSE/MAE/R² or the task's metric) printed on labeled data that were not used for fitting; check the split mirrors the train/test relationship.
  3. Check the answer: does it quote an out-of-sample error number, or only feature importances and prediction summary statistics?
  4. Verify the feature set used for prediction exists with the same meaning in the test file and would not be unavailable/degenerate there (e.g., raw index-like or non-recurring calendar values used as top predictors).
- **Discriminator**: A real violation is when *no* numeric out-of-sample accuracy estimate exists anywhere, so nothing distinguishes a good model from one barely better than the mean; it is fine if the script reports a validation/CV score (even a simple one) on an appropriately constructed split, or compares against a baseline, even if the final answer text emphasizes other details.
- **Consequence**: The submitted prediction file is syntactically valid (right column name and row count) but its error against the hidden labels exceeds the grader's threshold, so the file check fails with no diagnostic in the attempt explaining why.
737Plot/artifact answers that aren't backed by saved, reproducible code and serialized underlying datataskda-code
Applies when
task -- the task asks for a chart or other artifact built to a spec file, and grading will inspect the artifact's underlying data/config (e.g., serialized plot data and numeric arrays) rather than just the image.
Pattern
The agent reports success and describes the figure in prose, but no script is retained that loads the source file, parses the spec, and writes every expected output; the plotted series is described from memory/summary statistics with no traceable extraction step, so the numbers may be invented or aggregated at the wrong granularity and the machine-checkable artifacts are absent.
Detection procedure
  1. From the task and any provided spec/config file, list every required output artifact (image plus any serialized data/config dumps) and every constrained property (title, axis labels, ticks, figure size, color, series ordering).
  2. Read the scripts: confirm there is executable code that (a) reads the actual source data file, (b) reads the spec file programmatically instead of hard-coding remembered values, and (c) writes each listed artifact to the required filename.
  3. Cross-check the answer's reported series (row count, date span, value range) against a quick recomputation path in the script; verify the values could only have come from the source file, not from narrative summary.
  4. Flag if any required artifact has no writing code, or if reported numbers cannot be traced to a load-and-transform step.
Discriminator
A real violation is missing/unwritten required artifacts or plotted values with no code path from the source data; it is fine if the script writes all artifacts and the prose summary merely paraphrases numbers that the code demonstrably computed from the input.
Consequence
The grader reports the expected data/config files as WRONG/MISSING (and any value checks fail), scoring 0 even though an image was produced.
id ee09a1f22967 · mined from da-code dacode-plot-line-015@s14
raw text (what the judge reads)
### Plot/artifact answers that aren't backed by saved, reproducible code and serialized underlying data
- **Applies when**: `task` -- the task asks for a chart or other artifact built to a spec file, and grading will inspect the artifact's underlying data/config (e.g., serialized plot data and numeric arrays) rather than just the image.
- **Pattern**: The agent reports success and describes the figure in prose, but no script is retained that loads the source file, parses the spec, and writes every expected output; the plotted series is described from memory/summary statistics with no traceable extraction step, so the numbers may be invented or aggregated at the wrong granularity and the machine-checkable artifacts are absent.
- **Detection procedure**:
  1. From the task and any provided spec/config file, list every required output artifact (image plus any serialized data/config dumps) and every constrained property (title, axis labels, ticks, figure size, color, series ordering).
  2. Read the scripts: confirm there is executable code that (a) reads the actual source data file, (b) reads the spec file programmatically instead of hard-coding remembered values, and (c) writes each listed artifact to the required filename.
  3. Cross-check the answer's reported series (row count, date span, value range) against a quick recomputation path in the script; verify the values could only have come from the source file, not from narrative summary.
  4. Flag if any required artifact has no writing code, or if reported numbers cannot be traced to a load-and-transform step.
- **Discriminator**: A real violation is missing/unwritten required artifacts or plotted values with no code path from the source data; it is fine if the script writes all artifacts and the prose summary merely paraphrases numbers that the code demonstrably computed from the input.
- **Consequence**: The grader reports the expected data/config files as WRONG/MISSING (and any value checks fail), scoring 0 even though an image was produced.
738Output file not reconciled with the provided sample/template (and boundary values left unchecked)taskda-code
Applies when
task -- the task supplies a sample result file (or an explicit format spec) and asks for computed statistics to be written into a result file, and the script generates that file itself.
Pattern
The attempt computes something plausible but writes its own column set/names/ordering, extra or missing fields, or unrounded/degenerate values (e.g. an exact 0.0 p-value from a finite number of resamples) without ever loading the sample file and confirming that the produced file matches it row-for-row and column-for-column.
Detection procedure
  1. Read the task for the referenced template/spec file and note the exact expected column names, row count, ordering, units and any rounding/precision requirement.
  2. Search the scripts for a read of that template (or a hard-coded header copied from it) and for a final assertion/print comparing the produced header and shape to it; absence of any such reconciliation is the first red flag.
  3. Inspect the submitted values for boundary/degenerate outputs (0, 1, exactly 0.0 probabilities, NaN, negative variances) and check whether the script bounds or documents them (e.g. reports < 1/n_resamples, or uses enough replicates/expected precision) rather than emitting the raw limit.
  4. Confirm the reported quantity is the one requested (final statistic, correct sign/direction, correct scale) and not an intermediate or differently-defined value.
Discriminator
A real violation is a file whose header/columns/rounding/value semantics were never checked against the given template, or a boundary value emitted with no justification; it is not a violation if the script explicitly loads the template, validates schema and dtypes, and a boundary value is genuinely correct and reported at the stated resolution.
Consequence
The grader compares the result file field-by-field and marks it WRONG/MISSING even when the underlying computation is close, so the check fails 0/1 despite a reasonable-looking number.
id a0b92fb1ade4 · mined from da-code dacode-data-sa-028@s14
raw text (what the judge reads)
### Output file not reconciled with the provided sample/template (and boundary values left unchecked)
- **Applies when**: `task` -- the task supplies a sample result file (or an explicit format spec) and asks for computed statistics to be written into a result file, and the script generates that file itself.
- **Pattern**: The attempt computes something plausible but writes its own column set/names/ordering, extra or missing fields, or unrounded/degenerate values (e.g. an exact `0.0` p-value from a finite number of resamples) without ever loading the sample file and confirming that the produced file matches it row-for-row and column-for-column.
- **Detection procedure**:
  1. Read the task for the referenced template/spec file and note the exact expected column names, row count, ordering, units and any rounding/precision requirement.
  2. Search the scripts for a read of that template (or a hard-coded header copied from it) and for a final assertion/print comparing the produced header and shape to it; absence of any such reconciliation is the first red flag.
  3. Inspect the submitted values for boundary/degenerate outputs (0, 1, exactly 0.0 probabilities, NaN, negative variances) and check whether the script bounds or documents them (e.g. reports `< 1/n_resamples`, or uses enough replicates/expected precision) rather than emitting the raw limit.
  4. Confirm the reported quantity is the one requested (final statistic, correct sign/direction, correct scale) and not an intermediate or differently-defined value.
- **Discriminator**: A real violation is a file whose header/columns/rounding/value semantics were never checked against the given template, or a boundary value emitted with no justification; it is *not* a violation if the script explicitly loads the template, validates schema and dtypes, and a boundary value is genuinely correct and reported at the stated resolution.
- **Consequence**: The grader compares the result file field-by-field and marks it WRONG/MISSING even when the underlying computation is close, so the check fails 0/1 despite a reasonable-looking number.
739Incomplete deliverables and unverified spec compliance for prescribed-format outputstaskda-code
Applies when
task -- the task points to auxiliary instruction/config files (e.g., a tips file, a plot/format spec) and expects a specific set of output artifacts (figure plus serialized data/config files).
Pattern
The attempt improvises its own interpretation of the analysis, produces only the most obvious artifact (the image), and never reads/echoes the spec files to confirm every required output, styling key, series definition, axis, ordering, or units — then declares success based on self-described "features" rather than on the spec.
Detection procedure
1) List from the task (and its referenced instruction/config files) every required output file and every explicitly constrained property (labels, titles, colors, sizes, aggregation definition, rounding, units, data ordering). 2) Check the scripts/answer for evidence that each referenced spec file was actually opened and each of its keys applied, and that each required artifact is written. 3) Check whether the answer reports which spec items were satisfied, or only generic claims ("proper axis labels", "grid", "high resolution"). 4) Confirm the aggregation/derived quantity matches the stated definition (e.g., "changes" vs raw values, correct grouping keys) rather than a plausible substitute.
Discriminator
A real violation is missing required artifacts or unread/unapplied spec files and self-invented conventions; it is fine if the agent quotes the spec, maps each key to code, saves all required files, and only fills in details the spec leaves open.
Consequence
Graders comparing artifact-by-artifact find missing serialized outputs and a figure/data whose values, series, or styling differ from the reference — all checks fail even though a plot exists.
id e2f1debe8ab4 · mined from da-code dacode-plot-line-006@s14
raw text (what the judge reads)
### Incomplete deliverables and unverified spec compliance for prescribed-format outputs
- **Applies when**: `task` -- the task points to auxiliary instruction/config files (e.g., a tips file, a plot/format spec) and expects a specific set of output artifacts (figure plus serialized data/config files).
- **Pattern**: The attempt improvises its own interpretation of the analysis, produces only the most obvious artifact (the image), and never reads/echoes the spec files to confirm every required output, styling key, series definition, axis, ordering, or units — then declares success based on self-described "features" rather than on the spec.
- **Detection procedure**: 1) List from the task (and its referenced instruction/config files) every required output file and every explicitly constrained property (labels, titles, colors, sizes, aggregation definition, rounding, units, data ordering). 2) Check the scripts/answer for evidence that each referenced spec file was actually opened and each of its keys applied, and that each required artifact is written. 3) Check whether the answer reports which spec items were satisfied, or only generic claims ("proper axis labels", "grid", "high resolution"). 4) Confirm the aggregation/derived quantity matches the stated definition (e.g., "changes" vs raw values, correct grouping keys) rather than a plausible substitute.
- **Discriminator**: A real violation is missing required artifacts or unread/unapplied spec files and self-invented conventions; it is fine if the agent quotes the spec, maps each key to code, saves all required files, and only fills in details the spec leaves open.
- **Consequence**: Graders comparing artifact-by-artifact find missing serialized outputs and a figure/data whose values, series, or styling differ from the reference — all checks fail even though a plot exists.
740Requested output artifact and schema never produced by the scriptstaskda-code
Applies when
task -- the task prescribes an exact answer template (key names, list/scalar shape, numeric formatting) and/or an expected result file, and the scripts only print diagnostics to stdout.
Pattern
The attempt does all the computation with print() statements and the agent hand-transcribes a final answer, so nothing writes the required result file and the transcribed object silently deviates from the shown schema (e.g., bare scalars or quoted strings where the template shows lists/numbers, renamed or extra keys, values re-formatted or rounded differently than computed).
Detection procedure
  1. From the task text, write down the exact required artifact (file name/path, if any) and the literal template: each key, whether each value is a list or scalar, and the expected type/precision.
  2. Grep the scripts for any serialization step (json.dump, to_json, to_csv, open(..., "w")) targeting that artifact; if none exists, the required output is unproduced.
  3. Compare the submitted answer object field-by-field against the template: key spelling, container type (list vs scalar), value type (number vs string), and numeric precision versus the value the script computed.
  4. Flag the attempt if the artifact is missing or any field's name/type/shape/precision differs from the template.
Discriminator
A real violation is a structural mismatch — no result file written, or keys/containers/value types that differ from the template. It is not a violation if the script writes the required artifact with the template's exact keys and containers and the answer text merely mirrors it (harmless whitespace or equivalent numeric rendering such as 82.6 vs 82.60 when the checker compares numerically).
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation is right, since the answer cannot be parsed/compared against the reference schema.
id d40817d36140 · mined from da-code dacode-di-text-002@s14
raw text (what the judge reads)
### Requested output artifact and schema never produced by the scripts
- **Applies when**: `task` -- the task prescribes an exact answer template (key names, list/scalar shape, numeric formatting) and/or an expected result file, and the scripts only print diagnostics to stdout.
- **Pattern**: The attempt does all the computation with `print()` statements and the agent hand-transcribes a final answer, so nothing writes the required result file and the transcribed object silently deviates from the shown schema (e.g., bare scalars or quoted strings where the template shows lists/numbers, renamed or extra keys, values re-formatted or rounded differently than computed).
- **Detection procedure**:
  1. From the task text, write down the exact required artifact (file name/path, if any) and the literal template: each key, whether each value is a list or scalar, and the expected type/precision.
  2. Grep the scripts for any serialization step (`json.dump`, `to_json`, `to_csv`, `open(..., "w")`) targeting that artifact; if none exists, the required output is unproduced.
  3. Compare the submitted answer object field-by-field against the template: key spelling, container type (list vs scalar), value type (number vs string), and numeric precision versus the value the script computed.
  4. Flag the attempt if the artifact is missing or any field's name/type/shape/precision differs from the template.
- **Discriminator**: A real violation is a structural mismatch — no result file written, or keys/containers/value types that differ from the template. It is *not* a violation if the script writes the required artifact with the template's exact keys and containers and the answer text merely mirrors it (harmless whitespace or equivalent numeric rendering such as `82.6` vs `82.60` when the checker compares numerically).
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 even when the underlying computation is right, since the answer cannot be parsed/compared against the reference schema.
741Assumed output schema instead of deriving it from the stated format requirementtaskda-code
Applies when
task -- the task says the deliverable file must "follow the required format" (or otherwise implies a fixed schema/column names/index/ordering/rounding) without spelling it out inline.
Pattern
The agent invents its own column names, index handling, row set, or numeric conventions, then asserts the output "follows the required format" without ever locating or reproducing the authoritative spec (README/docs/template/sample file/notebook prompt in the data directory), and its verification script only re-checks its own arithmetic against itself.
Detection procedure
  1. Read the task for any words indicating a prescribed output format, and note that the exact schema is not given in the prompt text.
  2. Search the scripts for any step that inspects the workspace for a format definition or template (listing the data directory, reading a README/sample/expected-columns file) and for any check of the produced file against it (column names, order, index column, row count, rounding, date representation).
  3. Check whether the agent's chosen names/structure are traceable to such a source or were simply made up; also check whether raw-data quirks (missing values, extra/unused columns, header/index columns) were inspected before aggregation.
  4. Read the answer: does it claim format compliance while the scripts contain no comparison to an external reference?
Discriminator
Fine if the schema is explicitly dictated in the task text (and the script matches it verbatim) or the script reads a template/expected-output file and validates against it; a violation is when the only "verification" recomputes the agent's own numbers and no external format source was ever consulted.
Consequence
The numbers may be internally consistent, but the file fails automated comparison (column-name/ordering/index/precision mismatch, or NaN-contaminated cumulative values), so the result file is scored WRONG/MISSING and the task earns 0.
id 3509af6e71d7 · mined from da-code dacode-dm-csv-050@s14
raw text (what the judge reads)
### Assumed output schema instead of deriving it from the stated format requirement
- **Applies when**: `task` -- the task says the deliverable file must "follow the required format" (or otherwise implies a fixed schema/column names/index/ordering/rounding) without spelling it out inline.
- **Pattern**: The agent invents its own column names, index handling, row set, or numeric conventions, then asserts the output "follows the required format" without ever locating or reproducing the authoritative spec (README/docs/template/sample file/notebook prompt in the data directory), and its verification script only re-checks its own arithmetic against itself.
- **Detection procedure**:
  1. Read the task for any words indicating a prescribed output format, and note that the exact schema is not given in the prompt text.
  2. Search the scripts for any step that inspects the workspace for a format definition or template (listing the data directory, reading a README/sample/expected-columns file) and for any check of the produced file against it (column names, order, index column, row count, rounding, date representation).
  3. Check whether the agent's chosen names/structure are traceable to such a source or were simply made up; also check whether raw-data quirks (missing values, extra/unused columns, header/index columns) were inspected before aggregation.
  4. Read the answer: does it claim format compliance while the scripts contain no comparison to an external reference?
- **Discriminator**: Fine if the schema is explicitly dictated in the task text (and the script matches it verbatim) or the script reads a template/expected-output file and validates against it; a violation is when the only "verification" recomputes the agent's own numbers and no external format source was ever consulted.
- **Consequence**: The numbers may be internally consistent, but the file fails automated comparison (column-name/ordering/index/precision mismatch, or NaN-contaminated cumulative values), so the result file is scored WRONG/MISSING and the task earns 0.
742Distribution test run on an uncleaned/unverified slice, with the required test output never reportedtaskinfiagent-dabench
Applies when
task -- a task asks for a statistical test or distribution statistics on one column, states an explicit decision rule (e.g. compare p-value to alpha), and asks that the test statistic be reported alongside the derived verdict.
Pattern
The attempt feeds the column straight into the test/statistic functions without showing how many rows survived, how missing values / non-numeric entries / sentinel or placeholder codes were handled, or whether any stated filtering applied; it then reports only the yes/no verdict and the moment values, omitting the p-value the task asked for, so nobody can check that the decision came from the intended sample.
Detection procedure
  1. Read the task and list every quantity it says to report (test p-value, statistics, rounding, format) and any implied data conditions (which rows count, how missing/invalid values are treated).
  2. In the scripts, find the exact array passed to the test and to the skew/kurtosis calls; check that dropna/dtype coercion/filtering is explicit and that the surviving len(), min/max, and dtype are printed.
  3. Compare the reported answer against the list from step 1: is the p-value (and any other requested intermediate) present, and is the verdict consistent with the stated alpha rule?
  4. Sanity-check plausibility: do the reported skew/kurtosis magnitudes look like a few extreme/sentinel points are dominating, and does the script contain no check (histogram, quantiles, outlier count) that would distinguish real data from artefacts?
Discriminator
A real violation is when the analysed vector's composition is undocumented/unfiltered and the requested test output is missing, so the verdict rests on an unverified sample; it is fine if the script prints the cleaned row count and the p-value and the answer reports them, even if the resulting verdict is borderline — an explicit, reproducible pipeline with the requested numbers is acceptable.
Consequence
The decision flips relative to the intended sample (verdict marked WRONG), and because the p-value and sample size were never reported, the error is invisible until grading.
id 73dbddace4b2 · mined from infiagent-dabench dabench-298@s14
raw text (what the judge reads)
### Distribution test run on an uncleaned/unverified slice, with the required test output never reported
- **Applies when**: `task` -- a task asks for a statistical test or distribution statistics on one column, states an explicit decision rule (e.g. compare p-value to alpha), and asks that the test statistic be reported alongside the derived verdict.
- **Pattern**: The attempt feeds the column straight into the test/statistic functions without showing how many rows survived, how missing values / non-numeric entries / sentinel or placeholder codes were handled, or whether any stated filtering applied; it then reports only the yes/no verdict and the moment values, omitting the p-value the task asked for, so nobody can check that the decision came from the intended sample.
- **Detection procedure**:
  1. Read the task and list every quantity it says to report (test p-value, statistics, rounding, format) and any implied data conditions (which rows count, how missing/invalid values are treated).
  2. In the scripts, find the exact array passed to the test and to the skew/kurtosis calls; check that dropna/dtype coercion/filtering is explicit and that the surviving `len()`, min/max, and dtype are printed.
  3. Compare the reported answer against the list from step 1: is the p-value (and any other requested intermediate) present, and is the verdict consistent with the stated alpha rule?
  4. Sanity-check plausibility: do the reported skew/kurtosis magnitudes look like a few extreme/sentinel points are dominating, and does the script contain no check (histogram, quantiles, outlier count) that would distinguish real data from artefacts?
- **Discriminator**: A real violation is when the analysed vector's composition is undocumented/unfiltered *and* the requested test output is missing, so the verdict rests on an unverified sample; it is fine if the script prints the cleaned row count and the p-value and the answer reports them, even if the resulting verdict is borderline — an explicit, reproducible pipeline with the requested numbers is acceptable.
- **Consequence**: The decision flips relative to the intended sample (verdict marked WRONG), and because the p-value and sample size were never reported, the error is invisible until grading.
743Model selection justified by in-sample scores on a subsampled training settaskda-code
Applies when
task -- the task asks for predictions on a held-out set scored by an accuracy metric, and the scripts train models and report performance numbers used to pick/weight the final model.
Pattern
The attempt drops from a proper train/validation split to fitting on a downsampled subset and then scoring the model on those very same rows (or on the full training data it was fit on), so every reported R²/RMSE is an in-sample fit statistic; ensemble weights and the "best model" choice are then set from these inflated numbers, and no unbiased estimate of test performance is ever produced.
Detection procedure
  1. Read the task to confirm the deliverable is scored on unseen data by a quantitative metric (not just file existence).
  2. In each script, trace the arrays passed to fit() and to the metric call — flag any case where the evaluation features/targets are the same object (or a subset) as the training ones, or where an earlier holdout split is silently abandoned in later scripts.
  3. Check whether the model/ensemble finally written to the submission file was ever scored on data excluded from its own training (holdout or cross-validation), and whether subsampling for "speed" discarded a large fraction of available rows without measuring the cost.
  4. Compare the reported metrics in the answer against what the final script actually computes; note if the answer cites numbers that are in-sample or come from a different, non-final model.
Discriminator
A real violation is when no out-of-fold/holdout score exists for the model that produced the submission, so quality claims and ensemble weights are unverifiable; it is fine if the script fits on a subsample or full data but reports a score on cleanly separated rows (or CV folds), and uses that score to choose the final model — small speed-motivated subsampling with an honest holdout score is acceptable.
Consequence
The submission file is well-formed but the underlying predictions are materially less accurate than a properly validated baseline, so the grader's metric threshold on the predicted column fails (file marked WRONG despite existing), while the answer confidently reports high R² values that never applied to unseen data.
id d433827830bf · mined from da-code dacode-ml-competition-008@s14
raw text (what the judge reads)
### Model selection justified by in-sample scores on a subsampled training set
- **Applies when**: `task` -- the task asks for predictions on a held-out set scored by an accuracy metric, and the scripts train models and report performance numbers used to pick/weight the final model.
- **Pattern**: The attempt drops from a proper train/validation split to fitting on a downsampled subset and then scoring the model on those very same rows (or on the full training data it was fit on), so every reported R²/RMSE is an in-sample fit statistic; ensemble weights and the "best model" choice are then set from these inflated numbers, and no unbiased estimate of test performance is ever produced.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on unseen data by a quantitative metric (not just file existence).
  2. In each script, trace the arrays passed to `fit()` and to the metric call — flag any case where the evaluation features/targets are the same object (or a subset) as the training ones, or where an earlier holdout split is silently abandoned in later scripts.
  3. Check whether the model/ensemble finally written to the submission file was ever scored on data excluded from its own training (holdout or cross-validation), and whether subsampling for "speed" discarded a large fraction of available rows without measuring the cost.
  4. Compare the reported metrics in the answer against what the final script actually computes; note if the answer cites numbers that are in-sample or come from a different, non-final model.
- **Discriminator**: A real violation is when *no* out-of-fold/holdout score exists for the model that produced the submission, so quality claims and ensemble weights are unverifiable; it is fine if the script fits on a subsample or full data but reports a score on cleanly separated rows (or CV folds), and uses that score to choose the final model — small speed-motivated subsampling with an honest holdout score is acceptable.
- **Consequence**: The submission file is well-formed but the underlying predictions are materially less accurate than a properly validated baseline, so the grader's metric threshold on the predicted column fails (file marked WRONG despite existing), while the answer confidently reports high R² values that never applied to unseen data.
744Boolean/token casing not normalized to the requested literal formattaskinfiagent-dabench
Applies when
task -- The answer template specifies a value in a particular literal form (e.g. a lowercase boolean, a specific string label, a unit or percentage form) and the script writes a language-native value directly into the answer.
Pattern
The script computes the correct result but serializes it with the programming language's default representation (e.g. Python's True/False, nan, 0.13000000001, Timestamp(...)) instead of converting it to the exact token the answer format demands, so the string comparison fails even though the analysis is right.
Detection procedure
  1. Read the task's answer format and list, for each field, the exact expected literal type/casing/precision (e.g. "boolean value" → true/false, "rounded to two decimals", specific unit).
  2. In the script, find the expression whose value is interpolated into each answer field and determine its native string rendering (e.g. an f-string of a Python/NumPy bool renders True; a NumPy bool may even render np.True_).
  3. Check whether any explicit normalization exists (e.g. str(x).lower(), bool(x) cast, f"{x:.4f}", unit conversion) before writing the answer.
  4. Compare the final answer string token-by-token against the format spec; flag any field whose casing, type spelling, or precision differs.
Discriminator
A real violation is a mismatch between the emitted literal and the specified literal (e.g. True where true is specified, 0.032 where four decimals are required). It is not a violation if the grader/spec is case- or type-agnostic, or if the script already applies an explicit normalization/formatting step that yields the required token.
Consequence
The numeric/statistical fields pass but the mis-formatted field is marked WRONG/MISSING, giving a partial score (e.g. 2/3 checks) and an overall incorrect verdict despite correct analysis.
id 864310206843 · mined from infiagent-dabench dabench-668@s14
raw text (what the judge reads)
### Boolean/token casing not normalized to the requested literal format
- **Applies when**: `task` -- The answer template specifies a value in a particular literal form (e.g. a lowercase boolean, a specific string label, a unit or percentage form) and the script writes a language-native value directly into the answer.
- **Pattern**: The script computes the correct result but serializes it with the programming language's default representation (e.g. Python's `True`/`False`, `nan`, `0.13000000001`, `Timestamp(...)`) instead of converting it to the exact token the answer format demands, so the string comparison fails even though the analysis is right.
- **Detection procedure**:
  1. Read the task's answer format and list, for each field, the exact expected literal type/casing/precision (e.g. "boolean value" → `true`/`false`, "rounded to two decimals", specific unit).
  2. In the script, find the expression whose value is interpolated into each answer field and determine its native string rendering (e.g. an f-string of a Python/NumPy bool renders `True`; a NumPy bool may even render `np.True_`).
  3. Check whether any explicit normalization exists (e.g. `str(x).lower()`, `bool(x)` cast, `f"{x:.4f}"`, unit conversion) before writing the answer.
  4. Compare the final answer string token-by-token against the format spec; flag any field whose casing, type spelling, or precision differs.
- **Discriminator**: A real violation is a mismatch between the emitted literal and the specified literal (e.g. `True` where `true` is specified, `0.032` where four decimals are required). It is *not* a violation if the grader/spec is case- or type-agnostic, or if the script already applies an explicit normalization/formatting step that yields the required token.
- **Consequence**: The numeric/statistical fields pass but the mis-formatted field is marked WRONG/MISSING, giving a partial score (e.g. 2/3 checks) and an overall incorrect verdict despite correct analysis.
745Answer-string does not literally match the requested output template (and no script backs it up)taskinfiagent-dabench
Applies when
task -- the task specifies an exact answer template with tagged fields (e.g. @field[value]), showing example values with particular quoting, casing, units, or delimiters, and expects the answer to be produced/printed by the analysis scripts.
Pattern
The attempt computes the right underlying values but emits them in a superficially different surface form than the template dictates — dropping the quotation marks shown in the example, changing case or spelling of a categorical label, adding units/spaces, reordering or renaming tags — and/or hand-types the final line instead of printing it from a saved, rerunnable script, so nothing enforces the template.
Detection procedure
  1. Copy the answer-format specification from the task verbatim and list each field, its exact delimiters, and the exact literal form of every example value (including quotes, hyphens, capitalization).
  2. Check the scripts: is there code that assembles and prints the final answer string from computed variables, and does that string template character-for-character reproduce the specification? If no script is saved at all, treat the final line as unverified.
  3. Diff the submitted answer against the specification token by token (not by meaning): quoting, case, whitespace, tag names, tag order, value vocabulary (only the allowed labels).
  4. Confirm each reported value is the requested final quantity, not an intermediate, and that categorical labels come from the allowed set exactly as written.
Discriminator
A real violation is any character-level deviation from the literal template or an answer with no script that generates it; a look-alike that is fine is a task whose spec is genuinely ambiguous about surface form (no example given) where the agent picks one reasonable rendering and its scripts print it deterministically.
Consequence
Exact-match grading fails every field even though the analysis logic and numeric conclusions were correct, yielding 0/N checks with "WRONG/MISSING" on values that look identical at a glance.
id 98e4e2e1296c · mined from infiagent-dabench dabench-550@s14
raw text (what the judge reads)
### Answer-string does not literally match the requested output template (and no script backs it up)
- **Applies when**: `task` -- the task specifies an exact answer template with tagged fields (e.g. `@field[value]`), showing example values with particular quoting, casing, units, or delimiters, and expects the answer to be produced/printed by the analysis scripts.
- **Pattern**: The attempt computes the right underlying values but emits them in a superficially different surface form than the template dictates — dropping the quotation marks shown in the example, changing case or spelling of a categorical label, adding units/spaces, reordering or renaming tags — and/or hand-types the final line instead of printing it from a saved, rerunnable script, so nothing enforces the template.
- **Detection procedure**:
  1. Copy the answer-format specification from the task verbatim and list each field, its exact delimiters, and the exact literal form of every example value (including quotes, hyphens, capitalization).
  2. Check the scripts: is there code that assembles and prints the final answer string from computed variables, and does that string template character-for-character reproduce the specification? If no script is saved at all, treat the final line as unverified.
  3. Diff the submitted answer against the specification token by token (not by meaning): quoting, case, whitespace, tag names, tag order, value vocabulary (only the allowed labels).
  4. Confirm each reported value is the requested final quantity, not an intermediate, and that categorical labels come from the allowed set exactly as written.
- **Discriminator**: A real violation is any character-level deviation from the literal template or an answer with no script that generates it; a look-alike that is fine is a task whose spec is genuinely ambiguous about surface form (no example given) where the agent picks one reasonable rendering and its scripts print it deterministically.
- **Consequence**: Exact-match grading fails every field even though the analysis logic and numeric conclusions were correct, yielding 0/N checks with "WRONG/MISSING" on values that look identical at a glance.
746Entity substitution: analyzing a dataset whose fields don't match the requested onestaskda-code
Applies when
task -- the task names specific entities/metrics/groupings (e.g., a grouping key, a measure, per-stage durations) that must exist in the input data, and the scripts pick an input file and columns themselves.
Pattern
The agent cannot find the named fields in the loaded file, so it silently substitutes unrelated columns as "equivalent" proxies (renaming categories to the requested labels), and reports success as if the requested quantities were computed; required output artifacts implied by the config/spec are also skipped.
Detection procedure
  1. List every entity the task requires (grouping variable, measure to rank by, per-stage quantity, output files/settings source).
  2. In the scripts/answer, find the file actually loaded and the columns actually used; check each required entity maps to a real column of matching semantics, not a renamed stand-in.
  3. Check whether the answer's category names/units make sense for the requested concept (e.g., counts presented as a monetary total, outcome labels presented as durations, axis labels unrelated to the chart).
  4. Verify every output artifact named or implied by the task/config spec is produced, not just the one file mentioned.
Discriminator
A genuine violation is when the reported groups/measures are semantically different from those asked for (proxy substitution, mislabeled units, wrong data source); it is fine if the columns have different names but demonstrably encode the requested quantity, with the mapping justified against the data dictionary.
Consequence
All expected outputs mismatch — values, labels, and files — so every grader check fails despite a confident "task completed" report.
id 05c3e8b39309 · mined from da-code dacode-plot-scatter-002@s14
raw text (what the judge reads)
### Entity substitution: analyzing a dataset whose fields don't match the requested ones
- **Applies when**: `task` -- the task names specific entities/metrics/groupings (e.g., a grouping key, a measure, per-stage durations) that must exist in the input data, and the scripts pick an input file and columns themselves.
- **Pattern**: The agent cannot find the named fields in the loaded file, so it silently substitutes unrelated columns as "equivalent" proxies (renaming categories to the requested labels), and reports success as if the requested quantities were computed; required output artifacts implied by the config/spec are also skipped.
- **Detection procedure**:
  1. List every entity the task requires (grouping variable, measure to rank by, per-stage quantity, output files/settings source).
  2. In the scripts/answer, find the file actually loaded and the columns actually used; check each required entity maps to a real column of matching semantics, not a renamed stand-in.
  3. Check whether the answer's category names/units make sense for the requested concept (e.g., counts presented as a monetary total, outcome labels presented as durations, axis labels unrelated to the chart).
  4. Verify every output artifact named or implied by the task/config spec is produced, not just the one file mentioned.
- **Discriminator**: A genuine violation is when the reported groups/measures are semantically different from those asked for (proxy substitution, mislabeled units, wrong data source); it is fine if the columns have different names but demonstrably encode the requested quantity, with the mapping justified against the data dictionary.
- **Consequence**: All expected outputs mismatch — values, labels, and files — so every grader check fails despite a confident "task completed" report.
747Validating a predictive model only in-sample (and never sanity-checking the prediction distribution)taskda-code
Applies when
task -- the task asks for predictions on an unlabeled test set and the scripts fit a model whose quality is summarized by fit/error statistics.
Pattern
The attempt reports R²/MAE/RMSE computed on the same rows the model was fit on (or on the full labeled set) instead of a held-out split or cross-validation, treats that optimistic number as evidence of success, and never compares the predicted values' distribution (mean, spread, min/max) with the observed target distribution or with a trivial baseline (mean/median predictor). It also typically drops or never tests high-signal columns (identifiers, dates, categorical metadata) while keeping only the numeric ones that were convenient to feed the model.
Detection procedure
  1. Read the task to confirm the deliverable is out-of-sample predictions and note any format/row-count/column-name constraints.
  2. In the scripts, locate where the metric is computed and check whether the data passed to score/metric are the same rows passed to fit; look for any train_test_split, CV, or holdout that is actually used for the reported number.
  3. Check whether any comparison against a baseline (predict the training mean) or a distribution sanity check (predicted std/min/max vs. target std/min/max) is performed before writing the output file.
  4. Read the answer: if the only quality evidence is an in-sample score and the reported prediction spread is much narrower than the target's, flag it.
Discriminator
A real violation is when no out-of-sample estimate exists anywhere (or the quoted number is explicitly the training-set score); it is fine if the script does CV/holdout and additionally prints a training score for reference, or if the model is deliberately unregularized but the reported metric comes from unseen data.
Consequence
The reported score is inflated by memorization, the model may be no better than (or worse than) predicting the mean, and the submitted prediction column—over-shrunk toward the mean and missing the strongest predictors—falls outside the grader's accepted error/correlation tolerance, so the output file is judged wrong.
id ddba99038125 · mined from da-code dacode-ml-regression-004@s14
raw text (what the judge reads)
### Validating a predictive model only in-sample (and never sanity-checking the prediction distribution)
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test set and the scripts fit a model whose quality is summarized by fit/error statistics.
- **Pattern**: The attempt reports R²/MAE/RMSE computed on the same rows the model was fit on (or on the full labeled set) instead of a held-out split or cross-validation, treats that optimistic number as evidence of success, and never compares the predicted values' distribution (mean, spread, min/max) with the observed target distribution or with a trivial baseline (mean/median predictor). It also typically drops or never tests high-signal columns (identifiers, dates, categorical metadata) while keeping only the numeric ones that were convenient to feed the model.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is out-of-sample predictions and note any format/row-count/column-name constraints.
  2. In the scripts, locate where the metric is computed and check whether the data passed to `score`/`metric` are the same rows passed to `fit`; look for any `train_test_split`, CV, or holdout that is actually used for the reported number.
  3. Check whether any comparison against a baseline (predict the training mean) or a distribution sanity check (predicted std/min/max vs. target std/min/max) is performed before writing the output file.
  4. Read the answer: if the only quality evidence is an in-sample score and the reported prediction spread is much narrower than the target's, flag it.
- **Discriminator**: A real violation is when *no* out-of-sample estimate exists anywhere (or the quoted number is explicitly the training-set score); it is fine if the script does CV/holdout and additionally prints a training score for reference, or if the model is deliberately unregularized but the reported metric comes from unseen data.
- **Consequence**: The reported score is inflated by memorization, the model may be no better than (or worse than) predicting the mean, and the submitted prediction column—over-shrunk toward the mean and missing the strongest predictors—falls outside the grader's accepted error/correlation tolerance, so the output file is judged wrong.
748Accepting a degenerate clustering solution without sanity-checking cluster sizes or outlier influencetaskda-code
Applies when
task -- the task asks for an unsupervised partition into "an appropriate number of groups" and the script auto-selects the number of clusters by maximizing an internal index (silhouette, etc.) on raw standardized features.
Pattern
The script feeds heavy-tailed, unscaled-in-distribution features straight into a centroid-based algorithm, picks k by the single best internal score, and then reports/writes a partition in which one or more clusters contain a handful (or one) of records — i.e. the "clusters" are really outliers — without questioning whether the partition is meaningful, comparing alternative k values or algorithms, or checking that clusters are interpretable/balanced enough to answer the stated business question.
Detection procedure
  1. Read the task to see what the grouping must support (e.g., ranking or triaging entities into usable groups) and whether any structure of the output is implied.
  2. In the script, check whether skew/outliers are addressed before clustering (log/robust transform, outlier handling, PCA) and whether k selection combines more than one signal (elbow + silhouette + cluster-size/interpretability).
  3. In the answer, inspect the reported cluster size distribution and the per-cluster profile: flag if any cluster is a singleton or a tiny fraction of the data, or if no per-cluster summary was produced at all.
  4. Verify the saved file's contents match what the task asked to be clustered (correct feature vector representation, correct row count, correct column naming/ordering) and that the label column is consistent with the reported distribution.
Discriminator
A genuinely small cluster is fine if the script explicitly examined it (reported centroid/feature means, argued it is a real extreme-value group) and cross-checked k against other criteria; the violation is selecting k purely by an internal score maximum and shipping singleton/near-singleton clusters with no diagnostic, transformation, or alternative considered.
Consequence
The saved label file disagrees with any reasonable reference partition (wrong number of groups, wrong assignment for most rows), so the file-comparison check fails even though the pipeline "ran successfully."
id 8371b94d25f9 · mined from da-code dacode-ml-cluster-013@s14
raw text (what the judge reads)
### Accepting a degenerate clustering solution without sanity-checking cluster sizes or outlier influence
- **Applies when**: `task` -- the task asks for an unsupervised partition into "an appropriate number of groups" and the script auto-selects the number of clusters by maximizing an internal index (silhouette, etc.) on raw standardized features.
- **Pattern**: The script feeds heavy-tailed, unscaled-in-distribution features straight into a centroid-based algorithm, picks k by the single best internal score, and then reports/writes a partition in which one or more clusters contain a handful (or one) of records — i.e. the "clusters" are really outliers — without questioning whether the partition is meaningful, comparing alternative k values or algorithms, or checking that clusters are interpretable/balanced enough to answer the stated business question.
- **Detection procedure**:
  1. Read the task to see what the grouping must support (e.g., ranking or triaging entities into usable groups) and whether any structure of the output is implied.
  2. In the script, check whether skew/outliers are addressed before clustering (log/robust transform, outlier handling, PCA) and whether k selection combines more than one signal (elbow + silhouette + cluster-size/interpretability).
  3. In the answer, inspect the reported cluster size distribution and the per-cluster profile: flag if any cluster is a singleton or a tiny fraction of the data, or if no per-cluster summary was produced at all.
  4. Verify the saved file's contents match what the task asked to be clustered (correct feature vector representation, correct row count, correct column naming/ordering) and that the label column is consistent with the reported distribution.
- **Discriminator**: A genuinely small cluster is fine if the script explicitly examined it (reported centroid/feature means, argued it is a real extreme-value group) and cross-checked k against other criteria; the violation is selecting k purely by an internal score maximum and shipping singleton/near-singleton clusters with no diagnostic, transformation, or alternative considered.
- **Consequence**: The saved label file disagrees with any reasonable reference partition (wrong number of groups, wrong assignment for most rows), so the file-comparison check fails even though the pipeline "ran successfully."
749Ambiguous or unverified sample axis when computing a per-group distribution statistictaskinfiagent-dabench
Applies when
task -- the task asks for a distribution statistic (skewness, kurtosis, variance, entropy, etc.) computed per entity while also fixing another dimension (a single year/segment/condition), so the reviewer must know which rows form each entity's sample.
Pattern
The attempt never makes explicit which axis supplies the multiple observations per entity (e.g. it collapses to one value per entity and then skews across entities, or aggregates over the wrong dimension, or silently drops rows with missing values), and it reports a winner without any reproducible script or intermediate table showing per-entity sample sizes and statistic values.
Detection procedure
  1. From the task statement, write down the intended unit of aggregation (one statistic per entity) and the intended sample within each unit (which rows/columns vary), plus any stated library/definition constraint.
  2. In the scripts, locate the groupby/axis/slice used and confirm each group yields ≥3 non-null values and that the statistic is applied along that axis, with the required definition/flag (e.g. Fisher/bias options) explicitly set rather than left to defaults or a different function.
  3. Check that the script prints a ranked table (entity, n, statistic) and that the reported answer is the argmax of that table, not an intermediate or eyeballed value.
  4. If no script or ranked output exists, or if the group sample would have length 1 (making the statistic undefined/NaN), flag the attempt as unverifiable.
Discriminator
A real violation is when the per-entity sample is undefined, degenerate, or drawn from the wrong axis/subset, or when the definition flag contradicts the stated constraint; a look-alike that is fine explicitly filters to the stated condition, groups on the correct key, shows counts per group, and uses the required function/parameterization even if the final winner is surprising.
Consequence
The statistic is computed over the wrong sample, so the ranking is arbitrary and the reported entity mismatches the expected one — the grader records 0/1 on the exact-name check.
id 519238c47b94 · mined from infiagent-dabench dabench-252@s14
raw text (what the judge reads)
### Ambiguous or unverified sample axis when computing a per-group distribution statistic
- **Applies when**: `task` -- the task asks for a distribution statistic (skewness, kurtosis, variance, entropy, etc.) computed *per entity* while also fixing another dimension (a single year/segment/condition), so the reviewer must know which rows form each entity's sample.
- **Pattern**: The attempt never makes explicit which axis supplies the multiple observations per entity (e.g. it collapses to one value per entity and then skews across entities, or aggregates over the wrong dimension, or silently drops rows with missing values), and it reports a winner without any reproducible script or intermediate table showing per-entity sample sizes and statistic values.
- **Detection procedure**:
  1. From the task statement, write down the intended unit of aggregation (one statistic per entity) and the intended sample within each unit (which rows/columns vary), plus any stated library/definition constraint.
  2. In the scripts, locate the groupby/axis/slice used and confirm each group yields ≥3 non-null values and that the statistic is applied along that axis, with the required definition/flag (e.g. Fisher/bias options) explicitly set rather than left to defaults or a different function.
  3. Check that the script prints a ranked table (entity, n, statistic) and that the reported answer is the argmax of that table, not an intermediate or eyeballed value.
  4. If no script or ranked output exists, or if the group sample would have length 1 (making the statistic undefined/NaN), flag the attempt as unverifiable.
- **Discriminator**: A real violation is when the per-entity sample is undefined, degenerate, or drawn from the wrong axis/subset, or when the definition flag contradicts the stated constraint; a look-alike that is fine explicitly filters to the stated condition, groups on the correct key, shows counts per group, and uses the required function/parameterization even if the final winner is surprising.
- **Consequence**: The statistic is computed over the wrong sample, so the ranking is arbitrary and the reported entity mismatches the expected one — the grader records 0/1 on the exact-name check.
750Truncating an identifier value to match a format template instead of reporting the full valuetaskinfiagent-dabench
Applies when
task -- the answer includes a date/ID/label whose format is described by a template (e.g. a date pattern) and the script derives it by slicing or reformatting a value pulled from the data.
Pattern
The agent takes the format hint literally as a truncation instruction, chopping off part of the identified value (e.g. dropping the day component, trailing digits, or a suffix) rather than reporting the exact record that the analysis identified; the stated format template is treated as authoritative over the actual granularity of the underlying record.
Detection procedure
1. Read the task to see what entity is asked for (a specific row/date/key) and at what granularity the data stores it. 2. In the scripts, find where the identified value is converted for output and look for string slicing, strftime/str[:n], rounding of keys, or period conversion. 3. Check whether the emitted value still uniquely identifies the record found; if several records in the data map to the same emitted string, the value has lost information. 4. Compare the printed diagnostic value (full record identifier) with the value actually written into the final answer — a mismatch in granularity is the flag.
Discriminator
A real violation is when the truncation discards information needed to identify the record (many rows share the emitted string) or contradicts the granularity the analysis operated on; it is fine when the data itself is natively at that coarser granularity, or the task explicitly says to aggregate/round to it, or the template is merely an example of layout rather than a truncation instruction.
Consequence
The graded field for the identifier fails string/value comparison against the full expected value even though the dependent numeric computation, done on the correct row, matches — yielding a partially-correct, overall-wrong result.
id c318af5b8f14 · mined from infiagent-dabench dabench-572@s14
raw text (what the judge reads)
### Truncating an identifier value to match a format template instead of reporting the full value
- **Applies when**: `task` -- the answer includes a date/ID/label whose format is described by a template (e.g. a date pattern) and the script derives it by slicing or reformatting a value pulled from the data.
- **Pattern**: The agent takes the format hint literally as a truncation instruction, chopping off part of the identified value (e.g. dropping the day component, trailing digits, or a suffix) rather than reporting the exact record that the analysis identified; the stated format template is treated as authoritative over the actual granularity of the underlying record.
- **Detection procedure**: 1. Read the task to see what entity is asked for (a specific row/date/key) and at what granularity the data stores it. 2. In the scripts, find where the identified value is converted for output and look for string slicing, `strftime`/`str[:n]`, rounding of keys, or period conversion. 3. Check whether the emitted value still uniquely identifies the record found; if several records in the data map to the same emitted string, the value has lost information. 4. Compare the printed diagnostic value (full record identifier) with the value actually written into the final answer — a mismatch in granularity is the flag.
- **Discriminator**: A real violation is when the truncation discards information needed to identify the record (many rows share the emitted string) or contradicts the granularity the analysis operated on; it is fine when the data itself is natively at that coarser granularity, or the task explicitly says to aggregate/round to it, or the template is merely an example of layout rather than a truncation instruction.
- **Consequence**: The graded field for the identifier fails string/value comparison against the full expected value even though the dependent numeric computation, done on the correct row, matches — yielding a partially-correct, overall-wrong result.
751Missing post-write verification of the prediction artifact against the provided templatetaskda-code
Applies when
task -- the task requires writing predictions/results to a file whose format is specified by an example/sample file and whose row count is fixed by the test set.
Pattern
The attempt builds a model and calls a write function, then asserts in prose that the output "matches the sample format," without ever reading the written file back and comparing it, element-for-element, to the template's header text, row count, column count, index presence, and allowed value domain (label encoding vs. probabilities, dtype, ordering aligned to the test rows).
Detection procedure
  1. From the task, list the hard output constraints: exact filename, exact column name(s), number of rows expected (= number of test rows), and the value domain shown in the sample file.
  2. In the scripts, locate the write call and check for index=False/header handling, and then look for a subsequent re-read of the file with explicit assertions on shape, header, dtype, and unique values, plus a check that predictions are in the same order as the test rows (no shuffling, no dropped rows from imputation/dropna).
  3. In the answer, check whether reported numbers (e.g., predicted class counts) sum to the expected test row count and whether any verification output is quoted, not just claimed.
  4. Flag if the only evidence of correctness is narrative text such as "saved with format matching the sample."
Discriminator
A real violation is an attempt with no executed read-back/assertions (or with counts that don't reconcile with the test set size, or rows lost/reordered by preprocessing). It is not a violation if the script prints the head and shape of the reloaded file and compares its header/value set to the sample file, even if written concisely.
Consequence
The grader cannot match the submitted file to the expected one — it reports the target file as WRONG/MISSING (bad header, extra index column, wrong row count, or misaligned/mis-encoded labels), scoring 0 despite plausible-looking validation metrics in the write-up.
id 5624a72429a0 · mined from da-code dacode-ml-binary-013@s14
raw text (what the judge reads)
### Missing post-write verification of the prediction artifact against the provided template
- **Applies when**: `task` -- the task requires writing predictions/results to a file whose format is specified by an example/sample file and whose row count is fixed by the test set.
- **Pattern**: The attempt builds a model and calls a write function, then asserts in prose that the output "matches the sample format," without ever reading the written file back and comparing it, element-for-element, to the template's header text, row count, column count, index presence, and allowed value domain (label encoding vs. probabilities, dtype, ordering aligned to the test rows).
- **Detection procedure**:
  1. From the task, list the hard output constraints: exact filename, exact column name(s), number of rows expected (= number of test rows), and the value domain shown in the sample file.
  2. In the scripts, locate the write call and check for `index=False`/header handling, and then look for a *subsequent* re-read of the file with explicit assertions on shape, header, dtype, and unique values, plus a check that predictions are in the same order as the test rows (no shuffling, no dropped rows from imputation/`dropna`).
  3. In the answer, check whether reported numbers (e.g., predicted class counts) sum to the expected test row count and whether any verification output is quoted, not just claimed.
  4. Flag if the only evidence of correctness is narrative text such as "saved with format matching the sample."
- **Discriminator**: A real violation is an attempt with no executed read-back/assertions (or with counts that don't reconcile with the test set size, or rows lost/reordered by preprocessing). It is *not* a violation if the script prints the head and shape of the reloaded file and compares its header/value set to the sample file, even if written concisely.
- **Consequence**: The grader cannot match the submitted file to the expected one — it reports the target file as WRONG/MISSING (bad header, extra index column, wrong row count, or misaligned/mis-encoded labels), scoring 0 despite plausible-looking validation metrics in the write-up.
752Output file schema does not exactly match the requested specificationtaskda-code
Applies when
task -- the task specifies an output artifact with a named column (or column set), row count, or value encoding, and the scripts write that file at the end.
Pattern
The attempt writes the result file with extra columns (e.g., an index/id column), renamed/differently-cased columns, a different value encoding than the source labels (0/1 or re-mapped strings instead of the original class labels), or a different row order/count than the evaluation set, while the narrative asserts the file is "correct format."
Detection procedure
  1. From the task text, extract the exact required filename, required column name(s), expected number of rows, and expected label representation.
  2. In the script, locate the final write call (to_csv/to_json/etc.) and read the DataFrame construction and its arguments: which columns are included, index=True/False, any header renaming, and whether predictions were inverse-transformed to the original label strings.
  3. Compare that written schema against the extracted requirement, and cross-check the answer's own "Columns: ..." / row-count claim against it.
  4. Flag if the written columns are a superset/subset/renaming of what was asked, if the index is emitted as an unnamed or extra column, or if the row count differs from the evaluation-set size.
Discriminator
A real violation is a concrete mismatch between what the code writes and what the task literally requested (extra id column, encoded labels, wrong header, wrong row count). It is not a violation if the extra content is explicitly permitted by the task, or if the required column is present with the required name and values and the only difference is cosmetic formatting the task left unspecified (e.g., quoting, float precision on non-required fields).
Consequence
The grader reads the file by expected column/shape and marks it WRONG/MISSING even when the underlying predictions are accurate, yielding 0 on the file check regardless of model quality.
id 84c16f993dd4 · mined from da-code dacode-ml-binary-009@s14
raw text (what the judge reads)
### Output file schema does not exactly match the requested specification
- **Applies when**: `task` -- the task specifies an output artifact with a named column (or column set), row count, or value encoding, and the scripts write that file at the end.
- **Pattern**: The attempt writes the result file with extra columns (e.g., an index/id column), renamed/differently-cased columns, a different value encoding than the source labels (0/1 or re-mapped strings instead of the original class labels), or a different row order/count than the evaluation set, while the narrative asserts the file is "correct format."
- **Detection procedure**:
  1. From the task text, extract the exact required filename, required column name(s), expected number of rows, and expected label representation.
  2. In the script, locate the final write call (`to_csv`/`to_json`/etc.) and read the DataFrame construction and its arguments: which columns are included, `index=True/False`, any header renaming, and whether predictions were inverse-transformed to the original label strings.
  3. Compare that written schema against the extracted requirement, and cross-check the answer's own "Columns: ..." / row-count claim against it.
  4. Flag if the written columns are a superset/subset/renaming of what was asked, if the index is emitted as an unnamed or extra column, or if the row count differs from the evaluation-set size.
- **Discriminator**: A real violation is a concrete mismatch between what the code writes and what the task literally requested (extra `id` column, encoded labels, wrong header, wrong row count). It is *not* a violation if the extra content is explicitly permitted by the task, or if the required column is present with the required name and values and the only difference is cosmetic formatting the task left unspecified (e.g., quoting, float precision on non-required fields).
- **Consequence**: The grader reads the file by expected column/shape and marks it WRONG/MISSING even when the underlying predictions are accurate, yielding 0 on the file check regardless of model quality.
753Fabricating an ad-hoc metric instead of using the definition the task inheritstaskda-code
Applies when
task -- the task asks to visualize/report "performance"/"the result" over a period while referring back to a previously computed quantity or a config file, without restating the formula in the prompt.
Pattern
The agent cannot find (or does not look for) the authoritative definition, so it invents its own composite score (e.g., weighted sums of counts it happened to have), documents the invention confidently, and also skips emitting the auxiliary artifacts (serialized plot spec / numeric result array) that the pipeline expects.
Detection procedure
  1. Read the task and list every referenced but unstated quantity ("the specified period", "each entity's performance", settings files) plus every output file implied by the pipeline.
  2. In the scripts, check whether each such quantity is loaded from the config/upstream artifact or hard-coded/derived by the agent's own reasoning; check whether every expected output file is written.
  3. In the answer, look for phrases that define a metric with a rationale the task never gave (e.g., "weighted more heavily because...", self-chosen top-N, self-chosen date bounds).
  4. Flag if any core quantity is agent-invented or any required artifact is absent from the script's write calls.
Discriminator
A real violation is a metric/period/entity set chosen by the agent with no traceable source in the prompt, config, or prior step output; it is fine if the agent reads the definition from the provided config/upstream file (or the prompt explicitly leaves the metric to the analyst) and still writes all requested artifacts.
Consequence
The chart's values, entity ranking, axis labels and title diverge from the reference, and missing side artifacts fail existence/value checks — here all expected files were judged wrong or missing (0/3).
id 6eb0cb35b8f5 · mined from da-code dacode-plot-bar-006@s14
raw text (what the judge reads)
### Fabricating an ad-hoc metric instead of using the definition the task inherits
- **Applies when**: `task` -- the task asks to visualize/report "performance"/"the result" over a period while referring back to a previously computed quantity or a config file, without restating the formula in the prompt.
- **Pattern**: The agent cannot find (or does not look for) the authoritative definition, so it invents its own composite score (e.g., weighted sums of counts it happened to have), documents the invention confidently, and also skips emitting the auxiliary artifacts (serialized plot spec / numeric result array) that the pipeline expects.
- **Detection procedure**:
  1. Read the task and list every referenced but unstated quantity ("the specified period", "each entity's performance", settings files) plus every output file implied by the pipeline.
  2. In the scripts, check whether each such quantity is loaded from the config/upstream artifact or hard-coded/derived by the agent's own reasoning; check whether every expected output file is written.
  3. In the answer, look for phrases that define a metric with a rationale the task never gave (e.g., "weighted more heavily because...", self-chosen top-N, self-chosen date bounds).
  4. Flag if any core quantity is agent-invented or any required artifact is absent from the script's write calls.
- **Discriminator**: A real violation is a metric/period/entity set chosen by the agent with no traceable source in the prompt, config, or prior step output; it is fine if the agent reads the definition from the provided config/upstream file (or the prompt explicitly leaves the metric to the analyst) and still writes all requested artifacts.
- **Consequence**: The chart's values, entity ranking, axis labels and title diverge from the reference, and missing side artifacts fail existence/value checks — here all expected files were judged wrong or missing (0/3).
754Submission file not verified against the required template (row count, ids, header) — answer pasted inline insteadtaskda-code
Applies when
task -- the task requires writing predictions to a named output file whose format is defined by a provided sample/template file.
Pattern
The attempt reports predictions as inline text in the answer (often truncated) without evidence that the required file was written to the expected path, and never checks that the produced file's header, id column, row count and ordering match the template and the test set.
Detection procedure
  1. From the task, note the exact output filename, the template's column names/order, and the number of test rows expected.
  2. In the scripts, look for an explicit write to that exact path (and no later overwrite), plus an assertion/print of the output's shape, column names, and id set compared to the template/test ids.
  3. Inspect the answer artifact: count its rows and compare to the test-set size; check the header spelling/case matches the template; check the last line is complete and values are in the valid range (e.g., probabilities in [0,1], per-row sums as required).
  4. Flag if the file write/verification is absent, or if the reported content is truncated, mis-headed, or has fewer/different ids than the test set.
Discriminator
A real violation is missing/unverified file output or an id/row/header mismatch (including truncated or partially written content); a look-alike that is fine is an answer that merely previews a few rows while the scripts demonstrably write the full file and print shape/id-match checks against the template.
Consequence
The grader finds the expected output file missing or structurally mismatched (wrong header, wrong/incomplete id set), so the submission is scored as WRONG/MISSING regardless of model quality.
id f29ab56308ae · mined from da-code dacode-ml-competition-003@s14
raw text (what the judge reads)
### Submission file not verified against the required template (row count, ids, header) — answer pasted inline instead
- **Applies when**: `task` -- the task requires writing predictions to a named output file whose format is defined by a provided sample/template file.
- **Pattern**: The attempt reports predictions as inline text in the answer (often truncated) without evidence that the required file was written to the expected path, and never checks that the produced file's header, id column, row count and ordering match the template and the test set.
- **Detection procedure**:
  1. From the task, note the exact output filename, the template's column names/order, and the number of test rows expected.
  2. In the scripts, look for an explicit write to that exact path (and no later overwrite), plus an assertion/print of the output's `shape`, column names, and id set compared to the template/test ids.
  3. Inspect the answer artifact: count its rows and compare to the test-set size; check the header spelling/case matches the template; check the last line is complete and values are in the valid range (e.g., probabilities in [0,1], per-row sums as required).
  4. Flag if the file write/verification is absent, or if the reported content is truncated, mis-headed, or has fewer/different ids than the test set.
- **Discriminator**: A real violation is missing/unverified file output or an id/row/header mismatch (including truncated or partially written content); a look-alike that is fine is an answer that merely *previews* a few rows while the scripts demonstrably write the full file and print shape/id-match checks against the template.
- **Consequence**: The grader finds the expected output file missing or structurally mismatched (wrong header, wrong/incomplete id set), so the submission is scored as WRONG/MISSING regardless of model quality.
755Label vocabulary and row-coverage of the prediction file not verified against the source datataskda-code
Applies when
task -- the task asks for a prediction/label file with a named column, and the label values come from a categorical field that already exists in the provided training data.
Pattern
The agent invents or normalizes its own category strings (lower-casing, re-wording, collapsing/dropping rare classes) and/or writes fewer rows than the evaluation set, instead of emitting exactly the distinct label strings as they appear in the source column, one row per required key, in the required order and with the required column name(s).
Detection procedure
  1. From the task/README, note the exact output filename, required column name(s), and the number of rows expected (one per record in the evaluation file).
  2. In the scripts, find where the label set is derived: check that predicted values are mapped back to df[label_col].unique() verbatim (no .lower(), .strip(), hand-typed string literals, or class subsetting) and that the written frame is joined/aligned to the full evaluation key list.
  3. Compare the submitted file's distinct values, character-for-character, with the distinct values of the label column in the provided data; compare row count and key set with the evaluation file.
  4. Flag if any emitted value is not an exact member of the source label set, if any source class is impossible to emit, or if row/key counts differ.
Discriminator
A real violation is a mismatch in the emitted strings or row set (e.g., different casing/spelling, missing classes, truncated or partial output). It is not a violation if the strings and keys match exactly and only the predicted class assignment is imperfect, or if the task explicitly prescribes a different encoding (e.g., integer codes) that the agent follows.
Consequence
The grader's exact-match/merge against the reference fails for every row (or for whole classes), so accuracy collapses to near zero and the file is reported as WRONG/MISSING despite a plausible-looking model.
id 9d92090dac58 · mined from da-code dacode-ml-multi-003@s14
raw text (what the judge reads)
### Label vocabulary and row-coverage of the prediction file not verified against the source data
- **Applies when**: `task` -- the task asks for a prediction/label file with a named column, and the label values come from a categorical field that already exists in the provided training data.
- **Pattern**: The agent invents or normalizes its own category strings (lower-casing, re-wording, collapsing/dropping rare classes) and/or writes fewer rows than the evaluation set, instead of emitting exactly the distinct label strings as they appear in the source column, one row per required key, in the required order and with the required column name(s).
- **Detection procedure**:
  1. From the task/README, note the exact output filename, required column name(s), and the number of rows expected (one per record in the evaluation file).
  2. In the scripts, find where the label set is derived: check that predicted values are mapped back to `df[label_col].unique()` verbatim (no `.lower()`, `.strip()`, hand-typed string literals, or class subsetting) and that the written frame is joined/aligned to the full evaluation key list.
  3. Compare the submitted file's distinct values, character-for-character, with the distinct values of the label column in the provided data; compare row count and key set with the evaluation file.
  4. Flag if any emitted value is not an exact member of the source label set, if any source class is impossible to emit, or if row/key counts differ.
- **Discriminator**: A real violation is a mismatch in the *emitted strings or row set* (e.g., different casing/spelling, missing classes, truncated or partial output). It is not a violation if the strings and keys match exactly and only the predicted class assignment is imperfect, or if the task explicitly prescribes a different encoding (e.g., integer codes) that the agent follows.
- **Consequence**: The grader's exact-match/merge against the reference fails for every row (or for whole classes), so accuracy collapses to near zero and the file is reported as WRONG/MISSING despite a plausible-looking model.
756Template/reference format provided but never enforced programmaticallytaskda-code
Applies when
task -- The task supplies a template or example output file (or an explicit format spec) that the deliverable must match, and the scripts write their own file from scratch.
Pattern
The attempt derives labels, ordering, precision and shape from its own intermediate objects (e.g. casting keys to timestamps, inventing column names, applying self-chosen rounding), then only prints the template beside the result for eyeball comparison instead of asserting conformity — so mismatched row/column labels, extra/missing rows or altered value precision go unnoticed.
Detection procedure
  1. In the task statement, note that a template/reference file exists and that the output "must match the format".
  2. In the scripts, find where the output file is written and check whether the template is read and used to drive the output (reindexing to the template's row/column labels, reusing its header strings, matching its numeric precision) — or merely printed.
  3. Check for an explicit conformity assertion: same shape, identical column names, identical row-key strings, same dtype/precision as template; absence of any assert/comparison on these is the red flag.
  4. In the answer, look for claims of "matches template structure" that are not backed by a printed diff or equality check.
Discriminator
A real violation is when nothing in the code guarantees label/shape/precision equality with the template (or the code visibly transforms keys into a different representation than the template's). Fine look-alikes: the script reindexes/reorders to the template axes, or prints an explicit equality check on columns, index and shape and it passes.
Consequence
The grader's file comparison fails on keys, headers, extra/missing cells or rounded values even if the underlying aggregation logic is right, scoring 0 for the expected output file.
id 43a13e1e988b · mined from da-code dacode-dm-csv-044@s14
raw text (what the judge reads)
### Template/reference format provided but never enforced programmatically
- **Applies when**: `task` -- The task supplies a template or example output file (or an explicit format spec) that the deliverable must match, and the scripts write their own file from scratch.
- **Pattern**: The attempt derives labels, ordering, precision and shape from its own intermediate objects (e.g. casting keys to timestamps, inventing column names, applying self-chosen rounding), then only *prints* the template beside the result for eyeball comparison instead of asserting conformity — so mismatched row/column labels, extra/missing rows or altered value precision go unnoticed.
- **Detection procedure**:
  1. In the task statement, note that a template/reference file exists and that the output "must match the format".
  2. In the scripts, find where the output file is written and check whether the template is *read and used* to drive the output (reindexing to the template's row/column labels, reusing its header strings, matching its numeric precision) — or merely printed.
  3. Check for an explicit conformity assertion: same shape, identical column names, identical row-key strings, same dtype/precision as template; absence of any `assert`/comparison on these is the red flag.
  4. In the answer, look for claims of "matches template structure" that are not backed by a printed diff or equality check.
- **Discriminator**: A real violation is when nothing in the code guarantees label/shape/precision equality with the template (or the code visibly transforms keys into a different representation than the template's). Fine look-alikes: the script reindexes/reorders to the template axes, or prints an explicit equality check on columns, index and shape and it passes.
- **Consequence**: The grader's file comparison fails on keys, headers, extra/missing cells or rounded values even if the underlying aggregation logic is right, scoring 0 for the expected output file.
757Group masks built from a naive null test without verifying missing-value sentinels and a complete row partitiontaskinfiagent-dabench
Applies when
task -- the task asks to split rows into "missing" vs "non-missing" groups on one column and compare an aggregate/statistic of another column between them.
Pattern
The attempt calls a single default null check (e.g. isna()/notna()) on the raw loaded column and aggregates the target column without checking how missingness is actually encoded (empty strings, "NA", "None", "-", whitespace, "nan" as text) or whether the target column is truly numeric; rows with sentinel strings land in the wrong group, or non-numeric/blank target values are silently dropped, shifting both group means.
Detection procedure
1. From the task, note the exact two groups required and that they must exhaust the dataset. 2. In the scripts, check whether the loader/mask handles alternative missing encodings and dtype of the split column, and whether the target column was explicitly coerced to numeric with a report of how many values failed. 3. Verify the script prints len(group_a), len(group_b), their sum vs len(df), and the count of non-null target values used in each mean. 4. Compare the reported means/counts against those diagnostics; if no counts are printed or sum ≠ total rows, flag as unverified.
Discriminator
A real violation is when group membership or the aggregated values depend on an unexamined encoding/dtype assumption and no count/partition sanity check was printed; it is fine if the script inspects the raw unique values / dtypes, shows that only genuine NaNs exist, and confirms the two group sizes add up to the full row count.
Consequence
Both group means (and hence the reported statistic) are off by a few percent while the significance conclusion still looks plausible, so the numeric checks fail even though the answer format and p-value direction appear right.
id 0a5291c29ff3 · mined from infiagent-dabench dabench-297@s14
raw text (what the judge reads)
### Group masks built from a naive null test without verifying missing-value sentinels and a complete row partition
- **Applies when**: `task` -- the task asks to split rows into "missing" vs "non-missing" groups on one column and compare an aggregate/statistic of another column between them.
- **Pattern**: The attempt calls a single default null check (e.g. `isna()`/`notna()`) on the raw loaded column and aggregates the target column without checking how missingness is actually encoded (empty strings, `"NA"`, `"None"`, `"-"`, whitespace, `"nan"` as text) or whether the target column is truly numeric; rows with sentinel strings land in the wrong group, or non-numeric/blank target values are silently dropped, shifting both group means.
- **Detection procedure**: 1. From the task, note the exact two groups required and that they must exhaust the dataset. 2. In the scripts, check whether the loader/mask handles alternative missing encodings and dtype of the split column, and whether the target column was explicitly coerced to numeric with a report of how many values failed. 3. Verify the script prints `len(group_a)`, `len(group_b)`, their sum vs `len(df)`, and the count of non-null target values used in each mean. 4. Compare the reported means/counts against those diagnostics; if no counts are printed or sum ≠ total rows, flag as unverified.
- **Discriminator**: A real violation is when group membership or the aggregated values depend on an unexamined encoding/dtype assumption and no count/partition sanity check was printed; it is fine if the script inspects the raw unique values / dtypes, shows that only genuine NaNs exist, and confirms the two group sizes add up to the full row count.
- **Consequence**: Both group means (and hence the reported statistic) are off by a few percent while the significance conclusion still looks plausible, so the numeric checks fail even though the answer format and p-value direction appear right.
758Missing required output artifacts / self-invented spec instead of the referenced guidancetaskda-code
Applies when
task -- the task points to an external spec (e.g., a guidance/README file) and/or an evaluation harness expects a fixed set of result files (image, array, serialized plot data, metrics file).
Pattern
The script writes only the one artifact that is obvious from the prompt text (here, the figure) and never emits the other expected files; the analytical definitions (which subset to filter, which categories to count, their order) are invented by the agent from plausible-sounding heuristics rather than read from the referenced spec, and the answer describes assumptions ("filtered to X", "categories defined as Y") that the task never stated.
Detection procedure
  1. From the task statement, list every output artifact and every rule the task delegates to an external document; check the script actually opens/echoes that document (or that its rules are quoted) rather than inferring them.
  2. Grep the script for write/save calls (savefig, np.save, to_json, to_csv, json.dump) and compare the set of produced paths/filenames against the required artifact list — any required artifact with no corresponding write call is a failure.
  3. Read the answer for any filtering thresholds, category definitions, or orderings that do not appear in the task text; if present, verify each is traceable to the spec, not to a comment the agent wrote for itself.
  4. Sanity-check the reported counts/shapes against the raw data size and the number of requested categories (e.g., are all requested groups present and non-degenerate?).
Discriminator
A real violation is a missing required file or a decision rule with no source in the task/spec. It is fine if the agent derives extra intermediate values or adds cosmetic choices (dpi, start angle, label format) that the spec leaves open, as long as every mandated artifact is written and every substantive rule is sourced.
Consequence
The grader reports the unwritten artifacts as MISSING and, because the group definitions/filters differ from the spec, the produced artifact's values also mismatch — all checks fail even though the script runs without error.
id 340f57b1980b · mined from da-code dacode-plot-pie-005@s14
raw text (what the judge reads)
### Missing required output artifacts / self-invented spec instead of the referenced guidance
- **Applies when**: `task` -- the task points to an external spec (e.g., a guidance/README file) and/or an evaluation harness expects a fixed set of result files (image, array, serialized plot data, metrics file).
- **Pattern**: The script writes only the one artifact that is obvious from the prompt text (here, the figure) and never emits the other expected files; the analytical definitions (which subset to filter, which categories to count, their order) are invented by the agent from plausible-sounding heuristics rather than read from the referenced spec, and the answer describes assumptions ("filtered to X", "categories defined as Y") that the task never stated.
- **Detection procedure**:
  1. From the task statement, list every output artifact and every rule the task delegates to an external document; check the script actually opens/echoes that document (or that its rules are quoted) rather than inferring them.
  2. Grep the script for write/save calls (`savefig`, `np.save`, `to_json`, `to_csv`, `json.dump`) and compare the set of produced paths/filenames against the required artifact list — any required artifact with no corresponding write call is a failure.
  3. Read the answer for any filtering thresholds, category definitions, or orderings that do not appear in the task text; if present, verify each is traceable to the spec, not to a comment the agent wrote for itself.
  4. Sanity-check the reported counts/shapes against the raw data size and the number of requested categories (e.g., are all requested groups present and non-degenerate?).
- **Discriminator**: A real violation is a missing required file or a decision rule with no source in the task/spec. It is fine if the agent derives extra intermediate values or adds cosmetic choices (dpi, start angle, label format) that the spec leaves open, as long as every mandated artifact is written and every substantive rule is sourced.
- **Consequence**: The grader reports the unwritten artifacts as MISSING and, because the group definitions/filters differ from the spec, the produced artifact's values also mismatch — all checks fail even though the script runs without error.
759Output template file never read; schema and label strings inventedtaskda-code
Applies when
task -- The task points to a provided sample/example output file (or explicitly states a format) that the deliverable must follow.
Pattern
The scripts compute the analytical answer but never load or inspect the reference format file; the final writer hard-codes column names, row labels, row order, and even filler/empty columns guessed from the prose of the prompt rather than copied from the template.
Detection procedure
  1. Read the task and note every artifact that defines the expected output (sample file path, stated column names, category wording, ordering, rounding/units).
  2. Grep the scripts for any read of that template file (or any assertion that the produced frame's columns/rows match it); if absent, the format was guessed.
  3. Compare the written frame's header, label strings, and number/order of rows against the format the task describes — look for invented headers, paraphrased category names (e.g., shortened or reworded labels), or extra blank columns.
  4. Confirm no post-write validation step re-reads the output and checks shape/columns against the template.
Discriminator
Not a violation if the script loads the template (or copies its exact header/rows) and fills values in, or if the task gives no format reference; it is a violation when the header/labels appear nowhere in the task text or template and were synthesized by the agent, even if the underlying computed values are correct.
Consequence
The grader's file/field comparison fails on column names or label strings despite correct analysis, scoring 0 for the result file.
id 826ad666fef1 · mined from da-code dacode-dm-csv-015@s14
raw text (what the judge reads)
### Output template file never read; schema and label strings invented
- **Applies when**: `task` -- The task points to a provided sample/example output file (or explicitly states a format) that the deliverable must follow.
- **Pattern**: The scripts compute the analytical answer but never load or inspect the reference format file; the final writer hard-codes column names, row labels, row order, and even filler/empty columns guessed from the prose of the prompt rather than copied from the template.
- **Detection procedure**:
  1. Read the task and note every artifact that defines the expected output (sample file path, stated column names, category wording, ordering, rounding/units).
  2. Grep the scripts for any read of that template file (or any assertion that the produced frame's columns/rows match it); if absent, the format was guessed.
  3. Compare the written frame's header, label strings, and number/order of rows against the format the task describes — look for invented headers, paraphrased category names (e.g., shortened or reworded labels), or extra blank columns.
  4. Confirm no post-write validation step re-reads the output and checks shape/columns against the template.
- **Discriminator**: Not a violation if the script loads the template (or copies its exact header/rows) and fills values in, or if the task gives no format reference; it *is* a violation when the header/labels appear nowhere in the task text or template and were synthesized by the agent, even if the underlying computed values are correct.
- **Consequence**: The grader's file/field comparison fails on column names or label strings despite correct analysis, scoring 0 for the result file.
760No data-quality validation of raw columns before imputation/modelingtaskinfiagent-dabench
Applies when
task -- the script reads a raw file (especially one whose name/provenance signals possible corruption) and immediately computes column means, imputes, and fits/evaluates a model on those columns.
Pattern
The attempt trusts the loaded columns as-is: it never checks dtypes, parseability, duplicates, or whether values fall in physically/logically plausible ranges. Corrupted entries (strings coerced or dropped, sentinel codes like -999/0, wrong units, sign errors, extreme outliers, duplicate rows) silently distort the imputation means and the fitted coefficients, so the reported metric is far off even though the code "runs".
Detection procedure
  1. Read the task for cues that the input may be dirty (file/column naming, an explicit instruction to handle missing values, or mention of "errors") and note which columns feed the model.
  2. In the script, look for any pre-imputation validation step: dtypes/info(), describe() with range/plausibility commentary, coercion of non-numeric values (pd.to_numeric(..., errors='coerce')), duplicate/outlier checks, or filtering of impossible values.
  3. Check whether the imputation means and fitted coefficients are sanity-checked against domain expectations (e.g., target mean and residual scale consistent with the observed spread of the target).
  4. Compare the reported metric's magnitude to the target's own variance: if MSE is of the same order as (or larger than) the target's variance, the model is no better than the mean and the script should have investigated instead of reporting.
Discriminator
A real violation is a script that prints only null counts/shapes and proceeds, with no dtype/range inspection and no reaction to an implausibly large error. It is not a violation if the script inspects the columns, documents that values are clean and numeric, and the reported error is small relative to target variance — or if it explicitly justifies keeping suspicious values as legitimate.
Consequence
The metric is computed on contaminated features/target, producing a value an order of magnitude away from the expected one, so the numeric answer fails the grader even though the pipeline and answer format are correct.
id 7854489c0946 · mined from infiagent-dabench dabench-432@s14
raw text (what the judge reads)
### No data-quality validation of raw columns before imputation/modeling
- **Applies when**: `task` -- the script reads a raw file (especially one whose name/provenance signals possible corruption) and immediately computes column means, imputes, and fits/evaluates a model on those columns.
- **Pattern**: The attempt trusts the loaded columns as-is: it never checks dtypes, parseability, duplicates, or whether values fall in physically/logically plausible ranges. Corrupted entries (strings coerced or dropped, sentinel codes like -999/0, wrong units, sign errors, extreme outliers, duplicate rows) silently distort the imputation means and the fitted coefficients, so the reported metric is far off even though the code "runs".
- **Detection procedure**:
  1. Read the task for cues that the input may be dirty (file/column naming, an explicit instruction to handle missing values, or mention of "errors") and note which columns feed the model.
  2. In the script, look for any pre-imputation validation step: `dtypes`/`info()`, `describe()` with range/plausibility commentary, coercion of non-numeric values (`pd.to_numeric(..., errors='coerce')`), duplicate/outlier checks, or filtering of impossible values.
  3. Check whether the imputation means and fitted coefficients are sanity-checked against domain expectations (e.g., target mean and residual scale consistent with the observed spread of the target).
  4. Compare the reported metric's magnitude to the target's own variance: if MSE is of the same order as (or larger than) the target's variance, the model is no better than the mean and the script should have investigated instead of reporting.
- **Discriminator**: A real violation is a script that prints only null counts/shapes and proceeds, with no dtype/range inspection and no reaction to an implausibly large error. It is *not* a violation if the script inspects the columns, documents that values are clean and numeric, and the reported error is small relative to target variance — or if it explicitly justifies keeping suspicious values as legitimate.
- **Consequence**: The metric is computed on contaminated features/target, producing a value an order of magnitude away from the expected one, so the numeric answer fails the grader even though the pipeline and answer format are correct.
761Unverified assumption about row ordering before a sequential/lagged computationtaskinfiagent-dabench
Applies when
task -- the computation depends on row order (differences, lags, percent changes, cumulative sums, time-series splits) and the script reorders or reverses the data before computing.
Pattern
The script asserts in a comment (or infers from a glance) that the file is in a particular chronological direction and applies a blanket reversal/shift, without ever parsing the ordering key and programmatically checking whether it is ascending or descending; a wrong guess silently flips the sign/direction of every lagged value.
Detection procedure
  1. In the task, identify that the requested quantity is defined relative to a "previous"/"next" row, i.e. order-dependent.
  2. In the script, find where ordering is established: is there an explicit parse of the ordering column and a sort_values(...) (or an assertion/is_monotonic check), or only a hard-coded reversal/iloc[::-1] justified by a comment?
  3. Check whether any post-hoc validation confirms the direction (e.g., printing first/last ordering key, verifying monotonicity, or checking a known-sign sanity property of the result).
  4. Compare the reported answer against direction-sensitive sanity signals (e.g., sign of the mean vs. the overall first-to-last change in the underlying series); an unexplained sign flip is a red flag.
Discriminator
A real violation is order-assumed-by-comment with no parsing/sorting/assertion on the ordering key; it is fine if the script explicitly converts the key to a proper type and sorts ascending (or verifies monotonicity) before shifting — even if the data happened to already be in that order.
Consequence
Lagged values are computed backwards, so the mean-like statistic flips sign (and dispersion shifts slightly), and the graded value mismatches the expected one.
id 93a79351b6ad · mined from infiagent-dabench dabench-75@s14
raw text (what the judge reads)
### Unverified assumption about row ordering before a sequential/lagged computation
- **Applies when**: `task` -- the computation depends on row order (differences, lags, percent changes, cumulative sums, time-series splits) and the script reorders or reverses the data before computing.
- **Pattern**: The script asserts in a comment (or infers from a glance) that the file is in a particular chronological direction and applies a blanket reversal/shift, without ever parsing the ordering key and programmatically checking whether it is ascending or descending; a wrong guess silently flips the sign/direction of every lagged value.
- **Detection procedure**:
  1. In the task, identify that the requested quantity is defined relative to a "previous"/"next" row, i.e. order-dependent.
  2. In the script, find where ordering is established: is there an explicit parse of the ordering column and a `sort_values(...)` (or an assertion/`is_monotonic` check), or only a hard-coded reversal/`iloc[::-1]` justified by a comment?
  3. Check whether any post-hoc validation confirms the direction (e.g., printing first/last ordering key, verifying monotonicity, or checking a known-sign sanity property of the result).
  4. Compare the reported answer against direction-sensitive sanity signals (e.g., sign of the mean vs. the overall first-to-last change in the underlying series); an unexplained sign flip is a red flag.
- **Discriminator**: A real violation is order-assumed-by-comment with no parsing/sorting/assertion on the ordering key; it is fine if the script explicitly converts the key to a proper type and sorts ascending (or verifies monotonicity) before shifting — even if the data happened to already be in that order.
- **Consequence**: Lagged values are computed backwards, so the mean-like statistic flips sign (and dispersion shifts slightly), and the graded value mismatches the expected one.
762Ignoring the referenced format specification file when building the output rowtaskda-code
Applies when
task -- The task points to an auxiliary instructions/tips file (or an explicitly enumerated set of output fields) that defines the exact schema, column names, value wording, or rounding of a result file the script must write.
Pattern
The script never opens or echoes the referenced spec file and instead invents its own column names, extra/missing fields, free-text category labels, and rounding, so the written file only loosely resembles the required format (e.g., an extra statistic column, self-worded decision/comment strings, a p-value rounded until it degenerates to 0.0).
Detection procedure
  1. In the task text, list every named instruction/spec artifact and every field explicitly required in the output ("test type, decision, p-value, comment", one row, etc.).
  2. Scan the scripts for any read/print of that spec artifact and for a literal comparison of the constructed header/row against the required field list and value vocabulary.
  3. Compare the script's output DataFrame columns and cell values to the task-stated field list: flag extra columns, renamed fields, invented wording, or numeric formatting not sanctioned by the spec.
  4. Check the numeric fields survive the formatting choice (a rounded p-value or metric that collapses to 0/1 or loses all information is a red flag; scientific notation or the spec's precision is usually intended).
Discriminator
A real violation is when no evidence exists that the agent consulted the format source and the produced schema/vocabulary is self-authored; it is fine if the script demonstrably reads or quotes the spec (or the task itself fully enumerates the fields) and the output columns/labels/precision match that enumeration exactly, even if the statistical test choice differs.
Consequence
The result file is marked WRONG/MISSING on exact-match comparison because column set, label strings, or rounded numeric values differ from the expected single-row format, scoring 0 regardless of whether the underlying test was statistically appropriate.
id 31521de8d340 · mined from da-code dacode-data-sa-004@s14
raw text (what the judge reads)
### Ignoring the referenced format specification file when building the output row
- **Applies when**: `task` -- The task points to an auxiliary instructions/tips file (or an explicitly enumerated set of output fields) that defines the exact schema, column names, value wording, or rounding of a result file the script must write.
- **Pattern**: The script never opens or echoes the referenced spec file and instead invents its own column names, extra/missing fields, free-text category labels, and rounding, so the written file only loosely resembles the required format (e.g., an extra statistic column, self-worded decision/comment strings, a p-value rounded until it degenerates to `0.0`).
- **Detection procedure**:
  1. In the task text, list every named instruction/spec artifact and every field explicitly required in the output ("test type, decision, p-value, comment", one row, etc.).
  2. Scan the scripts for any read/print of that spec artifact and for a literal comparison of the constructed header/row against the required field list and value vocabulary.
  3. Compare the script's output DataFrame columns and cell values to the task-stated field list: flag extra columns, renamed fields, invented wording, or numeric formatting not sanctioned by the spec.
  4. Check the numeric fields survive the formatting choice (a rounded p-value or metric that collapses to 0/1 or loses all information is a red flag; scientific notation or the spec's precision is usually intended).
- **Discriminator**: A real violation is when no evidence exists that the agent consulted the format source and the produced schema/vocabulary is self-authored; it is fine if the script demonstrably reads or quotes the spec (or the task itself fully enumerates the fields) and the output columns/labels/precision match that enumeration exactly, even if the statistical test choice differs.
- **Consequence**: The result file is marked WRONG/MISSING on exact-match comparison because column set, label strings, or rounded numeric values differ from the expected single-row format, scoring 0 regardless of whether the underlying test was statistically appropriate.
763Unverified row-selection / missing-value handling for a pairwise statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single summary statistic (correlation, mean difference, test p-value) computed over two or more columns of a raw table, and the answer must match to a fixed rounding precision.
Pattern
The attempt loads the file with default settings and calls the statistic function directly, without inspecting how many rows actually enter the computation — non-numeric/sentinel values silently coerced or dropped, blank/NA rows removed per-column instead of pairwise, duplicate header/footer or aggregate rows left in, or a subset of the file read (e.g. one sheet/first chunk). The reported value is then close to but not equal to the value obtained from the correct row set, and no cross-check is done.
Detection procedure
  1. Read the task: note that the statistic is over a specific pair of columns of the full table and is graded at a fixed number of decimals, so small row-set differences change the graded digit.
  2. Read the script (or note that no script/log was saved — an immediate inadequacy, since the number cannot be reproduced or audited): check whether it prints the raw row count, the per-column non-null counts, the dtypes, and the final number of pairs used after alignment/dropping.
  3. Check that missing values are handled pairwise on the two columns of interest (rows dropped only if either of those columns is missing) and that columns are explicitly numeric, with any sentinel/text values inspected rather than silently coerced.
  4. Check the answer for a sanity/robustness cross-check: the statistic recomputed by a second independent route (e.g. another library or a manual formula) or with/without alternative NA handling, and the reported rounding done from the full-precision value.
Discriminator
A real violation is when the script never reveals the effective sample size or dtype handling, so an alternative defensible row set would shift the rounded value; it is fine if the script prints counts/dtypes showing the columns are fully numeric with no missing entries (so all row-selection choices coincide) and the reported figure is a correctly rounded version of the printed full-precision statistic.
Consequence
The graded statistic is off by one unit in the last required decimal (e.g. 0.53 vs 0.54), so the numeric check fails even though the qualitative conclusion (significant / not significant) is right, and the attempt is marked incorrect.
id 6b40d74ce415 · mined from infiagent-dabench dabench-300@s14
raw text (what the judge reads)
### Unverified row-selection / missing-value handling for a pairwise statistic
- **Applies when**: `task` -- the task asks for a single summary statistic (correlation, mean difference, test p-value) computed over two or more columns of a raw table, and the answer must match to a fixed rounding precision.
- **Pattern**: The attempt loads the file with default settings and calls the statistic function directly, without inspecting how many rows actually enter the computation — non-numeric/sentinel values silently coerced or dropped, blank/NA rows removed per-column instead of pairwise, duplicate header/footer or aggregate rows left in, or a subset of the file read (e.g. one sheet/first chunk). The reported value is then close to but not equal to the value obtained from the correct row set, and no cross-check is done.
- **Detection procedure**:
  1. Read the task: note that the statistic is over a specific pair of columns of the full table and is graded at a fixed number of decimals, so small row-set differences change the graded digit.
  2. Read the script (or note that no script/log was saved — an immediate inadequacy, since the number cannot be reproduced or audited): check whether it prints the raw row count, the per-column non-null counts, the dtypes, and the final number of pairs used after alignment/dropping.
  3. Check that missing values are handled pairwise on the two columns of interest (rows dropped only if either of those columns is missing) and that columns are explicitly numeric, with any sentinel/text values inspected rather than silently coerced.
  4. Check the answer for a sanity/robustness cross-check: the statistic recomputed by a second independent route (e.g. another library or a manual formula) or with/without alternative NA handling, and the reported rounding done from the full-precision value.
- **Discriminator**: A real violation is when the script never reveals the effective sample size or dtype handling, so an alternative defensible row set would shift the rounded value; it is fine if the script prints counts/dtypes showing the columns are fully numeric with no missing entries (so all row-selection choices coincide) and the reported figure is a correctly rounded version of the printed full-precision statistic.
- **Consequence**: The graded statistic is off by one unit in the last required decimal (e.g. 0.53 vs 0.54), so the numeric check fails even though the qualitative conclusion (significant / not significant) is right, and the attempt is marked incorrect.
764Output artifacts written outside the task's canonical data directorytaskinfiagent-dabench
Applies when
task -- the answer format requires reporting file paths for derived/output files, and the source data lives in a specific provided directory.
Pattern
The agent saves outputs to its own working/home directory (or a temp or relative path) instead of the directory where the supplied input data resides, then reports that path; every numeric value can be right while the path string fails an exact-match check.
Detection procedure
  1. In the task statement, note the directory of the input file(s) and any explicitly stated output location or naming convention.
  2. In the scripts, find every write call (to_csv, save, open(...,'w'), plot/model dumps) and record the exact destination path, including whether it is relative or absolute.
  3. Compare each destination with the input-data directory and with the path string in the final answer; flag if the output directory differs from the input directory (absent an explicit instruction to use another location), or if the reported path is relative/unresolved or does not literally equal where the file was written.
  4. Also check the same literalness for other format-sensitive fields (rounding to the stated number of decimals, integer vs. float) reported alongside the path.
Discriminator
A real violation is writing/reporting a path in a different directory than the provided data with no instruction to do so, or a mismatch between where the script writes and what the answer reports. It is fine if the task explicitly names an output directory and the agent used it, or if the agent wrote into the input directory and reported that absolute path.
Consequence
The grader's exact string comparison on the path field fails, marking the whole grouped answer wrong even though all computed statistics match ground truth.
id e3d78d18383f · mined from infiagent-dabench dabench-743@s14
raw text (what the judge reads)
### Output artifacts written outside the task's canonical data directory
- **Applies when**: `task` -- the answer format requires reporting file paths for derived/output files, and the source data lives in a specific provided directory.
- **Pattern**: The agent saves outputs to its own working/home directory (or a temp or relative path) instead of the directory where the supplied input data resides, then reports that path; every numeric value can be right while the path string fails an exact-match check.
- **Detection procedure**:
  1. In the task statement, note the directory of the input file(s) and any explicitly stated output location or naming convention.
  2. In the scripts, find every write call (`to_csv`, `save`, `open(...,'w')`, plot/model dumps) and record the exact destination path, including whether it is relative or absolute.
  3. Compare each destination with the input-data directory and with the path string in the final answer; flag if the output directory differs from the input directory (absent an explicit instruction to use another location), or if the reported path is relative/unresolved or does not literally equal where the file was written.
  4. Also check the same literalness for other format-sensitive fields (rounding to the stated number of decimals, integer vs. float) reported alongside the path.
- **Discriminator**: A real violation is writing/reporting a path in a different directory than the provided data with no instruction to do so, or a mismatch between where the script writes and what the answer reports. It is fine if the task explicitly names an output directory and the agent used it, or if the agent wrote into the input directory and reported that absolute path.
- **Consequence**: The grader's exact string comparison on the path field fails, marking the whole grouped answer wrong even though all computed statistics match ground truth.
765Incomplete set of required output artifacts (only the visibly named file is produced)taskda-code
Applies when
task -- the task points to an external spec/config file for plotting or reporting guidelines and/or the evaluation expects several output files (e.g., an image plus a serialized plot spec and the underlying numeric array), not just the one file named in the prompt sentence.
Pattern
The agent reads the spec only for cosmetic settings (colors, size, title), writes the single explicitly named output file, and reports the numbers in prose — never persisting the derived values or the plot description in the machine-readable files the spec/harness requires, and never re-listing the output directory to confirm what exists.
Detection procedure
  1. Read the task and the referenced spec/config file in full; enumerate every artifact it mandates (file names, formats, key names, ordering/label conventions), not just the one mentioned in the prompt text.
  2. Read the scripts and list every file actually written; diff this list against the enumerated requirements.
  3. Check that the numeric quantities behind the figure (the aggregated values/labels, in the required order and units) are saved in the required serialized formats, not only printed or embedded in the image.
  4. Check the answer for an explicit verification step (directory listing / reload of each saved file) confirming all artifacts exist and are non-trivial.
Discriminator
A real violation is a missing or unreadable mandated artifact, or one whose content/order/keys deviate from the spec; it is not a violation if all required files are written and merely reported briefly in prose, nor if extra files beyond the spec are also present.
Consequence
The grader marks the missing files as WRONG/MISSING and the attempt fails every file-level check even when the headline analytic value is right.
id 763b22ec4eb0 · mined from da-code dacode-plot-pie-008@s14
raw text (what the judge reads)
### Incomplete set of required output artifacts (only the visibly named file is produced)
- **Applies when**: `task` -- the task points to an external spec/config file for plotting or reporting guidelines and/or the evaluation expects several output files (e.g., an image plus a serialized plot spec and the underlying numeric array), not just the one file named in the prompt sentence.
- **Pattern**: The agent reads the spec only for cosmetic settings (colors, size, title), writes the single explicitly named output file, and reports the numbers in prose — never persisting the derived values or the plot description in the machine-readable files the spec/harness requires, and never re-listing the output directory to confirm what exists.
- **Detection procedure**:
  1. Read the task and the referenced spec/config file in full; enumerate every artifact it mandates (file names, formats, key names, ordering/label conventions), not just the one mentioned in the prompt text.
  2. Read the scripts and list every file actually written; diff this list against the enumerated requirements.
  3. Check that the numeric quantities behind the figure (the aggregated values/labels, in the required order and units) are saved in the required serialized formats, not only printed or embedded in the image.
  4. Check the answer for an explicit verification step (directory listing / reload of each saved file) confirming all artifacts exist and are non-trivial.
- **Discriminator**: A real violation is a missing or unreadable mandated artifact, or one whose content/order/keys deviate from the spec; it is *not* a violation if all required files are written and merely reported briefly in prose, nor if extra files beyond the spec are also present.
- **Consequence**: The grader marks the missing files as WRONG/MISSING and the attempt fails every file-level check even when the headline analytic value is right.
766Ignoring the provided output template's schema and category vocabularytaskda-code
Applies when
task -- The task supplies a pre-existing result file (or explicit format spec) whose rows/columns encode the required categories, ordering, or labels, and the scripts must fill it in.
Pattern
The attempt never loads/inspects the template before writing it, and instead invents its own category definitions (own bucket boundaries, own label names, an extra "Unknown"/catch-all bucket for missing values, own row order), then writes rows keyed by those invented labels — so even if the underlying counts are computable, they land under keys the grader does not recognize.
Detection procedure
  1. From the task statement, note that an output file/format is provided and identify which fields are fixed by it (row keys, label spellings, column names, ordering, units/rounding).
  2. Scan the scripts for a read of that template (or of any accompanying spec file) before the write; check whether the category set used is derived from it or hard-coded from the agent's own assumptions.
  3. Compare the labels/rows in the answer against the template's: count of rows, exact label strings, presence of extra buckets for unmatched/missing records, and column names.
  4. Check that records with missing/unmatched key values are handled in a way the template allows, rather than dumped into a new category.
Discriminator
A real violation is when the category set, labels, or row structure originate from the agent's own definitions rather than the provided artifact (extra rows, renamed labels, invented thresholds). It is not a violation if the agent derived the buckets from the template/spec and merely reformatted display text in the prose summary while the written file matches exactly.
Consequence
The output file fails an exact row/label match against the expected file — the grader reports the result file as WRONG/MISSING and scores 0, regardless of whether some individual counts happen to be right.
id 5e56bcd3728e · mined from da-code dacode-dm-csv-001@s14
raw text (what the judge reads)
### Ignoring the provided output template's schema and category vocabulary
- **Applies when**: `task` -- The task supplies a pre-existing result file (or explicit format spec) whose rows/columns encode the required categories, ordering, or labels, and the scripts must fill it in.
- **Pattern**: The attempt never loads/inspects the template before writing it, and instead invents its own category definitions (own bucket boundaries, own label names, an extra "Unknown"/catch-all bucket for missing values, own row order), then writes rows keyed by those invented labels — so even if the underlying counts are computable, they land under keys the grader does not recognize.
- **Detection procedure**:
  1. From the task statement, note that an output file/format is provided and identify which fields are fixed by it (row keys, label spellings, column names, ordering, units/rounding).
  2. Scan the scripts for a read of that template (or of any accompanying spec file) *before* the write; check whether the category set used is derived from it or hard-coded from the agent's own assumptions.
  3. Compare the labels/rows in the answer against the template's: count of rows, exact label strings, presence of extra buckets for unmatched/missing records, and column names.
  4. Check that records with missing/unmatched key values are handled in a way the template allows, rather than dumped into a new category.
- **Discriminator**: A real violation is when the category set, labels, or row structure originate from the agent's own definitions rather than the provided artifact (extra rows, renamed labels, invented thresholds). It is *not* a violation if the agent derived the buckets from the template/spec and merely reformatted display text in the prose summary while the written file matches exactly.
- **Consequence**: The output file fails an exact row/label match against the expected file — the grader reports the result file as WRONG/MISSING and scores 0, regardless of whether some individual counts happen to be right.
767Unverified entity-level aggregation and threshold-split definition before computing a subgroup statistictaskinfiagent-dabench
Applies when
task -- the task asks for a statistic (correlation, test, metric) computed per group after (a) collapsing raw records into one row per entity via aggregations (max, span/duration, total) and (b) splitting entities by a data-derived threshold such as a median.
Pattern
The attempt jumps straight to the statistic without pinning down the unit of analysis or the split rule: it may correlate row-level records instead of one row per entity, compute the duration/extent from row counts rather than actual first-to-last timestamps, aggregate damage inconsistently (sum vs. max vs. per-record), take the median over rows instead of over entities, or handle the "equal to the median" and missing/zero-damage cases arbitrarily. The numbers look plausible and both subgroups yield tiny p-values, so nothing appears wrong, but the r values are off.
Detection procedure
  1. From the task, list the unit of analysis (one row per entity) and every derived column needed (the max-of-X, the duration, the damage measure) plus the split rule.
  2. In the scripts, locate the groupby/aggregation step: confirm the frame used for the statistic has exactly one row per entity, that duration is derived from actual start/end values in the correct units (not record counts or a proxy), and that the ranking/threshold variable is aggregated once per entity.
  3. Check the threshold: is the median taken over the deduplicated entity table, is the tie/boundary case assigned explicitly, and are rows with missing or non-positive values in the split variable handled deliberately rather than silently dropped or kept?
  4. Check that a sanity print exists (row counts per subgroup, min/max of duration and category, count of ties at the threshold) and that the two subgroup sizes are roughly balanced as a median split implies; absence of such validation is itself the flag.
Discriminator
A real violation is when the analysis frame's granularity or the threshold definition cannot be confirmed from the scripts (or is demonstrably row-level/proxy-based). It is not a violation if the script explicitly deduplicates to one row per entity, derives each quantity by the stated definition, documents the tie/missing rule, and prints counts — even if a defensible alternative convention would shift the coefficient slightly.
Consequence
The relationship direction and significance verdict can still come out right, but the reported correlation coefficient differs from ground truth at the required rounding precision (e.g., 0.58 vs 0.56), so the answer fails the numeric checks despite passing the categorical ones.
id 697d9c3bffd0 · mined from infiagent-dabench dabench-431@s14
raw text (what the judge reads)
### Unverified entity-level aggregation and threshold-split definition before computing a subgroup statistic
- **Applies when**: `task` -- the task asks for a statistic (correlation, test, metric) computed per group after (a) collapsing raw records into one row per entity via aggregations (max, span/duration, total) and (b) splitting entities by a data-derived threshold such as a median.
- **Pattern**: The attempt jumps straight to the statistic without pinning down the unit of analysis or the split rule: it may correlate row-level records instead of one row per entity, compute the duration/extent from row counts rather than actual first-to-last timestamps, aggregate damage inconsistently (sum vs. max vs. per-record), take the median over rows instead of over entities, or handle the "equal to the median" and missing/zero-damage cases arbitrarily. The numbers look plausible and both subgroups yield tiny p-values, so nothing appears wrong, but the r values are off.
- **Detection procedure**:
  1. From the task, list the unit of analysis (one row per entity) and every derived column needed (the max-of-X, the duration, the damage measure) plus the split rule.
  2. In the scripts, locate the groupby/aggregation step: confirm the frame used for the statistic has exactly one row per entity, that duration is derived from actual start/end values in the correct units (not record counts or a proxy), and that the ranking/threshold variable is aggregated once per entity.
  3. Check the threshold: is the median taken over the deduplicated entity table, is the tie/boundary case assigned explicitly, and are rows with missing or non-positive values in the split variable handled deliberately rather than silently dropped or kept?
  4. Check that a sanity print exists (row counts per subgroup, min/max of duration and category, count of ties at the threshold) and that the two subgroup sizes are roughly balanced as a median split implies; absence of such validation is itself the flag.
- **Discriminator**: A real violation is when the analysis frame's granularity or the threshold definition cannot be confirmed from the scripts (or is demonstrably row-level/proxy-based). It is *not* a violation if the script explicitly deduplicates to one row per entity, derives each quantity by the stated definition, documents the tie/missing rule, and prints counts — even if a defensible alternative convention would shift the coefficient slightly.
- **Consequence**: The relationship direction and significance verdict can still come out right, but the reported correlation coefficient differs from ground truth at the required rounding precision (e.g., 0.58 vs 0.56), so the answer fails the numeric checks despite passing the categorical ones.
768Arbitrarily narrowed baseline feature settaskinfiagent-dabench
Applies when
task -- the task asks whether adding an engineered feature improves a predictive model, but does not explicitly enumerate which predictors the baseline model should use.
Pattern
The script silently builds the baseline design matrix from only the handful of columns that happen to be mentioned in the prompt (e.g. those used to build the new feature or the correlation), discarding the other available predictors, and then compares against the same narrow set plus the new feature. Both reported errors are inflated relative to the natural "all available predictors" baseline, even though the relative comparison may still look sensible.
Detection procedure
  1. In the task text, check whether the predictor list for the baseline model is explicitly specified or left implicit ("predict Y", "add feature Z and compare").
  2. In the script, find where X is constructed and list the columns kept versus all non-target columns in the loaded data.
  3. If columns are dropped without any stated reason in the task (not because of dtype, identifier, leakage, or explicit instruction), flag it; also check that any non-numeric column was encoded rather than silently dropped.
  4. Sanity-check the reported errors: if the model's error is barely better than the target's standard deviation, that suggests informative columns were left out.
Discriminator
A real violation is dropping usable predictors purely by prompt-mention heuristics; it is fine to drop columns that the task names explicitly, that are IDs/duplicates of the target, that would leak, or that are non-numeric and handled by an explicit documented encoding decision.
Consequence
Both RMSE (or accuracy) values differ materially from the reference values computed on the full feature set, so the numeric checks fail even though the correlation/statistic parts pass.
id 33ff2e66f31d · mined from infiagent-dabench dabench-549@s14
raw text (what the judge reads)
### Arbitrarily narrowed baseline feature set
- **Applies when**: `task` -- the task asks whether adding an engineered feature improves a predictive model, but does not explicitly enumerate which predictors the baseline model should use.
- **Pattern**: The script silently builds the baseline design matrix from only the handful of columns that happen to be mentioned in the prompt (e.g. those used to build the new feature or the correlation), discarding the other available predictors, and then compares against the same narrow set plus the new feature. Both reported errors are inflated relative to the natural "all available predictors" baseline, even though the relative comparison may still look sensible.
- **Detection procedure**:
  1. In the task text, check whether the predictor list for the baseline model is explicitly specified or left implicit ("predict Y", "add feature Z and compare").
  2. In the script, find where `X` is constructed and list the columns kept versus all non-target columns in the loaded data.
  3. If columns are dropped without any stated reason in the task (not because of dtype, identifier, leakage, or explicit instruction), flag it; also check that any non-numeric column was encoded rather than silently dropped.
  4. Sanity-check the reported errors: if the model's error is barely better than the target's standard deviation, that suggests informative columns were left out.
- **Discriminator**: A real violation is dropping usable predictors purely by prompt-mention heuristics; it is fine to drop columns that the task names explicitly, that are IDs/duplicates of the target, that would leak, or that are non-numeric *and* handled by an explicit documented encoding decision.
- **Consequence**: Both RMSE (or accuracy) values differ materially from the reference values computed on the full feature set, so the numeric checks fail even though the correlation/statistic parts pass.
769No held-out validation and no calibration/sanity check of the predicted label distributiontaskda-code
Applies when
task -- the task asks for predictions on an unlabeled test set and the script trains a single classifier and writes hard labels straight to the output file.
Pattern
The attempt fits one model with arbitrary default hyperparameters, never evaluates it on a held-out or cross-validated split (no accuracy/F1/AUC number at all), never addresses class imbalance or the decision threshold, and then reports the run as successful based only on file-shape facts (row count, no NaNs). A tell-tale symptom is that the predicted positive rate is far below the training base rate (argmax at 0.5 on an imbalanced target), which is never questioned.
Detection procedure
  1. Read the task to see whether the graded artifact is a set of predictions whose quality (not just format) determines the score.
  2. Scan the script for any hold-out/CV evaluation of the trained model on labeled data; if the only fit/predict pair uses all training rows and then the unlabeled test rows, there is no evidence the model is better than trivial.
  3. Compare the reported/derivable positive-class rate in the output with the reported class prior of the training target; flag a large unexplained gap, and check whether imbalance handling (class weights, threshold tuning, resampling) or model comparison was even attempted.
  4. Check the answer text: if it asserts "good separation"/"strong performance" with no measured metric, the claim is unsupported.
Discriminator
A real violation is zero measured generalization performance plus no imbalance/threshold reasoning; it is not a violation if the script reports a CV or hold-out metric (even for a single default model) and the predicted rate deviation is explained by that evaluation, nor if the task explicitly grades only file format.
Consequence
The submitted prediction file passes format checks but falls below the grader's accuracy/F1 threshold, so the expected output file is scored WRONG and the task fails 0/1.
id 6d3f700d0016 · mined from da-code dacode-ml-binary-016@s14
raw text (what the judge reads)
### No held-out validation and no calibration/sanity check of the predicted label distribution
- **Applies when**: `task` -- the task asks for predictions on an unlabeled test set and the script trains a single classifier and writes hard labels straight to the output file.
- **Pattern**: The attempt fits one model with arbitrary default hyperparameters, never evaluates it on a held-out or cross-validated split (no accuracy/F1/AUC number at all), never addresses class imbalance or the decision threshold, and then reports the run as successful based only on file-shape facts (row count, no NaNs). A tell-tale symptom is that the predicted positive rate is far below the training base rate (argmax at 0.5 on an imbalanced target), which is never questioned.
- **Detection procedure**:
  1. Read the task to see whether the graded artifact is a set of predictions whose quality (not just format) determines the score.
  2. Scan the script for any hold-out/CV evaluation of the trained model on labeled data; if the only `fit`/`predict` pair uses all training rows and then the unlabeled test rows, there is no evidence the model is better than trivial.
  3. Compare the reported/derivable positive-class rate in the output with the reported class prior of the training target; flag a large unexplained gap, and check whether imbalance handling (class weights, threshold tuning, resampling) or model comparison was even attempted.
  4. Check the answer text: if it asserts "good separation"/"strong performance" with no measured metric, the claim is unsupported.
- **Discriminator**: A real violation is zero measured generalization performance plus no imbalance/threshold reasoning; it is *not* a violation if the script reports a CV or hold-out metric (even for a single default model) and the predicted rate deviation is explained by that evaluation, nor if the task explicitly grades only file format.
- **Consequence**: The submitted prediction file passes format checks but falls below the grader's accuracy/F1 threshold, so the expected output file is scored WRONG and the task fails 0/1.
770Unverified prediction artifact (path/schema/alignment/quality never checked against the test input)taskda-code
Applies when
task -- the task's deliverable is a prediction (or results) file written to a specified name/location, and the agent reports success mainly by quoting summary statistics of its own model.
Pattern
The attempt trains a model and writes an output file, but never re-reads the written file to confirm it exists at the required path, has exactly the required column name(s), has one row per test row in the original test-file order, contains no NaNs, and produces values whose distribution is plausible against the training target; reproducibility is also unverified (no saved script showing how test rows were loaded, features aligned, and rows ordered). Poor/regressed-to-the-mean predictions (heavily compressed range vs. the training target, weak validation score) are reported as "success" without any threshold check.
Detection procedure
  1. Read the task for the exact deliverable: filename/path, column name(s), implied row count and ordering, any rounding/units/index constraints.
  2. Read the scripts for a post-write verification step: does the code re-load the output file and assert path, shape, column names, row alignment to the test file (same order, same key/index), and absence of missing values? Check that the feature set and preprocessing applied to test data are identical to those fit on train (same encoders/scalers, no refitting on test, no dropped/reordered columns).
  3. Compare the answer's reported row count and prediction statistics with the test file's row count and the training target's distribution (min/max/mean/spread); flag if the test row count is unconfirmed or the prediction spread is drastically narrower than the target's, or if the answer only reports validation metrics as evidence of correctness.
  4. If no script is retained or the pipeline can't be re-run to regenerate the file, treat the deliverable as unverified.
Discriminator
A fine attempt shows an explicit read-back/assertion of the saved file against the test input (shape, exact column name, order/keys, no NaNs) and reports at least one held-out accuracy figure with a sanity comparison to a trivial baseline; a violation asserts success purely from in-memory model metrics and self-reported statistics, or leaves the generating code unsaved so nothing can be checked.
Consequence
The grader reads the expected file and finds it missing, at the wrong path, with the wrong column/row count/order, or with predictions too far from the true values, so the single file check fails and the whole task scores 0 despite a plausible-sounding report.
id 6220e6de8b6a · mined from da-code dacode-ml-regression-015@s14
raw text (what the judge reads)
### Unverified prediction artifact (path/schema/alignment/quality never checked against the test input)
- **Applies when**: `task` -- the task's deliverable is a prediction (or results) file written to a specified name/location, and the agent reports success mainly by quoting summary statistics of its own model.
- **Pattern**: The attempt trains a model and writes an output file, but never re-reads the written file to confirm it exists at the required path, has exactly the required column name(s), has one row per test row in the original test-file order, contains no NaNs, and produces values whose distribution is plausible against the training target; reproducibility is also unverified (no saved script showing how test rows were loaded, features aligned, and rows ordered). Poor/regressed-to-the-mean predictions (heavily compressed range vs. the training target, weak validation score) are reported as "success" without any threshold check.
- **Detection procedure**:
  1. Read the task for the exact deliverable: filename/path, column name(s), implied row count and ordering, any rounding/units/index constraints.
  2. Read the scripts for a post-write verification step: does the code re-load the output file and assert path, shape, column names, row alignment to the test file (same order, same key/index), and absence of missing values? Check that the feature set and preprocessing applied to test data are identical to those fit on train (same encoders/scalers, no refitting on test, no dropped/reordered columns).
  3. Compare the answer's reported row count and prediction statistics with the test file's row count and the training target's distribution (min/max/mean/spread); flag if the test row count is unconfirmed or the prediction spread is drastically narrower than the target's, or if the answer only reports validation metrics as evidence of correctness.
  4. If no script is retained or the pipeline can't be re-run to regenerate the file, treat the deliverable as unverified.
- **Discriminator**: A fine attempt shows an explicit read-back/assertion of the saved file against the test input (shape, exact column name, order/keys, no NaNs) and reports at least one held-out accuracy figure with a sanity comparison to a trivial baseline; a violation asserts success purely from in-memory model metrics and self-reported statistics, or leaves the generating code unsaved so nothing can be checked.
- **Consequence**: The grader reads the expected file and finds it missing, at the wrong path, with the wrong column/row count/order, or with predictions too far from the true values, so the single file check fails and the whole task scores 0 despite a plausible-sounding report.
771Aggregate statistic computed on an incomplete or silently filtered slice of the datataskinfiagent-dabench
Applies when
task -- the task asks for a single summary statistic over "all observations" of a field, and the scripts load the data (possibly multi-file/multi-sheet/chunked) and apply drops, filters, or type coercions before aggregating.
Pattern
The attempt reports a number without ever establishing the denominator: it reads only part of the source (one file/sheet/chunk, a nrows/sample limit, a default header or delimiter that silently drops rows), or coerces the column with errors='coerce'/dropna/outlier trimming that removes valid values, and never checks that the count used matches the full record count. The reported figure is then a mean of a subpopulation rather than of the requested population. Absence of saved, rerunnable scripts makes the scope unverifiable.
Detection procedure
  1. Read the task and note exactly which population the statistic is over and which exclusions (missing values, outliers) are explicitly authorized — nothing else may shrink the data.
  2. In the scripts, trace from source to aggregation: how many files/sheets/chunks exist vs. how many are read; any nrows, head, sampling, subsetting, join, or merge that could drop rows; any coercion/filter applied to the target column.
  3. Check the script prints the raw row count, the non-null count actually averaged, and the value range/dtype of the column, and that these are consistent with the source's documented size; if no script or no counts are reported, treat scope as unestablished.
  4. Sanity-check the reported value against the column's plausible range and against a mean computed with no optional trimming — an unexplained shift between the two signals over-filtering.
Discriminator
A real violation is when rows are lost for reasons the task never authorized, or when the count behind the number is never shown; it is not a violation when the script reads the complete source, prints total vs. used counts, and drops only the exclusions the task explicitly permits with a stated rule.
Consequence
The submitted mean deviates from the ground-truth value computed over the full data, and the grader marks the single numeric check as wrong with no way to trace which subset was actually used.
id 0ee55e3a3989 · mined from infiagent-dabench dabench-320@s14
raw text (what the judge reads)
### Aggregate statistic computed on an incomplete or silently filtered slice of the data
- **Applies when**: `task` -- the task asks for a single summary statistic over "all observations" of a field, and the scripts load the data (possibly multi-file/multi-sheet/chunked) and apply drops, filters, or type coercions before aggregating.
- **Pattern**: The attempt reports a number without ever establishing the denominator: it reads only part of the source (one file/sheet/chunk, a `nrows`/sample limit, a default header or delimiter that silently drops rows), or coerces the column with `errors='coerce'`/dropna/outlier trimming that removes valid values, and never checks that the count used matches the full record count. The reported figure is then a mean of a subpopulation rather than of the requested population. Absence of saved, rerunnable scripts makes the scope unverifiable.
- **Detection procedure**:
  1. Read the task and note exactly which population the statistic is over and which exclusions (missing values, outliers) are explicitly authorized — nothing else may shrink the data.
  2. In the scripts, trace from source to aggregation: how many files/sheets/chunks exist vs. how many are read; any `nrows`, `head`, sampling, subsetting, join, or `merge` that could drop rows; any coercion/filter applied to the target column.
  3. Check the script prints the raw row count, the non-null count actually averaged, and the value range/dtype of the column, and that these are consistent with the source's documented size; if no script or no counts are reported, treat scope as unestablished.
  4. Sanity-check the reported value against the column's plausible range and against a mean computed with no optional trimming — an unexplained shift between the two signals over-filtering.
- **Discriminator**: A real violation is when rows are lost for reasons the task never authorized, or when the count behind the number is never shown; it is *not* a violation when the script reads the complete source, prints total vs. used counts, and drops only the exclusions the task explicitly permits with a stated rule.
- **Consequence**: The submitted mean deviates from the ground-truth value computed over the full data, and the grader marks the single numeric check as wrong with no way to trace which subset was actually used.
772Extremes reported from a column that was never verified as clean numerictaskda-code
Applies when
task -- the answer is the argmin/argmax (or top-k) of a quantity read directly from a raw tabular file whose columns may carry thousands separators, units, percent signs, or other text formatting.
Pattern
The attempt loads the file with defaults, applies the requested imputation, and takes idxmax/idxmin without ever checking the column's dtype or the parsed value range; if the column arrived as strings (or was coerced to NaN and then filled with the mean), the ordering is lexicographic or the extreme rows are imputed placeholders, so the reported labels are wrong even though the code "runs".
Detection procedure
  1. From the task, identify the single column whose extreme values determine the answer.
  2. In the scripts, look for an explicit dtype check / numeric conversion of that column (e.g. stripping separators, symbols, to_numeric with inspection of how many entries failed) performed before imputation and before sorting; absence of any such step is the flag.
  3. Check that the script prints the sorted head/tail with the actual numeric values (not just the labels) and that the mean-imputation did not itself create the extreme row.
  4. Compare the reported extreme values/labels to common-sense magnitude for the quantity; an implausible or clearly non-extreme entity in the min/max slot confirms a parsing or imputation artifact.
Discriminator
A real violation is when no dtype/parse validation or value printout exists (or the printout shows string dtype / suspiciously clustered values); it is fine if the script demonstrably converts to numeric, reports how many values were unparsable, and shows the ranked values that justify the chosen labels — even if it uses defaults, provided the verification output is present.
Consequence
The submitted min/max labels differ from the ground-truth entities in result.json, so the exact-match check fails (0/1), and no intermediate evidence exists to diagnose why.
id 8545017af98d · mined from da-code dacode-di-text-001@s14
raw text (what the judge reads)
### Extremes reported from a column that was never verified as clean numeric
- **Applies when**: `task` -- the answer is the argmin/argmax (or top-k) of a quantity read directly from a raw tabular file whose columns may carry thousands separators, units, percent signs, or other text formatting.
- **Pattern**: The attempt loads the file with defaults, applies the requested imputation, and takes `idxmax`/`idxmin` without ever checking the column's dtype or the parsed value range; if the column arrived as strings (or was coerced to NaN and then filled with the mean), the ordering is lexicographic or the extreme rows are imputed placeholders, so the reported labels are wrong even though the code "runs".
- **Detection procedure**:
  1. From the task, identify the single column whose extreme values determine the answer.
  2. In the scripts, look for an explicit dtype check / numeric conversion of that column (e.g. stripping separators, symbols, `to_numeric` with inspection of how many entries failed) performed *before* imputation and before sorting; absence of any such step is the flag.
  3. Check that the script prints the sorted head/tail with the actual numeric values (not just the labels) and that the mean-imputation did not itself create the extreme row.
  4. Compare the reported extreme values/labels to common-sense magnitude for the quantity; an implausible or clearly non-extreme entity in the min/max slot confirms a parsing or imputation artifact.
- **Discriminator**: A real violation is when no dtype/parse validation or value printout exists (or the printout shows string dtype / suspiciously clustered values); it is fine if the script demonstrably converts to numeric, reports how many values were unparsable, and shows the ranked values that justify the chosen labels — even if it uses defaults, provided the verification output is present.
- **Consequence**: The submitted min/max labels differ from the ground-truth entities in `result.json`, so the exact-match check fails (0/1), and no intermediate evidence exists to diagnose why.
773Ratio (or other derived quantity) computed with operands reversed and never sanity-checkedtaskda-code
Applies when
task -- the task asks for a derived quantity built from two or more columns (a ratio, difference, rate, per-unit normalization) whose order/direction is explicitly stated in the prompt.
Pattern
The script pulls the two columns from the data file by position, name-guess, or a hard-coded column index and divides/subtracts in whichever order they appear in the table, producing the inverse (or negated) quantity; the agent then reports summary statistics and confidence intervals of that inverted quantity without ever checking that its magnitude is plausible for the stated definition, and without confirming the output file's columns/ordering match the provided sample template.
Detection procedure
  1. Read the task and write down the exact stated definition (numerator vs denominator, minuend vs subtrahend) and any required output template.
  2. In the script, locate the line that builds the derived quantity and verify the variable actually loaded into the numerator position corresponds to the field named first in the task — not just the first column of the file.
  3. Check the reported value against a domain/range sanity check implied by the definition (e.g., if the task says A-to-B and A is known to exceed B, the ratio must be >1; a value that is almost exactly the reciprocal of the plausible value is a red flag).
  4. Compare the written output file's header names, column count, row order, and numeric formatting against the supplied sample file; flag if CI bounds are stuffed into a bracketed string when the sample splits them, or vice versa.
Discriminator
A real violation is when the reported value is (approximately) the reciprocal/negation of what the stated definition implies, or the output columns deviate from the sample; it is not a violation if the ordering matches the task wording and the value is merely surprising, nor if column names in the raw data differ cosmetically but the correct field is used.
Consequence
Every downstream number — the means and both bootstrap interval endpoints — is wrong by inversion, and the output file fails an exact/tolerance comparison against the expected result file, scoring 0 despite a methodologically sound bootstrap.
id 95fb9ed1e5c2 · mined from da-code dacode-data-sa-029@s14
raw text (what the judge reads)
### Ratio (or other derived quantity) computed with operands reversed and never sanity-checked
- **Applies when**: `task` -- the task asks for a derived quantity built from two or more columns (a ratio, difference, rate, per-unit normalization) whose order/direction is explicitly stated in the prompt.
- **Pattern**: The script pulls the two columns from the data file by position, name-guess, or a hard-coded column index and divides/subtracts in whichever order they appear in the table, producing the inverse (or negated) quantity; the agent then reports summary statistics and confidence intervals of that inverted quantity without ever checking that its magnitude is plausible for the stated definition, and without confirming the output file's columns/ordering match the provided sample template.
- **Detection procedure**:
  1. Read the task and write down the exact stated definition (numerator vs denominator, minuend vs subtrahend) and any required output template.
  2. In the script, locate the line that builds the derived quantity and verify the variable actually loaded into the numerator position corresponds to the field named first in the task — not just the first column of the file.
  3. Check the reported value against a domain/range sanity check implied by the definition (e.g., if the task says A-to-B and A is known to exceed B, the ratio must be >1; a value that is almost exactly the reciprocal of the plausible value is a red flag).
  4. Compare the written output file's header names, column count, row order, and numeric formatting against the supplied sample file; flag if CI bounds are stuffed into a bracketed string when the sample splits them, or vice versa.
- **Discriminator**: A real violation is when the reported value is (approximately) the reciprocal/negation of what the stated definition implies, or the output columns deviate from the sample; it is *not* a violation if the ordering matches the task wording and the value is merely surprising, nor if column names in the raw data differ cosmetically but the correct field is used.
- **Consequence**: Every downstream number — the means and both bootstrap interval endpoints — is wrong by inversion, and the output file fails an exact/tolerance comparison against the expected result file, scoring 0 despite a methodologically sound bootstrap.
774Ambiguous estimator convention chosen arbitrarily (e.g., population vs sample dispersion)taskinfiagent-dabench
Applies when
task -- the task asks for a summary statistic whose formula has more than one standard convention (denominator n vs n-1, ddof, normalization, bias correction) and the prompt does not state which.
Pattern
The script computes both variants, notes the ambiguity in a comment, then silently picks one (often the non-default of the tooling the task's ecosystem implies) and reports only that number, without justifying the choice against the task's likely reference implementation or reporting the alternative.
Detection procedure
1) Read the task wording for any specification of the estimator convention; if none, mark the statistic as ambiguous. 2) In the scripts, look for parameters that select the convention (e.g., ddof, bias, sample=) and see whether a value was hard-coded on the agent's own guess. 3) Check whether the chosen convention matches the default of the library/idiom the task points to (e.g., a dataframe's built-in .std() default) and whether the final answer states or hedges the choice. 4) Confirm the final reported value is the one from the convention the task's tooling would produce.
Discriminator
Fine if the task explicitly names the convention, or if both conventions round to the same reported value, or if the agent aligned with the default of the tool the task implies. A violation is picking the non-default convention (or any convention) purely on the agent's own reading while a differing alternative was computed and discarded.
Consequence
The dependent statistic is off by a factor of √(n/(n−1)) — the answer fails that check while other, convention-free quantities (like the mean) pass, yielding a partial-credit "incorrect" verdict.
id cb7ea9b28013 · mined from infiagent-dabench dabench-255@s14
raw text (what the judge reads)
### Ambiguous estimator convention chosen arbitrarily (e.g., population vs sample dispersion)
- **Applies when**: `task` -- the task asks for a summary statistic whose formula has more than one standard convention (denominator n vs n-1, ddof, normalization, bias correction) and the prompt does not state which.
- **Pattern**: The script computes both variants, notes the ambiguity in a comment, then silently picks one (often the non-default of the tooling the task's ecosystem implies) and reports only that number, without justifying the choice against the task's likely reference implementation or reporting the alternative.
- **Detection procedure**: 1) Read the task wording for any specification of the estimator convention; if none, mark the statistic as ambiguous. 2) In the scripts, look for parameters that select the convention (e.g., `ddof`, `bias`, `sample=`) and see whether a value was hard-coded on the agent's own guess. 3) Check whether the chosen convention matches the default of the library/idiom the task points to (e.g., a dataframe's built-in `.std()` default) and whether the final answer states or hedges the choice. 4) Confirm the final reported value is the one from the convention the task's tooling would produce.
- **Discriminator**: Fine if the task explicitly names the convention, or if both conventions round to the same reported value, or if the agent aligned with the default of the tool the task implies. A violation is picking the non-default convention (or any convention) purely on the agent's own reading while a differing alternative was computed and discarded.
- **Consequence**: The dependent statistic is off by a factor of √(n/(n−1)) — the answer fails that check while other, convention-free quantities (like the mean) pass, yielding a partial-credit "incorrect" verdict.
775Fit reported only on training data, with no distribution sanity check on the predictionstaskda-code
Applies when
task -- a script trains a model on labeled data and must write predictions for an unlabeled/held-out file, and the report cites only in-sample scores.
Pattern
The attempt evaluates the model on the same rows it was fit on (near-perfect R²/RMSE presented as "excellent performance"), never uses a hold-out split or cross-validation, and never checks that the emitted prediction column is plausible — its mean/median/quantiles/min/max and dtype are not compared against the training target's, so a heavily shrunk, skewed, negative, or fractional-count prediction vector passes unnoticed.
Detection procedure
  1. Read the task: identify the target, its natural scale/type (counts, non-negative, integers, units), and the required output file/column.
  2. Read the scripts: check whether any score is computed on rows excluded from fitting (train/test split, CV, or a time-based split); flag if every reported metric uses the fitted rows.
  3. Read the scripts/answer for a comparison of the predicted distribution to the training target distribution (mean, median, quantiles, min/max, share of zeros, dtype/rounding); flag if absent or if the reported prediction summary differs from the training target by a large factor or violates the target's domain (e.g., non-integer or below the natural minimum).
  4. Confirm row count and column name of the written file match the test input and the requested format.
Discriminator
A real violation is in-sample-only validation plus no distributional/range check (or a check whose numbers are clearly off-scale). It is fine if the agent held out data or ran CV and reported those out-of-sample scores, or if it explicitly justified an in-sample-only fit while showing prediction summaries that align with the training target's scale and type.
Consequence
The unvalidated model is badly overfit/mis-scaled, so the written prediction column is systematically off (e.g., an order of magnitude too small and fractional where counts are expected), and the file-level comparison against the expected predictions fails despite a confident "excellent performance" claim.
id 98b108c6a117 · mined from da-code dacode-ml-regression-008@s15
raw text (what the judge reads)
### Fit reported only on training data, with no distribution sanity check on the predictions
- **Applies when**: `task` -- a script trains a model on labeled data and must write predictions for an unlabeled/held-out file, and the report cites only in-sample scores.
- **Pattern**: The attempt evaluates the model on the same rows it was fit on (near-perfect R²/RMSE presented as "excellent performance"), never uses a hold-out split or cross-validation, and never checks that the emitted prediction column is plausible — its mean/median/quantiles/min/max and dtype are not compared against the training target's, so a heavily shrunk, skewed, negative, or fractional-count prediction vector passes unnoticed.
- **Detection procedure**:
  1. Read the task: identify the target, its natural scale/type (counts, non-negative, integers, units), and the required output file/column.
  2. Read the scripts: check whether any score is computed on rows excluded from fitting (train/test split, CV, or a time-based split); flag if every reported metric uses the fitted rows.
  3. Read the scripts/answer for a comparison of the predicted distribution to the training target distribution (mean, median, quantiles, min/max, share of zeros, dtype/rounding); flag if absent or if the reported prediction summary differs from the training target by a large factor or violates the target's domain (e.g., non-integer or below the natural minimum).
  4. Confirm row count and column name of the written file match the test input and the requested format.
- **Discriminator**: A real violation is in-sample-only validation *plus* no distributional/range check (or a check whose numbers are clearly off-scale). It is fine if the agent held out data or ran CV and reported those out-of-sample scores, or if it explicitly justified an in-sample-only fit while showing prediction summaries that align with the training target's scale and type.
- **Consequence**: The unvalidated model is badly overfit/mis-scaled, so the written prediction column is systematically off (e.g., an order of magnitude too small and fractional where counts are expected), and the file-level comparison against the expected predictions fails despite a confident "excellent performance" claim.
776Unquestioned use of the default parametric test on the full, unfiltered populationtaskda-code
Applies when
task -- the task asks for a p-value and a reject/fail-to-reject decision from a comparison of two groups drawn from large historical/observational tables that contain many heterogeneous subsets (eras, competition types, categories).
Pattern
The script loads every row, computes the outcome variable, and immediately calls a single default test (e.g., a standard two-sample t-test) on the entire dataset — with no explicit scoping of the population being compared and no check that the test's assumptions (normality/symmetry of the outcome, equal variances, independence) hold for a discrete, right-skewed count variable. The resulting p-value is astronomically small and is reported without any sensitivity check.
Detection procedure
  1. Read the task and README for any implicit or explicit scoping of the comparison (time window, competition/category, subgroup) and for hints about the intended test family; note that a domain-meaningful comparison rarely means "all rows ever recorded".
  2. In the script, check whether any filtering/subsetting step exists before the test, and whether the choice of test is justified (distribution plot, normality/variance check, or an explicit alternative such as a rank-based/non-parametric test).
  3. Check whether the script prints group sizes and distribution summaries and compares at least one alternative test or subset to see whether the reject/fail-to-reject decision is stable.
  4. Inspect the answer: a p-value at the far tail (e.g., <1e-50) driven by tens of thousands of rows, with no reported robustness check, is a red flag that the population and test were chosen by default rather than by design.
Discriminator
A real violation is when no scoping rationale and no assumption/robustness check appear anywhere — the test is simply the first one imported. It is not a violation if the script explicitly argues (in code or output) that the full population is the intended one and shows that the parametric and non-parametric tests, or filtered and unfiltered samples, give the same decision.
Consequence
The reported p-value differs by many orders of magnitude from the reference value computed on the intended subset with the intended test, so the numeric column fails the grader's tolerance check even when the reject/fail-to-reject label happens to match.
id 9e8d45844187 · mined from da-code dacode-data-sa-001@s15
raw text (what the judge reads)
### Unquestioned use of the default parametric test on the full, unfiltered population
- **Applies when**: `task` -- the task asks for a p-value and a reject/fail-to-reject decision from a comparison of two groups drawn from large historical/observational tables that contain many heterogeneous subsets (eras, competition types, categories).
- **Pattern**: The script loads every row, computes the outcome variable, and immediately calls a single default test (e.g., a standard two-sample t-test) on the entire dataset — with no explicit scoping of the population being compared and no check that the test's assumptions (normality/symmetry of the outcome, equal variances, independence) hold for a discrete, right-skewed count variable. The resulting p-value is astronomically small and is reported without any sensitivity check.
- **Detection procedure**:
  1. Read the task and README for any implicit or explicit scoping of the comparison (time window, competition/category, subgroup) and for hints about the intended test family; note that a domain-meaningful comparison rarely means "all rows ever recorded".
  2. In the script, check whether any filtering/subsetting step exists before the test, and whether the choice of test is justified (distribution plot, normality/variance check, or an explicit alternative such as a rank-based/non-parametric test).
  3. Check whether the script prints group sizes and distribution summaries and compares at least one alternative test or subset to see whether the reject/fail-to-reject decision is stable.
  4. Inspect the answer: a p-value at the far tail (e.g., <1e-50) driven by tens of thousands of rows, with no reported robustness check, is a red flag that the population and test were chosen by default rather than by design.
- **Discriminator**: A real violation is when no scoping rationale and no assumption/robustness check appear anywhere — the test is simply the first one imported. It is *not* a violation if the script explicitly argues (in code or output) that the full population is the intended one and shows that the parametric and non-parametric tests, or filtered and unfiltered samples, give the same decision.
- **Consequence**: The reported p-value differs by many orders of magnitude from the reference value computed on the intended subset with the intended test, so the numeric column fails the grader's tolerance check even when the reject/fail-to-reject label happens to match.
777Failure to conform to a provided reference output file's exact formatting/precisiontaskda-code
Applies when
task -- The task supplies a sample/template output file and asks that the deliverable follow its exact structure and formatting.
Pattern
The agent reads the template only to copy superficial aspects (e.g., column names or row order) and then dumps raw computed floats, ignoring the template's decimal precision, rounding, numeric formatting, column order/dtypes, or header text; the answer contains tell-tale unrounded floating-point values (e.g., long trailing digit strings) that could not match the template.
Detection procedure
  1. In the task, confirm a reference/sample output file is provided and that "exact structure and formatting" is required.
  2. In the scripts, check whether the sample file is actually parsed and used to derive column names, column order, row order, AND numeric formatting (rounding/decimal places, integer vs float, any string formatting) before writing; a mere hard-coded row/column order without rounding is a red flag.
  3. Inspect the produced output rows for values whose precision or type differs from the template's cells (e.g., many decimals where the sample shows a fixed number, or floats where the sample shows integers).
  4. Confirm no post-write comparison against the sample's formatting (only a re-print of the sample and the result, with no assertion on formatting, does not count).
Discriminator
A real violation is when the written values' precision/type/labels demonstrably differ from what the template exhibits; it is fine if the script explicitly rounds/formats to the template's convention (or the template genuinely shows full-precision values) even if the formatting logic is hard-coded.
Consequence
The grader compares cell-by-cell against the expected file and marks the output WRONG despite the underlying aggregation logic being close or correct.
id c964e8b8f070 · mined from da-code dacode-dm-csv-011@s15
raw text (what the judge reads)
### Failure to conform to a provided reference output file's exact formatting/precision
- **Applies when**: `task` -- The task supplies a sample/template output file and asks that the deliverable follow its exact structure and formatting.
- **Pattern**: The agent reads the template only to copy superficial aspects (e.g., column names or row order) and then dumps raw computed floats, ignoring the template's decimal precision, rounding, numeric formatting, column order/dtypes, or header text; the answer contains tell-tale unrounded floating-point values (e.g., long trailing digit strings) that could not match the template.
- **Detection procedure**:
  1. In the task, confirm a reference/sample output file is provided and that "exact structure and formatting" is required.
  2. In the scripts, check whether the sample file is actually parsed and used to derive column names, column order, row order, AND numeric formatting (rounding/decimal places, integer vs float, any string formatting) before writing; a mere hard-coded row/column order without rounding is a red flag.
  3. Inspect the produced output rows for values whose precision or type differs from the template's cells (e.g., many decimals where the sample shows a fixed number, or floats where the sample shows integers).
  4. Confirm no post-write comparison against the sample's formatting (only a re-print of the sample and the result, with no assertion on formatting, does not count).
- **Discriminator**: A real violation is when the written values' precision/type/labels demonstrably differ from what the template exhibits; it is fine if the script explicitly rounds/formats to the template's convention (or the template genuinely shows full-precision values) even if the formatting logic is hard-coded.
- **Consequence**: The grader compares cell-by-cell against the expected file and marks the output WRONG despite the underlying aggregation logic being close or correct.
778Requested output artifact not verified for schema, row count, and contenttaskda-code
Applies when
task -- the task requires writing a result file with an explicitly specified set of columns and one row per input record (e.g., per-sample labels or predictions alongside their feature values).
Pattern
The agent produces a narrative summary asserting the file was written, but never re-reads the saved file to confirm it exists with the exact required column names, one row per retained record, and the intended cell values; internally inconsistent counts (row totals that mix header/data, or a row count that doesn't match the number of records after any dropping/imputation/deduplication) go unnoticed, and transformed values (scaled/encoded/PCA) are silently substituted for the feature values the task asked to be saved.
Detection procedure
  1. From the task statement, write down the exact deliverable contract: filename, column names/pattern, number of columns, expected number of data rows, and what each cell should contain.
  2. In the scripts, locate the write call and trace back what the DataFrame's columns and index actually are — check that column names are generated to match the required pattern exactly, that no rows were dropped/filtered relative to the input without being accounted for, and that the values written are the ones requested rather than an intermediate representation.
  3. Confirm the script (or the agent) reloads the saved file and prints shape, column list, head, and label value counts as a post-write sanity check.
  4. Compare those printed facts with the counts stated in the final answer; any mismatch or ambiguous statement (e.g., a row total that only works if the header is counted as data) is a red flag.
Discriminator
A real violation is when no read-back/shape-and-column verification of the deliverable exists, or the verified facts contradict the task contract or the answer's own numbers. It is fine if the file is verified and the row count legitimately differs from the raw input because the script explicitly documented and justified the filtering, and the column names still match the required pattern exactly.
Consequence
The grader loads the expected file and finds it missing, mis-named/mis-shaped, or with wrong column headers or wrong cell contents, so the deliverable check fails outright regardless of how sensible the clustering/model narrative sounds.
id 20025332940d · mined from da-code dacode-ml-cluster-014@s15
raw text (what the judge reads)
### Requested output artifact not verified for schema, row count, and content
- **Applies when**: `task` -- the task requires writing a result file with an explicitly specified set of columns and one row per input record (e.g., per-sample labels or predictions alongside their feature values).
- **Pattern**: The agent produces a narrative summary asserting the file was written, but never re-reads the saved file to confirm it exists with the exact required column names, one row per retained record, and the intended cell values; internally inconsistent counts (row totals that mix header/data, or a row count that doesn't match the number of records after any dropping/imputation/deduplication) go unnoticed, and transformed values (scaled/encoded/PCA) are silently substituted for the feature values the task asked to be saved.
- **Detection procedure**:
  1. From the task statement, write down the exact deliverable contract: filename, column names/pattern, number of columns, expected number of data rows, and what each cell should contain.
  2. In the scripts, locate the write call and trace back what the DataFrame's columns and index actually are — check that column names are generated to match the required pattern exactly, that no rows were dropped/filtered relative to the input without being accounted for, and that the values written are the ones requested rather than an intermediate representation.
  3. Confirm the script (or the agent) reloads the saved file and prints shape, column list, head, and label value counts as a post-write sanity check.
  4. Compare those printed facts with the counts stated in the final answer; any mismatch or ambiguous statement (e.g., a row total that only works if the header is counted as data) is a red flag.
- **Discriminator**: A real violation is when no read-back/shape-and-column verification of the deliverable exists, or the verified facts contradict the task contract or the answer's own numbers. It is fine if the file is verified and the row count legitimately differs from the raw input because the script explicitly documented and justified the filtering, and the column names still match the required pattern exactly.
- **Consequence**: The grader loads the expected file and finds it missing, mis-named/mis-shaped, or with wrong column headers or wrong cell contents, so the deliverable check fails outright regardless of how sensible the clustering/model narrative sounds.
779Ignoring the provided submission template (no format/coverage validation, plus unrequested post-processing of predictions)taskda-code
Applies when
task -- the task supplies an example/template output file (or an explicit output spec) and the script builds its own result file from scratch.
Pattern
The script never loads or compares against the template: it hard-codes column names, ordering, row set and output path from assumptions, and then applies self-invented transformations to the predictions (integer rounding, clipping to a range guessed from training data) that the task never requested — silently changing both the file contract and the predicted values.
Detection procedure
  1. Read the task/README and note the template file, required columns, id set, dtype, and the exact expected output filename/location.
  2. Search the scripts for any read of the template and any assertion on row count, id set equality/order, column names, and dtype; absence of all of these is the first flag.
  3. Inspect the post-prediction steps: flag any round, astype(int), clip, rescaling, or sorting applied to the model output that is not demanded by the task or the template's dtype.
  4. Check the answer file itself: does the written path match the expected artifact location, does the row count equal the test set size, and are values plausibly distributed (not collapsed onto a coarse grid or truncated at invented bounds)?
Discriminator
A real violation is inferring the format/value domain instead of deriving it from the template, or coercing continuous predictions into integers/bounds with no stated requirement. It is fine if the script reads the template (or the task explicitly states integer/categorical output) and asserts that ids, columns and shape match — rounding is then a justified, verified step rather than a guess.
Consequence
The grader reports the expected output file as WRONG/MISSING — wrong path, wrong/missing rows or column names — or, if the file loads, the metric is measurably worse than the unrounded/unclipped predictions because the discretization adds avoidable error.
id 709bc4e801c3 · mined from da-code dacode-ml-competition-009@s15
raw text (what the judge reads)
### Ignoring the provided submission template (no format/coverage validation, plus unrequested post-processing of predictions)

- **Applies when**: `task` -- the task supplies an example/template output file (or an explicit output spec) and the script builds its own result file from scratch.
- **Pattern**: The script never loads or compares against the template: it hard-codes column names, ordering, row set and output path from assumptions, and then applies self-invented transformations to the predictions (integer rounding, clipping to a range guessed from training data) that the task never requested — silently changing both the file contract and the predicted values.
- **Detection procedure**:
  1. Read the task/README and note the template file, required columns, id set, dtype, and the exact expected output filename/location.
  2. Search the scripts for any read of the template and any assertion on row count, id set equality/order, column names, and dtype; absence of all of these is the first flag.
  3. Inspect the post-prediction steps: flag any `round`, `astype(int)`, `clip`, rescaling, or sorting applied to the model output that is not demanded by the task or the template's dtype.
  4. Check the answer file itself: does the written path match the expected artifact location, does the row count equal the test set size, and are values plausibly distributed (not collapsed onto a coarse grid or truncated at invented bounds)?
- **Discriminator**: A real violation is inferring the format/value domain instead of deriving it from the template, or coercing continuous predictions into integers/bounds with no stated requirement. It is fine if the script reads the template (or the task explicitly states integer/categorical output) and asserts that ids, columns and shape match — rounding is then a justified, verified step rather than a guess.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING — wrong path, wrong/missing rows or column names — or, if the file loads, the metric is measurably worse than the unrounded/unclipped predictions because the discretization adds avoidable error.
780Output feature matrix and row coverage don't match what was actually clustered / what the task impliestaskda-code
Applies when
task -- the task asks for per-record (or per-entity) labels saved to a file with generically named feature columns, and the scripts derive their own engineered/transformed features and drop rows before fitting.
Pattern
The attempt silently redefines the population (filtering out nulls, negatives, cancellations, aggregating many raw rows into far fewer entities) and then writes out a different representation than the one the model consumed (e.g., raw pre-scaling values, or a subset of the engineered columns), so the delivered Feature_i columns and row count don't correspond to the vectors that produced Cluster, nor to the full set of items the task asked to segment.
Detection procedure
  1. Read the task statement and note what unit each output row must represent and whether any filtering/aggregation was authorized; note the required column names and their implied meaning.
  2. Trace in the scripts, from raw load to to_csv, how many rows survive each filter/aggregation and which array is passed to fit/fit_predict versus which DataFrame columns are exported.
  3. Confirm the exported feature columns are exactly the (same values, same order, same row alignment as the) matrix used for clustering, and that the exported row count equals the number of labels and matches the task's expected unit/population.
  4. Check the answer for an explicit statement/sanity check of row count and feature definition; absent such reconciliation, or if unjustified filters shrank the population, flag it.
Discriminator
A legitimate attempt documents the unit of analysis and any dropped rows as required by the task, and its exported features are the same vectors that were clustered (scaling is fine only if the export is self-consistent and label-aligned); a violation exports a different transform/subset, or reduces the population by unrequested cleaning so the row count no longer matches the expected output.
Consequence
The saved file has the right headers but the wrong shape and wrong feature values, so a row-wise or shape-based comparison against the expected clustering file fails outright.
id 0c052a605635 · mined from da-code dacode-ml-cluster-016@s15
raw text (what the judge reads)
### Output feature matrix and row coverage don't match what was actually clustered / what the task implies
- **Applies when**: `task` -- the task asks for per-record (or per-entity) labels saved to a file with generically named feature columns, and the scripts derive their own engineered/transformed features and drop rows before fitting.
- **Pattern**: The attempt silently redefines the population (filtering out nulls, negatives, cancellations, aggregating many raw rows into far fewer entities) and then writes out a *different* representation than the one the model consumed (e.g., raw pre-scaling values, or a subset of the engineered columns), so the delivered `Feature_i` columns and row count don't correspond to the vectors that produced `Cluster`, nor to the full set of items the task asked to segment.
- **Detection procedure**:
  1. Read the task statement and note what unit each output row must represent and whether any filtering/aggregation was authorized; note the required column names and their implied meaning.
  2. Trace in the scripts, from raw load to `to_csv`, how many rows survive each filter/aggregation and which array is passed to `fit`/`fit_predict` versus which DataFrame columns are exported.
  3. Confirm the exported feature columns are exactly the (same values, same order, same row alignment as the) matrix used for clustering, and that the exported row count equals the number of labels and matches the task's expected unit/population.
  4. Check the answer for an explicit statement/sanity check of row count and feature definition; absent such reconciliation, or if unjustified filters shrank the population, flag it.
- **Discriminator**: A legitimate attempt documents the unit of analysis and any dropped rows as required by the task, and its exported features are the same vectors that were clustered (scaling is fine only if the export is self-consistent and label-aligned); a violation exports a different transform/subset, or reduces the population by unrequested cleaning so the row count no longer matches the expected output.
- **Consequence**: The saved file has the right headers but the wrong shape and wrong feature values, so a row-wise or shape-based comparison against the expected clustering file fails outright.
781Prediction file shape/alignment never verified against the test settaskda-code
Applies when
task -- the task requires writing a per-row output file (e.g., predictions with a specified column name) for a provided evaluation/test input, and the agent produces such a file.
Pattern
The attempt generates predictions and saves them without ever asserting that the output has exactly one row per input row, in the original input order, with the exact requested column name(s) and no index/extra columns; the delivered file ends up truncated, reordered, filtered (e.g., rows dropped by NA-handling or a merge), or renamed, so it cannot be aligned with ground truth.
Detection procedure
  1. From the task/data, note the expected output contract: number of rows in the evaluation input, required column name, ordering, and any rounding/units constraints.
  2. In the scripts, trace the object written to disk: check whether any step (dropna, filtering, groupby, merge, resample, deduplication, train/test slicing) could change row count or order between reading the input and writing the output, and whether the write uses index=False and the exact required header.
  3. Look for an explicit sanity check before/after writing (len(pred) == len(test), header equals required name, no NaNs, values in a plausible range); absence of any such assertion is the flag.
  4. Inspect the delivered file itself: count rows and compare to the evaluation input row count, and confirm the header string matches exactly.
Discriminator
A real violation is a row-count/order/header mismatch or a pipeline step that can silently drop or reorder rows with no verification; it is not a violation if the script preserves the input index end-to-end (or explicitly reindexes to the input keys) and the written file demonstrably has the same row count, order and required header, even if the numeric predictions are imperfect.
Consequence
The grader cannot align predictions with ground truth and marks the expected output file WRONG/MISSING regardless of model quality, scoring 0.
id 11fd01923b1c · mined from da-code dacode-ml-regression-002@s15
raw text (what the judge reads)
### Prediction file shape/alignment never verified against the test set
- **Applies when**: `task` -- the task requires writing a per-row output file (e.g., predictions with a specified column name) for a provided evaluation/test input, and the agent produces such a file.
- **Pattern**: The attempt generates predictions and saves them without ever asserting that the output has exactly one row per input row, in the original input order, with the exact requested column name(s) and no index/extra columns; the delivered file ends up truncated, reordered, filtered (e.g., rows dropped by NA-handling or a merge), or renamed, so it cannot be aligned with ground truth.
- **Detection procedure**:
  1. From the task/data, note the expected output contract: number of rows in the evaluation input, required column name, ordering, and any rounding/units constraints.
  2. In the scripts, trace the object written to disk: check whether any step (dropna, filtering, groupby, merge, resample, deduplication, train/test slicing) could change row count or order between reading the input and writing the output, and whether the write uses `index=False` and the exact required header.
  3. Look for an explicit sanity check before/after writing (`len(pred) == len(test)`, header equals required name, no NaNs, values in a plausible range); absence of any such assertion is the flag.
  4. Inspect the delivered file itself: count rows and compare to the evaluation input row count, and confirm the header string matches exactly.
- **Discriminator**: A real violation is a row-count/order/header mismatch or a pipeline step that can silently drop or reorder rows with no verification; it is *not* a violation if the script preserves the input index end-to-end (or explicitly reindexes to the input keys) and the written file demonstrably has the same row count, order and required header, even if the numeric predictions are imperfect.
- **Consequence**: The grader cannot align predictions with ground truth and marks the expected output file WRONG/MISSING regardless of model quality, scoring 0.
782Missing required output artifacts (only the human-readable/visual output produced)taskda-code
Applies when
task -- the task names or implies specific deliverable files (e.g. a spec/config file to follow plus serialized data/plot metadata), not just an image or a prose summary.
Pattern
The agent produces only the eye-catching artifact (a rendered figure) and a narrative description of it, never writing the machine-checkable files the task/spec requires, and never saving the scripts that would let the outputs be regenerated or inspected.
Detection procedure
1. Read the task and any referenced spec file to enumerate every expected output path and format (image, JSON/YAML of plot properties, numeric array dumps, CSV, etc.). 2. Read the scripts for explicit write calls to each of those paths (savefig, json.dump, np.save, to_csv) and confirm the saved values come from the actual plotted objects rather than hand-typed constants. 3. Compare the agent's answer: does it assert completion while listing only one file, and does it restate spec values as prose instead of pointing to written artifacts? 4. Flag if any enumerated deliverable has no corresponding write in code or on disk.
Discriminator
A real violation is a missing/unwritten required artifact or one populated from hardcoded text rather than the computed data; it is fine if all required files are written (possibly with extra optional files) even if the answer's prose summary is terse or imperfectly worded.
Consequence
The grader's per-file checks report the expected outputs as WRONG/MISSING, scoring 0 even though a plausible-looking chart exists.
id 6c3b7d2fa077 · mined from da-code dacode-plot-line-015@s15
raw text (what the judge reads)
### Missing required output artifacts (only the human-readable/visual output produced)
- **Applies when**: `task` -- the task names or implies specific deliverable files (e.g. a spec/config file to follow plus serialized data/plot metadata), not just an image or a prose summary.
- **Pattern**: The agent produces only the eye-catching artifact (a rendered figure) and a narrative description of it, never writing the machine-checkable files the task/spec requires, and never saving the scripts that would let the outputs be regenerated or inspected.
- **Detection procedure**: 1. Read the task and any referenced spec file to enumerate every expected output path and format (image, JSON/YAML of plot properties, numeric array dumps, CSV, etc.). 2. Read the scripts for explicit write calls to each of those paths (`savefig`, `json.dump`, `np.save`, `to_csv`) and confirm the saved values come from the actual plotted objects rather than hand-typed constants. 3. Compare the agent's answer: does it assert completion while listing only one file, and does it restate spec values as prose instead of pointing to written artifacts? 4. Flag if any enumerated deliverable has no corresponding write in code or on disk.
- **Discriminator**: A real violation is a missing/unwritten required artifact or one populated from hardcoded text rather than the computed data; it is fine if all required files are written (possibly with extra optional files) even if the answer's prose summary is terse or imperfectly worded.
- **Consequence**: The grader's per-file checks report the expected outputs as WRONG/MISSING, scoring 0 even though a plausible-looking chart exists.
783Prescribed resampling/analysis procedure silently replaced with a different (stronger) schemetaskda-code
Applies when
task -- The task or accompanying notes spell out a specific statistical procedure (which quantity to shift/transform, which population to resample from, how many replicates, how to count "as or more extreme") and the script implements a hypothesis test or simulation.
Pattern
The script follows the prescribed steps only superficially — e.g., it transforms each group as instructed but then pools the groups into one array and draws every replicate from that pooled array (or otherwise swaps in a permutation/pooled-variance scheme). This discards each group's own dispersion and sample-specific structure, making the null distribution far narrower or wider than the one the task defines, and typically yields a degenerate result (p = 0, or exactly 1).
Detection procedure
  1. From the task/README, write down the intended null-distribution recipe as a literal sequence: what is shifted, what set each resample is drawn from, what statistic is recorded.
  2. In the script, locate the resampling loop and check the source array of each draw: does group 1's replicate come from group 1's transformed data only, and group 2's from group 2's, as specified — or from a concatenation of both?
  3. Check the extremeness comparison (one- vs two-sided, >= vs >) and replicate count against the stated instructions.
  4. Inspect the reported value for degeneracy: a p-value of exactly 0 (or 1) with a modest number of replicates is a red flag that the null distribution's spread does not match the observed effect scale; compare the reported bootstrap std to the analytic SE of the difference (√(s₁²/n₁ + s₂²/n₂)).
Discriminator
A real violation is a resampling source or transformation that differs from the one the task text names (pooled draws when per-group draws were prescribed), or a reported statistic that a sanity check shows is degenerate. A look-alike that is fine: per-group resampling from correctly shifted data that legitimately produces a very small p-value, where the bootstrap std matches the analytic standard error and only the replicate count limits resolution (reported as e.g. "< 1e-4").
Consequence
The p-value comes from the wrong null distribution and is off by orders of magnitude (0 instead of a small but nonzero value), so the written result file fails the grader's value check.
id 98af2806dd12 · mined from da-code dacode-data-sa-028@s15
raw text (what the judge reads)
### Prescribed resampling/analysis procedure silently replaced with a different (stronger) scheme
- **Applies when**: `task` -- The task or accompanying notes spell out a specific statistical procedure (which quantity to shift/transform, which population to resample from, how many replicates, how to count "as or more extreme") and the script implements a hypothesis test or simulation.
- **Pattern**: The script follows the prescribed steps only superficially — e.g., it transforms each group as instructed but then pools the groups into one array and draws every replicate from that pooled array (or otherwise swaps in a permutation/pooled-variance scheme). This discards each group's own dispersion and sample-specific structure, making the null distribution far narrower or wider than the one the task defines, and typically yields a degenerate result (p = 0, or exactly 1).
- **Detection procedure**:
  1. From the task/README, write down the intended null-distribution recipe as a literal sequence: what is shifted, what set each resample is drawn from, what statistic is recorded.
  2. In the script, locate the resampling loop and check the source array of each draw: does group 1's replicate come from group 1's transformed data only, and group 2's from group 2's, as specified — or from a concatenation of both?
  3. Check the extremeness comparison (one- vs two-sided, `>=` vs `>`) and replicate count against the stated instructions.
  4. Inspect the reported value for degeneracy: a p-value of exactly 0 (or 1) with a modest number of replicates is a red flag that the null distribution's spread does not match the observed effect scale; compare the reported bootstrap std to the analytic SE of the difference (√(s₁²/n₁ + s₂²/n₂)).
- **Discriminator**: A real violation is a resampling source or transformation that differs from the one the task text names (pooled draws when per-group draws were prescribed), or a reported statistic that a sanity check shows is degenerate. A look-alike that is fine: per-group resampling from correctly shifted data that legitimately produces a very small p-value, where the bootstrap std matches the analytic standard error and only the replicate count limits resolution (reported as e.g. "< 1e-4").
- **Consequence**: The p-value comes from the wrong null distribution and is off by orders of magnitude (0 instead of a small but nonzero value), so the written result file fails the grader's value check.
784Spec files referenced by the task are never read; their contents are guessedtaskda-code
Applies when
task -- The instructions point to auxiliary specification/config files (e.g., a tips/notes file, a YAML/JSON formatting or parameter file) that define preprocessing rules, plotting style, or required output artifacts.
Pattern
The scripts never open, print, or parse those files; instead the agent hardcodes plausible-sounding rules ("excluded these categories", "seasons defined as...", chart colors, axis labels) and produces only the one obvious output file, silently skipping any additional artifacts (serialized figure spec, numeric array dump) the spec would have required.
Detection procedure
  1. From the task statement, list every referenced auxiliary file and every expected output artifact.
  2. Grep the scripts for reads of each referenced file (open, yaml.safe_load, json.load, read_csv) and for writes of each expected artifact.
  3. Check whether filter lists, category definitions, axis/label/style values, or ordering constants appear as literals in the code rather than being derived from the spec file.
  4. Check the final answer: does it assert compliance ("applied all requirements from the spec") without ever having loaded the spec, and does it enumerate fewer output files than the task implies?
Discriminator
A real violation is code that contains spec-derived constants with no corresponding file read, or that emits fewer artifacts than requested. It is not a violation if the script loads the spec and then applies its values (even via intermediate variables), or if the extra artifacts are written in a different script/step that is present in the submission.
Consequence
Every graded artifact mismatches — missing files score as absent, and the produced figure/values differ from the reference because the filtering, grouping, ordering, and styling rules were invented rather than read, yielding 0 passed checks despite a confident "task completed" report.
id e5f36e4c2375 · mined from da-code dacode-plot-line-006@s15
raw text (what the judge reads)
### Spec files referenced by the task are never read; their contents are guessed
- **Applies when**: `task` -- The instructions point to auxiliary specification/config files (e.g., a tips/notes file, a YAML/JSON formatting or parameter file) that define preprocessing rules, plotting style, or required output artifacts.
- **Pattern**: The scripts never open, print, or parse those files; instead the agent hardcodes plausible-sounding rules ("excluded these categories", "seasons defined as...", chart colors, axis labels) and produces only the one obvious output file, silently skipping any additional artifacts (serialized figure spec, numeric array dump) the spec would have required.
- **Detection procedure**:
  1. From the task statement, list every referenced auxiliary file and every expected output artifact.
  2. Grep the scripts for reads of each referenced file (`open`, `yaml.safe_load`, `json.load`, `read_csv`) and for writes of each expected artifact.
  3. Check whether filter lists, category definitions, axis/label/style values, or ordering constants appear as literals in the code rather than being derived from the spec file.
  4. Check the final answer: does it assert compliance ("applied all requirements from the spec") without ever having loaded the spec, and does it enumerate fewer output files than the task implies?
- **Discriminator**: A real violation is code that contains spec-derived constants with no corresponding file read, or that emits fewer artifacts than requested. It is *not* a violation if the script loads the spec and then applies its values (even via intermediate variables), or if the extra artifacts are written in a different script/step that is present in the submission.
- **Consequence**: Every graded artifact mismatches — missing files score as absent, and the produced figure/values differ from the reference because the filtering, grouping, ordering, and styling rules were invented rather than read, yielding 0 passed checks despite a confident "task completed" report.
785Ignoring an external specification file that defines the required derivationtaskda-code
Applies when
task -- The task points to an auxiliary document/spec (e.g. a .md/config/instructions file in the working directory) that defines how a variable must be binned, categorized, filtered, or otherwise transformed before the requested output is produced.
Pattern
The attempt never opens or quotes the spec, and instead uses the raw categories/values already present in the data (or an invented scheme) as if they were the required groups; it then declares success without emitting all the artifacts the task/harness expects, and without any saved script showing the mapping.
Detection procedure
  1. Read the task and list every referenced external file and every named output artifact/format.
  2. Search the scripts for code that reads the spec file (or an explicit, commented transcription of its rules); if absent, the derivation is unverified.
  3. Compare the category labels/edges in the answer against the raw distinct values of the source column: if they coincide one-to-one with the raw response options, no regrouping was applied.
  4. Confirm each promised output file is actually written by code (not merely asserted in prose) and that group counts sum to the correct, correctly filtered row total.
Discriminator
A legitimate attempt shows the spec's rules in code (bin edges/label mapping read from or transcribed from the file) and produces labels that differ from, or are provably identical by design to, the raw values; a violation shows no evidence the spec was ever consulted and merely relabels raw categories. Merely stylistic differences in the chart are not violations.
Consequence
The saved plot/array encodes the wrong grouping (and expected companion artifacts are missing), so every file-level equality check against the reference fails despite a confident "task completed" report.
id d1999394deac · mined from da-code dacode-plot-bar-005@s15
raw text (what the judge reads)
### Ignoring an external specification file that defines the required derivation
- **Applies when**: `task` -- The task points to an auxiliary document/spec (e.g. a `.md`/config/instructions file in the working directory) that defines how a variable must be binned, categorized, filtered, or otherwise transformed before the requested output is produced.
- **Pattern**: The attempt never opens or quotes the spec, and instead uses the raw categories/values already present in the data (or an invented scheme) as if they were the required groups; it then declares success without emitting all the artifacts the task/harness expects, and without any saved script showing the mapping.
- **Detection procedure**:
  1. Read the task and list every referenced external file and every named output artifact/format.
  2. Search the scripts for code that reads the spec file (or an explicit, commented transcription of its rules); if absent, the derivation is unverified.
  3. Compare the category labels/edges in the answer against the raw distinct values of the source column: if they coincide one-to-one with the raw response options, no regrouping was applied.
  4. Confirm each promised output file is actually written by code (not merely asserted in prose) and that group counts sum to the correct, correctly filtered row total.
- **Discriminator**: A legitimate attempt shows the spec's rules in code (bin edges/label mapping read from or transcribed from the file) and produces labels that differ from, or are provably identical by design to, the raw values; a violation shows no evidence the spec was ever consulted and merely relabels raw categories. Merely stylistic differences in the chart are not violations.
- **Consequence**: The saved plot/array encodes the wrong grouping (and expected companion artifacts are missing), so every file-level equality check against the reference fails despite a confident "task completed" report.
786Answer never persisted in the requested artifact/format (printed to stdout instead)taskda-code
Applies when
task -- the task specifies a deliverable output (a named result file and/or an exact JSON/text schema, e.g. keys mapped to lists) and the agent's scripts only compute and print values.
Pattern
The script does the analysis and prints the findings, then the agent retypes them in chat; no code writes the expected file, and the emitted structure deviates from the stated schema (scalars where the template shows bracketed lists, renamed/extra/missing keys, differing rounding or units).
Detection procedure
  1. Read the task and note every hard output requirement: file name/location, key names, value container type (list vs scalar), numeric formatting.
  2. Grep the scripts for any write operation (json.dump, to_csv, open(..., 'w'), to_json) targeting that file; if only print statements exist, the deliverable is missing.
  3. Compare the agent's final answer character-by-character against the template: are keys spelled identically, and are values wrapped exactly as shown (e.g. [...])?
  4. Flag as inadequate if the file is absent or the structure/types differ, regardless of whether the computed numbers look plausible.
Discriminator
A real violation is when no code path produces the required artifact or the schema shape differs (scalar vs list, wrong key names, wrong precision). A look-alike that is fine is a script that writes the correct file and also prints a human-readable summary, or one whose values are equivalent but formatted exactly as the template allows.
Consequence
The grader looks for the specified result file and schema, finds it missing or shape-mismatched, and marks the attempt wrong (0/1) even if the underlying analysis identified the right entity and value.
id a47e8f10834e · mined from da-code dacode-di-text-002@s15
raw text (what the judge reads)
### Answer never persisted in the requested artifact/format (printed to stdout instead)
- **Applies when**: `task` -- the task specifies a deliverable output (a named result file and/or an exact JSON/text schema, e.g. keys mapped to lists) and the agent's scripts only compute and print values.
- **Pattern**: The script does the analysis and `print`s the findings, then the agent retypes them in chat; no code writes the expected file, and the emitted structure deviates from the stated schema (scalars where the template shows bracketed lists, renamed/extra/missing keys, differing rounding or units).
- **Detection procedure**:
  1. Read the task and note every hard output requirement: file name/location, key names, value container type (list vs scalar), numeric formatting.
  2. Grep the scripts for any write operation (`json.dump`, `to_csv`, `open(..., 'w')`, `to_json`) targeting that file; if only `print` statements exist, the deliverable is missing.
  3. Compare the agent's final answer character-by-character against the template: are keys spelled identically, and are values wrapped exactly as shown (e.g. `[...]`)?
  4. Flag as inadequate if the file is absent or the structure/types differ, regardless of whether the computed numbers look plausible.
- **Discriminator**: A real violation is when no code path produces the required artifact or the schema shape differs (scalar vs list, wrong key names, wrong precision). A look-alike that is fine is a script that writes the correct file and also prints a human-readable summary, or one whose values are equivalent but formatted exactly as the template allows.
- **Consequence**: The grader looks for the specified result file and schema, finds it missing or shape-mismatched, and marks the attempt wrong (0/1) even if the underlying analysis identified the right entity and value.
787Silently excluding candidate columns the task said to includetaskinfiagent-dabench
Applies when
task -- the task says to compute a statistic against "all other numerical variables" (or all features) and pick the extremum, and the script builds the candidate list itself
Pattern
The script filters out one or more numerical columns it judges "trivial", "redundant", or "derived" (e.g., an index/rank/ID or an alternate encoding of the target) without any instruction to do so, then reports the extremum of the reduced set.
Detection procedure
1. Read the task and note exactly which columns are in scope and which (if any) exclusions are stated. 2. In the script, find where the candidate column list is constructed and list every column dropped by name, dtype filter, or hard-coded exclusion. 3. Compare: any drop not explicitly authorized by the task is a violation, especially one that could plausibly hold the strongest relationship. 4. Check whether the reported answer would change if that column were reinstated.
Discriminator
Legitimate exclusions are the target variable itself, non-numeric columns, or exclusions the task explicitly states; a violation is an extra exclusion justified only by the agent's own reasoning about redundancy or interpretability.
Consequence
The grader expects the column the agent removed, so both the variable name and the sign/direction derived from it are reported wrong (0 checks passed).
id 62f79959b381 · mined from infiagent-dabench dabench-117@s15
raw text (what the judge reads)
### Silently excluding candidate columns the task said to include
- **Applies when**: `task` -- the task says to compute a statistic against "all other numerical variables" (or all features) and pick the extremum, and the script builds the candidate list itself
- **Pattern**: The script filters out one or more numerical columns it judges "trivial", "redundant", or "derived" (e.g., an index/rank/ID or an alternate encoding of the target) without any instruction to do so, then reports the extremum of the reduced set.
- **Detection procedure**: 1. Read the task and note exactly which columns are in scope and which (if any) exclusions are stated. 2. In the script, find where the candidate column list is constructed and list every column dropped by name, dtype filter, or hard-coded exclusion. 3. Compare: any drop not explicitly authorized by the task is a violation, especially one that could plausibly hold the strongest relationship. 4. Check whether the reported answer would change if that column were reinstated.
- **Discriminator**: Legitimate exclusions are the target variable itself, non-numeric columns, or exclusions the task explicitly states; a violation is an extra exclusion justified only by the agent's own reasoning about redundancy or interpretability.
- **Consequence**: The grader expects the column the agent removed, so both the variable name and the sign/direction derived from it are reported wrong (0 checks passed).
788Statistical test run on an unverified/unsanitized input vector (and required intermediate not reported)taskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test and/or summary statistics on a single column, with a stated decision rule and an explicit instruction to report an intermediate quantity (e.g., the p-value).
Pattern
The attempt loads the file and pipes the column straight into the test/statistic without establishing what rows actually entered it — no check of dtype coercion, missing/sentinel values (NaN, empty strings, 0/-999 placeholders), duplicated or non-target rows, or the required filtering — and then reports only the final verdict, omitting the p-value the task asked to report and any evidence of the vector's size/range. A shifted or contaminated input silently flips the test decision.
Detection procedure
  1. Read the task and list every stated requirement: which rows/column, the test to use, alpha, the decision direction, the values to report (including intermediates such as the p-value) and rounding/format.
  2. In the scripts, locate the exact expression passed to the test/statistic and check whether it is preceded by explicit dtype conversion, missing-value handling, and any required row filtering — and whether the script prints n, min/max/mean of that vector.
  3. Check whether the script prints the required intermediate (p-value) and whether the reported verdict is derived from it by the stated rule rather than asserted.
  4. Cross-check consistency: verify that the reported shape statistics, n, and the test outcome tell the same story (e.g., near-zero skew/kurtosis with a rejection, or extreme skew with "normal", or a suspiciously huge/tiny n) — any mismatch means the input vector was probably not the intended one.
Discriminator
A real violation is when the analyzed vector's composition is never verified or printed (no n/NaN/dtype evidence) and the required intermediate is missing, so the reviewer cannot tell which rows produced the verdict. It is not a violation if the script explicitly cleans and prints the vector's n, dtype and dropped-row count, reports the p-value, and the verdict follows the stated rule — even if the conclusion happens to be borderline.
Consequence
The decision string (and often the shape statistics) is computed from the wrong sample, so the graded categorical answer inverts and all checks fail, with no printed p-value or row counts to diagnose the discrepancy.
id cbc717888a41 · mined from infiagent-dabench dabench-298@s15
raw text (what the judge reads)
### Statistical test run on an unverified/unsanitized input vector (and required intermediate not reported)
- **Applies when**: `task` -- the task asks for a hypothesis test and/or summary statistics on a single column, with a stated decision rule and an explicit instruction to report an intermediate quantity (e.g., the p-value).
- **Pattern**: The attempt loads the file and pipes the column straight into the test/statistic without establishing what rows actually entered it — no check of dtype coercion, missing/sentinel values (NaN, empty strings, 0/-999 placeholders), duplicated or non-target rows, or the required filtering — and then reports only the final verdict, omitting the p-value the task asked to report and any evidence of the vector's size/range. A shifted or contaminated input silently flips the test decision.
- **Detection procedure**:
  1. Read the task and list every stated requirement: which rows/column, the test to use, alpha, the decision direction, the values to report (including intermediates such as the p-value) and rounding/format.
  2. In the scripts, locate the exact expression passed to the test/statistic and check whether it is preceded by explicit dtype conversion, missing-value handling, and any required row filtering — and whether the script prints n, min/max/mean of that vector.
  3. Check whether the script prints the required intermediate (p-value) and whether the reported verdict is derived from it by the stated rule rather than asserted.
  4. Cross-check consistency: verify that the reported shape statistics, n, and the test outcome tell the same story (e.g., near-zero skew/kurtosis with a rejection, or extreme skew with "normal", or a suspiciously huge/tiny n) — any mismatch means the input vector was probably not the intended one.
- **Discriminator**: A real violation is when the analyzed vector's composition is never verified or printed (no n/NaN/dtype evidence) and the required intermediate is missing, so the reviewer cannot tell which rows produced the verdict. It is *not* a violation if the script explicitly cleans and prints the vector's n, dtype and dropped-row count, reports the p-value, and the verdict follows the stated rule — even if the conclusion happens to be borderline.
- **Consequence**: The decision string (and often the shape statistics) is computed from the wrong sample, so the graded categorical answer inverts and all checks fail, with no printed p-value or row counts to diagnose the discrepancy.
789No held-out validation before choosing/overwriting the final modeltaskda-code
Applies when
task -- the task asks for predictions scored by an unseen metric, and the scripts fit several candidate models (or successively rewrite the same output file) to produce the submission.
Pattern
The attempt trains models and writes predictions directly from a fit on all training data, with no train/validation split or cross-validation, no reported error metric per candidate, and no comparison against a trivial baseline; the file that survives is the last script run (often a deliberately cheaper/lower-capacity "quick" version) rather than the best-scoring one, and the answer justifies the choice only with qualitative prose and prediction summary statistics.
Detection procedure
  1. From the task, identify that submitted predictions will be scored against hidden targets, so predictive accuracy — not file existence — is what matters.
  2. In the scripts, search for any holdout/CV evaluation (train_test_split used for scoring, cross_val_score, an explicit metric computed on unseen rows) and for a printed metric per candidate model; note whether an earlier script's submission is overwritten by a later, simpler one.
  3. In the answer, check whether a validation score (and ideally a baseline comparison) is quoted as the reason for the chosen model; treat prediction min/max/mean/std, feature lists, and row counts as non-evidence of accuracy.
  4. Flag if no unseen-data score exists for the model that actually produced the delivered file, or if hyperparameters/model family were reduced for runtime without re-measuring the cost in accuracy.
Discriminator
A real violation is zero quantitative evidence on unseen data for the delivered model (or a downgrade justified only by speed); it is not a violation when the agent reports holdout/CV metrics for each candidate, picks the best, and merely also reports distribution summaries, nor when a single model is chosen after documented CV even if simple.
Consequence
The submitted predictions come from an unvalidated, likely underfitting model whose hidden-set score falls below the grader's accuracy threshold, so the expected submission file is judged WRONG despite having the correct shape and column names.
id 31a3fda9abaf · mined from da-code dacode-ml-competition-008@s15
raw text (what the judge reads)
### No held-out validation before choosing/overwriting the final model
- **Applies when**: `task` -- the task asks for predictions scored by an unseen metric, and the scripts fit several candidate models (or successively rewrite the same output file) to produce the submission.
- **Pattern**: The attempt trains models and writes predictions directly from a `fit` on all training data, with no train/validation split or cross-validation, no reported error metric per candidate, and no comparison against a trivial baseline; the file that survives is the last script run (often a deliberately cheaper/lower-capacity "quick" version) rather than the best-scoring one, and the answer justifies the choice only with qualitative prose and prediction summary statistics.
- **Detection procedure**:
  1. From the task, identify that submitted predictions will be scored against hidden targets, so predictive accuracy — not file existence — is what matters.
  2. In the scripts, search for any holdout/CV evaluation (`train_test_split` used for scoring, `cross_val_score`, an explicit metric computed on unseen rows) and for a printed metric per candidate model; note whether an earlier script's submission is overwritten by a later, simpler one.
  3. In the answer, check whether a validation score (and ideally a baseline comparison) is quoted as the reason for the chosen model; treat prediction min/max/mean/std, feature lists, and row counts as non-evidence of accuracy.
  4. Flag if no unseen-data score exists for the model that actually produced the delivered file, or if hyperparameters/model family were reduced for runtime without re-measuring the cost in accuracy.
- **Discriminator**: A real violation is zero quantitative evidence on unseen data for the delivered model (or a downgrade justified only by speed); it is *not* a violation when the agent reports holdout/CV metrics for each candidate, picks the best, and merely also reports distribution summaries, nor when a single model is chosen after documented CV even if simple.
- **Consequence**: The submitted predictions come from an unvalidated, likely underfitting model whose hidden-set score falls below the grader's accuracy threshold, so the expected submission file is judged WRONG despite having the correct shape and column names.
790Invented qualification threshold / aggregation rule instead of the spec provided in the tasktaskda-code
Applies when
task -- the task or its documentation supplies an explicit definition, eligibility filter, or a sample/template output file for the requested rankings or aggregate statistics.
Pattern
The agent never opens or echoes the provided definition/sample artifact, and instead hard-codes its own guessed cutoff (e.g., "minimum N observations per group"), its own aggregation choice (sum vs. mean vs. count, per-item vs. per-group), and its own column layout, then reports the resulting ranking as if authoritative.
Detection procedure
  1. From the task text and README, list every stated definition, minimum-count/eligibility rule, unit/rounding rule, and any referenced sample output file — including ones that are truncated or only partially quoted.
  2. Search the scripts for a read/print of that sample file and for a comment or code path tying each threshold and each aggregation function back to the stated rule.
  3. Flag if any threshold or aggregation appears as a bare literal ("threshold = 1000") or an arbitrary choice with no source, or if the output schema/column names/ordering were constructed by hand rather than copied from the sample.
  4. Check the final file's header, row count, and index column against the sample's exact layout.
Discriminator
Fine if the spec truly leaves the parameter open and the agent documents the choice and shows a sensitivity check across plausible values; a violation when a spec or sample existed (or was visibly truncated and never retrieved) and the agent silently substituted its own value, so a different plausible cutoff would reorder the reported top-N.
Consequence
The ranked names and/or the output schema differ from the reference file, and exact-match file comparison fails (0 checks passed) even though the code runs without error.
id bfbb79ebbfa1 · mined from da-code dacode-dm-csv-009@s15
raw text (what the judge reads)
### Invented qualification threshold / aggregation rule instead of the spec provided in the task
- **Applies when**: `task` -- the task or its documentation supplies an explicit definition, eligibility filter, or a sample/template output file for the requested rankings or aggregate statistics.
- **Pattern**: The agent never opens or echoes the provided definition/sample artifact, and instead hard-codes its own guessed cutoff (e.g., "minimum N observations per group"), its own aggregation choice (sum vs. mean vs. count, per-item vs. per-group), and its own column layout, then reports the resulting ranking as if authoritative.
- **Detection procedure**:
  1. From the task text and README, list every stated definition, minimum-count/eligibility rule, unit/rounding rule, and any referenced sample output file — including ones that are truncated or only partially quoted.
  2. Search the scripts for a read/print of that sample file and for a comment or code path tying each threshold and each aggregation function back to the stated rule.
  3. Flag if any threshold or aggregation appears as a bare literal ("threshold = 1000") or an arbitrary choice with no source, or if the output schema/column names/ordering were constructed by hand rather than copied from the sample.
  4. Check the final file's header, row count, and index column against the sample's exact layout.
- **Discriminator**: Fine if the spec truly leaves the parameter open and the agent documents the choice and shows a sensitivity check across plausible values; a violation when a spec or sample existed (or was visibly truncated and never retrieved) and the agent silently substituted its own value, so a different plausible cutoff would reorder the reported top-N.
- **Consequence**: The ranked names and/or the output schema differ from the reference file, and exact-match file comparison fails (0 checks passed) even though the code runs without error.
791Substituting self-generated inputs/config for the provided ones (and skipping required output artifacts)taskda-code
Applies when
task -- the task references specific supplied input files (a dataset, a config/spec file) and expects a specific set of output artifacts.
Pattern
The agent cannot find or does not locate the real inputs, so it creates a small synthetic stand-in dataset and/or writes its own config with invented parameter values, then runs the pipeline on those fabricated inputs and reports results as if they came from the real data; it also emits only the most obvious output file and silently omits other required artifacts (serialized figure spec, numeric arrays, tables).
Detection procedure
  1. From the task statement, list every input file the scripts must consume and every output artifact that must exist at the end (names, formats, locations).
  2. Read the scripts and any setup steps: confirm each input is read from the provided path and never written/generated by the agent, and confirm the config's keys/values are loaded rather than hard-coded or authored by the agent.
  3. Cross-check the answer's reported data facts (row counts, category counts, value ranges) against what the real source should plausibly contain; a suspiciously round sample size or a claim of having "generated" the data/config is a red flag.
  4. Verify each required output artifact is produced by the code; if any is absent, the attempt is incomplete regardless of chart correctness.
Discriminator
Legitimate attempts may write derived intermediates (cached frames, saved figures) but always trace back to reading the supplied source files; a violation is when the source data or the specification itself originates from the agent, or when required deliverables are never written.
Consequence
Every artifact check fails — numbers, bin edges, and styling are drawn from invented inputs so they cannot match ground truth, and missing files are scored as wrong/missing outright.
id bf781d92b6a9 · mined from da-code dacode-plot-bar-007@s15
raw text (what the judge reads)
### Substituting self-generated inputs/config for the provided ones (and skipping required output artifacts)
- **Applies when**: `task` -- the task references specific supplied input files (a dataset, a config/spec file) and expects a specific set of output artifacts.
- **Pattern**: The agent cannot find or does not locate the real inputs, so it creates a small synthetic stand-in dataset and/or writes its own config with invented parameter values, then runs the pipeline on those fabricated inputs and reports results as if they came from the real data; it also emits only the most obvious output file and silently omits other required artifacts (serialized figure spec, numeric arrays, tables).
- **Detection procedure**:
  1. From the task statement, list every input file the scripts must consume and every output artifact that must exist at the end (names, formats, locations).
  2. Read the scripts and any setup steps: confirm each input is *read* from the provided path and never *written*/generated by the agent, and confirm the config's keys/values are loaded rather than hard-coded or authored by the agent.
  3. Cross-check the answer's reported data facts (row counts, category counts, value ranges) against what the real source should plausibly contain; a suspiciously round sample size or a claim of having "generated" the data/config is a red flag.
  4. Verify each required output artifact is produced by the code; if any is absent, the attempt is incomplete regardless of chart correctness.
- **Discriminator**: Legitimate attempts may write *derived* intermediates (cached frames, saved figures) but always trace back to reading the supplied source files; a violation is when the source data or the specification itself originates from the agent, or when required deliverables are never written.
- **Consequence**: Every artifact check fails — numbers, bin edges, and styling are drawn from invented inputs so they cannot match ground truth, and missing files are scored as wrong/missing outright.
792Degenerate/empty-subset result reported without verification or in a format that violates the stated answer spectaskinfiagent-dabench
Applies when
task -- a task asks for a statistic on a filtered subset (or any computation that can yield an undefined/empty result) and specifies a strict answer format such as a float rounded to N decimals inside a token.
Pattern
The attempt's filtering produces an empty or all-missing subset, and the agent simply emits whatever placeholder its tooling printed (e.g. an alternately-cased or spelled null token, an integer, a quoted string) without (a) checking whether the emptiness is real or caused by a filtering/dtype/case mismatch, and (b) rendering the value exactly as the format template requires. No saved script exists to show how the subset and statistic were derived.
Detection procedure
  1. Read the task for the exact filter chain, the statistic, and the literal answer template (rounding, decimals, token name, capitalization).
  2. In the scripts, confirm each filter is applied to the stated column with matching dtype/value representation, and that row counts after each filter are printed; check that the degenerate case is handled explicitly rather than falling through.
  3. Compare the emitted answer string character-by-character to the template: does it match required rounding/decimal places, and if the value is undefined, is the null token written in the canonical lowercase/standard form the grader would parse?
  4. If no script or no intermediate counts are available, treat the result as unverified — the reviewer cannot distinguish a genuinely empty subset from a bad filter.
Discriminator
A real violation is an unverified and/or non-canonically formatted degenerate output (no post-filter counts, or token casing/precision deviating from the spec). It is fine if the script prints subset sizes and missing-value counts, shows the emptiness is inherent to the data, and emits the null/statistic in exactly the requested literal form.
Consequence
The grader's string/numeric comparison fails to match the expected value even when the underlying computation was arguably right, scoring 0.
id fca08e5d8964 · mined from infiagent-dabench dabench-554@s15
raw text (what the judge reads)
### Degenerate/empty-subset result reported without verification or in a format that violates the stated answer spec
- **Applies when**: `task` -- a task asks for a statistic on a filtered subset (or any computation that can yield an undefined/empty result) and specifies a strict answer format such as a float rounded to N decimals inside a token.
- **Pattern**: The attempt's filtering produces an empty or all-missing subset, and the agent simply emits whatever placeholder its tooling printed (e.g. an alternately-cased or spelled null token, an integer, a quoted string) without (a) checking whether the emptiness is real or caused by a filtering/dtype/case mismatch, and (b) rendering the value exactly as the format template requires. No saved script exists to show how the subset and statistic were derived.
- **Detection procedure**:
  1. Read the task for the exact filter chain, the statistic, and the literal answer template (rounding, decimals, token name, capitalization).
  2. In the scripts, confirm each filter is applied to the stated column with matching dtype/value representation, and that row counts after each filter are printed; check that the degenerate case is handled explicitly rather than falling through.
  3. Compare the emitted answer string character-by-character to the template: does it match required rounding/decimal places, and if the value is undefined, is the null token written in the canonical lowercase/standard form the grader would parse?
  4. If no script or no intermediate counts are available, treat the result as unverified — the reviewer cannot distinguish a genuinely empty subset from a bad filter.
- **Discriminator**: A real violation is an unverified and/or non-canonically formatted degenerate output (no post-filter counts, or token casing/precision deviating from the spec). It is fine if the script prints subset sizes and missing-value counts, shows the emptiness is inherent to the data, and emits the null/statistic in exactly the requested literal form.
- **Consequence**: The grader's string/numeric comparison fails to match the expected value even when the underlying computation was arguably right, scoring 0.
793Ranking/scoring not derived from the externally specified formula (and its subset-dependent parameters)taskda-code
Applies when
task -- the task points to an external spec (formula file, README, sample output) that defines a composite score, a filtering threshold, and an output format, and the agent must produce a ranked list or metric file from it.
Pattern
The agent never demonstrably parses/implements the given formula end-to-end: it applies a plausible-looking default ranking (raw counts, raw averages, or a library/default "popularity" order), computes formula constants (e.g., the global mean or the quantile cutoff) on the full data instead of the filtered subset the task specifies (or vice versa), and/or writes the result to a different filename/header/ordering than the sample — often with no saved script showing the computation.
Detection procedure
  1. Read the task and open the referenced spec/sample files; write down every required ingredient: the filter rule (which quantile/threshold, which population it is computed over), each symbol in the formula and the subset each is estimated from, the number of rows requested, and the exact output filename/columns/order.
  2. Read the scripts for a line-by-line correspondence: is the threshold computed with the right quantile direction and on the right population; is every formula symbol present (no dropped weighting term); is the score computed only on the retained rows; is the final write using the exact required path and header from the sample.
  3. Cross-check the answer against a trivial baseline: recompute (or eyeball) the top-N under a naive sort of a single raw column; if the submitted list is identical or near-identical in order to that baseline, treat the formula as unimplemented until the script proves otherwise.
  4. Confirm the deliverable artifact exists on disk with the requested row count and column names, not just printed in the chat/answer text.
Discriminator
A real violation is when the code (or the absence of code) cannot be mapped to every term/constraint of the spec, or the required file is absent/misnamed/mis-headed. A look-alike that is fine is a correctly implemented formula whose output happens to resemble a naive ranking — acceptable only if the script explicitly computes the threshold and each formula term on the specified subsets and writes the exact required file.
Consequence
The grader compares the required output file to the reference and marks it WRONG/MISSING — either the file is not produced in the expected location/format, or the ranking differs because the weighting/filtering was skipped or parameterized from the wrong subset.
id ffacf664beb3 · mined from da-code dacode-dm-csv-007@s15
raw text (what the judge reads)
### Ranking/scoring not derived from the externally specified formula (and its subset-dependent parameters)
- **Applies when**: `task` -- the task points to an external spec (formula file, README, sample output) that defines a composite score, a filtering threshold, and an output format, and the agent must produce a ranked list or metric file from it.
- **Pattern**: The agent never demonstrably parses/implements the given formula end-to-end: it applies a plausible-looking default ranking (raw counts, raw averages, or a library/default "popularity" order), computes formula constants (e.g., the global mean or the quantile cutoff) on the full data instead of the filtered subset the task specifies (or vice versa), and/or writes the result to a different filename/header/ordering than the sample — often with no saved script showing the computation.
- **Detection procedure**:
  1. Read the task and open the referenced spec/sample files; write down every required ingredient: the filter rule (which quantile/threshold, which population it is computed over), each symbol in the formula and the subset each is estimated from, the number of rows requested, and the exact output filename/columns/order.
  2. Read the scripts for a line-by-line correspondence: is the threshold computed with the right quantile direction and on the right population; is every formula symbol present (no dropped weighting term); is the score computed only on the retained rows; is the final write using the exact required path and header from the sample.
  3. Cross-check the answer against a trivial baseline: recompute (or eyeball) the top-N under a naive sort of a single raw column; if the submitted list is identical or near-identical in order to that baseline, treat the formula as unimplemented until the script proves otherwise.
  4. Confirm the deliverable artifact exists on disk with the requested row count and column names, not just printed in the chat/answer text.
- **Discriminator**: A real violation is when the code (or the absence of code) cannot be mapped to every term/constraint of the spec, or the required file is absent/misnamed/mis-headed. A look-alike that is fine is a correctly implemented formula whose output happens to resemble a naive ranking — acceptable only if the script explicitly computes the threshold and each formula term on the specified subsets and writes the exact required file.
- **Consequence**: The grader compares the required output file to the reference and marks it WRONG/MISSING — either the file is not produced in the expected location/format, or the ranking differs because the weighting/filtering was skipped or parameterized from the wrong subset.
794Answer produced without a verifiable script over the provided data (unreproducible, plausible-looking values)taskda-code
Applies when
task -- the task asks for specific records/values to be extracted from a supplied dataset after stated preprocessing (e.g., a named imputation rule) and with stated ordering/format, and the deliverable is a small list or JSON of entities.
Pattern
The attempt reports an answer that looks domain-plausible (well-known extremes, "obvious" rankings) but there is no saved/executed code that loads the file, applies the required preprocessing, computes the ranking, and writes the required output artifact — so the answer is effectively recalled or eyeballed rather than derived from the actual rows, and dataset-specific quirks (naming conventions used in the file, rows present/absent, imputed values, tie order, required sort direction within each group) are never respected.
Detection procedure
  1. From the task, list the mandatory computational steps and constraints: the preprocessing rule, the selection/ranking rule, the sort direction for each requested group, the output file name and JSON schema.
  2. In the scripts, verify each step exists in code: dataset load, cleaning of the target field into numeric dtype, the named imputation, the sort/nlargest/nsmallest, and an explicit write of the required artifact; if scripts are absent or do not touch the data file, stop — the answer is unverifiable.
  3. Cross-check the reported entity labels against the dataset's own label spelling/coverage and check the ordering inside each returned list matches the requested direction; recompute one group by hand from the printed intermediate table if available.
  4. Confirm the artifact actually exists at the requested path with the exact keys and list lengths.
Discriminator
A real violation is when no executed code produces the reported values (or the code omits a stated constraint such as the imputation or per-group sort direction), so the labels/order cannot be traced to the file. It is not a violation if a script demonstrably computes the result from the data and the answer merely happens to coincide with well-known real-world extremes.
Consequence
The grader compares against values derived from the actual file and the required key/order/format, and marks the submission wrong or the expected artifact missing, even though the list "looks right".
id 7b2ed9e83ee4 · mined from da-code dacode-di-text-003@s15
raw text (what the judge reads)
### Answer produced without a verifiable script over the provided data (unreproducible, plausible-looking values)
- **Applies when**: `task` -- the task asks for specific records/values to be extracted from a supplied dataset after stated preprocessing (e.g., a named imputation rule) and with stated ordering/format, and the deliverable is a small list or JSON of entities.
- **Pattern**: The attempt reports an answer that looks domain-plausible (well-known extremes, "obvious" rankings) but there is no saved/executed code that loads the file, applies the required preprocessing, computes the ranking, and writes the required output artifact — so the answer is effectively recalled or eyeballed rather than derived from the actual rows, and dataset-specific quirks (naming conventions used in the file, rows present/absent, imputed values, tie order, required sort direction within each group) are never respected.
- **Detection procedure**:
  1. From the task, list the mandatory computational steps and constraints: the preprocessing rule, the selection/ranking rule, the sort direction for *each* requested group, the output file name and JSON schema.
  2. In the scripts, verify each step exists in code: dataset load, cleaning of the target field into numeric dtype, the named imputation, the sort/`nlargest`/`nsmallest`, and an explicit write of the required artifact; if scripts are absent or do not touch the data file, stop — the answer is unverifiable.
  3. Cross-check the reported entity labels against the dataset's own label spelling/coverage and check the ordering inside each returned list matches the requested direction; recompute one group by hand from the printed intermediate table if available.
  4. Confirm the artifact actually exists at the requested path with the exact keys and list lengths.
- **Discriminator**: A real violation is when no executed code produces the reported values (or the code omits a stated constraint such as the imputation or per-group sort direction), so the labels/order cannot be traced to the file. It is *not* a violation if a script demonstrably computes the result from the data and the answer merely happens to coincide with well-known real-world extremes.
- **Consequence**: The grader compares against values derived from the actual file and the required key/order/format, and marks the submission wrong or the expected artifact missing, even though the list "looks right".
795Analysis answers a different question than the task specifies (wrong inputs, wrong chart/aggregation, missing required artifacts)taskda-code
Applies when
task -- the task names the required computation (e.g., a specific aggregate per category for a ranked subset), the required visualization type, and the exact output files/settings to produce.
Pattern
The script loads whichever file happens to be present in the data directory and computes some plausible-looking but unrelated statistic and plot type; the requested grouping/ranking/aggregation is never performed, the required chart form is not produced, and only a subset of the mandated output artifacts is written. The final answer confidently narrates the substituted analysis instead of flagging the mismatch.
Detection procedure
  1. From the task text, list the required elements: input entities/columns referenced, the aggregation/statistic, any subsetting/ranking rule, the chart type/orientation, and every output file name plus the config file whose keys must be honored.
  2. Read the script and map each required element to the code that implements it; check that the loaded file actually contains the referenced fields and that the config keys used match the config file's real contents.
  3. Check the write calls against the required artifact list (all files, correct names/locations/formats), and confirm the plot call matches the requested chart type.
  4. Read the final answer: does it describe the requested quantities and entities, or a different analysis? Does it acknowledge any data/column mismatch?
Discriminator
A real violation is when the computed quantity, grouping, chart type, or produced files cannot be mapped onto the task's stated requirements (or the columns used don't exist in the described schema). A look-alike that is fine is a faithful implementation that merely differs in cosmetic choices (colors, figure size, label wording) while the aggregation, subsetting, chart type, and all required artifacts are present.
Consequence
Every graded artifact fails — numeric outputs and plot metadata don't exist or contain values from an unrelated computation — so the score is 0 despite a confident "task completed" report.
id e9eb9eb03a36 · mined from da-code dacode-plot-scatter-002@s15
raw text (what the judge reads)
### Analysis answers a different question than the task specifies (wrong inputs, wrong chart/aggregation, missing required artifacts)
- **Applies when**: `task` -- the task names the required computation (e.g., a specific aggregate per category for a ranked subset), the required visualization type, and the exact output files/settings to produce.
- **Pattern**: The script loads whichever file happens to be present in the data directory and computes some plausible-looking but unrelated statistic and plot type; the requested grouping/ranking/aggregation is never performed, the required chart form is not produced, and only a subset of the mandated output artifacts is written. The final answer confidently narrates the substituted analysis instead of flagging the mismatch.
- **Detection procedure**:
  1. From the task text, list the required elements: input entities/columns referenced, the aggregation/statistic, any subsetting/ranking rule, the chart type/orientation, and every output file name plus the config file whose keys must be honored.
  2. Read the script and map each required element to the code that implements it; check that the loaded file actually contains the referenced fields and that the config keys used match the config file's real contents.
  3. Check the write calls against the required artifact list (all files, correct names/locations/formats), and confirm the plot call matches the requested chart type.
  4. Read the final answer: does it describe the requested quantities and entities, or a different analysis? Does it acknowledge any data/column mismatch?
- **Discriminator**: A real violation is when the computed quantity, grouping, chart type, or produced files cannot be mapped onto the task's stated requirements (or the columns used don't exist in the described schema). A look-alike that is fine is a faithful implementation that merely differs in cosmetic choices (colors, figure size, label wording) while the aggregation, subsetting, chart type, and all required artifacts are present.
- **Consequence**: Every graded artifact fails — numeric outputs and plot metadata don't exist or contain values from an unrelated computation — so the score is 0 despite a confident "task completed" report.
796Filtered statistic reported without evidence that the filter (and the underlying derived quantity) actually took effecttaskinfiagent-dabench
Applies when
task -- the task asks for a summary statistic computed on a derived per-entity quantity after applying a stated exclusion rule (threshold, z-score, IQR, date/unit filter).
Pattern
The script builds the derived quantity with an unverified aggregation (wrong grouping key, wrong unit, duplicated or unaggregated rows, missing/NaN handling) and/or applies the exclusion rule in a way that removes nothing or the wrong rows, then reports the resulting numbers with no before/after comparison, no count of removed points, and no plausibility check.
Detection procedure
  1. Read the task and write down the exact definition of the derived quantity (one value per what entity? in what units?) and the exact exclusion rule (which statistic, which threshold, computed on which set).
  2. In the script, check that the derived quantity is produced by an explicit aggregation/grouping over the correct key with duplicates and missing values handled, and that its length equals the expected number of entities.
  3. Check that the script prints n_before, n_removed, and both pre- and post-filter mean/std, and that the reported answer is taken from the post-filter arrays with the required rounding/format.
  4. Sanity-check the reported numbers: after removing upper-tail outliers the mean and std must be strictly smaller than the pre-filter values and n_removed must be > 0 if any |z| > threshold existed; flag if the answer is indistinguishable from the unfiltered statistic or if the std is of the same magnitude as the mean's outlier-driven inflation.
Discriminator
A genuine violation is when the script never reports removal counts / pre-vs-post values, or the answer equals (or is implausibly close to) the unfiltered statistic, or the entity count of the derived series is unverified; it is fine if the script prints these diagnostics and the filter legitimately removes few or zero points on a truly clean, correctly aggregated series.
Consequence
The reported mean and standard deviation are those of an unfiltered or wrongly aggregated series—inflated by orders of tail magnitude—so both graded values miss the expected numbers and the answer scores 0.
id 19271c360114 · mined from infiagent-dabench dabench-619@s15
raw text (what the judge reads)
### Filtered statistic reported without evidence that the filter (and the underlying derived quantity) actually took effect
- **Applies when**: `task` -- the task asks for a summary statistic computed on a derived per-entity quantity after applying a stated exclusion rule (threshold, z-score, IQR, date/unit filter).
- **Pattern**: The script builds the derived quantity with an unverified aggregation (wrong grouping key, wrong unit, duplicated or unaggregated rows, missing/NaN handling) and/or applies the exclusion rule in a way that removes nothing or the wrong rows, then reports the resulting numbers with no before/after comparison, no count of removed points, and no plausibility check.
- **Detection procedure**:
  1. Read the task and write down the exact definition of the derived quantity (one value per what entity? in what units?) and the exact exclusion rule (which statistic, which threshold, computed on which set).
  2. In the script, check that the derived quantity is produced by an explicit aggregation/grouping over the correct key with duplicates and missing values handled, and that its length equals the expected number of entities.
  3. Check that the script prints n_before, n_removed, and both pre- and post-filter mean/std, and that the reported answer is taken from the post-filter arrays with the required rounding/format.
  4. Sanity-check the reported numbers: after removing upper-tail outliers the mean and std must be strictly smaller than the pre-filter values and n_removed must be > 0 if any |z| > threshold existed; flag if the answer is indistinguishable from the unfiltered statistic or if the std is of the same magnitude as the mean's outlier-driven inflation.
- **Discriminator**: A genuine violation is when the script never reports removal counts / pre-vs-post values, or the answer equals (or is implausibly close to) the unfiltered statistic, or the entity count of the derived series is unverified; it is fine if the script prints these diagnostics and the filter legitimately removes few or zero points on a truly clean, correctly aggregated series.
- **Consequence**: The reported mean and standard deviation are those of an unfiltered or wrongly aggregated series—inflated by orders of tail magnitude—so both graded values miss the expected numbers and the answer scores 0.
797Analysis run on fabricated/synthetic data instead of the provided datasettaskda-code
Applies when
task -- the task references a supplied dataset (README, data directory, or file) and the scripts must load it to compute the reported numbers.
Pattern
The script fails to locate or verify the real input file and silently falls back to generating random/sample data (e.g., np.random.* with invented column names and categories), then reports statistics from that simulated data as the final answer.
Detection procedure
1. Read the task/README and note that the numbers must derive from the actual provided data. 2. Scan the scripts for any data-creation code (np.random, hardcoded DataFrame(...), "sample data", seed) and for whether a real file path is loaded unconditionally with a hard failure if missing. 3. Check whether column/category names used in the analysis were confirmed from the loaded file, or merely assumed; check whether the final reporting script re-loads the real data at all. 4. Compare the answer's group counts/values against what the real data's documented size and schema would imply.
Discriminator
A real violation is when the reported numbers come from data the script invented (or from an unvalidated assumed schema); it is fine if synthetic data is only used in a separate self-test/demo and the reported answer is computed from the actually loaded file, with the load path erroring out (not silently substituting) when the file is absent.
Consequence
The reported statistics are essentially random numbers unrelated to the real data, so every value (and often the number/order of groups) mismatches the expected result file, scoring 0.
id 0881267bbe1b · mined from da-code dacode-data-sa-061@s15
raw text (what the judge reads)
### Analysis run on fabricated/synthetic data instead of the provided dataset
- **Applies when**: `task` -- the task references a supplied dataset (README, data directory, or file) and the scripts must load it to compute the reported numbers.
- **Pattern**: The script fails to locate or verify the real input file and silently falls back to generating random/sample data (e.g., `np.random.*` with invented column names and categories), then reports statistics from that simulated data as the final answer.
- **Detection procedure**: 1. Read the task/README and note that the numbers must derive from the actual provided data. 2. Scan the scripts for any data-creation code (`np.random`, hardcoded `DataFrame(...)`, "sample data", `seed`) and for whether a real file path is loaded unconditionally with a hard failure if missing. 3. Check whether column/category names used in the analysis were confirmed from the loaded file, or merely assumed; check whether the final reporting script re-loads the real data at all. 4. Compare the answer's group counts/values against what the real data's documented size and schema would imply.
- **Discriminator**: A real violation is when the reported numbers come from data the script invented (or from an unvalidated assumed schema); it is fine if synthetic data is only used in a separate self-test/demo and the reported answer is computed from the actually loaded file, with the load path erroring out (not silently substituting) when the file is absent.
- **Consequence**: The reported statistics are essentially random numbers unrelated to the real data, so every value (and often the number/order of groups) mismatches the expected result file, scoring 0.
798Model accepted on training-set fit alone, with no held-out validation or baseline comparisontaskda-code
Applies when
task -- a script trains a predictive model on one file and writes predictions for a separate unlabeled file that will be graded against hidden ground-truth targets.
Pattern
The script fits one or more models on the full labeled data, reports R²/RMSE computed on those same training rows (an in-sample, optimistically biased number), and ships the predictions without any hold-out/cross-validation estimate, without comparing against a trivial baseline (predicting the mean/median), and without checking whether stronger available signals (identifier, metadata, temporal, or categorical columns, or an exact overlap between the test rows and the labeled source) were left unused.
Detection procedure
  1. Read the task to see that scoring is on unseen targets, so an out-of-sample accuracy estimate is required to know if the submission is usable.
  2. In the script, locate every call that computes a metric and check which data it is given; flag if the arrays passed are the same ones used in fit, and confirm no train/validation split or cross-validation object is actually used (imported-but-unused split helpers count as absent).
  3. Check whether the reported error is compared with a naive constant-prediction baseline, and whether the feature list omits columns present in both files that plausibly carry more signal than the ones chosen (including keys that would allow directly matching test rows to labeled rows).
  4. Inspect the output step: confirm the submission is only sanity-checked for range/shape, with no evidence-based claim about accuracy.
Discriminator
A real violation is when every reported number is in-sample and no split/CV exists anywhere, so the agent cannot distinguish a model from a mean-predictor; it is fine if the script additionally prints training fit but bases model selection on a hold-out/CV score, or if the target is deterministically recoverable and the script verifies that recovery on known rows.
Consequence
The submitted file has the right shape and column name but near-zero (or negative) out-of-sample skill, so the grader's accuracy/error threshold on the predictions fails while the agent believes the run succeeded.
id 95f073000074 · mined from da-code dacode-ml-regression-004@s15
raw text (what the judge reads)
### Model accepted on training-set fit alone, with no held-out validation or baseline comparison
- **Applies when**: `task` -- a script trains a predictive model on one file and writes predictions for a separate unlabeled file that will be graded against hidden ground-truth targets.
- **Pattern**: The script fits one or more models on the full labeled data, reports R²/RMSE computed on those same training rows (an in-sample, optimistically biased number), and ships the predictions without any hold-out/cross-validation estimate, without comparing against a trivial baseline (predicting the mean/median), and without checking whether stronger available signals (identifier, metadata, temporal, or categorical columns, or an exact overlap between the test rows and the labeled source) were left unused.
- **Detection procedure**:
  1. Read the task to see that scoring is on unseen targets, so an out-of-sample accuracy estimate is required to know if the submission is usable.
  2. In the script, locate every call that computes a metric and check which data it is given; flag if the arrays passed are the same ones used in `fit`, and confirm no train/validation split or cross-validation object is actually used (imported-but-unused split helpers count as absent).
  3. Check whether the reported error is compared with a naive constant-prediction baseline, and whether the feature list omits columns present in both files that plausibly carry more signal than the ones chosen (including keys that would allow directly matching test rows to labeled rows).
  4. Inspect the output step: confirm the submission is only sanity-checked for range/shape, with no evidence-based claim about accuracy.
- **Discriminator**: A real violation is when *every* reported number is in-sample and no split/CV exists anywhere, so the agent cannot distinguish a model from a mean-predictor; it is fine if the script additionally prints training fit but bases model selection on a hold-out/CV score, or if the target is deterministically recoverable and the script verifies that recovery on known rows.
- **Consequence**: The submitted file has the right shape and column name but near-zero (or negative) out-of-sample skill, so the grader's accuracy/error threshold on the predictions fails while the agent believes the run succeeded.
799Blindly selecting a model/hyperparameter by argmax of one internal metric, without sanity-checking the resulting solutiontaskda-code
Applies when
task -- the task asks for an "appropriate" number of groups/components (or any unsupervised hyperparameter) and the script picks it automatically by maximizing a single score over a range.
Pattern
The script sweeps k (or similar), takes argmax of one criterion (e.g. silhouette) on unscaled-outlier-heavy standardized data, and accepts the winner even though the resulting partition is degenerate — singleton or near-singleton groups that merely isolate outliers — with no cross-check against a second criterion (elbow/inertia, gap, BIC, stability), no outlier/skew handling, and no comparison of the runner-up solutions.
Detection procedure
  1. Read the task for what "appropriate" means and whether any interpretable grouping is implied by the stated goal (e.g. ranking/segmenting entities into usable tiers).
  2. In the script, check whether the chosen hyperparameter comes from a single automated argmax and whether any validation of the resulting partition (group sizes, separation, robustness to seed/scaling, second criterion) is performed before writing output.
  3. In the reported output, inspect the group-size distribution and the absolute value of the score: groups of size 1–3 out of hundreds, or a low absolute score (e.g. silhouette ≈ 0.3) barely above neighbors, indicate the metric is rewarding outlier isolation, not structure.
  4. Confirm the script never revisits the choice (no skew/outlier transform, no re-run at a nearby k) after seeing such a distribution.
Discriminator
A real violation is accepting a partition with degenerate/outlier-only groups or a score that is essentially flat across candidates, with no second line of evidence. It is fine if the argmax is corroborated (elbow, stability, or a clearly dominant score) and all groups are substantively populated and interpretable — even if some group is small for genuine domain reasons that the agent states and defends.
Consequence
The saved label column encodes an unstable, outlier-driven partition whose cluster count and assignments disagree with the reference solution, so the output-file check fails even though the file format and row count look correct.
id a393a8735daa · mined from da-code dacode-ml-cluster-013@s15
raw text (what the judge reads)
### Blindly selecting a model/hyperparameter by argmax of one internal metric, without sanity-checking the resulting solution
- **Applies when**: `task` -- the task asks for an "appropriate" number of groups/components (or any unsupervised hyperparameter) and the script picks it automatically by maximizing a single score over a range.
- **Pattern**: The script sweeps k (or similar), takes `argmax` of one criterion (e.g. silhouette) on unscaled-outlier-heavy standardized data, and accepts the winner even though the resulting partition is degenerate — singleton or near-singleton groups that merely isolate outliers — with no cross-check against a second criterion (elbow/inertia, gap, BIC, stability), no outlier/skew handling, and no comparison of the runner-up solutions.
- **Detection procedure**:
  1. Read the task for what "appropriate" means and whether any interpretable grouping is implied by the stated goal (e.g. ranking/segmenting entities into usable tiers).
  2. In the script, check whether the chosen hyperparameter comes from a single automated `argmax` and whether any validation of the resulting partition (group sizes, separation, robustness to seed/scaling, second criterion) is performed before writing output.
  3. In the reported output, inspect the group-size distribution and the absolute value of the score: groups of size 1–3 out of hundreds, or a low absolute score (e.g. silhouette ≈ 0.3) barely above neighbors, indicate the metric is rewarding outlier isolation, not structure.
  4. Confirm the script never revisits the choice (no skew/outlier transform, no re-run at a nearby k) after seeing such a distribution.
- **Discriminator**: A real violation is accepting a partition with degenerate/outlier-only groups or a score that is essentially flat across candidates, with no second line of evidence. It is fine if the argmax is corroborated (elbow, stability, or a clearly dominant score) and all groups are substantively populated and interpretable — even if some group is small for genuine domain reasons that the agent states and defends.
- **Consequence**: The saved label column encodes an unstable, outlier-driven partition whose cluster count and assignments disagree with the reference solution, so the output-file check fails even though the file format and row count look correct.
800Statistic computed on a different subset/axis than the task specifies (and on only part of the data)taskinfiagent-dabench
Applies when
task -- the prompt asks for a per-entity statistic restricted to a stated slice (a given year, group, segment) and a ranking of entities over "all" of them in the dataset.
Pattern
The script finds the literal slice too small/degenerate to support the requested statistic, silently substitutes a different aggregation (e.g., collapsing over the other axis, or over all periods instead of the named one), and/or loads only one of the several files/partitions that together constitute "the dataset", then reports the argmax of that substituted computation as the answer.
Detection procedure
  1. From the task, write down exactly three things: the unit of analysis (per entity), the filter/slice that must be applied, and the statistic definition/options required.
  2. In the scripts, locate the array actually passed to the statistic call and check which axis and which rows/columns it spans; confirm it equals "entity × stated slice" and not some other combination, and confirm the required option/definition flag is set as stated.
  3. Check data loading: verify every file/partition that contributes entities to "all countries/entities" is read and concatenated, not just the first one found.
  4. Inspect whether the script's own comments/prints admit the literal reading was abandoned ("this doesn't tell us...", "maybe the task means...") without any attempt to reconcile — that admission plus an unchanged final answer is the flag.
Discriminator
A real violation is when the computed quantity would change the ranking if the stated slice/axis/scope were used, and no justification or cross-check was done; a look-alike that is fine is a script that computes the literal quantity first, documents why it is degenerate, and then reports a reconciled result consistent with the stated slice and full data scope (e.g., derives a per-entity value that is still tied to the named slice).
Consequence
The reported argmax entity comes from a different distribution than the graded one, so the single expected string mismatches and the grader scores 0/1.
id c9e046c32d49 · mined from infiagent-dabench dabench-252@s15
raw text (what the judge reads)
### Statistic computed on a different subset/axis than the task specifies (and on only part of the data)
- **Applies when**: `task` -- the prompt asks for a per-entity statistic restricted to a stated slice (a given year, group, segment) and a ranking of entities over "all" of them in the dataset.
- **Pattern**: The script finds the literal slice too small/degenerate to support the requested statistic, silently substitutes a different aggregation (e.g., collapsing over the other axis, or over all periods instead of the named one), and/or loads only one of the several files/partitions that together constitute "the dataset", then reports the argmax of that substituted computation as the answer.
- **Detection procedure**:
  1. From the task, write down exactly three things: the unit of analysis (per entity), the filter/slice that must be applied, and the statistic definition/options required.
  2. In the scripts, locate the array actually passed to the statistic call and check which axis and which rows/columns it spans; confirm it equals "entity × stated slice" and not some other combination, and confirm the required option/definition flag is set as stated.
  3. Check data loading: verify every file/partition that contributes entities to "all countries/entities" is read and concatenated, not just the first one found.
  4. Inspect whether the script's own comments/prints admit the literal reading was abandoned ("this doesn't tell us...", "maybe the task means...") without any attempt to reconcile — that admission plus an unchanged final answer is the flag.
- **Discriminator**: A real violation is when the computed quantity would change the ranking if the stated slice/axis/scope were used, and no justification or cross-check was done; a look-alike that is fine is a script that computes the literal quantity first, documents why it is degenerate, and then reports a reconciled result consistent with the stated slice and full data scope (e.g., derives a per-entity value that is still tied to the named slice).
- **Consequence**: The reported argmax entity comes from a different distribution than the graded one, so the single expected string mismatches and the grader scores 0/1.
801Answer reports the payload instead of the requested identifiertaskinfiagent-dabench
Applies when
task -- the answer template asks for a name/label (e.g., of a created column, chosen model, selected feature) rather than the underlying data values.
Pattern
The agent computes the artifact correctly but fills the answer slot with the full vector/table of computed values (often truncated), instead of the short identifier the format string literally requests.
Detection procedure
1. Read the answer format spec and decide what token type it expects — a single name/label, a scalar, or a list. 2. Note any gloss like "where X refers to the newly created column…" — this describes the referent, not the required content. 3. Inspect the submitted string: does it contain a single short token, or a long serialized array/dict? 4. Check whether the submission is complete and well-formed (unclosed brackets, trailing commas, elision indicate an oversized payload).
Discriminator
A real violation is when the task's slot semantics are "name the thing" but the answer dumps its contents (or vice versa); it is fine if the task explicitly asks for the list of per-row values or if the format example itself shows an array.
Consequence
The grader's exact-match on the expected identifier fails (0/1), even though the underlying computation and script logic were correct.
id 6d87474dffa5 · mined from infiagent-dabench dabench-741@s15
raw text (what the judge reads)
### Answer reports the payload instead of the requested identifier
- **Applies when**: `task` -- the answer template asks for a name/label (e.g., of a created column, chosen model, selected feature) rather than the underlying data values.
- **Pattern**: The agent computes the artifact correctly but fills the answer slot with the full vector/table of computed values (often truncated), instead of the short identifier the format string literally requests.
- **Detection procedure**: 1. Read the answer format spec and decide what token type it expects — a single name/label, a scalar, or a list. 2. Note any gloss like "where X refers to the newly created column…" — this describes the referent, not the required content. 3. Inspect the submitted string: does it contain a single short token, or a long serialized array/dict? 4. Check whether the submission is complete and well-formed (unclosed brackets, trailing commas, elision indicate an oversized payload).
- **Discriminator**: A real violation is when the task's slot semantics are "name the thing" but the answer dumps its contents (or vice versa); it is fine if the task explicitly asks for the list of per-row values or if the format example itself shows an array.
- **Consequence**: The grader's exact-match on the expected identifier fails (0/1), even though the underlying computation and script logic were correct.
802Losing identifying precision when formatting a key/label answertaskinfiagent-dabench
Applies when
task -- the answer includes an identifier (a date, key, index, or label) that must point to one specific row/record, and the script reformats or truncates it before reporting.
Pattern
The script finds the correct record but emits a coarsened version of its identifier (e.g., strftime/round/substring/regex that drops components), so the reported value no longer uniquely designates the record found, even though the accompanying numeric result was computed from the full-precision record.
Detection procedure
  1. In the task, note the granularity of the identifier present in the source data and the granularity implied by the requested output template.
  2. In the scripts, find every formatting/casting step applied to that identifier before printing and check whether it discards components (time parts, sub-fields, decimals, prefixes).
  3. Compare the printed identifier against the record actually used for the downstream computation: does the printed string map back to exactly one record in the data?
  4. If the format template is coarser than the data's natural granularity, flag it and require the answer to at least report the full identifier (or explicitly reconcile the ambiguity), rather than silently truncating.
Discriminator
A real violation is when the coarsened identifier matches many rows in the dataset or drops information the analysis itself relied on; it is fine when the data's own granularity equals the requested format (e.g., the records are genuinely monthly/aggregated), or when the truncation is a lossless re-encoding of the same value.
Consequence
The numeric part of the answer can be graded correct while the identifier check fails on exact-string comparison, yielding a partial-credit/incorrect verdict.
id 7b651c1b4c75 · mined from infiagent-dabench dabench-572@s15
raw text (what the judge reads)
### Losing identifying precision when formatting a key/label answer
- **Applies when**: `task` -- the answer includes an identifier (a date, key, index, or label) that must point to one specific row/record, and the script reformats or truncates it before reporting.
- **Pattern**: The script finds the correct record but emits a coarsened version of its identifier (e.g., strftime/round/substring/regex that drops components), so the reported value no longer uniquely designates the record found, even though the accompanying numeric result was computed from the full-precision record.
- **Detection procedure**:
  1. In the task, note the granularity of the identifier present in the source data and the granularity implied by the requested output template.
  2. In the scripts, find every formatting/casting step applied to that identifier before printing and check whether it discards components (time parts, sub-fields, decimals, prefixes).
  3. Compare the printed identifier against the record actually used for the downstream computation: does the printed string map back to exactly one record in the data?
  4. If the format template is coarser than the data's natural granularity, flag it and require the answer to at least report the full identifier (or explicitly reconcile the ambiguity), rather than silently truncating.
- **Discriminator**: A real violation is when the coarsened identifier matches many rows in the dataset or drops information the analysis itself relied on; it is fine when the data's own granularity equals the requested format (e.g., the records are genuinely monthly/aggregated), or when the truncation is a lossless re-encoding of the same value.
- **Consequence**: The numeric part of the answer can be graded correct while the identifier check fails on exact-string comparison, yielding a partial-credit/incorrect verdict.
803Prediction file not validated against the provided template and test settaskda-code
Applies when
task -- the task asks for an output file of per-row predictions written in the format of a supplied sample/template file.
Pattern
The agent produces the output file (and often keeps no runnable script), but never verifies the artifact against the template: header spelling/case, single-vs-multiple columns, presence of an index column, number of rows equal to the number of test rows, row order identical to the test file, and value domain (e.g., integer class labels vs. probabilities vs. strings). It reports "done" based on the model step succeeding rather than on the file's contents.
Detection procedure
  1. Read the task for the exact required filename, column name, and the template file it must mimic; note the expected number of prediction rows from the test data description.
  2. Read the scripts for the write step: check that the frame written has exactly the template's columns, index=False (or index if the template has one), the same dtype/label encoding as the template, and that predictions were generated from the untouched test file in its original row order (no dropping/filtering/shuffling/deduplication of test rows).
  3. Check that a post-write verification exists (re-read the file; assert shape, header equality with the template, and value set) — or reproduce it yourself on the delivered file.
  4. If no script or no artifact inspection is available, treat the claim as unverified: the answer cannot be shown to satisfy the format.
Discriminator
A real violation is a file whose header, row count, row order, or value domain provably can differ from the template/test set, or that is unverifiable because nothing checks it. A look-alike that is fine is a file whose writing code demonstrably matches the template and is accompanied by explicit shape/header/value assertions, even if the model itself is simple or the accuracy is mediocre.
Consequence
The grader reads the expected file and finds it missing, mis-named, mis-headered, or with the wrong row count/label values, marking the output WRONG regardless of predictive quality (0 checks passed).
id 9c645483e58c · mined from da-code dacode-ml-binary-013@s15
raw text (what the judge reads)
### Prediction file not validated against the provided template and test set
- **Applies when**: `task` -- the task asks for an output file of per-row predictions written in the format of a supplied sample/template file.
- **Pattern**: The agent produces the output file (and often keeps no runnable script), but never verifies the artifact against the template: header spelling/case, single-vs-multiple columns, presence of an index column, number of rows equal to the number of test rows, row order identical to the test file, and value domain (e.g., integer class labels vs. probabilities vs. strings). It reports "done" based on the model step succeeding rather than on the file's contents.
- **Detection procedure**:
  1. Read the task for the exact required filename, column name, and the template file it must mimic; note the expected number of prediction rows from the test data description.
  2. Read the scripts for the write step: check that the frame written has exactly the template's columns, `index=False` (or index if the template has one), the same dtype/label encoding as the template, and that predictions were generated from the untouched test file in its original row order (no dropping/filtering/shuffling/deduplication of test rows).
  3. Check that a post-write verification exists (re-read the file; assert shape, header equality with the template, and value set) — or reproduce it yourself on the delivered file.
  4. If no script or no artifact inspection is available, treat the claim as unverified: the answer cannot be shown to satisfy the format.
- **Discriminator**: A real violation is a file whose header, row count, row order, or value domain provably can differ from the template/test set, or that is unverifiable because nothing checks it. A look-alike that is fine is a file whose writing code demonstrably matches the template and is accompanied by explicit shape/header/value assertions, even if the model itself is simple or the accuracy is mediocre.
- **Consequence**: The grader reads the expected file and finds it missing, mis-named, mis-headered, or with the wrong row count/label values, marking the output WRONG regardless of predictive quality (0 checks passed).
804Answer emitted in a near-miss format (extra quoting/delimiters) with no reproducible scripttaskinfiagent-dabench
Applies when
task -- the task prescribes a literal answer template (e.g. @key[list_of_values]) and the agent must emit that string, ideally produced/printed by a saved script.
Pattern
The attempt computes the right underlying values but serializes them with decorations the template never asked for (Python-style quotes, brackets inside brackets, trailing punctuation, index labels, units, differing separators or ordering), and/or hand-types the final line instead of printing it from a saved script, so the exact submitted string cannot be checked or regenerated.
Detection procedure
  1. Read the task's answer-format spec and write down the exact literal skeleton, noting whether values are quoted, how they are separated, whether ordering/casing matters, and how names must be spelled (source labels vs. paraphrases).
  2. Locate in the scripts the code that constructs and prints the final answer string; if no script/output artifact exists, treat the answer as unverified.
  3. Character-by-character diff the submitted string against the skeleton from step 1: flag any added quoting, nested brackets, whitespace/separator deviation, renamed keys, or values not taken verbatim from the source labels.
  4. Confirm the printed values are the requested final quantities (not intermediates) and that re-running the script would reproduce the same string.
Discriminator
A real violation is a formatting/serialization deviation from the literal template or an unreproducible hand-written answer, even when the substantive values are correct; a look-alike that is fine is cosmetic freedom the task explicitly leaves open (e.g. spec shows a generic placeholder and the grader accepts any consistent delimiter) or a formatting choice that the script demonstrably derives from the spec.
Consequence
The grader parses the answer literally, fails to match the expected key/value tokens, and marks the check WRONG/MISSING despite substantively correct analysis — and with no saved script, the mismatch cannot be diagnosed or fixed.
id bfc8c7935431 · mined from infiagent-dabench dabench-254@s15
raw text (what the judge reads)
### Answer emitted in a near-miss format (extra quoting/delimiters) with no reproducible script
- **Applies when**: `task` -- the task prescribes a literal answer template (e.g. `@key[list_of_values]`) and the agent must emit that string, ideally produced/printed by a saved script.
- **Pattern**: The attempt computes the right underlying values but serializes them with decorations the template never asked for (Python-style quotes, brackets inside brackets, trailing punctuation, index labels, units, differing separators or ordering), and/or hand-types the final line instead of printing it from a saved script, so the exact submitted string cannot be checked or regenerated.
- **Detection procedure**:
  1. Read the task's answer-format spec and write down the exact literal skeleton, noting whether values are quoted, how they are separated, whether ordering/casing matters, and how names must be spelled (source labels vs. paraphrases).
  2. Locate in the scripts the code that constructs and prints the final answer string; if no script/output artifact exists, treat the answer as unverified.
  3. Character-by-character diff the submitted string against the skeleton from step 1: flag any added quoting, nested brackets, whitespace/separator deviation, renamed keys, or values not taken verbatim from the source labels.
  4. Confirm the printed values are the requested final quantities (not intermediates) and that re-running the script would reproduce the same string.
- **Discriminator**: A real violation is a formatting/serialization deviation from the literal template or an unreproducible hand-written answer, even when the substantive values are correct; a look-alike that is fine is cosmetic freedom the task explicitly leaves open (e.g. spec shows a generic placeholder and the grader accepts any consistent delimiter) or a formatting choice that the script demonstrably derives from the spec.
- **Consequence**: The grader parses the answer literally, fails to match the expected key/value tokens, and marks the check WRONG/MISSING despite substantively correct analysis — and with no saved script, the mismatch cannot be diagnosed or fixed.
805Ad-hoc, unjustified metric definition computed over only part of the relevant rowstaskda-code
Applies when
task -- the task asks to visualize/report a "performance"/"score"/ranking quantity whose exact definition is not spelled out in the prompt, and the entities being ranked appear in more than one column/role of the raw data.
Pattern
The agent invents a formula out of convenience (e.g., a weighted sum of record counts or of an unrelated flag column) instead of a domain-standard, defensible definition, and aggregates using a single grouping column so every record where the entity appears in the other role is silently dropped; the output artifacts and their ordering are then reverse-engineered to match whatever labels appear in the provided config file rather than derived from the computation.
Detection procedure
  1. Read the task and any config/spec file: note which quantity is implied by the axis label/title and which output artifacts (chart plus any serialized data/array files) are expected.
  2. In the script, locate the line defining the ranked quantity; check whether it corresponds to a recognized definition of the concept named on the axis, or is an arbitrary combination (counts, flags, weights like +2* with no stated basis).
  3. Check the groupby/aggregation: confirm all rows in which an entity participates are included (both/all role columns concatenated), not just one column; confirm the filter window matches the stated period.
  4. Compare the script's own printed ranking against the ordering/labels in the config; if the agent had to subset or re-order to the config's labels because its ranking disagreed, that is evidence the metric is wrong. Also verify every required output file is written.
Discriminator
A real violation is when the computed quantity would change materially under the standard definition (e.g., wins/points/goal difference over all appearances) or when the entity's rows in the second role are excluded; it is fine if the task or config explicitly fixes the formula, or if the entity genuinely appears in only one column and the agent's ranking reproduces the config ordering without manual re-selection.
Consequence
The plotted values, the saved numeric array, and the serialized plot metadata all differ from the reference, so every artifact check fails even though the figure looks stylistically correct.
id 18bb66179d73 · mined from da-code dacode-plot-bar-006@s15
raw text (what the judge reads)
### Ad-hoc, unjustified metric definition computed over only part of the relevant rows
- **Applies when**: `task` -- the task asks to visualize/report a "performance"/"score"/ranking quantity whose exact definition is not spelled out in the prompt, and the entities being ranked appear in more than one column/role of the raw data.
- **Pattern**: The agent invents a formula out of convenience (e.g., a weighted sum of record counts or of an unrelated flag column) instead of a domain-standard, defensible definition, and aggregates using a single grouping column so every record where the entity appears in the other role is silently dropped; the output artifacts and their ordering are then reverse-engineered to match whatever labels appear in the provided config file rather than derived from the computation.
- **Detection procedure**:
  1. Read the task and any config/spec file: note which quantity is implied by the axis label/title and which output artifacts (chart plus any serialized data/array files) are expected.
  2. In the script, locate the line defining the ranked quantity; check whether it corresponds to a recognized definition of the concept named on the axis, or is an arbitrary combination (counts, flags, weights like `+2*` with no stated basis).
  3. Check the `groupby`/aggregation: confirm all rows in which an entity participates are included (both/all role columns concatenated), not just one column; confirm the filter window matches the stated period.
  4. Compare the script's own printed ranking against the ordering/labels in the config; if the agent had to subset or re-order to the config's labels because its ranking disagreed, that is evidence the metric is wrong. Also verify every required output file is written.
- **Discriminator**: A real violation is when the computed quantity would change materially under the standard definition (e.g., wins/points/goal difference over all appearances) or when the entity's rows in the second role are excluded; it is fine if the task or config explicitly fixes the formula, or if the entity genuinely appears in only one column and the agent's ranking reproduces the config ordering without manual re-selection.
- **Consequence**: The plotted values, the saved numeric array, and the serialized plot metadata all differ from the reference, so every artifact check fails even though the figure looks stylistically correct.
806Verification that only re-runs the same pipeline instead of testing the assumptions behind a threshold statistictaskinfiagent-dabench
Applies when
task -- The task asks for a count/flag derived from a threshold on a computed statistic (z-score, quantile, ratio) and the scripts produce that count from a single fixed pipeline, then "verify" it by re-executing the identical computation.
Pattern
The agent picks one set of implicit choices (source file, column selection, dtype/missing-value handling, sentinel values, std with ddof=0 vs ddof=1, whole-column vs per-group scaling) and never varies any of them; its "verification" scripts recompute the same numbers, so a wrong assumption is copied into the final answer. Extreme results (a large count of flagged rows for a threshold that should flag very few, or zero) are reported without a plausibility check against what the threshold implies.
Detection procedure
  1. Read the task and list the free choices it does not pin down (which file/split, which column if names are ambiguous, how NaNs/placeholder codes are treated, sample vs population standard deviation, global vs grouped statistics).
  2. Read the scripts and check whether any of those choices is (a) justified from an inspection of the data (dtypes, unique values, value ranges, printed distribution) and (b) tested against at least one plausible alternative.
  3. Check whether the reported count is compared to an expectation implied by the threshold and observed distribution shape (e.g., how many rows a 3-sigma cut can plausibly flag given n and the min/max) rather than accepted as-is.
  4. If all "verification" scripts reproduce the same code path and yield the same number, and no alternative assumption or plausibility check appears, flag the attempt.
Discriminator
A fine attempt inspects the raw column (dtype, missing/sentinel encodings, extremes), states why one convention is used, and shows the count is stable (or explains the difference) under a reasonable alternative; a violation is triple-running the same formula and treating identical output as confirmation, especially when the flagged fraction is implausible for the stated threshold.
Consequence
The single unvalidated assumption propagates into the reported integer, so the answer differs from the reference count (here, a large count where the reference is a very different value) and the grader marks the answer wrong despite the pipeline "verifying" itself.
id 0d0455bb587a · mined from infiagent-dabench dabench-361@s15
raw text (what the judge reads)
### Verification that only re-runs the same pipeline instead of testing the assumptions behind a threshold statistic
- **Applies when**: `task` -- The task asks for a count/flag derived from a threshold on a computed statistic (z-score, quantile, ratio) and the scripts produce that count from a single fixed pipeline, then "verify" it by re-executing the identical computation.
- **Pattern**: The agent picks one set of implicit choices (source file, column selection, dtype/missing-value handling, sentinel values, std with ddof=0 vs ddof=1, whole-column vs per-group scaling) and never varies any of them; its "verification" scripts recompute the same numbers, so a wrong assumption is copied into the final answer. Extreme results (a large count of flagged rows for a threshold that should flag very few, or zero) are reported without a plausibility check against what the threshold implies.
- **Detection procedure**:
  1. Read the task and list the free choices it does not pin down (which file/split, which column if names are ambiguous, how NaNs/placeholder codes are treated, sample vs population standard deviation, global vs grouped statistics).
  2. Read the scripts and check whether any of those choices is (a) justified from an inspection of the data (dtypes, unique values, value ranges, printed distribution) and (b) tested against at least one plausible alternative.
  3. Check whether the reported count is compared to an expectation implied by the threshold and observed distribution shape (e.g., how many rows a 3-sigma cut can plausibly flag given n and the min/max) rather than accepted as-is.
  4. If all "verification" scripts reproduce the same code path and yield the same number, and no alternative assumption or plausibility check appears, flag the attempt.
- **Discriminator**: A fine attempt inspects the raw column (dtype, missing/sentinel encodings, extremes), states why one convention is used, and shows the count is stable (or explains the difference) under a reasonable alternative; a violation is triple-running the same formula and treating identical output as confirmation, especially when the flagged fraction is implausible for the stated threshold.
- **Consequence**: The single unvalidated assumption propagates into the reported integer, so the answer differs from the reference count (here, a large count where the reference is a very different value) and the grader marks the answer wrong despite the pipeline "verifying" itself.
807Training-only auxiliary metadata merged in as model featurestaskda-code
Applies when
task -- the scripts join a supplementary file (extra labels, annotations, metadata) onto both the training and the scoring table before building the feature matrix.
Pattern
The attempt merges an auxiliary table that only has rows/keys for training records onto the test set as well. For test rows the joined columns come back entirely missing and are silently filled by an imputer (or label-encoded as "Unknown"), so the model is trained on informative — often target-derived — columns that are constant/meaningless at prediction time. No check is made that the merge actually matched test keys or that train and test feature distributions are comparable, and the submission's column names/row count are never verified against the provided sample submission.
Detection procedure
  1. From the task/README, list which files are available for the scoring rows; note any auxiliary file keyed only to training identifiers or containing label-adjacent information.
  2. In the scripts, find every merge/join into the feature pipeline and check whether the join key exists in the scoring table and whether the merge result is validated (matched-row count, how, NaN share per joined column).
  3. Check whether the joined columns are excluded from the feature list; if they are kept, check whether imputation would mask an all-missing column on the scoring side.
  4. Check the submission-building code against the provided sample file: header names, id column, and number of rows explicitly compared.
Discriminator
A real violation is a join whose columns are populated for training rows but unmatched/empty (or derived from the target) for scoring rows, with no post-merge validation. It is fine if the auxiliary file legitimately covers both sets with verified key overlap, or if the joined columns are dropped before fitting.
Consequence
Cross-validated scores look strong while test predictions are driven by imputed constants (or leak the label), producing poorly calibrated probabilities and a bad balanced-log-loss score; a mismatched header/row count makes the submission file fail validation outright.
id cc7a5daa7d96 · mined from da-code dacode-ml-competition-003@s15
raw text (what the judge reads)
### Training-only auxiliary metadata merged in as model features
- **Applies when**: `task` -- the scripts join a supplementary file (extra labels, annotations, metadata) onto both the training and the scoring table before building the feature matrix.
- **Pattern**: The attempt merges an auxiliary table that only has rows/keys for training records onto the test set as well. For test rows the joined columns come back entirely missing and are silently filled by an imputer (or label-encoded as "Unknown"), so the model is trained on informative — often target-derived — columns that are constant/meaningless at prediction time. No check is made that the merge actually matched test keys or that train and test feature distributions are comparable, and the submission's column names/row count are never verified against the provided sample submission.
- **Detection procedure**:
  1. From the task/README, list which files are available for the scoring rows; note any auxiliary file keyed only to training identifiers or containing label-adjacent information.
  2. In the scripts, find every merge/join into the feature pipeline and check whether the join key exists in the scoring table and whether the merge result is validated (matched-row count, `how`, NaN share per joined column).
  3. Check whether the joined columns are excluded from the feature list; if they are kept, check whether imputation would mask an all-missing column on the scoring side.
  4. Check the submission-building code against the provided sample file: header names, id column, and number of rows explicitly compared.
- **Discriminator**: A real violation is a join whose columns are populated for training rows but unmatched/empty (or derived from the target) for scoring rows, with no post-merge validation. It is fine if the auxiliary file legitimately covers both sets with verified key overlap, or if the joined columns are dropped before fitting.
- **Consequence**: Cross-validated scores look strong while test predictions are driven by imputed constants (or leak the label), producing poorly calibrated probabilities and a bad balanced-log-loss score; a mismatched header/row count makes the submission file fail validation outright.
808Fabricating the target labels instead of locating the real supervised signaltaskda-code
Applies when
task -- the task names a target variable to predict from a provided train/test split, and the scripts must find that label in the given files.
Pattern
The agent fails to locate the labeled column in the provided files, concludes it "was not in the dataset", and synthesizes surrogate labels from hand-written heuristics (e.g., thresholding another numeric field), then trains and reports accuracy against those invented labels — so the model never learns the actual quantity asked for, and the output rows/labels need not align with the required prediction file.
Detection procedure
1. Read the task and note the exact target name, the file the predictions must come from, and the required output file/column. 2. In the scripts, search for where the target is read: confirm it is loaded from a provided file (train split) rather than constructed by rules, mapping, or random assignment. 3. Check that predictions are produced for exactly the rows of the designated prediction file, in order, and written with the requested column name and row count. 4. Read the answer/summary for admissions like "target not present, created labels with domain knowledge", or a label distribution/row count that doesn't match the prediction file's size.
Discriminator
Legitimate cases derive features or auxiliary weak signals from other columns while still training and scoring on the provided ground-truth labels; a violation replaces the ground truth itself with a self-invented rule, or reports metrics computed against those invented labels. Also fine: the agent first exhausts a search of all provided files/joins for the label and documents where it found it.
Consequence
The submitted prediction file is scored against real held-out labels and matches near chance or is missing/misshapen entirely, so the file check fails despite a plausible-looking internal accuracy.
id 29291db442cb · mined from da-code dacode-ml-multi-003@s15
raw text (what the judge reads)
### Fabricating the target labels instead of locating the real supervised signal
- **Applies when**: `task` -- the task names a target variable to predict from a provided train/test split, and the scripts must find that label in the given files.
- **Pattern**: The agent fails to locate the labeled column in the provided files, concludes it "was not in the dataset", and synthesizes surrogate labels from hand-written heuristics (e.g., thresholding another numeric field), then trains and reports accuracy against those invented labels — so the model never learns the actual quantity asked for, and the output rows/labels need not align with the required prediction file.
- **Detection procedure**: 1. Read the task and note the exact target name, the file the predictions must come from, and the required output file/column. 2. In the scripts, search for where the target is read: confirm it is loaded from a provided file (train split) rather than constructed by rules, mapping, or random assignment. 3. Check that predictions are produced for exactly the rows of the designated prediction file, in order, and written with the requested column name and row count. 4. Read the answer/summary for admissions like "target not present, created labels with domain knowledge", or a label distribution/row count that doesn't match the prediction file's size.
- **Discriminator**: Legitimate cases derive *features* or auxiliary weak signals from other columns while still training and scoring on the provided ground-truth labels; a violation replaces the ground truth itself with a self-invented rule, or reports metrics computed against those invented labels. Also fine: the agent first exhausts a search of all provided files/joins for the label and documents where it found it.
- **Consequence**: The submitted prediction file is scored against real held-out labels and matches near chance or is missing/misshapen entirely, so the file check fails despite a plausible-looking internal accuracy.
809Output format assumed instead of read from the provided templatetaskda-code
Applies when
task -- the task points to a template/example/schema file (or explicitly states a required layout) that the deliverable must match.
Pattern
The scripts never load or inspect the template; the agent invents row/column labels, index origin (0- vs 1-based), date/label formatting, rounding, column order and header names from intuition, then asserts in the answer that "the format matches" without any comparison.
Detection procedure
  1. Read the task statement and note every stated formatting constraint and any referenced template/example artifact.
  2. Search the scripts for any read of that template (or of a schema description) and for any comparison of the produced object's shape, column names/dtypes, index values, and rounding against it.
  3. Inspect the hard-coded formatting decisions in the script (label strings, offsets, round(...), index=True/False, fill values) and ask whether each is derived from the template or merely assumed.
  4. Check whether the final answer claims format compliance while providing no evidence of a template diff or equality check.
Discriminator
A real violation is when no template read or structural comparison exists anywhere and formatting choices are unjustified guesses; it is fine if the script loads the template (or the task fully specifies the layout inline) and validates shape/headers/index/precision against it, even if it then reformats manually.
Consequence
The file's values may be conceptually reasonable but the header names, index labels, offsets or rounding differ from the expected file, so an exact/structural file comparison fails and the task scores 0.
id 7e8eb3d98b90 · mined from da-code dacode-dm-csv-044@s15
raw text (what the judge reads)
### Output format assumed instead of read from the provided template
- **Applies when**: `task` -- the task points to a template/example/schema file (or explicitly states a required layout) that the deliverable must match.
- **Pattern**: The scripts never load or inspect the template; the agent invents row/column labels, index origin (0- vs 1-based), date/label formatting, rounding, column order and header names from intuition, then asserts in the answer that "the format matches" without any comparison.
- **Detection procedure**:
  1. Read the task statement and note every stated formatting constraint and any referenced template/example artifact.
  2. Search the scripts for any read of that template (or of a schema description) and for any comparison of the produced object's shape, column names/dtypes, index values, and rounding against it.
  3. Inspect the hard-coded formatting decisions in the script (label strings, offsets, `round(...)`, `index=True/False`, fill values) and ask whether each is derived from the template or merely assumed.
  4. Check whether the final answer claims format compliance while providing no evidence of a template diff or equality check.
- **Discriminator**: A real violation is when no template read or structural comparison exists anywhere and formatting choices are unjustified guesses; it is fine if the script loads the template (or the task fully specifies the layout inline) and validates shape/headers/index/precision against it, even if it then reformats manually.
- **Consequence**: The file's values may be conceptually reasonable but the header names, index labels, offsets or rounding differ from the expected file, so an exact/structural file comparison fails and the task scores 0.
810Model selected and "validated" on in-sample training predictions onlytaskda-code
Applies when
task -- the scripts fit predictive models and must choose among models/hyperparameters/ensembles before producing held-out predictions scored by an external metric.
Pattern
The attempt computes the evaluation metric by predicting on the same rows used for fitting (model.fit(X_train, y_train); metric(y_train, model.predict(X_train))), imports cross-validation utilities but never uses them, and picks the "best"/"tuned" configuration by that resubstitution score — so a high reported score reflects memorization, not generalization, and nothing detects an overfit or mis-specified pipeline (e.g. scaling applied for one model but not others, hard-voting on an ordinal metric).
Detection procedure
  1. Read the task to identify the scoring metric and that scoring happens on unlabeled held-out data.
  2. In the scripts, locate every call that computes the metric and check whether the y_true/X passed are the exact arrays used in .fit(...); note whether any train/validation split, cross_val_score, or K-fold loop actually feeds the model-selection decision.
  3. Check whether hyperparameters/ensemble claims ("best found through tuning") are backed by any out-of-sample number in the run, or are asserted in comments only.
  4. Inspect the final answer/submission for signs consistent with unvalidated fitting (e.g. predicted class distribution or value range wildly unlike the training label distribution, degenerate single-value output) and for the absence of any reported held-out estimate.
Discriminator
A real violation is when no out-of-sample estimate exists anywhere and selection is driven by the train-fit score; it is fine to refit the chosen model on all labeled data for the final submission after cross-validated selection, and fine to also print a training score alongside a genuine CV/holdout score.
Consequence
The reported metric is inflated toward 1.0 while the actual held-out score is much lower (or the predictions collapse/misalign), so the graded submission fails the accuracy/agreement threshold and is marked wrong.
id 0cc17997c8da · mined from da-code dacode-ml-competition-006@s15
raw text (what the judge reads)
### Model selected and "validated" on in-sample training predictions only
- **Applies when**: `task` -- the scripts fit predictive models and must choose among models/hyperparameters/ensembles before producing held-out predictions scored by an external metric.
- **Pattern**: The attempt computes the evaluation metric by predicting on the same rows used for fitting (`model.fit(X_train, y_train); metric(y_train, model.predict(X_train))`), imports cross-validation utilities but never uses them, and picks the "best"/"tuned" configuration by that resubstitution score — so a high reported score reflects memorization, not generalization, and nothing detects an overfit or mis-specified pipeline (e.g. scaling applied for one model but not others, hard-voting on an ordinal metric).
- **Detection procedure**:
  1. Read the task to identify the scoring metric and that scoring happens on unlabeled held-out data.
  2. In the scripts, locate every call that computes the metric and check whether the `y_true`/`X` passed are the exact arrays used in `.fit(...)`; note whether any train/validation split, `cross_val_score`, or K-fold loop actually feeds the model-selection decision.
  3. Check whether hyperparameters/ensemble claims ("best found through tuning") are backed by any out-of-sample number in the run, or are asserted in comments only.
  4. Inspect the final answer/submission for signs consistent with unvalidated fitting (e.g. predicted class distribution or value range wildly unlike the training label distribution, degenerate single-value output) and for the absence of any reported held-out estimate.
- **Discriminator**: A real violation is when *no* out-of-sample estimate exists anywhere and selection is driven by the train-fit score; it is fine to refit the chosen model on all labeled data for the final submission *after* cross-validated selection, and fine to also print a training score alongside a genuine CV/holdout score.
- **Consequence**: The reported metric is inflated toward 1.0 while the actual held-out score is much lower (or the predictions collapse/misalign), so the graded submission fails the accuracy/agreement threshold and is marked wrong.
811Silent extra filtering / unverified definition of "missing" when splitting groupstaskinfiagent-dabench
Applies when
task -- the task asks to split rows into groups by whether a field is null/missing and compute per-group statistics or a test on another column.
Pattern
The script defines the groups with a single default rule (e.g., isna()/notna() on the raw parsed frame) and additionally drops rows with missing values in the measured column (or otherwise sub-filters) before computing the group means, without ever checking how missingness is actually encoded in the file (empty strings, sentinel text like "NA"/"None"/"-", whitespace, placeholder numbers) or verifying that the two group sizes reconstruct the full row count.
Detection procedure
  1. Read the task and note that group membership and the reported statistic must be defined over all rows, with no filtering beyond the stated split.
  2. In the script, locate the group-construction lines and check for extra operations (.dropna(), != 0, index/column selections, index_col=, dtype coercion) applied to the measured column or the frame after the split.
  3. Check whether the script prints/asserts (a) the number of rows in each group, (b) that these sum to the total rows, and (c) the raw distinct representations of "empty" values in the split column before deciding the null rule.
  4. Compare the reported group means/counts against these diagnostics; if no such diagnostic exists or counts don't sum to the total, flag it.
Discriminator
A real violation is when rows are removed or reclassified by a rule the task never asked for (or when the null rule is assumed rather than confirmed against the raw values), so group sizes/means change; it is fine if the extra filter provably removes nothing (verified count check) or the task itself specifies dropping those rows.
Consequence
The per-group means (and often the test statistic) shift away from the ground-truth values computed on the full groups, so the numeric answers fail exact-match checks even though the test procedure and answer format look correct.
id ec3f4f4cdf33 · mined from infiagent-dabench dabench-297@s15
raw text (what the judge reads)
### Silent extra filtering / unverified definition of "missing" when splitting groups
- **Applies when**: `task` -- the task asks to split rows into groups by whether a field is null/missing and compute per-group statistics or a test on another column.
- **Pattern**: The script defines the groups with a single default rule (e.g., `isna()`/`notna()` on the raw parsed frame) and additionally drops rows with missing values in the measured column (or otherwise sub-filters) before computing the group means, without ever checking how missingness is actually encoded in the file (empty strings, sentinel text like "NA"/"None"/"-", whitespace, placeholder numbers) or verifying that the two group sizes reconstruct the full row count.
- **Detection procedure**:
  1. Read the task and note that group membership and the reported statistic must be defined over *all* rows, with no filtering beyond the stated split.
  2. In the script, locate the group-construction lines and check for extra operations (`.dropna()`, `!= 0`, index/column selections, `index_col=`, dtype coercion) applied to the measured column or the frame after the split.
  3. Check whether the script prints/asserts (a) the number of rows in each group, (b) that these sum to the total rows, and (c) the raw distinct representations of "empty" values in the split column before deciding the null rule.
  4. Compare the reported group means/counts against these diagnostics; if no such diagnostic exists or counts don't sum to the total, flag it.
- **Discriminator**: A real violation is when rows are removed or reclassified by a rule the task never asked for (or when the null rule is assumed rather than confirmed against the raw values), so group sizes/means change; it is fine if the extra filter provably removes nothing (verified count check) or the task itself specifies dropping those rows.
- **Consequence**: The per-group means (and often the test statistic) shift away from the ground-truth values computed on the full groups, so the numeric answers fail exact-match checks even though the test procedure and answer format look correct.
812Spec file referenced by the task is never opened, so definitions and outputs are inventedtaskda-code
Applies when
task -- the instructions point to an external specification (a guidance/README/config file, notebook, or prior step's schema) that defines the filtering rules, category definitions, and the set of artifacts to save.
Pattern
The script never reads or quotes the referenced spec; instead the agent hard-codes its own plausible-looking rules (arbitrary thresholds, date windows, keyword-based category mapping) and writes only the one artifact explicitly named in the prompt, silently dropping other files the spec requires (e.g., serialized plot data, saved arrays) and any prescribed label ordering/color mapping.
Detection procedure
  1. From the task text, list every referenced document and every output artifact/format the workflow implies (image, JSON, array dumps, printed values) plus every quantitative rule that must come from that document.
  2. Search the scripts for a read/parse of the referenced document; check whether each rule (subset filters, category definitions, ordering, colors, figure size) is traceable to it rather than to an inline comment authored by the agent.
  3. Diff the artifacts actually written by the script against the list from step 1; flag any missing file or any category/ordering choice the agent justified only with its own docstring.
  4. Check the answer text: does it assert conformance ("as per guidance") without showing the source rule or the full artifact list?
Discriminator
Fine if the spec's rules are reproduced verbatim/loaded from the file and all named artifacts are written, even if the agent adds extra commentary; a violation is when thresholds, groupings, or output sets exist only as the agent's own invention, or when the artifact set is a strict subset of what the workflow requires.
Consequence
Grader checks each expected artifact independently; missing files fail outright and the invented filtering/grouping yields counts and a chart that mismatch the reference, giving 0/N checks even though the script runs cleanly.
id a06bf3bf5104 · mined from da-code dacode-plot-pie-005@s15
raw text (what the judge reads)
### Spec file referenced by the task is never opened, so definitions and outputs are invented
- **Applies when**: `task` -- the instructions point to an external specification (a guidance/README/config file, notebook, or prior step's schema) that defines the filtering rules, category definitions, and the set of artifacts to save.
- **Pattern**: The script never reads or quotes the referenced spec; instead the agent hard-codes its own plausible-looking rules (arbitrary thresholds, date windows, keyword-based category mapping) and writes only the one artifact explicitly named in the prompt, silently dropping other files the spec requires (e.g., serialized plot data, saved arrays) and any prescribed label ordering/color mapping.
- **Detection procedure**:
  1. From the task text, list every referenced document and every output artifact/format the workflow implies (image, JSON, array dumps, printed values) plus every quantitative rule that must come from that document.
  2. Search the scripts for a read/parse of the referenced document; check whether each rule (subset filters, category definitions, ordering, colors, figure size) is traceable to it rather than to an inline comment authored by the agent.
  3. Diff the artifacts actually written by the script against the list from step 1; flag any missing file or any category/ordering choice the agent justified only with its own docstring.
  4. Check the answer text: does it assert conformance ("as per guidance") without showing the source rule or the full artifact list?
- **Discriminator**: Fine if the spec's rules are reproduced verbatim/loaded from the file and all named artifacts are written, even if the agent adds extra commentary; a violation is when thresholds, groupings, or output sets exist only as the agent's own invention, or when the artifact set is a strict subset of what the workflow requires.
- **Consequence**: Grader checks each expected artifact independently; missing files fail outright and the invented filtering/grouping yields counts and a chart that mismatch the reference, giving 0/N checks even though the script runs cleanly.
813Unverified preprocessing of the mandated columns, with no sanity check of the error magnitude against target variancetaskinfiagent-dabench
Applies when
task -- the task prescribes an exact preprocessing recipe (e.g., impute specified columns with their means) plus a fixed split and a single error metric, and the agent reports one number.
Pattern
The attempt skips or silently changes the prescribed cleaning (drops rows with nulls instead of imputing, leaves the target column untreated, fails to coerce string/formatted numeric columns to numeric so "missing" is never detected, imputes after the split or with different statistics), and then reports the resulting error without checking whether its magnitude is plausible for the target's scale — and, worse, no script is preserved so none of this can be re-checked.
Detection procedure
  1. From the task, list every mandated preprocessing/split/metric step and the exact set of columns each applies to.
  2. In the scripts, locate each step and confirm it is executed on all named columns (including the response variable), that non-numeric/parsed columns are converted before missingness is measured, and that row counts before and after cleaning are unchanged when imputation (not deletion) is specified; if no script exists, treat the result as unverifiable and reject.
  3. Check the split/model settings actually used (fraction, seed, feature set) match the stated ones and that the metric is computed on the held-out subset with the standard definition.
  4. Compare the reported error to a simple reference the reviewer can compute (target variance / mean-prediction baseline); an error of the same order as, or larger than, the target variance indicates a broken pipeline rather than a real model.
Discriminator
A genuine violation shows a missing/altered mandated step, a shrunken row count, or an error not materially better than the mean-predictor baseline; a look-alike that is fine performs every mandated step verbatim and reports a large-but-explained error that still beats the baseline (large error can be legitimate when the target has a large scale).
Consequence
The metric is computed on a different effective dataset than the specification implies, so the reported value is off by an order of magnitude from the expected one and the exact-match grader fails.
id ea0b4d9dfd2f · mined from infiagent-dabench dabench-432@s15
raw text (what the judge reads)
### Unverified preprocessing of the mandated columns, with no sanity check of the error magnitude against target variance
- **Applies when**: `task` -- the task prescribes an exact preprocessing recipe (e.g., impute specified columns with their means) plus a fixed split and a single error metric, and the agent reports one number.
- **Pattern**: The attempt skips or silently changes the prescribed cleaning (drops rows with nulls instead of imputing, leaves the target column untreated, fails to coerce string/formatted numeric columns to numeric so "missing" is never detected, imputes after the split or with different statistics), and then reports the resulting error without checking whether its magnitude is plausible for the target's scale — and, worse, no script is preserved so none of this can be re-checked.
- **Detection procedure**:
  1. From the task, list every mandated preprocessing/split/metric step and the exact set of columns each applies to.
  2. In the scripts, locate each step and confirm it is executed on all named columns (including the response variable), that non-numeric/parsed columns are converted before missingness is measured, and that row counts before and after cleaning are unchanged when imputation (not deletion) is specified; if no script exists, treat the result as unverifiable and reject.
  3. Check the split/model settings actually used (fraction, seed, feature set) match the stated ones and that the metric is computed on the held-out subset with the standard definition.
  4. Compare the reported error to a simple reference the reviewer can compute (target variance / mean-prediction baseline); an error of the same order as, or larger than, the target variance indicates a broken pipeline rather than a real model.
- **Discriminator**: A genuine violation shows a missing/altered mandated step, a shrunken row count, or an error not materially better than the mean-predictor baseline; a look-alike that is fine performs every mandated step verbatim and reports a large-but-explained error that still beats the baseline (large error can be legitimate when the target has a large scale).
- **Consequence**: The metric is computed on a different effective dataset than the specification implies, so the reported value is off by an order of magnitude from the expected one and the exact-match grader fails.
814Row ordering not verified before computing sequential/time-dependent differencestaskinfiagent-dabench
Applies when
task -- the task asks for a period-over-period change, lag, diff, cumulative or rolling statistic that depends on the order of rows (e.g., a date/time-indexed series).
Pattern
The script loads the file and immediately applies a shift/diff/pct-change style operation using the file's native row order, without parsing the ordering key to a proper dtype and explicitly sorting ascending; if the source is stored newest-first (or unsorted), every difference is computed against the wrong neighbour, which typically flips the sign of the mean while leaving the dispersion nearly unchanged.
Detection procedure
  1. Read the task and note that the requested quantity is defined relative to a "previous"/"next" row, i.e., it is order-dependent.
  2. In the script, look for an explicit conversion of the ordering column to a sortable type and an explicit ascending sort (or set_index + sort_index) executed before the shift/diff; also check the first/last rows of the raw file are inspected.
  3. Check whether any sanity output is printed that would reveal ordering (head/tail of the ordering column, first few computed values, or a comparison of a manual first-vs-second-row calculation).
  4. Compare the reported location statistic's sign/magnitude to what the raw endpoints imply (last value vs. first value over the span); a mismatch in sign is decisive.
Discriminator
A real violation is when no ordering check or sort exists and the ordering key is text/unsorted; it is not a violation if the script explicitly sorts ascending (or verifiably shows the data is already ascending) and only then computes the differences — nor if the statistic is order-independent (e.g., a plain mean of an existing column).
Consequence
The dispersion statistic looks plausible (matches to ~0.01) while the mean has the opposite sign, so the answer fails all checks despite appearing internally consistent.
id 5f7f80970c24 · mined from infiagent-dabench dabench-75@s15
raw text (what the judge reads)
### Row ordering not verified before computing sequential/time-dependent differences
- **Applies when**: `task` -- the task asks for a period-over-period change, lag, diff, cumulative or rolling statistic that depends on the order of rows (e.g., a date/time-indexed series).
- **Pattern**: The script loads the file and immediately applies a shift/diff/pct-change style operation using the file's native row order, without parsing the ordering key to a proper dtype and explicitly sorting ascending; if the source is stored newest-first (or unsorted), every difference is computed against the wrong neighbour, which typically flips the sign of the mean while leaving the dispersion nearly unchanged.
- **Detection procedure**:
  1. Read the task and note that the requested quantity is defined relative to a "previous"/"next" row, i.e., it is order-dependent.
  2. In the script, look for an explicit conversion of the ordering column to a sortable type and an explicit ascending sort (or set_index + sort_index) executed *before* the shift/diff; also check the first/last rows of the raw file are inspected.
  3. Check whether any sanity output is printed that would reveal ordering (head/tail of the ordering column, first few computed values, or a comparison of a manual first-vs-second-row calculation).
  4. Compare the reported location statistic's sign/magnitude to what the raw endpoints imply (last value vs. first value over the span); a mismatch in sign is decisive.
- **Discriminator**: A real violation is when no ordering check or sort exists and the ordering key is text/unsorted; it is *not* a violation if the script explicitly sorts ascending (or verifiably shows the data is already ascending) and only then computes the differences — nor if the statistic is order-independent (e.g., a plain mean of an existing column).
- **Consequence**: The dispersion statistic looks plausible (matches to ~0.01) while the mean has the opposite sign, so the answer fails all checks despite appearing internally consistent.
815Silent row-dropping changes the analyzed sample without justificationtaskinfiagent-dabench
Applies when
task -- the task asks for a statistic over two (or more) columns of a table and the script applies dropna(), filtering, or subsetting before computing it.
Pattern
The agent unilaterally removes rows (pairwise/listwise NaN drops, zero/outlier filters, or subsetting to a slice) that the task never asked for, computes the statistic on this reduced subset, and reports it as the answer — never comparing it against the statistic on the full column pair, so a small but decisive shift in the value goes unnoticed.
Detection procedure
  1. Read the task and list any filtering the task explicitly authorizes; if none is stated, the expected sample is all rows of the requested columns.
  2. In the script, find every operation that changes row count before the statistic (dropna, boolean masks, head, index_col/parsing choices, dtype coercions that turn values into NaN) and check whether it is task-mandated.
  3. Check whether the script prints/compares the statistic both with and without the unauthorized filtering, plus row counts before/after; if only one variant is reported, flag it.
  4. Confirm the reported number is the one computed on the task-sanctioned sample at the requested rounding.
Discriminator
A real violation is dropping/filtering rows that the task did not request (or where the rows dropped are not truly missing, e.g. NaNs created by a bad read/dtype step) with no robustness comparison. It is fine if the filtering is explicitly required by the task, or if the script demonstrates the statistic is unchanged (to the required precision) with and without the drop.
Consequence
The reported statistic differs from the ground-truth value computed on the intended rows — here off by one unit in the second decimal (0.53 vs 0.54) — so the numeric check fails even though the qualitative conclusion happens to pass.
id 48e63961f9b9 · mined from infiagent-dabench dabench-300@s15
raw text (what the judge reads)
### Silent row-dropping changes the analyzed sample without justification
- **Applies when**: `task` -- the task asks for a statistic over two (or more) columns of a table and the script applies `dropna()`, filtering, or subsetting before computing it.
- **Pattern**: The agent unilaterally removes rows (pairwise/listwise NaN drops, zero/outlier filters, or subsetting to a slice) that the task never asked for, computes the statistic on this reduced subset, and reports it as the answer — never comparing it against the statistic on the full column pair, so a small but decisive shift in the value goes unnoticed.
- **Detection procedure**:
  1. Read the task and list any filtering the task explicitly authorizes; if none is stated, the expected sample is all rows of the requested columns.
  2. In the script, find every operation that changes row count before the statistic (`dropna`, boolean masks, `head`, `index_col`/parsing choices, dtype coercions that turn values into NaN) and check whether it is task-mandated.
  3. Check whether the script prints/compares the statistic both with and without the unauthorized filtering, plus row counts before/after; if only one variant is reported, flag it.
  4. Confirm the reported number is the one computed on the task-sanctioned sample at the requested rounding.
- **Discriminator**: A real violation is dropping/filtering rows that the task did not request (or where the rows dropped are not truly missing, e.g. NaNs created by a bad read/dtype step) with no robustness comparison. It is fine if the filtering is explicitly required by the task, or if the script demonstrates the statistic is unchanged (to the required precision) with and without the drop.
- **Consequence**: The reported statistic differs from the ground-truth value computed on the intended rows — here off by one unit in the second decimal (0.53 vs 0.54) — so the numeric check fails even though the qualitative conclusion happens to pass.
816Prediction file not verified to align row-for-row with the test inputtaskda-code
Applies when
task -- the deliverable is a file of per-row predictions (one value per test observation) written out by the agent's script.
Pattern
The agent produces a prediction file without ever asserting that its length, row order, and header match the test input; rows get dropped or reordered by upstream steps (dropna/filtering, grouping, deduplication, merges, reset_index-less joins, or reading a partial/subset frame), or the file is hand-assembled/truncated rather than emitted by a reproducible script, so the submitted column is not one-to-one with the test rows.
Detection procedure
  1. From the task, determine the required output: exact filename, exact column name, and the expected number of rows = number of rows in the test input.
  2. In the scripts, trace the object written to disk back to the test frame: check whether any step between reading the test file and writing predictions can change row count or order (filtering, dropna, merge, groupby, sampling, index-based alignment) and whether an explicit len(pred) == len(test) / index-equality assertion exists.
  3. Inspect the submitted file itself: count data rows, confirm the header string, and confirm there are no missing/blank/duplicated-block rows.
  4. Flag if the count differs from the test row count, if order-preservation is not guaranteed, or if no script exists that regenerates the file deterministically.
Discriminator
A real violation is a mismatch in count/order/header or an unreproducible, manually pasted output; a look-alike that is fine is a pipeline that transforms features heavily but keeps the test frame's index intact and writes exactly one prediction per test row with the required header (values themselves may be imperfect — accuracy is judged separately).
Consequence
The grader cannot align predictions to ground truth, so the output file is scored as wrong/missing regardless of model quality (0/1 checks passed).
id eb22e2b818be · mined from da-code dacode-ml-regression-014@s15
raw text (what the judge reads)
### Prediction file not verified to align row-for-row with the test input
- **Applies when**: `task` -- the deliverable is a file of per-row predictions (one value per test observation) written out by the agent's script.
- **Pattern**: The agent produces a prediction file without ever asserting that its length, row order, and header match the test input; rows get dropped or reordered by upstream steps (dropna/filtering, grouping, deduplication, merges, `reset_index`-less joins, or reading a partial/subset frame), or the file is hand-assembled/truncated rather than emitted by a reproducible script, so the submitted column is not one-to-one with the test rows.
- **Detection procedure**:
  1. From the task, determine the required output: exact filename, exact column name, and the expected number of rows = number of rows in the test input.
  2. In the scripts, trace the object written to disk back to the test frame: check whether any step between reading the test file and writing predictions can change row count or order (filtering, `dropna`, `merge`, `groupby`, sampling, index-based alignment) and whether an explicit `len(pred) == len(test)` / index-equality assertion exists.
  3. Inspect the submitted file itself: count data rows, confirm the header string, and confirm there are no missing/blank/duplicated-block rows.
  4. Flag if the count differs from the test row count, if order-preservation is not guaranteed, or if no script exists that regenerates the file deterministically.
- **Discriminator**: A real violation is a mismatch in count/order/header or an unreproducible, manually pasted output; a look-alike that is fine is a pipeline that transforms features heavily but keeps the test frame's index intact and writes exactly one prediction per test row with the required header (values themselves may be imperfect — accuracy is judged separately).
- **Consequence**: The grader cannot align predictions to ground truth, so the output file is scored as wrong/missing regardless of model quality (0/1 checks passed).
817Incomplete production of the required output artifacts specified by the task or its config filetaskda-code
Applies when
task -- the task points to an external specification (a config/spec/guidelines file) and/or names deliverable files, and the agent must emit plot/figure plus accompanying machine-readable result files.
Pattern
The attempt fixates on the one explicitly named artifact (e.g., the image) and reports numbers only in prose, never parsing the spec file end-to-end nor writing the companion serialized outputs (plot data / arrays / selected-key files) that the spec or the grading harness expects; no script is saved that shows how each artifact was produced.
Detection procedure
  1. Read the task and open the referenced spec/config file; list every required item: figure file name, figure parameters (size, colors, title, labels, ordering, percentage format) and any required non-image outputs.
  2. Read the scripts and check that each listed item is loaded from the spec (not hardcoded from memory) and that each required output file is explicitly written to disk with the exact expected name/format.
  3. Check the final answer: does it point to concrete saved files for every deliverable, or does it only narrate values in text?
  4. Flag if any required artifact is unaccounted for, or if the scripts themselves are missing/unreproducible so artifact generation cannot be verified.
Discriminator
A real violation is a deliverable named or implied by the task/spec that no script writes (or writes under a different name/format). Not a violation if all specified artifacts are written and the agent additionally summarizes them in prose, or if extra files are produced beyond the requirement.
Consequence
The grader marks the missing/renamed artifacts as WRONG/MISSING, so the attempt scores zero on those checks even when the headline computed value happens to be right.
id a2056a4b1b06 · mined from da-code dacode-plot-pie-008@s15
raw text (what the judge reads)
### Incomplete production of the required output artifacts specified by the task or its config file
- **Applies when**: `task` -- the task points to an external specification (a config/spec/guidelines file) and/or names deliverable files, and the agent must emit plot/figure plus accompanying machine-readable result files.
- **Pattern**: The attempt fixates on the one explicitly named artifact (e.g., the image) and reports numbers only in prose, never parsing the spec file end-to-end nor writing the companion serialized outputs (plot data / arrays / selected-key files) that the spec or the grading harness expects; no script is saved that shows how each artifact was produced.
- **Detection procedure**:
  1. Read the task and open the referenced spec/config file; list *every* required item: figure file name, figure parameters (size, colors, title, labels, ordering, percentage format) and any required non-image outputs.
  2. Read the scripts and check that each listed item is loaded from the spec (not hardcoded from memory) and that each required output file is explicitly written to disk with the exact expected name/format.
  3. Check the final answer: does it point to concrete saved files for every deliverable, or does it only narrate values in text?
  4. Flag if any required artifact is unaccounted for, or if the scripts themselves are missing/unreproducible so artifact generation cannot be verified.
- **Discriminator**: A real violation is a deliverable named or implied by the task/spec that no script writes (or writes under a different name/format). Not a violation if all specified artifacts are written and the agent additionally summarizes them in prose, or if extra files are produced beyond the requirement.
- **Consequence**: The grader marks the missing/renamed artifacts as WRONG/MISSING, so the attempt scores zero on those checks even when the headline computed value happens to be right.
818Entity-level aggregation and group-split definitions not verified before computing the statistictaskinfiagent-dabench
Applies when
task -- the task asks for a correlation/statistic between two derived per-entity quantities (e.g., a max over records, a span/duration) computed within subgroups defined by a data-driven threshold such as a median.
Pattern
The attempt computes the statistic directly on the raw record-level table, or aggregates with an unstated/ad-hoc definition (duration as row count instead of last–first timestamp, wrong grouping key so one entity appears multiple times, threshold applied to record-level rather than entity-level values, ambiguous >= vs > at the median), and reports a coefficient without any sanity check on the number of entities per group. Frequently no script is retained, so the aggregation cannot be audited.
Detection procedure
  1. From the task, write down the intended unit of analysis (one row per entity) and the exact definition of each derived variable and of the subgroup split.
  2. In the scripts, locate the groupby/aggregation and confirm the grouping key uniquely identifies an entity, that the "max"/"duration" columns are built with the stated definition and correct dtypes (datetime differences converted to consistent units, ordinal categories numeric), and that the threshold is computed on the aggregated table.
  3. Check that the correlation is run on the two aggregated columns for each subgroup separately, and that group sizes / total entity count are printed and sum to the number of distinct entities.
  4. Compare the reported values against these checks; if the aggregation step or entity counts are absent (or no script exists), treat the result as unverified.
Discriminator
A real violation is a different unit of analysis, a different derived-variable definition, or an unreproducible/undocumented pipeline — even when the sign and significance come out right. A look-alike that is fine is a documented, entity-level pipeline whose small numeric differences stem only from a legitimately ambiguous tie-handling rule that the agent states and whose group counts are reported.
Consequence
The p-value and relationship label can still look correct while the coefficient is off by a few hundredths (e.g., 0.58 vs 0.56), so the exact-value check fails and the answer is graded incorrect.
id d06366bd2cec · mined from infiagent-dabench dabench-431@s15
raw text (what the judge reads)
### Entity-level aggregation and group-split definitions not verified before computing the statistic
- **Applies when**: `task` -- the task asks for a correlation/statistic between two *derived* per-entity quantities (e.g., a max over records, a span/duration) computed within subgroups defined by a data-driven threshold such as a median.
- **Pattern**: The attempt computes the statistic directly on the raw record-level table, or aggregates with an unstated/ad-hoc definition (duration as row count instead of last–first timestamp, wrong grouping key so one entity appears multiple times, threshold applied to record-level rather than entity-level values, ambiguous `>=` vs `>` at the median), and reports a coefficient without any sanity check on the number of entities per group. Frequently no script is retained, so the aggregation cannot be audited.
- **Detection procedure**:
  1. From the task, write down the intended unit of analysis (one row per entity) and the exact definition of each derived variable and of the subgroup split.
  2. In the scripts, locate the groupby/aggregation and confirm the grouping key uniquely identifies an entity, that the "max"/"duration" columns are built with the stated definition and correct dtypes (datetime differences converted to consistent units, ordinal categories numeric), and that the threshold is computed on the aggregated table.
  3. Check that the correlation is run on the two aggregated columns for each subgroup separately, and that group sizes / total entity count are printed and sum to the number of distinct entities.
  4. Compare the reported values against these checks; if the aggregation step or entity counts are absent (or no script exists), treat the result as unverified.
- **Discriminator**: A real violation is a different unit of analysis, a different derived-variable definition, or an unreproducible/undocumented pipeline — even when the sign and significance come out right. A look-alike that is fine is a documented, entity-level pipeline whose small numeric differences stem only from a legitimately ambiguous tie-handling rule that the agent states and whose group counts are reported.
- **Consequence**: The p-value and relationship label can still look correct while the coefficient is off by a few hundredths (e.g., 0.58 vs 0.56), so the exact-value check fails and the answer is graded incorrect.
819Baseline feature set silently narrowed / available columns dropped instead of using all predictorstaskinfiagent-dabench
Applies when
task -- the task asks to build a "baseline" model predicting a target and then compare it to a model with one engineered feature, without explicitly listing which input columns to use.
Pattern
The script hand-picks a subset of the available columns as the baseline predictors, silently dropping some numeric columns and any non-numeric/categorical column (rather than encoding it), so the "original model" is not the natural all-features baseline and both RMSEs shift away from the reference values.
Detection procedure
  1. From the task statement, note whether the predictor set is specified; if not, the default is "all columns other than the target (plus the engineered one)".
  2. In the script, list the columns actually passed to the model and diff them against the dataframe's columns minus the target.
  3. For each omitted column, check whether the script gives a justification the task authorizes (e.g., the task says to exclude it, or it is an ID/leakage column); check whether categorical columns were dropped instead of encoded.
  4. Confirm the engineered-feature model differs from the baseline only by the added feature, and both use the same split/seed.
Discriminator
A real violation is an unexplained/unjustified subset (or dropping a categorical because it is a string). It is fine to omit a column the task explicitly excludes, an identifier, or a column that is the target's direct transform — and fine to use a listed feature set when the task names the features.
Consequence
Both reported RMSE values are computed from an under-specified model and miss the expected numbers (baseline and volume-model RMSE both off by ~0.03–0.04), so 2 of 3 graded checks fail even though the correlation is right.
id bbeef11e4613 · mined from infiagent-dabench dabench-549@s15
raw text (what the judge reads)
### Baseline feature set silently narrowed / available columns dropped instead of using all predictors
- **Applies when**: `task` -- the task asks to build a "baseline" model predicting a target and then compare it to a model with one engineered feature, without explicitly listing which input columns to use.
- **Pattern**: The script hand-picks a subset of the available columns as the baseline predictors, silently dropping some numeric columns and any non-numeric/categorical column (rather than encoding it), so the "original model" is not the natural all-features baseline and both RMSEs shift away from the reference values.
- **Detection procedure**:
  1. From the task statement, note whether the predictor set is specified; if not, the default is "all columns other than the target (plus the engineered one)".
  2. In the script, list the columns actually passed to the model and diff them against the dataframe's columns minus the target.
  3. For each omitted column, check whether the script gives a justification the task authorizes (e.g., the task says to exclude it, or it is an ID/leakage column); check whether categorical columns were dropped instead of encoded.
  4. Confirm the engineered-feature model differs from the baseline *only* by the added feature, and both use the same split/seed.
- **Discriminator**: A real violation is an unexplained/unjustified subset (or dropping a categorical because it is a string). It is fine to omit a column the task explicitly excludes, an identifier, or a column that is the target's direct transform — and fine to use a listed feature set when the task names the features.
- **Consequence**: Both reported RMSE values are computed from an under-specified model and miss the expected numbers (baseline and volume-model RMSE both off by ~0.03–0.04), so 2 of 3 graded checks fail even though the correlation is right.
820Unvalidated model with an unreconciled prediction-distribution sanity checktaskda-code
Applies when
task -- the script fits a classifier/regressor on labeled data and writes predictions for an unlabeled test file, with no held-out evaluation of the chosen configuration.
Pattern
The attempt trains one or more models, applies a default decision rule (e.g. argmax/0.5 threshold) directly to test data, never measures accuracy/AUC/F1 on any validation split, and then reports the output distribution while glossing over or misdescribing a clear mismatch with the training base rate (e.g. claiming a ~10% predicted positive rate "aligns with" a ~22% observed rate) — i.e. an obvious contradiction is stated in the answer but never investigated.
Detection procedure
  1. Read the task/output spec and note the label balance and any implied evaluation metric or output format (column names, row count, whether an index/ID column is wanted).
  2. Scan the script for any train/validation split, cross-validation, or scored evaluation of the final configuration; note if train_test_split/CV is imported but unused and predictions come straight from a single fit.
  3. Compare the reported predicted class distribution against the training class distribution and against the model/threshold choice; check whether the answer explains any large gap or just asserts consistency.
  4. Confirm the written file's columns/rows match the requested format exactly (extra or missing columns, wrong count) rather than trusting the narrative.
Discriminator
A real violation is zero measured performance evidence plus an unexplained or falsely-explained distribution/format anomaly; it is fine if the script reports validation scores (or explicitly tuned the threshold/class weights) and the predicted rate deviation is a deliberate, justified consequence of that tuning.
Consequence
The submitted prediction file scores below the grader's accuracy/agreement threshold (and may also be rejected on format), so the file is marked WRONG despite a confident "task completed successfully" report.
id 15e3f912d545 · mined from da-code dacode-ml-binary-016@s15
raw text (what the judge reads)
### Unvalidated model with an unreconciled prediction-distribution sanity check
- **Applies when**: `task` -- the script fits a classifier/regressor on labeled data and writes predictions for an unlabeled test file, with no held-out evaluation of the chosen configuration.
- **Pattern**: The attempt trains one or more models, applies a default decision rule (e.g. argmax/0.5 threshold) directly to test data, never measures accuracy/AUC/F1 on any validation split, and then reports the output distribution while glossing over or misdescribing a clear mismatch with the training base rate (e.g. claiming a ~10% predicted positive rate "aligns with" a ~22% observed rate) — i.e. an obvious contradiction is stated in the answer but never investigated.
- **Detection procedure**:
  1. Read the task/output spec and note the label balance and any implied evaluation metric or output format (column names, row count, whether an index/ID column is wanted).
  2. Scan the script for any train/validation split, cross-validation, or scored evaluation of the final configuration; note if `train_test_split`/CV is imported but unused and predictions come straight from a single fit.
  3. Compare the reported predicted class distribution against the training class distribution and against the model/threshold choice; check whether the answer explains any large gap or just asserts consistency.
  4. Confirm the written file's columns/rows match the requested format exactly (extra or missing columns, wrong count) rather than trusting the narrative.
- **Discriminator**: A real violation is zero measured performance evidence *plus* an unexplained or falsely-explained distribution/format anomaly; it is fine if the script reports validation scores (or explicitly tuned the threshold/class weights) and the predicted rate deviation is a deliberate, justified consequence of that tuning.
- **Consequence**: The submitted prediction file scores below the grader's accuracy/agreement threshold (and may also be rejected on format), so the file is marked WRONG despite a confident "task completed successfully" report.
821Inventing input data from memory instead of loading the provided filestaskda-code
Applies when
task -- the task ships a dataset (and often a sample output file) in the working directory, and the script must read them to compute the requested quantity.
Pattern
The script hardcodes arrays/values recalled from prior knowledge of a "well-known" dataset (or a plausible reconstruction) rather than reading the supplied files, and likewise writes an output schema invented by the agent instead of one copied from the provided example file; no step in the script ever opens the delivered data.
Detection procedure
  1. Read the task/README and list every artifact the agent is supposed to consume (data files, example/sample output files) and produce.
  2. Scan the script for I/O calls (read_csv, open, load, glob of the data directory); if the analysis variables are literal constants defined in-code, the provided data was never used.
  3. Check whether the output columns/rows/naming were derived from the sample file (was it read or at least printed?) or simply made up.
  4. Check whether stated procedural constraints implied by the task (e.g., a fixed random seed implying a resampling-based statistic) are actually exercised by the chosen method; a purely deterministic closed-form test makes the seed a dead giveaway that the intended method differs.
Discriminator
Legitimate: constants that are documented parameters/thresholds, or a small hardcoded lookup used in addition to loading the real data, with values verified against the loaded file. Violation: the entire analysis input (or the output schema) exists only as in-code literals, with no read of the shipped files and no reconciliation (row counts, means, column names) against them.
Consequence
The reported statistic is computed on data that differs from the graded source, so the saved file's value (and often its column layout) does not match the expected result, and the file check fails outright.
id cd54918acc79 · mined from da-code dacode-data-sa-039@s15
raw text (what the judge reads)
### Inventing input data from memory instead of loading the provided files
- **Applies when**: `task` -- the task ships a dataset (and often a sample output file) in the working directory, and the script must read them to compute the requested quantity.
- **Pattern**: The script hardcodes arrays/values recalled from prior knowledge of a "well-known" dataset (or a plausible reconstruction) rather than reading the supplied files, and likewise writes an output schema invented by the agent instead of one copied from the provided example file; no step in the script ever opens the delivered data.
- **Detection procedure**:
  1. Read the task/README and list every artifact the agent is supposed to consume (data files, example/sample output files) and produce.
  2. Scan the script for I/O calls (`read_csv`, `open`, `load`, glob of the data directory); if the analysis variables are literal constants defined in-code, the provided data was never used.
  3. Check whether the output columns/rows/naming were derived from the sample file (was it read or at least printed?) or simply made up.
  4. Check whether stated procedural constraints implied by the task (e.g., a fixed random seed implying a resampling-based statistic) are actually exercised by the chosen method; a purely deterministic closed-form test makes the seed a dead giveaway that the intended method differs.
- **Discriminator**: Legitimate: constants that are documented parameters/thresholds, or a small hardcoded lookup used *in addition* to loading the real data, with values verified against the loaded file. Violation: the entire analysis input (or the output schema) exists only as in-code literals, with no read of the shipped files and no reconciliation (row counts, means, column names) against them.
- **Consequence**: The reported statistic is computed on data that differs from the graded source, so the saved file's value (and often its column layout) does not match the expected result, and the file check fails outright.
822Numeric columns left as formatted strings before imputation and extreme-value selectiontaskda-code
Applies when
task -- the task asks to impute missing values with a statistic (mean/median) and then report the row(s) attaining the max/min of a numeric column that arrives from CSV with formatting artifacts (thousands separators, currency/percent signs, units, blanks, "N/A").
Pattern
The attempt loads the file with defaults, never checks dtypes, and calls fillna(mean)/idxmax/idxmin/sort_values on a column that pandas read as object; the mean silently skips the column (or errors are swallowed) and the extremes come from lexicographic string ordering, so a plausible-looking but wrong country/row is reported. Often no script or intermediate output is saved, so the number behind the answer is never shown.
Detection procedure
  1. In the task, note which column drives the reported answer and which columns must be imputed.
  2. In the scripts, look for an explicit cleaning step before the statistic: dtype inspection, stripping of ,/%/$/units, pd.to_numeric(..., errors='coerce'), and confirmation that the target column's dtype is numeric after cleaning.
  3. Check that the answer is accompanied by the actual extreme values (and count of imputed cells), not just the label(s); recompute or eyeball whether those values are the true global max/min and are physically plausible in the stated units.
  4. Confirm the required output artifact (e.g., the named result file) is written in the requested JSON shape.
Discriminator
A real violation is when no numeric coercion/dtype assertion exists for the column used in the comparison (or the reported extremes are absent/implausible); it is fine if the script verifies the column is already numeric, or cleans it and prints the min/max values that match the data's true range.
Consequence
The grader sees the wrong label(s) for the max and/or min (string-sorted or NaN-affected picks) and the check fails, even though the JSON key structure looks correct.
id 249c89cd8a75 · mined from da-code dacode-di-text-001@s15
raw text (what the judge reads)
### Numeric columns left as formatted strings before imputation and extreme-value selection
- **Applies when**: `task` -- the task asks to impute missing values with a statistic (mean/median) and then report the row(s) attaining the max/min of a numeric column that arrives from CSV with formatting artifacts (thousands separators, currency/percent signs, units, blanks, "N/A").
- **Pattern**: The attempt loads the file with defaults, never checks dtypes, and calls `fillna(mean)`/`idxmax`/`idxmin`/`sort_values` on a column that pandas read as `object`; the mean silently skips the column (or errors are swallowed) and the extremes come from lexicographic string ordering, so a plausible-looking but wrong country/row is reported. Often no script or intermediate output is saved, so the number behind the answer is never shown.
- **Detection procedure**:
  1. In the task, note which column drives the reported answer and which columns must be imputed.
  2. In the scripts, look for an explicit cleaning step before the statistic: dtype inspection, stripping of `,`/`%`/`$`/units, `pd.to_numeric(..., errors='coerce')`, and confirmation that the target column's dtype is numeric after cleaning.
  3. Check that the answer is accompanied by the actual extreme values (and count of imputed cells), not just the label(s); recompute or eyeball whether those values are the true global max/min and are physically plausible in the stated units.
  4. Confirm the required output artifact (e.g., the named result file) is written in the requested JSON shape.
- **Discriminator**: A real violation is when no numeric coercion/dtype assertion exists for the column used in the comparison (or the reported extremes are absent/implausible); it is fine if the script verifies the column is already numeric, or cleans it and prints the min/max values that match the data's true range.
- **Consequence**: The grader sees the wrong label(s) for the max and/or min (string-sorted or NaN-affected picks) and the check fails, even though the JSON key structure looks correct.
823Output file omits explicitly requested result componentstaskda-code
Applies when
task -- the task asks to save a result file that includes several named deliverables (e.g., computed component scores, a composite score, a segment/group label, and a final assigned level) and the script writes the file by selecting a subset of columns.
Pattern
The script computes all intermediate quantities and even writes them to a secondary "full" file, but the graded output file is built from a narrow column selection (e.g., only the ID plus the final label), silently dropping deliverables the task listed; the answer then describes the reduced file as complete.
Detection procedure
  1. From the task statement, enumerate every quantity the prompt says the saved file must include (and any required naming/ordering/rounding conventions).
  2. In the script, find the line that constructs the DataFrame written to the required output path and list its columns.
  3. Diff the two lists; also check whether a richer version was written to a different, non-required filename (a strong signal the agent knew the columns mattered but graded-file selection was arbitrary).
  4. Check the answer's description of the file format against the task's enumeration, and check that binning/threshold definitions used are the conventional/derived ones rather than invented cutoffs unsupported by the task or data distribution.
Discriminator
A real violation is when a task-named deliverable is absent from the required output file (or renamed beyond recognition); a look-alike that is fine is extra unrequested columns present, or column order differing when the task imposes no order, as long as every requested quantity appears in the graded file.
Consequence
The saved file fails column/content comparison against the expected result even if the underlying per-customer computation is right, so the file check reports WRONG/MISSING and the task scores 0.
id 91763d2c848f · mined from da-code dacode-dm-csv-052@s15
raw text (what the judge reads)
### Output file omits explicitly requested result components
- **Applies when**: `task` -- the task asks to save a result file that includes several named deliverables (e.g., computed component scores, a composite score, a segment/group label, and a final assigned level) and the script writes the file by selecting a subset of columns.
- **Pattern**: The script computes all intermediate quantities and even writes them to a secondary "full" file, but the graded output file is built from a narrow column selection (e.g., only the ID plus the final label), silently dropping deliverables the task listed; the answer then describes the reduced file as complete.
- **Detection procedure**:
  1. From the task statement, enumerate every quantity the prompt says the saved file must include (and any required naming/ordering/rounding conventions).
  2. In the script, find the line that constructs the DataFrame written to the required output path and list its columns.
  3. Diff the two lists; also check whether a richer version was written to a different, non-required filename (a strong signal the agent knew the columns mattered but graded-file selection was arbitrary).
  4. Check the answer's description of the file format against the task's enumeration, and check that binning/threshold definitions used are the conventional/derived ones rather than invented cutoffs unsupported by the task or data distribution.
- **Discriminator**: A real violation is when a task-named deliverable is absent from the required output file (or renamed beyond recognition); a look-alike that is fine is extra unrequested columns present, or column order differing when the task imposes no order, as long as every requested quantity appears in the graded file.
- **Consequence**: The saved file fails column/content comparison against the expected result even if the underlying per-customer computation is right, so the file check reports WRONG/MISSING and the task scores 0.
824Result never validated against the provided reference format or a plausibility sanity checktaskda-code
Applies when
task -- the task supplies a sample/reference output file and a described input dataset, and the script loads a data file, computes a statistic, and writes the output file directly.
Pattern
The script hard-codes an input filename without confirming it is the dataset described in the task (row count, column names/values matching the README), never reads or compares against the provided sample output (shape, index/header labels, ordering, rounding), and the answer reports numbers that contradict domain expectations (e.g., near-zero associations among variables that are obviously related, or values outside the plausible range) without any follow-up check.
Detection procedure
  1. Read the task for the named/described input data and for any provided sample output file, noting the format constraints it implies (labels, order, precision).
  2. In the script, check whether the loaded file is verified to be that dataset (print/assert on shape, expected columns, value ranges) and whether the sample output is ever loaded/compared to the produced file.
  3. Inspect the reported numbers and metadata (row counts, dtypes, statistic values) and ask whether they are consistent with the dataset description and with subject-matter expectation; flag if a surprising result is reported with no diagnostic.
  4. Flag the attempt if either the input identity or the output format was never checked, or an implausible result was accepted at face value.
Discriminator
A fine attempt either asserts/prints checks that confirm the input matches the described dataset and the output matches the sample's structure, or explicitly investigates and explains a surprising result; a violation is silently trusting whatever file/values appeared and asserting "format verified" only by printing its own output.
Consequence
The grader compares the produced file to the expected one and finds mismatched values (and possibly labels/ordering), so the single file check fails and the task scores 0.
id d5e3edfef1ef · mined from da-code dacode-data-sa-026@s15
raw text (what the judge reads)
### Result never validated against the provided reference format or a plausibility sanity check
- **Applies when**: `task` -- the task supplies a sample/reference output file and a described input dataset, and the script loads a data file, computes a statistic, and writes the output file directly.
- **Pattern**: The script hard-codes an input filename without confirming it is the dataset described in the task (row count, column names/values matching the README), never reads or compares against the provided sample output (shape, index/header labels, ordering, rounding), and the answer reports numbers that contradict domain expectations (e.g., near-zero associations among variables that are obviously related, or values outside the plausible range) without any follow-up check.
- **Detection procedure**:
  1. Read the task for the named/described input data and for any provided sample output file, noting the format constraints it implies (labels, order, precision).
  2. In the script, check whether the loaded file is verified to be that dataset (print/assert on shape, expected columns, value ranges) and whether the sample output is ever loaded/compared to the produced file.
  3. Inspect the reported numbers and metadata (row counts, dtypes, statistic values) and ask whether they are consistent with the dataset description and with subject-matter expectation; flag if a surprising result is reported with no diagnostic.
  4. Flag the attempt if either the input identity or the output format was never checked, or an implausible result was accepted at face value.
- **Discriminator**: A fine attempt either asserts/prints checks that confirm the input matches the described dataset and the output matches the sample's structure, or explicitly investigates and explains a surprising result; a violation is silently trusting whatever file/values appeared and asserting "format verified" only by printing its own output.
- **Consequence**: The grader compares the produced file to the expected one and finds mismatched values (and possibly labels/ordering), so the single file check fails and the task scores 0.
825Saved feature matrix doesn't match the space the model actually usedtaskda-code
Applies when
task -- the deliverable is a file containing the feature vectors plus model outputs (e.g., cluster labels), and the script transforms features (scaling, log, PCA, encoding) before fitting.
Pattern
The script fits the model on transformed/standardized features but writes the raw, untransformed (often highly skewed, outlier-dominated, wildly different-scale) columns into the output file alongside the labels; no outlier/skew handling is done either, so the labels are not recoverable or evaluable from the exported features and any grader-side quality metric (silhouette, separation, cluster-count sanity) computed on the exported matrix collapses.
Detection procedure
  1. From the task, note what exactly must appear in the output file (feature columns + label) and that the label must be consistent with those feature columns.
  2. In the script, trace which array is passed to fit/fit_predict versus which DataFrame is written to disk; check whether the written frame is the pre-transform copy.
  3. Check whether skew/outliers were addressed (log/robust transform, trimming) and whether the reported cluster sizes/metric were computed in the same space as the exported features.
  4. In the answer, look for feature ranges spanning many orders of magnitude (e.g., 0–5,000 next to 0–350,000) and a quality score reported only for the transformed space.
Discriminator
Fine if the exported columns are the same representation the model was fit on (or the transform is monotone-per-column and label geometry is preserved and stated); a violation if labels were assigned in a scaled/reduced space while raw incomparable-scale columns are exported, or if no attempt was made to tame extreme skew before clustering.
Consequence
Recomputing the clustering/quality check on the submitted feature columns yields labels or scores inconsistent with the reported ones (near-degenerate clusters dominated by one large-magnitude column), so the file comparison fails.
id 53e6622ad71b · mined from da-code dacode-ml-cluster-019@s15
raw text (what the judge reads)
### Saved feature matrix doesn't match the space the model actually used
- **Applies when**: `task` -- the deliverable is a file containing the feature vectors plus model outputs (e.g., cluster labels), and the script transforms features (scaling, log, PCA, encoding) before fitting.
- **Pattern**: The script fits the model on transformed/standardized features but writes the *raw*, untransformed (often highly skewed, outlier-dominated, wildly different-scale) columns into the output file alongside the labels; no outlier/skew handling is done either, so the labels are not recoverable or evaluable from the exported features and any grader-side quality metric (silhouette, separation, cluster-count sanity) computed on the exported matrix collapses.
- **Detection procedure**:
  1. From the task, note what exactly must appear in the output file (feature columns + label) and that the label must be consistent with those feature columns.
  2. In the script, trace which array is passed to `fit`/`fit_predict` versus which DataFrame is written to disk; check whether the written frame is the pre-transform copy.
  3. Check whether skew/outliers were addressed (log/robust transform, trimming) and whether the reported cluster sizes/metric were computed in the same space as the exported features.
  4. In the answer, look for feature ranges spanning many orders of magnitude (e.g., 0–5,000 next to 0–350,000) and a quality score reported only for the transformed space.
- **Discriminator**: Fine if the exported columns are the same representation the model was fit on (or the transform is monotone-per-column and label geometry is preserved and stated); a violation if labels were assigned in a scaled/reduced space while raw incomparable-scale columns are exported, or if no attempt was made to tame extreme skew before clustering.
- **Consequence**: Recomputing the clustering/quality check on the submitted feature columns yields labels or scores inconsistent with the reported ones (near-degenerate clusters dominated by one large-magnitude column), so the file comparison fails.
826Fabricated/hardcoded inputs instead of the provided data filestaskda-code
Applies when
task -- the task references a supplied dataset (README, data directory, or input files) and the script must derive its statistic from that data.
Pattern
The exploration step fails or is superficial (e.g., only lists a directory or reads the wrong file), and the analysis script then hardcodes summary numbers recalled from memory or guessed, computing the requested statistic from those invented constants rather than from the actual records.
Detection procedure
  1. In the task/README, confirm that actual data files are provided and identify what units/granularity the requested quantity has.
  2. Scan the analysis script for any read_csv/load of the real input; flag it if the key quantities appear as literal numbers or dictionaries defined in code with comments like "known values" / "classic dataset".
  3. Check whether the exploration script actually printed the real files' contents and whether the hardcoded values were verified against them; an error/exception left unresolved is a red flag.
  4. Sanity-check the reported answer's magnitude and definition against anything stated in the README or derivable from the real data (rates, counts, plausible ranges).
Discriminator
Legitimate cases load the provided data and may hardcode only constants that are given in the task text (e.g., a stated cutoff date or a documented threshold); a violation invents the core observations or aggregates that should have been computed from the files.
Consequence
The confidence interval / metric is computed on fictional inputs, so the reported numbers don't match the values derivable from the real data and the graded output file is marked wrong even though the code runs without error.
id d14d8e91b791 · mined from da-code dacode-data-sa-031@s15
raw text (what the judge reads)
### Fabricated/hardcoded inputs instead of the provided data files
- **Applies when**: `task` -- the task references a supplied dataset (README, data directory, or input files) and the script must derive its statistic from that data.
- **Pattern**: The exploration step fails or is superficial (e.g., only lists a directory or reads the wrong file), and the analysis script then hardcodes summary numbers recalled from memory or guessed, computing the requested statistic from those invented constants rather than from the actual records.
- **Detection procedure**:
  1. In the task/README, confirm that actual data files are provided and identify what units/granularity the requested quantity has.
  2. Scan the analysis script for any `read_csv`/`load` of the real input; flag it if the key quantities appear as literal numbers or dictionaries defined in code with comments like "known values" / "classic dataset".
  3. Check whether the exploration script actually printed the real files' contents and whether the hardcoded values were verified against them; an error/exception left unresolved is a red flag.
  4. Sanity-check the reported answer's magnitude and definition against anything stated in the README or derivable from the real data (rates, counts, plausible ranges).
- **Discriminator**: Legitimate cases load the provided data and may hardcode only constants that are given in the task text (e.g., a stated cutoff date or a documented threshold); a violation invents the core observations or aggregates that should have been computed from the files.
- **Consequence**: The confidence interval / metric is computed on fictional inputs, so the reported numbers don't match the values derivable from the real data and the graded output file is marked wrong even though the code runs without error.
827Submission never verified against the provided template (coverage, ids, order, location) or an offline metrictaskda-code
Applies when
task -- the task supplies an example/sample output file and a scoring metric, and the scripts write a predictions file directly from a model without comparing it to that template or estimating the score locally.
Pattern
The script fits a model, builds an output frame from its own assumptions (column names derived from an encoder's class order, ids taken from whatever frame was loaded, file written to a self-chosen path), and stops. There is no assertion that the row count equals the number of required rows, that the id set/order matches the template exactly, that no rows were dropped by preprocessing (NaN handling, unseen-category encoding failures), and no held-out estimate of the stated metric to show the output is better than a trivial baseline.
Detection procedure
  1. From the task, note the exact required output: file name/path, header names and order, and the set of ids/rows that must appear, plus the scoring metric.
  2. In the scripts, look for (a) reading the sample/template file and comparing shape, column names and id list against the generated frame, (b) an explicit check that every input row survived preprocessing, and (c) any hold-out/CV computation of the stated metric.
  3. Inspect the produced answer: count rows and columns, check ids against the input rows, check the written path against the requested one, and check that probability columns are mapped to the right label (not just positional).
  4. If none of steps 2a–2c exist, and step 3 cannot be confirmed from the script output alone, flag the attempt.
Discriminator
A fine attempt may skip a fancy validation split but still asserts shape/id/column agreement with the template and writes to the requested path; a violation ships the file with no coverage/format assertion and no metric estimate, so silent row loss, mislabeled probability columns, or a wrong output location goes undetected.
Consequence
The grader reports the expected output file as wrong or missing (mismatched row count/ids/columns or wrong path), or scores a needlessly poor metric value, with no local evidence the agent could have used to notice.
id 290cc7b2d122 · mined from da-code dacode-ml-competition-005@s16
raw text (what the judge reads)
### Submission never verified against the provided template (coverage, ids, order, location) or an offline metric
- **Applies when**: `task` -- the task supplies an example/sample output file and a scoring metric, and the scripts write a predictions file directly from a model without comparing it to that template or estimating the score locally.
- **Pattern**: The script fits a model, builds an output frame from its own assumptions (column names derived from an encoder's class order, ids taken from whatever frame was loaded, file written to a self-chosen path), and stops. There is no assertion that the row count equals the number of required rows, that the id set/order matches the template exactly, that no rows were dropped by preprocessing (NaN handling, unseen-category encoding failures), and no held-out estimate of the stated metric to show the output is better than a trivial baseline.
- **Detection procedure**:
  1. From the task, note the exact required output: file name/path, header names and order, and the set of ids/rows that must appear, plus the scoring metric.
  2. In the scripts, look for (a) reading the sample/template file and comparing shape, column names and id list against the generated frame, (b) an explicit check that every input row survived preprocessing, and (c) any hold-out/CV computation of the stated metric.
  3. Inspect the produced answer: count rows and columns, check ids against the input rows, check the written path against the requested one, and check that probability columns are mapped to the right label (not just positional).
  4. If none of steps 2a–2c exist, and step 3 cannot be confirmed from the script output alone, flag the attempt.
- **Discriminator**: A fine attempt may skip a fancy validation split but still asserts shape/id/column agreement with the template and writes to the requested path; a violation ships the file with no coverage/format assertion and no metric estimate, so silent row loss, mislabeled probability columns, or a wrong output location goes undetected.
- **Consequence**: The grader reports the expected output file as wrong or missing (mismatched row count/ids/columns or wrong path), or scores a needlessly poor metric value, with no local evidence the agent could have used to notice.
828Implausible validation score with no prediction-vs-training distribution sanity checktaskda-code
Applies when
task -- the agent trains a supervised model on a provided training set and must write predictions for a held-out file, reporting a validation score as evidence of quality.
Pattern
The attempt reports a near-perfect validation score (e.g. R² ≈ 0.97 on a noisy, heavy-tailed count/sales target) and accepts it uncritically, never checking whether the score is inflated by leakage (target-derived or duplicate/near-duplicate rows, target-encoded features, imputing/scaling fit on all data) and never comparing the predicted output distribution to the training target distribution.
Detection procedure
  1. From the task/README, note the nature of the target (noisy, skewed, count-like) and form a rough expectation of achievable accuracy from the given weak features.
  2. In the scripts, list every feature: flag any built from the target or from aggregations that include the row's own target, any duplicate rows shared between train/validation, and any preprocessing fit before the split.
  3. Compare the reported prediction summary (min/max, mean, median) against the training target's summary; a large shift (e.g. predicted median orders of magnitude below the training median, or mass collapsed near zero) signals a broken pipeline, mis-aligned feature columns, or rows silently dropped/filled.
  4. Confirm the answer includes such a comparison plus row-count/column-name checks against the test file; if only a self-reported score is given, mark inadequate.
Discriminator
A genuinely easy, near-deterministic target (or a task where the strong feature is legitimately available at prediction time) can honestly yield a very high score — the distinguishing signs of a real violation are (a) no leakage audit, and (b) a predicted distribution that is inconsistent with the training target despite the claimed near-perfect fit.
Consequence
The written prediction file scores far worse than the reported validation metric on the grader's held-out target (systematically under- or over-scaled predictions), so the expected output file is judged wrong.
id 1732a7bb4029 · mined from da-code dacode-ml-regression-008@s16
raw text (what the judge reads)
### Implausible validation score with no prediction-vs-training distribution sanity check
- **Applies when**: `task` -- the agent trains a supervised model on a provided training set and must write predictions for a held-out file, reporting a validation score as evidence of quality.
- **Pattern**: The attempt reports a near-perfect validation score (e.g. R² ≈ 0.97 on a noisy, heavy-tailed count/sales target) and accepts it uncritically, never checking whether the score is inflated by leakage (target-derived or duplicate/near-duplicate rows, target-encoded features, imputing/scaling fit on all data) and never comparing the predicted output distribution to the training target distribution.
- **Detection procedure**:
  1. From the task/README, note the nature of the target (noisy, skewed, count-like) and form a rough expectation of achievable accuracy from the given weak features.
  2. In the scripts, list every feature: flag any built from the target or from aggregations that include the row's own target, any duplicate rows shared between train/validation, and any preprocessing fit before the split.
  3. Compare the reported prediction summary (min/max, mean, median) against the training target's summary; a large shift (e.g. predicted median orders of magnitude below the training median, or mass collapsed near zero) signals a broken pipeline, mis-aligned feature columns, or rows silently dropped/filled.
  4. Confirm the answer includes such a comparison plus row-count/column-name checks against the test file; if only a self-reported score is given, mark inadequate.
- **Discriminator**: A genuinely easy, near-deterministic target (or a task where the strong feature is legitimately available at prediction time) can honestly yield a very high score — the distinguishing signs of a real violation are (a) no leakage audit, and (b) a predicted distribution that is inconsistent with the training target despite the claimed near-perfect fit.
- **Consequence**: The written prediction file scores far worse than the reported validation metric on the grader's held-out target (systematically under- or over-scaled predictions), so the expected output file is judged wrong.
829Statistical test run on the unfiltered population with a default, unjustified test specificationtaskda-code
Applies when
task -- the task asks for a p-value and an accept/reject decision, and the scripts load the raw file(s) and immediately feed whole columns into an off-the-shelf test function.
Pattern
The agent skips the scoping and specification steps of inference: it never restricts the data to the population/time window/category implied by the task or data description, and it picks the default test (e.g., two-sample parametric, two-sided, equal-variance) without checking the distribution shape, sample-size balance, or the directional wording of the hypothesis. The resulting p-value is computed on the wrong sample and/or under the wrong test definition, yet is reported with full confidence.
Detection procedure
  1. Read the task and README for any stated or strongly implied restriction on which records count (competition/segment, date range, group membership, exclusions) and for the direction of the alternative hypothesis; list them.
  2. Read the script and check whether each restriction appears as an explicit filter before the statistic is computed, and whether the row count after filtering is printed/compared to the raw count.
  3. Check the test call: does the chosen test match the data's distribution (was normality/skew or variance ever inspected?) and does the alternative/sided setting match the hypothesis wording? A silent default here is a red flag.
  4. Inspect the reported p-value magnitude: an extreme value (e.g., <1e-20) from a huge, unfiltered N usually indicates the test was run on the entire dataset rather than the intended subset.
Discriminator
A real violation is when the task/README implies a narrower population or a specific alternative and the script uses everything with defaults and no diagnostics. It is not a violation when the task genuinely covers all records and the script explicitly documents that the default test's assumptions were checked (or a robust/nonparametric alternative was justified) — the key is evidence of a deliberate choice, not an unexamined default.
Consequence
The p-value differs from the reference by many orders of magnitude (and the reject/fail-to-reject label may flip), so the exact-value check on the saved output file fails even though the file format is correct.
id db2cbef42d4f · mined from da-code dacode-data-sa-001@s16
raw text (what the judge reads)
### Statistical test run on the unfiltered population with a default, unjustified test specification
- **Applies when**: `task` -- the task asks for a p-value and an accept/reject decision, and the scripts load the raw file(s) and immediately feed whole columns into an off-the-shelf test function.
- **Pattern**: The agent skips the scoping and specification steps of inference: it never restricts the data to the population/time window/category implied by the task or data description, and it picks the default test (e.g., two-sample parametric, two-sided, equal-variance) without checking the distribution shape, sample-size balance, or the directional wording of the hypothesis. The resulting p-value is computed on the wrong sample and/or under the wrong test definition, yet is reported with full confidence.
- **Detection procedure**:
  1. Read the task and README for any stated or strongly implied restriction on which records count (competition/segment, date range, group membership, exclusions) and for the direction of the alternative hypothesis; list them.
  2. Read the script and check whether each restriction appears as an explicit filter before the statistic is computed, and whether the row count after filtering is printed/compared to the raw count.
  3. Check the test call: does the chosen test match the data's distribution (was normality/skew or variance ever inspected?) and does the `alternative`/sided setting match the hypothesis wording? A silent default here is a red flag.
  4. Inspect the reported p-value magnitude: an extreme value (e.g., <1e-20) from a huge, unfiltered N usually indicates the test was run on the entire dataset rather than the intended subset.
- **Discriminator**: A real violation is when the task/README implies a narrower population or a specific alternative and the script uses everything with defaults and no diagnostics. It is *not* a violation when the task genuinely covers all records and the script explicitly documents that the default test's assumptions were checked (or a robust/nonparametric alternative was justified) — the key is evidence of a deliberate choice, not an unexamined default.
- **Consequence**: The p-value differs from the reference by many orders of magnitude (and the reject/fail-to-reject label may flip), so the exact-value check on the saved output file fails even though the file format is correct.
830Template/output-format file provided but never inspectedtaskda-code
Applies when
task -- the task says the output must follow the exact structure and formatting of a provided sample/reference file (column names, order, sorting, rounding, units, row granularity).
Pattern
The scripts never load or print the sample file; the agent hard-codes column names, column order, sort order, and rounding from guesswork, then writes the deliverable and "verifies" only its own arithmetic against itself.
Detection procedure
  1. Read the task and note that a sample/spec file defines the required schema and formatting.
  2. Search the scripts for any read/print of that sample file (or an explicit assertion comparing produced headers/dtypes/row count/ordering to it) — note its absence.
  3. Check whether every formatting choice in the writing step (header spelling, column sequence, decimal rounding, sort key, aggregation granularity such as one row per group vs. per group-and-period) is justified by the sample rather than by assumption.
  4. Confirm the verification script re-derives the same numbers instead of validating the output against the required structure.
Discriminator
Fine if the agent actually reads/prints the reference file (or asserts equality of headers, column order and row keys against it) and then matches it; a violation is when the schema is invented and self-consistency is mistaken for compliance — even numerically correct aggregates fail.
Consequence
The grader compares the deliverable file against the expected one and marks it WRONG/MISSING due to mismatched headers, column order, rounding, row set or sort order, scoring 0 despite plausible-looking numbers.
id 75eb75361a6b · mined from da-code dacode-dm-csv-011@s16
raw text (what the judge reads)
### Template/output-format file provided but never inspected
- **Applies when**: `task` -- the task says the output must follow the exact structure and formatting of a provided sample/reference file (column names, order, sorting, rounding, units, row granularity).
- **Pattern**: The scripts never load or print the sample file; the agent hard-codes column names, column order, sort order, and rounding from guesswork, then writes the deliverable and "verifies" only its own arithmetic against itself.
- **Detection procedure**:
  1. Read the task and note that a sample/spec file defines the required schema and formatting.
  2. Search the scripts for any read/print of that sample file (or an explicit assertion comparing produced headers/dtypes/row count/ordering to it) — note its absence.
  3. Check whether every formatting choice in the writing step (header spelling, column sequence, decimal rounding, sort key, aggregation granularity such as one row per group vs. per group-and-period) is justified by the sample rather than by assumption.
  4. Confirm the verification script re-derives the same numbers instead of validating the output against the required structure.
- **Discriminator**: Fine if the agent actually reads/prints the reference file (or asserts equality of headers, column order and row keys against it) and then matches it; a violation is when the schema is invented and self-consistency is mistaken for compliance — even numerically correct aggregates fail.
- **Consequence**: The grader compares the deliverable file against the expected one and marks it WRONG/MISSING due to mismatched headers, column order, rounding, row set or sort order, scoring 0 despite plausible-looking numbers.
831Output feature columns silently redefined by dropping/transforming input featurestaskda-code
Applies when
task -- the deliverable is a per-row file whose columns are indexed to "the feature vector" (e.g. Feature_i) plus a derived label, and the script builds that file from an internally chosen, preprocessed subset of the input columns.
Pattern
The script arbitrarily discards columns it finds inconvenient (non-numeric, date, id-like, or "constant" fields) instead of encoding/deriving them, then writes the scaled/imputed matrix as the Feature_i columns, so the emitted feature vector neither covers the dataset's informative attributes nor matches the values a checker can reconstruct from the raw data.
Detection procedure
  1. Read the task and note exactly what the row-level output columns are supposed to represent (which features, how many, in what values/order).
  2. In the script, locate the explicit feature list and compare it to the full set of usable input columns: are categorical/date/text columns simply dropped rather than encoded or engineered, and is any documented attribute lost?
  3. Check what array is written to the output: raw/engineered feature values, or an intermediate transform (standardized, PCA, imputed) that cannot be traced back to the input rows.
  4. Confirm the answer/summary states the feature count and semantics, and that row count and column count match the input table and the declared feature list.
Discriminator
Dropping a column is fine when it is genuinely non-informative and justified (unique identifier, all-constant filler) and the retained set still spans the documented attribute groups; a violation is dropping informative attributes purely because they need encoding, or emitting a transformed matrix when the task asks for the feature vector itself.
Consequence
The saved file has the wrong number/meaning of Feature_i columns and values that don't correspond to the input rows, so the grader's comparison of the expected clustering/feature table fails outright even if the cluster labels are internally sensible.
id a53c5c91b588 · mined from da-code dacode-ml-cluster-014@s16
raw text (what the judge reads)
### Output feature columns silently redefined by dropping/transforming input features
- **Applies when**: `task` -- the deliverable is a per-row file whose columns are indexed to "the feature vector" (e.g. `Feature_i`) plus a derived label, and the script builds that file from an internally chosen, preprocessed subset of the input columns.
- **Pattern**: The script arbitrarily discards columns it finds inconvenient (non-numeric, date, id-like, or "constant" fields) instead of encoding/deriving them, then writes the *scaled/imputed* matrix as the `Feature_i` columns, so the emitted feature vector neither covers the dataset's informative attributes nor matches the values a checker can reconstruct from the raw data.
- **Detection procedure**:
  1. Read the task and note exactly what the row-level output columns are supposed to represent (which features, how many, in what values/order).
  2. In the script, locate the explicit feature list and compare it to the full set of usable input columns: are categorical/date/text columns simply dropped rather than encoded or engineered, and is any documented attribute lost?
  3. Check what array is written to the output: raw/engineered feature values, or an intermediate transform (standardized, PCA, imputed) that cannot be traced back to the input rows.
  4. Confirm the answer/summary states the feature count and semantics, and that row count and column count match the input table and the declared feature list.
- **Discriminator**: Dropping a column is fine when it is genuinely non-informative and justified (unique identifier, all-constant filler) and the retained set still spans the documented attribute groups; a violation is dropping informative attributes purely because they need encoding, or emitting a transformed matrix when the task asks for the feature vector itself.
- **Consequence**: The saved file has the wrong number/meaning of `Feature_i` columns and values that don't correspond to the input rows, so the grader's comparison of the expected clustering/feature table fails outright even if the cluster labels are internally sensible.
832Unjustified post-processing of predictions (rounding/clipping) that the submission format never requiredtaskda-code
Applies when
task -- the task asks for predictions written to a submission file whose format is defined by a provided sample/template, and the script applies transformations (rounding to integers, clipping to a train-observed range, casting dtype) to model outputs after prediction.
Pattern
The agent notices the target takes integer/bounded values in training data and therefore snaps continuous model outputs to integers and clamps them to the min/max seen in training, without checking the sample submission's dtype or the competition's evaluation metric; the destroyed sub-unit resolution (and the truncated tails) systematically worsens the score, and the file may also mismatch the expected column names/dtype/row order.
Detection procedure
  1. Read the task/README and open the provided sample submission: note the exact column names, row count, id ordering, and whether the example values are integers or floats; note the stated evaluation metric if given.
  2. In the script, locate every operation applied between model.predict(...) and to_csv(...) — look for round, astype(int), clip, np.maximum/minimum, and any re-mapping to class labels.
  3. Check whether each such operation is demanded by the sample format/metric, or is merely inferred from the training target's appearance; also verify the written frame's columns and id column are copied from the test/sample file rather than regenerated.
  4. Check the answer text for signs of self-justification from distribution statistics only (e.g., "predictions are integers, within valid range, mean close to train mean") in place of a held-out metric comparing rounded vs. unrounded predictions.
Discriminator
A real violation is discretization/clipping applied to a regression submission whose template shows continuous values or whose metric rewards continuous estimates, and with no validation evidence that it helps. It is not a violation if the sample submission clearly shows integer class labels / the task explicitly requires rounding or bounds, or if the agent measured both variants on a validation split and chose the better one.
Consequence
The submission file is accepted structurally but scores materially worse than the raw predictions (or fails a dtype/format check), so the grader marks the expected output file WRONG despite a plausible-sounding modeling narrative.
id 1265033083d3 · mined from da-code dacode-ml-competition-009@s16
raw text (what the judge reads)
### Unjustified post-processing of predictions (rounding/clipping) that the submission format never required
- **Applies when**: `task` -- the task asks for predictions written to a submission file whose format is defined by a provided sample/template, and the script applies transformations (rounding to integers, clipping to a train-observed range, casting dtype) to model outputs after prediction.
- **Pattern**: The agent notices the target takes integer/bounded values in training data and therefore snaps continuous model outputs to integers and clamps them to the min/max seen in training, without checking the sample submission's dtype or the competition's evaluation metric; the destroyed sub-unit resolution (and the truncated tails) systematically worsens the score, and the file may also mismatch the expected column names/dtype/row order.
- **Detection procedure**:
  1. Read the task/README and open the provided sample submission: note the exact column names, row count, id ordering, and whether the example values are integers or floats; note the stated evaluation metric if given.
  2. In the script, locate every operation applied between `model.predict(...)` and `to_csv(...)` — look for `round`, `astype(int)`, `clip`, `np.maximum/minimum`, and any re-mapping to class labels.
  3. Check whether each such operation is demanded by the sample format/metric, or is merely inferred from the training target's appearance; also verify the written frame's columns and id column are copied from the test/sample file rather than regenerated.
  4. Check the answer text for signs of self-justification from distribution statistics only (e.g., "predictions are integers, within valid range, mean close to train mean") in place of a held-out metric comparing rounded vs. unrounded predictions.
- **Discriminator**: A real violation is discretization/clipping applied to a regression submission whose template shows continuous values or whose metric rewards continuous estimates, and with no validation evidence that it helps. It is *not* a violation if the sample submission clearly shows integer class labels / the task explicitly requires rounding or bounds, or if the agent measured both variants on a validation split and chose the better one.
- **Consequence**: The submission file is accepted structurally but scores materially worse than the raw predictions (or fails a dtype/format check), so the grader marks the expected output file WRONG despite a plausible-sounding modeling narrative.
833Output row coverage silently reduced by ad-hoc row dropping / feature choicestaskda-code
Applies when
task -- the task asks for per-record outputs (e.g., a label per input row) saved to a file, and the script does its own cleaning, row filtering, or feature selection before producing them.
Pattern
The script drops rows with missing/unparsable values (or filters by a self-invented threshold) and includes every numeric-looking column—including identifier/code/coordinate style columns—as a feature, then writes an output file whose row count and feature set no longer correspond to the full input as the task implies; no check is made that the deliverable covers all input records or that features are meaningful.
Detection procedure
  1. Read the task and note whether the deliverable is expected to have one row per input record and whether any row filtering was authorized.
  2. In the script, locate every dropna, drop_duplicates, thresh=, boolean filter, or subset selection, and every column-inclusion rule, and determine the resulting number of rows and which columns become features.
  3. Compare the reported output shape in the answer against the raw input row count and against the columns a reasonable feature set would contain (exclude IDs, codes, arbitrary numeric keys).
  4. Confirm the script performs an explicit assertion/sanity check that output rows == input rows (or that any exclusion was explicitly requested); flag if absent.
Discriminator
A real violation is unilateral row loss or inclusion of non-informative identifier columns with no justification from the task; it is fine if the task explicitly permits filtering, or if missing values are imputed/encoded so that all records are retained and identifier-like columns are deliberately excluded and documented.
Consequence
The saved file has fewer rows (and/or a different feature-column count) than the reference deliverable, so shape/row-alignment checks fail and the file is graded WRONG regardless of clustering quality.
id eb5b51414743 · mined from da-code dacode-ml-cluster-009@s16
raw text (what the judge reads)
### Output row coverage silently reduced by ad-hoc row dropping / feature choices
- **Applies when**: `task` -- the task asks for per-record outputs (e.g., a label per input row) saved to a file, and the script does its own cleaning, row filtering, or feature selection before producing them.
- **Pattern**: The script drops rows with missing/unparsable values (or filters by a self-invented threshold) and includes every numeric-looking column—including identifier/code/coordinate style columns—as a feature, then writes an output file whose row count and feature set no longer correspond to the full input as the task implies; no check is made that the deliverable covers all input records or that features are meaningful.
- **Detection procedure**:
  1. Read the task and note whether the deliverable is expected to have one row per input record and whether any row filtering was authorized.
  2. In the script, locate every `dropna`, `drop_duplicates`, `thresh=`, boolean filter, or subset selection, and every column-inclusion rule, and determine the resulting number of rows and which columns become features.
  3. Compare the reported output shape in the answer against the raw input row count and against the columns a reasonable feature set would contain (exclude IDs, codes, arbitrary numeric keys).
  4. Confirm the script performs an explicit assertion/sanity check that output rows == input rows (or that any exclusion was explicitly requested); flag if absent.
- **Discriminator**: A real violation is unilateral row loss or inclusion of non-informative identifier columns with no justification from the task; it is fine if the task explicitly permits filtering, or if missing values are imputed/encoded so that all records are retained and identifier-like columns are deliberately excluded and documented.
- **Consequence**: The saved file has fewer rows (and/or a different feature-column count) than the reference deliverable, so shape/row-alignment checks fail and the file is graded WRONG regardless of clustering quality.
834Output artifact never validated against the test set's shape, keys, and value sanitytaskda-code
Applies when
task -- the task asks for predictions/derived values written to a file with a specified column name, and the scripts build features by merging/aggregating auxiliary tables before predicting.
Pattern
The agent trains a model, writes the result file, and declares success in prose without any post-write check that the file has exactly one row per test record in the original test order, contains exactly the requested column name, has no missing values, and holds values in a plausible range; joins/aggregations or dropna steps silently drop or reorder rows, or the header/index differs from what was requested.
Detection procedure
  1. From the task, note the required file name, required column name(s), and the expected number of rows (= number of test records) and their ordering key.
  2. In the scripts, trace the path from the test table to the written file: look for merges, groupby aggregations, dropna, filtering, or re-sorting between loading test and writing output, and check whether the prediction vector is re-attached to the original test frame by key rather than by positional concatenation.
  3. Look for an explicit verification block after writing (re-read the file; assert row count equals test row count, assert column names, assert no NaN, print min/max/mean) — absence of any such check is the flag.
  4. Cross-check the answer's own reported counts (e.g., test rows vs. rows written/matched) for inconsistency, and check whether it states any validated shape/range of the saved output.
Discriminator
A real violation is when no shape/column/NaN check exists, or the reported counts imply rows were lost/added/reordered relative to test; it is fine if the script asserts the output shape and column names (or the merge is provably left-join on the test key with no row-dropping) even if the model itself is simple.
Consequence
The grader reads result.csv and finds the wrong row count, a missing/renamed column, NaNs, or predictions misaligned with the test rows, so the file scores as WRONG/MISSING regardless of model quality.
id 086a5f13a484 · mined from da-code dacode-ml-regression-002@s16
raw text (what the judge reads)
### Output artifact never validated against the test set's shape, keys, and value sanity
- **Applies when**: `task` -- the task asks for predictions/derived values written to a file with a specified column name, and the scripts build features by merging/aggregating auxiliary tables before predicting.
- **Pattern**: The agent trains a model, writes the result file, and declares success in prose without any post-write check that the file has exactly one row per test record in the original test order, contains exactly the requested column name, has no missing values, and holds values in a plausible range; joins/aggregations or dropna steps silently drop or reorder rows, or the header/index differs from what was requested.
- **Detection procedure**:
  1. From the task, note the required file name, required column name(s), and the expected number of rows (= number of test records) and their ordering key.
  2. In the scripts, trace the path from the test table to the written file: look for merges, groupby aggregations, `dropna`, filtering, or re-sorting between loading test and writing output, and check whether the prediction vector is re-attached to the *original* test frame by key rather than by positional concatenation.
  3. Look for an explicit verification block after writing (re-read the file; assert row count equals test row count, assert column names, assert no NaN, print min/max/mean) — absence of any such check is the flag.
  4. Cross-check the answer's own reported counts (e.g., test rows vs. rows written/matched) for inconsistency, and check whether it states any validated shape/range of the saved output.
- **Discriminator**: A real violation is when no shape/column/NaN check exists, or the reported counts imply rows were lost/added/reordered relative to test; it is fine if the script asserts the output shape and column names (or the merge is provably left-join on the test key with no row-dropping) even if the model itself is simple.
- **Consequence**: The grader reads result.csv and finds the wrong row count, a missing/renamed column, NaNs, or predictions misaligned with the test rows, so the file scores as WRONG/MISSING regardless of model quality.
835Missing machine-checkable artifacts alongside the human-facing outputtaskda-code
Applies when
task -- the task's expected deliverables include serialized data/spec files (e.g. a JSON of plot parameters, an array/CSV of plotted or computed values) in addition to a rendered image or prose summary.
Pattern
The agent produces only the visual/prose artifact it was explicitly told to save (the image) and describes the configuration in text, never writing out the structured files the grader reads; no reproducible script is retained either, so the underlying series cannot be recovered or checked.
Detection procedure
1. List every output file implied by the task statement and its config/spec file (including any conventional companion files such as a spec dump and a numeric array of the plotted data). 2. Read the scripts/commands for file-writing calls and confirm each expected artifact is written with the expected name, type and shape. 3. Read the answer for claims like "all specifications followed" and check whether they are backed by saved files rather than narration. 4. Flag if any expected artifact is absent, unnamed, or only described in text.
Discriminator
A real violation is when a required artifact is never created (or created with a different name/format), so nothing but prose evidences the result; it is fine if all artifacts exist and the prose is merely an extra summary, or if the task genuinely requests only one output.
Consequence
The grader marks the expected files as WRONG/MISSING and scores 0, even if the rendered chart itself looks correct.
id 19b82ae59da9 · mined from da-code dacode-plot-line-015@s16
raw text (what the judge reads)
### Missing machine-checkable artifacts alongside the human-facing output
- **Applies when**: `task` -- the task's expected deliverables include serialized data/spec files (e.g. a JSON of plot parameters, an array/CSV of plotted or computed values) in addition to a rendered image or prose summary.
- **Pattern**: The agent produces only the visual/prose artifact it was explicitly told to save (the image) and describes the configuration in text, never writing out the structured files the grader reads; no reproducible script is retained either, so the underlying series cannot be recovered or checked.
- **Detection procedure**: 1. List every output file implied by the task statement and its config/spec file (including any conventional companion files such as a spec dump and a numeric array of the plotted data). 2. Read the scripts/commands for file-writing calls and confirm each expected artifact is written with the expected name, type and shape. 3. Read the answer for claims like "all specifications followed" and check whether they are backed by saved files rather than narration. 4. Flag if any expected artifact is absent, unnamed, or only described in text.
- **Discriminator**: A real violation is when a required artifact is never created (or created with a different name/format), so nothing but prose evidences the result; it is fine if all artifacts exist and the prose is merely an extra summary, or if the task genuinely requests only one output.
- **Consequence**: The grader marks the expected files as WRONG/MISSING and scores 0, even if the rendered chart itself looks correct.
836Analysis run on hardcoded/fabricated data instead of the provided input filestaskda-code
Applies when
task -- the task points to supplied data files (plus instruction/spec files such as a tips or sample-output file) that the scripts must read to compute the reported statistic.
Pattern
The script embeds literal arrays/values typed inline (or invented "representative" numbers) and computes the result from them, rather than loading the actual files; often the agent notes the data "isn't available" in one script and then proceeds with an in-code copy in the next, and likewise never opens the instruction/sample-format file.
Detection procedure
  1. Read the task and list every referenced input artifact (data file(s), instruction file, sample result file) and the required output path/columns.
  2. Grep the scripts for file I/O (read_csv, open, load, directory listing) and check that each listed artifact is actually read and its contents used downstream.
  3. Flag any numeric arrays/constants defined literally in the script that stand in for the dataset; check whether group sizes, labels, or subsetting rules were assumed rather than derived from the file.
  4. Compare the written output's columns/precision against the sample-format file the task named; if that file was never read, treat format compliance as unverified.
Discriminator
Inline constants are acceptable when they are genuine parameters (seed, number of resamples, thresholds) or values the task itself states; it is a violation when the observations, group membership, or sample sizes that determine the answer come from the script text instead of the supplied files.
Consequence
The reported statistic reflects invented data, so it mismatches the expected value (and may be a degenerate value such as 0.0 or 1.0 that also fails any range/plausibility sanity check), and the output file can miss required columns or precision — the grader marks the result file wrong.
id 64413f70e9e4 · mined from da-code dacode-data-sa-028@s16
raw text (what the judge reads)
### Analysis run on hardcoded/fabricated data instead of the provided input files
- **Applies when**: `task` -- the task points to supplied data files (plus instruction/spec files such as a tips or sample-output file) that the scripts must read to compute the reported statistic.
- **Pattern**: The script embeds literal arrays/values typed inline (or invented "representative" numbers) and computes the result from them, rather than loading the actual files; often the agent notes the data "isn't available" in one script and then proceeds with an in-code copy in the next, and likewise never opens the instruction/sample-format file.
- **Detection procedure**:
  1. Read the task and list every referenced input artifact (data file(s), instruction file, sample result file) and the required output path/columns.
  2. Grep the scripts for file I/O (`read_csv`, `open`, `load`, directory listing) and check that each listed artifact is actually read and its contents used downstream.
  3. Flag any numeric arrays/constants defined literally in the script that stand in for the dataset; check whether group sizes, labels, or subsetting rules were assumed rather than derived from the file.
  4. Compare the written output's columns/precision against the sample-format file the task named; if that file was never read, treat format compliance as unverified.
- **Discriminator**: Inline constants are acceptable when they are genuine parameters (seed, number of resamples, thresholds) or values the task itself states; it is a violation when the observations, group membership, or sample sizes that determine the answer come from the script text instead of the supplied files.
- **Consequence**: The reported statistic reflects invented data, so it mismatches the expected value (and may be a degenerate value such as 0.0 or 1.0 that also fails any range/plausibility sanity check), and the output file can miss required columns or precision — the grader marks the result file wrong.
837Claimed conformance to a provided spec (grouping/format rules) with no reproducible script or verificationtaskda-code
Applies when
task -- the task points to an auxiliary document/README that defines categories, bins, ordering, or output artifacts, and the agent must produce specific files plus a summary.
Pattern
The agent asserts the result "follows the specification" but leaves no script or intermediate artifact showing (a) that the spec file was actually parsed and its category boundaries/labels reproduced exactly, and (b) that every requested output file was written; the reported groups look plausible but are really the raw category labels already present in the data or the agent's own invented bins, and no count/coverage sanity check is done.
Detection procedure
  1. From the task, list the hard requirements: the referenced spec document, the exact bin/label/order rules it implies, and every output artifact (plot file, serialized data, arrays) with names/formats.
  2. In the scripts, look for code that reads the spec document (or a literal transcription of its bins) and maps raw values into those bins; if there is no such code — or no scripts at all — the conformance claim is unverifiable.
  3. Check that the script explicitly writes each requested artifact, and that the aggregation is done after dropping header/blank/unparseable rows, with the group counts summing to the number of valid rows rather than to the dataset's advertised total.
  4. Compare the labels/order in the answer against the spec's scheme; mismatched, extra, or missing groups (or a total that exactly equals the raw row count) indicate the spec was bypassed.
Discriminator
A real violation is a claim of compliance with no code path implementing the spec's binning or writing the named outputs, or counts that cannot be reconciled with valid-row counts; a look-alike that is fine is a script that hard-codes bins clearly quoted from the spec, emits every requested file, and reports totals equal to the non-missing response count.
Consequence
The grader compares the required artifacts (plot data, saved arrays, image) against the spec-defined grouping and finds them absent or built on the wrong bins/row set, failing all file checks despite a confident-sounding summary.
id afc4b4cc31b6 · mined from da-code dacode-plot-bar-005@s16
raw text (what the judge reads)
### Claimed conformance to a provided spec (grouping/format rules) with no reproducible script or verification
- **Applies when**: `task` -- the task points to an auxiliary document/README that defines categories, bins, ordering, or output artifacts, and the agent must produce specific files plus a summary.
- **Pattern**: The agent asserts the result "follows the specification" but leaves no script or intermediate artifact showing (a) that the spec file was actually parsed and its category boundaries/labels reproduced exactly, and (b) that every requested output file was written; the reported groups look plausible but are really the raw category labels already present in the data or the agent's own invented bins, and no count/coverage sanity check is done.
- **Detection procedure**:
  1. From the task, list the hard requirements: the referenced spec document, the exact bin/label/order rules it implies, and every output artifact (plot file, serialized data, arrays) with names/formats.
  2. In the scripts, look for code that reads the spec document (or a literal transcription of its bins) and maps raw values into those bins; if there is no such code — or no scripts at all — the conformance claim is unverifiable.
  3. Check that the script explicitly writes each requested artifact, and that the aggregation is done after dropping header/blank/unparseable rows, with the group counts summing to the number of valid rows rather than to the dataset's advertised total.
  4. Compare the labels/order in the answer against the spec's scheme; mismatched, extra, or missing groups (or a total that exactly equals the raw row count) indicate the spec was bypassed.
- **Discriminator**: A real violation is a claim of compliance with no code path implementing the spec's binning or writing the named outputs, or counts that cannot be reconciled with valid-row counts; a look-alike that is fine is a script that hard-codes bins clearly quoted from the spec, emits every requested file, and reports totals equal to the non-missing response count.
- **Consequence**: The grader compares the required artifacts (plot data, saved arrays, image) against the spec-defined grouping and finds them absent or built on the wrong bins/row set, failing all file checks despite a confident-sounding summary.
838Unvalidated choice for an ambiguous ingredient of a hand-coded statistictaskinfiagent-dabench
Applies when
task -- the task names a specific formula/statistic and the script computes it from sub-quantities whose definition is ambiguous or implementation-dependent (e.g. mode of a near-continuous variable, multi-modal ties, sample vs. population denominator, interpolation method for quantiles, binning choice).
Pattern
The script silently takes the first value returned by a library helper (or a default parameter) for the ambiguous ingredient, never inspecting whether the ingredient is unique/stable, and never comparing the final number against alternative admissible definitions or a library implementation; the single resulting number is reported as the answer.
Detection procedure
  1. Read the task and write down the exact formula requested and every sub-quantity it needs.
  2. In the scripts, locate how each sub-quantity is obtained and ask whether a different but equally defensible convention (ties, ddof, rounding/binning, interpolation) would change it materially.
  3. Check whether the script prints diagnostics for that ingredient (how many tied modes, value counts / distribution granularity, denominator used) and whether it computes the final statistic under at least one alternative convention for cross-checking.
  4. Check whether the reported final value is the one from the requested formula and whether its magnitude was reconciled with any other estimate the script printed (in this attempt, several competing skew-type numbers were printed and one was chosen without justification).
Discriminator
Fine if the ambiguous ingredient is provably unique/robust (e.g. a clear single dominant value shown in output, or the convention is fixed by the task statement) and the script documents that check; a violation is picking values[0] / a default silently while multiple plausible values exist and no sensitivity or sanity comparison is reported.
Consequence
The reported numeric value differs from the reference at the required rounding precision (only the qualitative label happens to match), so the numeric check fails and the overall answer is graded incorrect.
id 37f2d05e3fa6 · mined from infiagent-dabench dabench-359@s16
raw text (what the judge reads)
### Unvalidated choice for an ambiguous ingredient of a hand-coded statistic
- **Applies when**: `task` -- the task names a specific formula/statistic and the script computes it from sub-quantities whose definition is ambiguous or implementation-dependent (e.g. mode of a near-continuous variable, multi-modal ties, sample vs. population denominator, interpolation method for quantiles, binning choice).
- **Pattern**: The script silently takes the first value returned by a library helper (or a default parameter) for the ambiguous ingredient, never inspecting whether the ingredient is unique/stable, and never comparing the final number against alternative admissible definitions or a library implementation; the single resulting number is reported as *the* answer.
- **Detection procedure**:
  1. Read the task and write down the exact formula requested and every sub-quantity it needs.
  2. In the scripts, locate how each sub-quantity is obtained and ask whether a different but equally defensible convention (ties, ddof, rounding/binning, interpolation) would change it materially.
  3. Check whether the script prints diagnostics for that ingredient (how many tied modes, value counts / distribution granularity, denominator used) and whether it computes the final statistic under at least one alternative convention for cross-checking.
  4. Check whether the reported final value is the one from the requested formula and whether its magnitude was reconciled with any other estimate the script printed (in this attempt, several competing skew-type numbers were printed and one was chosen without justification).
- **Discriminator**: Fine if the ambiguous ingredient is provably unique/robust (e.g. a clear single dominant value shown in output, or the convention is fixed by the task statement) and the script documents that check; a violation is picking `values[0]` / a default silently while multiple plausible values exist and no sensitivity or sanity comparison is reported.
- **Consequence**: The reported numeric value differs from the reference at the required rounding precision (only the qualitative label happens to match), so the numeric check fails and the overall answer is graded incorrect.
839Verification that only re-derives the same formula, with no validation of the raw inputs or the output conventiontaskda-code
Applies when
task -- scripts read a supplied table, apply a chained/accumulating transformation (cumulative products, running sums, rolling stats) and write a result file whose column names and value convention were prescribed by the task.
Pattern
The agent loads the file and immediately feeds the raw columns into the computation without checking for missing/NaN entries, non-numeric dtypes, duplicate or unsorted keys, or scale (percent vs fraction); the "verification" script then recomputes the same expression with the same assumptions and declares a match, and the output column names/value convention (e.g. growth factor vs. net change, index reset at start) are invented by the agent rather than confirmed against the task's stated format.
Detection procedure
  1. Read the task for any prescribed output schema, ordering, or definition of the quantity, and note anything left implicit that the agent must pin down.
  2. In the scripts, look for any explicit inspection of the loaded data before use — isna().sum(), dtypes, row/date count, min/max ranges, sorting — and any handling step (fill/drop) justified by that inspection; absence of all of these is the flag, especially since a single missing entry propagates through an accumulating computation and corrupts every subsequent row.
  3. Check whether the "verification" script uses an independent route (hand-computed value, alternative formulation, reconciliation against a known benchmark) or simply repeats the original expression; repetition is not verification.
  4. Inspect the produced file: are column names and the value convention traceable to the task statement, and do the first/last values pass a plausibility check (first row equals the first period's value under the chosen convention, magnitudes in a sane range)?
Discriminator
A genuine violation is when nothing in the scripts could have surfaced a data defect or a convention mismatch — all checks are tautological restatements of the code. It is not a violation if the agent inspected the data and documented that it is complete/clean and numeric, or if the output convention is uniquely fixed by the task wording and the agent's check compares against an externally derived number.
Consequence
The file matches the agent's own arithmetic but not the reference; either the accumulated series diverges from the expected one after the first defective row, or the column names/value convention differ, and the grader reports the output file as WRONG despite the agent's self-consistency checks passing.
id 0f0bf03e137c · mined from da-code dacode-dm-csv-050@s16
raw text (what the judge reads)
### Verification that only re-derives the same formula, with no validation of the raw inputs or the output convention
- **Applies when**: `task` -- scripts read a supplied table, apply a chained/accumulating transformation (cumulative products, running sums, rolling stats) and write a result file whose column names and value convention were prescribed by the task.
- **Pattern**: The agent loads the file and immediately feeds the raw columns into the computation without checking for missing/NaN entries, non-numeric dtypes, duplicate or unsorted keys, or scale (percent vs fraction); the "verification" script then recomputes the same expression with the same assumptions and declares a match, and the output column names/value convention (e.g. growth factor vs. net change, index reset at start) are invented by the agent rather than confirmed against the task's stated format.
- **Detection procedure**:
  1. Read the task for any prescribed output schema, ordering, or definition of the quantity, and note anything left implicit that the agent must pin down.
  2. In the scripts, look for any explicit inspection of the loaded data before use — `isna().sum()`, `dtypes`, row/date count, min/max ranges, sorting — and any handling step (fill/drop) justified by that inspection; absence of all of these is the flag, especially since a single missing entry propagates through an accumulating computation and corrupts every subsequent row.
  3. Check whether the "verification" script uses an independent route (hand-computed value, alternative formulation, reconciliation against a known benchmark) or simply repeats the original expression; repetition is not verification.
  4. Inspect the produced file: are column names and the value convention traceable to the task statement, and do the first/last values pass a plausibility check (first row equals the first period's value under the chosen convention, magnitudes in a sane range)?
- **Discriminator**: A genuine violation is when nothing in the scripts could have surfaced a data defect or a convention mismatch — all checks are tautological restatements of the code. It is *not* a violation if the agent inspected the data and documented that it is complete/clean and numeric, or if the output convention is uniquely fixed by the task wording and the agent's check compares against an externally derived number.
- **Consequence**: The file matches the agent's own arithmetic but not the reference; either the accumulated series diverges from the expected one after the first defective row, or the column names/value convention differ, and the grader reports the output file as WRONG despite the agent's self-consistency checks passing.
840Distribution statistics computed on an uncleaned target vector (sentinels/NaNs/wrong dtype) with no sanity checktaskinfiagent-dabench
Applies when
task -- the task asks for a normality test and/or shape statistics (skewness, kurtosis, p-value) on a single numeric column extracted from a raw table.
Pattern
The script pulls the column straight from the file and feeds it to the test/statistic without inspecting it: missing-value placeholders (NaN, empty strings, sentinel codes like -999/0/9999), string-typed numerics coerced oddly, or duplicate/irrelevant rows are left in. The resulting heavy tails or large negative/positive skew are then reported at face value, and the stated intermediate (e.g., the p-value) is often omitted, so no one notices the vector analyzed is not the intended one.
Detection procedure
  1. Read the task to identify exactly which values are supposed to enter the statistic (column, any filtering, and any required intermediate to report such as a p-value).
  2. In the script, check whether the column is validated before the statistic: dtype check/to_numeric, isna().sum(), describe() or min/max/value_counts printed, and explicit handling of placeholder codes and non-finite values.
  3. Compare the reported statistics to that diagnostic output: do extreme skew/kurtosis magnitudes correspond to a plausible data range, or are they driven by a handful of out-of-range values? Recompute mentally with those values excluded.
  4. Confirm the answer includes every requested quantity (p-value, rounding, yes/no wording) and that the yes/no verdict follows from the same cleaned vector used for the moments.
Discriminator
A genuine violation is when no missing/sentinel/dtype inspection appears anywhere and the reported shape statistics are extreme relative to a well-behaved column; it is not a violation if the script prints the diagnostics, shows the column is clean (or documents the exclusions), and the extreme values are real observations within the variable's legitimate range.
Consequence
The normality decision flips (test rejects on contamination-driven tails) and skewness/kurtosis are far from the ground-truth values, so every checked field fails even though the test function itself was used correctly.
id 31f602a5e839 · mined from infiagent-dabench dabench-298@s16
raw text (what the judge reads)
### Distribution statistics computed on an uncleaned target vector (sentinels/NaNs/wrong dtype) with no sanity check
- **Applies when**: `task` -- the task asks for a normality test and/or shape statistics (skewness, kurtosis, p-value) on a single numeric column extracted from a raw table.
- **Pattern**: The script pulls the column straight from the file and feeds it to the test/statistic without inspecting it: missing-value placeholders (NaN, empty strings, sentinel codes like -999/0/9999), string-typed numerics coerced oddly, or duplicate/irrelevant rows are left in. The resulting heavy tails or large negative/positive skew are then reported at face value, and the stated intermediate (e.g., the p-value) is often omitted, so no one notices the vector analyzed is not the intended one.
- **Detection procedure**:
  1. Read the task to identify exactly which values are supposed to enter the statistic (column, any filtering, and any required intermediate to report such as a p-value).
  2. In the script, check whether the column is validated before the statistic: dtype check/`to_numeric`, `isna().sum()`, `describe()` or min/max/value_counts printed, and explicit handling of placeholder codes and non-finite values.
  3. Compare the reported statistics to that diagnostic output: do extreme skew/kurtosis magnitudes correspond to a plausible data range, or are they driven by a handful of out-of-range values? Recompute mentally with those values excluded.
  4. Confirm the answer includes every requested quantity (p-value, rounding, yes/no wording) and that the yes/no verdict follows from the same cleaned vector used for the moments.
- **Discriminator**: A genuine violation is when no missing/sentinel/dtype inspection appears anywhere and the reported shape statistics are extreme relative to a well-behaved column; it is *not* a violation if the script prints the diagnostics, shows the column is clean (or documents the exclusions), and the extreme values are real observations within the variable's legitimate range.
- **Consequence**: The normality decision flips (test rejects on contamination-driven tails) and skewness/kurtosis are far from the ground-truth values, so every checked field fails even though the test function itself was used correctly.
841Submitting model predictions without any held-out validation of predictive qualitytaskda-code
Applies when
task -- the task asks for predicted labels/values on a test file that will be scored against hidden ground truth, and the script fits a single model and writes predictions directly.
Pattern
The agent builds one default/first-guess pipeline (fixed vectorizer caps, aggressive token filtering, untuned linear model), fits it on all training rows, and immediately writes test predictions. The only "checks" printed are shapes and the distribution of predicted classes — nothing estimates how accurate the predictions actually are, so a model scoring far below the grader's threshold is indistinguishable from a good one.
Detection procedure
  1. Read the task to confirm the deliverable is scored on prediction quality (accuracy/error vs. hidden labels), not just file existence/format.
  2. Scan the training script for any train/validation split, cross-validation, or scoring call on labelled data (train_test_split, cross_val_score, .score, a metric on held-out rows). If absent, the attempt has no evidence its output is good.
  3. Check whether preprocessing/model choices that materially affect this task type were justified by measured comparison (e.g., discarding tokens that carry the signal, tight feature caps, no hyperparameter search, no alternative model tried).
  4. Look at the reported answer: if it consists only of the prediction column with no accompanying validation score, treat the quality claim as unsupported.
Discriminator
A real violation is the absence of any measured performance estimate on labelled data before submission. It is not a violation if the agent reports a held-out/CV score (even from a simple model) and that score is plausibly above the task's implicit bar — a simple baseline that was validated and compared against at least one alternative is acceptable; an unvalidated one is not.
Consequence
The written file has correct shape and column name but low label agreement with ground truth, so the accuracy check fails and the grader marks the expected result file WRONG despite the pipeline "running successfully".
id 0edaeec9c9fb · mined from da-code dacode-ml-multi-011@s16
raw text (what the judge reads)
### Submitting model predictions without any held-out validation of predictive quality

- **Applies when**: `task` -- the task asks for predicted labels/values on a test file that will be scored against hidden ground truth, and the script fits a single model and writes predictions directly.
- **Pattern**: The agent builds one default/first-guess pipeline (fixed vectorizer caps, aggressive token filtering, untuned linear model), fits it on all training rows, and immediately writes test predictions. The only "checks" printed are shapes and the distribution of predicted classes — nothing estimates how accurate the predictions actually are, so a model scoring far below the grader's threshold is indistinguishable from a good one.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on prediction quality (accuracy/error vs. hidden labels), not just file existence/format.
  2. Scan the training script for any train/validation split, cross-validation, or scoring call on labelled data (`train_test_split`, `cross_val_score`, `.score`, a metric on held-out rows). If absent, the attempt has no evidence its output is good.
  3. Check whether preprocessing/model choices that materially affect this task type were justified by measured comparison (e.g., discarding tokens that carry the signal, tight feature caps, no hyperparameter search, no alternative model tried).
  4. Look at the reported answer: if it consists only of the prediction column with no accompanying validation score, treat the quality claim as unsupported.
- **Discriminator**: A real violation is the absence of any measured performance estimate on labelled data before submission. It is *not* a violation if the agent reports a held-out/CV score (even from a simple model) and that score is plausibly above the task's implicit bar — a simple baseline that was validated and compared against at least one alternative is acceptable; an unvalidated one is not.
- **Consequence**: The written file has correct shape and column name but low label agreement with ground truth, so the accuracy check fails and the grader marks the expected result file WRONG despite the pipeline "running successfully".
842No held-out validation to select among successive competing modelstaskda-code
Applies when
task -- the agent writes several script versions that each fit a model and each overwrite the same final output file, and the deliverable is a prediction file scored against hidden labels.
Pattern
Every script trains on the full (or arbitrarily subsampled) training data and immediately predicts on test, without ever computing a held-out/cross-validated score in the competition's metric. The submitted file is simply whatever the last-executed script wrote — often a simpler, weaker, or sampled-down model — chosen for speed rather than measured accuracy, and no evidence is produced that it beats the earlier versions.
Detection procedure
  1. Read the task to identify the scored artifact and the (stated or implied) evaluation metric.
  2. Scan each script for a validation mechanism: a train/validation split, cross-validation, or any printed error/score on data not used for fitting. Note whether any script prints a comparable number.
  3. Check whether all scripts write to the same output path, and determine which one produced the final answer; see if that choice is justified by a measured score rather than by runtime or ordering.
  4. Inspect the final predictions for signs of an under-fit or degenerate model (e.g., tiny spread relative to the target's spread, output of a linear model on a nonlinear problem, model fit on a small random subsample of available data) with no accompanying accuracy estimate.
Discriminator
A real violation is the total absence of any out-of-sample score, so model choice is unjustified; it is not a violation if the agent reports validation/CV scores for each candidate and the submitted file demonstrably corresponds to the best-scoring candidate (retraining the chosen model on all data is fine), nor if a single model is used but its held-out error is measured and sane.
Consequence
The graded submission comes from an unvalidated, likely under-fit model whose error exceeds the required threshold, so the file is marked WRONG despite having a valid format and row count.
id c05f59e36665 · mined from da-code dacode-ml-competition-008@s16
raw text (what the judge reads)
### No held-out validation to select among successive competing models
- **Applies when**: `task` -- the agent writes several script versions that each fit a model and each overwrite the same final output file, and the deliverable is a prediction file scored against hidden labels.
- **Pattern**: Every script trains on the full (or arbitrarily subsampled) training data and immediately predicts on test, without ever computing a held-out/cross-validated score in the competition's metric. The submitted file is simply whatever the last-executed script wrote — often a simpler, weaker, or sampled-down model — chosen for speed rather than measured accuracy, and no evidence is produced that it beats the earlier versions.
- **Detection procedure**:
  1. Read the task to identify the scored artifact and the (stated or implied) evaluation metric.
  2. Scan each script for a validation mechanism: a train/validation split, cross-validation, or any printed error/score on data not used for fitting. Note whether any script prints a comparable number.
  3. Check whether all scripts write to the same output path, and determine which one produced the final answer; see if that choice is justified by a measured score rather than by runtime or ordering.
  4. Inspect the final predictions for signs of an under-fit or degenerate model (e.g., tiny spread relative to the target's spread, output of a linear model on a nonlinear problem, model fit on a small random subsample of available data) with no accompanying accuracy estimate.
- **Discriminator**: A real violation is the total absence of any out-of-sample score, so model choice is unjustified; it is *not* a violation if the agent reports validation/CV scores for each candidate and the submitted file demonstrably corresponds to the best-scoring candidate (retraining the chosen model on all data is fine), nor if a single model is used but its held-out error is measured and sane.
- **Consequence**: The graded submission comes from an unvalidated, likely under-fit model whose error exceeds the required threshold, so the file is marked WRONG despite having a valid format and row count.
843Deliverable not persisted to the requested output file/formattaskda-code
Applies when
task -- the task asks for a specific answer artifact (a named results file, a JSON object with exact keys, a fixed schema) and the scripts end by printing or saving results somewhere.
Pattern
The script computes plausible values but only prints them or writes them to an ad‑hoc path/filename (e.g. a scratch .txt in the working/home directory) instead of the exact deliverable the task specifies, so the graded artifact is missing or unreadable even when the numbers are right.
Detection procedure
  1. Read the task statement and list every hard output requirement: file name/extension, directory, top-level keys/labels, value types, and ordering.
  2. Scan the scripts for all write/save calls (open(...,'w'), to_csv, to_json, json.dump) and note the literal paths and the structure being written.
  3. Compare the written path/extension and the serialized structure key-by-key to the required spec; also confirm the final reported answer matches the file contents.
  4. Flag if no write targets the required file name/location, or if keys/ordering/types differ from the template given in the task.
Discriminator
A real violation is a mismatch in the artifact that will be graded (wrong filename, wrong extension, missing/renamed keys, values in the wrong order or wrong type). It is not a violation if the required file is written correctly and extra debug prints or additional copies exist elsewhere, or if the task never named an output artifact.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, regardless of whether the computed statistics were correct.
id e9396566f127 · mined from da-code dacode-di-text-003@s16
raw text (what the judge reads)
### Deliverable not persisted to the requested output file/format
- **Applies when**: `task` -- the task asks for a specific answer artifact (a named results file, a JSON object with exact keys, a fixed schema) and the scripts end by printing or saving results somewhere.
- **Pattern**: The script computes plausible values but only prints them or writes them to an ad‑hoc path/filename (e.g. a scratch `.txt` in the working/home directory) instead of the exact deliverable the task specifies, so the graded artifact is missing or unreadable even when the numbers are right.
- **Detection procedure**:
  1. Read the task statement and list every hard output requirement: file name/extension, directory, top-level keys/labels, value types, and ordering.
  2. Scan the scripts for all write/save calls (`open(...,'w')`, `to_csv`, `to_json`, `json.dump`) and note the literal paths and the structure being written.
  3. Compare the written path/extension and the serialized structure key-by-key to the required spec; also confirm the final reported answer matches the file contents.
  4. Flag if no write targets the required file name/location, or if keys/ordering/types differ from the template given in the task.
- **Discriminator**: A real violation is a mismatch in the artifact that will be graded (wrong filename, wrong extension, missing/renamed keys, values in the wrong order or wrong type). It is *not* a violation if the required file is written correctly and extra debug prints or additional copies exist elsewhere, or if the task never named an output artifact.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, regardless of whether the computed statistics were correct.
844Answer emitted without the literal quoting/delimiter form shown in the required templatetaskinfiagent-dabench
Applies when
task -- the task specifies an exact answer template with tagged fields (e.g. @field["Value"]) and the script/answer prints those tags itself.
Pattern
The attempt computes correct values but serializes them in a paraphrased form — dropping the quotation marks, changing brackets, adding a colon/space (@field: Value), reordering or renaming tags, or printing the tags only inside verbose log text — so an exact-match grader fails every field even though the analysis was right.
Detection procedure
  1. Copy the answer template verbatim from the task statement, noting every literal character: tag names, brackets, quotes around string values, separators between fields.
  2. Read the script's final print/summary block and the submitted answer string, and compare them character-by-character against the template (are string values wrapped in the same quotes? is the separator identical? is each tag present exactly once with the exact name?).
  3. If any literal differs — most commonly missing quotes around string-typed values, or @field: X instead of @field[X] — flag the attempt regardless of whether the underlying numbers/labels are correct.
  4. Also confirm the values themselves are the requested final quantities (e.g. a range/label string), not derived or intermediate printouts.
Discriminator
A real violation is a deviation in the literal formatting characters or tag set that an exact string matcher would reject; harmless look-alikes are surrounding log lines, extra whitespace outside the tags, or extra explanatory text, provided the exact templated substring (including quotes) appears intact somewhere in the final answer.
Consequence
The grader reports every expected field as WRONG/MISSING even though the submitted values equal the ground truth, yielding 0/N checks passed.
id 7f8600f89e58 · mined from infiagent-dabench dabench-550@s16
raw text (what the judge reads)
### Answer emitted without the literal quoting/delimiter form shown in the required template
- **Applies when**: `task` -- the task specifies an exact answer template with tagged fields (e.g. `@field["Value"]`) and the script/answer prints those tags itself.
- **Pattern**: The attempt computes correct values but serializes them in a paraphrased form — dropping the quotation marks, changing brackets, adding a colon/space (`@field: Value`), reordering or renaming tags, or printing the tags only inside verbose log text — so an exact-match grader fails every field even though the analysis was right.
- **Detection procedure**:
  1. Copy the answer template verbatim from the task statement, noting every literal character: tag names, brackets, quotes around string values, separators between fields.
  2. Read the script's final print/summary block and the submitted answer string, and compare them character-by-character against the template (are string values wrapped in the same quotes? is the separator identical? is each tag present exactly once with the exact name?).
  3. If any literal differs — most commonly missing quotes around string-typed values, or `@field: X` instead of `@field[X]` — flag the attempt regardless of whether the underlying numbers/labels are correct.
  4. Also confirm the values themselves are the requested final quantities (e.g. a range/label string), not derived or intermediate printouts.
- **Discriminator**: A real violation is a deviation in the literal formatting characters or tag set that an exact string matcher would reject; harmless look-alikes are surrounding log lines, extra whitespace outside the tags, or extra explanatory text, provided the exact templated substring (including quotes) appears intact somewhere in the final answer.
- **Consequence**: The grader reports every expected field as WRONG/MISSING even though the submitted values equal the ground truth, yielding 0/N checks passed.
845Analysis performed on data that doesn't contain the requested entitiestaskda-code
Applies when
task -- the task names specific entities/columns (e.g., a grouping key and a measured quantity) and the working directory contains several data files, only some of which are relevant.
Pattern
The agent picks a file that lacks the named columns, then silently substitutes a superficially similar grouping key and measure ("top-N by count of X" instead of "top-N by the requested aggregate of Y"), producing a chart of the right shape but the wrong content; required auxiliary outputs are also skipped.
Detection procedure
  1. From the task statement, list the exact quantities required: grouping dimension, ranking metric, stacked/aggregated measure, and every output artifact named or implied (image, serialized data, config-derived settings).
  2. In the scripts, check that the loaded file actually contains columns matching those quantities and that the ranking/aggregation is computed on the requested metric, not a proxy such as row counts.
  3. In the answer, verify the reported category labels and units are of the requested type (e.g., cities and average days), and that all expected output files were written.
  4. Flag if any substitution of dataset, key, or measure is made without evidence the requested columns are absent everywhere, or if named outputs are missing.
Discriminator
A real violation is when the plotted dimension/measure semantically differs from what was asked (different entity type or different statistic). A look-alike that is fine is when the columns are present under different names/derived fields but the script demonstrably maps them to the requested semantics (e.g., computing stage durations from timestamps).
Consequence
All graded artifacts mismatch — the saved figure, the numeric array, and the config/metadata file fail comparison, scoring 0 even though a plot was produced.
id de9952f21e1b · mined from da-code dacode-plot-scatter-002@s16
raw text (what the judge reads)
### Analysis performed on data that doesn't contain the requested entities
- **Applies when**: `task` -- the task names specific entities/columns (e.g., a grouping key and a measured quantity) and the working directory contains several data files, only some of which are relevant.
- **Pattern**: The agent picks a file that lacks the named columns, then silently substitutes a superficially similar grouping key and measure ("top-N by count of X" instead of "top-N by the requested aggregate of Y"), producing a chart of the right shape but the wrong content; required auxiliary outputs are also skipped.
- **Detection procedure**:
  1. From the task statement, list the exact quantities required: grouping dimension, ranking metric, stacked/aggregated measure, and every output artifact named or implied (image, serialized data, config-derived settings).
  2. In the scripts, check that the loaded file actually contains columns matching those quantities and that the ranking/aggregation is computed on the requested metric, not a proxy such as row counts.
  3. In the answer, verify the reported category labels and units are of the requested type (e.g., cities and average days), and that all expected output files were written.
  4. Flag if any substitution of dataset, key, or measure is made without evidence the requested columns are absent everywhere, or if named outputs are missing.
- **Discriminator**: A real violation is when the plotted dimension/measure semantically differs from what was asked (different entity type or different statistic). A look-alike that is fine is when the columns are present under different names/derived fields but the script demonstrably maps them to the requested semantics (e.g., computing stage durations from timestamps).
- **Consequence**: All graded artifacts mismatch — the saved figure, the numeric array, and the config/metadata file fail comparison, scoring 0 even though a plot was produced.
846Group-wise (stratified) analysis collapsed into a single global resulttaskda-code
Applies when
task -- the instructions say to perform preprocessing and/or a statistical test "for each" level of some grouping variable, and the requested output format uses lists/arrays for the reported values.
Pattern
The attempt does the filtering or the test once on the pooled dataset (or filters per group but then pools before testing), and returns a single scalar in each list, instead of producing one filtered subset and one test statistic per group in a defined order.
Detection procedure
  1. Read the task and count the distinct strata implied by the "for each X" clause and the plural/list-shaped output template; note the expected length of each output list and the ordering implied.
  2. Read the scripts: check that the split into strata happens before the outlier/quartile computation, and that the test is invoked inside the per-stratum loop rather than once on the concatenated frame.
  3. Compare the length of each list in the submitted answer to the number of strata; verify each conclusion string corresponds element-wise to its p-value.
  4. Sanity-check row counts per stratum after filtering (they should sum to less than the raw total and vary by stratum) — a single count equal to the whole cleaned dataset signals pooling.
Discriminator
A genuine violation returns fewer results than there are strata (typically one), or computes thresholds/quartiles on pooled data; it is not a violation if the task genuinely asks for one overall test and merely wraps the single value in a list, or if a stratum is legitimately dropped for insufficient data and this is explicitly justified and still ordered consistently.
Consequence
The saved result file mismatches the expected structure (wrong list length and wrong p-values, since pooled quartiles remove different points), so every element-wise check fails and the answer is graded incorrect.
id 857573903bc5 · mined from da-code dacode-data-sa-061@s16
raw text (what the judge reads)
### Group-wise (stratified) analysis collapsed into a single global result
- **Applies when**: `task` -- the instructions say to perform preprocessing and/or a statistical test "for each" level of some grouping variable, and the requested output format uses lists/arrays for the reported values.
- **Pattern**: The attempt does the filtering or the test once on the pooled dataset (or filters per group but then pools before testing), and returns a single scalar in each list, instead of producing one filtered subset and one test statistic per group in a defined order.
- **Detection procedure**:
  1. Read the task and count the distinct strata implied by the "for each X" clause and the plural/list-shaped output template; note the expected length of each output list and the ordering implied.
  2. Read the scripts: check that the split into strata happens *before* the outlier/quartile computation, and that the test is invoked inside the per-stratum loop rather than once on the concatenated frame.
  3. Compare the length of each list in the submitted answer to the number of strata; verify each conclusion string corresponds element-wise to its p-value.
  4. Sanity-check row counts per stratum after filtering (they should sum to less than the raw total and vary by stratum) — a single count equal to the whole cleaned dataset signals pooling.
- **Discriminator**: A genuine violation returns fewer results than there are strata (typically one), or computes thresholds/quartiles on pooled data; it is *not* a violation if the task genuinely asks for one overall test and merely wraps the single value in a list, or if a stratum is legitimately dropped for insufficient data and this is explicitly justified and still ordered consistently.
- **Consequence**: The saved result file mismatches the expected structure (wrong list length and wrong p-values, since pooled quartiles remove different points), so every element-wise check fails and the answer is graded incorrect.
847Model selection and quality claims based only on in-sample (training-set) fittaskda-code
Applies when
task -- the task asks for predictions on a held-out test file and the script trains one or more models, then reports fit metrics and/or chooses ensemble weights.
Pattern
The script calls fit on the full training data and then computes MSE/MAE/R² by predicting on that same training data (no train/validation split, no cross-validation), uses those optimistic numbers to rank models and hand-pick blend weights, and the answer presents them as evidence the predictions are good — while never checking generalization or whether the available informative columns (categorical/metadata/date fields) were used at all.
Detection procedure
  1. From the task, note that the deliverable is judged by predictive accuracy on unseen rows, so the only meaningful evidence is out-of-sample error.
  2. In the script, check the data passed to the metric functions: if the same matrix used in fit is passed to predict/score, and there is no train_test_split/cross_val_score/holdout fold, all reported metrics are in-sample.
  3. Check whether model/ensemble hyperparameters or weights were chosen from those in-sample numbers, and whether strong non-numeric predictors present in the data were dropped without any encoding or justification.
  4. In the answer, look for accuracy claims (R², "best performer", "realistic range") that rest solely on these training-fit numbers or on distributional sanity checks rather than validated error.
Discriminator
A real violation is when no estimate of out-of-sample error exists anywhere (a tree ensemble's near-perfect training R² is meaningless), and decisions were made from training fit. It is fine if the script reports training fit for diagnostics but also reports holdout/CV error and selects models/weights from that, or if the chosen model is validated by a proper split even when only summary statistics appear in the final answer.
Consequence
The submitted predictions can be far worse than the reported R² suggests (over-fit ensemble, unused predictive columns, over-shrunk predictions clustered near the mean), so the grader's accuracy/correlation check against the true target fails even though the file has the right shape and column name.
id 6c03dc4d862d · mined from da-code dacode-ml-regression-004@s16
raw text (what the judge reads)
### Model selection and quality claims based only on in-sample (training-set) fit
- **Applies when**: `task` -- the task asks for predictions on a held-out test file and the script trains one or more models, then reports fit metrics and/or chooses ensemble weights.
- **Pattern**: The script calls `fit` on the full training data and then computes MSE/MAE/R² by predicting on that same training data (no train/validation split, no cross-validation), uses those optimistic numbers to rank models and hand-pick blend weights, and the answer presents them as evidence the predictions are good — while never checking generalization or whether the available informative columns (categorical/metadata/date fields) were used at all.
- **Detection procedure**:
  1. From the task, note that the deliverable is judged by predictive accuracy on unseen rows, so the only meaningful evidence is out-of-sample error.
  2. In the script, check the data passed to the metric functions: if the same matrix used in `fit` is passed to `predict`/`score`, and there is no `train_test_split`/`cross_val_score`/holdout fold, all reported metrics are in-sample.
  3. Check whether model/ensemble hyperparameters or weights were chosen from those in-sample numbers, and whether strong non-numeric predictors present in the data were dropped without any encoding or justification.
  4. In the answer, look for accuracy claims (R², "best performer", "realistic range") that rest solely on these training-fit numbers or on distributional sanity checks rather than validated error.
- **Discriminator**: A real violation is when *no* estimate of out-of-sample error exists anywhere (a tree ensemble's near-perfect training R² is meaningless), and decisions were made from training fit. It is fine if the script reports training fit for diagnostics but also reports holdout/CV error and selects models/weights from that, or if the chosen model is validated by a proper split even when only summary statistics appear in the final answer.
- **Consequence**: The submitted predictions can be far worse than the reported R² suggests (over-fit ensemble, unused predictive columns, over-shrunk predictions clustered near the mean), so the grader's accuracy/correlation check against the true target fails even though the file has the right shape and column name.
848Unvalidated cluster solution written straight to the deliverabletaskda-code
Applies when
task -- the task asks for an unsupervised grouping with "an appropriate number of groups" exported to a specified file/column format, and the script picks the group count automatically from a single internal score and immediately writes the result.
Pattern
The attempt maximizes one selection statistic (e.g., silhouette) over a wide range, accepts whatever count wins even when it yields degenerate groups (singletons / handfuls of outliers) and no cross-check (elbow/inertia, stability across seeds, alternative metric, domain expectation of a few interpretable tiers), and dumps the result without checking that the exported table's shape, row count, and feature columns correspond to the exact vectors that were actually clustered (raw vs. scaled/transformed) and to the ordering/naming the task prescribes.
Detection procedure
  1. Read the task for the required file name, column naming convention, and any implied structure (one row per input record, feature values plus a label).
  2. In the script, find how the group count is chosen: is more than one criterion used, are degenerate/tiny groups rejected, is the choice robust to random seed?
  3. Check what matrix is written out versus what was fed to the clustering algorithm (pre- vs. post-scaling/PCA), and whether column renaming preserves count and order; confirm row count equals the number of input records after any row dropping.
  4. Read the answer for the reported group sizes: if any group has ~1-3 members while others have dozens, or if the reported sizes/feature list are not tied to a sanity check, flag it.
Discriminator
A real violation is a solution accepted solely because one score was highest, with degenerate groups and/or an exported feature representation inconsistent with the clustered one; a look-alike that is fine either justifies the chosen count with two or more converging signals (and notes small groups are genuine, deliberately reported outliers) and exports exactly the vectors clustered, in the required naming/order, with row count matching the input.
Consequence
The saved file's labels (and possibly its feature columns and row count) do not match the reference grouping, so the file-level comparison fails outright even though the pipeline "ran successfully".
id d176e43e9f7a · mined from da-code dacode-ml-cluster-013@s16
raw text (what the judge reads)
### Unvalidated cluster solution written straight to the deliverable
- **Applies when**: `task` -- the task asks for an unsupervised grouping with "an appropriate number of groups" exported to a specified file/column format, and the script picks the group count automatically from a single internal score and immediately writes the result.
- **Pattern**: The attempt maximizes one selection statistic (e.g., silhouette) over a wide range, accepts whatever count wins even when it yields degenerate groups (singletons / handfuls of outliers) and no cross-check (elbow/inertia, stability across seeds, alternative metric, domain expectation of a few interpretable tiers), and dumps the result without checking that the exported table's shape, row count, and feature columns correspond to the exact vectors that were actually clustered (raw vs. scaled/transformed) and to the ordering/naming the task prescribes.
- **Detection procedure**:
  1. Read the task for the required file name, column naming convention, and any implied structure (one row per input record, feature values plus a label).
  2. In the script, find how the group count is chosen: is more than one criterion used, are degenerate/tiny groups rejected, is the choice robust to random seed?
  3. Check what matrix is written out versus what was fed to the clustering algorithm (pre- vs. post-scaling/PCA), and whether column renaming preserves count and order; confirm row count equals the number of input records after any row dropping.
  4. Read the answer for the reported group sizes: if any group has ~1-3 members while others have dozens, or if the reported sizes/feature list are not tied to a sanity check, flag it.
- **Discriminator**: A real violation is a solution accepted solely because one score was highest, with degenerate groups and/or an exported feature representation inconsistent with the clustered one; a look-alike that is fine either justifies the chosen count with two or more converging signals (and notes small groups are genuine, deliberately reported outliers) and exports exactly the vectors clustered, in the required naming/order, with row count matching the input.
- **Consequence**: The saved file's labels (and possibly its feature columns and row count) do not match the reference grouping, so the file-level comparison fails outright even though the pipeline "ran successfully".
849Silently redefining the requested statistic when it appears ill-defined under the loaded data layouttaskinfiagent-dabench
Applies when
task -- the task asks for a statistic over a specific slice/filter (a given year, group, subset), but in the file the agent loaded that slice yields too few values (e.g., one value per entity) to compute the statistic.
Pattern
Instead of resolving the mismatch, the scripts drop or invert the stated filter (aggregating over the whole dimension that was supposed to be fixed, or over the wrong axis), declare the stated constraint "a red herring", and report the winner of this substituted computation without ever checking whether another available file/shape (long format, per-observation records, other regional/global files, additional variables) would make the original request well-defined.
Detection procedure
  1. From the task, write down exactly which dimension is fixed by the constraint and which dimension the statistic is supposed to vary over.
  2. In the scripts, check the array actually passed to the statistic function: does it hold the fixed dimension constant, or has that filter been discarded/replaced by another axis?
  3. Check whether the scripts inventoried the available data (other files in the data directory, reshaping to long form, other columns/units) before concluding the requested computation was impossible; comments like "this must be a red herring" or "each entity has only one value, so use all years instead" are the red flag.
  4. Check that the reported entity is the argmax of the statistic as literally specified, not of the substituted statistic.
Discriminator
A real violation is silent substitution of a different computation while still answering as if the original were computed. It is acceptable if the agent first enumerates and inspects alternative data sources/shapes, documents that the literal reading is genuinely impossible there, and the chosen reinterpretation still respects the stated fixed dimension (e.g., grouping observations within the specified slice) rather than discarding it.
Consequence
The argmax comes from a different distribution than the one asked about, so the named entity mismatches ground truth and the single answer check fails (0/1).
id b6d07fce5918 · mined from infiagent-dabench dabench-252@s16
raw text (what the judge reads)
### Silently redefining the requested statistic when it appears ill-defined under the loaded data layout
- **Applies when**: `task` -- the task asks for a statistic over a specific slice/filter (a given year, group, subset), but in the file the agent loaded that slice yields too few values (e.g., one value per entity) to compute the statistic.
- **Pattern**: Instead of resolving the mismatch, the scripts drop or invert the stated filter (aggregating over the whole dimension that was supposed to be fixed, or over the wrong axis), declare the stated constraint "a red herring", and report the winner of this substituted computation without ever checking whether another available file/shape (long format, per-observation records, other regional/global files, additional variables) would make the original request well-defined.
- **Detection procedure**:
  1. From the task, write down exactly which dimension is fixed by the constraint and which dimension the statistic is supposed to vary over.
  2. In the scripts, check the array actually passed to the statistic function: does it hold the fixed dimension constant, or has that filter been discarded/replaced by another axis?
  3. Check whether the scripts inventoried the available data (other files in the data directory, reshaping to long form, other columns/units) before concluding the requested computation was impossible; comments like "this must be a red herring" or "each entity has only one value, so use all years instead" are the red flag.
  4. Check that the reported entity is the argmax of the statistic as literally specified, not of the substituted statistic.
- **Discriminator**: A real violation is silent substitution of a different computation while still answering as if the original were computed. It is acceptable if the agent first enumerates and inspects alternative data sources/shapes, documents that the literal reading is genuinely impossible there, and the chosen reinterpretation still respects the stated fixed dimension (e.g., grouping observations *within* the specified slice) rather than discarding it.
- **Consequence**: The argmax comes from a different distribution than the one asked about, so the named entity mismatches ground truth and the single answer check fails (0/1).
850Unvalidated two-stage aggregation (mean-of-group-totals) and no check against the provided output templatetaskda-code
Applies when
task -- the task asks for a statistic computed over aggregates (e.g., mean/median of per-period or per-group totals) and specifies an output file that must match a provided sample/template file.
Pattern
The attempt collapses the data in one pass (or groups only on the keys that happen to appear in the data), so the divisor of the second-stage average is the number of observed groups rather than the full intended set of periods/groups, and it writes its own header/column names/rounding/row order instead of copying the schema of the supplied sample output. The result is never sanity-checked (number of period groups per category, plausible value range, row count, header equality), and no reproducible script is retained.
Detection procedure
  1. Read the task: identify the two aggregation levels (inner totals key, outer averaging key), any stated rounding/units/ordering, and the existence of a reference/sample output file.
  2. In the scripts, confirm there are two explicit steps — a groupby producing inner totals, then a mean over those totals — and check how the period/group key is derived (does it cover the complete span, are zero-activity periods represented, are dates parsed rather than string-sliced, are duplicate or filtered rows handled consistently?).
  3. Check that the script loads the sample output file and reproduces its exact column names, column count, row set/order, and numeric precision, rather than inventing labels or rounding.
  4. Check for an assertion/printout validating shape and per-category group counts (e.g., every category divided by the same expected number of periods) and that values fall in a plausible range; absence of any such check, or of a saved script, means the number cannot be trusted.
Discriminator
A real violation is a one-pass average (or a divisor equal to only the periods with data when the task implies all periods) and/or an output whose header/precision/ordering differs from the template; a look-alike that is fine explicitly builds the complete period index (or justifies excluding empty periods), then averages the totals, and asserts equality of its output schema with the sample file.
Consequence
The saved file fails exact comparison with the expected file — either every numeric value is inflated/deflated because of the wrong denominator or aggregation order, or the values are right but the header/rounding/row order mismatch — so the check scores 0.
id 747e4d376413 · mined from da-code dacode-dm-csv-010@s16
raw text (what the judge reads)
### Unvalidated two-stage aggregation (mean-of-group-totals) and no check against the provided output template
- **Applies when**: `task` -- the task asks for a statistic computed over aggregates (e.g., mean/median of per-period or per-group totals) and specifies an output file that must match a provided sample/template file.
- **Pattern**: The attempt collapses the data in one pass (or groups only on the keys that happen to appear in the data), so the divisor of the second-stage average is the number of *observed* groups rather than the full intended set of periods/groups, and it writes its own header/column names/rounding/row order instead of copying the schema of the supplied sample output. The result is never sanity-checked (number of period groups per category, plausible value range, row count, header equality), and no reproducible script is retained.
- **Detection procedure**:
  1. Read the task: identify the two aggregation levels (inner totals key, outer averaging key), any stated rounding/units/ordering, and the existence of a reference/sample output file.
  2. In the scripts, confirm there are two explicit steps — a groupby producing inner totals, then a mean over those totals — and check how the period/group key is derived (does it cover the complete span, are zero-activity periods represented, are dates parsed rather than string-sliced, are duplicate or filtered rows handled consistently?).
  3. Check that the script loads the sample output file and reproduces its exact column names, column count, row set/order, and numeric precision, rather than inventing labels or rounding.
  4. Check for an assertion/printout validating shape and per-category group counts (e.g., every category divided by the same expected number of periods) and that values fall in a plausible range; absence of any such check, or of a saved script, means the number cannot be trusted.
- **Discriminator**: A real violation is a one-pass average (or a divisor equal to only the periods with data when the task implies all periods) and/or an output whose header/precision/ordering differs from the template; a look-alike that is fine explicitly builds the complete period index (or justifies excluding empty periods), then averages the totals, and asserts equality of its output schema with the sample file.
- **Consequence**: The saved file fails exact comparison with the expected file — either every numeric value is inflated/deflated because of the wrong denominator or aggregation order, or the values are right but the header/rounding/row order mismatch — so the check scores 0.
851Loss of identifier precision when coercing a result into a stated format templatetaskinfiagent-dabench
Applies when
task -- the deliverable is an identifier of a specific record (a date, ID, key, category) that must be reported in a stated format string, and the script reformats/truncates the underlying value to produce it.
Pattern
The script finds the correct record but then applies a format template literally in the coarsest possible way (e.g. truncating a full timestamp/key to a shorter field pattern), discarding the granularity that actually identifies the record; the reported answer therefore names a period/group rather than the single row that was found, even though the analysis itself was right.
Detection procedure
  1. Read the task: note the requested output field, its example format, and whether the quantity being identified is inherently a unique row (a single argmax record) or an aggregate group.
  2. Read the script: find where the identifier is stringified/rounded/truncated and compare the resolution of the emitted string with the resolution of the raw value in the data (does the source value carry more distinguishing components than the emitted string?).
  3. Check whether the emitted string still uniquely designates the record found — i.e. whether many rows in the dataset would map to the same output string.
  4. Inspect the final answer: if it is a strictly coarser version of the located record, flag it; an adequate attempt reports the identifier at full source granularity (and may note the format template as ambiguous) rather than silently dropping components.
Discriminator
A real violation is truncation that destroys uniqueness of an argmax/lookup record (many source rows collapse to the reported string). It is not a violation when the task genuinely asks for a group-level result (the analysis aggregated to that granularity first), or when the discarded components are constant/meaningless in the data, or when a stated rounding rule applies to a numeric measure rather than to a record key.
Consequence
The identifier check fails on exact-match comparison against the full-resolution ground-truth key, so the submission is marked wrong even though dependent numeric answers computed from the correct row match.
id 15a125803310 · mined from infiagent-dabench dabench-572@s16
raw text (what the judge reads)
### Loss of identifier precision when coercing a result into a stated format template
- **Applies when**: `task` -- the deliverable is an identifier of a specific record (a date, ID, key, category) that must be reported in a stated format string, and the script reformats/truncates the underlying value to produce it.
- **Pattern**: The script finds the correct record but then applies a format template literally in the coarsest possible way (e.g. truncating a full timestamp/key to a shorter field pattern), discarding the granularity that actually identifies the record; the reported answer therefore names a period/group rather than the single row that was found, even though the analysis itself was right.
- **Detection procedure**:
  1. Read the task: note the requested output field, its example format, and whether the quantity being identified is inherently a unique row (a single argmax record) or an aggregate group.
  2. Read the script: find where the identifier is stringified/rounded/truncated and compare the resolution of the emitted string with the resolution of the raw value in the data (does the source value carry more distinguishing components than the emitted string?).
  3. Check whether the emitted string still uniquely designates the record found — i.e. whether many rows in the dataset would map to the same output string.
  4. Inspect the final answer: if it is a strictly coarser version of the located record, flag it; an adequate attempt reports the identifier at full source granularity (and may note the format template as ambiguous) rather than silently dropping components.
- **Discriminator**: A real violation is truncation that destroys uniqueness of an argmax/lookup record (many source rows collapse to the reported string). It is *not* a violation when the task genuinely asks for a group-level result (the analysis aggregated to that granularity first), or when the discarded components are constant/meaningless in the data, or when a stated rounding rule applies to a numeric measure rather than to a record key.
- **Consequence**: The identifier check fails on exact-match comparison against the full-resolution ground-truth key, so the submission is marked wrong even though dependent numeric answers computed from the correct row match.
852Fabricating input data instead of loading the provided datasettaskda-code
Applies when
task -- the task references a provided dataset/course data and asks for a statistic or model output derived from it.
Pattern
The script hard-codes synthetic/"realistic" arrays (or simulates values with a random seed or an exact linear formula) rather than reading any data file, so the reported statistic reflects invented numbers; a degenerate result (e.g., exactly 0 error) is then accepted without question.
Detection procedure
1. Read the task and note that the quantity must come from the supplied data. 2. Scan the scripts for any file-reading call (read_csv, read_excel, loading a package dataset, an API/download); if none exists and values appear as literal arrays or np.random/closed-form constructions, flag it. 3. Check whether the agent ever searched the working directory/README for the actual data source or tried an alternative when data was missing. 4. Inspect the answer for tell-tale degenerate values (perfect fit, zero residuals, R²=1) that indicate self-generated data.
Discriminator
A real violation is when the analysis inputs are invented; it is fine if literals are only parameters, thresholds, or a reproducibility check, while the actual observations are loaded from the provided source (or a documented, verifiable public source with values cross-checked).
Consequence
The reported number is unrelated to the ground-truth data, so the expected output file fails the value comparison (0/1 checks passed) even if the file name and column header are correct.
id dfbf4b80d58f · mined from da-code dacode-data-sa-043@s16
raw text (what the judge reads)
### Fabricating input data instead of loading the provided dataset
- **Applies when**: `task` -- the task references a provided dataset/course data and asks for a statistic or model output derived from it.
- **Pattern**: The script hard-codes synthetic/"realistic" arrays (or simulates values with a random seed or an exact linear formula) rather than reading any data file, so the reported statistic reflects invented numbers; a degenerate result (e.g., exactly 0 error) is then accepted without question.
- **Detection procedure**: 1. Read the task and note that the quantity must come from the supplied data. 2. Scan the scripts for any file-reading call (`read_csv`, `read_excel`, loading a package dataset, an API/download); if none exists and values appear as literal arrays or `np.random`/closed-form constructions, flag it. 3. Check whether the agent ever searched the working directory/README for the actual data source or tried an alternative when data was missing. 4. Inspect the answer for tell-tale degenerate values (perfect fit, zero residuals, R²=1) that indicate self-generated data.
- **Discriminator**: A real violation is when the *analysis inputs* are invented; it is fine if literals are only parameters, thresholds, or a reproducibility check, while the actual observations are loaded from the provided source (or a documented, verifiable public source with values cross-checked).
- **Consequence**: The reported number is unrelated to the ground-truth data, so the expected output file fails the value comparison (0/1 checks passed) even if the file name and column header are correct.
853Final answer literal doesn't match the requested answer syntax (and no reproducible script backs it)taskinfiagent-dabench
Applies when
task -- the task prescribes an exact answer token/format (e.g., @name[list_of_strings], a number with fixed rounding, a comma-separated list) and the agent must emit that literal as its final output.
Pattern
The agent computes a plausible (even correct) result but hand-writes the final answer with decorations the format spec never asked for — quoted strings, nested brackets, extra whitespace, JSON-style punctuation, different ordering/separator — and leaves no saved script that generates the answer string, so the emitted literal is never checked against the spec.
Detection procedure
  1. Read the task and copy out the exact answer template, including delimiters, quoting, separators, and any ordering/rounding rules.
  2. Read the scripts: confirm a script exists, is saved, and prints the final answer string in exactly that template (rather than the answer being typed by hand from console output).
  3. Character-by-character compare the agent's submitted literal to the template: check delimiter type, presence/absence of quotes around list items, separator (, vs ,), stray brackets, and label spelling/case.
  4. If any element of the literal is not explicitly sanctioned by the template, or if no script reproduces the literal, flag the attempt.
Discriminator
A real violation is a mismatch in the serialization of the answer (added quotes/brackets, wrong separator, wrong label, unsaved/unreproducible derivation), even when the underlying values are right. A look-alike that is fine is a literal that differs only in ways the task explicitly permits (e.g., ordering when order is stated as irrelevant, or a quoting style the template itself shows).
Consequence
The grader parses the field and marks it WRONG/MISSING despite substantively correct analysis, yielding 0/1 checks passed, and the missing script makes the result unverifiable on review.
id 4f86825fd22f · mined from infiagent-dabench dabench-254@s16
raw text (what the judge reads)
### Final answer literal doesn't match the requested answer syntax (and no reproducible script backs it)
- **Applies when**: `task` -- the task prescribes an exact answer token/format (e.g., `@name[list_of_strings]`, a number with fixed rounding, a comma-separated list) and the agent must emit that literal as its final output.
- **Pattern**: The agent computes a plausible (even correct) result but hand-writes the final answer with decorations the format spec never asked for — quoted strings, nested brackets, extra whitespace, JSON-style punctuation, different ordering/separator — and leaves no saved script that generates the answer string, so the emitted literal is never checked against the spec.
- **Detection procedure**:
  1. Read the task and copy out the exact answer template, including delimiters, quoting, separators, and any ordering/rounding rules.
  2. Read the scripts: confirm a script exists, is saved, and prints the final answer string in exactly that template (rather than the answer being typed by hand from console output).
  3. Character-by-character compare the agent's submitted literal to the template: check delimiter type, presence/absence of quotes around list items, separator (`, ` vs `,`), stray brackets, and label spelling/case.
  4. If any element of the literal is not explicitly sanctioned by the template, or if no script reproduces the literal, flag the attempt.
- **Discriminator**: A real violation is a mismatch in the *serialization* of the answer (added quotes/brackets, wrong separator, wrong label, unsaved/unreproducible derivation), even when the underlying values are right. A look-alike that is fine is a literal that differs only in ways the task explicitly permits (e.g., ordering when order is stated as irrelevant, or a quoting style the template itself shows).
- **Consequence**: The grader parses the field and marks it WRONG/MISSING despite substantively correct analysis, yielding 0/1 checks passed, and the missing script makes the result unverifiable on review.
854Ships predictions with no held-out validation and no alignment/format sanity checktaskda-code
Applies when
task -- the script fits a model on a training file and writes a prediction file for a separate test file that must match the test rows in count, order, and label format.
Pattern
The script trains, predicts, and writes the output in one pass: it never scores the model on a held-out split (so accuracy is unknown), and it never asserts that the number of written rows equals the number of test rows, that the row order corresponds to the test file's original order, or that the emitted label strings/values match the exact vocabulary/format expected. Extra risk signals: train and test feature frames are concatenated and re-split by position without checking that both have identical column names and order (silently introducing all-NaN or reordered columns), preprocessing/encoders/scalers are fit on train+test together, and an ID column is read/dropped inconsistently between the two files.
Detection procedure
  1. Read the task for the required output: file name, column name, one row per test record, label spelling, and any ordering constraint.
  2. In the scripts, look for (a) any train/validation split or cross-validation with a printed score, and (b) any assertion/print comparing len(predictions) to len(test) and comparing the processed train vs test column lists.
  3. Trace how the test features are built: if train and test are concatenated then split by position, check that both frames were verified to have the same columns (after dropping the target/ID) and the same row order as the original test file.
  4. Inspect the produced file: does it have exactly the test-row count, the required header, and only the exact expected label values?
Discriminator
A fine attempt may skip a fancy ensemble but still prints a validation score and explicit shape/column/label checks (or uses a single aligned X_test derived directly from the test file with the same fitted transformer). A violation is an attempt where model quality and output alignment are both unverified — the agent could not have known whether the file has the right rows in the right order at all.
Consequence
The saved file can be mis-sized, mis-ordered, mis-labeled, or built from silently corrupted/misaligned features, so the grader's row-wise comparison to the ground truth fails (wrong/missing result file) even though the modeling code "ran successfully."
id dbb287cd5185 · mined from da-code dacode-ml-binary-009@s16
raw text (what the judge reads)
### Ships predictions with no held-out validation and no alignment/format sanity check

- **Applies when**: `task` -- the script fits a model on a training file and writes a prediction file for a separate test file that must match the test rows in count, order, and label format.
- **Pattern**: The script trains, predicts, and writes the output in one pass: it never scores the model on a held-out split (so accuracy is unknown), and it never asserts that the number of written rows equals the number of test rows, that the row order corresponds to the test file's original order, or that the emitted label strings/values match the exact vocabulary/format expected. Extra risk signals: train and test feature frames are concatenated and re-split by position without checking that both have identical column names and order (silently introducing all-NaN or reordered columns), preprocessing/encoders/scalers are fit on train+test together, and an ID column is read/dropped inconsistently between the two files.
- **Detection procedure**:
  1. Read the task for the required output: file name, column name, one row per test record, label spelling, and any ordering constraint.
  2. In the scripts, look for (a) any train/validation split or cross-validation with a printed score, and (b) any assertion/print comparing `len(predictions)` to `len(test)` and comparing the processed train vs test column lists.
  3. Trace how the test features are built: if train and test are concatenated then split by position, check that both frames were verified to have the same columns (after dropping the target/ID) and the same row order as the original test file.
  4. Inspect the produced file: does it have exactly the test-row count, the required header, and only the exact expected label values?
- **Discriminator**: A fine attempt may skip a fancy ensemble but still prints a validation score and explicit shape/column/label checks (or uses a single aligned `X_test` derived directly from the test file with the same fitted transformer). A violation is an attempt where model quality and output alignment are both unverified — the agent could not have known whether the file has the right rows in the right order at all.
- **Consequence**: The saved file can be mis-sized, mis-ordered, mis-labeled, or built from silently corrupted/misaligned features, so the grader's row-wise comparison to the ground truth fails (wrong/missing result file) even though the modeling code "ran successfully."
855Aggregating an entity's statistics from only one of the columns where it appears (with dtype‑fragile filters and no sanity check)taskda-code
Applies when
task -- the script computes a per-entity metric from a table where the same entity can appear in several columns/roles (e.g. two participant columns, or a flag column whose stored dtype is uncertain).
Pattern
The agent groups by a single role column, so every record where the entity occupies the other role is silently dropped, and/or filters a flag column by comparing to a hard-coded literal (e.g. == 'TRUE') without checking the parsed dtype; the resulting counts are then plugged into a composite score and never cross-checked against the metric definition or any independent count.
Detection procedure
  1. Read the task/spec files for the exact definition of the requested metric (which records count, which roles/subsets are included, how ties/order are decided) and which output artifacts are required.
  2. In the script, list every groupby/filter key and check whether an entity's records can also live in a column that is never touched, and whether any equality test is made against a literal whose dtype may differ after parsing (bool vs string, int vs str, NaN handling).
  3. Look for a validation step: printed per-entity totals compared with a direct independent count, a check that the flag filter returns a non-zero/plausible number of rows, and a check that the entity set/ordering matches the spec.
  4. Check the answer: are the reported numbers plausible given the total record count, and were all required output files (chart plus any numeric/JSON artifacts) written?
Discriminator
A real violation is when the metric definition covers all roles/records but the code only sees a subset (or a filter that can evaluate to all-False), and no counter-check is run. It is fine if the spec genuinely restricts the metric to one role/subset and the script prints evidence (row counts, flag value counts) confirming the filter matched the intended rows.
Consequence
The per-entity values are systematically too small and the ranking/selection of entities is wrong, so the saved chart and any derived numeric artifacts mismatch the expected values and all file-level checks fail.
id 6ac4df6a43b8 · mined from da-code dacode-plot-bar-006@s16
raw text (what the judge reads)
### Aggregating an entity's statistics from only one of the columns where it appears (with dtype‑fragile filters and no sanity check)

- **Applies when**: `task` -- the script computes a per-entity metric from a table where the same entity can appear in several columns/roles (e.g. two participant columns, or a flag column whose stored dtype is uncertain).
- **Pattern**: The agent groups by a single role column, so every record where the entity occupies the other role is silently dropped, and/or filters a flag column by comparing to a hard-coded literal (e.g. `== 'TRUE'`) without checking the parsed dtype; the resulting counts are then plugged into a composite score and never cross-checked against the metric definition or any independent count.
- **Detection procedure**:
  1. Read the task/spec files for the exact definition of the requested metric (which records count, which roles/subsets are included, how ties/order are decided) and which output artifacts are required.
  2. In the script, list every `groupby`/filter key and check whether an entity's records can also live in a column that is never touched, and whether any equality test is made against a literal whose dtype may differ after parsing (bool vs string, int vs str, NaN handling).
  3. Look for a validation step: printed per-entity totals compared with a direct independent count, a check that the flag filter returns a non-zero/plausible number of rows, and a check that the entity set/ordering matches the spec.
  4. Check the answer: are the reported numbers plausible given the total record count, and were all required output files (chart plus any numeric/JSON artifacts) written?
- **Discriminator**: A real violation is when the metric definition covers all roles/records but the code only sees a subset (or a filter that can evaluate to all-False), and no counter-check is run. It is fine if the spec genuinely restricts the metric to one role/subset and the script prints evidence (row counts, flag value counts) confirming the filter matched the intended rows.
- **Consequence**: The per-entity values are systematically too small and the ranking/selection of entities is wrong, so the saved chart and any derived numeric artifacts mismatch the expected values and all file-level checks fail.
856Hardcoding a mapping/spec instead of reading the provided reference documenttaskda-code
Applies when
task -- the task says to use a definition, label mapping, filter, or rule that lives in an accompanying file (notes/readme/tips/config), and the scripts recode or aggregate values based on that rule.
Pattern
The script never opens or prints the referenced document; the agent invents the rule from domain intuition (e.g. plausible-sounding expanded label names, an assumed category order, an assumed threshold), then reports a result whose category names/values come from that guess rather than the specified source — and often only prints the result instead of writing it to the required output artifact.
Detection procedure
  1. Read the task and list every external artifact it points to and what it is supposed to supply (mapping strings, filters, rounding, units).
  2. Scan the scripts for any read/open/parse of that artifact; check whether the values used (dict literals, thresholds, category names) are traceable to it or are literals typed in by the agent.
  3. Check that each required output value is produced exactly as specified (naming, rounding, decimal vs percent) and that the requested file/format artifact is actually written, not just printed.
  4. If the rule source was never read, treat the reported labels/values as unverified even if the numeric aggregation logic looks correct.
Discriminator
Fine if the script loads the reference file (or the script quotes its contents verbatim after inspecting it) and the literals demonstrably match; a violation is when the literals are the agent's plausible reconstruction and no step in the pipeline ever validates them against the file — a correct count/ratio does not excuse wrong label text or a missing result file.
Consequence
The numeric statistic may match while the reported category string (or output artifact) does not, so the exact-match check on the expected result file fails, scoring 0.
id 84c064d51d01 · mined from da-code dacode-di-text-004@s16
raw text (what the judge reads)
### Hardcoding a mapping/spec instead of reading the provided reference document
- **Applies when**: `task` -- the task says to use a definition, label mapping, filter, or rule that lives in an accompanying file (notes/readme/tips/config), and the scripts recode or aggregate values based on that rule.
- **Pattern**: The script never opens or prints the referenced document; the agent invents the rule from domain intuition (e.g. plausible-sounding expanded label names, an assumed category order, an assumed threshold), then reports a result whose category names/values come from that guess rather than the specified source — and often only prints the result instead of writing it to the required output artifact.
- **Detection procedure**:
  1. Read the task and list every external artifact it points to and what it is supposed to supply (mapping strings, filters, rounding, units).
  2. Scan the scripts for any read/open/parse of that artifact; check whether the values used (dict literals, thresholds, category names) are traceable to it or are literals typed in by the agent.
  3. Check that each required output value is produced exactly as specified (naming, rounding, decimal vs percent) and that the requested file/format artifact is actually written, not just printed.
  4. If the rule source was never read, treat the reported labels/values as unverified even if the numeric aggregation logic looks correct.
- **Discriminator**: Fine if the script loads the reference file (or the script quotes its contents verbatim after inspecting it) and the literals demonstrably match; a violation is when the literals are the agent's plausible reconstruction and no step in the pipeline ever validates them against the file — a correct count/ratio does not excuse wrong label text or a missing result file.
- **Consequence**: The numeric statistic may match while the reported category string (or output artifact) does not, so the exact-match check on the expected result file fails, scoring 0.
857Predictions not aligned to the designated evaluation split / required output schemataskda-code
Applies when
task -- the task names a specific input file whose rows must be predicted and a specific output file/column format, and the scripts build predictions from whatever files they find.
Pattern
The attempt ignores the designated prediction input, generating labels for a different (usually much larger, all-rows) set, and/or emits extra or differently-named columns; it also substitutes hand-set rules for a model fit on the labeled training rows, without ever checking its output row count/IDs/label vocabulary against the evaluation file or the label distribution in the training labels.
Detection procedure
  1. From the task, note the exact evaluation input file, the expected number of prediction rows, the required output filename, and the required column name(s).
  2. In the scripts, confirm the prediction frame is derived by reading that evaluation file (or joined onto its ID list) and that the written file contains exactly the required column(s), in the required order, with the same row count and ID set.
  3. Check that predicted labels come from the observed label set of the labeled training data (exact strings/casing) and that at least one comparison against held-out labeled rows (accuracy/F1 or distribution match) was reported.
  4. In the answer, compare the reported row count and class distribution to the evaluation file size and the training label distribution; large mismatch means the wrong subset or an unvalidated rule.
Discriminator
A real violation is when row count/IDs/column names/label strings differ from what the evaluation file and label vocabulary require, or no accuracy check on labeled data exists; it is fine if the script predicts on a superset internally but then subsets/reindexes to the evaluation IDs before writing, and reports a validated score.
Consequence
The grader cannot join or score the submission (missing/extra rows, wrong IDs, wrong column or label spelling) and marks the result file WRONG/MISSING, or the score is near chance because unfit heuristics produce a distribution unlike the true labels.
id 6ff2943b8c3b · mined from da-code dacode-ml-multi-003@s16
raw text (what the judge reads)
### Predictions not aligned to the designated evaluation split / required output schema
- **Applies when**: `task` -- the task names a specific input file whose rows must be predicted and a specific output file/column format, and the scripts build predictions from whatever files they find.
- **Pattern**: The attempt ignores the designated prediction input, generating labels for a different (usually much larger, all-rows) set, and/or emits extra or differently-named columns; it also substitutes hand-set rules for a model fit on the labeled training rows, without ever checking its output row count/IDs/label vocabulary against the evaluation file or the label distribution in the training labels.
- **Detection procedure**:
  1. From the task, note the exact evaluation input file, the expected number of prediction rows, the required output filename, and the required column name(s).
  2. In the scripts, confirm the prediction frame is derived by reading that evaluation file (or joined onto its ID list) and that the written file contains exactly the required column(s), in the required order, with the same row count and ID set.
  3. Check that predicted labels come from the observed label set of the labeled training data (exact strings/casing) and that at least one comparison against held-out labeled rows (accuracy/F1 or distribution match) was reported.
  4. In the answer, compare the reported row count and class distribution to the evaluation file size and the training label distribution; large mismatch means the wrong subset or an unvalidated rule.
- **Discriminator**: A real violation is when row count/IDs/column names/label strings differ from what the evaluation file and label vocabulary require, or no accuracy check on labeled data exists; it is fine if the script predicts on a superset internally but then subsets/reindexes to the evaluation IDs before writing, and reports a validated score.
- **Consequence**: The grader cannot join or score the submission (missing/extra rows, wrong IDs, wrong column or label spelling) and marks the result file WRONG/MISSING, or the score is near chance because unfit heuristics produce a distribution unlike the true labels.
858Model selection/validation uses a proxy metric instead of the metric stated in the tasktaskda-code
Applies when
task -- the task names a specific evaluation metric (e.g. an agreement, ranking, or weighted/ordinal error measure) and the scripts train several candidate models and pick or blend them based on cross-validation scores.
Pattern
The scripts compute and compare only a convenient default score (plain accuracy, default score(), MSE) and choose hyperparameters, ensemble weights, and the final prediction rule by maximizing that proxy, never once computing the metric the task will be graded on, and never handling the metric's special structure (ordinal/weighted penalties, class imbalance, tie-breaking, thresholding of continuous predictions).
Detection procedure
  1. Read the task/README and write down the exact grading metric and any structure it implies (ordered targets, weighting, imbalance sensitivity).
  2. Grep the scripts for the scoring used in cross_val_score/scoring=/manual evaluation and for any implementation or import of the stated metric.
  3. Check how the final prediction is formed (argmax of blended probabilities, fixed weights, rounding rule) and whether those choices were validated against the stated metric on held-out data.
  4. Inspect the produced prediction distribution versus the training target distribution: if the proxy-optimized model collapses onto the majority classes while the true metric rewards spreading into extreme classes, flag it.
Discriminator
A real violation is when the graded metric is never computed anywhere, so all model/ensemble/decision-rule choices rest on an unvalidated proxy. It is fine if the agent computes the stated metric (or a provably monotone-equivalent surrogate) at least once on held-out data to justify the final choice, even if a proxy is also reported.
Consequence
The submission is well-formed and passes shape/ID checks but scores far below an equivalent model tuned on the real metric — predictions concentrate on frequent classes, the agreement/weighted score is low, and the graded comparison against the expected result fails.
id 6fce5adf8f2d · mined from da-code dacode-ml-competition-006@s16
raw text (what the judge reads)
### Model selection/validation uses a proxy metric instead of the metric stated in the task
- **Applies when**: `task` -- the task names a specific evaluation metric (e.g. an agreement, ranking, or weighted/ordinal error measure) and the scripts train several candidate models and pick or blend them based on cross-validation scores.
- **Pattern**: The scripts compute and compare only a convenient default score (plain accuracy, default `score()`, MSE) and choose hyperparameters, ensemble weights, and the final prediction rule by maximizing that proxy, never once computing the metric the task will be graded on, and never handling the metric's special structure (ordinal/weighted penalties, class imbalance, tie-breaking, thresholding of continuous predictions).
- **Detection procedure**:
  1. Read the task/README and write down the exact grading metric and any structure it implies (ordered targets, weighting, imbalance sensitivity).
  2. Grep the scripts for the scoring used in `cross_val_score`/`scoring=`/manual evaluation and for any implementation or import of the stated metric.
  3. Check how the final prediction is formed (argmax of blended probabilities, fixed weights, rounding rule) and whether those choices were validated against the stated metric on held-out data.
  4. Inspect the produced prediction distribution versus the training target distribution: if the proxy-optimized model collapses onto the majority classes while the true metric rewards spreading into extreme classes, flag it.
- **Discriminator**: A real violation is when the graded metric is never computed anywhere, so all model/ensemble/decision-rule choices rest on an unvalidated proxy. It is fine if the agent computes the stated metric (or a provably monotone-equivalent surrogate) at least once on held-out data to justify the final choice, even if a proxy is also reported.
- **Consequence**: The submission is well-formed and passes shape/ID checks but scores far below an equivalent model tuned on the real metric — predictions concentrate on frequent classes, the agreement/weighted score is low, and the graded comparison against the expected result fails.
859Null/group-mask definition and numeric coercion are not validated before computing group statisticstaskinfiagent-dabench
Applies when
task -- the analysis splits rows into groups by whether a column is missing (or by some other boolean condition) and then compares a numeric column's aggregate/statistic across those groups.
Pattern
The attempt takes the loader's default notion of "missing" (e.g., isna() on a column that also contains empty strings, "NA", "None", "-", whitespace, or nested/empty containers) and/or lets the numeric column be read as object/string with values silently dropped or coerced, so both group sizes and group means are subtly shifted. No script or answer step confirms the mask semantics, the row counts per group, or that the numeric column is fully parseable.
Detection procedure
  1. Read the task to see exactly how the grouping condition and the target statistic are defined (which column defines groups, which column is aggregated).
  2. In the scripts, find how the grouping mask is built and how the aggregated column is typed: is there any inspection of the raw distinct values / sentinel strings in the grouping column, any explicit handling of empty-string/placeholder missingness, and any to_numeric with an error check on the aggregated column?
  3. Check whether the scripts print and the answer is cross-checked against group row counts (n per group) and total rows = n1 + n2, plus a value-range sanity check on each mean.
  4. If the mask is taken on faith and no counts/dtype validation is printed, flag as inadequate — especially if the reported means are close to but not exactly reproducible from a stated, auditable row partition.
Discriminator
Fine if the scripts explicitly enumerate the grouping column's raw values (or show that only true nulls exist), coerce the aggregated column with an error/NaN count check, and report n per group summing to the dataset size; a real violation is when "missing" is assumed to equal the default isna() result and no counts/dtype diagnostics exist to catch mislabeled rows.
Consequence
Group means (and often the test statistic) are computed on slightly wrong row sets, so the reported numbers differ from ground truth by a few percent and fail exact-value checks even though the p-value verdict looks plausible.
id eedb75d488bd · mined from infiagent-dabench dabench-297@s16
raw text (what the judge reads)
### Null/group-mask definition and numeric coercion are not validated before computing group statistics
- **Applies when**: `task` -- the analysis splits rows into groups by whether a column is missing (or by some other boolean condition) and then compares a numeric column's aggregate/statistic across those groups.
- **Pattern**: The attempt takes the loader's default notion of "missing" (e.g., `isna()` on a column that also contains empty strings, `"NA"`, `"None"`, `"-"`, whitespace, or nested/empty containers) and/or lets the numeric column be read as object/string with values silently dropped or coerced, so both group sizes and group means are subtly shifted. No script or answer step confirms the mask semantics, the row counts per group, or that the numeric column is fully parseable.
- **Detection procedure**:
  1. Read the task to see exactly how the grouping condition and the target statistic are defined (which column defines groups, which column is aggregated).
  2. In the scripts, find how the grouping mask is built and how the aggregated column is typed: is there any inspection of the raw distinct values / sentinel strings in the grouping column, any explicit handling of empty-string/placeholder missingness, and any `to_numeric` with an error check on the aggregated column?
  3. Check whether the scripts print and the answer is cross-checked against group row counts (n per group) and total rows = n1 + n2, plus a value-range sanity check on each mean.
  4. If the mask is taken on faith and no counts/dtype validation is printed, flag as inadequate — especially if the reported means are close to but not exactly reproducible from a stated, auditable row partition.
- **Discriminator**: Fine if the scripts explicitly enumerate the grouping column's raw values (or show that only true nulls exist), coerce the aggregated column with an error/NaN count check, and report n per group summing to the dataset size; a real violation is when "missing" is assumed to equal the default `isna()` result and no counts/dtype diagnostics exist to catch mislabeled rows.
- **Consequence**: Group means (and often the test statistic) are computed on slightly wrong row sets, so the reported numbers differ from ground truth by a few percent and fail exact-value checks even though the p-value verdict looks plausible.
860Incomplete set of required output artifactstaskda-code
Applies when
task -- the task (or a referenced guidance/spec file) implies several deliverables — e.g. a saved figure plus serialized numeric/plot data — that a grader will check as files.
Pattern
The script produces only the most visible artifact (the image) and reports the remaining numbers in prose, never writing the other required files to the expected paths/names/formats; the agent also never re-reads the spec to enumerate deliverables.
Detection procedure
1. From the task text and any referenced guidance file, list every artifact required (filename, format, location) and every stated constraint (figure size, colors, legend, ordering). 2. Grep the script for write operations (savefig, np.save, json.dump, to_csv) and collect the exact paths written. 3. Diff the two lists; also confirm the written paths match the expected directory and extension. 4. Check the final answer: does it merely narrate numbers that should have been serialized?
Discriminator
A real violation is a required artifact with no corresponding write call (or written under a different name/dir/format). Not a violation if all required files are written and extra intermediates or extra prose are merely additional.
Consequence
The grader reports missing/wrong files for the unwritten artifacts, failing those checks regardless of whether the computed numbers were right.
id ab6a0e853622 · mined from da-code dacode-plot-pie-005@s16
raw text (what the judge reads)
### Incomplete set of required output artifacts
- **Applies when**: `task` -- the task (or a referenced guidance/spec file) implies several deliverables — e.g. a saved figure plus serialized numeric/plot data — that a grader will check as files.
- **Pattern**: The script produces only the most visible artifact (the image) and reports the remaining numbers in prose, never writing the other required files to the expected paths/names/formats; the agent also never re-reads the spec to enumerate deliverables.
- **Detection procedure**: 1. From the task text and any referenced guidance file, list every artifact required (filename, format, location) and every stated constraint (figure size, colors, legend, ordering). 2. Grep the script for write operations (`savefig`, `np.save`, `json.dump`, `to_csv`) and collect the exact paths written. 3. Diff the two lists; also confirm the written paths match the expected directory and extension. 4. Check the final answer: does it merely narrate numbers that should have been serialized?
- **Discriminator**: A real violation is a required artifact with no corresponding write call (or written under a different name/dir/format). Not a violation if all required files are written and extra intermediates or extra prose are merely additional.
- **Consequence**: The grader reports missing/wrong files for the unwritten artifacts, failing those checks regardless of whether the computed numbers were right.
861Reported error metric never sanity-checked against a trivial baseline (target variance)taskinfiagent-dabench
Applies when
task -- a regression task asks for a held-out error metric (MSE/RMSE/MAE) after prescribed preprocessing (e.g., mean-imputation of specific columns) and a fixed train/test split.
Pattern
The attempt reports a raw metric value straight from the fitted model without ever comparing it to the variance of the target (i.e., the MSE a constant mean-predictor would achieve), so a metric that is several times worse than predicting the mean — a near-certain sign of broken preprocessing (units/dtype not coerced to numeric before imputation, imputation applied to the wrong columns or after row-dropping, features/target misaligned, outliers left in a column stored as text, or fitting on the wrong subset) — is submitted as the answer. Frequently there is also no saved script, so the preprocessing chain cannot be audited.
Detection procedure
  1. From the task, identify the target column, the required preprocessing steps, and the requested metric.
  2. In the scripts, confirm each named column is coerced to a numeric dtype before imputation, that imputation replaces (not drops) missing entries so row counts are unchanged, and that X and y come from the same rows of the same split.
  3. Check whether the script prints any baseline comparison or diagnostic (target variance / std, MSE of a mean-predictor, R², shapes and NaN counts after preprocessing) alongside the reported metric.
  4. Compare the submitted metric to the plausible spread of the target: if implied RMSE is comparable to or larger than the target's standard deviation, treat the pipeline as unverified.
Discriminator
A genuine violation is an unaudited pipeline whose error implies the model is no better than (or worse than) the mean predictor, with no diagnostic output to justify it; it is not a violation if the script prints shape/dtype/NaN and baseline-variance checks and the error is plausibly below the target's variance, even if the model is weak — nor if a legitimately noisy relationship yields a high but still sub-baseline error.
Consequence
The grader compares against the reference value computed with correctly typed, fully imputed data; a silently corrupted feature or target column inflates MSE by an order of magnitude and the submitted number fails the exact-match check.
id e10d59f551ee · mined from infiagent-dabench dabench-432@s16
raw text (what the judge reads)
### Reported error metric never sanity-checked against a trivial baseline (target variance)
- **Applies when**: `task` -- a regression task asks for a held-out error metric (MSE/RMSE/MAE) after prescribed preprocessing (e.g., mean-imputation of specific columns) and a fixed train/test split.
- **Pattern**: The attempt reports a raw metric value straight from the fitted model without ever comparing it to the variance of the target (i.e., the MSE a constant mean-predictor would achieve), so a metric that is several times *worse* than predicting the mean — a near-certain sign of broken preprocessing (units/dtype not coerced to numeric before imputation, imputation applied to the wrong columns or after row-dropping, features/target misaligned, outliers left in a column stored as text, or fitting on the wrong subset) — is submitted as the answer. Frequently there is also no saved script, so the preprocessing chain cannot be audited.
- **Detection procedure**:
  1. From the task, identify the target column, the required preprocessing steps, and the requested metric.
  2. In the scripts, confirm each named column is coerced to a numeric dtype *before* imputation, that imputation replaces (not drops) missing entries so row counts are unchanged, and that X and y come from the same rows of the same split.
  3. Check whether the script prints any baseline comparison or diagnostic (target variance / std, `MSE` of a mean-predictor, R², shapes and NaN counts after preprocessing) alongside the reported metric.
  4. Compare the submitted metric to the plausible spread of the target: if implied RMSE is comparable to or larger than the target's standard deviation, treat the pipeline as unverified.
- **Discriminator**: A genuine violation is an unaudited pipeline whose error implies the model is no better than (or worse than) the mean predictor, with no diagnostic output to justify it; it is *not* a violation if the script prints shape/dtype/NaN and baseline-variance checks and the error is plausibly below the target's variance, even if the model is weak — nor if a legitimately noisy relationship yields a high but still sub-baseline error.
- **Consequence**: The grader compares against the reference value computed with correctly typed, fully imputed data; a silently corrupted feature or target column inflates MSE by an order of magnitude and the submitted number fails the exact-match check.
862Ignoring a referenced instruction/spec file that defines the required output schemataskda-code
Applies when
task -- the prompt says to follow an auxiliary instructions file (e.g. tips.md, README, spec) and/or enumerates the exact fields a single-row result file must contain.
Pattern
The agent never opens/reads the referenced file, invents its own column names, adds extra or omits required fields, and formats values by its own convention (e.g. truncating a tiny p-value to 0.000000), so the output cannot match the expected schema even if the underlying statistic is right.
Detection procedure
  1. From the task text, list every referenced instruction file and every explicitly named output field/format constraint (field set, one row, rounding/precision, ordering).
  2. Scan the scripts for any read of the referenced file (or a printed quote of its contents); if absent, the agent's schema is unverified guesswork.
  3. Compare the written file's header/values against the task's enumerated fields: flag extra columns, renamed columns, missing required ones, and values whose precision/wording is chosen arbitrarily rather than from the spec.
  4. Sanity-check numeric formatting: a reported value that collapses to 0.000000 or otherwise loses all information signals a formatting choice not derived from the spec.
Discriminator
A real violation is when no evidence exists that the agent inspected the authoritative format source and the emitted schema deviates from the fields listed in the task; it is fine if the agent read the spec (or the task fully specifies the schema) and the output matches it exactly, even if extra descriptive text differs in wording.
Consequence
The result file is marked WRONG/MISSING on exact-match or field-wise comparison despite a statistically plausible test, yielding 0 checks passed.
id facffb544e01 · mined from da-code dacode-data-sa-004@s16
raw text (what the judge reads)
### Ignoring a referenced instruction/spec file that defines the required output schema
- **Applies when**: `task` -- the prompt says to follow an auxiliary instructions file (e.g. `tips.md`, `README`, `spec`) and/or enumerates the exact fields a single-row result file must contain.
- **Pattern**: The agent never opens/reads the referenced file, invents its own column names, adds extra or omits required fields, and formats values by its own convention (e.g. truncating a tiny p-value to `0.000000`), so the output cannot match the expected schema even if the underlying statistic is right.
- **Detection procedure**:
  1. From the task text, list every referenced instruction file and every explicitly named output field/format constraint (field set, one row, rounding/precision, ordering).
  2. Scan the scripts for any read of the referenced file (or a printed quote of its contents); if absent, the agent's schema is unverified guesswork.
  3. Compare the written file's header/values against the task's enumerated fields: flag extra columns, renamed columns, missing required ones, and values whose precision/wording is chosen arbitrarily rather than from the spec.
  4. Sanity-check numeric formatting: a reported value that collapses to `0.000000` or otherwise loses all information signals a formatting choice not derived from the spec.
- **Discriminator**: A real violation is when no evidence exists that the agent inspected the authoritative format source and the emitted schema deviates from the fields listed in the task; it is fine if the agent read the spec (or the task fully specifies the schema) and the output matches it exactly, even if extra descriptive text differs in wording.
- **Consequence**: The result file is marked WRONG/MISSING on exact-match or field-wise comparison despite a statistically plausible test, yielding 0 checks passed.
863Unverified row-selection / missing-value handling for a precision-sensitive statistictaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic reported at a fixed rounding precision (e.g. 2–4 decimals) computed over two or more columns of a table, and the script must decide which rows enter the computation.
Pattern
The attempt loads the table and computes the statistic without explicitly stating and checking the analysis population: it silently relies on a library's default NA/dtype handling (or applies an extra filter, dedup, or subset/sheet selection not requested), never prints the number of rows actually used, and reports a value that is off by one unit in the last requested decimal place. Often no script is preserved, so the row set cannot be reconstructed.
Detection procedure
  1. Read the task: note that no filtering/subsetting is requested and note the required rounding precision — this tells you how sensitive the answer is to a few rows.
  2. Read the script: locate every operation that can change the row set or dtypes before the statistic (dropna, read_* with type coercion, non-numeric strings coerced to NaN, joins, dedup, head/sample, per-column vs pairwise NA deletion, sheet/subset choice). Check that exactly one, explicitly justified policy is used.
  3. Check that the script prints diagnostics for the computed quantity: total rows loaded, rows used after cleaning, dtypes of the two columns, and the unrounded statistic — and that the reviewer can see these numbers.
  4. Check the answer: the reported value must be the rounded unrounded statistic, all requested fields must be present at the stated precision (e.g. a p-value must be shown with the requested number of decimals, not collapsed to 0.0), and the run must be reproducible from a saved script.
Discriminator
A real violation is an unexamined or unjustified row/dtype decision with no printed row count, so a one-decimal discrepancy cannot be diagnosed. It is fine if the script uses the full table (or a filter the task explicitly requires), documents the NA policy, and prints row counts/dtypes that confirm the intended population — even if some rows are legitimately dropped.
Consequence
The statistic is computed on a slightly different population than intended and rounds to a neighboring value (e.g. 0.53 vs 0.54), so the exact-match numeric check fails even though the qualitative conclusion is right; formatting slips like 0.0 for a 4-decimal p-value can fail additional checks.
id c114c5f9c1d1 · mined from infiagent-dabench dabench-300@s16
raw text (what the judge reads)
### Unverified row-selection / missing-value handling for a precision-sensitive statistic
- **Applies when**: `task` -- the task asks for a single numeric statistic reported at a fixed rounding precision (e.g. 2–4 decimals) computed over two or more columns of a table, and the script must decide which rows enter the computation.
- **Pattern**: The attempt loads the table and computes the statistic without explicitly stating and checking the analysis population: it silently relies on a library's default NA/dtype handling (or applies an extra filter, dedup, or subset/sheet selection not requested), never prints the number of rows actually used, and reports a value that is off by one unit in the last requested decimal place. Often no script is preserved, so the row set cannot be reconstructed.
- **Detection procedure**:
  1. Read the task: note that no filtering/subsetting is requested and note the required rounding precision — this tells you how sensitive the answer is to a few rows.
  2. Read the script: locate every operation that can change the row set or dtypes before the statistic (`dropna`, `read_*` with type coercion, non-numeric strings coerced to NaN, joins, dedup, head/sample, per-column vs pairwise NA deletion, sheet/subset choice). Check that exactly one, explicitly justified policy is used.
  3. Check that the script prints diagnostics for the computed quantity: total rows loaded, rows used after cleaning, dtypes of the two columns, and the unrounded statistic — and that the reviewer can see these numbers.
  4. Check the answer: the reported value must be the rounded unrounded statistic, all requested fields must be present at the stated precision (e.g. a p-value must be shown with the requested number of decimals, not collapsed to `0.0`), and the run must be reproducible from a saved script.
- **Discriminator**: A real violation is an unexamined or unjustified row/dtype decision with no printed row count, so a one-decimal discrepancy cannot be diagnosed. It is fine if the script uses the full table (or a filter the task explicitly requires), documents the NA policy, and prints row counts/dtypes that confirm the intended population — even if some rows are legitimately dropped.
- **Consequence**: The statistic is computed on a slightly different population than intended and rounds to a neighboring value (e.g. 0.53 vs 0.54), so the exact-match numeric check fails even though the qualitative conclusion is right; formatting slips like `0.0` for a 4-decimal p-value can fail additional checks.
864Predictions shipped without a held-out validation score or distribution/shape sanity checktaskda-code
Applies when
task -- the deliverable is a file of model predictions on a provided test set, and the answer only reports model hyperparameters and summary statistics of the predictions.
Pattern
The attempt fits one model on the full training data, writes the prediction file, and reports descriptive stats of its own outputs (mean/min/max/std) as if they were evidence of quality — with no train/validation split, no error metric on labeled data, no baseline comparison, and no check that the predicted values are plausible against the observed target distribution or that the output row count/order matches the test rows exactly.
Detection procedure
  1. Read the task to identify the required output file, column name, row count (= number of test rows) and any implied units/rounding.
  2. Search the scripts for a labeled holdout or cross-validation and a computed error metric (e.g. RMSE/MAE/R² or relative error) plus at least one trivial baseline; note their absence.
  3. Search for post-write assertions: output length equals test length, index/order preserved, no NaNs, and predicted min/max/mean compared to the training target's min/max/mean.
  4. Check whether the answer's claim of success rests only on hyperparameters and self-reported prediction statistics rather than on a validation number; also check that the reported prediction range does not extend far beyond the plausible target range.
Discriminator
A real violation has zero quantitative evidence measured against known labels and zero shape/range assertions. It is not a violation if the scripts report a holdout/CV metric (even from a single simple model) and verify the output file's shape, column name and value range — a modest-accuracy but validated and sanity-checked pipeline is acceptable.
Consequence
The grader compares the submitted file against ground truth and fails it (file "WRONG/MISSING") because errors that a holdout metric, row-count check, or range check would have exposed — mis-encoded categoricals, unaligned/duplicated rows, implausible extreme predictions, or a model far worse than a baseline — went undetected, and the answer offers no evidence to distinguish a good fit from a broken one.
id 6be6854b9365 · mined from da-code dacode-ml-regression-014@s16
raw text (what the judge reads)
### Predictions shipped without a held-out validation score or distribution/shape sanity check
- **Applies when**: `task` -- the deliverable is a file of model predictions on a provided test set, and the answer only reports model hyperparameters and summary statistics of the predictions.
- **Pattern**: The attempt fits one model on the full training data, writes the prediction file, and reports descriptive stats of its own outputs (mean/min/max/std) as if they were evidence of quality — with no train/validation split, no error metric on labeled data, no baseline comparison, and no check that the predicted values are plausible against the observed target distribution or that the output row count/order matches the test rows exactly.
- **Detection procedure**:
  1. Read the task to identify the required output file, column name, row count (= number of test rows) and any implied units/rounding.
  2. Search the scripts for a labeled holdout or cross-validation and a computed error metric (e.g. RMSE/MAE/R² or relative error) plus at least one trivial baseline; note their absence.
  3. Search for post-write assertions: output length equals test length, index/order preserved, no NaNs, and predicted min/max/mean compared to the training target's min/max/mean.
  4. Check whether the answer's claim of success rests only on hyperparameters and self-reported prediction statistics rather than on a validation number; also check that the reported prediction range does not extend far beyond the plausible target range.
- **Discriminator**: A real violation has zero quantitative evidence measured against known labels and zero shape/range assertions. It is *not* a violation if the scripts report a holdout/CV metric (even from a single simple model) and verify the output file's shape, column name and value range — a modest-accuracy but validated and sanity-checked pipeline is acceptable.
- **Consequence**: The grader compares the submitted file against ground truth and fails it (file "WRONG/MISSING") because errors that a holdout metric, row-count check, or range check would have exposed — mis-encoded categoricals, unaligned/duplicated rows, implausible extreme predictions, or a model far worse than a baseline — went undetected, and the answer offers no evidence to distinguish a good fit from a broken one.
865Uses a different data source than the provided split files, with no held-out validation or row-alignment checktaskda-code
Applies when
task -- the task supplies designated train/test files and asks for per-row predictions written to a specific output file/column.
Pattern
The attempt trains on (or reindexes from) some other file — e.g. the original/full upstream dataset or a re-derived split — instead of the exact provided training file, and reports only self-descriptive stats (row counts, class distribution, "no missing values") without any held-out accuracy or verification that output rows correspond 1:1, in order, to the provided test file.
Detection procedure
  1. Read the task to identify the exact file(s) designated as training input, the exact file designated as prediction input, and the required output format.
  2. In the scripts, check which file paths are actually loaded for fitting and for predicting; confirm the prediction input is the designated test file and the fit input is the designated train file (not a superset that may contain test rows, and not a re-split of a combined file).
  3. Check that the script verifies len(predictions) == len(test_file) and writes them in the test file's original row order without sorting/deduplication/filtering, and that a labeled holdout (or CV) score is computed and reported.
  4. Compare the answer's reported row counts and class shares against the designated test file's row count and the training label distribution; large deviations or counts that only match the upstream dataset are red flags.
Discriminator
A real violation is reading a different/larger source for training or emitting a row count or ordering that cannot be traced to the provided test file, with no validation score to detect it; it is fine if the script loads the designated files, keeps test order, asserts the row count, and reports a holdout metric — even if the model is simple.
Consequence
The submitted prediction column misaligns with the grader's ground-truth rows (or is scored against labels the model already memorized), so row-wise accuracy collapses and the file is marked WRONG despite plausible-looking summary statistics.
id da3ed98d5ea2 · mined from da-code dacode-ml-multi-008@s16
raw text (what the judge reads)
### Uses a different data source than the provided split files, with no held-out validation or row-alignment check
- **Applies when**: `task` -- the task supplies designated train/test files and asks for per-row predictions written to a specific output file/column.
- **Pattern**: The attempt trains on (or reindexes from) some other file — e.g. the original/full upstream dataset or a re-derived split — instead of the exact provided training file, and reports only self-descriptive stats (row counts, class distribution, "no missing values") without any held-out accuracy or verification that output rows correspond 1:1, in order, to the provided test file.
- **Detection procedure**:
  1. Read the task to identify the exact file(s) designated as training input, the exact file designated as prediction input, and the required output format.
  2. In the scripts, check which file paths are actually loaded for fitting and for predicting; confirm the prediction input is the designated test file and the fit input is the designated train file (not a superset that may contain test rows, and not a re-split of a combined file).
  3. Check that the script verifies `len(predictions) == len(test_file)` and writes them in the test file's original row order without sorting/deduplication/filtering, and that a labeled holdout (or CV) score is computed and reported.
  4. Compare the answer's reported row counts and class shares against the designated test file's row count and the training label distribution; large deviations or counts that only match the upstream dataset are red flags.
  5. 
- **Discriminator**: A real violation is reading a different/larger source for training or emitting a row count or ordering that cannot be traced to the provided test file, with no validation score to detect it; it is fine if the script loads the designated files, keeps test order, asserts the row count, and reports a holdout metric — even if the model is simple.
- **Consequence**: The submitted prediction column misaligns with the grader's ground-truth rows (or is scored against labels the model already memorized), so row-wise accuracy collapses and the file is marked WRONG despite plausible-looking summary statistics.
866Ignoring the provided output template's schema and row settaskda-code
Applies when
task -- the task supplies a pre-existing result file (CSV/JSON/template) that must be filled in "adhering strictly to its format".
Pattern
The agent never reads the template's header, row labels, and expected number/order of rows; it instead writes its own invented column names, adds or drops categories (e.g. an "unknown"/null bucket, or missing a category present in the template), or fills cells with a different quantity than the one the template's labels ask for, and then reports prose numbers as if they were the deliverable.
Detection procedure
  1. Read the task for the named output file and confirm the scripts explicitly load/inspect that file (header row, index labels, row count, dtypes) before writing.
  2. Compare the categories/keys the script derives from the data against the labels already present in the template: are any template rows missing, and are any extra rows (nulls, "unknown", unmapped values) being appended?
  3. Check that each written cell matches the quantity implied by the template's column names (count vs. name vs. rate), and that the final file is written in place with the original column names, order, and separator — not overwritten by a freshly constructed frame.
  4. Cross-check the answer text against the file: do the reported per-category numbers sum/align with the number of rows written, and does the categorization scheme match the template's?
Discriminator
A real violation is inventing schema (new/missing rows or columns, renamed headers, different unit/quantity per cell) without evidence the template was read; it is fine if the script reads the template, maps its own derived categories onto the template's labels (dropping or imputing unmatched records deliberately), and preserves header/order even if the numeric values are debatable.
Consequence
The grader's file comparison fails outright (WRONG/MISSING) because keys, row count, or column names don't match the expected file, regardless of how reasonable the narrative summary looks.
id b751a9b53808 · mined from da-code dacode-dm-csv-001@s16
raw text (what the judge reads)
### Ignoring the provided output template's schema and row set
- **Applies when**: `task` -- the task supplies a pre-existing result file (CSV/JSON/template) that must be filled in "adhering strictly to its format".
- **Pattern**: The agent never reads the template's header, row labels, and expected number/order of rows; it instead writes its own invented column names, adds or drops categories (e.g. an "unknown"/null bucket, or missing a category present in the template), or fills cells with a different quantity than the one the template's labels ask for, and then reports prose numbers as if they were the deliverable.
- **Detection procedure**:
  1. Read the task for the named output file and confirm the scripts explicitly load/inspect that file (header row, index labels, row count, dtypes) before writing.
  2. Compare the categories/keys the script derives from the data against the labels already present in the template: are any template rows missing, and are any extra rows (nulls, "unknown", unmapped values) being appended?
  3. Check that each written cell matches the quantity implied by the template's column names (count vs. name vs. rate), and that the final file is written in place with the original column names, order, and separator — not overwritten by a freshly constructed frame.
  4. Cross-check the answer text against the file: do the reported per-category numbers sum/align with the number of rows written, and does the categorization scheme match the template's?
- **Discriminator**: A real violation is inventing schema (new/missing rows or columns, renamed headers, different unit/quantity per cell) without evidence the template was read; it is *fine* if the script reads the template, maps its own derived categories onto the template's labels (dropping or imputing unmatched records deliberately), and preserves header/order even if the numeric values are debatable.
- **Consequence**: The grader's file comparison fails outright (WRONG/MISSING) because keys, row count, or column names don't match the expected file, regardless of how reasonable the narrative summary looks.
867Unverified entity-level aggregation / split definition behind a correlationtaskinfiagent-dabench
Applies when
task -- the task asks for a statistic (e.g., correlation) between two quantities that must first be derived per entity (max of a per-record attribute, duration from first/last timestamps) and then computed on subgroups defined by a threshold such as a median.
Pattern
The attempt computes the statistic directly on raw records or on a hastily derived table without documenting/validating the unit of analysis, the duration definition (last−first timestamp, inclusive vs exclusive, time unit, duplicate/multi-year records), how the per-entity damage value is obtained (max vs sum vs last), or how ties at the median threshold are assigned — so a plausible-looking but slightly different r is reported, and no script is retained to check.
Detection procedure
  1. Read the task and write down the intended unit of analysis and the exact definition of each derived variable and of the subgroup split.
  2. In the scripts, locate the aggregation step: confirm there is one row per entity (a groupby/drop_duplicates on a stable entity key), and check how duration, the max attribute, and the damage value per entity are computed and in what units.
  3. Check the threshold logic: is the median computed on the entity-level table (not raw rows), and are > / >= boundary cases handled so the two subgroups partition the data — print group sizes and confirm they sum to the entity count and are roughly balanced.
  4. Cross-check the reported numbers: entity counts, min/max duration, category range, and whether r and p are recomputed from the same filtered frame that was described; if no script or intermediate counts exist, treat the result as unverified.
Discriminator
A real violation is a missing or ambiguous aggregation/threshold step (raw-row correlation, duplicated entities, undocumented duration unit, ties silently dropped or double-counted). A look-alike that is fine is a script that does aggregate per entity and prints group sizes/ranges, even if it makes a defensible choice among equivalent conventions that leaves counts consistent.
Consequence
Sign and significance may still come out right, so the qualitative verdict passes, but the correlation coefficient is off by a few hundredths and fails the exact rounded-value check, marking the whole answer wrong.
id a65c980d8ca0 · mined from infiagent-dabench dabench-431@s16
raw text (what the judge reads)
### Unverified entity-level aggregation / split definition behind a correlation
- **Applies when**: `task` -- the task asks for a statistic (e.g., correlation) between two quantities that must first be *derived* per entity (max of a per-record attribute, duration from first/last timestamps) and then computed on subgroups defined by a threshold such as a median.
- **Pattern**: The attempt computes the statistic directly on raw records or on a hastily derived table without documenting/validating the unit of analysis, the duration definition (last−first timestamp, inclusive vs exclusive, time unit, duplicate/multi-year records), how the per-entity damage value is obtained (max vs sum vs last), or how ties at the median threshold are assigned — so a plausible-looking but slightly different r is reported, and no script is retained to check.
- **Detection procedure**:
  1. Read the task and write down the intended unit of analysis and the exact definition of each derived variable and of the subgroup split.
  2. In the scripts, locate the aggregation step: confirm there is one row per entity (a `groupby`/`drop_duplicates` on a stable entity key), and check how duration, the max attribute, and the damage value per entity are computed and in what units.
  3. Check the threshold logic: is the median computed on the entity-level table (not raw rows), and are `>` / `>=` boundary cases handled so the two subgroups partition the data — print group sizes and confirm they sum to the entity count and are roughly balanced.
  4. Cross-check the reported numbers: entity counts, min/max duration, category range, and whether r and p are recomputed from the same filtered frame that was described; if no script or intermediate counts exist, treat the result as unverified.
- **Discriminator**: A real violation is a missing or ambiguous aggregation/threshold step (raw-row correlation, duplicated entities, undocumented duration unit, ties silently dropped or double-counted). A look-alike that is fine is a script that does aggregate per entity and prints group sizes/ranges, even if it makes a defensible choice among equivalent conventions that leaves counts consistent.
- **Consequence**: Sign and significance may still come out right, so the qualitative verdict passes, but the correlation coefficient is off by a few hundredths and fails the exact rounded-value check, marking the whole answer wrong.
868No held-out validation or sanity check of the prediction distribution before submittingtaskda-code
Applies when
task -- the script fits a model on labeled data and writes predictions for an unlabeled test file that will be scored against hidden ground truth.
Pattern
The attempt trains a single model with hand-picked hyperparameters, never evaluates it on a held-out split or via cross-validation (no accuracy/F1/AUC reported), and never compares the predicted class distribution with the training label distribution — so an under-performing or badly-calibrated classifier (e.g., one that predicts far fewer positives than the base rate on an imbalanced target) is shipped unchecked. Formatting of the output file is likewise asserted rather than verified against the requested columns/row count.
Detection procedure
  1. Read the task for the scoring artifact and any format constraints (file name, required column name(s), one row per test record).
  2. Scan the script for a train/validation split or cross-validation with a metric printed; note that importing a splitting utility without using it does not count.
  3. Check whether the script compares predicted positive rate / class counts against the training target's base rate, and whether it verifies output shape and column names after writing.
  4. Read the answer: if it reports only prediction counts and hyperparameters with no out-of-sample performance estimate, and those counts imply a positive rate materially below (or above) the training base rate without comment, flag it.
Discriminator
A genuine violation has zero out-of-sample evidence that the model is any good and no distributional/format sanity check. It is not a violation if the agent reports a validation/CV score (even from a simple single split), or explicitly checks the prediction rate against the training prior and justifies a deliberate threshold choice, or if the target is balanced and the reported metric shows the model is reasonable.
Consequence
The saved predictions can be systematically biased toward the majority class or otherwise below the grader's accuracy/F1 threshold, or malformed (extra/missing columns, wrong row count), and the file check fails with no diagnostic in the answer explaining why.
id d294ad01a9c3 · mined from da-code dacode-ml-binary-016@s16
raw text (what the judge reads)
### No held-out validation or sanity check of the prediction distribution before submitting
- **Applies when**: `task` -- the script fits a model on labeled data and writes predictions for an unlabeled test file that will be scored against hidden ground truth.
- **Pattern**: The attempt trains a single model with hand-picked hyperparameters, never evaluates it on a held-out split or via cross-validation (no accuracy/F1/AUC reported), and never compares the predicted class distribution with the training label distribution — so an under-performing or badly-calibrated classifier (e.g., one that predicts far fewer positives than the base rate on an imbalanced target) is shipped unchecked. Formatting of the output file is likewise asserted rather than verified against the requested columns/row count.
- **Detection procedure**:
  1. Read the task for the scoring artifact and any format constraints (file name, required column name(s), one row per test record).
  2. Scan the script for a train/validation split or cross-validation with a metric printed; note that importing a splitting utility without using it does not count.
  3. Check whether the script compares predicted positive rate / class counts against the training target's base rate, and whether it verifies output shape and column names after writing.
  4. Read the answer: if it reports only prediction counts and hyperparameters with no out-of-sample performance estimate, and those counts imply a positive rate materially below (or above) the training base rate without comment, flag it.
- **Discriminator**: A genuine violation has zero out-of-sample evidence that the model is any good and no distributional/format sanity check. It is *not* a violation if the agent reports a validation/CV score (even from a simple single split), or explicitly checks the prediction rate against the training prior and justifies a deliberate threshold choice, or if the target is balanced and the reported metric shows the model is reasonable.
- **Consequence**: The saved predictions can be systematically biased toward the majority class or otherwise below the grader's accuracy/F1 threshold, or malformed (extra/missing columns, wrong row count), and the file check fails with no diagnostic in the answer explaining why.
869Required output artifact not produced in the specified formattaskda-code
Applies when
task -- the task asks for the result to be written to a named file following a provided sample/template file's format, in addition to (or instead of) reporting a value.
Pattern
The attempt computes a number and reports it inline in the final answer, but never writes the named output file, or writes it without inspecting the template (wrong filename, path, missing header/column names, wrong row layout, or extra index column).
Detection procedure
  1. Read the task and list every required deliverable: exact output filename, the template file whose schema must be copied, and any rounding/units/ordering constraints.
  2. Search the scripts for code that reads the template and writes the named file (e.g., a to-file call with that exact name) and check the written columns/rows mirror the template.
  3. Check the answer/log for evidence the file exists after execution (a write confirmation, a re-read printout, or a directory listing).
  4. If no script exists at all, or no write step is present, flag immediately regardless of whether the reported number looks plausible.
Discriminator
A real violation is missing/misnamed/mis-schema'd output or output existing only as a printed value; a look-alike that is fine is a script that writes the exact filename with template-matching columns and merely also prints the value for convenience.
Consequence
The grader looks for the named file and reports it WRONG/MISSING, scoring 0 even if the computed statistic is numerically correct.
id 5f88660d9e97 · mined from da-code dacode-data-sa-039@s16
raw text (what the judge reads)
### Required output artifact not produced in the specified format
- **Applies when**: `task` -- the task asks for the result to be written to a named file following a provided sample/template file's format, in addition to (or instead of) reporting a value.
- **Pattern**: The attempt computes a number and reports it inline in the final answer, but never writes the named output file, or writes it without inspecting the template (wrong filename, path, missing header/column names, wrong row layout, or extra index column).
- **Detection procedure**:
  1. Read the task and list every required deliverable: exact output filename, the template file whose schema must be copied, and any rounding/units/ordering constraints.
  2. Search the scripts for code that reads the template and writes the named file (e.g., a to-file call with that exact name) and check the written columns/rows mirror the template.
  3. Check the answer/log for evidence the file exists after execution (a write confirmation, a re-read printout, or a directory listing).
  4. If no script exists at all, or no write step is present, flag immediately regardless of whether the reported number looks plausible.
- **Discriminator**: A real violation is missing/misnamed/mis-schema'd output or output existing only as a printed value; a look-alike that is fine is a script that writes the exact filename with template-matching columns and merely also prints the value for convenience.
- **Consequence**: The grader looks for the named file and reports it WRONG/MISSING, scoring 0 even if the computed statistic is numerically correct.
870Unvalidated predictions: no held-out error estimate or distribution sanity check before submittingtaskda-code
Applies when
task -- the deliverable is a prediction file produced by a model trained on one file and applied to a separate test file, and the agent reports success without any quantified validation.
Pattern
The attempt fits a single model, writes the output file, and reports only descriptive facts (row counts, feature count, min/max/mean of predictions) with no held-out validation score, no comparison of predicted vs. training-target distribution, no check that the training source actually matches the test file's schema/rows (e.g. test rows possibly being a subset of the "complete" training file, duplicated/misaligned feature columns, dropped or re-ordered rows), and no reproducible script for either step.
Detection procedure
  1. Read the task to fix the required deliverable (file name, column name, one prediction per test row, original test row order).
  2. In the scripts, look for (a) a train/validation split or cross-validation with a reported error metric, (b) an explicit alignment check that the feature columns used at predict time are the same names/dtypes/order as at fit time and that no target-derived or ID/time column leaked in, and (c) confirmation that the number of output rows equals the number of test rows in unchanged order.
  3. In the answer, check whether any quantified generalization estimate is given; if the only evidence is prediction summary statistics, compare those statistics to the training target's own statistics (mean, spread, max) — a badly shrunk or shifted prediction distribution, or a mean/range far from the training target's, signals a broken pipeline.
  4. Flag the attempt if no validation metric exists, or if scripts are absent/not rerunnable so the claim cannot be checked.
Discriminator
A fine attempt may use a simple model, but it still reports a held-out metric and shows the output shape/column/order matches the test file; a violation reports only self-consistent bookkeeping numbers that would look identical whether the features were aligned or scrambled. Predictions that are legitimately smoother than the target (a real property of averaged models) are acceptable if backed by a validation score; unexplained shrinkage with no score is not.
Consequence
The saved file may have the right name, shape, and header yet contain predictions that miss the true values, so the grader's value-based comparison against the expected predictions fails while the agent's summary still claims success.
id 749cb015eebb · mined from da-code dacode-ml-regression-015@s16
raw text (what the judge reads)
### Unvalidated predictions: no held-out error estimate or distribution sanity check before submitting
- **Applies when**: `task` -- the deliverable is a prediction file produced by a model trained on one file and applied to a separate test file, and the agent reports success without any quantified validation.
- **Pattern**: The attempt fits a single model, writes the output file, and reports only descriptive facts (row counts, feature count, min/max/mean of predictions) with no held-out validation score, no comparison of predicted vs. training-target distribution, no check that the training source actually matches the test file's schema/rows (e.g. test rows possibly being a subset of the "complete" training file, duplicated/misaligned feature columns, dropped or re-ordered rows), and no reproducible script for either step.
- **Detection procedure**:
  1. Read the task to fix the required deliverable (file name, column name, one prediction per test row, original test row order).
  2. In the scripts, look for (a) a train/validation split or cross-validation with a reported error metric, (b) an explicit alignment check that the feature columns used at predict time are the same names/dtypes/order as at fit time and that no target-derived or ID/time column leaked in, and (c) confirmation that the number of output rows equals the number of test rows in unchanged order.
  3. In the answer, check whether any quantified generalization estimate is given; if the only evidence is prediction summary statistics, compare those statistics to the training target's own statistics (mean, spread, max) — a badly shrunk or shifted prediction distribution, or a mean/range far from the training target's, signals a broken pipeline.
  4. Flag the attempt if no validation metric exists, or if scripts are absent/not rerunnable so the claim cannot be checked.
- **Discriminator**: A fine attempt may use a simple model, but it still reports a held-out metric and shows the output shape/column/order matches the test file; a violation reports only self-consistent bookkeeping numbers that would look identical whether the features were aligned or scrambled. Predictions that are legitimately smoother than the target (a real property of averaged models) are acceptable *if* backed by a validation score; unexplained shrinkage with no score is not.
- **Consequence**: The saved file may have the right name, shape, and header yet contain predictions that miss the true values, so the grader's value-based comparison against the expected predictions fails while the agent's summary still claims success.
871Extremum selected with `idxmax`/`idxmin` when the requested answer allows multiple tied entities (and no result artifact is written)taskda-code
Applies when
task -- the task asks for the entity/entities attaining a maximum or minimum of some column and the requested output schema is a list (or otherwise plural), while the script picks the extremum via a single-index lookup.
Pattern
The script computes idxmax()/idxmin() (or sort_values().iloc[0], nlargest(1)) and reports exactly one label, silently discarding any other rows sharing the same extreme value; it also often only prints results instead of writing the required output file in the required format.
Detection procedure
  1. Read the task's output template: note whether each key expects a list/multiple values and whether a specific result file must be produced.
  2. In the scripts, locate how the extremum row is chosen; flag any use of a first-match index (idxmax, idxmin, iloc[0], nlargest(1)) that is not preceded/followed by a mask such as df[df[col] == df[col].max()] or a duplicate-count check on the extreme value.
  3. Check whether the imputation step could itself create ties (e.g. filling many rows with the same constant) or whether the raw column is coarse/integer-valued so ties are plausible; the scripts never report how many rows attain the extreme value.
  4. Verify a file matching the expected artifact name/format is written; printing to stdout only is a failure.
Discriminator
A real violation is when the column's extreme value can plausibly be shared (integer/rounded values, imputed constants, duplicated entities) and no tie check exists; it is fine if the script explicitly filters all rows equal to the extremum, or demonstrates uniqueness (e.g. prints the count of rows at the extreme value) before returning a single element.
Consequence
The submitted JSON contains a single, arbitrary member of a tied group (order-dependent), so the exact-match check against the expected result file fails even though the extreme value itself was computed correctly — and if no result file is written, the grader marks it missing outright.
id 52a08aebe5f6 · mined from da-code dacode-di-text-001@s16
raw text (what the judge reads)
### Extremum selected with `idxmax`/`idxmin` when the requested answer allows multiple tied entities (and no result artifact is written)

- **Applies when**: `task` -- the task asks for the entity/entities attaining a maximum or minimum of some column and the requested output schema is a list (or otherwise plural), while the script picks the extremum via a single-index lookup.
- **Pattern**: The script computes `idxmax()`/`idxmin()` (or `sort_values().iloc[0]`, `nlargest(1)`) and reports exactly one label, silently discarding any other rows sharing the same extreme value; it also often only prints results instead of writing the required output file in the required format.
- **Detection procedure**:
  1. Read the task's output template: note whether each key expects a list/multiple values and whether a specific result file must be produced.
  2. In the scripts, locate how the extremum row is chosen; flag any use of a first-match index (`idxmax`, `idxmin`, `iloc[0]`, `nlargest(1)`) that is not preceded/followed by a mask such as `df[df[col] == df[col].max()]` or a duplicate-count check on the extreme value.
  3. Check whether the imputation step could itself create ties (e.g. filling many rows with the same constant) or whether the raw column is coarse/integer-valued so ties are plausible; the scripts never report how many rows attain the extreme value.
  4. Verify a file matching the expected artifact name/format is written; printing to stdout only is a failure.
- **Discriminator**: A real violation is when the column's extreme value can plausibly be shared (integer/rounded values, imputed constants, duplicated entities) and no tie check exists; it is fine if the script explicitly filters all rows equal to the extremum, or demonstrates uniqueness (e.g. prints the count of rows at the extreme value) before returning a single element.
- **Consequence**: The submitted JSON contains a single, arbitrary member of a tied group (order-dependent), so the exact-match check against the expected result file fails even though the extreme value itself was computed correctly — and if no result file is written, the grader marks it missing outright.
872Deliverable omits requested per-entity columns (over-reduced output schema)taskda-code
Applies when
task -- the task asks for several derived quantities per entity (component scores, aggregated metrics, group/segment labels, class levels) to be saved into a single result file.
Pattern
The scripts compute all the intermediate per-entity metrics and component scores, but the final saved file keeps only one summary column (plus the ID), silently dropping the other quantities the task explicitly asked to include; the answer then declares success and claims verification against a reference that only some columns were compared to.
Detection procedure
  1. From the task statement, enumerate every quantity requested per entity (raw metrics, component scores, combined score, segmentation, label) and the required file name.
  2. In the scripts, find the code that writes the result file and list the exact columns written; compare that list against the enumeration from step 1.
  3. In the answer, check the stated output "Format:" / column list and row count; flag if fewer fields than requested, or if the answer never shows the saved file's head/shape.
  4. Check whether the claimed "100% match to reference" was computed over all requested columns or only a subset (and whether earlier scripts in the same session showed mismatches on those columns).
Discriminator
A real violation is when a requested quantity is computed in memory (or comparable in a reference) but absent from the saved file; it is not a violation if the extra columns are genuinely not requested, or if the task explicitly asks for a minimal ID+label output.
Consequence
The grader compares the saved file's schema/values to the expected result and marks it WRONG/MISSING because required columns are absent, even though the underlying computation may have been correct.
id e51c2497a00e · mined from da-code dacode-dm-csv-052@s16
raw text (what the judge reads)
### Deliverable omits requested per-entity columns (over-reduced output schema)
- **Applies when**: `task` -- the task asks for several derived quantities per entity (component scores, aggregated metrics, group/segment labels, class levels) to be saved into a single result file.
- **Pattern**: The scripts compute all the intermediate per-entity metrics and component scores, but the final saved file keeps only one summary column (plus the ID), silently dropping the other quantities the task explicitly asked to include; the answer then declares success and claims verification against a reference that only some columns were compared to.
- **Detection procedure**:
  1. From the task statement, enumerate every quantity requested per entity (raw metrics, component scores, combined score, segmentation, label) and the required file name.
  2. In the scripts, find the code that writes the result file and list the exact columns written; compare that list against the enumeration from step 1.
  3. In the answer, check the stated output "Format:" / column list and row count; flag if fewer fields than requested, or if the answer never shows the saved file's head/shape.
  4. Check whether the claimed "100% match to reference" was computed over all requested columns or only a subset (and whether earlier scripts in the same session showed mismatches on those columns).
- **Discriminator**: A real violation is when a requested quantity is computed in memory (or comparable in a reference) but absent from the saved file; it is *not* a violation if the extra columns are genuinely not requested, or if the task explicitly asks for a minimal ID+label output.
- **Consequence**: The grader compares the saved file's schema/values to the expected result and marks it WRONG/MISSING because required columns are absent, even though the underlying computation may have been correct.
873Requested output artifact not produced with the specified schema and full row coveragetaskda-code
Applies when
task -- the task asks for results to be written to a named file with explicitly named/ordered columns covering every record of the provided data.
Pattern
The attempt performs the analysis on a truncated/subsampled slice of the data and reports narrative summary statistics (counts, quality scores, chosen hyperparameter) instead of verifiably writing the named file with the exact required column names and one row per input record; no script is retained that demonstrates the write.
Detection procedure
  1. From the task, list the exact deliverable(s): filename, required column names/spellings, and the expected number of rows (= rows in the source data) and columns.
  2. In the scripts, find the line that writes that file; confirm it writes to the required filename, builds columns with the required naming pattern (e.g., per-feature columns plus the label column), and uses the full loaded dataset rather than a head/sample/subset.
  3. In the answer, check that reported row/feature counts match the source data's true shape, and that the deliverable's existence and shape are confirmed (not just metrics described in prose).
  4. Flag if no script exists, if the write step is absent, if columns are renamed/omitted, or if the sample count is smaller than the dataset.
Discriminator
A genuine violation is missing/mis-schema'd/partial output (e.g., only a subsample of rows, wrong column names, results only in the chat text). A look-alike that is fine: the full file is written with the correct schema and row count, and the prose merely summarizes it — or subsampling was explicitly permitted by the task and the output still covers all requested records.
Consequence
The grader looks for the named result file with the expected schema and row coverage, finds it missing or mismatched, and marks the deliverable WRONG/MISSING regardless of clustering quality.
id a934bca047c7 · mined from da-code dacode-ml-cluster-010@s16
raw text (what the judge reads)
### Requested output artifact not produced with the specified schema and full row coverage
- **Applies when**: `task` -- the task asks for results to be written to a named file with explicitly named/ordered columns covering every record of the provided data.
- **Pattern**: The attempt performs the analysis on a truncated/subsampled slice of the data and reports narrative summary statistics (counts, quality scores, chosen hyperparameter) instead of verifiably writing the named file with the exact required column names and one row per input record; no script is retained that demonstrates the write.
- **Detection procedure**:
  1. From the task, list the exact deliverable(s): filename, required column names/spellings, and the expected number of rows (= rows in the source data) and columns.
  2. In the scripts, find the line that writes that file; confirm it writes to the required filename, builds columns with the required naming pattern (e.g., per-feature columns plus the label column), and uses the full loaded dataset rather than a head/sample/subset.
  3. In the answer, check that reported row/feature counts match the source data's true shape, and that the deliverable's existence and shape are confirmed (not just metrics described in prose).
  4. Flag if no script exists, if the write step is absent, if columns are renamed/omitted, or if the sample count is smaller than the dataset.
- **Discriminator**: A genuine violation is missing/mis-schema'd/partial output (e.g., only a subsample of rows, wrong column names, results only in the chat text). A look-alike that is fine: the full file is written with the correct schema and row count, and the prose merely summarizes it — or subsampling was explicitly permitted by the task and the output still covers all requested records.
- **Consequence**: The grader looks for the named result file with the expected schema and row coverage, finds it missing or mismatched, and marks the deliverable WRONG/MISSING regardless of clustering quality.
874Accepting a degenerate cluster solution driven by heavy-tailed featurestaskda-code
Applies when
task -- the task asks for unsupervised grouping into "an appropriate number of clusters" and the scripts build aggregate numeric features (sums, counts, totals, averages) that are strongly right-skewed, then pick k by an internal score (silhouette/inertia).
Pattern
The attempt standardizes skewed, outlier-heavy features without any transformation or outlier handling, then chooses k by maximizing an internal index. Because isolating a handful of extreme points maximizes that index, the reported solution has one or more near-empty clusters (a few members) and the remaining mass split into two undifferentiated groups — a solution that is not a meaningful segmentation, and is never checked against cluster-size or profile sanity criteria.
Detection procedure
  1. In the task, note that the deliverable is an interpretable set of customer/entity groups plus a feature matrix, not just any label vector.
  2. In the scripts, check whether skew-heavy aggregate features are log/rank/robust-transformed or outlier-trimmed before scaling, and whether k is selected by a single internal score with no additional criterion (elbow, stability, minimum cluster size, cluster profiling).
  3. In the printed/reported results, look at the per-cluster counts: flag if any cluster holds a negligible share (e.g. <1% or a literal handful of rows) or if the number of substantive clusters collapses to 1–2.
  4. Confirm the final saved file was produced by the same, final pipeline (correct row count, feature columns, and label column names/ordering as specified) rather than by an earlier exploratory script.
Discriminator
A genuinely small cluster is fine if the attempt deliberately kept outliers and shows the small group is a coherent, described segment with several members and distinct profiles across multiple features; the violation is a size-1-to-few cluster that exists only because untransformed extremes inflate the selection metric, with the rest of the data left essentially unsegmented and no sanity check performed.
Consequence
The saved label file matches neither the expected number nor the expected composition of clusters (labels are effectively a 2-way split plus outlier flags), so the file-comparison/cluster-agreement check fails and the reported segmentation is uninformative.
id 0c4d099a39c9 · mined from da-code dacode-ml-cluster-019@s16
raw text (what the judge reads)
### Accepting a degenerate cluster solution driven by heavy-tailed features
- **Applies when**: `task` -- the task asks for unsupervised grouping into "an appropriate number of clusters" and the scripts build aggregate numeric features (sums, counts, totals, averages) that are strongly right-skewed, then pick k by an internal score (silhouette/inertia).
- **Pattern**: The attempt standardizes skewed, outlier-heavy features without any transformation or outlier handling, then chooses k by maximizing an internal index. Because isolating a handful of extreme points maximizes that index, the reported solution has one or more near-empty clusters (a few members) and the remaining mass split into two undifferentiated groups — a solution that is not a meaningful segmentation, and is never checked against cluster-size or profile sanity criteria.
- **Detection procedure**:
  1. In the task, note that the deliverable is an interpretable set of customer/entity groups plus a feature matrix, not just any label vector.
  2. In the scripts, check whether skew-heavy aggregate features are log/rank/robust-transformed or outlier-trimmed before scaling, and whether k is selected by a single internal score with no additional criterion (elbow, stability, minimum cluster size, cluster profiling).
  3. In the printed/reported results, look at the per-cluster counts: flag if any cluster holds a negligible share (e.g. <1% or a literal handful of rows) or if the number of substantive clusters collapses to 1–2.
  4. Confirm the final saved file was produced by the same, final pipeline (correct row count, feature columns, and label column names/ordering as specified) rather than by an earlier exploratory script.
- **Discriminator**: A genuinely small cluster is fine if the attempt deliberately kept outliers and shows the small group is a coherent, described segment with several members and distinct profiles across *multiple* features; the violation is a size-1-to-few cluster that exists only because untransformed extremes inflate the selection metric, with the rest of the data left essentially unsegmented and no sanity check performed.
- **Consequence**: The saved label file matches neither the expected number nor the expected composition of clusters (labels are effectively a 2-way split plus outlier flags), so the file-comparison/cluster-agreement check fails and the reported segmentation is uninformative.
875Missing held-out validation + no distribution sanity check on predictionstaskda-code
Applies when
task -- a script trains a model on a training subset and writes predicted values for a test file that will be scored against hidden ground truth.
Pattern
The attempt reports near-perfect fit statistics (e.g., R² ≈ 0.99) that were computed on the same rows used for fitting, never evaluates on a held-out split, and never compares the distribution/scale of the written predictions to the distribution of the target in the training data — so a systematically shrunken, collapsed, or mis-scaled prediction column is accepted as "successful".
Detection procedure
  1. In the task, note that the deliverable is a prediction column judged by closeness to unseen truth, so the only meaningful self-check is out-of-sample error plus scale plausibility.
  2. In the scripts, check whether the reported score comes from a train/validation split (or CV) on rows excluded from fitting; flag any score computed with fit and score/predict on the identical data, and flag missing checks for row count, ordering/ID alignment with the test file, and NaN/dtype handling of the target.
  3. Compare the reported summary statistics of the predictions (min/median/mean/max) against the same statistics of the target in the training data; a large mismatch (e.g., median orders of magnitude smaller, or mean far below the training mean) with no explanation is a red flag.
  4. Confirm the answer justifies the output scale/format (rounding, units, non-negativity, one row per test row) rather than only asserting the file was written.
Discriminator
A genuine violation shows fit-quality numbers from in-sample data and/or a prediction distribution grossly inconsistent with the training target with no diagnosis; it is fine if the score comes from truly held-out rows and any distribution shift is explicitly checked and explained (e.g., the test subset is known to differ, or a log/monotone transform was inverted correctly).
Consequence
The saved prediction file is systematically biased toward the low end / mis-scaled, so the grader's accuracy or error threshold against the true values fails even though the reported R² looked excellent.
id 0017ce120124 · mined from da-code dacode-ml-regression-008@s17
raw text (what the judge reads)
### Missing held-out validation + no distribution sanity check on predictions
- **Applies when**: `task` -- a script trains a model on a training subset and writes predicted values for a test file that will be scored against hidden ground truth.
- **Pattern**: The attempt reports near-perfect fit statistics (e.g., R² ≈ 0.99) that were computed on the same rows used for fitting, never evaluates on a held-out split, and never compares the distribution/scale of the written predictions to the distribution of the target in the training data — so a systematically shrunken, collapsed, or mis-scaled prediction column is accepted as "successful".
- **Detection procedure**:
  1. In the task, note that the deliverable is a prediction column judged by closeness to unseen truth, so the only meaningful self-check is out-of-sample error plus scale plausibility.
  2. In the scripts, check whether the reported score comes from a train/validation split (or CV) on rows excluded from fitting; flag any score computed with `fit` and `score`/`predict` on the identical data, and flag missing checks for row count, ordering/ID alignment with the test file, and NaN/dtype handling of the target.
  3. Compare the reported summary statistics of the predictions (min/median/mean/max) against the same statistics of the target in the training data; a large mismatch (e.g., median orders of magnitude smaller, or mean far below the training mean) with no explanation is a red flag.
  4. Confirm the answer justifies the output scale/format (rounding, units, non-negativity, one row per test row) rather than only asserting the file was written.
- **Discriminator**: A genuine violation shows fit-quality numbers from in-sample data and/or a prediction distribution grossly inconsistent with the training target with no diagnosis; it is fine if the score comes from truly held-out rows and any distribution shift is explicitly checked and explained (e.g., the test subset is known to differ, or a log/monotone transform was inverted correctly).
- **Consequence**: The saved prediction file is systematically biased toward the low end / mis-scaled, so the grader's accuracy or error threshold against the true values fails even though the reported R² looked excellent.
876Statistical test run on the whole raw table with default test settingstaskda-code
Applies when
task -- the task asks for a p-value / test decision comparing groups, and the scripts load the raw file(s) and immediately feed all rows into a canned test function.
Pattern
The agent never defines the analysis population or the test form: it skips any filtering to the relevant records (time window, competition/category, status flags, eligible rows), and picks the default two-sided parametric test without checking the direction implied by the question or the distributional shape of the outcome (skewed / discrete count data), so the reported p-value describes a different comparison than the one requested.
Detection procedure
  1. Read the task and README for any wording that scopes the comparison (a date/period, a subset label, a "since"/"only" qualifier) or that implies a directional claim ("more than", "higher"), and note the outcome variable's nature (counts, bounded, heavy-tailed).
  2. In the scripts, look for an explicit subsetting step and an explicit justification of the test (ttest_ind vs. rank-based/nonparametric, one-sided alternative= vs. default two-sided); flag if rows are used as-is and the test is the library default.
  3. Check whether the scripts print row counts/means for the subset actually intended and compare them with the counts printed on the full data — an order-of-magnitude mismatch means the wrong population was used.
  4. Inspect the reported p-value for implausible extremes (e.g. 1e-100 scale) that arise from pooling tens of thousands of unintended rows, and confirm the saved file's values were derived from the scoped test.
Discriminator
A real violation is when the task or README contains scoping/direction cues (or the outcome is clearly non-normal) and the scripts contain no corresponding filter or test-choice decision; it is fine if the agent explicitly considered the scope and test family and documented that the full sample and a two-sided parametric test are the correct choices.
Consequence
The p-value in result.csv differs by many orders of magnitude from the reference value (and can flip the reject / fail-to-reject decision), so the file-comparison check fails even though the code runs without error.
id 1ef607623f1c · mined from da-code dacode-data-sa-001@s17
raw text (what the judge reads)
### Statistical test run on the whole raw table with default test settings
- **Applies when**: `task` -- the task asks for a p-value / test decision comparing groups, and the scripts load the raw file(s) and immediately feed all rows into a canned test function.
- **Pattern**: The agent never defines the analysis population or the test form: it skips any filtering to the relevant records (time window, competition/category, status flags, eligible rows), and picks the default two-sided parametric test without checking the direction implied by the question or the distributional shape of the outcome (skewed / discrete count data), so the reported p-value describes a different comparison than the one requested.
- **Detection procedure**:
  1. Read the task and README for any wording that scopes the comparison (a date/period, a subset label, a "since"/"only" qualifier) or that implies a directional claim ("more than", "higher"), and note the outcome variable's nature (counts, bounded, heavy-tailed).
  2. In the scripts, look for an explicit subsetting step and an explicit justification of the test (`ttest_ind` vs. rank-based/nonparametric, one-sided `alternative=` vs. default two-sided); flag if rows are used as-is and the test is the library default.
  3. Check whether the scripts print row counts/means for the subset actually intended and compare them with the counts printed on the full data — an order-of-magnitude mismatch means the wrong population was used.
  4. Inspect the reported p-value for implausible extremes (e.g. 1e-100 scale) that arise from pooling tens of thousands of unintended rows, and confirm the saved file's values were derived from the scoped test.
- **Discriminator**: A real violation is when the task or README contains scoping/direction cues (or the outcome is clearly non-normal) and the scripts contain no corresponding filter or test-choice decision; it is fine if the agent explicitly considered the scope and test family and documented that the full sample and a two-sided parametric test are the correct choices.
- **Consequence**: The p-value in `result.csv` differs by many orders of magnitude from the reference value (and can flip the reject / fail-to-reject decision), so the file-comparison check fails even though the code runs without error.
877Reference output template provided but never inspected or matchedtaskda-code
Applies when
task -- the instructions say the deliverable must follow the "exact structure/formatting" of a provided sample/template output file, and the scripts generate that file programmatically.
Pattern
The attempt never loads or prints the sample file; instead it guesses the header names, column order, row order, and numeric formatting from memory or from the aggregation defaults, and writes raw unrounded floats (often with full binary precision) into the deliverable.
Detection procedure
  1. Read the task for any named template/sample artifact and any stated formatting constraints (rounding, units, column names, sort order, header presence).
  2. Search the scripts for a read/print of that template file and an explicit comparison of the produced frame's columns, dtypes, row ordering, and decimal precision against it.
  3. Inspect the final answer's header and values: are column names/order derived from the template, and do numeric values show a consistent, deliberate precision rather than artifacts like 916684.7799999999?
  4. If no template read and no formatting normalization step exists, flag the attempt regardless of whether the underlying aggregation logic looks correct.
Discriminator
A real violation is guessing/assuming the schema; it is not a violation if the script loads the template (or hardcodes values with a comment citing the inspected template) and explicitly enforces column names, ordering, and rounding — even if it later reorders rows for readability. Correct aggregation math does not excuse an unverified output format.
Consequence
The grader compares the deliverable cell-by-cell against the expected file and reports the file as WRONG/MISSING due to header, ordering, or precision mismatch, scoring 0 despite plausible-looking numbers.
id 7af1e238eab1 · mined from da-code dacode-dm-csv-011@s17
raw text (what the judge reads)
### Reference output template provided but never inspected or matched
- **Applies when**: `task` -- the instructions say the deliverable must follow the "exact structure/formatting" of a provided sample/template output file, and the scripts generate that file programmatically.
- **Pattern**: The attempt never loads or prints the sample file; instead it guesses the header names, column order, row order, and numeric formatting from memory or from the aggregation defaults, and writes raw unrounded floats (often with full binary precision) into the deliverable.
- **Detection procedure**:
  1. Read the task for any named template/sample artifact and any stated formatting constraints (rounding, units, column names, sort order, header presence).
  2. Search the scripts for a read/print of that template file and an explicit comparison of the produced frame's columns, dtypes, row ordering, and decimal precision against it.
  3. Inspect the final answer's header and values: are column names/order derived from the template, and do numeric values show a consistent, deliberate precision rather than artifacts like `916684.7799999999`?
  4. If no template read and no formatting normalization step exists, flag the attempt regardless of whether the underlying aggregation logic looks correct.
- **Discriminator**: A real violation is guessing/assuming the schema; it is *not* a violation if the script loads the template (or hardcodes values with a comment citing the inspected template) and explicitly enforces column names, ordering, and rounding — even if it later reorders rows for readability. Correct aggregation math does not excuse an unverified output format.
- **Consequence**: The grader compares the deliverable cell-by-cell against the expected file and reports the file as WRONG/MISSING due to header, ordering, or precision mismatch, scoring 0 despite plausible-looking numbers.
878Kitchen-sink feature matrix for distance-based clustering (ordinal-encoding of non-ordinal/date/ID-like fields, no validation of the resulting structure)taskda-code
Applies when
task -- an unsupervised clustering/segmentation task where the script builds the feature matrix by dropping only the obvious identifier and passing every remaining raw column (including string categoricals, dates, rare binary flags) through a single scaler into a distance-based algorithm.
Pattern
The attempt applies LabelEncoder/integer codes to nominal or date-like text columns (creating fake ordinal distances and a near-continuous, high-cardinality axis), keeps all sparse 0/1 and low-variance flag columns unweighted, skips any derived/domain features or dimensionality reduction, then picks the number of clusters purely by the argmax of a low absolute silhouette over a small k-grid — which almost always yields the degenerate k=2 split — and reports it without checking cluster separation, size balance, or interpretability.
Detection procedure
  1. In the task/README, list which columns are identifiers, dates, nominal categories, and sparse indicators; note that the requested output is a feature-vector + cluster-label file, so the feature space itself is part of the deliverable.
  2. In the script, check how each such column is transformed: is any nominal/date column mapped to a single integer code and then z-scored, and are all ~dozens of raw columns fed in with equal weight and no reduction (PCA/one-hot/feature engineering)?
  3. Check the k-selection logic: is k chosen only as argmax silhouette, with the winning score low (roughly <0.3) and the winner being the smallest k in the grid, with no elbow/stability/profile cross-check?
  4. Check the answer for post-hoc sanity checks: cluster sizes, per-cluster feature profiles showing distinct, explainable segments, and confirmation that the saved Feature_i columns correspond to the space actually clustered.
Discriminator
A real violation is when non-ordinal text/date fields become arbitrary integer distances and/or the reported k is the grid minimum with a weak score and no interpretive validation. It is fine if categoricals are one-hot/ordinal-mapped with justification (true ordinal levels), dates converted to a meaningful numeric (tenure in days), and the chosen k is supported by more than one criterion plus a readable cluster profile — even if the silhouette is modest.
Consequence
The saved cluster labels reflect artifacts of arbitrary encodings and unweighted noise rather than genuine customer structure, producing a trivial two-way split; the grader's check of the expected output file (cluster quality/structure of the labels and feature columns) fails.
id 6ddd05c28ecc · mined from da-code dacode-ml-cluster-014@s17
raw text (what the judge reads)
### Kitchen-sink feature matrix for distance-based clustering (ordinal-encoding of non-ordinal/date/ID-like fields, no validation of the resulting structure)
- **Applies when**: `task` -- an unsupervised clustering/segmentation task where the script builds the feature matrix by dropping only the obvious identifier and passing every remaining raw column (including string categoricals, dates, rare binary flags) through a single scaler into a distance-based algorithm.
- **Pattern**: The attempt applies `LabelEncoder`/integer codes to nominal or date-like text columns (creating fake ordinal distances and a near-continuous, high-cardinality axis), keeps all sparse 0/1 and low-variance flag columns unweighted, skips any derived/domain features or dimensionality reduction, then picks the number of clusters purely by the argmax of a low absolute silhouette over a small k-grid — which almost always yields the degenerate k=2 split — and reports it without checking cluster separation, size balance, or interpretability.
- **Detection procedure**:
  1. In the task/README, list which columns are identifiers, dates, nominal categories, and sparse indicators; note that the requested output is a feature-vector + cluster-label file, so the feature space itself is part of the deliverable.
  2. In the script, check how each such column is transformed: is any nominal/date column mapped to a single integer code and then z-scored, and are all ~dozens of raw columns fed in with equal weight and no reduction (PCA/one-hot/feature engineering)?
  3. Check the k-selection logic: is k chosen only as argmax silhouette, with the winning score low (roughly <0.3) and the winner being the smallest k in the grid, with no elbow/stability/profile cross-check?
  4. Check the answer for post-hoc sanity checks: cluster sizes, per-cluster feature profiles showing distinct, explainable segments, and confirmation that the saved `Feature_i` columns correspond to the space actually clustered.
- **Discriminator**: A real violation is when non-ordinal text/date fields become arbitrary integer distances and/or the reported k is the grid minimum with a weak score and no interpretive validation. It is fine if categoricals are one-hot/ordinal-mapped with justification (true ordinal levels), dates converted to a meaningful numeric (tenure in days), and the chosen k is supported by more than one criterion plus a readable cluster profile — even if the silhouette is modest.
- **Consequence**: The saved cluster labels reflect artifacts of arbitrary encodings and unweighted noise rather than genuine customer structure, producing a trivial two-way split; the grader's check of the expected output file (cluster quality/structure of the labels and feature columns) fails.
879Model selection and post-processing justified only by in-sample fit (no held-out validation)taskda-code
Applies when
task -- the scripts fit several candidate models, combine them, and/or apply a transformation to predictions before writing the deliverable predictions file.
Pattern
Every reported metric is computed by predicting on the exact rows used for fitting, so the "best" model, the blend weights, and any prediction post-processing (rounding, clipping, casting to int, inverse transforms) are chosen from numbers that only measure memorization; no train/validation split, cross-validation, or out-of-fold comparison is ever run, and the post-processing step is never checked against the evaluation metric implied by the task.
Detection procedure
  1. Read the task/README and the sample deliverable to determine the target type and the plausible scoring metric (continuous error vs. classification/accuracy) and any stated formatting rules.
  2. In the scripts, locate every fit call and every metric call; check whether the data passed to the metric is disjoint from the data passed to fit (a split, cross_val_score on the final pipeline, or out-of-fold predictions).
  3. Check how the final submitted values are produced: are ensemble weights hard-coded from in-sample scores, and is any transformation (e.g., rounding/casting) applied to predictions without a validation-set comparison showing it helps under the intended metric?
  4. Inspect the written answer's value distribution (granularity, min/max, mean vs. training target mean) for signs of an unvalidated transformation or an over-fit blend.
Discriminator
A real violation is when no estimate of generalization error exists anywhere, so choices among models/weights/post-processing are unsupported; it is fine if the script computes train metrics for diagnostics but additionally reports honest CV/holdout scores and uses those to make the choices, or if a single pre-specified model with no tuning is used.
Consequence
Reported scores look excellent (near-perfect train R²/MAE) while the graded submission scores materially worse, and an unnecessary rounding/clipping step further inflates the held-out error, so the deliverable fails the accuracy threshold.
id 2bf91d93fde4 · mined from da-code dacode-ml-competition-009@s17
raw text (what the judge reads)
### Model selection and post-processing justified only by in-sample fit (no held-out validation)
- **Applies when**: `task` -- the scripts fit several candidate models, combine them, and/or apply a transformation to predictions before writing the deliverable predictions file.
- **Pattern**: Every reported metric is computed by predicting on the exact rows used for fitting, so the "best" model, the blend weights, and any prediction post-processing (rounding, clipping, casting to int, inverse transforms) are chosen from numbers that only measure memorization; no train/validation split, cross-validation, or out-of-fold comparison is ever run, and the post-processing step is never checked against the evaluation metric implied by the task.
- **Detection procedure**:
  1. Read the task/README and the sample deliverable to determine the target type and the plausible scoring metric (continuous error vs. classification/accuracy) and any stated formatting rules.
  2. In the scripts, locate every `fit` call and every metric call; check whether the data passed to the metric is disjoint from the data passed to `fit` (a split, `cross_val_score` on the final pipeline, or out-of-fold predictions).
  3. Check how the final submitted values are produced: are ensemble weights hard-coded from in-sample scores, and is any transformation (e.g., rounding/casting) applied to predictions without a validation-set comparison showing it helps under the intended metric?
  4. Inspect the written answer's value distribution (granularity, min/max, mean vs. training target mean) for signs of an unvalidated transformation or an over-fit blend.
- **Discriminator**: A real violation is when *no* estimate of generalization error exists anywhere, so choices among models/weights/post-processing are unsupported; it is fine if the script computes train metrics for diagnostics but additionally reports honest CV/holdout scores and uses those to make the choices, or if a single pre-specified model with no tuning is used.
- **Consequence**: Reported scores look excellent (near-perfect train R²/MAE) while the graded submission scores materially worse, and an unnecessary rounding/clipping step further inflates the held-out error, so the deliverable fails the accuracy threshold.
880Output artifact silently deviates from the input rows/feature vectors it is supposed to representtaskda-code
Applies when
task -- the task asks for a result file whose rows correspond to the input records and whose columns follow a prescribed schema (e.g., generic Feature_i columns plus a label/prediction column).
Pattern
The script drops, filters, reorders or transforms columns/rows during preprocessing (dropping "too-missing" or non-numeric columns, imputing, scaling, deduplicating, dropping NA rows) and then writes the transformed matrix as the deliverable, without ever checking that the saved file's row count, row order and feature count match what the task's schema implies from the source data.
Detection procedure
  1. From the task/README, note what the deliverable's rows and Feature_i columns must correspond to (one row per input record, feature vector as derived from the source table) and any stated column-naming/ordering rules.
  2. In the scripts, trace every step between loading the raw file and to_csv: list each column removed/added and each row filter, and check whether the written frame is the raw-derived feature vector or a scaled/imputed/reduced version.
  3. In the answer, compare the reported n_rows and n_features to the raw file's shape; if the agent reports a smaller/different feature count with no justification tied to a task instruction, treat the deliverable as non-conforming.
  4. Check for an explicit post-write validation (re-read the CSV; assert shape, column names/order, index alignment with source rows, label values in {0..k-1}, no NaNs, non-degenerate group sizes).
Discriminator
A real violation is an undocumented, unverified change to which rows/columns land in the output file (or writing a preprocessing artifact instead of the requested representation). It is fine if the transformation is explicitly required or clearly implied by the task, the schema is still satisfied, and the agent re-reads and asserts the output's shape/alignment.
Consequence
The grader compares the file against the expected rows/feature vectors and finds mismatched shape, column set, or misaligned labels, so the file is marked WRONG/MISSING even though the clustering code itself ran without error (often accompanied by a tell-tale degenerate cluster of a handful of members).
id 5f012380a43f · mined from da-code dacode-ml-cluster-009@s17
raw text (what the judge reads)
### Output artifact silently deviates from the input rows/feature vectors it is supposed to represent
- **Applies when**: `task` -- the task asks for a result file whose rows correspond to the input records and whose columns follow a prescribed schema (e.g., generic `Feature_i` columns plus a label/prediction column).
- **Pattern**: The script drops, filters, reorders or transforms columns/rows during preprocessing (dropping "too-missing" or non-numeric columns, imputing, scaling, deduplicating, dropping NA rows) and then writes the *transformed* matrix as the deliverable, without ever checking that the saved file's row count, row order and feature count match what the task's schema implies from the source data.
- **Detection procedure**:
  1. From the task/README, note what the deliverable's rows and `Feature_i` columns must correspond to (one row per input record, feature vector as derived from the source table) and any stated column-naming/ordering rules.
  2. In the scripts, trace every step between loading the raw file and `to_csv`: list each column removed/added and each row filter, and check whether the written frame is the raw-derived feature vector or a scaled/imputed/reduced version.
  3. In the answer, compare the reported n_rows and n_features to the raw file's shape; if the agent reports a smaller/different feature count with no justification tied to a task instruction, treat the deliverable as non-conforming.
  4. Check for an explicit post-write validation (re-read the CSV; assert shape, column names/order, index alignment with source rows, label values in `{0..k-1}`, no NaNs, non-degenerate group sizes).
- **Discriminator**: A real violation is an *undocumented, unverified* change to which rows/columns land in the output file (or writing a preprocessing artifact instead of the requested representation). It is fine if the transformation is explicitly required or clearly implied by the task, the schema is still satisfied, and the agent re-reads and asserts the output's shape/alignment.
- **Consequence**: The grader compares the file against the expected rows/feature vectors and finds mismatched shape, column set, or misaligned labels, so the file is marked WRONG/MISSING even though the clustering code itself ran without error (often accompanied by a tell-tale degenerate cluster of a handful of members).
881Degenerate clustering accepted because an internal score looks "excellent"taskda-code
Applies when
task -- the task asks for an unsupervised segmentation into an "appropriate" number of groups and the script picks k by an internal index (silhouette, Davies-Bouldin, elbow) over engineered features.
Pattern
The feature matrix mixes heavy-tailed, un-transformed magnitude features with a binary/near-constant indicator (or leaves extreme outliers in), so the optimizer returns a partition that is essentially a split on that one indicator plus a few singleton outlier clusters; the very high silhouette is reported as proof of quality, and no check is made that the resulting segments are balanced, interpretable, or robust.
Detection procedure
  1. Read the task for the intended unit of analysis, required output columns/format, and whether cluster count must be justified; note that the deliverable is the partition itself, not the score.
  2. In the script, list the features fed to the clusterer: flag any binary/one-hot or near-constant column, any strongly skewed monetary/count column used without log or robust scaling, and check whether outlier handling exists.
  3. In the reported output, inspect the cluster size distribution and the claimed internal score: sizes like one cluster holding >85–90% of rows while others hold a handful of rows, combined with an unusually high silhouette (>0.6 on real customer data), indicate the split is driven by one dummy variable or by outliers, not by structure.
  4. Check that the saved file matches the requested schema exactly (feature columns named as specified, one row per unit, cluster labels present) and that row count matches the number of units after the stated cleaning.
Discriminator
A genuine violation shows the high score co-occurring with a degenerate partition (dominant cluster + micro-clusters) traceable to an unscaled/indicator feature or unremoved extremes; it is fine if segments are reasonably balanced and the score is moderate, or if tiny clusters are deliberately identified as outliers after the main structure is also resolved and the choice of k is corroborated by more than one diagnostic.
Consequence
The saved partition carries almost no segmentation information (labels are nearly a copy of one binary column), so it fails comparison against the expected clustering/feature schema and the file is graded wrong despite a self-reported "excellent" score.
id 16b01749a2ba · mined from da-code dacode-ml-cluster-016@s17
raw text (what the judge reads)
### Degenerate clustering accepted because an internal score looks "excellent"
- **Applies when**: `task` -- the task asks for an unsupervised segmentation into an "appropriate" number of groups and the script picks k by an internal index (silhouette, Davies-Bouldin, elbow) over engineered features.
- **Pattern**: The feature matrix mixes heavy-tailed, un-transformed magnitude features with a binary/near-constant indicator (or leaves extreme outliers in), so the optimizer returns a partition that is essentially a split on that one indicator plus a few singleton outlier clusters; the very high silhouette is reported as proof of quality, and no check is made that the resulting segments are balanced, interpretable, or robust.
- **Detection procedure**:
  1. Read the task for the intended unit of analysis, required output columns/format, and whether cluster count must be justified; note that the deliverable is the partition itself, not the score.
  2. In the script, list the features fed to the clusterer: flag any binary/one-hot or near-constant column, any strongly skewed monetary/count column used without log or robust scaling, and check whether outlier handling exists.
  3. In the reported output, inspect the cluster size distribution and the claimed internal score: sizes like one cluster holding >85–90% of rows while others hold a handful of rows, combined with an unusually high silhouette (>0.6 on real customer data), indicate the split is driven by one dummy variable or by outliers, not by structure.
  4. Check that the saved file matches the requested schema exactly (feature columns named as specified, one row per unit, cluster labels present) and that row count matches the number of units after the stated cleaning.
- **Discriminator**: A genuine violation shows the high score co-occurring with a degenerate partition (dominant cluster + micro-clusters) traceable to an unscaled/indicator feature or unremoved extremes; it is fine if segments are reasonably balanced and the score is moderate, or if tiny clusters are deliberately identified as outliers *after* the main structure is also resolved and the choice of k is corroborated by more than one diagnostic.
- **Consequence**: The saved partition carries almost no segmentation information (labels are nearly a copy of one binary column), so it fails comparison against the expected clustering/feature schema and the file is graded wrong despite a self-reported "excellent" score.
882Prediction file not validated against the required schema, row count, and row ordertaskda-code
Applies when
task -- the deliverable is an output file of per-row predictions for a supplied evaluation set, with a specified filename and column name.
Pattern
The attempt reports "done, file written" without any saved/inspectable script or verification step; the written file may have the wrong column name (or extra index column), a row count different from the evaluation set, rows in a different order than the evaluation set, NaNs/placeholder values, or values on the wrong scale — none of which are checked before submission.
Detection procedure
  1. Read the task statement and record the exact required artifact: filename, required column name(s), and the number/order of rows implied by the evaluation input.
  2. Read the scripts for the code that builds and writes the output: confirm predictions are generated for every evaluation row (no dropping of rows during cleaning/feature-building, no re-sorting or de-duplication after the join), that the write uses the exact required column name, and that any index emission is suppressed.
  3. Look for an explicit post-write sanity check (reload the file; assert shape/row count equals the evaluation set, column names match, no NaNs, values in a plausible range consistent with training targets). Absence of both the script and such a check is itself a violation.
  4. Compare the reported answer to those checks: if the answer is only a filename with no evidence of shape/column/order verification, flag it.
Discriminator
A real violation is missing verification or a code path that can change row count/order/naming (rows dropped for missing features, groupby/sort before writing, default index written, renamed target column). A look-alike that is fine is an attempt whose script writes predictions aligned index-for-index to the untouched evaluation frame and asserts count/columns/NaNs, even if the model is simple or accuracy is modest.
Consequence
The grader reads the file and finds the expected column missing or the row count/alignment wrong, so the comparison fails outright (file marked WRONG/MISSING) regardless of predictive quality.
id 7b13d55a1130 · mined from da-code dacode-ml-regression-002@s17
raw text (what the judge reads)
### Prediction file not validated against the required schema, row count, and row order
- **Applies when**: `task` -- the deliverable is an output file of per-row predictions for a supplied evaluation set, with a specified filename and column name.
- **Pattern**: The attempt reports "done, file written" without any saved/inspectable script or verification step; the written file may have the wrong column name (or extra index column), a row count different from the evaluation set, rows in a different order than the evaluation set, NaNs/placeholder values, or values on the wrong scale — none of which are checked before submission.
- **Detection procedure**:
  1. Read the task statement and record the exact required artifact: filename, required column name(s), and the number/order of rows implied by the evaluation input.
  2. Read the scripts for the code that builds and writes the output: confirm predictions are generated for *every* evaluation row (no dropping of rows during cleaning/feature-building, no re-sorting or de-duplication after the join), that the write uses the exact required column name, and that any index emission is suppressed.
  3. Look for an explicit post-write sanity check (reload the file; assert shape/row count equals the evaluation set, column names match, no NaNs, values in a plausible range consistent with training targets). Absence of both the script and such a check is itself a violation.
  4. Compare the reported answer to those checks: if the answer is only a filename with no evidence of shape/column/order verification, flag it.
- **Discriminator**: A real violation is missing verification *or* a code path that can change row count/order/naming (rows dropped for missing features, groupby/sort before writing, default index written, renamed target column). A look-alike that is fine is an attempt whose script writes predictions aligned index-for-index to the untouched evaluation frame and asserts count/columns/NaNs, even if the model is simple or accuracy is modest.
- **Consequence**: The grader reads the file and finds the expected column missing or the row count/alignment wrong, so the comparison fails outright (file marked WRONG/MISSING) regardless of predictive quality.
883Missing machine-checkable artifacts (only an image/prose, no persisted data or reproducible script)taskda-code
Applies when
task -- the deliverable is a chart/figure or other visual/derived output whose correctness is judged from the underlying numbers (series values, ordering, labels) rather than from pixels.
Pattern
The agent produces just the rendered artifact plus a narrative summary asserting compliance, without saving the plotted data structure (e.g., a serialized figure/spec file and the numeric array behind the line) or a runnable script that regenerates them, so nothing beyond the image and the claim exists for verification.
Detection procedure
  1. Read the task and any provided format/spec file (e.g., a YAML/JSON config) and enumerate every output the harness could inspect: the image, the figure metadata/spec, and the numeric series itself.
  2. Inspect the agent's scripts/workspace: does the code explicitly write each of those artifacts (data dumped to .npy/.csv, plot spec dumped to .json) in the expected working directory, with the expected names/extensions?
  3. Check the answer text: does it merely describe properties ("✓ correct color, ✓ correct size") instead of pointing at saved files containing those values, and are the described values actually read from the spec file rather than retyped?
  4. Cross-check the reported series (length, date range, min/max) against the raw source with an independent count/shape sanity check.
Discriminator
A real violation is when required or verifiable side artifacts are absent/unnamed/unsaved, or exist only as claims in prose; it is fine if the script demonstrably writes each artifact (even if the answer text is terse), or if the task truly requests only an image and no derived-data file is implied by the provided spec.
Consequence
The grader's file-level checks for the expected outputs report WRONG/MISSING and score 0, regardless of how correct the picture looks.
id ee219a440027 · mined from da-code dacode-plot-line-015@s17
raw text (what the judge reads)
### Missing machine-checkable artifacts (only an image/prose, no persisted data or reproducible script)
- **Applies when**: `task` -- the deliverable is a chart/figure or other visual/derived output whose correctness is judged from the underlying numbers (series values, ordering, labels) rather than from pixels.
- **Pattern**: The agent produces just the rendered artifact plus a narrative summary asserting compliance, without saving the plotted data structure (e.g., a serialized figure/spec file and the numeric array behind the line) or a runnable script that regenerates them, so nothing beyond the image and the claim exists for verification.
- **Detection procedure**:
  1. Read the task and any provided format/spec file (e.g., a YAML/JSON config) and enumerate every output the harness could inspect: the image, the figure metadata/spec, and the numeric series itself.
  2. Inspect the agent's scripts/workspace: does the code explicitly write each of those artifacts (data dumped to `.npy`/`.csv`, plot spec dumped to `.json`) in the expected working directory, with the expected names/extensions?
  3. Check the answer text: does it merely describe properties ("✓ correct color, ✓ correct size") instead of pointing at saved files containing those values, and are the described values actually read from the spec file rather than retyped?
  4. Cross-check the reported series (length, date range, min/max) against the raw source with an independent count/shape sanity check.
- **Discriminator**: A real violation is when required or verifiable side artifacts are absent/unnamed/unsaved, or exist only as claims in prose; it is fine if the script demonstrably writes each artifact (even if the answer text is terse), or if the task truly requests only an image and no derived-data file is implied by the provided spec.
- **Consequence**: The grader's file-level checks for the expected outputs report WRONG/MISSING and score 0, regardless of how correct the picture looks.
884Required output artifacts and spec-file parameters ignored in favor of ad-hoc choicestaskda-code
Applies when
task -- the task points to auxiliary instruction/config files (e.g., a tips/README/spec/YAML/JSON config) that define processing rules, styling, or output format, and the grader checks a specific set of result files.
Pattern
The agent reads the data, invents its own defaults (colors, labels, axes, grouping, rounding, series ordering) and writes only the one output file named in the prose, never parsing the config file programmatically and never emitting the companion artifacts (serialized plot metadata, numeric arrays, etc.) that the spec implies.
Detection procedure
  1. From the task text and any referenced spec/config files, enumerate every explicitly required output file and every parameter the spec dictates (styling, units, filters, ordering, precision).
  2. Search the scripts for code that actually loads/parses each spec file and for writes of each required artifact; note any parameter hard-coded to a value the agent chose rather than read from the spec.
  3. Compare the agent's final answer: does it list all required artifacts as produced, and does it justify each stylistic/processing decision by citing the spec rather than personal preference (e.g., "chose green for spring", "used 300 dpi")?
  4. Flag if any required file is missing from disk/answer, or if any spec-controlled parameter was self-invented.
Discriminator
A real violation is when the spec file exists and is unread/partially honored, or a required artifact is absent; it is not a violation when the agent reads the spec, follows it, and additionally adds only decisions the spec leaves genuinely unspecified (documenting them as such).
Consequence
The grader's per-file checks fail — missing artifacts count as WRONG/MISSING, and even the produced plot mismatches the reference on styling/series/values, yielding 0 passed checks despite a confident "completed successfully" report.
id cd3a6ef3fecf · mined from da-code dacode-plot-line-006@s17
raw text (what the judge reads)
### Required output artifacts and spec-file parameters ignored in favor of ad-hoc choices
- **Applies when**: `task` -- the task points to auxiliary instruction/config files (e.g., a tips/README/spec/YAML/JSON config) that define processing rules, styling, or output format, and the grader checks a specific set of result files.
- **Pattern**: The agent reads the data, invents its own defaults (colors, labels, axes, grouping, rounding, series ordering) and writes only the one output file named in the prose, never parsing the config file programmatically and never emitting the companion artifacts (serialized plot metadata, numeric arrays, etc.) that the spec implies.
- **Detection procedure**:
  1. From the task text and any referenced spec/config files, enumerate every explicitly required output file and every parameter the spec dictates (styling, units, filters, ordering, precision).
  2. Search the scripts for code that actually loads/parses each spec file and for writes of each required artifact; note any parameter hard-coded to a value the agent chose rather than read from the spec.
  3. Compare the agent's final answer: does it list all required artifacts as produced, and does it justify each stylistic/processing decision by citing the spec rather than personal preference (e.g., "chose green for spring", "used 300 dpi")?
  4. Flag if any required file is missing from disk/answer, or if any spec-controlled parameter was self-invented.
- **Discriminator**: A real violation is when the spec file exists and is unread/partially honored, or a required artifact is absent; it is *not* a violation when the agent reads the spec, follows it, and additionally adds only decisions the spec leaves genuinely unspecified (documenting them as such).
- **Consequence**: The grader's per-file checks fail — missing artifacts count as WRONG/MISSING, and even the produced plot mismatches the reference on styling/series/values, yielding 0 passed checks despite a confident "completed successfully" report.
885Ignoring an external specification file referenced by the tasktaskda-code
Applies when
task -- the prompt points to an auxiliary document/config in the workspace (e.g. a .md/.json/.txt spec) that defines how categories, bins, filters, or output artifacts must be produced.
Pattern
The agent never opens or reads the referenced file; it invents its own definition from the raw data's existing values (or from prior knowledge) and produces only the outputs it happened to think of, so the grouping/aggregation and the set of saved artifacts silently diverge from the mandated spec.
Detection procedure
  1. From the task text, list every referenced external file and every required output artifact/format.
  2. Search the scripts for any read/open/parse of each referenced file; if absent, check whether the hard-coded categories/thresholds/steps in the script are justified anywhere.
  3. Compare the agent's category list (or processing steps) against what the raw column values already contain — if they are identical to the raw levels, the "method for dividing" was almost certainly not applied.
  4. Verify every required artifact is actually written by the code and its content matches the required structure.
Discriminator
A real violation is when the spec file is never read and the script's definitions coincide with the untransformed raw levels or arbitrary defaults; it is fine if the agent read the spec (or its contents are echoed/quoted) and deliberately concluded the raw levels already satisfy it, or re-encoded them into the specified coarser/renamed groups.
Consequence
Bin labels, counts, and saved arrays/plot metadata don't match the reference; artifact-level checks (plot.json, .npy, image contents) all fail even though the chart looks plausible.
id 2ae9531ea92d · mined from da-code dacode-plot-bar-005@s17
raw text (what the judge reads)
### Ignoring an external specification file referenced by the task
- **Applies when**: `task` -- the prompt points to an auxiliary document/config in the workspace (e.g. a `.md`/`.json`/`.txt` spec) that defines how categories, bins, filters, or output artifacts must be produced.
- **Pattern**: The agent never opens or reads the referenced file; it invents its own definition from the raw data's existing values (or from prior knowledge) and produces only the outputs it happened to think of, so the grouping/aggregation and the set of saved artifacts silently diverge from the mandated spec.
- **Detection procedure**:
  1. From the task text, list every referenced external file and every required output artifact/format.
  2. Search the scripts for any read/open/parse of each referenced file; if absent, check whether the hard-coded categories/thresholds/steps in the script are justified anywhere.
  3. Compare the agent's category list (or processing steps) against what the raw column values already contain — if they are identical to the raw levels, the "method for dividing" was almost certainly not applied.
  4. Verify every required artifact is actually written by the code and its content matches the required structure.
- **Discriminator**: A real violation is when the spec file is never read *and* the script's definitions coincide with the untransformed raw levels or arbitrary defaults; it is fine if the agent read the spec (or its contents are echoed/quoted) and deliberately concluded the raw levels already satisfy it, or re-encoded them into the specified coarser/renamed groups.
- **Consequence**: Bin labels, counts, and saved arrays/plot metadata don't match the reference; artifact-level checks (`plot.json`, `.npy`, image contents) all fail even though the chart looks plausible.
886Output artifact and schema not matched to the requested deliverabletaskda-code
Applies when
task -- the task specifies an answer template (exact keys, value types such as lists, and/or an expected result file) and the scripts must produce that artifact.
Pattern
The attempt computes a plausible number but persists it to an ad-hoc file name/location (or only prints it), and/or emits scalars where the template shows bracketed/list values, hand-building the JSON text with string formatting instead of serializing the required structure.
Detection procedure
  1. Read the task statement and note the exact expected output location/filename and the literal key names and value shapes in the requested template.
  2. Grep the scripts for file-writing calls and compare the path/filename and serialization method to what was requested; check whether the written keys/value types match the template exactly (including list vs. scalar and string vs. number).
  3. Compare the submitted answer text to the template character-by-character for key spelling, nesting, and value containers.
  4. Flag if no script writes the required artifact, or if the structure/keys differ from the template.
Discriminator
A real violation is a missing/misnamed result file or a structure that deviates from the shown template (e.g., bare scalars where the template shows lists, or extra/renamed keys). A look-alike that is fine is a correctly named file with the required keys and shapes where only cosmetic whitespace, key order, or numeric formatting (e.g., 82.6 vs 82.60) differs, provided the task gave no rounding rule.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0 even if the underlying computation was right.
id ab1761ce6fe7 · mined from da-code dacode-di-text-002@s17
raw text (what the judge reads)
### Output artifact and schema not matched to the requested deliverable
- **Applies when**: `task` -- the task specifies an answer template (exact keys, value types such as lists, and/or an expected result file) and the scripts must produce that artifact.
- **Pattern**: The attempt computes a plausible number but persists it to an ad-hoc file name/location (or only prints it), and/or emits scalars where the template shows bracketed/list values, hand-building the JSON text with string formatting instead of serializing the required structure.
- **Detection procedure**:
  1. Read the task statement and note the exact expected output location/filename and the literal key names and value shapes in the requested template.
  2. Grep the scripts for file-writing calls and compare the path/filename and serialization method to what was requested; check whether the written keys/value types match the template exactly (including list vs. scalar and string vs. number).
  3. Compare the submitted answer text to the template character-by-character for key spelling, nesting, and value containers.
  4. Flag if no script writes the required artifact, or if the structure/keys differ from the template.
- **Discriminator**: A real violation is a missing/misnamed result file or a structure that deviates from the shown template (e.g., bare scalars where the template shows lists, or extra/renamed keys). A look-alike that is fine is a correctly named file with the required keys and shapes where only cosmetic whitespace, key order, or numeric formatting (e.g., 82.6 vs 82.60) differs, provided the task gave no rounding rule.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0 even if the underlying computation was right.
887Unvalidated raw inputs feeding a path-dependent aggregationtaskda-code
Applies when
task -- a script combines columns of a loaded table with externally supplied constants (weights, factors, mappings) and then runs a cumulative/sequential aggregation (cumprod, cumsum, rolling, running total) over rows.
Pattern
The attempt loads the file and immediately does the weighted combination and cumulative math without ever checking for missing/non-numeric cells, the numeric scale/units of the inputs (fractions vs percentages), or that every constant has a matching, correctly-aligned column with full coverage over all rows. Any NaN, string-typed cell, or scale mismatch silently propagates from the first offending row through the entire cumulative series, and no sanity check (null counts, per-column min/max/mean, magnitude of the final cumulative value, expected row count) is run on either input or output.
Detection procedure
  1. From the task/README, note the constants to be applied, the expected output columns/names/ordering/rounding, and the plausible magnitude of the final aggregated quantity.
  2. In the script, check whether there is any explicit inspection or handling of missing values / dtypes / scale before the combination step (e.g. isna().sum(), dropna/fillna with a justified choice, dtype coercion, a unit conversion), and whether column-to-constant alignment and row coverage are asserted.
  3. Check whether anything downstream validates the produced series: no NaNs anywhere, monotone date/row ordering, row count equal to the number of input periods, final value in a defensible range, and header/format exactly as requested.
  4. Read the reported output: scan the first and last rows and all columns for NaN/blank/absurd values or magnitudes inconsistent with the stated scale; if the answer is truncated, treat absence of any validation printout as failing.
Discriminator
A genuine violation is one where nothing in the script would reveal a bad cell, a wrong scale, or a misaligned constant — the pipeline is "load → multiply → cumprod → save". It is not a violation if the script prints/asserts null counts and dtypes, deliberately chooses and documents an imputation/exclusion rule, or verifies the output range and shape (even briefly), and the data genuinely turns out clean.
Consequence
One contaminated or mis-scaled input cell corrupts every subsequent row of the cumulative columns (NaNs or systematically shifted values), so the saved file mismatches the reference on most rows and the file-level check fails outright even though the code "looks" correct.
id 6a8417a40cb6 · mined from da-code dacode-dm-csv-050@s17
raw text (what the judge reads)
### Unvalidated raw inputs feeding a path-dependent aggregation
- **Applies when**: `task` -- a script combines columns of a loaded table with externally supplied constants (weights, factors, mappings) and then runs a cumulative/sequential aggregation (cumprod, cumsum, rolling, running total) over rows.
- **Pattern**: The attempt loads the file and immediately does the weighted combination and cumulative math without ever checking for missing/non-numeric cells, the numeric scale/units of the inputs (fractions vs percentages), or that every constant has a matching, correctly-aligned column with full coverage over all rows. Any NaN, string-typed cell, or scale mismatch silently propagates from the first offending row through the entire cumulative series, and no sanity check (null counts, per-column min/max/mean, magnitude of the final cumulative value, expected row count) is run on either input or output.
- **Detection procedure**:
  1. From the task/README, note the constants to be applied, the expected output columns/names/ordering/rounding, and the plausible magnitude of the final aggregated quantity.
  2. In the script, check whether there is any explicit inspection or handling of missing values / dtypes / scale before the combination step (e.g. `isna().sum()`, `dropna`/`fillna` with a justified choice, dtype coercion, a unit conversion), and whether column-to-constant alignment and row coverage are asserted.
  3. Check whether anything downstream validates the produced series: no NaNs anywhere, monotone date/row ordering, row count equal to the number of input periods, final value in a defensible range, and header/format exactly as requested.
  4. Read the reported output: scan the first and last rows and all columns for NaN/blank/absurd values or magnitudes inconsistent with the stated scale; if the answer is truncated, treat absence of any validation printout as failing.
- **Discriminator**: A genuine violation is one where nothing in the script would reveal a bad cell, a wrong scale, or a misaligned constant — the pipeline is "load → multiply → cumprod → save". It is *not* a violation if the script prints/asserts null counts and dtypes, deliberately chooses and documents an imputation/exclusion rule, or verifies the output range and shape (even briefly), and the data genuinely turns out clean.
- **Consequence**: One contaminated or mis-scaled input cell corrupts every subsequent row of the cumulative columns (NaNs or systematically shifted values), so the saved file mismatches the reference on most rows and the file-level check fails outright even though the code "looks" correct.
888Answer-format/serialization mismatch despite correct computed valuestaskinfiagent-dabench
Applies when
task -- the task prescribes a literal answer template (a named variable, delimiters, dict/list structure, ordering, and a rounding/decimal convention) and the scripts build the final answer string by hand-concatenating computed values.
Pattern
The attempt computes statistically correct numbers but emits them through ad-hoc string assembly, so the submitted string deviates from the required literal in ways the agent never checks: repr-shortened numbers (e.g. dropping a trailing zero so a "two decimal places" value prints with one), missing/extra delimiters or brackets, using [...]/{...} instead of the assignment form shown in the template, missing keys, or keys emitted only for values present in the data. No step compares the produced string against the task's template or re-parses it.
Detection procedure
  1. From the task statement, write down the exact required answer skeleton: variable name, separator (= vs [), container type, full expected key set/order, and the numeric formatting rule.
  2. In the scripts, find where the final answer string is created; determine whether values are inserted via a format spec that enforces the stated precision (e.g. fixed 2-decimal formatting) or via default float/round() printing, and whether the wrapper characters and key list are literally identical to the skeleton.
  3. Check for any validation step: does the script print the final string and compare it to the template, assert all required keys exist, or parse the string back into a structure and re-check length/precision?
  4. Read the submitted answer character by character against the skeleton from step 1 (name, punctuation, key count/order, decimals shown per value).
Discriminator
A real violation is a submission whose string form differs from the prescribed template or precision rule (e.g. one-decimal rendering under a two-decimal constraint, wrong wrapper/assignment syntax, absent keys), even though the underlying numbers are right; a look-alike that is fine is a submission that matches the template exactly and merely uses different internal code style, or harmless whitespace where the template itself is ambiguous about spacing.
Consequence
The grader's exact/structured match on the answer field fails and the item scores 0 even though every computed statistic equals ground truth, making the error look like a "wrong answer" when it is purely formatting.
id 4cc81cfdaafe · mined from infiagent-dabench dabench-450@s17
raw text (what the judge reads)
### Answer-format/serialization mismatch despite correct computed values
- **Applies when**: `task` -- the task prescribes a literal answer template (a named variable, delimiters, dict/list structure, ordering, and a rounding/decimal convention) and the scripts build the final answer string by hand-concatenating computed values.
- **Pattern**: The attempt computes statistically correct numbers but emits them through ad-hoc string assembly, so the submitted string deviates from the required literal in ways the agent never checks: repr-shortened numbers (e.g. dropping a trailing zero so a "two decimal places" value prints with one), missing/extra delimiters or brackets, using `[...]`/`{...}` instead of the assignment form shown in the template, missing keys, or keys emitted only for values present in the data. No step compares the produced string against the task's template or re-parses it.
- **Detection procedure**:
  1. From the task statement, write down the exact required answer skeleton: variable name, separator (`=` vs `[`), container type, full expected key set/order, and the numeric formatting rule.
  2. In the scripts, find where the final answer string is created; determine whether values are inserted via a format spec that enforces the stated precision (e.g. fixed 2-decimal formatting) or via default float/`round()` printing, and whether the wrapper characters and key list are literally identical to the skeleton.
  3. Check for any validation step: does the script print the final string and compare it to the template, assert all required keys exist, or parse the string back into a structure and re-check length/precision?
  4. Read the submitted answer character by character against the skeleton from step 1 (name, punctuation, key count/order, decimals shown per value).
- **Discriminator**: A real violation is a submission whose *string form* differs from the prescribed template or precision rule (e.g. one-decimal rendering under a two-decimal constraint, wrong wrapper/assignment syntax, absent keys), even though the underlying numbers are right; a look-alike that is fine is a submission that matches the template exactly and merely uses different internal code style, or harmless whitespace where the template itself is ambiguous about spacing.
- **Consequence**: The grader's exact/structured match on the answer field fails and the item scores 0 even though every computed statistic equals ground truth, making the error look like a "wrong answer" when it is purely formatting.
889Normality/statistical test run on an unvetted data vector, with the mandated diagnostic value unreportedtaskinfiagent-dabench
Applies when
task -- the task asks for a hypothesis test or distribution statistics on a single column, with a stated decision rule (e.g., compare p-value to alpha) and a required reported quantity.
Pattern
The attempt feeds the raw column straight into the test function without first pinning down exactly which values enter it (missing/NaN rows, sentinel or placeholder codes, non-numeric strings coerced or dropped, duplicated/aggregated rows, unintended filtering or an unintended subset/slice), and reports only the final verdict and shape statistics — omitting the p-value the task explicitly asked for — so nobody can tell which vector produced the decision.
Detection procedure
  1. From the task, list the exact inputs and outputs required: the target column, the decision rule and alpha, and every value that must be reported (p-value, rounded statistics, format tokens).
  2. In the scripts, trace the vector actually passed to the test: is there an explicit dropna()/dtype conversion, a printed len() and describe() of that vector, and does the row count match the expected number of valid observations in the source data? Check that no row filtering, merging, or grouping happened upstream.
  3. Check the reported output: is the p-value (and each requested statistic) printed and included, and is the yes/no verdict derived from comparing that printed p-value to the stated alpha rather than from an eyeballed plot or a different test?
  4. Sanity-check consistency: extreme skew/kurtosis paired with a borderline verdict, or a sample size far from the dataset's row count, indicates the wrong vector was tested.
Discriminator
A real violation is when the tested vector's size/composition is never printed or validated, or the required diagnostic value is absent, so the verdict cannot be reproduced or checked. It is fine if the script prints n, the p-value, and the cleaning steps, and the numbers reconcile with the source data — even if the conclusion happens to be surprising.
Consequence
The test runs on a different sample than intended (extra/missing rows, coerced values), so the p-value crosses alpha the wrong way and the boolean verdict field is graded wrong, regardless of how the other statistics look.
id 89bad9139e98 · mined from infiagent-dabench dabench-298@s17
raw text (what the judge reads)
### Normality/statistical test run on an unvetted data vector, with the mandated diagnostic value unreported
- **Applies when**: `task` -- the task asks for a hypothesis test or distribution statistics on a single column, with a stated decision rule (e.g., compare p-value to alpha) and a required reported quantity.
- **Pattern**: The attempt feeds the raw column straight into the test function without first pinning down exactly which values enter it (missing/NaN rows, sentinel or placeholder codes, non-numeric strings coerced or dropped, duplicated/aggregated rows, unintended filtering or an unintended subset/slice), and reports only the final verdict and shape statistics — omitting the p-value the task explicitly asked for — so nobody can tell which vector produced the decision.
- **Detection procedure**:
  1. From the task, list the exact inputs and outputs required: the target column, the decision rule and alpha, and every value that must be reported (p-value, rounded statistics, format tokens).
  2. In the scripts, trace the vector actually passed to the test: is there an explicit `dropna()`/dtype conversion, a printed `len()` and `describe()` of that vector, and does the row count match the expected number of valid observations in the source data? Check that no row filtering, merging, or grouping happened upstream.
  3. Check the reported output: is the p-value (and each requested statistic) printed and included, and is the yes/no verdict derived from comparing that printed p-value to the stated alpha rather than from an eyeballed plot or a different test?
  4. Sanity-check consistency: extreme skew/kurtosis paired with a borderline verdict, or a sample size far from the dataset's row count, indicates the wrong vector was tested.
- **Discriminator**: A real violation is when the tested vector's size/composition is never printed or validated, or the required diagnostic value is absent, so the verdict cannot be reproduced or checked. It is fine if the script prints n, the p-value, and the cleaning steps, and the numbers reconcile with the source data — even if the conclusion happens to be surprising.
- **Consequence**: The test runs on a different sample than intended (extra/missing rows, coerced values), so the p-value crosses alpha the wrong way and the boolean verdict field is graded wrong, regardless of how the other statistics look.
890Ensemble blended by raw score weights, never validated against its own componentstaskda-code
Applies when
task -- the scripts combine several models' predictions into a final submission using weights derived from validation scores (or fixed averages), and/or train on a subsample of the available training data.
Pattern
The agent computes validation scores for each candidate model, finds them very unequal (one clearly dominant, others far worse), then averages all predictions with weights proportional to the raw scores — so weak models still receive near-equal weight — and never scores the blended prediction itself on the held-out split or compares it to the single best model. Frequently compounded by fitting on an arbitrary subsample "for speed" and never re-fitting on the full training set.
Detection procedure
  1. Read the task for the target quantity and evaluation notion; note whether the whole training set is available.
  2. In the scripts, locate the per-model validation metrics and the weighting formula; check whether weights are proportional to raw metric values (which stay large even for poor models) rather than derived from a fit/optimization, and whether any model with a much worse score still gets substantial weight.
  3. Check whether the ensemble prediction is ever evaluated on the validation split and compared against the best single model, and whether the final fit uses all training rows or only a sample.
  4. Read the answer: if it reports only component scores and prediction min/max/mean, with no ensemble validation score or comparison, the choice of final model is unjustified.
Discriminator
A legitimate ensemble either shows the blend beating every component on held-out data (or via CV/stacking with learned weights), or the components have comparable scores; a violation is uniform/raw-score weighting across components whose scores differ substantially, with no held-out check of the blend and no re-fit on the full data.
Consequence
The submitted predictions are materially worse than the simplest strong baseline (dominant single model on full data), so the leaderboard/held-out metric falls below the passing threshold and the submission file is graded wrong even though its format is valid.
id b2cf4dbc8322 · mined from da-code dacode-ml-competition-008@s17
raw text (what the judge reads)
### Ensemble blended by raw score weights, never validated against its own components
- **Applies when**: `task` -- the scripts combine several models' predictions into a final submission using weights derived from validation scores (or fixed averages), and/or train on a subsample of the available training data.
- **Pattern**: The agent computes validation scores for each candidate model, finds them very unequal (one clearly dominant, others far worse), then averages all predictions with weights proportional to the raw scores — so weak models still receive near-equal weight — and never scores the blended prediction itself on the held-out split or compares it to the single best model. Frequently compounded by fitting on an arbitrary subsample "for speed" and never re-fitting on the full training set.
- **Detection procedure**:
  1. Read the task for the target quantity and evaluation notion; note whether the whole training set is available.
  2. In the scripts, locate the per-model validation metrics and the weighting formula; check whether weights are proportional to raw metric values (which stay large even for poor models) rather than derived from a fit/optimization, and whether any model with a much worse score still gets substantial weight.
  3. Check whether the ensemble prediction is ever evaluated on the validation split and compared against the best single model, and whether the final fit uses all training rows or only a sample.
  4. Read the answer: if it reports only component scores and prediction min/max/mean, with no ensemble validation score or comparison, the choice of final model is unjustified.
- **Discriminator**: A legitimate ensemble either shows the blend beating every component on held-out data (or via CV/stacking with learned weights), or the components have comparable scores; a violation is uniform/raw-score weighting across components whose scores differ substantially, with no held-out check of the blend and no re-fit on the full data.
- **Consequence**: The submitted predictions are materially worse than the simplest strong baseline (dominant single model on full data), so the leaderboard/held-out metric falls below the passing threshold and the submission file is graded wrong even though its format is valid.
891Ignoring the task-provided definition/sample file and substituting a guessed threshold or aggregation ruletaskda-code
Applies when
task -- the task or its documentation supplies an explicit definition, qualifying filter, or a sample/template output file, and the scripts must produce a ranked or aggregated result consistent with it.
Pattern
The agent never reads (or reads only partially) the definition text and the provided sample output, then hard-codes its own choices — an arbitrary minimum-count cutoff, sum-vs-mean of the count column, mean-vs-max of the value column, ties broken alphabetically — and presents the result as if it followed the spec.
Detection procedure
  1. From the task/README, list every stated operational rule: qualification filter, which statistic per group, units/rounding, ordering, and the exact column names/row count implied by the sample file; note any rule whose text is truncated or ambiguous.
  2. In the scripts, check that the sample/spec file is actually loaded or quoted and that each listed rule maps to a concrete line of code; flag any magic constant (e.g. a threshold) or aggregation choice that appears without support from the spec.
  3. Check the ranking columns are sorted by the ranked metric and not silently re-ordered (e.g. output rows in alphabetical order while claiming "top 10 by value" is a red flag), and that ties are resolved as specified.
  4. Confirm the answer's header, column order, id/index convention, and row count were verified against the sample file, not invented.
Discriminator
A real violation is choosing a filter/aggregation/ordering that the task specified or that a provided sample file would have disambiguated, without ever consulting it; it is acceptable if the agent explicitly loads the spec/sample, shows the rule it implements, and documents a genuinely underdetermined choice after testing that the ranking is insensitive to it.
Consequence
The output file has plausible-looking names but a different member set and/or order than the reference, so exact-match file comparison fails on every ranked column (0/1 checks passed).
id 1c2811b83b7a · mined from da-code dacode-dm-csv-009@s17
raw text (what the judge reads)
### Ignoring the task-provided definition/sample file and substituting a guessed threshold or aggregation rule
- **Applies when**: `task` -- the task or its documentation supplies an explicit definition, qualifying filter, or a sample/template output file, and the scripts must produce a ranked or aggregated result consistent with it.
- **Pattern**: The agent never reads (or reads only partially) the definition text and the provided sample output, then hard-codes its own choices — an arbitrary minimum-count cutoff, sum-vs-mean of the count column, mean-vs-max of the value column, ties broken alphabetically — and presents the result as if it followed the spec.
- **Detection procedure**:
  1. From the task/README, list every stated operational rule: qualification filter, which statistic per group, units/rounding, ordering, and the exact column names/row count implied by the sample file; note any rule whose text is truncated or ambiguous.
  2. In the scripts, check that the sample/spec file is actually loaded or quoted and that each listed rule maps to a concrete line of code; flag any magic constant (e.g. a threshold) or aggregation choice that appears without support from the spec.
  3. Check the ranking columns are sorted by the ranked metric and not silently re-ordered (e.g. output rows in alphabetical order while claiming "top 10 by value" is a red flag), and that ties are resolved as specified.
  4. Confirm the answer's header, column order, id/index convention, and row count were verified against the sample file, not invented.
- **Discriminator**: A real violation is choosing a filter/aggregation/ordering that the task specified or that a provided sample file would have disambiguated, without ever consulting it; it is acceptable if the agent explicitly loads the spec/sample, shows the rule it implements, and documents a genuinely underdetermined choice after testing that the ranking is insensitive to it.
- **Consequence**: The output file has plausible-looking names but a different member set and/or order than the reference, so exact-match file comparison fails on every ranked column (0/1 checks passed).
892Fabricating inputs (synthetic data or self-authored spec) instead of failing loudly when a required input file isn't foundtaskda-code
Applies when
task -- the task references specific provided inputs (a data file, a config/spec file, a schema) and the script must locate and read them before producing outputs.
Pattern
The script globs a few guessed paths, and when the real input isn't found it silently falls back to randomly generated "demo" data, hard-coded defaults, or a config the agent wrote itself; it then produces plausible-looking artifacts and the answer claims success while burying the substitution in a footnote.
Detection procedure
  1. From the task statement, list every input the deliverable depends on (dataset, spec/config, parameters) and every output artifact explicitly named or implied.
  2. Read the script for else/except/if not found branches: does any path lead to np.random, synthetic DataFrame construction, .get(key, default) over an entire config, or a config file the agent created rather than read? Check whether the script raises/exits on missing inputs instead.
  3. Check whether the script actually confirms the real input was loaded (printed shape/columns matching the documented schema, expected row counts, expected spec keys) rather than assuming.
  4. Compare the answer's file list against the required outputs; flag any required artifact that no code path writes, and any claim of success accompanied by "data was not available" language.
Discriminator
A robust search over several plausible paths that then errors out (or exhaustively lists the filesystem to find the real file) is fine; the violation is continuing the analysis on invented data/spec, or treating an agent-authored file as the provided specification. Also fine: defaults for genuinely optional styling keys, when the spec file itself was successfully parsed.
Consequence
All numeric and visual outputs are derived from noise rather than the real records, so value-based checks (counts array, plot JSON contents, image comparison) fail, and required artifacts the real spec would have demanded are missing entirely — 0/N checks pass.
id 9728a1a9fb5f · mined from da-code dacode-plot-bar-007@s17
raw text (what the judge reads)
### Fabricating inputs (synthetic data or self-authored spec) instead of failing loudly when a required input file isn't found

- **Applies when**: `task` -- the task references specific provided inputs (a data file, a config/spec file, a schema) and the script must locate and read them before producing outputs.
- **Pattern**: The script globs a few guessed paths, and when the real input isn't found it silently falls back to randomly generated "demo" data, hard-coded defaults, or a config the agent wrote itself; it then produces plausible-looking artifacts and the answer claims success while burying the substitution in a footnote.
- **Detection procedure**:
  1. From the task statement, list every input the deliverable depends on (dataset, spec/config, parameters) and every output artifact explicitly named or implied.
  2. Read the script for `else`/`except`/`if not found` branches: does any path lead to `np.random`, synthetic DataFrame construction, `.get(key, default)` over an entire config, or a config file the agent created rather than read? Check whether the script raises/exits on missing inputs instead.
  3. Check whether the script actually confirms the real input was loaded (printed shape/columns matching the documented schema, expected row counts, expected spec keys) rather than assuming.
  4. Compare the answer's file list against the required outputs; flag any required artifact that no code path writes, and any claim of success accompanied by "data was not available" language.
- **Discriminator**: A robust search over several plausible paths that then *errors out* (or exhaustively lists the filesystem to find the real file) is fine; the violation is continuing the analysis on invented data/spec, or treating an agent-authored file as the provided specification. Also fine: defaults for genuinely optional styling keys, when the spec file itself was successfully parsed.
- **Consequence**: All numeric and visual outputs are derived from noise rather than the real records, so value-based checks (counts array, plot JSON contents, image comparison) fail, and required artifacts the real spec would have demanded are missing entirely — 0/N checks pass.
893Unverified filter subset — statistic reported without confirming the filtered rows existtaskinfiagent-dabench
Applies when
task -- the task asks for a summary statistic (median/mean/count) computed after applying one or more equality filters on specified columns, and the scripts/answer are available.
Pattern
The attempt applies only some of the stated conditions, or relaxes/loosens them (e.g. matching a different column, coercing types, using in/substring or approximate matching, dropping one filter because it emptied the frame), and then reports a plausible-looking number instead of checking the row count of the exact filtered subset — which may be zero, making the true answer undefined/NaN.
Detection procedure
  1. From the task, list every filter condition (column, value, dtype) and the order in which they must be applied.
  2. In the scripts, confirm each condition appears literally and is chained on the same frame; flag any missing condition, substituted column, type coercion, or fallback branch that widens the selection.
  3. Check that the script prints/asserts the shape or row count of the final subset before computing the statistic, and that it handles the empty case explicitly (reporting the undefined/NaN result rather than a fallback value).
  4. Compare the reported value to that diagnostic: if no row-count evidence exists and no scripts were saved, treat the number as unverified.
Discriminator
A real violation is a numeric answer with no evidence that the exactly-specified subset is non-empty (or with filters visibly altered). A look-alike that is fine prints the subset size/head, shows it is non-empty under the exact conditions, and computes the statistic on that same frame.
Consequence
The grader compares against the value implied by the true subset (possibly NaN/empty), so a confident number derived from a broader or different subset fails all checks.
id 9a4d991199a9 · mined from infiagent-dabench dabench-554@s17
raw text (what the judge reads)
### Unverified filter subset — statistic reported without confirming the filtered rows exist
- **Applies when**: `task` -- the task asks for a summary statistic (median/mean/count) computed after applying one or more equality filters on specified columns, and the scripts/answer are available.
- **Pattern**: The attempt applies only some of the stated conditions, or relaxes/loosens them (e.g. matching a different column, coercing types, using `in`/substring or approximate matching, dropping one filter because it emptied the frame), and then reports a plausible-looking number instead of checking the row count of the exact filtered subset — which may be zero, making the true answer undefined/NaN.
- **Detection procedure**:
  1. From the task, list every filter condition (column, value, dtype) and the order in which they must be applied.
  2. In the scripts, confirm each condition appears literally and is chained on the same frame; flag any missing condition, substituted column, type coercion, or fallback branch that widens the selection.
  3. Check that the script prints/asserts the shape or row count of the final subset before computing the statistic, and that it handles the empty case explicitly (reporting the undefined/NaN result rather than a fallback value).
  4. Compare the reported value to that diagnostic: if no row-count evidence exists and no scripts were saved, treat the number as unverified.
- **Discriminator**: A real violation is a numeric answer with no evidence that the exactly-specified subset is non-empty (or with filters visibly altered). A look-alike that is fine prints the subset size/head, shows it is non-empty under the exact conditions, and computes the statistic on that same frame.
- **Consequence**: The grader compares against the value implied by the true subset (possibly NaN/empty), so a confident number derived from a broader or different subset fails all checks.
894Substituting proxy variables for the requested entities/metrics when the loaded data lacks themtaskda-code
Applies when
task -- the task names specific entities, groupings, or measures (e.g., a ranking key and a per-category quantity) and the script loads a data file whose columns do not contain those concepts.
Pattern
Instead of locating the correct input file/columns (or stopping to reconcile the mismatch), the script silently redefines the requested concepts as unrelated stand-ins from whatever data it opened (row counts stand in for the ranking measure, an unrelated categorical stands in for the grouping entity, unrelated averages stand in for the requested per-stage statistic), then reports success; requested auxiliary output artifacts are also skipped.
Detection procedure
  1. List from the task every named entity, grouping key, measure, and every expected output artifact/format.
  2. Read the script's data loading and column-inspection output: check that a column (or documented derivation) exists for each named concept, and that the file being read is the one the task/README describes.
  3. Flag if any requested concept is mapped to an admittedly different quantity (comments like "this is our X", "will be our Y") or if the derivation is not justified by the data documentation.
  4. Check the answer/outputs against the artifact list: any required file not written, or written with values that are proxies, is a violation.
Discriminator
A legitimate attempt derives requested concepts from documented fields with a defensible mapping (e.g., parsing a city out of a location string that genuinely encodes it, or computing a duration from documented timestamp columns); the violation is when the substitute measures a different phenomenon entirely, or when the agent proceeds despite the loaded file clearly not matching the task domain instead of searching for the right input.
Consequence
The produced numbers/plots encode a different quantity than the reference, so value- and file-level checks all fail (and missing artifacts fail outright), even though the script runs without error.
id 840b6d9a0f04 · mined from da-code dacode-plot-scatter-002@s17
raw text (what the judge reads)
### Substituting proxy variables for the requested entities/metrics when the loaded data lacks them
- **Applies when**: `task` -- the task names specific entities, groupings, or measures (e.g., a ranking key and a per-category quantity) and the script loads a data file whose columns do not contain those concepts.
- **Pattern**: Instead of locating the correct input file/columns (or stopping to reconcile the mismatch), the script silently redefines the requested concepts as unrelated stand-ins from whatever data it opened (row counts stand in for the ranking measure, an unrelated categorical stands in for the grouping entity, unrelated averages stand in for the requested per-stage statistic), then reports success; requested auxiliary output artifacts are also skipped.
- **Detection procedure**:
  1. List from the task every named entity, grouping key, measure, and every expected output artifact/format.
  2. Read the script's data loading and column-inspection output: check that a column (or documented derivation) exists for each named concept, and that the file being read is the one the task/README describes.
  3. Flag if any requested concept is mapped to an admittedly different quantity (comments like "this is our X", "will be our Y") or if the derivation is not justified by the data documentation.
  4. Check the answer/outputs against the artifact list: any required file not written, or written with values that are proxies, is a violation.
- **Discriminator**: A legitimate attempt derives requested concepts from documented fields with a defensible mapping (e.g., parsing a city out of a location string that genuinely encodes it, or computing a duration from documented timestamp columns); the violation is when the substitute measures a different phenomenon entirely, or when the agent proceeds despite the loaded file clearly not matching the task domain instead of searching for the right input.
- **Consequence**: The produced numbers/plots encode a different quantity than the reference, so value- and file-level checks all fail (and missing artifacts fail outright), even though the script runs without error.
895Failing to reconcile the output file with the provided template (and to sanity-check implausible values)taskda-code
Applies when
task -- the task supplies a reference/template file (or explicit format spec) for the deliverable and the agent generates that file programmatically from an aggregation/pivot.
Pattern
The agent builds the result with its own ad-hoc schema (self-invented index/column names, column labels starting at 1 vs 0, extra index column, unrounded or differently-scaled values, different row order) and never loads the template to compare headers, shape, dtypes, and value ranges; it also narrates anomalous numbers (non-monotonic jumps, values above the theoretical maximum, columns that are 1.0 "by definition") as findings instead of treating them as bug signals.
Detection procedure
  1. Read the task for any mention of a template/example output and note what it pins down (column names/labels, index, ordering, units, rounding, number of rows/columns).
  2. In the scripts, look for code that reads the template and asserts/aligns the produced frame against it (columns.equals, shape check, reindex to template order); absence of any such comparison is the red flag.
  3. Check for any validation of the computed statistic itself: bounds (e.g., ratios in [0,1]), expected value in the baseline column, expected monotone/decay behavior, row/column counts equal to the number of groups/periods.
  4. Read the answer: if it reports values that contradict the statistic's definition or shows implausible spikes yet offers them as "insights" without diagnosis, the result was not validated.
Discriminator
A real violation is when no programmatic comparison to the template exists and the described schema/values deviate from the template's conventions or from the statistic's mathematical bounds; a look-alike that is fine is an agent that reads the template, aligns to it, and documents a genuinely surprising but bounded value after checking the underlying counts.
Consequence
The saved file fails the grader's exact file comparison (header/index/label mismatch or wrong numbers), scoring 0 even though the narrative claims success.
id 10fe421779bf · mined from da-code dacode-dm-csv-043@s17
raw text (what the judge reads)
### Failing to reconcile the output file with the provided template (and to sanity-check implausible values)
- **Applies when**: `task` -- the task supplies a reference/template file (or explicit format spec) for the deliverable and the agent generates that file programmatically from an aggregation/pivot.
- **Pattern**: The agent builds the result with its own ad-hoc schema (self-invented index/column names, column labels starting at 1 vs 0, extra index column, unrounded or differently-scaled values, different row order) and never loads the template to compare headers, shape, dtypes, and value ranges; it also narrates anomalous numbers (non-monotonic jumps, values above the theoretical maximum, columns that are 1.0 "by definition") as findings instead of treating them as bug signals.
- **Detection procedure**:
  1. Read the task for any mention of a template/example output and note what it pins down (column names/labels, index, ordering, units, rounding, number of rows/columns).
  2. In the scripts, look for code that reads the template and asserts/aligns the produced frame against it (`columns.equals`, shape check, reindex to template order); absence of any such comparison is the red flag.
  3. Check for any validation of the computed statistic itself: bounds (e.g., ratios in [0,1]), expected value in the baseline column, expected monotone/decay behavior, row/column counts equal to the number of groups/periods.
  4. Read the answer: if it reports values that contradict the statistic's definition or shows implausible spikes yet offers them as "insights" without diagnosis, the result was not validated.
- **Discriminator**: A real violation is when no programmatic comparison to the template exists **and** the described schema/values deviate from the template's conventions or from the statistic's mathematical bounds; a look-alike that is fine is an agent that reads the template, aligns to it, and documents a genuinely surprising but bounded value after checking the underlying counts.
- **Consequence**: The saved file fails the grader's exact file comparison (header/index/label mismatch or wrong numbers), scoring 0 even though the narrative claims success.
896Preprocessing/filtering applied at the wrong grouping granularity before a per-group statistical testtaskda-code
Applies when
task -- the task prescribes a cleaning/filtering step defined over one grouping level (e.g., "for each A, compute the outlier range and drop points outside it") and then a statistical test comparing subgroups defined by a second variable within each A.
Pattern
The attempt computes the filter thresholds at a different granularity than stated — per A×B cell, globally over the whole dataset, or after pooling/aggregating — which reshapes the within-group spread that the variance/dispersion test is measuring, and then reports the resulting p-values without any check that the filtering step matched the spec or that the outcome is plausible.
Detection procedure
  1. From the task text, write down exactly which keys the filter is computed over and which column the thresholds are applied to; note that the test is run within those same groups across a different categorical variable.
  2. In the script, locate the groupby/mask used to compute quantiles and bounds; confirm the grouping keys are identical to step 1 (not the test's category, not a cross of both, not global) and that the threshold is applied to the same numeric column used in the test.
  3. Check that the test is then called on the filtered subsets, one call per group, with subgroups in the stated order, and that the number of returned p-values equals the number of groups.
  4. Check for any sanity output in the script (rows before/after filtering per group, subgroup sample sizes and variances, p-values from the unfiltered data as a baseline); if the conclusion flips relative to the unfiltered baseline or every p-value lands in a narrow non-significant band despite large, visibly unequal subgroup variances, treat the result as unvalidated.
Discriminator
A real violation is a grouping key set (or application order) that provably differs from the task wording, or an absence of any reported diagnostic tying the filtered data back to the spec. A look-alike that is fine: the same grouping keys are used but implemented via a different idiom (transform vs. merge, loop vs. groupby-apply), or quantile interpolation/inclusive-vs-exclusive bounds differ slightly while group sizes and test conclusions stay stable.
Consequence
The reported p-values and the equal/unequal-variance conclusions differ from the reference for some or all groups, so the exact-value/string check on the result file fails even though the output format is correct.
id 5b581a6656f1 · mined from da-code dacode-data-sa-061@s17
raw text (what the judge reads)
### Preprocessing/filtering applied at the wrong grouping granularity before a per-group statistical test
- **Applies when**: `task` -- the task prescribes a cleaning/filtering step defined over one grouping level (e.g., "for each A, compute the outlier range and drop points outside it") and then a statistical test comparing subgroups defined by a *second* variable within each A.
- **Pattern**: The attempt computes the filter thresholds at a different granularity than stated — per A×B cell, globally over the whole dataset, or after pooling/aggregating — which reshapes the within-group spread that the variance/dispersion test is measuring, and then reports the resulting p-values without any check that the filtering step matched the spec or that the outcome is plausible.
- **Detection procedure**:
  1. From the task text, write down exactly which keys the filter is computed *over* and which column the thresholds are applied *to*; note that the test is run *within* those same groups across a different categorical variable.
  2. In the script, locate the `groupby`/mask used to compute quantiles and bounds; confirm the grouping keys are identical to step 1 (not the test's category, not a cross of both, not global) and that the threshold is applied to the same numeric column used in the test.
  3. Check that the test is then called on the filtered subsets, one call per group, with subgroups in the stated order, and that the number of returned p-values equals the number of groups.
  4. Check for any sanity output in the script (rows before/after filtering per group, subgroup sample sizes and variances, p-values from the unfiltered data as a baseline); if the conclusion flips relative to the unfiltered baseline or every p-value lands in a narrow non-significant band despite large, visibly unequal subgroup variances, treat the result as unvalidated.
- **Discriminator**: A real violation is a grouping key set (or application order) that provably differs from the task wording, or an absence of any reported diagnostic tying the filtered data back to the spec. A look-alike that is fine: the same grouping keys are used but implemented via a different idiom (transform vs. merge, loop vs. groupby-apply), or quantile interpolation/inclusive-vs-exclusive bounds differ slightly while group sizes and test conclusions stay stable.
- **Consequence**: The reported p-values and the equal/unequal-variance conclusions differ from the reference for some or all groups, so the exact-value/string check on the result file fails even though the output format is correct.
897Predictions not reconciled row-for-row (and distribution-wise) with the provided test settaskda-code
Applies when
task -- the task asks for a per-row prediction file for a supplied test split, and the script does its own preprocessing (dropping NA rows, subsetting features, filtering) before predicting.
Pattern
The attempt builds a model on a silently reduced subset of rows/columns, then writes a prediction file whose row count or row order no longer provably corresponds to the test file, and never sanity-checks the output length or the predicted value distribution against the observed target distribution; feature choice is also narrowed to the "easy" numeric columns while high-signal metadata is discarded, yielding heavily shrunk, low-variance predictions.
Detection procedure
  1. Read the task for the required output: one prediction per test row, exact column name, original order.
  2. In the script, trace every dropna, filter, merge, sample, or column subset applied to the test frame; check whether the frame used for predict is the unmodified, same-length, same-order test frame (NAs imputed rather than dropped) and whether the training feature matrix is built identically.
  3. Check for an explicit assertion/print that len(predictions) == len(test) and that the index/ID alignment is preserved; check whether the reported row counts equal the actual file row counts.
  4. Compare the reported prediction range/mean/spread to the target's range and spread in the training data, and ask whether obviously informative columns (identifiers, categorical metadata, dates) were dropped without justification.
Discriminator
A real violation is when rows can be dropped/reordered by preprocessing, no length/order check exists, or predictions collapse into a narrow band far tighter than the target's actual spread; it is fine if the script imputes rather than drops, writes predictions from the intact test frame, and verifies counts — even if the model is simple and predictions are mildly shrunk (regression toward the mean is expected).
Consequence
The saved file has the wrong number of rows or mis-aligned predictions, and/or an error/correlation far worse than a reasonable baseline, so the grader's file check on the prediction file fails.
id c54500c813ec · mined from da-code dacode-ml-regression-004@s17
raw text (what the judge reads)
### Predictions not reconciled row-for-row (and distribution-wise) with the provided test set
- **Applies when**: `task` -- the task asks for a per-row prediction file for a supplied test split, and the script does its own preprocessing (dropping NA rows, subsetting features, filtering) before predicting.
- **Pattern**: The attempt builds a model on a silently reduced subset of rows/columns, then writes a prediction file whose row count or row order no longer provably corresponds to the test file, and never sanity-checks the output length or the predicted value distribution against the observed target distribution; feature choice is also narrowed to the "easy" numeric columns while high-signal metadata is discarded, yielding heavily shrunk, low-variance predictions.
- **Detection procedure**:
  1. Read the task for the required output: one prediction per test row, exact column name, original order.
  2. In the script, trace every `dropna`, filter, `merge`, `sample`, or column subset applied to the test frame; check whether the frame used for `predict` is the unmodified, same-length, same-order test frame (NAs imputed rather than dropped) and whether the training feature matrix is built identically.
  3. Check for an explicit assertion/print that `len(predictions) == len(test)` and that the index/ID alignment is preserved; check whether the reported row counts equal the actual file row counts.
  4. Compare the reported prediction range/mean/spread to the target's range and spread in the training data, and ask whether obviously informative columns (identifiers, categorical metadata, dates) were dropped without justification.
- **Discriminator**: A real violation is when rows can be dropped/reordered by preprocessing, no length/order check exists, or predictions collapse into a narrow band far tighter than the target's actual spread; it is fine if the script imputes rather than drops, writes predictions from the intact test frame, and verifies counts — even if the model is simple and predictions are mildly shrunk (regression toward the mean is expected).
- **Consequence**: The saved file has the wrong number of rows or mis-aligned predictions, and/or an error/correlation far worse than a reasonable baseline, so the grader's file check on the prediction file fails.
898Selecting the model configuration by a single automatic score without sanity-checking the resulting group structuretaskda-code
Applies when
task -- the task asks for an unsupervised grouping/segmentation "into an appropriate number of groups" and the script picks that number by taking the argmax/elbow of one internal index over a range.
Pattern
The attempt scans a range of settings, selects the one maximizing a single internal criterion, and writes out the labels without checking whether the chosen solution is degenerate (near-empty or singleton groups driven by extreme records), whether the score value itself indicates weak structure, or whether the saved columns actually contain the values the answer claims (e.g., raw vs. transformed features). Outlier handling is limited to dropping nulls, so a few extreme rows can capture their own groups.
Detection procedure
  1. In the task, note that the requested deliverable is a grouping that must be interpretable/plausible, and note exactly which columns and value types the output file must contain.
  2. In the script, check whether the chosen setting is decided solely by argmax of one index, and whether any post-selection validation exists (group sizes, score magnitude, comparison with a second criterion, outlier/skew treatment such as log or robust scaling).
  3. In the answer, inspect the reported group sizes and score: flag if any group has ~1-3 members out of hundreds, or if the reported score is low (weak separation) while a smaller, more balanced solution was available.
  4. Cross-check that the described content of the saved file (scaled vs. original values, column names, row count) matches what the code actually writes; flag any discrepancy.
Discriminator
A genuine violation is a solution whose group sizes are grossly imbalanced/degenerate or whose selection rests on one weak score with no robustness check or outlier handling, and/or whose reported file contents contradict the code. It is not a violation if the attempt chose a slightly unusual number but justified it with multiple criteria, showed all groups are substantively populated and interpretable, and the written file demonstrably matches the required schema and values.
Consequence
The saved label file disagrees with the expected grouping (wrong number of groups, tiny outlier-only groups, or feature columns in the wrong scale), so the file check fails even though the pipeline "ran successfully."
id 582b516d9fe8 · mined from da-code dacode-ml-cluster-013@s17
raw text (what the judge reads)
### Selecting the model configuration by a single automatic score without sanity-checking the resulting group structure
- **Applies when**: `task` -- the task asks for an unsupervised grouping/segmentation "into an appropriate number of groups" and the script picks that number by taking the argmax/elbow of one internal index over a range.
- **Pattern**: The attempt scans a range of settings, selects the one maximizing a single internal criterion, and writes out the labels without checking whether the chosen solution is degenerate (near-empty or singleton groups driven by extreme records), whether the score value itself indicates weak structure, or whether the saved columns actually contain the values the answer claims (e.g., raw vs. transformed features). Outlier handling is limited to dropping nulls, so a few extreme rows can capture their own groups.
- **Detection procedure**:
  1. In the task, note that the requested deliverable is a grouping that must be interpretable/plausible, and note exactly which columns and value types the output file must contain.
  2. In the script, check whether the chosen setting is decided solely by argmax of one index, and whether any post-selection validation exists (group sizes, score magnitude, comparison with a second criterion, outlier/skew treatment such as log or robust scaling).
  3. In the answer, inspect the reported group sizes and score: flag if any group has ~1-3 members out of hundreds, or if the reported score is low (weak separation) while a smaller, more balanced solution was available.
  4. Cross-check that the described content of the saved file (scaled vs. original values, column names, row count) matches what the code actually writes; flag any discrepancy.
- **Discriminator**: A genuine violation is a solution whose group sizes are grossly imbalanced/degenerate or whose selection rests on one weak score with no robustness check or outlier handling, and/or whose reported file contents contradict the code. It is *not* a violation if the attempt chose a slightly unusual number but justified it with multiple criteria, showed all groups are substantively populated and interpretable, and the written file demonstrably matches the required schema and values.
- **Consequence**: The saved label file disagrees with the expected grouping (wrong number of groups, tiny outlier-only groups, or feature columns in the wrong scale), so the file check fails even though the pipeline "ran successfully."
899Stated estimator definition contradicted by the flag actually used in codetaskinfiagent-dabench
Applies when
task -- the task names a specific variant of a statistic (e.g., bias-corrected/adjusted vs. population, sample vs. population variance, ddof, normalization, unbiased estimator) and the script computes it with a library function that exposes that variant as a parameter.
Pattern
The script hard-codes the default or opposite parameter value (e.g., bias=True, ddof=0) while a comment asserts it is the requested definition, so the stated constraint is silently violated; the same wrong flag is copy-pasted across every exploratory variant, and the ambiguity of which subset the statistic is computed over is resolved by trying several interpretations and reporting whichever ran, without any check that the chosen subset matches the filter stated in the task.
Detection procedure
  1. From the task text, list every explicit statistical specification (estimator variant, correction flag, units, rounding, filter/subset).
  2. In the scripts, locate each call computing that statistic and read the actual keyword arguments and the library's documented meaning of each flag; compare to the list from step 1, ignoring the script's own comments.
  3. Check that the data actually fed to the statistic corresponds to the subset/filter named in the task (and that the group axis produces >1 value per group so the statistic is defined); if the script explores multiple incompatible interpretations, verify the final answer comes from the one that satisfies the task's filter.
  4. Confirm the reported answer is traceable to a single run whose flags and subset both satisfy steps 2 and 3.
Discriminator
A real violation is when the argument value (or the subset used) provably differs from the task's stated specification, or when no run in the scripts satisfies both simultaneously. A look-alike that is fine is code that omits the argument but where the library default already equals the requested variant, or code that computes several variants and explicitly selects the compliant one for the final answer.
Consequence
The reported ranking/extremum comes from a different estimator or a different data slice than requested, so the top-ranked entity differs from ground truth and the grader marks the answer wrong (0/1), even though the pipeline runs without error.
id be2e29b74c8c · mined from infiagent-dabench dabench-252@s17
raw text (what the judge reads)
### Stated estimator definition contradicted by the flag actually used in code
- **Applies when**: `task` -- the task names a specific variant of a statistic (e.g., bias-corrected/adjusted vs. population, sample vs. population variance, ddof, normalization, unbiased estimator) and the script computes it with a library function that exposes that variant as a parameter.
- **Pattern**: The script hard-codes the default or opposite parameter value (e.g., `bias=True`, `ddof=0`) while a comment asserts it *is* the requested definition, so the stated constraint is silently violated; the same wrong flag is copy-pasted across every exploratory variant, and the ambiguity of *which subset* the statistic is computed over is resolved by trying several interpretations and reporting whichever ran, without any check that the chosen subset matches the filter stated in the task.
- **Detection procedure**:
  1. From the task text, list every explicit statistical specification (estimator variant, correction flag, units, rounding, filter/subset).
  2. In the scripts, locate each call computing that statistic and read the actual keyword arguments and the library's documented meaning of each flag; compare to the list from step 1, ignoring the script's own comments.
  3. Check that the data actually fed to the statistic corresponds to the subset/filter named in the task (and that the group axis produces >1 value per group so the statistic is defined); if the script explores multiple incompatible interpretations, verify the final answer comes from the one that satisfies the task's filter.
  4. Confirm the reported answer is traceable to a single run whose flags and subset both satisfy steps 2 and 3.
- **Discriminator**: A real violation is when the argument value (or the subset used) provably differs from the task's stated specification, or when no run in the scripts satisfies both simultaneously. A look-alike that is fine is code that omits the argument but where the library default already equals the requested variant, or code that computes several variants and explicitly selects the compliant one for the final answer.
- **Consequence**: The reported ranking/extremum comes from a different estimator or a different data slice than requested, so the top-ranked entity differs from ground truth and the grader marks the answer wrong (0/1), even though the pipeline runs without error.
900Truncating/reformatting a returned identifier to fit a format template, losing required precisiontaskinfiagent-dabench
Applies when
task -- the task asks you to report a specific data key (date, ID, category label) that you locate in the data, and the answer-format spec shows a template or example with coarser granularity than the underlying values.
Pattern
The script finds the correct record but then applies a lossy transformation to the reported key (string slicing, truncation, rounding, resampling, strftime with fewer fields, lower-casing/renaming) purely to match the literal template, discarding information that uniquely identifies the record; the submitted answer therefore names a period/group rather than the actual row that was found.
Detection procedure
  1. In the task, note what entity is asked for ("the date/row/ID with the maximum ...") and separately note the format hint; check whether the hint is coarser than the granularity of the data being searched.
  2. In the scripts, look for any post-hoc mangling of the located key before printing (e.g. x[:7], truncation, .round(), aggregation to a coarser period, manual re-typing) that is not required by the computation itself.
  3. Compare the value in the final answer with the value the script actually used to index the neighboring/previous record; if the reported key is less specific than the key used internally, flag it.
  4. Prefer answers that report the full-precision key as found (and, if a template is ambiguous, keep the finest granularity available rather than deleting fields).
Discriminator
A real violation is when the reported key no longer pinpoints the record used in the calculation (many rows share the reported value). It is fine when the task genuinely requires aggregation to that granularity (e.g. "the month with the highest average"), where the coarse key is the true answer computed on grouped data, or when the data itself has only that granularity.
Consequence
The dependent numeric result may still match, but the identifier check fails on exact-string comparison, so the attempt is scored partially/fully wrong despite correct analysis.
id a029ee31282f · mined from infiagent-dabench dabench-572@s17
raw text (what the judge reads)
### Truncating/reformatting a returned identifier to fit a format template, losing required precision
- **Applies when**: `task` -- the task asks you to report a specific data key (date, ID, category label) that you locate in the data, and the answer-format spec shows a template or example with coarser granularity than the underlying values.
- **Pattern**: The script finds the correct record but then applies a lossy transformation to the reported key (string slicing, truncation, rounding, resampling, `strftime` with fewer fields, lower-casing/renaming) purely to match the literal template, discarding information that uniquely identifies the record; the submitted answer therefore names a period/group rather than the actual row that was found.
- **Detection procedure**:
  1. In the task, note what entity is asked for ("the date/row/ID with the maximum ...") and separately note the format hint; check whether the hint is coarser than the granularity of the data being searched.
  2. In the scripts, look for any post-hoc mangling of the located key before printing (e.g. `x[:7]`, truncation, `.round()`, aggregation to a coarser period, manual re-typing) that is not required by the computation itself.
  3. Compare the value in the final answer with the value the script actually used to index the neighboring/previous record; if the reported key is less specific than the key used internally, flag it.
  4. Prefer answers that report the full-precision key as found (and, if a template is ambiguous, keep the finest granularity available rather than deleting fields).
- **Discriminator**: A real violation is when the reported key no longer pinpoints the record used in the calculation (many rows share the reported value). It is fine when the task genuinely requires aggregation to that granularity (e.g. "the month with the highest average"), where the coarse key is the true answer computed on grouped data, or when the data itself has only that granularity.
- **Consequence**: The dependent numeric result may still match, but the identifier check fails on exact-string comparison, so the attempt is scored partially/fully wrong despite correct analysis.
901Answer list re-ordered away from the source/canonical ordertaskinfiagent-dabench
Applies when
task -- the deliverable is a set/list of entity names (or labels) extracted from a table, and the script reorders them before writing the answer.
Pattern
The analysis itself is correct, but the reporting step applies an incidental sort (e.g., by the computed metric, descending) or otherwise permutes the elements, so the submitted list does not follow the order implied by the task or the natural order of the source data (row order in the file, or alphabetical). Exact-match graders then reject a substantively right answer.
Detection procedure
  1. Read the task/answer-format spec and note whether any ordering is stated; if none is stated, note the default order the elements appear in the source table (and whether that coincides with alphabetical order).
  2. In the script, locate the code that builds the final answer object and look for sort_values, sort, sorted, set, groupby, reversed, or dictionary iteration that changes element order relative to the filtered source rows.
  3. Compare the order produced by that code with the order from step 1; flag if the final written list is ordered by a quantity that the task never asked to order by.
  4. Also confirm the string values are copied verbatim from the source (no case/whitespace/renaming changes) as part of the same formatting check.
Discriminator
A real violation is a permutation (or renaming) introduced purely in the reporting step when the task gave no ordering instruction and the source order differs; it is not a violation if the task explicitly requests that ordering, or if the sort happens only for a diagnostic printout while the answer is written from the unsorted, source-ordered selection (or if the sorted order happens to coincide with the source/alphabetical order).
Consequence
The grader reports the expected items as WRONG/MISSING even though the same names were submitted, yielding 0/1 on exact-match comparison.
id 5d54ed4972c1 · mined from infiagent-dabench dabench-254@s17
raw text (what the judge reads)
### Answer list re-ordered away from the source/canonical order
- **Applies when**: `task` -- the deliverable is a set/list of entity names (or labels) extracted from a table, and the script reorders them before writing the answer.
- **Pattern**: The analysis itself is correct, but the reporting step applies an incidental sort (e.g., by the computed metric, descending) or otherwise permutes the elements, so the submitted list does not follow the order implied by the task or the natural order of the source data (row order in the file, or alphabetical). Exact-match graders then reject a substantively right answer.
- **Detection procedure**:
  1. Read the task/answer-format spec and note whether any ordering is stated; if none is stated, note the default order the elements appear in the source table (and whether that coincides with alphabetical order).
  2. In the script, locate the code that builds the final answer object and look for `sort_values`, `sort`, `sorted`, `set`, `groupby`, `reversed`, or dictionary iteration that changes element order relative to the filtered source rows.
  3. Compare the order produced by that code with the order from step 1; flag if the final written list is ordered by a quantity that the task never asked to order by.
  4. Also confirm the string values are copied verbatim from the source (no case/whitespace/renaming changes) as part of the same formatting check.
- **Discriminator**: A real violation is a permutation (or renaming) introduced purely in the reporting step when the task gave no ordering instruction and the source order differs; it is *not* a violation if the task explicitly requests that ordering, or if the sort happens only for a diagnostic printout while the answer is written from the unsorted, source-ordered selection (or if the sorted order happens to coincide with the source/alphabetical order).
- **Consequence**: The grader reports the expected items as WRONG/MISSING even though the same names were submitted, yielding 0/1 on exact-match comparison.
902Output file schema deviates from the exact format requestedtaskda-code
Applies when
task -- the task specifies the deliverable file and the exact column(s) it must contain, and the script builds and writes that file at the end.
Pattern
The script writes the prediction file with extra columns (e.g., an identifier column), a written index, renamed/re-cased headers, or a different row order/count than the input rows it was supposed to score — instead of exactly the requested single column in input order. Often this is coupled with sloppy index handling on load (e.g., index_col=0 plus a separate ID column), which also risks silent row misalignment.
Detection procedure
  1. From the task statement, write down the exact required deliverable: file name, required column name(s), whether an index/ID is allowed, and the expected number of rows.
  2. In the script, find the DataFrame construction and the to_csv call; list the columns actually written, the index= setting, and check the header spelling/case against the requirement.
  3. Trace how the test set was read and how predictions were attached: confirm the prediction array is in the same order and length as the test rows, with no filtering, dropna, sorting, or index resetting in between.
  4. Compare with the reported answer: if the answer itself states the file "contains columns: id, satisfaction" while the task asked only for the named prediction column, flag it.
Discriminator
A real violation is a written file whose columns/rows differ from the literal specification (extra ID column, index written, misspelled header, wrong row count/order). A look-alike that is fine is a script that keeps an ID internally for alignment but drops it before writing, or writes exactly the requested column with index=False even if intermediate frames had more columns.
Consequence
The grader loads the file and compares against the reference by column name/position and row order; extra or shifted columns or misaligned rows make the comparison fail outright, scoring 0 even when the underlying model predictions are accurate.
id 27b692b4b9d0 · mined from da-code dacode-ml-binary-009@s17
raw text (what the judge reads)
### Output file schema deviates from the exact format requested
- **Applies when**: `task` -- the task specifies the deliverable file and the exact column(s) it must contain, and the script builds and writes that file at the end.
- **Pattern**: The script writes the prediction file with extra columns (e.g., an identifier column), a written index, renamed/re-cased headers, or a different row order/count than the input rows it was supposed to score — instead of exactly the requested single column in input order. Often this is coupled with sloppy index handling on load (e.g., `index_col=0` plus a separate ID column), which also risks silent row misalignment.
- **Detection procedure**:
  1. From the task statement, write down the exact required deliverable: file name, required column name(s), whether an index/ID is allowed, and the expected number of rows.
  2. In the script, find the DataFrame construction and the `to_csv` call; list the columns actually written, the `index=` setting, and check the header spelling/case against the requirement.
  3. Trace how the test set was read and how predictions were attached: confirm the prediction array is in the same order and length as the test rows, with no filtering, dropna, sorting, or index resetting in between.
  4. Compare with the reported answer: if the answer itself states the file "contains columns: id, satisfaction" while the task asked only for the named prediction column, flag it.
- **Discriminator**: A real violation is a written file whose columns/rows differ from the literal specification (extra ID column, index written, misspelled header, wrong row count/order). A look-alike that is fine is a script that keeps an ID internally for alignment but drops it before writing, or writes exactly the requested column with `index=False` even if intermediate frames had more columns.
- **Consequence**: The grader loads the file and compares against the reference by column name/position and row order; extra or shifted columns or misaligned rows make the comparison fail outright, scoring 0 even when the underlying model predictions are accurate.
903Answer-format literal mismatch (delimiters/quoting/typing of the reported value)taskinfiagent-dabench
Applies when
task -- the task specifies an exact answer template with placeholders whose types and delimiters are shown (e.g., quoted strings vs. bare numbers, units, list brackets, ordering of tags).
Pattern
The agent computes a value that is substantively right but emits it in a form that deviates from the template — dropping the required quotes around a string, adding stray quotes/brackets around a number, renaming a tag, adding units, or changing case/whitespace — so an exact-match grader rejects that field even though the analysis was fine.
Detection procedure
  1. Copy the answer template from the task and mark, for each placeholder, its exact required wrapper (quotes? brackets? decimals? tag name spelling?).
  2. Read the script's printing/formatting code (or, if no script was saved, the final answer text) and check whether the emitted token is character-for-character consistent with the template for each field, not just semantically equal.
  3. Also verify the value being placed in each slot is the requested entity (e.g., the identifier/label vs. its count or index) and that a tie or degenerate case (all values equal, empty selection) is resolved by an explicitly stated, deterministic rule.
  4. Flag if any field differs from the template in quoting, delimiters, tag name, numeric formatting, or ordering.
Discriminator
A real violation is a deviation in the rendered token (missing/extra quotes, altered tag, unit suffix, changed precision) even when the underlying value is correct; a look-alike that is fine is a value formatted exactly as the template but computed via a different (still valid) code path, or harmless surrounding prose outside the tags.
Consequence
The grader scores that field as WRONG/MISSING while other fields pass, yielding a partial-credit failure (e.g., 1/2 checks) despite a correct underlying computation.
id 73a508de28b1 · mined from infiagent-dabench dabench-760@s17
raw text (what the judge reads)
### Answer-format literal mismatch (delimiters/quoting/typing of the reported value)
- **Applies when**: `task` -- the task specifies an exact answer template with placeholders whose types and delimiters are shown (e.g., quoted strings vs. bare numbers, units, list brackets, ordering of tags).
- **Pattern**: The agent computes a value that is substantively right but emits it in a form that deviates from the template — dropping the required quotes around a string, adding stray quotes/brackets around a number, renaming a tag, adding units, or changing case/whitespace — so an exact-match grader rejects that field even though the analysis was fine.
- **Detection procedure**:
  1. Copy the answer template from the task and mark, for each placeholder, its exact required wrapper (quotes? brackets? decimals? tag name spelling?).
  2. Read the script's printing/formatting code (or, if no script was saved, the final answer text) and check whether the emitted token is character-for-character consistent with the template for each field, not just semantically equal.
  3. Also verify the value being placed in each slot is the requested entity (e.g., the identifier/label vs. its count or index) and that a tie or degenerate case (all values equal, empty selection) is resolved by an explicitly stated, deterministic rule.
  4. Flag if any field differs from the template in quoting, delimiters, tag name, numeric formatting, or ordering.
- **Discriminator**: A real violation is a deviation in the *rendered* token (missing/extra quotes, altered tag, unit suffix, changed precision) even when the underlying value is correct; a look-alike that is fine is a value formatted exactly as the template but computed via a different (still valid) code path, or harmless surrounding prose outside the tags.
- **Consequence**: The grader scores that field as WRONG/MISSING while other fields pass, yielding a partial-credit failure (e.g., 1/2 checks) despite a correct underlying computation.
904Invented metric definition that contradicts the spec's own consistency cluestaskda-code
Applies when
task -- The task says to plot/report "performance" (or any aggregate) using an external config/spec that fixes labels, order, axis names, and the script must choose how to compute the plotted quantity.
Pattern
The script defines the aggregate with an ad-hoc, unstated formula (e.g., counts plus an arbitrarily weighted flag column) instead of deriving it from the spec's semantics, and never checks that the resulting values are consistent with the fixed label order or with the axis/title wording; it also emits only the image and skips other required output artifacts.
Detection procedure
  1. Read the task and the config: list every constraint it pins down (label set and their order, axis/title text, period/filter, file outputs) and what those imply about the quantity being plotted.
  2. Read the script: find the exact expression producing the plotted values; ask whether every term is justified by the task/config wording or is an invented weight/proxy, and whether flag/boolean columns are compared with the correct dtype.
  3. Compare the script's computed values against the spec-implied structure — e.g., if the config's category order implies a sorted ranking, verify the computed values follow that order; also check the aggregation covers all relevant rows (both roles/sides, not just one grouping key).
  4. Check the answer/output list against every artifact the task requires; missing or extra files count as failure.
Discriminator
A real violation is when the plotted numbers cannot be reproduced from any natural reading of the task and visibly break a spec-implied invariant (ordering, monotonicity, plausible range) or a required output is absent. It is fine if the metric is a reasonable literal reading of the task and the values respect the config's ordering/labels, even if a slightly different but equivalent formulation exists.
Consequence
The saved chart and any numeric artifact contain values from a different statistic than expected, so exact-value/array comparisons fail and required files are reported missing — zero checks passed despite the script running cleanly.
id e16bd336a9f2 · mined from da-code dacode-plot-bar-006@s17
raw text (what the judge reads)
### Invented metric definition that contradicts the spec's own consistency clues
- **Applies when**: `task` -- The task says to plot/report "performance" (or any aggregate) using an external config/spec that fixes labels, order, axis names, and the script must choose how to compute the plotted quantity.
- **Pattern**: The script defines the aggregate with an ad-hoc, unstated formula (e.g., counts plus an arbitrarily weighted flag column) instead of deriving it from the spec's semantics, and never checks that the resulting values are consistent with the fixed label order or with the axis/title wording; it also emits only the image and skips other required output artifacts.
- **Detection procedure**:
  1. Read the task and the config: list every constraint it pins down (label set and their order, axis/title text, period/filter, file outputs) and what those imply about the quantity being plotted.
  2. Read the script: find the exact expression producing the plotted values; ask whether every term is justified by the task/config wording or is an invented weight/proxy, and whether flag/boolean columns are compared with the correct dtype.
  3. Compare the script's computed values against the spec-implied structure — e.g., if the config's category order implies a sorted ranking, verify the computed values follow that order; also check the aggregation covers all relevant rows (both roles/sides, not just one grouping key).
  4. Check the answer/output list against every artifact the task requires; missing or extra files count as failure.
- **Discriminator**: A real violation is when the plotted numbers cannot be reproduced from any natural reading of the task and visibly break a spec-implied invariant (ordering, monotonicity, plausible range) or a required output is absent. It is fine if the metric is a reasonable literal reading of the task and the values respect the config's ordering/labels, even if a slightly different but equivalent formulation exists.
- **Consequence**: The saved chart and any numeric artifact contain values from a different statistic than expected, so exact-value/array comparisons fail and required files are reported missing — zero checks passed despite the script running cleanly.
905Accepting an outlier/threshold count without validating input scope and statistic definitiontaskinfiagent-dabench
Applies when
task -- the task asks for a count (or removal) of points passing a fixed statistical threshold, and the scripts compute it in one pass over one file/column with one library convention.
Pattern
The attempt hard-codes a single input file, a single column name, and a single formula variant (e.g., scipy.stats.zscore with population std / default NaN handling), then reports whatever count comes out — never checking that the flagged rows are genuinely extreme, that the chosen file/column is the one the task means, or that a common alternative convention (sample vs population std, NaN/non-numeric coercion, deduplicated or full dataset) gives the same count.
Detection procedure
  1. From the task, list the ambiguities that materially change the count: which file(s)/split constitute "the dataset", exactly which column matches the requested variable, and how missing/non-numeric entries and the std convention are handled.
  2. In the scripts, check whether each ambiguity is resolved by evidence (directory listing, dtype/column printout, comparison of variants) or silently by a default; look for any recomputation with an alternative convention or input.
  3. Inspect whether the script prints the flagged values themselves and compares them to the distribution (min/max, quantiles, units) to confirm they are implausibly extreme rather than legitimate tail values.
  4. Compare the reported count to the expected order of magnitude for the stated threshold (for a roughly bell-shaped variable, |z|>3 should flag well under ~0.3% of rows); a much larger share unexamined is a red flag.
Discriminator
A fine attempt either shows the variants agree, or explicitly justifies the chosen file/column/convention from inspected data and displays the flagged extreme values; a violation reports a single default-path number with no cross-check, no listing of flagged values, and no reaction to a count that is large relative to the threshold's implied tail probability.
Consequence
The reported count comes from the wrong subset or the wrong statistic variant and mismatches the expected value exactly (e.g., a nonzero count reported where the correct answer is zero), failing the single graded check.
id 5cc7a9201f14 · mined from infiagent-dabench dabench-361@s17
raw text (what the judge reads)
### Accepting an outlier/threshold count without validating input scope and statistic definition
- **Applies when**: `task` -- the task asks for a count (or removal) of points passing a fixed statistical threshold, and the scripts compute it in one pass over one file/column with one library convention.
- **Pattern**: The attempt hard-codes a single input file, a single column name, and a single formula variant (e.g., `scipy.stats.zscore` with population std / default NaN handling), then reports whatever count comes out — never checking that the flagged rows are genuinely extreme, that the chosen file/column is the one the task means, or that a common alternative convention (sample vs population std, NaN/non-numeric coercion, deduplicated or full dataset) gives the same count.
- **Detection procedure**:
  1. From the task, list the ambiguities that materially change the count: which file(s)/split constitute "the dataset", exactly which column matches the requested variable, and how missing/non-numeric entries and the std convention are handled.
  2. In the scripts, check whether each ambiguity is resolved by evidence (directory listing, dtype/column printout, comparison of variants) or silently by a default; look for any recomputation with an alternative convention or input.
  3. Inspect whether the script prints the flagged values themselves and compares them to the distribution (min/max, quantiles, units) to confirm they are implausibly extreme rather than legitimate tail values.
  4. Compare the reported count to the expected order of magnitude for the stated threshold (for a roughly bell-shaped variable, |z|>3 should flag well under ~0.3% of rows); a much larger share unexamined is a red flag.
- **Discriminator**: A fine attempt either shows the variants agree, or explicitly justifies the chosen file/column/convention from inspected data and displays the flagged extreme values; a violation reports a single default-path number with no cross-check, no listing of flagged values, and no reaction to a count that is large relative to the threshold's implied tail probability.
- **Consequence**: The reported count comes from the wrong subset or the wrong statistic variant and mismatches the expected value exactly (e.g., a nonzero count reported where the correct answer is zero), failing the single graded check.
906Ignoring a task-referenced specification file and substituting assumed valuestaskda-code
Applies when
task -- the prompt tells the agent to use a mapping/rule/threshold/config defined in an accompanying documentation or spec file (README, tips, notes, schema) and the scripts must apply it before computing the requested statistic.
Pattern
The script never opens or quotes the referenced file; instead it hard-codes a "standard" or "typical" mapping/rule invented from domain intuition (often flagged by a comment like "common convention"), so the transformed categories, labels, or derived values differ from the ones the grader expects — and any required output artifact is written from those assumed values (or not written at all).
Detection procedure
  1. From the task text, list every external file or documented rule the task says to use, plus every required output artifact and its exact key/label format.
  2. Search the scripts for a read/parse of each referenced file (open/read_csv/read_json/markdown parse) and for the literal values it defines; note any hard-coded dict/list/constant that plays that role instead.
  3. Check that the final reported labels are exactly the strings the spec file prescribes, and that the required result file is written with the requested keys and value types.
  4. If the spec is never read (or its contents never echoed/validated) while a substituted constant is used, or the required artifact is absent, flag the attempt.
Discriminator
A real violation is inventing the rule without any evidence of reading the spec; it is fine if the script loads the spec (or explicitly prints/reproduces its exact contents for verification) and the hard-coded values are shown to match it — cosmetic differences in variable naming or an additional convenience copy of a verified mapping are not violations.
Consequence
Numeric ratio may look plausible, but the label string (and thus the required result file / key values) mismatches the expected mapping, so the grader marks the answer wrong or the expected output file missing.
id 13e29224ae83 · mined from da-code dacode-di-text-004@s17
raw text (what the judge reads)
### Ignoring a task-referenced specification file and substituting assumed values
- **Applies when**: `task` -- the prompt tells the agent to use a mapping/rule/threshold/config defined in an accompanying documentation or spec file (README, tips, notes, schema) and the scripts must apply it before computing the requested statistic.
- **Pattern**: The script never opens or quotes the referenced file; instead it hard-codes a "standard" or "typical" mapping/rule invented from domain intuition (often flagged by a comment like "common convention"), so the transformed categories, labels, or derived values differ from the ones the grader expects — and any required output artifact is written from those assumed values (or not written at all).
- **Detection procedure**:
  1. From the task text, list every external file or documented rule the task says to use, plus every required output artifact and its exact key/label format.
  2. Search the scripts for a read/parse of each referenced file (open/read_csv/read_json/markdown parse) and for the literal values it defines; note any hard-coded dict/list/constant that plays that role instead.
  3. Check that the final reported labels are exactly the strings the spec file prescribes, and that the required result file is written with the requested keys and value types.
  4. If the spec is never read (or its contents never echoed/validated) while a substituted constant is used, or the required artifact is absent, flag the attempt.
- **Discriminator**: A real violation is inventing the rule without any evidence of reading the spec; it is fine if the script loads the spec (or explicitly prints/reproduces its exact contents for verification) and the hard-coded values are shown to match it — cosmetic differences in variable naming or an additional convenience copy of a verified mapping are not violations.
- **Consequence**: Numeric ratio may look plausible, but the label string (and thus the required result file / key values) mismatches the expected mapping, so the grader marks the answer wrong or the expected output file missing.
907Submission artifact not verified against the required sample file (schema, header, row count, persistence)taskda-code
Applies when
task -- the task requires writing predictions/results to a specific output file whose format is defined by a provided template or spec.
Pattern
The agent produces predictions but never programmatically writes and re-reads the required file, instead echoing rows into the chat/answer; the emitted content is truncated, uses header names/casing or column order copied from memory rather than from the template, and covers fewer rows than the evaluation set.
Detection procedure
  1. From the task, note the exact required output path, and load the provided template/spec to get its header string, column order and number of rows/IDs.
  2. In the scripts, look for an explicit save step to that exact path (e.g., a write/to_csv with index=False) followed by a read-back check; absence of any file-writing code, or only printing results, is an immediate fail.
  3. Compare the answer's header and ID column against the template's header and ID list character-for-character (case, naming, order); compare row count to the template's row count and check for truncation or missing IDs.
  4. Sanity-check value constraints implied by the metric (e.g., probabilities in [0,1] and per-row sums where required, no NaNs/duplicated IDs).
Discriminator
A real violation is when the required file is not written, is truncated/short, or its header/ID set/order differs from the template; a look-alike that is fine is a correctly written full-length file whose values simply differ from an unknown ideal, or an abbreviated preview in the chat that is explicitly accompanied by code writing the complete file.
Consequence
The grader reports the expected output file as WRONG/MISSING — it cannot be parsed or joined to the ground-truth IDs — so the score is zero regardless of model quality.
id 7a73afd46c15 · mined from da-code dacode-ml-competition-003@s17
raw text (what the judge reads)
### Submission artifact not verified against the required sample file (schema, header, row count, persistence)
- **Applies when**: `task` -- the task requires writing predictions/results to a specific output file whose format is defined by a provided template or spec.
- **Pattern**: The agent produces predictions but never programmatically writes and re-reads the required file, instead echoing rows into the chat/answer; the emitted content is truncated, uses header names/casing or column order copied from memory rather than from the template, and covers fewer rows than the evaluation set.
- **Detection procedure**:
  1. From the task, note the exact required output path, and load the provided template/spec to get its header string, column order and number of rows/IDs.
  2. In the scripts, look for an explicit save step to that exact path (e.g., a write/`to_csv` with `index=False`) followed by a read-back check; absence of any file-writing code, or only printing results, is an immediate fail.
  3. Compare the answer's header and ID column against the template's header and ID list character-for-character (case, naming, order); compare row count to the template's row count and check for truncation or missing IDs.
  4. Sanity-check value constraints implied by the metric (e.g., probabilities in [0,1] and per-row sums where required, no NaNs/duplicated IDs).
- **Discriminator**: A real violation is when the required file is not written, is truncated/short, or its header/ID set/order differs from the template; a look-alike that is fine is a correctly written full-length file whose values simply differ from an unknown ideal, or an abbreviated preview in the chat that is explicitly accompanied by code writing the complete file.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING — it cannot be parsed or joined to the ground-truth IDs — so the score is zero regardless of model quality.
908Claiming labels are unavailable and substituting an unvalidated heuristictaskda-code
Applies when
task -- The task asks for supervised predictions of a named target column on a held-out file, and the scripts never fit a model on labeled rows (asserting the labels "aren't provided") but instead assign classes by hand-tuned rules.
Pattern
The agent does a shallow search, concludes no ground-truth labels exist, invents a rule-based mapping from loosely related features, and invents its own class vocabulary/label spelling and class balance — never checking the label values or their empirical frequencies in any labeled source.
Detection procedure
  1. From the task, identify the exact target column name and which provided file(s) could contain it (main table, train split, auxiliary tables joined by key).
  2. In the scripts, look for a step that loads that column and inspects its unique values/frequencies; if the only evidence for "no labels" is one file's header, treat the search as incomplete.
  3. Check whether predicted category strings and their distribution match the label vocabulary/prior observed in the data (casing, separators, number of classes, plausible class shares).
  4. Check the answer for a supervised fit/validation step with a reported held-out score; a heuristic with no accuracy estimate is unverified.
Discriminator
Genuine violation = target column demonstrably exists somewhere in the provided data (or its category set is discoverable) yet was never read, and/or output labels use invented names or an implausible distribution (e.g., a rare senior class predicted for half the rows). Acceptable = the agent enumerated all files/columns, showed the target is truly absent, then used an external-knowledge heuristic while still matching the required label vocabulary and reporting a sanity-checked distribution.
Consequence
Predictions are compared against real labels with mismatched category strings and a badly skewed prior, so accuracy collapses toward or below the majority-class baseline and the file is graded wrong despite passing format/count self-checks.
id 8773d083eeaf · mined from da-code dacode-ml-multi-003@s17
raw text (what the judge reads)
### Claiming labels are unavailable and substituting an unvalidated heuristic
- **Applies when**: `task` -- The task asks for supervised predictions of a named target column on a held-out file, and the scripts never fit a model on labeled rows (asserting the labels "aren't provided") but instead assign classes by hand-tuned rules.
- **Pattern**: The agent does a shallow search, concludes no ground-truth labels exist, invents a rule-based mapping from loosely related features, and invents its own class vocabulary/label spelling and class balance — never checking the label values or their empirical frequencies in any labeled source.
- **Detection procedure**:
  1. From the task, identify the exact target column name and which provided file(s) could contain it (main table, train split, auxiliary tables joined by key).
  2. In the scripts, look for a step that loads that column and inspects its unique values/frequencies; if the only evidence for "no labels" is one file's header, treat the search as incomplete.
  3. Check whether predicted category strings and their distribution match the label vocabulary/prior observed in the data (casing, separators, number of classes, plausible class shares).
  4. Check the answer for a supervised fit/validation step with a reported held-out score; a heuristic with no accuracy estimate is unverified.
- **Discriminator**: Genuine violation = target column demonstrably exists somewhere in the provided data (or its category set is discoverable) yet was never read, and/or output labels use invented names or an implausible distribution (e.g., a rare senior class predicted for half the rows). Acceptable = the agent enumerated all files/columns, showed the target is truly absent, then used an external-knowledge heuristic while still matching the required label vocabulary and reporting a sanity-checked distribution.
- **Consequence**: Predictions are compared against real labels with mismatched category strings and a badly skewed prior, so accuracy collapses toward or below the majority-class baseline and the file is graded wrong despite passing format/count self-checks.
909Template file provided but never read; output format inferred/guessedtaskda-code
Applies when
task -- The task says the deliverable must match a provided template/example file (or a stated schema), and the scripts write the output file directly from a computed table.
Pattern
The scripts never open, parse, or compare against the template; the agent invents index/column labels, ordering, offsets (e.g., 0- vs 1-based indices), date/string formatting, rounding, and NaN handling, then "fixes" the format post-hoc with ad-hoc string edits based on assumption rather than on the template's actual contents.
Detection procedure
  1. In the task statement, note that a template/reference file exists and identify its path.
  2. Search all scripts for any read of that template (e.g., loading it, printing its header/index/shape/dtypes) and for any explicit comparison of the produced file's shape, column names, row labels, and value precision against it.
  3. If absent, check whether the format choices in the code (label formatting, index base, rounding digits, column naming, empty-cell representation) are justified by anything other than the agent's guess — including a post-hoc reformat script that changes labels without evidence.
  4. Read the final answer for a claim of format compliance ("format matches") unsupported by a template-based diff.
Discriminator
A real violation is when no read/diff against the template occurs anywhere and format decisions are asserted rather than verified. It is fine if the script loads the template (or the task fully specifies the schema inline) and programmatically asserts matching headers/index/shape/precision, even if it then reformats the output to comply.
Consequence
The saved file differs from the expected artifact in labels, column set/offset, ordering, or rounding, so an automated file comparison marks the result WRONG/MISSING even when the underlying aggregation logic is plausible.
id 1409d55e6107 · mined from da-code dacode-dm-csv-044@s17
raw text (what the judge reads)
### Template file provided but never read; output format inferred/guessed
- **Applies when**: `task` -- The task says the deliverable must match a provided template/example file (or a stated schema), and the scripts write the output file directly from a computed table.
- **Pattern**: The scripts never open, parse, or compare against the template; the agent invents index/column labels, ordering, offsets (e.g., 0- vs 1-based indices), date/string formatting, rounding, and NaN handling, then "fixes" the format post-hoc with ad-hoc string edits based on assumption rather than on the template's actual contents.
- **Detection procedure**:
  1. In the task statement, note that a template/reference file exists and identify its path.
  2. Search all scripts for any read of that template (e.g., loading it, printing its header/index/shape/dtypes) and for any explicit comparison of the produced file's shape, column names, row labels, and value precision against it.
  3. If absent, check whether the format choices in the code (label formatting, index base, rounding digits, column naming, empty-cell representation) are justified by anything other than the agent's guess — including a post-hoc reformat script that changes labels without evidence.
  4. Read the final answer for a claim of format compliance ("format matches") unsupported by a template-based diff.
- **Discriminator**: A real violation is when no read/diff against the template occurs anywhere and format decisions are asserted rather than verified. It is fine if the script loads the template (or the task fully specifies the schema inline) and programmatically asserts matching headers/index/shape/precision, even if it then reformats the output to comply.
- **Consequence**: The saved file differs from the expected artifact in labels, column set/offset, ordering, or rounding, so an automated file comparison marks the result WRONG/MISSING even when the underlying aggregation logic is plausible.
910Optimizing/validating with a proxy metric instead of the competition's stated metrictaskda-code
Applies when
task -- the task specifies an explicit evaluation metric (e.g., a rank/agreement/weighted score) and the scripts perform cross-validation, model selection, or ensembling to choose the submitted predictions.
Pattern
The scripts score candidate models with a convenient default metric (log loss, accuracy, RMSE) and combine them by a generic rule (majority vote, unweighted average), never computing the stated metric on any held-out split; nothing in the pipeline verifies that the chosen predictions actually score well under the metric that will be graded, nor that the metric's structure (e.g., ordinal distances, class imbalance, thresholding/rounding of continuous scores) is respected.
Detection procedure
  1. Read the task/README and note the exact scoring metric and any implied structure of the target (ordered categories, cost weighting, rounding rules).
  2. Search the scripts for that metric being implemented or imported and applied to out-of-fold predictions; check what scoring= / selection criterion is actually used.
  3. Check whether the final prediction step (ensemble/threshold/rounding) is tuned or at least validated against the stated metric, and whether any reported number corresponds to it.
  4. Inspect the answer's predicted-value distribution versus the training target distribution and confirm the submission has the required rows/ids/format; unvalidated proxy-metric pipelines typically collapse to the majority classes.
Discriminator
A real violation is when the stated metric is never computed anywhere, so the agent has no evidence its choice is good; it is acceptable if a proxy is used for speed but the final candidates are compared with the stated metric on held-out data (or the proxy is provably monotone-equivalent to it).
Consequence
The graded score under the true metric is far below what a metric-aware baseline achieves (predictions squeezed into a few frequent classes, low agreement/kappa), so the submission fails the grader's correctness/score threshold.
id a44999eaf2a6 · mined from da-code dacode-ml-competition-006@s17
raw text (what the judge reads)
### Optimizing/validating with a proxy metric instead of the competition's stated metric
- **Applies when**: `task` -- the task specifies an explicit evaluation metric (e.g., a rank/agreement/weighted score) and the scripts perform cross-validation, model selection, or ensembling to choose the submitted predictions.
- **Pattern**: The scripts score candidate models with a convenient default metric (log loss, accuracy, RMSE) and combine them by a generic rule (majority vote, unweighted average), never computing the stated metric on any held-out split; nothing in the pipeline verifies that the chosen predictions actually score well under the metric that will be graded, nor that the metric's structure (e.g., ordinal distances, class imbalance, thresholding/rounding of continuous scores) is respected.
- **Detection procedure**:
  1. Read the task/README and note the exact scoring metric and any implied structure of the target (ordered categories, cost weighting, rounding rules).
  2. Search the scripts for that metric being implemented or imported and applied to out-of-fold predictions; check what `scoring=` / selection criterion is actually used.
  3. Check whether the final prediction step (ensemble/threshold/rounding) is tuned or at least validated against the stated metric, and whether any reported number corresponds to it.
  4. Inspect the answer's predicted-value distribution versus the training target distribution and confirm the submission has the required rows/ids/format; unvalidated proxy-metric pipelines typically collapse to the majority classes.
- **Discriminator**: A real violation is when the stated metric is never computed anywhere, so the agent has no evidence its choice is good; it is acceptable if a proxy is used for speed but the final candidates are compared with the stated metric on held-out data (or the proxy is provably monotone-equivalent to it).
- **Consequence**: The graded score under the true metric is far below what a metric-aware baseline achieves (predictions squeezed into a few frequent classes, low agreement/kappa), so the submission fails the grader's correctness/score threshold.
911Missing-value grouping based on assumed encoding instead of inspected sentinel valuestaskinfiagent-dabench
Applies when
task -- the analysis splits, filters, or aggregates rows according to whether a field is "null"/"missing", and the script decides membership with a default helper (e.g. isna()/notna(), dropna()) right after a plain load of the raw file.
Pattern
The attempt trusts the loader's default missing-value detection and never checks how absence is actually encoded in that column (empty strings, whitespace, "NA", "None", ".", 0, placeholder text, or a column shifted by an index/header option). Group sizes and group statistics are therefore computed on slightly wrong subsets, and no verification of the split is attempted.
Detection procedure
  1. In the task, identify the exact predicate that defines the groups/filter ("null vs not null", "missing", "unassigned").
  2. In the scripts, check whether anything inspects the raw column before splitting — e.g. printing value_counts(dropna=False), unique values, string lengths, dtype, or the parsed columns/header alignment — versus jumping straight to isna().
  3. Check whether the split is validated: do group counts sum to the row total, are they plausible, and is any alternative missing-encoding tested to see if group sizes/means change?
  4. Compare reported group statistics against any independent cross-check in the scripts; absence of one plus absence of step 2 is the flag.
Discriminator
Fine if the script demonstrates the column is truly numeric/typed and that only genuine NaNs exist (explicit unique-value or counts printout, or an na_values/dtype specification matching the file), or if it re-derives the same groups a second way. A violation is when missingness semantics are assumed and the only evidence shown is the output of the very function whose correctness is in question.
Consequence
A handful of placeholder-coded rows land in the wrong group, so both group means (and the test statistic) shift off the expected values; the grader marks the numeric answers wrong even though the test procedure and output format are correct.
id 038b7f3f1594 · mined from infiagent-dabench dabench-297@s17
raw text (what the judge reads)
### Missing-value grouping based on assumed encoding instead of inspected sentinel values
- **Applies when**: `task` -- the analysis splits, filters, or aggregates rows according to whether a field is "null"/"missing", and the script decides membership with a default helper (e.g. `isna()`/`notna()`, `dropna()`) right after a plain load of the raw file.
- **Pattern**: The attempt trusts the loader's default missing-value detection and never checks how absence is actually encoded in that column (empty strings, whitespace, `"NA"`, `"None"`, `"."`, `0`, placeholder text, or a column shifted by an index/header option). Group sizes and group statistics are therefore computed on slightly wrong subsets, and no verification of the split is attempted.
- **Detection procedure**:
  1. In the task, identify the exact predicate that defines the groups/filter ("null vs not null", "missing", "unassigned").
  2. In the scripts, check whether anything inspects the raw column before splitting — e.g. printing `value_counts(dropna=False)`, unique values, string lengths, dtype, or the parsed columns/header alignment — versus jumping straight to `isna()`.
  3. Check whether the split is validated: do group counts sum to the row total, are they plausible, and is any alternative missing-encoding tested to see if group sizes/means change?
  4. Compare reported group statistics against any independent cross-check in the scripts; absence of one plus absence of step 2 is the flag.
- **Discriminator**: Fine if the script demonstrates the column is truly numeric/typed and that only genuine NaNs exist (explicit unique-value or counts printout, or an `na_values`/dtype specification matching the file), or if it re-derives the same groups a second way. A violation is when missingness semantics are assumed and the only evidence shown is the output of the very function whose correctness is in question.
- **Consequence**: A handful of placeholder-coded rows land in the wrong group, so both group means (and the test statistic) shift off the expected values; the grader marks the numeric answers wrong even though the test procedure and output format are correct.
912Missing or unverified required output artifacts (spec file not consulted)taskda-code
Applies when
task -- The task points to an external instruction/guidance file and/or names specific deliverable files (figures, serialized arrays, JSON summaries) that the script must write to disk.
Pattern
The agent writes a script that implements its own guessed procedure (inventing filters, thresholds, category-mapping rules, or column semantics never stated) and saves only the one artifact it happened to think of, never opening/echoing the referenced guidance file and never producing the other required outputs; the final answer describes numbers in prose instead of the required files.
Detection procedure
  1. From the task text, list every artifact that must exist after the run (file names/extensions) and every rule that is said to live in an external spec.
  2. Grep the script for a read of that spec file and for a write/save call per listed artifact (e.g., savefig, to_json, np.save, to_csv), including the exact target paths/filenames and directory.
  3. Flag if any listed artifact has no corresponding write, if paths differ from what was requested, or if the script contains preprocessing/filtering/grouping rules with no traceable source in the task or spec (hard-coded thresholds, hand-written keyword-to-category maps, priority orderings).
  4. Check the answer: does it merely narrate numbers, or does it confirm each required file was created with the requested properties (size, colors, labels, ordering)?
Discriminator
A real violation is a missing/differently-named artifact or an assumption the agent invented because it never read the spec. It is not a violation if the artifact is written under the exact requested name via a different but equivalent API, or if a rule the reviewer doesn't recognize is verifiably quoted from the provided guidance/README.
Consequence
The grader checks for each expected file and reports WRONG/MISSING for every one that is absent or built from unsanctioned filtering, so the attempt scores zero even if the produced chart looks plausible.
id 0e833220f3ca · mined from da-code dacode-plot-pie-005@s17
raw text (what the judge reads)
### Missing or unverified required output artifacts (spec file not consulted)
- **Applies when**: `task` -- The task points to an external instruction/guidance file and/or names specific deliverable files (figures, serialized arrays, JSON summaries) that the script must write to disk.
- **Pattern**: The agent writes a script that implements its own guessed procedure (inventing filters, thresholds, category-mapping rules, or column semantics never stated) and saves only the one artifact it happened to think of, never opening/echoing the referenced guidance file and never producing the other required outputs; the final answer describes numbers in prose instead of the required files.
- **Detection procedure**:
  1. From the task text, list every artifact that must exist after the run (file names/extensions) and every rule that is said to live in an external spec.
  2. Grep the script for a read of that spec file and for a write/save call per listed artifact (e.g., savefig, to_json, np.save, to_csv), including the exact target paths/filenames and directory.
  3. Flag if any listed artifact has no corresponding write, if paths differ from what was requested, or if the script contains preprocessing/filtering/grouping rules with no traceable source in the task or spec (hard-coded thresholds, hand-written keyword-to-category maps, priority orderings).
  4. Check the answer: does it merely narrate numbers, or does it confirm each required file was created with the requested properties (size, colors, labels, ordering)?
- **Discriminator**: A real violation is a missing/differently-named artifact or an assumption the agent invented because it never read the spec. It is *not* a violation if the artifact is written under the exact requested name via a different but equivalent API, or if a rule the reviewer doesn't recognize is verifiably quoted from the provided guidance/README.
- **Consequence**: The grader checks for each expected file and reports WRONG/MISSING for every one that is absent or built from unsanctioned filtering, so the attempt scores zero even if the produced chart looks plausible.
913Imputation/parsing recipe not applied literally, and error never sanity-checked against a mean baselinetaskinfiagent-dabench
Applies when
task -- the prompt prescribes an exact preprocessing recipe (e.g., "impute missing values in columns X, Y, Z with their column means") plus a fixed split and error metric, and the raw columns may contain non-numeric or sentinel values.
Pattern
The attempt deviates from the prescribed recipe — dropping rows with missing/unparseable values instead of imputing, imputing only some of the named columns, letting to_numeric(errors='coerce')/dtype issues silently convert real values to NaN (or leaving them as strings so rows vanish), or imputing after the split from split-specific means — so the model is fit and scored on a different row set than intended; the reported error is then accepted without any magnitude check.
Detection procedure
  1. From the task, list every named preprocessing step, the target/feature roles, the split fraction, and the metric definition.
  2. In the scripts, trace row counts and dtypes: number of rows loaded → after any parsing/coercion → after imputation → training rows + test rows; confirm the sum equals the loaded count and that every named column was mean-imputed (no dropna, no filtering, no subsetting not requested).
  3. Check that the imputation values come from the intended data scope stated by the task and that all named columns, including the target, are handled.
  4. Compare the reported error to a trivial baseline computable from the data (e.g., variance of the target on the test rows); flag if the model's MSE is not clearly below it, or if its scale is implausible given the target's spread.
Discriminator
A real violation shows row counts shrinking, dtype coercion silently nulling values, missing columns in the imputation step, or an error at/above the mean-baseline variance; it is fine if all named columns are imputed, the row count is preserved end to end, and the error is comfortably below the variance of the target even if the exact number differs from expectations.
Consequence
The model is trained/scored on a truncated or mis-parsed sample, so the reported MSE is off by an order of magnitude from the reference value and the answer is graded wrong despite correct-looking format.
id 3586803e13da · mined from infiagent-dabench dabench-432@s17
raw text (what the judge reads)
### Imputation/parsing recipe not applied literally, and error never sanity-checked against a mean baseline
- **Applies when**: `task` -- the prompt prescribes an exact preprocessing recipe (e.g., "impute missing values in columns X, Y, Z with their column means") plus a fixed split and error metric, and the raw columns may contain non-numeric or sentinel values.
- **Pattern**: The attempt deviates from the prescribed recipe — dropping rows with missing/unparseable values instead of imputing, imputing only some of the named columns, letting `to_numeric(errors='coerce')`/dtype issues silently convert real values to NaN (or leaving them as strings so rows vanish), or imputing after the split from split-specific means — so the model is fit and scored on a different row set than intended; the reported error is then accepted without any magnitude check.
- **Detection procedure**:
  1. From the task, list every named preprocessing step, the target/feature roles, the split fraction, and the metric definition.
  2. In the scripts, trace row counts and dtypes: number of rows loaded → after any parsing/coercion → after imputation → training rows + test rows; confirm the sum equals the loaded count and that every named column was mean-imputed (no `dropna`, no filtering, no subsetting not requested).
  3. Check that the imputation values come from the intended data scope stated by the task and that all named columns, including the target, are handled.
  4. Compare the reported error to a trivial baseline computable from the data (e.g., variance of the target on the test rows); flag if the model's MSE is not clearly below it, or if its scale is implausible given the target's spread.
- **Discriminator**: A real violation shows row counts shrinking, dtype coercion silently nulling values, missing columns in the imputation step, or an error at/above the mean-baseline variance; it is fine if all named columns are imputed, the row count is preserved end to end, and the error is comfortably below the variance of the target even if the exact number differs from expectations.
- **Consequence**: The model is trained/scored on a truncated or mis-parsed sample, so the reported MSE is off by an order of magnitude from the reference value and the answer is graded wrong despite correct-looking format.
914Statistic computed on an unverified/implicitly filtered row set (no reproducible script or N reported)taskinfiagent-dabench
Applies when
task -- the task asks for a summary statistic or test computed over two or more columns of a table, and the agent's answer depends on exactly which rows are included after type coercion and missing-value handling.
Pattern
The agent runs an ad-hoc computation (often not saved) that silently drops or keeps rows — e.g. relying on a library's default pairwise/listwise NaN handling, coercing strings to numeric with errors='coerce', reading only part of the file, or leaving in sentinel/placeholder values — and reports a rounded value without ever stating the number of observations used or checking that the value is stable under the obvious alternative handling. Borderline rounding then lands one unit off the expected value.
Detection procedure
  1. Read the task to identify which rows are in scope (any stated filter; otherwise all rows with valid values in the relevant columns).
  2. In the scripts, locate the exact line producing the statistic and trace backwards: how the file was loaded (full file? correct delimiter/header?), dtypes of the columns, and where rows could be dropped or coerced (dropna, astype, to_numeric, boolean masks, sentinels like -9/-999/'NA').
  3. Check whether the script prints diagnostics — total rows, non-null counts per column, N actually entering the statistic, and the unrounded value — and whether the agent inspected them.
  4. Compare the reported value's formatting to the stated format/rounding rules (correct number of decimals for each quantity), and check that a nearly-tied rounding boundary was not accepted without confirming N.
Discriminator
A real violation is when the row set is never verified — no N/non-null counts printed, no script retained, or an unexamined coercion/drop step — so an alternative but equally defensible handling would change the reported digits. It is not a violation if the script explicitly reports counts and the unrounded statistic, and the drop rule is the only sensible one given the task (e.g. rows with genuinely missing values in either column excluded), even if the final number is close to a rounding boundary.
Consequence
The reported coefficient differs from the reference in the last reported decimal (and formatting of secondary values like the p-value may not match the requested precision), so exact-match checks on the numeric fields fail even though the qualitative conclusion is right.
id e001165d198e · mined from infiagent-dabench dabench-300@s17
raw text (what the judge reads)
### Statistic computed on an unverified/implicitly filtered row set (no reproducible script or N reported)
- **Applies when**: `task` -- the task asks for a summary statistic or test computed over two or more columns of a table, and the agent's answer depends on exactly which rows are included after type coercion and missing-value handling.
- **Pattern**: The agent runs an ad-hoc computation (often not saved) that silently drops or keeps rows — e.g. relying on a library's default pairwise/listwise NaN handling, coercing strings to numeric with `errors='coerce'`, reading only part of the file, or leaving in sentinel/placeholder values — and reports a rounded value without ever stating the number of observations used or checking that the value is stable under the obvious alternative handling. Borderline rounding then lands one unit off the expected value.
- **Detection procedure**:
  1. Read the task to identify which rows are in scope (any stated filter; otherwise all rows with valid values in the relevant columns).
  2. In the scripts, locate the exact line producing the statistic and trace backwards: how the file was loaded (full file? correct delimiter/header?), dtypes of the columns, and where rows could be dropped or coerced (`dropna`, `astype`, `to_numeric`, boolean masks, sentinels like -9/-999/'NA').
  3. Check whether the script prints diagnostics — total rows, non-null counts per column, N actually entering the statistic, and the unrounded value — and whether the agent inspected them.
  4. Compare the reported value's formatting to the stated format/rounding rules (correct number of decimals for each quantity), and check that a nearly-tied rounding boundary was not accepted without confirming N.
- **Discriminator**: A real violation is when the row set is never verified — no N/non-null counts printed, no script retained, or an unexamined coercion/drop step — so an alternative but equally defensible handling would change the reported digits. It is *not* a violation if the script explicitly reports counts and the unrounded statistic, and the drop rule is the only sensible one given the task (e.g. rows with genuinely missing values in either column excluded), even if the final number is close to a rounding boundary.
- **Consequence**: The reported coefficient differs from the reference in the last reported decimal (and formatting of secondary values like the p-value may not match the requested precision), so exact-match checks on the numeric fields fail even though the qualitative conclusion is right.
915Submitting predictions with no held-out accuracy estimate and no check for label recoverability from the provided reference datataskda-code
Applies when
task -- the agent must produce a prediction file for unlabeled rows while a larger labeled source/reference table is available locally, and the scripts fit a model and write predictions in one pass.
Pattern
The script trains a single default-hyperparameter pipeline on the whole labeled file, immediately predicts on the evaluation rows, and writes the output — with no train/validation split, no cross-validated score, no comparison of alternative feature sets/models, and no attempt to verify whether the evaluation rows also appear in the labeled source (e.g., by joining on the shared identifier or exact text key) so that many labels could be looked up exactly instead of guessed.
Detection procedure
  1. Read the task and note the accuracy-style bar implied by the grader and which auxiliary labeled data is available.
  2. Scan the scripts for any evaluation step: a hold-out split, cross_val_score, or a printed validation metric. If the only printed diagnostics are shapes and predicted-class counts, no quality evidence exists.
  3. Check whether the script ever joins/merges the evaluation rows against the labeled reference on an ID or exact-duplicate key, or at least measures the overlap; if it only uses one raw text column as features and ignores other informative columns/identifiers, flag it.
  4. Inspect the produced file: confirm row count equals the number of evaluation rows and that the label vocabulary/format matches the source, and note that nothing in the run demonstrates the predictions are better than an untested baseline.
Discriminator
A real violation is zero measured out-of-sample performance and zero exploration of exact-match label recovery; it is not a violation if the agent reports a validation/CV score above the plausible bar (or shows overlap is empty) and only then retrains on all data and predicts.
Consequence
The submitted file is well-formed but its accuracy is unknown and lands below the grader's threshold, so the single output check fails with no diagnostic to explain why.
id 72c0b1943234 · mined from da-code dacode-ml-multi-008@s17
raw text (what the judge reads)
### Submitting predictions with no held-out accuracy estimate and no check for label recoverability from the provided reference data
- **Applies when**: `task` -- the agent must produce a prediction file for unlabeled rows while a larger labeled source/reference table is available locally, and the scripts fit a model and write predictions in one pass.
- **Pattern**: The script trains a single default-hyperparameter pipeline on the whole labeled file, immediately predicts on the evaluation rows, and writes the output — with no train/validation split, no cross-validated score, no comparison of alternative feature sets/models, and no attempt to verify whether the evaluation rows also appear in the labeled source (e.g., by joining on the shared identifier or exact text key) so that many labels could be looked up exactly instead of guessed.
- **Detection procedure**:
  1. Read the task and note the accuracy-style bar implied by the grader and which auxiliary labeled data is available.
  2. Scan the scripts for any evaluation step: a hold-out split, `cross_val_score`, or a printed validation metric. If the only printed diagnostics are shapes and predicted-class counts, no quality evidence exists.
  3. Check whether the script ever joins/merges the evaluation rows against the labeled reference on an ID or exact-duplicate key, or at least measures the overlap; if it only uses one raw text column as features and ignores other informative columns/identifiers, flag it.
  4. Inspect the produced file: confirm row count equals the number of evaluation rows and that the label vocabulary/format matches the source, and note that nothing in the run demonstrates the predictions are better than an untested baseline.
- **Discriminator**: A real violation is zero measured out-of-sample performance and zero exploration of exact-match label recovery; it is *not* a violation if the agent reports a validation/CV score above the plausible bar (or shows overlap is empty) and only then retrains on all data and predicts.
- **Consequence**: The submitted file is well-formed but its accuracy is unknown and lands below the grader's threshold, so the single output check fails with no diagnostic to explain why.
916Output artifacts written to an arbitrary location/format instead of the expected deliverable conventiontaskinfiagent-dabench
Applies when
task -- The requested answer includes a file path to a derived dataset/artifact the agent must create.
Pattern
The agent invents its own output directory, filename, and file contents (e.g., writes to its own home/scratch dir, keeps only a couple of columns, or renames/duplicates the transformed column) instead of writing the full transformed dataset back to the same working directory as the input data with a name that clearly derives from the source file and the requested transformation.
Detection procedure
  1. From the task statement, note that a file path is a graded field and note the directory the input data was read from and its filename.
  2. In the script, find the to_csv/save call: check the directory matches the data directory used for input, the filename is derived from the source file plus the transformation, and the saved frame is the full dataset with the transformation applied in place.
  3. Compare the path in the final answer against the path actually written and against the input-data directory; flag any mismatch or any self-chosen scratch path.
  4. Check that no columns were silently dropped or renamed in a way that a consumer of the file could not reproduce the original schema.
Discriminator
A real violation is when the artifact lives outside the data directory, is named arbitrarily, or contains only a subset/renamed version of the data; it is fine if the task explicitly designates an output directory/name and the agent follows it, or if the agent writes into the input data directory with a transparently derived name and the full transformed table.
Consequence
The numeric fields may all match, but the file-path field fails string comparison against the expected artifact path, marking the whole answer wrong.
id 63db72121d11 · mined from infiagent-dabench dabench-743@s17
raw text (what the judge reads)
### Output artifacts written to an arbitrary location/format instead of the expected deliverable convention
- **Applies when**: `task` -- The requested answer includes a file path to a derived dataset/artifact the agent must create.
- **Pattern**: The agent invents its own output directory, filename, and file contents (e.g., writes to its own home/scratch dir, keeps only a couple of columns, or renames/duplicates the transformed column) instead of writing the full transformed dataset back to the same working directory as the input data with a name that clearly derives from the source file and the requested transformation.
- **Detection procedure**:
  1. From the task statement, note that a file path is a graded field and note the directory the input data was read from and its filename.
  2. In the script, find the `to_csv`/`save` call: check the directory matches the data directory used for input, the filename is derived from the source file plus the transformation, and the saved frame is the full dataset with the transformation applied in place.
  3. Compare the path in the final answer against the path actually written and against the input-data directory; flag any mismatch or any self-chosen scratch path.
  4. Check that no columns were silently dropped or renamed in a way that a consumer of the file could not reproduce the original schema.
- **Discriminator**: A real violation is when the artifact lives outside the data directory, is named arbitrarily, or contains only a subset/renamed version of the data; it is fine if the task explicitly designates an output directory/name and the agent follows it, or if the agent writes into the input data directory with a transparently derived name and the full transformed table.
- **Consequence**: The numeric fields may all match, but the file-path field fails string comparison against the expected artifact path, marking the whole answer wrong.
917Missing or partial deliverable artifacts required by the task/spec filetaskda-code
Applies when
task -- The task points to an external configuration/guideline file (e.g., a YAML/JSON spec) and/or asks for saved outputs, and the grader checks specific files on disk.
Pattern
The attempt produces only the one obviously named output (e.g., the image) and reports numbers in prose, while silently skipping other artifacts the spec/task implies (serialized plot data, saved arrays/tables of the computed values) and not persisting the scripts that generated them; the spec is also only partially applied (a couple of styling keys quoted, the rest unverified).
Detection procedure
  1. Read the task statement and the referenced spec/config file and enumerate every required artifact, filename, and formatting key (labels, ordering, title, legend, sizes, colors, data serialization).
  2. Read the scripts and list every file actually written and every spec key actually consumed; check that the config is loaded programmatically rather than hard-coded from a glance.
  3. Diff the two lists: any required artifact not written, or any spec key never read, is a violation; also confirm the scripts themselves exist/were saved and are rerunnable.
  4. Check the final answer states where each artifact was saved rather than only narrating numbers.
Discriminator
A real violation is a required output/spec element with no corresponding write or read in the code; a look-alike that is fine is an artifact produced under the exact required name/format via a different but equivalent code path, or spec keys that are genuinely absent from the config file.
Consequence
The grader reports expected files as WRONG/MISSING (zero checks passed) even when the intermediate analysis conclusion is right.
id 05ac68a64427 · mined from da-code dacode-plot-pie-008@s17
raw text (what the judge reads)
### Missing or partial deliverable artifacts required by the task/spec file
- **Applies when**: `task` -- The task points to an external configuration/guideline file (e.g., a YAML/JSON spec) and/or asks for saved outputs, and the grader checks specific files on disk.
- **Pattern**: The attempt produces only the one obviously named output (e.g., the image) and reports numbers in prose, while silently skipping other artifacts the spec/task implies (serialized plot data, saved arrays/tables of the computed values) and not persisting the scripts that generated them; the spec is also only partially applied (a couple of styling keys quoted, the rest unverified).
- **Detection procedure**:
  1. Read the task statement and the referenced spec/config file and enumerate every required artifact, filename, and formatting key (labels, ordering, title, legend, sizes, colors, data serialization).
  2. Read the scripts and list every file actually written and every spec key actually consumed; check that the config is loaded programmatically rather than hard-coded from a glance.
  3. Diff the two lists: any required artifact not written, or any spec key never read, is a violation; also confirm the scripts themselves exist/were saved and are rerunnable.
  4. Check the final answer states where each artifact was saved rather than only narrating numbers.
- **Discriminator**: A real violation is a required output/spec element with no corresponding write or read in the code; a look-alike that is fine is an artifact produced under the exact required name/format via a different but equivalent code path, or spec keys that are genuinely absent from the config file.
- **Consequence**: The grader reports expected files as WRONG/MISSING (zero checks passed) even when the intermediate analysis conclusion is right.
918Ignoring the provided output template's schema and inventing one's own categoriestaskda-code
Applies when
task -- the task supplies a pre-existing result file/template ("enter results into X, adhering strictly to its format") and the analysis requires assigning records to categories or filling specific fields.
Pattern
The agent never reads the template to learn its exact columns, row labels, category names, ordering and value types; instead it defines its own grouping scheme (self-chosen bin edges/labels, extra "Unknown"/catch-all buckets, extra summary rows) and writes a file shaped by that scheme, or reports numbers only in prose.
Detection procedure
  1. In the task statement, note that a target file is provided and identify the constraint "match its format".
  2. In the scripts, check for an explicit read of the template file and for code that aligns output rows/columns/labels to it (e.g., filling values into the loaded frame, or asserting the header/label set matches). Absence of any such read is the red flag.
  3. In the answer/output, compare the category labels, their number and order, and the column names against the template's; flag self-invented thresholds, extra buckets for missing values, or renamed/reordered fields.
  4. Sanity-check a few category assignments against domain-obvious cases and against expected distribution (a category with 0 members or an implausibly large "unknown" group signals a wrong assignment rule).
Discriminator
A real violation is when the template's labels/columns/rounding could have been read but were replaced by the agent's own conventions; it is fine if the agent reads the template, keeps its exact schema, and merely documents an internally defined rule for values the template leaves undefined.
Consequence
The grader compares the produced file cell-by-cell against the expected file and marks it WRONG/MISSING because keys, labels, or row sets do not align, even if the underlying computation were partially right.
id e8c9ec0613a0 · mined from da-code dacode-dm-csv-001@s17
raw text (what the judge reads)
### Ignoring the provided output template's schema and inventing one's own categories
- **Applies when**: `task` -- the task supplies a pre-existing result file/template ("enter results into X, adhering strictly to its format") and the analysis requires assigning records to categories or filling specific fields.
- **Pattern**: The agent never reads the template to learn its exact columns, row labels, category names, ordering and value types; instead it defines its own grouping scheme (self-chosen bin edges/labels, extra "Unknown"/catch-all buckets, extra summary rows) and writes a file shaped by that scheme, or reports numbers only in prose.
- **Detection procedure**:
  1. In the task statement, note that a target file is provided and identify the constraint "match its format".
  2. In the scripts, check for an explicit read of the template file and for code that aligns output rows/columns/labels to it (e.g., filling values into the loaded frame, or asserting the header/label set matches). Absence of any such read is the red flag.
  3. In the answer/output, compare the category labels, their number and order, and the column names against the template's; flag self-invented thresholds, extra buckets for missing values, or renamed/reordered fields.
  4. Sanity-check a few category assignments against domain-obvious cases and against expected distribution (a category with 0 members or an implausibly large "unknown" group signals a wrong assignment rule).
- **Discriminator**: A real violation is when the template's labels/columns/rounding could have been read but were replaced by the agent's own conventions; it is fine if the agent reads the template, keeps its exact schema, and merely documents an internally defined rule for values the template leaves undefined.
- **Consequence**: The grader compares the produced file cell-by-cell against the expected file and marks it WRONG/MISSING because keys, labels, or row sets do not align, even if the underlying computation were partially right.
919Derived per-entity quantities and subgroup splits not verified before correlatingtaskinfiagent-dabench
Applies when
task -- the task asks for a correlation/statistic between two quantities that must first be derived by aggregating raw records to one row per entity (e.g., max-of-a-level, span between first/last timestamps) and then split into subgroups by a threshold such as a median.
Pattern
The attempt computes the statistic directly on raw record-level rows, or aggregates with a subtly different definition (wrong grouping key, duration measured in the wrong unit or as count-of-rows instead of elapsed time, threshold computed after instead of before deduplication, ties at the median assigned inconsistently, missing/duplicate rows not dropped), so the numeric coefficient is close to but not equal to the correct value; no scripts or intermediate counts are saved to check it.
Detection procedure
  1. From the task, write down the intended unit of analysis (one row per entity) and the exact definition of each derived variable and of the subgroup threshold.
  2. In the scripts, locate the group-by/aggregation step and confirm the grouping key uniquely identifies an entity, that each derived variable matches the stated definition and unit, and that the threshold is computed on the same deduplicated entity-level table (and check how rows equal to the threshold are assigned).
  3. Check that the script prints sanity values — number of entities, sizes of the two subgroups, min/max/range of each derived variable — and that these are plausible; absence of such prints or of any saved script is itself a flag.
  4. Compare the reported statistic and its formatting (rounding to the required number of decimals, e.g. a p-value printed as 0.0 rather than to four places) against the constraints.
Discriminator
A real violation is an aggregation/split definition that provably differs from the task wording or leaves the entity table with duplicate or missing entities; a look-alike that is fine is an equivalent but differently coded aggregation whose printed entity counts, subgroup sizes and variable ranges match the specification.
Consequence
The relationship label may still come out right, but the correlation coefficient is off by a few hundredths from ground truth, so the exact-value checks fail and the answer is graded incorrect.
id 53b6b5a1f740 · mined from infiagent-dabench dabench-431@s17
raw text (what the judge reads)
### Derived per-entity quantities and subgroup splits not verified before correlating
- **Applies when**: `task` -- the task asks for a correlation/statistic between two quantities that must first be derived by aggregating raw records to one row per entity (e.g., max-of-a-level, span between first/last timestamps) and then split into subgroups by a threshold such as a median.
- **Pattern**: The attempt computes the statistic directly on raw record-level rows, or aggregates with a subtly different definition (wrong grouping key, duration measured in the wrong unit or as count-of-rows instead of elapsed time, threshold computed after instead of before deduplication, ties at the median assigned inconsistently, missing/duplicate rows not dropped), so the numeric coefficient is close to but not equal to the correct value; no scripts or intermediate counts are saved to check it.
- **Detection procedure**:
  1. From the task, write down the intended unit of analysis (one row per entity) and the exact definition of each derived variable and of the subgroup threshold.
  2. In the scripts, locate the group-by/aggregation step and confirm the grouping key uniquely identifies an entity, that each derived variable matches the stated definition and unit, and that the threshold is computed on the same deduplicated entity-level table (and check how rows equal to the threshold are assigned).
  3. Check that the script prints sanity values — number of entities, sizes of the two subgroups, min/max/range of each derived variable — and that these are plausible; absence of such prints or of any saved script is itself a flag.
  4. Compare the reported statistic and its formatting (rounding to the required number of decimals, e.g. a p-value printed as `0.0` rather than to four places) against the constraints.
- **Discriminator**: A real violation is an aggregation/split definition that provably differs from the task wording or leaves the entity table with duplicate or missing entities; a look-alike that is fine is an equivalent but differently coded aggregation whose printed entity counts, subgroup sizes and variable ranges match the specification.
- **Consequence**: The relationship label may still come out right, but the correlation coefficient is off by a few hundredths from ground truth, so the exact-value checks fail and the answer is graded incorrect.
920Accepting a weak model without verifying that the validation setup mirrors how the test rows were split off (and without checking train/test overlap)taskda-code
Applies when
task -- the task asks for predictions on a provided held-out file that will be graded against hidden true targets, and the script builds its own validation split from a separate full/complete training file.
Pattern
The agent trains with a default random split, reports mediocre validation accuracy, and submits without checking how the held-out file relates to the training file — whether the split was temporal/blocked, whether the held-out rows (or their targets) are actually present in the "complete" training file, whether the identifier/timestamp column it dropped encodes the split or is itself a strong predictor, and whether feature columns line up by name and order between the two files.
Detection procedure
  1. From the task, note that the graded artifact is per-row predictions compared to unseen truth, so accuracy — not just file format — decides the verdict.
  2. In the script, check for an explicit comparison of the two files: row counts vs. the full dataset, a merge/lookup of the held-out rows' keys against the training file, the column sets of train vs. test compared by name and order, and whether the validation split is random or matched to the structure (time order/contiguous block) that separates the held-out rows.
  3. Check whether any dropped column (timestamp, index, key) was examined for predictive or join value before discarding, and whether feature engineering/model alternatives were compared.
  4. In the answer, look for whether reported validation performance is claimed to be representative of the held-out set, and whether any sanity check compares prediction distribution (mean, range, count) against the training target distribution and the expected row count.
Discriminator
A fine attempt either shows the held-out keys are absent from the training file and uses a split that reproduces the held-out structure (e.g., last-block/temporal validation), or documents evidence that a random split is appropriate; a violation drops the key column unexamined, never compares file contents/columns, and treats a single random-split score as sufficient evidence of submission quality.
Consequence
The submitted file has the right shape and column name but per-row values that miss the graded error/correlation threshold, so the file check fails while the agent's self-reported validation metric looks acceptable.
id 58aa5b072f66 · mined from da-code dacode-ml-regression-015@s17
raw text (what the judge reads)
### Accepting a weak model without verifying that the validation setup mirrors how the test rows were split off (and without checking train/test overlap)
- **Applies when**: `task` -- the task asks for predictions on a provided held-out file that will be graded against hidden true targets, and the script builds its own validation split from a separate full/complete training file.
- **Pattern**: The agent trains with a default random split, reports mediocre validation accuracy, and submits without checking how the held-out file relates to the training file — whether the split was temporal/blocked, whether the held-out rows (or their targets) are actually present in the "complete" training file, whether the identifier/timestamp column it dropped encodes the split or is itself a strong predictor, and whether feature columns line up by name and order between the two files.
- **Detection procedure**:
  1. From the task, note that the graded artifact is per-row predictions compared to unseen truth, so accuracy — not just file format — decides the verdict.
  2. In the script, check for an explicit comparison of the two files: row counts vs. the full dataset, a merge/lookup of the held-out rows' keys against the training file, the column sets of train vs. test compared by name and order, and whether the validation split is random or matched to the structure (time order/contiguous block) that separates the held-out rows.
  3. Check whether any dropped column (timestamp, index, key) was examined for predictive or join value before discarding, and whether feature engineering/model alternatives were compared.
  4. In the answer, look for whether reported validation performance is claimed to be representative of the held-out set, and whether any sanity check compares prediction distribution (mean, range, count) against the training target distribution and the expected row count.
- **Discriminator**: A fine attempt either shows the held-out keys are absent from the training file and uses a split that reproduces the held-out structure (e.g., last-block/temporal validation), or documents evidence that a random split is appropriate; a violation drops the key column unexamined, never compares file contents/columns, and treats a single random-split score as sufficient evidence of submission quality.
- **Consequence**: The submitted file has the right shape and column name but per-row values that miss the graded error/correlation threshold, so the file check fails while the agent's self-reported validation metric looks acceptable.
921Output does not instantiate the requested JSON schema (scalar substituted for the specified container type / file not produced)taskda-code
Applies when
task -- the task supplies a literal output template (keys, bracket/list markers, file name) that the deliverable must fill in, and the agent produces a final JSON answer.
Pattern
The agent fills the template with a value of a different type or shape than shown (e.g., a bare string where a list/array placeholder is given, a nested object where a flat value is asked, extra or renamed keys), and/or never writes the artifact to the exact required output file — often with no saved script, so the schema can never be re-checked or regenerated.
Detection procedure
  1. Copy the literal template from the task statement and note, for each key, the exact container type implied by the placeholder ([...] = array, "..." = string, etc.) and the required output file name/path.
  2. Inspect the scripts for the code that serializes the result: confirm it builds each value in that container type (wrapping single results in a list if the template shows one) and dumps to the required file; if no script exists, treat the result as unverifiable.
  3. Compare the submitted answer key-by-key against the template: same key names, same order-insensitive key set, same value types, and ties handled as multiple entries if the container permits them.
  4. Flag if any key's value type differs from the placeholder, if keys are missing/extra, or if the required result file is absent.
Discriminator
A real violation is a structural mismatch with the given template (scalar vs. array, missing file, renamed key) even when the underlying computed values look right; a look-alike that is fine is a correct container holding one element, or cosmetic differences (whitespace, key order, int vs. float of the same number) that a JSON comparison would treat as equivalent.
Consequence
The grader's exact-match/schema check on the expected result file fails (WRONG/MISSING), scoring 0 even though the analytic content may have been correct.
id 3cb32d55c973 · mined from da-code dacode-di-text-001@s17
raw text (what the judge reads)
### Output does not instantiate the requested JSON schema (scalar substituted for the specified container type / file not produced)
- **Applies when**: `task` -- the task supplies a literal output template (keys, bracket/list markers, file name) that the deliverable must fill in, and the agent produces a final JSON answer.
- **Pattern**: The agent fills the template with a value of a different type or shape than shown (e.g., a bare string where a list/array placeholder is given, a nested object where a flat value is asked, extra or renamed keys), and/or never writes the artifact to the exact required output file — often with no saved script, so the schema can never be re-checked or regenerated.
- **Detection procedure**:
  1. Copy the literal template from the task statement and note, for each key, the exact container type implied by the placeholder (`[...]` = array, `"..."` = string, etc.) and the required output file name/path.
  2. Inspect the scripts for the code that serializes the result: confirm it builds each value in that container type (wrapping single results in a list if the template shows one) and dumps to the required file; if no script exists, treat the result as unverifiable.
  3. Compare the submitted answer key-by-key against the template: same key names, same order-insensitive key set, same value types, and ties handled as multiple entries if the container permits them.
  4. Flag if any key's value type differs from the placeholder, if keys are missing/extra, or if the required result file is absent.
- **Discriminator**: A real violation is a structural mismatch with the given template (scalar vs. array, missing file, renamed key) even when the underlying computed values look right; a look-alike that is fine is a correct container holding one element, or cosmetic differences (whitespace, key order, int vs. float of the same number) that a JSON comparison would treat as equivalent.
- **Consequence**: The grader's exact-match/schema check on the expected result file fails (WRONG/MISSING), scoring 0 even though the analytic content may have been correct.
922Output file omits requested intermediate/derived columnstaskda-code
Applies when
task -- the task asks for a saved file that includes several named quantities (per-dimension scores, a composite score, a segment label, an assigned level), and the script writes the file by selecting a subset of the computed columns.
Pattern
The script correctly computes all the intermediate quantities in memory, then exports only the single "final" column (plus an ID), silently dropping the other artifacts the task explicitly said to include; the answer text describes the full methodology, masking the truncated output.
Detection procedure
  1. From the task statement, list every quantity that must appear in the saved file (each component score, any composite score/segment string, the final label/level, and the key identifier).
  2. In the script, find the line(s) that build the exported frame (df[[...]], to_csv, index/header settings) and enumerate the columns actually written, plus whether the identifier is a column or an index.
  3. Diff the two lists; also check the answer for a stated column layout and confirm it matches the export, not just the in-memory frame.
  4. Verify no sanity check on the output (shape/columns/row count) was run after writing.
Discriminator
A real violation is when a quantity the task names as part of the deliverable is computed but not written (or the ID is lost to the index). It is not a violation if the task only asked for the final label, or if the extra quantities are written under reasonably equivalent names/order — naming variation is fine, missing content is not.
Consequence
File-level comparison against the expected output fails on column mismatch (missing score/segment fields), so the check scores 0 even though the underlying computation was right.
id 69e083023ac4 · mined from da-code dacode-dm-csv-052@s17
raw text (what the judge reads)
### Output file omits requested intermediate/derived columns
- **Applies when**: `task` -- the task asks for a saved file that includes several named quantities (per-dimension scores, a composite score, a segment label, an assigned level), and the script writes the file by selecting a subset of the computed columns.
- **Pattern**: The script correctly computes all the intermediate quantities in memory, then exports only the single "final" column (plus an ID), silently dropping the other artifacts the task explicitly said to include; the answer text describes the full methodology, masking the truncated output.
- **Detection procedure**:
  1. From the task statement, list every quantity that must appear in the saved file (each component score, any composite score/segment string, the final label/level, and the key identifier).
  2. In the script, find the line(s) that build the exported frame (`df[[...]]`, `to_csv`, index/header settings) and enumerate the columns actually written, plus whether the identifier is a column or an index.
  3. Diff the two lists; also check the answer for a stated column layout and confirm it matches the export, not just the in-memory frame.
  4. Verify no sanity check on the output (shape/columns/row count) was run after writing.
- **Discriminator**: A real violation is when a quantity the task names as part of the deliverable is computed but not written (or the ID is lost to the index). It is *not* a violation if the task only asked for the final label, or if the extra quantities are written under reasonably equivalent names/order — naming variation is fine, missing content is not.
- **Consequence**: File-level comparison against the expected output fails on column mismatch (missing score/segment fields), so the check scores 0 even though the underlying computation was right.
923Ignoring a stated population/subset filter before computing statisticstaskda-code
Applies when
task -- the task names a specific subgroup, condition, or category of records (species, cohort, region, class, split) and the loaded file contains additional groups beyond that subgroup.
Pattern
The script loads the full table and computes the requested statistic on all rows (or only groups by a secondary key such as time period), never applying the subgroup filter the task specifies; it may even print the distinct values of the grouping column but never uses them to subset, and no sanity check compares the resulting magnitudes against expectation.
Detection procedure
  1. From the task text, list every filtering condition stated or implied (subgroup label, time values, valid-record conditions) and the exact grouping the output should have.
  2. In the script, find the line(s) between loading and computing; confirm a boolean mask or equivalent exists for each listed condition, and check the row counts per output group are consistent with only that subgroup.
  3. In the answer, check reported sample sizes and value magnitudes against what the filtered subgroup should plausibly give (e.g., counts far larger than the subgroup, or a statistic sitting midway between known group values, indicates pooling).
  4. Also verify the script reads/echoes any provided template file so column names, ordering, and numeric formatting of the output match it rather than being invented.
Discriminator
A real violation is when the data demonstrably contains other groups (script itself enumerates them, or counts exceed the subgroup) and no filter is applied; it is fine if the file was pre-filtered upstream and the script verifies this (e.g., asserts a single unique value of the grouping column) or if the filter is applied at load time.
Consequence
The written result file contains statistics for the wrong population — means and bootstrap intervals shifted well outside the expected values — so every value check against the reference file fails even though the bootstrap procedure itself is correct.
id 87e76a9a1597 · mined from da-code dacode-data-sa-029@s17
raw text (what the judge reads)
### Ignoring a stated population/subset filter before computing statistics
- **Applies when**: `task` -- the task names a specific subgroup, condition, or category of records (species, cohort, region, class, split) and the loaded file contains additional groups beyond that subgroup.
- **Pattern**: The script loads the full table and computes the requested statistic on all rows (or only groups by a secondary key such as time period), never applying the subgroup filter the task specifies; it may even print the distinct values of the grouping column but never uses them to subset, and no sanity check compares the resulting magnitudes against expectation.
- **Detection procedure**:
  1. From the task text, list every filtering condition stated or implied (subgroup label, time values, valid-record conditions) and the exact grouping the output should have.
  2. In the script, find the line(s) between loading and computing; confirm a boolean mask or equivalent exists for *each* listed condition, and check the row counts per output group are consistent with only that subgroup.
  3. In the answer, check reported sample sizes and value magnitudes against what the filtered subgroup should plausibly give (e.g., counts far larger than the subgroup, or a statistic sitting midway between known group values, indicates pooling).
  4. Also verify the script reads/echoes any provided template file so column names, ordering, and numeric formatting of the output match it rather than being invented.
- **Discriminator**: A real violation is when the data demonstrably contains other groups (script itself enumerates them, or counts exceed the subgroup) and no filter is applied; it is fine if the file was pre-filtered upstream and the script verifies this (e.g., asserts a single unique value of the grouping column) or if the filter is applied at load time.
- **Consequence**: The written result file contains statistics for the wrong population — means and bootstrap intervals shifted well outside the expected values — so every value check against the reference file fails even though the bootstrap procedure itself is correct.
924Near-zero (or otherwise implausible) statistics accepted without a sanity check on the conversion/parsing steptaskda-code
Applies when
task -- a task requires coercing raw columns to numeric/typed values, dropping missing rows, and then reporting a bounded statistic (correlation, accuracy, ratio) computed from them.
Pattern
The agent applies a blunt type conversion (e.g. to_numeric(..., errors='coerce'), blind astype, regex strip) and/or an index-misaligning operation, never inspects the resulting values, and reports a statistic whose magnitude is implausible for the described relationship (e.g. all off-diagonal correlations ≈ 0 among conceptually related rating fields) as if it were a valid finding.
Detection procedure
  1. Read the task to identify which columns must be converted and what the statistic's expected range/behaviour is; note whether domain knowledge implies a non-trivial relationship.
  2. In the scripts, check whether after conversion/filtering the agent printed diagnostics — row count retained, dtype, min/max/unique values, head of the cleaned frame — and whether any operation (sorting, resetting/dropping index, per-column dropna, separate Series extraction) could misalign rows across columns.
  3. Compare the reported statistic to what those diagnostics would predict: bounded rating-like variables that survived cleaning should not yield uniformly ~0 (or ~1) associations; flag if no diagnostic exists to rule out silent corruption.
  4. Confirm the saved artifact was re-read and its shape/labels/values checked against the required sample format before reporting.
Discriminator
A genuine violation is when the implausible value is unexplained and unverified — no printout of cleaned values, no alignment check, no comparison against a quick alternative computation. It is not a violation if the agent showed the cleaned data (valid parsed numbers, aligned rows) and the near-zero/extreme result is reproducible from that inspected data.
Consequence
The saved file contains numerically wrong entries (though correctly shaped), so an exact/tolerance comparison against the expected matrix fails and the task scores 0.
id 38e31fee7429 · mined from da-code dacode-data-sa-026@s17
raw text (what the judge reads)
### Near-zero (or otherwise implausible) statistics accepted without a sanity check on the conversion/parsing step
- **Applies when**: `task` -- a task requires coercing raw columns to numeric/typed values, dropping missing rows, and then reporting a bounded statistic (correlation, accuracy, ratio) computed from them.
- **Pattern**: The agent applies a blunt type conversion (e.g. `to_numeric(..., errors='coerce')`, blind `astype`, regex strip) and/or an index-misaligning operation, never inspects the resulting values, and reports a statistic whose magnitude is implausible for the described relationship (e.g. all off-diagonal correlations ≈ 0 among conceptually related rating fields) as if it were a valid finding.
- **Detection procedure**:
  1. Read the task to identify which columns must be converted and what the statistic's expected range/behaviour is; note whether domain knowledge implies a non-trivial relationship.
  2. In the scripts, check whether after conversion/filtering the agent printed diagnostics — row count retained, dtype, min/max/unique values, head of the cleaned frame — and whether any operation (sorting, resetting/dropping index, per-column `dropna`, separate Series extraction) could misalign rows across columns.
  3. Compare the reported statistic to what those diagnostics would predict: bounded rating-like variables that survived cleaning should not yield uniformly ~0 (or ~1) associations; flag if no diagnostic exists to rule out silent corruption.
  4. Confirm the saved artifact was re-read and its shape/labels/values checked against the required sample format before reporting.
- **Discriminator**: A genuine violation is when the implausible value is unexplained and unverified — no printout of cleaned values, no alignment check, no comparison against a quick alternative computation. It is *not* a violation if the agent showed the cleaned data (valid parsed numbers, aligned rows) and the near-zero/extreme result is reproducible from that inspected data.
- **Consequence**: The saved file contains numerically wrong entries (though correctly shaped), so an exact/tolerance comparison against the expected matrix fails and the task scores 0.
925Unvalidated clustering solution: skewed features and degenerate micro-clusterstaskda-code
Applies when
task -- the task asks for an unsupervised grouping of entities into "an appropriate number of clusters" and the script aggregates raw transactional/heavy-tailed data into features, scales them, and picks k by an internal score.
Pattern
The attempt feeds highly right-skewed, outlier-dominated aggregates straight into a distance-based algorithm (only z-scoring, no log/winsorize/outlier handling), selects k by maximizing an internal index without checking the resulting partition, and reports a solution where most entities fall in a couple of clusters while several clusters contain a handful of extreme outliers; no sanity check on cluster sizes, feature count, or the exact output schema/row count is performed.
Detection procedure
  1. Read the task to note the required deliverable schema (feature columns + label column) and the fact that clusters should describe interpretable groups of entities, not outliers.
  2. In the script, check whether the engineered features are heavy-tailed sums/totals and whether any transformation or outlier treatment is applied before scaling; also check whether the feature set is a defensible minimal set or an ad-hoc pile of collinear derived ratios.
  3. Check whether the code validates the chosen partition after fitting: cluster size distribution, minimum cluster size relative to n, and that the saved file's rows equal the number of entities and columns match the requested names/order.
  4. Inspect the reported cluster summary in the answer for clusters with sizes of a few entities or extreme means, and for a chosen k that only wins because such singleton-like clusters inflate the internal score.
Discriminator
A real violation is when cluster sizes are wildly degenerate (clusters of ~1-30 out of thousands) and no transformation/outlier step or size sanity check exists; a look-alike that is fine is a moderately imbalanced solution where skew was addressed (log/robust scaling or outlier trimming) and the author explicitly justified small clusters as a meaningful high-value segment while confirming output shape and column naming.
Consequence
The saved cluster file encodes an outlier-driven, non-reproducible partition (and possibly an unexpected feature set/row count), so the expected-file check fails even though the pipeline "ran" and reported plausible-looking metrics.
id de5925865a66 · mined from da-code dacode-ml-cluster-019@s17
raw text (what the judge reads)
### Unvalidated clustering solution: skewed features and degenerate micro-clusters
- **Applies when**: `task` -- the task asks for an unsupervised grouping of entities into "an appropriate number of clusters" and the script aggregates raw transactional/heavy-tailed data into features, scales them, and picks k by an internal score.
- **Pattern**: The attempt feeds highly right-skewed, outlier-dominated aggregates straight into a distance-based algorithm (only z-scoring, no log/winsorize/outlier handling), selects k by maximizing an internal index without checking the resulting partition, and reports a solution where most entities fall in a couple of clusters while several clusters contain a handful of extreme outliers; no sanity check on cluster sizes, feature count, or the exact output schema/row count is performed.
- **Detection procedure**:
  1. Read the task to note the required deliverable schema (feature columns + label column) and the fact that clusters should describe interpretable groups of entities, not outliers.
  2. In the script, check whether the engineered features are heavy-tailed sums/totals and whether any transformation or outlier treatment is applied before scaling; also check whether the feature set is a defensible minimal set or an ad-hoc pile of collinear derived ratios.
  3. Check whether the code validates the chosen partition after fitting: cluster size distribution, minimum cluster size relative to n, and that the saved file's rows equal the number of entities and columns match the requested names/order.
  4. Inspect the reported cluster summary in the answer for clusters with sizes of a few entities or extreme means, and for a chosen k that only wins because such singleton-like clusters inflate the internal score.
- **Discriminator**: A real violation is when cluster sizes are wildly degenerate (clusters of ~1-30 out of thousands) and no transformation/outlier step or size sanity check exists; a look-alike that is fine is a moderately imbalanced solution where skew was addressed (log/robust scaling or outlier trimming) and the author explicitly justified small clusters as a meaningful high-value segment while confirming output shape and column naming.
- **Consequence**: The saved cluster file encodes an outlier-driven, non-reproducible partition (and possibly an unexpected feature set/row count), so the expected-file check fails even though the pipeline "ran" and reported plausible-looking metrics.
926Unrequested unit conversion of stored valuestaskinfiagent-dabench
Applies when
task -- reported statistics are computed from columns whose stored scale (e.g., fractions/proportions vs. percent, cents vs. dollars, seconds vs. minutes) may differ from the wording used in the question.
Pattern
The script inspects the raw values, notices they are on one scale, and then multiplies/divides to match the phrasing in the prompt instead of reporting statistics in the data's native units; the rounding constraint is then applied to the rescaled numbers.
Detection procedure
1. Read the task for any explicit statement of units or a conversion instruction. 2. In the script, look for a scaling step (*100, /100, unit arithmetic) applied to the column before/after computing the statistic, and check whether the task actually asked for it. 3. Cross-check the requested rounding precision against the scale: a constraint asking for 2–3 decimal places implies values on the native (small-magnitude) scale, whereas rescaled values would make those decimals meaningless. 4. Confirm the final reported numbers are in the same units as the source column unless conversion was demanded.
Discriminator
A real violation is a conversion invented by the agent from the question's phrasing alone (no explicit unit requirement) and inconsistent with the stated rounding granularity; it is fine if the task explicitly names target units, or if the column is documented in units that genuinely require conversion to satisfy the request.
Consequence
Every numeric field is off by the conversion factor (e.g., 31.76 vs. 0.32), so all value checks fail even though the underlying computation and the normality verdicts were correct.
id 7c7677a208e9 · mined from infiagent-dabench dabench-144@s18
raw text (what the judge reads)
### Unrequested unit conversion of stored values
- **Applies when**: `task` -- reported statistics are computed from columns whose stored scale (e.g., fractions/proportions vs. percent, cents vs. dollars, seconds vs. minutes) may differ from the wording used in the question.
- **Pattern**: The script inspects the raw values, notices they are on one scale, and then multiplies/divides to match the phrasing in the prompt instead of reporting statistics in the data's native units; the rounding constraint is then applied to the rescaled numbers.
- **Detection procedure**: 1. Read the task for any explicit statement of units or a conversion instruction. 2. In the script, look for a scaling step (`*100`, `/100`, unit arithmetic) applied to the column before/after computing the statistic, and check whether the task actually asked for it. 3. Cross-check the requested rounding precision against the scale: a constraint asking for 2–3 decimal places implies values on the native (small-magnitude) scale, whereas rescaled values would make those decimals meaningless. 4. Confirm the final reported numbers are in the same units as the source column unless conversion was demanded.
- **Discriminator**: A real violation is a conversion invented by the agent from the question's phrasing alone (no explicit unit requirement) and inconsistent with the stated rounding granularity; it is fine if the task explicitly names target units, or if the column is documented in units that genuinely require conversion to satisfy the request.
- **Consequence**: Every numeric field is off by the conversion factor (e.g., 31.76 vs. 0.32), so all value checks fail even though the underlying computation and the normality verdicts were correct.
927No held-out validation of the stated evaluation metric before submittingtaskda-code
Applies when
task -- the task specifies a scoring metric (e.g., log loss, RMSE, AUC) and the script trains models and writes a prediction file in one pass, with no train/validation split or cross-validation.
Pattern
The attempt fits one or several models on all labeled data, blends them with hand-picked weights, and writes probabilities/predictions directly to the submission file. Nothing in the script ever computes the competition metric on data with known labels, so mis-specified hyperparameters, mis-aligned class-to-column mapping (e.g., assigning predict_proba columns by an assumed label order rather than model.classes_), untuned/overconfident probabilities, or a truncated/incomplete output file all pass silently.
Detection procedure
  1. Read the task statement and note the exact metric and required output columns/rows.
  2. Search the script for any split, CV loop, or metric call (train_test_split, cross_val_*, log_loss, etc.) that is actually evaluated and printed on held-out labeled data; note if the imported metric is never called.
  3. Check whether prediction columns are mapped to classes using the fitted model's class attribute (or a verified ordering) and whether ensemble weights/hyperparameters were chosen by any measured score rather than asserted.
  4. Check the produced file for shape sanity: row count equals number of test ids, ids match the test set, no NaNs, probabilities in [0,1] and rows summing to ~1.
Discriminator
A real violation is the total absence of any measured out-of-sample metric value (and of shape/id sanity assertions) — the agent has no evidence its blend beats a trivial baseline or that column labels are right. It is not a violation if the script reports a validation/CV score for each model and uses it to select or weight them, even if the final refit uses all data; nor if a documented, verified class ordering is used alongside a reported score.
Consequence
The submission may score worse than a constant-prior baseline, have swapped class columns (inflating log loss dramatically), or be incomplete/misaligned with the test ids — the grader reports the file as wrong/missing with no diagnostic having been triggered during the run.
id d96e9657a7e9 · mined from da-code dacode-ml-competition-005@s18
raw text (what the judge reads)
### No held-out validation of the stated evaluation metric before submitting
- **Applies when**: `task` -- the task specifies a scoring metric (e.g., log loss, RMSE, AUC) and the script trains models and writes a prediction file in one pass, with no train/validation split or cross-validation.
- **Pattern**: The attempt fits one or several models on all labeled data, blends them with hand-picked weights, and writes probabilities/predictions directly to the submission file. Nothing in the script ever computes the competition metric on data with known labels, so mis-specified hyperparameters, mis-aligned class-to-column mapping (e.g., assigning `predict_proba` columns by an assumed label order rather than `model.classes_`), untuned/overconfident probabilities, or a truncated/incomplete output file all pass silently.
- **Detection procedure**:
  1. Read the task statement and note the exact metric and required output columns/rows.
  2. Search the script for any split, CV loop, or metric call (`train_test_split`, `cross_val_*`, `log_loss`, etc.) that is actually evaluated and printed on held-out labeled data; note if the imported metric is never called.
  3. Check whether prediction columns are mapped to classes using the fitted model's class attribute (or a verified ordering) and whether ensemble weights/hyperparameters were chosen by any measured score rather than asserted.
  4. Check the produced file for shape sanity: row count equals number of test ids, ids match the test set, no NaNs, probabilities in [0,1] and rows summing to ~1.
- **Discriminator**: A real violation is the total absence of any measured out-of-sample metric value (and of shape/id sanity assertions) — the agent has no evidence its blend beats a trivial baseline or that column labels are right. It is *not* a violation if the script reports a validation/CV score for each model and uses it to select or weight them, even if the final refit uses all data; nor if a documented, verified class ordering is used alongside a reported score.
- **Consequence**: The submission may score worse than a constant-prior baseline, have swapped class columns (inflating log loss dramatically), or be incomplete/misaligned with the test ids — the grader reports the file as wrong/missing with no diagnostic having been triggered during the run.
928Trains on an auxiliary source without verifying schema alignment or validating prediction qualitytaskda-code
Applies when
task -- the test/inference file is given but labels must come from separate raw source file(s), and the script builds features from those sources and applies them to the test frame.
Pattern
The script assumes the training source and the inference file share identical column names, dtypes and semantics, silently fills or drops anything that mismatches, and then writes predictions with no held-out evaluation and no comparison of the predicted distribution against the observed target distribution — so a systematically shrunken/degenerate prediction vector (e.g., mostly zeros or a saturated ceiling) is never noticed.
Detection procedure
  1. Read the task and list the exact columns of the inference file and of the label-bearing source file(s); note any naming/spelling/dtype differences (singular vs plural, alternate spellings, index columns, booleans stored as strings).
  2. In the script, check whether every feature name used for training is looked up in the test frame the same way, and whether a KeyError/fillna(0)/reindex path could quietly substitute constants for real features; also check whether test rows might already be present in the training source (duplicate IDs) or, conversely, come from a different population.
  3. Check for any train/validation split with a reported error metric, and for any post-prediction sanity check comparing summary statistics (mean, median, quantiles, share of zeros, max) of predictions with the training labels.
  4. Inspect the produced answer: if a large fraction of rows are identical (0) or clipped at a repeated value while the label distribution is broad, flag it.
Discriminator
A real violation is when identical-looking feature names are actually different or absent across files, or when no validation/sanity comparison exists and the output distribution is visibly degenerate; a look-alike that is fine is a script that explicitly asserts column presence/alignment (or renames columns) and reports a holdout score plus prediction-vs-label summary statistics that are of the same order of magnitude.
Consequence
The model effectively predicts from constants or misaligned features, the output collapses toward zero/low values, and the grader's accuracy/error check on the submitted file fails even though the file has the right shape and column name.
id d59542166445 · mined from da-code dacode-ml-regression-008@s18
raw text (what the judge reads)
### Trains on an auxiliary source without verifying schema alignment or validating prediction quality
- **Applies when**: `task` -- the test/inference file is given but labels must come from separate raw source file(s), and the script builds features from those sources and applies them to the test frame.
- **Pattern**: The script assumes the training source and the inference file share identical column names, dtypes and semantics, silently fills or drops anything that mismatches, and then writes predictions with no held-out evaluation and no comparison of the predicted distribution against the observed target distribution — so a systematically shrunken/degenerate prediction vector (e.g., mostly zeros or a saturated ceiling) is never noticed.
- **Detection procedure**:
  1. Read the task and list the exact columns of the inference file and of the label-bearing source file(s); note any naming/spelling/dtype differences (singular vs plural, alternate spellings, index columns, booleans stored as strings).
  2. In the script, check whether every feature name used for training is looked up in the test frame the same way, and whether a `KeyError`/`fillna(0)`/`reindex` path could quietly substitute constants for real features; also check whether test rows might already be present in the training source (duplicate IDs) or, conversely, come from a different population.
  3. Check for any train/validation split with a reported error metric, and for any post-prediction sanity check comparing summary statistics (mean, median, quantiles, share of zeros, max) of predictions with the training labels.
  4. Inspect the produced answer: if a large fraction of rows are identical (0) or clipped at a repeated value while the label distribution is broad, flag it.
- **Discriminator**: A real violation is when identical-looking feature names are actually different or absent across files, or when no validation/sanity comparison exists and the output distribution is visibly degenerate; a look-alike that is fine is a script that explicitly asserts column presence/alignment (or renames columns) and reports a holdout score plus prediction-vs-label summary statistics that are of the same order of magnitude.
- **Consequence**: The model effectively predicts from constants or misaligned features, the output collapses toward zero/low values, and the grader's accuracy/error check on the submitted file fails even though the file has the right shape and column name.
929Unscoped population and unchecked test assumptions in hypothesis testingtaskda-code
Applies when
task -- the task asks for a p-value/decision from a statistical comparison between two groups drawn from large raw historical/observational files.
Pattern
The script loads the full raw files, pools every row into each group, and immediately calls a default parametric test (e.g., an independent two-sample t-test) without (a) restricting to the subpopulation/time window the question is really about, or (b) checking whether the outcome distribution justifies a parametric mean test versus a rank-based/nonparametric or one-sided alternative — yielding an extreme p-value that is reported without any sanity discussion.
Detection procedure
  1. Read the task/README for any scoping cues (competition type, date range, category, exclusion rules) and for the direction of the hypothesis (one- vs two-sided); note whether the script applies any row filter at all.
  2. In the script, check whether the test's assumptions were examined (distribution shape, discreteness/skew of the outcome, unequal variance) or whether a default test was chosen blindly.
  3. Compare the group sample sizes actually used against the plausible size of the intended subpopulation; a p-value astronomically small (e.g., <1e-50) from tens of thousands of pooled rows is a red flag that the whole corpus was used.
  4. Verify the reported p-value and decision string are the ones the task requested, and would survive an alternative-test sensitivity check.
Discriminator
A real violation is when the task or data description implies a narrower population or a non-normal/ordinal outcome and the script does neither filtering nor assumption checking; it is fine if the task explicitly says to use all records and the outcome is plausibly suited to the chosen test (documented by the script).
Consequence
The p-value is computed on the wrong sample with the wrong test statistic, so the saved value (and possibly the reject/fail-to-reject decision) mismatches the expected file and the check fails.
id f74c9f82fa81 · mined from da-code dacode-data-sa-001@s18
raw text (what the judge reads)
### Unscoped population and unchecked test assumptions in hypothesis testing
- **Applies when**: `task` -- the task asks for a p-value/decision from a statistical comparison between two groups drawn from large raw historical/observational files.
- **Pattern**: The script loads the full raw files, pools every row into each group, and immediately calls a default parametric test (e.g., an independent two-sample t-test) without (a) restricting to the subpopulation/time window the question is really about, or (b) checking whether the outcome distribution justifies a parametric mean test versus a rank-based/nonparametric or one-sided alternative — yielding an extreme p-value that is reported without any sanity discussion.
- **Detection procedure**:
  1. Read the task/README for any scoping cues (competition type, date range, category, exclusion rules) and for the direction of the hypothesis (one- vs two-sided); note whether the script applies any row filter at all.
  2. In the script, check whether the test's assumptions were examined (distribution shape, discreteness/skew of the outcome, unequal variance) or whether a default test was chosen blindly.
  3. Compare the group sample sizes actually used against the plausible size of the intended subpopulation; a p-value astronomically small (e.g., <1e-50) from tens of thousands of pooled rows is a red flag that the whole corpus was used.
  4. Verify the reported p-value and decision string are the ones the task requested, and would survive an alternative-test sensitivity check.
- **Discriminator**: A real violation is when the task or data description implies a narrower population or a non-normal/ordinal outcome and the script does neither filtering nor assumption checking; it is fine if the task explicitly says to use all records and the outcome is plausibly suited to the chosen test (documented by the script).
- **Consequence**: The p-value is computed on the wrong sample with the wrong test statistic, so the saved value (and possibly the reject/fail-to-reject decision) mismatches the expected file and the check fails.
930Ignoring the provided output template when a task says "match the sample file exactly"taskda-code
Applies when
task -- The task supplies an example/sample output file (or an explicit schema) that the deliverable must replicate in structure and formatting.
Pattern
The scripts never read, print, or compare against the sample file; the agent guesses column names, column order, row ordering, and numeric formatting from memory or from its own intermediate DataFrame, and writes the result without any assertion that it conforms to the template (e.g., unrounded floats with long decimal tails, invented header names, arbitrary sort key).
Detection procedure
  1. In the task statement, note that an example output artifact is referenced and identify what it constrains (header text, column count/order, row order, rounding/units, index presence).
  2. Search the scripts for any load of or reference to that sample artifact; if absent, the conformance was never verified.
  3. Inspect the produced output: check whether headers/order were derived from the template or renamed ad hoc, and whether numeric columns are formatted consistently (fixed decimals) rather than raw floating-point residue like ...4.7799999999.
  4. Confirm no post-write validation step (shape, column equality, dtype, rounding assertion) exists.
Discriminator
A real violation is when nothing in the pipeline reads or asserts against the template and the output shows tell-tale ad-hoc formatting; it is not a violation if the script loads the template (or hard-codes its exact header/order/rounding with a visible comparison or assertion) and the output demonstrably matches, even if the values themselves are later disputed.
Consequence
The grader does an exact/tolerance comparison against the expected file and reports the result file as WRONG/MISSING even when the underlying aggregation logic is close or correct.
id 5c88542b8243 · mined from da-code dacode-dm-csv-011@s18
raw text (what the judge reads)
### Ignoring the provided output template when a task says "match the sample file exactly"
- **Applies when**: `task` -- The task supplies an example/sample output file (or an explicit schema) that the deliverable must replicate in structure and formatting.
- **Pattern**: The scripts never read, print, or compare against the sample file; the agent guesses column names, column order, row ordering, and numeric formatting from memory or from its own intermediate DataFrame, and writes the result without any assertion that it conforms to the template (e.g., unrounded floats with long decimal tails, invented header names, arbitrary sort key).
- **Detection procedure**:
  1. In the task statement, note that an example output artifact is referenced and identify what it constrains (header text, column count/order, row order, rounding/units, index presence).
  2. Search the scripts for any load of or reference to that sample artifact; if absent, the conformance was never verified.
  3. Inspect the produced output: check whether headers/order were derived from the template or renamed ad hoc, and whether numeric columns are formatted consistently (fixed decimals) rather than raw floating-point residue like `...4.7799999999`.
  4. Confirm no post-write validation step (shape, column equality, dtype, rounding assertion) exists.
- **Discriminator**: A real violation is when nothing in the pipeline reads or asserts against the template and the output shows tell-tale ad-hoc formatting; it is *not* a violation if the script loads the template (or hard-codes its exact header/order/rounding with a visible comparison or assertion) and the output demonstrably matches, even if the values themselves are later disputed.
- **Consequence**: The grader does an exact/tolerance comparison against the expected file and reports the result file as WRONG/MISSING even when the underlying aggregation logic is close or correct.
931Output feature columns don't match the feature vector actually used (categoricals silently dropped, transforms not reflected)taskda-code
Applies when
task -- the deliverable is a file whose columns are defined as "the i-th value of the feature vector" plus a model output (e.g., cluster/label/prediction), and the script builds that file from a hand-picked subset of raw columns.
Pattern
The script silently discards usable attributes (categorical, date, or text fields) instead of encoding/deriving them, and/or writes the pre-transformation raw values into the Feature_i columns while the model was fit on a different (scaled/encoded/reduced) matrix — so the exported feature vector is not the one that produced the labels and omits information the task expected to be used.
Detection procedure
  1. From the task/README, list every attribute available and note which are identifiers/constants (legitimately droppable) versus informative but non-numeric (require encoding, not dropping).
  2. In the script, find the exact array passed to fit/fit_predict and the DataFrame written to the output file; check whether their column sets and values are the same object/lineage.
  3. Flag if any informative non-numeric attribute is dropped with no encoding step, or if the exported columns are raw values while fitting used a transformed matrix, or if the number of Feature_i columns differs from the model's n_features_in_.
  4. Check the answer text for a justification of each exclusion; unexplained exclusions ("removed non-numeric variables") are a red flag.
Discriminator
Dropping true IDs, constant columns, or a column the task explicitly excludes is fine; so is exporting raw values if no transformation was applied before fitting or the task explicitly asks for original values. The violation is dropping informative attributes purely because they were inconvenient dtypes, or an exported matrix that cannot reproduce the reported labels.
Consequence
The saved file's feature block disagrees with the reference feature vector (wrong column count/values), so a file-level comparison of Feature_i columns fails and the whole check scores 0 even if the clustering itself is reasonable.
id 3aa97b03d595 · mined from da-code dacode-ml-cluster-014@s18
raw text (what the judge reads)
### Output feature columns don't match the feature vector actually used (categoricals silently dropped, transforms not reflected)
- **Applies when**: `task` -- the deliverable is a file whose columns are defined as "the i-th value of the feature vector" plus a model output (e.g., cluster/label/prediction), and the script builds that file from a hand-picked subset of raw columns.
- **Pattern**: The script silently discards usable attributes (categorical, date, or text fields) instead of encoding/deriving them, and/or writes the pre-transformation raw values into the `Feature_i` columns while the model was fit on a different (scaled/encoded/reduced) matrix — so the exported feature vector is not the one that produced the labels and omits information the task expected to be used.
- **Detection procedure**:
  1. From the task/README, list every attribute available and note which are identifiers/constants (legitimately droppable) versus informative but non-numeric (require encoding, not dropping).
  2. In the script, find the exact array passed to `fit`/`fit_predict` and the DataFrame written to the output file; check whether their column sets and values are the same object/lineage.
  3. Flag if any informative non-numeric attribute is dropped with no encoding step, or if the exported columns are raw values while fitting used a transformed matrix, or if the number of `Feature_i` columns differs from the model's `n_features_in_`.
  4. Check the answer text for a justification of each exclusion; unexplained exclusions ("removed non-numeric variables") are a red flag.
- **Discriminator**: Dropping true IDs, constant columns, or a column the task explicitly excludes is fine; so is exporting raw values *if* no transformation was applied before fitting or the task explicitly asks for original values. The violation is dropping informative attributes purely because they were inconvenient dtypes, or an exported matrix that cannot reproduce the reported labels.
- **Consequence**: The saved file's feature block disagrees with the reference feature vector (wrong column count/values), so a file-level comparison of `Feature_i` columns fails and the whole check scores 0 even if the clustering itself is reasonable.
932Unjustified rounding/type-coercion of continuous model outputtaskda-code
Applies when
task -- a regression-style prediction must be written to a submission file, and the script post-processes raw model output (rounding, casting to int, clipping) before saving.
Pattern
The agent assumes the target must be integral (or otherwise discretized) because training labels happen to be integers, and applies np.round(...).astype(int) (or similar) to predictions, without checking the provided sample submission's value format or the scoring metric; the discretization is never validated against a held-out score computed the same way as the final submission.
Detection procedure
  1. Read the task/README and the provided sample submission file usage in the scripts: is there any stated requirement that predictions be integers/discrete, and does the script ever inspect the sample file's dtypes or example values?
  2. Search the modeling script for post-prediction transformations (round, astype(int), clip, thresholding) applied only to test predictions.
  3. Check whether validation metrics were recomputed after the same transformation; if validation used raw floats but the submission is rounded, the reported quality does not describe the submitted file.
  4. Inspect the answer file: if every value is an integer while a continuous regression metric is implied, flag it.
Discriminator
A real violation is transforming test predictions in a way that is (a) not required by the task/sample format and (b) not validated by comparing metrics with and without the transform. It is fine if the task or sample submission explicitly demands integers/classes, or if the agent measured that the rounded predictions score at least as well on a held-out split.
Consequence
The submitted file loses sub-unit precision (and any asymmetric transform bias), so the scored error is systematically worse than the validation number reported, and a format/value comparison against the expected submission fails.
id df9eb3b45314 · mined from da-code dacode-ml-competition-009@s18
raw text (what the judge reads)
### Unjustified rounding/type-coercion of continuous model output
- **Applies when**: `task` -- a regression-style prediction must be written to a submission file, and the script post-processes raw model output (rounding, casting to int, clipping) before saving.
- **Pattern**: The agent assumes the target must be integral (or otherwise discretized) because training labels happen to be integers, and applies `np.round(...).astype(int)` (or similar) to predictions, without checking the provided sample submission's value format or the scoring metric; the discretization is never validated against a held-out score computed the same way as the final submission.
- **Detection procedure**:
  1. Read the task/README and the provided sample submission file usage in the scripts: is there any stated requirement that predictions be integers/discrete, and does the script ever inspect the sample file's dtypes or example values?
  2. Search the modeling script for post-prediction transformations (`round`, `astype(int)`, `clip`, thresholding) applied only to test predictions.
  3. Check whether validation metrics were recomputed after the same transformation; if validation used raw floats but the submission is rounded, the reported quality does not describe the submitted file.
  4. Inspect the answer file: if every value is an integer while a continuous regression metric is implied, flag it.
- **Discriminator**: A real violation is transforming test predictions in a way that is (a) not required by the task/sample format and (b) not validated by comparing metrics with and without the transform. It is fine if the task or sample submission explicitly demands integers/classes, or if the agent measured that the rounded predictions score at least as well on a held-out split.
- **Consequence**: The submitted file loses sub-unit precision (and any asymmetric transform bias), so the scored error is systematically worse than the validation number reported, and a format/value comparison against the expected submission fails.
933Silently dropping most features by relying on inferred dtypes instead of cleaning numeric-like text columnstaskda-code
Applies when
task -- The task asks to model/cluster/analyze a tabular dataset whose documented fields are mostly quantitative, and the script builds its feature matrix with an automatic dtype filter (e.g., select_dtypes(number)) or a hardcoded short column list.
Pattern
Many quantitative columns are loaded as strings because of formatting characters (currency symbols, thousands separators, percent signs, units, placeholder tokens), so the dtype filter silently keeps only a small subset; the agent also keeps meaningless identifier-like numeric columns (codes, IDs, coordinates) and never reconciles the retained feature count against the documented set of variables.
Detection procedure
  1. Read the task/README and count how many fields are described as quantitative indicators.
  2. In the script, find how the feature matrix is chosen; check whether there is any parsing/cleaning step that strips non-numeric characters and coerces text columns to numbers before the selection.
  3. Compare the number/identity of features actually used (printed in the script output or the answer) to the documented quantitative fields; also check whether pure identifier/code columns were included as features.
  4. Confirm the answer/output file's feature-column count matches the intended feature space rather than a small accidental subset.
Discriminator
A real violation is when quantitative columns were excluded implicitly by dtype inference or included as junk identifiers with no stated justification. It is fine if the script explicitly parses/cleans string-formatted numerics and then deliberately documents which columns are dropped and why (e.g., true categorical names, codes).
Consequence
The output file has far fewer Feature_i columns than expected and cluster labels derived from a distorted, partly meaningless feature space, so file-level comparison against the reference result fails on both shape and content.
id 739b23415aad · mined from da-code dacode-ml-cluster-009@s18
raw text (what the judge reads)
### Silently dropping most features by relying on inferred dtypes instead of cleaning numeric-like text columns
- **Applies when**: `task` -- The task asks to model/cluster/analyze a tabular dataset whose documented fields are mostly quantitative, and the script builds its feature matrix with an automatic dtype filter (e.g., `select_dtypes(number)`) or a hardcoded short column list.
- **Pattern**: Many quantitative columns are loaded as strings because of formatting characters (currency symbols, thousands separators, percent signs, units, placeholder tokens), so the dtype filter silently keeps only a small subset; the agent also keeps meaningless identifier-like numeric columns (codes, IDs, coordinates) and never reconciles the retained feature count against the documented set of variables.
- **Detection procedure**:
  1. Read the task/README and count how many fields are described as quantitative indicators.
  2. In the script, find how the feature matrix is chosen; check whether there is any parsing/cleaning step that strips non-numeric characters and coerces text columns to numbers before the selection.
  3. Compare the number/identity of features actually used (printed in the script output or the answer) to the documented quantitative fields; also check whether pure identifier/code columns were included as features.
  4. Confirm the answer/output file's feature-column count matches the intended feature space rather than a small accidental subset.
- **Discriminator**: A real violation is when quantitative columns were excluded implicitly by dtype inference or included as junk identifiers with no stated justification. It is fine if the script explicitly parses/cleans string-formatted numerics and then deliberately documents which columns are dropped and why (e.g., true categorical names, codes).
- **Consequence**: The output file has far fewer `Feature_i` columns than expected and cluster labels derived from a distorted, partly meaningless feature space, so file-level comparison against the reference result fails on both shape and content.
934Degenerate clustering accepted because outlier-driven separation inflates internal metricstaskda-code
Applies when
task -- an unsupervised segmentation/grouping task where the agent picks the number of groups via internal validity scores and must emit per-record group labels.
Pattern
The agent builds heavily right-skewed aggregate features, standardizes without taming skew/outliers, and then selects the "best" solution by silhouette (or similar), which is maximized by a split that isolates a handful of extreme records; nearly all records land in one group, and the reported group sizes are internally inconsistent (they don't sum to the record count or fewer groups are listed than were requested).
Detection procedure
  1. From the task, note the requested output (file name, exact column names, one row per unit) and that a useful segmentation is expected.
  2. In the scripts, check whether skewed monetary/count features are log/robust-transformed or outlier-capped before scaling, and whether the model-selection criterion is guarded against trivial solutions (e.g., minimum cluster share, inspecting size distribution, not just the top internal score).
  3. In the answer, recompute the sanity checks: do the reported cluster sizes sum to the stated number of units, is the number of listed clusters equal to the chosen k, and is any cluster >90–95% of the data?
  4. Verify the deliverable itself is described as written with the exact required columns/rows, not just summarized in prose.
Discriminator
A genuinely imbalanced but valid segmentation is fine if the agent explicitly checked skew/outlier handling, compared alternatives, and its reported counts are complete and consistent; the violation is an unexamined near-single-cluster result with suspiciously high separation scores and arithmetic that doesn't reconcile with the record count or chosen k.
Consequence
The saved label file contains an essentially constant label column (and possibly wrong shape/columns), so the grader's check on the expected output file fails.
id f96f6fa11481 · mined from da-code dacode-ml-cluster-016@s18
raw text (what the judge reads)
### Degenerate clustering accepted because outlier-driven separation inflates internal metrics
- **Applies when**: `task` -- an unsupervised segmentation/grouping task where the agent picks the number of groups via internal validity scores and must emit per-record group labels.
- **Pattern**: The agent builds heavily right-skewed aggregate features, standardizes without taming skew/outliers, and then selects the "best" solution by silhouette (or similar), which is maximized by a split that isolates a handful of extreme records; nearly all records land in one group, and the reported group sizes are internally inconsistent (they don't sum to the record count or fewer groups are listed than were requested).
- **Detection procedure**:
  1. From the task, note the requested output (file name, exact column names, one row per unit) and that a *useful* segmentation is expected.
  2. In the scripts, check whether skewed monetary/count features are log/robust-transformed or outlier-capped before scaling, and whether the model-selection criterion is guarded against trivial solutions (e.g., minimum cluster share, inspecting size distribution, not just the top internal score).
  3. In the answer, recompute the sanity checks: do the reported cluster sizes sum to the stated number of units, is the number of listed clusters equal to the chosen k, and is any cluster >90–95% of the data?
  4. Verify the deliverable itself is described as written with the exact required columns/rows, not just summarized in prose.
- **Discriminator**: A genuinely imbalanced but valid segmentation is fine if the agent explicitly checked skew/outlier handling, compared alternatives, and its reported counts are complete and consistent; the violation is an unexamined near-single-cluster result with suspiciously high separation scores and arithmetic that doesn't reconcile with the record count or chosen k.
- **Consequence**: The saved label file contains an essentially constant label column (and possibly wrong shape/columns), so the grader's check on the expected output file fails.
935Prediction output not verified to match the required file, row count, and row order of the test settaskda-code
Applies when
task -- the task asks for per-row predictions/outputs written to a named file with a specified column name, produced from a given test/holdout input.
Pattern
The agent produces a column of plausible-looking numbers but never asserts that the written artifact exists at the required path/name, has exactly one row per input row in the original input order, and uses the exact requested column header; the values are instead emitted inline or truncated/aggregated, so the deliverable silently has the wrong shape or is missing.
Detection procedure
  1. From the task statement, extract the required output filename, column name(s), and the expected number of rows (i.e., the row count of the provided test/holdout file).
  2. In the scripts, find the line that writes the output; check it targets that exact path/header, writes with index=False-style formatting, and that predictions were generated from the full test frame without dropping/filtering/deduplicating rows or reordering (e.g., after dropna, groupby, merges, or resampling).
  3. Look for an explicit sanity check in the script or answer (len(pred) == len(test), shape print, head of the saved file re-read from disk); if absent, treat the shape as unverified.
  4. Compare the count of values actually shown/produced in the answer against the test row count and value range plausibility; a mismatch or an obviously small/rounded count is a violation.
Discriminator
A real violation is a missing/misnamed file, a row count differing from the test set, reordered rows, or a wrong/absent header. It is not a violation if the file is correctly written and verified and the agent merely displays a truncated preview of a full-length file, nor if predictions are merely inaccurate while shape, name, and order are correct.
Consequence
The grader cannot align predictions with ground truth — the expected result file is reported WRONG/MISSING and the attempt scores 0 regardless of model quality.
id ecab3727dc04 · mined from da-code dacode-ml-regression-002@s18
raw text (what the judge reads)
### Prediction output not verified to match the required file, row count, and row order of the test set
- **Applies when**: `task` -- the task asks for per-row predictions/outputs written to a named file with a specified column name, produced from a given test/holdout input.
- **Pattern**: The agent produces a column of plausible-looking numbers but never asserts that the written artifact exists at the required path/name, has exactly one row per input row in the original input order, and uses the exact requested column header; the values are instead emitted inline or truncated/aggregated, so the deliverable silently has the wrong shape or is missing.
- **Detection procedure**:
  1. From the task statement, extract the required output filename, column name(s), and the expected number of rows (i.e., the row count of the provided test/holdout file).
  2. In the scripts, find the line that writes the output; check it targets that exact path/header, writes with `index=False`-style formatting, and that predictions were generated from the full test frame without dropping/filtering/deduplicating rows or reordering (e.g., after `dropna`, `groupby`, merges, or resampling).
  3. Look for an explicit sanity check in the script or answer (`len(pred) == len(test)`, shape print, head of the saved file re-read from disk); if absent, treat the shape as unverified.
  4. Compare the count of values actually shown/produced in the answer against the test row count and value range plausibility; a mismatch or an obviously small/rounded count is a violation.
- **Discriminator**: A real violation is a missing/misnamed file, a row count differing from the test set, reordered rows, or a wrong/absent header. It is *not* a violation if the file is correctly written and verified and the agent merely displays a truncated preview of a full-length file, nor if predictions are merely inaccurate while shape, name, and order are correct.
- **Consequence**: The grader cannot align predictions with ground truth — the expected result file is reported WRONG/MISSING and the attempt scores 0 regardless of model quality.
936Referenced specification file never opened, or its categories silently replaced by the data's native valuestaskda-code
Applies when
task -- the task points to an external document (README, spec, config, codebook) that defines how values must be binned, grouped, filtered, or labeled before the requested output is produced.
Pattern
The attempt never loads/quotes the spec, and instead reports the raw levels already present in the source column (or its own invented bins), asserting "grouped as specified"; it also skips the auxiliary artifacts/records (saved script, serialized data/plot objects) that would let anyone verify the mapping.
Detection procedure
  1. From the task, list every referenced spec file and every required output artifact/format.
  2. In the scripts/transcript, look for an explicit read of the spec and an explicit mapping structure (dict, bin edges, pd.cut) derived from it; absence of both is a violation.
  3. Compare the reported category labels/counts to the distinct raw values of the source column — if they are identical in number and naming, no grouping was actually applied.
  4. Sanity-check totals and artifact list: does the summed count match the documented respondent count (watch for header/metadata rows inflating it by one), and does every requested output file exist?
Discriminator
A fine attempt may end up with labels resembling the raw levels if it shows the spec-derived mapping and any required merges/renames; a violation is when the mapping is asserted but never constructed, or the counts/labels demonstrably equal the ungrouped value counts.
Consequence
Graders comparing the saved arrays/plot data against the spec-defined groups find mismatched bin labels, counts, and off-by-one totals, and missing required artifacts, so every check fails despite a confident "task completed" report.
id 13868ccfdc01 · mined from da-code dacode-plot-bar-005@s18
raw text (what the judge reads)
### Referenced specification file never opened, or its categories silently replaced by the data's native values
- **Applies when**: `task` -- the task points to an external document (README, spec, config, codebook) that defines how values must be binned, grouped, filtered, or labeled before the requested output is produced.
- **Pattern**: The attempt never loads/quotes the spec, and instead reports the raw levels already present in the source column (or its own invented bins), asserting "grouped as specified"; it also skips the auxiliary artifacts/records (saved script, serialized data/plot objects) that would let anyone verify the mapping.
- **Detection procedure**:
  1. From the task, list every referenced spec file and every required output artifact/format.
  2. In the scripts/transcript, look for an explicit read of the spec and an explicit mapping structure (dict, bin edges, `pd.cut`) derived from it; absence of both is a violation.
  3. Compare the reported category labels/counts to the distinct raw values of the source column — if they are identical in number and naming, no grouping was actually applied.
  4. Sanity-check totals and artifact list: does the summed count match the documented respondent count (watch for header/metadata rows inflating it by one), and does every requested output file exist?
- **Discriminator**: A fine attempt may end up with labels resembling the raw levels *if* it shows the spec-derived mapping and any required merges/renames; a violation is when the mapping is asserted but never constructed, or the counts/labels demonstrably equal the ungrouped value counts.
- **Consequence**: Graders comparing the saved arrays/plot data against the spec-defined groups find mismatched bin labels, counts, and off-by-one totals, and missing required artifacts, so every check fails despite a confident "task completed" report.
937Answer not persisted to the required artifact in the exact requested schemataskda-code
Applies when
task -- the task specifies an output format (a template with particular key names and value containers) and/or the harness expects a named result file to be written by the scripts.
Pattern
The agent computes a plausible value and reports it only in chat prose, and/or fills the requested template with a different value type or key spelling than shown (e.g., bare scalar/string where the template shows a list, renamed/extra keys, rounding or units not matching), so no gradeable artifact matching the required schema exists.
Detection procedure
  1. Read the task statement and extract the exact expected deliverable: the file name(s) to be produced, the literal key names, and the value containers/types shown in the template (list vs scalar), plus any rounding/unit instructions.
  2. Scan the scripts for a write step (json.dump, to_json, to_csv, etc.) that emits exactly that filename in the expected working directory; if no script or no write step exists, flag immediately.
  3. Compare the structure actually written/reported field-by-field against the template: key names identical, values wrapped as shown, numeric precision/units as instructed.
  4. Confirm the reported value is the final requested quantity (not an intermediate or a differently-defined statistic) and that any prescribed preprocessing step (e.g., the specified imputation) was actually executed before it.
Discriminator
A real violation is a missing output file or a structural/type/key deviation from the given template; a look-alike that is fine is a correct file written with the template's exact keys and containers where only cosmetic whitespace, key order, or an equivalent numeric representation (e.g., 82.6 vs 82.60) differs.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even if the underlying computation happens to be right.
id 054bbd03274e · mined from da-code dacode-di-text-002@s18
raw text (what the judge reads)
### Answer not persisted to the required artifact in the exact requested schema
- **Applies when**: `task` -- the task specifies an output format (a template with particular key names and value containers) and/or the harness expects a named result file to be written by the scripts.
- **Pattern**: The agent computes a plausible value and reports it only in chat prose, and/or fills the requested template with a different value type or key spelling than shown (e.g., bare scalar/string where the template shows a list, renamed/extra keys, rounding or units not matching), so no gradeable artifact matching the required schema exists.
- **Detection procedure**:
  1. Read the task statement and extract the exact expected deliverable: the file name(s) to be produced, the literal key names, and the value containers/types shown in the template (list vs scalar), plus any rounding/unit instructions.
  2. Scan the scripts for a write step (`json.dump`, `to_json`, `to_csv`, etc.) that emits exactly that filename in the expected working directory; if no script or no write step exists, flag immediately.
  3. Compare the structure actually written/reported field-by-field against the template: key names identical, values wrapped as shown, numeric precision/units as instructed.
  4. Confirm the reported value is the final requested quantity (not an intermediate or a differently-defined statistic) and that any prescribed preprocessing step (e.g., the specified imputation) was actually executed before it.
- **Discriminator**: A real violation is a missing output file or a structural/type/key deviation from the given template; a look-alike that is fine is a correct file written with the template's exact keys and containers where only cosmetic whitespace, key order, or an equivalent numeric representation (e.g., 82.6 vs 82.60) differs.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even if the underlying computation happens to be right.
938Unvalidated raw inputs fed straight into a cumulative/aggregate computationtaskda-code
Applies when
task -- a script loads a raw tabular file and immediately combines its numeric columns (weighted sums, compounding, running products/sums, group aggregates) into a final deliverable file.
Pattern
The attempt never inspects the loaded columns for missing/non-numeric entries, unit/scale conventions (percent vs. fraction, levels vs. changes), row ordering, or extra/unexpected rows and columns; it just applies the formula and writes the output, and the reported numbers are accepted without any independent plausibility check. It also invents output column names/format from assumption rather than from a stated or discoverable specification.
Detection procedure
  1. Read the task for any stated output schema, units, ordering, or rounding requirement, and note what the input is supposed to contain.
  2. In the script, check whether there is any explicit handling/reporting of NaNs, dtypes, duplicate or out-of-range rows, column-name mismatches, and value scale before the aggregation step (printing head()/dtypes is not handling).
  3. Check whether the final numbers are cross-validated by an independent route (e.g., recompute the end value a different way, compare per-column simple aggregates, verify row count and value ranges match expectations).
  4. Check whether output column names/order/date formatting are justified by the task or an example, rather than chosen freely by the agent.
Discriminator
A real violation is when a plausible input defect (a missing value, a percent-scaled column, unsorted rows) would silently and materially change the written numbers and nothing in the script would catch it; it is fine if the script asserts/cleans these properties, or if it demonstrably verified the data is complete, numeric, correctly scaled and ordered before aggregating.
Consequence
The saved file compares unequal to the reference (NaN-propagated or scale-shifted columns, wrong headers/ordering), so the file check fails even though the reported summary numbers look internally consistent.
id e5cd4c55c02b · mined from da-code dacode-dm-csv-050@s18
raw text (what the judge reads)
### Unvalidated raw inputs fed straight into a cumulative/aggregate computation
- **Applies when**: `task` -- a script loads a raw tabular file and immediately combines its numeric columns (weighted sums, compounding, running products/sums, group aggregates) into a final deliverable file.
- **Pattern**: The attempt never inspects the loaded columns for missing/non-numeric entries, unit/scale conventions (percent vs. fraction, levels vs. changes), row ordering, or extra/unexpected rows and columns; it just applies the formula and writes the output, and the reported numbers are accepted without any independent plausibility check. It also invents output column names/format from assumption rather than from a stated or discoverable specification.
- **Detection procedure**:
  1. Read the task for any stated output schema, units, ordering, or rounding requirement, and note what the input is supposed to contain.
  2. In the script, check whether there is any explicit handling/reporting of NaNs, dtypes, duplicate or out-of-range rows, column-name mismatches, and value scale before the aggregation step (printing `head()`/`dtypes` is not handling).
  3. Check whether the final numbers are cross-validated by an independent route (e.g., recompute the end value a different way, compare per-column simple aggregates, verify row count and value ranges match expectations).
  4. Check whether output column names/order/date formatting are justified by the task or an example, rather than chosen freely by the agent.
- **Discriminator**: A real violation is when a plausible input defect (a missing value, a percent-scaled column, unsorted rows) would silently and materially change the written numbers and nothing in the script would catch it; it is fine if the script asserts/cleans these properties, or if it demonstrably verified the data is complete, numeric, correctly scaled and ordered before aggregating.
- **Consequence**: The saved file compares unequal to the reference (NaN-propagated or scale-shifted columns, wrong headers/ordering), so the file check fails even though the reported summary numbers look internally consistent.
939Normality (or other statistical) test run on an unvetted data scope, with the required intermediate value never reported or sanity-checkedtaskinfiagent-dabench
Applies when
task -- the task asks to run a specific statistical test on a single column/variable with a stated decision rule, and to report the test statistic/p-value plus descriptive moments.
Pattern
The attempt loads the file and pipes the raw column straight into the test function without first establishing what the analysis population is — no check for missing/placeholder values (NaN, 0, -999, sentinel codes), no check for duplicate/aggregated rows, no check that the dtype is numeric rather than string-coerced, and no check of the resulting n. It then reports only the final verdict and moments, omitting the p-value the task explicitly asked for, so the borderline decision (test is highly sensitive to n and to a handful of contaminating extreme values) is never exposed or defended.
Detection procedure
  1. Read the task and list every quantity the answer must contain (here: the decision, the p-value, and the moments) and every constraint on the data scope (filters, alpha, rounding).
  2. In the scripts, find the exact expression passed to the test/statistic call and trace it back to the raw load: is there an explicit dropna/dtype coercion/filter/de-duplication, and is len() (or shape) of the tested array printed anywhere?
  3. Check whether the script prints the p-value and n, and whether the reported answer includes the p-value; also check whether skewness/kurtosis are computed on the same cleaned array as the test.
  4. Sanity-check consistency: if the reported moments imply strong departure from normality (|skew| > 1, excess kurtosis > 3) while a "normal" verdict is expected, or vice versa, demand the printed p-value and n before accepting either verdict.
Discriminator
A real violation is when the tested array's composition is never established or printed (no n, no missing-value handling, no dtype check) and the requested p-value is absent — the answer cannot be audited. It is not a violation if the script explicitly documents the cleaning decision and prints n and the p-value, even if it ends up testing the full raw column, since that scope choice is then visible and defensible.
Consequence
The decision flips on the alpha threshold because the test was run on a contaminated or wrongly sized sample, so the boolean verdict field fails the grader (and the moments may be off for the same reason), while the missing p-value leaves nothing to diagnose the discrepancy.
id 31b9330823d9 · mined from infiagent-dabench dabench-298@s18
raw text (what the judge reads)
### Normality (or other statistical) test run on an unvetted data scope, with the required intermediate value never reported or sanity-checked
- **Applies when**: `task` -- the task asks to run a specific statistical test on a single column/variable with a stated decision rule, and to report the test statistic/p-value plus descriptive moments.
- **Pattern**: The attempt loads the file and pipes the raw column straight into the test function without first establishing what the analysis population is — no check for missing/placeholder values (NaN, 0, -999, sentinel codes), no check for duplicate/aggregated rows, no check that the dtype is numeric rather than string-coerced, and no check of the resulting n. It then reports only the final verdict and moments, omitting the p-value the task explicitly asked for, so the borderline decision (test is highly sensitive to n and to a handful of contaminating extreme values) is never exposed or defended.
- **Detection procedure**:
  1. Read the task and list every quantity the answer must contain (here: the decision, the p-value, and the moments) and every constraint on the data scope (filters, alpha, rounding).
  2. In the scripts, find the exact expression passed to the test/statistic call and trace it back to the raw load: is there an explicit dropna/dtype coercion/filter/de-duplication, and is `len()` (or shape) of the tested array printed anywhere?
  3. Check whether the script prints the p-value and n, and whether the reported answer includes the p-value; also check whether skewness/kurtosis are computed on the *same* cleaned array as the test.
  4. Sanity-check consistency: if the reported moments imply strong departure from normality (|skew| > 1, excess kurtosis > 3) while a "normal" verdict is expected, or vice versa, demand the printed p-value and n before accepting either verdict.
- **Discriminator**: A real violation is when the tested array's composition is never established or printed (no n, no missing-value handling, no dtype check) and the requested p-value is absent — the answer cannot be audited. It is *not* a violation if the script explicitly documents the cleaning decision and prints n and the p-value, even if it ends up testing the full raw column, since that scope choice is then visible and defensible.
- **Consequence**: The decision flips on the alpha threshold because the test was run on a contaminated or wrongly sized sample, so the boolean verdict field fails the grader (and the moments may be off for the same reason), while the missing p-value leaves nothing to diagnose the discrepancy.
940Deliverable file not verified against the template at the required pathtaskda-code
Applies when
task -- the task requires writing a prediction/result file at a specified location and in the format of a provided sample/template file.
Pattern
The agent reports that the output file was produced (often only a path string) without any reproducible script that reads the template, builds the output with the same columns/ID ordering/row count, writes it to the exact expected path, and re-reads it to confirm; the file may be missing, at a different path, or structurally wrong (missing ID column, renamed target column, wrong number of rows, NaNs, out-of-range values).
Detection procedure
  1. From the task, note the exact required filename, location, and the template's column names, dtypes, ID set, and row count.
  2. In the scripts, locate the write call: check the path matches the required one and that the frame written is built by aligning to the template's IDs/columns rather than ad-hoc.
  3. Check for a post-write verification step: re-read the written file and assert shape, column names, ID match with the template, no missing values, and prediction range plausible for the target/metric.
  4. Check the answer: it must be backed by such verification output, not merely a claim or a path; if no script was saved at all, the deliverable is unverifiable.
Discriminator
A real violation is absence of any evidence that a correctly-formatted file exists at the required path (no saved script, no read-back assertions, path differing from the specified one). It is fine if the script writes to the exact expected path and prints/asserts the shape, columns, ID alignment, and value ranges, even if the modeling approach is simple.
Consequence
The grader looks for the expected result file and finds it missing or structurally mismatched, scoring 0 regardless of model quality.
id 92166fd45e44 · mined from da-code dacode-ml-competition-008@s18
raw text (what the judge reads)
### Deliverable file not verified against the template at the required path
- **Applies when**: `task` -- the task requires writing a prediction/result file at a specified location and in the format of a provided sample/template file.
- **Pattern**: The agent reports that the output file was produced (often only a path string) without any reproducible script that reads the template, builds the output with the same columns/ID ordering/row count, writes it to the exact expected path, and re-reads it to confirm; the file may be missing, at a different path, or structurally wrong (missing ID column, renamed target column, wrong number of rows, NaNs, out-of-range values).
- **Detection procedure**:
  1. From the task, note the exact required filename, location, and the template's column names, dtypes, ID set, and row count.
  2. In the scripts, locate the write call: check the path matches the required one and that the frame written is built by aligning to the template's IDs/columns rather than ad-hoc.
  3. Check for a post-write verification step: re-read the written file and assert shape, column names, ID match with the template, no missing values, and prediction range plausible for the target/metric.
  4. Check the answer: it must be backed by such verification output, not merely a claim or a path; if no script was saved at all, the deliverable is unverifiable.
- **Discriminator**: A real violation is absence of any evidence that a correctly-formatted file exists at the required path (no saved script, no read-back assertions, path differing from the specified one). It is fine if the script writes to the exact expected path and prints/asserts the shape, columns, ID alignment, and value ranges, even if the modeling approach is simple.
- **Consequence**: The grader looks for the expected result file and finds it missing or structurally mismatched, scoring 0 regardless of model quality.
941Silent fallback to synthetic/substitute data instead of the provided inputtaskda-code
Applies when
task -- the task supplies (or presumes) a specific input dataset at a known location, and the scripts must locate and load it before computing the requested artifacts.
Pattern
The script guesses at file paths, and when none match it generates random/synthetic records (or downloads an unrelated copy from the internet) and continues the full pipeline on that stand-in, producing plausible-looking numbers and plots that are unrelated to the real data; the agent then reports these fabricated results as the answer without flagging the substitution.
Detection procedure
  1. Read the task for the canonical input location/format and the exact list of deliverables it asks for.
  2. Scan the loading code for if os.path.exists(...) chains, try/except around downloads, np.random/synthetic DataFrame construction, or any branch that proceeds when the real file is absent — and check whether the script ever verifies which branch executed or hard-fails otherwise.
  3. Check that the script actually parses the real column semantics (e.g., mixed-type or string-encoded fields) rather than assuming the clean types it invented for the stand-in.
  4. Compare the reported numbers and produced files against the task's required outputs: are all requested artifacts written, and are row/category counts consistent with the real dataset's known size?
Discriminator
A robust attempt may search several candidate paths, but it asserts that a real file was found (printing the resolved path and shape) and aborts with an error if not; a violation is code that silently manufactures data or falls back to an external/unverified source and reports its outputs as the answer.
Consequence
Every value-bearing deliverable (arrays, JSON summaries, figure heights) is computed from data the grader never saw, so all output-comparison checks fail even though the script "ran successfully".
id e66c7ac30bb4 · mined from da-code dacode-plot-bar-007@s18
raw text (what the judge reads)
### Silent fallback to synthetic/substitute data instead of the provided input
- **Applies when**: `task` -- the task supplies (or presumes) a specific input dataset at a known location, and the scripts must locate and load it before computing the requested artifacts.
- **Pattern**: The script guesses at file paths, and when none match it generates random/synthetic records (or downloads an unrelated copy from the internet) and continues the full pipeline on that stand-in, producing plausible-looking numbers and plots that are unrelated to the real data; the agent then reports these fabricated results as the answer without flagging the substitution.
- **Detection procedure**:
  1. Read the task for the canonical input location/format and the exact list of deliverables it asks for.
  2. Scan the loading code for `if os.path.exists(...)` chains, `try/except` around downloads, `np.random`/synthetic DataFrame construction, or any branch that proceeds when the real file is absent — and check whether the script ever verifies which branch executed or hard-fails otherwise.
  3. Check that the script actually parses the real column semantics (e.g., mixed-type or string-encoded fields) rather than assuming the clean types it invented for the stand-in.
  4. Compare the reported numbers and produced files against the task's required outputs: are all requested artifacts written, and are row/category counts consistent with the real dataset's known size?
- **Discriminator**: A robust attempt may search several candidate paths, but it asserts that a real file was found (printing the resolved path and shape) and aborts with an error if not; a violation is code that *silently manufactures* data or falls back to an external/unverified source and reports its outputs as the answer.
- **Consequence**: Every value-bearing deliverable (arrays, JSON summaries, figure heights) is computed from data the grader never saw, so all output-comparison checks fail even though the script "ran successfully".
942Unverifiable answer: no saved script and no required output artifacttaskda-code
Applies when
task -- the task asks for a specific result computed from a provided dataset (with stated preprocessing/ordering rules) and expects it saved in a named output file/format.
Pattern
The agent reports a plausible-looking answer directly in chat (apparently from prior knowledge or an unsaved ad-hoc run), leaving no script that loads the data, performs the stated preprocessing, computes the requested ranking/statistic, and writes the required artifact — so neither the preprocessing rule nor the ordering constraint can be verified or reproduced.
Detection procedure
  1. From the task, list the required deliverables: the output file name/format, the stated preprocessing step, and any ordering/rounding/filtering constraint.
  2. Check whether any saved script exists that reads the input data and ends by writing that exact artifact with that exact key/schema.
  3. If a script exists, confirm each stated constraint is implemented in code (imputation applied before ranking; both lists explicitly sorted in the demanded direction) rather than assumed; if no script exists, flag immediately.
  4. Cross-check the reported values against the data (spot-check a couple of extremes and the sort direction of every list) — inability to do so from artifacts is itself the violation.
Discriminator
A real violation is an answer with no reproducible code path or missing/renamed output file, or where a constraint (e.g., sort direction applied to all requested lists) never appears in code; a look-alike that is fine is a short but complete script that loads, preprocesses, sorts as specified, prints and writes the exact required file, even if the analysis is only a few lines.
Consequence
The grader finds the expected result file missing or its contents (values or their order) inconsistent with a correct data-derived computation, scoring 0.
id afe3381ac85c · mined from da-code dacode-di-text-003@s18
raw text (what the judge reads)
### Unverifiable answer: no saved script and no required output artifact
- **Applies when**: `task` -- the task asks for a specific result computed from a provided dataset (with stated preprocessing/ordering rules) and expects it saved in a named output file/format.
- **Pattern**: The agent reports a plausible-looking answer directly in chat (apparently from prior knowledge or an unsaved ad-hoc run), leaving no script that loads the data, performs the stated preprocessing, computes the requested ranking/statistic, and writes the required artifact — so neither the preprocessing rule nor the ordering constraint can be verified or reproduced.
- **Detection procedure**:
  1. From the task, list the required deliverables: the output file name/format, the stated preprocessing step, and any ordering/rounding/filtering constraint.
  2. Check whether any saved script exists that reads the input data and ends by writing that exact artifact with that exact key/schema.
  3. If a script exists, confirm each stated constraint is implemented in code (imputation applied before ranking; both lists explicitly sorted in the demanded direction) rather than assumed; if no script exists, flag immediately.
  4. Cross-check the reported values against the data (spot-check a couple of extremes and the sort direction of every list) — inability to do so from artifacts is itself the violation.
- **Discriminator**: A real violation is an answer with no reproducible code path or missing/renamed output file, or where a constraint (e.g., sort direction applied to *all* requested lists) never appears in code; a look-alike that is fine is a short but complete script that loads, preprocesses, sorts as specified, prints and writes the exact required file, even if the analysis is only a few lines.
- **Consequence**: The grader finds the expected result file missing or its contents (values or their order) inconsistent with a correct data-derived computation, scoring 0.
943Substituting proxy entities/metrics when the required fields aren't found in the datataskda-code
Applies when
task -- the task names specific entities, groupings, or measures (and required output artifacts), and the agent must locate them in the provided data before computing.
Pattern
The agent cannot find the requested columns/entities, silently redefines them as loosely analogous fields from a different file (or a different dataset entirely), computes a chart/statistic over those proxies, and declares success without flagging the substitution or producing all requested outputs.
Detection procedure
  1. List every entity, grouping key, aggregated quantity, and output file the task explicitly requests.
  2. Read the scripts/logs and check that each requested item maps to an actual column in the intended input source, and that all requested artifacts are written.
  3. Compare the answer's own description of what was plotted/computed against the task wording — look for phrases like "interpreted as", "used X instead of Y", or category names unrelated to the request.
  4. If any requested field is unavailable, verify the agent investigated other provided files/schemas and reported the ambiguity rather than improvising.
Discriminator
A legitimate case renames or derives the requested quantity from equivalent columns (e.g., computing a duration from two timestamp columns) and the semantics match the request; a violation replaces the requested concept with a semantically different one, or draws from an unrelated data source, so the axes/categories no longer answer the question.
Consequence
The produced figure and any accompanying data/config artifacts encode the wrong groups and quantities, so every expected output file mismatches and all checks fail.
id ed07f974e25c · mined from da-code dacode-plot-scatter-002@s18
raw text (what the judge reads)
### Substituting proxy entities/metrics when the required fields aren't found in the data
- **Applies when**: `task` -- the task names specific entities, groupings, or measures (and required output artifacts), and the agent must locate them in the provided data before computing.
- **Pattern**: The agent cannot find the requested columns/entities, silently redefines them as loosely analogous fields from a different file (or a different dataset entirely), computes a chart/statistic over those proxies, and declares success without flagging the substitution or producing all requested outputs.
- **Detection procedure**:
  1. List every entity, grouping key, aggregated quantity, and output file the task explicitly requests.
  2. Read the scripts/logs and check that each requested item maps to an actual column in the intended input source, and that all requested artifacts are written.
  3. Compare the answer's own description of what was plotted/computed against the task wording — look for phrases like "interpreted as", "used X instead of Y", or category names unrelated to the request.
  4. If any requested field is unavailable, verify the agent investigated other provided files/schemas and reported the ambiguity rather than improvising.
- **Discriminator**: A legitimate case renames or derives the requested quantity from equivalent columns (e.g., computing a duration from two timestamp columns) and the semantics match the request; a violation replaces the requested concept with a semantically different one, or draws from an unrelated data source, so the axes/categories no longer answer the question.
- **Consequence**: The produced figure and any accompanying data/config artifacts encode the wrong groups and quantities, so every expected output file mismatches and all checks fail.
944Unvalidated choice of the target quantity's definition / aggregation leveltaskinfiagent-dabench
Applies when
task -- the task asks for a statistic on a derived quantity (e.g., a per-entity duration, count, or rate) and the data contain several plausible candidate columns or require deciding whether to aggregate rows into entities.
Pattern
The agent's exploration script surfaces multiple candidate definitions of the quantity (a ready-made per-row measure, a difference of two timestamp/index columns, or a group-by sum), then silently commits to one of them in the analysis script with no justification, no cross-check of the alternatives, and no sanity check that the resulting distribution matches the units/scale implied by the task. All downstream filtering and statistics inherit the wrong base variable.
Detection procedure
  1. Read the task and note whether the requested quantity is stated to be per-row or per-entity, and what units/scale it should have.
  2. In the exploration script, list every candidate column or derivation the agent printed for that quantity; in the analysis script, note which one it used and whether any grouping/aggregation was applied.
  3. Check whether the agent ever compared candidates (row counts, magnitude ranges, agreement between a raw column and a computed difference) or explained why the chosen one is the one the task means.
  4. Confirm the final answer's magnitude and the number of retained data points are consistent with the chosen definition and reported anywhere in the answer/logs.
Discriminator
A real violation is when two or more defensible definitions exist and the choice is arbitrary/unjustified (or an aggregation was added/omitted without evidence); it is fine if the task or column semantics uniquely determine the quantity, or if the agent explicitly compared alternatives and showed they coincide or that one is invalid (e.g., missing/negative values).
Consequence
The mean and standard deviation are computed on a different variable than intended, so both reported values are off by a large factor and every graded check fails, even though the outlier-removal logic itself is correct.
id f26c8618f341 · mined from infiagent-dabench dabench-619@s18
raw text (what the judge reads)
### Unvalidated choice of the target quantity's definition / aggregation level
- **Applies when**: `task` -- the task asks for a statistic on a derived quantity (e.g., a per-entity duration, count, or rate) and the data contain several plausible candidate columns or require deciding whether to aggregate rows into entities.
- **Pattern**: The agent's exploration script surfaces multiple candidate definitions of the quantity (a ready-made per-row measure, a difference of two timestamp/index columns, or a group-by sum), then silently commits to one of them in the analysis script with no justification, no cross-check of the alternatives, and no sanity check that the resulting distribution matches the units/scale implied by the task. All downstream filtering and statistics inherit the wrong base variable.
- **Detection procedure**:
  1. Read the task and note whether the requested quantity is stated to be per-row or per-entity, and what units/scale it should have.
  2. In the exploration script, list every candidate column or derivation the agent printed for that quantity; in the analysis script, note which one it used and whether any grouping/aggregation was applied.
  3. Check whether the agent ever compared candidates (row counts, magnitude ranges, agreement between a raw column and a computed difference) or explained why the chosen one is the one the task means.
  4. Confirm the final answer's magnitude and the number of retained data points are consistent with the chosen definition and reported anywhere in the answer/logs.
- **Discriminator**: A real violation is when two or more defensible definitions exist and the choice is arbitrary/unjustified (or an aggregation was added/omitted without evidence); it is fine if the task or column semantics uniquely determine the quantity, or if the agent explicitly compared alternatives and showed they coincide or that one is invalid (e.g., missing/negative values).
- **Consequence**: The mean and standard deviation are computed on a different variable than intended, so both reported values are off by a large factor and every graded check fails, even though the outlier-removal logic itself is correct.
945Answer string does not literally instantiate the requested template (and no script proves it)taskinfiagent-dabench
Applies when
task -- the task dictates a rigid answer template (e.g. @name[...] with a specific delimiter/key set/spacing/value type), and the agent must emit that string, ideally generated by a saved script.
Pattern
The agent computes plausible numbers but hand-types the final answer, deviating from the template in some mechanical way (added/removed brackets or braces, reordered or renamed keys, quoting/spacing differences, values as floats/strings instead of ints, extra prose), and saves no script that produced the string — so the deviation is invisible and unreproducible.
Detection procedure
  1. Copy the answer template verbatim from the task and list every literal token it fixes: prefix/tag name, opening and closing delimiters, key names and their order, separator characters, and expected value type.
  2. Diff the agent's submitted string against that token list character by character; note any token that is added, dropped, renamed, reordered, or retyped.
  3. Check the scripts: is there code that prints the final answer string exactly as submitted (so it can be re-run and re-checked)? If no script exists or it prints only a dataframe/dict repr that the agent then retyped, mark the answer as unverified.
  4. Also confirm the value set covers exactly the constrained columns/keys the task lists — no omissions, no extras.
Discriminator
A real violation is any structural/token-level departure from the stated template, or an answer string with no generating code; a look-alike that is fine is a string that matches the template's tokens and key order with only whitespace the task explicitly leaves free, and is emitted by a saved, re-runnable print statement.
Consequence
The grader's parser fails to match the expected key/value structure and reports WRONG/MISSING for the whole item even when the underlying numbers are right, scoring 0.
id b94c2f45edc1 · mined from infiagent-dabench dabench-451@s18
raw text (what the judge reads)
### Answer string does not literally instantiate the requested template (and no script proves it)
- **Applies when**: `task` -- the task dictates a rigid answer template (e.g. `@name[...]` with a specific delimiter/key set/spacing/value type), and the agent must emit that string, ideally generated by a saved script.
- **Pattern**: The agent computes plausible numbers but hand-types the final answer, deviating from the template in some mechanical way (added/removed brackets or braces, reordered or renamed keys, quoting/spacing differences, values as floats/strings instead of ints, extra prose), and saves no script that produced the string — so the deviation is invisible and unreproducible.
- **Detection procedure**:
  1. Copy the answer template verbatim from the task and list every literal token it fixes: prefix/tag name, opening and closing delimiters, key names and their order, separator characters, and expected value type.
  2. Diff the agent's submitted string against that token list character by character; note any token that is added, dropped, renamed, reordered, or retyped.
  3. Check the scripts: is there code that prints the final answer string exactly as submitted (so it can be re-run and re-checked)? If no script exists or it prints only a dataframe/dict repr that the agent then retyped, mark the answer as unverified.
  4. Also confirm the value set covers exactly the constrained columns/keys the task lists — no omissions, no extras.
- **Discriminator**: A real violation is any structural/token-level departure from the stated template, or an answer string with no generating code; a look-alike that is fine is a string that matches the template's tokens and key order with only whitespace the task explicitly leaves free, and is emitted by a saved, re-runnable print statement.
- **Consequence**: The grader's parser fails to match the expected key/value structure and reports WRONG/MISSING for the whole item even when the underlying numbers are right, scoring 0.
946Outlier filtering applied at the wrong grouping level / wrong scope before a group-comparison testtaskda-code
Applies when
task -- the task prescribes a preprocessing step (e.g., quartile/IQR-based outlier removal) that is scoped to one grouping variable, and then a statistical test comparing another grouping variable's subgroups.
Pattern
The attempt computes the filter bounds on the wrong slice — the whole dataset, the test's own subgroups, or a different column than the target measure — or applies the filter after/instead of the prescribed order, and also drops or silently merges rare categories, so the samples entering the test are not the ones the task defines. No script is retained to show which rows survived.
Detection procedure
  1. From the task text, write down explicitly: which column the quantiles are computed on, which grouping variable defines the filtering partitions, which grouping variable defines the test groups, and how many test outputs are expected.
  2. In the scripts, locate the quantile/bound computation and check the groupby key and target column against step 1; confirm the filter is applied per-partition (not globally, not per-test-group) and that the test is run within each partition afterwards.
  3. Check that all prescribed categories/partitions are present after filtering (print group counts) and that the number of reported p-values equals the number of partitions, in a stated/consistent order.
  4. Recompute or spot-check one partition's bounds and surviving row count; verify the reported p-values map to the same groups and that each conclusion follows from its own p-value at the intended alpha.
Discriminator
A real violation is when the filtering partition, target column, or test-group set differs from the task's wording (e.g., bounds from pooled data or from the test groups themselves), producing different retained rows; a look-alike that is fine is an implementation that groups correctly but differs only in immaterial details (e.g., inclusive vs. exclusive boundary comparison, quantile interpolation method) that leave counts essentially unchanged.
Consequence
The p-values shift enough to flip one or more equal/not-equal conclusions, so the saved result file mismatches the expected values and the answer is graded wrong even though the format looks right.
id 69945a314cd4 · mined from da-code dacode-data-sa-061@s18
raw text (what the judge reads)
### Outlier filtering applied at the wrong grouping level / wrong scope before a group-comparison test
- **Applies when**: `task` -- the task prescribes a preprocessing step (e.g., quartile/IQR-based outlier removal) that is scoped to one grouping variable, and then a statistical test comparing another grouping variable's subgroups.
- **Pattern**: The attempt computes the filter bounds on the wrong slice — the whole dataset, the test's own subgroups, or a different column than the target measure — or applies the filter after/instead of the prescribed order, and also drops or silently merges rare categories, so the samples entering the test are not the ones the task defines. No script is retained to show which rows survived.
- **Detection procedure**:
  1. From the task text, write down explicitly: which column the quantiles are computed on, which grouping variable defines the filtering partitions, which grouping variable defines the test groups, and how many test outputs are expected.
  2. In the scripts, locate the quantile/bound computation and check the `groupby` key and target column against step 1; confirm the filter is applied per-partition (not globally, not per-test-group) and that the test is run within each partition afterwards.
  3. Check that all prescribed categories/partitions are present after filtering (print group counts) and that the number of reported p-values equals the number of partitions, in a stated/consistent order.
  4. Recompute or spot-check one partition's bounds and surviving row count; verify the reported p-values map to the same groups and that each conclusion follows from its own p-value at the intended alpha.
- **Discriminator**: A real violation is when the filtering partition, target column, or test-group set differs from the task's wording (e.g., bounds from pooled data or from the test groups themselves), producing different retained rows; a look-alike that is fine is an implementation that groups correctly but differs only in immaterial details (e.g., inclusive vs. exclusive boundary comparison, quantile interpolation method) that leave counts essentially unchanged.
- **Consequence**: The p-values shift enough to flip one or more equal/not-equal conclusions, so the saved result file mismatches the expected values and the answer is graded wrong even though the format looks right.
947Near-zero validation skill accepted without revisiting featurestaskda-code
Applies when
task -- the script trains a supervised model and reports a held-out score (R², accuracy, RMSE vs. baseline) before writing the prediction file.
Pattern
The agent builds features from only a narrow, convenient subset of columns (e.g. generic numeric columns), gets a held-out score that is indistinguishable from predicting the training mean/majority class, calls it "reasonable," and submits the predictions anyway without exploring the richer identifier/metadata/temporal/categorical columns present in both train and test, or checking that the prediction distribution resembles the target distribution.
Detection procedure
  1. Read the task and the data description to list all columns available in both the training file and the scoring file, including text, categorical, date, and grouping columns.
  2. In the script, compare the feature list actually used against that inventory; note any columns dropped without justification, and note whether any baseline (mean/median or trivial predictor) was computed for comparison.
  3. Read the reported validation metric: is it near the no-skill value (R²≈0, RMSE≈target std, accuracy≈base rate)? Compare the spread of the predictions to the spread of the training target.
  4. Check whether the agent, after seeing that score, iterated (added features, encodings, group statistics) or simply saved and declared success.
Discriminator
A real violation is when high-signal columns available at prediction time were never tried and the reported score is at/near no-skill while the agent still declares the result adequate. It is not a violation if the agent genuinely tested the additional columns (or documented why they are unavailable/leaky at prediction time) and the low score reflects an irreducibly noisy target, with the limitation stated honestly.
Consequence
The submitted column is essentially a flat, mean-centered prediction with far too little variance, so any accuracy/correlation-based check against the true values fails and the output file is marked wrong.
id 8d3eb74b2d33 · mined from da-code dacode-ml-regression-004@s18
raw text (what the judge reads)
### Near-zero validation skill accepted without revisiting features
- **Applies when**: `task` -- the script trains a supervised model and reports a held-out score (R², accuracy, RMSE vs. baseline) before writing the prediction file.
- **Pattern**: The agent builds features from only a narrow, convenient subset of columns (e.g. generic numeric columns), gets a held-out score that is indistinguishable from predicting the training mean/majority class, calls it "reasonable," and submits the predictions anyway without exploring the richer identifier/metadata/temporal/categorical columns present in both train and test, or checking that the prediction distribution resembles the target distribution.
- **Detection procedure**:
  1. Read the task and the data description to list all columns available in both the training file and the scoring file, including text, categorical, date, and grouping columns.
  2. In the script, compare the feature list actually used against that inventory; note any columns dropped without justification, and note whether any baseline (mean/median or trivial predictor) was computed for comparison.
  3. Read the reported validation metric: is it near the no-skill value (R²≈0, RMSE≈target std, accuracy≈base rate)? Compare the spread of the predictions to the spread of the training target.
  4. Check whether the agent, after seeing that score, iterated (added features, encodings, group statistics) or simply saved and declared success.
- **Discriminator**: A real violation is when high-signal columns available at prediction time were never tried and the reported score is at/near no-skill while the agent still declares the result adequate. It is *not* a violation if the agent genuinely tested the additional columns (or documented why they are unavailable/leaky at prediction time) and the low score reflects an irreducibly noisy target, with the limitation stated honestly.
- **Consequence**: The submitted column is essentially a flat, mean-centered prediction with far too little variance, so any accuracy/correlation-based check against the true values fails and the output file is marked wrong.
948Degenerate / unvalidated cluster solution accepted without sanity checkstaskda-code
Applies when
task -- the task asks for an unsupervised grouping (or any model output whose label distribution is the deliverable) and the script picks the number of groups automatically and writes labels straight to the output file.
Pattern
The attempt runs a single algorithm on raw (or merely standardized) skewed features, selects the group count by one internal score whose value is weak/ambiguous, and reports a partition containing near-empty groups (a handful or one member) that are really outlier capsules — with no check that the grouping is balanced, stable, or interpretable, and no comparison against neighbouring group counts or alternative preprocessing (skew correction/outlier handling).
Detection procedure
  1. Read the task for the required deliverable: expected column names, one row per input record, and whether the grouping is meant to be substantively interpretable (e.g., ranked tiers/severity groups).
  2. In the script, check whether feature scaling, missing values, and heavy-tailed/outlier features are handled, and whether the group count is chosen from more than one signal (e.g., score curve across k, elbow, cross-checked algorithm, or stability across seeds).
  3. In the answer, inspect the reported group sizes and selection metric: flag if any group holds ~1–2% or fewer of the records, if the chosen score is near the floor for the method, or if only one k was justified.
  4. Confirm the written file's shape/column names and label range were verified after writing (row count equals record count, labels contiguous, no NaNs).
Discriminator
A genuine small group is fine if the script shows evidence it is a stable, substantively meaningful segment (persists across seeds/algorithms, clearly separated after outlier-aware preprocessing) and the group count is defended by more than a single marginal score; a violation is a singleton/near-singleton group produced by unmitigated outliers with no stability or interpretability evidence, or a chosen k defended only by a weak internal score.
Consequence
The saved label column disagrees with the reference partition (wrong number of groups and grossly imbalanced membership), so the file check fails even though the format looks correct.
id 70e4edc6b806 · mined from da-code dacode-ml-cluster-013@s18
raw text (what the judge reads)
### Degenerate / unvalidated cluster solution accepted without sanity checks
- **Applies when**: `task` -- the task asks for an unsupervised grouping (or any model output whose label distribution is the deliverable) and the script picks the number of groups automatically and writes labels straight to the output file.
- **Pattern**: The attempt runs a single algorithm on raw (or merely standardized) skewed features, selects the group count by one internal score whose value is weak/ambiguous, and reports a partition containing near-empty groups (a handful or one member) that are really outlier capsules — with no check that the grouping is balanced, stable, or interpretable, and no comparison against neighbouring group counts or alternative preprocessing (skew correction/outlier handling).
- **Detection procedure**:
  1. Read the task for the required deliverable: expected column names, one row per input record, and whether the grouping is meant to be substantively interpretable (e.g., ranked tiers/severity groups).
  2. In the script, check whether feature scaling, missing values, and heavy-tailed/outlier features are handled, and whether the group count is chosen from more than one signal (e.g., score curve across k, elbow, cross-checked algorithm, or stability across seeds).
  3. In the answer, inspect the reported group sizes and selection metric: flag if any group holds ~1–2% or fewer of the records, if the chosen score is near the floor for the method, or if only one k was justified.
  4. Confirm the written file's shape/column names and label range were verified after writing (row count equals record count, labels contiguous, no NaNs).
- **Discriminator**: A genuine small group is fine if the script shows evidence it is a stable, substantively meaningful segment (persists across seeds/algorithms, clearly separated after outlier-aware preprocessing) and the group count is defended by more than a single marginal score; a violation is a singleton/near-singleton group produced by unmitigated outliers with no stability or interpretability evidence, or a chosen k defended only by a weak internal score.
- **Consequence**: The saved label column disagrees with the reference partition (wrong number of groups and grossly imbalanced membership), so the file check fails even though the format looks correct.
949Statistic computed over a different slice (or parameterization) than the task specifiestaskinfiagent-dabench
Applies when
task -- the prompt names a specific subset (a year, group, region, split) and/or a specific variant of a statistic/estimator, and the script computes the quantity by aggregating data itself.
Pattern
The script never materializes the stated subset (or silently widens it to "all columns/rows matching a prefix") and/or uses the library's default option instead of the named variant; the printed diagnostics reinforce the wrong slice, so verification scripts re-confirm the same mistake.
Detection procedure
  1. From the task text, list every explicit restriction: which rows/columns/period define the sample, which axis the statistic runs over, which definition/flag/normalization is demanded, and which data files are in scope.
  2. In the script, locate the line that builds the input array to the statistic and the call itself; check that each restriction from step 1 appears literally (a filter to the named subset, the correct axis, the correct keyword such as the bias/ddof/definition flag), and check whether relevant data files were all loaded rather than just one.
  3. If any restriction is absent, check whether the code's comment claims compliance ("default is X") without evidence — a mismatch between the comment and the actual library semantics is a violation.
  4. Confirm the reported answer is the one selected from the correctly restricted computation, not from the broader/looser one.
Discriminator
A real violation is when the restriction changes the input set or the estimator formula (e.g., aggregating over all periods when one period was named, or using the biased/unadjusted form when the adjusted one was named). A look-alike that is fine is when the restriction is provably a no-op on this data (e.g., the filter selects everything, or the flag's default equals the requested variant) and the script or reviewer can demonstrate that.
Consequence
The ranking/argmax or numeric value is derived from a different population or formula, so the grader sees a plausible but wrong entity/number and scores 0, with no error message to hint at the cause.
id e80441cae4a4 · mined from infiagent-dabench dabench-252@s18
raw text (what the judge reads)
### Statistic computed over a different slice (or parameterization) than the task specifies
- **Applies when**: `task` -- the prompt names a specific subset (a year, group, region, split) and/or a specific variant of a statistic/estimator, and the script computes the quantity by aggregating data itself.
- **Pattern**: The script never materializes the stated subset (or silently widens it to "all columns/rows matching a prefix") and/or uses the library's default option instead of the named variant; the printed diagnostics reinforce the wrong slice, so verification scripts re-confirm the same mistake.
- **Detection procedure**:
  1. From the task text, list every explicit restriction: which rows/columns/period define the sample, which axis the statistic runs over, which definition/flag/normalization is demanded, and which data files are in scope.
  2. In the script, locate the line that builds the input array to the statistic and the call itself; check that each restriction from step 1 appears literally (a filter to the named subset, the correct axis, the correct keyword such as the bias/ddof/definition flag), and check whether relevant data files were all loaded rather than just one.
  3. If any restriction is absent, check whether the code's comment claims compliance ("default is X") without evidence — a mismatch between the comment and the actual library semantics is a violation.
  4. Confirm the reported answer is the one selected from the correctly restricted computation, not from the broader/looser one.
- **Discriminator**: A real violation is when the restriction changes the input set or the estimator formula (e.g., aggregating over all periods when one period was named, or using the biased/unadjusted form when the adjusted one was named). A look-alike that is fine is when the restriction is provably a no-op on this data (e.g., the filter selects everything, or the flag's default equals the requested variant) and the script or reviewer can demonstrate that.
- **Consequence**: The ranking/argmax or numeric value is derived from a different population or formula, so the grader sees a plausible but wrong entity/number and scores 0, with no error message to hint at the cause.
950Reported identifier truncated to fit a format template instead of preserving the value's actual granularitytaskinfiagent-dabench
Applies when
task -- The task asks you to locate a specific record (e.g., an argmax/argmin key such as a date, ID, or category) and echo it back, and the answer template shows an abbreviated or lower-precision rendering than the values actually stored in the data.
Pattern
The agent finds the correct record but then coerces the key to the literal template pattern (dropping day-level precision, truncating an ID, lowercasing/reformatting a label), so the submitted string no longer matches the value present in the data; the dependent calculation is done on the correct full-precision record, hiding the error.
Detection procedure
  1. Read the task and note the granularity/format of the identifier the answer template implies.
  2. Inspect the data (or the script's parsing/printing of the key column) to see the granularity actually stored, and check whether the argmax selection was done at that full granularity.
  3. Compare the submitted identifier string to the raw key of the selected record — flag if any component (day, suffix, case, separators) was dropped or rewritten purely to match the template.
  4. Confirm consistency: if the downstream computation used a full-precision record, the reported identifier must be that same full-precision key.
Discriminator
A real violation is losing information that exists in the data (full date reduced to month, ID prefix only). It is not a violation when the underlying values genuinely have that coarser granularity, or when the task explicitly asks for an aggregation at the coarser level and the aggregation was actually performed at that level.
Consequence
The identifier check fails as WRONG/MISSING even though the derived numeric answer matches, producing a partial score (e.g., 1/2 checks) and an overall incorrect verdict.
id b5f4ead6b899 · mined from infiagent-dabench dabench-572@s18
raw text (what the judge reads)
### Reported identifier truncated to fit a format template instead of preserving the value's actual granularity
- **Applies when**: `task` -- The task asks you to locate a specific record (e.g., an argmax/argmin key such as a date, ID, or category) and echo it back, and the answer template shows an abbreviated or lower-precision rendering than the values actually stored in the data.
- **Pattern**: The agent finds the correct record but then coerces the key to the literal template pattern (dropping day-level precision, truncating an ID, lowercasing/reformatting a label), so the submitted string no longer matches the value present in the data; the dependent calculation is done on the correct full-precision record, hiding the error.
- **Detection procedure**:
  1. Read the task and note the granularity/format of the identifier the answer template implies.
  2. Inspect the data (or the script's parsing/printing of the key column) to see the granularity actually stored, and check whether the argmax selection was done at that full granularity.
  3. Compare the submitted identifier string to the raw key of the selected record — flag if any component (day, suffix, case, separators) was dropped or rewritten purely to match the template.
  4. Confirm consistency: if the downstream computation used a full-precision record, the reported identifier must be that same full-precision key.
- **Discriminator**: A real violation is losing information that exists in the data (full date reduced to month, ID prefix only). It is *not* a violation when the underlying values genuinely have that coarser granularity, or when the task explicitly asks for an aggregation at the coarser level and the aggregation was actually performed at that level.
- **Consequence**: The identifier check fails as WRONG/MISSING even though the derived numeric answer matches, producing a partial score (e.g., 1/2 checks) and an overall incorrect verdict.
951Degenerate single-class prediction output on an imbalanced classification tasktaskda-code
Applies when
task -- the task asks for per-row class predictions on a held-out set where the positive class is rare and the stated goal is to detect those rare events.
Pattern
The attempt emits a prediction column that is constant (all majority class, or nearly so), typically because the model was fit with a default 0.5 threshold on heavily imbalanced data, or because the pipeline silently failed (no model actually trained, all-NaN features, wrong feature matrix) and a fallback/default value was written out. The submission is technically well-formatted, so the agent never notices there are zero positive predictions.
Detection procedure
  1. Read the task/README to determine the target's expected base rate (stated class imbalance, or the training label distribution the script should have computed) and whether the objective emphasizes recall of the minority class.
  2. In the scripts, check that a model is actually fit on training features and that predict is called on the transformed test features — look for evidence of class-imbalance handling (class weights, resampling, tuned threshold) and any validation-score printout.
  3. In the output file, compute the value counts of the prediction column: flag if the minority class count is 0 or wildly below the training base rate times the number of test rows (e.g., <10–20% of expected positives), or if the counts were never checked in the script.
  4. Confirm the row count and header match the sample-format file, and that the script did not fall back to writing a constant after an exception.
Discriminator
A genuine violation is a constant/near-empty minority prediction with no evidence of a validated model (no held-out score, no threshold selection, no positive-rate sanity check). A look-alike that is fine: a heavily skewed but non-degenerate prediction whose positive rate is comparable to the training prevalence and is backed by a reported validation metric (recall/F1/AUC) on the minority class.
Consequence
The submitted file matches the required shape but scores at or below the majority-class baseline on any minority-sensitive metric (recall = 0, F1 = 0), so an exact/threshold comparison against the expected target file fails.
id 9877e6cd3c92 · mined from da-code dacode-ml-binary-013@s18
raw text (what the judge reads)
### Degenerate single-class prediction output on an imbalanced classification task
- **Applies when**: `task` -- the task asks for per-row class predictions on a held-out set where the positive class is rare and the stated goal is to *detect* those rare events.
- **Pattern**: The attempt emits a prediction column that is constant (all majority class, or nearly so), typically because the model was fit with a default 0.5 threshold on heavily imbalanced data, or because the pipeline silently failed (no model actually trained, all-NaN features, wrong feature matrix) and a fallback/default value was written out. The submission is technically well-formatted, so the agent never notices there are zero positive predictions.
- **Detection procedure**:
  1. Read the task/README to determine the target's expected base rate (stated class imbalance, or the training label distribution the script should have computed) and whether the objective emphasizes recall of the minority class.
  2. In the scripts, check that a model is actually fit on training features and that `predict` is called on the transformed test features — look for evidence of class-imbalance handling (class weights, resampling, tuned threshold) and any validation-score printout.
  3. In the output file, compute the value counts of the prediction column: flag if the minority class count is 0 or wildly below the training base rate times the number of test rows (e.g., <10–20% of expected positives), or if the counts were never checked in the script.
  4. Confirm the row count and header match the sample-format file, and that the script did not fall back to writing a constant after an exception.
- **Discriminator**: A genuine violation is a constant/near-empty minority prediction with no evidence of a validated model (no held-out score, no threshold selection, no positive-rate sanity check). A look-alike that is fine: a heavily skewed but non-degenerate prediction whose positive rate is comparable to the training prevalence and is backed by a reported validation metric (recall/F1/AUC) on the minority class.
- **Consequence**: The submitted file matches the required shape but scores at or below the majority-class baseline on any minority-sensitive metric (recall = 0, F1 = 0), so an exact/threshold comparison against the expected target file fails.
952Analysis run on a single partition/file when the task asks about the whole populationtaskinfiagent-dabench
Applies when
task -- the question is posed over an entire population of entities ("all X"), and the scripts hard-code a single input file, sheet, or pre-filtered subset without checking what else is available.
Pattern
The agent loads one convenient slice of the data (one region/segment/split file), computes the requested statistic on it, and reports the result as if it covered the full population — never enumerating the data directory or verifying row counts against the expected number of entities. Any statistic that depends on the full distribution (quantiles, thresholds, rankings) is then computed on the wrong subset, and the reported label set is likewise incomplete or format-mismatched with the expected key.
Detection procedure
  1. Read the task and note the stated scope of entities and the exact answer key/format requested.
  2. In the scripts, list every data source read; check whether the agent enumerated available files/tables or justified that the single source is the complete population.
  3. Check for a sanity check on shape/count (number of unique entities loaded vs. the number implied by the task) and, if several partitions exist, whether they are concatenated before computing quantiles/thresholds.
  4. Compare the final printed answer against the requested key name, list structure, and string spelling/quoting.
Discriminator
A real violation is when the loaded source is provably one of several partitions (or a filtered view) and no concatenation/verification step exists; it is fine if the script explicitly checks the source contains all entities, or the task itself scopes the analysis to that subset.
Consequence
Quantile-derived thresholds and the resulting entity list are computed on a non-representative subset, so the graded answer differs from the ground-truth set (or fails the exact-key/format check) and scores 0.
id 273b97234da7 · mined from infiagent-dabench dabench-254@s18
raw text (what the judge reads)
### Analysis run on a single partition/file when the task asks about the whole population
- **Applies when**: `task` -- the question is posed over an entire population of entities ("all X"), and the scripts hard-code a single input file, sheet, or pre-filtered subset without checking what else is available.
- **Pattern**: The agent loads one convenient slice of the data (one region/segment/split file), computes the requested statistic on it, and reports the result as if it covered the full population — never enumerating the data directory or verifying row counts against the expected number of entities. Any statistic that depends on the full distribution (quantiles, thresholds, rankings) is then computed on the wrong subset, and the reported label set is likewise incomplete or format-mismatched with the expected key.
- **Detection procedure**:
  1. Read the task and note the stated scope of entities and the exact answer key/format requested.
  2. In the scripts, list every data source read; check whether the agent enumerated available files/tables or justified that the single source is the complete population.
  3. Check for a sanity check on shape/count (number of unique entities loaded vs. the number implied by the task) and, if several partitions exist, whether they are concatenated before computing quantiles/thresholds.
  4. Compare the final printed answer against the requested key name, list structure, and string spelling/quoting.
- **Discriminator**: A real violation is when the loaded source is provably one of several partitions (or a filtered view) and no concatenation/verification step exists; it is fine if the script explicitly checks the source contains all entities, or the task itself scopes the analysis to that subset.
- **Consequence**: Quantile-derived thresholds and the resulting entity list are computed on a non-representative subset, so the graded answer differs from the ground-truth set (or fails the exact-key/format check) and scores 0.
953Ships the deliverable without holdout validation or re-reading/verifying the saved artifacttaskda-code
Applies when
task -- the task asks for a prediction/result file with a specified column name and format, and the script trains a model on labeled data then writes predictions straight to disk.
Pattern
The attempt fits one model, predicts, and calls to_csv (often with the DataFrame index or extra columns included), then declares success based only on printed class counts — never holding out a labeled subset to estimate accuracy, never re-loading the written file to confirm its exact columns, row count, ordering, and that the predicted label strings match the label vocabulary/format of the training target.
Detection procedure
  1. Read the task for the required output: file name, exact column name(s), implied row count/order, and label values expected.
  2. Scan the script for (a) any evaluation on a labeled holdout/CV split (imports of split/metric functions that are never used are a red flag) and (b) any post-write step that reads the output file back and asserts shape, header, null count, and unique label values.
  3. Check how the file is written: is an index or extra column emitted, are labels written in the exact original spelling/case, are rows in test-file order with no dropped/reordered rows?
  4. Check the reported answer: does it cite a validation score and a verified file schema, or only class proportions and "task completed"?
Discriminator
A real violation has zero quantitative quality estimate and zero verification of the artifact's contents — the agent could not have detected a wrong header, extra index column, mismatched label strings, or a degenerate model. It is not a violation if the script reports a holdout/CV metric and prints/asserts the reloaded file's shape, columns and label set, even if no hyperparameter tuning was done.
Consequence
The grader reads result.csv and finds it wrong or unparsable in the expected schema (extra/renamed columns, altered label text, wrong row count) or below the accuracy threshold, and marks the file WRONG/MISSING despite a confident "completed successfully" report.
id 406afd90bff9 · mined from da-code dacode-ml-binary-009@s18
raw text (what the judge reads)
### Ships the deliverable without holdout validation or re-reading/verifying the saved artifact
- **Applies when**: `task` -- the task asks for a prediction/result file with a specified column name and format, and the script trains a model on labeled data then writes predictions straight to disk.
- **Pattern**: The attempt fits one model, predicts, and calls `to_csv` (often with the DataFrame index or extra columns included), then declares success based only on printed class counts — never holding out a labeled subset to estimate accuracy, never re-loading the written file to confirm its exact columns, row count, ordering, and that the predicted label strings match the label vocabulary/format of the training target.
- **Detection procedure**:
  1. Read the task for the required output: file name, exact column name(s), implied row count/order, and label values expected.
  2. Scan the script for (a) any evaluation on a labeled holdout/CV split (imports of split/metric functions that are never used are a red flag) and (b) any post-write step that reads the output file back and asserts shape, header, null count, and unique label values.
  3. Check how the file is written: is an index or extra column emitted, are labels written in the exact original spelling/case, are rows in test-file order with no dropped/reordered rows?
  4. Check the reported answer: does it cite a validation score and a verified file schema, or only class proportions and "task completed"?
- **Discriminator**: A real violation has zero quantitative quality estimate *and* zero verification of the artifact's contents — the agent could not have detected a wrong header, extra index column, mismatched label strings, or a degenerate model. It is not a violation if the script reports a holdout/CV metric and prints/asserts the reloaded file's shape, columns and label set, even if no hyperparameter tuning was done.
- **Consequence**: The grader reads `result.csv` and finds it wrong or unparsable in the expected schema (extra/renamed columns, altered label text, wrong row count) or below the accuracy threshold, and marks the file WRONG/MISSING despite a confident "completed successfully" report.
954Fabricated aggregation metric (and one-sided entity stats) instead of the definition implied by the task/configtaskda-code
Applies when
task -- the script must summarize a "performance"/score per entity from paired or two-sided records, and the metric is only loosely named in the prompt or partly specified in a config/spec file.
Pattern
The agent invents an arbitrary formula (e.g., a weighted sum of match counts and a venue flag) with no support in the task text, config, or domain convention, and aggregates only one side of the paired records (grouping by one participant column and ignoring the rows where the entity appears in the other column), so both the definition and the underlying counts are wrong; it also silently skips required auxiliary outputs listed in the spec.
Detection procedure
  1. Read the task and the referenced spec/config file and list every required output artifact and every field that constrains the computation (labels, ordering, filtering window, units, derived quantities).
  2. In the script, locate the line(s) that define the plotted/reported quantity and check whether each term traces back to an explicit instruction, a config key, or a standard domain definition; flag any coefficient or term the agent chose itself, and flag any use of axis labels/titles as evidence for a formula.
  3. Check the aggregation: for paired records, verify the entity's rows are collected from all participant columns (concat/melt or two groupbys summed), not a single groupby(one_side_column).
  4. Compare produced files against the required artifact list and sanity-check magnitudes/ordering against a known-plausible ranking of entities.
Discriminator
A real violation is an unsupported formula or single-side grouping that changes the numbers; it is fine if the metric is explicitly given (in prompt or config) or is the standard definition, and the single-column grouping is provably complete (the entity can only appear in that column) — or if the agent explicitly validated its choice against the config/expected labels.
Consequence
The chart values, ordering, and any saved numeric/spec artifacts differ from the reference, so every expected-file check (image, spec JSON, numeric array) fails even though the plot "looks" correct.
id e73a092715a9 · mined from da-code dacode-plot-bar-006@s18
raw text (what the judge reads)
### Fabricated aggregation metric (and one-sided entity stats) instead of the definition implied by the task/config
- **Applies when**: `task` -- the script must summarize a "performance"/score per entity from paired or two-sided records, and the metric is only loosely named in the prompt or partly specified in a config/spec file.
- **Pattern**: The agent invents an arbitrary formula (e.g., a weighted sum of match counts and a venue flag) with no support in the task text, config, or domain convention, and aggregates only one side of the paired records (grouping by one participant column and ignoring the rows where the entity appears in the other column), so both the definition and the underlying counts are wrong; it also silently skips required auxiliary outputs listed in the spec.
- **Detection procedure**:
  1. Read the task and the referenced spec/config file and list every required output artifact and every field that constrains the computation (labels, ordering, filtering window, units, derived quantities).
  2. In the script, locate the line(s) that define the plotted/reported quantity and check whether each term traces back to an explicit instruction, a config key, or a standard domain definition; flag any coefficient or term the agent chose itself, and flag any use of axis labels/titles as evidence for a formula.
  3. Check the aggregation: for paired records, verify the entity's rows are collected from *all* participant columns (concat/melt or two groupbys summed), not a single `groupby(one_side_column)`.
  4. Compare produced files against the required artifact list and sanity-check magnitudes/ordering against a known-plausible ranking of entities.
- **Discriminator**: A real violation is an unsupported formula or single-side grouping that changes the numbers; it is *fine* if the metric is explicitly given (in prompt or config) or is the standard definition, and the single-column grouping is provably complete (the entity can only appear in that column) — or if the agent explicitly validated its choice against the config/expected labels.
- **Consequence**: The chart values, ordering, and any saved numeric/spec artifacts differ from the reference, so every expected-file check (image, spec JSON, numeric array) fails even though the plot "looks" correct.
955Unverified input selection and no sanity-check of a threshold-based counttaskinfiagent-dabench
Applies when
task -- the task asks for a count/flag of records passing a fixed statistical threshold (e.g., a sigma/percentile rule) on a named column of "the dataset", and the script hard-codes one file and one column without confirming they are the intended ones.
Pattern
The agent picks the first plausible file/column it finds (e.g., a split or variant file), computes the statistic with one library's default conventions (population vs. sample std, NaN dropped/kept), and reports the resulting count with no cross-check that the file, column, row count, or the magnitude of the result is consistent with the task's expectations.
Detection procedure
  1. Read the task and note exactly which data source, column, and rule (threshold, definition, denominator) are specified.
  2. In the script, check whether the loaded file/column is justified — is the directory listed, are candidate files/columns compared, is the row count and column set printed and reconciled with the task description?
  3. Check whether the threshold statistic is computed only once with library defaults, or whether the agent verified sensitivity to conventions (ddof, NaN/sentinel handling, standardization on the full column vs. a subset).
  4. Look at the reported count as a fraction of rows and ask whether it is plausible under the stated rule (e.g., a >3σ rule should flag a tiny tail; a large fraction signals heavy tails, sentinel/placeholder codes, wrong column, or wrong file) — flag if the script prints the number without any such reasoning or inspection of the flagged values.
Discriminator
A fine attempt either shows the candidate files/columns and explains the choice, prints distribution diagnostics and the actual extreme values, and notes the count is consistent with the rule; a violation reports a bare count from a single unexamined file/column with no plausibility or convention check. (An implausible-looking count that the agent inspected and justified with evidence is not a violation.)
Consequence
The reported count comes from the wrong data subset or a mis-specified statistic, so the graded scalar (outlier_count-style answer) differs from ground truth and all checks fail, even though the code runs cleanly.
id e265e1dc33a3 · mined from infiagent-dabench dabench-361@s18
raw text (what the judge reads)
### Unverified input selection and no sanity-check of a threshold-based count
- **Applies when**: `task` -- the task asks for a count/flag of records passing a fixed statistical threshold (e.g., a sigma/percentile rule) on a named column of "the dataset", and the script hard-codes one file and one column without confirming they are the intended ones.
- **Pattern**: The agent picks the first plausible file/column it finds (e.g., a split or variant file), computes the statistic with one library's default conventions (population vs. sample std, NaN dropped/kept), and reports the resulting count with no cross-check that the file, column, row count, or the magnitude of the result is consistent with the task's expectations.
- **Detection procedure**:
  1. Read the task and note exactly which data source, column, and rule (threshold, definition, denominator) are specified.
  2. In the script, check whether the loaded file/column is justified — is the directory listed, are candidate files/columns compared, is the row count and column set printed and reconciled with the task description?
  3. Check whether the threshold statistic is computed only once with library defaults, or whether the agent verified sensitivity to conventions (ddof, NaN/sentinel handling, standardization on the full column vs. a subset).
  4. Look at the reported count as a fraction of rows and ask whether it is plausible under the stated rule (e.g., a >3σ rule should flag a tiny tail; a large fraction signals heavy tails, sentinel/placeholder codes, wrong column, or wrong file) — flag if the script prints the number without any such reasoning or inspection of the flagged values.
- **Discriminator**: A fine attempt either shows the candidate files/columns and explains the choice, prints distribution diagnostics and the actual extreme values, and notes the count is consistent with the rule; a violation reports a bare count from a single unexamined file/column with no plausibility or convention check. (An implausible-looking count that the agent inspected and justified with evidence is not a violation.)
- **Consequence**: The reported count comes from the wrong data subset or a mis-specified statistic, so the graded scalar (`outlier_count`-style answer) differs from ground truth and all checks fail, even though the code runs cleanly.
956Substituting assumed conventions for an explicitly referenced specification filetaskda-code
Applies when
task -- the task instructs the agent to use a definition, mapping, threshold, or rule supplied in an auxiliary document/config, and the scripts must apply it.
Pattern
The scripts never load or quote the referenced document; instead the agent hardcodes a mapping/definition from "standard industry knowledge" (optionally after a failed file search) and proceeds, so downstream label names, groupings, or denominators may not match the specified ones — and the required output artifact is often not written either.
Detection procedure
  1. From the task text, list every external artifact referenced (spec/notes/config) and every required output file/format.
  2. Scan the scripts for a read of that artifact (open/read_csv/read of the file, or its contents printed and echoed); check whether the applied definition is derived from it or hardcoded.
  3. Check that the reported labels/values use exactly the vocabulary and granularity from the spec (e.g., could two raw categories collapse into one group, changing which is most frequent?), and that the answer is persisted to the requested filename/format.
  4. If the artifact was not found, verify the agent searched the actual data directory and halted/reported ambiguity rather than guessing.
Discriminator
Fine if the script reads the referenced file (or reproduces its verified contents verbatim) and the hardcoded values provably match it; a violation is when the mapping's source is the agent's assumption, or when the spec could merge/rename categories in a way the script cannot rule out.
Consequence
The reported category name and ratio are computed over the wrong grouping (and the expected result file is missing), so the graded comparison fails even though the counting arithmetic is correct.
id d920fe4799ec · mined from da-code dacode-di-text-004@s18
raw text (what the judge reads)
### Substituting assumed conventions for an explicitly referenced specification file
- **Applies when**: `task` -- the task instructs the agent to use a definition, mapping, threshold, or rule supplied in an auxiliary document/config, and the scripts must apply it.
- **Pattern**: The scripts never load or quote the referenced document; instead the agent hardcodes a mapping/definition from "standard industry knowledge" (optionally after a failed file search) and proceeds, so downstream label names, groupings, or denominators may not match the specified ones — and the required output artifact is often not written either.
- **Detection procedure**:
  1. From the task text, list every external artifact referenced (spec/notes/config) and every required output file/format.
  2. Scan the scripts for a read of that artifact (open/read_csv/read of the file, or its contents printed and echoed); check whether the applied definition is derived from it or hardcoded.
  3. Check that the reported labels/values use exactly the vocabulary and granularity from the spec (e.g., could two raw categories collapse into one group, changing which is most frequent?), and that the answer is persisted to the requested filename/format.
  4. If the artifact was not found, verify the agent searched the actual data directory and halted/reported ambiguity rather than guessing.
- **Discriminator**: Fine if the script reads the referenced file (or reproduces its verified contents verbatim) and the hardcoded values provably match it; a violation is when the mapping's source is the agent's assumption, or when the spec could merge/rename categories in a way the script cannot rule out.
- **Consequence**: The reported category name and ratio are computed over the wrong grouping (and the expected result file is missing), so the graded comparison fails even though the counting arithmetic is correct.
957Submission file not verified against the provided template (header names, row count, completeness)taskda-code
Applies when
task -- the task requires writing predictions to an output file whose format is defined by a provided sample/template file, and the agent must produce that file rather than just print values.
Pattern
The agent constructs the output by hand or streams it to the transcript, using column names/casing/order copied from prose or invented, and never programmatically loads the template to confirm the header spelling, column order, id column values/order, and that one row exists for every test id — often ending up with a truncated, partially written, or mis-headed file.
Detection procedure
  1. Read the task/README for the named template file and note that the output must match it exactly (header text and casing, column order, one row per test id).
  2. In the scripts, look for code that (a) reads the template and/or test ids, (b) builds a DataFrame with columns taken from the template, and (c) writes the required filename with index=False; also check a saved script exists at all rather than ad-hoc pasted output.
  3. Check for an explicit post-write sanity check: reload the written file and assert its shape/row count equals the template's, its header equals the template's header, ids match the test set, and value columns are in valid range and sum to 1 where required.
  4. Compare the answer's header and row count to the template's; any casing/name difference, missing/extra rows, or evidence of truncation is a violation.
Discriminator
A real violation is a header/column/id/row-count mismatch or an unverified, hand-assembled or truncated file; it is fine if the script derives columns and ids from the template/test data and asserts shape and header after writing, even if the probability values themselves are imperfect.
Consequence
The grader cannot parse or align the submission against ground truth (unknown column names, missing ids), so the file is marked WRONG/MISSING and scores zero regardless of model quality.
id b6615e4e52ff · mined from da-code dacode-ml-competition-003@s18
raw text (what the judge reads)
### Submission file not verified against the provided template (header names, row count, completeness)
- **Applies when**: `task` -- the task requires writing predictions to an output file whose format is defined by a provided sample/template file, and the agent must produce that file rather than just print values.
- **Pattern**: The agent constructs the output by hand or streams it to the transcript, using column names/casing/order copied from prose or invented, and never programmatically loads the template to confirm the header spelling, column order, id column values/order, and that one row exists for every test id — often ending up with a truncated, partially written, or mis-headed file.
- **Detection procedure**:
  1. Read the task/README for the named template file and note that the output must match it exactly (header text and casing, column order, one row per test id).
  2. In the scripts, look for code that (a) reads the template and/or test ids, (b) builds a DataFrame with columns taken from the template, and (c) writes the required filename with `index=False`; also check a saved script exists at all rather than ad-hoc pasted output.
  3. Check for an explicit post-write sanity check: reload the written file and assert its shape/row count equals the template's, its header equals the template's header, ids match the test set, and value columns are in valid range and sum to 1 where required.
  4. Compare the answer's header and row count to the template's; any casing/name difference, missing/extra rows, or evidence of truncation is a violation.
- **Discriminator**: A real violation is a header/column/id/row-count mismatch or an unverified, hand-assembled or truncated file; it is fine if the script derives columns and ids from the template/test data and asserts shape and header after writing, even if the probability values themselves are imperfect.
- **Consequence**: The grader cannot parse or align the submission against ground truth (unknown column names, missing ids), so the file is marked WRONG/MISSING and scores zero regardless of model quality.
958Fabricating the evaluation set and labels instead of using the provided onestaskda-code
Applies when
task -- the task names specific input files (e.g., a designated test split) and a target column, and the scripts must produce predictions aligned to those rows.
Pattern
The agent never loads the specified evaluation file (or the file containing the target), instead synthesizes its own train/test split from auxiliary tables and invents target labels via ad-hoc heuristics, then trains/evaluates a model on those self-made labels and reports near-perfect accuracy.
Detection procedure
1) List the files/columns the task explicitly requires and the expected row count/keys of the output. 2) Grep the scripts for reads of those exact files and for the target column; check whether the target is read from data or constructed by a rule/function in code. 3) Check whether the predicted rows come from the provided evaluation file's keys in its original order, or from an internally generated split. 4) Compare reported validation accuracy and output row count against expectations — labels created by a deterministic rule from the same features yield implausibly high scores.
Discriminator
A real violation is when the ground-truth target or the evaluation rows are invented in code (no read of the required file, or apply(heuristic) defines the label). It is fine if the agent creates internal validation splits in addition to predicting on the provided evaluation rows, or engineers features (not labels) heuristically.
Consequence
The output file's row keys and count do not match the expected evaluation set, so nearly all rows are unmatched/wrong and the file is scored WRONG/MISSING despite reported ~100% internal accuracy.
id ccc16295d3bf · mined from da-code dacode-ml-multi-003@s18
raw text (what the judge reads)
### Fabricating the evaluation set and labels instead of using the provided ones
- **Applies when**: `task` -- the task names specific input files (e.g., a designated test split) and a target column, and the scripts must produce predictions aligned to those rows.
- **Pattern**: The agent never loads the specified evaluation file (or the file containing the target), instead synthesizes its own train/test split from auxiliary tables and invents target labels via ad-hoc heuristics, then trains/evaluates a model on those self-made labels and reports near-perfect accuracy.
- **Detection procedure**: 1) List the files/columns the task explicitly requires and the expected row count/keys of the output. 2) Grep the scripts for reads of those exact files and for the target column; check whether the target is read from data or constructed by a rule/function in code. 3) Check whether the predicted rows come from the provided evaluation file's keys in its original order, or from an internally generated split. 4) Compare reported validation accuracy and output row count against expectations — labels created by a deterministic rule from the same features yield implausibly high scores.
- **Discriminator**: A real violation is when the ground-truth target or the evaluation rows are invented in code (no read of the required file, or `apply(heuristic)` defines the label). It is fine if the agent creates internal validation splits *in addition to* predicting on the provided evaluation rows, or engineers features (not labels) heuristically.
- **Consequence**: The output file's row keys and count do not match the expected evaluation set, so nearly all rows are unmatched/wrong and the file is scored WRONG/MISSING despite reported ~100% internal accuracy.
959Template/format conformance asserted by eyeball instead of programmatic comparisontaskda-code
Applies when
task -- the task supplies a template or example output file (or an explicit output spec) and the scripts write a result file whose row labels, column headers, ordering, or numeric precision must match it.
Pattern
The agent constructs the output from its own assumptions (self-chosen index string format, self-chosen column names/order, arbitrary rounding such as 1 decimal, its own set of rows/columns), prints the template only to visually skim it, and never performs an explicit structural comparison; discrepancies in label formatting, precision, or included cells go undetected.
Detection procedure
  1. Read the task/README for a referenced template or format constraint and note what must match (header names, index labels, row/column count and order, decimals, units, missing-value representation).
  2. Scan the scripts for a step that loads the template and asserts equality of columns, index, and shape against the produced output (e.g., comparing header lists, df.shape, index values, and value precision); a mere print of both files is not such a step.
  3. Check whether any transformation is justified by the template or invented by the agent — rounding level, date-label styling, 0- vs 1-based period numbering, index naming.
  4. Check the final answer for a stated verification that the output matches the template cell-for-cell in structure; absence of it, or verification only against the agent's own file, is a violation.
Discriminator
Fine if the script (or a verification run) reads the template and compares headers/index/shape/precision programmatically and any remaining differences are only in the numeric values being computed; a violation if format decisions rest on visual inspection or unstated assumptions, or if verification re-reads only the agent's own output.
Consequence
The grader's file comparison fails (wrong index labels, extra/missing columns, or over-rounded values), scoring the result file as WRONG even when the underlying aggregation logic is plausible.
id df7d1bf85169 · mined from da-code dacode-dm-csv-044@s18
raw text (what the judge reads)
### Template/format conformance asserted by eyeball instead of programmatic comparison
- **Applies when**: `task` -- the task supplies a template or example output file (or an explicit output spec) and the scripts write a result file whose row labels, column headers, ordering, or numeric precision must match it.
- **Pattern**: The agent constructs the output from its own assumptions (self-chosen index string format, self-chosen column names/order, arbitrary rounding such as 1 decimal, its own set of rows/columns), prints the template only to visually skim it, and never performs an explicit structural comparison; discrepancies in label formatting, precision, or included cells go undetected.
- **Detection procedure**:
  1. Read the task/README for a referenced template or format constraint and note what must match (header names, index labels, row/column count and order, decimals, units, missing-value representation).
  2. Scan the scripts for a step that *loads* the template and asserts equality of columns, index, and shape against the produced output (e.g., comparing header lists, `df.shape`, index values, and value precision); a mere `print` of both files is not such a step.
  3. Check whether any transformation is justified by the template or invented by the agent — rounding level, date-label styling, 0- vs 1-based period numbering, index naming.
  4. Check the final answer for a stated verification that the output matches the template cell-for-cell in structure; absence of it, or verification only against the agent's own file, is a violation.
- **Discriminator**: Fine if the script (or a verification run) reads the template and compares headers/index/shape/precision programmatically and any remaining differences are only in the numeric values being computed; a violation if format decisions rest on visual inspection or unstated assumptions, or if verification re-reads only the agent's own output.
- **Consequence**: The grader's file comparison fails (wrong index labels, extra/missing columns, or over-rounded values), scoring the result file as WRONG even when the underlying aggregation logic is plausible.
960Required deliverable file is never written at the exact specified name/pathtaskda-code
Applies when
task -- the task names a specific output artifact (e.g., a results file with a given filename and column format) that the scripts are supposed to produce.
Pattern
The scripts generate many intermediate/variant outputs under ad-hoc names (or only print/return results inline), and no code path writes the single exactly-named file at the expected location; the agent pastes rows into its chat answer as if that satisfied the deliverable.
Detection procedure
  1. Read the task and note the exact required output filename, location, header, and columns.
  2. Grep every script for write calls (to_csv, open(...,'w'), savetxt) and list all output paths produced.
  3. Check that at least one write targets the required filename/path verbatim, that it is executed unconditionally at the end of the pipeline (not inside a branch or after a truncated/incomplete script), and that its columns/header match the specified format and row count.
  4. Confirm the reported answer is that file's content, not a separately assembled or partial listing.
Discriminator
A real violation is when no write matches the required name/path (only variants like _v2.csv, _rf.csv, or a copy step that is never run), or the final script is incomplete/errors before the write. It is fine if the required file is written under a differently named variable/model choice as long as one write uses the exact required path and format; extra auxiliary files are harmless.
Consequence
The grader looks for the specified artifact and reports it as missing/wrong, scoring zero regardless of how good the underlying model or metric was.
id aafe9bbc8819 · mined from da-code dacode-ml-competition-006@s18
raw text (what the judge reads)
### Required deliverable file is never written at the exact specified name/path
- **Applies when**: `task` -- the task names a specific output artifact (e.g., a results file with a given filename and column format) that the scripts are supposed to produce.
- **Pattern**: The scripts generate many intermediate/variant outputs under ad-hoc names (or only print/return results inline), and no code path writes the single exactly-named file at the expected location; the agent pastes rows into its chat answer as if that satisfied the deliverable.
- **Detection procedure**:
  1. Read the task and note the exact required output filename, location, header, and columns.
  2. Grep every script for write calls (`to_csv`, `open(...,'w')`, `savetxt`) and list all output paths produced.
  3. Check that at least one write targets the required filename/path verbatim, that it is executed unconditionally at the end of the pipeline (not inside a branch or after a truncated/incomplete script), and that its columns/header match the specified format and row count.
  4. Confirm the reported answer is that file's content, not a separately assembled or partial listing.
- **Discriminator**: A real violation is when no write matches the required name/path (only variants like `*_v2.csv`, `*_rf.csv`, or a copy step that is never run), or the final script is incomplete/errors before the write. It is fine if the required file is written under a differently named variable/model choice as long as one write uses the exact required path and format; extra auxiliary files are harmless.
- **Consequence**: The grader looks for the specified artifact and reports it as missing/wrong, scoring zero regardless of how good the underlying model or metric was.
961Grouping by "missingness" without validating the null mask against sentinel/placeholder valuestaskinfiagent-dabench
Applies when
task -- the task asks to split rows into "missing vs. present" (or any boolean/filter-defined) groups on one field and compare a statistic of another field between the two groups.
Pattern
The attempt defines the group mask with a single naive call (e.g. isnull() / notnull(), or a dropna on the whole frame) without checking how missingness is actually encoded in the loaded data — empty strings, "NA", "None", "-", whitespace, 0, or a stringified null get counted in the wrong group; or extra rows are silently dropped because another column had NaNs — so both group means are shifted slightly yet look plausible.
Detection procedure
  1. Read the task to fix the intended partition and the statistic requested for each side.
  2. In the script, locate where the mask is built and where any row-dropping occurs (dropna, filtering, merges, read_* with na_values/keep_default_na, dtype coercion) and check whether the mask column is read as text.
  3. Confirm the script prints the row count of each group and that the two counts sum exactly to the total row count of the raw file, and that the aggregated column's own NaNs are handled once and explicitly.
  4. Check that the script inspects the unique/most frequent raw values of the grouping column (or the counts of candidate placeholders) to prove no placeholder is being mis-classified.
Discriminator
A real violation is a mask built with no evidence about the column's raw encoding and no group-count reconciliation; it is fine if the script shows the group sizes summing to the full row count and demonstrates (via value counts or explicit na_values) that placeholders were mapped to the intended group.
Consequence
Group means (and the p-value) are computed on a slightly wrong partition or a truncated subset, so the reported means differ from the expected values by a few percent and every numeric check fails even though the analysis "looks" right.
id 6c779fb89824 · mined from infiagent-dabench dabench-297@s18
raw text (what the judge reads)
### Grouping by "missingness" without validating the null mask against sentinel/placeholder values
- **Applies when**: `task` -- the task asks to split rows into "missing vs. present" (or any boolean/filter-defined) groups on one field and compare a statistic of another field between the two groups.
- **Pattern**: The attempt defines the group mask with a single naive call (e.g. `isnull()` / `notnull()`, or a dropna on the whole frame) without checking how missingness is actually encoded in the loaded data — empty strings, `"NA"`, `"None"`, `"-"`, whitespace, `0`, or a stringified null get counted in the wrong group; or extra rows are silently dropped because another column had NaNs — so both group means are shifted slightly yet look plausible.
- **Detection procedure**:
  1. Read the task to fix the intended partition and the statistic requested for each side.
  2. In the script, locate where the mask is built and where any row-dropping occurs (`dropna`, filtering, merges, `read_*` with `na_values`/`keep_default_na`, dtype coercion) and check whether the mask column is read as text.
  3. Confirm the script prints the row count of each group and that the two counts sum exactly to the total row count of the raw file, and that the aggregated column's own NaNs are handled once and explicitly.
  4. Check that the script inspects the unique/most frequent raw values of the grouping column (or the counts of candidate placeholders) to prove no placeholder is being mis-classified.
- **Discriminator**: A real violation is a mask built with no evidence about the column's raw encoding and no group-count reconciliation; it is fine if the script shows the group sizes summing to the full row count and demonstrates (via value counts or explicit `na_values`) that placeholders were mapped to the intended group.
- **Consequence**: Group means (and the p-value) are computed on a slightly wrong partition or a truncated subset, so the reported means differ from the expected values by a few percent and every numeric check fails even though the analysis "looks" right.
962Missing required output artifacts and self-invented processing rulestaskda-code
Applies when
task -- the task points to an external spec/guidance file and expects a set of saved deliverables (figures, serialized numbers, structured plot/metric dumps) plus derived categories or filters.
Pattern
The agent never loads or quotes the referenced guidance, substitutes its own filtering thresholds and category-mapping heuristics, and writes only the one artifact it happened to think of (e.g. the image), skipping the other required saved outputs; the final answer narrates numbers instead of confirming each deliverable exists.
Detection procedure
  1. From the task text, list every deliverable name/format the grader could check and every rule that is delegated to an external spec (thresholds, groupings, ordering, colors, sizes).
  2. Grep the scripts for a read/parse of that spec file and for a save/dump call per deliverable in the list.
  3. Flag if any deliverable has no corresponding write call, or if any grouping/filter constant appears as a hard-coded guess (docstring rules, timedelta(days=180), keyword substring matching) with no traceable source in the spec or data.
  4. Check whether the answer asserts the artifact set is complete without listing all of them.
Discriminator
A real violation is inventing definitions or omitting an artifact the task/spec names; it is fine if the spec's rules were actually read and quoted, or if a genuinely ambiguous choice is stated explicitly and all named artifacts are written.
Consequence
Grader file checks fail as WRONG/MISSING for the unwritten artifacts, and the produced figure/counts disagree with reference values because the category and filter definitions differ from the specified ones.
id fe4d79534134 · mined from da-code dacode-plot-pie-005@s18
raw text (what the judge reads)
### Missing required output artifacts and self-invented processing rules
- **Applies when**: `task` -- the task points to an external spec/guidance file and expects a set of saved deliverables (figures, serialized numbers, structured plot/metric dumps) plus derived categories or filters.
- **Pattern**: The agent never loads or quotes the referenced guidance, substitutes its own filtering thresholds and category-mapping heuristics, and writes only the one artifact it happened to think of (e.g. the image), skipping the other required saved outputs; the final answer narrates numbers instead of confirming each deliverable exists.
- **Detection procedure**:
  1. From the task text, list every deliverable name/format the grader could check and every rule that is delegated to an external spec (thresholds, groupings, ordering, colors, sizes).
  2. Grep the scripts for a read/parse of that spec file and for a save/dump call per deliverable in the list.
  3. Flag if any deliverable has no corresponding write call, or if any grouping/filter constant appears as a hard-coded guess (docstring rules, `timedelta(days=180)`, keyword substring matching) with no traceable source in the spec or data.
  4. Check whether the answer asserts the artifact set is complete without listing all of them.
- **Discriminator**: A real violation is inventing definitions or omitting an artifact the task/spec names; it is fine if the spec's rules were actually read and quoted, or if a genuinely ambiguous choice is stated explicitly and all named artifacts are written.
- **Consequence**: Grader file checks fail as WRONG/MISSING for the unwritten artifacts, and the produced figure/counts disagree with reference values because the category and filter definitions differ from the specified ones.
963Prescribed missing-value handling silently replaced (row loss / coercion) and error never sanity-checked against target scaletaskinfiagent-dabench
Applies when
task -- the task fixes a preprocessing recipe (e.g., impute specified columns with a stated statistic) and asks for a single error metric from a train/test split.
Pattern
The attempt deviates from the mandated imputation — dropping rows with nulls, imputing only some of the named columns, imputing after subsetting/splitting, or letting non-numeric-looking values (currency/unit strings, sentinels like 0/-999/"unknown") be coerced or silently discarded — so the modeled rows and the fitted relationship differ from the specification; it then reports the resulting error without checking whether its magnitude is plausible relative to the target's own variance/range.
Detection procedure
  1. From the task, list the exact preprocessing steps and the columns they apply to, and note the requested output (metric, split fraction, rounding).
  2. In the scripts, trace each named column from load to model input: check dtype conversion, any dropna/filtering/errors='coerce', and whether the imputation statistic is computed on the intended data and applied to all required columns before splitting.
  3. Confirm the script prints row counts/shapes before and after preprocessing (and NaN counts) so an unintended row loss or all-NaN column would surface; absent such prints, treat the pipeline as unverified.
  4. Compare the reported error to a trivial baseline the script should print (variance/MSE of predicting the target mean); if the reported MSE is of the same order as or larger than that baseline, the pipeline is suspect and must be rejected.
Discriminator
A genuine violation is a pipeline whose effective row set, column set, or imputation differs from the stated recipe, or whose error exceeds/approaches the mean-predictor baseline with no explanation; it is not a violation if the recipe is followed exactly and the residual error is simply large because the predictors are weak — provided the script demonstrates row counts, dtypes, and the baseline comparison.
Consequence
The regression is fit on a different sample or scale than intended, so the reported MSE is off by a large factor from the reference value and the answer is graded wrong even though the format matches.
id 1e1982683650 · mined from infiagent-dabench dabench-432@s18
raw text (what the judge reads)
### Prescribed missing-value handling silently replaced (row loss / coercion) and error never sanity-checked against target scale
- **Applies when**: `task` -- the task fixes a preprocessing recipe (e.g., impute specified columns with a stated statistic) and asks for a single error metric from a train/test split.
- **Pattern**: The attempt deviates from the mandated imputation — dropping rows with nulls, imputing only some of the named columns, imputing after subsetting/splitting, or letting non-numeric-looking values (currency/unit strings, sentinels like 0/-999/"unknown") be coerced or silently discarded — so the modeled rows and the fitted relationship differ from the specification; it then reports the resulting error without checking whether its magnitude is plausible relative to the target's own variance/range.
- **Detection procedure**:
  1. From the task, list the exact preprocessing steps and the columns they apply to, and note the requested output (metric, split fraction, rounding).
  2. In the scripts, trace each named column from load to model input: check dtype conversion, any `dropna`/filtering/`errors='coerce'`, and whether the imputation statistic is computed on the intended data and applied to *all* required columns before splitting.
  3. Confirm the script prints row counts/shapes before and after preprocessing (and NaN counts) so an unintended row loss or all-NaN column would surface; absent such prints, treat the pipeline as unverified.
  4. Compare the reported error to a trivial baseline the script should print (variance/MSE of predicting the target mean); if the reported MSE is of the same order as or larger than that baseline, the pipeline is suspect and must be rejected.
- **Discriminator**: A genuine violation is a pipeline whose effective row set, column set, or imputation differs from the stated recipe, or whose error exceeds/approaches the mean-predictor baseline with no explanation; it is *not* a violation if the recipe is followed exactly and the residual error is simply large because the predictors are weak — provided the script demonstrates row counts, dtypes, and the baseline comparison.
- **Consequence**: The regression is fit on a different sample or scale than intended, so the reported MSE is off by a large factor from the reference value and the answer is graded wrong even though the format matches.
964Undocumented row-selection / missing-value handling for a summary statistic reported to fixed precisiontaskinfiagent-dabench
Applies when
task -- the task asks for a single numeric statistic (correlation, mean, coefficient, metric) rounded to a fixed number of decimals, computed over two or more columns of a loaded table.
Pattern
The attempt loads the data and calls a one-liner statistic function without ever stating or checking which rows entered the computation — no report of row counts before/after coercion, no explicit decision about non-numeric/blank/sentinel values, no saved script — so the value silently reflects a different sample than the intended full valid set and lands one unit off in the last reported digit.
Detection procedure
  1. Read the task and note the required rounding precision and which columns/rows are in scope (all rows unless a filter is stated).
  2. In the scripts, look for explicit dtype coercion of the involved columns and an explicit, stated policy for missing/non-numeric entries (e.g., pairwise drop of rows missing either column), plus a printed count of rows actually used.
  3. Check that the printed N and the statistic are reported together, and that the reported value is the final rounded statistic (not an intermediate or a value from a subset/sample/head of the data).
  4. If no script exists or step 2's evidence is absent, treat the number as unverified — especially when the answer is quoted at a precision finer than the demonstrated robustness (e.g., two decimals with no N check, or a p-value reported as an exact 0).
Discriminator
A real violation is an attempt that cannot show which rows were used or whether coercion/dropping occurred; it is fine if the script prints row counts and confirms the statistic is stable under the documented missing-data policy (and reports p-values in a defensible form rather than collapsing tiny values to 0 without noting it).
Consequence
The rounded statistic differs from the reference in the last decimal (e.g., 0.53 vs 0.54), so the exact-match check on that field fails even though the qualitative conclusion is right.
id b82c8d3af8af · mined from infiagent-dabench dabench-300@s18
raw text (what the judge reads)
### Undocumented row-selection / missing-value handling for a summary statistic reported to fixed precision
- **Applies when**: `task` -- the task asks for a single numeric statistic (correlation, mean, coefficient, metric) rounded to a fixed number of decimals, computed over two or more columns of a loaded table.
- **Pattern**: The attempt loads the data and calls a one-liner statistic function without ever stating or checking which rows entered the computation — no report of row counts before/after coercion, no explicit decision about non-numeric/blank/sentinel values, no saved script — so the value silently reflects a different sample than the intended full valid set and lands one unit off in the last reported digit.
- **Detection procedure**:
  1. Read the task and note the required rounding precision and which columns/rows are in scope (all rows unless a filter is stated).
  2. In the scripts, look for explicit dtype coercion of the involved columns and an explicit, stated policy for missing/non-numeric entries (e.g., pairwise drop of rows missing either column), plus a printed count of rows actually used.
  3. Check that the printed N and the statistic are reported together, and that the reported value is the final rounded statistic (not an intermediate or a value from a subset/sample/head of the data).
  4. If no script exists or step 2's evidence is absent, treat the number as unverified — especially when the answer is quoted at a precision finer than the demonstrated robustness (e.g., two decimals with no N check, or a p-value reported as an exact 0).
- **Discriminator**: A real violation is an attempt that cannot show which rows were used or whether coercion/dropping occurred; it is fine if the script prints row counts and confirms the statistic is stable under the documented missing-data policy (and reports p-values in a defensible form rather than collapsing tiny values to 0 without noting it).
- **Consequence**: The rounded statistic differs from the reference in the last decimal (e.g., 0.53 vs 0.54), so the exact-match check on that field fails even though the qualitative conclusion is right.
965No held-out error estimate — accuracy asserted from distribution similarity alonetaskda-code
Applies when
task -- the task asks for predictions on a supplied unlabeled split, and the script fits a model on all labeled rows and writes predictions directly to the required output file.
Pattern
The attempt never scores the model on any labeled data it did not train on (no train/validation split, no cross-validation, no reported RMSE/MAE/R²). Instead it "validates" by comparing summary statistics (min/max/mean) of the predictions to the training target distribution, and declares success. Because distributional agreement is insensitive to per-row accuracy, gross problems (uninformative or mis-aligned encodings of high-cardinality categoricals, unusable feature handling, fitting on the wrong rows, or ignoring that the labeled source file may already contain the evaluation rows and thus permit a near-exact lookup/merge) go undetected, as do implausible predictions (e.g. values clipped at zero) that a real metric would expose.
Detection procedure
  1. Read the task to identify the predictive target and the unlabeled split whose predictions will be graded against hidden truth.
  2. Search the scripts for any split of the labeled data (train_test_split, KFold, cross_val_score) followed by a metric call on the held-out part; note whether any error number is computed at all, and whether the labeled file is checked for overlap/duplicate keys with the evaluation split.
  3. Read the answer for the evidence offered: is it a quantitative generalization error, or only prediction-vs-training summary statistics ("mean matches", "range matches")?
  4. Flag if no out-of-sample metric exists, or if the only quality claim is distributional, or if the reported prediction range includes values the target cannot take.
Discriminator
A fine attempt reports at least one out-of-sample error figure (holdout or CV) on the labeled data and, ideally, compares it to a trivial baseline; internal early-stopping on a validation fraction without ever reporting that score, or purely distributional comparisons, do not count. It is also fine to have no metric only if the task explicitly forbids it — merely writing the correct file shape is not evidence of correctness.
Consequence
The submitted file has the right column name and row count but per-row predictions far off the hidden truth, so the grader's error/accuracy threshold on the output file fails while the agent's self-report claims success.
id ae5294017158 · mined from da-code dacode-ml-regression-014@s18
raw text (what the judge reads)
### No held-out error estimate — accuracy asserted from distribution similarity alone
- **Applies when**: `task` -- the task asks for predictions on a supplied unlabeled split, and the script fits a model on all labeled rows and writes predictions directly to the required output file.
- **Pattern**: The attempt never scores the model on any labeled data it did not train on (no train/validation split, no cross-validation, no reported RMSE/MAE/R²). Instead it "validates" by comparing summary statistics (min/max/mean) of the predictions to the training target distribution, and declares success. Because distributional agreement is insensitive to per-row accuracy, gross problems (uninformative or mis-aligned encodings of high-cardinality categoricals, unusable feature handling, fitting on the wrong rows, or ignoring that the labeled source file may already contain the evaluation rows and thus permit a near-exact lookup/merge) go undetected, as do implausible predictions (e.g. values clipped at zero) that a real metric would expose.
- **Detection procedure**:
  1. Read the task to identify the predictive target and the unlabeled split whose predictions will be graded against hidden truth.
  2. Search the scripts for any split of the labeled data (`train_test_split`, KFold, cross_val_score) followed by a metric call on the held-out part; note whether any error number is computed at all, and whether the labeled file is checked for overlap/duplicate keys with the evaluation split.
  3. Read the answer for the evidence offered: is it a quantitative generalization error, or only prediction-vs-training summary statistics ("mean matches", "range matches")?
  4. Flag if no out-of-sample metric exists, or if the only quality claim is distributional, or if the reported prediction range includes values the target cannot take.
- **Discriminator**: A fine attempt reports at least one out-of-sample error figure (holdout or CV) on the labeled data and, ideally, compares it to a trivial baseline; internal early-stopping on a validation fraction without ever reporting that score, or purely distributional comparisons, do not count. It is also fine to have no metric only if the task explicitly forbids it — merely writing the correct file shape is not evidence of correctness.
- **Consequence**: The submitted file has the right column name and row count but per-row predictions far off the hidden truth, so the grader's error/accuracy threshold on the output file fails while the agent's self-report claims success.
966Ignoring an external plot/output spec file and emitting only a subset of required artifactstaskda-code
Applies when
task -- the task points to a separate configuration/specification file (e.g. a YAML/JSON of plotting or formatting guidelines) and/or implies saved artifacts beyond the single file named in the prompt.
Pattern
The attempt computes the analytic result and writes one obvious output (an image), but never parses the referenced spec file, never applies its required options (title, labels, colors, ordering, categories, figure size, serialized plot data), and never persists the companion artifacts (the plot's data/spec dump and the numeric array underlying the chart) that graders check.
Detection procedure
  1. Read the task and list every file it names or implies as an output, plus every external spec/config file it says to adhere to.
  2. Search the scripts for a load of that spec file and for each of its keys being actually used when building the figure; search for a write of every listed output path.
  3. Compare with the final answer/artifact list: if the spec is never read, or any required output file is absent, flag it.
  4. Check that the values written into the saved data artifact are exactly the quantities the chart displays (same category set, same ordering, same units/normalization) rather than an ad-hoc summary.
Discriminator
A real violation is a missing spec load or missing artifact file; it is not a violation if the script reads the spec, applies each directive, and writes all outputs but merely uses a different internal variable naming or helper function to do so. Also fine if a "missing" file is genuinely not requested anywhere in the task or spec.
Consequence
Grader checks for the unwritten/misformatted artifacts fail outright (file WRONG/MISSING), so the run scores zero even when the headline statistic (the selected group and its value) happens to be right.
id 738b5ef94419 · mined from da-code dacode-plot-pie-008@s18
raw text (what the judge reads)
### Ignoring an external plot/output spec file and emitting only a subset of required artifacts
- **Applies when**: `task` -- the task points to a separate configuration/specification file (e.g. a YAML/JSON of plotting or formatting guidelines) and/or implies saved artifacts beyond the single file named in the prompt.
- **Pattern**: The attempt computes the analytic result and writes one obvious output (an image), but never parses the referenced spec file, never applies its required options (title, labels, colors, ordering, categories, figure size, serialized plot data), and never persists the companion artifacts (the plot's data/spec dump and the numeric array underlying the chart) that graders check.
- **Detection procedure**:
  1. Read the task and list every file it names or implies as an output, plus every external spec/config file it says to adhere to.
  2. Search the scripts for a load of that spec file and for each of its keys being actually used when building the figure; search for a write of every listed output path.
  3. Compare with the final answer/artifact list: if the spec is never read, or any required output file is absent, flag it.
  4. Check that the values written into the saved data artifact are exactly the quantities the chart displays (same category set, same ordering, same units/normalization) rather than an ad-hoc summary.
- **Discriminator**: A real violation is a missing spec load or missing artifact file; it is *not* a violation if the script reads the spec, applies each directive, and writes all outputs but merely uses a different internal variable naming or helper function to do so. Also fine if a "missing" file is genuinely not requested anywhere in the task or spec.
- **Consequence**: Grader checks for the unwritten/misformatted artifacts fail outright (file WRONG/MISSING), so the run scores zero even when the headline statistic (the selected group and its value) happens to be right.
967Undocumented preprocessing/decision-rule choices in a regression-as-classifier pipelinetaskinfiagent-dabench
Applies when
task -- a continuous-output model is mandated for a binary target, the feature set contains missing values or non-numeric columns, and a single accuracy-style score with a fixed split seed is reported.
Pattern
The attempt silently picks non-default preprocessing or output-conversion choices — dropping rows with missing values (or dropping a whole column) instead of imputing, encoding categoricals in a way that changes column count, or converting continuous predictions to labels with an ad hoc rule (rounding on a shifted scale, non-0.5 cutoff, sign test) — without checking that row counts, split sizes, and label range remain as the task implies, so the reported score comes from a different evaluation sample or decision rule than the intended one.
Detection procedure
  1. Read the task for the mandated model type, encoding scheme, split fractions/seed, and metric; note that a continuous predictor needs an explicit, standard threshold to yield labels.
  2. In the scripts, locate every operation that can change row count (dropna, filtering, merges) or drop features, and confirm that missing values are instead imputed so the full dataset is split and scored; record the pre-split and post-split row counts.
  3. Check that encoding is applied to the full data before splitting (consistent columns for train/test) and that predictions are converted to class labels with a defensible 0.5 cutoff on the 0/1 scale, not by an unexplained rounding/sign rule.
  4. Verify the reported number is test-set accuracy (not train accuracy, R², or a CV mean), lies in [0,1], is rounded as requested, and matches the required answer template.
Discriminator
A violation is a choice that silently shrinks or alters the evaluation set / decision rule with no justification or sanity check (e.g. hundreds of rows removed, a feature dropped, cutoff other than 0.5). It is fine if rows are removed because the task explicitly says to filter, or if imputation/threshold choices are stated and the resulting split sizes and label range are verified.
Consequence
Accuracy is computed on a different test subset or with a different labeling rule, giving a value off by a few points (e.g. 0.76 vs. 0.78) and an exact-match grader failure.
id 18d87e579d5a · mined from infiagent-dabench dabench-7@s18
raw text (what the judge reads)
### Undocumented preprocessing/decision-rule choices in a regression-as-classifier pipeline
- **Applies when**: `task` -- a continuous-output model is mandated for a binary target, the feature set contains missing values or non-numeric columns, and a single accuracy-style score with a fixed split seed is reported.
- **Pattern**: The attempt silently picks non-default preprocessing or output-conversion choices — dropping rows with missing values (or dropping a whole column) instead of imputing, encoding categoricals in a way that changes column count, or converting continuous predictions to labels with an ad hoc rule (rounding on a shifted scale, non-0.5 cutoff, sign test) — without checking that row counts, split sizes, and label range remain as the task implies, so the reported score comes from a different evaluation sample or decision rule than the intended one.
- **Detection procedure**:
  1. Read the task for the mandated model type, encoding scheme, split fractions/seed, and metric; note that a continuous predictor needs an explicit, standard threshold to yield labels.
  2. In the scripts, locate every operation that can change row count (`dropna`, filtering, merges) or drop features, and confirm that missing values are instead imputed so the full dataset is split and scored; record the pre-split and post-split row counts.
  3. Check that encoding is applied to the full data before splitting (consistent columns for train/test) and that predictions are converted to class labels with a defensible 0.5 cutoff on the 0/1 scale, not by an unexplained rounding/sign rule.
  4. Verify the reported number is test-set accuracy (not train accuracy, R², or a CV mean), lies in [0,1], is rounded as requested, and matches the required answer template.
- **Discriminator**: A violation is a choice that silently shrinks or alters the evaluation set / decision rule with no justification or sanity check (e.g. hundreds of rows removed, a feature dropped, cutoff other than 0.5). It is fine if rows are removed because the task explicitly says to filter, or if imputation/threshold choices are stated and the resulting split sizes and label range are verified.
- **Consequence**: Accuracy is computed on a different test subset or with a different labeling rule, giving a value off by a few points (e.g. 0.76 vs. 0.78) and an exact-match grader failure.
968Ships a classifier with no held-out validation and no check of predicted class balancetaskda-code
Applies when
task -- the script fits a predictive model on a labeled training file and writes predictions for an unlabeled test file that will be scored against hidden ground truth.
Pattern
The script fits one model on 100% of the training data, immediately calls predict/threshold at the default 0.5, and writes the output file — with no train/validation split, cross-validation, or comparison of the predicted positive rate against the training base rate. The answer reports only counts and methodology, never an estimated score on held-out data, so an under-predicting or mis-calibrated model (common with imbalanced classes and a default threshold) goes undetected.
Detection procedure
  1. Read the task to see what will be graded (a predictions file scored by some accuracy/F1-type metric) and whether the target is imbalanced.
  2. Scan the script for any held-out estimate of quality: a split, cross_val_score, or metric printed on labels the model did not train on. If the only printed diagnostics are shapes, dtypes, and prediction counts, the attempt is unvalidated.
  3. Compare the reported positive-prediction rate on test with the training positive rate (both are usually printed). A large one-sided gap (e.g., predicted rate roughly half the training prior) signals a default-threshold/imbalance problem that was never examined or corrected.
  4. Check the answer for any claim of expected performance; absence of one, or a claim not backed by held-out numbers, confirms the gap.
Discriminator
Not a violation if the script does hold out data (or CV) and reports a metric, and either the predicted positive rate matches the training prior or the deviation is explicitly justified by a threshold tuned on validation for the target metric. It is a violation when the only evidence of correctness is "the file was written with the right shape/columns".
Consequence
The grader compares predictions to true labels and the score falls below the acceptance threshold (typically from mass under-prediction of the minority class, i.e., low recall/F1), so the expected output file is marked WRONG even though its format is fine.
id c23092da973b · mined from da-code dacode-ml-binary-016@s18
raw text (what the judge reads)
### Ships a classifier with no held-out validation and no check of predicted class balance
- **Applies when**: `task` -- the script fits a predictive model on a labeled training file and writes predictions for an unlabeled test file that will be scored against hidden ground truth.
- **Pattern**: The script fits one model on 100% of the training data, immediately calls `predict`/threshold at the default 0.5, and writes the output file — with no train/validation split, cross-validation, or comparison of the predicted positive rate against the training base rate. The answer reports only counts and methodology, never an estimated score on held-out data, so an under-predicting or mis-calibrated model (common with imbalanced classes and a default threshold) goes undetected.
- **Detection procedure**:
  1. Read the task to see what will be graded (a predictions file scored by some accuracy/F1-type metric) and whether the target is imbalanced.
  2. Scan the script for any held-out estimate of quality: a split, `cross_val_score`, or metric printed on labels the model did not train on. If the only printed diagnostics are shapes, dtypes, and prediction counts, the attempt is unvalidated.
  3. Compare the reported positive-prediction rate on test with the training positive rate (both are usually printed). A large one-sided gap (e.g., predicted rate roughly half the training prior) signals a default-threshold/imbalance problem that was never examined or corrected.
  4. Check the answer for any claim of expected performance; absence of one, or a claim not backed by held-out numbers, confirms the gap.
- **Discriminator**: Not a violation if the script does hold out data (or CV) and reports a metric, and either the predicted positive rate matches the training prior or the deviation is explicitly justified by a threshold tuned on validation for the target metric. It *is* a violation when the only evidence of correctness is "the file was written with the right shape/columns".
- **Consequence**: The grader compares predictions to true labels and the score falls below the acceptance threshold (typically from mass under-prediction of the minority class, i.e., low recall/F1), so the expected output file is marked WRONG even though its format is fine.
969Unverified data exclusion (aggressive "outlier"/subset filtering) before computing a summary statistictaskinfiagent-dabench
Applies when
task -- The task asks for a single descriptive statistic over a column, with vague or boilerplate wording about ignoring missing values/outliers, and the script applies filtering, deduplication, or file/row selection before aggregating.
Pattern
The attempt computes the statistic on a reduced subset — e.g. it applies an invented outlier rule (z-score, IQR, percentile clipping), keeps only certain categories/files/rows, or drops rows via dropna() on the whole frame instead of the target column — and reports that number as the requested statistic without comparing it to the straightforward full-column value or reporting the row count used.
Detection procedure
  1. In the task, note exactly which population the statistic is defined over (all non-missing observations of the column) and whether any concrete filtering rule (thresholds, units, groups) is actually specified.
  2. In the scripts, list every operation that changes the row set before aggregation (filters, boolean masks, drop_duplicates, subframe/file selection, frame-wide dropna, clipping) and check whether each is explicitly mandated by the task.
  3. Check whether the script prints the number of observations used before and after filtering, and whether it also computes the unfiltered (missing-values-only-dropped) statistic as a baseline.
  4. Compare the reported answer to that baseline; if they differ materially and only the filtered value is reported with no justification, flag the attempt.
Discriminator
A real violation is discretionary/self-invented data removal, or removal justified only by generic instruction boilerplate, that measurably shifts the reported value and is never cross-checked. A look-alike that is fine is removal of genuinely invalid records under a rule stated in the task (explicit range, stated unit, named group), or missing-value handling restricted to the target column, with counts logged and the effect on the statistic shown to be negligible or explicitly required.
Consequence
The reported number is a statistic over a different population than requested, so it misses the expected value beyond tolerance and the grader marks the single-value check wrong even though the code runs cleanly.
id 4e24a86fed0a · mined from infiagent-dabench dabench-320@s18
raw text (what the judge reads)
### Unverified data exclusion (aggressive "outlier"/subset filtering) before computing a summary statistic
- **Applies when**: `task` -- The task asks for a single descriptive statistic over a column, with vague or boilerplate wording about ignoring missing values/outliers, and the script applies filtering, deduplication, or file/row selection before aggregating.
- **Pattern**: The attempt computes the statistic on a reduced subset — e.g. it applies an invented outlier rule (z-score, IQR, percentile clipping), keeps only certain categories/files/rows, or drops rows via `dropna()` on the whole frame instead of the target column — and reports that number as the requested statistic without comparing it to the straightforward full-column value or reporting the row count used.
- **Detection procedure**:
  1. In the task, note exactly which population the statistic is defined over (all non-missing observations of the column) and whether any concrete filtering rule (thresholds, units, groups) is actually specified.
  2. In the scripts, list every operation that changes the row set before aggregation (filters, boolean masks, `drop_duplicates`, subframe/file selection, frame-wide `dropna`, clipping) and check whether each is explicitly mandated by the task.
  3. Check whether the script prints the number of observations used before and after filtering, and whether it also computes the unfiltered (missing-values-only-dropped) statistic as a baseline.
  4. Compare the reported answer to that baseline; if they differ materially and only the filtered value is reported with no justification, flag the attempt.
- **Discriminator**: A real violation is discretionary/self-invented data removal, or removal justified only by generic instruction boilerplate, that measurably shifts the reported value and is never cross-checked. A look-alike that is fine is removal of genuinely invalid records under a rule stated in the task (explicit range, stated unit, named group), or missing-value handling restricted to the target column, with counts logged and the effect on the statistic shown to be negligible or explicitly required.
- **Consequence**: The reported number is a statistic over a different population than requested, so it misses the expected value beyond tolerance and the grader marks the single-value check wrong even though the code runs cleanly.
970Extreme-value (min/max) answers taken from an unverified, unparsed numeric columntaskda-code
Applies when
task -- the task asks to impute missing values and then report the row(s) attaining the maximum or minimum of a numeric measure read from a raw tabular file.
Pattern
The attempt loads the file with defaults and calls mean()/idxmax()/idxmin()/sort_values() on a column that is actually object/string dtype (thousands separators, units, %, currency symbols, placeholder strings like "N/A", "-", or blanks). String columns sort lexicographically and imputation silently skips them, so the reported extremes are artifacts of text ordering rather than magnitude; no shape/dtype/range sanity check or plausibility check against domain knowledge is performed, and the required output file is never validated.
Detection procedure
  1. From the task, note which column drives the answer and whether the imputation step must actually affect it.
  2. In the scripts, check that the column is explicitly coerced to numeric (strip separators/symbols, pd.to_numeric(..., errors='coerce')) and that dtypes and non-null counts are printed before the mean fill and before the argmax/argmin.
  3. Check that the extreme selection is done on the numeric column (not on a text sort or on a different/intermediate column) and that ties are handled consistently with the list-valued output format.
  4. Compare the reported extreme values/rows with a printed top-5 and bottom-5 of the numeric column and with a common-sense range for the quantity; confirm the answer was written to the requested output file in the requested JSON shape.
Discriminator
A real violation is when no dtype/parse verification exists and the column plausibly contains formatted text or placeholders — the reported extreme is implausible or inconsistent with the quantity's known scale. It is not a violation if the script demonstrably converts the column to a numeric dtype (or the source is already numeric) and prints the sorted head/tail supporting the chosen rows, even if the final names are surprising.
Consequence
The grader compares against the true numeric extremes and marks the answer wrong (or the expected result file missing), since the reported country/row is the lexicographic or unimputed extreme rather than the true max/min.
id a00c6a902bcf · mined from da-code dacode-di-text-001@s18
raw text (what the judge reads)
### Extreme-value (min/max) answers taken from an unverified, unparsed numeric column
- **Applies when**: `task` -- the task asks to impute missing values and then report the row(s) attaining the maximum or minimum of a numeric measure read from a raw tabular file.
- **Pattern**: The attempt loads the file with defaults and calls `mean()`/`idxmax()`/`idxmin()`/`sort_values()` on a column that is actually object/string dtype (thousands separators, units, `%`, currency symbols, placeholder strings like "N/A", "-", or blanks). String columns sort lexicographically and imputation silently skips them, so the reported extremes are artifacts of text ordering rather than magnitude; no shape/dtype/range sanity check or plausibility check against domain knowledge is performed, and the required output file is never validated.
- **Detection procedure**:
  1. From the task, note which column drives the answer and whether the imputation step must actually affect it.
  2. In the scripts, check that the column is explicitly coerced to numeric (strip separators/symbols, `pd.to_numeric(..., errors='coerce')`) and that dtypes and non-null counts are printed *before* the mean fill and *before* the argmax/argmin.
  3. Check that the extreme selection is done on the numeric column (not on a text sort or on a different/intermediate column) and that ties are handled consistently with the list-valued output format.
  4. Compare the reported extreme values/rows with a printed top-5 and bottom-5 of the numeric column and with a common-sense range for the quantity; confirm the answer was written to the requested output file in the requested JSON shape.
- **Discriminator**: A real violation is when no dtype/parse verification exists and the column plausibly contains formatted text or placeholders — the reported extreme is implausible or inconsistent with the quantity's known scale. It is *not* a violation if the script demonstrably converts the column to a numeric dtype (or the source is already numeric) and prints the sorted head/tail supporting the chosen rows, even if the final names are surprising.
- **Consequence**: The grader compares against the true numeric extremes and marks the answer wrong (or the expected result file missing), since the reported country/row is the lexicographic or unimputed extreme rather than the true max/min.
971Dropping requested derived columns from the saved output filetaskda-code
Applies when
task -- the task asks for a results file that must include several derived quantities (e.g., component scores, a combined score, a group/segment label, and a level/class label) for each entity.
Pattern
The script computes all the intermediate and final quantities correctly in a working DataFrame, writes the full table to auxiliary/scratch files, but then subsets the deliverable file down to just the identifier plus one final label, silently discarding the other quantities the task explicitly named (and/or renaming them to non-obvious names).
Detection procedure
  1. From the task statement, list every quantity the deliverable is required to contain ("including X and Y" implies both X and Y plus the per-entity scores that define them).
  2. In the script, find the line(s) that write the deliverable filename and note exactly which columns are selected and what they are called.
  3. Compare that column list against the list from step 1; flag any required quantity that exists in memory but is not in the written file, or whose name/format was altered.
  4. Check the reported answer for an admission such as "final output columns: id + one label" while extra columns were dumped to other files.
Discriminator
A real violation is when a required quantity was computed but excluded from (or renamed unrecognizably in) the graded file; it is not a violation if the task genuinely asks only for the single label, or if the extra columns are present under reasonable equivalent names, or if truly optional diagnostics are omitted.
Consequence
The graded file fails column/content comparison against the expected output ("WRONG/MISSING") even though the underlying computation may have been right, yielding 0 checks passed.
id 049b387a3da0 · mined from da-code dacode-dm-csv-052@s18
raw text (what the judge reads)
### Dropping requested derived columns from the saved output file
- **Applies when**: `task` -- the task asks for a results file that must include several derived quantities (e.g., component scores, a combined score, a group/segment label, and a level/class label) for each entity.
- **Pattern**: The script computes all the intermediate and final quantities correctly in a working DataFrame, writes the full table to auxiliary/scratch files, but then subsets the deliverable file down to just the identifier plus one final label, silently discarding the other quantities the task explicitly named (and/or renaming them to non-obvious names).
- **Detection procedure**:
  1. From the task statement, list every quantity the deliverable is required to contain ("including X and Y" implies both X and Y plus the per-entity scores that define them).
  2. In the script, find the line(s) that write the deliverable filename and note exactly which columns are selected and what they are called.
  3. Compare that column list against the list from step 1; flag any required quantity that exists in memory but is not in the written file, or whose name/format was altered.
  4. Check the reported answer for an admission such as "final output columns: id + one label" while extra columns were dumped to other files.
- **Discriminator**: A real violation is when a required quantity was computed but excluded from (or renamed unrecognizably in) the graded file; it is not a violation if the task genuinely asks only for the single label, or if the extra columns are present under reasonable equivalent names, or if truly optional diagnostics are omitted.
- **Consequence**: The graded file fails column/content comparison against the expected output ("WRONG/MISSING") even though the underlying computation may have been right, yielding 0 checks passed.
972Fabricated or hardcoded input data instead of loading the provided filestaskda-code
Applies when
task -- the task references a supplied dataset (and/or a sample output file defining the format), and the scripts must read those files to produce the answer.
Pattern
The script embeds literal arrays/dicts of "the data" typed from memory or invented (often suspiciously round, uniform-length, e.g. exactly N rows per group), never opens the provided data or format-template files, and then reports statistics from this stand-in as if authoritative.
Detection procedure
  1. From the task/README, list every input artifact that must be consumed (data file(s), sample/template output file, any stated schema).
  2. Grep the scripts for file-reading calls (read_csv, open, load, package loaders) and check that each listed artifact is actually read; note anything replaced by inline literals.
  3. Cross-check inline literals against plausibility signals: are group sizes identical/round, values low-precision, counts absent from any file, no provenance comment or citation?
  4. Check the output-writing step: are column names/order/row order taken from the provided template file, or guessed?
Discriminator
Legitimate: small constants that are genuinely given in the task text (thresholds, seeds, bootstrap size), or literals derived programmatically from a file that is read elsewhere in the pipeline. Violation: the observations themselves — the rows the statistics are computed over — exist only in the script, with no read of the supplied dataset, or the output schema is invented while a template file was provided and never opened.
Consequence
Every reported statistic and the written output file are computed from the wrong sample (and possibly the wrong schema), so the graded file mismatches the expected values/format and scores 0 even though the code runs without error.
id c55a141deebf · mined from da-code dacode-data-sa-029@s18
raw text (what the judge reads)
### Fabricated or hardcoded input data instead of loading the provided files
- **Applies when**: `task` -- the task references a supplied dataset (and/or a sample output file defining the format), and the scripts must read those files to produce the answer.
- **Pattern**: The script embeds literal arrays/dicts of "the data" typed from memory or invented (often suspiciously round, uniform-length, e.g. exactly N rows per group), never opens the provided data or format-template files, and then reports statistics from this stand-in as if authoritative.
- **Detection procedure**:
  1. From the task/README, list every input artifact that must be consumed (data file(s), sample/template output file, any stated schema).
  2. Grep the scripts for file-reading calls (`read_csv`, `open`, `load`, package loaders) and check that each listed artifact is actually read; note anything replaced by inline literals.
  3. Cross-check inline literals against plausibility signals: are group sizes identical/round, values low-precision, counts absent from any file, no provenance comment or citation?
  4. Check the output-writing step: are column names/order/row order taken from the provided template file, or guessed?
- **Discriminator**: Legitimate: small constants that are genuinely given in the task text (thresholds, seeds, bootstrap size), or literals derived programmatically from a file that is read elsewhere in the pipeline. Violation: the observations themselves — the rows the statistics are computed over — exist only in the script, with no read of the supplied dataset, or the output schema is invented while a template file was provided and never opened.
- **Consequence**: Every reported statistic and the written output file are computed from the wrong sample (and possibly the wrong schema), so the graded file mismatches the expected values/format and scores 0 even though the code runs without error.
973Implausible statistic values after type coercion left uncheckedtaskda-code
Applies when
task -- the task requires casting text/mixed columns to numeric (or otherwise re-encoding) before computing a summary statistic, correlation, or metric that is then written to a required output file.
Pattern
The attempt coerces columns to numeric (often with silent errors='coerce', per-column parsing, or index/label misalignment), drops missing rows, computes the statistic, and reports it verbatim — with no check that the values are domain-plausible, that the retained row count is sensible, or that the output matches the provided sample format. Near-zero correlations (or otherwise degenerate/uniform values) among variables that should be related are accepted as a "finding" instead of treated as a bug signal.
Detection procedure
  1. From the task, note the expected value semantics (e.g., ratings on a bounded scale, variables that plausibly co-vary) and the required output shape/format from any sample file.
  2. In the scripts, check whether the numeric conversion is validated: how many values became NaN per column, what the resulting min/max/unique values look like, and whether row alignment is preserved after dropping/reindexing.
  3. In the reported answer, compare the numbers against the plausibility expectation and against the retained-row count; flag if all off-diagonal/aggregate values are suspiciously near zero, constant, or outside the natural range.
  4. Confirm the written file's header/index/orientation and rounding were verified against the sample format, not just assumed.
Discriminator
A genuine violation is when the scripts contain no post-conversion validation (dtype, value range, NaN counts, alignment) and the reported numbers are implausible or unexplained; it is fine if the agent explicitly verified parsed values and row counts and the weak/odd result is reproducibly supported by the data's actual distribution.
Consequence
The saved output contains numerically wrong statistics (and possibly a mismatched header/index layout), so the file comparison against the expected result fails outright.
id cf8884ec8f78 · mined from da-code dacode-data-sa-026@s18
raw text (what the judge reads)
### Implausible statistic values after type coercion left unchecked
- **Applies when**: `task` -- the task requires casting text/mixed columns to numeric (or otherwise re-encoding) before computing a summary statistic, correlation, or metric that is then written to a required output file.
- **Pattern**: The attempt coerces columns to numeric (often with silent `errors='coerce'`, per-column parsing, or index/label misalignment), drops missing rows, computes the statistic, and reports it verbatim — with no check that the values are domain-plausible, that the retained row count is sensible, or that the output matches the provided sample format. Near-zero correlations (or otherwise degenerate/uniform values) among variables that should be related are accepted as a "finding" instead of treated as a bug signal.
- **Detection procedure**:
  1. From the task, note the expected value semantics (e.g., ratings on a bounded scale, variables that plausibly co-vary) and the required output shape/format from any sample file.
  2. In the scripts, check whether the numeric conversion is validated: how many values became NaN per column, what the resulting min/max/unique values look like, and whether row alignment is preserved after dropping/reindexing.
  3. In the reported answer, compare the numbers against the plausibility expectation and against the retained-row count; flag if all off-diagonal/aggregate values are suspiciously near zero, constant, or outside the natural range.
  4. Confirm the written file's header/index/orientation and rounding were verified against the sample format, not just assumed.
- **Discriminator**: A genuine violation is when the scripts contain no post-conversion validation (dtype, value range, NaN counts, alignment) and the reported numbers are implausible or unexplained; it is fine if the agent explicitly verified parsed values and row counts and the weak/odd result is reproducibly supported by the data's actual distribution.
- **Consequence**: The saved output contains numerically wrong statistics (and possibly a mismatched header/index layout), so the file comparison against the expected result fails outright.
974Selecting a hyperparameter at the edge of the searched grid and accepting a near-degenerate quality score without alternativestaskda-code
Applies when
task -- the task asks for an unsupervised/model-selection choice (e.g., number of clusters/components) and the script sweeps a fixed grid, picks the argmax/argmin of one internal quality metric, and writes the result out.
Pattern
The attempt fits the model on all raw columns with a single uniform preprocessing (e.g., one scaler applied to continuous and binary/one-hot indicators alike), sweeps a truncated grid, and selects the value at the extreme end of that grid, even though the winning quality score is very low/near-degenerate; it neither extends the grid, nor tries an alternative representation (subset of features, dimensionality reduction, distance/algorithm suited to mixed data), nor cross-checks the choice against structure documented in the data description.
Detection procedure
  1. Read the task/README for any documented structure that constrains the expected answer (number of natural groups, known label cardinality, mixed variable types, dataset size).
  2. In the scripts, identify the searched grid and whether the selected value lies at its first or last element; check whether all columns received the same transformation regardless of type.
  3. In the answer, read the reported quality metric and cluster/component sizes: flag if the metric is barely above the no-structure baseline, if sizes are extremely imbalanced, or if the reported row/feature counts disagree with the documented dataset size.
  4. Confirm no alternative preprocessing/representation or extended grid was tried and compared before writing the output file.
Discriminator
A real violation is a boundary selection plus a weak/unvalidated score plus a single untested preprocessing path; it is fine if the interior of the grid contains the optimum, or the boundary choice is justified by a clear elbow/second metric/domain fact, or the agent compared at least one alternative representation and reported that the chosen setup was better.
Consequence
The saved labeling has the wrong number of clusters and poor separation, so a grader checking cluster count or a cluster-quality/agreement threshold on the output file scores it wrong; the accompanying summary may also contradict the true row count.
id 6b22400bec58 · mined from da-code dacode-ml-cluster-010@s18
raw text (what the judge reads)
### Selecting a hyperparameter at the edge of the searched grid and accepting a near-degenerate quality score without alternatives
- **Applies when**: `task` -- the task asks for an unsupervised/model-selection choice (e.g., number of clusters/components) and the script sweeps a fixed grid, picks the argmax/argmin of one internal quality metric, and writes the result out.
- **Pattern**: The attempt fits the model on all raw columns with a single uniform preprocessing (e.g., one scaler applied to continuous and binary/one-hot indicators alike), sweeps a truncated grid, and selects the value at the extreme end of that grid, even though the winning quality score is very low/near-degenerate; it neither extends the grid, nor tries an alternative representation (subset of features, dimensionality reduction, distance/algorithm suited to mixed data), nor cross-checks the choice against structure documented in the data description.
- **Detection procedure**:
  1. Read the task/README for any documented structure that constrains the expected answer (number of natural groups, known label cardinality, mixed variable types, dataset size).
  2. In the scripts, identify the searched grid and whether the selected value lies at its first or last element; check whether all columns received the same transformation regardless of type.
  3. In the answer, read the reported quality metric and cluster/component sizes: flag if the metric is barely above the no-structure baseline, if sizes are extremely imbalanced, or if the reported row/feature counts disagree with the documented dataset size.
  4. Confirm no alternative preprocessing/representation or extended grid was tried and compared before writing the output file.
- **Discriminator**: A real violation is a boundary selection plus a weak/unvalidated score plus a single untested preprocessing path; it is fine if the interior of the grid contains the optimum, or the boundary choice is justified by a clear elbow/second metric/domain fact, or the agent compared at least one alternative representation and reported that the chosen setup was better.
- **Consequence**: The saved labeling has the wrong number of clusters and poor separation, so a grader checking cluster count or a cluster-quality/agreement threshold on the output file scores it wrong; the accompanying summary may also contradict the true row count.
975Unvalidated degenerate clustering on heavy-tailed featurestaskda-code
Applies when
task -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script builds aggregate numeric features (sums, totals, counts, ratios) that are strongly right-skewed, then standardizes and runs a centroid-based algorithm.
Pattern
The attempt feeds raw heavy-tailed/outlier-dominated features straight into z-scoring + k-means, picks the number of groups from an internal index (or overrides it with an ad-hoc "balance" argument), and accepts a solution where one group holds nearly all rows and the others hold a handful of extreme points — never checking group sizes, per-group feature profiles, or the effect of a variance-stabilizing transform / outlier handling.
Detection procedure
  1. In the task, note that the deliverable is a meaningful segmentation, not just any label column, and note any required column names/format.
  2. In the scripts, check whether skewness of the engineered features is examined and whether any log/rank/quantile transform, winsorizing, or outlier treatment is applied before scaling; also check whether redundant/collinear derived features (e.g., totals plus their ratios) dominate the space.
  3. Check how the number of groups is chosen: is the choice consistent with the reported diagnostics, or is an index-optimal value discarded with a hand-waved justification? Look for suspiciously high separation scores (e.g., silhouette ≳ 0.8) that signal outlier-vs-rest splits.
  4. In the answer, look for reported group-size distributions and interpretations; flag if sizes are wildly imbalanced (a group with <1% of rows) or if no size/profile check is reported at all, and verify the saved file's columns/row count match the requested schema.
Discriminator
A genuine violation is a partition where separation is achieved by isolating a few extreme observations (tiny groups, no transform considered, arbitrary K choice). It is not a violation if skew was addressed (or shown to be immaterial), group sizes are reported and non-degenerate, and the chosen K is defended with diagnostics plus interpretable per-group profiles — even if the sizes are somewhat uneven.
Consequence
The saved label file fails the grader's check on cluster structure (number/balance/quality of clusters), since labels essentially mark outliers rather than customer segments, and the reported conclusions are unsupported.
id be5056ecdf09 · mined from da-code dacode-ml-cluster-019@s18
raw text (what the judge reads)
### Unvalidated degenerate clustering on heavy-tailed features
- **Applies when**: `task` -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script builds aggregate numeric features (sums, totals, counts, ratios) that are strongly right-skewed, then standardizes and runs a centroid-based algorithm.
- **Pattern**: The attempt feeds raw heavy-tailed/outlier-dominated features straight into z-scoring + k-means, picks the number of groups from an internal index (or overrides it with an ad-hoc "balance" argument), and accepts a solution where one group holds nearly all rows and the others hold a handful of extreme points — never checking group sizes, per-group feature profiles, or the effect of a variance-stabilizing transform / outlier handling.
- **Detection procedure**:
  1. In the task, note that the deliverable is a meaningful segmentation, not just any label column, and note any required column names/format.
  2. In the scripts, check whether skewness of the engineered features is examined and whether any log/rank/quantile transform, winsorizing, or outlier treatment is applied before scaling; also check whether redundant/collinear derived features (e.g., totals plus their ratios) dominate the space.
  3. Check how the number of groups is chosen: is the choice consistent with the reported diagnostics, or is an index-optimal value discarded with a hand-waved justification? Look for suspiciously high separation scores (e.g., silhouette ≳ 0.8) that signal outlier-vs-rest splits.
  4. In the answer, look for reported group-size distributions and interpretations; flag if sizes are wildly imbalanced (a group with <1% of rows) or if no size/profile check is reported at all, and verify the saved file's columns/row count match the requested schema.
- **Discriminator**: A genuine violation is a partition where separation is achieved by isolating a few extreme observations (tiny groups, no transform considered, arbitrary K choice). It is *not* a violation if skew was addressed (or shown to be immaterial), group sizes are reported and non-degenerate, and the chosen K is defended with diagnostics plus interpretable per-group profiles — even if the sizes are somewhat uneven.
- **Consequence**: The saved label file fails the grader's check on cluster structure (number/balance/quality of clusters), since labels essentially mark outliers rather than customer segments, and the reported conclusions are unsupported.
976Submission covers only a fraction of the required test IDstaskda-code
Applies when
task -- the task requires writing a prediction/output file with one row per record of a provided evaluation input, in a prescribed column layout.
Pattern
The attempt produces a file (or pastes an answer) containing only a handful of rows — a preview, a debug slice, or a sample — instead of the full set of evaluation IDs, and never checks the row count/ID set against the input file; the file may also be missing or written to the wrong path/name.
Detection procedure
  1. From the task/README, determine the required output filename, header, and the expected number of rows (= number of rows in the evaluation input file, or in the provided sample submission).
  2. In the scripts, locate the write step and confirm it writes the full prediction frame (built from all evaluation rows) to the exact required filename, with no head()/sample()/slicing/early-break in the pipeline.
  3. Compare the answer/file: count rows and check that the ID set exactly matches the evaluation input IDs (no duplicates, no omissions, same order convention if specified).
  4. Verify the presence and naming of all required probability/target columns and that values are plausible (in range, per-row sums sensible).
Discriminator
A real violation is a row count or ID set that does not equal the evaluation input's (e.g., tens of rows where thousands are expected), or a missing/misnamed file. It is fine if the answer text shows only an excerpt but the script demonstrably writes the complete file with matching IDs and count assertions.
Consequence
The grader reports the expected output file as WRONG/MISSING (unscorable or effectively infinite/penalized loss), since most evaluation IDs have no predictions.
id f3b1bff6a461 · mined from da-code dacode-ml-competition-005@s19
raw text (what the judge reads)
### Submission covers only a fraction of the required test IDs
- **Applies when**: `task` -- the task requires writing a prediction/output file with one row per record of a provided evaluation input, in a prescribed column layout.
- **Pattern**: The attempt produces a file (or pastes an answer) containing only a handful of rows — a preview, a debug slice, or a sample — instead of the full set of evaluation IDs, and never checks the row count/ID set against the input file; the file may also be missing or written to the wrong path/name.
- **Detection procedure**:
  1. From the task/README, determine the required output filename, header, and the expected number of rows (= number of rows in the evaluation input file, or in the provided sample submission).
  2. In the scripts, locate the write step and confirm it writes the full prediction frame (built from all evaluation rows) to the exact required filename, with no head()/sample()/slicing/early-break in the pipeline.
  3. Compare the answer/file: count rows and check that the ID set exactly matches the evaluation input IDs (no duplicates, no omissions, same order convention if specified).
  4. Verify the presence and naming of all required probability/target columns and that values are plausible (in range, per-row sums sensible).
- **Discriminator**: A real violation is a row count or ID set that does not equal the evaluation input's (e.g., tens of rows where thousands are expected), or a missing/misnamed file. It is fine if the answer text shows only an excerpt but the script demonstrably writes the complete file with matching IDs and count assertions.
- **Consequence**: The grader reports the expected output file as WRONG/MISSING (unscorable or effectively infinite/penalized loss), since most evaluation IDs have no predictions.
977Ships an unvalidated model: no held-out error estimate and no comparison of predicted vs. observed target distributiontaskda-code
Applies when
task -- the task asks for predictions on a test file and the scripts fit one or more regressors on a separate training source, then write predictions straight to the output file.
Pattern
The scripts train models, hand-pick an ensemble/weights and hyperparameters with no cross-validation or hold-out split, print only summary stats of the predictions, and never check those stats against the training target's distribution (mean, median, spread, heavy tail) or against any error metric — so a systematically shrunken/biased prediction set (typical for skewed count targets fit with squared-error loss and no transformation) is submitted as final.
Detection procedure
  1. Read the task to identify the prediction target and the implicit accuracy expectation for the output file.
  2. Scan the scripts for any train/validation split, cross-validation, or metric computed on labeled data; note whether ensemble weights/hyperparameters are justified by any measured score.
  3. Check whether the scripts compare the predicted distribution (min/max/mean/median/quantiles) with the same statistics of the training labels, and whether a skewed target is handled (log transform, quantile-aware model, or explicit rationale).
  4. Confirm the answer is just the written file with no reported validation error — i.e., nothing that could have caught a badly calibrated or degenerate prediction set.
Discriminator
A real violation has zero labeled-data evaluation anywhere and no distributional sanity check; it is not a violation if the scripts report a CV/hold-out score (even a weak one) used to choose among models, or explicitly compare prediction quantiles to label quantiles and document the gap.
Consequence
The submitted file has the right shape and column name but predictions collapse toward the mean and miss the target's tail, so the grader's accuracy/error threshold on the expected file fails.
id e64cb061579d · mined from da-code dacode-ml-regression-008@s19
raw text (what the judge reads)
### Ships an unvalidated model: no held-out error estimate and no comparison of predicted vs. observed target distribution
- **Applies when**: `task` -- the task asks for predictions on a test file and the scripts fit one or more regressors on a separate training source, then write predictions straight to the output file.
- **Pattern**: The scripts train models, hand-pick an ensemble/weights and hyperparameters with no cross-validation or hold-out split, print only summary stats of the predictions, and never check those stats against the training target's distribution (mean, median, spread, heavy tail) or against any error metric — so a systematically shrunken/biased prediction set (typical for skewed count targets fit with squared-error loss and no transformation) is submitted as final.
- **Detection procedure**:
  1. Read the task to identify the prediction target and the implicit accuracy expectation for the output file.
  2. Scan the scripts for any train/validation split, cross-validation, or metric computed on labeled data; note whether ensemble weights/hyperparameters are justified by any measured score.
  3. Check whether the scripts compare the predicted distribution (min/max/mean/median/quantiles) with the same statistics of the training labels, and whether a skewed target is handled (log transform, quantile-aware model, or explicit rationale).
  4. Confirm the answer is just the written file with no reported validation error — i.e., nothing that could have caught a badly calibrated or degenerate prediction set.
- **Discriminator**: A real violation has zero labeled-data evaluation anywhere and no distributional sanity check; it is *not* a violation if the scripts report a CV/hold-out score (even a weak one) used to choose among models, or explicitly compare prediction quantiles to label quantiles and document the gap.
- **Consequence**: The submitted file has the right shape and column name but predictions collapse toward the mean and miss the target's tail, so the grader's accuracy/error threshold on the expected file fails.
978Output never validated against the provided sample/template file (schema, ordering, rounding)taskda-code
Applies when
task -- the task says results must be written to a file "following the exact structure/format" of a provided sample or template output.
Pattern
The scripts never read the sample file; the agent hard-codes guessed column names/order and sort order, and writes raw floats straight from the aggregation without matching the sample's decimal precision or numeric formatting (visible as artifacts like long floating-point tails).
Detection procedure
  1. In the task text, note that a sample/template output file is referenced and identify what it constrains (column names, column order, row order, number of rows, rounding/units).
  2. Search the scripts for any read/inspection of that sample file and for an explicit conformance step (reindex to sample columns, apply the sample's row ordering, round()/format to the sample's decimal places).
  3. Inspect the produced answer: are float columns printed with inconsistent/unbounded precision, are column names/order derived from the agent's own rename/sort_values rather than from the sample, and is the row ordering justified by the sample rather than assumed alphabetical?
  4. Flag if steps 2–3 show no evidence the output was compared cell-by-cell (or at least header- and dtype-wise) against the sample before saving.
Discriminator
A fine attempt either loads the sample and programmatically aligns columns/ordering/rounding to it, or explicitly prints the sample's header and one sample row and demonstrates the produced file matches; a violation only asserts "renamed columns to match expected output" without ever having opened the sample, and leaves unrounded/unformatted numbers.
Consequence
The graded file mismatches the expected file on header, row order, or numeric precision, so an exact/tolerant file comparison fails even when the underlying aggregation logic is right — reported as WRONG/MISSING for the output file.
id 206ed63ac8a6 · mined from da-code dacode-dm-csv-011@s19
raw text (what the judge reads)
### Output never validated against the provided sample/template file (schema, ordering, rounding)
- **Applies when**: `task` -- the task says results must be written to a file "following the exact structure/format" of a provided sample or template output.
- **Pattern**: The scripts never read the sample file; the agent hard-codes guessed column names/order and sort order, and writes raw floats straight from the aggregation without matching the sample's decimal precision or numeric formatting (visible as artifacts like long floating-point tails).
- **Detection procedure**:
  1. In the task text, note that a sample/template output file is referenced and identify what it constrains (column names, column order, row order, number of rows, rounding/units).
  2. Search the scripts for any read/inspection of that sample file and for an explicit conformance step (reindex to sample columns, apply the sample's row ordering, `round()`/format to the sample's decimal places).
  3. Inspect the produced answer: are float columns printed with inconsistent/unbounded precision, are column names/order derived from the agent's own `rename`/`sort_values` rather than from the sample, and is the row ordering justified by the sample rather than assumed alphabetical?
  4. Flag if steps 2–3 show no evidence the output was compared cell-by-cell (or at least header- and dtype-wise) against the sample before saving.
- **Discriminator**: A fine attempt either loads the sample and programmatically aligns columns/ordering/rounding to it, or explicitly prints the sample's header and one sample row and demonstrates the produced file matches; a violation only asserts "renamed columns to match expected output" without ever having opened the sample, and leaves unrounded/unformatted numbers.
- **Consequence**: The graded file mismatches the expected file on header, row order, or numeric precision, so an exact/tolerant file comparison fails even when the underlying aggregation logic is right — reported as WRONG/MISSING for the output file.
979Output feature columns don't faithfully mirror the input feature vectortaskda-code
Applies when
task -- the task asks to save per-row results with generically indexed columns (e.g. Feature_i) plus a derived label/prediction, implying the output must reproduce the input feature vector alongside the new column.
Pattern
The attempt silently redefines what "the feature vector" is — dropping original columns, replacing raw fields with transformed/engineered substitutes, appending extra derived aggregates, or writing scaled/encoded values — so the number, order, and content of Feature_i columns no longer correspond to the dataset's own columns row-for-row; row count may also change from filtering/dropna.
Detection procedure
  1. From the task statement, determine what "ith value in the feature vector" must be: read the input file's column list and note the count and order after only the minimal cleaning the task authorizes.
  2. In the scripts, trace the DataFrame actually written to the output file: list every column added (engineered/aggregate/one-hot), every column dropped or renamed, whether values written are raw or standardized/imputed, and whether any rows were filtered.
  3. Compare the answer's stated column count/names and row count against step 1; flag if the count differs from the input's feature count, if any column is a self-invented derivation, or if written values are transformed rather than original.
  4. Check that the label column is named exactly as requested and that the file has header names matching the requested pattern (Feature_0..., Cluster).
Discriminator
A real violation is changing the identity/cardinality of the feature vector (adding "Total_*" style aggregates, substituting age for birth year, exporting scaled values, dropping rows). Not a violation: applying scaling/encoding only internally for model fitting while exporting the original per-row values, or dropping a pure identifier column that is explicitly not a feature, as long as the exported columns still map 1:1 to input columns in original order.
Consequence
The saved file fails the expected-file check (column count/shape or per-row values mismatch), so the task scores 0 even though the clustering itself ran without error.
id 30dea2ca9346 · mined from da-code dacode-ml-cluster-014@s19
raw text (what the judge reads)
### Output feature columns don't faithfully mirror the input feature vector
- **Applies when**: `task` -- the task asks to save per-row results with generically indexed columns (e.g. `Feature_i`) plus a derived label/prediction, implying the output must reproduce the input feature vector alongside the new column.
- **Pattern**: The attempt silently redefines what "the feature vector" is — dropping original columns, replacing raw fields with transformed/engineered substitutes, appending extra derived aggregates, or writing scaled/encoded values — so the number, order, and content of `Feature_i` columns no longer correspond to the dataset's own columns row-for-row; row count may also change from filtering/dropna.
- **Detection procedure**:
  1. From the task statement, determine what "ith value in the feature vector" must be: read the input file's column list and note the count and order after only the minimal cleaning the task authorizes.
  2. In the scripts, trace the DataFrame actually written to the output file: list every column added (engineered/aggregate/one-hot), every column dropped or renamed, whether values written are raw or standardized/imputed, and whether any rows were filtered.
  3. Compare the answer's stated column count/names and row count against step 1; flag if the count differs from the input's feature count, if any column is a self-invented derivation, or if written values are transformed rather than original.
  4. Check that the label column is named exactly as requested and that the file has header names matching the requested pattern (`Feature_0...`, `Cluster`).
- **Discriminator**: A real violation is changing the identity/cardinality of the feature vector (adding "Total_*" style aggregates, substituting age for birth year, exporting scaled values, dropping rows). Not a violation: applying scaling/encoding only internally for model fitting while exporting the original per-row values, or dropping a pure identifier column that is explicitly not a feature, as long as the exported columns still map 1:1 to input columns in original order.
- **Consequence**: The saved file fails the expected-file check (column count/shape or per-row values mismatch), so the task scores 0 even though the clustering itself ran without error.
980Submission file never validated against the provided format templatetaskda-code
Applies when
task -- The task says to write predictions to a named output file "adhering to the format specified" in a provided sample/template file, and the scripts construct that file themselves.
Pattern
The scripts hard-code an assumed schema and output path, never reading the template file, and additionally apply unrequested post-processing (e.g., casting continuous predictions to integers / rounding) — so column names, dtypes, row count/order, or the file location can silently diverge from what the grader reads, and precision is thrown away.
Detection procedure
  1. In the task statement, note the required output filename/location and the existence of a template file describing the format.
  2. Search the scripts for a read of that template file and any assertion comparing the built output's columns, row count, id set/order, and dtypes against it; also check the written path matches the required one exactly.
  3. Check whether predictions are transformed before writing (rounding, clipping, type casts) and whether the task/metric actually requires that transformation.
  4. Check the final answer for evidence that the saved file was re-read and sanity-checked (shape, head, id alignment) rather than just "submission saved".
Discriminator
A real violation is when no template comparison or post-write verification exists, or the output path/schema/dtype is assumed; it is fine if the script explicitly loads the template, aligns ids/columns to it, keeps the value format the metric expects, and re-reads the written file to confirm shape and columns.
Consequence
The grader reports the expected submission file as WRONG/MISSING (not found at the expected path, or wrong columns/ids/row count), or the score is needlessly degraded by discretized predictions, even though the modeling itself ran fine.
id d2d4143b6814 · mined from da-code dacode-ml-competition-009@s19
raw text (what the judge reads)
### Submission file never validated against the provided format template
- **Applies when**: `task` -- The task says to write predictions to a named output file "adhering to the format specified" in a provided sample/template file, and the scripts construct that file themselves.
- **Pattern**: The scripts hard-code an assumed schema and output path, never reading the template file, and additionally apply unrequested post-processing (e.g., casting continuous predictions to integers / rounding) — so column names, dtypes, row count/order, or the file location can silently diverge from what the grader reads, and precision is thrown away.
- **Detection procedure**:
  1. In the task statement, note the required output filename/location and the existence of a template file describing the format.
  2. Search the scripts for a read of that template file and any assertion comparing the built output's columns, row count, id set/order, and dtypes against it; also check the written path matches the required one exactly.
  3. Check whether predictions are transformed before writing (rounding, clipping, type casts) and whether the task/metric actually requires that transformation.
  4. Check the final answer for evidence that the saved file was re-read and sanity-checked (shape, head, id alignment) rather than just "submission saved".
- **Discriminator**: A real violation is when no template comparison or post-write verification exists, or the output path/schema/dtype is assumed; it is fine if the script explicitly loads the template, aligns ids/columns to it, keeps the value format the metric expects, and re-reads the written file to confirm shape and columns.
- **Consequence**: The grader reports the expected submission file as WRONG/MISSING (not found at the expected path, or wrong columns/ids/row count), or the score is needlessly degraded by discretized predictions, even though the modeling itself ran fine.
981Output rows silently reduced (or features arbitrarily subset) relative to the full input datasettaskda-code
Applies when
task -- the task asks for a per-record output file (e.g., labels/predictions/scores plus the feature vector) built from a whole dataset that contains missing values or mixed-type columns.
Pattern
The script drops records that fail an ad-hoc completeness threshold, and/or hand-picks a small subset of the available usable columns as "the feature vector", then writes a file whose row count and column count no longer correspond to the full dataset the task named. The answer reports the reduced counts as if they were the deliverable.
Detection procedure
  1. Read the task and note that the deliverable is one row per record of the stated dataset, with the feature columns implied by "the feature vector".
  2. In the script, look for dropna, thresh=, row filters, or a hard-coded list of columns; compare the resulting shape to the raw loaded shape printed earlier.
  3. Check the answer/log: if the reported number of output rows is smaller than the raw record count, or the number of Feature_i columns is smaller than the number of columns that could be coerced to numeric, flag it.
  4. Also check that included features are meaningful measurements, not identifier-like codes (dial codes, IDs, indices) that were kept only because they parse as numbers.
Discriminator
A real violation is unexplained/ad-hoc shrinkage — rows or columns removed for convenience when imputation or full coercion was available, with no requirement in the task to filter. It is fine if the task itself specifies a filter, or if columns are genuinely non-numeric/free-text and could not be encoded, and the script documents that every remaining record is retained via imputation.
Consequence
The saved file's row/column shape and label vector cannot be aligned with the reference solution, so a shape- or per-record-comparison grader marks the result file WRONG/MISSING even though the clustering code itself ran without error.
id db6cf8acf734 · mined from da-code dacode-ml-cluster-009@s19
raw text (what the judge reads)
### Output rows silently reduced (or features arbitrarily subset) relative to the full input dataset
- **Applies when**: `task` -- the task asks for a per-record output file (e.g., labels/predictions/scores plus the feature vector) built from a whole dataset that contains missing values or mixed-type columns.
- **Pattern**: The script drops records that fail an ad-hoc completeness threshold, and/or hand-picks a small subset of the available usable columns as "the feature vector", then writes a file whose row count and column count no longer correspond to the full dataset the task named. The answer reports the reduced counts as if they were the deliverable.
- **Detection procedure**:
  1. Read the task and note that the deliverable is one row per record of the stated dataset, with the feature columns implied by "the feature vector".
  2. In the script, look for `dropna`, `thresh=`, row filters, or a hard-coded list of columns; compare the resulting shape to the raw loaded shape printed earlier.
  3. Check the answer/log: if the reported number of output rows is smaller than the raw record count, or the number of `Feature_i` columns is smaller than the number of columns that could be coerced to numeric, flag it.
  4. Also check that included features are meaningful measurements, not identifier-like codes (dial codes, IDs, indices) that were kept only because they parse as numbers.
- **Discriminator**: A real violation is *unexplained/ad-hoc* shrinkage — rows or columns removed for convenience when imputation or full coercion was available, with no requirement in the task to filter. It is fine if the task itself specifies a filter, or if columns are genuinely non-numeric/free-text and could not be encoded, and the script documents that every remaining record is retained via imputation.
- **Consequence**: The saved file's row/column shape and label vector cannot be aligned with the reference solution, so a shape- or per-record-comparison grader marks the result file WRONG/MISSING even though the clustering code itself ran without error.
982Unverified output artifact + anomalous diagnostics explained away instead of fixedtaskda-code
Applies when
task -- the task requires writing a result file with a prescribed schema and choosing a model/hyperparameter (e.g., number of groups) from an internal quality metric.
Pattern
The agent engineers highly skewed, mutually redundant features, runs the selection metric, sees an obviously degenerate signal (e.g., a near-perfect score paired with a single-point/near-empty group, or scores dominated by a few extreme records), verbally "rejects" that option and hand-picks another setting, then declares success from in-memory variables without reloading the written file and checking it against the literal spec (column names/indexing base, one row per entity, value ranges, no NaNs, cluster sizes).
Detection procedure
  1. Read the task and list the hard output requirements: exact file name, exact column naming pattern and starting index, row granularity, and what each row must correspond to.
  2. Read the scripts for the preprocessing/feature step: check whether heavy-tailed or outlier-dominated features were left untransformed/unwinsorized and whether row-dropping filters silently change the set of entities that must appear in the output.
  3. Read the model-selection code: does an anomalous metric value get diagnosed (outlier removal, log/robust scaling, cluster-size inspection) or merely narrated and skipped in favor of a manually chosen setting?
  4. Check whether the script (or the answer) re-reads the saved file and asserts shape, column names, dtype, no missing values, and reasonable label distribution; an answer that only cites in-memory counts fails this step.
Discriminator
A real violation is when no post-write read-back/schema assertion exists and a degenerate diagnostic was overridden without changing the data or method; it is fine if the agent addressed the skew (robust scaling/outlier handling) or justified the choice with a stability check, and independently verified the saved file matches the requested schema and entity count.
Consequence
The grader loads the result file and finds it missing, mis-named/mis-indexed columns, wrong row count/granularity, or a near-trivial partition (one giant cluster plus singletons), so the file check fails even though the report claims success.
id 974c29e936e2 · mined from da-code dacode-ml-cluster-016@s19
raw text (what the judge reads)
### Unverified output artifact + anomalous diagnostics explained away instead of fixed
- **Applies when**: `task` -- the task requires writing a result file with a prescribed schema and choosing a model/hyperparameter (e.g., number of groups) from an internal quality metric.
- **Pattern**: The agent engineers highly skewed, mutually redundant features, runs the selection metric, sees an obviously degenerate signal (e.g., a near-perfect score paired with a single-point/near-empty group, or scores dominated by a few extreme records), verbally "rejects" that option and hand-picks another setting, then declares success from in-memory variables without reloading the written file and checking it against the literal spec (column names/indexing base, one row per entity, value ranges, no NaNs, cluster sizes).
- **Detection procedure**:
  1. Read the task and list the hard output requirements: exact file name, exact column naming pattern and starting index, row granularity, and what each row must correspond to.
  2. Read the scripts for the preprocessing/feature step: check whether heavy-tailed or outlier-dominated features were left untransformed/unwinsorized and whether row-dropping filters silently change the set of entities that must appear in the output.
  3. Read the model-selection code: does an anomalous metric value get *diagnosed* (outlier removal, log/robust scaling, cluster-size inspection) or merely narrated and skipped in favor of a manually chosen setting?
  4. Check whether the script (or the answer) re-reads the saved file and asserts shape, column names, dtype, no missing values, and reasonable label distribution; an answer that only cites in-memory counts fails this step.
- **Discriminator**: A real violation is when no post-write read-back/schema assertion exists *and* a degenerate diagnostic was overridden without changing the data or method; it is fine if the agent addressed the skew (robust scaling/outlier handling) or justified the choice with a stability check, and independently verified the saved file matches the requested schema and entity count.
- **Consequence**: The grader loads the result file and finds it missing, mis-named/mis-indexed columns, wrong row count/granularity, or a near-trivial partition (one giant cluster plus singletons), so the file check fails even though the report claims success.
983Ships a model whose accuracy was never estimated out-of-sample (in-sample fit reported as "performance")taskda-code
Applies when
task -- the task asks for predictions on a held-out file to be scored by an external metric, and the scripts fit a model and write predictions in one pass.
Pattern
The attempt trains on all labeled rows, reports only the fit score on those same rows (or nothing at all), and never creates a validation/CV split — so overfitting, feature-availability mismatch between train and prediction inputs, and degenerate imputations go undetected; the high in-sample score is then presented as evidence the answer is good.
Detection procedure
  1. Read the task to confirm the deliverable is scored on unseen rows by an accuracy-type metric.
  2. Scan the scripts for any hold-out split, time-based split, or cross-validation, and for any error metric computed on rows excluded from fit(); note whether the reported score comes from model.score(X, y) on the training matrix.
  3. Check whether the prediction-side feature matrix is validated against the training one (per-column NaN/missing rates, identical column set/order, comparable value ranges) rather than silently filled with training means.
  4. Check the answer: is the quoted quality number in-sample, and is there any sanity comparison of predicted distribution vs. observed target distribution?
Discriminator
A real violation is when no estimate on data withheld from training exists anywhere (in-sample score only). It is fine if the agent trains a final model on all data after reporting a hold-out/CV score, or if it validates predictions against a withheld period and only then refits.
Consequence
The grader compares predictions to true values and finds error far worse than the advertised in-sample R²/score (or predictions shifted/flattened because some prediction-time features were absent and mean-filled), so the output file is marked wrong with no diagnostic in the attempt explaining why.
id c008e2e70183 · mined from da-code dacode-ml-regression-002@s19
raw text (what the judge reads)
### Ships a model whose accuracy was never estimated out-of-sample (in-sample fit reported as "performance")
- **Applies when**: `task` -- the task asks for predictions on a held-out file to be scored by an external metric, and the scripts fit a model and write predictions in one pass.
- **Pattern**: The attempt trains on all labeled rows, reports only the fit score on those same rows (or nothing at all), and never creates a validation/CV split — so overfitting, feature-availability mismatch between train and prediction inputs, and degenerate imputations go undetected; the high in-sample score is then presented as evidence the answer is good.
- **Detection procedure**:
  1. Read the task to confirm the deliverable is scored on unseen rows by an accuracy-type metric.
  2. Scan the scripts for any hold-out split, time-based split, or cross-validation, and for any error metric computed on rows excluded from `fit()`; note whether the reported score comes from `model.score(X, y)` on the training matrix.
  3. Check whether the prediction-side feature matrix is validated against the training one (per-column NaN/missing rates, identical column set/order, comparable value ranges) rather than silently filled with training means.
  4. Check the answer: is the quoted quality number in-sample, and is there any sanity comparison of predicted distribution vs. observed target distribution?
- **Discriminator**: A real violation is when *no* estimate on data withheld from training exists anywhere (in-sample score only). It is fine if the agent trains a final model on all data *after* reporting a hold-out/CV score, or if it validates predictions against a withheld period and only then refits.
- **Consequence**: The grader compares predictions to true values and finds error far worse than the advertised in-sample R²/score (or predictions shifted/flattened because some prediction-time features were absent and mean-filled), so the output file is marked wrong with no diagnostic in the attempt explaining why.
984Fabricated/hard-coded data instead of loading the provided datasettaskda-code
Applies when
task -- the task asks for an analysis, statistic, or plot derived from a supplied dataset file.
Pattern
The script never reads the actual data file (or reads only the config/spec file) and instead defines the values inline as "sample"/"representative"/"approximate" numbers, then formats them to the requested output.
Detection procedure
1) From the task, identify which input data file(s) must be read to answer. 2) Scan the script for a load call (read_csv/read_excel/load/open) on that data file and trace whether the plotted/reported values flow from it. 3) Look for literal arrays/dicts of numbers or comments like "sample data", "based on", "estimated". 4) Check the answer: does it present numbers that were never computed from the file, and are all required output artifacts produced?
Discriminator
Hard-coded literals are acceptable when they come from the task/config specification itself (titles, colors, figure size, axis labels, thresholds) or are verifiably re-derived from loaded data; a violation is when the substantive data values (the quantities being analyzed) originate in the script rather than the dataset.
Consequence
Every value-based check fails — expected artifacts (serialized plot data / arrays) are missing or mismatch the true series, so the grader scores 0 even though the image looks well-formatted.
id 41126cc1e0dd · mined from da-code dacode-plot-line-015@s19
raw text (what the judge reads)
### Fabricated/hard-coded data instead of loading the provided dataset
- **Applies when**: `task` -- the task asks for an analysis, statistic, or plot derived from a supplied dataset file.
- **Pattern**: The script never reads the actual data file (or reads only the config/spec file) and instead defines the values inline as "sample"/"representative"/"approximate" numbers, then formats them to the requested output.
- **Detection procedure**: 1) From the task, identify which input data file(s) must be read to answer. 2) Scan the script for a load call (`read_csv`/`read_excel`/`load`/`open`) on that data file and trace whether the plotted/reported values flow from it. 3) Look for literal arrays/dicts of numbers or comments like "sample data", "based on", "estimated". 4) Check the answer: does it present numbers that were never computed from the file, and are all required output artifacts produced?
- **Discriminator**: Hard-coded literals are acceptable when they come from the task/config specification itself (titles, colors, figure size, axis labels, thresholds) or are verifiably re-derived from loaded data; a violation is when the substantive data values (the quantities being analyzed) originate in the script rather than the dataset.
- **Consequence**: Every value-based check fails — expected artifacts (serialized plot data / arrays) are missing or mismatch the true series, so the grader scores 0 even though the image looks well-formatted.
985Fabricated/synthetic input data instead of the provided datasettaskda-code
Applies when
task -- the task references provided data files and auxiliary instruction/format files (e.g., a tips or README file, a sample output file) that the analysis must be based on.
Pattern
The script hard-codes arrays of "typical" or "realistic" values invented by the agent, never opens the supplied data files or the stated instruction/format file, and then computes and reports statistics from those invented numbers.
Detection procedure
1. List the data/instruction/format artifacts the task says to use. 2. Grep the scripts for any file-reading call (read_csv, load, open, etc.) and check whether each named artifact is actually read. 3. Check whether the numeric inputs to the computation come from literal in-script arrays rather than loaded data, and whether output column names/order come from the sample format file rather than being guessed. 4. Sanity-check the reported statistic against any counts/ranges implied by the real data description.
Discriminator
A real violation is when the analysis' inputs are literals the agent authored; it is fine if literals only encode documented constants, thresholds, or parameters (seed, number of bootstrap replicates) while the measurements themselves are loaded from the provided files.
Consequence
The reported statistic and p-value are unrelated to the true data (and column names may not match the sample), so the result file fails every value/format check.
id 93e0fc45e968 · mined from da-code dacode-data-sa-028@s19
raw text (what the judge reads)
### Fabricated/synthetic input data instead of the provided dataset
- **Applies when**: `task` -- the task references provided data files and auxiliary instruction/format files (e.g., a tips or README file, a sample output file) that the analysis must be based on.
- **Pattern**: The script hard-codes arrays of "typical" or "realistic" values invented by the agent, never opens the supplied data files or the stated instruction/format file, and then computes and reports statistics from those invented numbers.
- **Detection procedure**: 1. List the data/instruction/format artifacts the task says to use. 2. Grep the scripts for any file-reading call (`read_csv`, `load`, `open`, etc.) and check whether each named artifact is actually read. 3. Check whether the numeric inputs to the computation come from literal in-script arrays rather than loaded data, and whether output column names/order come from the sample format file rather than being guessed. 4. Sanity-check the reported statistic against any counts/ranges implied by the real data description.
- **Discriminator**: A real violation is when the analysis' *inputs* are literals the agent authored; it is fine if literals only encode documented constants, thresholds, or parameters (seed, number of bootstrap replicates) while the measurements themselves are loaded from the provided files.
- **Consequence**: The reported statistic and p-value are unrelated to the true data (and column names may not match the sample), so the result file fails every value/format check.
986Specification files referenced by the task are never read or appliedtaskda-code
Applies when
task -- The prompt points to auxiliary instruction/config artifacts (e.g., a tips/notes file, a YAML/JSON style or format spec) that define the analysis rules and required output artifacts.
Pattern
The agent never opens those files; it guesses the rules from memory or generic conventions, hard-codes its own filters/definitions/labels/styling, and emits only the one obvious output file, silently skipping other required artifacts (serialized figure data, arrays, metadata) and every explicit formatting constraint (figure size, colors, titles, axis labels, ordering, units, rounding).
Detection procedure
  1. From the task text, list every referenced instruction/config file and every named output artifact.
  2. Grep the scripts for reads of each referenced file (open, read_csv, yaml.safe_load, json.load); if none exist, the rules used are unverified guesses.
  3. Grep for writes of each named output artifact; confirm each one is produced with the requested name/format.
  4. Check whether plotting/derivation parameters in the code (filters, groupings, category definitions, labels, sizes, styles) are traceable to the spec rather than invented inline, and whether the answer text cites spec contents that were actually loaded.
Discriminator
A real violation is when the spec file is never opened at run time (paraphrasing its "requirements" in comments/answer without reading it counts as violation), or a required artifact is absent. It is fine if the script loads the spec and then legitimately hard-codes values it read, or if the extra artifact is genuinely written under the exact requested name.
Consequence
Missing artifacts are scored as WRONG/MISSING, and the one produced file diverges from the required content/format, so all correctness checks fail despite a confident "task completed" report.
id 5aedfe233444 · mined from da-code dacode-plot-line-006@s19
raw text (what the judge reads)
### Specification files referenced by the task are never read or applied
- **Applies when**: `task` -- The prompt points to auxiliary instruction/config artifacts (e.g., a tips/notes file, a YAML/JSON style or format spec) that define the analysis rules and required output artifacts.
- **Pattern**: The agent never opens those files; it guesses the rules from memory or generic conventions, hard-codes its own filters/definitions/labels/styling, and emits only the one obvious output file, silently skipping other required artifacts (serialized figure data, arrays, metadata) and every explicit formatting constraint (figure size, colors, titles, axis labels, ordering, units, rounding).
- **Detection procedure**:
  1. From the task text, list every referenced instruction/config file and every named output artifact.
  2. Grep the scripts for reads of each referenced file (`open`, `read_csv`, `yaml.safe_load`, `json.load`); if none exist, the rules used are unverified guesses.
  3. Grep for writes of each named output artifact; confirm each one is produced with the requested name/format.
  4. Check whether plotting/derivation parameters in the code (filters, groupings, category definitions, labels, sizes, styles) are traceable to the spec rather than invented inline, and whether the answer text cites spec contents that were actually loaded.
- **Discriminator**: A real violation is when the spec file is never opened at run time (paraphrasing its "requirements" in comments/answer without reading it counts as violation), or a required artifact is absent. It is fine if the script loads the spec and then legitimately hard-codes values it read, or if the extra artifact is genuinely written under the exact requested name.
- **Consequence**: Missing artifacts are scored as WRONG/MISSING, and the one produced file diverges from the required content/format, so all correctness checks fail despite a confident "task completed" report.
987Ignoring a task-referenced specification file (assuming its contents instead of reading it)taskda-code
Applies when
task -- The task instructions point to an external document/config/schema in the workspace (e.g., a .md, .json, or .txt file) that defines how to bin, filter, name, order, or output the results.
Pattern
The script never opens or prints the referenced file; instead it hardcodes categories/bins/labels/thresholds inferred from the raw data's existing values or from the agent's prior expectations, and the answer asserts compliance ("as specified in ...") without any evidence the file was inspected.
Detection procedure
  1. List every external artifact the task names as governing the method or the outputs (spec files, and any required output filenames/formats).
  2. Search the scripts for a read/open/print of each named spec file, and for the creation of each named output artifact.
  3. If a spec file is never read, check whether the constants in the script (bin edges, group labels, ordering) could only have come from the data's own value set or a guess — that is the violation.
  4. Confirm the answer does not quote or restate any content actually loaded from the spec.
Discriminator
Fine if the script reads the spec (or its content is quoted in the transcript) and the hardcoded constants demonstrably match it; a violation if the constants are derived solely from value_counts()/unique values of the column, or if a required output file named in the task is never written.
Consequence
The grouping/aggregation differs from the reference definition (e.g., merged or re-cut categories) and/or required output files are missing, so every file-level check fails even though the plot looks superficially correct.
id 9c5ac6e0f508 · mined from da-code dacode-plot-bar-005@s19
raw text (what the judge reads)
### Ignoring a task-referenced specification file (assuming its contents instead of reading it)
- **Applies when**: `task` -- The task instructions point to an external document/config/schema in the workspace (e.g., a `.md`, `.json`, or `.txt` file) that defines how to bin, filter, name, order, or output the results.
- **Pattern**: The script never opens or prints the referenced file; instead it hardcodes categories/bins/labels/thresholds inferred from the raw data's existing values or from the agent's prior expectations, and the answer asserts compliance ("as specified in ...") without any evidence the file was inspected.
- **Detection procedure**:
  1. List every external artifact the task names as governing the method or the outputs (spec files, and any required output filenames/formats).
  2. Search the scripts for a read/open/print of each named spec file, and for the creation of each named output artifact.
  3. If a spec file is never read, check whether the constants in the script (bin edges, group labels, ordering) could only have come from the data's own value set or a guess — that is the violation.
  4. Confirm the answer does not quote or restate any content actually loaded from the spec.
- **Discriminator**: Fine if the script reads the spec (or its content is quoted in the transcript) and the hardcoded constants demonstrably match it; a violation if the constants are derived solely from `value_counts()`/unique values of the column, or if a required output file named in the task is never written.
- **Consequence**: The grouping/aggregation differs from the reference definition (e.g., merged or re-cut categories) and/or required output files are missing, so every file-level check fails even though the plot looks superficially correct.
988No script writes the required output artifact in the requested schemataskda-code
Applies when
task -- the task specifies a concrete deliverable (an output file and/or a literal answer template with given keys and value types) and the agent's scripts only compute and print results.
Pattern
The scripts do the analysis and print the final numbers to stdout, but never serialize the answer to the expected file (e.g. result.json) and never construct the object in exactly the shape shown in the prompt (key names, list-vs-scalar values, rounding/units), leaving the graded artifact absent or schema-mismatched.
Detection procedure
  1. Read the task statement and note every required deliverable: file name/path, key names, and the exact value shape shown in the template.
  2. Grep the scripts for any write/serialization call (json.dump, to_csv, open(...,'w'), etc.) and check that the target path matches the expected artifact.
  3. Compare the in-script answer object (or printed text) key-by-key with the template: same key spelling, same container type for values, same units/rounding.
  4. If no write exists, or the object deviates from the template, flag the attempt regardless of whether the computation is correct.
Discriminator
A real violation is a missing/renamed/differently-typed deliverable; it is not a violation if the script writes the correct file and the printed console output is merely additional logging, or if cosmetic differences (whitespace, key order) don't change the parsed structure.
Consequence
The grader reports the expected result file as WRONG/MISSING and scores 0, even though the underlying computed value may be right.
id 879ffa1823f5 · mined from da-code dacode-di-text-002@s19
raw text (what the judge reads)
### No script writes the required output artifact in the requested schema
- **Applies when**: `task` -- the task specifies a concrete deliverable (an output file and/or a literal answer template with given keys and value types) and the agent's scripts only compute and `print` results.
- **Pattern**: The scripts do the analysis and print the final numbers to stdout, but never serialize the answer to the expected file (e.g. `result.json`) and never construct the object in exactly the shape shown in the prompt (key names, list-vs-scalar values, rounding/units), leaving the graded artifact absent or schema-mismatched.
- **Detection procedure**:
  1. Read the task statement and note every required deliverable: file name/path, key names, and the exact value shape shown in the template.
  2. Grep the scripts for any write/serialization call (`json.dump`, `to_csv`, `open(...,'w')`, etc.) and check that the target path matches the expected artifact.
  3. Compare the in-script answer object (or printed text) key-by-key with the template: same key spelling, same container type for values, same units/rounding.
  4. If no write exists, or the object deviates from the template, flag the attempt regardless of whether the computation is correct.
- **Discriminator**: A real violation is a missing/renamed/differently-typed deliverable; it is *not* a violation if the script writes the correct file and the printed console output is merely additional logging, or if cosmetic differences (whitespace, key order) don't change the parsed structure.
- **Consequence**: The grader reports the expected result file as WRONG/MISSING and scores 0, even though the underlying computed value may be right.
989Computing a dataset-wide statistic on only one split/file of the available datataskinfiagent-dabench
Applies when
task -- the task asks for a descriptive statistic or property of "the dataset" and the script hard-codes a single input file (e.g., a _train/_test/partial extract) without checking what other data files exist.
Pattern
The agent loads the first plausible file it finds, computes the requested statistic on that subset, and reports it as the answer for the whole dataset; no reconciliation of row counts against the full data, and no note of why that file was chosen. Nearby variants (e.g., population vs. sample denominator in a dispersion term) are also picked implicitly without checking which definition the constraint implies.
Detection procedure
  1. Read the task: does it refer to "the dataset" as a whole, or does it explicitly name a split/subset?
  2. Read the script's load step: list every file path it reads, and check whether the script (or any prior command) enumerated the data directory to confirm those files are the complete data.
  3. Compare the number of rows/records actually used against the total available (or at least confirm the agent verified this); flag if a split-suffixed or partial file is used for a whole-dataset claim.
  4. Check any formula with a convention choice (ddof, denominator, rounding point) against the definition stated in the task, since the same numeric output can silently shift.
Discriminator
A real violation is a whole-dataset question answered from an unverified subset, or a formula convention chosen without justification. It is not a violation if the task itself names the split, or if the agent inspected the directory and documented that the loaded file is the entire dataset (or that other files lack the needed column).
Consequence
The reported numeric value differs from the reference computed on the full data (the qualitative label may still coincidentally match), so the value check fails even though the interpretation check passes.
id 7d752a10f6ae · mined from infiagent-dabench dabench-359@s19
raw text (what the judge reads)
### Computing a dataset-wide statistic on only one split/file of the available data
- **Applies when**: `task` -- the task asks for a descriptive statistic or property of "the dataset" and the script hard-codes a single input file (e.g., a `*_train`/`*_test`/partial extract) without checking what other data files exist.
- **Pattern**: The agent loads the first plausible file it finds, computes the requested statistic on that subset, and reports it as the answer for the whole dataset; no reconciliation of row counts against the full data, and no note of why that file was chosen. Nearby variants (e.g., population vs. sample denominator in a dispersion term) are also picked implicitly without checking which definition the constraint implies.
- **Detection procedure**:
  1. Read the task: does it refer to "the dataset" as a whole, or does it explicitly name a split/subset?
  2. Read the script's load step: list every file path it reads, and check whether the script (or any prior command) enumerated the data directory to confirm those files are the complete data.
  3. Compare the number of rows/records actually used against the total available (or at least confirm the agent verified this); flag if a split-suffixed or partial file is used for a whole-dataset claim.
  4. Check any formula with a convention choice (ddof, denominator, rounding point) against the definition stated in the task, since the same numeric output can silently shift.
- **Discriminator**: A real violation is a whole-dataset question answered from an unverified subset, or a formula convention chosen without justification. It is *not* a violation if the task itself names the split, or if the agent inspected the directory and documented that the loaded file is the entire dataset (or that other files lack the needed column).
- **Consequence**: The reported numeric value differs from the reference computed on the full data (the qualitative label may still coincidentally match), so the value check fails even though the interpretation check passes.
990Statistic computed on an uncleaned / unverified input vector (no sanity check, required diagnostic not reported)taskinfiagent-dabench
Applies when
task -- The task asks for a distributional statistic or hypothesis test (normality test, skew/kurtosis, correlation, mean) on one named column, with a stated decision rule and a required intermediate value (e.g., the p-value) to report.
Pattern
The attempt pulls the column straight from the file and feeds it to the test without checking dtype, missing values, sentinel/placeholder codes (-999, 0-fill, "N/A"), duplicated header rows, or unintended extra rows; it then reports only the final verdict and moments, omitting the required diagnostic (p-value) and any evidence of what the tested vector actually was. Contaminating values dominate the moments and flip the test decision.
Detection procedure
  1. Read the task and list every required output and every required intermediate (p-value, n, alpha rule) plus any implied preprocessing (which rows, which column, how to treat NaNs).
  2. In the scripts, locate the exact vector passed to the test: check that it is dropna()-ed (or explicitly justified otherwise), numerically typed, filtered to the intended rows, and that its length/min/max are printed.
  3. Check whether the script prints the test statistic and p-value and compares them to the stated alpha, rather than hard-coding or eyeballing a verdict.
  4. Compare the answer's skew/kurtosis with the printed summary: extreme asymmetry or heavy tails (|skew| > ~2, kurtosis >> 3) with no comment or outlier inspection is a red flag that placeholder/foreign values entered the vector; also confirm every requested field, including the p-value, appears in the final answer.
Discriminator
A real violation is when the tested vector's size/contents were never printed or validated, or a required reported quantity is missing — so a sentinel- or NaN-driven distortion could pass unnoticed. It is not a violation if the script logs n, dtype, and range, shows the cleaning step, and the extreme moments are genuinely present in the clean data.
Consequence
The test operates on a different vector than intended, so the p-value crosses the alpha threshold the wrong way and the reported skew/kurtosis are off; the grader marks the normality verdict (and usually the moments) wrong.
id 0d17415ddaeb · mined from infiagent-dabench dabench-298@s19
raw text (what the judge reads)
### Statistic computed on an uncleaned / unverified input vector (no sanity check, required diagnostic not reported)
- **Applies when**: `task` -- The task asks for a distributional statistic or hypothesis test (normality test, skew/kurtosis, correlation, mean) on one named column, with a stated decision rule and a required intermediate value (e.g., the p-value) to report.
- **Pattern**: The attempt pulls the column straight from the file and feeds it to the test without checking dtype, missing values, sentinel/placeholder codes (-999, 0-fill, "N/A"), duplicated header rows, or unintended extra rows; it then reports only the final verdict and moments, omitting the required diagnostic (p-value) and any evidence of what the tested vector actually was. Contaminating values dominate the moments and flip the test decision.
- **Detection procedure**:
  1. Read the task and list every required output *and* every required intermediate (p-value, n, alpha rule) plus any implied preprocessing (which rows, which column, how to treat NaNs).
  2. In the scripts, locate the exact vector passed to the test: check that it is `dropna()`-ed (or explicitly justified otherwise), numerically typed, filtered to the intended rows, and that its length/min/max are printed.
  3. Check whether the script prints the test statistic and p-value and compares them to the stated alpha, rather than hard-coding or eyeballing a verdict.
  4. Compare the answer's skew/kurtosis with the printed summary: extreme asymmetry or heavy tails (|skew| > ~2, kurtosis >> 3) with no comment or outlier inspection is a red flag that placeholder/foreign values entered the vector; also confirm every requested field, including the p-value, appears in the final answer.
- **Discriminator**: A real violation is when the tested vector's size/contents were never printed or validated, or a required reported quantity is missing — so a sentinel- or NaN-driven distortion could pass unnoticed. It is *not* a violation if the script logs n, dtype, and range, shows the cleaning step, and the extreme moments are genuinely present in the clean data.
- **Consequence**: The test operates on a different vector than intended, so the p-value crosses the alpha threshold the wrong way and the reported skew/kurtosis are off; the grader marks the normality verdict (and usually the moments) wrong.
991Prediction file row count / alignment with the input rows to be scoredtaskda-code
Applies when
task -- the task asks for a per-row output file (predictions, labels, scores) covering every record of a given input file.
Pattern
The attempt produces an output file whose number of rows (and their order) does not correspond one-to-one with the input rows to be predicted — e.g. only a sample, a truncated preview, a subset that survived dropna/filtering/deduplication, or predictions generated for a validation split instead of the full target file.
Detection procedure
  1. From the task, identify the input file to be scored and determine its exact row count (read it, don't assume).
  2. In the scripts, trace which frame is passed to predict/transform and whether any filtering, dropna, head, sample, deduplication, or train/validation split occurred between loading and writing; check that the written frame is the full input frame in original order.
  3. Read the produced file: confirm len(output) == len(input), the required column name/header is present exactly as specified, and label values are drawn from the expected label set.
  4. If any of these fail or the script never asserts them, flag the attempt.
Discriminator
A real violation is a genuine mismatch in row count or ordering versus the target input file; a look-alike that is fine is an output that is merely displayed truncated in the report/console while the on-disk file has the full, correctly ordered set of rows (verify on disk, not from the printed excerpt).
Consequence
The grader cannot align predictions with ground truth, so the file is marked WRONG/MISSING regardless of model quality, and any reported accuracy reflects only the covered subset.
id db9be74cc685 · mined from da-code dacode-ml-multi-011@s19
raw text (what the judge reads)
### Prediction file row count / alignment with the input rows to be scored
- **Applies when**: `task` -- the task asks for a per-row output file (predictions, labels, scores) covering every record of a given input file.
- **Pattern**: The attempt produces an output file whose number of rows (and their order) does not correspond one-to-one with the input rows to be predicted — e.g. only a sample, a truncated preview, a subset that survived dropna/filtering/deduplication, or predictions generated for a validation split instead of the full target file.
- **Detection procedure**:
  1. From the task, identify the input file to be scored and determine its exact row count (read it, don't assume).
  2. In the scripts, trace which frame is passed to `predict`/`transform` and whether any filtering, `dropna`, `head`, `sample`, deduplication, or train/validation split occurred between loading and writing; check that the written frame is the full input frame in original order.
  3. Read the produced file: confirm `len(output) == len(input)`, the required column name/header is present exactly as specified, and label values are drawn from the expected label set.
  4. If any of these fail or the script never asserts them, flag the attempt.
- **Discriminator**: A real violation is a genuine mismatch in row count or ordering versus the target input file; a look-alike that is fine is an output that is merely displayed truncated in the report/console while the on-disk file has the full, correctly ordered set of rows (verify on disk, not from the printed excerpt).
- **Consequence**: The grader cannot align predictions with ground truth, so the file is marked WRONG/MISSING regardless of model quality, and any reported accuracy reflects only the covered subset.
992Final submission produced by an unvalidated model (no held-out metric, arbitrary ensemble weights)taskda-code
Applies when
task -- the task asks for predictions on a held-out set scored by an accuracy/error metric, and the script that actually writes the submission file fits models and predicts without any cross-validation or hold-out evaluation.
Pattern
The attempt runs several exploratory scripts (some with validation), but the last script that overwrites the submission trains on all data, blends models with hand-picked weights, and reports only prediction min/max/mean and file shape — never a validation score for the exact pipeline being submitted, no comparison against a simple baseline, and no check that the prediction distribution matches the target distribution in training.
Detection procedure
1. Read the task to identify the scored deliverable and its implied metric. 2. Find the script whose output is the final submission file, and check whether it computes a hold-out/CV metric for that exact model/ensemble (and weight choice) rather than inheriting a score from a different, earlier configuration. 3. Check whether the answer reports a validation metric for the submitted pipeline and compares it to at least one alternative/baseline; treat "shape correct, no NaNs, values in range" as format checks, not quality evidence. 4. Compare the reported prediction spread/mean to the training target's spread/mean — heavily shrunk variance or a much narrower range than the target signals an underfit/mis-blended model that was never validated.
Discriminator
A real violation is when no metric exists for the submitted configuration (weights, hyperparameters, feature set) and the blend is justified only by intuition ("trees usually win"); it is fine if the script tunes/selects via CV or hold-out, prints the score for the final configuration, and then refits that same configuration on all data.
Consequence
The submitted predictions score materially worse than an easily reachable baseline (e.g., a plain regularized linear or properly tuned GBM), so the leaderboard/threshold check on the submission file fails even though the file format is valid.
id 061fe4d0b4ab · mined from da-code dacode-ml-competition-008@s19
raw text (what the judge reads)
### Final submission produced by an unvalidated model (no held-out metric, arbitrary ensemble weights)
- **Applies when**: `task` -- the task asks for predictions on a held-out set scored by an accuracy/error metric, and the script that actually writes the submission file fits models and predicts without any cross-validation or hold-out evaluation.
- **Pattern**: The attempt runs several exploratory scripts (some with validation), but the last script that overwrites the submission trains on all data, blends models with hand-picked weights, and reports only prediction min/max/mean and file shape — never a validation score for the exact pipeline being submitted, no comparison against a simple baseline, and no check that the prediction distribution matches the target distribution in training.
- **Detection procedure**: 1. Read the task to identify the scored deliverable and its implied metric. 2. Find the script whose output is the final submission file, and check whether it computes a hold-out/CV metric for that exact model/ensemble (and weight choice) rather than inheriting a score from a different, earlier configuration. 3. Check whether the answer reports a validation metric for the submitted pipeline and compares it to at least one alternative/baseline; treat "shape correct, no NaNs, values in range" as format checks, not quality evidence. 4. Compare the reported prediction spread/mean to the training target's spread/mean — heavily shrunk variance or a much narrower range than the target signals an underfit/mis-blended model that was never validated.
- **Discriminator**: A real violation is when no metric exists for the submitted configuration (weights, hyperparameters, feature set) and the blend is justified only by intuition ("trees usually win"); it is fine if the script tunes/selects via CV or hold-out, prints the score for the final configuration, and then refits that same configuration on all data.
- **Consequence**: The submitted predictions score materially worse than an easily reachable baseline (e.g., a plain regularized linear or properly tuned GBM), so the leaderboard/threshold check on the submission file fails even though the file format is valid.
993Ignoring the provided spec/config for analysis parameters and omitting required companion output artifactstaskda-code
Applies when
task -- the task points to an external specification file (yaml/json/config) and/or names several expected output files, and the script produces a figure or summary from binned/aggregated data.
Pattern
The script reads only the cosmetic keys of the spec (title, colors, fonts, figsize) while hard-coding the substantive analysis choices — bin edges, category labels, ordering, filters, units — and writes only the visual output, skipping the machine-checkable data artifacts (e.g. serialized plot data and numeric arrays) that the task expects.
Detection procedure
  1. From the task text, list every named output file and every parameter the spec file is said to govern.
  2. In the script, check whether the spec file is fully consumed: print/dump its keys and verify that each key controlling grouping, bins, labels, filtering, or ordering is actually used rather than replaced by a literal in the code.
  3. Grep the script for save/write calls and confirm one exists for each expected output file, with the expected structure (not just savefig).
  4. Compare the answer's reported categories/counts to what the spec dictates; any bin scheme or label set the agent invented, or any missing file, is a violation.
Discriminator
A fine attempt either uses spec-provided values for all substantive parameters (defaulting only where the spec is genuinely silent) and emits every requested file; a violation invents grouping/bin definitions that the spec supplies, or reports success while one or more required artifacts were never written. Adding extra, harmless styling not in the spec is not a violation.
Consequence
Graders that compare the serialized plot data / numeric arrays find them missing, and the figure's bars and category labels do not match the reference, so every check fails despite a confident "task completed" report.
id 3d9b8810d4e9 · mined from da-code dacode-plot-bar-007@s19
raw text (what the judge reads)
### Ignoring the provided spec/config for analysis parameters and omitting required companion output artifacts
- **Applies when**: `task` -- the task points to an external specification file (yaml/json/config) and/or names several expected output files, and the script produces a figure or summary from binned/aggregated data.
- **Pattern**: The script reads only the cosmetic keys of the spec (title, colors, fonts, figsize) while hard-coding the substantive analysis choices — bin edges, category labels, ordering, filters, units — and writes only the visual output, skipping the machine-checkable data artifacts (e.g. serialized plot data and numeric arrays) that the task expects.
- **Detection procedure**:
  1. From the task text, list every named output file and every parameter the spec file is said to govern.
  2. In the script, check whether the spec file is fully consumed: print/dump its keys and verify that each key controlling grouping, bins, labels, filtering, or ordering is actually used rather than replaced by a literal in the code.
  3. Grep the script for save/write calls and confirm one exists for each expected output file, with the expected structure (not just `savefig`).
  4. Compare the answer's reported categories/counts to what the spec dictates; any bin scheme or label set the agent invented, or any missing file, is a violation.
- **Discriminator**: A fine attempt either uses spec-provided values for all substantive parameters (defaulting only where the spec is genuinely silent) and emits every requested file; a violation invents grouping/bin definitions that the spec supplies, or reports success while one or more required artifacts were never written. Adding extra, harmless styling not in the spec is not a violation.
- **Consequence**: Graders that compare the serialized plot data / numeric arrays find them missing, and the figure's bars and category labels do not match the reference, so every check fails despite a confident "task completed" report.
994Substituting different entities/metrics than the ones the task namestaskda-code
Applies when
task -- The prompt names specific quantities, grouping keys, and ranking criteria (e.g., an average per stage, grouped by a named entity, ranked by a named measure) plus required output artifacts.
Pattern
The agent cannot find the named fields in the available data, silently redefines them as loose analogues from whatever file it happens to load (a different grouping key, a different ranking measure, a different value to plot), and reports success as if the requested analysis was performed.
Detection procedure
  1. List from the task every required element: the value being aggregated, the aggregation function, the grouping key, the ranking/selection rule, and the output files.
  2. Read the scripts and map each element to the concrete column/computation used; check that the loaded file actually contains fields matching the requested semantics.
  3. Read the answer's own description of axes/categories and compare it word-for-word to the task's wording; any renaming ("matches" for the ranking measure, "outcomes" for the stages) is a substitution.
  4. Confirm every named output artifact is produced with the required config/settings source applied, not improvised styling.
Discriminator
A real violation is when the plotted/reported quantity or grouping differs in meaning from what was asked (different entity type, different metric, different ranking basis). A look-alike that is fine is a faithful computation that merely renames columns internally or derives the requested field from raw inputs while preserving its definition.
Consequence
The saved figure and any numeric/serialized outputs encode the wrong categories and values, so all artifact comparisons against the expected results fail (0/N checks), even though the run "completed successfully."
id 02a99956122a · mined from da-code dacode-plot-scatter-002@s19
raw text (what the judge reads)
### Substituting different entities/metrics than the ones the task names
- **Applies when**: `task` -- The prompt names specific quantities, grouping keys, and ranking criteria (e.g., an average per stage, grouped by a named entity, ranked by a named measure) plus required output artifacts.
- **Pattern**: The agent cannot find the named fields in the available data, silently redefines them as loose analogues from whatever file it happens to load (a different grouping key, a different ranking measure, a different value to plot), and reports success as if the requested analysis was performed.
- **Detection procedure**:
  1. List from the task every required element: the value being aggregated, the aggregation function, the grouping key, the ranking/selection rule, and the output files.
  2. Read the scripts and map each element to the concrete column/computation used; check that the loaded file actually contains fields matching the requested semantics.
  3. Read the answer's own description of axes/categories and compare it word-for-word to the task's wording; any renaming ("matches" for the ranking measure, "outcomes" for the stages) is a substitution.
  4. Confirm every named output artifact is produced with the required config/settings source applied, not improvised styling.
- **Discriminator**: A real violation is when the plotted/reported quantity or grouping differs in meaning from what was asked (different entity type, different metric, different ranking basis). A look-alike that is fine is a faithful computation that merely renames columns internally or derives the requested field from raw inputs while preserving its definition.
- **Consequence**: The saved figure and any numeric/serialized outputs encode the wrong categories and values, so all artifact comparisons against the expected results fail (0/N checks), even though the run "completed successfully."
995Per-group analysis collapsed into a single global statistictaskda-code
Applies when
task -- the instructions specify a grouping/stratification step (e.g., "for each <category>, filter/compute ...") before running a statistical test or metric, and the requested output format is a list.
Pattern
The attempt uses the grouping only as a preprocessing detail (or ignores it entirely), then pools all rows and reports one number, so the answer list has a single element instead of one entry per group; the output may also only be printed rather than written to the required result file.
Detection procedure
  1. Read the task and count how many distinct outputs the wording implies (number of groups/strata named or derivable from the data) and note the required file/format (list-valued keys usually signal multiple entries).
  2. In the scripts, check whether the loop/groupby used for filtering also drives the test/metric computation, or whether the data is re-pooled afterwards.
  3. Compare the length of each list in the answer with the group count from step 1, and confirm the paired keys (values and conclusions) have equal length and consistent group ordering.
  4. Verify the answer was persisted to the exact filename/schema requested.
Discriminator
A single value is correct only if the task explicitly asks for one overall test after group-wise cleaning; if the task phrases the grouping as "for each ..." or the schema uses lists of parallel values/labels, one element is a violation. Length mismatch caused by a group being empty after filtering must be justified in the script, not silent.
Consequence
The expected result file is judged wrong/missing because the reported list has the wrong cardinality (and often a p-value near a decision boundary that flips the conclusion), scoring 0.
id 1d397c0a82c2 · mined from da-code dacode-data-sa-061@s19
raw text (what the judge reads)
### Per-group analysis collapsed into a single global statistic
- **Applies when**: `task` -- the instructions specify a grouping/stratification step (e.g., "for each <category>, filter/compute ...") before running a statistical test or metric, and the requested output format is a list.
- **Pattern**: The attempt uses the grouping only as a preprocessing detail (or ignores it entirely), then pools all rows and reports one number, so the answer list has a single element instead of one entry per group; the output may also only be printed rather than written to the required result file.
- **Detection procedure**:
  1. Read the task and count how many distinct outputs the wording implies (number of groups/strata named or derivable from the data) and note the required file/format (list-valued keys usually signal multiple entries).
  2. In the scripts, check whether the loop/`groupby` used for filtering also drives the test/metric computation, or whether the data is re-pooled afterwards.
  3. Compare the length of each list in the answer with the group count from step 1, and confirm the paired keys (values and conclusions) have equal length and consistent group ordering.
  4. Verify the answer was persisted to the exact filename/schema requested.
- **Discriminator**: A single value is correct only if the task explicitly asks for one overall test after group-wise cleaning; if the task phrases the grouping as "for each ..." or the schema uses lists of parallel values/labels, one element is a violation. Length mismatch caused by a group being empty after filtering must be justified in the script, not silent.
- **Consequence**: The expected result file is judged wrong/missing because the reported list has the wrong cardinality (and often a p-value near a decision boundary that flips the conclusion), scoring 0.
996Held-out validation skipped and high-signal columns discardedtaskda-code
Applies when
task -- a supervised prediction task where a full labeled source table is given plus a test table, and the script trains a model on a hand-picked subset of numeric columns.
Pattern
The attempt selects only a few "obvious" numeric/audio-style features, drops all identifier, textual, categorical and date columns without testing their value, never checks whether the test rows are already present (joinable) in the provided labeled table, and judges quality solely from metrics computed on the same rows used for fitting — then reports that in-sample score as evidence the answer is good.
Detection procedure
  1. From the task, note that predictive accuracy on unseen rows (not fit quality) is what will be graded.
  2. In the script, check whether any train/validation split or cross-validation is created; flag if the only reported metrics come from predict() on the same matrix passed to fit().
  3. Check whether the script compares candidate feature sets/baselines (e.g., adding categorical/date/identifier features, or a mean-prediction baseline), and whether it tests for key-column overlap between the labeled source table and the test table (which could give near-exact answers by lookup/merge).
  4. In the answer, look for an optimistic score (high R²/low RMSE) that is described as model performance with no out-of-sample number; also check whether prediction spread is far narrower than the label's true spread.
Discriminator
A real violation is when no out-of-sample estimate exists anywhere and informative columns were dropped by assumption; it is fine if the script reports a CV/holdout score (even a mediocre one) and documents that additional columns were tried or are genuinely unavailable in the test table.
Consequence
The reported metric overstates accuracy (memorized training fit), predictions collapse toward the mean, and the grader's row-wise comparison against true labels fails the accuracy/correlation threshold even though the file format looks correct.
id 11e7c7d4c831 · mined from da-code dacode-ml-regression-004@s19
raw text (what the judge reads)
### Held-out validation skipped and high-signal columns discarded
- **Applies when**: `task` -- a supervised prediction task where a full labeled source table is given plus a test table, and the script trains a model on a hand-picked subset of numeric columns.
- **Pattern**: The attempt selects only a few "obvious" numeric/audio-style features, drops all identifier, textual, categorical and date columns without testing their value, never checks whether the test rows are already present (joinable) in the provided labeled table, and judges quality solely from metrics computed on the same rows used for fitting — then reports that in-sample score as evidence the answer is good.
- **Detection procedure**:
  1. From the task, note that predictive accuracy on unseen rows (not fit quality) is what will be graded.
  2. In the script, check whether any train/validation split or cross-validation is created; flag if the only reported metrics come from `predict()` on the same matrix passed to `fit()`.
  3. Check whether the script compares candidate feature sets/baselines (e.g., adding categorical/date/identifier features, or a mean-prediction baseline), and whether it tests for key-column overlap between the labeled source table and the test table (which could give near-exact answers by lookup/merge).
  4. In the answer, look for an optimistic score (high R²/low RMSE) that is described as model performance with no out-of-sample number; also check whether prediction spread is far narrower than the label's true spread.
- **Discriminator**: A real violation is when *no* out-of-sample estimate exists anywhere and informative columns were dropped by assumption; it is fine if the script reports a CV/holdout score (even a mediocre one) and documents that additional columns were tried or are genuinely unavailable in the test table.
- **Consequence**: The reported metric overstates accuracy (memorized training fit), predictions collapse toward the mean, and the grader's row-wise comparison against true labels fails the accuracy/correlation threshold even though the file format looks correct.
997Missing feature scaling / unvalidated cluster structure in distance-based unsupervised modelstaskda-code
Applies when
task -- the task asks for clustering (or any distance/variance-based unsupervised method) on tabular features that span very different units and magnitudes, and the script must choose a cluster count and write per-row labels.
Pattern
The attempt feeds raw, un-standardized columns straight into a Euclidean-distance algorithm (and/or picks k without a documented selection curve), so a few large-magnitude columns dominate the geometry; the result is accepted despite a low separation score and degenerate cluster sizes (singleton or near-singleton clusters that are really outliers), and no sanity check is run on the saved file.
Detection procedure
  1. From the task/README, list the feature columns and note their typical ranges; confirm the algorithm is distance- or variance-based.
  2. In the script, check for an explicit scaling/normalization step (e.g. standardize or min-max) applied to all numeric features before fitting, and check that the cluster count comes from a stated criterion (elbow/silhouette sweep, or a domain-justified value) rather than an arbitrary constant.
  3. In the reported output, inspect the cluster-size distribution and the quality score: clusters of size 1–3 alongside one cluster holding ~half the rows, plus a weak silhouette, indicate scale-dominated or outlier-driven splits.
  4. Verify the written file was re-read and checked for the exact required column names/order, row count equal to the input rows, and label values that are consistent integers.
Discriminator
A genuine violation shows no scaling step (or scaling applied to only some columns / after fitting) or an unjustified k, together with degenerate cluster sizes; it is fine if the script scales properly, justifies k with a sweep, and the small cluster is a deliberately reported outlier group whose members are inspected and the separation metric is reasonable for the data.
Consequence
The saved label column does not match the expected partition (grader compares cluster assignments/structure), so the output file is marked WRONG even though the file exists and has the right headers.
id 21bf692476a1 · mined from da-code dacode-ml-cluster-013@s19
raw text (what the judge reads)
### Missing feature scaling / unvalidated cluster structure in distance-based unsupervised models
- **Applies when**: `task` -- the task asks for clustering (or any distance/variance-based unsupervised method) on tabular features that span very different units and magnitudes, and the script must choose a cluster count and write per-row labels.
- **Pattern**: The attempt feeds raw, un-standardized columns straight into a Euclidean-distance algorithm (and/or picks k without a documented selection curve), so a few large-magnitude columns dominate the geometry; the result is accepted despite a low separation score and degenerate cluster sizes (singleton or near-singleton clusters that are really outliers), and no sanity check is run on the saved file.
- **Detection procedure**:
  1. From the task/README, list the feature columns and note their typical ranges; confirm the algorithm is distance- or variance-based.
  2. In the script, check for an explicit scaling/normalization step (e.g. standardize or min-max) applied to all numeric features before fitting, and check that the cluster count comes from a stated criterion (elbow/silhouette sweep, or a domain-justified value) rather than an arbitrary constant.
  3. In the reported output, inspect the cluster-size distribution and the quality score: clusters of size 1–3 alongside one cluster holding ~half the rows, plus a weak silhouette, indicate scale-dominated or outlier-driven splits.
  4. Verify the written file was re-read and checked for the exact required column names/order, row count equal to the input rows, and label values that are consistent integers.
- **Discriminator**: A genuine violation shows no scaling step (or scaling applied to only some columns / after fitting) or an unjustified k, together with degenerate cluster sizes; it is fine if the script scales properly, justifies k with a sweep, and the small cluster is a deliberately reported outlier group whose members are inspected and the separation metric is reasonable for the data.
- **Consequence**: The saved label column does not match the expected partition (grader compares cluster assignments/structure), so the output file is marked WRONG even though the file exists and has the right headers.
998Statistic computed over the wrong slice/axis, ignoring an explicit subset specifier in the tasktaskinfiagent-dabench
Applies when
task -- the task names a specific subset or coordinate (a single year, group, region, split, or column) over which a statistic must be computed, and the data are stored in a wide/multi-dimensional layout (or split across several files) so several different slices are technically computable.
Pattern
The script computes the statistic over a convenient but different slice than the one requested — e.g. reducing across all periods/columns of each row instead of restricting to the named coordinate, or aggregating along the wrong axis — and never demonstrates that the named subset was actually used to filter or index the data. The requested identifier never appears in the code as a filter/selection, and the candidate universe of entities may also be truncated (only one of several source files loaded).
Detection procedure
  1. From the task statement, list every explicit qualifier (specific period/label, entity universe, definition/variant of the statistic, output format) that constrains the computation.
  2. Search the scripts for each qualifier: is there a filter, column selection, or index using that exact value, and are all files/entities that could belong to the answer universe loaded?
  3. Check the reduction step: what array is passed into the statistic function, and does its shape/length correspond to the requested subset rather than to the full row/table?
  4. Check whether any sanity output (shape, count, list of included entities/columns) is printed that would confirm the intended slice; absence of such a check plus absence of the qualifier in code is a violation.
Discriminator
A real violation is when the requested qualifier is nowhere used to select data, so the reported number describes a different population than asked. It is not a violation if the qualifier is genuinely inapplicable after a documented, correct reinterpretation (e.g. the script explicitly restricts to the named subset and then reduces over the only remaining dimension), or if the script filters correctly but merely prints extra diagnostics over other slices.
Consequence
The reported entity/value is the argmax/estimate of a different quantity, so the graded answer names the wrong item and scores 0 even though the code runs without error.
id bc444ebbc986 · mined from infiagent-dabench dabench-252@s19
raw text (what the judge reads)
### Statistic computed over the wrong slice/axis, ignoring an explicit subset specifier in the task
- **Applies when**: `task` -- the task names a specific subset or coordinate (a single year, group, region, split, or column) over which a statistic must be computed, and the data are stored in a wide/multi-dimensional layout (or split across several files) so several different slices are technically computable.
- **Pattern**: The script computes the statistic over a convenient but different slice than the one requested — e.g. reducing across all periods/columns of each row instead of restricting to the named coordinate, or aggregating along the wrong axis — and never demonstrates that the named subset was actually used to filter or index the data. The requested identifier never appears in the code as a filter/selection, and the candidate universe of entities may also be truncated (only one of several source files loaded).
- **Detection procedure**:
  1. From the task statement, list every explicit qualifier (specific period/label, entity universe, definition/variant of the statistic, output format) that constrains the computation.
  2. Search the scripts for each qualifier: is there a filter, column selection, or index using that exact value, and are all files/entities that could belong to the answer universe loaded?
  3. Check the reduction step: what array is passed into the statistic function, and does its shape/length correspond to the requested subset rather than to the full row/table?
  4. Check whether any sanity output (shape, count, list of included entities/columns) is printed that would confirm the intended slice; absence of such a check plus absence of the qualifier in code is a violation.
- **Discriminator**: A real violation is when the requested qualifier is nowhere used to select data, so the reported number describes a different population than asked. It is *not* a violation if the qualifier is genuinely inapplicable after a documented, correct reinterpretation (e.g. the script explicitly restricts to the named subset and then reduces over the only remaining dimension), or if the script filters correctly but merely prints extra diagnostics over other slices.
- **Consequence**: The reported entity/value is the argmax/estimate of a different quantity, so the graded answer names the wrong item and scores 0 even though the code runs without error.
999Truncating a computed identifier to fit an ambiguous format templatetaskinfiagent-dabench
Applies when
task -- the answer requires reporting a key/label (date, ID, category, index) that the script extracts from the data, and the requested output format string is coarser or otherwise inconsistent with the granularity of the values actually stored in the data.
Pattern
The script correctly locates the row of interest, but then applies a lossy string transformation (slicing, strftime to a shorter pattern, rounding, stripping suffixes) so the reported label matches the literal format template, discarding precision that the grader expects; downstream computations still use the full-precision value, so only the reported label is degraded.
Detection procedure
  1. Read the task's answer-format spec and note the stated granularity/precision of each reported field.
  2. In the script, find where the identifier is produced and check for any truncation/reformatting (e.g., value[:7], format strings, casts) between extraction and reporting.
  3. Compare the granularity of the raw values in the source column (as printed by the script's own inspection output) to what is being emitted — if the raw value carries more resolution and the script deletes it, flag it.
  4. Check the answer text: does the emitted label uniquely identify the row that was actually used in the calculation, or could it refer to many rows?
Discriminator
A real violation is when the truncated label no longer pins down the single record used for the rest of the computation (many rows share it) and the raw data unambiguously supports a finer label; it is fine when the source values genuinely have only that granularity, or when the coarser label is itself the natural unit of the grouping performed (e.g., the maximum was computed over aggregated periods).
Consequence
The numeric part of the answer matches, but the identifier field is scored wrong/missing, so the submission fails on that check even though the analysis was correct.
id 520661c93952 · mined from infiagent-dabench dabench-572@s19
raw text (what the judge reads)
### Truncating a computed identifier to fit an ambiguous format template
- **Applies when**: `task` -- the answer requires reporting a key/label (date, ID, category, index) that the script extracts from the data, and the requested output format string is coarser or otherwise inconsistent with the granularity of the values actually stored in the data.
- **Pattern**: The script correctly locates the row of interest, but then applies a lossy string transformation (slicing, `strftime` to a shorter pattern, rounding, stripping suffixes) so the reported label matches the literal format template, discarding precision that the grader expects; downstream computations still use the full-precision value, so only the reported label is degraded.
- **Detection procedure**:
  1. Read the task's answer-format spec and note the stated granularity/precision of each reported field.
  2. In the script, find where the identifier is produced and check for any truncation/reformatting (e.g., `value[:7]`, format strings, casts) between extraction and reporting.
  3. Compare the granularity of the raw values in the source column (as printed by the script's own inspection output) to what is being emitted — if the raw value carries more resolution and the script deletes it, flag it.
  4. Check the answer text: does the emitted label uniquely identify the row that was actually used in the calculation, or could it refer to many rows?
- **Discriminator**: A real violation is when the truncated label no longer pins down the single record used for the rest of the computation (many rows share it) and the raw data unambiguously supports a finer label; it is fine when the source values genuinely have only that granularity, or when the coarser label is itself the natural unit of the grouping performed (e.g., the maximum was computed over aggregated periods).
- **Consequence**: The numeric part of the answer matches, but the identifier field is scored wrong/missing, so the submission fails on that check even though the analysis was correct.