MLE-bench · 100 rubrics · first 100 of all
HF EdwardoSunny/mle-rubrics-onpolicy-100 · local data/libraries/mle-rubrics-onpolicy-100.json
task -- the task specifies an explicit evaluation metric (e.g., log loss, AUC, RMSE) for predictions on an unlabeled test set, and the script trains models and writes a submission file.cross_val_score/KFold, or explicit computation of that metric on labeled data; check whether printed diagnostics include a metric value at all.83f41d76ebb5 · mined from da-code dacode-ml-competition-005### No held-out validation against the competition's stated metric before submitting - **Applies when**: `task` -- the task specifies an explicit evaluation metric (e.g., log loss, AUC, RMSE) for predictions on an unlabeled test set, and the script trains models and writes a submission file. - **Pattern**: The script fits one or more models on 100% of the labeled data, blends them with hand-picked weights, and writes predictions straight to disk — never computing the stated metric on a validation split, cross-validation folds, or even against a trivial baseline (e.g., class-prior probabilities). Hyperparameters, imputation choices, and ensemble weights are therefore unjustified, and there is no evidence the output is better than random or than a constant prediction. - **Detection procedure**: 1. Read the task and note the exact scoring metric and its inputs (probabilities vs. labels, per-row normalization, clipping). 2. Scan the script for any train/validation split, `cross_val_score`/`KFold`, or explicit computation of that metric on labeled data; check whether printed diagnostics include a metric value at all. 3. Check whether model/ensemble choices (weights, depth, learning rate) are tied to any measured score, or are literal constants written by the author. 4. Inspect the answer/submission: is any score, baseline comparison, or sanity check (row count equal to test rows, ids matching test ids, probability ranges/sums, no degenerate constant column) reported? - **Discriminator**: A real violation is when *no* estimate of the stated metric exists anywhere for any candidate model, so the attempt cannot distinguish a good submission from a broken one. It is *not* a violation if the script measures the metric via CV/holdout (even briefly) and uses it to pick among options, nor if the metric is unmeasurable because labels genuinely do not exist for any subset. - **Consequence**: The submission may be systematically miscalibrated, mis-ordered, mislabeled by class column, or simply far worse than a simple baseline; the grader reports a failing/incorrect submission with no diagnostic trail, since the agent had no internal score to catch it.
task -- a script trains a regressor/classifier on one file and writes predictions for another file as the deliverable, with no ground truth available for the predicted rows.4e7f8b275c9f · mined from da-code dacode-ml-regression-008### Ships model predictions with no held-out validation and no distribution sanity check - **Applies when**: `task` -- a script trains a regressor/classifier on one file and writes predictions for another file as the deliverable, with no ground truth available for the predicted rows. - **Pattern**: The attempt fits a single model on 100% of the training rows, immediately predicts on the target rows, and writes the output without (a) any hold-out/CV error estimate, (b) a comparison of the predicted value distribution against the training target distribution, or (c) a check that the feature columns/dtypes/encodings used at prediction time are the same as those learned at fit time. Degenerate output (e.g. the vast majority of predictions collapsed to the same value, or a range far narrower than the target's) is accepted as-is. - **Detection procedure**: 1. Read the task: identify the deliverable (file, required column names/order, row alignment, rounding/units) and note that no labels exist for the predicted rows. 2. Read the script: check whether a train/validation split or cross-validation with a printed error metric exists, and whether the same column names, imputation, and category encodings are applied consistently to both datasets (watch for near-identical but non-identical column spellings, or one file's schema being assumed for the other). 3. Read the script's output stage: check whether the predicted values' summary statistics are compared to the training target's summary statistics, and whether any explicit guard rejects a degenerate/off-scale prediction vector. 4. Inspect the submitted answer: compute the share of identical values and the min/max; if predictions are dominated by a single value or their spread is an order of magnitude off the target's documented spread, and no validation metric was reported, flag it. - **Discriminator**: A real violation is an unvalidated pipeline whose output is visibly degenerate or whose feature handling silently differs between fit and predict; it is *not* a violation if the script reports a hold-out/CV score and the prediction distribution plausibly matches the training target (a skewed target legitimately yields many small values, provided the reported validation error supports it). - **Consequence**: The written file passes format checks but its values are near-constant/mis-scaled, so the grader's accuracy or error tolerance against the reference targets fails, marking the deliverable WRONG.
task -- the task asks for a statistical test, metric, or comparison whose population is qualified by filters (date range, category/tournament type, subgroup, one-sided direction) stated in the prompt, README, or problem context.e6246376b83e · mined from da-code dacode-data-sa-001### Hypothesis test run on the full table instead of the task-specified population/subset - **Applies when**: `task` -- the task asks for a statistical test, metric, or comparison whose population is qualified by filters (date range, category/tournament type, subgroup, one-sided direction) stated in the prompt, README, or problem context. - **Pattern**: The attempt loads the raw files and computes the statistic on every row of both groups, silently dropping the qualifying conditions (and/or defaulting to a two-sided test when a directional hypothesis was specified), so the reported p-value/metric describes a different population than the one asked about. - **Detection procedure**: 1. Read the task and any README, and list every explicit or implied restriction on the rows/columns to be analyzed (time window, subset of categories, groups compared) plus the alternative hypothesis direction and significance level. 2. Read the scripts and check that each restriction appears as a concrete filter/dtype conversion (e.g., date parsing then range filter, category equality filter) before the statistic is computed, and that the test call matches the stated alternative and test type (paired vs independent, equal-variance assumption). 3. Compare the row counts used in the test against the raw file row counts; if the script never prints or reduces counts, treat the population as unverified. 4. Check the reported statistic's magnitude for plausibility given the intended (usually much smaller) subset — extreme p-values (e.g., 1e-100 or smaller) usually signal a far larger n than intended. - **Discriminator**: A real violation is a missing or incorrect filter/direction that changes which rows enter the computation; a look-alike that is fine is a script that applies the filters in a different but equivalent way (e.g., filtering at load time, using a query string) and can show the reduced counts/subset consistent with the task description. - **Consequence**: The reported p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample, so it will not match the expected value in the output file and the check fails even though the file format is correct.
task -- the task says to write results into an output file "following the exact structure/format of" a provided sample/template file, and the scripts build that output.375545aa1e68 · mined from da-code dacode-dm-csv-011### Output template file never actually inspected - **Applies when**: `task` -- the task says to write results into an output file "following the exact structure/format of" a provided sample/template file, and the scripts build that output. - **Pattern**: The attempt never loads or prints the sample/template file; it hard-codes guessed column names, column order, row ordering, rounding/units and header text from the prose of the task, then asserts in the answer that the output "matches the required format." - **Detection procedure**: 1. In the task, note that a reference/sample output file is supplied and is the authority on schema and formatting. 2. Search the scripts for any read/print/comparison of that sample file (e.g., loading it, checking its columns, dtypes, row count, decimal places, sort order). 3. If absent, check whether the emitted frame's column names, column order, sort order and numeric formatting are instead invented in code or copied from the task wording. 4. Check the answer for unverified claims of format compliance (no printed diff against the template). - **Discriminator**: A real violation is when no code path ever reads the template, so agreement with it is pure luck; it is fine if the script reads the template (or explicitly reindexes/renames/rounds/sorts to the template's columns and formatting) and prints a shape/column/dtype comparison, even if the final naming happens to be hard-coded afterwards. - **Consequence**: The graded file is judged WRONG/MISSING on a strict file comparison — mismatched header names or order, wrong row ordering, or unrounded/differently scaled values — even when the underlying aggregation logic is right.
task -- the task requires writing results to a named file with an explicitly specified column naming scheme and one row per input record, and the scripts generate that file programmatically.2fad5c23094f · mined from da-code dacode-ml-cluster-014### Unverified output artifact: schema/content of the saved file never checked against the requested spec - **Applies when**: `task` -- the task requires writing results to a named file with an explicitly specified column naming scheme and one row per input record, and the scripts generate that file programmatically. - **Pattern**: The attempt builds a feature matrix after silently dropping/imputing/encoding columns, fits a model, writes the file, and then reports a prose summary (counts, chosen hyperparameter, feature legend) without ever re-reading the written file to confirm it exists at the expected path and that its header names, column count, row count, and label values match the requested format; degenerate or minimal-complexity results (e.g., the smallest possible number of groups) are accepted without a sanity check. - **Detection procedure**: 1. From the task, list the required artifact path, the exact column-naming convention, the expected number of rows (records) and the expected value semantics of the label column. 2. In the scripts, locate the write call and check whether the DataFrame passed in has columns generated to match the convention exactly (no index column, no leftover original names, no extra/missing feature columns relative to the matrix actually clustered) and whether every input record survives preprocessing (no silent row drops from NaN handling or filtering). 3. Check for a read-back/assert step after writing: does any code load the file and print/verify shape, header, row count, and label distribution? Compare that to what the final answer claims. 4. Inspect the reported result for degeneracy or implausibility (a single dominant group, the minimum possible number of clusters, row count ≠ dataset size) that no validation step addressed. - **Discriminator**: A real violation is when the answer's claims about the file are asserted from in-memory variables or narrative only, with no post-write verification and no shape/format assertion — or when preprocessing changed the row/column set without reconciling it to the spec. It is *not* a violation if the script (or a follow-up run) reloads the artifact and asserts the header pattern, row count equal to the number of input records, and valid label values, even if the modeling choices are debatable. - **Consequence**: The grader reads the artifact and finds it missing, misnamed, mis-headered, or with the wrong number of rows/columns (or a degenerate labeling), scoring the file check WRONG/MISSING despite a confident-sounding summary.
task -- the task requires writing a prediction/result file whose format is defined by a sample/template file, and the scripts generate that file at a hard-coded location.head() or shapes without comparison does not count.5c50253d9364 · mined from da-code dacode-ml-competition-009### Submission artifact never validated against the provided template (path, columns, ids, dtype) - **Applies when**: `task` -- the task requires writing a prediction/result file whose format is defined by a sample/template file, and the scripts generate that file at a hard-coded location. - **Pattern**: The attempt builds the output frame from its own assumptions (its own column names, its own output directory, its own value dtype such as forcibly rounding/casting continuous predictions), prints a few rows as "verification", and never programmatically loads the template to confirm the file is written where the task expects, with the same column names/order, the same number and set of identifiers, and value types consistent with the evaluation metric. - **Detection procedure**: 1. In the task/README, note the required output filename, its expected location relative to the working directory, and the template file that defines its schema. 2. In the scripts, find every write of the output file: check the path string, the constructed column names/order, and any post-processing of predictions (rounding, clipping, int casting, sorting). 3. Check whether any script actually reads the template and asserts equality of columns, row count, and id set/order against the written file — printing `head()` or shapes without comparison does not count. 4. In the answer, check whether it states the verified output location and schema match, or only narrates modeling choices and CV scores. - **Discriminator**: A real violation is when no code compares the produced file to the template/required path, or when values are transformed in a way the metric does not ask for (e.g., integer rounding of a continuous/log-scale target). It is fine if the script asserts column equality, id alignment and row counts against the template (even implicitly by copying the template and overwriting the prediction column) and writes to the location the task specifies. - **Consequence**: The grader reports the expected result file as missing or wrong (not found at the expected path, mismatched columns/ids, or degraded score from unnecessary value transformation), so the attempt scores 0 despite a plausible-looking model and CV metrics.
task -- the task asks for a per-record output file (labels, predictions, scores) covering a dataset that contains missing values or mixed dtypes.dropna() (or restricts to rows/columns that are fully populated), so the produced file contains far fewer records than the input dataset, and the answer reports the reduced count as if it were the full result.dropna/subsetting before the model fit and for whether the output is written from the reduced frame.e8f839fe2e4e · mined from da-code dacode-ml-cluster-009### Dropping rows with missing values instead of imputing, shrinking the required output - **Applies when**: `task` -- the task asks for a per-record output file (labels, predictions, scores) covering a dataset that contains missing values or mixed dtypes. - **Pattern**: The script handles missing data with a blanket `dropna()` (or restricts to rows/columns that are fully populated), so the produced file contains far fewer records than the input dataset, and the answer reports the reduced count as if it were the full result. - **Detection procedure**: 1. Read the task for the required output granularity — does it implicitly require one row per input record, and does it state any filtering? If no filtering is stated, every input record must appear. 2. In the script, look for `dropna`/subsetting before the model fit and for whether the output is written from the reduced frame. 3. Compare the row count claimed in the answer (or in the written file) against the raw dataset's row count; also check the number of feature columns against the number of usable numeric columns. 4. Flag if rows were silently lost and no imputation (mean/median/etc.) or justification was applied. - **Discriminator**: A real violation is unrequested row loss that changes the output's coverage; it is fine if the task explicitly asks to filter/subset, or if only a couple of records are dropped for a documented reason **and** the task does not require complete coverage. - **Consequence**: The output file has the wrong shape/row count and cluster labels that cannot be aligned to the expected per-record results, so the file comparison fails outright.
task -- the task specifies an exact output file with an exact set/naming of columns, and the script builds that file from an intermediate DataFrame.insert, copy of source columns, reset_index) and every rename mapping, plus the index= argument.columns.tolist() and compare to the spec, or merely print a head and declare success? Also check the answer's described pipeline matches the script actually run.de25d1ca3a10 · mined from da-code dacode-ml-cluster-016### Output schema deviation: extra/renamed columns in the required result file - **Applies when**: `task` -- the task specifies an exact output file with an exact set/naming of columns, and the script builds that file from an intermediate DataFrame. - **Pattern**: The script writes the file but adds identifier/bookkeeping columns, keeps original names, mis-indexes the required suffix numbering, or otherwise emits a superset/variant of the requested schema instead of exactly the stated columns (and does not assert the schema before saving). - **Detection procedure**: 1. From the task statement, write down the exact required column list (names, naming convention, and whether anything else is allowed) and any index/ordering requirement. 2. In the script, trace the DataFrame that is passed to the write call: list every column added (`insert`, `copy` of source columns, `reset_index`) and every rename mapping, plus the `index=` argument. 3. Compare that final column list to the required list; also check the numbering convention starts/increments as the task implies and that no ID/label leftovers survive. 4. Check the answer/verification output: does it print the saved file's `columns.tolist()` and compare to the spec, or merely print a head and declare success? Also check the answer's described pipeline matches the script actually run. - **Discriminator**: A real violation is a saved file whose column set differs from the specification (extra column, unrenamed column, off-by-one naming, index written as a column). A look-alike that is fine: extra columns exist only in in-memory/intermediate frames or in separate diagnostic files, while the required file contains exactly the specified columns. - **Consequence**: The grader reads the result file, fails the schema/column check (or mis-aligns the feature vector), and marks the expected file WRONG/MISSING regardless of the clustering quality.
task -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation.exclude list, select_dtypes, or all-NaN filtering) despite being predictive of the target.train_test_split (shuffled) / random CV on rows that are sequential in time, rather than a chronological or block split matching the test period?fa642c18c0f2 · mined from da-code dacode-ml-regression-002### Validation split and feature set that don't mirror the actual prediction setting - **Applies when**: `task` -- a model must produce predictions for a supplied test file whose rows come from an ordered (e.g., time-indexed) process, and the scripts choose which of the shared columns to use as features and how to hold out data for validation. - **Pattern**: The attempt builds features by blanket-excluding columns (dropping some that exist in *both* train and test and are strongly related to the target) and then measures quality with a random/shuffled train-test split of ordered data, reporting the resulting optimistic error/R² as evidence the submission is good — without ever checking predictions against a simple baseline or against the target's own distribution on a held-out block that resembles the test rows. - **Detection procedure**: 1. From the task and test file schema, list the columns actually available at prediction time; then read the scripts' feature-selection code and note any available column that is excluded or silently dropped (e.g., via an `exclude` list, `select_dtypes`, or all-NaN filtering) despite being predictive of the target. 2. Inspect the holdout logic: does it use `train_test_split` (shuffled) / random CV on rows that are sequential in time, rather than a chronological or block split matching the test period? 3. Check whether the answer's reported metric is the only evidence of correctness — i.e., no baseline comparison (persistence/mean/available forecast column) and no sanity check that predicted distribution (mean, range, count, ordering) matches the training target and the expected output rows. 4. Flag if (1) or (2) holds and (3) holds. - **Discriminator**: A real violation excludes usable, task-legitimate predictors and/or validates in a way that leaks temporally adjacent rows, so the quoted metric cannot be trusted; a look-alike that is fine excludes only columns genuinely absent/unusable in the test file (true leakage or all-missing), and validates on a chronological holdout that reproduces the test-time information set, with a baseline comparison reported. - **Consequence**: Validation metrics look strong (low MAE, high R²) while the submitted predictions are systematically off on the real test rows, so the graded file fails the accuracy/tolerance check despite the correct file name and row count.
task -- the task points to an external format spec (e.g. a YAML/JSON config) and grading depends on machine-readable artifacts (plot data/series arrays) in addition to a rendered image.cd5c52df7363 · mined from da-code dacode-plot-line-015### Incomplete deliverables for a spec-driven plotting task - **Applies when**: `task` -- the task points to an external format spec (e.g. a YAML/JSON config) and grading depends on machine-readable artifacts (plot data/series arrays) in addition to a rendered image. - **Pattern**: The attempt produces only the image and narrates that it "follows the spec", without saving the accompanying data/spec artifacts the grader expects, and without keeping reproducible scripts; internal inconsistencies (e.g. a title naming one date range while the described data covers another) are left unresolved. - **Detection procedure**: 1. Read the task and any referenced spec file; enumerate every required output artifact (image, serialized plot spec, numeric array of plotted values) and every named property (figsize, color, title, axis labels, ticks, ordering, aggregation level). 2. Check the working directory / scripts for each enumerated artifact actually being written, and check that a script exists that reproduces them from the raw data. 3. Cross-check the answer's claimed properties against the spec verbatim (string equality of title/labels, tick list, series length) and against the described data range for contradictions. 4. Verify the plotted series is derived with the stated granularity and ordering (e.g. correct time aggregation, sorted x-values, no dropped or duplicated periods) rather than asserted. - **Discriminator**: A real violation is missing required artifacts or a property that mismatches the spec text (including inconsistent date ranges/labels); it is fine if all required files are written and the only differences are cosmetic extras (markers, gridlines, legend) not constrained by the spec. - **Consequence**: Grader checks on the expected data artifacts report WRONG/MISSING and the attempt scores 0 even though an image exists.
task -- the task names a specific output file/format (e.g., a result file matching a provided sample) and points to input data that the scripts must locate and load.to_csv, open(..., 'w'), etc.) and check the written schema against the provided sample; also check whether the sample file was read/inspected at all.except: pass / if df is None: branch that generates random or literal numbers? Confirm the script would fail loudly rather than proceed on fake data.np.random.normal values, a p-value of exactly 0) to see whether the answer came from the fallback.c0919b027b92 · mined from da-code dacode-data-sa-028### Missing required output artifact (and fabricated fallback data instead of the real inputs) - **Applies when**: `task` -- the task names a specific output file/format (e.g., a result file matching a provided sample) and points to input data that the scripts must locate and load. - **Pattern**: The scripts implement the statistical/modeling logic but never write the requested file; input loading is done by guessing candidate paths with a silent fallback that synthesizes or hard-codes plausible data, so the reported numbers come from invented inputs and the deliverable file is absent. The final answer is a prose/console report instead of the specified artifact. - **Detection procedure**: 1. From the task, list every required deliverable: exact output filename, expected columns/rows/ordering, rounding/units, and the exact input file(s) to be used. 2. Grep the scripts for any write of that deliverable (`to_csv`, `open(..., 'w')`, etc.) and check the written schema against the provided sample; also check whether the sample file was read/inspected at all. 3. Trace the data-loading path: does it read the actual provided file(s), or try a list of guessed paths / `except: pass` / `if df is None:` branch that generates random or literal numbers? Confirm the script would fail loudly rather than proceed on fake data. 4. Compare the reported numbers to the branch that could have produced them (e.g., suspiciously round sample sizes, seeded `np.random.normal` values, a p-value of exactly 0) to see whether the answer came from the fallback. - **Discriminator**: A real violation is when no code path writes the named file with the sample's schema, or the reported values could only come from synthetic/hard-coded inputs. It is fine if the script robustly searches for the input but raises on failure, and does write the deliverable with the required columns/precision — even if the analysis logic is only in a helper module or the console also prints a summary. - **Consequence**: The grader finds the expected result file missing or containing values derived from invented data, so all output checks fail regardless of whether the statistical method was correct.
task -- The task points to auxiliary instruction/config files (e.g., a tips/README/spec text file, a YAML/JSON formatting config) and/or expects a set of deliverable files beyond the obvious main output.open, read_csv, yaml.safe_load, json.load) and for writes of each expected output; note any file mentioned in the task but absent from the code.06f5882f6230 · mined from da-code dacode-plot-line-006### Ignoring referenced specification files and required output artifacts - **Applies when**: `task` -- The task points to auxiliary instruction/config files (e.g., a tips/README/spec text file, a YAML/JSON formatting config) and/or expects a set of deliverable files beyond the obvious main output. - **Pattern**: The scripts never open, parse, or echo the referenced instruction/config files; the agent instead hardcodes its own assumptions (grouping definitions, columns, titles, colors, axis labels, figure size, ranges) and writes only the single most obvious artifact, then declares success by paraphrasing the spec it never actually read. - **Detection procedure**: 1. List every external file the task text references (instruction files, config/format files) and every output file implied or named. 2. Grep the scripts for reads of each referenced file (`open`, `read_csv`, `yaml.safe_load`, `json.load`) and for writes of each expected output; note any file mentioned in the task but absent from the code. 3. Check whether formatting/derivation choices in the plotting or aggregation code trace back to values loaded from the config, or are literals invented by the agent. 4. Compare the answer's claims (e.g., "formatted according to the spec", "definitions from the tips file") against step 2 — claims about content of unread files are unsupported. - **Discriminator**: A real violation is code that contains no read of a task-referenced file, or omits an expected artifact entirely. It is *not* a violation if the script reads the config and then legitimately falls back to defaults for keys the config doesn't specify, or if the extra artifacts are written by a separate step visible elsewhere in the run. - **Consequence**: Graders that compare each required artifact (image, serialized plot spec, numeric array) find them missing or mismatched in styling, labels, series definitions, or data values, so all checks fail even if the underlying computation looks plausible.
task -- the instructions point to an external document (README/spec/config) that defines how to bin, filter, map, or order values before producing the requested output.pd.cut bins, list of labels) that materializes those rules; if no script exists, treat the derivation as unverifiable.59487d5f27e9 · mined from da-code dacode-plot-bar-005### Spec file referenced by the task is never read or implemented - **Applies when**: `task` -- the instructions point to an external document (README/spec/config) that defines how to bin, filter, map, or order values before producing the requested output. - **Pattern**: The agent skips or paraphrases the spec, reuses the raw categories/values already present in the data (or its own invented grouping), and asserts in the answer that "the spec was followed" without any script step that opens the spec file or encodes its rules; often no reproducible script is saved at all. - **Detection procedure**: 1. From the task, list every referenced auxiliary document and the exact rules it is supposed to supply (bin edges, labels, exclusions, ordering, required output artifacts). 2. Search the scripts for a read of that document and for an explicit mapping/binning structure (dict, `pd.cut` bins, list of labels) that materializes those rules; if no script exists, treat the derivation as unverifiable. 3. Compare the category labels/counts in the answer against the raw distinct values of the source field: if they are identical (or the totals equal the unfiltered row count when the spec implies grouping/filtering), the spec was not applied. 4. Check that every artifact the task/spec implies (plot plus any data dumps) is actually written by the script, not just described in prose. - **Discriminator**: A genuine violation is when no code path encodes the spec's rules and the output categories coincide with the raw field's own levels; it is fine if the code hard-codes bins that are demonstrably transcribed from the spec (labels/edges match) even without re-reading the file at runtime. - **Consequence**: Saved outputs (image and any numeric/JSON dumps) have wrong bin labels, wrong bin counts, or are missing entirely, so every file-level check fails despite a confident "task completed" report.
task -- the task states a specific preprocessing step and a specific answer format/output file, and the agent supplies a final answer with no (or non-runnable) script that produces it.225b4ad6590e · mined from da-code dacode-di-text-002### Missing reproducible script and required output artifact - **Applies when**: `task` -- the task states a specific preprocessing step and a specific answer format/output file, and the agent supplies a final answer with no (or non-runnable) script that produces it. - **Pattern**: The agent reports a value it believes it read off the data (or recalls), without a saved, end-to-end script that loads the raw file, coerces the relevant column to numeric (stripping symbols/thousand separators/percent signs), applies the stated imputation, computes the requested extremum/statistic, and writes the result to the exact requested artifact (e.g. the named JSON/CSV) in the exact requested schema. Nothing in the deliverables lets a reviewer re-derive the number, and the required file is never created. - **Detection procedure**: 1. List every deliverable the task demands: the output file name/location, the key names, the value types, and any mandated preprocessing step. 2. Look for a script whose code path visibly does each of those: raw load → dtype cleaning of the target column → the mandated imputation → the selection/aggregation → serialization to the named file. 3. If any link is absent (no script at all, or the script prints to stdout only, or never touches dtypes/imputation), mark inadequate; also check the reported number is a plausible value of the stated unit (e.g. a percentage inside 0–100) and that it belongs to the row that a numeric — not lexicographic — comparison would select. 4. Confirm the emitted structure matches the requested shape exactly (same keys, list-vs-scalar, rounding). - **Discriminator**: A real violation is when the answer cannot be traced to executed code that both preprocesses as instructed and writes the named artifact. A look-alike that is fine: a short but complete script that does all steps and dumps the file, even if the agent additionally quotes the answer in prose. - **Consequence**: The grader finds the expected result file missing or containing a value derived from uncleaned/unimputed data (e.g. a string-max or a pre-imputation row), so the check fails with 0/1 even though the prose answer looks well-formatted.
task -- The task names a specific, textbook-named statistic, coefficient, or metric variant (e.g. a particular "first/second coefficient", a specific averaging or normalization convention) that the script must compute by hand.a6e90dea273c · mined from infiagent-dabench dabench-359### Named statistic implemented with the wrong formula variant - **Applies when**: `task` -- The task names a specific, textbook-named statistic, coefficient, or metric variant (e.g. a particular "first/second coefficient", a specific averaging or normalization convention) that the script must compute by hand. - **Pattern**: The script hard-codes one plausible variant of the named quantity (or calls a library default) without checking that it matches the exact named definition, often with a comment that asserts the mapping rather than verifying it — e.g. implementing the median-based sibling formula while labeling it as the mode-based one, or using sample vs. population normalization inconsistently with the definition. - **Detection procedure**: 1. From the task statement, write down the exact canonical definition of the named statistic, including which central-tendency/normalization terms it uses. 2. Read the script's arithmetic expression (not its comments or variable names) and map each term to the canonical definition. 3. Check whether an alternative, similarly named variant exists that the script may have substituted; if so, require evidence in the script that both were computed/compared or that the chosen one was justified. 4. Confirm downstream classification/rounding uses the value from the correct formula. - **Discriminator**: A real violation is when the coded expression is a *different* named quantity (different terms, different denominator convention) than the one requested; a look-alike that is fine is an algebraically equivalent rewriting, or a negligible convention choice (e.g. ddof) that provably cannot change the reported rounded value or the qualitative conclusion. - **Consequence**: The qualitative label (sign/direction/category) may still match by luck, but the numeric field differs from the reference, so the answer fails the value check while passing the type check — partial credit at best.
task -- a script aggregates several raw columns row-wise (weighted sums, means, cumulative products) and a second "verification" script is used to confirm the saved output..sum(axis=1), .mean()), then "verifies" by recomputing with the identical code path and reporting a match — so any shared assumption error (NaNs treated as zero, percent vs. decimal returns, dropped/misaligned column, wrong starting baseline or compounding convention) is confirmed rather than caught.isna().sum(), dtype check, min/max/scale check, row/column count check, or explicit NaN policy in the aggregation — and whether the aggregation is done with an operator that skips NaNs by default.8499bdc410cc · mined from da-code dacode-dm-csv-050### Circular "verification" that re-runs the same code instead of validating input assumptions - **Applies when**: `task` -- a script aggregates several raw columns row-wise (weighted sums, means, cumulative products) and a second "verification" script is used to confirm the saved output. - **Pattern**: The attempt reads the raw file with no inspection of missing values, dtypes, units/scale, or a leading initialization row, aggregates with a reducer that silently ignores NaNs (e.g. `.sum(axis=1)`, `.mean()`), then "verifies" by recomputing with the identical code path and reporting a match — so any shared assumption error (NaNs treated as zero, percent vs. decimal returns, dropped/misaligned column, wrong starting baseline or compounding convention) is confirmed rather than caught. - **Detection procedure**: 1. In the task/README, list the assumptions the computation depends on: units/scale of the input values, expected number of rows/columns, and the exact definition and starting point of the requested output quantity. 2. In the scripts, check whether the raw input is ever validated — `isna().sum()`, dtype check, min/max/scale check, row/column count check, or explicit NaN policy in the aggregation — and whether the aggregation is done with an operator that skips NaNs by default. 3. Check whether the verification script does anything independent (recompute one row by hand from raw numbers, compare against a known benchmark, check output ranges/row counts/column names against the required format) or merely re-executes the same formula. 4. Inspect the answer for tell-tale signs: implausibly smooth/small magnitudes, no NaN or zero at the series start, row count differing from the source file, or values that would be off by ~100x under a unit mismatch. - **Discriminator**: A real violation is when *no* check exists that could fail independently of the main computation (data-quality assertions absent and the verifier duplicates the formula). It is fine if the script asserts input completeness/scale, handles NaNs explicitly, or cross-checks at least one output value against an externally derived or hand-computed reference — even if the code is otherwise reused. - **Consequence**: All internal checks "pass" while the saved file diverges from the reference (shifted baseline, NaN-as-zero contamination, wrong scale or row alignment), and the grader marks the output file WRONG on numeric comparison.
task -- the task prescribes an exact answer template containing a structured literal (dict/list/tuple) with string keys or labels, e.g. @name[{'k_1':v_1, ...}].name= prefix, or otherwise deviates from the template, so an exact/parse-based grader cannot match it even though the numbers are right.= or [], brace type, quoting of keys, separators, and key naming scheme.repr()/print(dict) with string keys, or a hard-coded template) rather than a loop/f-string that strips quotes.ast.literal_eval; if it raises or yields different key types than the template, flag it.6c69bc6e9f3b · mined from infiagent-dabench dabench-450### Dict/collection answers emitted without literal-syntax fidelity (unquoted keys, altered delimiters)
- **Applies when**: `task` -- the task prescribes an exact answer template containing a structured literal (dict/list/tuple) with string keys or labels, e.g. `@name[{'k_1':v_1, ...}]`.
- **Pattern**: The agent computes correct values but serializes the container by printing a Python object, f-string, or hand-typed text that drops the quotes around keys, swaps bracket/brace types, omits the `name=` prefix, or otherwise deviates from the template, so an exact/parse-based grader cannot match it even though the numbers are right.
- **Detection procedure**:
1. Copy the answer template from the task statement verbatim and note every literal character: variable name, `=` or `[]`, brace type, quoting of keys, separators, and key naming scheme.
2. In the scripts, find the line that builds the final answer string and check whether it uses a faithful literal serialization (e.g. `repr()`/`print(dict)` with string keys, or a hard-coded template) rather than a loop/f-string that strips quotes.
3. Compare the submitted answer token-by-token against the template; flag any mismatch in quoting, key spelling/order, brackets, or prefix.
4. Optionally, attempt to parse the submitted payload with `ast.literal_eval`; if it raises or yields different key types than the template, flag it.
- **Discriminator**: A real violation is a syntactic/format deviation from the stated template (unquoted or renamed keys, wrong bracket, missing prefix, wrong ordering, wrong rounding presentation). A look-alike that is fine is cosmetic whitespace or trailing-zero differences that still parse to the same literal object with the same key strings and values.
- **Consequence**: The grader reports the expected mapping as WRONG/MISSING and scores 0 despite numerically identical values.task -- the task asks for a distributional statistic or hypothesis test on a single numeric column (normality test, skewness, kurtosis, moments) and prescribes a specific test, threshold, or value to report.fbe935ae176c · mined from infiagent-dabench dabench-298### Missing evidence that the statistic was computed on the correctly cleaned data vector - **Applies when**: `task` -- the task asks for a distributional statistic or hypothesis test on a single numeric column (normality test, skewness, kurtosis, moments) and prescribes a specific test, threshold, or value to report. - **Pattern**: The attempt reports only the final verdict/numbers with no saved script and no intermediate diagnostics, so the vector actually fed to the test is unverifiable: NaNs, missing-value sentinels (e.g. -999, 0, blanks), strings coerced to numbers, or extra rows outside the intended subset can silently enter and dominate the moments, flipping the test outcome. - **Detection procedure**: 1. Read the task and list every required output and diagnostic (e.g. the test name, alpha, the p-value to be reported, rounding rules). 2. Inspect the scripts for an explicit, reproducible chain: load → select column → coerce dtype → drop/handle missing and sentinel values → print n before and after cleaning → run the prescribed test → print p-value, skewness, kurtosis. 3. Check the answer against these prints: is the reported p-value present, is the verdict consistent with the stated alpha, and are the moment values plausible for the printed n and value range? 4. Flag if any script is missing, if n and cleaning steps are never printed, or if a required diagnostic (p-value) is absent from the report. - **Discriminator**: A genuine violation is an answer with no reproducible cleaning/diagnostic trail, or one whose extreme moment values (e.g. large |skew|, heavy kurtosis) are asserted without any check for sentinels/outliers that would explain them; it is *not* a violation if the script prints row counts, shows the column is clean numeric, reports the p-value, and the extreme moments are corroborated by printed summary statistics. - **Consequence**: The test is run on a contaminated or wrong-length vector, so the normality verdict and the skewness/kurtosis values differ from ground truth and every graded field fails, with no artifact left to diagnose the discrepancy.
task -- the task asks for predictions on an unlabeled evaluation file and the scripts train a model on a labeled training file, with grading presumably based on prediction quality against hidden labels.c162df6dc85b · mined from da-code dacode-ml-multi-011### No held-out validation of predictive quality before submitting predictions - **Applies when**: `task` -- the task asks for predictions on an unlabeled evaluation file and the scripts train a model on a labeled training file, with grading presumably based on prediction quality against hidden labels. - **Pattern**: The attempt fits a single default/lightly-tuned baseline (e.g., one vectorizer + one simple classifier with hand-picked hyperparameters), never splits off a labeled validation set or runs cross-validation, and reports only formatting facts (row count, column name) and the predicted class distribution as evidence of success — so no one knows whether the model clears the accuracy bar the grader uses. - **Detection procedure**: 1. Read the task: confirm the deliverable is per-row predictions judged against hidden ground truth, not just a file with the right shape. 2. Read the scripts: look for any train/validation split, cross-validation, or scoring call on labeled data; also check whether more than one model/feature setting was compared. 3. Read the answer: check whether it states a measured quality estimate (accuracy/F1 on held-out labeled data) rather than only shapes, column names, and label frequencies. 4. If steps 2–3 find no measured score, flag: the attempt has no evidence its predictions are better than chance-level or an unaccepted baseline. - **Discriminator**: A real violation is the total absence of any quantitative held-out estimate (or a score computed on the same rows used for training, which is equally uninformative). It is *not* a violation if the attempt reports a legitimate held-out/CV score and simply chose a simple model, nor if the task explicitly only checks file format; comparing predicted vs. training label distributions alone is not a substitute for a score. - **Consequence**: The output file has the correct shape and column name but low agreement with the hidden labels, so an accuracy-threshold check marks the result file WRONG while the agent's self-report claims success.
task -- the task asks for predictions on a held-out set written to a submission file, and the scripts fit a model and write the file.5d53bf8aed41 · mined from da-code dacode-ml-competition-008### Declaring success on format checks alone, with no held-out performance validation - **Applies when**: `task` -- the task asks for predictions on a held-out set written to a submission file, and the scripts fit a model and write the file. - **Pattern**: The attempt validates only superficial properties of the output (row count, column names, value range, no NaNs, mean roughly matching the target mean) and reports model hyperparameters as evidence of quality, while never computing the competition-relevant metric on a held-out/validation split, never comparing to a trivial baseline (e.g., predicting the target mean), and often silently training on a truncated subsample of the available rows or a reduced feature set. - **Detection procedure**: 1. Read the task to identify the evaluation target and the implied scoring metric (e.g., regression error/rank correlation/AUC) and whether all training rows/features are available. 2. In the scripts, look for (a) a train/validation split with the metric computed on the validation part, (b) a comparison against a constant/naive baseline, and (c) whether the fit uses the full training data and the same feature set as the test-time transform. 3. In the answer, check whether any quantitative out-of-sample score is reported, or whether only distributional/format statistics and hyperparameters are cited as "good". 4. Compare the reported spread of predictions to the spread of the training target; a much narrower prediction range with no validation score is a red flag for severe underfitting/undertrained model. - **Discriminator**: A real violation reports no out-of-sample metric at all (or only format/range checks) so predictive quality is unknown; a look-alike that is fine reports a validation score computed on data excluded from fitting, ideally alongside a baseline score, even if the final answer text also mentions format checks. - **Consequence**: The submission file is well-formed but scores at or near a naive baseline (or below the grader's accuracy threshold), so the correctness check on the predictions fails despite all self-reported checks passing.
task -- the task asks for a metric/statistic computed separately for groups defined by numeric ranges (below X, between X and Y, above Y) of some column.<, <=, between), the dtype of the grouping column (numeric vs string/categorical), and whether the boundary categories are assigned to exactly one group.c93e41af5a54 · mined from infiagent-dabench dabench-513### Ambiguous or unverified subgroup boundaries when a statistic is requested per range-defined bin - **Applies when**: `task` -- the task asks for a metric/statistic computed separately for groups defined by numeric ranges (below X, between X and Y, above Y) of some column. - **Pattern**: The attempt picks one plausible binning without making inclusivity explicit or checking it — e.g. treating "between X and Y" as an open interval that drops rows exactly at the boundaries (or as closed, double-counting them), grouping on a column whose values are strings/half-steps so comparisons or matches silently misclassify rows, and never printing per-group row counts so that the groups do not sum to the filtered population. - **Detection procedure**: 1. Read the task and list the requested groups and their boundary values; note that boundary rows are typically a large share of a coarse rating/score scale. 2. In the scripts, locate the filtering/binning code: check the comparison operators (`<`, `<=`, `between`), the dtype of the grouping column (numeric vs string/categorical), and whether the boundary categories are assigned to exactly one group. 3. Verify the script prints, per group, the row count after dropping nulls in both correlated columns, and that these counts sum to the total non-null population of the filtered set (no dropped or duplicated boundary rows). 4. Check the answer: if the middle/edge group's value is implausibly close to a neighbouring group's value, or group sizes were never reported, treat the binning as unvalidated. - **Discriminator**: A real violation is a binning where boundary-valued rows are excluded or misassigned, or where group counts are never shown; it is fine if the script documents the chosen convention, assigns every row in the filtered population to exactly one group, and reports counts that reconcile with the total (even if the convention is debatable, the reconciliation exposes it). - **Consequence**: Groups adjacent to the boundaries are computed on the wrong subset, so their correlation coefficients differ materially from the reference values while the unambiguous group matches — a partial-credit failure like 1/3 checks passed.
task -- the task references an explicit definition, eligibility rule, or a provided sample/template output file, and part of that specification is incomplete, truncated, or not obviously reproduced in the scripts.162f575169db · mined from da-code dacode-dm-csv-009### Invented qualification thresholds / unverified output spec when the task's definition is ambiguous or truncated - **Applies when**: `task` -- the task references an explicit definition, eligibility rule, or a provided sample/template output file, and part of that specification is incomplete, truncated, or not obviously reproduced in the scripts. - **Pattern**: The attempt guesses a cutoff (e.g., a made-up minimum count applied on a self-chosen basis such as per-item mean instead of total), applies that same filter to every sub-ranking regardless of whether the rule pertains to it, and never opens the supplied sample/template to confirm column names, ordering, key type, and row conventions. - **Detection procedure**: 1. Read the task/README and list every explicitly stated rule (qualification criteria, aggregation level, rounding, units, ordering) and every referenced auxiliary file (sample/template, dictionary, docs). 2. Search the scripts for code that loads or prints each referenced auxiliary file and for constants that encode the stated rules; flag any numeric threshold or filter that appears without a documented source. 3. Check whether each requested output quantity is computed with the definition scoped to it, i.e. whether the eligibility filter is applied only where the specification says it applies, and on the stated statistic (sum vs. mean vs. count). 4. Compare the emitted file's header/ordering/index against the sample file's actual contents (not an assumed layout); flag if never compared. - **Discriminator**: A real violation is a threshold, aggregation choice, or output layout that the scripts never derive from the data, README, or sample file — it is hard-coded on intuition. It is fine if the agent inspects the sample/spec, documents the source of the rule, tests the sensitivity of the ranking to plausible alternative interpretations, or the rule genuinely has a single unambiguous reading in the provided text. - **Consequence**: The rankings are computed on a differently filtered population (and/or with mismatched column names/ordering) than the reference, so exact-match file comparison fails on all rows even though the underlying aggregation code is arithmetically correct.
task -- the task points to an external specification file (config/YAML/JSON) for output formatting or binning and/or implies a set of output artifacts beyond the headline file.be758efe9c22 · mined from da-code dacode-plot-bar-007### Spec file / required artifacts asserted rather than actually read and produced - **Applies when**: `task` -- the task points to an external specification file (config/YAML/JSON) for output formatting or binning and/or implies a set of output artifacts beyond the headline file. - **Pattern**: The attempt hard-codes its own titles, labels, bin edges, styling and category set, describes them in the final answer as if they came from the spec, and emits only the single obvious output file while silently skipping the other expected artifacts (serialized plot data, saved arrays); the derived group counts are never checked against the total row count of the filtered subset. - **Detection procedure**: 1. From the task text, list (a) every file that must be *read* as a specification and (b) every artifact that must be *written*. 2. In the scripts, confirm the spec file is actually loaded and its values are used to drive the plot (bins, labels, title, order, figure params) instead of literals appearing in the code; confirm each expected artifact is written by an explicit save call. 3. Compare the answer's stated categories/counts with the spec's own categories and with the dataset: does the number of groups match the spec, and do the group counts sum to the size of the filtered subset (and to a plausible fraction of the raw file)? 4. Flag if any spec value is invented, any artifact is unwritten, or the counts fail the sum/magnitude sanity check. - **Discriminator**: A fine attempt reads the spec and its literals coincide with the spec values (verifiable by comparing code/output to the file contents) and writes all requested artifacts; a violation is code whose formatting/binning choices exist nowhere in the spec, or an answer that claims "all specifications applied" without the spec ever being parsed, or an obviously truncated subset count with no reconciliation. - **Consequence**: The grader checks each expected artifact (plot metadata, image, saved array); missing files and mismatched bins/counts make every check fail even though the script ran without error.
task -- the task specifies how the final answer must be ordered/formatted (e.g., "sort from highest to lowest") and/or written to a named result file, and the scripts produce ranked lists or dictionaries.5623989d04a7 · mined from da-code dacode-di-text-003### Stated output ordering/serialization constraint not enforced in code - **Applies when**: `task` -- the task specifies how the final answer must be ordered/formatted (e.g., "sort from highest to lowest") and/or written to a named result file, and the scripts produce ranked lists or dictionaries. - **Pattern**: The scripts extract the right records but leave them in whatever order the selection helper returns (e.g., an ascending "smallest-n" query for the low group while the high group is descending), and only print results instead of emitting the requested artifact/JSON structure — so the stated ordering/format constraint is never explicitly applied. - **Detection procedure**: 1. Read the task and list every explicit output constraint: ordering direction, grouping keys, rounding/units, and the required file name/JSON schema. 2. In the scripts, locate each list-producing step and check whether an explicit sort (or reverse) with the required direction is applied to *every* output list, not just some of them; check whether the final structure is serialized to the required file. 3. Compare the printed/reported lists against the constraint: for a "highest to lowest" requirement, verify the associated values decrease monotonically in each list. 4. Flag if any list's implied value order contradicts the requirement, or if no code writes the required output artifact. - **Discriminator**: A real violation is when the code contains no ordering/serialization step matching the stated requirement (or the reported values run the wrong way); it is fine if a selection helper already returns the required direction and the reported values verifiably follow it, and the required file is written with the requested keys. - **Consequence**: The grader compares the expected result file element-by-element and marks the answer wrong (or missing) even though the correct set of records was identified.
task -- The task specifies a rigid answer template with tagged fields (e.g. @field[value]) and shows the exact literal form the values should take (quoted strings, units, ranges, allowed vocabulary).6703d17fd1ed · mined from infiagent-dabench dabench-550### Answer string doesn't literally match the requested output template (quoting/delimiters/tokens) - **Applies when**: `task` -- The task specifies a rigid answer template with tagged fields (e.g. `@field[value]`) and shows the exact literal form the values should take (quoted strings, units, ranges, allowed vocabulary). - **Pattern**: The agent computes plausible or even correct values but emits them in a different surface form than the template demands — dropping the quotation marks shown in the spec, changing separators or spacing, reordering fields, adding extra prose/units, or paraphrasing an allowed label — so a literal string matcher scores every field wrong. This is often compounded by no saved script, so nothing else in the submission can be checked or re-derived. - **Detection procedure**: 1. Copy the answer-format line from the task verbatim and list each field, its delimiters, and exactly how the sample value is written (with or without quotes, capitalization, allowed token set). 2. Copy the agent's final answer string and diff it character-by-character against that template, field by field and in order. 3. Check that each value is drawn from the task's stated vocabulary/format (e.g. the literal range string, the enumerated category labels) rather than a synonym, computed variant, or extra-annotated form. 4. Confirm a script exists that produces those exact strings (or that the answer is at least reproducible), rather than being hand-typed with no artifact. - **Discriminator**: A real violation is any deviation in the characters a strict matcher sees — missing/added quotes, wrong field order, extra text inside brackets. A look-alike that is fine is a difference the task explicitly leaves free (e.g. whitespace between fields when the spec shows none but no strictness is implied) or a value that is spelled exactly as the task's enumerated option even if worded differently from the agent's internal computation. - **Consequence**: The grader reports 0/N checks passed with "WRONG/MISSING" for fields whose semantic content is actually right, so a substantively correct analysis scores zero.
task -- the task names specific entities, groupings, measures, config files, or output artifacts, and the scripts operate on a file whose columns do not contain them.7e32c9b270d6 · mined from da-code dacode-plot-scatter-002### Substituting a "creative interpretation" for the requested entities and metrics - **Applies when**: `task` -- the task names specific entities, groupings, measures, config files, or output artifacts, and the scripts operate on a file whose columns do not contain them. - **Pattern**: Instead of locating the correct input data (or the correct columns) and producing every requested artifact, the attempt declares a "data mismatch" and maps the requested concepts onto unrelated available fields as proxies, then reports success on that redefined problem; stated config/settings and secondary output files are ignored or only nominally loaded. - **Detection procedure**: 1. From the task, list the required inputs (entities to group by, measure to rank by, quantities to aggregate), any settings file that must be applied, and every output file expected. 2. In the scripts, check whether each listed item is read from a real column/field with matching semantics, or is invented/renamed as a stand-in. 3. Check that the settings file's contents actually drive the output (figure size, labels, colors, ordering) rather than being loaded and discarded, and that all expected output artifacts are written. 4. In the answer, look for language such as "proxy", "interpretation", "mismatch", "mapped X to Y" — a self-declared redefinition of the task. - **Discriminator**: A genuine violation redefines what is being measured or grouped, or silently drops required outputs. A look-alike that is fine is when the correct fields exist but must be derived by documented, semantics-preserving steps (parsing dates into durations, joining tables, standardizing codes) with the requested measure and grouping preserved; also fine is exhausting the available data files first and documenting that the needed field is genuinely absent, rather than assuming it after inspecting one file. - **Consequence**: The saved figure and any numeric export encode the wrong entities, ordering, and units, so every value-level and file-level check fails (missing expected artifacts, mismatched arrays), even though a plot was produced.
task -- the task names a specific output file (or template/schema) that the deliverable must be saved as.to_csv, to_excel, savefig, etc.) and compare the literal string and the written frame's columns/index to step 1, character for character. 3. In the answer, check which filename and structure the agent claims to have produced. 4. Flag if any of the three differ, or if the agent never verifies the saved file by reading it back and comparing to the template.fbf0ce2a8103 · mined from da-code dacode-dm-csv-043### Output artifact name/format does not match the exact specification - **Applies when**: `task` -- the task names a specific output file (or template/schema) that the deliverable must be saved as. - **Pattern**: The agent computes plausible numbers but writes them to a file whose name, spelling, location, or column/index layout differs from the one literally requested (e.g., "corrected" spelling, different directory, missing index column or header names from the template), and the answer text describes the self-chosen name as if it were the requirement. - **Detection procedure**: 1. Copy the exact filename/path and any template column/index names from the task statement. 2. In the scripts, find every write call (`to_csv`, `to_excel`, `savefig`, etc.) and compare the literal string and the written frame's columns/index to step 1, character for character. 3. In the answer, check which filename and structure the agent claims to have produced. 4. Flag if any of the three differ, or if the agent never verifies the saved file by reading it back and comparing to the template. - **Discriminator**: A real violation is any deviation in the literal artifact name/path or in the template's column/index structure (including a "fixed" typo). Not a violation if the requested file is written exactly as specified and extra helper files also exist, or if the agent additionally saves an alias copy under the required name. - **Consequence**: The grader looks for the specified file and reports it as WRONG/MISSING, scoring 0 regardless of whether the underlying computation was correct.
task -- the task references a supplied dataset (with a README/data dictionary) and the script must load it to compute the reported numbers.read_csv, read_excel, load_dataset, DB query, path strings); if the only data construction is np.random.*, seed(...) loops, or literal lists, flag it. 3. Cross-check reported category names, group counts and sample sizes against the README/actual file — invented data usually shows suspiciously round or uniform group sizes. 4. Verify the answer is written to the exact requested output file/format, not an ad-hoc filename.f43f43c20432 · mined from da-code dacode-data-sa-061### Fabricated/synthetic data substituted for the provided dataset - **Applies when**: `task` -- the task references a supplied dataset (with a README/data dictionary) and the script must load it to compute the reported numbers. - **Pattern**: The script never reads any input file; instead it generates data with a random generator (or hard-codes values) that merely mimics the expected schema, then runs the requested statistics on that invented data and reports the results as the answer. - **Detection procedure**: 1. Read the task to confirm an external data source is expected. 2. Scan the script for any file-reading call (`read_csv`, `read_excel`, `load_dataset`, DB query, path strings); if the only data construction is `np.random.*`, `seed(...)` loops, or literal lists, flag it. 3. Cross-check reported category names, group counts and sample sizes against the README/actual file — invented data usually shows suspiciously round or uniform group sizes. 4. Verify the answer is written to the exact requested output file/format, not an ad-hoc filename. - **Discriminator**: Legitimate uses of random generation are auxiliary (bootstrap resampling, simulation baselines, train/test shuffling) applied *on top of* loaded real data; the violation is when the reported statistics derive entirely from generated values with no real data path. - **Consequence**: The p-values and conclusions are unrelated to the true data, so every value mismatches the expected result file (and the output file may be missing/misnamed), yielding 0 checks passed.
task -- a predictive task where the provided table contains identifier, categorical, and date/text columns alongside a handful of continuous measurements, and the scripts feed only the continuous measurements into the model.201ddc2f86f9 · mined from da-code dacode-ml-regression-004### Discarding high-signal columns and accepting a near-baseline validation score - **Applies when**: `task` -- a predictive task where the provided table contains identifier, categorical, and date/text columns alongside a handful of continuous measurements, and the scripts feed only the continuous measurements into the model. - **Pattern**: The attempt builds an ensemble of tree models on a hand-picked list of "numeric" columns, drops or never inspects the entity/name/date/genre/text columns (and never checks whether test rows can be matched back to rows of the reference table by identifier), and then ships predictions even though holdout error is essentially the same as predicting the target mean. - **Detection procedure**: 1. From the task/README and the exploration script, list all columns present in the reference table; from the modeling script, list the columns actually used as features. 2. Flag every dropped column that plausibly carries signal (identifiers, entity/creator names, dates or derived time features, categorical labels, text fields) and check whether the script attempted any encoding, aggregation, target-derived grouping, or an identifier join/overlap check between test rows and the reference table. 3. Read the reported validation metric and compare it to the trivial baseline computable from the printed target statistics (e.g., RMSE vs. the target's standard deviation, or the printed R²); also compare prediction spread (min/max/std) to the target's spread. 4. Confirm whether the attempt reported this gap as acceptable and stopped, without trying richer features or a deduplication/lookup path. - **Discriminator**: A real violation is when informative columns were available and untried *and* the achieved error is at or near the mean-baseline (R² ≈ 0, prediction std ≪ target std, heavily shrunk range). It is not a violation if the dropped columns are genuinely uninformative (free-form IDs with no reuse, constant columns) or if the agent tested encodings/joins and documented that they did not help while the model clearly beats the baseline. - **Consequence**: Predictions collapse toward the training mean, so any accuracy/correlation threshold the grader applies to the submitted file fails even though the file has the right name, shape, and column header.
task -- the task asks for an "appropriate" number of groups/hyperparameter and the script picks it by taking the single best value of one unsupervised score over a grid.argmax(silhouette) (or argmin(DB)) alone, and accepts the result even though the chosen solution is a poor/degenerate partition — e.g. clusters containing 1–3 points that are really outliers, an absolute score near 0.3 that is barely above the neighbouring k values, and no cross-check against the elbow curve, alternative metrics, robustness to seed, or outlier/skew handling (log-transform, winsorizing) of heavy-tailed features.915988065c7c · mined from da-code dacode-ml-cluster-013### Blind argmax of one internal metric for cluster count, with no sanity check on resulting cluster sizes - **Applies when**: `task` -- the task asks for an "appropriate" number of groups/hyperparameter and the script picks it by taking the single best value of one unsupervised score over a grid. - **Pattern**: The script sweeps k, selects `argmax(silhouette)` (or `argmin(DB)`) alone, and accepts the result even though the chosen solution is a poor/degenerate partition — e.g. clusters containing 1–3 points that are really outliers, an absolute score near 0.3 that is barely above the neighbouring k values, and no cross-check against the elbow curve, alternative metrics, robustness to seed, or outlier/skew handling (log-transform, winsorizing) of heavy-tailed features. - **Detection procedure**: 1. Read the task for how the number of groups is to be justified and what the output must represent. 2. In the script, check whether the selection uses more than one criterion / any agreement check, and whether a minimum-cluster-size or stability check gates the final choice. 3. In the reported output, inspect the cluster size distribution and the score table: flag singleton/near-singleton clusters, or a winning score that differs from the runner-up by a trivial margin. 4. Check whether skewed/outlier-heavy variables were transformed or outliers handled before distance-based clustering; if not, the "winning" k is likely just isolating outliers. - **Discriminator**: A fine attempt either shows the metric has a clear, well-separated optimum with balanced, interpretable groups, or explicitly justifies small clusters as genuine outlier groups after robustness checks; a violation is accepting a marginal metric win that yields degenerate clusters with no corroborating evidence or interpretation. - **Consequence**: The saved label column encodes an outlier-driven partition whose cluster count and assignments disagree with the expected grouping, so the file-level comparison (number of clusters / label agreement) fails even though the pipeline runs without error.
task -- the task asks for the entity (country, user, product…) that maximizes a distribution statistic (skewness, variance, mean, etc.) computed within each entity, so each entity's statistic depends on which rows/columns are collapsed into its sample.bc7e5fbfddb6 · mined from infiagent-dabench dabench-252### Per-group statistic computed over an unverified grouping/axis - **Applies when**: `task` -- the task asks for the entity (country, user, product…) that maximizes a distribution statistic (skewness, variance, mean, etc.) computed within each entity, so each entity's statistic depends on which rows/columns are collapsed into its sample. - **Pattern**: The attempt computes the statistic without an explicit, inspectable definition of the per-entity sample: it aggregates along the wrong axis (e.g., across entities instead of within one), collapses a filter it should have kept (or keeps rows it should have filtered), silently drops/keeps NaNs and non-numeric values, and then reports only the single argmax — no group sizes, no ranked list, no reproducible script — so an off-by-one-grouping error is invisible. - **Detection procedure**: 1. From the task, write down exactly what one entity's sample should be: which rows are selected, which column(s)/periods supply the values, and how many values are expected per entity. 2. In the script, locate the groupby/loop and the statistic call; confirm the grouping key is the entity, the value vector is the intended one (not transposed, not a row across entities), the required definition/flag is set (e.g., Fisher/bias options), and NaN/dtype handling is explicit. 3. Check the script prints diagnostics: per-entity sample counts, number of entities, and the top-5 ranked statistic values — not just the argmax. 4. If scripts are missing or the answer is a bare label with no printed intermediate table, treat the result as unverifiable and reject. - **Discriminator**: A real violation is when the reviewer cannot reconstruct, from the code, which values form each entity's sample (or can see the axis/filter/NaN handling is inconsistent with the task wording), or when entities with degenerate samples (n≤2, all-NaN, constant) are ranked alongside valid ones. It is fine if the grouping and value selection are explicit and matching, and diagnostics show plausible counts and a stable margin between the top candidates — even if the chosen library call differs stylistically. - **Consequence**: The argmax shifts to a neighboring entity whose sample was built differently, so the reported name mismatches ground truth and the grader records 0/1 with no trace to diagnose.
task -- the task asks the agent to create/derive something (a new column, a chosen feature, a selected model) and specifies an answer template whose placeholder names an identifier or label rather than the underlying data.69d360903ae5 · mined from infiagent-dabench dabench-741### Answer payload mismatch: dumping computed values instead of the requested identifier
- **Applies when**: `task` -- the task asks the agent to create/derive something (a new column, a chosen feature, a selected model) and specifies an answer template whose placeholder names an identifier or label rather than the underlying data.
- **Pattern**: The agent correctly performs the computation but fills the answer slot with the full vector of computed values (often truncated or with inconsistent rounding), rather than the single requested identifier/name described by the placeholder.
- **Detection procedure**: 1. Read the answer-format spec and identify exactly what the placeholder denotes (a name/label vs. a numeric list vs. a single statistic). 2. Check the placeholder's descriptive text and the task verb ("create a feature called…", "which variable…") to infer arity: one token or many. 3. Compare the agent's submitted payload to that arity/type; flag if it substitutes a long value list, an intermediate table, or extra commentary for a single identifier (or vice versa). 4. Verify any stated formatting constraints (rounding, ordering, delimiter, units) are applied to whatever is submitted.
- **Discriminator**: A real violation is a type/arity mismatch with the placeholder's stated meaning (e.g., hundreds of numbers where a column name is described, or truncated output for a required full list). It is *not* a violation if the task genuinely requests every per-row value and the agent supplies them completely, in order, with the specified precision.
- **Consequence**: The grader's exact-match on the expected identifier fails, scoring 0 even though the underlying computation was correct.task -- the task says the deliverable file must match the format of a supplied sample/example result file (or otherwise specifies an exact output schema).read_csv of the sample, printing its header/rows) and any comparison of the produced columns to it..round(2), self-chosen labels, self-chosen sort key).9662e0737271 · mined from da-code dacode-dm-csv-010### Failure to read and conform to the provided output-format template - **Applies when**: `task` -- the task says the deliverable file must match the format of a supplied sample/example result file (or otherwise specifies an exact output schema). - **Pattern**: The scripts never load or print the sample file; the agent invents column names, column order, row ordering, and rounding/precision from intuition, then saves the result and "verifies" only its own arithmetic, so the output can be numerically plausible yet schematically wrong. - **Detection procedure**: 1. Read the task and note every stated output constraint (template file, file name, column names, units, rounding, sort order). 2. Search the scripts for any read/inspection of the template file (e.g., `read_csv` of the sample, printing its header/rows) and any comparison of the produced columns to it. 3. If absent, check whether the header, column order, row order, and numeric formatting in the submitted answer are asserted anywhere or merely hard-coded by guess (e.g., an arbitrary `.round(2)`, self-chosen labels, self-chosen sort key). 4. Flag when the deliverable's schema/precision has no traceable source in the provided template. - **Discriminator**: Not a violation if the script actually loads the template (or explicitly reproduces its exact header/ordering/precision verified against it) and only then writes the output; it *is* a violation when the format is inferred from the agent's own phrasing of the task, even if the underlying aggregation logic is correct. - **Consequence**: The grader compares the file against the expected schema/values and marks it WRONG/MISSING because headers, ordering, or rounded values do not match, despite correct intermediate computations.
task -- the answer includes a key/label (date, ID, category, index) that must be located in the data, and the answer template shows a format or example whose granularity may be coarser than the data's actual granularity.a16ef15bdf8e · mined from infiagent-dabench dabench-572### Coarsening an identifier value to fit a format hint, losing required precision - **Applies when**: `task` -- the answer includes a key/label (date, ID, category, index) that must be located in the data, and the answer template shows a format or example whose granularity may be coarser than the data's actual granularity. - **Pattern**: The attempt correctly locates the record but then truncates, rounds, or re-renders the identifier (e.g., drops components of a composite/hierarchical key) to match a literal format hint, reporting a value that no longer uniquely identifies the record it found — even though downstream computations were done on the full-precision value. - **Detection procedure**: 1. In the task, note the granularity of the identifier the data actually contains and whether any other constraint (e.g., "previous record", per-record lookup) implies the full-precision key is needed. 2. In the scripts/output, find the raw identifier of the located record and compare it to the string finally emitted; check for truncation, reformatting, or aggregation of the key. 3. Check internal consistency: does the emitted identifier, if fed back into the data, select exactly one record — and the same record used for the dependent calculation? 4. If the emitted key is coarser than the located record (maps to many rows) while the dependent value came from a single row, flag it. - **Discriminator**: A real violation is when the reported key is ambiguous or lower-precision than the record actually used, so it cannot be verified against the data; it is *not* a violation when the data genuinely has that coarse granularity (one row per that key), or when the task explicitly asks for an aggregate over the coarser grouping and the dependent computation was done at that same level. - **Consequence**: The identifier check fails on exact-string comparison against the ground-truth full-precision key, so the submission is marked partially correct at best (dependent numeric value may pass while the key fails).
task -- the task asks for a statistic from a model built on specific named quantities from a provided dataset, and the script instead pulls data from an external source or swaps in a "close enough" stand-in series.b74b966f23d1 · mined from da-code dacode-data-sa-043### Substituting proxy/external data for the variables the task actually names - **Applies when**: `task` -- the task asks for a statistic from a model built on specific named quantities from a provided dataset, and the script instead pulls data from an external source or swaps in a "close enough" stand-in series. - **Pattern**: The agent cannot locate (or does not look for) the required variables in the supplied data, so it downloads unrelated series or redefines a variable as a proxy (e.g., a rate instead of the named quantity, a different index as the "portfolio"), sometimes with a different frequency, aggregation, or date coverage, then reports the statistic from that substitute model as the answer while acknowledging the substitution in the write-up. - **Detection procedure**: 1. From the task text, list the exact dependent and independent quantities, the aggregation (e.g., per-period min/mean) and the date range required. 2. In the scripts, trace where each modeling column comes from: is it derived from the provided dataset files, or fetched/invented externally? 3. Check whether the aggregation and filtering literally match the task wording (correct reduction function, correct period grouping, correct start/end). 4. Read the answer/comments for hedging language such as "proxy", "alternative interpretation", "actual variable not available", or a fallback second model — this signals the reported number is not from the requested regression. - **Discriminator**: A real violation is when a modeling variable is not derivable from the provided data and no equivalence is demonstrated, or the required aggregation is replaced by a different one. It is fine if the agent uses a documented column alias/rename within the provided dataset, or reconstructs the exact quantity (same units, frequency, and reduction) from raw provided columns. - **Consequence**: The saved statistic comes from a different model on different inputs, so the value in the output file fails the exact/tolerance comparison against ground truth even though the code runs cleanly.
task -- the task states an asymmetric cost/priority (e.g., missed positives are far more expensive) or a specific evaluation metric for a rare-class classification, and the scripts must emit hard 0/1 labels.predict() at the implicit 0.5 probability cutoff — never optimizing recall / the stated cost function, never tuning or even reporting the threshold, and never checking the resulting positive rate against the training base rate.predict() and unrelated scores?161f08488349 · mined from da-code dacode-ml-binary-013### Model selection and decision threshold not aligned with the task's stated objective on imbalanced classes - **Applies when**: `task` -- the task states an asymmetric cost/priority (e.g., missed positives are far more expensive) or a specific evaluation metric for a rare-class classification, and the scripts must emit hard 0/1 labels. - **Pattern**: The attempt trains several off-the-shelf classifiers, compares them with generic metrics (accuracy, ROC-AUC, F1) on one arbitrary hold-out split, then calls `predict()` at the implicit 0.5 probability cutoff — never optimizing recall / the stated cost function, never tuning or even reporting the threshold, and never checking the resulting positive rate against the training base rate. - **Detection procedure**: 1. Read the task statement and write down the metric or cost trade-off it actually asks to optimize, plus the class balance seen in the training data. 2. In the scripts, check how the final model and cutoff are chosen: is selection driven by that metric (e.g., threshold sweep / cost curve / recall-oriented CV), or by default `predict()` and unrelated scores? 3. Compare the predicted positive count/rate in the answer with the training-set positive rate and with the recall implied on validation; flag if positives are predicted at or below base rate while the task penalizes false negatives. 4. Check the selection evidence: single random split vs. repeated/stratified CV, and whether the chosen model is justified by numbers actually printed. - **Discriminator**: A genuine violation uses the default cutoff and off-target metrics with no sensitivity analysis; an acceptable attempt either explicitly tunes the threshold/class weights against the stated cost or metric and shows the trade-off, or documents that the default cutoff is already optimal under that metric. - **Consequence**: The saved label column has too few positives (low recall), so the graded score under the cost/recall-based check falls below the required threshold and the prediction file is marked wrong even though the file format is fine.
task -- the question asks for a statistic or detection over "all" units (all countries, all users, all records) and the data directory contains multiple files/partitions or the script filters/loads a single subset.5f28c3b3b22a · mined from infiagent-dabench dabench-254### Analyzing only one partition of the data when the task asks for the whole population
- **Applies when**: `task` -- the question asks for a statistic or detection over "all" units (all countries, all users, all records) and the data directory contains multiple files/partitions or the script filters/loads a single subset.
- **Pattern**: The script hard-codes a single input file (or one group/region/segment) and computes quartiles, thresholds, or aggregates on that partial sample, then reports the result as if it covered the full population — the reference quantities (Q1/Q3, means, ranks) are therefore derived from the wrong subset, even if some of the returned items happen to coincide with the expected ones.
- **Detection procedure**:
1. Read the task statement and note the stated scope of the population ("for all …") and any required grouping or filtering.
2. Read the script's data-loading lines: check whether it enumerates/concatenates every relevant file or subset, or loads exactly one and never verifies that this covers the requested scope.
3. Compare the row/unit count printed or implied by the script against the count implied by the task scope (e.g., number of files in the directory, total known units); a much smaller count signals a partial load.
4. Check that the threshold/statistic itself (not just the final list) was computed on the full-scope data, and that the answer's labels and format match the requested output exactly.
- **Discriminator**: A real violation is loading/filtering a subset when the task's scope is broader and no justification or coverage check exists. It is fine if the task explicitly restricts scope to that subset, or if the script demonstrates (by listing directory contents or comparing counts) that the single file *is* the complete population.
- **Consequence**: Quartiles/thresholds computed on a partial sample yield a different outlier/selection set than the ground truth built from the full population, so the graded answer list mismatches and scores 0.task -- the deliverable is a prediction file with a specified column name whose values are the original class labels (or a specified numeric format) taken from a categorical target in the training data.to_csv: check whether an inverse transform back to the original strings is applied, whether index=False is used, and whether the header equals the requested name.value_counts() keys character-for-character identical to the training label values.index=False and prints a verification of shape, column name, and label set — even if the prose summary paraphrases class names loosely.6039bfdfbd4a · mined from da-code dacode-ml-binary-009### Output label/schema fidelity not verified against the training target's exact categories - **Applies when**: `task` -- the deliverable is a prediction file with a specified column name whose values are the original class labels (or a specified numeric format) taken from a categorical target in the training data. - **Pattern**: The attempt trains a model, encodes the target internally, then writes predictions using re-derived or re-formatted labels (title-cased, renamed, 0/1 codes, extra index column, added ID column, wrong row count/order) without ever asserting that the written values and file schema exactly reproduce the label strings and column layout requested; the report even describes classes with capitalization/wording that differs from the raw data. - **Detection procedure**: 1. From the task/README, note the exact required file name, column name(s), and the exact set of label strings as they appear in the training target column. 2. In the scripts, trace the target from load → encoding → inverse mapping → `to_csv`: check whether an inverse transform back to the original strings is applied, whether `index=False` is used, and whether the header equals the requested name. 3. Check for any explicit sanity check in the script or answer: row count equals the test set size, single expected column, and `value_counts()` keys character-for-character identical to the training label values. 4. Read the answer's class-distribution/summary text: if the class names or column layout it reports differ in spelling, case, or count from the raw data, treat the output as unverified. - **Discriminator**: A real violation is the absence of an inverse mapping/format assertion, or evidence of relabeled/re-cased/coded values, extra index or ID columns, or a row count not matching the test rows; a look-alike that is fine is a script that writes original label strings (or the explicitly requested encoding) with `index=False` and prints a verification of shape, column name, and label set — even if the prose summary paraphrases class names loosely. - **Consequence**: The grader cannot match the expected column values and marks the required output file WRONG/MISSING, so the attempt scores 0 regardless of how good the underlying model is.
task -- the task asks for a chart/table summarizing a substantive quantity (e.g., "performance", "top N"), and a settings/config file supplies cosmetic details plus axis/category labels.590ec2f02edf · mined from da-code dacode-plot-bar-006### Invented metric definition + entity list back-filled from the plot config instead of derived from the data - **Applies when**: `task` -- the task asks for a chart/table summarizing a substantive quantity (e.g., "performance", "top N"), and a settings/config file supplies cosmetic details plus axis/category labels. - **Pattern**: The script treats the config's label list as the source of truth for *which* entities to include and invents an ad-hoc, unjustified formula for the plotted value (arbitrary weights, mixing unrelated counts), rather than computing the substantive quantity from the data and letting the ranking/selection fall out of it. Filters (e.g., the "specified period") are also guessed from a title string. No check is made that the computed ordering/selection reproduces the config's labels, and required auxiliary output artifacts implied by the task/settings are never written. - **Detection procedure**: 1. Read the task and the config: list every constraint (which quantity is requested, which subset/period, ordering, required output files/format). 2. In the scripts, locate the formula for the plotted values and ask whether each term and weight is traceable to the task, README, or config — or was chosen by the agent. 3. Check whether the set/order of plotted categories is computed from the data (e.g., sorted by the metric and truncated) or merely intersected with the config's label list; if the latter, check whether the agent verified that its metric reproduces those labels in that order. 4. Check the answer/output for all requested artifacts and for a sanity check on the plotted numbers (plausible ranges, monotonic ordering, counts matching N). - **Discriminator**: Fine if the metric is a standard/derivable definition for the requested quantity and the config labels are used only as a consistency check that the independently computed ranking matches; a violation if the metric contains agent-chosen weights/terms, the plotted bars are visibly non-monotonic with respect to a "top/best" label ordering, or the category list could not have been reproduced from the data by the script's own logic. - **Consequence**: The chart's bar values (and any saved numeric/JSON companion outputs) differ from the reference, so all value/plot checks fail even though the figure looks well-formed, and missing required output files fail outright.
task -- the task asks for a count of rows meeting a statistical threshold (z-score, IQR, quantile, etc.) computed on a single column, and the script relies on a library helper plus a boolean mask to both count and filter.(|x - mean| > kstd).sum(), or comparing min/max to mean ± kstd) or checking the column's dtype, NaN count, and units.max and min lie inside mean ± k·std, the count must be 0; if the count is non-zero, the script should print actual flagged values whose distance from the mean exceeds the bound.nan_policy='omit' shrinking the array, ddof differences, positional mask on a non-default index) could shift which rows are flagged.04b18e867611 · mined from infiagent-dabench dabench-361### Outlier/threshold counts reported without an independent cross-check of the flagging logic - **Applies when**: `task` -- the task asks for a count of rows meeting a statistical threshold (z-score, IQR, quantile, etc.) computed on a single column, and the script relies on a library helper plus a boolean mask to both count and filter. - **Pattern**: The script computes the statistic with one code path (e.g. a library function with special NaN/ddof handling), builds a positional boolean/index mask, applies it to the dataframe, and reports the resulting count without ever verifying it against a simple, independent recomputation (`(|x - mean| > k*std).sum()`, or comparing `min`/`max` to `mean ± k*std`) or checking the column's dtype, NaN count, and units. - **Detection procedure**: 1. From the task, note the exact rule and threshold and the exact quantity requested (count of flagged rows). 2. In the script, check whether the column is inspected first (dtype, non-null count, min/max/describe) and whether any non-numeric/missing values would be silently dropped or coerced, and whether the mask indices (positional vs. label) match the dataframe index used for filtering. 3. Check whether the flag count is recomputed a second, independent way (manual mean/std formula, or comparing extreme values to the mean ± k·std bounds) and whether the two agree. 4. Compare the reported count to the reported summary stats: if `max` and `min` lie inside mean ± k·std, the count must be 0; if the count is non-zero, the script should print actual flagged values whose distance from the mean exceeds the bound. - **Discriminator**: A fine attempt prints the column's summary statistics and the flagged values, and the flagged values are visibly beyond the stated threshold from the mean; a violation reports a count that is never reconciled with the printed distribution, or whose mask alignment/NaN handling (e.g. `nan_policy='omit'` shrinking the array, ddof differences, positional mask on a non-default index) could shift which rows are flagged. - **Consequence**: The reported integer count (and the size of the filtered dataframe) is off — often a large spurious count where the true answer is zero or vice versa — so the graded numeric answer fails outright.
task -- the prompt points to an auxiliary document/config (readme, notes, mapping file, data dictionary) that defines how values must be transformed, named, filtered, or reported.open, read_csv, read_text, path string); if absent, the transformation rules are unverified assumptions.a1b122fce003 · mined from da-code dacode-di-text-004### Substituting assumed conventions for an explicitly referenced specification file - **Applies when**: `task` -- the prompt points to an auxiliary document/config (readme, notes, mapping file, data dictionary) that defines how values must be transformed, named, filtered, or reported. - **Pattern**: The scripts never open or parse the referenced file; instead the agent hardcodes a "standard"/guessed mapping or definition from prior knowledge, then computes and reports statistics using those invented labels (and often skips writing the specified output artifact). - **Detection procedure**: 1. Read the task and list every external resource it names and every constraint that resource is said to carry (label names, categories, filters, output file). 2. Grep the scripts for any read of that resource (`open`, `read_csv`, `read_text`, path string); if absent, the transformation rules are unverified assumptions. 3. Compare the literal strings/values used in the script's hardcoded dictionary or constants against what the task says should come from the file — no in-script evidence of agreement is a violation. 4. Check that the answer is emitted in the required form/artifact (e.g., written to the named result file with the named keys), not only printed to stdout. - **Discriminator**: Fine if the script actually loads the referenced file (or quotes its verified contents after inspecting it) and derives the mapping from it; a violation if the mapping's provenance is only a comment like "these are the standard codes" or general domain knowledge. - **Consequence**: The numeric ratio may be right while the category label differs from the expected vocabulary (or the required output file is missing), so exact-match grading on the result file fails.
task -- the task supplies a sample/template output file (or explicit format spec) and the scripts write a predictions file to disk.a1442e208ec3 · mined from da-code dacode-ml-competition-003### Submission schema invented instead of copied from the provided template - **Applies when**: `task` -- the task supplies a sample/template output file (or explicit format spec) and the scripts write a predictions file to disk. - **Pattern**: The scripts never read the template file; column names, column order, ID column casing, and row ordering are hardcoded from the agent's guess (e.g. self-chosen probability column labels), and the "verification" step only re-checks the file against the agent's own assumptions rather than against the template. - **Detection procedure**: 1. In the task/README, note that a template output file is provided and that the output must match it. 2. Grep the scripts for a read of that template file; if absent, inspect how the output DataFrame's columns/order are constructed. 3. Compare the agent's produced header and row order to the template's header and ID order (names, count, case, spelling, ordering of rows). 4. Check whether the verification script asserts equality with the template header/ID sequence, or merely checks internal consistency (probabilities summing to 1, ID set membership). - **Discriminator**: Not a violation if the hardcoded header/order provably equals the template (agent printed the template header and matched it exactly, including ID ordering); it is a violation when the header labels or ordering are guessed from prose/memory and never diffed against the actual file. - **Consequence**: The grader cannot parse/align the submission and marks the file WRONG/MISSING regardless of predictive quality, yielding a 0 score.
task -- The task names specific input/output files (e.g., a training file with a target column and a test file to predict on) and the scripts must produce predictions keyed to those exact rows.0b8517bdd1a1 · mined from da-code dacode-ml-multi-003### Fabricating labels instead of using the provided supervised target and evaluation split - **Applies when**: `task` -- The task names specific input/output files (e.g., a training file with a target column and a test file to predict on) and the scripts must produce predictions keyed to those exact rows. - **Pattern**: The attempt declares the ground-truth target "unavailable", builds heuristic pseudo-labels from unrelated auxiliary tables, trains/validates on those self-generated labels (yielding an implausibly high self-consistent CV score), and writes predictions for a row set it invented rather than the rows of the specified test file. - **Detection procedure**: 1) From the task statement, list the required input file(s), the required prediction file, its required column(s), and the required row set/ordering. 2) Grep the scripts for reads of those files and for the target column; note whether the label used in training comes from the data or from a hand-written rule function. 3) Compare the row count/keys of the produced output against the specified test file's row count/keys, and check the reported column names match exactly what was asked. 4) Check whether the reported validation score is measured against real labels or against the same rule that generated them. - **Discriminator**: A real violation is when the supervised target exists in the provided data (or the test row set is explicitly given) but the script never reads it and instead invents labels/rows; it is acceptable if the task is genuinely unsupervised, or if heuristics are used only as extra features alongside the true labels and the output still matches the specified test keys and columns. - **Consequence**: The output file has the wrong number of rows, wrong keys, and/or extra/misnamed columns, and its label distribution bears no relation to the true target, so the grader marks the result file wrong/missing despite a near-perfect self-reported CV accuracy.
task -- the task supplies a template/example output file that fixes the row and column labels of a computed table (e.g., group keys by period index), and the script builds its own table then renames or shifts the labels to make them look like the template.f25e75f5bcdb · mined from da-code dacode-dm-csv-044### Relabeling output axes to "match" a template instead of verifying value-to-cell alignment
- **Applies when**: `task` -- the task supplies a template/example output file that fixes the row and column labels of a computed table (e.g., group keys by period index), and the script builds its own table then renames or shifts the labels to make them look like the template.
- **Pattern**: The script computes a grouped statistic with its own natural index (e.g., offsets starting at 0), then cosmetically renames columns/rows (adding an offset, reformatting dates) so the header string-matches the template, without establishing which underlying group each template cell actually represents. Verification scripts then only compare shapes/labels, or spot-check cells while openly speculating ("column 4 might mean offset 3"), and the ambiguity is never resolved before writing the file.
- **Detection procedure**:
1. Read the task/template: note the exact expected row labels, column labels, ordering, and rounding/format of the deliverable.
2. In the scripts, find every place where labels are renamed, offset, reindexed, or reformatted after the computation, and ask whether the mapping from computed group → template cell is derived from a stated definition or merely assumed to make headers line up.
3. Check whether any script independently confirms the mapping (e.g., recomputing one cell from raw rows filtered by the definition the template implies, and matching a non-empty template value or a documented convention) rather than comparing an empty template's headers.
4. Read the answer: look for statements that reveal unresolved ambiguity about what a column/row means, or claims of correctness backed only by "shape and headers match".
- **Discriminator**: A fine attempt derives the label convention from the task/template (or from filled example values) and verifies at least one cell end-to-end against raw data under that convention; a violation performs an arbitrary shift/rename purely so the header text matches, and its own debug output shows two competing interpretations left unsettled.
- **Consequence**: Every value lands one position off (or the whole grid is shifted/transposed), so the file has the right shape and labels but wrong contents, and the grader marks the expected output file WRONG.task -- the task states an explicit scoring metric (e.g., a rank/agreement/ordinal or cost-weighted metric) and the scripts must choose a model, tune it, or decide when the answer is good enough.db6067bf7360 · mined from da-code dacode-ml-competition-006### Optimizing/validating with a metric other than the one the task specifies - **Applies when**: `task` -- the task states an explicit scoring metric (e.g., a rank/agreement/ordinal or cost-weighted metric) and the scripts must choose a model, tune it, or decide when the answer is good enough. - **Pattern**: The scripts never implement or compute the stated metric; they select the "best" model and ensemble weights using a generic proxy (accuracy, weighted F1, plain log-loss) and treat the target as unordered classes, so the reported/validated numbers say nothing about the actual grading score. Often compounded by a contaminated check (the final ensemble is fit on all training rows and then "evaluated" on a subset of those same rows). - **Detection procedure**: 1. Read the task and write down the exact evaluation metric and any structure it implies (ordering of labels, weighting, thresholding, rounding). 2. Grep the scripts for that metric's computation (or an equivalent implementation) and for the objects used in model comparison/selection; note which score drives the choice of final model. 3. Check that the score used for selection is computed on data held out from every fitting step (scaler, feature transforms, base models, ensemble weights) — not on rows the final model was refit on. 4. Check the answer/prediction distribution against what the metric rewards (e.g., under an ordinal-agreement metric, whether extreme classes and label ordering are handled at all rather than collapsed to majority classes). - **Discriminator**: A real violation is when no run-time estimate of the stated metric exists anywhere, or the metric-relevant structure is ignored, so the agent cannot tell a good submission from a bad one. It is *not* a violation if the agent computes the stated metric on a clean holdout/CV and additionally reports proxies, or if it can show the proxy is a monotone stand-in for the stated metric on that holdout. - **Consequence**: The submission looks well-formed and plausible, but the graded score (the stated metric) is far below what a metric-aware baseline achieves — predictions cluster on majority labels and the attempt fails the correctness check despite high reported accuracy/F1.
task -- the task asks to split rows into "null" vs "non-null" (or otherwise filtered) groups on one column and compare an aggregate of another column between the groups.isna() after a loader that coerced sentinels like "NA", "None", "", "-" to real values or vice versa, dropna() on the whole frame, filtering on a different column, or reading only a subset/sample of the file) and/or lets the aggregated column be non-numeric so rows are silently dropped or coerced — then reports group means without checking that the two groups reconstruct the full dataset.na_values=/converters, whole-frame dropna(), astype/to_numeric(errors='coerce'), string-vs-NaN handling, and any row subsetting or sampling; confirm the aggregate is computed on the intended column for each mask.len(group_A) + len(group_B) == len(df), per-group non-null counts of the aggregated column, and total row count matching the raw file; if absent, the attempt is unverified.91e8c2fe7699 · mined from infiagent-dabench dabench-297### Missingness mask defined inconsistently with the raw data, without a partition sanity check - **Applies when**: `task` -- the task asks to split rows into "null" vs "non-null" (or otherwise filtered) groups on one column and compare an aggregate of another column between the groups. - **Pattern**: The attempt defines the group mask with an ad-hoc or lossy rule (e.g., `isna()` after a loader that coerced sentinels like `"NA"`, `"None"`, `""`, `"-"` to real values or vice versa, `dropna()` on the whole frame, filtering on a different column, or reading only a subset/sample of the file) and/or lets the aggregated column be non-numeric so rows are silently dropped or coerced — then reports group means without checking that the two groups reconstruct the full dataset. - **Detection procedure**: 1. From the task, note the exact grouping rule (null vs not-null on the specified column) and the column to be aggregated; note that no other filtering was authorized. 2. In the scripts, locate the load step and the mask: check for `na_values=`/converters, whole-frame `dropna()`, `astype`/`to_numeric(errors='coerce')`, string-vs-NaN handling, and any row subsetting or sampling; confirm the aggregate is computed on the intended column for each mask. 3. Require an explicit printed sanity check: `len(group_A) + len(group_B) == len(df)`, per-group non-null counts of the aggregated column, and total row count matching the raw file; if absent, the attempt is unverified. 4. Check the reported means are consistent with those counts (e.g., recompute the pooled mean from group means and sizes and compare to the overall column mean). - **Discriminator**: A real violation is when rows are added/lost or reclassified relative to the raw column's true missingness (counts don't partition the file, or sentinel strings are treated inconsistently). It is fine if the aggregated column itself has NaNs excluded by the aggregation *within* each group, provided the grouping mask still partitions all rows and this is stated/verified. - **Consequence**: Both group means shift by a few percent — plausible-looking numbers that fail exact-value checks (the t-test p-value can still look tiny, masking the error), so the answer is graded wrong.
task -- the task (or a referenced guidance/spec file) asks for computed results plus saved outputs, and the script's only persisted product is one figure or console prints while the numeric results live in the chat answer.savefig, to_csv, np.save, json.dump, to_json, etc.) and record the exact paths produced.print/the chat answer.767c30e1213e · mined from da-code dacode-plot-pie-005### Missing or unsaved required output artifacts - **Applies when**: `task` -- the task (or a referenced guidance/spec file) asks for computed results plus saved outputs, and the script's only persisted product is one figure or console prints while the numeric results live in the chat answer. - **Pattern**: The agent computes the requested quantities but never serializes them to the expected files/paths/formats (e.g., no saved array/table of the tallied values, no structured record of the plotted data), and instead reports them only as prose; it also does not re-read the spec file to enumerate the full list of deliverables. - **Detection procedure**: 1. From the task statement and any referenced guidance/README file, list every deliverable: each output file name, its format, its location, and the values it must contain. 2. Grep the script for every write/save call (`savefig`, `to_csv`, `np.save`, `json.dump`, `to_json`, etc.) and record the exact paths produced. 3. Compare the two lists; flag if any required deliverable has no corresponding write, is written to a different directory/name/extension, or if a required numeric result exists only in `print`/the chat answer. 4. Check the saved figure/data actually encodes the requested numbers (counts, labels, ordering) rather than a stylistic approximation. - **Discriminator**: A real violation is a deliverable that is never written, or written under a name/format the spec did not ask for; it is *not* a violation if the file exists with the right content and only cosmetic details (title text, dpi, autopct) differ, or if the harness itself relocates correctly named outputs. - **Consequence**: The grader checks each expected artifact independently and marks every missing/mismatched file WRONG/MISSING, so the attempt scores 0 even when the reported numbers are plausible.
task -- the task supplies an example output file (or explicit column/row spec) and asks that results be saved in that exact format.0b75529b4da3 · mined from da-code dacode-dm-csv-015### Output schema not copied from the provided sample/format template - **Applies when**: `task` -- the task supplies an example output file (or explicit column/row spec) and asks that results be saved in that exact format. - **Pattern**: The attempt invents its own header names, column count, or row labels (often adding stray/empty columns or renaming fields) instead of reading the sample file and reproducing its schema, so the content may be right while the file fails automated comparison; frequently the values are also hard-coded/recalled rather than produced by an aggregation over the data, with no script retained to show how they were derived. - **Detection procedure**: 1. In the task statement, locate the referenced sample/format artifact and note its exact headers, column order, and row-label wording. 2. In the scripts, check that the sample file is actually loaded or its header literally reproduced, and that the reported values come from a computed aggregation over the full dataset (grouping/summing per entity) rather than typed-in constants. 3. Diff the submitted file's header and row labels against the sample: same number of columns, same names/spelling/case, no extra or empty trailing fields, expected number of rows. 4. Sanity-check one value by re-deriving it from the data (e.g., the max of the aggregated quantity) to confirm the ranking logic, not memory, produced it. - **Discriminator**: A real violation is a structural/label deviation from the given template (extra column, renamed header, invented row wording) or values with no code path producing them; a look-alike that is fine is a file whose schema matches the template exactly and differs only in benign ways the spec leaves open (row order when unspecified, quoting/whitespace). - **Consequence**: The grader's file/field comparison marks the expected output file WRONG/MISSING and scores 0, even if the underlying leaders were conceptually right.
task -- the script fits a model or computes a statistic on numeric columns from a source that is flagged (by filename, docs, or dtype) as containing errors, and the required preprocessing is only stated at a high level (e.g., "impute missing with the mean").8008c46c5afc · mined from infiagent-dabench dabench-432### Unvalidated raw values in a knowingly "dirty" dataset before imputation/modeling - **Applies when**: `task` -- the script fits a model or computes a statistic on numeric columns from a source that is flagged (by filename, docs, or dtype) as containing errors, and the required preprocessing is only stated at a high level (e.g., "impute missing with the mean"). - **Pattern**: The attempt takes the columns at face value: it prints dtypes/describe output but never checks whether values are actually parseable numerics within physically plausible ranges (non-numeric strings, sentinel codes, sign errors, unit/magnitude mismatches, duplicated or blank rows). Corrupt entries silently either become NaN-and-mean-imputed or remain as extreme outliers, so the fitted model and the reported error metric are driven by garbage rows. - **Detection procedure**: 1. Read the task for signals that the input is noisy/erroneous, and note which columns feed the model and which is the target. 2. In the script, look for any explicit validation or cleaning step on those columns: coercion to numeric with inspection of what failed, range/plausibility checks, outlier or duplicate handling, verification that the imputed mean is computed on clean values. 3. Check whether the script sanity-checks the final metric against the target's own scale (e.g., compare RMSE to the target's standard deviation or observed range) and reacts if it is implausibly large. 4. If neither cleaning nor a scale sanity-check exists, and the reported metric implies typical errors comparable to or larger than the target's spread, flag the attempt. - **Discriminator**: A real violation is when no inspection of value validity occurred at all, or the reported error is out of line with the target's natural variability and is accepted without investigation. It is fine if the script explicitly verified the columns are clean (ranges, dtype coercion with zero failures) and the metric is consistent with the target's scale — even if no rows were removed. - **Consequence**: The regression is fit on contaminated values, inflating the reported error by roughly an order of magnitude relative to the reference value, so the single numeric answer fails the grader's exact-match check.
task -- the requested quantity depends on row order (differences, lags, percent changes, cumulative sums, rolling stats) and the script explicitly sorts, reverses, or otherwise re-indexes the table before computing it.sort_values on a datetime/index) or prints head/tail of that key and states the observed direction; a violation if the direction is only assumed from a glance or a comment, or if raw row position is used as the order without any key check. Merely reversing data is not itself wrong — the absence of verification is.992d72ba6673 · mined from infiagent-dabench dabench-75### Unverified reordering of rows before a sequence-dependent computation - **Applies when**: `task` -- the requested quantity depends on row order (differences, lags, percent changes, cumulative sums, rolling stats) and the script explicitly sorts, reverses, or otherwise re-indexes the table before computing it. - **Pattern**: The agent asserts an ordering (e.g., "the file is newest-first, so reverse it") in a comment and applies the transformation without ever parsing the ordering key, printing the first/last key values, or checking monotonicity — so the sequence may be flipped relative to the intended one, silently inverting signs and shifting the result. - **Detection procedure**: 1. Read the task and note that the requested statistic is defined relative to a "previous"/"next" element, i.e. it is order-sensitive and sign-sensitive. 2. In the scripts, locate any reversal/sort/reset_index step and check whether it is justified by evidence produced in the same run: is the ordering column converted to a proper dtype (e.g., datetime) and is its monotonic direction verified or sorted on explicitly? 3. Check whether any post-hoc sanity check on the ordering-sensitive output exists (e.g., comparing the sign/magnitude of the statistic computed both as-is and reordered, or spot-checking one difference against two identified rows by their key values). 4. Compare the reported answer to what would result under the opposite ordering; if the two differ mainly by sign, the answer is unvalidated. - **Discriminator**: Fine if the script sorts by the parsed ordering key (`sort_values` on a datetime/index) or prints head/tail of that key and states the observed direction; a violation if the direction is only assumed from a glance or a comment, or if raw row position is used as the order without any key check. Merely reversing data is not itself wrong — the absence of verification is. - **Consequence**: The lag is taken in the wrong direction, so the mean flips sign (and dispersion shifts slightly), and the graded values fail even though the formula and rounding look correct.
task -- The task points to an auxiliary instructions file (e.g., a tips/spec/README note) that defines the required test procedure and the exact fields/values of a result file the agent must write.e6cd929607a6 · mined from da-code dacode-data-sa-004### Output schema invented instead of read from the task's referenced specification - **Applies when**: `task` -- The task points to an auxiliary instructions file (e.g., a tips/spec/README note) that defines the required test procedure and the exact fields/values of a result file the agent must write. - **Pattern**: The script never opens or echoes the referenced spec file; the agent guesses the column names, ordering, value wording (e.g., decision phrasing, test-type label), rounding, and number of fields, then declares the output "in the required format" without any comparison to the stated format. - **Detection procedure**: 1. Read the task and list every referenced instruction artifact and every explicitly constrained property of the deliverable (file name, one row vs. many, field set, allowed value strings, precision). 2. Scan the scripts for any read/print of that artifact and for a literal mapping from its stated fields to the written columns; note whether the header/values in the code are copied from the spec or invented. 3. Check the agent's answer for evidence it quoted the spec's required fields/values; absence of such a quote, or extra/renamed fields (e.g., an added statistic column, custom decision wording), is the flag. 4. Also verify the chosen test and its preconditions match what the spec prescribes rather than the agent's own judgement (including how degenerate groups, e.g., a group with almost no observations, are handled). - **Discriminator**: Fine if the script demonstrably reads or verbatim reproduces the spec's field names/values and any deviation is justified; a violation is when the format is asserted from memory or from the prose of the main prompt only, with no trace of the spec content anywhere in scripts or output. - **Consequence**: The result file exists and looks plausible, but exact-match/field-wise grading of the deliverable fails (wrong or extra columns, non-matching decision strings or precision), scoring 0 despite a statistically defensible analysis.
task -- The task asks to "normalize"/"scale" columns and then report summary statistics (e.g., means) of those transformed columns.934d706c6c1c · mined from infiagent-dabench dabench-28### Normalization method chosen so the reported statistic is trivially degenerate - **Applies when**: `task` -- The task asks to "normalize"/"scale" columns and then report summary statistics (e.g., means) of those transformed columns. - **Pattern**: The agent applies z-score standardization (subtract mean, divide by std), which forces every reported mean to 0 (or -0.0), instead of a scaling that preserves informative means (e.g., min-max to [0,1]); it then reports these degenerate zeros without questioning that the task would not ask for values that are known in advance. - **Detection procedure**: 1. Read the task to see which columns are to be scaled and which statistic must be reported afterwards. 2. Inspect the script's scaling step: check whether the transform makes the requested statistic constant by construction (standardization → mean 0; centering → mean 0). 3. Check the answer: if the reported values for all scaled columns are 0.0/-0.0 (or otherwise identical trivial constants) while unscaled columns carry real information, flag it. 4. Confirm the answer format expectation (four-decimal distinct floats per column) is inconsistent with an all-zero result, and require the alternative scaling (min-max) or an explicit justification. - **Discriminator**: A real violation is when the requested statistic is mathematically fixed by the chosen transform, rendering the report uninformative; it is fine if standardization is explicitly named in the task, or if the reported statistic (e.g., std, min, max, or a downstream model score) is still informative under standardization. - **Consequence**: All scaled columns' reported means are 0.0 instead of the expected in-range values, so every check on those columns fails while unscaled/binary columns coincidentally pass.
task -- the task asks for a single summary statistic or test result computed over two or more columns of a supplied table, and the scripts (or lack of them) do not show how rows with missing/non-numeric entries were handled or how many rows entered the computation.dropna() without subset=, boolean filters, head/nrows, dtype coercion, deduplication) and check whether any of them can remove rows that have valid values in both target columns.0.0 for a p-value that should be given to four decimals).c07951b6377d · mined from infiagent-dabench dabench-300### Unverified row set / missing-value handling for a whole-dataset statistic - **Applies when**: `task` -- the task asks for a single summary statistic or test result computed over two or more columns of a supplied table, and the scripts (or lack of them) do not show how rows with missing/non-numeric entries were handled or how many rows entered the computation. - **Pattern**: The attempt computes the statistic after an implicit or convenience row reduction — e.g. dropping all rows with any NA anywhere in the table instead of only in the two relevant columns, coercing/filtering out non-numeric values, subsetting to a preview/sample of the file, or reading with a wrong delimiter/header so some rows are lost — and reports the resulting number without stating the sample size or checking it against the file's row count. It also reports derived values with less precision than requested (e.g. a truncated p-value) instead of the stated rounding. - **Detection procedure**: 1. From the task, note that the statistic is defined over the full table on the pair of named columns, and note the required output precision/format for each reported quantity. 2. In the scripts, locate the load step and every row-removing operation (`dropna()` without `subset=`, boolean filters, `head`/`nrows`, dtype coercion, deduplication) and check whether any of them can remove rows that have valid values in both target columns. 3. Check that the script prints the number of observations actually used (and non-null counts per column) and that this is compared with the raw file row count; if no script or no such print exists, treat the number as unverified. 4. Compare the reported figures against the required rounding/format (e.g. two vs four decimals, no bare `0.0` for a p-value that should be given to four decimals). - **Discriminator**: A real violation is any row exclusion that is not required by the task and not justified/reported (or an unstated sample size), or a reported value whose precision differs from the spec. It is *not* a violation if the script restricts to pairwise-complete cases on exactly the two target columns, prints the resulting n, and matches it to the expected count — even if a few rows are legitimately dropped as genuinely missing. - **Consequence**: The coefficient shifts in the second decimal place (or the p-value string is malformed), so the graded numeric field mismatches the expected value even though the qualitative conclusion is right, and the attempt is scored wrong.
task -- the script trains a model on a labeled file and writes predictions for an unlabeled file, with no accuracy metric required in the deliverable.bf4dfa0fddbd · mined from da-code dacode-ml-regression-014### No held-out validation or distribution sanity check on the predictions - **Applies when**: `task` -- the script trains a model on a labeled file and writes predictions for an unlabeled file, with no accuracy metric required in the deliverable. - **Pattern**: The attempt fits one model configuration on 100% of the labeled data, immediately predicts on the target file, and reports only self-described summary statistics of the predictions — never computing an error estimate on a held-out split, never comparing the predicted distribution (min/median/mean/max, skew) against the training label distribution, and never checking that the test rows were preprocessed identically (same feature construction, same category vocabulary, same handling of unseen/missing categories, same target scale/units). - **Detection procedure**: 1. Read the task: identify the required output file, its column, expected row count, and the plausible scale/units of the quantity being predicted. 2. Read the script: check whether any split (train/validation, CV) is used to produce a numeric error estimate, and whether unseen categories/missing values in the target file are mapped in a way consistent with training (e.g., silently collapsed to an existing valid code rather than an explicit "unknown"). 3. Read the answer: compare the reported prediction statistics with the label statistics of the training data; flag if the mean/median/max are implausible or off by a large factor, or if a floor/clip was applied without justification. 4. Confirm whether any validation number is reported at all; if the only evidence of quality is "predictions were saved", the attempt is unverified. - **Discriminator**: A real violation is an attempt with zero out-of-sample error estimate *and* no comparison of predicted vs. observed label distributions; it is fine if the agent reports a CV/holdout score (even from a simple baseline) and shows the prediction distribution matching the training label distribution, or explains any deliberate shift. - **Consequence**: The submitted file has the right shape and column name but systematically miscalibrated values (inflated mean/median, absurd extremes), so the grader's tolerance/error check on the predicted values fails while the agent claims success.
task -- the task asks for a per-row output file (one prediction per input record) to be written for a held-out input set.head()/sampled subset, a debug run, or rows dropped by filtering/dropna — and/or in an order that no longer aligns with the input rows, while still claiming completeness.nrows=, .head(), .sample(), dropna/filtering, or partial-batch loop) and that row order is preserved (no sort, groupby, or reindex before writing).4e931ce66f66 · mined from da-code dacode-ml-multi-008### Prediction file row count/order doesn't match the evaluation input - **Applies when**: `task` -- the task asks for a per-row output file (one prediction per input record) to be written for a held-out input set. - **Pattern**: The attempt produces an output file whose number of rows is far smaller (or larger) than the number of records in the provided input file — e.g. only a handful of rows from a truncated preview, a `head()`/sampled subset, a debug run, or rows dropped by filtering/`dropna` — and/or in an order that no longer aligns with the input rows, while still claiming completeness. - **Detection procedure**: 1. From the task and data files, determine the expected output length = number of records in the held-out input file (read its shape, don't assume), and the required column name/header. 2. In the scripts, trace the object that is written out: check that it was built from the *full* input frame (no `nrows=`, `.head()`, `.sample()`, `dropna`/filtering, or partial-batch loop) and that row order is preserved (no `sort`, `groupby`, or reindex before writing). 3. Count the rows in the produced answer/file and compare to the expected length; also check the header and allowed label/value set. 4. Flag if the counts differ, if the count equals a suspicious small/round number, or if scripts are absent so the count cannot be traced to the full input. - **Discriminator**: A genuine violation is a length/alignment mismatch with the held-out input (or an untraceable pipeline); it is *not* a violation if the row count equals the input record count and the task itself legitimately requested an aggregated or filtered subset with that smaller size. - **Consequence**: The grader cannot align predictions to ground truth, so the submission is scored as wrong/missing regardless of model quality (or the score reflects only a tiny fraction of rows).
task -- the answer must include file paths to artifacts the agent creates (cleaned/transformed CSVs, model files, plots) that a grader will match or open.to_csv, savefig, dump) and record the literal or constructed output path; check whether it is derived from the input path or hardcoded to an unrelated directory.915e7b36967f · mined from infiagent-dabench dabench-743### Output artifacts written to an arbitrary directory instead of the task's working/data directory - **Applies when**: `task` -- the answer must include file paths to artifacts the agent creates (cleaned/transformed CSVs, model files, plots) that a grader will match or open. - **Pattern**: The agent saves outputs to its own home/current directory (or a temp path) rather than the directory the input data lives in / the directory implied by the task environment, and reports that path verbatim, so all path-bearing checks fail even though the computed numbers are right. - **Detection procedure**: 1. From the task statement, note where the input file(s) are located and any stated or conventional output location; treat the input's directory as the default expected output location when none is stated. 2. In the scripts, find every write call (`to_csv`, `savefig`, `dump`) and record the literal or constructed output path; check whether it is derived from the input path or hardcoded to an unrelated directory. 3. Compare the paths reported in the final answer to the input directory and to the paths actually written; flag any mismatch in directory, filename, or extension. 4. Confirm the reported paths are absolute and refer to files that the script actually creates (name spelled identically, no later overwrite/rename). - **Discriminator**: A real violation is an output path whose directory differs from where the data was read (or from an explicitly required location), or a reported path that differs from what the script writes; it is fine if the task explicitly permits any path and the reported path exactly matches a file the script created in the same data directory. - **Consequence**: Path-containing checks are marked WRONG/MISSING even when the numeric fields match, so the item scores 0.
task -- the task points to an external configuration/specification file (e.g., a YAML/JSON style guide) and/or names specific output files that the deliverable must consist of.371bb26355c6 · mined from da-code dacode-plot-pie-008### Incomplete compliance with an external plot/output specification (missing required artifacts) - **Applies when**: `task` -- the task points to an external configuration/specification file (e.g., a YAML/JSON style guide) and/or names specific output files that the deliverable must consist of. - **Pattern**: The agent reads only a few keys of the spec file (or hardcodes assumptions about it), produces just the one visually obvious artifact (an image), and skips the other mandated outputs (serialized plot data / numeric arrays / config echo), or applies the spec's directives (titles, labels, order, colors, units, size) only partially. The final answer is prose about the finding rather than a checklist of produced files matching the requested format. - **Detection procedure**: 1. From the task text, list every required output artifact (file names/extensions) and every constraint the referenced spec file is said to govern. 2. In the scripts, list every file actually written and every spec key actually consumed; compare against step 1, including whether the spec is fully enumerated/validated rather than accessed by a few hardcoded keys. 3. Check the answer/report: does it confirm each required artifact exists with the right content type (e.g., underlying counts/values saved, not just an image), and in the requested location? 4. Flag if any required artifact is never written, or any spec section is never read (e.g., keys present in the config that no code path uses). - **Discriminator**: A real violation is a required output file or spec directive with no corresponding code path. It is *not* a violation if the script writes all named artifacts and defensively reads the spec generically (e.g., iterating keys, with a printed dump verifying which directives were applied), even if some spec keys are simply unset/empty in the file. - **Consequence**: The grader checks each expected file independently; missing serialized data/config artifacts and partially styled figures score as WRONG/MISSING, so the run fails all checks even when the intermediate analytical finding is right.
task -- The task prescribes a fixed pipeline (given features, fixed split fraction and random seed, one metric) and the script must decide how to handle missing/invalid entries in the feature columns.dropna()) on the feature/target frame before splitting, so the number of rows — and therefore the exact train/test partition produced by the seeded split — differs from the intended full-dataset pipeline; no imputation or justification is given, and the resulting row/split counts are never reconciled against the raw data.dropna, boolean masks, drop_duplicates, index slicing) applied before train_test_split. 3. Check whether the script reports the raw row count vs. the post-cleaning count and whether the loaded file/columns actually still contain missing values that require handling; check whether an imputation (mean/median/mode) alternative was considered. 4. Confirm the answer was computed on this reduced set with no sanity comparison to the full-data variant.7eb80ae7a4ed · mined from infiagent-dabench dabench-7### Dropping rows with missing values silently changes the evaluated sample - **Applies when**: `task` -- The task prescribes a fixed pipeline (given features, fixed split fraction and random seed, one metric) and the script must decide how to handle missing/invalid entries in the feature columns. - **Pattern**: The script calls a blanket row-drop (e.g. `dropna()`) on the feature/target frame before splitting, so the number of rows — and therefore the exact train/test partition produced by the seeded split — differs from the intended full-dataset pipeline; no imputation or justification is given, and the resulting row/split counts are never reconciled against the raw data. - **Detection procedure**: 1. Read the task for any statement (or absence of a statement) about missing-value handling and note that the split is seed-fixed, so the row set fully determines the answer. 2. In the script, locate every operation that removes or filters rows (`dropna`, boolean masks, `drop_duplicates`, index slicing) applied before `train_test_split`. 3. Check whether the script reports the raw row count vs. the post-cleaning count and whether the loaded file/columns actually still contain missing values that require handling; check whether an imputation (mean/median/mode) alternative was considered. 4. Confirm the answer was computed on this reduced set with no sanity comparison to the full-data variant. - **Discriminator**: A real violation is unrequested row deletion in *feature* columns that shrinks the modeling sample (imputation would preserve it); it is fine to drop rows lacking the *target* (unlabeled rows cannot be scored), or to drop nothing because the columns are already complete — and it is fine if the script verifies the row count is unchanged after cleaning. - **Consequence**: The seeded split covers a different subset than the reference pipeline, so the reported accuracy is off by a few points (e.g. 0.76 vs. the expected 0.78) and the exact-value check fails.
task -- The task supplies a pre-existing output file (template/skeleton) and asks that results be entered into it "in the same format", and the scripts generate that file's contents.b9bcbaa11c91 · mined from da-code dacode-dm-csv-001### Ignoring the provided output template when filling a required results file - **Applies when**: `task` -- The task supplies a pre-existing output file (template/skeleton) and asks that results be entered into it "in the same format", and the scripts generate that file's contents. - **Pattern**: The attempt never reads or inspects the supplied template; it invents its own schema — column names, row labels/category names, category definitions (e.g. self-chosen bin boundaries), row ordering, and extra catch-all rows — and writes a fresh file from scratch, then "verifies" the format only against its own assumptions. - **Detection procedure**: 1. In the task text, identify the exact output artifact required and any statement that its existing format must be followed. 2. Search the scripts for any read/inspection of that template file (e.g. loading it, printing its rows/columns/label values) before writing. 3. Check whether the labels, categories, ordering, and column headers written out are derived from the template or hard-coded from the agent's own reasoning; check whether extra/renamed rows (e.g. an "unknown/other" bucket) or self-invented category thresholds were introduced. 4. Check whether the final answer is a narrative report rather than confirmation that the template file was filled in place with matching keys. - **Discriminator**: A real violation is inventing the key set/schema without ever reading the template; it is fine if the script loads the template, preserves its exact headers and row keys/order, and only populates the value cells — even if the agent additionally prints a human-readable summary. - **Consequence**: The required file fails key-by-key comparison against the expected file (mismatched row labels, extra rows, or values computed under different category definitions), so the grader reports the output file as WRONG/MISSING and scores 0.
task -- the agent fits a predictive model and must produce predictions for an unlabeled evaluation set that will be scored externally.predict on the training frame), reporting a high in-sample score as evidence the submission is good; no cross-validation or hold-out estimate, no comparison against a trivial baseline, and no check that the predicted positive rate / class balance is plausible relative to the training labels. Hyperparameters (depth, class weights, thresholds) are therefore chosen with no honest signal, and overfit or badly calibrated predictions go out unvalidated.fit, or a split/CV fold withheld from training?23ae9ba0fd0b · mined from da-code dacode-ml-binary-016### Self-reported performance measured on the training data instead of a held-out split - **Applies when**: `task` -- the agent fits a predictive model and must produce predictions for an unlabeled evaluation set that will be scored externally. - **Pattern**: The script trains on all labeled rows and then computes the quality metric with the same rows (or via `predict` on the training frame), reporting a high in-sample score as evidence the submission is good; no cross-validation or hold-out estimate, no comparison against a trivial baseline, and no check that the predicted positive rate / class balance is plausible relative to the training labels. Hyperparameters (depth, class weights, thresholds) are therefore chosen with no honest signal, and overfit or badly calibrated predictions go out unvalidated. - **Detection procedure**: 1. In the task, note that the scored artifact is predictions on data whose labels the agent never sees, so the only defensible quality claim is an out-of-sample estimate. 2. In the scripts, locate the metric computation and check which rows are passed to it: are they the same rows used in `fit`, or a split/CV fold withheld from training? 3. Check whether any baseline (majority class, simple logistic/default hyperparameters) is scored on that same held-out data for comparison, and whether the predicted label distribution on the evaluation set is compared to the training base rate. 4. In the answer, see whether the headline score is labeled as training/in-sample; if it is, or if its provenance is unstated and step 2 shows no split, treat the quality claim as unsupported. - **Discriminator**: A real violation is a metric whose inputs are provably a subset of the fitted rows with no CV/hold-out anywhere; it is *not* a violation if the agent held out or cross-validated (even with a modest score), or if it reports an in-sample number *in addition to* a clearly identified out-of-sample estimate used for model selection. - **Consequence**: The reported score is inflated and uninformative (e.g., near-perfect in-sample vs. much weaker true performance), so the submitted prediction file can fall below the grader's accuracy/AUC threshold — and its class mix be far off the true base rate — while the answer confidently claims success.
task -- the task points to a sample/reference output file for the result format and/or states an execution constraint (e.g., fix a random seed) that implies a specific computational approach.5b9026b242e9 · mined from da-code dacode-data-sa-039### Ignoring the provided output template and the method implied by stated constraints - **Applies when**: `task` -- the task points to a sample/reference output file for the result format and/or states an execution constraint (e.g., fix a random seed) that implies a specific computational approach. - **Pattern**: The attempt never loads or inspects the sample output file and instead invents its own column names/rows, and it substitutes a closed-form/analytic shortcut for the randomized procedure implied by the seed requirement (so the seed is never actually used), then reports that value as the requested statistic. - **Detection procedure**: 1. Read the task for (a) a named format/template file and (b) constraints such as a random seed, rounding, units, or ordering. 2. In the scripts, check that the template is actually read (or its exact header/row layout reproduced) and that every stated constraint appears in code — e.g., the seed is set and consumed by a stochastic step (resampling/permutation/bootstrap/model init). 3. Compare the written file's schema (column names, count, order, row keys) and the value's provenance against the template and the requested quantity; flag if the schema is self-invented or the seed is decorative/unused. 4. Check the reported answer restates the same schema and value as the saved file. - **Discriminator**: A real violation is inventing a schema or using a deterministic method when the task's constraints (template + seed) demand a specific format and a randomized computation; a look-alike that is fine is a deterministic method used when no randomized procedure is implied, or extra explanatory prose alongside a file that provably matches the template schema. - **Consequence**: The saved file fails exact-format/value comparison against the expected result (column names or the numeric p-value differ), scoring 0 even though the narrative conclusion may sound plausible.
task -- the deliverable is a prediction file for a held-out split that will be scored against hidden ground truth, and the agent reports only self-computed validation metrics.3d32ead079fb · mined from da-code dacode-ml-regression-015### Predictions not validated against the reference target distribution - **Applies when**: `task` -- the deliverable is a prediction file for a held-out split that will be scored against hidden ground truth, and the agent reports only self-computed validation metrics. - **Pattern**: The agent fits one default model, reports a mediocre in-house score (or none), and declares success from file-writing alone — never checking that predicted values plausibly reproduce the training target's distribution (mean, spread, min/max, skew) or that the row count/order/index matches the test input, and never trying to improve a clearly weak fit. - **Detection procedure**: 1) Read the task for the required file name, column name, row count and any ordering/format constraint. 2) In the scripts, check whether the same feature construction and preprocessing are applied to train and test, and whether predictions are written in the test file's original row order with no extra index column. 3) Compare the reported prediction summary statistics (mean, min, max, variance) with the training target's statistics — flag if the predictions are visibly compressed, shifted, clipped, or far narrower/wider than the target. 4) Flag if the reported goodness-of-fit is weak and no alternative model, tuning, or error analysis was attempted before finalizing. - **Discriminator**: A genuine violation shows missing sanity checks *and* evidence of a distribution/shape mismatch or an unimproved weak model; it is fine if the agent explicitly compared prediction stats to the target, verified row count/order/column name, and justified the model choice with a comparison — shrinkage toward the mean alone is normal for regression and not by itself a violation. - **Consequence**: The submitted file scores below the grader's accuracy/error threshold (or fails format/shape validation), so the expected output is marked WRONG despite the script running without error.
task -- the task asks to impute missing values and then rank/extremize a quantity that can either be read from one raw column or be recomputed from other columns in the data (e.g., a ratio of two available fields).6918c1c8dc9a · mined from da-code dacode-di-text-001### Trusting a single pre-existing column for a requested derived quantity, with a vacuous imputation step - **Applies when**: `task` -- the task asks to impute missing values and then rank/extremize a quantity that can either be read from one raw column or be recomputed from other columns in the data (e.g., a ratio of two available fields). - **Pattern**: The script picks the one column whose name resembles the requested quantity, cleans/imputes only that column, and reports its argmax/argmin — never checking that the stated preprocessing actually did anything (the column may have zero missing values, making the imputation a no-op) and never cross-validating the column against the quantity implied by its definition computed from the underlying fields. - **Detection procedure**: 1. From the task wording, write down the definition of the requested quantity and which raw fields it depends on; note the explicitly required preprocessing step. 2. In the scripts, check whether preprocessing is applied to all fields the quantity depends on (or dataset-wide as stated) or only to one convenient column, and whether the script prints/asserts that the imputation changed at least one value. 3. Check whether the script recomputes the quantity from its components and compares with the raw column (or at least verifies the raw column's units/scale and its extremes against known plausible ranges). 4. Inspect the reported extremes: if the reported min/max come solely from the raw column with no consistency check, and the imputation was a no-op, flag the attempt. - **Discriminator**: A fine attempt either shows that the raw column is consistent with the definition (recomputed values match, or the task unambiguously names that single column) and reports how many values the imputation filled; a violation silently assumes the column is correct and complete, leaving the mandated imputation unexercised and the definition unverified. - **Consequence**: The reported extreme entities come from an inconsistent/uncorrected column, so the submitted min and/or max labels differ from the expected ones and the result file fails the equality check.
task -- the task asks to compute several per-entity quantities (scores, groups/segments, labels) and save "the results, including X and Y" to one named output file, and the script builds a full result table plus a trimmed export.aa0e6f15023f · mined from da-code dacode-dm-csv-052### Requested output file omits the computed quantities the task asked to include - **Applies when**: `task` -- the task asks to compute several per-entity quantities (scores, groups/segments, labels) and save "the results, including X and Y" to one named output file, and the script builds a full result table plus a trimmed export. - **Pattern**: The script computes all the intermediate per-entity metrics and derived scores in a working dataframe, but writes only a minimal subset of columns (e.g., ID + final label) to the required filename, dumping the complete table to a differently-named auxiliary file; the answer then claims completeness because the auxiliary file exists. - **Detection procedure**: 1. From the task text, list every quantity that must appear in the named output file (each component metric, each derived score, the grouping/segment, the final level). 2. In the script, find the write call to that exact filename and enumerate the columns actually selected there. 3. Diff the two lists; also check whether a richer table was written to a *different* filename. 4. Check the answer's description of the file's contents against the task's required contents. - **Discriminator**: A real violation is when quantities the task explicitly named (or the per-entity scores the task's phrasing requires) are absent from the required file, even if present elsewhere on disk. It is fine if the file has extra columns, different column ordering/names, or if the task genuinely asked only for the final label and the extra table is a bonus. - **Consequence**: The grader compares the named file against a reference containing the full set of columns and marks it WRONG/MISSING (0 checks passed) despite the underlying computation being reasonable.
task -- the task asks for a statistic over a specific subgroup/condition (e.g., one category, species, region, cohort) while the raw files contain multiple groups, and/or a template file defines the required output columns/values.df[df[group_col] == ...]) on the group column and for any inspection/validation of the actual columns present in each input file; check whether the template/sample output file is ever read or its columns replicated.98d4b8fe741f · mined from da-code dacode-data-sa-029### Ignoring a required subgroup filter (and unverified schema) when aggregating a statistic - **Applies when**: `task` -- the task asks for a statistic over a specific subgroup/condition (e.g., one category, species, region, cohort) while the raw files contain multiple groups, and/or a template file defines the required output columns/values. - **Pattern**: The script reads each file and aggregates over every row, assuming the file already contains only the requested subgroup and that the column names it hard-codes exist as written; no filtering step, no schema inspection, and no comparison against the provided sample output file. - **Detection procedure**: 1. Read the task and note every restriction on rows (group/category, year, condition) and every stated output constraint (column names, ordering, rounding, format template). 2. In the script, look for an explicit filtering operation (`df[df[group_col] == ...]`) on the group column and for any inspection/validation of the actual columns present in each input file; check whether the template/sample output file is ever read or its columns replicated. 3. Check whether the script prints or asserts row counts / value ranges per group after loading, so a mismatch between the filtered subgroup size and the full file size would be visible. 4. Compare the reported numbers to a rough domain sanity expectation (e.g., counts identical across heterogeneous files, or a ratio suspiciously near 1 when the two measured quantities are known to differ substantially) — an unexplained coincidence signals the filter or column mapping was wrong. - **Discriminator**: A real violation is when the raw file plausibly contains rows outside the requested subgroup (or differently named columns) and the script never filters/validates; it is fine if the script demonstrably verifies the subgroup (asserts unique group values, prints per-group counts, or the file is documented as already restricted) and reproduces the template's exact columns. - **Consequence**: The reported means/CIs are computed on the wrong row subset (or wrong columns) and the output file fails an exact/tolerance comparison to the expected result.csv, scoring 0.
task -- the task asks for a statistic/model over a whole dataset file, and the script reads the data (possibly with nrows, a preview/head, a single chunk, one of several files, or a hand-made sample) before computing the requested output.nrows, head, slicing, chunk iteration, sampling, reading only one of multiple input files).f5f9da84cf97 · mined from da-code dacode-data-sa-026### Loaded only a truncated slice of the source data (and never sanity-checked the result) - **Applies when**: `task` -- the task asks for a statistic/model over a whole dataset file, and the script reads the data (possibly with `nrows`, a preview/head, a single chunk, one of several files, or a hand-made sample) before computing the requested output. - **Pattern**: The attempt computes the requested statistic on a small subset of the real data (e.g., a first-N-rows preview kept from an exploratory step) while reporting it as the full-dataset result, and accepts numerically implausible values (near-zero or wrong-signed correlations among conceptually related variables, tiny counts) without any cross-check against the file's true size or expected direction. - **Detection procedure**: 1. From the task, note that the statistic must cover the entire dataset; find the loading call in the script and check for row limits (`nrows`, `head`, slicing, chunk iteration, sampling, reading only one of multiple input files). 2. Compare the row count printed/reported in the answer against the actual size of the source file (line count / shape); flag suspiciously round or small numbers (e.g., exactly 100) as a preview artifact. 3. Check whether the script prints/validates shape, non-null counts, and value ranges of the analyzed columns, and whether the answer's numbers are directionally plausible for related variables. 4. Flag if the computation runs on anything other than all rows of the intended source, or if no shape/plausibility check exists. - **Discriminator**: A real violation is a row limit or partial read that affects the *final* computation, or a reported N that is far below the file's true row count; it is fine if the limited read is used only for schema inspection/debugging while the final computation reloads the complete file, or if the reduced N is fully explained by the required missing-value filtering on a full read. - **Consequence**: The saved output file contains statistics from an unrepresentative subsample, so every cell differs from the expected values and the file-comparison check fails (0/1), even though the format looks correct.
task -- the task asks for an "appropriate" number of groups/components (or any hyperparameter) and the script picks it by scanning a bounded grid and taking the arg-max/arg-min of an internal metric.a6874f279501 · mined from da-code dacode-ml-cluster-010### Cluster count chosen at the edge of the search grid, ignoring known group structure - **Applies when**: `task` -- the task asks for an "appropriate" number of groups/components (or any hyperparameter) and the script picks it by scanning a bounded grid and taking the arg-max/arg-min of an internal metric. - **Pattern**: The attempt sweeps k over a hard-coded range, selects the value with the best score, and that winner lands on the last (or first) value of the range with a weak absolute score — so the "optimum" is an artifact of where the sweep was truncated, not a real structural optimum. It also disregards strong external evidence of the natural group count available in the data (e.g., a categorical/label column that was dropped, or a documented set of classes), and never re-runs with a wider range or a second criterion (elbow, stability, comparison against the known cardinality). - **Detection procedure**: 1. In the task/README, note any stated or implied number of natural groups (a target/label column the script drops, documented categories, domain description). 2. In the script, find the candidate range and the selection rule; check whether the selected value can be an endpoint of that range and whether any secondary check (elbow, stability, agreement with known cardinality) gates the choice. 3. In the answer, compare the reported chosen value to the range bounds and look at the absolute value of the selection metric (e.g., a silhouette near 0.1–0.2 means almost no separation, so the arg-max is noise). 4. Flag if the choice equals a range endpoint, or if the metric is essentially flat/weak, and no justification ties the choice to the data's known structure. - **Discriminator**: Fine if the winning value lies strictly inside the range with a clear peak, or if an endpoint win is confirmed by extending the range or by an independent criterion / the known group count; a violation is an endpoint (or noise-level) arg-max accepted as-is with no widening and no cross-check. - **Consequence**: The saved cluster labels use a partition count that disagrees with the reference grouping, so label-count / agreement checks (e.g., number of distinct clusters, ARI/NMI or cluster-quality thresholds) on the output file fail even though the file itself is well formed.
task -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script picks the number of groups by maximizing an internal score over raw aggregated, heavy-tailed features.e76127f08a60 · mined from da-code dacode-ml-cluster-019### Degenerate cluster solution driven by unhandled outliers/skew - **Applies when**: `task` -- the task asks for an unsupervised grouping into "an appropriate number of groups" and the script picks the number of groups by maximizing an internal score over raw aggregated, heavy-tailed features. - **Pattern**: The script builds skewed count/monetary aggregates, applies only mean/variance scaling (no log/robust transform, no outlier handling), then selects the configuration with the best silhouette — which is almost always the smallest number of groups that isolates a handful of extreme records, yielding one giant group plus a few singletons and no interpretable segmentation. - **Detection procedure**: 1) Read the task to confirm the deliverable is a meaningful multi-group segmentation, not just any label column. 2) In the script, check whether heavy-tailed features are log/rank/robust-transformed or winsorized before scaling, and whether the model-selection loop has any balance/interpretability guard beyond a single internal index. 3) Read the reported group sizes and score: flag if the near-perfect score coincides with a group containing a negligible fraction (e.g. <1%) of rows and the remainder in a single group. 4) Check whether the agent noticed and diagnosed this degeneracy or accepted it as "excellent quality". - **Discriminator**: A genuine violation is a mathematically extreme score produced by outlier isolation with essentially no partition of the bulk of the data; it is fine if a small group exists alongside several substantively sized groups, or if the agent explicitly justifies the small group after outlier treatment and shows the bulk is still meaningfully partitioned. - **Consequence**: The saved label file has near-zero effective segmentation (one dominant label), so any grader check on cluster count, label distribution, or segment-quality/consistency fails even though the file format is correct.
task -- the task supplies a dataset (with a README) and asks for a statistic or confidence interval computed from it, and the script defines the inputs inline rather than loading any file.read_csv, read_excel, load, path strings) and check that the quantity is derived from those loaded objects; flag literal constants that stand in for data.4a57796ac07a · mined from da-code dacode-data-sa-031### Hardcoded/recalled numbers instead of resampling the provided data - **Applies when**: `task` -- the task supplies a dataset (with a README) and asks for a statistic or confidence interval computed from it, and the script defines the inputs inline rather than loading any file. - **Pattern**: The attempt never reads the provided data files; it types in summary counts or rates from memory/background knowledge and then simulates from a parametric assumption (e.g., drawing from a Binomial with the assumed rate) instead of resampling the actual observed records, so both the point estimate and the interval reflect invented inputs and an unrequested method. - **Detection procedure**: 1. Read the task/README and list the data files that are expected to be the source of the requested quantity. 2. Scan the script for any I/O (`read_csv`, `read_excel`, `load`, path strings) and check that the quantity is derived from those loaded objects; flag literal constants that stand in for data. 3. Check that the resampling/estimation step draws with replacement from the loaded observations (or their groups) rather than from a hand-specified distribution or parameter. 4. Compare the reported numbers against what the actual files would yield (row counts, group sizes, the observed effect size) — if the script prints counts that cannot be traced to a file, treat the result as unverified. - **Discriminator**: A real violation is when key inputs (counts, rates, group definitions, filters) originate from the agent's assumptions and are never cross-checked against the delivered files; it is fine to hardcode constants that are explicitly stated in the task/README, or to compute summaries in one script and pass them along, as long as they were derived from the loaded data. - **Consequence**: The output file contains an interval centered on the wrong effect size (and with the wrong width from a parametric rather than empirical bootstrap), so the graded value falls outside the expected tolerance and the result file is marked WRONG.
task -- the task requires writing an output file whose rows/keys and column layout must match a provided sample/template file over an entire held-out set.assert len(sub) == len(test), set(sub[key]) == set(template[key]), column-name/order equality, and a check that written file re-reads with the expected shape. Absence of all of these is the violation.06ba55fe8222 · mined from da-code dacode-ml-competition-005@s2### Missing verification that the prediction file exactly matches the required submission template - **Applies when**: `task` -- the task requires writing an output file whose rows/keys and column layout must match a provided sample/template file over an entire held-out set. - **Pattern**: The scripts build predictions and write the output directly from model probabilities, but never load the template, never assert that the number of rows equals the number of held-out records, that the key column covers exactly the same identifiers (same set, same order), and that the column names/order match; the final answer is accepted without any row-count or key-coverage sanity check (and can silently be truncated, reordered, subset, or column-swapped). - **Detection procedure**: 1. Read the task for the stated output contract: required file name, header names, key column, and the fact that every held-out record must appear exactly once. 2. Scan the scripts for an explicit comparison against that contract — e.g. reading the sample/template file, `assert len(sub) == len(test)`, `set(sub[key]) == set(template[key])`, column-name/order equality, and a check that written file re-reads with the expected shape. Absence of all of these is the violation. 3. Inspect the produced answer/file: count rows and compare to the held-out record count; check the key column for duplicates, missing ids, and that class-probability columns are in the specified order and non-degenerate. 4. Also confirm the class-index-to-column mapping is derived from the fitted model's class ordering rather than assumed positionally. - **Discriminator**: A real violation is the total absence of any template/row-count/key check (or an actual mismatch in the emitted file). It is not a violation if the script reconstructs the submission from the template's key column (e.g. merges predictions onto the template ids) or asserts shape/ids/columns before saving — the checks may be terse, but they must exist and cover keys, count, and column layout. - **Consequence**: The grader reports the expected output file as WRONG/MISSING regardless of model quality, since rows are missing/misaligned or columns are mismapped, and the log-loss cannot be computed over the full held-out set.
task -- the agent trains a regressor to produce a predicted count/quantity/price column for held-out rows and writes it to an output file.6ffdf2b3a3f1 · mined from da-code dacode-ml-regression-008@s2### Predicted values not sanity-checked against the training target's scale and type - **Applies when**: `task` -- the agent trains a regressor to produce a predicted count/quantity/price column for held-out rows and writes it to an output file. - **Pattern**: The script transforms the target (log/sqrt/scaling) or fits on a filtered/atypical subset, then writes raw model output without inverting the transform or checking it against the observed target distribution — producing predictions whose scale, dtype (e.g., fractional values for integer counts), or row count is inconsistent with the training labels, with no assertion or comparison step. - **Detection procedure**: 1. From the task/README, note what the target represents (units, integer vs continuous, plausible range) and how many output rows are required. 2. In the scripts, find every target transformation, clipping, subset filter, or scaling; verify an explicit inverse transform is applied before writing, and that the output is built from the full held-out set in original order. 3. Compare the answer's summary statistics (min, median, max, count, share of near-zero values) against the training target's statistics; flag if the central tendency differs by an order of magnitude or the value type is incompatible (e.g., mostly sub-unit values where labels are large integers). 4. Check the script contains an explicit sanity assertion on shape/range before saving; absence plus a mismatch in step 3 confirms the violation. - **Discriminator**: A genuine violation shows a systematic distribution mismatch (e.g., predictions concentrated far below the label median, or row count ≠ required rows). It is *not* a violation if predictions are merely smoother/less dispersed than labels — regression shrinkage toward the mean is expected as long as the central tendency and units match the labels. - **Consequence**: The output file fails the grader's value/format comparison — error metrics are dominated by a constant scale bias (or the file has the wrong number/type of entries), so the answer is marked wrong despite a plausibly reasonable model.
task -- the task asks for a p-value / significance decision about a statistic of some group comparison, and the scripts run an off-the-shelf test on whole raw tables straight from disk.query/boolean mask/date parse restrict the rows before the statistic is computed? If the row counts used are essentially the full file sizes, the scoping step is missing.ttest_ind on a bounded count variable with no diagnostic is a red flag.519e6a3e5f2c · mined from da-code dacode-data-sa-001@s2### Unfiltered population and unchecked test assumptions in hypothesis testing - **Applies when**: `task` -- the task asks for a p-value / significance decision about a statistic of some group comparison, and the scripts run an off-the-shelf test on whole raw tables straight from disk. - **Pattern**: The script loads every row of each source file, computes the derived quantity, and immediately applies a default parametric test (e.g., independent-samples t-test) without (a) restricting to the subpopulation, time window, or competition/segment that the question is actually about, and (b) checking whether the test's distributional assumptions (normality, symmetry, equal variance, independence) hold for the derived quantity — typically a small-integer, skewed, count-like variable. - **Detection procedure**: 1. Read the task/README and list every scoping qualifier implied by the question (which subset of records, which period, which category the claim concerns) plus the significance level and output format. 2. Read the script for filtering steps: does any `query`/boolean mask/date parse restrict the rows before the statistic is computed? If the row counts used are essentially the full file sizes, the scoping step is missing. 3. Inspect the test choice: is there any diagnostic (histogram, skew, normality test, variance comparison) or justification, and does the test's one/two-sided direction match the stated hypothesis? A default `ttest_ind` on a bounded count variable with no diagnostic is a red flag. 4. Check the reported p-value's magnitude for plausibility given the intended (much smaller) subset — an astronomically small p-value is a symptom of using far more rows than the question intended. - **Discriminator**: A real violation is when the task or README implies a narrower population/assumption-appropriate test and the script silently uses all rows with a default parametric test; it is fine if the task genuinely asks about all records and the script either documents the assumption check or the chosen test is robust/appropriate for the variable's distribution. - **Consequence**: The p-value (and possibly the reject/fail-to-reject decision) is computed on the wrong sample with the wrong test, so the saved value does not match the expected one and the file fails the grader even though the format is correct.
task -- the task supplies a sample/expected output file (or explicit format spec) and asks for results "following the exact structure and formatting" of it.round(), float_format, sort key derived from the template, column list copied from the template header). Absence of both is the red flag.0d14093b2bf3 · mined from da-code dacode-dm-csv-011@s2### Ignoring the provided output template (format/rounding/ordering never verified) - **Applies when**: `task` -- the task supplies a sample/expected output file (or explicit format spec) and asks for results "following the exact structure and formatting" of it. - **Pattern**: The scripts compute aggregates but never read or parse the template file; column names, column order, row order, and numeric precision are guessed from intuition, and raw floating-point sums are written verbatim (e.g. values with 10+ decimal digits or inconsistent decimal counts across rows), so the output can't byte-match the reference even when the underlying math is right. - **Detection procedure**: 1. In the task text, note whether a template/sample output artifact or explicit formatting rules (rounding, units, ordering, header names) are mentioned. 2. Search the scripts for any read of that template (e.g. loading the sample file) and for any explicit formatting step (`round()`, `float_format`, sort key derived from the template, column list copied from the template header). Absence of both is the red flag. 3. Inspect the produced answer: check that all numeric cells share a consistent, plausible precision, that header names/order match the template exactly, and that row order follows a rule that is stated or derivable from the template rather than an ad-hoc choice. 4. Confirm no post-hoc comparison of the generated file against the template's shape/columns/dtypes was performed. - **Discriminator**: A real violation is when the template is never inspected and the formatting choices (precision, ordering, header) are unverified guesses. It is *not* a violation if the script reads the template (or the task states the format literally) and demonstrably enforces the same columns, order, and rounding — even if it does so with hardcoded values copied from the template. - **Consequence**: The grader's exact/tolerance file comparison fails on formatting alone (extra decimal digits, wrong row order, or mismatched headers), reporting the result file as WRONG despite arguably correct aggregation logic.
task -- the task asks for a result file whose columns are the per-record feature values (e.g., Feature_i) plus a derived label, and the script builds that file from an ad-hoc, self-chosen preprocessing pipeline.to_csv: dropped columns, imputations, encodings, scaling, row filtering — and check whether the written matrix is the raw feature values, the transformed ones, or an inconsistent mix, and whether row count still equals the number of input records.7fbf54e450eb · mined from da-code dacode-ml-cluster-014@s2### Output feature matrix silently redefined and never sanity-checked against the requested spec - **Applies when**: `task` -- the task asks for a result file whose columns are the per-record feature values (e.g., `Feature_i`) plus a derived label, and the script builds that file from an ad-hoc, self-chosen preprocessing pipeline. - **Pattern**: The agent drops, imputes, one-hot expands and reorders columns for modeling convenience, then dumps that transformed matrix as the "feature vector" without checking that the saved file's row count, column count/order, value scale and label column are a defensible, non-degenerate representation of the original records; the choice of the number of clusters/labels is also taken straight from an unexamined metric on that arbitrary space (often collapsing to the minimum allowed value, e.g. 2, with a very low score) and never sanity-checked. - **Detection procedure**: 1. Read the task statement and note exactly what each output column is supposed to contain (which records, which feature values, what label) and where the file must be written. 2. In the script, trace every transformation between loading the raw data and `to_csv`: dropped columns, imputations, encodings, scaling, row filtering — and check whether the written matrix is the raw feature values, the transformed ones, or an inconsistent mix, and whether row count still equals the number of input records. 3. Check for any post-write validation: re-reading the file, asserting shape/column names/dtypes, label counts and no NaNs, and a check that the produced grouping is non-trivial (more than a minimal split, balanced enough, score meaningfully above the alternatives). 4. Read the answer text: does it merely restate the pipeline and the selection metric, or does it show these validation numbers and justify the discarded/derived columns? - **Discriminator**: A real violation is when transformations that change the feature vector's identity (dropping informative columns, expanding categoricals, imputation) are made without justification *and* no shape/content/degeneracy check on the saved file appears anywhere; it is *not* a violation if the script documents a principled feature definition and verifies the written file (rows = records, expected column naming/ordering, sensible label distribution), even if the exact preprocessing differs from a reference. - **Consequence**: The graded file has a feature matrix and/or label column that cannot be matched to the expected output (wrong width, transformed values, trivial or collapsed clustering), so the file check fails outright even though the script ran without error.
task -- the task supplies a sample/template output file and asks for predictions written to a submission file in that exact format.set(ids) equality, column-name equality, ordering, null check) before writing the output.3c6582a36cc3 · mined from da-code dacode-ml-competition-009@s2### Never validating the output file against the provided submission template - **Applies when**: `task` -- the task supplies a sample/template output file and asks for predictions written to a submission file in that exact format. - **Pattern**: The scripts build a submission purely from the agent's own assumptions (hand-written column names, ids taken from an intermediate frame, predictions coerced/rounded to a chosen dtype) and never load the template to check column names, row count, id set/order, or value type; a later "improved" script may also be truncated or fail, silently leaving an earlier or partial file on disk. - **Detection procedure**: 1. Read the task/README for the named template file and any stated format or metric constraints (column names, one row per test id, continuous vs. integer targets). 2. Search the scripts for a read of that template and for an explicit comparison against it (shape equality, `set(ids)` equality, column-name equality, ordering, null check) before writing the output. 3. Check that the final produced file comes from the last script that ran to completion, and that any post-processing (rounding, clipping, casting) is justified by the stated metric rather than applied by default. 4. Inspect the answer itself: count rows and compare with the expected test size, confirm the header matches the template, and confirm no duplicate/missing ids. - **Discriminator**: A real violation is the absence of any template-based verification *and* an output whose shape/ids/dtype cannot be shown to match; it is fine if the script derives ids directly from the template or test file and asserts shape/id agreement (even without printing), or if rounding is explicitly required by the task. - **Consequence**: The grader reports the submission as WRONG/MISSING — row count or id set mismatched, wrong header, or unnecessary integer rounding inflating the error metric — even though the model itself may be reasonable.
task -- the task asks to run an unsupervised/modeling procedure on "the dataset" and to output the feature vector alongside results (e.g., columns named Feature_i), without naming which columns to use.Feature_0..Feature_k) against the expected dimensionality from step 2 and against the row count of the output.Feature_i columns) and the cluster assignments diverge from the reference, so an exact/структural file comparison fails even though the pipeline ran without error.3a42457fcc6b · mined from da-code dacode-ml-cluster-009@s2### Arbitrary feature subsetting when the task implies the full feature vector - **Applies when**: `task` -- the task asks to run an unsupervised/modeling procedure on "the dataset" and to output the feature vector alongside results (e.g., columns named `Feature_i`), without naming which columns to use. - **Pattern**: The script silently keeps only a handful of hand-picked columns (often a thematically related few, or only those with no missing values) instead of using all usable columns after standard cleaning, so the emitted feature matrix has far fewer dimensions than the data provides and the labels come from a different feature space than the reference solution. - **Detection procedure**: 1. Read the task for any statement restricting the columns; if none, the default is "all informative columns after cleaning" (drop pure identifiers/text keys, coerce numeric strings, encode or drop categoricals, impute or drop missing values in a documented way). 2. In the script, find the column-selection step and count how many columns enter the model; compare with the number of candidate columns in the raw file. 3. Check whether each exclusion is justified by a stated rule (non-numeric, identifier, all-missing) rather than by topical judgment or convenience; check whether missing values were dropped/imputed instead of used as a reason to discard whole columns. 4. Compare the answer's reported feature count/columns (`Feature_0..Feature_k`) against the expected dimensionality from step 2 and against the row count of the output. - **Discriminator**: A real violation is dropping numeric, mostly-populated columns for subjective reasons (or to avoid handling NaNs/dtype cleanup); it is fine to drop columns that are identifiers, codes, free text, or unparseable, or to reduce dimensionality by an explicit method the task allows, when the reasoning is stated and applied uniformly. - **Consequence**: The output file's schema (number of `Feature_i` columns) and the cluster assignments diverge from the reference, so an exact/структural file comparison fails even though the pipeline ran without error.
task -- the deliverable is a per-record/per-entity output file (labels, predictions, scores) covering the units derived from the input data, and the script applies filtering or aggregation before producing it.dropna/threshold applied after the entity table is built, and note how many units each removes; check whether any is driven purely by a self-created feature being undefined rather than by data validity.Feature_0… vs Feature_1…, presence/absence of an ID column, ordering).f78080af2b84 · mined from da-code dacode-ml-cluster-016@s2### Silently dropping entities that must appear in the required output file - **Applies when**: `task` -- the deliverable is a per-record/per-entity output file (labels, predictions, scores) covering the units derived from the input data, and the script applies filtering or aggregation before producing it. - **Pattern**: The script removes a substantial subset of the entities as a side effect of preprocessing or feature construction (e.g. dropping rows where a derived statistic is undefined/NaN, requiring a minimum count, trimming "outliers"), so the saved file has far fewer rows than the natural population of entities in the cleaned data — and the answer never reconciles this row count against the input. - **Detection procedure**: 1. From the task, identify the unit of the requested output file and the expected row count implied by the cleaned input (e.g. number of distinct entities after only the clearly-justified cleaning steps). 2. In the scripts, list every filter/`dropna`/threshold applied *after* the entity table is built, and note how many units each removes; check whether any is driven purely by a self-created feature being undefined rather than by data validity. 3. Compare the row count written to the output file with the count from step 1; also check the column naming/indexing convention matches the literal spec (e.g. `Feature_0…` vs `Feature_1…`, presence/absence of an ID column, ordering). 4. Check the answer text: does it state the final row count and explain the gap, or does it present coverage loss as a methodological choice without validating it against the requirement? - **Discriminator**: A real violation is dropping units that are valid members of the target population purely for algorithmic convenience (undefined std, "needs ≥2 events", outlier removal) with no instruction to do so; a look-alike that is fine is removing records that cannot be attributed to any unit at all (missing entity key) or that the task/README explicitly designates as invalid (cancellations), where the remaining set is still the full population. - **Consequence**: The saved file has the wrong shape/coverage (thousands of entities missing) and, if the grader compares row counts, column names, or per-entity labels, it fails the file check outright even though the clustering itself may be reasonable; degenerate singleton clusters from unremoved extremes are a further symptom that no sanity check on cluster sizes was run.
task -- the task asks for predictions (or derived values) written to a named file with an explicitly stated column name, one row per input record.to_csv) and check what DataFrame is written: which columns it contains, whether the header spelling matches exactly, whether an index is written, and whether the frame was built from the full, unfiltered, un-reordered test input (no dropped rows from NaN handling, no re-sorting, no deduplication).eac24519e6c2 · mined from da-code dacode-ml-regression-002@s2### Output file schema/row-alignment not verified against the required submission spec - **Applies when**: `task` -- the task asks for predictions (or derived values) written to a named file with an explicitly stated column name, one row per input record. - **Pattern**: The agent writes the file with its own convention — extra index/key columns, renamed or extra headers, a different row count than the test input, or rows in an order other than the input's — and reports only model quality statistics, never checking that the artifact matches the literal requested format. - **Detection procedure**: 1. From the task statement, extract the exact required artifact: filename, required column name(s), expected number of rows (= rows in the provided test input), and implied row order. 2. In the scripts, find the write step (e.g., `to_csv`) and check what DataFrame is written: which columns it contains, whether the header spelling matches exactly, whether an index is written, and whether the frame was built from the full, unfiltered, un-reordered test input (no dropped rows from NaN handling, no re-sorting, no deduplication). 3. In the answer, look for an explicit verification of shape/columns against the test file (row count equality, column list equality, no NaNs); absence of such a check, or a reported row/column set that differs from the input's, is the flag. 4. Cross-check the reported prediction count and column list against the test input's row count and the required header. - **Discriminator**: A real violation is a mismatch in the artifact itself — missing/renamed/extra columns, row count ≠ test rows, or rows realigned by dropping/sorting. It is *not* a violation if extra columns are explicitly permitted by the task, or if the agent kept an auxiliary column but verified counts and order and the required column is present and correctly named; nor is strong/weak model accuracy relevant here. - **Consequence**: The grader loads the file, fails schema or index alignment (or compares mismatched rows), and marks the result wrong/missing regardless of how good the underlying model was.
task -- the task asks for a chart/output built to a spec file and the grading depends on the underlying plotted data being recoverable (e.g., a serialized plot description and/or a numeric array), not just an image file..npy/CSV of plotted values) and every constrained property (labels, ticks, order, units, aggregation level).ffe44a540300 · mined from da-code dacode-plot-line-015@s2### Unverifiable, artifact-incomplete plotting claims (no saved script, no data dump) - **Applies when**: `task` -- the task asks for a chart/output built to a spec file and the grading depends on the underlying plotted data being recoverable (e.g., a serialized plot description and/or a numeric array), not just an image file. - **Pattern**: The agent reports a checklist of "spec met ✓" items and an image file, but leaves behind no runnable script and no machine-readable dump of the plotted series; the numbers quoted appear to be typed in or eyeballed rather than aggregated from the source file, and the auxiliary artifacts the grader reads are never written. - **Detection procedure**: 1. Read the task/spec and list every expected output artifact (image, JSON/spec dump, `.npy`/CSV of plotted values) and every constrained property (labels, ticks, order, units, aggregation level). 2. Check the saved scripts: is there code that loads the raw source, performs the stated aggregation, plots, and then explicitly serializes the plotted x/y values and figure metadata to the required artifact names? 3. Compare the numbers in the answer to the aggregation the script would produce — do they trace to a computed object, or are they hard-coded/suspiciously patterned (e.g., a value equal to the group label, implausibly small counts for the stated unit)? 4. Confirm the answer asserts nothing that isn't demonstrably read back from the produced artifacts (e.g., re-reading the dump and printing shape/range). - **Discriminator**: A real violation is when required non-image artifacts are absent or the plotted values cannot be reproduced from code over the raw data; a look-alike that is fine is a script that computes and dumps the series programmatically and merely *summarizes* it in prose, even if only the image is visually inspected. - **Consequence**: The grader's checks on the expected artifacts (plot spec JSON, numeric array) report WRONG/MISSING, and any hard-coded series that differs from the true aggregation fails value comparison — scoring 0 despite a confident "task completed" report.
task -- the task/README explicitly dictates the inference procedure (e.g., shift each group to a common mean, then bootstrap each group separately and count replicates at least as extreme as the observed statistic) and the answer is a single p-value written to a template file.48a055917476 · mined from da-code dacode-data-sa-028@s2### Prescribed resampling procedure replaced by a different (more "extreme") test - **Applies when**: `task` -- the task/README explicitly dictates the inference procedure (e.g., shift each group to a common mean, then bootstrap each group separately and count replicates at least as extreme as the observed statistic) and the answer is a single p-value written to a template file. - **Pattern**: The attempt computes a p-value with a substitute method (parametric t-test, label-permutation test, pooled single-sample bootstrap, resampling only one group, or a one-sided/two-sided convention different from the one implied), and/or reports a Monte-Carlo p-value at or below the resolution floor of its replicate count (e.g., a handful of hits out of 10^5 draws) without checking that the number is stable or plausible. - **Detection procedure**: 1. From the task/README, write down the required steps: what statistic is observed, how the null is imposed (shifting vs. relabeling), which units are resampled, replicate count, and the tail definition. 2. In the script, locate the null-generation code and check line-by-line that each group is resampled from its own shifted copy (not pooled/permuted), that the observed statistic is recomputed identically, and that the extremity comparison matches the stated tail and sign. 3. Check the reported value against the simulation's granularity (p ≥ 1/replicates) and rerun-stability with a different seed; also compare order of magnitude to what the prescribed method plausibly yields versus a parametric shortcut. 4. Verify the output file exists with the exact column name/format/precision from the template, containing the final p-value (not a count, a test statistic, or an unrounded scientific-notation surprise) — and that a script was saved so the procedure is reproducible. - **Discriminator**: A real violation is a null model or tail convention structurally different from the one specified (pooling/permuting instead of shifting, resampling one group, wrong comparison direction), or a p-value whose magnitude is an artifact of too few replicates; it is *not* a violation if the prescribed procedure is implemented and the p-value merely differs in the last digits due to random seed, provided it is well above 1/replicates and formatted as requested. - **Consequence**: The written file's p-value differs from the reference by orders of magnitude (or fails tolerance/format matching), so the result-file check fails even though the pipeline "ran successfully".
task -- The prompt points to auxiliary instruction or configuration files (e.g., a tips/notes file, a YAML/JSON spec) that define methodology, styling, and/or required deliverables.01a98f571698 · mined from da-code dacode-plot-line-006@s2### Ignoring provided specification/config files and their required output artifacts - **Applies when**: `task` -- The prompt points to auxiliary instruction or configuration files (e.g., a tips/notes file, a YAML/JSON spec) that define methodology, styling, and/or required deliverables. - **Pattern**: The agent never opens or parses those files; it guesses the methodology and plotting parameters from the task wording, hard-codes its own choices (variable selection, grouping, labels, colors, figure size, axis ranges), and produces only the single obviously-named output while skipping companion artifacts the spec implies (serialized plot data, arrays, config echoes). - **Detection procedure**: 1. List every instruction/config file and every output artifact named or implied by the task statement. 2. Search the scripts for code that reads each of those files (open/read/yaml.safe_load/json.load) and for code that writes each expected artifact. 3. Check whether formatting/aggregation choices in the plotting or computation code are traceable to values loaded from the config, or are literals invented by the agent; check whether the answer text quotes the actual file contents versus paraphrasing assumptions. 4. Verify the answer reports the same set of deliverables the task requires, not just one. - **Discriminator**: A real violation is when no code path ever reads the spec file(s) or when required outputs are missing entirely; it is fine if the agent read the files and then legitimately inlined their values (evidence: printed/quoted contents, parameter names matching the spec) and produced all requested artifacts. - **Consequence**: Graders comparing each expected artifact fail on missing files, and the produced figure mismatches the reference on styling and on the underlying aggregated series, giving 0 passed checks.
task -- The prompt points to an external document/notes file (e.g. a markdown/readme/spec in the working directory) that defines how to bin, group, filter, or compute the requested quantity.value_counts() of a raw column.cb2210c59cb1 · mined from da-code dacode-plot-bar-005@s2### Ignoring a task-referenced specification file that defines the required method - **Applies when**: `task` -- The prompt points to an external document/notes file (e.g. a markdown/readme/spec in the working directory) that defines how to bin, group, filter, or compute the requested quantity. - **Pattern**: The scripts never open or echo the referenced file; the agent instead assumes the data's own native categories/defaults (e.g. pre-existing bucket labels or an ad-hoc ordering it invents) and produces outputs whose granularity, labels, or aggregation cannot be verified against the stated method. - **Detection procedure**: 1. List every external file, rule, or convention the task text references, plus every explicitly required output artifact/name. 2. Grep the scripts for a read/print of each referenced file; confirm the derived categories/groups in the code are traceable to that file's content rather than to `value_counts()` of a raw column. 3. Check the answer for evidence the specified method was applied (group boundaries/labels quoted from the spec) and that all required output files were written to the expected location. 4. If the spec was never loaded, or produced groups differ in number/boundaries from anything in the spec, flag it. - **Discriminator**: Fine if the script demonstrably reads the spec (or quotes its rules) and the resulting groups match it — even if the raw column happens to be pre-binned; a violation is using dataset-native or self-invented categories with no reference to the mandated definition. - **Consequence**: Aggregated counts/labels differ from the reference binning, so the saved figure and any numeric arrays mismatch the expected artifacts and all checks fail, despite the plot having correct title/axis labels.
task -- the prompt shows a literal output template (e.g., keys whose values are bracketed lists) and/or the grading setup expects a specific result file produced by the agent's code.17f8edabbb6b · mined from da-code dacode-di-text-002@s2### Answer not persisted to the expected artifact with the exact requested schema - **Applies when**: `task` -- the prompt shows a literal output template (e.g., keys whose values are bracketed lists) and/or the grading setup expects a specific result file produced by the agent's code. - **Pattern**: The attempt computes a plausible-looking value but only reports it in prose/chat, and/or reshapes the template (scalars instead of lists, renamed/extra/missing keys, added units or rounding not asked for), leaving no saved script or result file that a grader can read. - **Detection procedure**: 1. Re-read the task and write down the exact required deliverable: file name/location (if any), key names, and the value type shown in the template (list vs scalar, string vs number). 2. Search the scripts for the code that serializes the final answer (e.g., a dump/write to the named file); if no script or no write step exists, the deliverable is missing. 3. Diff the submitted object against the template key-by-key and type-by-type; confirm every key is present, spelled identically, and wrapped in the same container type. 4. Check that the reported number is the final requested quantity in the requested units/precision, not an intermediate or reformatted variant. - **Discriminator**: A real violation is a structural mismatch (no file written, scalar where a list is shown, renamed/absent keys); harmless look-alikes are cosmetic differences the template does not constrain, such as key ordering, whitespace, or a value that is genuinely scalar because the template shows a scalar. - **Consequence**: The grader finds the expected result file missing or fails schema/type comparison, marking the submission wrong even if the underlying computation was right.
task -- the script reads a raw table and immediately aggregates it (weighted sums, compounding, cumulative products) into an output series that is graded against exact expected values.isna().sum(), dtype checks, min/max or magnitude checks, row-count/date-continuity checks, handling of the first period. If aggregation happens with none of these, flag it.4442748969bb · mined from da-code dacode-dm-csv-050@s2### Skipping input-integrity validation before building a derived/cumulative series - **Applies when**: `task` -- the script reads a raw table and immediately aggregates it (weighted sums, compounding, cumulative products) into an output series that is graded against exact expected values. - **Pattern**: The attempt loads the file and jumps straight to arithmetic without checking for missing/blank cells, non-numeric dtypes, duplicate or unsorted keys, extra/renamed columns, or the scale/units of the values (e.g., fractions vs. percentages, levels vs. period-over-period changes). Any NaN, string column, off-by-100 scaling, or first-row convention silently propagates through the cumulative operation and corrupts every downstream row, and the attempt's "verification" script only re-derives its own numbers instead of testing them against an independent expectation. - **Detection procedure**: 1. Read the task/README for what the input columns are supposed to represent (units, scale, whether a base/first row exists) and what the output must contain (column names, ordering, index, rounding, row count). 2. Scan the scripts for explicit integrity checks before the aggregation: `isna().sum()`, dtype checks, min/max or magnitude checks, row-count/date-continuity checks, handling of the first period. If aggregation happens with none of these, flag it. 3. Check whether the "verification" step compares to anything independent (a known ground-truth magnitude, a hand-computed row, a second method) or merely recomputes the same formula. 4. Inspect the reported output: are the values in a plausible range and magnitude for the stated quantity, is the row count/first row consistent with the input period, and do the column names/order match the required format exactly? - **Discriminator**: A real violation is when nothing in the pipeline could have detected a NaN, a mis-scaled column, or a wrong first-row/base convention, and no external cross-check exists. It is *not* a violation if the script explicitly inspects/handles missing values and units (or documents that the data is clean after checking) and validates at least one output value against an independent reference, even if the code is otherwise terse. - **Consequence**: The saved file has the right shape and looks superficially reasonable, but every value is systematically off (shifted, scaled, or NaN-contaminated from the first defective row onward), so an exact/tolerance comparison against the expected file fails on all checks.
task -- the task names a specific statistical test/metric to run on a specified column or dataset, and the script contains a branch that resamples, truncates, or otherwise changes the input before computing it (often to dodge a library size limit or runtime cost).np.random.*, .sample(, slicing, head/tail, dropna-with-different-scope, or size-based if branches sitting between data loading and the mandated computation; check whether a seed is set and whether the same subset feeds every reported number.faa875efbf32 · mined from infiagent-dabench dabench-298@s2### Silently substituting a random subsample (or otherwise altered input) for the statistic the task specified - **Applies when**: `task` -- the task names a specific statistical test/metric to run on a specified column or dataset, and the script contains a branch that resamples, truncates, or otherwise changes the input before computing it (often to dodge a library size limit or runtime cost). - **Pattern**: The script computes the mandated test on a random subset (frequently with no fixed seed) or a differently-filtered set than the one used for the other reported statistics, then reports the resulting p-value/decision as if it came from the full specified data — making the answer non-reproducible and potentially opposite to the intended result. - **Detection procedure**: 1. Read the task and note exactly which data (which column, which rows, all of them?) the named test/statistic must be computed on, and any stated thresholds. 2. Scan the script for any `np.random.*`, `.sample(`, slicing, head/tail, dropna-with-different-scope, or size-based `if` branches sitting between data loading and the mandated computation; check whether a seed is set and whether the same subset feeds every reported number. 3. Check whether the script prints diagnostics that let one verify the input to the test (n used, missing count, min/max/unique values) and whether the script would raise/flag rather than silently deviate when the library's constraints are hit. 4. Compare the reported decision with the other reported numbers and with the data's nature (e.g., few distinct integer values, small n, mild moment values) — if the decision hinges on a resampled subset or contradicts the descriptive statistics, treat it as unverified. - **Discriminator**: A real violation is a *silent, unrequested* change of the analysis input (or a nondeterministic one) affecting the reported figure; it is fine if the task itself authorizes subsetting/filtering, or if the full specified data is used and any subsetting is only for auxiliary plots/exploration and is seeded and documented. - **Consequence**: The reported test decision (and its p-value) reflects a different, randomly varying sample than the ground-truth computation, so the categorical answer flips (e.g., "no" instead of "yes") and the grader scores 0 even when the accompanying descriptive statistics look plausible.
task -- the deliverable is a prediction file whose rows must correspond one-to-one, in order, with the rows of a provided evaluation input file.to_csv(..., index=False)) with the exact header, and that predictions were produced from the full, unfiltered, unshuffled input (no dropna, sample, head, sorting, or partial-batch loop that would change row count/order).len(output) == len(input), header matches, and no nulls/unexpected label values.7fd1355dd264 · mined from da-code dacode-ml-multi-011@s2### Prediction file not verified against the test set (row count / alignment / persistence) - **Applies when**: `task` -- the deliverable is a prediction file whose rows must correspond one-to-one, in order, with the rows of a provided evaluation input file. - **Pattern**: The agent pastes predictions into the chat answer instead of (or in addition to) writing the required file, and never checks that the written file has exactly as many rows as the input, in the same order, with the exact required column name and no extra index column. Truncated, reordered, deduplicated, or dropped-row outputs (e.g., after dropping missing/blank text) go unnoticed. - **Detection procedure**: 1. Read the task for the required output filename, column name(s), and the input file that defines the number and order of predictions. 2. In the scripts, confirm the file is actually written to the required path (`to_csv(..., index=False)`) with the exact header, and that predictions were produced from the full, unfiltered, unshuffled input (no `dropna`, `sample`, `head`, sorting, or partial-batch loop that would change row count/order). 3. Look for an explicit post-write sanity check: re-read the file and assert `len(output) == len(input)`, header matches, and no nulls/unexpected label values. 4. Inspect the final answer: if it consists of the predictions dumped as text (especially cut off mid-word) rather than a confirmation of a validated saved file, treat the artifact as unverified. - **Discriminator**: A real violation is missing/unwritten file, or absent verification of row count, order, header, and label vocabulary. It is fine if the agent writes the file and shows a check of shape/head/value counts, even if it also prints a preview of the predictions in the answer. - **Consequence**: The grader looks for the expected file and compares row-by-row; a missing, truncated, misaligned, or wrongly-headed file scores 0 regardless of model quality.
task -- the deliverable is a prediction/result file and the agent runs several successive scripts, each rewriting the same output file, with the last one being a simplified/faster variant.ae915c5e4760 · mined from da-code dacode-ml-competition-008@s2### Final artifact produced by an unvalidated "fallback" model that overwrites better-validated work - **Applies when**: `task` -- the deliverable is a prediction/result file and the agent runs several successive scripts, each rewriting the same output file, with the last one being a simplified/faster variant. - **Pattern**: Earlier scripts hold out data and report validation scores for stronger models, but the final script that actually writes the deliverable drops validation entirely, fits a much weaker/simpler estimator on all data, and silently overwrites the previous output; the answer reports only descriptive statistics of the predictions (mean/min/max) as if they were evidence of quality, with no held-out error estimate or comparison against the earlier candidates. - **Detection procedure**: 1. Read the task to identify the required deliverable file and the implied evaluation criterion (accuracy/error of the predictions, not merely file existence). 2. Identify which script last writes that file, and check whether it (a) computes any held-out or cross-validated score for the exact model/pipeline whose predictions are saved, and (b) is at least as strong as models validated in earlier scripts. 3. Check the reported answer for a quantitative generalization metric tied to the shipped predictions and an explicit reason for choosing that model over the alternatives. 4. Sanity-check the shipped predictions against the training target's distribution/range and the required id count, ordering, column names, and output path. - **Discriminator**: A real violation is when the shipped model was never scored on unseen data, or was demonstrably weaker than an already-validated alternative and the switch is justified only by speed/convenience. It is fine if a simpler model is shipped after being compared on the same validation protocol and shown to be competitive, or if refitting a validated configuration on the full data without re-scoring. - **Consequence**: The submission file is well-formed but its predictive quality falls below the grader's threshold (poor R²/high error versus the hidden labels), so the deliverable is marked wrong despite the run "completing successfully."
task -- the task asks for the top N entities by some aggregate statistic, and the prompt/README supplies a definition, eligibility threshold, or tie-breaking/ordering convention for that statistic.sort_values(..., ascending=False) on the aggregate with a defined tie-breaker before head(N).938bf48c33a6 · mined from da-code dacode-dm-csv-009@s2### Ranked "top-N" lists produced without applying the stated qualification rule and rank ordering - **Applies when**: `task` -- the task asks for the top N entities by some aggregate statistic, and the prompt/README supplies a definition, eligibility threshold, or tie-breaking/ordering convention for that statistic. - **Pattern**: The attempt groups rows and sorts by a raw aggregate without honoring the stated qualification rule (e.g., minimum number of underlying records per entity, deduplication of repeated items, numeric parsing/cleaning of the value column), and/or emits the N names in an order that is not the ranking order (alphabetical or arbitrary), so ties at the ceiling value are resolved incorrectly. - **Detection procedure**: 1. Read the task/README and write down every constraint attached to the ranking: how the aggregate is defined, any eligibility filter, how ties are broken, and whether row 0 must be the top-ranked entity. 2. In the scripts, locate the groupby/aggregation and check for (a) cleaning/casting of the ranked column, (b) an explicit filter implementing the eligibility rule, (c) `sort_values(..., ascending=False)` on the aggregate with a defined tie-breaker before `head(N)`. 3. Inspect the answer file: check whether any column's entries are in alphabetical order or whether the top entries look like one-off/low-volume entities (a sign the eligibility filter and rank ordering were skipped). 4. Confirm the emitted aggregate values were sanity-checked (range, count of rows per entity, N rows exactly, column names/ordering matching the sample format). - **Discriminator**: A real violation is when no code implements the stated threshold/ordering, or the output order cannot be reproduced by sorting on the aggregate; a look-alike that is fine is a list that merely *happens* to be near-alphabetical because genuine ties were broken by a documented, implemented rule. - **Consequence**: The exact-match check on the saved file fails because the entity set and/or their row positions differ from the reference ranking.
task -- the task refers to a supplied dataset (with a README/spec) and the agent's scripts include a step that generates, simulates, or hard-codes the input data.np.random.*, range(...) fillers, manual dicts/lists written to the input path, seed) or writes to the same path later read as input. 3. Check whether the script instead locates and loads the real provided file (and would fail loudly if absent) and whether all required outputs are written. 4. Cross-check the answer's reported counts/ranges against the dataset description (row counts, units, realistic value ranges) and against the mentioned spec/config keys.bbec7f72c429 · mined from da-code dacode-plot-bar-007@s2### Fabricating input data instead of using the provided dataset - **Applies when**: `task` -- the task refers to a supplied dataset (with a README/spec) and the agent's scripts include a step that generates, simulates, or hard-codes the input data. - **Pattern**: The agent cannot find or fails to load the real input file, so it synthesizes a random/placeholder table with invented columns and value ranges, then runs the full analysis on that fake data and reports the resulting numbers as if they came from the real source. Related symptom: only a subset of required output artifacts is produced, and stated spec files are only partially honored. - **Detection procedure**: 1. Read the task and note which input files and which output artifacts (files, formats, names) are expected. 2. Scan every script for data creation calls (`np.random.*`, `range(...)` fillers, manual dicts/lists written to the input path, `seed`) or writes to the same path later read as input. 3. Check whether the script instead locates and loads the real provided file (and would fail loudly if absent) and whether all required outputs are written. 4. Cross-check the answer's reported counts/ranges against the dataset description (row counts, units, realistic value ranges) and against the mentioned spec/config keys. - **Discriminator**: A real violation is when the analytical result reported comes from data the agent itself invented, or when the real file was never read. Legitimate look-alikes: creating tiny synthetic fixtures for unit-testing a plotting/utility function while the reported result still comes from the real dataset, or augmenting real data in a documented, task-sanctioned way. - **Consequence**: All value-based and file-based checks fail: derived arrays/JSON summaries and the figure encode arbitrary random counts (and required artifacts may be missing entirely), so the grader reports 0 of the expected outputs correct.
task -- a task prescribes a specific answer template (e.g. @name[value]) for a statistic computed on a filtered subset, and the filter may match zero rows.c6a27ce88566 · mined from infiagent-dabench dabench-554@s2### Empty-subset result reported as prose instead of the required answer format
- **Applies when**: `task` -- a task prescribes a specific answer template (e.g. `@name[value]`) for a statistic computed on a filtered subset, and the filter may match zero rows.
- **Pattern**: The agent discovers the filter yields no rows (or an undefined statistic) and replaces the required formatted answer with an explanatory sentence ("no data available", "filter value does not exist"), instead of emitting the statistic's defined degenerate value (NaN/empty) inside the requested format.
- **Detection procedure**:
1. Read the task and note the exact required output token/format and rounding rules.
2. Read the scripts for a branch that short-circuits when the filtered frame is empty (or when the aggregate returns NaN) and prints a message rather than the formatted value; check whether the filter/dtype (e.g. string vs numeric key) was verified before concluding the subset is empty.
3. Compare the submitted answer against the required template: does it literally contain the token and a value slot?
4. If it does not, flag it — an empty subset is a valid result to report, not a reason to abandon the format.
- **Discriminator**: A real violation is any answer that abandons the template; it is fine to report an unusual value (NaN, 0, empty) *inside* the template, and it is also fine to note the emptiness as extra commentary alongside a properly formatted answer. Also not a violation if the emptiness stemmed from a genuine miscoded filter that, once fixed, yields data — that is a different (filtering) error.
- **Consequence**: The grader's field-by-field check finds no parsable value for the requested key, so the expected result (including NaN) is scored WRONG/MISSING and the task fails at 0/1.task -- the task asks for specific records (top/bottom-N names, IDs, categories) or statistics to be extracted from a supplied data file after a prescribed preprocessing step.6f4b40ed7ee2 · mined from da-code dacode-di-text-003@s2### Unverifiable answer: entity names/values not traced back to the provided data via runnable code - **Applies when**: `task` -- the task asks for specific records (top/bottom-N names, IDs, categories) or statistics to be extracted from a supplied data file after a prescribed preprocessing step. - **Pattern**: The submission presents a plausible-looking list that reflects general/world knowledge or an unsaved ad-hoc computation, with no script that (a) loads the given file, (b) applies the stated preprocessing (e.g., the specified imputation), (c) sorts by the requested field in the stated direction, and (d) writes the answer to the required output file; entity labels are also not copied verbatim from the data's key column. - **Detection procedure**: 1. Read the task and note the required output artifact/format, the mandated preprocessing, and the ordering/sorting constraint for *every* requested list. 2. Look for a script that reads the provided file and prints/dumps exactly the reported values; if no script exists, or the script's output was never shown to match the submitted answer, treat the answer as unverified. 3. Cross-check each reported label against the data's identifier column spelling/format (e.g., official vs. colloquial names, punctuation, abbreviations) and check that each requested list is ordered as instructed, not by a default or intuitive order. 4. Sanity-check the numbers behind the picks (count = N, values within plausible range, no rows dropped instead of imputed) — if the underlying values are not reported at all, that is itself a flag. - **Discriminator**: A real violation is an answer whose values cannot be reproduced from the file by any shown code, or whose labels differ from the dataset's own strings/ordering rule; a look-alike that is fine is an answer that happens to match common knowledge *but* is accompanied by a script whose printed output and written result file match it exactly and whose labels are taken verbatim from the data. - **Consequence**: The grader's exact comparison against the expected result file fails on missing/misnamed entities, wrong list order, or a missing output file, scoring 0 even though the list looks superficially reasonable.
task -- The task specifies an exact answer template with named tags, delimiters, and example value formatting, and expects the values to be produced by saved, re-runnable analysis code.3722a8c2393a · mined from infiagent-dabench dabench-550@s2### Answer string not emitted in the literal template (verbatim tokens, quoting, order) with no reproducible script backing it - **Applies when**: `task` -- The task specifies an exact answer template with named tags, delimiters, and example value formatting, and expects the values to be produced by saved, re-runnable analysis code. - **Pattern**: The agent computes conceptually correct values but serializes them in a variant of the requested template — dropping the quotation marks/units shown in the example, renaming or reordering tags, changing separators, adding prose around the tags — and/or leaves no script that prints the final answer string, so nothing ever validated the emitted text against the template. A literal string-matching grader then fails every field even though the underlying analysis was right. - **Detection procedure**: 1. Copy the answer template exactly as given in the task, including every quote, bracket, tag name, separator, and the formatting shown in any example value. 2. Check the scripts: is there a step that constructs and prints the final answer string, so the submitted text is a program output rather than hand-typed? If no script is saved at all, the answer is unverifiable by construction — flag it. 3. Diff the submitted answer character-by-character against the template: tag names and order, presence/absence of quotes around each value, capitalization and spelling of the allowed value vocabulary, and the exact form of any range/unit string. 4. Confirm each value is one of the permitted options and is the final requested quantity, not an intermediate. - **Discriminator**: A real violation is any character-level deviation from the specified template (missing quotes, extra text, altered tag names/order), or an answer with no code that produced it; a look-alike that is fine is a template where the task itself shows the value unquoted or explicitly allows whitespace variation, and the answer matches that shown form exactly. - **Consequence**: A regex/exact-match grader reports every field as WRONG/MISSING even though the submitted values are semantically identical to ground truth, yielding a 0/N score with no partial credit.
task -- the prompt asks for specific entities, groupings, or measures (e.g., "top N of X by Y, broken down by stage Z") and the scripts must locate those fields in the provided data files.7d4eac7f01cd · mined from da-code dacode-plot-scatter-002@s2### Substituting proxy data/entities for the ones the task names - **Applies when**: `task` -- the prompt asks for specific entities, groupings, or measures (e.g., "top N of X by Y, broken down by stage Z") and the scripts must locate those fields in the provided data files. - **Pattern**: The agent cannot find the requested fields in the files it opened, so it silently swaps in a different data source and re-interprets each requested concept as a loose "analogy" (different grouping key, different ranking measure, different segment definitions), then declares success instead of resolving the mismatch. - **Detection procedure**: 1. List from the task the exact entity to rank, the ranking measure, and the segments/values to plot. 2. In the scripts, identify which file and which columns supply each of those three items. 3. Flag if any is drawn from a different domain or is described in the answer as an "analogy"/"proxy"/"equivalent", or if the reported category labels are not instances of the requested entity type. 4. Check the answer's reported units/axis labels match the requested measure (e.g., a duration in days, not a count or probability). - **Discriminator**: A real violation is using a different variable or dataset than the task names because the requested one was not found; acceptable look-alikes are cases where the requested concept genuinely exists under a differently-spelled column name and the mapping is stated and verifiable (same semantics, same units). - **Consequence**: Every derived artifact (figure, saved arrays, config-driven outputs) encodes the wrong categories and wrong quantities, so all value- and label-level checks fail even though a chart of the right visual type was produced.
task -- the prompt specifies an exact answer string template (a tag, an = or : separator, and specific wrapping brackets/braces) for reporting one or more computed values.@/=/: separators, and each opening/closing bracket or brace in order.dict/list repr.=, wrong bracket type, missing wrapper, renamed/reordered keys); harmless look-alikes are cosmetic differences the template does not constrain, such as inner whitespace, quote style, or int-vs-float rendering of the same value.f83a462bf083 · mined from infiagent-dabench dabench-451@s2### Answer-format template not reproduced literally (delimiters/separators dropped) - **Applies when**: `task` -- the prompt specifies an exact answer string template (a tag, an `=` or `:` separator, and specific wrapping brackets/braces) for reporting one or more computed values. - **Pattern**: The attempt computes the values correctly but emits them in a paraphrased wrapper — omitting or substituting the required leading/trailing delimiters, the key-name-to-value separator, or the bracket nesting — so an exact-match grader fails even though the analysis is right. - **Detection procedure**: 1. Copy the answer template exactly as given in the task and mark every literal token: tag name, `@`/`=`/`:` separators, and each opening/closing bracket or brace in order. 2. Read the scripts (or final message construction) and check whether the output string is built from that literal template — e.g. an f-string/print that hard-codes the tag and delimiters — rather than from a bare `dict`/`list` repr. 3. Compare the submitted answer token-by-token against the template: same tag, same separator, same bracket sequence and nesting, same key spelling/order. 4. Flag if any literal token is missing, added, or swapped, even when the numeric values look plausible. - **Discriminator**: A real violation is a structural mismatch in the required literal tokens (missing `=`, wrong bracket type, missing wrapper, renamed/reordered keys); harmless look-alikes are cosmetic differences the template does not constrain, such as inner whitespace, quote style, or int-vs-float rendering of the same value. - **Consequence**: The grader reports the expected variable as WRONG/MISSING and scores 0 despite numerically correct values, because it cannot parse the submitted string against the required pattern.
task -- the deliverable is a prediction/output file with a specified column name and one row per record of a held-out input file.5d3480fe4243 · mined from da-code dacode-ml-regression-004@s2### Unverified prediction file (no reproducible script, no shape/format sanity check) - **Applies when**: `task` -- the deliverable is a prediction/output file with a specified column name and one row per record of a held-out input file. - **Pattern**: The attempt produces the output file without any saved, runnable script that reads the held-out input, fits on the training portion, and writes predictions — and never checks that the written file has exactly the required column name, the same number of rows as the held-out input, the original row order, and values in a plausible range/dtype for the target. - **Detection procedure**: 1. Read the task statement and note the exact required file name, column name(s), row count (from the held-out input), and any ordering/rounding/units constraints. 2. Look for a script that is complete end-to-end (load train + held-out data → preprocess → fit → predict → write file); if scripts are absent or only partially cover this path, the result cannot be reproduced or audited and must be treated as unverified. 3. Inspect the produced file's header and shape: does it contain the requested column spelled exactly as asked (no extra index column, no renamed/extra columns), and does its row count equal the held-out input's row count? 4. Check value sanity: dtype numeric (or as required), no NaNs/empty cells, values inside the target's observed range, and non-constant/non-degenerate predictions aligned to the input row order. - **Discriminator**: A real violation is missing/unreproducible code *or* a file whose header, row count, ordering, or value range does not match the specification. A look-alike that is fine: a script that differs stylistically or uses a simple baseline model, but demonstrably writes the exact requested column, one row per held-out record in input order, with valid values. - **Consequence**: The grader compares the submitted file against the expected schema/row alignment and marks it WRONG/MISSING (0 checks passed) even though a file with the right name exists.
task -- the task asks for an unsupervised grouping (k-means/hierarchical/DBSCAN-style) of records whose numeric columns have wildly different units and ranges, and the deliverable is a label file.a56e2e2c47e9 · mined from da-code dacode-ml-cluster-013@s2### Feature scaling omitted before distance-based clustering, yielding degenerate outlier clusters - **Applies when**: `task` -- the task asks for an unsupervised grouping (k-means/hierarchical/DBSCAN-style) of records whose numeric columns have wildly different units and ranges, and the deliverable is a label file. - **Pattern**: The attempt feeds raw, unstandardized columns straight into a distance-based algorithm (and/or picks *k* by a score computed on those raw features), so one or two large-magnitude columns dominate the distance metric; the resulting partition contains singleton or near-singleton clusters that merely isolate extreme values, while the bulk of the records collapse into one or two huge groups. The attempt reports this as the "optimal" solution without a sanity check on cluster sizes or on whether the grouping reflects all features. - **Detection procedure**: 1. In the task/README, note the feature columns and their plausible scales (percentages, per-capita monetary values, rates) — flag if magnitudes differ by orders of magnitude. 2. In the scripts, check whether a scaler/normalizer (or a distance metric that is scale-invariant) is applied to the feature matrix *before* fitting and before any cluster-count selection metric; also check that model selection uses the same transformed space. 3. In the reported answer, inspect the cluster size distribution and cluster descriptions: singleton/2–3-member clusters described as "outliers" or clusters characterized by a single dominant variable are strong evidence of unscaled distances. 4. Confirm the saved label file was actually verified (row count, column names, label range) rather than only asserted in prose. - **Discriminator**: A real violation is scaling never applied (or applied only after clustering / only for plots) together with a lopsided partition dominated by high-variance columns. It is *not* a violation if the script standardizes (or the algorithm/metric is scale-free) and small clusters are then justified by inspection — genuinely extreme records can legitimately form small clusters in a properly scaled space. - **Consequence**: The submitted label column encodes an outlier split rather than the intended socio-economic-style grouping, so the expected output file comparison (cluster structure/agreement with reference labels) fails even though the file format looks plausible.
task -- The question restricts the statistic to a specific slice (a given year, group, region, or subset of columns/rows) and the scripts compute an aggregate statistic per entity.e8a3e51f01ca · mined from infiagent-dabench dabench-252@s2### Stated scope/filter in the task is not reflected in the computed statistic (wrong axis/subset) - **Applies when**: `task` -- The question restricts the statistic to a specific slice (a given year, group, region, or subset of columns/rows) and the scripts compute an aggregate statistic per entity. - **Pattern**: The script ignores the stated restriction and aggregates over the full set of columns/rows (e.g., every period rather than the specified one), or aggregates along the wrong axis, so the reported ranking answers a different question than the one asked. Extra "verification" code re-checks the same wrong computation, giving false confidence. - **Detection procedure**: 1. From the task text, list every explicit scope word (year, subset, grouping, units, estimator/definition flag) that must appear as a filter or parameter in the code. 2. Read the script and locate where each scope word is applied: is there a selection of the specific column/rows before the statistic, and is the statistic taken along the axis implied by the question? 3. If a scope word has no corresponding filter (e.g., all period columns are passed into the per-row statistic, or all rows into a per-column statistic), flag it; also check whether the resulting sample size/shape is what the stated slice would produce. 4. Confirm the sanity check: does the intermediate printout show data only from the requested slice, and is the final reported quantity the one named in the question (not an intermediate or differently-scoped one)? - **Discriminator**: A real violation is when the requested slice is never selected anywhere in the pipeline (or is selected but then discarded/overwritten). It is *not* a violation when the slice is implicit because the loaded file/frame is already restricted to it, or when the statistic legitimately requires the broader sample and the slice only defines the grouping key — in those cases the code should still show an explicit filter or a comment plus a shape/count check matching the slice. - **Consequence**: The ranking/argmax is computed over the wrong sample, so the reported entity differs from the ground truth and the grader marks the single expected value as WRONG, even though the estimator flag and output format are correct.
task -- The task asks to identify a specific record/key (a date, ID, category) from row-level data and the answer template shows a format string that is coarser or ambiguous relative to the data's granularity.strftime to a shorter pattern, rounding, taking a prefix/substring) that discards the identifying detail, so the reported key no longer uniquely designates the row actually used for the downstream computation.835c6e408b56 · mined from infiagent-dabench dabench-572@s2### Reformatting an identifier to a coarser granularity than the data (losing precision in the reported key) - **Applies when**: `task` -- The task asks to identify a specific record/key (a date, ID, category) from row-level data and the answer template shows a format string that is coarser or ambiguous relative to the data's granularity. - **Pattern**: The script correctly locates the extremum/target row, but then applies a formatting/truncation step (e.g., `strftime` to a shorter pattern, rounding, taking a prefix/substring) that discards the identifying detail, so the reported key no longer uniquely designates the row actually used for the downstream computation. - **Detection procedure**: 1. Read the task: determine the granularity of the entity being identified (one row / one timestamp / one ID) and note whether the downstream calculation depends on that exact row. 2. Read the script: find where the identified key is converted for output; check whether the emitted string contains strictly less information than the key stored in the data. 3. Compare the emitted key against the value used internally for the dependent computation — if the dependent computation uses the full-precision key but the answer reports a truncated one, flag it; prefer reporting the key exactly as it appears in the source (or, if the template is genuinely ambiguous, report the full-precision value rather than truncating). 4. Check for collisions: would the truncated key match multiple rows in the data? If yes, it cannot be the intended unique answer. - **Discriminator**: A real violation is when truncation destroys uniqueness or drops detail present in the source key while the rest of the answer depends on that detail. It is fine when the task explicitly aggregates at the coarser level (e.g., the maximum is computed over monthly aggregates) or when the source key itself has only that granularity. - **Consequence**: The numeric part of the answer can be correct while the identifier check fails, giving a partial-credit/incorrect verdict (e.g., 1 of 2 checks passed).
task -- the task explicitly instructs that results be saved to a named output file (e.g., result.csv) with the computed statistic(s).to_csv/open(...).write/etc.) targeting that exact filename; if no scripts exist at all, the deliverable is unverifiable by definition.30ce0fb02dee · mined from da-code dacode-data-sa-043@s2### Missing/unverifiable output artifact — answer reported inline instead of written to the required file - **Applies when**: `task` -- the task explicitly instructs that results be saved to a named output file (e.g., `result.csv`) with the computed statistic(s). - **Pattern**: The attempt computes a single number and reports it in prose/chat, but no saved script or code path demonstrably writes the named file (with the expected column/row layout, precision, and the exact requested quantity); the artifact is absent, empty, or contains an intermediate value rather than the requested one. - **Detection procedure**: 1. Read the task and list every required deliverable: file name, location, and what it must contain (which statistic, how labelled, any rounding/format constraints). 2. Search the submitted scripts for an explicit write call (`to_csv`/`open(...).write`/etc.) targeting that exact filename; if no scripts exist at all, the deliverable is unverifiable by definition. 3. Trace which variable is written and confirm it is the final requested quantity computed on the specified subset/aggregation (not an intermediate, unfiltered, or differently-aggregated value), and that the file layout would parse as a table. 4. Cross-check the reported number against the file-writing code: if the only evidence is a bare float in the answer text, flag it. - **Discriminator**: A real violation is when no reproducible code writes the named artifact, or the artifact holds a different quantity/format than requested; a look-alike that is fine is a script that writes the correct file under the right name and merely also echoes the value in the answer text. - **Consequence**: The grader looks for the named result file and finds it missing or containing a mismatched value/format, scoring 0 regardless of whether the number quoted in the chat happens to be close.
task -- the deliverable is a per-row prediction file for a supplied test set, especially with a stated cost/recall asymmetry between error types.len(predictions) == len(test) and value_counts(); absence of any saved script is itself a failure of verifiability.cfb02ca4e7e0 · mined from da-code dacode-ml-binary-013@s2### Missing reproducible pipeline + unvalidated prediction file (shape / class-rate sanity check) - **Applies when**: `task` -- the deliverable is a per-row prediction file for a supplied test set, especially with a stated cost/recall asymmetry between error types. - **Pattern**: The agent produces the output file with no saved, re-runnable script that reads the test file, aligns rows, and writes the required column, and never checks that the emitted file has exactly one prediction per test row, in the original test-row order, with the exact header/values shown in the sample; it also never compares the predicted positive rate against the training base rate or the asymmetric-cost objective, so a near-all-negative (default-threshold, imbalance-collapsed) prediction is submitted uninspected. - **Detection procedure**: 1. From the task, note the required output format (header text, single column, allowed values) and the exact number of rows in the test file and in the provided sample output. 2. In the scripts, look for code that (a) loads the test set, (b) predicts, and (c) writes the file preserving test-row order, plus an explicit assertion/print of `len(predictions) == len(test)` and `value_counts()`; absence of any saved script is itself a failure of verifiability. 3. Compare the answer file: count rows (excluding header), confirm header string and value domain match the sample, and compute the fraction of positive predictions. 4. Compare that fraction to the minority-class prevalence in training and to the stated cost preference (e.g., missed positives more costly than false alarms); flag if it is far below prevalence or if any threshold/class-weight decision was never justified or validated on a held-out split. - **Discriminator**: A genuine violation is a file whose row count/format cannot be shown to match the test set, or a positive rate materially below the training prevalence with no validation evidence (no held-out recall/F1/cost comparison, no threshold tuning). It is *not* a violation if the script asserts shape and ordering, the sparse positive rate is backed by held-out metrics on an imbalanced problem, and the format matches the sample exactly. - **Consequence**: The grader compares the submitted file row-by-row against the reference; a wrong row count, wrong header/order, or a degenerate mostly-negative prediction fails the file check outright (0/1) even though the answer "looks like" valid predictions.