STAT 841 / CM 763 - Assignment 1 - Data package
Fall 2026

Extract all four CSV files before starting Question 6. Keep their filenames
unchanged and place them in the directory used by your Q6 program. Follow
the assignment PDF for all modeling, evaluation, and submission requirements.

FILES

A1_train.csv
  600 labeled observations. Row y contains labels 0 or 1; rows x1 through
  x20 contain the original features. Observation IDs are obs001 to obs600.

A1_test.csv
  300 unlabeled observations with the same 20 feature rows.
  Observation IDs are test001 to test300.

A1_noise_train.csv
  100 additional noise-feature rows, z1 through z100, for the 600 labeled
  training observations. These are used only in Question 6(c).

A1_resampling.csv
  The fold row supplies the fixed five-fold assignment for Question 6(b).
  The early_stop row supplies fit/validation membership for Question 6(c).

FORMAT AND USE

Observations are columns. The first row contains observation IDs, and the
first column contains row names. Match columns by their observation IDs.

Use the supplied folds and fit/validation split as instructed. Fit all
preprocessing using the relevant training subset only. Do not use test
features to fit preprocessing or select models.

Question 3 uses simulated data that you generate yourself, as specified in
the assignment. The four supplied CSVs are for Question 6.
