Skip to content

Clinic 22

Data Cleaning Choice

A missing ticket-age field may encode never having a ticket. Represent it explicitly and compare policies on the same validation data.

Situation

A churn model includes days_since_last_ticket. For 27% of customers it is absent because they have never opened a ticket. Churn prevalence is 4% among these customers and 15% among those with a ticket. The overall prevalence is 0.27 × 0.04 + 0.73 × 0.15 = 0.1203.

Here, ticket age is structurally undefined for never-ticketed customers. Verify this with data owners; an ingestion failure could also produce a missing value and would need separate handling. A has_ever_ticketed indicator plus an explicitly imputed age can represent this distinction.

Artifact Packet

These AUC values are illustrative measurements on the same complete held-out population. Each policy is fitted using training data only. For drop_rows, evaluate the fitted model on all validation customers with a training-fitted fallback for missing age; do not silently remove the difficult validation rows. The runner calculates prevalence from the stated group rates; it does not fit classifiers.

policy validation AUC training churn prevalence
drop_column 0.701 0.1203
drop_rows 0.724 0.1500
mean_impute_only 0.735 0.1203
value_plus_missing_flag 0.764 0.1203

Dropping a column or imputing a feature while retaining all rows and labels leaves label prevalence unchanged. Dropping missing rows changes the training population and raises prevalence to 0.15 here.

Decision Prompt

  1. What does a missing days-since-ticket value mean operationally?
  2. Which policies change label prevalence, and why?
  3. Which policy would you select from this packet?
  4. Can label association or an AUC gain establish MNAR?

Strong Reasoning Looks Like

  • inspect the reason for missingness and preserve its operational meaning
  • fit imputers and encoders inside each training fold
  • compare candidates on the same validation rows, including the never-ticketed slice
  • distinguish predictive usefulness from identifying a missing-data mechanism

MCAR means missingness is independent of observed and unobserved data. MAR allows dependence on observed data, with no remaining dependence on missing values after conditioning on observed information. MNAR violates that conditional independence. Association between missingness and an observed label, or an AUC increase after adding a flag, does not identify MNAR. See Seaman et al. on missing-data definitions.

Run The Clinic In Browser

The runner prints this fixed illustrative packet and recalculates any derived columns. It does not run a new training experiment. Edit its PACKET values to explore the decision.

Reference Reveal

Open after writing your note Select **`value_plus_missing_flag`** from this packet because it has the strongest held-out AUC. Check calibration and slice performance before release; the packet contains no calibration measurement. The result is conditional on these measurements, not a universal promise that an indicator improves AUC by a fixed amount. Retain the flag to distinguish a real age from the imputed placeholder. For example, a mean age with a missingness indicator is a reasonable candidate; fit the mean on training rows with known age. If missingness has several causes, encode the causes where reliable and available at inference.

What To Do Next

  1. open Tabular Feature Engineering for the missingness-as-feature pattern in depth
  2. open Leakage Patterns for prior shift and filter-based leakage
  3. open Feature Selection Or Regularize — the adjacent clinic on what to do after cleaning is handled
  4. run the P(y|missing) diagnostic on every column in your dataset with ≥ 5% missing; most cleaning decisions come straight out of that table