Build a leakage-safe benchmark
Design a synthetic data utility benchmark with an untouched real test set.
Design a benchmark to test whether synthetic fraud transactions can train a useful model without leaking test information. Split by leakage unit first; fit preprocessing and synthesis only on train; evaluate once on untouched real test data. The expensive mistake is fitting encoders, rare-category grouping, or the synthesizer on the full dataset before splitting. That can leak test-set structure even without exact row copies. Step 1 Split real data by merchant_id into train, validation, and test before any preprocessing. Merchant-level grouping prevents the model from seeing synthetic versions of merchants later used for test. Step 2 Fit imputation, category bucketing,…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in