Find the quasi-identifier trap
Apply a quasi-identifier review to reduce re-identification risk in a synthetic data specification.
Review a synthetic HR dataset for an internal performance dashboard demo. The source population is 480 employees, and the proposed fields are fake_name, exact_age, office_city, job_title, hire_month, salary_band, and performance_rating. Separate direct identifiers, quasi-identifiers, sensitive fields, and task utility before deciding what to keep. The novice move is to remove fake_name risk and keep the rest because the rows are synthetic. That ignores the identifying power of rare combinations. Step 1 Name the task: dashboard demo of filters and review-state UI, not compensation analysis. The task decides which fields need fidelity. Salary and exact identity details do not support this…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in