Worked Walkthrough: Dedupe without losing provenance
Build a dedupe key and provenance record for repeated scraped observations.
Deduplicate job postings that appear through multiple category URLs without merging distinct roles. Define the observation key, normalize only the fields used in that key, and preserve every source appearance as provenance. Dedupe by title alone feels simple but merges different roles with the same title and hides where each listing appeared. Name the observation One row in the canonical table represents one job posting, not one page appearance. This decides what counts as duplicate. Page appearances are evidence; the job posting is the observation. Build the key Use employer_id, normalized_title, location, and apply_url when no stable job_id exists. A…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in