Make every retry land on the same partition
Design retryable pipeline tasks around deterministic inputs, partitioned outputs, and idempotent writes.
The move: make retries deterministic. A retryable task needs a stable input and a stable output. That means the task should know exactly which slice of data it owns: a partition, interval, file batch, source offset range, or snapshot ID. It should not discover the slice by asking for "latest" during execution. The write side matters just as much. Blind inserts create duplicates when a retry happens after a partial success. Safer patterns include replace-partition, merge on a natural key, write-then-atomic-swap, or commit only after validation passes. The exact method depends on the warehouse and table format, but the principle…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in