The migration nobody noticed
The best outcome for an infrastructure migration is that no one can tell it happened. No incident channel, no apology post, no spike in the dashboard anyone screenshots. Just a slow handover and then a Tuesday where you delete the old thing.
Here's the shape of one that went that way: replacing the push delivery layer under a notification system serving roughly ninety million buyers.
The problem with "just switch it"
A notification pipeline has two failure modes, and they pull in opposite directions:
- Drop a message. Someone doesn't learn their order shipped.
- Send it twice. Someone gets woken up twice, which is worse than either of us wants to admit.
A cutover optimises for one of these at the expense of the other. Flip the switch and anything in flight is dropped. Dual-write naively and anything in flight goes out twice.
So we didn't cut over. We ran both for six weeks.
Shadow writes first
Phase one sent every notification to the new provider with delivery disabled. Nothing reached a device. What we got was a log we could compare.
# Shadow comparison, day 9
events_in 4,218,904
old_path_delivered 4,218,671 (99.994%)
new_path_would_send 4,217,902 (99.976%)
divergent 769
└─ template_missing 712 ← our bug, not theirs
└─ locale_fallback 48
└─ unexplained 9Those nine unexplained divergences took eleven days to chase down. They were a race in how we resolved a user's device token when they'd reinstalled the app during the window. We'd have shipped that bug to ninety million people.
Idempotency is the whole trick
Once we started real dual delivery, the only thing standing between us and double-sends was a key:
// Stable across retries, across both providers, and across a redeploy.
// Deliberately NOT time-based: a retry 40s later must collide.
func dedupeKey(userID, templateID string, triggeredAt time.Time) string {
bucket := triggeredAt.Truncate(5 * time.Minute).Unix()
sum := sha256.Sum256([]byte(fmt.Sprintf("%s:%s:%d", userID, templateID, bucket)))
return hex.EncodeToString(sum[:16])
}The five-minute bucket is the compromise. Too narrow and a legitimate retry after a provider timeout looks like a new message. Too wide and a user who genuinely triggers the same notification twice gets one. Five minutes was right for our templates; it would be wrong for a chat app.
Ramping
We moved in percentages, by template, lowest-stakes first:
- Marketing digests at 1% — the ones where a miss costs nothing
- Same templates to 50%, then 100%
- Transactional at 1% — order shipped, payment confirmed
- Transactional to 100%, two weeks later than planned
Step four slipping was the right call and it was not a popular one. We'd found a latency regression on a single template and the pressure to call it "good enough" was real. It wasn't good enough. It took six more days.
Deleting the old path
Six weeks in, the old provider was serving zero traffic and we turned it off. Then we left the code in place for another month, because a rollback you can't execute isn't a rollback.
The delete PR was 3,000 lines red and nobody reviewed it carefully. That's fine. By then there was nothing to review.
What I'd tell myself
Make the comparison deeper than feels necessary. Decide the dedupe window on purpose. And protect the slow ramp from the people — including you — who want to be done.