Default Settings Are Production: Lessons from a Relay Rollout
Our team wanted to try our new WebSocket relay in production. The quickest way was to point the default transport settings at it. Within a day, users at several customers couldn’t use their live sessions, and it took a while until the last customer was back to normal.
I wasn’t the one who made the change, but I followed the whole thing closely. These are my notes: what happened, why it was hard to see, and what I’d do differently.
A bit of context
One of our products is a real-time app: a host runs live sessions with a group of participants, and all their clients talk through a WebSocket relay. If the relay can’t connect, nothing works.
The relay address and related connection settings live in a shared settings preset. We can assign a preset per customer, and there’s a default one for everyone else. Most customers use the default, so changing it changes everyone at once. The smallest scope we had was a whole customer.
We had rewritten the relay, and it ran on a new hostname under a brand-new domain that customers had never seen.
How it unfolded
The test setting went into the default preset. Every customer on the default picked it up.
The first report came in on day one: hosts got an error when starting a session. When I saw it, I reverted the default settings to the old relay. We reproduced it and found a real bug. We fixed it and deployed it the same day, and it felt like we were done. But it was a different bug, and it had masked the real culprit. Later that day, the default was switched back to the new relay.
That afternoon another customer reported a blank screen on session start, on the fixed release. The next day more came in: every participant device was stuck waiting to connect. Their IT teams all said the same thing: they hadn’t changed anything.
I always say I like days when I don’t need coffee to wake up. The next few days weren’t those days.
It took until day four to see that every reporting customer used the same preset: the default, pointing at the new relay. It didn’t help that it was a quiet day with fewer sessions to look at. Support asked customers whether they blocked the new relay, and allowing a single URL wasn’t always enough.
On day seven we gave the new relay a DNS record under our existing domain, the one customers already allowed. On day eight the root cause was clear: some customers blocked the new domain entirely. We moved affected customers back to the old settings, added a fallback relay URL, and told customers which hostnames to allow, and not to remove the old ones yet.
Eventually the last customer confirmed everything worked.
They had been right all along. They hadn’t changed anything. We had.
What I took from it
A default is production. The default preset looked like a convenient place to test. But a setting almost everyone inherits is the widest blast radius you have. There was no smaller scope, so “testing” meant “everyone”.
Your domain is part of your API. Plenty of customer networks run strict allowlists. A hostname on a domain nobody has allowed yet is blocked by default. The protocol was identical, but for those customers it was a breaking change.
Don’t close an incident because the first fix shipped. An unrelated bug showed similar symptoms on the same day. Fixing it made us believe the problem was gone while the real one was still running. Now I’d keep it open until the reports actually stop.
Silent failures cost days. A blocked WebSocket looked like a device waiting forever, not like “can’t reach the server”. Nobody, not the host and not the IT team, knew where to look. A clear message and a quick diagnostic check would have saved a lot of back-and-forth.
The fix I’d like to see: settings per group
Rolling back fixed the incident, not the reason it could happen. What I’d like to see is a transport preset per group, which only internal staff can set:
- Group preset
- Customer preset
- Default preset
Customers never see it. It would let us test new infrastructure on our own test groups first, and then on one group a friendly customer agrees to, instead of on everyone.
The preset travels with each session, and the server picks it once when the session starts, so a running session is never switched halfway through.
If I could do it again
- Treat any change to a shared default like a production deploy, because that’s what it is.
- Never ship a new domain together with new infrastructure. Launch on an already-allowed domain first and move domains later, announced in advance.
- Keep the incident open until the reports stop, not until the first fix ships.
- Make blocked connections say so.
- Before a migration, make sure you can see which settings each session actually used.
- Let clients report why they can’t connect, through a channel that still works.
- Watch the migration by the setting it changed: successful connections per relay, from day one.
- Slow is smooth, smooth is fast - as they say.
The short version: if you want to test something new in production, you need a scope smaller than “everyone”. Percentage canaries are one option. When your product is built around groups, one group is an even better one.
Related: Designing for Graceful Failure by Separating Failure Domains and When Is a Feature Really Done?