Most systems fail in boring ways. A connection times out. A node goes down. A partition appears and the network splits in two. The boring part is that these failures are predictable. What matters is how you prepare for them.
I spent the last few years building and operating distributed systems at various scales. The single biggest lesson: resilience isn't something you add later. It has to be part of the architecture from the start. Retry logic and circuit breakers on top of a fragile system just give you a fragile system that retries.
The patterns that held up best in production were almost always the simplest ones. Redundancy with clear leader election. Idempotent operations so replaying a message is safe. Careful timeouts with exponential backoff and jitter. None of these are new ideas, but they're easy to skip when you're moving fast.
One thing that surprised me: monitoring tells you something is wrong, but it doesn't tell you why. I've learned to invest more in observability. Structured logging, distributed tracing, and the ability to ask arbitrary questions about the system's state. When a node fails at 3am, you don't want to guess. You want to query.
The systems I admire most are boring. They don't fail in interesting ways because they've been designed to handle the interesting failures before they happen. That's the goal: make distributed systems as uneventful as a single machine. Hard to achieve, but worth aiming for.