Serverless has grown up. Functions that run on demand, managed queues and databases that scale themselves now power everything from marketing sites to busy SaaS products. You pay for what you use, and there are no servers to patch at 2am. But serverless doesn't make failure go away. It just moves it somewhere else, and resilient systems plan for that from day one.
Design every step to be retried
In a serverless system, things will run twice. A queue delivers a message again, a webhook is resent, a function times out halfway through. The fix is idempotency: make every operation safe to repeat.
- Give every job or payment an idempotency key, and skip work you've already done.
- Retry with exponential backoff and a little randomness, so retries don't all land at once.
- Send messages that keep failing to a dead-letter queue, and alert someone, instead of retrying forever.
Keep state out of your functions
Functions are short-lived and may run on a different machine every time. Anything that needs to survive, such as sessions, uploads and progress through a long job, belongs in a database, object storage or a queue. For long workflows, use a workflow or step-function service that records each step, rather than one giant function that has to finish in one go.
Know your limits before your customers find them
Every platform has limits on execution time, memory, payload size and concurrent executions. Write them down for the services you use. Cold starts matter too: keep functions small, trim dependencies, and move latency-sensitive work such as redirects, auth checks and personalisation to edge functions that start almost instantly.
Make it observable
When a request passes through five services, logs alone won't tell you what went wrong. Add structured logs with a request ID that follows the request everywhere, trace calls between services, and set alerts on error rates and queue depth, not just on outages.
When serverless isn't the answer
Steady, heavy workloads such as video processing or long-running connections can be cheaper and simpler on containers. Many of the best architectures are hybrid: serverless for the spiky, event-driven parts, and a small, always-on service where it makes sense.




