Today I'm excited to announce @nestjs/resilience, a new package that brings retries, timeouts, circuit breakers, bulkheads, and fallbacks to Nest. You apply them as decorators on controllers, GraphQL resolvers, microservice handlers, and WebSocket gateways, or as policy objects inside your services, and slow or failing dependencies stop taking the rest of your application down with them.
In case you're not familiar with NestJS, it is a TypeScript Node.js framework that helps you build enterprise-grade efficient and scalable Node.js applications.
Let's dive right in! 🐈
Why resilience?
Sooner or later, every application that calls another system meets one that is slow or down. A shipping carrier's API starts taking 30 seconds to answer, requests pile up, the event loop fills with pending promises, and suddenly your own API is the one timing out.
The patterns that prevent this are well known: put a deadline on every call, retry what is safe to retry, stop calling a dependency that is clearly broken, cap how much concurrent work a single endpoint can take, and have a plan B. What most teams end up with, though, is a hand-rolled mix of setTimeout(), retry loops, and flags scattered across services. @nestjs/resilience gives you all of these patterns as building blocks that fit the way you already write Nest code.
Getting started
First, install the package:
$ npm i --save @nestjs/resilienceThen register the module in your root module:
import { Module } from '@nestjs/common';import { ResilienceModule } from '@nestjs/resilience';@Module({ imports: [ ResilienceModule.forRoot({ defaults: { timeout: '5s', retry: { attempts: 3, backoff: { delay: '200ms', maxDelay: '2s' } }, }, presets: { carrier: { timeout: '2s', retry: { attempts: 2 }, circuitBreaker: { failureRateThreshold: 50, minimumCalls: 10, openDuration: '30s', }, }, }, }), ],})export class AppModule {}There are two things worth noting here. defaults only provide fallback values for fields you leave out, they never enable a stage on their own. presets, on the other hand, are named configurations you share across every handler that talks to the same dependency (here, our shipping carrier).
Timeouts
Let's start with the simplest one:
@Get('quotes')@Timeout('1.5s')getQuotes(@Query('orderId') orderId: string, @Signal() signal: AbortSignal) { return this.shippingService.getQuotes(orderId, signal);}When the time runs out, the client receives 504 Gateway Timeout. But that's only half of the story. The @Signal() decorator injects an AbortSignal for the current attempt, and when the timeout elapses, that signal aborts. Pass it down to fetch() (or to @nestjs/http-client) and the work is actually cancelled, instead of quietly running in the background after nobody is waiting for it anymore.
Retries
@Get('quotes')@Retry({ attempts: 3, backoff: { delay: '200ms', maxDelay: '2s' } })getQuotes(@Query('orderId') orderId: string) { return this.shippingService.getQuotes(orderId);}Retries are conservative by default. Safe HTTP methods (GET, HEAD, OPTIONS) and GraphQL queries are retried, but a POST or a mutation is not, because repeating it could, for example, create the same shipment twice. When you know an operation is safe to repeat, say so explicitly:
@Post('shipments')@Retry({ idempotent: true })createShipment(@Body('orderId') orderId: string, @Signal() signal: AbortSignal) { return this.shippingService.createShipment(orderId, signal);}Client errors (a 4xx HttpException, or any error with a status or statusCode in the 400s) are never retried, since sending the same bad request again won't make it any better. The exceptions are 408 and 429 coming from a dependency, which mean "slow down" and are treated as failures. Also, keep in mind that the timeout applies per attempt, not to the whole sequence.
Circuit breakers
Retrying a dependency that's down only makes things worse. A circuit breaker notices when calls keep failing and stops making them for a while:
@Get('quotes')@CircuitBreaker({ failureRateThreshold: 50, minimumCalls: 10, openDuration: '30s',})getQuotes(@Query('orderId') orderId: string) { return this.shippingService.getQuotes(orderId);}Once the breaker has seen at least 10 calls and half of them failed, it opens. While it's open, calls fail immediately with 503 Service Unavailable and a Retry-After header, so your callers know exactly when to come back. After openDuration, the breaker becomes half-open and lets a single probe call through. If it succeeds, traffic flows again.
Fallbacks
Sometimes a degraded answer beats no answer at all. If the carrier is unavailable, we can still show flat-rate shipping:
@Get('quotes')@Fallback('flatRateQuotes', { handleIf: (error) => error instanceof CircuitOpenError,})getQuotes(@Query('orderId') orderId: string) { return this.shippingService.getQuotes(orderId);}flatRateQuotes() { return this.shippingService.flatRateQuotes();}The fallback method receives the error and the ExecutionContext. And if you misspell its name, the application fails at startup (and not at 3 a.m. in production 🙃). The same goes for invalid durations and unknown presets.
Bulkheads
Some endpoints are expensive: a CSV export, a report, an image conversion. A bulkhead caps how many of them can run at once, so a handful of heavy requests can't starve everything else:
@Get('export')@Bulkhead({ maxConcurrent: 2, maxQueue: 3, queueTimeout: '10s' })export() { return this.ordersService.exportCsv();}Two exports run, three more wait in line for up to 10 seconds, and everything beyond that gets 503 with a "Server is at capacity" message. Bulkheads are per-handler by default, but you can name one to share it across several handlers.
Presets
Writing the same five decorators on every handler that calls the carrier gets old quickly. That's what presets are for:
@Get('quotes')@Resilience('carrier')@Timeout('1.5s')getQuotes(@Query('orderId') orderId: string, @Signal() signal: AbortSignal) { return this.shippingService.getQuotes(orderId, signal);}@Resilience('carrier') applies every stage from the preset, and any individual decorator next to it overrides just that stage. Here, quotes get a tighter timeout than the rest of the carrier calls.
Not only HTTP
The very same decorators work on microservice handlers:
@MessagePattern('shipping.quotes')@Resilience('carrier')@Retry({ attempts: 3 })getQuotes(@Payload() data: { orderId: string }, @Signal() signal: AbortSignal) { return this.shippingService.getQuotes(data.orderId, signal);}and GraphQL resolvers:
@Query(() => [ShippingQuoteModel], { nullable: true })@Resilience('carrier')@Timeout('1.5s')shippingQuotes(@Args('orderId') orderId: string, @Signal() signal: AbortSignal) { return this.shippingService.getQuotes(orderId, signal);}Errors are mapped to whatever the transport understands. Over HTTP, that's a 504 or a 503 with a code such as TIMEOUT, CIRCUIT_OPEN, or BULKHEAD_FULL in the body. In GraphQL, the same code ends up in the error's extensions.
‼️ In a hybrid application, connect the microservice with { inheritAppConfig: true }, so it picks up the same configuration as your HTTP app. ‼️
Policies in services
Not everything is an entrypoint. Background jobs and scheduled tasks need resilience too, and for them, there's the ResilienceService:
@Injectable()export class RepricingService { private readonly policy: ResiliencePolicy; constructor(resilience: ResilienceService) { this.policy = resilience.preset('carrier'); } async repricePendingOrders(): Promise<number> { const pending = await this.ordersService.findAll(); let repriced = 0; for (const order of pending) { try { await this.policy.execute( ({ signal }) => this.carrierClient.getQuotes(order, signal), { source: 'RepricingService.repricePendingOrders' }, ); repriced++; } catch (error) { if (error instanceof CircuitOpenError) { this.logger.warn(`Carrier unavailable, stopping after ${repriced} orders`); break; } } } return repriced; }}Note that the policy shares its breaker with every handler using the carrier preset. Once the HTTP traffic has tripped it, the job stops immediately instead of hammering a dependency that's already down. Besides preset(), the service exposes create() for ad-hoc policies, plus circuitBreaker(name) and bulkhead(name) to inspect their current state.
Pairs well with @nestjs/idempotency
Retrying a POST on the server is only half of the job. The client may retry too (a flaky mobile connection, an impatient user, a load balancer). Combine @Retry({ idempotent: true }) with @Idempotent() from @nestjs/idempotency, and repeated requests are deduplicated before they ever reach your handler:
@Post('shipments')@Idempotent()@Resilience('carrier')@Retry({ idempotent: true })createShipment(@Body('orderId') orderId: string, @Signal() signal: AbortSignal) { return this.shippingService.createShipment(orderId, signal);}Just make sure to import IdempotencyModule before ResilienceModule, so its interceptor runs outermost.
Observability
You want to know when a breaker opens, preferably before your customers do. Every transition is published on the ResilienceEvents.events$ stream (and on diagnostics channels such as nestjs:resilience:circuit-open):
@Injectable()export class ResilienceLogger implements OnModuleInit { private readonly logger = new Logger('Resilience'); constructor(private readonly resilienceEvents: ResilienceEvents) {} onModuleInit() { this.resilienceEvents.events$.subscribe((event) => { if (event.type === 'circuit-open') { this.logger.warn(`Circuit "${event.policy}" opened`); } }); }}Testing
You don't have to make a dependency fail 10 times in a row to test what happens when the breaker is open. Just trip it:
it('returns 503 while the carrier breaker is open', async () => { app.get(ResilienceService).circuitBreaker('carrier').trip(); const response = await request(app.getHttpServer()).get('/shipping/quotes?orderId=1001'); expect(response.status).toBe(503); expect(response.get('Retry-After')).toBeDefined();});And to test timeouts and backoff without actually waiting, fake setTimeout, clearTimeout, and Date, while keeping real I/O for HTTP.
Try it today!
$ npm i --save @nestjs/resilienceThe new Resilience chapter in the official documentation covers everything above in depth, along with a production checklist (budgeting attempts × timeout, retrying at one layer only, tuning breakers per dependency, and more). Give it a try, and please open an issue if anything gets in your way!
Happy coding! 🐈
Learn NestJS - Official NestJS Courses 📚
Level-up your NestJS and Node.js ecosystem skills in these incremental workshop-style courses, from the NestJS Creator himself, and help support the NestJS framework! 🐈🚀 The NestJS Fundamentals Course is now LIVE and 25% off for a limited time!
🎉 NEW - NestJS Course Extensions now live!
- NestJS Advanced Concepts Course now LIVE!
- NestJS Advanced Bundle (Advanced Architecture and Advanced Concepts) now 22% OFF!
- NestJS Microservices now LIVE!
- NestJS Authentication / Authorization Course now LIVE!
- NestJS GraphQL Course (code-first & schema-first approaches) are now LIVE!
- NestJS Authentication / Authorization Course now LIVE!
