Introducing @nestjs/resilience.

Kamil Mysliwiec | Trilon Consulting
Kamil Mysliwiec

Today I'm excited to announce @nestjs/resilience, a new package that brings retries, timeouts, circuit breakers, bulkheads, and fallbacks to Nest. You apply them as decorators on controllers, GraphQL resolvers, microservice handlers, and WebSocket gateways, or as policy objects inside your services, and slow or failing dependencies stop taking the rest of your application down with them.

In case you're not familiar with NestJS, it is a TypeScript Node.js framework that helps you build enterprise-grade efficient and scalable Node.js applications.

Let's dive right in! 🐈

Why resilience?

Sooner or later, every application that calls another system meets one that is slow or down. A shipping carrier's API starts taking 30 seconds to answer, requests pile up, the event loop fills with pending promises, and suddenly your own API is the one timing out.

The patterns that prevent this are well known: put a deadline on every call, retry what is safe to retry, stop calling a dependency that is clearly broken, cap how much concurrent work a single endpoint can take, and have a plan B. What most teams end up with, though, is a hand-rolled mix of setTimeout(), retry loops, and flags scattered across services. @nestjs/resilience gives you all of these patterns as building blocks that fit the way you already write Nest code.

Getting started

First, install the package:

$ npm i --save @nestjs/resilience

Then register the module in your root module:

import { Module } from '@nestjs/common';
import { ResilienceModule } from '@nestjs/resilience';
@Module({
imports: [
ResilienceModule.forRoot({
defaults: {
timeout: '5s',
retry: { attempts: 3, backoff: { delay: '200ms', maxDelay: '2s' } },
},
presets: {
carrier: {
timeout: '2s',
retry: { attempts: 2 },
circuitBreaker: {
failureRateThreshold: 50,
minimumCalls: 10,
openDuration: '30s',
},
},
},
}),
],
})
export class AppModule {}

There are two things worth noting here. defaults only provide fallback values for fields you leave out, they never enable a stage on their own. presets, on the other hand, are named configurations you share across every handler that talks to the same dependency (here, our shipping carrier).

Timeouts

Let's start with the simplest one:

@Get('quotes')
@Timeout('1.5s')
getQuotes(@Query('orderId') orderId: string, @Signal() signal: AbortSignal) {
return this.shippingService.getQuotes(orderId, signal);
}

When the time runs out, the client receives 504 Gateway Timeout. But that's only half of the story. The @Signal() decorator injects an AbortSignal for the current attempt, and when the timeout elapses, that signal aborts. Pass it down to fetch() (or to @nestjs/http-client) and the work is actually cancelled, instead of quietly running in the background after nobody is waiting for it anymore.

Retries

@Get('quotes')
@Retry({ attempts: 3, backoff: { delay: '200ms', maxDelay: '2s' } })
getQuotes(@Query('orderId') orderId: string) {
return this.shippingService.getQuotes(orderId);
}

Retries are conservative by default. Safe HTTP methods (GET, HEAD, OPTIONS) and GraphQL queries are retried, but a POST or a mutation is not, because repeating it could, for example, create the same shipment twice. When you know an operation is safe to repeat, say so explicitly:

@Post('shipments')
@Retry({ idempotent: true })
createShipment(@Body('orderId') orderId: string, @Signal() signal: AbortSignal) {
return this.shippingService.createShipment(orderId, signal);
}

Client errors (a 4xx HttpException, or any error with a status or statusCode in the 400s) are never retried, since sending the same bad request again won't make it any better. The exceptions are 408 and 429 coming from a dependency, which mean "slow down" and are treated as failures. Also, keep in mind that the timeout applies per attempt, not to the whole sequence.

Circuit breakers

Retrying a dependency that's down only makes things worse. A circuit breaker notices when calls keep failing and stops making them for a while:

@Get('quotes')
@CircuitBreaker({
failureRateThreshold: 50,
minimumCalls: 10,
openDuration: '30s',
})
getQuotes(@Query('orderId') orderId: string) {
return this.shippingService.getQuotes(orderId);
}

Once the breaker has seen at least 10 calls and half of them failed, it opens. While it's open, calls fail immediately with 503 Service Unavailable and a Retry-After header, so your callers know exactly when to come back. After openDuration, the breaker becomes half-open and lets a single probe call through. If it succeeds, traffic flows again.

Fallbacks

Sometimes a degraded answer beats no answer at all. If the carrier is unavailable, we can still show flat-rate shipping:

@Get('quotes')
@Fallback('flatRateQuotes', {
handleIf: (error) => error instanceof CircuitOpenError,
})
getQuotes(@Query('orderId') orderId: string) {
return this.shippingService.getQuotes(orderId);
}
flatRateQuotes() {
return this.shippingService.flatRateQuotes();
}

The fallback method receives the error and the ExecutionContext. And if you misspell its name, the application fails at startup (and not at 3 a.m. in production 🙃). The same goes for invalid durations and unknown presets.

Bulkheads

Some endpoints are expensive: a CSV export, a report, an image conversion. A bulkhead caps how many of them can run at once, so a handful of heavy requests can't starve everything else:

@Get('export')
@Bulkhead({ maxConcurrent: 2, maxQueue: 3, queueTimeout: '10s' })
export() {
return this.ordersService.exportCsv();
}

Two exports run, three more wait in line for up to 10 seconds, and everything beyond that gets 503 with a "Server is at capacity" message. Bulkheads are per-handler by default, but you can name one to share it across several handlers.

Presets

Writing the same five decorators on every handler that calls the carrier gets old quickly. That's what presets are for:

@Get('quotes')
@Resilience('carrier')
@Timeout('1.5s')
getQuotes(@Query('orderId') orderId: string, @Signal() signal: AbortSignal) {
return this.shippingService.getQuotes(orderId, signal);
}

@Resilience('carrier') applies every stage from the preset, and any individual decorator next to it overrides just that stage. Here, quotes get a tighter timeout than the rest of the carrier calls.

Not only HTTP

The very same decorators work on microservice handlers:

@MessagePattern('shipping.quotes')
@Resilience('carrier')
@Retry({ attempts: 3 })
getQuotes(@Payload() data: { orderId: string }, @Signal() signal: AbortSignal) {
return this.shippingService.getQuotes(data.orderId, signal);
}

and GraphQL resolvers:

@Query(() => [ShippingQuoteModel], { nullable: true })
@Resilience('carrier')
@Timeout('1.5s')
shippingQuotes(@Args('orderId') orderId: string, @Signal() signal: AbortSignal) {
return this.shippingService.getQuotes(orderId, signal);
}

Errors are mapped to whatever the transport understands. Over HTTP, that's a 504 or a 503 with a code such as TIMEOUT, CIRCUIT_OPEN, or BULKHEAD_FULL in the body. In GraphQL, the same code ends up in the error's extensions.

‼️ In a hybrid application, connect the microservice with { inheritAppConfig: true }, so it picks up the same configuration as your HTTP app. ‼️

Policies in services

Not everything is an entrypoint. Background jobs and scheduled tasks need resilience too, and for them, there's the ResilienceService:

@Injectable()
export class RepricingService {
private readonly policy: ResiliencePolicy;
constructor(resilience: ResilienceService) {
this.policy = resilience.preset('carrier');
}
async repricePendingOrders(): Promise<number> {
const pending = await this.ordersService.findAll();
let repriced = 0;
for (const order of pending) {
try {
await this.policy.execute(
({ signal }) => this.carrierClient.getQuotes(order, signal),
{ source: 'RepricingService.repricePendingOrders' },
);
repriced++;
} catch (error) {
if (error instanceof CircuitOpenError) {
this.logger.warn(`Carrier unavailable, stopping after ${repriced} orders`);
break;
}
}
}
return repriced;
}
}

Note that the policy shares its breaker with every handler using the carrier preset. Once the HTTP traffic has tripped it, the job stops immediately instead of hammering a dependency that's already down. Besides preset(), the service exposes create() for ad-hoc policies, plus circuitBreaker(name) and bulkhead(name) to inspect their current state.

Pairs well with @nestjs/idempotency

Retrying a POST on the server is only half of the job. The client may retry too (a flaky mobile connection, an impatient user, a load balancer). Combine @Retry({ idempotent: true }) with @Idempotent() from @nestjs/idempotency, and repeated requests are deduplicated before they ever reach your handler:

@Post('shipments')
@Idempotent()
@Resilience('carrier')
@Retry({ idempotent: true })
createShipment(@Body('orderId') orderId: string, @Signal() signal: AbortSignal) {
return this.shippingService.createShipment(orderId, signal);
}

Just make sure to import IdempotencyModule before ResilienceModule, so its interceptor runs outermost.

Observability

You want to know when a breaker opens, preferably before your customers do. Every transition is published on the ResilienceEvents.events$ stream (and on diagnostics channels such as nestjs:resilience:circuit-open):

@Injectable()
export class ResilienceLogger implements OnModuleInit {
private readonly logger = new Logger('Resilience');
constructor(private readonly resilienceEvents: ResilienceEvents) {}
onModuleInit() {
this.resilienceEvents.events$.subscribe((event) => {
if (event.type === 'circuit-open') {
this.logger.warn(`Circuit "${event.policy}" opened`);
}
});
}
}

Testing

You don't have to make a dependency fail 10 times in a row to test what happens when the breaker is open. Just trip it:

it('returns 503 while the carrier breaker is open', async () => {
app.get(ResilienceService).circuitBreaker('carrier').trip();
const response = await request(app.getHttpServer()).get('/shipping/quotes?orderId=1001');
expect(response.status).toBe(503);
expect(response.get('Retry-After')).toBeDefined();
});

And to test timeouts and backoff without actually waiting, fake setTimeout, clearTimeout, and Date, while keeping real I/O for HTTP.

Try it today!

$ npm i --save @nestjs/resilience

The new Resilience chapter in the official documentation covers everything above in depth, along with a production checklist (budgeting attempts × timeout, retrying at one layer only, tuning breakers per dependency, and more). Give it a try, and please open an issue if anything gets in your way!

Happy coding! 🐈


Learn NestJS - Official NestJS Courses 📚

Level-up your NestJS and Node.js ecosystem skills in these incremental workshop-style courses, from the NestJS Creator himself, and help support the NestJS framework! 🐈

🚀 The NestJS Fundamentals Course is now LIVE and 25% off for a limited time!

🎉 NEW - NestJS Course Extensions now live!
#NestJS
#NodeJS

Share this Post!

📬 Trilon Newsletter

Stay up to date with all the latest Articles & News!

More from the Trilon Blog .

Kamil Mysliwiec | Trilon Consulting
Kamil Mysliwiec

Announcing @nestjs/storage

Meet @nestjs/storage, one API for local disks and S3-compatible storage in NestJS, with safe uploads, signed URLs, direct uploads, and an in-memory disk for tests.

Read More
Kamil Mysliwiec | Trilon Consulting
Kamil Mysliwiec

Distributed Locks with @nestjs/locks

Meet @nestjs/locks, distributed locks for NestJS: run scheduled jobs on one instance, prevent overlapping runs, and elect a leader, with leases and fencing tokens.

Read More
Kamil Mysliwiec | Trilon Consulting
Kamil Mysliwiec

Sending Email with @nestjs/mail

Meet @nestjs/mail, a new package for transactional email in NestJS with typed mail classes, safe templates, SMTP and HTTP transports, and outbox integration.

Read More

What we do at Trilon .

At Trilon, our goal is to help elevate teams - giving them the push they need to truly succeed in today's ever-changing tech world.

Trilon - Consulting

Consulting .

Let us help take your Application to the next level - planning the next big steps, reviewing architecture, and brainstorming with the team to ensure you achieve your most ambitious goals!

Trilon - Development and Team Augmentation

Development .

Trilon can become part of your development process, making sure that you're building enterprise-grade, scalable applications with best-practices in mind, all while getting things done better and faster!

Trilon - Workshops on NestJS, Node, and other modern JavaScript topics

Workshops .

Have a Trilon team member come to YOU! Get your team up to speed with guided workshops on a huge variety of topics. Modern NodeJS (or NestJS) development, JavaScript frameworks, Reactive Programming, or anything in between! We've got you covered.

Trilon - Open-source contributors

Open-source .

We love open-source because we love giving back to the community! We help maintain & contribute to some of the largest open-source projects, and hope to always share our knowledge with the world!

Explore more

Write us a message .

Let's talk about how Trilon can help your next project get to the next level.

Rather send us an email? Write to:

hello@trilon.io
© 2019-2026 Trilon.