Today I'm excited to announce @nestjs/locks, a new package that brings distributed locks to Nest. With a single decorator, a scheduled job runs on one instance instead of all of them, a slow job never overlaps with itself, and a long-running background task always has exactly one leader. Under the hood, it uses leases and fencing tokens, so it stays correct even when an instance pauses at the worst possible moment.
In case you're not familiar with NestJS, it is a TypeScript Node.js framework that helps you build enterprise-grade efficient and scalable Node.js applications.
Let's dive right in! 🐈
Why distributed locks?
Here's a story I've heard more times than I can count. You add a nightly job with @nestjs/schedule, everything works perfectly on your machine, and then you deploy it behind a load balancer with three replicas. The next morning, every invoice has been exported three times. 🙃
@Cron() runs on every instance, because every instance is the same application. The usual workarounds (an environment variable that marks one replica as "the special one", a separate worker deployment, a SELECT ... FOR UPDATE someone wrote in a hurry) are either fragile or break the moment that one instance goes down. What you actually want is a lock that every instance agrees on.
Getting started
First, install the package:
$ npm i --save @nestjs/locksThen register the module, along with a lock store:
@Module({ imports: [ DrizzleModule.forRootAsync({ /* ... */ }), ScheduleModule.forRoot(), LocksModule.forRoot(), ], providers: [DrizzleLockStore],})export class AppModule {}Locks are only as good as the place they're stored in, so they live in a store shared by all instances. A store is a class with three methods (acquire(), renew(), and release()), each a single atomic operation. The documentation includes complete, ready-to-copy implementations for PostgreSQL (with Drizzle) and Redis (with Lua scripts), so you don't have to come up with these from scratch.
‼️ Without a shared store, the application refuses to start in production. An in-memory lock only protects a single process, which is exactly the situation where you don't need one. You can opt in with allowInMemoryStorage: true, but you have to really mean it. ‼️
Run on one instance
@Cron('0 2 * * *')@OnOneInstance({ key: 'invoices:nightly-export' })async exportInvoices() { // Runs only on the instance that holds the lease}At 2 a.m., every instance wakes up and tries to take the lock. One wins and exports the invoices. The others simply skip this tick. That's it.
Note the explicit key. You could leave it out, but then renaming the method or the class would quietly change the lock's identity. And during a rolling deployment, the old and the new version would happily run the job side by side.
Never overlap
Some jobs run often, and occasionally take longer than their interval:
@Cron('*/5 * * * *')@WithoutOverlapping()async reconcileStock() { // Skips a tick if the previous run is still in progress}If reconciling stock usually takes 1 minute but today takes 7, the next tick is skipped instead of starting a second run that races the first one.
Leader election
Scheduled jobs start and finish. Some work, though, needs to run continuously, such as a subscription to a warehouse feed, a queue consumer that must be unique, or a long-polling connection. For those, there's @LeaderElection():
@Injectable()@LeaderElection('inventory:warehouse-feed')export class WarehouseFeed implements OnLeadershipAcquired { onLeadershipAcquired(lock: Lock) { this.warehouseClient.subscribe({ signal: lock.signal }); }}One instance becomes the leader and subscribes. If it crashes or gets shut down, its lease expires (or is released right away, if you call app.enableShutdownHooks()), and another instance takes over. When leadership is lost, lock.signal aborts, so the subscription stops on its own.
Leases and fencing tokens
So far, this all sounds simple, and that's on purpose. But there's a classic trap in distributed locking that's worth explaining.
Every lock is a lease with a TTL (30s by default) that the holder keeps renewing in the background. If an instance dies, it stops renewing, and the lock frees itself. So far, so good. Now imagine an instance that doesn't die, but pauses: a long GC pause, a frozen container, a VM migration. Its lease expires, another instance acquires the lock and starts writing, and then the first one wakes up, still believing it holds the lock, and overwrites everything with stale data.
The fix is a fencing token: a number that grows every time the lock is acquired. The holder sends it along with its writes, and the database rejects any write carrying a lower token than the last one it accepted:
async writeStock(counts: StockCount[], fencingToken: number) { for (const { productId, count } of counts) { await this.db .update(products) .set({ inStock: count, stockFencingToken: fencingToken }) .where( and(eq(products.id, productId), lte(products.stockFencingToken, fencingToken)), ); }}Inside a locked method, you can read the current token and an AbortSignal from the LocksContext:
async reconcileStock() { const { fencingToken, signal } = this.locksContext; const counts = await this.warehouseClient.stockCounts({ signal }); await this.stockRepository.writeStock(counts, fencingToken!);}The signal aborts as soon as the lease is lost, and the fencing token takes care of whatever was already in flight. Together, a paused instance can't corrupt your data, even in the worst case.
Locking by hand
Decorators cover scheduled work, but sometimes the same job can also be triggered from an endpoint. In that case, use withLock():
@Post('admin/invoice-exports')async exportNow() { try { return await this.locks.withLock( 'invoices:nightly-export', () => this.invoiceExportJob.export(), { wait: '2s' }, ); } catch (error) { if (error instanceof LockNotAcquiredError) { throw new ConflictException('Export already running'); } throw error; }}Since it shares its key with the nightly job, an admin clicking "Export now" at 2 a.m. can't start a second export. Instead, they get a 409 Conflict.
Observability
Losing a lock unexpectedly is something you want to know about. A LockLostError carries a detectedBy field that tells you how the loss was noticed, and the lock-lost, leadership-acquired, and leadership-lost events are published on diagnostics channels (nestjs:locks:lock-lost, and so on), so you can hook them into your logs and alerts.
Testing
Testing time-based behavior shouldn't require waiting. Pass a ManualLockClock as the clock option and move time forward yourself, for example, to check what happens when a lease expires.
And if you write your own store, there's a ready-made contract test suite that verifies it behaves correctly, including under concurrent access:
const cases = lockStoreContract( () => new DrizzleLockStore(db, new LocksStorage()), { concurrent: true },);Try it today!
$ npm i --save @nestjs/locksThe new Distributed locks chapter in the official documentation covers everything above in depth, including the complete PostgreSQL and Redis stores and a production checklist (picking a TTL, Redis persistence, clock synchronization, and more). Give it a try, and please open an issue if anything gets in your way!
Happy coding! 🐈
Learn NestJS - Official NestJS Courses 📚
Level-up your NestJS and Node.js ecosystem skills in these incremental workshop-style courses, from the NestJS Creator himself, and help support the NestJS framework! 🐈🚀 The NestJS Fundamentals Course is now LIVE and 25% off for a limited time!
🎉 NEW - NestJS Course Extensions now live!
- NestJS Advanced Concepts Course now LIVE!
- NestJS Advanced Bundle (Advanced Architecture and Advanced Concepts) now 22% OFF!
- NestJS Microservices now LIVE!
- NestJS Authentication / Authorization Course now LIVE!
- NestJS GraphQL Course (code-first & schema-first approaches) are now LIVE!
- NestJS Authentication / Authorization Course now LIVE!
