The Deadlock-Detection Storm That Brings PostgreSQL to a Halt

One of the more fun issues I recently had to debug was a PostgreSQL failure mode in which there is no deadlock, the database is not out of CPU (in fact, oddly low CPU usage was observed in our case), and the storage is healthy. And yet, transaction throughput falls all the way to practically zero. Even more confusingly, our database telemetry stopped reporting just when we needed it most.

All it takes is one heavyweight lock held for too long, enough backends queued behind it, and enough continuing arrival pressure.

The proximate cause is the code meant to save PostgreSQL from deadlocks. Each waiter whose deadlock_timeout expires runs the deadlock detector. To obtain a consistent view of the waits-for graph, that detector takes every partition lock protecting PostgreSQL's regular lock table in exclusive mode. There are 16 of these lightweight locks in a standard build. While one backend searches the graph, ordinary heavyweight-lock acquisition and release across the cluster is obstructed. With enough waiters, detectors line up behind detectors and the partitioned lock table behaves like one global lock.

The result is a positive-feedback loop: the queue makes deadlock checks more expensive; expensive checks keep the lock manager unavailable for longer; slower lock processing grows the queue; and the larger queue makes the next round of checks still more expensive.

This is not a theoretical curiosity. A 2017 PostgreSQL hackers thread reported a benchmark falling from roughly 50,000 TPS to intervals of exactly 0 TPS after hundreds of clients queued on a transaction-level advisory lock. A proposed two-pass detector avoided the collapse in that benchmark, but the patch was returned with feedback and never merged. [1] [2] The current PostgreSQL source still takes all 16 lock-manager partition locks exclusively before running DeadLockCheck. [3] [4]

Two kinds of locks

The terminology matters here because two different kinds of locks participate in the failure.

Heavyweight locks, also called regular or lock-manager locks, are the locks visible through pg_locks. They represent locks on relations, transaction IDs, advisory-lock keys, and other database objects. A backend that cannot be granted one of these locks joins that lock object's wait queue and sleeps.

Lightweight locks, or LWLocks, protect PostgreSQL's shared-memory data structures. PostgreSQL divides its shared regular-lock table into 16 hash partitions. Each partition has an LWLock. Normal lock-manager operations need only the LWLock for the partition containing the lock object they are changing, so unrelated heavyweight locks can normally be processed concurrently. [5]

Deadlock detection breaks that partitioning temporarily. A waits-for graph can cross any number of lock objects and therefore any number of hash partitions. PostgreSQL takes the simple, conservative approach: acquire all the partition LWLocks, in partition-number order, and keep them until the check is complete.

The relevant code in CheckDeadLock() is almost a literal description:

for (i = 0; i < NUM_LOCK_PARTITIONS; i++)
    LWLockAcquire(LockHashPartitionLockByIndex(i), LW_EXCLUSIVE);

result = DeadLockCheck(MyProc);

for (i = NUM_LOCK_PARTITIONS; --i >= 0;)
    LWLockRelease(LockHashPartitionLockByIndex(i));

NUM_LOCK_PARTITIONS is 1 << 4, or 16, in current PostgreSQL source. The locks are released in reverse order so another backend that needs all of them is not awakened until it has a chance to acquire the complete set. [3] [4]

How an innocent lock wait becomes a storm

Suppose backend B0 holds heavyweight lock L. It might be an advisory lock, a relation lock, or a transaction-ID lock reached through a row-lock wait. For the failure described here, the exact lock type is less important than its being represented in the regular lock manager.

Backend B1 requests a conflicting mode on L. PostgreSQL puts B1 in L's wait queue, arms a timer, and puts the backend to sleep. PostgreSQL does not run deadlock detection immediately because the check is expensive and most lock waits are not deadlocks.

The timer is deadlock_timeout. Its default is one second. If B0 releases L within that second, B1 wakes normally and no deadlock check occurs. If the timer expires first, B1 wakes, acquires all 16 lock-manager partition LWLocks exclusively, and searches for a cycle involving itself. If there is no cycle, it releases the LWLocks and goes back to sleep on L. [6]

Now put B2 through Bn behind B1. Each backend has its own timeout. Once their waits cross the one-second threshold, they each attempt the same operation:

  1. Acquire all 16 lock-manager partition LWLocks exclusively.
  2. Traverse the waits-for graph starting from itself.
  3. Find no deadlock.
  4. Release all 16 LWLocks and resume waiting for L.

The detectors cannot run concurrently because they all request the same LWLocks exclusively. They form a convoy. During each check, normal operations that need any regular lock-table partition must wait too. A backend releasing a heavyweight lock also needs to enter the lock manager, so the machinery that would drain the original queue now competes with the machinery repeatedly examining it.

Why the work grows with the queue

PostgreSQL's detector searches a waits-for graph. A process points to processes that block it. Some edges are hard edges, created by already granted conflicting locks. Others are soft edges, created by conflicting requests ahead of a process in the same wait queue. PostgreSQL may resolve a cycle containing soft edges by reordering a wait queue rather than aborting a transaction. [5]

Even on a single hot queue with no actual deadlock, a check can do much more than walk the queue once. The detector scans the lock's process-lock list for holders, then examines earlier waiters for conflicting requests. When an earlier waiter blocks the current one, it follows that dependency recursively. Those visits can scan the same lock's lists again, and checking whether a process has already been visited takes another scan. No cycle is required for this work to add up. [11]

Then the next timed-out backend starts its own search. There is no shared "we just checked this queue" result. With enough waiters, PostgreSQL spends more and more time rediscovering that the same congested queue has no deadlock.

The exact cost depends on the lock modes and the shape of the waits-for graph. If the detector finds soft cycles, it can also recursively test hypothetical wait-queue reorderings. The important property for this failure is simpler: larger queues make checks more expensive, each timed-out waiter does the work for itself, and all of that work occurs inside an exclusive all-partitions critical section.

The collapse condition

There is a service-rate problem hiding inside the lock problem.

Let lambda be the rate at which lock waits survive past deadlock_timeout. Let T(n) be the time one deadlock check takes with a queue of size n. Because the checks serialize, the detector convoy can service at most about 1 / T(n) checks per second. Meanwhile T(n) grows with n.

Once the arrival rate exceeds that shrinking service rate, the system becomes self-amplifying:

more timed-out waiters -> larger graph -> longer exclusive hold -> lower lock-manager throughput -> more timed-out waiters

On a many-core server this can look bizarre. Roughly one core's worth of work is busy running the serialized detectors while many other backend processes sleep on LWLock:LockManager or on their original heavyweight locks. CPU can be available in aggregate while the lock-manager path that gates useful work is effectively single-threaded. Existing work that never touches the regular lock manager may continue, but ordinary transactions acquire and release locks often enough that measured TPS can reach literally zero.

Why database telemetry goes dark

The missing telemetry was not an unrelated monitoring failure. It was another client of the subsystem that had stopped making progress.

Most database telemetry is collected from inside PostgreSQL. An exporter opens a connection and periodically queries views such as pg_stat_activity, pg_stat_database, pg_stat_statements, and pg_locks. The graphs may live in a separate monitoring system, but the observations still have to be produced by SQL queries running on the affected database.

Lock telemetry has an especially direct dependency on the contended LWLocks. To produce pg_locks, GetLockStatusData() first inspects the fast-path lock arrays and then acquires all 16 regular lock-manager partition locks in LW_SHARED mode so it can copy a consistent snapshot of the lock table. pg_blocking_pids() follows a similar path: it takes ProcArrayLock and then all 16 lock-manager partitions in shared mode. [8] [10]

Those shared acquisitions cannot pass a detector holding the same partitions in exclusive mode. During a detector storm, the telemetry backend therefore joins the LWLock convoy. If the exporter collects several metrics as one scrape, one stuck lock-inspection query can delay or time out the whole scrape. The monitoring system then receives no fresh database samples and draws a gap, even though PostgreSQL's internal counters and backend activity still exist in shared memory.

This explains the eerie split we observed: host-level telemetry still showed a live machine with plenty of idle CPU, while database-level telemetry vanished. The observer had not remained outside the failure domain. It was querying the frozen system from within it.

No deadlock is required

The most misleading part of the incident is its name. deadlock_timeout is not a maximum lock-wait duration, and crossing it does not imply that a deadlock exists. It is only the delay before a waiting backend pays the cost of checking.

A single long but legitimate critical section is sufficient. An application can manufacture the failure with a transaction-level advisory lock around slow work. A migration holding a strong table lock, a transaction left open after an update, or a sufficiently hot row can provide the same shape: one blocker, a growing queue, and waiters that outlive the timer.

The detector will correctly report "no deadlock" every time. Correctness of each individual decision does not prevent the collection of decisions from destroying throughput.

Recognizing it in production

The signature is a combination of symptoms:

  • TPS and completed queries abruptly fall while offered load remains high.
  • One backend, or roughly one core's worth of backend CPU, remains busy while many other backends sleep.
  • Host telemetry continues while SQL-scraped database telemetry develops gaps or disappears altogether.
  • The database log contains lock-wait messages if log_lock_waits is enabled, but few or no "deadlock detected" errors.

If a diagnostic connection and query can still get through, pg_stat_activity shows the original Lock waits plus a growing number of LWLock waits whose event is LockManager. The same evidence may be available in samples captured while the system was approaching collapse. pg_locks shows many ungranted requests, often concentrated on one lock identity, with waitstart values older than deadlock_timeout.

These SQL views are not dependable during the fully developed failure. pg_locks in particular needs all 16 lock-manager partitions to produce its snapshot, so the query used to confirm the diagnosis can be trapped behind the same detector convoy. At that point, missing database telemetry and diagnostic queries that do not return are themselves part of the signature; use host-level process and CPU observations plus PostgreSQL logs as the out-of-band evidence.

PostgreSQL documents LWLock:LockManager as waiting to read or update heavyweight-lock information. It also recommends pg_blocking_pids() instead of trying to reconstruct blocking relationships by self-joining pg_locks. [7] [8]

A compact first query is:

SELECT wait_event_type, wait_event, count(*)
FROM pg_stat_activity
WHERE state = 'active'
GROUP BY 1, 2
ORDER BY 3 DESC;

If the server is still responsive, find old ungranted locks and their blockers:

SELECT l.pid,
       l.locktype,
       l.mode,
       l.waitstart,
       pg_blocking_pids(l.pid) AS blocking_pids,
       a.query
FROM pg_locks AS l
JOIN pg_stat_activity AS a USING (pid)
WHERE NOT l.granted
ORDER BY l.waitstart;

Be restrained while diagnosing a collapse. Reading lock state is not free: a consistent pg_locks snapshot must itself interact with the lock manager. Repeatedly polling a large lock table can add pressure to the subsystem already in distress. It may simply block, which is why the regular telemetry path may have already gone dark. [8] [10]

Breaking the loop

The immediate objective is to stop the queue from growing and let the lock manager drain.

First, shed or queue new work outside PostgreSQL. Configure the connection pool to limit how many requests may actively use the database at once. Excess requests then wait in the application or pool instead of becoming PostgreSQL backends and joining the lock queue. Thousands of client requests can wait cheaply outside the database; thousands of PostgreSQL backends that have each armed a deadlock timer become active participants in the storm.

Second, identify and end the root blocker if doing so is safe. Canceling work that is merely running does not necessarily end its transaction; terminating the session or explicitly rolling back may be required to release its locks. Lock cleanup still has to make progress through the contended lock manager, so recovery need not be instantaneous.

For a database under heavy use, increase deadlock_timeout above the normal transaction duration. PostgreSQL's documentation calls the one-second default "probably about the smallest value you would want in practice," recommends a higher value on heavily loaded servers, and says the setting should ideally exceed the typical transaction time. A longer timeout gives a healthy but slow lock holder more time to finish before every waiter starts an expensive deadlock check. [6]

This is a tradeoff. Real deadlocks take longer to report. Raising deadlock_timeout is also a fuse adjustment, not a root-cause fix. A blocker that outlasts the larger window, or an arrival rate high enough to fill it, can still produce the same collapse.

Then fix the workload that made a lock live so long:

  • Keep transactions short and never perform remote calls or unrelated slow work while holding a contended lock.
  • Prefer nonblocking advisory-lock APIs such as pg_try_advisory_xact_lock when dropping or retrying work is acceptable.
  • Put a finite lock_timeout on workloads that can safely fail and retry. If it expires before deadlock_timeout, the statement leaves the wait rather than joining the detector wave. PostgreSQL cautions against setting lock_timeout globally because that applies to every session. [9]
  • Cap database concurrency according to the maximum tolerable number of simultaneous lock waiters, not only available CPU and memory.

Why PostgreSQL still works this way

The lock-manager README states the design assumption plainly: in a properly functioning system, deadlock checks should not happen often enough to be performance-critical. Acquiring all partitions is much simpler than safely walking a changing cross-partition graph. [5]

That assumption is sound under ordinary contention. It fails discontinuously under overload. The detector was designed as an exceptional operation, but deadlock_timeout turns a sufficiently old queue into a scheduler for that exceptional operation, once per waiting backend.

The abandoned two-pass patch proposed taking all partitions in shared mode for the common no-deadlock case and repeating the check under exclusive locks only when resolution might mutate a wait queue. The benchmark attached to the patch is compelling evidence for the mechanism, but it is not a fix users can rely on: the proposal was never committed, and current PostgreSQL still uses the exclusive single-pass design. [1] [2] [3]

The operational lesson is broader than this one function. Backpressure must be applied before work becomes expensive shared state. Once every request owns a backend, occupies a lock wait queue, and independently schedules global coordination work, "more concurrency" is no longer concurrency. It is a queue that makes its own server slower.


[1]Yura Sokolov, Two pass CheckDeadlock in contentent case, pgsql-hackers, October 2017.
[2]PostgreSQL Commitfest, Two pass check for deadlock, returned with feedback in February 2019.
[3]PostgreSQL source, CheckDeadLock in proc.c.
[4]PostgreSQL source, lock partition count in lwlock.h.
[5]PostgreSQL source, Locking overview and deadlock-detection algorithm.
[6]PostgreSQL documentation, Lock Management: deadlock_timeout.
[7]PostgreSQL documentation, The Cumulative Statistics System: wait events.
[8]PostgreSQL documentation, The pg_locks View.
[9]PostgreSQL documentation, Client Connection Defaults: lock_timeout.
[10]PostgreSQL source, GetLockStatusData and GetBlockerStatusData in lock.c.
[11]PostgreSQL 18 source, FindLockCycleRecurse and FindLockCycleRecurseMember in deadlock.c.

Comments