An inactive replication slot is filling your disk

· Vladimir Chemeris, the developer of HeapDeck

Disk usage on the primary climbs in a straight line. Checkpoints run, nothing is in the logs, the tables aren’t growing that fast, and yet pg_wal is tens of gigabytes and getting bigger. The usual cause is a replication slot whose consumer went away.

What a slot holds on to

A replication slot records how far its consumer has read: restart_lsn. The primary promises to keep every WAL segment from that point on, so the consumer can always resume where it stopped. Physical slots serve streaming replicas; logical slots serve logical replication subscriptions and change-data-capture tools such as Debezium.

While the consumer is connected, restart_lsn moves forward and old WAL is recycled at each checkpoint. When the consumer disappears, a replica that was decommissioned, a CDC connector that was switched off, a subscription dropped on the subscriber but not on the publisher, the slot stays. Its restart_lsn stops moving and the primary keeps every byte of WAL written since. By default there is no limit: the server keeps WAL until the disk is full, and a primary that can’t write WAL stops.

A slot can hold back more than WAL. A physical slot used with hot_standby_feedback and a logical slot (through catalog_xmin) keep vacuum from removing rows that their consumer might still need. An abandoned slot can therefore also mean growing bloat and a transaction ID age that never goes down.

Finding the slot

Run this on the primary:

SELECT slot_name,
       slot_type,
       database,
       active,
       wal_status,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal
FROM pg_replication_slots
ORDER BY pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) DESC NULLS LAST;

Look for active = false with a large retained_wal. On PostgreSQL 17 and later, inactive_since tells you when the slot was last used, which usually settles who it belonged to.

wal_status says how much danger the slot is in:

  • reserved: the WAL it needs fits within max_wal_size;
  • extended: it needs more than that, and the server is keeping the extra because of the slot;
  • unreserved: the slot will lose required WAL at the next checkpoint;
  • lost: the WAL it needs is already gone, and the slot can’t be used to resume.

Deciding what to do

First find the owner. Slot names usually help: a replica’s name, debezium, the name of a subscription. A physical consumer that is connected shows up in pg_stat_replication; a logical subscriber has the subscription in its own pg_subscription.

If the consumer is gone for good, drop the slot.

SELECT pg_drop_replication_slot('debezium_orders');

The retained WAL is released and disk space comes back at the next checkpoint. If the consumer still exists and comes back later, it won’t be able to resume: a replica has to be rebuilt from a base backup, and a CDC pipeline needs a fresh snapshot.

PostgreSQL refuses to drop a slot that is in use and says which process holds it, for example replication slot "x" is active for PID 182. Stop the consumer first.

If the consumer is coming back, keep the slot and watch the disk. Since PostgreSQL 13, max_slot_wal_keep_size caps how much WAL any slot can hold. A slot that would need more is invalidated (wal_status = lost) instead of filling the disk. You trade an outage for a rebuild of that one consumer, which is usually the better deal.

Keeping it from happening again

  • Alert on slots that are inactive or retain more than a few gigabytes of WAL, not only on disk usage, which comes too late.
  • Set max_slot_wal_keep_size to a size your disk can afford.
  • Add “drop its replication slot” to the checklist for decommissioning a replica or a CDC connector.

On the phone

In HeapDeck, the Replication tile on Health turns amber as soon as any slot is inactive, usually days before the disk alert. The Replication screen lists inactive slots first with the WAL each retains and shows a red banner: An inactive slot makes PostgreSQL keep WAL forever. This is the most common cause of a disk filling up silently. Slot details show wal_status, how long ago it was last active on PostgreSQL 17 and later, and for HeapDeck Pro a Drop slot button that only works after you type the slot name.

More about the screen: replication lag and slots on your iPhone.