better-push
Operate in production

Retries and failures

Which provider errors are retried, how backoff works, and what dead-lettering leaves behind.

A queue is only worth having if it retries the right things. better-push decides that from the normalized ProviderResultCode every provider already returns - there is no new provider surface and nothing to configure per transport.

This page assumes you have async delivery configured. Without a queue there is exactly one attempt, so every failure below is final.

What is retried

CodeRetriedWhy
rate_limitedyesthe service asked you to slow down, not to stop
network_erroryesa connection fault says nothing about the token
provider_erroryesa 5xx or a credential fault you can fix and re-run
invalid_tokennothe token is dead; the device is disabled
expired_tokennosame, and it will not come back
payload_too_largenothe payload stays too large on the next try
provider_not_configurednonever enqueued in the first place

A job is retried when any delivery in it failed with a retryable code.

A run of provider_error means check your config

provider_error on every delivery for one provider is usually a credential problem - a wrong p8 key, an expired service account, the wrong APNs gateway. Retries will keep failing until you fix it, then the next attempt succeeds. Nothing was disabled in the meantime; see Token Lifecycle.

Retries never re-send

The retry job carries only the deliveries that still need sending. That is not bookkeeping the queue does - it falls out of the delivery status:

  • a delivery that succeeded is sent,
  • one that failed permanently is failed,
  • one waiting for a retry stays queued, with error holding the last code and attempts counting the tries.

When the retry runs, it reloads only the queued rows. So a batch of three where one sent, one hit a dead token, and one was rate-limited retries exactly one delivery - and the device that already got the notification never gets it twice.

Delivery statuses

StatusMeaning
queuedno successful attempt yet, including between retries
senthanded to the push service
failedpermanently failed: a non-retryable code, or retries exhausted
suppresseda preference switched the channel off; nothing was sent
deliveredin-app rows (the row in your database is the delivery)

Backoff

The default schedule is 30s, 2m, 8m, 32m, capped at an hour, each with ±20% jitter. The jitter matters: without it, every worker that failed during the same provider outage would retry at the same instant and cause a second one.

Override it if your traffic wants something else:

dbQueue({
  maxAttempts: 8,
  backoff: (attempt) => Math.min(5_000 * 2 ** attempt, 600_000),
});

attempt is 1 on the first failure. The returned value is milliseconds until the job is due again.

Dead-lettering

A job dies when it hits a non-retryable failure, or when it has been claimed maxAttempts times (5 by default). It is marked status = 'dead', its last_error is kept, and it is never claimed again. Its deliveries are written as failed with the last provider code.

SELECT id, kind, attempts, last_error, updated_at
FROM bp_job
WHERE status = 'dead'
ORDER BY updated_at DESC;

Dead rows are kept for inspection - they are the only place a permanently failed job's history lives, since successful jobs are deleted. Clear them out when you have read them:

These queries are the dbQueue view. On bullmq() the same state lives in Redis - a dead-lettered job is a BullMQ failed job - and the studio's queue screen shows either one identically, because it reads through the backend's own inspector rather than querying bp_job directly.

DELETE FROM bp_job WHERE status = 'dead' AND updated_at < now() - interval '30 days';

To retry one by hand after fixing the underlying problem, put it back:

UPDATE bp_job
SET status = 'pending', attempts = 0, run_at = now(), locked_at = NULL
WHERE id = '...';

Its deliveries are failed by then, so also set the ones you want re-sent back to queued - the job only reloads rows in that state.

Attempts are counted at claim time

attempts increments when a worker claims a job, not when it reports a failure. That is deliberate: a job that crashes its worker every time - a poison job - would otherwise never burn an attempt, be reclaimed after every visibility timeout, and loop forever. Counting at claim guarantees it eventually dead-letters and shows up in the query above.

The same counter is what recovers a crashed worker. A claim that is never settled ages out after visibilityTimeoutMs (5 minutes by default) and another worker picks the job up.

Failures that are not the provider's fault

SituationWhat happens
the notification was deleted after enqueuethe job is retired, not retried
a device row was deleted after enqueuethat delivery is failed with device_deleted
the job payload is unreadabledead-lettered immediately - no retry can fix it
the handler throwstreated as retryable; the worker loop survives
the worker is killed mid-jobreclaimed after the visibility timeout

On this page