Retries and failures
Which provider errors are retried, how backoff works, and what dead-lettering leaves behind.
A queue is only worth having if it retries the right things. better-push decides
that from the normalized ProviderResultCode every provider already returns -
there is no new provider surface and nothing to configure per transport.
This page assumes you have async delivery configured. Without a queue there is exactly one attempt, so every failure below is final.
What is retried
| Code | Retried | Why |
|---|---|---|
rate_limited | yes | the service asked you to slow down, not to stop |
network_error | yes | a connection fault says nothing about the token |
provider_error | yes | a 5xx or a credential fault you can fix and re-run |
invalid_token | no | the token is dead; the device is disabled |
expired_token | no | same, and it will not come back |
payload_too_large | no | the payload stays too large on the next try |
provider_not_configured | no | never enqueued in the first place |
A job is retried when any delivery in it failed with a retryable code.
A run of provider_error means check your config
provider_error on every delivery for one provider is usually a credential
problem - a wrong p8 key, an expired service account, the wrong APNs gateway.
Retries will keep failing until you fix it, then the next attempt succeeds.
Nothing was disabled in the meantime; see
Token Lifecycle.
Retries never re-send
The retry job carries only the deliveries that still need sending. That is not bookkeeping the queue does - it falls out of the delivery status:
- a delivery that succeeded is
sent, - one that failed permanently is
failed, - one waiting for a retry stays
queued, witherrorholding the last code andattemptscounting the tries.
When the retry runs, it reloads only the queued rows. So a batch of three where
one sent, one hit a dead token, and one was rate-limited retries exactly one
delivery - and the device that already got the notification never gets it twice.
Delivery statuses
| Status | Meaning |
|---|---|
queued | no successful attempt yet, including between retries |
sent | handed to the push service |
failed | permanently failed: a non-retryable code, or retries exhausted |
suppressed | a preference switched the channel off; nothing was sent |
delivered | in-app rows (the row in your database is the delivery) |
Backoff
The default schedule is 30s, 2m, 8m, 32m, capped at an hour, each with ±20% jitter. The jitter matters: without it, every worker that failed during the same provider outage would retry at the same instant and cause a second one.
Override it if your traffic wants something else:
dbQueue({
maxAttempts: 8,
backoff: (attempt) => Math.min(5_000 * 2 ** attempt, 600_000),
});attempt is 1 on the first failure. The returned value is milliseconds until
the job is due again.
Dead-lettering
A job dies when it hits a non-retryable failure, or when it has been claimed
maxAttempts times (5 by default). It is marked status = 'dead', its
last_error is kept, and it is never claimed again. Its deliveries are written
as failed with the last provider code.
SELECT id, kind, attempts, last_error, updated_at
FROM bp_job
WHERE status = 'dead'
ORDER BY updated_at DESC;Dead rows are kept for inspection - they are the only place a permanently failed job's history lives, since successful jobs are deleted. Clear them out when you have read them:
These queries are the dbQueue view. On bullmq()
the same state lives in Redis - a dead-lettered job is a BullMQ failed job -
and the studio's queue screen shows either one identically, because it reads
through the backend's own inspector rather than querying bp_job directly.
DELETE FROM bp_job WHERE status = 'dead' AND updated_at < now() - interval '30 days';To retry one by hand after fixing the underlying problem, put it back:
UPDATE bp_job
SET status = 'pending', attempts = 0, run_at = now(), locked_at = NULL
WHERE id = '...';Its deliveries are failed by then, so also set the ones you want re-sent back
to queued - the job only reloads rows in that state.
Attempts are counted at claim time
attempts increments when a worker claims a job, not when it reports a failure.
That is deliberate: a job that crashes its worker every time - a poison job -
would otherwise never burn an attempt, be reclaimed after every visibility
timeout, and loop forever. Counting at claim guarantees it eventually
dead-letters and shows up in the query above.
The same counter is what recovers a crashed worker. A claim that is never settled
ages out after visibilityTimeoutMs (5 minutes by default) and another worker
picks the job up.
Failures that are not the provider's fault
| Situation | What happens |
|---|---|
| the notification was deleted after enqueue | the job is retired, not retried |
| a device row was deleted after enqueue | that delivery is failed with device_deleted |
| the job payload is unreadable | dead-lettered immediately - no retry can fix it |
| the handler throws | treated as retryable; the worker loop survives |
| the worker is killed mid-job | reclaimed after the visibility timeout |