Metrics, rollups, and retention
Hourly rollups that keep charts cheap, opt-in pruning, worker heartbeats, and the onEvent hook.
Notification tables grow forever, and the obvious dashboard query - "deliveries in the last 24h grouped by status" - is a full scan of the largest table. This page is about the machinery that stops both being true.
All of it rides on the queue: a worker schedules and runs it, with no cron and no scheduler.
Rollups
bp_metric_rollup holds hourly counts keyed by
(bucket, type, channel, provider, status). Studio charts read it instead of
scanning bp_delivery.
They are on by default whenever a queue is configured, because they are additive, cheap, and the only thing keeping charts affordable:
betterPush({
// …
queue: dbQueue(),
metrics: { rollups: true, rollupLookbackHours: 6 },
});rollupLookbackHours is the interesting one. A delivery's status keeps
changing after the hour it was created in - queued becomes sent, a retry
becomes failed - so recent buckets have to be revisited rather than written
once. Each run recomputes the last N hours and replaces them, which is also what
makes a re-run idempotent rather than doubling counts.
The current, incomplete hour is never rolled up: it would be wrong the moment it was written. The stats layer reads raw rows for that live edge and stitches the two together.
No worker means no rollups, and that is fine
Without a queue there is no worker, so the watermark never advances. The stats layer notices and falls back to aggregating raw rows, which is affordable at the volumes inline delivery implies. The studio says so with a banner rather than silently showing an empty chart.
Retention
Nothing is ever deleted unless you configure it. That is the only safe default for someone else's notification history.
betterPush({
// …
retention: {
deliveries: "30d",
notifications: "90d",
deadJobs: "14d",
rollups: "13mo",
audit: "1y",
digestWindows: "30d",
},
});digestWindows prunes flushed digest windows only. An open
window is live state, not history: it holds items a user is still owed, so it is
never deleted no matter how old it is.
Durations are <number><unit> with ms, s, m, h, d, w, mo, or y.
A typo is rejected at construction with the valid units listed - the one place a
permissive parser would be genuinely dangerous.
notifications must outlast deliveries
Deleting a notification cascades to its deliveries, so a shorter notification window would silently win and your configured delivery retention would be a lie. That configuration is rejected:
invalid betterPush() configuration (at retention.notifications): "7d" is shorter
than retention.deliveries "30d", but deleting a notification also deletes its
deliveries - so the delivery window could never be honoured.How pruning runs
A prune job is scheduled once a day, and only when retention is set. It
deletes in bounded batches (5,000 rows per statement) inside a time budget,
so it never takes a long lock on a live table. Whatever is left is picked up by
the next day's run.
Scheduling without a scheduler
Maintenance needs no cron, no leader election, and no new machinery. Every
worker simply tries to enqueue jobs whose ids come from the clock -
rollup-2026-07-29T15, prune-2026-07-30 - and the primary key does the rest:
inserting an id that already exists is a no-op, so ten workers ticking at once
produce exactly one row, and SKIP LOCKED means exactly one of them runs it.
Each tick schedules the next period, due at that boundary. A completed job's row is deleted, so scheduling the current period would re-enqueue it the moment it finished.
Worker heartbeats
bp_job.locked_by only names workers currently holding a claim, so an idle
worker - the normal state of a healthy queue - would be invisible. Every worker
upserts its own bp_worker row every ~15 seconds with its hostname, uptime, and
how many jobs it has processed.
The studio shows a worker as live while its heartbeat is under 60 seconds old, and the same data drives the health endpoint's "jobs are pending but no worker has heartbeated recently".
Set BETTER_PUSH_WORKER_VERSION in your deployment to see a half-finished
rollout at a glance.
The onEvent hook
better-push deliberately has no HTTP request log table. A feed polling every 30 seconds would write more rows than the notification data itself, and anyone who wants that retention already has a system built for it.
Instead, the lifecycle events it already produces are available as a callback:
betterPush({
// …
onEvent: (event) => {
switch (event.type) {
case "notification.created":
case "delivery.attempted":
case "job.dead_lettered":
telemetry.record(event);
}
},
});It is fire-and-forget: never awaited on the send path, and errors - thrown or rejected - are logged rather than propagated. A telemetry sink must never be able to slow a notification down or fail one.
The indexes
The canonical schema includes the time-range indexes these queries need:
CREATE INDEX bp_delivery_created_idx ON bp_delivery (created_at DESC);
CREATE INDEX bp_delivery_status_created_idx ON bp_delivery (status, created_at DESC);
CREATE INDEX bp_delivery_device_idx ON bp_delivery (device_id);
CREATE INDEX bp_notification_created_idx ON bp_notification (created_at DESC);drizzle-kit generate includes them with the better-push tables.
Maintenance options
Prop
Type
Prop
Type