I made a Durable Object alarm handler throw seven times in a row on a local runtime. It ran at +0, 2, 6, 14.5, 32, 68 and 145 seconds. Then nothing woke the object again, and the lease it was supposed to expire sat in the table, overdue.
That is what the alarms documentation says will happen, in three sentences: an object has one alarm, alarms run at least once, and a failing handler is retried only a limited number of times. They matter if the object owns time-limited things (leases, sessions, reservations) that must be cleaned up when they expire, which is what an alarm looks made for. Below is a small expiry ledger built to survive all three, with what each part did when I ran it.
TL;DR
- One alarm per object.
setAlarmreplaces whatever alarm is set. If one object owns many expirations, the alarm is only a wake-up call for the earliest one; the list of expirations belongs in a table. - At least once. If the handler throws after doing the external work, it runs again from the beginning. The work must tolerate a repeat. In my test it ran twice, as predicted.
- Retries are bounded. The docs say exponential backoff from a 2-second delay, up to 6 retries. On local workerd I saw 7 attempts (the first plus 6), spaced roughly 2, 4, 8, 18, 36 and 77 seconds apart, and then the alarm was gone:
getAlarm()returnednulland the overdue lease stayed in the table. - Re-arm on wake-up. A constructor that sets an alarm only when none is set and a row is pending brought the expiry back after a runtime restart. A constructor that sets an alarm unconditionally can interfere with an alarm that was already set, which the docs warn about.
- These are local results. I did not run on Cloudflare’s network.
What the documentation promises
From the alarms page (checked on the day of writing):
- Each Durable Object can have one alarm at a time. Calling
setAlarm()when one is already set overrides it. - Alarms have “guaranteed at-least-once execution” and are retried automatically if
alarm()throws, with exponential backoff starting at a 2-second delay and up to 6 retries. This applies only to the most recentsetAlarm()call. - The handler receives
retryCountandisRetry. Only onealarm()runs at a time per object. If the object is terminated unexpectedly,alarm()may be re-run from the beginning on another machine. - Inside
alarm(),getAlarm()returnsnullunlesssetAlarm()was called since the handler started. - If the object wakes up, the constructor runs before
alarm(). The docs warn that asetAlarm()in the constructor can interfere with an alarm that is already set, so check first. - If you want retries beyond the built-in ones, the docs advise catching exceptions in
alarm()and rescheduling.
From the storage API page: setAlarm() with a time at or before now schedules the alarm for the immediate future, and alarms usually start within milliseconds “but can be delayed by up to a minute due to maintenance or failures while failover takes place”.
The storage API (ctx.storage.sql) is SQLite-backed, and writes that happen with no await in between are committed together. I rely on that only loosely in the code below.
The ledger, annotated
The whole object is about 35 lines. I will go through it in the order it executes.
export class LeaseLedger extends DurableObject {
constructor(ctx, env) {
super(ctx, env);
this.sql = ctx.storage.sql;
this.sql.exec(`CREATE TABLE IF NOT EXISTS leases (id TEXT PRIMARY KEY, expire_at INTEGER NOT NULL)`);
this.sql.exec(`CREATE INDEX IF NOT EXISTS leases_by_expiry ON leases (expire_at)`);
ctx.blockConcurrencyWhile(() => this.ensureAlarm()); // (1)
}
(1) Every time the object is built (first request, after eviction, after a restart), it makes sure a wake-up is armed. blockConcurrencyWhile holds incoming events until that finishes, so no request sees a half-initialized object.
async ensureAlarm() {
const next = this.sql.exec(`SELECT MIN(expire_at) AS t FROM leases`).one().t; // (2)
if (next === null) return;
const current = await this.ctx.storage.getAlarm();
if (current === null || next < current) await this.ctx.storage.setAlarm(next); // (3)
}
(2) The table is the truth; the alarm is derived from it. (3) The alarm moves only earlier, never later, and is untouched if it is already at the right time. This is the guard the docs ask for. The function is idempotent, so it is safe to call from the constructor, from grant, and from alarm.
async grant(id, ttlMs) {
this.sql.exec(`INSERT OR REPLACE INTO leases (id, expire_at) VALUES (?, ?)`, id, Date.now() + ttlMs);
await this.ensureAlarm();
}
A new lease is one insert and one ensureAlarm. If the new lease expires sooner than the pending alarm, the alarm moves forward.
async alarm() {
const due = this.sql.exec(`SELECT id FROM leases WHERE expire_at <= ? ORDER BY expire_at LIMIT 100`, Date.now()).toArray(); // (4)
for (const { id } of due) {
await this.revoke(id); // (5)
this.sql.exec(`DELETE FROM leases WHERE id = ?`, id); // (6)
}
await this.ensureAlarm(); // (7)
}
async revoke(_id) { /* the external effect */ }
}
(4) Read what is due now; do not trust the number of times you were called. toArray() finishes the cursor before the first await. (5) and (6) are the order that matters. Effect first, then remove the row. If the process dies between them, the row survives and the next run repeats the effect. The reverse order would lose the effect on a crash; “at least once” means you choose duplicates over losses and make duplicates harmless. (7) Chains to the next expiry, or to the rows beyond the LIMIT. Since getAlarm() is null inside the handler unless you set one, ensureAlarm simply sets it.
What workerd did
I used wrangler dev 4.147.0 (local workerd, compatibility date 2026-09-01) and a lab subclass that overrides revoke to count calls and to throw on request. All numbers below are from that local runtime.
Normal path. Two leases with 300 ms and 1500 ms lifetimes. The alarm fired once for each, the second 1.2 s after the first, and the table ended up empty with getAlarm() equal to null.
Crash after the effect. I made revoke throw once after recording its call. The first attempt ran the effect and failed before the DELETE; the platform retried about two seconds later with retryCount=1, ran the effect again, and removed the row. The effect ran twice for one lease.
Alarm moves only earlier. I granted a 5 s lease, then a 60 s lease (the alarm did not move), then a 1 s lease (the alarm moved earlier).
Exhausting the retries. I made the handler throw 7 times. The journal shows attempts at +0, 2.0, 6.1, 14.5, 32.1, 67.9 and 144.6 seconds, with retryCount from 0 to 6. After that no further attempt happened over the next 100 seconds of polling; getAlarm() returned null and the lease was still in the table, overdue. This matches the docs: once the retries are spent, nothing runs the alarm until the next setAlarm.
Revival. I cleared the failure injection, stopped and restarted wrangler dev with the same persisted state, and made a single request. The object’s constructor ran ensureAlarm, found an overdue row and no alarm, and set one. The alarm fired within about a second and the lease was removed.
The lab object has a switch that keeps the constructor from re-arming during the exhaustion test; without it, even a status request would have revived the object. The clean result above was produced with the switch working.
Making the effect tolerate a repeat
The ledger makes repeats rare. It cannot make them impossible, so the effect needs a shape that survives them. Two shapes work:
- Naturally idempotent: “revoke credential X” or “set status to expired” leaves the same state when run twice. My test counts the calls, but the observable state is the same after one call or two.
- Keyed: when the effect is a call to something else (send, charge, create), pass the lease
idas the key so the other side can recognize the repeat.
I kept the effect a stub in the lab, so this section is a description of the shape and not something I tested end to end.
Run it on workerd
mkdir alarm-lab && cd alarm-lab
# put wrangler.jsonc, src/ledger.js, src/index.js, lab.mjs, probe.mjs from the lab here
npm i -D [email protected] # needs Node.js 22
npx wrangler dev --port 8799 --persist-to ./state &
node lab.mjs main # happy path, crash after the effect, alarm only moves earlier
node probe.mjs # seven failures; about four minutes
The probe prints attempts=7 and alarm=null while the lease is still listed. To see revival, stop wrangler, clear the lab switch (/noboot?on=0 on that object before stopping), start it again with the same --persist-to, and run node lab.mjs revive <object>.
Ways an alarm handler fails silently
- Treating the alarm as the data. Symptom: two expirations, one alarm, and the later one is lost when the earlier
setAlarmreplaces it. Fix: keep expirations in a table and set the alarm to the minimum. - Assuming a failing handler is retried forever. Symptom: a transient outage or a bug lasting a few minutes leaves overdue rows that nothing processes, with no error afterwards. Fix: re-arm on wake-up as above, and add a scheduled check from outside (a cron-style job that touches objects with overdue work). Catching the exception and rescheduling yourself is the docs’ other option; I did not test it, and it needs a per-row attempt counter or a bad row will block the rest.
- Setting the alarm unconditionally in the constructor. Symptom: the alarm is pushed later every time the object wakes, and some expirations run late. Fix: read
getAlarm()and move it only earlier. I followed the docs’ warning here; I did not reproduce the harmful version. - Deleting before the effect. Symptom: a crash loses the effect permanently. Fix: effect first, then the row, and an effect that tolerates repetition.
- Reading the clock instead of the table. Symptom: alarms fire early or late and the handler does nothing, or handles the wrong rows. Fix: query due rows by
expire_at <= nowevery time. Docs say an alarm may be late, and an early or duplicate run is harmless if the query decides. - Calling
setAlarmfromalarm()and expectinggetAlarm()to show the old one. Symptom: a re-arm decision is based onnull. Fix: remember that inside the handlergetAlarm()isnullunless you have set a new one, as the docs say.
When an alarm fits
Use this shape when one object owns many time-based obligations and “eventually, at least once, within about a minute” is acceptable. The ledger plus alarm pair is small and everything lives in the object’s own storage.
Do not use it for work that must happen at an exact moment, or for cleanup that nobody notices if it is missed for a day unless you also add the external check. If the effect cannot be made repeat-safe, the alarm is the wrong tool until you can.
What workerd showed, and what only Cloudflare can
Verified on local wrangler dev 4.147.0 (workerd) with Node.js 22.23.3: the normal path, the double execution after a crash, the alarm moving only earlier, seven attempts with the spacing above followed by null and an overdue row, and revival by the constructor after a restart. The documentation statements were read on the Cloudflare pages named above.
Not verified: the Cloudflare production network. The retry spacing above is the local runtime’s; production timing may differ. The “up to a minute” delay, behavior when an object is moved to another machine, and the effect of deleteAlarm() inside the handler are the docs’ statements, not mine. I did not test the catch-and-reschedule alternative, a constructor that sets the alarm unconditionally, or an external poller.
What an alarm promises
An alarm is a promise to call you, not a promise to finish your work. Keep the work in a table, let every wake-up re-check it, and write the effect so that being called twice costs nothing.