Apple’s APNs documentation says: “APNs stores only one notification per bundle ID.” Google’s FCM documentation says an offline Android device keeps at most 100 non-collapsible messages and discards all of them once that limit is reached. Neither sentence fits the way push is usually treated, as a message queue that carries each change to the phone.

That design works on a desk. Then somebody turns on airplane mode for an hour, force-quits the app, or the server sends twenty updates in a minute. Below is what the two services document, and a simulation of what a client has to do about it.

TL;DR

  • Apple documents that APNs is best-effort, may reorder notifications, and stores only one notification per bundle ID for an offline device. Google documents that FCM does not guarantee order, keeps at most 100 non-collapsible messages for an offline Android device, and discards all of them when the limit is exceeded.
  • A background (silent) notification on iOS is documented as low priority, not guaranteed, throttled when frequent, and discarded if the user force-quits the app.
  • So a push can be lost, replaced by a newer one, duplicated by your own retries, or delivered late. A client that treats the payload as the data cannot recover from any of these.
  • Design the push as a hint: “something changed, here is a cursor”. The client then pulls from its own cursor, in the same code that it uses when the app opens. Add at least one trigger for that pull that does not depend on push, such as app foreground or a periodic sync.
  • In a simulation with a lossy channel, applying payloads converged in 77 of 1000 runs; pull-after-hint converged in 855 of 1000; pull-after-hint plus one foreground sync converged in 1000 of 1000. The numbers describe my simulated channel only. They are not delivery rates for APNs or FCM.

What the docs say can happen

SituationWhat the documentation saysWhat it means for the client
Device offlineAPNs may store a notification, for up to 30 days depending on apns-expiration, and delivers it the next time the device is online. It stores only one per bundle ID, usually the latest, “not always guaranteed”.You can receive 1 push for 5 changes.
Device offline (Android)FCM stores messages until the device connects. Non-collapsible messages are limited to 100 on Android; past that, all stored messages are discarded and the app later gets a special signal to do a full sync.You can receive nothing for 100 changes, then a “you missed things” callback.
CollapsingCollapsible messages replace undelivered ones. FCM keeps at most four collapse keys per registration token. Notification messages are always collapsible.Older payloads vanish by design.
OrderingAPNs “may reorder notifications you send to the same device token”. FCM “doesn’t guarantee the order of delivery”.Applying payloads in arrival order can apply them backwards.
Expiryapns-expiration of 0 means attempt once and do not store; FCM ttl of 0 means drop if it cannot be delivered immediately. FCM’s default lifetime is four weeks; a device offline longer gets its messages discarded.A short expiry is a deliberate drop.
Background updates on iOSBackground notifications are low priority, “the system doesn’t guarantee their delivery”, may be throttled, and Apple says not to send more than two or three per hour. A new one replaces the held one. If the app is force-quit, the held one is discarded.A silent push is a nudge that a force-quit can discard.
Priority and Doze (Android)Normal-priority messages can be delayed while the device is in Doze; high-priority ones that do not result in visible notifications may be deprioritized. The handler gets only a few seconds of work.Timing is not predictable, and long work belongs in a job.
App removedFCM discards the message and invalidates the token (UNREGISTERED in the HTTP v1 API, NotRegistered in the legacy one).Dead tokens must be pruned.

I collected these from Apple’s and Google’s current docs. The point is not any one row. It is that every row is a way for “the payload I sent” and “the payload the app saw” to differ.

Treat the push as a hint

If the push can be lost, replaced, delayed or reordered, then correctness cannot depend on it. What is left to a push is the cheapest job: waking the app and saying “ask the server”. The server owns the truth and an ordered log of changes; the client holds a cursor.

// Server: the only place the truth lives.
changes(since) {
  if (log.length && since < log[0].seq - 1) return { reset: true, ...this.snapshot() };   // log no longer reaches back
  return { cursor: seq, events: log.filter((e) => e.seq > since) };
}

// Client: a push is only a reason to pull.
const hintClient = (server) => {
  const s = { items: {}, cursor: 0 };
  const pull = () => {
    const r = server.changes(s.cursor);
    if (r.reset) { s.items = { ...r.items }; s.cursor = r.cursor; return; }
    for (const e of r.events) if (e.seq === s.cursor + 1) { applyOp(s, e); s.cursor = e.seq; }   // contiguous only
  };
  return { s, onPush: (msg) => { if (msg.hintCursor > s.cursor) pull(); }, sync: pull };
};

The push payload carries only { hintCursor }. The client compares it with its own cursor and, if it is behind, pulls. A duplicate push does nothing, an old push does nothing, and a push that overtakes another does one pull that covers both. Events are applied only when they are the next sequence number, so a gap can never be skipped over silently. If the server has trimmed its log past the client’s cursor, it answers with a reset and a snapshot.

The simulation

I did not have access to APNs or FCM for this, so the channel is a model. It implements only behaviors that the table above quotes from the docs: loss, duplication and random delay; “keep only the newest” for an offline device; and “discard everything past a cap”. The server applies 30 changes and keeps 8 of them in its log. I compared two clients against 1000 seeds per channel.

channel      payload-applied    hint + pull    hint + pull + 1 foreground sync
lossy        77/1000 converged  855/1000       1000/1000
keepLatest   4/1000 converged   1000/1000      1000/1000
overflow     0/1000 converged   0/1000         1000/1000

The payload client here is deliberately naive: it applies what arrives, with no sequence check and no pull. Adding sequence numbers would let it notice a gap, but it would still need a way to fetch what is missing, which is the pull. Read the results as follows:

  • Payload-applied almost never reaches the server’s state. Loss, duplication and reordering each break it in different ways.
  • Hint + pull is nearly always right on the lossy channel and exactly right on keepLatest: the single surviving push is enough to trigger a pull of everything. It fails when the last pushes were lost (145 runs on the lossy channel) and when nothing arrives (overflow), because no pull ever started.
  • One foreground sync closes every gap. This is the practical result: push makes the app fresh quickly when it works, and an independent trigger makes it correct whenever it does not.

The exact counts depend on the loss rate and delays I chose (30% loss, 10% duplication, delays up to four steps). Change them in push-hint.mjs and the hint-only column will move. The pattern should not.

Push as the hint, push as the product

Use push-as-hint whenever app state lives on a server and a stale screen is a bug: messaging, collaboration, task lists, sync of any kind. It also keeps your push volume low, since you can coalesce.

If the push is the product, for example a one-time code or a time-sensitive alert whose text is the whole point, the payload has to carry content and you should mark expiry and priority deliberately. Even then, do not make the app’s state depend on it. A visible notification and a state change are two different jobs.

Run the simulation

mkdir push-sim && cd push-sim
# put push-hint.mjs and push-hint.test.mjs from the lab here
node push-hint.mjs
node --test push-hint.test.mjs

To see the failure of the payload approach, edit onPush in payloadClient to log what it receives and run one seed with run({ channel: 'lossy', seed: 7 }).

Ways a push design fails

  • Carrying state in the payload. Symptom: two devices show different data and neither is stale in the usual sense, because they applied different subsets. Fix: send a cursor or a version, never the new value, and let the pull return it.
  • No pull trigger independent of push. Symptom: the simulation’s overflow row: after a burst while offline, nothing arrives and the app stays stale until something else prompts it. Fix: sync on foreground, on reconnect, and on a timer you can justify.
  • Applying an event that is not the next one. Symptom: after a reordered pair, state is subtly wrong and never repaired. Fix: apply only cursor + 1, otherwise pull again.
  • A log that cannot reach back, and no snapshot path. Symptom: a client offline for a while asks for events the server has already trimmed and gets an error or an empty answer. Fix: return a reset with a snapshot, and test it by shrinking retention in a test.
  • Sending a silent push per change. Symptom: iOS throttles them and Apple says not to exceed two or three per hour. Fix: coalesce on the server; one hint per burst is as good as ten.
  • Trusting onMessageReceived for long work. Symptom: the handler is cut off. Google says you get a few seconds and recommends WorkManager for more. Fix: in the handler, schedule a sync job; do the pull there.
  • Keeping dead tokens. Symptom: growing send failures and wasted work. Fix: delete a token when FCM reports it as unregistered, and look at the delivery data FCM offers for dropped messages.

What is measured, and what is modelled

Verified: the simulation, run with Node.js 22.23.3 (node push-hint.mjs, and three tests with node --test that check convergence of hint-plus-sync on every channel and seed, non-convergence of payload application, and staleness without any pull trigger). I read the Apple and Firebase pages for the table above, including apns-expiration, one stored notification per bundle ID, the 100-message limit and the force-quit sentence.

Not verified: any real APNs or FCM behavior. The channel is a model of the documented statements, not a measurement of delivery. I did not run an iOS or Android app. If you want the real numbers for your app, FCM provides delivery data you can query, and in APNs the apns-id and the metrics Apple points to are the places to look. I did not use them.

Design the app for the missing push

The documentation of both services describes a best-effort nudge. Build the app so that a nudge is all it needs: a cursor from the server, a pull the app can run on any trigger, and a snapshot for the case where the log has moved on.