Mobile Shell — Architecture
Feature: 051-mobile-app-shell (T083) · Written: 2026-09-03
Source: apps/mobile/ · Spec: specs/051-mobile-app-shell/
The FlowPOS mobile app is an Expo/React Native container that hosts the
unmodified frontend-pwa build in a WebView. The web application is the
product; the container gives it an installed-app lifecycle, an offline cold start,
a durable session, and a release channel that does not go through app review.
The container deliberately owns very little. Everything it does own is here.
To install a debug build on a simulator, emulator, or USB device, and to see backend / PWA / shell changes locally, in staging, and in production: Run locally.
The one rule that must never change
The web build is served from
http://127.0.0.1:47285, and that origin is permanent.
Not a preference — a data-retention boundary. Browser storage is keyed by origin, so changing the host or the port makes every merchant's stored data unreachable: their Firebase session, their IndexedDB, their cached queries. Every device would silently sign out at once, and nothing in the logs would say why.
Two related decisions follow from it, and both are load-bearing:
- The port is pinned, never
0. An ephemeral port produces a new origin on every launch, which means storage is wiped on every launch. The failure looks like "the app randomly signs people out", which is nearly impossible to diagnose from a merchant report. - It is loopback HTTP, not
file://.file://is not a secure context, so IndexedDB andlocalStorageboth fail their probes and Firebase Auth silently falls back to in-memory persistence — signing the merchant out on every cold start (research D2). Serving over loopback HTTP is what makes the session durable at all.
This costs an Android API 28 floor, because loopback HTTP needs
usesCleartextTraffic and cleartext is blocked by default from API 28 onward.
That cost was accepted knowingly; the alternative was an app that cannot stay
signed in.
If the origin must ever change, it is a data migration, not a config edit.
Startup sequence
App.tsx
├─ initObservability() Sentry, tagged with update id / channel / runtime version
├─ startUpdateSync() fire-and-forget; applies at the NEXT cold start
├─ evaluateSessionGate() device-lock check — runs BEFORE the WebView mounts
└─ resolvePackagedBuild()
├─ ensureWebBundle() extract www.zip → documentDirectory/www-<version>/
├─ server.start() bind the pinned origin
└─ <WebViewHost> mounts only once the server is listening
└─ onTransportReady()
├─ startOutbox() schema → orphan recovery → prune → ledger → drain
├─ registerSession() attaches listeners SYNCHRONOUSLY (see below)
├─ registerOutbox() bridge request handlers
└─ checkVersionFloor() raises a condition; never blocks
Two orderings in that diagram are correctness requirements, not style:
- The session gate runs before the WebView mounts. Mounting first would let a device with no OS lock persist a session before anything checked (FR-018/FR-019).
- Session listeners attach synchronously inside
onTransportReady. The transport does not buffer notifications, and the web application restores its session as soon as its bundle runs — which beats SQLite opening. Registering after anawaitloses thesession.startedthe container exists to react to. This cost a day to find once; do not reintroduce it.
Packaged web build
scripts/build-mobile-web.sh builds frontend-pwa in --mode mobile and emits one
archive:
| Output | What it is |
|---|---|
www.zip | the whole Vite output, ~3.5 MB compressed / 11.1 MB raw |
manifest.json | webBuildVersion, sha256, sizes, git SHA |
Why one archive rather than loose files. EAS Update delivers assets through
Metro's asset graph, which requires every file to be statically require()d. Vite
emits content-hashed filenames, so that list cannot be maintained by hand. A single
archive sidesteps it while still travelling inside the signed update manifest, so it
inherits atomicity, hash verification, code signing and rollback (research D1).
webBuildVersion is YYYY.MM.DD+<pwaVersion>.<gitSha>. The floor check
orders only the zero-padded date prefix — a false positive there locks a merchant
out, and pwaVersion is not monotonic across branches. The OTA gate (embedded
binary vs channel head) also orders the dotted PWA version, because refusing an
older zip leaves the till on what it already has: 2026.09.11+0.0.383… must not
replace …+0.0.384….
The build mode is configuration only — no frontend-pwa source changes (FR-002).
It disables the service worker, targets es2017 for Sunmi System WebViews that never
update, and strips external <head> resources. The browser build is byte-identical
before and after (verified in T031).
Two guards in the build script fail the build rather than shipping a subtle problem:
a service worker in the output (it would become a second update mechanism racing EAS
Update over the same assets, and would win silently — research D3), and any remaining
fonts.googleapis.com / maps.googleapis.com / cdn.gpteng.co reference (they
would block an offline cold start, and a third-party script would execute inside the
session origin).
Extraction is atomic
web-bundle.ts extracts into a .tmp sibling and commits with a single rename,
writing the completeness marker last. A directory carrying a valid marker is
therefore complete by construction, and a crash mid-extract leaves only a
discardable .tmp — never a half-populated served root (FR-028). Extraction targets
documentDirectory, never the cache directory, which the OS may reclaim.
The cache buster is not optional
The WebView loads /index.html?b=<sha256[0:16]>. Because the origin and path are
pinned forever, /index.html is a permanent HTTP cache key: without the buster the
WebView serves the first build it ever fetched, and every delivered update
silently never runs, with no error anywhere (research D16).
The bridge
The only contract between container and web application is
packages/global/types/mobile-bridge.types.ts plus its companion
consts/mobile-bridge.consts.ts. Both sides compile against the same declarations —
two hand-synced copies is how version-skew guarantees rot. The mobile- prefix is
historical: since 054-desktop-runtime-shell the same contract is spoken by the
desktop Tauri host as well. Neither host is privileged, and the PWA never branches
on which one injected the bridge. How that generalises past printing:
Native Hardware Runtime.
Requests (web → native): handshake, diagnostics.echo, outbox.enqueue,
outbox.query.
Notifications: session.started, session.ended (web → native), condition
(native → web).
Errors: UNSUPPORTED_CAPABILITY, TIMEOUT, MALFORMED, REJECTED_ORIGIN,
STORAGE_UNAVAILABLE, INTERNAL.
Rules that make it survivable:
- Every envelope carries a correlation id (
cid), which continues intooutbox_job.correlation_idand every Sentry event. - Messages from any origin other than the pinned one are rejected.
- Every request has a deadline (10 s default) and answers exactly once — a handler settling after the deadline still produces no second response, because resolving a settled promise twice is worse than a timeout.
- Skew degrades in both directions: an unknown request type answers
UNSUPPORTED_CAPABILITY; an unrecognised incoming message type is ignored safely. Either side can therefore add message types without breaking the other, and adding one does not bumpBRIDGE_PROTOCOL_VERSION— only a breaking envelope change does. - The PWA-side client (
apps/frontend-pwa/src/lib/mobile-bridge.ts) is fully inert in a browser: no global installed, capabilities report null, notifications no-op, and a request rejects withNO_SHELLrather than hanging.
The container never draws merchant-facing text
Conditions — update_required, web_build_unsupported, device_lock_absent,
durable_storage_unavailable, work_needs_attention — are notified to the web
application, which renders them through its existing i18next catalogue. The PWA
already has the merchant's language, the design system and the accessibility work.
The container's own catalogue (apps/mobile/src/i18n/) holds exactly two strings —
the startup spinner and the "could not start" title — because those are the only
messages that can appear before the web application exists to show anything. Keep it
that way: anything the PWA could have rendered is rendered worse here.
Sessions
The container never holds a credential (FR-012). The web application owns sign-in; the container's only job is making that session durable, and only when the OS says the device is protected.
- Device lock is read with
getEnrolledLevelAsync(), notisEnrolledAsync()— the latter reports biometrics only, so a passcode-only device would wrongly read as unlocked. The probe fails closed: an unreadable level reportsnone. - The sentinel is deliberately unprotected. It records only that a protected
session exists and the
SecurityLevelseen at write time. That asymmetry is the whole mechanism: it survives the OS invalidation that erases its protected counterpart, which is what makes lock removal observable.SecureStorereturnsnullfor both "never stored" and "lock was removed", so the protected read alone cannot tell them apart. - No lock means degraded, not blocked. The session is not persisted, the
device_lock_absentcondition is raised, and the app stays fully usable for that run. - Teardown clears the sentinel last, so an interrupted teardown fails toward "re-authenticate" rather than silently restoring a half-deleted session.
- At rest:
NSFileProtectionCompleteUntilFirstUserAuthenticationon iOS. The stronger-soundingNSFileProtectionCompleteis wrong here — it needs an entitlement a free provisioning team cannot grant (so dev and release builds would diverge exactly where security behaviour is tested), and it would make the served 11 MB bundle unreadable after a screen lock, turning a routine lock into a failed launch for no threat-model gain.
Outbox
A durable, producer-agnostic queue proving at-most-once execution. In this
feature it accepts exactly one job type, diagnostic.noop — there is no real
producer and no printing whatsoever.
SQLite with journal_mode=WAL (not the default — set explicitly) and
synchronous=FULL. FULL rather than NORMAL because a WAL commit can be lost to an OS
crash under NORMAL, which would defeat the durable-before-acknowledge guarantee the
queue exists to provide.
- Every mutation uses
withExclusiveTransactionAsync. PlainwithTransactionAsyncswallows unrelated concurrent queries into the transaction — with a worker draining while the WebView enqueues, that is a live correctness bug. - De-duplication leans on a
UNIQUE(idempotency_key)constraint, not a priorSELECT: check-then-insert has a race a concurrent caller can lose, and losing it means the same work queued twice. - An interrupted job is failed, never retried. A job found
in_progressat startup died mid-attempt, and the container cannot know whether the side effect happened. Retrying is exactly what at-most-once forbids. It surfaces through thework_needs_attentioncondition instead. - Retry: 5 attempts, ×4 backoff (1 s, 4 s, 16 s, 64 s, 256 s), 10 s per-attempt timeout. Jobs older than 60 minutes are skipped with a visible count, never silently dropped.
- Retention:
doneafter 7 days or 200 rows,failedafter 30 days — pruned at startup, the one point where nothing is claimed and a delete cannot race a drain.
The durability ledger, and why row counts cannot prove anything
OUTBOX_LEDGER reports per-state counts and lifetime counters from an
outbox_counter table: accepted, completed, pruned, cleared,
double_completions. Each is bumped inside the same transaction as the row change it
describes.
The counters exist because per-state counts alone cannot detect loss. state is
CHECK-constrained to four values, so rows-present is identically the sum of the
four counts — a job destroyed by a crash disappears from both sides and cancels out.
An earlier revision of the durability matrix compared exactly those two numbers and
could therefore only ever report zero lost, including on a run that lost everything.
The invariants that do work:
accepted - pruned - cleared == rows present SC-010, no work lost
double_completions == 0 SC-011, at-most-once
markDone only matches a row that is not already done, so a second completion of
the same job — the sole trace a double execution leaves, since the row looks
identical either way — is counted rather than absorbed.
Releases
Container and packaged web build ship as one release. scripts/publish-mobile-release.sh
rebuilds the archive and publishes in a single command, so the two cannot diverge.
runtimeVersion: { policy: "fingerprint" }— native drift automatically forces a store build rather than depending on someone remembering to bump a string.fallbackToCacheTimeout: 0— the app launches on the build already on disk, and a downloaded update applies at the next cold start.checkAutomatically: NEVER— JS owns the check, so an embedded binary can refuse a channel head whosewebBuildVersionis older than the zip it already ships (the desktoppackagedvs SHA hole). NativeON_LOADwould download first and apply on the next cold start regardless. Rollback to embedded, and republish onto a device already on an OTA, still follow Expo.Updates.reloadAsync()is not called unattended. Reloading tears down the WebView, and the WebView holds the merchant's in-progress work. The unattended path therefore waits for the next cold start. A merchant who is ready to restart can tap Check for updates on the About page; that is a bridge request (updates.apply), not a timer in the container. A source-level guard test still forbidsreloadAsyncanywhere except that apply path.- Channels (
development,internal,merchant) are baked into the binary at build time, so publishing to one cannot reach another.
Signing: see certificate rotation. An expired certificate permanently bricks OTA for that binary.
Daily frontend-pwa deploys reach installed devices this way — they do not
need a store release. Native changes do. Product-level picture:
Printing topology.
First binaries onto TestFlight and Play:
Store submission.
Backend surface
Exactly two changes, both additive:
GET /api/v1/status/mobile-shell(@IsPublic()) — returnsminimumWebBuildVersion,minimumRuntimeVersion,storeUrlsfrom configuration. No table, no repository. A null floor means no constraint, and an unreachable endpoint means proceed — an unknown floor is not evidence of an unsupported build, and this app must work offline (FR-030).main.tsCORS allowlist gains the single literal originhttp://127.0.0.1:47285.isPrivateDevOriginis not enabled outside development; doing so would let any local server on any port make credentialed calls to the FlowPOS API.
What the container is not allowed to do
Feature 051 forbade printing so the shell could ship without it (verified in
T080). Feature 052 added native discovery and print transports under
apps/mobile/src/printing/; those are not the WebView. The rules that remain:
- No sign-in surface, no account picker, no token refresh, no session resurrection.
- No copy of any authentication credential (FR-012). Station claims are registered by the web application, not the container.
- No coupling to
frontend-pwainternals — routes, markup, styling, storage keys. The only contract is the bridge module and the pinned origin (verified in T086). - Nothing requiring the container to act on the merchant's behalf while the web application is not running (FR-068). Today's print jobs arrive from the WebView. A future Android agent mode would register as a device, like Print Bridge — that is a new feature, not a weakening of this rule. See Printing topology.
- Biome only.
create-expo-appscaffolds ESLint; it was removed and must stay removed.
Related
- Run locally — simulators and devices, logs, Metro, shipping backend / PWA / shell changes
- Certificate rotation runbook
- Store submission — App Store, Play, Sunmi
- Manual verification procedures — MV-1…MV-5
- Printing topology
- Vendor printing — Star SDK, Android SPP/USB, cash drawer
- Native Hardware Runtime — shared bridge with desktop; capabilities beyond printing
specs/051-mobile-app-shell/— spec, plan, research decisions D1–D16, tasksapps/mobile/CLAUDE.md— the standing Apple-account rule