Skip to main content

Mobile Shell — Manual Verification Procedures

Feature: 051-mobile-app-shell · Decided: 2026-09-03

Six checks (T028, T029, T042, T043, T078-iOS, T052) were specified as automated e2e tests. They are documented manual procedures instead, deliberately.

Why manual is the right call here, and where it stops being right. Each of these runs a handful of times, needs a physical device, and several need a human to change device state mid-run — setting and removing a passcode, toggling a radio — which no harness automates well. The cost of a harness is not repaid at this repetition count. MV-6 is manual for a second reason: automating a tool whose job is to publish mis-signed payloads at a release channel is a worse trade than writing the steps down.

That reasoning does not extend to everything, and four cases are scripted instead:

ScriptedWhy not manual
scripts/session-durability-android.sh (T038, 20 cycles)Twenty repetitions is past the point where a human runs it honestly
scripts/outbox-durability-android.sh (T078, 100+ cycles)Purely observational at high repetition — the case a harness is for
scripts/update-atomicity-android.sh (T053)Repeated publish/adopt cycles, read from one log line
scripts/update-rollback-android.sh (T057)A fixed sequence with exact assertions, including that queued work survives

The line between the two is repetition and human judgement, not importance.

Run these before any pilot release, and after any change to the shell's serving, session, or update paths. Getting a debug build onto a simulator or device first: Run locally. Physical hardware is still required for these procedures.


Before any run: confirm the build under test is the one running​

The shell serves a pinned origin and a pinned path, which makes /index.html a permanent HTTP cache key. A WebView that has cached an earlier build will keep serving it, with a healthy-looking app and no error anywhere — so a procedure run against a stale build proves nothing (research D16).

The entry URL carries ?b=<sha256[0:16]> of the packaged archive to defeat this. To confirm, compare the hash in the launch log with the archive you built:

adb logcat -d | grep SHELL_BUNDLE_READY      # version and extraction directory
node -p "require('./apps/mobile/assets/web/manifest.json').sha256.slice(0,16)"

On iOS, read the same line from the Xcode console.

MV-1 — Browser/app parity (T028, SC-001, FR-006)​

Asserts: every merchant task achievable in a mobile browser is achievable in the app, with equivalent outcomes.

  1. On a desktop or mobile browser, sign in to the same environment the app points at.
  2. Complete this journey, noting each outcome: sign in → select business/location → open the Sale module → search → open Dashboard → sign out.
  3. Repeat the identical journey in the app on a physical device.
  4. Compare outcome by outcome.

Pass: identical outcomes at every step. Fail: any step that succeeds in one and not the other, or produces different data.

Presentation differences (layout, safe-area insets, native keyboard) are not failures. This checks capability and outcome, not pixels.


MV-2 — Offline cold start (T029, SC-004, FR-003)​

Asserts: the packaged build renders with no network, and only genuinely network-dependent operations are unavailable.

  1. Launch the app once online so the web bundle is extracted.
  2. Enable airplane mode — Android: adb shell svc wifi disable; iOS: Control Centre.
  3. Force-close the app, then launch it.

Pass: the interface renders, and an unreachable API shows the app's own actionable offline state — "Can't reach FlowPOS… Retry". Fail: a blank screen, an indefinite spinner, or a native error surface competing with the web application's own.

A blank screen here previously meant an infinite /loading ↔ /sign-in redirect loop (research D12). If it recurs, that is a frontend-pwa regression, not a shell one.

⚠️ Use a release build. A debug build pulls its JS bundle from Metro — over USB on Android — so an airplane-mode test against a debug build is not genuinely offline.


MV-3 — Account-switch isolation (T042, SC-017, FR-021, FR-022)​

Asserts: signing in as a different account leaves nothing of the previous session reachable.

  1. Sign in as account A. Note the business and location shown on Dashboard.
  2. Sign out, then sign in as account B (a different business).
  3. Inspect: business/location shown, cached lists, and any queued work.

Pass: nothing from A is visible or reachable under B; the log records the previous session's teardown and any discarded queued items. Fail: A's business name, location, or cached data appears under B, or discarded work vanishes without a recorded reason.


MV-4 — Device-lock matrix (T043, SC-016, FR-018, FR-019, FR-020)​

Asserts: session durability is gated on an OS device lock, and lock removal ends the session.

Physical devices only — simulators and emulators do not reproduce keychain/keystore invalidation, which is the entire mechanism under test.

#Device stateActionExpected
1No passcode/biometricLaunch, sign inApp fully usable. Session not persisted. Merchant told why they will sign in each time. Log: persist=false condition=device_lock_absent
2Passcode setSign in, force-close, relaunchStill signed in. Log: persist=true
3Passcode set, session persisted, then remove the passcodeRelaunchSession ended, device-held data cleared. Log: end-session … device-lock-removed

Step 3 is the reason the unprotected sentinel exists. SecureStore returns null both for "never stored" and "invalidated because the lock was removed"; only the sentinel distinguishes them (data-model.md §2). If step 3 instead behaves like a fresh start, the sentinel is broken.

Restore the device passcode afterwards.


MV-5 — Outbox durability, iOS (T078 iOS half, SC-010, SC-011)​

Asserts: no acknowledged work is lost, and none is executed twice, across terminations.

Android runs this automatically at 100+ cycles via scripts/outbox-durability-android.sh. iOS has no adb equivalent — no scripted force-stop, reboot or airplane-mode toggle — so it is 20 cycles by hand, using the same checklist. Twenty is chosen to cover each termination mode several times, not because the guarantee is weaker on iOS.

Per cycle, rotating the termination mode as the Android script does (force-close ×7, reboot ×1, airplane mode ×1, app-update reload ×1 per ten):

  1. Note the ledger line the app logs at startup:

    OUTBOX_LEDGER present=N done=N failed=N pending=N in_progress=N \
    accepted=N completed=N pruned=N cleared=N doubles=N
  2. Terminate by this cycle's mode.

  3. Relaunch and read the ledger again.

  4. Check, every cycle — an end-state-only check would miss a job created and lost within a single cycle:

CheckMeaning
accepted - pruned - cleared equals presentNothing was lost (SC-010)
doubles is 0Nothing ran twice (SC-011)
completed never exceeds acceptedNothing ran twice (SC-011, second angle)
accepted never decreasesThe app was not reinstalled mid-run, which would reset the baseline
in_progress is 0 after startupOrphan recovery ran
Banner appears when failed > 0Interrupted work is visible, not silent (FR-054)

Do not substitute done + failed + pending + in_progress for present. An earlier revision of this procedure said exactly that, and the check could never fail: state is constrained to those four values, so their sum is the row count, and a job destroyed by a crash vanishes from both sides. The four lifetime counters exist because they do not move when a row is destroyed — which is the only reason a shortfall means anything.

Where the work comes from. This feature has no real producer (FR-058), so a development build enqueues one diagnostic.noop job per launch. Without it the ledger reads all zeros every cycle and the run reports a pass that proves nothing — check the first cycle shows a non-zero accepted before trusting any later one. Release builds never seed, so this is a development-build procedure by construction.

Pass: 20 cycles with no lost work and no double execution.

The banner check is the one most easily skipped and most worth keeping. Orphan recovery deliberately never retries an interrupted job, so if the banner does not appear, that work is invisible to the merchant — the silent loss FR-054 exists to prevent, and it would look identical to everything working.


MV-6 — Update signature rejection (T052, SC-018, FR-032, FR-033)​

Asserts: a tampered or wrongly-signed update payload is rejected, the device stays on its last known-good release, and the rejection is reported.

Why this one is a procedure and not a script. The other durability checks repeat dozens of times and observe passively; this runs a handful of times and requires publishing a deliberately bad release, generating a throwaway key pair, and then undoing all of it. Automating a tool whose job is to push mis-signed payloads at a release channel is a worse trade than writing the steps down — a script left in the repo is a script someone eventually runs against the wrong channel.

Blocked until T047. Code signing must actually be configured, which requires the paid plan tier decided in T046. Until then FR-032/FR-033 are knowingly unmet and this procedure cannot run at all — there is no signature to reject.

Safety rules​

  • internal channel only. Never publish an intentionally bad payload to merchant, and never to a channel any merchant-owned device is enrolled on.
  • Use a device you are willing to reinstall. A device left on a rejected-update loop is recovered with a fresh install from the channel.
  • The rogue key pair is throwaway. Generate it outside the repo, and delete it when finished. It must never reach Doppler or GCP Secret Manager, where it would sit indistinguishable from the real one.

Procedure​

  1. Establish known-good. Publish a normal release and let the device adopt it (two launches — an update downloads on one and applies on the next cold start). Record its SHELL_RELEASE line:

    SHELL_RELEASE updateId=… embedded=false runtimeVersion=… channel=internal \
    webBuildVersion=… webSha=…

    Everything below compares against this line. Note both halves: a rejection that somehow left the container on the good release but swapped the web build would be a worse outcome than a clean rejection, and only the paired line shows it.

  2. Generate a rogue key pair, in a scratch directory outside the repo:

    cd "$(mktemp -d)"
    npx expo-updates codesigning:generate \
    --key-output-directory keys --certificate-output-directory certs \
    --certificate-validity-duration-years 1 \
    --certificate-common-name "ROGUE — MV-6 test only"
  3. Publish signed with the rogue key. Point the publish at the rogue private key rather than the configured one — the certificate embedded in the installed binary is the real one, so the device must refuse this payload:

    cd apps/mobile
    eas update --channel internal --message "MV-6 rogue signature" \
    --private-key-path /path/to/scratch/keys/private-key.pem
  4. Attempt adoption. Launch, wait ~15 s, force-close, launch again.

  5. Check all four:

CheckWhereMeaning
SHELL_RELEASE still reports the step-1 updateId and webBuildVersionlogcat / ConsoleDevice stayed on last known-good — neither half moved (SC-018)
The app launched normally and is usableon screenA rejected update must not brick the app; it is a non-event to the merchant
A rejection is reported in Sentry, tagged with this releaseSentryFR-033 — an unreported rejection is indistinguishable from no attack
SHELL_UPDATE_SYNC error appearslogcat / ConsoleThe failure surfaced rather than being swallowed
  1. Also test outright tampering (optional but cheap): repeat with the signature header stripped entirely rather than wrong. Both must be rejected — "unsigned" and "wrongly signed" are different code paths, and only one of them is the obvious one.

  2. Clean up. Republish the known-good release to internal, confirm the device adopts it, then delete the scratch directory and its keys.

Pass: the device never runs the rogue payload, stays fully usable on the prior release, and the rejection is visible in telemetry.

The check most likely to be skipped is the Sentry one, and it is the one that matters most in production. A device that silently rejects updates looks exactly like a device with nothing to update — so without the report, a compromised update path and a healthy one are the same picture.