Skip to content

RFC: Hardening the SQLite persistence / Electric sync stack #1659

Description

@KyleAMathews

RFC: Hardening the SQLite persistence / Electric sync stack

Status: Closed after the focused persistence correctness and readiness sweep. All 15 issues open in the original audit are now closed: #1416 was addressed by the opt-in readiness implementation in #1955/#1960, and the old #1615 readiness PR was closed unmerged as superseded. #82/#865 were closed after the existing offline-transactions outbox and persisted synced data shipped; #1883 was fixed by #1936. #1486 remains closed pending a current-path reproduction and #1567 remains closed as a separate recovery-feature choice. The acknowledgment, cross-tab ownership, and corrupt-file recovery questions below are preserved as explicit follow-up policies or capabilities, not unfinished fixes in this RFC.
Status audit: 2026-09-30, published main at c879ba6d855914e5c0cac49af7a007eb5cabe7e4 (includes #1955, #1936, and #1960).
Original audit baseline: 2026-09-16, 3b991173f828f28b7e32203decc8b4bad0f258c9.
Scope: Correctness, durability, recovery, driver conformance, and type composition. A merged fix is not by itself proof that every older reporter environment is covered. Product-policy changes remain decision-gated.

1. Current state

The September 16 inventory had 15 open issue IDs and four open PRs. All 15 original issue IDs are now closed. #1416 closed as addressed by the opt-in offline-rendering path in #1955/#1960; its older exact reporter environment was not rerun, and the fix does not make Collection readiness local-first. The related #1615 PR closed unmerged because its unconditional markReady mechanism conflicts with the chosen separate-readiness contract. Other reports were closed after fixes, bounded current-path evidence, or explicit scope decisions; closure does not establish every older reporter environment.

Area Landed since the original audit What remains
D1 — reset and resume integrity, #1589 #1846 atomically resets baseline-coupled metadata, certifies SQLite resume snapshots, and requests a fresh Electric snapshot when evidence is uncertain. #1910 adds a live Chromium two-tab OPFS + Electric witness with divergent schema versions, follower-to-leader handoff, and poisoned empty/stale-resume upgrades. #1589 closed after #1910. This bounded browser fixture does not prove React Native, Firefox/Zen, React rendering, large datasets, or exclusive OPFS ownership.
D2/D3/D13 — per-collection coordinator routing, clone-safe subset requests, and exact subset ownership, #1589, #1753, #1498 #1845 routes complete transactions through the elected adapter for the exact collection, validates wire values, and tracks exact subset leases. #1868 repairs adapter registration across unsubscribe/restart. #1905 adds a real Chromium two-tab OPFS oracle: the follower posts a clone-safe remote-subset request, makes no upstream Query call, receives the public row, and rejects an invalid nested wire value before another post. #1498 was closed after #1905. #1905 does not run React, Zen/Firefox, or a live Electric backend; it does not prove those host paths. #1753 is closed: routing, reset, and clone faults are repaired, and the reported Electric-backed mutation path does not append coordinator persistence to its user handler. A distinct local-only partial-success contract remains in D8. #1910 now covers the reported Chromium two-tab live Electric reset/resume history. Exclusive-handle topology is separately D10.
D4 — op-sqlite result decoding, #1499 #1848 decodes supported row and columnar carriers and rejects unknown shapes. #1499 closed. Native iOS/Android bridge receipts remain an assurance gap, not a reason to relabel the closed issue as open.
D6 — readiness and first coordinated write, #1416, #1443 #1899 removes the startup election/scheduler cycle; #1443 closed. #1860 restores warm on-demand query readiness. #1955 adds separate persistedStatus, isPersistedReady, and persistedError through the live-query observer and all framework integrations, plus collection-level opt-in network-first initial rendering for React and Solid Suspense. #1960 lets React Suspense render completed persisted query data after a client-stream failure while still surfacing derived-query failures. #1416 closed as addressed; #1615 closed unmerged as superseded. Collection isReady and ordinary preload() remain upstream-gated. Every eager persisted source in a query must opt in; an empty restore counts as ready. The older reporter's exact Solid environment was not rerun, and mixed queries without all-source opt-in remain unavailable for fallback.
D7 — source transaction durability and hydration races, #1456, #1754 #1822 fences stale offline replay reads. #1853 gives authoritative source transactions FIFO publication/durability, observable failure, lifecycle fencing, and explicit rejection of unsupported hydration-crossing schedules. #1911 adds a live Chromium/OPFS/Electric hydration-straddle witness and a late-buffer mutant; #1754 is closed. #1914 adds a test-only live Chromium immediate-reload witness: the first page closes before its OPFS write, the reopened raw OPFS snapshot is empty before Electric starts, and the exact row returns from Electric without a new user insert. #1456 is closed. #1911 does not cover every Electric schedule. #1914 does not prove the reporter's Edge/Windows host, offline recovery, or a durable-acknowledgment contract for awaitTxId; those limits are recorded separately from the closed report. Durable pending mutations are D8.
Adjacent core Collection settlement #1907 repairs accepted-delete ownership after an accepted dependent edit and covers both delete/edit settlement orders, both reinsert outcomes, and truncate/no-truncate in an eight-case RED/GREEN oracle. This core-state repair does not prove a persistence-host or live Electric path. Its full-Transaction retention concern is recorded but remains unquantified.
D8 — offline mutation persistence, #82, #865, and local-only partial success (historical #1456/#1753 context) @tanstack/offline-transactions already provides an offline outbox, persisted synced data has shipped, and #82/#865 were closed on 2026-09-29. #1837 repaired offline-runtime replay and serializer paths. This RFC does not need a new outbox feature. Separate contract questions remain in the existing offline-transactions path: provider success followed by failed durable acknowledgment, retry/idempotency, unknown encodings, and restored optimistic overlays. In sync-absent/local-only mode, a successful app handler followed by coordinator failure currently rejects the mutation and rolls back optimism; changing that caller-result policy needs a separate decision. Do not treat these as reasons to reopen #82/#865 by default.
D9 — SQLite expression indexes, old PR #1487 #1867 makes index DDL and runtime predicates use the same validated expression shape and rebuilds stale physical indexes. #1487 was closed unmerged as superseded by #1867, with a note crediting its author. Native-host planner behavior remains unverified.
D10 — browser multi-tab OPFS ownership, #1486 (historical #1753 context) #1845 strengthens elected write ownership; #1844 releases OPFS workers on pagehide. A current-main investigation did not reproduce #1486, so it was closed pending a better reproduction. #1936 bounds openBrowserWASQLiteOPFSDatabase() with a default timeout and optional abort signal; #1883 closed after that fix. Bounded opening does not establish a single cross-tab owner for every OPFS read and write. Revisit exclusive-handle topology only with a current reproduction or a chosen new ownership contract. #1883's stalled-open report is resolved; it is not proof of general ownership.
D11 — persistence type composition, #1452, old PR #1560 #1866 preserves schema inference and aligns the Expo SQLite driver with vendor overloads. #1452 closed; #1560 closed without merge. No known D11 code task remains. Retain the type matrix as regression coverage.
D12 — corruption recovery, #1567 Baseline certification in #1846 avoids trusting incomplete data after supported resets; it is not a corrupt-file recovery API. A diagnostic reproduced SQLITE_CORRUPT/SQLITE_NOTADB startup failures; #1567 was closed because automatic quarantine/rebuild needs a new path-owning API and an application-level rebuildability contract. No automatic corrupt-file recovery is claimed. Revisit as a separately chosen feature, not an outstanding fix to the existing caller-owned-handle factory.
Migration idempotence, #1711 #1900 accepts narrowly matched Tauri string or Error duplicate-column failures and preserves data on reopen. #1711 closed. The test simulates Tauri's error boundary with Better SQLite; a native Tauri-host run remains an assurance gap.
Shared-driver cold-start scheduling, #1752 #1868 provides a fair shared-driver hydrate lane (one regular operation between queued hydrations) and tests Browser/OPFS ordering. #1910 routes ordinary source commits through that scheduler before the database-wide writer lock. #1916 batches distinct-key, non-delete cold full replacements while preserving transactional rows, metadata, resume position, and sequential fallback; its Cloudflare path respects the 100-parameter limit. #1752 is closed because this removes the identified per-row worker-call cost. The bounded test shows 1,033 database calls before versus at most 40 after for a 205-row replacement. Chromium/OPFS multi-collection latency was not measured; closure does not claim a latency number or preempt an already-running write.

2. Release-gating invariants

  1. A destructive reset clears the rows, tombstones, applied transactions, and sync metadata it invalidates in one transaction.
  2. A sync adapter never trusts a resume point without a compatible, complete local baseline.
  3. Collection A's configuration or traffic cannot read, write, reset, or acknowledge work for collection B.
  4. Every coordinator message is clone-safe, routed to one collection, and reports success only after the requested work happened.
  5. A known driver result shape is decoded losslessly; an unknown shape is an error, never an empty result.
  6. Readiness has one explicit policy and settles exactly once across hydration, upstream readiness, failure, and cleanup.
  7. An accepted transaction becomes durable or replayable, or produces an observable failure. It cannot silently disappear on restart.
  8. Index DDL and indexed runtime predicates have the same SQLite expression shape.
  9. Exactly one supported owner accesses an exclusive OPFS database, including during leadership changes.
  10. Schema migrations are idempotent and corruption recovery cannot inherit stale sync state.

The landed PRs establish deterministic evidence for many of these laws. Invariant 6 now has a separate persisted-readiness signal and an opt-in network-first initial-render policy; it does not redefine Collection readiness. The pending-mutation portion of invariant 7 still needs explicit acknowledgment/retry policy in the existing offline-transactions system. Exclusive OPFS ownership (9) and automatic corrupt-file recovery (10) are unclaimed, deferred product capabilities rather than active fixes under closed #1486/#1567. Do not infer a real-host guarantee from a simulated driver boundary.

3. Oracle-first rule for remaining work

Follow docs/contributing/oracle-tests.md and the coverage map. A new fix starts with a reproducible RED witness on current main and leaves that witness plus generalized histories GREEN. Name and demonstrate: (1) the authoritative law, (2) legal/adversarial history, (3) an independent oracle, (4) the production path or fixture, (5) the exact observation, and (6) reach, fault-control, replay, and cleanup evidence. Do not rewrite a passing oracle's expected behavior to match a policy change before the maintainer chooses that policy.

For open reporter issues, rerun the claimed failure mechanism against current packages. Close the original report when a relevant current-path witness and its surrounding checks establish the repair; record adjacent host and policy limits separately instead of leaving every historical report open. A host-specific guarantee still needs evidence on that host. #1443, #1498, and #1456 are closed after focused witnesses; their closures do not claim every browser or native environment.

4. Completed work and bounded follow-ups

  1. Track closed-report limits and the separate caller-result decision: fix(sqlite): avoid writer-lock inversion during live Electric recovery #1910 covers Persisted Electric collection ends up permanently empty: diverging schemaVersions wipe rows, schema reset leaves the resume point behind #1589's divergent-schema-version and poisoned stale-resume histories in two Chromium tabs with OPFS and live Electric; Persisted Electric collection ends up permanently empty: diverging schemaVersions wipe rows, schema reset leaves the resume point behind #1589 is closed. test(browser-sqlite): cover live Electric hydration straddle #1911 covers one live Electric+OPFS begin-during-hydration/commit-after-close history; Persisted collections: sync transaction straddling a hydrate-window close is buffered but never flushed (lost update) #1754 is closed. test(browser-sqlite): cover Electric insert across immediate reload #1914 covers one immediate-reload history with an empty pre-Electric OPFS checkpoint and no reinsert; Data not persisted locally when using @tanstack/browser-db-sqlite-persistence + @tanstack/electric-db-collection doesn't persists data locally #1456 is closed. BrowserCollectionCoordinator: per-collection schemaVersions cause reset loops, and sync-ingested writes bypass leadership #1753 was closed after a current-main probe separated sync-absent from sync-present behavior: only the local-only wrapper appends coordinator persistence to the user handler. The original Electric-backed allegation was not reproduced. Decide the separate local-only caller-result contract in D8 before changing behavior. Track Edge/Windows and other host limits without reopening Data not persisted locally when using @tanstack/browser-db-sqlite-persistence + @tanstack/electric-db-collection doesn't persists data locally #1456 by default.
  2. Record the completed cold-replacement fix: perf(sqlite): batch cold replacement writes #1916 reduces the identified per-row SQLite worker-call cost with atomic batching and a sequential fallback; Cold-start persist storm blocks hydrates for seconds on the shared SQLite driver (browser/OPFS) #1752 is closed. Real Chromium/OPFS latency remains unmeasured, but no latency measurement is required to close that root-cost report. Revisit a distinct long-write delay only with a fresh reproduction.
  3. Separate shipped readiness from remaining policy choices: feat(db): add network-first initial rendering for persisted collections #1955/fix: render persisted data in Suspense after client stream errors #1960 provide opt-in persisted readiness and React/Solid initial-render fallback without changing ordinary Collection readiness. useLiveQuery blocks rendering of locally persisted collection when Electric is unavailable #1416 is closed as addressed and the unmerged fix(db-sqlite-persistence-core): persisted preload() hangs when upstream sync never calls markReady #1615 is closed as superseded. Offline-First Support #82/Persistence of synced data #865 are closed; do not add a second outbox. Decide any existing offline-transactions acknowledgment and local-only partial-success contract in separate follow-up work before changing caller behavior. fix(browser-sqlite): bound OPFS database opening #1936 fixed the bounded-open bug and openBrowserWASQLiteOPFSDatabase() never settles while a frozen tab holds the database #1883 is closed. Cross-tab ownership and corrupt-file quarantine/rebuild remain deferred feature/API choices under closed BrowserCollectionCoordinator: secondary tabs fail to acquire leadership in multi-tab sessions #1486/SQLite persistence: recover or provide helper for corrupt local DB files #1567.
  4. Track proof limits without reopening closed issues by default: fix(react-native): decode op-sqlite results losslessly #1848 lacks native op-sqlite device receipts, fix(sqlite-persistence): accept Tauri string migration errors #1900 lacks a native Tauri-host run, and fix(sqlite): enforce crash-only persistence coordination #1845/Fix persisted collection durability and lifecycle races #1853 rely partly on controlled Browser/Electron/Electric fixtures. test(browser-sqlite): cover two-tab remote subset transport #1905 and fix(sqlite): avoid writer-lock inversion during live Electric recovery #1910 do not test the reporter's Zen setup; fix(sqlite): avoid writer-lock inversion during live Electric recovery #1910 also does not establish React Native reset recovery. Add host coverage when making a host-specific claim or when a current reproduction demands it.

5. Resolved and deferred maintainer decisions

  1. Resolved opt-in readiness policy (closed useLiveQuery blocks rendering of locally persisted collection when Electric is unavailable #1416; superseded, closed fix(db-sqlite-persistence-core): persisted preload() hangs when upstream sync never calls markReady #1615): feat(db): add network-first initial rendering for persisted collections #1955 exposes per-query persistedStatus, isPersistedReady, and persistedError through all framework integrations. Each eager persisted source Collection opts into initialRender: { strategy: 'network-first', networkTimeoutMs }; React and Solid Suspense prefer Collection readiness, then permit a completed restore after the network deadline or source failure. networkTimeoutMs is a network-preference period, not a total restore timeout. fix: render persisted data in Suspense after client stream errors #1960 covers failed React client streams and preserves derived-query errors. Empty restored results count; mixed queries with an unconfigured source remain unavailable for fallback. This supplies an official offline-rendering path without changing Collection.isReady or ordinary preload(). useLiveQuery blocks rendering of locally persisted collection when Electric is unavailable #1416 was closed on that basis; fix(db-sqlite-persistence-core): persisted preload() hangs when upstream sync never calls markReady #1615's unconditional-ready mechanism was not merged.
  2. Existing offline outbox and partial-success acknowledgments (closed Offline-First Support #82/Persistence of synced data #865; historical Data not persisted locally when using @tanstack/browser-db-sqlite-persistence + @tanstack/electric-db-collection doesn't persists data locally #1456/BrowserCollectionCoordinator: per-collection schemaVersions cause reset loops, and sync-ingested writes bypass leadership #1753 context): @tanstack/offline-transactions exists, so no integrated outbox feature is pending under those closed issues. The separate unresolved contract is what caller success, retry, and durable bytes mean when provider work succeeds but outbox deletion fails. Decide idempotency, rejection/conflict, unknown encodings, and restored optimistic overlays in that package's own follow-up work. The RFC discussion's server-success/failed-deletion history and the sync-absent handler-success/coordinator-failure history remain useful acceptance cases.
  3. Deferred browser multi-tab ownership (BrowserCollectionCoordinator: secondary tabs fail to acquire leadership in multi-tab sessions #1486; historical BrowserCollectionCoordinator: per-collection schemaVersions cause reset loops, and sync-ingested writes bypass leadership #1753 context): BrowserCollectionCoordinator: secondary tabs fail to acquire leadership in multi-tab sessions #1486 is closed pending a better current-path reproduction. fix(browser-sqlite): bound OPFS database opening #1936 separately fixed the frozen-holder stall by bounding database opening; openBrowserWASQLiteOPFSDatabase() never settles while a frozen tab holds the database #1883 is closed. Neither establishes a single owner for all OPFS access. If a new ownership contract is prioritized, choose a SharedWorker owner or leader-RPC ownership for all access, including follower initialization and reads; exercise open/close, leader loss, delayed messages, and concurrent writes in real multi-context fixtures.
  4. Deferred corruption recovery (SQLite persistence: recover or provide helper for corrupt local DB files #1567): SQLite persistence: recover or provide helper for corrupt local DB files #1567 is closed. A future opt-in quarantine/rebuild feature needs an owner of path opening, validation, handle closing, file moves and sidecars, plus an application declaration that the whole database is rebuildable. Its contract should retry once, propagate a second failure, and leave unrelated errors untouched.

6. Deferred architecture

A replica manifest, broader orthogonal storage/hydration/sync/durability status beyond the shipped persisted-readiness signal, generation-fenced reset operations, a database-level storage broker, authority classification, operational repair APIs, and a typed persistence-composition facade remain future directions. They are not prerequisites for the already-landed correctness work.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions