5.3 KiB
SITE-31 — Server Decision Record (ADR)
Status: PROPOSED — transport/server unbuilt; client seams implemented and tested
Date: 2026-09-27
Decides: which server the site app syncs against, what authority it has
per data class, and what the client may assume before it exists.
1. Decision
Extend the tracked nimanyatta/ tree (pinned sibling commit
84fe56de61ee39f464a18f4aa575307a2966d3b6, see
nigig-site-core/src/interop.rs::SERVER_PIN_COMMIT) into the scope §5
backend: Rust API service + PostgreSQL + S3-compatible object storage, with
delta sync, resumable media upload, server arbitration for shared-document
conflicts, and organisation backup/export.
2. Why this server and not another
- The tree is already in this repository, with an auth model, a database layer, broadcast/realtime paths, and REST coverage repaired in-branch.
- Its WebSocket chat protocol (
nimanyatta/crates/nimanyatta-protocol, postcard) is a workspace member here and its wire format is pinned by characterization tests — the only app/server surface with that property. - Its HTTP sync DTOs (
SyncRequest/SyncResponse/timelines) are readable in-tree vianigig-lite/crates/common; the field-level contract (body, externally-taggedMessage) is verified by the extracted harness, not assumed.
3. Per-data-class authority (the E2EE divergence, resolved)
Scope §6 demands server-side authority for approvals, signatures, locking and roles; SITE-08/09 demand end-to-end protection where the server must not read payloads. Both hold, on different data:
| Data class | Authority | Rationale |
|---|---|---|
| Approvals, signatures, locks, roles, membership | Server-arbitrated. Server enforces, both sides check. | Tamper-evidence requires a party both endpoints can blame. |
| Report/worker/PII application payloads in transit | E2EE where reviewed (e2ee envelopes); otherwise local-only. |
Server sees routing metadata + ciphertext lengths only. |
| Chat transport until E2EE review passes | Absent. Local cache + offline outbox only. | No TLS-only/plain-room substitute ships (SITE-P0-05). |
| Conflicts on shared documents | Server arbitration; loser flagged for the Report Master. | Deterministic order beats wall-clock voting. |
4. Explicit non-decisions (still open, owned elsewhere)
- Hosting model and data residency (scope §18 Q1): option per organisation; no default asserted here.
- AI/STT provider and budget (scope §18 Q3): no provider named; no audio leaves a device without recorded consent regardless of provider.
- Workspace membership of
nimanyatta/nigig-common: both carry deliberate independent[workspace]tables; absorbing them is the server team's decision, not a mechanical change. Until then the client pins by file-level drift detection (interop::assert_pinned_tree), never by importing the server workspace. - Client portal timing (scope §18 Q5): digest-by-share first (R4), portal after this server exists (R5).
5. Consequences for the client
integrations.rsintents (share, calendar, webhooks, API scopes, backup manifest, remote lookup) are validated locally and dispatched nowhere: every dispatch names its missing provider. No caller can mistake validation for delivery.- Webhook/API traffic authenticates against the SITE-20 capability matrix on both ends; the server enforces it independently.
- Backup/restore drills run against this server once it serves the API; until then the local encrypted operation journal is the durability story and no sync claim ships.
6. Reconnect discipline (thundering-herd defense)
A server restart that drops 100,000 WebSocket clients must not be answered by 100,000 immediate reconnects — that is a self-inflicted traffic attack, and "reconnect immediately" in any client is a bug.
Client rules (implemented and tested in net_client::RetryPolicy):
- Every retry waits an exponential backoff (
500 msbase,30 scap, 8-attempt budget) with deterministic equal jitter, so a fleet that failed together does not retry together. No RNG dependency: the spread derives from the operation id, exact in tests. - A server
Retry-Afteralways wins within the cap; a missing or absurd value can neither park a client forever nor be ignored. - Bulk replays drain paced: journal catch-up and outbox flush go out in
SYNC_DRAIN_BATCH-sized batches (chat::drain_outbox_paced), spaced by the retry policy — never as one burst per reconnect. - Reconnect reuses the existing session/token; it never re-runs login or re-fetches full state first (no AUTH stampede, no thundering snapshot).
Server expectations (specified here, owned by the server team):
- Refuse new upgrades while warming, with
Retry-After, instead of accepting connections that cannot be served yet (readiness gate). - Shed gracefully past capacity: bound connection count, evict oldest-idle first, and say so in metrics — never accept-then-blackhole.
- Recover staggered: per-device sequence resumption (already in the envelope design) rather than mass snapshot pushes.
- Load-test scenario to add once the server harness runs here: drop N clients simultaneously, assert p99 reconnect delay spread covers the backoff window, AUTH rate stays flat, and no client retries before its computed delay.