# SITE-31 — Server Decision Record (ADR) **Status:** `PROPOSED — transport/server unbuilt; client seams implemented and tested` **Date:** 2026-09-27 **Decides:** which server the site app syncs against, what authority it has per data class, and what the client may assume before it exists. ## 1. Decision Extend the **tracked `nimanyatta/` tree** (pinned sibling commit `84fe56de61ee39f464a18f4aa575307a2966d3b6`, see `nigig-site-core/src/interop.rs::SERVER_PIN_COMMIT`) into the scope §5 backend: Rust API service + PostgreSQL + S3-compatible object storage, with delta sync, resumable media upload, server arbitration for shared-document conflicts, and organisation backup/export. ## 2. Why this server and not another - The tree is already in this repository, with an auth model, a database layer, broadcast/realtime paths, and REST coverage repaired in-branch. - Its WebSocket chat protocol (`nimanyatta/crates/nimanyatta-protocol`, postcard) is a workspace member here and its wire format is pinned by characterization tests — the only app/server surface with that property. - Its HTTP sync DTOs (`SyncRequest`/`SyncResponse`/timelines) are readable in-tree via `nigig-lite/crates/common`; the field-level contract (`body`, externally-tagged `Message`) is verified by the extracted harness, not assumed. ## 3. Per-data-class authority (the E2EE divergence, resolved) Scope §6 demands server-side authority for approvals, signatures, locking and roles; SITE-08/09 demand end-to-end protection where the server must not read payloads. Both hold, on different data: | Data class | Authority | Rationale | |---|---|---| | Approvals, signatures, locks, roles, membership | **Server-arbitrated.** Server enforces, both sides check. | Tamper-evidence requires a party both endpoints can blame. | | Report/worker/PII application payloads in transit | **E2EE where reviewed** (`e2ee` envelopes); otherwise local-only. | Server sees routing metadata + ciphertext lengths only. | | Chat transport until E2EE review passes | **Absent.** Local cache + offline outbox only. | No TLS-only/plain-room substitute ships (SITE-P0-05). | | Conflicts on shared documents | Server arbitration; loser flagged for the Report Master. | Deterministic order beats wall-clock voting. | ## 4. Explicit non-decisions (still open, owned elsewhere) - **Hosting model and data residency** (scope §18 Q1): option per organisation; no default asserted here. - **AI/STT provider and budget** (scope §18 Q3): no provider named; no audio leaves a device without recorded consent regardless of provider. - **Workspace membership of `nimanyatta`/`nigig-common`:** both carry deliberate independent `[workspace]` tables; absorbing them is the server team's decision, not a mechanical change. Until then the client pins by file-level drift detection (`interop::assert_pinned_tree`), never by importing the server workspace. - **Client portal timing** (scope §18 Q5): digest-by-share first (R4), portal after this server exists (R5). ## 5. Consequences for the client - `integrations.rs` intents (share, calendar, webhooks, API scopes, backup manifest, remote lookup) are validated locally and dispatched nowhere: every dispatch names its missing provider. No caller can mistake validation for delivery. - Webhook/API traffic authenticates against the SITE-20 capability matrix on both ends; the server enforces it independently. - Backup/restore drills run against this server once it serves the API; until then the local encrypted operation journal is the durability story and no sync claim ships. ## 6. Reconnect discipline (thundering-herd defense) A server restart that drops 100,000 WebSocket clients must not be answered by 100,000 immediate reconnects — that is a self-inflicted traffic attack, and "reconnect immediately" in any client is a bug. **Client rules (implemented and tested in `net_client::RetryPolicy`):** - Every retry waits an exponential backoff (`500 ms` base, `30 s` cap, 8-attempt budget) with deterministic equal jitter, so a fleet that failed together does not retry together. No RNG dependency: the spread derives from the operation id, exact in tests. - A server `Retry-After` always wins within the cap; a missing or absurd value can neither park a client forever nor be ignored. - Bulk replays drain paced: journal catch-up and outbox flush go out in `SYNC_DRAIN_BATCH`-sized batches (`chat::drain_outbox_paced`), spaced by the retry policy — never as one burst per reconnect. - Reconnect reuses the existing session/token; it never re-runs login or re-fetches full state first (no AUTH stampede, no thundering snapshot). **Server expectations (specified here, owned by the server team):** - Refuse new upgrades while warming, with `Retry-After`, instead of accepting connections that cannot be served yet (readiness gate). - Shed gracefully past capacity: bound connection count, evict oldest-idle first, and say so in metrics — never accept-then-blackhole. - Recover staggered: per-device sequence resumption (already in the envelope design) rather than mass snapshot pushes. - Load-test scenario to add once the server harness runs here: drop N clients simultaneously, assert p99 reconnect delay spread covers the backoff window, AUTH rate stays flat, and no client retries before its computed delay.