nigig-org/crates/apps/nigig-site/SITE_31_SERVER_ADR.md

5.3 KiB

SITE-31 — Server Decision Record (ADR)

Status: PROPOSED — transport/server unbuilt; client seams implemented and tested Date: 2026-09-27 Decides: which server the site app syncs against, what authority it has per data class, and what the client may assume before it exists.

1. Decision

Extend the tracked nimanyatta/ tree (pinned sibling commit 84fe56de61ee39f464a18f4aa575307a2966d3b6, see nigig-site-core/src/interop.rs::SERVER_PIN_COMMIT) into the scope §5 backend: Rust API service + PostgreSQL + S3-compatible object storage, with delta sync, resumable media upload, server arbitration for shared-document conflicts, and organisation backup/export.

2. Why this server and not another

  • The tree is already in this repository, with an auth model, a database layer, broadcast/realtime paths, and REST coverage repaired in-branch.
  • Its WebSocket chat protocol (nimanyatta/crates/nimanyatta-protocol, postcard) is a workspace member here and its wire format is pinned by characterization tests — the only app/server surface with that property.
  • Its HTTP sync DTOs (SyncRequest/SyncResponse/timelines) are readable in-tree via nigig-lite/crates/common; the field-level contract (body, externally-tagged Message) is verified by the extracted harness, not assumed.

3. Per-data-class authority (the E2EE divergence, resolved)

Scope §6 demands server-side authority for approvals, signatures, locking and roles; SITE-08/09 demand end-to-end protection where the server must not read payloads. Both hold, on different data:

Data class Authority Rationale
Approvals, signatures, locks, roles, membership Server-arbitrated. Server enforces, both sides check. Tamper-evidence requires a party both endpoints can blame.
Report/worker/PII application payloads in transit E2EE where reviewed (e2ee envelopes); otherwise local-only. Server sees routing metadata + ciphertext lengths only.
Chat transport until E2EE review passes Absent. Local cache + offline outbox only. No TLS-only/plain-room substitute ships (SITE-P0-05).
Conflicts on shared documents Server arbitration; loser flagged for the Report Master. Deterministic order beats wall-clock voting.

4. Explicit non-decisions (still open, owned elsewhere)

  • Hosting model and data residency (scope §18 Q1): option per organisation; no default asserted here.
  • AI/STT provider and budget (scope §18 Q3): no provider named; no audio leaves a device without recorded consent regardless of provider.
  • Workspace membership of nimanyatta/nigig-common: both carry deliberate independent [workspace] tables; absorbing them is the server team's decision, not a mechanical change. Until then the client pins by file-level drift detection (interop::assert_pinned_tree), never by importing the server workspace.
  • Client portal timing (scope §18 Q5): digest-by-share first (R4), portal after this server exists (R5).

5. Consequences for the client

  • integrations.rs intents (share, calendar, webhooks, API scopes, backup manifest, remote lookup) are validated locally and dispatched nowhere: every dispatch names its missing provider. No caller can mistake validation for delivery.
  • Webhook/API traffic authenticates against the SITE-20 capability matrix on both ends; the server enforces it independently.
  • Backup/restore drills run against this server once it serves the API; until then the local encrypted operation journal is the durability story and no sync claim ships.

6. Reconnect discipline (thundering-herd defense)

A server restart that drops 100,000 WebSocket clients must not be answered by 100,000 immediate reconnects — that is a self-inflicted traffic attack, and "reconnect immediately" in any client is a bug.

Client rules (implemented and tested in net_client::RetryPolicy):

  • Every retry waits an exponential backoff (500 ms base, 30 s cap, 8-attempt budget) with deterministic equal jitter, so a fleet that failed together does not retry together. No RNG dependency: the spread derives from the operation id, exact in tests.
  • A server Retry-After always wins within the cap; a missing or absurd value can neither park a client forever nor be ignored.
  • Bulk replays drain paced: journal catch-up and outbox flush go out in SYNC_DRAIN_BATCH-sized batches (chat::drain_outbox_paced), spaced by the retry policy — never as one burst per reconnect.
  • Reconnect reuses the existing session/token; it never re-runs login or re-fetches full state first (no AUTH stampede, no thundering snapshot).

Server expectations (specified here, owned by the server team):

  • Refuse new upgrades while warming, with Retry-After, instead of accepting connections that cannot be served yet (readiness gate).
  • Shed gracefully past capacity: bound connection count, evict oldest-idle first, and say so in metrics — never accept-then-blackhole.
  • Recover staggered: per-device sequence resumption (already in the envelope design) rather than mass snapshot pushes.
  • Load-test scenario to add once the server harness runs here: drop N clients simultaneously, assert p99 reconnect delay spread covers the backoff window, AUTH rate stays flat, and no client retries before its computed delay.