nigig-org/REVIEWS/adr/0038-pdf-text-reflow-cache-documentstructure.md
andodeki ac8f8aa002
Some checks failed
p2p-intel / engine (push) Waiting to run
p2p-intel / notifications (push) Waiting to run
p2p-intel / coverage (push) Waiting to run
p2p-intel / makepad-app (push) Waiting to run
p2p-intel / exchange-tab (push) Waiting to run
Payment domain, storage, platform and UI / isolated-payment-tests (push) Waiting to run
Payment domain, storage, platform and UI / payment-ui-tests (push) Waiting to run
repo hygiene / hygiene (push) Has been cancelled
PDF engine / engine (push) Has been cancelled
PDF engine / makepad-integration (push) Has been cancelled
PDF engine / fuzz (push) Has been cancelled
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
traffic / gates (push) Has been cancelled
traffic / nigig-traffic (push) Has been cancelled
traffic / supply-chain (push) Has been cancelled
doc-engine / engine (push) Has been cancelled
doc-engine / coverage (push) Has been cancelled
doc-engine / consumer (push) Has been cancelled
email / gates (push) Has been cancelled
email / email-domain (push) Has been cancelled
email / nigig-email (push) Has been cancelled
email / supply-chain (push) Has been cancelled
nigig-map / test (push) Has been cancelled
sms / gates (push) Has been cancelled
sms / robius-sms (push) Has been cancelled
sms / android (push) Has been cancelled
sms / nigig-sms (push) Has been cancelled
sms / supply-chain (push) Has been cancelled
spreadsheet / engine-coverage (push) Has been cancelled
spreadsheet / ui-controller-coverage (push) Has been cancelled
chore: sync full working tree to gitdab
Whole-tree sync: cad-core/cad-ui split sources, nigig-build
construction_frame migration, pdf port progress, mpesa/pay/uikit/doc
updates, workspace members/profiles/lock, CI workflows and reviews.
See individual file history for details.
2026-09-12 07:15:24 +03:00

171 lines
8.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# ADR 0038: text reflow, caching and document structure
- **Status:** Accepted
- **Date:** 2026-09-04
- **Review item:** `NIGIG_PDF_FEATURE_PARITY_PLAN.md` Phase 8, "Content
reflow", "Text diff/cache for cheap re-renders, plus disk/page cache", and
"Document AI surface: structure/layout understanding"
- **Supersedes:** nothing; these are new surfaces on the ADR 0034 substrate
- **Related:** ADR 0034 (text search and layout, which this builds on), ADR
0033 (the wire codec the disk cache reuses), ADR 0031 (per-glyph offsets)
## Context
Phase 8's search and selection bullets are done (ADR 0034). Three bullets
remain, each "not started" or "partial" in the plan at the time of writing:
1. **Content reflow** — a re-readable, re-wrappable rendering of the text the
reader can show at a different width, on top of the same extracted runs.
2. **Text diff/cache** — cheap re-renders: the reader should not recompute a
page's text when nothing changed, and should not lose it on restart.
3. **Document AI** — answering "what *is* this page?" — headings, body, list
items, columns — rather than presenting an undifferentiated run soup.
The substrate is already there. `reading_order`, `detect_lines` and
`detect_columns_from_segments` (ADR 0034) already decide the geometry of
reading: which runs belong to which line, and which lines to which column.
The new work is the *semantics* on top: re-wrap the runs at a width, hash
and diff the extracted text, cache it, and classify it.
Three constraints shaped the decisions below:
- **No new dependencies.** The crate has no serde/bincode/postcard. Anything
that must be serialized uses the existing `wire::{encode, decode}`
`RenderCommand` codec (ADR 0033).
- **One layout truth.** The extracted text, the searched text and the
reflowed text must all be the same story, or a reader will show one thing
and search another — the exact defect ADR 0034 was written to kill.
- **`f64` is not `Hash`.** Hashable keys (text digests) must go through
`.to_bits()`, and equality of geometry predicates with `f64` must be
reasoned about, not assumed.
## Decision
### Reflow = reading order + paragraph grouping + word-wrap
`reflow` (and `reflow_from_commands`, its full-pipeline sibling) is three
stages:
1. **Reading order**`reading_order` from ADR 0034, so a two-column page
re-reads down each column rather than across.
2. **Paragraph grouping** — geometric: consecutive reading-order runs whose
baselines are within `PARAGRAPH_DETACH_EM = 2.0` ems belong to one
paragraph; a larger baseline step opens a new one. This is the spell-check
style of paragraphing — the reader has no markup, only coordinates.
3. **Word-wrap at the target width** — runs are re-laid onto lines of at
most `max_line_width` using a uniform per-em advance model
(`CHAR_ADVANCE_EM = 0.5`, `SPACE_ADVANCE_EM = 0.25`). This is an *advance
estimate, not a glyph metric*: exact widths would require the font's width
table, and reflow already has the real origins to preserve for
selection/hit-testing via `ReflowPage::segments()`.
The output is a `ReflowPage``ReflowParagraph`s of `ReflowLine`s —
with `plain_text()` (the "layout-friendly export") and `segments()` (which
feeds selection, hit-testing and search unchanged over the reflowed layout).
### Text digest + run diff, with order as part of the identity
`text_digest` is an FNV-1a over the page's runs: the text **and** the origin,
advance, font size and glyph offsets. Order matters and is documented as such
— two runs swapped is a different page, so a swapped page must not be
reported "unchanged". This is deliberately conservative: it errs toward
"changed" (recompute) rather than "unchanged" (stale highlight).
`diff_page_text` turns a pair of digests into a `TextDelta` of added /
removed / unchanged run sets, which a reader uses to decide whether a
selection or highlight needs moving. `is_noop()` is the cheap "nothing
changed" check.
### Disk cache on top of the wire codec, not a new serialization
`DiskCache` persists `DiskEntry { commands, images, page_height, text }` in
a file per page, with an `NGPD` magic and a version byte, generated from the
existing `wire::{encode, decode}`. A `Generation` guard makes a stale
generation a **miss**, never a corrupt read; truncated or corrupt bytes are a
miss, never a panic. The reader's render pipeline therefore stays
in-memory-first (the Phase 3 `PageCache` with its LRU budget) with the disk
as the persistence layer, both keyed by the same generation.
No serde, no `f64` hashing: the file is the wire codec's byte stream plus a
versioned envelope.
### Document structure by per-page comparison, not absolute rules
`analyze_structure` walks the reading order and groups runs into
`StructureBlock`s, classifying each `BlockKind` by signals a text-first
renderer actually has, *relative to the page's own statistics*:
- **Heading** — font size ≥ 1.45× the page's lower-median run size.
- **Caption** — ≤ 0.7× the median.
- **ListItem** — body-sized but starts with a bullet or an Arabic numeral
prefix.
- **Body** — everything in between.
- **Other** — by preservation of the run system, not effectively used.
The lower median (not the mean, not the upper median) is deliberate: with
sizes `[12, 24]` a single heading must not drag the "body size" up to be a
heading too. The whole-page-is-one-size case must read as body, not
heading-on-heading — asserted directly.
Because a PDF does not declare headings, this is a heuristic and is
documented as one. It is a single, tested place for that judgement; a reader
can act on it or override it.
## Consequences
- A reader can show a page re-wrapped at any width, and can present structure
(headings/list items) without re-deriving geometry.
- Re-renders are cheap when text is unchanged (digest), and survive restart
(disk cache) with corrupt/stale entries degrading to a miss rather than a
crash.
- Selection, hit-testing and search work over the reflowed layout too,
because `segments()` emits the same `TextSegment` shape — the "one layout
truth" invariant is preserved rather than a second, inconsistent layout
being introduced.
- The disk cache adds one file per page; no eviction policy beyond the
in-memory `PageCache`'s LRU budget and the generation guard.
- Reflow's advance model is an estimate. Charges are that a reflow at a very
narrow width may wrap more or fewer words than glyph-accurate shaping
would. This is acceptable for a transportable reflow export and is stated
on `ReflowOptions`.
- Block classification is heuristic and English/left-to-right biased. Right-
to-left reading order is not handled (as in ADR 0034), and unusual layouts
may mis-group. Errs toward `Body`.
## Merge criteria
Enumerated from the plan bullets first, per ADR 0021.
| Criterion | State |
|---|---|
| Reflow re-lays runs at a target width | ✅ wrap and re-layout unit tests |
| Reflow reads a two-column page down each column | ✅ reading order reused; real-page integration test |
| Reflow preserves every glyph of a real page | ✅ `reflow.rs` character-multiset integration test |
| Reflow output is layout-friendly text | ✅ `ReflowPage::plain_text` unit + integration |
| Reflowed layout feeds selection/search | ✅ `ReflowPage::segments()` emits `TextSegment`s |
| Empty/single-line pages handled | ✅ `reflow` returns `None` for empty; unit tests |
| Text digest catches any text change | ✅ covers content + origin + advance + size + offsets; order documented |
| Run diff reports added/removed/unchanged | ✅ unit tests |
| `is_noop` lets the reader skip re-render | ✅ unit tests |
| Disk cache round-trips commands and text | ✅ `put_then_get_round_trips`, `text_round_trips_through_the_disk` |
| Unknown page is a miss, not an error | ✅ |
| Stale generation is a miss | ✅ `a_stale_generation_is_a_miss` |
| Corrupt/truncated bytes are a miss, not a panic | ✅ `corrupt_bytes_are_a_miss_not_a_panic` |
| `clear_generation` removes only that generation | ✅ |
| Structure classifies a heading by size | ✅ |
| … a list item by bullet/number | ✅ |
| … the whole-page-one-size case as body | ✅ (lower median) |
| Reading order preserved in structure | ✅ top-to-bottom integration of emitted-bottom-up runs |
| Reflow unit tests (plan exit criterion) | ✅ unit + integration in `reflow.rs` |
| Search + selection integration tests (plan exit criterion) | ✅ already satisfied by ADR 0034 |
## Non-goals (stated so nobody mistakes them for gaps)
- **No glyph-accurate reflow shaping.** The wrap uses per-em advance
estimates; exact metrics are out of scope and would need per-font width
tables threaded through reflow.
- **No disk eviction.** The disk cache is a persistence layer; the in-memory
LRU budget does eviction.
- **No OCR, no tagged-PDF semantics.** Structure here is derived from
geometry and size, not from the PDF structure tree (a later-phase concern);
a scanned page has no text to classify.