Some checks failed
p2p-intel / engine (push) Waiting to run
p2p-intel / notifications (push) Waiting to run
p2p-intel / coverage (push) Waiting to run
p2p-intel / makepad-app (push) Waiting to run
p2p-intel / exchange-tab (push) Waiting to run
Payment domain, storage, platform and UI / isolated-payment-tests (push) Waiting to run
Payment domain, storage, platform and UI / payment-ui-tests (push) Waiting to run
repo hygiene / hygiene (push) Has been cancelled
PDF engine / engine (push) Has been cancelled
PDF engine / makepad-integration (push) Has been cancelled
PDF engine / fuzz (push) Has been cancelled
nigig-build (CAD) / supply-chain (push) Has been cancelled
nigig-build (CAD) / cad-module (push) Has been cancelled
nigig-build (CAD) / full-crate-check (push) Has been cancelled
nigig-build (CAD) / cad-engine-coverage (push) Has been cancelled
nigig-build (CAD) / doc-workspace-coverage (push) Has been cancelled
nigig-build (CAD) / cad-widget-coverage (push) Has been cancelled
traffic / gates (push) Has been cancelled
traffic / nigig-traffic (push) Has been cancelled
traffic / supply-chain (push) Has been cancelled
doc-engine / engine (push) Has been cancelled
doc-engine / coverage (push) Has been cancelled
doc-engine / consumer (push) Has been cancelled
email / gates (push) Has been cancelled
email / email-domain (push) Has been cancelled
email / nigig-email (push) Has been cancelled
email / supply-chain (push) Has been cancelled
nigig-map / test (push) Has been cancelled
sms / gates (push) Has been cancelled
sms / robius-sms (push) Has been cancelled
sms / android (push) Has been cancelled
sms / nigig-sms (push) Has been cancelled
sms / supply-chain (push) Has been cancelled
spreadsheet / engine-coverage (push) Has been cancelled
spreadsheet / ui-controller-coverage (push) Has been cancelled
Whole-tree sync: cad-core/cad-ui split sources, nigig-build construction_frame migration, pdf port progress, mpesa/pay/uikit/doc updates, workspace members/profiles/lock, CI workflows and reviews. See individual file history for details.
171 lines
8.9 KiB
Markdown
171 lines
8.9 KiB
Markdown
# ADR 0038: text reflow, caching and document structure
|
||
|
||
- **Status:** Accepted
|
||
- **Date:** 2026-09-04
|
||
- **Review item:** `NIGIG_PDF_FEATURE_PARITY_PLAN.md` Phase 8, "Content
|
||
reflow", "Text diff/cache for cheap re-renders, plus disk/page cache", and
|
||
"Document AI surface: structure/layout understanding"
|
||
- **Supersedes:** nothing; these are new surfaces on the ADR 0034 substrate
|
||
- **Related:** ADR 0034 (text search and layout, which this builds on), ADR
|
||
0033 (the wire codec the disk cache reuses), ADR 0031 (per-glyph offsets)
|
||
|
||
## Context
|
||
|
||
Phase 8's search and selection bullets are done (ADR 0034). Three bullets
|
||
remain, each "not started" or "partial" in the plan at the time of writing:
|
||
|
||
1. **Content reflow** — a re-readable, re-wrappable rendering of the text the
|
||
reader can show at a different width, on top of the same extracted runs.
|
||
2. **Text diff/cache** — cheap re-renders: the reader should not recompute a
|
||
page's text when nothing changed, and should not lose it on restart.
|
||
3. **Document AI** — answering "what *is* this page?" — headings, body, list
|
||
items, columns — rather than presenting an undifferentiated run soup.
|
||
|
||
The substrate is already there. `reading_order`, `detect_lines` and
|
||
`detect_columns_from_segments` (ADR 0034) already decide the geometry of
|
||
reading: which runs belong to which line, and which lines to which column.
|
||
The new work is the *semantics* on top: re-wrap the runs at a width, hash
|
||
and diff the extracted text, cache it, and classify it.
|
||
|
||
Three constraints shaped the decisions below:
|
||
|
||
- **No new dependencies.** The crate has no serde/bincode/postcard. Anything
|
||
that must be serialized uses the existing `wire::{encode, decode}`
|
||
`RenderCommand` codec (ADR 0033).
|
||
- **One layout truth.** The extracted text, the searched text and the
|
||
reflowed text must all be the same story, or a reader will show one thing
|
||
and search another — the exact defect ADR 0034 was written to kill.
|
||
- **`f64` is not `Hash`.** Hashable keys (text digests) must go through
|
||
`.to_bits()`, and equality of geometry predicates with `f64` must be
|
||
reasoned about, not assumed.
|
||
|
||
## Decision
|
||
|
||
### Reflow = reading order + paragraph grouping + word-wrap
|
||
|
||
`reflow` (and `reflow_from_commands`, its full-pipeline sibling) is three
|
||
stages:
|
||
|
||
1. **Reading order** — `reading_order` from ADR 0034, so a two-column page
|
||
re-reads down each column rather than across.
|
||
2. **Paragraph grouping** — geometric: consecutive reading-order runs whose
|
||
baselines are within `PARAGRAPH_DETACH_EM = 2.0` ems belong to one
|
||
paragraph; a larger baseline step opens a new one. This is the spell-check
|
||
style of paragraphing — the reader has no markup, only coordinates.
|
||
3. **Word-wrap at the target width** — runs are re-laid onto lines of at
|
||
most `max_line_width` using a uniform per-em advance model
|
||
(`CHAR_ADVANCE_EM = 0.5`, `SPACE_ADVANCE_EM = 0.25`). This is an *advance
|
||
estimate, not a glyph metric*: exact widths would require the font's width
|
||
table, and reflow already has the real origins to preserve for
|
||
selection/hit-testing via `ReflowPage::segments()`.
|
||
|
||
The output is a `ReflowPage` — `ReflowParagraph`s of `ReflowLine`s —
|
||
with `plain_text()` (the "layout-friendly export") and `segments()` (which
|
||
feeds selection, hit-testing and search unchanged over the reflowed layout).
|
||
|
||
### Text digest + run diff, with order as part of the identity
|
||
|
||
`text_digest` is an FNV-1a over the page's runs: the text **and** the origin,
|
||
advance, font size and glyph offsets. Order matters and is documented as such
|
||
— two runs swapped is a different page, so a swapped page must not be
|
||
reported "unchanged". This is deliberately conservative: it errs toward
|
||
"changed" (recompute) rather than "unchanged" (stale highlight).
|
||
|
||
`diff_page_text` turns a pair of digests into a `TextDelta` of added /
|
||
removed / unchanged run sets, which a reader uses to decide whether a
|
||
selection or highlight needs moving. `is_noop()` is the cheap "nothing
|
||
changed" check.
|
||
|
||
### Disk cache on top of the wire codec, not a new serialization
|
||
|
||
`DiskCache` persists `DiskEntry { commands, images, page_height, text }` in
|
||
a file per page, with an `NGPD` magic and a version byte, generated from the
|
||
existing `wire::{encode, decode}`. A `Generation` guard makes a stale
|
||
generation a **miss**, never a corrupt read; truncated or corrupt bytes are a
|
||
miss, never a panic. The reader's render pipeline therefore stays
|
||
in-memory-first (the Phase 3 `PageCache` with its LRU budget) with the disk
|
||
as the persistence layer, both keyed by the same generation.
|
||
|
||
No serde, no `f64` hashing: the file is the wire codec's byte stream plus a
|
||
versioned envelope.
|
||
|
||
### Document structure by per-page comparison, not absolute rules
|
||
|
||
`analyze_structure` walks the reading order and groups runs into
|
||
`StructureBlock`s, classifying each `BlockKind` by signals a text-first
|
||
renderer actually has, *relative to the page's own statistics*:
|
||
|
||
- **Heading** — font size ≥ 1.45× the page's lower-median run size.
|
||
- **Caption** — ≤ 0.7× the median.
|
||
- **ListItem** — body-sized but starts with a bullet or an Arabic numeral
|
||
prefix.
|
||
- **Body** — everything in between.
|
||
- **Other** — by preservation of the run system, not effectively used.
|
||
|
||
The lower median (not the mean, not the upper median) is deliberate: with
|
||
sizes `[12, 24]` a single heading must not drag the "body size" up to be a
|
||
heading too. The whole-page-is-one-size case must read as body, not
|
||
heading-on-heading — asserted directly.
|
||
|
||
Because a PDF does not declare headings, this is a heuristic and is
|
||
documented as one. It is a single, tested place for that judgement; a reader
|
||
can act on it or override it.
|
||
|
||
## Consequences
|
||
|
||
- A reader can show a page re-wrapped at any width, and can present structure
|
||
(headings/list items) without re-deriving geometry.
|
||
- Re-renders are cheap when text is unchanged (digest), and survive restart
|
||
(disk cache) with corrupt/stale entries degrading to a miss rather than a
|
||
crash.
|
||
- Selection, hit-testing and search work over the reflowed layout too,
|
||
because `segments()` emits the same `TextSegment` shape — the "one layout
|
||
truth" invariant is preserved rather than a second, inconsistent layout
|
||
being introduced.
|
||
- The disk cache adds one file per page; no eviction policy beyond the
|
||
in-memory `PageCache`'s LRU budget and the generation guard.
|
||
- Reflow's advance model is an estimate. Charges are that a reflow at a very
|
||
narrow width may wrap more or fewer words than glyph-accurate shaping
|
||
would. This is acceptable for a transportable reflow export and is stated
|
||
on `ReflowOptions`.
|
||
- Block classification is heuristic and English/left-to-right biased. Right-
|
||
to-left reading order is not handled (as in ADR 0034), and unusual layouts
|
||
may mis-group. Errs toward `Body`.
|
||
|
||
## Merge criteria
|
||
|
||
Enumerated from the plan bullets first, per ADR 0021.
|
||
|
||
| Criterion | State |
|
||
|---|---|
|
||
| Reflow re-lays runs at a target width | ✅ wrap and re-layout unit tests |
|
||
| Reflow reads a two-column page down each column | ✅ reading order reused; real-page integration test |
|
||
| Reflow preserves every glyph of a real page | ✅ `reflow.rs` character-multiset integration test |
|
||
| Reflow output is layout-friendly text | ✅ `ReflowPage::plain_text` unit + integration |
|
||
| Reflowed layout feeds selection/search | ✅ `ReflowPage::segments()` emits `TextSegment`s |
|
||
| Empty/single-line pages handled | ✅ `reflow` returns `None` for empty; unit tests |
|
||
| Text digest catches any text change | ✅ covers content + origin + advance + size + offsets; order documented |
|
||
| Run diff reports added/removed/unchanged | ✅ unit tests |
|
||
| `is_noop` lets the reader skip re-render | ✅ unit tests |
|
||
| Disk cache round-trips commands and text | ✅ `put_then_get_round_trips`, `text_round_trips_through_the_disk` |
|
||
| Unknown page is a miss, not an error | ✅ |
|
||
| Stale generation is a miss | ✅ `a_stale_generation_is_a_miss` |
|
||
| Corrupt/truncated bytes are a miss, not a panic | ✅ `corrupt_bytes_are_a_miss_not_a_panic` |
|
||
| `clear_generation` removes only that generation | ✅ |
|
||
| Structure classifies a heading by size | ✅ |
|
||
| … a list item by bullet/number | ✅ |
|
||
| … the whole-page-one-size case as body | ✅ (lower median) |
|
||
| Reading order preserved in structure | ✅ top-to-bottom integration of emitted-bottom-up runs |
|
||
| Reflow unit tests (plan exit criterion) | ✅ unit + integration in `reflow.rs` |
|
||
| Search + selection integration tests (plan exit criterion) | ✅ already satisfied by ADR 0034 |
|
||
|
||
## Non-goals (stated so nobody mistakes them for gaps)
|
||
|
||
- **No glyph-accurate reflow shaping.** The wrap uses per-em advance
|
||
estimates; exact metrics are out of scope and would need per-font width
|
||
tables threaded through reflow.
|
||
- **No disk eviction.** The disk cache is a persistence layer; the in-memory
|
||
LRU budget does eviction.
|
||
- **No OCR, no tagged-PDF semantics.** Structure here is derived from
|
||
geometry and size, not from the PDF structure tree (a later-phase concern);
|
||
a scanned page has no text to classify.
|