Tonight's failures all traced to one blind spot: the converter could
only reason about text within a single block. A position at a chapter
heading sends walk-up context (heading + the paragraphs below, joined by
the plugin's block capture); a selection can span several paragraphs.
Neither shape could be verified (containment compared one block against
a multi-block quote, so the CORRECT structural landing at the heading
was rejected) nor matched by text search (it never crossed block
boundaries). The ladder then fell to the percentage rung — which labeled
its char-count guess Precision "exact" — and that confidently-wrong CFI
was stored: reading positions reopened paragraphs away from the true
spot, and a highlight echo overwrote the row's good web CFIs with a
garbage start anchor that made the highlight unpaintable ("disappeared").
Four changes, all in the forward converter and its consumers:
- Quote verification: after the structural walk lands, read the
whitespace-normalized document text forward from the landing point
(crossing block boundaries; inline spans join directly so drop-cap
splits still read as one word). A usable context must be a prefix of
that stream — which is exactly what device captures are: the text from
the position onward, or the selection between two anchors. The old
single-block containment checks remain as secondary acceptance.
- Cross-block text search: the search rung matches against the whole
document flattened in reading order, with every rune mapped back to
its source node and offset. A context spanning blocks now matches, and
the matched extent yields a true range end (EndEPUBCFI) that
highlights use as their end anchor, threaded through the facade as
CanonicalLocator.EndCFI.
- Honest labels: the percentage rung returns Precision "percentage" —
a char-count estimate must never masquerade as an exact anchor.
- Confident-only storage: progress adopts a converted locator solely at
structural/exact precision (section hrefs keep their legacy handling;
anything lower stores percentage only), and highlight conversion
returns CFIs only at structural/exact precision — a low-confidence
echo yields empty, which applyLWW coalescing turns into preservation
of the row's existing web CFIs instead of clobbering them.
Tests: walk-up context at a heading verifies structurally and lands in
the heading; a block-spanning context is found by search with a range
end landing in the following paragraph; the percentage rung is honestly
labeled; all drop-cap guards stay green.
Per-annotation percentage derivation (deriveAnnotationPercentage) built a
fresh CFIConverter for every highlight/note/bookmark, re-reading and
re-parsing the whole EPUB each time. Export SectionPercentageCached so
handlers reach the same bounded cache ConvertToCanonical already uses
(8 books, insertion-order eviction): one parse per book per push instead
of one per annotation.
ConvertToCanonical/ConvertFromCanonical built a fresh CFIConverter
per call, and each annotation converts twice (pos0+pos1) — a book
with 200 highlights re-opened and re-parsed the EPUB 400+ times per
sync, and again per metadata pull. A bounded 8-entry cache keyed by
path now shares converters (the parsing work belongs on the server;
clients stay thin). CFIConverter gained a mutex around its lazily
built spine/doc caches since instances are now shared between
concurrent requests.
Adds CFIConverter.SectionPercentage: book-wide percentage for a CRE
xpointer from the spine char distribution (midpoint of its document)
— the server-side counterpart to dropping per-annotation
getPageFromXPointer lookups from the plugin.
AnnotationService is the central service for cross-device annotation sync.
It provides SaveHighlight, SaveNote, and SaveBookmark methods that handle
the full sync lifecycle:
Identity (3-layer):
1. Server UUID (primary key)
2. Per-device native ID stored in device_sync_data JSONB
3. Content dedup_key: sha1(normalize(selection_text) + bucket_position)
- CFI character offsets are stripped for bucketing so the same
highlight at slightly different offsets still deduplicates
- Raw positions are preserved in the DB for precise restoration
Resolution policy (LWW):
- When the incoming annotation has an explicit ModifiedAt timestamp,
last_modified_at wins
- When the device sends zero ModifiedAt (creation time only), field-diff
mode compares content fields (text/color/note/percentage) — if all
match, the save is skipped; if any differ, the save is applied with
server-receive-time as the new last_modified_at
Conflict detection:
- When incoming and existing annotations have different sources (e.g.
koreader vs kobo) and content differs, an auto_resolved sync_conflict
is recorded with both sides' data for audit trail
- Broadcasts a WebSocket conflict notification for real-time UI updates
Tombstone management:
- Delete-wins: tombstoned annotations block recreation from stale pushes
- 30-day TTL before physical purge
- PurgeExpiredTombstones method + StartTombstonePurger goroutine (24h ticker)
Add locators.go with unified bidirectional CFI conversion:
ConvertToCanonical / ConvertFromCanonical
- CRE XPointer <-> standard EPUB CFI (for KOReader)
- KEPUB CFI passthrough (for Kobo)
- Skips non-reflowable formats (PDF, CBZ, fixed-layout EPUBs)
Add 25 unit tests covering:
- Dedup key determinism, text normalization, position sensitivity
- Offset insensitivity (CFI char-offset bucketing)
- Device sync data merge (preserves existing, overwrites same source)
- Cross-source detection
- LWW comparison (newer wins, older skipped, fallback to updated_at)
- Field-diff mode (identical content skipped, changes applied)
- Tombstone TTL constant
- CRE XPointer parsing and classification
- Standard EPUB CFI classification