The test-runner stage copied only /app from the builder, but the Go module
cache lives in /go/pkg/mod (GOMODCACHE) — the cache mounts used by the
builder target /root/go/pkg/mod, which the go tool ignores, so those mounts
never held anything. Every `compose run tests` / make test-integration
invocation therefore re-downloaded all dependencies from the network,
making integration runs slow and timeout-prone.
The builder now materializes /go/pkg/mod into an image layer
(go-module-cache) after the binary build, and the test-runner stage
restores it to /go/pkg/mod before running tests. Test-only stage; the
production final stage is unaffected.
Three fixes to the Delete Library flow on /admin/library:
1. Honest warning copy. The old text ("Media files will not be deleted")
read as reassurance while deleting a library actually cascade-destroys
every book record with it: reading progress, highlights, notes,
bookmarks, and ratings are gone with no archive window and no undo.
The modal now states this as a scannable list:
- Permanently removed, no archive or undo: the library, its folder
mappings, and all book records — with their reading progress,
highlights, notes, bookmarks, and ratings.
- Not touched: media files on disk.
Admins skim danger dialogs; the irreversible part now leads.
2. The dialog stayed open after confirming. The confirm button swaps
#libraries-container via htmx (so the deleted row vanished) but the
modal lives outside the swapped container and nothing closed it. An
htmx:afterRequest listener now hides the modal on successful deletes
and leaves it open on errors.
3. Latent ReferenceError in openFolderBrowser: the function parameter is
targetInputId but the body referenced an undefined targetInput, so
opening the folder browser threw and the picker never loaded.
The device catalog flattened every visible library into one list, so
duplicate copies of the same book each got an entry and search had no
library context. The root feed is now a standard OPDS 1.1 navigation feed
that mirrors the web UI's library model:
- Root /catalog: an "All Books" entry first (the previous flat cross-
library behavior, also still served at ?all=1 for clients that want
the single flat list), followed by one folder entry per visible
library with live book counts from GetVisibleLibraryMediaCounts.
- New GET /library/:libraryId/catalog: acquisition feed scoped to one
visible library (403 when the device owner cannot see it), paginated,
with an up-link back to the root.
- Scoped search: each library feed's rel="search" OpenSearch template
pins &library_id=<id>, so opening search from inside a library folder
searches only that library — clients substitute only {searchTerms},
so no client-side changes are required. Root search stays global.
- Global search entries now carry the owning library as a category (and
as a fallback summary when the book has no description), so duplicate
copies are distinguishable in unscoped result lists.
The acquisition entry builder is extracted into addAcquisitionEntries and
shared by the all-books and per-library feeds. bookhoard.koplugin needs
no changes: it only registers the root URL, and KOReader's stock OPDS
client renders navigation feeds natively.
Tests: TestOPDSLibraryFolders covers the nav-feed shape, flat ?all=1
mode, scoped catalog isolation, scoped/global search behavior, and 403s
for libraries outside the device's visibility. Library names avoid the
word "test" on purpose — setupDeviceTest re-runs setupTestServer's
setup-time cleanup, which deletes every library whose name contains it.
Covers the copy-then-delete-later workflow: a user copies books to a new
library, deletes the originals, and ends up with an active copy in the new
library plus an archived twin holding the real reading history. Until now
those twins could only be purged (destroying the history) or restored
(showing a permanently broken entry).
- POST /api/media-items/:id/merge (admin only, body {target_id}):
validates the source is archived/missing, the target is active, and
both share the same file_sha256; then re-parents every child row onto
the target via the existing reparent_media_item_children function (the
same machinery as hash-conflict resolution) and deletes the source row.
Per-user collisions keep the active copy's data, mirroring that flow.
Affected data: reading progress, speed, ratings, highlights, notes,
bookmarks (including tombstoned deleted-annotation history), formats,
collections, kobo shelves/entitlements, sync rows, panel data,
processing issues, and device aliases.
- GET /admin/archived: ListHiddenMediaItems' match_* columns surface each
row's best active twin; rows with a twin get a confirm-guarded "Merge"
button (data-merge-source/-target) next to Restore/Delete, wired in
web/src/admin.ts like the existing unarchive/delete handlers.
- templates.ArchivedItem gains MatchID/MatchTitle/MatchLibraryName;
frontend.go populates them from the listing row.
After a merge the archived row is gone, so the retention purge can never
destroy the merged data. Deleted-annotation tombstones carry over and
remain restorable from the target book's "recently deleted" history.
Previously the scanner's SHA-256 dedup was library-scoped: moving a book
between libraries created a duplicate row (new ID) while the old row went
missing and archived after two scans, orphaning reading progress and
annotations from the file that users still see.
processMediaFile now falls through to cross-library detection when the
same-library hash lookup misses:
- GetMediaItemsBySHA256AndSameLibraryType returns candidates in other
same-type libraries; each candidate's file is stat'd through ITS OWN
library's folders (the old existence check stat'd against the current
scan's folder tree, which is meaningless across libraries).
- File gone -> it is a move: the candidate must pass the target library's
type rules (ValidateMediaItemForLibrary), then the row is repointed via
MoveMediaItemToLibrary with the recomputed relative path. Reading
history, annotations, and collections follow automatically because the
row keeps its ID. Incompatible formats (e.g. a reflowable EPUB into a
manga library) are rejected with a processing issue instead of being
force-imported; the upsert on (media_item_id, issue_type) keeps
repeated scans from spamming duplicates.
- File still present -> deliberate multi-library copy: fall through to
normal import so both libraries keep independent rows.
Candidate selection is factored into selectMoveCandidate (pure function,
unit-tested in media_scanner_move_test.go): input arrives pre-ordered
(archived first, then most missing scans, then oldest) and only
candidates whose file is verifiably gone qualify, so deliberate copies
are never repointed.
Three related query groups in queries.sql (sqlc regenerated; no generated
signatures changed, so no caller edits were needed):
1. Dashboard/collection archive filters. The smart-section and collection
queries returned archived and missing items (visible only as broken
covers once their files vanished). Added the codebase-standard
"AND archived_at IS NULL AND missing_scan_count = 0" predicate to:
GetContinueReadingItems, GetRecentlyAddedItems, GetRecentlyReadItems,
GetNotStartedItems, GetCollectionItemsForDashboard, GetLibraryItems.
Matches the existing convention in ListMediaItems, SearchMediaItems,
the library counts, and GetContinueSeriesItems.
2. Cross-library SHA move detection support.
- GetMediaItemsBySHA256AndSameLibraryType: finds identical content
(same file_sha256) registered in a DIFFERENT library of the SAME
library type, ordered archived-first, most-missing, oldest. Type
scoping keeps ebooks from merging into manga/comics libraries.
- MoveMediaItemToLibrary: repoints an existing row to the new library
(library_id + recomputed file_path + file_size) and clears
archive/missing state. The row keeps its ID, so reading progress,
annotations, collections, and kobo shelves follow the book
automatically; the new set_library_type_name_on_library_change
trigger refreshes the denormalized type column.
3. Archived-twin matching. ListHiddenMediaItems (admin archived-items
page) now LEFT JOIN LATERALs the best active same-SHA twin per hidden
row — preferring same library type, then same library, then most
reading progress, then oldest — exposing match_id / match_title /
match_library_name so the UI can offer merging the archived row's
reading data into its active duplicate. COALESCE on match_title keeps
the no-match case NULL-safe.
schema.sql is re-executed in full on every application startup, so every
statement in it must be idempotent. Two related problems introduced with the
cross-library move feature, plus their fix:
Problem 1: media_items.library_type_name was only populated by a BEFORE
INSERT trigger. When the scanner repoints an existing row to a different
library (cross-library move detection), the denormalized library_type_name
went stale, mislabeling the item's type for validation and display.
Fix: add a second trigger, set_library_type_name_on_library_change, that
fires BEFORE UPDATE OF library_id and re-runs the same population function.
Problem 2 (outage): the new trigger EXECUTE FUNCTIONs set_library_type_name(),
but the pre-existing idempotency dance drops that function on every startup
before recreating it. On the second and later boots, DROP FUNCTION failed
with SQLSTATE 2BP01 (function still depended on by the trigger created by
the previous boot), aborting the whole schema transaction and crash-looping
the container.
Fix: drop BOTH triggers before dropping the function, and create both after
it. The ordering now survives any number of restarts on any database state.
Adds TestSchemaInitializationIsIdempotent, which replays the production
database.Initialize twice against the same database so this class of
"second startup" regression fails in CI instead of in production. Note the
test uses a dedicated pgxpool: Initialize holds the advisory lock on one
connection while executing the schema on another, which deadlocks against
the shared test pool's MaxConns=1.
go mod download all added sums for modules only reachable from test
builds (chroma, bluemonday, compress families); without them the
integration-test binary build fetches missing sums at compile time.
Two pre-existing harness breaks found while running the device tests:
- RegisterRoutes dereferences cfg.Settings (auth rate limit) but
setupTestServer never set it — every integration test nil-panicked at
route registration. Wire database.NewSettingsRegistry(queries) the
same way main.go does.
- The cleanup deliberately PRESERVES testuser@tests.bookhoard.internal
(shared dev admin), but setup then blindly re-INSERTed that user, so
every run after the first failed on users_email_key. Reuse the user
when it exists; default collections are created only for a NEW user
(the preserved admin already has its set).
Also documented the offline runner used on slow links: build the test
binary on the host (CGO_ENABLED=0 go test -c) and execute it inside the
bookhoard-tests image against the compose network — no in-container
module downloads, no image rebuild.
The mutex hardening added a fall-through 'pending' response path that
returned while still holding pendingMu — one status poll before approval
leaked the lock and every later register/approve/status request hung
forever (caught by TestDeviceRegistrationFlow's approve step hanging).
App reinstalls that preserve data (Android Studio installDebug over an
existing install) re-register with the same device_identifier, but the
devices row from the previous install still exists — device_identifier
is UNIQUE, so ApproveDevice's blind INSERT failed with a unique
violation and returned 500 'failed to create device' (reproduced via
curl: second approve with the same identifier = instant 500; the ~98s
in the original report was app-side retry/polling, not server wait).
ApproveDevice is now idempotent: look the device up by identifier
first; a row owned by the approving user gets its auth token rotated
via UpdateDeviceAuthToken (row id unchanged, so synced highlights/
bookmarks/progress anchored to it stay valid; fresh install = fresh
credentials, old token invalidated); a row owned by another user gets
409; unknown identifiers INSERT as before, with the 23505 race falling
through to the rotate path. DB failures are logged (they were silent).
Also guard the in-memory pendingRegistrations map with a mutex —
register/approve/reject/status/list all touch it from HTTP goroutines,
and a racing write is a Go runtime fatal, not an error. The approver's
credential publication and the status poller's approved-branch snapshot
now run under the lock so the token can never be read half-written.
Regression test: TestApproveDeviceReapprovalRotatesToken — register →
approve → re-register same identifier → approve (must be 200) → token
rotated, exactly one devices row, row carries the new token.
User-facing documentation for the rebuilt web reader (§6.5 of the app
handoff): swipe-to-page / tap-for-menu touch navigation, long-press word
selection with drag extension and handles, the selection popover (color
dots, notes with save-together-on-create, copy incl. plain-HTTP LAN,
edit/delete by tapping a painted highlight), below-the-selection popover
placement rationale, PDF/comic behavior, and annotations drawer.
Linked from the user documentation portal and the docs index quick
links / quick-find tables.
The lockfile is gitignored (each machine keeps its own), but
.dockerignore did not exclude it, so any stale local lock rode into
every docker build via `COPY package*.json`. Because the forked
foliate-js declares "version": "0.0.0" on every commit, npm treats
the git pin as already satisfied by name@version and never
re-resolves the new commit hash — silently installing and bundling
the old code. This bit both the host npm cache mount (documented at
Dockerfile:18-20) and, today, `make rebuild-app-force`: a fresh
no-cache image was built with the pre-feature 1305a52 foliate-js
(chunk fixed-layout-B8-qRQLl.js) despite package.json pinning
e16530a, while the Gitea runner (fresh checkout, no lockfile, cold
cache) built correctly.
With no lockfile in the context, npm install resolves git pins
fresh from package.json each build (tarballs are cached by
commit-specific URLs), so the persistent npm cache mount cannot
serve old commits across pin bumps. Local lockfiles can no longer
poison builds even if regenerated on the host.
Verified: after evicting the poisoned cache mounts
(docker builder prune --filter type=exec.cachemount) and rebuilding,
the container serves fixed-layout-BE0KdOql.js with both dblclick
handlers present, matching the reference build.
Cancel only collapsed the editor section, and edit-mode saves only set
noteOpen=false — the minimized popover (colors/pencil/copy/delete) stayed
up. Explicit Save and Cancel now call hideSelectionPopover; color tweaks
with the editor open still keep it open.
- The in-book pointerdown dismiss listener lived in the reflowable-only
block, so PDF taps never closed the selection popover — hoist it so
every content doc (EPUB, PDF, comics) dismisses on tap.
- createHighlight hardcoded note_text: '' — 'Highlight with note' saved
the highlight but silently dropped the note. Send p.note (and pass it
to the EPUB overlayer).
- saveHighlightChanges collapsed the note editor on every save; color
dots in create mode created a highlight outright. With the editor
open, color dots now just recolor (create) or recolor-and-save with
the editor kept open (edit).
Chromium's touch selection takeover swallows pointerup in the fx iframe,
so the PDF popover's only trigger never fired on real phones — it only
appeared when timing happened to deliver the event. Mirror the EPUB
path: a selectionchange debounce (400ms settle) opens the popover, with
the pointerup path kept for desktop plus the quick-tap word-select guard.
Also place the popover BELOW the selection on coarse pointers (+90px
clearing the native selection menu and drag handles), matching the
reflowable path; desktop keeps above-placement.
GitHub-pinned builds silently reused a stale foliate-js tarball from
Docker's npm cache mount (name@version never changed from 0.0.0);
every build since the GitHub pin shipped pre-94bb384 code. Repin to
the mirror and bust the cache mount.
The fx renderer re-rendered the PDF on every viewport resize;
mobile browser toolbar transitions (7-17% height change) during a
text selection drag caused the old canvas to be cleared for
re-rendering while the async pdf.js render raced with the next
resize, blanking the page. Now gated by the same 25% threshold
as the paginator.
The fx renderer's gesture classifier only checked whether the touch
started on a .textLayer span; Chromium's long-press selects the
nearest word even when the finger landed between spans, so the
classifier saw a non-selectable target and classified the drag as
swipe/pan — preventDefault on the iframe's touchmove then cancelled
the native selection extension mid-drag and could blank the PDF
canvas. If any frame already has a non-collapsed selection, the
drag is now classified as native.
pokeChrome's anySelection gate only prevents the chrome from being
raised — if it was already visible when the selection began, it
stayed up. noteSelectionActivity now actively hides it the moment
a selection appears.
- Custom drag handles (start/end) for post-lift selection adjustment:
the native handles are disabled with the rest of the native touch
selection controller; these are positioned at the selection's
boundary carets and driven through the same clamped caret mapping
- Copy button falls back to execCommand via a transient textarea on
plain HTTP (navigator.clipboard is unavailable on LAN addresses);
both paths clear the selection and dismiss the popover
- Only dismiss the selection popover on actual position changes in
relocate (detail.section is a fresh object on every relocate, so
reference comparison always saw a move); a tap's no-op relocate
was closing the just-opened highlight edit popover
- Pin foliate-js f872a01 (synthetic click dispatch on quick taps)
and 422e8e0 (don't snap taps that never panned)
- tap zones: map in-iframe taps into the visible page slice (the
iframe is laid out at full section width; every tap previously
computed as the left margin) and shrink the default zone size to 12%
- hyphens: manual on coarse pointers: Chrome's touch word-selection
walks hyphen fragments, re-anchoring line-start drags and
overshooting line ends into the next column
- selection popover: opens for settled touch selections (the gesture
takeover swallows pointerup), positioned below the selection with
on-screen clamping; creating a highlight clears the selection so
the browser's own menu follows
- settle-time recovery: return the view to the selection's anchor
page and clamp the selection to the visible page after the
browser's selection auto-scroll wanders off the page grid
- overscroll-behavior: none on the reader page (pull-to-refresh
during downward drags)
Non-touch gate on selection auto-paging (drag-selecting no longer
pages away mid-gesture) and small height-only viewport resizes no
longer re-wrap the book (mobile browser toolbar transitions).
Fixed-layout EPUBs lean on the reading_direction column (the web reader
forces book.dir = rtl from it when the file didn't set direction
itself), but the scanner never populated it for EPUBs - only ComicInfo
fed it. Meanwhile real Japanese EPUBs declare page-progression-direction
on the OPF spine, which foliate reads client-side but nothing stored.
Read the spine attribute in the structured OPF parser and map it into
ReadingDirection in parseOPFContent (EPUB2/3, case-insensitive, plus
'right-to-left'/'left-to-right' spellings); undeclared stays empty
rather than forcing ltr, preserving the editor's Auto default. Sidecar
OPFs are metadata-only documents without spines, so the Calibre path is
a no-op. The merge gap-fill copies an embedded-only direction into a
blank sidecar field, and hand-set values keep winning through the
existing OverrideReadingDirection protection.
Tests: declared rtl/RTL/ltr, undeclared and unknown values staying
empty, plus sidecar-wins vs embedded-fills merge cases. Existing manga
EPUBs declaring rtl (verified live in-library) pick the value up on
their next scan, feeding the API and reader config mobile clients
consume.
Same gap-fill rule as the EPUB path, against the PDF's embedded Info
dictionary via a readPDFInfoDict helper reusing extractPDFMetadata's
field conventions (creator falls back to author, subject maps to
description, producer to publisher, keywords to tags, plus page count).
Sidecar values always win; unopenable PDFs skip silently.
TestMergeMetadataPDFGapFill uses an Info-bearing hand-built PDF fixture
(shared with the sidecar-cover test) and asserts the sidecar title is
kept while author/description/publisher/tags/page count fill in.
A sparse metadata.opf/metadata.json (title only, no description) left
books thin even when the file itself carried rich data: the embedded
extractors only ran when no sidecar existed at all. Now mergeMetadata
fills blanks from the book's own OPF - title, author, description,
publisher, language, ISBN, ASIN, series/number, publish date, tags,
contributors - while sidecar values always win and unparseable files
skip silently. Also covers .kepub, which the merge previously ignored
while the extractor already supported it.
Adds TestMergeMetadataEPUBGapFill asserting both directions: sidecar
title/author survive, embedded description/publisher/language fill in.
Admins need to see what the archive lifecycle is holding: a dedicated
/admin/archived page listing every hidden item (archived or still in
the missing grace window) with library, file path, status, and - when
retention is enabled - the exact date it will be permanently deleted
(archived_at + ARCHIVE_RETENTION_DAYS).
Per row: Restore (POST /api/media-items/:id/unarchive, clears the
archive state so it reappears; if the file is still gone the next scan
hides it again) and Delete (existing DELETE endpoint for single-row
purge with its reading history). Purge All Archived reuses the existing
bulk button. The library admin banner links to the page and keeps its
purge button; row actions use data attributes with delegated listeners
in admin.ts since this templ version has no JSFunctionCall helper.
Verified live: page 200 with purge dates shown, unarchive returned 204
and reset the row, per-item delete and bulk purge ({purged:1}) both
removed their rows with no leftovers.
Two visibility fixes so the UI reflects the shelf's real state on the
next scan instead of only after the archive gate:
- Missing items disappear at once: the user-facing filters already hid
archived rows; the same listings now also require missing_scan_count =
0. A deleted or moved-then-not-yet-repointed file vanishes from the
UI immediately, while purge timing stays gated on archived_at plus the
retention window - grace protects data, not visibility. Restored
automatically when the file returns.
- Steam Deck SD-card model for unmounted storage: resolveLibrary stats
each library's folder roots and flags libraries with no live folder
as Offline (LibraryData gains the field, computed at request time so
mounts/unmounts react instantly). The bookshelf shows an empty shelf
plus a 'storage is not connected' notice for an offline selected
library, and both the shared LibrarySwitcher and the bookshelf's
inline select label offline libraries with their true holding counts.
Nothing is marked or purged while offline.
- TotalMediaCount skips offline libraries, so the 'All Books/Libraries'
totals match what is actually visible.
Verified live: renaming uploads/Manga away produced the notice, an
empty shelf, and the offline dropdown label with a corrected total;
renaming it back restored all 38 cards with zero dirty rows.
When the SHA-256 dedup found identical content already in the library,
the scan skipped the file as a duplicate - and after files moved
between folders the old row kept its stale path, cycled missing ->
archived, and the new path never took. The archive feature turned the
old destructive move behavior into a stuck move instead.
Now the dedup branch stats the old location: if it is gone, the book
was MOVED, so the row is repointed (file_path, file_size) with archive
state cleared and reading history intact. 'Skip as duplicate' only
applies when the old path still exists (a true copy). Verified live: a
moved EPUB kept its single row, followed the file, and never entered
the missing/archive cycle.
Storage behind the scan lifecycle fixes:
- MoveMediaItemFilePath: repoints a row (path, size) and clears archive
state when identical content reappears at a new location.
- ListHiddenMediaItems: archived OR missing rows for the admin
archived-items page; ListMediaItemsByLibraryIncludingArchived (prior
commit) already fed the scanner.
- All user-facing listings now hide missing items immediately, not just
archived ones: ListMediaItems, ListMediaItemsByLibrary,
ListMediaItemsSorted, SearchMediaItems, SearchMediaItemsUnified, the
next_books CTE, library media counts, and the five search autocomplete
value queries gained AND mi.missing_scan_count = 0 next to the
archived_at filter. Purge timing is unchanged (archived_at +
ARCHIVE_RETENTION_DAYS), so the grace period keeps protecting data
while the UI reflects removals on the first scan.
- Detail lookups by id/path/sha stay unfiltered on purpose.
A full rescan wiped cover_image_path for every PDF/EPUB living in a
metadata.json (or metadata.opf-less) folder with no cover.jpg next to
the book: the sidecar branches returned early after findSidecarCover
missed, and mergeMetadata has no PDF/EPUB cover logic of its own. Four
books lost their thumbnails while their {file}.cover.jpg files still sat
on disk - most visibly the Audiobookshelf-managed No Starch titles.
Both sidecar branches now fall through to an embedded-cover fallback
(PDF via extractPDFCover, EPUB/KEPUB via extractEPUBCover) whenever no
sidecar cover file exists. Comics are untouched: mergeMetadata already
extracts their covers from the archive.
Adds TestSidecarCoverFallback with a hand-built one-page PDF carrying a
JPEG XObject plus a metadata.json sidecar, asserting the sidecar title
wins while the cover still comes from the file. Verified live: rescans
restored all four dereferenced covers with no leftover rows.
Reset to Scanned cleared the overrides row (successfully) but then set
the in-memory copy to nil before handing it to updateMediaItem. pgx
encodes a nil []string parameter as SQL NULL, so the follow-up UPDATE
wrote metadata_overrides = NULL into the column's NOT NULL constraint
and the whole rescan failed with:
failed to update media item: ERROR: null value in column
"metadata_overrides" of relation "media_items" violates not-null
constraint (SQLSTATE 23502)
Two changes:
- RescanMediaItem's reset path assigns []string{} instead of nil, with a
comment explaining the pgx nil-to-NULL encoding trap.
- updateMediaItem routes the override set through utils.MergeOverrides,
whose contract guarantees a non-nil slice, so no caller can write
NULL into that column again (verified against pgx v5.9.2 source: a
scanned '{}' round-trips as non-nil in both directions; the nil could
only come from our own assignment).
The plain Rescan path never hit this - only Reset did. Worse, the reset
is the remedy when a book's cover_image_path override pins an empty
cover, so the crash also blocked the way out of that state. After this
fix, a plain rescan on an already-reset book repopulates scanned
metadata and extracts the cover.
Adopt Calibre's reading conventions for the Dublin Core metadata that
parseOPFContent now pulls from the structured OPF parse:
- Titles: EPUB3 title-type selection (prefer 'main', join a distinct
subtitle with ': ' exactly as Calibre stores it). There is no separate
subtitle column by design - Calibre-sidecar books arrive pre-joined,
so a column would stay empty for most libraries and force every client
to reimplement concatenation.
- Genre: first dc:subject, mirroring the existing processGenresAndTags
behavior of the Calibre-sidecar path; the embedded path never
populated Genre before. Subjects stay one-element-one-tag - Library
of Congress headings legitimately contain commas ("Holmes, Sherlock
(Fictitious character) -- Fiction") and must not be split.
- Identifiers: urn:isbn:/urn:asin: prefixed values parse in addition to
opf:scheme attributes, and the scheme-less fallback now requires an
ISBN-shaped value (10/13 digits, optional separators/trailing X) so
URIs like the Gutenberg identifiers cannot masquerade as ISBNs -
observed live on 'A Study in Scarlet'.
- Series: EPUB3 belongs-to-collection with collection-type=series and
group-position refines, ahead of the classic calibre:series metas.
- Audiobookshelf metadata.json sidecars join their subtitle field into
the title the same way.
Tests cover title-type main+subtitle joining, belongs-to-collection
series with fractional group-position, urn:isbn extraction, genre/tag
parity, and comma preservation inside subject headings.
The cover lookup scraped the OPF with attribute-order-sensitive regexes.
Real books serialize attributes in any order - Grand Central's '3 Days to
Live' puts href before id on manifest items and content before name on
the cover meta - so all three regex paths missed and the book fell
through to filename guessing, extracting no cover at all. Attribute
order is meaningless in XML; the regexes were never safe.
Replace them with a structured parse (encoding/xml, namespace and
attribute-order agnostic; see the new media_scanner_opf.go) and follow
Calibre's read_raster_cover resolution order:
1. manifest item with properties=cover-image (non-(X)HTML media only)
2. <meta name=cover> resolved through the manifest, same media guard
3. first spine item that is itself a raster image (store manga)
4. NEW cover-page fallback: books declaring no raster cover at all -
the classic EPUB2/Adobe cover.xhtml wrapper - are mined for
<img src> / SVG <image xlink:href> references (Calibre renders the
page with Qt; extracting the referenced image covers the practical
cases without a rendering engine)
5. existing zip filename guessing stays as the last resort, and the old
regex chain survives as findCoverInOPFLegacy for OPFs too malformed
for a real XML parse.
Hrefs are now URL-decoded and posix-normalized against the OPF's own
path (path.Join semantics), so '../art/cover.jpg' from a nested cover
page and %20-encoded names resolve correctly.
Tests: attribute-order chaos modeled on the failing Patterson book,
SVG-wrapped cover pages via guide references, image-first spines, and
path resolution edge cases. Verified live against the real
'3 Days to Live' EPUB, which previously produced no cover.
Days an archived item is kept with its reading history before library
scans purge it for good (default 90). Set 0 to keep archived items until
purged manually from the library admin page.
Bulk escape hatch for archived rows (files missing from disk for 2+
scans) so a mass external deletion never has to wait out the retention
window or be clicked away row by row:
- POST /api/media-items/purge-archived (admin only) hard-deletes all
archived items and returns the purged count; reading history goes with
the rows, so the call is confirmed in the UI first.
- The library admin page shows an 'Archived items: N' card (only when
non-zero) with a Purge Archived Now button that calls the endpoint,
toasts the result, and reloads.
- Frontend admin JS exposes window.purgeArchivedItems following the
existing localStorage-bearer-token pattern.
Replace the silent hard-delete orphan cleanup (which logged only through
ScannerLogger file logs and whose failure paths left rows undetected)
with an archive lifecycle that preserves reading history:
- A file missing in one scan is marked (missing_scan_count = 1); missing
in a second consecutive scan archives it (archived_at, hidden from
browsing, progress/notes/highlights survive). Every branch logs to
stdout with an [ARCHIVE] prefix so skips are always visible.
- When a file reappears - same path, or identical content at a new path
via the SHA-256 dedup match - the archived state clears automatically
and the item returns with its history intact.
- Archived rows older than ARCHIVE_RETENTION_DAYS are hard-purged at
scan time (cascading deletes); 0 disables auto-purge for manual-only
management. Retention is read from the environment in NewMediaScanner.
Add the storage behind the archive-instead-of-delete lifecycle:
- media_items.missing_scan_count (INT, default 0) and archived_at
(TIMESTAMPTZ, partial index), both as idempotent ADD COLUMN IF NOT
EXISTS backfills for existing installs.
- MarkMediaItemMissing / ArchiveMediaItem / ClearMediaItemArchive plus
PurgeExpiredArchivedMediaItems (retention cutoff) and
PurgeAllArchivedMediaItems (manual bulk), with CountArchivedMediaItems
for the admin UI.
- Archived items are hidden from every user-facing listing:
ListMediaItems, ListMediaItemsByLibrary, ListMediaItemsSorted,
SearchMediaItems, SearchMediaItemsUnified, the next_books CTE, the
library media counts, and the search autocomplete value lists. Detail
lookups by id/path/sha are intentionally unfiltered, and a dedicated
ListMediaItemsByLibraryIncludingArchived feeds the scanner so the
lifecycle pass can see and restore archived rows.
Libraries managed by Audiobookshelf keep a metadata.json next to each
book (title, authors, series+sequence, genres/tags, publisher,
description, isbn/asin, language, published year/date) - and no
metadata.opf. The scanner silently ignored those files: deleting them
changed nothing, and their data never reached the database.
Parse them as a first-class sidecar in extractMetadata, priority
metadata.opf -> metadata.json -> embedded media. Only fields with a
matching media_items column are mapped; narrators, subtitle, explicit,
abridged, and chapters are deliberately skipped.
Cover handling is unchanged: the existing findSidecarCover priority
(cover.jpg / folder.jpg / {basename}.jpg) applies to the sidecar branch
exactly as it does for Calibre.
go-epub's ReadBook parses every spine chapter and fails the entire call
if any single chapter (or the TOC) is malformed, discarding already-parsed
OPF metadata. For books like Pragmatic's 'A Common-Sense Guide' the OPF
holds good title/author/publisher/ISBN metadata that rescans then wrote
as blanks - success toast, no (visible) change.
Refactor to parse the EPUB's own OPF document with the same Dublin Core
machinery used for Calibre sidecars: parseCalibreMetadataOPF is now a thin
file wrapper around a reusable parseOPFContent([]byte), and the OPF lookup
previously inline in extractEPUBCover is shared via findOPFPathInZip. Since
metadata never touches chapter bodies, chapter damage cannot blank it.
Also picked up along the way: dc:language mapping and scheme-less
dc:identifier values that normalize to a valid ISBN (EPUB3 style).
Tests cover a Pragmatic-style EPUB (dc namespace on <metadata>, no
identifier scheme, deliberately malformed chapter) that must still yield
full metadata, plus series/date/subject OPF parsing.
Complete the metadata editor modal to match the backend override work:
- File Path: read-only, monospace field showing the resolved on-disk
location (MediaDetail.FileLocation), next to Format and File Size, so
admins can see exactly which file backs the record without leaving the
editor. Renders empty for non-admins, who never receive the path.
- Reset to Scanned: new button beside Rescan. It confirms, then calls
POST /api/media-items/:id/rescan?reset_overrides=true to drop all
per-field user overrides and re-extract scanned defaults. Rescan alone
keeps customizations; Reset discards them.
Generate Cover has never worked: it fetched the book file using a URL
scraped from the cover preview <img> tag (so it downloaded either the
existing cover JPEG or, when no cover existed, the detail page HTML),
then handed it to foliate-js, which rejects both. Its fixed-layout path
also called view.renderer.renderPage(), a method that does not exist in
the pinned foliate fork. Every click ended in the same generic 'Cover
generation failed' toast.
The working alternative already exists server-side: the scanner's
PDF/EPUB cover extraction plus the per-book Rescan button, now that the
rasterizer renders the CropBox. Users who want a specific image can
upload one.
Delete web/src/cover-generator.ts, the modal buttons, and the dead
generateCover()/coverGenerating/fileUrl plumbing in book-detail.ts.
The book detail page exposed server internals and unusable controls to
everyday users: it now computes and renders the book's absolute on-disk
location, and gates all of it behind the admin role.
- The page handler resolves library folder + relative path (verified
with os.Stat; falls back to the relative path when the file is not
found on disk) into the new MediaDetail.FileLocation field - only for
admins, so the absolute path never leaves the server for regular
users. This also makes it possible to locate sparse entries whose
metadata rows are largely blank.
- The Metadata grid gains a full-width, monospace, click-selectable
Location row (admins only).
- The Edit button and the MetadataEditorModal markup render only for
admins. The modal drives admin-only endpoints (metadata PUT, rescan,
reset), so non-admins previously saw a button and a form that could
only ever fail with 403s.
Previously both the metadata editor (PUT /api/media-items/:id) and the
scanner (library scans, force rescans, per-book rescans) wrote through
the same unconditional UPDATE media_items query, so any rescan wiped
user-written descriptions, tags, and uploaded covers. Custom and scanned
values were indistinguishable, and custom cover uploads even wrote to
the same {file}.cover.jpg sidecar path the scanner generates, so each
side silently clobbered the other.
Introduce metadata_overrides, a TEXT[] column on media_items listing the
column names the user has customized:
- Saving metadata records overrides per field by diffing the submitted
values against the stored row (an untouched save records nothing);
overrides accumulate until an explicit reset. Cover upload/removal
always marks cover_image_path. Bulk updates mark each applied field.
- The scanner merges: updateMediaItem() now takes the existing row and
restores every overridden column (including derived *_search arrays)
before writing, and preserves the override set itself.
- Uploaded covers move to a dedicated {file}.custom_cover.{jpg|png|webp}
sidecar so the scanner can never overwrite a user cover on disk.
- RescanMediaItem gains resetOverrides: POST /api/media-items/:id/rescan?
reset_overrides=true clears the set first, returning the item to pure
scanned defaults.
Shared detection/restore helpers live in internal/utils
(metadata_overrides.go) with unit tests covering detection, accumulation,
unset-form equality, and restore-with-derived-fields. Schema change is
an idempotent ADD COLUMN IF NOT EXISTS applied on startup. Also includes
incidental gofmt of NewMediaScanner literals in media_scanner.go.
parseTableNames() regex-scanned every line of schema.sql, comments
included, with 'CREATE TABLE (?:IF NOT EXISTS )?(?:\w+\.)?(\w+)'. A doc
comment containing that phrase in prose - e.g. 'declared in CREATE TABLE
above' - registered a phantom table ('above'), and startup verification
then failed with 'missing tables: above', crash-looping the app
container on every restart.
Skip lines whose trimmed form starts with '--' so comments can never
contribute table names, and add a regression test asserting every parsed
table maps back to a real CREATE TABLE statement.
pdftoppm defaults to rasterizing the MediaBox, while PDF viewers (pdf.js
in the reader, and every other viewer) display the CropBox. For PDFs
whose page 1 is the full print cover wrap (back cover + spine + front
cover in one landscape page) with a CropBox covering only the front
cover - e.g. No Starch's XeTeX-built 'Algorithmic Thinking' - the
fallback stored the entire spread as a squashed landscape cover, while
the reader correctly showed just the front cover.
Pass -cropbox so the rendered cover always matches what the reader
displays. Poppler falls back to the MediaBox when a PDF defines no
CropBox, so PDFs with identical boxes (the common case) render exactly
as before.