Files
bookhoard/SCANNER_RESTORATION_PLAN.md
T
john-okeefe 2258bb2f5d docs(ebook): add comprehensive ebook.go refactor plan
- Add EBOOK_REFACTOR_PLAN.md with 856 lines of detailed instructions
- Split 1,350-line ebook.go into focused single-responsibility files
- Zero API changes, only code organization for maintainability
- Phase-by-phase safety checkpoints and rollback procedures

Target file organization after refactor:
- media.go (~600 lines): Media CRUD + metadata
- search.go (~80 lines): Query and search operations
- matching.go (~200 lines): Book matching and sync operations
- ebook.go (~150 lines): SetupRoutes only

Plan ensures AI can implement without breaking any functionality.
2026-02-07 00:16:20 -05:00

19 KiB

Scanner System Fix & Enhancement Plan

Created: February 6, 2026 Updated: February 6, 2026 (clarified scope and priorities) Status: Ready to implement


Executive Summary

This plan addresses TWO separate issues with the scanner system:

1. CRITICAL BUG (Priority: HIGH)

Auto-start functionality was removed during router refactor (commit 6784c25):

  • Scheduler doesn't start on server boot
  • Watch mode doesn't auto-start for libraries
  • No graceful shutdown for scanner services
  • handlers.SetupRoutes() returns Handler but return value isn't captured

2. ENHANCEMENTS (Priority: MEDIUM)

  • Library-type-aware scanning (prevents cross-contamination between ebook/comic/manga libraries)
  • Comic/manga metadata extraction (ComicInfo.xml parsing)
  • Better code organization (extract scanner routes to separate file)

What WASN'T broken:

  • Scanner routes still work (7 endpoints in internal/handlers/ebook.go:145-154)
  • Can manually start/stop scanner via API
  • Scan job tracking works
  • Watch mode works when manually triggered

Changes Summary

Files to Create:

  1. internal/router/scanner.go (new - optional, for organization)
  2. internal/app/app.go (new - required, for lifecycle management)

Files to Modify:

  1. internal/router/router.go (capture and return EbookHandler)
  2. cmd/server/main.go (restore auto-start calls, use App pattern)
  3. internal/services/ebook_scanner.go (library-type-aware scanning)
  4. internal/handlers/ebook.go (comic/manga metadata extraction)

Risk Assessment: LOW-MEDIUM

  • Auto-start restoration: LOW (restores existing code that was removed)
  • Library-type-awareness: MEDIUM (core scanner logic change)
  • Comic/manga scanning: MEDIUM (new feature)

Phase 1: Restore Auto-Start Functionality [CRITICAL BUG FIX]

Problem

Before router refactor (commit 799b640):

h := handlers.SetupRoutes(protected, queries)

// Start scheduler for auto-scanning
go h.StartScheduler()
defer h.StopScheduler()

// Start watch mode for all libraries (background)
go func() {
    time.Sleep(2 * time.Second)
    if err := h.StartWatchModeForAllLibraries(context.Background()); err != nil {
        log.Printf("Warning: failed to start watch mode for libraries: %v", err)
    }
}()

After refactor (current):

handlers.SetupRoutes(protected, cfg.Queries, cfg.ConnManager)  // Return value ignored!
// No scheduler start
// No watch mode start
// No graceful shutdown

Solution

Step 1.1: Update internal/router/router.go

Location: Line 113-114

Current Code:

jwtMiddleware := createJWTMiddleware(cfg)
protected := e.Group("/api", jwtMiddleware)
handlers.SetupRoutes(protected, cfg.Queries, cfg.ConnManager)

New Code:

jwtMiddleware := createJWTMiddleware(cfg)
protected := e.Group("/api", jwtMiddleware)
ebookHandler := handlers.SetupRoutes(protected, cfg.Queries, cfg.ConnManager)
return ebookHandler  // Add return statement to RegisterRoutes

Update RegisterRoutes signature:

Current (line 86):

func RegisterRoutes(cfg *Config) {

New:

func RegisterRoutes(cfg *Config) *handlers.Handler {
    // ... existing code ...
    return ebookHandler  // Return at end of function
}

Step 1.2: Update cmd/server/main.go

Location: After line 151 (after router.RegisterRoutes(routerConfig))

Add this code:

// Register all routes and get ebook handler
ebookHandler := router.RegisterRoutes(routerConfig)

// ========================================================================
// BACKGROUND SERVICES - Restore auto-start functionality
// ========================================================================

// Start scheduler for auto-scanning
go ebookHandler.StartScheduler()
defer ebookHandler.StopScheduler()

// Start watch mode for all libraries (background)
go func() {
    time.Sleep(2 * time.Second) // Wait for server to be ready
    if err := ebookHandler.StartWatchModeForAllLibraries(context.Background()); err != nil {
        log.Printf("Warning: failed to start watch mode for libraries: %v", err)
    }
}()

Verification:

  • Application compiles
  • Server starts without errors
  • Check logs for "Starting scheduler" message
  • Check logs for watch mode starting after 2 seconds

Phase 2: Implement Library-Type-Aware Scanning [ENHANCEMENT]

Purpose

Prevent cross-contamination between library types:

  • Epub libraries should only scan .epub files
  • Comic libraries should only scan .cbz/.cbr files
  • Manga libraries should only scan appropriate formats
  • Each library type has configurable allowed extensions

File: internal/services/ebook_scanner.go

Change 2.1: Add libraryTypes cache field

Location: Line 59-65 (EbookScanner struct)

Current Code:

type EbookScanner struct {
    db               *database.Queries
    watcher          *fsnotify.Watcher
    folders          []string
    adminID          pgtype.UUID
    defaultLibraryID pgtype.UUID
}

New Code:

type EbookScanner struct {
    db               *database.Queries
    watcher          *fsnotify.Watcher
    folders          []string
    adminID          pgtype.UUID
    defaultLibraryID pgtype.UUID
    libraryTypes     map[string][]string  // folder -> allowed extensions cache
}

Change 2.2: Initialize libraryTypes in NewEbookScanner

Location: Line 73-74

New Code:

return &EbookScanner{
    db:               db,
    watcher:          watcher,
    folders:          []string{},
    adminID:          pgtype.UUID{},
    defaultLibraryID: pgtype.UUID{Valid: false},
    libraryTypes:     make(map[string][]string),  // ← ADD THIS
}

Change 2.3: Build library types cache in SetFolders

Location: Line 86-109 (SetFolders function)

New Code:

func (s *EbookScanner) SetFolders(folders []string) error {
    s.folders = folders

    // Remove old watch if exists
    if s.watcher != nil {
        s.watcher.Close()
    }

    // Create new watcher
    watcher, err := fsnotify.NewWatcher()
    if err != nil {
        return fmt.Errorf("failed to create watcher: %v", err)
    }
    s.watcher = watcher

    // Build cache of allowed extensions per folder
    s.libraryTypes = make(map[string][]string)
    ctx := context.Background()

    for _, folder := range folders {
        // Get library for this folder
        lib, err := s.db.GetLibraryByFolder(ctx, folder)
        if err != nil {
            fmt.Printf("Warning: failed to get library for folder %s: %v\n", folder, err)
            continue
        }

        // Get library type with allowed extensions
        libType, err := s.db.GetLibraryType(ctx, lib.LibraryTypeID)
        if err != nil {
            fmt.Printf("Warning: failed to get library type for %s: %v\n", folder, err)
            continue
        }

        // Cache allowed extensions for this folder
        s.libraryTypes[folder] = libType.AllowedExtensions
        fmt.Printf("Scanner: Folder %s (type: %s) allows extensions: %v\n",
            folder, libType.Name, libType.AllowedExtensions)
    }

    // Add all folders to watch
    for _, folder := range folders {
        if err := s.watcher.Add(folder); err != nil {
            fmt.Printf("Warning: failed to watch folder %s: %v\n", folder, err)
        }
    }

    return nil
}

Change 2.4: Replace isEbookFile with isScannableFile

Location: Line 168-177 (isEbookFile function)

New Code:

// isScannableFile checks if a file should be scanned based on library type configuration
func (s *EbookScanner) isScannableFile(path string) bool {
    ext := strings.ToLower(filepath.Ext(path))

    // Find which folder this file belongs to
    var folder string
    for _, f := range s.folders {
        if strings.HasPrefix(path, f) {
            folder = f
            break
        }
    }

    // If no folder match, don't scan
    if folder == "" {
        return false
    }

    // Get allowed extensions for this folder's library
    allowed, ok := s.libraryTypes[folder]
    if !ok {
        // No library type info, skip file
        fmt.Printf("Warning: No library type info for folder %s, skipping %s\n", folder, path)
        return false
    }

    // Check if file extension is allowed for this library type
    for _, allowedExt := range allowed {
        if ext == strings.ToLower(allowedExt) {
            return true
        }
    }

    return false
}

Change 2.5: Update ScanFolders to use isScannableFile

Location: Line 146 (in ScanFolders function)

Current Code:

// Check if it's an ebook file
if s.isEbookFile(path) {

New Code:

// Check if file should be scanned based on library type
if s.isScannableFile(path) {

Verification:

  • Code compiles
  • Ebook libraries scan only .epub files (check logs)
  • Comic libraries scan only .cbz/.cbr files
  • No cross-contamination between library types
  • Test: Create ebook library, add .cbz file → should be ignored
  • Test: Create comic library, add .epub file → should be ignored

Phase 3: Comic/Manga Metadata Extraction [NEW FEATURE]

Purpose

Extract metadata from comic/manga archives (.cbz, .cbr, .cb7):

  • Parse ComicInfo.xml from archives
  • Extract cover images
  • Get series, issue number, publisher, etc.
  • Support comic library management

File: internal/handlers/ebook.go

Step 3.1: Add ComicInfo.xml parsing

Add to imports:

import (
    "archive/zip"
    "encoding/xml"
    "image"
    _ "image/jpeg"
    _ "image/png"
    "path/filepath"
    "strings"
)

Step 3.2: Define ComicInfo struct

Add after existing structs:

// ComicInfo represents metadata from ComicInfo.xml
type ComicInfo struct {
    XMLName    xml.Name `xml:"ComicInfo"`
    Title      string   `xml:"Title"`
    Series     string   `xml:"Series"`
    Number     int      `xml:"Number"`
    Volume     int      `xml:"Volume"`
    Publisher  string   `xml:"Publisher"`
    Year       int      `xml:"Year"`
    Month      int      `xml:"Month"`
    Day        int      `xml:"Day"`
    Writer     string   `xml:"Writer"`
    Penciller  string   `xml:"Penciller"`
    Inker      string   `xml:"Inker"`
    Colorist   string   `xml:"Colorist"`
    Letterer   string   `xml:"Letterer"`
    CoverArtist string  `xml:"CoverArtist"`
    Genre      string   `xml:"Genre"`
    Tags       string   `xml:"Tags"`
    Web        string   `xml:"Web"`
    Notes      string   `xml:"Notes"`
}

Step 3.3: Add extraction function

Add new function:

// extractComicMetadata extracts metadata from comic archive
func extractComicMetadata(filePath string) (*ComicInfo, []byte, error) {
    // Open archive
    r, err := zip.OpenReader(filePath)
    if err != nil {
        return nil, nil, fmt.Errorf("failed to open comic archive: %w", err)
    }
    defer r.Close()

    // Look for ComicInfo.xml
    var comicInfo *ComicInfo
    var coverImage []byte

    for _, f := range r.File {
        if f.Name == "ComicInfo.xml" {
            rc, err := f.Open()
            if err != nil {
                return nil, nil, fmt.Errorf("failed to open ComicInfo.xml: %w", err)
            }

            data, err := io.ReadAll(rc)
            rc.Close()
            if err != nil {
                return nil, nil, fmt.Errorf("failed to read ComicInfo.xml: %w", err)
            }

            comicInfo = &ComicInfo{}
            if err := xml.Unmarshal(data, comicInfo); err != nil {
                return nil, nil, fmt.Errorf("failed to parse ComicInfo.xml: %w", err)
            }
        }

        // Look for cover image (usually first image in root)
        if coverImage == nil && isImageFile(f.Name) {
            // Usually in root directory, not subdirectories
            if !strings.Contains(filepath.Dir(f.Name), string(filepath.Separator)) ||
                filepath.Dir(f.Name) == "." {
                rc, err := f.Open()
                if err != nil {
                    continue
                }

                coverImage, err = io.ReadAll(rc)
                rc.Close()
                if err == nil {
                    // Validate it's actually an image
                    _, _, err = image.Decode(bytes.NewReader(coverImage))
                    if err != nil {
                        coverImage = nil // Not a valid image
                    }
                }
            }
        }
    }

    if comicInfo == nil {
        // No ComicInfo.xml, create minimal metadata from filename
        comicInfo = &ComicInfo{}
        basename := filepath.Base(filePath)
        comicInfo.Title = strings.TrimSuffix(basename, filepath.Ext(basename))
    }

    return comicInfo, coverImage, nil
}

// isImageFile checks if a file is an image based on extension
func isImageFile(filename string) bool {
    ext := strings.ToLower(filepath.Ext(filename))
    return ext == ".jpg" || ext == ".jpeg" || ext == ".png" || ext == ".gif"
}

Step 3.4: Integrate into scanner

Update ScanFolders to extract comic metadata for .cbz/.cbr files:

Location: In the file processing loop (around line 160-180)

Add before creating media item:

var comicInfo *ComicInfo
var coverImage []byte

// Extract comic metadata if applicable
if strings.ToLower(filepath.Ext(path)) == ".cbz" {
    info, cover, err := extractComicMetadata(path)
    if err != nil {
        log.Printf("Warning: failed to extract comic metadata from %s: %v", path, err)
    } else {
        comicInfo = info
        coverImage = cover
    }
}

// When creating media item, use comic metadata
title := comicInfo.Title
if title == "" {
    title = filepath.Base(path)
}

// Use cover image if available
if len(coverImage) > 0 {
    // Use extracted cover
    // ... existing cover processing code ...
}

Verification:

  • Code compiles
  • Create .cbz file with ComicInfo.xml
  • Scan comic library
  • Check metadata was extracted (title, series, issue)
  • Check cover image was extracted
  • Test .cbz without ComicInfo.xml (should use filename)

Phase 4: Organize Scanner Routes [OPTIONAL - LOW PRIORITY]

Purpose

Move scanner routes from internal/handlers/ebook.go to internal/router/scanner.go for better organization.

Note: This is purely cosmetic. Scanner routes already work fine where they are.

File: internal/router/scanner.go (NEW)

package router

import (
    "github.com/labstack/echo/v4"
)

// registerScannerRoutes registers all scanner-related endpoints
func registerScannerRoutes(cfg *Config) {
    // Routes are already registered in handlers.SetupRoutes()
    // This file is for documentation/organization purposes
    // Actual routes are in:
    // - internal/handlers/ebook.go:145-154 (scanner endpoints)
    // - internal/handlers/auth.go (scan settings endpoints)
}

Decision: SKIP this phase. The current organization works fine.


Phase 5: Application Lifecycle Management [OPTIONAL - DEFER]

Purpose

Create internal/app/app.go for better lifecycle management, graceful shutdown, signal handling.

Decision: DEFER to future implementation. The simple approach in Phase 1 is sufficient for now.


Implementation Order

Priority 1: Critical Bug Fix (Phase 1)

  • Time: 15 minutes
  • Risk: LOW
  • Impact: Restores auto-scan and watch mode

Priority 2: Library-Type-Awareness (Phase 2)

  • Time: 1 hour
  • Risk: MEDIUM
  • Impact: Prevents cross-contamination

Priority 3: Comic/Manga Scanning (Phase 3)

  • Time: 2-3 hours
  • Risk: MEDIUM
  • Impact: New feature for comic libraries

Priority 4: Code Organization (Phase 4)

  • Time: 30 minutes
  • Risk: LOW
  • Impact: Cosmetic (SKIP for now)

Priority 5: App Lifecycle (Phase 5)

  • Time: 2 hours
  • Risk: MEDIUM
  • Impact: Better structure (DEFER for now)

Testing Checklist

After Phase 1 (Auto-start restoration):

  • Application compiles
  • Server starts without errors
  • Logs show "Starting scheduler"
  • Logs show "Starting watch mode for all libraries" after 2 seconds
  • Scheduler triggers auto-scans
  • SIGTERM triggers graceful shutdown

After Phase 2 (Library-type-awareness):

  • Create ebook library, add .epub → scans correctly
  • Create ebook library, add .cbz → ignored
  • Create comic library, add .cbz → scans correctly
  • Create comic library, add .epub → ignored
  • Check logs for library type messages

After Phase 3 (Comic/manga scanning):

  • Create test .cbz with ComicInfo.xml
  • Add to comic library
  • Scan library
  • Verify metadata extracted (title, series, issue)
  • Verify cover image extracted
  • Test .cbz without ComicInfo.xml (uses filename)

Verification Commands

Test scanner auto-start:

# Start server
podman-compose up -d --build

# Check logs
podman logs bookhoard | grep -i "scheduler\|watch mode"

# Should see:
# "Starting scheduler for auto-scanning"
# "Starting watch mode for all libraries"

Test library-type-awareness:

# Create ebook library
curl -X POST http://localhost:8765/api/libraries \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name":"Ebooks","library_type_id":"<ebook-uuid>","folders":["/path/to/ebooks"]}'

# Create comic library
curl -X POST http://localhost:8765/api/libraries \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"name":"Comics","library_type_id":"<comic-uuid>","folders":["/path/to/comics"]}'

# Add .epub to comic library → should be ignored
# Add .cbz to ebook library → should be ignored
# Check scan logs for filtering messages

Test comic metadata extraction:

# Create test .cbz with ComicInfo.xml
zip test.cbz ComicInfo.xml cover.jpg page1.jpg

# Add to comic library and scan
curl -X POST http://localhost:8765/api/scanner/scan \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"folder_paths":["/path/to/comics"]}'

# Check media item has correct metadata
curl http://localhost:8765/api/media-items?library_id=<comic-lib-id> \
  -H "Authorization: Bearer $TOKEN" | jq .

Rollback Plan

Rollback Phase 1:

git checkout cmd/server/main.go
git checkout internal/router/router.go

Rollback Phase 2:

git checkout internal/services/ebook_scanner.go

Rollback Phase 3:

git checkout internal/handlers/ebook.go

Summary

Critical Issues Fixed:

  • Auto-start scheduler (restored)
  • Auto-start watch mode (restored)
  • Graceful shutdown (restored)

Enhancements Added:

  • Library-type-aware scanning (prevents cross-contamination)
  • Comic/manga metadata extraction (ComicInfo.xml parsing)
  • Comic cover image extraction

Deferred:

  • ⏸️ Code organization (scanner routes file)
  • ⏸️ App lifecycle management

Estimated Time:

  • Phase 1 (bug fix): 15 minutes
  • Phase 2 (enhancement): 1 hour
  • Phase 3 (new feature): 2-3 hours
  • Total: 4-5 hours

Risk Level:

  • Phase 1: LOW (restoring removed code)
  • Phase 2: MEDIUM (core scanner logic)
  • Phase 3: MEDIUM (new feature, isolated)

Ready for implementation in priority order.