Files
immich-photo-flow/docs/ROADMAP.md
T
m038 7745f05323 docs(roadmap): correct M1.5 spike claims (spike now run + verified)
Fixes both false claims the spec flagged: M1.5 was never 'verified in M1' and M1
did not 'run a read-only feasibility spike' (it deferred it). The spike has now
run separately and passed, so both are updated to point at the findings spec
(smart_search.embedding, 1152-dim, cosine, clean assetId->asset.id join) and the
spike is listed under Related specs.
2026-06-27 20:40:26 +02:00

6.7 KiB

immich-photo-flow — Roadmap

Last updated: 2026-06-27

This is the authoritative, living roadmap for the project. Each milestone gets its own design spec → implementation plan → build cycle under docs/superpowers/. This document is the map; the specs are the territory.

Vision

A monorepo of small, independent Python/Flask tools that clean up and structure a ~40k-image, 15-year Immich library into travel "memories." It is the upstream stage of an existing pipeline (all repos under /home/mischa/Projects/):

immich-photo-flow (trips + tags)  →  image-rater (pick best → 4+ album)  →  travel-memories (album → Grav blog)

Every handoff is through Immich itself (tags / albums / ratings) — no app calls another directly.

Core principle — division of labor: AI/algorithms propose; the human does QA at the cluster/group level (mostly in bulk, attention concentrated where confidence is low) and owns content creation. Nothing is silently auto-applied to a 15-year library of memories.

End state

One monorepo:

shared/{immich, core (SQLite), ai, ui}
apps/{trip-cluster, tag-verify, enrich, image-rater, travel-memories}

Each app is independently containerized (its own Dockerfile + port). shared/* are local editable packages providing the common foundation.

Milestones

# Name Goal Status
M1 Foundation + trip-cluster Stand up the monorepo, shared packages (shared/ai deferred to M3), SQLite store, Immich ingest, and the trip-cluster app end-to-end. Validate on a hard, GPS-poor sample against a quantified acceptance bar (not the easy last trip) before widening. This is the POC that proves the foundation. Shipped — implemented, reviewed, merged to main (2026-06-27); 66 tests green. Validation-gate run on a hard sample still pending.
M1.5 Visual similarity Trip-level CLIP clustering over Immich's pgvector embeddings (readability/shape verified by the pgvector spike — see findings) — the rescue signal for the GPS-poor old library where timestamp/GPS signals are weakest. Not started; dependency spike done (2026-06-27, embeddings readable + joinable — see findings spec)
M2 tag-verify Verify/normalize existing tags, dedupe the tag vocabulary, find outliers — on the proven foundation. Not started
M3 enrich Geocode existing location tags ("Kiev" → coordinates) and backfill GPS into Immich to improve future trip detection; AI captioning; content tags; noise detection. Not started
M4 Migrate image-rater Move the existing app onto the shared packages (+ optional SQLite). Largely mechanical (delete local copies, import shared). Not started
M5 Migrate travel-memories Same migration for the blog-export app. Not started
M6 Non-trip categorization (future/optional) The deferred full-library taxonomy for everyday/non-trip photos. Deferred

Sequencing rationale: Build the new categorizer first as the POC that earns the shared foundation (Option 1), then migrate the existing apps onto it (Option 2 end-state) — in milestones, not all at once. Doing shared/ui + shared/immich well in M1 makes M4/M5 mostly mechanical; the existing apps already duplicate exactly what those packages absorb.

Cross-cutting decisions & conventions

These apply across all milestones (decided during the 2026-06-27 brainstorm):

  • Immich is the source of truth. SQLite is a rebuildable working/review layer; all durable results are written back to Immich as tags.
  • State store: SQLite (shared/core), not the JSON-file pattern of the existing apps — chosen for easy cross-trip filtering at 40k scale. Designed cleanly enough that the other apps can migrate onto it.
  • Stack/conventions mirror the sibling apps: Python 3.12 + Flask + argparse CLI; no-build frontend (Jinja + DaisyUI/Tailwind/Alpine/HTMX); one module talks to Immich, one to Anthropic; pytest + pytest-httpserver + Playwright; TDD; one Docker container per app; user: ${UID}:${GID}; state on a mounted volume.
  • Ports: travel-memories 8082, image-rater 8083, trip-cluster 8084.
  • Tag conventions:
    • Content/organizational tags (trip names, locations, people) are user-facing, follow the existing convention, and are never namespaced.
    • Pipeline meta-tags are nested under a single parent _pipeline/ so they can be removed wholesale and never clutter the tag list: _pipeline/processed, _pipeline/non-trip, _pipeline/ai-rating/<0-5> (image-rater after M4), etc. Defined as a shared convention in shared/immich.
    • Always reconcile with the live Immich instance before writing tags (e.g. image-rater's existing un-namespaced ai-rating/<n>).
  • Trip detection is timestamp-first, anchored by existing trip/location tags, refined by GPS where present. GPS is a minority signal (old camera photos lack it).
  • Resumable / idempotent / scopeable throughout: this is a long-running, vet-first, work-over-time effort; _pipeline/processed + writeback_log make work re-derivable and re-runnable; ingest is scopeable (date range / tag / subset) so tools are validated on one trip before the backlog.

Enablement notes (things to set up to unlock later milestones)

  • Visual similarity (M1.5; dependency verified by the pgvector spike, 2026-06-27): Immich's REST API doesn't cleanly expose CLIP embeddings. Viable path = read-only access to Immich's Postgres pgvector embeddings. A read-only feasibility spike (scripts/pgvector_spike.py, tracked separately from M1 — not run inside M1) confirmed embeddings are readable and pinned the shape against the live DB: smart_search.embedding, 1152-dim, cosine <=>, assetId → asset.id join clean. See the findings spec for the version-pinned contract. The trip-level clustering itself is M1.5. User provides DB access (IMMICH_DB_URL) and is re-running CLIP with a stronger model — both pure upside, especially for the GPS-poor old library.
  • GPS facts: Immich's reverse-geocoding only labels coordinates a photo already has; it does not invent GPS, and Immich cannot infer GPS from image content. For GPS-less old photos, coordinates come from manual map placement or the M3 enrich step (geocode location tags → write back as GPS). Reverse-geocoding + metadata-extraction jobs can be run anytime to strengthen location anchors for the GPS-having subset.
  • M1: docs/superpowers/specs/2026-06-27-immich-photo-flow-design.md
  • M1.5 pgvector spike — design: docs/superpowers/specs/2026-06-27-pgvector-embedding-spike-design.md
  • M1.5 pgvector spike — findings (M1.5's contract): docs/superpowers/specs/2026-06-27-pgvector-embedding-findings.md