Fixes both false claims the spec flagged: M1.5 was never 'verified in M1' and M1 did not 'run a read-only feasibility spike' (it deferred it). The spike has now run separately and passed, so both are updated to point at the findings spec (smart_search.embedding, 1152-dim, cosine, clean assetId->asset.id join) and the spike is listed under Related specs.
6.7 KiB
immich-photo-flow — Roadmap
Last updated: 2026-06-27
This is the authoritative, living roadmap for the project. Each milestone gets its own design spec → implementation plan → build cycle under docs/superpowers/. This document is the map; the specs are the territory.
Vision
A monorepo of small, independent Python/Flask tools that clean up and structure a ~40k-image, 15-year Immich library into travel "memories." It is the upstream stage of an existing pipeline (all repos under /home/mischa/Projects/):
immich-photo-flow (trips + tags) → image-rater (pick best → 4+ album) → travel-memories (album → Grav blog)
Every handoff is through Immich itself (tags / albums / ratings) — no app calls another directly.
Core principle — division of labor: AI/algorithms propose; the human does QA at the cluster/group level (mostly in bulk, attention concentrated where confidence is low) and owns content creation. Nothing is silently auto-applied to a 15-year library of memories.
End state
One monorepo:
shared/{immich, core (SQLite), ai, ui}
apps/{trip-cluster, tag-verify, enrich, image-rater, travel-memories}
Each app is independently containerized (its own Dockerfile + port). shared/* are local editable packages providing the common foundation.
Milestones
| # | Name | Goal | Status |
|---|---|---|---|
| M1 | Foundation + trip-cluster |
Stand up the monorepo, shared packages (shared/ai deferred to M3), SQLite store, Immich ingest, and the trip-cluster app end-to-end. Validate on a hard, GPS-poor sample against a quantified acceptance bar (not the easy last trip) before widening. This is the POC that proves the foundation. |
Shipped — implemented, reviewed, merged to main (2026-06-27); 66 tests green. Validation-gate run on a hard sample still pending. |
| M1.5 | Visual similarity | Trip-level CLIP clustering over Immich's pgvector embeddings (readability/shape verified by the pgvector spike — see findings) — the rescue signal for the GPS-poor old library where timestamp/GPS signals are weakest. | Not started; dependency spike done (2026-06-27, embeddings readable + joinable — see findings spec) |
| M2 | tag-verify |
Verify/normalize existing tags, dedupe the tag vocabulary, find outliers — on the proven foundation. | Not started |
| M3 | enrich |
Geocode existing location tags ("Kiev" → coordinates) and backfill GPS into Immich to improve future trip detection; AI captioning; content tags; noise detection. | Not started |
| M4 | Migrate image-rater |
Move the existing app onto the shared packages (+ optional SQLite). Largely mechanical (delete local copies, import shared). | Not started |
| M5 | Migrate travel-memories |
Same migration for the blog-export app. | Not started |
| M6 | Non-trip categorization (future/optional) | The deferred full-library taxonomy for everyday/non-trip photos. | Deferred |
Sequencing rationale: Build the new categorizer first as the POC that earns the shared foundation (Option 1), then migrate the existing apps onto it (Option 2 end-state) — in milestones, not all at once. Doing shared/ui + shared/immich well in M1 makes M4/M5 mostly mechanical; the existing apps already duplicate exactly what those packages absorb.
Cross-cutting decisions & conventions
These apply across all milestones (decided during the 2026-06-27 brainstorm):
- Immich is the source of truth. SQLite is a rebuildable working/review layer; all durable results are written back to Immich as tags.
- State store: SQLite (
shared/core), not the JSON-file pattern of the existing apps — chosen for easy cross-trip filtering at 40k scale. Designed cleanly enough that the other apps can migrate onto it. - Stack/conventions mirror the sibling apps: Python 3.12 + Flask + argparse CLI; no-build frontend (Jinja + DaisyUI/Tailwind/Alpine/HTMX); one module talks to Immich, one to Anthropic; pytest + pytest-httpserver + Playwright; TDD; one Docker container per app;
user: ${UID}:${GID}; state on a mounted volume. - Ports: travel-memories 8082, image-rater 8083, trip-cluster 8084.
- Tag conventions:
- Content/organizational tags (trip names, locations, people) are user-facing, follow the existing convention, and are never namespaced.
- Pipeline meta-tags are nested under a single parent
_pipeline/so they can be removed wholesale and never clutter the tag list:_pipeline/processed,_pipeline/non-trip,_pipeline/ai-rating/<0-5>(image-rater after M4), etc. Defined as a shared convention inshared/immich. - Always reconcile with the live Immich instance before writing tags (e.g. image-rater's existing un-namespaced
ai-rating/<n>).
- Trip detection is timestamp-first, anchored by existing trip/location tags, refined by GPS where present. GPS is a minority signal (old camera photos lack it).
- Resumable / idempotent / scopeable throughout: this is a long-running, vet-first, work-over-time effort;
_pipeline/processed+writeback_logmake work re-derivable and re-runnable; ingest is scopeable (date range / tag / subset) so tools are validated on one trip before the backlog.
Enablement notes (things to set up to unlock later milestones)
- Visual similarity (M1.5; dependency verified by the pgvector spike, 2026-06-27): Immich's REST API doesn't cleanly expose CLIP embeddings. Viable path = read-only access to Immich's Postgres pgvector embeddings. A read-only feasibility spike (
scripts/pgvector_spike.py, tracked separately from M1 — not run inside M1) confirmed embeddings are readable and pinned the shape against the live DB:smart_search.embedding, 1152-dim, cosine<=>,assetId → asset.idjoin clean. See the findings spec for the version-pinned contract. The trip-level clustering itself is M1.5. User provides DB access (IMMICH_DB_URL) and is re-running CLIP with a stronger model — both pure upside, especially for the GPS-poor old library. - GPS facts: Immich's reverse-geocoding only labels coordinates a photo already has; it does not invent GPS, and Immich cannot infer GPS from image content. For GPS-less old photos, coordinates come from manual map placement or the M3
enrichstep (geocode location tags → write back as GPS). Reverse-geocoding + metadata-extraction jobs can be run anytime to strengthen location anchors for the GPS-having subset.
Related specs
- M1:
docs/superpowers/specs/2026-06-27-immich-photo-flow-design.md - M1.5 pgvector spike — design:
docs/superpowers/specs/2026-06-27-pgvector-embedding-spike-design.md - M1.5 pgvector spike — findings (M1.5's contract):
docs/superpowers/specs/2026-06-27-pgvector-embedding-findings.md