Files
immich-photo-flow/docs/superpowers/specs/2026-06-27-immich-photo-flow-design.md
T
2026-06-27 15:04:44 +02:00

209 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# immich-photo-flow — Design (M1: Foundation + Trip-Cluster)
- **Date:** 2026-06-27
- **Status:** Approved design, pending implementation plan
- **Scope of this spec:** Milestone M1 only (shared foundation + the `trip-cluster` app). Later milestones are documented as a roadmap, each to get its own spec→plan→build cycle.
## Purpose
A monorepo of small, independent tools that clean up and structure a ~40k-image, 15-year [Immich](https://immich.app) library so it can become travel "memories." This is the **upstream** stage of an existing pipeline:
```
[NEW categorizer: trips + tags] → image-rater (pick best → 4+ album) → travel-memories (album → Grav blog)
```
Every handoff is through **Immich itself** (tags, albums, ratings) — no app talks to another directly. The categorizer's job is to leave clean **trip tags** (and explicit **non-trip** marks) in Immich for `image-rater` to consume by tag.
### Division of labor (the core principle)
AI/algorithms do what they do well; the human does QA and (later, downstream) content creation.
- The machine **proposes** groupings with a confidence score.
- The human **reviews at the cluster level** (a few dozen candidate trips), not photo-by-photo, mostly in bulk, with attention concentrated where confidence is low.
- Nothing is silently auto-applied to the library. Trust matters with 15 years of memories.
### What M1 explicitly does NOT do
- It does **not** decide which photos *within* a trip are keepers — that is `image-rater`'s job downstream.
- It does **not** fully categorize everyday/non-trip photos beyond marking them as non-trip (deferred to a future milestone).
- It does **not** draft any narrative text (the user owns content creation).
## Context & constraints (from the existing ecosystem)
The two existing apps (`claude-image-rater`, `claude-travel-memories`, both under `/home/mischa/Projects/`) establish strong, consistent conventions this project follows and builds upon:
- **Stack:** Python 3.12 + Flask (factory `create_app()`) + an argparse CLI sharing one package.
- **No-build frontend:** Jinja templates extending `base.html`, **DaisyUI 4 + Tailwind + Alpine.js** (CDN), HTMX for navigation, one `static/app.js`.
- **Single-responsibility modules:** one module is the *only* thing that talks to Immich; one is the *only* thing that talks to Anthropic.
- **Testing:** pytest + `pytest-httpserver` (mock Immich) + Playwright UI tests, **TDD** ("failing test first").
- **Docker:** one container per app, distinct port, `user: ${UID}:${GID}`, state on a mounted volume.
- **Proven UI:** travel-memories' triage screen (grid + ring-selection + keyboard navigation + fullscreen lightbox) was reused by image-rater. This becomes the canonical shared component (see §UI).
Both apps currently **duplicate** near-identical code (`immich.py`, `config.py`, `slugify`, `base.html`, thumbnail proxy route, JSON-state pattern). That duplication is the evidence that extracting a shared foundation is high-value, and that migrating those apps onto it later (M4/M5) is largely mechanical.
### Decisions locked during brainstorming
- **State store: SQLite** (not the JSON-file pattern of the existing apps). A single DB makes the cross-trip filtering this project needs ("all assets in a trip window but missing the trip tag", "everything still unreviewed") far easier. `shared/core` is designed cleanly enough that the other apps *could* migrate onto it later.
- **Immich is the source of truth.** SQLite is the working/review layer; all durable results are written back to Immich as tags.
- **Monorepo with shared packages + per-app containers** (Approach A on a shared foundation). Build the categorizer first as the POC that proves the foundation, then migrate the other apps onto it in later milestones (Option 2 end-state), in milestones — not all at once.
- **GPS is a minority signal.** Most old camera photos lack GPS, so trip detection is **timestamp-first**, anchored by existing tags, refined by GPS where present.
- **Everyday/non-trip photos:** assign trips AND positively mark non-trip so they stop resurfacing as "unreviewed." Full non-trip categorization is deferred.
## Roadmap
The full multi-milestone roadmap (M1M6), end-state, and cross-cutting decisions live in the authoritative **[`docs/ROADMAP.md`](../../ROADMAP.md)**. In brief: **this spec is M1** — Foundation + `trip-cluster`, the POC that proves the shared foundation. M2 (`tag-verify`), M3 (`enrich`), M4/M5 (migrate the existing apps onto the foundation), and M6 (non-trip categorization) follow, each with its own spec→plan→build cycle.
---
# M1 Design
## Monorepo layout
```
immich-photo-flow/
pyproject.toml # workspace; shared/* as editable path deps
docker-compose.yml # one service per app (just trip-cluster for M1)
.env.example # IMMICH_URL, IMMICH_API_KEY, ANTHROPIC_API_KEY, UID, GID
README.md
CLAUDE.md
shared/
immich/ # the one true Immich client
core/ # SQLite store + domain models + persistence
ai/ # Anthropic batch wrapper (scaffolded; light/unused in M1)
ui/ # base.html, Tailwind/DaisyUI/Alpine/HTMX, Jinja macros, app.js
apps/
trip-cluster/
Dockerfile
categorize.py # CLI entry point -> app.cli.main()
app/ # Flask factory, routes, app-specific logic
tests/
docs/superpowers/specs/ ...
```
- `shared/*` are local, editable Python packages so apps `import` them directly. Packaging via a workspace (`pyproject.toml` path deps; `uv` workspace optional). This stays close to the existing pip+venv practice while enabling the monorepo.
- Each app owns its Dockerfile and a distinct port. **trip-cluster serves on 8084** (8082 = travel-memories, 8083 = image-rater).
- The SQLite DB and downloaded thumbnails live on a mounted volume so state persists across container restarts and is reachable from the host. `user: ${UID}:${GID}`.
## Shared packages
### `shared/immich`
The single Immich REST client (consolidates the duplicated `immich.py` from both existing apps). `requests.Session` with `x-api-key`. Capabilities needed for M1:
- `resolve_tag_id(name)`, list tags / tag inventory.
- Search assets (paginated `POST /api/search/metadata`) by date range, by tag, and broadly — returning at least: `id`, `originalFileName`, `localDateTime`, GPS lat/lon (when present), city/country (when present), `type`, existing tags, rating.
- `download_thumbnail(asset_id)` (`…/thumbnail?size=preview`).
- `upsert_tag(name)`, `tag_assets(tag_id, asset_ids)` (write-back of trip / `_pipeline/*` tags), plus a helper for the shared `_pipeline/` meta-tag namespace (see Tag conventions).
### `shared/core`
SQLite store + domain dataclasses + persistence (connection, schema/migrations, data-access helpers). Owns the on-disk schema. Designed as a clean data-access layer so other apps can adopt it later.
**Domain model (central unit = Cluster / candidate trip):**
- **Asset** — mirror of an Immich asset: `immich_id`, `taken_at`, `gps_lat`/`gps_lon` (nullable), `place_city`/`place_country` (nullable), `type`, `has_gps`, `thumb_path`, `ingested_at`.
- **Tag** — inventory of Immich tags: `name`, `immich_tag_id`, usage `count` (ingested now; used heavily in M2).
- **AssetTag** — existing Immich tags per asset (denormalized for filtering).
- **Cluster** — a candidate trip: `id`, `start_at`, `end_at`, `count`, `suggested_name`, `confidence`, `kind_guess` (`trip`|`everyday`), `status` (`pending`|`approved`|`non_trip`|`merged`|`split`|`skipped`), `decided_name`, `reviewed_at`, `notes`.
- **ClusterMember** — asset↔cluster link with `member_confidence` and `is_outlier`.
- **WritebackLog** — every change pushed to Immich (`asset_id`, `action`, `tag`, `result`, `applied_at`) for idempotency.
- **Meta** — last-ingest timestamp, run params, schema version.
**SQLite tables:** `assets`, `asset_tags`, `tags`, `clusters`, `cluster_members`, `writeback_log`, `meta`. Single DB file on the mounted volume.
### `shared/ai`
Anthropic **batch** Messages API wrapper using the official `anthropic` SDK, carrying over image-rater's proven conventions (per-criterion judgments, `confidence`, score floors, chunking ≤50/batch, default model `claude-haiku-4-5`). **Scaffolded but barely used in M1** — trip-cluster is algorithm-first (near-zero AI cost). It earns its keep in M3 (enrich).
### `shared/ui`
The shared visual foundation, extracted from the proven travel-memories/image-rater templates:
- `base.html` (DaisyUI 4 + Tailwind + Alpine.js + HTMX via CDN, navbar, content block).
- The canonical **grid + fullscreen lightbox** component: thumbnail grid with ring-selection, **arrow-key navigation**, **full-screen view** (arrows = prev/next, Esc = close). *(First-class requirement.)*
- Reusable Jinja macros (image grid, lightbox, confidence/status badges, approve/reject controls) and shared `app.js` Alpine components.
## The `trip-cluster` app (M1 deliverable)
### Workflow
```
categorize ingest [--from DATE --to DATE | --tag NAME | --subset N] # 1. pull → SQLite + thumbs
categorize cluster [--gap-threshold ...] # 2. build candidate trips
categorize serve # 3. review UI on :8084
# (approve / non-trip / split / merge)
categorize apply # 4. write approved tags back to Immich
```
(`serve` may also trigger `apply` per-cluster via a UI button, plus a batch "apply all approved" action; CLI `apply` is the headless equivalent.)
### 1. Ingest (automatic, scopeable)
Pull assets from Immich via `shared/immich`, upsert metadata into SQLite, download thumbnails. **Incremental & idempotent** (fetch only changed-since-last via Immich `updatedAt`; upsert by `immich_id`).
**Read-back of pipeline state:** ingest also reads each asset's existing tags, including `_pipeline/*`. Assets carrying `_pipeline/processed` are marked already-adjudicated in SQLite, so the working DB can be **rebuilt from Immich** after a loss and review resumes without redoing finished work.
**Scopeable from day one** — the mechanism for the vet-first plan. Instead of forcing a full 40k pull, `ingest` accepts a bounded scope:
- `--from / --to` (date window) → e.g. the last trip plus surrounding everyday photos;
- `--tag NAME` → a single already-tagged trip;
- `--subset N` → a cap.
This lets the whole tool be validated on the last trip for near-zero cost before widening to the backlog over time.
### 2. Cluster (automatic) — signal hierarchy
Pure, algorithmic, **no API calls** (keeps the POC nearly free). Produces candidate trips, each with a `suggested_name`, `confidence`, and `kind_guess`:
1. **Existing trip tag** → authoritative seed; its assets form a confirmed cluster (still shown, to verify *completeness*). Respects the user's existing trip-tag convention.
2. **Timestamp gap clustering** → primary structure: sort by `taken_at`, split where the inter-photo gap exceeds a tunable threshold.
3. **Location anchors** (existing location tags like "Kiev" + GPS when present) → refine boundaries, propose names.
4. **Coverage detection** → assets *inside* a confirmed trip's time window but *missing* its tag are flagged "likely belongs here" (the completeness gap); assets carrying a trip tag but *outside* their cluster are flagged as outliers.
5. **Visual similarity****deferred from M1; leading post-M1 enhancement.** Trips are defined by time + place, not visual likeness, so the signals above resolve the large majority; visual similarity only helps a narrow case (an ambiguous time gap where two bursts may be one trip). pHash is the wrong tool (it finds near-duplicates, not trip-level similarity). CLIP is the right tool, but Immich's REST API does not cleanly expose raw embedding vectors. **Viable path: read-only access to Immich's Postgres pgvector embeddings** (do nearest-neighbor/clustering ourselves), pending a feasibility check against the live Immich version. Especially valuable here because the old library is GPS-poor, so visual continuity may be one of the few secondary signals for those photos.
**Confidence & kind_guess** (echoing image-rater's confidence/floor approach): tight time window + existing trip tag + consistent location → high confidence; sparse, untagged, no GPS → low ("needs your eye"). Low-volume scattered clusters → `kind_guess = everyday` (suggested non-trip).
### 3. Review (human, cluster-level)
A review screen lists **candidate trips sorted by "needs attention"** (low confidence first). Per cluster: thumbnail grid (with the shared grid+lightbox / arrow-key / full-screen component), editable suggested name, and actions:
- **Approve trip** (confirm + tweak name)
- **Mark non-trip**
- **Split** (break one cluster into two)
- **Merge adjacent** (combine with a neighbor)
- **Skip**
To keep it fast: high-confidence clusters arrive **pre-filled**, and an **"approve all high-confidence"** bulk action lets the user rubber-stamp the obvious cases, concentrating attention on the fuzzy ones. Every decision **persists immediately** to SQLite (resumable: closing and reopening resumes exactly where the user left off).
### 4. Write-back (to Immich, idempotent)
On approval (per-cluster or batch `apply`):
- **Approve** → write the trip tag (via `upsert_tag` + `tag_assets`) to all member assets, respecting the existing **content** trip-tag convention (trip names are user-facing, not namespaced).
- **Mark non-trip** → apply `_pipeline/non-trip` so those assets are filtered out and never resurface as unreviewed.
- **Mark processed** → every adjudicated asset (trip-assigned, non-trip, **or** reviewed-and-skipped) gets `_pipeline/processed`. This is the durable "done" flag in the source of truth: it captures the reviewed-but-untagged case, and lets a fresh `ingest` re-derive what's already handled even if the SQLite working DB is lost (see Ingest read-back).
- **Idempotent**: `writeback_log` records applied changes; re-runs skip what's already done. **Explicit confirmation before any write** (mirrors image-rater's export safety).
All durable state lands in **Immich**; SQLite remains the working/review layer that can be rebuilt from Immich tags.
### Tag conventions (shared across all apps)
Two clearly separated kinds of tags:
- **Content / organizational tags** — trip names ("Italy 2019"), locations ("Kiev"), people. User-facing, part of the existing convention, **never namespaced**. trip-cluster must **read and respect** the existing trip-tag convention rather than invent a parallel one.
- **Pipeline meta-tags** — everything the tooling generates as machinery, nested under a single parent **`_pipeline/`** so the whole set can be removed by deleting the parent and never clutters the tag list. Defined as a **shared convention in `shared/immich`** and used by every app in the monorepo:
- `_pipeline/processed` — categorizer: asset adjudicated (any outcome)
- `_pipeline/non-trip` — categorizer: everyday/noise
- `_pipeline/ai-rating/<0-5>` — image-rater's `ai-rating/<n>` moves under this root when migrated (M4)
- room for future meta-tags (review-state, etc.)
**Reconciliation caveat:** before writing, **inspect the live Immich instance** and reconcile with anything already present (e.g. image-rater's existing un-namespaced `ai-rating/<n>`) rather than blindly creating duplicates.
## App shape & conventions
- `categorize` CLI (argparse subcommands: `ingest` · `cluster` · `serve` · `apply`) + Flask factory `create_app()`.
- Jinja + DaisyUI/Tailwind/Alpine/HTMX via `shared/ui`, **no build step**, Immich thumbnail proxy route, serves on **8084**.
- Env vars (same names as existing apps): `IMMICH_URL`, `IMMICH_API_KEY`, `ANTHROPIC_API_KEY` (Anthropic optional for M1), loaded from `.env`.
## Testing strategy
TDD throughout, mirroring the existing apps:
- **Unit:** clustering algorithm as a pure function (timestamps/tags/GPS → clusters) — deterministic and easy to assert; SQLite store; config; Immich client (mocked via `pytest-httpserver`).
- **Route/UI:** the review flow (approve / non-trip / split / merge, immediate persistence, idempotent write-back) with Immich mocked.
- **Playwright UI tests:** grid + arrow-key navigation + full-screen lightbox behavior, mirroring travel-memories/image-rater.
- Shared packages tested independently of the apps.
## Cost notes
M1 is algorithm-first, so trip-cluster makes **essentially no Anthropic calls** — clustering is local timestamp/tag/GPS math. AI cost becomes relevant only in M3 (enrich), where image-input tokens dominate and the batch API (~50% cheaper, chunked ≤50) plus the Haiku default keep it modest (image-rater reference: ~700 images on Haiku ≈ $1).
## Out of scope for M1
- Choosing keepers within a trip (image-rater).
- Full non-trip/everyday categorization (M6).
- Migrating the existing apps onto the shared foundation (M4/M5).
- Geocoding / GPS backfill / AI captioning (M3).
- Narrative text drafting (human-owned; downstream).