216 lines
18 KiB
Markdown
216 lines
18 KiB
Markdown
# immich-photo-flow — Design (M1: Foundation + Trip-Cluster)
|
|
|
|
- **Date:** 2026-06-27
|
|
- **Status:** Approved design, pending implementation plan
|
|
- **Scope of this spec:** Milestone M1 only (shared foundation + the `trip-cluster` app). Later milestones are documented as a roadmap, each to get its own spec→plan→build cycle.
|
|
|
|
## Purpose
|
|
|
|
A monorepo of small, independent tools that clean up and structure a ~40k-image, 15-year [Immich](https://immich.app) library so it can become travel "memories." This is the **upstream** stage of an existing pipeline:
|
|
|
|
```
|
|
[NEW categorizer: trips + tags] → image-rater (pick best → 4+ album) → travel-memories (album → Grav blog)
|
|
```
|
|
|
|
Every handoff is through **Immich itself** (tags, albums, ratings) — no app talks to another directly. The categorizer's job is to leave clean **trip tags** (and explicit **non-trip** marks) in Immich for `image-rater` to consume by tag.
|
|
|
|
### Division of labor (the core principle)
|
|
|
|
AI/algorithms do what they do well; the human does QA and (later, downstream) content creation.
|
|
|
|
- The machine **proposes** groupings with a confidence score.
|
|
- The human **reviews at the cluster level** (a few dozen candidate trips), not photo-by-photo, mostly in bulk, with attention concentrated where confidence is low.
|
|
- Nothing is silently auto-applied to the library. Trust matters with 15 years of memories.
|
|
|
|
### What M1 explicitly does NOT do
|
|
|
|
- It does **not** decide which photos *within* a trip are keepers — that is `image-rater`'s job downstream.
|
|
- It does **not** fully categorize everyday/non-trip photos beyond marking them as non-trip (deferred to a future milestone).
|
|
- It does **not** draft any narrative text (the user owns content creation).
|
|
|
|
## Context & constraints (from the existing ecosystem)
|
|
|
|
The two existing apps (`claude-image-rater`, `claude-travel-memories`, both under `/home/mischa/Projects/`) establish strong, consistent conventions this project follows and builds upon:
|
|
|
|
- **Stack:** Python 3.12 + Flask (factory `create_app()`) + an argparse CLI sharing one package.
|
|
- **No-build frontend:** Jinja templates extending `base.html`, **DaisyUI 4 + Tailwind + Alpine.js** (CDN), HTMX for navigation, one `static/app.js`.
|
|
- **Single-responsibility modules:** one module is the *only* thing that talks to Immich; one is the *only* thing that talks to Anthropic.
|
|
- **Testing:** pytest + `pytest-httpserver` (mock Immich) + Playwright UI tests, **TDD** ("failing test first").
|
|
- **Docker:** one container per app, distinct port, `user: ${UID}:${GID}`, state on a mounted volume.
|
|
- **Proven UI:** travel-memories' triage screen (grid + ring-selection + keyboard navigation + fullscreen lightbox) was reused by image-rater. This becomes the canonical shared component (see §UI).
|
|
|
|
Both apps currently **duplicate** near-identical code (`immich.py`, `config.py`, `slugify`, `base.html`, thumbnail proxy route, JSON-state pattern). That duplication is the evidence that extracting a shared foundation is high-value, and that migrating those apps onto it later (M4/M5) is largely mechanical.
|
|
|
|
### Decisions locked during brainstorming
|
|
|
|
- **State store: SQLite** (not the JSON-file pattern of the existing apps). A single DB makes the cross-trip filtering this project needs ("all assets in a trip window but missing the trip tag", "everything still unreviewed") far easier. `shared/core` is designed cleanly enough that the other apps *could* migrate onto it later.
|
|
- **Immich is the source of truth.** SQLite is the working/review layer; all durable results are written back to Immich as tags.
|
|
- **Monorepo with shared packages + per-app containers** (Approach A on a shared foundation). Build the categorizer first as the POC that proves the foundation, then migrate the other apps onto it in later milestones (Option 2 end-state), in milestones — not all at once.
|
|
- **GPS is a minority signal.** Most old camera photos lack GPS, so trip detection is **timestamp-first**, anchored by existing tags, refined by GPS where present.
|
|
- **Everyday/non-trip photos:** assign trips AND positively mark non-trip so they stop resurfacing as "unreviewed." Full non-trip categorization is deferred.
|
|
|
|
## Roadmap (sequenced; each milestone independently shippable)
|
|
|
|
**End state:** one monorepo — `shared/{immich, core, ai, ui}` + `apps/{trip-cluster, tag-verify, enrich, image-rater, travel-memories}`, each its own container.
|
|
|
|
- **M1 — Foundation + trip-cluster (THIS SPEC).** Stand up the monorepo, shared packages, SQLite store, Immich ingest, and the `trip-cluster` app end-to-end. Validate on the last trip's photos. This is the POC that proves the foundation.
|
|
- **M2 — tag-verify** app on the proven foundation (verify/normalize existing tags, dedupe vocabulary, find outliers).
|
|
- **M3 — enrich** app (geocode location tags like "Kiev" → coordinates and optionally backfill GPS into Immich to improve future trip detection; AI captioning; content tags; noise detection).
|
|
- **M4 — migrate image-rater** onto shared packages (+ optional SQLite).
|
|
- **M5 — migrate travel-memories** onto shared packages (+ optional SQLite).
|
|
- **M6 (future/optional) — non-trip categorization** (the deferred full-library taxonomy).
|
|
|
|
---
|
|
|
|
# M1 Design
|
|
|
|
## Monorepo layout
|
|
|
|
```
|
|
immich-photo-flow/
|
|
pyproject.toml # workspace; shared/* as editable path deps
|
|
docker-compose.yml # one service per app (just trip-cluster for M1)
|
|
.env.example # IMMICH_URL, IMMICH_API_KEY, ANTHROPIC_API_KEY, UID, GID
|
|
README.md
|
|
CLAUDE.md
|
|
shared/
|
|
immich/ # the one true Immich client
|
|
core/ # SQLite store + domain models + persistence
|
|
ai/ # Anthropic batch wrapper (scaffolded; light/unused in M1)
|
|
ui/ # base.html, Tailwind/DaisyUI/Alpine/HTMX, Jinja macros, app.js
|
|
apps/
|
|
trip-cluster/
|
|
Dockerfile
|
|
categorize.py # CLI entry point -> app.cli.main()
|
|
app/ # Flask factory, routes, app-specific logic
|
|
tests/
|
|
docs/superpowers/specs/ ...
|
|
```
|
|
|
|
- `shared/*` are local, editable Python packages so apps `import` them directly. Packaging via a workspace (`pyproject.toml` path deps; `uv` workspace optional). This stays close to the existing pip+venv practice while enabling the monorepo.
|
|
- Each app owns its Dockerfile and a distinct port. **trip-cluster serves on 8084** (8082 = travel-memories, 8083 = image-rater).
|
|
- The SQLite DB and downloaded thumbnails live on a mounted volume so state persists across container restarts and is reachable from the host. `user: ${UID}:${GID}`.
|
|
|
|
## Shared packages
|
|
|
|
### `shared/immich`
|
|
The single Immich REST client (consolidates the duplicated `immich.py` from both existing apps). `requests.Session` with `x-api-key`. Capabilities needed for M1:
|
|
- `resolve_tag_id(name)`, list tags / tag inventory.
|
|
- Search assets (paginated `POST /api/search/metadata`) by date range, by tag, and broadly — returning at least: `id`, `originalFileName`, `localDateTime`, GPS lat/lon (when present), city/country (when present), `type`, existing tags, rating.
|
|
- `download_thumbnail(asset_id)` (`…/thumbnail?size=preview`).
|
|
- `upsert_tag(name)`, `tag_assets(tag_id, asset_ids)` (write-back of trip / `_pipeline/*` tags), plus a helper for the shared `_pipeline/` meta-tag namespace (see Tag conventions).
|
|
|
|
### `shared/core`
|
|
SQLite store + domain dataclasses + persistence (connection, schema/migrations, data-access helpers). Owns the on-disk schema. Designed as a clean data-access layer so other apps can adopt it later.
|
|
|
|
**Domain model (central unit = Cluster / candidate trip):**
|
|
|
|
- **Asset** — mirror of an Immich asset: `immich_id`, `taken_at`, `gps_lat`/`gps_lon` (nullable), `place_city`/`place_country` (nullable), `type`, `has_gps`, `thumb_path`, `ingested_at`.
|
|
- **Tag** — inventory of Immich tags: `name`, `immich_tag_id`, usage `count` (ingested now; used heavily in M2).
|
|
- **AssetTag** — existing Immich tags per asset (denormalized for filtering).
|
|
- **Cluster** — a candidate trip: `id`, `start_at`, `end_at`, `count`, `suggested_name`, `confidence`, `kind_guess` (`trip`|`everyday`), `status` (`pending`|`approved`|`non_trip`|`merged`|`split`|`skipped`), `decided_name`, `reviewed_at`, `notes`.
|
|
- **ClusterMember** — asset↔cluster link with `member_confidence` and `is_outlier`.
|
|
- **WritebackLog** — every change pushed to Immich (`asset_id`, `action`, `tag`, `result`, `applied_at`) for idempotency.
|
|
- **Meta** — last-ingest timestamp, run params, schema version.
|
|
|
|
**SQLite tables:** `assets`, `asset_tags`, `tags`, `clusters`, `cluster_members`, `writeback_log`, `meta`. Single DB file on the mounted volume.
|
|
|
|
### `shared/ai`
|
|
Anthropic **batch** Messages API wrapper using the official `anthropic` SDK, carrying over image-rater's proven conventions (per-criterion judgments, `confidence`, score floors, chunking ≤50/batch, default model `claude-haiku-4-5`). **Scaffolded but barely used in M1** — trip-cluster is algorithm-first (near-zero AI cost). It earns its keep in M3 (enrich).
|
|
|
|
### `shared/ui`
|
|
The shared visual foundation, extracted from the proven travel-memories/image-rater templates:
|
|
- `base.html` (DaisyUI 4 + Tailwind + Alpine.js + HTMX via CDN, navbar, content block).
|
|
- The canonical **grid + fullscreen lightbox** component: thumbnail grid with ring-selection, **arrow-key navigation**, **full-screen view** (arrows = prev/next, Esc = close). *(First-class requirement.)*
|
|
- Reusable Jinja macros (image grid, lightbox, confidence/status badges, approve/reject controls) and shared `app.js` Alpine components.
|
|
|
|
## The `trip-cluster` app (M1 deliverable)
|
|
|
|
### Workflow
|
|
|
|
```
|
|
categorize ingest [--from DATE --to DATE | --tag NAME | --subset N] # 1. pull → SQLite + thumbs
|
|
categorize cluster [--gap-threshold ...] # 2. build candidate trips
|
|
categorize serve # 3. review UI on :8084
|
|
# (approve / non-trip / split / merge)
|
|
categorize apply # 4. write approved tags back to Immich
|
|
```
|
|
|
|
(`serve` may also trigger `apply` per-cluster via a UI button, plus a batch "apply all approved" action; CLI `apply` is the headless equivalent.)
|
|
|
|
### 1. Ingest (automatic, scopeable)
|
|
Pull assets from Immich via `shared/immich`, upsert metadata into SQLite, download thumbnails. **Incremental & idempotent** (fetch only changed-since-last via Immich `updatedAt`; upsert by `immich_id`).
|
|
|
|
**Read-back of pipeline state:** ingest also reads each asset's existing tags, including `_pipeline/*`. Assets carrying `_pipeline/processed` are marked already-adjudicated in SQLite, so the working DB can be **rebuilt from Immich** after a loss and review resumes without redoing finished work.
|
|
|
|
**Scopeable from day one** — the mechanism for the vet-first plan. Instead of forcing a full 40k pull, `ingest` accepts a bounded scope:
|
|
- `--from / --to` (date window) → e.g. the last trip plus surrounding everyday photos;
|
|
- `--tag NAME` → a single already-tagged trip;
|
|
- `--subset N` → a cap.
|
|
|
|
This lets the whole tool be validated on the last trip for near-zero cost before widening to the backlog over time.
|
|
|
|
### 2. Cluster (automatic) — signal hierarchy
|
|
Pure, algorithmic, **no API calls** (keeps the POC nearly free). Produces candidate trips, each with a `suggested_name`, `confidence`, and `kind_guess`:
|
|
|
|
1. **Existing trip tag** → authoritative seed; its assets form a confirmed cluster (still shown, to verify *completeness*). Respects the user's existing trip-tag convention.
|
|
2. **Timestamp gap clustering** → primary structure: sort by `taken_at`, split where the inter-photo gap exceeds a tunable threshold.
|
|
3. **Location anchors** (existing location tags like "Kiev" + GPS when present) → refine boundaries, propose names.
|
|
4. **Coverage detection** → assets *inside* a confirmed trip's time window but *missing* its tag are flagged "likely belongs here" (the completeness gap); assets carrying a trip tag but *outside* their cluster are flagged as outliers.
|
|
5. **Visual similarity** → **deferred from M1; leading post-M1 enhancement.** Trips are defined by time + place, not visual likeness, so the signals above resolve the large majority; visual similarity only helps a narrow case (an ambiguous time gap where two bursts may be one trip). pHash is the wrong tool (it finds near-duplicates, not trip-level similarity). CLIP is the right tool, but Immich's REST API does not cleanly expose raw embedding vectors. **Viable path: read-only access to Immich's Postgres pgvector embeddings** (do nearest-neighbor/clustering ourselves), pending a feasibility check against the live Immich version. Especially valuable here because the old library is GPS-poor, so visual continuity may be one of the few secondary signals for those photos.
|
|
|
|
**Confidence & kind_guess** (echoing image-rater's confidence/floor approach): tight time window + existing trip tag + consistent location → high confidence; sparse, untagged, no GPS → low ("needs your eye"). Low-volume scattered clusters → `kind_guess = everyday` (suggested non-trip).
|
|
|
|
### 3. Review (human, cluster-level)
|
|
A review screen lists **candidate trips sorted by "needs attention"** (low confidence first). Per cluster: thumbnail grid (with the shared grid+lightbox / arrow-key / full-screen component), editable suggested name, and actions:
|
|
- **Approve trip** (confirm + tweak name)
|
|
- **Mark non-trip**
|
|
- **Split** (break one cluster into two)
|
|
- **Merge adjacent** (combine with a neighbor)
|
|
- **Skip**
|
|
|
|
To keep it fast: high-confidence clusters arrive **pre-filled**, and an **"approve all high-confidence"** bulk action lets the user rubber-stamp the obvious cases, concentrating attention on the fuzzy ones. Every decision **persists immediately** to SQLite (resumable: closing and reopening resumes exactly where the user left off).
|
|
|
|
### 4. Write-back (to Immich, idempotent)
|
|
On approval (per-cluster or batch `apply`):
|
|
- **Approve** → write the trip tag (via `upsert_tag` + `tag_assets`) to all member assets, respecting the existing **content** trip-tag convention (trip names are user-facing, not namespaced).
|
|
- **Mark non-trip** → apply `_pipeline/non-trip` so those assets are filtered out and never resurface as unreviewed.
|
|
- **Mark processed** → every adjudicated asset (trip-assigned, non-trip, **or** reviewed-and-skipped) gets `_pipeline/processed`. This is the durable "done" flag in the source of truth: it captures the reviewed-but-untagged case, and lets a fresh `ingest` re-derive what's already handled even if the SQLite working DB is lost (see Ingest read-back).
|
|
- **Idempotent**: `writeback_log` records applied changes; re-runs skip what's already done. **Explicit confirmation before any write** (mirrors image-rater's export safety).
|
|
|
|
All durable state lands in **Immich**; SQLite remains the working/review layer that can be rebuilt from Immich tags.
|
|
|
|
### Tag conventions (shared across all apps)
|
|
Two clearly separated kinds of tags:
|
|
|
|
- **Content / organizational tags** — trip names ("Italy 2019"), locations ("Kiev"), people. User-facing, part of the existing convention, **never namespaced**. trip-cluster must **read and respect** the existing trip-tag convention rather than invent a parallel one.
|
|
- **Pipeline meta-tags** — everything the tooling generates as machinery, nested under a single parent **`_pipeline/`** so the whole set can be removed by deleting the parent and never clutters the tag list. Defined as a **shared convention in `shared/immich`** and used by every app in the monorepo:
|
|
- `_pipeline/processed` — categorizer: asset adjudicated (any outcome)
|
|
- `_pipeline/non-trip` — categorizer: everyday/noise
|
|
- `_pipeline/ai-rating/<0-5>` — image-rater's `ai-rating/<n>` moves under this root when migrated (M4)
|
|
- room for future meta-tags (review-state, etc.)
|
|
|
|
**Reconciliation caveat:** before writing, **inspect the live Immich instance** and reconcile with anything already present (e.g. image-rater's existing un-namespaced `ai-rating/<n>`) rather than blindly creating duplicates.
|
|
|
|
## App shape & conventions
|
|
- `categorize` CLI (argparse subcommands: `ingest` · `cluster` · `serve` · `apply`) + Flask factory `create_app()`.
|
|
- Jinja + DaisyUI/Tailwind/Alpine/HTMX via `shared/ui`, **no build step**, Immich thumbnail proxy route, serves on **8084**.
|
|
- Env vars (same names as existing apps): `IMMICH_URL`, `IMMICH_API_KEY`, `ANTHROPIC_API_KEY` (Anthropic optional for M1), loaded from `.env`.
|
|
|
|
## Testing strategy
|
|
TDD throughout, mirroring the existing apps:
|
|
- **Unit:** clustering algorithm as a pure function (timestamps/tags/GPS → clusters) — deterministic and easy to assert; SQLite store; config; Immich client (mocked via `pytest-httpserver`).
|
|
- **Route/UI:** the review flow (approve / non-trip / split / merge, immediate persistence, idempotent write-back) with Immich mocked.
|
|
- **Playwright UI tests:** grid + arrow-key navigation + full-screen lightbox behavior, mirroring travel-memories/image-rater.
|
|
- Shared packages tested independently of the apps.
|
|
|
|
## Cost notes
|
|
M1 is algorithm-first, so trip-cluster makes **essentially no Anthropic calls** — clustering is local timestamp/tag/GPS math. AI cost becomes relevant only in M3 (enrich), where image-input tokens dominate and the batch API (~50% cheaper, chunked ≤50) plus the Haiku default keep it modest (image-rater reference: ~700 images on Haiku ≈ $1).
|
|
|
|
## Out of scope for M1
|
|
- Choosing keepers within a trip (image-rater).
|
|
- Full non-trip/everyday categorization (M6).
|
|
- Migrating the existing apps onto the shared foundation (M4/M5).
|
|
- Geocoding / GPS backfill / AI captioning (M3).
|
|
- Narrative text drafting (human-owned; downstream).
|