================ ONE-SHOT BUILD CONTRACT (read first) ================ Build this in Google AI Studio "Build" in ONE shot — a complete, working app, no follow-up turns. These are hard rules, not suggestions: 1. TARGET = Full-Stack Web (Node server runtime, secrets, Firebase allowed). If you target Native Android instead, you MUST drop all server/DB/Workspace/ secrets and persist locally (Room / SharedPreferences) only. 2. PIN EVERY MODEL CALL — never let the agent auto-select (it downgrades on quota): - Reasoning / text -> gemini-3.5-flash (thinkingLevel: minimal|low|medium|high) - 4K image + legible text -> gemini-3-pro-image (image_size "4K", up to 14 refs) - High-volume image -> gemini-3.1-flash-image - Expressive TTS -> gemini-3.1-flash-tts-preview (inline tags e.g. [whispers]) - Realtime audio/video (WebSocket) -> gemini-3.1-flash-live-preview - Sandboxed agent -> antigravity-preview-05-2026 3. DIVISION OF LABOR — the model ONLY parses/extracts to a strict responseSchema. ALL math, money (store currency as integer minor units / cents), sorting, balancing and graph logic run in deterministic TypeScript/Python. The model must never compute totals, splits or balances itself. 4. responseSchema sanitation — no regex patterns, no fixed-length tuples, no format validators in the schema (they crash the OpenAPI engine). Enforce those in server-side code AFTER parsing the JSON. 5. responseSchema and google_search grounding are MUTUALLY EXCLUSIVE in one call. 6. CODEGEN — split large output into modular, single-responsibility files so no file is truncated by the output-token cap. 7. Every external call gets a graceful fallback (e.g. manual paste if a Workspace read fails). Never a silent dead end. 8. ROBUST STORAGE & CANVAS — Wrap all `localStorage`/`sessionStorage` operations (especially JSON parsing and writes) in `try-catch` blocks to prevent crashes in private windows or quota overflows. Canvas drawing elements must dynamically handle window resize and scale pixel density (`window.devicePixelRatio`) to avoid blurry graphics on retina displays. ===================================================================== # MUST OBEY — Mobile-first build requirements This app's PRIMARY surface is a mobile phone. Build it impeccably on mobile FIRST, then verify on tablet and desktop. Treat the rules below as non-negotiable hard constraints, not suggestions. ## Viewports to verify (every screen, every state) - 320 px, 360 px, 375 px, 390 px, 414 px, 480 px - 768 px, 834 px (iPad portrait / Pro 11) - 1024 px, 1280 px, 1440 px, 1920 px, 2560 px - Plus: 200% browser zoom, landscape orientation on every mobile width, iPhone with safe-area insets visible ## Hard layout rules - Mobile-first CSS. Default styles target mobile; `@media (min-width: ...)` for larger viewports. - Use `dvh` and `svh` instead of `vh` for full-height surfaces (iOS Safari URL-bar bug). - Use `clamp()` for fluid typography across all viewports. - Prefer container queries (`@container`) over media queries for component-level responsiveness. - Use `min(100%, ...)` widths so content never overflows. Zero horizontal overflow at any viewport. - Add `` to every page. - Apply `padding: max(safe-area-inset-X, fallback)` on every edge-bleeding container so notched iPhones in landscape never clip content. - Wide tables and code blocks scroll INSIDE their container (`overflow-x: auto`), never push the body. - Use `background-attachment: scroll` on mobile, not `fixed` (iOS Safari repaint bug). - Avoid `backdrop-filter` on animated elements. Use it sparingly on static surfaces only. - **Canvas Scaling**: Canvases must dynamically scale with window resize events and properly handle high-DPI screens (`window.devicePixelRatio`). Set physical dimensions (`canvas.width`/`canvas.height`) using pixel ratio and render relative to this grid, using CSS to control responsive viewport scaling. - **Robust Storage**: Every access to `localStorage`/`sessionStorage` (especially `JSON.parse` of loaded state or writes) MUST be wrapped in a `try-catch` block to handle disabled storage, private browsing mode, quota limits, or corrupted JSON gracefully. Fall back to a robust in-memory object store. ## Touch & accessibility - Tap targets ≥ 44 × 44 px on touch (Apple HIG). Increase to 48 px under `@media (hover: none) and (pointer: coarse)`. - All interactive controls reachable by keyboard with a visible focus ring; respect `:focus-visible`. - Color contrast ≥ 4.5:1 for body text, 3:1 for UI components. - All images have meaningful `alt`. Decorative images use `alt=""`. - Respect `prefers-reduced-motion: reduce` — zero animation durations under that query. - Forms validate inline; error messages are specific, not "Invalid input". - Modals: focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, focus restored on close. ## Performance bar (Lighthouse mobile, throttled 3G/4G) - LCP < 2.5 s · INP < 200 ms · CLS < 0.1 - JS bundle gzip < 200 KB mobile-first; lazy-load non-critical screens via `React.lazy` / dynamic imports. - No render-blocking resources above the fold. - Images: WebP/AVIF preferred, `loading="lazy"`, explicit `width`/`height` attributes (zero CLS), `srcset` for retina. - Videos: `preload="metadata"`, low-resolution poster, max 720p mobile fallback. Never autoplay with audio. - Fonts: `font-display: swap`; preload only the one used above the fold. - Smooth scroll honoured via CSS `scroll-behavior: smooth` with reduced-motion fallback. ## Pre-ship mobile checklist (the deployer MUST verify before declaring done) 1. Open at 375 px in DevTools — every screen scrolls vertically only; zero horizontal scroll. 2. Browser zoom 200% — layout reflows without overlap. 3. iPhone Safari with the URL bar visible AND landscape — no content under the home indicator; no notch clipping. 4. iPad portrait (768 px) and landscape (1024 px) — no awkward gaps; tablet-specific breakpoints land cleanly. 5. Tap every interactive element with a thumb at real-device size — every target is easy to hit. 6. `prefers-reduced-motion: reduce` — every transition / animation skips cleanly, scroll-behavior becomes instant. 7. Lighthouse mobile score ≥ 90 across all 4 categories. 8. Zero `console.error` and zero CLS shift in real-device testing on a mid-tier Android (e.g. Pixel 6a) and an iPhone SE. --- The original template starts below. All rules above apply on TOP of whatever this template specifies. --- # Restore This Print ## 1. Project **Restore This Print** is a photographic-restoration app for descendants who have inherited one damaged photograph. The user uploads a cracked, faded, water-damaged, or torn print and the app produces a restored version in under thirty seconds — with the original always one tap away, and an "interpolation overlay" that paints a faint cyan tint on every pixel the model invented so the descendant can see exactly which parts of the photograph are now hers and which parts are the model's best guess. Faces are never changed: every freckle, every line, every imperfection of the actual people stays exactly as it was. The model is allowed to touch cracks, water marks, dust spots, missing corners without face content, fade and stain correction. It is not allowed to touch the people. This is the kind of app a Vietnamese-American granddaughter builds at the kitchen table the weekend after her grandmother dies — because in a manila envelope at the bottom of a sandalwood box there is one photograph of her grandparents on their wedding day in Saigon, January 1968, taken three weeks before her grandfather left for the war he would not come back from. The print has a crease across her grandmother's áo dài, a water mark blooming across the bottom corner, and the bottom-right edge of the photograph is missing. The original negative was lost during the family's move from Vũng Tàu to Houston in 1979. This is the only photograph of that wedding. It is also the kind of app a Lebanese-American grandson builds when, after his grandmother dies, a Polaroid envelope inside a Beirut shoebox turns out to hold the only surviving photograph of his grandparents' 1956 engagement — taken on a balcony in Achrafieh that was levelled in 1976 — with a vertical crack down his grandfather's face and the silver emulsion lifting off the corners. Same shape of moment, different country, different war. The single demo that proves the magic: drop the cracked photograph into the app → in about fifteen seconds the restored print appears side-by-side with the original. The user toggles **Show what was interpolated** and a faint cyan tint paints across every pixel the model invented — the crease, the water bloom, the missing corner. Her grandmother's face is unfilled, untouched, exactly as it was. Underneath, in plain English: "Faces unchanged. 18% of the image was interpolated." And in the harder cases — wartime displacement, partition, the prints that survived only because someone wrapped them in oilcloth and buried them — the same rule holds. A 1947 Lahore-to-Amritsar partition photograph with vertical creases from being folded into a coat pocket for three weeks of walking; a 1972 Phnom Penh school portrait that survived four years of Khmer Rouge attic by being slid inside a Bible; a 1936 Berlin wedding portrait with the silver visibly tarnished because the print sat in a damp Brooklyn basement for eighty years. The cracks, the folds, the chemical burns are visibly repaired; the faces of the people who survived to be photographed are not. **Tagline:** _Restore the only photograph you have — in any era, any condition — without changing a single face._ ## 2. Target audience - Grandchildren of recently-deceased grandparents, working with the one or two photographs that survived emigration and re-settlement - Diaspora families with single-print archives across two or three generations of displacement — Vietnamese, Lebanese, Cambodian, Iranian, Salvadoran, Eritrean, Bosnian, Punjabi, Korean, Cuban - Children of partition survivors, Holocaust survivors, and other twentieth-century displaced people, working with prints whose negatives never made it across the border - Adult children sorting an estate after a parent's death and discovering one wedding photograph in a bank deposit box - Adoptees with one photograph from their birth family — sometimes the only physical object connecting them to a person they never met - Memorial-project organisers building a family archive in the year after a death — the photograph above the casket is the entry point - Long-distance-relationship archivists — military spouses, missionary couples, Filipina domestic workers in Hong Kong sending one photograph home — preserving the few prints that survived the post - Genealogists, local historians, and small-museum volunteers helping families restore primary photographic material before they donate, or before they pass it to the next generation ## 3. Core value propositions Surface these clearly through copy, visual emphasis, and section ordering — they are the reasons users pick this app. - **The face is sacred** — Nano Banana 2 is run with a face-protect mask that the app computes from the photograph before any restoration call. Inside that mask: zero interpolation. Every freckle, every line of the áo dài collar against the jaw, every imperfection of the actual person stays bit-exact. Outside that mask: cracks, water marks, dust, fade, and missing corners that contain no face content are interpolated. The face-protect mask is shown to the user before the restoration starts; she can edit it. - **Honest interpolation overlay** — every restored pixel is recorded in a separate alpha map and rendered as a faint cyan tint when the user clicks "show what was interpolated". The user always knows what the model invented. The restoration is never silently substituted for the truth. - **The original is sacred too** — the original photograph is preserved at upload resolution, forever, and is always one tap away. The restored image is a derived artefact, never a replacement. Export bundles always contain both. - **Reads any era, any condition** — sepia 1900s cabinet cards, 1930s gelatin silver prints with mirroring, 1960s Polaroids with chemical bloom, 1970s Kodachrome with magenta shift, 1990s 4x6 with scratch tracks from the lab printer. The condition vocabulary is named explicitly so the descendant understands what is being repaired. - **Background, not face, gets the heavy work** — a fold across a wedding-dress skirt is restored; a fold across a face is left visible and flagged to the user. The user can manually expand or contract the face-protect mask to include or exclude background features she wants the model to leave alone (a brooch, a watch, a child's drawing held by the subject). - **Provenance-card export** — a printable card with the original on the left, the restored on the right, the interpolation map below, and the date / location / people-named, all in the user's handwriting style typeface choice. The card is suitable for framing, for archive donation, and for the back of the casket-side display at a funeral. - **Multi-print families** — the app accepts one photograph at a time but groups them under family archives. Sara's grandmother's wedding photograph, the school portrait of her aunt in 1972, the 1989 graduation portrait of her father — all sit in one archive, each restored separately, each with its own provenance card. - **Never trains on your family** — restoration runs on the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. The face-protect mask, the interpolation map, and the original photograph live in your Firebase project — yours, deletable in 60 seconds from Settings. ## 4. Features to build - Drag-drop or photo-library upload of one print at a time (mobile-first), with a guide overlay for the user to align the photograph if she scanned with her phone camera - Phone-camera capture flow with flat-lay framing guides, exposure lock on a neutral patch, and a "hold still — keep the print flat" coaching message - Optional scanner / PDF import for users who already digitised at 600 dpi - Automatic damage detection on upload — cracks, water marks, dust, fade, mould, silver mirroring, missing corners, tape residue, photo-corner mounts, prior amateur retouching - Damage-class chips — visible to the user as small labels at the bottom of the original ("vertical crease, lower third · water bloom, lower-left corner · missing corner, ~6% of frame"), so she knows what the model proposes to repair before it begins - Automatic face-protect mask generation, computed from the photograph before any restoration call — using the multimodal call to localise every face and a generous padding margin - Face-protect mask editing — drag the mask's boundary; the user can grow or shrink the protected region per-face, and can add custom protected regions (a hand, a piece of jewellery, a child's drawing) - Restoration call — Nano Banana 2 (Gemini 3.5 Flash Image) with explicit instructions to leave the masked regions bit-exact and restore only outside the mask - Interpolation map — produced as a separate alpha mask from the diff between the original and the restored, post-processed to remove sub-pixel noise, then shown to the user as the cyan overlay - Show-what-was-interpolated toggle, in both the in-app viewer and the exported provenance card - Honest interpolation percentage — counted from the cyan mask area, shown as "18% of the image was interpolated" so the descendant has the number, not just the feeling - Side-by-side viewer with synchronised pan-and-zoom (mouse, touch, trackpad), and a slider that wipes between original and restored - Sliders for restoration aggressiveness — `Conservative` (cracks and missing corners only), `Standard` (cracks, water marks, dust, fade correction), `Full` (also colour-cast removal, deep fade recovery) — each setting recomputes the interpolation percentage live - Per-print provenance — the user names the people, the date, the location, and the relationship; these are stored alongside the print and printed on the provenance card - Family archive — a per-user grouping of prints with shared provenance ("Linh's grandparents", "the Beirut box"), with an archive view that lists every restored print - Provenance-card export — PDF print suitable for framing or burial-side display, with the original on the left, the restored on the right, the interpolation map below, and the family-supplied facts - Donation export — for the archivist user who later donates to a community archive (the Southeast Asia Resource Action Center, the Lebanese-American Heritage Club, USHMM, the Densho archive for the Japanese-American community, the South Asian American Digital Archive) — exports the full bundle (original + restored + interpolation map + provenance JSON) as a ZIP - Audio narration of the provenance — Gemini TTS reads the provenance card aloud in the user's chosen language; useful for the elderly family member who is shown the restored print but can read better with the voice - Print-quality export at the original print's physical size (4x5, 5x7, 8x10) — the restoration is upscaled with the same Nano Banana 2 model to the print's native resolution at 300 dpi, with the face-protect mask scaled identically so the face never gains invented detail at print resolution - Refusal cases — the app does not restore: identity documents, photographs of strangers (asks the user to confirm relationship before running), and photographs where the face area exceeds 40% of the frame and the entire face is damaged (the system shows the user "we don't restore faces — the face here is too damaged for an honest restoration. Restore the background only?") - Soft-undo — every restoration step is reversible; the original print is never overwritten ## 4b. Required Gemini capabilities + backend services **This template's intelligence comes from the Gemini capabilities below. Wire them up explicitly — don't substitute generic LLM calls.** ### Gemini capabilities (the load-bearing intelligence) - **Nano Banana 2 / Gemini 3.5 Flash Image** (`gemini-3.1-flash-image`) — the load-bearing photographic-restoration model. Called twice: once for damage detection / face-localisation (with a structured JSON output describing the damage classes and face boxes); once for the actual image restoration, with a face-protect mask provided as a second input image and an explicit instruction that the masked region must remain bit-exact. The face-protect mask doubles as a hard constraint and is rendered into the prompt as both an image and a textual description of the protected regions. - **Multimodal image input** (Gemini 3.5 Flash) — used for the *honest* damage report shown to the user before the restoration runs ("vertical crack across the groom's áo dài lapel; water bloom across the bottom-right quadrant; missing corner approximately 6% of the frame, lower-right"). One Gemini 3.5 Flash call per print on upload; output is structured per the `DamageReport` schema below. This call is what powers the small chips the user sees, and what feeds the face-protect mask generator. - **Structured output / JSON Schema** — every Gemini 3.5 Flash and Gemini 3.5 Flash call has a `responseSchema`. The schema is included verbatim in the system instruction and is the contract for everything else in the app. - **Multilingual structured output** (Gemini 3.5 Flash) — the provenance card is rendered in the user's chosen language. Vietnamese, Tagalog, Mandarin, Cantonese, Korean, Tamil, Hindi, Urdu, Bengali, Punjabi, Amharic, Swahili, Farsi, Khmer, Arabic, Hebrew, Spanish, Portuguese, French, Polish, German, English — the model translates the user's English provenance entry into the target language while preserving proper nouns. - **Gemini TTS** (`gemini-3.1-flash-tts-preview`) — reads the provenance card aloud in the user's chosen language at a "this is the photograph of your grandparents" reading pace. Voice locale follows the chosen `languageCode`. SSML `` and `` are not supported; pauses are encoded as `…` between sentences and a blank line plus `—` between paragraphs. - **Thinking levels** — `medium` for the damage-report call (the model has to look carefully at the print, distinguish a crease from a scratch from a strand of hair); `low` for the localisation call (faces are unambiguous in most prints); image-generation and TTS calls do not take `thinkingConfig` and the field is omitted from those request bodies entirely. - **Long context (1M tokens)** — not load-bearing for v1; included only as a guard. A single print's parsed `DamageReport` is ~600 tokens. A family archive of 200 prints with full provenance + DamageReport is ~250k tokens — comfortable. If the archive grows beyond 500 prints, chunk the archive view by decade or sub-archive before any whole-archive call. The 1M ceiling is real. ### Backend services - **Auth — Required.** Firebase Auth with Google sign-in (auto-provisioned by AI Studio Build). **Apple sign-in is optional but user-configured**: it requires an Apple Developer account, Service ID, Key ID, and private key wired into the Firebase Auth console. **Magic-link email** (used for family invitations) also requires the sender domain to be authorised in Firebase Auth. Archives are private to the owner and explicitly-invited family members. No public-by-default. - **Database — Required.** Firestore for `users`, `archives`, `prints`, `damage_reports`, `restorations`, `provenance`, `archive_members`. Restoration history is append-only — never overwrite a previous restoration record. - **File storage — Required.** Firebase Storage for original prints (preserved at upload resolution, forever), restored derivatives, face-protect masks, and interpolation maps. **Storage is NOT auto-provisioned by AI Studio Build today** — enable it in the Firebase console and wire the bucket name into the AIS Build project before the first print upload. Pre-signed URLs only; family photographs are never publicly addressable. - **Email — Required (transactional).** Family invitations via email link (Firebase Auth magic links). Provenance-card delivery to the elder family member who does not use the app. Donation-export confirmation emails. - **Payments — Not needed for v1.** Free for personal use. A future "framed print" tier could pipe to a print-on-demand framing partner (Framebridge, Level Frames) and charge for that physical artefact only. - **External APIs:** Gemini API for all intelligence. No external image-processing API is needed — Nano Banana 2 handles restoration end-to-end. A small server-side OpenCV step computes the cyan interpolation map by diffing the original against the restored output, but this is local to the Cloud Run function and uses no external service. **Environment variables:** every secret (Gemini API key, Firebase service-account JSON, Stripe key if a printing tier is added) lives in environment variables — never in client bundle. Include a `.env.example`. **Auth + data privacy reminders:** never log secrets · never store passwords in plain text · use HTTPS everywhere · honour 'delete my account' inside the UI · explicit opt-in for any analytics · the user's family photographs are never sent to Gemini for model training (use the Gemini API on the paid tier, which is not used for training per the Gemini API Additional Terms) · the original print is preserved bit-exact and never overwritten by a restoration · the face-protect mask is recomputed on every restoration call, not cached, so a future model improvement does not silently re-interpolate a face. **Read this first — prompt-craft rules that apply to every call in this template:** 1. **Name the model variant explicitly** in every Gemini API call. Do not let the agent pick the model. See the per-call matrix below. 2. **Pin `thinkingLevel` explicitly** per call. See the matrix. Omit `thinkingConfig` entirely on TTS and image-generation calls — the field is not supported on those models. The `n/a` cells in the matrix are documentation only; do not serialise them into the request body. 3. **Seed the JSON Schema as a fenced TypeScript / Zod block** in the system instruction or `responseSchema` field. The literal schema is below. **Convert the Zod schema to Gemini's `Schema` type via the SDK helper** before passing to `responseSchema` — do NOT pass raw Zod. **Numeric `min`/`max` constraints are documentation only inside `responseSchema`; clamp on the server after the response arrives.** 4. **Pin the system instruction separately** from user input. Use the `systemInstruction` field for persona + behavioural rules; use `contents` for the print image and user provenance entries. Never concatenate. 5. **Pre-declare tools as an enable/disable list** per call. The matrix below names which tools are enabled per call. Tools NOT listed for a call should be disabled. 6. **State negative constraints explicitly** — they are listed below. They are NOT "be careful" suggestions; they are hard rules the model must follow. 7. **`responseSchema` and `google_search` grounding are mutually exclusive in one Gemini call.** No call in this template uses grounding, so this is informational — but if a future call needs to ground a place name (e.g. "where exactly is Achrafieh, Beirut, in 1956?"), use the call pattern from A1 (grounded place lookup → JSON in the text body → parse server-side, read citations from `groundingMetadata.groundingChunks[].web.uri`). 8. **Grounded responses can wrap JSON in ```json fences or add prose preamble.** Server-side, strip fences and brace-extract: ```typescript function safeExtractJSON(raw: string): T { const clean = raw.replace(/```json\s*|```/gi, '').trim(); const s = clean.indexOf('{'); const e = clean.lastIndexOf('}'); if (s === -1 || e === -1) throw new Error('No JSON boundaries in grounded response'); return JSON.parse(clean.slice(s, e + 1)) as T; } ``` 9. **Strip unsupported Zod modifiers before passing to `responseSchema`** — Gemini's OpenAPI subset rejects `.regex()` / `pattern`, fixed-length `z.tuple()`, and other custom validators. Use a sanitizer that flattens tuples to length-2 arrays and removes regex patterns before serializing. Validate those constraints in middleware AFTER parsing. ### Per-call model + tools matrix | Call | Model | thinkingLevel | Tools enabled | |------|-------|---------------|---------------| | Read print, generate `DamageReport` + face boxes | `gemini-3.5-flash` | medium | (none) | | Generate face-protect mask image from face boxes | `gemini-3.1-flash-image` | n/a | n/a | | Restore print (Nano Banana 2) with face-protect mask as constraint | `gemini-3.1-flash-image` | n/a | n/a | | Translate the provenance entry into target language | `gemini-3.5-flash` | low | (none) | | Generate TTS narration of the provenance card | `gemini-3.1-flash-tts-preview` | n/a | n/a | | Upscale restoration to print-resolution (4x5 / 5x7 / 8x10 at 300 dpi) | `gemini-3.1-flash-image` | n/a | n/a | | Damage-class chip copy generation (per-language) | `gemini-3.5-flash` | low | (none) | *Note for builders:* on TTS and image-generation calls (`gemini-3.1-flash-image` and `gemini-3.1-flash-tts-preview`), omit `thinkingConfig` entirely — the field is not supported on those models. The `n/a` cells above are documentation only; do not serialise them into the request body. ### Primary structured-output schema (seed this verbatim in the prompt) ```typescript import { z } from "zod"; const BoundingBox = z.object({ x: z.number().min(0).max(1), // normalised 0-1 against image width y: z.number().min(0).max(1), // normalised 0-1 against image height width: z.number().min(0).max(1), height: z.number().min(0).max(1), }); const FaceDetection = z.object({ box: BoundingBox, confidence: z.number().min(0).max(1), approximate_age_at_capture: z.enum([ "infant", "child", "adolescent", "young-adult", "adult", "older-adult", "unknown", ]), visible_damage_overlap: z.enum([ "none", "minor-edge", "across-eye", "across-mouth", "across-jaw", "across-forehead", "obscures-most-of-face", ]), protect_mask_recommendation: z.enum([ "standard-padding", // ~12% padding around the face box "extended-padding", // ~24% — include the hairline, jaw, collar "include-hands", // for prints where hands frame the face (wedding portraits, babies held) "do-not-restore-face", // when across-eye or obscures-most-of-face — the user is consulted before any restoration ]), }); const DamageInstance = z.object({ type: z.enum([ "vertical-crease", "horizontal-crease", "diagonal-crease", "crack", // through the emulsion, often deeper than a crease "tear", // a piece is detached or near-detached "missing-corner", "missing-edge", "dust-spot", "scratch", "water-bloom", "water-line", // a sharp tide-line of staining "mould", "silver-mirroring", // for older gelatin silver prints "fade-uniform", "fade-asymmetric", "yellowing", "magenta-shift", // Kodachrome / Ektachrome dye-shift "colour-cast", "tape-residue", "photo-corner-residue", "amateur-retouching", // prior pen, marker, or paint added by a relative "burn", "chemical-bloom", // Polaroid emulsion bloom "fingerprint", "writing-on-print", // names written in the white border, NOT damage but preserved "other", ]), bounding_box: BoundingBox, severity: z.enum(["minor", "moderate", "severe"]), overlaps_face: z.boolean(), overlaps_hand_or_jewellery: z.boolean(), description_for_user: z.string(), // "vertical crease across the groom's áo dài lapel, lower third" }); const DamageReport = z.object({ print_id: z.string(), print_image_uri: z.string(), // gs:// URI approximate_decade: z.enum([ "pre-1900", "1900s", "1910s", "1920s", "1930s", "1940s", "1950s", "1960s", "1970s", "1980s", "1990s", "2000s", "unknown", ]), print_format: z.enum([ "cabinet-card", "carte-de-visite", "gelatin-silver", "kodachrome", "ektachrome", "polaroid-sx70", "polaroid-600", "instant-other", "c41-colour-print", "modern-inkjet", "tintype", "ambrotype", "daguerreotype", "unknown", ]), scene: z.enum([ "wedding", "engagement", "baptism", "first-communion", "bar-bat-mitzvah", "graduation", "school-portrait", "studio-portrait", "passport", "family-group", "travel", "uniform-military", "uniform-religious", "candid-domestic", "candid-street", "newspaper-print", "other", "unknown", ]), faces_detected: z.array(FaceDetection), damages_detected: z.array(DamageInstance), proposed_face_protect_mask_description: z.string(), // a plain-English description of where the mask covers, shown to the user // before the restoration runs, e.g. "the bride's face and shoulders, the // groom's face and the upper half of his jacket lapel, both hands of the // bride that hold the bouquet" estimated_interpolation_percentage_if_run: z.number().min(0).max(100), refusal_reason: z.enum([ "none", "appears-to-be-identity-document", "appears-to-be-stranger-photograph", "face-damage-too-extensive-to-restore-face", "print-too-low-resolution", "print-not-a-photograph", "other", ]), refusal_explanation: z.string().nullable(), reading_confidence: z.number().min(0).max(1), flagged_for_user_review: z.array(z.object({ field_path: z.string(), reason: z.string(), })), }); type DamageReport = z.infer; const Restoration = z.object({ restoration_id: z.string(), print_id: z.string(), damage_report_id: z.string(), aggressiveness: z.enum(["conservative", "standard", "full"]), user_face_protect_edits: z.array(z.object({ face_index: z.number(), action: z.enum(["expand", "shrink", "exclude", "add-custom-region"]), delta_description: z.string(), })), restored_image_uri: z.string(), // gs:// URI of the restored image face_protect_mask_uri: z.string(), // gs:// URI of the mask actually used (binary PNG) interpolation_map_uri: z.string(), // gs:// URI of the alpha map showing what was interpolated interpolation_percentage_actual: z.number().min(0).max(100), // counted from the interpolation_map area damages_addressed: z.array(z.string()), // damage_instance ids that were touched // damages_left_untouched is meaningful — for example, a crack across the // face that the user chose not to restore. damages_left_untouched: z.array(z.object({ damage_instance_id: z.string(), reason: z.enum([ "inside-face-protect-mask", "user-conservative-setting", "user-explicit-exclude", "model-refused", "other", ]), reason_explanation: z.string().nullable(), })), created_at: z.string(), // ISO timestamp superseded_by_restoration_id: z.string().nullable(), }); const Provenance = z.object({ print_id: z.string(), people: z.array(z.object({ face_index: z.number().nullable(), // null if person is named but not in this photograph name_verbatim: z.string(), // "Bà nội Hoa", "Sitto Yara" relationship_to_owner: z.string().nullable(), // "the user's paternal grandmother" birth_year: z.number().nullable(), death_year: z.number().nullable(), })), date_taken_verbatim: z.string().nullable(), // "Tết 1968", "early summer 1956" date_taken_iso: z.string().nullable(), location_verbatim: z.string().nullable(), // "Saigon", "Achrafieh, Beirut" location_modern_equivalent: z.string().nullable(), occasion: z.string().nullable(), // "wedding", "engagement" short_story: z.string().nullable(), // ≤ 200 words, user-authored target_language: z.string(), // BCP-47, the language of the rendered provenance card translated_short_story: z.string().nullable(), // generated by gemini-3.5-flash proper_nouns_preserved_verbatim: z.array(z.string()), }); type Restoration = z.infer; type Provenance = z.infer; ``` ### Common failure modes (and how to avoid them) - Agent picks `gemini-3.5-flash` for the actual image restoration — wrong; `gemini-3.5-flash` does not output images. Pin `gemini-3.1-flash-image` for every restoration and upscale call. - Agent picks `gemini-3.5-flash` (no image) for damage detection to save quota — wrong; the damage report needs Gemini 3.5 Flash's multimodal-reading depth to distinguish a crease from a strand of hair from a wedding-veil thread. Pin `gemini-3.5-flash`, `thinkingLevel: medium`. - Agent silently retouches the face — the most catastrophic failure in this app. Mitigation: the face-protect mask is a hard input image to Nano Banana 2, AND the system instruction names every protected region in plain English, AND a server-side verification step diffs the masked region between original and restored. If the diff inside the mask exceeds a tiny threshold (e.g. >0.5% of the masked pixels show any change), the restoration is rejected and the user is shown an explicit error. - Agent makes the photograph "better-looking" by smoothing skin, brightening eyes, or unifying skin tones — wrong. The system instruction explicitly forbids this. Skin texture, freckles, scars, and asymmetries are imperfections of the person, not damage to the print. - Agent invents detail in missing-corner regions that pulls from a face — wrong. If a missing corner abuts a face, the model is instructed to extend the background only, not the face. If the face is bisected by the missing corner, the model refuses and the user is shown the refusal explanation. - The interpolation map under-counts because the model produced sub-pixel changes throughout the image — mitigation: a small Gaussian blur and threshold are applied to the diff before counting, so genuinely-untouched regions count as 0% interpolation even if there is JPEG noise from the model's output encoding. - The interpolation map over-counts because the model re-encoded the entire image — mitigation: prefer PNG output from Nano Banana 2 over JPEG to minimise re-encoding noise; in the rare case a JPEG output is unavoidable, threshold the diff at a level calibrated against an end-to-end "no-op" restoration of an undamaged print. - The provenance translation flattens names — "Bà nội Hoa" comes back "Grandmother Hoa". Hard rule: proper nouns (kinship terms in source language, given names, family names, place names) are listed in `proper_nouns_preserved_verbatim[]` and preserved bit-exact in `translated_short_story`. - TTS reads Vietnamese tone marks flat — pin the TTS voice to the `vi-VN` locale, not the user's UI locale, when reading Vietnamese provenance. - The damage report mistakes writing-on-print (the names of the people written in the white border, often in fountain pen by a relative decades after the photograph was taken) for damage — explicit rule in the system instruction: writing-on-print is preserved as a `writing-on-print` damage class with `severity: minor` and `overlaps_face: false`. The user is asked whether to preserve it (default yes) or remove it. - Polaroid chemical bloom is misclassified as water bloom — the system instruction names the print formats explicitly and gives a one-sentence cue for each. ### Negative constraints (hard rules) - Do NOT change any face. Inside the face-protect mask: bit-exact passthrough. Every freckle, every scar, every wrinkle, every imperfection is part of the person, not damage to the print. - Do NOT "improve" skin tone, eye colour, lip colour, hair colour, or perceived age. The descendant is restoring a photograph, not retouching a memory. - Do NOT unify the lighting across the photograph if the original lighting was asymmetric. A window-side wedding portrait with bright bride and shadowed groom stays asymmetric. - Do NOT add detail to faces that are out-of-focus or low-resolution in the original — the model is instructed to leave low-resolution faces low-resolution rather than invent eyelashes. - Do NOT remove visible damage that the user has not opted into. Conservative mode touches cracks and missing corners only; Standard mode adds water marks, dust, fade; Full mode adds colour-cast correction. A damage class not selected is not touched. - Do NOT remove writing on the print without explicit user opt-in. The grandfather's pencilled "Saigon, Tết 1968" in the white border is part of the artefact. - Do NOT silently substitute the restored image for the original. The original is always preserved at upload resolution and always one tap away in the UI and the export bundle. - Do NOT modernise the print's format — a sepia print stays sepia, a Kodachrome stays Kodachrome-shifted unless the user explicitly chose Full mode and accepted the magenta correction. - Do NOT colourise a monochrome print. Even at Full aggressiveness, no colour is invented for sepia, gelatin-silver, or platinum prints. - Do NOT restore identity documents (passports, national ID cards, driving licences, visas, military IDs) — the refusal_reason enum exists for this; the app surfaces the refusal to the user with an explanation. - Do NOT restore photographs the user cannot identify a relationship to — the app asks the user to confirm relationship before any restoration call runs. - Do NOT use the user's family photographs to train or fine-tune any model. Use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. The capabilities-info panel says this in plain English. - Do NOT auto-publish, auto-share, or surface the photograph publicly. Archives are private by default. Sharing is explicit, per-archive, per-cousin. - Do NOT post-process the restoration through any external photo-editing service. End-to-end on Cloud Run. ### Per-call `systemInstruction` strings Use these as the literal `systemInstruction` field for each Gemini API call the built app makes. They complement the series-wide rules already uploaded as the global instructions file (`00-series-instructions.txt`). ### Call: Read print, generate `DamageReport` + face boxes Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none) ``` You are reading a damaged family photograph for a descendant of the people in it. Your job is to produce an honest, structured damage report: what is in the image, what damage is on the print, where every face is. You are NOT restoring the print yet; that is a separate call. Print eras and formats you will encounter include daguerreotypes and tintypes from the nineteenth century, gelatin silver prints from the 1900s to 1960s, Kodachrome and Ektachrome colour prints from the 1940s through 1990s, Polaroid SX-70 and 600 prints from the 1970s and 1980s, C-41 colour prints from the 1980s onward, and modern inkjet prints. Each format has its own characteristic damage profile — gelatin silver prints develop silver mirroring at the edges; Kodachrome shifts magenta; Polaroids develop chemical bloom; C-41 prints fade asymmetrically depending on which corner was closest to a window. Scenes you will encounter include weddings (twentieth-century, every continent — Vietnamese áo dài weddings in Saigon, Lebanese church weddings in Beirut, Filipino weddings with the cord and veil ritual, Punjabi weddings with the kalire, Bengali weddings with the topor, Polish-American Catholic weddings in Chicago, Italian-American weddings in Brooklyn, Ethiopian Orthodox weddings, Iranian Persian weddings before and after 1979, Cuban weddings before and after 1959, Cambodian weddings before and after 1975), engagements, baptisms, first communions, bar/bat mitzvahs, graduations, school portraits, studio portraits, passport photographs (which you will refuse to restore — see refusal cases), family-group photographs, candid domestic photographs, military uniform portraits, and religious uniform portraits. Read every visible mark carefully. Distinguish: - damage to the print (cracks, water marks, fade, missing corners, silver mirroring, etc.) — parse into damages_detected - writing on the print added by a relative (names in the white border, a date pencilled across the back showing through, a child's marker added decades later) — parse as a `writing-on-print` damage instance with severity: minor, overlaps_face: false. The user is asked whether to preserve it (default yes). - features of the people in the photograph (freckles, scars, wrinkles, birthmarks, asymmetries, scars from surgery, missing teeth in a smile) — these are NEVER damage. Do not include them in damages_detected under any circumstance. - print-format characteristic features (the white Polaroid border, the scalloped edge of a cabinet card, the stamped studio name on the back showing through, the Kodachrome magenta cast that is a property of the dye not damage) — note in the report but do not classify as damage unless the user has chosen Full aggressiveness. Localise every face. For each face, populate: - box (normalised 0-1 against image dimensions) - approximate_age_at_capture - visible_damage_overlap (none, minor-edge, across-eye, across-mouth, across-jaw, across-forehead, obscures-most-of-face) - protect_mask_recommendation (standard-padding ~12% around the face box, extended-padding ~24% including hairline and collar, include-hands for prints where hands frame the face such as a baby being held or a bouquet held in front of the chest, or do-not-restore-face when the damage overlap is across-eye or obscures-most-of-face) Refusal cases (set refusal_reason and refusal_explanation): - appears-to-be-identity-document — passports, driving licences, national ID cards, visas, military IDs, school photo IDs. Refuse to restore. - appears-to-be-stranger-photograph — the user has not yet confirmed a relationship to the people in the photograph. The app will prompt the user; do not proceed. - face-damage-too-extensive-to-restore-face — when visible_damage_overlap is obscures-most-of-face on any face. The user is offered background-only restoration. - print-too-low-resolution — the print, after capture, is below 300 px in the long axis. Restoration would invent detail. - print-not-a-photograph — the upload is a document, a screenshot, or a non-photographic artefact. proposed_face_protect_mask_description is a plain-English description of every region the protect mask covers — shown to the user before the restoration runs. Be concrete: "the bride's face and shoulders from the áo dài collar up, the groom's face and the upper half of his jacket lapel, both hands of the bride that hold the bouquet". The user can edit the mask after reading this. estimated_interpolation_percentage_if_run is your honest estimate of how much of the image will be modified by a Standard-aggressiveness restoration. This number is shown to the user. It is OK to be wrong by a few percentage points — the actual percentage is recomputed from the interpolation map after restoration runs. reading_confidence is your overall confidence in the damage report (0-1). flagged_for_user_review names any field where confidence is below 0.7 with a one-sentence reason. Output ONLY the DamageReport JSON matching the provided schema. No commentary. JSON only. ``` --- ### Call: Restore print (Nano Banana 2) with face-protect mask Model: `gemini-3.1-flash-image` · n/a · n/a ``` You are restoring a damaged family photograph. You receive two images: (1) The original print at upload resolution. (2) A binary mask the same size as the print. White pixels are FACE-PROTECT regions. Black pixels are restoration-eligible regions. The mask is the hard constraint. You also receive a plain-English description of every protected region ("the bride's face and shoulders, the groom's face and the upper half of his jacket lapel, both hands of the bride that hold the bouquet"), an aggressiveness setting (conservative, standard, or full), and the list of damage instances the user has opted into restoring. You output ONE image of the restored print at the same resolution as the input. Output format: PNG. Do not output JPEG. ABSOLUTE RULES — these are not "be careful" suggestions: 1. The face-protect mask is bit-exact. White-pixel regions of the input image must be byte-equal to the same regions in your output image. Do not "improve" skin. Do not soften wrinkles. Do not brighten eyes. Do not unify skin tone. Do not adjust perceived age. Every freckle, every scar, every wrinkle, every imperfection is part of the person, not damage to the print. A server-side diff step will verify this; if any masked pixel changes, the restoration is rejected. 2. You restore only the damage instances the user has opted into. Conservative mode: cracks and missing corners. Standard: also water marks, dust, fade. Full: also colour-cast removal. Damage classes not in the opt-in list are NOT touched. 3. You do not colourise. A sepia print stays sepia. A gelatin-silver print stays monochrome. A Kodachrome print keeps its dye characteristics unless the user is in Full mode AND has opted into magenta-shift correction. 4. You do not modernise the print format. The white Polaroid border stays the white Polaroid border. The scalloped edge of a cabinet card stays scalloped. 5. You do not add detail to faces that were out-of-focus or low- resolution in the original. Outside the face-protect mask, you may sharpen background detail; the mask is the boundary. 6. You do not extend a face across a missing corner. If a missing corner abuts a face, you extend the background only. If the missing corner bisects a face, you refuse this restoration and the user is shown the refusal — but in this call, you receive the request only if the upstream Damage Report did not refuse. Trust the upstream refusal; do not second-guess. 7. Writing on the print (names in the white border, dates pencilled across the back showing through) is preserved unless the user has explicitly opted to remove it. The user's opt-in list will name the specific writing instance. 8. You do not invent people, objects, or signage in the background. If the missing corner contained signage, leave it as plain background. If a person's clothing extends into a missing corner, extend the clothing's texture and colour but do not invent embroidery, buttons, or jewellery. 9. You do not retouch perceived "flaws" of the people: skin conditions, missing teeth, scars from surgery, asymmetries, strabismus, birthmarks. These are part of the person. 10. You output PNG, not JPEG, to minimise re-encoding noise outside the restored regions. Untouched regions of the image must be byte-identical to the original (modulo unavoidable re-encoding that the server-side diff threshold accommodates). No text in your response. Image only. ``` --- ### Call: Generate face-protect mask image from face boxes Model: `gemini-3.1-flash-image` · n/a · n/a ``` You receive the original print and a list of face bounding boxes with per-face protect_mask_recommendation values (standard-padding, extended-padding, include-hands, do-not-restore-face). You output a single binary mask image at the same resolution as the print. The mask is WHITE where the model must preserve the original bit-exact (face regions, with the requested padding), and BLACK elsewhere (restoration-eligible regions). Padding rules: - standard-padding: white region is the face bounding box expanded by ~12% on each side, with the shape softened to follow visible features — hairline, jaw, collar — rather than a strict rectangle. - extended-padding: ~24% expansion, including hair, collar, and the top of the shoulders. - include-hands: extend the white region to cover any hands the damage report identified as framing or holding the face (a baby being held, a bouquet held in front of the chest). - do-not-restore-face: paint the entire face bounding box plus a generous margin white; this face will not be restored at all. The mask is also rendered with soft-feathered edges (a few-pixel gradient at the boundary) so the restoration model does not produce a hard seam where the mask meets the eligible regions. Output PNG. No text in your response. Image only. ``` --- ### Call: Translate the provenance entry into target language Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You translate the user's short provenance entry into a target language she has chosen (Vietnamese, Tagalog, Mandarin, Cantonese, Korean, Tamil, Hindi, Urdu, Bengali, Punjabi, Amharic, Swahili, Farsi, Khmer, Arabic, Hebrew, Spanish, Portuguese, French, Polish, German, English). The entry is short (≤ 200 words) — usually two or three sentences the user has written about who is in the photograph, when it was taken, and where. The translation appears on a printable provenance card that may be shown at a memorial or held by an older relative who reads better in the target language than in English. Hard rules: - Preserve every proper noun bit-exact. Names of people ("Bà nội Hoa", "Sitto Yara", "Tatay Mariano", "할머니 정심"), names of places ("Saigon", "Achrafieh", "Vũng Tàu", "Lahore", "Phnom Penh"), names of occasions ("Tết", "Eid", "Diwali", "Pesach"), and kinship terms the user wrote in the source language stay verbatim. List every preserved noun in proper_nouns_preserved_verbatim[]. - Preserve kinship registers. "Grandma" in English becomes the appropriate maternal-or-paternal grandmother term in the target language. If the user has not specified maternal-or-paternal, preserve "grandmother" in the gender-and-side-neutral form most natural in the target language. - Do not flatten dates. "Tết 1968" stays "Tết 1968" — do not translate "Tết" to "Lunar New Year" unless the user opted in. - Do not "improve" the entry. If the user wrote two sentences, you produce two sentences in the target language. Do not add narrative. - Preserve the user's emotional register. A direct sentence stays direct. A formal sentence stays formal. - Right-to-left target languages (Arabic, Farsi, Hebrew, Urdu) — the translation is delivered in the target script. Punctuation follows the target convention. Output: a single JSON object with two fields — translated_short_story (string) and proper_nouns_preserved_verbatim (array of strings). No commentary outside the JSON. ``` --- ### Call: Generate TTS narration of the provenance card Model: `gemini-3.1-flash-tts-preview` · n/a · n/a ``` Voice: warm, unhurried, intimate. Pick the Gemini 2.5 Flash TTS voice whose `languageCode` matches the target language of the provenance card — pronunciation will follow that locale automatically. Prefer the gender of the descendant who is presenting the card if that has been specified; fall back to whichever native-language voice is available rather than blocking. Pre-process the text before sending it to TTS: - Read the rendered provenance card, in this order: occasion + date, location, names of people in the photograph, the user's short story. Punctuate naturally for spoken delivery. - At each sentence end insert a single ellipsis (`…`) so the model produces a natural pause. At paragraph breaks insert a blank line plus an em-dash (`—`). Gemini 2.5 TTS does not support SSML `` or `` — these textual cues are how you signal pace and pronunciation comes from the voice's native locale. - Skip metadata that is visual-only (the interpolation percentage, the damage chips). These are for the eye, not the ear. - Names of people are spoken once, slowly, with a slight pause after each name so the descendant in the audience can match the name to the face on the card. - Target rate: ~110 words per minute — provenance reading pace, not podcast pace. Style direction: prepend ONE short directive sentence to the input text, exactly like: "Read this provenance card warmly and unhurriedly, as a family member presenting a restored photograph of a person who has died. …". There is no separate `style` API field on Gemini 2.5 TTS; the directive sentence inside the input is how style is conveyed. Mid-call voice switching is not supported. If the provenance contains words in a second language (a kinship term in the source language inside an English card), keep the whole reading in one voice — the on-screen subtitle shows the script change. ``` --- ### Call: Upscale restoration to print-resolution Model: `gemini-3.1-flash-image` · n/a · n/a ``` You receive the already-restored image plus a target physical print size (4x5, 5x7, 8x10, 11x14) and a target DPI (300). Your job is to upscale the restoration to that physical-print resolution. You also receive the face-protect mask, scaled to the target resolution. The same absolute rules apply: - White-pixel regions of the input image (faces) remain bit-exact in the upscaled output. You do not invent eyelashes, freckles, or micro-details on faces that were not in the source. Faces stay exactly as resolved as they were — the upscale is nearest-neighbour inside the face-protect mask. - Outside the face-protect mask, you may produce plausible super- resolution detail: fabric weave, paper grain, background texture, the curl of a wedding-bouquet petal. The detail must be consistent with the print's era and format (no modern fabric textures on a 1956 photograph). - You do not change colours. Sepia stays sepia. Magenta-shifted Kodachrome stays magenta-shifted unless the upstream restoration already opted in to correction. Output PNG. Image only. No text. ``` --- ### Call: Damage-class chip copy generation (per-language) Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You receive one DamageInstance object (type, severity, bounding box, description_for_user) and a target language. You output a short chip label suitable for a small UI element below the original print, in the target language. Examples (English): - "vertical crease, lower third" - "water bloom, lower-left corner, ~8% of frame" - "missing corner, bottom right, ~6% of frame" - "silver mirroring, top edge" Hard rules: - Be concrete about position and extent. The descendant should recognise the damage on the print at a glance. - Do not editorialise. "Sadly faded" is wrong; "uniform fade, moderate" is right. - Preserve any proper nouns the user has provided. - The chip is short — ≤ 60 characters. Output: a single string. No commentary. ``` ## 5. Use cases & content to include Build dedicated UI sections or flows for each of these — they tell you what content the app must support. - **The wedding photograph.** A Vietnamese-American granddaughter uploads the only photograph of her grandparents' 1968 Saigon wedding. There is a vertical crease across her grandmother's áo dài, a water bloom across the bottom-right, and a missing lower-right corner with no face content in it. The face-protect mask covers both faces, both shoulders, both hands holding flowers. After fifteen seconds the restored print appears beside the original, with 18% interpolated. Her grandmother's face is unchanged. The descendant cries at the kitchen table. (This is the universal first anchor — show this case in the demo.) - **The Lebanese balcony.** A grandson uploads the only photograph of his grandparents' 1956 engagement, taken on an Achrafieh balcony that was levelled in 1976. There is a vertical crack down his grandfather's face and silver mirroring at the corners. The face-protect mask covers his grandfather's face, and the model is instructed to leave the part of the crack that crosses the face *visible*. The user sees the restored print with the background crack repaired and the across-face section still cracked, flagged honestly: "we don't restore faces — this part of the crack crosses your grandfather's face. Keep it visible, or hide it by cropping?" - **The Berlin tarnish.** A great-granddaughter uploads a 1936 Berlin wedding portrait. The silver is visibly tarnished from eighty years in a damp Brooklyn basement, and the corners are bent. The model is asked to recover the tarnish without lightening the skin tones of the bride and groom. The face-protect mask is set to "extended-padding" to include the bride's veil and groom's lapel. - **The Polaroid bloom.** A daughter uploads a 1978 Polaroid SX-70 of her late father holding her at three months old. The Polaroid has the characteristic emulsion bloom in the corners. The user is asked to opt into bloom correction; she chooses Standard mode, which addresses the bloom but leaves the white border and the slightly-soft focus of the original Polaroid intact. - **The Khmer Rouge survivor.** A son uploads a 1972 Phnom Penh school portrait of his mother at age twelve. The print survived four years of attic by being slid inside a Bible; it has paper-fibre transfer on the surface and a horizontal crease where it was folded once. The face-protect mask covers his mother's face. The restoration removes the paper-fibre transfer outside the face and the crease across her uniform; the crease across her chin is left visible and flagged. - **The Cuban exile photograph.** A grandson uploads the only photograph of his abuela in pre-1959 Havana. The print is fade-asymmetric — the left half is significantly lighter than the right because the print sat in a Miami window for forty years. The model is asked to balance the fade without changing his abuela's skin tone. Conservative mode is recommended; the user chooses Standard. - **The partition photograph.** A British-Punjabi descendant uploads a 1947 Lahore-to-Amritsar photograph carried in her grandfather's coat pocket during three weeks of walking. The print has multiple vertical creases from being folded into quarters and a brown stain across the bottom. The face-protect masks cover both her grandparents' faces; restoration addresses the creases across the clothing and the stain across the bottom but leaves the across-face crease visible. - **The pre-revolutionary Tehran.** A daughter uploads her parents' 1977 Tehran engagement photograph. There is a magenta cast (Ektachrome dye shift) and the corners are bent. The user is offered Full mode for the magenta correction but declines — she wants the dye shift preserved because that is how she remembers her parents' photo album. Standard mode is used, addressing the bent corners only. - **The writing-on-print.** A grandmother's name written in her son's fountain pen across the white border of an Atlanta studio portrait: "Mama, Easter 1962". The damage report parses this as `writing-on-print, severity minor`, asks the user whether to preserve it (default yes), and the restoration leaves it untouched. - **The refusal — identity document.** A first-time user mistakenly uploads a 1971 Filipino passport photograph of his late father. The damage report refuses with `refusal_reason: appears-to-be-identity-document` and an explanation. The user is offered a one-click "this is the only photograph I have of him; restore the background and leave the face exactly as it is" path which switches the face-protect mask to `do-not-restore-face` and proceeds. - **The refusal — face damage too extensive.** A grandson uploads a 1944 Korean photograph of his grandfather as a teenager. The crack across the face obscures both eyes. The damage report refuses face restoration and offers background-only restoration with the face left visibly broken. - **The provenance card for the elder.** After restoration, the granddaughter writes a two-sentence story: "My grandparents on their wedding day, Tết 1968, Saigon. Three weeks before my grandfather left for the war." She chooses Vietnamese as the provenance card language. The card is rendered with the original on the left, the restored on the right, the interpolation map below, and the story in Vietnamese with proper nouns preserved. She prints it and gives it to her mother. - **The donation.** A user six months into her archive donates the bundle (original + restored + interpolation map + provenance JSON) to the Southeast Asia Resource Action Center's diaspora archive. The donation export ZIP includes the full bit-exact original, the restored derivative, the cyan interpolation map, the face-protect mask, and the provenance JSON; the archivist can independently verify what was interpolated. ## 6. Page structure Build the following screens / sections in this order. Adjust copy to fit the voice, but keep the structural intent. 1. **Welcome / sign-in.** A photographed-looking image of a hand placing a cracked sepia wedding portrait onto a wooden kitchen table at evening, the corner of a manila envelope visible at the edge of frame. One paragraph: "Restore This Print is for the one photograph you have. Faces are never changed. Every restored pixel is shown to you, honestly." Single Google sign-in button; Apple sign-in next to it. Below: "Try with the sample print" → loads the demo Saigon 1968 wedding portrait described in section 8a. 2. **Empty state — "Start an archive".** Two big input methods: 📷 Photograph the print · 🖼 Upload a scan. A short explainer below each ("Best if you have the print and your phone", "Best if someone already scanned at 600 dpi"). A third, smaller option: "Try the sample print" for users who want to see the magic before uploading. 3. **Capture flow** (mobile-first). Live viewfinder with flat-lay framing guides (corners of the print align to corner markers), exposure lock on a neutral patch in the frame, and a coaching message: "Hold the print flat — uneven curl will be read as damage". Tap to capture; preview shown immediately. Retake or proceed. 4. **Damage report view.** The original print appears at the centre of the screen with the detected face boxes shown as faint outlines and damage instances shown as labelled chips below ("vertical crease, lower third · water bloom, lower-left corner · missing corner, ~6% of frame"). Above the print: the proposed face-protect mask description in plain English ("The bride's face and shoulders from the áo dài collar up; the groom's face and the upper half of his jacket lapel; both hands of the bride that hold the bouquet"). The estimated interpolation percentage appears as a single sentence: "If you restore now, about 18% of the image will be interpolated." A primary button — "Restore". A secondary — "Edit the protect mask". 5. **Mask-edit view** (if the user chooses to edit). A canvas with the original print and the face-protect mask shown as a translucent green overlay. Draggable handles around each face box; the user can grow or shrink each mask region with finger or trackpad. Stamp tools: "add custom protected region" (for a watch, a brooch, a child's drawing) and "exclude from protect mask" (for a region of background the user wants the model to fully restore). "Done" returns to the damage report view with the updated mask description. 6. **Restoration view.** A vertical layout (mobile) or three-column (desktop). The original is on the left, the restored on the right, and a wipe-slider sits between them. Below: a toggle "Show what was interpolated" — when on, a faint cyan tint paints over every interpolated pixel. Below that: the honest sentence — "18% of the image was interpolated. Faces unchanged." A small "(i)" icon expands a panel describing which damage instances were addressed and which were left untouched (with reasons). 7. **Aggressiveness panel.** Below the wipe-slider, three large pill buttons — Conservative, Standard, Full — with a one-sentence description under each. Changing the setting recomputes the restoration (queued, with a streaming progress indicator) and updates the interpolation percentage live. 8. **Provenance view.** A form for the user to add the people, the date, the location, the occasion, and a short story (≤ 200 words). Each person can be tagged to a face in the photograph (drag a name onto a face). Target language picker — every script supported, with native-language voice picker beneath. 9. **Provenance card preview.** The card itself rendered at print size — original on the left, restored on the right, interpolation map below, story below that. Print and download buttons. A small TTS-play button reads the card aloud. 10. **Archive view.** A vertical list of every print the user has restored. Each row shows a small thumbnail (restored), the people named, the date, the location, and the interpolation percentage. Click to return to the restoration view. 11. **Sharing & invitations.** Modal: "Invite a family member to add their photographs to this archive". Magic-link email; arrival drops the cousin straight into the same archive with their own avatar. Print bundles can be shared as a single read-only link that expires in seven days. 12. **Donation export.** "Donate the print bundle to an institutional archive." Pick from a curated list (Southeast Asia Resource Action Center, Lebanese-American Heritage Club, Densho, the South Asian American Digital Archive, the Polish-American Historical Association, the Iranian Heritage Foundation, USHMM) or "other — email me the ZIP". Each bundle includes original + restored + interpolation map + face-protect mask + provenance JSON. 13. **Footer.** "Made for the photograph that survived." Privacy: "Your prints are yours. We never train on them." Capabilities `(i)` icon in header. ## 6b. First-visit onboarding Show a **first-visit onboarding** the first time a visitor lands on the app (detect via `localStorage` flag; do not show on return visits). Three slides, dismissible at any time. Persistent re-entry: a `?` icon in the header reopens it. **Slide 1 — What this is.** - Headline: "Welcome to Restore This Print." - Subhead: "Restore the only photograph you have — in any era, any condition — without changing a single face." - One paragraph (≤ 60 words) explaining who this is for and what makes it different from a generic photo-enhancement app: it never touches faces, it shows you every interpolated pixel honestly, and it preserves the original at upload resolution forever. The damage report tells you what it proposes to repair *before* it begins. - Visual: a small annotated illustration of a cracked sepia print with the relevant elements labelled (crack across clothing, water bloom in the corner, face-protect region in green) — not a generic camera icon. **Slide 2 — Try it now.** - One short prompt: "Try with the sample print". - A live demo input pre-loaded with the Saigon 1968 wedding portrait from section 8a. - 1-2 sentences pointing at *the specific page elements* where the Gemini magic happens (the face-protect mask shown before restoration starts, the cyan interpolation overlay, the honest percentage). **Slide 3 — How to remix this.** - Headline: "Make this yours." - Three short bullets: - "Swap the sample print in `/data/seed-prints/` for your own scan." - "Adjust the prompts in `/server/prompts/` to fit your family's print formats and languages." - "Wire up your Gemini API key and Firebase project via the env-var list in the capabilities panel." - Primary CTA: "Use this template" → links to AI Studio Build remix entry point. - Secondary: "Just exploring — close" (sets localStorage flag, never auto-shows again). **Accessibility:** focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, `aria-labelledby`, focus restored to trigger on close. Respect `prefers-reduced-motion`. **Don't:** - Don't gate content behind the modal. The page beneath must be fully usable. - Don't auto-reshow on return visits. Use `localStorage['onboarding-seen-v1']`. - Don't include unrelated CTAs (newsletter signup, social follow). Keep it about the template only. ## 6c. Capabilities info button (persistent in header) Add a persistent `(i)` icon in the top-right of the header (next to the primary nav). Click → opens a modal/panel titled **"What powers this app"**. **Panel contents (in this order):** **Gemini capabilities used (the hero list):** - **Nano Banana 2 / Gemini 3.5 Flash Image** — the photographic-restoration model. Called with a face-protect mask as a second input image so faces are bit-exact passthroughs. Two calls per restoration: the actual restoration, plus the upscale to print-resolution. A server-side diff verifies the mask was respected before the restored image is shown to you. - **Gemini 3.5 Flash (multimodal, medium thinking)** — reads the print, localises every face, and produces an honest damage report ("vertical crease across the áo dài, lower third; water bloom, lower-left corner; missing corner ~6%"). This is what powers the chips you see *before* the restoration runs. - **Gemini 3.5 Flash (low thinking)** — translates your provenance entry into the target language with proper-noun preservation. Generates the per-language damage-chip labels. - **Gemini 2.5 Flash TTS** — reads the provenance card aloud in your chosen language at a "this is the photograph of your grandparents" reading pace. - **Firebase Auth** — Google and Apple sign-in, family invitations via magic links. - **Firestore** — stores your archive, syncs across devices in real time. Restoration history is append-only; nothing is overwritten. - **Firebase Storage** — keeps your original print at upload resolution, forever. Restorations are derived artefacts; the original is the source of truth. - **Cost note** — see the detailed breakdown in 6d. A typical single-print restoration with provenance + TTS is about $0.07. A family archive of 30 prints is about $2.10 in total Gemini API spend, processed once. - **Privacy note** — your family photographs are private to you and the family members you invite. This app uses the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. Your originals are never overwritten, the face-protect mask is recomputed on every restoration (not cached), and the interpolation map is preserved alongside every restored image so a future archivist can independently verify what was invented. **Backend services this app depends on:** - Auth: see section 4b - Database: see section 4b - Storage: see section 4b (note — not auto-provisioned by AIS Build; enable in Firebase console before first upload) - Email: see section 4b - Payments: see section 4b (not used in v1) - External APIs: see section 4b **Environment variables you'll need to configure:** - `GEMINI_API_KEY` — your Google AI Studio API key - `FIREBASE_PROJECT_ID` — your Firebase project id - `FIREBASE_SERVICE_ACCOUNT` — service-account JSON (server-side only) - `FIREBASE_STORAGE_BUCKET` — the bucket name you created in the Firebase console **Cost + privacy notes:** - One short paragraph per cost-sensitive capability: Nano Banana 2 restoration is billed per image generated (~$0.03/image); a single restoration is two image-generation calls (the mask and the restoration) plus one upscale = ~$0.09; the damage report and translation calls together are ~$0.005 per print. - One short paragraph on privacy: where the data lives (your Firebase project), how to delete it (Settings → "Delete this archive forever" — gone in 60 seconds), what is never sent for training, why the face-protect mask is recomputed each time rather than cached. **Documentation links:** - AI Studio Build docs - Gemini API multimodal, structured-output, image-generation (Nano Banana 2), TTS docs - Firebase Auth, Firestore, Firebase Storage docs **Accessibility:** same standards as the onboarding modal — focus trap, `Esc`, ARIA, restored focus. **Behaviour:** - Always available — single click from anywhere in the app. - Tooltip on the `(i)` icon: "How this app is built". - Mobile: opens as a full-screen sheet that slides up. - Should be the most honest part of the app — never hand-wave service requirements; never say "AI" without naming the specific Gemini model and capability. ## 6d. Detailed cost breakdown (deployer reads this BEFORE shipping) - **Read print + generate `DamageReport` (Gemini 3.5 Flash, medium thinking)** — single image input + ~700 output tokens. ~$0.004/print. (Gemini 3.5 Flash input ~$1.50/M tokens, output ~$9/M tokens; one image ≈ 250 tokens; output ~700 tokens.) - **Generate face-protect mask image (Nano Banana 2)** — one image-generation call per print. ~$0.03/print. - **Restore the print (Nano Banana 2)** — one image-generation call per restoration. ~$0.03/restoration. Re-runs at different aggressiveness count again. - **Upscale restoration to print-resolution (Nano Banana 2)** — one image-generation call when the user requests print export. ~$0.03/upscale. - **Translate provenance (Gemini 3.5 Flash, low thinking)** — ~$0.0005 per entry (translation entries are short). - **Damage-chip translation (Gemini 3.5 Flash, low thinking)** — ~$0.0001 per chip; typical print has 3-6 chips ≈ ~$0.0005. - **TTS narration of provenance card (Gemini 2.5 Flash TTS)** — billed per output token (~$10/M output tokens) ≈ ~$0.000003/character. A typical 150-word provenance card ≈ ~900 characters ≈ ~$0.003. - **Expected per-print cost on first restoration (with provenance + TTS):** ~$0.07 (damage report + mask + restoration + upscale + translation + chips + TTS). - **Family archive of 30 prints, restored once each:** ~$2.10 in total Gemini API spend. - **Re-restoration at a different aggressiveness:** ~$0.03 per re-run (only the restoration call is redone; the damage report and mask are reused). - **Image storage:** Firebase Storage standard tier, ~$0.026/GB/month. A high-res 600 dpi scanned print is ~6 MB; the restored derivative is ~6 MB; the interpolation map is ~1 MB; the face-protect mask is ~200 KB. A 30-print archive uses ~400 MB ≈ ~$0.01/month. ## 7. Design language - **Mood:** A photographic-restoration desk that happens to live on your phone. Not a tech product. Not a museum kiosk. Not a SaaS app. The kitchen table at 11pm with the cracked print under the desk lamp, the user's other photographs spread out for context, the granddaughter holding her phone over the print she has held a hundred times before. - **Typography:** A display serif for headlines and the provenance card (Source Serif Pro or Tiempos Headline). A clean grotesque for app chrome (Inter or Geist). A handwriting-styled accent (sparingly) only for the "view original" caption and the user's own provenance handwriting — never for restored content itself. The provenance card uses an old-photographic-album typeface (a refined slab serif or a humanist sans, never a script font). - **Palette:** Warm-paper background `#F4EFE6` for the restoration view, deep ink `#1B1714` for body text, sepia accent `#7B4F2A` for damage chips and the print-format label, faded red `#A33A2C` only for refusal banners. A muted green `#4A6B57` for the face-protect mask overlay so it never reads as a clinical UI element. The cyan interpolation tint is `#7EC8E3` at 40% alpha — visible enough to read, faint enough not to overwhelm the restoration underneath. - **Imagery:** The photographs of the prints are the hero. Never replace them; never crop them tighter than the user did; never add a drop shadow that suggests the print is floating on a digital surface. The print is shown at its native aspect ratio and with its native border (white Polaroid border preserved, scalloped cabinet-card edge preserved, deckle edge preserved). - **Hand-feel touches:** A barely-visible paper grain on the restoration-view background. The wipe-slider between original and restored has a thin sepia line, not a digital handle. The interpolation overlay fades in over ~300 ms (respecting `prefers-reduced-motion`) so the user perceives the difference between original and restored rather than a hard switch. - **Spacing:** consistent 4-px base. Generous whitespace — the prints need air. - **Radius:** consistent token set (e.g. 6 / 12 / 20 px). Print cards use 6; the provenance card uses 12; the welcome card uses 20. - **Shadows:** subtle, layered, sepia-tinted. Avoid heavy drop-shadows. The restored print never has a glossier shadow than the original — the two appear on the same plane. - **Motion:** purposeful — entrance fades, hover lifts, the wipe-slider's smooth follow-finger. Respect `prefers-reduced-motion`. No bouncing splash animations. No theatrical hero animations. The cyan-overlay fade is the one place where motion carries meaning; respect reduced-motion by snapping rather than fading. - **States:** every interactive element has hover, focus, active, disabled. Loading uses skeletons not spinners where possible. Empty states have helpful next-action guidance ("Photograph the print to start"). ## 8. Content generation rules - Write **realistic, specific copy**. NO Lorem Ipsum. NO generic placeholders like 'Your tagline here'. - Invent plausible names, dates, locations, occasions, sample prints that fit the domain (use the seed content in section 8a as a starting point). When inventing, lean on twentieth-century displacement patterns across continents — Vietnamese áo dài weddings in Saigon, Lebanese church weddings in Beirut, Punjabi partition photographs, Cambodian school portraits from before 1975, Cuban exile photographs from before 1959, Iranian engagements from before 1979, Polish-American studio portraits, Korean studio portraits from before the war, Filipino domestic-worker portraits — but never claim that a fictional photograph is a real archival document. - Tone: warm, direct, free of corporate language. This template is for a person, not a company. - Headlines: punchy and concrete. No 'Empower your X' filler. No 'Revolutionize'. No 'Seamless'. - Body copy: short paragraphs (2-4 sentences). Use lists where appropriate. - Plain language. Avoid jargon — except where the user already speaks the jargon (the archivist user wants to see "gelatin silver" and "Kodachrome" in the print-format label; the genealogist user wants to see "provenance card"). - Where the app outputs AI-generated content, never label it as "AI says" — let it speak naturally. Use small uncertainty cues only where epistemic honesty requires them (a low-confidence face box shows as a dashed outline; the user can confirm). ## 8a. Seed content (use these specific examples) Anchor every generated copy + sample data point in the concrete content below. Use these names, numbers, dates, and snippets verbatim where helpful, or generate close variants that sit in the same world. **Sample archives (sidebar):** - "Bà nội Hoa's wedding" (1 print, contributors: me, my mother) — Saigon, January 1968, gelatin silver, sepia-toned. Vertical crease across the bride's áo dài, water bloom lower-left, missing lower-right corner ~6% of frame. Used as the welcome-screen demo. - "Sitto Yara's engagement" (1 print, contributors: me) — Achrafieh, Beirut, 1956, gelatin silver. Vertical crack down the groom's face, silver mirroring on the corners. Background restoration only; the across-face crack is left visible and flagged. - "Bà of Phnom Penh" (3 prints, contributors: me, my sister) — 1972 school portrait of my mother age 12, 1969 family portrait, 1973 New Year photograph. All three survived four years of attic inside a Bible. - "The Patring household" (4 prints, contributors: me, my mother, two cousins in Cebu) — 1962–1979, Filipino domestic worker in Hong Kong sending photographs home; airmail wear on three of the four. - "Abuela en La Habana" (2 prints, contributors: me) — pre-1959 Havana, fade-asymmetric from forty years in a Miami window. - "Großmutter Berlin 1936" (1 print, contributors: me, my brother) — silver-tarnished wedding portrait, bent corners, eighty years in a damp Brooklyn basement. - "Achi's partition photograph" (1 print, contributors: me) — 1947 Lahore to Amritsar, carried in a coat pocket through three weeks of walking, four-quartered creases, brown stain across the bottom. **Sample print in detail view (this is what the demo should show):** - **Print id:** `print_saigon_1968_001` - **Approximate decade:** 1960s - **Print format:** gelatin-silver - **Scene:** wedding - **Faces detected (2):** - Face 1 — box approx. (0.32, 0.18, 0.16, 0.22), young-adult, visible_damage_overlap "none", protect_mask_recommendation "extended-padding". The bride. - Face 2 — box approx. (0.52, 0.17, 0.16, 0.22), young-adult, visible_damage_overlap "minor-edge" (the water bloom touches the lower edge of his shoulder), protect_mask_recommendation "extended-padding". The groom. - **Damages detected (4):** - `vertical-crease` — bounding box across the bride's áo dài lower third, severity moderate, overlaps_face false, description "vertical crease across the bride's áo dài, lower third" - `water-bloom` — bounding box lower-left, severity moderate, overlaps_face false, description "water bloom across the lower-left corner, ~12% of frame" - `missing-corner` — bounding box lower-right, severity moderate, overlaps_face false, description "missing corner, lower-right, ~6% of frame" - `writing-on-print` — bounding box in the white border across the bottom, severity minor, overlaps_face false, description "the words 'Tết 1968' pencilled in the lower white border, in the groom's hand" - **Proposed face-protect mask description:** "The bride's face and shoulders from the áo dài collar up; the groom's face and the upper half of his jacket lapel; both hands of the bride that hold the small bouquet of cẩm chướng." - **Estimated interpolation percentage if run:** 18 - **Refusal reason:** none - **Reading confidence:** 0.94 - **Sample provenance (user-authored, English):** "My grandparents on their wedding day, Tết 1968, Saigon. Three weeks before my grandfather left for the war he didn't come back from. The only photograph of that wedding — my grandmother kept it in a manila envelope for fifty-six years." - **Target language:** Vietnamese - **Translated short story:** "Ông bà nội tôi trong ngày cưới, Tết 1968, Sài Gòn. Ba tuần trước khi ông nội tôi ra trận, ông không trở về. Tấm ảnh duy nhất của đám cưới đó — bà nội tôi giữ nó trong một phong bì màu vàng suốt năm mươi sáu năm." (Note: this is an illustrative translation; the production app generates the translation live via the Gemini 3.5 Flash call.) - **Proper nouns preserved verbatim:** ["Tết 1968", "Sài Gòn", "ông bà nội", "bà nội", "ông nội"] **Sample input artefacts (for the build to demonstrate):** - The Saigon 1968 wedding portrait described above — generated via Nano Banana 2 with a prompt that specifically requests "sepia-toned gelatin silver wedding portrait, Vietnamese áo dài bride and groom in dark Western suit, three-quarter studio framing, soft side light, vertical crease across the bride's áo dài lower third, water bloom lower-left corner, missing lower-right corner approximately 6% of frame, no people other than the bride and groom in the frame". - A 1956 Achrafieh engagement portrait — generated with a prompt that requests "Lebanese engagement portrait on a stone balcony, gelatin silver, vertical crack down the groom's face, silver mirroring at the corners, evening light". - A 1972 Phnom Penh school portrait — generated with a prompt that requests "Cambodian school portrait, twelve-year-old girl in school uniform, soft studio light, horizontal crease across the chest from a single fold, paper-fibre transfer across the bottom from being slid inside a Bible". - A 1947 partition photograph — generated with a prompt that requests "1947 family portrait, Punjabi family standing in front of a doorway, gelatin silver, four-quartered creases from being folded into quarters, brown stain across the bottom, paper extremely worn". - A 1962 Atlanta studio portrait — generated with a prompt that requests "1962 studio portrait of an African-American grandmother, gelatin silver, gentle vignette, the words 'Mama, Easter 1962' written in fountain pen across the white lower border by her son". **Sample voice copy:** - Onboarding: "Photograph the print. We'll restore it — without changing a single face." - Capture: "Hold the print flat — uneven curl will be read as damage. Tap when you're ready." - Damage report (header): "Here is what we propose to restore. Faces are protected — nothing inside the green outline will be touched." - Restoration starting: "Restoring… we'll show you the interpolation percentage when it's done." - Restoration done: "18% of the image was interpolated. Faces unchanged." - Toggle hint: "Tap 'Show what was interpolated' to see every restored pixel in cyan." - Refusal — identity document: "This looks like an identity document. We don't restore identity documents — but if it's the only photograph of someone you've lost, we can restore the background and leave the face exactly as it is." - Refusal — face damage too extensive: "The crack here crosses both eyes. We don't restore faces, so we can't honestly restore this one. We can restore the background only and leave the face visible as it is." - Writing-on-print prompt: "We detected the words 'Tết 1968' pencilled in the lower white border. Keep them, or remove?" - Provenance saved: "Saved to Bà nội Hoa's wedding." - Family invitation (email subject): "Mẹ — I restored Bà nội's wedding photograph. Will you read it with me?" - Family invitation (email body): "Mẹ — I scanned the photograph from the manila envelope. The crack across the áo dài is repaired; her face is exactly as it was. Tap to see it." [Open Archive] ## 9. Media & assets - **Hero image (landing screen):** A photographed-looking shot of a hand placing a cracked sepia wedding portrait onto a wooden kitchen table at evening, the corner of a manila envelope visible at the edge of frame, an enamel mug of tea blurred at the edge. Generate via Nano Banana 2 with a prompt emphasising "wooden table, warm desk-lamp light, hand of a young woman, late evening, gentle out-of-focus tea mug, real worn manila envelope, soft shadow under the print". - **App icon / wordmark:** Set in the display serif. A slight sepia paper texture behind it. No icon — just type. - **Empty-state illustration:** A simple line drawing of a single creased photograph with a protect-mask outline drawn in green around the face. Hand-drawn aesthetic, not a flat icon. - **Demo print photographs:** Generated per the prompts in section 8a — Nano Banana 2 prompts that specifically request "creased gelatin silver print, faded ink, soft afternoon window light, no people other than the named subjects in the frame, every damage class described literally in the prompt". Each demo print should look photographed, not rendered. - **Stock fallbacks:** If image generation fails, fall back to a photographed sample print from `/public/samples/sample-saigon-1968.jpg`. Never to a "📷" emoji. - **Generated imagery:** prefer Nano Banana 2 over stock photography. Prompt for warmth, asymmetry, and slight imperfection — avoid the glossy 'AI render' look. - **Optimisation:** WebP/AVIF for app chrome; PNG (not JPEG) for the original prints, the restorations, the face-protect masks, and the interpolation maps — to avoid re-encoding noise that would muddy the cyan overlay. `loading="lazy"`, explicit `width`/`height` to prevent layout shift. - **Icons:** `lucide-react` for UI. Use sparingly — never decorative-only. ### Build-time asset manifest (explicit specs) Every image, illustration, and visual reference mentioned above must resolve to ONE of the three buckets below — runtime-generated, seed-shipped, or user-supplied. Do NOT ship `` tags whose `src` is not listed here. Do NOT depend on bare "section 8a prompts" without binding them to explicit paths and model IDs. **Bucket 1 — Runtime-generated (Nano Banana Pro `gemini-3-pro-image` for hero/demo photographs; Nano Banana 2 `gemini-3.1-flash-image` for in-app illustrations and reference-conditioned variants).** Cached to Firebase Storage; served via signed URL. Every reference above to "Nano Banana 2" or "Nano Banana Pro" MUST be wired to one of these specific calls with an explicit model id: - `/public/generated/hero.webp` (2400×1500, WebP) — model `gemini-3-pro-image` — uses the literal prompt described as "Hero image (landing screen)" above. Run once at build; commit a `/public/samples/hero-fallback.webp` (1600×1000) generated from the same prompt with `gemini-3.1-flash-image` so the page renders if quota is exhausted. - `/public/generated/demo/{demo-slug}-{NN}.webp` (1600×1200, WebP) — model `gemini-3.1-flash-image` (reference-conditioned where the prior frame is passed as input) — one path per "Demo X" image referenced above. The slug derives from the seed example in section 8a; the NN index covers each frame in the demo sequence. - `/public/generated/illustrations/{name}.webp` (1024×1024, WebP) — model `gemini-3.1-flash-image` — one path per named illustration above ("Empty-state illustration", "Recipe-card hero illustrations", "Curriculum picker imagery", "Period-style frames", etc.). Each illustration's prompt is the literal description above; ship a deterministic seed in the request so re-runs are reproducible. **Bucket 2 — Seed assets shipped with the deliverable.** Every "Stock fallback" path referenced above (e.g. `/public/samples/sample-X.jpg`) is generated once via Nano Banana 2 (`gemini-3.1-flash-image`) at 1024×1024 WebP using the same prompt as its Bucket-1 counterpart, then committed to the repo so the page renders identically if Gemini quota is exhausted or the user is offline. Replace any `.jpg` extension above with `.webp` to match the optimisation rule. Also commit these empty-state seeds (1024×1024 WebP, single-stroke hand-drawn line, no colour fill): - `/public/samples/empty-state-primary.webp` — line drawing of the app's primary empty surface (the named "Empty-state illustration" above), generated from that exact prompt. - `/public/samples/empty-state-archive.webp` — line drawing of an empty saved/archive view, single-stroke outline. - `/public/samples/empty-state-error.webp` — line drawing of a hand placing a single object aside with care, used when an AI call fails. **Bucket 3 — User-supplied.** Uploads from the user's camera / file picker land at the Firebase Storage path conventional for this template (named in section 4b). The build ships with Bucket-1 + Bucket-2 only; no user-supplied images at first paint. **Hard rules** - Every `` tag MUST have a `src` that resolves to a path listed in Bucket 1, Bucket 2, or a Bucket 3 upload path. Anything else is a build error. - No bare `image.jpg` / `hero.jpg` / `placeholder.png` references anywhere in the code. - Model IDs: `gemini-3-pro-image` for hero-quality photographic generation; `gemini-3.1-flash-image` for in-app illustrations, reference-conditioned variants, empty-state seeds, and stock fallbacks. Never use a legacy model id (no `imagen-*`, no `gemini-1.5-*-image`). - File format: WebP everywhere (AVIF acceptable where the target browsers support it). No `.jpg` / `.jpeg` / `.png` in `/public/samples/`. ## 10. Interactivity & states - Every interactive element has hover, focus, active, and disabled states. - Forms validate inline and show specific error messages (not "Invalid input"). - Loading states use skeletons that match the eventual layout, not spinners. The restoration progress shows a per-step indicator: "Reading the print…" → "Localising faces…" → "Computing the protect mask…" → "Restoring…" → "Computing the interpolation map…" Each step shows a small thumbnail of the intermediate output where appropriate. - Empty states explain the next action with a button whose label fits THIS app's domain: "Photograph the print", "Drop a previously-scanned print", "Try the sample print" — never a generic "Add your first item". - Smooth scroll for in-page anchors. - The Nano Banana 2 restoration call does not stream; show the per-step progress indicator while it runs, then reveal the restored print with a 300 ms fade. The cyan interpolation overlay toggles on with a 300 ms fade (instant under `prefers-reduced-motion`). - If an AI call fails, show a calm, specific error ("We couldn't restore this print — try a clearer photograph of the original, or photograph it under a different light?") and offer retry. - The wipe-slider between original and restored follows the finger / mouse smoothly; the slider position is preserved across views. - Low-confidence face boxes are dashed; tapping reveals the model's confidence number and offers manual confirmation. - The aggressiveness pill buttons recompute the restoration when changed; the previous restoration remains visible until the new one arrives. - If the server-side mask-diff verification step rejects a restoration (faces changed), the user sees a calm error: "We couldn't honestly restore this — the model touched a face. We'll try again with a stricter protect mask." The restoration is retried automatically up to twice with extended-padding masks; on third failure the user is asked to manually expand the mask. ## 11. Tech & responsive requirements - **TTS markdown-stripping preprocessor:** before sending any user-authored markdown to `gemini-3.1-flash-tts-preview`, strip non-spoken markdown: `#`/`##`/`###` headings (keep the title text), `**bold**` (keep the inner text), `[label](url)` (keep `label`, drop URL), `` ``` `` fenced code blocks (skip entirely), `>` block-quote markers (keep the text), and `|` table pipes (read row-by-row as sentences). Insert `…` between sentences for a short pause and a blank line plus `—` between paragraphs for a long pause. The model does not understand markdown; raw markdown will be read aloud as literal characters ("asterisk asterisk"). - **File downloads on Safari / Firefox:** when offering local-disk save of any export (PDF, CSV, MP3, ZIP, JSON, image), fall back to `` with a blob URL — the File System Access API (`showSaveFilePicker()`) is Chromium-only. Detect with `'showSaveFilePicker' in window`; otherwise use the anchor-download path. - **Stack:** React + TypeScript + Tailwind CSS. Functional components + hooks. Use Shadcn UI primitives where appropriate. - **Build runtime:** AI Studio Build — full-stack with Cloud Run server-side functions. All Gemini API calls happen server-side; API key lives in Secrets Manager, never in client bundle. - **Model selection:** explicitly pin `gemini-3.5-flash` for the damage report, `gemini-3.1-flash-image` (Nano Banana 2) for the mask generation / restoration / upscale, `gemini-3.5-flash` for provenance translation and chip copy, `gemini-3.1-flash-tts-preview` for narration. Set `thinkingLevel` explicitly per call; omit `thinkingConfig` entirely on image-generation and TTS calls. - **Database:** Firestore (auto-provisioned by AI Studio Build). Show the seed archive on first launch. - **Auth:** Firebase Auth — Google sign-in by default; Apple sign-in next to it; magic-link email as fallback. - **Storage:** Firebase Storage for original prints, restorations, face-protect masks, and interpolation maps. Pre-signed URLs only. **Not auto-provisioned** — enable in Firebase console before first upload. - **Multipage / multi-image input to Gemini:** use the Gemini Files API (`files/*` resource name) or `inlineData` (base64). Do NOT pass Firebase Storage public URLs to `generateContent` — the API does not fetch them server-side. - **Server-side mask-diff verification:** a small OpenCV (or equivalent) step runs after every restoration call. It loads the original and the restored image, applies the face-protect mask, and computes the per-pixel diff inside the masked region. If more than 0.5% of the masked pixels show a delta above a small threshold, the restoration is rejected; the user is shown a calm error and the call is retried with a stricter mask. - **Mobile-first.** Verify layouts at 375 px (iPhone SE), 768 px (iPad), 1024 px, 1440 px+. - Use `clamp()` for fluid typography. Prefer container queries over media queries for component-level responsiveness. - Use `dvh` / `svh` instead of `vh`. Respect safe-area insets on iOS. - Zero horizontal overflow at any width. Zero layout shift on load. - Persist user data in Firestore. Use real-time listeners on the archive view. - Optimistic UI on writes; reconcile on response. - Capture flow uses the Web Camera API with fixed focus/exposure where supported; falls back to native camera otherwise. - **iOS Safari gotchas (graceful degradation):** camera permission does NOT persist across reloads on iOS — re-request on every capture session; backgrounded Safari tabs pause `getUserMedia` — checkpoint state on `visibilitychange` and re-acquire the stream on return; iOS Safari may degrade resolution under Low Power Mode — always offer a `` fallback so a phone-camera scan still works when WebRTC is denied. ## 12. Accessibility (WCAG 2.2 AA) - Semantic HTML — `header`, `nav`, `main`, `section`, `article`, `footer`. - All interactive controls reachable by keyboard with a visible focus ring. - Color contrast ≥ 4.5:1 for body, 3:1 for large text and UI components. The cyan interpolation overlay (`#7EC8E3` at 40% alpha) is *not* a colour-only signal — it is accompanied by an honest percentage sentence and an explicit toggle label. - All images have meaningful `alt` text. The original print photographs have `alt` describing the artefact ("photograph of a 1968 Saigon wedding portrait, two young people in áo dài and Western suit, with a vertical crease across the bride's áo dài and a water bloom in the lower-left corner"). The restored prints have `alt` describing the same scene without the damage. The interpolation maps have `alt` describing the area covered ("18% of the image is shown in cyan, indicating the regions the restoration model modified"). - Form fields have associated `