# MUST OBEY — Mobile-first build requirements This app's PRIMARY surface is a mobile phone. Build it impeccably on mobile FIRST, then verify on tablet and desktop. Treat the rules below as non-negotiable hard constraints, not suggestions. ## Viewports to verify (every screen, every state) - 320 px, 360 px, 375 px, 390 px, 414 px, 480 px - 768 px, 834 px (iPad portrait / Pro 11) - 1024 px, 1280 px, 1440 px, 1920 px, 2560 px - Plus: 200% browser zoom, landscape orientation on every mobile width, iPhone with safe-area insets visible ## Hard layout rules - Mobile-first CSS. Default styles target mobile; `@media (min-width: ...)` for larger viewports. - Use `dvh` and `svh` instead of `vh` for full-height surfaces (iOS Safari URL-bar bug). - Use `clamp()` for fluid typography across all viewports. - Prefer container queries (`@container`) over media queries for component-level responsiveness. - Use `min(100%, ...)` widths so content never overflows. Zero horizontal overflow at any viewport. - Add `` to every page. - Apply `padding: max(safe-area-inset-X, fallback)` on every edge-bleeding container so notched iPhones in landscape never clip content. - Wide tables and code blocks scroll INSIDE their container (`overflow-x: auto`), never push the body. - Use `background-attachment: scroll` on mobile, not `fixed` (iOS Safari repaint bug). - Avoid `backdrop-filter` on animated elements. Use it sparingly on static surfaces only. - **Canvas Scaling**: Canvases must dynamically scale with window resize events and properly handle high-DPI screens (`window.devicePixelRatio`). Set physical dimensions (`canvas.width`/`canvas.height`) using pixel ratio and render relative to this grid, using CSS to control responsive viewport scaling. - **Robust Storage**: Every access to `localStorage`/`sessionStorage` (especially `JSON.parse` of loaded state or writes) MUST be wrapped in a `try-catch` block to handle disabled storage, private browsing mode, quota limits, or corrupted JSON gracefully. Fall back to a robust in-memory object store. ## Touch & accessibility - Tap targets ≥ 44 × 44 px on touch (Apple HIG). Increase to 48 px under `@media (hover: none) and (pointer: coarse)`. - All interactive controls reachable by keyboard with a visible focus ring; respect `:focus-visible`. - Color contrast ≥ 4.5:1 for body text, 3:1 for UI components. - All images have meaningful `alt`. Decorative images use `alt=""`. - Respect `prefers-reduced-motion: reduce` — zero animation durations under that query. - Forms validate inline; error messages are specific, not "Invalid input". - Modals: focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, focus restored on close. ## Performance bar (Lighthouse mobile, throttled 3G/4G) - LCP < 2.5 s · INP < 200 ms · CLS < 0.1 - JS bundle gzip < 200 KB mobile-first; lazy-load non-critical screens via `React.lazy` / dynamic imports. - No render-blocking resources above the fold. - Images: WebP/AVIF preferred, `loading="lazy"`, explicit `width`/`height` attributes (zero CLS), `srcset` for retina. - Videos: `preload="metadata"`, low-resolution poster, max 720p mobile fallback. Never autoplay with audio. - Fonts: `font-display: swap`; preload only the one used above the fold. - Smooth scroll honoured via CSS `scroll-behavior: smooth` with reduced-motion fallback. ## Pre-ship mobile checklist (the deployer MUST verify before declaring done) 1. Open at 375 px in DevTools — every screen scrolls vertically only; zero horizontal scroll. 2. Browser zoom 200% — layout reflows without overlap. 3. iPhone Safari with the URL bar visible AND landscape — no content under the home indicator; no notch clipping. 4. iPad portrait (768 px) and landscape (1024 px) — no awkward gaps; tablet-specific breakpoints land cleanly. 5. Tap every interactive element with a thumb at real-device size — every target is easy to hit. 6. `prefers-reduced-motion: reduce` — every transition / animation skips cleanly, scroll-behavior becomes instant. 7. Lighthouse mobile score ≥ 90 across all 4 categories. 8. Zero `console.error` and zero CLS shift in real-device testing on a mid-tier Android (e.g. Pixel 6a) and an iPhone SE. --- The original template starts below. All rules above apply on TOP of whatever this template specifies. --- # Headshot Studio ## 1. Project **Headshot Studio** is the thirty-second alternative to a $400 studio appointment and the $150 AI-headshot subscription. Upload one ordinary phone selfie — taken at your kitchen table, in the back of a car, in the bathroom mirror before a coffee chat — and walk away with six 4K studio-grade portraits across six explicit style presets (Corporate, Editorial, Academic, Artisan, Tech, Creative). Lighting, wardrobe, crop ratio, and background change between presets. Your face does not. Identity is preserved exactly — geometry, freckles, the gap between your teeth, the scar above your left eyebrow, the shape of your nose, the colour of your skin in the lighting it was photographed in. The app shows the original alongside every generation so you can verify likeness before you download. The job-to-be-done is universal and recognisable: a user needs a studio-quality professional headshot **today**. For LinkedIn. For a CV. For a speaker bio at a conference next week. For a wedding announcement going out tomorrow. For an Airbnb host page. For a press release that quoted them this morning. For a Slack avatar that does not embarrass them at a new job. They have one casual selfie taken on a phone, often hastily, often in bad light. They want studio-grade output in under a minute, at a price they can ignore. **The 30-second demo that proves the magic:** drop one casual selfie onto the canvas → eighteen seconds later, six 4K studio-quality variations come back, one per preset. Each preset has its own lighting (warm key + soft fill for Corporate; high-contrast rim lighting for Editorial; library/window light for Academic; workshop light for Artisan; flat clean studio for Tech; coloured-gel mixed light for Creative). Each preset has its own wardrobe (charcoal two-piece for Corporate; black turtleneck for Editorial; herringbone blazer for Academic; denim apron for Artisan; clean black tee for Tech; experimental fabric for Creative). Each preset has its own crop ratio (1:1 for LinkedIn, 3:4 for portrait CVs, 5:7 for press release, 16:9 for speaker decks). Every face is the user's face, unchanged. Pick one. Download a 4K PNG. Done. **Why it beats the alternatives:** - **Faster than a studio shoot.** A photographer's calendar opens in three weeks. Headshot Studio opens in three seconds. - **Cheaper than a paid AI-headshot subscription.** Services that charge $30–$150 a pop take 12–48 hours and email back a folder of near-faces, not the user's face. This app costs ~$0.18 of Gemini-API spend per session and runs in eighteen seconds. - **Same identity quality as Nano Banana Pro's reference-image feature.** Post-I/O 2026, `gemini-3-pro-image` accepts up to 14 reference images and renders 4K output with legible in-image text. That is studio-grade identity preservation at studio-grade resolution, and Headshot Studio is the cleanest expression of it. - **No watermark. No subscription. The user owns the output.** **Tagline:** _One ordinary selfie in. Six studio-grade you's out — at 4K, in 30 seconds._ ## 2. Target audience - Professionals updating LinkedIn after a job change — first impression with a recruiter, a hiring panel, a new team - Freelancers and consultants whose CV / portfolio / website needs a headshot that does not say "I shot this in the bathroom mirror" - Speakers submitting bios to conferences, podcasts, panel events — the headshot the event poster will use - Authors with a book launch, a press release, a publisher bio page - PhD candidates and early-career academics building a faculty page or a thesis dust-jacket photo - Founders going into a fundraising round — the photo the deck pitches alongside the founder bio - Real-estate agents whose for-sale listings need a face on every yard sign - Wedding announcements — the engaged couple needs studio-portrait versions of casual phone pictures, fast, often the night before the announcement goes out - Airbnb hosts whose listing photo is currently a Snapchat selfie - Press-release subjects who got quoted in the morning and need a professional portrait by the afternoon - Artists, makers, musicians, designers — anyone whose creative bio page needs a portrait that matches their work, not a corporate template ## 3. Core value propositions Surface these clearly through copy, visual emphasis, and section ordering — they are the reasons users pick this app over a studio shoot, over a paid subscription, over their own phone's portrait mode. - **Identity is preserved exactly.** Face geometry, eye colour, skin tone, freckles, dimples, scars, the gap between the teeth, the shape of the nose, the natural asymmetry every face has — all preserved. The app never smooths skin beyond what natural studio lighting would produce. The app never lightens or darkens skin tone. The app never narrows a nose, lifts an eye, slims a jaw, or thins a chin. The reference photo is shown alongside every generation so the user can verify, side by side, that the person in the generated portrait is the person who uploaded the selfie. - **Six explicit presets, not one vague "professional" filter.** Corporate is not Editorial. Editorial is not Academic. The user picks the situation (LinkedIn for a finance job → Corporate; Substack author photo → Editorial; faculty page → Academic; maker portfolio → Artisan; engineering team page → Tech; gallery bio → Creative). The preset moves wardrobe, lighting, background, and crop ratio in one step. The face stays. - **4K output with the right crop ratio for the destination.** The Corporate preset outputs a 1:1 square cropped for LinkedIn at 4K. The Editorial preset outputs a 3:4 portrait at 4K for press release. The Speaker preset outputs a 16:9 landscape at 4K for conference decks. The user does not have to know which dimension goes where — the presets handle it. - **Thirty seconds, eighteen if your wi-fi is good.** One upload, one tap, six 4K results streamed back as they finish. The first result lands in five seconds; all six are in by twenty. - **No watermark. No subscription. Yours forever.** Download a 4K PNG immediately. The user owns the output, no Google watermark, no Headshot Studio logo, no upsell to "premium" for the unbranded version. - **The original alongside the result, always.** The reference photo sits in a strip beneath every generation. The user verifies likeness with their own eyes. If a generation drifts, the user can say so with a tap ("this doesn't look like me — regenerate this preset") and the prompt is reshaped to weight identity higher. - **Diversity-aware lighting and wardrobe.** Studio portrait conventions historically over-lit fair skin and under-lit darker skin. The Corporate preset's lighting is tuned to read the user's actual skin tone correctly — soft fill on darker complexions, no blown highlights on fairer ones, accurate skin colour rendering in every case. Wardrobe is offered in cuts that work for the user's apparent build and presentation, not a single default body type. - **Honest about what it is.** The footer never claims the image was photographed. The download includes EXIF metadata flagging the image as AI-generated. The capabilities panel names the model (`gemini-3-pro-image`, Nano Banana Pro) and the SynthID watermark that ships with every Gemini image. - **No "beautify" toggle.** No skin smoothing slider. No teeth- whitener. No eye enlargement. Those are the failure modes of consumer photo apps and the reason AI headshots have a bad reputation; the app does not offer them. ## 4. Features to build - **Single-selfie upload** — drag-and-drop the source photo onto the canvas, or tap to pick from camera roll, or take a fresh selfie in-app via the device camera - **Source-photo quality check** — server-side check before any generation: face must occupy ≥25% of the frame, eyes must be open, primary face must be in focus, no extreme tilt (>30° head roll), no heavy occlusion (sunglasses, masks). Failures show a specific, kind recovery message ("we need to see your eyes — try a photo without sunglasses"). - **Six preset cards** on the canvas — Corporate, Editorial, Academic, Artisan, Tech, Creative. Each card shows a preview of the lighting/wardrobe/background combination using a generic placeholder face until the user uploads. Tap one to generate just that preset; tap "Generate all six" to run the full set in parallel. - **4K output per preset** — each generation is requested at the native 4K resolution `gemini-3-pro-image` supports. The download filename includes the preset name and timestamp (`headshot-studio-corporate-2026-06-01.png`). - **Crop ratio per preset** — Corporate is 1:1 (LinkedIn), Editorial is 3:4 (press release), Academic is 3:4 (faculty page), Artisan is 4:5 (portfolio grid), Tech is 1:1 (team page), Creative is 4:5 (gallery bio), Speaker (the bonus 7th preset, opt-in) is 16:9. - **Side-by-side compare strip** — the source photo sits in a persistent strip beneath every generation. Each result is shown alongside the source at matching height so the user can compare facial features directly. A "swap" button shows the source full- size for a moment when held. - **"Doesn't look like me" feedback** — under each generation, a small tap target says "this doesn't look like me". One tap regenerates with the identity-preservation weight pushed higher in the prompt. The system instruction's identity-anchor phrasing changes accordingly. - **Wardrobe override** — each preset has a default wardrobe; an optional dropdown opens four alternatives within the preset (Corporate offers: charcoal two-piece, navy suit + open collar, smart-casual blazer, cream knit). Wardrobe override does not change the lighting or background. - **Background override** — each preset has a default background; a small palette underneath offers three alternatives within the preset's mood. Editorial's defaults are graphite seamless paper, rim-lit charcoal, or muted indigo wall. - **Speaker-deck variant** — opt-in 7th preset that outputs the same six lighting/wardrobe combinations at 16:9 with the user offset to the right of frame, leaving room for a name + title overlay. Text is generated by Nano Banana Pro's legible-typography capability — the user types "Dr. Asha Devarajan / Head of Research, Calliope Labs" and the model renders it crisply at 4K alongside the portrait. - **Bulk re-render** — if the user updates the source photo (e.g. a clearer selfie), tap "regenerate all" → all currently-displayed presets re-render in parallel with the new source. - **Download** — per-result download as 4K PNG. A "download all six" button packages them as a zip with a small README.txt naming the preset, crop ratio, and intended use case. - **History** — every generation persists for 30 days in the user's account (or session, for guests). The user can return, regenerate with the same source, swap presets without re-uploading. - **Guest mode** — the first three generations are free, no sign-in required. After three generations the app asks the user to sign in with Google to keep going. No paywall; sign-in is for history and rate-limit accounting. - **Save reference photo for next time** — opt-in. The user's source photo is stored encrypted in Firebase Storage and reused if they return within 30 days. Opt-out by default. ## 4b. Required Gemini capabilities + backend services **This template's intelligence comes from the Gemini capabilities below. Wire them up explicitly — don't substitute generic LLM calls.** ### Gemini capabilities (the load-bearing intelligence) - **Nano Banana Pro identity-preserving portrait generation** (`gemini-3-pro-image`, GA 2026-05-28) — accepts the user's source selfie as the primary reference image plus up to 13 additional style-guide images per preset (the preset's lighting reference, wardrobe reference, background reference, crop reference) for a total of up to 14 reference images per call. Outputs a single 4K portrait at the preset's crop ratio with the user's face geometry preserved exactly. This is the post-I/O 2026 hero capability and the reason this template exists. *Note: Google announced the 14-reference style-guide capability at I/O 2026 but has not publicly pinned the exact request payload shape — verify the SDK call structure against the live `gemini-3-pro-image` docs before shipping.* - **Legible in-image typography** (`gemini-3-pro-image`) — for the Speaker preset only, the user's name + title is rendered crisply at 4K alongside the portrait. Pre-I/O 2026 Nano Banana models could not render legible text; the I/O 2026 Pro model can, and this is the demonstration. - **Source-photo quality check** (`gemini-3.5-flash`, multimodal image input) — before any generation, the source photo is sent to 3.5 Flash to assess face size, eye-openness, focus, tilt, occlusion, and lighting condition. Returns a typed `SourcePhotoCheck` JSON with one of: `proceed`, `proceed_with_caveat`, `recoverable`, `block`. Recoverable failures generate a kind recovery message. - **Preset prompt assembly** (TypeScript on the server) — the literal prompt sent to `gemini-3-pro-image` is assembled from the preset's component fragments (identity anchor, wardrobe directive, lighting directive, background directive, crop directive, negative constraints). The assembly is deterministic and version-pinned; the Gemini call does not write its own prompt. - **Skin-tone & lighting-fairness lint** (`gemini-3.5-flash`, multimodal image input) — after each generation, the result is sent to 3.5 Flash with the source photo for a pairwise comparison on three axes: skin-tone fidelity (the generated skin colour matches the source within ΔE < 5), facial-feature geometry (no silent narrowing or slimming), and lighting fairness (the generated image does not under-light darker complexions or over- expose fairer ones). Returns a `LikenessAudit` JSON. If the audit fails, the result is auto-regenerated once with a higher identity anchor weight before being shown to the user. - **Thinking levels** — the source-photo quality check and the likeness audit both run at `thinkingLevel: low` on `gemini-3.5-flash`. The portrait generation itself is on `gemini-3-pro-image`, which does not accept a `thinkingConfig` field — omit it entirely from that call. - **No grounded search.** This template does not use grounded search on any call. All prompts and audits run from the model's multimodal vision capability alone. - **No long-context.** The longest call is the per-preset image generation (one source photo + up to 13 reference images + a short prompt). Total payload is well under 100k tokens. ### Backend services - **Auth — Required.** Firebase Auth with Google sign-in (auto- provisioned by AI Studio Build, post-I/O 2026 still). **Guest mode** allows three free generations without sign-in (rate-limited by IP + browser fingerprint). After three generations the app prompts for Google sign-in to continue. Apple sign-in is optional and requires the user-configured Apple Developer wiring. - **Database — Required.** Firestore for `users`, `sessions`, `source_photos_meta`, `generations`, `feedback_flags`. Generation blobs themselves live in Storage; Firestore holds metadata (preset, crop, timestamp, likeness-audit score, user feedback). - **File storage — Required.** Firebase Storage for source photos (encrypted at rest, deleted after 30 days unless the user opts in to save) and for 4K PNG outputs (kept for 30 days). **Storage is NOT auto-provisioned by AI Studio Build today** — enable it in the Firebase console and wire the bucket name into the AIS Build project before first upload. Pre-signed URLs only; no public-by- default access. - **Email — Optional.** Not used for v1. A future "send these to your inbox" feature would need a sender-domain-authorised email configuration; not built in v1. - **Payments — Not needed for v1.** The app is free; the per- session Gemini API cost (~$0.18) is borne by the operator. A future "team / studio" tier (bulk corporate headshots for a 200- person company) might charge via Stripe; not built in v1. - **External APIs:** Gemini API for all intelligence. No external API beyond Gemini. **Environment variables:** every secret (Gemini API key, Firebase service-account JSON) lives in environment variables — never in client bundle. Include a `.env.example`. **Auth + data privacy reminders:** never log secrets · never store passwords in plain text · use HTTPS everywhere · honour 'delete my account' inside the UI · explicit opt-in for any analytics · source photos are encrypted at rest and deleted after 30 days unless the user explicitly opts in to "save my reference photo for next time" · generated outputs are the user's property, no Google watermark, no Headshot Studio branding, SynthID watermark ships embedded as it ships on every `gemini-3-pro-image` output and is disclosed in the capabilities panel · this app uses the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. **Read this first — prompt-craft rules that apply to every call in this template:** 1. **Name the model variant explicitly** in every Gemini API call. Do not let the agent pick the model. See the per-call matrix below. Specifically: `gemini-3-pro-image` for portrait generation (this is the post-I/O 2026 Pro image model, not `gemini-3.5-flash- image` and not the deprecated `gemini-3-pro-image`); and `gemini-3.5-flash` for source-photo checks + likeness audits (this is the post-I/O 2026 default, not the legacy `gemini-3.5-flash` or `gemini-3.5-flash` strings). 2. **Pin `thinkingLevel` explicitly** per call where supported. See the matrix. `gemini-3-pro-image` does NOT support `thinkingConfig` — omit the field entirely from those calls. 3. **Seed the JSON Schema as a fenced TypeScript / Zod block** in the system instruction or `responseSchema` field. The literal schemas are below. **Convert the Zod schema to Gemini's `Schema` type via the SDK helper** before passing to `responseSchema` — do NOT pass raw Zod. **Numeric `min`/`max` constraints are documentation only inside `responseSchema`; clamp on the server after the response arrives.** 4. **Pin the system instruction separately** from user input. Use the `systemInstruction` field for persona + behavioural rules; use `contents` for the user's photo + the assembled preset prompt. Never concatenate. 5. **Pre-declare tools as an enable/disable list** per call. This template enables NO tools on any call — no `google_search`, no `code_execution`, no function calling. Tools NOT listed for a call should be disabled. 6. **State negative constraints explicitly** — they are listed below. They are NOT "be careful" suggestions; they are hard rules the model must follow. 7. **Reference-image upload via Files API.** Source photos and preset reference images go to the Gemini Developer API Files API and are referenced via `fileData: { fileUri: "files/abc123xyz", mimeType }` (the `files/*` resource name returned by the Files API `upload` response — NOT a `gs://` URI; `gs://` belongs to Vertex AI / Cloud Storage, a different surface). Do NOT pass Firebase Storage public URLs — the API does not fetch them server-side. For source photos under 5 MB, base64 `inlineData` is acceptable but the Files API is preferred for the 4K reference style guides. 8. **Strip unsupported Zod modifiers before passing to `responseSchema`** — Gemini's OpenAPI subset rejects `.regex()` / `pattern`, fixed-length `z.tuple()`, and other custom validators. Use a sanitizer that flattens tuples to arrays and removes regex patterns before serializing. Validate those constraints in middleware AFTER parsing. ### Per-call model + tools matrix | Call | Model | thinkingLevel | Tools enabled | |------|-------|---------------|---------------| | Source-photo quality check (face size, eyes open, focus, tilt, occlusion, lighting) | `gemini-3.5-flash` | low | (none) | | Portrait generation per preset (Corporate / Editorial / Academic / Artisan / Tech / Creative / Speaker) | `gemini-3-pro-image` | n/a | (none) | | Likeness audit (post-generation, pairwise vs source) | `gemini-3.5-flash` | low | (none) | | "Doesn't look like me" regeneration (identity-anchor weight bumped) | `gemini-3-pro-image` | n/a | (none) | | Welcome / empty-state illustration generation | `gemini-3.1-flash-image` (Nano Banana 2) | n/a | (none) | *Note for builders:* `gemini-3-pro-image` and `gemini-3.1-flash-image` do NOT accept `thinkingConfig` — the `n/a` cells in this matrix are documentation only; do not serialise them into the request body. No call in this template uses `google_search` grounding, so the `responseSchema` + grounding mutual-exclusion issue does not arise here. `responseSchema` is used on the source-photo check and the likeness audit calls (both on `gemini-3.5-flash`, both tool-free). ### Primary structured-output schemas (seed verbatim in the prompt) ```typescript import { z } from "zod"; const PresetId = z.enum([ "corporate", "editorial", "academic", "artisan", "tech", "creative", "speaker", ]); const CropRatio = z.enum([ "1:1", // LinkedIn, Slack, team page "3:4", // press release, faculty page, CV portrait "4:5", // portfolio grid, gallery bio "5:7", // book jacket, magazine "16:9", // speaker deck, conference banner ]); const SourcePhotoCheck = z.object({ decision: z.enum([ "proceed", "proceed_with_caveat", "recoverable", "block", ]), face_fraction_of_frame: z.number().min(0).max(1), primary_face_in_focus: z.boolean(), eyes_open: z.boolean(), head_roll_degrees: z.number(), // 0 = level, signed head_yaw_degrees: z.number(), head_pitch_degrees: z.number(), occlusions_detected: z.array(z.enum([ "sunglasses", "mask", "hand", "hair_over_eyes", "hat_brim_shadow", "phone_in_frame", "other", ])), lighting_condition: z.enum([ "balanced", "front_blown", "backlit_silhouette", "side_lit_strong", "low_light_grainy", "fluorescent_colour_cast", "mixed_colour_cast", "unknown", ]), apparent_skin_tone_bin: z.enum([ // for fairness-aware lighting "very_fair", "fair", "light_medium", "medium", "medium_deep", "deep", "very_deep", "uncertain", ]), recovery_message_to_user: z.string().nullable(), // human-readable caveat_for_generation: z.string().nullable(), // model-readable check_confidence: z.number().min(0).max(1), }); const PresetSpec = z.object({ preset_id: PresetId, preset_name_display: z.string(), // "Corporate" crop_ratio: CropRatio, output_width_px: z.number(), // 4096 for the long edge output_height_px: z.number(), intended_use_case: z.string(), // "LinkedIn profile photo" lighting_directive: z.string(), // model-readable wardrobe_directive: z.string(), // model-readable background_directive: z.string(), // model-readable identity_anchor_weight: z.enum([ // controls how hard to lock face "default", "elevated", "max", ]), }); const GenerationResult = z.object({ generation_id: z.string(), preset_id: PresetId, source_photo_uri: z.string(), // files/* (Developer API resource name) output_uri: z.string(), // files/* (Developer API resource name) output_width_px: z.number(), output_height_px: z.number(), crop_ratio_used: CropRatio, prompt_used_verbatim: z.string(), // what we sent to Gemini latency_ms: z.number(), synthid_watermark_present: z.boolean(), // always true; surfaced in UI generated_at_iso: z.string(), }); const LikenessAudit = z.object({ generation_id: z.string(), source_photo_uri: z.string(), output_uri: z.string(), skin_tone_match_delta_e: z.number(), // 0 = identical, lower = better skin_tone_match_pass: z.boolean(), // delta_e < 5 facial_geometry_pass: z.boolean(), // no silent narrowing/slimming lighting_fairness_pass: z.boolean(), // dark skin not under-lit, fair not blown identity_likeness_score: z.number().min(0).max(1), audit_notes: z.string(), // 1-2 sentences needs_regeneration: z.boolean(), regeneration_reason: z.string().nullable(), }); const UserLikenessFeedback = z.object({ generation_id: z.string(), user_said_doesnt_look_like_me: z.boolean(), user_free_text: z.string().nullable(), submitted_at_iso: z.string(), }); type SourcePhotoCheck = z.infer; type PresetSpec = z.infer; type GenerationResult = z.infer; type LikenessAudit = z.infer; type UserLikenessFeedback = z.infer; ``` ### Common failure modes (and how to avoid them) - Agent picks `gemini-3.1-flash-image` (Nano Banana 2) for the portrait call to save quota — pin `gemini-3-pro-image` (Nano Banana Pro) explicitly. Nano Banana 2 cannot render legible text for the Speaker preset and produces noticeably softer identity preservation. Use it only for the welcome / empty-state illustrations. - Agent picks the deprecated `-preview` IDs (`gemini-3-pro-image- preview`, `gemini-3.1-flash-image`) — those endpoints shut down 2026-06-25. Use the GA IDs (`gemini-3-pro-image`, `gemini-3.1-flash-image`) without the `-preview` suffix. - Agent silently lightens or darkens skin tone toward a studio- portrait "default" — the lighting-fairness lint catches this post-hoc; the system instruction on the portrait call also forbids it explicitly. - Agent thins nose / lifts eyes / narrows jaw — the identity- anchor weight in the prompt and the negative constraints forbid this. The likeness audit checks for it pairwise. If detected, one regeneration is attempted with `identity_anchor_weight: max`; if it fails again, the result is shown to the user with a "this may not look like you — try a clearer source photo" caveat, never silently shipped. - Crop ratio mismatch — the preset declares a crop ratio, the prompt requests it, the output is verified server-side after generation. If the dimensions returned do not match the requested ratio (within 2% tolerance), the output is centre- cropped to the declared ratio before being shown. - Wardrobe defaults that don't fit the user's apparent presentation — the wardrobe directive in each preset is offered in two cuts by default (a structured cut and a relaxed cut); the user can switch via the wardrobe override. The model is never told the user's gender; wardrobe is offered, not assumed. - Speaker-preset typography hallucinated — the user-supplied name and title are passed verbatim to the prompt with explicit "render this text exactly as typed" instruction. Post- generation, the rendered text is OCR'd by `gemini-3.5-flash` and compared character-by-character to the input; if they don't match, the generation is rejected and re-requested. - Source photo passed as a Firebase Storage public URL — the Gemini API does not fetch arbitrary public URLs. Source photos upload to the Gemini Developer API Files API and are referenced via `fileData.fileUri: "files/abc123xyz"` (the `files/*` resource name, NOT a `gs://` URI). - Three free guest generations enforced client-side only — the rate-limit must enforce server-side (IP + browser fingerprint + Firebase anonymous-auth UID) or it is trivially bypassable. - User's reference photo persisted without opt-in — the default is delete-after-session for the source photo; "save my reference photo for 30 days" is an explicit, unticked checkbox in the upload flow. - Identity-preservation prompt drifts across presets — the identity-anchor fragment of each preset's prompt is identical verbatim across all six presets; only the lighting / wardrobe / background / crop fragments differ. This guarantees the user's face is locked the same way in every preset. ### Negative constraints (hard rules) - Do NOT alter the user's face geometry. Nose shape, eye spacing, jaw width, chin shape, cheekbone height, forehead height, ear shape, lip thickness, the natural asymmetry every face has — all preserved exactly. The likeness audit checks this; the system instruction states it explicitly; the negative constraints in the prompt list every facial feature by name. - Do NOT lighten or darken the user's skin tone beyond what natural studio lighting variation would produce (ΔE < 5 vs source). Do NOT use a default "studio-portrait" skin tone that the model has been trained to expect. Match the source. - Do NOT smooth skin beyond what natural studio lighting would produce. Pores, fine lines, scars, freckles, blemishes, beard shadow, makeup — all visible at 4K, as they are visible in real studio portraits. - Do NOT remove glasses unless the user explicitly opts out of the glasses-on-face preservation toggle. - Do NOT remove or change visible religious or cultural items (kippah, hijab, turban, cross necklace, mangalsutra, tilak, bindi, kufi, etc.) — these stay exactly as in the source unless the user explicitly opts out via the wardrobe override. - Do NOT alter the user's hairstyle, hair colour, or hair length beyond what the wardrobe preset's hair-styling note specifies. Hair is part of identity. - Do NOT alter visible tattoos, piercings, or other body modifications — preserved as in the source. - Do NOT generate a portrait of a different person and ship it with the source for comparison. The likeness audit must pass; if it fails twice, surface the failure to the user honestly. - Do NOT render commercial branding on wardrobe or background — no visible logos, no brand names. Wardrobe is generic studio wardrobe. - Do NOT remove SynthID watermarks. They ship embedded in every `gemini-3-pro-image` output by Google and the app surfaces this honestly in the capabilities panel and the download README. - Do NOT moralise about the user's appearance — no "you'd look more professional if". The presets offer; they do not judge. - Do NOT auto-share or publish any generation. Downloads are manual; the user owns them and decides where they go. - Do NOT use the user's source photo, generations, or feedback to train or fine-tune any model. Use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. The capabilities panel says this in plain English. - Do NOT claim the output is a photograph. The download EXIF flags AI-generation; the on-screen label reads "AI-generated portrait, reference photo: yours". - Do NOT generate portraits of a person other than the uploader. Source photos must be the user's own face; the source-photo quality check rejects multi-face inputs and requires the user to crop to a single primary face before proceeding. ### Per-call `systemInstruction` strings Use these as the literal `systemInstruction` field for each Gemini API call the built app makes. They complement the series-wide rules already uploaded as the global instructions file. ### Call: Source-photo quality check Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You receive a single photograph that a user wants to use as the source for a studio-quality portrait generation. Your job is to assess whether the photograph is suitable, and if not, to explain to the user in one short kind sentence what they could do differently. Inspect the photograph for: - Face size in frame. The primary face should occupy at least 25 percent of the frame area. If less, the generation will lack the resolution to preserve identity. - Number of faces. Exactly one primary face is required. If there are multiple faces, identify the primary (largest, most centred) and report that the source must be cropped to a single face. - Eyes. Both eyes must be visible and open. Squinting against sunlight is acceptable; closed eyes are not. - Focus. The primary face must be in focus. Motion blur over the face is disqualifying. - Tilt. Head roll within 30 degrees; head yaw within 45 degrees; head pitch within 30 degrees. Beyond these, generation quality degrades. - Occlusions. Flag sunglasses, masks, hands covering the face, hair over eyes, hat brims casting shadow over eyes, phones in frame. Sunglasses and masks are disqualifying. Hand on chin or thoughtful pose is fine. - Lighting. Identify front-blown, backlit silhouette, strong side light, low-light grainy, fluorescent or mixed colour casts. Generation can usually correct, but flag it. - Apparent skin tone bin. Estimate to one of seven bins (very_fair / fair / light_medium / medium / medium_deep / deep / very_deep) or "uncertain". The downstream generation prompt uses this to set fairness-aware lighting parameters. If you cannot estimate confidently, return "uncertain" rather than guessing — the generation prompt will use neutral lighting in that case. Output a SourcePhotoCheck JSON object matching the provided schema. decision logic: - "proceed" — all checks pass; generation will produce a good result. - "proceed_with_caveat" — minor issues (mild lighting cast, slight tilt) that the model can correct; include a `caveat_for_generation` string the downstream prompt should append. - "recoverable" — the user should retake the photo; include a `recovery_message_to_user` in plain, kind English ("we need to see your eyes — try a photo without sunglasses"). - "block" — the photo cannot be used (no face visible, multiple faces with no clear primary, extreme corruption). Provide a recovery_message_to_user. Hard rules: - Do NOT comment on the user's appearance. The check evaluates the photograph as input, not the person. - Do NOT estimate age, gender, presentation, or ethnicity beyond the skin-tone bin used for fairness-aware lighting. Wardrobe is offered, not assumed. - The recovery_message_to_user is short, kind, specific, and written for the user to read directly — never robotic ("INVALID INPUT") and never apologetic past one syllable ("try a photo without sunglasses" is the right shape; "we are deeply sorry but unfortunately your photograph does not meet our requirements" is not). - check_confidence reflects how certain you are about the decision. Below 0.7, prefer proceed_with_caveat over block. Output ONLY the SourcePhotoCheck JSON. No commentary. JSON only. ``` --- ### Call: Portrait generation per preset Model: `gemini-3-pro-image` (Nano Banana Pro) · thinkingLevel: n/a · Tools: (none) ``` You are generating a single 4K studio-quality portrait from a user's source photograph and a preset specification. You will receive: - The user's source photograph, as the primary reference image. Treat this as the IDENTITY ANCHOR. Every facial feature, every aspect of skin tone and texture, every visible cultural or religious item, hair colour and length, glasses, scars, freckles, tattoos, asymmetry — all must be preserved EXACTLY in the output. The face of the person in your output IS the face of the person in the source photo. - Up to 13 additional reference images per call, supplied by the preset: a lighting reference, a wardrobe reference, a background reference, a crop-ratio reference, and additional variant references. Use these for LIGHTING / WARDROBE / BACKGROUND / CROP only. Do NOT take facial features from these references. - A text directive specifying the preset's intended use case, lighting, wardrobe, background, and crop ratio. The text directive is structured as: PRESET: USE CASE: CROP RATIO: OUTPUT DIMENSIONS: LIGHTING: WARDROBE: BACKGROUND: IDENTITY ANCHOR WEIGHT: FAIRNESS-AWARE LIGHTING BIN: Hard rules — IDENTITY: - The face in the output IS the face in the source. Geometry of every feature preserved: nose shape and width, eye shape and spacing, jaw width, chin shape, cheekbone height, forehead height, ear shape, lip thickness, the natural asymmetry every face has. Do not silently regularise asymmetry. - Skin tone matches the source within natural studio-lighting variation. Do NOT lighten or darken to a default "studio portrait" tone. The fairness-aware lighting bin in the directive tunes the lighting setup for the user's actual skin tone — soft fill on darker complexions, no blown highlights on fairer ones. - Hair colour, length, and style preserved as in the source unless the preset's wardrobe directive specifically says otherwise (e.g. "hair pulled back from face" is an allowed styling note; "make hair shorter" is not). - Glasses preserved unless the source photo shows none. - Cultural and religious items preserved exactly: kippah, hijab, turban, cross necklace, mangalsutra, tilak, bindi, kufi, and every equivalent. These are part of identity. - Visible tattoos, piercings, scars, freckles, beard, moustache, stubble — all preserved. - Skin texture preserved: pores, fine lines, beard shadow, real skin imperfection. NO smoothing beyond what natural studio lighting would produce. NO "beautify" filter behaviour. Hard rules — OUTPUT: - Output at the exact crop ratio requested. If the user's source is portrait and the requested crop is landscape, generate fresh framing — do NOT stretch or letterbox. - Output at 4K (4096 px on the long edge). The Pro image model supports 4K natively. - Render any user-supplied text (Speaker preset only) at exactly the typography specified, with no character substitution and no hallucinated additional text. The Pro image model renders legible text at 4K; use that capability. - No commercial branding visible on wardrobe or background. No logos, no brand names, no recognisable luxury items. - No props the directive did not specify. - No moralising about the user's appearance. If the IDENTITY ANCHOR WEIGHT is "elevated" or "max" (set when the user has flagged a previous generation as "doesn't look like me"), lock the facial feature preservation harder still and relax the wardrobe / lighting freedom proportionally — when the two conflict, identity wins. Output: the 4K portrait image. No commentary. ``` --- ### Call: Likeness audit (post-generation) Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You receive two images: the user's source photograph and the generated portrait. Your job is a strict pairwise comparison on three axes: 1. SKIN TONE FIDELITY. Compare the average skin tone of the visible face in both images. The generated skin colour must match the source within a perceptual delta of ΔE < 5. Lighting variation is acceptable; tonal shift beyond ΔE 5 is not. 2. FACIAL FEATURE GEOMETRY. Compare facial feature geometry side by side. Check nose shape and width, eye shape and spacing, jaw width, chin shape, cheekbone height, lip thickness, the asymmetry of every face. The generated face must look like the same person as the source. A different-looking person fails. 3. LIGHTING FAIRNESS. For darker skin tones, the generated image must not be under-lit (face details and features clearly visible; shadows give depth, not erase information). For fairer skin tones, the generated image must not have blown highlights (no white-clipped patches on cheeks or forehead). Output a LikenessAudit JSON object matching the provided schema. Decision logic: - All three axes pass → needs_regeneration: false. - Any axis fails → needs_regeneration: true. Populate regeneration_reason with one specific sentence ("skin tone in output is two stops lighter than source — regenerate with elevated identity anchor weight"). The next generation will run with identity_anchor_weight: "elevated" or "max". identity_likeness_score is your overall best-effort score from 0 to 1. Score components: 0.4 for facial geometry, 0.3 for skin tone fidelity, 0.2 for lighting fairness, 0.1 for general plausibility. Hard rules: - Be strict. The cost of letting a bad-likeness generation through is the user looking at it and feeling that the app did not preserve their face. The cost of a false negative (one extra regeneration) is eighteen seconds and three cents. - Do NOT comment on the user's appearance in audit_notes — audit the resemblance to the source, not the user. - Do NOT moralise about the user's choices of wardrobe or background; the audit is on identity preservation only. Output ONLY the LikenessAudit JSON. No commentary. JSON only. ``` --- ### Call: "Doesn't look like me" regeneration Model: `gemini-3-pro-image` · thinkingLevel: n/a · Tools: (none) ``` The user has flagged a previous generation as "doesn't look like me". You are re-running the same preset with the same source photo but with the IDENTITY ANCHOR WEIGHT bumped from "default" to "elevated" (or "elevated" to "max"). Apply the same rules as the standard portrait generation call, with three modifications: - Identity wins over lighting/wardrobe/background freedom. If a realistic studio lighting setup would slightly shift the perceived face shape, choose the lighting setup that preserves the face shape, even if it is a slightly less typical studio lighting. - Skin tone fidelity is locked tighter. ΔE target drops from 5 to 3. - Facial geometry tolerance is tighter. Where the previous generation may have silently regularised asymmetry, the regeneration preserves it precisely. The user's free-text feedback (if provided) is appended to the prompt verbatim: "User said: ''. Take this feedback seriously and apply it to this regeneration." Output: the 4K portrait image. No commentary. ``` --- ### Call: Welcome / empty-state illustration generation Model: `gemini-3.1-flash-image` (Nano Banana 2) · n/a · Tools: (none) ``` You generate a single photographic-looking image for the welcome screen or the empty-state of Headshot Studio. Images depict a real studio environment: a softbox at golden hour, an empty seamless paper backdrop, a wooden stool in front of a graphite wall, a 50mm lens on a tripod beside a reflector. Prompt anchors that work well: - "soft empty studio with a charcoal seamless backdrop, a softbox at frame-left casting warm key light, an empty wooden stool centre-frame, no people in frame, shallow depth of field" - "studio table at golden hour through a window, a reflector leaning on the wall, an open notebook with a pencil beside a warm coffee mug, no people" - "a row of six framed portrait silhouettes on a graphite wall, empty silhouettes — no identifiable faces — meant to be filled, soft warm light from above" Hard rules: - Photographic, not cartoon, not illustration-style. - No people in frame. No faces. The welcome illustration is the empty studio waiting for the user. - No commercial branding visible. - Warm professional lighting, slight imperfection, no glossy AI- render look. - Aspect ratios: 3:2 for hero, 1:1 for empty states. This is the only call in the app that uses Nano Banana 2 instead of Nano Banana Pro — the welcome imagery doesn't need 4K typography or 14-reference style guides. ``` ## 5. Use cases & content to include Build dedicated UI sections or flows for each of these — they tell you what content the app must support. - **The Monday morning LinkedIn refresh.** A user changed jobs last Friday. Their new manager added them to the company LinkedIn page this morning. The current profile photo is a holiday selfie from three summers ago. They open Headshot Studio at 9:14am, upload the holiday selfie, tap "Corporate", and at 9:14:23 they download a 1:1 4K LinkedIn-ready portrait in a charcoal two-piece against a warm graphite background. By 9:16 it is uploaded to LinkedIn. - **The conference bio at midnight.** A speaker submits to a conference at 11:58pm. The bio form needs a 3:4 portrait, 4K. The speaker has one selfie on their phone from earlier this week. They upload it, tap "Editorial", and download the 3:4 high-contrast portrait against a graphite seamless paper backdrop at 11:58:30. - **The faculty page two days before semester.** A new lecturer joins a department on Monday. The faculty-page editor needs the portrait by Sunday night. The lecturer uploads a candid from a conference the prior month, taps "Academic", and downloads a 3:4 warm window-lit portrait in a herringbone blazer against a soft bookshelf-suggestive background. - **The press release at lunch.** A user is quoted in a press release going out at 2pm. The PR team asks for a 5:7 high- resolution headshot by 1:30. The user uploads a selfie taken in their kitchen, taps "Editorial", and downloads a 5:7 4K portrait in time. - **The wedding announcement the day before.** The engaged couple realises their announcement goes out tomorrow and the only portraits they have are casual phone pictures. They each upload a selfie, tap "Editorial" with the soft-paper background variant, and download two matching 3:4 portraits at 4K. - **The Airbnb host who just got a co-host onboarded.** Both hosts upload one selfie each. Each gets a Corporate-preset 1:1 portrait ready for the Airbnb host page. The listing visually upgrades inside ten minutes. - **The artisan portfolio refresh.** A potter in a small studio needs a portrait for their new portfolio site. They upload a selfie taken at their wheel, tap "Artisan", and download a 4:5 warm workshop-lit portrait in a denim apron against a brick wall with hints of pottery shelves in the background. - **The fundraising round headshot.** A founder is preparing the deck for a Series A pitch on Friday. The deck needs a 16:9 hero with the founder offset to the right and their name + title rendered legibly. The user uploads a selfie, taps the Speaker preset, types "Asha Devarajan / Co-founder, Calliope Labs", and downloads a 16:9 4K hero with the typography rendered crisply beside the portrait. - **The "doesn't look like me" recovery.** A user receives a generation that they feel softens their nose. They tap "this doesn't look like me" under the generation. The regeneration runs with elevated identity-anchor weight and the result lands in five seconds with the nose preserved exactly. The user verifies with the side-by-side compare strip. - **The source-photo recovery flow.** A user uploads a photo of themselves wearing sunglasses on a beach. The source-photo check returns `recoverable` with the message "we need to see your eyes — try a photo without sunglasses". The user uploads a different selfie. The check passes. - **The cultural / religious item preservation.** A user wearing a kippah uploads a selfie. Every generation preserves the kippah exactly. A user wearing a hijab uploads a selfie. Every generation preserves the hijab as worn in the source. A user with a tilak preserves the tilak. The wardrobe override does not override these by default. - **The reading-glasses kept on.** A user wears reading glasses in every professional context. Every generation preserves the glasses. The wardrobe override has a checkbox "remove glasses for this generation" — unchecked by default. ## 6. Page structure Build the following screens / sections in this order. Adjust copy to fit the voice, but keep the structural intent. 1. **Welcome / canvas (single page).** A photographed-looking empty studio image in the hero — a softbox at frame-left, an empty wooden stool centre-frame, a graphite wall behind. One paragraph beside it: "Upload one ordinary selfie. Walk away with six studio- grade you's in 4K. No appointment, no subscription, no Google watermark." Single upload zone (drag-drop or tap to pick or take-a-photo). Below the zone, the six preset cards as a 2×3 grid (Corporate, Editorial, Academic, Artisan, Tech, Creative), each with a small lighting-reference thumbnail and a 1-line intent ("Corporate — LinkedIn, team page, recruiter inbox"). A "Speaker preset" tile below the six, marked as optional. 2. **Source-photo check banner.** Once the upload finishes, `gemini-3.5-flash` runs the quality check. The banner above the preset grid surfaces the result: green check + "your photo's good — pick a preset" / amber + "your photo will work; we'll correct for the slight backlight" / red + the kind recovery message and a "try another" button. Block decisions disable the generation buttons until a new photo is uploaded. 3. **Generation in progress.** Tap a preset → a card slides into the results strip with a skeleton placeholder. The first preset's image streams in around the 5-second mark; subsequent presets land as each generation completes. A small "running likeness audit…" indicator appears on each generation for ~1 second before the result is finalised; if the audit fails, the indicator shifts to "regenerating for a closer match…" and the auto- regenerated result replaces the first. 4. **Results strip with side-by-side compare.** Each completed generation sits in a results strip beneath the preset grid. The source photo is shown as a small fixed tile at the strip's left edge. Each generation card has the preset name + crop ratio labelled, a "this doesn't look like me" tap target, a wardrobe override dropdown, a background override palette, and a "download 4K PNG" button. Holding "compare" momentarily replaces the generation with the source full-size for direct comparison. 5. **Wardrobe override drawer.** Tap the wardrobe label → a small drawer opens with three or four alternates for that preset, each shown as a thumbnail. Pick one → that single generation re-runs with the new wardrobe. 6. **Background override drawer.** Same pattern for background alternates. 7. **Speaker preset modal.** Tap the Speaker tile → a small modal asks for the user's name and title (two text fields, plain English). The user types "Asha Devarajan / Co-founder, Calliope Labs". The modal previews the typography style. Tap "generate" → a 16:9 4K Speaker portrait is produced with the typography rendered crisply beside the offset portrait. Six variants are produced (one per preset's lighting mood) so the user picks the match. 8. **Download all.** A "download all six" button at the top of the results strip packages every completed generation as a zip containing six 4K PNGs and a small README.txt naming the preset, crop ratio, and intended use case for each. (Pre-zipped server- side; signed URL.) 9. **History view (signed-in users).** A simple list of past sessions over the last 30 days, each session showing the source thumbnail and the preset thumbnails generated. Tap a session → the canvas reopens with that source and those generations restored. Useful for "I downloaded the LinkedIn one but I need the speaker one too". 10. **Settings & privacy.** Saved-source-photo toggle (default off), history retention (default 30 days), "delete all my data" control with 60-second cool-off. Privacy panel restates the not-trained-on policy in plain English and surfaces the SynthID watermark disclosure. 11. **Footer.** "Made for the user who needs a studio-grade headshot today." Privacy: "Your photos and your generations are yours. We never train on them." Capabilities `(i)` icon in header. ## 6b. First-visit onboarding Show a **first-visit onboarding** the first time a visitor lands on the app (detect via `localStorage` flag; do not show on return visits). Three slides, dismissible at any time. Persistent re-entry: a `?` icon in the header reopens it. **Slide 1 — What this is.** - Headline: "One selfie in. Six studio-grade you's out." - Subhead: "Studio-grade portraits in 30 seconds, at 4K, across six explicit presets. Your face — exactly. No subscription. No watermark." - One paragraph (≤ 60 words) explaining who this is for and what makes it different from a paid headshot service: the speed (eighteen seconds), the identity preservation (your face, exactly), the resolution (4K), the cost (~$0.18 of Gemini API per session, free to the user), and the model (Nano Banana Pro, the post-I/O 2026 image model that finally renders 4K with legible typography and 14-reference style guides). - Visual: a small annotated diagram of the six presets in a 2×3 grid with the source photo at the centre and arrows fanning out. **Slide 2 — Try it now.** - One short prompt: "Drop a selfie here, or pick one from your camera roll". - A live drop zone in the slide itself. - 1-2 sentences pointing at *the specific page elements* where the Gemini magic happens (the source-photo quality check banner, the results strip, the side-by-side compare). **Slide 3 — How to remix this.** - Headline: "Make this yours." - Three short bullets: - "Swap the six presets in `/data/presets/` for your own preset library (Boudoir? Maternity? Pet portrait?)." - "Adjust the per-call prompts in `/server/prompts/` to fit your domain — every word is in the repo, no hidden prompts." - "Wire up your Gemini API key and Firebase project via the env-var list in the capabilities panel." - Primary CTA: "Use this template" → links to AI Studio Build remix entry point. - Secondary: "Just exploring — close" (sets localStorage flag, never auto-shows again). **Accessibility:** focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, `aria-labelledby`, focus restored to trigger on close. Respect `prefers-reduced-motion`. **Don't:** - Don't gate content behind the modal. The page beneath must be fully usable. - Don't auto-reshow on return visits. Use `localStorage['onboarding-seen-v1']`. - Don't include unrelated CTAs (newsletter signup, social follow). Keep it about the template only. ## 6c. Capabilities info button (persistent in header) Add a persistent `(i)` icon in the top-right of the header (next to the primary nav). Click → opens a modal/panel titled **"What powers this app"**. **Panel contents (in this order):** **Gemini capabilities used (the hero list):** - **Nano Banana Pro (`gemini-3-pro-image`)** — the post-I/O 2026 image model. Generates the six 4K portraits per session. Accepts the user's source selfie as the identity anchor and up to 13 additional style-guide reference images per call. Renders legible text at 4K for the Speaker preset. SynthID watermark embedded in every output by Google (disclosed, never removed). - **Gemini 3.5 Flash multimodal (`gemini-3.5-flash`)** — runs the source-photo quality check before generation and the likeness audit after each generation. Both at `thinkingLevel: low`. The post-I/O 2026 default Flash model. - **Structured output / JSON Schema** — the source-photo check returns a typed `SourcePhotoCheck`; the likeness audit returns a typed `LikenessAudit`. Schemas live in the repo. - **Nano Banana 2 (`gemini-3.1-flash-image`)** — generates the welcome / empty-state illustrations (the empty studio, the softbox, the wooden stool). Pro is overkill for backgrounds. - **Firebase Auth** — Google sign-in (after three guest generations). Apple sign-in next to it (requires user-configured Apple Developer wiring). - **Firestore** — stores your generation history, syncs across devices. - **Firebase Storage** — keeps your source photos (encrypted, deleted after 30 days unless opted in) and 4K outputs (kept for 30 days). **Cost note** — a full session (one source photo, six presets, the quality check, six likeness audits, one welcome illustration if first-visit) costs about **$0.18 of Gemini API spend** end-to-end. See the detailed breakdown in 6d. **Privacy note** — your source photos, your generations, and your likeness-feedback are private to you. This app uses the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. Source photos are encrypted at rest and deleted after 30 days unless you explicitly opt in to "save my reference photo for next time". Generated outputs ship with the SynthID watermark Google embeds in every Nano Banana Pro output — this is disclosed; it cannot be removed. **Identity-preservation note** — this app's hard rule is to preserve your face geometry, your skin tone, your hair, your cultural and religious items, your glasses, your tattoos, your asymmetry. After every generation a Gemini 3.5 Flash audit runs pairwise against your source photo on three axes: skin-tone fidelity (ΔE < 5), facial- feature geometry (no silent narrowing), and lighting fairness (no under-lit darker tones, no blown highlights on fairer tones). Audit failures trigger one auto-regeneration with elevated identity preservation; if it fails twice, you see the result with an honest "this may not look like you" caveat rather than a quietly-bad portrait. **Backend services this app depends on:** - Auth: see section 4b - Database: see section 4b - Storage: see section 4b — REQUIRES manual enable in Firebase console; AIS Build does not auto-provision Storage today. - Email: not used in v1. - Apple sign-in: optional, requires an Apple Developer account and Service-ID config. - Payments: not used in v1 — the app is free to the user; per- session Gemini API cost is borne by the operator. - External APIs: only Gemini. **Environment variables you'll need to configure:** - `GEMINI_API_KEY` — your Google AI Studio API key - `FIREBASE_PROJECT_ID` — your Firebase project id - `FIREBASE_SERVICE_ACCOUNT` — service-account JSON (server-side only) - `FIREBASE_STORAGE_BUCKET` — your bucket name (post manual enable) **Documentation links:** - AI Studio Build docs - Gemini API multimodal image (Nano Banana Pro), multimodal image understanding (Gemini 3.5 Flash) docs - Firebase Auth, Firestore, Firebase Storage docs - SynthID watermarking documentation **Accessibility:** same standards as the onboarding modal — focus trap, `Esc`, ARIA, restored focus. **Behaviour:** - Always available — single click from anywhere in the app. - Tooltip on the `(i)` icon: "How this app is built". - Mobile: opens as a full-screen sheet that slides up. - Should be the most honest part of the app — never hand-wave service requirements; never say "AI" without naming the specific Gemini model and capability. ## 6d. Detailed cost breakdown (deployer reads this BEFORE shipping) - **Source-photo quality check (Gemini 3.5 Flash, low thinking)** — one image + a short system instruction → ~1,200 input tokens and ~200 output tokens. At Gemini 3.5 Flash pricing ($1.50/M input, $9.00/M output, global tier) that is ~$0.0036 per check. Once per session. - **Portrait generation per preset (Nano Banana Pro, `gemini-3-pro-image`)** — per Google's I/O 2026 pricing, Nano Banana Pro is ~$2.00/M input / $12.00/M output. A single 4K portrait call averages ~3,500 input tokens (source + style guide images encoded + short prompt) and the image output is billed per-image at the Pro tier. The per-image price is approximate (~$0.02 per generated 4K image as a working estimate — Google has not pinned an exact public per-image figure; verify on the pricing page before shipping). Six presets per session → ~$0.12 on the working estimate. - **Likeness audit (Gemini 3.5 Flash, low thinking)** — source + output + audit instruction → ~1,500 input tokens and ~300 output tokens → ~$0.0050 per audit. Six audits per session → ~$0.030. - **"Doesn't look like me" regeneration (Nano Banana Pro)** — same approximate cost as a standard generation (~$0.02 working estimate). Expected ~0.2 per session on average → ~$0.004. - **Welcome / empty-state image (Nano Banana 2, `gemini-3.1-flash-image`)** — approximate ~$0.005 per image (verify before shipping). Generated once per fresh deploy (cached client-side), or once per regenerate-empty-state action. - **Expected per-session cost (full lifecycle):** ~$0.18 — one source-photo check, six Pro portrait generations, six likeness audits, occasional regeneration. Generously rounded up to $0.20 for budgeting. - **Storage:** Firebase Storage standard tier ~$0.026/GB/month. A source photo at ~4 MB plus six 4K PNGs at ~5 MB each ≈ ~34 MB per session ≈ ~$0.0009/month per session retained. Default 30-day retention → ~$0.001 per session total storage cost. - **At scale (1,000 sessions/day):** ~$180/day Gemini API spend + ~$1/day storage. The operator decides whether to charge or to swallow the cost; per-session it remains cheaper than any paid AI-headshot service. ## 7. Design language - **Mood:** A real photographer's working surface, not a SaaS dashboard. Not a tech product. The empty studio at 9am before the first appointment: a softbox at frame-left, a wooden stool centre frame, a graphite seamless backdrop, a 50mm lens on a tripod in the corner, a reflector leaning against the wall, warm window light coming in from the right. Calm. Considered. The user is the subject; the app is the studio. - **Typography:** Clean grotesque for app chrome and data labels (Inter or Geist). Display serif for preset names and the hero headline (Source Serif Pro or Fraunces). A neutral monospace for technical metadata in the capabilities panel (file sizes, model IDs). - **Palette:** Warm paper background `#F4EFE6` for the canvas surface, deep ink `#1B1714` for body text, graphite `#2A2A2D` for the empty-studio dark accents, warm tungsten `#D6A55C` for active states and download CTAs, soft sage `#7A9486` for the source- photo-check green pass state, muted clay `#B66E5B` for the source- photo-check amber caveat state, restrained red `#A33A2C` for the source-photo-check red block state. Borrowed from a working photographer's studio — not from SaaS design systems. - **Imagery:** Photographic. The empty-studio hero image generated via Nano Banana 2. No flat illustrations. No camera-icon glyphs. No "smile" emoji. - **Hand-feel touches:** Each new generation slides into the results strip with a faint slide+fade as if a print were emerging from a darkroom tray. The source photo in the compare strip has a hand-crafted paper edge. The wardrobe override drawer opens with a sliding-fabric haptic note. - **Spacing:** consistent 4-px base. Generous whitespace — the preset grid and the results strip need air. - **Radius:** consistent token set (e.g. 6 / 12 / 20 px). Generation cards use 6; the preset grid tiles use 12; the welcome card uses 20. - **Shadows:** subtle, layered, warm-tinted. Avoid heavy drop- shadows. - **Motion:** purposeful — entrance fades, hover lifts, page transitions. Respect `prefers-reduced-motion`. The "generation in progress → result settles" animation is the one place where motion carries meaning; respect reduced-motion by jumping instead of animating. No bouncing splash animations. No theatrical hero animations. - **States:** every interactive element has hover, focus, active, disabled. Loading uses skeletons that match the eventual generation-card layout, not spinners. Empty states have helpful next-action guidance ("Upload one selfie to begin"). ## 8. Content generation rules - Write **realistic, specific copy**. NO Lorem Ipsum. NO generic placeholders like 'Your tagline here'. - Invent plausible names, preset descriptions, and use-case examples that fit the domain. The Speaker-preset example user "Dr. Asha Devarajan / Co-founder, Calliope Labs" is fictional; use names that feel real without naming a real person. - Tone: warm, direct, free of corporate language. This template is for a user who just needs the photo, not for a brand-marketing team. - Headlines: punchy and concrete. No 'Empower your X' filler. No 'Revolutionize'. No 'Seamless'. No 'AI-powered'. - Body copy: short paragraphs (2-4 sentences). Use lists where appropriate. - Plain language. Avoid jargon — except where the user already speaks the jargon (a designer knows "4K"; a presenter knows "16:9"; a recruiter knows "LinkedIn 1:1"). - Where the app outputs AI-generated content, never label it as "AI says" — let the result speak for itself. Use small uncertainty cues only where epistemic honesty requires them (a borderline likeness-audit result shows a small "we ran a second pass — let us know if this still doesn't look like you" inline note). ## 8a. Seed content (use these specific examples) Anchor every generated copy + sample data point in the concrete content below. Use these names, numbers, and snippets verbatim where helpful, or generate close variants that sit in the same world. **Sample preset library (the six the canvas shows on load):** - **Corporate.** Use case: LinkedIn profile, recruiter inbox, team page. Crop: 1:1 (4096×4096 px). Lighting: warm soft key from camera-left, soft fill from camera-right, faint hair light from above. Wardrobe default: charcoal two-piece with a soft-knit layer underneath for a modern feel; alternate offered as navy suit with open collar; smart-casual blazer over a clean tee; cream knit cardigan for a softer read. Background default: warm graphite seamless paper at a 4-foot distance from the subject. - **Editorial.** Use case: press release, magazine bio, podcast guest tile. Crop: 3:4 (3072×4096 px). Lighting: high-contrast rim light from camera-left with a deep shadow on the right side of the face, faint fill from a reflector. Wardrobe default: black turtleneck or black crew-neck; alternate offered as a cream button-down with rolled sleeves; a tailored leather jacket; a structured wool coat. Background: graphite seamless paper, slightly darker than Corporate. - **Academic.** Use case: faculty page, thesis dust-jacket, conference programme bio. Crop: 3:4 (3072×4096 px). Lighting: warm window light from camera-right, soft fill from camera- left, a faint amber spill suggesting a library or office. Wardrobe default: herringbone blazer over a button-down; alternate offered as a knit blazer over a tee; a wool jumper with a collared shirt underneath; a linen blazer for warmer- climate faculty. Background: a soft-blurred bookshelf suggestion, never a recognisable library. - **Artisan.** Use case: maker portfolio, craft-fair bio, small- studio website. Crop: 4:5 (3277×4096 px). Lighting: warm workshop light from above with a faint side spill from a window, shadow falling naturally on the work surface in front of the subject. Wardrobe default: denim apron over a crew-neck or a flannel; alternate offered as a leather apron; a chambray work shirt; a paint-spattered tee (with the spatter pattern preserved as a real working-clothes detail, not a costume). Background: brick wall with hints of workshop shelving in soft focus. - **Tech.** Use case: engineering team page, conference talk banner, GitHub profile. Crop: 1:1 (4096×4096 px). Lighting: flat clean studio light from above and front, no dramatic shadows. Wardrobe default: clean black tee or a clean grey hoodie (zipper, not pullover); alternate offered as a button- down rolled to the elbow; a knit sweater over a tee; a black turtleneck. Background: clean off-white seamless paper. - **Creative.** Use case: gallery bio, artist portfolio, music EPK. Crop: 4:5 (3277×4096 px). Lighting: coloured-gel mixed light — a soft teal gel from camera-left and a warm amber gel from camera-right, with the face caught in the warmer half. Wardrobe default: an experimental fabric (boucle, mohair, or textured linen) in a saturated colour the user chooses from a small palette; alternate offered as a vintage piece (denim jacket, jumpsuit, leather); a kimono or kurta as a culturally specific option offered to users who indicate they wear one in the source. Background: a muted colour wall (forest green, burnt sienna, or muted indigo). **Bonus preset:** - **Speaker.** Use case: conference deck hero, panel banner, webinar header. Crop: 16:9 (4096×2304 px). Lighting: any of the six preset moods (the user picks). Wardrobe: any of the six preset wardrobes (the user picks). Layout: the user is offset to the right of the frame; the left third is the empty space reserved for the user's name + title, rendered by Nano Banana Pro's legible 4K typography capability. The user types their name and title at the moment of generation. **Sample Speaker-preset input:** - Name: "Asha Devarajan" - Title: "Co-founder, Calliope Labs" **Sample source-photo check responses:** - **proceed** — "Your photo's good. Pick a preset." - **proceed_with_caveat** — "Your photo has a slight backlight — we'll correct for it." (caveat_for_generation: "source is mildly back-lit; lift the face by 1 stop in the generation pass") - **recoverable** — "We need to see your eyes — try a photo without sunglasses." (Or: "Try a clearer photo — this one's a little blurry." Or: "Try a photo where your face fills more of the frame — get a little closer to the camera.") - **block** — "We can't find a face in this photo — try uploading a different selfie." (Or: "We see more than one face — crop to just yours and re-upload.") **Sample likeness-audit pass:** - skin_tone_match_delta_e: 2.7 - skin_tone_match_pass: true - facial_geometry_pass: true - lighting_fairness_pass: true - identity_likeness_score: 0.91 - audit_notes: "Skin tone and facial geometry preserved within tight tolerance; warm key light reads accurately on the user's medium-deep complexion." - needs_regeneration: false **Sample likeness-audit fail (triggers auto-regeneration):** - skin_tone_match_delta_e: 7.4 - skin_tone_match_pass: false - facial_geometry_pass: true - lighting_fairness_pass: true - identity_likeness_score: 0.62 - audit_notes: "Skin tone in output reads about two stops lighter than source; regenerate with elevated identity anchor weight." - needs_regeneration: true - regeneration_reason: "skin tone shift > ΔE 5" **Sample voice copy:** - Welcome line: "One selfie in. Six studio-grade you's out — at 4K, in 30 seconds." - Upload zone empty state: "Drop a selfie here, or tap to pick from your camera roll." - Source-check pass: "Your photo's good. Pick a preset." - Generation in progress: "Lighting the set…" / "Setting the wardrobe…" / "Final touches…" / "Verifying it still looks like you…" - Result settled: "Corporate, 1:1, 4K — ready when you are." - Doesn't-look-like-me tap: "We'll run it again, locking your features tighter." - Regenerated: "Closer match — does this look like you?" - Download button: "Download 4K PNG" - Download all: "Download all six (zip)" - Identity reminder banner under every generation: "Reference photo: yours. AI-generated portrait, SynthID watermarked." - Footer: "Made for the user who needs a studio-grade headshot today. Your photo and your generations are yours. We never train on them." ## 9. Media & assets - **Hero image (welcome screen):** A photographed-looking shot of an empty studio at golden hour — a softbox at frame-left casting warm key light, a wooden stool centre-frame, a graphite seamless backdrop, no people. Generate via Nano Banana 2 (`gemini-3.1-flash-image`) with a prompt emphasising "warm afternoon light, no people in frame, a softbox at frame-left, a wooden stool centre-frame, a graphite seamless backdrop, real worn paper texture on the backdrop, soft shadow under the stool". - **App icon / wordmark:** Set in the display serif. Slightly worn paper texture behind it. No camera icon. Just type. - **Empty-state illustration (no photo yet):** A simple photographed-looking shot of an empty wooden stool in front of a graphite seamless backdrop with a softbox at frame-left, taken at golden hour. The empty seat is the invitation. - **Preset thumbnails:** Each preset card shows a 1:1 lighting/ wardrobe thumbnail rendered with a neutral placeholder face silhouette (no identifiable features) so the user sees the mood before uploading. Generated via Nano Banana 2 per preset. - **Reference style-guide images per preset:** Each preset includes a small style-guide library (lighting reference, wardrobe reference, background reference) that ships in the repo at `/data/presets//lighting.webp`, `/data/presets//wardrobe.webp`, `/data/presets//background.webp` (1024×1024 WebP each, 6 presets × 3 = 18 seed files). Regenerate via Nano Banana 2 (`gemini-3.1-flash-image`) using the lighting / wardrobe / background descriptions from section 8a per preset (e.g. for Corporate lighting: "warm soft key from camera-left, soft fill from camera-right, faint hair light from above, graphite seamless paper backdrop at 4-foot distance, neutral placeholder mannequin silhouette, no real face, photographic, no text"). These are passed as the up-to-13 additional reference images on each `gemini-3-pro-image` call. - **Stock fallbacks:** If image generation for the welcome screen fails, fall back to a hand-curated empty-studio photograph at `/public/samples/empty-studio.jpg` (3:2 WebP, 2048×1365 — ship as a seed asset; recreate via Nano Banana 2 `gemini-3.1-flash-image` with the prompt "photographic empty studio at golden hour, softbox at frame-left casting warm key light onto an empty wooden stool centre-frame, graphite seamless paper backdrop with worn paper texture, soft cast shadow under the stool, no people, no text, no commercial branding, slight imperfection"). Never to a "📸" emoji. - **Generated imagery:** prefer Nano Banana 2 for welcome and preset thumbnails. Prompt for warmth, asymmetry, and slight imperfection — avoid the glossy 'AI render' look. - **Optimisation:** WebP/AVIF for thumbnails, `loading="lazy"`, explicit `width`/`height` to prevent layout shift. The 4K PNG output downloads are not converted — the user gets the original 4K PNG. - **Icons:** `lucide-react` for UI. Use sparingly — never decorative-only. ## 10. Interactivity & states - Every interactive element has hover, focus, active, and disabled states. - Forms validate inline and show specific error messages (not "Invalid input"). "Your name is empty — Speaker preset needs a name to render" is the right shape. - Loading states use skeletons that match the eventual generation card layout, not spinners. Each generation card shows a shimmering placeholder at the correct aspect ratio (1:1, 3:4, 4:5, 16:9) before its image streams in. - Empty states explain the next action with a button whose label fits THIS app's domain: "Upload one selfie to begin", "Pick a preset to generate your first portrait", "Type your name to render the Speaker variant" — never a generic "Add your first item". - Smooth scroll for in-page anchors. - AI-generated content streams in token-by-token where supported (for the source-photo check JSON and the likeness-audit JSON); the image generation streams as a progressive-resolution placeholder until the 4K final lands. - If an AI call fails, show a calm, specific error ("We couldn't reach the model — your photo's still on the canvas. Try again, or pick a different preset.") and offer retry. - The "this doesn't look like me" tap target is small but prominent — it sits beneath every generation card with a fade-in on hover. The regeneration runs immediately, replaces the current card with a skeleton, and lands the new result in ~5 seconds. - Wardrobe and background override drawers slide in from the preset card; tap outside to close. They preserve the user's current selection if reopened. - The compare button is a hold-to-preview: hold to see the source full-size; release to return to the generation. Tap (without hold) opens a side-by-side modal at full resolution. - The download button shows a brief checkmark state on success rather than a toast notification. ## 11. Tech & responsive requirements - **Stack:** React + TypeScript + Tailwind CSS. Functional components + hooks. Use Shadcn UI primitives where appropriate. Image streaming via the Web Streams API. - **Build runtime:** AI Studio Build — full-stack with Cloud Run server-side functions. All Gemini API calls happen server-side; API key lives in Secrets Manager, never in client bundle. The AI Studio Build post-I/O 2026 free 2-app Cloud Run deploy is the intended deploy target. - **Model selection:** explicitly pin `gemini-3-pro-image` for portrait generation and the doesn't-look-like-me regeneration; `gemini-3.5-flash` for source-photo check and likeness audit; `gemini-3.1-flash-image` for welcome / empty-state imagery. Set `thinkingLevel: low` on the Flash calls; omit `thinkingConfig` entirely on the image-generation calls. - **Database:** Firestore (auto-provisioned by AI Studio Build). - **Auth:** Firebase Auth — Google sign-in by default after three guest generations; Apple sign-in next to it (optional, user- configured). - **Storage:** Firebase Storage for source photos (encrypted at rest, 30-day default retention) and 4K PNG outputs (30-day retention). Pre-signed URLs only. Storage requires manual Firebase-console enable before first use. - **Image upload:** the source photo uploads to Firebase Storage client-side, then a server-side function uploads it to the Gemini Developer API Files API and passes the resulting `files/*` resource name (e.g. `files/abc123xyz`) via `fileData.fileUri` to `generateContent`. Do NOT pass Firebase Storage public URLs to `generateContent` — the API does not fetch them server-side. - **Output handling:** the 4K PNG returned by `gemini-3-pro-image` is streamed to the client and simultaneously written to Firebase Storage at a signed URL. Download buttons hit the signed URL rather than re-running the generation. - **Rate limits:** guest mode → 3 generations per IP+fingerprint per 24h, enforced server-side. Signed-in mode → 100 generations per day (effectively unlimited for normal users); operator can raise via env var. - **Mobile-first.** Verify layouts at 375 px (iPhone SE), 768 px (iPad), 1024 px, 1440 px+. The preset grid collapses to a 2×3 on mobile, 3×2 on tablet, 6×1 on desktop ultrawide. - Use `clamp()` for fluid typography. Prefer container queries over media queries for component-level responsiveness. - Use `dvh` / `svh` instead of `vh`. Respect safe-area insets on iOS — the take-a-photo flow on iOS lives in the safe area. - Zero horizontal overflow at any width. Zero layout shift on load. - Persist user data in Firestore. Use real-time listeners on the results strip so each preset's generation lands the moment it completes, in parallel. - Optimistic UI on writes; reconcile on response. - The download button uses the browser's File System Access API where available (Chrome/Edge desktop) for a native save dialog; falls back to a standard download for Safari and mobile. ## 12. Accessibility (WCAG 2.2 AA) - Semantic HTML — `header`, `nav`, `main`, `section`, `article`, `footer`. - All interactive controls reachable by keyboard with a visible focus ring. - Color contrast ≥ 4.5:1 for body, 3:1 for large text and UI components. - All images have meaningful `alt` text. Each generation's `alt` describes the preset, crop ratio, and that it is AI-generated with the user's source photo as the reference ("AI-generated Corporate-preset portrait at 4K, 1:1 crop, generated from your uploaded reference photo"). - The source photo's `alt` reads "your uploaded reference photo". - Form fields have associated `