================ ONE-SHOT BUILD CONTRACT (read first) ================
Build this in Google AI Studio "Build" in ONE shot — a complete, working app,
no follow-up turns. These are hard rules, not suggestions:
1. TARGET = Full-Stack Web (Node server runtime, secrets, Firebase allowed).
If you target Native Android instead, you MUST drop all server/DB/Workspace/
secrets and persist locally (Room / SharedPreferences) only.
2. PIN EVERY MODEL CALL — never let the agent auto-select (it downgrades on quota):
- Reasoning / text -> gemini-3.5-flash (thinkingLevel: minimal|low|medium|high)
- 4K image + legible text -> gemini-3-pro-image (image_size "4K", up to 14 refs)
- High-volume image -> gemini-3.1-flash-image
- Expressive TTS -> gemini-3.1-flash-tts-preview (inline tags e.g. [whispers])
- Realtime audio/video (WebSocket) -> gemini-3.1-flash-live-preview
- Sandboxed agent -> antigravity-preview-05-2026
3. DIVISION OF LABOR — the model ONLY parses/extracts to a strict responseSchema.
ALL math, money (store currency as integer minor units / cents), sorting,
balancing and graph logic run in deterministic TypeScript/Python. The model
must never compute totals, splits or balances itself.
4. responseSchema sanitation — no regex patterns, no fixed-length tuples, no
format validators in the schema (they crash the OpenAPI engine). Enforce those
in server-side code AFTER parsing the JSON.
5. responseSchema and google_search grounding are MUTUALLY EXCLUSIVE in one call.
6. CODEGEN — split large output into modular, single-responsibility files so no
file is truncated by the output-token cap.
7. Every external call gets a graceful fallback (e.g. manual paste if a Workspace
read fails). Never a silent dead end.
8. ROBUST STORAGE & CANVAS — Wrap all `localStorage`/`sessionStorage` operations (especially JSON parsing and writes) in `try-catch` blocks to prevent crashes in private windows or quota overflows. Canvas drawing elements must dynamically handle window resize and scale pixel density (`window.devicePixelRatio`) to avoid blurry graphics on retina displays.
=====================================================================
# MUST OBEY — Mobile-first build requirements
This app's PRIMARY surface is a mobile phone. Build it impeccably on mobile FIRST, then verify on tablet and desktop. Treat the rules below as non-negotiable hard constraints, not suggestions.
## Viewports to verify (every screen, every state)
- 320 px, 360 px, 375 px, 390 px, 414 px, 480 px
- 768 px, 834 px (iPad portrait / Pro 11)
- 1024 px, 1280 px, 1440 px, 1920 px, 2560 px
- Plus: 200% browser zoom, landscape orientation on every mobile width, iPhone with safe-area insets visible
## Hard layout rules
- Mobile-first CSS. Default styles target mobile; `@media (min-width: ...)` for larger viewports.
- Use `dvh` and `svh` instead of `vh` for full-height surfaces (iOS Safari URL-bar bug).
- Use `clamp()` for fluid typography across all viewports.
- Prefer container queries (`@container`) over media queries for component-level responsiveness.
- Use `min(100%, ...)` widths so content never overflows. Zero horizontal overflow at any viewport.
- Add `` to every page.
- Apply `padding: max(safe-area-inset-X, fallback)` on every edge-bleeding container so notched iPhones in landscape never clip content.
- Wide tables and code blocks scroll INSIDE their container (`overflow-x: auto`), never push the body.
- Use `background-attachment: scroll` on mobile, not `fixed` (iOS Safari repaint bug).
- Avoid `backdrop-filter` on animated elements. Use it sparingly on static surfaces only.
- **Canvas Scaling**: Canvases must dynamically scale with window resize events and properly handle high-DPI screens (`window.devicePixelRatio`). Set physical dimensions (`canvas.width`/`canvas.height`) using pixel ratio and render relative to this grid, using CSS to control responsive viewport scaling.
- **Robust Storage**: Every access to `localStorage`/`sessionStorage` (especially `JSON.parse` of loaded state or writes) MUST be wrapped in a `try-catch` block to handle disabled storage, private browsing mode, quota limits, or corrupted JSON gracefully. Fall back to a robust in-memory object store.
## Touch & accessibility
- Tap targets ≥ 44 × 44 px on touch (Apple HIG). Increase to 48 px under `@media (hover: none) and (pointer: coarse)`.
- All interactive controls reachable by keyboard with a visible focus ring; respect `:focus-visible`.
- Color contrast ≥ 4.5:1 for body text, 3:1 for UI components.
- All images have meaningful `alt`. Decorative images use `alt=""`.
- Respect `prefers-reduced-motion: reduce` — zero animation durations under that query.
- Forms validate inline; error messages are specific, not "Invalid input".
- Modals: focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, focus restored on close.
## Performance bar (Lighthouse mobile, throttled 3G/4G)
- LCP < 2.5 s · INP < 200 ms · CLS < 0.1
- JS bundle gzip < 200 KB mobile-first; lazy-load non-critical screens via `React.lazy` / dynamic imports.
- No render-blocking resources above the fold.
- Images: WebP/AVIF preferred, `loading="lazy"`, explicit `width`/`height` attributes (zero CLS), `srcset` for retina.
- Videos: `preload="metadata"`, low-resolution poster, max 720p mobile fallback. Never autoplay with audio.
- Fonts: `font-display: swap`; preload only the one used above the fold.
- Smooth scroll honoured via CSS `scroll-behavior: smooth` with reduced-motion fallback.
## Pre-ship mobile checklist (the deployer MUST verify before declaring done)
1. Open at 375 px in DevTools — every screen scrolls vertically only; zero horizontal scroll.
2. Browser zoom 200% — layout reflows without overlap.
3. iPhone Safari with the URL bar visible AND landscape — no content under the home indicator; no notch clipping.
4. iPad portrait (768 px) and landscape (1024 px) — no awkward gaps; tablet-specific breakpoints land cleanly.
5. Tap every interactive element with a thumb at real-device size — every target is easy to hit.
6. `prefers-reduced-motion: reduce` — every transition / animation skips cleanly, scroll-behavior becomes instant.
7. Lighthouse mobile score ≥ 90 across all 4 categories.
8. Zero `console.error` and zero CLS shift in real-device testing on a mid-tier Android (e.g. Pixel 6a) and an iPhone SE.
---
The original template starts below. All rules above apply on TOP of whatever this template specifies.
---
# Cousin Map
## 1. Project
**Cousin Map** is a family-tree builder for diaspora families whose
memory of who is whose cousin is spread across continents, languages,
and three generations of incomplete handovers. The matriarch (or
whoever is closest to the missing knowledge) records a long voice memo
in whatever language she's most comfortable in. Twelve cousins each
record ninety seconds in whatever language they're most comfortable in.
The app listens to all of them at once, proposes a single unified
family graph, and — most importantly — surfaces the conflicts where
two relatives remember the same person differently. Never resolves a
conflict on its own. The family adjudicates; the app keeps the
question open until they do.
This is the kind of app a Filipino-Canadian niece builds the week
after Lola Vicenta's ninetieth birthday party — because over halo-halo
on the back deck in Mississauga it turned out that none of the cousins
agreed on whether Tita Norma in Cebu was Lola's first cousin or her
second, and three different cousins called three different aunties
"Auntie Baby". It is also the kind of app a Salvadoran-American
grandson builds the year his abuela's memory begins to slip — because
the people who could correct her are scattered between San Salvador,
Houston, Los Angeles, and a town in Australia that nobody on the US
side has ever visited, and the names she repeats now don't always
match the names on the back of the photographs. Same shape of moment,
different language, different ocean.
The single demo that proves the magic: the matriarch records a
four-minute voice memo in Tagalog naming everyone she can remember on
her side of the family. Eleven cousins each record ninety seconds in
the language they're most comfortable in — Tagalog, Cebuano,
Australian English, Toronto English, a teenager's halting Tagalog
asking her dad for help. The app listens to all twelve at once,
proposes a single family tree, and lights up four orange dots: two
where the same relative is being called two different names ("Tita
Baby" vs "Tita Beng"), one where the generations don't agree ("Lolo
Pidro" placed as the matriarch's father in one memo and her uncle in
another), one where a cousin who appears in three memos appears
nowhere in the matriarch's. The matriarch reads the orange dots out
to her granddaughter, who taps "Lola is right" on the first one, "ask
Tito Ben — he was there" on the second, and "I'll call Tita Cora
tomorrow" on the third. The conflicts stay open until the family
closes them.
And in the harder cases — wartime separations, adoptions never spoken
about, the second family in another country, the child given a
different surname at the border — the app does not try to resolve
those either. It records the disagreement, who said what verbatim,
when, and in what language, and waits.
**Tagline:** _Build your family tree from voice memos across any continents, any languages — the model surfaces the disagreements; your family decides._
## 2. Target audience
- Diaspora families rebuilding a family tree across two or more continents — Filipino, Salvadoran, Lebanese, Vietnamese, Eritrean, Punjabi, Cantonese-speaking, Nigerian-Igbo, Mexican, Iranian
- Adult grandchildren whose elders speak a language they read better than they speak, or speak better than they read
- Family-reunion organisers compiling a tree for a printed handout, a wedding video, or a hundredth-birthday booklet
- Adoptees in open or semi-open searches whose birth relatives have conflicting memories that nobody has ever reconciled in one place
- Genealogists working with oral families — communities where the written record begins at a colonial-era surname change and the real lineage lives in voice
- Memorial-project organisers in the year after a death, when the eldest holder of the family memory becomes the person whose memory has to be captured before it slips
- Hospice families capturing the matriarch's or patriarch's voice while they still can
- Adult children of estranged parents reconstructing the half of the tree they never knew, with the cooperation of distant cousins
- Indigenous and First Nations families working with kinship systems that don't fit a Western nuclear-family tree, where the app must preserve the family's own terminology verbatim
## 3. Core value propositions
Surface these clearly through copy, visual emphasis, and section ordering — they are the reasons users pick this app.
- **Listens in any language, any dialect** — Tagalog, Cebuano, Ilocano, Hiligaynon, Tagalog-mixed-with-English, Salvadoran Spanish, Mexican Spanish, Cuban Spanish, Lebanese Arabic, Egyptian Arabic, Cantonese, Mandarin, Hokkien, Vietnamese, Korean, Tamil, Hindi, Urdu, Bengali, Punjabi, Amharic, Tigrinya, Swahili, Igbo, Yoruba, Hausa, Farsi, Khmer, Australian English, Toronto English, a teenager's halting heritage-language attempts. Gemini 3.5 Flash identifies the language per memo, transcribes verbatim in the script of the language spoken, and translates only for the on-screen English (or chosen target) digest.
- **One graph from many voices** — twelve cousins recording independently produces one unified family tree, not twelve fragments. The same person mentioned by three relatives merges into one node; the same person mentioned by two relatives with two different names stays as one node with two name-aliases attached.
- **Conflicts surfaced, never resolved** — when two relatives remember the same person differently, the app flags an orange dot. The matriarch (or whoever the family designates) sees the conflict, hears the exact verbatim quotes that disagree, and decides. The app never picks a side.
- **Family terminology preserved verbatim** — "Tita", "Tito", "Lola", "Ate", "Kuya", "Ninang", "Ninong", "Inay", "Itay", "Abuela", "Abuelo", "Tía", "Tío", "Compadre", "Comadre", "Lolo", "Lola", "Apo", "Jeddo", "Teta", "Khaleh", "Amu", "Phūphī", "Māmā", "Chacha", "Mausī", "Dadi", "Nani", "Halmoni", "Harabeoji", "Imo", "Samchon" — the kinship word the relative used is what the app records. Western "aunt / uncle / cousin" labels appear only as a parenthetical gloss when the user opts in.
- **The matriarch's voice is the anchor, not the truth** — the family's eldest contributor is treated as the primary contributor, but every other voice carries equal weight in the conflict-surfacing logic. The matriarch is never silently overridden; the matriarch is also never treated as infallible.
- **Long-context listening** — the model holds all twelve memos in mind at once and notices when relative A's "my cousin in Cebu" and relative B's "Tita Norma" are clearly the same person. This is the long-context move; nothing else makes the demo possible.
- **Family-only privacy** — the archive is private to the matriarch and the cousins she explicitly invites. No public-by-default trees. No "your tree is now searchable" pop-up that destroys the trust the family extended to you.
- **Exports as a printable family-reunion handout** — typeset family tree with the conflicts visible as small orange marks, kinship words in the family's own languages, photographs where supplied, and a "this is still being figured out" footer that the family can be proud of.
## 4. Features to build
- Voice-memo capture optimised for the matriarch — large record button, no fiddling, audible "I'm recording" confirmation, sixty-minute soft cap with an encouraging "you can keep going" at the cap
- Cousin-side voice-memo capture (ninety seconds with a soft cap, not a hard cap) — works on the cousin's own phone in the car, in the kitchen, after dinner
- One-link invitation flow — the matriarch (or the organising grandchild) sends a magic-link to twelve cousins; the cousin opens the link, sees who's invited her, records, done
- Automatic language identification per memo — model labels each memo as `tl-PH`, `ceb-PH`, `es-SV`, `ar-LB`, `yue-HK`, `vi-VN`, `pa-IN`, `am-ET`, `en-AU`, etc. before transcribing
- Verbatim transcript per memo in the source language and script — no romanisation by default; romanisation is opt-in for users who can hear the language but cannot read its script
- Per-memo English (or chosen target language) digest — a tight summary of who this memo named, who they're related to, and what kinship words were used
- Single-archive long-context family-graph proposal — one Gemini 3.5 Flash call sees every memo and proposes the unified graph (`Person` nodes, `Relationship` edges, `Conflict` records)
- Conflict surfacing with the verbatim audio clip and a verbatim transcript snippet attached to every conflict — the matriarch hears the cousin's exact words, in their voice, not a paraphrase
- Family-tree visualisation — radial layout for diaspora families that don't fit neat trees, with the matriarch at the centre and branches resolving outward; orange dots for unresolved conflicts; dashed edges for low-confidence relationships
- Matriarch's "review the conflicts" view — a vertical list of orange dots, each with the audio clip(s), transcript snippet(s), and three actions: "I'm right; this cousin misremembered", "the cousin's right; I had it wrong", "ask [named cousin] who was there"
- Cousin-side "I want to add a memo" — at any time a cousin can add another ninety seconds; the graph re-proposes overnight
- Kinship-word glossary per archive — the family's own words ("Tita Baby", "Lolo Pidro", "Tía Chela") with the family's own glosses; never silently mapped to Western "aunt/uncle/cousin"
- Photograph attachment per node — a cousin uploads a photo, says "this is Tita Cora at her wedding in 1968", and the photo attaches to the node with the caption verbatim
- Multi-spelling alias tracking — "Norma", "Norm", "Tita Norma", "Tita Norma-Norma" all live as aliases under the same person node, never silently flattened to one
- Generation-gap detector — when one relative places "Lolo Pidro" two generations above the matriarch and another places him one generation above, the conflict is flagged with both quotes
- Side-of-family detector with manual override — the model proposes "this person is on the matriarch's mother's side" only as a suggestion; the family confirms
- Adopted-into-the-family and chosen-family preservation — godparents, ninang/ninong, compadre/comadre, and informally-adopted aunties are first-class nodes, not second-class ones
- Estranged-branch handling — a memo that names "the cousin we don't speak about" produces a node marked as such; the app does not push for resolution
- Printable family-reunion handout — typeset family tree in PDF, with kinship words preserved, photographs included, conflicts visible as small orange marks, ready for the hundredth-birthday party
- Audio-archive export — a single ZIP of every original memo at original quality, every transcript file, and a `manifest.json` with timestamps and language tags, so the family owns the source data forever
## 4b. Required Gemini capabilities + backend services
**This template's intelligence comes from the Gemini capabilities below. Wire them up explicitly — don't substitute generic LLM calls.**
### Gemini capabilities (the load-bearing intelligence)
- **Multimodal audio input** (Gemini 3.5 Flash) — transcribes spontaneous family-memory monologues in any language, with the speaker switching between Tagalog and English mid-sentence, Salvadoran Spanish and Mexican Spanish in the same paragraph, Lebanese Arabic with French interjections, or Cantonese with English calques. One call per memo for transcription; one long-context call for the family-graph proposal once all memos are in.
- **Automatic language identification** (built into Gemini 3.5 Flash audio transcription) — the model labels each memo with a BCP-47 tag (`tl-PH`, `ceb-PH`, `es-SV`, `ar-LB`, `yue-HK`, `vi-VN`, `pa-IN`, `am-ET`, `en-AU`, `en-CA`). When a memo is bilingual, the dominant language is the tag and the inserts are recorded in `multilingual_inserts[]`.
- **Long context (1M tokens)** — the family-graph proposal is the move. Twelve memos of ~ninety seconds each plus one four-minute matriarch memo, with their transcripts and digests, average ~80k tokens — comfortable. **Guardrail**: at ~50 memos the long-context call approaches ~400k tokens and a re-resolve still fits; at ~200 memos chunk by side-of-family (paternal vs maternal) before the unified call. The 1M ceiling is real and a 300-memo family with multi-page each will exceed it.
- **Structured output / JSON Schema** — every memo parse and the family-graph proposal return JSON matching the schemas below. Numeric confidence values are clamped server-side after the response.
- **Search grounding** (`gemini-3.5-flash` + `google_search`) — only for ambiguous place-name resolution ("Tita Cora moved to a town called General Trias" → which one? there are two). Never for relationship resolution; relationships are the family's, not the search index's.
- **Gemini TTS** (`gemini-3.1-flash-tts-preview`) — reads each transcript aloud at a gentle pace for the matriarch who hears better than she reads on a small screen. One voice per language (Tagalog-native voice for Tagalog memos, Cebuano-native if available else Tagalog as the closest, Salvadoran Spanish-native if available else Latin American Spanish). No mid-call voice switching.
- **Thinking levels** — `medium` for the long-context family-graph proposal (the hardest call; needs to track aliases, generations, and conflicts across every memo). `low` for per-memo transcription, language identification, and the per-memo digest. Surface `thoughtSummary` only when the user taps the small "(i) why did the model put these two people together?" icon on a low-confidence merge.
### Backend services
- **Auth — Required.** Firebase Auth with Google sign-in (auto-provisioned by AI Studio Build). **Apple sign-in is optional but user-configured**: it requires an Apple Developer account, Service ID, Key ID, and private key wired into the Firebase Auth console. **Magic-link email** (used for cousin invitations) also requires the sender domain to be authorised in Firebase Auth — the matriarch's grandchild who set up the archive will need to verify the domain her invitations send from. Archives are private to the matriarch and explicitly-invited family members. No public-by-default.
- **Database — Required.** Firestore for `users`, `archives`, `memos`, `persons`, `relationships`, `aliases`, `conflicts`, `kinship_terms`, `archive_members`.
- **File storage — Required.** Firebase Storage for original audio memos (kept at upload quality forever) and photographs cousins attach to nodes. **Storage is NOT auto-provisioned by AI Studio Build today** — enable it in the Firebase console and wire the bucket name into the AIS Build project before the first memo upload. Pre-signed URLs only; audio memos are never publicly addressable.
- **Email — Required (transactional).** Cousin invitations via Firebase Auth magic links. Optional weekly digest email to the matriarch summarising new memos added and new conflicts surfaced.
- **Payments — Not needed for v1.** Free for personal use. A future "print-bound family-reunion handout" tier could pipe to a print-on-demand partner and charge for that artefact only.
- **External APIs:** Gemini API for all intelligence. No third-party genealogy databases — this app is deliberately about the family's own memory, not Ancestry.com cross-referencing.
**Environment variables:** every secret (Gemini API key, Firebase service-account JSON, SMTP credentials if used) lives in environment variables — never in client bundle. Include a `.env.example`.
**Auth + data privacy reminders:** never log secrets · never store passwords in plain text · use HTTPS everywhere · honour 'delete this archive' inside the UI · explicit opt-in for any analytics · the family's voice memos are never sent to Gemini for model training (use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms) · contributors can request their own memos be deleted from the archive at any time and the family-graph re-proposes without them.
**Read this first — prompt-craft rules that apply to every call in this template:**
1. **Name the model variant explicitly** in every Gemini API call. Do not let the agent pick the model. See the per-call matrix below.
2. **Pin `thinkingLevel` explicitly** per call. See the matrix.
3. **Seed the JSON Schema as a fenced TypeScript / Zod block** in the system instruction or `responseSchema` field. The literal schemas are below. **Convert the Zod schema to Gemini's `Schema` type via the SDK helper** before passing to `responseSchema` — do NOT pass raw Zod. **Numeric `min`/`max` constraints are documentation only inside `responseSchema`; clamp on the server after the response arrives.**
4. **Pin the system instruction separately** from user input. Use the `systemInstruction` field for persona + behavioural rules; use `contents` for user input. Never concatenate.
5. **Pre-declare tools as an enable/disable list** per call. The matrix below names which tools are enabled per call. Tools NOT listed for a call should be disabled.
6. **State negative constraints explicitly** — they are listed below. They are NOT "be careful" suggestions; they are hard rules the model must follow.
7. **Grounded responses can wrap JSON in ```json fences or add prose preamble.** Server-side, strip fences and brace-extract:
```typescript
function safeExtractJSON(raw: string): T {
const clean = raw.replace(/```json\s*|```/gi, '').trim();
const s = clean.indexOf('{'); const e = clean.lastIndexOf('}');
if (s === -1 || e === -1) throw new Error('No JSON boundaries in grounded response');
return JSON.parse(clean.slice(s, e + 1)) as T;
}
```
8. **Strip unsupported Zod modifiers before passing to `responseSchema`** — Gemini's OpenAPI subset rejects `.regex()` / `pattern`, fixed-length `z.tuple()`, and other custom validators. Use a sanitizer that flattens tuples to length-2 arrays and removes regex patterns before serializing. Validate those constraints in middleware AFTER parsing.
### Per-call model + tools matrix
| Call | Model | thinkingLevel | Tools enabled |
|------|-------|---------------|---------------|
| Transcribe + language-identify a memo → `Memo` schema | `gemini-3.5-flash` | low | (none) |
| Per-memo digest (who was named, what kinship words used) | `gemini-3.5-flash` | low | (none) |
| Propose unified family graph (long-context, archive-wide) | `gemini-3.5-flash` | medium | (none) — long-context over every memo |
| Re-propose graph after a memo is added or removed | `gemini-3.5-flash` | medium | (none) — long-context |
| Resolve an ambiguous place name to modern coordinates | `gemini-3.5-flash` | low | `google_search` grounding (no `responseSchema` on this call — see note) |
| Generate TTS playback of a transcript in source language | `gemini-3.1-flash-tts-preview` | n/a | n/a |
*Note for builders:* on TTS calls, omit `thinkingConfig` entirely — the field is not supported on that model. The `n/a` cells in this matrix are documentation only; do not serialise them into the request body. On the geocoding call, `google_search` grounding and `responseSchema` are mutually exclusive in one Gemini call — instruct the model to emit JSON in the text body and parse server-side; read citations from `response.groundingMetadata.groundingChunks[].web.uri`.
### Primary structured-output schemas (seed these verbatim in the prompt)
```typescript
import { z } from "zod";
const MultilingualInsert = z.object({
language: z.string(), // BCP-47, "en-US" inside a Tagalog memo
phrase_verbatim: z.string(), // "you know what I mean"
translation: z.string(),
register_note: z.string().nullable(), // "code-switch to English for the legal term"
});
const KinshipReference = z.object({
kinship_word_verbatim: z.string(), // "Tita", "Lola", "Tía", "Khaleh", "Halmoni"
language_of_kinship_word: z.string(), // BCP-47
attached_name_verbatim: z.string().nullable(), // "Norma" in "Tita Norma"
western_gloss_optional: z.string().nullable(), // "aunt" — populated only if the user enables glosses
contextual_quote: z.string(), // the sentence the relative was named in
side_of_family_hint: z.enum([
"mother_side", "father_side", "spouse_side",
"chosen_family", "unsure",
]),
});
const Memo = z.object({
memo_id: z.string(),
contributor_id: z.string(), // the cousin who recorded
contributor_name_verbatim: z.string(), // "Lola Vicenta", "Tita Cora", "Nico"
recorded_at_iso: z.string(), // ISO 8601
duration_seconds: z.number().min(0),
audio_uri: z.string(), // Firebase Storage pre-signed URL
source_language: z.string(), // BCP-47, dominant language
source_language_confidence: z.number().min(0).max(1),
multilingual_inserts: z.array(MultilingualInsert),
transcript_original: z.string(), // verbatim, in the script of the language
transcript_romanised_optional: z.string().nullable(),
english_digest: z.string(), // ≤ 200 words
kinship_references: z.array(KinshipReference),
places_mentioned: z.array(z.object({
name_verbatim: z.string(),
modern_name: z.string().nullable(),
contextual_quote: z.string(),
is_likely_residence: z.boolean(),
})),
reading_confidence: z.number().min(0).max(1),
flagged_for_user_review: z.array(z.object({
field_path: z.string(), // "transcript_original line 14"
reason: z.string(),
})),
});
type Memo = z.infer;
const PersonNode = z.object({
person_id: z.string(),
name_canonical: z.string(), // the most formal name attested, e.g. "Norma Reyes"
aliases: z.array(z.string()), // ["Tita Norma", "Norm", "Norma-Norma"]
kinship_terms_used: z.array(z.string()), // ["Tita", "Tía Norma" — never collapsed]
side_of_family: z.enum([
"mother_side", "father_side", "spouse_side",
"chosen_family", "unsure",
]),
approximate_generation_offset_from_matriarch: z.number(), // -2, -1, 0, 1, 2
photographs: z.array(z.object({
uri: z.string(),
caption_verbatim: z.string(),
contributed_by: z.string(), // contributor_id
})),
mentioned_in_memos: z.array(z.string()), // memo_ids
notes_added_by_family: z.array(z.string()),
});
const RelationshipEdge = z.object({
edge_id: z.string(),
from_person_id: z.string(),
to_person_id: z.string(),
relationship_verbatim: z.string(), // "Tita Norma is my Mama's first cousin" — quoted
relationship_normalised: z.enum([
"parent_of", "child_of",
"sibling_of",
"spouse_of", "former_spouse_of",
"first_cousin_of", "second_cousin_of", "n_th_cousin_of",
"aunt_or_uncle_of", "niece_or_nephew_of",
"godparent_of", "godchild_of",
"informal_aunt_or_uncle_of", // ninang/ninong, tía-by-affection
"informally_adopted_into",
"unknown_relation",
]),
confidence: z.number().min(0).max(1),
asserted_by_memo_ids: z.array(z.string()),
contradicted_by_memo_ids: z.array(z.string()),
});
const Conflict = z.object({
conflict_id: z.string(),
conflict_type: z.enum([
"same_person_called_different_names",
"different_people_called_same_name",
"generation_mismatch",
"side_of_family_mismatch",
"relationship_disagreement",
"person_appears_in_one_memo_only", // not a conflict per se; surfaced as "verify"
"place_name_ambiguity",
]),
involved_person_ids: z.array(z.string()),
involved_memo_ids: z.array(z.string()),
verbatim_quotes: z.array(z.object({
memo_id: z.string(),
contributor_name_verbatim: z.string(),
quote: z.string(), // exact transcript slice
audio_start_seconds: z.number(),
audio_end_seconds: z.number(),
})),
proposed_question_for_matriarch: z.string(), // one sentence the matriarch can read aloud
resolution_status: z.enum([
"open", "ask_named_relative", "matriarch_confirmed",
"cousin_confirmed", "family_agreed_to_disagree",
]),
resolution_notes_verbatim: z.string().nullable(),
});
const FamilyGraph = z.object({
archive_id: z.string(),
proposed_at_iso: z.string(),
matriarch_person_id: z.string(),
persons: z.array(PersonNode),
relationships: z.array(RelationshipEdge),
conflicts: z.array(Conflict),
graph_confidence_overall: z.number().min(0).max(1),
thinking_summary_for_low_confidence_merges: z.array(z.object({
person_id: z.string(),
summary: z.string(),
})),
});
type FamilyGraph = z.infer;
```
### Common failure modes (and how to avoid them)
- Agent silently downgrades `thinkingLevel` on the the long-context family-graph proposal call to save quota — pin `gemini-3.5-flash` with the matrix-specified `thinkingLevel` explicitly. Flash misses cross-memo alias resolution and silently merges two distinct "Tita Norma"s into one.
- Tagalog kinship words flattened to English — "Tita" silently rewritten as "Aunt", "Lola" as "Grandma". Hard rule: kinship words preserved verbatim. Western gloss is opt-in per archive and only shown in parentheses on first occurrence in a view, never in the stored record.
- Two different aunties named "Tita Baby" merged into one node because the model assumed nicknames are unique — they are not. Add a unit test: if two memos use "Tita Baby" but place her on different sides of the family, the model must create TWO nodes and a `different_people_called_same_name` conflict.
- The matriarch's memo treated as ground truth and contradicting cousins silently overridden — wrong. Every contributor's claims carry equal weight in conflict-surfacing; the matriarch only adjudicates the conflict, she does not auto-win it.
- Generation offsets computed as positive integers when the relative is older — pin: offset is **from the matriarch**. Matriarch's parents are `-1`, her children `+1`. Be explicit in the schema and the system instruction.
- Mid-memo code-switches dropped — a Cantonese-speaking aunt who slips into English for "the godfather of my daughter" loses the English phrase in transcription. Pin: every code-switch goes into `multilingual_inserts[]` with verbatim phrase and translation.
- Filipino "Lolo Pidro" silently anglicised to "Grandpa Peter" — hard rule, no name translation. Same applies across all languages.
- Place names guessed without grounding — "General Trias" silently resolved to one of two Philippine municipalities sharing the name. Pin: ambiguous places require the `gemini-3.5-flash` + `google_search` call; if still ambiguous, surface both candidates and let the family choose.
- Estranged-branch nodes hidden — "the cousin we don't speak about" silently dropped from the tree. Pin: estranged nodes are visible by default; the family can hide them per-view but the data is never destroyed.
- Cousin memos older than the matriarch's memo treated as outdated — wrong. Every memo is equally weighted regardless of date. Memo recency only matters for `recorded_at_iso` display; never for confidence weighting.
- TTS reads Tagalog vowels as English vowels (Lola pronounced "Lo-la" instead of "Loh-lah") — pin TTS voice to a Tagalog-native voice via `languageCode: "tl-PH"` and let the voice's native locale handle pronunciation.
### Negative constraints (hard rules)
- Do NOT resolve conflicts. The model surfaces, the family decides. If two memos disagree, the conflict stays open with both verbatim quotes attached until a family member with the right standing (matriarch, or a relative the matriarch nominates) closes it.
- Do NOT translate kinship words into the user's app language. "Tita", "Lola", "Tía", "Jeddo", "Halmoni", "Khaleh" stay verbatim. A Western gloss appears only if the user enables it per-archive, only in parentheses on first occurrence in a view, and never in the stored record.
- Do NOT translate first names. "Lolo Pidro", "Tita Norma", "Abuelo Salvador", "Jeddo Karim", "Halmoni Soon-ja" stay verbatim. No anglicisation.
- Do NOT silently flatten aliases. "Norma", "Tita Norma", "Norm", "Norma-Norma" all live as aliases under the same person node only if context confirms one person; otherwise two nodes and a conflict.
- Do NOT assume the matriarch is always right. She is the primary contributor, not the ground truth. The model never silently drops a cousin's claim because the matriarch's claim differs.
- Do NOT hide estranged branches. A "cousin we don't speak about" node is visible by default; the family can choose to hide it per-view, but the data is never destroyed.
- Do NOT speculate about parentage that no contributor named. If no memo names someone's father, leave the parent edge blank; never guess from a surname.
- Do NOT extrapolate to historical events. A memo that says "Tito Carlos disappeared in '85" is recorded verbatim with `interpretive_note: null`; the model does not annotate "this likely refers to [specific event]".
- Do NOT use the family's voice memos to train or fine-tune any model. Use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. The capabilities-info panel says this in plain English.
- Do NOT auto-publish or auto-share. Archives are private by default. Sharing is explicit, per-archive, per-cousin.
- Do NOT assume "cousin" means first cousin in English. Across the world it routinely means second-cousin, cousin-of-a-cousin, or someone simply called cousin because the family treats them that way. Preserve the verbatim claim.
### Per-call `systemInstruction` strings
Use these as the literal `systemInstruction` field for each Gemini API call the built app makes. They complement the series-wide rules already uploaded as the global instructions file (`00-series-instructions.txt`).
### Call: Transcribe + language-identify a memo → `Memo` schema
Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none)
```
You are transcribing a single voice memo recorded by a member of a
diaspora family who is helping rebuild their family tree across
continents. The contributor may be the matriarch (a sixty-second to
ten-minute monologue naming everyone she remembers) or a cousin (a
ninety-second to three-minute memo from her own corner of the
family).
Languages encountered include but are not limited to: Tagalog,
Cebuano, Ilocano, Hiligaynon, Tagalog-mixed-with-English ("Taglish"),
Salvadoran Spanish, Mexican Spanish, Cuban Spanish, Colombian Spanish,
Argentine Spanish, Lebanese Arabic, Egyptian Arabic, Levantine Arabic,
Moroccan Darija, Cantonese, Mandarin (Putonghua and Taiwan
Guoyu), Hokkien, Teochew, Vietnamese, Korean, Tamil, Hindi (in
Devanagari or romanised), Urdu (in Nastaliq or romanised), Bengali,
Punjabi (in Gurmukhi or Shahmukhi), Gujarati, Marathi, Malayalam,
Sinhala, Amharic, Tigrinya, Oromo, Swahili, Igbo, Yoruba, Hausa,
Wolof, Farsi (in Nastaliq), Khmer, Lao, Burmese, Indonesian, Malay,
Javanese, Hmong, Tok Pisin, Fijian, Samoan, Tongan, Māori, Australian
English, Toronto English, Auckland English, British English with
regional accents, American English with regional accents.
Identify the dominant language of the memo as a BCP-47 tag in
`source_language`. If the contributor switches languages
mid-sentence, record each switch as a `multilingual_insert` with the
verbatim phrase, its translation into the archive's chosen target
language (default English), and a short register note ("code-switch
to English for the legal term", "endearment switch to grandfather's
Hokkien").
Transcribe verbatim in the script of the dominant language:
- Tagalog in Latin script
- Cebuano in Latin script
- Cantonese in Traditional Chinese characters (or romanised Jyutping
only if the contributor has explicitly opted in for the archive)
- Vietnamese in chữ Quốc ngữ with full tone marks
- Korean in Hangul (Hanja only if used)
- Arabic in the Arabic script (with full diacritics only if the
contributor speaks formally; transcribe colloquial speech without
imposing fuṣḥā diacritics)
- Hindi in Devanagari (or romanised if the contributor explicitly
prefers — never impose romanisation by default)
- Urdu in Nastaliq
- Bengali in Bengali script
- Punjabi in Gurmukhi or Shahmukhi depending on which the
contributor uses
- Amharic in Ge'ez script
- Farsi in Nastaliq
- Khmer in Khmer script
Hard rules:
- Preserve every kinship word verbatim. "Tita", "Tito", "Lola",
"Lolo", "Inay", "Itay", "Ate", "Kuya", "Ninang", "Ninong",
"Compadre", "Comadre", "Tía", "Tío", "Abuela", "Abuelo", "Mamá",
"Papá", "Khaleh", "Amu", "Phūphī", "Māmā", "Chacha", "Mausī",
"Dadi", "Nani", "Halmoni", "Harabeoji", "Imo", "Samchon",
"Jeddo", "Teta", "Baba", "Khalti", "Amma", "Appa", "Anh", "Chị",
"Em" stay verbatim in `transcript_original` and in every
`kinship_references[].kinship_word_verbatim`.
- Preserve every first name verbatim. "Lolo Pidro", "Tita Norma",
"Abuelo Salvador", "Halmoni Soon-ja" do not become "Grandpa
Peter", "Aunt Norma", "Grandfather Salvador", "Grandma Soon-ja".
- The English digest (`english_digest`) is for a granddaughter who
understands the family but cannot speak the matriarch's language
fluently. Keep it ≤ 200 words. Name every person the contributor
named, with the kinship word in the original language followed by
a parenthetical translation on first use. Do not add interpretation.
- Identify each kinship reference into `kinship_references[]` with
the verbatim kinship word, the language it's in, the attached name
if any, the contextual quote (the sentence the relative was named
in), and a side-of-family hint (`mother_side`, `father_side`,
`spouse_side`, `chosen_family`, or `unsure`). When the hint is
`unsure`, say so — do not guess.
- For place names, store the verbatim name as the contributor said
it. Do not resolve to modern names in this call; that is a
separate call. If a place is clearly someone's residence
(contributor said "she lives in [place]"), set `is_likely_residence:
true`.
- If you cannot make out a word, write `[inaudible]` in
`transcript_original` and add an entry to `flagged_for_user_review`
with the field path and a one-sentence reason. Do not invent.
- If the contributor mentions someone in an estranged or sensitive
way ("the cousin we don't speak about", "Tía Marisol who left us"),
record the verbatim quote in the kinship reference. Do not soften.
Do not omit. The family decides what to keep.
Output ONLY the Memo JSON matching the provided schema. No
commentary outside the structured output.
```
---
### Call: Per-memo digest (who was named, what kinship words used)
Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none)
```
You receive one Memo's transcript and kinship references. Your task:
produce a tight English (or chosen target language) digest of ≤ 200
words that the family's organising grandchild can scan in under a
minute.
The digest must:
- Open with the contributor's name and what side of the family they
speak from ("Lola Vicenta speaks from her own (maternal) side of
the family").
- Name every person the contributor named, with the kinship word
preserved verbatim in the original language and a parenthetical
gloss only on first occurrence in this digest. Example: "Tita
Norma (Lola's first cousin, mother's side, lives in Cebu)".
- Note any code-switches or multilingual inserts that change meaning
(a Cantonese-speaking auntie who slipped into English for the
word "godfather" — note it, because the choice of language matters
to who she is naming).
- End with a one-sentence honest note about what was hard to make
out ("the second cousin's name was inaudible — to verify with the
contributor").
Hard rules:
- Do NOT translate kinship words. Western gloss in parentheses on
first occurrence only; never as a replacement.
- Do NOT translate first names.
- Do NOT add interpretation, speculation, or commentary about the
family's history. The digest names; the family interprets.
- Do NOT exceed 200 words.
Output: the digest as a single string. No commentary.
```
---
### Call: Propose unified family graph (long-context, archive-wide)
Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none, long context over every memo)
```
You receive every Memo in the archive at once (long-context). Your
task: propose a unified family graph with PersonNode and
RelationshipEdge records, and a set of Conflict records surfacing
every disagreement between contributors.
You do NOT resolve conflicts. The family does. Your job is to
surface them clearly, with verbatim quotes from every contributor
whose memo contributed to the disagreement.
Hard rules on identity resolution:
- Two memos that both name "Tita Norma" refer to the same person
ONLY if multiple signals align: same side-of-family, same
approximate generation offset from the matriarch, same residence
hint, same relationship to a named anchor person. If only the
name aligns and other signals contradict, create TWO nodes and
a `different_people_called_same_name` conflict.
- A name like "Tita Baby" or "Tita Inday" or "Tía Chela" is a
nickname — common across many families and across multiple
generations. NEVER auto-merge two "Tita Baby"s without strong
multi-signal confirmation.
- Two memos that name the same person with different name spellings
("Norma" / "Norm" / "Norma-Norma" / "Tita Norma") and aligned
contextual signals merge into ONE node with the spellings as
`aliases[]`. Choose `name_canonical` as the most formal name
attested across the memos; if no formal name is attested, use the
longest spelling.
Hard rules on generations:
- `approximate_generation_offset_from_matriarch` is an integer with
the matriarch at 0. Her parents are `-1`, her grandparents `-2`,
her children `+1`, her grandchildren `+2`. A cousin on her own
generation is `0`. A cousin's child is `+1`.
- If two memos disagree on a generation offset for the same person,
emit a `generation_mismatch` conflict. Both verbatim quotes
attached.
Hard rules on relationships:
- `relationship_verbatim` is the exact phrase the contributor used
("Tita Norma is my Mama's first cousin"). Do not paraphrase.
- `relationship_normalised` is the closest enum value; if none fits
cleanly, use `unknown_relation`. Confidence reflects how cleanly
the verbatim phrase maps to the enum.
- Edges are asserted_by the memo IDs that named them. If another
memo contradicts an edge (e.g., one memo says "Tita Norma is my
Mama's first cousin" and another says "Tita Norma is my Mama's
niece"), record both: the edge `asserted_by` the first memo, the
same edge `contradicted_by` the second, and emit a
`relationship_disagreement` conflict.
Hard rules on conflicts:
- Every conflict carries the verbatim quotes from every memo that
asserts a side of the disagreement, with `audio_start_seconds`
and `audio_end_seconds` so the family can replay the exact moment.
- `proposed_question_for_matriarch` is ONE sentence the matriarch
(or organiser) can read aloud at the next family video call. Plain
language, no jargon. Example: "Lola — Cousin Marisol says Tita
Baby is on your sister Edith's side; Lola Vicenta said Tita Baby
is on your brother Romeo's side. Which is right?"
- Conflicts open by default. Do NOT pre-resolve.
Hard rules on kinship words:
- `kinship_terms_used` on each PersonNode lists every kinship word
any contributor used for that person, verbatim, never collapsed.
"Tita Norma", "Tía Norma", "Cousin Norma" are three different
kinship words for the same person if signals align — record all
three.
Hard rules on chosen family and informal adoption:
- Godparents, ninang/ninong, compadre/comadre, informal aunties and
uncles are first-class nodes with the relationship enum
`godparent_of`, `godchild_of`, `informal_aunt_or_uncle_of`, or
`informally_adopted_into`. They are not second-class to the
biological tree.
Hard rules on estranged and sensitive branches:
- An estranged or sensitive branch is recorded verbatim. Do NOT
omit. Do NOT soften. The family chooses what to hide; the data
is preserved.
Hard rules on confidence:
- `confidence` on every edge and merge is in `[0, 1]`. Below 0.6 is
always surfaced for review.
- `graph_confidence_overall` is the median of every edge confidence;
not an average, so a single very-low edge doesn't drag the whole
graph below the family's comfort threshold.
Output ONLY the FamilyGraph JSON matching the provided schema. No
commentary outside the structured output.
```
---
### Call: Re-propose graph after a memo is added or removed
Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none, long context)
```
You receive the previous FamilyGraph proposal plus every Memo in
the archive (including the newly-added or remaining-after-removal
memos). Your task: produce an updated FamilyGraph that incorporates
the change while preserving every family-resolved conflict.
Hard rules:
- Any Conflict whose `resolution_status` is `matriarch_confirmed`,
`cousin_confirmed`, or `family_agreed_to_disagree` MUST be
preserved verbatim in the new proposal. Do NOT re-open a resolved
conflict because the new memo provides new evidence; instead,
emit a NEW conflict referencing the resolved one and surface it
to the family.
- Any PersonNode whose family-confirmed name, side-of-family, or
generation has been adjudicated by the family MUST NOT be
silently overridden by the new memo. If the new memo contradicts
a family-confirmed value, emit a new conflict.
- If a contributor has been removed from the archive and their
memos deleted, every node and edge attributable only to those
memos goes too. Nodes and edges attested by other contributors
remain.
- Surface every new merge, every new edge, and every new conflict
in a top-level `changes_since_last_proposal[]` array so the
organiser can see what changed in one scroll. (Extend the schema
with this array; do not silently fold the changes into the body.)
Output ONLY the updated FamilyGraph JSON. No commentary outside the
structured output.
```
---
### Call: Resolve an ambiguous place name to modern coordinates
Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: `google_search` grounding
```
You resolve an ambiguous place name a relative mentioned in a voice
memo to modern coordinates, while preserving the verbatim name.
Given a verbatim place name and surrounding context ("Tita Norma
lives in General Trias", contributor speaks Tagalog and is based in
Cebu), return:
- `name_verbatim` (unchanged)
- `candidates[]`: an array of plausible modern resolutions, each
with `modern_name`, `modern_coordinates` (lat, lng), `country`,
`administrative_context`, and a `disambiguation_hint` ("the General
Trias in Cavite, Luzon — population ~450k" vs "the General Trias
in Davao region, smaller barangay-level locality")
- one citation URL per candidate from grounded search
Hard rules:
- Use `google_search` grounding for ANY place name that could be
ambiguous, or whose modern administrative context the family
might not know (a barangay that was renamed; a colonial-era
district that no longer exists; a town that crossed a border
during the twentieth century).
- Do NOT pick a winner. Surface every plausible candidate so the
family chooses.
- Preserve the period name. Do not silently substitute the modern
name.
- If the place no longer exists under any modern name (a village
destroyed in war, a barangay merged into another locality), say
so explicitly: `modern_name = "[name], no longer exists as a
distinct locality; approximate former location near [nearest
modern locality]"`.
Output the response as JSON in the text body (NOT via
`responseSchema` — `responseSchema` and `google_search` cannot be
combined in the same Gemini call today). Server-side: parse the
JSON, then read citation URLs from the response's
`groundingMetadata.groundingChunks[].web.uri` — do NOT ask the
model to include URLs in the JSON body; it will hallucinate them.
No commentary outside the JSON.
```
---
### Call: Generate TTS playback of a transcript in source language
Model: `gemini-3.1-flash-tts-preview` · n/a · n/a
```
Voice: warm, unhurried, the pace of a relative speaking to the
matriarch over a kitchen table. Pick the Gemini 2.5 Flash TTS voice
whose `languageCode` matches the memo's `source_language` —
pronunciation will follow that locale automatically. Prefer the
voice gender that matches the contributor's voice when both are
published for the locale; fall back to whichever is available
rather than blocking.
Pre-process the text before sending it to TTS:
- Read from `transcript_original` (the source-language version, in
the script of the language).
- At each sentence break, insert a single ellipsis (`…`) so the
TTS model produces a natural pause. At paragraph breaks, insert a
blank line plus an em-dash (`—`). Gemini 2.5 TTS does not support
SSML `` — these textual cues are how you signal pace.
- For `multilingual_inserts`, keep the whole memo in one voice. The
inserted phrase is rendered in italics in the on-screen subtitle —
the listener sees the language switch even if the audio is one
voice. Optional: stitch a second TTS call client-side for the
inserted phrase if the family asks for dramatic switching.
- Target rate: ~120 words per minute for English narration of a
digest; ~100 words per minute when reading a transcript aloud in
a language the matriarch hears better than she reads.
Style direction: prepend ONE short directive sentence to the
text input, exactly like: "Read warmly and unhurriedly, as a family
member speaking to the matriarch in her own language. …". There is
no separate `style` API field on Gemini 2.5 TTS; the directive
sentence inside the input is how style is conveyed.
Phoneme overrides (Tagalog "ng", Korean "ㄹ", Arabic emphatic
consonants, Hebrew final letters, Punjabi retroflex) are NOT
exposed by Gemini 2.5 TTS — no SSML `` tag. Pronunciation
comes from the chosen voice's native locale.
```
## 5. Use cases & content to include
Build dedicated UI sections or flows for each of these — they tell you what content the app must support.
- **The ninetieth birthday.** A Filipino-Canadian family hosts Lola Vicenta's ninetieth in Mississauga. Over halo-halo on the back deck, three cousins realise none of them agree on whether Tita Norma in Cebu is Lola's first cousin or second. The organising granddaughter sets up the archive, sends magic links to twelve cousins across Toronto, Cebu, Manila, Sydney, and Riyadh, and Lola records four minutes from her armchair. The conflicts appear by Tuesday.
- **The matriarch slipping.** A Salvadoran-American grandson notices his abuela's memory of who-is-whose-cousin is getting harder to pin down. He sets up the archive while she can still record. Four cousins in San Salvador, two cousins in Houston, one cousin in Los Angeles, and a great-aunt in a town in Australia each record ninety seconds. The graph proposes a unified tree with eleven orange dots. The grandson sits with abuela on Sundays for an hour and they close one dot per Sunday.
- **The Lebanese family across three countries.** A Lebanese-Australian family in Sydney has cousins in Beirut, Montréal, São Paulo, and Detroit. The matriarch records in Lebanese Arabic; cousins record in Lebanese Arabic, Australian English, Brazilian Portuguese, and Québec French. The app handles every transcription, surfaces a generation mismatch on Jeddo Karim (placed as great-grandfather by some, great-great-grandfather by one), and proposes one printable handout for the next reunion in 2027.
- **The Cantonese-only grandmother.** A Cantonese-speaking grandmother in Hong Kong records nine minutes in Cantonese naming everyone on her side of the family. Her grandchildren — born in Vancouver, fluent in English, conversational in Cantonese, illiterate in written Chinese — read the English digest first, then play the audio with the transcript in Traditional characters scrolling along. They add their cousins' memos in English and the graph merges across the languages.
- **The Vietnamese boat-family.** A widow in Houston records the names of every relative her husband used to mention before he died. Three nieces in Ho Chi Minh City and two cousins in Paris each record ninety seconds in Vietnamese, Vietnamese-English, and French. The graph proposes a tree, surfaces three conflicts (one about whether Bác Tâm crossed in '79 or '81, two about kinship terms used differently by the southern and central-region branches), and the widow forwards each conflict to the relative she trusts to adjudicate.
- **The Iranian family before and after revolution.** A grandson in Berkeley is rebuilding the family graph with relatives still in Tehran, two cousins in Toronto, and an estranged uncle in Stockholm whom the family rarely speaks to. Khaleh Mehri in Tehran records in Farsi, the Toronto cousins in English with Farsi inserts, the Stockholm uncle in English (he hasn't spoken Farsi since 1986). The graph correctly preserves the Stockholm branch as an estranged-but-visible node — the family chooses whether to hide it on the printable handout, but the data is never destroyed.
- **The Punjabi joint family.** A bride-to-be in Mississauga wants a printable family tree for her wedding handout in six weeks. She invites thirteen aunties and uncles across Punjab, the UK, and Canada. Some record in Punjabi (Gurmukhi), some in Punjabi (Shahmukhi), some in Hindi, some in English. The app surfaces seven cousin-relationships that nobody on her side had ever explicitly stated, the bride-to-be's mother adjudicates over four phone calls, and the handout prints on time.
- **The Korean family with adoptees.** A Korean-American adoptee with cooperation from her birth family records her own ninety seconds with her birth mother's birth family. Halmoni records in Korean; aunties record in Korean and English; the adoptee herself records in English with three Korean kinship words she's learning. The graph treats the adoptee as a first-class node — neither hidden nor flagged as an outlier — and surfaces one conflict over Imo Hyun-ja's husband's name, which Halmoni and Imo Hyun-ja remember differently.
- **The Salvadoran-American matriarch and the second family nobody talks about.** A grandson discovers, while reviewing the matriarch's memo against two cousin memos, that his abuelo had a second family in Mexico City whom abuela never named and the Mexico City cousins are not invited to the archive. The grandson does not act on this. The app surfaces the conflict (one cousin mentions "Tía Lupe" who appears nowhere else and references a Mexico City address), but does not pursue it. The matriarch is left to decide if and when to bring it up.
## 6. Page structure
Build the following screens / sections in this order. Adjust copy to fit the voice, but keep the structural intent.
1. **Welcome / sign-in.** A photographed-looking image of a family around a kitchen table, the matriarch in the foreground holding a phone close to her face to record, a young grandchild beside her listening. One paragraph: "Cousin Map turns voice memos from cousins on every continent into one family tree — with every kinship word kept in its own language, and every disagreement surfaced for the family to decide." Single Google sign-in button; Apple sign-in next to it. Below: "Try with the sample archive" → loads the demo archive in section 8a.
2. **Empty state — "Start an archive".** Two big inputs side by side: 🎙 Record the matriarch's memo (large, primary) · 🔗 Invite cousins (secondary). Short explainers below each ("Start with whoever holds the most family memory — usually an elder.", "Send a magic link to every cousin who can add ninety seconds."). A third small link below: "Add a memo later — start with cousins first" (for families where the matriarch needs more notice).
3. **Matriarch recording flow** (mobile-first). A single full-bleed record button — pressable with one finger held steady. Above it, a soft prompt: "Tell us about your side of the family. Anyone you remember. Take your time." Audible "I'm recording" confirmation. Visual waveform that pulses gently so the matriarch can see the phone is listening. Soft cap at sixty minutes; if she keeps going past, the cap dissolves with a calm "you can keep going". After stopping: "Listen back" / "Re-record" / "Save and add another later".
4. **Cousin invitation flow.** Modal: "Invite a cousin to add ninety seconds." Form: cousin's name (the kinship word + name the matriarch uses for her, "Tita Cora"), email, optional pre-filled note ("Lola is putting our family tree together. Will you tell her who you remember on your side?"). Sends a magic link; cousin opens it on her own phone in her own time. The matriarch sees a small green dot when each cousin records.
5. **Cousin recording flow** (mobile-first). Opens directly from the magic link — no sign-up gates. Shows the cousin who invited her ("Lola Vicenta invites you to add your memory of our family to the tree."). A single record button. Soft cap at three minutes. After stopping: "Listen back" / "Re-record" / "Send to Lola" (sends the memo to the archive immediately).
6. **Processing view.** A vertical list of every memo received. Each row shows the contributor's avatar, the kinship word + name the matriarch uses for her ("Tita Cora"), the language detected ("Cebuano"), the duration ("01:32"), and a step-by-step honest progress: "Transcribing the Cebuano…" → "Translating for the digest…" → "Identifying the relatives Tita Cora named…". Each step takes 6-15 seconds. The user can close the app and come back.
7. **Memo detail view.** A two-column layout on desktop, stacked on mobile. Left column: the audio player with the contributor's verbatim transcript in their own script (Cebuano in Latin script, Cantonese in Traditional Chinese, Korean in Hangul) scrolling in sync with playback. Right column: the English (or chosen target) digest — ≤ 200 words, naming every relative Tita Cora named, with kinship words preserved verbatim. Sticky header: contributor → language → date recorded. Below: a small "(i) why did the model digest it this way?" icon that surfaces the `thoughtSummary` only when tapped.
8. **Family-tree view.** A radial layout with the matriarch at the centre and branches resolving outward. Each node shows the kinship word + name ("Tita Norma"), the contributor whose memo first named her ("named by: Lola Vicenta"), and any photograph the family attached. Edges show the relationship in the verbatim phrase a contributor used ("Mama's first cousin"). Low-confidence edges are dashed. Orange dots mark unresolved conflicts. The radial layout pans and zooms by drag and pinch on mobile.
9. **Conflicts view.** A vertical list of orange dots. Each entry shows: the conflict type ("Same person, two names"), the involved persons, the verbatim quotes from each contributor with their audio clip and a play button, the model's `proposed_question_for_matriarch` ("Lola — Cousin Marisol says Tita Baby is on your sister Edith's side; you said Tita Baby is on your brother Romeo's side. Which is right?"), and three actions: "I'm right" / "The cousin's right" / "Ask [named cousin] who was there".
10. **Glossary view.** The family's kinship words — every "Tita", "Lola", "Lolo", "Tito", "Ate", "Kuya", "Ninang", "Tía", "Khaleh", "Halmoni" used by any contributor — with the family's own glosses (filled in by the matriarch or organiser, never auto-filled by the model). A separate section for chosen-family terms (ninang/ninong, compadre/comadre, godparent_of relationships).
11. **Sharing & invitations.** Modal: "Add another cousin to the archive". Magic-link email; arrival drops the cousin straight into the cousin recording flow. Below: a list of cousins already invited, with each one's status ("recorded yesterday", "magic link sent, hasn't opened it yet", "asked to be removed — memos deleted").
12. **Family-reunion handout export.** Side-by-side typeset preview. Choose: radial layout or vertical hierarchical layout. Toggle: include estranged branches, include photographs, include the orange-dot footnotes ("there are seven things our family is still figuring out"). Export PDF; future tier: print-bound family-reunion handout.
13. **Audio-archive export.** "Download every original memo." A single ZIP of every audio file at upload quality, every transcript file in the source language, every English digest, and a `manifest.json` with timestamps, language tags, and contributor metadata. The family owns the data forever.
14. **Footer.** "Made for the families spread across oceans." Privacy: "Your archive is yours. We never train on it. Any cousin can ask for her memos to be deleted at any time, and the tree re-proposes without her." Capabilities `(i)` icon in header.
## 6b. First-visit onboarding
Show a **first-visit onboarding** the first time a visitor lands on the app (detect via `localStorage` flag; do not show on return visits). Three slides, dismissible at any time. Persistent re-entry: a `?` icon in the header reopens it.
**Slide 1 — What this is.**
- Headline: "Welcome to Cousin Map."
- Subhead: "Build your family tree from voice memos across any continents, any languages — the model surfaces the disagreements; your family decides."
- One paragraph (≤ 60 words) explaining who this is for and what makes it different from a generic genealogy app: it listens to elders in their own language, it preserves every kinship word verbatim, and — most importantly — it never resolves a family disagreement on its own. The matriarch (or whoever the family designates) adjudicates.
- Visual: a small annotated illustration of a radial family tree with one orange dot on it, labelled "this is a question for your family — the app won't answer it" — not a generic tree icon.
**Slide 2 — Try it now.**
- One short prompt: "Try with the sample archive".
- A live demo input pre-loaded with the seed archive in section 8a — Lola Vicenta's memo and three cousin memos, with two surfaced conflicts and one resolved conflict.
- 1-2 sentences pointing at *the specific page elements* where the Gemini magic happens (the cross-memo merge of "Tita Norma" from three different cousins' memos, the surfaced generation conflict on Lolo Pidro, the kinship-word glossary showing four different relatives all called "Tita Baby" preserved verbatim).
**Slide 3 — How to remix this.**
- Headline: "Make this yours."
- Three short bullets:
- "Swap the sample archive in `/data/seed-archive/` for your own family's memos."
- "Adjust the prompts in `/server/prompts/` to fit your family's language(s) and kinship terms."
- "Wire up your Gemini API key and Firebase project via the env-var list in the capabilities panel."
- Primary CTA: "Use this template" → links to AI Studio Build remix entry point.
- Secondary: "Just exploring — close" (sets localStorage flag, never auto-shows again).
**Accessibility:** focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, `aria-labelledby`, focus restored to trigger on close. Respect `prefers-reduced-motion`.
**Don't:**
- Don't gate content behind the modal. The page beneath must be fully usable.
- Don't auto-reshow on return visits. Use `localStorage['onboarding-seen-v1']`.
- Don't include unrelated CTAs (newsletter signup, social follow). Keep it about the template only.
## 6c. Capabilities info button (persistent in header)
Add a persistent `(i)` icon in the top-right of the header (next to the primary nav). Click → opens a modal/panel titled **"What powers this app"**.
**Panel contents (in this order):**
**Gemini capabilities used (the hero list):**
- **Gemini 3.5 Flash (audio multimodal)** — transcribes spontaneous family-memory monologues in Tagalog, Cebuano, Cantonese, Vietnamese, Spanish, Arabic, Korean, Punjabi, Amharic, Farsi, Khmer, English, and dozens more. One call per memo for transcription. Automatic language identification per memo with BCP-47 tagging.
- **Gemini 3.5 Flash (long context, 1M tokens)** — the hero capability. The family-graph proposal call sees every memo at once and notices when relative A's "my cousin in Cebu" and relative B's "Tita Norma" are clearly the same person. Without long context, no unified graph; with long context, the demo works.
- **Gemini 3.5 Flash (structured output)** — every memo and the family-graph proposal return typed JSON matching the schemas in section 4b. Numeric confidences are clamped server-side.
- **Gemini 3.5 Flash + grounded search** — only for ambiguous place-name resolution ("General Trias" — which one?). Never for relationship resolution; the family decides those.
- **Gemini TTS (2.5 Flash Preview)** — reads each transcript aloud at the matriarch's pace, in the language of the memo, on a voice native to that locale. No mid-call voice switching; one voice per memo.
- **Firebase Auth** — Google and Apple sign-in. Magic-link email for cousin invitations.
- **Firestore** — stores your archive, syncs across devices in real time. Conflicts and resolutions persist across re-proposals.
- **Firebase Storage** — keeps the original voice memos at upload quality, forever. The family owns the source recordings.
- **Cost note** — see the detailed breakdown in 6d. A family of 12 cousins recording into an archive once costs about $0.45 of Gemini API spend, total. Weekly re-proposals of the family graph as new cousins join cost about $0.15 each.
- **Privacy note** — your family's voice memos are private to the matriarch and the cousins she invites. This app uses the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. Any cousin can ask to be removed from the archive at any time; her memos are deleted within 60 seconds and the family graph re-proposes without her.
**Backend services this app depends on:**
- Auth: see section 4b
- Database: see section 4b
- Storage: see section 4b
- Email: see section 4b
- Payments: see section 4b (not used in v1)
- External APIs: see section 4b
**Environment variables you'll need to configure:**
- `GEMINI_API_KEY` — your Google AI Studio API key
- `FIREBASE_PROJECT_ID` — your Firebase project id
- `FIREBASE_SERVICE_ACCOUNT` — service-account JSON (server-side only)
- `FIREBASE_STORAGE_BUCKET` — your Firebase Storage bucket name (must be enabled in the Firebase console first; not auto-provisioned)
**Cost + privacy notes:**
- One short paragraph per cost-sensitive capability: long-context family-graph proposal is billed per token of input — a 12-memo archive at ~80k tokens costs about $0.10 each time it runs. Weekly is plenty; the model will not produce meaningfully different output across hours.
- One short paragraph on privacy: where the data lives (your Firebase project), how to delete it (Settings → "Delete this archive forever" — gone in 60 seconds), what is never sent for training, and how a single cousin can remove herself from the archive without disrupting the rest of the family's work.
**Documentation links:**
- AI Studio Build docs
- Gemini API audio multimodal, long-context, structured-output, TTS docs
- Firebase Auth, Firestore, Firebase Storage docs
- A short note on the BCP-47 language tags the app uses and how to extend the kinship-word list
**Accessibility:** same standards as the onboarding modal — focus trap, `Esc`, ARIA, restored focus.
**Behaviour:**
- Always available — single click from anywhere in the app.
- Tooltip on the `(i)` icon: "How this app is built".
- Mobile: opens as a full-screen sheet that slides up.
- Should be the most honest part of the app — never hand-wave service requirements; never say "AI" without naming the specific Gemini model and capability.
## 6d. Detailed cost breakdown (deployer reads this BEFORE shipping)
- **Transcribe + language-identify a memo (Gemini 3.5 Flash, low thinking)** — typical 90-second memo ≈ ~1.5k input audio tokens, ~600 output tokens. ~$0.005/memo. A 4-minute matriarch memo ~$0.012.
- **Per-memo digest (Gemini 3.5 Flash, low thinking)** — typical 200-word digest ≈ ~250 output tokens over ~1k input. ~$0.0003/memo.
- **Propose unified family graph (Gemini 3.5 Flash, medium thinking, long-context, archive-wide)** — 12-memo archive ≈ ~80k input tokens, ~5k output tokens. ~$0.10/run. Weekly is plenty; default schedule: nightly during the first week of an archive's life, weekly thereafter.
- **Re-propose graph after a memo is added or removed (Gemini 3.5 Flash, medium thinking, long-context)** — same as above, ~$0.10/run. Triggered automatically on memo add/remove.
- **Geocode an ambiguous place name (Gemini 3.5 Flash + grounded search)** — ~$0.001/place. A typical 12-memo archive has 4-8 distinct places; runs once at first reference, never again unless the family edits.
- **TTS playback of a transcript (Gemini 2.5 Flash TTS)** — billed per output token (~$10/M output tokens), effectively ~$0.000003/character. A 90-second memo transcript ≈ ~600 characters ≈ ~$0.002 per full read. Cached per memo; charged once unless the transcript is edited.
- **Expected one-time cost for a 12-cousin family archive at first ingest:** ~$0.45.
- **Expected ongoing cost per weekly graph re-proposal as the archive grows:** ~$0.15 at 12 memos, ~$0.40 at 40 memos, ~$1.20 at 100 memos.
- **Audio storage:** Firebase Storage standard tier, ~$0.026/GB/month. A 90-second voice memo at 128kbps AAC is ~1.5 MB; a 12-memo archive uses ~20 MB ≈ < $0.01/month. A 100-memo archive ≈ ~150 MB ≈ < $0.01/month.
## 7. Design language
- **Mood:** A family living room across three time zones at once. Not a tech product. Not a genealogy database. The kitchen at 9 pm with the phone held close to grandmother's mouth, and the same phone at 4 am Sydney time as a cousin sits on her bed and tells the app who she remembers. Patient, warm, never rushed. The app waits as long as the matriarch needs.
- **Typography:** A humanist serif for transcript text and digest body (Source Serif Pro or PT Serif) — the typography of a letter, not a database. A clean grotesque for app chrome (Inter or Geist). The matriarch's transcript displays at a generously large default size (18-20px) with a one-tap zoom to 24-28px for elders. Scripts that aren't Latin — Cantonese, Korean, Arabic, Devanagari, Gurmukhi, Khmer — display in the script with a font stack that prioritises Noto Sans for the script in question, never falling back to a Latin font that mangles the characters.
- **Palette:** Warm parchment background `#F5EFE3` for the memo-detail view, deep ink `#1B1714` for body text, a quiet teal `#2A6F77` for verbatim kinship words (so they stand out as the family's own language), a soft amber `#C97A2A` only for conflict markers and orange dots on the radial tree, a muted indigo `#3A4D77` for the family's own annotations and the matriarch's adjudication notes. No saturated reds; this is not an alert UI.
- **Imagery:** The voice waveforms are the hero. Each memo's waveform displays inline beneath the contributor's name, scrubbable, with the section the user is currently reading highlighted. Photographs cousins attach to nodes appear at their uploaded aspect ratio — never cropped square by default. The radial family-tree layout uses thin lines and generous whitespace; nodes breathe.
- **Hand-feel touches:** Pressing the matriarch's record button produces a gentle low-frequency confirmation tone (not a beep). The audible "I'm recording" voice confirmation uses the same voice the app will later use for TTS playback in the matriarch's language. The orange dots on the radial tree pulse very slowly (3-second cycle) so the family notices them at a glance without feeling alarmed; pulse respects `prefers-reduced-motion`.
- **Spacing:** consistent 4-px base. Generous whitespace — the voice memos are listened to, not consumed.
- **Radius:** consistent token set (e.g. 8 / 14 / 22 px). Memo cards use 8; the family-tree node bubbles use 14; the welcome card uses 22.
- **Shadows:** subtle, layered, warm-tinted. Avoid heavy drop-shadows.
- **Motion:** purposeful — entrance fades, hover lifts, page transitions. Respect `prefers-reduced-motion`. No bouncing splash animations. No theatrical hero animations. The radial tree's re-layout when a new cousin's memo is added uses a 600 ms ease-out tween with reduced-motion falling back to instant.
- **States:** every interactive element has hover, focus, active, disabled. Loading uses skeletons not spinners where possible. Empty states have helpful next-action guidance ("Send a magic link to Tita Cora to add her ninety seconds").
## 8. Content generation rules
- Write **realistic, specific copy**. NO Lorem Ipsum. NO generic placeholders like 'Your tagline here'.
- Invent plausible names, dates, kinship terms, voice-memo snippets that fit the domain (use the seed content in section 8a as a starting point). When inventing, anchor in real diaspora patterns — Filipino-Canadian families in Toronto and Mississauga, Lebanese-Australian families in Sydney, Salvadoran-American families across San Salvador / Houston / Los Angeles, Vietnamese-American families in Houston with relatives in Ho Chi Minh City and Paris, Korean-American adoptee families, Punjabi families across Punjab / UK / Canada — but never claim that a fictional cousin is a real person.
- Tone: warm, direct, free of corporate language. This template is for a family, not a company.
- Headlines: punchy and concrete. No 'Empower your X' filler. No 'Revolutionize'. No 'Seamless'.
- Body copy: short paragraphs (2-4 sentences). Use lists where appropriate.
- Plain language. The matriarch's UI text is plain enough that an elder reads it without glasses. The cousin's UI text is plain enough that a teenager records ninety seconds without thinking about the app.
- Where the app outputs AI-generated content (a digest, a proposed family-tree layout), never label it as "AI says" — let it speak naturally. Use small uncertainty cues only where epistemic honesty requires them (a dashed edge for low-confidence, an orange dot for a conflict, a "the model thought this was…" panel only behind a tapped `(i)` icon).
## 8a. Seed content (use these specific examples)
Anchor every generated copy + sample data point in the concrete content below. Use these names, numbers, dates, and snippets verbatim where helpful, or generate close variants that sit in the same world.
**Sample archives (sidebar):**
- "Lola Vicenta's Tree" (12 cousins recorded, matriarch Lola Vicenta Reyes, 90 years old, Mississauga; cousins in Cebu, Manila, Toronto, Sydney, Riyadh) — primarily Tagalog and Cebuano, with Toronto English and Sydney English from the younger cousins. 1 matriarch memo (4:12) + 11 cousin memos (avg 1:38).
- "Abuela Marisol" (8 contributors recorded, matriarch Abuela Marisol Ramírez, 86 years old, San Salvador; cousins in San Salvador, Houston, Los Angeles, Adelaide) — Salvadoran Spanish, Mexican Spanish (the LA branch), Australian English (the Adelaide great-niece who barely speaks Spanish). 1 matriarch memo (5:48) + 7 cousin memos.
- "Jeddo Karim" (14 contributors recorded, patriarch Jeddo Karim Saliba, 87 years old, Sydney; cousins in Beirut, Montréal, São Paulo, Detroit, Sydney) — Lebanese Arabic from the elder generation, Lebanese Arabic + Australian English + Québec French + Brazilian Portuguese from the cousins. 1 patriarch memo (6:30) + 13 cousin memos.
- "Halmoni's Memory" (9 contributors recorded, Halmoni Soon-ja Park, 88 years old, Seoul; cousins in Seoul, Los Angeles, Vancouver, the Korean-American adoptee in Minneapolis with her birth family's cooperation) — Korean from Halmoni and three aunties, Korean+English from the LA and Vancouver cousins, English with three Korean kinship words from the adoptee. 1 Halmoni memo (3:55) + 8 cousin memos.
**Sample matriarch memo in detail view (this is what the demo should show):**
- **Contributor (verbatim):** "Lola Vicenta"
- **Contributor inferred full:** Vicenta Reyes (née Aquino), the matriarch, 90 years old, recorded in her armchair in Mississauga
- **Recorded at:** 2026-03-14T19:42:00-04:00
- **Duration:** 4:12
- **Source language:** Tagalog (tl-PH)
- **Source language confidence:** 0.94
- **Multilingual inserts:** one English insert "you know — first cousin once removed, that's what they call it in Canada" with register note "code-switch to English to name the Western kinship term"; one Cebuano insert "ang akong igsoon" → "my sister" with register note "the matriarch slips into Cebuano when naming her own siblings, a habit from her childhood in Cebu before the family moved to Manila"
- **Transcript original (excerpt):** "Anak ko, simulan natin sa nanay ko — Mama Carmen — siya ang nanganak sa amin lahat sa Cebu. Apat kami: ako, si Tita Pacita, si Tito Bert, at si Tita Baby — siya ang bunso, namatay siya noong 1962. … Si Tita Norma sa Cebu, siya ay first cousin ko sa nanay ko side — anak ng kapatid ni Mama Carmen na si Tita Inday. Pero may isa pa rin akong Tita Norma — sa tatay ko side, ibang Tita Norma 'yon, anak ng kuya ni Itay Pidro. Yan ang dapat ninyong itanong sa akin para hindi malito."
- **Transcript original (longer excerpt):** "Hindi ko na alaala ang lahat — pero subukan ko. Si Tita Baby — yung bunso, yung namatay — at may iba pang Tita Baby — yung anak nina Tita Pacita, siya 'yung pamangkin ko. Magkapareho ng pangalan kasi parehong Babette ang totoo nilang pangalan. Hindi ito conflict, anak — totoo lang na may dalawang Tita Baby sa pamilya natin."
- **English digest:** "Lola Vicenta speaks from her own (maternal) side of the family. She begins with her mother, Mama Carmen — who gave birth to all four of Lola's generation in Cebu: Lola Vicenta herself, Tita Pacita (older), Tito Bert (younger brother), and Tita Baby (youngest, died 1962). On her father's side she names Itay Pidro, his older brother (whose son's daughter is the second Tita Norma she keeps having to disambiguate), and two of his sisters whose names she struggles to recall ('the older one — let me think — Tita Lucing, I think'). She explicitly flags two relatives the family confuses: two Tita Normas (one on each side), and two Tita Babys (her own sister who died, and her niece — daughter of Tita Pacita — who shares the formal name Babette). The Cebuano slip-into for 'my sister' is consistent with Lola's pre-Manila childhood. The model digested this in 187 words; one passage at 2:48 was hard to make out and is flagged for review (the second sister of Itay Pidro)."
- **Kinship references (excerpt, 8 of 23):**
- "Mama Carmen" (Tagalog kinship: "Mama" + name "Carmen", language tl-PH, side: mother_side, quote: "siya ang nanganak sa amin lahat sa Cebu")
- "Tita Pacita" (Tagalog: "Tita" + "Pacita", side: mother_side, quote: "ako, si Tita Pacita, si Tito Bert, at si Tita Baby")
- "Tita Baby" #1 (Tagalog: "Tita" + "Baby", side: mother_side, quote: "siya ang bunso, namatay siya noong 1962"; alias_for_canonical: "Babette Aquino")
- "Tita Baby" #2 (Tagalog: "Tita" + "Baby", side: mother_side via Tita Pacita; quote: "may iba pang Tita Baby — yung anak nina Tita Pacita"; alias_for_canonical: "Babette Reyes")
- "Tita Norma" #1 (Tagalog: "Tita" + "Norma", side: mother_side, quote: "siya ay first cousin ko sa nanay ko side")
- "Tita Norma" #2 (Tagalog: "Tita" + "Norma", side: father_side, quote: "ibang Tita Norma 'yon, anak ng kuya ni Itay Pidro")
- "Itay Pidro" (Tagalog: "Itay" + "Pidro", side: father_side, quote: "sa tatay ko side")
- "Lolo Pidro" — NOT used by Lola; her father is "Itay Pidro". Cousin Marisol's memo uses "Lolo Pidro" for the same person. This is the source of the surfaced generation-mismatch conflict, because Marisol's "Lolo" implies one generation older than Lola Vicenta places him.
- **Reading confidence:** 0.91
- **Flagged for review:** one passage at 2:48 ("the second sister of Itay Pidro" — the audio was muffled by a delivery truck outside; matriarch can re-record this section)
**Sample cousin memo (one of eleven):**
- **Contributor (verbatim):** "Marisol"
- **Contributor inferred full:** Marisol Reyes-Tan, 34, Lola Vicenta's grandniece, recorded in her kitchen in Sydney
- **Recorded at:** 2026-03-15T20:18:00+11:00
- **Duration:** 2:04
- **Source language:** Australian English (en-AU) with Tagalog inserts
- **Multilingual inserts:** four Tagalog inserts ("Lolo Pidro", "Tita Baby", "kapitbahay" → "neighbour", "ate" → "older sister")
- **Transcript original (excerpt):** "Right, so on my side — my mum is Tita Pacita's daughter, so I'm Lola Vicenta's grandniece. I never met Lolo Pidro — he died before I was born — but my mum used to talk about him all the time, said he was a quiet man, a teacher in Cebu. And Tita Baby — the one we all loved — she's actually my mum's sister, Babette. So when Lola Vicenta says Tita Baby she means a different Tita Baby, the one who died in 1962. That one's Mum's auntie. I get them confused on phone calls."
- **Surfaced conflicts from this memo against the matriarch's:**
- **Generation mismatch on Itay Pidro / Lolo Pidro:** Lola Vicenta calls him "Itay Pidro" (her father, generation -1). Marisol calls him "Lolo Pidro" (her great-grandfather, generation -3 from Marisol → which would place him at generation -2 from Lola Vicenta, off by one). The conflict surfaces with both quotes, both audio clips, and the proposed question: "Lola — was Pidro your father, or your grandfather? Marisol says her mum called him Lolo Pidro."
- **Confirmation of two Tita Babys:** Marisol's memo confirms what Lola Vicenta said — there are two Tita Babys, both formally named Babette. This is logged as a `person_appears_in_one_memo_only` resolution (i.e., NOT a conflict; surfaced as a "verified by two memos" check mark on both nodes).
**Sample surfaced conflicts (from the demo archive):**
1. **Generation mismatch on Pidro Reyes** — Lola Vicenta places him as her father (Itay); Marisol places him as her great-grandfather (Lolo). Proposed question: "Lola — was Pidro your father, or your grandfather? Marisol says her mum called him Lolo Pidro." Open.
2. **Same kinship-and-name used for two people: "Tita Baby"** — confirmed by Lola Vicenta and Marisol that there are two; cousin Hector's memo from Cebu only mentions one and is ambiguous about which. Proposed question: "Hector — when you said Tita Baby in your memo, did you mean the one who died in 1962, or the one who's still with us in Mississauga?" Open.
3. **Side-of-family disagreement on Tita Norma** — Lola Vicenta names two Tita Normas (one maternal, one paternal); cousin Joanne's memo from Riyadh names "Tita Norma" without disambiguation. Proposed question: "Joanne — when you said Tita Norma is your favourite, did you mean the one in Cebu (Mama Carmen's niece) or the one in Iloilo (Itay Pidro's niece)?" Open.
4. **Person appears in one memo only — surfaced as 'verify'** — cousin Eugene's memo from Manila names a "Tita Cora" who appears nowhere else. The model surfaces this not as a conflict but as a friendly "ask whether Tita Cora belongs in the tree" check. Open.
**Sample voice copy:**
- Onboarding: "Tell us about your side of the family. Anyone you remember. Take your time."
- Cousin invitation email subject: "Lola is putting our family tree together. Will you add ninety seconds?"
- Cousin invitation email body: "Hi Marisol — Lola Vicenta is putting together a family tree. She's recorded who she remembers on her side; will you add who you remember on yours? Ninety seconds, in any language you'd like. Tap here." [Add to Lola's Tree]
- Processing: "Listening in Tagalog…" / "Writing it down…" / "Reading who Lola named…" / "Looking for who else mentioned them…"
- Empty archive: "Lola's Tree is waiting for its first voice. Tap the big red button when Lola is ready to start."
- Save confirmation: "Added to Lola's Tree — Marisol's memo from Sydney, 2 minutes 4 seconds."
- Conflict surfaced: "Two cousins remember Pidro differently. The model has set this aside as a question for the family — not an answer."
- Resolved conflict: "Lola decided: Pidro is her father, Marisol's great-grandfather. Generation locked. The tree updates tomorrow."
- Low confidence: "This passage was hard to hear — a delivery truck went past at 2:48. Want to re-record this section or leave it as is?"
- Cousin removal: "Tita Joanne has asked for her memos to be removed. They've been deleted; the tree will re-propose tonight without them."
**Sample multilingual diaspora content:**
- A Cantonese-only grandmother's nine-minute memo in Cantonese (Traditional Chinese characters): "我哋呢邊嘅人 — 由阿婆 阿娥 開始 — 佢係嗰啲早年由廣州去香港嗰批人之一…" with English digest naming Po Po Ngor, Goo Goo Mei-Ling, Kau Foo Yat-sing.
- A Lebanese matriarch's memo in Lebanese Arabic transcribed in Arabic script: "بدنا نبلش من جدّي كريم — هو من الناصرة بالأصل — تعرّف على تيتا فادية بالشام عَ ضو الحرب التانية…" with English digest naming Jeddo Karim, Teta Fadia, Khaleh Lina.
- A Vietnamese widow's memo in Vietnamese with full tone marks: "Bắt đầu từ ông nội — Ông Tâm — ông là người vượt biên năm bảy chín…" with English digest naming Ông Tâm, Bà Hoa, Bác Năm.
- A Korean Halmoni's memo in Hangul: "우리 집안 시작은 우리 아버지부터예요 — 박씨 가문 — 아버지께서 만주에서 돌아오셨을 때…" with English digest naming Harabeoji Park, Halmoni Park, Imo Hyun-ja.
- A Salvadoran matriarch's memo in Salvadoran Spanish: "Vamos a empezar con mi mamá — Mamá Lupita — ella nació en Sonsonate en mil novecientos veinte…" with English digest naming Mamá Lupita, Tío Beto, Tía Chela.
**Sample family-reunion handout export (PDF preview content):**
- Cover: "The Reyes-Aquino Family Tree — March 2026, prepared by Lola Vicenta and 11 cousins across 5 countries"
- Page 1: radial tree with Lola Vicenta at the centre, paternal branch left, maternal branch right, chosen-family (godparents, ninang/ninong) branch above, low-confidence edges dashed, conflict dots orange.
- Page 2-3: glossary of kinship words used in this family ("Tita", "Tito", "Lola", "Lolo", "Itay", "Inay", "Ate", "Kuya", "Ninang", "Ninong", "Apo"), with the family's own glosses contributed by Lola Vicenta and Marisol.
- Page 4-5: an honest "what we're still figuring out" section listing the four open conflicts and the named relatives the family wants to ask.
- Page 6: contributors page — every cousin who contributed a memo, with their photograph, location, and the duration of their contribution.
- Back cover: "This tree was built by our family over six weeks. Some things are still being figured out. That's part of it."
## 9. Media & assets
- **Hero image (landing screen):** A photographed-looking shot of a family kitchen at evening, the matriarch in the foreground holding her phone with both hands at chin height to record a voice memo, a grandchild beside her listening with one hand resting on her arm. Warm light from a hanging lamp; a steaming mug on the table. Generate via Nano Banana 2 with a prompt emphasising "wooden kitchen table, warm pendant-lamp light, grandmother in her late eighties of Filipino heritage holding her phone with both hands, a teenage grandchild beside her listening, late evening, soft window light from the left, a half-finished mug of milky tea, the phone screen glowing softly".
- **App icon / wordmark:** Set in the humanist serif. A single small radial-dot motif at the leading edge of the wordmark — three concentric rings with one orange dot off to one side suggesting an unresolved conflict.
- **Empty-state illustration:** A simple line drawing of an open phone receiving a voice memo waveform; behind it, the radial outline of a family tree with one node lit and the others waiting.
- **Demo memo waveforms:** Generated visually per memo. Each waveform is the actual amplitude envelope of the audio, never a decorative pattern. Use the open-source `WaveSurfer.js` library for scrubbable playback that highlights the section being read by the user.
- **Family-tree node bubbles:** Subtle warm-grey background, kinship word in the teal accent colour at the top, name in the deep-ink body colour below, contributor avatar at the bottom-left, photograph (if any) at the bottom-right. Orange dot at the upper-right corner of any node involved in an open conflict.
- **Stock fallbacks:** If image generation fails, fall back to the photographed sample kitchen-table image from `/public/samples/sample-kitchen.jpg`. Never to a "🧓" emoji.
- **Generated imagery:** prefer Nano Banana 2 over stock photography. Prompt for warmth, asymmetry, and slight imperfection — avoid the glossy 'AI render' look.
- **Optimisation:** WebP/AVIF, `loading="lazy"`, explicit `width`/`height` to prevent layout shift.
- **Icons:** `lucide-react` for UI. Use sparingly — never decorative-only.
### Build-time asset manifest (explicit specs)
Every image, illustration, and visual reference mentioned above must resolve to ONE of the three buckets below — runtime-generated, seed-shipped, or user-supplied. Do NOT ship `` tags whose `src` is not listed here. Do NOT depend on bare "section 8a prompts" without binding them to explicit paths and model IDs.
**Bucket 1 — Runtime-generated (Nano Banana Pro `gemini-3-pro-image` for hero/demo photographs; Nano Banana 2 `gemini-3.1-flash-image` for in-app illustrations and reference-conditioned variants).** Cached to Firebase Storage; served via signed URL. Every reference above to "Nano Banana 2" or "Nano Banana Pro" MUST be wired to one of these specific calls with an explicit model id:
- `/public/generated/hero.webp` (2400×1500, WebP) — model `gemini-3-pro-image` — uses the literal prompt described as "Hero image (landing screen)" above. Run once at build; commit a `/public/samples/hero-fallback.webp` (1600×1000) generated from the same prompt with `gemini-3.1-flash-image` so the page renders if quota is exhausted.
- `/public/generated/demo/{demo-slug}-{NN}.webp` (1600×1200, WebP) — model `gemini-3.1-flash-image` (reference-conditioned where the prior frame is passed as input) — one path per "Demo X" image referenced above. The slug derives from the seed example in section 8a; the NN index covers each frame in the demo sequence.
- `/public/generated/illustrations/{name}.webp` (1024×1024, WebP) — model `gemini-3.1-flash-image` — one path per named illustration above ("Empty-state illustration", "Recipe-card hero illustrations", "Curriculum picker imagery", "Period-style frames", etc.). Each illustration's prompt is the literal description above; ship a deterministic seed in the request so re-runs are reproducible.
**Bucket 2 — Seed assets shipped with the deliverable.** Every "Stock fallback" path referenced above (e.g. `/public/samples/sample-X.jpg`) is generated once via Nano Banana 2 (`gemini-3.1-flash-image`) at 1024×1024 WebP using the same prompt as its Bucket-1 counterpart, then committed to the repo so the page renders identically if Gemini quota is exhausted or the user is offline. Replace any `.jpg` extension above with `.webp` to match the optimisation rule. Also commit these empty-state seeds (1024×1024 WebP, single-stroke hand-drawn line, no colour fill):
- `/public/samples/empty-state-primary.webp` — line drawing of the app's primary empty surface (the named "Empty-state illustration" above), generated from that exact prompt.
- `/public/samples/empty-state-archive.webp` — line drawing of an empty saved/archive view, single-stroke outline.
- `/public/samples/empty-state-error.webp` — line drawing of a hand placing a single object aside with care, used when an AI call fails.
**Bucket 3 — User-supplied.** Uploads from the user's camera / file picker land at the Firebase Storage path conventional for this template (named in section 4b). The build ships with Bucket-1 + Bucket-2 only; no user-supplied images at first paint.
**Hard rules**
- Every `` tag MUST have a `src` that resolves to a path listed in Bucket 1, Bucket 2, or a Bucket 3 upload path. Anything else is a build error.
- No bare `image.jpg` / `hero.jpg` / `placeholder.png` references anywhere in the code.
- Model IDs: `gemini-3-pro-image` for hero-quality photographic generation; `gemini-3.1-flash-image` for in-app illustrations, reference-conditioned variants, empty-state seeds, and stock fallbacks. Never use a legacy model id (no `imagen-*`, no `gemini-1.5-*-image`).
- File format: WebP everywhere (AVIF acceptable where the target browsers support it). No `.jpg` / `.jpeg` / `.png` in `/public/samples/`.
## 10. Interactivity & states
- Every interactive element has hover, focus, active, and disabled states.
- Forms validate inline and show specific error messages (not "Invalid input").
- Loading states use skeletons that match the eventual layout, not spinners.
- Empty states explain the next action with a button whose label fits THIS app's domain: "Record Lola's first memo", "Send a magic link to Tita Cora", "Invite a cousin in Cebu" — never a generic "Add your first item".
- Smooth scroll for in-page anchors.
- All AI-generated content streams in token-by-token where supported, with a clear "transcribing in Tagalog…" or "listening…" indicator before content starts arriving.
- If an AI call fails, show a calm, specific error ("We had trouble hearing this passage — a truck went past around 2:48. Want to re-record this section?") and offer retry.
- Low-confidence transcript words are faintly underlined; tapping reveals the alternates the model considered and a "I know what she said — type it" option for the matriarch's grandchild.
- The radial tree's re-layout when a new cousin's memo is added takes 600 ms with `prefers-reduced-motion` falling back to instant.
- Orange conflict dots pulse very slowly (3-second cycle) on the radial tree; pulse respects `prefers-reduced-motion`.
- Audio playback is gapless when the user scrubs across the highlighted-passage boundary; no clicks, no pops.
## 11. Tech & responsive requirements
- **TTS markdown-stripping preprocessor:** before sending any user-authored markdown to `gemini-3.1-flash-tts-preview`, strip non-spoken markdown: `#`/`##`/`###` headings (keep the title text), `**bold**` (keep the inner text), `[label](url)` (keep `label`, drop URL), `` ``` `` fenced code blocks (skip entirely), `>` block-quote markers (keep the text), and `|` table pipes (read row-by-row as sentences). Insert `…` between sentences for a short pause and a blank line plus `—` between paragraphs for a long pause. The model does not understand markdown; raw markdown will be read aloud as literal characters ("asterisk asterisk").
- **File downloads on Safari / Firefox:** when offering local-disk save of any export (PDF, CSV, MP3, ZIP, JSON, image), fall back to `` with a blob URL — the File System Access API (`showSaveFilePicker()`) is Chromium-only. Detect with `'showSaveFilePicker' in window`; otherwise use the anchor-download path.
- **Stack:** React + TypeScript + Tailwind CSS. Functional components + hooks. Use Shadcn UI primitives where appropriate. Audio playback uses `WaveSurfer.js`; family-tree visualisation uses `d3-hierarchy` for the radial layout with custom SVG node rendering.
- **Build runtime:** AI Studio Build — full-stack with Cloud Run server-side functions. All Gemini API calls happen server-side; API key lives in Secrets Manager, never in client bundle.
- **Model selection:** explicitly pin `gemini-3.5-flash` for memo transcription and family-graph proposal, `gemini-3.5-flash` for per-memo digest and geocoding, `gemini-3.1-flash-tts-preview` for TTS. Set `thinkingLevel` explicitly per call. Omit `thinkingConfig` entirely on the TTS call.
- **Database:** Firestore (auto-provisioned by AI Studio Build). Show the seed archive on first launch.
- **Auth:** Firebase Auth — Google sign-in by default; Apple sign-in next to it (requires Apple Developer config); magic-link email as the cousin-invitation flow (requires sender-domain authorisation).
- **Storage:** Firebase Storage for original voice memos and attached photographs. Pre-signed URLs only. Storage is NOT auto-provisioned by AI Studio Build — enable it in the Firebase console first.
- **Mobile-first.** Verify layouts at 375 px (iPhone SE), 768 px (iPad), 1024 px, 1440 px+. The matriarch's recording flow is verified at 375 px first; the radial tree is verified at 1440 px first then adapted down.
- Use `clamp()` for fluid typography. Prefer container queries over media queries for component-level responsiveness.
- Use `dvh` / `svh` instead of `vh`. Respect safe-area insets on iOS.
- Zero horizontal overflow at any width. Zero layout shift on load.
- Persist user data in Firestore. Use real-time listeners on the archive view so the organiser sees a cousin's recording appear within seconds of upload.
- Optimistic UI on writes; reconcile on response.
- Voice recording uses the Web Audio API with `MediaRecorder`; falls back to native file-picker upload on browsers without `MediaRecorder` support. Records as AAC at 128kbps for size + quality; falls back to MP3 if AAC unavailable.
- **iOS Safari gotchas (graceful degradation):** Safari `MediaRecorder` is already AAC — no fallback dance needed, but feature-detect for sanity; mic permission does NOT persist across page reloads on iOS — re-request on every cousin's recording session; an incoming call interrupts the audio session (`MediaStreamTrack.onmute` fires) — auto-pause and offer a clean "resume / start over" choice; backgrounded Safari tabs pause `getUserMedia` — pair `visibilitychange` with a screen Wake Lock during long oral-history recordings.
## 12. Accessibility (WCAG 2.2 AA)
- Semantic HTML — `header`, `nav`, `main`, `section`, `article`, `footer`.
- All interactive controls reachable by keyboard with a visible focus ring.
- Color contrast ≥ 4.5:1 for body, 3:1 for large text and UI components. The teal accent on parchment background and the amber conflict marker on parchment both verified at AA.
- All audio has a synchronised transcript visible by default; no audio-only content.
- The voice waveform has a keyboard-accessible alternative — left/right arrow keys scrub by 1 second, shift+arrow scrubs by 10 seconds.
- All images have meaningful `alt` text. The cousin avatars use the cousin's chosen name in the alt ("Marisol, recorded from Sydney").
- Form fields have associated `