================ ONE-SHOT BUILD CONTRACT (read first) ================ Build this in Google AI Studio "Build" in ONE shot — a complete, working app, no follow-up turns. These are hard rules, not suggestions: 1. TARGET = Full-Stack Web (Node server runtime, secrets, Firebase allowed). If you target Native Android instead, you MUST drop all server/DB/Workspace/ secrets and persist locally (Room / SharedPreferences) only. 2. PIN EVERY MODEL CALL — never let the agent auto-select (it downgrades on quota): - Reasoning / text -> gemini-3.5-flash (thinkingLevel: minimal|low|medium|high) - 4K image + legible text -> gemini-3-pro-image (image_size "4K", up to 14 refs) - High-volume image -> gemini-3.1-flash-image - Expressive TTS -> gemini-3.1-flash-tts-preview (inline tags e.g. [whispers]) - Realtime audio/video (WebSocket) -> gemini-3.1-flash-live-preview - Sandboxed agent -> antigravity-preview-05-2026 3. DIVISION OF LABOR — the model ONLY parses/extracts to a strict responseSchema. ALL math, money (store currency as integer minor units / cents), sorting, balancing and graph logic run in deterministic TypeScript/Python. The model must never compute totals, splits or balances itself. 4. responseSchema sanitation — no regex patterns, no fixed-length tuples, no format validators in the schema (they crash the OpenAPI engine). Enforce those in server-side code AFTER parsing the JSON. 5. responseSchema and google_search grounding are MUTUALLY EXCLUSIVE in one call. 6. CODEGEN — split large output into modular, single-responsibility files so no file is truncated by the output-token cap. 7. Every external call gets a graceful fallback (e.g. manual paste if a Workspace read fails). Never a silent dead end. 8. ROBUST STORAGE & CANVAS — Wrap all `localStorage`/`sessionStorage` operations (especially JSON parsing and writes) in `try-catch` blocks to prevent crashes in private windows or quota overflows. Canvas drawing elements must dynamically handle window resize and scale pixel density (`window.devicePixelRatio`) to avoid blurry graphics on retina displays. ===================================================================== # MUST OBEY — Mobile-first build requirements This app's PRIMARY surface is a mobile phone. Build it impeccably on mobile FIRST, then verify on tablet and desktop. Treat the rules below as non-negotiable hard constraints, not suggestions. ## Viewports to verify (every screen, every state) - 320 px, 360 px, 375 px, 390 px, 414 px, 480 px - 768 px, 834 px (iPad portrait / Pro 11) - 1024 px, 1280 px, 1440 px, 1920 px, 2560 px - Plus: 200% browser zoom, landscape orientation on every mobile width, iPhone with safe-area insets visible ## Hard layout rules - Mobile-first CSS. Default styles target mobile; `@media (min-width: ...)` for larger viewports. - Use `dvh` and `svh` instead of `vh` for full-height surfaces (iOS Safari URL-bar bug). - Use `clamp()` for fluid typography across all viewports. - Prefer container queries (`@container`) over media queries for component-level responsiveness. - Use `min(100%, ...)` widths so content never overflows. Zero horizontal overflow at any viewport. - Add `` to every page. - Apply `padding: max(safe-area-inset-X, fallback)` on every edge-bleeding container so notched iPhones in landscape never clip content. - Wide tables and code blocks scroll INSIDE their container (`overflow-x: auto`), never push the body. - Use `background-attachment: scroll` on mobile, not `fixed` (iOS Safari repaint bug). - Avoid `backdrop-filter` on animated elements. Use it sparingly on static surfaces only. - **Canvas Scaling**: Canvases must dynamically scale with window resize events and properly handle high-DPI screens (`window.devicePixelRatio`). Set physical dimensions (`canvas.width`/`canvas.height`) using pixel ratio and render relative to this grid, using CSS to control responsive viewport scaling. - **Robust Storage**: Every access to `localStorage`/`sessionStorage` (especially `JSON.parse` of loaded state or writes) MUST be wrapped in a `try-catch` block to handle disabled storage, private browsing mode, quota limits, or corrupted JSON gracefully. Fall back to a robust in-memory object store. ## Touch & accessibility - Tap targets ≥ 44 × 44 px on touch (Apple HIG). Increase to 48 px under `@media (hover: none) and (pointer: coarse)`. - All interactive controls reachable by keyboard with a visible focus ring; respect `:focus-visible`. - Color contrast ≥ 4.5:1 for body text, 3:1 for UI components. - All images have meaningful `alt`. Decorative images use `alt=""`. - Respect `prefers-reduced-motion: reduce` — zero animation durations under that query. - Forms validate inline; error messages are specific, not "Invalid input". - Modals: focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, focus restored on close. ## Performance bar (Lighthouse mobile, throttled 3G/4G) - LCP < 2.5 s · INP < 200 ms · CLS < 0.1 - JS bundle gzip < 200 KB mobile-first; lazy-load non-critical screens via `React.lazy` / dynamic imports. - No render-blocking resources above the fold. - Images: WebP/AVIF preferred, `loading="lazy"`, explicit `width`/`height` attributes (zero CLS), `srcset` for retina. - Videos: `preload="metadata"`, low-resolution poster, max 720p mobile fallback. Never autoplay with audio. - Fonts: `font-display: swap`; preload only the one used above the fold. - Smooth scroll honoured via CSS `scroll-behavior: smooth` with reduced-motion fallback. ## Pre-ship mobile checklist (the deployer MUST verify before declaring done) 1. Open at 375 px in DevTools — every screen scrolls vertically only; zero horizontal scroll. 2. Browser zoom 200% — layout reflows without overlap. 3. iPhone Safari with the URL bar visible AND landscape — no content under the home indicator; no notch clipping. 4. iPad portrait (768 px) and landscape (1024 px) — no awkward gaps; tablet-specific breakpoints land cleanly. 5. Tap every interactive element with a thumb at real-device size — every target is easy to hit. 6. `prefers-reduced-motion: reduce` — every transition / animation skips cleanly, scroll-behavior becomes instant. 7. Lighthouse mobile score ≥ 90 across all 4 categories. 8. Zero `console.error` and zero CLS shift in real-device testing on a mid-tier Android (e.g. Pixel 6a) and an iPhone SE. --- The original template starts below. All rules above apply on TOP of whatever this template specifies. --- # Heritage Class ## 1. Project **Heritage Class** is a live conversation tutor for the third generation — the teenager whose grandmother still speaks Punjabi (or Tagalog, or Korean, or Amharic, or Gujarati) but whose own thread to the language has thinned to a few greetings and a stubborn embarrassment. The user opens the app, picks "Punjabi — Sydney Sikh community, the way Nani speaks", and has a real voice conversation with a model that calibrates to their actual level in real time, uses the family's specific dialect, never shames a stumble, and saves a transcript so the next session begins where this one left off. By the end of a six-month run, the teen can phone Nani and have the kind of conversation that needed an aunt to translate before. This is the kind of app a fourteen-year-old Sikh-Australian girl in Sydney's western suburbs opens the night before her grandmother flies in from Punjab for the first time in two years — because last visit the conversation stalled at "I'm fine, how are you?" and she promised herself that this time it would be different. It is also the kind of app a Filipino-American teen in Daly City uses on the bus to school to practise the Tagalog her late grandmother spoke at home, with the specific Visayan-tinged code-mixing his mother grew up around — not the textbook Manila register a generic course would teach. Same shape of want, different language, different continent. The single demo that proves the magic: the teen taps "Start", says "Hi Nani-jee, kiddan?" into the microphone, and a warm voice answers in the family's specific dialect of Punjabi spoken in Sydney's Sikh community — "Bilkul vadhia, beta. Tu das, school kiveñ chal riha?" The model has already detected from those four syllables that the user is roughly A1 — comfortable with greetings, not yet with the past tense — and it slows its delivery, picks short sentences, and uses the family-honorific "jee" not the formal "sahib". The teen fumbles a reply. The model never says "wrong"; it offers the phrase back gently and waits. Two minutes later, the user has said five sentences they have never said aloud before. The transcript saves to "Nani Conversations" with the new phrases marked, ready for tomorrow. And in the heavier cases — where the heritage language is endangered, where the grandparent has already died and the teen is teaching themselves from old voice memos, where the dialect is one a textbook will never cover (Yiddish as spoken in their great-grandmother's Krakow household, Western Armenian carried out of Aleppo in 2014, Hokkien from a Penang family three generations into Sydney) — the app reads the curriculum the family already has: a folder of voice memos, a list of phrases the user typed, a recipe in the grandmother's handwriting. It grounds the live conversation in *that* corpus, not a generic textbook. The heritage class meets the heritage where it actually lives. **Tagline:** _Real conversations in the language your family actually speaks — in any dialect, at any level, without ever being made to feel small._ ## 2. Target audience - Third-generation diaspora teenagers learning the family language their parents stopped speaking at home — Punjabi in Sydney, Tagalog in Daly City, Korean in Toronto, Hokkien in London, Tamil in Singapore, Amharic in Washington DC, Igbo in Houston, Western Armenian in Glendale, Yoruba in Brixton, Vietnamese in Houston - Adult heritage learners who can understand 60% of what their parents say but can barely produce a sentence — the "passive bilingual" gap - University students taking a heritage-language placement test next month and panicking that they will be put in beginner with monolingual English-speakers - Long-distance grandchildren preparing for a planned video call with a grandparent who does not switch to English easily - Parents of fourth-generation kids who want to seed the language earlier than they themselves received it - Adoptees in transracial families reclaiming a heritage language they were never raised with — Korean-Americans adopted in the 80s, Chinese-Americans adopted in the 90s, Ethiopian children adopted into white families - Diaspora adults who returned to the homeland for one year and want to keep the spoken fluency alive after coming back - Heritage-language community-school teachers (Saturday-school instructors) looking for a between-class practice tool calibrated to their students - Endangered-language families — Western Armenian, Ladino, Wukchumni, Aymara, Quechua, Hokkien — for whom a generic Duolingo course will never exist ## 3. Core value propositions Surface these clearly through copy, visual emphasis, and section ordering — they are the reasons users pick this app. - **The family's specific dialect, not the textbook standard.** Punjabi as spoken in Sydney's Sikh community sounds different from standard Indian Punjabi which sounds different from Pakistani Punjabi which sounds different from Doabi or Majhi or Pothohari. Tagalog in Daly City carries forty years of English code-mixing that a Manila textbook will pretend doesn't happen. Mandarin in a Singapore Chinese-Australian household carries Hokkien substrate. The Live API call picks dialect from a one-time setup conversation with the user about *their* family, not from a country dropdown. - **A1 means A1, honestly.** No "intermediate" labels that flatter. The opening conversation measures the user against the CEFR A1/A2/B1/B2/C1 ladder for *production* (not comprehension — the heritage gap is asymmetric). The model calibrates every sentence length, every vocabulary choice, and every grammar concept to the actual level. When the user grows, the level moves with them. The number is shown clearly and updated transparently. - **Never shames a stumble.** Hard rule baked into the system instruction: never "wrong", never "no", never "you mean…". The model gently restates and waits. The voice register is warm-aunt, not teacher. The transcript later shows what the user said, what the model offered, what the user could try next — but never a row of red crosses. - **Live conversation calibrated in real time.** The Gemini Live API streams two-way audio at conversational latency. The teen speaks, the model answers in 600 ms, the conversation holds together for ten minutes. The model adjusts pace mid-conversation: if the user falls quiet for four seconds twice in a row, sentences shorten and the model offers a fork ("Want to keep going on school, or switch to talking about food?"). - **Curriculum grounding from the family's own corpus.** The user can drop in: voice memos from a grandparent, a list of family phrases, a typed glossary of household words ("we call this *pakhi*, not *fan*"), even a handwritten recipe. The Live API references *that* corpus in conversation: "Your Nani called this *sabzi* in the voice memo you uploaded — let's use that word, not *vegetables*". Authentic vocabulary, no textbook drift. - **The transcript is the textbook.** Every session saves a transcript that doubles as a personalised reader. New words are highlighted, new grammar patterns are explained in plain English in a side margin, and the user can re-listen to any line in the model's voice. Six months in, the transcript collection *is* the textbook the user could never have bought. - **Goal-oriented sessions.** The user picks a goal before each session: "Survive a 5-minute call with Nani about school", "Order food at Pho Tau Bay without switching to English", "Ask Lolo about the Visayan side of the family", "Tell Halmoni I miss her". The conversation drives to that goal. The transcript ends with the rehearsed sentences, ready to use in life. - **Endangered-language friendly.** When the language is one a standard model is shaky on (Western Armenian, Ladino, certain regional Punjabi registers, Hokkien), the app routes more heavily through the family-corpus grounding and surfaces honest confidence flags: "I'm less sure about this verb form — does your family say it like this?" — instead of confidently hallucinating. ## 4. Features to build - Live voice-to-voice conversation via Gemini Live API at conversational latency, push-to-talk and hands-free modes - A short one-time setup conversation that detects the user's heritage language, dialect, family-region context, and approximate level - Per-language CEFR-aligned level calibration (A1 → C1) recalculated every session, shown transparently on the home screen - Goal-picker at session start: "Survive a 5-min call with Nani", "Order food in the language", "Tell Halmoni about my exam", "Just chat for 10 minutes about anything", "Practise the past tense" - Real-time difficulty adjustment — sentence length, vocabulary range, grammar complexity, speech rate adjust mid-conversation based on user response latency and accuracy signals - Dialect picker that goes deeper than country: "Punjabi — Sydney Sikh community", "Tagalog — Daly City code-mixed with English", "Korean — banmal with halmoni", "Mandarin — Taiwanese Mandarin with Hokkien tags", "Amharic — Addis colloquial", "Yiddish — Galicianer family register" - Family-corpus upload — drop voice memos, photos of handwritten recipes, typed phrase lists; the corpus grounds every Live API session - Transcript view per session, with new vocabulary highlighted, new grammar patterns explained in a side margin in plain English - Replay-the-line: tap any line of the model's transcript to hear it again at the same warm pace - Per-language CEFR progress chart over time, weeks not days — heritage learning is slow and the app honours that - Family voice library — record the actual grandparent reading 30 short phrases (optional, consent-gated); the Live API model uses these as pronunciation anchors for the dialect (NOT a voice clone — see negative constraints) - Phrase backpack — phrases the user explicitly wants to take into a real conversation get pinned, with the model offering increasingly confident drills until the phrase is ready - "Call mode" — a short pre-call rehearsal flow: the user names who they are about to phone (Nani, Lolo, Halmoni), names the goal of the call, picks two phrases from the backpack, and runs a 3-minute live rehearsal of the call's opening - Session pacing — the model slows when the user falls quiet, offers a topic fork when stuck for ten seconds, gently ends at the agreed session length rather than dragging - Honesty signals — the model surfaces "I'm less sure about this form in your dialect; does your family say it this way?" when the family corpus contradicts the standard - Shame-free correction patterns — the model never says "wrong" or "no"; it offers the phrase back gently, waits, and only after two organic repetitions surfaces the pattern in plain English in the post-session notes - Calendar plant — schedule the next session for the user's commute or before-school window, with notifications respecting school hours and bedtime - Family invitations — invite the grandparent to record voice memos remotely via a magic link; their recordings auto-enrich the family corpus - Export per-session transcript as PDF for the heritage-language Saturday-school teacher - Offline-light mode — recent transcripts and the phrase backpack stay available without network; only the live conversation requires network ## 4b. Required Gemini capabilities + backend services **This template's intelligence comes from the Gemini capabilities below. Wire them up explicitly — don't substitute generic LLM calls.** ### Gemini capabilities (the load-bearing intelligence) - **Gemini Live API** (the headline capability) — bidirectional streaming audio at conversational latency, voice activity detection, interruption handling, in-session context updates. The Live API is what makes "real conversation" feel like real conversation. Session config carries the heritage-language code, the dialect markers, the level, the family-corpus references, and the goal of the session. - **Multilingual production at the dialect level** (Gemini 3.5 Flash for pre-session prep + system-prompt construction; Live API for in-session production) — Punjabi in Gurmukhi script with Sydney-diaspora vocabulary; Tagalog with Daly City code-mixing; Korean banmal between grandparent and grandchild; Tamil with Singapore Hokkien borrowings; Amharic in Addis colloquial register; Yiddish in the family's specific (Galicianer / Litvak / Hungarian) flavour; Hokkien in the Penang or Singapore variant; Western Armenian as carried out of Aleppo; Ladino in Sephardic family registers. - **Long context (1M tokens)** for grounding the family corpus — once a user has uploaded thirty voice memos, ten typed phrase lists, and three handwritten recipe photos, the entire corpus rides in the Live API session prompt. The model picks vocabulary from that corpus over textbook vocabulary. **Guardrail**: most corpora stay under 100k tokens; cap the family-corpus injection at 200k tokens per session prompt to leave headroom for the conversation itself. Chunk by recency + relevance if the corpus grows past the cap. - **Structured output / JSON Schema** — every session ends with a `SessionDebrief` JSON object scored against the CEFR ladder, surfacing new vocabulary attempted, grammar patterns the user touched, and recommended next-session targets. Schema below. - **Multimodal image input** (Gemini 3.5 Flash) for parsing handwritten family recipes, phrase notebooks, and other artefacts uploaded to the family corpus. Handwriting in non-Latin scripts (Gurmukhi, Devanagari, Hangul, Hebrew, Ge'ez, Hanzi) is read with explicit script identification. - **Gemini TTS** (`gemini-3.1-flash-tts-preview`) for transcript playback in the dialect — used outside the Live API session for the "replay this line" feature and for the family-voice-library pronunciation samples. Live API itself produces audio in real time; TTS is for after-the-fact review. - **Thinking levels** — `medium` for pre-session level calibration and session-debrief scoring (the model has to weigh production evidence honestly). `low` for the in-session real-time turn responses (latency budget). `low` for the family-corpus parse calls. ### Backend services - **Auth — Required.** Firebase Auth with Google sign-in (auto-provisioned by AI Studio Build). **Apple sign-in is optional but user-configured**: requires an Apple Developer account, Service ID, Key ID, and private key wired into the Firebase Auth console. **Magic-link email** (used to invite a grandparent to contribute voice memos) requires the sender domain to be authorised in Firebase Auth. Sessions are private to the owner; family-corpus contributions are explicitly shared via invitation. - **Database — Required.** Firestore for `users`, `heritage_profiles`, `family_corpus_items`, `sessions`, `transcripts`, `phrase_backpack`, `level_history`, `goals`, `invitations`. - **File storage — Required.** Firebase Storage for uploaded voice memos (consent-gated retention), handwritten recipe photos, family voice library recordings. **Storage is NOT auto-provisioned by AI Studio Build today** — enable it in the Firebase console and wire the bucket name into the AIS Build project before first audio upload. Pre-signed URLs only; voice memos are never publicly addressable. - **Email — Required (transactional).** Family-corpus invitations via magic link (Firebase Auth magic links). Weekly progress summary emails (opt-in). Saturday-school teacher reports (opt-in, attachment). - **Payments — Not needed for v1.** Free for personal use. A future "school edition" with classroom dashboards could justify a paid tier, but the v1 product never paywalls a teen's access to their own heritage language. - **External APIs:** Gemini API for all intelligence. The Live API is the long-running session driver; standard `generateContent` for pre-session prep and post-session scoring. **Environment variables:** every secret (Gemini API key, Firebase service-account JSON, Stripe key if a paid tier is added) lives in environment variables — never in client bundle. Include a `.env.example`. **Auth + data privacy reminders:** never log secrets · never store passwords in plain text · use HTTPS everywhere · honour 'delete my account' inside the UI · explicit opt-in for any analytics · the user's voice and conversation transcripts are never sent to Gemini for model training (use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms) · grandparent voice memos require explicit recorded consent before upload, with a clear "delete forever" option that wipes them from Storage within 60 seconds · the family voice library is NOT used to clone a grandparent's voice into TTS, only as pronunciation reference inside Live API session prompts (see negative constraints). **Read this first — prompt-craft rules that apply to every call in this template:** 1. **Name the model variant explicitly** in every Gemini API call. Do not let the agent pick the model. See the per-call matrix below. 2. **Pin `thinkingLevel` explicitly** per call. See the matrix. 3. **Seed the JSON Schema as a fenced TypeScript / Zod block** in the system instruction or `responseSchema` field for structured-output calls. The literal schema is below. **Convert the Zod schema to Gemini's `Schema` type via the SDK helper** before passing to `responseSchema` — do NOT pass raw Zod. **Numeric `min`/`max` constraints are documentation only inside `responseSchema`; clamp on the server after the response arrives.** 4. **Pin the system instruction separately** from user input. Use the `systemInstruction` field for persona + behavioural rules; use `contents` for user input. Never concatenate. For the Live API, system instruction is set at session start in the session config — not per turn. 5. **Pre-declare tools as an enable/disable list** per call. The matrix below names which tools are enabled per call. Tools NOT listed for a call should be disabled. 6. **State negative constraints explicitly** — they are listed below. They are NOT "be careful" suggestions; they are hard rules the model must follow. 7. **Strip unsupported Zod modifiers before passing to `responseSchema`** — Gemini's OpenAPI subset rejects `.regex()` / `pattern`, fixed-length `z.tuple()`, and other custom validators. Use a sanitizer that flattens tuples to length-2 arrays and removes regex patterns before serializing. Validate those constraints in middleware AFTER parsing. ### Per-call model + tools matrix | Call | Model | thinkingLevel | Tools enabled | |------|-------|---------------|---------------| | Setup conversation — detect language, dialect, level | `gemini-3.1-flash-live-preview` (Live API) | n/a | (none — Live API does not accept thinkingConfig) | | Heritage profile finalisation → `HeritageProfile` schema | `gemini-3.5-flash` | medium | (none) | | Live conversation session — dialect + level calibrated | `gemini-3.1-flash-live-preview` (Live API) | n/a | (none — Live API does not accept thinkingConfig) | | Post-session debrief → `SessionDebrief` schema | `gemini-3.5-flash` | medium | (none) | | Parse handwritten family recipe / phrase notebook | `gemini-3.5-flash` | medium | (none) | | Family corpus chunk + embed for in-session grounding | `gemini-3.5-flash` | low | (none) | | Replay-the-line TTS playback | `gemini-3.1-flash-tts-preview` | n/a | n/a | | Phrase-backpack drill scoring | `gemini-3.5-flash` | low | (none) | *Note for builders:* on Live API and TTS calls, omit `thinkingConfig` entirely — the field is not supported on those models. The `n/a` cells in this matrix are documentation only; do not serialise them into the request body. ### Primary structured-output schemas (seed these verbatim in the prompt) ```typescript import { z } from "zod"; const CefrLevel = z.enum(["A1", "A2", "B1", "B2", "C1", "C2"]); const DialectMarker = z.object({ marker_type: z.enum([ "lexical", // vocabulary choice "phonological", // sound shift "morphological", // word-form change "syntactic", // sentence structure "code_mixed_with", // diaspora code-mixing target "honorific_register", // who uses what to whom ]), description: z.string(), // "uses 'jee' as warm honorific not formal 'sahib'" example_phrase_verbatim: z.string(), // in native script example_phrase_romanised: z.string().nullable(), }); const FamilyCorpusItem = z.object({ item_id: z.string(), source_type: z.enum([ "voice_memo_from_relative", "user_typed_phrase_list", "handwritten_recipe", "handwritten_letter", "transcribed_phone_call", "other", ]), contributor_name: z.string().nullable(), // "Nani Harpreet" relationship_to_user: z.string().nullable(), // "maternal grandmother" storage_uri: z.string(), // gs:// URI duration_seconds: z.number().nullable(), text_extracted: z.string().nullable(), // transcribed or OCR'd language: z.string(), // BCP-47, "pa-IN" with dialect notes elsewhere dialect_notes: z.string().nullable(), consent_recorded_at: z.string().nullable(), // ISO timestamp of explicit consent recording }); const HeritageProfile = z.object({ profile_id: z.string(), user_id: z.string(), language_canonical: z.string(), // "Punjabi" language_bcp47: z.string(), // "pa-IN" script: z.enum([ "gurmukhi", "shahmukhi", "latin", "hangul", "hanzi_traditional", "hanzi_simplified", "devanagari", "hebrew", "geez", "arabic_nastaliq", "tamil", "khmer", "thai", "cyrillic", "other", ]), dialect_label_user_facing: z.string(), // "Punjabi — Sydney Sikh community" dialect_markers: z.array(DialectMarker), family_region_context: z.string(), // "Sikh diaspora, Sydney western suburbs; family from Doaba region originally" // ASYMMETRIC HERITAGE LEVELS production_level: CefrLevel, // can the user speak? comprehension_level: CefrLevel, // can the user understand? reading_level: CefrLevel.nullable(), // script literacy is often null production_level_confidence: z.number().min(0).max(1), comprehension_level_confidence: z.number().min(0).max(1), setup_evidence_quotes: z.array(z.object({ user_said_verbatim: z.string(), interpretation: z.string(), contributes_to: z.enum(["production_level", "comprehension_level", "dialect_marker"]), })), goal_horizon: z.enum([ "weekend_with_grandparent", "weekly_video_calls", "moving_to_homeland_in_months", "saturday_school_placement", "heritage_language_class_at_university", "keep_a_dying_language_alive", "no_specific_goal_yet", ]), family_corpus_item_count: z.number().min(0), endangered_or_low_resource: z.boolean(), // routes more heavily through family corpus }); type HeritageProfile = z.infer; const AttemptedUtterance = z.object({ user_said_verbatim: z.string(), // exact transcription of what the user said attempted_meaning: z.string(), // what they were trying to say level_of_attempt: CefrLevel, what_landed: z.string(), // verbatim portion that was correct what_drifted: z.string().nullable(), // verbatim portion that drifted from target model_gentle_restatement: z.string(), // what the model said back, no commentary model_did_NOT_say: z.literal("nothing in the 'wrong/no/incorrect' family"), }); const NewPhraseEncountered = z.object({ phrase_in_script: z.string(), phrase_romanised: z.string().nullable(), english_gloss: z.string(), source: z.enum([ "from_family_corpus_voice_memo", "from_family_corpus_recipe", "from_family_corpus_typed_list", "model_offered_for_dialect_fit", "user_introduced", ]), source_quote: z.string().nullable(), // "your Nani used this word in the chicken curry memo" in_phrase_backpack: z.boolean(), // user pinned for drill? }); const SessionDebrief = z.object({ session_id: z.string(), started_at_iso: z.string(), ended_at_iso: z.string(), duration_seconds: z.number().min(0), goal_set_at_start: z.string(), // "Survive 5-min call with Nani about school" goal_reached: z.boolean(), goal_progress_summary: z.string(), // one paragraph, warm voice starting_production_level: CefrLevel, ending_production_level: CefrLevel, // may equal starting; that's fine level_movement_evidence: z.array(z.string()), attempted_utterances: z.array(AttemptedUtterance), new_phrases_encountered: z.array(NewPhraseEncountered), family_corpus_items_referenced: z.array(z.string()), // item_ids next_session_recommendation: z.object({ suggested_goal: z.string(), suggested_horizon_days: z.number().min(1).max(14), rationale_for_user: z.string(), // shown to user, ≤ 2 sentences, no jargon }), shame_check: z.literal("verified: no 'wrong/no/incorrect' language used in this session"), }); type SessionDebrief = z.infer; const PhraseBackpackEntry = z.object({ entry_id: z.string(), phrase_in_script: z.string(), phrase_romanised: z.string().nullable(), english_gloss: z.string(), reason_user_added: z.string(), // "I want to ask Nani about her trip" drill_count: z.number().min(0), last_drilled_iso: z.string().nullable(), confidence_score: z.number().min(0).max(1), // user-self-rated, not model ready_for_real_use: z.boolean(), }); ``` ### Common failure modes (and how to avoid them) - Agent picks `gemini-3.5-flash` for the Live conversation — wrong tool. Live API is its own model family. Pin `gemini-3.1-flash-live-preview` for the session. Standard `generateContent` is for prep, scoring, and grounding — never for the live turn. - Model uses "wrong" or "no" or "incorrect" or "actually it's…" during the live conversation. Hard rule pinned in the session-config system instruction: never use that family of words. The model offers the phrase back gently and waits. - Model defaults to standard textbook variant (Indian standard Punjabi, Manila Tagalog, Seoul standard Korean) when the user specified a diaspora dialect. The setup-call must elicit dialect markers and pin them in the live session config; the live system instruction must reference the specific markers, not just "Punjabi". - Model rushes — speaks at podcast pace when the user is A1. The session config must set a target speech rate of ~110 words per minute for A1, faster only as level rises. Long sentences are broken into short ones. Pauses are honoured. - The session "improves" the user's level prematurely. The CEFR movement happens in the post-session debrief, scored from production evidence — not optimistically pumped because the user laughed once. - Family-corpus grounding ignored mid-conversation. The session config must inject the corpus chunks the user explicitly tagged with the session's goal. The model must prefer corpus vocabulary over generic vocabulary. Add a unit test: a user with a recipe memo using "*sabzi*" should never hear the model say "vegetable" or "*subzi*" with the textbook spelling. - Family voice memos used to clone the grandparent's voice into TTS. Hard rule: the family voice library is a *pronunciation reference for the Live API session prompt*, NEVER a voice clone. The TTS replay-the-line uses a generic-locale voice, not the grandparent's voice. Sample-based voice cloning of a real living or recently-living person without explicit, recorded, ongoing consent is not in scope for this product. - Model continues past the agreed session length. The session config carries the agreed end time. The model gently wraps two minutes before, summarises, and ends — does not push for an extra ten minutes. - Endangered-language hallucination. For low-resource languages (Western Armenian, certain Ladino registers, regional Hokkien variants), the model invents grammar with full confidence. The session config must require honesty flags: "I'm less sure about this form in your dialect — does your family say it this way?" — not silent confidence. - The post-session debrief shames. The schema enforces a `shame_check` literal that requires no "wrong/no/incorrect" language in the transcript. The debrief speaks in warm voice: "Today you said five sentences you hadn't said before. The past tense came close twice." ### Negative constraints (hard rules) - Do NOT use the words "wrong", "no", "incorrect", "actually", "you mean", or any near-synonym at any point in the live conversation. The model gently restates and waits. - Do NOT default to textbook standard when a diaspora dialect is specified. The dialect markers in the heritage profile are load-bearing. - Do NOT exceed conversational pace for the user's level. A1 ≈ 110 wpm with short sentences. B1 ≈ 140 wpm with mid-length sentences. Speeding up against the user's level is a worse failure than going too slowly. - Do NOT extend the session past the agreed length. Gently wrap two minutes before; never push. - Do NOT clone a grandparent's voice into TTS. The family voice library is a pronunciation reference inside the Live API session prompt only. The TTS replay-the-line uses a generic locale voice. - Do NOT score the user's level upward without production evidence. The CEFR ladder moves on what the user *said*, not on what they understood. - Do NOT translate proper nouns: family names, dish names, neighbourhood names, festival names. "Nani-jee", "*sabzi*", "*halo-halo*", "Diwali", "Halmoni" stay in the heritage language. An English gloss appears in the transcript margin only the first time, never thereafter. - Do NOT use the user's voice or conversation transcripts to train any model. Use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. The capabilities-info panel says this in plain English. - Do NOT auto-publish or auto-share session transcripts. They are private to the user. The Saturday-school teacher export is opt-in, per-session. - Do NOT collect voice memos from a grandparent without explicit recorded consent. The family-corpus upload flow records consent before any audio is sent to Gemini. - Do NOT continue conversation when the user is silent for >15 seconds. The model offers one fork ("Want to switch to talking about food instead?") and then waits, then ends gracefully. Heritage learning is asymmetric: comprehension can comfortably exceed production by a level or two, and the model should never punish a long thinking pause. - Do NOT flag a "mistake" that is in fact the family's correct dialect form. If the user says *pakhi* and the textbook says *fan*, the model uses *pakhi*. The family corpus is the ground truth, not the textbook. ### Per-call `systemInstruction` strings Use these as the literal `systemInstruction` field for each Gemini API call the built app makes. They complement the series-wide rules already uploaded as the global instructions file (`00-series-instructions.txt`). ### Call: Setup conversation — detect language, dialect, level Model: `gemini-3.1-flash-live-preview` (Live API) · n/a · Tools: (none) ``` You are running a one-time setup conversation with a heritage-language learner. Your job is to detect: which language they are learning, which specific dialect their family speaks, their approximate production level, their approximate comprehension level, their goal horizon. The conversation is in English (the user's dominant language) by default; you will switch to the heritage language only briefly to test comprehension when the user is comfortable. Tone: warm aunt, not teacher. Curious about their family. Never clinical, never assessment-test. Open with a short, easy question: "Whose language are you here to learn?" Wait. Listen carefully. Distinguish: - a named relative ("my Nani", "my Halmoni", "my Lola", "my Yiayia") — this is the heritage anchor. Note the relationship word — the relationship word itself is dialect data. - a country name without a relative — gently follow up: "is this from a grandparent, a great-grandparent, or further back? It helps me pick the right register." - an endangered or low-resource language (Western Armenian, Ladino, Hokkien, Wukchumni, Aymara) — flag internally that you will need to lean heavily on family-corpus grounding once provided. Build up the dialect picture from at least three signals: 1. The honorific the user uses for the relative ("Nani-jee", "Lolo ko", "Halmoni" with -ya, "Yiayia mou"). Record verbatim. 2. The neighbourhood / city where the relative learned the language. Do not assume modern country borders — ask gently. "And was this the language she grew up with at home, or one she picked up later?" 3. One dish name or festival name the user knows in the language. The dish name carries enormous dialect signal — "halo-halo" vs "ginataang bilo-bilo" vs "kakanin"; "saag" vs "sarson da saag" vs "haak"; "kimchi" vs "kkakdugi" vs "yangbaechu kimchi". Now estimate production. Ask: "Could you say a sentence — anything, even one you're embarrassed by — in the language?" Wait. What the user says is gold-standard A1/A2/B1 evidence. Do not push for more if the user shows discomfort; one sentence is enough. Praise honestly and specifically: "you got the verb-final word order right" not "great job". Estimate comprehension separately. Switch to the heritage language briefly and ask one short, friendly question — a greeting plus a simple "how are you?" — and listen to how the user reacts. If they answer in English, that's data. If they answer in the heritage language with hesitation, that's data. If they answer fluidly, revise production estimate upward. Set goal_horizon by asking: "Why now? Is there a person, a trip, a conversation you have in mind?" End the setup conversation after no more than 6 minutes. Hand off to the structured-output call that will populate the HeritageProfile schema. Hard rules: - Never use "wrong", "no", "incorrect", "actually", "you mean". - Never sound like a test. The setup is a chat with a curious aunt who wants to know the family. - Never default to a country dropdown. Always ground in the specific relative and their specific community. - For endangered or low-resource languages: tell the user honestly that "for this language, your family's voice memos and phrases will be even more important than usual — I'll lean on whatever you upload". No commentary outside the live conversation. ``` --- ### Call: Heritage profile finalisation → `HeritageProfile` schema Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none) ``` You receive the full transcript of a setup conversation between a heritage learner and a warm Live API guide. Your job is to populate the HeritageProfile schema with the evidence in front of you. Hard rules: - Asymmetric levels. Heritage learners almost always have comprehension ≥ production by one or two CEFR levels. Score them separately. Set production_level honestly low if the user only produced greetings. - Dialect markers come from the transcript, never from a country default. If the user said "Nani-jee", that's a dialect marker (north Indian Punjabi diaspora warm honorific). If the user said "halo-halo", that's a dialect marker (Filipino diaspora dessert vocabulary). Populate dialect_markers[] with at least one marker per signal-type the transcript supports; leave empty rather than invent. - setup_evidence_quotes ties each level estimate to a verbatim user utterance. If you cannot find a verbatim quote for a level estimate, lower confidence accordingly. - family_region_context combines the relative, the city / neighbourhood, and the inferred ancestral region. Example: "Sikh diaspora, Sydney western suburbs; family from Doaba region originally". Never claim a region that wasn't named in the transcript. - endangered_or_low_resource is true for: Western Armenian, Ladino, Hokkien in Penang / Singapore registers, regional Punjabi (Pothohari, Multani), Wukchumni, Quechua, Aymara, Garifuna, Yiddish in any living family variant, certain creoles. When true, the in-session prompt will lean more heavily on the family corpus. - goal_horizon comes from the user's stated "why now". If the user did not state a why, set "no_specific_goal_yet" rather than guess. Output ONLY the HeritageProfile JSON matching the provided schema. No commentary. JSON only. ``` --- ### Call: Live conversation session — dialect + level calibrated Model: `gemini-3.1-flash-live-preview` (Live API) · n/a · Tools: (none) ``` You are a warm, patient conversation partner in {language_canonical} ({dialect_label_user_facing}). The learner you are speaking with is {user_name}, who set production level {production_level} and comprehension level {comprehension_level} in their setup. Today's session goal: {goal_set_at_start}. Today's session length: {duration_minutes} minutes. DIALECT BINDING — these markers are non-negotiable for this session: {dialect_markers_block — bullet list of verbatim markers from HeritageProfile.dialect_markers[]} FAMILY CORPUS — these items are available; PREFER their vocabulary over textbook vocabulary in your responses: {family_corpus_chunks_block — selected items relevant to the goal, each with contributor name and source quote in the heritage language} Tone: warm, like a great-aunt at a kitchen table. Curious about the user, never clinical. The user is here to speak their family's language with you; you are honoured to be the practice partner. Pace at {target_words_per_minute} wpm for level {production_level}. Sentences are short. The user is allowed long pauses. If the user falls quiet for more than 6 seconds, wait. If they fall quiet for more than 12 seconds, gently offer a topic fork: "Want to keep going on this, or switch to something else?" — and wait. Hard rules (these are not "guidelines"; they are constraints): 1. NEVER use the words "wrong", "no", "incorrect", "actually", "you mean", or any near-synonym. NEVER. When the user says something that drifts from the target, you offer the phrase back gently: "Ah — I'd say: [phrase]. Want to try that again?" Then you wait. If the user does not pick it up, you let it pass — the same phrase will come around again later. There are no red marks in this conversation. 2. NEVER speak faster than the target WPM for this level. NEVER use sentences longer than the level allows: A1 ≤ 8 words, A2 ≤ 12, B1 ≤ 18, B2 ≤ 25, C1 + natural length. 3. PREFER vocabulary from the family corpus. If the corpus has the user's Nani calling something "sabzi", you say "sabzi" — never "vegetable" in English, never the textbook standard spelling. 4. HONOUR the asymmetric level. The user may understand more than they can say. If they answer in the heritage language with a fragment, take the fragment seriously and respond in full sentences they can still parse — do not match their fragment length. 5. SURFACE HONESTY for endangered or low-resource forms. If you are uncertain whether a verb form is the family's dialect, say so: "I'd say it like this — does your Nani say it this way too?" 6. PROPER NOUNS stay in the heritage language. Names of family, dishes, neighbourhoods, festivals are never translated mid- conversation. If the user uses a proper noun unfamiliar to the corpus, ask gently: "Is that a family name, or a place?" 7. END GRACEFULLY at the agreed session length. Two minutes before the end, name it: "We've about two minutes — want to finish on the [goal] sentence?" Wrap warmly. Never push for an extra ten minutes. 8. SHAME CHECK. If you find yourself about to say anything that feels like correction, STOP. Reframe: "Here's how I'd say it — want to try?" Switch to English ONLY when: - the user explicitly asks ("how do I say…") - a safety / medical / emergency situation arises - the user expresses distress Never switch to English to make a grammar point. Stay in the heritage language; offer the phrase, wait, move on. No commentary outside the live conversation. ``` --- ### Call: Post-session debrief → `SessionDebrief` schema Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none) ``` You receive the full audio transcript of a heritage-language conversation session. Your job: populate the SessionDebrief schema with honest, warm-voiced scoring. Hard rules: - Level movement requires production evidence. If the user only produced A1 utterances during the session, ending_production_level ≤ A2 and only if there is concrete evidence of A2 production. Do not pump levels to feel good. - attempted_utterances captures every meaningful production attempt. user_said_verbatim is the exact transcription, including hesitations and false starts. what_landed quotes the portion that matched the target. what_drifted is null when nothing drifted — do not invent drift. - model_did_NOT_say is a literal constant. If the session contains any of the forbidden words ("wrong", "no", "incorrect", "actually", "you mean"), set the shame_check literal to a fail string and flag it for review — do not silently pass. - new_phrases_encountered prefers family-corpus sources. When a phrase came from the corpus, include the source_quote verbatim ("your Nani used this in the chicken curry memo on 2026-04-12"). - next_session_recommendation is concrete and warm. Bad: "continue practising". Good: "Next time, try the call-mode rehearsal — pick one phrase from today and aim to use it with Nani this weekend." - goal_progress_summary speaks to the user directly, in their preferred name. Tone: "Today you said five sentences you hadn't said before. The past tense came close twice. The sentence about your sister was beautiful — say it again next time and it will stick." - shame_check is a literal constant when the session is clean. If the audit fails, set the failure flag and explain in level_movement_evidence which utterance violated. Output ONLY the SessionDebrief JSON matching the provided schema. No commentary. JSON only. ``` --- ### Call: Parse handwritten family recipe / phrase notebook Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none) ``` You receive one or more photographs of a handwritten artefact a heritage learner uploaded to enrich their family corpus. The artefact might be: a recipe in the grandmother's handwriting, a typed phrase list, a child's handwritten greeting card, a postcard in the heritage language, a page from a family prayer book. Identify the script: Gurmukhi, Shahmukhi, Devanagari, Hangul, Hanzi (traditional or simplified), Hebrew, Ge'ez, Arabic Nastaliq, Tamil, Khmer, Thai, Cyrillic, or Latin (romanised). State the script explicitly in the response. Transcribe verbatim. Preserve every diacritic. Polish ł, ą, ę; Hindi candrabindus; Korean batchim consonants in their original Hangul positions; Yiddish hekher pasekh; Arabic shadda and tanween; Vietnamese tone marks. Render in the exact Unicode character. For each transcribed phrase, provide: - the phrase verbatim in its native script - a romanisation appropriate for the language (Hunterian for Hindi, Revised for Korean, pinyin for Mandarin, YIVO for Yiddish, IAST for Sanskrit, jyutping for Cantonese, Pe̍h-ōe-jī for Hokkien) - an English gloss for the descendant who may not read the script - dialect notes if the spelling or word choice is non-standard Do NOT modernise. If the recipe says "ghee" with a regional spelling variant, keep the spelling. If a phrase uses an honorific form the textbook would consider archaic, keep the honorific. Do NOT translate proper nouns: dish names, family names, place names, festival names. Add a parenthetical English gloss on first occurrence only. If parts of the page are illegible (faded ink, fold creases, torn edges), mark them as [illegible: 3 words] inline rather than inventing. Output a structured JSON list of FamilyCorpusItem-compatible phrases. No commentary outside the JSON. ``` --- ### Call: Family corpus chunk + embed for in-session grounding Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You receive a FamilyCorpusItem and the goal of an upcoming conversation session. Your job: extract the chunks of the corpus item most relevant to the goal, format them for injection into the Live API session prompt, and return them as a tight list. Hard rules: - Quote, do not summarise. The Live API session needs the verbatim phrasing — that is the dialect signal. - Preserve the contributor's name with each chunk: "Nani Harpreet, voice memo 2026-04-12, talking about Vaisakhi preparations". - Cap total injected corpus text at 8,000 tokens per session prompt. If more is relevant, rank by recency × goal-relevance and trim. - Honour endangered_or_low_resource flag: when true, include more corpus, even if the goal-relevance is marginal — the corpus is the dialect's only honest reference. - Strip any consent-revoked items. If consent_recorded_at is null for a voice-memo-from-relative, do NOT include it. Output a JSON list of {speaker, source_label, quote_verbatim, goal_relevance_score} objects. No commentary outside the JSON. ``` --- ### Call: Replay-the-line TTS playback Model: `gemini-3.1-flash-tts-preview` · n/a · n/a ``` Voice: warm, unhurried, in the locale of the heritage language. Pick the Gemini 2.5 Flash TTS voice whose `languageCode` matches the heritage language's BCP-47 code — pronunciation will follow that locale automatically. Prefer a female voice for grandmother dialogue contexts, a male voice for grandfather contexts, where both are published for the locale; fall back to whichever is available rather than blocking. Pre-process the text before sending it to TTS: - Render in the native script (Gurmukhi, Hangul, Hanzi, etc.) — TTS pronunciation is tied to the script + locale combination. - At each sentence break, insert a single ellipsis (`…`) so the TTS model produces a natural pause. At paragraph breaks, insert a blank line plus an em-dash (`—`). Gemini 2.5 TTS does not support SSML `` — these textual cues are how you signal pace. - Target rate: ~110 words per minute for A1 / A2 learners, ~140 wpm for B1+ — letter-reading pace, not podcast pace. Style direction: prepend ONE short directive sentence to the text input, exactly like: "Read this warmly and slowly, as a grandmother might say it to a grandchild who is still learning. …". There is no separate `style` API field on Gemini 2.5 TTS; the directive sentence inside the input is how style is conveyed. DO NOT use this call to clone a grandparent's voice. The family voice library is a pronunciation reference inside the Live API session prompt only. This TTS call uses a generic locale voice. Phoneme overrides (Punjabi tonal markers, Korean batchim, Mandarin tones, Vietnamese tone marks) are NOT exposed by Gemini 2.5 TTS — no SSML `` tag. Pronunciation comes from the chosen voice's native locale. ``` --- ### Call: Phrase-backpack drill scoring Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You receive: a PhraseBackpackEntry, the user's spoken attempt (verbatim transcription), and the entry's prior drill_count. Your job: return a small JSON object with: - portion_matched: verbatim portion of the user's attempt that matched the target phrase - portion_drifted: verbatim portion that drifted (or null) - gentle_restatement: how a warm aunt would offer the phrase back - ready_for_real_use: boolean, true only when the user has produced the phrase fluidly across at least 3 drills without drift Hard rules: - Never use "wrong", "no", "incorrect", "actually", "you mean". - gentle_restatement is in the heritage language, not English, unless the user explicitly asks for English. - ready_for_real_use is honest. Do NOT pump it just to feel good. Heritage learners do not need false confidence; they need the real thing. Output JSON only. No commentary. ``` ## 5. Use cases & content to include Build dedicated UI sections or flows for each of these — they tell you what content the app must support. - **Nani is flying in this weekend.** A Sikh-Australian fourteen-year-old in Sydney opens the app on Thursday night. Her grandmother lands Sunday. She picks goal: "Survive a 30-minute conversation with Nani about school". The app runs three short sessions Thursday, Friday, Saturday. By Sunday, she has rehearsed seven sentences in the Sydney Sikh community register of Punjabi — including the "jee" honorific Nani prefers and the specific word *pakhi* for "fan" that the family uses. - **The Daly City Tagalog session on the bus.** A Filipino-American sophomore takes the 38R to school. Twenty-five minutes each way. He plugs in his AirPods and runs a ten-minute morning session and a ten-minute afternoon session. The conversations use Daly City code-mixed Tagalog — English for the technical words his late Lola used, Tagalog for the family terms. The family corpus he uploaded includes a recipe his Lola wrote out for kare-kare and a voicemail she left him three years ago. - **Halmoni's birthday call.** A Korean-Canadian teen in Toronto prepares for her grandmother's seventieth birthday video call. She picks goal: "Sing happy birthday in Korean and ask Halmoni about her trip to Jeju". She uses banmal correctly (the casual register a grandchild uses with a grandmother in many families, which textbooks under-cover). The app schedules a 5-minute rehearsal the hour before the call. - **The Saturday-school placement.** A Sri Lankan-Tamil teenager in London is starting Tamil Saturday school next month. She does not want to be placed in the absolute beginner class, but she also does not want to overshoot. She runs a one-time placement session and exports the SessionDebrief PDF to the Saturday-school teacher. The teacher sees: comprehension B1, production A2, dialect Jaffna Tamil with British English code-mixing. - **The endangered-language family.** A Western Armenian family in Glendale; the grandmother passed away two years ago; the teen has fifteen voice memos of his Medzmama telling stories. He uploads the memos. The Live API sessions are heavily grounded in his Medzmama's specific vocabulary — including the family's specific pronunciation of certain consonants that differs from standard Western Armenian. Honesty flags appear often ("I'd say it like this — but your Medzmama said it slightly differently in the wedding memo; want to use her version?"). - **The Filipino-American adoptee.** A 19-year-old adopted at six months by a white American family is reclaiming Tagalog with no living family link. She picks "Tagalog — Manila standard register" because she has no family dialect to anchor to. The app honours that. It does not invent a family. It runs straight Manila-standard Tagalog and does not push for a heritage-corpus upload she does not have. The shame-free rule applies just as hard. - **The Vietnamese-Australian dinner table.** A 16-year-old in Melbourne whose parents arrived as boat people from Saigon in 1981. He wants to speak Vietnamese at the family Tet meal without his older sister translating. He picks Southern Vietnamese (not the Hanoi register most apps default to). The app calibrates to Southern-register tones and lexical choices. The family corpus includes recordings of his uncle telling the family escape story. - **The Mandarin-with-Hokkien-substrate household.** A Chinese-Singaporean-Australian teen in Sydney whose grandparents speak Mandarin sprinkled with Hokkien tags. He picks "Taiwanese / Singapore Mandarin with Hokkien substrate". The app honours the substrate words his Ah Gong uses for food. - **The post-call cool-down.** After a five-minute live session, the user gets a SessionDebrief that names the five sentences they said for the first time. The phrase backpack gets two new pinned phrases. The next-session recommendation lands in their calendar: "Friday morning, 8:00, on the bus — try the past tense about your weekend." - **The Yiddish family corpus.** A young adult in Brooklyn is rebuilding her great-grandmother's Galicianer Yiddish from a folder of audio cassettes her grandfather digitised before he died. She uploads twenty-three transcribed memos. The app runs sessions in the Galicianer register — not standard YIVO Yiddish, not Hungarian Yiddish — and surfaces honesty flags often. - **The Amharic-DC commute.** A first-generation Ethiopian-American student in Washington DC speaks comfortable Amharic at home but can never produce it outside the kitchen. She runs short sessions on her Metro commute. The dialect is Addis colloquial. The app uses the honorific *ato* and *weizero* correctly across generations. - **The Igbo-Houston Saturday morning.** A Nigerian-American mother of two wants to seed Igbo for her kids before they're old enough to refuse. She runs her own sessions in the early morning. The dialect is her parents' Anambra Igbo. Within three weeks she has a phrase backpack of forty sentences she can use at the family WhatsApp call. ## 6. Page structure Build the following screens / sections in this order. Adjust copy to fit the voice, but keep the structural intent. 1. **Welcome / sign-in.** A photographed-looking image of a teenager at a kitchen table with AirPods in, a phone on the table, a teapot in the background and a wall calendar showing a date circled (the grandmother's arrival date). One paragraph: "Heritage Class is real conversation in the language your family actually speaks — at your real level, in your family's specific dialect, without ever being made to feel small." Single Google sign-in button; Apple sign-in next to it. Below: "Try with the sample family" → loads the demo profile in section 8a (Sydney Punjabi A1). 2. **First-visit setup.** A short live conversation (see section 6b). The app talks the user through whose language they're learning, with the warm-aunt tone — never a form, never a country dropdown. 3. **Home — Today.** Top card: the user's current production level (e.g. "A1 → reaching for A2") and comprehension level, with a one-line warm explanation. Below: today's suggested session (one card with goal + duration + suggested time) and a small row of recent transcripts. A "Start a session now" primary button. A "Plant a session in my calendar" secondary button. 4. **Goal picker (session start).** Four-by-two grid of goal cards: "Survive a 5-min call with [relative name]", "Order food in the language", "Tell [relative] about my [event]", "Just chat about anything", "Practise the past tense", "Practise asking questions", "Roll the call-mode rehearsal", "Pick from phrase backpack". Each card carries a small icon and a one-line explanation. Default duration slider: 5 / 10 / 15 minutes. 5. **Live session.** Full-bleed minimal UI. A breathing audio-level orb at centre, a small timer in the top-right counting down to the agreed end, a small "switch to text" affordance for moments the user cannot speak aloud (on the bus, in class). Captions stream below the orb in the heritage language, with a translation toggle. The model's pace and sentence length adjust live; visible UI never says "you got it wrong" — ever. A subtle pause-to-think indicator appears when the model is waiting on the user; never aggressive. 6. **Session debrief.** Three sections, scrollable. *What you said today* — verbatim utterances in the heritage language with English glosses on tap. *What you reached for* — phrases the user attempted and what landed vs what drifted, in warm voice. *Where you are* — the CEFR level chart with the new point plotted, with the evidence behind the movement explained in one sentence ("today's past-tense sentence about your sister moved your production estimate from A1 to A1+"). At the bottom: "Add 2 phrases to backpack" and "Plant next session". 7. **Transcript view.** A reading-mode rendering of any past session. Each line in the heritage language with the English gloss in a side margin; new vocabulary highlighted; grammar patterns explained in plain English in marginal notes. Tap any line to replay it in the model's voice. A small "export PDF for my teacher" affordance. 8. **Phrase backpack.** A vertical list of pinned phrases, sorted by readiness. Each entry: phrase in script, romanisation, English gloss, the reason the user added it, a confidence pill ("ready", "almost there", "still drilling"). Tap any entry to run a 60-second drill. A "Ready for real use" badge appears only when the user has produced the phrase fluidly across at least three drills. 9. **Family corpus.** A view of every voice memo, recipe photo, phrase list, and transcribed call the user has added. Each item shows: contributor name + relationship, source type, language + dialect notes, consent status, and a "delete forever" button. Top: "Invite a family member to add a voice memo" → magic-link flow. 10. **Family voice library** (sub-section of corpus). A small grid of recorded pronunciation samples from real relatives (30 short phrases is the recommended set). Each sample carries a clear consent timestamp. A persistent note: "These recordings are used as a pronunciation reference for the Live API session prompt. They are never used to clone a voice for TTS." 11. **Progress.** A CEFR ladder for the heritage language, with production and comprehension plotted separately on a weeks-not-days timeline. Below: a list of the most recent next-session recommendations and how many landed. 12. **Settings.** Heritage profile editor (language, dialect markers, family region context); session length default; calendar planting preferences; school-hours / bedtime windows the app must respect; data privacy ("delete all my transcripts forever", "delete all family corpus forever", "export everything as a zip"). 13. **Footer.** "Made for the call you're about to make." Privacy: "Your conversations are yours. We never train on them. Family voice memos require explicit consent and can be deleted in 60 seconds." Capabilities `(i)` icon in header. ## 6b. First-visit onboarding Show a **first-visit onboarding** the first time a visitor lands on the app (detect via `localStorage` flag; do not show on return visits). Three slides, dismissible at any time. Persistent re-entry: a `?` icon in the header reopens it. **Slide 1 — What this is.** - Headline: "Welcome to Heritage Class." - Subhead: "Real conversations in the language your family actually speaks — in any dialect, at any level, without ever being made to feel small." - One paragraph (≤ 60 words) explaining who this is for and what makes it different from a generic language app: it calibrates to A1 honestly instead of flattering you, it uses your family's specific dialect not the textbook standard, and it never says "wrong". - Visual: a small annotated illustration of a phone on a kitchen table next to a teapot — not a generic globe icon. **Slide 2 — Try it now.** - One short prompt: "Try with the sample family". - A live demo pre-loaded with the Sydney Sikh community Punjabi A1 profile from section 8a. - 1-2 sentences pointing at *the specific page elements* where the Gemini magic happens (the dialect picker that goes deeper than country, the family-corpus upload that grounds vocabulary, the shame-free correction pattern). **Slide 3 — How to remix this.** - Headline: "Make this yours." - Three short bullets: - "Swap the sample profile in `/data/seed-profile/` for your own family language." - "Drop voice memos and phrases into `/data/family-corpus/` to ground the conversation in your own dialect." - "Wire up your Gemini API key and Firebase project via the env-var list in the capabilities panel." - Primary CTA: "Use this template" → links to AI Studio Build remix entry point. - Secondary: "Just exploring — close" (sets localStorage flag, never auto-shows again). **Accessibility:** focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, `aria-labelledby`, focus restored to trigger on close. Respect `prefers-reduced-motion`. **Don't:** - Don't gate content behind the modal. The page beneath must be fully usable. - Don't auto-reshow on return visits. Use `localStorage['onboarding-seen-v1']`. - Don't include unrelated CTAs (newsletter signup, social follow). Keep it about the template only. ## 6c. Capabilities info button (persistent in header) Add a persistent `(i)` icon in the top-right of the header (next to the primary nav). Click → opens a modal/panel titled **"What powers this app"**. **Panel contents (in this order):** **Gemini capabilities used (the hero list):** - **Gemini Live API** — bidirectional streaming audio at conversational latency. This is the headline capability. The model speaks, you speak, the conversation flows. Pace and sentence length calibrate to your level mid-conversation. - **Gemini 3.5 Flash (multilingual at the dialect level)** — pre-session prep and post-session scoring honour your family's specific dialect: Sydney Sikh community Punjabi, Daly City Tagalog, Toronto Korean banmal with halmoni, Glendale Western Armenian, Brixton Yoruba, Houston Vietnamese, Galicianer Yiddish, Singapore Mandarin with Hokkien substrate. - **Gemini 3.5 Flash (long context)** — your family corpus (voice memos, recipes, phrase lists) rides in the Live API session prompt so the model uses *your Nani's* words, not a textbook's. - **Gemini 3.5 Flash (multimodal)** — reads handwritten recipes and phrase notebooks in Gurmukhi, Hangul, Hanzi, Devanagari, Hebrew, Ge'ez, Arabic Nastaliq, Tamil, Khmer. - **Gemini TTS** — replays lines from your past transcripts at the warm, unhurried pace you actually need. Uses a generic locale voice — NEVER a clone of a real relative. - **Firebase Auth** — Google and Apple sign-in, family invitations via magic links. - **Firestore** — stores your profile, sessions, transcripts, phrase backpack — syncs across your phone and laptop in real time. - **Firebase Storage** — keeps voice memos and recipe photos at upload resolution, with explicit per-item consent and a delete-forever option. - **Cost note** — see the detailed breakdown in 6d. A 10-minute Live API session costs about $0.50. A typical learner runs ~10 sessions a month, ~$5/month of Gemini API spend, total. - **Privacy note** — your conversations and voice memos are private to you and the family members you invite. This app uses the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. Grandparent voice memos require explicit recorded consent before upload and can be deleted in 60 seconds. **Backend services this app depends on:** - Auth: see section 4b - Database: see section 4b - Storage: see section 4b - Email: see section 4b - Payments: see section 4b (not used in v1) - External APIs: see section 4b **Environment variables you'll need to configure:** - `GEMINI_API_KEY` — your Google AI Studio API key - `FIREBASE_PROJECT_ID` — your Firebase project id - `FIREBASE_SERVICE_ACCOUNT` — service-account JSON (server-side only) - `FIREBASE_STORAGE_BUCKET` — your Storage bucket (enable Storage in console first; not auto-provisioned) **Cost + privacy notes:** - The Live API session is billed per minute of streaming audio (input + output). A 10-minute A1 session at the recommended pace costs about $0.50; a 15-minute B1 session costs about $0.85. - The post-session debrief call (Gemini 3.5 Flash, medium thinking) costs ~$0.02 per session. - The family-corpus parse calls (one per uploaded artefact) cost ~$0.01-$0.03 per item depending on length. - TTS replay-the-line: ~$0.0005 per replay, cached per line. - One short paragraph on privacy: where the data lives (your Firebase project), how to delete it (Settings → "Delete all my transcripts forever" — gone in 60 seconds), what is never sent for training. Grandparent voice memos have their own per-item consent record. **Documentation links:** - AI Studio Build docs - Gemini Live API docs - Gemini API multilingual + long-context + multimodal docs - Firebase Auth, Firestore, Firebase Storage docs - CEFR self-assessment grid (for users who want to understand the level ladder) **Accessibility:** same standards as the onboarding modal — focus trap, `Esc`, ARIA, restored focus. **Behaviour:** - Always available — single click from anywhere in the app. - Tooltip on the `(i)` icon: "How this app is built". - Mobile: opens as a full-screen sheet that slides up. - Should be the most honest part of the app — never hand-wave service requirements; never say "AI" without naming the specific Gemini model and capability. ## 6d. Detailed cost breakdown (deployer reads this BEFORE shipping) - **Setup conversation (Live API, ~6 min once)** — ~$0.30/user, billed at first launch only. - **Heritage profile finalisation (Gemini 3.5 Flash, medium thinking)** — ~150 input tokens of setup transcript summary + ~800 output tokens of structured profile. ~$0.005/user, once. - **Live conversation session (Live API)** — billed per minute of streaming audio (input + output). A 10-minute A1 session at the recommended pace (~110 wpm, conversational turn-taking) costs ~$0.45-$0.55. A 15-minute B1 session costs ~$0.75-$0.90. Sessions are the dominant cost line; budget accordingly. - **Post-session debrief (Gemini 3.5 Flash, medium thinking)** — ~3,000 input tokens (full session transcript) + ~1,200 output tokens (structured debrief). ~$0.02/session. - **Parse handwritten family recipe (Gemini 3.5 Flash, medium thinking)** — typical 1-page recipe ≈ 2 images, ~600 output tokens. ~$0.015/item. - **Family corpus chunk + embed (Gemini 3.5 Flash, low thinking)** — runs at upload time and again before each session. ~$0.001/item per chunk pass. - **TTS replay-the-line (Gemini 2.5 Flash TTS)** — billed per output token (~$10/M output tokens), effectively ~$0.000003/character. A 12-word line ≈ $0.0002 per replay. Cached per line; charged once. - **Phrase-backpack drill scoring (Gemini 3.5 Flash, low thinking)** — ~$0.0008/drill. - **Expected per-active-user monthly cost:** A learner who runs 10 sessions per month at 10 minutes each, with a small family corpus and ~30 phrase-backpack drills, runs ~$5-$6 of Gemini API spend per month. Scaling to a thousand active users: ~$5-6k/month. - **Voice memo storage:** Firebase Storage standard tier, ~$0.026/GB/month. A 60-second voice memo at 128 kbps ≈ 1 MB; a family corpus of 50 memos ≈ 50 MB ≈ negligible per user. ## 7. Design language - **Mood:** A kitchen table at 7:30 pm, after the homework is done, before the grandmother phones. Warm, unhurried, low-tech-feeling. Not a language-learning app's gamified rush; not a tutor's clinical clipboard. The app should feel like sitting down with a great-aunt who has all evening for you. - **Typography:** Display serif for headings and the live session captions (Source Serif Pro or Adobe Caslon Pro) — heritage languages deserve dignity. A clean grotesque for app chrome and small UI (Inter or Geist). Native-script fonts loaded carefully: Noto Sans Gurmukhi for Punjabi, Noto Sans Korean for Hangul, Noto Sans Tamil, Noto Sans Devanagari, Noto Sans Hebrew, Noto Sans Ethiopic, Noto Sans Arabic, Noto Sans Khmer — never let the script fall back to a tofu glyph. - **Palette:** Warm cream background `#F5EFE3` for the live session and reading views, deep ink `#1B1714` for body text, muted teal accent `#2F5B5B` for the user's own input and pinned phrases, warm terracotta `#B8593B` for the family-corpus references (so vocabulary from Nani's voice memo is visually marked as "from family"). A muted gold `#A88547` for the CEFR level ladder points. No saturation-heavy palette; no SaaS purple. - **Imagery:** Real kitchen-table photography. The grandmother on a video call. The teenager with AirPods on the bus. The handwritten recipe under a lamp. Never stock photos of "diverse young people smiling at laptops". The family-corpus items are shown as photographs of the originals (recipe in handwriting, voice memo as a waveform with the contributor's name). - **Hand-feel touches:** A barely-visible paper grain on the transcript-view background. The live-session orb pulses subtly with audio level, not theatrically. The "phrase added to backpack" confirmation is an inline ink-on-paper tick, never a confetti animation. The map of CEFR progress over weeks looks like a hand-drawn growth chart, not a SaaS dashboard. - **Spacing:** consistent 4-px base. Generous whitespace — the user needs room to think between turns. - **Radius:** consistent token set (e.g. 6 / 12 / 20 px). Goal cards use 6; the session-debrief panel uses 12; the welcome card uses 20. - **Shadows:** subtle, layered, warm-tinted. Avoid heavy drop-shadows. - **Motion:** purposeful — the live-session orb breathes, the transcript lines fade in as the model speaks, the CEFR level chart animates a new point in over 600 ms. Respect `prefers-reduced-motion`. No bouncing splash animations. The phrase-backpack readiness-pill transition (from "almost there" to "ready") is the one place where motion carries emotional weight; respect reduced-motion by jumping rather than animating. - **States:** every interactive element has hover, focus, active, disabled. Loading uses skeletons not spinners where possible. Empty states have helpful next-action guidance ("Run your first conversation — pick a goal"). ## 8. Content generation rules - Write **realistic, specific copy**. NO Lorem Ipsum. NO generic placeholders like 'Your tagline here'. - Invent plausible names, dialects, family contexts, voice-memo content, phrases that fit the domain (use the seed content in section 8a as a starting point). When inventing, lean on real diaspora communities and real dialect markers — but never claim a fictional voice memo is a real recording or attribute a fictional phrase to a real person. - Tone: warm, direct, free of corporate language. This template is for a teenager who is nervous about being made to feel small. Not for a company. - Headlines: punchy and concrete. No 'Empower your X' filler. No 'Revolutionize'. No 'Seamless'. - Body copy: short paragraphs (2-4 sentences). Use lists where appropriate. - Plain language. Avoid jargon — except where the user already speaks the jargon (the heritage-language teacher user wants to see "CEFR A1/A2/B1" in the export; the linguistics student user wants to see "lexical / phonological / morphological" in the dialect-markers panel). - Where the app outputs AI-generated content, never label it as "AI says" — let it speak naturally. Use small uncertainty cues only where epistemic honesty requires them ("I'm less sure about this form in your dialect — does your family say it this way?" appears as a small inline note, not a banner). ## 8a. Seed content (use these specific examples) Anchor every generated copy + sample data point in the concrete content below. Use these names, numbers, dates, and snippets verbatim where helpful, or generate close variants that sit in the same world. **Sample heritage profiles (sidebar):** - **"Nani's Punjabi — Sydney"** (the demo profile) — Punjabi in Gurmukhi script, Sydney Sikh community register, family from Doaba region originally. Production A1, comprehension A2. Family corpus: 12 voice memos from Nani Harpreet (recorded in Sydney, ages 68-72), a handwritten recipe for sarson da saag, a typed phrase list of household terms ("pakhi = fan", "chappal = sandals", "thand = cold"). Dialect markers: "Nani-jee" warm honorific, "pakhi" for fan (not "pankha"), English code-mixing on technical words. - **"Lola's Tagalog — Daly City"** — Tagalog in Latin script, Daly City code-mixed register, family from Cebu originally (so Visayan-tinged). Production A2, comprehension B1. Family corpus: 8 voice memos from Lola Patring (recorded 2019-2022 before she died), a handwritten kare-kare recipe, a typed list of family vocatives ("Anak ko", "Ate", "Kuya"). Dialect markers: English code-mixing for technical words, Visayan substrate vocabulary in food, the diminutive "-iting" attached to children's names. - **"Halmoni's Korean — Toronto"** — Korean in Hangul script, banmal between grandmother and grandchild, family from Busan originally. Production A1, comprehension A2. Family corpus: 15 voice memos from Halmoni Sunhee, a handwritten kimbap recipe, a typed list of food vocabulary. Dialect markers: banmal endings used affectionately, Busan satoori intonation, "halmoni" with -ya attached for closeness. - **"Medzmama's Western Armenian — Glendale"** — Western Armenian in Armenian script, Aleppo family register (the Medzmama left Aleppo in 1962), endangered_or_low_resource flag is true. Production A1, comprehension A1+. Family corpus: 22 voice memos from before Medzmama died in 2024 (consent recorded by the family while she was alive), a handwritten lahmajun recipe, a typed list of household terms. Dialect markers: specific consonant pronunciations from the Aleppo register, Arabic substrate loan-words for certain foods. - **"Yiayia's Greek — Brooklyn"** — Modern Greek in Greek script, Kalymnos island family register. Production A2, comprehension B1. Family corpus: 6 voice memos from Yiayia Maria, a handwritten recipe for koulourakia, a typed list of dinner-table vocatives. Dialect markers: Kalymnos island pronunciation, sponge-fishing family vocabulary, "Yiayia mou" as the warm vocative. - **"Jiddo's Amharic — Addis to London"** — Amharic in Ge'ez script, Addis colloquial register, family relocated to London in 1991 after the Derg. Production A1, comprehension A2. Family corpus: 9 voice memos from Jiddo Yonas (recorded in London 2024-2025), a handwritten injera recipe, a typed list of coffee-ceremony vocabulary. Dialect markers: Addis colloquial verb endings, code-mixing with Italian loan-words for kitchen tools (a 1930s legacy), the diaspora honorific "ato" attached to elders. - **"Abuelita's Spanish — Mexico City to Houston"** — Mexican Spanish in Latin script, Chilango register from Mexico City's Roma Sur, family in Houston since 1979. Production B1, comprehension B2 (this user's heritage gap is narrower than most). Family corpus: 4 voice memos from Abuelita Lucía, a handwritten mole verde recipe, a typed list of *mexicanismos*. Dialect markers: Chilango "ándale" and "neta" as conversational fillers, diminutive "-ito" attached to objects and family names, no Spain-Spanish "vosotros". - **"A Pó's Hokkien — Penang to Sydney"** — Penang Hokkien in Pe̍h-ōe-jī romanisation (the user does not read Hanzi), endangered_or_low_resource flag is true. Production A1, comprehension A2. Family corpus: 11 voice memos from A Pó Beng (recorded in Sydney 2025), a typed list of food vocabulary, a transcribed phone call. Dialect markers: Penang Hokkien sandhi tones, Malay loan-words for food and household items, the diminutive "a-" prefixed to nicknames. **Sample live session in detail (this is what the demo should show):** - **User:** 14-year-old Sikh-Australian girl, name Simrin, attending Year 9 in Sydney's western suburbs - **Heritage profile:** "Nani's Punjabi — Sydney" (above) - **Goal set at start:** "Survive a 30-minute conversation with Nani when she arrives on Sunday — at least about school" - **Duration:** 10 minutes - **Family corpus items referenced this session:** voice memo from Nani about Simrin's school (2026-04-02), phrase list term "*pakhi*" (fan, not pankha), voice memo from Nani about Vaisakhi preparations (2026-04-15) - **Attempted utterances (3, in shorthand):** - User said verbatim: "Sat sri akal Nani-jee, school theek si aaj" (Sat sri akal Nani-jee, school was fine today) - What landed: greeting + honorific + topic + simple-past verb agreement - What drifted: nothing - Model gentle restatement: "Sat sri akal beta. Sona, school theek si — bahut khushi hoyi" (Hello, dear. Lovely, school was fine — I'm so happy) - User said verbatim: "Math vich mainu… mainu… achha lagda" (In math I… I… like it) - What landed: topic + verb attempt - What drifted: hesitation between subject and verb, present-tense agreement slightly off - Model gentle restatement: "Math vich tainu achha lagda hai? Bahut vadhia." (You like math? Wonderful.) — note: model offers correct form back gently, never says "wrong" - User said verbatim: "Mainu pakhi de neeche baith ke padhna achha lagda" (I like sitting under the fan and studying — using family word *pakhi*) - What landed: full sentence + family-corpus vocabulary - What drifted: nothing - Model gentle restatement: "Pakhi de neeche, ji bilkul — main vi os tarah karda si jab main chhoti si" (Under the fan, exactly — I used to do that too when I was little) — note: model picked up *pakhi* from corpus, did not switch to textbook *pankha* - **New phrases encountered (3, in shorthand):** - "School theek si aaj" (school was fine today) — model offered, source: model_offered_for_dialect_fit, in backpack: yes - "Pakhi de neeche baith ke" (sitting under the fan) — source: user_introduced via family corpus, in backpack: yes - "Bahut khushi hoyi" (I am so happy) — model offered, source: model_offered_for_dialect_fit, in backpack: no - **Starting production level:** A1 - **Ending production level:** A1 (no concrete evidence of A2 production yet — the past-tense verb agreement came close but not consistent enough) - **Goal progress summary (in warm voice):** "Today you said five sentences to Nani in Punjabi that you've never said before. The school sentence landed beautifully on the first try. The math sentence reached for the present tense and almost got there — let's give that one a little more time. The fan sentence, using *pakhi* the way Nani says it, was real Punjabi — Nani will recognise it instantly when you say it on Sunday." - **Next session recommendation:** "Saturday morning — let's rehearse one more time, focused on the present tense for things you like and don't like. You'll be ready for Sunday." - **Shame check:** verified — no "wrong/no/incorrect" language used in this session **Sample input artefacts (for the build to demonstrate):** - A 90-second voice memo of Nani Harpreet recorded in Sydney April 2026, talking about Simrin's school in Punjabi with a few English words mixed in - A handwritten recipe for sarson da saag in Gurmukhi script, photographed under kitchen-table light - A typed phrase list (text file) of family-specific household vocabulary - A 75-second voice memo of Lola Patring (recorded 2021, before she died, with consent timestamped at the time of recording) talking about Cebu food in Tagalog with Visayan tags **Sample voice copy:** - Onboarding: "Whose language are you here to learn?" - Live session pause indicator: "Take your time." - Live session topic fork (after 12s silence): "Want to keep going on this, or switch to something else?" - Empty session list: "This is where your past conversations will live. Run your first one — pick a goal." - Error (Live API connection drop): "We lost the line for a moment. Pick up where you left off?" - Phrase added to backpack: "Added *pakhi de neeche baith ke* to your backpack — we'll drill it next time." - Level movement: "Your production moved from A1 to A1+ this session — based on the past-tense sentence you said about your sister." - Honesty flag: "I'd say this form like this — does your family say it this way too?" - Session wrapping: "We've about two minutes left — want to finish on the sentence about Sunday?" **Sample family invitation email subject + body:** - Subject: "Nani — Simrin wants to learn your Punjabi. Will you send a voice memo?" - Body: "Dear Nani — Simrin is using an app to practise Punjabi before you arrive Sunday. It works better when it knows how you actually speak. Would you record a 1-minute voice memo about anything — what you cooked today, what the weather is like — so the app can hear your voice? Tap to record." [Open Family Corpus] ## 9. Media & assets - **Hero image (landing screen):** A photographed-looking shot of a teenager at a kitchen table, AirPods in one ear, phone on the table next to a chai cup, with a wall calendar in the background showing a circled date. Generate via Nano Banana 2 with a prompt emphasising "kitchen table at 7:30 pm, warm desk-lamp light, teenager focused but at ease, real chai cup with steam, soft shadow under the phone, calendar slightly out-of-focus in background". Avoid the glossy AI render look — ask for asymmetry and slight imperfection. - **App icon / wordmark:** Set in the display serif. A small warm-cream tinted background. No icon — just type. - **Empty-state illustration:** A simple line drawing of a single phone next to a teapot. Hand-drawn aesthetic, not a flat icon. - **Demo profile imagery:** Generated per the prompts in section 8a — Nano Banana 2 prompts for each sample heritage profile (Sydney kitchen, Daly City kitchen, Toronto kitchen, Glendale kitchen, Brooklyn kitchen). Each demo should look photographed at home, not rendered in a studio. - **Live session UI:** A breathing audio-level orb at centre, no decorative chrome. The orb's idle state is a slow 4-second breath; it speeds and brightens with the model's audio output. - **Stock fallbacks:** If image generation fails, fall back to the photographed sample kitchen-table image from `/public/samples/sample-kitchen.jpg`. Never to a generic "👵" emoji. - **Generated imagery:** prefer Nano Banana 2 over stock photography. Prompt for warmth, asymmetry, and slight imperfection — avoid the glossy 'AI render' look. - **Optimisation:** WebP/AVIF, `loading="lazy"`, explicit `width`/`height` to prevent layout shift. - **Icons:** `lucide-react` for UI. Use sparingly — never decorative-only. ### Build-time asset manifest (explicit specs) Every image, illustration, and visual reference mentioned above must resolve to ONE of the three buckets below — runtime-generated, seed-shipped, or user-supplied. Do NOT ship `` tags whose `src` is not listed here. Do NOT depend on bare "section 8a prompts" without binding them to explicit paths and model IDs. **Bucket 1 — Runtime-generated (Nano Banana Pro `gemini-3-pro-image` for hero/demo photographs; Nano Banana 2 `gemini-3.1-flash-image` for in-app illustrations and reference-conditioned variants).** Cached to Firebase Storage; served via signed URL. Every reference above to "Nano Banana 2" or "Nano Banana Pro" MUST be wired to one of these specific calls with an explicit model id: - `/public/generated/hero.webp` (2400×1500, WebP) — model `gemini-3-pro-image` — uses the literal prompt described as "Hero image (landing screen)" above. Run once at build; commit a `/public/samples/hero-fallback.webp` (1600×1000) generated from the same prompt with `gemini-3.1-flash-image` so the page renders if quota is exhausted. - `/public/generated/demo/{demo-slug}-{NN}.webp` (1600×1200, WebP) — model `gemini-3.1-flash-image` (reference-conditioned where the prior frame is passed as input) — one path per "Demo X" image referenced above. The slug derives from the seed example in section 8a; the NN index covers each frame in the demo sequence. - `/public/generated/illustrations/{name}.webp` (1024×1024, WebP) — model `gemini-3.1-flash-image` — one path per named illustration above ("Empty-state illustration", "Recipe-card hero illustrations", "Curriculum picker imagery", "Period-style frames", etc.). Each illustration's prompt is the literal description above; ship a deterministic seed in the request so re-runs are reproducible. **Bucket 2 — Seed assets shipped with the deliverable.** Every "Stock fallback" path referenced above (e.g. `/public/samples/sample-X.jpg`) is generated once via Nano Banana 2 (`gemini-3.1-flash-image`) at 1024×1024 WebP using the same prompt as its Bucket-1 counterpart, then committed to the repo so the page renders identically if Gemini quota is exhausted or the user is offline. Replace any `.jpg` extension above with `.webp` to match the optimisation rule. Also commit these empty-state seeds (1024×1024 WebP, single-stroke hand-drawn line, no colour fill): - `/public/samples/empty-state-primary.webp` — line drawing of the app's primary empty surface (the named "Empty-state illustration" above), generated from that exact prompt. - `/public/samples/empty-state-archive.webp` — line drawing of an empty saved/archive view, single-stroke outline. - `/public/samples/empty-state-error.webp` — line drawing of a hand placing a single object aside with care, used when an AI call fails. **Bucket 3 — User-supplied.** Uploads from the user's camera / file picker land at the Firebase Storage path conventional for this template (named in section 4b). The build ships with Bucket-1 + Bucket-2 only; no user-supplied images at first paint. **Hard rules** - Every `` tag MUST have a `src` that resolves to a path listed in Bucket 1, Bucket 2, or a Bucket 3 upload path. Anything else is a build error. - No bare `image.jpg` / `hero.jpg` / `placeholder.png` references anywhere in the code. - Model IDs: `gemini-3-pro-image` for hero-quality photographic generation; `gemini-3.1-flash-image` for in-app illustrations, reference-conditioned variants, empty-state seeds, and stock fallbacks. Never use a legacy model id (no `imagen-*`, no `gemini-1.5-*-image`). - File format: WebP everywhere (AVIF acceptable where the target browsers support it). No `.jpg` / `.jpeg` / `.png` in `/public/samples/`. ## 10. Interactivity & states - Every interactive element has hover, focus, active, and disabled states. - Forms validate inline and show specific error messages (not "Invalid input"). The goal-picker validates that a goal is selected before "Start session" enables. - Loading states use skeletons that match the eventual layout, not spinners. - Empty states explain the next action with a button whose label fits THIS app's domain: "Run your first conversation", "Drop a voice memo from your grandmother", "Invite a family member" — never a generic "Add your first item". - Smooth scroll for in-page anchors. - The live-session UI streams captions in real time, with the heritage-language caption arriving first and the English gloss optionally fading in below after a 400 ms delay (so the user reaches for the heritage language first). - If a Live API call fails, show a calm, specific error ("We lost the line for a moment. Pick up where you left off?") and offer a one-tap reconnect. - Low-confidence pronunciation flags in the transcript are marked with a small inline question mark; tapping reveals the model's honesty note ("I'm less sure about this form in your dialect; does your family say it this way?"). - The CEFR level chart animates a new point in over 600 ms; respect `prefers-reduced-motion` by jumping instead of animating. - The phrase-backpack readiness-pill transition uses a fade with motion; respect reduced-motion by switching instantly. - The session timer respects user-set school-hours and bedtime windows: if a session would run past bedtime, the goal-picker offers a shorter duration default and explains why. **Session-resume snippet (Live API 2-min cycle):** the Live API audio+video session caps at 2 minutes. On every Live tick, persist a `SessionSyncState` to `sessionStorage`; on reconnect, pass a concise context-summary block as the first system message of the next handshake so the model continues without losing thread. ```typescript interface SessionSyncState { activeSessionId: string; accumulatedSegments: Array<{ speaker: string; text: string; timestamp: number }>; // ...template-specific cursor state (current page, turn index, etc.) } ``` ## 11. Tech & responsive requirements - **TTS markdown-stripping preprocessor:** before sending any user-authored markdown to `gemini-3.1-flash-tts-preview`, strip non-spoken markdown: `#`/`##`/`###` headings (keep the title text), `**bold**` (keep the inner text), `[label](url)` (keep `label`, drop URL), `` ``` `` fenced code blocks (skip entirely), `>` block-quote markers (keep the text), and `|` table pipes (read row-by-row as sentences). Insert `…` between sentences for a short pause and a blank line plus `—` between paragraphs for a long pause. The model does not understand markdown; raw markdown will be read aloud as literal characters ("asterisk asterisk"). - **File downloads on Safari / Firefox:** when offering local-disk save of any export (PDF, CSV, MP3, ZIP, JSON, image), fall back to `` with a blob URL — the File System Access API (`showSaveFilePicker()`) is Chromium-only. Detect with `'showSaveFilePicker' in window`; otherwise use the anchor-download path. - **Stack:** React + TypeScript + Tailwind CSS. Functional components + hooks. Use Shadcn UI primitives where appropriate. - **Build runtime:** AI Studio Build — full-stack with Cloud Run server-side functions. All Gemini API calls happen server-side; API key lives in Secrets Manager, never in client bundle. The Live API session is mediated server-side; the client streams audio over WebSocket to the server which streams to Gemini. - **Model selection:** explicitly pin `gemini-3.1-flash-live-preview` for the Live API session and setup call, `gemini-3.5-flash` for profile finalisation, debrief, and corpus parsing, `gemini-3.5-flash` for chunking and drill scoring, `gemini-3.1-flash-tts-preview` for replay-the-line. Set `thinkingLevel` explicitly per call. - **Database:** Firestore (auto-provisioned by AI Studio Build). Show the seed profile on first launch. - **Auth:** Firebase Auth — Google sign-in by default; Apple sign-in next to it; magic-link email as fallback (also used for family invitations). - **Storage:** Firebase Storage for voice memos and recipe photos. Pre-signed URLs only. - **Mobile-first.** Verify layouts at 375 px (iPhone SE), 768 px (iPad), 1024 px, 1440 px+. The live-session UI is designed for one-handed use on a phone with AirPods in. - Use `clamp()` for fluid typography. Prefer container queries over media queries for component-level responsiveness. - Use `dvh` / `svh` instead of `vh`. Respect safe-area insets on iOS. - Zero horizontal overflow at any width. Zero layout shift on load. - Persist user data in Firestore. Use real-time listeners on the home view (today's session card updates as the day passes). - Optimistic UI on writes; reconcile on response. - The live-session WebSocket reconnects within 3 seconds on network drop; the user does not lose the session state. - The microphone permission request is contextual — asked at the start of the first session, not at app launch. - **iOS Safari gotchas (graceful degradation):** the Live API session must survive iOS audio-session interruption (incoming call, Siri, alarm) — listen for `MediaStreamTrack.onmute` and pause the session calmly ("take a second — I'll wait"); resume on `onunmute`. Microphone permission does NOT persist across page reloads on iOS — re-request at the start of every session, not just the first. Backgrounded Safari tabs throttle WebSocket and kill `getUserMedia` — pair `visibilitychange` with a screen Wake Lock during a session so a 12-minute conversation with the grandmother's vocabulary is not silently dropped. PCM streaming must go via `AudioWorklet` (Safari `MediaRecorder` is AAC-only). With AirPods, register `MediaSession` action handlers so a single tap can pause/resume the session. ## 12. Accessibility (WCAG 2.2 AA) - Semantic HTML — `header`, `nav`, `main`, `section`, `article`, `footer`. - All interactive controls reachable by keyboard with a visible focus ring. - Color contrast ≥ 4.5:1 for body, 3:1 for large text and UI components. - All images have meaningful `alt` text. The hero image carries `alt` describing the artefact ("photograph of a teenager at a kitchen table with AirPods in one ear and a chai cup, a wall calendar with a circled date in the background"). - Form fields have associated `