# MUST OBEY — Mobile-first build requirements This app's PRIMARY surface is a mobile phone. Build it impeccably on mobile FIRST, then verify on tablet and desktop. Treat the rules below as non-negotiable hard constraints, not suggestions. ## Viewports to verify (every screen, every state) - 320 px, 360 px, 375 px, 390 px, 414 px, 480 px - 768 px, 834 px (iPad portrait / Pro 11) - 1024 px, 1280 px, 1440 px, 1920 px, 2560 px - Plus: 200% browser zoom, landscape orientation on every mobile width, iPhone with safe-area insets visible ## Hard layout rules - Mobile-first CSS. Default styles target mobile; `@media (min-width: ...)` for larger viewports. - Use `dvh` and `svh` instead of `vh` for full-height surfaces (iOS Safari URL-bar bug). - Use `clamp()` for fluid typography across all viewports. - Prefer container queries (`@container`) over media queries for component-level responsiveness. - Use `min(100%, ...)` widths so content never overflows. Zero horizontal overflow at any viewport. - Add `` to every page. - Apply `padding: max(safe-area-inset-X, fallback)` on every edge-bleeding container so notched iPhones in landscape never clip content. - Wide tables and code blocks scroll INSIDE their container (`overflow-x: auto`), never push the body. - Use `background-attachment: scroll` on mobile, not `fixed` (iOS Safari repaint bug). - Avoid `backdrop-filter` on animated elements. Use it sparingly on static surfaces only. - **Canvas Scaling**: Canvases must dynamically scale with window resize events and properly handle high-DPI screens (`window.devicePixelRatio`). Set physical dimensions (`canvas.width`/`canvas.height`) using pixel ratio and render relative to this grid, using CSS to control responsive viewport scaling. - **Robust Storage**: Every access to `localStorage`/`sessionStorage` (especially `JSON.parse` of loaded state or writes) MUST be wrapped in a `try-catch` block to handle disabled storage, private browsing mode, quota limits, or corrupted JSON gracefully. Fall back to a robust in-memory object store. ## Touch & accessibility - Tap targets ≥ 44 × 44 px on touch (Apple HIG). Increase to 48 px under `@media (hover: none) and (pointer: coarse)`. - All interactive controls reachable by keyboard with a visible focus ring; respect `:focus-visible`. - Color contrast ≥ 4.5:1 for body text, 3:1 for UI components. - All images have meaningful `alt`. Decorative images use `alt=""`. - Respect `prefers-reduced-motion: reduce` — zero animation durations under that query. - Forms validate inline; error messages are specific, not "Invalid input". - Modals: focus trap, `Esc` closes, `role="dialog"`, `aria-modal="true"`, focus restored on close. ## Performance bar (Lighthouse mobile, throttled 3G/4G) - LCP < 2.5 s · INP < 200 ms · CLS < 0.1 - JS bundle gzip < 200 KB mobile-first; lazy-load non-critical screens via `React.lazy` / dynamic imports. - No render-blocking resources above the fold. - Images: WebP/AVIF preferred, `loading="lazy"`, explicit `width`/`height` attributes (zero CLS), `srcset` for retina. - Videos: `preload="metadata"`, low-resolution poster, max 720p mobile fallback. Never autoplay with audio. - Fonts: `font-display: swap`; preload only the one used above the fold. - Smooth scroll honoured via CSS `scroll-behavior: smooth` with reduced-motion fallback. ## Pre-ship mobile checklist (the deployer MUST verify before declaring done) 1. Open at 375 px in DevTools — every screen scrolls vertically only; zero horizontal scroll. 2. Browser zoom 200% — layout reflows without overlap. 3. iPhone Safari with the URL bar visible AND landscape — no content under the home indicator; no notch clipping. 4. iPad portrait (768 px) and landscape (1024 px) — no awkward gaps; tablet-specific breakpoints land cleanly. 5. Tap every interactive element with a thumb at real-device size — every target is easy to hit. 6. `prefers-reduced-motion: reduce` — every transition / animation skips cleanly, scroll-behavior becomes instant. 7. Lighthouse mobile score ≥ 90 across all 4 categories. 8. Zero `console.error` and zero CLS shift in real-device testing on a mid-tier Android (e.g. Pixel 6a) and an iPhone SE. --- The original template starts below. All rules above apply on TOP of whatever this template specifies. --- # Compare Two Things ## 1. Project **Compare Two Things** is the structured comparison tool for the person who already has 14 browser tabs open. You drop in two products, two cities, two flats, two job offers, two car insurance policies, two strollers, two laptops, two language-learning apps — and in about thirty seconds you get back a clean side-by-side table where every cell links to a real, current source. Not specs the model "remembers". Not a SEO comparison page seeded by an affiliate network. The page the manufacturer published this month, the spec sheet PDF, the Reddit thread where eight owners weighed in last week, the price the retailer is charging today. The shape of the table is the same every time: a row per dimension, a column per option, a verdict paragraph below. The intelligence is in three places — picking the right dimensions for the kind of thing being compared, grounding every cell in a real fetched source with a citation URL, and writing a verdict paragraph that says "for your stated priorities the second option wins on these three rows because…" rather than "you should buy the second one". The distinction matters: this app makes recommendations, it does not give advice. The decision is the user's. The single demo that proves the magic: a parent in a kitchen at 9pm types *"Bugaboo Fox 5 vs UPPAbaby Cruz V2, city walking + weekly grocery runs, two-year-old + a baby on the way"*. In under forty-five seconds the app comes back with a twelve-row table covering folded weight, fold dimensions, recline angle, max child weight, basket capacity in litres, wheel diameter, suspension type, current price across three retailers, parts-replacement cost, what owners say in the last sixty days of reviews (three positives, three negatives per option, each one linked to the source thread), and a one-paragraph verdict that names which option wins for *city walking + grocery runs* specifically and which compromises that choice carries. Every spec links back to the manufacturer page or the retailer listing that justifies it. When a number couldn't be verified the cell shows "couldn't verify" with the search attempts listed on hover — never an invented number. And it works the same way at higher stakes. Two job offers: salary, equity strike, cliff and vest, healthcare premium, PTO, parental leave, remote policy from the actual handbook (where the company has published one), Glassdoor + Levels.fyi range for the title, recent layoff signals from press coverage, commute time from your stated home postcode. Two cities: cost-of-living index from a current source, school ratings, climate, broadband availability, the actual rental listings on the market this week. Two language apps: pricing tier, free-tier limits, languages offered, the specific pedagogy approach (spaced repetition? CEFR- aligned? AI tutor?), and what app-store reviews from the last sixty days actually complain about. **Tagline:** _Two options, one honest table — every cell sourced, no invented specs._ ## 2. Target audience The audience is universal — every adult makes comparison decisions several times a year and most of them currently happen with 14 browser tabs. The app is not segmented by demographic. It is segmented by the *kind of decision being made*. The same app serves: - **Major-purchase shoppers** weighing two strollers, two car seats, two ovens, two e-bikes, two espresso machines, two laptops, two cameras, two mattresses, two snow tyres, two electric kettles, two robot vacuums. - **Big-ticket service shoppers** comparing two car insurance policies, two health insurance plans, two mortgage offers, two broadband plans, two private school applications, two pension funds, two energy tariffs. - **Career-decision makers** with two job offers on the table and a deadline this week — compensation, benefits, leave, equity, the actual employer culture as it shows up in current press + reviews. - **Relocation researchers** comparing two cities or two neighbourhoods on cost-of-living, schools, climate, healthcare, broadband, commute, and the actual rental market today. - **Software / subscription evaluators** comparing two SaaS tools, two language-learning apps, two newsletter platforms, two AI coding assistants — features, pricing tiers, free-tier limits, user-reported pain points from the last sixty days. - **Travel planners** comparing two hotels in the same area, two airline routings, two car-rental options, two tour operators — current price, current reviews, current cancellation policy. - **Property buyers and renters** comparing two flats or two houses on price, square footage, EPC rating, council tax band, estate agent reputation, the school catchment, the commute by the actual transport options. - **Vehicle shoppers** comparing two cars on MSRP, reliability, recall history, current dealer inventory, total cost of ownership over five years. - **Equipment-driven professionals** — a film-maker comparing two lenses, a chef comparing two ovens, a tattoo artist comparing two rotary machines, a carpenter comparing two track saws. - **Anyone who has been told by the AI tool they already use that the spec on a product is X**, then discovered the manufacturer's actual spec page says Y. This app exists because that experience is a betrayal of trust and the fix is grounded search with a visible citation on every cell. ## 3. Core value propositions Surface these clearly in copy, in visual emphasis, and in section ordering — they are the reasons a visitor stops scrolling and taps. - **Every cell is sourced.** Nothing in the table is invented. The weight of the stroller is the number the manufacturer publishes, with the URL on hover. The current price is the live retailer page, with the timestamp of the fetch. The reliability score comes from the publication that produced it. If the model could not find a current, citeable source for a row, the cell shows "couldn't verify" and lists the queries it tried — it never guesses. - **The dimensions match the decision.** The app picks the rows that matter for *this* kind of decision. Two strollers ask for folded weight, basket capacity, recline angle, wheel diameter. Two job offers ask for base, equity, cliff, healthcare premium, PTO. Two cities ask for cost-of-living index, school ratings, climate type. The user can edit the row list — add, remove, reorder — and the comparison re-grounds in place. - **The verdict is a recommendation, not advice.** The verdict paragraph below the table names which option wins for the user's stated priorities, on which rows, and which compromises that choice carries. It does not say "you should buy this". It says "for *city walking + grocery runs* the Cruz wins on basket capacity and folded weight; the Fox 5 wins on suspension and legacy parts availability; the trade is comfort over compactness". - **Two becomes three or four when you need it.** Most decisions start binary and grow. Drag a third or a fourth option into the comparison and the table re-grounds — same rows, same sources, same verdict logic, just three columns instead of two. - **The reviews are recent and they are themed.** The "what owners say" rows do not summarise marketing copy. They surface three recurring positive themes and three recurring negative themes from real-user reviews dated in the last sixty days, each one linked back to the thread or review listing it came from. The model is told explicitly to ignore review pages older than six months for fast-moving categories. - **The price is current and it is the price *you* would pay.** Where the country and currency are detectable from the visit (or the user has set them in onboarding), the live retailer prices come back in that currency, including any country- specific availability gaps ("not currently sold in IE — closest retailer ships from UK with €18 customs"). - **Long-context across many options.** When the user adds a fourth option, the model sees all four in the same call so the rows stay comparable. When the user wants "the same comparison but optimised for stairs not flat pavement", the entire prior comparison is in context — the model adjusts the verdict without re-discovering the specs. - **The disclaimers are honest and they are visible.** Every monetary comparison surfaces a "not financial advice" note in the verdict. Every safety-adjacent comparison (car seats, helmets, child-product weight limits) surfaces a "not a substitute for the manufacturer's safety guidance" note. These are not buried — they are part of the verdict block. ## 4. Features to build - **The comparison entry box** — a single input that accepts free text ("Bugaboo Fox 5 vs UPPAbaby Cruz V2 for city walking + grocery runs"), structured fields (option A, option B, the job-to-be-done), or a paste of two URLs. - **The dimension picker** — the model proposes the rows that matter for this kind of decision; the user can add, remove, reorder, or pin "must-have" dimensions. - **The comparison table itself** — a two-or-more-column layout, one row per dimension, every cell containing a value, a unit, and a citation chip linking to the source. - **The verdict paragraph** — below the table, a 5-8 sentence recommendation block naming which option wins for the stated priorities, on which rows, with which compromises. - **Drag-add a third option** — the user can drop a third product/city/offer into the comparison and the table re-grounds with the same rows. - **Edit the priorities** — change the "for what" line and the verdict re-runs without re-fetching the underlying specs (uses long context). - **Currency + country detection** — country inferred from visitor signals (with explicit confirmation on first run), prices fetched in that currency; gaps surfaced honestly. - **"Why this row" inline explainer** — tap any row label, the model returns a one-sentence explanation of why that dimension matters for this kind of decision. - **"Couldn't verify" mode** — when a row has no current citeable source, the cell shows "couldn't verify" with the attempted queries listed on hover. - **Citation chips with hover preview** — every cell with a value has a chip; hover shows the source title, the publication date, and the fetched timestamp. - **Stale-source warning** — citations older than a configurable window (six months for products, twelve months for service policies, two years for cities) are flagged with a yellow chip and the model is asked to re-search. - **Compare history** — past comparisons live in the user's profile; re-run a comparison with one tap and the table updates with current data. - **Share a comparison** — generate a static, citation-preserving link to a comparison so a partner can read it (no edits unless they sign in). - **Export as PDF or markdown** — both formats preserve the citation URLs as footnotes; PDF is print-clean. - **Side-by-side hero illustration** — Nano Banana Pro composes a single 4K hero image with both options visually side-by-side, used for the share card and PDF cover. - **Read it to me** — TTS playback of the verdict paragraph, for the user who is decision-fatigued and wants to hear the recommendation while walking the dog. - **Disclaimer block in the verdict** — automatically inserted for financial, medical, safety-adjacent, and legal-adjacent comparisons; never buried. - **"What's the catch" row** — the model is required to surface one "what's the catch" row per option (a non-obvious downside or trade-off it found in reviews or in fine print). - **First-visit onboarding** — three-step set-up: pick the kind of decision (or "let the model guess"), confirm country + currency, enter the two options. - **Capabilities-info button** — visible in the header; opens a panel naming the exact Gemini calls used, the cost per comparison, what data leaves the device, and which data is never used for model training. ## 4b. Required Gemini capabilities + backend services **This template's intelligence comes from the Gemini capabilities below. Wire them up explicitly — don't substitute generic LLM calls.** ### Gemini capabilities (the load-bearing intelligence) - **Grounded comparison via Google Search** (`gemini-3.5-flash`, thinkingLevel `medium`, `google_search` tool enabled). This is the load-bearing call. The model is given the two-or-more option names, the kind of decision, and the user's stated priorities, and emits a structured `ComparisonTable` object as JSON in the *text body* (not via `responseSchema` — `responseSchema` and `google_search` are mutually exclusive in one Gemini call). The server parses the JSON, then walks `response.groundingMetadata.groundingChunks[]` to attach a `web.uri` citation to every cell that has a value. Cells with no matching citation collapse to `"couldn't verify"` with the searched queries logged in the cell's metadata. - **Dimension-picker call** (`gemini-3.5-flash`, thinkingLevel `low`, no tools, `responseSchema` enabled). Given the two option names and the kind of decision, returns the list of rows to populate — ordered by importance to the user's stated priorities. Schema is `DimensionList`. This call runs first so the grounded comparison call already knows which dimensions to source. Splitting the two calls keeps the grounded call focused and prevents the model from inventing dimensions on the fly inside the grounded response. - **Verdict-writer call** (`gemini-3.5-flash`, thinkingLevel `medium`, no tools, `responseSchema` enabled). Given the fully-populated `ComparisonTable` (including citations) and the user's priorities, emits a `Verdict` JSON object: recommendation, supporting rows, compromises, required disclaimers. This is split from the grounded call so the verdict-writer reads only the *verified* cells — never an unverified or hallucinated cell. - **Long context across the table** (Gemini 3.5 Flash, 1M-token context). When the user adds a third or fourth option, the whole prior table + citations + verdict are placed in the context of the new grounded call so the dimensions stay comparable and the model can directly reference the prior rows. **Guardrail:** a typical 12-row, 4-column table with citations averages ~30k tokens. A user with 200 past comparisons in history could exceed 500k tokens of memory; the history loader caps at the 20 most recent comparisons and summarises older ones to ~500 tokens each. Never send more than 800k tokens of comparison payload per call. - **Stale-citation re-check** (`gemini-3.5-flash` grounded, low thinking). When a citation chip is flagged stale, this call fetches a fresh source for just that one cell and updates the cell's `citation_url`, `citation_fetched_at`, and `citation_dated` fields. The server batches all stale cells in one row-by-row sweep to minimise grounded-call cost. - **"Why this row" explainer** (`gemini-3.5-flash`, thinkingLevel `low`, no tools). Returns one sentence explaining why a given dimension matters for this kind of decision. Used by the inline explainer in the table. - **Hero side-by-side image** (`gemini-3-pro-image` / Nano Banana Pro). One 4K hero image with the two (or more) options visually side-by-side, used for share cards and PDF cover. Nano Banana Pro is required — the legible-text-at-4K capability lets the model render each option's name as a caption inside the image. **Guardrail:** the model is told explicitly NOT to render brand logos it does not have a reference image for; descriptive captions instead of logos. - **Read-it-to-me TTS** (`gemini-3.1-flash-tts-preview`). Reads the verdict paragraph aloud. No SSML — pauses are encoded as `…` at sentence boundaries and a blank line plus em-dash `—` at paragraph boundaries; a one-sentence style directive is prepended to the input. - **"What's the catch" extractor** (`gemini-3.5-flash`, thinkingLevel `low`, `google_search` enabled). Surfaces one recurring non-obvious downside per option from review threads in the last sixty days; emits the catch as a free-text string in the text body, with the citation URL coming from groundingMetadata. ### Backend services - **Auth — Required.** Firebase Auth with Google sign-in (auto-provisioned by AI Studio Build). Apple sign-in is optional and requires a configured Apple Developer account. Magic-link email for share invites requires the sender domain to be authorised in Firebase Auth. - **Database — Required.** Firestore for `users`, `comparisons`, `dimensions`, `cells`, `citations`, `verdicts`, `shared_links`, `compare_history`. - **File storage — Optional but recommended.** Firebase Storage for generated hero images and exported PDFs. Storage is NOT auto-provisioned by AI Studio Build today — enable it in the Firebase console and wire the bucket name into the AIS Build project before first PDF export. Pre-signed URLs only. - **Workspace integration — Optional.** With the post-I/O 2026 Workspace integration baked into AI Studio Build, signed-in users can push a comparison directly into a Google Sheet without an OAuth dance. Off by default, opt-in per user. - **Server functions — Required.** Cloud Run server functions for the four orchestrated Gemini calls (dimension picker → grounded comparison → verdict writer → optional hero image). The free 2-app Cloud Run deploy from AI Studio Build covers this template's needs out of the box. - **External APIs:** Gemini API only. Optional ip-based geolocation for currency hinting (the user can override). No required external API beyond Gemini. **Environment variables:** every secret (Gemini API key, Firebase service-account JSON, any geolocation key) lives in environment variables — never in the client bundle. Include a `.env.example`. **Auth + data privacy reminders:** never log secrets · never store passwords in plain text · use HTTPS everywhere · honour "delete my account" inside the UI · explicit opt-in for any analytics · the user's comparison history, priorities, and citations are never sent to Gemini for model training (use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms) · sharing a comparison is per-comparison and revocable. **Read this first — prompt-craft rules that apply to every call in this template:** 1. **Name the model variant explicitly** in every Gemini API call. Do not let the agent pick the model. See the per-call matrix below. `gemini-3.5-flash` is the new default flagship as of Google I/O 2026 — it beats the previous Pro tier on grounded reasoning at 4× the speed. 2. **Pin `thinkingLevel` explicitly** per call. See the matrix. On image-gen and TTS calls, omit `thinkingConfig` entirely — the field is unsupported on those models. 3. **`responseSchema` and `google_search` are mutually exclusive in one call.** Every grounded call (the comparison call, the "what's the catch" extractor, the stale-source re-check) emits JSON in the text body; the server parses it and walks `response.groundingMetadata.groundingChunks[]` for citation URLs. **Do NOT ask the model to emit URLs in the JSON body — it will hallucinate them.** 4. **Seed the JSON schema as a fenced TypeScript / Zod block** in the system instruction (for grounded calls) or as `responseSchema` (for ungrounded calls). The literal schemas are below. **Convert the Zod schema to Gemini's `Schema` type via the SDK helper** before passing as `responseSchema` — do NOT pass raw Zod. **Numeric `min`/`max` constraints in `responseSchema` are documentation only; clamp server-side.** 5. **Pin the `systemInstruction` separately** from user input. Use `systemInstruction` for the persona and the rules; use `contents` for the option names, kind-of-decision, and priorities. Never concatenate. 6. **Pre-declare tools as an enable/disable list** per call. The matrix names which tools are enabled per call; tools NOT listed should be disabled. Do not enable `google_search` on calls that need a `responseSchema`. 7. **State negative constraints explicitly** — they are listed below. They are NOT "be careful" suggestions; they are hard rules the model must follow. 8. **Grounded responses can wrap JSON in ```json fences or add prose preamble.** Server-side, strip fences and brace-extract: ```typescript function safeExtractJSON(raw: string): T { const clean = raw.replace(/```json\s*|```/gi, '').trim(); const s = clean.indexOf('{'); const e = clean.lastIndexOf('}'); if (s === -1 || e === -1) throw new Error('No JSON boundaries in grounded response'); return JSON.parse(clean.slice(s, e + 1)) as T; } ``` 9. **Strip unsupported Zod modifiers before passing to `responseSchema`** — Gemini's OpenAPI subset rejects `.regex()` / `pattern`, fixed-length `z.tuple()`, and other custom validators. Use a sanitizer that flattens tuples to arrays and removes regex patterns before serializing. Validate those constraints in middleware AFTER parsing. 10. **Files API uses `files/*` resource names, not `gs://` URIs.** The AI Studio Build runtime uses the Gemini Developer API (`@google/genai` SDK). Files API `upload` returns a resource name like `files/abc123xyz`, passed via `fileData: { fileUri, mimeType }`. `gs://` URIs belong to Vertex AI / Cloud Storage — a different surface, not accepted here. ### Per-call model + tools matrix | Call | Model | thinkingLevel | Tools enabled | |------|-------|---------------|---------------| | Dimension picker → `DimensionList` | `gemini-3.5-flash` | low | (none); `responseSchema` enabled | | Grounded comparison → `ComparisonTable` (JSON in text body) | `gemini-3.5-flash` | medium | `google_search` — no `responseSchema` | | Verdict writer → `Verdict` | `gemini-3.5-flash` | medium | (none); `responseSchema` enabled | | "What's the catch" extractor (per option) | `gemini-3.5-flash` | low | `google_search` — no `responseSchema` | | Stale-citation re-check (per cell, batched) | `gemini-3.5-flash` | low | `google_search` — no `responseSchema` | | "Why this row" inline explainer | `gemini-3.5-flash` | low | (none); `responseSchema` optional | | Hero side-by-side image | `gemini-3-pro-image` | n/a | n/a | | Verdict TTS playback | `gemini-3.1-flash-tts-preview` | n/a | n/a | *Note for builders:* on TTS and image-generation calls, omit `thinkingConfig` entirely — the field is not supported on those models. The `n/a` cells in this matrix are documentation only; do not serialise them into the request body. Grounded search calls emit JSON in the text body — `responseSchema` and `google_search` cannot be combined in the same Gemini call; parse the JSON server-side and read citation URLs from `response.groundingMetadata.groundingChunks[].web.uri`. As of 2026-06-01 `gemini-3.5-flash` is announced but pre-GA (Sundar's keynote: "give us until next month"); do NOT wire it. The post-I/O 2026 Computer Use endpoint is still pinned to `gemini-2.5-computer-use-preview-10-2025`; it is not used in this template. Gemini Omni Flash (the I/O 2026 video model) has no public developer API as of 2026-06-01 and is also not used here. ### Primary structured-output schemas (seed verbatim in the prompt) ```typescript import { z } from "zod"; const DecisionKind = z.enum([ "physical_product", // strollers, laptops, espresso machines "service_policy", // car insurance, broadband, energy tariff "job_offer", "city_or_neighbourhood", "property_listing", // flat, house "software_subscription", "travel_option", // hotel, flight, tour operator "vehicle", "financial_product", // mortgage, pension fund, savings "education_option", // course, school, university programme "other", ]); const Currency = z.enum([ "USD", "EUR", "GBP", "CAD", "AUD", "NZD", "JPY", "KRW", "CHF", "SEK", "NOK", "DKK", "MXN", "BRL", "ARS", "ZAR", "INR", "SGD", "HKD", "TWD", "AED", "SAR", "CNY", "PHP", "THB", "VND", "IDR", "MYR", "TRY", "PLN", "CZK", "HUF", "other", ]); const DimensionSpec = z.object({ dimension_id: z.string(), // stable id for the row label: z.string(), // "Folded weight" unit_hint: z.string().nullable(), // "kg" / "lb" / "$ / month" / null importance: z.enum([ "must_have", // user pinned as critical "recommended", // model picked as central "context", // useful background ]), why_this_row: z.string().nullable(), // one sentence the inline // explainer can show }); const DimensionList = z.object({ decision_kind: DecisionKind, dimensions: z.array(DimensionSpec), // ordered by importance options_named: z.array(z.string()), // verbatim option names // (no model normalisation) priorities_verbatim: z.string(), // user's stated "for what" flagged_for_user_review: z.array(z.string()), // dimensions the // model is unsure // are worth including }); const Citation = z.object({ citation_url: z.string().nullable(), // populated by server // from groundingMetadata citation_title: z.string().nullable(), citation_publisher: z.string().nullable(), // e.g. "Manufacturer", // "Reddit r/strollers", // "Consumer Reports" citation_dated_iso: z.string().nullable(), // publication date if // surfaced citation_fetched_at_iso: z.string(), // server timestamp stale: z.boolean(), // server flag }); const Cell = z.object({ cell_id: z.string(), option_id: z.string(), // which column dimension_id: z.string(), // which row value_string: z.string().nullable(), // human-readable cell // ("8.7 kg", "$1,099", // "13–22% per year") value_numeric: z.number().nullable(), // when a single number value_unit: z.string().nullable(), value_range_low: z.number().nullable(), value_range_high: z.number().nullable(), citation: Citation, verified: z.boolean(), // true only when citation // URL is set AND fresh unverified_reason: z.string().nullable(),// "no current source // found for this spec" searched_queries: z.array(z.string()), // the queries the model // tried, for the hover notes_one_line: z.string().nullable(), // optional one-line note // ("varies by retailer") }); const ReviewTheme = z.object({ polarity: z.enum(["positive", "negative"]), theme_one_line: z.string(), // "owners praise the // suspension on cobbled // streets" recurring_phrase_verbatim: z.string().nullable(), // a representative quote // from a review (no full // sentences from one user // — paraphrase or short // common phrasing only) source_count_at_least: z.number(), // "seen in at least N // distinct reviews" oldest_source_dated_iso: z.string().nullable(), newest_source_dated_iso: z.string().nullable(), }); const OptionReviewBundle = z.object({ option_id: z.string(), positives: z.array(ReviewTheme), // exactly three negatives: z.array(ReviewTheme), // exactly three whats_the_catch: z.string().nullable(), // single non-obvious // downside the model // found }); const ComparisonOption = z.object({ option_id: z.string(), option_name_verbatim: z.string(), // the user's wording, // preserved (no // normalisation) option_name_canonical: z.string().nullable(), // model's canonical // name once grounded // (e.g. "UPPAbaby Cruz // V2 (2024)") option_manufacturer: z.string().nullable(), option_url_canonical: z.string().nullable(), }); const ComparisonTable = z.object({ comparison_id: z.string(), decision_kind: DecisionKind, options: z.array(ComparisonOption), // 2-4 entries dimensions: z.array(DimensionSpec), cells: z.array(Cell), reviews: z.array(OptionReviewBundle), // one bundle per option country_hint: z.string().nullable(), // ISO-3166-1 alpha-2 currency: Currency.nullable(), fetched_at_iso: z.string(), unverified_cell_count: z.number(), // for the UI summary }); const VerdictRow = z.object({ dimension_id: z.string(), winner_option_id: z.string().nullable(),// null = tie or // not decided by this // row margin: z.enum(["clear", "close", "tie"]), rationale_one_line: z.string(), }); const Verdict = z.object({ comparison_id: z.string(), recommended_option_id: z.string().nullable(), // null = honest "no // recommendation" when // priorities are too // ambiguous or the // data is too thin recommendation_strength: z.enum(["clear", "lean", "close_call"]), priorities_verbatim: z.string(), supporting_rows: z.array(VerdictRow), // rows where the // recommended option // wins compromise_rows: z.array(VerdictRow), // rows where the // user is giving // something up verdict_paragraph: z.string(), // 5-8 sentences required_disclaimers: z.array(z.enum([ "financial", "medical", "safety_adjacent", "legal_adjacent", "stale_sources", "not_advice", ])), cited_cell_ids: z.array(z.string()), // every cell_id the // verdict relied on // — STRICT subset of // ComparisonTable.cells honesty_note: z.string().nullable(), // optional note if the // model wants to flag // weak evidence }); type DimensionList = z.infer; type ComparisonTable = z.infer; type Verdict = z.infer; ``` ### Common failure modes (and how to avoid them) - **Agent silently swaps `gemini-3.5-flash` for the announced but pre-GA `gemini-3.5-flash`.** As of 2026-06-01, 3.5 Pro is Vertex preview only ("give us until next month"). Pin `gemini-3.5-flash` in every call. It already beats the previous 3.1 Pro on grounded reasoning. - **Agent enables `responseSchema` and `google_search` on the same call.** They are mutually exclusive. The grounded comparison call MUST emit JSON in the text body; the server parses it. If the agent wires `responseSchema` on a grounded call, the request returns 400 (`INVALID_ARGUMENT`); the caller surfaces a clear error to the developer logs and retries without `responseSchema`. - **Model emits URLs directly in the JSON body of a grounded call.** It will hallucinate them. The system instruction must say explicitly: do NOT include URLs in the JSON; citations are attached server-side from `response.groundingMetadata.groundingChunks[].web.uri`. - **Model invents a spec when no source is found.** The system instruction must require `verified: false` and `value_string: null` whenever no grounding chunk matched the cell. The server validates: any cell where `verified: true` but `citation.citation_url == null` is rejected and the cell is forcibly set to `verified: false` with an `unverified_reason: "no citation attached"`. - **Citations get attached to the wrong cell.** The system instruction asks the model to label each fact it states with a `` token referencing the `(option_id, dimension_id)` pair; the server pairs grounding chunks to cells by matching the labels in the text body to the chunks the model emitted alongside them. Unlabelled chunks are dropped. - **Reviews older than six months get treated as current.** The reviews bundle has `oldest_source_dated_iso` and `newest_source_dated_iso` fields the model must populate. The server clamps: if a `ReviewTheme.newest_source_dated_iso` is older than 180 days for a product category, the theme is dropped and re-requested. - **The verdict-writer fabricates a row that wasn't in the table.** The system instruction requires `cited_cell_ids` to be a strict subset of the comparison table's `cells[]`. The server validates this set membership; if any cell_id is not in the table, the verdict is rejected and re-requested. - **The verdict mixes "recommendation" with "advice".** The system instruction explicitly bans imperative second-person language ("you should buy"); requires conditional recommendation language ("for *your stated priorities*, the Cruz wins on basket capacity and folded weight"). The server runs a simple regex check for banned imperative phrases ("you should", "buy the", "I recommend you", "go with") and rejects if found. - **Currency conversion done by the model.** Currency conversion drifts and the model has no real-time FX. The system instruction must say: report prices in the currency the retailer charges; do NOT convert; surface the original currency in `value_unit`. The UI handles conversion display via a published FX feed. - **Long-context dump exceeds 1M tokens.** When a user has many prior comparisons in their history, the history loader caps at the 20 most recent comparisons in full, plus per-comparison summary stubs (~500 tokens each) for older ones. Never send more than 800k tokens of payload per call. - **The hero image renders fake brand logos.** Nano Banana Pro is excellent at typography but will invent plausible-looking brand marks. The system instruction must explicitly forbid logos for brands the user has NOT provided as reference images, and require descriptive name captions in the image instead. - **"What's the catch" surfaces a 2019 forum thread.** The catch extractor is grounded with `google_search` and the system instruction caps the source date at the last sixty days for fast-moving categories (products, software, travel) and the last twelve months for slow-moving (cities, educational programmes, pension funds). - **TTS playback reads citation URLs aloud.** The verdict-TTS pre-processor must strip the `[1]`, `[2]` reference markers from the paragraph before sending to TTS; the URLs themselves never enter the TTS input. - **Three-option mode breaks the two-column UI.** The table must render 2-4 columns natively from day one; the schema supports it and the grid uses CSS subgrid (or a min-width column rail with horizontal scroll on mobile). ### Negative constraints (hard rules) - Do NOT invent specs, prices, ratings, or review themes. Every cell that asserts a value must be backed by a citation URL attached server-side from `groundingMetadata`. Unverified cells render as "couldn't verify" with the queries the model tried. - Do NOT include URLs in the JSON body of a grounded call. The model will hallucinate them. Citations come from `response.groundingMetadata.groundingChunks[].web.uri`. - Do NOT combine `responseSchema` with `google_search` in one Gemini call. They are mutually exclusive. The grounded comparison call emits JSON in the text body; the server parses it. - Do NOT write the verdict as advice. Banned phrases include "you should", "you ought to", "I recommend you", "go with", "buy the", "pick the". The server enforces a regex check. Allowed: "for your stated priorities, the Cruz wins on basket capacity"; "for *city walking + grocery runs*, the Fox 5 is the more comfortable choice and the Cruz is the more compact one". - Do NOT make financial recommendations on financial-product comparisons (mortgages, pensions, insurance) without the "not financial advice" disclaimer in the verdict. The `required_disclaimers` array enforces this; the UI surfaces the disclaimer in a prominent callout below the verdict. - Do NOT make medical or safety recommendations on safety- adjacent comparisons (car seats, child products, helmets, bicycles, medications, pharmacy products, supplements, diet products) without the relevant safety disclaimer. - Do NOT cite a source older than six months for a product comparison, twelve months for a service-policy comparison, or two years for a city/neighbourhood comparison. Stale sources are flagged and re-fetched. - Do NOT report prices in any currency other than the retailer's charging currency. No model-side currency conversion. UI handles display conversion via a published FX rate, never the model. - Do NOT normalise option names away from what the user typed. Preserve `option_name_verbatim` exactly. The canonical name lives in `option_name_canonical`, separately. - Do NOT render brand logos in the hero image for brands the user has not provided as a reference image. Use descriptive text captions in the image instead. - Do NOT auto-publish comparisons. Sharing is explicit, per- comparison, and revocable. - Do NOT use the user's comparisons, priorities, citations, or history to train or fine-tune any model. Use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms. - Do NOT skip the "couldn't verify" state. A pretty table with invented numbers is worse than an honest table with three unverified cells. - Do NOT recommend more than one option in the verdict. The schema enforces a single `recommended_option_id` (or null for honest "no recommendation"). - Do NOT moralise. The app does not lecture the user on the decision being made ("do you really need a stroller this expensive?"). The user's choices are the user's. ### Per-call `systemInstruction` strings Use these as the literal `systemInstruction` field for each Gemini API call the built app makes. They complement the series-wide rules already uploaded as the global instructions file (`00-series-instructions.txt`). ### Call: Dimension picker → `DimensionList` Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none); `responseSchema` enabled ``` You are picking the dimensions (rows) that matter for a two-or-more option comparison. You receive: - The verbatim option names the user typed. - The kind of decision (one of DecisionKind). - The user's stated priorities verbatim ("for city walking and grocery runs"). - The user's country hint, if known. Your job: return the ordered list of dimensions a careful buyer would want to see in this comparison, weighted by the user's priorities. Hard rules: - 8-14 dimensions for product comparisons; 6-10 for job offers, cities, and financial products; 5-8 for travel options and software subscriptions. - Pin the dimensions the user named in their priorities at the top with importance="must_have". - Use plain-English labels. "Folded weight" not "stowage mass". "Monthly premium" not "MRPM". - Pick units that match the user's country hint. kg for the EU/UK/CA/AU; lb for the US. Prices in the retailer's currency. - Always include at least one "what people actually say" dimension for product, software, and travel comparisons (this maps to the OptionReviewBundle later). - For job offers: base, equity (strike, cliff, vest schedule), signing bonus, healthcare premium, PTO, parental leave, remote policy, commute, employer reviews. Do NOT include "culture fit" as a row — it is not measurable. - For financial products: APR, fees, term, penalty for early exit, FSCS/SIPC/DGS guarantee status, customer complaints ratio if published. - For cities: cost-of-living index, median rent, climate type, walkability, public transport coverage, school ratings (with the rating publisher named), broadband availability, healthcare access, crime index from a named source. - Include why_this_row for every dimension that is not immediately obvious to a layperson. - flagged_for_user_review: name any dimension you are unsure is worth including for this decision; the UI will offer the user a one-tap toggle. - Preserve option names verbatim in options_named — do NOT normalise capitalisation or spelling. Output ONLY the DimensionList JSON matching the provided schema. No commentary. JSON only. ``` --- ### Call: Grounded comparison → `ComparisonTable` (JSON in text body) Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: `google_search` — no `responseSchema` ``` You receive: - A ComparisonTable scaffold pre-populated with the options and dimensions from the DimensionList call. - The user's stated priorities verbatim. - The user's country and currency hint, if known. - The current ISO date. Your job: fill in every cell with a current, sourced fact, and surface three positive + three negative review themes per option from sources dated in the last sixty days (for products, software, travel) or the last twelve months (for slow-moving categories like cities and education). Use `google_search` grounding for every factual cell and every review theme. Do NOT answer from training-data memory for any number, price, rating, or current spec. OUTPUT FORMAT — read carefully: The Gemini API does NOT allow `responseSchema` and `google_search` in the same call. You will output JSON in the *text body* of your response, matching the ComparisonTable schema seeded below. The server will parse that JSON, then attach citation URLs to each cell from `response.groundingMetadata.groundingChunks[].web.uri`. You will help the server attach citations correctly by emitting, immediately after each fact statement in your internal reasoning, a label of the form: or for a review theme: The labels are short tokens; the server uses them to pair each grounding chunk with the cell or review theme it belongs to. Hard rules: - DO NOT include URLs in the JSON body. Citation URLs come from groundingMetadata server-side. If you include a URL in a value_string or notes_one_line field, the server will reject the cell and re-request it. - For every cell, populate value_string with the human-readable answer ("8.7 kg", "$1,099 at REI", "5 years / 50,000 miles"). When the value is a single number, also populate value_numeric and value_unit. When the value is a range, populate value_range_low and value_range_high. - If no current sourced answer was found for a cell, set verified: false, value_string: null, value_numeric: null, and unverified_reason to a short honest explanation ("manufacturer's spec page does not publish folded weight for this model in the EU SKU"). Populate searched_queries with the queries you tried. - For prices, report the retailer's charging currency in value_unit (e.g. "USD", "GBP", "EUR"). Do NOT convert currencies. If no retailer in the user's country sells the product, surface that explicitly in notes_one_line ("not currently sold in IE; closest retailer ships from UK with ~€18 customs"). - For each option, surface EXACTLY three positive themes and EXACTLY three negative themes in the reviews bundle. Themes must be recurring (seen in at least 3 distinct reviews) and dated in the last sixty days (products, software, travel) or last twelve months (cities, education, financial products). - For each theme, populate oldest_source_dated_iso and newest_source_dated_iso. Do NOT paraphrase a single reviewer's full sentence — capture a recurring short phrasing or a normalised theme description. - For the whats_the_catch field per option, find ONE recurring non-obvious downside (something the marketing copy does not surface but multiple owners flag). If you cannot find one with high confidence, leave it null. - Preserve option_name_verbatim exactly as the user typed it. Populate option_name_canonical with the manufacturer's current canonical product name (e.g. "UPPAbaby Cruz V2 - 2024 Edition") if you can find it. - For safety-adjacent comparisons (car seats, helmets, child products, medications), surface recall history if any. - For financial products, surface the regulator-protection status (FSCS in UK, SIPC in US, DGS in EU) if published. - Populate fetched_at_iso with the current ISO date you receive in the input. - Populate unverified_cell_count by counting cells with verified: false. Voice: - Cells contain facts, not opinions. "8.7 kg folded" not "lightweight at 8.7 kg". - Review themes are descriptive, not evaluative. "Owners report the basket sags when filled past ~10 kg" not "the basket is bad". Output ONLY the ComparisonTable JSON in the text body. No commentary outside the JSON. No URLs in the JSON body. ``` --- ### Call: Verdict writer → `Verdict` Model: `gemini-3.5-flash` · thinkingLevel: medium · Tools: (none); `responseSchema` enabled ``` You receive a fully-populated ComparisonTable including attached citations and a verified flag per cell. You also receive the user's stated priorities verbatim. Your job: write the verdict. Pick the option that best matches the user's stated priorities, name the supporting rows, name the compromise rows, and produce the verdict paragraph. Hard rules: - Read ONLY verified cells. Cells where verified is false are not evidence; you may mention them in honesty_note but not in supporting_rows or compromise_rows. - Populate cited_cell_ids with ONLY cell_ids that exist in the input ComparisonTable.cells[]. Do NOT fabricate cell_ids. The server will reject any cell_id not present in the table. - recommended_option_id is a single option_id, OR null when: - the priorities are too ambiguous to break a tie, OR - the verified data is too thin (fewer than 6 verified cells across the table), OR - the options are functionally equivalent on every must-have row. Set recommendation_strength accordingly: "clear", "lean", or "close_call". When recommended_option_id is null, recommendation_strength is "close_call". - The verdict paragraph is 5-8 sentences. It MUST: - Open by re-stating the user's priorities verbatim or near-verbatim, so the reader sees what the recommendation is anchored to. - Name which rows the recommended option wins on (and why those rows matter for these priorities). - Name which rows the user is giving up something on, in plain language. - Surface the required disclaimers in the same paragraph (not as a postscript) — financial, medical, safety- adjacent, legal-adjacent, stale-sources as relevant. - End with a one-sentence honesty note if the evidence is thin or the recommendation is a close call. - BANNED LANGUAGE — do not use: - "you should", "you ought to", "you need to" - "I recommend you", "I'd go with", "my pick is" - "buy the X", "go with the X", "choose the X" - generic superlatives without a row reference ("the Cruz is better", "the Fox 5 is the best stroller") - ALLOWED LANGUAGE — use: - "for your stated priorities ([priorities_verbatim]), the Cruz wins on basket capacity and folded weight" - "the Fox 5 leads on suspension and parts availability; the compromise is folded weight, which is 1.4 kg heavier" - "if your priorities shift toward [X], the Fox 5 becomes the closer match" - required_disclaimers — set the right enum members: - "financial" — any financial-product comparison - "medical" — any medication, supplement, or health-product comparison - "safety_adjacent" — car seats, helmets, child products, vehicles, electrical appliances - "legal_adjacent" — insurance terms, contracts, tax products - "stale_sources" — if any verified cell's citation is older than the category's freshness window - "not_advice" — always set; the app makes recommendations not advice - honesty_note: optional one or two sentences flagging any limitation in the comparison — "review evidence for the Cruz V2 is thinner on cobbled streets than for the Fox 5; if cobbles are a daily concern, weigh that qualitatively". Output ONLY the Verdict JSON matching the provided schema. No commentary. JSON only. ``` --- ### Call: "What's the catch" extractor (per option) Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: `google_search` — no `responseSchema` ``` You receive one option (name + manufacturer + canonical URL) and the decision kind. Your job: find ONE recurring non-obvious downside — the "catch" — and return it as JSON in the text body. Use `google_search` grounding. Cap sources at the last sixty days for products / software / travel; last twelve months for cities / education / financial products. Output JSON: { "catch_one_line": "", "catch_detail": "<2-3 sentences of supporting context>", "source_count_at_least": , "oldest_source_dated_iso": "", "newest_source_dated_iso": "" } Hard rules: - The catch must be a recurring downside (seen in at least 3 distinct sources), not a one-off complaint. - The catch must NOT be something the manufacturer prominently discloses (e.g. "weighs 9 kg" is not a catch if the spec sheet says so on page 1). - If no recurring catch is found, return null for catch_one_line and catch_detail — do NOT invent one. - Do NOT include URLs in the JSON body. Citations come from groundingMetadata server-side. No commentary outside the JSON. ``` --- ### Call: Stale-citation re-check (per cell, batched) Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: `google_search` — no `responseSchema` ``` You receive one (option_id, dimension_id, last_known_value) and the decision kind. Your job: find the current value from a source dated within the relevant freshness window (60 days for products / software / travel; 12 months for slow-moving categories; 24 months for cities) and return it as JSON in the text body. Output JSON: { "option_id": "", "dimension_id": "", "value_string": "", "value_numeric": , "value_unit": "", "value_range_low": , "value_range_high": , "value_changed_since_last_check": , "fetched_at_iso": "" } Hard rules: - Use `google_search` grounding. - If no fresh source is found, return value_string: null and value_changed_since_last_check: false. - Do NOT include URLs in the JSON body. Citations come from groundingMetadata server-side. - Do NOT convert currencies. No commentary outside the JSON. ``` --- ### Call: "Why this row" inline explainer Model: `gemini-3.5-flash` · thinkingLevel: low · Tools: (none) ``` You receive (decision_kind, dimension_label, priorities_verbatim) and return ONE sentence explaining why this dimension matters for this kind of decision, given these priorities. Output JSON: { "why_this_row": "" } Hard rules: - One sentence. Plain English. No jargon. - Anchor to the user's priorities if relevant. - Do NOT moralise ("you really should care about basket capacity"); explain ("basket capacity matters for weekly grocery runs because larger baskets carry more bags without slowing the push"). - If the dimension is not actually important for this decision, say so honestly ("a stroller's wheel diameter matters less on smooth pavement than on cobbles or grass"). No commentary outside the JSON. ``` --- ### Call: Hero side-by-side image Model: `gemini-3-pro-image` (Nano Banana Pro) · n/a · n/a ``` You generate a single 4K photographic-looking image showing the comparison's options visually side-by-side, for use as the share card and the PDF cover. Image structure: - Aspect ratio 16:9. - A central divider (a thin vertical line, or a soft vignette) separates two halves; for 3-option mode use thirds; for 4-option mode use a 2×2 grid. - Each half/third/quadrant shows ONE option in its natural context. - Each section has a discreet caption at the bottom rendering the option's verbatim name in clean 4K- legible typography. Prompt anchors that work well: - "two strollers side by side in a park at golden hour, one with a black frame and grey fabric, one with a navy frame and cream fabric, soft natural light, no people, no visible brand marks; bottom-left caption reads 'Bugaboo Fox 5', bottom-right caption reads 'UPPAbaby Cruz V2', both in clean sans-serif" - "two cities at dawn, left half a coastal European cityscape with red roof tiles, right half a North American downtown with glass towers, soft pastel sky, no people, captions reading 'Lisbon' and 'Boston' in clean sans-serif" Hard rules: - Photographic, not cartoon, not illustration-style. - Each option clearly distinguishable from the other. - The verbatim option names are rendered as captions inside the image — Nano Banana Pro is required for this; 4K typography must be legible. - Do NOT render brand logos for any brand the caller has not provided as a reference image. The user's typed option name is the only "brand reference" — render it as a typographic caption, never as a logo mark. - No people in frame unless the comparison is about a service that requires depicting a person (rare). - Warm, neutral lighting; avoid the glossy AI-render look. - Avoid offensive juxtapositions (e.g. when comparing two cities, do NOT lean on stereotypes; show architecture, landscape, and time-of-day). ``` --- ### Call: Verdict TTS playback Model: `gemini-3.1-flash-tts-preview` · n/a · n/a ``` Voice: calm, even, unhurried — like a friend reading you the recommendation while you walk the dog. Pick the Gemini 3.1 Flash TTS voice whose languageCode matches the verdict's locale (typically en-US, en-GB, en-AU, de-DE, fr-FR, es-ES, pt-PT, pt-BR, it-IT, etc.). Use case: read the verdict_paragraph aloud, plus an optional list of the supporting_rows and compromise_rows. Pre-process the text before sending to TTS: - Source the text from Verdict.verdict_paragraph. - Strip any inline citation markers like "[1]" or "[2]" before sending to TTS — they do not belong in spoken output. - At sentence boundaries, insert an ellipsis ("…") for a natural pause. At paragraph boundaries (if the verdict block has more than one paragraph), insert a blank line plus an em-dash ("—"). Gemini 3.1 Flash TTS does NOT support SSML — these textual cues are how pace is conveyed. - Pronounce option names verbatim. Mid-call voice switching is not supported; if the option names are in a different language than the verdict paragraph, keep the whole reading in the dominant-language voice. - Target rate: ~140 words per minute — natural conversation pace. Style direction: prepend ONE short directive sentence to the text input, exactly like: "Read this comparison verdict calmly and unhurriedly, like reading it to a friend who is deciding what to buy. …". There is no separate `style` API field on Gemini 3.1 Flash TTS; the directive sentence inside the input is how style is conveyed. Phoneme overrides for brand names and product names are NOT exposed by Gemini 3.1 Flash TTS — no SSML `` tag. Pronunciation follows the chosen voice's native locale. ``` ## 5. Use cases & content to include Build dedicated UI flows or example seeds for each of these — they tell you what content the app must support out of the box. - **Two strollers, city walking + grocery runs.** A parent at 9pm types "Bugaboo Fox 5 vs UPPAbaby Cruz V2 for city walking + weekly grocery runs, two-year-old + a baby on the way". The table comes back with folded weight, fold dimensions, recline angle, max child weight, basket capacity in litres, suspension type, wheel diameter, current price across three retailers, parts-replacement cost, and three positive + three negative review themes per option from the last sixty days. Verdict paragraph names the Cruz V2 as the closer match for the stated priorities, with the suspension trade-off named explicitly. - **Two laptops, code + photo edit.** A freelance developer- photographer compares "MacBook Pro 14 M5 base vs Framework 16 Ryzen AI HX 370" for "Lightroom + Xcode + occasional Premiere". Rows: chip + GPU, RAM, storage, panel (size, brightness, refresh, colour gamut), ports, weight, battery life under sustained load, repairability score, current price, parts availability, owner-reported thermals. - **Two cities, family relocation.** A couple comparing "Lisbon vs Valencia" for "remote-work + two kids in primary school + weekend hiking access". Rows: cost-of-living index, median rent for a 3-bed, climate type, walkability, public-transport coverage, school options (international + state, with a named source for the ratings), broadband availability, healthcare access, crime index, residency-permit path for the family's nationality, and the typical sunshine hours. - **Two job offers.** A senior engineer comparing two SF-bay offers: "Acme Corp Sr. SWE vs Beta Inc Staff SWE". Rows: base, equity (strike, cliff, vest schedule, refresh policy from the handbook), signing bonus, healthcare premium (and out-of-pocket max), PTO, parental leave, remote policy from the handbook, total comp range from Levels.fyi for the title, recent layoff signals from the press, Glassdoor ratings (with date), and commute time from the user's stated home postcode. Verdict mentions the "not financial advice" disclaimer and surfaces equity volatility as a compromise. - **Two car-insurance policies.** A driver in Cork comparing two annual policies. Rows: annual premium, excess, no-claims protection, breakdown cover, courtesy car, replacement-glass cover, foreign-use cover (days), tracker discount, customer complaints ratio from the regulator's most recent published report, and verbatim policy-document highlights for the driving-other-cars clause. - **Two flats to rent.** A renter in Lyon comparing two flats. Rows: monthly rent, charges, surface area, number of rooms, floor + lift, DPE energy rating, deposit, agency fees, the neighbourhood's average rental per m², distance to the nearest tram, supermarkets within 500 m, and a "what's the catch" surfaced from the listing's fine print or recent street-level signals. - **Two language-learning apps.** A learner comparing two apps for "conversational Japanese, 30 min/day, six-month goal". Rows: pricing tier, free-tier limits, languages offered, pedagogy (spaced repetition? CEFR-aligned? AI tutor?), speaking practice mode, offline mode, parent-account controls, three positive + three negative themes from app-store reviews in the last sixty days, and the typical user-reported time-to-A2 from the community. - **Two ovens.** A home cook comparing two built-in ovens. Rows: cavity volume, energy class, steam mode, pyrolytic clean, fan + grill modes, telescopic rails, hob-pairing compatibility, current price, installation cost estimate, replacement-parts availability, warranty length, and recall history from the manufacturer. - **Two e-bikes.** A commuter comparing two e-bikes for "12 km flat commute with one steep hill, locked outside overnight". Rows: motor type + position, battery Wh, claimed range (and observed range from reviews), charging time, weight, frame size availability, hydraulic vs mechanical brakes, theft-rating from a named source (e.g. Sold Secure or ART), price, parts network, and warranty. - **Two mortgage offers.** A buyer comparing two 5-year fixes from named lenders. Rows: APR, arrangement fee, early-repayment charge, max LTV, valuation fee, broker fee, the lender's standard variable rate (so the user sees the cliff at the end of the fix), the lender's affordability calculator output for the user's stated income, and the lender's customer-complaints ratio from the regulator's most recent published report. - **Two car-rental options for the same trip.** A traveller comparing two car-rental deals for a one-week Crete trip. Rows: total price all-in, mileage limit, fuel policy, insurance (and excess), deposit on collection, one-way fee, customer-rating from a named aggregator in the last sixty days, and the depot's actual location vs the airport (sometimes the "airport" depot is a 20-minute shuttle ride). - **Two pension funds.** A 35-year-old comparing two passive global-equity funds. Rows: TER, tracking difference, fund size, dividend treatment (accumulating vs distributing), domicile, tax efficiency in the user's country, regulator status, and the fund's factsheet-dated month. The verdict surfaces the "not financial advice" disclaimer prominently. - **Two newsletter platforms.** A writer comparing two newsletter SaaS. Rows: free-tier subscriber limit, pricing curve at 1k / 5k / 25k subscribers, sponsor- marketplace status, custom-domain support, paid- subscriber payouts (and platform cut), import tool, export tool, deliverability rating from a named third-party tester, and the platform's TOS clause for content ownership. - **Two private-school applications.** A parent comparing two named secondary schools. Rows: most recent inspection rating from the relevant authority, fees, pupil-to-teacher ratio, GCSE/IB results for the most recent published year, university-destination breakdown (with the year), extra-curricular offerings, transport options, bursary availability. The verdict surfaces the "not advice" disclaimer plus a stale-sources warning if the inspection report is older than 24 months. - **Two snow tyres.** A driver in Helsinki comparing two studded winter tyres. Rows: TUV / ADAC most-recent- year ranking, wet braking, dry braking, ice braking, noise, rolling resistance, price, EU tyre label class, speed rating, and the manufacturer's tread-life guarantee. ## 6. Page structure The app is one continuous flow. Each section below is a discrete visual block on the same page (or screen on mobile). Build them in this order top-to-bottom. - **The compare entry block** — at the top, a single text input with placeholder "Two things, comma-separated — and what matters to you. e.g. *Bugaboo Fox 5 vs UPPAbaby Cruz V2 for city walking + grocery runs*". Below the input, three small pill buttons: *Pick decision kind*, *Set country + currency*, *Use my history*. - **Below the input, a one-line how-it-works strip** — "Every cell is sourced. Nothing is invented. Two becomes three when you need it." Toned-down, monospace caption. - **The comparison table** — appears once the user submits the entry. Sticky header with the option columns; left column is the dimension labels; cells are the values + citation chips. The table is scrollable horizontally on mobile (CSS subgrid, min column widths). Each dimension label is tappable to surface the "why this row" inline explainer. - **The reviews bundle below the table** — for each option, a card with three positive themes and three negative themes, each linked to its source thread. Below that, the "what's the catch" callout in a yellow-tinted block. - **The verdict block** — large card below the table. Verdict paragraph in serif type at 18-19 px. Below the paragraph, the supporting-rows and compromise-rows in a two-column mini- grid. Below that, the required-disclaimers strip — never buried, always visible. - **The action row** — beneath the verdict: *Add a third option*, *Export PDF*, *Send to Sheet*, *Share link*, *Read it to me*. - **The compare history strip** — at the bottom, a horizontal scroller of past comparisons; each card shows the option names, the decision kind, and the date. Tap to re-run. ## 6b. First-visit onboarding A three-step onboarding the first time a user lands on the app. Keep it tight; the goal is "tap once, you're comparing". 1. **What kind of decision?** A grid of 11 decision-kind tiles (`physical_product`, `service_policy`, `job_offer`, `city_or_neighbourhood`, `property_listing`, `software_subscription`, `travel_option`, `vehicle`, `financial_product`, `education_option`, `other`) — or "let the model guess". One tap. 2. **Country + currency.** Pre-filled from IP geolocation, the user can override. Two dropdowns. One tap each to confirm. 3. **Your two options.** Two text fields and a one-line "what matters to you" field. Tap *Compare*. The dimension picker call fires and the table starts populating. After onboarding, the same flow lives at the top of every page as the compare-entry block; the onboarding never re-shows unless the user signs out. ## 6c. Capabilities info button Visible in the header on every page. Opens a side panel naming the exact Gemini calls used in this app: - "Each comparison runs four Gemini calls: a dimension picker (`gemini-3.5-flash`, low thinking), a grounded comparison (`gemini-3.5-flash`, medium thinking, Google Search), a verdict writer (`gemini-3.5-flash`, medium thinking), and optionally a hero image (`gemini-3-pro-image` / Nano Banana Pro). Every cell with a value links to the source the model found via Google Search." - "Your comparison history, your priorities, and the citations the model fetches are not used to train any AI model. We use the Gemini API on the paid tier, where Google does not use your content for model training, per the Gemini API Additional Terms." - A small "what each call costs at today's Gemini pricing" line with a link to the detailed cost breakdown below. - A small footnote: "We use Gemini 3.5 Flash as the default — it shipped at Google I/O 2026 and beats the previous Pro tier on grounded reasoning at about 4× the speed." ## 6d. Detailed cost breakdown A second panel from the capabilities button, surfacing the honest per-comparison cost so the developer can reason about unit economics. Numbers are illustrative; refresh them against the live Gemini pricing page on deploy. - **Dimension picker** (`gemini-3.5-flash`, low thinking, no tools, `responseSchema`): ~2k input + ~1.5k output → ~$0.017 per call at $1.50 / $9 per 1M. - **Grounded comparison** (`gemini-3.5-flash`, medium thinking, Google Search): ~10k input + ~6k output → ~$0.069 per call. Google Search grounding adds the per-call grounding cost per the current pricing page. - **Verdict writer** (`gemini-3.5-flash`, medium thinking, `responseSchema`): ~25k input (the populated table) + ~1k output → ~$0.046 per call. - **"What's the catch" extractor** (`gemini-3.5-flash`, low thinking, Google Search, per option): ~3k input + ~1k output → ~$0.014 × 2 options = ~$0.028 per comparison. - **Hero image** (`gemini-3-pro-image`, Nano Banana Pro, optional): approximate ~$0.02 per image generated (Google has not pinned an exact public per-image figure; verify before shipping). - **TTS** (`gemini-3.1-flash-tts-preview`, optional): approximate ~$0.005 per verdict read (the exact TTS character-token price was not pinned at I/O 2026; verify before shipping). - **Cached input** (`gemini-3.5-flash`, $0.15/M cached): when a user runs the same comparison with different priorities, the table cells are cached on the verdict re-write, dropping the re-verdict cost to ~$0.005. Total typical comparison: **~$0.16-$0.18 fully-loaded** including hero image and TTS. Without optional extras: **~$0.13**. The "add a third option" path re-runs the grounded comparison call for the third column only and re-runs the verdict; ~$0.10 additional. ## 7. Design language The aesthetic is **product**, not memoir. This is a tool a visitor uses once and immediately understands. Voice is jobs- to-be-done universal — not a named persona, not a story. - **Typography:** SF Pro / Inter for UI; New York / Newsreader for the verdict paragraph (one serif moment to give the recommendation weight). Body 16 px, line-height 1.55. Verdict paragraph 18-19 px, line-height 1.6. - **Colour:** mostly white / near-white background (#ffffff, #f5f5f7 cards). Single accent: cobalt violet for the citation chips and the primary CTA (#5b3df2 base, #ede9fe tint). Disclaimer strip: warm sand (#fef9c3 background, #ca8a04 text). Status: success #16a34a / #dcfce7; danger #dc2626 / #fee2e2. - **Table structure:** the table is the hero. Wide on desktop, horizontally scrollable on mobile with sticky first-column dimension labels. Each cell has a 12 px citation chip in the bottom-right corner; tap or hover reveals the source preview card. - **Citation chips:** small, low-contrast by default; on hover/tap they pop with the source title, the publisher domain, the published date, and a "fetched at" timestamp. Stale citations show a soft yellow border. - **Verdict block:** larger card with a thin left-side accent bar in cobalt violet. Verdict paragraph in serif. Below that, the supporting-rows / compromise-rows mini-grid in monospace small caps for the row labels (visual marker that these are dimension-references, not prose). Disclaimers strip is full-width below in the warm-sand tint. - **No 3D, no glass, no soft shadows on cards.** Subtle hairline borders (1 px, 8% opacity) define structure; cards rely on background-tint contrast. - **Empty states:** small hand-drawn-looking dotted outlines with the line "couldn't verify — try a more specific name or a manufacturer URL". - **Motion:** restrained. Citation chips fade in at 120 ms as they attach to cells; the verdict block slides up 12 px on reveal. Reduced-motion preference disables both. - **Iconography:** lucide-react, mono-stroke, no fills. - **Mobile-first:** the entry block is full-width with a big tap target; the table is horizontally scrollable; the verdict block stacks the supporting/compromise grid into a single column. Test at 375 px, 768 px, 1440 px. ## 8. Content generation rules The content is the rows, the cells, the verdict. The rules: - Every cell value must come from a sourced fetch via `google_search`. If no source was found, the cell shows "couldn't verify" — never an invented number, never a "typically around" hedge. - Currency in the retailer's charging unit. No model-side conversion. - Review themes capture *recurring* patterns (≥3 distinct sources), not single-reviewer complaints. - The verdict paragraph re-states the user's priorities verbatim near the start so the reader sees the anchor. - The verdict paragraph uses recommendation language ("for your stated priorities, X wins on rows A and B") not advice language ("you should buy X"). - Disclaimers (financial, medical, safety-adjacent, legal- adjacent, stale sources, not-advice) are surfaced in the verdict paragraph itself, not as a footnote. - The "what's the catch" callout is required to be a recurring pattern surfaced from reviews, not from marketing copy. - Brand names, model names, city names, school names, fund names are not translated. "UPPAbaby" stays "UPPAbaby" in every locale. - When adding a third option, the table re-grounds for the new column only; existing columns keep their citations unless the user explicitly asks for a refresh. ## 8a. Seed content Ship the app with five pre-built comparison templates the visitor can tap to see the magic without typing anything: 1. **"Bugaboo Fox 5 vs UPPAbaby Cruz V2 — city walking + grocery runs"** — two strollers; the demo comparison. 2. **"MacBook Pro 14 M5 base vs Framework 16 Ryzen AI HX 370 — code + photo edit"** — two laptops. 3. **"Lisbon vs Valencia — remote-work + two kids in primary school + weekend hiking"** — two cities. 4. **"Acme Corp Sr SWE vs Beta Inc Staff SWE — 5-year horizon"** — two job offers (fictional company names for the seed; the real-decision use case is the same shape). 5. **"Duolingo Max vs Busuu Premium — conversational Japanese, 30 min/day, six-month goal"** — two language apps. Each seed comparison can be re-run in one tap, refreshing the underlying citations against current sources. ## 9. Media & assets - **Hero side-by-side image per comparison** — generated with Nano Banana Pro at the time of comparison; stored in Firebase Storage; reused on the share card and PDF cover. Regenerate on demand if the user dislikes the first attempt. - **App icon and favicon** — a flat two-column glyph (left column violet, right column sand) suggesting the comparison table. Vector SVG, no shadows. - **Open Graph card** — generated per comparison with the two option names + the verdict's one-line recommendation. Uses the hero image as background, with a 60% white overlay so the typography stays legible at small sizes. - **PDF export** — paginated, citation footnotes preserved as numbered references with URLs at the bottom of each page; hero image on the cover; verdict paragraph in serif on page 2; table on pages 3+. - **No stock photography.** The app is data-dense; the hero image is the only generated visual. - **Loading state:** skeleton rows in the table, then a per-cell fade-in as each cell's value + citation chip resolve. Avoid spinners — let the user see the table fill in one cell at a time. ## 10. Interactivity & states - **Empty state:** entry block + the five seed comparisons as pre-built cards. - **Loading state:** table scaffold visible immediately (dimension labels + option columns + empty cells); cells fill in one by one as the grounded call streams. - **Streaming state:** as the grounded call streams, cells light up one at a time with their value + citation chip. The verdict block stays greyed out until every cell has resolved or been marked unverified. - **Hover state on a citation chip:** a small preview card shows the source title, publisher, publication date, and fetched-at timestamp; tapping the chip opens the source URL in a new tab. - **Hover state on a dimension label:** a small tooltip surfaces the "why this row" one-liner. - **Edit-priorities state:** a slim editor below the verdict lets the user rewrite the priorities; on save, only the verdict-writer call re-runs (no re-fetch of the cells); the verdict block animates a 200 ms cross-fade. - **Add-an-option state:** a "+ add option" tile appears to the right of the rightmost column on hover; tap to surface a small modal with the option name + URL field; on submit, the grounded call re-runs for the new column only. - **Stale state:** any citation older than the freshness window shows a soft yellow border on the chip; a table-header strip says "3 cells have sources older than 60 days — refresh?". One tap re-grounds those cells only. - **Couldn't-verify state:** cells show the literal text "couldn't verify" in muted small-caps; tap to reveal the queries the model tried and a manual-input fallback. - **Error state:** if the grounded call fails (rate limit, network), the table scaffold remains; a retry button appears in the table header; existing verified cells are preserved. - **Reduced-motion state:** all animations replaced by instant transitions; the streaming reveal becomes a single "table populated" reveal. - **Offline state:** the app falls back to the user's last cached comparison; a banner says "you're offline — the cells you see are from . Online refresh available when reconnected." ## 11. Tech & responsive requirements - **Stack:** React + TypeScript on the front-end. Cloud Run server functions for the four orchestrated Gemini calls. Firestore + Firebase Auth + (optional) Firebase Storage. Free 2-app Cloud Run deploy from AI Studio Build. - **State:** server-side state in Firestore (`comparisons`, `cells`, `citations`, `verdicts`). Front-end uses React Query / TanStack Query for cache and optimistic UI. - **Streaming:** the grounded comparison call uses `generateContent` with streaming enabled; the server forwards cell-level updates via SSE to the client so the table populates one cell at a time. - **Schema validation:** every Gemini response is parsed through the matching Zod schema server-side. Unverified cells (no citation attached) are forced to `verified: false`. - **Long-context guardrails:** at >500k tokens of payload, the request is rejected with an honest error before hitting the API. - **Performance:** first cell must paint within 1.5 s of submit (dimension picker + scaffold render). Full table populated within 30 s on a typical product comparison. Verdict block within 5 s of the last cell resolving. - **Responsive:** test at 375 px (iPhone SE), 768 px (iPad), 1440 px+ (desktop). Mobile: vertical scroll, horizontal table scroll with sticky first column. Tablet: same as mobile but with the verdict block beside the table on landscape. Desktop: side-by-side layout with the entry block sticky on the left rail. - Use `clamp()` for fluid typography (e.g. `clamp(0.95rem, 0.85rem + 0.4vw, 1.05rem)` for body, `clamp(1.6rem, 1.2rem + 2.4vw, 2.6rem)` for H1). Prefer container queries over media queries for component-level responsiveness (the comparison table, the verdict card, the entry block all use `container-type: inline-size`). - Use `dvh` / `svh` instead of `vh`. Respect safe-area insets on iOS — the sticky entry rail on desktop and the bottom action bar on mobile must clear the home indicator via `env(safe-area-inset-bottom)`. - Images: WebP / AVIF with `` and `srcset` for the optional hero; every `` carries explicit `width`/`height` + `loading="lazy"` to prevent CLS. - Performance budget: LCP < 2.5 s, INP < 200 ms, CLS < 0.1 on mobile Safari iPhone 12+. Initial JS bundle ≤ 200 KB gzipped; the table-streaming view lazy-loads. - **Browser support:** latest two versions of Safari, Chrome, Firefox, and Edge. Verify all visual changes on Safari AND Chrome before shipping. - **Caching:** Gemini cached input ($0.15/M) is used on the verdict-writer call when the same comparison runs more than once. - **Native Android (post-I/O 2026):** the same comparison flow exports as a Kotlin + Jetpack Compose app via AI Studio Build's native-Android target. The Android build pulls citations live from the same Cloud Run functions. ## 12. Accessibility (WCAG 2.2 AA) - **Semantic HTML:** the table is a real `` with ``, ``, `
` for option columns, `` for dimension labels. - **Keyboard navigation:** every interactive element (citation chip, dimension label, action button, add-option tile) reachable by Tab in source order; focus styles are visible (2 px violet outline + 1 px white halo). - **Focus not obscured:** the sticky table header never covers a focused cell — when focus crosses under the header, the table scrolls to keep focus visible. - **Target size:** every chip, button, and tap target is at least 44 × 44 px (48 × 48 px under coarse pointer per WCAG 2.2 AA + the mobile preamble). Citation chips can render as 24 px squares visually but their hit area extends to 44 × 44 px via transparent padding so touch users never miss. - **Contrast:** body text against background ≥ 7:1; muted text ≥ 4.5:1; the cobalt-violet accent against white meets ≥ 4.5:1 for non-text UI components and ≥ 7:1 for body text uses. - **Screen-reader labels:** - Each cell announces "row label, option name, value, source publisher". - The verdict block announces "Recommendation for your priorities: …" before reading the paragraph. - The "couldn't verify" cells announce "no current source found for this cell — queries tried: …". - **Live regions:** the streaming table populates inside an `aria-live="polite"` region so screen-reader users hear cells as they resolve, without being interrupted. - **Reduced motion:** `prefers-reduced-motion: reduce` disables the cell-fade animation, the verdict slide-in, and the citation-chip transition; everything snaps in place. - **Colour-blind safety:** the soft yellow "stale" border is paired with an icon glyph; the violet citation chip is paired with a small "[1]"-style numeric reference; never rely on colour alone. - **Language:** the `` attribute reflects the app's runtime locale; per-cell language is `lang`-tagged when the cell value is in a different language from the page (e.g. a Japanese product name). - **TTS controls:** the verdict TTS playback has explicit play / pause / stop buttons, a visible progress bar, and a captioned transcript (which is the verdict paragraph itself, so no extra work). ## 13. Quality bar — avoid AI clichés - **No invented numbers.** Ever. "Couldn't verify" beats "approximately around 8 kg" every time. - **No advice voice.** "For your stated priorities, the Cruz wins on basket capacity" beats "you should buy the Cruz". - **No generic verdicts.** A verdict that does not name specific rows is no verdict at all — it must cite the evidence. - **No emoji-laden cell content.** Cells contain values and citations, not "✅", "❌", or "⭐". Visual indicators belong in the UI chrome, not the data. - **No "as an AI" hedges.** The verdict is direct: a recommendation grounded in the table. - **No "in conclusion" or "overall" filler.** The verdict paragraph opens by anchoring to the priorities, names the supporting rows, names the compromises, and ends on a one-sentence honesty note when relevant. - **No infinite "it depends" outputs.** When the data is thin, the verdict says so honestly and either makes a conditional lean or returns null for the `recommended_option_id` with `recommendation_strength: close_call`. - **No SEO-style comparison prose.** The table is the artefact. Paragraphs surrounding the table are kept short (the verdict block is the only prose moment). - **No "best of 2026" framing.** This app makes a comparison for the user, not a ranking for the internet. - **No fake review quotes.** Review themes capture patterns; the `recurring_phrase_verbatim` field uses short common phrasings, never a full sentence from a single review. - **No marketing-deck-style "by the numbers" strip** at the top of the page. The data lives in the table; the table is the marketing. ## 14. Deliverables A built AI Studio Build app at the end of this template includes: - **The front-end app** — React + TypeScript, deployed to Cloud Run via the free 2-app Cloud Run deploy from AI Studio Build. Lighthouse score ≥ 95 on Performance, Accessibility, Best Practices, SEO. - **The Cloud Run server functions** — four orchestrated Gemini calls (dimension picker → grounded comparison → verdict writer → optional hero image + TTS + catch extractor). All secrets in env vars; `.env.example` shipped. - **The Firestore schema** — `users`, `comparisons`, `dimensions`, `cells`, `citations`, `verdicts`, `shared_links`, `compare_history`. Security rules locked to the owner; shared links are read-only and revocable. - **The five seed comparisons** — pre-populated in Firestore so a fresh-install visitor sees the magic on first paint. - **The capabilities + cost panels** — wired to the current Gemini pricing page; numbers refreshed at deploy time. - **The PDF + Sheet export** — Cloud Run function renders a print-clean PDF with citation footnotes; Workspace integration pushes the same comparison to a Sheet. - **The native Android target** — Kotlin + Jetpack Compose app built from the same AI Studio Build project, publishable to Play Internal Test (post-I/O 2026 capability). - **A test plan** — manual checklist covering the five seed comparisons end-to-end, the add-third-option flow, the stale-citation re-check, the offline state, the reduced-motion state, the screen-reader walkthrough, and the contrast audit at 100% / 200% / 400% zoom. - **Honest open caveats:** Gemini 3.5 Pro is announced but pre-GA as of 2026-06-01 — pin 3.5 Flash. Computer Use is not used in this template; if a future "click the retailer's add-to-cart" flow is added, it requires `gemini-2.5-computer-use-preview-10-2025`. Gemini Omni Flash (the I/O 2026 video model) has no public developer API yet and is not used here. Managed Agents (`antigravity-preview-05-2026`) could replace the orchestrated four-call chain with a single agent call, but adds preview-status risk; pinned-as-orchestrated for v1.