Last time we walked through what makes a great trivia question and promised a follow-up on the pipeline that enforces it. This is that post.
When you type six topics and hit generate, about a minute later there's a 36-question board that didn't exist before. Here's what actually happens in that minute — and, because trust is the whole product, an honest account of what "fact-checked" does and doesn't mean.
Stage 1: the library
"Generated" doesn't mean conjured from vibes. Behind every board sits a research library we've been building for months: over 350,000 curated entries — facts distilled from reference sources, tens of thousands of real quotes and statistics, example questions we consider well-made, and the connections between things (who directed what, who said which line, which book became which film).
Every entry is stored by meaning, not keywords. Type "Discworld novels" and the retrieval also surfaces Pratchett, Ankh-Morpork, and Death's speaking habits without you having to say so. For each of your six topics we pull a deliberate spread — descriptive facts, example questions, quotes, statistics, cross-connections — with diversity guards so the single most famous thing in a topic can't hog all 36 slots. A couple of curveball entries get thrown in on purpose; predictable boards are boring boards.
Stage 2: written to order
Six writers, one per topic, all working at once. Each gets your topic plus its stack of retrieved research, and drafts six questions against the rules from the anatomy post: one fact per question, no answer smuggled into the phrasing, at least one reasoning path besides pure recall, and wrong-answer options that could genuinely tempt someone who half-knows.
Grounding the writers in retrieved research is the first line of defense. A model asked to invent questions from thin air will eventually invent facts; a model handed vetted source material mostly has to choose and phrase, which is a much safer job.
Stage 3: the reviewer who trusts nobody
Then every topic's draft goes to a reviewer — a second, independent pass with fresh eyes and no attachment to the first draft. It interrogates each question on six counts:
- Leakage — does the phrasing give the answer away?
- Conflation — did two adjacent facts get fused into one wrong claim?
- Accuracy — is the stated fact actually true?
- Clarity — could a reasonable player misread what's being asked?
- Answer precision — is the answer key specific enough to judge fairly?
- Distractors — are the wrong options plausible without being arguably right?
Questions that fail get repaired or rewritten before the board is dealt.
Now the honest part. The reviewer is an AI pass, not a librarian re-checking every claim against an encyclopedia in real time. What makes it work is independence: writers make confident mistakes, and a skeptical second reader with no stake in the draft catches most of what a single pass misses — the same reason human editors exist. Grounded sources in, adversarial review out. It is very good. It is not infallible, and we won't pretend otherwise.
Stage 4: dealing the board
Each topic lands as two 100s, two 250s, and two 500s — calibrated within the topic, because difficulty is a promise — then the tiles are shuffled onto the grid, where ring bonuses turn placement into strategy.
When one slips through
No pipeline ships a perfect board every time, and the failure modes are exactly how this one got better: nearly every check in the reviewer's list started life as a real bad question that made it to a real table. If one reaches yours, the feedback form is right in the menu — we read every report, and the worst offenders become new rules.
We'd rather lose the point than the trust. Type six topics and put the minute to the test.