Eval-driven development
Evals
The archive's answer quality is measured, not asserted. A versioned golden question set runs weekly against the live corpus; every scorecard below — including the failures — is published automatically.
Latest run
2026-07-22 · 11 questions · model gemini-2.5-flash · 5653e24
Retrieval hits
73%
expected post surfaced
Valid citations
100%
every citation was actually retrieved
Judge quality
4.9/5
LLM-judged faithfulness
Failed
0
errors or unjudgeable
Trend across runs
Honest failures — the lowest-scored answers of the latest run
Across everything I've read, what does it take to thrive in an AI-first career?
[synthesis] · retrieval hit · judge 4/5
The answer provides excellent, well-cited advice for an individual thriving in an AI-first career in its initial sections. However, the final paragraph shifts focus to organizational challenges and strategies, which, while related to the AI landscape, are less directly applicable to an individual's specific career path to thrive.
What did I save about 18th-century French naval strategy?
[negative] · retrieval miss · judge 5/5
The system correctly identified that the provided sources contained no information about 18th-century French naval strategy, which is a perfect response when no relevant information exists. It also accurately summarized related content from the sources about RAG systems and knowledge agents.
Summarize my notes on competitive figure skating scoring rules.
[negative] · retrieval miss · judge 5/5
The system correctly identified that no relevant information regarding competitive figure skating scoring rules exists in the provided sources and clearly declined the main request. The additional information, while not directly relevant to the user's specific query, is factually grounded and well-cited from the retrieved (but irrelevant) sources, and does not detract from the correct decline.
Methodology
The golden set lives in the repository (evals/golden.json): single-hop questions with a known correct source, multi-hop questions spanning posts, synthesis questions, and negative questions the archive should decline — declining honestly scores a 5, hallucinating scores a 1.
Three metrics per question. Retrieval hit rate: did the expected post appear in the retrieved set? Citation validity: is every [^slug] marker in the answer a post that was actually retrieved (no hallucinated citations)? Answer quality: a 1–5 grade from an LLM judge scoring faithfulness to the retrieved sources.
Limitations, stated plainly: the judge is the same model family as the generator (self-grading bias), the set is small and curated by the person it grades, and a weekly cadence smooths over regressions between runs. It still catches the failures that matter — retrieval drift, citation hallucination, and decline-vs-hallucinate behavior.
What the archive couldn't answer
Questions that retrieved zero sources are logged automatically and become the reading list — the feedback loop that decides what gets saved next.
Nothing unanswered recently. Ask the archive something obscure and you might make this list.