David's Archive
HomeLibraryStatsAbout
Log in

Eval-driven development

Evals

The archive's answer quality is measured, not asserted. A versioned golden question set runs weekly against the live corpus; every scorecard below — including the failures — is published automatically.

Latest run

2026-07-22 · 11 questions · model gemini-2.5-flash · 5653e24

Retrieval hits

73%

expected post surfaced

Valid citations

100%

every citation was actually retrieved

Judge quality

4.9/5

LLM-judged faithfulness

Failed

0

errors or unjudgeable

Trend across runs

Retrieval hit rateCitation validity

Honest failures — the lowest-scored answers of the latest run

Across everything I've read, what does it take to thrive in an AI-first career?

[synthesis] · retrieval hit · judge 4/5

The answer provides excellent, well-cited advice for an individual thriving in an AI-first career in its initial sections. However, the final paragraph shifts focus to organizational challenges and strategies, which, while related to the AI landscape, are less directly applicable to an individual's specific career path to thrive.

What did I save about 18th-century French naval strategy?

[negative] · retrieval miss · judge 5/5

The system correctly identified that the provided sources contained no information about 18th-century French naval strategy, which is a perfect response when no relevant information exists. It also accurately summarized related content from the sources about RAG systems and knowledge agents.

Summarize my notes on competitive figure skating scoring rules.

[negative] · retrieval miss · judge 5/5

The system correctly identified that no relevant information regarding competitive figure skating scoring rules exists in the provided sources and clearly declined the main request. The additional information, while not directly relevant to the user's specific query, is factually grounded and well-cited from the retrieved (but irrelevant) sources, and does not detract from the correct decline.

Methodology

The golden set lives in the repository (evals/golden.json): single-hop questions with a known correct source, multi-hop questions spanning posts, synthesis questions, and negative questions the archive should decline — declining honestly scores a 5, hallucinating scores a 1.

Three metrics per question. Retrieval hit rate: did the expected post appear in the retrieved set? Citation validity: is every [^slug] marker in the answer a post that was actually retrieved (no hallucinated citations)? Answer quality: a 1–5 grade from an LLM judge scoring faithfulness to the retrieved sources.

Limitations, stated plainly: the judge is the same model family as the generator (self-grading bias), the set is small and curated by the person it grades, and a weekly cadence smooths over regressions between runs. It still catches the failures that matter — retrieval drift, citation hallucination, and decline-vs-hallucinate behavior.

What the archive couldn't answer

Questions that retrieved zero sources are logged automatically and become the reading list — the feedback loop that decides what gets saved next.

Nothing unanswered recently. Ask the archive something obscure and you might make this list.

David's Archive

One person's reading, compressed into a queryable knowledge base. Built with retrieval, evals, and a learning loop — and instrumented end to end.

Explore

  • Library
  • Stats
  • Evals
  • About

Project

  • RSS feed
  • Source on GitHub
  • Case study

Weekly digest

New posts, once a week. No tracking pixels.

© 2026 David's Archive · answers cite their sources