RAG Eval Dataset: 200 Hand-Written Questions Beat Synthetic Ones

Hand-write about 200 RAG eval questions from real traffic, with gold answers owned by one domain expert. Score retrieval and generation separately, gate every PR, and use synthetic generators only for retriever coverage.

By Rajesh Beri·October 4, 2026·13 min read
Share:
A wooden desk with a printed stack of question sheets, a red pen marking pass and fail ticks beside each line, and an open binder of policy documents with coloured tabs, lit by a desk lamp.

Illustration generated using AI

If your team argues about RAG quality in meetings, the fix is about 200 questions written by hand from real user traffic, each with a gold answer and the document that proves it. Score retrieval and generation separately, and run the set on every pull request. Buy nothing until that set exists. Once it does, keep it in the tracing tool you already pay for; if you have none, use Langfuse, which is MIT-licensed to self-host and $29 a month on its Core cloud plan.

A RAG evaluation dataset (also called a golden set) is a fixed list of questions, each paired with the answer a domain expert accepts and the source passages that support it, used to score a retrieval-augmented system the same way every time it changes.

Synthetic generators are useful for one job: stress-testing the retriever. The loser here is the fully synthetic set, and the product that loses on price for this workload is Confident AI's hosted platform, whose free tier caps you at five test runs a week.

Option What it does for your dataset Price (checked October 4, 2026) Pick it if Do NOT pick it if
Hand-written set in a spreadsheet + git You write questions, gold answers, source IDs Free; costs expert hours Always, as the source of truth Never skip it
Ragas Synthesises questions from a knowledge graph of your docs Free, Apache-2.0; you pay judge-model tokens You need retrieval coverage of a big corpus fast You want it to rank generator models
DeepEval / Confident AI Synthesises "goldens"; pytest-style runner Library free (Apache-2.0); platform free tier 5 runs/week, then $200/month Your engineers live in pytest You need a hosted CI gate on a small budget
Langfuse Versioned datasets, links items to production traces Hobby free (50k units, 30-day retention); Core $29/month (100k units) Default pick; self-host required You need a turnkey PR comment out of the box
Braintrust Datasets, experiments, GitHub Action posts PR diffs Starter $0 (10k scores, 14-day retention); Pro $249/month You want the best PR review loop and can pay You must self-host without an enterprise deal
LangSmith Versioned, tagged datasets with splits Developer $0 (1 seat, 5k traces); Plus $39/seat/month (10k traces) You already run LangChain or LangGraph You pay per seat for a large review team

How Many Questions Is Enough?

Two hundred questions is enough to see a regression of about five points, and 50 is not enough to see anything smaller than a coin flip's worth of noise. The math is the ordinary standard error of a pass rate, which Anthropic's guide to statistics for model evals asks every eval to report alongside the score.

With binary pass/fail grading and a system passing around 80%, the 95% interval on the score is roughly ±11 points at 50 questions, ±8 at 100, ±5.5 at 200 and ±4 at 400 (my arithmetic, using 1.96 × √(p(1−p)/n)). A team arguing whether a new chunking strategy moved quality from 78% to 82% on a 50-question set is arguing about noise.

Two corrections make 200 go further than it looks. First, compare versions on the same questions. Anthropic's paper recommends paired differences because models "get the same questions right and wrong" with per-question correlations of 0.3 to 0.7, which it calls "a 'free' variance reduction". At a correlation of 0.5, the smallest change 200 paired questions can resolve drops from about 8 points to about 5.5. Second, watch for clustering: if you wrote ten questions off the same policy PDF, they are not independent, and the paper warns that ignoring this "will lead us to underestimate the standard error." Cap any single source document at a handful of questions.

We covered the same arithmetic for model bake-offs in Only 3 of 36 Model Gaps Were Real. The rule carries over: size the set to the smallest difference you would act on.

Where Should the 200 Questions Come From?

Take them from your own logs first, then from the people who will be blamed when the system is wrong, and only then from a generator. A workable mix for a system that already has users:

  • About 120 from production traffic, sampled across your query segments rather than the most frequent queries. Hamel Husain's evals FAQ says to "start with 100 diverse traces and annotate at least the first 30 yourself," continuing until new reviews stop revealing new failure modes.
  • About 50 from the domain owner: the questions they know the system gets wrong, the ones with regulatory consequences, and the ones where the right answer is "the policy doesn't say."
  • About 30 adversarial or out-of-scope questions, so that refusing correctly is graded instead of assumed.

Before launch, when there are no logs, have three or four intended users write questions in their own words without looking at the documents. That last condition matters. A question written while reading a chunk borrows its vocabulary, and that makes retrieval look better than it will be.

The evidence on synthetic questions is specific. A 2025 study by van Elburg, van der Putten and Marx compared synthetic and human-labelled benchmarks on four datasets of about 100 questions each, so weigh it as one small study. Synthetic sets "reliably rank the RAGs varying in terms of retriever configuration," with Kendall's τ of 0.67 to 0.84 on BLEU, ROUGE and semantic similarity (context precision agreed at only 0.08). When the systems differed by generator model, rankings showed "substantial inconsistencies" and sometimes inverted, partly because synthetic questions were "more specific and technical" than real ones and partly from a possible stylistic bias toward the GPT-4o model that generated them.

So use Ragas testset generation or the DeepEval Synthesizer to produce a second, larger retrieval set: Ragas defaults to 50% single-hop specific questions and 25% each of two multi-hop types, and DeepEval runs Evol-Instruct-style "evolutions" to make inputs harder. Jason Liu's version of this is to generate questions for every chunk and treat it as a floor: "Synthetic data should just be around 97% recall." If you can't hit that on questions written from your own chunks, real users will do worse. Keep that set out of model-selection decisions.


Who Writes the Gold Answers?

One named domain expert owns the gold answers, and engineers do not write them. Husain's FAQ recommends "a single domain expert as a 'benevolent dictator'" because one consistent standard beats a committee's average, and outsourcing labelling breaks the loop between what the expert sees and what the team fixes.

Each item needs four fields: the question, a short gold answer stating the facts that must appear, the IDs of the documents that support it, and a pass/fail rule in one sentence. Grade pass/fail. The same FAQ argues that "binary evaluations force clearer thinking and more consistent labeling," and a 1-to-5 scale lets a reviewer park every hard case at 3.

Expect the rubric to change once the expert sees outputs. Shankar et al. call this criteria drift: "users need criteria to grade outputs, but grading outputs helps users define criteria." Budget a second pass over the first 50 answers after the expert has graded 100 system outputs. At roughly three minutes per item, 200 gold answers plus that revision is two or three working days of one expert's time, which is the real price of this whole exercise.

What Should You Measure, and Why Separately?

Score retrieval and generation as two numbers, because a single "answer quality" score cannot tell you which half broke. Ragas splits its metrics the same way: context precision and context recall on the retrieval side, faithfulness and response relevancy on the generation side.

Retrieval: because every item carries its supporting document IDs, recall@k is a deterministic count with no judge model involved. Liu warns that "most people are spending too much time on the actual synthesis" before checking whether the right data was retrieved. Report it per segment too: an overall 70% can hide a 30% on the questions that matter.

Generation: score two things on items where retrieval succeeded. One is correctness against the gold answer (does it state the required facts). The other is faithfulness (is every claim supported by the retrieved passages). Both need an LLM judge, and the judge needs its own check: have the domain expert grade 50 items blind and compare. If the judge agrees on fewer than about nine in ten, fix the judge prompt before you trust a trend line.

Skip off-the-shelf "helpfulness" or similarity scores as the headline number. Husain's warning is blunt: "Generic evaluations waste time and create false confidence when you use them as quality measures."


How Do You Catch Regressions in CI?

Run the full 200 on every pull request that touches prompts, chunking, embeddings, the retriever or the model, and fail the build on a drop larger than your interval. With paired comparison against the main branch, a five-point drop in retrieval recall or generation pass rate is a reasonable first threshold for a 200-question set.

Here is the workload the prices below are normalised to: 200 questions, 40 gated pull requests a month, so 8,000 eval runs a month, each logging one trace, three spans and two judge scores.

  • Langfuse counts billable units as "Count of Traces + Count of Observations + Count of Scores". That is six units per run, or 48,000 a month, which fits the free Hobby tier's 50k with almost no headroom and sits comfortably in Core's 100k for $29 a month (prices checked October 4, 2026). Its datasets create a new version on every add, update, delete or archive, and items can link to the production trace they came from. The repo is MIT licensed "except for the ee folders," so self-hosting the dataset features is free. You write the CI comparison script yourself.
  • Braintrust is the strongest PR loop: its GitHub Action will "post a live summary comment on the associated pull request" with score regressions and improvements. The workload produces 16,000 scores a month. The Starter tier includes 10k, then $2.50 per 1k, so about $15 in overage, but retention is 14 days, which is too short to compare against last month. Pro is $249 a month with 50k scores and 30-day retention (checked October 4, 2026). If you need on-premises hosting, you are into an enterprise contract.
  • LangSmith versions datasets on every change, lets you tag versions and evaluate on splits. The Plus plan is $39 per seat a month with 10k base traces; usage beyond that bills in LangChain Standard Units at $1.00 per LSU, and the page does not state a per-trace conversion we could normalise (checked October 4, 2026). Base traces keep 14 days. Five reviewers cost $195 a month before usage. Choose it if your pipeline is already LangGraph; don't if your labelling team is large, because seats are the meter.
  • DeepEval is the best fit for engineers who think in tests: it is "similar to Pytest but specialized for unit testing LLM apps," and deepeval test run drops into any CI runner for free. Its hosted platform, Confident AI, allows 5 test runs a week on the free tier. Forty gated PRs a month needs roughly ten a week, which puts you on Starter at $200 a month (checked October 4, 2026). That is seven times Langfuse Core for this workload, which is why it loses here. Use the library and skip the platform unless you want its dashboards.
  • Ragas is a library, Apache-2.0, now maintained by Vibrant Labs. It has no hosting and no gate. Pair it with any of the above for metrics and the synthetic retrieval set.

We compared these platforms as regression gates in more depth in Braintrust vs Langfuse vs Promptfoo. On storage and retention costs at scale, see AI Observability Pricing.

How Do You Keep the Set Current as the Corpus Moves?

Treat the dataset as code with an owner and a quarterly review, because a gold answer that cites a superseded policy will fail a correct system and teach the team to ignore red builds. Three rules hold up:

  1. Version every change. Langfuse and LangSmith both create a new dataset version per edit; if you keep the set in git, the commit is the version. Record which version each CI score came from, or a score jump after an edit looks like a model improvement.
  2. Retire by archiving, never by deleting. When a source document is replaced, archive the items that cite it (Langfuse supports an ARCHIVED status for exactly this) and write replacements from the new document. Keep the archived items for audit.
  3. Add from failures. Every production incident and every thumbs-down the expert agrees with becomes a candidate item. Keep a fixed core of about 150 that never changes between quarters, so long-run trends stay comparable, and rotate the remaining 50.

Watch the corpus side too. If you re-embed or re-chunk, retrieval scores will move even if nothing else changed; our embedding model comparison covers what a re-index costs.


The Decision: What Predicts Regret?

The decision that teams regret is buying an eval platform before anyone has written a gold answer. The platform will happily generate 1,000 synthetic questions, the dashboard will show 91%, and the meeting arguments will continue, because nobody in the room believes questions a model wrote from your own chunks.

Three criteria predict regret more than any feature list:

  • Retention versus your release cycle. If you compare against last quarter's baseline, a 14-day retention tier (Braintrust Starter, LangSmith base traces) will drop the comparison before you need it. Export results to your own store or pay for longer retention.
  • Who owns the gold answers. If the answer is "the RAG team," the set will grade what the team already believes is correct.
  • Whether your data may leave your network. If it may not, Langfuse self-hosted or the DeepEval and Ragas libraries are the realistic options without an enterprise negotiation.

What changes the answer: if you already pay for LangSmith or Braintrust, store the set there and skip the migration. If you have no users yet, a synthetic set is a reasonable bootstrap for retrieval only, and every synthetic question should be replaced by a real one within the first few months of traffic. If your corpus is small (a few hundred documents), 100 hand-written questions with paired comparison may be enough. Accept ±8 points and say so in the report.

This Week: name the domain owner, pull 300 production queries across your segments, and have the owner pick and answer the first 50 with source document IDs.

This Month: finish 200 items, compute recall@k against the document IDs with no judge involved, and calibrate an LLM judge against 50 blind expert grades.

Before Q1 Close: wire the 200-question paired comparison into CI on every retrieval or prompt change, with a failure threshold set from your measured interval, and schedule the first quarterly refresh.

Start with the 50 questions your domain owner already knows the system gets wrong.

Continue Reading

Share:

Frequently Asked Questions

How many questions does a RAG evaluation dataset need?

About 200 for most teams. At an 80% pass rate, 200 binary-graded questions give a 95% interval of roughly ±5.5 points, versus ±11 at 50. Comparing versions on the same questions (paired differences) tightens it further.

Can I use synthetic questions from Ragas or DeepEval instead of writing them?

Use them for retrieval testing only. A 2025 study found synthetic benchmarks ranked retriever configurations consistently with human-labelled sets but gave inconsistent, sometimes inverted rankings when comparing generator models.

Who should write the gold answers for a RAG eval set?

One named domain expert, not the engineers building the system. Each item needs the question, a short gold answer listing required facts, the supporting document IDs and a one-sentence pass/fail rule.

Should retrieval and generation be scored separately?

Yes. Score retrieval with recall@k against the gold document IDs, which needs no judge model, and score generation for correctness and faithfulness with an LLM judge calibrated against expert grades. One combined score cannot show which half broke.

How often should a RAG eval dataset be refreshed?

Review it quarterly and whenever source documents change. Archive items that cite superseded documents rather than deleting them, keep a fixed core of about 150 questions for trend comparison, and add new items from production failures.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →