Measuring Jeopardy Question Difficulty with Jev

data-science
NLP
AI
projects
Scoring 500k clues with a structured-output LLM, then testing whether AI adds predictive information beyond board position
Author

Nick Tacik

Published

October 1, 2026

Introduction

  • In my previous Jeopardy post, I asked how nicely we can cluster the clues into different types of categories. Now I want to turn to the query of “how difficult are the questions in jeopardy, and how does that vary?”
  • There is not much signal in our jeopardy data regarding this. Obviously, if no one buzzes in with the correct answer, that constitutes evidence that the clue was more difficult than others. If one person buzzes in correctly, then we don’t know if the other two players would have known the answer, or how hard it was for them to actually come up with that answer. Fans of the show will intuit that the questions generally get harder as you go from row 1 to row 5, but what if we want some analysis of how much harder the questions are in the tournament of champions, or how much easier they are on celebrity jeopardy? Are daily doubles and final jeopardy harder than average? Were clues easier or harder in the past?
  • Questions like this are why I wanted to get an LLM to judge question difficulty. Once we are confident we can reasonably trust what it’s doing, we can begin getting these answers.
  • Jev is a new structured output model that can take text input like an LLM, but only gives structured model output - like a traditional neural network with a soft-max head. It is quite cheap ($42 per billion input tokens, and free output tokens) and fast (70-500ms response times). It’s appropriate for so-called “System 1 Shaped Queries” - queries that require fast, reactive answers, rather than deep research or multi-step reasoning. “How easy is this question on a scale of 1-10?” seems like it fits the bill, and I have a lot of questions in my dataset to evaluate, so I thought it would be a great chance to test Jev out!

Analysis

Establishing a Baseline

  • To begin our analysis, we should actually verify that board position does indeed correlate with some measure of difficulty. Starting with ~524k clues in our dataset, we produced a data set of ~412k board clues from regular season games, excluding final jeopardy clues, media clues, daily doubles, and clues without an observed outcome. In Figure 1 we plot the fraction of clues that were answered correctly. We see that board position does indeed predict correctness - the cheapest clues are answered correctly about 96% of the time, while the most expensive clues are answered correctly about 72% of the time. The relation is clearly monotonic, but not linear.
Figure 1: Fraction of clues answered correctly by row position (1=cheapest, 5=most expensive), with 95% Wilson confidence intervals. Regular-season Jeopardy and Double Jeopardy board clues.

Querying Jev

  • What query should we actually use to evaluate the difficulty of the jeopardy questions? We have several hundred thousand to evaluate, so we should fix the query before running the whole batch.
  • Jev doesn’t take a freeform prompt - it takes structured questions with a description for each score level. I sampled 500 clues stratified by board position and scored them with three rubric formulations. They all use the same instruction - You are an expert Jeopardy analyst. Rate how difficult this clue is for a typical contestant. - but we tried varying the per-level descriptions.
  • The winner was chosen on two criteria - strongest negative correlation with actual player outcomes, and median difficulty closest to 5.
Variant Description style Corr Median
Descriptive Narrative anchors (“almost anyone would know this…”) −0.182 5.0
Expected solve rate Concrete correctness rates (“~45% of contestants answer correctly”) −0.227 5.0
Concise Short labels (“median — typical for a Jeopardy contestant”) −0.191 4.0
  • The expected-solve-rate rubric had the strongest correlation and a centered median. Grounding each level in a concrete expected correctness rate appears to give the model a clearer signal than qualitative labels. Note that the rubric’s percentages refer to individual contestant success rates, while our validation measures whether any contestant answered correctly — so the rubric framing and the outcome variable aren’t measuring exactly the same thing.

  • We ran the full query set in parallel. It took about an hour and $10 worth of tokens to get everything labeled. The 500 spike clues were drawn from games set aside before any rubric development, and those games are excluded from both the training and test sets used in the validation below.

Jev Difficulty vs. Correctness

  • Now that we’ve labeled all the questions in our dataset, does the AI’s difficulty rating track actual player outcomes? In Figure 2 we plot the fraction of questions answered correctly vs. the 1-10 Jev difficulty. Indeed, we see a monotone relation. Clues rated 1 are answered correctly 98.8% of the time, while difficulty 9 clues are answered correctly only 67.1% of the time. The outer buckets are sparsely populated, while the inner buckets are more heavily populated. The overall correlation is r=-0.204.
Figure 2: Fraction of clues answered correctly by Jev difficulty (1–10). Bubble size shows clue count in each bucket. Linear trend shown in dashes.
  • Below are three different questions, to give us a taste of some of Jev’s labels. The first one is a low dollar value, but high AI difficulty question that indeed stumped everyone. The second is a high dollar value, low AI difficulty that was answered correctly. Finally, an easy AI difficulty that stumped everyone. Personally, I feel like these particular ratings seem totally accurate!
Low dollar value, high AI difficulty — Triple Stumper
SORORITY NOW!  ·  $200  ·  Jev difficulty: 9/10
In 1998 the eyes of this state univ. were upon Kappa Phi Gamma, founded as the USA's first women's South Asian Greek-lettered org.
Reveal answer
UT Austin
Nobody answered correctly.
High dollar value, low AI difficulty — Answered correctly
SCIENCE TEST  ·  $1,600  ·  Jev difficulty: 1/10
Newton's second law, F = ma, says that net force acting on a body is equal to its mass times its this, abbreviated a
Reveal answer
acceleration
1 contestant(s) answered correctly.
AI rated easy (difficulty ≤ 3) — still a Triple Stumper
MYTHS & LEGENDS  ·  $200  ·  Jev difficulty: 3/10
It's not a piece of fairy jewelry, it's a circle fairies like to dance inside
Reveal answer
a fairy ring
Nobody answered correctly despite a low AI difficulty rating.

Does Jev Add Predictive Value?

The real test. I split games into training and test sets (20% held out, selected before any model or rubric development). Two models:

  • Baseline: lookup P(any correct | row, round) from training games
  • Full model: lookup P(any correct | row, round, difficulty) from training games

Both models produce a probability estimate for each clue. I evaluate them using the Brier score: the average squared difference between the predicted probability and the actual outcome (0 or 1). Formally, \(BS = \frac{1}{N}\sum_{i=1}^{N}(p_i - o_i)^2\), where \(p_i\) is the predicted probability and \(o_i \in \{0,1\}\) is whether at least one contestant answered correctly. Lower is better — a perfect predictor scores 0, and a model that always guesses 50% scores 0.25.

I computed confidence intervals by bootstrapping by game (resampling whole games with replacement 500 times), which accounts for the fact that clues within a game aren’t independent.

Brier: baseline = 0.1119, +Jev = 0.1092
Improvement: 0.0027 (2.4%) [95% CI: 0.0024, 0.0030]
(a) Brier score (lower = better) for the baseline model and the model augmented with Jev difficulty. Improvement and bootstrap 95% CI printed above.
(b)
Figure 3

Adding Jev difficulty reduces Brier score from 0.1119 to 0.1092 — an absolute improvement of 0.0027 (2.4% relative), with a bootstrap 95% CI of [0.0024, 0.0030]. The confidence interval doesn’t include zero, so the improvement is statistically clear on this split. The effect is modest — board position provides a useful baseline — but Jev adds real information beyond what dollar value alone provides.

To put this in plain terms: before each clue, the model guesses the probability that at least one contestant answers correctly. The baseline model uses only board position — it predicts something like 96% for row-1 clues and 72% for row-5 clues, but treats every clue in the same cell identically. Adding the Jev difficulty lets the model distinguish a suspiciously hard row-2 clue from an easy one. The 2.4% improvement means the model’s guesses are slightly less wrong on average. Small, but real: the AI is picking up on something in the clue text that dollar value doesn’t capture.

Difficulty Correlations

Over Time

How has the difficulty changed over time? The data shows that it’s been getting easier since the 80’s, with an increased difficulty spike in recent years. There could be numerous explanations for this, however. I hypothesize that questions in the 80’s asking about 80’s figures read as niche historical knowledge today, thereby seeming harder in retrospect.

Figure 4: Mean Jev difficulty by air year, regular-season board clues only. Dashed line marks Ken Jennings beginning as guest host (January 2021). This reflects how today’s model rates archived clues, not a direct measurement of how hard contestants found them at the time.

Regular Jeopardy vs. Double Jeopardy

Are clues harder in the second round?

Figure 5: Mean Jev difficulty by row and round. Double Jeopardy clues are harder at every row position.

Double Jeopardy clues are harder at every row position — the gap is around 0.25–0.36 difficulty points. This makes sense: DJ clues cost twice as much and are intentionally set harder by the writers.

Daily Doubles vs. Normal Clues

Do the cluemakers put the daily doubles on harder clues than average?

Figure 6: Difficulty distribution for daily double clues vs. regular board clues, regular-season games only.
Mean difficulty — Daily Doubles: 5.28, Regular clues: 4.82
Regular clues reweighted to DD row/round distribution: 5.13
Adjusted DD-vs-regular gap: 0.15

Daily doubles average about 0.46 points harder than regular clues on the raw comparison, but most of that gap reflects where on the board Daily Doubles tend to appear — deeper rows and Double Jeopardy, which are harder regardless. After reweighting regular clues to match the Daily Double row/round distribution, the remaining gap is only about 0.14 points. There may be a small deliberate placement effect, but it’s much more modest than the raw numbers suggest. (Note: contestants wager on Daily Doubles before the clue is read, so the wager itself doesn’t reflect knowing the answer.)

Final Jeopardy vs. Normal Clues

Are the final jeopardy clues harder than an average clue?

Figure 7: Difficulty distribution for Final Jeopardy clues vs. regular board clues.
Mean difficulty — Final Jeopardy: 6.07, Regular clues: 4.82

Final Jeopardy clues are rated meaningfully harder on average (6.07 vs 4.82 for regular board clues), with the distribution shifted toward the upper end.

Game Type

Figure 8: Difficulty distribution by game type, shown as smoothed lines (Gaussian kernel, σ=0.6). Regular-season, TOC, Celebrity, College, and Teen editions.

Tournament and celebrity games skew toward the extremes compared to regular play — tournaments concentrate clues at mid-to-high difficulty, while celebrity games have more easy clues. The distributions share a broadly similar bell shape across game types, with the main differences at the tails.

Category Type

Categories cluster into recognizable types — for the clustering methodology, see the companion post. Here are the 20 most common category types, sorted by mean AI difficulty.

Figure 9: Mean Jev difficulty (±95% CI) for the 20 most common category types, sorted from easiest to hardest.

The results have some surprises. Wordplay & Vocabulary is the easiest category type at 4.32, while Movies (5.27), Television (5.29), and Notable People & Awards (5.30) are the hardest among the top 20 — Movies and Television receive higher mean AI difficulty ratings than the geography and history clusters shown here. The spread across category types is about one point (4.32 to 5.30), which is smaller than the spread across individual clues within any one category — topic alone isn’t destiny.

Try it yourself

Explore clues by difficulty, round, and category type. The sampler draws from 16,040 clues (rated 1–9) across the full difficulty range.