LLM Demand Index · live benchmark
counting…
loading the count…
best fit in the picks% = share of all tasks
Every task typed here is ranked by Jev over the whole catalog, weighing fit for the task against price. The index counts, per model, how often it came out the best fit and how often it made the picks. No fixed exam and no votes: the tasks are what people actually need done, and every new one updates the count.
Why a running count holds up: it is a poll. A few thousand voters predict an election because a sample carries the population's shares; a few hundred real tasks already carry which models win most of the work. The margin of error shrinks as the count grows.
This is how the field measures demand: LMArena ranks models by summing real prompts and votes, OpenAI and NBER classified 1.5 million real conversations, Anthropic's Economic Index maps a million conversations to work tasks, and OpenRouter with a16z summed 100 trillion tokens of real requests. All of them aggregate real usage rather than score a fixed test set.
What it is not: MMLU or SWE-bench measure skill on a fixed exam, arenas measure chat preference. This index measures demand: which model a price-aware judge picks for the mix of tasks people bring. Counting rules: best fit is Jev's top pick; in the picks means within 15% of the best; a visitor re-typing the same task counts once; only Jev-ranked tasks count.