What will it cost to screen your question bank with an LLM?

You own a bank of closed items. You want to know which of them a student with a chatbot can now answer without knowing anything. This page prices the obvious plan — run every item past a cheap model a few times, then past an expensive one — and tells you how wrong the answer will still be.

Prices are dated. Every list price below is what the provider published in August 2026, and every one of them is an editable field. Check them against your own invoice before you plan a purchase. Two of the models here have no verified price in our records; those fields start empty and are marked.

Your bank

Count the rendered question with its options and markup, not just the stem.
This sets the guessing floor, which turns out to matter more than the model does.

The two models

What one call actually costs

Almost everyone who plans a screening run of this kind budgets it wrong, and the error is never small. Here is the arithmetic, in the order it bites.

Step 1 — the input side, which behaves

Models bill by token, not by character. Across roughly 100,000 of our own logged calls, HTML question bodies in Ukrainian ran about 2.1 characters per token on Gemini and Claude models, and about 2.7 on the GPT family, which uses a different tokenizer. English runs a little longer per token; Ukrainian is the tighter case because Cyrillic costs more tokens per character.

The per-call input figures used below are measured averages from our logs at roughly 640 characters per question, and they already contain the instruction and the wrapper around the question. We scale them by your length.

Step 2 — the output side, where people go wrong

You ask the model to answer a four-option item. It replies B. One character. If you budget for that, you will be out by a factor of thirty.

First, "B" is never one token: wrapped in whatever JSON or scaffolding your harness asks for, the visible reply lands around a dozen tokens. That part is harmless. The part that is not harmless is that a thinking model also bills reasoning tokens it never returns to you. They are not in the response body. They are on the invoice.

On Gemini 3.7 Flash we measured 441 hidden reasoning tokens against 18 visible ones, averaged over thousands of calls answering exactly this kind of item. 96% of the output you pay for is text nobody ever sees. A non-thinking model like Gemini 3.1 Flash-Lite reports its hidden count as exactly zero, so its bill is what it looks like.

18 visible tokens 441 hidden reasoning tokens Gemini 3.7 Flash, per call, measured
Step 3 — a second trap, and a different mechanism

Some thinking models also multiply the input. Our prompt for one item is about 240 tokens long. Running it against GPT-5.6-luna-pro, the provider billed us 2,966 input tokens for it — 12.4 times the prompt we sent — because the model re-reads its own accumulated context on every internal reasoning step, and every re-read is billed as fresh input.

The plain and the pro variant of that model carry the same published list price ($0.20 in / $1.20 out per million). Same prompt, same price sheet:

So: you cannot derive a thinking model's cost from the length of your prompt. Neither multiplier is visible on a price page, neither is visible in the response, and the two of them compound. Both figures on this page are measured from our own logged usage records, not estimated.

What your two calls cost

When do we say a model has solved an item?

Not when it passes once. On a four-option item a model that knows nothing at all passes one run in four, and passes at least one of eight runs 90% of the time. Scoring a single pass as knowledge would condemn most of your bank on the strength of luck.

So after every run we ask one question of the item's record so far: could a pure guesser plausibly have produced this? If a guesser would produce a record this good less than one time in twenty, we call the item solved and stop paying for it. Otherwise we keep going. That threshold — one in twenty — is fixed here; it is the conventional line, and moving it is not the lever worth pulling.

The consequence is that the format of your items, not the model you buy, sets the minimum number of runs. A five-option item needs two consecutive passes before the record can beat a guesser. A four-option item needs three. True/false needs five. Short text entry needs one, because a guesser cannot type the right answer by accident.

The staircase for

How many of the runs so far must have passed before the record clears the one-in-twenty line — and which runs, therefore, cannot change any decision at all.

After runPasses neededA guesser's odds of thatCan this run certify anyone new?

The grid: how many runs of each to buy

Rows are runs of the cheap model, columns are runs of the frontier model. An item is dropped from the working stock the moment its record beats the guessing floor; whatever survives the cheap stage goes to the frontier stage; whatever survives both is what you keep and go on using. Each cell shows the two numbers you have to trade off: the share of your remaining working stock that a model can in fact solve — items you kept but should not have — and what the screening costs per thousand items.

Runs of the frontier model →
Each cell: top = share of the stock you keep that a model can actually solve; bottom = dollars per 1,000 items screened. Rows are runs of the cheap model.

What each extra run buys

Both charts read off the same grid above; switch to the table view for the numbers.

What we would tell you to buy


Where the numbers come from, and what they assume