You own a bank of closed items. You want to know which of them a student with a chatbot can now answer without knowing anything. This page prices the obvious plan — run every item past a cheap model a few times, then past an expensive one — and tells you how wrong the answer will still be.
Almost everyone who plans a screening run of this kind budgets it wrong, and the error is never small. Here is the arithmetic, in the order it bites.
Models bill by token, not by character. Across roughly 100,000 of our own logged calls, HTML question bodies in Ukrainian ran about 2.1 characters per token on Gemini and Claude models, and about 2.7 on the GPT family, which uses a different tokenizer. English runs a little longer per token; Ukrainian is the tighter case because Cyrillic costs more tokens per character.
The per-call input figures used below are measured averages from our logs at roughly 640 characters per question, and they already contain the instruction and the wrapper around the question. We scale them by your length.
You ask the model to answer a four-option item. It replies B. One character. If you budget for that, you will be out by a factor of thirty.
First, "B" is never one token: wrapped in whatever JSON or scaffolding your harness asks for, the visible reply lands around a dozen tokens. That part is harmless. The part that is not harmless is that a thinking model also bills reasoning tokens it never returns to you. They are not in the response body. They are on the invoice.
On Gemini 3.7 Flash we measured 441 hidden reasoning tokens against 18 visible ones, averaged over thousands of calls answering exactly this kind of item. 96% of the output you pay for is text nobody ever sees. A non-thinking model like Gemini 3.1 Flash-Lite reports its hidden count as exactly zero, so its bill is what it looks like.
Some thinking models also multiply the input. Our prompt for one item is about 240 tokens long. Running it against GPT-5.6-luna-pro, the provider billed us 2,966 input tokens for it — 12.4 times the prompt we sent — because the model re-reads its own accumulated context on every internal reasoning step, and every re-read is billed as fresh input.
The plain and the pro variant of that model carry the same published list price ($0.20 in / $1.20 out per million). Same prompt, same price sheet:
So: you cannot derive a thinking model's cost from the length of your prompt. Neither multiplier is visible on a price page, neither is visible in the response, and the two of them compound. Both figures on this page are measured from our own logged usage records, not estimated.
Not when it passes once. On a four-option item a model that knows nothing at all passes one run in four, and passes at least one of eight runs 90% of the time. Scoring a single pass as knowledge would condemn most of your bank on the strength of luck.
So after every run we ask one question of the item's record so far: could a pure guesser plausibly have produced this? If a guesser would produce a record this good less than one time in twenty, we call the item solved and stop paying for it. Otherwise we keep going. That threshold — one in twenty — is fixed here; it is the conventional line, and moving it is not the lever worth pulling.
The consequence is that the format of your items, not the model you buy, sets the minimum number of runs. A five-option item needs two consecutive passes before the record can beat a guesser. A four-option item needs three. True/false needs five. Short text entry needs one, because a guesser cannot type the right answer by accident.
How many of the runs so far must have passed before the record clears the one-in-twenty line — and which runs, therefore, cannot change any decision at all.
| After run | Passes needed | A guesser's odds of that | Can this run certify anyone new? |
|---|
Rows are runs of the cheap model, columns are runs of the frontier model. An item is dropped from the working stock the moment its record beats the guessing floor; whatever survives the cheap stage goes to the frontier stage; whatever survives both is what you keep and go on using. Each cell shows the two numbers you have to trade off: the share of your remaining working stock that a model can in fact solve — items you kept but should not have — and what the screening costs per thousand items.