A model call is easy to price. A completed piece of work is harder. Retrieval, routing, validation, retries and human review can all change the economics of the system.
That matters when comparing an architecture. A smaller model may cost less per call and still require more attempts or more correction. A more capable route may earn its cost on difficult tasks while adding little on routine work.
Begin with a definition of useful.
An evaluation needs to represent the actual task. What counts as a correct result? What must be cited? When should the system abstain? Which actions need approval? A price is only meaningful alongside those conditions.
The unit of value is the completed task.
Test the combination.
There are several places to change the result: better source preparation, a different retrieval method, deterministic checks, a specialist component, a more selective model route or a clearer human decision point.
An experiment should help choose between those alternatives. Measure the whole route, including the cases that fail. Record the operating assumptions so a comparison can be repeated.
The frontier moves with the work.
A configuration that earns its place on one workload may be a poor choice on another. Cost, response time, quality and control need to be considered together. The architecture should reflect that trade-off.
Our cost–quality instrument illustrates the selection process with synthetic values. It is a way to examine the idea, not a provider benchmark.