“What does the AI cost?” is only useful once we know what it is supposed to accomplish.

For a flooring company, the unit might be an inquiry with the right next step. For a contractor, a job request ready for someone to act on. For an app creator, enough information to make the next useful improvement.

We want a short scorecard that answers three questions: Did it get the task right? What did it cost? How much work did it leave for a person?

Change one part, compare the same work

Suppose a workflow reads an inquiry, checks what is missing, and prepares a review packet. We can change the model used for the missing-information check while keeping the inquiry, source material, and expected answers fixed.

That lets us compare a cheaper model with a more capable one without quietly changing the task. If the cheaper option performs well enough on representative cases, it has earned its place. If it misses important facts or creates more correction work, the API saving may not be worth it.

The Built Correct AI scorecard makes that comparison the starting point. The detailed lab keeps the underlying controls and run records available for inspection.

Keep the scorecard small

Measure Business question
Correct outcomes Did the workflow choose the expected next step?
Missed details Did it overlook information the team needed?
Unnecessary questions Did it create avoidable follow-up work?
API cost What did the recorded model usage cost at the stated rates?
Human review How much time did someone spend checking or correcting it?
Completion time How long did the full attempt take?

The workflow determines which of these matter most. A search task may also need a retrieval measure. A quantity calculation may only need ordinary correctness checks. A completeness decision needs missing-detail and unnecessary-question counts more than a fashionable retrieval metric.

Keep denominators visible. “Three out of four examples” means something different from “75% accurate” presented without context. A failed attempt belongs in the record. A missing cost is unknown, not zero.

Give expected answers their own place

Reference documents tell the system what it can know. Evaluation labels tell us what it should have done on a particular case. Mixing them makes a test easier to pass without making the system more useful.

Our public examples have engineering-authored expected answers. They are useful fixtures for checking behavior; they have not been independently validated as a representative customer benchmark.

For a business trial, a domain expert should label real examples or carefully scrubbed equivalents. Include missing facts, contradictions, unfamiliar wording, and cases that should reach a person. Keep a separate group of cases out of prompt and threshold tuning. Put related examples in the same group so near-duplicates do not leak into the final evaluation.

A reviewer correction is feedback first. Confirm it before freezing it into a new benchmark case. Keep the original prediction and the correction separate so the record still shows what the system actually did.

Human time can outweigh token cost

At an assumed $60 per hour, one minute of review costs $1. That is arithmetic for illustration, not a measured Built Correct saving or a recommended labor rate.

The useful comparison is therefore wider than API cost per call. Track the API estimate and actual recorded review minutes separately. If infrastructure, retrieval, or other processing costs have not been allocated, say so. Do not label a partial total as the complete cost of the business outcome.

Likewise, “accepted by a reviewer” and “correct against a benchmark” are different observations. Both can be useful; one should not silently stand in for the other.

A review setting changes the policy

A stricter review threshold sends more uncertain cases to a person. It does not make the underlying model smarter. Show the errors alongside the work passed to reviewers. Accepting no cases is not evidence of perfect automation.

Applying a different threshold to saved predictions is a policy calculation. Calling a different model is new inference, with new usage. Keep those two operations distinct in the record. For Jev specifically, TypeSafe describes confidence as a statistic of the returned probability distribution; test how it relates to correctness in the actual task.

We want the public experience to make the tradeoff easy to grasp, with the engineering detail one click away. Compare the examples, or read the flooring workflow.

Built with perspective.
The Built Correct team.

More from the journal