Applied AI · 4 min read
Choose the model after you define the job
A model comparison becomes useful when it tests the task, the failure you cannot accept, and the conditions in which the system will run.
A new model arrives with an impressive score. The obvious next move is to try it. The less exciting move is to write down what would make it useful for your work before seeing how it performs.
I have been evaluating local models for coding and review workflows, and the second move has become much more valuable to me. A general ranking cannot decide whether a particular model is a good reviewer under my time limits, with my tools, on the kinds of changes I need it to examine.
The job needs a definition before it gets a winner.
Name the failure that matters
For a code reviewer, finding plausible things to comment on is only part of the task. It also needs to catch the defect that makes a change unsafe to rely on. A polished review that misses that defect can leave the next person more confident than the evidence warrants.
In a small local comparison in July, I tested two model variants, each with two template configurations, across three code changes. One candidate returned reviews faster and found useful issues, but missed a known issue on a change where the other caught it. The slower candidate also ran into the review tool’s deadlines. There was no clean winner for the job.
That was a limited experiment, not a ranking of those models for every task. Its value was identifying which tradeoff I actually faced: useful findings, missed defects, and completion under the conditions of the workflow.
Check the test before trusting the score
The same comparison exposed a problem in the machinery running the test. One review was reported as successful with zero comments even though its subtasks had timed out. The process had ended; the review had not been completed. Those two facts had been collapsed into one status.
An empty review can mean the model found no issues. It can also mean the model never finished examining the change. Until the test distinguishes those outcomes, its summary gives you a misleading comparison.
Inspect what the system actually received, attempted, and returned. Record failed subtasks and partial completion. If one candidate keeps timing out, decide whether you are measuring its reasoning, the serving setup, or the practical performance of the whole workflow. Each can be a useful question. They lead to different conclusions.
Write the decision rule in advance
For a later comparison, I recorded the question, configurations, predictions, and disqualifying behavior before collecting the results. The question was narrow: could a challenger replace the reviewer already serving this machine?
That kept a broad score from becoming the whole decision. A model could perform well overall and still fail a requirement of the reviewer role. The existing model had to pass the same scrutiny. Familiarity does not earn an exemption.
For your own workflow, choose representative tasks and examples of failures you need to detect. Write down what counts as correct, what counts as incomplete, and which result would prevent you from using the model in that role. Then run the comparison without rewriting the criteria around the outcome you hoped for.
Compare the system you will actually use
The model name does not fully describe the experiment. Inputs, instructions, software versions, resource limits, and simultaneous requests can all change what the workflow delivers. Record those conditions so you can tell what changed when the result changes.
There is a cost to this care. Small teams cannot build a research program around every tool choice, and they should not have to. A small, repeatable set of relevant tasks can answer a bounded question. Be equally precise about what it leaves unanswered.
If you use an example to tune the setup repeatedly, add fresh examples before treating the result as evidence that the system handles unfamiliar work. Keep the difficult cases around for regression checks, but avoid teaching to the entire test.
I want an evaluation to end with a usable decision: keep the current setup, change a constraint, assign a narrower role, or keep testing. ‘No clean winner’ is useful when it explains the next move. A leaderboard position, on its own, rarely does.