Skip to content
All posts

3 min readVAKYN

Your work is not a benchmark

Models are tuned to raise general scores. The only score that predicts how a model will do on your work is a score on your work.

Every few weeks a new model tops a leaderboard. The announcement has a table, the table has bold numbers, and the numbers are a point or two higher than last month's. Then a team tries the model on their own queue, and it is about as good as the one they had. Sometimes worse.

This is not bad luck. It is what happens when a number becomes the goal.

A score is a sample of someone else's work

A benchmark is a fixed set of questions someone wrote to stand in for "general ability". It is useful when it first appears: it lets you compare models on the same footing. But it samples a particular kind of task, written in a particular style, with particular traps.

Your work samples something else. A support team's tickets mix three languages and a lot of anger. A security team's alerts are mostly noise from a scanner nobody turned off. A staffing team's CVs list skills in tables, paragraphs and screenshots of tables. None of that is in the benchmark, and none of the benchmark is in your queue.

So the question "which model is best?" has no general answer. The useful question is narrower: best at what, on whose cases?

Goodhart was right about models too

The economist Charles Goodhart noticed in the 1970s that a measure used to steer a system stops measuring well. Marilyn Strathern later put it in the form people quote: when a measure becomes a target, it ceases to be a good measure.

Public benchmarks are targets now. Models are trained, tuned and selected to raise them. Some of that effort makes models better in general. Some of it makes them better at the benchmark: its format, its topics, its habits, and sometimes its exact questions, which leak into training data. From the outside you cannot tell the two apart. The score goes up either way.

The result is a field where the headline numbers rise faster than the usefulness, and where the model that wins the table is not reliably the model that wins on your desk.

Measure on the work instead

The fix is not a better leaderboard. It is to stop asking a general score to answer a specific question.

For the decisions a company makes thousands of times a day, a good test looks like this:

  1. Take real cases from the decision you want to automate. Not synthetic examples: the tickets, alerts or documents your team actually handles, with the answers your best people gave.
  2. Score the model on those cases, and on its confidence. Being right matters. Knowing when it might be wrong matters as much, because that decides how many cases you can automate safely and how many go to a person.
  3. Look at where it fails, case by case. Failures cluster. A model that is fine on refunds may be lost on partial refunds in another currency.
  4. Fix exactly there, then measure again. Train on the kind of case it got wrong, check that it now gets them right, and check that nothing else got worse.

This is less exciting than a leaderboard. It is also the only number that predicts what will happen on Monday morning when the model meets your queue.

How we work

VAKYN does not chase a public leaderboard, and we do not tune for one. We train decision models across nearly 1,000 occupations in 22 domains, and then we prove them on your cases. When we show a customer results, the results are on their work: where VAKYN is right, how sure it was, and the list of cases it still gets wrong.

Then we run the VAKYN loop on that list. Find where it fails, train exactly there, check again, and repeat when your work changes. The model gets better at your work, not at someone else's test.

If you want to see what that looks like on your hardest decision, talk to us.