Why semantic scoring is the control surface for real-world AI systems

Why semantic scoring is the control surface for real-world AI systems

By 2028, most organizations will have faced a simple divide:

  • The AI-aligned: Those that have encoded the quality standards of their knowledge work into AI systems that can be measured, improved, audited, and safely changed.
  • The Red Tape Machines: Those drowning in complexity and confused about why AI initiatives keep under-delivering. In comparison to their peers.

It is obvious that scoring is useful. The non-obvious part is the mental model of score-driven AI adoption.

AI won't always execute. But it should always measure what matters for your organization, your customers and your product.

Anyone can ask ChatGPT or Claude for a score - but without proper harness tooling, those scores become sloppy, inconsistent, or non-actionable.

"Average chatbot answer relevance went from 0.70 to 0.80" is not helpful in itself!

"The 7-day average Upsales Signals Judge score across 3 chatbots dropped from 0.95 to 0.70 after the system prompt changed" - is a whole other matter.

The world is a continuum, not a pass/fail switch

The questions we face once AI enters real operations are nuanced:

  • Is the system improving?
  • Is the improvement in one dimension degrading another?
  • Did the new model preserve all quality dimensions of the agent responses?

Some data scientists urge AI practitioners to "avoid Likert scores" and stick to a string of binary pass/fail judgements, to avoid AI slop and vanity metrics. Sometimes this works - but often it misses key information, and ultimately it should still be used to make a score, to track the total rate of change.

That 3% you rounded out actually matters. Often, that's the most important thing.

A 0.97 faithfulness score and a 1.00 are not the same product. Consider the effect of mismanaging 3% of refunds or 3% of medication recommendations!

Three things change once scoring is continuous

A simple yes/no result tells you whether something passed. A continuous score tells you how well it performed, where it is improving, and where risk is accumulating.

  • Signal: Improvement becomes visible while it is still small.
  • De-risking with proxies: Risk shows up before actual user incidents do.
  • Self-improvement: optimization becomes straight-forward.

The last item is the key to improvement - both for an AI customer support agent and the human manager overseeing a sales team.

What we are launching today

That is also why today we are launching our new brand look, symbolized by our signature progress bars.

Long-term AI evaluation needs versioned AI evaluators based on:

  • expert expectations
  • complete context
  • policy constraints
  • process KPIs
  • known failure modes

Each such evaluator needs:

  • a rubric
  • provenance
  • calibration examples
  • calibration methodology consistent over time
  • known limits
  • per-dimension thresholds

It runs repeatedly across releases, datasets, and production traffic. It can be compared across model changes. It can explain what changed, where, and why it matters. It tracks both regression and accumulation of real business value.

Start thinking of taking your high-stakes AI systems from that 97% to 99.999%.

How to get started?

Related reading