Sleak
People Development

Philipp HeidekerSeptember 15, 202615 min read

What Is a Scorecard? Criteria, Levels, Limits

A conversation scorecard is an evaluation rubric set before the call: named criteria, described levels, and evidence from the transcript for every rating.

What Is a Scorecard? Criteria, Levels, Limits

TL;DR. A scorecard in training is an evaluation rubric that defines in advance what a good conversation looks like: named criteria, clear levels for each criterion, and evidence from the transcript. This makes feedback repeatable, regardless of which manager listened to the conversation. In English, the term has several meanings. Here, it does not refer to a training scorecard for programme metrics or an interview scorecard for ranking candidates. It also has a clear limit: a scorecard evaluates what was said, not whether the deal closed. This post explains its structure, its levels, who owns it, and the most common mistakes.

Key Takeaways

  • A scorecard has three parts: named criteria, a description of what each level looks like, and evidence from the transcript behind every rating.
  • In English the term is occupied twice over. A training scorecard usually means a dashboard of programme metrics, and an interview scorecard ranks candidates. Neither evaluates a single conversation against a standard.
  • Behaviour is measured more often than the industry admits, but it is mostly asked about rather than observed: 54 percent of organisations evaluate at Kirkpatrick level 3 (ATD, 2019, n = 779), while Will Thalheimer estimates that well over 80 percent of transfer studies capture learner perception instead of behaviour (Thalheimer, 2020).
  • More than 2,200 scorecards on the Sleak platform have produced over 1.9 million individual criterion ratings. Rating happens per criterion, not as a single overall grade.
  • The practice behind the rubric is measured: at Schwäbisch Hall, the group that trained with two practice conversations scored 72.5 percent action competence against 61.4 percent for the course-only group, an improvement of 11.1 percentage points.

Ask a sales organisation what a good discovery call sounds like and every manager will give you a different answer. Ask to see that standard in writing and there is often nothing to show. Organisations evaluate conversations all the time, but few have documented what they are evaluating them against. In competitive diving, the standard is published before anyone steps onto the board, and every deduction relates to a specific part of the dive. In professional development, the standard often exists only in the listener's head.

A scorecard puts that standard on paper. This post first clarifies what the term means, because in English it is used for several different instruments, only one of which evaluates conversations. It then looks at what a scorecard contains, why levels need descriptions rather than just numbers, who should write it, how it works outside sales, and where its limits lie. Those limits matter when deciding whether the instrument is worth the effort.

What is a conversation scorecard?

A scorecard in training is an evaluation rubric that defines in advance what a good conversation looks like: named criteria, clear levels for each criterion, and evidence from the transcript. It makes feedback repeatable instead of dependent on the manager who happened to listen. All three elements are essential.

A scorecard has three components. First, the criteria: the specific aspects that matter in this type of conversation, such as qualifying the need, handling disagreement, or closing with a commitment. Second, the levels: a description of what full, partial, or no achievement looks like for each criterion. Third, the evidence requirement: every rating refers to a specific moment in the conversation. Without evidence, a rating is simply an opinion with a number attached. Without described levels, the standard changes with every evaluator.

At Sleak this artefact is called a Scorecard, or Standard of Excellence. It is the yardstick a practice conversation is scored against: after a voice-based AI role play, the AI Coach, meaning the always-available coaching instance each person talks to, returns a rating per criterion with a quote from the transcript attached, rather than general praise. Scorecards are versioned, so a draft and published snapshots exist side by side, and 18 ready-made templates cover the common conversation types.

Training scorecard, interview scorecard, balanced scorecard: which one is this?

In English, four different instruments share the word scorecard. This article focuses on the one used to evaluate conversations. The distinction matters. Search results tend to show programme dashboards and hiring rubrics, so buyers may approach a conversation scorecard expecting something quite different.

InstrumentWhat it evaluatesWhen it is usedWhat it produces
Balanced scorecardthe organisationquarterlystrategic KPIs across four perspectives
Training scorecarda training programmeper programme or per quartercompletion, satisfaction and cost metrics
Interview scorecarda person as a candidateonce per hiring processa selection decision
Conversation scorecardone conversation performanceafter every practice or real conversationa rating per criterion, with evidence

The key difference is what gets evaluated. A training scorecard shows how a programme is performing. An interview scorecard helps decide whether to hire someone. A conversation scorecard assesses whether a specific person demonstrated the behaviours the organisation considers important in a specific conversation. The same standard can then be used again the following week. Its closest common relative is the customer-service QA scorecard, which also evaluates individual conversations. However, it typically names categories without defining the levels or specifying what counts as evidence.

What is a scorecard made of?

A usable scorecard maps its criteria onto the phases of the conversation and describes, for each criterion, what achievement looks like. Without the phase structure you get a list of virtues, and virtues cannot be rated. Empathetic is not a criterion. Acknowledged the stated objection and tested it with a question before answering it: that is one.

The number of criteria follows the same logic. What belongs in a scorecard is what a manager can actually distinguish while listening, not what sounds comprehensive in a planning document. Twenty criteria per phase produce ratings nobody reads, and they dilute the judgement, because the three criteria that decide the conversation disappear into the list. Working scorecards run few criteria per phase, each with a precise description.

The evidence requirement is what separates a rubric from a questionnaire. Every rating names the criterion, the value, and the point in the transcript the judgement rests on. That makes it checkable, and it lets the person being rated disagree with something specific. This is not a presentation detail. It is the precondition for a rating turning into a conversation rather than a notification.

Why described levels instead of a 1 to 10 scale?

A number without a description does not remove subjectivity. A score of 7 out of 10 establishes nothing that will necessarily hold next week. One manager's 7 may be another's 5, and neither can explain the rating without defining the scale afterwards.

Anchored levels solve this by putting the description before the number. For each criterion it is settled in advance what full achievement looks like and how you can tell it is missing. The number afterwards is only shorthand for that description. This is also why a few well-described levels work better than ten fine gradations: the finer the scale, the larger the share of it that nobody can describe in words.

On the Sleak platform, ratings run on a 0 to 100 scale and conversations can be filtered by score range afterwards. Across every criterion rating produced so far, more than 1.9 million of them, the three most frequent single values are 0, 100 and 50, together just under 38 percent. The anchors get hit in practice without the range between them going unused. How that rating logic turns into a working coaching rhythm is a separate question, and the related posts at the end cover it.

Who writes the scorecard?

The manager accountable for the outcome writes the scorecard, not L&D and not the vendor. This is not an organisational preference. A rubric decides what counts as good in this team, and that decision belongs to whoever answers for the result.

In the Sleak model the frame for this is an Initiative, a development goal set by a leader covering both KNOW and DO. The manager does not need instructional design training for it. They need to be able to name what makes a conversation good, and that is exactly where most rollouts stall: not on the technology, but on whether an organisation can describe its five most important conversations at all. In leadership development this writing work is the real cost of the programme, and it is paid once.

L&D and the vendor still have a role. Templates supply the structure, meaning the phases of a conversation type and the criteria that usually belong in it. What goes into the level descriptions has to come from the organisation. A bought scorecard nobody internally has touched evaluates somebody else's expectations.

What does a scorecard look like outside sales?

The structure stays the same and the criteria change with the conversation type. A scorecard is not a sales instrument. It fits any recurring conversation whose quality somebody has to judge.

In procurement, a scorecard rates whether the walk-away number was set before the first price discussion, whether concessions were traded rather than given, and whether the outcome was summarised in writing. What that looks like in negotiation training for sourcing teams then comes down to which negotiation line the organisation has decided on. In leadership the criteria are different: whether observation was separated from evaluation, whether the other person got to speak, whether the conversation ended with a commitment somebody can check. In service it is de-escalation, the quality of the promises made, and how policy that contradicts the customer's wish gets handled.

What is striking is how similar the criteria look structurally across departments whose content has nothing in common. They always describe an observable action inside the conversation, never a trait of the person. That is the part of the method that actually transfers, and it is why the same structure holds in sales, procurement, leadership and service.

Is behaviour actually measured, or just asked about?

Behaviour is measured more often than the industry's reputation suggests, and it is mostly asked about rather than observed. That is the more precise version of a popular complaint, and it is the more uncomfortable one.

Kirkpatrick levelShare of organisations
1 Reactionaround 80 percent
2 Learningaround 80 percent
3 Behaviour54 percent
4 Results38 percent
5 ROI16 percent

Those figures come from an ATD survey of 779 talent development professionals (ATD, 2019). So more than half of organisations do evaluate at level 3. The catch is in the method: level 3 data is typically collected through follow-up surveys, and the established approach is to ask 25 to 30 randomly selected participants, at least 30 days after the programme, to estimate how much of it they apply (Training Industry, 2022). Will Thalheimer estimates that well over 80 percent of transfer studies therefore measure not the behaviour but the learner's perception of it (Thalheimer, 2020).

Somebody assessing their own behaviour change is reporting satisfaction under a different name. This is exactly where the scorecard sits: a standard defined in advance, plus a repeatable rating of the same action by somebody other than the person performing it. The effect of that is measurable. At Schwäbisch Hall, the group that trained with two practice conversations reached 72.5 percent action competence against 61.4 percent for the course-only group, and the gap between what participants knew and what they could apply fell from 32.9 points to 13.9. What the study measured was practice against a standard, not the standard on its own.

Isn't this just a checklist with a better name?

A scorecard differs from a checklist in two ways: it describes the levels and requires evidence. The objection is understandable, because many evaluation forms really are checklists with a more impressive name. A checklist asks whether something happened. A rubric describes how well it happened and asks for proof.

One criterion makes the difference concrete. Checklist: objection handled, yes or no. Rubric: full achievement when the objection was acknowledged, tested with a question, and only then answered. Partial when it was answered but not tested. None when it was talked over. Same moment in the conversation, three distinguishable outcomes, each of them anchored to a line in the transcript. Vinzenz D., a small-business user reviewing Sleak on G2, named the practical effect: "I especially appreciate the depth of the AI-generated feedback. I can now train all scenarios as I need them with less time effort and am no longer dependent on managers." That is a review, not a controlled study.

For many purposes, a checklist is still the right tool. Mandatory disclosures, regulatory statements, and anything that was either said or not said only require a binary answer. Describing levels would add unnecessary work. An experienced manager who has listened to the same person for two years will also notice things that no rubric captures. A scorecard does not replace that judgement. It ensures that the manager's impression is not the only basis for feedback and that each review does not start from scratch.

What does a scorecard not tell you?

A scorecard evaluates conversational behaviour against a standard, not the commercial outcome of the conversation. A call can be run cleanly and still not close, and a badly run call can end in an order. Reading the scorecard as a revenue forecast is a category error.

Four mistakes recur. First, criteria that describe traits instead of actions, so empathetic rather than asked a follow-up question. Second, too many criteria, which buries the ones that decide the conversation. Third, levels made of numbers alone, which puts the standard back in people's heads. Fourth, and most consequential, a scorecard that rewards the wrong thing. A rubric reliably trains exactly what it rates, and that holds just as firmly when the criteria have drifted away from what the market actually responds to.

There is also a limit to machine evaluation itself. What gets evaluated is the text of the conversation, not tone, pace, or presence in the room. That boundary defines where a human reviewer still has to look. The honest summary: a scorecard ranks more reliably than it grades, so it is good at showing which conversations differed and why, and weaker at fixing an absolute quality bar that holds in every context.

FAQ

What is the difference between a scorecard and a feedback form?

A feedback form collects impressions after the conversation, a scorecard fixes the standard before it. The form asks how it went. The scorecard checks whether defined criteria were met and requires evidence for each rating. That is why two ratings from the same scorecard can be compared and two completed feedback forms cannot.

How many criteria should a scorecard have?

As many as can actually be distinguished while the conversation runs, mapped onto its phases. A few precisely described criteria per phase produce ratings that get read and accepted. A long list produces completeness on paper and buries the three criteria that decide the outcome.

Who defines the criteria?

The manager accountable for the result. Templates can supply the structure, meaning the phases and the criteria that usually apply, but the level descriptions have to come from inside the organisation. A scorecard nobody internally has edited is evaluating somebody else's expectations.

Does every conversation need its own scorecard?

No, every conversation type does. A first meeting, a price negotiation and a difficult feedback conversation have different phases and therefore different criteria. Within one conversation type the scorecard stays fixed, because that is what makes several attempts comparable.

How can you check whether the rating is fair?

Through the procedure rather than the feeling. Two tests go a long way: rate the same recording twice against the same scorecard and compare the results, then check each rating to see whether the quoted evidence actually supports the level. Where no evidence is given, the rating is not checkable, regardless of whether a human or a machine produced it.

Can a scorecard be tied to a methodology like SPIN, MEDDIC, SBI or GROW?

Yes, and it is the most common use. A methodology supplies the phases and often the vocabulary, and the scorecard translates them into observable actions with levels. Without that translation a methodology stays a piece of training content everybody can name and nobody can be shown to apply.


Related articles

More on practice, feedback and getting better at the conversations the job actually turns on.

People Development

Practice Without an Audience: Role Plays

Role plays in training rarely fail on method. They fail because safety and repetition compete for the same scarce resource in a training room.

Philipp Heideker12 min read
Read article
People Development

Learning Transfer: Why Training Rarely Sticks

Learning transfer fails because nobody schedules the practice phase between the classroom and the real conversation. Here is what the evidence supports.

Philipp Heideker12 min read
Read article
People Development

Deliberate Practice at Work: Why One Training Session Rarely Changes Behavior

Deliberate practice turns knowledge into behavior through application, feedback and another attempt. The hard part is creating enough chances to practise.

Philipp Heideker13 min read
Read article