arrow_backBack to the overview
Calibration & coaching

A worker that measurably gets better.

Score everything. Classify the failures. Prove the fix.

Most AI deployments improve by anecdote: someone hears a bad call, someone edits a prompt, nobody can say whether it helped. The Hub treats improvement as an engineering discipline. Every conversation is scored, low scores become classified defects with an owner, revisions are proposed with the evidence attached, a person approves them, and nothing reaches production until it passes regression against a saved suite.

Every callscored automatically, not sampled
Evidence boundfindings quoted from the transcript
Human approvalno prompt changes itself
Scoring

The outcome carries the score.

The rubric is derived from the outcome contract, which means any worker with a contract can be scored without bespoke work.

ComponentContribution
Primary outcomeCarries the score. Each rung is worth materially less than the one above it, from full marks for the top rung down to nothing for no outcome at all
Outcome statusA tentative outcome earns partial credit, rising when a concrete confirmation path was agreed. Honesty is rewarded, and an inflated claim of confirmed is a defect rather than a win
Partial upgradeA recorded, transcript supported path to a better outcome adds credit without allowing the higher rung to be claimed
Standing objectivesBonuses, capped in aggregate. Judged on whether the worker took what the conversation naturally offered, not on whether it ground through a checklist
Behavioural complianceCascade violations, restating, stacked questions, manufactured urgency, premature endings, missed disclosure and unresolved identity are deductions regardless of outcome

What scoring is for

A score exists to direct attention. It tells you which conversations to look at, which prompt layer to examine, which vertical is underperforming, and which skill someone should practise. That is the whole of its purpose, and the narrowness is deliberate.

What it is never used for

  • Not a reward signal. We do not optimise a model against its own rubric. Automated optimisation against a proxy is how a system learns to score well rather than to work well
  • Not training data. Scored conversations improve prompts, knowledge and skills, all artefacts you own and can inspect. They do not become weights
  • Not a performance measure. Attaching a score to someone's pay converts an honest diagnostic into an incentive to game the record
Failures

A bad call becomes a defect with an owner.

Alongside the score, the evaluator returns typed findings, each anchored to a specific quote or turn reference. A number tells you a call went badly. A quoted finding tells you which line of which prompt layer caused it.

ClassMeaningWhere it goes
BehaviouralThe worker knew what to do and did not do itPrompt revision queue
KnowledgeThe worker lacked information it neededCuration task in the knowledge base
AssetThe worker needed collateral that did not existThe request board
CapabilityThe worker needed a tool or channel it does not haveRoadmap
ConfigurationWrong engine, thinking level, voice or scenario bindingControl plane
ComplianceA guardrail was approached or crossedImmediate escalation, never queued
query_stats
One bad call is noise. The same finding across forty calls in one vertical is a defect with a business case. Recurring signatures are what drive prioritisation, which is why classification matters more than volume.

The loop

Conversation Score Classify Coach proposes 3 to 6 surgical edits A person approves Regression suite Promote

The coach never rewrites a prompt wholesale, and that restraint is deliberate: wholesale rewrites destroy the accumulated calibration of a mature prompt and make regression impossible to attribute. Every version carries its author, its motivating findings, its regression result and its approver, so a prompt in production can always be traced back to the conversations that shaped it.

Simulation

Harden a worker before it speaks to anyone.

The behaviours that matter most appear rarely and cost the most when they do: an abrupt ending while a question is still open, a settled topic re-asked, a referral offered and not captured. Catching those needs volume that no human review can supply and no live pipeline can safely generate.

grid_on

The matrix

Every scenario against every contact archetype, attitude, difficulty, language and interaction history. Every cell populated rather than sampled, because the point is to find where a worker fails.

mood

The attitude range

Friendly, concerned, busy, sceptical, hostile, gatekeeping, confused, monosyllabic, digressive, the wrong person entirely, and a contact who asks to be removed.

key

Seeded cues

The simulated contact holds facts it will disclose only under stated conditions, so the harness knows the ground truth and can measure what the worker actually extracted.

psychology_alt
Seeded cues are what make simulation measurable rather than merely voluminous. A referral is given only if permission is asked properly. A funding stage is named only if the worker probes past the first mention. A correction to the briefing appears only if the worker treats its intelligence as hypothesis. And a decoy, being a plausible but incorrect statement, tests whether the worker records what it was told or what it verified.

Two measures, not one

Cue recall is the proportion of available intelligence the worker actually elicited. Contract fidelity is the proportion of what it elicited that reached the outcome record correctly.

Separating them matters because two very different failures look identical from outside. A worker that never asked the right question has a conversation problem. A worker that heard the answer and did not record it has a contract problem. Only ground truth tells them apart.

Voice simulation is required, not optional

Text mode carries scale. Voice carries realism. Neither substitutes for the other.

  • Agent versus agent in voice tests barge-in, overtalk, silence handling, pacing, mishearing and voicemail. Repeatable, so it sits in the promotion gate
  • Agent versus human in voice adds the unscripted awkwardness a model does not produce: talking over the worker, going quiet, changing their mind. A judgement rather than a threshold, and it doubles as rehearsal
rule
Gated on floors, not on an average. Results are reported per cell, because a revision can lift the overall mean while collapsing against hostile contacts, and an average that hides a total failure is worse than no measurement. The gate requires no compliance findings anywhere, an abrupt ending rate of zero, and cue recall above threshold in every cell.
The domain coach

The first sales coach in the building with the data.

A human sales coach normally works from anecdote, from the calls they happened to sit in on, and from the seller's own account of what happened. A domain coach on the Hub works from every scored conversation, every typed finding, every cue that was available and missed, and the sentiment trajectory of each call.

ModeWhat happens
Deal coachingWorks the qualification gaps on a live opportunity: which elements are unfilled, which rest on weak evidence, who the economic buyer likely is given the stakeholder map, and what single question would most reduce uncertainty next
Call preparationThe briefing before a human takes a call: what the account intelligence supports, the objections most likely for this role and industry, the insight worth leading with, and what the last conversation left open
Post call debriefThe same analysis machinery pointed at a person's call: what earned progress, what was left on the table, which cues were missed, and where the conversation turned
Skill developmentLongitudinal rather than per call. Talk and listen ratio, question type distribution, discovery depth, objection handling success by category, and whether insights precede asks
Objection rehearsalLive practice against a seeded simulated contact, at a chosen attitude, in a chosen language, as many times as you want, with no prospect at risk
Methodology teachingExplaining a framework in the context of your own live deal rather than in the abstract, which is the only form of methodology training that survives a quarter end

Rehearsal reuses the simulation harness, pointed at a person instead of a worker. Role play normally needs a manager's time, happens rarely, and is scored on impression. Here it is on demand, scored on the same rubric as production conversations, and repeatable until the pattern changes.

privacy_tip
The governance line. Coaching records about a named person are restricted data, collected to develop them, visible to the person they describe, and contestable. Development and performance management are different purposes, and the platform keeps them technically distinguishable rather than separated by policy alone. The coach recommends. It does not assign, does not escalate about a person, and does not write to a human resources system.
Next

Improvement you can audit is improvement you can sell internally.

Every conversation your workers hold makes them measurably better, and the improvement is expressed in versioned artefacts you own and can inspect.