arrow_backBack to the overview
Calibration & coaching

A worker that measurably gets better.

Score everything. Classify the failures. Prove the fix.

Most AI deployments improve by anecdote: someone hears a bad call, someone edits a prompt, nobody can say whether it helped. The Hub treats improvement as an engineering discipline. Every conversation is scored, low scores become classified defects with an owner, revisions are proposed with the evidence attached, a person approves them, and nothing reaches production until it passes regression against a saved suite.

Every callscored automatically, not sampled
Evidence boundfindings quoted from the transcript
Human approvalno prompt changes itself
Scoring

The outcome carries the score.

The rubric is derived from the outcome contract, which means any worker with a contract can be scored without bespoke work. The model produces the evidence. The code produces the number. A score computed by a model is not reproducible, cannot be regression tested and cannot be compared across prompt versions, so the scorer is deliberately dumb arithmetic and nothing about that is negotiable.

ComponentContribution
Primary outcomeCarries the score. Each rung is worth materially less than the one above it, from full marks for the top rung down to nothing for no outcome at all
Outcome statusA tentative outcome earns partial credit, rising when a concrete confirmation path was agreed. Honesty is rewarded, and an inflated claim of confirmed is a defect rather than a win
Partial upgradeA recorded, transcript supported path to a better outcome adds credit without allowing the higher rung to be claimed
Standing objectivesA fixed bonus for each one carrying evidence, with the total capped. Judged on whether the worker took what the conversation naturally offered, not on whether it ground through a checklist. The ceiling differs per contract, because the objective sets differ in size, and any report that hides that invites the conclusion that one set of callers is underperforming
Behavioural failuresNot deductions. Restating, stacked questions, premature endings, missed disclosure and unresolved identity are certification cells, and a blocking cell fails regardless of score. Trading them against a number is exactly what the separate gate exists to prevent

What scoring is for

A score exists to direct attention. It tells you which conversations to look at, which prompt layer to examine, which vertical is underperforming, and which skill someone should practise. That is the whole of its purpose, and the narrowness is deliberate.

What it is never used for

  • Not a reward signal. We do not optimise a model against its own rubric. Automated optimisation against a proxy is how a system learns to score well rather than to work well
  • Not training data. Scored conversations improve prompts, knowledge and skills, all artefacts you own and can inspect. They do not become weights
  • Not a performance measure. Attaching a score to someone's pay converts an honest diagnostic into an incentive to game the record
Agent against agent, with the answers seeded in advance.
Agent against agent, with the answers seeded in advance. Because the harness knows which facts were available to be found, recall and fidelity are computed rather than judged. The decoy row is a fact that was never true: if it reaches the structured record, that is fabrication, measured directly. One cell below its floor fails the run even though the average improved.
Two gates

Good enough to deploy is a different question from allowed to run.

Calibration asks whether this worker is good enough to deploy here, and failure looks like a low score, a regressed cell or a missing cue. Certification asks whether this pack may be listed for this channel in this jurisdiction, and failure looks like a compliant-seeming transcript that breaks a rule. They are separate gates because the failures look nothing alike, and a scoring threshold passes both.

The failures a score cannot see

A conversation that delivers every required identification line, but delivers one of them in turn three when the rule requires it before the first question, is non-compliant and scores well. A conversation that continues in English on a non-English turn scores well too, and produces something the contact cannot consent within. Both are blocking cells. Neither is reducible to a number.

The grid is per cell, not per pack

Every channel crossed with every jurisdiction, each cell passed independently. Disclosure requirements, consent bases and permitted hours differ by jurisdiction, and a channel changes what a ladder rung can even mean, so a pass in one cell says nothing about any other. A jurisdiction with an unanswered dimension cannot be certified at all, because there is nothing complete to certify against.

balance
Property adds a fair treatment review, and it gates go-live. Six cells, all blocking. Five check specific behaviours: the boundary between a lawful question and an unlawful one, volunteered characteristics not being recorded, a discriminatory instruction being declined and escalated rather than quietly complied with or covered by an invented reason. The sixth checks what no transcript review can see. Two runs differing only in an inferable characteristic must produce the same properties offered, the same slots proposed and the same qualification path. A worker that treats people differently produces conversations that read perfectly well one at a time, so the only check that finds it is mechanical and runs every build. There is no waiver for this group.
split_scene
Baselines are recorded per contract, and never pooled. Inside sales tracks its distributions separately: mass text simulation, voice simulation and live conversations. Property records a baseline for each of its five contracts. Synthetic contacts are more cooperative than real ones, so pooling simulation with live produces a number that describes neither population, and a buyer score and a vendor score measure different work. Every aggregate carries its channel mix and its distribution mix as a visible label, and an unlabelled aggregate must not be producible. This is the rule that gets quietly broken first, and the failure is invisible because the arithmetic stays correct.
Failures

A bad call becomes a defect with an owner.

Alongside the score, the evaluator returns typed findings, each anchored to a specific quote or turn reference. A number tells you a call went badly. A quoted finding tells you which line of which prompt layer caused it.

ClassMeaningWhere it goes
BehaviouralThe worker knew what to do and did not do itPrompt revision queue
KnowledgeThe worker lacked information it neededCuration task in the knowledge base
AssetThe worker needed collateral that did not existThe request board
CapabilityThe worker needed a tool or channel it does not haveRoadmap
ConfigurationWrong engine, thinking level, voice or scenario bindingControl plane
ComplianceA guardrail was approached or crossedImmediate escalation, never queued
query_stats
One bad call is noise. The same finding across forty calls in one vertical is a defect with a business case. Recurring signatures are what drive prioritisation, which is why classification matters more than volume.

The loop

Conversation Score Classify Coach proposes 3 to 6 surgical edits A person approves Regression suite Promote

The coach never rewrites a prompt wholesale, and that restraint is deliberate: wholesale rewrites destroy the accumulated calibration of a mature prompt and make regression impossible to attribute. Every version carries its author, its motivating findings, its regression result and its approver, so a prompt in production can always be traced back to the conversations that shaped it.

Simulation

Harden a worker before it speaks to anyone.

The behaviours that matter most appear rarely and cost the most when they do: an abrupt ending while a question is still open, a settled topic re-asked, a referral offered and not captured. Catching those needs volume that no human review can supply and no live pipeline can safely generate.

grid_on

The matrix

Every scenario against every contact archetype, attitude, difficulty, language and interaction history. Every cell populated rather than sampled, because the point is to find where a worker fails.

mood

The attitude range

Friendly, concerned, busy, sceptical, hostile, gatekeeping, confused, monosyllabic, digressive, the wrong person entirely, and a contact who asks to be removed.

key

Seeded cues

The simulated contact holds facts it will disclose only under stated conditions, so the harness knows the ground truth and can measure what the worker actually extracted.

psychology_alt
Seeded cues are what make simulation measurable rather than merely voluminous. A referral is given only if permission is asked properly. A funding stage is named only if the worker probes past the first mention. A correction to the briefing appears only if the worker treats its intelligence as hypothesis. And a decoy, being a plausible but incorrect statement, tests whether the worker records what it was told or what it verified.

Two measures, not one

Cue recall is the proportion of available intelligence the worker actually elicited. Contract fidelity is the proportion of what it elicited that reached the outcome record correctly.

Separating them matters because two very different failures look identical from outside. A worker that never asked the right question has a conversation problem. A worker that heard the answer and did not record it has a contract problem. Only ground truth tells them apart.

Voice simulation is required, not optional

Text mode carries scale. Voice carries realism. Neither substitutes for the other.

  • Agent versus agent in voice tests barge-in, overtalk, silence handling, pacing, mishearing and voicemail. Repeatable, so it sits in the promotion gate
  • Agent versus human in voice adds the unscripted awkwardness a model does not produce: talking over the worker, going quiet, changing their mind. A judgement rather than a threshold, and it doubles as rehearsal
rule
Gated on floors, not on an average. Results are reported per cell, because a revision can lift the overall mean while collapsing against hostile contacts, and an average that hides a total failure is worse than no measurement. The gate requires no compliance findings anywhere, an abrupt ending rate of zero, and cue recall above threshold in every cell.
Regression

Every version is tested before anyone reads the proposal.

The internal use of simulation is the one that never gets demonstrated and matters most.

Every candidate prompt version enters a regression run without anybody triggering it. Whether the revision came from the coach, from a steward's own edit or from a pack update, the run starts and the result is attached before a human looks at it. So an approval queue presents a proposed edit with its regression result already in hand, not a proposal somebody has to remember to test. A revision that fails never reaches a steward as a decision. It reaches them as a finding, which is a different and much cheaper conversation.

The cost of promoting a change is set by the cost of undoing one. A platform where reverting is hard develops a culture of not changing anything, and a worker nobody dares improve stops improving.

undo

Rollback is a pointer change

Every layer is independently versioned and every conversation records the version it ran under, so reverting takes seconds and is itself a recorded event with a reason.

monitor_heart

A promoted version is watched

Regression proves a version against the suite, not against the world. Live scores, abrupt endings, compliance findings and outcome distribution are compared against the pre-promotion window, and one that degrades past a declared threshold reverts on its own. Conservative on purpose: an unnecessary rollback costs a day, an unnoticed regression costs a quarter.

compress

Consolidation is scheduled

Revisions accumulate until the prompt is sediment nobody dares edit. A steward periodically merges them, proves equivalence by regression, prunes what no longer earns its place and re-baselines. Where a prompt gets shorter rather than longer.

rule
The suite is versioned too, and comparisons across versions are refused rather than caveated. The most common way a regression harness lies is by comparing a result against a baseline produced under a different suite, which makes an improvement and a regression indistinguishable. Adding cells, which happens every time a failure is classified, costs a re-baseline run against the current production version. That cost is what keeps the numbers meaningful.
Campaign simulation

Will this campaign work, and what will it cost?

The same harness pointed at a campaign rather than a worker, which answers a commercially larger question than prompt quality.

Resolve the campaign's real membership. Generate synthetic counterparties matched to the actual distribution in that list rather than a generic matrix, being the real mix of roles, seniorities, languages, industries and interaction histories it contains. Run at volume in text with a voice sample for realism. Then report.

OutputThe question it answers
Projected contact and conversion distributionIs the coverage target achievable with the capacity allocated?
Projected cost, per outcome and in totalIs the cost per outcome viable before we spend anything?
Projected duration against the cycle windowWill this finish inside the window or run past it?
Failure modes, with the cells that produced themWhere does the ladder break for this population?
Messaging findingsWhich openings and objection handlers landed against this audience specifically
Compliance dry runHow many of these people are actually contactable
fact_check
The cheapest output is frequently the most useful. The compliance dry run resolves the membership, puts every subject through the precheck, and reports the eligible count before anybody commits to a target. That single step prevents the most common campaign failure in this category, which is discovering three days in that ten thousand names were three thousand two hundred contactable ones and the target was never reachable.

Variants compared before launch

Two messages, two ladders, two list definitions, run against the same synthetic population. That makes it a real experiment because the population is held constant, and it happens before money is spent rather than after half of it is.

And one caveat we put next to the number

Synthetic counterparties are more cooperative, articulate and patient than real ones. So a simulated conversion rate is a relative comparison and an upper bound, never a forecast, and we say so where the number appears rather than in a footnote. Failure modes are treated as real, because one that occurs against a cooperative synthetic contact will certainly occur against a real one. The asymmetry runs one way.

insights
Projections calibrate against actuals. Once a campaign has run, its simulated and observed distributions are compared and the ratio is retained per worker, per market and per list origin. Over time the projection becomes a genuine estimate rather than a hope, and until it does, the platform tells you which it is currently producing.
The domain coach

The first sales coach in the building with the data.

A human sales coach normally works from anecdote, from the calls they happened to sit in on, and from the seller's own account of what happened. A domain coach on the Hub works from every scored conversation, every typed finding, every cue that was available and missed, and the sentiment trajectory of each call.

ModeWhat happens
Deal coachingWorks the qualification gaps on a live opportunity: which elements are unfilled, which rest on weak evidence, who the economic buyer likely is given the stakeholder map, and what single question would most reduce uncertainty next
Call preparationThe briefing before a human takes a call: what the account intelligence supports, the objections most likely for this role and industry, the insight worth leading with, and what the last conversation left open
Post call debriefThe same analysis machinery pointed at a person's call: what earned progress, what was left on the table, which cues were missed, and where the conversation turned
Skill developmentLongitudinal rather than per call. Talk and listen ratio, question type distribution, discovery depth, objection handling success by category, and whether insights precede asks
Objection rehearsalLive practice against a seeded simulated contact, at a chosen attitude, in a chosen language, as many times as you want, with no prospect at risk
Methodology teachingExplaining a framework in the context of your own live deal rather than in the abstract, which is the only form of methodology training that survives a quarter end

Rehearsal reuses the simulation harness, pointed at a person instead of a worker. Role play normally needs a manager's time, happens rarely, and is scored on impression. Here it is on demand, scored on the same rubric as production conversations, and repeatable until the pattern changes.

privacy_tip
The governance line. Coaching records about a named person are restricted data, collected to develop them, visible to the person they describe, and contestable. Development and performance management are different purposes, and the platform keeps them technically distinguishable rather than separated by policy alone. The coach recommends. It does not assign, does not escalate about a person, and does not write to a human resources system.
Next

Improvement you can audit is improvement you can sell internally.

Every conversation your workers hold makes them measurably better, and the improvement is expressed in versioned artefacts you own and can inspect.