arrow_backBack to the overview
Operations & cost

You cannot run what you cannot see or price.

A trace per conversation. A cost per interaction. A budget that acts.

This is the layer that makes a fleet of virtual workers operable rather than experimental, and it is the layer most often retrofitted. Retrofitting is expensive: telemetry added after the fact never covers the paths that turn out to be interesting, and entitlement added after the fact is unenforceable because the code already assumes everything is permitted. The Hub emits a structured trace and a cost record from the first conversation it handles.

Span levellatency captured per step, not per call
Per interactioncost attributed along every dimension
Never silentlimits warn, act and raise an event
Telemetry

A conversation is a trace, not a log line.

Context assembly, each turn, each tool call, each skill invocation, retrieval, write back and evaluation, with latency captured per span. That is what makes a voice worker debuggable at all. When a call feels slow, the answer is almost never "the model" and almost always one specific span.

Dimension groupCaptured on every interaction
IdentityTenant, worker, the worker definition version, and the version of each of the six prompt layers actually in force
ConfigurationThe engine bindings actually used, being model and thinking level per task class, voice provider and voice, and any fallback that was triggered
ContextScenario, engagement goal, channel, language spoken, record language, jurisdiction, capture level, and the account and contact under the required pseudonymisation
ConductTurn count, duration, talk and listen ratio, interruptions, tools called, skills invoked, retrievals issued, escalations and requests raised
ResultOutcome level and status, score, typed findings, cue recall where simulated, transaction outcomes, write backs performed
CostEvery priced unit consumed, attributed to the task class that consumed it
monitoring

Operations asks: is it working?

Voice first token latency, turn latency at median and tail, tool error rates, channel delivery failures, degradation events, queue depth, concurrency against capacity.

payments

Finance asks: what did it cost?

Per interaction cost attributed along every dimension, and the ratios that matter commercially rather than raw token counts.

lightbulb

Product asks: what next?

Which failure classes dominate, which simulation cells are weak, which knowledge gaps recur, which requests are raised most often.

link
A correlation identifier threads from the originating event through the flow, the conversation, the outcome contract and the system of record write back, so a booked meeting traces back to its trigger and a bad write back traces forward from the turn that caused it. Telemetry and audit stay separate systems: audit is immutable and compliance scoped, telemetry is operational and carries identifiers and measures, never conversation content.
Cost capture

A conversation with a known price.

Cost is captured per interaction and attributed along every telemetry dimension, so one record answers what a worker, a channel, a language, a scenario or a campaign costs.

Cost classNotes
Live conversation tokensAudio dominates by a wide margin, which has a design consequence worth stating
Voice engine timePer provider, which is one of the axes on which two providers actually get compared
CarriageTelephony minutes and messaging fees, per channel and per destination
Deliberative model callsAnalysis, coaching, research and extraction. Individually large, but bounded and predictable per conversation
Retrieval and embeddingKnowledge queries during and after a conversation
StorageThe sleeper cost. Retained audio at scale, in many languages, under a long retention period
SimulationThe mass suite has a real price, and it is visible before a run rather than after

A worker that talks less is both better and cheaper. The rule requiring one short sentence per turn exists because it makes the worker sound like a competent colleague. It also happens to be the single largest cost lever in the system, so conversation design and unit economics point the same way.

forum

Cost per conversation, by worker, channel, language and scenario

event_available

Cost per booked meeting or completed conversion, which is the number that matters

trending_up

Cost per qualified opportunity, traced through to the system of record

badge

Projected cost per worker per month, so a buyer can hold it against a salary

savings
Budgets act, they do not merely report. Ceilings at tenant, worker and campaign level, with a soft threshold that warns and a hard threshold that does something configured: degrade to a cheaper binding, queue non urgent work, or stop and escalate. The one prohibited behaviour is silent failure. And a live conversation is never terminated by a budget: quota is evaluated when a conversation starts, never mid turn, because hanging up on a prospect to save tokens is worse than any overspend it prevents.
Entitlements, limits and caps

What you bought, enforced where it matters.

Entitlement is a commercial statement. Limits are the technical enforcement. Keeping them distinct is what stops a sales concession from requiring a deployment.

DimensionTypical shape
WorkersWhich worker packs are licensed, and how many instances may be active at once
ChannelsWhich are enabled, since telephony and messaging carry both cost and compliance weight
LanguagesWhich of the supported set are enabled for your tenant
Engine tiersWhether the deliberative and elevated thinking tiers are available, or only standard bindings
Knowledge and assetsNumber and size of knowledge bases, asset volume, retention periods
RecordingWhether audio capture is available at all, and for how long recordings are retained
SimulationMass runs per period, and whether the voice tier is included
TransactionsWhether transactional capability is enabled, and the value ceiling that applies
SeatsHow many users of each role, which is how the coach and the studio are packaged
Interface rateAPI and MCP call rates, and concurrency

Enforced at two points, never one

The control plane refuses configuration that exceeds entitlement, so you cannot enable a channel you have not bought and the option is visibly unavailable rather than failing later. The runtime enforces consumption limits, being rate, concurrency and spend, at the point of use.

Behaviour at a limit follows the same principle as budget: warn before the ceiling, act at it in a configured way, name the limit and the remedy in language an operator can act on, and raise it as a system event.

Capped pilots are a first class shape

A pilot is an entitlement configuration: a few workers, one channel, one language, a fixed simulation allowance, a spend ceiling and a time bound. Making that a configuration rather than a manual arrangement is what lets a pilot be provisioned in a day and converted without a rebuild.

Metering is the same record that feeds cost capture, deliberately. A usage based model requires a meter you can audit, and a meter that disagrees with the cost attribution is worse than no meter at all. One record, two readings.

The control plane

Workers are supervised, and the supervisor may be a worker.

The layer that defines, configures, assigns and supervises workers is kept separate from the data plane in which conversation happens, because the two have different availability requirements. A control plane outage must never take a live call down.

tune

Definitions and templates

Authoring, layered prompt composition, versioning, staging, promotion gates, agent templates, voice profile mappings, scenario and flow libraries.

groups

Rosters and assignment

Which workers exist, which are active, which people each serves, and which accounts, territories, languages and channels each covers.

speed

Capacity and pacing

Daily volumes, calling windows, concurrency and spend ceilings, all adjustable without a release.

Supervisory responsibilityHuman managerVirtual manager
Monitor score distributions and flag driftReviews dashboards on a cadenceContinuous, alerts on threshold breach
Triage the failure ledgerJudgement on ambiguous casesClassifies and clusters recurring signatures
Approve prompt revisionsHolds the authorityMay propose only. Approval stays with a named person
Reallocate objectives and prioritiesStrategic judgementWithin pre-approved bounds only
Handle escalations from workersOwns the outcomeRoutes and prepares context. Never resolves a compliance escalation
Coach individual workersReviews sampled conversationsRuns the calibration loop and surfaces the exceptions
gpp_maybe
What makes a virtual supervisor safe is that its authority is a strict subset, and the boundary is explicit. It may observe, classify, cluster, prioritise, propose and reallocate within bounds. It may not approve changes to a kernel or client prompt layer, alter compliance configuration, expand its own authority, or resolve an escalation that reached it because a guardrail was approached. The practical case for it is simple: the calibration loop generates far more signal than a human manager can read, and a virtual manager is the filter that turns thousands of scored conversations into the dozen decisions a person should actually make this week.
The request board

The worker asks for what it does not have.

When a worker reaches a moment it cannot serve, such as a contact asking for a case study in a vertical that has none, the correct behaviour is neither to improvise nor to fail silently. It is to tell the contact honestly that it will follow up, log the gap, and raise a request.

Raised

New requests, with the conversation, contact and account attached as context.

Triaged

Classified, prioritised and owned. Priority follows how many conversations are blocked on the same gap.

In progress

Being produced, by a person or by another virtual worker.

In review

Awaiting approval before entering the asset library or the knowledge base.

Available

Published. The requesting worker is notified and a proactive event fires to every relevant contact.

Declined

With a reason, which is itself intelligence. A repeatedly declined request is a strategy question.

The board also carries the calibration queue of proposed prompt edits and the escalation queue, so everything needing a human decision about the worker fleet lives in one place. Two properties make it more than a task list: requests carry their originating context, so whoever picks one up sees the actual conversation that needed it, and an available asset closes the loop backwards, because the contact whose question could not be answered becomes a queued proactive touchpoint the moment it is published.

Next

See how it is packaged and priced.

A platform licence, workers as the unit of value, domain content on subscription, and usage passed through with a meter you can audit.