A Chorus transcript accuracy test compares the contracted ZoomInfo Chorus output with a blinded human reference on the calls, languages, audio paths, speakers, summaries, actions and CRM associations your team uses. Measure WER, critical tokens, speaker attribution, proposition/action precision and recall, critical errors, association, duplicate/retry behavior, privacy and deletion before rollout.
This is not the Gong versus Chorus comparison, a Chorus review, or a general notetaker guide. The Gong and Avoma accuracy tests own other systems. Every worked number below is synthetic.
Define the Chorus validation boundary
This page owns buyer-controlled measurement of one configured pipeline: calendar/telephony and consent → recording → transcript → speaker → summary/topic/action → review → CRM association/write → audit/correction. It does not rank conversation-intelligence products or assert Chorus supports an artifact merely because another tool does.
ZoomInfo’s public Chorus product listing describes recording, transcription and analysis. A historical ZoomInfo announcement described one Salesforce/Engage-to-Chorus recording path. Neither proves the current behavior of your contract. Treat demonstrations and sales answers as claims to convert into written, testable acceptance rows.
The sales call recording guide owns capture-tool selection; this protocol starts after Chorus is already the candidate. Include the competing manual process as a baseline so a feature can be rejected when review and correction work exceeds its value.
Freeze the documented configuration
Create a capability contract before collecting audio. Record edition/package, region, workspace, users, administrators, calendar/meeting/telephony sources, bot or native capture, supported languages/dialects, transcript/diarization behavior, summary/topics/actions, templates/trackers, CRM and objects, integration identity, field mappings, permissions, export formats, retention, deletion, support access, audit and processing date.
For each artifact label documented, observed, unavailable, disabled, or unresolved. The evidence ledger has zero unresolved editorial claims; your procurement packet may have unresolved product questions, which must remain gates rather than being filled by inference. If ZoomInfo cannot document current language, export, retry or CRM behavior, exclude it from automatic use until a sandbox test establishes it.
Build a stratified, consented call corpus
Sample production distribution and risky edges. Include discovery, demo, negotiation, renewal, onboarding, support and internal calls as relevant. Label language, dialect/operational speech condition, code-switching, number of speakers, roles, duration, platform, headset/room/mobile/PSTN, noise, echo, overlap, low bandwidth, hold music, late joiner, transfer and recording source.
Add critical content: person/company/product names, acronyms, amounts, currencies, dates/time zones, quantities, negation, uncertainty, conditions, promises, decisions, actions, owners and deadlines. Add CRM cases: one/multiple contacts, leads, accounts and opportunities, duplicate people, parent/subsidiary, changed owner, unknown attendee and two concurrent deals.
Use only recordings approved for the stated purpose, consent/notice, access and retention. Log consent path, opt-out, excluded participant and stopped recording. Qualified privacy/legal owners must define requirements; a prohibited capture fails before accuracy scoring.
Create and calibrate the blind reference
Two trained annotators create the reference without seeing Chorus output. Lock normalization for punctuation, casing, fillers, numbers, partial words, overlap, inaudible spans and speaker turns. Define propositions, summary-required facts, topics, actions, explicit versus inferred commitments, owner, due date, CRM target and critical error.
Calibrate on a held-out subset, report raw and suitable chance-adjusted agreement, revise the codebook, then freeze it. Adjudicate disagreements while preserving original labels and unresolved audio. Each fact needs speaker, timestamp/evidence span, normalized value, polarity/negation, certainty, criticality and expected CRM destination.
Measure transcript, critical-token, and speaker accuracy
Use WER for transcript comparability, not as the whole verdict. The NIST OpenASR evaluation plan defines WER=(substitutions+deletions+insertions)/reference words. Publish numerator, denominator and distribution by language, audio, call type and speaker condition. Keep unscorable audio separate.
Also score critical-token accuracy for names, products, amounts, currencies, dates, negation and commitments. For speaker attribution, classify each scored utterance correct, wrong, unknown or improperly split/merged. Report attributed-utterance accuracy and critical speaker reversals. A perfect sentence assigned to the wrong party can invert commercial meaning.
Do not infer accent from identity or appearance. Use responsibly collected operational labels only where needed, preserve unlabeled cases and suppress small-group reporting that could identify speakers.
Score summaries, topics, and actions
Score generated artifacts as atomic propositions. Summary precision is supported generated propositions divided by generated propositions; recall is required reference propositions recovered divided by required reference propositions. Label unsupported, contradicted, wrong speaker, wrong certainty, wrong topic and duplicate separately.
For actions, score action, owner and due date separately and jointly. An invented deadline or buyer commitment can fail a critical gate even if the prose looks useful. If the contracted product exposes topics, trackers or next steps, define expected labels before processing and measure precision/recall. If it does not, do not invent the feature for the test.
Test CRM association, retries, and idempotency
An accurate artifact on the wrong record is a critical failure. In a CRM sandbox, seed duplicate contacts, contact plus lead, parent/subsidiary, unknown attendee, two active opportunities, merged/deleted record, deactivated owner, invalid required field, permission rejection, rate limit, timeout before/after write, replay, out-of-order event and partial batch failure.
Association precision equals correctly targeted written artifacts divided by artifacts written. Association recall equals required approved artifacts correctly written divided by required reference artifacts. Track wrong object, owner, account, opportunity, activity type, timestamp, source link, field and privacy state.
Test idempotency by repeating the same accepted event and retrying after a simulated timeout. Final state should contain one intended artifact. Log source call/event ID, attempt, target, prior value, result, error, retry and audit. Default integration writes off until this test passes.
Test recording, permissions, retention, export, and deletion
Test capture, access and lifecycle controls. Cover calendar connected/disconnected, assistant admitted/denied, duplicate bot, canceled/rescheduled call, recording disabled, unsupported source, wrong language, late joiner, user deprovisioned and permission removed. Measure missed/extra captures and time to artifact.
Create test users for seller, manager, enablement, admin, unrelated team and external participant. Verify recording, transcript, summary, action, search result, clip, export, shared link and CRM artifact access. Test revocation and support access. A restricted meeting leaking through search or export fails regardless of accuracy.
Obtain the current ZoomInfo DPA, Chorus customer-content terms, security evidence, retention schedules, subprocessors, regions, export/deletion mechanics and incident commitments. The public ZoomInfo privacy policy is not a substitute. Test export completeness and delete a test meeting/user/workspace as contracted, recording actual propagation and documented retention without promising immediate erasure.
Apply hard gates, scorecard, and worked example
Hard gates precede the weighted score: approved capture/consent; zero restricted leaks; zero unauthorized CRM writes; zero critical wrong-record associations; maximum critical fact/action error; required export/deletion; least privilege; visible retries; idempotency; audit; rollback; and incident ownership.
| Dimension | Weight |
|---|---|
| Transcript/critical tokens | 20 |
| Speaker attribution | 10 |
| Summary/topics | 20 |
| Actions/owners/dates | 15 |
| CRM/reliability | 15 |
| Privacy/governance | Hard gate + 10 |
| Economics/adoption | 10 |
Synthetic example: a 20-word reference with one substitution and one deletion has WER 2/20=10%. Ten required propositions exist; output contains nine, seven supported, six required: precision 7/9=77.8%, recall 6/10=60%. Of eight required CRM artifacts, seven land once correctly and one lands on the wrong deal: reconciliation 7/8=87.5% plus one critical association error. These are not Chorus results.
Run a bounded pilot and calculate TCO
Run baseline, artifact-only, sandbox integration, failure injection and bounded production phases. Use ordinary users and the same definitions. Track latency, review/correction minutes, overrides, missed captures, incidents, adoption and drift. Approve only the languages, capture paths, call types, artifacts and CRM behaviors that pass.
Term TCO equals Chorus/ZoomInfo licenses and required bundles + meeting/CRM dependencies + implementation + templates/trackers + security/privacy/legal + corpus/double annotation + integration/failure tests + human review/correction + administration/monitoring + storage/export/retention + support/incidents + parallel tools and exit − retired labor/tools. Use current quotes; no price is fabricated here.
Revalidate after model, language, template, CRM, permission, retention, capture or contract changes. The defensible conclusion is bounded: this documented configuration passed these predeclared tests on this corpus and workflow, with these limitations.