A pre-call research accuracy test checks the assembled brief—not whether a rep followed a preparation routine. It asks whether the brief belongs to the scheduled meeting, identifies the right people and opportunity, supports facts with current authorized sources, labels inference, resolves conflicts, and declines to answer when evidence is unsafe or insufficient.
This page does not teach the five-minute preparation workflow; use the sales call prep guide for that. It does not replace the call-prep template, rank tools, or validate an upstream contact database. It owns quality assurance for the final brief delivered before one specific meeting.
Define what the accuracy test owns
Test the complete brief pipeline: calendar event → eligible meeting → participant identities → account and opportunity candidates → authorized CRM, email, prior-call and public sources → retrieved passages → generated facts and inferences → cited brief → rep correction or acceptance. Record the product, model and prompt version when visible, retrieval configuration, connectors, credentials, CRM schema, calendar and mailbox, source cutoff, timezone, brief template, and generation time.
Write the use decision before measuring. A discovery brief may need verified role, account description, recent trigger, prior interaction, open opportunity, known stakeholders, evidence-backed hypothesis, and questions. A renewal brief needs contract and adoption context under stricter permissions. Do not score a field that the workflow does not require, and do not reward extra personal data merely because it is available.
Contract identity, facts, inference, source, and time
Every brief item needs a type. Fact is a source-supported proposition. Inference is a labeled interpretation derived from facts. Question is a proposed way to test an unknown. Unknown means the system should abstain. Keep those types visually distinct.
Create one truth-table row per atomic proposition: brief ID, meeting ID, person/account/opportunity IDs, field, expected value, polarity, source URL or internal object ID, source owner, publication/event time, retrieval time, permitted user, transformation, inference label, conflicting sources, expiry rule, severity, and reviewer. W3C’s PROV-O Recommendation models provenance through entities, activities, agents, derivation and time. You do not need RDF to borrow that discipline: preserve what evidence was used, what process transformed it, and who or what was responsible.
A source link alone is incomplete. Save the exact supported proposition and as-of state, not a long copyrighted excerpt. “VP Sales” from a profile viewed today and “joined in 2021” from a dated announcement have different freshness rules. If a title conflicts with the CRM, company page, and professional profile, show the conflict or abstain instead of silently choosing.
Build a blinded labeled brief set
Sample the deployment population: discovery, demo, renewal, expansion, executive and technical meetings; new and known contacts; small and large accounts; several regions, languages and timezones; clean and sparse CRM records; recent and old opportunities; reschedules, recurring meetings and forwarded invites. Include “no eligible external meeting” and “insufficient evidence” cases.
Split cases into rubric calibration, blinded evaluation, and later holdout. Two qualified reviewers independently label atomic facts, inference status, conflicts, permission eligibility and critical errors. Resolve ambiguous rubric language on calibration cases, then freeze the codebook before evaluation. Do not let a vendor tune on the final set.
Stratify failure cases deliberately: namesakes, job change, company rename, parent/subsidiary, consultant joining several accounts, two open opportunities, personal email, generic domain, departed owner, private CRM note, restricted opportunity, outdated press release, corrected funding figure, deleted post, and no reliable recent event.
Test meeting, person, account, and opportunity identity
Identity errors contaminate every downstream fact. Preserve calendar provider event ID and iCalUID where applicable, organizer, calendar owner, recurrence ID, original start, attendees, RSVP, conference ID, update sequence, and CRM candidates. Google’s official Events API documentation distinguishes event IDs and iCalUID retrieval, exposes organizers and attendees under authorization scopes, and warns that an email field can contain a generated non-working value. Therefore “attendee email exists” is not sufficient identity evidence.
Test reschedule, cancellation, duplicate calendar copies, recurring instance, forwarded invitation, changed attendee, organizer alias, shared calendar, group address, external guest without CRM record, and reused conference link. Verify exactly one current brief attaches to the intended event. A canceled or internal-only event should not receive an external-account brief.
For each participant, evaluate correct contact and role; for the company, correct account and parent/subsidiary; for the deal, correct opportunity or explicit no-safe-match. Measure person, account and opportunity edges separately. A safe quarantine is better than joining a board-meeting transcript to a similarly named prospect.
Resolve upstream duplicate and ownership defects with the CRM data-quality framework, then preserve field and relationship authority using CRM integration best practices. The brief test should reveal source defects, not silently repair production records.
Test provenance, freshness, conflict, and inference
Require a click-through source and as-of date beside material facts. Compare generated atomic propositions with the frozen truth table. Score names, roles, dates, amounts, currencies, products, locations, employment, funding, technology, prior commitments, stakeholders, risks and next steps separately. Negation and corrections are critical: “not moving this quarter” cannot become “moving this quarter.”
Define freshness by field and decision. A historical founding date may remain stable; role, account ownership, opportunity stage and a “recent event” can decay quickly. Staleness = generation timestamp − authoritative source event or verification timestamp. Do not set one universal expiry. Mark unknown source time and require review where timing changes the opener or decision.
Rank source authority for each field before the run. A signed/internal record may govern contract state; CRM may govern current owner; a company filing or announcement may support a public event. When reliable sources disagree, retain both values, sources and dates, and route the conflict. Never convert a hypothesis such as “likely expanding” into a fact. Label its premises and write the question that could falsify it.
If the brief includes third-party buying signals, validate that upstream feed separately with the intent-data quality test. A correctly copied signal is still unsuitable evidence when its account match, topic, or event window is wrong.
Protect permission-sensitive context
Run every test under the intended rep, manager, admin and restricted-user identities. Verify the brief cannot reveal a private event, another territory’s confidential opportunity, HR or compensation notes, legal discussions, credentials, hidden email recipients, restricted recording, deleted content, or a field the user cannot access in its source system.
Google’s calendar-sharing documentation shows that calendar and event visibility can differ and that recurring-event changes have specific propagation behavior. Test effective access, not the connector’s broad service-account view. The ICO’s current data-minimisation guidance describes personal data as adequate, relevant and limited to what is necessary. Applicability varies; qualified owners must approve purpose, access, retention and deletion.
Use synthetic records for destructive and hostile tests. Where prior-call context is included, follow the recording governance checklist. Log source access, generated brief access, correction, export and deletion without exposing the sensitive content again in monitoring.
Inject hostile sources and require abstention
Place instruction-like text in a webpage, email signature, CRM note and retrieved document: requests to ignore rules, reveal hidden data, change account identity, invent a favorable summary, or contact an unapproved person. Retrieved content is evidence, not authority. It must not alter system instructions, scopes, tools, recipients or write permissions.
Also seed unsupported superlatives, sarcasm, quoted competitor claims, ambiguous pronouns, future plans stated conditionally, copied email history, corrected dates, stale cached pages, inaccessible sources, and missing provenance. Expected behavior may be exclusion, explicit uncertainty, conflict display, or abstention. Penalize a confident unsupported claim more heavily than an omitted optional fact.
NIST’s Generative AI Profile discusses confabulation and validity/reliability risks. It does not certify a sales product. Use it to justify adversarial evaluation, documented limits, human review, and incident controls—not a claim that a model is accurate.
Calculate precision, recall, usable coverage, and critical errors
- Fact precision = correct included atomic facts ÷ all included atomic facts.
- Fact recall = correct included atomic facts ÷ eligible expected facts.
- Source precision = included citations that support the adjacent claim ÷ included citations.
- Freshness pass rate = eligible time-sensitive facts inside their field rule ÷ eligible time-sensitive facts.
- Identity edge accuracy = correct person/account/opportunity edges ÷ eligible expected edges.
- Usable coverage = briefs passing all required identity, evidence and permission fields ÷ eligible briefs.
- Critical-error rate = briefs with at least one predefined critical error ÷ eligible briefs.
- Abstention precision = correct abstentions ÷ all abstentions; also report missed required abstentions.
Publish raw counts and results by call type, source availability, region, role, ambiguity and freshness band. Precision can be inflated by outputting one easy fact; recall can be inflated by including noisy claims. Usable coverage requires both. Do not combine identity, privacy and ordinary typo errors into one flattering score.
Apply hard gates to the worked example
Fictional example—not a Gangly or vendor result: 100 eligible briefs contain 800 expected facts and 250 expected identity edges. The system includes 760 facts, 722 correct; creates 252 edges, 244 correct; and 84 briefs pass every required field. It produces one private-note disclosure and two wrong-opportunity associations classified critical.
- Fact precision = 722 ÷ 760 = 95.0%.
- Fact recall = 722 ÷ 800 = 90.25%.
- Identity edge precision = 244 ÷ 252 = 96.83%.
- Identity edge recall = 244 ÷ 250 = 97.6%.
- Usable coverage = 84 ÷ 100 = 84.0%.
- Critical-error rate = 3 ÷ 100 = 3.0%.
If gates require zero unauthorized disclosure and zero wrong-opportunity critical associations, the run fails. Remove the unsafe retrieval path, repair association logic, regenerate a fresh blinded set, and retain both reports. Do not average away a hard-gate failure.
Run a canary, roll back, and calculate TCO
After offline gates pass, canary a consenting team with read-only or draft briefs. Compare to the prior process on the same eligible meeting mix. Track reviewed briefs, correction categories and minutes, source clicks, abstentions, late/missing briefs, permission incidents, identity corrections and rep-reported usefulness. Treat meetings or revenue as exploratory; the short pilot does not establish causality.
Keep the old prep path available. Rollback disables generation and delivery, revokes excess scopes, removes cached sensitive content under policy, preserves incident evidence, restores the previous workflow, corrects contaminated CRM or calendar associations, notifies owners, and verifies no queued brief is delivered. Retest after source, model, prompt, schema, permission or association changes.
Term TCO = licenses and usage + source/data access + calendar/CRM/mail/call connectors + security/privacy/legal review + labeling and adjudication + implementation + review/correction time + monitoring/incidents + storage/retention + support + overlap + export/deletion/exit. Use current quotes and observed pilot labor. Compare cost per usable, gate-passing brief—not cost per generated brief.
Print the pre-call brief QA checklist
☐ Freeze meeting population, sources, permissions, prompt/model, schema, cutoff and timezone
☐ Define required fields, fact/inference/question/unknown types, freshness and severity
☐ Preserve meeting, person, account, opportunity, source and transformation IDs
☐ Build calibration, blinded evaluation and untouched holdout sets
☐ Test reschedules, recurrence, forwarding, aliases, namesakes and multiple deals
☐ Score source support, as-of time, conflicts, polarity, corrections and inference labels
☐ Verify effective rep permissions and data minimisation
☐ Inject hostile instructions, stale pages, missing sources and ambiguous evidence
☐ Require abstention and quarantine when no safe answer exists
☐ Report precision, recall, usable coverage, critical errors and raw counts by stratum
☐ Apply hard gates before averages
☐ Canary read-only, rehearse rollback, calculate TCO and define retest triggers
The best brief is not the longest or most confident. It is attached to the right meeting, shows what is known and when, separates inference from evidence, respects access, makes uncertainty visible, and gives the rep a fast way to correct it before speaking with a buyer.