An ICP can be strategically sensible and operationally unusable. “Mid-market SaaS with a modern stack and urgent need” is not testable until each criterion has a definition, source, as-of date, allowable inference, exclusion rule, owner, and missing-data treatment. An ICP data-quality audit measures whether accounts can be classified consistently from defensible evidence.
This is an audit method, not a claim that any particular ICP improves revenue. Primary sources were reviewed August 8, 2026. Results apply only to the frozen population, criteria version, sources, labeling policy and as-of date.
Keep ICP evidence quality distinct from definition and allocation
Use the ICP template to define the profile. Use account selection criteria to rank fit, value, access, signals and feasibility. Use prospecting-data audits to evaluate provider/contact records and territory planning to allocate accounts. This audit asks whether the account evidence supporting an existing ICP is reproducible and current.
Freeze ICP version, eligible universe, segments, countries, data sources, as-of date, scoring/exclusion logic, sample, annotators, CRM snapshot, planned corrections and review window. Name ICP business owner, data steward, source owner, CRM admin, analyst, privacy reviewer, adjudicator and change approver.
Turn ICP assumptions into a versioned criteria register
Create one row per criterion: ID, plain-language definition, field/API name, data type, allowed values/range, gating or scoring role, weight, positive rule, negative rule, unknown treatment, exclusions, evidence type, preferred/fallback source, as-of rule, freshness window, inference method, confidence, owner, approver, effective date and change reason.
Separate firmographic criteria such as industry, location, size and ownership; technographic criteria such as installed systems and version; operating criteria such as sales motion, team structure and regulatory environment; need criteria such as observable workflow/problem evidence; and exclusions such as unsupported region, incompatible model, conflict or prohibited category.
Write acceptance examples and counterexamples. “50–500 employees” needs the unit, source and as-of date. “Uses Salesforce” needs evidence that distinguishes current production use from an old job post. “Has need” must state which observable evidence qualifies and which remains a hypothesis.
Map facts, inferences, hypotheses, provenance, and as-of dates
Label every value fact when directly supported by an accepted source, inference when derived by a documented rule, or hypothesis when unverified. Preserve source URL/record ID, publisher, retrieval time, source effective date, quoted/structured evidence, transformation, confidence and reviewer.
The Census Bureau’s NAICS system is a versioned industry classification. Record the NAICS version and distinguish establishment from company classification; a code does not prove need. The SEC’s EDGAR APIs provide structured submissions and company facts for covered public issuers, but do not cover every private company and require filing/as-of context.
Define source precedence and conflict handling. Never overwrite conflicting credible sources with whichever value arrived last. Store both observations, their dates and adjudication. For inferred employee band, tech use or need, preserve inputs and rule version so another reviewer can reproduce the label.
Build a stratified account sample, not a winner-only sample
Sample the eligible universe across current ICP-in, ICP-out and unknown; size/industry/region bands; source providers; recent/stale records; customers, open opportunities, closed-lost, unworked and excluded accounts; high/low scores; and conflicts/missingness. Oversample rare exclusions and boundary values, then report both stratified results and population-weighted estimates.
Closed-won-only sampling creates survivorship and selection bias. Include accounts the team never pursued and accounts rejected for good reasons. Freeze stable account IDs and dedupe parent/subsidiary and domain aliases before sampling. Record inclusion probability and do not swap difficult accounts after labels begin.
Choose sample size from required segment visibility and feasible review, not a universal number. Report raw denominators and uncertainty. A segment with five accounts cannot support the same confidence as one with hundreds.
Calibrate independent reviewers and adjudicate disagreements
Two reviewers independently label every sampled criterion from the frozen evidence pack. They assign value, fact/inference/hypothesis, confidence, as-of date, freshness, conflict, usability and ICP include/exclude/unknown decision. Hide current CRM score and downstream sales outcome during first-pass labeling where practical.
Run a calibration batch, resolve policy ambiguity, then lock the codebook before the main set. Keep original labels, disagreements, adjudicated result and rule changes. Calculate percent agreement and a chance-adjusted statistic where appropriate, by criterion and segment. High overall agreement can hide an unusable need or exclusion field.
Reviewers must be allowed to abstain. Forced guesses inflate apparent coverage and create segment leakage. Treat systematic disagreement as a criterion-definition defect, not reviewer underperformance.
Measure completeness, accuracy, usable coverage, conflict, and staleness
- Completeness = populated required criterion cells ÷ eligible required cells.
- Field accuracy = correct tested values ÷ eligible tested values.
- Usable coverage = accounts with all gating fields accurate/current enough to decide ÷ eligible accounts.
- Conflict rate = criterion cells with unresolved credible conflicts ÷ tested cells.
- Staleness rate = cells older than criterion freshness window ÷ cells requiring freshness.
- Segment leakage = adjudicated out-of-segment accounts selected in ÷ all adjudicated out-of-segment accounts.
- Exclusion error = adjudicated eligible accounts excluded ÷ all adjudicated eligible accounts.
- Agreement = reviewer matches ÷ independently double-labeled decisions, reported with denominators.
Report by criterion, source, segment, region, account age and fact/inference/hypothesis. Do not average a critical exclusion into a broad score. Set hard gates for prohibited regions, explicit incompatibilities, privacy restrictions and other business-critical exclusions before observing results.
Test selection stability under source and time changes
Re-run classification with source A versus accepted source B; fresh versus prior snapshot; strict unknown versus neutral unknown; and reasonable threshold changes. Measure selection overlap, accounts entering/leaving, rank movement, segment mix and exclusion changes. A profile whose selected list changes radically with one defensible source choice is operationally fragile.
Document sensitivity by criterion. If employee count near a boundary drives most movement, introduce a review band rather than false precision. If inferred technology use changes frequently, shorten its freshness window or remove it as a hard gate. Preserve old versions so past routing decisions remain explainable.
Stage CRM corrections and prevent unsafe bulk writes
Do not write adjudicated labels directly into production. Stage stable account ID, old value, proposed value, source/as-of, evidence class, confidence, criterion/version, owner, reason and approval. Validate types, picklists, permissions, parent/child edges, ownership, automation and downstream routing in a sandbox or shadow fields.
Canary a bounded segment. Reconcile row counts, fields, edges, changed selections, assignments and exclusions; inject duplicate/replay, late source, merged account, changed owner, field rename, permission loss and rollback. Keep CRM authority explicit through the CRM data-quality guide. Rollback restores prior versioned values and selections without replaying assignments.
Compare descriptive outcomes without causal claims
After deployment, compare selected versus nonselected and old-version versus new-version accounts on descriptive coverage, contactability, outreach, replies, meetings, qualification, opportunities, win/loss and retention where available. Freeze definitions, observation windows and denominators. Report selection volume and missing outcome data.
These groups are not randomized. Rep skill, effort, channel, offer, timing, market, account familiarity and survivorship can drive differences. Write “the selected group showed X in this period,” not “the ICP caused X.” Use holdouts or stronger experimental designs only when operationally and ethically appropriate; do not withhold necessary compliance exclusions.
Govern review cadence, privacy, changes, and TCO
Review volatile sources more often than stable classifications. Trigger review on taxonomy/source changes, product/market shift, region launch, pricing/packaging change, material leakage/exclusion error, source outage or persistent reviewer disagreement. Version criteria, codebook, source policy, thresholds, approval, rollout and rollback. Add the accepted checks to the CRM hygiene playbook so stale or conflicting ICP fields re-enter an owned exception workflow.
Apply purpose limitation, minimization, access, retention and correction through the NIST Privacy Framework plus actual law/contracts/policy. Avoid sensitive personal data that is unnecessary for account fit. Federal information-quality guidance is a useful quality lens, not a commercial ICP standard.
TCO = source licenses/credits + enrichment + data engineering + CRM administration + labeling/adjudication + privacy/security/legal review + QA + remediation + monitoring + review cycles + change management + archive. Compare sources on accepted usable coverage and error cost, not raw records alone.
Print the ICP evidence register and audit checklist
| Control | Evidence | Status |
|---|---|---|
| ICP version, universe, sources, as-of and owners frozen | Audit charter | □ |
| Each criterion has definition, rule, source, freshness, exclusion | Criteria register | □ |
| Facts, inferences and hypotheses preserve provenance | Evidence table | □ |
| Stratified sample includes in/out/unknown and rare errors | Sample manifest | □ |
| Independent labels calibrated/adjudicated | Codebook/log | □ |
| Quality, agreement, leakage and exclusion metrics reported | Audit report | □ |
| Selection stability tested | Sensitivity matrix | □ |
| CRM changes staged, canaried, reconciled and reversible | Change pack | □ |
| Outcomes labeled descriptive; governance/TCO approved | Decision memo | □ |
Gangly may help operationalize a reviewed account workflow, but it does not make an ICP true. The accountable team must own criteria, evidence, exclusions, uncertainty, corrections and review.