Define what the Lavender calibration test owns
This page owns one buyer decision: whether the exact Lavender configuration gives useful, repeatable coaching under an approved message-quality policy. It does not replace the AI sales email writer guide, the personalization guide, the outcome definitions in cold email metrics, or Gangly’s separate Outreach Grader.
Freeze tenant, plan, integration, browser/client, language, features, settings, score version when visible, test dates, user roles, and evaluation population. State whether you are testing score only, reason codes, rewrite suggestions, personalization assistance, manager reporting, or historical coaching. A result applies only to that configuration and population.
Lavender’s official Email Coach page documents real-time scoring, coaching, and a personalization assistant. It recommends 90 or above and links high score with replies. Treat that as vendor performance language, not independent proof. This test evaluates alignment with buyer-approved quality and risk; a separate experimental design would be needed to establish effects on replies.
Freeze a stratified message set and context packet
Sample actual approved-to-use drafts before showing them to Lavender. Stratify by first touch/follow-up/reply, cold/warm/customer, persona, seniority, industry, region, language, sender experience, channel origin, message length, personalization depth, proof type, CTA, mobile formatting, and required compliance treatment. Include strong, borderline, and deliberately unsafe cases.
Each message receives an immutable ID and an approved context packet: recipient role and relationship, account facts, triggering evidence, source URLs and dates, CRM state, prior thread, offer, proof library, prohibited claims, consent or lawful-use decision where applicable, suppression state, jurisdiction, sender identity, and desired next step. Remove unnecessary personal data and use synthetic recipients for dangerous examples.
Freeze the original draft, context, Lavender score/reasons, proposed edits, rater labels, timestamps, and versions separately. Never let a current score overwrite the original. Keep a holdout set untouched while raters and thresholds are refined.
Build a behaviorally anchored quality rubric
Define observable behaviors at each level instead of “good,” “clear,” or “personalized.” Use dimensions such as:
| Dimension | Pass anchor | Fail anchor | Critical error |
|---|---|---|---|
| Recipient/context fit | Reason and offer match approved role evidence | Generic or unsupported relevance | Wrong person/account or prohibited targeting |
| Claim truth | Every factual claim maps to approved evidence | Vague or unverifiable assertion | Fabricated result, identity, relationship, or sensitive inference |
| Clarity | Recipient can identify reason, value, and ask on one read | Ambiguous referent or buried purpose | Deceptive subject/header or materially misleading statement |
| Personalization | Specific, relevant, proportionate, sourceable | Token merge or irrelevant detail | Intrusive, sensitive, or disallowed personal data |
| CTA | One proportionate, answerable next step | Multiple competing or vague asks | Coercive or misrepresented obligation |
| Policy/compliance | Required identity, preference, suppression, and jurisdiction controls pass | Correctable formatting issue | Suppressed recipient or prohibited send |
Assign dimension weights only after owners explain the tradeoff. Define overall pass, borderline/manual review, and fail. Critical errors override the numeric total. Link narrow craft questions to the canonical body-copy guide and CTA guide rather than hiding subjective writing preferences inside calibration.
Train and blind independent human raters
Use at least two raters who understand the audience, evidence policy, and compliance boundary but did not write the messages. Train on a separate set, discuss disagreements, revise ambiguous anchors, then freeze the rubric. Do not train on the final holdout.
Hide Lavender score, suggestions, sender, outcome, and system identity during human rating. Randomize order. Have raters label every dimension and critical error independently, including evidence references and a concise reason. Adjudicate only after initial labels are locked, preserving both original judgments.
Repeat a hidden subset to check within-rater stability. If reviewers cannot apply the rubric consistently, the tool cannot be judged against it. Improve the reference before changing the product threshold.
Measure agreement without confusing it with validity
Start with the human reference. Exact agreement = messages receiving the same categorical label from both raters ÷ messages rated by both. For two categorical raters, Cohen’s κ = (observed agreement − chance-expected agreement) ÷ (1 − chance-expected agreement). For more raters or ordinal scores, choose an appropriate statistic with a qualified analyst and declare it before analysis.
The peer-reviewed inter-rater reliability explainer distinguishes percent agreement from kappa and warns that correlation can poorly reflect agreement. Kappa is also affected by label prevalence and marginal distributions. Publish the confusion matrix, counts, agreement, kappa, prevalence, and uncertainty; never turn a conventional kappa label into proof that the rubric is valid.
Correlation between Lavender’s numeric score and a human numeric total only shows co-movement. A system could correlate while remaining systematically ten points too generous or passing critical errors. Inspect absolute differences, rank reversals, threshold decisions, strata, and concrete disagreements.
Calibrate score thresholds and bins
Freeze meaningful Lavender score bins before unblinding—for example, buyer-chosen ranges with enough records—and report human pass, borderline, fail, and critical-error counts inside each. Human pass rate in bin = human-pass messages ÷ rated messages in that bin. Calibration means observed buyer-defined pass rates behave as expected for the operational use, not that the score predicts replies.
Compare candidate operating thresholds using sensitivity and specificity only if the buyer’s pass label is reliable: sensitivity = tool pass among human passes ÷ all human passes; specificity = tool fail among human fails ÷ all human fails. Also show positive predictive value, which changes with prevalence.
Choose threshold and review band from error costs. A team may allow automated low-risk style coaching but require human review for evidence, personalization, compliance, executive contacts, or sensitive contexts regardless of score.
Gate false passes, false fails, and critical claims
False-pass rate = messages Lavender passes that the human reference fails ÷ all human-fail messages. False-fail rate = messages Lavender fails that humans pass ÷ all human-pass messages. Also publish the operational false-discovery view: failed human messages among all tool-passed messages. Name the denominator every time.
Create a critical-error taxonomy: invented or unsupported claim; wrong recipient/account; deceptive subject or identity; sensitive inference; suppressed recipient; missing required preference mechanism; unsafe promise; fabricated relationship; wrong proof attribution; or privacy-policy violation. Report error count, message count, severity, coaching action, whether the suggestion introduced or removed it, and owner.
No weighted average offsets a critical error. Precommit zero-tolerance classes and escalation. Audit the score reason as well as the score: a passing number with an unsafe rationale can train reps toward the wrong behavior.
Test coaching edits, acceptance, and overrides
For each original draft, capture every suggested edit as a versioned diff. Label the targeted rubric dimension, whether the edit is accepted, modified, rejected, or unavailable, why, and the after-edit human label. Never infer “helpful” from acceptance alone; reps may accept quickly or reject a sound suggestion for contextual reasons.
Calculate safe acceptance rate = accepted suggestions that preserve or improve all hard gates ÷ reviewed suggestions. Track critical errors introduced, corrected, or missed; evidence removed; factual meaning changed; voice altered; and time to review. Separate mechanical edits from substantive claims.
Define override authority. Reps may override style preferences with a reason; compliance or fabricated-claim failures require an approved reviewer. Monitor repeated overrides by rule, rep, persona, and context to identify a poor score rule, weak training, or noncompliance.
Review privacy, access, and compliance
Inventory OAuth scopes, current-message access, optional historical access, recipient and account data, attachments, prompts, telemetry, model/provider access, subprocessors, locations, encryption, retention, training use, deletion, data-subject requests, admin roles, deprovisioning, audit logs, and incident terms. Verify live consent screens, admin console, security materials, DPA, and contract.
Lavender’s official privacy and security summary says it accesses only the email being worked on by default, with optional historic-email access for benchmarking and coaching, and says analyzed email is deleted. Its separate privacy policy supplies broader terms. These are vendor statements; test tenant scopes and obtain contractual answers.
Score coaching cannot approve a send. The FTC’s CAN-SPAM guide says US commercial email, including B2B, must use accurate headers, non-deceptive subjects, required disclosures/address and opt-out controls, and honor opt-outs. Other jurisdictions differ. Qualified owners must define applicable rules, consent, legitimate-use decisions, suppression, and retention.
Run a matched live pilot without causal overclaim
After offline gates pass, randomize or otherwise match eligible senders/messages within approved strata. Freeze eligibility, assignment, primary outcome, observation window, exclusions, sample plan, and stopping rules. Keep list, deliverability setup, sender reputation, timing, offer, sequence, and follow-up policy as comparable as practical.
Measure rubric pass, critical errors, edit acceptance, override, review time, sent volume, delivery, reply classification, meetings, opt-outs, complaints, and incidents. Blind outcome classification to treatment when possible. Analyze assignment groups, not only messages reps chose to score.
A difference in replies during a small operational pilot is an association unless the design, adherence, power, uncertainty, and interference support a causal conclusion. Do not claim the score predicts replies merely because high-scored messages had more replies; message quality, account fit, sender, signal, offer, and selection can cause both.
Apply the worked example and hard gates
This fictional example is not Lavender performance. Human adjudication passes 30 of 50 frozen messages and fails 20. Lavender passes 31: 27 human passes and four human fails. It fails 19: 16 human fails and three human passes. Two of the four false passes contain critical unsupported claims.
| Measure | Calculation | Result |
|---|---|---|
| Exact pass/fail agreement | (27 + 16) ÷ 50 | 86% |
| False-pass rate | 4 ÷ 20 human fails | 20% |
| False-fail rate | 3 ÷ 30 human passes | 10% |
| Tool-pass precision | 27 ÷ 31 tool passes | 87.1% |
| Critical false passes | 2 ÷ 50 messages | 4% |
Suppose hard gates require zero critical false passes, human-rater agreement above a precommitted floor, false-pass rate at or below 5%, complete privacy approval, and no prohibited send. The run fails. Do not raise the threshold after viewing this set and call it validated; tune on a development set, then test once on the untouched holdout.
Calculate TCO and write the decision packet
Term TCO = licenses + identity/integration + security/privacy/procurement + rubric design + rater training and labeling + administration + rep review + enablement + monitoring + incident/remediation + renewal change + exit. In a fictional term, $24,000 licenses plus $6,000 security/integration, $12,000 calibration labor, $10,000 administration/review, $5,000 training, and $3,000 monitoring/exit equals $60,000. Replace every input with the quote and loaded buyer labor.
The decision packet should state tenant/configuration, message strata, context contract, rubric/version, rater reliability, score distributions, bins, agreement, false-pass/false-fail, critical errors, edit results, privacy findings, live-pilot design and uncertainty, TCO, exceptions, owner, and retest trigger.
A defensible conclusion is bounded: “This configuration met our quality rubric and hard gates for these message types.” Public documentation was evaluated, but no Lavender tenant, proprietary scoring logic, independent reply-prediction evidence, or customer dataset was inspected. Rankings, reply improvement, and causal effects are not claimed.