Define the Clari backtest boundary
This page owns a Clari-specific procurement or renewal test. It does not replace the category-level sales forecasting tools guide, the reusable sales forecast template, or the general forecast accuracy operating guide. Its decision is narrower: does the exact Clari edition, configuration, hierarchy, forecast type, data connection, user process, and evaluation population improve a named forecasting job enough to justify renewal or purchase?
Write a one-page test charter before exporting data. Name the legal customer, Clari modules and edition, CRM and other sources, revenue motion, fiscal calendar, currency treatment, hierarchy, forecast types, categories, horizons, covered teams, exclusions, historical window, shadow period, decision owner, data owner, evaluator, and signed-quote term. Record whether outputs are algorithmic, rep submissions, manager calls, rollups, or scenarios. Never pool unlike outputs under one “Clari accuracy” percentage.
Clari’s official Forecast product page documents configurable forecasting, automated rollups, drill-downs, scenario modeling, and integrations. Those descriptions identify test cases; they do not establish performance in your environment. Ask the vendor to map each tested function to the contracted edition and live admin configuration.
Freeze the configuration and historical snapshots
Create a configuration manifest: version/export time, fiscal periods, hierarchy and ownership rules, included opportunity types, currency and exchange-rate date, forecast categories, rollup logic, fields and filters, probability inputs, model/output name, override permissions, integrations, refresh schedule, and any material change date. Hash or otherwise make the export tamper-evident.
For every forecast origin, preserve the forecast value, category or probability, opportunity membership, amount, expected close period, owner, hierarchy, submission and override values, timestamps, configuration version, CRM state, and any available model version. A current dashboard is not a historical snapshot. Rebuilding March’s forecast from August’s records introduces won/lost outcomes, changed amounts, pushed dates, deleted opportunities, and later activities that the March forecast could not have known.
Maintain a change log beside the snapshots. If a category mapping or inclusion rule changed mid-window, either evaluate versions separately or apply a documented common representation without importing future state. A renewal claim should say “configuration A on population B at horizon C,” not “the platform was accurate.”
Contract horizons, cohorts, categories, amounts, and actuals
Define the unit before computing error. A horizon is the distance between frozen origin and fiscal close—for example, 90, 60, 30, 14, and 7 days. A cohort may be region, segment, product, new versus expansion, subscription versus consumption, team, manager, currency, or forecast type. Report only cohorts large and stable enough to interpret, while retaining failed material cohorts rather than hiding them in a global total.
Specify what the number means: gross bookings, net bookings, annual contract value, recognized revenue, consumption, renewal amount, or another finance-controlled measure. Define cancellations, partial wins, down-sells, multi-currency conversion, split credit, reopenings, amendments, and transactions posted after close. The actuals source and freeze date belong in the contract.
For categorical evaluation, map every opportunity at origin to an allowed category, then map the realized outcome to won in-period, lost, slipped, reduced, expanded, or excluded under a prewritten rule. Keep amount accuracy and category classification separate. An accurate total created by offsetting over- and under-forecasts can still be operationally misleading.
Prevent leakage and reconcile actuals
Leakage control is the hard gate. Each forecast may use only data present at its origin. Ban current-state joins, retroactively corrected stages, later close dates, post-origin activities, outcomes, future exchange rates, and labels generated from downstream systems after the cutoff. Keep the evaluator blind to system identity when practical, and freeze all transformation code before unblinding results.
Reconcile the final actual from finance to CRM and Clari with a bridge: opening candidate amount + additions − removals ± amount changes ± currency adjustments = evaluated actual. Report unmatched records and dollars. Reconciliation rate = amount with a documented one-to-one or approved bridge ÷ total evaluated actual amount. Do not score accuracy until reconciliation meets its precommitted gate.
Also audit duplicates, deleted records, ownership changes, hierarchy moves, null amounts, dates outside the fiscal calendar, and late postings. The CRM data-quality framework covers permanent ownership; this test records how much missing or corrected source data changes the verdict.
Build matched naive, manual, and CRM baselines
Clari must beat an honest alternative, not zero. Freeze at least three baselines where available: a naive rule such as the recent comparable-period actual or unchanged current run rate; the contemporaneous rep/manager submission; and a CRM-native rollup using the same category, probability, or stage rules. State exactly which rule produced each number.
Match every baseline on origin, horizon, cohort, actual definition, opportunity set, exchange rates, and allowed information. Do not compare a week-one Clari output with a final-week manager call. Do not give the incumbent corrected history while withholding it from Clari. If a baseline is absent at an origin, mark the pair missing rather than substituting a later value.
Report paired deltas per origin and cohort. A global average can conceal a method that improves enterprise renewal forecasts while degrading new-business SMB forecasts. The procurement decision may be “use for one job, not all jobs.”
Calculate error, bias, calibration, and slippage
No single metric is “forecast accuracy.” Hyndman and Koehler’s peer-reviewed analysis of accuracy measures explains why scale and percentage measures have different limitations. Publish denominators and multiple measures:
- MAE = Σ|forecast − actual| ÷ number of origins. It is readable in revenue units but not comparable across differently scaled cohorts.
- WAPE = Σ|forecast − actual| ÷ Σ|actual|. It weights large periods heavily and is undefined when total actual is zero.
- sMAPE = mean[2|forecast − actual| ÷ (|actual| + |forecast|)]. State the convention for both-zero pairs and do not call it perfectly symmetric in interpretation.
- Signed bias = Σ(forecast − actual) ÷ Σ|actual|. Positive means aggregate over-forecasting under this sign convention; show currency bias too.
- Calibration: group stated close probabilities into frozen bins, then compare mean stated probability with observed win frequency. Publish bin counts; a small bin is not a reliable verdict.
- Category confusion: cross-tab origin category against won in-period, lost, slipped, or other realized class. Report row counts and amount-weighted results.
- Slippage rate = origin amount expected in-period that closes later or remains open after cutoff ÷ origin amount expected in-period. Keep loss and slip separate.
Show every result by horizon and critical cohort, with the same calculation for each baseline. Review the dedicated forecast bias guide when signed errors persist even if absolute error looks acceptable.
Run a rolling-origin backtest
Use successive historical origins in chronological order. At each origin, freeze the available configuration and records, generate or retrieve each method’s forecast, then score it only against the later frozen actual. Advance the origin and repeat. This mirrors the rolling-forecast-origin method described in Forecasting: Principles and Practice and implemented by the authors’ time-series cross-validation reference.
Do not randomly shuffle quarters. Preserve fiscal sequence, configuration changes, acquisitions, territory redesigns, seasonality, and macro shifts. Use both an expanding view, which shows all available origins, and a recent fixed window when old operating conditions are no longer comparable. Declare these views before results.
Keep the final eligible periods untouched while designing thresholds. If you repeatedly tune mappings and thresholds on the same periods, they become development data, not independent evidence. Version the evaluation code and rerun all methods together after an approved correction.
Handle missing or unusable history
If immutable Clari snapshots or contemporaneous baselines do not exist, a historical accuracy verdict is unavailable. Do not approximate them with current CRM state, screenshots without underlying membership, or memories of forecast calls. State which origins, fields, configurations, and overrides are missing and whether missingness favors a method.
Recover only exports that preserve the cutoff state and can be reconciled: archived forecast submissions, approved data-warehouse snapshots, finance packets, and time-stamped API extracts. Keep recovered origins separate from natively immutable snapshots and sensitivity-test exclusions.
Then move the decision to a prospective shadow period. Procurement can be conditional on export access, snapshot retention, and a future acceptance test. Missing history is itself an operational and exit-risk cost, not permission to manufacture precision.
Run a live shadow quarter
Operate Clari beside the incumbent without letting either output change labels, scope, or the other method’s inputs. Freeze forecasts at every agreed horizon, capture late refreshes and overrides, and reconcile actuals after finance close. Users should follow a documented submission cadence so adoption differences are visible rather than silently attributed to algorithms.
Inject controlled failures: CRM refresh delay, API outage, owner reassignment, hierarchy change, currency update, duplicate opportunity, amount reduction, close-date push, revoked permission, missing field, and late finance adjustment. Record detection time, visible status, retry behavior, recovery, lost history, and whether a manual override survives or is overwritten.
Run the weekly governance through the forecast review meeting framework, but keep evaluators from changing acceptance rules mid-quarter. Log every exception with owner, evidence, disposition, and retest date.
Set acceptance thresholds before seeing results
Finance, RevOps, and sales leadership should sign a gate sheet before unblinding. Set maximum WAPE, MAE in currency, absolute bias, critical category errors, amount slippage, reconciliation gap, refresh latency, unrecovered failure time, and protected-permission breaches by horizon and material cohort. Require a minimum paired improvement over the chosen incumbent or explain why equivalent accuracy still creates enough workflow value.
Separate hard gates from scored tradeoffs. Data leakage, unreconciled material actuals, missing audit history, prohibited access, and inability to reproduce a result should normally stop the verdict. Accuracy, workflow time, adoption evidence, administration, and cost can then be scored with declared weights.
A threshold is a buyer decision, not an industry benchmark. Record who approved it, why it is tolerable for the business, the sample behind it, and what change triggers retesting. Do not average away a failed regulated region, revenue type, or executive reporting horizon.
Govern failures, overrides, and retests
Every override needs original value, new value, actor, time, reason code, evidence, affected hierarchy, and whether the override enters later modeling or calibration. Score native and overridden outputs separately. Otherwise strong managers can make the platform appear accurate, or a useful platform can be blamed for an undisciplined manual call.
Create a failure register with severity, affected origins and cohorts, source system, detection, containment, correction, backfill decision, and recurrence test. Never silently edit a frozen forecast. If a defect invalidates an origin, preserve the original result, publish the exclusion rule, and show sensitivity with and without it.
Retest after material changes to edition, model/output, forecast categories, hierarchy, CRM mapping, fiscal calendar, revenue model, integrations, major territory structure, or acquisition. Renewal governance should require reproducible exports and documented exit rights, not merely dashboard access.
Apply the worked example and calculate TCO
This fictional example teaches the math; it is not a Clari result. Four quarter-end actuals are $100, $120, $80, and $100. The same-horizon Clari snapshots are $110, $100, $90, and $120. Absolute errors are $10, $20, $10, and $20, totaling $60.
| Measure | Calculation | Result |
|---|---|---|
| MAE | $60 ÷ 4 | $15 |
| WAPE | $60 ÷ $400 actual | 15% |
| Signed bias | ($20 net over-forecast) ÷ $400 | +5% |
| sMAPE | mean of 2|F−A|/(|A|+|F|) | 14.4% |
| Naive WAPE | $40 absolute error ÷ $400 for forecasts [100,100,100,100] | 10% |
Here, Clari fails an acceptance rule requiring WAPE below the matched naive baseline, even though 15% might sound acceptable without context. The evaluator must still inspect cohort, category, calibration, bias, reconciliation, and failure gates before deciding whether remediation can change the result.
Term TCO = contract + modules/seats + implementation + integrations + administration + snapshot storage/recovery + forecast-review labor + validation + remediation + security/procurement + renewal change + exit/migration. In a fictional term, a $72,000 contract plus $18,000 implementation/integration, $20,000 loaded administration/review, $8,000 validation/remediation, and $7,000 exit reserve yields $125,000 TCO. Replace all inputs with the signed quote and loaded buyer costs. Report TCO beside matched error improvement, critical gates, workflow evidence, and the configuration to which the verdict applies.
The defensible final sentence is conditional: “configuration X passed or failed gates Y on cohorts Z over these frozen origins and this shadow period.” Public documentation was reviewed for product-test scope; no hands-on Clari environment, proprietary model, customer dataset, or vendor accuracy result was independently verified.