A sales admin time study is a measurement protocol, not a universal benchmark. It estimates how a defined group allocates a defined denominator of time under a frozen activity taxonomy. A defensible report states who was eligible, how days were sampled, what counted as work, how simultaneous activities were handled, what was missing, and how uncertainty was calculated.
This page owns study design. The admin-time reduction guide owns interventions after baseline measurement. The rep time audit is a lightweight individual exercise; this protocol is for aggregate team research. Use sales productivity KPIs for the broader measurement system rather than relabeling time share as productivity.
Freeze the study question and denominator
Write the estimand before collecting data. “How much admin do reps do?” is incomplete. A useful question is: “Among eligible mid-market account executives employed by Entity A during four study weeks, what share of recorded working minutes on sampled weekdays fell into the frozen administrative taxonomy?”
Name one primary denominator: recorded work minutes, scheduled work minutes, or a fixed elapsed window. Do not mix them. For recorded work, admin share = eligible admin minutes ÷ eligible recorded work minutes. Predeclare treatment of breaks, leave, travel, after-hours work, simultaneous activity, and unclassified gaps. Research records must not replace legally required time or pay records.
Build a sample you can describe honestly
Define population, sampling frame, selection unit, observation unit, and analysis unit separately. People may be sampled, days observed, episodes labeled, and estimates reported per person-day. Calling 2,000 episodes a 2,000-person study is wrong.
| Design field | Freeze | Failure to avoid |
|---|---|---|
| Population | Entity, role, segment, geography, tenure, status, dates | Calling one convenience team “sales reps” |
| Frame | Eligible roster and exclusions before selection | Dropping low-activity reps after collection |
| Selection | Random, stratified, census, or voluntary method | Presenting volunteers as representative |
| Coverage | Weekdays, time zones, month/quarter end, travel, holidays | Sampling only quiet days |
| Response | Invited, consented, started, completed, usable counts | Hiding incomplete diaries |
Stratify on predeclared factors relevant to the decision. Preserve nonresponse and partial diaries. Compare observable respondent and nonrespondent characteristics where permitted, while acknowledging that this cannot eliminate nonresponse bias.
Define the activity taxonomy before collection
The taxonomy determines the estimate. Publish definitions, boundary examples, exclusions, and a version before labeling. The Bureau of Labor Statistics publishes American Time Use Survey methods, activity lexicons, and coding rules. A company study is not ATUS, but separating diary collection from documented coding is a useful pattern.
| Class | Example inclusion | Boundary |
|---|---|---|
| Buyer interaction | Live prospect/customer conversation | Waiting room and internal debrief |
| Prospecting execution | Calling or composing approved outreach | Research versus message creation |
| Opportunity administration | CRM entry, notes, corrections, approvals | Buyer-facing next-step email |
| Preparation/research | Account research and meeting preparation | General training |
| Internal coordination | Forecast, pipeline, deal desk, handoff | Buyer present for part of meeting |
| Other work | Training, hiring, travel, other duties | Work versus interruption |
| Uncodable/unrecorded | Insufficient or missing evidence | Never force a desired label |
For multitasking, either assign each minute to the reported primary activity and retain secondary activity, or use fractional allocation. Predeclare the rule and test sensitivity. Never count the same minute twice in an elapsed-time share.
Combine diaries with bounded instrumentation
Diaries capture intent and invisible work; instrumentation supplies timestamps and prompts. BLS describes ATUS activity files containing codes, start/stop times, and locations in its microdata documentation. Borrow episode-level structure, not its national sampling claims.
- Train participants with neutral examples.
- Collect start, stop, verbatim activity, primary/secondary activity, context, and confidence.
- Prompt near the event or use a consistent recall window; record collection lag.
- Import only approved calendar, meeting, CRM, telephony, or application events.
- Keep source events separate from labels and corrections.
- Allow identity and timestamp correction without silent deletion.
A calendar block may be canceled, an open app may be idle, and manual work may leave no event. Application foreground time is not ground truth. Retain provenance and a conflict state when diary and instrument evidence disagree.
Calibrate labels and adjudicate disagreement
Freeze a manual and blind coders to performance and desired conclusions. Have two reviewers independently label a stratified subset including short episodes, multitasking, sparse descriptions, and boundary categories.
Raw agreement = exact matches ÷ double-labeled eligible episodes. Publish a confusion matrix. If prevalence is imbalanced, a chance-corrected statistic such as Cohen's kappa can add context, but kappa depends on prevalence and label design. Agreement does not establish validity.
Use an independent adjudicator. Preserve both initial labels, the adjudicated label, reason, manual version, and timestamp. Recurring disagreements should trigger a versioned manual clarification and controlled relabeling.
Calculate time shares with explicit exclusions
Analyze at the unit promised. Pooling all minutes weights longer recorded days more. Averaging participant shares weights people equally but needs a sparse-diary rule. Publish a primary estimator and sensitivity estimator.
Keep an exclusion ledger: ineligible participant, out-of-window day, leave, duplicate, impossible duration, missing denominator, or withdrawn data. Report records and minutes removed at every step. Do not silently discard incomplete days because they move the result. CRM and sequence event counts belong to sales activity tracking; they become time-use evidence only through the study's frozen mapping and limitations.
Report uncertainty and sensitivity
Publish participant count, eligible/completed days, episode count, total minutes, missing minutes, response, and distribution—not one percentage. Use survey weights only when the design supports them. Label volunteer or convenience samples honestly.
Estimate uncertainty at the independent sampling unit. Minutes within one person are correlated; treating every minute as independent makes intervals too narrow. Depending on design, use participant-level bootstrap resampling, cluster-robust methods, or a transparent descriptive range reviewed by a statistician.
Test primary versus fractional multitasking, pooled versus participant-equal weighting, complete versus partial diaries, uncertain episodes included versus excluded, and alternate boundaries. If the conclusion changes, call it definition-sensitive.
Protect employee privacy and employment boundaries
Time-use data can reveal behavior, relationships, location, health, leave, and performance-adjacent information. The NIST Privacy Framework is voluntary and not employment law, but its privacy-risk perspective supports purpose limitation, minimization, access control, retention, and deletion.
Document purpose, lawful basis where applicable, notice/consent, worker consultation, required or voluntary status, fields, allowed/prohibited uses, access, retention, deletion, correction, withdrawal, complaint, and incident response. Aggregate outputs and predeclare small-cell suppression. Never promise anonymity if reidentification remains possible.
The Department of Labor's FLSA recordkeeping guidance describes accurate hours-worked and wage-record obligations for covered employers. This study is not a timekeeping substitute. HR, privacy, employment-law, labor-relations, security, and payroll owners must review each worker and jurisdiction.
Validate AI-assisted classification separately
AI may suggest labels but adds another measurement system. NIST describes its AI Risk Management Framework as voluntary guidance for managing AI risks. Using an AI model is not validation.
Create a frozen human-labeled test set stratified by role, category, duration, language, ambiguity, and sensitive content. Report per-category precision, recall, abstention, critical privacy errors, and confusion. Test hostile text, namesakes, truncation, private meetings, after-hours events, and changed taxonomy versions. Keep consequential employment decisions outside automated classification and retain human review and rollback.
Publish a reproducible study packet
A skeptical reader should reproduce definitions and arithmetic even when raw employee data stay restricted. Publish the question, decisions, dates, population, frame, recruitment, diary prompts, taxonomy, coding manual, instrumentation manifest, version hashes, participant flow, counts, missingness, exclusions, label checks, formulas, weighting, uncertainty, sensitivity, privacy controls, conflicts, and limitations. Share analysis code where safe.
Do not call a convenience result “the average sales rep.” Do not convert descriptive role or period differences into causal claims. After measurement, use a sales workflow audit to locate the process boundary, apply the operational reduction playbook in a controlled intervention, then rerun the unchanged protocol.
Sales admin time study checklist
☐ Freeze population, estimand, denominator, dates, and exclusions
☐ Document frame, selection, strata, response, and observation units
☐ Version taxonomy, examples, multitasking rule, and uncodable state
☐ Separate diary, instrumentation, labels, and corrections
☐ Double-label a blinded stratified subset
☐ Publish participant, day, episode, and minute denominators
☐ Calculate weighting, missingness, uncertainty, and sensitivity
☐ Minimize data, limit access, suppress small cells, retain briefly, delete
☐ Validate AI labels against frozen human truth
☐ Separate time allocation from productivity and causal claims