Task success measures whether a participant reaches a defined correct outcome; task time measures elapsed time under a stated start, end, and assistance policy. Both require realistic tasks, eligible participants, consistent setup, and interpretation with errors and qualitative evidence. Faster is not always better, and small studies should not be presented with population precision they cannot support.
Who this is for: UX researchers and product teams comparing the usability of a defined workflow across iterations or relevant user groups.
- Predefine correct, partial, failed, assisted, and abandoned outcomes for one realistic task.
- Set timing boundaries and report distributions and context instead of relying on an unexplained average.
- Interpret success and time with errors, confidence, strategy, accessibility, sample limits, and customer consequence.
Define the task and outcome
Write the customer goal, starting state, supplied information, environment, and correct end state. Separate technical completion from correct customer outcome. A participant may submit a transfer successfully but choose the wrong account, which should not count as success merely because the system accepted it.
Define full success, meaningful partial success, critical error, noncritical error, assistance, and abandonment before sessions. Decide how alternative valid paths count. Pilot definitions against observed behavior and revise ambiguity before formal comparison, not after seeing which version appears to perform better.
Measure time consistently
Choose the timing start and end from the user goal, such as from receiving the scenario to visible confirmation. State whether reading instructions, system delay, interruptions, and facilitator assistance are included. Use the same policy across conditions and preserve raw timing plus notes about exceptional events.
Task times are often uneven, so inspect individual values and distributions rather than relying only on a mean. Report medians or other summaries when appropriate to the analysis, with sample and context. Do not remove slow attempts automatically; they may reveal severe interaction barriers or legitimate user variation.
Design a credible comparison
Recruit participants relevant to the workflow and record experience, device, input method, and access needs that may affect interpretation. Keep task wording, data, environment, facilitation, and assistance policy consistent. Counterbalance order where learning or fatigue could favor one design condition.
Plan the analysis according to the decision and study design. A small formative test can reveal mechanisms and major failures but may not estimate a stable success rate. Larger benchmark work needs justified sample planning and uncertainty reporting. Avoid universal targets detached from task consequence, baseline, and business context.
Interpret behavior responsibly
Read success and time beside wrong turns, errors, requests for help, confidence, and observed strategy. A fast task can reflect accidental action, skipped review, or premature abandonment. A slower task can be appropriate when people deliberately verify a consequential decision and avoid costly mistakes.
Compare with a baseline or alternative and state practical consequence, not only statistical output. Segment only where planned or clearly exploratory, keeping counts visible. Document instrumentation, exclusions, and limits. Use findings to choose a design action and later verify the complete workflow rather than optimizing one number in isolation.
Compare a hypothetical transfer review
Hypothetical study compares two confirmation designs for moving funds between sample accounts.
- Define success as selecting the intended hypothetical source, destination, and amount, then confirming the correct review summary.
- Start timing at scenario display and stop at confirmation, with system delay and assistance recorded separately.
- Counterbalance designs and capture critical wrong-account errors, reversals, help, and verification behavior.
- Interpret hypothetical time differences with accuracy and strategy, avoiding a claim that the fastest design is automatically best.
Task metric protocol
Define measurement and interpretation before observing comparative results.
- Decision, participant profile, task goal, starting state, supplied data, environment, and correct outcome.
- Full success, partial success, critical error, noncritical error, assistance, abandonment, and alternative path.
- Timer start, timer end, system delay, interruption, outlier policy, raw record, and summary choice.
- Study design, order, counterbalancing, sample rationale, baseline, uncertainty, segments, and limitations.
- Qualitative observations, confidence, customer consequence, guardrails, decision rule, and follow-up check.
Common mistakes
- Counting system submission as success even when the participant completed the wrong customer outcome.
- Comparing average times from inconsistent starting points, assistance policies, or participant populations.
- Declaring the fastest design superior without examining errors, accidental actions, and task consequence.
Try one
Design B is slower than Design A but has fewer critical mistakes in a high-consequence review task. Which should ship?
A strong answer refuses to decide from time alone. It examines predefined success and error criteria, practical magnitude, uncertainty, participant strategies, and the cost of delay versus a critical mistake. The team may prefer B, revise it for clarity, or gather more evidence. The answer should state sample limits and avoid inventing a universal acceptable time.
Sources
- Nielsen Norman Group usability metrics guideGuidance on selecting behavioral and attitudinal measures for a defined usability question.
- Nielsen Norman Group task-time guidanceGuidance on interpreting task time with distributions, task definitions, and user context.