- Plan
- Explore
- Concept
- Evaluate
- Launch
NASA-TLX
Overview
How much workload did a task actually impose on the person who just performed it?
NASA-TLX (NASA Task Load Index) is a subjective, multidimensional questionnaire for rating how much workload a task imposed on the person who just performed it — the workload-focused counterpart to Usability Evaluation Methods‘s satisfaction- and loyalty-focused standardized instruments (SUS, NPS). Developed at NASA Ames Research Center over a three-year cycle of more than 40 laboratory simulations, it has since been cited in over 4,400 studies and is widely treated as the standard subjective-workload measure across aviation, healthcare, and other complex, high-demand settings.
NASA-TLX (NASA Task Load Index) is a subjective, multidimensional questionnaire for rating how much workload a task imposed on the person who just performed it — the workload-focused counterpart to Usability Evaluation Methods‘s satisfaction- and loyalty-focused standardized instruments (SUS, NPS).
Rate workload on six subscales
Each of six subscales is rated on a 0–100 scale in 5-point steps, using a fixed descriptive question:
- Mental Demand — how much mental and perceptual activity the task required; was it easy or demanding, simple or complex?
- Physical Demand — how much physical activity the task required; was it easy or demanding, slack or strenuous?
- Temporal Demand — how much time pressure the pace of the task or its elements created; was the pace slow or rapid?
- (Own) Performance — how successful the participant felt they were at the task, and how satisfied they were with that performance.
- Effort — how hard the participant had to work, mentally and physically, to reach their level of performance.
- Frustration — how irritated, stressed, and annoyed versus content, relaxed, and complacent the participant felt during the task.
No numbers on the track itself, just two endpoints — the participant marks wherever their own sense of that dimension actually falls.
Weight the subscales before combining them
The full procedure has two parts. First, the participant rates all six subscales for the task just completed. Second — done once per task type, not on every administration — the participant works through all fifteen possible pairs of the six subscales and picks, for each pair, whichever one felt more relevant to the workload just experienced; how often a subscale wins across those fifteen pairings becomes its weight. The overall score sums each subscale’s (rating × weight), then divides by 15, giving one 0–100 workload figure that reflects which dimensions actually drove that participant’s experience rather than treating all six as equally important by default.
The overall score sums each subscale’s (rating × weight), then divides by 15, giving one 0–100 workload figure that reflects which dimensions actually drove that participant’s experience rather than treating all six as equally important by default.
Raw TLX, a widely-used shortcut, skips the pairwise-weighting step and simply averages the six unweighted subscale ratings — evidence suggests this shortened version can even increase experimental validity by removing an extra, error-prone judgment task — and lets an irrelevant subscale be dropped from a specific task rather than forcing a rating on it.
Administer it consistently
Available as the original paper-and-pencil form or as an iOS app; the app’s rating control (a continuous “Subjective Analogue Equivalent Rating” slider) is built to reproduce the paper form’s unanchored feel, unlike the discrete, locking scale steps common to unofficial computerized versions. Format isn’t neutral: one study found a paper form produced lower measured workload than the identical content on a screen, though other computer- and wearable-based versions still track relative workload changes reliably. When measuring the same task type repeatedly, only the six subscale ratings need to be redone on each repeat administration — the pairwise weighting step is reused unless the task itself changes in kind.
Apply it alongside behavioral metrics
Used the same way SUS is used in a formal usability study — as a subjective self-report layered on top of behavioral bottom-line data — but answering a different question: not whether the interface felt usable, but how much it cost the person using it. This matters most for an interface embedded in an inherently demanding task (a cockpit, a clinical charting system, a control room), where two designs might produce identical completion times and error rates while asking very different amounts of the person operating them — a difference NASA-TLX surfaces that a purely behavioral metric would miss. The construct it rates — workload on limited mental and physical resources — is the same one Cognitive Load describes theoretically; NASA-TLX puts a number on it directly, from the person who just experienced it, rather than inferring it from an interface’s design.
This matters most for an interface embedded in an inherently demanding task (a cockpit, a clinical charting system, a control room), where two designs might produce identical completion times and error rates while asking very different amounts of the person operating them — a difference NASA-TLX surfaces that a purely behavioral metric would miss.
Related Concepts
Principles
Processes
Sources
NASA Task Load Index (TLX) is the source for NASA-TLX’s origin, its six subscale names, its application settings, and its open-source paper/iOS-app availability.
NASA-TLX (Wikipedia) is the source for the exact subscale question wording and 0–100/5-point rating scale, the two-part pairwise-weighting scoring procedure and its divide-by-15 formula, the Raw TLX variant, and the paper-vs-screen administration format-effect finding.