Module 6: Listen and Measure (Listen)
Numbers that speak: complementing the qualitative story
Estimated time: 2.5 - 3 hours
In the Listen phase of the cycle you stop looking at a single session and start listening at scale: surveys, analytics, A/B tests. It doesn't replace the qualitative work of the problem space (M4) or the evaluation in Module 5; it complements them. Here you learn to measure the "how much" and "how often" to go with the "why" you already know how to surface.
Table of Contents
- Introduction: Why do numbers matter in UX Research?
- Surveys vs. Interviews: When to use each method
- Designing unbiased questions
- Usability metrics: putting a number on what you evaluated
- Fundamentals of A/B Testing and Analytics
- Mixed Methods: Combining Qual + Quant
- Practical exercise
- References and resources
Learning Objectives
At the end of this module, you will be able to:
- Distinguish when to use surveys versus interviews according to research objectives.
- Design survey questions minimizing cognitive and cultural biases.
- Measure a task-based test with the four standard metrics, and know whether the number you got is good or bad.
- Understand the fundamentals of A/B testing, heatmaps, and conversion funnels.
- Apply mixed methodologies (Mixed Methods) to obtain more robust insights.
1. Introduction: Why do numbers matter?
By training and personal preference, I have a bias toward qualitative methods: moderated testing, interviews, direct observation.
But over time (and several projects where I had to defend my findings to stakeholders), I learned something crucial: quantitative data doesn't replace qualitative data, it complements it. And when you combine them well, you have a much more powerful story.
"Conversations with users tell you the WHY. Quantitative data tells you the HOW MUCH and HOW OFTEN. Together, they tell the complete story."
As Adaptive Path says in their Experience Mapping guide: quantitative data can help validate what you learn in qualitative studies, prioritize the focus of your interviews, and make stakeholders feel more comfortable with a larger sample size.
The key? Knowing when to use each approach and how to combine them strategically.
2. Surveys vs. Interviews: When to use each method
This is probably the question I get most from junior researchers: Do I do interviews or a survey? The short answer is: it depends. The long answer... here we go.
2.1 Interviews: The power of "why"
Interviews are your tool when you need depth. When you want to understand motivations, emotions, context, frustrations... all that internal world of the user that doesn't appear in metrics.
Use interviews when:
- You're exploring a new problem and don't know what questions to ask
- You need to understand the "why" behind a behavior
- The topic is sensitive and requires rapport (e.g., personal finances, health)
- You want to capture user stories and narratives
- Your population is small or hard to reach (e.g., B2B, experts)
2.2 Surveys: The power of "how many"
Surveys shine when you need scale. When you already have clear hypotheses and want to validate them with a larger sample, or when you need data you can generalize.
Use surveys when:
- You already understand the problem and want to quantify its magnitude
- You need representative data from a large population
- You want to prioritize features or pain points
- Stakeholders need numbers to make decisions
- You have limited budget or time for research
2.3 Comparison table
| Aspect | Interviews | Surveys |
|---|---|---|
| Objective | Explore, understand depth | Validate, quantify, generalize |
| Sample size | 5-15 users (per segment) | 30+ for basic statistics; 100+ for robust analysis |
| Type of data | Qualitative (narratives, quotes) | Quantitative (numbers, %) |
| Flexibility | High (you can explore new topics) | Low (predefined questions) |
| Cost/time | High per participant | Low per participant |
| Analysis | Thematic coding, synthesis | Descriptive, inferential statistics |
"In my experience, the most common mistake is jumping straight to surveys without having done interviews first. You end up asking the wrong questions to many people."
3. Designing Unbiased Questions
This is where many projects go wrong (sorry for the language, but it's true). A survey with poorly designed questions gives you worthless data. And the worst part is you don't realize it until it's too late.
3.1 The respondent's cognitive process
Before writing questions, you need to understand how the brain processes a survey. According to Tourangeau, Rips, and Rasinski, there are five stages the respondent goes through:
- Perception: "What is this?" - They see or hear the stimulus
- Comprehension: "What are they asking me?" - They interpret the question
- Retrieval: "What do I know about this?" - They search their memory
- Judgment: "What is my answer?" - They formulate an opinion
- Response: "How do I report it?" - They translate to the survey format
Important: The order of questions activates information in memory that affects subsequent responses. Priming is real!
3.2 Closed vs. open questions
Closed questions (with predefined options):
- Easier to analyze
- Lower cognitive effort for the respondent
- Risk: you may omit important options or force responses
Open questions (free response):
- Capture perspectives you didn't anticipate
- More difficult to analyze (require coding)
- Lower response rate (more effort)
3.3 Common biases and how to avoid them
Acquiescence bias
The tendency to agree with everything. Solution: alternate the direction of questions.
❌ Bad: "Do you agree that our product is easy to use?"
✅ Better: "How easy or difficult is it for you to use our product?" (balanced scale)
Social desirability bias
Responding with what they believe is socially "correct." Common in sensitive topics.
❌ Bad: "How often do you exercise?" (everyone will exaggerate)
✅ Better: "In a typical week, how many days do you do at least 30 minutes of physical activity?"
Double-barreled questions
Asking two things in one. The respondent doesn't know what to answer.
❌ Bad: "How satisfied are you with the price and quality of the product?"
✅ Better: Separate into two distinct questions.
Primacy/recency bias
In long lists, options at the beginning and end are selected more. Solution: rotate options or use visual scales.
3.4 Cultural considerations for Latin America
This topic is crucial if you work in our region. Research is not simply translating materials from English.
- Tú vs. Usted: Affects the tone and comfort of the participant
- Questions about income: May be considered intrusive in Mexico and other countries
- Alternative for socioeconomic level: Ask number of light bulbs/lights in home or access to services
- Idioms: Spanish from Chile ≠ Argentina ≠ Mexico
"A perfectly calculated sample is useless if data quality is poor due to lack of cultural sensitivity."
4. Usability metrics: putting a number on what you evaluated
In Module 5 you evaluated an interface and found problems. This section is the other side of that coin: putting a number on what you found, so you can compare it against a previous version, a competitor, or a target.
It is the same distinction you already saw between interviews and surveys, but inside a usability test. A formative test looks for what breaks; a quantitative one measures how much it breaks. They do not compete: in practice the order that works is qualitative first, to learn what to measure, and quantitative afterwards, to size it.
4.1 The four metrics of a task-based test
The ISO 9241 standard defines usability as the intersection of effectiveness, efficiency and satisfaction. In practice that translates into four metrics:
- Task completion rate. What share of participants finish the task. It is coded binary: 1 if they completed it, 0 if not. Sauro recommends avoiding partial percentages (25%, 50%, 75%) because they introduce ambiguity exactly where you need clarity.
- Time on task. How long the ones who finish take. Watch the stopwatch: start it when the person orients toward the screen and begins to act, not while they read the scenario. The 20 seconds someone spends reading your prompt are not part of your product's usability.
- Errors. Unintended actions: the wrong field, the skipped step, the forgotten checkbox. Worth recording even when the task is completed, because they explain the long times: they correlate strongly with time and account for around 25% of the differences between participants (Sauro & Lewis, 2009, as cited in Sauro, 2010).
- Satisfaction. What the person reports, after each task and at the end of the test. See 4.4.
"The success criterion is defined before testing, not while watching the recording. If you do not know what counts as completing the task, any result can be argued after the fact."
4.2 Is your number good or bad?
A 72% completion rate means nothing on its own. Sauro published percentile distributions from his database of usability tests, and they work as a reference point when you have neither your own data nor a measured competitor:
| Metric | Poor (low percentile) | Median | Good (high percentile) |
|---|---|---|---|
| Completion rate | 42% (p20) | 78% (p50) | 97% (p80) |
| Errors per task | 1.6 (p20) | 0.66 (p50) | 0.18 (p80) |
| SUS | 53 (p25) | 66 (p50) | 80 (p75) |
The percentiles do not line up exactly across rows because each metric comes from a different dataset: completion rate from 1,200 tasks across more than 120 tests, and errors from 719 tasks (Sauro, 2010).
About the SUS 66: that is the median of the 2010 dataset. The figure you will see quoted today is 68, from a later and larger database of 500 studies (MeasuringU), and it is the one our own calculator uses. The gap between the two is smaller than the standard deviation of SUS scores (around 19 points): either one works as a reference line for reading your result.
How to read it: if your task has a 90% completion rate, it sits at the 70th percentile — better than 70% of the tasks in that database. Below 56% it drops to the 30th percentile. And more than 2.4 errors per task is, in Sauro's words, "both unusable and unusual".
The caveat that matters: these numbers are a starting point, not a target. What counts as acceptable depends on the cost of failure. If getting it wrong means losing money or abandoning the process, you aim for 100%; if it is an application people use daily, 70% is simply unacceptable even though it sits above the database median.
4.3 Task times are deceptive
Three things almost nobody tells you about time:
- The distribution is skewed. A few participants get stuck and stretch the tail to the right, so the plain average is inflated by them. With real data it is better to transform to a logarithmic scale and report the geometric mean, which represents the typical participant more faithfully.
- Times do not compare across tasks. Thirty seconds can be blazing fast for one task and an eternity for another. Time only means something against a reference point.
- That reference point has to be declared. Sauro calls them specification limits, and they can come from: the previous version of the product, a competitor, an industry standard, or a multiple of an expert user's time. Any of them beats none.
4.4 Satisfaction: two questions at two moments
After each task, a single question is enough. Literally enough: one 7-point question ("Overall, this task was… very easy / very difficult") performs as well as the three-question questionnaire that used to be standard (Sauro & Dumas, 2009, as cited in Sauro, 2010). It costs seconds and tells you which task to focus the analysis on.
At the end of the test, the SUS (System Usability Scale), ten statements producing a score from 0 to 100. It is not a percentage: a SUS of 68 does not mean "68% usability", it means the average of 500 studies. Above 68 you are above average, below 68 you are under it.
SUS has been around since 1986 and was not designed for websites, so there is some debate about how well it measures them. It stays in use because it is short, free, and there is something to compare against.
One figure worth having at hand when someone tells you "but users liked it": around 14% of people rate satisfaction highly after failing the task. Reported satisfaction and observed performance are not the same thing, and that is exactly why both are measured.
4.5 The single score, and why to distrust it
Sooner or later someone will ask you for "one number" for usability. It exists: it is called SUM (Single Usability Metric) and averages the four metrics after standardizing them onto a common scale.
Sauro puts it well: an aggregate score is to the set of metrics what an abstract is to the full paper. It is useful for communicating and for comparing products, but it does not replace the metrics underneath. If you hand over only the single score, you have taken away the one actionable thing you had: knowing which of the four is broken.
4.6 How many participants?
It depends on which of the two questions you are asking, and the difference is large:
- To find problems (formative), the probability of seeing at least once a problem affecting a proportion
pof users acrossnsessions is1 - (1 - p)^n. For frequent problems 5 to 8 sessions are enough; for a problem affecting 10% of users, five sessions give you only a 41% chance of seeing it. - To measure (quantitative), the usual floor is 40 participants, which is what holds a margin of error near ±15 points on a binary metric such as completion rate.
You can work out both cases with the site's Sample Calculator, which also shows you what margin of error each sample size buys.
5. Fundamentals of A/B Testing and Analytics
OK, now let's enter the world of behavioral data. This is the part where many UX Researchers feel out of their comfort zone, but it's more accessible than you think. You don't need to be a data scientist to understand these concepts :)
5.1 What is A/B Testing?
A/B testing (or split testing) is an optimization methodology that compares two or more versions of a digital element to determine which performs better according to predefined metrics.
Classic example:
- Version A (control): Blue button that says "Sign up"
- Version B (variant): Green button that says "Start free"
- Metric: Conversion rate (% of users who click)
Key concepts:
- Baseline Conversion Rate: Your current metric (e.g., 3% of visitors sign up)
- MDE (Minimum Detectable Effect): The minimum change you want to detect (e.g., +10% relative)
- Statistical Power: Probability of detecting a real effect (typically 80%)
- Confidence Level: How sure you are of the result (typically 95%)
5.2 Analytics Tools
Heatmaps
Visualizations that show where users click and how they interact with the page. Warm colors (red, orange) indicate higher activity.
Types of heatmaps:
- Click maps: Where they click
- Scroll maps: How far down the page they scroll
- Move maps: Where they move the cursor (visual attention)
Tools: Microsoft Clarity (free), Hotjar, Crazy Egg
Clickstreams
The sequential journey users take through your site. Shows you the most common paths and where they abandon.
Conversion funnels
Funnels that show the % of users who complete each step of a process. They identify leak points where users are lost.
E-commerce funnel example:
| Stage | Users | Rate |
|---|---|---|
| Visit product page | 10,000 | 100% |
| Add to cart | 2,000 | 20% |
| Start checkout | 800 | 8% |
| Complete purchase | 300 | 3% |
In this example, the biggest leak point is between "Visit product" and "Add to cart." Why? That's where you need qualitative research to understand the "why."
5.3 Limitations of A/B Testing
Before you run off to A/B test everything, an important disclaimer:
- You need sufficient traffic: Without volume, there's no statistical significance
- It tells you WHAT works, not WHY it works
- Only optimizes within the current paradigm (doesn't find radically new solutions)
- The statistical winner may not be the experience winner
6. Mixed Methods: Combining Qual + Quant
We've arrived at what I consider the highest level of UX Research: knowing how to strategically combine qualitative and quantitative methodologies.
6.1 Why mix methods?
As I mentioned at the beginning, qualitative and quantitative methods answer different questions:
- Qualitative: Why? How? What is the experience?
- Quantitative: How many? How often? How big?
When you combine them, you get triangulation: multiple viewpoints on the same phenomenon, which increases the validity of your findings.
6.2 Mixed Methods design patterns
Pattern 1: Exploration → Validation (Sequential)
Flow: Interviews → Survey
First you explore the problem with interviews (discover themes, hypotheses, user language). Then you design a survey to validate and quantify those findings.
Example: "In interviews we discovered that 6/8 users mentioned frustration with the payment process. We designed a survey to validate with n=500 and confirm that 72% have this problem."
Pattern 2: Quantification → Deepening (Sequential)
Flow: Analytics/Survey → Interviews
First you identify patterns in data (e.g., 80% abandon at step 3 of the funnel). Then you do interviews to understand why.
Example: "Data shows that young users convert 3x more than older users. We do interviews with both groups to understand the difference."
Pattern 3: Convergent (Parallel)
Flow: Interviews + Survey simultaneously
You collect both types of data at the same time and compare/contrast results at the end.
6.3 Practical example: Experience Map project
Following the Adaptive Path guide, this is what a real Experience Mapping project with Mixed Methods would look like:
- Desk Research: Review existing analytics, previous research, NPS data
- Exploratory survey: Identify segments and prioritize touchpoints
- In-depth interviews: 8-12 users per segment, using directed storytelling
- Synthesis: Create journey map with quali + quanti data
- Validation: Survey to confirm main pain points
"Customer conversations and observations are your primary tool to learn, identify patterns, and capture the richness of human experience. But quantitative data validates and prioritizes." - Adaptive Path
7. Practical Exercise
Situation: Your team is redesigning the onboarding flow of a personal finance app. Analytics data shows that 65% of users abandon before completing registration.
Part 1: Plan your Mixed Methods
- What quantitative data do you need to review first? (hint: funnels, heatmaps)
- What hypotheses would you generate from that data?
- How would you design the interviews to deepen?
- What follow-up survey would you propose?
Part 2: Design 5 survey questions
Write 5 questions for a post-abandonment survey. Include at least:
- 2 closed questions (scale or multiple choice)
- 1 open question
- Avoid the biases discussed in section 3
Part 3: Define your metrics
The team wants to run a usability test of the new onboarding before shipping it. Define:
- What exactly counts as "completing registration" (your success criterion, written before testing)
- Which of the four metrics you will collect, and why not the others
- What you will compare time on task against (your specification limit)
- Whether it is a formative or a quantitative test, and how many participants that decision implies
Part 4: Propose an A/B test
Based on a hypothesis about abandonment, design an A/B test. Define:
- Variable to test (what you change)
- Success metric (what you measure)
- How you would interpret the results
8. References and Resources
Recommended readings
- Adaptive Path. (n.d.). Guide to Experience Mapping.
- Sauro, J. (2010). A Practical Guide to Measuring Usability: 72 Answers to the Most Common Questions about Quantifying the Usability of Websites and Software. Measuring Usability LLC. ISBN 1453806563.
- MeasuringU. Measuring Usability with the System Usability Scale (SUS). https://measuringu.com/sus/
- Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The Psychology of Survey Response. Cambridge University Press.
- MeasuringU. Schools of Thought on Sample Sizes in UX Research. https://measuringu.com/schools-of-thought-on-ux-sample-sizes/
- Maze. What is UX Research, Why it Matters, and Key Methods. https://maze.co/guides/ux-research/
Tools mentioned
- Microsoft Clarity (free heatmaps): https://clarity.microsoft.com/
- Google Analytics: https://analytics.google.com/
- UXR sample calculator: https://uxr.cl/en/learn/tools/sample-calculator/
Resources in Spanish
- Justinmind. ¿Qué es la investigación UX? Métodos y buenas prácticas. https://www.justinmind.com/es/ux-diseno/ux-investigacion
- ATLAS.ti. Dominio de las entrevistas semiestructuradas. https://atlasti.com/es/research-hub/entrevistas-semiestructuradas
- Design Toolkit UOC. Diarios de usuario. https://design-toolkit.recursos.uoc.edu/es/diarios-de-usuario/
Measurement and behavioral data
You can also explore the Sample Calculator we've developed to help you determine sample sizes for surveys, usability testing and A/B testing in the Latin American context.
See you in the next module! :)