Testing conversational AI with real customers works differently from testing a website. A website shows every person the same screen. An AI assistant writes a new answer every time. Some of those answers will be wrong, off-brand, or outside what the business can stand behind.
This guide covers what to watch, how to set up sessions, what to measure, and a checklist. It is for product teams preparing a launch and investors judging an AI feature.
Why testing conversational AI is different
A standard usability test checks whether people can find and use things. With an AI assistant, you also check what it says.
Answers vary. The same question can get a different answer on a second try.
Wrong answers sound right. Nielsen Norman Group describes AI hallucinations as output that seems plausible but is incorrect, delivered with confidence.
The assistant's errors belong to the company. In Moffatt v. Air Canada (2024), a British Columbia tribunal held the airline liable after its website chatbot gave a customer wrong information about bereavement fares. The tribunal said the chatbot was still just a part of Air Canada's website.
Better technology reduces the risk without removing it. Stanford RegLab researchers tested legal research tools that use retrieval-augmented generation. Two of the tools hallucinated more than 17% of the time, and a third more than 34% of the time.
What to look for when you test an AI assistant
What to watch | What good looks like | Warning sign |
|---|---|---|
Answer accuracy | Matches current policy, pricing, and product facts | A confident answer the business can't support |
Trust calibration | People check high-stakes answers and act on simple ones | People act on a wrong answer without checking |
Recovery | The person spots a problem, rephrases, and the assistant corrects itself | The person loops, gives up, or leaves with the wrong answer |
Handoff to a person | A clear route to a human, with the conversation carried over | The person can't find a human or has to repeat everything |
Guardrails and refusals | Declines out-of-scope requests plainly and offers a next step | Refuses safe questions, or answers ones it should decline |
Tone | Short, direct, and on-brand | Small talk, flattery, or long blocks of text |
Trust calibration deserves extra attention. NN/g's Pavel Samsonov notes that confident tone and polished formatting discourage people from checking AI output, and that users often trust it enough to skip reading it end to end. Note whether participants verify an answer before acting on it.
Tone is easy to get wrong. In an NN/g study of 9 participants using 8 site chatbots, people typed short, search-style queries, skipped pleasantries, and pushed back on flattering lines like "great question." The same research found people will not experiment to discover what a chatbot is for. If the value isn't obvious right away, they move on.
How to set up sessions with real customers
Recruit the right people. Plan 5 to 8 participants per customer group, per round. Our guide on how many user interviews you need explains why this sample size holds up. A good screener keeps out people who don't match the group.
Gather real prompts first. Before sessions, pull a sample of real questions from support tickets, site search logs, and call notes. Ask each participant to bring one question they actually have.
Write goal-based tasks. Give people a goal and let them use their own words. A hypothetical health plan task might read: "You need urgent care this weekend and the closest clinic may be out of network. Find out what you would pay." Avoid scripting the exact question.
Include the hard cases on purpose. Add at least one task where the right response is a refusal, and one that should end with a person. Add one where the honest answer is "I don't know."
Repeat critical prompts. Outside the sessions, run each high-stakes question several times and compare the answers.
Consider a Wizard of Oz test. If the assistant isn't built yet, a team member can write the replies behind the scenes while the participant uses a realistic chat screen. NN/g recommends this method for conversational interfaces and for testing ideas before investing in costly technology like generative AI. Give the "wizard" tone-of-voice guidelines and a list of answers the business can support.
What to measure
Measure | How to capture it | How to report it |
|---|---|---|
Task success | Did the person end with a correct, usable answer? Check against your source of truth. | Counts, such as "5 of 8 completed" |
Answer accuracy | Grade each answer as correct, partly correct, or wrong | Counts per task |
Unsupported answers | Flag any answer that conflicts with policy, pricing, or legal terms | Every instance, with the transcript |
Recovery | After a wrong answer, did the person notice and get back on track? | Counts and quotes |
Handoff | Number of turns to reach a person, and whether context carried over | Turns per task |
Trust | After key answers, ask "Would you act on this?" Compare with accuracy. | Count of mismatches |
Report small samples as counts. Keep full transcripts for every flagged answer. Our guide to writing a research readout shows how to turn these findings into decisions.
A checklist for testing conversational AI
[ ] Real customer questions gathered from support, search, and sales
[ ] 5 to 8 participants per customer group, screened
[ ] Goal-based tasks written in plain language
[ ] At least one task each for a refusal, a handoff, and an "I don't know"
[ ] Critical prompts run several times to check for variation
[ ] Handoff tested end to end, including what the human agent sees
[ ] Transcripts saved for every wrong or unsupported answer
[ ] A second round planned to confirm fixes
What investors should ask about an AI product
Ask to see transcripts from sessions with real customers. Ask how often the assistant is wrong on the questions that matter most, and how the team knows. Ask how handoff works and who keeps the source content current. Weak answers point to support costs and legal exposure after close. Our guide to testing a growth plan before close covers how customer evidence fits into diligence.
Frequently asked questions
How many people do you need to test an AI assistant?
Plan 5 to 8 participants per customer group, per round. Because answers vary, also run each critical question several times outside the sessions.
What is Wizard of Oz testing for chatbots?
A person writes the assistant's replies behind the scenes while the participant uses a realistic chat screen. It lets you test the conversation before you build the model.
How do you test AI guardrails with users?
Include tasks that should trigger a refusal or a handoff. Watch whether the assistant declines plainly and offers a next step, and check that it still answers safe questions.
Should participants know they are talking to an AI?
Yes, in most tests, because that matches what customers will see at launch. In a Wizard of Oz test, explain the setup to participants afterward.
What is the most common mistake in chatbot user testing?
Testing only with prompts the team wrote. Real customers phrase questions differently and ask things the team never planned for.
Get real evidence before your AI assistant goes live
Habib Innovation Partners tests AI assistants with the customers who will use them. Through the Innovation Retainer, we run small, fast rounds that show where your assistant helps, where it fails, and what to fix first. For investors evaluating a company's AI product, Investor Validation delivers fixed-fee customer research in 2 to 4 weeks. See how it works or start a conversation.
Sources
Page Laubheimer, "AI Hallucinations: What Designers Need to Know," Nielsen Norman Group, 2025
Pavel Samsonov, "AI Chatbots Discourage Error Checking," Nielsen Norman Group, 2025
Maria Rosala, Georgia Kenderova, and Tanner Kohler, "Less Chat, More Answer: Site AI Chatbots Need to Get to the Point," Nielsen Norman Group, 2026
Maria Rosala, Georgia Kenderova, and Tanner Kohler, "What Is Your Site's AI Chatbot for? Users Can't Tell," Nielsen Norman Group, 2026
Sara Paul and Maria Rosala, "The Wizard of Oz Method in UX," Nielsen Norman Group, 2024
Stanford HAI, "AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries," Stanford University, 2024
Kirsten Thompson, "Airline ordered to compensate a B.C. man because its chatbot provided inaccurate information," Dentons Data, 2024 (on Moffatt v. Air Canada, 2024 BCCRT 149)
