Testing AI Assistants With Real Customers: What to Look For

Testing AI Assistants With Real Customers: What to Look For

Testing conversational AI with real customers reveals problems that internal demos miss: misplaced trust, wrong answers, weak recovery and awkward handoffs to people. Learn how to set up chatbot user testing sessions, what to measure, how to probe guardrails, and use our checklist for your next round.

Testing conversational AI with real customers reveals problems that internal demos miss: misplaced trust, wrong answers, weak recovery and awkward handoffs to people. Learn how to set up chatbot user testing sessions, what to measure, how to probe guardrails, and use our checklist for your next round.

Reba Habib

Reba Habib

var(--variable-r8pjYFD68)

Testing conversational AI with real customers works differently from testing a website. A website shows every person the same screen. An AI assistant writes a new answer every time. Some of those answers will be wrong, off-brand, or outside what the business can stand behind.

This guide covers what to watch, how to set up sessions, what to measure, and a checklist. It is for product teams preparing a launch and investors judging an AI feature.

Why testing conversational AI is different

A standard usability test checks whether people can find and use things. With an AI assistant, you also check what it says.

  • Answers vary. The same question can get a different answer on a second try.

  • Wrong answers sound right. Nielsen Norman Group describes AI hallucinations as output that seems plausible but is incorrect, delivered with confidence.

  • The assistant's errors belong to the company. In Moffatt v. Air Canada (2024), a British Columbia tribunal held the airline liable after its website chatbot gave a customer wrong information about bereavement fares. The tribunal said the chatbot was still just a part of Air Canada's website.

Better technology reduces the risk without removing it. Stanford RegLab researchers tested legal research tools that use retrieval-augmented generation. Two of the tools hallucinated more than 17% of the time, and a third more than 34% of the time.

What to look for when you test an AI assistant

What to watch

What good looks like

Warning sign

Answer accuracy

Matches current policy, pricing, and product facts

A confident answer the business can't support

Trust calibration

People check high-stakes answers and act on simple ones

People act on a wrong answer without checking

Recovery

The person spots a problem, rephrases, and the assistant corrects itself

The person loops, gives up, or leaves with the wrong answer

Handoff to a person

A clear route to a human, with the conversation carried over

The person can't find a human or has to repeat everything

Guardrails and refusals

Declines out-of-scope requests plainly and offers a next step

Refuses safe questions, or answers ones it should decline

Tone

Short, direct, and on-brand

Small talk, flattery, or long blocks of text

Trust calibration deserves extra attention. NN/g's Pavel Samsonov notes that confident tone and polished formatting discourage people from checking AI output, and that users often trust it enough to skip reading it end to end. Note whether participants verify an answer before acting on it.

Tone is easy to get wrong. In an NN/g study of 9 participants using 8 site chatbots, people typed short, search-style queries, skipped pleasantries, and pushed back on flattering lines like "great question." The same research found people will not experiment to discover what a chatbot is for. If the value isn't obvious right away, they move on.

How to set up sessions with real customers

Recruit the right people. Plan 5 to 8 participants per customer group, per round. Our guide on how many user interviews you need explains why this sample size holds up. A good screener keeps out people who don't match the group.

Gather real prompts first. Before sessions, pull a sample of real questions from support tickets, site search logs, and call notes. Ask each participant to bring one question they actually have.

Write goal-based tasks. Give people a goal and let them use their own words. A hypothetical health plan task might read: "You need urgent care this weekend and the closest clinic may be out of network. Find out what you would pay." Avoid scripting the exact question.

Include the hard cases on purpose. Add at least one task where the right response is a refusal, and one that should end with a person. Add one where the honest answer is "I don't know."

Repeat critical prompts. Outside the sessions, run each high-stakes question several times and compare the answers.

Consider a Wizard of Oz test. If the assistant isn't built yet, a team member can write the replies behind the scenes while the participant uses a realistic chat screen. NN/g recommends this method for conversational interfaces and for testing ideas before investing in costly technology like generative AI. Give the "wizard" tone-of-voice guidelines and a list of answers the business can support.

What to measure

Measure

How to capture it

How to report it

Task success

Did the person end with a correct, usable answer? Check against your source of truth.

Counts, such as "5 of 8 completed"

Answer accuracy

Grade each answer as correct, partly correct, or wrong

Counts per task

Unsupported answers

Flag any answer that conflicts with policy, pricing, or legal terms

Every instance, with the transcript

Recovery

After a wrong answer, did the person notice and get back on track?

Counts and quotes

Handoff

Number of turns to reach a person, and whether context carried over

Turns per task

Trust

After key answers, ask "Would you act on this?" Compare with accuracy.

Count of mismatches

Report small samples as counts. Keep full transcripts for every flagged answer. Our guide to writing a research readout shows how to turn these findings into decisions.

A checklist for testing conversational AI

  • [ ] Real customer questions gathered from support, search, and sales

  • [ ] 5 to 8 participants per customer group, screened

  • [ ] Goal-based tasks written in plain language

  • [ ] At least one task each for a refusal, a handoff, and an "I don't know"

  • [ ] Critical prompts run several times to check for variation

  • [ ] Handoff tested end to end, including what the human agent sees

  • [ ] Transcripts saved for every wrong or unsupported answer

  • [ ] A second round planned to confirm fixes

What investors should ask about an AI product

Ask to see transcripts from sessions with real customers. Ask how often the assistant is wrong on the questions that matter most, and how the team knows. Ask how handoff works and who keeps the source content current. Weak answers point to support costs and legal exposure after close. Our guide to testing a growth plan before close covers how customer evidence fits into diligence.

Frequently asked questions

How many people do you need to test an AI assistant?

Plan 5 to 8 participants per customer group, per round. Because answers vary, also run each critical question several times outside the sessions.

What is Wizard of Oz testing for chatbots?

A person writes the assistant's replies behind the scenes while the participant uses a realistic chat screen. It lets you test the conversation before you build the model.

How do you test AI guardrails with users?

Include tasks that should trigger a refusal or a handoff. Watch whether the assistant declines plainly and offers a next step, and check that it still answers safe questions.

Should participants know they are talking to an AI?

Yes, in most tests, because that matches what customers will see at launch. In a Wizard of Oz test, explain the setup to participants afterward.

What is the most common mistake in chatbot user testing?

Testing only with prompts the team wrote. Real customers phrase questions differently and ask things the team never planned for.

Get real evidence before your AI assistant goes live

Habib Innovation Partners tests AI assistants with the customers who will use them. Through the Innovation Retainer, we run small, fast rounds that show where your assistant helps, where it fails, and what to fix first. For investors evaluating a company's AI product, Investor Validation delivers fixed-fee customer research in 2 to 4 weeks. See how it works or start a conversation.

Sources

Testing conversational AI with real customers works differently from testing a website. A website shows every person the same screen. An AI assistant writes a new answer every time. Some of those answers will be wrong, off-brand, or outside what the business can stand behind.

This guide covers what to watch, how to set up sessions, what to measure, and a checklist. It is for product teams preparing a launch and investors judging an AI feature.

Why testing conversational AI is different

A standard usability test checks whether people can find and use things. With an AI assistant, you also check what it says.

  • Answers vary. The same question can get a different answer on a second try.

  • Wrong answers sound right. Nielsen Norman Group describes AI hallucinations as output that seems plausible but is incorrect, delivered with confidence.

  • The assistant's errors belong to the company. In Moffatt v. Air Canada (2024), a British Columbia tribunal held the airline liable after its website chatbot gave a customer wrong information about bereavement fares. The tribunal said the chatbot was still just a part of Air Canada's website.

Better technology reduces the risk without removing it. Stanford RegLab researchers tested legal research tools that use retrieval-augmented generation. Two of the tools hallucinated more than 17% of the time, and a third more than 34% of the time.

What to look for when you test an AI assistant

What to watch

What good looks like

Warning sign

Answer accuracy

Matches current policy, pricing, and product facts

A confident answer the business can't support

Trust calibration

People check high-stakes answers and act on simple ones

People act on a wrong answer without checking

Recovery

The person spots a problem, rephrases, and the assistant corrects itself

The person loops, gives up, or leaves with the wrong answer

Handoff to a person

A clear route to a human, with the conversation carried over

The person can't find a human or has to repeat everything

Guardrails and refusals

Declines out-of-scope requests plainly and offers a next step

Refuses safe questions, or answers ones it should decline

Tone

Short, direct, and on-brand

Small talk, flattery, or long blocks of text

Trust calibration deserves extra attention. NN/g's Pavel Samsonov notes that confident tone and polished formatting discourage people from checking AI output, and that users often trust it enough to skip reading it end to end. Note whether participants verify an answer before acting on it.

Tone is easy to get wrong. In an NN/g study of 9 participants using 8 site chatbots, people typed short, search-style queries, skipped pleasantries, and pushed back on flattering lines like "great question." The same research found people will not experiment to discover what a chatbot is for. If the value isn't obvious right away, they move on.

How to set up sessions with real customers

Recruit the right people. Plan 5 to 8 participants per customer group, per round. Our guide on how many user interviews you need explains why this sample size holds up. A good screener keeps out people who don't match the group.

Gather real prompts first. Before sessions, pull a sample of real questions from support tickets, site search logs, and call notes. Ask each participant to bring one question they actually have.

Write goal-based tasks. Give people a goal and let them use their own words. A hypothetical health plan task might read: "You need urgent care this weekend and the closest clinic may be out of network. Find out what you would pay." Avoid scripting the exact question.

Include the hard cases on purpose. Add at least one task where the right response is a refusal, and one that should end with a person. Add one where the honest answer is "I don't know."

Repeat critical prompts. Outside the sessions, run each high-stakes question several times and compare the answers.

Consider a Wizard of Oz test. If the assistant isn't built yet, a team member can write the replies behind the scenes while the participant uses a realistic chat screen. NN/g recommends this method for conversational interfaces and for testing ideas before investing in costly technology like generative AI. Give the "wizard" tone-of-voice guidelines and a list of answers the business can support.

What to measure

Measure

How to capture it

How to report it

Task success

Did the person end with a correct, usable answer? Check against your source of truth.

Counts, such as "5 of 8 completed"

Answer accuracy

Grade each answer as correct, partly correct, or wrong

Counts per task

Unsupported answers

Flag any answer that conflicts with policy, pricing, or legal terms

Every instance, with the transcript

Recovery

After a wrong answer, did the person notice and get back on track?

Counts and quotes

Handoff

Number of turns to reach a person, and whether context carried over

Turns per task

Trust

After key answers, ask "Would you act on this?" Compare with accuracy.

Count of mismatches

Report small samples as counts. Keep full transcripts for every flagged answer. Our guide to writing a research readout shows how to turn these findings into decisions.

A checklist for testing conversational AI

  • [ ] Real customer questions gathered from support, search, and sales

  • [ ] 5 to 8 participants per customer group, screened

  • [ ] Goal-based tasks written in plain language

  • [ ] At least one task each for a refusal, a handoff, and an "I don't know"

  • [ ] Critical prompts run several times to check for variation

  • [ ] Handoff tested end to end, including what the human agent sees

  • [ ] Transcripts saved for every wrong or unsupported answer

  • [ ] A second round planned to confirm fixes

What investors should ask about an AI product

Ask to see transcripts from sessions with real customers. Ask how often the assistant is wrong on the questions that matter most, and how the team knows. Ask how handoff works and who keeps the source content current. Weak answers point to support costs and legal exposure after close. Our guide to testing a growth plan before close covers how customer evidence fits into diligence.

Frequently asked questions

How many people do you need to test an AI assistant?

Plan 5 to 8 participants per customer group, per round. Because answers vary, also run each critical question several times outside the sessions.

What is Wizard of Oz testing for chatbots?

A person writes the assistant's replies behind the scenes while the participant uses a realistic chat screen. It lets you test the conversation before you build the model.

How do you test AI guardrails with users?

Include tasks that should trigger a refusal or a handoff. Watch whether the assistant declines plainly and offers a next step, and check that it still answers safe questions.

Should participants know they are talking to an AI?

Yes, in most tests, because that matches what customers will see at launch. In a Wizard of Oz test, explain the setup to participants afterward.

What is the most common mistake in chatbot user testing?

Testing only with prompts the team wrote. Real customers phrase questions differently and ask things the team never planned for.

Get real evidence before your AI assistant goes live

Habib Innovation Partners tests AI assistants with the customers who will use them. Through the Innovation Retainer, we run small, fast rounds that show where your assistant helps, where it fails, and what to fix first. For investors evaluating a company's AI product, Investor Validation delivers fixed-fee customer research in 2 to 4 weeks. See how it works or start a conversation.

Sources

© 2026 Habib Innovation Partners LLC

Meaning, made usable.

habibinnovation.com