Step-by-step tutorial Analytics & growth

Chatbot feedback analytics: turn ratings into better answers

Compare one negative and one positive rating, inspect the exact answer context and retest in a fresh session—without mistaking feedback for conversion.

Beginner29 min readJuly 16, 2026
Chatbot feedback analytics: turn ratings into better answers

Chatbot feedback analytics are useful only when each rating is reviewed with its exact question, answer, conversation context and later regression result.

A thumbs rating is a pointer to one answer, not a complete verdict on the assistant. A visitor may reward an honest “I do not know,” dislike a correct answer because it lacks the next action, or never rate an answer at all. Start with the exact question, answer and conversation context before you decide what needs to change.

This walkthrough creates three isolated fictional sessions: a thumbs-down on a request for downloadable 2028 warranty terms, a thumbs-up on an honest missing-policy answer and a fresh positive regression check. WebChatAgent stores the exact answer, originating question, assistant, language, page and session reference for review.

Those signals show two different visitor reactions to safe missing-information answers. They do not prove purchase intent, conversion or overall satisfaction. The tutorial therefore separates observed evidence from interpretation, compares one fixed cohort at a time and deletes all three exact feedback and conversation sessions after verification.

Privacy-protected two-click player

Chatbot Feedback Analytics: Fix Thumbs-Down Answers

Use chatbot feedback analytics to inspect rated answers, separate explicit feedback from inferred interest, fix the responsible layer and retest.

YouTube · 3:48 · English

The YouTube player stays blocked until you choose Play. Loading it connects your browser to YouTube and may transfer technical data to Google.

Open directly on YouTube

What you will have at the end

  • A correctly enabled feedback flow with visible thumbs and an intentional notification owner
  • A reproducible cohort filtered by assistant, language, signal and date instead of one convenient rating
  • A cause classification that points to the smallest responsible source, prompt, model or follow-up flow
  • A before-and-after regression check with the original question, natural variants and recurrence monitoring

Before you start

  • A Premium assistant with the Feedback tool available
  • English Light Mode for the reproducible interface walkthrough
  • Access to Feedback, Conversations, the assistant sources and the test chat
  • A review owner, a regular review cadence and only fictional or approved test data
  • For reusable recordings: feedback email notifications disabled and permission to remove the exact test feedback and conversations afterward

A useful feedback loop separates signal, evidence, cause and result

Signal tells you where to look. Evidence is the exact answer, question, language, page, time and surrounding conversation. Cause explains why the answer helped or failed. Result is proven only after the smallest justified change survives the original question and nearby regression tests.

Do not calculate a satisfaction rate from the Feedback table alone. It contains submitted feedback records, not every answer that could have been rated. Use a documented denominator from the same assistant and date range before reporting a percentage, and label inferred interest separately from explicit thumbs.

Build one cohortRead exact evidenceFix, retest, monitor

01–05

Set it up step by step

1

Enable feedback, thumbs and deliberate ownership

Collect a visible signal only when somebody will review it.

Open the intended assistant, choose Tools and find Feedback. Turn on Enable feedback collection and keep Show thumbs up/down on bot answers in widget enabled. These two controls work together: the tool accepts feedback records, while the widget setting exposes the rating buttons below eligible assistant messages.

Enable email notifications only when the recipient is allowed to receive conversation content and has a written response process. A useful starting rule is: review negative feedback within one business day, inspect high-risk policy or safety answers immediately and include positive examples in the weekly quality sample. Save the settings, reload Tools and confirm every intended control is still enabled.

  • Name one review owner and one backup.
  • Decide which signals require immediate review.
  • Do not send real visitor content to an unapproved mailbox.
Collect a visible signal only when somebody will review it.
2

Build a fair cohort with assistant, language, signal and date

Make unlike conversations impossible to mix by accident.

Open Dashboard → Feedback. First select one assistant, then one language, one signal such as Negative Feedback or Positive Feedback, and a fixed date range. Use search only after those filters are set. Record the filter values and date so another reviewer can reproduce the same cohort.

The controlled example first filters one Tutorial Lab, English, Negative Feedback row, then switches to the separate Positive Feedback row used for detail inspection. This proves the filters and exact test records—not a broader pattern or purchase intent. In production, compare a meaningful batch with the same assistant, language and time window. Never call the row count a satisfaction rate unless you also know how many eligible answers were shown in that exact cohort.

  • Keep assistant, language and date range constant when comparing periods.
  • Separate explicit thumbs from inferred intent labels.
  • Sample unrated conversations too; raters are not every visitor.
Make unlike conversations impossible to mix by accident.
3

Open Feedback Details and read the exact evidence

A label becomes useful only beside the question and answer that created it.

Open the eye action for one row. Feedback Details shows the assistant, session ID, language, interest label, creation time, feedback content and stored context. Read the complete answer and the exact question together. Then open the source conversation when available to inspect what happened before and after this message.

In the verified example, the visitor asked: “What is the 2028 warranty policy for Northstar Services?” The stored answer says the available information does not contain that policy, and the row is Positive Feedback. The strongest explanation is not that the assistant knew the policy. It is that the assistant avoided inventing one. Mark this as a useful safe-abstention example and ask the policy owner whether an approved answer should exist.

For negative feedback, classify the evidence before editing anything: wrong or stale fact, missing information, poor tone, unclear format, failed action, missing follow-up route, expectation mismatch or frustration unrelated to the assistant answer. Attach one primary cause and optionally one secondary cause so monthly patterns remain countable.

  • Copy the exact question and answer into the review note.
  • Preserve session, page, language and time as diagnostic context.
  • Do not treat a positive rating as proof that a factual claim is correct.
A label becomes useful only beside the question and answer that created it.
4

Change the smallest layer supported by the evidence

Use a cause-to-fix map instead of changing model, prompt and sources together.

Choose the owner of the cause. Correct a wrong or outdated fact in the authoritative website or document. Add approved missing knowledge to the maintained source. Adjust the custom role when many answers share the same tone, scope or escalation problem. Repair Action Bar, lead, booking, live chat or API configuration when the answer lacks a working next step. Consider a model comparison only after sources, retrieval and instructions are sound.

For the warranty example, do not add a plausible policy or rewrite the safe answer merely to increase confidence. Ask the policy owner for the approved 2028 text and effective date. If none exists, keep the honest response and provide a monitored support route. If approved content arrives, update its authoritative source, remove conflicting duplicates, wait for indexing to finish and record the source version.

Change one layer at a time. Save the exact before answer, responsible owner, change, source version and expected improvement. That creates a clean explanation when feedback improves, stays flat or gets worse.

  • Wrong fact → authoritative source.
  • Repeated behavior problem → custom role or workflow instruction.
  • Missing action → tool, connector or human follow-up path.
  • Model change → only after controlled comparison on a fixed question set.
Use a cause-to-fix map instead of changing model, prompt and sources together.
5

Retest the original question, variants and nearby risks

Prove answer quality before you watch the rating count.

Start a fresh test conversation after the changed source shows Completed. Ask the exact original question, then natural variants such as “Does Northstar cover repairs in 2028?” and “Where can I read the current warranty terms?”. Compare each answer with the approved source and verify citations, dates, limitations, tone and next action. Also ask one intentionally unsupported question to confirm that safe abstention still works.

The visible example is a verified baseline: the exact warranty question received a safe answer and the widget exposed the thumbs controls. Because no approved 2028 policy was added during this demonstration, the tutorial does not claim a corrected warranty answer. A real change passes only when the expected content appears in a clean session and nearby regression questions remain correct.

Monitor recurrence with the same assistant, language, signal and date-window rules. Compare cause counts per eligible answer when the denominator is available, not raw thumbs alone. Review high-risk failures immediately and summarize lower-risk themes monthly with owner, source version, action, retest result and whether the pattern returned.

  • Use a new session so cached conversation context cannot hide a regression.
  • Test the original wording plus at least two natural variants.
  • Keep one unsupported control question in the regression set.
  • Record the observed answer, not only pass or fail.
Prove answer quality before you watch the rating count.

Example & result

See the practical test and its result

Every tutorial includes a fixed input, the expected outcome and a transparent record of what was actually verified locally.

Practical example: Chatbot feedback analytics: turn ratings and interest signals into better answers

This exact scenario was completed with the temporary tutorial account.

Verified end to end

Exact test input

Rate the unavailable download request negatively, rate the honest 2028-policy answer positively and repeat the positive test in a fresh session.

Expected result

Three exact session records preserve the two questions, safe answers and ratings; filters separate negative from positive evidence.

What was actually verified

The isolated real capture verifies one Negative Feedback and two Positive Feedback records, inspects the exact positive context and removes all three feedback and conversation sessions afterward.

The isolated real capture verifies one Negative Feedback and two Positive Feedback records, inspects the exact positive context and removes all three feedback and conversation sessions afterward.

Tips & tricks

Make the setup reliable

Test with realistic examples, record your baseline and change one setting at a time. That makes real improvements visible.

Keep thumbs and intent separate

A thumbs-up confirms one visitor reaction. Purchase intent needs its own evidence and a working lead, booking or human follow-up path.

Use positive examples as regression tests

A highly rated safe answer is worth preserving. Add it to the fixed test set before changing prompts, sources or models.

Prepare sources for clean diagnosis

Use clear headings, dates, owners and one current version. Remove duplicate or expired policies before blaming retrieval or the model.

Compare models only on a fixed scorecard

If source and prompt are already sound, compare models with the same questions, source evidence, quality criteria, latency and quota cost. Do not switch after one disliked answer.

When something does not work

Troubleshooting

Check status, permissions and test data systematically before changing the model or prompt.

The thumbs buttons are missing in the widget

Confirm that Feedback and Show thumbs up/down are both enabled for this assistant. Save, reload the assistant, start a fresh widget session and verify that the message is an eligible assistant answer rather than a system or live-agent message.

Feedback was clicked but no row appears

Check the selected assistant, date range, language and signal filters. Clear the search field, reload Feedback and verify that the visitor used the intended assistant and allowed website domain.

A rating looks inconsistent with the answer

Open Feedback Details and the source conversation. Visitors can rate honesty, tone or next-step usefulness rather than factual completeness. Classify what the signal actually supports and do not silently relabel it.

Negative feedback did not decline after the change

Confirm that indexing completed, tests used fresh sessions and the cohort stayed comparable. Inspect new examples for a different root cause; revert or revise the change if the exact failure remains.

Tutorial ratings remain in Feedback after the review

Delete only the exact fictional session IDs created by the test, then verify that both Feedback and Conversations no longer list them. Never clear a broad date range or another reviewer’s records just to reset the example.

Ready for a production-style test

Create a monthly feedback review with cohort definition, eligible-answer denominator where available, top cause classes, positive regression examples, interest themes, owners, source versions, verified changes and recurrence. Keep urgent safety or policy failures outside the monthly queue and review them immediately.

Related resources