The 60-second version
- Flag a bad AI reply: point at it and click the thumbs-down. Say what went wrong in your own words; AI assist drafts the correction and you submit it.
- Triage in AI Feedback. Each flag is shown with its conversation. Keep it, edit it, or dismiss it.
- Review & improve: four screens (keep or dismiss, see why, review the fix, choose tests). You approve a before-and-after diff before anything is saved.
- Run the tests it added, so the fix stays fixed.
Fact, person or behaviour?
“Teaching the AI” is three different jobs, and picking the right one saves an afternoon. Ask what the reply got wrong:
How the feedback loop works
Four moments, read left to right. Each is a different screen and usually a different person: whoever is working the Inbox flags, and an admin does the rest.
The customer gets an answer that is not true.
Anyone reading the chat clicks the thumbs-down.
An admin keeps what is right and dismisses what is not.
A fix is written, you approve it, and a test guards it.
- A flag belongs to one reply and one person. It stores the reply, the conversation around it, what the reply should have said and a category.
- A fix is a reviewed edit, not a retrain. The same copilot you use in the agent editor writes it, in a cautious mode that prefers small surgical changes. Nothing is saved until you press Apply the fix.
- The diagnosis comes first. Before you see a fix, each kept flag is traced to the part of the setup that caused it, so you judge whether the fix removes the cause.
- A test guards the fix. Any kept flag can become a test that replays the conversation and checks the new reply.
Flag a reply
Anyone reading the conversation can react to an AI agent’s reply. Only replies sent by an AI agent have the thumbs; messages from people, workflows and broadcasts do not.
Point at the reply and click the thumbs-down
Say what went wrong
A drawer titled What went wrong? slides in with the reply quoted at the top. Type into the box at the bottom, in plain words, for example “It quoted a 25% goodwill credit but our SOP says 50%.” Press Enter to send. AI assist may ask a question or two. You can stop it at any time with Skip the questions - draft it now.
The draft has three parts. What the AI should have said and Why are in your own words and editable. Filed under is the category AI assist picked; click Change to pick another. In the drawer the four categories read as Said something untrue, Wrong tone, Ignored the conversation and Something else. The admin later sees the same four as Hallucination, Wrong tone, Missed context and Other. If you add more detail after the draft appears, Done - review the draft takes you straight back to it.
Or mark a good one
A thumbs-up is one click, with no drawer: you get a “Marked as a good response” message with Undo for a few seconds. Either way a small pill, Flagged or Marked good, now sits under the reply.
- One feedback per person per reply. Click your own thumb again, or the pill, to edit it. A drawer titled Feedback on this response lists every teammate’s feedback, and offers Edit my feedback (or Mark good / Flag if you haven’t voted).
- You can switch a verdict only while it is Pending. Once an admin has trained on or rejected it, the other thumb is greyed out with the tooltip “Already used for training - this verdict is locked”.
Triage the queue
Every flag lands in the AI Feedback page of the Admin Dashboard. Getting there takes two clicks.
Open the Admin Dashboard
Click AI Feedback
Read the queue
Search looks through the what it should have said and why text. The waiting line counts only Pending thumbs-down items for that agent; good examples are listed but never counted.
Open one item and decide
Click a row. A drawer opens with Feedback on {agent} response and, top right, Agent (open the agent) and Chat (open the conversation).
- Feedback: the flagged message, category, what it should have said, the reviewer’s reason, and a collapsed Clarification chat (the reviewer’s own words often hold a detail the summary dropped).
- Conversation context: the messages around the flagged turn, so you can see what the customer really asked.
- LLM call: the tokens used, the prompt the model was sent and its raw response. It doesn’t show a model name, and it is greyed out when nothing was captured for that message.
The moves on an item: Edit any field; Reject with a reason (Duplicate, Not actionable, Not the AI’s fault, Off-topic or Other) and an optional note; Move back to Pending for a rejected item; Promote to test case; and Delete when nothing depends on it yet. If you edit an item that was already used in a fix, a banner reminds you the edit won’t change that past run.
Review & improve
This is where flags change the agent. It is one full-screen flow with four steps, on top of the agent’s editor, so the sidebar is gone on purpose. A step rail across the top shows where you are, and Leave is always top right. It works on the agent’s Pending thumbs-down items only.
Start it
The Feedback tab is the same list as the AI Feedback page, filtered to this agent. If nothing is waiting the flow says so and there is nothing to do.
Step 1 of 4: keep or dismiss each flag
Not every flag is right, and this is where you say so. The counter tracks to fix · dismissed · left, and Change my mind undoes a decision. Dismissals are saved as soon as you move on.
Step 2 of 4: see why it happened
At the top, The pattern sums up what the flags have in common. Each finding then names where the cause lives: Instructions, Actions, Arguments, Exit reasons, Required resources, Output data, Knowledge or Outside the agent. It quotes the wording responsible and gives a Direction for the fix, marked “a best guess - check it” when the diagnosis is unsure.
Step 3 of 4: review the fix
I added one rule so trial slots only come from the booking action, and softened the follow-up.
If the fix touches more than the instructions, each other change appears as its own before-and-after card that you can edit or switch off. Continue stays greyed out while the fix is being written, or if nothing is left to apply.
Step 4 of 4: choose tests and apply
Tests show up in the agent’s Tests tab, ready to run after every change.
- Instructions updated
- 2 pieces of feedback marked as fixed
- 1 dismissed
- 1 regression test added
Applying does four things: writes the change to the agent, moves the kept flags to the Trained tab (labelled with the run), leaves the dismissed ones on Rejected, and creates a test for each one you ticked. If you leave before Apply, you are asked “Leave without applying the fix?”
Tests keep it fixed
A fix that is never re-checked quietly comes undone the next time someone edits the instructions for another reason. A test freezes the situation: it replays the conversation up to the flagged reply (with up to five messages before it) and an AI judge checks the agent’s new reply against what it should have said.
- From the review flow, ticked items become tests when you apply the fix.
- From one queue item, click Promote to test case in its drawer. It opens the new test in the Tests tab, and the button becomes View test case. The caret beside it offers Test as a simulated contact, which reads the whole conversation into a persona the agent must handle live on every run; it takes a few seconds longer.
- Good replies can be promoted too. The note you left is what the test checks against, so write one.
Setups to copy
- 1Each time it happens, click the thumbs-down and write a concrete What the AI should have said.
- 2When three or four have piled up, click Review & improve on that agent.
- 3Fix this the real ones, read the diagnosis, then read the diff as if you were publishing it.
- 4Tick the concrete ones as tests, Apply the fix, then Run the tests.
- 1When the agent nails a tricky reply, click the thumbs-up. One click, with Undo.
- 2Click the Marked good pill, then Edit my feedback, and say why it was good.
- 3In AI Feedback, open it and click Promote to test case. It is checked on every run.
- 1Flagged but actually fine? Dismiss it as Not the AI’s fault with a one-line note.
- 2Same issue twice? Dismiss the second as Duplicate.
- 3Fixed the FAQ in the Knowledge Hub instead? Dismiss as Not actionable and say so in the note.
Good to know & pitfalls
- Nothing is fine-tuned. The fix is a reviewed edit to the agent’s own setup: its instructions and, where the cause is there, its actions, arguments, exit conditions, required resources or output data. It is fast, visible, and yours to edit.
- Apply the fix saves the agent straight away. There is no separate draft step, and the fix does not touch the Draft / Published switch. A Published agent uses the new setup on its next reply, so Step 3 is your only review.
- Only Pending thumbs-down items feed a fix. Good examples never do; their job is to become tests. Dismissed and already-trained items are left out too.
- Concrete beats vague. What the AI should have said is required to submit, and it is what the fix and the test lean on hardest. “Tone was off” teaches little; the exact sentence you wanted teaches a lot.
- Dismissals save as you go; the fix does not. The moment you move to Step 2, your dismissals are saved. Leave after that and the diagnosis and the written fix are discarded, the agent stays as it was, and dismissed flags stay dismissed.
- Every item carries its context. Open one and you get Feedback, Conversation context and LLM call (tokens in and out, the exact prompt sent, the raw response). The LLM call tab is greyed out when nothing was captured.
- Reject keeps a record; Delete does not. Rejected items stay on the Rejected tab with their reason. Delete is permanent and offered only while an item is untouched: not trained, not used in a fix, not promoted to a test.
- Only an AI agent’s replies can be flagged. An AI Model step in a workflow computes things for other steps and never sends a message, so there is nothing to flag.
- Library-managed agents can’t be refined this way. Their instructions belong to the library; the service refuses with “Managed agents cannot be refined directly - edit the agent configuration instead”.
- Tests do not run themselves. A test is checked when you click Run All or run it from its row. Run them after every fix.
Need a hand?
Our Singapore-based team is one message away - happy to help you get set up.