Test and improve

Use Preview, simulations, conversations, and gaps to improve a Viber agent.

Viber provides one improvement loop across manual testing, automated tests, and real conversations:

  1. test or observe the agent;
  2. inspect the answer, sources, Playbook, and tool activity;
  3. correct the agent with Vibe or the structured screens;
  4. replay or rerun the scenario; and
  5. publish the verified change.

Preview the working version

Open Preview from the main agent navigation to have a conversation with the current working version. Preview can include unpublished Instructions, Playbooks, Knowledge, and tool references.

Use Start new session to start a fresh Preview conversation. This clears the conversation context but does not publish or discard the working version.

Correct and replay Preview

Flag an incorrect Preview answer to describe what was wrong and send it to Vibe. When the correction changes the working version, Viber can offer Replay conversation to rerun the user turns against the updated draft.

Replay is useful when the correction affects an earlier turn or the path through the conversation. Start a new session instead when you want a different scenario.

Flagging an inaccurate Acme Support answer and describing the correction

Run simulations

Open Monitor > Simulations to create repeatable tests. Viber supports question tests and flow tests.

Question test

Sends one question, captures one answer and its cited sources, then rates the answer.

Flow test

Runs a multi-turn scenario against one Playbook and evaluates the resulting behavior.

Question tests

A question test checks an important knowledge answer. The question is the only required content; an optional expected answer can help rate the first run.

Rate the result as:

  • Good when the answer is ready for the user;
  • Acceptable when the answer is usable but should improve; or
  • Poor when the answer is wrong or unsuitable.

The result includes cited sources so you can distinguish missing knowledge from incorrect retrieval, use, tone, or formatting. Later runs can use earlier ratings as a reference, but you can always review and override a rating.

Flow tests

A flow test exercises one Playbook through a multi-turn conversation with a virtual user. Define:

  • the user scenario;
  • an optional opening message;
  • how external tools should behave; and
  • the conditions that determine whether the run passes.

Mock external tool results when a real call would be unsafe, expensive, or inconsistent. Evaluations can check the agent’s reply, the full timeline, tool calls and outputs, or Playbook events.

Diagnose a failed run

A failed simulation does not automatically mean the agent is wrong. Inspect the run before choosing what to change:

  1. Read the transcript and the failure rationale.
  2. Confirm that a flow test targets the intended Playbook.
  3. Check whether mocked or real tool behavior matches the scenario.
  4. Check whether the evaluation describes the desired outcome accurately.
  5. For question tests, inspect the answer and cited sources.

Use Fix with Vibe to send the finished run and its diagnostic context to the builder conversation. After the change, rerun the simulation.

Learn from real conversations

Open Monitor > Conversations to review conversations from Preview and live channels. Use the channel and status context to distinguish a draft test from published traffic.

Open Monitor > Gaps to review identified weaknesses and send an actionable correction to Vibe. A gap can originate from feedback, a knowledge rating, or another monitored signal.

Connect channels

Choose where customers can talk to the published agent.