Control the pilot scope
Select a few support intents with approved sources and a clear human owner. Define the channels, languages and hours included. Keep policy exceptions and unsupported actions outside the automated scope. A US–Saudi pilot should contain both languages and representative regional scenarios, but its results only describe that sample. Save the question set before tuning so the evaluation cannot drift toward easier cases.
Decide what success and stopping mean
Measure factual accuracy, appropriate handoff, reviewer effort and customer outcomes alongside speed. Set your own acceptance thresholds based on the consequence of errors; there is no universal score that proves readiness. Specify critical failures that pause the pilot. Keep a separate held-out set to check whether improvements generalize beyond the questions used during tuning.
Fictional worked example
Fictional pilot: 30 delivery-policy questions and 20 handoff cases, reviewed in Arabic and English. If replies become faster but miss an exception requiring approval, investigate before expansion. A report should show the counts, language mix, errors and review time rather than advertising a single 'automation rate.' These numbers illustrate a design and are not measured Anti Sandbox results.
Action checklist
- Name the included intents, sources, languages, channels and owner.
- Save baseline and held-out questions before making changes.
- Define acceptance thresholds and critical stop conditions.
- Publish sample counts and errors with any improvement claim.
