Skip to content
anti sandbox.
Measurement and pilots

Plan a small AI support pilot with clear exit criteria

Choose a narrow set of questions, establish a manual baseline and decide what evidence is needed before expanding usage.

Control the pilot scope

Select a few support intents with approved sources and a clear human owner. Define the channels, languages and hours included. Keep policy exceptions and unsupported actions outside the automated scope. A US–Saudi pilot should contain both languages and representative regional scenarios, but its results only describe that sample. Save the question set before tuning so the evaluation cannot drift toward easier cases.

Decide what success and stopping mean

Measure factual accuracy, appropriate handoff, reviewer effort and customer outcomes alongside speed. Set your own acceptance thresholds based on the consequence of errors; there is no universal score that proves readiness. Specify critical failures that pause the pilot. Keep a separate held-out set to check whether improvements generalize beyond the questions used during tuning.

Fictional worked example

Fictional pilot: 30 delivery-policy questions and 20 handoff cases, reviewed in Arabic and English. If replies become faster but miss an exception requiring approval, investigate before expansion. A report should show the counts, language mix, errors and review time rather than advertising a single 'automation rate.' These numbers illustrate a design and are not measured Anti Sandbox results.

Action checklist

  1. Name the included intents, sources, languages, channels and owner.
  2. Save baseline and held-out questions before making changes.
  3. Define acceptance thresholds and critical stop conditions.
  4. Publish sample counts and errors with any improvement claim.
Related pages
See it yourself

A workspace worth exploring.

Open the demo