Define critical failures first
List the errors that make a reply unacceptable even if the prose is excellent: a fabricated order status, an unsupported commitment, the wrong customer reference or a missed required handoff. Then score factual accuracy, completeness, clarity and tone separately. Keep critical failures visible instead of averaging them into a high overall score.
Calibrate reviewers with the same cases
Give two reviewers a shared set of Arabic and English examples, the approved sources and the scoring definitions. Compare disagreements before evaluating a larger sample. Record why a score differs and refine the rubric. Reviewers should evaluate the same customer intent across languages, not reward a literal translation that changes the practical meaning.
Fictional worked example
In a fictional scoring exercise, a reply has natural Arabic and a friendly tone but states that an unverified refund was completed. Mark the unsupported completion as a critical failure. A second reply that clearly requests a human decision may be less elegant but operationally correct. The scorecard should make that distinction visible to whoever approves rollout.
Action checklist
- Define critical failures and do not hide them in an average.
- Score facts, completeness, handoff, clarity and tone separately.
- Calibrate at least two reviewers on shared examples when possible.
- Report disagreement and sample composition with the final scores.
