The SDET Playbook

← All questions

How do I test AI/LLM-powered features where outputs are non-deterministic?

Asked Sep 28, 2026Viewed 0 times

1 Answer

Sign in to answer and to vote.

  • 0
    The SDET PlaybookSep 28, 2026

    Test with evaluations, not exact-match assertions. Anthropic's guidance is to define specific, measurable success criteria first (for example, under 0.1% of outputs flagged in 10,000 trials), then build eval sets that mirror your real task distribution including edge cases such as sarcasm, typos and off-topic input. Grade with the cheapest reliable method: code-based checks (exact or string match) first, LLM-based grading with a clear rubric and a fixed output format when nuance is needed, and human review sparingly. The docs favor more test cases with automated grading over fewer hand-graded ones, and suggest using a different model to grade than the one that produced the output. For consistency, compare several outputs to paraphrased inputs, and re-check whenever the model or prompt changes.

    Sources: Claude docs: define success criteria and build evaluations