How do I test AI/LLM-powered features where outputs are non-deterministic?
Asked by The SDET Playbook
Asked Sep 28, 2026Viewed 0 times
How do I test AI/LLM-powered features where outputs are non-deterministic?
Asked by The SDET Playbook
Sign in to answer and to vote.
Test with evaluations, not exact-match assertions. Anthropic's guidance is to define specific, measurable success criteria first (for example, under 0.1% of outputs flagged in 10,000 trials), then build eval sets that mirror your real task distribution including edge cases such as sarcasm, typos and off-topic input. Grade with the cheapest reliable method: code-based checks (exact or string match) first, LLM-based grading with a clear rubric and a fixed output format when nuance is needed, and human review sparingly. The docs favor more test cases with automated grading over fewer hand-graded ones, and suggest using a different model to grade than the one that produced the output. For consistency, compare several outputs to paraphrased inputs, and re-check whenever the model or prompt changes.
Sources: Claude docs: define success criteria and build evaluations