A successful demonstration shows that a system can produce a good result. An evaluation asks how consistently it does so, under what conditions, and with what limitations.

Define the job

Write down the task in plain language. Who uses the output? What decision follows? What should happen when the system cannot provide a reliable answer? These questions make evaluation criteria more concrete.

Choose representative scenarios

Include typical inputs, ambiguous requests, incomplete information, and situations with meaningful consequences. For a retrieval application, test what happens when the source material is missing or contradictory.

Separate kinds of quality

Accuracy, groundedness, helpfulness, instruction following, and appropriate refusal are different questions. A single aggregate score can hide a serious weakness in one of them.

Decide what evidence is enough

Set acceptance criteria before interpreting results. Document the test conditions and evaluation method, including the role of human review. Record unresolved issues and conditions attached to deployment.

Plan to revisit

Changing the model, prompt, retrieval source, permissions, or intended audience can change behavior. Make reassessment part of the workflow rather than a one-time exercise.

An evaluation should support a specific decision and make the limits of its evidence easy to understand.

A practical perspective from Blue Lotus. The right approach depends on your system, context, and responsibilities.

Talk through your situation