A successful demonstration shows that a system can produce a good result. An evaluation asks how consistently it does so, under what conditions, and with what limitations.
Define the job
Write down the task in plain language. Who uses the output? What decision follows? What should happen when the system cannot provide a reliable answer? These questions make evaluation criteria more concrete.
Choose representative scenarios
Include typical inputs, ambiguous requests, incomplete information, and situations with meaningful consequences. For a retrieval application, test what happens when the source material is missing or contradictory.
Separate kinds of quality
Accuracy, groundedness, helpfulness, instruction following, and appropriate refusal are different questions. A single aggregate score can hide a serious weakness in one of them.
Decide what evidence is enough
Set acceptance criteria before interpreting results. Document the test conditions and evaluation method, including the role of human review. Record unresolved issues and conditions attached to deployment.
Plan to revisit
Changing the model, prompt, retrieval source, permissions, or intended audience can change behavior. Make reassessment part of the workflow rather than a one-time exercise.
An evaluation should support a specific decision and make the limits of its evidence easy to understand.
A practical perspective from Blue Lotus. The right approach depends on your system, context, and responsibilities.
Talk through your situation