Instructions
Quality dimensions tell you what matters — evaluation criteria tell you how to judge it. This step turns your quality framework into a practical testing tool: specific criteria you can apply consistently to every output, with clear definitions of what passing and failing look like for each one.
The common mistake is evaluating AI outputs subjectively and inconsistently. If you look at an output and think "that seems pretty good," you'll never know if your prompt changes are actually improving things or just producing different outputs. Consistent evaluation requires criteria that are specific enough to apply the same way every time — and clear enough that another person, or a future version of you, could apply them too.
Good evaluation criteria have a few properties. They're o...