Pick your feature's task type and the failure modes you're worried about, and get a starter eval-question scaffold plus a rubric structure — edge cases, adversarial inputs, format-compliance checks, and more — from authored patterns. You fill in the expected answers for your feature; nothing here runs an evaluation for you. Free.
Pick the task type and any failure modes you want the scaffold to specifically cover.
The scaffold, every category and the rubric are free and stay free. The Worked Example adds what a scaffold can't: what a filled-in eval set actually looks like.
Team licence for an engineering org? Email for a group rate.
Unlocks the Worked Example on this device and survives reloads. No account, nothing transmitted.