LLM feature testing ·

QE discipline, applied to your AI feature.

Pick your feature's task type and the failure modes you're worried about, and get a starter eval-question scaffold plus a rubric structure — edge cases, adversarial inputs, format-compliance checks, and more — from authored patterns. You fill in the expected answers for your feature; nothing here runs an evaluation for you. Free.

5task types
5optional concern categories
0bytes transmitted

Describe your feature

Pick the task type and any failure modes you want the scaffold to specifically cover.

Pricing

The scaffold, every category and the rubric are free and stay free. The Worked Example adds what a scaffold can't: what a filled-in eval set actually looks like.

Free
$0
  • The full scaffold — happy-path, edge-case and format categories always included
  • Optional concern categories (adversarial, hallucination, bias, safety, latency)
  • A task-type-specific rubric structure
  • 100% client-side, nothing transmitted
Worked Example
$39 one-time
  • A fully filled-in eval set for a sample classifier feature, expected answers included
  • Grading-consistency tips — keeping multiple reviewers aligned on the same rubric
  • Common scaffold-filling mistakes and how to avoid them
Get the Worked Example

Team licence for an engineering org? Email for a group rate.

Straight answers

Enter your licence code

Unlocks the Worked Example on this device and survives reloads. No account, nothing transmitted.