AI coding · Code quality
AI Evaluation
Also called: Evals, Eval set, Model evaluation, Quality regression
Use a fixed set of cases and scoring rules to see whether a prompt or model change got better or worse—not just two gut-feel tries.
In detail
Evals turn "feels smarter" into comparable numbers or checklists. Keep 20 real user questions, a few tricky UIs, or fields that must be extracted, and rerun them whenever you change prompts or models.
For Vibe Coding products (diagnose, generate, upgrade), evals matter: swapping models, rules, or temperature can silently regress. Scoring can be scripts (exact match, structured fields) or human rubrics.
Tiny or fake eval sets make you overfit the exam. Grow the set from real failures.
Developer infoTerm ID, DOM cues, match priority
- Term ID
ai-eval- DOM selectors
- No DOM cues. This concept isn't detected directly on a page.
- Priority
- 1 · when several match at the same level, the higher priority wins
- Version
- v1 · updated Oct 5, 2026