---
name: honest-grading
description: Grade completion from inspected evidence without inflating mocked, partial, or reported results.
version: 1.0.0
updated: 2026-08-14
---

# Honest Grading

Use these states:

- **Done:** required behavior is integrated and personally verified.
- **Built, unverified:** implementation exists but required runtime or visual proof is missing.
- **Partial:** some rubric items pass; list each failing item.
- **Blocked:** a specific external dependency or authority prevents progress.
- **Failed:** an attempted required check failed; include the literal useful error.

An agent report, green unit suite, mocked provider, screenshot not opened, or unconfirmed deployment cannot by itself earn “done.” Name which tests ran, which flows ran, which evidence you inspected, and which boundaries remain. Use exact counts where possible. An honest lower score is better than a persuasive fiction.

Measure served-model identity from the upstream provider response; never treat an echo of the
requested model as evidence. A buffered full response replayed in fragments is not streaming—prove
progressive arrival with chunk timing. Verify routing and middleware against the production start
command and production entrypoint because development servers can mask upstream interception.

When the deploy source is in scope, “done” also requires the focused change to be committed and
synchronized there. Publishing or deployment still requires separate explicit authority.

Keep workflow state separate from claim truth. Grade claims `PROVED`, `CHECKED`, `CONDITIONAL(on what)`, `OBSERVED`, or `SPECULATION`; research may add `REFEREED`, `REFUTED`, `BROKEN`, and `GAP`. Retain negative results and literal evidence ceilings.

A validator, manifest, route consumer, historical receipt, scripted transcript, or loose media artifact proves only that artifact exists. It does not prove the current target contains producing code or an executed end-to-end proof package. Pin independent verdicts to exact identical bytes; the builder cannot self-green.

Any receipt used to claim build completion must conform to the one canonical
[evidence receipt schema](../tutorial-verified-done/references/evidence-receipt-schema.md). A
historical or domain-specific receipt cannot substitute for its current prompt, story, E2E,
regression, distinct-frame, and twin fields.
