Product

Deterministic graders, not a model judge

The decision

D-1 recorded

Evalset grades with deterministic graders, not LLM-as-judge. What was cut to get there: LLM-as-judge grading. The reason is not recorded yet.

Source: Bryan, product-facts kickoff request, 2026-09-21. Last verified 2026-09-21.

That is what is on the record. Everything below is what the record does and does not support.

What got cut

Model-as-judge grading. That was the alternative on the table and it is the one that was dropped.

It is worth being precise about what that costs, because it is not nothing. A model judge scales to criteria a deterministic grader cannot express: tone, whether an explanation actually explains, whether a summary is faithful in the way a reader would mean it. Giving that up means the eval suite can only test what can be checked by a rule.

What I have not written down

The reason. The fact ledger has the decision and what was cut, and it has TODO(bryan) where the reasoning should be.

I could write a convincing paragraph here about determinism, reproducibility and not grading a model with a model. It would probably even be close. It would also be me reconstructing a decision from the outside and presenting it as a record, which is the exact failure mode this desk exists to avoid. Meetings and recollection are sources of intent. They are not sources of fact.

So the line stays open until I fill it in, and this page stays short.

What has to be in the finished memo

  1. Why deterministic. The specific failure with a model judge that made this not worth trying, or the constraint that made it impossible.
  2. What that costs. Which criteria the suite now cannot test at all, named.
  3. What would change it. The condition under which I would revisit this, written before I am tempted to revisit it.
  4. When. The date the decision was actually made, not the date it was written down.

Point four matters more than it looks.

T-3 recorded

2026-09-21, company/product-facts.md drafted for Bryan's review.

Source: the product-facts drafting task. Last verified 2026-09-21.

is when this got drafted. It is not when the decision was made, and the ledger currently has nothing for that.

The eval suite this describes has no shipped code or spec yet:

Not written yet, waiting on P-6

What evalset actually is, in one line, and where it is specified. No spec or shipped code was found in the calabrodesign repo or in the linked calabro-mcp pointer as of 2026-09-21, so there is nothing to link the decision to.

Checked 2026-09-21.