Typed verdicts and field kinds¶
An evaluator's verdict is an instance of a Pydantic model: the verdict type. It is usually the same type people use to give feedback, so an evaluator predicts exactly what a person would have said, and the two can be compared field by field (ADR-0001). evalr needs nothing else from the type: an application's own model, a feedback type registered with artifactr or reflexr, or an ad-hoc model all work.
A verdict type¶
from enum import Enum
from typing import Annotated, Literal
from pydantic import BaseModel, Field
class Topic(Enum):
BILLING = "billing"
BUG = "bug"
OTHER = "other"
class Helpfulness(BaseModel):
"""Whether the reply helped the person."""
rating: Annotated[int, Field(ge=1, le=5, description="How much the reply helped")]
resolved: bool = Field(description="The request was fully addressed")
topic: Topic = Field(default=Topic.OTHER, description="What the request was about")
tone: Literal["warm", "neutral", "cold"] | None = None
share_done: float = Field(default=1.0, ge=0.0, le=1.0, description="How much was done")
reason: str | None = Field(default=None, description="Why, in a sentence")
The docstring says what the verdict is for, and each field's description is the instruction an evaluator reads for that field. Language-model judges see both; decision models turn each field into a question with the description as its text.
Field kinds¶
The type of each field decides how it is judged and how it is scored. evalr calls this the field's kind:
| Field type | Kind | Filled by | Agreement with people | Score |
|---|---|---|---|---|
bool |
binary | any evaluator | accuracy, Cohen's kappa | BOOLEAN |
Literal, Enum |
categorical | any evaluator | accuracy, Cohen's kappa | CATEGORICAL |
int bounded on both sides |
ordinal | any evaluator | mean absolute error, Spearman | NUMERIC |
any other int or float |
numeric | any evaluator | mean absolute error, Spearman | NUMERIC |
str |
text | language-model judges only | not compared | TEXT |
A few rules complete the table:
X | Noneis judged asX, and may be left empty. An empty field gives no score and is left out of agreement.- Bounds come from
Field(ge=, le=, gt=, lt=)orannotated_typesconstraints, at any level ofAnnotated, including inside an optional. An integer's exclusive bounds become inclusive:gt=0is a lower bound of 1. - Anything else (a list, a nested model, a union of several types) cannot be judged. Evaluators, datasets and metrics raise
UnsupportedFieldas soon as they are given such a type, rather than failing on the first input.
verdict_fields describes a type the way every evaluator, metric and integration sees it. It is the one place in evalr that reads field types:
from evalr import verdict_fields
for field in verdict_fields(Helpfulness):
print(field.name, field.kind.value, field.lower, field.upper, field.optional)
rating ordinal 1.0 5.0 False
resolved binary None None False
topic categorical None None False
tone categorical None None True
share_done numeric 0.0 1.0 False
reason text None None True
Each VerdictField also has the field's description, its choices (a Literal's values or an Enum's members, and (False, True) for a binary field) and whether the model required it.
Verdicts¶
An evaluator returns a Verdict[V], an immutable Pydantic model that wraps the value with what it takes to trust it and trace it:
| Field | Holds |
|---|---|
value |
The verdict type's instance |
confidence |
The probability that each field's value is right, from 0 to 1, for the fields the evaluator has one for. Decision models report them; most language-model judges do not. Its keys must be fields of the verdict type. |
evaluator, version |
Who judged, and which version of it |
latency |
Wall-clock seconds the evaluation took |
cost |
US dollars, when the evaluator knows |
trace_id |
The OpenTelemetry trace the evaluation ran in, as 32 hex digits, when there was one |
from evalr import Verdict
verdict = Verdict(
value=Helpfulness(rating=4, resolved=True, reason="It refunded the charge"),
confidence={"resolved": 0.92},
evaluator="helpfulness-decision",
version="3f2a9c0d1e4b",
)
print(verdict.value.rating, verdict.confidence["resolved"])
You rarely build one by hand: evaluators do, through judging, which fills in the latency and the trace.
Versions keep verdicts apart¶
Every evaluator has a name and a version, and every verdict records both. When an evaluator changes (its instructions are retrained, its thresholds recalibrated, its function rewritten), its version changes, so verdicts from before and after are never mixed in a metric, a score or an experiment. Trained evaluators derive their version from what they are: a DSPy judge hashes its program and types, and a decision evaluator its model, types, instructions and thresholds. A function evaluator's version is given, so bump it when the function changes.
Designing a verdict type¶
- Describe every field. The description is the only instruction a judge has for the field, and the text of a decision model's question.
- Prefer bounded ratings (
Annotated[int, Field(ge=1, le=5)]) to open-ended numbers. They are ordinal, their agreement is measured over the scale, and a decision model can answer them as a choice. - Give text fields a default. A decision model cannot write text, so it can only judge a type whose text fields can be left empty.
reason: str | None = Noneworks with every kind of evaluator; a requiredstrlimits the type to language-model judges. - Keep it flat. One level of fields of the kinds above. A verdict that needs structure is usually two verdict types.
- Avoid clashes with the input type. A DSPy judge's signature has a field for every field of the input and of the verdict, so their names must differ, and
reasoningis DSPy's.
Feedback types from artifactr and reflexr¶
artifactr's and reflexr's feedback types are Pydantic models, so they are verdict types as they are: an evaluator can give the same Helpfulness feedback people give in a thread. Both libraries register a feedback type under a name (helpfulness), and scores are named {type}.{field} after it, the convention the three libraries share. evalr names scores after the class, in snake case, unless told otherwise: pass the registered name as type_name where it differs (Scores and score sinks).