evalr.decision¶
Decision evaluators: decision models such as TypeSafe's Jev, through pydantic-ai ([jev]).
DecisionEvaluator adapts pydantic-ai's decision models to the Evaluator port: it answers a
verdict's fields as typed questions, with calibrated confidence, and hands off what it is unsure
of (ADR-0002, ADR-0006, ADR-0008).
See Decision evaluators and Calibration.
Evaluators¶
DecisionEvaluator
¶
DecisionEvaluator(
verdict_type: type[VerdictT],
*,
inputs: type[InputT],
model: Model | str = DEFAULT_MODEL,
boolean_threshold: float | None = None,
min_confidence: float | None = None,
instructions: str | None = None,
name: str | None = None,
formatter: Formatter[InputT] | None = None,
tracer_provider: TracerProvider | None = None,
)
A decision model as an evaluator: fast, cheap typed answers with calibrated confidence.
It is a pydantic-ai Agent on a decision model (TypeSafe's Jev by default) whose output
type is the verdict type's decision-only view (decision_view): each field becomes a typed
question, and the answers carry probabilities. Fields the model cannot fill, such as free
text, keep their defaults; Fallback hands inputs to a language-model judge that fills
them.
It hands off (raises HandOff) when pydantic-ai does (DecisionHandOff), and when any
field's confidence is below min_confidence. Both thresholds, and pydantic-ai's
decision_boolean_threshold, are what ThresholdCalibration tunes.
Example
Build a decision evaluator.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
verdict_type
|
type[VerdictT]
|
The verdict it gives. |
required |
inputs
|
type[InputT]
|
The input it reads, rendered as the model's state by the formatter. |
required |
model
|
Model | str
|
A pydantic-ai decision model, or its name. Credentials are needed only
when it first runs ( |
DEFAULT_MODEL
|
boolean_threshold
|
float | None
|
The probability of yes at which a yes-or-no field is yes:
pydantic-ai's |
None
|
min_confidence
|
float | None
|
Hand off when any field's confidence is below this. |
None
|
instructions
|
str | None
|
Background the model reads with every question. |
None
|
name
|
str | None
|
The evaluator's name; |
None
|
formatter
|
Formatter[InputT] | None
|
Renders the input as the model's state; within 30,000 estimated tokens by default, under Jev's 32K limit. |
None
|
tracer_provider
|
TracerProvider | None
|
Where evaluation spans go; the global provider by default. |
None
|
Raises:
| Type | Description |
|---|---|
UnsupportedField
|
A field the model cannot fill has no default. |
ValueError
|
A threshold is out of range. |
version
property
¶
version: str
A hash of the model's name, the types, the instructions and the thresholds.
boolean_threshold
property
¶
boolean_threshold: float
The probability of yes at which a yes-or-no field is yes.
min_confidence
property
¶
min_confidence: float | None
The confidence below which it hands off, if any.
decide
async
¶
decide(input: InputT) -> Decision[VerdictT]
Ask the decision model, without handing off.
Raises:
| Type | Description |
|---|---|
DecisionHandOff
|
pydantic-ai handed the step off. |
Decision
dataclass
¶
Decision(
value: VerdictT,
confidence: dict[str, float],
probabilities: dict[str, float],
model: str | None,
cost: float | None,
)
What a decision model answered for one input, before any hand-off.
Attributes:
| Name | Type | Description |
|---|---|---|
value |
VerdictT
|
The verdict, with the fields the model cannot fill at their defaults. |
confidence |
dict[str, float]
|
The probability that each field's value is right, where the model gives one: a choice's probability, and a yes-or-no answer's probability of the answer given. |
probabilities |
dict[str, float]
|
For each yes-or-no field, the model's probability of yes, from which the answer follows by the boolean threshold. |
model |
str | None
|
The model that answered, as it named itself ( |
cost |
float | None
|
US dollars, when known. |
DEFAULT_MODEL
module-attribute
¶
TypeSafe's Jev, the latest version. Pin a version (typesafe:jev-1.13.0) once calibrated.
DEFAULT_BOOLEAN_THRESHOLD
module-attribute
¶
pydantic-ai's threshold for a yes, when none is set.
The decision-only view¶
decision_view
¶
decision_view(
verdict_type: type[BaseModel],
) -> DecisionView
Derive the fields of a verdict type that a decision model can fill.
Raises:
| Type | Description |
|---|---|
UnsupportedField
|
A field the view leaves out has no default, so no verdict could be made; give it a default, or judge the type with a language model. |
DecisionView
dataclass
¶
A verdict type as a decision model sees it.
Attributes:
| Name | Type | Description |
|---|---|---|
model |
type[BaseModel]
|
The output type the decision model fills, named |
decided |
tuple[str, ...]
|
The verdict fields the view has, in order. |
left_out |
tuple[str, ...]
|
The verdict fields it leaves out, which keep their defaults. |
MAX_CHOICES
module-attribute
¶
The most options a decision model's choice question offers.
Calibration¶
ThresholdCalibration
¶
ThresholdCalibration(
*,
target_agreement: float = 0.9,
grid: Sequence[float] = DEFAULT_GRID,
max_concurrency: int = 4,
)
Calibration as an Optimizer of decision evaluators.
The fitted evaluator records a Training whose scores are its agreement with people on
the validation examples it keeps, before and after, with the share it keeps (its coverage)
in results: a higher hand-off threshold keeps fewer examples, and the fallback judges
the rest.
Configure calibration.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
target_agreement
|
float
|
The agreement with people, from 0 to 1, the evaluator should reach on what it keeps. |
0.9
|
grid
|
Sequence[float]
|
Candidate thresholds, each strictly between 0 and 1. |
DEFAULT_GRID
|
max_concurrency
|
int
|
How many examples are asked at once. |
4
|
Raises:
| Type | Description |
|---|---|
ValueError
|
The target or a candidate threshold is out of range. |
settings
property
¶
The settings, as recorded on a calibrated evaluator.
optimize
async
¶
optimize(
evaluator: DecisionEvaluator[InputT, VerdictT],
/,
*,
train: Dataset[InputT, VerdictT],
validate: Dataset[InputT, VerdictT],
) -> DecisionEvaluator[InputT, VerdictT]
Fit the thresholds on train, and measure before and after on validate.
Examples the model hands off whatever the thresholds (pydantic-ai's own hand-offs) are left out of the fitting.