Skip to content

evalr.core

evalr's core, the hexagon: values, pure functions over them, and the ports.

The core depends on pydantic and the OpenTelemetry API only, and performs no I/O. The adapters (DSPy, decision models, Langfuse, Hugging Face, files, in-memory) implement its ports from their own packages (ADR-0006).

Verdicts

What an evaluator returns. See Typed verdicts and field kinds.

Verdict pydantic-model

Bases: BaseModel

An evaluator's judgement of one input: a value of the verdict type, and how it was made.

The value is an instance of the verdict type, which can be any Pydantic model, typically a feedback type people also give. Everything else records how far to trust it and where it came from, so verdicts from different evaluators, or different versions of one, never mix.

Attributes:

Name Type Description
value V

The verdict itself.

confidence dict[str, Confidence]

The probability that each field's value is right, for the fields the evaluator has one for: decision models report them, most judges do not.

evaluator str

The evaluator's name.

version str

The evaluator's version. A trained judge's version is a hash of its program.

latency float

Wall-clock seconds the evaluation took.

cost Annotated[float, Field(ge=0.0)] | None

What the evaluation cost in US dollars, when the evaluator knows.

trace_id Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None

The OpenTelemetry trace the evaluation ran in, as 32 hex digits, when there was one.

Fields:

Validators:

  • _confidence_is_per_field

Confidence

Confidence = Annotated[float, Field(ge=0.0, le=1.0)]

A probability that a field's value is right, from 0 to 1.

Verdict fields

How each field of a verdict type is judged and scored. See Field kinds.

FieldKind

Bases: StrEnum

How a verdict field is judged and scored.

BINARY class-attribute instance-attribute

BINARY = 'binary'

A bool: yes or no.

CATEGORICAL class-attribute instance-attribute

CATEGORICAL = 'categorical'

A Literal or Enum: one of a fixed set of choices.

ORDINAL class-attribute instance-attribute

ORDINAL = 'ordinal'

An int bounded on both sides: a rating on a scale.

NUMERIC class-attribute instance-attribute

NUMERIC = 'numeric'

Any other int or float.

TEXT class-attribute instance-attribute

TEXT = 'text'

A str: free text, which only language-model judges fill. It is scored as TEXT, but the agreement metrics do not compare it.

VerdictField dataclass

VerdictField(
    name: str,
    kind: FieldKind,
    description: str | None = None,
    choices: tuple[object, ...] = (),
    lower: float | None = None,
    upper: float | None = None,
    optional: bool = False,
    required: bool = True,
)

One field of a verdict type, as evaluators and metrics see it.

Attributes:

Name Type Description
name str

The field's name on the model.

kind FieldKind

How the field is judged and scored.

description str | None

The field's description, which judges read as an instruction.

choices tuple[object, ...]

The allowed values of a categorical field, in declaration order: the Literal's arguments or the Enum's members. (False, True) for a binary field, and empty for other kinds.

lower float | None

The inclusive lower bound of a numeric or ordinal field, if it has one. An ordinal field's exclusive bound is converted (gt=0 becomes 1); a float's is kept as is.

upper float | None

The upper bound, likewise.

optional bool

Whether the field accepts None.

required bool

Whether the model requires a value for the field.

compared property

compared: bool

Whether the agreement metrics compare the field: every kind but text.

verdict_fields

verdict_fields(
    verdict_type: type[BaseModel],
) -> tuple[VerdictField, ...]

Describe how each field of a verdict type is judged, in declaration order.

Parameters:

Name Type Description Default
verdict_type type[BaseModel]

Any Pydantic model.

required

Returns:

Type Description
tuple[VerdictField, ...]

One VerdictField per field of the model.

Raises:

Type Description
UnsupportedField

A field's type cannot be judged, such as a list or a nested model.

canonical_fields

canonical_fields(
    verdict_type: type[BaseModel],
) -> list[JsonValue]

How each field of a verdict type is judged, as JSON no Python or pydantic upgrade changes.

Each field is its name, kind, choices (an Enum's values), bounds and whether it may be empty. Evaluators hash it into their versions, so a change to how a field is judged changes them.

UnsupportedField

Bases: TypeError

A verdict field has a type that evaluators cannot judge.

Evaluators

The evaluator port, function evaluators and the hand-off composition. See Function evaluators and Hand-off and fallback.

Evaluator

Bases: Protocol

Judges an input and returns a typed verdict.

DSPy judges, decision evaluators and function evaluators all implement it, and so can an application's own. Evaluators are compared by name and version: a changed evaluator must change its version, so that verdicts from before and after never mix.

An evaluator that declines an input, for another to judge, raises HandOff.

name property

name: str

The evaluator's name, recorded on every verdict.

version property

version: str

The evaluator's version, recorded on every verdict.

verdict_type property

verdict_type: type[VerdictT]

The Pydantic model the evaluator's verdicts are instances of.

evaluate async

evaluate(input: InputT) -> Verdict[VerdictT]

Judge one input.

FunctionEvaluator

FunctionEvaluator(
    function: Callable[
        [InputT], VerdictT | Awaitable[VerdictT]
    ],
    *,
    verdict_type: type[VerdictT],
    name: str | None = None,
    version: str = "1",
    tracer_provider: TracerProvider | None = None,
)

An evaluator that is a function of its input: a deterministic measure.

The function may be sync or async. Its version is given, not derived: change it whenever the function's behaviour changes.

Example
def resolved(thread: Thread) -> Resolution:
    return Resolution(resolved=thread.messages[-1].author == "user")


evaluator = FunctionEvaluator(resolved, verdict_type=Resolution, version="2")
verdict = await evaluator.evaluate(thread)

Wrap a function as an evaluator.

Parameters:

Name Type Description Default
function Callable[[InputT], VerdictT | Awaitable[VerdictT]]

Computes the verdict's value from the input.

required
verdict_type type[VerdictT]

The Pydantic model the function returns.

required
name str | None

The evaluator's name; the function's name by default.

None
version str

The evaluator's version; bump it when the function changes.

'1'
tracer_provider TracerProvider | None

Where evaluation spans go; the global provider by default.

None

Raises:

Type Description
UnsupportedField

The verdict type has a field evaluators cannot judge.

name property

name: str

The evaluator's name.

version property

version: str

The evaluator's version.

verdict_type property

verdict_type: type[VerdictT]

The verdict type.

evaluate async

evaluate(input: InputT) -> Verdict[VerdictT]

Run the function on the input.

Raises:

Type Description
TypeError

The function returned something other than the verdict type.

Fallback

Fallback(
    primary: Evaluator[InputT, VerdictT],
    fallback: Evaluator[InputT, VerdictT],
    *,
    min_confidence: float | None = None,
    name: str | None = None,
    tracer_provider: TracerProvider | None = None,
)

Two evaluators of one verdict type: the primary judges, and hands off to the fallback.

The primary hands off by raising HandOff, or, when min_confidence is set, by giving a verdict with any field's confidence below it. The usual pairing is a fast, cheap decision model first and a language-model judge behind it:

evaluator = Fallback(DecisionEvaluator(...), DspyJudge(...), min_confidence=0.7)

Each verdict records the evaluator that actually gave it, so the two are measured apart. The composition runs in a span named evalr.fallback {name}, which records whether and why it handed off.

Compose two evaluators.

Parameters:

Name Type Description Default
primary Evaluator[InputT, VerdictT]

Judges first.

required
fallback Evaluator[InputT, VerdictT]

Judges what the primary hands off.

required
min_confidence float | None

Hand off a verdict with any field's confidence below this.

None
name str | None

The composition's name; {primary}+{fallback} by default.

None
tracer_provider TracerProvider | None

Where the composition's spans go; the global provider by default.

None

Raises:

Type Description
TypeError

The two evaluators' verdict types differ.

ValueError

min_confidence is outside [0, 1].

name property

name: str

The composition's name.

version property

version: str

A hash of both evaluators' names and versions, and the threshold.

verdict_type property

verdict_type: type[VerdictT]

The verdict type both evaluators give.

evaluate async

evaluate(input: InputT) -> Verdict[VerdictT]

Ask the primary, and the fallback if the primary hands off.

HandOff

Bases: Exception

An evaluator declines to judge an input, for another evaluator to judge instead.

A decision model raises it when it is unsure, or cannot fill a field; Fallback catches it and asks its fallback.

Examples and datasets

Inputs with what people said about them, and where they come from and go. See Datasets, splits and stores and Feedback sources.

Example pydantic-model

Bases: BaseModel

One input, with the verdict people gave it, a reference output, or both.

Attributes:

Name Type Description
id str

A stable identifier. Splits are hashed from it, and syncing uses it, so an example keeps its id for life: derive it from the feedback or item it came from.

input InputT

What an evaluator judges: a thread, a run, an artifact version.

verdict VerdictT | None

The verdict people gave the input, when there is one. Judges are trained and measured against it.

reference JsonValue

A reference output for the system being evaluated, such as the answer it should give, for evaluators that compare against one.

trace_id Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None

The OpenTelemetry trace the input came from, as 32 hex digits, so datasets and experiments link back to it.

metadata dict[str, JsonValue]

Anything else worth keeping with the example, as JSON.

Fields:

  • id (str)
  • input (InputT)
  • verdict (VerdictT | None)
  • reference (JsonValue)
  • trace_id (Annotated[str, Field(pattern='^[0-9a-f]{32}$')] | None)
  • metadata (dict[str, JsonValue])

Dataset

Dataset(
    name: str,
    examples: Iterable[Example[InputT, VerdictT]],
    *,
    input_type: type[InputT],
    verdict_type: type[VerdictT],
    description: str = "",
)

An immutable, named collection of examples of one input type and one verdict type.

Example
dataset = Dataset(
    "helpfulness",
    examples,
    input_type=Thread,
    verdict_type=Helpfulness,
)
train, validate = dataset.labelled().split(0.2)

Collect examples into a dataset.

Parameters:

Name Type Description Default
name str

The dataset's name, used when it is synced to Langfuse or Hugging Face.

required
examples Iterable[Example[InputT, VerdictT]]

The examples, in any order. Their ids must be unique.

required
input_type type[InputT]

The Pydantic model of every example's input.

required
verdict_type type[VerdictT]

The Pydantic model of every example's verdict.

required
description str

What the dataset holds, for people browsing it.

''

Raises:

Type Description
DuplicateExample

Two examples share an id.

UnsupportedField

The verdict type has a field evaluators cannot judge.

version cached property

version: str

A hash of the examples' content, independent of their order: 16 hex digits.

Any change to any example changes it, so a trained judge records the exact data it was trained on. Numbers are hashed by value, as JSON reads them (1.0 and 1 are one number), so a dataset keeps its version through stores that write whole numbers without a decimal point, as Langfuse does.

from_records staticmethod

from_records(
    name: str,
    records: Iterable[Mapping[str, JsonValue]],
    *,
    input_type: type[I],
    verdict_type: type[V],
    description: str = "",
) -> Dataset[I, V]

Load a dataset from JSON records, as written by records.

Raises:

Type Description
ValidationError

A record is not a valid example of these types.

records

records() -> list[dict[str, JsonValue]]

The examples as JSON records, in order: one object per example.

filter

filter(
    predicate: Callable[[Example[InputT, VerdictT]], bool],
) -> Self

The examples for which the predicate holds, as a dataset of the same name.

labelled

labelled() -> Self

The examples that have a verdict from people.

split

split(
    validate: float = 0.2, *, salt: str = ""
) -> tuple[Self, Self]

Split into training and validation examples, deterministically by id.

An example goes to validation when split_bucket(id, salt) < validate. So an example never moves between the two as examples are added or removed, and raising the fraction only moves examples from training to validation.

Parameters:

Name Type Description Default
validate float

The expected share of examples in validation, from 0 to 1.

0.2
salt str

Gives an independent split of the same examples.

''

Returns:

Type Description
tuple[Self, Self]

The training and validation datasets, each keeping this dataset's order.

Raises:

Type Description
ValueError

The fraction is outside [0, 1].

split_bucket

split_bucket(example_id: str, salt: str = '') -> float

Place an example in [0, 1) by a hash of its id.

The position depends on nothing but the id and the salt, so it is the same in every process and every run, and adding or removing other examples never moves it.

Parameters:

Name Type Description Default
example_id str

The example's id.

required
salt str

Changes every position at once, for an independent split of the same examples.

''

DatasetStore

Bases: Protocol

Saves datasets under their names, and loads them by name and revision.

Every save makes a revision that can be loaded later, even after further saves, so an experiment or a trained judge can name the exact data it used. Saving the same content again changes nothing a load can see.

save async

save(dataset: Dataset[InputT, VerdictT]) -> str

Store the dataset under its name, replacing what a load of the name returns.

Returns:

Type Description
str

The revision just saved, which load accepts.

load async

load(
    name: str,
    /,
    *,
    input_type: type[InputT],
    verdict_type: type[VerdictT],
    revision: str | None = None,
) -> Dataset[InputT, VerdictT]

Load a dataset, validating its examples as the given types.

Parameters:

Name Type Description Default
name str

The dataset's name.

required
input_type type[InputT]

The Pydantic model of the examples' inputs.

required
verdict_type type[VerdictT]

The Pydantic model of the examples' verdicts.

required
revision str | None

A revision save returned; the latest when None.

None

Raises:

Type Description
DatasetNotFound

No dataset has the name, or it has no such revision.

ValidationError

The stored examples are not of the given types.

DatasetNotFound

Bases: LookupError

A store has no dataset of that name, or no such revision of it.

DuplicateExample

Bases: ValueError

Two examples in one dataset share an id.

FeedbackSource

Bases: Protocol

Yields examples from people's typed feedback.

The libraries implement it in their [evals] extras, turning their feedback of one type, with the context of its target (a thread, a run), into examples; evalr never imports them. Every example has a verdict, its id is stable (derived from the feedback), and iterating again yields the same examples.

input_type property

input_type: type[InputT]

The Pydantic model of the examples' inputs.

verdict_type property

verdict_type: type[VerdictT]

The Pydantic model of the examples' verdicts: the feedback type.

examples

examples() -> AsyncIterator[Example[InputT, VerdictT]]

Yield the examples.

collect async

collect(
    name: str,
    source: FeedbackSource[InputT, VerdictT],
    *,
    description: str = "",
) -> Dataset[InputT, VerdictT]

Gather a feedback source's examples into a dataset.

Parameters:

Name Type Description Default
name str

The dataset's name.

required
source FeedbackSource[InputT, VerdictT]

Where the feedback comes from, such as a library's [evals] extra.

required
description str

What the dataset holds.

''

Raises:

Type Description
DuplicateExample

The source yielded two examples with one id.

Formatters

The text a judge reads for an input, within a token budget. See What a judge reads.

Formatter

Bases: Protocol

Anything that renders an input as the text a judge reads.

InputFormatter dataclass

InputFormatter(
    max_tokens: int = 30000,
    count_tokens: TokenCounter = estimate_tokens,
)

Renders any Pydantic model within a token budget.

Each field is rendered as a section headed by its name and description. Strings are used as they are, lists one item to a line, and anything else as JSON. When the whole exceeds the budget, the longest list loses its oldest items (after the first, which is usually the request), and then the longest remaining text is shortened in the middle, until it fits.

Attributes:

Name Type Description
max_tokens int

The budget for the rendered text.

count_tokens TokenCounter

How to count tokens; a conservative estimate by default.

fields

fields(input: BaseModel) -> dict[str, str]

Render each field of the input separately, within the budget together.

Judges that read an input field by field (such as a DSPy signature with one input field per field of the model) use this instead of the single text.

TokenCounter

TokenCounter = Callable[[str], int]

Counts the tokens in a text, for a particular model's tokenizer.

estimate_tokens

estimate_tokens(text: str) -> int

Estimate a text's tokens without a tokenizer: one per three bytes of UTF-8, rounded up.

This overestimates for English prose (about four characters per token) and is close for text in scripts that take three bytes a character, so a budget measured with it is rarely exceeded. Pass a real tokenizer's counter to a formatter when the limit is exact.

Scores

Verdicts and people's feedback as named values, how each field is scored, and the ports they leave through. artifactr and reflexr build their feedback mirrors on them (ADR-0011). See Scores and score sinks.

Score pydantic-model

Bases: BaseModel

One field of a verdict, or of a piece of people's feedback, as a named value.

A score is attached to a trace (and, within it, a span) or to a session. evalr's scores come from verdicts and name the evaluator; the libraries' feedback mirrors make scores from people's feedback, with no evaluator, and describe where it came from in source.

Attributes:

Name Type Description
id str

Derived from what the score is about and its name, so a sink that upserts by id keeps one score per field however often it is recorded.

name str

{type}.{field}.

value bool | float | str

A bool for BOOLEAN, a float for NUMERIC, and a string for CATEGORICAL and TEXT; a value of another type is refused.

data_type ScoreDataType

How the value is to be read.

trace_id str | None

The trace the score is attached to, when there is one.

span_id Annotated[str, Field(pattern='^[0-9a-f]{16}$')] | None

The span within that trace the score judges, as 16 hex digits, when known; only with a trace.

session_id str | None

The session the score is attached to, such as a thread or a causal chain, when it is about the session rather than one trace.

timestamp AwareDatetime | None

When the score was given, with its time zone; when it is recorded, if None.

evaluator str | None

The evaluator that gave the verdict; None for people's feedback.

version str | None

The evaluator's version.

confidence float | None

The evaluator's confidence in the field's value, when it has one.

source Mapping[str, str]

Where the score came from, beyond an evaluator: the tenant, workspace and person that gave a piece of feedback, for example. Its keys cannot be evaluator, version or confidence, which metadata adds.

Fields:

Validators:

  • _fields_agree

metadata property

metadata: dict[str, JsonValue]

The source, then the evaluator, its version and the confidence, as score metadata.

ScoreDataType

ScoreDataType = Literal[
    "NUMERIC", "BOOLEAN", "CATEGORICAL", "TEXT"
]

A score's data type, as Langfuse and the OpenTelemetry conventions name them.

ScoreSink

Bases: Protocol

Records scores: verdicts and feedback as named values, on the traces or sessions they judge.

Recording is idempotent by score id: recording a score again replaces it. evalr's online evaluation records verdicts' scores through it, and the libraries' feedback mirrors people's feedback.

record async

record(scores: Sequence[Score]) -> None

Record the scores.

scores

Scores: verdicts and people's feedback as the named values observability backends record.

Every field of a verdict, or of a piece of feedback, becomes one score named {type}.{field}. The mapping is the one artifactr and reflexr use for feedback, and they build on it (ADR-0011). A field's kind decides the score's data type (score_configs):

  • binary fields are BOOLEAN
  • ordinal and numeric fields are NUMERIC
  • categorical fields are CATEGORICAL, with the choice as a string
  • text fields are TEXT

A field left empty (None or "") gives no score, and a category or text longer than MAX_TEXT characters is cut. A verdict's scores have ids derived from what the verdict is about, the evaluator and the field, so recording a verdict again replaces its scores rather than adding more.

SCORE_NAMESPACE module-attribute

SCORE_NAMESPACE = uuid.uuid5(
    uuid.NAMESPACE_URL,
    "https://github.com/alexnodeland/evalr/scores",
)

The namespace of the ids scores gives.

MAX_TEXT module-attribute

MAX_TEXT = 500

The longest category or text a score holds, in characters; longer ones are cut.

Score pydantic-model

Bases: BaseModel

One field of a verdict, or of a piece of people's feedback, as a named value.

A score is attached to a trace (and, within it, a span) or to a session. evalr's scores come from verdicts and name the evaluator; the libraries' feedback mirrors make scores from people's feedback, with no evaluator, and describe where it came from in source.

Attributes:

Name Type Description
id str

Derived from what the score is about and its name, so a sink that upserts by id keeps one score per field however often it is recorded.

name str

{type}.{field}.

value bool | float | str

A bool for BOOLEAN, a float for NUMERIC, and a string for CATEGORICAL and TEXT; a value of another type is refused.

data_type ScoreDataType

How the value is to be read.

trace_id str | None

The trace the score is attached to, when there is one.

span_id Annotated[str, Field(pattern='^[0-9a-f]{16}$')] | None

The span within that trace the score judges, as 16 hex digits, when known; only with a trace.

session_id str | None

The session the score is attached to, such as a thread or a causal chain, when it is about the session rather than one trace.

timestamp AwareDatetime | None

When the score was given, with its time zone; when it is recorded, if None.

evaluator str | None

The evaluator that gave the verdict; None for people's feedback.

version str | None

The evaluator's version.

confidence float | None

The evaluator's confidence in the field's value, when it has one.

source Mapping[str, str]

Where the score came from, beyond an evaluator: the tenant, workspace and person that gave a piece of feedback, for example. Its keys cannot be evaluator, version or confidence, which metadata adds.

Fields:

Validators:

  • _fields_agree
metadata property
metadata: dict[str, JsonValue]

The source, then the evaluator, its version and the confidence, as score metadata.

score_values

score_values(
    verdict_type: type[BaseModel],
    value: Mapping[str, object],
    *,
    type_name: str | None = None,
) -> list[tuple[ScoreConfig, bool | float | str]]

Pair each field of a validated value that has a value with its config and score value.

The libraries keep feedback as JSON, so value may be a model's JSON (an Enum as its value) as well as its fields.

Parameters:

Name Type Description Default
verdict_type type[BaseModel]

The value's type, which score_configs reads.

required
value Mapping[str, object]

The value's fields, by name, validated as verdict_type.

required
type_name str | None

The {type} in the scores' names; score_type_name of the type by default.

None

Returns:

Type Description
list[tuple[ScoreConfig, bool | float | str]]

A bool for each BOOLEAN score, a float for each NUMERIC one, and a string, cut at

list[tuple[ScoreConfig, bool | float | str]]

MAX_TEXT characters, for each CATEGORICAL and TEXT one, in the order of the

list[tuple[ScoreConfig, bool | float | str]]

type's fields. Fields that are None, "" or missing are left out.

scores

scores(
    verdict: Verdict[BaseModel],
    *,
    type_name: str | None = None,
    subject: str | None = None,
    trace_id: str | None = None,
    span_id: str | None = None,
) -> list[Score]

Turn a verdict into one score per field that has a value.

Parameters:

Name Type Description Default
verdict Verdict[BaseModel]

The verdict.

required
type_name str | None

The {type} in the scores' names; score_type_name of the verdict type by default. Pass a library's registered feedback name where it differs.

None
subject str | None

What the verdict is about, such as a run or an example id. It keys the scores' ids, so pass it whenever the verdict has no trace.

None
trace_id str | None

The trace to attach the scores to; the verdict's own by default.

None
span_id str | None

The span the verdict judges, within that trace, when known.

None

Returns:

Type Description
list[Score]

The scores, in the order of the verdict type's fields, with values as score_values

list[Score]

gives them.

score_values

score_values(
    verdict_type: type[BaseModel],
    value: Mapping[str, object],
    *,
    type_name: str | None = None,
) -> list[tuple[ScoreConfig, bool | float | str]]

Pair each field of a validated value that has a value with its config and score value.

The libraries keep feedback as JSON, so value may be a model's JSON (an Enum as its value) as well as its fields.

Parameters:

Name Type Description Default
verdict_type type[BaseModel]

The value's type, which score_configs reads.

required
value Mapping[str, object]

The value's fields, by name, validated as verdict_type.

required
type_name str | None

The {type} in the scores' names; score_type_name of the type by default.

None

Returns:

Type Description
list[tuple[ScoreConfig, bool | float | str]]

A bool for each BOOLEAN score, a float for each NUMERIC one, and a string, cut at

list[tuple[ScoreConfig, bool | float | str]]

MAX_TEXT characters, for each CATEGORICAL and TEXT one, in the order of the

list[tuple[ScoreConfig, bool | float | str]]

type's fields. Fields that are None, "" or missing are left out.

score_type_name

score_type_name(verdict_type: type[BaseModel]) -> str

The name scores give a verdict type: its class name in snake case.

Helpfulness is helpfulness and TaskCompletion is task_completion, as the libraries name their feedback types by default.

SCORE_NAMESPACE module-attribute

SCORE_NAMESPACE = uuid.uuid5(
    uuid.NAMESPACE_URL,
    "https://github.com/alexnodeland/evalr/scores",
)

The namespace of the ids scores gives.

MAX_TEXT module-attribute

MAX_TEXT = 500

The longest category or text a score holds, in characters; longer ones are cut.

Score configs

How each field of a type is scored, and the port that keeps it. See Score configs.

ScoreConfig dataclass

ScoreConfig(
    name: str,
    type_name: str,
    field: str,
    data_type: ScoreDataType,
    description: str | None = None,
    minimum: float | None = None,
    maximum: float | None = None,
    categories: tuple[str, ...] = (),
)

How one field of a verdict or feedback type is scored, for a backend to read its scores by.

A backend that knows a score's config can check its values and offer its choices to people scoring by hand. sync_score_configs creates them in a ScoreConfigStore.

Attributes:

Name Type Description
name str

The score's name, {type}.{field}.

type_name str

The {type}: the type's score name, or the name a library registers its feedback type under.

field str

The field's name.

data_type ScoreDataType

How the field is scored.

description str | None

The field's description, if it has one.

minimum float | None

The lowest value of a numeric field, if it is bounded below. The bound is kept as declared: gt=0 is 0, even for an int, where VerdictField.lower is 1.

maximum float | None

The highest value of a numeric field, if it is bounded above, likewise.

categories tuple[str, ...]

A categorical field's choices as strings, in declaration order: the Literal's arguments or the Enum's values. Empty for other kinds.

score_configs

score_configs(
    verdict_type: type[BaseModel],
    *,
    type_name: str | None = None,
) -> tuple[ScoreConfig, ...]

Describe how each field of a verdict or feedback type is scored, in declaration order.

Each field is scored by its kind, as verdict_fields reads it: binary fields are BOOLEAN, ordinal and numeric ones NUMERIC, categorical ones CATEGORICAL and text ones TEXT. Unlike verdict_fields, a field of a type that cannot be judged (a list, a nested model, a union of several types) raises nothing: it is not scored. A library's feedback type may have such fields, where a verdict type may not.

Parameters:

Name Type Description Default
verdict_type type[BaseModel]

Any Pydantic model.

required
type_name str | None

The {type} in the scores' names; score_type_name of the type by default. Pass a library's registered feedback name where it differs.

None

Returns:

Type Description
tuple[ScoreConfig, ...]

One config per field that can be scored.

ScoreConfigStore

Bases: Protocol

Keeps score configs, so a backend knows each score's data type, range and choices.

Configs are found by name and only ever created: sync_score_configs creates the ones a store lacks, and leaves the others as they are.

names async

names() -> Collection[str]

Return the names of the configs the store has, including any it has archived.

create async

create(config: ScoreConfig) -> None

Create a config whose name the store does not have.

sync_score_configs async

sync_score_configs(
    store: ScoreConfigStore, configs: Iterable[ScoreConfig]
) -> list[str]

Create the configs whose names a store does not have yet.

A config the store has by name is left as it is, even if its definition differs, so a backend's links from scores to configs never break:

await sync_score_configs(store, score_configs(Helpfulness))

Parameters:

Name Type Description Default
store ScoreConfigStore

Where the configs live.

required
configs Iterable[ScoreConfig]

The configs to have, such as score_configs of each verdict or feedback type.

required

Returns:

Type Description
list[str]

The names of the configs created, in the order given. A name given twice is created once.

Experiments

A task run over a dataset, with every output judged. See Experiments.

ExperimentTracker

Bases: Protocol

Runs a task over a dataset, judges every output, and keeps the results.

A failing task or evaluator fails only its own item, which records the error; the run goes on. Each example runs in its own trace, which the verdicts record.

run_experiment async

run_experiment(
    name: str,
    /,
    *,
    dataset: Dataset[InputT, VerdictT],
    task: Task[InputT, VerdictT, OutputT],
    evaluators: Sequence[Evaluator[OutputT, BaseModel]],
    run_name: str | None = None,
    max_concurrency: int = 4,
    metadata: Mapping[str, str] | None = None,
) -> ExperimentResult[OutputT]

Run the task on every example, and judge each output with every evaluator.

Parameters:

Name Type Description Default
name str

The experiment's name, shared by its runs.

required
dataset Dataset[InputT, VerdictT]

The examples to run.

required
task Task[InputT, VerdictT, OutputT]

The system being evaluated.

required
evaluators Sequence[Evaluator[OutputT, BaseModel]]

What judges each output.

required
run_name str | None

This run's name; the tracker chooses one by default.

None
max_concurrency int

How many examples run at once.

4
metadata Mapping[str, str] | None

Anything to keep with the run, such as the model or prompt under test.

None

Returns:

Type Description
ExperimentResult[OutputT]

One item per example, in the dataset's order.

Task

Task = Callable[
    [Example[InputT, VerdictT]], Awaitable[OutputT]
]

What an experiment runs for each example: the system being evaluated, producing an output.

The output is what the experiment's evaluators judge, so it holds whatever they need, such as the request and the new reply. A task that judges the example's input as it is (to measure an evaluator against people's verdicts) returns the input.

ExperimentResult dataclass

ExperimentResult(
    name: str,
    run_name: str,
    dataset: str,
    dataset_version: str,
    items: tuple[ItemResult[OutputT], ...],
    url: str | None = None,
)

An experiment's results: one item per example of the dataset, in the dataset's order.

Attributes:

Name Type Description
name str

The experiment's name, shared by its runs.

run_name str

This run's name.

dataset str

The dataset's name.

dataset_version str

The dataset's content hash.

items tuple[ItemResult[OutputT], ...]

One result per example.

url str | None

Where the tracker shows the run, when it has a UI.

verdicts

verdicts(evaluator: str) -> dict[str, Verdict[BaseModel]]
verdicts(
    evaluator: str, *, verdict_type: type[V]
) -> dict[str, Verdict[V]]
verdicts(
    evaluator: str,
    *,
    verdict_type: type[BaseModel] = BaseModel,
) -> Mapping[str, Verdict[BaseModel]]

One evaluator's verdicts, by example id.

An experiment can have evaluators of several verdict types, so verdicts are typed as BaseModel unless the evaluator's type is given:

verdicts = result.verdicts("helpfulness-judge", verdict_type=Helpfulness)
ratings = [verdict.value.rating for verdict in verdicts.values()]

Parameters:

Name Type Description Default
evaluator str

The evaluator's name, as its verdicts record it.

required
verdict_type type[BaseModel]

The verdict type it gives. Each verdict is checked to be one.

BaseModel

Raises:

Type Description
TypeError

One of the evaluator's verdicts is not of verdict_type.

ItemResult dataclass

ItemResult(
    example_id: str,
    output: OutputT | None,
    verdicts: tuple[Verdict[BaseModel], ...] = (),
    errors: tuple[str, ...] = (),
    trace_id: str | None = None,
)

What happened to one example in an experiment.

Attributes:

Name Type Description
example_id str

The example's id.

output OutputT | None

The task's output, or None when the task failed.

verdicts tuple[Verdict[BaseModel], ...]

One verdict per evaluator that succeeded, in the evaluators' order.

errors tuple[str, ...]

What failed: the task, or an evaluator by name, with its message.

trace_id str | None

The trace the example ran in, when there was one.

Measuring and optimizing

An evaluator against people's verdicts, and fitted to them. See Agreement and calibration metrics and DSPy judges and GEPA.

measure async

measure(
    evaluator: Evaluator[InputT, VerdictT],
    dataset: Dataset[InputT, VerdictT],
    *,
    max_concurrency: int = 4,
) -> Measurement[VerdictT]

Judge every labelled example of a dataset, and compare with people's verdicts.

An example the evaluator fails on (or hands off) counts as a missing prediction.

Parameters:

Name Type Description Default
evaluator Evaluator[InputT, VerdictT]

The evaluator to measure.

required
dataset Dataset[InputT, VerdictT]

Examples with people's verdicts; unlabelled ones are skipped.

required
max_concurrency int

How many examples are judged at once.

4

Measurement dataclass

Measurement(
    evaluator: str,
    version: str,
    dataset: str,
    dataset_version: str,
    verdicts: dict[str, Verdict[VerdictT]],
    errors: dict[str, str],
    agreement: Agreement,
    calibration: dict[str, FieldCalibration],
    stats: list[EvaluatorStats],
)

How an evaluator did on a dataset's labelled examples.

Attributes:

Name Type Description
evaluator str

The evaluator's name.

version str

Its version.

dataset str

The dataset's name.

dataset_version str

The dataset's content hash.

verdicts dict[str, Verdict[VerdictT]]

The verdicts it gave, by example id.

errors dict[str, str]

What went wrong where it gave none, by example id.

agreement Agreement

Its agreement with people's verdicts.

calibration dict[str, FieldCalibration]

How well its confidence predicts being right, per field.

stats list[EvaluatorStats]

Latency and cost, per evaluator version that gave verdicts (a composition's verdicts come from more than one).

Optimizer

Bases: Protocol

Fits an evaluator to people's verdicts on a training set, checked on a validation set.

GEPA fits a DSPy judge's instructions; threshold calibration fits a decision evaluator's thresholds. The evaluator given is not changed: the fitted one is returned, with a new version if it differs.

optimize async

optimize(
    evaluator: EvaluatorT,
    /,
    *,
    train: Dataset[InputT, VerdictT],
    validate: Dataset[InputT, VerdictT],
) -> EvaluatorT

Fit the evaluator.

Parameters:

Name Type Description Default
evaluator EvaluatorT

The evaluator to start from.

required
train Dataset[InputT, VerdictT]

Labelled examples to fit on.

required
validate Dataset[InputT, VerdictT]

Labelled examples, none of them in train, to check the fit on.

required

optimize async

optimize(
    evaluator: EvaluatorT,
    *,
    train: Dataset[InputT, VerdictT],
    validate: Dataset[InputT, VerdictT],
    optimizer: Optimizer[InputT, VerdictT, EvaluatorT],
) -> EvaluatorT

Fit an evaluator to people's verdicts with an optimizer.

GEPA fits a DSPy judge; threshold calibration fits a decision evaluator. Only labelled examples are used, and no example may be in both sets, so the validation score is honest. A small dataset can split with nothing on one side; until it has enough examples to split, measure the evaluator on all of them instead.

Parameters:

Name Type Description Default
evaluator EvaluatorT

The evaluator to start from. It is not changed.

required
train Dataset[InputT, VerdictT]

Examples to fit on.

required
validate Dataset[InputT, VerdictT]

Examples to check the fit on.

required
optimizer Optimizer[InputT, VerdictT, EvaluatorT]

How to fit this kind of evaluator.

required

Returns:

Type Description
EvaluatorT

The fitted evaluator, with a new version if it changed.

Raises:

Type Description
ValueError

A set has no labelled examples, or the sets share an example.

Training pydantic-model

Bases: BaseModel

How a fitted evaluator was fitted, kept with it so its version can be explained.

Attributes:

Name Type Description
optimizer str

The optimizer's name, such as gepa or thresholds.

settings dict[str, JsonValue]

The optimizer's settings, as JSON.

base_version str

The version of the evaluator it started from.

train DatasetRef

The data it was fitted on.

validation DatasetRef

The data it was checked on.

score_before float | None

Agreement with people on the validation data before fitting, from 0 to 1.

score_after float | None

Agreement after fitting.

results dict[str, JsonValue]

What the optimizer found, as JSON, such as calibrated thresholds.

Fields:

DatasetRef pydantic-model

Bases: BaseModel

Which data was used: a dataset's name, content hash and size.

Fields:

of classmethod

of(dataset: Dataset[InputT, VerdictT]) -> Self

Refer to a dataset.

Metrics

Agreement with people, calibration, and cost and latency. See Agreement and calibration metrics.

agreement

agreement(
    expected: Sequence[V],
    predicted: Sequence[V | None],
    *,
    verdict_type: type[V],
) -> Agreement

Compare predicted verdicts with the expected ones, pair by pair.

Parameters:

Name Type Description Default
expected Sequence[V]

People's verdicts.

required
predicted Sequence[V | None]

The evaluator's verdicts for the same inputs, in the same order; None where it gave none.

required
verdict_type type[V]

The verdict type, whose fields are compared.

required

Raises:

Type Description
ValueError

The sequences differ in length.

Agreement pydantic-model

Bases: BaseModel

How far an evaluator agrees with people over a set of verdicts.

Attributes:

Name Type Description
n int

The pairs compared.

score float | None

The mean agreement_score over pairs that have one, from 0 to 1.

fields dict[str, FieldAgreement]

Per compared field, in the verdict type's order.

Fields:

FieldAgreement pydantic-model

Bases: BaseModel

Agreement on one field, over the pairs whose expected value is set.

Attributes:

Name Type Description
field str

The field's name.

kind FieldKind

The field's kind, which decides the measures.

n int

Pairs whose expected value is set.

missing int

Of those, pairs with no predicted value.

score float | None

The mean of the per-pair agreement, from 0 to 1.

accuracy float | None

For binary and categorical fields; a missing prediction is wrong.

kappa float | None

Cohen's kappa, for binary and categorical fields.

mean_absolute_error float | None

For ordinal and numeric fields, over the pairs with a prediction.

spearman float | None

Spearman's rank correlation, likewise.

Fields:

agreement_score

agreement_score(
    expected: V, predicted: V | None
) -> float | None

The mean of field_agreement: one number from 0 to 1 per pair of verdicts.

None when the expected verdict has no compared field with a value.

field_agreement

field_agreement(
    expected: V, predicted: V | None
) -> dict[str, float]

How far a predicted verdict agrees with the expected one, per compared field, from 0 to 1.

Only fields with an expected value count. Binary and categorical fields agree fully or not at all. Ordinal and bounded numeric fields lose agreement in proportion to the distance over their range (1 - |e - p| / (upper - lower)); unbounded ones as 1 / (1 + |e - p|). A missing prediction (the whole verdict, or the field) agrees not at all.

calibration

calibration(
    expected: Sequence[V],
    verdicts: Sequence[Verdict[V] | None],
    *,
    verdict_type: type[V],
    bins: int = 10,
) -> dict[str, FieldCalibration]

Measure how well each field's confidence predicts that its value is right.

Parameters:

Name Type Description Default
expected Sequence[V]

People's verdicts.

required
verdicts Sequence[Verdict[V] | None]

The evaluator's verdicts for the same inputs, in the same order; None where it gave none.

required
verdict_type type[V]

The verdict type, whose fields are measured.

required
bins int

The bins of the expected calibration error.

10

Returns:

Type Description
dict[str, FieldCalibration]

One entry per field that any verdict has a confidence for, in the verdict type's order.

Raises:

Type Description
ValueError

The sequences differ in length.

FieldCalibration pydantic-model

Bases: BaseModel

Calibration of one field's confidence, over the pairs that have one.

Attributes:

Name Type Description
field str

The field's name.

n int

Pairs with an expected value and a confidence.

accuracy float | None

The share of those whose predicted value was right.

confidence float | None

The mean confidence.

expected_calibration_error float | None

See expected_calibration_error.

brier_score float | None

See brier_score.

Fields:

  • field (str)
  • n (int)
  • accuracy (float | None)
  • confidence (float | None)
  • expected_calibration_error (float | None)
  • brier_score (float | None)

evaluator_stats

evaluator_stats(
    verdicts: Iterable[Verdict[BaseModel]],
) -> list[EvaluatorStats]

Latency and cost per evaluator version, sorted by name and version.

EvaluatorStats pydantic-model

Bases: BaseModel

Latency and cost of one version of one evaluator.

Attributes:

Name Type Description
evaluator str

The evaluator's name.

version str

Its version.

n int

Verdicts measured.

mean_latency float

Mean seconds per verdict.

p50_latency float

Median seconds, by nearest rank.

p95_latency float

95th percentile seconds, by nearest rank.

total_cost float | None

US dollars over the verdicts that report a cost; None if none do.

mean_cost float | None

The mean over those verdicts.

Fields:

accuracy

accuracy(
    expected: Sequence[Hashable],
    predicted: Sequence[Hashable],
) -> float | None

The share of pairs that are equal; None for no pairs.

cohen_kappa

cohen_kappa(
    expected: Sequence[Hashable],
    predicted: Sequence[Hashable],
) -> float | None

Cohen's kappa: agreement beyond what the two sides' label frequencies give by chance.

1 is perfect agreement and 0 is chance. None for no pairs, or when chance agreement is already perfect (both sides always give the one same label).

mean_absolute_error

mean_absolute_error(
    expected: Sequence[float], predicted: Sequence[float]
) -> float | None

The mean absolute difference between pairs; None for no pairs.

spearman

spearman(
    expected: Sequence[float], predicted: Sequence[float]
) -> float | None

Spearman's rank correlation, with tied values given their average rank.

None for fewer than two pairs, or when either side is constant.

brier_score

brier_score(
    confidences: Sequence[float], correct: Sequence[bool]
) -> float | None

The mean squared difference between confidence and correctness; None for no pairs.

0 is perfect; always answering with confidence 0.5 scores 0.25.

expected_calibration_error

expected_calibration_error(
    confidences: Sequence[float],
    correct: Sequence[bool],
    *,
    bins: int = 10,
) -> float | None

The gap between confidence and accuracy, averaged over bins of equal width.

Confidences fall into bins intervals of [0, 1], each closed below and open above, except the last, which includes 1. The result weights each bin's gap between its mean confidence and its accuracy by its share of the pairs. None for no pairs.

Raises:

Type Description
ValueError

bins is less than 1, or a confidence is outside [0, 1].

Tracing

OpenTelemetry spans for evaluations, through the API only. See Your own evaluators.

judging

judging(
    tracer: Tracer,
    *,
    evaluator: str,
    version: str,
    verdict_type: type[BaseModel],
) -> Generator[Judging]

Run an evaluation in a span named evalr.evaluate {evaluator}.

The span is current inside the block, so the spans of whatever the evaluator calls (a language model, an agent) nest under it. An exception is recorded on the span and re-raised.

Parameters:

Name Type Description Default
tracer Tracer

evalr's tracer, from get_tracer.

required
evaluator str

The evaluator's name.

required
version str

The evaluator's version.

required
verdict_type type[BaseModel]

The verdict type the evaluator returns.

required

Yields:

Type Description
Generator[Judging]

The evaluation in progress, whose verdict builds the result.

Judging dataclass

Judging(
    evaluator: str, version: str, span: Span, started: float
)

One evaluation in progress: its span and its clock.

Evaluators build their verdict through verdict, which stamps it with the evaluator, the latency so far and the trace id.

verdict

verdict(
    value: V,
    *,
    confidence: dict[str, float] | None = None,
    cost: float | None = None,
) -> Verdict[V]

Wrap a value in a verdict from this evaluation.

Parameters:

Name Type Description Default
value V

The verdict type's instance.

required
confidence dict[str, float] | None

The probability that each field's value is right, where known.

None
cost float | None

What the evaluation cost in US dollars, where known.

None

get_tracer

get_tracer(
    tracer_provider: TracerProvider | None = None,
) -> Tracer

Return evalr's tracer from the given provider, or the global one.

Parameters:

Name Type Description Default
tracer_provider TracerProvider | None

The provider to use. None uses the global provider, which is a no-op until the application configures the SDK.

None

current_trace_id

current_trace_id() -> str | None

Return the current span's trace id as 32 hex digits, or None outside any trace.

SCOPE module-attribute

SCOPE = 'evalr'

The instrumentation scope of evalr's spans.