Skip to content

RFC-0001: v0.1 implementation plan

Status: Implemented Author: Alex Nodeland Created: 2026-09-28 Discussion: accepted on 2026-09-28 Siblings: artifactr RFC-0002 and reflexr RFC-0002 define the typed feedback evalr learns from, and the [evals] extras that use it. stackr RFC-0001 runs the infrastructure.

Summary

evalr is a Python library for typed evaluation of agent systems. An evaluator judges an input (a chat thread, a workflow run, an artifact version) and returns a verdict: an instance of a Pydantic type, typically one of the feedback types people also give. Because people and evaluators produce the same types, evaluators can be trained on people's feedback and measured against it.

Two kinds of evaluator are equal citizens:

  • DSPy judges: language-model judges whose signatures are derived from the verdict type, optimized with GEPA on datasets seeded by feedback.
  • Decision models: TypeSafe's Jev, through pydantic-ai's decision-model support. It answers the verdict's fields as typed questions with calibrated probabilities, and hands a field it is unsure of to a language model.

Around them, evalr provides:

  • datasets built from feedback, and synced to Langfuse and Hugging Face
  • experiments run in Langfuse
  • agreement and calibration metrics
  • end-to-end measures
  • OpenTelemetry spans

evalr imports neither artifactr nor reflexr. The libraries depend on it through their [evals] extras.

Motivation

artifactr and reflexr record typed feedback from people. Turning that into better agents needs:

  • Offline evaluators that predict the feedback people would give, so every change (a prompt, a model, a tool) can be scored before people see it.
  • Online evaluators cheap enough to run on live traffic, so quality is measured continuously, not just sampled by people.
  • Datasets and experiments that are versioned, reproducible and visible next to the traces they judge.

Writing this in each library would duplicate it, and putting DSPy, Hugging Face and Langfuse in the libraries' core would burden users who never evaluate.

Design

Types

class Helpfulness(BaseModel):  # in the application, or a library's Feedback subclass
    rating: Annotated[int, Field(ge=1, le=5, description="How much the reply helped")]
    resolved: bool = Field(description="The request was fully addressed")
    reason: str | None = None


judge = DspyJudge(Helpfulness, inputs=ThreadTranscript)  # a language-model judge
decider = DecisionEvaluator(Helpfulness, inputs=ThreadTranscript)  # Jev, with an LLM fallback
verdict = await decider.evaluate(transcript)  # Verdict[Helpfulness]
verdict.value.rating, verdict.confidence["resolved"]
  • Evaluator[InputT, VerdictT] is a protocol: async evaluate(input) -> Verdict[VerdictT]. A Verdict wraps the value with:
    • per-field confidence, where the evaluator has it
    • the evaluator's name and version
    • latency and cost
    • the trace id
  • Verdict types are any Pydantic model, so evalr needs nothing from the libraries. Field types decide how a field is judged and scored:
    • bounded numbers are ordinal or numeric
    • bool is yes/no
    • Literal and Enum are categorical
    • str is free text: only language-model judges fill it
  • Inputs are Pydantic models too. Formatters turn them into the text a judge reads, within a budget. Jev's state is limited to 32K tokens, so long threads are summarized or windowed.

Evaluator kinds

Kind How Strengths
DspyJudge A DSPy module whose signature comes from the input and verdict types, with instructions from field descriptions Reasons in free text; optimizable with GEPA
DecisionEvaluator A pydantic-ai agent on a decision model (typesafe:jev-latest) with output_type set to the verdict type; FallbackModel hands unsure or unfillable fields to a language-model judge Fast and cheap, with calibrated probabilities; the fallback keeps coverage
FunctionEvaluator A pure function of the input Deterministic measures, such as rewrite rate

Training judges with GEPA

  • optimize(judge, train=..., validate=..., optimizer=GEPA(...)) fits a DSPy judge to people's feedback. The metric is per-field agreement, with the people's reason text as GEPA's textual feedback, so the optimizer learns from why people judged as they did.
  • A trained judge is versioned: its compiled program, dataset version, metric scores and optimizer settings are saved together. Its version string is a hash of the program. The version is recorded on every verdict, so scores from different judges never mix.
  • For decision evaluators, the "training" is threshold calibration: decision_boolean_threshold and route thresholds are tuned against the same datasets. Versioned Jev model ids (jev-1.13.0) pin behavior once tuned.

Datasets

  • Dataset[InputT, VerdictT] holds examples of an input with a verdict from people, a reference, or both.
  • Sources:
    • feedback exported by the libraries' [evals] extras
    • Hugging Face datasets at a pinned revision
    • existing Langfuse datasets
  • Sync: evalr writes datasets to Langfuse, with deterministic item ids so re-syncing updates rather than duplicates, and to the Hugging Face Hub, with a dataset card recording the source and version.
  • Splits are deterministic (hashed by example id), so an example never moves between train and validation.

Experiments and metrics

  • run_experiment(dataset, task, evaluators=[...]) runs over Langfuse's experiment API. The task's own spans (agents, tools, database) nest under each item's trace, and every evaluator's verdict becomes Langfuse scores.
  • Metrics:
    • agreement with people, per field: accuracy and Cohen's kappa for categories and yes/no, mean absolute error and rank correlation for ordinals
    • calibration for decision models: expected calibration error and Brier score
    • cost and latency per evaluator
  • End-to-end measures are defined generically here, with each library's extra supplying the data: task completion (judged), drop-off and rewrite rate (computed from logs).

Observability

evalr emits OpenTelemetry spans under the evalr scope (API only). An online verdict can also be emitted as a gen_ai.evaluation.result event on the span it judged, so evaluation data isn't tied to one backend. pydantic-ai traces decision evaluators itself. DSPy is traced through OpenInference's instrumentor, and runs with num_threads=1 inside evalr, which handles concurrency itself so trace context is kept.

Packages and dependencies

Package Depends on Responsibility
evalr.core pydantic, opentelemetry-api Verdicts, the evaluator protocol, datasets, splits, metrics. Pure.
evalr.dspy ([dspy]) dspy DspyJudge, signature derivation, optimize with GEPA
evalr.decision ([jev]) pydantic-ai-slim with the typesafe extra DecisionEvaluator, fallback, calibration
evalr.langfuse ([langfuse]) langfuse Dataset sync, experiments, scores
evalr.hf ([hf]) datasets Hugging Face import and export

Python 3.12+, with the same quality gates as artifactr and reflexr: pyright strict with no suppressions, 100% branch coverage, and no network in tests (Jev is tested through the SDK's transport, DSPy with its dummy language model, and Langfuse with an in-memory exporter).

Phases

Phase Deliverable Exit criteria
0. Foundation Packaging, tooling, CI, process, license CI green at 100% coverage
1. Core Verdicts, the evaluator protocol, datasets, splits, agreement and calibration metrics, function evaluators Metrics property-tested against reference implementations
2. DSPy judges DspyJudge, signature derivation, formatters, optimize with GEPA, versioned judges A judge trained on a synthetic dataset beats its untrained self, offline, with a scripted language model
3. Decision evaluators DecisionEvaluator on Jev with a fallback, threshold calibration Typed verdicts with confidence from a mocked Jev API; unsure fields handed off
4. Langfuse and Hugging Face Dataset sync, experiments, scores, Hugging Face import and export Round trips against in-memory fakes; one manual run against local Langfuse
5. End-to-end measures Task completion, drop-off, rewrite rate; online evaluation helpers Measures computed on recorded fixtures from both libraries
6. Docs site and brand A documentation site in the family's style Strict build in CI

Drawbacks

  • A third library to keep aligned with the other two. The shared-conventions table in the libraries' RFCs is the reference.
  • DSPy and TypeSafe's SDK are young: evalr pins lower bounds and tests against pinned versions.

Alternatives

  • One kind of evaluator: language-model judges alone are slow and costly online; decision models alone cannot explain themselves. Both, with a fallback between them, covers each other's weakness.
  • Langfuse's managed evaluators: convenient, but they are prompts in a product, not typed, trained, versioned artifacts in code.

Unresolved questions

  • Where trained judges live: files in the application's repository, the Hugging Face Hub, or Langfuse's prompt management. v0.1 saves files, and records the version on every verdict. Settled by ADR-0007: one JSON file per judge, loaded against the types in code, with no pickles and no model from the file.
  • Free-text fields in decision evaluators: pydantic-ai raises UnfillableRoute for them. evalr either derives a decision-only view of the verdict type, or relies on the fallback. Settled by ADR-0008: both. A decision-only view leaves text out, and the core's Fallback composes a language-model judge that fills it.

Tracking

  • Phase 0: foundation
  • Phase 1: core
  • Phase 2: DSPy judges
  • Phase 3: decision evaluators
  • Phase 4: Langfuse and Hugging Face (the manual run against a local Langfuse 4.46 is done: #26)
  • Phase 5: end-to-end measures
  • Phase 6: docs site and brand, published at evalr.alexnodeland.com (ADR-0010)