API reference¶
The reference is generated from the library's docstrings and type annotations. It documents each package's public API: the names in its __all__. Anything else is internal and may change without notice.
Packages¶
evalr is built as ports and adapters (Concepts). evalr.core holds the values, the pure functions over them and the ports; every other package implements or uses the ports, depends only on the core, and needs its extra where it adapts a third-party library.
| Package | What it holds | Install |
|---|---|---|
evalr.core |
Verdicts and field kinds, the ports, function evaluators and Fallback, examples, datasets and splits, formatters, scores and score configs, experiments' results, measuring and optimizing, the metrics, and tracing |
core |
evalr.memory |
In-memory adapters of every port, and BestOf |
core |
evalr.contracts |
A contract suite per port, which every adapter passes | core |
evalr.jsonl |
A dataset store on JSON Lines files | core |
evalr.measures |
End-to-end workflow measures: task completion, drop-off and rewrites | core |
evalr.online |
Evaluators on live traffic, budgets, and OpenTelemetry evaluation events | core |
evalr.dspy |
DSPy judges, GEPA, and saved judges | dspy extra |
evalr.decision |
Decision evaluators, the decision-only view, and threshold calibration | jev extra |
evalr.langfuse |
Langfuse datasets, scores and experiments | langfuse extra |
evalr.hf |
Hugging Face datasets: publishing, pinning and importing | hf extra |
The top-level package¶
evalr re-exports every type in evalr.core, and the functions most applications call, so from evalr import Dataset, measure works. The rest (the individual metrics, the tracing helpers, the parts of the score mapping, split_bucket and the constants) is imported from evalr.core, where every name is documented:
| Name | Documented in |
|---|---|
Verdict, Confidence |
evalr.core: Verdicts |
FieldKind, VerdictField, verdict_fields, UnsupportedField |
evalr.core: Verdict fields |
Evaluator, FunctionEvaluator, Fallback, HandOff |
evalr.core: Evaluators |
Example, Dataset, DatasetStore, DatasetNotFound, DuplicateExample, FeedbackSource, collect |
evalr.core: Examples and datasets |
Formatter, InputFormatter, TokenCounter |
evalr.core: Formatters |
Score, ScoreDataType, ScoreSink, scores |
evalr.core: Scores |
ScoreConfig, ScoreConfigStore, score_configs, sync_score_configs |
evalr.core: Score configs |
ExperimentTracker, Task, ExperimentResult, ItemResult |
evalr.core: Experiments |
measure, Measurement, Optimizer, optimize, Training, DatasetRef |
evalr.core: Measuring and optimizing |
agreement, Agreement, FieldAgreement, calibration, FieldCalibration, evaluator_stats, EvaluatorStats |
evalr.core: Metrics |
Judging |
evalr.core: Tracing |
evalr.__version__ is the installed version.