Scores and score sinks¶
Observability backends record judgements as scores: named values attached to the traces or sessions they judge. evalr turns every verdict into scores, one per field, and records them through a score sink: in memory, in Langfuse, or as OpenTelemetry events. artifactr and reflexr turn people's feedback into scores with the same mapping, from evalr, so people's scores and evaluators' scores for the same field sit side by side (ADR-0011).
The examples on this page use the Helpfulness type from Getting started.
A verdict as scores¶
scores(verdict) gives one Score per field of the verdict that has a value, in the order of the type's fields:
from evalr import Verdict, scores
verdict = Verdict(
value=Helpfulness(rating=4, resolved=True, reason="It refunded the charge"),
confidence={"rating": 0.8, "resolved": 0.93},
evaluator="helpfulness-decision",
version="7cb26270254b",
trace_id="4bf92f3577b34da6a3ce929d0e0e4736",
)
for score in scores(verdict):
print(score.name, score.data_type, repr(score.value), score.metadata)
helpfulness.rating NUMERIC 4.0 {'evaluator': 'helpfulness-decision', 'version': '7cb26270254b', 'confidence': 0.8}
helpfulness.resolved BOOLEAN True {'evaluator': 'helpfulness-decision', 'version': '7cb26270254b', 'confidence': 0.93}
helpfulness.reason TEXT 'It refunded the charge' {'evaluator': 'helpfulness-decision', 'version': '7cb26270254b'}
The field's kind decides the score's data type and value:
| Field kind | Data type | Value |
|---|---|---|
| binary | BOOLEAN |
True or False |
| ordinal, numeric | NUMERIC |
a float |
| categorical | CATEGORICAL |
the choice as a string (an Enum's value) |
| text | TEXT |
the text |
A field left empty (None or "") gives no score, and a choice or text longer than MAX_TEXT (500 characters) is cut. Each score carries its evaluator and version, the field's confidence where the evaluator had one (the metadata above), and the trace it is attached to: the verdict's own unless you pass trace_id=. span_id= names the span within that trace that the verdict judges, where you know it.
Names and ids¶
A score is named {type}.{field}. The type is the verdict type's class name in snake case (Helpfulness is helpfulness, TaskCompletion is task_completion), which is also how artifactr and reflexr name feedback types by default. Where a library registers the feedback type under another name, pass it, so the evaluator's scores land beside people's:
A score's id is a UUID derived from what the verdict is about, the evaluator, its version and the score's name. What it is about is the subject you pass (a run id, a turn id, an example id), or else the verdict's trace, or else a hash of the verdict's value. So:
- Recording a verdict again replaces its scores rather than adding more: a store that keeps one score per id stays clean however often an evaluation is retried.
- A new evaluator version adds new scores rather than overwriting the old version's, so the two can be compared.
Pass subject= whenever the verdict has no trace, so two verdicts on different inputs cannot collide.
Score configs¶
A backend that knows how to read a score (its data type, its range, its choices) can check the values it receives and offer the choices to people scoring by hand. score_configs describes each field of a type as a ScoreConfig, in the order of its fields:
from evalr import score_configs
for config in score_configs(Helpfulness):
print(config.name, config.data_type, config.minimum, config.maximum, config.categories)
helpfulness.rating NUMERIC 1.0 5.0 ()
helpfulness.resolved BOOLEAN None None ()
helpfulness.reason TEXT None None ()
- The data type follows the field's kind, as in the table above, and
categoriesholds a categorical field's choices as strings. - The bounds are kept as declared:
Field(gt=0)gives a minimum of 0, even on anint, whereverdict_fieldsreads 1. It is what the libraries have always created in Langfuse, and a config, once created, is never changed. - A field that cannot be scored is skipped, where
verdict_fieldsraises: a list, a nested model or a union of several types. A library's feedback type may have such fields. type_name=names the scores as a library registered the type, as forscores.
A ScoreConfigStore keeps configs by name: names() lists the ones it has, and create(config) adds one. sync_score_configs creates the configs a store lacks and leaves the others as they are, so a config's definition never changes under the scores that use it:
from evalr import sync_score_configs
from evalr.memory import InMemoryScoreConfigStore
store = InMemoryScoreConfigStore()
print(await sync_score_configs(store, score_configs(Helpfulness)))
print(await sync_score_configs(store, score_configs(Helpfulness))) # nothing new
LangfuseScoreConfigStore keeps them as Langfuse's score configs (Scores in Langfuse), and artifactr and reflexr sync every feedback type they register into it.
Feedback as scores¶
artifactr and reflexr keep people's feedback as JSON in their logs, and mirror it to a score sink. Their mirrors use evalr's mapping, so a person's rating and an evaluator's rating become the same score. score_values pairs each field of a validated value with its config and its score value, and takes a model's JSON as well as its fields:
from evalr.core import score_values
feedback = {"rating": 2, "resolved": False, "reason": ""}
for config, value in score_values(Helpfulness, feedback, type_name="reply_helpfulness"):
print(config.name, config.data_type, repr(value))
What the mapping does not decide stays in each library: which trace or session a piece of feedback is about, and the score's id. A mirror builds each Score itself, with no evaluator, and describes where the feedback came from in source, which a sink records with the rest of the score's metadata:
from datetime import UTC, datetime
from evalr import Score
score = Score(
id="0f6d2c1e-5b8a-5d3c-9e7f-1a2b3c4d5e6f", # derived from the feedback's event and the name
name="reply_helpfulness.rating",
value=2.0,
data_type="NUMERIC",
session_id="thread_01", # feedback on a whole thread; a turn's feedback names its trace
timestamp=datetime(2026, 9, 29, 9, 30, tzinfo=UTC), # when it was given
source={"workspace_id": "support", "actor": "alice"},
)
print(score.metadata)
A score's metadata is its source, then the evaluator, its version and the confidence, where it has them. A Score refuses what does not fit together: a value not of its data type (a bool for BOOLEAN, a float for NUMERIC, a string for CATEGORICAL and TEXT), a source key that would clash with the rest of its metadata (evaluator, version or confidence), and a span without a trace (span_id without trace_id).
Score sinks¶
A ScoreSink records scores, idempotently by id: recording a score again replaces it, and the latest value wins. It is one method:
from evalr.memory import InMemoryScoreSink
sink = InMemoryScoreSink()
await sink.record(scores(verdict))
await sink.record(scores(verdict)) # replaces, does not add
print(len(sink.scores))
| Sink | Package | Records scores |
|---|---|---|
InMemoryScoreSink |
evalr.memory |
In memory, by id, for tests |
LangfuseScoreSink(client) |
evalr.langfuse |
In Langfuse, on the traces or sessions they judge |
OtelEventSink() |
evalr.online |
As OpenTelemetry gen_ai.evaluation.result events (Online evaluation) |
Online evaluation sends every verdict's scores to the sinks you give it, and the libraries' feedback mirrors send people's. A sink of your own is any class with an async record(scores); check it with check_score_sink.
Scores in Langfuse¶
LangfuseScoreSink records each score with the Langfuse client's create_score, on the score's trace or session. It needs the langfuse extra:
from langfuse import Langfuse
from evalr.langfuse import LangfuseScoreSink
langfuse_scores = LangfuseScoreSink(Langfuse())
await langfuse_scores.record(scores(verdict))
await langfuse_scores.flush() # before a short-lived process exits
- It records evaluators' scores and the libraries' mirrors of people's feedback alike.
- Each score keeps its id as Langfuse's
score_id, so recording a verdict again replaces its scores in Langfuse too, and its span (as Langfuse's observation) andtimestamp, where it has them. A yes or no is sent as 1 or 0. - Evalr's online scores attach to the judged span as a Langfuse observation; if the Langfuse exporter filters that span out (v4's default keeps only LLM spans), the score names an observation Langfuse lacks.
- The score's
metadata(its source, the evaluator, its version and the confidence) goes in Langfuse's metadata. Langfuse reads back a metadata value that parses as JSON as that JSON, so a version such as"1"(aFunctionEvaluator's default) reads back as the number1. - Langfuse attaches scores to traces or sessions, so if any score has neither,
recordraisesValueErrorand queues none of them. - The client queues scores and sends them in the background;
flush()waits until they are sent.
LangfuseScoreConfigStore keeps score configs as Langfuse's own, so Langfuse knows each score's type, range and choices:
from evalr.langfuse import LangfuseScoreConfigStore
await sync_score_configs(LangfuseScoreConfigStore(Langfuse()), score_configs(Helpfulness))
- A config keeps its bounds, its description, and a categorical field's choices as Langfuse's categories, valued by their order.
- Langfuse takes a config name of at most 35 letters, digits, spaces and
_.()-. Any other name is refused withValueErrorbefore anything is sent: shorten the type's name (type_name=) or the field's. - Its calls run in a worker thread, and
names()reads every page of Langfuse's configs, archived ones included.
Scores from experiments in Langfuse are recorded by Langfuse's experiment API instead, as each item's evaluations.