evalr.langfuse¶
Langfuse datasets, scores, score configs and experiments (the [langfuse] extra).
Langfuse holds working datasets and experiments, next to the traces they judge (ADR-0003), and
the scores of evaluators and people alike (ADR-0012). This package adapts it to the
DatasetStore, ScoreSink, ScoreConfigStore and ExperimentTracker ports
(ADR-0006), through the application's own Langfuse client.
See Langfuse datasets, Scores in Langfuse and Experiments in Langfuse.
Datasets¶
LangfuseDatasetStore
¶
Saves datasets to Langfuse, and loads them from it, through a Langfuse client.
The client is the application's, configured as it likes; the store makes no other connection. Its calls are synchronous, so they run in a worker thread.
Use a Langfuse client.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Langfuse
|
The application's client. |
required |
save
async
¶
load
async
¶
load(
name: str,
/,
*,
input_type: type[InputT],
verdict_type: type[VerdictT],
revision: str | None = None,
) -> Dataset[InputT, VerdictT]
Read the dataset's active items, as they are or as they were at revision.
Items not written by evalr are read as they are: the input as the input, and the expected output as the verdict.
Raises:
| Type | Description |
|---|---|
DatasetNotFound
|
Langfuse has no dataset of that name, or the revision is not a time. |
ValidationError
|
The items are not examples of these types. |
Scores¶
LangfuseScoreSink
¶
Records scores in Langfuse, on the traces or sessions they judge.
It records evaluators' scores and the libraries' mirrors of people's feedback alike. Each
score keeps its id (Langfuse's score_id), so recording a score again replaces it, and its
span (Langfuse's observation) and timestamp, when it has them. A yes or no is sent as 1 or 0,
and the score's metadata (its source, the evaluator, its version and the confidence)
becomes Langfuse's. Langfuse sends scores in the background; flush waits for them.
Use a Langfuse client.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Langfuse
|
The application's client. |
required |
record
async
¶
Queue the scores for Langfuse.
Raises:
| Type | Description |
|---|---|
ValueError
|
A score has neither a trace nor a session, one of which Langfuse needs to attach it to; none of the scores is queued. |
LangfuseScoreConfigStore
¶
Keeps score configs in Langfuse, so it knows each score's type, range and choices.
Langfuse's API is synchronous, so its calls run in a worker thread.
Use a Langfuse client.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Langfuse
|
The application's client. |
required |
create
async
¶
create(config: ScoreConfig) -> None
Create a score config in Langfuse.
Raises:
| Type | Description |
|---|---|
ValueError
|
The name is not one Langfuse accepts: at most 35 letters, digits, spaces
and |
Experiments¶
LangfuseExperimentTracker
¶
LangfuseExperimentTracker(
client: Langfuse,
*,
type_names: Mapping[type[BaseModel], str] | None = None,
)
Runs experiments in Langfuse, where they show beside the traces of each item.
Each example's task and evaluators run in the item's trace, so the spans of whatever they call nest under it, and each verdict records it. Text fields, which Langfuse's evaluations cannot hold, become the comment of the verdict's other scores.
Use a Langfuse client.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
client
|
Langfuse
|
The application's client. |
required |
type_names
|
Mapping[type[BaseModel], str] | None
|
Score names for verdict types, where a library registers its feedback under a name other than the class name in snake case, so that evaluators' scores and people's line up. |
None
|
run_experiment
async
¶
run_experiment(
name: str,
/,
*,
dataset: Dataset[InputT, VerdictT],
task: Task[InputT, VerdictT, OutputT],
evaluators: Sequence[Evaluator[OutputT, BaseModel]],
run_name: str | None = None,
max_concurrency: int = 4,
metadata: Mapping[str, str] | None = None,
) -> ExperimentResult[OutputT]
Run the task on every example in Langfuse, and judge each output.
A run is named by Langfuse ({name} - {time}) unless named.