Skip to content

evalr.memory

In-memory adapters of every port, for tests and prototypes.

They behave like the real adapters, as the contract suites in evalr.contracts check, but keep everything in the process. Applications and the sibling libraries use them to test their own evaluation code without a network.

See Testing with the contracts.

Datasets and feedback

InMemoryDatasetStore

InMemoryDatasetStore()

Keeps datasets as JSON records in memory, one revision per distinct content.

A revision is the dataset's content hash (Dataset.version), so saving the same examples again makes no new revision. Loading validates the records again, as a real store does.

Start empty.

save async

save(dataset: Dataset[InputT, VerdictT]) -> str

Store the dataset's records under its name, and return its content hash.

load async

load(
    name: str,
    /,
    *,
    input_type: type[InputT],
    verdict_type: type[VerdictT],
    revision: str | None = None,
) -> Dataset[InputT, VerdictT]

Load the latest revision of a dataset, or the one given.

Raises:

Type Description
DatasetNotFound

No dataset has the name, or it has no such revision.

InMemoryFeedbackSource

InMemoryFeedbackSource(
    examples: Iterable[Example[InputT, VerdictT]],
    *,
    input_type: type[InputT],
    verdict_type: type[VerdictT],
)

Yields a fixed list of examples, as a library's feedback source would.

Hold the examples to yield.

Parameters:

Name Type Description Default
examples Iterable[Example[InputT, VerdictT]]

The feedback, as examples with verdicts.

required
input_type type[InputT]

The Pydantic model of the inputs.

required
verdict_type type[VerdictT]

The feedback type.

required

input_type property

input_type: type[InputT]

The Pydantic model of the inputs.

verdict_type property

verdict_type: type[VerdictT]

The feedback type.

examples async

examples() -> AsyncIterator[Example[InputT, VerdictT]]

Yield the examples, in order.

Scores and experiments

InMemoryScoreSink

InMemoryScoreSink()

Keeps the latest score of each id.

Attributes:

Name Type Description
scores dict[str, Score]

The recorded scores, by id.

Start empty.

record async

record(scores: Sequence[Score]) -> None

Keep the scores, replacing any with the same id.

InMemoryScoreConfigStore

InMemoryScoreConfigStore()

Keeps score configs by name.

Attributes:

Name Type Description
configs dict[str, ScoreConfig]

The configs created, by name.

Start empty.

names async

names() -> set[str]

Return the names of the configs created.

create async

create(config: ScoreConfig) -> None

Keep a config.

Raises:

Type Description
ValueError

The store has a config of that name, so the caller created one it should have looked for first.

InMemoryExperimentTracker

InMemoryExperimentTracker(
    *, tracer_provider: TracerProvider | None = None
)

Runs experiments in the process, each example in its own span, and keeps every run.

Attributes:

Name Type Description
runs list[ExperimentResult[BaseModel]]

Every run so far, oldest first.

Start with no runs.

Parameters:

Name Type Description Default
tracer_provider TracerProvider | None

Where the examples' spans go; the global provider by default.

None

run_experiment async

run_experiment(
    name: str,
    /,
    *,
    dataset: Dataset[InputT, VerdictT],
    task: Task[InputT, VerdictT, OutputT],
    evaluators: Sequence[Evaluator[OutputT, BaseModel]],
    run_name: str | None = None,
    max_concurrency: int = 4,
    metadata: Mapping[str, str] | None = None,
) -> ExperimentResult[OutputT]

Run the task on every example, and judge each output with every evaluator.

A run is named {name} #{n} by default, counting this tracker's runs of the experiment. Each example runs in a span named evalr.experiment.item {name}.

Optimizers

BestOf

BestOf(
    candidates: Sequence[Evaluator[InputT, VerdictT]],
    *,
    max_concurrency: int = 4,
)

Picks the evaluator that agrees best with people on the training set.

It works for any kind of evaluator, and trains nothing: it measures the evaluator it is given and each candidate on train, keeps the best, and measures that one on validate. A tie keeps the earlier, so the evaluator given wins unless a candidate is strictly better.

Attributes:

Name Type Description
measurements list[Measurement[VerdictT]]

Every measurement taken, in order; the last is the choice on validate.

Hold the candidates.

Parameters:

Name Type Description Default
candidates Sequence[Evaluator[InputT, VerdictT]]

The evaluators to compare with the one being optimized.

required
max_concurrency int

How many examples are judged at once.

4

optimize async

optimize(
    evaluator: Evaluator[InputT, VerdictT],
    /,
    *,
    train: Dataset[InputT, VerdictT],
    validate: Dataset[InputT, VerdictT],
) -> Evaluator[InputT, VerdictT]

Return whichever of the evaluator and the candidates scores best on train.