evalr.memory¶
In-memory adapters of every port, for tests and prototypes.
They behave like the real adapters, as the contract suites in evalr.contracts check, but keep
everything in the process. Applications and the sibling libraries use them to test their own
evaluation code without a network.
See Testing with the contracts.
Datasets and feedback¶
InMemoryDatasetStore
¶
Keeps datasets as JSON records in memory, one revision per distinct content.
A revision is the dataset's content hash (Dataset.version), so saving the same examples
again makes no new revision. Loading validates the records again, as a real store does.
Start empty.
save
async
¶
Store the dataset's records under its name, and return its content hash.
load
async
¶
load(
name: str,
/,
*,
input_type: type[InputT],
verdict_type: type[VerdictT],
revision: str | None = None,
) -> Dataset[InputT, VerdictT]
Load the latest revision of a dataset, or the one given.
Raises:
| Type | Description |
|---|---|
DatasetNotFound
|
No dataset has the name, or it has no such revision. |
InMemoryFeedbackSource
¶
InMemoryFeedbackSource(
examples: Iterable[Example[InputT, VerdictT]],
*,
input_type: type[InputT],
verdict_type: type[VerdictT],
)
Yields a fixed list of examples, as a library's feedback source would.
Hold the examples to yield.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
examples
|
Iterable[Example[InputT, VerdictT]]
|
The feedback, as examples with verdicts. |
required |
input_type
|
type[InputT]
|
The Pydantic model of the inputs. |
required |
verdict_type
|
type[VerdictT]
|
The feedback type. |
required |
Scores and experiments¶
InMemoryScoreSink
¶
InMemoryScoreConfigStore
¶
Keeps score configs by name.
Attributes:
| Name | Type | Description |
|---|---|---|
configs |
dict[str, ScoreConfig]
|
The configs created, by name. |
Start empty.
create
async
¶
create(config: ScoreConfig) -> None
Keep a config.
Raises:
| Type | Description |
|---|---|
ValueError
|
The store has a config of that name, so the caller created one it should have looked for first. |
InMemoryExperimentTracker
¶
InMemoryExperimentTracker(
*, tracer_provider: TracerProvider | None = None
)
Runs experiments in the process, each example in its own span, and keeps every run.
Attributes:
| Name | Type | Description |
|---|---|---|
runs |
list[ExperimentResult[BaseModel]]
|
Every run so far, oldest first. |
Start with no runs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tracer_provider
|
TracerProvider | None
|
Where the examples' spans go; the global provider by default. |
None
|
run_experiment
async
¶
run_experiment(
name: str,
/,
*,
dataset: Dataset[InputT, VerdictT],
task: Task[InputT, VerdictT, OutputT],
evaluators: Sequence[Evaluator[OutputT, BaseModel]],
run_name: str | None = None,
max_concurrency: int = 4,
metadata: Mapping[str, str] | None = None,
) -> ExperimentResult[OutputT]
Run the task on every example, and judge each output with every evaluator.
A run is named {name} #{n} by default, counting this tracker's runs of the
experiment. Each example runs in a span named evalr.experiment.item {name}.
Optimizers¶
BestOf
¶
Picks the evaluator that agrees best with people on the training set.
It works for any kind of evaluator, and trains nothing: it measures the evaluator it is
given and each candidate on train, keeps the best, and measures that one on
validate. A tie keeps the earlier, so the evaluator given wins unless a candidate is
strictly better.
Attributes:
| Name | Type | Description |
|---|---|---|
measurements |
list[Measurement[VerdictT]]
|
Every measurement taken, in order; the last is the choice on |
Hold the candidates.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
candidates
|
Sequence[Evaluator[InputT, VerdictT]]
|
The evaluators to compare with the one being optimized. |
required |
max_concurrency
|
int
|
How many examples are judged at once. |
4
|